![]()
Measuring Language Transfer in Robot Policies
Adding Greek to a Cosmos3 vision-language-action policy
Ayoub Kirouane1 Georgios Giaples1 Christos Petrocheilos1
1Sophea AI, KIEFER SA, Athens, Greece
{a.kirouane, g.giaples, c.petrocheilos}@kiefer.gr
models@sophea.ai
September 2026
Robot foundation models are trained and evaluated almost entirely in English, and no robot demonstration corpus exists for most languages. We ask what it takes to add one, using Greek and an open vision-language-action stack whose Greek is machine rephrased from English under a mandated glossary, none of it human-authored. That part is easy. The hard part is knowing whether it worked, and this paper is mostly about that. Five instruments that a practitioner would reach for first all report success where there is none: a colour-histogram coherence metric doubled twice on generations that were pure noise; a standard single-goal benchmark scored under Greek instructions and under deliberately wrong ones; a ten-goal suite credited a policy with Greek instruction-following that a ninety-task suite shows to be marginal at best, under three points over its own control on every seed; training loss ranks six policies within of each other while their Greek ability spans a factor of ; and single-run comparisons between recipes are uninterpretable, because target-language success moves points on the random seed while English moves . Five of our own conclusions did not survive contact with these controls; we report each retraction with the evidence that forced it, because a claim that survives a control it could have failed is worth more than one that was never tested. What survives is a small, checkable core: demonstrations in the target language are necessary but not sufficient. A multilingual tower transfers nothing to the action pathway without them (English-only training leaves Greek at its own wrong-instruction floor, across three seeds), and target-language demonstrations on their own are little better: on ninety tasks, across three seeds, a Greek-only policy’s margin over its own wrong-instruction floor never exceeds points while a bilingual policy’s never falls below . A bilingual policy follows Greek on roughly half the episodes English reaches on ten goals and two fifths of them on ninety tasks, with no architecture change; much of the ten-goal number is that pipeline’s glossary-bound phrasing rather than the language, a penalty that training on seven phrasings per task halves; and warm-starting from a language-adapted world model (one run) or unfreezing the text tower (three seeds) both make things worse. We also translated a real-robot corpus an order of magnitude larger in episodes than every simulated suite we had combined, and explain precisely why we cannot score it. Our recommendations are technical and cheap: build the null before the metric, and replicate before believing.
. Introduction
Robot foundation models are built, trained, and evaluated in English. Their demonstration corpora are English-annotated, their benchmarks issue English instructions, and the language towers they inherit from vision-language pretraining are strongest in English. For a speaker of any other language this is a hard wall: there is no Greek robot demonstration dataset, and collecting one means teleoperating a robot for months in order to reproduce capabilities that already exist behind an English interface.
This paper asks what the cheap path buys. We translate an existing corpus by machine, change nothing about the architecture, and measure what transfers. The translation took hours and worked. Measuring it took the rest of the study, and that asymmetry is our main finding: for a claim of the form “the policy follows language ”, the bottleneck is not the data or the model but the instrument.
We therefore report this as a study rather than a system. Its spine is a set of controls and what they destroyed. Every positive claim below is paired with a condition under which it would have failed, and where the control won, we say so: five conclusions we had drawn and written up did not survive replication or a null, and they appear here as retractions with the evidence that forced them rather than as omissions. We think this is the more useful artifact. The recipe we ended with is small enough to state in a sentence, and the reader who copies it without copying the controls will not know whether it worked for them.
What survives is layered, and the layers are the contribution: the choice of base model decides whether the project is possible at all; a multilingual tower is necessary but transfers nothing by itself; demonstrations in the target language are what convert tower knowledge into behaviour; and what the target language actually learns is largely the training translator’s phrasing, a penalty that training on several phrasings per task halves on the ten-goal suite. Two interventions that seem obviously helpful (initializing from a world model adapted to the target language, and unfreezing the tower so it can adapt) make things worse, and we report them as such.
Measuring any of this turns out to require care, because the benchmarks that would normally answer the question cannot. Where each scene admits exactly one trained goal, success rates are insensitive to the instruction: prior work finds that vision-language-action policies largely ignore language on such suites [10], and our own policies reproduce that pathology (84.6% under Greek instructions, 82.6% under deliberately wrong ones). Working cross-lingually gives us an unusually clean instrument for this. A language the policy provably cannot follow (established by an identical policy trained without that language, which performs at chance) provides a guaranteed null, with a null guaranteed by construction rather than assumed, at the cost of training a second policy to establish it.
Contributions.
The tower is the bottleneck. Supervised fine-tuning cannot teach Greek grounding to a video world model whose text tower was pretrained (effectively) on English: across a dose ladder from frozen tower to full learning-rate tower training at triple duration, Greek-conditioned generation remains noise. Swapping to a base model with a multilingual tower, the same stock recipe yields coherent Greek-conditioned scenes; the judged result below comes from a subsequent -clip run (70/30 Greek/English) built on those captions plus LIBERO renders.
Grounding transfers by level (single judge, validated for coherence only). The localized world model produces coherent in-domain scenes for 64% of Greek prompts, where the base model produced none that we judged coherent on inspection, but matches the specific prompt content in 0% of cases, reaching a partial match on 30% (English: 100% coherent, 90% good). Greek conditioning transmits domain reliably and specifics only partially.
Negative transfer from world model to policy (one run per arm). Warm-starting the policy from the Greek-adapted world model, the intuitive two-stage pipeline, degrades both languages (English 79.2% vs. 96.4% from base; Greek 13.6% vs. 48.6%). Video-generation fine-tuning appears to erode action-relevant representations of the base model.
A working, honestly-scoped Greek policy, and its mechanism. Bilingual demonstration sampling yields 48.6% Greek task success on ten goals, and 27.4% (three-seed mean) on ninety tasks, from machine-produced Greek alone. Per goal this is bimodal rather than uniform, and the wrong-instruction residue is command-dependent rather than a fixed fallback, though we could not identify the feature that governs it (Section 5).
A cross-lingual attribution control. Prior work shows that vision-language-action policies largely ignore language on standard suites [10]; we reproduce it in a new regime (84.6% Greek versus 82.6% wrong instructions on a single-goal suite) and contribute a variant whose null is guaranteed by construction: an identically trained policy that provably cannot follow the probe language. This complements perturbation-based controls rather than replacing them; it buys a null we do not have to assume, at the cost of needing a second trained policy.
Replication at scale, three seeds per arm. On a ninety-task suite (uniform chance , but for a policy that reads the scene and ignores the instruction) the bilingual policy’s margin over its own wrong-instruction control is – points on every seed, while a policy trained on Greek demonstrations alone reaches at most ; every bilingual seed exceeds every Greek-only seed. The seed instability that dominates the ten-goal results shrinks from to points, so it was largely an artifact of ten clusters.
. Related Work
Benchmarks passable without the modality they advertise.
That a multimodal agent can score well without using one of its modalities is an old and repeatedly rediscovered result. Balanced VQA was built because image-blind question priors solved the original benchmark [11, 1]; unimodal ablations matched full models in vision-and-language navigation [26]; prompt-based classifiers perform nearly as well with irrelevant prompts [28]. In imitation learning the mechanism has a name: a policy latches onto whichever observed variable best predicts the demonstrated action, which is causal confusion [7], and scene identity is a far better predictor than instruction text whenever a scene determines its goal. Benchmark design has responded: CALVIN chains instructions in a shared scene so affordance cannot substitute for language [21], and the LIBERO suite family varies goal independently of scene [19].
Most directly, prior work established our methodological conclusion in English. LIBERO-Plus perturbs seven factors across VLA models and reports that they are “largely insensitive to language variations” and “tend to ignore language instructions completely” [10]. We therefore do not claim the observation as novel. Our contribution on this axis is a complementary instrument and a confirmation in a different regime: rather than deleting or paraphrasing an instruction, we issue it in a language the policy is independently proven not to follow (an otherwise identical policy trained without target-language demonstrations scores at chance). The null is then guaranteed by construction rather than assumed, which removes the question of how much a paraphrase ought to matter, at the cost of a second trained policy. It does not supersede perturbation designs, which need no such policy and cover factors a language swap cannot.
Vision-language-action models and language conditioning.
Success rates on LIBERO and similar suites are routinely reported as evidence of instruction following [5, 15, 24, 4], building on language-conditioned imitation learning [20]. Our policies use the Cosmos platform [22]; the localization question we ask is orthogonal to model scale. Our own search did not surface prior work that trains and evaluates a vision-language-action policy in a language other than English, but this is a fast-moving area and we make no priority claim: we expect concurrent work and would welcome correction. We frame the contribution as instruments rather than as a result that beats a baseline.
Multilingual transfer and machine-translated localization.
Zero-shot cross-lingual transfer from multilingual encoders is well characterized [6, 12], and in instruction tuning a small multilingual admixture over a strong high-resource base outperforms monolingual target-language tuning [25]. We set out to test the embodied analogue of that result and can report only that our data point the same way and, on the larger suite, resolve it in paired form: on ten goals bilingual training beats Greek-only on average across three seeds per arm but not separably; on ninety tasks, again at three seeds per arm, every bilingual seed’s margin over its own control exceeds every Greek-only seed’s (Section 5). What we can state cleanly in either case is that without demonstrations in the target language a policy does not follow that language at all, and that with them alone it barely does. Multilingual grounding has been studied in navigation [17] but not, as far as we know, in manipulation. Because our entire target-language corpus is machine generated, machine artifacts are a live confound [2], though the pipeline rephrases rather than translates; we return to this in Section 8.
What fine-tuning costs a pretrained representation.
Full fine-tuning distorts pretrained features and degrades out-of-distribution performance [18, 29], a specific case of catastrophic forgetting [16]. For VLAs specifically, prior work argues that action-training gradients degrade the backbone’s semantic knowledge and proposes insulating the backbone from them [8]. Our tower-unfreezing result is consistent with that prediction and extends it to the cross-lingual case, where the pretrained representation is the only source of target-language competence and its degradation is therefore unusually visible.
World models as policy initializers.
Video-generative pretraining has been reported to help manipulation policies [30, 9], and generated trajectories with inverse-dynamics labels are an established route to policy data [3, 13]. Our negative-transfer result sits in tension with that literature: warm-starting from a video-generation checkpoint degraded our policies in both languages. We do not claim to overturn those results, whose setups differ from ours in objective, scale, and architecture; we report the discrepancy and the one-variable comparison that produced it.
. Setup
Model stack.
We use the Cosmos3 open robot-learning stack (Figure 1): a video world model in two variants, one with an English-centric 2B text tower (“English-tower model”) and one whose tower is a multilingual 8B vision-language model (“multilingual-tower model”), and an action policy architecture that conditions on instructions through the same tower and decodes action chunks through dedicated adapter modules. Policies are trained by imitation on demonstration datasets in LeRobot format; evaluation is closed-loop in the LIBERO simulator, with binary task success judged by the simulator’s goal predicates.
Why the stack looks like this.
Cosmos3 is an omnimodal world model built on a unified Mixture-of-Transformers architecture: an autoregressive transformer for reasoning and a diffusion transformer for generation share one set of multimodal attention layers and a single 3D rotary position embedding over space and time, so language, vision, audio and action are handled as modalities of one model rather than by separate encoders joined at the output (Figure 2) [23]. In reasoner mode, text and visual tokens run through causal self-attention; in generator mode, noisy image, video, audio and action tokens are denoised under full attention. This matters for what follows in one specific way: the instruction pathway a policy conditions on is the same pathway the world model reads, which is why Section 4 can interrogate the tower through generated video and Section 5 can interrogate it through actions, and why a result about one is evidence about the other. We take this description from the vendor’s documentation and do not verify it; nothing we measure opens the backbone or localizes where inside it a language is represented.
Greek data, produced by machine.
No Greek robot data exists and we authored none by hand. An LLM pipeline rephrases English into Greek imperatives “as if originally authored in Greek” under a mandatory glossary, with anti-calque rules and a review pass: (a) 1,273 rich structured scene captions of a BridgeData V2 subset [27], and (b) all 53,207 unique task instructions of the two suites: 53,096 from DROID [14], which carries them across 57,639 success episodes, and 111 from LIBERO [19]. Quality was audited by Greek-character ratio (mean ), structure preservation, and spot review. The translated instruction set is dominated by DROID; what we can do with it is limited by measurement rather than by data, and Section 6 states that limit precisely. Bilingual training requires no loader changes: instruction fields hold pipe-separated variants from which the loader samples uniformly at each step, so appending the Greek translation to the English variants yields 50/50 language sampling for free.
Tokenization.
The multilingual tower’s tokenizer is far less efficient in Greek than in English. Across the ten evaluation instructions it produces 274 Greek tokens against 72 English ones, an inflation of (Figure 3), and the segmentation is close to character-level: put the bowl on top of the cabinet is eight word-like tokens, while its Greek translation is twenty-four pieces, most of them single letters. We report this here because it is a candidate explanation for results below, and because it is a property of the substrate rather than of our method.
Training configuration.
Both studies use the stock Cosmos3 SFT trainer on B200. Policies train for iterations from the base action-policy checkpoint, AdamW at learning rate , weight decay , warm-up steps on a cosine schedule, batch size one per device, text tower frozen, action chunks of sixteen at 20 FPS with a ten-dimensional frame-wise-relative action space and 6-D rotations, agentview and wrist cameras at . The multilingual tower is Qwen3-VL-8B-Instruct; the English-centric tower is the 2B model of the same family. The instruction and caption translations were produced by kimi-k3; the independent Greek renderings and the English rewordings by gemini-3.7-flash and kimi-k3 respectively; and the coherence and caption-match judgements of Study 1 by gemini-3.7-flash. We name these because a reader reproducing the study needs to know that the “independent” translator and the judge are themselves particular models, and because two of them recur in more than one role. World-model runs use the same trainer for iterations with the tower frozen and only the generation pathway and its projections in the optimizer. Seeds are , and wherever three are reported. Configuration files accompany the paper.
Evaluation protocol.
Every policy is evaluated three ways on the same tasks, seeds, and initial states: correct English instructions, correct Greek instructions, and wrong instructions (each task commanded with a different task’s Greek instruction). All headline numbers use 50 trials per task (500 episodes per condition); exploratory readings at 10 trials per task are noted where relevant and differed from the replication by at most 3.4 points. For world-model outputs we report judged coherence and caption match (a vision-language model judged each clip against the English source caption regardless of generation language). We flag a weakness in this instrument: its calibration set labelled Greek-prompted clips as noise by assumption, which is part of what it is used to measure, and it validated only the binary coherence judgement, not the three-way content-match judgement we report in Section 4. We use it after finding that a color-histogram similarity metric was fooled twice by noise whose palette drifted toward the reference (Section 7).
Statistical treatment.
Each evaluation condition comprises ten goals 50 trials. Success is strongly clustered by goal (per-goal rates span the full range within a single condition), so episodes are not independent draws and a binomial interval over 500 episodes badly understates uncertainty. We therefore treat the goal as the sampling unit and report a paired task-level bootstrap (20,000 resamples of the ten goals, applied to both arms of a comparison simultaneously). With only ten clusters the percentile bootstrap is mildly anti-conservative (nominal 95% intervals cover roughly 89–91% under our own per-goal rate profiles), so we corroborate every conclusion that matters with an exact paired sign-flip permutation test, whose smallest attainable two-sided at is . Between-seed variance is not negligible and we have measured part of it. Retraining the reference policy with only the seed changed moves Greek success across 48.6/80.2/55.2% at three seeds, a range of 31.6 points, while English moves 1.0 (96.4/95.6/96.6%) and the wrong-instruction floor stays low (4.4/2.2/0.8%). On the ninety-task suite the same three-seed comparison moves Greek by 2.4 points, so much of that spread belongs to the ten-cluster suite rather than to the recipe (Section 5). The high-resource language is stable under reseeding; the low-resource language is not. That is itself a result, and a caution for anyone reporting single-run low-resource numbers. Its consequence here is that our cross-policy comparisons cannot be read at the precision their point estimates suggest. We have since replicated four arms at three seeds each, which is what allows the contrasts in Section 5 to be stated at all; arms still at one run are marked where they appear. Even at three seeds, a difference of thirty points between two singly-trained policies is not clearly larger than the spread we observe between seeds of one policy. We mark those claims as directional throughout, and we replicated the arms that carry practical advice. Comparisons within a policy (the wrong-instruction conditions, the phrasing ladder, the per-goal profile) share a checkpoint, so the additive seed effect cancels; what does not cancel is any seed-by-condition interaction, and we measure one (the translator penalty itself spans points across seeds of the single-phrasing arm). They are far less affected, not unaffected. Counting every inferential comparison in the text, not only those we tabulate, we make at least twenty, and we apply no multiplicity correction; Appendix A lists them all with their sampling units so a reader can apply one. Only the within-policy contrasts, whose task-level tests reach , would survive one; no run-level contrast can, because the smallest attainable at three runs per arm is . Small differences should be read as exploratory.
Our cross-policy claims therefore rest on a single decision rule, which we state once here: every seed of one arm exceeds every seed of the other on the paired within-policy quantity. At three runs per arm this is the exact permutation test’s floor, , so it carries a one-in-ten false-positive rate per comparison and we do not report it as significance. We use it as an ordering statement, always alongside the run-level effect size (for the ninety-task margins: a -point gap between the smallest bilingual margin, , and the largest target-only one, ).
. Study 1: Localizing the World Model
Targets in Figure 1: the text tower, swapped between its two variants, and the video world model it conditions.
An English tower cannot learn Greek by SFT.
We fine-tuned the English-tower world model on 855 Greek and 367 English rich captions, the of the that survived caption-quality filtering, under a dose ladder: generation pathway only (tower frozen); tower at the base learning rate; tower at the full rate for triple the iterations; and a 100%-Greek continuation. In every configuration, Greek-conditioned generation remained structureless noise while English-conditioned generation improved markedly (judged frame inspection; the histogram metric incident of Section 7 occurred here). We conclude that moderate-scale caption SFT cannot induce a new language in a tower that lacks it.
A multilingual tower localizes cheaply, but partially.
The multilingual-tower base model also produces noise from Greek prompts, so tower comprehension alone is insufficient: the generation pathway has never seen Greek-derived conditioning. However, the stock fine-tuning recipe (tower frozen, generation pathway trained) on the same captions, followed by a larger mixed run (6,836 clips, 70/30 Greek/English), produces genuinely coherent Greek-conditioned robot scenes (Figure 4).
Quantifying the transfer level.
Judged on 51 held-out scenes per language, English generation is fully coherent with a good content match on 90% of clips (95% Wilson interval ; the Greek intervals below are for coherence and for good match, so “none” is an upper bound of seven percent rather than a demonstrated zero). Greek generation is coherent on 64% of the 50 clips that returned an evaluable judgement (of 51 submitted) but achieves a good content match on none, and a partial match on 30%. Greek conditioning reliably selects the right kind of scene and not the described content: grounding transferred at the scene level, not the content level.
. Study 2: Does the Policy Read?
Targets in Figure 1: the text tower, held frozen, and the action adapters, which are the only modules trained.
Single-goal suites cannot attribute language.
We trained a policy on the LIBERO-10 suite with bilingual instruction sampling, warm-started from the localized world model. It performs well: 87.8% success under English instructions and 84.6% under Greek. Read naively, this is a bilingual robot. The wrong-instruction control refutes that reading: commanded to perform a different task, the policy still succeeds 82.6% of the time in Greek and 83.4% in English (Table 1; Figure 5). The English column carries this argument. This policy is the warm-started recipe, which the discriminative suite later shows reads Greek at chance (), so its Greek column cannot distinguish a suite that hides language from a policy that never had it; its English column can, because the same policy reads English at on the discriminative multi-goal suite, well clear of that suite’s floor. All four conditions lie within six points of one another. A task-level interval does exclude zero for a small effect of the Greek instruction on this suite ( points over wrong instructions, 95% CI , though the exact paired test gives ) and not for English (, CI ), so we do not claim the instruction is entirely unused; we claim it is nearly irrelevant, and negligible beside the 44.2-point separation the multi-goal suite produces. Each LIBERO-10 scene admits one trained goal, so the policy executes what the scene affords and the instruction is ignored. No language claim survives on such a suite without this control.
A discriminative protocol.
The LIBERO-Goal suite places ten goals in a single scene, making the instruction the only goal-selection signal. Here the three-way evaluation becomes discriminative, and it immediately falsified two of our own assumptions.
Negative transfer from the world model.
The policy warm-started from the Greek-adapted world model reaches 79.2% English but only 13.6% Greek, statistically at the 10% chance level (and 9.4% under wrong instructions). Retraining the same recipe from the base model, discarding the world-model warm start, improves both languages dramatically: 96.4% English and 48.6% Greek (wrong: 4.4%). The intuitive transfer pipeline was not merely unnecessary but harmful. We offer video-generation fine-tuning eroding action-relevant representations as one explanation, but a second is not excluded by this design: the world model was adapted on BridgeData V2, a real WidowX corpus, while the policy is trained and evaluated on a simulated Franka in LIBERO, so the warm start also carries a domain shift. Separating the two needs a world model adapted in-domain, which we did not run. The world model remains valuable as a synthetic-data engine, but its weights should not seed the policy.
| Suite | Policy | English | Greek | Wrong |
| LIBERO-10 (1 goal/scene) | bilingual, warm-started | 87.8% | 84.6% | 82.6% |
| same policy, wrong instruction in English | 83.4% | |||
| LIBERO-Goal (10 goals/scene) | bilingual, warm-started | 79.2% | 13.6% | 9.4% |
| LIBERO-Goal (10 goals/scene) | bilingual, from base | 96.4% | 48.6%† | 4.4% |
| same policy, wrong instruction in English | 0.2% | |||
| LIBERO-Goal | English-only, from base | 98.2% | 9.0%† | 7.6% |
| LIBERO-Goal | Greek-only, from base | 10.0% | 20.4%† | 2.2% |
| LIBERO-Goal | bilingual, tower unfrozen | 82.4% | 18.0%† | 4.8% |
| LIBERO-Goal | bilingual, 7 Greek phrasings/task | 92.8% | 52.0%† | 3.0% |
†Single run. We later replicated these four rows at three seeds each; Greek is highly seed-dependent while English is not. Greek across seeds: bilingual , English-only , Greek-only , tower-unfrozen , seven-phrasings . Rows in this table should not be compared against one another on their Greek column: a 31.6-point seed range swamps most of the differences. Only the English-only row is separable from the baseline, and it lands on its own wrong-instruction floor. See Section 3.
What bilingual sampling buys, and what the mean hides.
The from-base bilingual policy succeeds on 48.6% of Greek-commanded episodes (task-level bootstrap 95% CI ; ten goals is ten clusters, and the interval is correspondingly wide), with zero architecture changes and purely machine-translated instructions. The mean is a poor description of the behaviour. Per goal, Greek success is sharply bimodal, with five goals at – and five at – and nothing in between, while English is at least on all ten (Table 2). The split reproduces in an independent smaller run at the same seed. It does not survive reseeding: profiling all three seeds gives five, nine and six of ten goals above , and no goal fails in all three, so the split is a property of the run and not of the recipe (Conclusion).
Two observations bear on the mechanism. First, the instruction set contains a near-minimal pair: put the bowl on top of the cabinet and put the wine bottle on top of the cabinet differ in Greek only in the object noun phrase, and score and respectively. Both content words are independently grounded elsewhere, so the failure is not vocabulary but discrimination within a shared referent. Second, the wrong-instruction condition leaves a small, structured residue rather than a uniform floor, and that residue is command-dependent. Re-running it with every distractor reassigned (a different rotation of the same instruction set) moves the residue entirely: the single goal carrying all of it under the first assignment drops to zero, and five different goals become non-zero. A policy falling back on one dominant behaviour whenever Greek is uninformative would produce the same cell under both assignments; it does not. The policy is therefore processing the content of a Greek instruction even when that instruction is wrong.
We tested one specific account of what it extracts (latching onto a shared head noun), and the data do not support it: of the six goals whose reassigned distractor shares a noun stem with them, only three produce any success, and two of the five non-zero cells share no stem at all. We therefore report the structure without claiming its cause. Greek conditioning selects behaviour that depends on the instruction’s content, discriminates poorly between goals that share a referent, and we could not identify the feature that decides which wrong command elicits which behaviour.
A third observation reframes all of these numbers. Our English-only policy is a measured per-goal null: it cannot follow Greek, yet its Greek success is not a flat 10% but on put the bowl on the plate, on put the bowl on the stove, and near zero elsewhere. Read against that null rather than against chance, the bilingual policy’s profile changes materially: its largest genuine gains are on two goals, while on three goals (including one the paragraph above would count as a Greek failure) it is at or below a policy with no Greek at all. Pooled chance rates conceal this; per-goal nulls should be reported wherever a guaranteed-null probe is available. The defensible claim is therefore not 48.6% Greek competence but partial, non-compositional referent-level grounding, and a caution that aggregate success rates conceal mechanism even after the wrong-instruction control has been passed.
| Goal | English | Greek | Wrong |
|---|---|---|---|
| put the bowl on the plate | 1.00 | 0.98 | 0.00 |
| put the bowl on top of the cabinet | 0.96 | 0.94 | 0.44 |
| turn on the stove | 1.00 | 0.90 | 0.00 |
| put the wine bottle on the rack | 0.88 | 0.74 | 0.00 |
| push the plate to the front of the stove | 0.98 | 0.72 | 0.00 |
| put the bowl on the stove | 0.98 | 0.24 | 0.00 |
| open the top drawer and put the bowl inside | 0.92 | 0.24 | 0.00 |
| open the middle drawer of the cabinet | 1.00 | 0.08 | 0.00 |
| put the cream cheese in the bowl | 0.94 | 0.02 | 0.00 |
| put the wine bottle on top of the cabinet | 0.98 | 0.00 | 0.00 |
How much of this is Greek, and how much is our translator?
Every Greek string the policy was trained on came from one machine-translation pipeline, so its 48.6% may reflect that pipeline’s idiolect rather than the language.
We test this with three progressively more independent phrasings of the same ten commands, at three seeds each. Held-out paraphrases from the same generator score 31.6/52.6/28.4% and 26.4/62.4/24.6% — ranges of 24 and 38 points, so the levels are seed artifacts and we quote them only to show their spread. An independent translation (a different language model, not a native speaker, prompted to phrase each command the way a Greek speaker would actually say it, sharing no string with the training set) costs 23.3 points on average: three seeds of this recipe fall 48.6/80.2/55.2% to 17.8/55.4/40.8%, drops of 30.8, 24.8 and 14.4 points. We report the drop rather than the level deliberately. The level is not stable across seeds (a 37.6-point range), whereas both of its terms come from the same checkpoint, so the seed variance that dominates cross-policy comparisons here cancels within a policy. What the degradation is not is monotone in distance from the training translator. An earlier single-seed reading of this ladder showed a clean descent, and we took it as evidence of a gradient running outward from the training translator’s idiolect. Replicated, the ordering holds in two of five paired comparisons and inverts by 36 points in one: the diversity policy at seed 43 scores 42.6% on the mildly reworded set and 78.6% on the structurally distant one. Because both terms come from one checkpoint, seed variance cannot explain this away. The honest statement is that changing the phrasing costs accuracy, and that how much it costs is not predicted by how distant the phrasing looks to us (Figure 6).
Is it the translator, or is it one phrasing?
The drop above has an alternative reading that the design so far does not exclude. The policy sees exactly one English phrasing per task as well, so it may simply be brittle to any rewording, in which case the effect has nothing to do with translation. The control is cheap and we ran it: the same ten commands rephrased by a different English speaker (turn on the stove becomes switch the stove on), the same three checkpoints, the same 50 trials per task. English success falls from to , drops of , and points, a mean of . The Greek drops on the same checkpoints are , and , a mean of . Every Greek drop exceeds every English drop, the smallest Greek being against the largest English , and both terms of each drop come from one checkpoint, so seed variance cancels within a row. The policy handles unseen English phrasings nearly as well as trained ones and loses roughly six times more when the rewording is in the low-resource language (Figure 8).
That control is not sufficient on its own, and a reader should see why before believing it. The two perturbations are not matched in strength. The English rewriter changed syntax and left every content noun in place, none of ten; the independent Greek set changes a content word in five of ten, including the term the training glossary had fixed (Figure 7). A perturbation that swaps the noun naming the target object is larger than one that reorders a verb phrase, and our own analysis says the noun is what the policy latches onto, so the comparison as it stands confounds language with perturbation size.
We therefore ran a second English control matched on exactly that dimension. It substitutes an everyday synonym for the object or surface in seven of the ten commands (stove to hob, cabinet to cupboard, rack to shelf), chosen so that the physical referent is unchanged and the substitute names no other object in the scene. This is a more aggressive noun perturbation than the Greek set’s five of ten. English success falls from to , drops of , and points, a mean of against in Greek. Every Greek drop still exceeds every matched-English drop, though the margin is now against rather than the comfortable gap the syntax-only control suggested.
Two things follow, and we state both. Substituting the referring noun is not free in English: one seed lost points, so part of what the Greek set costs is the size of the perturbation and not the language. But English absorbs the harder perturbation at points where Greek loses under a milder one, roughly five times more, and the English cost barely moves as the perturbation goes from touching no nouns () to touching seven of ten () while Greek at an intermediate strength is five times worse than either. Generic brittleness to rewording therefore cannot explain the Greek drop. What we claim is that bounded statement, not that language accounts for all of it.
Two conclusions follow, and the first is uncomfortable. A large fraction of the headline number is translator-specific: the three seeds lose , and of their trained-phrasing level when the translator changes, a mean of , against for a rewording in English that substitutes the referring noun more often than the Greek set does. A policy trained on machine-translated instructions learns, to a substantial degree, the phrasing habits of the system that produced them; reported in the target language’s name, this overstates what was learned. Anyone localizing a stack this way should evaluate with an independently produced instruction set. We previously reported that paraphrases from our own generator understate this gap by more than ten points, and we withdraw that: it held at the seed we first measured and does not survive replication. Across three seeds our own held-out paraphrase sets cost and points against the independent set’s , which is the same penalty, not a smaller one. What the independent set buys is therefore not a larger measured drop but an instrument we did not author; the argument for using one is provenance, not magnitude, and the penalty is not a model copying one translator’s words but brittleness to any unseen Greek phrasing. Second, what remains is real: every one of the six policies we measured across both recipes scores above its instruction-blind floor on the independent set, so the policies do follow Greek they have never seen phrased that way. The honest headline is therefore not “48.6% Greek instruction-following” but “Greek instruction-following that loses roughly a third of its value when the translator changes.”
What failure looks like.
Success rates say how often a policy succeeds, never what it does instead. We therefore scored all ten goal predicates at the end of every episode, for a hundred episodes per condition, which distinguishes a policy that attempted the right task and fumbled it from one that competently performed the wrong one. The two languages fail in different ways. Of the English failures, none ended with any goal satisfied: the policy tried the commanded task and missed. Of the Greek failures, ended with the robot having successfully completed a different task in the scene. The wrong-instruction control fails the same way, at (Figure 9). Greek failure is therefore not a motor deficit but a comprehension one, and its signature is statistically indistinguishable from having been given an unrelated command.
The misdirection also has a single attractor. One goal, put the bowl on top of the cabinet, absorbs of the Greek diversions and of the wrong-instruction ones, so the fallback is one over-practised behaviour rather than random flailing; it is also among the goals the policy performs best when commanded. The clearest case is again a minimal pair: commanded put the wine bottle on top of the cabinet in Greek, the policy performs the bowl variant eight times, preserving the destination and substituting the object. This is the referent-collapse signature above, now read from a mechanism rather than inferred from an aggregate. We report it for one policy at one seed, so the attractor’s identity should not be assumed stable across runs; the asymmetry between motor and comprehension failure is the transferable part.
Phrasing diversity halves the translator penalty.
Because the preceding policies see a single Greek phrasing per task, we retrained with seven distinct Greek phrasings per task while holding the language sampling ratio fixed, isolating phrasing variety from the amount of target-language gradient. Comparing the two recipes by their success levels does not work here: with three seeds each they overlap (– against – on the independent set), because seed variance is larger than the effect. Comparing each policy against itself does work, because both terms of a within-policy drop come from one checkpoint and the seed term cancels (Figure 10). Measured that way the effect is unambiguous: moving to the independent translator costs the single-phrasing policies , and points, and the diverse policies , and . Every single-phrasing drop exceeds every diverse one, an exact run-level permutation test returns (the smallest value three runs per arm can produce), and the spread of the penalty collapses from points to . Diversity roughly halves the penalty and makes it predictable. It is the only intervention we tried that survives replication, and it is worth the cost we also measure: English success falls points, 95% CI . This is the cost of training on seven phrasings, and should not be confused with the -point cost of evaluating on an unseen English phrasing reported below.
Locating the residual gap.
Three probes test the obvious explanations for why Greek trails English (Table 1, lower half; Figure 11), and all three fail in the same direction. Training on English alone yields 98.2% English and 9.0% Greek: chance. A multilingual tower transfers nothing to the action pathway without demonstrations in the target language, and bilingual training costs only 1.8 points of English for 39.6 points of Greek. Training on Greek alone, which doubles Greek gradient share from one half to all of it, scored 20.4% in our first run, and we initially read that -point deficit as evidence that English demonstrations scaffold Greek. Replication withdraws it. Across three seeds the Greek-only arm scores 20.4/66.6/23.8% against the bilingual arm’s 48.6/80.2/55.2%: the ranges overlap, the best Greek-only run beats two of three bilingual runs, and an exact run-level permutation test gives where is the smallest value three runs per arm can attain. The means still favour the mixture (61.3% vs. 36.9%) and Greek-only is the most seed-unstable arm we measured (46.2-point range), so we report a direction this design cannot resolve rather than a mechanism. What does replicate is that every Greek-only run reads its instructions on this suite (20.4/66.6/23.8% against 2.2/3.6/8.4% under wrong ones). On the ninety-task suite the same recipe clears its floor by at most points on any of three seeds (Figure 12); we return to the point below. Finally, unfreezing the text tower, so that it may adapt to Greek directly, degrades both languages (82.4% / 18.0%) despite normal convergence; small-data action training damages the pretrained multilingual representation rather than improving it.
These failures were each significant under a task-level test, and we resist unifying them; replication later showed that test to be the wrong one for comparing policies, and withdrew one of the three (Section 5). An earlier draft of this paper proposed the slogan “do not touch the tower”; auditing our own configurations refuted it. The world model whose checkpoint we warm-start from was itself trained with the tower frozen (only the generation pathway and its projections were in the optimizer), and the monolingual probes likewise use the stock frozen-tower recipe. In two of the three failing interventions the multilingual representation is bit-identical to the one in the winning recipe. The failures therefore have three distinct locations: warm-starting damages the generation pathway that the policy recipe also trains; monolingual training changes the language composition of the data; and only the third actually perturbs the tower. Of the three, only the tower intervention and the no-Greek ablation survive seed replication as separable from the baseline. We report them separately, and note that the tower result replicates across three seeds but tests a single configuration, maximal aggression (every attention and MLP projection, the embedding table, and both layer norms unfrozen at a flat learning rate), which says nothing about gentler adaptation such as low-rank adapters or a strongly damped tower learning rate.
We can therefore say what did not work and where, but our evidence does not identify the cause of the residual English–Greek gap. Three probes, two of them replicated and one at a single run, do not close an intervention space. Two explanations remain open and we cannot separate them. The first is representational: the tower encodes Greek less richly than English, and fine-tuning at this scale cannot supply what pretraining did not. The second is mechanical, and cheaper to fix: the tower’s tokenizer fragments these Greek instructions into as many tokens as their English originals, most of them single characters, so the action pathway must learn goal selection from twenty-four sub-lexical pieces where English offers eight words. That would also predict the failure we observe, since character-level segmentations of the bowl on top of the cabinet and the wine bottle on top of the cabinet share most of their token -grams. Distinguishing the two inside one tower requires either a direct probe of the frozen tower’s encodings or a retrained policy on transliterated or vocabulary-extended Greek; both are future work.
Comparing across the two towers, however, is free, and it cuts against the mechanical explanation. The English-centric 2B tower, the one that never learns Greek at all in Study 1, fragments the same ten Greek instructions less than the multilingual 8B tower that learns it: tokens against , an inflation over English of rather than , and tokens per character rather than . If fragmentation were what blocks a language, the tower with the coarser Greek segmentation should be the more teachable one, and it is the opposite. We take this as evidence that fragmentation is not sufficient to explain a tower’s failure to acquire a language. It does not establish the converse, that representational depth carries the weight instead: the two towers differ in parameter count and in pretraining corpus as well as in fertility, so this is one counterexample against a sufficiency claim, not an attribution of cause at . It does not settle the residual gap within the multilingual tower, where the two remain entangled. We report the diagnostic because it costs nothing: anyone choosing a base model can compute both towers’ fertility in the target language before spending a GPU-hour, and should not read a low count as a green light.
Does any of this survive a larger suite?
Every number above rests on ten goals, and ten goals is ten clusters. We therefore repeated the protocol on the ninety-task suite, both arms, three training seeds each, four instruction conditions, trials per task, so episodes per cell and in total. Uniform chance falls from to , but that figure misdescribes the problem the policy faces: the ninety tasks are twenty scenes carrying two to seven goals each (mean ), so a policy that recognises the scene and performs a goal afforded by it scores about without reading the instruction. The measured wrong-instruction floor is , points below that, which is why we read margins over the measured floor throughout rather than distance from uniform chance. Three things change, and only the first is comfortable.
The bilingual policy reads Greek on every seed, and the margin is what shrinks. It scores under correct Greek ( across seeds) against a wrong-instruction floor (), a separation of points that varies by only points across seeds (; Figure 12). The effect is real and stable and the ordering is intact, but Greek success as a fraction of English falls from roughly one half on ten goals to here ( against ), and the absolute English–Greek gap widens to points. Both ratios count the wrong-instruction floor as Greek following. Corrected for it, Greek retains of English instruction-following at ninety tasks (), and we prefer that number. A reader who took “roughly half of English performance” from our ten-goal result should read it as an upper bound obtained on an easier suite.
The Greek-only policy barely reads at all, on any seed. Its margin over its own wrong-instruction floor is , and points across the three seeds (mean ), and its four conditions span , and points. On ten goals this arm cleared its floor on all three seeds, by to points, and we read that as weak but genuine instruction-following; here the largest margin any seed produces is . Every bilingual margin (minimum ) exceeds every Greek-only margin (maximum ): perfect separation at three seeds per arm, for which an exact run-level permutation test returns , the floor of this design rather than a small . The stronger evidence is within each run, where ninety tasks give a real interval. A paired task-level bootstrap over the ninety tasks puts the bilingual margin at , and across seeds, excluding zero in all three. The target-only arm gives , and : zero falls inside the interval for two of its three seeds and outside for the third, so the honest statement is that target-only training reads Greek weakly and inconsistently, not that it never reads it. Because both terms of each margin come from one checkpoint, the seed term cancels within a policy, so this compares within-policy quantities rather than levels. It is the strongest evidence we have that target-language demonstrations alone are not sufficient. It is weaker evidence for the stronger reading, that the English demonstrations are what make Greek legible: that is a between-policy claim, and a within-policy statistic constrains it only through the ordering of the two arms’ margins. In this paired form and on this suite it reinstates the scaffolding reading we withdrew on ten goals; we still do not report it as a level. The per-task view says the same thing in a different way (Figure 13): for the bilingual policy, Greek success clears the task’s own wrong-instruction rate by more than ten points on of tasks and correlates with English success at ; for the Greek-only policy it does so on , and its Greek and English per-task success correlate at , which is what a policy that does not read the instruction language looks like.
The translator-idiolect penalty largely disappears, and we do not think that is good news. Moving the bilingual policy from the training translator’s phrasings to an independently produced Greek rendering costs , and points across seeds (mean ) here, against points on ten goals. The tempting reading is that idiolect matters less at scale. The likelier one is that the instrument has run out of room: on ten goals Greek sat points above its floor, and a penalty had somewhere to go, whereas here it sits points above and cannot fall far before hitting the motion prior. A penalty measured against a compressed dynamic range is not evidence of a smaller effect. We report the number and decline to interpret it as a reduction.
Two further things follow. First, the seed instability that dominates the ten-goal results is largely a small-suite artifact: Greek success for the bilingual recipe varies by points across seeds on ninety tasks where it varied by on ten. Ninety tasks buy not only a tighter estimate of each policy but a stable one. Second, the suite is no longer what binds this paper; what binds every cross-policy claim is now the floor of the run-level test at three seeds per arm (), and pushing below it costs one training run per seed per arm.
. The Real-Robot Corpus We Cannot Score
Ninety-one percent of our translated instructions belong to DROID [14]: unique instructions carried by success episodes of real teleoperation, against LIBERO’s tasks. It is by far the larger asset, and every number in this paper comes from the smaller one. The reason is not that the data is unusable but that we cannot score it. DROID is real-robot data with no simulator, so the closed-loop success protocol on which every result above rests requires a physical Franka arm we do not have.
We report the position honestly rather than omitting the corpus. A bilingual Greek policy is trained on it, using the recipe this paper’s evaluations selected: from base, tower frozen, an equal mixture of English and Greek instructions. Instruction-level holdout is enforced rather than episode-level, because the same command recurs across episodes and an episode split would leak held-out phrasings back into training through a different episode. The accounting is: success episodes carry unique instructions; of those instructions, appearing in episodes, are held out; episodes remain for training. The figure quoted in Section 3 is the translated instruction set, from this corpus plus from LIBERO, and is a count of strings rather than episodes. Absent a simulator, the available proxy is action-prediction error on held-out episodes under the same three-way contrast used throughout, and we require it to clear its own gate before reading anything else from it: if it cannot separate a correct English instruction from an unrelated one, it is not measuring language and none of its Greek values mean anything.
It clears the gate, and the result is negative for Greek. Over held-out windows per condition, English instructions predict the ground-truth action chunk better than an unrelated instruction by (MSE against ), from a single run at windows per condition and reported without an interval; we read the gate as directional evidence that the metric responds to language at all, not as a measured effect size. Against that working metric, Greek is indistinguishable from the null: , a separation of . An independently produced Greek translation scores worse than the null. The gate passing is what makes this readable. Had every condition been flat, a blind metric and an untrained policy would be indistinguishable; because English separates on the same checkpoint, the instrument is working and the flat Greek is a property of the policy.
We state the scope carefully. The English signal is itself small, and this checkpoint is iterations of a recipe written for , so instruction conditioning is only beginning to emerge in either language. The claim we can defend is that at five percent of the intended schedule, English conditioning is already detectable on real-robot data and Greek conditioning is not. Whether Greek emerges later is untested and would cost roughly three days of compute to answer on a proxy that still cannot produce a success rate; we judged that a poor trade and say so rather than leaving the reader to assume we ran it.
Two things follow for anyone repeating this. First, the translation is cheap and the evaluation is not: at a cost of hours we produced Greek for real-robot episodes, ten times every simulated suite we had combined and a hundred and thirty times the one our results come from, and then measured none of it. The bottleneck was never the language. Second, a proxy metric needs its null before it needs its result, which is the same discipline the wrong-instruction control enforces above and the same one a colour-histogram metric evaded twice (Section 7). We describe the held-out construction and the independent re-translation in enough detail to be repeated, so the measurement can be done by anyone with the hardware and a translation pipeline of their own.
. Lessons for Evaluation
Five instruments failed us informatively. First, a color-histogram similarity metric against reference videos doubled, twice, while the underlying generations remained pure noise whose palette had drifted toward the reference (Figure 14); coherence must be judged structurally or by a calibrated judge, never by color statistics. Second, single-goal benchmarks silently cannot falsify language claims; the wrong-instruction control costs one extra evaluation run and would have exposed the 84.6% Greek illusion immediately.
Third, and the instrument that cost us the most, a single training run per arm is not a measurement in the low-resource language. Retraining one recipe with only the seed changed moves Greek success across while English moves a single point. Any cross-recipe difference smaller than about thirty points is therefore unresolvable at one run per arm, and three of our five retractions are of exactly that shape: an effect that looked clean at and vanished when the arm was replicated. The asymmetry is the practical warning. Nothing in our English numbers would have suggested this instability existed, so a practitioner validating a pipeline in the high-resource language will conclude, wrongly, that single runs are informative.
Fourth, and most cheaply overlooked: training loss carries no usable signal about instruction-following (Figure 15). Across six completed policies the late-training loss spans –, a range of , while their Greek success spans –, a factor of . Two policies at identical loss () differ by points of Greek, and the best Greek policy of the six carries the highest loss. The cause is structural rather than incidental: behaviour-cloning loss is dominated by action prediction from the visual and proprioceptive prior, so language conditioning is a small enough term that a policy at chance in the target language and one at are indistinguishable to the optimiser. This is the same blindness as the single-goal suites, one layer up, and it means a training curve can never be used to rank recipes on the property this paper is about.
Fifth, a small multi-goal suite is itself an instrument, and it misled us in two ways we saw only by scaling it. On ten goals the Greek-only policy cleared its wrong-instruction floor on every seed and we read that as weak but genuine instruction-following; on ninety tasks, three seeds per arm, its margin is at most points. And the -point seed instability we had treated as a property of low-resource training is largely a property of ten clusters: the same recipe varies by points across seeds on ninety tasks (Figure 16). Ten goals does not only bound every interval; it manufactures variance that a larger suite removes. The same caveat we apply to the shrunken idiolect penalty applies here in the authors’ favour and we state it too: a quantity that sits points above its floor has less room to vary than one that sits above, and the two suites were trained on different data, so part of this contraction is range, not stability.
What the five share is that each reports success where there is none, and each is cheaper to consult than the control that exposes it. That is why we pair every positive claim in this paper with a condition under which it would have failed: success-rate parity is not understanding, and five of our own preliminary conclusions were retracted when a control or a replication contradicted them; Table 3 lists them.
| Claim we drew | What refuted it | Status now |
|---|---|---|
| “Do not touch the tower”: every intervention that helped left it frozen | Audit of our own configurations: two of the three failing interventions leave the tower bit-identical | Withdrawn as a unifying rule; the tower result survives on its own |
| English demonstrations scaffold Greek (bilingual vs target-only ) | Seed replication: ranges overlap, run-level | Withdrawn as a level; reinstated as a paired margin comparison at ninety tasks (Section 5) |
| Phrasing diversity raises the independent-translator level ( vs ) | Seed replication: – against –, overlapping | Withdrawn as a level; survives as a within-policy drop |
| Held-out-paraphrase success decays monotonically with distance from the training translator | Seed replication: the ordering holds in two of five paired comparisons and inverts by points in one | Withdrawn; we report that phrasing change costs accuracy, not how much |
| Our own paraphrases understate the translator gap by more than ten points | Seed replication: own-generator drops average and against the independent set’s | Withdrawn; the case for an independent instrument is provenance, not effect size |
. Limitations
One language, one model family, simulation only. The Greek corpus is machine generated, rephrased rather than literally translated, and carries machine-artifact risk. No human-authored Greek appears anywhere in this study. Every Greek string a policy trained on or was evaluated against was written by a language model, including the “independent” set: independence there means a different model with a different, glossary-free prompt, not a native speaker. This bounds the idiolect result specifically. What we measure is the cost of moving between two machine renderings of the same commands, which is a lower bound on the cost of moving to Greek as a person would actually speak it, and it leaves open the possibility that both renderings share machine-translation artifacts that a human set would not. Collecting native phrasings for these ten commands is an afternoon’s work for a Greek speaker and is the first thing we would add. World-model binding was judged by a single VLM judge (calibrated, but one model). Policies generalize partially to unseen phrasings of their trained goals, well above the wrong-instruction floor in every run, but the level is strongly seed-dependent (24–62% across three seeds on one paraphrase set) and we can only report it as a range. Every goal a policy is asked to perform was seen in training; we measure no generalization to novel tasks. We do not identify the cause of the residual Greek–English gap: our probes eliminate three explanations but leave representational depth and tokenizer fragmentation (, measured) entangled, and we cannot say which dominates. Ten goals is ten clusters, and it binds every interval we report on that suite; we therefore repeated the protocol on ninety tasks, which removes that ceiling and changes two of our readings (Section 5). What binds the paper now is the run-level test rather than the suite: the ninety-task arms are three training runs each, enough for perfect separation between recipes but not to push an exact permutation test below its floor. Each further seed costs one training run per arm, and that is the most valuable next step for anyone repeating this. Our cross-policy comparisons originally rested on one training run per arm against a measured Greek seed spread of over thirty points. We have since replicated four arms at three seeds each, which resolved two comparisons and dissolved one; the arms still at (the warm-start policy, and the world-model dose ladder of Study 1) remain directional, and we mark them as such wherever they appear.
What remains open.
We state these as questions rather than retiring them, because each is answerable and none is answered here. (i) Does the language mixture matter beyond the presence of target-language data? Bilingual training beats target-only on average across three seeds each ( against ) but the ranges overlap and an exact run-level test returns , where is the floor at this design; we first read the single-run version of this comparison as evidence that the high-resource language scaffolds the low-resource one, and withdrew that reading. The ninety-task suite at three seeds per arm answers it in paired form: every bilingual seed’s margin over its own wrong-instruction floor ( to points) exceeds every target-only seed’s ( to ), at this design, and the Greek levels no longer overlap ( to against to ). On this evidence the mixture matters, though at a design whose run-level test cannot go below we state it as a direction rather than a demonstrated effect. What stays open is how much: three seeds per arm cannot push the run-level test below , and we do not report a level difference. (ii) Does any of this transfer beyond the simulator? Our translated corpus is dominated by a real-robot dataset that no experiment here consumes, because closed-loop success on it requires physical hardware; the action-prediction proxy on held-out episodes is reported in Section 6 and is negative for Greek at five percent of the intended training schedule. (iii) What sets the residual target–English gap? Our probes eliminate three explanations but leave representational depth and tokenizer fragmentation entangled. (iv) Does the phrasing-diversity benefit extend past ten goals? It replicates cleanly as a within-policy effect on the ten-goal suite. Our ninety-task runs do not test it: neither arm is phrasing-diverse, and building one needs a diverse ninety-task training set we have not built. It remains untested, and the compressed idiolect penalty we measure at ninety tasks makes it harder, not easier, to detect.
. Conclusion
| Do | What it bought, measured | Runs | |
|---|---|---|---|
| 1 | Check that the base model’s text tower already handles the language. It is free to test and decides the project. Tokenizer fertility is a cheap first look but not a sufficient test. | English tower: noise from Greek on every rung of a dose ladder. Multilingual tower, identical recipe: coherent Greek scenes on 64% of prompts. The failing tower fragments Greek less ( against ). | 1 each |
| 2 | Train the policy from the base checkpoint, not from a language-adapted world model. | Warm start 79.2 / 13.6% (EN / EL); from base 96.4 / 48.6%. | 1 |
| 3 | Append the translated instructions as pipe-separated variants beside the English ones. The loader samples 50/50; no code changes. | Greek margin over the wrong-instruction floor 6.7–7.1 points on every seed; Greek demonstrations alone: at most 2.7. | 3 per arm |
| 4 | Leave the text tower frozen. | Unfrozen: 82.4 / 18.0%, worse in both languages, replicated. | 3 |
| 5 | Train on several phrasings per task, not one. | Independent-translator penalty 23.3 12.3 points; its spread 16.4 2.2. | 3 per arm |
| 6 | Evaluate on a suite where each scene admits several goals, and always run the wrong-instruction control. | Single-goal suite: 84.6% Greek against 82.6% wrong. The control costs one evaluation run. | n/a |
| 7 | Evaluate on an independently produced translation. Run the same rewording in the high-resource language as a control, matched for how much it changes. | Our own paraphrases cost the same as the independent set (, vs ), so the case for an independent set is provenance, not a larger number. English rewording costs points with syntax alone and when it also substitutes the referring noun in seven of ten commands, against in Greek. | 3 |
| 8 | Train at least three seeds per arm; compare within-policy margins, never levels. | Greek moved 31.6 points on the seed at ten goals (2.4 at ninety); English moved 1.0. | n/a |
Adding a language to a robot foundation model is cheap where the base model’s tower already speaks it, and out of reach of every caption fine-tuning configuration we could afford where it does not. Given the right substrate, a corpus of machine-produced Greek and a data change that requires no code carry a policy from chance to roughly half of English performance on a ten-goal suite and two fifths of it on a ninety-task one, but only with demonstrations in both languages: target-language data is necessary, and on the larger suite a policy trained on it alone barely follows it, under three points over its own control on every seed. The two interventions that look most promising, initializing from a language-adapted world model and adapting the tower directly, both hurt.
What survives of the target language is partial and non-compositional, though less uniformly than we first reported. Greek success is sharply bimodal across goals at the seed we profiled: five goals at – and five at –. Profiling all three seeds shows the split is not a property of the recipe (Appendix B). The per-goal counts above are , and of ten, and only one goal (goal 0) stays below it in every seed, so the identity of the failing goals otherwise moves with the seed while four goals succeed in every one. We could not identify a lexical property separating the modes: every instruction in the set shares a content noun with at least one other. We would not have seen that structure from aggregate success rates, and we would not have trusted the aggregates at all without a control that current single-goal benchmarks make mandatory. For anyone localizing such a stack, the practical order (Table 4) is: check the tower first, since it is free to test and decides the project; supply demonstrations in the target language, since nothing else substitutes for them, and keep the high-resource ones alongside, since on our larger suite the target language alone bought almost nothing; leave the pretrained representation alone; and measure with a null you can guarantee rather than one you assume.
We close on the part we expect to outlast the numbers. Language adaptation of vision-language-action models was, when we began, thinly covered from where we could see: the multilingual literature largely stops where the output is text, and the robotics literature is written in English throughout. We expect that to change quickly. That means the first groups to work here will be calibrating instruments rather than beating baselines, and will do it without a body of prior negative results to warn them. This study is an attempt to supply some. Of the conclusions we drew over its course, five did not survive a control or a replication, and each failed in a way that would have been invisible to the measurement a reasonable practitioner would have chosen first. We would rather publish that record than a cleaner one we trust less, and we think the controls that produced it transfer further than any of our numbers do.
. Availability
The policy studied in Section 5 is available as Sophea-Nano-Policy-LIBERO-Greek-v1 at https://huggingface.co/KIEFERSA/Sophea-Nano-Policy-LIBERO-Greek-v1: the bilingual recipe at the seed whose Greek evaluated best of three ( against a three-seed mean of and a range of – on the ten-goal suite), as a Hugging Face safetensors export that loads directly into the Cosmos3 policy server, with a model card that reports the wrong-instruction control alongside the headline numbers and the commands to reproduce the three-way evaluation. A reader should treat that checkpoint as the maximum of a three-seed draw rather than the recipe’s expectation; the expectation is the mean, and Section 5 gives the spread. Scripts, dose-ladder configurations, training configurations and the evaluation harness accompany the paper. We are precise about what that reproduces: the pipeline and every evaluation, but not our exact numbers, because a reader’s own translation of the instructions is a different translator, and Section 5 measures what changing translator costs. Reproducing the numbers requires our instruction sets, which we will supply on request for research use. Two artifacts are deliberately not released: the Greek-adapted world model of Section 4, which matches the specific content of a Greek caption on none of the held-out scenes and should not ship on a coherence number alone, and the translated corpora.
Appendix A Every inferential comparison
The body makes twenty inferential comparisons and applies no multiplicity correction. Table 5 lists them so a reader can apply one. Two properties of this design matter more than any individual . First, the two families answer different questions and have different power: within-policy contrasts share a checkpoint, take the goal or task as the sampling unit, and reach at ten clusters; cross-policy contrasts take the training run as the sampling unit, and at three runs per arm the smallest attainable two-sided is . Second, no cross-policy contrast can therefore survive any correction, and we do not claim one does. Under Bonferroni at over twenty comparisons () only the within-policy task-level contrasts survive.
| Comparison | Family | Test / unit | Result | Survives? |
|---|---|---|---|---|
| Greek vs wrong instruction, single-goal suite | within | paired task-level, | , | no |
| Greek vs wrong, multi-goal suite | within | paired task-level, | yes | |
| English vs wrong, multi-goal suite | within | paired task-level, | yes | |
| Greek level, from-base bilingual | within | task-level bootstrap | CI | n/a |
| Greek vs wrong, LIBERO-90, bilingual, 3 seeds | within | paired task-level, | , CIs exclude | yes |
| Greek vs wrong, LIBERO-90, target-only, 3 seeds | within | paired task-level, | , of include | partly |
| Trained vs independent phrasings, single-phrasing arm | within | paired, per seed | yes | |
| Trained vs independent phrasings, diverse arm | within | paired, per seed | yes | |
| Trained vs held-out paraphrase (two sets) | within | paired, per seed | non-monotone; retracted | no |
| Trained vs paraphrased English, syntax only | within | paired, per seed | yes | |
| Trained vs paraphrased English, nouns substituted | within | paired, per seed | yes | |
| English vs Greek rewording penalty | within, paired | every-seed ordering, pairs | EN EL ; exact floor here is , not | no |
| Failure-mode split (English vs Greek) | within | per-episode goal predicates | vs | yes |
| Single-phrasing vs diverse penalty | cross | exact run-level, v | (floor) | no |
| Bilingual vs target-only, ten goals | cross | exact run-level, v | no | |
| Bilingual vs target-only margin, LIBERO-90 | cross | exact run-level, v | (floor) | no |
| Bilingual vs English-only (no-Greek ablation) | cross | exact run-level, v | separates | no |
| Frozen vs unfrozen tower | cross | exact run-level, v | separates | no |
| Warm-started vs from-base | cross | one run per arm | directional only | no |
| English cost of phrasing diversity | cross | task-level CI | CI | no |
| Loss vs Greek success across six policies | cross | correlation, | , exploratory | no |
Appendix B Per-goal profiles and the Greek instructions
| English | Greek, by seed | wrong | ||||
|---|---|---|---|---|---|---|
| # | Goal | ref. run | 42 | 43 | 44 | ref. run |
| 0 | open the middle drawer | 100 | 8 | 8 | 0 | 0 |
| 1 | bowl on the stove | 98 | 24 | 60 | 58 | 0 |
| 2 | wine bottle on cabinet | 98 | 0 | 94 | 0 | 0 |
| 3 | open top drawer, bowl inside | 92 | 24 | 86 | 92 | 0 |
| 4 | bowl on top of cabinet | 96 | 94 | 100 | 94 | 44 |
| 5 | push plate to front of stove | 98 | 72 | 96 | 96 | 0 |
| 6 | cream cheese in the bowl | 94 | 2 | 94 | 36 | 0 |
| 7 | turn on the stove | 100 | 90 | 98 | 78 | 0 |
| 8 | bowl on the plate | 100 | 98 | 98 | 98 | 0 |
| 9 | wine bottle on the rack | 88 | 74 | 68 | 0 | 0 |
| goals | 10 | 5 | 9 | 6 | 0 | |
Section 5 reports that Greek success is bimodal across goals and that the split moves with the seed. Table 6 gives the underlying profile for all three seeds of the bilingual recipe, together with the English and wrong-instruction columns of the reference run, and Figure 7 gives the ten Greek instructions so a reader can judge translation adequacy rather than take our word for it.
Two things in these tables constrain explanations we would otherwise be free to offer. First, the number of goals above is five, nine and six across seeds and only goal 0 stays below it in every seed, so “five succeed and five fail” describes one run, not the recipe. Second, the obvious mechanical explanation for a persistent per-goal failure, that the translation of that goal is wrong, does not hold for the one goal that persistently fails: open the middle drawer of the cabinet renders literally and unambiguously (Figure 7, row 0), and the independent translator produces the same sentence. Meanwhile the two instructions where the glossary does introduce a real ambiguity, goals 1 and 5, where stove becomes a word that also means kitchen rather than hob, score and : one weak, one strong. Translation adequacy therefore does not predict the per-goal profile in either direction here. We report this because it was the first explanation we reached for and it did not survive the strings.
References
- [1] Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. Don’t just assume; look and answer: Overcoming priors for visual question answering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. arXiv:1712.00377.
- [2] Mikel Artetxe, Gorka Labaka, and Eneko Agirre. Translation artifacts in cross-lingual transfer learning. In Empirical Methods in Natural Language Processing (EMNLP), 2020. arXiv:2004.04721.
- [3] Bowen Baker et al. Video PreTraining (VPT): Learning to act by watching unlabeled online videos. In Advances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2206.11795.
- [4] Kevin Black et al. : A vision-language-action flow model for general robot control. arXiv preprint, 2024.
- [5] Anthony Brohan et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023. Proceedings version (CoRL 2023, PMLR v229) lists Zitkovich et al.
- [6] Alexis Conneau et al. Unsupervised cross-lingual representation learning at scale. In Association for Computational Linguistics (ACL), 2020. arXiv:1911.02116.
- [7] Pim de Haan, Dinesh Jayaraman, and Sergey Levine. Causal confusion in imitation learning. In Advances in Neural Information Processing Systems (NeurIPS), 2019. arXiv:1905.11979.
- [8] Danny Driess et al. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better. In Advances in Neural Information Processing Systems (NeurIPS), spotlight, 2025. arXiv:2505.23705, verified 2026-08-26.
- [9] Yilun Du et al. Learning universal policies via text-guided video generation. In Advances in Neural Information Processing Systems (NeurIPS), 2023. arXiv:2302.00111.
- [10] Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. LIBERO-Plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626, 2025. Author list verified 2026-08-26.
- [11] Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. arXiv:1612.00837.
- [12] Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning (ICML), 2020. arXiv:2003.11080.
- [13] Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang, Johan Bjorck, et al. DreamGen: Unlocking generalization in robot learning through video world models. arXiv preprint arXiv:2505.12705, 2025.
- [14] Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, et al. DROID: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945, 2024.
- [15] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, et al. OpenVLA: An open-source vision-language-action model. In Conference on Robot Learning (CoRL), 2024. arXiv:2406.09246.
- [16] James Kirkpatrick et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 2017. arXiv:1612.00796.
- [17] Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding. In Empirical Methods in Natural Language Processing (EMNLP), 2020. arXiv:2010.07954.
- [18] Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. Fine-tuning can distort pretrained features and underperform out-of-distribution. In International Conference on Learning Representations (ICLR), 2022. arXiv:2202.10054.
- [19] Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. LIBERO: Benchmarking knowledge transfer for lifelong robot learning. In Advances in Neural Information Processing Systems, Datasets and Benchmarks Track, 2023. arXiv:2306.03310.
- [20] Corey Lynch and Pierre Sermanet. Language conditioned imitation learning over unstructured data. In Robotics: Science and Systems (RSS), 2021. arXiv:2005.07648.
- [21] Oier Mees, Lukas Hermann, Erick Rosete-Beas, and Wolfram Burgard. CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters, 2022. arXiv:2112.03227.
- [22] NVIDIA. Cosmos world foundation model platform for physical AI. arXiv preprint, 2025.
- [23] NVIDIA. Cosmos 3: Omnimodal world models for physical AI. Model release and technical report, 2026. https://research.nvidia.com/labs/cosmos-lab/cosmos3/.
- [24] Octo Model Team. Octo: An open-source generalist robot policy. In Robotics: Science and Systems (RSS), 2024. arXiv:2405.12213.
- [25] Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. Multilingual instruction tuning with just a pinch of multilinguality. In Findings of the Association for Computational Linguistics: ACL 2024, 2024. arXiv:2401.01854.
- [26] Jesse Thomason, Daniel Gordon, and Yonatan Bisk. Shifting the baseline: Single modality performance on visual navigation & QA. In North American Chapter of the Association for Computational Linguistics (NAACL), 2019. arXiv:1811.00613.
- [27] Homer Walke, Kevin Black, Abraham Lee, Moo Jin Kim, Max Du, Chongyi Zheng, Tony Zhao, Philippe Hansen-Estruch, Quan Vuong, Andre He, Vivek Myers, Kuan Fang, Chelsea Finn, and Sergey Levine. BridgeData V2: A dataset for robot learning at scale. In Conference on Robot Learning (CoRL), 2023.
- [28] Albert Webson and Ellie Pavlick. Do prompt-based models really understand the meaning of their prompts? In North American Chapter of the Association for Computational Linguistics (NAACL), 2022. arXiv:2109.01247.
- [29] Mitchell Wortsman et al. Robust fine-tuning of zero-shot models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. arXiv:2109.01903.
- [30] Hongtao Wu et al. Unleashing large-scale video generative pre-training for visual robot manipulation. In International Conference on Learning Representations (ICLR), 2024. arXiv:2312.13139.