Paper deep dive
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
Varvara Arzt, Allan Hanbury, Terra Blevins
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/18/2026, 5:40:32 AM
Summary
This study investigates word order preferences in decoder-only language models by comparing performance on 192 artificial languages and typologically diverse natural languages. The authors find that models exhibit a strong left-branching preference on artificial languages, which contradicts cross-linguistic universals. In contrast, on natural languages, models develop a preference for right-branching SVO languages as training data scales, despite SOV being the most frequent order cross-linguistically. This SVO advantage correlates with language resource levels and data quality rather than inherent architectural biases, suggesting that observed word order biases are data-driven. The widespread adoption of LLMs trained on imbalanced, SVO-heavy data risks reducing global word order diversity.
Entities (16)
Relation Signals (8)
Word Order Bias → isdrivenby → Data
confidence 95% · establishing that word order biases observed in practice are data-driven.
SOV → ismostfrequentcrosslinguistically → True
confidence 95% · SOV falls behind despite being the most frequent order cross-linguistically.
SVO → correlateswith → High Resource Level
confidence 90% · This SVO advantage... correlates with language resource level and data quality rather than word order.
Goldfish → evaluatedon → Natural Languages
confidence 90% · monolingual Goldfish models... on natural languages
GPT-2 → evaluatedon → Artificial Languages
confidence 90% · GPT-2 on artificial languages (PPL)
Transformer → exhibitspreferencefor → Left-Branching
confidence 90% · On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases.
Transformer → exhibitspreferencefor → SVO
confidence 85% · On natural languages... a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind
LLM → risksreducing → Word Order Diversity
confidence 85% · these biases risk gradually reducing word order diversity... with the widespread adoption of LLMs.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.
Tags
Links
- Source: https://arxiv.org/abs/2608.15129v1
- Canonical: https://arxiv.org/abs/2608.15129v1
Trouble viewing inline? Open PDF directly →
Full Text
93,502 characters extracted from source content.
Expand or collapse full text
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models Varvara Arzt Allan Hanbury Terra Blevins Affiliation: Khoury College of Computer Sciences, Northeastern University[0.3em] Correspondence:varvara.arzt@tuwien.ac.at [0.5em] Faculty of Informatics TU Wien Abstract We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.11 1 Code and data are available at https://github.com/kleines-gespenst/word-order-preferences 1 Introduction Figure 1: Head directionality vs. model performance across attested, artificial, and natural languages. x-axis: right-branching switch count (§3.1), from fully left-branching (0R/5L) to fully right-branching (5R/0L). ↓ = better in (b)–(d); colours denote base word order (legend in (a)). (a) Attested language frequency by base word order. (b) GPT-2 on artificial languages (PPL). (c, d) Monolingual models on natural languages at two training data scales (BPEC; §3.4). Word order is one of the most fundamental dimensions of cross-linguistic variation. Languages vary widely in how they linearise both clause-level constituents (e.g., SVO in English, SOV in Japanese) and internal constituent orderings, such as the placement of adpositions and adjectives relative to their heads; however, many of these ordering choices tend to correlate (20; 15; 14; 53) with, for instance, SOV languages typically using postpositions and SVO languages prepositions, placing languages on a scale of directionality (Figure 1a) from left-branching, where dependents precede their heads (e.g., postpositions), to right-branching, where heads come first (e.g., prepositions). Beyond dominant orders, many languages exhibit productive word order variation that serves discourse-pragmatic functions such as topicalisation and focus marking (23). As language models (LMs) are increasingly deployed across typologically diverse languages, a natural question arises: do these models develop systematic word order preferences, and if so, are such preferences architectural or data-driven? This question is especially pressing due to the potential for LLM-mediated language homogenisation (27; 18) and LLM bias towards more formal writing styles (1), which tend to favour rigid word orders. Prior work has used artificial languages generated from formal grammars (55, e.g.,) and synthetic variants of natural languages (47, e.g.,) to study the word order preferences of LMs. However, artificial languages lack the semantic, morphological, and discourse-level complexity of natural languages, so findings from controlled settings may not generalise; conversely, evaluation on natural languages alone cannot isolate word order from confounds such as morphology, data quality, and script. We combine both approaches at multiple training data scales to address this gap, providing three methodological contributions: (1) We train models on 192 artificial languages spanning all six base word orders as well as 32 internal constituent orderings, analysing directionality preferences, alignment with cross-linguistic universals, and constituent-level surprisal at multiple training data scales. (2) We similarly evaluate monolingual and multilingual models trained on typologically diverse natural languages, analysing how word order preferences and constituent surprisal scale with growing data. (3) By mapping artificial language configurations to natural language typological features and comparing models trained at matched data scales across both paradigms, we unify artificial and natural evaluations of word order preference into a single, controlled comparison, allowing us to disentangle architectural and data-driven biases. We find that on artificial languages, models exhibit a left-branching preference (Figure 1b) that aligns with neither cross-linguistic universals nor human learners’ preference for consistent directionality (13). Constituent surprisal analyses suggest this preference may partly reflect the entropy structure of a semantics-free grammar rather than a model-inherent architectural prior. However, these base word order rankings can shift with data volume, and directionality preferences reverse early in training. In contrast, while no word order preference is visible at 5 MB (Figure 1c) for natural languages, we find that SVO languages are easiest and SOV hardest in monolingual models at 1000 MB (Figure 1d), implicating data quality rather than quantity alone. This SVO advantage extends to multilingual models and disappears only when this word order is dominated by very-low-resource languages, as in BLOOM (7). These findings show that, in practice, models consistently favour the word orders of highly-resourced languages, particularly at scale. Since highly-resourced languages are overwhelmingly SVO (e.g., English and Chinese; Appendix A.4), this creates a precondition for silent typological homogenisation: a gradual reduction of word order diversity driven by imbalanced training data. Our results suggest that this precondition is real, robust across model types and training regimes, and in principle addressable through data curation. 2 Related Work LM word order preferences have mainly been studied through artificial languages and synthetic variants of natural languages. 55 train LMs on languages generated from a probabilistic context-free grammar (PCFG) varying in branching directionality, finding a preference for left-branching configurations and best performance on both widely attested SOV and rarely attested OVS, suggesting no alignment with typological frequency; our artificial language experiments build upon their approach (§3.1). 32 extend this to cognitively-motivated LMs, finding that models with parsing strategies and memory limitations achieve lower perplexity (PPL) on typologically frequent word orders. 17 adopt Generalized Categorial Grammar, covering all six S/V/O linearisations, and find that typologically plausible orders facilitate length generalisation. In emergent communication, 35 show that neural agents converge from an equal-probability mixed-order language toward a single dominant order through communicative pressure alone, an effect that strengthens with larger agent group size (36). A complementary line of work creates synthetic variants of attested languages. 47 create synthetic English variants with altered word order and case marking, finding that RNN agreement prediction reflects recency bias. 8 similarly alter word order flexibility and case marking in English to study neural machine translation, finding that translating flexible-order languages is harder in low-resource settings. 25 show that GPT-2 assigns higher PPL to English variants with disrupted hierarchical structure, though 59 find this does not generalise across nine diverse languages. 10 find that real word orders distribute information more uniformly than counterfactual alternatives, with the effect strongest for SVO. In concurrent work, 57 find that disharmonic (i.e., typologically inconsistent in directionality) variants of English and Japanese based on five Greenbergian universals22 2 Cross-linguistic generalisations about word order co-occurrences (20), e.g., SOV ↔ postpositions. are learned more slowly but reach comparable final performance; we test a partly overlapping set of universals (§3.4). No prior work, to our knowledge, combines controlled artificial and natural language experiments to analyse word order preferences; we do so here. 3 Methodology To disentangle potential architectural biases from effects of training data, we perform parallel experiments on artificial and natural languages, using statistics on attested33 3 Documented natural languages languages as a reference for whether model preferences correlate with human biases. Artificial languages allow us to analyse all possible word order permutations across syntactic constructions in isolation, avoiding confounds from typological dimensions like morphological agreement (55). Experiments with natural languages address whether findings on artificial languages generalise, and help establish whether word order preferences can be reliably studied with artificial languages (16; 59). 3.1 Artificial Languages Each language is encoded as [Base]-[5-char], where the base specifies the S/V/O order and each character selects a binary parameter: L = left-branching, R = right-branching. Switch L (left-branching) R (right-branching) 1 VP_Comp SCompVerbCompS_Comp\ Verb_Comp VerbCompSCompVerb_Comp\ S_Comp 2 Comp SCompS\ Comp CompSComp\ S 3 P NPPostpNP\ Postp PrepNPPrep\ NP 4 NP AdjNounAdj\ Noun NounAdjNoun\ Adj 5 Rel VPRelNounVP\ Rel\ Noun NounRelVPNoun\ Rel\ VP Example: SOV-L ≈ Japanese Feature Value Base (SOV) Subject–Object–Verb Switch 1 = L complement clause before verb Switch 2 = L clause before complementiser Switch 3 = L postpositions Switch 4 = L adjective before noun Switch 5 = L prenominal relative clause 6 base orders × 252^5 configurations = 192 languages. Figure 2: Language encoding scheme forestforest Figure 3: Same sentence as SOV-L (≈ Japanese, top) and SVO-RRRLR (≈ English, bottom): same hierarchical structure, but word order differs. Verbs & complementiser highlighted; subj/obj = overt case markers. We generate artificial languages building upon 55. Each language is defined by a base S/V/O order and five binary head-direction switches controlling the ordering of complement clauses, complementisers, adpositions, adjectives, and relative clauses (Figure 2). Each switch selects between an L (left-branching, head-final) variant, in which the dependent precedes the head, and an R (right-branching, head-initial) variant, in which the head precedes the dependent. A base PCFG with a Zipf-distributed lexicon generates bracketed parse trees; main clause constituents are then deterministically reordered and head-direction switches applied, so that all 192 variants share identical derivation probabilities and differ only in word order. Figure 3 illustrates this with the same sentence under SOV-L (≈ Japanese) and SVO-RRRLR (≈ English): the hierarchical structure is preserved while the surface linearisation changes. We extend 55 in three ways: (1) we add an explicit switch for clausal complement order, which 55 tie to nominal object order, making them inseparable by design;44 4 Nominal and clausal object placement are typologically correlated (15) but not always identical; e.g., the Agob-Ende-Kawam language has preverbal nominal but postverbal clausal objects (50, GB135). (2) we add a six-way base word order parameter covering all six S/V/O linearisations, as does 17 but via a different formalism. Combining six base orders with 252^5 switch configurations yields 192 artificial languages with fixed word order; and (3) we scale the vocabulary from ∼ 1,400 to 50k pseudowords generated by Wuggy (26) from the most frequent English words in wordfreq (51), with Zipfian sampling weights (α=1.0α=1.0) per part-of-speech class (details in Appendix A.1). 3.2 Natural Languages Word Order Clustering To enable comparison between artificial and natural language experiments, we classify each natural language using the same encoding scheme applied to our artificial languages. We obtain word order labels from the typological databases WALS (14), Grambank (49), and APiCS (39) and map them to our artificial language labels (e.g., English → SVO-RRRLR, Japanese → SOV-L). We supplement database labels with manual verification grounded in peer-reviewed publications devoted to individual languages, and use lang2vec (38) and URIEL+ (28) for language clustering and feature extraction. Languages with non-binary or missing values for a given feature were excluded from our analysis. Full feature mapping procedure and coverage statistics are given in Appendix A.2. We note that typological database classifications are typically based on grammar descriptions rather than corpus frequencies (34; 23), and a single categorical label may obscure gradient word order variation (34); see § Limitations for further details. Data We evaluate on two datasets. Our primary evaluation set is FLORES-200 (43), which provides parallel sentences across a typologically broad set of languages. We also evaluate on the Parallel Universal Dependencies (PUD) treebanks (58), whose languages are all covered by FLORES-200, which both verify our FLORES-200 findings and provide gold-standard syntactic annotations for the per-token surprisal analysis, complementing automatic Stanza (46) parses on FLORES-200. Full list of languages used in §5 is given in Appendix A.2. 3.3 Models For artificial languages, we train a separate GPT-2-style decoder-only model on each of the 192 languages to evaluate whether word order preferences arise from the transformer architecture itself. For natural languages, we primarily focus on monolingual Goldfish models (9), GPT-2-based models with a 50k-token vocabulary, scaled in parameter count with training data volume, trained on one of up to 350 languages in a maximally comparable setup at four training data volumes (5 MB, 10 MB, 100 MB, and 1000 MB of content-equivalent text per language after byte premium scaling; 9)55 5 Byte premium scaling adjusts raw data sizes by UTF-8 byte ratio relative to English (4): e.g., the 5 MB Burmese model was trained on 25 MB raw text (https://huggingface.co/goldfish-models/mya_mymr_5mb)., enabling systematic analysis of how word order preferences develop with growing amounts of data. To verify whether these patterns generalise to multilingual training, we additionally evaluate mGPT (1.3B and 13B parameters, ∼ 400B tokens across 61 languages with no exact per-language statistics available; (48)), BLOOM (560M, 3B, and 7.1B parameters, 341B tokens across 46 languages, median 527 MB per language; (7; 33)) and XGLM (564M and 1.7B parameters, ∼ 500B tokens across 30 languages with exponential upsampling of low-resource languages; (37)). All three multilingual models were designed to broaden language coverage beyond English-centric corpora and document their training data composition, allowing us to verify which languages each model was trained on and to assess potential contamination with FLORES-200 and PUD. 3.4 Evaluation We evaluate word order preferences at two levels: clause-level S/V/O linearisation and internal constituent order (the five switches in Figure 2), aggregating the latter into a directionality scale. For artificial languages, we report test perplexity (PPL). Following 25 and 59, we additionally analyse training dynamics by tracking test perplexity every 50 steps and computing the area under the PPL curve (AUC); this is not possible for Goldfish models as no intermediate checkpoints are available. We also test whether model preferences align with a subset of Greenbergian universals (20) testable in our setup: subject-object ordering, base word order and adposition type, adjective placement, and verb-object order with relative clause placement (15) (details in Appendix A.3). For natural languages, where different tokenisers make raw PPL incomparable across languages, we use bits-per-English-character (BPEC; Appendix A.3) as defined in 12. Throughout, a language is easier (harder) for a model if it achieves lower (higher) PPL/BPEC. For both artificial and natural languages, we evaluate per-token surprisal (56) to analyse whether tokens with different syntactic functions are differentially predictable across word orders. 4 Experimental Setup 4.1 Artificial Languages We train GPT-2-style decoder-only models using the HuggingFace Transformers library66 6 https://huggingface.co/docs/transformers (GPT2LMHeadModel) at three training data volumes: the original 55 setup (∼ 0.5 MB, 10K sentences), 5 MB, and 10 MB. As 0.5 MB results are comparable to 5 MB, and we aim for comparability with the monolingual Goldfish models, we focus on the 5 MB and 10 MB settings. The model architecture follows the Goldfish configuration at comparable sizes: 4 layers and 8 attention heads (9). Unlike 55, who used a whitespace tokeniser (under which the vocabulary is effectively arbitrary), we train a BPE word-boundary SentencePiece tokeniser (31) with a vocabulary size of 15k (details in Appendix A.1). Following 55, we train each model on 10 fully disjoint train/dev/test splits generated from independent data and report averaged test perplexity to ensure robustness of results. Implementation details, hyperparameters and compute cost reported in Appendix A.1. 4.2 Natural Languages We evaluate Goldfish models, available on HuggingFace,77 7 https://huggingface.co/goldfish-models at all four training data sizes and the multilingual models introduced in §3.3 on FLORES-200 (43) (Appendix A.2) and PUD (58). Syntactic annotations for the surprisal analysis come from Stanza (46) on FLORES-200 and gold UD annotations on PUD. For multilingual models, we restrict evaluation to the subset of each model’s training languages overlapping with FLORES-200 (language counts per model in Table 2). 5 Results We compare LM word order preferences on artificial languages (§5.1), where models show a strong left-branching preference misaligned with cross-linguistic universals, and natural languages (§5.2), where training data composition increasingly shapes model preferences. Figure 4: LM preferences on 192 artificial languages (6 base orders × 32 branching configurations), averaged over 10 runs on 5 MB data. (A) PPL per base order across configurations, sorted by across-order mean (dotted line). Shaded bands: ±1± 1 SD. Vertical markers (RRRLL, LLRRR): SV and VS base orders respond inversely to clause-level switches. (B) Paired switch effects (ΔPPL ) for configuration pairs differing only in the target switch. ⧫ : overall mean, where >0=>0= left-branching is preferred. Switch effect sizes dzd_z are all p<.001p<.001. 5.1 Artificial Languages On artificial languages, models consistently prefer left-branching configurations, conflicting with cross-linguistic universals, and base word order rankings shift unstably with data volume. Directionality Preferences Figure 1b shows that the mean PPL across settings increases consistently with the number of right-branching switches in an artificial language across all six base word orders, with only slight decreases in PPL for majority right-branching settings in SVO and SOV languages. This preference for left-branching configurations corroborates prior work (55; 32; 16). However, the top-ranked configuration for each base order has no close natural language counterpart,88 8 A partial exception is SVO-RRRLL (≈ Mandarin Chinese), though the prenominal relative clause (switch 5 = L) is very rare in VO languages (15; 11). suggesting that transformers do not mirror human word order learning biases. In contrast, attested languages cluster at both extremes of the directionality scale (Figure 1a), consistent with the human harmony bias toward consistent directionality (13). The magnitude of the left-branching preference also varies substantially across individual switches (Figure 4B). Switch 5 (relative clause) dominates: all sixteen left-branching relative clause configurations are easier than the right-branching ones, at Δ of 4–7 points on 5 MB. This strong preference for prenominal relatives runs counter to the typological distribution, where postnominal relatives outnumber prenominal ones approximately 4:1 (15; 14). Other switches show smaller but consistent trends; for example, Switch 4 (pre- or postnominal adjectives) shows a moderate preference for adjectives preceding the noun. All five paired t-tests on the effect of internal constituent switches are significant even after a strict Holm–Bonferroni correction to remove false positives (p<.001p<.001). More broadly, switches that reorder clause-level constituents (relative clauses, complement clauses) produce larger perplexity differences than switches that rearrange phrase-internal elements (adjective–noun, adposition–NP), suggesting greater sensitivity to the direction of long-range dependencies. Base Word Order Model preferences of base word orders are not stable across training data volumes and also do not reflect typological frequency (Appendix Figure 8). At 5 MB, SOV ranks first and VSO last in mean PPL across configurations (Figure 4A), consistent with 55. At 10 MB, the three VS orders rank 1–3 and the three SV orders rank 4–6, with SVO the hardest (Appendix Figure 6), which is opposite the trend we see on natural languages (where SVO is easiest to model at scale; §5.2). Internal constituents also interact with base word order, with individual base order curves crossing over for specific switch settings (Figure 4A): SV bases (SVO, SOV, OSV) and VS bases (VSO, VOS, OVS) respond differently to individual switches. Right-branching verb complements and complementisers (switches 1–2 = R) selectively benefit SV bases by up to 5.5 PPL points (vertical marker RRRLL in Figure 4A), while right-branching relative clauses (switch 5 = R) reverse this, favouring VS bases by up to 3.8 points (vertical marker LLRRR). This observed SV/VS split does not align with the VO/OV distinction underlying most implicational word order universals (20; 15; 14), and direct tests of typologically expected combinations confirm this: the tested universals are either not reflected (adposition, adjective placement) or significant in the opposite direction (subject–object order and relative clause placement, at 10 MB). These results indicate that the inductive biases demonstrated by transformers on artificial languages do not match cross-linguistic word order universals. Training Dynamics Figure 5: PPL Δ (right −- left branching) during training on 5 MB artificial data. SVO: right-branching is initially easier before reversing to a left-branching advantage by step 100. VSO: left-branching always dominates. The model’s directionality preference shifts in the earliest phase of training, again following the SV/VS split (Figure 5 for SVO and VSO, with other orders following this trend in Appendix Figure 7). For SV base orders, right-branching configurations are initially easier (by up to 28 PPL points) before fully reversing to a left-branching advantage by step 100 (epoch ∼ 2.6); for VS orders, left-branching dominates from the first checkpoint. Throughout training, the internal constituent order switches rather than the base S/V/O order drive the PPL spread, with convergence speed and final PPL strongly correlated (ρ>0.89ρ>0.89; Appendix Figure 11). Moreover, the early reversal for SV languages echoes the preference shift observed by 16 and indicates that directionality preferences may be sensitive to training dynamics and data quantity, which should be considered when comparing results across studies. 5.2 Natural Languages While LMs show no word order preference in small training data scales on natural languages, an SVO advantage emerges at scale in both mono- and multilingual models, correlating with language resource level and data quality rather than word order type. Base order 5 MB 10 MB 100 MB 1 GB SVO (32) 2.13 (.15) 2.00 (.15) 1.55 (.08) 1.32 (.06) SOV (22) 2.13 (.10) 2.02 (.10) 1.64 (.11) 1.45 (.12) VSO (3) 2.20 (.07) 2.06 (.09) 1.64 (.06) 1.42 (.06) NoDom (11) 2.13 (.11) 2.00 (.10) 1.56 (.08) 1.33 (.05) SVO–SOV δ 0.000.00 −0.10-0.10 −0.60∗-0.60^*** −0.76∗-0.76^*** At 5 MB: SVO = 2.1288, SOV = 2.1289, NoDom = 2.1284. Table 1: Median BPEC and (interquartile range; IQR) for base word orders of 68 monolingual Goldfish models (evaluated on FLORES-200, lang. count in parentheses). ↓ = better, bold = best per column. Bottom row: SVO–SOV effect size (Cliff’s δ); order preference is absent at 5 MB (Mann–Whitney U=353U=353, p=.99p=.99) and large by 1 GB (U=84U=84, p<10−5p<10^-5). ∗p<.001^***p<.001. Directionality Preferences Figure 1c demonstrates that at 5 MB, LMs trained on natural languages exhibit no clear directionality preference; however, by 1000 MB, right-branching SVO languages become easiest (Figure 1d; intermediate scales in Appendix Figure 9).99 9 Here we restrict to the 68 languages with Goldfish models at all four training data sizes. When all available languages at each size are included (Appendix Figure 10), SVO even underperforms SOV at small scales. This may reflect the same left-branching inductive bias observed in artificial languages, overridden by data effects in natural languages as training data grows. Base Word Order Preferences across Scale This shift across data scales is also visible at the level of base word order. At 5 MB, monolingual Goldfish models show no preference: SVO, SOV, and NoDominant all centre around BPEC 2.13, with no significant difference between orders (Table 1). As data grows, an SVO advantage emerges and strengthens monotonically (SVO–SOV Cliff’s δ=−0.76δ=-0.76 at 1 GB): by 1 GB, SOV incurs ∼ 10% higher compression cost (Δ ≈ 0.13), despite being the most frequent order cross-linguistically (Appendix Figure 8). All four model sizes are evaluated on the same 68 languages, so the shift is driven solely by training data volume, not language selection. The same pattern holds on PUD (details in Appendix Table 7). Because word order correlates with other language properties, the observed SVO advantage may also be an artefact of a confound such as morphological complexity (3) rather than a genuine word order preference. To test this, we fit a linear mixed-effects model on the BPEC of the SVO/SOV languages in Figure 1c,d, with a word order × size interaction, per-language random intercepts, and three covariates: morphological complexity (subword MATTR; 52), language resourcedness (24), and training data composition (OSCAR web-crawl vs. curated share; Appendix A.5). With all three covariates included, the growth of the SVO–SOV gap with data survives every control (β: +0.049→+0.038+0.049→+0.038, p≤.001p≤.001; Appendix Table 6). Language resourcedness is the only covariate with an observable effect on this growth (∼ 20%), but we note it is often strongly correlated with other, likely important factors we did not control for, such as data quality, which disproportionately hurts low-resource languages (44; 54)1010 10 Within the SVO group of Goldfish models, the easiest language shifts with scale: Indonesian has the lowest BPEC at 5 and 10 MB, with English just behind, but English leads at larger scales, consistent with higher-resource languages benefiting more from cleaner data as data grows. (see § Limitations). Script is a further possible confound, since SOV and SVO languages differ systematically in writing system: in FLORES-200, SOV languages are non-alphabetic far more often than SVO (71% vs. 19%). Restricting to segmental alphabets, however, the gap still grows with data (29 SVO / 7 SOV: near zero at 5 MB to +0.057+0.057 at 1 GB, 95% CI [+0.011,+0.131][+0.011,+0.131]). Beyond the aggregate BPEC results, per-constituent surprisal further confirms that preferences are data-driven: the ranking reverses from S << O << V in semantic-free artificial to V << O << S in natural languages both on FLORES-200 and PUD data, with the same architecture yielding opposite profiles depending on whether semantic dependencies are present. Multilingual Models Model SVO SOV NoDom δ BLOOM-560m 3.62 (2.86) 2.59 (.59) — +.26+.26 BLOOM-3b 2.99 (2.20) 1.88 (.34) — +.27+.27 BLOOM-7b1 2.81 (2.06) 1.82 (.29) — +.25+.25 XGLM-564M 1.46 (1.06) 1.89 (4.69) 1.47 (.63) −.49∗-.49^* XGLM-1.7B 1.41 (.88) 1.78 (5.12) 1.42 (.54) −.53∗-.53^* mGPT-1.3B 1.53 (.28) 1.74 (.25) 1.64 (.31) −.43∗-.43^* mGPT-13B 1.44 (.25) 1.68 (.18) 1.54 (.24) −.44∗-.44^* Table 2: Median FLORES-200 BPEC and (IQR) by base word order for multilingual models. Language counts (SVO/SOV/NoDom): BLOOM 28/13/0, XGLM 14/10/3, mGPT 25/22/9. ↓ = better, bold = best per model. Last column: SVO–SOV effect size (Cliff’s δ); SVO is preferred by XGLM and mGPT (δ<0δ<0) but not BLOOM (δ>0δ>0), whose SVO group is very-low-resource and heterogeneous (wide SVO IQR). ∗p<.05^*p<.05. The SVO advantage and the right-branching directionality preference both extend to XGLM and mGPT (Table 2; Appendix Figure 12), where SVO groups are dominated by highly-resourced Indo-European languages.1111 11 On an alphabetic script subset the multilingual SVO advantage persists in mGPT (22 SVO / 10 SOV, SVO vs. SOV Mann–Whitney p=.027p=.027 at 1.3B, .030.030 at 13B; Cliff’s δ≈−0.50δ≈-0.50). BLOOM and XGLM have too few alphabetic SOV languages (1–2) to test. BLOOM is the exception: 21 of its 28 SVO languages are Niger-Congo languages, most very-low-resource, with a median training data volume of 1.7 MB, compared to 3.7 GB for all other languages1212 12 The median for the remaining non-Niger-Congo SVO languages in BLOOM is 79 GB., inflating the SVO group’s median BPEC and masking the advantage. More broadly, SVO languages receive substantially more training data than SOV languages across all three model families: in BLOOM, the median SVO/SOV ratio is 11× (33), reduced to 5× in XGLM1313 13 https://huggingface.co/facebook/xglm-564M after upsampling. The preference thus reflects resource scale rather than word order. All three model families (mGPT, XGLM, and BLOOM) were explicitly designed to broaden language coverage beyond English-centric corpora, yet the SVO advantage persists wherever the SVO group is not dominated by very-low-resource languages; in BLOOM, the wide interquartile range (IQR) of SVO languages (Table 2, Appendix Figure 12) reflects the heterogeneity between the high-resource non-Niger-Congo and the very-low-resource Niger-Congo languages within this group. 5.3 Comparison of Artificial and Natural Language Results At small training scales, artificial and natural language results partially align: models show a left-branching advantage and no clear base word order preference in both settings. As data grows, however, the two settings consistently diverge across base word orders, internal constituent switches, and constituent-level surprisal trends. The left-branching advantage is stable across data sizes in artificial languages but vanishes in natural languages at scale, where right-branching SVO languages become easiest. Base word order rankings also reverse; SVO is hardest in artificial languages at 10 MB (Appendix Figure 6) but easiest in natural languages at 1 GB (Table 1). Per-constituent surprisal also flips, consistent with semantic selectional restrictions reshaping predictability in natural languages (19). These contrasts across artificial and natural languages establish that word order preferences in transformer LMs are shaped by data rather than architecture; our multilingual results further reinforce this, with the models’ SVO preference disappearing only when very-low-resource languages dominate this word order. Even monolingual Goldfish models trained on content-equivalent amounts of data per language (4) favour the SVO order most commonly seen in highly-resourced languages at scale (Appendix A.4), suggesting that data quantity alone does not fully explain this preference. Instead, the SVO advantage arises from properties of the training data beyond word order: it survives controls for morphological complexity, script, and, in the coarse form we can measure, training data composition, with only language resourcedness partly explaining it (§5.2), which is itself a proxy for data quality (44; 54); finer aspects of data quality we leave to future work (see § Limitations). 6 Conclusion We present a systematic analysis of word order preferences in decoder LMs across artificial and natural languages at multiple training data scales, combining both to disentangle architectural biases from data-driven preferences. On artificial languages, models prefer left-branching configurations with no alignment with cross-linguistic universals, and SVO languages become hardest at scale; on natural languages, no preference emerges at small scales, but LMs exhibit a preference for SVO languages as data grows while SOV falls behind. These preferences are data-driven: the same architecture exhibits opposite rankings on artificial and natural languages, and the preference for SVO order in natural languages correlates with resource level and data quality rather than word order type. Crucially, even where architectural biases exist (e.g., the left-branching preference on artificial languages), they are overridden by data effects at scale. Since highly-resourced languages are overwhelmingly SVO, this data-driven advantage has real typological consequences. Given that 1 GB of training data suffices for the SVO preference to emerge even in monolingual models, the effects we study are likely more pronounced in frontier models where the training data quantity and quality gaps between high- and low-resource languages are even larger. This creates a precondition for what we term silent typological homogenisation: a hypothesised gradual reduction of word order diversity in languages that productively use multiple base word orders towards that of the dominant languages. To confirm this hypothesis and better understand the implications of LM word order preferences, we argue that extending this analysis to larger models and spoken language data, testing whether finetuning or downstream task application alters word order preferences (5), and applying mechanistic interpretability to identify how the observed order preferences are encoded (45) remain important directions for future work. Limitations The scope of our experiments is necessarily limited. Our PCFG production rules, following 55 and 32, cover a limited subset of the typological landscape: relative clause placement, for instance, cannot be fully captured by a binary switch given the range of attested strategies (14). Our experiments focus on the GPT-2 architecture, consistent with prior work (55; 25) and motivated by direct comparability with Goldfish models (9), which share the same architecture and hyperparameters. We did not evaluate larger multilingual models beyond mGPT, BLOOM, and XGLM due to insufficient transparency regarding per-language training data composition, which is essential for our analysis of resource-level effects. Our language sample may also be biased toward Indo-European languages, which are overrepresented in both typological databases and training data. Our evaluation data and metrics also carry certain limitations. We evaluate on written, relatively formal corpora (FLORES-200, PUD), but spoken language shows substantially greater word order flexibility (23), and register and genre strongly influence constituent ordering, so our results may not generalise to spoken or informal registers. FLORES-200 and PUD consist of translated text, which may introduce translationese biases (29). Metrics such as PPL and BPEC, while standard, measure overall compression rather than word order preferences directly; behavioural analysis of model outputs would provide more direct evidence. Our data quality analysis also relies on the observed correlation between language resourcedness (24, which we find the SVO advantage tracks) and data quality. While higher-resource languages tend to have cleaner, more diverse corpora (30; 44; 54), and web-crawled data is noisiest for low-resource languages (30), our training data composition variable (OSCAR web-crawl share) remains a coarse proxy, as we do not have direct annotations for data genre or quality. Isolating which finer aspects of data quality drive the effect requires per-language annotation of register, source, and corpus cleanliness, which we leave to future work. Finally, our analysis relies on categorical word order labels from databases such as WALS (14) and Grambank (49), which encode features in simplified, often binary ways. WALS feature 81A classifies dominant word order based on main clause order alone, with no representation of variability across clause types or the pragmatic conditions under which ordering alternations arise. This can introduce noise: Modern Standard Arabic is classified as VSO based on a 1958 source, whereas more recent work documents flexible VSO/SVO alternation (2; 21); German instead is classified as having no dominant order in WALS despite arguably SVO-dominant main clauses, the very criterion WALS uses for classification, while exhibiting SOV in subordinate clauses and clause-final non-finite verbs in auxiliary and modal constructions. Similarly, Russian is classified as SVO-dominant and Belarusian as having no dominant order in WALS, despite comparable word order flexibility in UD corpora (42). Word order variation is arguably gradient rather than categorical (34), and a single label may obscure the flexibility of a language’s actual word order distribution; even languages labeled with a single dominant order contain sentences with alternative orderings in the evaluation data. Resources such as APiCS partially address this by encoding word order frequencies, but the underlying data are themselves difficult to obtain: reliable estimates require representative corpora, consistent annotation standards across languages, and large-scale annotation by linguists with near-native competence, none of which is straightforward to achieve. More broadly, many typological features are not binary: Afrikaans, for example, has prepositions, postpositions, and circumpositions. References Abdulhai et al. (2026) M. Abdulhai, I. White, Y. Wan, I. Qureshi, J. Leibo, M. Kleiman-Weiner, and N. Jaques How llms distort our written language. External Links: 2603.18161, Link Cited by: §1. Arad Greshler et al. (2017) T. Arad Greshler, N. Melnik, and S. Wintner Seeking control in Modern Standard Arabic. Glossa: a journal of general linguistics 2 (1), p. 90. External Links: Document, Link Cited by: Limitations. Arnett and Bergen (2025) C. Arnett and B. Bergen Why do language models perform worse for morphologically complex languages?. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, p. 6607–6623. External Links: Link Cited by: §5.2. Arnett et al. (2024) C. Arnett, T. A. Chang, and B. K. Bergen A bit of a problem: measurement disparities in dataset sizes across languages. In Proceedings of the Annual Meeting of the Special Interest Group on Under-Resourced Languages, External Links: Link Cited by: §5.3, footnote 5. Belinkov and Glass (2019) Y. Belinkov and J. Glass Analysis methods in neural language processing: a survey. Transactions of the Association for Computational Linguistics 7, p. 49–72. External Links: Document Cited by: §6. Bentz et al. (2016) C. Bentz, T. Ruzsics, A. Koplenig, and T. Samardžić A comparison between morphological complexity measures: typological data vs. language corpora. In Proceedings of the Workshop on Computational Linguistics for Linguistic Complexity (CL4LC), D. Brunato, F. Dell’Orletta, G. Venturi, T. François, and P. Blache (Eds.), Osaka, Japan, p. 142–153. External Links: Link Cited by: §A.5. BigScience Workshop et al. (2023) BigScience Workshop, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, M. Gallé, J. Tow, A. M. Rush, S. Biderman, A. Webson, P. S. Ammanamanchi, T. Wang, B. Sagot, N. Muennighoff, A. V. del Moral, O. Ruwase, R. Bawden, S. Bekman, A. McMillan-Major, I. Beltagy, H. Nguyen, L. Saulnier, S. Tan, P. O. Suarez, V. Sanh, H. Laurençon, Y. Jernite, J. Launay, M. Mitchell, C. Raffel, A. Gokaslan, A. Simhi, A. Soroa, A. F. Aji, A. Alfassy, A. Rogers, A. K. Nitzav, C. Xu, C. Mou, C. Emezue, C. Klamm, C. Leong, D. van Strien, D. I. Adelani, D. Radev, E. G. Ponferrada, E. Levkovizh, E. Kim, E. B. Natan, F. D. Toni, G. Dupont, G. Kruszewski, G. Pistilli, H. Elsahar, H. Benyamina, H. Tran, I. Yu, I. Abdulmumin, I. Johnson, I. Gonzalez-Dios, J. de la Rosa, J. Chim, J. Dodge, J. Zhu, J. Chang, J. Frohberg, J. Tobing, J. Bhattacharjee, K. Almubarak, K. Chen, K. Lo, L. V. Werra, L. Weber, L. Phan, L. B. allal, L. Tanguy, M. Dey, M. R. Muñoz, M. Masoud, M. Grandury, M. Šaško, M. Huang, M. Coavoux, M. Singh, M. T. Jiang, M. C. Vu, M. A. Jauhar, M. Ghaleb, N. Subramani, N. Kassner, N. Khamis, O. Nguyen, O. Espejel, O. de Gibert, P. Villegas, P. Henderson, P. Colombo, P. Amuok, Q. Lhoest, R. Harliman, R. Bommasani, R. L. López, R. Ribeiro, S. Osei, S. Pyysalo, S. Nagel, S. Bose, S. H. Muhammad, S. Sharma, S. Longpre, S. Nikpoor, S. Silberberg, S. Pai, S. Zink, T. T. Torrent, T. Schick, T. Thrush, V. Danchev, V. Nikoulina, V. Laippala, V. Lepercq, V. Prabhu, Z. Alyafeai, Z. Talat, A. Raja, B. Heinzerling, C. Si, D. E. Taşar, E. Salesky, S. J. Mielke, W. Y. Lee, A. Sharma, A. Santilli, A. Chaffin, A. Stiegler, D. Datta, E. Szczechla, G. Chhablani, H. Wang, H. Pandey, H. Strobelt, J. A. Fries, J. Rozen, L. Gao, L. Sutawika, M. S. Bari, M. S. Al-shaibani, M. Manica, N. Nayak, R. Teehan, S. Albanie, S. Shen, S. Ben-David, S. H. Bach, T. Kim, T. Bers, T. Fevry, T. Neeraj, U. Thakker, V. Raunak, X. Tang, Z. Yong, Z. Sun, S. Brody, Y. Uri, H. Tojarieh, A. Roberts, H. W. Chung, J. Tae, J. Phang, O. Press, C. Li, D. Narayanan, H. Bourfoune, J. Casper, J. Rasley, M. Ryabinin, M. Mishra, M. Zhang, M. Shoeybi, M. Peyrounette, N. Patry, N. Tazi, O. Sanseviero, P. von Platen, P. Cornette, P. F. Lavallée, R. Lacroix, S. Rajbhandari, S. Gandhi, S. Smith, S. Requena, S. Patil, T. Dettmers, A. Baruwa, A. Singh, A. Cheveleva, A. Ligozat, A. Subramonian, A. Névéol, C. Lovering, D. Garrette, D. Tunuguntla, E. Reiter, E. Taktasheva, E. Voloshina, E. Bogdanov, G. I. Winata, H. Schoelkopf, J. Kalo, J. Novikova, J. Z. Forde, J. Clive, J. Kasai, K. Kawamura, L. Hazan, M. Carpuat, M. Clinciu, N. Kim, N. Cheng, O. Serikov, O. Antverg, O. van der Wal, R. Zhang, R. Zhang, S. Gehrmann, S. Mirkin, S. Pais, T. Shavrina, T. Scialom, T. Yun, T. Limisiewicz, V. Rieser, V. Protasov, V. Mikhailov, Y. Pruksachatkun, Y. Belinkov, Z. Bamberger, Z. Kasner, A. Rueda, A. Pestana, A. Feizpour, A. Khan, A. Faranak, A. Santos, A. Hevia, A. Unldreaj, A. Aghagol, A. Abdollahi, A. Tammour, A. HajiHosseini, B. Behroozi, B. Ajibade, B. Saxena, C. M. Ferrandis, D. McDuff, D. Contractor, D. Lansky, D. David, D. Kiela, D. A. Nguyen, E. Tan, E. Baylor, E. Ozoani, F. Mirza, F. Ononiwu, H. Rezanejad, H. Jones, I. Bhattacharya, I. Solaiman, I. Sedenko, I. Nejadgholi, J. Passmore, J. Seltzer, J. B. Sanz, L. Dutra, M. Samagaio, M. Elbadri, M. Mieskes, M. Gerchick, M. Akinlolu, M. McKenna, M. Qiu, M. Ghauri, M. Burynok, N. Abrar, N. Rajani, N. Elkott, N. Fahmy, O. Samuel, R. An, R. Kromann, R. Hao, S. Alizadeh, S. Shubber, S. Wang, S. Roy, S. Viguier, T. Le, T. Oyebade, T. Le, Y. Yang, Z. Nguyen, A. R. Kashyap, A. Palasciano, A. Callahan, A. Shukla, A. Miranda-Escalada, A. Singh, B. Beilharz, B. Wang, C. Brito, C. Zhou, C. Jain, C. Xu, C. Fourrier, D. L. Periñán, D. Molano, D. Yu, E. Manjavacas, F. Barth, F. Fuhrimann, G. Altay, G. Bayrak, G. Burns, H. U. Vrabec, I. Bello, I. Dash, J. Kang, J. Giorgi, J. Golde, J. D. Posada, K. R. Sivaraman, L. Bulchandani, L. Liu, L. Shinzato, M. H. de Bykhovetz, M. Takeuchi, M. Pàmies, M. A. Castillo, M. Nezhurina, M. Sänger, M. Samwald, M. Cullan, M. Weinberg, M. D. Wolf, M. Mihaljcic, M. Liu, M. Freidank, M. Kang, N. Seelam, N. Dahlberg, N. M. Broad, N. Muellner, P. Fung, P. Haller, R. Chandrasekhar, R. Eisenberg, R. Martin, R. Canalli, R. Su, R. Su, S. Cahyawijaya, S. Garda, S. S. Deshmukh, S. Mishra, S. Kiblawi, S. Ott, S. Sang-aroonsiri, S. Kumar, S. Schweter, S. Bharati, T. Laud, T. Gigant, T. Kainuma, W. Kusa, Y. Labrak, Y. S. Bajaj, Y. Venkatraman, Y. Xu, Y. Xu, Y. Xu, Z. Tan, Z. Xie, Z. Ye, M. Bras, Y. Belkada, and T. Wolf BLOOM: a 176b-parameter open-access multilingual language model. External Links: 2211.05100, Link Cited by: §A.7, §1, §3.3. Bisazza et al. (2021) A. Bisazza, A. Üstün, and S. Sportel On the difficulty of translating free-order case-marking languages. Transactions of the Association for Computational Linguistics 9, p. 1233–1248. External Links: Link Cited by: §2. Chang et al. (2026) T. A. Chang, C. Arnett, Z. Tu, and B. K. Bergen Goldfish: monolingual language models for 350 languages. In Proceedings of the 15th Language Resources and Evaluation Conference (LREC), External Links: Link Cited by: §A.1, §A.1, §A.2, §A.7, §3.3, §4.1, Limitations. Clark et al. (2023) T. H. Clark, C. Meister, T. Pimentel, M. Hahn, R. Cotterell, R. Futrell, and R. Levy A cross-linguistic pressure for Uniform Information Density in word order. Transactions of the Association for Computational Linguistics 11, p. 1048–1065. External Links: Link, Document Cited by: §2. Comrie (2008) B. Comrie Prenominal relative clauses in verb-object languages. Language and Linguistics 9 (4), p. 723–733. Cited by: footnote 8. Cotterell et al. (2018) R. Cotterell, S. J. Mielke, J. Eisner, and B. Roark Are all languages equally hard to language-model?. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), M. Walker, H. Ji, and A. Stent (Eds.), New Orleans, Louisiana, p. 536–541. External Links: Link, Document Cited by: §A.3, §3.4. Culbertson and Newport (2015) J. Culbertson and E. L. Newport Harmonic biases in child learners: in support of language universals. Cognition 139, p. 71–82. External Links: Document, ISSN 0010-0277, Link Cited by: §1, §5.1. M. S. Dryer and M. Haspelmath (Eds.) (2013) M. S. Dryer and M. Haspelmath (Eds.) WALS online (v2020.3). Max Planck Institute for Evolutionary Anthropology, Leipzig. External Links: Link Cited by: §A.2, §A.7, §1, §3.2, §5.1, §5.1, Limitations, Limitations. Dryer (1992) M. S. Dryer The Greenbergian word order correlations. Language 68 (1), p. 81–138. Cited by: §A.3, §1, §3.4, §5.1, §5.1, footnote 4, footnote 8. El-Naggar et al. (2025a) N. El-Naggar, T. Kuribayashi, and T. Briscoe GCG-based artificial languages for evaluating inductive biases of neural language models. In Proceedings of the 29th Conference on Computational Natural Language Learning, G. Boleda and M. Roth (Eds.), Vienna, Austria, p. 540–556. External Links: Link, Document, ISBN 979-8-89176-271-8 Cited by: §3, §5.1, §5.1. El-Naggar et al. (2025b) N. El-Naggar, T. Kuribayashi, and T. Briscoe Which word orders facilitate length generalization in LMs? an investigation with GCG-based artificial languages. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 35599–35613. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §A.1, §2, §3.1. Fitterer et al. (2025) S. Fitterer, D. Gangl, and J. Ulbrich Testing English news articles for lexical homogenization due to widespread use of large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), J. Zhao, M. Wang, and Z. Liu (Eds.), Vienna, Austria, p. 1239–1245. External Links: Link, Document, ISBN 979-8-89176-254-1 Cited by: §1. Futrell (2019) R. Futrell Information-theoretic locality properties of natural language. In Proceedings of the First Workshop on Quantitative Syntax (Quasy, SyntaxFest 2019), X. Chen and R. Ferrer-i-Cancho (Eds.), Paris, France, p. 2–15. External Links: Link, Document Cited by: §5.3. Greenberg (1963) J. H. Greenberg Some universals of grammar with particular reference to the order of meaningful elements. In Universals of Language, J. H. Greenberg (Ed.), p. 73–113. Cited by: §A.3, §1, §3.4, §5.1, footnote 2. Himmelreich (2023) A. Himmelreich Feature deletion by head movement – a new solution to agreement asymmetries in Modern Standard Arabic. Glossa: a journal of general linguistics 8 (1). External Links: Document, Link Cited by: Limitations. Hudson (1994) R. Hudson About 37% of word-tokens are nouns. Language 70 (2), p. 331–339. External Links: ISSN 00978507, 15350665, Link Cited by: §A.1. Hül and Dobrovoljc (2025) N. Hül and K. Dobrovoljc Word order variation in spoken and written corpora: a cross-linguistic study of SVO and alternative orders. In Proceedings of the Eighth International Conference on Dependency Linguistics (Depling, SyntaxFest 2025), E. Hajičová and S. Kahane (Eds.), Ljubljana, Slovenia, p. 150–155. External Links: Link, ISBN 979-8-89176-290-9 Cited by: §1, §3.2, Limitations. Joshi et al. (2020) P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, p. 6282–6293. External Links: Link, Document Cited by: §A.4, §A.5, Table 4, §5.2, Limitations. Kallini et al. (2024) J. Kallini, I. Papadimitriou, R. Futrell, K. Mahowald, and C. Potts Mission: impossible language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 14691–14714. External Links: Link, Document Cited by: §2, §3.4, Limitations. Keuleers and Brysbaert (2010) E. Keuleers and M. Brysbaert Wuggy: a multilingual pseudoword generator. Behavior Research Methods 42 (3), p. 627–633. External Links: Document, Link Cited by: §A.1, §A.7, §3.1. Kew et al. (2024) T. Kew, F. Schottmann, and R. Sennrich Turning English-centric LLMs into polyglots: how much multilinguality is needed?. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 13097–13124. External Links: Link, Document Cited by: §1. Khan et al. (2025) A. Khan, M. Shipton, D. Anugraha, K. Duan, P. H. Hoang, E. Khiu, A. S. Doğruöz, and E. A. Lee URIEL+: enhancing linguistic inclusion and usability in a typological and multilingual knowledge base. In Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Schockaert (Eds.), Abu Dhabi, UAE, p. 6937–6952. External Links: Link Cited by: §A.2, §3.2. Koppel and Ordan (2011) M. Koppel and N. Ordan Translationese and its dialects. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, D. Lin, Y. Matsumoto, and R. Mihalcea (Eds.), Portland, Oregon, USA, p. 1318–1326. External Links: Link Cited by: Limitations. Kreutzer et al. (2022) J. Kreutzer, I. Caswell, L. Wang, A. Wahab, D. van Esch, N. Ulzii-Orshikh, A. Tapo, N. Subramani, A. Sokolov, C. Sikasote, M. Setyawan, S. Sarin, S. Samb, B. Sagot, C. Rivera, A. Rios, I. Papadimitriou, S. Osei, P. O. Suarez, I. Orife, K. Ogueji, A. N. Rubungo, T. Q. Nguyen, M. Müller, A. Müller, S. H. Muhammad, N. Muhammad, A. Mnyakeni, J. Mirzakhalov, T. Matangira, C. Leong, N. Lawson, S. Kudugunta, Y. Jernite, M. Jenny, O. Firat, B. F. P. Dossou, S. Dlamini, N. de Silva, S. Çabuk Ballı, S. Biderman, A. Battisti, A. Baruwa, A. Bapna, P. Baljekar, I. A. Azime, A. Awokoya, D. Ataman, O. Ahia, O. Ahia, S. Agrawal, and M. Adeyemi Quality at a glance: an audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics 10, p. 50–72. External Links: Link, Document Cited by: Limitations. Kudo and Richardson (2018) T. Kudo and J. Richardson SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing. External Links: 1808.06226, Link Cited by: §A.1, §4.1. Kuribayashi et al. (2024) T. Kuribayashi, R. Ueda, R. Yoshida, Y. Oseki, T. Briscoe, and T. Baldwin Emergent word order universals from cognitively-motivated language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 14522–14543. External Links: Link, Document Cited by: §A.1, §A.1, §2, §5.1, Limitations. Laurençon et al. (2023) H. Laurençon, L. Saulnier, T. Wang, C. Akiki, A. Villanova del Moral, T. Le Scao, L. Von Werra, C. Mou, E. González Ponferrada, H. Nguyen, J. Frohberg, M. Šaško, Q. Lhoest, A. McMillan-Major, G. Dupont, S. Biderman, A. Rogers, L. Ben Allal, F. De Toni, G. Pistilli, O. Nguyen, S. Nikpoor, M. Masoud, P. Colombo, J. de la Rosa, P. Villegas, T. Thrush, S. Longpre, S. Nagel, L. Weber, M. Muñoz, J. Zhu, D. Van Strien, Z. Alyafeai, K. Almubarak, M. C. Vu, I. Gonzalez-Dios, A. Soroa, K. Lo, M. Dey, P. Ortiz Suarez, A. Gokaslan, S. Bose, D. Adelani, L. Phan, H. Tran, I. Yu, S. Pai, J. Chim, V. Lepercq, S. Ilic, M. Mitchell, S. A. Luccioni, and Y. Jernite The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset. External Links: 2303.03915, Link Cited by: §3.3, §5.2. Levshina et al. (2023) N. Levshina, S. Namboodiripad, M. Allassonnière-Tang, M. Kramer, L. Talamo, A. Verkerk, S. Wilmoth, G. G. Rodriguez, T. M. Gupton, E. Kidd, Z. Liu, C. Naccarato, R. Nordlinger, A. Panova, and N. Stoynova Why we need a gradient approach to word order. Linguistics 61 (4), p. 825–883. External Links: Link, Document Cited by: §3.2, Limitations. Lian et al. (2023) Y. Lian, A. Bisazza, and T. Verhoef Communication drives the emergence of language universals in neural agents: evidence from the word-order/case-marking trade-off. Transactions of the Association for Computational Linguistics 11, p. 1033–1047. External Links: Link, Document Cited by: §2. Lian et al. (2024) Y. Lian, T. Verhoef, and A. Bisazza NeLLCom-X: a comprehensive neural-agent framework to simulate language learning and group communication. In Proceedings of the 28th Conference on Computational Natural Language Learning, L. Barak and M. Alikhani (Eds.), Miami, FL, USA, p. 243–258. External Links: Link, Document Cited by: §2. Lin et al. (2022) X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, R. Pasunuru, S. Shleifer, P. S. Koura, V. Chaudhary, B. O’Horo, J. Wang, L. Zettlemoyer, Z. Kozareva, M. Diab, V. Stoyanov, and X. Li Few-shot learning with multilingual language models. External Links: 2112.10668, Link Cited by: §A.7, §3.3. Littell et al. (2017) P. Littell, D. R. Mortensen, K. Lin, K. Kairis, C. Turner, and L. Levin URIEL and lang2vec: representing languages as typological, geographical, and phylogenetic vectors. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, M. Lapata, P. Blunsom, and A. Koller (Eds.), Valencia, Spain, p. 8–14. External Links: Link Cited by: §A.2, §3.2. S. M. Michaelis, P. Maurer, M. Haspelmath, and M. Huber (Eds.) (2013) S. M. Michaelis, P. Maurer, M. Haspelmath, and M. Huber (Eds.) The atlas of pidgin and creole language structures. Oxford University Press, Oxford. External Links: Link Cited by: §A.2, §3.2. Mielke et al. (2019) S. J. Mielke, R. Cotterell, K. Gorman, B. Roark, and J. Eisner What kind of language is hard to language-model?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L. Màrquez (Eds.), Florence, Italy, p. 4975–4989. External Links: Link, Document Cited by: §A.5. Muennighoff et al. (2023) N. Muennighoff, A. M. Rush, B. Barak, T. L. Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. Raffel Scaling data-constrained language models. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §A.1. Nivre et al. (2020) J. Nivre, M. de Marneffe, F. Ginter, J. Hajič, C. D. Manning, S. Pyysalo, S. Schuster, F. Tyers, and D. Zeman Universal Dependencies v2: an evergrowing multilingual treebank collection. In Proceedings of the Twelfth Language Resources and Evaluation Conference, N. Calzolari, F. Béchet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis (Eds.), Marseille, France, p. 4034–4043 (eng). External Links: Link, ISBN 979-10-95546-34-4 Cited by: §A.2, Limitations. NLLB Team et al. (2022) NLLB Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, S. Spruit, C. Tran, P. Andrews, N. F. Ayan, S. Bhosale, S. Edunov, A. Fan, C. Gao, V. Goswami, F. Guzmán, P. Koehn, A. Mourachko, C. Ropers, S. Saleem, H. Schwenk, and J. Wang No language left behind: scaling human-centered machine translation. External Links: 2207.04672, Link Cited by: §A.7, §3.2, §4.2. Oladipo et al. (2023) A. Oladipo, M. Adeyemi, O. Ahia, A. T. Owodunni, O. Ogundepo, D. I. Adelani, and J. Lin Better quality pre-training data and t5 models for African languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 158–168. External Links: Link, Document Cited by: §5.2, §5.3, Limitations. Olah (2022) C. Olah Mechanistic interpretability, variables, and the importance of interpretable bases. Note: Transformer Circuits Thread External Links: Link Cited by: §6. Qi et al. (2020) P. Qi, Y. Zhang, Y. Zhang, J. Bolton, and C. D. Manning Stanza: a python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, A. Celikyilmaz and T. Wen (Eds.), Online, p. 101–108. External Links: Link, Document Cited by: §A.7, §3.2, §4.2. Ravfogel et al. (2019) S. Ravfogel, Y. Goldberg, and T. Linzen Studying the inductive biases of RNNs with synthetic variations of natural languages. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, p. 3532–3542. External Links: Link, Document Cited by: §1, §2. Shliazhko et al. (2024) O. Shliazhko, A. Fenogenova, M. Tikhonova, A. Kozlova, V. Mikhailov, and T. Shavrina MGPT: few-shot learners go multilingual. Transactions of the Association for Computational Linguistics 12, p. 58–79. External Links: Link, Document Cited by: §A.7, §3.3. Skirgård et al. (2023a) H. Skirgård, H. J. Haynie, D. E. Blasi, H. Hammarström, J. Collins, J. J. Latarche, J. Lesage, T. Weber, A. Witzlack-Makarevich, S. Passmore, A. Chira, L. Maurits, R. Dinnage, M. Dunn, G. Reesink, R. Singer, C. Bowern, P. Epps, J. Hill, O. Vesakoski, M. Robbeets, N. K. Abbas, D. Auer, N. A. Bakker, G. Barbos, R. D. Borges, S. Danielsen, L. Dorenbusch, E. Dorn, J. Elliott, G. Falcone, J. Fischer, Y. Ghanggo Ate, H. Gibson, H. Göbel, J. A. Goodall, V. Gruner, A. Harvey, R. Hayes, L. Heer, R. E. Herrera Miranda, N. Hübler, B. Huntington-Rainey, J. K. Ivani, M. Johns, E. Just, E. Kashima, C. Kipf, J. V. Klingenberg, N. König, A. Koti, R. G. A. Kowalik, O. Krasnoukhova, N. L.M. Lindvall, M. Lorenzen, H. Lutzenberger, T. R.A. Martins, C. Mata German, S. van der Meer, J. Montoya Samamé, M. Müller, S. Muradoglu, K. Neely, J. Nickel, M. Norvik, C. A. Oluoch, J. Peacock, I. O.C. Pearey, N. Peck, S. Petit, S. Pieper, M. Poblete, D. Prestipino, L. Raabe, A. Raja, J. Reimringer, S. C. Rey, J. Rizaew, E. Ruppert, K. K. Salmon, J. Sammet, R. Schembri, L. Schlabbach, F. W.P. Schmidt, A. Skilton, W. D. Smith, H. de Sousa, K. Sverredal, D. Valle, J. Vera, J. Voß, T. Witte, H. Wu, S. Yam, J. Ye 葉婧婷, M. Yong, T. Yuditha, R. Zariquiey, R. Forkel, N. Evans, S. C. Levinson, M. Haspelmath, S. J. Greenhill, Q. D. Atkinson, and R. D. Gray Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss. Science Advances 9 (16). External Links: Document Cited by: §A.2, §A.7, §3.2, Limitations. Skirgård et al. (2023b) H. Skirgård, H. J. Haynie, D. E. Blasi, H. Hammarström, and et al. Grambank. Zenodo. Note: Current release version of the Grambank data External Links: Document Cited by: footnote 4. Speer (2022) R. Speer Rspeer/wordfreq: v3.0. Zenodo. External Links: Document, Link Cited by: §A.1, §A.7, §3.1. Tatariya et al. (2025) K. Tatariya, W. Poelman, and M. de Lhoneux On the interplay between positional encodings, morphological complexity, and word order flexibility. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, K. Inui, S. Sakti, H. Wang, D. F. Wong, P. Bhattacharyya, B. Banerjee, A. Ekbal, T. Chakraborty, and D. P. Singh (Eds.), Mumbai, India, p. 1761–1778. External Links: Link, Document, ISBN 979-8-89176-298-5 Cited by: §A.5, §5.2. Verkerk et al. (2026) A. Verkerk, O. Shcherbakova, H. J. Haynie, H. Skirgård, C. Rzymski, Q. D. Atkinson, S. J. Greenhill, and R. D. Gray Enduring constraints on grammar revealed by Bayesian spatiophylogenetic analyses. Nature Human Behaviour 10 (1), p. 126–136. External Links: Document Cited by: §1. Wang et al. (2025) J. Wang, Y. Lu, M. Weber, M. Ryabinin, D. I. Adelani, Y. Chen, R. Tang, and P. Stenetorp Multilingual language model pretraining using machine-translated data. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 28087–28107. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §5.2, §5.3, Limitations. White and Cotterell (2021) J. C. White and R. Cotterell Examining the inductive bias of neural language models with artificial languages. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), C. Zong, F. Xia, W. Li, and R. Navigli (Eds.), Online, p. 454–463. External Links: Link, Document Cited by: §A.1, §A.1, §A.1, §1, §2, §3.1, §3.1, §3, §4.1, §4.1, §5.1, §5.1, Limitations. Wilcox et al. (2023) E. G. Wilcox, T. Pimentel, C. Meister, R. Cotterell, and R. P. Levy Testing the predictions of surprisal theory in 11 languages. Transactions of the Association for Computational Linguistics 11, p. 1451–1470. External Links: Link, Document Cited by: §3.4. Xu et al. (2026) T. Xu, T. Kuribayashi, Y. Oseki, R. Cotterell, and A. Warstadt Can language models learn typologically implausible languages?. Transactions of the Association for Computational Linguistics 14, p. 588–611. External Links: Document Cited by: §2. Zeman et al. (2017) D. Zeman, M. Popel, M. Straka, J. Hajič, J. Nivre, F. Ginter, J. Luotolahti, S. Pyysalo, S. Petrov, M. Potthast, F. Tyers, E. Badmaeva, M. Gokirmak, A. Nedoluzhko, S. Cinková, J. Hajič jr., J. Hlaváčová, V. Kettnerová, Z. Urešová, J. Kanerva, S. Ojala, A. Missilä, C. D. Manning, S. Schuster, S. Reddy, D. Taji, N. Habash, H. Leung, M. de Marneffe, M. Sanguinetti, M. Simi, H. Kanayama, V. de Paiva, K. Droganova, H. Martínez Alonso, Ç. Çöltekin, U. Sulubacak, H. Uszkoreit, V. Macketanz, A. Burchardt, K. Harris, K. Marheinecke, G. Rehm, T. Kayadelen, M. Attia, A. Elkahky, Z. Yu, E. Pitler, S. Lertpradit, M. Mandl, J. Kirchner, H. F. Alcalde, J. Strnadová, E. Banerjee, R. Manurung, A. Stella, A. Shimada, S. Kwak, G. Mendonça, T. Lando, R. Nitisaroj, and J. Li CoNLL 2017 shared task: multilingual parsing from raw text to Universal Dependencies. In Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, J. Hajič and D. Zeman (Eds.), Vancouver, Canada, p. 1–19. External Links: Link, Document Cited by: §A.7, §3.2, §4.2. Ziv et al. (2026) I. Ziv, N. Lan, and E. Chemla Biasless language models learn unnaturally: how LLMs fail to distinguish the possible from the impossible. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 5393–5403. External Links: Link, Document, ISBN 979-8-89176-380-7 Cited by: §2, §3.4, §3. Appendix A Appendix A.1 Artificial Language Details Vocabulary Source words were drawn from the most frequent English words in the wordfreq library (51) and used to generate pseudowords with Wuggy (26). The resulting vocabulary contains ∼ 50k wordforms across five open-class categories: nouns inflecting for number (50%), adjectives (10%), and three verb classes (intransitive, transitive, complement-taking) inflecting for number and tense (40% combined). Closed-class items comprise 4 prepositions, 7 pronouns, and single invariant tokens for the subordinating complementiser, relativiser, coordinating conjunction, and overt subject and object case markers. This POS inventory follows 55 and subsequent work (32; 17); fully replicating natural language POS complexity is infeasible within a controlled PCFG setup. The noun-heavy distribution (50%) reflects this restricted inventory and aligns with corpus-based POS statistics (22). Within each open-class category, items are assigned Zipfian sampling weights (α=1.0α=1.0) by frequency rank; closed-class items are sampled uniformly. Training Corpus Each language variant was used to generate 5 MB and 10 MB of training text by sampling sentences from the PCFG, compared to 10,000 sentences in 55. The PCFG allows recursive expansion, with a maximum of 400 expansions per sentence to bound recursion depth, following 55. Generated sentences have an average length of 12 tokens. Tokeniser We train a shared SentencePiece BPE tokeniser (31) (15k vocabulary) on the generated corpus in word-boundary mode, so merges never cross word boundaries. Because all 192 languages are permutations of the same sentences with identical vocabulary and word frequencies, the subword inventory depends only on the shared word list, not on word order: retraining the tokeniser separately on each grammar yields identical 15,000-piece vocabularies. We also enable add_dummy_prefix so that sentence-initial words receive the same token IDs as mid-sentence occurrences, avoiding position-based bias. To confirm our results are not tokeniser-driven, we additionally retrained all 192 models (single seed) with a unigram tokeniser (as used in the Goldfish models; 9) and obtained comparable results: the difficulty ordering is nearly unchanged (Spearman ρ=0.94ρ=0.94 at 5 MB, 0.910.91 at 10 MB), all five switch effects keep their direction, and the base-order extremes are preserved. Model and Hyperparameters We fully reproduced the results of 55 before conducting our experiments. We then made the following modifications. We adopt the Goldfish architecture (9) to match the model size used in our natural language experiments, enabling direct comparison (full hyperparameters in Table 3). Our models match the Goldfish configuration in layers, attention heads, and embedding dimensions, differing only in vocabulary size (15k vs 50k); prior work with a much smaller vocabulary (∼ 1,400 pseudowords; 55; 32) recovers the same trends, so our results are not vocabulary-size-driven. We set weight decay to 0, as the default value of 0.01 led to over-regularisation at our model size and data volume, roughly doubling perplexity in pilot experiments. We reduce warmup steps to 10% of total training steps. We cap training at 20 epochs rather than the 10,000 update steps used by 55, which on their data volume results in over 100 epochs. In pilot experiments on 5 MB and 10 MB corpora, we observed train–validation loss crossover at epochs ∼ 10 and ∼ 16, respectively, with validation loss plateauing or increasing beyond epoch 50. Limiting training to 20 epochs is consistent with 41, who find diminishing gains beyond 16 epochs of data repetition; with 9, who train Goldfish models for 10 epochs; and with 32, who use 10 epochs in a similar setup with a smaller training corpus. We evaluate at step level and select the checkpoint with the lowest validation loss. We limit artificial language experiments to 5 and 10 MB to match the Goldfish architecture at these sizes; at larger scales, Goldfish uses more layers, precluding direct comparison. Artificial languages are also easier to model than natural. Hyperparameter Value Model Architecture GPT-2 (decoder-only) Layers 4 Attention heads 8 Hidden size 512 FFN inner size 2,048 Dropout 0.1 Activation ReLU Tied embeddings Yes Parameters 20.6M Tokeniser Type SentencePiece BPE (word-boundary) Vocabulary size 15,000 Data Train / Dev / Test 80 / 10 / 10 Optimisation Framework HuggingFace Transformers Optimiser Adam (β1=0.9 _1=0.9, β2=0.999 _2=0.999, ε= =1e-6) Learning rate 1e-4 LR schedule Linear decay with warmup Warmup 10% of total steps Weight decay 0 Gradient clipping 1.0 Sequence length 512 tokens Batch size 64 sequences (4 × 16 grad. accum.) Max epochs 20 Precision bf16 Table 3: Model and training hyperparameters for artificial language experiments. Compute Cost For artificial languages, we trained 1,920 models (192 grammars × 10 splits) at each of two data sizes (5 MB and 10 MB) on NVIDIA A100 GPUs. A single model takes approximately 2.5 minutes (5 MB) and 5 minutes (10 MB), totalling approximately 252 A100 GPU-hours with a wall-clock time of approximately 2 days using parallelised Slurm array jobs. A.2 Natural Language Details Word Order Labeling We map each switch in our artificial grammar (Figure 2) to attested typological features: Base word order: WALS 81A; Switch 1 (complement clause order): Grambank GB135; Switch 2 (complementiser position): Grambank GB421; Switch 3 (adposition order): WALS 85A; Switch 4 (adjective–noun order): WALS 87A, Grambank GB193; Switch 5 (relative clause–noun order): WALS 90A. Features were drawn from lang2vec (38), which aggregates WALS, SSWL, and Ethnologue. Where lang2vec lacked coverage, we consulted WALS (14) directly, then Grambank (49) and APiCS (39) via URIEL+ (28). FLORES-200 covers 200 languages, with script alternatives for four (204 total). Of these, 52 lacked a base word order label in WALS, Grambank, and APiCS; we supplemented these using peer-reviewed articles for specific languages, corroborated where possible by corpus-based labels from UD treebanks (42), bringing coverage to 194 languages. Of the 194, 132 have binary (L/R) values for all five switches. Languages with non-binary or missing values were excluded from the corresponding analysis. Evaluation Languages Tables 4–5 list all 103 FLORES-200 languages with a base word order label that appear in at least one evaluation set and form the basis of all analyses in §5: 68 Goldfish languages available at all four training sizes,1414 14 Larger sets are available at individual sizes (e.g., 130 at 5 MB) but are not used in §5 to ensure that the same languages are compared across all four training scales, enabling attribution to data volume rather than language sample changes. 46 BLOOM, 30 XGLM, and 56 mGPT training languages; 20 of the 103 also have a PUD treebank. Goldfish Model Coverage on FLORES-200 We evaluated Goldfish models on 183 of FLORES-200 languages, requiring a strict one-to-one match between FLORES codes and dedicated Goldfish models; we do not adopt the cross-variety substitutions in 9 (e.g. awa_Deva → hin_Deva), as word order differences between varieties could confound our analysis. The 21 excluded varieties (ace_Arab, acm_Arab, acq_Arab, aeb_Arab, ajp_Arab, arb_Latn, ars_Arab, ary_Arab, awa_Deva, bjn_Arab, kam_Latn, kas_Arab, min_Arab, mni_Beng, npi_Deva, nus_Latn, ory_Orya, sat_Beng, taq_Latn, taq_Tfng, tzm_Tfng) lack a dedicated model due to missing script coverage or absence from Goldfish entirely. Model Language FLORES Label Family Res. Goldfish BLOOM XGLM mGPT PUD Afrikaans afr_Latn NoDom-RRRLR Indo-European 3 ✓ × × ✓ × Akan aka_Latn SVO-RNLR? Niger-Congo 1 × ✓∗ × × × Amharic amh_Ethi SOV-LLRLL Afro-Asiatic 2 ✓ × × × × Armenian hye_Armn NoDom-RRLLN Indo-European 1 ✓ × × ✓ × Assamese asm_Beng SOV-L??L? Indo-European 1 × ✓∗ × × × Ayacucho Quechua quy_Latn SOV-L?L Quechuan – × × ✓∗ × × Bambara bam_Latn SOV-LLLRL Niger-Congo 1 × ✓ × × × Bashkir bak_Cyrl SOV-LRLLL Turkic 1 × × × ✓ × Basque eus_Latn SOV-LLLRL Isolate 4 ✓ ✓ ✓ ✓ × Belarusian bel_Cyrl SVO-RRRLR Indo-European 3 ✓ × × ✓ × Bengali ben_Beng SOV-LRLLL Indo-European 3 ✓ ✓ ✓ ✓ × Bosnian bos_Latn SVO-RRRLR Indo-European 3 ✓ × × × × Bulgarian bul_Cyrl SVO-RRRLR Indo-European 3 ✓ × ✓ ✓ × Burmese mya_Mymr SOV-LNLRL Sino-Tibetan 1 × × ✓ ✓ × Catalan cat_Latn SVO-R Indo-European 4 ✓ ✓ ✓ × × Chinese (Simp.) zho_Hans SVO-RRRLL Sino-Tibetan 5 ✓ ✓ ✓ × ✓ Chinese (Trad.) zho_Hant SVO-RRRLL Sino-Tibetan 5 × ✓ × × × Croatian hrv_Latn SVO-RRRLR Indo-European 4 ✓ × × × × Czech ces_Latn SVO-RRRLR Indo-European 4 ✓ × × × ✓ Danish dan_Latn SVO-RRRLR Indo-European 3 ✓ × × ✓ × Dutch nld_Latn NoDom-RRRLR Indo-European 4 ✓ × × × × Eastern Panjabi pan_Guru SOV-RRLLN Indo-European 2 ✓ ✓ × × × English eng_Latn SVO-RRRLR Indo-European 5 ✓ ✓ ✓ ✓ ✓ Esperanto epo_Latn NoDom-RRRLR Constructed 1 ✓ × × × × Estonian est_Latn SVO-RRLLR Uralic 3 ✓ × ✓ ✓ × Finnish fin_Latn SVO-RRLLR Uralic 4 ✓ × ✓ ✓ ✓ Fon fon_Latn SVO-RRNRR Niger-Congo 0 × ✓ × × × French fra_Latn SVO-R Indo-European 5 ✓ ✓ ✓ ✓ ✓ Galician glg_Latn NoDom-R Indo-European 3 ✓ × × × × Ganda lug_Latn SVO-R Niger-Congo 1 × ✓ × × × Georgian kat_Geor NoDom-RRLLR Kartvelian 3 ✓ × × ✓ × German deu_Latn NoDom-RRRLR Indo-European 5 ✓ × ✓ ✓ ✓ Greek ell_Grek NoDom-RRRLR Indo-European 3 ✓ × ✓ ✓ × Gujarati guj_Gujr SOV-LRLLL Indo-European 1 ✓ ✓ × × × Haitian Creole hat_Latn SVO-R Creole 0 × × ✓ × × Halh Mongolian khk_Cyrl SOV-L Mongolic – × × × ✓ × Hausa hau_Latn SVO-RRRLR Afro-Asiatic 2 ✓ × × × × Hebrew heb_Hebr SVO-R Afro-Asiatic 3 ✓ × × ✓ × Hindi hin_Deva SOV-LRLLL Indo-European 4 ✓ ✓ ✓ ✓ ✓ Hungarian hun_Latn NoDom-RRLLN Uralic 4 ✓ × × ✓ × Icelandic isl_Latn SVO-RRRLR Indo-European 2 ✓ × × × ✓ Igbo ibo_Latn SVO-R Niger-Congo 1 × ✓ × × × Indonesian ind_Latn SVO-R Austronesian 3 ✓ ✓ ✓ ✓ ✓ Italian ita_Latn NoDom-R Indo-European 4 ✓ × ✓ ✓ ✓ Japanese jpn_Jpan SOV-L Japonic 5 ✓ × ✓ ✓ ✓ Javanese jav_Latn SVO-R Austronesian 1 × × × ✓ × Kannada kan_Knda SOV-L Dravidian 1 ✓ ✓ × × × Kazakh kaz_Cyrl SOV-L???? Turkic 3 ✓∗ × × ✓∗ × Kikuyu kik_Latn SVO-R Niger-Congo 1 × ✓ × × × Kinyarwanda kin_Latn SVO-R?R? Niger-Congo 1 × ✓∗ × × × Korean kor_Hang SOV-L Koreanic 4 ✓ × ✓ ✓ ✓ Kyrgyz kir_Cyrl SOV-L?L? Turkic 1 ✓∗ × × ✓∗ × Table 4: Evaluation languages (part 1 of 2): the 103 FLORES-200 languages with a base word order label appearing in at least one evaluation set. Label: dominant order + five head-direction parameters following the encoding in Figure 2 (NoDom = No Dominant Order). Res.: resource class 0–5 (24); – = not classified. Model columns: presence in training data (Goldfish = available at all four training sizes, 68/103; BLOOM, 46; XGLM, 30; mGPT, 56). PUD: PUD treebank available (20/103; all PUD languages fall within this set). ✓∗: included for base word order analysis but excluded from directionality analysis (§5) due to missing parameters (?). Model Language FLORES Label Family Res. Goldfish BLOOM XGLM mGPT PUD Lingala lin_Latn SVO-R?R Niger-Congo 1 × ✓∗ × × × Lithuanian lit_Latn NoDom-RRRLR Indo-European 3 ✓ × × ✓ × Macedonian mkd_Cyrl SVO-RRRLR Indo-European 1 ✓ × × × × Malayalam mal_Mlym SOV-L Dravidian 1 ✓ ✓ × ✓ × Maltese mlt_Latn SVO-R Afro-Asiatic 2 ✓ × × × × Marathi mar_Deva SOV-RNLLL Indo-European 2 ✓ ✓ × ✓ × MSA arb_Arab VSO-R Afro-Asiatic – ✓ ✓ ✓ ✓ ✓ Nepali npi_Deva SOV-L? Indo-European 1 × ✓∗ × × × Northern Sotho nso_Latn SVO-R???? Niger-Congo 1 × ✓∗ × × × Northern Uzbek uzn_Latn SOV-LRLLL Turkic – × × × ✓ × Norwegian Bokmål nob_Latn SVO-RRRLR Indo-European 1 ✓ × × × × Nyanja nya_Latn SVO-R Niger-Congo 1 × ✓ × × × Odia ory_Orya SOV-?NLLN Indo-European 1 × ✓∗ × × × Polish pol_Latn SVO-RRRLR Indo-European 4 ✓ × × ✓ ✓ Portuguese por_Latn SVO-R Indo-European 4 ✓ ✓ ✓ ✓ ✓ Romanian ron_Latn SVO-R Indo-European 3 ✓ × × ✓ × Rundi run_Latn SVO-R Niger-Congo 0 × ✓ × × × Russian rus_Cyrl SVO-RRRLR Indo-European 4 ✓ × ✓ ✓ ✓ Serbian srp_Cyrl SVO-RRRLR Indo-European 4 ✓ × × × × Shona sna_Latn SVO-R Niger-Congo 1 × ✓ × × × Sinhala sin_Sinh SOV-L Indo-European 0 ✓ × × × × Slovak slk_Latn SVO-RRRLR Indo-European 3 ✓ × × × × Slovenian slv_Latn SVO-RRRLR Indo-European 3 ✓ × × × × Somali som_Latn SOV-LRNRR Afro-Asiatic 1 ✓ × × × × Southern Sotho sot_Latn SVO-R Niger-Congo 1 × ✓ × × × Spanish spa_Latn SVO-R Indo-European 5 ✓ ✓ ✓ ✓ ✓ Std. Latvian lvs_Latn SVO-RRRLR Indo-European – × × × ✓ × Std. Malay zsm_Latn SVO-R Austronesian – × × × ✓ × Swahili swh_Latn SVO-R Niger-Congo – ✓ ✓ ✓ ✓ × Swedish swe_Latn SVO-RRRLR Indo-European 4 ✓ × × ✓ ✓ Tagalog tgl_Latn VSO-RRRNR Austronesian 3 ✓ × × × × Tajik tgk_Cyrl SOV-LRRRR Indo-European 1 ✓ × × × × Tamil tam_Taml SOV-L Dravidian 3 ✓ ✓ ✓ × × Tatar tat_Cyrl SOV-LRLLL Turkic 1 ✓ × × ✓ × Telugu tel_Telu SOV-L Dravidian 1 ✓ ✓ ✓ ✓ × Thai tha_Thai SVO-R Tai-Kadai 3 ✓ × ✓ ✓ ✓ Tsonga tso_Latn SVO-R?R? Niger-Congo 1 × ✓∗ × × × Tswana tsn_Latn SVO-R Niger-Congo 2 × ✓ × × × Tumbuka tum_Latn SVO-R???? Niger-Congo 1 × ✓∗ × × × Turkish tur_Latn SOV-L Turkic 4 ✓ × ✓ ✓ ✓ Turkmen tuk_Latn SOV-L Turkic 1 × × × ✓ × Twi twi_Latn SVO-RNLR? Niger-Congo 1 × ✓∗ × × × Ukrainian ukr_Cyrl SVO-RRRLR Indo-European 3 ✓ × × ✓ × Urdu urd_Arab SOV-LRLLL Indo-European 3 ✓ ✓ ✓ × × Vietnamese vie_Latn SVO-R Austroasiatic 4 ✓ ✓ ✓ ✓ × Welsh cym_Latn VSO-R Indo-European 1 ✓ × × × × Western Persian pes_Arab SOV-?R Indo-European – ✓∗ × × ✓∗ × Wolof wol_Latn SVO-R Niger-Congo 2 × ✓ × × × Xhosa xho_Latn SVO-R Niger-Congo 2 × ✓ × × × Yoruba yor_Latn SVO-R Niger-Congo 2 × ✓ × ✓ × Zulu zul_Latn SVO-R Niger-Congo 2 × ✓ × × × Table 5: Evaluation languages (part 2 of 2, continued from Table 4). Column definitions are identical. MSA = Modern Standard Arabic. A.3 Evaluation BPEC For artificial languages, raw perplexity is directly comparable because all 192 variants share the same vocabulary, tokeniser, and derivation probabilities. For natural languages, per-token PPL is not comparable across scripts and tokenisers. We therefore follow 12 and normalise corpus negative log-likelihood by the English character count of the parallel corpus to obtain bits per English character (BPEC): BPECℓ=−∑s∈Sℓ∑t=1nslog2pθ(xt(s)∣x<t(s))CengBPEC_ = - _s∈ S_ \; _t=1^n_s _2p_θ\! (x_t^(s) x_<t^(s) )C_eng (1) where SℓS_ are the evaluation sentences in language ℓ , nsn_s the scored subword tokens per sentence, and CengC_eng the total character count of the English portion of the respective corpus (FLORES-200/PUD). We prepend the model’s start token ([CLS] for Goldfish, [BOS] for multilingual models) as conditioning context, but exclude it from the scored tokens. Greenbergian Universals From Greenberg’s 45 universals (20), we select the subset testable with syntactic constructions in our artificial languages. Since our PCFG generates only declarative sentences, we are limited to universals involving base S/V/O order, adpositions, adjective–noun order, and relative clauses. This yields four implicational universals: Universal 1 (subject precedes object in the dominant orders SOV, SVO, VSO), Universal 3 (VSO languages are prepositional), Universal 4 (SOV languages are postpositional), and Universal 17 (VSO languages place the adjective after the noun). We also test Dryer’s observation that OV languages correlate with prenominal relative clauses and VO languages with postnominal ones (15). For each universal, we compare mean test PPL between the configuration predicted by the universal and its counterpart. A.4 Language Resource Level and Word Order Distribution We use the resource taxonomy of 24, which classifies languages into six classes (0–5) based on the availability of labeled and unlabeled NLP data, as a proxy for resourcedness (column Res. in Tables 4–5). We report statistics on the 68 Goldfish languages listed in Tables 4–5, tracked across all four training scales, as these are the languages for which we directly observe the emerging SVO advantage (65 with a resource class assignment in 24). SVO languages have substantially higher resourcedness than SOV (mean resource class 3.35 vs. 2.24; Mann–Whitney U=472U=472, p=0.003p=0.003, r=0.72r=0.72). Of the 19 highly-resourced (Class 4–5) SVO, SOV languages, 14 are SVO (74%), and a Cochran–Armitage trend test confirms that the proportion of SVO rises monotonically with resource class (z=3.03z=3.03, p=0.002p=0.002). A.5 Mixed-Effects Analysis of the SVO–SOV Gap In line with prior work that uses mixed-effects models to separate language-modeling difficulty from confounding language properties (40), we fit linear mixed-effects models to the Goldfish BPEC of the 50 SVO/SOV languages (32 SVO, 18 SOV) in our full directionality-annotated set (Figure 1c,d), across all four training sizes (200 language-by-size observations; Table 6); VSO and NoDominant languages also shown in Figure 1c,d are excluded, as the model contrasts SVO against SOV. The model is ∼ bpec _×+covariates is\_SOV× size+covariates +(1∣), +(1 language), with training size entered as z-scored log10 _10(MB), so main effects are read at the centre of the size range (∼ 47 MB) and interactions per SD of log-size (∼ 8× data). The covariates (all reported in Table 6) are morphological complexity (mattr_z), language resourcedness (resource_lvl, z-scored resourcedness class from 24), and training data composition (data_comp, OSCAR web-crawl share).1515 15 The share of each language’s Goldfish training data drawn from OSCAR rather than more curated corpora, computed from the per-language proportions column of the Goldfish data documentation: https://github.com/tylerachang/goldfish/blob/main/data/goldfish_data_info.tsv. We measure morphological complexity as subword MATTR (1000-token window) under each language’s own tokeniser (52, as in), since corpus-based measures such as TTR have been shown to correlate with, and be interchangeable with, typological complexity measures derived from databases such as WALS (6), and it is available for all languages, so it does not limit the sample. Models are fit by maximum likelihood so the nested models M1–M9 are comparable; the robustness block in Table 6 refits key models with a maximal random-effects structure (random intercept + random size slope per language). Our finding is the is_SOV × size interaction (the growth of the SVO–SOV performance gap with training data scale), and each covariate is tested as a predictor of it, both as a level term and as an interaction. Table 6 reports all coefficients across the nine nested models (M1–M9) and the robustness checks. Term β SE p M1 base (no covariates) is_SOV +0.071 0.024 .003 is_SOV×size +0.049 0.011 <<.001 M2 + morphological complexity (level) is_SOV +0.050 0.024 .041 is_SOV×size +0.049 0.011 <<.001 mattr_z +0.027 0.012 .021 M3 + language resourcedness (level) is_SOV +0.053 0.025 .033 is_SOV×size +0.049 0.011 <<.001 resource_lvl −0.023-0.023 0.012 .057 M4 + training data composition (level) is_SOV +0.052 0.025 .034 is_SOV×size +0.049 0.011 <<.001 data_comp −0.021-0.021 0.012 .076 M5 + all three levels is_SOV +0.030 0.025 .224 is_SOV×size +0.049 0.011 <<.001 mattr_z +0.025 0.011 .022 resource_lvl −0.011-0.011 0.014 .441 data_comp −0.014-0.014 0.014 .317 M6 + morphological complexity × word order is_SOV +0.055 0.025 .029 is_SOV×size +0.049 0.011 <<.001 mattr_z +0.034 0.014 .018 mattr_z×is_SOV −0.022-0.022 0.025 .386 M7 + language resourcedness × size is_SOV +0.036 0.025 .147 is_SOV×size +0.039 0.012 <<.001 resource_lvl×size −0.012-0.012 0.006 .037 mattr_z +0.024 0.012 .037 M8 + training data composition × size is_SOV +0.032 0.025 .202 is_SOV×size +0.040 0.012 <<.001 data_comp −0.021-0.021 0.011 .064 data_comp×size −0.009-0.009 0.006 .098 mattr_z +0.027 0.011 .015 M9 full model (resourcedness & composition × size) is_SOV +0.030 0.025 .224 is_SOV×size +0.038 0.012 .001 mattr_z +0.025 0.011 .022 resource_lvl×size −0.010-0.010 0.007 .165 data_comp×size −0.003-0.003 0.007 .634 Robustness: maximal random effects M1: is_SOV×size +0.049 0.011 <<.001 M7: is_SOV×size +0.039 0.012 .001 M9: is_SOV×size +0.038 0.012 .002 binary res. is_SOV×size +0.037 0.011 .001 binary res. high_res×size −0.032-0.032 0.012 .007 Table 6: Mixed-effects models of Goldfish BPEC (50 SVO/SOV languages × 4 sizes). Bold = p<.05p<.05. Each covariate (morphological complexity, resourcedness, and training-data composition, OSCAR share) is entered as a level term (M2–M4), jointly (M5), and in interaction (M6–M9). Our finding, the word order × size interaction (is_SOV×size, the SOV gap growing with data), survives every control (M2–M9) and a maximal random-effects structure. The main effect (is_SOV, the SVO–SOV gap at the average training size, ∼ 47 MB) is absorbed once all three covariate levels are added together (M5). Our finding, the is_SOV×size interaction, survives every control: it stays significant across M2–M9 and under the maximal random-effects structure, shrinking only ∼ 20% (from +0.049+0.049 to +0.038+0.038) when resourcedness is allowed to scale with data (M7, M9). The main effect (is_SOV, the average SVO–SOV gap), by contrast, is absorbed: it is significant on its own (M1) and with any single covariate (M2–M4) but non-significant once all three are added together (M5), so the covariates jointly explain the average gap, but not its growth with data. Among the covariates, only resourcedness has a significant standalone size interaction (resource_lvl×size, β=−0.012β=-0.012, p=.037p=.037; M7), in the same direction as our effect: higher-resource languages improve faster with data, widening the gap for the lower-resourced SOV group, which accounts for the ∼ 20% reduction. Morphological complexity does not act differently on SOV (mattr_z×is_SOV, p=.39p=.39; M6) and its own effect shrinks with data (within SVO, MATTR × log-size β=−0.017β=-0.017, p=.014p=.014), so it cannot explain an effect that grows with data. Training data composition shows no size interaction on its own (data_comp×size, p=.098p=.098; M8); in the full model (M9) it is collinear with resourcedness (r=0.64r=0.64), so neither size interaction is individually significant there, though the word order interaction is unchanged. A.6 Additional Results Additional artificial language results appear in Figures 6, 7, and 11; natural language results in Figures 8–10, 12, and Table 7. Figure 6: Replication of Figure 4 with 10 MB training samples. Setup and notation are identical. All five paired t-tests remain significant after Holm–Bonferroni correction (p<.001p<.001). Figure 7: Extended version of Figure 5 showing all six base orders. SV orders: right-branching initially easier, flips after ∼ 100 steps. VS orders: left-branching preferred throughout. Same trend at 10 MB. Figure 8: Base word order distribution across attested languages. Figure 1a uses only the N=526N=526 subset with complete switch annotations for all five parameters, which skews toward well-documented, right-branching Indo-European languages. Order n 5 MB 10 MB 100 MB 1 GB SVO 13 2.20 2.08 1.60 1.38 SOV 4 2.16 2.05 1.67 1.49 VSO 1 2.26 2.15 1.72 1.50 NoDominant 2 2.21 2.08 1.61 1.38 Table 7: Median BPEC of Goldfish models on PUD treebank languages by base word order. SOV leads at small scales; SVO leads at large scales, mirroring FLORES-200 (Table 1). Per-order samples are small (n=1n=1–1313), so we omit IQR and test the ordering directly: the SVO advantage reaches significance by 1 GB (SVO vs. SOV Mann–Whitney p=.045p=.045), and the SOV×size interaction replicates FLORES-200 (mixed-effects β=+0.04β=+0.04, p=.04p=.04). Figure 9: Directionality vs. BPEC at all four Goldfish scales, same 62 languages (cf. Figure 1c–d for 5 and 1 GB only). SVO separates from SOV at 100 MB–1 GB. Figure 10: As Figure 9 but with all available languages per scale. Same pattern, ruling out selection artefacts. Figure 11: Convergence AUC (÷103 10^3; lower = faster) for all 32 artificial language configurations at 5 MB (a) and 10 MB (b), ranked by cross-order mean. Rankings stable across sizes (Spearman ρ=0.884ρ=0.884) and strongly correlated with final PPL (ρ>0.89ρ>0.89). Figure 12: BPEC by base word order (a–c) and head directionality (d–f) for the largest model in each multilingual family (FLORES-200); smaller sizes show the same patterns. SVO compresses better than SOV in XGLM and mGPT; BLOOM reverses because 21 of 28 SVO languages are very-low-resource Niger-Congo languages (§5.2). The right-branching SVO advantage is visible even in BLOOM’s directionality panels (d–f), where it is masked at the base word order level. A.7 Artefact Use and Licensing All artefacts are used for research consistent with their intended purpose. FLORES-200 (43) and PUD (58): C-BY-SA 4.0; WALS (14) and Grambank (49): C-BY 4.0; BLOOM (7): BigScience RAIL v1.0; XGLM (37): MIT; mGPT (48) and Stanza (46): Apache 2.0; Wuggy (26): GPL; wordfreq (51): MIT. Goldfish models (9) are publicly available but unlicensed.