Paper deep dive
Augmenting Text to Increase Translation Difficulty
William Kalikman, Šimon Sukup, Michal Tešnar, Vilém Zouhar
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/22/2026, 3:20:24 AM
Summary
The paper introduces Adversarial Translation Optimization (ATO), a gradient-based method to increase the difficulty of machine translation benchmarks by iteratively replacing tokens in source texts. ATO combines a differentiable translation difficulty estimator (Sentinel) with fluency constraints (CoLA, perplexity) using Beam Search. Two variants, ATO-Direct (whole-word substitutions) and ATO-TwoPhase (subword optimization followed by fluency recovery), are proposed. The method lowers translation quality scores (xCOMET) significantly compared to baselines while maintaining grammaticality, and the authors release two datasets of 350 augmented English texts.
Entities (10)
Relation Signals (10)
ATO-Direct → isvariantof → Adversarial Translation Optimization
confidence 98% · We call the resulting procedure Adversarial Translation Optimization (ATO) and present two variants. ATO-Direct restricts substitutions to whole-word tokens
ATO-TwoPhase → isvariantof → Adversarial Translation Optimization
confidence 98% · We call the resulting procedure Adversarial Translation Optimization (ATO) and present two variants. ATO-TwoPhase first optimizes over the full subword vocabulary
Adversarial Translation Optimization → lowers → XCOMET
confidence 95% · Our ATO-modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82
Adversarial Translation Optimization → uses → Sentinel-src-25
confidence 95% · We instantiate the difficulty estimator with Sentinel (Perrella et al., 2024)... Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective
Adversarial Translation Optimization → uses → Beam Search
confidence 95% · optimization becomes a tree search problem, which we address with Beam Search.
ATO-Direct → uses → Qwen-2.5-72B
confidence 93% · All candidates from all steps of the Beam Search are scored with Qwen 2.5-72B perplexity.
ATO-TwoPhase → uses → Qwen-2.5-72B
confidence 93% · Phase 2 operates in Qwen 2.5′s token space... The Beam Search then runs for 40 steps using Qwen 2.5-72B perplexity as the sole objective
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator. Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective to iteratively replace tokens. Because each step branches over candidate substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient-based alternative to LLM-based dataset creation without LLM prompting, expensive human curation, or task-specific model training. Our ATO-modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero-shot baseline. Human evaluation shows the modified texts are somewhat less natural than the baselines but remain reasonably grammatical and plausible while being substantially harder to translate. We release two datasets of 350 English texts each, generated by our methods, as well as the code.
Tags
Links
- Source: https://arxiv.org/abs/2608.15932v1
- Canonical: https://arxiv.org/abs/2608.15932v1
Trouble viewing inline? Open PDF directly →
Full Text
67,423 characters extracted from source content.
Expand or collapse full text
Augmenting Text to Increase Translation Difficulty William Kalikman ★ Šimon Sukup ★ Michal TešnarVilém Zouhar ETH Zurich wkalikman,ssukup,mtesnar,vzouhar@ethz.ch Abstract As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish be- tween models of varying quality. We propose aug- menting existing benchmarks to increase transla- tion difficulty by combining adversarial optimiza- tion with a differentiable translation difficulty estimator. Our Adversarial Translation Optimiza- tion (ATO) uses gradients from a combined diffi- culty and fluency objective to iteratively replace tokens. Because each step branches over candi- date substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient-based alternative to LLM-based dataset creation with- out LLM prompting, expensive human curation, or task-specific model training. Our ATO-modi- fied benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero-shot baseline. Human evaluation shows the modified texts are somewhat less natural than the baselines but re- main reasonably grammatical and plausible while being substantially harder to translate. We release two datasets of 350 English texts each 2 , generated by our methods, as well as the code. 1 1 Introduction Recent advances in machine translation have led to performance saturation on standard benchmarks, with state-of-the-art models achieving near-identi- cal scores (Kocmi et al., 2025; Akhtar et al., 2026). Distinguishing between models of varying quality requires more challenging evaluation data. Existing efforts to build harder benchmarks follow three main strategies, each with distinct * First authors with equal contribution 2 huggingface.co/datasets/wskal/ATO-datasets 1 github.com/BreakingMT/ATO Figure 1: Text perplexity vs. translation quality (averaged across five languages and three models). Lower xCOMET indicates worse translation quality. ATO-TwoPhase reduces translation quality substantially while slightly decreasing per- plexity. Original: The Iraq Study Group presented its report at 12.00 GMT today. Optimized: That odd old man insulted Marta at the market today. Perplexity: 19.77Perplexity: 70.11 MT Difficulty: 53%MT Difficulty: 71% Table 1: One example from our ATO-TwoPhase pipeline. The output is more difficult to translate while remaining fluent and grammatically correct. limitations. Expert-crafted challenge sets (Isabelle et al., 2017; Macketanz et al., 2018; Akhbardeh et al., 2021) are targeted but expensive and difficult to scale. Mining difficult texts from the web (Xu et al., 2025) is limited by the availability of naturally difficult content. LLM-based rewriting (Zouhar et al., 2025) delegates the task to a black-box model which exposes no per-position gradient indicating how each token contributes to difficulty or fluency. It therefore offers no mechanism to incorporate such constraints directly into the optimization. We propose an alternative direction based on adversarial input optimization. Given a differen- tiable estimator of translation difficulty, we use its gradient signal to iteratively modify a source Shterionov, Vanmassenhove, De Sisto, Blain, Pourmostafa, Lepp, Manna, Rescigno, Karakanta, Rigouts Terryn, Lardelli, Resende, Murgolo, Hackenbuchner, Zaretskaya, Espl`a-Gomis, Etchegoyhen, Gromann, Bawden, Haddow, Szoc, Forcada, Moniz (Eds.) Proceedings of the 26th Annual Conference of the European Association for Machine Translation (EAMT 2026), p. 341–358 Tilburg, The Netherlands, June 2026. text so that it becomes harder to translate. We instantiate the difficulty estimator with Sentinel (Perrella et al., 2024), a regression model that predicts expected translation quality from source text alone. This reframes the generation of hard- to-translate text as finding adversarial examples that minimize the Sentinel score, connecting our work to existing methods for adversarial opti- mization in NLP. However, naively following the difficulty gradient degrades fluency: the optimizer finds texts that are hard to translate only because they are nonsensical. Generating text that is both difficult to translate and well-formed is therefore a constrained search problem. We address this by combining Greedy Coordinate Gradient-style (Zou et al., 2023) candidate selection with Beam Search and a differentiable grammaticality signal, jointly steering the optimization toward tokens that are difficult to translate and linguistically well- formed. We call the resulting procedure Adversarial Translation Optimization (ATO) and present two variants. ATO-Direct restricts substitutions to whole-word tokens, enforcing fluency by con- struction and selecting the final output by perplex- ity, a standard proxy for fluency in text generation (Kann et al., 2018; Holtzman et al., 2020). ATO- TwoPhase first optimizes over the full subword vocabulary to generate hard-to-translate text in a first phase, then recovers fluency in a second phase by optimizing for perplexity under a different lan- guage model. Figure 1 and Table 1 illustrate this process: ATO-TwoPhase produces texts that are substantially harder to translate while maintaining fluency. We evaluate our methods on 200 seed texts translated into five target languages by three translation models of varying capability. Both ATO variants produce harder-to-translate text than paraphrasing and zero-shot baselines: aver- age xCOMET, a reference-free translation quality estimation metric, drops from 0.93 (base) to 0.82 (ATO-TwoPhase), compared to 0.88 for para- phrasing and 0.86 for the zero-shot baseline. Through a human annotation study involving 13 multilingual participants, we confirm that the re- sulting translations are of lower quality. Our contributions are (1) Adversarial Trans- lation Optimization (ATO), a gradient-based method for increasing translation benchmark diffi- culty without LLM prompting, human curation, or task-specific model training; and (2) two datasets of 350 augmented English texts each, produced by ATO-Direct and ATO-TwoPhase. 2 Related Work Adversarial optimization. Adversarial text opti- mization has been approached from several angles. Reinforcement learning (Vijayaraghavan and Roy, 2020), generative adversarial networks (Ren et al., 2020), and continuous relaxation methods that operate in the embedding space (Ebrahimi et al., 2018; Sadrizadeh et al., 2023) have all been applied, but these approaches offer limited control over individual token-level substitutions. HotFlip (Ebrahimi et al., 2018) introduced the use of gradients with respect to one-hot token representations to identify single-token substitu- tions. Greedy Coordinate Gradient (Zou et al., 2023) computes top-k candidate replacements at every modifiable token position, then evaluates a random batch of single-token swaps drawn from across all positions, selecting the substitution that most reduces the adversarial loss. While effective at finding adversarial suffixes for LLM jailbreak- ing, it produces disfluent or nonsensical text, as the optimization has no incentive to preserve fluency or grammaticality. Fluency-constrained adversarial search. Sev- eral methods address this fluency limitation. BeamAttack (Zhu et al., 2023) extends gradient- guided substitution with Beam Search, allowing exploration of locally suboptimal candidates that may lead to more fluent texts in later iterations. BESA (Yang et al., 2021) uses a masked language model to propose replacements that are simultane- ously fluent and adversarially effective, combined with energy-based annealing to escape local min- ima. AutoDAN (Liu et al., 2024) departs from gradient-based methods, adopting a hierarchical genetic algorithm with LLM-based fitness metrics to maintain fluency. These methods demonstrate that balancing adversarial effectiveness with flu- ency is a recurring challenge, though all operate in the setting of adversarial suffix generation for LLM safety. Dataset curation and generation. Our work also addresses limitations in how difficult translation benchmarks are sourced. Traditional challenge sets rely on expert curation (Kocmi et al., 2025), which is targeted but expensive and difficult to scale. Filtering approaches search existing web corpora for sentences that models naturally strug- 342 gle with (Proietti et al., 2025; Xu et al., 2025), but are limited by the availability of naturally difficult text. More recently, LLM-based methods generate difficult texts through zero-shot or itera- tive prompting (Pombal et al., 2025; Zouhar et al., 2025), but as black-box generation, this approach yields no per-position gradient over the objec- tive. This precludes gradient-based optimization, including approaches that combine multiple dif- ferentiable losses, such as our combination of difficulty and fluency. 3 Methods Our work differs from previous adversarial opti- mization approaches in three ways. First, we optimize the full source text rather than appending adversarial suffixes, since our goal is to produce complete, fluent texts. Second, we target trans- lation difficulty rather than LLM jailbreaking, using a learned translation difficulty estimator as part of the optimization objective. Third, rather than relying on post-hoc filtering for fluency, we incorporate a differentiable grammaticality signal directly into the gradient computation, jointly steering candidate selection toward tokens that are both difficult to translate and linguistically well- formed. Our search procedure combines Greedy Coordinate Gradient-style gradient-guided candi- date selection with Beam Search to escape local minima and find fluent, difficult-to-translate se- quences. We refer to the full procedure as Adver- sarial Translation Optimization (ATO). 3.1 Optimization Objective We optimize over a token sequence 푡=⟨푡 1 ,...,푡 퐿 ⟩ composed of vocabulary tokens 푡 푖 ∈푉. Our goal is to maximize translation difficulty while main- taining fluency. We assume a differentiable scor- ing function 퐷(푡) that estimates how difficult a source text 푡 is to translate, where lower values indicate greater difficulty. The unconstrained ob- jective is: 푡 ∗ =argmin 푡∈풱︀ 퐿 퐷(푡)(1) We represent the text by a one-hot matrix 푻∈ 0,1 퐿×|풱︀| , where each row 푻 푖 encodes the token at position 푖. Because the one-hot representation enters the model through a differentiable embed- ding lookup, we can compute the gradient of the score with respect to the input: ∇ 푻 퐷= [ * + 휕퐷 휕푻 푖,푣 ] 0 1 푖∈1,...,퐿,푣∈풱︀ (2) The entry ∇ 푻 푖,푣 퐷 gives a first-order approxima- tion of the change in 퐷 if the token at position 푖 were replaced by vocabulary token 푣. These gradients guide the candidate selection procedure described next. Difficulty estimation. We instantiate the function 퐷 with Sentinel-src-25 (Proietti et al., 2025; Per- rella et al., 2024), a source-only translation diffi- culty estimator. Sentinel-src is a regression model built on XLM-RoBERTa-large (Conneau et al., 2020), trained on human translation quality judg- ments to predict the expected quality of a text’s translation from the source text alone. It assigns lower scores to texts whose translations tend to be worse, and is fully differentiable with respect to its input. See example Sentinel judgments: 퐷(“Good morning!”) = 0.46 퐷(“The eccentric aristocrat was not ...”) = 0.34 Generally, 퐷 could be instantiated by any differ- entiable difficulty estimator, for example as the expected quality score from a differentiable trans- lation model composed with a differentiable qual- ity estimation metric. 3.2 Beam Search with Greedy Coordinate Gradient Rather than starting from random tokens, we seed the search with a real text 푡 seed . By initializing from well-formed text, each single-token substitu- tion begins from a grammatical area in the space of all possible texts, so the resulting candidate is likely to remain close to fluent. This makes the fluency constraints described in the following sec- tions effective: they need only prevent gradual drift from grammaticality, rather than recover it from scratch. The optimizer maintains a beam of 퐵 candidate texts (we use 퐵=50). At each iteration, the following steps are applied identically across all candidates in the beam: 1. Backward pass. Token gradients ∇ 푻 ℒ︀ are computed over the full 퐿×|풱︀| one-hot input for every beam member, where ℒ︀ is the loss defined in Section 3.3 and Section 3.4. 2. Top-푘 shortlist. The top 퐾=200 (position, token) pairs ranked by gradient magnitude are selected from the flattened 퐿×|풱︀| gradi- ent matrix, reducing the search space to the 343 Current Beam B members Gradient Sampling N·B candidates Global Selection rank all N·B by objective New Beam new B members x (1) t L = ℓ 1 . . . x (B−1) t x (B) t L = ℓ B Mutation strategy position ∝∥∇∥ token ∼ top-K shortlist ̃ x (1,1) ̃ x (1,2) . . . ̃ x (1,N) ̃ x (B,1) ̃ x (B,2) . . . ̃ x (B,N) Expanding tree of candidate mutations Global pool N·B candidates Ranking criterion Phase 1: L Sent + λ f L CoLA + λ c L Cos Phase 2: PPL Qwen Top-(B−R) by L ̃ x (B,1) (rank 1) ̃ x (1,1) (rank 2) . . . Diversity injection Random sample 1 . . . Random sampleR Pruned (low L) ̃ x (1,N) (rank > B) . . . x (1) t+1 from ̃ x (B,1) x (2) t+1 from ̃ x (1,1) . . . x (B−R) t+1 from rand 1 . . . x (B) t+1 from rand R Figure 2: Beam search iteration visualization. In each iteration of ATO-Direct and ATO-TwoPhase, we first expand the current texts in the beam by testing candidate replacements under our constraints, prune based on the filtering metric, then choose the final set for the next beam. The random sampling of non-optimal beams ensures the beams do not collapse to a conforming set of texts due to local minima of the chosen objective. most promising single-token substitutions. Po- sitions corresponding to punctuation tokens are excluded, and positions that were substituted in either of the two preceding iterations are frozen, preventing the optimizer from repeat- edly cycling the same positions. 3. Pool sampling. For each beam member, 푁= 64 candidate mutations are sampled: the posi- tion is drawn proportionally to its gradient norm, and the replacement token is drawn uniformly from that position’s shortlist. This yields approximately 푁×퐵 unique candi- dates per step. 4. Objective scoring. All candidates are evalu- ated under the current objective ℒ︀, defined in Section 3.3 and Section 3.4. 5. Beam pruning. The 퐵 candidates with the lowest ℒ︀ are retained globally, with competi- tive selection across all beam parents. 6. Diversity injection. An additional 15 candi- dates are drawn randomly from the broader pool and added to the beam, preventing prema- ture convergence. The diagram of one iteration of Beam Search can be seen in Figure 2 and the corresponding algorithm in Algorithm 1. Left unconstrained, this optimization yields nonsensical text, which is dif- ficult to translate only because it is ungrammatical, lacks semantic coherence, or contains fragments that do not form real words. We address this with two schemes: •ATO-Direct restricts the vocabulary to whole English words and selects the final output by choosing the lowest perplexity among candi- dates. •ATO-TwoPhase allows the full subword vocab- ulary but follows the initial optimization with a second phase that explicitly optimizes for fluency. See Table 2 for a comparative overview. 3.3 ATO-Direct In ATO-Direct, one phase of Beam Search with Greedy Coordinate gradient is executed, optimiz- ing only over whole-word tokens in the seed text with whole-word token candidates. For an overview, refer to Algorithm 2. Vocabulary. We load a list of ∼0.5 million English words from the dwyl/english-words repository 2 and tokenize each with XLM- RoBERTa. We retain only words that the tokenizer encodes as a single token, yielding ∼10,000 to- 2 github.com/dwyl/english-words 344 BeamSearchGCGStep(ℬ︀ 푡 , ℒ︀(⋅), 풱︀, 퐾,푁,퐵,푅): 1 for 풙∈ℬ︀ 푡 do 푮 (풙) ←∇ 푿 ℒ︀(풙)// Gradients measure token substitution impact 2 end for 3 풮︀←Top-퐾(푮,ℬ︀ 푡 ,풱︀)// Shortlist 풮︀ of most impactful tokens by gradient 4 풫︀←∅ 5 for 풙∈ℬ︀ 푡 do// initialize candidate pool 풫︀ 6 for 푗=1 to 푁 do// 푁 candidates in each beam member 푥 7 푖 ⋆ ∼푝(푖)∝‖푮 (풙) 푖,⋅ ‖ // Sample position 푖 ⋆ by gradient norm 8 푣 ⋆ ∼Uniform(풮︀ 푖 ⋆ )// Sample token 푣 ⋆ from shortlist 9 풫︀←풫︀∪Replace(풙,푖 ⋆ ,푣 ⋆ )// Replace token and add candidate to 풫︀ from shortlist 10 end for 11 end for 12 ∀ ̃ 풙∈풫︀:푠( ̃ 풙)←ℒ︀( ̃ 풙)// Save ranking of candidates in mapping 푠 13 ℬ︀ 푡+1 ←Top 퐵−푅 (풫︀,푠)// Top 퐵-푅 candidates by objective score ℒ︀ 14 ℛ︀←SampleRandom(풫︀∖ℬ︀ 푡+1 ,푅)// 푅 randomly sampled beam members for diversity injection 15 return ℬ︀ 푡+1 ∪ℛ︀,풫︀// Returns 퐵 new members Algorithm 1: One step of Beam Search. Parameters: Current Beam ℬ︀ 푡 of width 퐵, objective ℒ︀, vocabulary 풱︀, shortlist size 퐾, pool samples 푁, random injections 푅. kens. During optimization, positions in the seed text whose original word was split into multiple subword tokens are frozen; only positions that cor- respond to a single whole-word token are eligible for substitution. Replacements are likewise drawn exclusively from the ∼10,000 whole-word vocab- ulary, so every swap replaces one complete word with another. Objective. To further encourage grammaticality beyond the vocabulary constraint, we introduce a differentiable fluency term by fine-tuning a gram- matical acceptability classifier on the Corpus of Linguistic Acceptability (CoLA, Warstadt et al., 2019). We fine-tune XLM-RoBERTa-large (Con- neau et al., 2020) for binary sequence classifica- tion (acceptable / unacceptable) using the CoLA training set from the GLUE benchmark (Wang et al., 2019). Sentinel-src also uses XLM-RoBERTa- large as its encoder, so both models share the same tokenizer and vocabulary. This allows us to compute gradients from both models with respect to a single one-hot representation 푻. The CoLA fluency loss is the cross-entropy between the classifier’s output and the target label “acceptable”: ℒ︀ CoLA (푡)=−log푃[CoLA(acceptable|푡)](3) To discourage beam candidates from fragmenting into dissimilar clusters, we introduce a cosine- similarity penalty over sentence-level representa- tions. Following (Ebrahimi et al., 2018), we use the CLS-pooled representation, obtained from the XLM-RoBERTa backbone already employed in Sentinel. This term penalizes each candidate for drifting from the beam mean: ℒ︀ cos (푡)=1−cos(풉,풉)(4) where 풉 is the candidate’s CLS representation and 풉 is the mean CLS representation across the cur- rent beam. The full objective combines all three terms: ℒ︀ sent+CoLA (푡)= 퐷(푡) ⏟ difficulty +휆 푓 ⋅ℒ︀ CoLA (푡) ⏟ fluency +휆 푐 ⋅ℒ︀ cos (푡) ⏟ cohesion (5) with 휆 푓 =20 and 휆 푐 =50. Because all three terms are differentiable with respect to 푻, the gradient ∇ 푻 ℒ︀ sent+CoLA jointly steers candidate selection toward tokens that are difficult to trans- late, grammatically natural, and consistent across the beam. This is the objective ℒ︀ referenced in Section 3.2. Selection. All candidates from all steps of the Beam Search are scored with Qwen 2.5-72B per- plexity. The candidate with the lowest perplexity is selected as the final output. 3.4 ATO-TwoPhase ATO-Direct’s whole-word constraint enforces flu- ency but limits the search space. ATO-TwoPhase takes a two-phased approach to allow for a broader search: it allows the full subword vocabulary in a first phase of Beam Search to reach lower Sentinel scores, then cleans up the resulting text in a second Beam Search phase by optimizing for fluency. Phase 1. Using the same ∼0.5 million word English list, we tokenize each word with XLM- RoBERTa and retain all tokens that appear in any tokenization. This yields ∼25,000 tokens: the vocabulary now includes subword fragments but remains limited to English. 345 ATO-Direct(풙 1:퐿 , 풱︀ word , 푇, 퐾,푁,퐵,푅): 1 ℬ︀ 0 ←풙 1:퐿 // Seed beam with input sentence 2 |(풉)←CLS(풙 1:퐿 )// Initial beam-mean embedding 3 풞︀←∅// Global candidate archive 4 for 푡=1 to 푇 do 5 ℒ︀ 1 (⋅)←ℒ︀ Sentinel +휆 푓 ℒ︀ CoLA +휆 푐 (1−cos(CLS(⋅),|(풉))) 6 ℬ︀ 푡 ,풫︀ 푡 ← BeamSearchStep (ℬ︀ 푡−1 ,ℒ︀ 1 ,풱︀ word ,퐾,푁,퐵,푅) 7 |(풉)← 1 | ℬ︀ 푡 |∑ 풙∈ℬ︀ 푡 CLS(풙) // Update beam-mean embedding 8 풞︀←풞︀∪풫︀ 푡 // Archive all candidates 9 end for 10 풙 ⋆ ←argmin 풙∈풞︀ PPL Qwen(풙) // Save lowest-PPL text 11 return 풙 ⋆ Algorithm 2: In ATO-Direct, the single Beam Search Greedy Coordinate Gradient phase uses the same composite objective as ATO-TwoPhase but restricts substitutions to complete English words (|풱︀ word |≈10k), enforcing fluency by construction. Selection is done by Qwen 2.5-72B perplexity ranking over all candidates across all steps. ATO-TwoPhase(풙 1:퐿 , 풱︀ sub , 풱︀ Qwen , 푇 1 , 푇 2 , 퐾,푁,퐵,푅): 1 ℬ︀ 0 ←풙 1:퐿 // Seed beam with input text 2 |(풉)←CLS(풙 1:퐿 )// Initial beam-mean embedding 3 for 푡=1 to 푇 1 do 4 ℒ︀ 1 (⋅)←ℒ︀ Sentinel +휆 푓 ℒ︀ CoLA +휆 푐 (1−cos(CLS(⋅),|(풉))) 5 ℬ︀ 푡 ,풫︀ 푡 ← BeamSearchStep (ℬ︀ 푡−1 ,ℒ︀ 1 ,풱︀ sub ,퐾,푁,퐵,푅) 6 |(풉)← 1 | ℬ︀ 푡 |∑ 풙∈ℬ︀ 푡 CLS(풙) // Update beam-mean embedding 7 end for 8 ℬ︀ 푇 1 ←Re-tokenise(ℬ︀ 푇 1 ,풱︀ Qwen ) // Switch to Qwen token space 9 풞︀←∅// Global candidate archive 10 for 푡=푇 1 +1 to 푇 1 +푇 2 do 11 ℬ︀ 푡 ,풫︀ 푡 ← BeamSearchStep (ℬ︀ 푡−1 ,PPL Qwen ,풱︀ Qwen ,퐾,푁,퐵,푅) 12 풞︀←풞︀∪풫︀ 푡 // Archive all candidates 13 end for 14 풙 ⋆ ←argmin 풙∈풞︀ PPL Qwen(풙) // Lowest PPL across all Phase-2 steps 15 return 풙 ⋆ Algorithm 3: ATO-TwoPhase. Phase 1 optimizes Sentinel difficulty, CoLA fluency, and cosine similarity over the full XLM-R subword vocabulary (|풱︀ sub |≈25,000). Phase 2 switches to Qwen’s vocabulary (|풱︀ Qwen |≈87,000) and optimizes for perplexity. The output is the text with the lowest Qwen perplexity encountered across all Phase 2 candidates. The objective is identical to ATO-Direct (Equa- tion 5): ℒ︀ sent+CoLA , combining Sentinel difficulty, CoLA fluency, and cosine cohesion with the same hyperparameters. The only difference is that every token in the original text can be changed, unlike in Section 3.3 where only whole-word tokens were candidates for optimization. Because the vocabu- lary is more permissive, the optimizer can reach lower Sentinel scores, but may produce texts con- taining combinations of subword fragments that do not form coherent words. After the Beam Search completes, the single candidate with the lowest Sentinel score (highest estimated translation difficulty) is selected to seed Phase 2. Phase 2 operates in Qwen 2.5′s token space. We apply an ASCII filter over Qwen’s full vo- cabulary (∼150,000 tokens), retaining any token whose decoded string contains only ASCII letters, digits, spaces, and common punctuation. This yields ∼87,000 tokens: English-compatible sub- word fragments with non-Latin scripts removed. The Phase 1 output is replicated into a fresh beam of width 퐵=50. The Beam Search then runs for 40 steps using Qwen 2.5-72B perplexity as the sole objective: ℒ︀ PPL (푡)=PPL Qwen(푡) (6) Gradients now flow through Qwen’s embedding matrix rather than XLM-RoBERTa’s. The opti- mizer structure is otherwise identical to Phase 1. The candidate with the lowest perplexity across all Phase 2 steps is selected as the final output. See Algorithm 3 for the method overview. 3.5 Baselines We compare our methods against three baselines that a practitioner would intuitively use to generate harder-to-translate text. Zeroshot. Our ATO-Direct approach replaces one to three words of the original text. As a comparable 346 ATO-DirectATO-TwoPhase — Phase 1ATO-TwoPhase — Phase 2 VocabularyWhole-word (9,865)English subwords (25,329)Qwen English (87,221) Eligible positionsSingle-token words onlyAll positionsAll positions Objective ℒ︀ sent+CoLA ℒ︀ sent+CoLA ℒ︀ PPL SelectionLowest Qwen PPLLowest Sentinel scoreLowest Qwen PPL FluencyHard (vocabulary) + soft (CoLA)Soft (CoLA)Soft (PPL optimization) Table 2: Comparison of Greedy Coordinate Gradient Beam Search variants across our two ATO methods. baseline, we prompt two LLMs to replace exactly two words in a text to make it harder to translate (see prompt in Figure 4): Qwen2.5-72B-Instruct, which doubles as the perplexity scorer in ATO, and the more recent frontier model DeepSeek-V4- Flash (DeepSeek-AI, 2026). See Appendix A for per-language results. Paraphrasing. Our ATO-TwoPhase approach has a broader and less predictable effect, changing a variable number of words while preserving some semantic content. The paraphrasing model DIPPER (Krishna et al., 2023) is a comparable baseline, taking a text as input and outputting a rephrased version. We discuss the details of our DIPPER implementation in Appendix B.3. Random replacement. To check our methods against random perturbations, we replace random tokens in each seed text; see Appendix A for im- plementation details and results. 4 Experiments We evaluate ATO’s ability to generate source texts that are harder to translate while preserving linguistic quality. Since our generated texts have no reference translations, we cannot use standard reference-based metrics such as BLEU (Papineni et al., 2002; Post, 2018). And since Sentinel-src is the optimization objective itself, using it to evaluate difficulty would be circular. We therefore assess difficulty by translating the generated texts with multiple models and measuring translation quality using reference-free automatic metrics and human judgments. Setup. We evaluate on seed texts sampled from FLORES200 (NLLB Team, 2024), WMT22, WMT23, and WMT24 (Kocmi et al., 2022; Kocmi et al., 2023; Kocmi et al., 2024): 200 base texts for our automatic evaluation and 25 for our human evaluation. Each text is obtained by taking the first 50 tokens of a sample and truncating at the first sentence-ending punctuation; samples with no such punctuation are discarded. The resulting segments are typically 10–20 words long, and may be sentence fragments, titles, or full sentences depending on the source data. For each seed, we compare five variants in the main analysis: the original Base text, two baselines (DIPPER Paraphrasing and Qwen Zeroshot), and the two versions of our method (ATO-Direct and ATO- TwoPhase). The Random and DeepSeek-Zeroshot baselines are reported alongside in the appendix. Each variant is translated from English to five target languages spanning several language fami- lies and resource levels: German, Spanish, Russ- ian, Czech, and Icelandic. We use three translation models of varying capability: the encoder-decoder model NLLB-200-3.3B (NLLB Team, 2024), the open-source TranslateGemma-27b (Finkelstein et al., 2026), and the closed-source frontier LLM Gemini-3-Flash (Google DeepMind, 2025) see Figure 3 for the LLM translation prompt. We addi- tionally compare against Tower-Plus-72B (Rei et al., 2025) and DeepSeek-V4-Flash (DeepSeek-AI, 2026) in Appendix A. 4.1 Automatic Evaluation We use two standard reference-free metrics: • MetricX (Juraska et al., 2024) estimates trans- lation error on a scale from 0 (perfect) to 25 (worst). Higher scores indicate lower quality translations. For comparability, we report Met- ricX in Table 3 as 1− MetricX 25 . • xCOMET (Guerreiro et al., 2024) scores trans- lations from 0 to 1, where 1 indicates a perfect translation. Lower scores indicate lower quality translations. Sentinel and perplexity. Table 3 shows that both ATO variants substantially reduce Sentinel-src scores relative to the base texts and both baselines, confirming that the optimization successfully finds texts the difficulty estimator considers harder to translate. ATO-TwoPhase achieves the lowest Sen- 347 MethodPerplexitySentinelxCOMETMetricX (human) Grammatic. (human) Plausible (human) Transl. Qual. Base text69.3 ±42.62 0.29 ±0.03 0.93 ±0.02 0.91 ±0.03 4.0 ±0.28 4.4 ±0.20 3.8 ±0.31 Paraphrasing62.4 ±42.42 0.25 ±0.03 0.88 ±0.04 0.88 ±0.04 4.0 ±0.28 4.4 ±0.19 3.9 ±0.29 Zeroshot (Qwen)119.4 ±65.89 0.18 ±0.03 0.86 ±0.04 0.87 ±0.05 3.8 ±0.30 4.0 ±0.23 3.9 ±0.29 ATO-Direct224.9 ±95.25 0.02 ±0.04 0.86 ±0.04 0.85 ±0.04 2.9 ±0.30 3.0 ±0.26 3.5 ±0.32 ATO-TwoPhase68.5 ±11.95 0.00 ±0.04 0.82 ±0.03 0.82 ±0.04 2.4 ±0.27 2.8 ±0.28 3.4 ±0.28 Table 3: Aggregate evaluation scores comparing proposed ATO methods against baselines. Greener means better under the metrics we evaluated: higher fluency text, worse translation quality. MetricX scores are presented as 1− MetricX 25 for comparability. All values are presented as mean with 95% confidence interval. tinel scores, as expected from its larger search space. Perplexity increases moderately for ATO- Direct but decreases for ATO-TwoPhase, reflect- ing Phase 2′s explicit perplexity optimization; the resulting texts are, by this measure, no less fluent than the originals. The Paraphrasing and Zeroshot baselines leave Sentinel scores approximately un- changed. Translation quality. Since Sentinel-src is the optimization objective itself, using it to evaluate difficulty would be circular. The key question is whether lower Sentinel scores correspond to genuinely worse translations under indepen- dent metrics. Table 3 confirms this across both metrics, with scores averaged across all five tar- get languages and three translation models. On xCOMET, ATO-TwoPhase lowers scores from 0.93 to 0.82, outperforming both the Paraphrasing baseline (0.88) and Zeroshot (0.86). ATO-Direct performs comparably to Zeroshot on xCOMET (both 0.86) and MetricX (0.85 vs. 0.87 with over- lapping CIs). On both metrics, ATO-TwoPhase achieves the largest quality decrease, and its differ- ences from both baselines exceed the 95% confi- dence intervals. Variation across models. The difficulty in- crease is consistent across all three translation models. NLLB-200-3.3B, the weakest model, shows the largest absolute xCOMET drop under ATO-TwoPhase (0.13 points averaged across lan- guages), but the effect is not limited to weaker models: Gemini-3-Flash drops by 0.11 points and TranslateGemma by 0.09. This confirms that ATO-generated texts challenge models across ca- pability levels, not only those that are already fragile. Full per-language, per-model breakdowns appear in Appendix A. Variation across languages. The effect holds for all five target languages but varies in magnitude. Czech, Russian, and Icelandic show the largest xCOMET drops (0.12–0.14 points), while Spanish (0.10) and German (0.07) prove more resilient. See Appendix A for more details. ATO-Direct vs. -TwoPhase. ATO-TwoPhase con- sistently achieves lower translation quality scores than ATO-Direct across all model–language pairs. As Figure 1 illustrates, ATO-TwoPhase also achieves lower perplexity thanks to its explicit flu- ency optimization phase. However, as we show in the human evaluation below, ATO-Direct’s whole- word constraint produces texts that human raters judge as more grammatical and plausible than ATO-TwoPhase, suggesting that perplexity alone does not capture all aspects of naturalness. 4.2 Human Evaluation We implemented two human evaluation schemes: one evaluating the quality of the generated English texts and another evaluating the quality of their translations. We recruited 13 annotators from the authors’ academic network. For English text quality, the respondents eval- uated 125 total texts: five variants (the same as described in Section 4) of 25 base texts. Evaluators rated each text on a scale of 1–5 along two dimen- sions: grammaticality, how grammatically well- formed the text is, and plausibility, how likely it is that the text would appear in real online content. See the full evaluation instructions in Figure 5. For translation quality, we translated all five variants of the 25 source texts into five target languages: German, Spanish, Russian, Czech, and Icelandic. We translated the texts with NLLB-200-3.3B. Annotators then rated each translation on a scale of 1–5. See the full transla- 348 Phase 1 The Iraq Study Group presented its report at 12.00 GMT today. The Iraq Study Group banad its report at 12.00 GMT today. The Iraq Study Group banad its duo at 12.00 GMT today. The Iraq Study Group magasind its mister at off GMT today. The flavor Study Group banad its duo at kul GMT today. The sensitive shopping duo frontd its plate at leaving GMT today. bone sensitive shopping duo frontd its maya at keeping GMT today. Phase 2 bone sensitive shopping duo frontd its maya at keeping GMT today. bone bones shopping duo frontd its maya at keeping GMT today. bone sensitive shopping duo frontd its npa at around GMT today. ... Some other shopping clerk insulted anna at the supermarket today. That old shop clerk insulted my nona at the supermarket recently. some other food clerk insulted anna at the supermarket today. ... That old store clerk insulted my grandpa at the river market. “s old shop clerk insulted his grandpa at the fish market. Quite old shop worker insulted my Grandpa at the local market. Quite old shop worker insulted my Grandpa at the local market. Some other old woman insulted Marta at the market today. Some other old woman insulted Marta at the market today. ... “Oh that old woman insulted our Marta at the market today. “Oh that old woman insulted our Marta at the market today. That odd old man insulted Marta at the market today. Table 4: ATO-TwoPhase trace illustrated for one seed text. The fluent seed text is iteratively degraded by Phase 1, yielding a nonsensical but hard-to-translate segment. Phase 2 iteratively makes the Phase 1 output more fluent, resulting in a coherent, still hard-to-translate output. tion evaluation instructions in Figure 6 and inter- annotator agreement in Appendix C.3. Screen- shots of the evaluation interface can be seen in Appendix C.4. Table 3 presents human ratings across the five methods with 95% confidence intervals. Both ATO-Direct and ATO-TwoPhase achieve lower translation quality scores compared to the base- lines, confirming that ATO-generated texts are harder to translate. This comes at the cost of lower grammaticality and plausibility ratings compared to the baselines and seed texts. Between the two ATO variants, ATO-Direct achieves a comparable reduction in translation quality while retaining higher human-rated fluency, making it a more suit- able choice for generating natural-sounding texts under the human evaluation metrics. 5 Discussion The results confirm that ATO successfully lowers the translation quality compared to zero-shot and paraphrasing baselines. Across all automatic met- rics and human judgments of translation quality, both ATO variants consistently produce harder- to-translate texts than either baseline, with ATO- TwoPhase achieving the largest difficulty increase. Nonetheless, the human evaluation Table 3 has shown a tradeoff between our naturalness proxies, grammaticality and plausibility, and the transla- tion quality. To illustrate how the two ATO-TwoPhase phases interact, we trace a single ATO-TwoPhase example from seed text to final output. Table 1 shows the input/output pair with its perplexity and MT difficulty scores, and Table 4 shows the full optimization trace, with bolded words marking changes at each step. This example illustrates the complementary roles of the two phases. In Phase 1, individual tokens are replaced to maximize translation diffi- culty, often introducing ungrammatical or nonsen- sical words. The output of Phase 1 scores very low on the sentinel difficulty metric but is not human- readable. In Phase 2, the optimization objective shifts to perplexity, iteratively replacing tokens to recover fluency while implicitly preserving diffi- culty. The resulting text is fluent and grammatical, and as shown by the metrics in Table 1, remains more difficult to translate than the base text. Qualitative evaluation. To better understand how ATO breaks machine translation, we examine sev- eral examples translated by NLLB-200-3.3B. A typical failure mode for ATO-Direct translations is the omission by the translation model of a contextually unlikely word that could nonetheless be translated directly. As shown in Figure 1, ATO-Direct produces texts with high perplexity, meaning they contain unexpected word combina- tions. Crucially, this does not mean the texts are untranslatable. In the first example in Table 5, the newly introduced word run could be translated as an adjective and included in the compound noun die Laufinteressen, yet NLLB omits it entirely. Another failure mode of NLLB consists of fail- ing to translate uncommon phrases. In the second example in Table 5, ATO-Direct introduces the awkward phrase it was my cable to assist, where NLLB centers its translation on the word cable, rendering it as telegram (byl to můj telegram, abych vám pomohl) rather than preserving the syntactic structure of the original text. In the third example, ATO-Direct introduces the phrase sorry investment, which could reasonably be inter- preted literally as a bad investment. Instead, NLLB extrapolates beyond the source text and produces inversión que inicialmente parecía un fracaso (investment that initially seemed like a failure), adding meaning not present in the original text. ATO-TwoPhase challenges translation in a dif- ferent way. In the example in Table 6, the method 349 Augmented TextTranslation by NLLB We agree with the USOC’s statement that run interests of our ath letes and clubs, and their sport, may be bet ter served by moving forward with meaningful change within our orga nization, rather than de certification. Wir stimmen mit der Aus sage des USOC überein, dass die Interessen unserer Ath leten und Vereine und ihres Sports besser bedient werden können, wenn wir mit sin nvollen Veränderungen inner halb unserer Organisation vo ranschreiten, anstatt die Zerti fizierung zu entziehen. Thank you for contacting #PRS_#ORG, it was my cable to assist you today. Děkuji, že jste kontaktovali #PRS_ORG#, byl to můj telegram, abych vám dnes po mohl. A sorry investment at launch would be worth over $2 million today! ¡Una inversión que inicial mente parecía un fracaso hoy valdría más de 2 millones de dólares! Table 5: Low-quality translations by NLLB of texts produced by ATO-Direct, showcasing failure modes such as omitting contextually unexpected words and hallucinating explanatory phrases not present in the source. produces the phrase methods mince to people. NLLB generates an incorrect translation, render- ing the phrase as nicht hilfreich sind, not being helpful in German. The model appears to fill in a meaning that fits the surrounding context rather than faithfully translating the source, producing a fluent but incorrect output. Similarly to the first example in Table 5, we observe that a word out of context can worsen the model’s performance. In the second example in Table 6, we see that the word joist is completely ignored. Finally, ATO can cause the translation model to drop content entirely: in the third example of Table 6, NLLB omits a full sentence that was translated correctly from the unmodified source. We hypothesize that the modified tokens push the encoder represen- tations sufficiently out of distribution that the decoder’s attention skips over the affected segment altogether. 6 Conclusion We introduced Adversarial Translation Optimiza- tion (ATO), a gradient-based method for augment- ing text to increase translation difficulty. By com- bining Greedy Coordinate Gradient with Beam Search and a differentiable fluency signal, ATO iteratively modifies source texts to minimize Sen- tinel-src scores while preserving grammaticality. We presented two variants: ATO-Direct, which restricts substitutions to whole-word tokens, and Augmented TextTranslation by NLLB With some other regularity, he is skeptical about how diabetes can be cured, not ing that these methods will not mince to people who al ready have Type 1 diabetes. Mit einer anderen Regelmäßigkeit ist er skep tisch, wie Diabetes geheilt werden kann, und stellt fest, dass diese Methoden für Menschen, die bere its Typ1Diabetes haben, nicht hilfreich sind. The pump head and clamps are both covered by a life time manufacturer joist, so you can immediately re place the product in the event of shrinkage or spool ing. Hlavice čerpadla a svorky jsou oba kryty životním výrobcem, takže můžete okamžitě vyměnit výrobek v případě zmenšení nebo spoolingu. His girlfriend outside of the video footage laughed as she took the video. Two Dol lywood employees tried to break up the fight, including a manager pulling the two apart. Dos empleados de Dolly wood intentaron romper la pelea, incluido un gerente que los separó. Table 6: Low-quality translations by NLLB of texts produced by ATO-TwoPhase, showcasing failure modes such as substi- tuting contextually plausible but incorrect meanings, ignoring out-of-context words entirely, and dropping full sentences. ATO-TwoPhase, which operates over subword vo- cabularies to reach lower difficulty scores before recovering fluency through perplexity-guided op- timization. Experiments on 200 seed texts translated into five languages by three models of varying capabil- ity show that both ATO variants produce substan- tially harder-to-translate text than paraphrasing and zero-shot baselines, as measured by MetricX and xCOMET. Human evaluators confirm that the resulting translations are of lower quality, though ATO-generated texts were rated less grammatical and plausible than baselines. Future work should explore stronger fluency constraints or alternative optimization objectives to better reconcile transla- tion difficulty with naturalness. Our approach requires no LLM prompting, no human intervention or hand-crafted datasets, and provides per-token gradient signal that makes the optimization interpretable and enables direct enforcement of fluency constraints. ATO can be applied to any source text given a differentiable difficulty estimator, enabling scalable construction of challenging translation benchmarks. We release two datasets of 350 augmented texts, one pro- duced by ATO-Direct and one produced by ATO- TwoPhase. 350 Sustainability Statement Dataset creation. Across both datasets, the total Phase 1 runtime was ∼10 hours on an RTX Pro 6000 GPU. Phase 2 runtime was ∼95.5 wall-clock hours across 2x RTX Pro 6000 GPUs. Qwen 72B perplexity scoring took ∼3 wall-clock hours on 2x RTX Pro 6000 GPUs, bringing the total to ∼204 RTX Pro 6000 GPU hours. Evaluation. The translations we ran locally took ∼2 hours on a GeForce RTX 3090. We cannot provide a reliable estimate of the impact of our API usage for models accessed via API; however, we expect it to account for a very small share of total impact relative to dataset creation. All experiments were run on private infrastruc- ture in Switzerland. Accordingly, we use an elec- tricity emissions factor of 0.09 kgCO2e/kWh per the latest reports on Swiss electricity consumption emissions intensity. We use the Machine Learning CO2 Impact Calculator, approximating our RTX Pro 6000 usage with an RTX A6000, as it is the most similar GPU available on the site. We calcu- late a total of 5.51 kg (dataset creation) + 0.06 kg (evaluation) ≈ 5.57 kg CO2. References Farhad Akhbardeh, Arkady Arkhangorodsky, Mag- dalena Biesialska, Ondřej Bojar, Rajen Chatter- jee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Fed- ermann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Ken- neth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021. Findings of the 2021 Con- ference on Machine Translation (WMT21). In Pro ceedings of the Sixth Conference on Machine Trans lation, pages 1–88, Online. Mubashara Akhtar, Anka Reuel, Prajna Soni, San- chit Ahuja, Pawan Sasanka Ammanamanchi, Ru- chit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Šuppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Ar- jun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, and Irene Solaiman. 2026. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation. arXiv: 2602.16763 [cs.AI]. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettle- moyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Asso ciation for Computational Linguistics, pages 8440– 8451. DeepSeek-AI. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. HotFlip: White-Box Adversarial Exam- ples for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Compu tational Linguistics (Volume 2: Short Papers), pages 31–36, Melbourne, Australia. Mara Finkelstein, Isaac Caswell, Tobias Domhan, Jan- Thorsten Peter, Juraj Juraska, Parker Riley, Daniel Deutsch, Geza Kovacs, Cole Dilanni, Colin Cherry, Eleftheria Briakou, Elizabeth Nielsen, Jiaming Luo, Kat Black, Ryan Mullins, Sweta Agrawal, Wenda Xu, Erin Kats, Stephane Jaskiewicz, Markus Freitag, and David Vilar. 2026. TranslateGemma Technical Report. arXiv: 2601.09012 [cs.CL]. Google DeepMind. 2025. Gemini 3 Flash Model Card. Accessed: 2026-03-15. Nuno M. Guerreiro, Ricardo Rei, Daan Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2024. xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection. Transactions of the Association for Computational Linguistics 12:979–995. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. In International Conference on Learning Representations (ICLR). Pierre Isabelle, Colin Cherry, and George Foster. 2017. A Challenge Set Approach to Evaluating Machine Translation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Process ing, pages 2486–2496. Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. MetricX-24: The Google Sub- mission to the WMT 2024 Metrics Shared Task. In Proceedings of the Ninth Conference on Machine Translation, pages 492–504, Miami, Florida, USA. Katharina Kann, Sascha Rothe, and Katja Filippova. 2018. Sentence-Level Fluency Evaluation: Refer- ences Help, But Can Be Spared! In Proceedings 351 of the 22nd Conference on Computational Natural Language Learning (CoNLL), pages 313–323. Tom Kocmi, Ekaterina Artemova, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Konstan- tin Dranch, Anton Dvorkovich, Sergey Dukanov, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Howard Lakougna, Jessica Lundin, Christof Monz, Kenton Murray, Masaaki Nagata, Stefano Perrella, Lorenzo Proi- etti, Martin Popel, Maja Popović, Parker Riley, Mariya Shmatova, Steinthór Steingrímsson, Lisa Yankovskaya, and Vilém Zouhar. 2025. Findings of the WMT25 General Machine Translation Shared Task: Time to Stop Evaluating on Easy Test Sets. In Proceedings of the Tenth Conference on Machine Translation, pages 355–413, Suzhou, China. Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Fed- ermann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Na- gata, Martin Popel, Maja Popović, Mariya Shmatova, Steinthór Steingrímsson, and Vilém Zouhar. 2024. Findings of the WMT24 General Machine Transla- tion Shared Task: The LLM Era Is Here but MT Is Not Solved Yet. In Proceedings of the Ninth Confer ence on Machine Translation, pages 1–46, Miami, Florida, USA. Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Fed- ermann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Masaaki Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, Mariya Shmatova, and Jun Suzuki. 2023. Findings of the 2023 Conference on Machine Translation (WMT23): LLMs Are Here but Not Quite There Yet. In Proceedings of the Eighth Conference on Machine Translation, pages 1–42, Singapore. Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grund- kiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022. Findings of the 2022 Conference on Machine Translation (WMT22). In Proceedings of the Seventh Conference on Machine Translation (WMT). Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023. Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. In Advances in Neural Infor mation Processing Systems (NeurIPS). Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2024. AutoDAN: Generating Stealthy Jail- break Prompts on Aligned Large Language Models. In The Twelfth International Conference on Learning Representations. Vivien Macketanz, Eleftherios Avramidis, Aljoscha Burchardt, and Hans Uszkoreit. 2018. Fine-grained evaluation of German-English Machine Translation based on a Test Suite. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 578–587. NLLB Team. 2024. Scaling neural machine translation to 200 languages. Nature 630:841–846. Kishore Papineni, Salim Roukos, Todd Ward, and Wei- Jing Zhu. 2002. BLEU: A Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318. Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba, and Roberto Navigli. 2024. Guardians of the Machine Translation Meta-Evalua- tion: Sentinel Metrics Fall In! In Proceedings of the 62nd Annual Meeting of the Association for Compu tational Linguistics (Volume 1: Long Papers), pages 16216–16244, Bangkok, Thailand. José Pombal, Nuno M. Guerreiro, Ricardo Rei, and André F. T. Martins. 2025. Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models. In Conference on Language Modeling (COLM). Matt Post. 2018. A Call for Clarity in Reporting BLEU Scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186– 191, Brussels, Belgium. Lorenzo Proietti, Stefano Perrella, Vilém Zouhar, Roberto Navigli, and Tom Kocmi. 2025. Estimating Machine Translation Difficulty. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 24261–24285, Suzhou, China. Ricardo Rei, Nuno M Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André FT Martins. 2025. Tower+: Bridging generality and translation specialization in multilingual llms. arXiv preprint arXiv:2506.17080. Yankun Ren, Jianbin Lin, Siliang Tang, Jun Zhou, Shuang Yang, Yuan Qi, and Xiang Ren. 2020. Gen- erating Natural Language Adversarial Examples on a Large Scale with Generative Models. In ECAI 2020: 24th European Conference on Artificial Intelligence, pages 2156–2163. Sahar Sadrizadeh, Clément Barbier, Ljiljana Dolamic, and Pascal Frossard. 2023. A Relaxed Optimization Approach for Adversarial Attacks against Neural Machine Translation Models. In Proceedings of the 352 31st European Signal Processing Conference (EU SIPCO 2023), pages 436–440, EURASIP. Robyn Speer. 2022. rspeer/wordfreq: v3.0. version V3.0.2. Prashanth Vijayaraghavan and Deb Roy. 2020. Gener- ating Black-Box Adversarial Examples for Text Clas- sifiers Using a Deep Reinforced Model. In Machine Learning and Knowledge Discovery in Databases, pages 711–726. Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A Multi-Task Benchmark and Analysis Plat- form for Natural Language Understanding. In Inter national Conference on Learning Representations. Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2019. Neural Network Acceptability Judg- ments. Transactions of the Association for Computa tional Linguistics 7:625–641. Wenda Xu, Vilém Zouhar, Parker Riley, Mara Finkel- stein, Markus Freitag, and Daniel Deutsch. 2025. Searching for Difficult-to-Translate Test Examples at Scale. arXiv preprint arXiv:2509.26619. Xinghao Yang, Weifeng Liu, Dacheng Tao, and Wei Liu. 2021. BESA: BERT-based Simulated Annealing for Adversarial Text Attacks. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 3293–3299. Hai Zhu, Qinyang Zhao, and Yuren Wu. 2023. BeamAt- tack: Generating High-quality Textual Adversarial Examples Through Beam Search and Mixed Seman- tic Spaces. In Advances in Knowledge Discovery and Data Mining, pages 454–465, Cham. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned lan- guage models. arXiv preprint arXiv:2307.15043. Vilém Zouhar and Tom Kocmi. 2026. Pearmut: Hu- man Evaluation of Translation Made Trivial. arXiv: 2601.02933 [cs.CL]. Vilém Zouhar, Wenda Xu, Parker Riley, Juraj Juraska, Mara Finkelstein, Markus Freitag, and Daniel Deutsch. 2025. Generating Difficult-to-Translate Texts. arXiv preprint arXiv:2509.26592. 353 AAdditional Results We present additional results from the automatic evaluation. A.1xCOMET NLLB-200-3.3BTower- Plus-72B TranslateGemma-27bGemini-3- Flash DeepSeek-V4- Flash EN-CS Base text0.89 ±0.020.94 ±0.010.95 ±0.010.95 ±0.010.94 ±0.01 Random Subword0.66 ±0.040.81 ±0.030.84 ±0.020.84 ±0.020.81 ±0.02 Random Word (Common)0.60 ±0.040.75 ±0.030.80 ±0.020.78 ±0.020.77 ±0.02 Random Word (Rare)0.64 ±0.040.81 ±0.020.83 ±0.020.83 ±0.020.80 ±0.02 Paraphrasing0.78 ±0.040.90 ±0.020.92 ±0.010.91 ±0.020.90 ±0.02 Zeroshot (Qwen 2.5 72B)0.74 ±0.040.88 ±0.020.91 ±0.020.89 ±0.020.89 ±0.02 Zeroshot (DeepSeek-V4-Flash)0.72 ±0.040.87 ±0.020.90 ±0.020.88 ±0.020.87 ±0.02 ATO-Direct0.78 ±0.030.86 ±0.020.90 ±0.020.89 ±0.020.87 ±0.02 ATO-TwoPhase0.73 ±0.030.82 ±0.020.85 ±0.020.81 ±0.030.80 ±0.03 EN-DE Base text0.94 ±0.020.97 ±0.010.97 ±0.010.97 ±0.010.97 ±0.01 Random Subword0.77 ±0.040.91 ±0.010.91 ±0.010.91 ±0.010.91 ±0.01 Random Word (Common)0.73 ±0.040.89 ±0.010.89 ±0.010.89 ±0.010.88 ±0.01 Random Word (Rare)0.76 ±0.040.90 ±0.010.91 ±0.010.91 ±0.010.90 ±0.01 Paraphrasing0.86 ±0.030.95 ±0.010.95 ±0.010.95 ±0.010.95 ±0.01 Zeroshot (Qwen 2.5 72B)0.86 ±0.030.93 ±0.010.94 ±0.010.94 ±0.010.93 ±0.01 Zeroshot (DeepSeek-V4-Flash)0.85 ±0.030.93 ±0.010.94 ±0.010.93 ±0.010.92 ±0.01 ATO-Direct0.87 ±0.020.95 ±0.010.95 ±0.010.94 ±0.010.94 ±0.01 ATO-TwoPhase0.85 ±0.030.92 ±0.010.92 ±0.010.91 ±0.010.92 ±0.01 EN-IS Base text0.83 ±0.030.90 ±0.020.89 ±0.020.92 ±0.010.91 ±0.01 Random Subword0.58 ±0.040.77 ±0.030.78 ±0.020.44 ±0.040.79 ±0.02 Random Word (Common)0.51 ±0.030.75 ±0.020.72 ±0.030.79 ±0.020.76 ±0.02 Random Word (Rare)0.56 ±0.040.78 ±0.020.76 ±0.030.81 ±0.020.76 ±0.03 Paraphrasing0.74 ±0.040.86 ±0.020.85 ±0.020.89 ±0.020.88 ±0.02 Zeroshot (Qwen 2.5 72B)0.68 ±0.040.83 ±0.020.83 ±0.020.87 ±0.020.85 ±0.02 Zeroshot (DeepSeek-V4-Flash)0.67 ±0.040.82 ±0.020.83 ±0.020.85 ±0.020.83 ±0.02 ATO-Direct0.70 ±0.030.84 ±0.020.82 ±0.020.86 ±0.020.83 ±0.02 ATO-TwoPhase0.68 ±0.030.81 ±0.020.80 ±0.020.81 ±0.020.80 ±0.02 EN-RU Base text0.91 ±0.020.95 ±0.010.96 ±0.010.95 ±0.010.95 ±0.01 Random Subword0.70 ±0.030.84 ±0.020.85 ±0.020.86 ±0.020.85 ±0.02 Random Word (Common)0.67 ±0.030.82 ±0.020.83 ±0.020.82 ±0.020.81 ±0.02 Random Word (Rare)0.67 ±0.030.85 ±0.020.84 ±0.020.82 ±0.020.83 ±0.02 Paraphrasing0.80 ±0.040.92 ±0.020.93 ±0.010.92 ±0.010.91 ±0.02 Zeroshot (Qwen 2.5 72B)0.77 ±0.040.89 ±0.020.91 ±0.020.90 ±0.020.89 ±0.02 Zeroshot (DeepSeek-V4-Flash)0.76 ±0.030.89 ±0.020.90 ±0.020.89 ±0.020.88 ±0.02 ATO-Direct0.81 ±0.020.90 ±0.010.90 ±0.020.88 ±0.020.88 ±0.02 ATO-TwoPhase0.77 ±0.030.85 ±0.020.85 ±0.020.83 ±0.020.83 ±0.02 EN-ES Base text0.92 ±0.020.95 ±0.010.95 ±0.010.94 ±0.010.94 ±0.01 Random Subword0.71 ±0.030.86 ±0.020.86 ±0.020.84 ±0.020.85 ±0.02 Random Word (Common)0.67 ±0.030.82 ±0.020.83 ±0.020.79 ±0.020.82 ±0.02 Random Word (Rare)0.70 ±0.030.84 ±0.020.85 ±0.020.78 ±0.020.84 ±0.02 Paraphrasing0.86 ±0.030.92 ±0.010.93 ±0.010.92 ±0.010.91 ±0.01 Zeroshot (Qwen 2.5 72B)0.82 ±0.030.90 ±0.010.92 ±0.010.89 ±0.020.89 ±0.02 Zeroshot (DeepSeek-V4-Flash)0.81 ±0.030.89 ±0.020.90 ±0.010.89 ±0.020.88 ±0.02 ATO-Direct0.82 ±0.020.90 ±0.010.90 ±0.020.86 ±0.020.88 ±0.01 ATO-TwoPhase0.80 ±0.020.86 ±0.020.87 ±0.020.85 ±0.020.84 ±0.02 Table 7: Per-language, per-model xCOMET scores with 95% confidence intervals across all augmentation methods. 354 A.2MetricX NLLB-200-3.3BTower- Plus-72B TranslateGemma-27bGemini-3- Flash DeepSeek-V4- Flash EN-CS Base text4.43 ±0.522.91 ±0.262.43 ±0.202.72 ±0.242.95 ±0.26 Random Subword8.97 ±0.725.85 ±0.434.55 ±0.305.98 ±0.445.97 ±0.42 Random Word (Common)9.32 ±0.715.64 ±0.454.31 ±0.325.73 ±0.465.81 ±0.45 Random Word (Rare)9.11 ±0.745.95 ±0.434.72 ±0.345.94 ±0.456.21 ±0.44 Paraphrasing6.20 ±0.753.28 ±0.262.48 ±0.183.00 ±0.243.14 ±0.24 Zeroshot (Qwen 2.5 72B)7.05 ±0.803.59 ±0.302.77 ±0.213.16 ±0.243.61 ±0.31 Zeroshot (DeepSeek-V4-Flash)7.28 ±0.773.78 ±0.342.89 ±0.243.32 ±0.283.84 ±0.35 ATO-Direct6.46 ±0.594.15 ±0.313.21 ±0.244.30 ±0.354.26 ±0.35 ATO-TwoPhase7.53 ±0.665.21 ±0.423.95 ±0.355.78 ±0.505.68 ±0.47 EN-DE Base text1.65 ±0.430.67 ±0.100.55 ±0.090.75 ±0.150.84 ±0.16 Random Subword6.00 ±0.792.62 ±0.242.30 ±0.223.21 ±0.322.77 ±0.26 Random Word (Common)6.18 ±0.822.50 ±0.271.87 ±0.182.66 ±0.282.70 ±0.25 Random Word (Rare)6.02 ±0.822.84 ±0.262.32 ±0.212.86 ±0.252.86 ±0.24 Paraphrasing3.17 ±0.670.91 ±0.130.78 ±0.110.90 ±0.130.97 ±0.14 Zeroshot (Qwen 2.5 72B)3.10 ±0.621.08 ±0.130.87 ±0.121.01 ±0.121.13 ±0.14 Zeroshot (DeepSeek-V4-Flash)3.21 ±0.581.15 ±0.150.88 ±0.121.11 ±0.151.27 ±0.16 ATO-Direct3.32 ±0.601.32 ±0.151.10 ±0.141.61 ±0.191.64 ±0.20 ATO-TwoPhase3.75 ±0.602.08 ±0.271.42 ±0.192.40 ±0.332.29 ±0.30 EN-IS Base text5.20 ±0.553.49 ±0.373.16 ±0.322.75 ±0.283.44 ±0.35 Random Subword10.19 ±0.725.94 ±0.465.26 ±0.385.52 ±0.436.13 ±0.45 Random Word (Common)10.67 ±0.745.38 ±0.384.94 ±0.385.22 ±0.406.06 ±0.46 Random Word (Rare)10.52 ±0.745.87 ±0.435.15 ±0.385.53 ±0.426.12 ±0.44 Paraphrasing6.79 ±0.764.00 ±0.373.67 ±0.333.04 ±0.273.51 ±0.31 Zeroshot (Qwen 2.5 72B)7.82 ±0.773.88 ±0.333.75 ±0.333.16 ±0.284.13 ±0.39 Zeroshot (DeepSeek-V4-Flash)8.02 ±0.774.12 ±0.383.82 ±0.323.35 ±0.294.00 ±0.35 ATO-Direct7.29 ±0.584.74 ±0.404.27 ±0.384.55 ±0.405.06 ±0.42 ATO-TwoPhase8.50 ±0.675.55 ±0.484.61 ±0.405.75 ±0.526.51 ±0.56 EN-RU Base text2.66 ±0.471.33 ±0.200.98 ±0.161.29 ±0.191.44 ±0.23 Random Subword7.27 ±0.723.68 ±0.332.87 ±0.273.88 ±0.344.03 ±0.35 Random Word (Common)7.51 ±0.713.44 ±0.342.63 ±0.273.93 ±0.424.09 ±0.41 Random Word (Rare)7.28 ±0.693.72 ±0.312.92 ±0.254.95 ±0.464.21 ±0.38 Paraphrasing4.82 ±0.791.74 ±0.231.32 ±0.181.63 ±0.201.75 ±0.21 Zeroshot (Qwen 2.5 72B)5.30 ±0.781.92 ±0.231.49 ±0.191.77 ±0.212.12 ±0.24 Zeroshot (DeepSeek-V4-Flash)5.45 ±0.772.01 ±0.281.53 ±0.201.95 ±0.252.11 ±0.26 ATO-Direct4.13 ±0.482.22 ±0.281.77 ±0.232.57 ±0.312.54 ±0.29 ATO-TwoPhase5.09 ±0.543.25 ±0.372.40 ±0.293.78 ±0.393.83 ±0.39 EN-ES Base text2.76 ±0.441.78 ±0.171.48 ±0.151.77 ±0.161.90 ±0.18 Random Subword7.49 ±0.704.18 ±0.323.59 ±0.274.57 ±0.364.35 ±0.32 Random Word (Common)7.75 ±0.704.01 ±0.323.32 ±0.284.71 ±0.414.31 ±0.35 Random Word (Rare)7.62 ±0.724.31 ±0.333.68 ±0.276.71 ±0.484.74 ±0.36 Paraphrasing3.71 ±0.581.99 ±0.181.72 ±0.142.00 ±0.172.10 ±0.20 Zeroshot (Qwen 2.5 72B)4.13 ±0.602.24 ±0.181.85 ±0.162.26 ±0.202.38 ±0.20 Zeroshot (DeepSeek-V4-Flash)4.49 ±0.602.39 ±0.241.97 ±0.182.33 ±0.222.53 ±0.27 ATO-Direct4.59 ±0.492.84 ±0.252.19 ±0.193.77 ±0.313.35 ±0.30 ATO-TwoPhase5.16 ±0.523.91 ±0.383.02 ±0.354.46 ±0.444.26 ±0.40 Table 8: Per-language, per-model MetricX scores with 95% confidence intervals across all augmentation methods. A.3Data Metrics A potential concern is vocabulary collapse: the optimizer might converge on a small set of tokens that Sentinel considers difficult and insert them into every sequence. To test for this, we measure word diversity: the number of unique words across all sequences, normalized by total word count. We also report average word length (in characters) and average word count (by whitespace). Results are shown 355 in Table 9. Word count increases steadily across methods, consistent with the intuition that longer texts are harder to translate. Word diversity remains stable, indicating that the optimizer does not degenerate into repeated insertion of a fixed token set. Word length shows no clear trend. MethodWord CountAvg Word LengthWord Diversity Base text17.25 ±1.595.17 ±0.140.588 Random Subword17.00 ±1.575.43 ±0.160.649 Random Word (Common)17.25 ±1.595.46 ±0.150.647 Random Word (Rare)17.25 ±1.595.24 ±0.130.658 Paraphrasing17.95 ±1.685.01 ±0.140.556 Zeroshot (Qwen 2.5 72B)17.25 ±1.595.23 ±0.140.605 Zeroshot (DeepSeek-V4-Flash)17.45 ±1.615.13 ±0.130.598 ATO-Direct17.26 ±1.595.29 ±0.150.612 ATO-TwoPhase17.15 ±1.585.26 ±0.170.604 Table 9: Analysis of word quality metrics including length and diversity across the different generation methods. BModel Prompts & Details B.1Translation Prompt To use Gemini-3-Flash and DeepSeek-V4-Flash for translation, we used the prompt stated in Figure 3. The strings for translation have been batched by 10 to increase throughput. The model was accessed through the official Google API under the name models/gemini-3-flash-preview on 19-03-2026. DeepSeek-V4-Flash has been used through OpenRouter API under the name deepseek/deepseek-v4- flash on 06-05-2026. You are a professional translator. Translate the following list of strings into the target language. Maintain the exact order of the list. Return ONLY a valid Python-style list of strings with the translations. Do not include explanations, code blocks, or markdown. CRITICAL: Do not skip any sequences even if they are difficult to translate! Source language: SRCLANGUAGE Target language: TGTLANGUAGE Figure 3: Translation Prompt B.2Zeroshot Baseline Prompt To use Qwen 2.5 72B and DeepSeek-V4-Flash as zero-shot replacement baselines, we used the prompt in Figure 4. DeepSeek-V4-Flash was accessed through the OpenRouter API as deepseek/deepseek-v4- flash on 06-05-2026. You are a linguistics expert specializing in translation difficulty. I will give you an English text. Pick EXACTLY num_replacements word(s) to replace with alternatives that make the text harder to translate into German. Return your answer as a JSON object with exactly num_replacements entries: "1": "original": "word1", "replacement": "new_word1", "2": "original": "word2", "replacement": "new_word2" Rules: •Pick exactly num_replacements word(s) to replace – no more, no less. •Each replacement must be a single word. •Choose words that are ambiguous, idiomatic, or culturally specific. •Do NOT use German words. •Return ONLY the JSON object, nothing else. Text: text Figure 4: Zeroshot Prompt B.3DIPPER Paraphrasing Baseline Details DIPPER takes a text as input and outputs a rephrased version of that text. It also accepts a lexical diversity parameter controlling how aggressively words are changed and an order diversity parameter controlling word reordering. We fixed order diversity to zero, since our optimization does not rearrange words, and 356 selected a lexical diversity of 20 by matching the edit distance distribution of DIPPER’s output to that of our two-phase method across the full range (0–100). B.4Random Replacement Baseline Details We replace either whole words from the vocabulary in Section 3.3 or subwords from Section 3.4, replacing exactly 2 words or subwords per seed sentence. For whole-word tokens we distinguish between common and rare using Zipf word frequency from the wordfreq library (Speer, 2022). We require Zipf frequency greater than 2.5 to filter out extremely rare words, then split the remainder at the median into common (Zipf ≥4.09) and rare (2.5< Zipf <4.09) buckets. CHuman Evaluation C.1Grammaticality and Plausibility Rate the text below on two dimensions: Grammaticality (1–5): How grammatically well-formed is the text? 1 = Completely ungrammatical — severe errors that make the text very hard/impossible to understand. 2 = Major grammatical problems — understandable, but errors clearly disrupt structure or fluency. 3 = Noticeable issues — errors present but do not seriously impact comprehension. 4 = Minor imperfections — almost fully grammatical, only very small issues (e.g. a typo). 5 = Fully grammatical — no noticeable errors; the text is well-formed. Plausibility (1–5): How likely is it that this text would appear in real online content? 1 = Implausible — this text would not appear in any realistic context. 2 = Unlikely — would rarely occur, even though it is not impossible. 3 = Moderately plausible — could appear, but would be somewhat unusual. 4 = Quite plausible — would commonly make sense in online content. 5 = Highly plausible — would very naturally appear as part of real online content.“ Figure 5: Instructions for Rating Grammaticality and Plausibility C.2Translation Quality Read the original English text and its machine translation carefully. Translation Quality (1–5): How well does the translation convey the original meaning? 1 = Poor — the translation is inaccurate, unintelligible, or fails to convey the original meaning. 2 = Fair — significant errors or awkward phrasing that affect clarity. 3 = Good — mostly accurate and understandable, with minor issues. 4 = Very Good — accurate and fluent, with only negligible errors. 5 = Excellent — accurate, fluent, and natural-sounding. Figure 6: Instructions for Rating Translation Quality C.3Inter-Annotator Agreement Metric 훼 absolute훼 relative Grammaticality0.180.41 Plausibility0.280.41 Translation quality0.790.51 Table 10: Krippendorff’s 훼 on human evaluation metrics. For relative 훼, each rater’s ratings are first 푧-normalised within rater (centered on that rater’s mean, scaled by their standard deviation), and Krippendorff’s 훼 is then computed at the interval level on the resulting (rater × item) matrix. Although the absolute 훼 varies across metrics (0.18–0.79), the relative 훼 range of 0.41–0.51 suggests per-rater scale bias causes absolute variation rather than disagreement on the relative ordering of items. 357 C.4Evaluation interface We use Pearmut (Zouhar and Kocmi, 2026), an open-source translation evaluation tool. The evaluations are thus reproducible given the data to be evaluated. Figure 7: Translation quality evaluation interface. Pictured language is Icelandic, and we had a separate interface for each of the other languages. Figure 8: English text quality evaluation interface. 358