Paper deep dive
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
Hongbin Zhang, Junhao Liu, Xuefeng Bai, Youcheng Pan, Yang Xiang, Kehai Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/1/2026, 1:52:23 AM
Summary
The paper introduces DualAnchor, a gloss-free training framework for Sign Language Translation (SLT) that addresses language-prior degradation and lexical fidelity gaps in LLM-based methods. It employs Token-level Prior Anchoring (TPA) to preserve linguistic fluency by regularizing next-token distributions against a frozen LLM, and Optimal Transport Alignment (OTA) to improve lexical accuracy by aligning visual tokens with textual content tokens via entropy-regularized partial optimal transport. DualAnchor achieves state-of-the-art performance on PHOENIX-2014T and CSL-Daily benchmarks.
Entities (10)
Relation Signals (8)
DualAnchor → usescomponent → Token-level Prior Anchoring
confidence 95% · DualAnchor... couples two complementary anchors... (i) Token-level Prior Anchoring (TPA)
DualAnchor → usescomponent → Optimal Transport Alignment
confidence 95% · DualAnchor... couples two complementary anchors... (ii) Optimal Transport Alignment (OTA)
DualAnchor → evaluatedon → PHOENIX-2014T
confidence 92% · DualAnchor achieves strong overall performance on both PHOENIX-2014T
DualAnchor → evaluatedon → CSL-Daily
confidence 92% · DualAnchor achieves strong overall performance on both... CSL-Daily
Token-level Prior Anchoring → addresses → language-prior degradation
confidence 90% · TPA preserves the LLM's language prior... TPA improves fluency
Optimal Transport Alignment → addresses → lexical fidelity gap
confidence 90% · OTA improves lexical fidelity... OTA reduces fine-grained lexical errors
DualAnchor → solvesproblem → Sign Language Translation
confidence 85% · DualAnchor, a gloss-free LLM-based SLT training framework
Sinkhorn optimization → usedin → Optimal Transport Alignment
confidence 85% · with Sinkhorn optimization inducing a soft alignment... under a cosine cost
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.
Tags
Links
- Source: https://arxiv.org/abs/2607.27614v1
- Canonical: https://arxiv.org/abs/2607.27614v1
Trouble viewing inline? Open PDF directly →
Full Text
49,928 characters extracted from source content.
Expand or collapse full text
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation Hongbin Zhang1,2 , Junhao Liu1 , Xuefeng Bai1, Youcheng Pan2, Yang Xiang2, Kehai Chen1,2 Abstract Recent advances in large language models (LLMs) have led sign language translation (SLT)—the task of converting sign-language videos into spoken-language text—to increasingly adopt LLMs as textual backbones. However, despite LLMs’ strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations—a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level; such alignment does not ensure accurate lexical details, creating a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation: (i) Token-level Prior Anchoring (TPA) preserves the LLM’s language prior by regularizing the multimodal decoder, at each decoding step, toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. (i) Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual–textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses link the gains to TPA’s fluency improvement and OTA’s reduction of fine-grained lexical errors. Introduction Sign Language Translation (SLT) translates sign-language videos into spoken-language text by recovering linguistic content from continuous visual signals and expressing it in target-language grammar (Camgoz et al. 2018; Camgöz et al. 2020; Yin and Read 2020). Recent advances in LLMs have led gloss-free SLT methods to adopt a generation framework in which the model encodes the sign video, maps the resulting visual representation into the LLM’s language space, and uses the pretrained LLM as the textual backbone for translation (Wong et al. 2024; Gong et al. 2024; Chen et al. 2024; Hwang et al. 2025). Within this framework, cross-modal alignment is usually learned at the sentence level through contrastive objectives over paired video and text representations or transferred semantic similarity (Zhou et al. 2023; Yin et al. 2023; Jiang et al. 2024). Figure 1: Motivating examples of the two issues in LLM-based SLT. Top: language-prior degradation produces disfluent text despite an LLM decoder. Bottom: strong global alignment can still coexist with weak lexical grounding, illustrating the lexical fidelity gap. However, this paradigm leaves two complementary issues unresolved. Despite the strong language modeling capabilities of their LLM backbones, existing methods often undermine rather than exploit the pretrained language prior during visual adaptation. Multimodal adaptation can shift the decoder’s next-token distribution away from the linguistic structure acquired during text pretraining and reduce output fluency; we term this issue language-prior degradation. Sentence-level objectives align an entire sign video with its paired sentence, but they do not directly tie individual content words to local sign evidence. A model can therefore achieve strong global video–text alignment while mistranslating actions, entities, or attributes; we call this discrepancy the lexical fidelity gap. The two issues affect complementary dimensions of translation quality: language-prior degradation weakens linguistic form, while the lexical fidelity gap compromises fine-grained content accuracy. Figure 1 provides an intuitive illustration. The upper output is ungrammatical despite its LLM decoder. The lower example shows strong global semantic alignment, yet replaces refrigerator and apple with visually unsupported cabinet and orange. To address these issues, we propose DualAnchor, a gloss-free LLM-based SLT framework with two anchors: token-level prior anchoring and optimal transport alignment. Token-level Prior Anchoring (TPA) works in distribution space. Under the same autoregressive prefix, it regularizes the multimodal next-token distribution toward a frozen LLM prior and scales the constraint with the prior’s confidence. This design preserves pretrained decoding behavior while still letting visual evidence influence uncertain positions. Optimal Transport Alignment (OTA) works in representation space. Entropy-regularized partial transport matches visual tokens to content tokens, while an unmatched sink absorbs unreliable correspondences. The partial formulation handles visual–lexical granularity mismatch without forcing every token to align. TPA preserves linguistic form, while OTA grounds lexical choice in fine-grained visual evidence. We evaluate DualAnchor on PHOENIX-2014T (Camgoz et al. 2018) and CSL-Daily (Zhou et al. 2021). Across the two benchmarks, DualAnchor attains the best BLEU-4 among the compared gloss-free methods and the highest arithmetic mean across the ten reported metrics. Targeted analyses associate the gains with more fluent generation and fewer fine-grained lexical errors, supporting the intended roles of TPA and OTA over generic model capacity. Our contributions are threefold: • We identify two complementary LLM-based SLT issues: language-prior degradation, which hurts fluency during multimodal adaptation, and the lexical fidelity gap, where sentence-level alignment still leaves content-word errors. • We introduce DualAnchor, coupling confidence-aware token-level prior anchoring with entropy-regularized partial transport to preserve the pretrained language prior and ground lexical choices in visual evidence. • On PHOENIX-2014T and CSL-Daily, DualAnchor achieves the best BLEU-4 among the compared gloss-free methods, and targeted analyses show better fluency and fewer fine-grained lexical errors. (a) CSL-Daily: fluency gap (b) CSL-Daily: prior divergence (c) PHOENIX-2014T: global alignment vs. lexical fidelity Figure 2: Preliminary diagnostics of prior drift and lexical fidelity. (a) On CSL-Daily, 81.9%81.9\% of the adapted baseline’s hypotheses have higher frozen-LM PPL than their paired references. (b) The student-to-prior KL divergence increases during multimodal adaptation and remains above its initial level. (c) On PHOENIX-2014T, 31.7%31.7\% of baseline hypotheses in the highest video–reference alignment quartile have IDF-weighted content-word F1 below 50%50\%, showing that high global compatibility can coexist with lexical translation errors. Preliminary Analysis To characterize the two issues that motivate our method, we conduct two preliminary diagnostics. We first examine whether multimodal adaptation is accompanied by an output-level fluency gap and increased divergence from the pretrained language prior. We next ask whether high global video–reference compatibility can coexist with lexical errors. Diagnostic I: Does multimodal adaptation coincide with an output-level fluency gap and greater prior divergence? Given a sign-language video Vi=vi,tt=1TiV_i=\v_i,t\_t=1^T_i and its spoken-language reference Yi=yi,nn=1NiY_i=\y_i,n\_n=1^N_i, the controlled baseline generates Y^i∼pθ(Y∣Vi) Y_i p_θ(Y V_i) autoregressively. We evaluate its hypotheses on the held-out sets of CSL-Daily (Chinese) (Zhou et al. 2021) and PHOENIX-2014T (German) (Camgoz et al. 2018). We measure fluency with a frozen external language model qϕq_φ, using Baichuan2-7B (Yang et al. 2023) for Chinese and Qwen2.5-7B (Yang et al. 2024) for German. For a tokenized sentence S=snn=1|S|S=\s_n\_n=1^|S|, we compute its perplexity and paired log-space fluency gap as PPLϕ(S) _φ(S) =exp(−1|S|∑n=1|S|logqϕ(sn∣s<n)), = \! (- 1|S| _n=1^|S| q_φ(s_n s_<n) ), (1) Δiflu _i^flu =logPPLϕ(Y^i)−logPPLϕ(Yi). = _φ( Y_i)- _φ(Y_i). (2) Lower PPL indicates higher target-language likelihood, while Δiflu>0 _i^flu>0 means that the hypothesis is less probable than its paired reference under the same evaluator. Figure 2(a) reveals a pronounced output-level fluency gap on CSL-Daily: 81.9%81.9\% of the adapted baseline’s hypotheses have higher PPL than their paired references, and the mean PPL differs by 3.4×3.4× (335.8335.8 versus 98.398.3). We next examine whether this output-level fluency gap co-occurs with increased divergence from the pretrained language prior by tracing the student-to-prior KL divergence throughout multimodal training. At epoch e, let pe,i,nsp^s_e,i,n denote the baseline’s next-token distribution and pi,npp^p_i,n the frozen prior distribution under the same gold autoregressive prefix. We average reverse KL over all valid target positions Ω : Deprior=1|Ω|∑(i,n)∈ΩDKL(pe,i,ns∥pi,np).D_e^prior= 1| | _(i,n)∈ D_KL\! (p^s_e,i,n\,\|\,p^p_i,n ). (3) Larger values indicate greater deviation from the frozen language prior under synchronized decoding contexts. Figure 2(b) shows that the KL divergence increases from 2.502.50 to a peak of 2.772.77 near epoch 65 and stabilizes around 2.632.63 after epoch 220, remaining above its initial level. Finding I. Multimodal adaptation is associated with an output-level fluency gap and increased prior divergence: the adapted baseline’s hypotheses receive substantially higher external-LM PPL than their paired references, while student-to-prior KL increases during training and remains above its initial level. This coupled pattern motivates preserving the pretrained language prior during visual adaptation. Diagnostic I: Can high global video–reference compatibility coexist with lexical translation errors? Using the same baseline outputs and held-out samples as in Diagnostic I, we examine whether gold video–reference pairs that are well matched at the sentence level can still have corresponding hypotheses with weak content-word fidelity. Let gvg_v and gtg_t denote the video and text encoders used by the sentence-level retrieval diagnostic. For each gold video–reference pair, we measure global compatibility as sisent=sim(gv(Vi),gt(Yi))s_i^sent=sim\! (g_v(V_i),g_t(Y_i) ), where larger values indicate stronger sentence-level compatibility between the video and its reference translation in the diagnostic representation space. To measure fine-grained lexical fidelity, let Kih=K(Y^i)K_i^h=K( Y_i) and Kir=K(Yi)K_i^r=K(Y_i) denote the normalized content-word sets of the hypothesis and reference, respectively, and define W(A)=∑w∈Aidf(w)W(A)= _w∈ Aidf(w). The IDF-weighted content-word F1 is F1,iidf=2W(Kih∩Kir)W(Kih)+W(Kir).F_1,i^idf= 2W(K_i^h∩ K_i^r)W(K_i^h)+W(K_i^r). (4) The resulting score measures content-word agreement between the baseline hypothesis and its reference. IDF assigns greater weight to rare and informative words, while F1 captures both missing reference content and unmatched content introduced by the hypothesis. Content-word-aware translation likewise treats lexical units as unequally important to sentence meaning (Chen et al. 2020a). Figure 2(c) shows that lexical errors remain common even among globally well-matched gold video–reference pairs: within the highest alignment quartile, 31.7%31.7\% of baseline hypotheses have IDF-weighted content-word F1 below 50%50\%. Thus, fine-grained lexical errors can persist despite high global video–reference compatibility. Finding I. High global video–reference compatibility can coexist with substantial lexical translation errors: even within the highest alignment quartile, 31.7%31.7\% of baseline hypotheses have IDF-weighted content-word F1 below 50%50\%. This coexistence motivates visual–textual alignment at a finer lexical granularity than sentence-level matching. Methodology Problem Formulation Given a sign video V=vℓ=1TV=\v_ \_ =1^T and its spoken-language translation Y=ytt=1NY=\y_t\_t=1^N, an LLM-based SLT model maps V to decoder-compatible visual tokens, v=Fψ(V)=ivi=1n,iv∈ℝd,Z^v=F_ψ(V)=\z^v_i\_i=1^n, ^v_i ^d, (5) where FψF_ψ comprises the backbone’s visual encoder and modality projection. The multimodal student factorizes the translation probability autoregressively as pθs(Y∣V)=∏t=1Npθs(yt∣v,y<t),p_ _s(Y V)= _t=1^Np_ _s\! (y_t ^v,y_<t ), (6) and learns from the teacher-forced translation loss ℒCE=−1|Ω|∑(b,t)∈Ωlogpθs(yb,t∣bv,yb,<t),L_CE=- 1| | _(b,t)∈ p_ _s\! (y_b,t ^v_b,y_b,<t ), (7) where Ω contains the non-padding target positions in a minibatch. Cross-entropy supervises the observed target but does not explicitly preserve the pretrained trajectory of next-token distributions or impose local geometry between visual evidence and lexical units. DualAnchor therefore adds two complementary training constraints (Figure 3): TPA anchors linguistic form in distribution space, and OTA anchors lexical content in representation space. Together, TPA and OTA target a translation’s linguistic form and visual content. Figure 3: Overview of DualAnchor. TPA anchors the student next-token distribution to a frozen language prior under the same autoregressive prefix, while OTA learns partial correspondences between visual tokens and textual content tokens. Both modules operate only during training. Token-Level Prior Anchoring Knowledge distillation transfers teacher behavior to sequence models and multilingual translation systems (Kim and Rush 2016; Sun et al. 2020). TPA adapts this principle to token-level preservation during multimodal training. Prefix-synchronized language prior. TPA retains a frozen copy of the language model, pθp_ _p, to preserve the generative structure acquired during pretraining. Let sbs_b be a context prompt and πb,t=[sb;yb,<t] _b,t=[s_b;y_b,<t] the context-augmented gold autoregressive prefix. The student and prior share πb,t _b,t and the vocabulary; only the student observes the sign video: pb,ts(k) p^s_b,t(k) =pθs(yt=k∣bv,πb,t), =p_ _s\! (y_t=k ^v_b, _b,t ), (8) pb,tp(k) p^p_b,t(k) =pθp(yt=k∣πb,t). =p_ _p\! (y_t=k _b,t ). (9) English Context prepends an English task instruction; Target Context, used by default, prepends its target-language rendering. In both settings, sbs_b is fixed before decoding and the target-dependent part of πb,t _b,t is exactly yb,<ty_b,<t; the prefix-only variant sets sb=∅s_b= . We freeze θp _p and detach its probabilities from gradient computation. The synchronized prefix isolates multimodal adaptation’s effect at each next-token position. Confidence-adaptive distribution anchoring. The frozen prior is not equally reliable at every position: strict imitation at uncertain positions can suppress useful visual evidence, whereas a sharp prior provides a stronger linguistic constraint. TPA measures this reliability with token entropy Hb,t H_b,t =−∑k∈pb,tp(k)logpb,tp(k), =- _k p^p_b,t(k) p^p_b,t(k), (10) H¯ H =1|Ω|∑(b,t)∈ΩHb,t, = 1| | _(b,t)∈ H_b,t, where V is the target vocabulary. TPA then applies the confidence gate ωb,t=λ0σ(γ(H¯−Hb,t)) _b,t= _0σ\! (γ( H-H_b,t) ), where λ0 _0 sets the maximum anchoring scale, γ controls gate sharpness, and σ(⋅)σ(·) is the sigmoid function. Low-entropy positions receive stronger prior guidance; high-entropy positions leave the student more freedom to follow the video. TPA averages reverse KL over valid target positions: ℒTPA=1|Ω|∑(b,t)∈Ωωb,tDKL(pb,ts∥pb,tp).L_TPA= 1| | _(b,t)∈ _b,tD_KL\! (p^s_b,t\,\|\,p^p_b,t ). (11) This direction penalizes probability mass that the adapted student assigns to tokens unsupported by the frozen prior. TPA regularizes the predictive distribution rather than prescribing a decoded sentence, preserving linguistic structure while retaining visual conditioning. Optimal Transport Alignment Content-focused cross-modal geometry. Sentence-level alignment captures global video-text correspondence but leaves individual actions, entities, attributes, and temporal expressions ungrounded. OTA addresses this gap by aligning the visual tokens in Equation (5) with content-token representations c=jcj=1mZ^c=\z^c_j\_j=1^m, where jc∈ℝdz^c_j ^d. The content selector excludes padding, special symbols, punctuation, and weakly semantic function tokens so that non-lexical units do not dominate the alignment. After ℓ2 _2 normalization, the pairwise cosine cost is Dij=1−(iv)⊤jc‖iv‖2‖jc‖2,∈ℝn×m.D_ij=1- (z^v_i) z^c_j\|z^v_i\|_2\,\|z^c_j\|_2, ^n× m. (12) The soft cost retains plausible correspondences because sign duration and text tokenization differ in granularity. Entropy-regularized partial transport. Entropy regularization makes optimal transport differentiable and efficiently solvable by Sinkhorn scaling (Cuturi 2013), while partial transport permits unmatched mass when the two supports contain unreliable correspondences (Chapel et al. 2020). Let =1nna= 1n1_n and =1mmb= 1m1_m denote uniform visual and textual masses, and let ρ∈(0,1]ρ∈(0,1] be the mass reserved for genuine cross-modal matching. We define the partial transport polytope Πρ(,)=∈ℝ+n×m: _ρ(a,b)= \P _+^n× m: m≤,⊤n≤, 1_m ,~P 1_n , (13) n⊤m=ρ. 1_n P1_m=ρ \. Introduce slack masses v=−ms^v=a-P1_m and c=−⊤ns^c=b-P 1_n, and set ¯=[v(c)⊤0] P= [ smallmatrixP&s^v\\ (s^c) &0 smallmatrix ], ¯=[;1−ρ] a=[a;1-ρ], and ¯=[;1−ρ] b=[b;1-ρ]. Let Π¯ρ(¯,¯) _ρ( a, b) contain balanced couplings with these marginals and P¯n+1,m+1=0 P_n+1,m+1=0. With ¯ D equal to D on the real block and constant on dustbin edges, we solve ¯⋆=argmin¯∈Π¯ρ(¯,¯)⟨¯,¯⟩−τℋ(¯), P = _ P∈ _ρ( a, b) P, D -τ\,H( P), (14) where τ>0τ>0 controls transport smoothness and ℋ(¯)=−∑i,jP¯ij(logP¯ij−1)H( P)=- _i,j P_ij( P_ij-1), with 0log0=00 0=0. Thus ℋ(¯)=ℋ()+ℋ(v)+ℋ(c)H( P)=H(P)+H(s^v)+H(s^c), explicitly regularizing unmatched capacity, while the real block transports exactly ρ mass. Let M be one except for Mn+1,m+1=0M_n+1,m+1=0; masked Sinkhorn scaling gives =⊙exp(−¯/τ),¯⋆=diag()diag(),K=M (- D/τ), P =diag(u)\,K\,diag(r), (15) where Sinkhorn iterations alternately rescale u and r to match ¯ a and ¯ b. The real block of ¯⋆ P gives ⋆P , while the dustbins absorb the remaining 1−ρ1-ρ mass. OTA then minimizes the average cost carried by genuine matches: ℒOTA=⟨⋆,⟩∑i=1n∑j=1mPij⋆.L_OTA= ,D _i=1^n _j=1^mP _ij. (16) The transport plan supplies token-level structure, while the dustbin prevents ambiguous or semantically empty units from creating spurious supervision. Joint Optimization and Inference The two anchors regularize complementary model objects and are combined with translation supervision as ℒ=ℒCE+αTPAℒTPA+αOTAℒOTA,L=L_CE+ _TPAL_TPA+ _OTAL_OTA, (17) where αTPA _TPA and αOTA _OTA balance linguistic stability and visual grounding. At inference, the frozen prior, content selector, and Sinkhorn solver are removed, and the student follows the original autoregressive decoding path. DualAnchor therefore changes training supervision without adding test-time teacher access or optimal-transport computation. Experiments PHOENIX-2014T Method B1 B2 B3 B4 R-L GFSLT-VLP 43.71 33.18 26.11 21.44 42.49 FLa-LLM 46.29 35.33 28.03 23.09 45.27 Sign2GPT 49.54 35.96 28.83 22.52 48.90 SignLLM 45.21 34.78 28.05 23.40 44.49 BeyondGloss 52.38 38.57 30.74 25.49 52.89 MMSLT 48.92 38.12 30.79 25.73 47.97 SpaMo 49.80 37.32 29.50 24.32 46.57 SCL-SLT 48.72 38.19 31.04 26.00 47.02 DualAnchor (ours) 53.93 41.28 33.18 27.60 50.62 CSL-Daily Method B1 B2 B3 B4 R-L GFSLT-VLP 39.37 24.93 16.26 11.00 36.44 FLa-LLM 37.13 25.12 18.38 14.20 37.25 Sign2GPT 41.75 28.73 20.60 15.40 42.36 SignLLM 39.55 28.13 20.07 15.75 39.91 BeyondGloss 53.12 38.63 27.82 21.53 53.46 MMSLT 49.87 36.37 27.29 21.11 48.92 SpaMo 48.90 36.90 26.78 20.55 47.46 SCL-SLT 52.81 39.28 29.82 23.25 51.08 DualAnchor (ours) 53.15 39.84 30.65 24.21 49.46 Table 1: Gloss-free SLT results on the official test sets. Bnn and R-L denote BLEU-n and ROUGE-L; bold and underlining mark the best and second-best scores. Datasets and metrics. We evaluate on PHOENIX-2014T and CSL-Daily. PHOENIX-2014T contains 8,257 German Sign Language weather videos with German translations and uses the standard 7,096/519/642 train/development/test split (Camgoz et al. 2018); CSL-Daily contains 20,654 Chinese Sign Language videos on daily-life topics with Chinese translations and uses the standard 18,401/1,077/1,176 split (Zhou et al. 2021). We report BLEU-1–BLEU-4 (Papineni et al. 2002) and ROUGE-L (Lin 2004). Baselines. We compare eight published gloss-free systems spanning visual-language pretraining (GFSLT-VLP (Zhou et al. 2023)), LLM adaptation (FLa-LLM (Chen et al. 2024), Sign2GPT (Wong et al. 2024), SignLLM (Gong et al. 2024), and SpaMo (Hwang et al. 2025)), enriched visual or MLLM supervision (BeyondGloss (Asasi et al. 2025) and MMSLT (Kim et al. 2025)), and contrastive alignment (SCL-SLT (Lai et al. 2026)). Implementation details. We fine-tune the language backbone with LoRA (Hu et al. 2022) using AdamW (Loshchilov and Hutter 2019). Unless otherwise specified, experiments use 4 GPUs, at most 500 epochs, and an early-stopping patience of 50 validation rounds. We use a learning rate of 3×10−43× 10^-4 with cosine decay, select checkpoints by the best validation BLEU-4, and decode with beam size 6 and length penalty 1.0; the auxiliary objectives are training-only. Main Results. DualAnchor outperforms SpaMo on all ten metrics and gives the strongest higher-order BLEU results on both datasets (Table 1). BLEU-4 rises from 24.32 to 27.60 on PHOENIX-2014T and from 20.55 to 24.21 on CSL-Daily, while the remaining BLEU and ROUGE-L scores improve consistently. Across all compared methods, DualAnchor achieves the highest arithmetic mean over the ten reported metrics. Gains span sign languages and domains and are strongest in multi-token correspondence. Backbone Method B4 R-L NLLB SpaMo 17.89 43.29 MMSLT 17.96 41.17 DualAnchor (ours) 24.21 49.46 mBART SpaMo 19.70 46.17 MMSLT 20.21 50.58 DualAnchor (ours) 22.35 49.75 mT0 SpaMo 19.59 42.12 MMSLT 19.92 43.32 DualAnchor (ours) 22.56 47.75 Table 2: CSL-Daily cross-backbone results. Bold marks the best completed score for each backbone. Rows share cached visual features, decoding settings, and the evaluation script. Cross-backbone generalization. Table 2 evaluates TPA and OTA with NLLB (NLLB Team 2024), mBART (Liu et al. 2020), and mT0 (Muennighoff et al. 2023) under the same protocol. Across completed runs, DualAnchor leads BLEU-4 on all three backbones and improves both metrics over SpaMo in every matched comparison. The gains are largest with NLLB and remain positive on the other two backbones. On mBART, MMSLT leads ROUGE-L but DualAnchor leads BLEU-4, making the ranking metric-dependent. Analysis This section asks four questions: (i) RQ1: How does TPA affect fluency and linguistic structure, and do these effects align with prior preservation? (i) RQ2: Does OTA improve fine-grained lexical grounding and use local visual evidence as intended? (i) RQ3: How sensitive are TPA and OTA to their loss weights? (iv) RQ4: Which TPA and OTA design choices govern translation quality and module behavior? RQ1: How does TPA affect fluency and linguistic structure, and do these effects align with prior preservation? Using SpaMo on CSL-Daily, we compare otherwise identical models with and without TPA to test how prior preservation relates to fluency and linguistic form. PPL and KL dynamics. At matched checkpoints, lower frozen-Baichuan2-7B corpus PPL indicates better external fluency, while lower mean student–prior KL under shared gold prefixes indicates less drift from the frozen prior. (a) External PPL (b) Student–prior KL Figure 4: CSL-Daily TPA diagnostics: (a) frozen-Baichuan2-7B PPL; (b) mean student–prior KL. Across training, lower PPL accompanies lower student–prior KL with TPA (Figure 4). Their agreement links the external fluency gain to reduced prior drift. Syntactic and grammatical analysis. Table 3 tests whether the pattern extends to linguistic structure. Stanza parses omit punctuation and the root relation. Dependency JSD compares hypothesis and reference relation histograms; its 0.30.3/0.50.5/0.70.7 exceedance rates capture mismatch severity, while normalized dependency edit divides relation-sequence Levenshtein distance by the longer sequence. Grammar metrics cover function-marker omissions per 100 sentences, question-particle precision, and structural-particle recall. Syntactic metrics (baseline vs. TPA) Metric Base TPA Δ Mean dependency JSD ↓ 0.4337 0.3955 ↓ 8.81% JSD >0.3>0.3 ↓ 73.89% 66.07% ↓ 7.82 p JSD >0.5>0.5 ↓ 35.12% 28.66% ↓ 6.46 p JSD >0.7>0.7 ↓ 11.48% 7.14% ↓ 4.34 p Dependency edit ↓ 0.6551 0.6139 ↓ 6.29% Grammatical metrics (baseline vs. TPA) Metric Base TPA Δ Function-marker omissions ↓ 62.1 56.9 ↓ 8.37% Question-particle precision ↑ 39.6% 46.8% ↑ 7.2 p Structural-particle recall ↑ 50.9% 53.7% ↑ 2.8 p Table 3: CSL-Daily TPA linguistic diagnostics (epoch 200; omissions per 100 sentences; p: percentage points). The distributional, sequence, and marker-level diagnostics in Table 3 agree: TPA reduces dependency mismatch and improves particle use. Together with Figure 4, they answer RQ1: prior preservation coincides with better-formed output in the matched comparison. RQ2: Does OTA improve grounding and evidence use? With SpaMo on CSL-Daily, Figures 5 and 6 compare matched OTA/no-OTA models on lexical discrimination, retrieval strata, and category/error behavior. Figure 7 evaluates the baseline and full DualAnchor with matched occlusion. Hard-negative separation and retrieval-quality analysis. Figure 5(a) compares sample-level lexical margins against semantically similar hard negatives that differ in action, object, attribute, time, or quantity. For each of 348 eligible examples, Mi=Slex(Y^i,Di+)−Slex(Y^i,Di−)M_i=S_lex( Y_i,D_i^+)-S_lex( Y_i,D_i^-), where SlexS_lex is content-token overlap and Di+D_i^+ and Di−D_i^- are the visually sensitive token sets of the correct and negative references. Positive margins favor the video-specific reference. Figure 5(b) partitions the 1,176 test examples into quartiles of sentence-level video–reference cosine similarity (Q1 lowest; Q4 highest) and reports the share with IDF-weighted content-token recall below 50%50\%. Recall is the recovered share of reference content-token IDF mass; lower low-recall rates are better. (a) Hard-negative margins (b) Low-recall rate by quartile Figure 5: OTA grounding diagnostics on CSL-Daily: (a) lexical margins against semantically similar hard negatives; (b) low-recall rates across video–text similarity quartiles. The two panels separate local grounding from coarse retrieval quality. OTA better distinguishes the correct content from matched confounders, and its low-recall advantage persists within every sentence-level similarity quartile. This advantage spans the observed similarity range. Token-level lexical analysis. At the analyzed checkpoint, Figure 6(a) groups reference content tokens into actions, objects, times/numbers, places, and attributes and reports IDF-weighted recall; annotations give absolute OTA–baseline differences in percentage points. Across all 1,176 test examples, Figure 6(b) reports the percentage exhibiting omission, substitution, over-generalization, or time/number errors; annotations give relative changes from the baseline. Higher recall and lower error rates are better. (a) Category-wise IDF recall (b) Sample-level error rates Figure 6: CSL-Daily OTA diagnostics: (a) category-wise IDF-weighted recall; (b) sample-level lexical-error rates. The category and error views agree: the recall advantage spans distinct content types and coincides with fewer examples exhibiting each measured lexical error. Together, the two views support broad lexical recovery. OTA visual occlusion analysis. Figure 7 tests whether OTA transport mass identifies visual evidence used during generation. With decoding fixed, the baseline and OTA are evaluated on the original video and with equal-length masks over random or highest-mass spans. We report absolute IDF-weighted recall and BLEU-4, not precomputed drops; larger changes indicate greater masking sensitivity. Figure 7: Absolute CSL-Daily IDF-weighted recall (left) and BLEU-4 (right) under original, random-mask, and Top-OTA-mask conditions for the baseline and OTA. Equal-length high-transport masks disrupt content recovery and sentence-level quality more than random masks for both models. The analyses show that OTA improves fine-grained grounding and locates generation-sensitive visual evidence via transport mass. RQ3: How sensitive are TPA and OTA to their respective loss weights? We vary one auxiliary-loss coefficient while fixing all other training and decoding settings. Figure 8 reports each run’s maximum corpus BLEU-4 over saved checkpoints; zero denotes the corresponding no-loss baseline. (a) TPA strength (b) OTA strength Figure 8: CSL-Daily best-checkpoint BLEU-4 versus (a) αTPA _TPA and (b) αOTA _OTA; zero denotes no auxiliary loss. The sweeps distinguish robustness from insensitivity: each objective remains effective across adjacent weights, but translation quality weakens when either term grows too dominant, especially OTA. Thus, the objectives tolerate moderate retuning but should remain auxiliary to translation loss. RQ4: How do core TPA and OTA design choices affect translation quality and module behavior? Table 4(a) compares TPA KL directions and distillation prefixes against no TPA, pairing BLEU-4 with frozen-model PPL. Panel (b) compares sentence-level, attention-based, hard, and OTA alignment plus two OTA subsettings, pairing BLEU-4 with IDF-weighted recall. TPA remains advantageous across the tested directions and prefixes despite a BLEU-4–PPL trade-off. Full OTA provides the strongest joint translation and recall result; removing the dustbin or transporting all tokens weakens this balance. Thus, TPA is robust across the tested implementations, whereas OTA benefits from selective partial matching that avoids unreliable correspondences. (a) TPA Design Setting B4 ↑ PPL ↓ Baseline 20.55 317.90 Direction of KL Optimization TPA w/ Forward KL 23.37 265.92 TPA w/ Reverse KL 23.75 250.84 Context Prefix TPA w/ English Context 22.10 222.90 TPA w/ Target Context 23.56 249.46 (b) OTA Design Design B4 ↑ IDF Recall ↑ Alignment Design Sentence-level 20.55 44.32 Attention alignment 22.27 47.88 Hard alignment 22.38 52.58 OTA 24.07 54.20 OTA Subsettings No dustbin 22.93 53.05 All tokens 23.50 52.32 Table 4: TPA and OTA design ablations on CSL-Daily: (a) KL direction and context prefix, with Target Context as the default; (b) alignment design, dustbin use, and token scope. Related Work Gloss-Free and LLM-Based Sign Language Translation. Gloss-based, unified, and iterative-prototype systems bridge video and text via intermediate recognition or shared representations (Camgöz et al. 2020; Yin and Read 2020; Zhang et al. 2023; Yao et al. 2023). Gloss-free models use visual-language pretraining, boundary-aware features, concept queries, or contrastive sign–text learning (Zhou et al. 2023; Yin et al. 2023; Lin et al. 2023; Jiang et al. 2024); unsupervised and multilingual models expand supervision (Guo et al. 2024; Tan et al. 2025). LLM systems condition pretrained decoders via adapters, discrete sign tokens, factorized adaptation, motion prompts, generated descriptions, or latent plans (Wong et al. 2024; Gong et al. 2024; Chen et al. 2024; Hwang et al. 2025; Kim et al. 2025; Jiang et al. 2026). Preserving Language Priors. Knowledge distillation transfers sequence behavior across models and multilingual translation tasks (Kim and Rush 2016; Sun et al. 2020), while InstructGPT constrains policy updates through per-token KL and pretraining gradients (Ouyang et al. 2022). Multimodal models preserve language competence via visual experts, replay, or selective cross-modal distillation (Wang et al. 2024; Lin et al. 2024; Irawan et al. 2026). TPA uses a shared gold prefix and confidence-weighted token-level KL to regularize conditional translation without an inference-time prior. Fine-Grained Sign-Text Alignment. Content words carry disproportionate semantic weight in translation (Chen et al. 2020a), motivating lexical-unit alignment rather than a single sentence target. Prior work spans global video–text objectives (Zhou et al. 2023; Yin et al. 2023; Jiang et al. 2024), sign–word retrieval, dense contrastive separation, segment supervision, selective negatives (Cheng et al. 2023; Ye et al. 2024; Low et al. 2025; Lai et al. 2026), and optimal-transport sign codebooks (Gong et al. 2024). Entropy-regularized partial transport is differentiable and robust to unmatched units, including in cross-domain alignment (Cuturi 2013; Chapel et al. 2020; Chen et al. 2020b). OTA directly learns soft many-to-many grounding between continuous visual and content-token representations without an inference-time alignment module. Conclusion LLM-based gloss-free SLT must preserve fluent generation and ground lexical choices in visual evidence. DualAnchor meets both requirements through complementary training-only objectives: TPA anchors next-token predictions to a frozen language prior, while OTA uses partial optimal transport to align visual and content tokens without forcing unreliable matches. On PHOENIX-2014T and CSL-Daily, DualAnchor achieves the best BLEU-4 among the compared gloss-free methods and remains effective across three language backbones. Analyses associate TPA with lower prior drift, stronger fluency, and better grammar; OTA recovers more content tokens, reduces lexical errors, and identifies generation-sensitive visual evidence. Thus, preserving linguistic form and fine-grained grounding jointly enable more faithful sign language translation. References S. Asasi, M. I. Lakhal, O. M. Sincan, and R. Bowden (2025) Beyond gloss: a hand-centric framework for gloss-free sign language translation. In 36th British Machine Vision Conference 2025, BMVC 2025, Sheffield, UK, November 24-27, 2025, External Links: Link Cited by: Experiments. N. C. Camgoz, S. Hadfield, O. Koller, H. Ney, and R. Bowden (2018) Neural sign language translation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), p. 7784–7793. Cited by: Introduction, Introduction, Preliminary Analysis, Experiments. N. C. Camgöz, O. Koller, S. Hadfield, and R. Bowden (2020) Sign language transformers: joint end-to-end sign language recognition and translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10020–10030. External Links: Document, Link Cited by: Introduction, Related Work. L. Chapel, M. Z. Alaya, and G. Gasso (2020) Partial optimal transport with applications on positive-unlabeled learning. In Advances in Neural Information Processing Systems, Vol. 33. External Links: Link Cited by: Optimal Transport Alignment, Related Work. K. Chen, R. Wang, M. Utiyama, and E. Sumita (2020a) Content word aware neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 358–364. External Links: Document, Link Cited by: Preliminary Analysis, Related Work. L. Chen, Z. Gan, Y. Cheng, L. Li, L. Carin, and J. Liu (2020b) Graph optimal transport for cross-domain alignment. In Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 119, p. 1542–1553. External Links: Link Cited by: Related Work. Z. Chen, B. Zhou, J. Li, J. Wan, Z. Lei, N. Jiang, Q. Lu, and G. Zhao (2024) Factorized learning assisted with large language model for gloss-free sign language translation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italia, p. 7071–7081. External Links: Link Cited by: Introduction, Experiments, Related Work. Y. Cheng, F. Wei, J. Bao, D. Chen, and W. Zhang (2023) CiCo: domain-aware sign language retrieval via cross-lingual contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 19016–19026. Cited by: Related Work. M. Cuturi (2013) Sinkhorn distances: lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems, C.J. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger (Eds.), Vol. 26, p. 2292–2300. External Links: Link Cited by: Optimal Transport Alignment, Related Work. J. Gong, L. G. Foo, Y. He, H. Rahmani, and J. Liu (2024) LLMs are good sign language translators. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 18362–18372. Cited by: Introduction, Experiments, Related Work, Related Work. Z. Guo, Z. He, W. Jiao, X. Wang, R. Wang, K. Chen, Z. Tu, Y. Xu, and M. Zhang (2024) Unsupervised sign language translation and generation. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 14041–14055. External Links: Link, Document Cited by: Related Work. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: Link Cited by: Experiments. E. J. Hwang, S. Cho, J. Lee, and J. C. Park (2025) An efficient gloss-free sign language translation using spatial configurations and motion dynamics with LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Albuquerque, New Mexico, p. 3901–3920. External Links: Document, Link Cited by: Introduction, Experiments, Related Work. P. A. Irawan, E. H. Fuadi, S. Kumar, A. F. Aji, and Y. Kementchedjhieva (2026) LinguDistill: recovering linguistic ability in vision-language models via selective cross-modal distillation. External Links: 2604.00829, Document, Link Cited by: Related Work. Y. Jiang, L. Zhang, X. Wei, and L. Qing (2026) Think in latent thoughts: a new paradigm for gloss-free sign language translation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, p. 9993–10012. External Links: Document, Link Cited by: Related Work. Z. Jiang, G. Sant, A. Moryossef, M. Müller, R. Sennrich, and S. Ebling (2024) SignCLIP: connecting text and sign language by contrastive learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 9171–9193. External Links: Document, Link Cited by: Introduction, Related Work, Related Work. J. Kim, H. Jeon, J. Bae, and H. Y. Kim (2025) Leveraging the power of MLLMs for gloss-free sign language translation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 21048–21058. Cited by: Experiments, Related Work. Y. Kim and A. M. Rush (2016) Sequence-level knowledge distillation. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, p. 1317–1327. External Links: Document, Link Cited by: Token-Level Prior Anchoring, Related Work. C. H. Lai, R. Zhao, X. Zhong, J. Su, and Y. Chen (2026) Selective contrastive learning for gloss free sign language translation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 45648–45660. External Links: Document, Link, ISBN 979-8-89176-390-6 Cited by: Experiments, Related Work. C. Lin (2004) ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, Barcelona, Spain, p. 74–81. External Links: Link Cited by: Experiments. J. Lin, H. Yin, W. Ping, P. Molchanov, M. Shoeybi, and S. Han (2024) VILA: on pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 26689–26699. Cited by: Related Work. K. Lin, X. Wang, L. Zhu, K. Sun, B. Zhang, and Y. Yang (2023) Gloss-free end-to-end sign language translation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 12904–12916. External Links: Document, Link Cited by: Related Work. Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov, M. Ghazvininejad, M. Lewis, and L. Zettlemoyer (2020) Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics 8, p. 726–742. External Links: Document, Link Cited by: Experiments. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: Link Cited by: Experiments. J. H. Low, O. M. Sincan, and R. Bowden (2025) SAGE: segment-aware gloss-free encoding for token-efficient sign language translation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, p. 5011–5020. Cited by: Related Work. N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. Le Scao, M. S. Bari, S. Shen, Z. X. Yong, H. Schoelkopf, X. Tang, D. Radev, A. F. Aji, K. Almubarak, S. Albanie, Z. Alyafeai, A. Webson, E. Raff, and C. Raffel (2023) Crosslingual generalization through multitask finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 15991–16111. External Links: Document, Link Cited by: Experiments. NLLB Team (2024) Scaling neural machine translation to 200 languages. Nature 630, p. 841–846. External Links: Document, Link Cited by: Experiments. L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35, p. 27730–27744. External Links: Link Cited by: Related Work. K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, Pennsylvania, USA, p. 311–318. External Links: Link, Document Cited by: Experiments. H. Sun, R. Wang, K. Chen, M. Utiyama, E. Sumita, and T. Zhao (2020) Knowledge distillation for multilingual unsupervised neural machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 3525–3535. External Links: Document, Link Cited by: Token-Level Prior Anchoring, Related Work. S. Tan, T. Miyazaki, and K. Nakadai (2025) Multilingual gloss-free sign language translation: towards building a sign language foundation model. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), p. 553–561. External Links: Document, Link Cited by: Related Work. W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, X. Song, J. Xu, K. Chen, B. Xu, J. Li, Y. Dong, M. Ding, and J. Tang (2024) CogVLM: visual expert for pretrained language models. In Advances in Neural Information Processing Systems, Vol. 37, p. 121475–121499. External Links: Document, Link Cited by: Related Work. R. Wong, N. C. Camgoz, and R. Bowden (2024) Sign2GPT: leveraging large language models for gloss-free sign language translation. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: Introduction, Experiments, Related Work. A. Yang, B. Xiao, B. Wang, B. Zhang, C. Bian, C. Yin, C. Lv, D. Pan, D. Wang, D. Yan, F. Yang, F. Deng, F. Wang, F. Liu, G. Ai, G. Dong, H. Zhao, H. Xu, H. Sun, H. Zhang, H. Liu, J. Ji, J. Xie, J. Dai, K. Fang, L. Su, L. Song, L. Liu, L. Ru, L. Ma, M. Wang, M. Liu, M. Lin, N. Nie, P. Guo, R. Sun, T. Zhang, T. Li, T. Li, W. Cheng, W. Chen, X. Zeng, X. Wang, X. Chen, X. Men, X. Yu, X. Pan, Y. Shen, Y. Wang, Y. Li, Y. Jiang, Y. Gao, Y. Zhang, Z. Zhou, and Z. Wu (2023) Baichuan 2: open large-scale language models. External Links: 2309.10305, Document, Link Cited by: Preliminary Analysis. A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2024) Qwen2.5 Technical Report. External Links: 2412.15115, Document, Link Cited by: Preliminary Analysis. H. Yao, W. Zhou, H. Feng, H. Hu, H. Zhou, and H. Li (2023) Sign language translation with iterative prototype. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 15592–15601. External Links: Link Cited by: Related Work. J. Ye, X. Wang, W. Jiao, J. Liang, and H. Xiong (2024) Improving gloss-free sign language translation by reducing representation density. In Advances in Neural Information Processing Systems, Vol. 37, p. 107379–107402. External Links: Document, Link Cited by: Related Work. A. Yin, T. Zhong, L. Tang, W. Jin, T. Jin, and Z. Zhao (2023) Gloss attention for gloss-free sign language translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 2551–2562. Cited by: Introduction, Related Work, Related Work. K. Yin and J. Read (2020) Better sign language translation with STMC-transformer. In Proceedings of the 28th International Conference on Computational Linguistics, p. 5975–5989. External Links: Document, Link Cited by: Introduction, Related Work. B. Zhang, M. Müller, and R. Sennrich (2023) SLTUNET: a simple unified model for sign language translation. In International Conference on Learning Representations, External Links: Link Cited by: Related Work. B. Zhou, Z. Chen, A. Clapés, J. Wan, Y. Liang, S. Escalera, Z. Lei, and D. Zhang (2023) Gloss-free sign language translation: improving from visual-language pretraining. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 20871–20881. Cited by: Introduction, Experiments, Related Work, Related Work. H. Zhou, W. Zhou, W. Qi, J. Pu, and H. Li (2021) Improving sign language translation with monolingual data by sign back-translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 1316–1325. Cited by: Introduction, Preliminary Analysis, Experiments.