Paper deep dive
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
Baban Gain, Asif Ekbal, Trilok Nath Singh
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 2:02:14 AM
Summary
This paper investigates the efficacy of weight-space model merging for multilingual machine translation. By fine-tuning Qwen-2.5-3B-Instruct on Indic-English language pairs, the authors demonstrate that merging independently fine-tuned models leads to significant performance degradation, particularly when target languages differ. The study uses span-conditioned neuron selectivity and layer-wise centered kernel alignment to show that fine-tuning redistributes language-specific neurons into upper transformer layers, creating geometric misalignment that hinders standard merging techniques.
Entities (6)
Relation Signals (2)
Qwen-2.5-3B-Instruct → finetunedon → Samanantar corpus
confidence 100% · We fine-tune the base Qwen-2.5-3B-Instruct model independently on eight bilingual translation tasks from the Samanantar corpus
Task Arithmetic → evaluatedon → Qwen-2.5-3B-Instruct
confidence 90% · We compare against several representative weight-space merging methods... Task Arithmetic
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succeeds in multitask settings, its behavior in multilingual contexts remains poorly understood. We systematically study weight-space merging for multilingual machine translation by fully fine-tuning language model on large-scale bilingual corpora and evaluating standard merging strategies. Our experiments reveal that merging degrades performance, especially when target languages differ. To explain this failure, we analyze internal representations using span-conditioned neuron selectivity and layer-wise centered kernel alignment. We find that language-specific neurons concentrate in embedding layers and upper transformer blocks, while intermediate layers remain largely shared across languages. Critically, fine-tuning redistributes rather than sharpens language selectivity: neurons for supervised and related languages become less exclusive, while those for unsupervised languages grow more isolated. This redistribution increases representational divergence in higher layers that govern generation. These findings suggest that multilingual fine-tuning may reshape geometry in ways that reduce compatibility with standard weight-space merging assumptions. Our work thus provides an explanation for why merging fails in multilingual translation scenarios.
Tags
Links
- Source: https://arxiv.org/abs/2604.02881v1
- Canonical: https://arxiv.org/abs/2604.02881v1
Trouble viewing inline? Open PDF directly →
Full Text
61,655 characters extracted from source content.
Expand or collapse full text
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging Baban GainAsif EkbalTrilok Nath Singh Indian Institute of Technology Patna gainbaban@gmail.com asif@iitp.ac.in tns@iitp.ac.in Abstract Weight-space model merging combines in- dependently fine-tuned models without ac- cessing original training data, offering a practical alternative to joint training. While merging succeeds in multitask settings, its behavior in multilingual contexts remains poorly understood. We systematically study weight-space merging for multilingual ma- chine translation by fully fine-tuning lan- guage model on large-scale bilingual cor- pora and evaluating standard merging strate- gies. Our experiments reveal that merging degrades performance, especially when tar- get languages differ. To explain this fail- ure, we analyze internal representations us- ing span-conditioned neuron selectivity and layer-wise centered kernel alignment. We find that language-specific neurons concen- trate in embedding layers and upper trans- former blocks, while intermediate layers re- main largely shared across languages. Crit- ically, fine-tuning redistributes rather than sharpens language selectivity: neurons for supervised and related languages become less exclusive, while those for unsupervised languages grow more isolated. This redis- tribution increases representational diver- gence in higher layers that govern genera- tion. These findings suggest that multilin- gual fine-tuning may reshape geometry in ways that reduce compatibility with stan- dard weight-space merging assumptions. Our work thus provides an explanation for why merging fails in multilingual transla- tion scenarios. 1 Introduction Large language models (LLMs) have shown re- markable progress across a range of tasks, in- cluding machine translation (Gain et al., 2026), summarization (Zhang et al., 2026), reasoning (Bandyopadhyay et al., 2025), code generation (Jiang et al., 2026), etc. While many modern sys- tems are designed to support multiple languages, training a single multilingual model that performs well across diverse language pairs remains compu- tationally expensive and data-intensive (He et al., 2024; Seto et al., 2025). Further, mixing datasets from multiple languages during training often leads to poor performance on low-resource lan- guages (Chang et al., 2024). A common alter- native is to fine-tune smaller monolingual or task specific models independently (Chouhan et al., 2024; Raihan and Zampieri, 2025), which is more feasible for low resource or domain specific set- tings but results in many specialized models that are costly to host, maintain, and deploy at scale (Yang et al., 2026). Consolidating independently fine tuned models is difficult because training data is often unavail- able due to privacy or licensing limits, making joint retraining unrealistic. Model merging ad- dresses this by combining weights directly with- out data, and has been known to be effective in multiple tasks. However, its behavior in multilin- gual settings remains underexplored for fully fine- tuned bilingual generative MT systems. Recently, some works have combined merging with continued pre-training to inject low-resource or code-mixed capabilities into a base model be- fore downstream fine-tuning (Tao et al., 2024; Ko- dali et al., 2025), where the goal is improved adap- tation or classification performance rather than consolidation of fully specialized generative sys- tems. Other approaches rely on merging language- specific adapters (Zhao et al., 2025; Dmonte et al., 2026), operating in a parameter-efficient setting; while modular and scalable, such adapter-only updates offers limited capacity (Biderman et al., 2024). Domain-focused studies further consider merging general and domain-specific models to enhance terminology retention across languages (Rousset et al., 2025), targeting vocabulary acqui- arXiv:2604.02881v1 [cs.CL] 3 Apr 2026 Traditional Fine-tuningModel Merging Hi-En Bn-En Hi-En Bn-En Hindi to English Bengali to English Ta-En Te-En Ta-En Te-En Tamil to English Telugu to English 4 GPUs 1 GPU Figure 1: Model merging requires only one GPU during deployment whereas individually fine- tuned models needs one GPU per language pair sition instead of end-to-end generative alignment. In contrast, our setup merges independently fine- tuned, full-parameter bilingual machine transla- tion models trained on million-scale corpora. In this paper, we systematically study model merging in multilingual settings using machine translation as a controlled testbed. MT naturally involves distinct source and target languages, al- lowing us to probe disparities between understand- ing and generation in LLMs. We fine tune mod- els on Indic–English pairs (Hindi, Bengali, Tamil, Telugu↔ English), which share a pivot language while differing in typological properties, inducing partially overlapping representational subspaces. We evaluate merging across three configurations: shared source language, shared target language, and merging unidirectional models to form bidi- rectional systems, enabling analysis of how lan- guage configuration affects weight space fusion. Our results show that multilingual merging be- haves differently from standard multitask merging and introduces unique challenges. Our contribu- tions can be summarized as follows: • To the best of our knowledge, this is one of the first systematic studies of weight space merging for multilingual MT using full pa- rameter fine tuning on million scale bilingual corpora, with checkpoints released for bench- marking. • Demonstration of strong directional asym- metry and larger degradation than multitask merging, especially when target languages differ or when forming bidirectional systems. • Analysis via span based activation analysis, Neuron Usage Alignment, and layer wise CKA showing that fine tuning redistributes language specialization to upper layers and creates geometric misalignment that harms merging. 2 Related Works 2.1 Model Merging Model merging has emerged as an alternative to joint multitask training, enabling composition of independently fine-tuned models directly in weight space. Early work on model soups (Worts- man et al., 2022) showed that simple weight av- eraging across fine-tuned checkpoints can im- prove robustness without additional training. Task Arithmetic (Ilharco et al., 2022) formalized this idea through task vectors, defined as the param- eter difference between a fine-tuned model and its pretrained initialization. By linearly combin- ing such vectors and adding them to the base model, multiple task capabilities can be com- posed post hoc.Naive addition of task vec- tors, however, introduces parameter interference when tasks induce conflicting updates. TIES (Ya- dav et al., 2023) addresses this by pruning low- magnitude updates and resolving sign conflicts be- fore merging, thereby reducing destructive can- cellation. DARE (Yu et al., 2024) instead ran- domly drops a large fraction of delta parameters and rescales the remainder to preserve expected magnitude, encouraging sparsity prior to fusion. Fisher merging (Matena and Raffel, 2022) weights parameters according to estimated sensitivity, per- forming a Fisher-weighted average to emphasize directions important for each task. More recently, geometric approaches have been proposed. Sub- space Boosting (Skorobogat et al., 2025) and TSV- Merging (Gargiulo et al., 2025) decompose task vectors via singular value decomposition and re- strict merging to dominant subspaces. By sup- pressing low-variance components, these methods aim to preserve coherent task-relevant structure while limiting interference. Despite this growing body of work, most evaluations have focused on vision backbones or English-centric multitask benchmarks (Huang et al., 2024; Qi et al., 2024).Recent system- atic analyses indicate that techniques effective in vision do not reliably transfer to large language models (Hitit et al., 2025). In particular, merg- ing behavior under multilingual specialization re- mains underexplored. When independently fine- tuned models specialize to different languages, representational shifts may be deeper and struc- turally asymmetric. Our work investigates this set- ting directly. 2.2 Neurons in LLMs Understanding multilingual specialization re- quires examining neuron-level behavior.Geva et al. (2021) characterized feed-forward layers in Transformers as key–value memories, where in- dividual neurons store contextual associations re- trieved during generation. ROME (Meng et al., 2022) later provided causal evidence for local- ized factual knowledge by identifying and edit- ing small subsets of responsible neurons in MLP layers. Subsequent studies have analyzed activa- tion patterns and selectivity. Voita et al. (2024) observed that many neurons remain largely in- active across inputs, while others exhibit highly selective responses to specific tokens or short n- grams. In multilingual models, Tang et al. (2024) introduced language activation metrics showing that small neuron subsets disproportionately con- tribute to language-specific processing. Mondal et al. (2025) further explored whether manipulat- ing such language-specific neurons, through test- time activation replacement or LoRA restricted to selected units, could improve cross-lingual trans- fer, though gains were inconsistent. While prior work emphasizes localization, spar- sity, and intervention within a single model, our study examines how neuron-level specialization evolves under independent fine-tuning and how it interacts with weight-space merging. 3 Methodology 3.1 Large-Scale MT Fine-tuning We fine-tune the base Qwen-2.5-3B-Instruct model independently on eight bilingual transla- tion tasks from the Samanantar corpus (Ramesh et al., 2022): English–Hindi, English–Bengali, English–Tamil, and English–Telugu. Fine-tuning is performed using the LLaMAFactory frame- work (Zheng et al., 2024) with full parameter up- dates allowing the entire model to specialize to each language pair. The selected language pairs are intentionally diverse. Hindi and Bengali belong to the Indo- Aryan language family, whereas Tamil and Telugu are Dravidian languages, enabling cross-family analysis. In addition, the pairs differ in training data scale and baseline performance of the base model, introducing heterogeneity along both lin- guistic and data axes. The training sizes are sub- stantial: 10.1M sentence pairs for English–Hindi, 8.6M for English–Bengali, 5.26M for English– Tamil, and 4.95M for English–Telugu, for each translation direction. Given the large-scale supervision and complete parameter updating, the resulting models consti- tute strong task-specialized systems. We therefore treat these independently fine-tuned checkpoints as practical upper bounds for their respective lan- guage pairs. They serve as competitive references when evaluating merging-based multilingual con- solidation and are subsequently used as inputs to our model merging experiments. 3.2 Baselines We compare against several representative weight- space merging methods, along with the pretrained backbone as a reference. Pretrained Model. The original instruction- tuned backbone prior to task-specific fine-tuning. It serves as a lower bound and quantifies the gains obtained through specialization and merging. Task Arithmetic (Ilharco et al., 2022). Fine- tuning is modeled as a task vector ∆ t = θ t − θ 0 , where θ 0 is the pretrained backbone. Multiple tasks are merged via linear combination, θ merged = θ 0 + P i α i ∆ i . Although simple, direct addition may introduce parameter interference when up- dates conflict. TIES (Yadav et al., 2023). TIES reduces de- structive interference by pruning low-magnitude updates, resolving sign conflicts per parameter, and merging the filtered deltas through normalized summation. DARE (Yu et al., 2024). DARE sparsifies task vectors by randomly dropping a proportion p of parameters and rescaling the remainder by 1 1−p . The sparsified deltas are then merged using stan- dard weight-space techniques. SCE-Merging (Wan et al., 2025).SCE se- lects significant weight changes relative to a pivot model, assigns layer-wise merging weights based on update strength, removes conflicting directions, and integrates the resulting deltas into the pivot. 4 Experimental Findings We evaluate weight-space merging under three controlled multilingual configurations: (i) shared BLEUCHRF ModelHindiBengaliTamilTeluguAverageHindiBengaliTamilTeluguAverage Base23.1218.986.509.5314.5319.9817.3610.776.9513.77 Finetuned - Hindi38.1019.342.075.4716.2562.4045.6814.8923.2136.55 Finetuned - Bengali23.1533.000.742.8814.9450.2159.3212.9223.0836.38 Finetuned - Tamil17.0411.0627.861.8314.4542.5636.2654.6018.7638.05 Finetuned - Telugu18.229.410.3731.4814.8743.5334.519.6557.3036.25 Merged Models Task Arithmetic28.0825.0411.1017.0520.3255.2153.2235.4047.1747.75 TIES28.0425.2211.7618.0020.7655.4253.4735.6245.6347.54 DARE28.5324.8611.1617.1520.4354.9952.6534.3244.1546.53 SCE-Merging33.6826.3511.3618.6022.5059.4553.2933.9943.7947.63 Table 1: BLEU and CHRF scores for different models across languages on Indic-English direction. BLEUCHRF ModelHindiBengaliTamilTeluguAverageHindiBengaliTamilTeluguAverage Base6.281.860.890.562.4028.6123.9024.0616.0223.15 Finetuned - Hindi29.150.310.380.407.5654.160.570.460.5713.94 Finetuned - Bengali0.0515.430.120.193.950.2349.980.150.2012.60 Finetuned - Tamil0.350.3912.871.093.680.640.4551.830.7813.35 Finetuned - Telugu0.460.470.8615.484.240.710.450.6650.0612.97 Merged Models Task Arithmetic6.452.100.400.432.3428.4024.0216.5812.4420.36 TIES6.602.070.360.442.3728.6224.1616.4112.3720.39 DARE6.501.980.340.432.3128.4423.9516.4112.6220.35 SCE-Merging0.760.080.070.080.2513.293.762.740.685.12 Table 2: BLEU and CHRF scores for different models across languages on English-Indic direction. target language (Many→One), (i) shared source language (One→Many), and (i) bidirectional construction by merging directionally opposite models. Across all settings, we compare merged checkpoints against both the pre-trained base model and the task-specific fine-tuned upper bounds. 4.1 Common Target Language: Many→One (Indic→English) Wemergeindependentlyfine-tuned Indic→English models that share a fixed tar- get language (English) but differ in source languages (Hindi, Bengali, Tamil, Telugu), iso- lating multilingual source aggregation under a constant generation space. As shown in Table 1, individual fine-tuned mod- els perform strongly on their supervised pairs but generalize poorly to unseen sources, generat- ing highly uneven performance profiles. Merg- ing reduces this imbalance: CHRF scores be- come consistently distributed across languages, with merged models reaching 46–48 on average, compared to 13.77 for the base model. Merged checkpoints also exceed the average cross-lingual performance of any single fine-tuned model. However, peak task performance is not pre- served. The gap to fine-tuned upper bounds is larger than typically reported in multitask merg- ing. Thus, while Many→One merging improves multilingual coverage and stability, it remains sub- optimal in retaining maximum task accuracy. 4.2 Common Source Language: One→Many (English→Indic) We next consider the complementary configura- tion where models share a common source lan- guage (English) but differ in target languages (Ta- ble 2). In this setting, merging does not yield gains over the base model. Both BLEU and CHRF show that target- language generation expertise is not effectively ag- gregated. Average performance falls below the pre-trained baseline, and degradation relative to fine-tuned upper bounds is substantially more se- vere than in the Many→One case. For example, in English→Hindi, merging retains roughly 22% of the BLEU score (6.5 vs. 29.15), with similar col- En->X->En ModelHindiBengaliTamilTeluguAverageHindiBengaliTamilTeluguAverage Fine-tuned29.1515.4312.8715.48-38.1033.0027.8631.48- Merged Models Task Arithmetic6.701.930.460.532.4128.2323.288.9818.0019.62 TIES7.052.280.560.592.6229.3924.289.1018.4520.31 DARE6.552.170.470.602.4528.9021.418.7118.6319.41 SCE-Merging0.580.511.071.320.8728.7612.5618.1021.7520.29 Table 3: BLEU Scores: Bidirectional Merging En->X->En ModelHindiBengaliTamilTeluguAverageHindiBengaliTamilTeluguAverage Fine-tuned54.1649.9851.8350.06-62.4059.3254.6057.30- Merged Models Task Arithmetic29.3223.0018.7713.3821.1255.9751.8334.7145.8847.10 TIES29.9825.7619.8314.2622.4656.6253.0134.9846.5847.80 DARE29.3424.5518.5613.5021.4956.2351.1234.4845.8346.92 SCE-Merging0.820.600.811.710.9956.6346.2941.8949.4648.57 Table 4: chrf Scores: Bidirectional Merging lapses for Bengali, Tamil, and Telugu. This reten- tion is far below the 70–80% commonly observed in multitask merging, indicating that combining distinct target-side generation spaces is consid- erably more destructive than aggregating diverse source encoders under a shared target. 4.3 Bidirectional Construction via Opposite Directions We finally evaluate whether merging can construct a bidirectional model by combining English→ ℓ and ℓ →English checkpoints for each ℓ ∈ hi, bn, ta, te (Tables 3 and 4). We observe pronounced directional asymme- try. English→Indic collapses under merging, with BLEU retention around 24% for Hindi, 15% for Bengali, and below 5% for Tamil and Telugu. CHRF exhibits the same qualitative pattern. In contrast, Indic→English is partially preserved: al- though BLEU decreases relative to upper bounds, CHRF indicates that English generation adequacy remains non-trivial. Overall, merging opposite directions leads to substantially greater degradation than unidirec- tional merging, and the effect is strongly direction- dependent. Conflicts between source–target role assignments in parameter space appear more se- vere than those arising from multilingual source aggregation alone. 5 Analysis 5.1 Language-Specific Neuron Behavior before and after fine-tuning To quantify language specialization at the neu- ron level, we analyze span-conditioned MLP gate activations measured during forward propagation. Our objective is to isolate source-side and target- side behavior within a single autoregressive se- quence and to identify neurons that are both se- lective and strongly activated for particular lan- guages.Our methodology is inspired by the LAPE framework (Tang et al., 2024), which moti- vates representation-level analysis for understand- ing language-specific specialization within multi- lingual models. Sequence construction. For each language ℓ ∈ L, let D ℓ = (u (ℓ) i ,t (ℓ) i ) N ℓ i=1 denote translation examples, where u (ℓ) i contains a source-language sentence embedded in an instruction prompt and t (ℓ) i is the corresponding target-language sentence. Each pair is rendered into a single token sequence z (ℓ) i = (z i1 ,...,z in ). Conceptually, the sequence decomposes as z (ℓ) i = (p i1 ,...,p ir i | z instruction ,x i1 ,...,x iS i | z source ,y i1 ,...,y iT i | z target ). Instruction tokens are excluded from all subse- quent measurements. Span masks. We define two disjoint masks over token positions: m src ij = 1[z ij ∈ x i1 ,...,x iS i ] and m tgt ij = 1[z ij ∈y i1 ,...,y iT i ]. Because the model is decoder-only with causal attention, the source-span activations are indepen- dent of target tokens, whereas target-span activa- tions are conditioned on the entire source segment. Span-conditioned activation rates. Let L de- note the number of layers and I the intermediate MLP width. Let G (l) ijk denote the post-nonlinearity gate activation of neuron k at position j in layer l. For span s∈src, tgt, define the positive activa- tion count C (ℓ,s) l,k = N ℓ X i=1 n X j=1 1[G (l) ijk > 0]m (s) ij , and the total number of masked tokens N (ℓ,s) = P N ℓ i=1 P n j=1 m (s) ij . The empirical activation prob- ability is then p (ℓ,s) l,k = C (ℓ,s) l,k N (ℓ,s) . The quantity p (ℓ,s) l,k represents the probability that neuron (l,k) produces a positive gate activation when processing span s under language ℓ. Cross-language selectivity. For each neuron (l,k), we normalize activation rates across lan- guages as q (ℓ,s) l,k = p (ℓ,s) l,k / P ℓ ′ ∈L p (ℓ ′ ,s) l,k , with P ℓ q (ℓ,s) l,k = 1. We quantify language selectivity via entropy H (s) l,k =− X ℓ∈L q (ℓ,s) l,k logq (ℓ,s) l,k . Low entropy indicates that activation is concen- trated in a small subset of languages, whereas high entropy indicates shared multilingual behav- ior. We select the fraction ρ of neurons with lowest entropy, i.e.,S (s) = arg min ⌊ρLI⌋ (l,k) H (s) l,k . High-activation criterion. Relative selectivity does not guarantee that a neuron is strongly en- gaged. To ensure that selected neurons exhibit substantial activation, we impose a global activa- tion threshold. Let P = p (ℓ,s) l,k : ∀l,k,ℓ denote the collec- tion of all span-conditioned activation probabili- ties. We define a threshold τ as a high percentile ofP . Indic→ EnEn→ Indic LanguageSource (src)Target (tgt)Source (src)Target (tgt) Hindi542→ 10041680→ 3472706→ 958697→ 743 Bengali701→ 8271693→ 3460686→ 328665→ 780 Tamil950→ 25622286→ 3444662→ 516557→ 1370 Telugu546→ 10942238→ 3553534→ 1498574→ 706 Table 5: Language-specific total selected neuron counts (summed over all layers). Each cell shows Instruct → Fine-tuned counts for the correspond- ing language and span. A selected neuron (l,k) ∈ S (s) is assigned to language ℓ only if p (ℓ,s) l,k > τ . This generates per-layer neuron index sets I (ℓ,s) l =k : (l,k)∈S (s) and p (ℓ,s) l,k > τ. Interpretation. For s = src, the selected neu- ron sets characterize source-language processing, whereas for s = tgt they characterize generation behavior conditioned on the source segment. Ta- ble 5 reports total selected neuron counts (summed over all layers) for both spans and translation di- rections. On the source side, a clear directional asymme- try emerges. In Indic→En, source-span totals in- crease consistently across languages (e.g., Hindi: 542 → 1004, Tamil: 950 → 2562). However, this increase is concentrated primarily in the em- bedding layer (Layer 0), while intermediate trans- former layers remain largely shared. This indi- cates that fine-tuning mainly strengthens lexical encoding for Indic input tokens rather than glob- ally expanding specialization across depth. The complementary En→Indic results display a different pattern. When English is the source language, source-span totals change modestly and may even decrease (e.g., Bengali: 686 → 328). Figure 5 confirms that embedding amplification is limited and that mid-layer overlap is largely preserved. Since English representations are al- ready well supported in the pretrained model, fine- tuning does not require substantial restructuring of source-side encoding. In contrast, generation-side specialization is substantially stronger. Target-span totals increase sharply in Indic→En (e.g., Hindi: 1680 → 3472, Telugu: 2238 → 3553). Figure 4 shows that these increases are concentrated in upper transformer layers (approximately layers 28–35), which directly govern autoregressive token predic- tion. Amplification is therefore not confined to the embedding layer but extends into late decod- ing blocks, indicating deeper reorganization of the generation subspace. A similar depth bias appears in En→Indic, as shown in Figure 6. Although magnitudes vary across languages, specialization again concentrates in upper layers, with hetero- geneous depth profiles across checkpoints. Some models allocate more capacity to embedding lay- ers, whereas others emphasize final transformer blocks. This variability is minimal on the source side but pronounced during generation. Importantly, the layerwise heatmaps reveal that these structural patterns are not restricted to the supervised language alone. Even when a fine- tuned model is probed on a language on which it was not explicitly trained, the activation pro- file often exhibits a qualitatively similar depth- dependent structure, particularly in upper layers. This suggests that fine-tuning redistributes repre- sentational geometry in a way that affects multiple languages, not solely the target pair. We analyze this cross-language generalization behavior more systematically in the next subsection. Taken together, the totals in Table 5 and the layerwise patterns in 6 reveal a consistent struc- tural asymmetry. Source-side specialization, when present, is largely embedding-driven and pre- serves substantial sharing across intermediate lay- ers. Generation-side specialization, however, re- shapes upper transformer blocks that determine the output distribution.Because weight-space merging assumes a degree of geometric compat- ibility across checkpoints, heterogeneous restruc- turing in these late decoding layers leads to mis- aligned generation subspaces. This asymmetry provides a plausible structural explanation for the fragility of weight-space merging in multilingual translation. 5.2 Neuron Usage Alignment To examine whether multilingual fine-tuning in- duces structural separation or shared special- ization, we introduce Neuron Usage Alignment (NUA). The goal is to determine whether indepen- dently fine-tuned bilingual models rely on distinct neuron subsets, or whether they modify largely overlapping computational units. Let C (ℓ,s) l,k denote the masked positive activation count of neuron (l,k) under span s ∈ src, tgt, and let p (ℓ,s) l,k be the corresponding activation rate. For a model M , we define the layer-wise neuron 05101520253035 Layer 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Average NUA (rate-based cosine) En->X: Average BaseFT vs FTFT Neuron Usage Alignment src base-ft src ft-ft tgt base-ft tgt ft-ft 05101520253035 Layer 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Average NUA (rate-based cosine) X->En: Average BaseFT vs FTFT Neuron Usage Alignment src base-ft src ft-ft tgt base-ft tgt ft-ft Figure 2: Neuron-level language selectivity aver- ages for both translation directions. Top: English to X. Bottom: X to English. usage vector u (M,s) l ∈R I whose entries are these activation rates across the I intermediate neurons. NUA between two models M a and M b at layer l is then computed as NUA (s) l (M a ,M b ) = u (M a ,s) l · u (M b ,s) l ∥u (M a ,s) l ∥ 2 ∥u (M b ,s) l ∥ 2 . High NUA indicates that two models activate largely the same neurons with similar frequencies under the masked span, while lower values indi- cate divergence in neuron utilization. Observations Across both source and target spans, NUA remains consistently high between in- dependently fine-tuned bilingual models. From Figure 2, mid and upper layers exhibit near-perfect cosine similarity (typically > 0.98), indicating that the same intermediate MLP neurons are ac- tivated at comparable relative frequencies across language pairs.This pattern holds especially strongly in the target span, where autoregressive generation relies on highly overlapping sets of neurons across fine-tuned checkpoints. Comparisons between the base instruction- tuned model and fine-tuned models show a mod- erate but systematic reduction in NUA, suggest- ing that fine-tuning reshapes neuron usage rel- ative to the pretrained baseline.However, the alignment among fine-tuned models themselves remains substantially higher than alignment be- tween fine-tuned and base models. This indicates that bilingual fine-tuning does not allocate dis- joint subnetworks for different languages, but in- stead modifies a largely shared set of computa- tional units. Taken together, these findings suggest that mul- tilingual specialization does not appear to emerge primarily through neuron partitioning.Rather, fine-tuned models engage similar neurons during both encoding and generation, implying that merg- ing degradation is unlikely to arise from struc- tural separation of neuron subsets. Instead, incom- patibility must originate from finer-grained dif- ferences within shared units, such as divergent weight directions or representational geometry. 5.3 Centered Kernel Analysis To quantify how fine-tuning alters internal repre- sentations, we employ Centered Kernel Alignment (CKA) as a layer-wise measure of representational similarity (Gretton et al., 2005; Kornblith et al., 2019). CKA evaluates whether two models induce similar similarity structures over the same set of inputs. Let x (ℓ) i N ℓ i=1 denote evaluation inputs for lan- guage pair ℓ. For transformer layer k, let H (m,ℓ) k ∈ R N ℓ ×d denote the mean-pooled hidden representa- tions produced by model m on those inputs. For a representation matrix H , define the Gram matrix K = H ⊤ . After centering with C = I − 1 N ℓ 11 ⊤ , we compute ̃ K = CKC. The linear CKA similarity between two representation matri- ces H a and H b is CKA(H a ,H b ) = ∥H ⊤ a H b ∥ 2 F ∥H ⊤ a H a ∥ F ∥H ⊤ b H b ∥ F . We compute CKA at the same layer index k un- der two comparisons. Base vs fine-tuned. CKA H (base,ℓ) k ,H (ft,ℓ) k , which measures how much geometry at depth k shifts after fine-tuning. Fine-tuned vs fine-tuned. CKA H (ft,ℓ) k ,H (ft,ℓ ′ ) k , ℓ̸= ℓ ′ , which measures cross-language geometric align- ment across independently fine-tuned checkpoints. Here, k indexes depth and ℓ indexes the lan- guage pair.High CKA indicates preserved or Early (0–11)Mid (12–27)Late (28–36) DirectionSpanI–FTFT–FTI–FTFT–FTI–FTFT–FT Indic→Ensrc0.9920.9950.9570.9820.740.88 Indic→Entgt0.9910.9940.9520.9780.710.86 En→Indicsrc0.9880.9930.9380.9720.690.83 En→Indictgt0.9850.9900.9210.9650.640.80 Table 6: Layer-banded masked CKA averages. I–FT denotes alignment between the pretrained Instruct model and each fine-tuned checkpoint. FT–FT denotes pairwise alignment among fine- tuned checkpoints. Early, mid, and late correspond to layers 0–11, 12–27, and 28–36, respectively. shared representational geometry, whereas low values reflect language-specific reorganization that may undermine weight-space compatibility. Observations Linear CKA exhibits a clear depth-dependent pattern, but the fine-grained late- layer results in Table 7 reveal that upper-layer di- vergence is far more structured than band averages alone suggest. In both translation directions, early layers remain almost perfectly aligned with the pretrained model, confirming that lexical encoding and shallow compositional structure are largely preserved after fine-tuning. Mid layers show mod- erate drift yet retain strong cross-model similar- ity. The decisive shift occurs in the final decod- ing block. In the English→Indic direction, while source-span alignment with the base model at lay- ers 34 and 35 remains moderate, target-span align- ment collapses sharply at layer 36, reaching values as low as 0.15–0.27 for several languages. The same pattern appears in fine-tuned to fine-tuned comparisons: source-side cross-language align- ment remains extremely high even at layer 36, often above 0.95, but target-side cross-language alignment drops dramatically, in some cases be- low 0.20.This indicates that English encod- ing geometry remains mutually compatible across checkpoints, whereas the target-language gener- ation subspace becomes highly language-specific and geometrically misaligned at the final layer. In contrast, the Indic→English direction shows substantially greater stability in the late decoding layers on the target span. Even at layer 36, align- ment between fine-tuned checkpoints and the base model remains relatively high for English gen- eration, and cross-language fine-tuned alignment frequently exceeds 0.85. Although some source- side divergence appears, the shared English gen- eration space remains structurally coherent across Direction Base vs ft: (src span)Base vs ft: (tgt span)Avg other lang: (src span)Avg other lang: (tgt span) L34L35L36L34L35L36L34L35L36L34L35L36 En→Hindi0.71580.78860.59720.93950.88440.27210.97780.98000.95240.92700.88700.2486 En→Bengali0.69900.77960.58550.92850.87730.28570.97480.98080.96210.91000.87190.2512 En→Tamil0.71480.79240.60510.89980.84750.14820.97130.97950.96580.73380.69470.1187 En→Telugu0.72160.79930.59540.89110.84610.14610.97390.98000.96770.80730.76240.0989 Hindi→En0.78450.74740.54930.93220.91590.78970.88160.87410.62030.97010.96400.8863 Bengali→En0.60410.63550.42990.94230.93320.82290.86550.86620.60270.96940.96510.8818 Tamil→En0.62720.61480.47040.90180.88790.70610.61350.63570.43100.92290.91950.7782 Telugu→En0.63610.60710.48360.91900.90780.68760.74590.74090.48270.95520.94890.8292 Table 7: Late-layer linear CKA for layers 34–36. We report (i) same-language checkpoint vs base model, and (i) average CKA between the language-specific checkpoint and the other three fine-tuned check- points in the same direction 05101520253035 Layer 5 10 15 20 25 30 35 40 Median Principal Angle (degrees) Target-Span Principal Angles Across Layers BaseHi BaseBn BaseTa BaseTe HiBn HiTe TaTe Figure 3: Angle between the representations across models finetuned on En→Indic. independently fine-tuned models. The late-layer shifts are also not smoothly monotonic; several checkpoints exhibit partial recovery at layer 35 before dropping at layer 36, suggesting that in- compatibility concentrates specifically in the fi- nal output-oriented block rather than uniformly across upper layers. These observations clarify the structural asymmetry underlying multilingual merging: early and intermediate layers preserve broadly shared multilingual representations, but the final decoding layer reorganizes in a target- specific manner. When target languages differ, in- dependently fine-tuned checkpoints no longer oc- cupy a compatible generative geometry, consis- tent with degradation observed in One→Many and bidirectional merging settings. 5.4 Principal Angle Analysis To measure geometric compatibility between in- dependently fine-tuned checkpoints, we compute layer-wise principal angles between their repre- sentation subspaces. Let H (m) k ∈R N×d denote the mean-pooled hid- den representations at layer k for model m over masked target tokens. After centering H (m) k , we obtain an orthonormal basis Q (m) k ∈R d×r from the top-r right singular vectors of its SVD. For two models a and b, we compute M k = Q (a)⊤ k Q (b) k . If σ i are the singular values of M k , the principal angles are θ (k) i = arccos(σ i ). We summarize each layer using the median principal angle Θ (a,b) k = median i θ (k) i . Observation. Neuron Usage Alignment (NUA) shows that independently fine-tuned bilingual models activate largely overlapping sets of neu- rons. However, from Figure 3, principal-angle analysis reveals that despite this overlap, the dom- inant representation directions in upper layers di- verge substantially on the target span. Thus, although similar neurons are being used, their collective activation geometry differs. Merg- ing therefore combines geometrically misaligned subspaces within shared computational units, pro- viding a representation-level explanation for mul- tilingual merging degradation. 5.5 Implications for Model Merging in Multilingual Setups Although not directly discussed previously, some findings of the paper have indirect connection with existing literature. Qu and Horváth (2025) show that weight-space interpolation can attenu- ate input-induced features as representations prop- agate through depth, leading to dominance of input-independent components and degraded per- formance. In our experiments, especially when merging models with different target languages or constructing bidirectional systems, we observe severe and direction-dependent performance col- lapse. Rather than a complete variance collapse, the multilingual setting exhibits selective suppres- sion of language-specific generative features, par- ticularly those concentrated in upper transformer layers after fine-tuning. This behavior is consis- tent with a depth-amplified feature attenuation ef- fect under interpolation. Moreover, our neuron-level analysis indicates that fine-tuning redistributes language specializa- tion across layers, increasing representational di- vergence in higher blocks. When such divergent representations are combined, scaling mismatches may dampen strongly language-conditioned ac- tivations, especially those governing target-side generation. Together, these parallels suggest that multilingual merging failures reflect a broader ge- ometric limitation of weight-space fusion when independently specialized models develop mis- aligned feature structures. Our results also resonate with the geometric perspective of Git Re-Basin (Ainsworth et al., 2023), which argues that successful merging re- lies on models occupying the same basin up to permutation symmetries. When internal fea- tures are not aligned modulo permutation, lin- ear interpolation combines incompatible sub- spaces and destroys functionality.In multilin- gual fine-tuning, we observe systematic redistribu- tion of language specialization across layers, with increased representational divergence in higher blocks. Such divergence suggests that indepen- dently fine-tuned translation models, particularly those targeting different output languages, may not remain permutation-equivalent in upper lay- ers. As a result, weight averaging aggregates mis- aligned generative features, producing the strong asymmetries and retention failures observed in our experiments. Together, these connections suggest that such failures may reflect a broader geometric limitation: independently specialized models re- shape their representational structure in ways that violate the alignment assumptions underpinning weight-space fusion. Recent work shows that multilingual LLMs rely on a separation between a semantic sub-circuit, which encodes language-agnostic meaning, and a language-specific sub-circuit, which governs out- put language control; unintended code-switching arises when the dominance of the language- specific pathway is weakened (Xiao et al., 2026). Our findings align with this perspective in a dif- ferent setting. While prior work analyzes compe- tition within a single multilingual model, we study the interaction of independently fine-tuned check- points under merging. We observe that merging is substantially more fragile when target languages differ. This suggests that shared semantic structure remains relatively compatible across checkpoints, but language-specific generation mechanisms are less stable when jointly combined. In this sense, multilingual merging can induce an effect analogous to weakened language-circuit dominance: degradation concentrates on target- language realization rather than semantic ade- quacy, providing a mechanistic interpretation of the asymmetries we observe. 6 Conclusion In this work, we have investigated neuron-level specialization in multilingual large language mod- els and analyzed how independent fine-tuning and weight-space merging affect internal language representations. Our results show that language specialization emerges in small but structured sub- sets of neurons, and this specialization exhibits directional asymmetry across translation settings. We further observed that independently fine-tuned models does not uniformly preserve these special- ized neurons. Representation similarity analysis reveals that early layers remain relatively stable, while middle and late layers undergo substantial reorganization, suggesting that language-specific behavior is concentrated in deeper representations. These findings suggest that weight-space merging interacts nontrivially with neuron-level structure. By linking behavioral performance with internal neuron dynamics, our study offers a perspective on why existing merging strategies degrade mul- tilingual competence. Overall, our work advances understanding of how multilingual knowledge is encoded, specialized, and transformed inside large language models. References Samuel Ainsworth, Jonathan Hayase, and Sid- dhartha Srinivasa. 2023. Git re-basin: Merg- ing models modulo permutation symmetries. In The Eleventh International Conference on Learning Representations. Dibyanayan Bandyopadhyay, Soham Bhattachar- jee, and Asif Ekbal. 2025. Thinking machines: A survey of llm based reasoning strategies. Dan Biderman, Jacob Portes, Jose Javier Gonzalez Ortiz, Mansheej Paul, Philip Greengard, Con- nor Jennings, Daniel King, Sam Havens, Vitaliy Chiley, Jonathan Frankle, Cody Blakeney, and John Patrick Cunningham. 2024. LoRA learns less and forgets less. Transactions on Machine Learning Research. Featured Certification. Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. 2024. When is mul- tilinguality a curse?language modeling for 250 high- and low-resource languages. In Pro- ceedings of the 2024 Conference on Empiri- cal Methods in Natural Language Processing, pages 4074–4096, Miami, Florida, USA. Asso- ciation for Computational Linguistics. Sanjay Chouhan, Shubha Brata Nath, and Apara- jita Dutta. 2024.Hindillm: Large language model for hindi. In Pattern Recognition: 27th International Conference, ICPR 2024, Kolkata, India, December 1–5, 2024, Proceedings, Part VI, page 255–270, Berlin, Heidelberg. Springer- Verlag. Alphaeus Dmonte, Vidhi Gupta, Daniel J Perry, and Mark Arehart. 2026. Improving training efficiency and reducing maintenance costs via language specific model merging. Baban Gain, Dibyanayan Bandyopadhyay, Asif Ekbal, and Trilok Nath Singh. 2026. Bridg- ing the linguistic divide: A survey on leveraging large language models for machine translation. Antonio Andrea Gargiulo, Donato Crisostomi, Maria Sofia Bucarelli, Simone Scardapane, Fabrizio Silvestri, and Emanuele Rodolà. 2025. Task singular vectors: Reducing task interfer- ence in model merging. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484– 5495, Online and Punta Cana, Dominican Re- public. Association for Computational Linguis- tics. Charles Goddard, Shamane Siriwardhana, Ma- likeh Ehghaghi,Luke Meyers,Vladimir Karpukhin, Brian Benedict, Mark McQuade, and Jacob Solawetz. 2024. Arcee’s MergeKit: A toolkit for merging large language models. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing: Industry Track, pages 477–485, Miami, Florida, US. Association for Computational Linguistics. Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Schölkopf. 2005.Measur- ing statistical dependence with hilbert-schmidt norms.In Proceedings of the 16th Inter- national Conference on Algorithmic Learning Theory, ALT’05, page 63–77, Berlin, Heidel- berg. Springer-Verlag. Daniil Gurgurov,Katharina Trinley,Yusser Al Ghussin, Tanja Baeumel, Josef Van Gen- abith, and Simon Ostermann. 2025. Language arithmetics: Towards systematic language neu- ron identification and manipulation.In Pro- ceedings of the 14th International Joint Confer- ence on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pages 2911–2937, Mumbai, India. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics. Yifei He, Alon Benhaim, Barun Patra, Praneetha Vaddamanu, Sanchit Ahuja, Parul Chopra, Vishrav Chaudhary, Han Zhao, and Xia Song. 2024. Scaling laws for multilingual language models. Yifei He, Siqi Zeng, Yuzheng Hu, Rui Yang, Tong Zhang, and Han Zhao. 2025.Mergebench: A benchmark for merging domain-specialized llms. O ̆ guz Ka ̆ gan Hitit, Leander Girrbach, and Zeynep Akata. 2025.A systematic study of model merging techniques in large language models. Chenyu Huang, Peng Ye, Tao Chen, Tong He, Xi- angyu Yue, and Wanli Ouyang. 2024. EMR- merging: Tuning-free high-performance model merging.In The Thirty-eighth Annual Con- ference on Neural Information Processing Sys- tems. Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman,SuchinGururangan,Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2022.Editing models with task arithmetic. arXiv preprint arXiv:2212.04089. Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2026. A survey on large lan- guage models for code generation. ACM Trans. Softw. Eng. Methodol., 35(2). Prashant Kodali, Vaishnavi Shivkumar, Swarang Joshi, Monojit Choudhary, Ponnurangam Ku- maraguru, and Manish Shrivastava. 2025. Adapting multilingual models to code-mixed tasks via model merging. Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019.Similarity of neural network representations revisited. In Proceedings of the 36th International Confer- ence on Machine Learning, volume 97 of Pro- ceedings of Machine Learning Research, pages 3519–3529. PMLR. Ilya Loshchilov and Frank Hutter. 2019. Decou- pled weight decay regularization. In Interna- tional Conference on Learning Representations. Michael S Matena and Colin A Raffel. 2022. Merging models with fisher-weighted averag- ing. Advances in Neural Information Process- ing Systems, 35:17703–17716. Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems. Soumen Kumar Mondal, Sayambhu Sen, Ab- hishek Singhania, and Preethi Jyothi. 2025. Language-specific neurons do not facilitate cross-lingual transfer. In The Sixth Workshop on Insights from Negative Results in NLP, pages 46–62, Albuquerque, New Mexico. Association for Computational Linguistics. Biqing Qi, Fangyuan Li, Zhen Wang, Junqi Gao, Dong Li, Peng Ye, and Bowen Zhou. 2024. Less is more: Efficient model merging with bi- nary task switch. Xingyu Qu and Samuel Horváth. 2025. Vanish- ing feature: Diagnosing model merging and be- yond. In Conference on Parsimony and Learn- ing, volume 280 of Proceedings of Machine Learning Research, pages 1051–1086. PMLR. Nishat Raihan and Marcos Zampieri. 2025. Tiger- LLM - a family of Bangla large language mod- els. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguis- tics (Volume 2: Short Papers), pages 887–896, Vienna, Austria. Association for Computational Linguistics. Gowtham Ramesh, Sumanth Doddapaneni, Ar- avinth Bheemaraj, Mayank Jobanputra, Ragha- van AK, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Mahalakshmi J, Divyanshu Kak- wani, Navneet Kumar, Aswin Pradeep, Sri- hari Nagaraj, Kumar Deepak, Vivek Raghavan, Anoop Kunchukuttan, Pratyush Kumar, and Mitesh Shantadevi Khapra. 2022. Samanantar: The largest publicly available parallel corpora collection for 11 Indic languages. Transactions of the Association for Computational Linguis- tics, 10:145–162. Thibault Rousset, Taisei Kakibuchi, Yusuke Sasaki, and Yoshihide Nomura. 2025. Merg- ing language and domain specific models: The impact on technical vocabulary acquisition. Skyler Seto,Maartje Ter Hoeve,Maureen de Seyssel, and David Grangier. 2025. Assess- ing the role of data quality in training bilingual language models. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2025, pages 22694–22720, Suzhou, China. As- sociation for Computational Linguistics. Ronald Skorobogat, Karsten Roth, and Mariana- Iuliana Georgescu. 2025.Subspace-boosted model merging. Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024. Language- specific neurons: The key to multilingual capa- bilities in large language models. In Proceed- ings of the 62nd Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers), pages 5701–5715, Bangkok, Thailand. Association for Computational Lin- guistics. Mingxu Tao, Chen Zhang, Quzhe Huang, Tianyao Ma, Songfang Huang, Dongyan Zhao, and Yan- song Feng. 2024. Unlocking the potential of model merging for low-resource languages. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8705–8720, Miami, Florida, USA. Association for Compu- tational Linguistics. NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gon- zalez, Prangthip Hansanti, John Hoffman, Se- marley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022.No language left behind: Scaling human-centered machine translation. Elena Voita, Javier Ferrando, and Christoforos Nalmpantis. 2024. Neurons in large language models: Dead, n-gram, positional. In Findings of the Association for Computational Linguis- tics: ACL 2024, pages 1288–1301, Bangkok, Thailand. Association for Computational Lin- guistics. Fanqi Wan, Longguang Zhong, Ziyi Yang, Rui- jun Chen, and Xiaojun Quan. 2025. FuseChat: Knowledge fusion of chat models.In Pro- ceedings of the 2025 Conference on Empiri- cal Methods in Natural Language Processing, pages 21618–21642, Suzhou, China. Associa- tion for Computational Linguistics. Mitchell Wortsman, Gabriel Ilharco, Samir Ya Gadre, Rebecca Roelofs, Raphael Gontijo- Lopes, Ari S Morcos, Hongseok Namkoong, Ali Farhadi, Yair Carmon, Simon Kornblith, et al. 2022. Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. In Interna- tional conference on machine learning, pages 23965–23998. PMLR. Yuxin Xiao, Zhen Huang, Wenxiao Wang, Binbin Lin, Xiaofei He, Xu Shen, and Jieping Ye. 2026. How do language models speak languages? a case study on unintended code-switching. Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023. Ties- merging: resolving interference when merging models. In Proceedings of the 37th Interna- tional Conference on Neural Information Pro- cessing Systems, NIPS ’23, Red Hook, NY, USA. Curran Associates Inc. Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie Zhang, and Dacheng Tao. 2026. Model merging in llms, mllms, and beyond: Methods, theories, applications, and opportunities. ACM Comput. Surv., 58(8). Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li. 2024. Language models are su- per mario: Absorbing abilities from homolo- gous models as a free lunch. In Forty-first In- ternational Conference on Machine Learning. Yang Zhang, Hanlei Jin, Dan Meng, Jun Wang, and Jinghua Tan. 2026. A comprehensive sur- vey on automatic text summarization with ex- ploration of llm-based methods. Neurocomput- ing, 663:131928. Yiran Zhao, Wenxuan Zhang, Huiming Wang, Kenji Kawaguchi, and Lidong Bing. 2025. AdaMergeX: Cross-lingual transfer with large language models via adaptive adapter merging. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the As- sociation for Computational Linguistics: Hu- man Language Technologies (Volume 1: Long Papers), pages 9785–9800, Albuquerque, New Mexico. Association for Computational Lin- guistics. Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024. LlamaFac- tory: Unified efficient fine-tuning of 100+ lan- guage models. In Proceedings of the 62nd An- nual Meeting of the Association for Computa- tional Linguistics (Volume 3: System Demon- strations), pages 400–410, Bangkok, Thailand. Association for Computational Linguistics. A Experimental Setup All models are fine-tuned with a learning rate of 5× 10 −5 for 3 epochs using AdamW (Loshchilov and Hutter, 2019) with β = (0.9, 0.999) and ε = 10 −8 . We use an inverse_sqrt sched- uler and set the random seed to 42. Training is performed on 8 GPUs with per-device batch size 8 and gradient accumulation of 16, resulting in a total train batch size of 1024. Validation is per- formed with FLORES dev set (Team et al., 2022) on corresponding language. For model merging, we use MergeBench (He et al., 2025) to implement TIES, DARE, and Task Arithmetic, and MergeKit (Goddard et al., 2024) for SCE-Merging.For TIES, we sweep K ∈ 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 and scaling ∈ 0.1, 0.2, 0.3, 0.4, 0.5.For DARE, we apply sparsity ∈ 0.6, 0.7, 0.8, 0.9 andscaling ∈ 0.1, 0.2, 0.3, 0.4, 0.5. ForTaskArithmetic,wevaryscaling ∈ 0.1, 0.2, 0.3, 0.4, 0.5.For SCE-Merging, we vary topk ∈ 0.1, 0.3, 0.5, 0.7, 0.9.For each configuration, we select the best merged checkpoint based on the average BLEU score over the languages involved. For language-level neuron analysis, we build on prior work on language specialization (Gurgurov et al., 2025). We select the lowest-entropy frac- tion ρ = 0.1 of neurons and use a high-percentile activation threshold τ = 0.8 when computing ac- tivation rates. Varying these values changes abso- lute counts but preserves the overall trends. B Limitations The empirical scope of this work is constrained along three main dimensions. First, language cov- erage is limited to four Indic languages paired with English. Although these languages span two families and differ in typological properties and data scale, they do not represent the full range of morphological complexity, script variation, and resource disparity found in broader multilingual settings.Extending the study to more distant language families or extremely low-resource sce- narios would require additional large-scale full- parameter fine-tuning, which is computationally demanding and outside the present experimental budget. Second, the core experiments are conducted using a single backbone and scale, Qwen-2.5- 3B-Instruct. This controlled setup enables clean mechanistic and representation-level comparisons, but merging behavior may vary with substantially larger parameter counts or different architectural families. We have conducted a similar study us- ing Llama-3.2-1B (for En→Indic) and observed consistent trends in merging performance degra- dation. These additional results are not included due to space constraints and to maintain a focused analysis, but they provide further evidence that the observed effects are not backbone-specific. Third, we do not include a jointly multilin- gual fine-tuned baseline trained on the union of all language pairs.Such a model would pro- vide a direct consolidation upper bound. How- ever, full-parameter multilingual training over the combined corpora would incur a cost comparable to the aggregate expense of training all bilingual models independently. Repeating this procedure across multiple backbones would multiply the al- ready substantial computational requirements and is therefore not pursued here. Target Masked Activations hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 189185251187 0000 0000 0000 0000 0000 0000 0000 3101 1111 10326 1882018 842115 17172323 731511 31303125 29302120 1211248696 106105132132 152155209187 262263326315 180200241232 174185245235 284267428396 214196304288 251238330317 179175232229 211201211208 128130156153 9799110106 95909488 71677876 56626466 9787109108 8774119120 40317361 0 50 100 150 200 250 300 350 400 Neuron count Bengali-English hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 575598706614 0000 0000 0000 0000 1111 1200 0000 3212 3202 15736 18181515 12342 5463 0103 13161513 910108 908610967 76837488 110116104146 173179186217 135151142190 143152147166 241264216369 202216170293 238262229355 182192170249 131140131162 838883118 57595671 64606571 38363945 31293333 43473353 41423676 41473161 0 100 200 300 400 500 600 700 Neuron count Tamil-English hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 604653908740 0000 0000 0000 0000 0000 1111 0000 0001 0000 2332 2223 54105 14141111 3333 22232123 16161715 69685054 48476865 6771121110 65669484 54505952 31314948 8083170160 6467106109 58709791 43428077 37375552 37386860 32313530 26273531 31272522 43452525 89827184 8279107117 29428075 0 200 400 600 800 Neuron count Instruct hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 195184274199 0000 0000 0000 0000 0000 0000 0000 1110 2100 4233 5142215 231711 14121612 4475 24272325 20191413 120113101104 98100130123 147163215191 226216272261 161151189175 152153188176 242251374334 184214306282 230230328305 189196245233 237228232225 141133152152 119120124128 123118120119 115989591 100919684 123137148143 115123159142 556711199 0 50 100 150 200 250 300 350 Neuron count Hindi-English hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 306298443339 0000 0000 0000 0000 1111 1110 0000 0000 0000 11434 26171118 10314 5566 0050 16141717 1212617 919464113 74819771 155161225141 230249306239 188197252191 185197221183 305319467278 228240342214 274294419270 224235283204 169175209163 9810414692 58598354 66727671 55485151 25234226 44517843 45508739 51547939 0 100 200 300 400 Neuron count Telugu-English Figure 4: Layer-wise neuron counts under target masking for Indic→En. Source Masked Activations hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 2436 0000 0000 1111 0000 0000 1111 0000 3332 1011 23192121 22272021 19192120 0000 5445 2222 1311 0000 4244 0000 5545 3222 2222 1111 1111 1111 1211 18181921 17171918 2433 6676 16171718 344348323335 469504479493 387426382403 851008596 0 100 200 300 400 500 Neuron count English-Bengali hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 5758 0000 0000 0000 0000 1010 1111 1111 2122 1113 101899098 72849275 33313029 5686 98910 4343 0000 1001 9777 3543 15171816 3776 1432 0111 1000 1111 0000 5436 5565 0100 1111 5445 22273130 27263130 21252721 8899 0 20 40 60 80 100 Neuron count English-Tamil hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 89910 0000 0000 0000 0000 1221 0101 1111 2321 25282722 75666570 219236224233 909698101 23202125 66710 7753 3334 14131316 29302824 4347 16161111 2123 1222 0000 2322 0000 1212 12131212 7776 3433 1223 9647 21191419 26302532 25272527 5656 0 50 100 150 200 Neuron count Instruct hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 15201520 0000 0000 1111 0000 2434 2222 0000 13141414 11121210 28212726 22232120 20232321 1222 2333 1211 1111 0000 4333 2000 14151616 6454 0300 2332 4443 7577 2322 21242427 11101212 7776 101198 6766 54565759 71666663 47474446 32383541 0 10 20 30 40 50 60 70 Neuron count English-Hindi hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 2111 0000 0000 0000 0000 0000 0000 0000 4444 4542 911109 881010 3444 0000 2222 0000 0000 0000 0000 0000 2222 0111 0000 2322 2321 2221 0000 29373125 981110 6755 5697 71776174 531572545574 781774778804 616620614644 141159150176 0 100 200 300 400 500 600 700 800 Neuron count English-Telugu Figure 5: Layer-wise neuron counts under source masking for En→Indic. Target Masked Activations hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 70102388248 0000 0000 0010 1141 1059 0004 0025 2010 3034 2254 232951 251733 211266 011264 121342 02947 251154 612937 161028 471948 461140 23724 591547 25629 45627 01420 8101331 141831 021233 10717 443342 33447571 46753846 15192121 6101515 0 50 100 150 200 250 300 350 Neuron count English-Bengali hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 132221586452 0000 0000 0011 1242 2137 2002 0035 2100 4124 8216 201033 31415 10017 10228 00022 00025 00018 11313 1137 22519 00112 1107 11319 22717 20515 0129 82521 01411 0003 1035 3498 22263640 20252028 1814825 1211612 0 100 200 300 400 500 Neuron count English-Tamil hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 68105314210 0000 0000 0010 1121 1058 0013 0033 0000 4111 6286 662732 971626 21417 00126 00113 00113 00110 0015 0013 1107 0006 0016 0006 0004 2115 0114 1091218 2341734 103927 311427 9125972 4453120129 1151557974 54705447 27281821 0 50 100 150 200 250 300 Neuron count Instruct hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 4590268163 0000 0000 0010 1031 1058 0013 0035 0010 2111 4186 333832 121721 202949 102046 211621 101832 602031 1512516 932717 1612219 1201915 401110 1113130 411612 123116 2179 1571622 1672417 721914 421314 1375835 42439863 80554937 26212621 22161511 0 50 100 150 200 250 Neuron count English-Hindi hi bn ta te Language 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 Layer 67109388167 0000 0000 0010 1120 2054 0010 0022 0000 3020 0040 00114 2241 1052 0151 0001 0090 141613 313628 354216 7117619 346312 01403 955016 1106322 386215 13289 645215 63525 10327 10305 10159516 324312953 54587547 30266025 22173516 0 50 100 150 200 250 300 350 Neuron count English-Telugu Figure 6: Layer-wise neuron counts under target masking for En→Indic.