Paper deep dive
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
Hailay Teklehaymanot, Dren Fazlija, Wolfgang Nejdl
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 1:37:09 AM
Summary
The paper introduces the Lexically Grounded Subword Embedding Initialization (LGSE) framework, designed to improve language model performance for morphologically rich, low-resource languages like Amharic and Tigrinya. By replacing arbitrary subword segmentation with morphologically informed units and initializing embeddings using FastText-based morpheme representations, LGSE preserves linguistic structure and reduces semantic fragmentation. The approach includes a regularization term during Language-Adaptive Pretraining to maintain alignment with the original embedding space, consistently outperforming baseline methods across Question Answering, Named Entity Recognition, and Text Classification tasks.
Entities (5)
Relation Signals (4)
LGSE → utilizes → FastText
confidence 98% · LGSE... constructs semantically coherent embeddings by averaging pretrained subword or FastText-based morpheme representations
LGSE → appliedto → XLM-R
confidence 95% · Our experiments are conducted using the multilingual encoder-based model XLM-R... LGSE consistently outperforms baseline methods
LGSE → improvesperformancefor → Amharic
confidence 95% · LGSE consistently outperforms baseline methods across all tasks... in two morphologically rich, low-resource languages: Amharic and Tigrinya
LGSE → improvesperformancefor → Tigrinya
confidence 95% · LGSE consistently outperforms baseline methods across all tasks... in two morphologically rich, low-resource languages: Amharic and Tigrinya
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Adapting pretrained language models to low-resource, morphologically rich languages remains a significant challenge. Existing vocabulary expansion methods typically rely on arbitrarily segmented subword units, resulting in fragmented lexical representations and loss of critical morphological information. To address this limitation, we propose the Lexically Grounded Subword Embedding Initialization (LGSE) framework, which introduces morphologically informed segmentation for initializing embeddings of novel tokens. Instead of using random vectors or arbitrary subwords, LGSE decomposes words into their constituent morphemes and constructs semantically coherent embeddings by averaging pretrained subword or FastText-based morpheme representations. When a token cannot be segmented into meaningful morphemes, its embedding is constructed using character n-gram representations to capture structural information. During Language-Adaptive Pretraining, we apply a regularization term that penalizes large deviations of newly introduced embeddings from their initialized values, preserving alignment with the original pretrained embedding space while enabling adaptation to the target language. To isolate the effect of initialization, we retain the original pre-trained model vocabulary and tokenizer and update only the new embeddings during adaptation. We evaluate LGSE on three NLP tasks: Question Answering, Named Entity Recognition, and Text Classification, in two morphologically rich, low-resource languages: Amharic and Tigrinya, where morphological segmentation resources are available. Experimental results show that LGSE consistently outperforms baseline methods across all tasks, demonstrating the effectiveness of morphologically grounded embedding initialization for improving representation quality in underrepresented languages. Project resources are available in the GitHub link.
Tags
Links
- Source: https://arxiv.org/abs/2603.22629v1
- Canonical: https://arxiv.org/abs/2603.22629v1
Trouble viewing inline? Open PDF directly →
Full Text
52,114 characters extracted from source content.
Expand or collapse full text
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation Hailay Kidu Teklehaymanot, Dren Fazlija, Wolfgang Nejdl L3S Research Center Hannover, Germany teklehaymanot, dren.fazlija, nejdl@L3S.de Abstract Adapting pretrained language models to low-resource, morphologically rich languages remains a significant challenge. Existing vocabulary expansion methods typically rely on arbitrarily segmented subword units, resulting in fragmented lexical representations and loss of critical morphological information. To address this limitation, we propose the Lexically Grounded Subword Embedding Initialization (LGSE) framework, which introduces morphologi- cally informed segmentation for initializing embeddings of novel tokens. Instead of using random vectors or arbitrary subwords, LGSE decomposes words into their constituent morphemes and constructs semantically coherent embeddings by averaging pretrained subword or FastText-based morpheme representations. When a token cannot be segmented into meaningful morphemes, its embedding is constructed using character n-gram representations to capture structural information. During Language-Adaptive Pretraining, we apply a regularization term that penalizes large deviations of newly introduced embeddings from their initialized values, preserving alignment with the original pretrained embedding space while enabling adaptation to the target language. To isolate the effect of initialization, we retain the original pre-trained model vocabulary and tokenizer and update only the new embeddings during adaptation. We evaluate LGSE on three NLP tasks: Question Answering, Named Entity Recognition, and Text Classification, in two morphologically rich, low-resource languages: Amharic and Tigrinya, where morphological segmentation resources are available. Experimental results show that LGSE consistently outperforms baseline methods across all tasks, demonstrating the effectiveness of morphologically grounded embedding initialization for improving representation quality in underrepresented languages. Project resources are available 1 . Keywords:Low-Resource Languages, Morphology-Aware Tokenization, Multilingual NLP 1.Introduction Pretrained multilingual language models (PLMs) have become foundational in modern natural lan- guage processing (NLP), leveraging token se- quences generated from word or subword-level units ( Liu et al.,2024). A representative example is XLM-R (Conneau et al.,2020), a transformer- based model trained on over 100 languages using the SentencePiece algorithm for subword segmen- tation. Although XLM-R employs a shared vocab- ulary of 250K subword units, the effective average token coverage per language is relatively limited, approximately 2.5K subwords compared to mono- lingual models such as GPT, which typically uti- lize vocabularies in the range of 40K tokens (Wang et al. ,2019). Despite their wide language coverage, PLMs tend to favor high-resource languages, especially those that are typologically or orthographically closer to English (e.g., French, Spanish). In contrast, morphologically rich languages such as German face heightened out-of-vocabulary (OOV) challenges due to their complex inflec- tional and derivational systems ( Ataman and Fed- erico ,2018;Lample et al.,2018;Wang et al., 2019). These issues are significantly exacerbated for low-resource languages, particularly those writ- Amharic Word for ”egg”: እንቁላል Subword-Based Embedding Initialization (BPE): እ ን ቁ ላ ል Embedding Composition: E1, E2, E3, E4, E5⇒Combined→ ⃗ V noise Lexically Grounded Subword Embedding (LGSE): እን ቁላል Embedding Composition: E_, E_⇒Combined→ ⃗ V meaningful Figure 1: Comparison of embedding initialization strategies: standard BPE subword splits vs. lin- guistically grounded morphemes. ten in non-Latin scripts. Languages based on the Ge’ez script, such as Amharic and Tigrinya, suffer from poor lexical coverage and unreliable token representations due to a combination of script-specific orthographic complexity and mini- mal training data. To mitigate out-of-vocabulary (OOV) issues, subword tokenization techniques such as Byte Pair Encoding (BPE), introduced by Sennrich et al.(2016), have become foundational in neural machine translation (NMT) and broader NLP pipelines (Hiraoka et al.,2019;Bostrom and Durrett,2020). However, BPE operates purely on character co-occurrence frequency and disre- gards linguistic structure, often fragmenting mor- phologically rich words into arbitrary subword units. This segmentation undermines semantic coher- ence, particularly in agglutinative or templatic lan- guages. The issue is especially pronounced in un- derrepresented languages using the Ge’ez script, such as Amharic and Tigrinya, where BPE fre- quently breaks full lexical units, including complete nouns, into semantically meaningless fragments. As illustrated in Figure 1, the Amharic word እንቁላል(‘Iənqulāl’, meaning “egg”) is decomposed into a series of subwords that fail to preserve its morphemic integrity. This fragmentation neg- atively affects subword embedding initialization by associating these noisy segments with ill- grounded or diluted vector representations, yield- ing embeddings that poorly capture the word’s meaning. Consequently, morphologically aware tokenization is essential not merely for better segmentation but as a prerequisite for reliable and linguistically grounded embedding initializa- tion in morphologically complex, low-resource lan- guages. Multilingual models like mBERT and XLM-R, which rely on shared vocabularies and embedding spaces across languages, often fail to encode the morphosyntactic nuances of such languages ( Ahia et al.,2023;Wang et al.,2019). While language- adaptive pretraining (LAPT) (Chau et al.,2020) and similar transfer learning techniques aim to bridge this representational gap, they still strug- gle when applied to typologically distinct scripts. In particular, expanding the vocabulary or retrain- ing the embedding matrix with newly introduced to- kens disrupts alignment with the pretrained distri- bution, complicating the integration of linguistically informed tokenizers (Dobler and de Melo,2023; de Vries and Nissim,2021). Although methods like vocabulary expansion and random or averaged embedding initialization provide partial relief, they fail to restore the struc- tural grounding that morpheme-level units pro- vide. This is especially critical for morphologi- cally rich languages, where tokenization and em- bedding decisions are tightly coupled. As high- lighted by Mofijul Islam et al.(2022) andLim- isiewicz et al.(2023), subword-based models of- ten produce semantically fragmented and unstable representations across languages, particularly in low-resource settings. To address these limitations, we advocate for embedding strategies that align with morphologi- cally aware tokenization. By respecting linguistic structure during both tokenization and embedding initialization, such methods promise not only im- proved representation quality but also fairer and more effective inclusion of underrepresented lan- guages in NLP systems (Hangya et al.,2023;Tek- lehaymanot et al.,2025b). Our contributions are: (1)We reveal that subword-based embeddings used in current multilingual pretrained models fail to capture the morphological structure of low- resource, morphologically rich languages, leading to fragmented and semantically weak representa- tions; (2)We propose Lexically Grounded Subword Embedding Initialization (LGSE), a novel strategy that respects linguistic boundaries by leveraging morpheme-aware segmentation for embedding ini- tialization. Unlike conventional methods that rely on arbitrary subword fragments, LGSE creates semantically coherent representations, enabling more accurate and robust representation learn- ing for underrepresented, morphologically rich lan- guages such as Amharic and Tigrinya. (3)We introduce the first human-annotated benchmark dataset for evaluating downstream NLP tasks and assessing model performance in identifying high-quality educational content for two morphologically rich, underrepresented lan- guages, Amharic and Tigrinya. This publicly avail- able resource 1 fills a critical gap for low-resource languages and provides a foundation for future research on cross-lingual transfer, morphological modeling, and educational AI. (4)We rigorously evaluate LGSE on three down- stream NLP tasks: Question Answering, Named Entity Recognition, and Text Classification, us- ing two morphologically complex, low-resource languages of Amharic and Tigrinya. Compared to strong multilingual baselines, LGSE achieves substantial and consistent improvements over conventional subword-based embedding methods, demonstrating the effectiveness of linguistically grounded initialization in challenging language set- tings. 2.Related Work 2.1.Subword-Based Tokenization in Low-Resource Languages Subword tokenization methods such as Byte-Pair Encoding (BPE) and SentencePiece are widely 1 https://hailaykidu.github.io/ LGSE-Project-/ used in multilingual pretrained language models. However, these approaches often cause exces- sive fragmentation when applied to morphologi- cally rich and low-resource languages (Rust et al., 2021;Muller et al.,2021;Teklehaymanot and Ne- jdl,2025). This over-segmentation leads to longer token sequences, which increase inference time (Hofmann et al.,2022;Sun et al.,2023), raise API costs (Ahia et al.,2023;Petrov et al.,2023), and degrade downstream task performance (Bostrom and Durrett,2020;Fujii et al.,2023). Tokenizers trained on high-resource languages often produce segmentation mismatches in low-resource lan- guages due to their lack of morphological aware- ness (Sun et al.,2023). 2.2.Morphologically Aware Tokenization Morphologically aware tokenization addresses the limitations of conventional subword methods, which often ignore morpheme boundaries, inflat- ing token counts and degrading performance on richly inflected languages. Approaches such as MorphBPE ( Asgari et al.,2025) and Morph- Piece (Jabbar,2023) integrate linguistic struc- ture with BPE, preserving morphemes and im- proving downstream accuracy across multiple lan- guages. Korean morpheme-aware tokenizers en- hance syntactic performance (Park et al.,2020), while MORSED ( Goot et al.,2025) and Turkish hy- brid tokenizers (Bayram et al.,2025) outperform BPE without sacrificing compression or seman- tic fidelity. MoVoC (Teklehaymanot et al.,2025a) demonstrates gains for low-resource Ge’ez-script languages. Token fertility analyses reveal a sys- tematic “token tax” on morphologically complex languages, underscoring the efficiency and equity advantages of morphology-aware tokenization. 2.3.Vocabulary Expansion and Embedding Initialization Vocabulary expansion is a common strategy for adapting pretrained models to underrepresented languages, particularly when the base vocabulary lacks coverage for non-Latin scripts or language- specific structures (Conneau et al.,2020;Downey et al. ,2024;Pfeiffer et al.,2020). Pretrained mod- els typically use a fixed vocabulary of approxi- mately 50K tokens (Ushio et al.,2023), which of- ten fails to represent morphologically rich or low- resource languages adequately. To address out-of-vocabulary (OOV) issues, em- bedding initialization methods aim to leverage pre- trained representations. Liu et al.(2021) pro- pose synthesizing OOV embeddings using sub- word and hyperword information. UniBridge aligns non-overlapping tokens across languages via syn- tactic and semantic embeddings to enhance cross- lingual transfer (Pham et al.,2024). EVALM mit- igates overfitting by initializing new tokens with high-resource language translations (InitHRL) and applying regularization during fine-tuning (Nag et al.,2023). Recent work has also explored vocabulary ex- pansion in decoder-only models such as LLaMA 2 and 3 to improve generative performance in low- resource settings (Balachandran,2023;Larcher et al.,2023;Lin et al.,2024;Cui et al.,2023; Fujii et al.,2024;Choi et al.,2024b;Nag et al., 2025). These approaches typically add subwords based on frequency and continue pretraining or fine-tune with instruction data. While effective in reducing token overhead and improving fluency, they often rely on naive subword addition and ini- tialization methods that do not account for linguis- tic structure. For example,Balachandran(2023) andCui et al.(2023) introduce additional tokens for Tamil and Chinese using simple initialization, while Fujii et al.(2024) adapts LLaMA 2 for Japanese through cross-lingual pretraining. Similarly,Choi et al.(2024a) andNguyen et al.(2024) extend cov- erage for Korean and Southeast Asian languages. However, these methods do not explicitly address fragmentation or incorporate morphological align- ment in their tokenization strategies. 2.4.Linguistically Informed Embedding Alignment Several studies have explored embedding reinitial- ization and alignment strategies that incorporate cross-lingual semantics. WECHSEL (Minixhofer et al.,2022) maps new subword embeddings to se- mantically similar words using multilingual vector alignment. Although it improves zero-shot transfer, it treats subwords as atomic units and overlooks morphological structure. OFA ( Liu et al.,2024) in- troduces matrix factorization to compress the em- bedding space for scalable adaptation, yet it also ignores language-internal patterns. Language- specific vocabulary augmentation has been shown to improve syntactic tasks in low-resource lan- guages (Chau et al.,2020), andMundra et al. (2024) provides a comparative analysis of embed- ding initialization methods. Nonetheless, existing approaches largely neglect morpheme-based seg- mentation and do not exploit morphological com- position for embedding initialization. 3.Problem Statement Multilingual pretrained models such as mBERT and XLM-R employ a shared subword vocabulary Vacross multiple languagesL=L 1 , L 2 , . . . , L m . For a wordw∈L i , a tokenizerTsegments it into subwordsT(w) = [s 1 , s 2 , . . . , s n ], where each sub- words i ∈Vis associated with a pretrained em- beddinge i ∈R d . However, this subword segmen- tation frequently fails to align with the word’s true morphemic structureM(w) = [m 1 , m 2 , . . . , m k ], where eachm j represents a linguistically meaning- ful morpheme. This misalignment is particularly problematic for morphologically rich and low-resource languages, leading to suboptimal semantic representations and poorer generalization on unseen or infrequent tokens. 4.Vocabulary Expansion and Initialization This section introduces two key approaches to en- hancing embedding initialization for morphologi- cally rich and low-resource languages. Section 4.1presents Lexically Grounded Subword Em- bedding Initialization Framework (LGSE) which leverages subword-level semantic representations from FastText (Bojanowski et al.,2017) to initialize embeddings for morpheme-aligned tokens. This method ensures that the initialized vectors cap- ture meaningful morphological and semantic pat- terns, aligning with the language’s internal struc- ture. In Section 4.2, we describe embedding ini- tialization for new morphologically grounded to- kens, which addresses out-of-vocabulary (OOV) scenarios by generating embeddings for novel morpheme-based units using composition strate- gies informed by morphological structure and dis- tributional semantics. Together, these strate- gies aim to improve vocabulary coverage, seman- tic coherence, and representation quality in low- resource, morphologically complex languages. 4.1.Lexically Grounded Subword Embedding Initialization (LGSE) Given access to a morphologically-aware tok- enizer and pretrained FastText embeddings, we represent a new tokentsegmented into mor- phemesM(t) = [m 1 , m 2 , . . . , m k ]. Each mor- phemem j is further represented by a set of char- actern-grams G j =g j1 , g j2 , . . . , g jn j . Eachn-gramghas an associated FastText embed- dingf g ∈R d . The embedding for morphemem j is computed as the average of its constituentn-gram embeddings m j = 1 |G j | ∑ g∈G j f g while the initial token embedding e t is obtained by averaging over all morpheme embeddings, i.e., e t = 1 k k ∑ j=1 m j = 1 k k ∑ j=1 1 |G j | ∑ g∈G j f g . To align the FastText embedding space with the pretrained model embedding space, a learned lin- ear projectionW∈R d×d is applied, i.e., e aligned t =We t . 4.2.Embedding Initialization for New Lexically Grounded Subword Tokens We initialize the embedding for a new token as the average of its morpheme embeddings computed via FastText-based pooling: e new = 1 k k ∑ j=1 m aligned j , wherem aligned j are morpheme embeddings after projection. For tokens without known morpheme segmentation or embeddings, we initialize by sam- pling from a multivariate normal distribution esti- mated from existing pretrained embeddings: e new ∼N(μ,Σ), whereμandΣare the mean and covariance matrix of pretrained embeddings. To prevent excessive deviation of new embeddings from their initializa- tion during continual pretraining or fine-tuning, we apply the regularization loss L reg =λ∥e new −μ∥ 2 , whereμis the initial embedding vector (e.g., from FastText projection), andλcontrols the regulariza- tion strength, balancing stability and adaptability. 5.Language-Adaptive Pretraining (LAPT) To enhance the XLM-R model’s performance on morphologically complex, low-resource languages such as Tigrinya and Amharic, we move away from subword-based approaches that utilize BPE vo- cabularies, such as FOCUS ( Dobler and de Melo, 2023). Instead, we initialize the embedding layer with lexically grounded representations derived from a morphology-aware tokenizer trained on lin- guistically annotated corpora. This tokenizer seg- ments text into morphemes, preserving the lan- guage’s meaningful lexical and grammatical struc- tures, unlike arbitrary subword units. We employLanguage-Adaptive Pretraining (LAPT)with a morpheme-level Masked Language Modeling (MLM) objective, initializing the embed- ding layer with morpheme-aware representations from annotated corpora to maintain the linguistic integrity of the target languages. For Amharic, we utilize the C100 corpus (133M tokens), previously used in XLM-R pretraining (Conneau et al.,2020), while for Tigrinya, we rely on data from (Gaim et al.,2021), totaling approximately 0.5GB. Hyperparameters for both languages are consistent, as detailed in Table1. We preserve the pretrained XLM-R encoder pa- rametersL 1 , L 2 , . . . , L n and adapt only the em- bedding layerE, initializing it with a language- specific vocabularyV morph tailored to each target language. Pretraining is performed on monolin- gual Tigrinya and Amharic corpora, applying a dy- namic masking probability of 15% to sequences that are either truncated or padded to a maximum length of 256 tokens. 6.Experimental Setup Our experiments are conducted using the mul- tilingual encoder-based model XLM-R ( Conneau et al. ,2020) as the foundational architecture. XLM- R is selected for its proven cross-lingual trans- fer performance and extensive use in multilin- gual NLP research. Its decoupled SentencePiece tokenizer enables straightforward integration of morpheme-level tokens without modifying the un- derlying model architecture. To ensure fair com- parison and reproducibility, all experiments utilize the base version of XLM-R and maintain consis- tent hyperparameters across both baseline and LGSE-enhanced models. The model contains approximately 125 million parameters. Training was performed on a single GPU with 4 CPU cores and 46 GB RAM, with each run allocated up to 24 GPU hours on an Ampere architecture GPU. The computational environment was managed using Anaconda to ensure consis- tency and reproducibility. 6.1.Linguistically Informed Hybrid Tokenization We adopt amorphologically informed tokeniza- tion strategyproposed in MoVoC ( Teklehay- manot et al.,2025a) for segmenting words into lexically grounded morphemes using supervised morphological analysis applied to monolingual cor- poraP am (Amharic) andP ti (Tigrinya). Unlike con- ventional tokenizers that rely solely on frequency- based subword segmentation, the approach re- spects linguistic boundaries to preserve morpho- logical integrity. To construct a vocabulary that is bothlinguistically meaningfulandcomputationally efficient, we combine high-frequency morpheme Table 1: Hyperparameter settings used for fur- ther pretraining with morpheme-aware tokeniza- tion and fine-tuning. HyperparameterValue Maximum sequence length 256 Batch size32 Number of training epochs 10 Learning rate5×10 −5 Learning rate scheduleConstant MLM probability0.15 Weight decay0.01 OptimizerAdam Adamε1×10 −8 Adamβ 1 0.9 Adamβ 2 0.999 Mixed precision (fp16)True tokens with subword units learned viaByte-Pair Encoding (BPE). A hyperparameterr∈[0,1]con- trols the ratio of morpheme tokens, yielding a hy- brid vocabulary: V=V BPE small ∪V morph , |V BPE small |=s(1−r), |V morph |=sr. (1) Tokenization proceeds in two stages:(i)words are first segmented into morphemes;(i)BPE is then appliedwithin each morpheme, preventing merges across morpheme boundaries. Formally, for a wordw=m 1 m 2 ·m k , the tok- enizer output is: Tokenizer(w) = k ∪ i=1 BPE small (m i ). This morphology-aware BPE forms the founda- tion of ourLexically Grounded Subword Embed- ding Initialization (LGSE)framework. By aligning embeddings with linguistically interpretable mor- phemes and subwords, LGSE mitigates seman- tic fragmentation and noise introduced by arbitrary subword splits, thereby producing representations that better capture the morphological richness of underrepresented languages. For practical effi- ciency, we apply the pre-trained MoVoC (Teklehay- manot et al. ,2025a) tokenizer to our parallel and monolingual corpora from theNo Language Left Behind (NLLB)project (Fan et al.,2021) for both Amharic and Tigrinya, generating tokenized se- quences with morphologically-informed subword units. 6.2.Baselines To evaluate the effectiveness of our proposed LGSE approach, we compare it against several strong baselines. In all cases, the original XLM-R encoder layers remain frozen during initialization to isolate the effect of embedding strategies. All models subsequently undergo Language Adaptive Pretraining (LAPT) under identical settings for fair- ness. •XLM-R Off-the-Shelf:The unmodified XLM- R model is used in a zero-shot setting without any additional training. This baseline provides a reference point for assessing the inherent transfer capabilities of the pretrained model in our target languages. •XLM-R + LAPT:The original XLM-R vocab- ulary and embeddings are preserved, and the model is further adapted using Language Adaptive Pretraining on monolingual target language data. This measures the gains from LAPT alone without modifying the tokenizer or embeddings. •Random Initialization for Newly Added To- kens + LAPT:When expanding the vocab- ulary with morphologically grounded tokens, only the embeddings for these new tokens are randomly initialized, while the pretrained em- beddings and encoder parameters remain un- changed. Each new embedding vector is sam- pled from a Gaussian distribution estimated from the original embedding matrix: e t ∼N(μ,Σ), whereμandΣare the empirical mean and covariance of the original embeddings. This baseline isolates the contribution of linguis- tically informed initialization by comparing against a purely random, statistically coherent initialization strategy. •Subword-Based Initialization (FOCUS) + LAPT:We adopt FOCUS ( Dobler and de Melo,2023), a subword-level embedding refinement method that computes weighted combinations of overlapping pretrained sub- word tokens using Sparsemax. This improves representations for rare or unseen tokens without modifying the original tokenizer or vocabulary. •Lexically Grounded Subword Embedding Initialization (LGSE) + LAPT:Our proposed approach combines morphology-aware tok- enization with embedding initialization based on FastText-derived morpheme embeddings, aligned via a learned projection layer. This linguistically informed strategy mitigates over- fragmentation and enhances coverage of mor- phologically rich words, improving representa- tion quality for low-resource languages. 7.Evaluation We evaluate our morphology-aware tokenization by comparing tokenization efficiency, token fertil- ity, and inference latency against standard BPE for Amharic and Tigrinya. Furthermore, we as- sess our proposed models on a range of down- stream NLP tasks across these two morpholog- ically rich, low-resource languages that use the Ge’ez script :AmharicandTigrinya. These lan- guages were selected due to the availability of su- pervised, morphologically annotated data as well as curated evaluation datasets. We conduct ex- periments on three key tasks: text classification, question answering, and named entity recognition. The hyperparameters used for all evaluation tasks are provided in Table 1. Text Classification:We address the task of as- signing predefined labels to input texts, with a spe- cific focus on evaluating the quality of educational content. Educational Quality Classification Dataset: To support this task, we introduce a new bench- mark dataset comprising 2,500 human-annotated samples inAmharicandTigrinya. The dataset was developed in close collaboration with local lin- guistic communities to ensure cultural and linguis- tic relevance. Data collection proceeded in two stages: initially, a diverse set of texts was sourced from publicly available educational materials, in- cluding manuals and blog posts; subsequently, each text was annotated on a 1-6 scale reflect- ing perceived educational quality. Comprehensive dataset statistics and illustrative examples will be provided in the Appendices in the camera-ready version. For model training and evaluation, the dataset was carefully curated and split into 80% for training, 10% for development, and 10% for test- ing. Named Entity Recognition (NER):We perform NER experiments using the balanced train-dev- test splits of theMasakhaNERdataset ( Adelani et al.,2021) for Amharic and for theTigrinya NER dataset(Yohannes and Amagasa,2022), where no official data split is provided, we create a con- sistent partition by randomly splitting the data into 80%for training,10%for development, and10% for testing. Model selection is based on perfor- mance on the development set, and final results are reported on the test set. Question Answering (QA):QA performance is evaluated on theTIGQAtrain-dev-test splits bal- anced dataset ( Teklehaymanot et al.,2024), which contains expert-annotated question-answer pairs in Tigrinya. For Amharic, we use theAmQA, train- dev-test splits dataset ( Taffa et al.,2024), devel- oped for low-resource QA benchmarking. The fi- nal results are reported on the test set for both Table 2: Performance of XLM-R across three NLP tasks in Tigrinya and Amharic. F1 score is used for QA and NER; Accuracy is used for TC. All results are reported as mean±standard deviation over five runs. The best performance per task is highlighted in bold. ModelTask CategoryTask Metric Tigrinya Amharic Avg XLM-R (off-the-shelf) Question AnsweringQAF1 61.3 ± 0.4 71.4 ± 0.9 66.35 Text ClassificationTCAC 63.2 ± 0.7 70.1 ± 0.6 66.65 Named Entity Recognition NER F1 66.4 ± 0.6 70.2 ± 0.8 68.3 XLM-R + LAPT Question AnsweringQAF1 70.5 ± 0.8 74.9 ± 0.5 72.7 Text ClassificationTCAC 69.4 ± 0.5 71.0 ± 0.4 67.8 Named Entity Recognition NER F1 69.8 ± 0.5 75.0 ± 0.6 70.4 XLM-R + Random + LAPT Question AnsweringQAF1 68.7 ± 0.6 71.3 ± 0.8 70 Text ClassificationTCAC 69.9 ± 0.6 70.8 ± 0.8 70.35 Named Entity Recognition NER F1 70.3 ± 0.7 74.0 ± 0.7 72.15 XLM-R + FOCUS + LAPT Question AnsweringQAF1 75.5 ± 0.3 77.8 ± 1.0 76.65 Text ClassificationTCAC 72.4 ± 0.4 76.5 ± 0.9 74.45 Named Entity Recognition NER F1 77.5 ± 0.4 78.1 ± 0.9 77.8 XLM-R + LGSE + LAPT Question AnsweringQAF178.0 ± 0.4 78.5 ± 0.4 78.25 Text ClassificationTCAC75.2 ± 0.5 77.8 ± 0.3 76.5 Named Entity Recognition NER F179.0 ± 0.3 79.4 ± 0.4 79.2 Amharic and Tigriyna QA datasets. We reportF1 scoresfor NER, QA, and Text classification. Each experiment is repeatedfive timeswith different random seeds. We report themean and standard deviationof results. The complete training con- figurations and hyperparameter settings are pre- sented in Table 1. We compare our approach against the baselines mentioned in Section6.2. Unlike these baselines, our method (LGSE) ex- plicitly incorporates morpheme-level structure , which we argue is essential for capturing the deep semantics of morphologically complex languages such as Amharic and Tigrinya. 8.Results and Discussion The results in Table2demonstrate a clear and con- sistent performance improvement when applying our proposed methods across all three tasks to the XLM-R model. 8.1.Baseline Performance The off-the-shelf XLM-R model yields the lowest performance across all tasks. This is expected, as the model has not been adapted to the specific languages or domains involved. For instance, it achieves an average QA F1 score of66.35and NER F1 score of68.30, indicating limited ability to generalize to Tigrinya and Amharic without further adaptation. 8.2.Impact of Language-Adaptive Pretraining (LAPT) ApplyingLanguage-Adaptive Pretraining (LAPT)substantially improves performance across all tasks. QA and NER scores increase by approximately 6-7 percentage points on average, confirming the benefit of continued pretrain- ing on language-specific data for low-resource scenarios. 8.3.Effect of Embedding Initialization Methods Beyond LAPT, we examine the impact of dif- ferent subword embedding initialization meth- ods: Random, FOCUS, and our proposed Lexi- cally Grounded Subword Embedding Initialization (LGSE). TheFOCUS + LAPTconfiguration outper- forms theRandom + LAPTbaseline, achieving a QA F1 score of76.65and NER F1 of77.80. This indicates that more informed subword representa- tions can lead to better convergence and improved performance. 8.4.Effectiveness of LGSE and Cross-Language Impact The proposed method,LGSE + LAPT, which in- tegrates Language-Adaptive Pretraining withLex- ically Grounded Subword Embedding Initial- ization (LGSE), achieves the best overall per- formance, obtaining QA F1 of78.25, TC accu- racy of76.50, and NER F1 of79.20. LGSE employs amorpheme-aware tokenizerthat cap- tures linguistically meaningful units, offering im- proved representations for morphologically rich and low-resource languages such as Amharic and Tigrinya. Unlike conventional subword-based ap- proaches, this method aligns with the underly- ing morphological structure of these languages, thereby enhancing semantic fidelity and reducing segmentation errors. Our analysis further reveals thatvocabulary overlapplays a non-trivial role in cross-lingual embedding transfer. Despite Tigrinya’s absence in pretraining corpora, we observe approximately 1,280 shared morphemes with Amharic, largely driven by code-mixing rather than strict linguis- tic similarity. While this overlap facilitates par- tial transfer, it also introduces potential seman- tic drift. To address rare and out-of-vocabulary morphemes, LGSE leveragesFastText-based character n-gram embeddings, enabling com- positional representations and robust initialization, which are crucial for improving generalization in low-resource settings. Cross-Language Impact.Although Amharic benefits from relatively larger resources, LGSE substantially reduces the performance gap with Tigrinya. This improvement underscores the effec- tiveness oflinguistically informed tokenization and embedding strategiesin supporting cross- lingual generalization under severe resource constraints, particularly for morphologically complex languages. 8.5.Tokenization Metrics and Efficiency As mentioned in Section6.1, to evaluate the practical utility of our tokenization approach, we adopt the pre-trained MoVoC tokenizer (Teklehay- manot et al.,2025a) for Amharic and Tigrinya corpora from theNo Language Left Behind (NLLB)project ( Fan et al.,2021). To highlight the advantages of our approach, we provide il- lustrative tokenization examples comparing our morphology-aware method with conventional BPE. For instance, the Tigrinya sentence: ”ሰላም ንኩሉ ፍጡር” (selam nəkulufət’ur) is tokenized into 21 tokens using standard BPE, whereas our morphology-aware approach produces only 6 to- kens, preserving morphemes and reducing over- segmentation. We defineToken Fertility (TF)as TF= Total tokens Total words . Lower TF indicates fewer redundant tokens per word. Additionally, we measureinference latency (IL) across different computational budgets to eval- uate efficiency trade-offs. Our analysis consis- tently shows that morphology-aware tokenization reduces sequence lengths, lowers TF, and de- creases IL, demonstrating both computational ef- ficiency and practical utility in morphologically rich, low-resource languages. 9.Conclusion We propose a Lexically Grounded Subword Em- bedding Initialization (LGSE) framework for mor- phologically rich, low-resource languages, fo- cusing on Amharic and Tigrinya. By combin- ing morpheme-aware tokenization with FastText- based compositional embeddings and Language- Adaptive Pretraining (LAPT), LGSE consistently improves performance across multiple down- stream tasks. These results underscore the ben- efits of incorporating lexical and morphological structure into multilingual NLP models. 10.Ethical Considerations and Limitations Limitations and Future Work.While the pro- posed framework demonstrates promising im- provements, it faces several limitations. First, it depends on morphologically annotated resources, which remain scarce for many low-resource lan- guages, constraining its applicability in truly mul- tilingual settings. Second, the current design tar- gets encoder-based architectures such as XLM- R, limiting direct integration with decoder-based or sequence-to-sequence models widely used in machine translation and other generative tasks. Third, the incorporation of Lexically Grounded Subword Embedding Initialization introduces ad- ditional computational overhead compared to frequency-driven subword segmentation methods, which may impact scalability for very large vocab- ularies or low-resource deployment environments. As future work, we plan to extend the framework to decoder-based and encoder–decoder architec- tures, enabling its use in machine translation and generative modeling. Additionally, we aim to in- vestigate vocabularyreplacementversusexpan- sionstrategies under these settings to better un- derstand their trade-offs in terms of efficiency and performance across diverse language families. Ethical Considerations.This work uses only publicly available datasets, with all sources prop- erly cited to ensure transparency. ChatGPT was used only for paraphrasing and language clarity no scientific content was generated. The Amharic and Tigrinya annotated datasets, models, and code will be released under an open-access li- cense to support research equity and inclusivity. No personally identifiable information (PII) or sen- sitive content is involved. All research activities adhere to established ethical guidelines for NLP, with attention to linguistic and cultural sensitivity in underrepresented language communities. Our goal is to promote responsible and inclusive cross- lingual NLP development. 11.Acknowledgements This research was supported by the German Academic Exchange Service (DAAD) through the Hilde Domin Programme (funding no. 57615863). 12.Bibliographical References David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Con- stantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, et al. 2021. Masakhaner: Named entity recog- nition for african languages.Transactions of the Association for Computational Linguistics, 9:1116–1131. Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, and Yulia Tsvetkov. 2023. Do all languages cost the same? tokenization in the era of commer- cial language models. InProceedings of the 2023 Conference on Empirical Methods in Nat- ural Language Processing, pages 9904–9923, Singapore. Association for Computational Lin- guistics. Ehsaneddin Asgari, Yassine El Kheir, and Moham- mad Ali Sadraei Javaheri. 2025. Morphbpe: A morpho-aware tokenizer bridging linguistic com- plexity for efficient llm training across morpholo- gies.arXiv preprint arXiv:2502.00894. Duygu Ataman and Marcello Federico. 2018.Com- positional representation of morphologically-rich input for neural machine translation. InProceed- ings of the 56th Annual Meeting of the Associ- ation for Computational Linguistics (Volume 2: Short Papers), pages 305–311, Melbourne, Aus- tralia. Association for Computational Linguistics. Abhinand Balachandran. 2023. Tamil-llama: A new tamil language model based on llama 2. arXiv preprint arXiv:2311.05845. M Ali Bayram, Ali Arda Fincan, Ahmet Semih Gümüş, Sercan Karakaş, Banu Diri, Savaş Yıldırım, and Demircan Çelik. 2025. Tokens with meaning: A hybrid tokenization approach for nlp. arXiv preprint arXiv:2508.14292. Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vec- tors with subword information.Transactions of the association for computational linguistics, 5:135–146. Kaj Bostrom and Greg Durrett. 2020.Byte pair encoding is suboptimal for language model pre- training. InFindings of the Association for Com- putational Linguistics: EMNLP 2020, pages 4617–4624, Online. Association for Computa- tional Linguistics. Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020. Parsing with multilingual BERT, a small corpus, and a small treebank. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 1324–1334, Online. Asso- ciation for Computational Linguistics. ChangSu Choi, Yongbin Jeong, Seoyoon Park, Inho Won, HyeonSeok Lim, SangMin Kim, Yejee Kang, Chanhyuk Yoon, Jaewan Park, Yiseul Lee, HyeJin Lee, Younggyun Hahm, Hansaem Kim, and KyungTae Lim. 2024a. Op- timizing language augmentation for multilingual large language models: A case study on Ko- rean . InProceedings of the 2024 Joint Interna- tional Conference on Computational Linguistics, Language Resources and Evaluation (LREC- COLING 2024), pages 12514–12526, Torino, Italia. ELRA and ICCL. ChangSu Choi, Yongbin Jeong, Seoyoon Park, Inho Won, HyeonSeok Lim, SangMin Kim, Yejee Kang, Chanhyuk Yoon, Jaewan Park, Yiseul Lee, et al. 2024b. Optimizing language aug- mentation for multilingual large language mod- els: A case study on korean.arXiv preprint arXiv:2403.10882. Kartikay Conneau, Alexis Workshop Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoy- anov. 2020.Unsupervised cross-lingual repre- sentation learning at scale. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8440–8451, Online. Association for Computational Linguis- tics. Yiming Cui, Ziqing Yang, and Xin Yao. 2023. Efficient and effective text encoding for chi- nese llama and alpaca.arXiv preprint arXiv:2304.08177. Wietse de Vries and Malvina Nissim. 2021.As good as new. how to successfully recycle En- glish GPT-2 to make models for other languages . InFindings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 836–846, Online. Association for Computational Linguis- tics. Konstantin Dobler and Gerard de Melo. 2023. FOCUS: Effective embedding initialization for monolingual specialization of multilingual mod- els. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Pro- cessing, pages 13440–13454, Singapore. Asso- ciation for Computational Linguistics. C. M. Downey, Terra Blevins, Dhwani Serai, Dwija Parikh, and Shane Steinert-Threlkeld. 2024.Tar- geted multilingual adaptation for low-resource language families. InFindings of the Association for Computational Linguistics: EMNLP 2024, pages 15647–15663, Miami, Florida, USA. As- sociation for Computational Linguistics. Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wen- zek, Vishrav Chaudhary, et al. 2021. Be- yond english-centric multilingual machine trans- lation.Journal of Machine Learning Research, 22(107):1–48. Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Hiroki Iida, Masanari Ohi, Kakeru Hattori, Hi- rai Shota, Sakae Mizuki, Rio Yokota, and Naoaki Okazaki. 2024. Continual pre-training for cross-lingual llm adaptation: Enhancing japanese language capabilities.arXiv preprint arXiv:2404.17790. Takuro Fujii, Koki Shibata, Atsuki Yamaguchi, Teru- fumi Morishita, and Yasuhiro Sogawa. 2023. How do different tokenizers perform on down- stream tasks in scriptio continua languages?: A case study in Japanese. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 39–49, Toronto, Canada. Association for Computational Linguis- tics. Fitsum Gaim, Wonsuk Yang, and Jong C. Park. 2021.Tlmd: Tigrinya language modeling dataset. [Dataset]. Rob van der Goot, Anette Jensen, Emil Allerslev Schledermann, Mikkel Wildner Kildeberg, Nico- laj Larsen, Mike Zhang, and Elisa Bassignana. 2025. MorSeD: Morphological segmentation of Danish and its effect on language modeling. InProceedings of the Joint 25th Nordic Con- ference on Computational Linguistics and 11th Baltic Conference on Human Language Tech- nologies (NoDaLiDa/Baltic-HLT 2025), pages 223–229, Tallinn, Estonia. University of Tartu Li- brary. Viktor Hangya, Silvia Severini, Radoslav Ralev, Alexander Fraser, and Hinrich Schütze. 2023. Multilingual word embeddings for low-resource languages using anchors and a chain of related languages. InProceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL), pages 95–105, Singapore. Association for Com- putational Linguistics. Tatsuya Hiraoka, Hiroyuki Shindo, and Yuji Mat- sumoto. 2019. Stochastic tokenization with a language model for neural text classification. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 1620–1629. Valentin Hofmann, Hinrich Schuetze, and Janet Pierrehumbert. 2022.An embarrassingly sim- ple method to mitigate undesirable properties of pretrained language model tokenizers. InPro- ceedings of the 60th Annual Meeting of the As- sociation for Computational Linguistics (Volume 2: Short Papers), pages 385–393, Dublin, Ire- land. Association for Computational Linguistics. Haris Jabbar. 2023. Morphpiece: A linguistic tok- enizer for large language models.arXiv preprint arXiv:2307.07262. Guillaume Lample, Myle Ott, Alexis Conneau, Lu- dovic Denoyer, and Marc’Aurelio Ranzato. 2018. Phrase-based & neural unsupervised machine translation . InProceedings of the 2018 Confer- ence on Empirical Methods in Natural Language Processing, pages 5039–5049, Brussels, Bel- gium. Association for Computational Linguistics. Celio Larcher, Marcos Piau, Paulo Finardi, Pe- dro Gengo, Piero Esposito, and Vinicius Caridá. 2023. Cabrita: closing the gap for foreign lan- guages.arXiv preprint arXiv:2308.11878. Tomasz Limisiewicz, Jiří Balhar, and David Mareček. 2023.Tokenization impacts multi- lingual language modeling: Assessing vocab- ulary allocation and overlap across languages. InFindings of the Association for Computa- tional Linguistics: ACL 2023, pages 5661–5681, Toronto, Canada. Association for Computational Linguistics. Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, An- dré FT Martins, and Hinrich Schütze. 2024. Mala-500: Massive language adaptation of large language models.arXiv preprint arXiv:2401.13303. Xin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, and Jinsong Su. 2021. Bridging subword gaps in pretrain-finetune paradigm for natural language generation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Con- ference on Natural Language Processing (Vol- ume 1: Long Papers), pages 6001–6011, On- line. Association for Computational Linguistics. Yihong Liu, Peiqin Lin, Mingyang Wang, and Hin- rich Schuetze. 2024.OFA: A framework of ini- tializing unseen subword embeddings for effi- cient large-scale multilingual continued pretrain- ing. InFindings of the Association for Compu- tational Linguistics: NAACL 2024, pages 1067– 1097, Mexico City, Mexico. Association for Com- putational Linguistics. Benjamin Minixhofer, Fabian Paischer, and Navid Rekabsaz. 2022.WECHSEL: Effective initial- ization of subword embeddings for cross-lingual transfer of monolingual language models. InPro- ceedings of the 2022 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Human Language Technolo- gies, pages 3992–4006, Seattle, United States. Association for Computational Linguistics. Md Mofijul Islam, Gustavo Aguilar, Pragaash Ponnusamy, Clint Solomon Mathialagan, Chengyuan Ma, and Chenlei Guo. 2022. A vocabulary-free multilingual neural tokenizer for end-to-end task learning. InProceedings of the 7th Workshop on Representation Learning for NLP, pages 91–99, Dublin, Ireland. Association for Computational Linguistics. Benjamin Muller, Antonios Anastasopoulos, Benoît Sagot, and Djamé Seddah. 2021.When being unseen from mBERT is just the begin- ning: Handling new languages with multilingual language models. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 448– 462, Online. Association for Computational Linguistics. Nandini Mundra, Aditya Nanda Kishore Khan- davally, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, and Mitesh M Khapra. 2024. An empirical comparison of vocabulary expansion and initialization approaches for language mod- els . InProceedings of the 28th Conference on Computational Natural Language Learning, pages 84–104, Miami, FL, USA. Association for Computational Linguistics. Arijit Nag, Soumen Chakrabarti, Animesh Mukher- jee, and Niloy Ganguly. 2025. Efficient continual pre-training of LLMs for low-resource languages. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associa- tion for Computational Linguistics: Human Lan- guage Technologies (Volume 3: Industry Track), pages 304–317, Albuquerque, New Mexico. As- sociation for Computational Linguistics. Arijit Nag, Bidisha Samanta, Animesh Mukher- jee, Niloy Ganguly, and Soumen Chakrabarti. 2023.Entropy-guided vocabulary augmenta- tion of multilingual language models for low- resource tasks. InFindings of the Association for Computational Linguistics: ACL 2023, pages 8619–8629, Toronto, Canada. Association for Computational Linguistics. Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Ma- hani Aljunied, Zhiqiang Hu, Chenhui Shen, Yew Ken Chia, Xingxuan Li, Jianyu Wang, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, and Lidong Bing. 2024.SeaLLMs - large language models for Southeast Asia . InPro- ceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol- ume 3: System Demonstrations), pages 294– 304, Bangkok, Thailand. Association for Com- putational Linguistics. Kyubyong Park, Joohong Lee, Seongbo Jang, and Dawoon Jung. 2020.An empirical study of to- kenization strategies for various Korean NLP tasks . InProceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th Inter- national Joint Conference on Natural Language Processing, pages 133–142, Suzhou, China. Association for Computational Linguistics. Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi. 2023. Language model tokeniz- ers introduce unfairness between languages. Advances in neural information processing sys- tems, 36:36963–36990. Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Se- bastian Ruder. 2020. Mad-x: An adapter-based framework for multi-task cross-lingual transfer. InProceedings of the 2020 Conference on Em- pirical Methods in Natural Language Processing (EMNLP), pages 7654–7673. Trinh Pham, Khoi Le, and Anh Tuan Luu. 2024. UniBridge: A unified approach to cross-lingual transfer learning for low-resource languages. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3168–3184, Bangkok, Thailand. Association for Computa- tional Linguistics. Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. 2021.How good is your tokenizer? on the monolingual perfor- mance of multilingual language models. InPro- ceedings of the 59th Annual Meeting of the As- sociation for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3118–3135, Online. Association for Com- putational Linguistics. Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016.Neural machine translation of rare words with subword units. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 1715–1725, Berlin, Germany. As- sociation for Computational Linguistics. Jimin Sun, Patrick Fernandes, Xinyi Wang, and Graham Neubig. 2023.A multi-dimensional evaluation of tokenizer-free multilingual pre- trained models . InFindings of the Associa- tion for Computational Linguistics: EACL 2023, pages 1725–1735, Dubrovnik, Croatia. Associa- tion for Computational Linguistics. Tilahun Abedissa Taffa, Ricardo Usbeck, and Yare- gal Assabie. 2024. Low resource question an- swering: An Amharic benchmarking dataset . InProceedings of the Fifth Workshop on Re- sources for African Indigenous Languages @ LREC-COLING 2024, pages 124–132, Torino, Italia. ELRA and ICCL. Hailay Kidu Teklehaymanot, Dren Fazlija, Niloy Ganguly, Gourab Kumar Patro, and Wolfgang Nejdl. 2024. TIGQA: An expert-annotated question-answering dataset in Tigrinya . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC- COLING 2024), pages 16142–16161, Torino, Italia. ELRA and ICCL. Hailay Kidu Teklehaymanot, Dren Fazlija, and Wolfgang Nejdl. 2025a. Movoc: Morphology- aware subword construction for ge’ez script lan- guages.arXiv preprint arXiv:2509.08812. Hailay Kidu Teklehaymanot, Gebrearegawi Gebre- mariam Gidey, and Wolfgang Nejdl. 2025b.Low- resource english–tigrinya mt: Leveraging mul- tilingual models, custom tokenizers, and clean evaluation benchmarks . In2025 3rd Interna- tional Conference on Foundation and Large Lan- guage Models (FLLM), pages 121–128. Hailay Kidu Teklehaymanot and Wolfgang Nejdl. 2025.Tokenization disparities as infrastructure bias: How subword systems create inequities in llm access and efficiency. In2025 3rd Interna- tional Conference on Foundation and Large Lan- guage Models (FLLM), pages 822–828. Asahi Ushio, Yi Zhou, and Jose Camacho- Collados. 2023.Efficient multilingual language model compression through vocabulary trim- ming. InFindings of the Association for Com- putational Linguistics: EMNLP 2023, pages 14725–14739, Singapore. Association for Com- putational Linguistics. Hai Wang, Dian Yu, Kai Sun, Jianshu Chen, and Dong Yu. 2019.Hai wang, dian yu, kai sun, jan- shu chen, and 791 dong yu. 2019. improving pre- trained multilingual 792 models with vocabulary expansion. arxiv preprint 793 arxiv:1909.12440. InProceedings of the 23rd Conference on Com- putational Natural Language Learning (CoNLL), pages 316–327, Hong Kong, China. Association for Computational Linguistics. Hailemariam Mehari Yohannes and Toshiyuki Am- agasa. 2022. A method of named entity recogni- tion for tigrinya.ACM SIGAPP Applied Comput- ing Review, 22(3):56–68.