Paper deep dive
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
Mehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira, Anubha Gupta
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/26/2026, 10:24:42 PM
Summary
The paper introduces MORPHOGEN, a new multilingual benchmark designed to evaluate the ability of Large Language Models (LLMs) to perform gender-aware morphological generation. The benchmark focuses on the GENFORM task, which requires models to rewrite first-person sentences from one gender to the opposite gender (masculine to feminine and vice versa) while preserving meaning and structure. The dataset covers three typologically diverse languages: French, Arabic, and Hindi, capturing various morphological strategies like suffixation and agreement. The authors benchmarked 15 LLMs (ranging from 2B to 70B parameters) using three new metrics: Sentence-Level Gender Accuracy (SGA), Gender IoU (GIoU), and Corpus-Level Gender Accuracy (CGA).
Entities (9)
Relation Signals (5)
MORPHOGEN → containstask → GENFORM
confidence 100% · The core task, GENFORM, requires models to rewrite a first-person sentence in the opposite gender
MORPHOGEN → coverslanguage → French
confidence 100% · covers three typologically diverse grammatically gendered languages: French, Arabic, and Hindi.
MORPHOGEN → coverslanguage → Arabic
confidence 100% · covers three typologically diverse grammatically gendered languages: French, Arabic, and Hindi.
MORPHOGEN → coverslanguage → Hindi
confidence 100% · covers three typologically diverse grammatically gendered languages: French, Arabic, and Hindi.
GENFORM → evaluatedby → SGA
confidence 90% · We propose three complementary metrics to assess the accuracy of gender transformations... (SGA, GIoU, CGA)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their ability to handle grammatical gender and morphological agreement remains underexplored. In morphologically rich languages, gender influences verb conjugation, pronouns, and even first-person constructions with explicit and implicit mentions of gender. We introduce MORPHOGEN, a morphologically grounded large-scale benchmark dataset for evaluating gender-aware generation in three typologically diverse grammatically gendered languages: French, Arabic, and Hindi. The core task, GENFORM, requires models to rewrite a first-person sentence in the opposite gender while preserving its meaning and structure. We construct a high-quality synthetic dataset spanning these three languages and benchmark 15 popular multilingual LLMs (2B-70B) on their ability to perform this transformation. Our results reveal significant gaps and interesting insights into how current models handle morphological gender. MORPHOGEN provides a focused diagnostic lens for gender-aware language modeling and lays the groundwork for future research on inclusive and morphology-sensitive NLP.
Tags
Links
- Source: https://arxiv.org/abs/2604.18914v1
- Canonical: https://arxiv.org/abs/2604.18914v1
Trouble viewing inline? Open PDF directly →
Full Text
64,543 characters extracted from source content.
Expand or collapse full text
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation Mehul Aggarwal ♡ * Aditya Agarwal ♡ * Arnav Goel ♡ * Medha Hira ♡ * Anubha Gupta ♡† ♡ SBILab, Indraprastha Institute of Information Technology Delhi anubha@iiitd.ac.in CodeDataset Abstract While multilingual large language models (LLMs) perform well on high-level tasks like translation and question answering, their abil- ity to handle grammatical gender and morpho- logical agreement remains underexplored. In morphologically rich languages, gender influ- ences verb conjugation, pronouns, and even first-person constructions with explicit and im- plicit mentions to gender. We thus introduce MORPHOGENa morphologically grounded large- scale benchmark dataset for evaluating gender- aware generation in three typologically diverse grammatically gendered languages i.e. French, Arabic and Hindi. The core task,GENFORM, requires models to rewrite a first-person sen- tence in the opposite gender while preserving its meaning and structure. We construct a high- quality synthetic dataset spanning French, Ara- bic, and Hindi, and benchmark 15 popular mul- tilingual LLMs (2B–70B) on their ability to perform this transformation. Our results re- veal gaps and interesting insights into the han- dling of morphological gender in current mod- els.MORPHOGENoffers a focused diagnostic lens for gender-aware language modeling and lays the groundwork for future research on inclusive and morphology-sensitive NLP. 1 Introduction Multilingual large language models (LLMs) demonstrate strong performance across tasks such as summarization, translation, and question answer- ing (Goel et al., 2023; Qin et al., 2024; Xu et al., 2025; Huang et al., 2025a; Anand et al., 2023; Kapuriya et al., 2025). Benchmark datasets like XTREME (Hu et al., 2020), Global-MMLU (Singh et al., 2024b), M-Eval (Son et al., 2025), Bench- MAX (Huang et al., 2025b), and IndicGenBench (Singh et al., 2024a) have become standard tools * Authors contributed equally † Corresponding author for evaluating task-specific performance of multi- lingual LLM models. However, they have been criticized for issues such as poor translation quality, data contamination, and an overwhelming empha- sis on high-level tasks that rely heavily on semantic and lexical cues, conflating linguistic competence with broader semantics, making it difficult to iso- late fine-grained weaknesses, particularly in mor- phologically rich or cross-cultural contexts (Wu et al., 2025). As LLMs are being increasing deployed across diverse linguistic settings, it becomes essential to evaluate their ability to apply morphological rules in a grammatically coherent manner (Piergentili et al., 2024; Savoldi et al., 2025). This is espe- cially critical for languages like French, Arabic, and Hindi, which feature rich grammatical gender constructs, where gender affects verb agreement, pronouns, adjectives, and even word order. For instance, first-person sentences in these languages often contain gendered verbs or adjective forms, even when the subject is implicit (Fig 1). Gender marking in such cases is morphologically subtle yet semantically significant (Gonen et al., 2019). Accu- rate modeling of gender morphology is thus crucial not only for inclusive applications like conversa- tional agents and machine translation, but also for probing how gender bias manifests in LLMs across gendered language structures (Sitaram et al., 2025; Zhao et al., 2024; Pikuliak et al., 2024). Despite the linguistic and practical importance of gender mor- phology, there is currently no benchmark that di- rectly evaluates multilingual LLMs on their ability to reason over and apply gender-specific grammati- cal rules in syntactically rich constructions. Exist- ing work (Joshi et al., 2024; Tang et al., 2025; Sant et al., 2024) has primarily tested morphological competence through tokenization or masked word prediction, but falls short of assessing whether mod- els can generate coherent, grammatical sentences conditioned on gender. To address this gap, we arXiv:2604.18914v1 [cs.CL] 20 Apr 2026 Figure 1: Example illustrating how gender-based morphology differs across the three languages introduceMORPHOGENa morphologically grounded benchmark dataset covering French, Arabic, and Hindi designed to evaluate gender-conditioned mor- phological reasoning of LLMs in first-person con- texts. On this benchmark, we define theGENFORMtask as: given a sentence and the speaker’s gender, the model must rewrite the sentence in the opposite gender while preserving grammatical correctness and meaning. To construct linguistically challeng- ing instances, we systematically exploit the rich morphological rules and gender-marking strate- gies in each language. This requires models to go beyond surface-level transformations and en- gage in compositional reasoning over linguistic structure. We evaluate 15 widely-used open- and closed-source multilingual LLMs on this task, span- ning model sizes from under 4 billion to 70 billion parameters. Our key contributions are as follows: (1) We present a new benchmark and dataset covering three typologically diverse, grammatically gen- dered languages: French, Arabic, and Hindi, along- side a parallel English corpus for each sentence (Section 3). This setup enables evaluation on our proposed task as well as on related NLP tasks such as machine translation and gender bias analy- sis. To the best of our knowledge, this is the first and most systematically constructed morphology- focused benchmark for these languages, which we plan to publicly release upon acceptance; (2) We introduce novel evaluation metrics to assess the ac- curacy of gender transformations. These are also applicable to downstream tasks such as translation and gender bias detection in natural language gener- ation (Section 4.2); and (3) We benchmark a range of multilingual LLMs on theGENFORMtask, provid- ing insights into their ability to model and reason about gendered morphological structures. 2 Related Work 2.1 Existing Benchmarks on Multilingual LLMs Recent advancements in multilingual LLM evalua- tion have produced several broad-coverage bench- marks. XTREME (Hu et al., 2020) emerged as a foundational multi-task benchmark spanning 40 languages and 9 tasks (e.g., NER, QA), though its focus on cross-lingual transfer left gaps in morphosyntactic evaluation. Subsequent works like M-Eval (Son et al., 2025) introduced meta- evaluation protocols for 18 languages, emphasiz- ing multilingual consistency in LLM-as-judge sce- narios, but remained task-agnostic to gender mor- phology. Resource-focused frameworks such as GlotEval (Luo et al., 2025) expanded coverage to hundreds of languages across seven NLP tasks, while mHumanEval (Raihan et al., 2024) ad- dressed code generation in 200+ languages via machine-translated prompts. Domain-specific ef- forts like MuST-SHE (Bentivogli et al., 2020). and WinoMT (Stanovsky et al., 2019) pioneered gender-disambiguated MT datasets for Romance languages, though their narrow scope (1k examples per language) limited utility for LLM evaluation. 2.2 Evaluating Gendered Languages in Multilingual LLMs and NLP Systems Grammatically gendered languages such as French, Arabic, and Hindi present unique evaluation chal- lenges due to their rich morphological systems. Prior work has shown that large language models (LLMs) often struggle with correctly realizing gen- der agreement across these languages. For instance, in Hindi, models exhibit errors in gender-inflected verb conjugations and occupational noun morphol- ogy (Hada et al., 2024). In Arabic, evaluations reveal gaps in handling gender agreement across dialectal variations (Rhel and Roussinov, 2025), while studies in French demonstrate a tendency for models to default to masculine forms despite contextual cues (Rescigno et al., 2020). Despite these findings, recent work (Mihaylov and Shtedritski, 2024) highlights that existing benchmarks do not systematically evaluate the ap- plication of gender morphology rules across diverse linguistic typologies. This limitation motivates our work. In contrast to prior datasets like Holistic Bias (Smith et al., 2022), which focus on English de- scriptors of gender and identity, our dataset directly targets the morphological realization of gender in multilingual, grammatically gendered settings. Complementary lines of research further exam- ine how cultural and speech-related factors influ- ence bias in LLMs and NLP systems, underscoring the importance of addressing gender bias through multiple perspectives and modalities (Goel et al., 2024a; Li et al., 2025; Goel et al., 2024b; Hira et al., 2024). 3 Dataset In this section, we introduce theMORPHOGENdataset. We first describe the dataset, the reason for choos- ing the specific languages, inform what we mean by gendered terms, and explain the rules. Next, we explain the dataset construction process, com- pare it to existing parallel corpora, and explain the GENFORM task formulation. 3.1 Dataset Description and Statistics Our proposed dataset,MORPHOGEN, covers three grammatically gendered languages: French, Ara- bic, and Hindi. For each language, we construct a corpus of sentence pairs. Each sentence pair ex- ists with the first person speaker as masculine and as feminine, along with a parallel English version. Thus, for each sentence, we have its gender coun- terfactual as a ground truth, which is used to define a Gendered Term. In other words, gendered terms refer to words that differ between a source sentence and its gender counterfactual. StatisticsArabicFrenchHindi Unique Sentences2,7199,9997,610 Number of Rules141213 Avg. Gender Terms*2.021.781.43 Max. Gender Terms*777 Avg. Word Count*12.3426.7615.46 Max. Word Count*386787 *computed per sentence Table 1: Statistics for MORPHOGEN dataset As shown in Table-1, the dataset includes 9,999 French, 2719 Arabic, and 7,610 Hindi sentence pairs. Figure 2 illustrates the distribution of gen- dered terms per sentence pairs, with some contain- ing up to seven gendered elements, highlighting the morphological complexity of our task. Figure 2: Gendered Terms Distribution in MORPHOGEN 3.2 Task Formulation For the proposedGENFORMtask onMORPHOGEN, we prompt a multilingual LLM with a first-person sen- tence to rewrite the sentence in the opposite gen- der, i.e., from masculine to feminine or vice versa, based on the original speaker’s gender. The model must correctly apply language-specific morpholog- ical rules while preserving the sentence’s meaning, fluency, and syntactic structure. Figure 3: General morphological rules for grammati- cally gendered languages 3.3 Gender Morphology for Chosen Languages MORPHOGEN comprises sentence pairs in three ty- pologically diverse, grammatically gendered lan- guages: French, Arabic, and Hindi. These were deliberately selected to capture a range of gender assignment strategies of semantic and morpholog- ical nature, offering a diverse testbed for evaluat- ing morphological behavior in multilingual LLMs across the selected languages. All three languages feature binary gender systems (masculine and fem- inine), but differ significantly in how gender is marked and propagated. This variation is depicted in Figure 1. French combines semantic, morphological, and phonological cues. While suffixes like -e often indicate feminine gender, exceptions are common. Gender agreement is mandatory across determiners, adjectives, and verbs, but variability in marking makes it typologically distinct. Arabic features a highly regular morphological system where gender is marked primarily via suf- fixation (e.g., -a for feminine). Agreement is strict and pervasive across verbs, adjectives, and pro- nouns, making it a consistent ground for evaluating morphological accuracy. Hindi employs a natural gender system with par- tial morphological marking. Gender is semantically assigned, especially for animate nouns, and com- monly marked via suffixes (e.g., - ̄ a for masculine, - ̄ ı for feminine). Agreement extends to verbs, adjec- tives, and pronouns, but with moderate regularity due to exceptions. Together, these languages exemplify distinct typological frameworks in gender morphology: French integrates phonological, morphological, and semantic gender assignment; Arabic employs regular morphological suffixation with strict agree- ment; and Hindi blends semantic natural gender with morphological suffixes. 3.4 Construction of Morphological Rules To evaluate the performance of multilingual mod- els on gender transformation in first-person con- texts, we constructed a set of language-specific morphological rules grounded in linguistic theory (shown in Figure 4 of Appendix). These rules are inspired by a general taxonomy of gender morphol- ogy across grammatically gendered languages (Fig- ure 3) and are illustrated with concrete examples (Table 3 of Appendix). We present an overview of our motivation behind constructing these rules as: (1) Verbs and Tenses. Gender inflection on verbs depends on both tense and aspect, varying across languages. For instance, French present-tense verbs are gender-invariant, while past participles in compound tenses agree in gender with the subject. Our rules capture such tense-specific patterns. (2) Adjectives and Role Nouns. Adjectives and identity-bearing nouns (e.g., occupations, national- ities) often mark speaker gender morphologically. We design transformation rules to reflect these reg- ular and predictable gendered forms. (3) Pronouns and Possessives. Gender marking in pronouns and possessives is language-dependent. Hindi marks the gender of the possessor, while French and Arabic express gender through gram- matical agreement. Our rules reflect these align- ment differences. (4) Clause-Level Effects. Gender agreement may be influenced by sentence structure, especially in constructions involving passives or object-fronting. We include rules to account for such syntactic in- teractions that affect gender realization. (5) Multiple Entities and Gender Interference. To evaluate a model’s sensitivity to speaker identity, we introduce sentences with two human referents. Only the speaker’s gender governs agreement, al- lowing us to test susceptibility to gender interfer- ence (Lee et al., 2024). We provide detailed rules with examples for each language in the following tables in the Appendix: French (Tables 4, 5), Arabic (Table 6) and Hindi (Tables 7, 8). 3.5 Dataset Construction We constructed theMORPHOGENdataset capturing sentence-level gender transformations in French, Arabic, and Hindi through a structured pipeline grounded in linguistic principles. For each lan- guage, we began by identifying grammatical phe- nomena where a speaker’s gender influences agree- ment or lexical choice, such as in tense and voice (e.g., active/passive), occupations and ad- jectives, pronouns and possessives, and multi- entity contexts prone to gender interference. As each language has its own gender-marking system and grammatical structures, we created language- specific templates (e.g., ‘I am a〈occupation〉’) and independently generated English source sen- tences for each language (i.e., the English inputs are not shared across languages), ensuring struc- tural consistency and systematic coverage across cases. Prompts specifying the rules, lexical argu- ments (e.g., occupation = doctor), and discourse contexts (e.g., politics, classroom, therapy) were used to generate English sentences via GPT-4o- mini (Hurst et al., 2024).These English sentences were translated into Hindi (using IndicTrans2 and GPT-4o-mini) (Gala et al., 2023; Hurst et al., 2024), Arabic (Grok-3) 1 , and French (NLLB-200) (Team et al., 2022). The dataset was refined by multiple bilingual an- notators proficient in English and their respective target languages. Each annotator was randomly as- signed a subset of the data and instructed to follow the refinement guidelines provided in the appendix, discarding any sentences that did not comply. Sub- 1 https://x.ai/grok sequently, each sentence was manually corrected into both masculine and feminine forms by the an- notators, strictly adhering to the correction guide- lines. For detailed annotator instructions, please refer to Appendix B. Finally, the validity of the dataset was verified by cross-validation among annotators. Every sentence pair was independently reviewed by two annota- tors. Two evaluation scores were used for this pro- cess: the Data Validation Score, which measures the overall proportion of valid samples, and the Inter-Annotator Agreement Score, which reports the fraction of entries where both annotators agreed on the validity judgment. The detailed validation procedure is provided in the Appendix B. Across all three languages, the average Data Validation Score and Inter-Annotator Agreement Score were 0.9705 and 0.9495, respectively. A total of eight annotators in the age group 18–21 participated in this process. Language-wise annotation details are presented in Table 9 in the Appendix. The resulting parallel gender-specific annotations form a high- quality gold-standard set for evaluating the model’s sensitivity to morphosyntactic gender variation. 3.6 Comparison with Existing Datasets Standard parallel corpora often default to mascu- line forms when gender is not explicitly marked. For instance, the EuroParl corpus includes speaker metadata but only 30% of its sentences are spoken by women, resulting in a male bias (Koehn, 2005). Such imbalance limits their suitability for evalu- ating gender accuracy. Specialized challenge sets exist but fall short for our speaker-gender restora- tion task: (1) WinoMT targets occupational stereotypes across languages, including English–Hindi, but relies on rigid templates that models may over- fit to (Stanovsky et al., 2019). (2) MT-GenEval improves diversity and realism for English–Hindi but lacks first-person sentences and speaker-gender labels (Currey et al., 2022). (3) MuST-SHE offers speaker annotations and first-person content, but is not publicly available (Bentivogli et al., 2020). (4) mGENTE supports gender-neutral generation across languages (Savoldi et al., 2025), but lacks speaker-grounded, first-person constructions. (5) Arabic Parallel Gender Corpus 2.0 provides first- and second-person gendered sentence pairs from English–Arabic OpenSubtitles (Lison and Tiede- mann, 2016; Al Khalifa et al., 2022), but its cover- age is limited to a few recurring gender-marking rules. Our dataset captures a broader range of morpho- syntactic phenomena, offering a stronger bench- mark across Arabic, Hindi, and French. To the best of our knowledge, no existing dataset: 1.Provides male and female translations for every sentence. 2. Aligns examples with grammatical triggers for gender inflection. 3.Ensures balanced ground truth for both genders. 4.Covers the full spectrum of gender-marking phe- nomena. Mining real transcripts is inefficient because most sentences are gender-neutral and only a few cover key structures. In contrast, prompting large language models under controlled templates en- ables efficient generation of diverse, balanced, and linguistically grounded examples across Hindi, Arabic, and French. 4 Experimental Setup 4.1 Models Benchmarked To effectively evaluate the performance of multilin- gual LLMs onMORPHOGEN, we conducted extensive benchmarking across 15 models spanning a diverse range of model families and parameter scales. The models evaluated include: •LLAMA: LLAMA-3.1-8B, LLAMA-3.2-3B, LLAMA-3.3-70B (Grattafiori et al., 2024) •Qwen: Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3-27B (Yang et al., 2025) • Gemma:Gemma2-2B,Gemma2-9B, Gemma3-4B, Gemma3-12B, Gemma3-27B (Team et al., 2024, 2025) • Phi: Phi4-14B (Abdin et al., 2024) Our goal was to cover a representative and prac- tical spectrum of contemporary multilingual LLMs, ranging from lightweight models (e.g., 2B–4B pa- rameters) suitable for deployment and industry use- cases, to high-capacity models (up to 70B parame- ters) that are expected to exhibit stronger multilin- gual generalization. These models were selected based on their widespread adoption, open-source availability, and explicit support for the three gen- dered languages under study. 4.2 Evaluation Metrics To evaluate model performance on theMORPHOGEN benchmark, we propose three complementary met- rics that measure an LLM’s ability to correctly per- form gender-aware morphological transformations at different granularities. Note that for any sen- tence, we collect gendered terms by referring to its gender-counterfactual as presented in Section 3. The proposed metrics are defined as follows: (1) Sentence-Level Gender Accuracy (SGA): This metric measures the proportion of correctly generated gendered terms in each sentence. For a given sentence, we compute the number of gen- dered words that were correctly modified (i.e., match the gold-standard target) and divide this by the total number of gendered terms in the reference sentence. SGA captures sentence-level precision in handling gendered terms, ensuring correctness at a fine-grained unit of evaluation. The final score is the average of this ratio across allNsentences in the corpora: SGA = 1 N N X i=1 |Gendered i ∩ Mismatch c i | |Gendered i | As described in Section 3.2, theGENFORMtask eval- uates bidirectional gender transformation: mascu- line to feminine and vice versa. We report dis- aggregated results for each direction, denoted as SGA M andSGA F , corresponding to masculine-to- feminine and feminine-to-masculine conversions, respectively. Additionally, to evaluate any perfor- mance gaps between the masculine and feminine disaggregation, we report the gaps between the masculine and feminine scores△SGA. △SGA = SGA M − SGA F (2) Gender IoU Score (GIoU): Inspired by the Intersection-over-Union (IoU) metric commonly used in object detection, GIoU metric provides a stricter and more comprehensive measure of mor- phological transformation quality. It penalizes both over-generation (modifying non-gendered terms or incorrect gendered entities) and under-generation (failing to modify gendered terms). For each sen- tence, we computed the ratio between the num- ber of correctly transformed gendered terms to the union of gendered and mismatched terms. The final score is the mean of sentence-level IOU values: GIoU = 1 N N X i=1 |Gendered i ∩ Mismatch c i | |Gendered i ∪ Mismatch i | This metric captures both precision and recall and is especially useful in sentences with multiple entities or partial gender relevance, where models may hallucinate or overlook certain terms. This metric is particularly useful for evaluating cases of gender interference, where the model incor- rectly alters the gender of words associated with entities other than the explicit speaker. In such cases, GIoU penalizes the model for transforming non-gendered terms, thereby ensuring that only valid gender-specific modifications are rewarded. Again, we report disaggregated results for each di- rection i.e.,GIoU M andGIoU F , corresponding to masculine-to-feminine and feminine-to-masculine conversions, respectively. (3) Corpus-Level Gender Accuracy (CGA) This is a corpus-level aggregation of gender cor- rectness. Instead of averaging per-sentence ratios, we computed the ratio of number of correctly gen- erated gendered terms across the entire test set to the total number of reference gendered terms in the corpus. This provides a holistic measure of overall transformation quality at an n-gram level: CGA = P N i=1 |Gendered i ∩ Mismatch c i | P N i=1 |Gendered i | Unlike SGA, which evaluates correctness at the sentence level, CGA extends evaluation to the word level across the entire corpus, which makes it es- pecially effective for longer or more complex sen- tences with multiple gendered terms. 5 Results and Discussion We evaluated 15 widely-used open-source and closed-source multilingual LLMs on the MORPHOGENbenchmark across French, Arabic, and Hindi, using the metrics defined in Section 4.2. The consolidated results are presented in Table 2, with detailed language-specific results provided in the Appendix: French in Table 10, Arabic in Table 11, and Hindi in Table 12. We analyse the variations in performance across different model families and sizes, and discuss the implications for gender bias in these models. 5.1 Smaller LMs can’t handle Complex Morphology Larger models consistently outperformed smaller ones across all languages, particularly in Arabic, where increased parameter size mitigated morpho- logical complexity. For example,Gemma3-27B (27B parameters) achieved a CGA of 74.74% in Arabic, markedly outperformingGemma2-2Bat 14.10%. In Hindi, smaller models remained viable due to simpler rules, withLLAMA-3.1-8Bscoring a CGA of 89.21%, compared toLLAMA-3.3-70B at 91.40%. French’s larger dataset challenged resource-constrained models, amplifying errors, asGemma2-2Brecorded a CGA of 37.54%, while Phi4-14Breached 87.70%. This suggests that pa- rameter size is critical for handling complex mor- phology but less impactful in simpler linguistic contexts like Hindi. 5.2 Masculine Bias in French and Arabic Gender bias varied notably across languages, as seen in the△SGAscores (Figure 5 in Appendix). In Hindi, bias was generally low but occasionally skewed toward feminine forms, with models like Gemma3-4Bshowing an△SGAof-14.32%, often preferring feminine outputs even when the target gender was male. In French, a stronger masculine bias was observed, particularly in larger models such asLLaMA3-70B, which exhibited an△SGA of15.15%due to consistent defaulting to masculine forms. Arabic showed persistent masculine bias, especially in plural constructions, withQwen3-32B recording an△SGAof11.94%, frequently gener- ating masculine outputs even in all-female contexts. These trends highlight the influence of gender bias of the training data used in these LLMs and under- score the need for targeted debiasing in morpholog- ically rich languages. 5.3 Significant Variance in Model Families Architectural differences influenced performance quality.Gemmamodels excelled in gender fairness, particularly in Arabic, maintaining balance in com- plex contexts.LLAMAmodels showed consistency in Hindi and French but struggled with bias in Ara- bic.Qwenmodels frequently exhibited masculine bias across languages, suggesting weaker gender handling.Phimodels achieved high consistency but faced challenges with entity recognition, espe- cially in Hindi. 5.4 Models Misapply Gender in Multi-Entity Sentences Gender interference occurs when a model incor- rectly alters words associated with all entities’ gen- ders instead of only the gendered terms in sentences with multiple human entities. To measure the cor- rect transformation of gendered terms, we use gen- der accuracy, which counts only the changes to the intended gendered words. To further penalize any modifications of non-gendered words, we introduce Gendered IoU (GIoU), which is a stricter metric ModelFrenchArabicHindi GIoU↑△SGA↓CGA↑GIoU↑△SGA↓CGA↑GIoU↑△SGA↓CGA↑ QWEN2.5-0.5B5.474.554.164.148.494.590.350.630.21 GEMMA2-2B39.73-5.1437.5414.73-0.8114.1071.417.3565.41 LLAMA-3.2-3B54.4911.4253.4818.31-27.6117.7548.54-64.6549.72 GEMMA3-4B52.70-14.1651.6045.68-8.2048.9367.50-14.3264.58 QWEN3-4B58.647.2553.2034.34-0.9035.9762.843.3368.51 LLAMA-3.1-8B67.893.6981.7643.510.9645.5183.12-0.4389.21 QWEN3-8B71.664.8669.9145.895.0351.0180.962.1687.82 GEMMA2-9B60.521.2655.5646.452.5545.2685.47-7.3484.39 GEMMA3-12B64.27-0.5874.2662.762.5065.5279.91-8.9380.99 PHI4-14B79.841.1787.7057.086.5866.1582.771.3895.10 QWEN3-14B74.2214.2373.9151.839.4556.0880.687.8185.80 GEMMA3-27B71.897.5379.6370.33-0.8374.7477.97-7.6182.56 QWEN3-32B76.2810.1074.7450.6911.9453.0083.215.1490.38 LLAMA-3.3-70B76.6815.1576.0859.167.5064.3793.333.6791.40 GPT-4O-MINI86.43-1.1190.2771.02-10.6180.2788.810.6293.36 Table 2: Cross-lingual comparison of Gender IoU (GIoU), Sentence-level Gender Accuracy Gap (△SGA), and Corpus-level Gender Accuracy (CGA) across 15 multilingual LLMs for French, Arabic, and Hindi. Higher GIoU and CGA indicate better gender understanding, while lower△SGA indicates reduced bias. that penalizes models for making unintended edits. These patterns are exemplified through results for the LLaMA family of models on multi-entity cases across languages. Illustrative cases for French, Ara- bic, and Hindi are presented in Figures 7, 8, and 9 in the Appendix, respectively. Thus, a large differ- ence between gender accuracy and GIoU indicates that models often transform non-gendered terms and suffer from gender interference and limited instruction following capability for this task. 5.5 French: Complex Morphology Amplifies Bias and Challenges Pronoun Agreement French’s larger dataset and complex morphology diluted performance, amplifying training imbal- ances, a trend evident in the GIoU scores presented in Figure 6a of appendix. Larger models exhibited masculine bias, while smaller models struggled sig- nificantly. Possessive pronoun agreement (e.g., son instructeur/son instructrice), requiring possession- based gender disambiguation, posed challenges. Smaller models lacked the morphological under- standing to handle this, whereas larger models per- formed more effectively, reflecting the impact of capacity on complex rule application. 5.6 Arabic: Lowest Scores with Persistent Masculine Bias in Plurals Arabic’s smaller, stricter dataset with intricate mor- phology yielded the lowest scores, as reflected in the GIoU scores in Figure 6b of appendix. Larger models mitigated complexity with balanced gen- der handling, while smaller models faltered, often showing masculine biases. Female plural agree- ment (e.g., ka-mumaththil ̄ at for “actresses”), de- faulting to masculine for female plural groups, highlighted inadequate training on gender-specific morphology, with most models over-applying mas- culine forms, even in all-female contexts. 5.7 Hindi: Feminine Skew and Entity Errors Models achieved higher performance on the Hindi dataset of theMORPHOGENbenchmark, reflecting its simpler morphology with fewer gender nuances, as illustrated in the GIoU scores in Figure 6c of appendix. Larger models demonstrated superior performance with minimal gender disparity, while smaller models remained competitive, underscor- ing Hindi’s accessibility. However, some models displayed a feminine bias in female-to-male conver- sions, and others showed weaker entity recognition due to erroneous gender modifications. Models in 8B–12B range exhibited stronger entity recog- nition abilities. Smaller models struggled on di- rect speech involving adjectives and occupations, and co-reference resolution (e.g., ́ siks . ak/ ́ siks . ik ̄ a for “teacher”) failing to resolve a speaker’s gender, un- like larger models with robust co-reference han- dling. 6 Conclusions and Future Work This paper introducedMORPHOGEN, a new multilin- gual benchmark for evaluating gender-aware mor- phological generation in LLMs, covering Hindi, French, and Arabic, three typologically diverse, gendered languages.MORPHOGENfocuses on a con- trolled first-person transformation task that isolates gender-sensitive morphological reasoning. We pro- posed novel evaluation metrics tailored to this set- ting and benchmarked 15 multilingual LLMs rang- ing from 2B to 70B parameters. Our results show models often confuse gendered forms, especially with multiple entities, and ex- hibit biased masculine-to-feminine vs. feminine- to-masculine transformations, with some models showing strong directional bias. This highlights persistent limitations in LLMs’ handling of gen- dered morphology. MORPHOGENoffers a foundation for studying mor- phological competence in multilingual models. Fu- ture work should expand it to include 2nd and 3rd person constructions, other gendered languages, and more complex discourse. Our work also en- ables developing gender-sensitive training and eval- uating bias in generative tasks like translation, sum- marization, and dialogue. 7 Limitations This work presentsMORPHOGEN, a large-scale, syn- thetic benchmark designed to evaluate multilingual language models on grammatical gender and mor- phological agreement across three typologically diverse and gendered languages: French, Arabic, and Hindi. While we believeMORPHOGENrepresents an important step toward more inclusive and lin- guistically grounded evaluation of LLMs, several limitations remain. First, the dataset currently covers only three lan- guages, each represented in a standardized form without accounting for dialectal variation. Specif- ically, we use Modern Standard Hindi, Standard Metropolitan French, and Modern Standard Ara- bic. French, Arabic, and Hindi each have dozens of dialects, many of which exhibit distinct gram- matical and lexical gender patterns, which are not yet included in this release. Second, our Arabic dataset is smaller than the others, primarily due to limited availability of high-quality source data and fewer native Arabic-speaking annotators. Third, both Hindi and Arabic are predominantly binary- gendered languages; consequently, our current dataset focuses only on male and female speaker forms. We recognize this binary framing as a limita- tion and aim to extend the dataset to better represent gender as a spectrum in future work. Finally, while we also introduce multi-entity scenarios to evaluate gender interference, these are currently limited to two human referents per sentence. Expanding to more complex discourse scenarios with multiple gendered entities remains an important direction for future research. Despite these limitations,MORPHOGENprovides a valuable and high-precision resource for advancing evaluation of how of LLMs across linguistically diverse settings. 8 Ethical Considerations WhileMORPHOGENaims to advance fairness and inclusivity by providing a gender-focused bench- mark for morphologically rich languages (French, Arabic, and Hindi), we recognize several ethical considerations regarding its development and ap- plication. First, our task formulation currently relies on the binary (masculine and feminine) grammatical cate- gories inherent to these languages, which does not encompass the full spectrum of gender identities. We plan to explore non-binary expansions in fu- ture iterations, guided by linguistic feasibility and community consultation. Additionally, to avoid reinforcing cultural or occupational stereotypes, we carefully curated prompts to actively challenge male-default biases (e.g., explicitly using feminine forms for roles like "doctor" or "leader"). Second, we acknowledge the general risk that improved grammatical coherence could be misused to generate harmful text. To mitigate this dual-use concern,MORPHOGENrelies exclusively on strictly synthetic prompts and neutral scenarios. Regarding data creation, annotations were com- pleted by undergraduate students (aged 18–21) who were fairly compensated and certified for their con- tributions. We prioritized annotator well-being by ensuring all tasks were completely free of sensitive, offensive, or personally identifiable content. Finally,MORPHOGENis released under a C BY- NC 4.0 license for research and non-commercial use, with the intent of helping the community build more equitable and linguistically inclusive NLP systems. 9 Acknowledgments We would like to acknowledge the Infosys Center for Artificial Intelligence (CAI) and IIIT-Delhi for their support during this research. We are also deeply grateful to Jagjot Singh, Akshit K Bansal, and Ankit Agarwal for their dedicated assistance in the annotation and creation of MORPHOGEN. References Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, and 1 others. 2024. Phi-4 technical re- port. arXiv preprint arXiv:2412.08905. Hend S. Al Khalifa, Abdulmohsen Al-Thubaity, Nora Al-Twairesh, Abdulrahman Alqahtani, and Abdullah Bahanshal. 2022. The arabic parallel gender cor- pus 2.0: Extensions and analyses. In Proceedings of the 13th International Conference on Language Resources and Evaluation (LREC). Avinash Anand, Arnav Goel, Medha Hira, Snehal Buldeo, Jatin Kumar, Astha Verma, Rushali Gupta, and Rajiv Ratn Shah. 2023. Sciphyrag - retrieval aug- mentation to improve llms on physics q & a. In Big Data and Artificial Intelligence: 11th International Conference, BDA 2023, Delhi, India, December 7–9, 2023, Proceedings, page 50–63, Berlin, Heidelberg. Springer-Verlag. Luisa Bentivogli, Beatrice Savoldi, Matteo Negri, Mat- tia Antonino Di Gangi, Roldano Cattoni, and Marco Turchi. 2020.Gender in danger?evaluating speech translation technology on the must-she corpus. Preprint, arXiv:2006.05754. Anna Currey, Maria Nadejde, Raghavendra Reddy Pap- pagari, Mia Mayer, Stanislas Lauly, Xing Niu, Ben- jamin Hsu, and Georgiana Dinu. 2022. MT-GenEval: A counterfactual and contextual dataset for evaluating gender accuracy in machine translation. In Proceed- ings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4287–4299, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Jay Gala, Pranjal A Chitale, Raghavan AK, Varun Gumma, Sumanth Doddapaneni, Aswanth Kumar, Janki Nawale, Anupama Sujatha, Ratish Puduppully, Vivek Raghavan, and 1 others. 2023. Indictrans2: Towards high-quality and accessible machine trans- lation models for all 22 scheduled indian languages. arXiv preprint arXiv:2305.16307. Arnav Goel, Medha Hira, Avinash Anand, Siddhesh Bangar, and Rajiv Ratn Shah. 2023. Advancements in scientific controllable text generation methods. Preprint, arXiv:2307.05538. Arnav Goel, Medha Hira, and Anubha Gupta. 2024a. Exploring multilingual unseen speaker emotion recognition: Leveraging co-attention cues in mul- titask learning. Preprint, arXiv:2406.08931. Arnav Goel, Medha Hira, and Anubha Gupta. 2024b. Multilingual prosody transfer: Comparing supervised & transfer learning. Preprint, arXiv:2406.00022. Hila Gonen, Yova Kementchedjhieva, and Yoav Gold- berg. 2019. How does grammatical gender affect noun representations in gender-marking languages? Preprint, arXiv:1910.14161. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulka- rni, Ujwal Gadiraju, Aditya Vashistha, Vivek Se- shadri, and Kalika Bali. 2024. Akal badi ya bias: An exploratory study of gender bias in hindi language technology. Preprint, arXiv:2405.06346. Medha Hira, Arnav Goel, and Anubha Gupta. 2024. Crossvoice: Crosslingual prosody preserving cascade-s2ST using transfer learning. In The Second Tiny Papers Track at ICLR 2024. Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task bench- mark for evaluating cross-lingual generalisation. In International conference on machine learning, pages 4411–4421. PMLR. Kaiyu Huang, Fengran Mo, Xinyu Zhang, Hongliang Li, You Li, Yuanchi Zhang, Weijian Yi, Yulong Mao, Jinchen Liu, Yuzhuang Xu, Jinan Xu, Jian-Yun Nie, and Yang Liu. 2025a. A survey on large language models with multilingualism: Recent advances and new frontiers. Preprint, arXiv:2405.10936. Xu Huang, Wenhao Zhu, Hanxu Hu, Conghui He, Lei Li, Shujian Huang, and Fei Yuan. 2025b. Benchmax: A comprehensive multilingual evaluation suite for large language models. Preprint, arXiv:2502.07346. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Ishika Joshi, Ishita Gupta, Adrita Dey, and Tapan Parikh. 2024. ’since lawyers are males..’: Examining implicit gender bias in hindi language generation by llms. arXiv preprint arXiv:2409.13484. Janak Kapuriya, Anwar Shaikh, Arnav Goel, Medha Hira, Apoorv Singh, Jay Saraf, Sanjana, Vaibhav Nauriyal, Avinash Anand, Zhengkui Wang, and Ra- jiv Ratn Shah. 2025. Enhancing scientific visual question answering via vision-caption aware super- vised fine-tuning. In Proceedings of the 2nd Interna- tional Workshop on Large Vision - Language Model Learning and Applications, LAVA ’25, page 13–30, New York, NY, USA. Association for Computing Machinery. P. Koehn. 2005. Europarl: A parallel corpus for statis- tical machine translation. In ACL Anthology, pages 79–86. Minwoo Lee, Hyukhun Koh, Minsung Kim, and Ky- omin Jung. 2024. Fine-grained gender control in machine translation with large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5416–5430, Mexico City, Mexico. Association for Computational Lin- guistics. Huihan Li, Arnav Goel, Keyu He, and Xiang Ren. 2025. Attributing culture-conditioned generations to pre- training corpora. Preprint, arXiv:2412.20760. Pierre Lison and Jörg Tiedemann. 2016. Opensub- titles2016: Extracting large parallel corpora from movie and tv subtitles. In Proceedings of the 10th In- ternational Conference on Language Resources and Evaluation (LREC). European Language Resources Association (ELRA). Hengyu Luo, Zihao Li, Joseph Attieh, Sawal Devkota, Ona de Gibert, Shaoxiong Ji, Peiqin Lin, Bhavani Sai Praneeth Varma Mantina, Ananda Sreenidhi, Raúl Vázquez, Mengjie Wang, Samea Yusofi, and Jörg Tiedemann. 2025. Gloteval: A test suite for mas- sively multilingual evaluation of large language mod- els. Preprint, arXiv:2504.04155. Viktor Mihaylov and Aleksandar Shtedritski. 2024. What an elegant bridge: Multilingual llms are bi- ased similarly in different languages.Preprint, arXiv:2407.09704. Andrea Piergentili, Beatrice Savoldi, Matteo Negri, and Luisa Bentivogli. 2024. Enhancing gender-inclusive machine translation with neomorphemes and large language models. Preprint, arXiv:2405.08477. Matúš Pikuliak, Andrea Hrckova, Stefan Oresko, and Marián Šimko. 2024. Women are beautiful, men are leaders: Gender stereotypes in machine translation and language modeling. Preprint, arXiv:2311.18711. Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S. Yu. 2024. Multilingual large language model: A survey of resources, taxonomy and fron- tiers. Preprint, arXiv:2404.04925. Nishat Raihan, Antonios Anastasopoulos, and Marcos Zampieri. 2024. mhumaneval–a multilingual bench- mark to evaluate large language models for code gen- eration. arXiv preprint arXiv:2410.15037. Argentina Anna Rescigno, Eva Vanmassenhove, Jo- hanna Monti, and Andy Way. 2020. A case study of natural gender phenomena in translation: A com- parison of google translate, bing microsoft translator and deepl for english to italian, french and spanish. CEUR Workshop Proceedings, 2769:62–90. Pub- lisher Copyright: Copyright © 2020 for this paper by its authors.; 7th Italian Conference on Computa- tional Linguistics, CLiC-it 2020 ; Conference date: 01-03-2021 Through 03-03-2021. Haneh Rhel and Dmitri Roussinov. 2025. Large lan- guage models and arabic content: A review. Preprint, arXiv:2505.08004. Aleix Sant, Carlos Escolano, Audrey Mash, Francesca De Luca Fornaciari, and Maite Melero. 2024. The power of prompts: Evaluating and mitigating gender bias in mt with llms. Preprint, arXiv:2407.18786. Beatrice Savoldi, Eleonora Cupin, Manjinder Thind, Anne Lauscher, Andrea Piergentili, Matteo Negri, and Luisa Bentivogli. 2025. mgente: A multilingual resource for gender-neutral language and translation. Preprint, arXiv:2501.09409. Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Di- nesh Tewari, and Partha Talukdar. 2024a. IndicGen- Bench: A multilingual benchmark to evaluate gen- eration capabilities of LLMs on Indic languages. In Proceedings of the 62nd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 11047–11073, Bangkok, Thai- land. Association for Computational Linguistics. Shivalika Singh, Angelika Romanou, Clémentine Four- rier, David I Adelani, Jian Gang Ngui, Daniel Vila- Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, and 1 others. 2024b. Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evalua- tion. arXiv preprint arXiv:2412.03304. Sunayana Sitaram, Adrian de Wynter, Isobel McCrum, Qilong Gu, and Si-Qing Chen. 2025. A multilingual, culture-first approach to addressing misgendering in llm applications. arXiv preprint arXiv:2503.20302. Eric Michael Smith, Isar Nejadgholi, Ahmad Beirami, and Byron C. Wallace. 2022. "i’m sorry to hear that": Finding new biases in language models with a holistic descriptor dataset. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2433–2448, Dublin, Ireland. Association for Computational Linguistics. Guijin Son, Dongkeun Yoon, Juyoung Suk, Javier Aula- Blasco, Mano Aslan, Vu Trong Kim, Shayekh Bin Islam, Jaume Prats-Cristià, Lucía Tormo-Bañuelos, and Seungone Kim. 2025. Mm-eval: A multilingual meta-evaluation benchmark for llm-as-a-judge and reward models. Preprint, arXiv:2410.17578. Gabriel Stanovsky, Noah A. Smith, and Luke Zettle- moyer. 2019. Evaluating gender bias in machine translation. In Proceedings of the 57th Annual Meet- ing of the Association for Computational Linguistics, pages 1679–1684, Florence, Italy. Association for Computational Linguistics. Kunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu, Gelei Deng, Shuai Li, Peigui Qi, Weiming Zhang, Tianwei Zhang, and Nenghai Yu. 2025. Gendercare: A comprehensive framework for assessing and reduc- ing gender bias in large language models. Preprint, arXiv:2408.12494. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupati- raju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, and 1 others. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118. NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Hef- fernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, and 20 others. 2022. No language left behind: Scal- ing human-centered machine translation. Preprint, arXiv:2207.04672. Minghao Wu, Weixuan Wang, Sinuo Liu, Huifeng Yin, Xintong Wang, Yu Zhao, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. 2025. The bitter lesson learned from 2,000+ multilingual bench- marks. Preprint, arXiv:2504.15521. Yuemei Xu, Ling Hu, Jiayi Zhao, Zihan Qiu, Kexin Xu, Yuqi Ye, and Hanwen Gu. 2025. A survey on multi- lingual large language models: corpora, alignment, and bias. Frontiers of Computer Science, 19(11). An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025.Qwen3 technical report.arXiv preprint arXiv:2505.09388. Jinman Zhao, Yitian Ding, Chen Jia, Yining Wang, and Zifan Qian. 2024. Gender bias in large language models across multiple languages. arXiv preprint arXiv:2403.00277. A Gender-Morphology Different languages express grammatical gender through distinct morphological patterns.An overview of these patterns is shown in Figure 3, with illustrative examples in Table 3. These pat- terns motivate our focus on three gendered lan- guages: French, Arabic, and Hindi. To support comparisons made in the results sec- tion, we define morphological complexity in terms of (i) the number of agreement targets (e.g., verbs, adjectives, determiners), (i) the regularity vs. ir- regularity of gender marking, and (i) the extent to which gender realization depends on syntactic context (e.g., tense, clause structure, or discourse configuration). Under this definition, French and Arabic both exhibit high morphological complex- ity, but for different reasons. French shows irregu- lar and context-dependent agreement (e.g., gender agreement in past participles but not in present tense), while Arabic exhibits a more systematic yet nuanced morphology with pervasive agreement across verbs, adjectives, pronouns, and construc- tions such as relative clauses, conditionals, and multi-entity contexts. In contrast, Hindi follows a comparatively more regular and semantically grounded system, with fewer context-dependent variations. For each of these languages, we provide example snippets along with the corresponding morphologi- cal rules in Tables 4, 5, 6, 7, and 8. B Annotator Guidelines Guidelines for Dataset Refinement •Off-Limits Language: No sentence may contain profanity, hate speech, slurs, or any other abusive or objectionable content. Annotators must ensure compliance with content-policy restrictions. •Naturalness: Sentences should reflect standard conversational phrasing that a fluent speaker would naturally use, avoiding stilted or machine- generated constructions. •Uniqueness: Identical or trivially paraphrased sentences are to be rejected and regenerated. •Template Fidelity: Sentences must follow the syntactic template exactly, without missing slots, extra words, or rearrangements. •Domain Coverage: For every template, sen- tences must span all conversational domains spec- ified (e.g., academic, healthcare, legal), ensuring diversity. •Gender Specificity: Each English sentence must be designed such that its translations differ be- tween masculine and feminine forms in the target language. Guidelines for Dataset Correction •Fidelity & Fluency: Translations must preserve meaning, tone, and register while being gram- matically correct and idiomatic in the target lan- guage. Annotators should check word choice, tense, punctuation, and readability. • Speaker-Gender Agreement:All gender- dependent morphology tied to the speaker (verbs, adjectives, pronouns, etc.) must appear in the cor- rect masculine form in the “male” version and the correct feminine form in the “female” version. (a) Hindi(b) French (c) Arabic Figure 4: Distribution of Sentence Frequency Per Morphological Rule for Each Language •Consistency for Implicit Gender Entities: Gen- dered terms referring to non-speaker entities must remain identical across male and female transla- tions. For instance, if friend is rendered in mas- culine form in the male version, it must remain masculine in the female version as well. Dataset Validation Process Each data sample was independently assigned a validity score of 1 or 0 by two annotator, indicating full compliance or non-compliance with the anno- tation guidelines, respectively. The Data Validation Score (DVS) and Inter-Annotator Agreement (IAA) were computed as follows: DVS = P N i=1 (s i1 + s i2 ) 2N (1) IAA = 1 N N X i=1 I (s i1 = s i2 )(2) where: • Ndenotes the total number of sentence pairs in the dataset. • s i1 ands i2 represent the binary validity scores (0or1) assigned to thei th sentence pair by Annotator 1 and Annotator 2, respectively. • I(·) is the indicator function, returning1if the condition inside is true and 0 otherwise. The Data Validation Score measures the overall proportion of valid samples across all annotators, while the Inter-Annotator Agreement quantifies the fraction of samples for which both annotators as- signed identical scores. C Model Hyperparameters and Compute Used For all models evaluated on theMORPHOGENbench- mark, we used a standardized inference config- uration to ensure consistency across generations. Table 3: Gender Morphology Overview Table 4: French Gendered Grammar Examples Across Rule Types [1] Table 5: French Gendered Grammar Examples Across Rule Types [2] Table 6: Arabic Gendered Grammar Examples Across Rule Types Table 7: Hindi Gendered Grammar Examples Across Rule Types [1] Table 8: Hindi Gendered Grammar Examples Across Rule Types [2] StatisticsArabicFrenchHindi Total Sentences Generated54131641510248 Sentences Discarded263864162694 Unique Sentences7610999910248 Data Validation Score0.97330.96510.9731 Inter-Annotator Score0.95260.93660.9594 Number of Annotators233 Table 9: Dataset Statistics Across Languages ModelGIoU↑ GIoU M ↑ GIoU F ↑ SGA↑ SGA M ↑ SGA F ↑△SGA↓ CGA↑ QWEN2.5-0.5B5.477.653.305.728.003.454.554.16 GEMMA2-2B39.7337.2942.1640.9038.3243.47-5.1437.54 LLAMA-3.2-3B54.4960.1948.8559.2064.9453.5211.4253.48 GEMMA3-4B52.7046.4959.0957.7250.7464.91-14.1651.60 QWEN3-4B58.6461.5955.7660.9864.6457.407.2553.20 LLAMA-3.1-8B67.8970.6764.6282.7584.4480.753.6981.76 QWEN3-8B71.6673.8969.3976.2578.6673.794.8669.91 GEMMA2-9B60.5262.0259.0265.4866.1164.841.2655.56 GEMMA3-12B64.2764.3464.2076.3376.0476.62-0.5874.26 PHI4-14B79.8481.4678.2289.6890.2689.091.1787.70 QWEN3-14B74.2280.6467.4878.7885.7371.4914.2373.91 GEMMA3-27B71.8975.4768.1183.1186.7879.257.5379.63 QWEN3-32B76.2880.8071.7679.3584.4074.3010.1074.74 LLAMA-3.3-70B76.6883.5369.8180.7688.3373.1715.1576.08 GPT-4O-MINI86.4386.6186.2591.7791.2292.33-1.1190.27 Table 10: Performance metrics of different models on French (% values; GIoU = Gender IoU, CGA = Corpus-Level Gender Accuracy, SGA = Sentence-Level Gender Accuracy,△SGA = Accuracy Gap, M = Male, F = Female) The input prompt was constructed using the model- specific chat template, and all models were queried in a zero-shot setting without any few-shot exam- ples. Generation Parameters.We used the following generation hyperparameters for all models, unless otherwise noted: •Sampling Strategy: Deterministic (no sam- pling) • do_sample: False • Max New Tokens: 256 •Temperature: 0.1 (low temperature for con- trolled and accurate generations) • Top-p: 0.95 • Top-k: Not used (default) • Num Return Sequences: 1 • Batch Size for Inference: 1 (due to varied token limits across models) All generations were performed with: • eos_token_id: Set to the tokenizer’s EOS to- ken •pad_token_id: Set to the tokenizer’s PAD token if defined, else fallback to EOS Compute Infrastructure.All experiments were run on an NVIDIA DGX A100 server equipped with 8 NVIDIA A100 GPUs, each with 40GB VRAM. While most models were executed using a single A100 GPU, larger models (e.g., mixture-of- experts or 65B+ parameter class) were distributed across multiple GPUs as needed via tensor or model parallelism. This setup ensured sufficient compute head- room for large-scale inference and supported par- allelized benchmarking across multiple languages and prompts. D Prompts D.1 Sentence Generation system_prompt ="Suppose you are an Expert English Sentence Generating System." user_prompt = ModelGIoU↑ GIoU M ↑ GIoU F ↑ SGA↑ SGA M ↑ SGA F ↑△SGA↓ CGA↑ QWEN2.5-0.5B4.147.310.726.2810.371.888.494.59 GEMMA2-2B14.7314.1415.3016.0415.6316.43-0.8114.10 LLAMA-3.2-3B18.315.9629.9620.956.7434.35-27.6117.75 GEMMA3-4B45.6845.3146.0655.3451.2359.43-8.2048.93 QWEN3-4B34.3434.0734.5937.6337.1738.07-0.9035.97 LLAMA-3.1-8B43.5144.5342.4950.6551.1350.170.9645.51 QWEN3-8B45.8947.9343.8951.4453.9948.965.0351.01 GEMMA2-9B46.4547.9244.9950.4351.7149.162.5545.26 GEMMA3-12B62.7664.6960.8269.3770.6268.122.5065.52 PHI4-14B57.0862.2452.2066.5169.8963.316.5866.15 QWEN3-14B51.8356.0747.7357.4862.2952.849.4556.08 GEMMA3-27B70.3371.4769.1977.1276.7077.53-0.8374.74 QWEN3-32B50.6957.3244.1056.5762.5650.6211.9453.00 LLAMA-3.3-70B59.1663.5054.8466.8470.6163.117.5064.37 GPT-4O-MINI71.0268.1373.9182.7677.4588.06-10.6180.27 Table 11: Performance metrics of different models on Arabic (% values; GIoU = Gender IoU, CGA = Corpus-Level Gender Accuracy, SGA = Sentence-Level Gender Accuracy,△SGA = Accuracy Gap, M = Male, F = Female) ModelGIoU↑ GIoU M ↑ GIoU F ↑ SGA↑ SGA M ↑ SGA F ↑△SGA↓ CGA↑ QWEN2.5-0.5B0.350.690.050.350.690.050.630.21 GEMMA2-2B71.4175.1367.8576.2880.0472.697.3565.41 LLAMA-3.2-3B48.5419.8576.9053.0820.5785.22-64.6549.72 GEMMA3-4B67.5060.2973.2171.7563.7678.08-14.3264.58 QWEN3-4B62.8461.9663.7073.7475.4172.083.3368.51 LLAMA-3.1-8B83.1284.0182.2391.6591.4491.87-0.4389.21 QWEN3-8B80.9682.3679.5791.5292.6090.432.1687.82 GEMMA2-9B85.4782.7888.2087.4283.7891.12-7.3484.39 GEMMA3-12B79.9175.6984.1684.9380.4889.41-8.9380.99 PHI4-14B82.7784.6980.8596.6997.3896.001.3895.10 QWEN3-14B80.6883.6277.7490.2294.1286.307.8185.80 GEMMA3-27B77.9775.6880.3183.9680.3487.96-7.6182.56 QWEN3-32B83.2185.8680.5693.8896.4691.315.1490.38 LLAMA-3.3-70B93.3395.0491.6294.0695.8992.223.6791.40 GPT-4O-MINI88.8190.0887.5495.7396.0495.420.6293.36 Table 12: Performance metrics of different models on Hindi (% values; GIoU = Gender IoU, CGA = Corpus-Level Gender Accuracy, SGA = Sentence-Level Gender Accuracy,△SGA = Accuracy Gap, M = Male, F = Female) Generate <Num_Sentences> English sentences. Strictly adhere to the format: <Template> Instructions: 1. Only output the sentences—do not include any additional text. 2. Each sentence must be unique in its context and the nouns used. 3. Vary the sentence lengths and ensure they sound natural and conversational. 4. Use a variety of creative contexts, including but not limited to [<Context_1>, <Context_2>, ..., <Context_n>]. The prompts are designed to guide a language model in generating diverse and natural-sounding English sentences. The system prompt establishes the model’s role, while the user prompt provides clear, structured instructions to ensure variety, con- textual relevance, and adherence to a specified for- mat. D.2 Zero Shot Prompts For zero-shot inference of the LLMs on the MORPHOGENbenchmark, we designed language- specific prompts to ensure precise gender-aware morphological transformations while preserving sentence structure. Although the prompts were provided to the models in the respective target lan- guages (Hindi, French, or Arabic), the structure and content of the system and user prompts were consistent across languages. The system prompt given to the model was: “You are a language assistant. Given a sentence in the target language and the gender of the speaker, adjust only the words that refer to the speaker to match the specified gender. Do not alter any other parts of the sentence. Return only Figure 5:△SGA(Accuracy Gap) across all models and languages (French, Arabic, Hindi) in theMORPHOGEN benchmark. Positive values indicate masculine bias, while negative values indicate feminine bias. the modified sentence with no explana- tions or extra words. If no change is required, return the sentence exactly as it is.“ The user prompt provided the transformation instruction, depending on the speaker’s gender: • For male speakers:Without changing the structure of the sentence, convert it as if it were spoken by a male speaker. • For female speakers:Without changing the structure of the sentence, convert it as if it were spoken by a female speaker. This was followed by the sentence to be trans- formed:Sentence to transform: [sentence]. These prompts were designed to enforce min- imal intervention, focusing solely on speaker- referring terms. This ensures the task evaluates the models’ ability to perform gender-specific transfor- mations without altering unrelated components of the sentence. The zero-shot setting tests the models’ inherent linguistic knowledge, aligning with the benchmark’s goal of assessing gender-aware mor- phological capabilities across diverse languages. E Detailed Results and Error Analysis This section provides a comprehensive breakdown of model performance on theMORPHOGENbench- mark, combining dataset statistics, rule-level quan- titative metrics, and a qualitative error analysis of the LLAMA model family. E.1 Quantitative Performance and Bias Analysis Table 9 details the dataset validation statistics, while the aggregate performance metrics across all models for French, Arabic, and Hindi are presented in the respective language tables. To better understand directional bias, Figure 5 vi- sualizes the Accuracy Gap (∆SGA) across all mod- els. A clear trend emerges regarding model scale and gender bias. Smaller models exhibit erratic and often extreme bias gaps. For instance, smaller architectures like LLAMA 3.2 3B demonstrate massive fluctuations, including severe feminine bias (∆SGA < 0) in Hindi and Arabic, contrasted with masculine bias in French. As model capacity increases (e.g., LLAMA 3.3 70B and GPT-4O- MINI), the∆SGAconverges closer to zero across all languages, indicating a more balanced, unbiased linguistic understanding rather than a reliance on statistical gender defaults. E.2 Rule-Level Morphological Competence To isolate where models succeed or fail, Fig- ure 6 presents the Gender Intersection over Union (GIoU) metrics broken down by specific grammati- cal rules for all 3 languages. Across all three languages, models generally per- form well on localized, simple rules (e.g., basic ad- jectives). However, performance sharply degrades on complex syntactical structures that require long- (a) French(b) Arabic (c) Hindi Figure 6: Rule-based and model-wise IoU metrics across all three languages. Figure 7: Example of results of LLAMA family of models on multiple entities in French dataset. Figure 8: Example of results of LLAMA family of models on multiple entities in Arabic dataset. Figure 9: Example of results of LLAMA family of models on multiple entities in Hindi dataset. range dependency tracking, such as indirect speech, pseudo-cleft sentences, and multiple entities. The gap between smaller and larger models is most pro- nounced in these complex categories, highlighting that parameter scale is crucial not just for vocabu- lary, but for maintaining morphological consistency across extended contexts. E.3 Qualitative Error Analysis: A LLAMA Case Study To ground these quantitative findings, Figures 7, 8, and 9 present a qualitative error analysis focus- ing on "gender interference" in complex sentences containing multiple entities. We compare the out- puts of LLAMA 3.2 3B, LLAMA 3.1 8B, and LLAMA 3.3 70B to illustrate the evolution of mor- phological control. Further structural examples of grammatical rules for each language can be found in Tables 5 through 8. LLaMA 3.2 3B.The smallest model consistently struggles to correctly apply gendered inflections, often defaulting to standard forms regardless of the intended speaker gender. This behavior is evident across all three languages, where the model fails to align verbs, adjectives, or participles with the correct grammatical gender. As seen in the quali- tative examples, these persistent errors in speaker gender realization heavily penalize its overall SGA and GIoU scores, and explain the extreme∆SGA variations observed in Figure 5. LLaMA 3.1 8B. The 8B model demonstrates a clear improvement in capturing basic gender mor- phology, particularly in correctly inflecting verbs and adjectives immediately adjacent to the speaker. However, it suffers from overgeneralization, lead- ing to severe gender interference. As illustrated in Figures 7 through 9, while the primary speaker’s gender is correctly realized, the model incorrectly alters the grammatical gender of other entities in the sentence to match the speaker. This indicates a partial understanding of agreement rules but insuf- ficient syntactic control over entity-specific bound- aries. Consequently, while its SGA scores improve relative to the 3B model, its GIoU remains sup- pressed due to these collateral inflection errors. LLaMA 3.3 70B. The largest model demon- strates robust performance across all cases, cor- rectly applying gender transformations while pre- serving agreement boundaries between entities. It maintains a clear distinction between speaker- specific and non-speaker-specific gender mark- ing, inflecting only relevant tokens without "bleed- ing" onto adjacent nouns, resulting in outputs that closely match reference sentences structurally and morphologically. Accordingly, it achieves the high- est SGA and GIoU scores (Fig 6) with minimal bias gaps, reflecting highly accurate gender realization across complex grammatical scenarios.