Paper deep dive
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models
Alaa Elobaid
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/14/2026, 1:59:26 AM
Summary
This paper evaluates demographic and linguistic biases in four omnimodal language models (Gemini 2.5 Flash, Gemma 3n, Qwen 2.5 Omni, and Phi-4 Multimodal) across text, image, audio, and video modalities. The study finds that while image and video tasks show relatively high performance and smaller demographic disparities, audio tasks exhibit significant performance gaps, substantial bias, and frequent prediction collapse, particularly across age, gender, and language groups.
Entities (6)
Relation Signals (3)
Gemini 2.5 Flash → evaluatedon → Casual Conversations V2
confidence 100% · Four omnimodal models are evaluated on tasks that include... CC2 dataset
Audio understanding tasks → exhibits → Demographic Bias
confidence 100% · audio understanding tasks exhibit significantly lower performance and substantial bias
Omnimodal Language Models → processes → Audio
confidence 100% · OLMs are capable of processing three or more modalities including text, vision, and audio
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper provides a comprehensive evaluation of demographic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their performance across different demographic groups and modalities is not well studied. Four omnimodal models are evaluated on tasks that include demographic attribute estimation, identity verification, activity recognition, multilingual speech transcription, and language identification. Accuracy differences are measured across age, gender, skin tone, language, and country of origin. The results show that image and video understanding tasks generally exhibit better performance with smaller demographic disparities. In contrast, audio understanding tasks exhibit significantly lower performance and substantial bias, including large accuracy differences across age groups, genders, and languages, and frequent prediction collapse toward narrow categories. These findings highlight the importance of evaluating fairness across all supported modalities as omnimodal language models are increasingly used in real-world applications.
Tags
Links
- Source: https://arxiv.org/abs/2604.10014v1
- Canonical: https://arxiv.org/abs/2604.10014v1
Trouble viewing inline? Open PDF directly →
Full Text
79,118 characters extracted from source content.
Expand or collapse full text
Demographic and Linguistic Bias Evaluation in Omnimodal Language Models Alaa Elobaid 1[0009-0009-2399-9754] Freie Universit ̈at Berlin, Berlin, Germany alaa.elobaid@fu-berlin.de Abstract. This paper provides a comprehensive evaluation of demo- graphic and linguistic biases in omnimodal language models that process text, images, audio, and video within a single framework. Although these models are being widely deployed, their performance across different de- mographic groups and modalities is not well studied. Four omnimodal models are evaluated on tasks that include demographic attribute es- timation, identity verification, activity recognition, multilingual speech transcription, and language identification. Accuracy differences are mea- sured across age, gender, skin tone, language, and country of origin. The results show that image and video understanding tasks generally exhibit better performance with smaller demographic disparities. In contrast, audio understanding tasks exhibit significantly lower performance and substantial bias, including large accuracy differences across age groups, genders, and languages, and frequent prediction collapse toward narrow categories. These findings highlight the importance of evaluating fair- ness across all supported modalities as omnimodal language models are increasingly used in real-world applications. Keywords: Multimodal models· Demographic bias· Biometrics. 1 Introduction Omnimodal Language Models (OLM) are a subset of Multimodal Language Models (MLM) that can process more than two modalities simultaneously. In contrast to unimodal Language Models (LM) that process exclusively textual input, or bimodal LMs such as Vision Language Models (VLM) (text and vi- sion) and audio LMs (text and audio), OLMs are capable of processing three or more modalities including text, vision, and audio within a unified framework. The term "omnimodal" was popularized by OpenAI with the release of GPT-4o [15] and has since been adopted by Amazon with Nova 2 Omni [17] and Alibaba with Qwen2.5-Omni [36]. However, bias evaluation research in MLMs has primarily focused on eval- uating bimodal models such as VLMs and Audio LMs on a single modality [22,23,20,21]. This paper aims to evaluate the demographic bias in the form of accuracy disparities between demographic groups such as age, skin colour, and gender in arXiv:2604.10014v1 [cs.CV] 11 Apr 2026 2A. Elobaid four OLMs by quantifying it using a combination of visual and audio tasks: de- mographic attribute estimation from both image and audio, multilingual speech transcription and language identification, image-based identity verification, and activity recognition from video. 2 Related work MLMs have evolved from modality specific (bimodal) models to omnimodal mod- els [18] capable of processing combinations of visual, auditory, textual, and other modalities within unified frameworks. While recent research has addressed tech- nical challenges in OLMs such as modality bias, where models disproportionately attend to dominant modalities due to training data imbalances [6], demographic bias has received far less attention. Existing demographic bias evaluation work in MLMs has focused predominantly on vision language models, examining social biases across gender, race, and age [22,23,7,14]. However, comparable evaluation for OLMs, particularly those with audio processing capabilities, is largely miss- ing. This gap is critical as OLMs capable of processing both vision and audio are increasingly deployed in real-world applications such as multimodal age verifica- tion [5] and multimodal healthcare assistants [2] without a clear understanding of how they perform across demographic groups. While this study focuses on OLMs that process multiple modalities simul- taneously, it is important to distinguish these from video LMs such as LLaVA- Video [37]. Video LMs employ explicit temporal modeling mechanisms including temporal attention modules and embeddings to capture inter-frame relation- ships, prioritizing depth of temporal visual understanding. However, when these video specialized models are extended to process audio, a recent study shows that they suffer from cross modal hallucinations and do not support the level of audio understanding needed for nuanced audio processing tasks [33]. In con- trast, OLMs are architecturally designed with native audio encoders that enable nuanced audio understanding tasks including multilingual speech transcription, speaker identification, and acoustic scene analysis alongside visual processing. Therefore, they are the focus of this work. The remainder of this literature review examines existing research on demo- graphic bias in VLMs and traditional automated speech recognition systems, establishing the foundation for understanding bias in OLMs. Narayan et al. [22] created the FaceXBench benchmark and evaluated 26 open-source and 2 proprietary VLMs on 14 face understanding tasks including demographic attribute estimation such as age, race, and gender. However, formu- lating bias and fairness solely as accuracy in demographic attribute estimation ignores how the model’s performance varies across different demographic groups. This age estimation, gender prediction, and race estimation approach to bias as- sessment was also adopted by Shahreza et al. [30], who developed a specialized VLM for face understanding tasks and evaluated it on the same FaceXBench benchmark [22]. Demographic and Linguistic Bias Evaluation in Omnimodal LMs3 Perera et al. [23] address this limitation by developing a benchmark of 10,000 VQA-based questions for attribute estimation and reporting model performance across demographic groups. Nonetheless, their study maintains the same single- task formulation of bias. Beyond demographic prediction accuracy, some works have investigated social representation bias in how VLMs portray different de- mographic groups [7,14], though such qualitative bias assessment falls outside the scope of this study, which focuses on quantifiable performance disparities across demographic groups. Bias in Automated Speech Recognition (ASR) is a well-documented phe- nomenon across traditional ASR models such as Whisper [26], wav2vec 2.0 [4], and Massively Multilingual Speech (MMS) [25]. Early work by Feng et al. [9] quantified bias in Dutch ASR systems across gender, age, regional accents, and non-native accents through phoneme-level error analysis. Subsequent research has included diverse linguistic contexts, with studies examining Portuguese [19], Italian dialects [31], and African low-resource languages [21,16]. Research has consistently revealed performance disparities, with minority dialects and non- native speakers experiencing higher Word Error Rates (WER) compared to stan- dard language counterparts [28,29,34]. However, despite extensive bias quantification research in ASR-specific mod- els, bias in MLMs with audio processing capabilities remains largely unexplored. Additionally, integrating audio processing capabilities into MLMs introduces bias manifestations that are fundamentally different from those in traditional ASR- specific models. For example, the phoneme-level error analysis methodology that has been central to traditional ASR bias research [10] is inapplicable to MLMs. Unlike traditional ASR architectures, MLMs operate as end-to-end systems that process audio through learned latent representations without explicit phoneme- level outputs, making fine-grained linguistic analysis unachievable. As a result, bias evaluation in OLMs and other audio-capable LMs must focus on down- stream task performance rather than intermediate linguistic representations, re- lying on higher-level metrics such as WER or Character Error Rate for ASR tasks. Despite this limitation, audio-capable LMs offer unique zero-shot speech understanding capabilities beyond traditional speech recognition [35], enabling tasks such as demographic attribute classification, language identification, and topic recognition without task-specific training. 3 Experimental setup Four representative OLMs are evaluated on four benchmark datasets spanning image, audio, and video modalities. The selected datasets provide comprehensive demographic attribute estimation, face verification, speech recognition, language identification, and activity recognition tasks, while enabling evaluation across diverse demographic groups and linguistic contexts. 4A. Elobaid 3.1 Datasets To test the models across all three modalities and the devised tasks, the following four datasets are selected: Casual Conversations V2 (C2) [24] is a large dataset for benchmarking multimodal AI systems on fairness and robustness. It contains 26,467 videos from 5,567 participants across seven countries, totaling 674 hours. The dataset pro- vides extensive annotations, including age, gender, language, geo-location, skin tone, activity, and audio transcriptions. Videos consist of either scripted readings from Dostoevsky’s The Idiot or nonscripted responses to preset questions. This breadth enables diverse tasks such as demographic attribute estimation, activity recognition, language identification and ASR across multiple languages. Casual Conversations V1 (C1) [13] consists of 45,186 videos from 3,011 participants, with an average video length of approximately 1 minute totaling 846 hours of content. The dataset is annotated for age, gender, skin type, and lighting conditions. C1 mainly involves English-language videos and is targeted at ASR and demographic attribute estimation tasks in vision and audio modalities. Balanced Faces in the Wild (BFW) [27] is a dataset comprising 20,000 facial images of 800 identities, annotated for race, age, and gender. BFW focuses on facial ID verification and demographic attribute estimation in images. Mozilla Common Voice (MCV) [3] is a large-scale multilingual speech dataset maintained by Mozilla. MCV 22.0 is used, which includes 33,815 hours of speech recordings from thousands of contributors across 137 languages, with annotations for transcription, language, age, and gender. Audio clips span 1 to 15 seconds. It is widely used for benchmarking ASR and language identification. 3.2 Tasks Image-Based Tasks: Models are evaluated on seven image understanding tasks using the C1, C2, and BFW datasets. Age classification requires models to predict age as an integer value, with accuracy computed using a±5 year toler- ance during evaluation, acknowledging the inherent difficulty of exact age pre- diction from visual appearance alone. Fitzpatrick skin tone classification also employs a±1 category tolerance during evaluation due to the challenging na- ture of skin type classification, where even trained dermatologists report only moderate agreement [12]. Gender classification and country prediction evaluate a model’s ability to infer binary gender and country of origin from faces. Face ver- ification, evaluated on the BFW dataset, tests the model’s ability to determine whether two facial images belong to the same person. Video-Based Tasks: Video tasks assess models’ understanding of human actions and body visibility using the C2 dataset. Action classification requires models to identify the person’s physical activity from six categories: rotating, standing, sitting, walking, laying, or waving. Visibility classification evaluates the model’s ability to determine body visibility from four categories: only head visible, upper body visible, full body visible, or lower body visible. Demographic and Linguistic Bias Evaluation in Omnimodal LMs5 To assess temporal understanding, evaluation clips that include behavior transitions are constructed. For each action or visibility annotated segment, an adjacent segment of equal length with a different label is identified and con- catenated into a single clip. This approach tests whether models can correctly distinguish when behaviors change over time. Audio-Based Tasks: Audio tasks evaluate the models’ multilingual speech understanding capabilities. Age and gender classification from audio requires models to predict those attributes solely from voice characteristics (with a±5 year tolerance for age). Language identification tests the model’s ability to rec- ognize the language spoken in audio across multiple languages including En- glish, Spanish, Portuguese, Hindi, Tagalog, Indonesian, Telugu, Tamil, and Viet- namese. Speech transcription evaluates ASR capabilities using Word Accuracy (WA) as the primary metric, which measures the proportion of words correctly recognized relative to a reference transcript. Word Accuracy is calculated as W A = 100× 1− S + D + I N where S is the number of substitutions, D the deletions, I the insertions, and N the total words in the reference transcript. 3.3 Models Gemini 2.5 Flash (proprietary) [8] is an efficient model from Google’s Gem- ini family of highly capable multimodal large language models, optimized for low-latency reasoning across text, images, audio, and video. It features a native multimodal encoder trained end-to-end for unified omnimodal understanding and rapid generation. Gemma 3n (open-weights) [11] is an efficient MLM developed by Google DeepMind that processes both visual and auditory inputs alongside text. The model employs a MobileNet-V5-300M encoder for visual feature extraction. For audio processing, Gemma 3n uses a 0.68B parameter encoder based on the Uni- versal Speech Model (USM) conformer-based architecture. The E2B configura- tion is selected for evaluation, with 5.44B total and 1.91B effective parameters. Qwen 2.5 Omni (open-weights) [36] is a MLM that integrates vision, au- dio, and text modalities in an end-to-end architecture. For visual encoding, the model utilizes an adapted version of the Qwen2.5-VL vision encoder, while audio processing is handled by a modified Whisper-large-v3 encoder with 1.55B pa- rameters. This study evaluates the 3B parameter variant, which offers a balance between performance and efficiency for multimodal understanding tasks. Phi-4 Multimodal (open-weights) [1] is Microsoft’s MLM that processes images and audio in addition to text using a mixture-of-LoRAs architecture. The model incorporates SigLIP-400M for image encoding, and employs a custom Conformer with 460M parameters for audio feature extraction. Built on the Phi- 4-Mini base, the complete Phi-4 Multimodal model totals 5.6B parameters. 6A. Elobaid 3.4 Evaluation Protocol All models are evaluated using structured JSON prompts that specify the ex- pected output format with prediction-confidence pairs for each task. Full-size original images and audio files are passed to the models without preprocessing. Complete prompts are provided in Appendix B. For categorical attributes (e.g., gender, language, country), prompts provide an enumerated list of valid discrete values from which models must select. For continuous attributes (e.g., age), prompts request integer predictions. Prompts constrain model outputs to response categories aligned with dataset annotations. For speech transcription tasks, prompts specify that transcribed text must use native script characters rather than romanized transliterations, ensuring proper evaluation of multilingual capabilities. Models are instructed to respond only with valid JSON, facilitating parsing and evaluation at scale. 3.5 Reproducibility and null prediction handling To promote reproducibility, greedy decoding is used by setting the sampling temperature to 0 and top-p to 1, ensuring that the model consistently selects the highest-probability token at each generation step [32]. Null predictions are defined as outputs where the model fails to generate valid responses or returns explicit refusal indicators. Null predictions are included in the total sample count, effectively counting as incorrect predictions. Notably, null predictions occur exclusively in audio tasks. 4 Results and discussion This section presents the demographic performance evaluation of the four OLMs presented in Section 3.3 across image, video, and audio understanding tasks. Ac- curacy is reported as the primary performance metric and bias is quantified as the standard deviation of accuracy across demographic groups. For each demo- graphic group, 120 samples are used for audio-based and image-based tasks, while 20 samples are used for video-based tasks. Demographic and Linguistic Bias Evaluation in Omnimodal LMs7 4.1 Overall Accuracy Table 1. Accuracy Across All Tasks, Models, and Datasets (All values are percentages) Task CategoryTaskDatasetMetricGeminiPhiQwenGemma Image-Based Age Classification C1Accuracy54.232.538.551.2 C2Accuracy63.142.846.552.5 Fitzpatrick skin C1Accuracy90.453.951.244.3 C2Accuracy87.849.950.043.4 Gender Classification C1Accuracy96.795.495.998.8 C2Accuracy97.548.895.094.4 Country PredictionCC2Accuracy83.217.538.935.9 Face VerificationBFWAccuracy90.069.570.575.8 Video-BasedActionCC2Accuracy84.261.676.477.9 VisibilityCC2Accuracy51.850.348.253.2 Audio-Based Age Classification C2Accuracy36.025.630.636.4 MCVAccuracy26.419.443.823.6 Gender Classification C2Accuracy99.655.591.580.9 MCVAccuracy87.521.988.737.9 Language ID C2Accuracy99.749.487.395.1 MCVAccuracy97.1100100100 Speech Transcription C2WA83.9734.8643.5769.35 MCVWA72.15-5.8152.1240.40 Models demonstrate variable performance across image tasks, with scores gen- erally exceeding 50% and reaching as high as 98.8% for gender classification. Although task-specific baseline results are not included, the relative differences in model performance across tasks, datasets, and modalities provide insights into task difficulty and model behaviour. Gemini leads in most image tasks, achiev- ing the highest scores in Fitzpatrick skin classification (87.8% on C2), gender classification (97.5% on C2), country prediction (83.2%), and face verification (90.0%). Gemma achieves results comparable to Gemini across age and gender classification, and face verification. The C2 dataset consistently yields better results than C1 for comparable tasks, suggesting dataset quality differences. Gender classification is the easiest vision task with all models except Phi con- sistently achieving over 92% accuracy, while country prediction on C2 proves challenging with scores ranging from 17.5% to 83.2%. For video tasks, Gemini maintains the highest performance in action recognition (84.2%), while visibility classification proves challenging for all models with accuracies around 50%. In contrast to vision tasks, audio tasks reveal significantly weaker results across models, with most scores falling below 50%. Gemini demonstrates sub- stantially stronger audio understanding capabilities compared to other models, achieving near-perfect language identification (99.7% on C2, 97.1% on MCV), high gender classification accuracy (99.6% on C2, 87.5% on MCV), and the best speech transcription performance (83.97% WA on C2, 72.15% on MCV). The dataset characteristics appear to strongly influence results, with Qwen, Gemma, 8A. Elobaid and Gemini performing better on C2’s longer audio clips than MCV’ short clips. However, Phi shows an opposite pattern, performing better on MCV for language identification (100%), which likely indicates data leakage from its train- ing on "20k hours selected public transcribed" data [1] that may have included MCV samples. Gender classification in audio reveals notable model performance differences. Qwen achieves consistent results across both datasets (91.5% C2, 88.7% MCV), whereas Gemma achieves lower results on C2 (80.9%) with a notable decline on MCV (37.9%). Phi’s results consistently remain below 56%. 4.2 Image-Based Tasks Face Verification - BFW (Table 2) Table 2. Face Verification Accuracy by Demographic Group ModelAFAMBFBMIFIMWFWMAvgStdStd (Race)Std (Gender) Gemini80.191.194.587.587.993.091.594.590.04.52.71.5 Gemma69.380.082.076.364.277.475.881.275.85.83.33.0 Qwen58.064.076.073.074.577.573.068.070.56.25.80.1 Phi466.075.072.561.555.071.573.581.069.57.75.22.8 All values are percentages. A: Asian. B: Black. I: Indian. W: White. F: Female. M: Male. In face verification, Gemini demonstrates the highest performance (90.0% average) with the lowest total Std (4.5%). The model shows particularly strong performance for Black females (94.5%) and White males (94.5%), while Asian females represent the weakest group (80.1%). The gender-based standard devi- ation is relatively low across models (ranging from 0.1% to 3.0%), with Qwen achieving near-parity (Std Gender: 0.1%). Race-based disparities are more pro- nounced than gender disparities across all models, as evidenced by higher Std (Race) values compared to Std (Gender). Asian females consistently represent the lowest-performing demographic group across three models (80.1% for Gem- ini, 69.3% for Gemma, 58.0% for Qwen). Phi exhibits the highest overall disparity (Std: 7.7%) with particularly poor performance on Indian females (55.0%) and Black males (61.5%). Qwen shows the most severe race-based bias (Std Race: 5.8%) but achieves near gender parity. Gemma demonstrates moderate bias with near-balanced race (3.3%) and gender (3.0%) standard deviations. C1 Image (Table 3) Age classification: Gemini has the lowest age bias (Std: 8.2%) compared to Phi (15.2%), Qwen (22.4%), and Gemma (16.8%). Open-source models exhibit systematic bias against older adults, with 70+ accuracy of only 8.3-31.7%, while Gemini maintains 40.2-65.8% across age groups. Qwen and Phi show extreme prediction concentration (36.2% at 30-39 and 57% at 18-29), whereas Gemini and Gemma distribute predictions more evenly (4.3-23.1% and 6.5-26.0%). Demographic and Linguistic Bias Evaluation in Omnimodal LMs9 Table 3. Per-Attribute Accuracy with Top Predictions (C1, 120 samples per group) Attr.GroupPhiQwenGemmaGemini AccPrdTopAccPrdTopAccPrdTopAccPrdTop Age 18–2957.557.530(53)60.89.930(49)74.215.325(54)55.020.426(47.5) 30–3943.343.330(68)66.736.230(68)64.226.035(39)48.311.843(23.3) 40–4921.721.730(48)42.521.845(38)38.312.635(22)60.022.543(28.3) 50–5937.537.550(45)28.310.045(50)53.319.355(39)55.817.254(20.0) 60–6922.522.550(43)24.219.960(38)45.820.360(29)65.823.167(17.5) 70+12.512.560(63)8.32.160(65)31.76.565(39)40.24.367(26.5) Overall32.5–38.5–51.2–54.2– Std15.2–22.4–16.8–8.2– Gender Female93.347.9F(93)94.248.0F(94)99.250.4F(99)94.247.7F(94.2) Male97.552.1M(98)97.552.0M(98)98.349.6M(98)99.252.3M(99.2) Overall95.4–95.9–98.8–96.7– Std3.0–2.3–0.6–2.5– Skin tone 154.20.02(54)0.00.03(100)60.046.41(47)78.30.02(78.3) 2100.029.63(60)100.00.03(100)98.311.51(59)90.828.92(63.3) 3100.070.43(64)100.096.23(100)39.224.41(61)96.718.23(40.8) 466.70.03(67)100.03.13(98)31.717.61(57)96.727.84(74.2) 50.00.03(91)5.00.73(95)36.70.04(37)95.818.55(49.2) 60.00.03(95)2.50.03(85)0.00.04(58)84.26.75(52.5) Overall53.9–51.2–44.3–90.4– Std42.8–49.1–35.2–7.0– Acc: Accuracy (%) with tolerance (±5 years for age,±1 for skin tone). Prd: Percentage of samples predicted as the attribute group. Top: Most frequent prediction with percentage of samples receiving that prediction in brackets. Gender classification: All models demonstrate minimal gender accuracy disparity (Std: 0.6-3.0%), indicating balanced performance. However, for every model except Gemma, accuracy is higher for males by 2 to 5%. Skin tone classification: Qwen, Phi, and Gemma achieve 0-5% accuracy for darker skin tones (Types 5-6). Qwen and Phi4 predict Type 3 for 96.2% and 70.4% of samples, respectively, while Gemma defaults to Type 1 with a 46.4% prediction rate. Gemini’s bias in the form of Std is much lower (Std: 7.0%), maintaining 78.3-96.7% accuracy across all Fitzpatrick types with more balanced predictions. C2 Image (Table 4) Age classification: Both Phi and Qwen predictions are concentrated in the 30-39 age range. Phi consistently defaults to age 30 for groups 18-49 (45-75% of predictions) as observed in the Top column, while Qwen predictions are more diverse, they are similarly concentrated most at ages 30 and 45. Gemma predicts age 20 for 98% of 18-29 samples and 51% of 30-39 samples. Gemini distributes predictions more evenly (18.8-35.4% per group) and exhibits the same inverted age pattern observed in C1: performing weakest on middle-age groups (40.0% for ages 30-39) while maintaining stronger performance at extremes. 10A. Elobaid Table 4. Per-Attribute Accuracy with Top Predictions (C2, 120 samples per group) Attr.GroupPhiQwenGemmaGemini AccPrdTopAccPrdTopAccPrdTopAccPrdTop Age 18–2965.046.330(45)81.255.025(49)97.544.120(98)88.335.425(37) 30–3952.575.030(75)66.278.730(64)45.028.420(51)40.024.228(24) 40–4917.526.230(61)22.528.735(38)18.813.830(58)50.020.845(27) 50+36.256.350(48)16.215.045(54)48.812.840(34)74.218.855(29) Overall42.8–46.5–52.5–63.1– Std17.8–27.8–28.4–19.1– Gender Female55.053.8F(51)93.848.8F(94)95.050.6F(95)98.350.8F(98) Male42.541.9F(56)96.251.2M(96)93.849.4M(94)96.748.8M(97) Overall48.8–95.0–94.4–97.5– Std6.2–1.2–0.6–0.8– Country Brazil13.813.8Ind(64)20.05.8USA(70)2.50.4USA(83)73.313.3Bra(73) India45.057.7Ind(45)26.25.0USA(68)88.837.7Ind(89)95.818.3Ind(96) Indonesia0.00.0Ind(61)51.214.4Ind(50)22.54.0Ind(65)78.313.5Ind(78) Mexico0.00.0Ind(68)23.84.6USA(75)5.01.2USA(88)71.712.6Mex(72) Philippines0.00.0Ind(56)13.82.3USA(46)7.50.8USA(50)89.218.9Phi(89) USA46.228.1Ind(53)98.867.3USA(99)88.855.4USA(89)90.821.2USA(91) Overall17.5–38.9–35.9–83.2– Std20.5–29.2–38.0–9.2– Skin tone 13.80.04(78)0.00.03(99)8.823.13(90)62.50.02(63) 26.23.84(79)1000.03(100)98.83.83(74)74.213.93(58) 382.51.54(75)10099.83(100)78.860.83(71)10024.74(55) 496.276.24(73)1000.23(98)67.512.13(66)99.243.54(84) 596.218.54(73)0.00.03(100)6.20.21(53)97.59.94(86) 615.00.04(81)0.00.03(98)0.00.04(59)93.38.16(47) Overall49.9–50.0–43.4–87.8– Std42.0–50.0–39.5–14.3– Abbreviations as in Table 3. Gender classification: Gender bias remains minimal across all models, be- tween 0.6% and 6.2%, with balanced prediction distributions. Phi4 demonstrates the highest inconsistency, misclassifying 56% of male samples as female as ob- served in the Top column. Country prediction: Phi and Qwen exhibit extreme prediction concentra- tion, defaulting to "India" or "USA" for the majority of samples regardless of actual country of origin. Gemini maintains more consistent performance across countries (Std: 9.2%) with balanced predictions (12.6-21.2% per country), while other models show near-complete failure for underrepresented countries (Brazil, Mexico, Philippines: 0-23.8%). Skin tone classification: Qwen exhibits near-total prediction collapse, con- centrating 99.8% of all predictions on Type 3 , resulting in 0% accuracy for Types 1, 5, and 6. Phi shows similar concentration toward Type 4 (76.2% of predic- tions), achieving near-zero accuracy for Types 1 and 2 and partial failure for Type 6. Gemma exhibits strong Type 3 bias (60.8% prediction rate) with failure at extremes (0% accuracy for Type 6 and only 6.2% to 8.8% for Types 1 and 5). Gemini maintains more consistent performance (Std: 14.3%), although higher than its 7.0% Std on C1, indicating dataset-specific variation. Demographic and Linguistic Bias Evaluation in Omnimodal LMs11 4.3 Video-Based Tasks - C2 (Table 5) Table 5. Demographic Group Performance on Video Understanding Tasks (C2) Attr.GroupPhiQwenGemmaGemini ActVisActVisActVisActVis AccTopAccTopAccTopAccTopAccTopAccTopAccTopAccTop Age 18–2961.8st(86)50.5fb(100)78.2st(81)50.7ub(100)80.5st(77)51.9fb(50)85.5st(68)51.4fb(100) 30–3959.2st(88)53.1fb(99)75.2st(82)46.3ub(100)76.5st(80)64.7ub(64)84.2st(72)53.8fb(98) 40–4963.1st(85)53.6fb(100)75.5st(83)46.4ub(100)77.4st(78)46.4ub(100)88.2st(66)50.0ub(96) 50+61.7st(88)–76.5st(82)–78.3st(76)–83.0st(70)– Overall61.5–52.4–76.3–47.8–78.2–54.3–85.2–51.7– Std1.4–1.4–1.2–1.9–1.3–7.5–2.1–1.6– Gender Female63.4st(87)51.4fb(100)76.2st(82)48.7ub(100)78.2st(77)49.5ub(90)85.3st(70)52.6fb(79) Male58.9st(88)47.1fb(97)77.4st(81)48.1ub(100)76.5st(79)54.8fb(57)82.6st(70)50.4fb(100) Overall61.1–49.2–76.8–48.4–77.3–52.1–83.9–51.5– Std2.3–2.2–0.6–0.3–0.9–2.7–1.4–1.1– Skin tone I–I64.3st(83)49.7fb(100)74.6st(83)51.2ub(100)78.7st(77)52.1ub(98)80.7st(71)47.8fb(75) I–IV61.3st(88)48.4fb(92)79.3st(81)45.0ub(100)77.0st(78)54.1ub(64)83.7st(71)59.1fb(98) V–VI60.9st(86)48.9fb(100)75.4st(82)49.5ub(100)78.3st(77)52.0ub(59)85.1st(68)49.6fb(100) Overall62.2–49.0–76.4–48.6–78.0–52.7–83.2–52.2– Std1.5–0.5–2.0–2.6–0.7–0.9–1.8–5.1– Acc: Accuracy (%). Act: Action. Vis: Visibility. st: standing. fb: full body. ub: upper body. Top: Most frequent prediction with percentage of samples receiving that prediction in brackets. Action recognition: In action recognition, models maintain low accuracy disparities with all standard deviations being below 2.3% across all demographic groups. However, models exhibit strong prediction concentration toward "stand- ing" (65.6%–88.3% of predictions) reflected in their moderate accuracies. Visibility classification: Visibility classification shows comparably low dis- parities across all groups (Std: 0.3%–7.5%), though Gemma exhibits relatively high age-based variance (7.5%) compared to other models (1.4%–1.9%). Over- all accuracy per group remains low across models (47.8%–54.3%), partially at- tributed to overlapping category definitions where "upper body visible" and "full body visible" are not mutually exclusive. The absence of visibility transitions in 50+ age group samples prevents complete bias assessment across age groups. 4.4 Audio-Based tasks MCV Audio (Table 6) Age classification: All models tend to concentrate predictions between 30 and 39. Phi returns high numbers of null predictions (38.3% to 44.2%). Qwen achieves the highest overall accuracy (43.8%) but exhibits significant accuracy disparity, 77.5% for ages 18 to 29 and 3.3% for ages 50+. Gemma’s predictions are concentrated on age 25 (33% to 36% of predictions as observed in the Top column) and has a considerable number of null predictions (6.7% to 16.7%). Gemini predicts age 35 for 45% to 50% of predictions hence achieving a 90% accuracy for ages 30 to 39 while showing failure for ages 50+ (0%) and a poor performance of 25% for ages 18 to 29, hence having the largest Std value of 35%. Gender classification: Phi demonstrates catastrophic failure for males, achieving 0% accuracy while misclassifying 67% of males as female, combined 12A. Elobaid Table 6. Per-Attribute Accuracy with Top Predictions and Null Values (MCV, 120 samples per group,±5 year tolerance for age) Attr.GroupPhiQwenGemmaGemini AccPrdNulTopAccPrdNulTopAccPrdNulTopAccPrdNulTop Age 18–2943.31.241.730(43)77.515.80.030(40)67.537.116.725(34)25.02.10.035(47) 30–3953.346.039.230(46)69.253.80.030(48)45.841.917.525(35)90.069.00.035(50) 40–496.79.644.230(47)28.320.60.030(43)5.03.511.725(36)31.728.30.035(47) 50+1.71.738.330(43)3.31.20.030(37)1.71.76.725(33)0.00.40.035(45) Overall19.4–43.8–23.6–26.4– Std22.4–30.2–28.5–35.8– Gender Female58.362.540.0F(58)90.853.30.0F(91)75.073.310.8F(75)80.045.00.0F(80) Male0.00.832.5F(67)84.246.70.0M(84)17.515.810.8F(72)90.055.00.0M(90) Overall21.9–88.7–37.9–87.5– Std29.1–3.3–28.8–5.0– Lang. English10011.110.0Eng(100)10011.110.0Eng(100)10011.110.0Eng(100)98.311.70.0Eng(98) Spanish10011.110.0Spa(100)10011.110.0Spa(100)10011.110.0Spa(100)99.211.60.0Spa(99) Portuguese10011.110.0Por(100)10011.110.0Por(100)10011.110.0Por(100)85.89.70.0Por(86) Hindi10011.110.0Hin(100)10011.110.0Hin(100)10011.110.0Hin(100)10011.90.0Hin(100) Italian10011.110.0Ita(100)10011.110.0Ita(100)10011.110.0Ita(100)95.010.60.0Ita(95) Indonesian10011.110.0Ind(100)10011.110.0Ind(100)10011.110.0Ind(100)99.211.00.0Ind(99) Tamil10011.110.0Tam(100)10011.110.0Tam(100)10011.110.0Tam(100)99.211.30.0Tam(99) Telugu10011.110.0Tel(100)10011.110.0Tel(100)10011.110.0Tel(100)98.311.40.0Tel(98) Vietnamese10011.110.0Vie(100)10011.110.0Vie(100)10011.110.0Vie(100)97.510.80.0Vie(98) Overall100–100–100–97.1– Std0.0–0.0–0.0–4.3– Acc: Accuracy (%). Prd: Percentage predicted as group. Nul: Null predictions (%). Top: Most frequent prediction with percentage of samples receiving that prediction in brackets. with high null prediction rates (32.5% to 40.0%). Gemma also frequently pre- dicts female (73.3% of all predictions) with considerable null predictions (10.8% average). Qwen and Gemini both predict gender more accurately with zero null predictions and much lower gender prediction accuracy disparities. Qwen is 6.6% more accurate on females, while Gemini is 10% more accurate on males. Language identification: Phi4, Qwen, and Gemma achieve perfect per- formance (100% accuracy, Std: 0.0%) across all nine languages with perfectly balanced prediction distributions (11.11% per language), suggesting likely leak- age of MCV samples during training. Gemini, while slightly less perfect (97.1% accuracy, 4.3% std), exhibits a more realistic performance with minor variability and slightly lower than average accuracy for Portuguese (85.8%). C2 Audio (Table 7) Age classification: Audio-based age classification on C2 shows poor per- formance across all models, with a high variance in bias: Phi (Std: 28.5%), Qwen (38.2%), Gemma (20.5%), and Gemini (Std: 39.9%). Both Phi and Qwen tend to concentrate predictions to age 30. Phi predicts age 30 for 45% to 75% of samples across all age groups, with particularly poor performance for the youngest (7.5% for 18-29) and oldest groups (7.5% for 50+). Additionally, Phi returns a high number of null predictions (up to 30%), indicating its frequent failure to gener- ate valid predictions. Qwen has a slightly better overall performance but suffers catastrophic failure for ages 40+ (0.8% to 1.7%), combined with substantial null rates (up to 14.2%) for older groups. Gemma shows the most balanced predic- tion distribution and lowest bias (Std: 20.5%), but still heavily concentrates on age 30 (52.5% to 59.2% of predictions across age groups). Gemini achieves a Demographic and Linguistic Bias Evaluation in Omnimodal LMs13 Table 7. Per-Attribute Accuracy with Top Predictions and Null Values (C2, 120 samples per group,±5 year tolerance for age) Attr.GroupPhiQwenGemmaGemini AccPrdNulTopAccPrdNulTopAccPrdNulTopAccPrdNulTop Age 18–297.53.130.030(45)8.34.20.030(49.2)35.824.60.030(59.2)95.056.00.020(45) 30–3967.565.213.330(65.8)80.880.610.835(42.5)59.260.60.030(55.8)25.825.20.020(36.7) 40–4911.713.310.030(75)1.72.713.335(45.8)13.39.60.030(55.8)10.015.40.035(25.8) 50+7.53.80.030(70)0.80.214.235(55)11.74.40.030(52.5)13.33.30.045(30.8) Overall25.6–30.6–36.4–36.0– Std28.5–38.2–20.5–39.9– Gender Female95.091.75.0F(95)95.851.72.5F(95.8)75.043.81.7F(75)10050.40.0F(100) Male1.70.810.0F(88.3)86.744.25.8M(86.7)85.054.22.5M(85)99.249.60.0M(99.2) Overall55.5–91.5–80.9–99.6– Std66.0–6.4–7.1–0.6– Lang. English65.038.10.0Eng(65)95.816.10.0Eng(95.8)57.59.90.0Eng(57.5)98.316.40.0Eng(98.3) Spanish81.719.05.8Spa(81.7)10019.30.0Spa(100)99.218.60.0Spa(99.2)10016.70.0Spa(100) Portuguese36.77.20.0Eng(56.7)10017.20.0Por(100)98.317.11.7Por(98.3)10016.90.0Por(100) Hindi58.39.99.2Hin(58.3)10016.80.0Hin(100)99.217.50.8Hin(99.2)10016.70.0Hin(100) Tagalog69.212.95.0Tag(69.2)1.70.30.0Ind(81.7)99.216.50.0Tag(99.2)10016.70.0Tag(100) Indonesian3.30.70.0Eng(44.2)10030.30.0Ind(100)98.316.71.7Ind(98.3)10016.70.0Ind(100) Overall49.4–87.3–95.1–99.7– Std28.0–43.2–16.9–0.7– Abbreviations as in Table 6. high accuracy of 95.0% for the youngest age group (ages 18 to 29), which de- clines drastically to values between 10.0% and 25.8% for ages 30+. The Pred. column reveals Gemini over-predicts the youngest group (56.0% of predictions are between 18 and 29) and under-predicts the 50+ group (3.3%). Gender classification: Phi demonstrates catastrophic failure (Std: 66.0%), with near-total prediction collapse toward female (91.7% of all predictions), achieving only 1.7% accuracy for male voices. The model also exhibits high null rates for male samples (10.0%). Qwen and Gemma maintain reasonable perfor- mance with low bias: Qwen (Std: 6.4%) and Gemma (7.1%), though Gemma shows weaker male performance (75.0% female vs 85.0% male). Gemini achieves near-perfect consistency (Std: 0.6%) with 99.6% overall accuracy and balanced prediction distribution (50.4% female, 49.6% male). Language identification: Phi defaults to English for a substantial portion of non-English samples (56.7% of Portuguese, 44.2% of Indonesian), achieving only 3.3% accuracy on Indonesian. Qwen often confuses Tagalog for Indonesian (1.7% accuracy for Tagalog, misclassifying 81.7% as Indonesian), resulting in over-prediction of the Indonesian class (30.3% of all samples predicted as In- donesian). Gemma maintains the most consistent cross-language performance (Std: 16.9%) with near-perfect accuracy for most languages (98.3% to 100%), but shows relative weakness for English (57.5%). Gemini demonstrates near- perfect consistency (Std: 0.7%) with 99.7% overall accuracy and nearly uniform prediction distribution across languages (16.4% to 16.9% per language). Speech Transcription - MCV and C2 (Table 8) Speech transcription performance varies significantly with clip length. MCV’s short clips reveal catastrophic failures in low-resource languages (e.g., Gemma’s Telugu: –23.77%, Phi’s Telugu: –113.00%, Qwen’s Telugu: –49.18% on MCV), whereas most models achieve higher overall accuracies on C2’s long-form scripted 14A. Elobaid Table 8. Word Accuracy (%) by Language, Model, and Dataset Language MCVCC2 GeminiPhiQwenGemmaGeminiPhiQwenGemma English82.4412.0877.0162.8084.5666.1861.4551.80 Hindi76.82-1.0470.3778.2677.630.0620.6967.82 Indonesian84.25-14.1484.0561.5781.600.0042.8577.05 Portuguese81.2162.1976.84-7.5287.3369.3956.4669.49 Spanish84.6831.1896.6176.2088.6571.3655.7481.74 Tagalog—84.032.1824.2068.20 Italian84.6111.6692.2244.97— Tamil43.30-39.17-36.1847.02— Telugu35.67-113.00-49.18-23.77— Vietnamese76.34-2.0457.3124.07— Overall72.15-5.8152.1240.4083.9734.8643.5769.35 Std18.8749.1555.0636.073.9837.4217.5110.24 readings (Gemini improves by 11.82%, Gemma by 28.95%, Phi by 40.67%). Gem- ini achieves the highest overall performance on both datasets (MCV: 72.15%, C2: 83.97%), outperforming other models under challenging conditions (low- resource, short-form). Gemma’s lower English performance on C2 (51.80% vs. 62.8% on MCV) reflects persistent language misclassification. In C2, Gemma incorrectly transcribes English audio into Russian text for 19.2% of samples, trig- gered by the mention of “Saint Petersburg” in Dostoevsky’s The Idiot readings (Appendix C). Representative examples of English audio being mistranscribed into Russian are provided in Appendix E. Similarly, Gemma’s performance drop in Portuguese on MCV (–7.52% vs. 69.49% on C2) stems from confusing Por- tuguese with Spanish. However, the orthographic similarity of the two languages makes it difficult to quantify misclassifications. Nevertheless, representative sam- ples are included in Appendix D. Qwen and Phi show high performance vari- ability on MCV (Std: 55.06% and 49.15%), indicating limited multilingual ASR capabilities. 5 Conclusion This work presents the first comprehensive evaluation of demographic and lin- guistic biases in OLMs across image, video, and audio modalities. The evalua- tion of four omnimodal models reveals significant performance differences across modalities, with distinct bias patterns emerging in each modality. Image understanding demonstrates variable performance overall, with consis- tently high accuracy in gender classification and face verification. However, age classification shows systematic bias against older adults, and skin tone classifica- tion and country prediction exhibit severe prediction collapse, with some models defaulting to a narrow set of categories. Video understanding shows relatively balanced performance across demographic groups for all models. Audio under- standing reveals the most severe demographic disparities, with large accuracy differences in age and gender classification from voice, language identification Demographic and Linguistic Bias Evaluation in Omnimodal LMs15 errors that disproportionately affect certain linguistic groups, and speech tran- scription failures particularly evident in low-resource languages. These findings underscore the critical need for comprehensive fairness eval- uation across all supported modalities before deploying OLMs in sensitive ap- plications such as multimodal age verification and healthcare assistants. Fu- ture research should investigate the architectural factors contributing to these modality-specific biases and develop unified mitigation strategies to achieve eq- uitable performance across image, video, and audio understanding. Acknowledgements I would like to thank the HPC Service of FUB-IT, Freie Universit ̈at Berlin, and Prof. Dr. Tim Landgraf for access to computing time. This work has also been supported by the Google Cloud Research Credits pro- gram with the award GCP19980904. References 1. Abouelenin, et al.: Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture of loras. arXiv preprint arXiv:2503.01743 (2025) 2. AlSaad, et al.: Multimodal large language models in health care: Applications, challenges, and future outlook. J Med Internet Res 26, e59505 (2024) 3. Ardila, et al.: Common voice: A massively-multilingual speech corpus. In: Proceed- ings of the Twelfth Language Resources and Evaluation Conference (2020) 4. Baevski, et al.: wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems (2020) 5. Biometric Update: How multimodal biometrics help with age verifica- tion and compliance (2024), https://w.biometricupdate.com/202410/ how-multimodal-biometrics-help-with-age-verification-and-compliance 6. Chen, et al.: Quantifying and mitigating unimodal biases in multimodal large lan- guage models: A causal perspective. In: Findings of the ACL: EMNLP 2024 (2024) 7. Cheng, et al.: Social debiasing for fair multi-modal llms. Proceedings of the IEEE/CVF International Conference on Computer Vision (2025) 8. Comanici, et al.: Gemini 2.5: Pushing the frontier with advanced reasoning, mul- timodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025) 9. Feng, et al.: Quantifying bias in automatic speech recognition. arXiv preprint arXiv:2103.15122 (2021) 10. Feng, et al.: Towards inclusive automatic speech recognition. Computer Speech & Language 84, 101567 (2024) 11. Gemma Team: Gemma 3 technical report. arXiv preprint arXiv:2503.19786 (2025) 12. Groh, et al.: Towards transparency in dermatology image datasets with skin tone annotations. Proceedings of the ACM on Human Computer Interaction (2022) 13. Hazirbas, et al.: Towards measuring fairness in ai: The casual conversations dataset. IEEE Trans. Biometrics, Behavior, and Identity Science 4(3), 324–332 (2021) 14. Huang, et al.: Visbias: Measuring explicit and implicit social biases in vision lan- guage models. Empirical Methods in Natural Language Processing (2025) 15. Hurst, et al.: Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024) 16. Imam, et al.: Automatic speech recognition for african low-resource languages: A systematic literature review. Proceedings of the 6th AfricaNLP Workshop (2025) 16A. Elobaid 17. Intelligence, A.A.G.: Amazon nova 2: Multimodal reasoning and generation models. Amazon Technical Reports (2025), https://w.amazon.science/publications/ amazon-nova-2-multimodal-reasoning-and-generation-models 18. Jiang, et al.: From specific-MLLMs to omni-MLLMs: A survey on MLLMs aligned with multi-modalities. In: Findings of ACL. p. 8617–8652 (2025) 19. Kulkarni, et al.: The balancing act: Unmasking and alleviating asr biases in por- tuguese. Proceedings of the 4th LT-EDI Workshop (2024) 20. Kulkarni, et al.: Unveiling biases while embracing sustainability: Assessing the dual challenges of automatic speech recognition systems. Interspeech 2024 (2025) 21. Nakatumba-Nabende, et al.: A systematic literature review on bias evaluation and mitigation in automatic speech recognition models for low-resource african lan- guages. ACM Computing Surveys (2025) 22. Narayan, et al.: Facexbench: Evaluating multimodal llms on face understanding. In: Proceedings of the 33rd ACM International Conference on Multimedia (2025) 23. Perera, et al.: Investigating social biases in multimodal llms. In: Proc. IEEE Int. Conf. Automatic Face and Gesture Recognition (FG). p. 1–10 (2025) 24. Porgali, et al.: The casual conversations v2 dataset. In: Proc. IEEE/CVF CVPR Workshops. p. 10–17 (2023) 25. Pratap, et al.: Scaling speech technology to 1000+ languages. Journal of Machine Learning Research 25 (2023) 26. Radford, et al.: Robust speech recognition via large-scale weak supervision. Pro- ceedings of the 40th International Conference on Machine Learning (2022) 27. Robinson, et al.: Face recognition: Too bias, or not too bias? In: Proc. IEEE/CVF CVPR Workshops. p. 0–1 (2020) 28. R ́ıo, et al.: Accents in speech recognition through the lens of a world englishes evaluation set. Research in Language 21(3), 225–244 (2023) 29. Serditova, et al.: Automatic speech recognition biases in newcastle english: an error analysis. Interspeech 2025 (2025) 30. Shahreza, et al.: Facellm: A multimodal large language model for face understand- ing. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (2025) 31. Shim, et al.: Dialetto, ma quanto dialetto? transcribing and evaluating dialects on a continuum. In: Proc. NAACL 2025 (2025) 32. Song, et al.: The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism. In: Proc. NAACL-HLT. p. 4195–4206 (2025) 33. Sung-Bin, et al.: Avhbench: A cross-modal hallucination benchmark for audio- visual large language models. ICRL (2025) 34. Torgbi, et al.: Adapting whisper for regional dialects: Enhancing public services for vulnerable populations in the united kingdom. Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects (2025) 35. Wang, et al.: A comprehensive review of multimodal large language models: Per- formance and challenges across different tasks. arXiv preprint arXiv:2408.01319 (2024) 36. Xu, et al.: Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215 (2025) 37. Zhang, et al.: Llava-video: Video instruction tuning with synthetic data. Transac- tions on Machine Learning Research (2025) Demographic and Linguistic Bias Evaluation in Omnimodal LMs17 Appendix A Overview This appendix provides the detailed specifications and examples that support the findings reported in the main text. Complete prompt specifications for all vision and audio tasks are provided in Appendix B. Appendix C shows the scripted reading from Dostoevsky’s The Idiot used in the C2 speech transcription task. Appendix D and Appendix E provide representative examples of systematic language misclassification failures in Gemma, where Portuguese audio is tran- scribed as Spanish and English audio from the Dostoevsky passage is transcribed as Russian, respectively. B Task Prompt Specifications B.1 Image-Based Task Prompts Casual Conversations V1 (C1) Image Tasks The following prompt was used for physical and demographic attribute estimation from images in the Ca- sual Conversations V1 dataset: Analyze this person’s appearance in the image and create a JSON response with these exact field names. Look carefully at the image and set each value based on what you actually see: "age": [estimate age as number between 15-80], "fitzpatrick_skin_tone": [rate 1-6: 1=very light, 6=very dark], "gender": [choose from: "male", "female"] Replace the bracketed instructions with actual values based on your analysis of the image. Return only the JSON object with no additional text. Casual Conversations V2 (C2) Image Tasks The following prompt was used for physical and demographic attribute estimation from images in the Ca- sual Conversations V2 dataset: Analyze this person’s appearance in the image and create a JSON response with these exact field names. Look carefully at the image and set each value based on what you actually see: "age": [estimate age as number between 15-80], 18A. Elobaid "fitzpatrick_skin_tone": [rate 1-6: 1=very light, 6=very dark], "gender": [choose from: "Male", "Female"], "country_of_nationality": [choose from: "Brazil", "India", "Indonesia", "Mexico", "Philippines", "USA", "Vietnam"] Replace the bracketed instructions with actual values based on your analysis of the image. Return only the JSON object with no additional text. Balanced Faces in the Wild (BFW) Face Verification For the face veri- fication task using the BFW dataset, models were provided with pairs of facial images and prompted to determine identity matching using the following prompt: Do those two images belong to the same person? Give the results in the form of a True or False and a score between 0 and 1 where 0 means they don’t belong to the same person and 1 they belong to the same person. B.2 Video-Based Task Prompts The following prompt was used for action classification and visibility classifica- tion from the Casual Conversations V2 dataset: Analyze each frame and classify the person’s action and visibility. For action, use one of: "rotating", "standing", "sitting", "walking", "laying", "waving" For visibility, use one of: "only head visible", "upper body visible", "full body visible", "lower body visible" Return your answer in valid JSON format like this: "action": "your_classification", "appearance": "your_classification" B.3 Audio-Based Task Prompts Casual Conversations V2 (C2) Audio Tasks The following prompt was used for demographic attribute estimation, transcription, and topic recognition from audio: Analyze this audio and provide a detailed assessment in the following JSON format. For each attribute, provide both a prediction and a confidence score (0-100): Demographic and Linguistic Bias Evaluation in Omnimodal LMs19 "spoken_language": "prediction": str [’English’, ’Spanish’, ’Hindi’, ’Indonesian’, ’Portuguese’, ’Tagalog’], "confidence": int, "transcribed_audio": "text": str, "confidence": int, "gender": "prediction": str [Male or Female], "confidence": int, "age": "prediction": int, "confidence": int For transcribed_audio.text: Provide the transcribed text in the characters of the detected language. No English characters. Numbers should be written out in the detected language rather than using Arabic numerals. Respond ONLY with valid JSON. Common Voice 22.0 Audio Tasks The following prompt was used for multi- lingual speech transcription, language identification, and demographic attribute estimation: Analyze this audio and provide a detailed assessment in the following JSON format. For each attribute, provide both a prediction and a confidence score (0-100): "spoken_language": "prediction": str [’English’, ’Spanish’, ’Hindi’, ’Indonesian’, ’Italian’, ’Portuguese’, ’Tagalog’, ’Tamil’, ’Telugu’, ’Vietnamese’], "confidence": int, "transcribed_audio": "text": str, "confidence": int, "gender": "prediction": str [Male or Female], "confidence": int, "age": "prediction": int, "confidence": int For transcribed_audio.text: Provide the transcribed text in the characters of the detected language. No English characters. Numbers should be written out in the detected language rather than using Arabic numerals. Respond ONLY with valid JSON. C C2 Scripted Reading: Dostoevsky’s The Idiot Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the lat- ter city at full speed. The morning was so damp and misty that it was 20A. Elobaid only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish anything more than a few yards away from the rail car windows. Some of the passengers by this particular train were returning from abroad; but the third-class carriages were the most filled up, mainly with insignificant persons of various occupations and degrees, picked up at the different stations nearer town. All of them seemed weary, and most of them had sleepy eyes and a shivering expression, while their complexions generally appeared to have taken on the color of the fog outside. One of them was a young man of about twenty-seven, not tall, with black curling hair, and small, gray, fiery eyes. He wore a large fur—or rather astrakhan—overcoat, which had kept him warm all night, while his neigh- bor had been obliged to bear the full severity of a Russian November night entirely unprepared. The wearer of this cloak was a young man, also of about twenty-six or twenty-seven years of age, slightly above average height, very fair, with a thin, pointed and very light-colored beard; his eyes were large and blue, and had an intent look about them. “Cold?” “Very,” said his neighbor, readily, “and this is a thaw, too. Imagine if it had been a hard frost! I never thought it would be so cold in the old country. I’ve gotten quite unaccustomed to it.” “What, been abroad, I suppose?” “Yes, straight from Switzerland.” “Wow! My goodness!” The young, black-haired man whistled, and then laughed. — Fyodor Dostoevsky, The Idiot D Representative Failure Cases: Portuguese-Spanish in Gemma Gemma consistently confuses Portuguese with Spanish, transcribing Portuguese audio in the wrong language. The following examples show 20 Portuguese tran- scription failures with Word Accuracy (WA) < 0.5: D.1 Portuguese Transcription Failures (WA< 0.5) 1. WA: 0.4286 Reference (Portuguese): O vinho de agosto n ̃ao faz suco. Gemma Transcription: El vino de agosto no hace suco. 2. WA: 0.4000 Reference (Portuguese): O que vocˆe estava quando veio aqui h ́a cinco anos? Gemma Transcription:¿Qu ́e estaba t ́u cuando ven ́ıa aqu ́ı hace cinco a ̃nos? 3. WA: 0.4000 Demographic and Linguistic Bias Evaluation in Omnimodal LMs21 Reference (Portuguese): Uma crian ̧ca est ́a de p ́e na frente de algumas ́arvores. Gemma Transcription: Una ni ̃na est ́a de pie en frente de algunas ́arboles. 4. WA: 0.3750 Reference (Portuguese): Ela n ̃ao conseguia encontrar um cesto de lixo. Gemma Transcription: No puedo encontrar un sitio de lixo. 5. WA: 0.3750 Reference (Portuguese): Um menino est ́a admirando um carro esportivo verde. Gemma Transcription: Un ni ̃no est ́a dibujando un carro deportivo verde. 6. WA: 0.3571 Reference (Portuguese): Ele precisava de algu ́em com quem conversar para evitar pensar na possibilidade de guerra. Gemma Transcription: ́ El necesitaba a alguien con quien conversar para evitar pensar en la posibilidad de guerra. 7. WA: 0.3333 Reference (Portuguese): No fogo de Costa, quem n ̃ao traz lenha, n ̃ao se aproxima dela. Gemma Transcription: No fago de costa, quien no trae le ̃na no se acerca a ella. 8. WA: 0.3333 Reference (Portuguese): Moju ́ı dos Campos Gemma Transcription: Moi juicios campos. 9. WA: 0.3333 Reference (Portuguese): N ̃ao adianta chorar pelo leite derramado Gemma Transcription: No aguanto escuchar pelo leito derramado. 10. WA: 0.3333 Reference (Portuguese): Por favor, tente entrar em contato conosco em setembro. Gemma Transcription: Por favor, tente entrar en contacto con nosotros en septiembre. 11. WA: 0.3000 Reference (Portuguese): Uma mulher cava uma tigela de comida e a come. Gemma Transcription: Una mujer cava una mat ‘@‘t‘e‘ de comida y come. 12. WA: 0.3000 Reference (Portuguese): Eu nunca poderia saber se eu fiz uma carta ruim. Gemma Transcription: No nunca pude saber si yo hice una carta ruin. 22A. Elobaid 13. WA: 0.3000 Reference (Portuguese): Komur n ̃ao pode ser avisado, ele n ̃ao pode ser ajudado. Gemma Transcription: C ́omo no puede ser avisado, ́el no puede ser ayudado. 14. WA: 0.2857 Reference (Portuguese): Duas senhoras dan ̧cam em sua vestimenta tradicional. Gemma Transcription: Dos se ̃noras danzan en su vestimenta tradicional. 15. WA: 0.2857 Reference (Portuguese): Como funciona a onisciˆencia e a onipresen ̧ca Gemma Transcription: C ́omo funciona la omnisciencia o omnipresencia 16. WA: 0.2857 Reference (Portuguese): Ela quebrou o bra ̧co em v ́arios lugares. Gemma Transcription: Ella quebr ́o el brazo en varios lugares. 17. WA: 0.2857 Reference (Portuguese): S ̃ao aves fascinadas pela ausˆencia de serpente Gemma Transcription: Son aves fascinadas por la presencia de serpiente. 18. WA: 0.2727 Reference (Portuguese): Um menino com uma camisola verde est ́a sentado em uma mesa. Gemma Transcription: Un ni ̃no con una camisa roja est ́a sentado en una mesa. 19. WA: 0.2727 Reference (Portuguese): As mulheres ainda tˆem o papel principal de cuidar da fam ́ılia Gemma Transcription: Las mujeres a ́un tienen un papel principal de cuidado de la familia. 20. WA: 0.2727 Reference (Portuguese): Pessoas que colocam em uma praia pequena, aproveitando o clima quente. Gemma Transcription: Personas que colocan en una playa peque ̃na, aprovechando el clima caliente. E Representative Failure Cases: English–Russian Confusion in Gemma The following examples demonstrate 20 cases where Gemma incorrectly tran- scribed English audio as Russian, revealing cross-lingual interference patterns similar to the Portuguese-Spanish confusion documented in Section D. English translations of the Russian transcriptions are provided using Google’s Neural Machine Translation (GNMT) to illustrate semantic drift from the original ref- erence text. Demographic and Linguistic Bias Evaluation in Omnimodal LMs23 E.1 Russian Transcriptions of English Audio 1. Russian Transcription: Тоа седа на ноември, дури иав. Аj, околу едната но ́к. Аj, троjка на воjски, докато им писбук хеjве. Беше приближуваj ́ки се до левиот ситед на фусби. Ден беше многу сумрак и мистериозен. Беше само со големо тешкост дека денот се пробуди. Беше невозможно да се разликува нешто пове ́ке од неколку часа от каде е карактеристичниот Translation (GNMT): Toa seeda na noemvri, duri iav. Ah, around the same time. Ah, three in the military, they’ve had enough of it. We are quickly approaching the Levit seated on the Fusby. The day is very dark and mysterious. Beshe samo so golemo teshkost deka denot se awaken. It’s impossible, but there’s nothing to be seen about it for a few hours Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 2. Russian Transcription: В конце ноября, во время затишья, в девять утра, поезд на польско-питерской железной дороге приближался к последнему городу на полной скорости. Утро было очень туманым и пасмурным, что потребовало больших усилий, чтобы разглядеть что-либо боле трех метров от окон вагона, и было невозможно различить что-либо боле трех метров от окон вагона. Некоторые из пасажиров этого Translation (GNMT): At the end of November, during a lull, at nine in the morning, a train on the Polish-St. Petersburg railway was approaching the last city at full speed. The morning was very foggy and cloudy, which required great effort to see anything more than three meters from the carriage windows, and it was impossible to distinguish Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 3. Russian Transcription: В конце ноября, в один из дней, около девяти часов утра, поезд на польско-питерской железной дороге приближался к второму дню. Второй город на полной скорости. Утро было настолько туманым и пасмурным, что только с большим трудом удалось разглядеть свет сквозь тучи. И было невозможно различить что-либо боле чем несколько метров впереди из окон вагонов. Некоторые Translation (GNMT): At the end of November, one day, at about nine o’clock in 24A. Elobaid the morning, the train on the Polish-St. Petersburg railway was approaching its second day. Second city at full speed. The morning was so foggy and cloudy that it was only with great difficulty that we managed to see the light through the clouds. And it was Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 4. Russian Transcription: Тогда в конце ноября, во время дождя, в девять утра, я ехал на машине. Петербургский трамвай приближался к городу на полной скорости. Утро было очень дождливым и туманым, что с большой трудностью позволило дню добиться того, чтобы он прошел без каких-либо нарушений, и было невозможно различить что-либо боле чем в нескольких ярдах от окон вагона. Translation (GNMT): Then at the end of November, during the rain, at nine in the morning, I was driving a car. The St. Petersburg tram was approaching the city at full speed. The morning was very rainy and foggy, which made it very difficult for the day to pass without any disturbance, and it was impossible to distinguish anything Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 5. Russian Transcription: В конце ноября, во время холода, в один из утрених часов, поезд на лини Варшава-Петербург двигался в сторону последнего города на полной скорости. Утро было таким пасмурным и туманым, что с большой трудностью удалось пробить солнце. И было невозможно различить что-либо боле чем несколько ярдов впереди от вагонов из окон. Некоторые из пасажиров этого поезда Translation (GNMT): At the end of November, during a cold spell, one morning, a train on the Warsaw-Petersburg line was moving towards the latter city at full speed. The morning was so cloudy and foggy that it was with great difficulty that the sun broke through. And it was impossible to distinguish anything more than a few yards ahead Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great Demographic and Linguistic Bias Evaluation in Omnimodal LMs25 difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 6. Russian Transcription: Почти в конце ноября, во время затишья, в девять утра по московскому времени поезд на польско-питерской железной дороге приближался к последнему городу на полной скорости. В этот день погода была очень туманой и пасмурной, поэтому с большой сложностью удалось разглядеть восход солнца. И было невозможно различить что-либо боле чем несколько ярд от вагонов из-за окон. Translation (GNMT): Almost at the end of November, during a lull, at nine in the morning Moscow time, a train on the Polish-St. Petersburg railway was approaching the last city at full speed. On this day the weather was very foggy and cloudy, so it was very difficult to see the sunrise. And it was impossible to distinguish anything Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 7. Russian Transcription: Тойсянд от ноября, дурня в 9 часов на утране. Поезд на Варшавский, Питерский рельс был продвигался вперед на полную скорость. Утро было солнечное и туманое, чтобы выйти из него было трудно. Но все же ону удалось пройти. И было невозможно различать що-то боле, чем на несколько ярдов от вагонов из окна. Некоторые пасажиры на этом Translation (GNMT): Toysyand from November, fool at 9 o’clock in the morning. The train to Warsaw, St. Petersburg rail was moving forward at full speed. The morning was sunny and foggy, so it was difficult to get out of it. But still he managed to get through. And it was impossible to distinguish anything more than a few yards from the Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 8. Russian Transcription: Туда, в конце ноября, во время поездки на электричке в Санкт-Петербурге, приближалась поздня ночь. Электричка двигалась на полной скорости. Утро было очень туманым и пасмурным, и только с большим трудом удалось разглядеть свет день. Ничего больше, чем несколько ярд от окон вагона, было невозможно различить. Некоторые из пасажиров этой конкретной электрички возвращались из заграничных поездок, 26A. Elobaid Translation (GNMT): There, at the end of November, during a train ride in St. Petersburg, late night was approaching. The train was moving at full speed. The morning was very foggy and cloudy, and it was only with great difficulty that we could see the light of day. Nothing more than a few yards from the carriage windows could be Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 9. Russian Transcription: Вторник, в конце ноября, во время шторма, в девять утра, поезд на польско-росийской железной дороге приближался к Санкт-Петербургу на полной скорости. Утро было настолько туманым и пасмурным, что только с большой тщательностью удалось разглядеть восход солнца. Ничего больше, чем несколько метров впереди, не было различимо с оконых мест. Некоторые из пасажиров этого поезда возвращались из Translation (GNMT): On a Tuesday at the end of November, during a storm, at nine in the morning, a train on the Polish-Russian railway was approaching St. Petersburg at full speed. The morning was so foggy and cloudy that only with great care was it possible to see the sunrise. Nothing more than a few meters ahead was visible Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 10. Russian Transcription: В тусклый ноябрьский день, около девяти утра, поезд на лини Варшава-Петербург приближался к станци «Питерская». Утро было очень туманым и пасмурным, и только с большим трудом день смог пробиться сквозь облака. Ничего боле чем несколько метров впереди вагонов было невозможно различить. Некоторые из пасажиров этого поезда возвращались из заграничных поездок, но третьи класы были заполнены, Translation (GNMT): On a dim November day, around nine in the morning, a train on the Warsaw-Petersburg line was approaching the Piterskaya station. The morning was very foggy and cloudy, and it was only with great difficulty that the day broke through the clouds. It was impossible to discern anything more than a few meters ahead of the carriages. Some Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, Demographic and Linguistic Bias Evaluation in Omnimodal LMs27 a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 11. Russian Transcription: Торс, день в ноябре, во время падения. В девять часов утра на Варшавском и Пирсовском трасе было приближающеся поезда на скорости полной. Утро было очень туманым и мутным, что сделало крайне трудным успешное торможение. Невозможно было различить что-либо боле чем в нескольких ярдах от поезда через окна. Некоторые из пасажиров этого поезда возвращались из заграничных Translation (GNMT): Torso, day in November, during the fall. At nine o’clock in the morning, on the Warsaw and Pirsovsky highways, there was an approaching train at full speed. The morning was very foggy and cloudy, making successful braking extremely difficult. It was impossible to make out anything more than a few yards from the train through the windows. Some of Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 12. Russian Transcription: Тогда здесь в конце ноября, дюрин тау, где на девять часов утра, траин о на де васа он питерсбург, рейо уае, уоал апрочинге латер сити, латер сити, ат фуспид. Де морнинг ваз со дамп и де мисти, дет уаз онли ин де грейт дификулт дет дей суксидин брейкн. И дет ваз импосибл ту дистингюш энитинг Translation (GNMT): Then here at the end of November, durin tau, where at nine o’clock in the morning, train o na de vassa on Petersburg, reyo uae, uoal approchinge later city, later city, at fuspid. De morning vaz so dump and de misty, det vaz only in de great difikult det dey suksidin breakn. And det vaz impossible tu Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 13. Russian Transcription: Тогда в конце ноября, во время того, что в девять часов утра один поезд на польско-питском железнодорожном пути приближался к последнему городу, в тумане. Утро было настолько туманым и мрачным, что было очень трудно даже успешно пройти через разрыв. И было невозможно различить 28A. Elobaid что-либо боле чем на несколько ярдов от вагонов с окнами. Некоторые из Translation (GNMT): Then at the end of November, at nine o’clock in the morning, one train on the Polish-Pitsky railway was approaching the last city in the fog. The morning was so foggy and gloomy that it was very difficult to even successfully navigate through the gap. And it was impossible to distinguish anything more than a few yards Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 14. Russian Transcription: В конце ноября, в начале декабря, в один из вечеров, поезд на польской железной дороге приближался к городу при полной скорости. Утро было очень пасмурным и туманым, что сделало крайне затруднительным даже то, чтобы разглядеть день. И было невозможно различить что-либо боле нескольких метров впереди от вагонов из окон. Некоторые пасажиры этого конкретного поезда возвращались Translation (GNMT): At the end of November, at the beginning of December, one evening, a train on the Polish railway was approaching the city at full speed. The morning was very cloudy and foggy, making it extremely difficult to even see the day. And it was impossible to distinguish anything more than a few meters ahead of the carriages Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 15. Russian Transcription: В суботу, в конце ноября, в девять часов утра на трасе возле города Питерхайв двигался грузовик на полной скорости. Утро было очень пасмурным и сырым, что сделало движение очень трудным. На нескольких милях впереди грузовика было трудно различить что-либо, кроме кабины автомобиля. Некоторые из пасажиров этого поезда возвращались из отпуска. В третьем класе вагонов было Translation (GNMT): On a Saturday in late November, at nine o’clock in the morning, a truck was moving at full speed on a highway near the town of Peterhive. The morning was very cloudy and damp, making travel very difficult. Several miles ahead of the truck it was difficult to make out anything other than the cab of the Demographic and Linguistic Bias Evaluation in Omnimodal LMs29 Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 16. Russian Transcription: То есть, я недавно узнал о ноябре, во время этого, в девять часов утра. Поезд на польско-пинской железной дороге приближался к последнему городу на полной скорости. Утро было очень туманым и пасмурным, что делало очень трудно разглядеть, и было невозможно различить что-либо боле нескольких ярдов от вагонов из-за окон. Некоторые из пасажиров этого необычного поезда Translation (GNMT): I mean, I recently found out about November, during this, at nine o’clock in the morning. The train on the Polish-Pinsk railway was approaching the last city at full speed. The morning was very foggy and overcast, making it very difficult to see, and it was impossible to distinguish anything more than a few yards from the Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 17. Russian Transcription: В один из морей в ноябре, во время сильного дождя, в девятый час утра, поезд на польско-росийской железной дороге приближался к станци Лесная на высокой скорости. Утро было таким туманым и влажным, что только с большой затруднительностью удалось добиться того, чтобы день начал пробиваться. И было невозможно различить что-либо боле чем несколько ярдов вдали от Translation (GNMT): On one of the seas in November, during heavy rain, at nine o’clock in the morning, a train on the Polish-Russian railway was approaching Lesnaya station at high speed. The morning was so foggy and humid that it was only with great difficulty that the day began to break through. And it was impossible to distinguish anything Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 18. Russian Transcription: В конце ноября, в пятницу, в девять часов утра поезд Варшавско-Плецской железной дороги приближался к последнему городу Польши. 30A. Elobaid Утро было очень туманым и пасмурным, поэтому с большой трудностью удалось запустить тормоза, и было невозможно различить что-либо боле чем несколько метров впереди вагонов. Некоторые из пасажиров этого поезда возвращались из поездок. Но третий вагон был самым Translation (GNMT): At the end of November, on a Friday, at nine o’clock in the morning, a train of the Warsaw-Pletsk Railway was approaching the last city in Poland. The morning was very foggy and cloudy, so it was with great difficulty that the brakes were applied, and it was impossible to distinguish anything more than a few Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 19. Russian Transcription: Во время позднего ноябрьского дня, около девяти часов утра, поезд на лини Варшава-Битова был приближался к станци Сидячий в полной скорости. Утро было очень туманым и пасмурным, что сделало невозможным разглядеть что-либо боле чем на несколько ярдов от окон вагона. Некоторые из пасажиров этого поезда возвращались из заграничных поездок, но в третьем вагоне находился самый... Translation (GNMT): On a late November day, around nine o’clock in the morning, a train on the Warsaw-Bitova line was approaching Sidyachy station at full speed. The morning was very foggy and overcast, making it impossible to see anything more than a few yards from the carriage windows. Some of the passengers on this train were returning from trips Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ... 20. Russian Transcription: Тогда день отправления в Новом Году, в девять часов утра, на вокзале Warsaw, в Петербурге рейс ушел в путь, приближаясь к ледяной городу, на полной скорости. Утро было очень пасмурным и туманым, что делало видимость крайне трудной, так что день следования был прерван. И было невозможно различить что-либо боле чем на несколько ярдов от вагонов. Translation (GNMT): Then the day of departure in the New Year, at nine o’clock in the morning, at the Warsaw station in St. Petersburg, the flight set off, approaching the icy city at full speed. The morning was very cloudy and foggy, making visibility extremely difficult, so the day’s voyage was interrupted. And Demographic and Linguistic Bias Evaluation in Omnimodal LMs31 it was impossible to distinguish anything Reference: Toward the end of November, during a thaw, at 9 o’clock one morning, a train on the Warsaw and Petersburg railway was approaching the latter city at full speed. The morning was so damp and misty that it was only with great difficulty that the day succeeded in breaking; and it was impossible to distinguish ...