Paper deep dive
Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs
Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.
Tags
Links
- Source: https://arxiv.org/abs/2608.23026v1
- Canonical: https://arxiv.org/abs/2608.23026v1
Trouble viewing inline? Open PDF directly â
Full Text
58,231 characters extracted from source content.
Expand or collapse full text
Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs Yuanjun Feng 1 , Tanzhou Liu 1 , Stefan Feuerriegel 2 , Yash Raj Shrestha 1 1 University of Lausanne, Switzerland 2 LMU Munich, Munich Center for Machine Learning (MCML), Germany yuanjun.feng, tanzhou.liu, yashraj.shrestha@unil.ch, feuerriegel@lmu.de Abstract Multilingual LLM outputs can vary across so- ciocultural contexts. However, evidence of cul- tural grounding can be misleading: identity la- bels may be inferred from explicit or indirect textual cues, while names and wording can re- veal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human- validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross- cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies system- atically across languages and tasks. Removing direct identity cues sharply reduces identity- label prediction in English and Chinese, but has a much smaller effect in French. Across all languageâgenre settings, the cultural context as- sociated with the source language receives the highest average relevance score, with moderate agreement between automated and human rat- ings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake sur- face cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns. 1 Introduction In the real world, bias rarely announces itself with a single, explicit sentence. Instead, it is deeply embedded within the complex fabric of narra- tives, manifesting in the allocation of roles, the logic of causality, and the subtlety of value judg- ments (Tabassum and Nayak, 2021; Santoniccolo et al., 2023). This invisible current continuously shapes readersâ commonsense assumptions about who seems naturally suited to which occupations. Rather than overt declarations of gender superior- ity, bias operates through the relentless repetition of assumptions about who is portrayed as a rational protagonist versus who is relegated to supportive, emotional labour (Allan et al., 2025; Rettberg and Wigers, 2025). Such inconspicuous yet cumulative narrative bias is difficult to detect and correct in digital media environments, and the assumptions it conveys can take hold early and shape subsequent interests (Bian et al., 2017). As large language models (LLMs) increasingly serve as the narrative infrastructure of human com- munication, they actively shape occupational and cultural narratives. In these settings, models do not merely âretrieve factsâ; they generate explanations that carry culturally specific assumptions across languages. These narratives can reinforce existing inequalities and risk propagating dominant cultural norms when producing content for multilingual audiences (Kirk et al., 2021). Many studies demonstrate that LLMs inherit and often amplify sociocultural bias in their training data (Bolukbasi et al., 2016; Zhao et al., 2018). However, existing evaluations predominantly fo- cus on explicit bias, assessed through sentence- completion prompts or sentence-level completions (Kotek et al., 2023). This focus on explicit bias, though useful, overlooks implicit and hidden bi- ases that arise in complex contexts such as task al- locations, persona settings, and nuanced language choices (Wilson and Caliskan, 2024). Identifying these subtler forms of bias is essential to develop- ing fairer and more reliable language models. Moreover, a fundamental tension persists be- tween the pursuit of universal fairness and robust- ness and the preservation of culturally grounded di- versity. Although recent discourse has shifted from identifying direct harms to examining how LLMs shape social perceptions (Weidinger et al., 2021; arXiv:2608.23026v1 [cs.CL] 24 Aug 2026 Lin and Li, 2025), their cultural impact remains insufficiently understood. In particular, it is unclear whether biases measured in explicit, controlled set- tings persist in implicit, long-form contexts, and how these dynamics vary across languages that encode distinct discursive norms and cultural tra- ditions (Ghosh and Wilson, 2025). As LLMs are deployed globally, there is still limited systematic evidence on how sociocultural patterns in gener- ated narratives vary across linguistic settings and whether apparent cross-cultural differences persist after surface cues are controlled. To bridge this gap in linguistic and cultural un- derstanding, we introduce a multi-agent framework to audit sociocultural patterns in LLM-generated narratives. Our framework connects three com- plementary analytical lenses: bias representation, identity-linked semantic separation, and cross- cultural patterns. We organize the audit around these lenses: Bias Representation: examines how socially patterned associations vary across language-linked contexts and task conditions. Identity-Linked Semantic Separation: uses full-text label recovery as a positive control and examines how identity- label recoverability changes after direct-cue deletion across languages. Cross-Cultural Patterns: examines context-specific rele- vance alongside the sensitivity of source-language separa- bility to surface-cue controls. Figure 1 provides an overview of the experiment setup and multi-agent auditing workflow. Appendix A maps each analytical lens to its study-specific focus and evidence. Contributions: (1) Multi-level sociocultural framework: We conceptualize multilingual socio- cultural evaluation as a multi-level problem span- ning bias representation, identity-linked semantic separation, and cross-cultural patterns, and opera- tionalize these levels in a unified audit. (2) Empir- ical findings: Our experiments show that language and task shape how LLMs represent social bias, identity, and cultural context, while cue-removal controls show that some apparent differences rely on names and language-specific wording; cultural grounding should therefore be evaluated through meaning and contextual relevance rather than sur- face cues alone. 2 Related Work As LLMs increasingly function as the narrative infrastructure of human communication, their re- liance on vast, uncurated textual data exposes them to deeply embedded gender and occupational stereotypes (Bolukbasi et al., 2016; Caliskan et al., 2017). This raises significant concerns regarding representational bias, particularly in socially im- pactful contexts such as career-path recommenda- tions and rĂ©sumĂ© screening (Wilson and Caliskan, 2024; Gaebler et al., 2024), and in the political framing of automatically generated content (Bang et al., 2024; Motoki et al., 2024). Existing research on LLM bias can be divided by task formulation. One stream uses controlled, explicit prompts and sentence-completion bench- marks to detect isolated bias patterns (Nangia et al., 2020; Nadeem et al., 2021). WinoBias, for exam- ple, tests associations between occupations and gendered pronouns in sentence completions (Zhao et al., 2018, 2019). However, these isolated metrics often miss how bias surfaces in long-form gener- ation. Another stream focuses on implicit tasks. Here, bias is not overt but woven into the logic and evaluative framing of stories or news reports (Hof- mann et al., 2024; Rettberg and Wigers, 2025). In these settings, prompts can shape the semantic con- struction of generated identities (Steinborn et al., 2022; Gnadt et al., 2025). The influence of language structure on the repro- duction of gender stereotypes is deeply connected to cultural context and the composition of train- ing data (Abid et al., 2021; Zhong et al., 2024). Language and culture intertwine, as cultural dimen- sions like gender norms shape institutions and civic participation (Alesina et al., 2013). Grammatical gender is a key linguistic driver. In gendered lan- guages such as French, gender is morphologically expressed in nouns and adjectives. Non-gendered languages (such as English) or character-based lan- guages (such as Chinese) employ different mor- phological or semantic strategies. Cross-country evidence suggests that gendered language structure correlates with higher societal gender inequality (Mavisakalyan, 2015). Evidence from immigrant households further suggests that speakers of more gender-marked mother tongues divide household labour along more traditional lines (Hicks et al., 2015). LLM development, especially Reinforcement Learning from Human Feedback (RLHF), adds complexity. RLHF can suppress overtly biased lan- guage while complying with safety policies (Dai et al., 2024). However, RLHF yields only limited improvements on bias benchmarks (Ouyang et al., 2022), and reduced overt bias need not imply the Experiment Setup Languages FrenchChinese English Task Formulations Implicit Task Story & News Representational Bias Stereotype Balance Multi-agent Auditing with Human-in-the-Loop ExtractionClassificationTranslation Identity ConstructionCross-Cultural Patterns Narrative Positioning Semantic Segregation Cultural Relevance Source Language & Culture Explicit Task Sentence Completion Humans provide validation and feedback 150+ native speakers recruited via Prolific Occupational gender norm alignment Skew in gender display Social role assignment Strength of native regional ties Protagonist identities, cultural markers Source-language separability Identity and cultural categories Models GPT-4o, Mistral-Large, LLaMA-3.3- 70B, GPT-4o-mini, Mistral-Small- 24B, Mistral-Nemo, Qwen3-32B, LLaMA-3.1-8B, DeepSeek-R1, DeepSeek-V3.1, Grok-4-Fast, Grok-4 Translate for aligned comparison Distinct semantic features by identity Shared criteria Scoring Relevance to a specific culture Figure 1: Overview of the experiment setup. absence of gender-associated differences in longer- form content. This motivates measuring both ag- gregate representation and separation in narrative embedding space. To evaluate and mitigate these biases, studies have shifted from static evaluations (Kurita et al., 2019; Cryan et al., 2020) toward dynamic and agen- tic approaches. Tools such as RUTEd (Lum et al., 2025), multi-agent debate frameworks (Feng et al., 2025), and simulations of interaction with stereo- typically biased AI (Allan et al., 2025) provide in- frastructure for diagnosing long-form behavior. Yet multilingual cultural audits introduce an additional identification problem. A marker may reveal its source language through original-language word- ing, names, institutions, or places; treating a lan- guage as a culture can also erase within-language heterogeneity (Hershcovich et al., 2022). Existing work documents cross-cultural preferences and cul- tural dominance, but rarely tests source-language separability under translation and named-entity masking. Our analysis pairs three-context rele- vance scores with those two controls. 3 Methodology We conduct a multilingual evaluation of socio- cultural patterns in LLM-generated outputs under different task conditions. The audit is organized around three complementary analytical lensesâ bias representation, identity-linked semantic sepa- ration, and cross-cultural patterns. 3.1 Multi-Agent Auditing Framework The framework coordinates four agents: extrac- tion, classification, cultural scoring, and translation. They identify protagonist labels and text-grounded cultural markers, select and score markers in three contexts, and prepare non-English markers for the surface-cue controls. Figure 1 summarizes output generation, framework processing, and the subse- quent three-lens analysis; full agent roles and scor- ing settings appear in Appendix D. Two Prolific human studies validate cultural- marker annotations (39 raters; 624 marker-level comparisons) and protagonist-label extraction (116 raters). Human endorsement of automated marker- inclusion decisions is 80.0%. After weighting to match the full-corpus distribution, automated and human protagonist labels agree for 91.3% of EN, 82.6% of FR, and 81.6% of ZH cases. Appendix D details sampling, aggregation, agreement metrics, and resampling procedures. 3.2 Materials 3.2.1 Task Formulations We use occupations as a focused and socially con- sequential case domain for measuring representa- tional bias. Diverse professions provide semantic variation and are intertwined with sociocultural fac- tors such as gender norms, power, and stereotypes (Kirk et al., 2021). Explicit Task (Sentence Completion): We prompt the model with âOccupation is [MASK]â and require it to select exactly one item from a fixed list of gendered labels (e.g., male, female, man, woman). 1 Implicit Task (Contextualized Generation): We instruct LLMs to generate long-form text in 1 We implement constrained decoding via schema enforce- ment (structured outputs) (Geng et al., 2025), reducing syn- tactic variance and eliminating out-of-distribution formatting errors during sentence-completion evaluation. two writing genres: Story and News. These genres serve as narrative settings where subtle biases may be embedded. Together, the Explicit Task isolates bias under controlled constraints, whereas the Implicit Task (Story and News) captures how bias manifests dur- ing contextualized generation. We refer to Explicit, Story, and News collectively as task conditions and reserve genre for Story and News. 3.2.2 Occupations We select 18 common occupations that span a range of gender-associated stereotypes (Zhao et al., 2018) and differ in occupational prestige (Hughes et al., 2024). A prespecified shared mapping in the exper- iment configuration designates eight occupations as female-associated and ten as male-associated, with translated equivalents used across EN, FR, and ZH; the complete mapping is provided with the accompanying materials. 3.2.3 Models We select 12 LLMs to balance training-language coverage and parameter scale. For each avail- able languageâtaskâmodelâoccupation configura- tion, we target 50 outputs and analyze the records available after parsing and processing. All 12 mod- els contribute EN and ZH cells, and nine contribute FR cells. Exact versions, language availability, and generation settings are listed in Appendix D. 3.2.4 Cultural Markers Following prior work (Masoud et al., 2025; Myung et al., 2024), the cross-cultural analysis extracts markers from the Implicit Task narratives using a ten-category taxonomy (Appendix B). An included cultural marker is an extracted candidate that sat- isfies the prespecified inclusion rule. A marker- bearing narrative contains at least one included cultural marker; only these narratives receive three- context relevance scores. 3.3 Metrics We compute probabilities (P) from occurrence frequencies among the available outputs for each languageâtaskâmodelâoccupation configuration. In the Explicit Task,Pis the proportion of times a specific label is selected from the candidate set. In the Implicit Task,Pis estimated from protagonist- label occurrence frequencies. 3.3.1 Bias Representation Metrics Building on prior work (Lum et al., 2025) and adapting it to our multilingual setting, we as- sess representational bias with two metrics. For each modelâlanguageâtaskâoccupation configura- tion, unidentified labels remain in the denominator but contribute to neither the female nor male nu- merator. Stereotype (M s,o ): Measures the degree to which the modelâs gender portrayals align with the occupationâs predefined gender stereotype. M s,o = P s o â P a o ,(1) whereP s o andP a o denote the probabilities of pro- ducing the stereotypical versus anti-stereotypical gender for occupation o. Balance (M b,o ): Measures skew in the modelâs gender portrayals, computed as the difference be- tween female and male assignment probabilities. M b,o = P f o â P m o ,(2) whereP f o andP m o are the probabilities of female versus male portrayals for occupation o. When reporting aggregate scores for a spe- cific model, language, and task, we average these occupation-level metrics across occupations with parsed outputs: M q = 1 |O m,l,t | X oâO m,l,t M q,o , q âs, b.(3) Here,qdenotes the stereotype or balance metric, andO m,l,t is the available occupation set for model m, language l, and task condition t. Within the Bias Representation lens, we fit separate Type-I ANOVAs with sum contrasts to aggregateM s andM b scores. Each modelâ languageâtask aggregate is one observation; predic- tors are language, task condition, their interaction, and model. We report partialÏ 2 and Benjaminiâ Hochberg-adjustedp-values. Alternative denomi- nators, common-model subsets, and trial-level mod- els appear in Appendix C. 3.3.2 Identity-Linked Semantic Separation Within the Identity-Linked Semantic Separation lens, we sample 3,000 implicit narratives from each languageâgenre cell (18,000 total; random seed 42) and encode the complete original-language con- tent with multilingual E5-large. 2 Primary inference remains in the original 1,024-dimensional space. 2 intfloat/multilingual-e5-large. Within each cell, a logistic-regression probe with balanced class weights predicts the female versus male protagonist label under five-fold stratified group cross-validation; modelâoccupation groups are kept within a single fold. We compute ROC AUC from all out-of-fold predictions and report 95% intervals from 2,000 bootstrap resamples of these groups. Because the labels are extracted from the same narratives and full texts retain names and direct gender terms, we use full-text performance as a positive control confirming that the probe can recover labels when direct cues remain. The main comparison repeats the probe on a paired, label-balanced subset after deleting protag- onist names and prespecified direct gender terms. Figure 3 uses a shared two-dimensional UMAP for visualization; all inference remains in the origi- nal E5 space. Full preprocessing, resampling, and visualization settings appear in Appendix D. 3.3.3 Cross-Cultural Pattern Measures The Cross-Cultural Patterns lens uses two comple- mentary quantities: self-context advantage derived from the three-context relevance scores and source- language separability derived from a classification probe. The design is motivated by work on cross- context preference and cultural adaptability (Wang et al., 2024; Naous et al., 2024; Rao et al., 2025; Li et al., 2024), as well as calls for distributional diagnostics and human verification in cultural eval- uation (Dai et al., 2025; Chiu et al., 2025). Narrative-weighted relevance: Letr m,c â [1, 7]be the automated relevance assigned to markermfor language-linked contextc. To pre- vent narratives with many extracted markers from receiving greater weight, we first average within each marker-bearing narrative d: Ìr d,c = 1 |M d | X mâM d r m,c .(4) We average Ìr d,c across narratives separately by source language and genre. This measure is condi- tional on|M d | > 0. Self-context advantage: For each marker- bearing narrative, we compare the relevance of its source-linked (âselfâ) context with the mean relevance of the other two contexts: A d = Ìr d,self â Ìr d,other1 + Ìr d,other2 2 .(5) Positive values indicate an automated self-context advantage: the included cultural markers in a nar- rative are scored as more relevant to the context linked to their source language than to the other two contexts. Source-language separability under surface- cue controls: We test how readily a markerâs source language can be recovered from (i) the orig- inal marker, (i) its English translation, and (i) the translation after detected named entities are re- placed by a common token. We encode the three versions with multilingual E5-large and evaluate a multinomial probe with balanced class weights under five-fold stratified group cross-validation. Balanced accuracy computed from all out-of-fold predictions is termed source-language separabil- ity; translation grouping, masking, and interval- estimation details appear in Appendix D. 4 Results We report findings through the three analytical lenses in turn: bias representation, identity-linked semantic separation, and cross-cultural patterns. 4.1 Bias Representation Across Languages and Tasks Figure 2 plots modelâtask combinations in the Balance and Stereotype space (Eq. 1 and Eq. 2). Higher Stereotype indicates stronger alignment with gender stereotypes, and Balance captures the overall female-to-male skew. In English, task conditions show clear separa- tion: the Explicit Task yields high Stereotype (0.5 to 0.8), while Story drops to 0.1 to 0.4. News fur- ther minimizes Stereotype (â 0) and shifts to a positive Balance (> 0.6). In French, differences appear along Balance rather than Stereotype. Story skews into negative Balance (< â0.4), whereas Explicit and News shift positive. In Chinese, task conditions overlap substantially: models consis- tently output moderate-to-high Stereotype (0.2 to 0.7). The primary all-output ANOVAs detect lan- guage, task condition, and their interaction for both metrics (all adjustedp < 0.001). The interaction is substantial for Balance (F (4, 79) = 21.63, par- tialÏ 2 = 0.495) and Stereotype (F (4, 79) = 9.07, partialÏ 2 = 0.278), matching Figure 2. Model- level variation is detected for Balance in the pri- mary specification (F (11, 79) = 1.98, adjusted p = 0.047, partialÏ 2 = 0.106) and for Stereo- type in complementary specifications; the languageâ task interaction remains detected throughout (Ap- pendix C). -1.0-0.50.00.51.0 -0.2 0.0 0.2 0.4 0.6 0.8 Stereotype English -1.0-0.50.00.51.0 Balance French -1.0-0.50.00.51.0 Chinese Explicit Story News Figure 2: task-level Balance and Stereotype scores by language. Each point corresponds to one model in one task condition. Shaded ellipses summarize dispersion across models within each condition. Some implicit narratives receive an unidenti- fied protagonist label, particularly in News; the languageâtask interaction remains detected when estimates are conditioned on female or male labels. Trial-level analyses likewise detect languageâgenre and model terms (allp < 0.001), providing con- sistent evidence of variation across settings and models (Appendix C). Overall, the representational metrics vary across languages and task conditions. We next turn to the Identity-Linked Semantic Separation lens, examin- ing how protagonist-gender recoverability changes after direct-cue deletion. 4.2 Identity-Linked Semantic Separation Across Languages and Genres Beyond aggregate representation, the Identity- Linked Semantic Separation lens uses full-text label recovery as a positive control and tests how recover- ability changes after direct-cue deletion on held-out modelâoccupation groups. The paired comparison in Table 1 shows that deleting names and direct gender terms reduces AUC by 0.230â0.365 in EN and 0.320â0.342 in ZH, but by only 0.003â0.037 in FR. Separately, the full-sample full-text probe yields near-ceiling out- of-fold AUCs of 0.992â1.000, with 95% interval lower bounds of at least 0.985, as expected when labels are extracted from the same narratives and direct cues remain. The remaining high French recoverability shows that identity-linked separation is less dependent on the removed direct lexical cues. Figure 3 provides a descriptive visualization of the full-text multilingual sample. Figure 4 provides an illustrative within-occupation contrast in narra- tive positioning. â Girlâs Story (Private Setting) ...Lila donned her grandmotherâs oversized spec- tacles and fashioned a cape from an old table- cloth. She transformed her treehouse into a grand courtroom, complete with a gavel made from a wooden spoon. Lilaâs first case was the Great Cookie Caper. Her little brother, Max, had been accused of sneaking cookies before dinner. Lila gathered evidence, questioned witnesses... â Boyâs Story (Public Setting) ...Leoâs school announced a âCareer Dayâ event. Each student was to dress up as their dream pro- fession. Excited, Leo donned a tiny suit and crafted a paper badge that read âLeo the Lawyer.â During the event, Leo conducted a mock trial about who ate the last cookie from the cookie jar. His classmates giggled as Leo presented his case with enthusiasm, listing clues and interviewing witnesses... Figure 4: Example of distinct narrative positioning for the same occupation. Both stories depict a child aspiring to become a lawyer. In the girlâs story, the scene unfolds in a private, family setting; in the boyâs story, it takes place in a public, school-based setting. Despite the shared occupation, the narratives associate gender with different social contexts. Direct-cue deletion causes large AUC losses in EN and ZH but much smaller losses in FR, reveal- ing language-dependent reliance on direct identity cues; near-ceiling full-text performance serves only as a positive control. We next turn to the Cross- Cultural Patterns lens, examining marker relevance and source-language separability under surface-cue controls. 4.3 Cross-Cultural Patterns Human raters endorse 80.0% of the automated marker-inclusion decisions (95% CI [76.3%, LanguageGenreFull textCue-deletedâ AUC ENStory.999 [.996, 1.000].634 [.555, .713] â.365 [â.444,â.286] ENNews1.000 [1.000, 1.000].771 [.703, .827] â.230 [â.297,â.173] FRStory.998 [.992, 1.000].994 [.988, .999] â.003 [â.010, .002] FRNews.997 [.989, 1.000].960 [.935, .981] â.037 [â.060,â.018] ZHStory.994 [.984, 1.000].674 [.601, .744] â.320 [â.390,â.252] ZHNews.988 [.976, .997].646 [.573, .717] â.342 [â.411,â.273] Table 1: Identity-label probe evaluated on held-out modelâoccupation groups in the paired subset (200 narratives per cell). Full text is the positive-control condition; values are ROC AUC with 95% intervals computed by resampling these groups, and â is cue-deleted minus full-text AUC. English - StoryFrench - StoryChinese - Story English - NewsFrench - NewsChinese - News FemaleFemale ContourMaleMale Contour Figure 3: Shared UMAP projection of the full-text multilingual narrative embeddings. Pink points and solid contours denote narratives associated with females, and blue points and dashed contours denote narratives associated with males; squares indicate Story and circles indicate News. Contours indicate density regions (20/50/80 percentiles). Display-only percentile trimming is described in the methodology; the UMAP axes are not used for inference. 83.5%]). Across 624 comparisons between human and automated relevance scores, agreement within two scale points is 0.641, quadratic-weightedÎș is 0.280, mean absolute error (MAE) is 2.09, and marker-level Spearman correlation is 0.555; all four metrics outperform their within-language permu- tation baselines (p < 0.001). Among 54 markers with at least two valid ratings whose human rel- evance scores span at most one scale point, the automated mean self-context advantage remains 3.28 [2.53, 3.95], and the source-linked-context score is strictly highest for 77.8% [64.8%, 88.9%]. Bootstrap intervals and humanâhuman agreement statistics appear in Appendix D. Our analysis comprises 4,562 included cultural markers from 3,926 marker-bearing implicit narra- tives. We first test whether the automated relevance scores favor the source-linked context. Figure 5 reports the narrative-weighted auto- mated relevance defined in Eq. 4. Each source- language row has its largest mean on the diagonal: included cultural markers are scored as most rele- vant to the context linked to the language in which the narrative was requested. The mean self-context advantages (Eq. 5) are 2.42, 4.26, and 2.94 for EN, FR, and ZH Story, respectively, and 2.64, 4.78, and 3.64 for News. Thus, all six cells show a positive automated self- context advantage. The surface-cue controls substantially reduce source-language separability. On markers with En- glish translations, balanced accuracy is 0.991 (95% CI [0.987, 0.994]) for original markers and falls to 0.721 [0.706, 0.736] after translation. Replacing detected named entities with a common token re- duces it further to 0.558 [0.542, 0.575], against a three-class chance level of 0.333. Overall, the two measures characterize comple- ENFRZH Evaluated language-linked context EN FR ZH Source language 4.913.101.87 2.516.101.16 2.932.325.56 Story ENFRZH Evaluated language-linked context EN FR ZH 4.912.841.70 2.156.381.06 2.552.176.00 News 1 2 3 4 5 6 7 Mean relevance (1-7) Figure 5: Cross-cultural relevance pattern among marker-bearing narratives. Each cell shows narrative-weighted mean automated relevance on the original 1â7 scale by source language, evaluated language-linked reference context, and genre. mentary aspects of the observed cross-cultural pat- tern: included cultural markers show a positive au- tomated self-context advantage, while their source- language separability is reduced by translation and named-entity masking. 5 Discussion First, the languageâtask interaction is our largest effect. In English, Explicit Task outputs strongly align with stereotypes, whereas News approaches zero Stereotype but shows a female skew; the Chi- nese conditions overlap more. Cue deletion like- wise reduces identity-label recoverability in En- glish and Chinese but not French, where grammati- cal agreement remains informative. Together, these results show that sociocultural scores are condi- tional on both elicitation and language, rather than stable model properties (Lum et al., 2025). Second, every languageâgenre cell shows a pos- itive self-context advantage, yet source-language separability falls after translation and entity mask- ing. This decline shows that wording and names contribute to apparent cultural grounding. Unlike closed-form benchmarks with answer keys (Chiu et al., 2025; Rao et al., 2025), our open-ended au- dit asks what remains after surface cues are weak- ened. Separating contextual relevance from lan- guage recognition therefore provides a more falsifi- able test without equating language with a bounded culture (Hershcovich et al., 2022). Finally, the appropriate evaluator depends on the construct. High human and automatedâhuman agreement for structured protagonist labels sup- ports automated scaling. Cultural relevance is more interpretive: humanâhuman agreement is moder- ate, and automatedâhuman agreement is lower. LLM judges offer consistent coverage but may share model priors, while human raters provide a more independent reference but are costly and necessarily partial. Human involvement is there- fore essential for defining constructs, revealing le- gitimate disagreement, and calibrating automated judges; LLMs can then extend the validated rubric at scale. Future evaluations should condition results on language and elicitation, use translation and entity controls, and report humanâhuman agree- ment, automatedâhuman agreement, and recruit- ment composition together. 6 Conclusion We introduced a human-validated, multi-agent au- dit that separates social bias, identity representa- tion, and cross-cultural patterns. Representation varies across languages and tasks, while cue re- moval reduces identity-label prediction in English and Chinese much more than in French. Source- linked contexts receive the highest average rele- vance scores, but translation and name masking substantially reduce source-language recognition. These signals are therefore not interchangeable: multilingual audits can mistake names and wording for cultural understanding. Our framework eval- uates them separately and tests whether apparent cross-cultural patterns remain after surface cues are weakened. Limitations 1.Our evidence is limited to three high-resource language-linked contexts (EN, FR, and ZH), not bounded cultures. The dataset covers 18 occupations, two implicit genres, unequal model pools, and three protagonist labels. Four aggregate cells have partial occupa- tion coverage, while unidentified labels affect some News estimates. 2. Some aggregate estimates depend on the ANOVA specification. The embedding probe measures identity-label recoverability, not causal effects; cue deletion may leave gram- matical or indirect cues, especially in French. 3. The cross-cultural analysis covers 3,926 marker-bearing narratives (6.6% of implicit outputs). Human validation examines 178 detected markers and only their relevance to the source-linked context, so it does not mea- sure extraction recall or validate other-context scores. The probe translates individual mark- ers rather than full narratives. Thus, our find- ings describe patterns in this corpus, not cul- tural authenticity, competence, causality, or broader generalizability. Ethical Considerations Our language-linked comparisons are analytical contrasts among model outputs, not definitions of authentic English, French, or Chinese culture. We do not attribute the observed patterns to inherent values of any community or prescribe what a cul- turally authentic narrative should contain. The find- ings should be used as diagnostic signals about model behavior without treating language groups or stereotypes as fixed categories. Participants in both Prolific studies provided in- formed consent; participants in the cultural-marker study were compensated at a rate of at least GBP 15 per hour. AI Usage Statement: We used generative AI to assist with writing and icon editing. All claims, analyses, and citations were verified by the authors. References Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In AIES â21, pages 298â306. Alberto Alesina, Paola Giuliano, and Nathan Nunn. 2013. On the origins of gender roles: Women and the plough. The Quarterly Journal of Economics, 128(2):469â530. Kevin Allan, Jacobo Azcona, Somayajulu Sripada, Georgios Leontidis, Clare A. M. Sutherland, Louise H. Phillips, and Douglas Martin. 2025. Stereotypical bias amplification and reversal in an experimental model of human interaction with gen- erative artificial intelligence. Royal Society Open Science, 12(4):241472. Yejin Bang, Delong Chen, Nayeon Lee, and Pascale Fung. 2024. Measuring Political Bias in Large Lan- guage Models: What Is Said and How It Is Said. In ACL â24, pages 11142â11159. Lin Bian, Sarah-Jane Leslie, and Andrei Cimpian. 2017. Gender stereotypes about intellectual ability emerge early and influence childrenâs interests. Science, 355(6323):389â391. Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam T. Kalai. 2016. Man is to computer programmer as woman is to home- maker? Debiasing word embeddings. In NeurIPS â16, pages 4349â4357. Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183â186. Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, and Yejin Choi. 2025. CulturalBench: A robust, diverse and challenging benchmark for mea- suring LMsâ cultural knowledge through human-AI red-teaming. In ACLâ25, pages 25663â25701. Jenna Cryan, Shiliang Tang, Xinyi Zhang, Miriam J. Metzger, Haitao Zheng, and Ben Y. Zhao. 2020. De- tecting gender stereotypes: Lexicon vs. Supervised learning methods. In CHI â20, pages 1â11. Juntao Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2024. Safe RLHF: Safe Reinforcement Learning from Human Feedback. In ICLR â24. Xunlian Dai, Li Zhou, Benyou Wang, and Haizhou Li. 2025. From word to world: Evaluate and mitigate culture bias in LLMs via word association test. In EMNLPâ25, pages 24510â24526. Zhaopeng Feng, Jiayuan Su, Jiamei Zheng, Jiahan Ren, Yan Zhang, Jian Wu, Hongwei Wang, and Zuozhu Liu. 2025. M-MAD: Multidimensional multi-agent debate for advanced machine translation evaluation. In ACLâ25, pages 7084â7107. Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. 2024. Auditing large language models for race & gender disparities: Implications for artificial intelligence-based hiring. Behavioral Science & Policy, 10(2):46â55. Saibo Geng, Hudson Cooper, MichaĆ Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, and Harsha Nori. 2025. JSON- SchemaBench: A rigorous benchmark of struc- tured outputs for language models.Preprint, arXiv:2501.10868. Sourojit Ghosh and Kyra Wilson. 2025. Bias is a math problem, ai bias is a technical problem: 10-year lit- erature review of ai/llm bias research reveals narrow [gender-centric] conceptions of âbiasâ, and academia- industry gap. Proceedings of the AAAI/ACM Confer- ence on AI, Ethics, and Society, 8(2):1091â1106. Kristin Gnadt, David Thulke, Simone Kopeinik, and Ralf SchlĂŒter. 2025. Exploring Gender Bias in Large Language Models: An In-depth Dive into the German Language. In GeBNLP â25, pages 427â450. Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Pi- queras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, and An- ders SĂžgaard. 2022. Challenges and strategies in cross-cultural NLP. In ACLâ22, pages 6997â7013. Daniel L. Hicks, Estefania Santacreu-Vasut, and Amir Shoham. 2015.Does mother tongue make for womenâs work? Linguistics, household labor, and gender identity. Journal of Economic Behavior & Organization, 110:19â44. Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect. Nature, 633(8028):147â154. Bradley T. Hughes, Sanjay Srivastava, Magdalena Leszko, and David M. Condon. 2024. Occupational Prestige: The Status Component of Socioeconomic Status. Collabra: Psychology, 10(1):92882. Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the- box: An empirical analysis of intersectional occupa- tional biases in popular generative language models. In NeurIPSâ21, volume 34, pages 2611â2624. Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in Large Language Mod- els. In CI â23: Collective Intelligence Conference, pages 12â24. Keita Kurita, Nidhi Vyas, Ayush Pareek, Alan W. Black, and Yulia Tsvetkov. 2019. Measuring Bias in Contex- tualized Word Representations. In GeBNLP @ ACL â19, pages 166â172. Cheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024. CultureLLM: Incor- porating Cultural Differences into Large Language Models. In NeurIPSâ24, volume 37, pages 84799â 84838. Xinru Lin and Luyang Li. 2025. Implicit bias in LLMs: A survey. Preprint, arXiv:2503.02776. Kristian Lum, Jacy Reese Anthis, Kevin Robinson, Chi- rag Nagpal, and Alexander Nicholas DâAmour. 2025. Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation. In ACL â25, pages 137â 161. Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2025. Cultural alignment in large language models: An explanatory analysis based on hofstedeâs cultural dimensions. In COLINGâ25, pages 8474â8503. Astghik Mavisakalyan. 2015. Gender in Language and Gender in Employment. Oxford Development Stud- ies, 43(4):403â424. Fabio Motoki, Valdemar Pinho Neto, and Victor Ro- drigues. 2024. More human than human: measuring ChatGPT political bias. Public Choice, 198(1-2):3â 23. Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, VĂctor GutiĂ©rrez-Basulto, YazmĂn Ibåñez-GarcĂa, Hwaran Lee, Shamsuddeen Hassan Muhammad, Kiwoong Park, Anar Sabuhi Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, and 3 others. 2024. BLEnD: A benchmark for LLMs on everyday knowledge in diverse cultures and languages. In NeurIPS â24, volume 37, pages 78104â78146. Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereotypical bias in pretrained language models. In ACL â21, pages 5356â5371. Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. CrowS-Pairs: A Chal- lenge Dataset for Measuring Social Biases in Masked Language Models. In EMNLP â20, pages 1953â 1967. Tarek Naous, Michael J. Ryan, Alan Ritter, and Wei Xu. 2024. Having beer after prayer? measuring cultural bias in large language models. In ACLâ24, pages 16366â16393. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPSâ22, volume 35, pages 27730â27744. Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2025. Nor- mAd: A framework for measuring the cultural adapt- ability of large language models. In NAACLâ25, pages 2373â2403. Jill Walker Rettberg and Hermann Wigers. 2025. AI- generated stories favour stability over change: homo- geneity and cultural stereotyping in narratives gener- ated by gpt-4o-mini. Open Research Europe, 5:202. FabrizioSantoniccolo,TommasoTrombetta, Maria Noemi Paradiso, and Luca RollĂš. 2023. Gender and media representations: A review of the literature on gender stereotypes, objectifica- tion and sexualization.International Journal of Environmental Research and Public Health, 20(10):5770. Victor Steinborn, Philipp Dufter, Haris Jabbar, and Hin- rich SchĂŒtze. 2022. An Information-Theoretic Ap- proach and Dataset for Probing Gender Stereotypes in Multilingual Masked Language Models. In Find- ings of NAACL â22, pages 921â932. Naznin Tabassum and Bhabani Shankar Nayak. 2021. Gender stereotypes and their impact on womenâs ca- reer progressions from a managerial perspective. IIM Kozhikode Society & Management Review, 10(2):192â 208. Wenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, and Michael R. Lyu. 2024. Not all countries celebrate thanksgiving: On the cultural dominance in large language models. In ACLâ24, pages 6349â6384. Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, and 4 others. 2021. Ethical and social risks of harm from language models. Preprint, arXiv:2112.04359. Kyra Wilson and Aylin Caliskan. 2024. Gender, race, and intersectional bias in resume screening via lan- guage model retrieval. In AIES â24, volume 7, pages 1578â1590. Jieyu Zhao, Tianlu Wang, Mark Yatskar, Ryan Cotterell, Vicente Ordonez, and Kai-Wei Chang. 2019. Gender bias in contextualized word embeddings. In NAACL â19, pages 629â634. Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Or- donez, and Kai-Wei Chang. 2018.Gender bias in coreference resolution: evaluation and debiasing methods. In NAACL â18, pages 15â20. Qishuai Zhong, Yike Yun, and Aixin Sun. 2024. Cul- tural value differences of LLMs: Prompt, language, and model size. Preprint, arXiv:2407.16891. A Evidence Map for the Analytical Lenses Bias representation: Focus: genderâoccupation asso- ciations. Evidence: Stereotype and Balance across language and task cells, followed by interaction, model-coverage, de- nominator, and trial-level robustness analyses. Identity-linked semantic separation: Focus: protagonist-gender labels. Evidence: a full-text positive- control probe and paired direct-cue-deletion comparison eval- uated on held-out modelâoccupation groups, with a display- only shared UMAP projection and an illustrative narrative contrast. Cross-cultural patterns: Focus: included cultural markers. Evidence: narrative-weighted relevance with human validation of source-linked scores, self-context advantage, and source-language separability measured with a classification probe under original, translated, and named-entity-masked inputs. B Experiment Materials B.1 Corpus Dimensions The study covers 18 occupations: Attendant, Cashier, Chef, Manager, Teacher, Librarian, Principal, School Bus Driver, Nurse, Surgeon, Secretary, Accountant, HR, Receptionist, CEO, Lawyer, Developer, and Salesperson. It uses three task conditionsâExplicit Task, Story, and Newsâin English (EN), French (FR), and Chinese (ZH). Outputs come from 12 mod- els; exact model versions and language coverage appear in Appendix D. B.2 Cultural-Marker Taxonomy and Examples The classification agent maps retained markers to the follow- ing ten-category taxonomy. Examples are illustrative items ob- served in the corpus; translations are provided for non-English examples. Toponyms.New York; Silicon Valley; Paris; Quartier Latin (Latin Quarter);ćäșŹ (Beijing). Institutions. MIT; UniversitĂ© de Paris-Sorbonne; Biblio- thĂšque Nationale de France (National Library of France);ćäœć€§ćŠ (Harvard University). Anthroponyms.Ms.Brown; Monsieur LĂ©on;æć„¶ (Grandma Li). Cuisine. Apple Pie; Macaroni and Cheese; Croissant; CrĂȘpes (Crepes); Coq au Vin (Chicken in Wine);ć·è(Sichuan Cuisine);ć «ćźéž (Eight Treasure Duck). Rituals.Thanksgiving; Christmas; NoĂ«l (Christmas);äžç§ (Moon Festival). Religion. Church; Ăglise (Church);ćșćź(Temple);ç„ćș (Shrine). Material culture.Baseball Cards; Montre Ă Gousset (Pocket Watch); Moulin Ă Vent (Windmill);çźç (Abacus). Arts and media.Shakespeare; Grimmâs Fairy Tales; Théùtre de lâĂtoile;ć°çć (The Little Prince). Mythology.Dragon; Fairy Godmother; Tooth Fairy; FĂ©e (Fairy);ć€ć° (Phoenix);éŸç (Dragon King). Value orientations.Individualism; LibertĂ© (Liberty);éäœ ć„çź (Collective Dedication). C Bias-Representation Robustness The primary Type-I ANOVAs use all available outputs. Com- plementary specifications condition the probabilities on fe- male or male labels and repeat both measure definitions on the nine models observed in every language. Trial-level binomial GLMs on the common-model subset model fe- male versus male protagonist labels and whether a protag- onist label is identified, using languageâgenre interactions, model, and occupation terms with cluster-robust covariance by modelâoccupation group. Both trial-level analyses detect the languageâgenre interaction and model terms (allp < 0.001). The languageâtask interaction is detected for both aggregate metrics in every specification. Table 2: Adjusted model-termp-values across the Type- I ANOVA specifications. The languageâtask interac- tion remains significant for both metrics in every speci- fication. SpecificationBalance Stereotype All available; all outputs.047.056 All available; identified only.144.023 Common nine; all outputs.398.048 Common nine; identified only.540.021 D Methodological and Validation Details D.1 Agent Roles and Generation Settings The extraction agent identifies protagonist attributes and pro- poses short, text-grounded candidate cultural markers while excluding occupation terms. The classification agent makes a binary inclusion decision (Keep/Remove) and maps included cultural markers to a ten-category taxonomy: place names (toponyms), institutions, personal names (anthroponyms), cui- sine, rituals, religion, material culture, arts and media, mythol- ogy, and value orientations. A candidate enters the analysis when it is classified as a non-generic, culturally specific en- tity, practice, or concept and has scores available for all three contexts. The cultural-scoring agent rates each included cul- tural marker separately against the EN-, FR-, and ZH-linked contexts on a 1â7 scale (1â2: not commonly associated; 3â 4: somewhat associated; 5â7: strongly associated), based on meaning rather than original-language wording. The trans- lation agent produces an English rendering of non-English markers. Marker-inclusion classification and three-context scoring useanthropic/claude-3.5-sonnet. Full prompt templates and the extraction and translation implementation are provided with the accompanying materials. Generation uses a unified Chat Completion API 3 with temperature 0.7 and a target of 50 outputs per occupationâ languageâtask cell. The FR pool comprises nine models; FR outputs are unavailable forQwen3-32B,DeepSeek-R1-0528, and DeepSeek-V3.1-Terminus. Exact model versions are: gpt-4o-2024-11-20 gpt-4o-mini-2024-07-18 mistral-large-2411 Mistral-Small-3.1-24B-Instruct-2503 Mistral-Nemo-Instruct-2407 Llama-3.3-70B-Instruct Llama-3.1-8B-Instruct Qwen3-32B DeepSeek-R1-0528 DeepSeek-V3.1-Terminus Grok-4 Grok-4-Fast 3 https://openrouter.ai D.2 Human Validation Recruitment and compensation: We recruited adult, source-language readers through Prolific and admin- istered the studies in Qualtrics. Recruitment was language- matched for EN, FR, and ZH; the cultural-marker study required native speakers, and the protagonist-label study used separate language-specific recruitment. Rewards were fixed and advertised before consent. Their hourly equiva- lents met or exceeded Prolificâs recommended rate at recruit- ment (GBP 9/hour). 4 The cultural-marker study paid at least GBP 15/hour and had a median completion time of 16.6 min- utes. The protagonist-label study was advertised as approxi- mately 15 minutes and had a retained-sample median of 10.5 minutes. Payment decisions were kept separate from analyti- cal inclusion and followed the advertised terms and Prolific policy; disagreement with the automated labels was never an exclusion criterion. Consent, ethics, and data handling: Before ei- ther task, participants saw a localized overview describing the research purpose, the AI-generated materials they would read, the task modules, expected duration, compensation, and data handling. They then chose either âYes, I consent to par- ticipateâ or âNo, I do not consentâ; selecting No ended the survey. Consent was followed by a short language-matched reading-comprehension check. The protocol was approved by the relevant institutional ethics-review body before recruit- ment; identifying institutional information is omitted during anonymous review. The surveys requested no names or con- tact details. They collected only country of upbringing and years lived in an environment where the source language is spoken, in addition to task responses and Prolific submission identifiers. Analysis used pseudonymous rater identifiers, and demographic information is reported only in aggregate. The occupation-association module explicitly stated that it con- cerned perceived social associations, not participantsâ personal beliefs or the occupationsâ actual gender distributions. Participant-facing instructions:All substantive in- structions were provided in the participantâs selected language; the complete localized wording is included in the accompa- nying Qualtrics survey files. For the cultural-marker task, participants were told that a marker is a specific word, phrase, or concept indicating a particular socio-cultural context, in- cluding localized artifacts, norms, idioms, slang, and culturally specific titles. They were instructed to judge the original text, to distinguish such markers from generic terms, and to rate relevance to the context they know on a 1â7 scale. For the protagonist task, the central individual was defined as the per- son whose actions, experience, or professional role organizes the text. Participants were explicitly instructed not to infer gender from occupation, name, nationality, or expectation, and instead to use only pronouns, titles, grammatical mark- ing, or direct identity statements expressed in the text. Six worked examples covered female, male, unclear, non-binary, multiple-protagonist, and no-individual cases. D.2.1 Cultural-Marker Study The study retains 39 raters: 15 EN, 11 FR, and 13 ZH native speakers. Participants completed the same localized three- stage procedure: 1. Consent and comprehension check.Participants provided informed consent and completed a reading- comprehension check. 2.Definitions and training. Participants received the definition of a cultural marker and the 1â7 relevance 4 Prolific payment guidance. scale, followed by three training items with immediate feedback. 3.Formal evaluation. Qualtrics randomly and evenly presented each participant with eight narratives from the relevant language-specific pool. Each case showed an AI-generated narrative and two extracted candidate markers. Participants made a Keep/Remove decision for each marker and rated its relevance to their own language-linked context from 1 (not at all) to 7 (very strongly), yielding 16 marker-level judgments per partic- ipant. A direct attention check and the two demographic questions followed. The resulting 624 comparisons cover 178 cultural markers sampled from the included set. Intervals use 2,000 bootstrap resamples at the marker level, and 10,000 within-language shuffles of automated scores across items provide permutation baselines. The pooled humanâhuman agreement baseline comprises 834 pairs over 176 markers (Îș = 0.376[0.282, 0.463]; MAE= 1.79[1.64, 1.95]). The high-agreement subset requires at least two valid ratings per marker and a human-rating range of at most one scale point. Table 3: Agreement between automated and human cultural-marker relevance scores. MetricEstimate 95% interval Within two scale points.641[.593, .688] Quadratic-weighted Îș.280[.213, .352] Mean absolute error2.09[1.91, 2.28] Spearman correlation.555[.466, .641] D.2.2 Protagonist-Label Study Each participant labeled 12 blinded narratives and then com- pleted one unambiguous gold item, an 18-occupation associ- ation matrix, a direct attention check, and the demographic questions. The formal sample contains 120 implicit narratives per language, stratified into the six genre-by-automated-label cells (Story/NewsĂfemale/male/unknown). Qualtrics evenly sampled two of the 20 items in each cell for every participant; automated labels were hidden. Raters chose among seven fine- grained outcomes: female, male, unstated/unclear, non-binary or another identity, multiple central individuals, no central individual, or unable to decide because the text was incom- plete or ill-formed. The last five outcomes were mapped to unknown only after annotation. We received 120 completed submissions and retained 116 raters (49 EN, 36 FR, and 31 ZH), yielding 1,392 formal judg- ments. Quality rules were applied before comparison with the automated labels: a response required consent, survey comple- tion, a correct comprehension check, all 12 formal judgments, a correct training-rule item, gold item, and attention check, and no joint speed flag (total duration below seven minutes and median narrative-page time below five seconds). Two EN responses failed the training-rule item, one EN response failed the gold item, and one FR response failed the gold item; no response was removed for incompleteness, comprehension, attention, speed, frequent use of unknown, low confidence, or disagreement with the automated label. An item-level human reference requires a clear majority among at least two valid ratings; 352 of 360 sampled items were resolved. We apply post-stratification, weighting items to match the corresponding genre-by-label distribution in the full implicit corpus. Human inter-rater reliability is Fleissâ Îș = 0.72â0.82, and weighted automatedâhumanÎș = 0.72â 0.86. Table 4: Geographic and language-environment char- acteristics of retained protagonist-label raters. EN, FR, and ZH responses span 5, 3, and 6 normalized country groups, respectively. Lang. nLargest upbringing group Years, median [range] EN49 UK: 19 (38.8%)36 [5, 64] (n=48) FR36 France: 29 (80.6%)26.5 [18, 58] ZH31 China: 26 (83.9%)24 [12, 43] We did not collect additional demographic attributes be- cause they were not required for the language-matched val- idation. The cultural-marker study analogously records the rater counts by language above and collected the same two background variables. D.3 Metric Illustration For a female-associated occupation with 40 female and 10 male protagonists, Eq. 1 givesM s,o = 0.8â 0.2 = 0.6. For an occupation with one female and 49 male protagonists, Eq. 2 gives M b,o = 0.02â 0.98 =â0.96. D.4Identity-Probe and Visualization Settings The paired cue-deletion analysis, which provides the main identity-probe comparison, samples 100 female- and 100 male- labeled narratives per languageâgenre cell. It deletes the ex- tracted protagonist name, reusable multi-character name com- ponents, and a prespecified multilingual list of direct gender terms without inserting placeholders; full-text and cue-deleted probes use the same folds. Paired AUC changes use 2,000 bootstrap resamples at the modelâoccupation-group level. The shared UMAP projection uses cosine distance, 15 neighbors, minimum distance 0.1, and seed 42. All sampled narratives enter the projection, whereas probes use the female- and male- labeled subset. For display, Figure 3 omits points outside the 15thâ90th percentile range on either coordinate within each label and cell. D.5 Cross-Cultural Probe Settings EN-source markers are left unchanged in the translation con- dition. Five-fold stratified group cross-validation groups iden- tical normalized English translations, preventing the same translated content from appearing in train and test folds; all three marker versions use the same folds. Intervals use 2,000 bootstrap resamples of translation groups. Named entities are detected with spaCyen_core_web_md; person, nationality or group, facility, organization, geopolitical, location, product, event, work-of-art, law, and language spans are replaced by the common tokenentity. The probe uses markers with an available English translation.