Paper deep dive
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
Alberto Testoni, Iacer Calixto
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/27/2026, 6:38:37 AM
Summary
This research investigates how social identity markers, specifically sexual orientation and religious affiliation, impact the accuracy and uncertainty calibration of Large Language Models (LLMs) in medical Question Answering (QA). Using 2,364 MedQA-USMLE questions and their counterfactual variants, the study demonstrates that 'homosexual' markers and intersectional identities (e.g., combining sexual orientation with religion) significantly degrade both model accuracy and the reliability of confidence signals (measured via Brier score and semantic entropy). These failures persist in both multiple-choice and open-ended generation settings, posing risks for clinical deployment where confidence-based deferral to clinicians is required.
Entities (8)
Relation Signals (5)
Intersectional identities â causesnonadditiveharmto â Calibration
confidence 100% · intersectional identities produce idiosyncratic, non-additive harms to calibration.
Sexual Orientation â distorts â Large Language Models
confidence 100% · investigates how social descriptors of a patient (specifically sexual orientation and religious affiliation) distort these uncertainty signals and model accuracy.
Religious Affiliation â distorts â Large Language Models
confidence 100% · specifically sexual orientation and religious affiliation) distort these uncertainty signals and model accuracy.
Homosexual markers â triggersdropin â Accuracy
confidence 100% · âHomosexualâ markers consistently trigger performance drops
MedQA-USMLE-4-options â usedtoevaluate â Large Language Models
confidence 100% · We use MedQA-USMLE-4-options... We evaluate nine LLMs on QA accuracy and semantic-entropy uncertainty calibration.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Safe clinical deployment of Large Language Models (LLMs) requires not only high accuracy but also robust uncertainty calibration to ensure models defer to clinicians when appropriate. Our paper investigates how social descriptors of a patient (specifically sexual orientation and religious affiliation) distort these uncertainty signals and model accuracy. Evaluating nine general-purpose and biomedical LLMs on 2,364 medical questions and their counterfactual variants, we demonstrate that identity markers cause a "calibration crisis". "Homosexual" markers consistently trigger performance drops, and intersectional identities produce idiosyncratic, non-additive harms to calibration. Moreover, a clinician-validated case study in an open-ended generation setting confirms that these failures are not an artifact of the multiple-choice format. Our results demonstrate that the presence of social identity cues does not merely shift predictions; it affects the reliability of confidence signals, posing a significant risk to equitable care and safe deployment in confidence-based clinical workflows.
Tags
Links
- Source: https://arxiv.org/abs/2604.17316v1
- Canonical: https://arxiv.org/abs/2604.17316v1
Trouble viewing inline? Open PDF directly â
Full Text
84,094 characters extracted from source content.
Expand or collapse full text
Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA Alberto Testoni1,2, Iacer Calixto1,2 1Department of Medical Informatics, Amsterdam University Medical Center, University of Amsterdam, Amsterdam, The Netherlands. 2Amsterdam Public Health, Methodology, Amsterdam, The Netherlands. Correspondence: a.testoni@amsterdamumc.nl; i.coimbra@amsterdamumc.nl Abstract Safe clinical deployment of Large Language Models (LLMs) requires not only high accuracy but also robust uncertainty calibration to ensure models defer to clinicians when appropriate. Our paper investigates how social descriptors of a patient (specifically sexual orientation and religious affiliation) distort these uncertainty signals and model accuracy. Evaluating nine general-purpose and biomedical LLMs on 2,364 medical questions and their counterfactual variants, we demonstrate that identity markers cause a âcalibration crisisâ. Homosexual markers consistently trigger performance drops, and intersectional identities produce idiosyncratic, non-additive harms to calibration. Moreover, a clinician-validated case study in an open-ended generation setting confirms that these failures are not an artifact of the multiple-choice format. Our results demonstrate that the presence of social identity cues does not merely shift predictions; it affects the reliability of confidence signals, posing a significant risk to equitable care and safe deployment in confidence-based clinical workflows. Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA Alberto Testoni1,2, Iacer Calixto1,2 1Department of Medical Informatics, Amsterdam University Medical Center, University of Amsterdam, Amsterdam, The Netherlands. 2Amsterdam Public Health, Methodology, Amsterdam, The Netherlands. Correspondence: a.testoni@amsterdamumc.nl; i.coimbra@amsterdamumc.nl 1 Introduction Large language models (LLMs) are increasingly integrated into clinical workflows, from patient-facing communication to decision support (Rajpurkar et al., 2022; Artsi et al., 2025). However, high benchmark accuracy alone does not ensure safe deployment. In practice, clinical systems often rely on a modelâs confidence score to triage cases, trigger escalation, or defer to clinicians (Dvijotham et al., 2023). Therefore, reliability becomes a first-class requirement: models must be accurate and well calibrated, so that higher confidence corresponds to higher likelihood of correctness, and their confidence signals should remain stable under benign input variations (Kuzucu et al., 2024). Healthcare further amplifies these concerns: if sensitive identity cues systematically affect clinical decisions or uncertainty estimates, they risk inequitable and unsafe patient care (Zack et al., 2024). Figure 1: Counterfactual clinical vignettes that differ only in social identity descriptors induce marked shifts in LLM outputs and confidence estimates; we further examine how these effects extend to open-ended QA. A growing body of research suggests that social descriptors can influence LLM-generated clinical recommendations. For example, incorporating sociodemographic attributes (e.g., race, sex) alters model outputs in clinical trial matching and QA (Ji et al., 2025). Moreover, recent work indicates that LLMs can propagate bias related to LGBTQIA+ and religious individuals (Hirsch et al., 2026; Chang et al., 2025; Abid et al., 2021), including along intersectional demographic axes (Ivetta et al., 2025). Recent studies further show that LLMs can infer demographic attributes from subtle cues and adapt their behavior accordingly, even without explicit identity information (Neplenbroek et al., 2025). Prior works, however, do not assess how identity markers shape model uncertainty. This issue becomes increasingly salient when descriptors are combined, as intersectional effects often exceed individual cues, a phenomenon well-documented in the social sciences (Collins, 2019). Addressing this gap requires tools that can reliably quantify model uncertainty. Recent advances in uncertainty estimation (UE) for LLMs, driven by hallucination detection and reliability concerns (Ye et al., 2024), provide such a framework. Among available methods, Semantic entropy is particularly well-suited to this setting: by quantifying predictive uncertainty over semantically distinct outputs rather than surface forms (Kuhn et al., 2023; Farquhar et al., 2024), it provides a principled proxy for model uncertainty. Although computationally demanding, it consistently shows competitive performance among UE methods (Lin et al., 2024; Vashurin et al., 2025; Testoni and Calixto, 2026) and has recently been validated in clinical settings (Penny-Dimri et al., 2025). Our paper brings these perspectives together by investigating how clinically non-diagnostic social identity markers affect LLM performance and semantic entropy calibration (see Figure 1 for an illustration). We focus on sexual orientation and religion, which remain underexplored despite appearing in clinical notes through histories and referrals (Lynch et al., 2020; Bragazzi et al., 2022). Both are salient axes in healthcare: sexual and gender minorities face well-documented disparities and barriers to care (Dahlhamer et al., 2016; Grasso et al., 2020), and religious affiliation (or its lack thereof) can reinforce disparities (Sinclair and Rosielle, 2020; Scheitle et al., 2023; Rahman et al., 2024). We consider 2.3k United States Medical Licensing Examination (USMLE) questions from MedQA (Jin et al., 2021) and construct counterfactual variants by inserting a single sentence specifying sexual orientation and/or religious affiliation into the vignette. We evaluate nine LLMs on QA accuracy and semantic-entropy uncertainty calibration. While prior work perturbs questions to study cognitive biases (Schmidgall et al., 2024), accuracy gaps across gender and ethnicity (Rawat et al., 2024), or overconfidence under ambiguity (Testoni et al., 2025), we instead examine how identity cues related to sexual orientation and religious affiliation influence correctness and, critically, the calibration of model uncertainty. Our results reveal a systematic degradation across all nine LLMs. âHeterosexualâ insertions act as a near-neutral baseline, whereas âhomosexualâ cues trigger consistent accuracy drops and, crucially, degrade uncertainty calibration and the reliability of confidence estimates. These effects compound under intersectional identities: combining sexual orientation with religious descriptors often induces harms that exceed the additive effects of each cue alone. To test whether these findings are specific to the multiple-choice format, we further introduce a clinician-validated evaluation setup that converts questions to open-ended generation. Results confirm that âhomosexualâ insertions reduce both accuracy and calibration in this setting as well, indicating that identity-driven distortions extend beyond constrained answer formats. 2 Methodology Data and counterfactual variants. We use MedQA-USMLE-4-options, a curated subset of MedQA (Jin et al., 2021) retaining only USMLE questions in English with four answer choices, as distributed by Baker (2023). We consider adult patient vignettes (patient age specified and â„ 18 years old), excluding questions already mentioning sexual orientation, religion, or psychiatry (the latter analyzed separately in Appendix A.4). We randomly sample 2,364 questions and generate counterfactual variants by adding a single templated sentence to the vignette (just before its final sentence) stating the patientâs (i) sexual orientation (The patient identifies as heterosexual / homosexual), (i) religion (The patient is Catholic / Muslim / atheist), or (i) both attributes jointly. While these traits exist on a broad spectrum, in this paper we focus on exemplar identities to maintain experimental control. As for religion, we focus on three illustrative categories reflecting distinct socio-cultural axes: a dominant Western religion and two identities often linked to LLM bias, as discussed in the Introduction. Broader identity coverage is an important direction that we leave for future work. In our main experiments, identity attributes are inserted as a stand-alone sentence placed before the final question. In Appendix B.1, we also show that an alternative embedded formulation (e.g., âA 45-year-old patient who identifies as homosexual comes to the physician âŠâ) yields similar patterns. (A) Accuracy (â is better) (Base in %, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 55.58 +0.72 -0.33 +1.15 +0.09 +0.60 !37-3.46 -1.31 !32-2.66 Qwen3-4B 56.77 -1.52 !36-3.39 +0.51 -1.06 -0.51 -1.10 !27-1.86 !30-2.46 Bio-Medical-Llama-3-8B 64.21 -1.60 !30-2.37 -1.10 !30-2.41 -0.72 !50-5.58 !43-4.44 !42-4.27 Llama-3.1-8B 57.57 -1.65 !31-2.62 -1.52 -1.31 -0.12 !41-4.19 !29-2.32 !36-3.38 Qwen3-30B 73.39 -0.25 -0.97 -0.97 !25-1.60 -0.21 -1.39 !29-2.28 !26-1.73 Llama-3.1-70B 84.31 !26-1.74 !33-2.92 -0.77 -1.48 -1.10 !37-3.47 !27-1.95 !33-2.84 OpenBioLLM-70B 77.44 !32-2.65 !60-7.21 -1.86 -2.06 -2.19 !47-5.10 !32-2.65 !39-3.78 GPT-4.1-mini 78.43 -1.40 !34-3.05 !25-1.53 -1.15 -0.55 !30-2.46 !37-3.56 !33-2.88 GPT-5.1 89.21 -0.80 !23-1.35 -0.84 -0.63 -0.67 -0.59 !24-1.44 !24-1.52 (B) Brier score (â is better) (Base raw, others: relative % change; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 0.24 +1.0% !20+3.7% -0.6% -0.8% +0.3% !23+5.4% !23+5.6% !22+4.9% Qwen3-4B 0.28 +1.1% +3.2% +0.4% +0.9% +1.2% +2.9% !21+4.1% +1.7% Bio-Medical-Llama-3-8B 0.21 !27+8.5% !35+14.1% !22+4.7% !25+7.3% !22+4.9% !31+11.2% !35+14.3% !30+10.4% Llama-3.1-8B 0.20 +0.4% !22+5.1% +1.1% +1.2% +1.0% !24+6.8% !25+7.2% !24+6.4% Qwen3-30B 0.17 +1.6% !23+6.0% +2.9% +1.3% !20+3.9% !23+5.4% !23+5.6% !23+5.5% Llama-3.1-70B 0.08 !34+13.6% !56+29.3% !25+7.4% !30+10.9% !35+14.5% !60+32.1% !53+26.8% !55+28.6% OpenBioLLM-70B 0.10 !34+13.5% !60+32.0% !30+11.0% !28+9.6% !37+15.9% !56+29.5% !53+26.8% !50+24.7% GPT-4.1-mini 0.12 !28+9.0% !30+10.8% !21+4.2% !21+3.9% +0.2% !26+7.8% !32+12.1% !32+11.9% GPT-5.1 0.07 -1.6% +6.0% !24-6.3% -1.5% -0.3% +1.1% +4.8% +2.7% (C) Confidence (Base is 1â1-normalized uncertainty, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 64.78 -0.53 !19-1.31 !37-2.53 !23-1.57 !16-1.08 !40-2.74 !15-1.01 !26-1.81 Qwen3-4B 74.56 -0.46 !40-2.73 -0.11 !14-0.97 -0.43 !27-1.82 !20-1.40 !30-2.09 Bio-Medical-Llama-3-8B 74.47 !21-1.43 !26-1.75 !15-1.04 !14-0.98 !12-0.84 !38-2.64 !33-2.24 !36-2.46 Llama-3.1-8B 59.98 !25-1.69 !50-3.43 !35-2.38 !30-2.06 !20-1.37 !59-4.02 !40-2.74 !38-2.61 Qwen3-30B 83.23 -0.62 !17-1.17 !19-1.31 !14-0.93 -0.36 -0.39 -0.75 !13-0.90 Llama-3.1-70B 85.88 !14-0.95 !37-2.56 !15-1.04 !23-1.59 !15-1.06 !32-2.19 !33-2.27 !30-2.07 OpenBioLLM-70B 74.93 !29-1.99 !80-5.49 !35-2.41 !64-4.41 !40-2.76 !72-4.96 !67-4.59 !53-3.65 GPT-4.1-mini 91.49 +0.14 -0.13 +0.20 -0.09 -0.08 -0.39 -0.18 -0.09 GPT-5.1 92.67 !12-0.81 !19-1.27 !10-0.70 !14-0.94 !8-0.55 !12-0.84 !17-1.14 !13-0.86 Table 1: Effects of identity insertions on multiple-choice QA accuracy and uncertainty. Columns correspond to counterfactual vignette insertions for sexual orientation, religion, and their combinations. Shaded cells denote paired differences that are statistically significant ( pâ€0.05p†0.05) relative to the original âBaseâ question, while underlined values denote statistically significant differences between ++homo and ++hetero within the same model (McNemar test for accuracy, paired bootstrap test for the others; pâ€0.05p†0.05). For accuracy and Brier score, green/red shading indicates improvement/worsening; confidence uses yellow shading only to mark significance. Accuracy (â is better) Brier score (â is better) Confidence Base +hetero +homo Base +hetero +homo Base +hetero +homo Llama-3.1-8B 37.12 -1.51 !37-3.56 0.32 -0.0% -3.3% 70.83 !21-1.43 !57-3.93 Qwen3-30B 50.47 +0.17 !26-1.82 0.38 -1.6% +1.0% 87.87 -0.46 !23-1.60 GPT-5.1 69.25 -0.64 !39-3.81 0.24 -4.4% !22+5.3% 90.15 !26-1.81 !33-2.29 Table 2: Open-ended case study. Accuracy, Brier score, and semantic-entropy confidence for base vs. sexual-orientation insertions across three representative models. Base shows absolute values; others are paired change. Shaded cells denote paired differences that are statistically significant relative to Base, while underlined values denote statistically significant differences between ++homo and ++hetero within the same model (pâ€0.05p†0.05). Inference, uncertainty, and metrics. We evaluate nine LLMs spanning both open-weight and closed-source models, including general-purpose and biomedical variants (for details, refer to Appendix A.1). For multiple-choice inference, models are prompted with the vignette (either original or counterfactual), question, four options, and asked to output only the chosen option letter in square brackets (e.g., [B]; see Appendix A.2 for more details). We extract the selected option using a regular expression and mark predictions that do not match the ground truth as incorrect. We use semantic entropy to quantify uncertainty, as it is the most consistently calibrated UE method in clinical QA (Penny-Dimri et al., 2025; Testoni and Calixto, 2026). For each question, we draw K=10K=10 model outputs using top-p=0.9p=0.9 and T=0.7T=0.7 (reshuffling the answer options at each generation), extract the selected answer option, and form an empirical option distribution p. We compute normalized entropy H~â(p)=Hâ(p)/logâĄ4 H(p)=H(p)/ 4 and define confidence c=1âH~â(p)c=1- H(p). We assess uncertainty calibration using the Brier score (Brier, 1950), which measures the mean squared error between predicted probabilities and binary correctness outcomes; lower values indicate better calibrated uncertainty estimates. Additional details on the evaluation metrics and supplementary results using the expected calibration error (ECE, Guo et al., 2017) and the area under the ROC curve (AUROC, Hanley and McNeil, 1982) are reported in Appendix A.3. The same evaluation framework is used in the open-ended case study, with task-specific adaptations detailed in Section 4. 3 Accuracy and Calibration Results Accuracy Shifts across Identity Cues. As shown in Table 1A, the insertion of benign identity cues leads to a consistent decline in multiple-choice accuracy across most models. While âheterosexualâ (+hetero) insertions result in negligible shifts, âhomosexualâ (+homo) insertions trigger significant performance drops. OpenBioLLM-70B, for instance, exhibits a substantial accuracy decrease of â7.21%-7.21\% (pâ€0.05p†0.05). Only Llama-3.2-3B and Qwen3-30B show no statistically significant drop. Religious descriptors show more varied but generally negative trends, with +Mus (Muslim) and +Cat (Catholic) cues inducing only occasional significant drops. A key finding is the compounding effect of intersectional identities. When sexual orientation and religion are combined, performance degradation often exceeds that observed for either perturbation in isolation. For Bio-Medical-Llama-3-8B, the accuracy drop for +homo+Cat (â5.58-5.58) is significantly larger than the drops for +homo (â2.37-2.37) or +Cat (â1.10-1.10) alone. Crucially, combining +hetero with religious cues yields smaller and less consistent shifts, as discussed in Appendix B.2. Our findings echo insights from the social sciences that overlapping identity categories can produce idiosyncratic, distinct effects (Collins, 2019), underscoring the importance of engaging with research outside traditional NLP. Impact on Uncertainty and Calibration. Beyond raw accuracy, identity insertions severely compromise the reliability of model confidence. Table 1B shows consistent increases in Brier scores. Under the +homo condition, Llama-3.1-70B and OpenBioLLM-70B see relative increases of +29.3%+29.3\% and +32.0%+32.0\% in Brier scores, respectively, indicating a sharp decline in predictive calibration. Crucially, identity insertions degrade calibration even when accuracy shifts are modest. Once again, calibration and confidence degrade more strongly when multiple identity cues are combined, across most models. Interestingly, Table 1C reveals that models respond to insertions with a decrease in confidence. While this suggests that models âdetectâ a change in the input, the corresponding rise in Brier scores proves this increased uncertainty is insufficient to maintain calibration. The resulting degradation in uncertainty calibration is a clear warning signal: confidence-based deferral rules may either fail to catch errors or trigger excessive deferrals, ultimately eroding system utility (Dvijotham et al., 2023). A lightweight logprob-based uncertainty baseline (Appendix B.3) shows similar trends, confirming that identity-driven distortion persists with token-level uncertainty estimates. Frontier Models. Despite being the most robust model, GPT-5.1 shows significant calibration loss under +homo vs. +hetero, indicating residual sensitivity to sociodemographic markers even at 89.21%89.21\% accuracy. Interestingly, GPT-5.1 differs from other models along specific dimensions, with improved calibration with Catholic insertions and reduced sensitivity to intersectional cues. GPT-4.1-mini exhibits minimal confidence change under +homo (â0.13-0.13) despite a significant accuracy drop (â3.05-3.05), a decoupling that is particularly concerning for real-world deployment. 4 Case Study: Open-Ended Generation Task Conversion. To test whether sensitivity to identity markers is an artifact of the multiple-choice format, we focus on sexual orientation and conduct an exploratory case study converting questions into open-ended formulations and removing answer options. We use GPT-5-mini for question reformulation (prompt in App. B.4) and subsequent semantic clustering for entropy evaluation (more details in App. B.5), and both stages undergo manual review to ensure the original intent is preserved without introducing noise. Accuracy is measured by comparing open-ended outputs to the corresponding multiple-choice ground-truth labels, using the same model as a judge to assess whether the free-form response matches the gold answer (Bavaresco et al., 2025). This automated evaluation is validated against annotations from an intensive care clinician with broad expertise, yielding high agreement (89% raw agreement; Cohenâs Îș=0.78Îș=0.78). These results confirm the LLM-based judge as a reliable proxy for clinical correctness within this simplified setting (more details in App. B.6). Figure 2: Illustrative failure cases. Top: In a multiple-choice setting, Bio-Medical-Llama-3-8B flips from the correct diagnosis (osteoarthritis) to an unrelated one (gout) when the +homo marker is inserted. Bottom: In open-ended QA, GPT-5.1 shifts from the correct mechanism (hepatorenal syndrome from cirrhosis) to a stereotype-driven explanation (HIV-associated nephropathy) under the +homo condition, while maintaining high confidence. Open-Ended Results. As shown in Table 2, the performance degradation associated with âhomosexualâ insertions persists in open-ended questions. All three representative models exhibit statistically significant accuracy drops under +homo, with GPT-5.1 showing the largest decrease (â3.81-3.81). Consistent with the multiple-choice findings, we observe a significant reduction in model confidence. Critically, for GPT-5.1, this shift was accompanied by a 5.3%5.3\% increase in the Brier score, highlighting a pronounced degradation of calibration in frontier models. These results indicate that the influence of clinically non-actionable social descriptors likely extends beyond specific task formulations. 5 Qualitative Examples In Figure 2, we highlight two illustrative failure cases. In a multiple-choice setting, Bio-Medical-Llama-3-8B flips from the correct diagnosis (osteoarthritis) to an unrelated one (gout) under +homo. In open-ended QA, GPT-5.1 shifts from the clinically supported mechanism (hepatorenal syndrome) to HIV-associated nephropathy, a stereotype-driven explanation unsupported by the vignette, while maintaining high confidence. This is the most dangerous failure mode for confidence-based deferral, as confidence fails to flag the shift. These cases suggest that identity cues can activate associative pathways that override clinical evidence, motivating future work on interpretability, targeted mitigation, and clinician-led evaluation of the prevalence and clinical impact of such failures. 6 Conclusion Our paper shows that social identity markers systematically degrade LLM accuracy and uncertainty calibration in medical QA. Crucially, the combined effects of sexual orientation and religious affiliation exceed their individual impacts, highlighting an intersectional vulnerability that persists across multiple-choice and open-ended settings. Because clinical deployment relies on calibrated confidence for safe deferral, these identity-driven distortions pose a direct risk for inequitable and unsafe patient care. Moreover, qualitative evidence suggests that these failures may reflect stereotype-driven reasoning that overrides clinical evidence, rather than random noise. Simply stripping identity attributes is not a robust fix: such information can be clinically relevant, appears in real notes, and may be implicit (as further discussed in Appendix C). Our findings motivate counterfactual calibration as a standard robustness check for clinical readiness, as well as calibration-aware training or post-hoc approaches that enforce stability of predictions and confidence under clinically irrelevant identity cues. More broadly, our findings highlight the need for deferral policies and deployment safeguards explicitly validated for reliability across identity groups, supported by extensive clinician-led evaluation. Limitations Some limitations should be considered to contextualize our findings and inform future research. Template-based counterfactuals. Identity cues are injected via a fixed, template-based sentence. Real clinical notes and patient histories mention social context in heterogeneous ways (timing, salience, narrative style, etc.). The observed sensitivities could be larger or smaller under alternative placements, paraphrases, or when identity is implied rather than explicitly stated. Appendix B.1 shows an alternative strategy to inject identity attributes via embedded formulation in the opening sentence of the vignette. As a robustness check on lexical choice, in Appendix B.7 we replace âhomosexualâ with the more contemporary and broader umbrella term âgayâ, obtaining the same qualitative patterns reported in Table 1. Residual clinical relevance. Although the insertions are intended to be clinically non-diagnostic in this setup (i.e., extraneous to the clinical decision), some identity cues may be statistically associated with particular conditions, risk factors, or behaviors. Disentangling these mechanisms is beyond the scope of this work. Open-ended pipeline. The open-ended case study uses GPT-5-mini for question conversion, clustering, and judging. We compare free-form answers against the original multiple-choice ground-truth, which can under-credit clinically plausible alternative formulations. While a clinicianâs validation supports the reliability of this pipeline in our small-scale study, its generality beyond this setting requires further investigation. Model Coverage. We do not benchmark reasoning-oriented LLMs. Recent work suggests that explicit reasoning and CoT-style generation can surface or even amplify social stereotypes, which could interact with identity cues in ways not captured by our current setup (Wu et al., 2025; Cantini et al., 2025). We leave this analysis for future work. Uncertainty operationalization. We focus on semantic-entropy-based uncertainty given its consistent advantage over other methods and its validation in clinical settings (Penny-Dimri et al., 2025). However, computing semantic entropy requires multiple generations per input, making it computationally expensive and potentially unsuitable for real-time use. While a simpler logprobs-based uncertainty signal shows similar calibration degradation in our experiments, other uncertainty measures may respond differently to identity perturbations. Intersectional Depth. We explored binary intersections (e.g., orientation and religion), but true intersectionality involves many more axes, including race, socioeconomic status, age, and disability. Clinical Impact vs. Statistical Significance. We report statistically significant drops in accuracy and calibration, but the direct impact of a 7.2 percentage point drop on patient outcomes depends heavily on the specific clinical workflow and the human-in-the-loop deferral threshold. Future work should involve more extensive human-AI collaboration studies to quantify actual clinical harm. Ethical Considerations We acknowledge several ethical considerations in line with the ACL Code of Ethics. Representational Risks: As noted in our methodology, we use discrete, templated identity markers (e.g., âhomosexual,â âMuslimâ) to maintain experimental control. We recognize that these labels oversimplify the fluid and intersectional nature of human identity. There is an inherent risk that by testing only these specific markers, we may overlook unique biases faced by individuals with identities not represented in our study. Risk of Misinterpretation in Clinical Settings: Our findings demonstrate that identity cues can destabilize model confidence and accuracy. A primary ethical concern is the potential for these biased uncertainty signals to lead to inequitable triage: a patient from a marginalized group might be more or less likely to have their case deferred to a human clinician based on an unreliable model confidence score. This poses a direct risk to the principle of distributive justice in healthcare. Biases in Automated Evaluation: Our case study utilizes GPT-5-mini as a judge for correctness and clustering. While we validated this against clinician annotations, LLM-based judges may harbor their own sociodemographic biases that could penalize or favor specific groups. We have attempted to mitigate this through manual checks and expert validation, but we caution against using such automated pipelines in high-stakes clinical deployment without extensive validation. Data and Geographic Scope: MedQA-USMLE reflects a Western-centric context. We caution that our results should not be used to make broad claims about LLM safety in global healthcare contexts with different clinical norms and saliency of identity markers. Future work should assess whether these effects generalize to non-Western datasets and different sociocultural contexts. Acknowledgments This publication is part of the project CaRe-NLP with file number NGF.1607.22.014 of the research programme AiNed Fellowship Grants, which is (partly) financed by the Dutch Research Council (NWO). We thank Merijn Reuland for the invaluable help throughout the project. We are also grateful to Marc van der Valk, VinĂcius Mendes, and Davide Cevenini for their feedback and discussions. References A. Abid, M. Farooqi, and J. Zou (2021) Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, p. 298â306. Cited by: §1. Y. Artsi, V. Sorin, B. S. Glicksberg, P. Korfiatis, G. N. Nadkarni, and E. Klang (2025) Large language models in real-world clinical workflows: a systematic review of applications and implementation. Frontiers in Digital Health 7, p. 1659134. Cited by: §1. G. Baker (2023) MedQA-usmle-4-options. Hugging Face. Note: https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options Cited by: §2. A. Bavaresco, R. Bernardi, L. Bertolazzi, D. Elliott, R. FernĂĄndez, A. Gatt, E. Ghaleb, M. Giulianelli, M. Hanna, A. Koller, A. Martins, P. Mondorf, V. Neplenbroek, S. Pezzelle, B. Plank, D. Schlangen, A. Suglia, A. K. Surikuchi, E. Takmaz, and A. Testoni (2025) LLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 238â255. External Links: Link, Document, ISBN 979-8-89176-252-7 Cited by: §4. N. L. Bragazzi, R. Khamisy-Farah, M. Converti, and I. W. on LGBTIQ Mental Health (2022) Ensuring equitable, inclusive and meaningful gender identity-and sexual orientation-related data collection in the healthcare sector: insights from a critical, pragmatic systematic review of the literature. International Review of Psychiatry 34 (3-4), p. 282â291. Cited by: §1. G. W. Brier (1950) Verification of forecasts expressed in terms of probability. Monthly Weather Review 78 (1), p. 1. Cited by: §2. R. Cantini, N. Gabriele, A. Orsino, and D. Talia (2025) Is reasoning all you need? probing bias in the age of reasoning language models. arXiv preprint arXiv:2507.02799. Cited by: Limitations. C. T. Chang, N. Srivathsa, C. Bou-Khalil, A. Swaminathan, M. R. Lunn, K. Mishra, S. Koyejo, and R. Daneshjou (2025) Evaluating anti-lgbtqia+ medical bias in large language models. PLOS Digital Health 4 (9), p. e0001001. Cited by: §1. S. D. Cochran, J. G. Sullivan, and V. M. Mays (2003) Prevalence of mental disorders, psychological distress, and mental health services use among lesbian, gay, and bisexual adults in the united states.. Journal of consulting and clinical psychology 71 (1), p. 53. Cited by: §A.4. P. H. Collins (2019) Intersectionality as critical social theory. Duke University Press. Cited by: §1, §3. ContactDoctor (2024) ContactDoctor-bio-medical: a high-performance biomedical language model. Note: https://huggingface.co/ContactDoctor/Bio-Medical-Llama-3-8B Cited by: 3rd item. J. M. Dahlhamer, A. M. Galinsky, S. S. Joestl, and B. W. Ward (2016) Barriers to health care among adults identifying as sexual minorities: a us national study. American journal of public health 106 (6), p. 1116â1122. Cited by: §1. K. Dvijotham, J. Winkens, M. Barsbey, S. Ghaisas, R. Stanforth, N. Pawlowski, P. Strachan, Z. Ahmed, S. Azizi, Y. Bachrach, L. Culp, M. Daswani, J. Freyberg, C. Kelly, A. Kiraly, T. Kohlberger, S. McKinney, B. Mustafa, V. Natarajan, K. Geras, J. Witowski, Z. Z. Qin, J. Creswell, S. Shetty, M. Sieniek, T. Spitz, G. Corrado, P. Kohli, T. Cemgil, and A. Karthikesalingam (2023) Enhancing the reliability and accuracy of ai-enabled diagnosis via complementarity-driven deferral to clinicians. Nature Medicine 29 (7), p. 1814â1820. External Links: ISSN 1546-170X, Document, Link Cited by: §1, §3. S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal (2024) Detecting hallucinations in large language models using semantic entropy. Nature 630 (8017), p. 625â630. Cited by: §1. C. Grasso, H. Goldhammer, R. J. Brown, and B. Furness (2020) Using sexual orientation and gender identity data in electronic health records to assess for disparities in preventive health screening services. International journal of medical informatics 142, p. 104245. Cited by: §1. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. GuzmĂĄn, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Ăelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. arXiv preprint arXiv:2407.21783. External Links: 2407.21783, Link Cited by: 4th item, 6th item. C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger (2017) On calibration of modern neural networks. In International conference on machine learning, p. 1321â1330. Cited by: §2. J. A. Hanley and B. J. McNeil (1982) The meaning and use of the area under a receiver operating characteristic (roc) curve.. Radiology 143 (1), p. 29â36. Cited by: §2. M. Hirsch, M. Elichiry, B. Radi, T. Quiroga, D. Restrepo, L. Benotti, V. Xhardez, J. Dunstan, and E. Ferrante (2026) Implicit bias in LLMs for transgender populations. arXiv preprint arXiv:2602.13253. Cited by: §1. G. Ivetta, M. J. Gomez, S. Martinelli, P. Palombini, M. E. Echeveste, N. C. Mazzeo, B. Busaniche, and L. Benotti (2025) HESEIA: a community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 25095â25117. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1. Y. Ji, W. Ma, S. Sivarajkumar, H. Zhang, E. M. Sadhu, Z. Li, X. Wu, S. Visweswaran, and Y. Wang (2025) Mitigating the risk of health inequity exacerbated by large language models. npj Digital Medicine 8 (1), p. 246. Cited by: §1. D. Jin, E. Pan, N. Oufattole, W. Weng, H. Fang, and P. Szolovits (2021) What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11 (14), p. 6421. Cited by: §1, §2. M. King, J. Semlyen, S. S. Tai, H. Killaspy, D. Osborn, D. Popelyuk, and I. Nazareth (2008) A systematic review of mental disorder, suicide, and deliberate self harm in lesbian, gay and bisexual people. BMC psychiatry 8 (1), p. 70. Cited by: §A.4. L. Kuhn, Y. Gal, and S. Farquhar (2023) Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, Cited by: §1. S. Kuzucu, J. Cheong, H. Gunes, and S. Kalkan (2024) Uncertainty as a fairness measure. Journal of Artificial Intelligence Research 81, p. 307â335. Cited by: §1. Z. Lin, S. Trivedi, and J. Sun (2024) Generating with confidence: uncertainty quantification for black-box large language models. Transactions on Machine Learning Research 2024. Cited by: §1. K. E. Lynch, P. R. Alba, O. V. Patterson, B. Viernes, G. Coronado, and S. L. DuVall (2020) The utility of clinical notes for sexual minority health research. American Journal of Preventive Medicine 59 (5), p. 755â763. Cited by: §1. A. Meta (2024) Llama 3.2: revolutionizing edge ai and vision with open, customizable models. Meta AI Blog. Retrieved December 20, p. 2024. Cited by: 1st item. V. Neplenbroek, A. Bisazza, and R. FernĂĄndez (2025) Reading between the prompts: how stereotypes shape LLMâs implicit personalization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 20367â20400. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1. A. Pal and M. Sankarasubbu (2024) OpenBioLLMs: advancing open-source large language models for healthcare and life sciences. Hugging Face. Note: https://huggingface.co/aaditya/OpenBioLLM-Llama3-70B Cited by: 7th item. J. C. Penny-Dimri, M. Bachmann, W. R. Cooke, S. Mathewlynn, S. Dockree, J. Tolladay, J. Kossen, L. Li, Y. Gal, and G. Davis Jones (2025) Measuring large language model uncertainty in womenâs health using semantic entropy and perplexity: a comparative study. The Lancet Obstetrics, Gynaecology, & Womenâs Health 1 (1), p. e47âe56. External Links: ISSN 3050-5038, Document, Link Cited by: §1, §2, Limitations. R. Rahman, J. Lapum, and N. Prendergast (2024) âTreat me like a personâ: unveiling healthcare narratives of muslim women who wear islamic head coverings through a poststructural narrative study. Canadian Journal of Nursing Research 56 (4), p. 377â387. Cited by: §1. P. Rajpurkar, E. Chen, O. Banerjee, and E. J. Topol (2022) AI in health and medicine. Nature medicine 28 (1), p. 31â38. Cited by: §1. R. Rawat, H. McBride, R. Ghosh, D. Nirmal, J. Moon, D. Alamuri, S. OâBrien, and K. Zhu (2024) DiversityMedQA: a benchmark for assessing demographic biases in medical diagnosis using large language models. In Proceedings of the Third Workshop on NLP for Positive Impact, D. Dementieva, O. Ignat, Z. Jin, R. Mihalcea, G. Piatti, J. Tetreault, S. Wilson, and J. Zhao (Eds.), Miami, Florida, USA, p. 334â348. External Links: Link, Document Cited by: §1. M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh (2020) Beyond accuracy: behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (Eds.), Online, p. 4902â4912. External Links: Link, Document Cited by: Appendix C. A. R. Sarkar, Y. Chuang, N. Mohammed, and X. Jiang (2024) De-identification is not enough: a comparison between de-identified and synthetic clinical notes. Scientific reports 14 (1), p. 29669. Cited by: Appendix C. C. P. Scheitle, J. Frost, and E. H. Ecklund (2023) The association between religious discrimination and health: disaggregating by types of discrimination experiences, religious tradition, and forms of health. Journal for the scientific study of religion 62 (4), p. 845â868. Cited by: §1. S. Schmidgall, C. Harris, I. Essien, D. Olshvang, T. Rahman, J. W. Kim, R. Ziaei, J. Eshraghian, P. Abadir, and R. Chellappa (2024) Evaluation and mitigation of cognitive biases in medical language models. npj Digital Medicine 7 (1), p. 295. Cited by: §1. C. T. Sinclair and D. A. Rosielle (2020) Avoid stigmatizing language about atheist patients. Journal of Pain and Symptom Management 60 (6), p. e30. Cited by: §1. C. G. Streed Jr, C. Grasso, S. L. Reisner, and K. H. Mayer (2020) Sexual orientation and gender identity data collection: clinical and public health importance. American Journal of Public Health 110 (7), p. 991â993. Cited by: Appendix C. A. Testoni and I. Calixto (2026) Mind the gap: benchmarking LLM uncertainty and calibration with specialty-aware clinical QA and reasoning-based behavioural features. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 2364â2382. External Links: Link, Document, ISBN 979-8-89176-380-7 Cited by: §1, §2. A. Testoni, B. Plank, and R. FernĂĄndez (2025) RAcQUEt: unveiling the dangers of overlooked referential ambiguity in visual LLMs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 23627â23647. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1. R. Vashurin, E. Fadeeva, A. Vazhentsev, L. Rvanova, D. Vasilev, A. Tsvigun, S. Petrakov, R. Xing, A. Sadallah, K. Grishchenkov, A. Panchenko, T. Baldwin, P. Nakov, M. Panov, and A. Shelmanov (2025) Benchmarking uncertainty quantification methods for large language models with LM-polygraph. Transactions of the Association for Computational Linguistics 13, p. 220â248. External Links: Link, Document Cited by: §1. C. Wittgens, M. M. Fischer, P. Buspavanich, S. Theobald, K. Schweizer, and S. Trautmann (2022) Mental health in people with minority sexual orientations: a meta-analysis of population-based studies. Acta Psychiatrica Scandinavica 145 (4), p. 357â372. Cited by: §A.4. X. Wu, J. Nian, T. Wei, Z. Tao, H. Wu, and Y. Fang (2025) Does reasoning introduce bias? a study of social bias evaluation and mitigation in LLM reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 18534â18555. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: Limitations. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: 2nd item, 5th item. F. Ye, M. Yang, J. Pang, L. Wang, D. Wong, E. Yilmaz, S. Shi, and Z. Tu (2024) Benchmarking llms via uncertainty quantification. Advances in Neural Information Processing Systems 37, p. 15356â15385. Cited by: §1. T. Zack, E. Lehman, M. Suzgun, J. A. Rodriguez, L. A. Celi, J. Gichoya, D. Jurafsky, P. Szolovits, D. W. Bates, R. E. Abdulnour, A. J. Butte, and E. Alsentzer (2024) Assessing the potential of gpt-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet Digital Health 6 (1), p. e12âe22. External Links: ISSN 2589-7500, Document, Link Cited by: §1. Appendix Appendix A Methodology Appendix A.1 Model Details and Licensing Information We provide links to the modelsâ Hugging Face repositories and licensing terms. âą Llama-3.2-3B: https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct (Llama 3.2 Community License Agreement). (Meta, 2024). âą Qwen3-4B: https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507 (Apache License 2.0). (Yang et al., 2025). âą Bio-Medical-Llama-3-8B: https://huggingface.co/ContactDoctor/Bio-Medical-Llama-3-8B (Bio-Medical-Llama-3-8B LLM License, Non-Commercial Use Only). (ContactDoctor, 2024). âą Llama-3.1-8B: https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct (Llama 3.1 Community License Agreement). (Grattafiori et al., 2024). âą Qwen3-30B: https://huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507 (Apache License 2.0). (Yang et al., 2025). âą Llama-3.1-70B: https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct (Llama 3.1 Community License Agreement). (Grattafiori et al., 2024). âą OpenBioLLM-70B: https://huggingface.co/aaditya/Llama3-OpenBioLLM-70B (Llama 3 Community License Agreement). (Pal and Sankarasubbu, 2024). GPT-4.1-mini, GPT-5.1 (standard, non-reasoning), and GPT-5-mini (reasoning) were accessed via the OpenAI API. As proprietary models, their weights and training data are not publicly available; usage is governed by OpenAIâs terms of service, and redistribution of the models is not permitted. Our experiments rely solely on model outputs obtained through the API. We use the MedQA-USMLE dataset available at https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options and distributed under Creative Commons Attribution 4.0 license. A.2 Prompt Template For multiple-choice inference, each model is prompted with the full clinical vignette, the question, and the four answer options labeled with letters. The model is instructed to select the single best answer and to output only the corresponding option letter enclosed in square brackets, with no additional text. This strict output constraint simplifies downstream parsing and ensures consistent behavior across models and decoding runs. <clinical vignette and question> A. <option text> B. <option text> C. <option text> D. <option text> Choose the single best answer to the multiple-choice question above. Output only the letter of the selected option in square brackets, with no other text, for example: [X]. A concrete example (with identity insertion related to religious affiliation): A 75-year-old woman presents to her physician with a cough and shortness of breath. She says that cough gets worse at night and her shortness of breath occurs with moderate exertion or when lying flat. She says these symptoms have been getting worse over the last 6 months. She mentions that she has to use 3 pillows while sleeping in order to relieve her symptoms. She denies any chest pain, chest tightness, or palpitations. Past medical history is significant for hypertension and diabetes mellitus type 2. Her medications are amiloride, glyburide, and metformin. Family history is significant for her father who also suffered diabetes mellitus type 2 before his death at 90 years old. The patient says she drinks alcohol occasionally but denies any smoking history. Her blood pressure is 130/95 m Hg, temperature is 36.5°C (97.7°F), and heart rate is 100/min. On physical examination, she has a sustained apical impulse, a normal S1 and S2, and a loud S4 without murmurs. There are bilateral crackles present bilaterally. A chest radiograph shows a mildly enlarged cardiac silhouette. A transesophageal echocardiogram is performed and shows a normal left ventricular ejection fraction. The patient is Catholic. Which of the following myocardial changes is most likely present in this patient? A. Ventricular hypertrophy with sarcomeres duplicated in series B. Ventricular hypertrophy with sarcomeres duplicated in parallel C. Asymmetric hypertrophy of the interventricular septum D. Granuloma consisting of lymphocytes, plasma cells and macrophages surrounding necrotic Choose the single best answer to the multiple-choice question above. Output only the letter of the selected option in square brackets, with no other text, for example: [X]. Model Parsing rate (%) Llama-3.2-3B 93.0 Qwen3-4B 99.1 Bio-Medical-Llama-3-8B 99.2 Llama-3.1-8B 91.0 Qwen3-30B 98.9 Llama-3.1-70B 100.0 OpenBioLLM-70B 99.7 GPT-4.1-mini 93.8 GPT-5.1 100.0 Table 3: MCQA option parsing under sampling. For each question, we sample 10 responses and attempt to extract the selected answer option via a regex-based parser. We report the percentage of questions for which the option is successfully extracted in at least half of the generations (â„5/10â„ 5/10). Higher values indicate more consistent adherence to the expected answer format. As shown in Table 3, across models, the regex-based option parser succeeds for the large majority of questions. The lower rates for some models (notably Llama-3.1-8B and Llama-3.2-3B, around 91â93%) correspond to deviations such as missing brackets, multiple options, or free-form explanations without an explicit option token. The near-ceiling performance for several larger models (up to 100%) supports the robustness of the extraction protocol for most settings. We do not observe significant differences in option parsing across different input perturbations. A.3 Evaluation Metrics and AUROC, ECE Results (A) AUROC (â) (Differences: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 76.77 +0.11 +0.49 +0.53 +0.48 -0.09 +0.14 +0.29 -0.40 Qwen3-4B 69.88 +0.44 +1.05 -1.40 +1.30 +0.35 +0.56 -1.11 +0.82 Bio-Medical-Llama-3-8B 75.95 -1.08 -1.85 +1.09 -0.62 +0.65 -0.52 -1.49 -0.84 Llama-3.1-8B 83.45 -0.98 -0.70 +0.47 -1.47 -1.30 !37-1.81 !49-2.78 -1.68 Qwen3-30B 75.97 -0.01 -0.89 +0.06 -1.01 -1.06 -0.12 -1.99 -0.95 Llama-3.1-70B 86.03 -0.37 -2.20 +1.21 +1.06 -1.77 !46-2.55 -0.99 !49-2.82 OpenBioLLM-70B 90.13 -1.73 !54-3.21 -1.27 -1.15 -1.42 !59-3.63 !60-3.68 !52-3.01 GPT-4.1-mini 72.54 -2.49 -0.93 !49-2.78 -0.67 +0.18 -1.99 -1.11 -1.95 GPT-5.1 80.07 !35+3.67 +0.26 +2.79 +2.78 +2.73 +3.18 +2.13 !30+3.39 (B) ECE (â) (Base: raw score x 100, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 21.17 +0.38 !38+1.91 -0.39 -0.23 -0.14 !42+2.27 !47+2.69 !36+1.74 Qwen3-4B 25.81 +0.79 !38+1.91 -0.50 +1.00 +0.90 +1.43 +1.08 +1.13 Bio-Medical-Llama-3-8B 17.47 !39+2.03 !58+3.57 !39+1.98 !38+1.90 !36+1.72 !51+3.00 !60+3.76 !46+2.61 Llama-3.1-8B 17.27 -1.33 +0.81 +0.40 -1.42 -1.24 +0.50 +0.14 +0.79 Qwen3-30B 14.69 +0.41 +1.02 +0.21 -0.43 +0.63 !34+1.55 +0.48 +1.12 Llama-3.1-70B 4.88 +1.13 !37+1.85 +0.80 +0.73 +0.94 !44+2.41 !39+2.02 !34+1.55 OpenBioLLM-70B 5.03 +0.26 !37+1.87 +0.68 +0.75 +0.86 +1.43 +1.62 +1.13 GPT-4.1-mini 9.84 !26+0.96 !33+1.49 +0.20 +0.40 +0.02 +0.73 !36+1.75 !31+1.36 GPT-5.1 5.53 -0.28 +0.01 -0.79 -0.23 -0.04 +0.09 +0.05 +0.29 Table 4: Uncertainty discrimination and calibration under counterfactual identity insertions. We report AUROC and expected calibration error (ECE) for using the modelâs semantic-entropy confidence to distinguish correct from incorrect answers. Base shows AUROC/ECE results using the original patient vignette, while other columns report the paired difference (Î ) to Base. Shading indicates changes that are statistically significant versus Base (pâ€0.05p†0.05). Evaluation metrics and comparisons. We mainly evaluate the quality of model uncertainty estimates using the Brier score in the main text, complemented by expected calibration error (ECE) in the Appendix. The Brier score measures the mean squared error between predicted probabilities and binary correctness outcomes. Lower values indicate better calibrated and sharper uncertainty estimates. ECE complements the Brier score by partitioning predictions into M=10M=10 confidence bins and computing the weighted average of the absolute difference between empirical accuracy and mean predicted confidence within each bin. To assess discrimination, we report the area under the receiver operating characteristic curve (AUROC), which captures whether confidence scores rank correct answers above incorrect ones, independently of any fixed decision threshold. AUROC values closer to 1 indicate better separability. For accuracy-based analyses, we distinguish between aggregated and single-sample evaluations. When computing Brier, ECE, and AUROC with semantic entropy estimates, we rely on majority voting across the K=10K=10 generations per question to obtain an estimate of the modelâs predictive behavior. In the results tables, we report accuracy using a single randomly selected generation per question. This choice is intended to better reflect a realistic deployment scenario, where a system typically produces one answer rather than an ensemble. Reporting single-sample accuracy avoids overestimating performance through implicit ensembling, while majority-vote accuracy is reserved for semantic entropy-based analyses. To compare accuracy between conditions, we use McNemarâs test, which is appropriate for paired binary outcomes and evaluates whether two models (or variants) differ significantly in their errors on the same set of questions. For continuous metrics such as Brier score, ECE, and AUROC, we employ paired bootstrap resampling with 1,000 resamples. We assess statistical significance at α=0.05α=0.05. Clarification on Brier score interpretation. In our setup, each question yields (i) a confidence score câ[0,1]câ[0,1], computed from the empirical option distribution over K=10K=10 samples via normalized entropy (c=1âH~â(p)c=1- H(p)), and (i) a binary correctness label yâ0,1yâ\0,1\ for the modelâs prediction. The Brier score is then computed as the mean squared error 1Nââi=1N(ciâyi)2 1N _i=1^N(c_i-y_i)^2. The Brier score is a proper scoring rule that captures both calibration and sharpness of probabilistic predictions. While we use it as our primary metric for uncertainty quality, we do not interpret it as a pure measure of calibration. To provide a more complete view, we additionally report expected calibration error (ECE) and AUROC, which capture complementary aspects of calibration and discrimination. Accuracy (â) Brier Score (â) Confidence base question +hetero +homo base question +hetero +homo base question +hetero +homo Llama-3.2-3B-Instruct 64.30 -3.31% -2.21% 0.212 +3.35% +7.92% 69.08 -0.54% +0.47% Llama-3.1-8B 61.70 -2.67% -2.30% 0.218 -5.59% -3.48% 63.28 -1.36% !39-2.89% Bio-Medical-Llama-3-8B 70.92 -4.33% -4.99% 0.203 +5.13% +7.98% 78.86 -0.43% !37-2.63% GPT-4.1-mini 81.09 +0.00% -0.95% 0.119 +11.26% +11.51% 91.34 +0.84 +0.92% Table 5: Effects of sexual-orientation insertions on multiple-choice psychiatry and substance use-related questions. Columns other than âBaseâ report relative absolute (Accuracy and Confidence) or relative (Brier) changes versus Base, using Semantic Entropy to extract uncertainty estimates. Highlighted cells mark statistically significant differences vs. base. Underlined cells signify statistically significant differences (p†0.05) of +homo vs. +hetero. AUROC and ECE results. Table 4A complements the main calibration analyses by showing how identity cues affect discrimination, i.e., whether confidence still ranks correct answers above incorrect ones. While some smaller open-weight models show only modest AUROC shifts (and occasional gains), several stronger models exhibit clear degradation when identity cues are introduced, especially under +homo and intersectional variants. In particular, OpenBioLLM-70B shows large, significant AUROC drops for +homo (â3.21-3.21) and for all +homo+religion combinations (down to â3.68-3.68), indicating that identity insertions can erode not only calibration (as measured by the Brier score) but also the ranking quality of confidence signals. A similar, though smaller, pattern appears for Llama-3.1-70B (e.g., â2.20-2.20 for +homo, â2.82-2.82 for +homo+Ath). Conversely, GPT-5.1 shows consistent AUROC improvements across conditions (including a significant gain for +hetero, +3.67+3.67), suggesting more robust uncertainty discrimination under these perturbations. Importantly, these AUROC trends do not contradict the main finding that identity cues can still harm reliability: discrimination and calibration capture different failure modes, so a model may maintain (or even improve) ranking ability while its probability estimates become miscalibrated. Overall, the AUROC results reinforce the paperâs central point that clinically non-essential identity markers can destabilize confidence-based decision signals, with the most pronounced discrimination failures emerging under intersectional insertions. Table 4B also reports calibration via expected calibration error (ECE, â). In the main text, we focus on Brier score because it is a proper scoring rule that directly evaluates probabilistic accuracy without binning choices, making it more stable and comparable across settings; here, we add ECE as a complementary view, computed with 10 equal-width confidence bins. The ECE results largely mirror our main findings: identity insertions often worsen calibration, with the largest and most consistent increases under +homo and especially under intersectional +homo+religion variants. Overall, ECE reinforces that medically non-informative identity cues can distort calibration, and that these effects are often amplified when identity cues are combined. (A) Accuracy (â) (Base in %, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero+Cat +hetero+Mus +hetero+Ath Llama-3.1-8B 57.57 -1.22 -0.51 -1.69 Bio-Medical-Llama-3-8B 64.21 !28-2.11 !29-2.20 !36-3.34 Qwen3-30B 73.39 -0.22 -1.06 -0.77 Llama-3.1-70B 84.31 !32-2.71 -1.15 !26-1.82 OpenBioLLM-70B 77.44 !31-2.53 !34-3.06 !36-3.39 (B) Brier score (â) (Base raw, others: relative % change; shaded cells pâ€0.05p†0.05) Base +hetero+Cat +hetero+Mus +hetero+Ath Llama-3.1-8B 0.20 +2.1% +1.2% +1.7% Bio-Medical-Llama-3-8B 0.21 !26+8.2% !30+10.9% !27+8.9% Qwen3-30B 0.17 +1.5% !20+3.8% !20+3.7% Llama-3.1-70B 0.08 !42+19.2% !43+20.1% !45+21.6% OpenBioLLM-70B 0.10 !42+19.1% !41+18.3% !44+21.0% (C) Confidence (Base is 1â1-normalized uncertainty, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero+Cat +hetero+Mus +hetero+Ath Llama-3.1-8B 59.98 !37-2.53 !28-1.91 !19-1.34 Bio-Medical-Llama-3-8B 74.47 !24-1.62 !29-1.96 !27-1.84 Qwen3-30B 83.20 !15-1.00 !13-0.92 !12-0.81 Llama-3.1-70B 85.88 !24-1.62 !17-1.17 !13-0.92 OpenBioLLM-70B 74.90 !25-1.69 !39-2.69 !25-1.72 Table 6: Results for joint hetero+religion identities. Shaded cells: pâ€0.05p†0.05 vs. Base. (A) Brier score (â) (Differences: relative % change; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 0.24 -2.8% +0.2% -0.9% -0.3% -1.2% !21+4.5% !16+1.0% +1.7% Bio-Medical-Llama-3-8B 0.23 !19+2.9% !23+5.5% !21+4.5% !23+6.0% !16+1.0% !30+10.6% !28+9.2% !25+7.4% Llama-3.1-70B 0.11 !32+12.1% !48+23.5% !21+3.9% !24+6.4% !28+9.1% !50+25.2% !39+17.2% !46+21.8% OpenBioLLM-70B 0.15 !21+4.3% !48+23.7% !24+6.1% !29+10.1% !26+8.0% !37+15.7% !32+12.3% !36+14.8% (B) Confidence (Differences: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 68.16 -0.45 !12+0.80 !11+0.74 -0.06 !19-1.32 !6+0.43 !19-1.33 !18-1.27 Bio-Medical-Llama-3-8B 75.07 !52-3.56 !58-4.01 !41-2.83 !52-3.57 !37-2.56 !68-4.69 !66-4.51 !61-4.20 Llama-3.1-70B 92.98 !5-0.35 !21-1.42 !11-0.75 !16-1.12 !5-0.36 !16-1.08 !14-0.97 !10-0.70 OpenBioLLM-70B 86.44 !15-1.00 !35-2.38 !16-1.13 !16-1.13 !20-1.37 !33-2.26 !35-2.43 !23-1.58 (C) AUROC (â) (Differences: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Llama-3.2-3B 67.68 +0.91 +1.04 +0.56 +0.44 -0.21 !23+0.93 -0.19 +0.67 Bio-Medical-Llama-3-8B 63.94 !44-3.28 !60-5.10 !59-4.94 !56-4.65 !35-2.28 !59-5.00 !59-5.04 !47-3.66 Llama-3.1-70B 87.01 -1.07 -3.58 -0.32 +0.00 -1.49 !40-2.82 -3.38 !33-2.03 OpenBioLLM-70B 82.59 !25+1.10 !19-0.48 !23-0.89 !40-2.87 !26-1.26 !22-0.80 !44-3.29 !31-1.84 Table 7: Logprobs results using the log-likelihood of the token associated with the selected option letter in the multiple-choice QA setup. A.4 Psychiatry and Substance Use Questions Prior work documents elevated rates of mental and psychiatric disorders, including substance use disorders, among sexual-minority populations relative to heterosexual populations (King et al., 2008; Cochran et al., 2003; Wittgens et al., 2022). This evidence motivates a focused analysis on clinical domains where disparities are well established and where miscalibrated uncertainty may carry additional risk. We define a psychiatry-related subset of MedQA-USMLE, comprising 423 questions identified via keywords spanning neurocognitive conditions and substance use. Interestingly, in Table 5, we observe that +hetero and +homo insertions yield broadly similar shifts: most models show accuracy drops under both conditions, with only small between-orientation differences that rarely reach statistical significance. Calibration and confidence exhibit the same pattern: Brier changes and confidence deltas are generally comparable across +hetero vs. +homo, with model-specific fluctuations but no systematic separation between the two cues. This apparent convergence is noteworthy, but it should be interpreted cautiously given the reduced sample size and corresponding lower power to detect small effects. Future work should test whether effects differ when sexual-orientation mentions are clinically relevant (e.g., in risk-factor or psychosocial contexts) versus clearly extraneous. Appendix B Results Appendix B.1 Alternative Injection Method: Embedded Identity Attributes To assess whether our findings depend on the specific attribute injection method, we compare the stand-alone approach used throughout the paper (where identity attributes appear in a separate sentence preceding the final question) with an embedded variant, where attributes are integrated into the opening phrase (e.g., âA 45-year-old patient who identifies as homosexualâŠâ). Table 8 reports accuracy for both methods across representative models and identity configurations. Both injection strategies produce performance drops for intersectional insertions that are comparable to the main experiments, and reproduce the same directional biases observed with single identifiers. In particular, +homo consistently degrades accuracy more than +hetero. Effect sizes under the embedded method tend to be smaller, suggesting that these phrased cues are somewhat less disruptive, but the qualitative patterns remain stable. For GPT-4.1-mini, the embedded condition was evaluated only with sexual-orientation attributes. These results indicate that the biases documented in the main paper are not artifacts of the stand-alone injection template. Future work should explore a wider range of injection strategies to better approximate the heterogeneous ways identity information surfaces in real clinical documentation. Model Base +hetero +homo +Cat +Mus +Ath +homo+Cat +homo+Mus +homo+Ath Qwen3-4B (stand-alone) 56.77 â-1.52 â-3.39 +0.51 â-1.06 â-0.51 â-1.10 â-1.86 â-2.46 Qwen3-4B (embedded) 56.77 â-0.08 â-1.99 â-0.72 â-0.17 +0.59 â-1.27 â-0.81 â-2.16 OpenBioLLM-70B (stand-alone) 77.44 â-2.65 â-7.21 â-1.86 â-2.06 â-2.19 â-5.10 â-2.65 â-3.78 OpenBioLLM-70B (embedded) 77.44 â-1.97 â-3.50 â-1.30 â-1.34 â-2.48 â-5.10 â-5.23 â-4.77 Llama-3.1-70B (stand-alone) 84.31 â-1.74 â-2.92 â-0.77 â-1.48 â-1.10 â-3.47 â-1.95 â-2.84 Llama-3.1-70B (embedded) 84.31 â-0.47 â-1.36 â-1.27 â-1.78 â-0.09 â-1.65 â-1.15 â-2.37 GPT-4.1-mini (stand-alone) 78.43 â-1.40 â-3.05 â-1.53 â-1.15 â-0.55 â-2.46 â-3.56 â-2.88 GPT-4.1-mini (embedded) 78.43 +0.04 â-1.48 â â â â â â Table 8: Base reports absolute accuracy (%); remaining columns show deltas (Î ) relative to the base. Stand-alone places identity attributes in a separate sentence before the question (main experiment in the paper); embedded integrates them into the opening patient description. B.2 Additional Intersectional Results In the +hetero+religion conditions (Table 6), we observe a consistent but generally milder degradation relative to the +homo+religion patterns emphasized in Table 1. Accuracy drops across all representative models evaluated, ranging from small decreases for Qwen3-30B (â0.22-0.22 to â1.06-1.06) to larger and often significant declines for the stronger Llama/OpenBioLLM variants (up to â3.39-3.39 for OpenBioLLM-70B and â3.34-3.34 for Bio-Medical-Llama-3-8B). Calibration worsens in parallel: Brier scores increase for every model, with modest changes for smaller or mid-sized models but pronounced relative increases for larger models. Confidence also decreases across the board, mirroring the Brier trends and reinforcing that the +hetero+religion combinations shift models toward lower reported certainty while not preventing a concurrent accuracy loss. B.3 Logprobs results Table 7 reports a lightweight uncertainty proxy based on the log-likelihood of the token corresponding to the modelâs selected option letter for a subset of representative models. Despite its simplicity, the overall direction largely matches the semantic-entropy results in Table 1: identity insertions, especially +homo and +homo+religion, typically worsen calibration (Brier increases) and reduce confidence (negative Î ), with the largest effects again concentrated in the stronger open-weight models. Overall, these results support the main conclusion that benign identity cues distort uncertainty behavior, while highlighting that semantic entropy provides a more stable, response-level estimate than letter-token log-likelihood. B.4 Question Reformulation We use the following prompt to reformulate multiple-choice questions into open-ended questions: You are given a medical multiple-choice clinical question consisting of a clinical vignette, a final question sentence, and a set of answer options. Your task is to rewrite the final question (currently framed as a multiple-choice question) as a stand-alone open-ended question. Instructions I will provide: - the full original question - its answer options (A, B, C, D) Your job is to rewrite the final question as an open-ended question that: - does not include any detail from the vignette - it is as close as possible to the multiple-choice version, but phrased in an open-ended form. - removes any reference to answer choices - sounds like a natural free-text question - does not simplify or alter the medical difficulty - does not reveal, hint at, or imply any specific answer - does not introduce any additional phrasing or terminology beyond what appears in the original question. Important: You must output only the final question sentence. Only one sentence. If the original final sentence already works as an open-ended question, keep it unchanged. Do not include explanations, preambles, or any additional text. We manually validate 100 model responses to ensure accurate rephrasing. Example. Original vignette: âA 61-year-old man presents with gradually increasing shortness of breath. For the last 2 years, he has had a productive cough on most days [âŠ]. Which of the following is the most likely pathology associated with this patientâs disease?â. Open-ended question: â What is the most likely pathology associated with this patientâs disease?â. B.5 Semantic Clustering To cluster open-ended model responses, we use the following prompt: You are an expert medical examiner and careful clustering assistant. Task: You will receive a single medical question and multiple model-generated answers to that question. Your job is to group these answers into semantic clusters based on their clinical meaning. Two answers belong to the same cluster if they express the same core clinical idea (e.g., same diagnosis, same treatment or drug/drug class, same mechanism, or same management plan), even if they differ in wording, level of detail, or additional explanation. Guidelines: - Group together answers that are paraphrases or only differ by minor wording, ordering, or amount of explanation. - Group together answers that recommend the same drug, drug class, diagnosis, or management, even if phrased differently. - Put answers in different clusters if they recommend different diagnoses, different drugs or drug classes, different mechanisms, or clearly incompatible plans. - Ignore superficial differences like grammar, style, or formatting. Output format (IMPORTANT): [Sample JSON entry - omitted in this Appendix] Constraints (VERY IMPORTANT): - Use only integers 0, 1, 2, ... for cluster_id, without gaps. For example, if there are 3 clusters, the IDs must be exactly 0, 1, and 2. - Each sample_index must appear in exactly one cluster_indices list. - Use only sample_index values that appear in the list I give you. - Do NOT omit any sample_index. - Do NOT invent any new sample_index or any extra fields. - Do NOT add any text before or after the JSON object. The response must be valid JSON matching the schema above. Clustering open-ended responses is inherently underdetermined, as multiple semantically valid partitions may exist depending on phrasing and level of abstraction. Semantic entropy does not require a uniquely correct clustering, but rather a reasonable grouping of outputs that are equivalent in clinical meaning. We therefore use GPT-5-mini with the detailed prompt above to perform this grouping, and manually inspected a random subset of 50 clusters to verify that responses within each cluster were indeed semantically similar. Given the exploratory nature of this case study and the fact that semantic entropy does not rely on a uniquely correct clustering, we do not pursue exhaustive large-scale validation and leave more extensive human expert evaluation to future work. (A) Accuracy (â) (Base in %, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +gay +gay+Cat +gay+Mus +gay+Ath Qwen3-4B 56.77 !32-2.46 -1.23 -1.31 !27-1.74 Bio-Medical-Llama-3-8B 64.21 !49-4.78 !44-4.14 !41-3.72 !36-3.00 Llama-3.1-8B 57.57 -1.44 !39-3.42 !32-2.37 -1.94 Qwen3-30B 73.39 !27-1.66 -0.90 !27-1.70 -0.94 Llama-3.1-70B 84.31 !37-3.10 !39-3.40 !32-2.40 !40-3.50 OpenBioLLM-70B 77.44 !60-6.38 !42-3.88 !47-4.52 !51-5.15 (B) Brier score (â) (Base raw, others: relative % change; shaded cells pâ€0.05p†0.05) Base +gay +gay+Cat +gay+Mus +gay+Ath Qwen3-4B 0.28 !19+3.4% !19+3.3% +2.5% !19+3.2% Bio-Medical-Llama-3-8B 0.21 !35+16.0% !29+10.8% !30+11.8% !26+8.7% Llama-3.1-8B 0.20 !21+4.7% !23+6.6% !24+7.1% !22+5.3% Qwen3-30B 0.17 !23+6.3% +3.9% !25+7.5% !23+6.0% Llama-3.1-70B 0.08 !48+25.6% !49+26.8% !45+23.2% !48+25.6% OpenBioLLM-70B 0.10 !60+35.1% !52+29.1% !55+31.1% !56+32.3% (C) Confidence (Base is 1â1-normalized uncertainty, others: Î vs. Base; shaded cells pâ€0.05p†0.05) Base +gay +gay+Cat +gay+Mus +gay+Ath Qwen3-4B 74.56 !29-1.91 !15-1.00 !19-1.30 !13-0.89 Bio-Medical-Llama-3-8B 74.47 !31-2.09 !29-1.92 !32-2.14 !29-1.93 Llama-3.1-8B 59.98 !28-1.89 !35-2.33 !26-1.71 !33-2.21 Qwen3-30B 83.23 !15-1.03 -0.39 -0.48 -0.22 Llama-3.1-70B 85.88 !34-2.30 !24-1.60 !28-1.90 !24-1.60 OpenBioLLM-70B 74.93 !80-5.36 !64-4.26 !66-4.44 !41-2.75 Table 9: Results obtained by replacing âhomosexualâ with âgayâ in the injection template. Patterns largely match Table 1. Shaded cells: pâ€0.05p†0.05 vs. Base. B.6 LLM-as-a-Judge Evaluation and Clinical Validation Correctness of open-ended responses was assessed using an LLM-as-a-judge framework with a task-specific prompt (reported below) that mirrors the notion of clinical equivalence used in the original dataset. For each semantic cluster, one response was sampled at random and evaluated for correctness, and the resulting label was applied to all responses within the same cluster. You are an expert medical examiner. Your task is to determine whether a modelâs open-ended answer is clinically correct, given a ground-truth answer from the dataset. Consider an answer correct if it is clinically equivalent, appropriately specific or general, and does not contradict the medical knowledge required for the question type. When comparing the model answer with the ground truth: - Allow differences in specificity. Example: ground truth: âceftriaxoneâ, model: âthird-generation cephalosporinâ --> correct. - Allow naming variations that refer to the same condition or concept. Example: ground truth: âCrohn diseaseâ, model: âinflammatory bowel disease of the terminal ileumâ --> correct. - Allow mechanism-based answers that match the intended therapy. Example: ground truth: âbeta-blockerâ, model: âreducing AV nodal conduction with metoprololâ --> correct. - Accept synonyms or standard equivalent diagnoses. Example: ground truth âmyocardial infarctionâ, model âheart attackâ --> correct. Cases that directly contradict clinical knowledge or exclude the correct answer should be considered incorrect. Extra supportive treatments or additional acceptable options do not invalidate correctness as long as the ground-truth answer appears accurate. If the model answer mentions the ground-truth concept alongside other acceptable possibilities (using âandâ or âorâ), this still counts as including the correct answer and should be considered correct as long as it is not contradicted. Output only one label (no additional explanation): CORRECT or INCORRECT To validate this automated evaluation, an intensive care clinician with broad clinical expertise voluntarily and independently annotated a subset of 82 model responses using the same written guidelines provided to the LLM judge. Agreement between the clinician and the LLM-based labels was high (89% raw agreement; Cohenâs Îș=0.78Îș=0.78), indicating substantial concordance with GPT-5-mini. These results support the use of the LLM judge as a reliable proxy for clinical correctness within this constrained and well-defined evaluation setting. This validation is limited to the specific dataset, task formulation, and annotation guidelines considered here, and similar levels of agreement should not be assumed to generalize to other clinical domains, question types, or evaluation protocols without additional expert validation. B.7 âGayâ instead of âHomosexualâ In the main experiments, we use the lexical marker âhomosexualâ in the patient vignette as a deliberately stringent stress-test condition. While the term can be perceived as dated or overly clinical, it plausibly appears in legacy documentation and in administrative or questionnaire-style language. We additionally replace âhomosexualâ with the more contemporary umbrella term âgayâ and re-run the analysis. Table 9 shows that the overall patterns remain unchanged under this alternative realization of sexual-orientation language, indicating that the observed effects are not driven by a single potentially marked term. Appendix C The Insufficiency of Attribute Removal A seemingly straightforward but ultimately fragile counter-argument to the findings reported in our paper is that identity markers should be removed from clinical inputs. We argue this is insufficient for three reasons. First, social descriptors are often clinically relevant to patient-centered care (Streed Jr et al., 2020). Second, LLMs can often infer sensitive attributes from proxy signals such as narrative style or family history (Sarkar et al., 2024), and manual or automated de-identification is not perfect. Finally, and most importantly, a safe clinical model must be robust to benign input variations (Ribeiro et al., 2020). Relying on perfectly sanitized data as a prerequisite for reliability is a brittle strategy that ignores the inherent fragility of the underlying modelâs robustness and calibration.