Paper deep dive
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/23/2026, 3:08:45 AM
Summary
This study presents a randomized algorithm audit of seven large language models (LLMs) to determine how reputation signals and demographic cues influence physician recommendations. Using a choice-based conjoint design with 40,068 responses, the authors found that reputation signals (rating, fee) dominate decision-making, while demographic signals (gender, ethnicity) create small but statistically significant biases favoring female and minority-signaled names, contrary to human discrimination patterns. Crucially, these biases are invisible in the models' self-reported explanations, highlighting a gap between revealed behavior and stated rationale.
Entities (10)
Relation Signals (7)
Visit Fee → negativelyimpacts → Choice Probability
confidence 98% · raising the fee from $90 to $190 lowers it by 20.0 pp
Patient Rating → positivelyimpacts → Choice Probability
confidence 98% · raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points
DeepSeek-R1 → failed → Auditability Gate
confidence 95% · One reasoning model failed the prespecified auditability gate outright.
LLM → performs → Algorithm Audit
confidence 95% · We report a prespecified randomized algorithm audit of what causally moves those recommendations.
LLM Explanations → failtodetect → Demographic Bias
confidence 92% · models mentioned gender or ethnicity in at most 0.03% of their stated reasons... these effects are invisible in the models' own explanations
Female-Signaled Names → positivelyimpacts → Choice Probability
confidence 90% · female-signaled names gain 2.5 pp
Minority-Signaled Names → positivelyimpacts → Choice Probability
confidence 90% · Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.
Tags
Links
- Source: https://arxiv.org/abs/2608.14399v1
- Canonical: https://arxiv.org/abs/2608.14399v1
Trouble viewing inline? Open PDF directly →
Full Text
56,071 characters extracted from source content.
Expand or collapse full text
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model–assisted physician choice Syeda Anshrah Gillani 1 Mirza Samad Ahmed Baig 2 1 Heidelberg University, Heidelberg, Germany 2 Fandaqah, Al Khobar, Saudi Arabia Corresponding author: Syeda Anshrah Gillani syeda.gilani@stud.uni-heidelberg.de Abstract Background: Patients increasingly ask large language model (LLM) assistants which doctor to see. These systems have become AI infomediaries: algorithms that interme- diate a person’s choice among other people and thereby decide, silently and at scale, which physicians become visible. What causally moves their recommendations, and whether physician names signaling gender or ethnicity play a role, is undocumented, because existing audits observe correlated real-world profiles from which causal weights cannot be identified. Objective: To estimate the causal effect of reputation signals (patient rating, review volume and recency, feedback response, hospital affiliation, visit fee, telehealth availabil- ity, years in practice), name-signaled demographics, and list position on the probability that an LLM recommends a physician. Methods: Prespecified algorithm audit using a randomized choice-based conjoint. Across three patient personas and nine prompt paraphrases, each audited model chose among five synthetic family-medicine physician cards whose attributes were independently randomized (3,024 choice sets; nine experimental arms probing decoding temperature, response format, and presentation order). Physician gender and ethnicity were signaled through names following correspondence-audit methodology. Average marginal component effects (AMCEs) were estimated by linear probability models with choice-set–clustered standard errors; effects are also expressed as visit-fee equivalents. 1 arXiv:2608.14399v1 [cs.CY] 14 Aug 2026 Whose doctor does the AI recommend?2 Results: Across seven models (six open-weight; gpt-4o-mini) and 40,068 scored re- sponses, reputation signals dominate: moving a rating from 3.9 to 4.7 raises choice probability by 31.4 p (95% CI 30.5–32.3) and a$90→$190 fee lowers it by 20.0 p (CI 19.1–21.0). Demographic parity is rejected in an unexpected direction: female-signaled names gain 2.5 p (CI 1.8–3.2) and Hispanic-, South-Asian–, and Black-signaled names gain 1.3–2.9 p over White-signaled names, tilts worth$7–$14 per visit in fee-equivalent terms; being listed first is worth$11. Yet models mentioned gender or ethnicity in ≤0.03% of their stated reasons and abstained in 0.39% of trials: the demographic and position effects were invisible in the models’ own explanations. A reasoning model (deepseek-r1:7b) failed the prespecified auditability gate outright. Conclusions: AI infomediaries neither reproduce the anti-minority discrimination documented in human audit studies nor achieve neutrality: they silently apply systematic demographic tilts of their own. The divergence between revealed weights and stated reasons means transparency obligations that rely on model self-report would not surface these effects; only behavioral audits can. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli with a single command, converting this study from a one-time snapshot into a recurring monitoring instrument for AI-mediated access to care. Keywords: large language models; algorithm audit; algorithmic fairness; health equity; physician ratings; conjoint analysis; patient choice; algorithmic gatekeeping; AI transparency; intersectional bias; digital determinants of health; correspondence audit 1 Introduction Choosing a doctor has always run on reputation. Patients weigh word of mouth, online ratings, credentials, and cost, and a large literature documents how these signals shape provider choice (Hanauer et al., 2014; Emmert et al., 2013; Yaraghi et al., 2018). A new intermediary has now inserted itself into that decision: conversational assistants built on large language models (LLMs). Instead of scanning a directory of physician profiles, a patient can simply ask an assistant, “Which of these doctors should I book?” A growing share do, as LLM assistants absorb search, comparison, and triage tasks that previously belonged to rating platforms and search engines (Ayers et al., 2023). When an assistant answers, it collapses a multi-attribute comparison into a single recommendation. The weights it applies, whether to a star rating, a consultation fee, a hospital affiliation, or, more troublingly, to a physician’s name, take effect silently, at scale, and without any of the professional accountability that governs human referral. The assistant has become an AI infomediary : an algorithm that Whose doctor does the AI recommend?3 intermediates one person’s choice among other people, and in doing so allocates visibility, and ultimately patients, across physicians. This paper asks a simple causal question: what moves an LLM’s physician recommen- dation? We answer it with a prespecified algorithm audit (Sandvig et al., 2014; Metaxa et al., 2021) built on a randomized choice-based conjoint design (Hainmueller et al., 2014). Audited models repeatedly choose among five synthetic family-medicine physician cards whose attributes are independently randomized: patient rating, review volume, review recency, whether the practice responds to feedback, hospital-system affiliation, new-patient visit fee, telehealth availability, years in practice, and gender and ethnicity, which are signaled through physician names using correspondence-audit methodology (Bertrand and Mullainathan, 2004; Gaddis, 2017). Because every attribute varies independently of every other and of display order, differences in choice probability identify the average marginal component effect (AMCE) of each signal, and the fee attribute converts any effect into an interpretable dollars-per-visit equivalent. The design speaks to three literatures at once. First, it extends the economics of online physician reputation (Gao et al., 2012; Yaraghi et al., 2018) from human choosers to algorithmic ones: prior conjoint evidence shows how patients trade off ratings and other attributes; we provide the matching causal estimates for the AI assistants now advising them. Second, it extends the algorithmic-fairness literature in health care (Obermeyer et al., 2019; Zack et al., 2024; Omiye et al., 2023), which has concentrated on clinical decisions (triage, diagnosis, treatment), to the consumer-facing question of which humans the algorithm makes visible. Existing audits of chatbot specialist recommendations are descriptive, asking models to name real physicians and characterizing who appears (Parikh et al., 2024); because real physicians’ attributes are correlated, such designs cannot separate a rating effect from a name effect. Our synthetic-profile randomization can. Third, it contributes to the emerging study of AI infomediaries and generative engine optimization (Aggarwal et al., 2024), where a companion audit in the travel domain (Baig et al., 2026) found that assistants reproduce human valence-price primacy while introducing content-free position effects; identical rating and recency levels here permit the first cross-domain comparison of LLM signal weights. Concretely, we make four contributions. (1) We provide, to our knowledge, the first randomized conjoint audit of LLM physician recommendation, across a panel of widely deployed models, three patient personas, nine prompt paraphrases, and nine prespecified experimental arms. (2) We estimate fee-equivalent values for every signal, including the dollar value of a first-listed position and of name-signaled demographic attributes, with uncertainty quantified by cluster-robust inference and Krinsky–Robb simulation. (3) We test demographic parity directly, at the group level and at the intersection of gender and ethnicity, Whose doctor does the AI recommend?4 using prespecified equivalence bounds, so that a null is informative rather than merely underpowered, and we characterize when models abstain from choosing or spontaneously flag demographic attributes. (4) We compare the models’ stated reasons against their revealed weights, quantifying how faithfully AI explanations track AI behavior in a health context. Beyond these estimates, the frozen-design instrument makes every finding falsifiable and repeatable: auditing any future model against identical stimuli is a single command, so the parity verdicts reported here can be re-issued, or revoked, with each model release. Across seven models and 40,068 scored responses, three findings emerge. First, LLM recommenders run on reputation: patient rating alone carries 37.7% of total attribute importance and fee 24.1%, with a 3.9→4.7 rating step worth$157 per visit in fee-equivalent terms. Second, demographic parity fails, but not in the direction human audit studies would predict: female-signaled and minority-signaled names gain one to three percentage points, tilts worth$7–$14 per visit, and no gender×ethnicity cell meets the prespecified equivalence bound. Third, none of this is visible in the models’ own explanations: gender and ethnicity appear in fewer than 0.03% of stated reasons, and models abstained from choosing in only 0.39% of trials, so the demographic and position effects operate entirely below the models’ self-reported rationale. The remainder of the paper reviews the relevant literatures and derives hypotheses (Section 2), details the design, model panel, and estimation strategy (Section 3), reports results (Section 4), and discusses implications for platforms, regulators, and health-services research (Section 5). 2 Background and hypotheses 2.1 Online reputation and physician choice Physician-rating websites have made service reputation legible in health care. Public awareness and use of these ratings grew rapidly through the 2010s (Hanauer et al., 2014), physician coverage on rating platforms expanded in parallel (Gao et al., 2012), and systematic reviews document that valence and volume of reviews influence patients’ stated willingness to select a provider (Emmert et al., 2013). The causal benchmark most relevant to our design is Yaraghi et al. (2018), a choice-based conjoint in which patients chose between physician profiles with randomized quality ratings: higher ratings raised choice odds substantially, and nonclinical (service) ratings mattered alongside clinical ones. Beyond ratings, patients weigh access attributes (wait time, fees, telehealth) and credential attributes (experience, affiliation) in ways that vary with their situation: an uninsured patient weighs fees differently from a Whose doctor does the AI recommend?5 chronically ill one. Our persona manipulation mirrors that heterogeneity. If LLM assistants have internalized the regularities of human choice documented in this literature, as their training corpora suggest they should, their revealed weights should reproduce the canonical ordering: valence first, then price, with secondary signals (volume, recency, responsiveness, affiliation, experience, telehealth) positive but smaller. Hypotheses H1–H8 formalize this expectation attribute by attribute (Table 4 lists the full prespecified family): higher patient ratings (H1), more reviews (H2), more recent reviews (H3), practices that respond to feedback (H4), university-hospital affiliation (H5), lower visit fees (H6), telehealth availability (H7), and more years in practice (H8) each increase the probability of recommendation. 2.2 Names as treatments: discrimination audits meet AI advice Correspondence audits established that names alone shift decisions: r ́esum ́es signed “Emily” or “Greg” received 50% more callbacks than identical r ́esum ́es signed “Lakisha” or “Jamal” (Bertrand and Mullainathan, 2004), and the methodology for selecting demographically distinctive names is now standard (Gaddis, 2017). Field experiments extended the finding to platform markets (Edelman et al., 2017). In health care specifically, algorithmic systems have repeatedly encoded group disparities: a widely used care-management algorithm systematically under-referred Black patients (Obermeyer et al., 2019), and LLMs have been shown to propagate race-based medicine (Omiye et al., 2023) and to generate clinical vignettes and treatment recommendations that vary with patient demographics (Zack et al., 2024). That literature concerns the algorithm’s treatment of patients. The question here is different and unexamined: when the algorithm chooses among physicians, does a name signaling gender or ethnicity move the recommendation, holding every reputational attribute fixed? Three outcomes are possible, and all are informative. The assistant may reproduce human discrimination absorbed from training data; it may be neutral, because alignment training suppresses demographic reasoning; or it may overcorrect. We therefore prespecify parity nulls with equivalence bounds rather than directional predictions: name-signaled gender (H9) and ethnicity (H10) have no effect on recommendation probability, tested jointly across the panel and bounded by two one-sided tests at±1.5 percentage points. Because demographic effects in health care are often intersectional, with clinical-bias audits reporting disparities for specific gender×ethnicity combinations that neither main effect anticipates (Zack et al., 2024), we additionally test whether the two signals combine additively (H13) and estimate all ten gender×ethnicity cell effects. Finally, an abstention is not necessarily a defect: when only names distinguish otherwise-comparable physicians, a model that declines Whose doctor does the AI recommend?6 to choose behaves more defensibly than one that silently lets names decide. We therefore treat abstention and whether responses explicitly acknowledge demographic attributes as outcomes in their own right. 2.3 Position, personas, and explanation faithfulness Two further properties of LLM choice behavior carry over from algorithm audits outside health. First, display position: retrieval and recommendation interfaces reward rank, and our companion travel-domain audit (Baig et al., 2026) found a causal first-listed advantage worth real money per booking. Because position is content-free, any position effect in physician recommendation is a pure artifact of the medium, consequential for provider- directory platforms whose orderings feed AI assistants. Second, context sensitivity: H11 predicts that the fee penalty is larger for a persona paying out of pocket, and H12 that review volume moderates the rating effect (a 4.7 average over 400 reviews is more diagnostic than over 12 (Yaraghi et al., 2018)). Finally, models explain their picks on request, but stated reasons need not track revealed weights; in the travel audit they tracked imperfectly (Baig et al., 2026). We quantify that gap here, with particular attention to whether demographic attributes that move choices are ever mentioned in explanations, the failure mode with the clearest governance implications. 3 Methods 3.1 Design overview We conducted a randomized choice-based conjoint embedded in an algorithm audit. Each trial presented one audited model with a persona, a prompt template, and five synthetic physician cards, and asked it to recommend one. Cards described board-certified family-medicine physicians accepting new patients with a practice near the user (all held constant), while eight reputational attributes varied independently and uniformly at random per card: patient rating (3.9, 4.3, or 4.7 of 5), number of reviews (12, 85, or 400), most recent review (3 days or 11 months ago), whether the practice responds to patient feedback (present or absent), affiliation (university hospital system vs. independent practice), new-patient visit fee ($90, $140, or$190), telehealth availability (present or absent), and years in practice (8, 18, or 28). Rating and recency levels replicate our travel-domain audit (Baig et al., 2026) to permit cross-domain comparison. Display order (slot 1–5) was randomized independently of content, identifying list-position effects. Figure 1 summarizes the pipeline from frozen design to inference; Figure 2 shows one card verbatim. Whose doctor does the AI recommend?7 1. Design seed + SHA-256 hash 3,024 choice sets 2. Stimuli 9 prompt paraphrases 3 patient personas 3. Execution 7 audited models 40,068 responses 4. Inference AMCEs, clustered SEs conditional logit 10 randomized attributes per card, including name-signaled gender×ethnicity 9 prespecified arms (temp-0, grounded, ranking, field order, retest, . . . ) Resume-safe JSONL: choice + stated reason per call H1–H13 (Holm) TOST equivalence (±1.5 p) fee-equivalents ($/visit) Figure 1: Audit pipeline. A design frozen (and hashed) before any data collection generates 3,024 five-card choice sets whose attributes vary independently; each audited model answers every choice set under nine prespecified arms; average marginal component effects and prespecified equivalence tests are estimated from the pooled choices. Because randomization is at the attribute level, any systematic difference in choice probability identifies the causal weight the model places on that signal. Dr. Priya Sharma - Family medicine physician Board-certified, accepting new patients Practice located within 20 minutes of you Patient rating: 4.7/5 (85 reviews) Most recent review: 3 days ago In practice for 18 years New-patient visit fee:$140 Affiliated with a major university hospital system Practice responds to patient feedback Telehealth visits available Randomized per card name (gender×ethnicity signal), patient rating, review volume, review recency, response to feedback, hospital affiliation, visit fee, telehealth, years in practice Held constant specialty, board certification, distance, accepting new patients Figure 2: One physician card as presented to the models (card layout shown; the grounded- arm snippet and reversed-layout variants are described in Methods). Black lines vary indepen- dently between cards; gray lines are identical on every card, so specialty, board certification, and distance cannot drive choices. Each choice set presents five such cards, and the model is asked to choose one for the persona and give a brief reason. Physician gender (female/male) and ethnicity (White, Black, Hispanic, East Asian, South Asian) were randomized per card and signaled through names (“Dr. Priya Sharma”), with four first names and four surnames per gender×ethnicity cell drawn from the correspondence-audit tradition (Bertrand and Mullainathan, 2004; Gaddis, 2017). Analysis is at the cell level; a prespecified placebo tests that the specific name exemplar within a cell has no effect. Within a choice set, no two cards shared a surname or a full attribute profile. 3.2 Stimuli, personas, and arms Prompts crossed three patient personas (new in town; managing a chronic condition; uninsured and paying out of pocket) with nine paraphrased templates, and requested a JSON response Whose doctor does the AI recommend?8 naming the chosen physician and a one-to-two-sentence reason. The design comprises 3,024 main-arm choice sets (112 per persona×template stratum), fixed across all audited models: every model sees identical stimuli, so model comparisons are within-trial. Eight prespecified secondary arms re-present stratified subsets of the same trials under altered conditions: temperature 0 (594 trials), grounded web-snippet formatting (297), top-3 ranking (297), reason-before-choice field order (189), five-repetition test–retest (108), a letter-labeled logprob arm separating position from letter-token preference (297), a persona-free control (297), and reversed card-line order (189). A 20-repetition manipulation check asked each model to define board certification. The design matrix was generated once (seed 20260702), hashed (SHA-256), and frozen before any non-pilot model call, and the analysis plan (hypotheses, estimators, equivalence bounds, and exclusion rules) was written and fixed before any confirmatory data were analyzed. 3.3 Model panel and execution The audited panel comprises six widely deployed open-weight instruction-tuned models executed locally via Ollama (llama3.2:3b, qwen2.5:3b, phi3:mini, mistral:7b-instruct-q4KM, gemma3:4b, and llama3.1:8b) and one proprietary production model (gpt-4o-mini, accessed through an Azure OpenAI deployment, API version 2025-01-01-preview) under the identical protocol, including seed control and, on the letter arm, token-level log probabilities; exact identifiers and inference dates are recorded in the design metadata. A seventh open-weight candidate, deepseek-r1:7b, was subject to a prespecified pilot gate and excluded before full-run collection (Results). Generation used temperature 0.7 (top-p0.9) except in the temperature arm, with per-call deterministic seeds derived from (model, arm, trial, repetition). Responses failing JSON parsing or naming no listed physician were retried up to twice with a stricter reminder; remaining failures are analyzed as missingness (prespecified exclusion of model×arm cells exceeding 5% failure). All calls are checkpointed and the pipeline is resume-safe. 3.4 Exploratory frontier-model pilot and archived forecasts Because API access to frontier models was not available at collection time, we additionally piloted four Claude-family models (claude-haiku-4-5, claude-sonnet-5, claude-opus-4-8, claude- fable-5) on the first 50 main-arm choice sets of the frozen design. These responses were not collected under the audit protocol above: they were gathered interactively inside an agentic coding environment (Claude Code, Anthropic), one fresh subagent per trial with no shared context, using the pipeline’s own stimulus renderer and parser (verbatim frozen prompts; one stricter re-ask on unparseable output; 0/200 parse failures). This collection mode leaves Whose doctor does the AI recommend?9 decoding parameters at the harness defaults rather than the protocol’s temperature 0.7/top-p 0.9, embeds each response in an agent system prompt rather than a bare API call, and is not seed-reproducible. The 200 resulting records are therefore quarantined from the analysis pipeline, enter no confirmatory test, and are reported only descriptively (Table 10). We put this pilot to a second, falsifiable use. For each piloted model we fit a conditional logit to its 50 choice sets and computed out-of-sample forecast probabilities of being chosen for every card in all 3,024 main-arm choice sets. These forecasts, which are model-based extrapolations rather than observations, were archived, with a timestamp, before any full API run of these models, so that a subsequent protocol-compliant run (claude-haiku-4-5 is the designated target) can grade them: agreement would suggest small harness-contaminated pilots cheaply approximate full audits; disagreement would itself measure the harness and decoding effects described above. 3.5 Estimation The estimand is the AMCE of each attribute level on the probability of being recommended (Hainmueller et al., 2014). Primary estimator: linear probability regression of an indicator for being chosen on attribute-level dummies and slot dummies, with standard errors clustered on choice set, estimated per model and pooled with model fixed effects. Secondary: conditional (McFadden) logit per choice set. Thirteen prespecified hypotheses are tested by joint Wald tests with Holm correction; demographic parity hypotheses (H9–H10) are additionally assessed by two one-sided tests with a smallest-effect-size-of-interest of±1.5 percentage points, and are pooled estimands, because a Monte-Carlo precision analysis showed per-model power for plausible demographic effect sizes is inadequate, whereas the pooled panel yields a minimum detectable effect below 1 percentage point. Because a linear-probability interaction can be nonzero under a purely additive logit process, H12 is corroborated only if the conditional-logit interaction agrees in sign and significance. H13 tests intersectionality, that is, whether name- signaled gender and ethnicity combine additively, as a joint test of their interaction terms; we additionally report all ten gender×ethnicity cell effects relative to the White-male reference, each with an equivalence verdict. Separately and descriptively, we classify each response as an abstention (declining to choose) and as demographic-aware (referring to a physician’s name, gender, ethnicity, or race), the latter measuring how often a model’s own words reveal that a name-signaled attribute was salient. Fee-equivalents divide each AMCE by the per-dollar fee AMCE (linearized over the$100 range), with 95% intervals by Krinsky–Robb simulation (10,000 draws), gated on monotonicity of the fee response. Stated reasons are coded for attribute mentions by a validated dictionary and an independent judge model (gemma2:2b, Whose doctor does the AI recommend?10 excluded from the audited panel), and compared with revealed importance shares. 3.6 Validation and reproducibility The complete pipeline (stimulus rendering, execution, parsing, and every analysis module) was validated end-to-end against a mock data-generating process with known conditional-logit effects, including planted name-signaled effects (female +0.10; Black−0.15 on the logit scale) and true-null ethnicity effects. The validation suite (18 assertions) confirms recovery of signs, magnitudes, true nulls, the fee-monotonicity gate, the name-exemplar placebo, and figure/table generation. Every call is checkpointed, and the frozen design matrices are hash-verified at load time, so a full audit is reproducible from the design seed. All physician identities are synthetic, and the demographic manipulation audits model behavior, not real physicians. 4 Results 4.1 Data quality and manipulation checks The confirmatory panel comprises seven models: six open-weight instruction-tuned models (llama3.2:3b, qwen2.5:3b, phi3:mini, mistral:7b-instruct, gemma3:4b, llama3.1:8b) and one proprietary production model (gpt-4o-mini, accessed via Azure OpenAI), yielding 40,068 scored responses across the nine arms. Parse-failure rates were 0% for gpt-4o-mini on every arm and below 2% for nearly all model–arm cells (Table 1); the exceptions were mistral on the retest (32.0%) and top-3 ranking (14.1%) arms, which enter the missingness analysis in Section 4.7. An eighth candidate, deepseek-r1:7b, failed its prespecified pilot gate (100% parse failures: its chain-of-thought exhausted the fixed output budget before emitting the required JSON) and was excluded before any full-run data were collected. This is itself a finding about the auditability of reasoning models under fixed-format protocols. Test–retest consistency was high: the modal choice was repeated in 73% (mistral) to 99% (gemma3) of five-fold repetitions (Table 2). 4.2 What moves the recommendation: AMCEs Reputation signals dominate (Figure 3; Table 3). Pooled across models, moving a card’s rating from 3.9 to 4.7 raises its choice probability by 31.4 percentage points (p) against a 20% baseline (95% CI 30.5–32.3); raising the fee from$90 to$190 lowers it by 20.0 p (CI 19.1–21.0). Review volume (400 vs. 12 reviews: +8.4 p), review recency (+3.8 p), Whose doctor does the AI recommend?11 Table 1: Response parse-failure rates. Modelcardreversed fieldorder grounded letter main nopersona ranktop3 retest temp0 gemma3:4b0.00.00.00.00.80.00.00.00.0 gpt-4o-mini0.00.00.00.00.00.00.00.00.0 llama3.1:8b0.00.00.00.00.00.00.00.00.0 llama3.2:3b0.00.00.01.70.00.00.70.00.0 mistral:7b-instruct-q4 KM0.00.00.00.00.00.014.132.00.0 phi3:mini0.00.00.00.00.00.07.10.00.0 qwen2.5:3b0.00.00.00.00.00.00.00.00.0 Note. Parse-failure rate (%) by model and arm; pilot and manipulation-check calls excluded. Table 2: Test–retest consistency of recommendations. ModelModal share Entropy gemma3:4b0.9930.015 gpt-4o-mini0.9670.064 llama3.1:8b0.9190.155 llama3.2:3b0.8560.295 mistral:7b-instruct-q4KM0.7320.560 phi3:mini0.8220.368 qwen2.5:3b0.8650.274 Note. Mean over test–retest choice sets (5 repetitions each). telehealth availability (+2.8 p), responding to feedback (+2.1 p), and 28 vs. 8 years in practice (+2.1 p) all move choices in the expected direction. Hospital-system affiliation does not (+0.3 p,P=.35), the lone prespecified reputation hypothesis (H5) not supported. In share-of-importance terms, rating accounts for 37.7% of the total attribute range, price 24.1%, volume 10.1%, and list position 8.2% (Table 5). Confirmatory hypotheses H1–H4 and H6–H8 were all supported after Holm correction (Table 4). 4.3 Name-signaled gender and ethnicity Demographic parity is rejected, in a direction the bias literature would not predict. Physician cards bearing female-signaled names were chosen 2.5 p more often than male-signaled ones (CI 1.8–3.2; H9P Holm < .001), and ethnicity effects are jointly significant (H10P Holm < .001): relative to White-signaled names, Hispanic-signaled (+2.8 p), South-Asian–signaled (+2.9 p), and Black-signaled (+1.3 p) names were chosen more often, with East-Asian–signaled names indistinguishable from White (+0.3 p,P=.65). The conditional-logit specification corroborates the linear estimates (pooled odds ratios 1.22 for female, 1.15–1.24 for the elevated ethnicity cells). Because the prespecified TOST bounds of±1.5 p are excluded by several of these intervals, the correct summary is not “parity holds” but “models exhibit a systematic Whose doctor does the AI recommend?12 302010010203040 AMCE: change in P(recommended), percentage points South Asian name (vs White) East Asian name (vs White) Hispanic name (vs White) Black name (vs White) Female name (vs male) 28 yrs in practice (vs 8) 18 yrs in practice (vs 8) Telehealth available $190 fee (vs $90) $140 fee (vs $90) Hospital-affiliated (vs independent) Responds to patient feedback Recent review (vs 11 mo.) 400 reviews (vs 12) 85 reviews (vs 12) Rating 4.7 (vs 3.9) Rating 4.3 (vs 3.9) Causal effect of physician signals on LLM recommendation pooled gemma3:4b gpt-4o-mini llama3.1:8b llama3.2:3b mistral:7b-instruct-q4_K_M phi3:mini qwen2.5:3b Figure 3: Average marginal component effects on choice probability (pooled and per model), with 95% choice-set–clustered confidence intervals. Baseline choice probability is 20% (five cards per set). pro-female, pro-minority tilt of one to three percentage points.” The prespecified name- exemplar placebo, however, is violated: choice shares differ across name exemplars within gender×ethnicity cells (LR = 370.0, df = 30,P < .001), so cell-level effects should be read as perceived-category effects entangled with name-specific familiarity, and the demographic estimates as audit-level signals rather than precise category parameters. 4.3.1 Intersectional effects The gender×ethnicity interaction is not jointly significant after correction (H13,P Holm = .099), indicating approximately additive effects, but the additive combination concentrates advantage: relative to White-male cards, female-Hispanic cards gain 5.7 p (CI 4.0–7.3), female–South-Asian 5.1 p, and female-Black 3.9 p (Table 6). No cell meets the prespecified equivalence criterion, so no gender×ethnicity combination can be certified as neutrally treated. Whose doctor does the AI recommend?13 Table 3: Average marginal component effects on P(recommended), pooled across the model panel (percentage points). AttributeAMCE (p)95% CIP Rating 4.3 (vs 3.9)8.06[7.34, 8.78] ¡.001 Rating 4.7 (vs 3.9)31.38[30.48, 32.29] ¡.001 85 reviews (vs 12)4.79[3.93, 5.64] ¡.001 400 reviews (vs 12)8.44[7.56, 9.31] ¡.001 Recent review (vs 11 mo.)3.80[3.09, 4.51] ¡.001 Responds to patient feedback2.09[1.37, 2.80] ¡.001 Hospital-affiliated (vs independent)0.34[-0.37, 1.04].349 \$140 fee (vs \$90)-14.46 [-15.45, -13.46] ¡.001 \$190 fee (vs \$90)-20.02 [-20.96, -19.08] ¡.001 Telehealth available2.78[2.08, 3.49] ¡.001 18 yrs in practice (vs 8)0.85[-0.01, 1.72].052 28 yrs in practice (vs 8)2.08[1.21, 2.96] ¡.001 Female name (vs male)2.51[1.81, 3.20] ¡.001 Black name (vs White)1.33[0.25, 2.42].016 Hispanic name (vs White)2.81[1.69, 3.93] ¡.001 East Asian name (vs White)0.25[-0.84, 1.34].653 South Asian name (vs White)2.88[1.74, 4.01] ¡.001 Position 2 (vs 1)0.68[-0.54, 1.89].274 Position 3 (vs 1)-0.23[-1.42, 0.97].711 Position 4 (vs 1)-3.31[-4.48, -2.14] ¡.001 Position 5 (vs 1)-6.15[-7.26, -5.04] ¡.001 4.3.2 Abstention and demographic-aware responses Abstention is rare and indiscriminate: models declined to choose in 0.39% of responses overall, and spontaneously flagged demographic attributes in only 0.01% (Table 7). The models that let names move their choices essentially never said so. 4.4 Position effects Position effects are content-free and economically meaningful: cards in slot 5 lose 6.2 p (CI 5.0–7.3) and slot 4 loses 3.3 p relative to slot 1, holding all content constant (Figure 5). The letter-labeled arm decomposes label from position using gpt-4o-mini token logprobs: both components are small in isolation (partialR 2 =.005 for slot given letter,.008 for letter given slot), with a residual anti-“E” letter preference (−7.3 p, P = .03) (Table 8). Whose doctor does the AI recommend?14 Table 4: Confirmatory hypothesis tests (joint Wald tests, Holm-corrected). HypothesisWald df P (raw) P (Holm) H1rating4636.252¡.001¡.001*** H2volume363.612¡.001¡.001*** H3recency111.311¡.001¡.001*** H4 response32.421¡.001¡.001*** H5network0.881.349.349 H6price1745.922¡.001¡.001*** H7telehealth59.801¡.001¡.001*** H8 experience22.082¡.001¡.001*** H9gender49.901¡.001¡.001*** H10ethnicity44.674¡.001¡.001*** H11 personaxprice834.636¡.001¡.001*** H12ratingxvolume170.904¡.001¡.001*** H13intersectional9.514.050.099 Note. Holm-corrected over the H1–H12 confirmatory family. ∗ p < .05, ∗ p < .01, ∗ p < .001. Table 5: Attribute importance: range of AMCEs and normalized share. AttributeRange (p) Share Patient rating31.380.377 Visit fee20.020.241 Review volume8.440.101 List position6.830.082 Recency3.800.046 Ethnicity (name-signaled)2.880.035 Telehealth2.780.033 Gender (name-signaled)2.510.030 Feedback response2.090.025 Experience2.080.025 Affiliation0.340.004 4.5 Fee-equivalents Converting effects through the fee gradient (monotonic, as required by the prespecified gate) prices each signal in dollars per visit (Table 9). The 3.9→4.7 rating step is worth$157 per visit (CI 147–167); recency$19; telehealth$14; responding to feedback$10. The demographic tilts are worth real money: a female-signaled name is equivalent to a$12.50 fee discount (CI 9–16), Hispanic- and South-Asian–signaled names to$14, and being listed first to$11 per visit, a pure interface artifact priced like a clinical credential. Whose doctor does the AI recommend?15 0246 AMCE vs White male (p) Black male Hispanic male East Asian male South Asian male White female Black female Hispanic female East Asian female South Asian female Intersectional gender × ethnicity effects Male Female Figure 4: Choice-probability advantage of each gender×ethnicity cell relative to White-male cards, with 95% confidence intervals and the prespecified ±1.5 p equivalence band. 4.6 Persona and interaction effects Fee sensitivity is persona-dependent (H11,P Holm < .001): the uninsured persona’s price penalty far exceeds the insured personas’ (Figure 6). Rating and volume interact (H12, P Holm < .001), with high volume amplifying the high-rating premium (Figure 7); the sign replicates in the conditional logit, though probability-scale interactions under a logit-additive process must be interpreted cautiously (Methods). 4.7 Stated versus revealed importance Stated reasons track reputation but conceal demographics and position (Figure 8; 21,143 coded reasons). Rating is mentioned in 57–91% of reasons across models and price in 42–56%, roughly commensurate with their revealed weights. Gender is mentioned in≤0.03% of reasons and ethnicity in≤0.03%, against revealed importance shares of 3.0% and 3.5%, and position in≤4% against a revealed share of 8.2%. An explanation mandate that relied on Whose doctor does the AI recommend?16 Table 6: Intersectional gender × ethnicity effects on P(recommended). Gender × ethnicity cell AMCE (p)95% CIPTOST P White male (ref.)0.00– Black male0.52 [-1.02, 2.05].510.104 Hispanic male1.74[0.19, 3.30].028.620 East Asian male0.77 [-0.79, 2.32].333.177 South Asian male2.41[0.81, 4.01].003.868 White female1.81[0.26, 3.35].022.651 Black female3.93[2.39, 5.46] ¡.001.999 Hispanic female5.65[4.04, 7.27] ¡.0011.000 East Asian female1.46 [-0.06, 2.99].060.482 South Asian female5.12[3.50, 6.74] ¡.0011.000 Note. Effect of each name-signaled cell vs. White male, from a single 10-level categorical (choice-set–clustered SEs). TOST p < .05 indicates statistical equivalence within ±1.5 p. Table 7: Abstention and demographic-aware response rates. Modeln Abstain (%) Demographic-aware (%) gemma3:4b30240.00.0 gpt-4o-mini30240.00.0 llama3.1:8b30240.10.0 llama3.2:3b30240.00.0 mistral:7b-instruct-q4 KM 30240.60.0 phi3:mini30241.20.1 qwen2.5:3b30240.20.0 Note. Main arm. Abstain = declines to choose or calls the options equal; demographic-aware = the response text refers to a physician’s name, gender, ethnicity, or race. model self-report would have detected none of the demographic or position effects measured here. 4.8 Robustness The reputation hierarchy is stable across every prespecified perturbation (Figure 9): grounded web-snippet formatting, reversed card layout, field order, and the plausibility-filtered sub- sample each shift headline AMCEs by at most a few points without sign changes, and temperature-0 decoding reproduces the ordering. Removing the persona raises rating weight (+8.4 p on the 4.7 step) and halves the fee penalty (−9.7 p), consistent with personas carrying real budget information. The missingness model finds parse failures unrelated to card content. The gate-excluded deepseek-r1 and the mistral retest arm are the only material Whose doctor does the AI recommend?17 64202 AMCE vs first-listed (p) Position 2 (vs 1) Position 3 (vs 1) Position 4 (vs 1) Position 5 (vs 1) Position bias in LLM physician recommendation Figure 5: List-position effects on choice probability by model, holding card content constant. Slot 1 is the reference. Table 8: Letter-arm decomposition of list-position effects from letter-token preferences (linear probability model on first-token letter probabilities, trial-clustered SEs). TermCoef. (p) SE (p)P Position 2 (vs 1)4.423.50 .206 Position 3 (vs 1)7.953.59 .027 Position 4 (vs 1)3.993.44 .246 Position 5 (vs 1)1.243.30 .707 Letter B (vs A)0.823.68 .824 Letter C (vs A)2.903.75 .438 Letter D (vs A)-0.733.61 .839 Letter E (vs A)-7.283.26 .025 Note. Position block incrementalR 2 = 0.0050; letter-token block = 0.0077. Slot and letter are independently randomized by design. data-quality caveats. 4.9 Exploratory: frontier-model session pilot To extend the audit’s reach beyond open-weight models, four frontier Claude-family models were piloted on the first 50 choice sets of the same frozen design, and quantitative forecasts of their full-run behavior were archived before any full-scale measurement (Methods). All four Whose doctor does the AI recommend?18 Table 9: Fee-equivalent trade-offs ($/visit) with Krinsky–Robb 95% CIs. AttributePooledNew in townChronic careBudget Rating 4.3 (vs 3.9)40.3 [36.1, 44.6]85.2 [72.5, 100.4]93.0 [78.8, 110.0]14.2 [10.3, 18.1] Rating 4.7 (vs 3.9)156.8 [147.4, 167.2] 366.4 [321.9, 425.1] 343.5 [301.3, 396.6] 46.5 [42.1, 51.1] 85 reviews (vs 12)23.9 [19.4, 28.5]66.6 [52.4, 83.1]57.3 [43.3, 73.3]4.3 [0.4, 8.2] 400 reviews (vs 12)42.2 [37.2, 47.3]106.5 [89.4, 127.1]105.9 [89.0, 126.2]7.6 [3.6, 11.7] Recent review (vs 11 mo.)19.0 [15.4, 22.7]42.6 [31.3, 55.6]41.8 [30.6, 54.2]6.8 [3.5, 10.0] Responds to patient feedback10.4 [6.9, 14.1]13.7 [3.0, 25.1]22.3 [11.7, 33.0]5.7 [2.3, 9.1] Hospital-affiliated (vs independent)1.7 [-1.8, 5.2]10.4 [0.0, 21.7]19.1 [8.5, 30.4]-5.8 [-8.9, -2.7] Telehealth available13.9 [10.3, 17.5]25.6 [14.9, 37.0]28.8 [18.8, 40.1]6.1 [2.9, 9.4] 18 yrs in practice (vs 8)4.3 [-0.1, 8.5]13.3 [0.4, 26.6]7.5 [-4.9, 20.3]2.8 [-1.0, 6.7] 28 yrs in practice (vs 8)10.4 [6.1, 14.9]21.0 [8.0, 34.9]31.3 [18.4, 45.0]3.1 [-0.9, 7.0] Female name (vs male)12.5 [9.0, 16.0]31.4 [20.9, 43.3]28.0 [18.0, 39.2]2.4 [-0.8, 5.6] Black name (vs White)6.7 [1.3, 12.2]5.0 [-11.0, 21.3]18.7 [3.0, 35.4]5.4 [0.5, 10.4] Hispanic name (vs White)14.1 [8.4, 19.8]28.6 [12.3, 46.0]17.4 [1.0, 34.7]8.9 [3.5, 14.1] East Asian name (vs White)1.2 [-4.3, 6.8]-0.2 [-15.7, 15.2]2.8 [-12.5, 18.6]1.0 [-4.3, 6.1] South Asian name (vs White)14.3 [8.7, 20.3]23.3 [6.6, 40.6]23.2 [6.7, 40.9]8.3 [3.0, 13.6] First-listed (vs avg 2-5)11.3 [6.5, 16.0]39.3 [24.9, 55.0]19.9 [7.0, 33.3]0.0 [-3.7, 3.7] $190 fee (vs $90)Telehealth available28 yrs in practice (vs 8) 40 30 20 10 0 AMCE (p) Persona-contingent signal weights budget chronic new_in_town Figure 6: Fee AMCEs by patient persona. The uninsured out-of-pocket persona shows the steepest price penalty. models completed the pilot with no parse failures (200/200 calls; collection caveats in Methods; numbers fromresults/tables/pilotfrontier.csv). Descriptively, the pilot reproduces the reputation-signal ordering: cards rated 4.7 were chosen 39–43% of the time against a 20% null while cards rated 3.9 were chosen 0–1%, the$90 fee roughly quintupled choice share relative to$190 (35–37% vs. 5–8%), and higher review volume and recency moved choice in the expected directions in all four models (Table 10). Two demographic patterns appear Whose doctor does the AI recommend?19 3.94.04.14.24.34.44.54.64.7 Displayed average rating 10 20 30 40 P(recommended), % Signal diagnosticity: rating effect by review volume Review volume 12 reviews 85 reviews 400 reviews Figure 7: Rating effects conditional on review volume: high volume amplifies the high-rating premium. consistently across the piloted models and merit confirmatory scrutiny in a protocol-compliant run: physicians with female-signaled names were chosen more often than male-signaled ones (22–24% vs. 16–19%), directionally matching the confirmatory panel’s pro-female tilt, and East-Asian–signaled names were chosen more often (31–36%) than White-signaled (12–19%) or Black-signaled (9–13%) names, a pattern the confirmatory panel does not show. Atn= 50 choice sets per model these are directional readings only; none is a hypothesis test. The archived pilot-based forecasts of full-run choice probabilities (Methods) could not be graded at analysis freeze because no protocol-compliant API run of a forecast-target model was available; they remain archived, timestamped, and ungraded, available for scoring the day such a run is executed. Whose doctor does the AI recommend?20 0.00.20.40.60.81.0 Revealed importance (AMCE share) 0.0 0.2 0.4 0.6 0.8 1.0 Stated importance (mention share) Stated vs revealed attribute importance gemma3:4b gpt-4o-mini llama3.1:8b llama3.2:3b mistral:7b-instruct-q4_K_M phi3:mini qwen2.5:3b stated = revealed Figure 8: Stated versus revealed importance by attribute and model: share of stated reasons mentioning each attribute (stated) against its share of total AMCE range (revealed). Demographic attributes and position sit far below the diagonal. Whose doctor does the AI recommend?21 1.00.50.00.5 AMCE difference (p) Rating 4.7 (vs 3.9) 400 reviews (vs 12) $190 fee (vs $90) 28 yrs in practice (vs 8) Recent review (vs 11 mo.) Responds to patient feedback Hospital-affiliated (vs independent) Telehealth available Female name (vs male) Black name (vs White) Temperature 21012 AMCE difference (p) Snippet format 1.51.00.50.00.51.0 AMCE difference (p) Field order Robustness of signal weights Figure 9: Headline AMCEs across robustness arms (grounded snippets, reversed layout, field order, temperature 0, persona-free, plausibility-filtered): signs and ordering are stable throughout. Table 10: Exploratory session pilot of four Claude-family models: share of cards chosen by attribute level (null = 0.20). Each model saw the first 50 main-arm choice sets of the frozen design (250 cards; 0 parse failures). Responses were collected interactively inside an agentic coding harness with default decoding parameters, not via the Messages API under the audit protocol; estimates are descriptive and exploratory, are quarantined from the confirmatory pipeline, and enter no confirmatory test. Haiku 4.5 Sonnet 5 Opus 4.8 Fable 5 Rating 3.90.010.000.010.01 Rating 4.30.140.180.180.16 Rating 4.70.430.400.390.40 Fee$900.360.370.350.35 Fee$1400.170.170.170.17 Fee$1900.060.050.080.08 12 reviews0.110.090.070.05 85 reviews0.180.200.180.23 400 reviews0.300.290.340.30 Recent review (3 days)0.270.260.260.27 Stale review (11 mo.)0.140.150.150.14 Female name0.220.220.230.24 Male name0.190.190.170.16 White name0.190.120.120.12 Black name0.090.110.130.13 Hispanic name0.160.140.160.16 East Asian name0.310.360.310.33 South Asian name0.240.250.250.24 Whose doctor does the AI recommend?22 5 Discussion 5.1 Principal findings Three results anchor the audit. First, LLM physician recommendation is reputation-driven to a degree exceeding human benchmarks: rating carries 37.7% of attribute importance and fee 24.1%, and the models weight the rating step at$157 per visit, signal weights directionally consistent with, but steeper than, those estimated from human physician-choice conjoints (Yaraghi et al., 2018). Second, demographic parity is rejected in the direction opposite to the discrimination documented in human audit studies: female-, Hispanic-, South-Asian–, and Black-signaled names gain one to three percentage points over male White-signaled names (worth$7–$14 per visit), and no gender×ethnicity cell meets the prespecified±1.5 p equivalence bound. Third, a content-free position effect persists (first-listed worth$11 per visit), and neither the demographic nor the position effects ever surface in the models’ stated reasons. 5.2 Comparison with prior work Where human patients discount female and minority physicians in correspondence and field settings (Bertrand and Mullainathan, 2004; Gaddis, 2017), the audited models tilt modestly the other way, a pattern more consistent with post-training fairness interventions overshooting neutrality than with learned representativeness. Relative to clinical-bias audits (Zack et al., 2024; Omiye et al., 2023), which find LLMs importing human clinical disparities, the consumer-facing recommendation layer studied here behaves differently: the risk is not replicated prejudice but unaudited, opaque tilts of either sign. Relative to descriptive chatbot- referral audits (Parikh et al., 2024), the randomized design shows what they cannot: that name effects exist net of every other attribute. And relative to the travel-domain companion audit (Baig et al., 2026), the same valence-price primacy and content-free position bias reappear with identical anchor levels, establishing cross-domain stability of the infomediary signal hierarchy. 5.3 An audit framework for AI infomediaries Beyond the physician-choice findings, the study demonstrates a reusable template for auditing infomediaries, that is, AI systems that intermediate a person’s choice among options or among other people. The same frozen-design, randomized-conjoint machinery produced causal signal weights in two domains (hotels (Baig et al., 2026); physicians here) with identical anchor levels, and the full pipeline spans design generation, execution, estimation, equivalence testing, and Whose doctor does the AI recommend?23 monetary equivalents: auditing a new model against the frozen design is a single command. Because model updates can silently reweight signals, parity verdicts expire; we propose treating audits of this form as recurring monitoring rather than one-shot evaluation. The observed heterogeneity underscores this: the pro-female tilt ranges from near-zero (gemma3) to clearly positive across models, one candidate model failed the auditability gate entirely, and the exploratory frontier pilot hints at ethnicity orderings the confirmatory panel does not show. Verdicts are therefore model-specific and version-specific, motivating per-release audits. 5.4 Implications For platforms: provider-directory orderings feed AI assistants; a content-free first-position effect worth$11 per visit means interface choices allocate patient flow. For health-equity monitoring: group-level parity testing with prespecified equivalence bounds detected tilts that favor minority physicians, a reminder that “bias” in deployed systems need not match the direction of historical discrimination, and that monitoring must test for departures in both directions. For physicians and practices: the fee-equivalent table prices controllable signals such as recent reviews ($19), telehealth ($14), and responding to feedback ($10), against the dominant rating step ($157). For regulators: the stated-vs-revealed gap is the paper’s sharpest governance result. Demographic attributes moved choices but were mentioned in ≤0.03% of reasons; transparency obligations that rely on model self-report would not have detected the effects measured here. For the alignment debate on refusals: abstention was rare (0.39%) and indiscriminate rather than concentrated on demographically contrastive choice sets: these models did not “know when not to answer.” 5.5 Limitations Synthetic profiles bound external validity: real directories embed correlated attributes and photographs; we audit text-only cards. Name signaling identifies perceived-category effects, not mechanisms, and the significant name-exemplar placebo shows specific names carry information beyond their gender×ethnicity cell, so cell-level estimates are audit-level signals rather than clean category parameters. The audited panel comprises six small open-weight models and one proprietary production model (gpt-4o-mini); open-weight models under- represent the proprietary assistants patients most use, and findings are a snapshot of specific model versions (identifiers and dates recorded). The frontier-model session pilot narrows but does not resolve this gap: four Claude-family models were probed on the identical frozen design, and quantitative forecasts of their full-run behavior were archived in advance of Whose doctor does the AI recommend?24 measurement, a falsifiable-prediction step we propose as standard practice for staged audits. The pilot remains exploratory by construction (harness-embedded collection, default decoding, n= 50 choice sets per model), and the forecasts carry evidential weight only once graded against a protocol-compliant API run. Choice sets held specialty, board certification, and distance constant; effects of those attributes are outside this design. Equivalence bounds of ±1.5 p are a prespecified judgment call; smaller systematic effects could persist below them. 5.6 Conclusions A prespecified randomized audit of seven LLMs establishes causally that AI physician recommendation runs on ratings and price, carries a content-free position premium, and applies small systematic demographic tilts, favoring female- and minority-signaled names, that the models never disclose in their own explanations. None of these properties is observable from model self-report; all of them are measurable, cheaply and repeatably, with a frozen randomized instrument. As AI infomediaries absorb the referral function, recurring behavioral audits of this form, rather than explanation mandates, are the monitoring technology fit for purpose, and a frozen-design instrument of this kind makes each new model release auditable in a day. References Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, and Ameet Deshpande. GEO: Generative engine optimization. In Pro- ceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 5–16, 2024. doi: 10.1145/3637528.3671900. John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, Zechariah Zhu, Jessica B. Kelley, Dennis J. Faix, Aaron M. Goodman, Christopher A. Longhurst, Michael Hogarth, and Davey M. Smith. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183(6): 589–596, 2023. Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, and Asher Ali. Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection. Working paper, under review, 2026. Marianne Bertrand and Sendhil Mullainathan. Are Emily and Greg more employable than Whose doctor does the AI recommend?25 Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4):991–1013, 2004. Benjamin Edelman, Michael Luca, and Dan Svirsky. Racial discrimination in the sharing economy: Evidence from a field experiment. American Economic Journal: Applied Economics, 9(2):1–22, 2017. Martin Emmert, Uwe Sander, and Frank Pisch. Eight questions about physician-rating websites: A systematic review. Journal of Medical Internet Research, 15(2):e24, 2013. S. Michael Gaddis. How black are Lakisha and Jamal? Racial perceptions from names used in correspondence audit studies. Sociological Science, 4:469–489, 2017. Guodong Gordon Gao, Jeffrey S. McCullough, Ritu Agarwal, and Ashish K. Jha. A changing landscape of physician quality reporting: Analysis of patients’ online ratings of their physicians over a 5-year period. Journal of Medical Internet Research, 14(1):e38, 2012. Jens Hainmueller, Daniel J. Hopkins, and Teppei Yamamoto. Causal inference in conjoint anal- ysis: Understanding multidimensional choices via stated preference experiments. Political Analysis, 22(1):1–30, 2014. David A. Hanauer, Kai Zheng, Dianne C. Singer, Achamyeleh Gebremariam, and Matthew M. Davis. Public awareness, perception, and use of online physician rating sites. JAMA, 311 (7):734–735, 2014. Dana ̈e Metaxa, Joon Sung Park, Ronald E. Robertson, Karrie Karahalios, Christo Wilson, Jeff Hancock, and Christian Sandvig. Auditing algorithms: Understanding algorithmic systems from the outside in. Foundations and Trends in Human–Computer Interaction, 14 (4):272–344, 2021. Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, 2019. Jesutofunmi A. Omiye, Jenna C. Lester, Simon Spichak, Veronica Rotemberg, and Roxana Daneshjou. Large language models propagate race-based medicine. npj Digital Medicine, 6: 195, 2023. Alomi O. Parikh, Michael C. Oca, Jordan R. Conger, Allison McCoy, Jessica Chang, and Sandy Zhang-Nunes. Accuracy and bias in artificial intelligence chatbot recommendations for oculoplastic surgeons. Cureus, 16(4):e57611, 2024. doi: 10.7759/cureus.57611. Whose doctor does the AI recommend?26 Christian Sandvig, Kevin Hamilton, Karrie Karahalios, and Cedric Langbort. Auditing algorithms: Research methods for detecting discrimination on internet platforms. In Data and Discrimination: Converting Critical Concerns into Productive Inquiry (preconference of the 64th Annual Meeting of the International Communication Association). Seattle, WA, 2014. Niam Yaraghi, Weiguang Wang, Guodong Gordon Gao, and Ritu Agarwal. How online quality ratings influence patients’ choice of medical providers: Controlled experimental survey study. Journal of Medical Internet Research, 20(3):e99, 2018. Travis Zack, Eric Lehman, Mirac Suzgun, Jorge A. Rodriguez, Leo Anthony Celi, Judy Gichoya, Dan Jurafsky, Peter Szolovits, David W. Bates, Raja-Elie E. Abdulnour, Atul J. Butte, and Emily Alsentzer. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. The Lancet Digital Health, 6(1): e12–e22, 2024.