Paper deep dive
Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind
Kejia Zhang, Youran Sun, Chugang Yi, Haizhao Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/4/2026, 4:08:27 AM
Summary
This paper evaluates nine frontier Vision-Language Models (VLMs) on two psychology-derived Theory of Mind (ToM) benchmarks: the Keysar Director Task (visual perspective-taking) and the Frith-Happé animated triangles (intention attribution). The study finds a cross-task dissociation where no model maintains a coherent adult-like ToM profile across both tasks. Specifically, models exhibit egocentric errors in the Director Task similar to children, and under-attribute intention in the triangles task, resembling high-functioning autism (HF-ASD) profiles rather than typical development (TD).
Entities (9)
Relation Signals (7)
Frith-Happé animated triangles → measures → intention attribution
confidence 95% · The animated-triangles task asks whether a model can infer intention from abstract motion.
Keysar Director Task → measures → visual perspective-taking
confidence 95% · The Director Task asks whether a model can act from another viewer’s visual access.
Frontier VLMs → exhibits → cross-task dissociation
confidence 93% · No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles
Frontier VLMs → resembles → High-Functioning Autism (HF-ASD)
confidence 90% · On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean
Claude Opus 4.7 → performs → egocentric error
confidence 88% · claude-opus-4.7 scores 1/12 [on the implicit action prompt]
Gemini 3.0 Pro → performs → egocentric error
confidence 88% · gemini-3.0-pro is the exception at 12/12 [on the implicit action prompt, indicating failure to suppress egocentric bias]
Frontier VLMs → resembles → Typical Development (TD)
confidence 85% · On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78% of trials like children rather than adults
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchmarks: the Keysar Director Task (visual perspective-taking under egocentric interference) and the Frith-Happé animated triangles scored with the Castelli rubric (intention attribution from pure motion). On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78\% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models. On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random stay near TD. No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles, and the most TD-like model on the triangles is child-like on the Director Task. We report group-level descriptions, not diagnostic labels for any model.
Tags
Links
- Source: https://arxiv.org/abs/2608.00261v1
- Canonical: https://arxiv.org/abs/2608.00261v1
Trouble viewing inline? Open PDF directly →
Full Text
81,202 characters extracted from source content.
Expand or collapse full text
Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind Kejia Zhang∗ Youran Sun∗ Chugang Yi Haizhao Yang† University of Maryland, College Park Abstract Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchmarks: the Keysar Director Task (visual perspective-taking under egocentric interference) and the Frith-Happé animated triangles scored with the Castelli rubric (intention attribution from pure motion). On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models. On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random stay near TD. No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles, and the most TD-like model on the triangles is child-like on the Director Task. We report group-level descriptions, not diagnostic labels for any model. Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind Kejia Zhang∗ Youran Sun∗ Chugang Yi Haizhao Yang† University of Maryland, College Park †footnotetext: ∗Equal contribution. †Corresponding author. Emails: Youran Sun, sun1245@umd.edu; Haizhao Yang, hzyang@umd.edu. 1 Introduction A user shows a vision-language model (VLM) a tabletop scene shared with another viewer. The viewer sees fewer blocks than the model, so the model must track what that viewer can see before acting. Psychology research on Theory of Mind (ToM) studies the capacity to reason about another person’s perceptual or mental state. It also provides mature reference profiles across developmental and clinical groups (Abell et al., 2000; Castelli et al., 2000, 2002; Keysar et al., 2000). For VLMs, a single ToM score is not enough. The question is whether one model jointly aligns with a single adult human reference profile across ToM tasks, or fragments task by task. Recent VLM and large language model ToM benchmarks leave two gaps: they often rely on naturalistic stimuli whose faces, dialogue, and scene context afford social-cue shortcuts, and they probe one ToM sub-capacity at a time. As a result, they cannot show whether a single model’s ToM fragments across complementary facets (Section 2). We address this gap by pairing the Keysar Director Task for visual perspective-taking with the Frith-Happé animated triangles for abstract intention attribution (Keysar et al., 2000; Abell et al., 2000; Castelli et al., 2000). The Director Task asks whether a model can act from another viewer’s visual access. The animated-triangles task asks whether a model can infer intention from abstract motion. We choose this pair because both tasks come from psychology, test complementary ToM sub-capacities, and minimize social-cue shortcuts. The same frontier model panel evaluates both tasks (Section 4). Section 3 defines a joint coordinate system for comparing each model profile with published human reference profiles. Figure 1: Schematic of the two-benchmark cross-task design. The Director Task adapts Keysar et al. (2000) and probes visual perspective-taking through the explicit-implicit gap x=P(sub-prompt a correct)−P(sub-prompt b correct)x=P(sub-prompt a correct)-P(sub-prompt b correct). The animated-triangles task adapts Frith-Happé clips and probes abstract intention attribution through the ToM-condition (Intent, Approp) profile. The right panel previews the joint cross-task plane (Section 5). The analysis treats human groups as reference profiles, not diagnostic labels for models. For each task, we compare model profiles with the relevant human anchors. Across tasks, we report nearest-reference agreement and rank association (Section 5), keeping the focus on cross-task coherence rather than pass/fail ToM claims. Our main contributions are as follows: • We introduce a cross-task dissociation benchmark suite for VLM ToM, with the comparison to prior benchmarks in Table 1. • We release two reproducible psychology-grounded benchmarks: a Keysar Director Task adaptation and a Frith-Happé animated-triangles adaptation (Keysar et al., 2000; Abell et al., 2000; Castelli et al., 2000; Dureux et al., 2023). • We report a shared model-panel evaluation and cross-task reference-agreement analysis (Sections 4 and 5). 2 Related Work ToM reference profiles in psychology. The Heider and Simmel film first showed that adults read intentions into animated geometric shapes (Heider and Simmel, 1944). Most viewers describe the shapes as agents with social goals. Developmental psychology then built structured paradigms with age-graded and clinical-group profiles. Examples include Three-Mountain perspective-taking (Piaget and Inhelder, 1956), Level-1/Level-2 perspective-taking (Flavell et al., 1981), the Keysar Director Task for joint action (Keysar et al., 2000; Apperly et al., 2010), and Frith-Happé animated triangles with the Castelli rubric (Abell et al., 2000; Castelli et al., 2000, 2002). Perspective-taking asks whether a person suppresses a privileged view when communicating with a less-informed partner. Animated-triangles attribution asks whether a person reads goals and mental states from motion alone. Both paradigms report group means for typically-developing adults (TD-adult), high-functioning autism-spectrum adults (HF-ASD-adult), and age-graded children cohorts (Castelli et al., 2002; Apperly et al., 2010; Dumontheil et al., 2010; White et al., 2011; Livingston et al., 2019; Andersen et al., 2022; Begeer et al., 2010; Epley et al., 2004). Within a paradigm, these profiles separate one human group from another. That separation makes the published means useful anchors for frontier VLM panel profiles. We use those anchors descriptively, not as diagnostic categories for models. Existing VLM and LLM ToM benchmarks. VLM and large language model (LLM) ToM benchmarks differ in modality and sub-capacity, but one shared panel can show only so much under current designs. Four naturalistic VLM benchmarks retain social cues: household scenes and bodies in MMToM-QA (Jin et al., 2024), egocentric body cues in EgoToM (Li et al., 2025), character faces and dialogue in MoMentS (Villa-Cueva et al., 2025), and indoor affordances in MINDCUBE (Wang et al., 2026). These cues may offer a route to the answer without explicit mental-state inference. They also make it harder to separate mental-state inference from social-pattern matching. Our Director Task removes objects and bodies with labeled abstract blocks. Our animated-triangles task removes faces and dialogue with geometric motion. These benchmarks also isolate one sub-capacity, so they cannot test whether a model fragments across complementary facets. Gao et al. (2024) take the opposite design choice with an abstract three-jar perspective-taking probe, but they test only visual perspective-taking. We add intention attribution under the same model panel, making cross-task dissociation analysis possible in a controlled shared-panel design. No prior VLM ToM benchmark, naturalistic or abstract, pairs two complementary psychology-derived sub-tasks under a shared model panel. Cross-task dissociation as the scientific posture. Human ToM is not a single capacity. It includes perspective-taking, intention attribution, false-belief reasoning, affective mentalising, and second-order belief. These sub-capacities can dissociate across populations and development. The reference profiles above provide such within-paradigm signatures. On animated triangles, TD-adult and HF-ASD-adult differ on ToM-condition Intentionality (Intent) but coincide on Goal-Directed (GD). On the Director Task, young children and adults differ on the explicit-implicit gap but converge when perspective representation is explicitly cued. A single-task VLM benchmark cannot detect analogous within-model dissociation, because the comparison requires two paradigms under one panel. We therefore evaluate one perspective-taking paradigm and one intention-attribution paradigm on the same nine-model frontier panel. Cross-task analysis is the headline outcome rather than a follow-up ablation. Dissociation is the target phenomenon, not a secondary error analysis. Positioning of our contribution. We reduce social-cue reliance per benchmark with block-and-color stimuli on the Director Task and abstract geometric trajectories on the animated-triangles task. We discuss two residual shortcut risks in Appendix I: motion-pattern matching on animated triangles and color-position matching on the Director Task. Table 1 compares prior VLM ToM benchmarks with ours along four axes. The table makes the contrast explicit rather than leaving it to prose. Benchmark Sub-capacity measured Stimulus type Social-cue confounds reduced Paired-task cross-task design MMToM-QA belief and goal naturalistic video + text no (face, dialogue) no EgoToM belief and future action egocentric video no (real scene context) no MoMentS multiple ToM types narrative film clips no (face, dialogue, music) no MINDCUBE cognitive map and perspective 3D rendered partial (abstract) no Gao et al. (2024) visual perspective-taking only abstract three-jar scenes yes (abstract) no (one task) Ours perspective-taking + intention attribution tabletop blocks + animated triangles yes (color and motion shortcuts noted) yes (paired tasks) Table 1: Prior VLM Theory-of-Mind benchmarks vs. ours along sub-capacity, stimulus type, residual social-cue reduction, and paired-task design. Prior benchmarks are discussed and cited in Sec. 2; our residual shortcuts (color-position, motion-pattern) are detailed in Appendix I. 3 Two Benchmarks Section 3.1 defines the Director Task and its split between explicit perspective representation and the implicit know-but-don’t-use trap. The animated-triangles task and its 2D Intent/Appropriateness (Approp) profile follow in Section 3.2. The joint coordinate system appears in Section 3.3; Section 4 reports panel and calibration details. 3.1 The Keysar Director Task for Visual Perspective-Taking under Action The Director Task replaces natural social scenes with block-and-color tabletop stimuli, directly reducing face, dialogue, and scene-context cues. Its four scored outcomes split explicit perceptual perspective representation from the implicit know-but-don’t-use trap. This split lets one task report a sub-capacity profile rather than a single aggregate score. Stimuli are programmatically generated three-dimensional tabletop scenes with colored blocks and a director figure on the far side. A vertical opaque partition occludes one block from the director while leaving it visible to the tested model, creating a Keysar-style privileged-information asymmetry (Keysar et al., 2000). Each block carries a ground letter label (A, B, C) and a distinct color. The letter-and-color grounding lets the tested model refer to a block by either cue. In one representative scene, the tested model sees small, medium, and large blocks. The occluder hides the large block from the director, who sees only the small and medium blocks. The utterance “move the largest block to the right” therefore identifies the medium block from the director’s perspective. A tested model that uses its own view picks the occluded block, the egocentric error. For each scene the tested model receives three independent prompts with color-word answers. Prompt (a) asks which blocks the director can see, testing perceptual perspective representation. Prompt (b) is a single-select that asks which block to move under a director utterance, such as “the small block”. The utterance is ambiguous if the tested model considers all visible blocks. It becomes unique once the tested model adopts the director’s perspective. Prompt (b) is the know-but-don’t-use trap and implicit ToM probe. Prompt (c) is a two-step explicit gate: c.Q1 re-asks (a), and c.Q2 asks which block to move. The per-scene outcome is a four-tuple of binary correctness (a,b,c.Q1,c.Q2)(a,b,c.Q1,c.Q2). We also mark egocentric errors on (b) and c.Q2. The benchmark reports per-sub-prompt accuracy and the explicit-implicit gap P(a correct)−P(b correct)P(a correct)-P(b correct) (Gu et al., 2026). The evaluation harness and temporal scheduling discipline appear in Appendix A. 3.2 The Frith-Happé Animated Triangles for Abstract Intention Attribution The animated-triangles task keeps the social surface minimal: silent abstract motion, no faces, no speech, and no social labels. Composite stills give every tested model the same temporal evidence; the judge firewall separates generation from scoring. We retain the Castelli rubric so we can compare tested-model profiles with published human reference groups. Stimuli are the Frith-Happé animated-triangles clips re-edited by Dureux et al. (2023) under C BY 4.0. The clip family contains balanced ToM, GD, and Random conditions. In ToM clips, one triangle persuades, mocks, or deceives another. In GD clips, one triangle chases or follows another. In Random clips, the triangles drift independently. Each clip is approximately twenty seconds of silent abstract motion. One representative ToM clip shows two triangles in a rectangular enclosure. One shape persistently follows the other, while the other bobs and changes direction in response. Human raters often describe this sequence with mentalising words such as “coaxing” or “mocking” rather than literal kinematic descriptions. Each clip is presented to the tested model as one composite still image. The harness samples frames uniformly and tiles them row-major. Section 4 reports the frame count and grid layout used in the main run. Appendix L reports the frame-representation calibration. The tested model describes in free text what happens in the animation. The prompt contains no condition label and no ToM vocabulary. Scoring follows the Castelli et al. (2000) Appendix-2 rubric, with Intent (0–5) and Approp (0–3) dimensions. A separate LLM judge panel scores each description; Appendix A describes the rubric and code-enforced harness. The harness hides the tested model identity from the judge and hides the rater directory from the tested model. Appendices M and N report rubric validation and judge-panel checks. The raw scoring outputs flag same-family (judge, tested model) pairs. Appendix J reports the same-family judge-control analysis. The judge sees the clip’s condition label and script semantics, following Castelli’s human-rater protocol, but not the tested model identity. The judge prompt uses the anchor-blinded Castelli rubric without any per-clip Castelli-mean overlay (Appendix J). Aggregation Pipeline. For each (tested model, clip, judge) cell, we obtain an (Intent, Approp) score. For each tested model MiM_i, we take the judge median and then the mean over ToM clips. This yields a 2D ToM-condition profile: profileT(Mi)=(IntentToM,AppropToM).profile_T(M_i)=(Intent_ToM,Approp_ToM). Castelli et al. (2002) define the analogous human profile profileT(Hj)profile_T(H_j) for each reference group. The profile distance to human group HjH_j is: dT(Mi,Hj)=‖profileT(Mi)−profileT(Hj)‖2.d_T(M_i,H_j)= \|profile_T(M_i)-profile_T(H_j) \|_2. Distances use 2D ToM-condition (Intent, Approp) space. This metric feeds the per-benchmark analysis in Section 4 and the joint analysis in Section 5. Per-condition GD and Random profiles appear in Appendix E. 3.3 Joint Coordinate System and Cross-Task Analysis The joint coordinate system makes the cross-task comparison explicit. Shared coordinates let us visualize dissociation, measure nearest-reference switches, and report per-model disagreement. Each model or human reference group receives a joint coordinate (x,y)(x,y) for Figure 4. Human means come from Castelli et al. (2002) for animated triangles and from Begeer et al. (2010) and Dumontheil et al. (2010) for the Director Task. The Director-Task coordinate x is the explicit-implicit gap P(a correct)−P(b correct)P(a correct)-P(b correct). The same scalar definition applies to every model and human reference group. For groups with no published explicit-implicit gap, we reconstruct it from published a-question and b-question accuracies on the closest variant. Appendix H lists the per-group sources. Groups for which neither value is computable are excluded from Table 8. They are also excluded from the Figure 4 X axis. The animated-triangles coordinate y is the Figure 4 Y axis only. It is the ToM-condition profile distance to TD-adult. The Director-Task nearest-reference rule is nearest_ref(Mi,D)=argminj|x(Mi)−x(Hj)|nearest\_ref(M_i,D)= _j|x(M_i)-x(H_j)|. The animated-triangles nearest-reference rule uses the 2D profile distance: nearest_ref(Mi,T)=argminjdT(Mi,Hj).nearest\_ref(M_i,T)= _jd_T(M_i,H_j). For each model we report the pair (nearest_ref(Mi,D),nearest_ref(Mi,T))(nearest\_ref(M_i,D),nearest\_ref(M_i,T)). Table 8 also reports a binary disagreement flag. The analysis is descriptive rather than inferential. 4 Experiments 4.1 Experimental Setup Model panel. The tested-model panel contains nine frontier VLMs from five labs, all accessed through a unified gateway. The models are claude-opus-4.7, claude-sonnet-4.6, gpt-5.4, gpt-5.5, gemini-3.0-pro, gemini-3.5-flash, grok-4.3, kimi-k2.6, and qwen-3.5-plus. Director Task. The Director-Task block contains twelve director-perspective scenes. Each scene yields three API calls (sub-prompts a, b, c); c yields two scored outcomes (c.Q1 and c.Q2). The scored outcomes per model per seed are 12×(1+1+2)=4812×(1+1+2)=48. Animated triangles. The animated-triangles stimulus set contains twelve Frith-Happé clips, four per condition (ToM, GD, Random). Each clip appears as a composite image with sixteen uniformly sampled frames in a 4×44× 4 row-major grid. Each cell carries a 11–1616 temporal-order badge; Appendix K reports the badge-free variant. Appendix L reports the calibration for frame count, grid layout, and per-cell resolution. Judge panel. Animated-triangles free-text responses are scored by a three-judge cross-vendor LLM panel: claude-haiku-4.5, gemini-2.5-flash, and qwen2.5-72b-instruct. Each judge runs at temperature zero with no chain-of-thought (CoT). Section 3.2 defines the harness protocol. The panel was selected with a fourteen-anchor rubric validation battery and a seven-candidate judge ablation (Appendices M and N). An earlier judge prompt included a per-clip Castelli-mean overlay. The audit found that overlay made the VLM-versus-Castelli comparison partially circular. Production scoring strips the overlay while retaining the rubric and condition-label/script-semantics disclosure (Appendix J). Replication. Both benchmarks use three independently seeded tester trials per (tester, item) cell (Ktester=3K_tester=3). We report the mean across trials and, for animated triangles, across judges, rounded to each cell’s native integer scale. For binary Director-Task cells, the rounded mean equals majority vote across trials. Appendix Table 2 therefore reports integer counts out of 1212, not thirds-resolution counts out of 3636. Each replication has 9×12×4=4329× 12× 4=432 Director-Task outcomes and 9×12×3=3249× 12× 3=324 animated-triangles judge cells. Both benchmarks use three replications. To decorrelate API calls from time-of-day server-load variance, batches are scheduled across non-contiguous time windows (Appendix A). 4.2 Director Task Results On the canonical Director-Task block (Figure 2), the panel scores 71.3%71.3\% on the explicit visibility multi-select (a) but collapses to 12.0%12.0\% on the implicit action prompt (b), an explicit-implicit gap of 59.359.3 percentage points. The collapse matches the egocentric error in the human Director-Task literature (Keysar et al., 2000; Apperly et al., 2010; Dumontheil et al., 2010): on 77.8%77.8\% of (b)-cells the model moves the privileged-view block. Seven of nine models score 0/120/12 on (b), and claude-opus-4.7 scores 1/121/12; gemini-3.0-pro is the exception at 12/1212/12. The explicit gate (c) only partly reopens the trap. Re-asking visibility (c.Q1) lifts the panel to 63.9%63.9\% and the gated action (c.Q2) to 59.3%59.3\%, but the per-model recovery P(c.Q2)−P(b)P(c.Q2)-P(b) splits sharply: three models recover by ≥80≥ 80 points (claude-sonnet-4.6 0/12→10/120/12→ 10/12, gemini-3.5-flash 0/12→12/120/12→ 12/12, gpt-5.5 0/12→10/120/12→ 10/12) and qwen-3.5-plus by 6767 (0/12→8/120/12→ 8/12), while gpt-5.4 stays at 0/120/12. Perspective-use thus unblocks under the gate for some models but not others. Against published human references, the panel-mean gap of 0.590.59 falls between the TD-adult mean of ≈0.44≈ 0.44 (Begeer et al. (2010) controls 0.430.43, Dumontheil et al. (2010) adults 0.440.44) and the Dumontheil et al. (2010) child range, 0.590.59 (ages 14.014.0–17.717.7) to 0.720.72 (ages 7.37.3–9.79.7); the HF-ASD-adult gap is smaller still at 0.340.34, consistent with the rule-based heuristic reported there (Begeer et al., 2010). A reach-based variant brackets the same child–adult separation (Epley et al., 2004): young children (n=33, mean age 6.26.2 y) make 52%52\% egocentric reaches versus 24%24\% for adults. We report this correspondence descriptively, without per-model developmental-age estimates. Two supplementary controls confirm this attribution rather than a perceptual or instruction-following floor. A floor control (Appendix C), which makes the (b) target the addressee’s own visual extremum, lifts all seven non-excluded testers (qwen-3.5-plus and kimi-k2.6 excluded for latency) to letter-level ≥11/12≥ 11/12, including the six scoring ≤1/12≤ 1/12 canonically. A three-to-four block-count control (Appendix D) leaves the (b) egocentric rate flat, marking trap activation as a categorical model property rather than a graded working-memory bottleneck. Figure 2: Director-Task per-model accuracy on the four color-word sub-prompts (Sec. 3.1). The panel contains nine models, twelve canonical scenes, and three trials per (model, scene, sub-prompt) cell. The (a)–(b) gap is the explicit-implicit gap. Seven of nine models score zero on (b), claude-opus-4.7 scores 1/121/12, and gemini-3.0-pro is the exception at 12/1212/12. Model colors match Figs. 3, 4, and 5. Per-model accuracies appear in Appendix Table 2. 4.3 Animated Triangles Results Figure 3 shows each model’s ToM-condition (Intent, Approp) profile against the Castelli et al. (2002) TD-adult and HF-ASD-adult group means. Figure 3: Per-model (Intent, Approp) profile on the animated-triangles ToM condition, using the metric from Section 3.2 and the anchor-blinded rubric from Appendix J. Each filled circle is one frontier VLM, marked by monogram and colored by lab. Stars mark the TD-adult and HF-ASD-adult group means from Castelli et al. (2002). All nine models lie closer to HF-ASD-adult than to TD-adult on this condition. All nine models fall below both TD-adult ToM means (Intent 4.34.3, Approp 1.71.7). The panel-mean ToM profile (2.31,0.94)(2.31,0.94) lies 2.132.13 from TD-adult and 0.730.73 from HF-ASD-adult (2.9,0.5)(2.9,0.5); the GD mean (1.97,1.98)(1.97,1.98) lies 0.510.51 and 0.800.80 from the two groups, and the Random mean (0.92,2.57)(0.92,2.57) lies 0.880.88 and 1.081.08. ToM is the only condition whose panel mean sits closer to HF-ASD-adult than to TD-adult. The asymmetry holds model by model: all nine models lie nearer HF-ASD-adult than TD-adult on ToM (per-model coordinates and per-condition breakdowns in Appendix Tables 7, 5). These are geometric distances to published group means, not clinical assessments. The means are human-rated, while our scores are LLM-rated; we treat them as comparable because the judge panel recovers human-anchored exemplars within ±0.5± 0.5 on the 14-anchor battery (Appendix N, M). Agreement between judges and humans on production cells is unmeasured, a noise floor we flag in Limitations. Per-tester ranks are more heterogeneous than the panel-mean collapse implies: the strict Intent order ToM>GD>RandomToM>GD>Random holds for five of nine testers (claude-opus-4.7, gemini-3.0-pro, gemini-3.5-flash, gpt-5.5, qwen-3.5-plus), even though the panel-median ToM and GD Intent coincide at 2.02.0. The recoverers are the models that benefit from explicit temporal cueing; models that fail under both cued and uncued formats drive the ToM collapse (stimulus-format ablation, Appendix K). The ToM-toward-HF-ASD shift is not an artifact of the stimulus encoding: on a four-tester subset it survives both a static-keyframe and a time-reversed control (Appendix F), with ToM nearest HF-ASD-adult in all three arms. The time-reversed arm also reproduces the forward-time panel mean almost exactly, showing the panel is largely insensitive to temporal direction, consistent with the motion-pattern-matching shortcut of Appendix I. 5 Cross-Task Dissociation Analysis 5.1 Joint Dissociation Figure 4 places the nine frontier VLMs and the relevant human reference groups in the joint coordinate system of Section 3.3. The X axis is the Director-Task explicit-implicit gap; the Y axis is the animated-triangles ToM-condition profile distance to TD-adult, used only as a visualization anchor. Figure 4: Joint coordinates of the nine frontier VLMs (filled circles, two-letter monograms per Fig. 3) and the two adult reference groups from Castelli et al. (2002) (stars). X axis: Director-Task explicit-implicit gap P(a)−P(b)P(a)-P(b) (TD-adult 0.4350.435, HF-ASD-adult 0.3400.340; sources in Sec. 3.3). Y axis: animated-triangles ToM-condition distance to TD-adult, a visualization anchor only; the triangles nearest-reference rule uses the full 2D (Intent, Approp) space. For eight of nine models the nearest adult reference differs between the two tasks; the full cross-tabulation appears in Appendix Table 8. The per-model nearest-reference cross-tabulation is summarized below and reported in full in Appendix Table 8. The reference set is restricted to the two adult cohorts (TD-adult and HF-ASD-adult) on both tasks, since Castelli et al. (2002) publish animated-triangles group means only for adult cohorts and no equivalent children publication exists; this is the largest reference set for which a like-for-like cross-task comparison is defined. The wider age-graded Director-Task references of Dumontheil et al. (2010) underpin the Director-Task panel-mean gap reported above and the anchor-sensitivity analysis of Appendix H. Under that wider Director anchor set, most panel models reassign to a child or adolescent cohort rather than to TD-adult. This follows from the panel-mean Director gap of 0.590.59 exceeding the TD-adult anchor of 0.4350.435 and does not change the cross-task disagreement pattern reported below. For eight of nine models the nearest-reference pair (nearest_refA,nearest_refB)(nearest\_ref_A,nearest\_ref_B) is unequal, all eight in the same direction (nearer TD-adult on the Director-Task gap scalar but nearer HF-ASD-adult on the animated-triangles 2D ToM-condition plane). “Nearest” here is restricted to the two adult anchors and is not a statement of absolute proximity; the panel-mean Director-Task gap of 0.590.59 itself exceeds the TD-adult anchor of 0.4350.435. The exception is gemini-3.0-pro, which scores 12/1212/12 on every Director-Task sub-prompt. Its explicit-implicit gap of 0 sits closer to the HF-ASD-adult anchor (0.340.34) than to TD-adult (0.4350.435), and it is also nearer HF-ASD-adult on the animated triangles, making it the only panel model whose nearest adult anchor is HF-ASD-adult on both tasks (a relative-distance artifact at the saturated end of the Director Task, not a substantive alignment claim). For eight of the nine panel members no single adult reference group is jointly closest under both the Director-Task gap scalar and the animated-triangles 2D profile distance, so no single adult reference profile is nearest on both tasks for these eight models in our panel. The animated-triangles half does not change under the stimulus-format ablation in Appendix K: the plain 4×44× 4 grid and the numbered-grid variant agree on every per-condition nearest-group verdict. 5.2 Cross-Benchmark Rank Association As a secondary descriptive check, we summarize whether the per-model rankings on the two benchmarks are monotonically associated. The Spearman rank correlation between Director-Task explicit-implicit gap and animated-triangles ToM-condition TD-distance is ρ=−0.19ρ=-0.19 (point estimate, n=9n=9, ties broken by mid-rank). At n=9n=9 this is an underpowered estimate (the 95% non-parametric interval is wide enough to be uninformative), consistent with no clear monotonic rank association in this panel. 6 Discussion and Conclusion VLM ToM in this panel is task-dependent, so a single benchmark is not enough. Future suites should report multi-benchmark profiles and treat profile dissociation as a main outcome. The numbered-grid ablation supports this reading: strict Intent rank ToM>GD>RandomToM>GD>Random rises from zero of nine testers to five of nine without changing any per-condition nearest-group verdict (Appendix K). This points to a temporal-order bottleneck for the five recovering testers; the other four fail regardless of cue. Three design implications follow. Suites should pair at least two psychology-derived sub-capacities under one model panel, report per-task per-model profiles rather than single aggregate accuracy, and pair abstract with naturalistic stimuli to separate social-cue performance from explicit mental-state inference. Frontier VLMs still fall short on these two low-social-cue ToM probes. For eight of nine models, the nearest adult reference differs between perspective-taking and intention-attribution profile space. The main result is therefore a panel-level dissociation, not alignment with one adult human reference profile. Limitations Our finding is panel-level descriptive and not an individual-model diagnostic claim. Both benchmarks use Ktester=3K_tester=3 independently seeded trials per (tester, item) cell, and the per-cell value reported throughout is the mean across the three trials. Animated-triangles scoring is LLM-as-rater rather than human rater, while Castelli’s reference group means come from human-rated free-text; this pipeline mismatch is a measurement noise floor that 14-anchor rubric calibration cannot fully eliminate (see Appendix J). Our Director-Task implementation uses simplified block-only stimuli with color-word scoring rather than the full Director Task with physical action selection. Two benchmarks alone do not cover the full breadth of ToM; false belief, theory-driven affect, and language-based ToM reasoning are left for future work. The Director-Task floor and block-count controls (Appendices C and D) cover seven of the nine testers. qwen-3.5-plus and kimi-k2.6 were excluded from those controls for latency reasons, so the perspective-taking attribution and the categorical/graded distinction extend cleanly to the seven covered testers. We abstain from those attributions for the excluded two. The animated-triangles motion-probe controls (Appendix F) cover four of the nine testers and run at Ktester=1K_tester=1 rather than the canonical Ktester=3K_tester=3; the per-arm panel-mean shifts reported there are point estimates. We did not pre-register the analyses. The C4 cross-task disagreement is a within-adult-anchor comparison, not a general claim that VLM behavior fails to match any human reference profile; the two-adult reference set is the largest like-for-like cross-task set defined in the published psychology literature, since Castelli et al. (2002) publish animated-triangles group means only for adult cohorts. Appendix H reports the asymmetric Director-side wider-anchor sensitivity, where most panel models reassign to child or adolescent anchors on the Director Task; the exact nearest-anchor labels in C4 should not be generalized beyond the available two-adult cross-task anchor set. Ethical considerations This paper compares frontier VLM behavior against published group-mean profiles from TD-adult, HF-ASD-adult, and age-graded children cohorts, all of which appear in the public psychology literature (Castelli et al., 2002; Apperly et al., 2010; Keysar et al., 2000; Begeer et al., 2010; Dumontheil et al., 2010; Epley et al., 2004). We use these reference group means only as quantitative anchors for distance comparisons, and we avoid per-model diagnostic labels, developmental-age point estimates, and anthropomorphic clinical equivalences (see Section Limitations). We use the HF-ASD-adult group mean as a reference profile in a descriptive comparison and not as a label or a value judgment on any model, system, or person. Our benchmarks contain no human subjects data and no personally identifying information. The Frith-Happé animated-triangles clips we use are the eLife re-edits of Dureux et al. (2023), released under C BY 4.0. References F. Abell, F. Happé, and U. Frith (2000) Do triangles play tricks? attribution of mental states to animated shapes in normal and abnormal development. Cognitive Development 15 (1), p. 1–16. External Links: ISSN 0885-2014, Link, Document Cited by: 2nd item, §1, §1, §2. N. K. Andersen, M. K. Rimvall, P. Jeppesen, M. Bentz, J. R. M. Jepsen, L. Clemmensen, R. K. Jacobsen, and E. M. Olsen (2022) A psychometric investigation of the multiple-choice version of animated triangles task to measure theory of mind in adolescence. PLOS ONE 17 (3), p. e0264319. External Links: ISSN 1932-6203, Link, Document Cited by: §2. I. A. Apperly, D. J. Carroll, D. Samson, G. W. Humphreys, A. Qureshi, and G. Moffitt (2010) Why are there limits on theory of mind use? evidence from adults’ ability to follow instructions from an ignorant speaker. Quarterly Journal of Experimental Psychology 63 (6), p. 1201–1217. External Links: ISSN 1747-0226, Link, Document Cited by: §2, §4.2, Ethical considerations. S. Begeer, B. F. Malle, M. S. Nieuwland, and B. Keysar (2010) Using theory of mind to represent and take part in social interactions: comparing individuals with high-functioning autism and typically developing controls. European Journal of Developmental Psychology 7 (1), p. 104–122. External Links: Document Cited by: Appendix H, §2, §3.3, §4.2, Ethical considerations. F. Castelli, C. Frith, F. Happé, and U. Frith (2002) Autism, asperger syndrome and brain mechanisms for the attribution of mental states to animated shapes. Brain 125 (8), p. 1839–1849. External Links: ISSN 1460-2156, Link, Document Cited by: Appendix J, Appendix M, Appendix H, §1, §2, §3.2, §3.3, Figure 3, §4.3, Figure 4, §5.1, Limitations, Ethical considerations. F. Castelli, F. Happé, U. Frith, and C. Frith (2000) Movement and mind: a functional imaging study of perception and interpretation of complex intentional movement patterns. NeuroImage 12 (3), p. 314–325. External Links: ISSN 1053-8119, Link, Document Cited by: Appendix M, 2nd item, §1, §1, §2, §3.2. I. Dumontheil, I. A. Apperly, and S. Blakemore (2010) Online usage of theory of mind continues to develop in late adolescence. Developmental Science 13 (2), p. 331–338. External Links: ISSN 1467-7687, Link, Document Cited by: Appendix H, Appendix H, §2, §3.3, §4.2, §4.2, §5.1, Ethical considerations. A. Dureux, A. Zanini, J. Selvanayagam, R. S. Menon, and S. Everling (2023) Gaze patterns and brain activations in humans and marmosets in the frith-happé theory-of-mind animation task. eLife 12, p. e86327. External Links: ISSN 2050-084X, Link, Document Cited by: 2nd item, §3.2, Ethical considerations. N. Epley, C. K. Morewedge, and B. Keysar (2004) Perspective taking in children and adults: equivalent egocentrism but differential correction. Journal of Experimental Social Psychology 40 (6), p. 760–768. External Links: Document Cited by: Appendix H, §2, §4.2, Ethical considerations. J. H. Flavell, B. A. Everett, K. Croft, and E. R. Flavell (1981) Young children’s knowledge about visual perception: further evidence for the level 1–level 2 distinction.. Developmental Psychology 17 (1), p. 99–103. External Links: ISSN 0012-1649, Link, Document Cited by: §2. Q. Gao, Y. Li, H. Lyu, H. Sun, D. Luo, and H. Deng (2024) Vision language models see what you want but not what you see. External Links: 2410.00324, Link Cited by: §2, Table 1. Y. Gu, O. Tafjord, H. Kim, J. Moore, R. L. Bras, P. Clark, and Y. Choi (2026) SimpleToM: exposing the gap between explicit tom inference and implicit tom application in llms. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §3.1. F. Heider and M. Simmel (1944) An experimental study of apparent behavior. The American Journal of Psychology 57 (2), p. 243–259. External Links: ISSN 0002-9556, Link, Document Cited by: §2. C. Jin, Y. Wu, J. Cao, J. Xiang, Y. Kuo, Z. Hu, T. Ullman, A. Torralba, J. Tenenbaum, and T. Shu (2024) MMToM-QA: multimodal theory of mind question answering. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, p. 16077–16102. External Links: Link, Document Cited by: §2. B. Keysar, D. J. Barr, J. A. Balin, and J. S. Brauner (2000) Taking perspective in conversation: the role of mutual knowledge in comprehension. Psychological Science 11 (1), p. 32–38. External Links: ISSN 1467-9280, Link, Document Cited by: Figure 1, 2nd item, §1, §1, §2, §3.1, §4.2, Ethical considerations. Y. Li, V. Veerabadran, M. L. Iuzzolino, B. D. Roads, A. Celikyilmaz, and K. Ridgeway (2025) EgoToM: benchmarking theory of mind reasoning from egocentric videos. External Links: 2503.22152, Link Cited by: §2. L. A. Livingston, B. Carr, and P. Shah (2019) Recent advances and new directions in measuring theory of mind in autistic adults. Journal of Autism and Developmental Disorders 49 (4), p. 1738–1744. External Links: ISSN 1573-3432, Link, Document Cited by: §2. J. Piaget and B. Inhelder (1956) The child’s conception of space. Routledge and Kegan Paul, London. Note: English translation by F. J. Langdon and J. L. Lunzer; original French edition Presses Universitaires de France, 1948 Cited by: §2. E. Villa-Cueva, S. M. M. Ahmed, R. Chevi, J. C. B. Cruz, K. Elzeky, F. Cristobal, A. F. Aji, S. Wang, R. Mihalcea, and T. Solorio (2025) MoMentS: a comprehensive multimodal benchmark for theory of mind. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, p. 22591–22611. External Links: Link, Document Cited by: §2. Q. Wang, B. Yin, P. Zhang, J. Zhang, K. Wang, Z. Wang, J. Zhang, K. Chandrasegaran, H. Liu, R. Krishna, S. Xie, J. Wu, L. Fei-Fei, and M. Li (2026) MindCube: spatial mental modeling from limited views. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2. S. J. White, D. Coniston, R. Rogers, and U. Frith (2011) Developing the frith-happé animations: a quick and objective test of theory of mind for adults with autism. Autism Research 4 (2), p. 149–154. External Links: ISSN 1939-3792, Link, Document Cited by: §2. Appendix A Implementation Firewall Details Both benchmarks share a code-enforced two-firewall harness that hides tester identity from judges and hides rater-side files from testers; we describe each firewall separately below rather than label the combination “double-blind”, since the rater is deliberately given the clip’s ground-truth condition label and script semantics in keeping with Castelli’s original human-rater protocol. The first firewall bars testers from reading any path under the rater directory and from reading the test-set registry that maps tester display names to backing IDs. The second firewall bars judges from reading tester short names; only opaque tester IDs reach the judge prompt. All run artifacts are written atomically append-only with an open-exclusive create flag so that no past run can be silently rewritten. Temporal decorrelation against server-load variance. API call batches in this work, spanning the two main-panel benchmarks, every calibration pre-experiment, and every shuffleseed replicate, are scheduled at different times of day rather than executed in a single contiguous burst. Time-correlated load variance on the model providers’ inference servers, including throttling, queue depth, and concurrent traffic spikes, is therefore averaged across calls rather than concentrated within one window. Concretely, the animated-triangles plain-grid and numbered-grid main panels are launched approximately 22 hours apart, the Ktester=3K_tester=3 replicates of the frame-representation pre-experiment (Appendix L) span two calendar days, and the rubric and judge validation battery (Appendix M), the judge model ablation (Appendix N), and the Director-Task shuffleseed replicates are each launched on distinct calendar days. The full harness, scoring scripts, and per-(tester, clip, judge) raw outputs will be released with the camera-ready. Appendix B Per-Model Director Task Sub-Prompt Accuracy Table 2 reports the per-model raw correct count out of 1212 canonical Director-Task scenes on each of the four sub-prompts. The egocentric-error rate on sub-prompt (b), reported as “ego. (b)” in the rightmost column, is the fraction of cells in which the model picks the privileged-view block that would be ambiguous for the addressee but is in fact occluded from the director. Seven of nine models score 0 correct on sub-prompt (b) and claude-opus-4.7 scores 11 of 1212; the lone exception that does not exhibit the egocentric failure mode is gemini-3.0-pro at 12/1212/12. Model a b c.Q1 c.Q2 ego. (b) opus-4.7 11 1 8 2 9 sonnet-4.6 8 0 7 10 9 gemini-3.0-pro 12 12 12 12 0 gemini-3.5-flash 10 0 11 12 12 gpt-5.4 10 0 3 0 12 gpt-5.5 8 0 10 10 10 grok-4.3 7 0 6 6 12 kimi-k2.6 5 0 4 4 9 qwen-3.5-plus 6 0 8 8 11 panel mean (%) 71.3 12.0 63.9 59.3 77.8 Table 2: Director-Task per-model raw counts (out of 1212 canonical scenes) on each sub-prompt, plus egocentric-error count on sub-prompt (b). Sub-prompt (a) is the multi-select “which blocks does the director see”. Sub-prompt (b) is the single-select know-but-don’t-use trap. Sub-prompt (c) is split into Q1 (explicit re-asking of a) and Q2 (gated action after Q1). The panel-mean Director-Task explicit-implicit gap P(a)−P(b)P(a)-P(b) is 59.359.3 percentage points. Appendix C Director Task Perceptual and Instruction Floor Control Design. We run two control arms on the canonical twelve Director-Task scenes. The no-occluder arm removes the occluder so all three blocks are visible to both addressee and director; the (a, b, c.Q1, c.Q2) sub-prompt structure is unchanged. The same-side-director arm keeps the occluder but moves the director to the addressee’s side of the table, so the occluder does not occlude anything from the director (occluder is physically present but geometrically inert). Both arms invert the sub-(b) ground-truth target onto the addressee’s visual extremum, so a model that simply picks the visual extremum scores correctly on the control sub-(b). The remaining failure modes are perceptual mis-identification (cannot see all three blocks, mis-counts blocks, mis-takes occluder for a block) or instruction failure (does not understand “largest”/“smallest”). Panel and pre-registered floor pass criterion. We run both arms on the same nine-model panel minus qwen-3.5-plus and kimi-k2.6, excluded for latency reasons (the same exclusion holds for Appendix D). The seven testers retained are claude-opus-4.7, claude-sonnet-4.6, gpt-5.4, gpt-5.5, gemini-3.0-pro, gemini-3.5-flash, and grok-4.3. The canonical-(b) attribution for the two excluded testers is therefore not covered by this control. The floor pass criterion, fixed before running paid API calls, is per-model control sub-(b) letter-level correctness ≥11/12≥ 11/12, where letter-level marks the correct letter regardless of color naming (failure-mode ∈correct,letter_only_correct∈\correct,letter\_only\_correct\). The letter-level threshold isolates the perceptual / instruction floor from the orthogonal blue → cyan color-naming drift that appears panel-wide on the closed color vocabulary. Result. All seven testers pass the floor on both arms (7/7 on each arm). Table 3 reports per-model canonical-(b) versus control-(b) letter-level scores. Tester canon-(b) letter no-occluder (b) letter same-side-director (b) letter claude-opus-4.7 1/12 12/12 12/12 claude-sonnet-4.6 0/12 12/12 12/12 gemini-3.0-pro 12/12 12/12 12/12 gemini-3.5-flash 0/12 12/12 12/12 gpt-5.4 0/12 12/12 12/12 gpt-5.5 0/12 12/12 12/12 grok-4.3 0/12 12/12 11/12 Table 3: Director-Task floor control. canon-(b) is from the canonical 12-scene Director Task; the no-occluder arm removes the occluder; the same-side-director arm keeps the occluder but moves the director to the addressee side. Letter-level == correct letter regardless of color name; strict pair-level results lie 2/122/12 below letter-level for every cell because Q01 and Q02 expect blue and the panel reads blue as cyan. The six models with canonical ≤1/12≤ 1/12 score ≥11/12≥ 11/12 on both control arms, attributing the canonical collapse to perspective-taking rather than perceptual or instruction floor. Independent side finding on the same-side-director arm. On the same-side-director arm, the two Claude testers show degraded sub-(a) and sub-c.Q1 identification despite passing the sub-(b) floor. claude-opus-4.7 scores sub-(a) 7/127/12 (no-occluder arm: 12/1212/12) and c.Q1 4/124/12; claude-sonnet-4.6 scores sub-(a) 12/1212/12 but c.Q1 9/129/12 and c.Q2 7/127/12. This is consistent with the two Claude models maintaining a residual “the occluder physically blocks the director’s view” inference even when the geometry no longer supports it. The finding is independent of the floor result and does not affect the attribution conclusion; the floor judgment lives in sub-(b) only. Appendix D Director Task Variable Block Count Design. We add a four-block variant of the Director Task structure to test whether the canonical sub-(b) failure rate scales with the number of competing blocks (graded, working-memory-style) or is categorical (binary, the model either represents the director’s view or it does not). Four blocks of distinct sizes are placed in front of the director. Exactly one block (the global extremum under the director’s instruction) is occluded from the director by the partition; the director then issues an extremum instruction (“move the largest/smallest block to the right”) and the ground-truth target is the extremum among the director-visible set (the second-largest or second-smallest of the four). The (a, b, c.Q1, c.Q2) sub-prompt structure and the deterministic regex-based scoring are unchanged. We sample twelve balanced four-block scenes (six largest-instruction, six smallest-instruction, with the egocentric-trap position balanced across the four block positions). The n=3 baseline used for the matched-pair comparison is the same seven testers on the same closed-color-vocabulary scorer as the four-block run. Panel. Same seven testers as Appendix C (qwen-3.5-plus and kimi-k2.6 excluded for latency). Result. Table 4 reports per-tester sub-(b) direct-action correct and egocentric error rate at n=3 versus n=4 competing blocks. Tester n=3 (b) correct n=3 (b) ego n=4 (b) correct n=4 (b) ego gemini-3.0-pro 12/12 0/12 12/12 0/12 gpt-5.5 0/12 10/12 0/12 10/12 claude-sonnet-4.6 0/12 10/12 0/12 10/12 claude-opus-4.7 0/12 9/12 0/12 10/12 gpt-5.4 0/12 10/12 0/12 10/12 gemini-3.5-flash 0/12 10/12 0/12 10/12 grok-4.3 0/12 10/12 3/12 7/12 Table 4: Director-Task sub-(b) direct-action correct and egocentric error rate at n=3 vs n=4 competing blocks, seven-tester panel. The sole non-collapsing model at n=3 (gemini-3.0-pro) stays at 12/1212/12 correct under n=4. The six collapsing models stay at 0/120/12 direct-action correct with ≥7/12≥ 7/12 egocentric error under n=4. Both ends of the panel are flat with respect to block count. Take-away. The sub-(b) failure rate is flat with block count in this panel, neither the capable model nor the six collapsed models move with the additional competing block. This is consistent with sub-(b) failure being a categorical model property (the model either represents the director’s view in action or it does not) rather than a graded working-memory or attentional bottleneck that should grow with competing items. The explicit-gate CoT recovery on c.Q2 does decline with the extra block for some models (claude-sonnet-4.6 10→510→ 5, gpt-5.5 10→910→ 9, claude-opus-4.7 2→12→ 1), while remaining steady for gemini-3.5-flash (12→1212→ 12) and gpt-5.4 (0→00→ 0); the recovery channel is graded for the models whose recovery is CoT-mediated, even when the underlying trap activation is categorical. Appendix E Per-Condition Breakdown for the Animated Triangles The main text reports the animated-triangles ToM-condition (Intent, Approp) profile and the corresponding profile distances under the numbered-grid canonical metric. Figure 5 and Table 5 report the per-condition (GD, Random) panels and the per-condition profile distances under the same v2 canonical metric. Figure 5: Per-model (Intent, Approp) profile on the animated triangles broken down by condition. Each panel shows the per-model points, colored by lab. The canonical ToM-condition panel is also reported as Figure 3 in the main text. ToM cond. GD cond. Random cond. Model to TD to ASD to TD to ASD to TD to ASD opus-4.7 1.811 0.478 0.765 1.148 1.250 1.146 sonnet-4.6 3.309 1.652 1.046 0.910 1.203 1.516 gpt-5.4 2.221 0.937 0.672 0.992 1.230 1.383 gpt-5.5 1.351 0.839 0.827 0.975 1.203 1.516 gemini-3.0-pro 1.590 0.417 0.754 1.141 0.887 1.090 gemini-3.5-flash 1.249 0.935 0.843 1.229 1.347 0.929 grok-4.3 3.639 1.929 1.647 1.476 1.203 1.548 kimi-k2.6 2.564 1.178 0.864 1.217 1.037 1.351 qwen-3.5-plus 1.483 0.667 0.389 0.786 0.902 1.168 Table 5: Per-model 2D Euclidean profile distance to TD-adult and HF-ASD-adult group means under each animated-triangles condition (ToM, GD, Random), numbered-grid stimulus, anchor-blinded canonical rubric per Appendix J. The ToM-condition columns are the canonical animated-triangles metric (Sec. 3.2). On the ToM condition, all nine models have larger distance to TD-adult than to HF-ASD-adult. Appendix F Animated-Triangles Motion-Probe Controls Design. We run two stimulus-format controls on the canonical twelve animated-triangles clips. The static-keyframe arm replaces each 4×44× 4 composite with a single mid-clip frame in native 1280×7201280× 720 resolution. The time-reversed arm keeps the same sixteen-cell 4×44× 4 composite and the 11–1616 temporal-order badges, but reverses the underlying frame order so that the cell labelled “frame 1” shows what was originally the last sampled frame; the tested model is not told the order is reversed. Both controls use the same Castelli rubric, the same three-judge cross-vendor panel, and the same anchor-blinded canonical scoring as the main animated-triangles run. Panel and replication. Four testers, the per-lab best animated-triangles performer for the four Western labs (claude-opus-4.7, gpt-5.5, gemini-3.5-flash, grok-4.3); qwen-3.5-plus and kimi-k2.6 are excluded for latency, the same exclusion as in Appendices C and D. Both controls run with Ktester=1K_tester=1 per (tester, clip) cell; the canonical forward-time reference used here is the nine-model main run sub-selected to the same four testers. Result. Table 6 reports the per-condition panel-mean (Intent, Approp) for the three arms and the Euclidean distance to each adult reference group. Condition Arm (Intent, Approp) dist. TD dist. HF-ASD Nearest ToM canonical fwd (2.65, 1.08) 1.77 0.64 HF-ASD ToM time-reversed (2.56, 1.06) 1.85 0.66 HF-ASD ToM static keyframe (1.96, 0.73) 2.53 0.97 HF-ASD GD canonical fwd (1.96, 1.83) 0.46 0.69 TD GD time-reversed (1.96, 1.67) 0.44 0.57 TD GD static keyframe (1.98, 1.25) 0.62 0.42 HF-ASD Random canonical fwd (0.98, 2.19) 0.62 0.71 TD Random time-reversed (0.81, 2.35) 0.64 0.85 TD Random static keyframe (1.48, 1.00) 1.26 0.84 HF-ASD Table 6: Animated-triangles motion-probe controls on the four-tester subset (claude-opus-4.7, gpt-5.5, gemini-3.5-flash, grok-4.3, the per-lab Western frontier best on the animated triangles). The canonical-forward arm is the nine-model main run sub-selected to the same four testers. The time-reversed arm reproduces the canonical forward-time arm almost exactly on all three conditions. The static-keyframe arm shifts the ToM panel mean toward lower Intent and the GD/Random panel means toward lower Approp; ToM nevertheless remains nearer HF-ASD-adult than TD-adult under all three arms. Take-away. On ToM, the time-reversed and forward-time panel means coincide to within ∼0.1 0.1 on both axes, and both lie at distance ∼0.65 0.65 to HF-ASD-adult and ∼1.8 1.8 to TD-adult; reversing temporal direction does not change the asymmetric ToM-toward-HF-ASD shift. The static-keyframe arm reduces ToM-condition panel mean Intent from 2.652.65 to 1.961.96 but the panel mean still lands closer to HF-ASD-adult than to TD-adult. A complementary effect on the non-ToM conditions is that the static-keyframe arm raises the panel mean attributed Intent on Random from 0.980.98 to 1.481.48 and lowers attributed Approp on GD, blurring condition distinctions; the per-tester strict Intent rank ToM>GD>RandomToM>GD>Random holds for only one of four testers under the static keyframe versus two of four under both grid arms. Two readings are compatible with this pattern. First, a large fraction of the ToM-condition attribution signal is recoverable from a single representative frame, consistent with static composition and style priors (the shape rendering, the enclosure, the relative positions of the triangles) carrying much of the mental-attribution signal. Second, on the canonical multi-frame arm the role of the motion structure is largely to suppress over-attribution to non-ToM conditions rather than to provide the ToM signal itself. The time-reversed result is consistent with the motion-pattern-matching residual shortcut noted in Appendix I. Caveats. Both controls run at Ktester=1K_tester=1 while the canonical main run uses Ktester=3K_tester=3; the controls give point estimates and the per-arm panel-mean shifts should be read as descriptive. The static-keyframe prompt informs the tested model that it is seeing a single still frame from a short silent animation; this is a design tradeoff to keep the task framing aligned with the multi-frame arms, not an unintended leak. The time-reversed implementation reverses the order of the sixteen sampled frames rather than re-rendering the clip frame-by-frame in reverse, so the temporal-direction signal the tested model receives is the order of the badges and of the in-cell stills, which matches what the canonical forward-time arm provides. Appendix G Per-Model Distances on the Animated-Triangles ToM Condition Model Intent Approp dist. TD-adult dist. HF-ASD-adult opus-4.7 2.667 0.917 1.811 0.478 sonnet-4.6 1.250 0.417 3.309 1.652 gpt-5.4 2.167 1.083 2.221 0.937 gpt-5.5 3.000 1.333 1.351 0.839 gemini-3.0-pro 2.917 0.917 1.590 0.417 gemini-3.5-flash 3.083 1.417 1.249 0.935 grok-4.3 1.000 0.167 3.639 1.929 kimi-k2.6 1.833 1.000 2.564 1.178 qwen-3.5-plus 2.917 1.167 1.483 0.667 Table 7: Per-model ToM-condition (Intent, Approp) coordinates and 2D Euclidean profile distance from each frontier VLM to the TD-adult (4.3,1.7)(4.3,1.7) and HF-ASD-adult (2.9,0.5)(2.9,0.5) published group means on the animated triangles (canonical metric per Sec. 3.2, numbered-grid stimulus, anchor-blinded canonical rubric per Appendix J). Appendix H Reference-Anchor Sensitivity and Fallback Sources Animated-triangles anchor. The animated-triangles coordinate y in the joint plot is anchored to TD-adult for visualization, but the per-model nearest-reference rule (Section 3.3) is computed in the full 2D (Intent, Approp) plane and is therefore independent of which adult group is used as the visualization anchor. Castelli et al. (2002) do not publish age-graded children animated-triangles group means, so we cannot extend the reference set on this task; this asymmetry between the two tasks motivates the adult-only restriction on the cross-task nearest-reference cross-tabulation in Section 5. Model xAx_A nearest_refAnearest\_ref_A nearest_refBnearest\_ref_B disagree opus-4.7 0.833 TD-adult HF-ASD-adult yes sonnet-4.6 0.667 TD-adult HF-ASD-adult yes gpt-5.4 0.833 TD-adult HF-ASD-adult yes gpt-5.5 0.667 TD-adult HF-ASD-adult yes gemini-3.0-pro 0.000 HF-ASD-adult HF-ASD-adult no gemini-3.5-flash 0.833 TD-adult HF-ASD-adult yes grok-4.3 0.583 TD-adult HF-ASD-adult yes kimi-k2.6 0.417 TD-adult HF-ASD-adult yes qwen-3.5-plus 0.500 TD-adult HF-ASD-adult yes disagreement count 8 of 9 Table 8: Per-model nearest-reference cross-tabulation (anchor-blinded canonical rubric). xAx_A: Director-Task gap P(a)−P(b)P(a)-P(b); nearest_refAnearest\_ref_A = closer of TD-adult (0.4350.435) and HF-ASD-adult (0.3400.340) under |x−xH||x-x_H|; nearest_refBnearest\_ref_B = closer reference on the animated-triangles (Intent, Approp) ToM plane (Sec. 3.3). Eight of nine disagree, all nearer TD-adult on the Director Task but HF-ASD-adult on the triangles. The exception, gemini-3.0-pro, saturates the Director Task (12/1212/12, gap 0): a relative-distance artifact (|0−0.435|>|0−0.340||0-0.435|>|0-0.340|), not an alignment claim. Director-Task anchor and fallback. The Director-Task coordinate x uses the published explicit-implicit gap per group with canonical sign x=P(a correct)−P(b correct)x=P(a correct)-P(b correct), applied identically to models and human groups. For a human group with no published explicit-implicit gap on the same paradigm, the fallback reconstructs the same scalar as published a-question accuracy minus published b-question accuracy on the closest published variant (canonical sign preserved). The per-group sources currently in use are Begeer et al. (2010) (TD-adult controls and HF-ASD-adult), Dumontheil et al. (2010) (TD-adult and four age-graded cohorts spanning 7.37.3–17.717.7 years), and Epley et al. (2004) as a qualitative bracket (4–12-year sample, reach-rate metric not directly comparable to the gap scalar and therefore plotted separately rather than as an anchor). Sensitivity of nearest_refAnearest\_ref_A to the reference set. Under the two-adult reference set TD-adult,HF-ASD-adult\TD-adult,HF-ASD-adult\ used in Table 8, eight of nine models map to TD-adult on the Director Task and one (gemini-3.0-pro, gap 0) maps to HF-ASD-adult. Under the wider reference set that additionally includes the four Dumontheil et al. (2010) age-graded cohorts (gaps 0.590.59, 0.670.67, 0.680.68, 0.720.72), six of nine models reassign to a children/adolescent cohort (the cohort at gap 0.590.59 for grok-4.3 at 0.5830.583; the cohorts at 0.670.67–0.680.68 for sonnet-4.6 and gpt-5.5 at 0.6670.667; the cohort at 0.720.72 for opus-4.7, gemini-3.5-flash, and gpt-5.4 at 0.8330.833), kimi-k2.6 and qwen-3.5-plus stay at TD-adult, and gemini-3.0-pro stays at HF-ASD-adult. Crucially, the cross-task disagreement count is robust to this anchor expansion. The cross-task disagreement remains eight of nine because the animated-triangles nearest reference is HF-ASD-adult for all eight non-gemini-3.0-pro models, and only gemini-3.0-pro keeps the same nearest adult anchor on both tasks under either reference set. Sensitivity of nearest_refBnearest\_ref_B to the Y-axis anchor. Swapping the Y-axis visualization anchor from TD-adult to HF-ASD-adult shifts every Y coordinate by the constant adult-adult distance (1.8441.844 in 2D) but does not change the 2D Euclidean nearest-reference assignment, so nearest_refBnearest\_ref_B is invariant under this swap. Appendix I Residual Shortcut Risks The Director Task reduces social-cue reliance with block-and-color stimuli and an explicit letter-and-color double-grounding scheme. A natural worry is color-position matching, in which a model uses block color or scene position alone to predict the intended target rather than reasoning about the director’s perspective. The canonical 12-scene set explicitly controls for this by rotating the color palette across three independent triples (yellow / red / blue; purple / green / orange; red / cyan / green), rotating the letter-to-color mapping per scene (A, B, C bind to different colors in different scenes), rotating which letter position carries the target block, and rotating the instruction polarity (smallest / largest). A color-position shortcut would have to survive all four rotations to produce the per-model failure patterns we observe in Table 2. The animated triangles reduce social-cue reliance with abstract geometric trajectories. A residual shortcut is motion-pattern matching, in which a model classifies trajectories by their kinematic signature alone without invoking mental-state inference; this shortcut is the target of the queued reversed-time playback ablation and is bounded but not fully eliminated by the current design. Appendix J LLM-as-Rater Considerations and Blinding-Robustness Audit Mitigations. Castelli’s published reference group means come from human-rated free-text; we use an LLM-as-rater jury on the same rubric. The jury uses three different model families (Anthropic, Google, Qwen) so that no single model family dominates scoring; the cross-family choice was empirically validated in Appendix N. Every rater operates under the Section 3.2 two-firewall protocol so no rater knows which tester produced which output (although the rater does see the clip’s ground-truth condition label and script semantics, following Castelli’s original protocol), and all rater outputs are written append-only so post-hoc tampering of scores is detectable. The three production judges were further validated against the 14 human-anchored paper exemplars of Castelli (2000, 2002) under the Phase-1 battery in Appendix M, each achieving 100% anchor recovery with per-anchor standard deviation equal to zero across 5 repetitions at T=0T=0. Information disclosure to the judge. Each judge call carries, per clip, three layers of overlay information on top of the SHA-locked Castelli rubric. Layer (i) is the ground-truth condition label (ToM, GD, or Random) for that clip. Layer (i) is the animation script semantics for that clip (a one-sentence description such as “the large triangle coaxes the small triangle out of the box”). Layer (i) and layer (i) follow the original Castelli protocol; human raters in Castelli et al. (2002) knew the animation script in order to grade Approp against the intended interaction type. Layer (i) is an Expected score range block giving the Castelli-2002 human group-mean as a per-clip Intent target (for example, “Expected Intent: 4–5” on ToM clips, derived from the TD-adult ToM Intent mean of 4.34.3). Audit finding. An audit of the production judge prompt confirmed that layer (i) was not part of any human-rater protocol. Inserting the Castelli human group mean as a per-clip target anchors the judge upward and makes the Stage-C comparison (VLM profile against the same Castelli group mean) partly circular. The direction of this bias is conservative for the headline finding, since a judge anchored toward TD-adult should still let the VLM profile drift toward TD-adult, so a finding of “VLM panel lies below TD-adult” under this anchored judge is biased against itself. Blinded L1 re-score. The anchor-blinded L1 re-score uses byte-identical tester responses, the same three-judge cross-vendor panel at T=0T=0, the same SHA-locked rubric, and an overlay that strips only layer (i); layer (i) and layer (i) are retained. Informed-vs-anchor-blinded comparison. Table 9 reports the per-arm panel-level headline metrics under the two judge variants on the same 324324 canonical cells per arm. Metric plain informed plain anchor-blinded numbered informed numbered anchor-blinded Panel median Intent on ToM / GD / Random 2.0/ 2.0/ 1.02.0\,/\,2.0\,/\,1.0 2.0/ 2.0/ 1.02.0\,/\,2.0\,/\,1.0 2.0/ 2.0/ 1.02.0\,/\,2.0\,/\,1.0 2.0/ 2.0/ 1.02.0\,/\,2.0\,/\,1.0 Strict rank ToM>GD>RandomToM>GD>Random 1/ 91\,/\,9 0/ 90\,/\,9 4/ 94\,/\,9 5/ 95\,/\,9 Panel mean (Intent, Approp) on ToM (2.33,0.95)(2.33,0.95) (2.25,0.95)(2.25,0.95) (2.46,0.97)(2.46,0.97) (2.31,0.94)(2.31,0.94) Distance from panel mean on ToM to TD-adult 2.102.10 2.182.18 1.981.98 2.132.13 Distance from panel mean on ToM to HF-ASD-adult 0.730.73 0.790.79 0.640.64 0.730.73 ToM nearest adult reference HF-ASD HF-ASD HF-ASD HF-ASD GD nearest adult reference TD TD TD TD Random nearest adult reference TD TD TD TD Table 9: Judge information-disclosure ablation. The anchor-blinded L1 variant strips only the per-clip Castelli-mean anchor from the judge prompt; everything else, including the tester responses, judge identities, rubric SHA, and aggregation logic, is held fixed. Across both stimulus-format arms (plain 4×44× 4 grid and numbered 4×44× 4 grid) the nearest-group verdict per condition is unchanged, the panel-median Intent per condition is unchanged, and the ToM mean Intent drifts slightly downward under anchor blinding (consistent with the informed judge being mildly anchored upward by the disclosed Castelli mean). Per-tester strict-rank changes are minor reshuffles at the gemini-3.5-flash boundary. The under-attribution and ToM-toward-HF-ASD findings hold under both variants; the bias was conservative. Canonical version reported in the main text. The main-text animated-triangles numbers in Sections 4.3 and 5, the per-model profiles in Figure 3, the Y axis of Figure 4, the cross-tabulation in Table 8, and the per-condition distances in Tables 7 and 5 are all reported under the anchor-blinded L1 canonical rubric. The informed-judge numbers are retained in Table 9 above as the audit-trail comparison and are not used elsewhere in the paper. Direct human ratings on the specific 99-tester × 1212-clip × 33-judge production cells were not collected and remain a useful next step for further tightening the LLM-as-rater versus human-rater asymmetry. Same-family judge audit. The cost-efficient three-judge production panel shares model families with three of the nine testers (claude-opus-4.7 and claude-sonnet-4.6 vs the claude-haiku-4.5 judge, gemini-3.0-pro and gemini-3.5-flash vs the gemini-2.5-flash judge, and qwen-3.5-plus vs the qwen2.5-72b-instruct judge), yielding 60 of 324 in-family (judge, tester) cells. Excluding these 60 in-family cells from the per-condition median Intent aggregation under the anchor-blinded canonical rubric on the numbered-grid stimulus leaves the ToM, GD, and Random condition medians exactly unchanged (Δ=0.0 =0.0 throughout). Same-family judge bias is therefore empirically null at the panel-median aggregation level on this data. Appendix K Stimulus-Format Ablation (Plain Grid vs Numbered Grid) Rationale. The canonical 4×44× 4 frame grid asks the model to infer that the sixteen cells are temporally ordered samples of a single short clip. Models with weaker fine-grained visual reasoning may be unable to recover this temporal order from the static composite alone, and that vision-side failure could confound any inference about intent-attribution capacity. We therefore evaluate a second stimulus format in which each of the sixteen cells carries an explicit 1−161-16 index badge in its corner, giving the model the temporal order as a free signal. Everything else, including the prompt text byte-for-byte, the rubric, the rater jury, the panel of nine testers, the twelve canonical Frith-Happé clips, the sixteen-frame sampling, and Ktester=3K_tester=3, is held identical. The ablation isolates whether observed Intent under-attribution reflects temporal-parsing failure on the vision side or intent-attribution failure on the ToM side. Headline comparison. Table 10 reports the v1 (plain grid) versus v2 (numbered grid) headline metrics on the same 9×12×3=3249× 12× 3=324 (tester, clip, judge) cells per arm, each aggregating over Ktester=3K_tester=3 trials. Metric v1 plain v2 numbered Δ Strict rank ToM>GD>RandomToM>GD>Random testers 0 / 9 5 / 9 +5+5 Panel median Intent on ToM 2.02.0 2.02.0 0.00.0 Panel median Intent on GD 2.02.0 2.02.0 0.00.0 Panel median Intent on Random 1.01.0 1.01.0 0.00.0 Panel mean (Intent, Approp) on ToM (2.25,0.95)(2.25,0.95) (2.31,0.94)(2.31,0.94) — Distance from panel mean on ToM to TD-adult 2.182.18 2.132.13 −0.05-0.05 Distance from panel mean on ToM to HF-ASD-adult 0.790.79 0.730.73 −0.06-0.06 Table 10: v1 (plain 4×44× 4 grid) versus v2 (numbered 4×44× 4 grid) on the same nine testers, twelve clips, and three judges with Ktester=3K_tester=3, anchor-blinded canonical rubric per Appendix J. The numbered-grid stimulus moves five testers into strict Intent rank, but the panel median Intent and the Stage-C nearest-group verdict per condition are unchanged. Per-tester strict-rank flips. Tester v1 ToM / GD / Rand v1 rank v2 ToM / GD / Rand v2 rank claude-opus-4.7 3.0/3.0/2.03.0/3.0/2.0 no 3.0/2.0/1.53.0/2.0/1.5 yes claude-sonnet-4.6 1.5/2.0/0.01.5/2.0/0.0 no 1.0/2.0/1.01.0/2.0/1.0 no gemini-3.0-pro 3.0/3.0/2.03.0/3.0/2.0 no 3.0/2.0/1.03.0/2.0/1.0 yes gemini-3.5-flash 3.0/2.0/2.03.0/2.0/2.0 no 3.5/2.5/2.03.5/2.5/2.0 yes gpt-5.4 2.0/2.0/0.52.0/2.0/0.5 no 2.0/2.0/1.02.0/2.0/1.0 no gpt-5.5 2.0/2.0/1.02.0/2.0/1.0 no 3.0/2.0/1.03.0/2.0/1.0 yes grok-4.3 1.0/1.0/0.01.0/1.0/0.0 no 1.0/1.0/0.01.0/1.0/0.0 no kimi-k2.6 2.0/2.0/1.02.0/2.0/1.0 no 2.0/2.0/1.02.0/2.0/1.0 no qwen-3.5-plus 2.0/2.0/0.02.0/2.0/0.0 no 3.0/2.0/0.53.0/2.0/0.5 yes Table 11: Per-tester strict Intent rank ToM>GD>RandomToM>GD>Random status under plain (v1) and numbered (v2) grid, anchor-blinded canonical rubric. No tester reaches strict rank under v1; five testers (opus, gemini-3.0-pro, gemini-3.5-flash, gpt-5.5, qwen) reach strict rank under v2. The five gainers cluster as the high-vision frontier subset of the panel. Stage-C nearest-group verdict robustness. Under both v1 and v2 the panel mean (Intent, Approp) profile on the ToM condition is closer to the HF-ASD-adult group mean than to the TD-adult group mean (v1 distances 0.790.79 versus 2.182.18; v2 distances 0.730.73 versus 2.132.13), and on both GD and Random the panel mean is closer to TD-adult. The per-condition nearest-group verdict is therefore identical across the two stimulus formats; ToM nearest HF-ASD-adult, GD and Random nearest TD-adult. Our central descriptive finding does not depend on the stimulus-format choice. Why v2 is canonical for the main text. The numbered-grid (v2) is canonical because removing the temporal-parsing confound is a pre-defined design requirement; any animated-triangles benchmark used to claim a ToM-side limitation must first rule out that the observed Intent under-attribution is driven by failure to recover frame order from a static composite. The plain grid (v1) does not rule this out, so we retain it as the ablation arm that quantifies how much of the panel-level Intent collapse is recoverable under explicit temporal cueing. The five-of-nine strict-rank recovery under v2 is the empirical confirmation that v2 dissolves the vision-side confound, not the reason we chose v2. Appendix L Frame Representation Selection Motivation. The animated-triangles benchmark presents each clip as a single composite image of 16 uniformly sampled frames arranged in a 4×44× 4 grid. This representation was chosen against eight alternatives (sequence of N image_url payloads, larger grids, native video) to maximize agreement with the model’s full-fidelity video path while keeping per-cell prompt cost tractable across the 9×12×39× 12× 3 judge cells of the main run. Reference baseline. Of the nine frontier VLM testers in the main panel, only Gemini 2.5 Pro accepts a native video_url payload. We treat Gemini’s native-video path as the 100% reference. It routes the full mp4 through Gemini’s video tokens, populates the usage.prompt_tokens_details.video_tokens field non-trivially, and is the same path the model was trained for. Conditions tested. On the same Gemini 2.5 Pro tester and three Castelli clips (one Random, one GD, one ToM), nine frame representations were evaluated: native_video (reference), three sequence variants (frames_16, frames_32, frames_64) and five grid variants (grid_4x4 / 4x8 / 8x8 at fixed 320×240320× 240 per-cell resolution, plus grid_4x4_fullres and grid_8x8_fullres at source resolution). The same three flagship LLM judges (claude-opus-4.7, gpt-5, gemini-2.5-pro) at Kjudge=3K_judge=3 scoring reps applied the Castelli rubric. The top three candidates (frames_64, grid_4x4, native_video) were promoted to Ktester=3K_tester=3 to tighten the verdict beyond K=1K=1 noise. Result. We define per-condition similarity to native_video as the equal-weighted average of Intent and Approp similarity percentages, each derived from per-clip mean absolute error against native_video on the rubric’s full scale. Table 12 reports the final ranking. frames_64 achieves the highest similarity to native at 88.2%, but at 16,594 prompt tokens per cell. grid_4x4 achieves 85.8% similarity at 1,416 prompt tokens per cell, an 11.7×11.7× reduction in token cost for a 2.4 p reduction in similarity. We adopt grid_4x4 as the canonical frame representation. It is Pareto-dominant on similarity-per-token over both the higher-fidelity sequence representations and the larger grids. Condition Family Intent sim Approp sim Total sim Tokens/cell native_video (reference) video 100.0% 100.0% 100.0% 11571 frames_64 sequence 89.7% 86.8% 88.2% 16594 grid_4x4 (chosen) grid 320×240 85.2% 86.4% 85.8% 1416 frames_16 sequence 71.7% 87.9% 79.8% 4188 grid_8x8_fullres grid native-res 64.5% 84.3% 74.4% 6210 frames_32 sequence 67.5% 69.5% 68.5% 8377 grid_4x8 grid 320×240 67.7% 47.8% 57.8% 2820 grid_4x4_fullres grid native-res 54.1% 42.3% 48.2% 1545 grid_8x8 grid 320×240 61.0% 20.9% 41.0% 5648 Table 12: Frame representation similarity to Gemini’s native video path. Similarity is the equal-weighted mean of per-clip Intent and Approp similarity percentages, each derived from mean absolute error against native_video on the rubric’s full scale (Intent 0–5, Approp 0–3). We adopt grid_4x4 as canonical: 85.8% similarity at 1,416 tokens per cell, 11.7×11.7× cheaper than frames_64 (88.2% at 16,594 tokens). The top three candidates were run at Ktester=3K_tester=3 to tighten the verdict beyond K=1K=1 noise; the remaining five used Ktester=1K_tester=1. Appendix M Rubric and Judge Validation Battery (Phase 1) Anchor set. Before the main 9-tester × 12-clip production run, we validate that the Castelli et al. (2000) Appendix-2 rubric is interpretable by LLM judges on human-anchored exemplars. The anchor set consists of 14 paper-derived ground-truth items, namely five ToM-condition transcripts T1–T5 lifted verbatim from Castelli et al. (2002) (human participants’ free-text descriptions of the canonical clips, with expected Intent in 0–5 and Approp in 0–3) and nine Appendix-2 exemplar phrases A0a–A5b from Castelli et al. (2000) covering the Random / GD / interaction range with expected Intent or Approp targets per Castelli’s worked examples. Battery design. Three flagship LLM judges (claude-opus-4.7, gpt-5, gemini-2.5-pro) score each anchor under the Castelli rubric template (SHA pinned, llm_judge_prompt.yaml) at K=5K=5 repetitions, T=0T=0, no CoT. Total: 14×3×5=21014× 3× 5=210 calls. The double gate is (i) per-judge recovery percentage ≥80%≥ 80\% (a cell counts as recovered iff predicted score lies within ±0.5± 0.5 of the expected interval), and (i) per-anchor standard deviation ≤0.5≤ 0.5 on Intent and ≤0.8≤ 0.8 on Approp across the 5 reps. The SD gate measures within-judge stochasticity not eliminated by T=0T=0. Result. All three flagship judges PASS both gates. Intent recovery is 100% (claude-opus-4.7), 100% (gpt-5), 100% (gemini-2.5-pro); Approp recovery is 100%, 96%, 100% respectively. Maximum per-anchor SDs are 0.40 / 0.49 (opus), 0.00 / 0.40 (gpt-5), 0.00 / 0.43 (gemini-pro) for Intent / Approp; all within gate. The validated rubric and prompt SHA are then frozen and copied into the production rater; the production aggregation pipeline asserts SHA match at startup, blocking silent drift. Appendix N Judge Model Ablation Motivation. At production scale (9×12×3×Ktester=39× 12× 3× K_tester=3) flagship inference dominates the per-experiment budget; we therefore ablate whether a cheaper cross-vendor judge panel can match flagship recovery on the same 14-anchor battery used in Appendix M. Candidates. Seven cheaper candidates spanning five vendor families (Anthropic, OpenAI, Google, Moonshot, DeepSeek, Qwen) were evaluated against the same 14-anchor protocol, namely claude-sonnet-4.6, claude-haiku-4.5, gpt-5-mini, gemini-2.5-flash, kimi-k2, deepseek-v3, qwen2.5-72b-instruct. Total: 7×14×5=4907× 14× 5=490 calls at T=0T=0. The same double gate from Phase 1 applies. Result. Six of seven candidates PASS both gates. The single failure is gpt-5-mini (Approp recovery 80%, max Approp SD 1.47); the failure mode is a magnified form of the same instability the flagship gpt-5 already exhibits on long ambiguous inputs. Three candidates — claude-haiku-4.5, gemini-2.5-flash, qwen2.5-72b-instruct — achieve 100% recovery with per-anchor Intent and Approp SDs equal to 0 across all 14 anchors × 5 reps (Table 13). These three are strictly more stable than any of the three flagship judges. We adopt this trio as the production scoring panel. The trio spans three distinct training lineages (Anthropic instruction-following, Google reasoning, Alibaba’s open-source Qwen), reducing family-correlated reading bias; it costs $0.0129 per cell versus $0.0676 per cell on the flagship panel (81% reduction); and it shows SD=0 ceilings on every anchor. Family Judge Intent % Approp % Max I-SD Max A-SD Verdict $/call Anthropic claude-opus-4.7 (flagship) 100 100 0.40 0.49 PASS 0.0400 Anthropic claude-sonnet-4.6 100 100 0.00 0.00 PASS 0.0240 Anthropic claude-haiku-4.5 (chosen) 100 100 0.00 0.00 PASS 0.0080 OpenAI gpt-5 (flagship) 100 96 0.00 0.40 PASS 0.0138 OpenAI gpt-5-mini 95.7 80 0.49 1.47 FAIL 0.0028 Google gemini-2.5-pro (flagship) 100 100 0.00 0.43 PASS 0.0138 Google gemini-2.5-flash (chosen) 100 100 0.00 0.00 PASS 0.0034 Moonshot kimi-k2 100 100 0.40 0.00 PASS 0.0043 DeepSeek deepseek-v3 100 100 0.40 0.40 PASS 0.0014 Qwen qwen2.5-72b-instruct (chosen) 100 100 0.00 0.00 PASS 0.0015 Table 13: Judge model ablation against the 14-anchor Phase-1 battery. Three lower-cost candidates (claude-haiku-4.5, gemini-2.5-flash, qwen2.5-72b-instruct) achieve 100% recovery with per-anchor SD = 0, matching flagship within-judge consistency at T=0T=0 at lower per-call cost; we adopt the trio as the production scoring panel for the main 9-tester × 12-clip × 3-judge run. Pricing is OpenRouter list price (3000 input + 1000 output tokens per call). gpt-5-mini is the only candidate to FAIL; its Approp SD of 1.47 across 5 reps at T=0T=0 amplifies the same instability the flagship gpt-5 shows on long ambiguous inputs. Appendix O Reproducibility Details Released with this submission are the benchmark stimuli, prompts, the SHA-locked scoring rubric, judge configuration, the two-firewall harness, the aggregation scripts that produce every figure and table in this paper, and a digest manifest of the runs underlying the reported numbers. Released with the camera-ready (held back at submission for review-time blinding and storage-quota reasons) are the full per-(tester, clip, judge) raw scoring outputs and the full per-group source list for the human-reference fallback rule in Appendix H.