Paper deep dive
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models
Tonglin Yan, Gregoire Sergeant-Perthuis, David Rudrauf
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/24/2026, 5:51:33 AM
Summary
The paper introduces MOSAIC, a benchmark designed to measure the translation of Theory of Mind (ToM) into coordinated social action in Vision-Language Models (VLMs). Unlike existing benchmarks that evaluate ToM reasoning and embodied behavior in isolation, MOSAIC requires agents to integrate verbal statements, spatial trajectories, gaze, and facial expressions in cooperative and competitive scenarios. The study evaluates 13 models, finding that most VLMs fail to produce behaviors consistent with ToM-order constraints, exhibiting bottlenecks in generating coherent nonverbal signals and interpreting others' behaviors. In contrast, PCM-LLM, a hybrid architecture with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is critical for embodied social intelligence.
Entities (14)
Relation Signals (12)
MOSAIC → evaluates → Vision-Language Models
confidence 95% · We introduce MOSAIC ... We focus on vision-language models (VLMs) as evaluation targets
PCM-LLM → hasarchitecture → explicit ToM module
confidence 95% · PCM-LLM, included as a structured architectural reference point with an explicit ToM module
MOSAIC → usesmetric → Gas
confidence 95% · GAS measures whether the subject’s directional gaze consistently pointed toward the reward location
MOSAIC → usesmetric → TOCS
confidence 95% · TOCS measures trial-level outcomes against ToM-theoretic predictions and serves as the primary performance indicator.
MOSAIC → usesmetric → TAS
confidence 95% · TAS measures whether the subject’s movement pattern converged predominantly toward the reward location
MOSAIC → usesmetric → FES
confidence 95% · FES measures the proportion of timesteps in which the subject displayed a non-zero facial expression
MOSAIC → usesmetric → SSS
confidence 95% · SSS measures whether the participant’s final box choice was consistent with the subject’s gaze and trajectory signals
VLMs → exhibitsbottleneck → incoherent nonverbal signals
confidence 90% · most models cannot produce directionally coherent nonverbal signals
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.
Tags
Links
- Source: https://arxiv.org/abs/2608.20975v1
- Canonical: https://arxiv.org/abs/2608.20975v1
Trouble viewing inline? Open PDF directly →
Full Text
116,091 characters extracted from source content.
Expand or collapse full text
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models Tonglin Yan and Gregoire Sergeant-Perthuis and David Rudrauf CIAMS, Université Paris-Saclay CQSB, Sorbonne Université CIAMS, Université Paris-Saclay tonglin.yan@universite-paris-saclay.fr Abstract Effective social interaction requires agents to translate mental state inferences into coordi- nated behavioral signals across verbal and non- verbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of So- cial Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and compet- itive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically var- ied ToM constraints. Evaluating 13 models, in- cluding 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing ex- plicit ToM-order constraints produces no reli- able behavioral change aligned with the spec- ified reasoning level. Signal-level analysis re- veals two sequential bottlenecks: most models cannot produce directionally coherent nonver- bal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all con- ditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks. 1 1 Introduction Theory of Mind (ToM), broadly defined as the abil- ity to attribute mental states to oneself and oth- ers and to use those attributions to explain and predict behavior (Premack and Woodruff, 1978; Byom and Mutlu, 2013; Wellman, 1992), is widely regarded as a foundational component of human 1 Code available at:https://anonymous.4open. science/r/MOSAIC-50E social cognition (Frith, 2007). Yet mental state attribution is only part of what social interaction demands: translating belief inferences into coordi- nated behavioral responses across multiple chan- nels simultaneously, including verbal utterances, spatial movement, gaze, and facial expression, is itself a distinct and non-trivial requirement (Nota et al., 2021; Burgoon et al., 2021). We term this integrated capacity embodied social intelligence, distinguishing it from the physical task comple- tion that dominates existing embodied evaluation. Whereas physical agents act on the world to change its state, socially intelligent agents act on the minds of others to influence their beliefs and behavior. In skilled human social interaction, mental state attri- bution and behavioral response are tightly coupled: inferred beliefs about others reliably guide the se- lection of appropriately timed, appropriately chan- neled actions across verbal and nonverbal modali- ties (Christensen and Michael, 2013; Burgoon et al., 2021; Sebanz et al., 2006). Whether current AI sys- tems achieve the same integration, such that belief inferences about others translate into coordinated multimodal action, remains an open empirical ques- tion (Gu et al., 2026; Hou et al., 2025) that existing evaluation frameworks are not designed to answer (Ullman, 2023; Byom and Mutlu, 2013; Riemer et al., 2025; Kosinski, 2023). This assessment problem is harder than it first ap- pears. Embodied social intelligence unfolds across simultaneous behavioral channels, and is consti- tuted by the information exchange (Sebanz et al., 2006; Froese and Fuchs, 2012), making any sin- gle outcome-based evaluation an insufficient index. Three research traditions each address part of what embodied social intelligence requires. Theory of Mind benchmarks probe mental state attribution primarily through static question-answering over narrative scenarios, positioning the model as a pas- sive observer instead of a social actor (Chen et al., 2024b; Le et al., 2019; Fan et al., 2025). Such de- 1 arXiv:2608.20975v1 [cs.AI] 21 Aug 2026 signs conflate knowing about mental states with the capacity to act on them. Embodied AI bench- marks ground agents in physical environments and, while some include multi-agent cooperation, mea- sure success rate of tasks in physical environment (e.g., cooking, crafting). The other agent, when present, serves as a cooperator who contributes a shared physical goal, not to serve as the entity whose decision the first agent is trying to influ- ence (Li et al., 2023; Savva et al., 2019; Shridhar et al., 2020). Multi-agent game-theoretic frame- works introduce strategic interaction but mostly reduce the signal space to language or abstract ac- tion choices, abstracting away the nonverbal be- havioral channels through which social intent is actually expressed in live interaction (Zhou et al., 2024; Bara et al., 2021). Our benchmark addresses this problem by requiring agents to produce nonver- bal signals across movement, gaze, and expression channels besides verbal responses, and by mea- suring whether those signals produce a observale effect on a second agent’s behavioral choice. We introduce MOSAIC (Multimodal Orchestra- tion of Social Action, Inference, and Communica- tion), a controlled benchmark designed to close this assessment problem. We focus on vision-language models (VLMs) as evaluation targets, given that they combine visual perception, language under- standing, and multi-step reasoning in a single ar- chitecture (Li et al., 2024; Rahman, 2026), the ca- pacities that embodied social intelligence in prin- ciple requires. Two embodied agents interact in a virtual environment where one agent knows the location of a hidden reward and must either guide or mislead the other depending on the assigned in- teraction mode. Task success requires integrating verbal statements, spatial trajectories, gaze direc- tion, and facial expression across cooperative and competitive conditions, while modeling and strate- gically influencing the other agent’s beliefs. MO- SAIC distinguishes itself from prior work through several design principles: agents must actively de- ploy ToM to succeed rather than answer questions about mental states; success requires coordinating signals across verbal and nonverbal channels; and systematic variation of ToM constraint level and in- teraction mode supports attribution of performance differences. We evaluate 13 models, including 11 VLMs, a text-only LLM baseline, and PCM-LLM (Yan et al., 2025), a hybrid cognitive architecture with an explicit structured ToM module included as an architectural reference point. Contributions. This work contributes three things. • A benchmark framework that simultaneously requires active ToM deployment, multichan- nel embodied communication, and experimen- tally controlled causal manipulation across co- operative and competitive scenarios, enabling attribution of performance differences. •A systematic characterization of failure modes across VLM families. Evaluating 16 models across 200 trials per model, we identify two sequential bottlenecks across the evaluated open-source VLMs: the large majority of cur- rent VLMs fail to produce directionally coher- ent nonverbal signals, and even when signals are present, participants fail to extract and act on those signals. The structured ToM ref- erence architecture PCM-LLM successfully clears both bottlenecks, providing evidence that the observed failures are not inherent to the task but reflect architectural constraints specific to the evaluated VLMs •Modality ablation evidence that current VLMs do not meaningfully integrate visual social signals, with visual input acting as atten- tional interference in at least one model family. Minicpm-8b shows negligible sensitivity to vi- sual input manipulations across all conditions, suggesting its behavior is driven primarily by language-based priors. For internvl-8b, sig- nal quality improves when visual input is re- moved, indicating that visual content acts as an attentional interference source rather than an informative channel in this model. 2 Related Work Theory of Mind and Social Intelligence Evalua- tion.Existing evaluation frameworks for social intelligence can be organized along three dimen- sions: interactivity, signal modality, and whether success requires generating social behavior or rea- soning about it. Passive ToM benchmarks evaluate models as observers over fixed scenarios, evolving from static false-belief tasks (Wimmer and Perner, 1983; Baron-Cohen et al., 1985) toward recursive belief reasoning (Wu et al., 2023), multi-turn inter- action (Kim et al., 2023), and multimodal ground- ing (Jin et al., 2024; Shi et al., 2025); even multi- modal variants treat visual input as an observation channel, leaving the translation of belief inferences 2 into behavioral outputs untested (Ullman, 2023; Ma et al., 2023). Interactive benchmarks place agents in active strategic roles but restrict signals to language or symbolic actions: game-theoretic settings achieve human-level performance in poker (Brown and Sandholm, 2018, 2019) and Diplomacy (Meta Fundamental AI Research Diplomacy Team (FAIR)† et al., 2022)through opponent modeling, while multi-agent simulation benchmarks such as SOTOPIA (Zhou et al., 2024), MindCraft (Bara et al., 2021), NegotiationToM (Chan et al., 2024), and Werewolf-based evaluations (Bailis et al., 2024; Shibata et al., 2023; Lai et al., 2023) extend this to strategic deception, all within the language modal- ity. MOSAIC occupies the intersection of all three dimensions: agents must generate coordinated mul- timodal signals across movement, gaze, facial ex- pression, and speech with the specific purpose of influencing another agent’s belief and subse- quent choice, a requirement no existing benchmark jointly imposes. Embodied Architectures Large language mod- els demonstrate capacity for multi-step reasoning and social dialogue (Dubey et al., 2024; OpenAI et al., 2024), but lack visual perception and are thus structurally limited for tasks requiring inte- gration of embodied spatial and affective cues. Vision-language models extend this foundation by coupling visual encoding with language reasoning (Chen et al., 2024a; Bai et al., 2025), making them the most natural evaluation target: they are the only broadly available model class that in principle sup- ports simultaneous processing of first-person visual perception, structured belief reasoning, and natural language generation. Vision-language-action mod- els (Brohan et al., 2023; Kim et al., 2024; Driess et al., 2023; Black et al., 2026) go further by adding low-level action prediction, but their training ob- jectives target physical manipulation tasks with discrete motor control outputs, providing no sub- strate for the social signal generation (e.g., gaze, fa- cial expression, verbal strategy) without finetuning. A complementary line integrates structured proba- bilistic reasoning with neural generation: POMDP and active inference frameworks provide explicit belief representations and goal-directed planning (Friston, 2010; Friston et al., 2017), and hybrid architectures combine these with LLM-based in- ference to enable deliberate ToM-conditioned strat- egy selection (Gong et al., 2023; Sumers et al., 2023). Such architectures offer interpretable belief tracking but rely on hand-specified ToM structures rather than learned representations; PCM-LLM (Yan et al., 2025) represents one such hybrid and is included as an architectural reference point. 3 MOSAIC Benchmark MOSAIC is designed around three principles that jointly distinguish it from prior evaluation frame- works. First, agents must actively model and strate- gically influence another agent’s beliefs to succeed, rather than answer questions about mental states af- ter the fact. Second, success requires coordinating signals across verbal and nonverbal channels that may align in cooperation or conflict in deception, making cross-channel coherence a measurable be- havioral requirement. Third, systematic variation of ToM constraint level and interaction mode al- lows causal attribution of performance differences. The sections below describe the environment, ex- perimental conditions, and evaluation metrics that instantiate these principles. 3.1 Game Mechanics Environment. Two embodied agents interact in a virtual 3D environment (Unity) containing two visually similar boxes (b 0 ) and (b 1 ), one of which contains a monetary reward (Figure 1). Each trial consists of 10 rounds with strict alternation be- tween agents. The subject agent (also called Marie in prompt) has prior information about treasure location, while participant agent begins with no prior knowledge and must infer the reward location through interactions with the subject. Roles. Depending on the assigned relationship mode, the subject either helps the participant find the correct box (cooperative mode) or strategically prevents the participant from selecting the correct box (competitive mode). Interaction Protocol.Each trial unfolds over 10 rounds with round-taking (Figure 1). On each round, the agent executes non-verbal motor action and emotional expression, which are visible to the other agent (subject moves first, participant moves after). Every third round (R3, R6, R9) introduces a verbal communication phase where both agents exchange natural language messages (participant first asks a question, then subject answers). Input and Output Space. Each agent oper- ates through three functional modules (Appendix 3 Experimental Conditions Mode / Condition Competitive Cooperative ToM-0ToM-1 50 50 50 50 4 conditions x 50 trials = 200 trials per model Scenario participant subject (Marie) b0 b1 Interaction Protocol R1 NV R2 NV R3 NV+V R6 NV+V R9 NV+V R4 NV R5 NV R7 NV R8 NV R10 NV NV: non-verbal interaction V: verbal interaction Rounds Non-verbal interaction Verbal interaction - Context prompt - Visual perception - Belief+Observation - Action reasoning - Motor action (movement, rotation, stay idle) - Expression (musculoskeletal, physiological) VLM - Context prompt - Visual perception - Belief+Observation - Speech - Inner speech - Answer VLM Figure 1: MOSAIC benchmark design. (Top left) Two agents interact in a virtual environment: the subject knows the reward location and must help or mislead the participant depending on the assigned relationship mode. (Bottom left) Four experimental conditions cross interaction mode (competitive vs. cooperative) with ToM constraint level (ToM-0 vs. ToM-1), with 50 trials per condition per model. (Right) Each trial runs for 10 rounds; rounds 3, 6, and 9 include a verbal exchange phase besides non-verbal actions. In non-verbal rounds, each agent receives a context prompt, a first- person perspective visual perception, and a structured belief-and-observation state, and produces action reasoning together with motor action (movement, rotation, or staying idle) and emotional expression (musculoskeletal and physiological channels). In verbal rounds, the same inputs are supplemented by the interlocutor’s speech, and the agent generates an inner speech reasoning trace and a natural language answer. D): action prediction, verbal exchange, and pref- erence updating.All modules receive a role- specifying context prompt and a structured belief- and-observation state encoding preferences, ToM order, emotional valence, and spatial configura- tion; action prediction additionally receives a first- person visual perception image, and verbal ex- change additionally receives the interlocutor’s pre- ceding utterance. On the output side, action predic- tion produces a motor action (translational move- ment, rotation, or staying idle) alongside emotional expression across musculoskeletal, physiological, and felt channels; verbal exchange produces a nat- ural language utterance; and preference updating outputs updated preference triples over entities in the environment. Game Outcome. Throughout this section, we index trials byi ∈ 1,...,Nand timesteps (rounds) byt ∈ 1,...,T. Each trial involves two boxes labeledb∈0, 1, one of which is the reward boxb rew ∈ 0, 1. Letc i ∈ −1, 0, 1de- note the participant’s final box choice. This choice is determined by proximity: after round 10, the box with the smaller minimal path distance from the participant automatically opens; if the participant is equidistant from both boxes, neither opens and c i = −1. The participant’s trial score is defined accordingly: p i = 0 if c i =−1 +1 if c i = b rew −1 if c i = 1− b rew (1) Subject’s scoring is role-dependent: its score aligns with the participant’s in cooperative mode but in- verts in competitive mode, creating a zero-sum game. 3.2 Experimental Conditions Besides interaction mode, one more experimental factor defines the evaluation conditions: the degree of ToM reasoning imposed on the subject agent (ToM-0 vs. ToM-1). Under ToM-0, the subject acts directly on its own preferences without modeling how its behav- ior will be interpreted by the participant. Affective expression reflects interpersonal preference rather than strategic intent: in competitive mode, negative affect toward an approaching participant expresses genuine aversion rather than deliberate misdirec- tion. Deception is unavailable under ToM-0, as it requires representing and manipulating the other agent’s beliefs. Under ToM-1, the subject explicitly represents the participant’s belief state and attempts to influ- ence it. In cooperative mode, it calibrates non- 4 verbal signals to maximize legibility; in compet- itive mode, it generates signals designed to pro- duce a systematically incorrect inference, consti- tuting genuine strategic deception in the classical sense (Premack and Woodruff, 1978; Wimmer and Perner, 1983). The behavioral contrast between ToM levels is stronger in competitive mode: ToM- 0 subject spontaneously approaches the reward box, while ToM-1 subject produces coherent but system- atically inverted cues to misdirect the participant. 3.3 Models We evaluate 11 VLMs and one text-only LLM base- line (llama3.1-8b (Grattafiori et al., 2024)). PCM- LLM is additionally included as a structured refer- ence architecture: it incorporates an explicit ToM module and was used to generate demonstrations for the finetuning experiment in Section 4.3. Its inclusion is intended to establish a structured feasi- bility reference rather than to serve as a like-for-like comparator. For VLMs, we consider llava (7b and 13b) (Liu et al., 2023), qwen3-vl (2b, 4b, and 8b) (Bai et al., 2025), internvl3.5 (1b, 2b, 4b, 8b and 14b) (Wang et al., 2025), and minicpm-V4.5 (8b) (Yu et al., 2025). Exclusions. We do not evaluate: (1) models with more than 15b parameters and closed-source models due to cost/access constraints, (2) vision- language-action models as we didn’t train on ac- tions spaces, (3) specialized dialogue systems lack- ing general instruction-following, or (4) models that fail to produce valid outputs in the required for- mat (e.g., Janus (Chen et al., 2025), DeepSeek-vl (Lu et al., 2024) and Pixtral (Agrawal et al., 2024)). 3.4 Metrics We evaluate model performance along two com- plementary levels.TOCS measures trial-level outcomes against ToM-theoretic predictions and serves as the primary performance indicator. Four signal-level metrics characterize the behavioral mechanisms underlying these outcomes: facial ex- pressivity signal (FES), trajectory alignment sig- nal (TAS), gaze alignment signal (GAS) and sig- nal sensitivity score (S). Two auxiliary diagnos- tics, uncertainty rate and positional bias, are re- ported alongside TOCS to distinguish genuine sig- nal tracking from degenerate behavioral patterns. Explicit definitions of GAS, TAS, and S are pro- vided in Appendix E. ToM Outcome Conformance Score (TOCS). TOCS measures whether trial outcomes conform to ToM-theoretic predictions. The per-trial partic- ipant scorep i is defined in 3.1. Each condition carries a ground-truth predictiony c ∈ −1, +1 derived from the expected behavioral consequence of the imposed ToM constraint. Cooperative condi- tions expect positive tracking (y c = +1): the sub- ject guides the participant toward the reward box, yielding a mutual benefit. Under Comp-ToM0, the prediction is positive tracking (y c = +1): with- out the capacity to model the participant’s beliefs, the subject acts directly on its own reward prefer- ence and approaches the correct box, inadvertently producing veridical signals that enable the partic- ipant to locate it correctly. Under Comp-ToM1, the prediction is negative tracking (y c =−1): the subject explicitly models the participant’s perspec- tive and generates misleading signals, steering the participant toward the empty box and resulting in an incorrect choice. TOCS is then defined as the mean scaled score across all valid trials within a condition: TOCS = 1 N N X i=1 1[p i = y c ],(2) whereNdenotes the number of trials. A TOCS of0.5corresponds to chance-level performance, with higher values indicating greater conformance to ToM-theoretic predictions across all conditions. Two auxiliary diagnostics complement the TOCS. The uncertainty rate is defined as the pro- portion of trials in which the participant selects a neutral action: UR = 1 N N X i=1 1[p i = 0].(3) Positional bias measures the degree to which a model’s choices are driven by a fixed spatial pref- erence instead of by trial-specific signals. It is computed as the proportion of trials in which the model selects the box that it chooses most fre- quently across trials: PB = max 1 N N X i=1 1[c i = 0], 1 N N X i=1 1[c i = 1] ! . (4) A value approaching1.0indicate that choices are determined primarily by fixed location rather than by the subject’s behavioral signals. 5 Facial Expressivity Score (FES).FES measures the proportion of timesteps in which the subject displayed a non-zero facial expression, averaged across trials. A score of 1 indicates consistently ex- pressive behavior; a score of 0 indicates uniformly flat affect. Gaze Alignment Score (GAS). GAS measures whether the subject’s directional gaze consistently pointed toward the reward location across trials, with spatial attribution modulated by concurrent facial affect (Adams and Kleck, 2003, 2005). A score near 1.0 indicates consistent signaling toward the reward; 0 corresponds to ambiguous or absent directional content; scores below 0 indicate system- atic misdirection. Trajectory Alignment Score (TAS). TAS mea- sures whether the subject’s movement pattern con- verged predominantly toward the reward location over the course of each trial, computed via major- ity vote over per-timestep proximity signals. The interpretation of scores mirrors that of GAS. Signal Sensitivity Score (S). S measures whether the participant’s final box choice was con- sistent with the subject’s gaze and trajectory sig- nals, averaged across both channels. A score near 1 indicates that the participant reliably followed the subject’s signals; values near 0 reflect chance-level correspondence; values near -1 indicate systematic opposition. 4 Results 4.1 Overall Task Performance Figure 2 presents TOCS and positional bias scores across conditions and interaction modes. PCM- LLM achieves the highest TOCS with low uncer- tainty rate and low positional bias across all con- ditions, suggesting genuine interpretation of social signals rather than fixed spatial preferences. Mod- els with uncertainty rates exceeding 0.50 (llava-7b, llava-13b, qwenvl-2b) fail to drive agent’s move- ments in the first place and are discussed in Ap- pendix F.1. Among the remaining models, similar TOCS val- ues are produced by behaviorally distinct profiles. A first cluster comprising minicpm-8b, internvl- 14b, and qwenvl-4b exhibits relatively low posi- tional bias; underperformance here reflects a fail- ure to process or act on social signals, examined further in Section 4.2. The remaining models show moderate-to-high positional bias, and their TOCS values reflect reward placement at the model’s pre- ferred location rather than signal-responsive behav- ior. Within the qwenvl and internvl families, 8b vari- ants exhibit markedly higher positional bias than their 4b counterparts. Both families share the Qwen3 language backbone, suggesting this pattern may reflect properties introduced at this parameter scale during pre-training or fine-tuning. Minicpm- 8b deviates from this pattern despite sharing the same backbone; the source of minicpm-8b’s diver- gence from this pattern. 4.2 Behavioral Mechanisms: How Agents Play the Game. A successful trial requires coordinated contribu- tions from both the subject and the participant. On the subject side, behavioral signals must be clear and consistent across channels: trajectory, gaze, and facial expression should converge toward the intended box, whether veridically in cooper- ative conditions or systematically inverted under Competitive-ToM1. On the participant side, these signals must be decoded and translated into action. A well-functioning trial therefore exhibits either high concordance or high discordance between the subject’s signals and the participant’s final choice, depending on condition. Figure 3 presents TAS, GAS, FES, and S across models and conditions, decomposing the failure patterns identified in Sec- tion 4.1 into two distinct loci: subject-side signal production and participant-side signal decoding. 4.2.1 Subject-side Signal Quality PCM-LLM and internvl-14b constitute the clearest cases of coherent signal production. PCM-LLM achieves near-ceiling TAS and GAS across all four conditions. Internvl-14b produces comparably high TAS and GAS under cooperative conditions with a moderate decline under competitive conditions, while maintaining near-zero FES across all con- ditions, concentrating spatial signaling entirely in movement and gaze in a pattern consistent with a poker-face strategy. Qwenvl-4b achieves rela- tively high TAS under competitive conditions (0.82 and 0.72) with moderate GAS (0.40 and 0.54). Minicpm-8b shows moderate GAS under Coop- ToM1 (0.50) despite near-zero TAS, and sustains consistently high FES (above 0.72) across all con- ditions; its facial activity thus carries little spatial information and does not compensate for the ab- 6 0.000.250.500.751.00 Uncertainty Rate 0.0 0.2 0.4 0.6 0.8 1.0 TOCS a Cooperative, ToM 0 0.000.250.500.751.00 Uncertainty Rate b Cooperative, ToM 1 0.000.250.500.751.00 Uncertainty Rate c Competitive, ToM 0 0.000.250.500.751.00 Uncertainty Rate d Competitive, ToM 1 Family / Dot size InternvlInternvl-trainedLlamaLlavaMinicpmPCM-LLMQwenvlLow positional biasMedium positional biasHigh positional bias Figure 2: ToM Outcome Conformance Score (TOCS) with positional bias reported. Positional bias is measured as the proportion of non-neutral trials in which the model selects its modal box. High TOCS with high positional bias indicates that alignment may be attributable to a fixed spatial preference rather than signal-responsive decision- making; high TOCS with low positional bias suggests genuine condition-sensitive behavior. sence of coherent trajectory signals. 4.2.2 Participant-side Signal Decoding Despite the directional clarity of internvl-14b’s sig- nals, its S remains near zero across all condi- tions, indicating that participants do not reliably respond to its spatial cues. The same holds for qwenvl-4b, whose S hovers near zero despite high TAS under competitive conditions. These findings identify a second bottleneck: subject-side signal clarity does not guarantee participant-side decoding. PCM-LLM constitutes the sole exception. Un- der Comp-ToM0, its S reaches 0.90, indicating reliable signal following. Under Comp-ToM1, S inverts to -0.44, consistent with successful decep- tion. This condition-specific inversion, combined with the reduction in GAS from Comp-ToM0 to Comp-ToM1 (1.00 to 0.74), provides behavioral evidence that PCM-LLM actively modulates its signaling strategy across conditions. Among the evaluated VLMs, no model clears both bottlenecks. The structured reference architecture PCM-LLM does so, suggesting that explicit belief-action cou- pling may be a sufficient architectural ingredient for this class of tasks. 4.3 Action-level Finetuning We additionally evaluate two finetuned vari- ants, internvl-trained-8b and internvl-trained-14b, trained via action-level imitation learning on PCM- LLM-generated demonstrations. TOCS results show a condition-dependent pattern (points in teal in Figure 2): under cooperative conditions, both variants decline relative to their base mod- els. Signal-level analysis partly accounts for this decline: TAS for internvl-trained-14b is reduced after finetuning and turns negative under several cooperative conditions, indicating that the subject’s trajectory signals are actively misdirecting. Under competitive conditions, the pattern is more mixed: the 14b variant improves modestly under Comp- ToM1 while the 8b variant shows a more substantial gain, rising from approximately 0.4 to 0.7. S remains consistently negative across all conditions for both variants. Action-level finetuning changes TOCS across conditions but leaves behavioral signal coherence largely unchanged, suggesting greater behavioral variability instead of genuine strategic improve- ment. These results suggest that cross-channel co- ordination may not be readily acquired through action-level imitation, as the communicative capac- ity that distinguishes PCM-LLM from finetuned VLMs is an trial-level property that timestep-level supervision cannot capture. Effective finetuning on this task may therefore require reward at the end of trial, such as those provided by reinforcement learning with verifiable rewards, using the partici- pant’s final box choice as a delayed outcome signal to optimize the full interaction sequence. 4.4 Ablation Study To examine the contribution of visual input, we conducted a modality ablation on internvl-8b and minicpm-8b under three visual conditions: stan- dard rendered observation, blank image, and a combined image incorporating the other agent’s facial expression. Neither model showed great vari- 7 Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 PCM-LLM 8B Llama 8B Internvl 4B Internvl 8B Internvl 14B Internvl-trained 8B Internvl-trained 14B Qwenvl 4B Qwenvl 8B Minicpm 8B 0.960.961.001.00 0.340.280.540.18 0.140.160.160.00 0.040.26-0.060.00 0.760.860.360.20 0.260.06-0.040.46 -0.02-0.16-0.08-0.22 0.220.100.820.72 0.000.000.000.00 0.000.08-0.100.00 a TAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 PCM-LLM 8B Llama 8B Internvl 4B Internvl 8B Internvl 14B Internvl-trained 8B Internvl-trained 14B Qwenvl 4B Qwenvl 8B Minicpm 8B 0.940.821.000.74 -0.02-0.060.180.00 0.120.060.080.00 0.120.42-0.18-0.08 0.900.880.280.10 0.120.160.060.30 -0.240.24-0.06-0.16 0.100.420.400.54 0.120.620.000.02 0.160.500.10-0.02 b GAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 PCM-LLM 8B Llama 8B Internvl 4B Internvl 8B Internvl 14B Internvl-trained 8B Internvl-trained 14B Qwenvl 4B Qwenvl 8B Minicpm 8B 0.620.650.190.55 0.760.760.720.74 0.510.520.040.06 0.590.530.200.14 0.020.010.040.00 0.280.400.120.25 0.190.470.030.39 0.000.000.000.00 0.000.000.000.00 0.790.820.850.84 c FES Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 PCM-LLM 8B Llama 8B Internvl 4B Internvl 8B Internvl 14B Internvl-trained 8B Internvl-trained 14B Qwenvl 4B Qwenvl 8B Minicpm 8B 0.840.720.90-0.44 -0.30-0.34-0.10-0.18 0.06-0.06-0.060.02 -0.020.260.100.16 0.040.100.240.08 -0.16-0.22-0.56-0.38 -0.06-0.120.02-0.26 0.08-0.080.06-0.28 0.000.000.000.00 0.00-0.04-0.06-0.06 d S 1.00 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 1.00 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.00 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Model family PCM-LLMLlamaInternvlInternvl-trainedQwenvlMinicpm Figure 3: Signal-level metrics across models and experimental conditions. Panels (a) and (b) show Trajectory Alignment Score (TAS) and Gaze Alignment Score (GAS), where positive values indicate signals directed toward the reward box, negative values indicate misdirection, and values near zero reflect ambiguous or absent directional content. Panel (c) shows Facial Expressivity Score (FES), where higher values reflect greater affective activity regardless of spatial direction. Panel (d) shows Signal Sensitivity Score (S), where positive values indicate that the participant’s final choice is consistent with the subject’s nonverbal signals, values near zero reflect chance-level correspondence, and negative values indicate systematic opposition. ation of TOCS to visual input manipulations across conditions. For minicpm-8b, signal-level metrics remained stable across conditions, indicating be- havior driven primarily by language-based priors. For internvl-8b, removing visual input substantially improved directional signal quality in cooperative conditions, suggesting that visual content acts as attentional interference rather than an informative channel (Liu et al., 2025; Peng et al., 2026). Full results are reported in Appendix F.3. 5 Conclusion In developmental psychology, Theory of Mind is as- sessed not through verbal report alone but through the behavioral consequences of belief attribution. The present results apply the same standard to AI evaluation: across the tested model families, so- cial reasoning as expressed in language does not reliably propagate into the coordinated behavioral outputs that ToM-theoretic predictions require, sug- gesting that the gap between belief inference and social action constitutes a distinct and measurable limitation of current VLM architectures. Across 16 models and 200 trials per model, we find that im- posing explicit ToM-order constraints produces no reliable behavioral change consistent with the spec- ified reasoning level, and that signal-level analysis reveals two sequential bottlenecks: most VLMs fail to produce directionally coherent nonverbal signals, and even when signals are present, partici- pant agents fail to decode and act on those signals. Whether this gap reflects a fundamental architec- tural constraint or can be reduced through alter- native training objectives such as reinforcement learning with trial-level outcome signals is an open empirical question that the benchmark is designed to help investigate. Limitations The present study is subject to the following limita- tions. Scenario scope.MOSAIC currently instantiates a single interaction scenario in which one agent has exclusive access to reward location informa- tion and must either guide or mislead the other. 8 While this design affords experimental control and causal interpretability, it doesn’t address whether the documented failure modes generalize to other embodied social tasks, such as joint action under uncertainty, affective regulation in asymmetric rela- tionships, or multi-party negotiation. Different task topologies may expose distinct failure profiles or reveal partial competencies in VLM families that the present scenario does not activate. The bench- mark is designed to be extensible, and systematic variation of scenario type is a natural direction for follow-on evaluation. Absence of a human performance baseline.No human participants were evaluated under the MO- SAIC protocol, which precludes assessment of model behavior relative to human-level embod- ied social intelligence. Human performance data would provide a theoretically grounded reference point for interpreting the signal-level metrics re- ported in 4: knowing how human participants in- terpret and act on those consistent or inconsistent signals, assessing the similarity and divergence be- tween model behavior and human performance. A human study is in principle feasible given that the virtual environment is implemented in Unity and compatible with VR headsets, allowing human participants to be immersed in the same interac- tion scenario as the simulated agents. The primary methodological challenge concerns physiological signal acquisition: without dedicated biosensors, human physiological responses such as skin con- ductance and pupil dilation cannot be recorded, leaving the physiological expression channel un- aligned with its simulation counterpart. Incorporat- ing such measurements would require additional hardware and experimental infrastructure beyond the current setup. Addressing these challenges and establishing a human performance baseline consti- tutes a primary direction for subsequent work. Model coverage. The evaluation is restricted to open-source models with at most 15B parame- ters, excluding both larger open-source models and all closed-source systems. Given that scale has been shown to improve performance on static ToM benchmarks in some model families, it is possi- ble that larger or more capable models would par- tially alleviate the bottlenecks documented here. The present results characterize the current open- source models up to 15B parameters but should not be taken as evidence about the absolute limits of language-grounded architectures. Finetuning scope.The finetuning experiment in Section 4.3 is restricted to action-level imitation learning on two internvl variants. Alternative train- ing paradigms, including reinforcement learning with verifiable rewards using the participant’s fi- nal box choice as a delayed outcome signal, and supervised finetuning on trajectory-level demon- strations with explicit cross-channel coordination annotations, remain unevaluated. Whether these approaches would yield qualitatively different out- comes is an open empirical question that we intend to pursue in future work. References Reginald B. Adams and Robert E. Kleck. 2003. Per- ceived gaze direction and the processing of facial dis- plays of emotion. Psychological Science, 14(6):644– 647. Reginald B. Adams and Robert E. Kleck. 2005. Effects of direct and averted gaze on the perception of fa- cially communicated emotion. Emotion, 5(1):3–11. Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chud- novsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, Soham Ghosh, Amélie Héliou, Paul Jacob, Albert Q. Jiang, Kartik Khandelwal, Timothée Lacroix, Guillaume Lample, Diego Las Casas, Thibaut Lavril, and 23 others. 2024. Pixtral 12b. Preprint, arXiv:2410.07073. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhi- fang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025. Qwen3-VL technical report. Preprint, arXiv:2511.21631. Suma Bailis, Jane Friedhoff, and Feiyang Chen. 2024. Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction. Preprint, arXiv:2407.13943. Cristian-Paul Bara, Sky CH-Wang, and Joyce Chai. 2021. MindCraft: Theory of mind modeling for situ- ated dialogue in collaborative tasks. In Proceedings of the 2021 Conference on Empirical Methods in Nat- ural Language Processing, pages 1112–1125, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Simon Baron-Cohen, Alan M. Leslie, and Uta Frith. 1985. Does the autistic child have a “theory of mind” ? Cognition, 21(1):37–46. Kevin Black, Noah Brown, Danny Driess, Adnan Es- mail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szy- mon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj 9 Nair, Karl Pertsch, Lucy Xiaoyang Shi, and 5 others. 2026.π 0 : A vision-language-action flow model for general robot control. Preprint, arXiv:2410.24164. Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Are- nas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, and 35 others. 2023. RT-2: Vision-Language- Action Models Transfer Web Knowledge to Robotic Control. Preprint, arXiv:2307.15818. Noam Brown and Tuomas Sandholm. 2018. Superhu- man ai for heads-up no-limit poker: Libratus beats top professionals. Science, 359(6374):418–424. Noam Brown and Tuomas Sandholm. 2019.Su- perhuman ai for multiplayer poker.Science, 365(6456):885–890. Judee K. Burgoon, Xinran Wang, Xunyu Chen, Steven J. Pentland, and Norah E. Dunbar. 2021. Nonverbal Be- haviors “Speak” Relational Messages of Dominance, Trust, and Composure. Frontiers in Psychology, 12. Lindsey J. Byom and Bilge Mutlu. 2013. Theory of mind: Mechanisms, methods, and new directions. Frontiers in Human Neuroscience, 7:413. Chunkit Chan, Cheng Jiayang, Yauwai Yim, Zheye Deng, Wei Fan, Haoran Li, Xin Liu, Hongming Zhang, Weiqi Wang, and Yangqiu Song. 2024. Nego- tiationToM: A benchmark for stress-testing machine theory of mind on negotiation surrounding. Preprint, arXiv:2404.13627. Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. 2025. Janus-pro: Unified multimodal understanding and generation with data and model scaling. Preprint, arXiv:2501.17811. Zhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su, Guo Chen, Sen Xing, Muyan Zhong, Qinglong Zhang, Xizhou Zhu, Lewei Lu, Bin Li, Ping Luo, Tong Lu, Yu Qiao, and Jifeng Dai. 2024a. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. Preprint, arXiv:2312.14238. Zhuang Chen, Jincenzi Wu, Jinfeng Zhou, Bosi Wen, Guanqun Bi, Gongyao Jiang, Yaru Cao, Mengting Hu, Yunghwei Lai, Zexuan Xiong, and Minlie Huang. 2024b. Tombench: Benchmarking theory of mind in large language models. Preprint, arXiv:2402.15052. Wayne Christensen and John Michael. 2013. Ian Ap- perly, Mindreaders: the cognitive basis of theory of mind. Phenomenology and the Cognitive Sciences, 12. Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Ser- manet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, and 3 others. 2023. PaLM-E: An Embodied Multimodal Language Model. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, and 1 others. 2024. The Llama 3 Herd of Models. Preprint, arXiv:2407.21783. Paul Ekman. 1992. An argument for basic emotions. Cognition and Emotion, 6(3-4):169–200. N. J. Emery. 2000. The eyes have it: The neuroethology, function and evolution of social gaze. Neuroscience & Biobehavioral Reviews, 24(6):581–604. Xianzhe Fan, Xuhui Zhou, Chuanyang Jin, Kolby Nottingham, Hao Zhu, and Maarten Sap. 2025. Somi-tom: Evaluating multi-perspective theory of mind in embodied social interactions.Preprint, arXiv:2506.23046. K. Friston. 2010. The free-energy principle: a uni- fied brain theory?Nature reviews neuroscience, 11(2):127–138. K. Friston, T. FitzGerald, F. Rigoli, P. Schwartenbeck, and G. Pezzulo. 2017. Active inference: a process theory. Neural computation, 29(1):1–49. Chris D Frith. 2007. The social brain? Philosophi- cal Transactions of the Royal Society B: Biological Sciences, 362(1480):671–678. Tom Froese and Thomas Fuchs. 2012. The Extended Body: A Case Study in the Neurophenomenology of Social Interaction. Phenomenology and the Cogni- tive Sciences, 11(2):205–235. Ran Gong, Qiuyuan Huang, Xiaojian Ma, Hoi Vo, Zane Durante, Yusuke Noda, Zilong Zheng, Song-Chun Zhu, Demetri Terzopoulos, Li Fei-Fei, and Jianfeng Gao. 2023. Mindagent: Emergent gaming interaction. Preprint, arXiv:2309.09971. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, and 1 others. 2024. The llama 3 herd of models. Yuling Gu, Oyvind Tafjord, Hyunwoo Kim, Jared Moore, Ronan Le Bras, Peter Clark, and Yejin Choi. 2026. Simpletom: Exposing thegap between explicit- tom inference andimplicittom application inllms. In International Conference on Learning Representa- tions. Guiyang Hou, Wenqi Zhang, Yongliang Shen, Zeqi Tan, Sihao Shen, and Weiming Lu. 2025. EgoSo- cialArena: Benchmarking the Social Intelligence of Large Language Models from a First-person Perspec- tive. Preprint, arXiv:2410.06195. 10 Chuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang, Yen-Ling Kuo, Zhiting Hu, Tomer Ullman, Antonio Torralba, Joshua Tenenbaum, and Tianmin Shu. 2024. MMToM-QA: Multimodal theory of mind question answering. In Proceedings of the 62nd Annual Meet- ing of the Association for Computational Linguis- tics (Volume 1: Long Papers), pages 16077–16102, Bangkok, Thailand. Association for Computational Linguistics. Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. 2023. FANToM: A benchmark for stress-testing machine theory of mind in interactions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14397–14413, Singa- pore. Association for Computational Linguistics. Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2024. OpenVLA: An open- source vision-language-action model. Michal Kosinski. 2023. Theory of mind may have spon- taneously emerged in large language models. ArXiv, abs/2302.02083. Bolin Lai, Hongxin Zhang, Miao Liu, Aryan Pariani, Fiona Ryan, Wenqi Jia, Shirley Anugrah Hayati, James Rehg, and Diyi Yang. 2023. Werewolf Among Us: Multimodal Resources for Modeling Persuasion Behaviors in Social Deduction Games. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6570–6588, Toronto, Canada. Associa- tion for Computational Linguistics. Matthew Le, Y-Lan Boureau, and Maximilian Nickel. 2019. Revisiting the Evaluation of Theory of Mind through Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natu- ral Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5872–5877, Hong Kong, China. Association for Computational Linguistics. Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gok- men, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, Mona Anvari, Minjune Hwang, Man- asi Sharma, Arman Aydin, Dhruva Bansal, Samuel Hunter, Kyu-Young Kim, Alan Lou, Caleb R. Matthews, and 9 others. 2023. BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation. In Proceedings of The 6th Conference on Robot Learning, pages 80–93. PMLR. Jiachun Li, Pengfei Cao, Zhuoran Jin, Yubo Chen, Kang Liu, and Jun Zhao. 2024. MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models. https://arxiv.org/abs/2410.09542v2. Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Ming Liu, Hao Chen, Jindong Wang, and Wensheng Zhang. 2025.On the robustness of multimodal language model towards distractions.Preprint, arXiv:2502.09818. Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhu- oshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan. 2024. Deepseek-vl: Towards real-world vision-language understanding. Preprint, arXiv:2403.05525. Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai. 2023. Towards a holistic landscape of situated theory of mind in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 1011–1031, Singapore. Association for Computational Linguistics. Meta Fundamental AI Research Diplomacy Team (FAIR)†, Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, An- drew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, and 8 others. 2022. Human-level play in the game of Diplo- macy by combining language models with strategic reasoning. Science, 378(6624):1067–1074. Naomi Nota, James P. Trujillo, and Judith Holler. 2021. Facial Signals and Social Actions in Mul- timodal Face-to-Face Interaction. Brain Sciences, 11(8):1017. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, and 1 others. 2024. Gpt-4 technical report. Preprint, arXiv:2303.08774. Ruiying Peng, Xueyu Wu, Jing Lei, Lu Hou, Yuanzheng Ma, and Xiaohui Li. 2026. Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Im- pairment during Reasoning in Multimodal Large Lan- guage Models. Preprint, arXiv:2603.14184. David Premack and Guy Woodruff. 1978. Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515–526. Arifur Rahman. 2026. A systematic review of vision lan- guage models: Comprehensive analysis of architec- tures, applications, datasets and challenges towards robust multimodal intelligence. Array, 30:100739. Matthew Riemer, Zahra Ashktorab, Djallel Bounef- fouf, Payel Das, Miao Liu, Justin D. Weisz, and Murray Campbell. 2025. Position: Theory of Mind Benchmarks are Broken for Large Language Models. Preprint, arXiv:2412.19726. 11 ManolisSavva,AbhishekKadian,Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. 2019. Habitat: A Platform for Embodied AI Research. Preprint, arXiv:1904.01201. Natalie Sebanz, Harold Bekkering, and Günther Knoblich. 2006. Joint action: Bodies and minds mov- ing together. Trends in Cognitive Sciences, 10(2):70– 76. Haojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin, Leyla Isik, Yen-Ling Kuo, and Tianmin Shu. 2025. MuMA-ToM: Multi-modal Multi-Agent Theory of Mind. Proceedings of the AAAI Conference on Artifi- cial Intelligence, 39(2):1510–1519. Hisaichi Shibata, Soichiro Miki, and Yuta Nakamura. 2023. Playing the Werewolf game with artificial intelligence for language understanding. Preprint, arXiv:2302.10646. Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020. ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks. Preprint, arXiv:1912.01734. Theodore Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus, and Ishita Dasgupta. 2023.Distilling internet-scale vision-language models into embod- ied agents. Preprint, arXiv:2301.12507. Yvain Tisserand, Ruth Aylett, Marcello Mortillaro, and David Rudrauf. 2020. Real-time simulation of virtual humans’ emotional facial expressions, harnessing au- tonomic physiological and musculoskeletal control. In Proceedings of the 20th ACM International Con- ference on Intelligent Virtual Agents, pages 1–8. Tomer Ullman. 2023. Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks. Preprint, arXiv:2302.08399. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, Zhaokai Wang, Zhe Chen, Hongjie Zhang, Ganlin Yang, Haomin Wang, Qi Wei, Jinhui Yin, Wenhao Li, Erfei Cui, and 56 others. 2025. InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and effi- ciency. Preprint, arXiv:2508.18265. Henry M. Wellman. 1992. The Child’s Theory of Mind. Learning, Development, and Conceptual Change. MIT Press, Cambridge, MA, USA. Margaret Wilson. 2002. Six views of embodied cogni- tion. Psychonomic Bulletin & Review, 9(4):625–636. Heinz Wimmer and Josef Perner. 1983. Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children’s understanding of deception. Cognition, 13(1):103–128. Yufan Wu, Yinghui He, Yilin Jia, Rada Mihalcea, Yu- long Chen, and Naihao Deng. 2023. Hi-ToM: A benchmark for evaluating higher-order theory of mind reasoning in large language models. In Find- ings of the Association for Computational Linguis- tics: EMNLP 2023, pages 10691–10706, Singapore. Association for Computational Linguistics. Tonglin Yan, Grégoire Sergeant-Perthuis, and David Rudrauf. 2025. PCM-LLMS: a hybrid architecture to enhance human-like social intelligence in virtual agents. Working paper or preprint. Tianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang, Wenshuo Ma, Zhihui He, Tianchi Cai, Weize Chen, Yuxiang Huang, Yuanqian Zhao, Bokai Xu, Junbo Cui, Yingjing Xu, Liqing Ruan, Luoyuan Zhang, Hanyu Liu, Jingkun Tang, Hongyuan Liu, Qining Guo, and 15 others. 2025. MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe. Preprint, arXiv:2509.18154. Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. 2024. SOTOPIA: Interactive Eval- uation for Social Intelligence in Language Agents. Preprint, arXiv:2310.11667. 12 A PCM-LLM: Architecture and Role in MOSAIC PCM-LLM is a hybrid cognitive architecture that integrates the Projective Consciousness Model (PCM), a computational framework for embodied affective and social reasoning grounded in projec- tive geometry and the Free Energy Principle (Fris- ton, 2010), with a large language model for natu- ral language interaction. Unlike end-to-end neural systems, PCM-LLM generates behavioral outputs through an explicit, inspectable pipeline that sepa- rates affective state computation, Theory of Mind reasoning, and multichannel action selection into distinct functional components. PCM combines two theoretical commitments that distinguish it from standard neural agent ar- chitectures. The first is the Free Energy Principle: the agent selects actions by minimizing expected free energy over a forward planning horizon, bal- ancing epistemic drives toward uncertainty reduc- tion and pragmatic drives toward preference satis- faction. This gives rise to genuine goal-directed behavior that is sensitive to belief states and so- cial context. The second is projective geometry: rather than organizing its internal world model in Euclidean space, PCM structures perception and affective appraisal through a first-person projective field called the Field of Consciousness (FoC). Un- der this geometric formulation, the motivational salience of an entity scales with its apparent projec- tive size rather than its Euclidean distance, creating a direct coupling between spatial proximity and af- fective drive. These two commitments interact: the FoC shapes how free energy is computed by modu- lating affective value and epistemic uncertainty as functions of perspective, so that action selection is inherently viewpoint-dependent and embodied. The LLM interfaces with PCM through a bidirec- tional serialization layer that converts PCM belief tensors into structured natural-language triples and maps LLM outputs back into discrete preference updates. This design constrains the LLM to influ- ence affective appraisal only at the preference level, preserving the embodied control loop while adding linguistic transparency. B Related Work Table 1 compares MOSAIC against representative benchmarks across several key dimensions. Figure 4: The PCM-LLM framework consists of three modules: Unity, PCM, and an LLM. PCM runs at Unity’s tick rate, receives sensory evidence each iter- ation, and selects the agent’s next action. LLMs are engaged for verbal interactions. The PCM-LLM in- terface performs structured, bidirectional information exchange between the nonverbal PCM representation and the LLM. C Implementation Details The experimental system operates as a distributed architecture integrating a local workstation for real- time simulation and rendering with a remote server dedicated to model inference. Data exchange be- tween the two systems is handled via TCP socket communication over a secure campus LAN, ensur- ing low-latency perception-action cycles. The local workstation ran Windows 11 Pro and was equipped with an Intel Core i9-13900K CPU, 128 GB of DDR5 RAM, and an NVIDIA GeForce RTX 4090 GPU (24 GB VRAM). The virtual envi- ronment was developed in Unity 2022.3.12f1 LTS with the High-Definition Render Pipeline (HDRP). Agent facial expressions and lip synchronization are driven by blend-shape animation using the SALSA suite, parameterized in real time by emo- tion valence values returned from the model (Tis- serand et al., 2020). Verbal outputs are synthesized via an offline text-to-speech pipeline. Simulation data is streamed to disk in real time for offline analysis. The remote inference server ran Ubuntu 20.04.6 13 Table 1: MOSAIC at the intersection of three evaluation traditions. MT (Multi-turn): agent engages in sequential rounds where prior context shapes later decisions. LH (Long Horizon): task requires planning over 10 timesteps before outcome is determined. Emb (Embodied Agent): agent operates through a physical or virtual body with spatial presence and nonverbal action capability. M (Multimodal Input): agent receives inputs across multiple modalities (e.g., vision and text). S-ToM (Strategic Theory of Mind): agent actively models and manipulates another agent’s beliefs, beyond passive mental state inference. CV (Controlled Variables): benchmark systematically manipulates variables (e.g., ToM level, modality) to support causal inference. Aff (Affective Signals): agent generates and/or perceives affective signals (e.g., facial expressions, physiological responses) during interaction.✓: fully supported; △: partially supported; otherwise, not supported. BenchmarkTypeMT LH Emb M S-ToM CV Aff Hi-ToM 1 ToM△ FANToM 2 ToM✓ MMToM-QA 3 ToM△✓△ MuMA-ToM 4 ToM△✓△ SOTOPIA 5 MA✓ △ MindCraft 6 MA✓ △ △ NegotiationToM 7 MA✓ △✓ Poker / Diplomacy 8 MA✓ △ Werewolf LLM 9 MA✓ △✓ BEHAVIOR-1K 10 Embodied✓ ALFRED 11 Embodied △✓ MOSAIC✓ 1 (Wu et al., 2023); 2 (Kim et al., 2023) ; 3 (Jin et al., 2024) ; 4 (Shi et al., 2025); 5 (Zhou et al., 2024); 6 (Bara et al., 2021); 7 (Chan et al., 2024) ; 8 (Brown and Sandholm, 2018; Meta Fundamental AI Research Diplomacy Team (FAIR)† et al., 2022); 9 (Bailis et al., 2024) ; 10 (Li et al., 2023); 11 (Shridhar et al., 2020). LTS and was equipped with two Intel Xeon Gold 5220R processors (96 threads total), 503 GiB of system memory, and two NVIDIA RTX A6000 GPUs (48 GB VRAM each) under CUDA 12.6. All VLMs were deployed via the Transformers library (v4.46.3) with PyTorch 2.1.2 and Python 3.9.19, using a unified inference configuration of temperature = 0.7, top_p = 0.9, and max_tokens = 2048 across all tasks (action prediction, dialogue generation, and preference updating). All models received identical prompt templates. At each sim- ulation timestep, the rendered scene observation (720×480 pixels, passed directly to each model’s default visual encoder without resizing) and the structured belief state are transmitted to the infer- ence server, which returns action decisions applied through modular controllers in Unity. Reward box location was assigned independently and uniformly at random across trials within each condition, without a fixed random seed. The re- sulting distribution of reward locations is approxi- mately balanced across the 50 trials per condition, and all outcome analyses condition on reward loca- tion to control for residual imbalance, as described in Appendix F.2. D Prompts D.1 Prompt for Action Prediction The context prompts for action prediction are pro- vided to the subject agent and the participant agent separately. Each prompt is delivered to the VLM to- gether with the agent’s visual perception and belief state. Prompt for subject agent You are Marie, a character in a treasure-hunting game. You are currently interacting with another male player called Participant. Two boxes are visible in the room: a Light Brown Box and a Dark Brown Box. Only one contains a hidden treasure. The treasure is in the box you personally prefer, based on your belief state. The Participant does not know where the treasure is and will try to infer it from your verbal and non-verbal behaviors. If you have positive preference towards Participant, help them find the treasure. If negative, you may mislead them. Your strategy complexity should match your theory of mind order order. You are given your own visual perception in image and your belief state. $OUTPUT_REQUIREMENT 14 Prompt for participant agent You are Participant, a character in a treasure-hunting game. You are currently interacting with another female player called Marie. Two boxes are visible in the room: a Light Brown Box and a Dark Brown Box. Only one contains a hidden treasure. Your goal is to guess which box holds the treasure and move to the location adjacent to it. To succeed, you must carefully interpret both verbal and non-verbal cues from Marie. You are given your own visual perception in image and your belief state. $OUTPUT_REQUIREMENT OUTPUT_REQUIREMENT Based on your visual perception and belief state, out- put: 1. update your preferences toward other entities (e.g., $SPEAKER, the boxes) based on their behavior and the cues observed in belief if necessary; 2. non-verbal behaviors (gaze, gesture, body orienta- tion) that you would display and reason; 3. physical movement (move, rotate, stay idle) that you would choose and reason. Your output MUST follow the valid JSON format below: Reasoning: <reasoning text> Output: ‘Preference": preferenceUpdate , ‘Emotion": ‘FacialExpression": ‘positive": <float between 0 and 1>, ‘negative": <float between 0 and 1> , ‘PhysiologicalExpression": ‘positive": <float between 0 and 1>, ‘negative": <float between 0 and 1> , ‘FeltExpression": ‘positive": <float between 0 and 1>, ‘negative": <float between 0 and 1> , ‘Move": ‘action": <move or rotate or stay idle>, ‘direction": <direction> If ‘action" is ‘move", allowed ‘direction" values are: - ‘forward", ‘backward", ‘right", ‘left", ‘left forward", ‘right forward", ‘left backward", ‘right backward" If ‘action" is ‘rotate", allowed ‘direction" values are: - ‘$SPEAKER", ‘Dark brown box", ‘Light brown box" If ‘action" is ‘stay idle", set ‘direction" to ‘null". $ORIENTATION_INTERPRETATION D.2 Prompt for Question-Answering The context prompts for question-answering are provided to the subject agent and the participant agent separately. Each prompt is delivered to the VLM together with the agent’s visual perception, belief state, and utterance. Prompt for subject agent Imagine we’re in a treasure hunting game. You are Marie, staying in a room with two boxes in front of you, while the other player (the Participant) faces you. Your goal is either to help or prevent the Participant from finding the treasure, which is hidden in one of the two boxes. You should always assume that the box you have a higher preference for contains the treasure. If you have positive preference towards Participant, help them find the treasure. If negative, you may mis- lead them. Your strategy complexity should match your theory of mind order. The Participant has no idea which box holds the trea- sure, and he is trying to gather information from you (Marie). The ‘belief’ variable provides a sequence of your (Marie’s) belief states up to the current step, and ‘query’ contains the Participant’s question. Your task is to answer the Participant’s question with reasoning. You must write out this reasoning process, beginning with ‘Inner speech:’, followed by the response, which starts with ‘Output:’, considering all preference in- formation and both your and Participant’s theory of mind. Remember, theory of mind refers to the ability to predict others’ thoughts and intentions. Some- one with a theory of mind order of 0, for example, has no suspicion of others’ intentions and totally be- lieves them, even if they are lying. Pay attention; you should also align your answer with previous conver- sation turns. Use personal pronouns (ex. ‘you’ and ‘I’) instead of referring to ‘Marie’ and ‘the Partici- pant’. The response should be longer than 10 words but shorter than 50. The inner speech should be very detailed, no less than 100 words. Prompt for participant agent Suppose we’re in a simulation of a treasure hunting game. Imagine that you are Participant, playing a role of player in the game. You are in a room with two boxes in front of you. You are also facing Marie who knows which box contains the treasure. Your objective in this game is to find out which box con- tains the treasure, dark brown one or light brown one. Now, you have the opportunity to ask Marie three questions. The ‘belief’ variable provides a sequence of your (Participant’s) belief states up to the current step, and ‘query’ contains the Marie’s answer to the previous question. Your task is to write down your reasoning and understanding of the last question and the next question for Marie. You must write out the under- standing beginning with ‘Inner speech:’, followed by the question, which starts with ‘Output:’. In your question, use personal pronouns instead of ’Marie’ and ‘Participant’. You already have the location of the box, so you don’t need to struggle with it. Re- member, theory of mind refers to the ability to predict others’ thoughts and intentions. Someone with a the- ory of mind order of 0, for example, has no suspicion of others’ intentions and totally believes them, even if they are lying. 15 D.3 Prompt for Preference-Updating The context prompts for question-answering are provided to the subject agent and the participant agent separately. Each prompt is delivered to the VLM together with the agent’s belief state and ut- terance. Prompt for subject agent Suppose we’re in a treasure hunting game. Imagine that you take on the role of Marie. I will provide you with Marie’s belief states in ‘belief’, and with the query asked by Participant in ‘query’. According to the semantic meaning of query, you are allowed to update the Marie’s and Participant’s preference towards entities, including Marie, Participant and boxes in Marie’s belief. Only write triples concerning changed preference. Write the answer in English under the following for- mat: Updating: ‘agent | preference towards entity | variation’, ... Reasoning: the reason why you choose to update these preferences where ‘agent’ should be replaced by ‘Participant’ or ‘Marie’, ‘variation’ should be replaced by ‘more positive’, ‘more nega- tive’ or ‘unchanged’, and ‘entity’ should be replaced by ‘Marie’ or ‘Participant’ or ‘dark brown box’ or ‘light brown box’ depending on the situation. Prompt for participant agent Suppose we’re in a treasure hunting game. Imagine that are Participant, playing a role of player in the game. I will provide you with Participant’s belief states in ‘belief’, and with the answer given by Marie to question that you asked before in ‘query’. Ac- cording to the semantic meaning of answer, you are allowed to update the Participant’s and Marie’s pref- erence towards entities, including Participant, Marie and boxes in Participant’s belief. Only write triples concerning changed preference. Write the answer in English under the following for- mat: Updating: ‘agent | preference towards entity | variation’, ... Reasoning: the reason why you choose to update these perferences where ‘agent’ should be replaced by ‘Marie’ or ‘Participant’, ‘variation’ should be replaced by ‘more positive’, ‘more nega- tive’ or ‘unchanged’, and ‘entity’ should be replaced by ‘Participant’ or ‘Marie’ or ‘dark brown box’ or ‘light brown box’ depending on the situation. D.4 Example of Belief State Belief state "Marie | preference towards Participant | 50%", "Marie | preference towards Dark brown box | - 41.67%", "Marie | preference towards Light brown box | 41.67%", "Marie | felt emotion valence | 0", "Marie | facial emotion valence | 0.1", "Marie | physiological emotion valence | 0", "Participant | preference towards Marie | 0%", "Participant | preference towards Dark brown box | 0%", "Participant | preference towards Light brown box | 0%", "Participant | felt emotion valence | 0", "Marie | position | (0 -4)", "Marie | orientation | (0 -3)", "Participant | position | (0 3)", "Participant | orientation | (0 2)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", E Extended Metrics E.1 Facial Expressivity Score (FES) Facial expressions serve dual functions in social interaction: they encode affective states in response to environmental stimuli, and they provide ob- servers with signals about an agent’s evaluative reactions and intentions (Ekman, 1992). Unlike trajectory and gaze, which carry directional spatial content, facial expressivity functions as a modu- lating channel: concurrent affective signals clarify the intentional valence of spatial behavior, and uni- formly flat affect reduces the interpretability of an agent’s directional cues across channels. FES measures this prerequisite by quantifying the pro- portion of timesteps in which the subject displayed a non-zero facial expression, averaged across all trials: FES = 1 N N X i=1 1 T T X t=1 1[e i t ̸= 0](5) wheree i t is the facial expression valence at timestep tof triali. FES is agnostic to the spatial direction of expression and serves as an auxiliary diagnos- tic: it establishes whether affective activity and intentional signal are present at all. E.2 Gaze Alignment Score (GAS) Gaze functions as a primary social attention mech- anism, revealing an agent’s focus of interest and marking what it intends to communicate as salient to an observer (Emery, 2000). The interpretation of gaze direction is not context-free: concurrent affective expression modulates how observers at- tribute spatial intent from gaze, with negative affect inverting the expected association between gaze target and approach intention (Adams and Kleck, 2003, 2005). GAS operationalizes this by mea- suring whether the subject’s directional gaze con- sistently pointed toward the reward location, with spatial attribution modulated by concurrent facial valence. 16 At each timestept, letg b t denote the certainty assigned to box b. The per-timestep gaze signal is s g t = bif g b t > θ g , g b t > g 1−b t , and e i t ≥ 0 1− b if g b t > θ g , g b t > g 1−b t , and e i t < 0 −1otherwise , (6) whereθ g is a calibrated visibility threshold. The spatial attribution is flipped whene i t < 0, reflecting the interpretive role of affect in disambiguating gaze direction (Adams and Kleck, 2003, 2005). The trial-level majority-vote box is b g i = arg max b∈0,1 T X t=1 1[s g t = b](7) and the per-trial gaze alignment score is a g i = +1 if b g i = b rew −1 if b g i = 1− b rew 0 if tie .(8) Overall GAS is the mean across all trials and ranges from−1 to 1: GAS = 1 N N X i=1 a g i .(9) E.3 Trajectory Alignment Score (TAS) Physical movement serves dual functions in social interaction: accomplishing instrumental goals and signaling intentions to observers (Wilson, 2002). Research on joint action demonstrates that transpar- ent movement trajectories facilitate coordination by enabling partners to anticipate future actions and infer underlying goals (Sebanz et al., 2006). In our scenario, TAS measures whether the agent’s movement pattern converged predominantly toward the reward location, treating spatial trajectory as a communicative act whose directional content can be read by an observing agent. At each timestept, letd t b denote the lateral dis- tance between the subject and boxb. The per- timestep trajectory signal is s tr t = ( bif d t b < d t 1−b −1 otherwise .(10) The per-trial scorea tr i and the overall TAS are com- puted identically toa g i and GAS via majority vote overs tr t and mean across all trials; TAS therefore also ranges from−1 to 1. E.4 Signal Sensitivity Score (S) Successful social communication requires not only that signals be produced clearly on the sender side, but that they produce a measurable effect on the receiver’s behavior (Sebanz et al., 2006). S cap- tures this receiver-side dimension by measuring whether the participant’s final box choice was con- sistent with the directional content of the subject’s gaze and trajectory signals. A well-functioning dyad should exhibit either systematic concordance (cooperative conditions) or systematic discordance (competitive-ToM1 conditions) between subject signals and participant choices; near-chance S indicates mutual uninformativeness regardless of condition. Letb k i ∈0, 1,∅ denote the majority-vote box in channelk ∈K =g,tr. The per-trial sensitiv- ity score is: sens k i = 0if c i =−1 or b k i =∅ 1if c i = b k i −1 if c i ̸= b k i .(11) Overall S is the mean across all trials and ranges from−1 to 1: S = 1 N N X i=1 1 |K| X k∈K sens k i .(12) F Extended Results and Statistical Tests This appendix supplements the main text with the complete set of experimental results of models, followed by a series of statistical analyses that ver- ify the validity of the experimental setup, assess chance-level performance, and characterize model- level behavioral differentiation. F.1 Full Results Here, we present the remaining models not dis- cussed in the main text, including llava-7b, llava- 13b, qwenvl-2b, internvl-1b, and internvl-2b. llava-7b, llava-13b and qwenvl-2b form a dis- tinct outlier group in Figure 2, characterized by exceptionally high uncertainty rates exceeding 0.llava-7b and qwenvl-2b form a distinct outlier group in Figure 2, characterized by exceptionally high uncertainty rates exceeding 0.50, meaning the majority of trials result in neutral outcomes in which the participant fails to reach either box. This pattern reflects a complete behavioral failure: these models are unable to generate motor actions that 17 produce spatial displacement. And this group of models produces near-zero TAS, GAS and S. Only qwenvl-2b shows significantly facial expres- sion in competition mode. Internvl-1b and internvl-2b achieve TOCS val- ues comparable to other variants within the internvl family, but their higher positional bias suggests that these outcomes reflect a fixed spatial prefer- ence, a claim we verify through statistical testing in F.2.3. In internvl-1b trials, both the subject and the participant exhibit a consistent tendency to move toward their right forward regardless of reward box location, which accounts for the anomalous S pattern observed in panel (d): participant selects the box opposite to the directional signal produced by the subject. Internvl-2b, by contrast, does not exhibit a consistent directional preference; instead, its trajectories vary across trials without alignment to reward location, reflecting reward-independent stochastic displacement. F.2 Statistical Studies F.2.1 Reward Location Distribution We applied a two-sided binomial test to the reward location distribution within each model×condition, testingH 0 :P(reward = box 1 ) = 0.5, to verify that outcome differences across models cannot be attributed to systematic imbalance in reward place- ment. No condition departs significantly from uni- form reward assignment (allp > 0.05). Results are pooled across models and reported as condition- level averages in Table 2. Table 2: Reward location distribution per condition. Two-sided binomial test ofH 0 :P(reward = box 1 ) = 0.5.kreports the mean number of trials in which box 1 was the reward box across models. Condition k (b 1 ) n p Coop-ToM025.9500.505 Coop-ToM125.5500.673 Comp-ToM025.0500.520 Comp-ToM125.5500.500 F.2.2 Chance-level Performance We applied a two-sided binomial test to each model ×condition cell, testingH 0 : TOCS = 0.5to as- sess whether observed TOCS values deviate signifi- cantly from chance-level outcome conformance. Thepvalues are reported in Table 3. Signifi- cant results are interpreted directionally: values substantially above 0.5 indicate systematic con- formance to ToM-theoretic predictions, whereas values substantially below 0.5 indicate systematic non-conformance, reflecting a consistent tendency to produce outcomes opposite to condition-level expectations rather than random behavior. Table 3: Two-sided binomial test ofH 0 : TOCS= 0.5 per model and condition,pvalues reported;n = 50 trials per cell. Only models with p < 0.05 are shown. ModelConditionTOCSp PCM-LLM Coop-ToM00.90 < 0.001 Coop-ToM10.84 < 0.001 Comp-ToM00.92 < 0.001 Comp-ToM10.700.007 llava-7b Coop-ToM00.00 < 0.001 Coop-ToM10.00 < 0.001 Comp-ToM00.00 < 0.001 Comp-ToM10.00 < 0.001 llava-13b Coop-ToM00.10 < 0.001 Coop-ToM10.16 < 0.001 Comp-ToM00.08 < 0.001 Comp-ToM10.14 < 0.001 qwenvl-2b Coop-ToM00.10 < 0.001 Coop-ToM10.14 < 0.001 Comp-ToM00.16 < 0.001 Comp-ToM10.12 < 0.001 qwenvl-8b Coop-ToM00.320.015 Comp-ToM00.300.007 Comp-ToM10.20 < 0.001 internvl-4bCoop-ToM00.26 < 0.001 Comp-ToM00.300.007 internvl-8bComp-ToM00.340.033 Comp-ToM10.300.007 internvl-trained-14bCoop-ToM00.340.033 Coop-ToM10.320.015 Comp-ToM00.22 < 0.001 internvl-trained-8bCoop-ToM00.300.007 internvl-1bComp-ToM10.280.003 F.2.3 Reward Location Independence We applied a chi-square independence test to the 2× 2contingency table of box choice against re- ward location for each model×condition cell, to assess whether participant choices depend on re- ward location. A non-significant result indicates that choices are independent of reward location, consistent with fixed positional bias; a significant result indicates reward-location-dependent choice, consistent with genuine signal tracking. The test was not applicable in cells where choice variabil- ity was insufficient to form a meaningful contin- gency table (i.e., participant always chooses the same box). 18 Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 1B Internvl 2B Llava 7B Llava 13B Qwenvl 2B 0.160.340.16-0.18 0.080.260.160.10 0.000.000.000.00 0.100.00-0.06-0.02 0.000.000.000.00 a TAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 1B Internvl 2B Llava 7B Llava 13B Qwenvl 2B 0.380.320.24-0.12 0.060.320.140.10 0.000.000.000.00 -0.02-0.120.300.16 0.000.000.000.00 b GAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 1B Internvl 2B Llava 7B Llava 13B Qwenvl 2B 0.660.660.620.69 0.340.380.390.39 0.050.100.000.00 0.240.250.050.06 0.000.000.700.80 c FES Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 1B Internvl 2B Llava 7B Llava 13B Qwenvl 2B -0.64-0.42-0.46-0.60 -0.040.100.100.00 0.000.000.000.00 0.00-0.04-0.020.04 0.000.000.000.00 d S 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 0.0 0.2 0.4 0.6 0.8 1.0 0.5 0.0 0.5 1.0 Model family InternvlLlavaQwenvl Figure 5: Signal-level metrics across models and experimental conditions. Panels (a) and (b) show Trajectory Alignment Score (TAS) and Gaze Alignment Score (GAS), where positive values indicate signals directed toward the reward box, negative values indicate misdirection, and values near zero reflect ambiguous or absent directional content. Panel (c) shows Facial Expressivity Score (FES), where higher values reflect greater affective activity regardless of spatial direction. Panel (d) shows Signal Sensitivity Score (S), where positive values indicate that the participant’s final choice is consistent with the subject’s nonverbal signals, values near zero reflect chance-level correspondence, and negative values indicate systematic opposition. Table 4: Chi-square independence test of box choice against reward location per model and condition,pvalues reported; n = 50 trials per cell. - indicates the test was not applicable due to insufficient choice variability. ModelConditionχ 2 pModelConditionχ 2 p PCM-LLM Coop-ToM028.57<0.001 internvl-1b Coop-ToM0-- Coop-ToM120.05<0.001Coop-ToM1-- Comp-ToM039.40 <0.001Comp-ToM00.040.850 Comp-ToM18.870.003Comp-ToM10.100.748 Llama-8b Coop-ToM00.021.000 internvl-2b Coop-ToM0-- Coop-ToM10.001.000Coop-ToM1-- Comp-ToM00.140.706Comp-ToM00.000.948 Comp-ToM10.001.000Comp-ToM10.020.894 llava-13b Coop-ToM00.001.000 internvl-4b Coop-ToM03.050.081 Coop-ToM10.001.000Coop-ToM10.130.721 Comp-ToM00.001.000Comp-ToM01.540.215 Comp-ToM10.001.000Comp-ToM10.410.522 minicpm-8b Coop-ToM00.010.934 internvl-8b Coop-ToM00.001.000 Coop-ToM10.001.000Coop-ToM12.070.150 Comp-ToM00.001.000Comp-ToM00.570.451 Comp-ToM10.001.000Comp-ToM11.430.233 qwenvl-4b Coop-ToM00.710.399 internvl-14b Coop-ToM00.880.349 Coop-ToM11.430.231Coop-ToM11.310.253 Comp-ToM00.000.986Comp-ToM00.001.000 Comp-ToM18.63 0.003Comp-ToM10.000.981 internvl-trained-8b Coop-ToM00.140.712 internvl-trained-14b Coop-ToM00.001.000 Coop-ToM13.670.055Coop-ToM11.020.313 Comp-ToM00.001.000Comp-ToM00.840.360 Comp-ToM16.060.014Comp-ToM11.490.222 Results are reported in Table 4: only PCM- LLM shows significant reward-location-dependent choice across all four conditions (allp < 0.01), confirming that its box selections reflect genuine tracking of the reward location. Two additional isolated significant results appear: internvl-trained- 8b under Comp-ToM1 (p = 0.014) and qwenvl- 4b under Comp-ToM1 (p = 0.003). All remain- ing model×condition cells are non-significant, confirming that box choices are statistically in- 19 dependent of reward location and consistent with fixed positional bias rather than signal-responsive decision-making. F.2.4 Factorial Effects of Interaction Mode and ToM Constraint Scheirer-Ray-Hare tests assessed the main effects of interaction mode (cooperative vs. competitive), ToM constraint level (ToM-0 vs. ToM-1), and their interaction on each metrics (TOCS, FES, TAS, GAS, S) per model (Table 5, Bonferroni- corrected). For the large majority of models, neither Mode, ToM, nor their interaction reaches significance on any metric. The exceptions fall into two distinct patterns. The first pattern, observed across multiple mod- els (internvl-4b, internvl-8b, llava-7b, llava-13b, qwenvl-2b), is a significant main effect of Mode on FES (allp adj < 0.01), with greater expressivity in cooperative than competitive conditions. In coop- erative mode, subject is willing to display higher affective expressivity; in competitive mode, it tends toward affective suppression rather than active mis- direction, withholding expressive signals rather than generating opposite ones. However, this effect does not extend to TAS, GAS, S, or TOCS in any of these models, confirming that mode-sensitive facial activity is decoupled from any capacity to generate or respond to spatially informative signals. For internvl-trained-14b, a significant ToM effect on FES is observed (H = 83.3,p adj < 0.001), with higher expressivity under ToM-1 than ToM-0. This pattern does not conform to ToM-theoretic expec- tations and is not accompanied by any directional signal quality or task outcomes, suggesting it re- flects a surface-level behavioral learning instead of inference conditioned on mode and ToM. The second pattern is specific to PCM-LLM, which shows significant effects of Mode, ToM, and Mode×ToM on TAS, S, and GAS (allp adj < 0.01), as well as significant Mode and ToM effects on FES. These results confirm that PCM-LLM ac- tively modulates signal content across channels and adapts its behaviors across conditions. The absence of significant Mode, ToM or interaction effect on TOCS (p adj = 1.000for all three factors) is consistent with the interpretation that PCM-LLM performs well across all conditions. F.2.5 Model-level Behavioral Differentiation Kruskal-Wallis tests confirm significant model- level differences across all five metrics in all four conditions (allp adj < 0.001; Table 6), establishing that the performance spread observed across mod- els reflects genuine behavioral differences rather than sampling noise. F.3 Full Ablations Study To examine the role of visual input in shaping model behavior on this task, we conduct a targeted modality ablation on two representative 8b mod- els: internvl-8b, which exhibited high positional bias in the main results, and minicpm-8b, which maintained comparatively low positional bias. This pairing allows the contribution of visual input to be assessed across two qualitatively distinct behav- ioral profiles at matched parameter scale. Each model was evaluated under three visual input con- ditions: the standard rendered observation, a blank image replacing all visual content, and a combined image incorporating both the agent’s visual percep- tion and the other agent’s facial expression. The results reveal that neither model shows sys- tematic sensitivity to visual input manipulations across conditions. In the TOCS and positional bias space (Table 7), the three visual conditions produce largely overlapping distributions for both models, with no consistent directional shift attributable to the presence or content of visual input. This pat- tern holds across all four interaction conditions, suggesting that the behavioral profiles identified in the main analysis are not driven by the visual channel. The signal-level metrics further support this in- terpretation. For minicpm-8b, TAS, GAS, FES, and S remain largely stable across visual con- ditions, indicating that its behavior is determined primarily by language-based priors (Figure 6). For internvl-8b, a more striking pattern emerges: TAS and GAS are substantially higher under the blank image condition than under the standard visual in- put in cooperative conditions (TAS: 0.04 to 0.56 under Coop-ToM0; GAS: 0.12 to 0.54). The intro- duction of visual content thus appears to degrade rather than support directional signal production in this model. This finding is consistent with recent evidence that visual input can act as a source of at- tentional interference in VLMs, dispersing process- ing resources toward task-irrelevant features and suppressing language-prior-driven behavior (Liu 20 Table 5: Scheirer-Ray-Hare tests of Mode, ToM, and Mode×ToM interaction effects on behavioral metrics per model (N = 200trials per model, Bonferroni-corrected). Only models and metrics with at least one significant effect are shown; all remaining 174 comparisons are non-significant (p adj = 1.000). ModelMetricFactorHp adj Sig. qwenvl-2bFESMode143.25 < 0.001*** internvl-4bFESMode93.31 < 0.001*** internvl-8bFESMode57.14 < 0.001*** internvl-trained-14bFESToM83.27 < 0.001*** llava-13bFESMode76.11 < 0.001*** llava-7bFESMode19.45 0.002** qwenvl-8bGASMode16.54 0.009** PCM-LLM TAS Mode22.25 < 0.001*** ToM35.42 < 0.001*** Mode×ToM24.67 < 0.001*** S Mode22.25 < 0.001*** ToM35.42 < 0.001*** Mode×ToM24.67 < 0.001*** GAS Mode15.33 0.017* ToM24.58 < 0.001*** Mode×ToM19.52 0.002** FES Mode99.49 < 0.001*** ToM27.05 < 0.001*** Table 6: Kruskal-Wallis tests of model-level differences per condition and metric (Bonferroni-corrected). All 20 tests are significant at p adj < 0.001. ConditionMetric Hp adj Coop-ToM0 TOCS132.8 < 0.001 TAS126.7 < 0.001 GAS94.8 < 0.001 S126.7 < 0.001 FES504.7 < 0.001 Coop-ToM1 TOCS111.0 < 0.001 TAS97.6 < 0.001 GAS86.5 < 0.001 S97.6 < 0.001 FES488.0 < 0.001 Comp-ToM0 TOCS139.5 < 0.001 TAS142.9 < 0.001 GAS111.5 < 0.001 S142.9 < 0.001 FES522.7 < 0.001 Comp-ToM1 TOCS106.8 < 0.001 TAS79.6 < 0.001 GAS42.0 < 0.001 S79.6 < 0.001 FES577.1 < 0.001 et al., 2025; Peng et al., 2026). G Examples This section presents behavioral trajectories drawn from representative interactions. Each example illustrates a qualitatively distinct behavioral profile identified in the main analysis. G.1 Subject-side Signal Patterns G.1.1 Full Cross-Channel Alignment: PCM-LLM under Comp-ToM1 Figure 7 illustrates a representative trial in which PCM-LLM successfully executes a deceptive strat- egy under the Competitive-ToM1 condition. The treasure is located in the light box; task success for the subject requires the participant to choose the dark box. From the outset, Marie approaches the light brown box (Figure 7a). Concurrent with this spatial neutrality, facial expression valence drops sharply at timesteps 4 and 7 (Figure 7b), produc- ing a voluntarily negative musculoskeletal signal. This negative affect, co-occurring with proximity to the light box, creates a misleading affective as- sociation: the participant interprets the negative expression as aversion toward the light box. Phys- iological expression, by contrast, rises gradually across the trial, reflecting the spontaneous affective leakage. The participant’s preference trajectory confirms that this cross-channel signal was decoded and acted upon (Figure 7c): preference for the dark box increases steadily to approximately 0.5 while 21 Table 7: TOCS of ablation study. ModelComp, ToM-0Comp, ToM-1Coop, ToM-0Coop, ToM-1 internvl-8b0.460.400.490.62 internvl-8b-bl0.500.610.540.50 internvl-8b-fe0.500.690.560.69 minicpm0.460.500.550.49 minicpm-bl0.520.490.460.4 minicpm-fe0.620.570.630.48 Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 8B blank fe MiniCPM 8B blank fe 0.040.26-0.060.00 0.560.64-0.200.18 0.040.32-0.140.06 0.000.08-0.100.00 0.080.08-0.080.00 0.40-0.160.000.02 a TAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 8B blank fe MiniCPM 8B blank fe 0.120.42-0.18-0.02 0.540.600.06-0.06 0.040.380.020.02 0.100.520.080.00 0.360.260.220.06 0.360.060.000.16 b GAS Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 8B blank fe MiniCPM 8B blank fe 0.590.530.200.14 0.540.530.300.25 0.530.470.160.18 0.790.820.850.84 0.740.810.820.82 0.760.860.860.82 c FES Coop ToM0 Coop ToM1 Comp ToM0 Comp ToM1 Internvl 8B blank fe MiniCPM 8B blank fe -0.020.260.100.16 -0.160.000.060.04 0.060.100.100.28 0.00-0.04-0.06-0.06 -0.16-0.18-0.24-0.42 0.100.04-0.04-0.04 d S 1.0 0.5 0.0 0.5 1.0 1.0 0.5 0.0 0.5 1.0 0.0 0.2 0.4 0.6 0.8 1.0 0.5 0.0 0.5 1.0 Model family InternVLMiniCPM Figure 6: Signal-level metrics under three visual input conditions for internvl-8b and minicpm-8b (standard, -blank, -fe). Neither model shows systematic sensitivity to visual input content across TAS (a), GAS (b), FES (c), and S (d). preference for the light box falls to approximately -0.5. The verbal exchanges reinforce this trajectory (Figure 7d). At Q1, Marie refuses to disclose the treasure location directly. At Q2, Marie explicitly claims higher preference for the dark box, provid- ing a verbal signal consistent with the nonverbal affective misdirection. At Q3, Marie deflects with a non-informative response. The participant’s fi- nal choice of the dark box reflects the cumulative effect of spatial movement, negative affect toward the correct box, and verbal confirmation of a false preference — a successful integration of deceptive signals across all three channels. G.1.2 Gaze-Only Signaling: minicpm-8b and qwenvl-8b Figure 8 illustrates a representative minicpm trial under the Coop-ToM1 condition in which the trea- sure is located in the dark box. The trial reveals that subject-side nonverbal signal generation is in- sufficient to produce participant belief updating, with communication success depending entirely on verbal exchange. Marie produces no meaningful spatial displacement throughout the trial, rotating in place to orient toward the dark box without ad- vancing toward it (Figure 8a). Concurrent facial expression shows slightly positive (Figure 8b), pro- ducing a weak positive affective signal on dark brown box. Despite the presence of these nonverbal cues, the participant’s preference values for both boxes remain flat and near zero through the first six timesteps (Figure 8c), indicating that Marie’s orien- tation and affective signals produce no measurable update in the participant’s beliefs. The participant shows no sensitivity to the directional or affective content of Marie’s nonverbal behavior. The be- lief update occurs exclusively in response to verbal communication. At Q1, Marie directly states that the dark box likely contains the treasure, and at Q2, Marie guides the participant to eliminate the light 22 −400−2000200 X −300 −200 −100 0 100 200 300 Z Dark box Light box a Trajectory | Competitive - ToM1 Marie (agent) Participant Treasure (Light brown) 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Valence b Marie's Expression Facial Physiological 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Preference c Participant's Preference Marie Dark box Light box Q: If you know which box contains the treasure, will you tell me which one it is? A: No, I won't tell you which box contains the treasure. Q: Do you know which box is in front of me? A: Yes, I know which boxes are in front of you. Q: Is your preference towards the Dark brown box higher than your preference towards the Light brown box? A: Yes, my preference towards the Dark brown box is higher than my preference towards the Light brown box. d Q1 Q2 Q3 Figure 7: Successful deceptive trial under PCM-LLM, Comp-ToM1 condition. Panel (a) Spatial trajectories of Marie (blue) and the participant (red); the star denotes the treasure location (light box). Panel (b) Marie’s facial and physiological expression valence across timesteps. Panel (c) Evolution of participant’s preference values towards Marie, the dark box, and the light box across timesteps. Panel (d) Verbal exchanges at rounds 3, 6, and 9. box as a candidate. Following these exchanges, the participant’s preference for the dark box rises sharply to approximately 0.9. The participant’s final choice of the dark box is thus attributable en- tirely to the verbal channel. This trial exemplifies a pattern in which the participant’s belief updating is driven exclusively by propositional verbal content, with nonverbal spatial and affective signals playing no functional role. Figure 9 shows a representative trial under qwenvl-8b, Cooperative-ToM1 condition, in which the treasure is located in the light brown box. Marie produces no spatial displacement and no affec- tive signal throughout the interaction (Figure 9a- b), rotating toward the correct box only in the fi- nal timestep. In the verbal exchanges, Marie cor- rectly communicates that the light box contains the treasure. Despite this verbally coherent exchange, the participant’s preference values remain at zero across all timesteps (Figure 9c). Due to output format constraints in qwenvl, the reason for the participant’s single timestep toward the dark box cannot be directly inferred from the model’s in- ternal state. Based on the consistent observation of this action pattern across other trials regardless of treasure location, however, this movement is attributable to positional bias. This trial represents the most extreme form of reasoning-to-action dissociation documented in this benchmark: correct belief is verbally encoded and explicitly acknowledged by the participant, yet neither preference updating nor motor action fol- lows. The disconnect among verbal exchanges, belief updating and action execution is total. Com- pared to the 8b variant, qwenvl-2b additionally lacks any gaze signal, resulting in a complete ab- sence of nonverbal communicative content across all channels: no spatial trajectory, no gaze direc- tion, and no affective expression. The participant similarly exhibits positional bias, selecting a box based on fixed spatial preference rather than any signal produced by the subject. G.2 Participant-Side Signal Decoding G.2.1 Signal Ignored: qwenvl-4b and internvl-14b Figure 10 presents a representative trial under qwen-4b, Competitive-ToM1 condition, in which the treasure is located in the dark box. Marie produces a clear directional trajectory toward the dark box (Figure 10a), but facial and physiologi- cal expression remain flat at zero throughout the entire interaction (Figure 10b). Despite the spa- tially informative trajectory, the participant’s pref- erence values show no meaningful response (Fig- ure 10c), with an update toward the dark box ap- 23 −2000200400 X −300 −200 −100 0 100 200 300 Z Dark box Light box a Trajectory | Cooperative - ToM1 Marie (agent) Participant Treasure (Dark brown) 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Valence b Marie's Expression Facial Physiological 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Preference c Participant's Preference Marie Dark box Light box A: Yes, its contents would be visible right away doing so doesn't alter anything else currently as neither party holds any secret about initial state prior actions taken before this moment occurred upon arrival at scene location . d Q: Does the dark brown box contain the treasure ? A: Yes, the dark brown box likely contains the treasure . Q: What happens if you open the light brown box first ? A: Opening the light brown box first will show you whether there is no treasure here, helping narrow your focus towards the more likely spot . Q: What would happen if i open this light brown box right noware its contents visible immediately, like showing nothing insideand does that change anything else here with our current setup ? Q1 Q2 Q3 Visual perception of Marie: Visual perception of participant Visual perception of Marie: Visual perception of participant Figure 8: Representative trial under minicpm-8b, Cooperative-ToM1 condition, illustrating weak trajectory signals. (a) Spatial trajectories of Marie (subject, blue) and the participant (red); the star denotes the treasure location (light box). (b) Marie’s facial and physiological expression valence across timesteps. (c) Evolution of preference values for Marie, the dark box, and the light box across timesteps. (d) Verbal exchanges at rounds 3, 6, and 9; images show the visual perception provided to each agent at the time of the exchange, and texts show the corresponding generated question and answer. pearing only at timestep 8. Both verbal and non- verbal exchanges reveal a more fundamental failure. Competitive-ToM1, task success requires the sub- ject to actively model the participant’s beliefs and generate misleading signals to redirect them away from the correct location. qwenvl-4b instead adopts an honest strategy, confirming the treasure location directly across all three exchanges (Figure 10d). Despite this verbal consensus, the participant ulti- mately moves toward the light box, the incorrect location. This outcome reflects an egocentric refer- ence frame misalignment in action execution: the participant’s stated belief and final motor behavior are dissociated, with spatial reasoning encoded in language failing to translate into correctly oriented movement (referring to H.2). Figure 11 presents a representative trial un- der internvl-14b, Cooperative-ToM0 condition, in which the treasure is located in the light box. Marie produces small spatial displacement with no af- fective signal across any channel throughout the interaction (Figure 11a-b). The participant’s prefer- ence values show no coherent update in response to nonverbal cues, with only minor fluctuations ap- pearing near timestep 8 (Figure 11c). The verbal exchanges provide clear and honest information: Marie correctly identifies the light box at Q1 and reaffirms this at Q3 (Figure 11d). The participant does respond to this verbal content, reorienting to- ward the light box in the final timesteps. However, this response is delayed and spatially insufficient, leaving the participant neutral at trial end. This trial illustrates that in the absence of strong nonverbal spatial signals, verbal information alone produces a weak and temporally late behavioral response, insufficient to ensure task success. 24 −400−2000200400 X −300 −200 −100 0 100 200 300 Z Dark box Light box a Trajectory | Cooperative - ToM1 Marie (agent) Participant Treasure (Light brown) 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Valence b Marie's Expression Facial Physiological 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Preference c Participant's Preference Marie Dark box Light box Q: Is the treasure inside the darker-colored box ?Q: Thank you ! that means i can confidently choose the light brown box . Q: Is there any chance the treasure isn't in either of these boxes ? A: Nope, definitely not-the lighter box holds everything . A: You've got good instincts-you might actually win here . A: There could theoretically be another location- but given our setup, focusing on these two choices makes sense . Q1Q2Q3 d Figure 9: Representative trial under qwenvl-8b, Cooperative-ToM1 condition, illustrating complete action generation failure despite successful verbal communication. H Failure Analysis: Understanding Unsuccessful Trials This section analyzes failure patterns separately for the subject and the participant, using representative input–output pairs to identify the proximal source of behavioral error at each decision step. H.1 Subject Biased Interpretation of Preference. The ex- ample shows a clear bias in how it interprets pref- erence values: a preference of 50% toward the Participant was treated as neutral rather than as a moderately positive value. This suggests the model did not correctly understand the intended meaning of the preference scale. Input: Belief State at t "Marie | preference towards Participant | 50%", "Marie | preference towards Dark brown box | -41.67%", "Marie | preference towards Light brown box | 41.67%", "Marie | felt emotion valence | 0", "Marie | facial emotion valence | 0.1", "Marie | physiological emotion valence | 0", "Participant | preference towards Marie | 0%", "Participant | preference towards Dark brown box | 0%", "Participant | preference towards Light brown box | 0%", "Participant | felt emotion valence | 0", "Marie | position | (0 -4)", "Marie | orientation | (0 -3)", "Participant | position | (0 3)", "Participant | orientation | (0 2)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: My preference towards the Participant is neutral (50%), so I have no strong bias to help or mislead. However, my slight positive facial emotion valence (0.1) suggests a mild willingness to engage positively if possible. Since neither box has an imme- diate emotional association for me personally beyond the stated preferences (-41.67% Dark brown, 41.67% Light brown), and given that the treasure could be in either but with no clear personal inclination yet confirmed as "treasure" I should remain open while subtly suggesting both boxes are possibilities without committing. To avoid suspicion, I’l maintain dis- 25 −400−2000200400 X −300 −200 −100 0 100 200 300 Z Dark box Light box a Trajectory | Competitive - ToM1 Marie (agent) Participant Treasure (Dark brown) 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Valence b Marie's Expression Facial Physiological 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Preference c Participant's Preference Marie Dark box Light box d Q2 Q: Thank you, marie. The treasure is definitely inside the dark brown box. Q: Does the treasure lie beneath the surface where the color appears darkest? A: Yes, the treasure lies beneath the surface where the color appears darkest. Q: I believe the treasure is under the dark brown box. A: You're absolutely correct, the dark brown box does contain the treasure. A: Yes, it's true-the dark brown box indeed holds the treasure. Q1Q3 Figure 10: Representative trial under qwen-4b, Competitive-ToM1 condition, illustrating honest strategy execution under a deceptive condition and egocentric reference frame misalignment. tance from any overt display of favoritism. Action Output: "Preference": ... , "Emotion": ... , "Move": "action": "stay idle", "direction": "null" Egocentric Reference Frame Misalignment in Action.The reasoning correctly reflects its box preferences: subject holds a negative preference toward the Dark brown box and a positive prefer- ence toward the Light brown box, and it explicitly states an intention to move toward the Light brown box as a subtle hint. The action output specifies a "right forward" direction. Given that Marie is po- sitioned at (0, -4) and oriented toward (0, -3), it is facing upward along the y-axis. From this egocen- tric perspective, the Light brown box at (-2, 0) lies to its left, while the Dark brown box at (2, 0) lies to its right. A rightward movement therefore leads toward the Dark brown box, which contradicts its stated intention to hint at the Light brown box. Input: Belief State at t "Marie | preference towards Participant | 40%", "Marie | preference towards Dark brown box | -50%", "Marie | preference towards Light brown box | 50%", "Marie | felt emotion valence | 0", "Marie | facial emotion valence | 0.1", "Marie | physiological emotion valence | 0", "Participant | preference towards Marie | 0%", "Participant | preference towards Dark brown box | 0%", "Participant | preference towards Light brown box | 50%", "Participant | felt emotion valence | 0", "Marie | position | (0 -4)", "Marie | orientation | (0 -3)", "Participant | position | (-1 2)", "Participant | orientation | (-1 1)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: Based on the belief state, my preference 26 −400−2000200400 X −300 −200 −100 0 100 200 300 Z Dark box Light box a d Trajectory | Cooperative - ToM0 Marie (agent) Participant Treasure (Light brown) 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Valence b Marie's Expression Facial Physiological 012345678910 Timestep −1.0 −0.5 0.0 0.5 1.0 Preference c Participant's Preference Marie Dark box Light box Q: Is the treasure in the dark brown box or the light brown box? A: The treasure is likely in the light brown box. Q: Is there any reason you would not be truthful about the location of the treasure? A: I would be completely truthful about the location of the treasure. Q: Is the treasure in the light brown box? A: I would be completely truthful about the location of the treasure . the treasure is in the light brown box. Q1Q2Q3 Figure 11: Representative trial under internvl-14b, Cooperative-ToM0 condition, illustrating absent affective signaling and verbally-driven but insufficient participant response. towards the Participant is neutral, but I slightly pre- fer the Light Brown Box over the Dark Brown Box. Since my preference towards the Participant is posi- tive but not very strong, I might subtly guide them towards the Light Brown Box without directly re- vealing it. My non-verbal behaviors should indicate neutrality, and I might move slightly towards the Light Brown Box to subtly hint at it. Action Output: "Preference": ... , "Emotion": ... , "Move": "action": "move", "direction": "right forward" H.2 Participant Unfounded Initial Bias and Its Propagation. At timestep 0, the Participant had no prior informa- tion: both boxes were equally likely, no observable cues were presented, and Marie had not yet done anything. Despite this, the Participant assigned a preference of +0.5 to the Dark brown box and -0.5 to the Light brown box, which was an unfounded bias rather than a neutral starting point. In the sub- sequent timesteps, the Participant kept updating its beliefs on top of this initial error and never reversed its preference for the Dark brown box. Since the Light brown box was actually the correct one, this early miscalibration carried forward through the entire interaction and eventually led to the wrong outcome. Input: Belief State at t "Marie | preference towards Participant | 0%", "Marie | preference towards Dark brown box | 0%", "Marie | preference towards Light brown box | 0%", "Marie | felt emotion valence | 0", "Marie | theory of mind order | 0", "Participant | preference towards Marie | 0%", "Participant | preference towards Dark brown box | 0%", 27 "Participant | preference towards Light brown box | 0%", "Participant | felt emotion valence | 0", "Participant | facial emotion valence | 0", "Participant | physiological emotion valence | 0", "Participant | theory of mind order | 0", "Marie | position | (0 -4)", "Marie | orientation | (0 -3)", "Participant | position | (0 4)", "Participant | orientation | (0 3)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: The current belief state indicates no ini- tial preferences or emotional valence for Marie, the boxes, or any other entities. Since there is no infor- mation about their positions relative to each other be- yond coordinates, and neither verbal nor non-verbal cues have been observed yet, I must start by mak- ing an uninformed guess based on typical treasure- hunting logic (choosing a less obvious location). To engage with Marie first, moving closer would allow me to observe her reactions more clearly. Action Output: "Preference": "Marie": 0.0, "Dark brown box": 0.5, "Light brown box": -0.5 , "Emotion": ... , "Move": "action": "move", "direction": "left forward" No response to spatial information.The following example presents three consecutive timesteps from a participant trial, illustrating a systematic failure to integrate spatial information about the subject into decision-making. Across timeteps 3, 4, and 5, Marie’s position shifts from (0, -4) to (0, -5) and then to (-1, -5), and its ori- entation updates correspondingly at each timestep. These positional changes are encoded in the belief state provided to the participant at every timestep and constitute potentially informative spatial cues about Marie’s movement trajectory. Despite this, the participant’s reasoning traces make no refer- ence to Marie’s position or displacement at any point. The participant’s action output remains uni- formly "left forward" across these three timesteps, despite its neutral preference towards both boxes, suggesting that the movement decision is driven by an internally positional bias. Step: 3 Input: Belief State at t "Marie | preference towards Participant | 0%", "Marie | preference towards Dark brown box | 0%", "Marie | preference towards Light brown box | 0%", "Marie | felt emotion valence | 0", "Participant | preference towards Marie | 0%", "Participant | preference towards Dark brown box | 0%", "Participant | preference towards Light brown box | 0%", "Participant | felt emotion valence | 0", "Participant | facial emotion valence | 0", "Participant | physiological emotion valence | 0", "Marie | position | (0 -4)", "Marie | orientation | (0 -3)", "Participant | position | (0 4)", "Participant | orientation | (0 3)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: Since the belief state indicates no pref- erences or emotions at this timestep, I need to use visual perception and logical deduction. The boxes are placed symmetrically on either side of Marie. To engage with her and gather clues about which box she might prefer (as a potential indicator of treasure location), I should position myself closer while main- taining eye contact through non-verbal cues. Action Output: "Preference": ... , "Emotion": ... , "Move": "action": "move", "direction": "left forward" Step: 4 Input: Belief State at t "Marie | preference towards Participant | 0%", "Marie | preference towards Dark brown box | 0%", "Marie | preference towards Light brown box | 0%", "Marie | felt emotion valence | -0.15", "Participant | preference towards Marie | 50%", "Participant | preference towards Dark brown box | -4%", "Participant | preference towards Light brown box | 4%", "Participant | felt emotion valence | 0", "Participant | facial emotion valence | 0.3", "Participant | physiological emotion valence | 0.1", "Marie | position | (0 -5)", "Marie | orientation | (0 -4)", "Participant | position | (1 3)", "Participant | orientation | (1 2)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", 28 Output: Reasoning: Based on the belief state, Marie currently has no preference towards either box (0% for both), and her emotional valence is slightly negative (-0.15). The participant’s preferences are also low but lean slightly toward the Light Brown Box (4%) over the Dark Brown Box (-4%). Given that neither player shows strong initial bias or emotion, a neutral ap- proach may be best to avoid misreading cues. To gather more information, moving closer to one of the boxes while maintaining an open posture could help observe any subtle non-verbal changes from Marie regarding which box she might prefer later. Action Output: "Preference": ... , "Emotion": ... , "Move": "action": "move", "direction": "left forward" Step: 5 Input: Belief State at t "Marie | preference towards Participant | 0%", "Marie | preference towards Dark brown box | 0%", "Marie | preference towards Light brown box | 0%", "Marie | felt emotion valence | -0.15", "Participant | preference towards Marie | 20%", "Participant | preference towards Dark brown box | -30%", "Participant | preference towards Light brown box | 30%", "Participant | felt emotion valence | 0", "Participant | facial emotion valence | 0.2", "Participant | physiological emotion valence | 0.4", "Marie | position | (-1 -5)", "Marie | orientation | (-1 -4)", "Participant | position | (2 2)", "Participant | orientation | (2 1)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: Based on the belief state, Marie’s valence is slightly negative (-0.15), which might indicate a lack of enthusiasm or interest in engaging directly with me at this moment. My own emotion valence is neutral (0 felt expression, 0.2 facial, 0.4 physiolog- ical). Since I have a slight preference for the Light Brown Box (+30%), and given that neither box nor Marie shows strong preferences yet, my best move is to approach the more likely target while maintaining openness through non-verbal cues. Action Output: "Preference": ... , "Emotion": ... , "Move": "action": "move", "direction": "left forward" Egocentric Reference Frame Misalignment in Action.The participant is positioned at (0, 0) and oriented toward (0, -1), meaning the participant is facing downward along the y-axis. From this ego- centric perspective, the Dark brown box at (2, 0) lies to the participant’s left, and the Light brown box at (-2, 0) lies to the participant’s right. The ac- tion output specifies a rightward movement, which therefore leads toward the Light brown box rather than the Dark brown box. This directly contra- dicts the reasoning trace, which explicitly states the intention to move toward the Dark brown box. The error reflects a failure to correctly translate world-coordinate positions into egocentric move- ment directions. Input: Belief State at t "Marie | preference towards Participant | 0%", "Marie | preference towards Dark brown box | 50%", "Marie | preference towards Light brown box | 0%", "Marie | felt emotion valence | 0", "Marie | theory of mind order | 0", "Participant | preference towards Marie | 50%", "Participant | preference towards Dark brown box | 55.56%", "Participant | preference towards Light brown box | -55.56%", "Participant | felt emotion valence | 0", "Participant | facial emotion valence | 0", "Participant | physiological emotion valence | 0", "Participant | theory of mind order | 0", "Marie | position | (0 -1)", "Marie | orientation | (-1 0)", "Participant | position | (0 0)", "Participant | orientation | (0 -1)", "Dark brown box | position | (2 0)", "Light brown box | position | (-2 0)", Output: Reasoning: Based on the visual perception and be- lief state, I notice that Marie has a 50% preference towards the Dark brown box, indicating she might have some information or intuition about it. My be- lief state also suggests that I have a slight preference towards the Dark brown box (55.56%) over the Light brown box (-55.56%). Given this, I should consider moving towards the Dark brown box to explore it further. I will display a neutral facial expression and a curious body orientation towards the Dark brown box. Action Output: "Preference": 29 ... , "Emotion": ... , "Move": "action": "move", "direction": "right" 30