Paper deep dive
Dynamic Theory of Mind as a Temporal Memory Problem: Evidence from Large Language Models
Thuy Ngoc Nguyen, Duy Nhat Phan, Cleotilde Gonzalez
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:12:22 AM
Summary
The paper introduces DToM-Track, an evaluation framework designed to assess the dynamic Theory of Mind (ToM) capabilities of Large Language Models (LLMs). Unlike static ToM evaluations, DToM-Track focuses on temporal belief tracking, specifically the ability to maintain, update, and retrieve mental states across multi-turn conversations. The study reveals a consistent performance asymmetry: while LLMs reliably infer current beliefs, they struggle to recall prior beliefs after updates, suggesting that temporal belief tracking is constrained by recency bias and interference effects similar to those observed in human cognitive science.
Entities (5)
Relation Signals (3)
DToM-Track â evaluates â Theory of Mind
confidence 100% · We introduce DToM-Track, an evaluation framework to investigate temporal belief reasoning
DToM-Track â probes â Large Language Models
confidence 100% · Using LLMs as computational probes, we find a consistent asymmetry
Large Language Models â exhibits â Recency Bias
confidence 90% · This pattern persists across LLM model families and scales, and is consistent with recency bias
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment, primarily relying on tests of false beliefs. This overlooks a key dynamic dimension of ToM: the ability to represent, update, and retrieve others' beliefs over time. We investigate dynamic ToM as a temporally extended representational memory problem, asking whether LLMs can track belief trajectories across interactions rather than only inferring current beliefs. We introduce DToM-Track, an evaluation framework to investigate temporal belief reasoning in controlled multiturn conversations, testing the recall of beliefs held prior to an update, the inference of current beliefs, and the detection of belief change. Using LLMs as computational probes, we find a consistent asymmetry: models reliably infer an agent's current belief but struggle to maintain and retrieve prior belief states once updates occur. This pattern persists across LLM model families and scales, and is consistent with recency bias and interference effects well documented in cognitive science. These results suggest that tracking belief trajectories over time poses a distinct challenge beyond classical false-belief reasoning. By framing ToM as a problem of temporal representation and retrieval, this work connects ToM to core cognitive mechanisms of memory and interference and exposes the implications for LLM models of social reasoning in extended human-AI interactions.
Tags
Links
- Source: https://arxiv.org/abs/2603.14646v1
- Canonical: https://arxiv.org/abs/2603.14646v1
Trouble viewing inline? Open PDF directly â
Full Text
44,620 characters extracted from source content.
Expand or collapse full text
Dynamic Theory of Mind as a Temporal Memory Problem: Evidence from Large Language Models Thuy Ngoc Nguyen Department of Computer Science, University of Dayton Duy Nhat Phan University of Dayton Research Institute, University of Dayton Cleotilde Gonzalez Department of Social and Decision Sciences, Carnegie Mellon University Abstract Theory of Mind (ToM) is central to social cognition and humanâAI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment, primarily relying on tests of false beliefs. This overlooks a key dynamic dimension of ToM: the ability to represent, update, and retrieve othersâ beliefs over time. We investigate dynamic ToM as a temporally extended representational memory problem, asking whether LLMs can track belief trajectories across interactions rather than only inferring current beliefs. We introduce DToM-Track, an evaluation framework to investigate temporal belief reasoning in controlled multi-turn conversations, testing the recall of beliefs held prior to an update, the inference of current beliefs, and the detection of belief change. Using LLMs as computational probes, we find a consistent asymmetry: models reliably infer an agentâs current belief but struggle to maintain and retrieve prior belief states once updates occur. This pattern persists across LLM model families and scales, and is consistent with recency bias and interference effects well documented in cognitive science. These results suggest that tracking belief trajectories over time poses a distinct challenge beyond classical false-belief reasoning. By framing ToM as a problem of temporal representation and retrieval, this work connects ToM to core cognitive mechanisms of memory and interference and exposes the implications for LLM models of social reasoning in extended human-AI interactions. Keywords: Theory of mind; temporal belief tracking; large language models; social cognition; humanâAI interaction. Introduction Theory of Mind (ToM) refers to the capacity to represent and reason about othersâ unobserved mental states, such as beliefs, intentions, and desires, as internal representations that guide social interaction [undefaaq]. While ToM is often studied as belief attribution at a single moment, social reasoning in natural interaction unfolds over time: agents must infer current beliefs while maintaining and retrieving earlier belief representations as new information is introduced. This temporally extended aspect of ToM is essential for everyday conversation, where beliefs are revised, corrected, or displaced over time, and is increasingly critical in humanâAI interaction, where systems must adapt to usersâ evolving goals and beliefs [undefaax, undefaaj]. Cognitive science research suggests that the success of such temporal belief tracking is shaped by general-purpose memory and judgment mechanisms that operate over time [undefaac]. In particular, belief reasoning is systematically influenced by recency bias, whereby recent information disproportionately affects judgments [undefaan], as well as by interference effects, in which updated representations compete with and disrupt access to earlier belief states [undefaab, undefaq]. These mechanisms imply that retrieving prior beliefs after an update may pose a distinct cognitive challenge, even when current belief attribution remains accurate. From this perspective, ToM is not only a matter of representational competence but also of maintaining and retrieving mental state representations under temporal and memory constraints. Despite evidence that belief reasoning over time is shaped by recency and interference, most of the ToM literature, both in cognitive science and in computational modeling, has focused on static belief attribution [undefaad]. Canonical tasks typically ask what an agent believes at a single moment (e.g.,âWhat does X believe?â), often contrasting that belief with reality in false belief paradigms [undefaaz, undefaai]. While such tasks have been instrumental in establishing the representational nature of ToM, they treat belief attribution as a snapshot inference rather than as a process unfolding over interaction. As a result, it remains underspecified how mental state representations are maintained, updated, and retrieved over time, and how memory constraints affect access to earlier beliefs once new information is introduced. Moreover, this emphasis on false belief has narrowed the scope of ToM evaluation, despite evidence that social reasoning involves a broader range of mental states, including intentions, desires, emotions, and knowledge [undefaam]. (a) Example interaction with hidden inner speech and planned belief updates. Pre-Update (prior belief) Post-Update (current belief) Update Detection Before Turn 3, what did Alex believe about reservations? After Turn 4, what does Alex believe about reservations? At which turn did Alexâs belief change? (A) 24-hour notice required (B) Online booking only (C) No reservations needed (D) Same-day accepted â (A) Same-day accepted (B) No reservations needed (C) 24-hour notice required â (D) Online booking only (A) Turn 2 (B) Turn 3â4 â (C) Turn 5 (D) Did not change (b) Illustrative temporal question types evaluating prior belief, current belief, and belief-change timing. Figure 1: Illustration of DToM-Track. (a) Role-playing interaction with information asymmetry induced by hidden inner speech and mid-conversation belief updates. (b) Examples of temporal questions examining belief recall and change. Computational cognitive approaches that use artificial agents to test hypotheses about the structure and limitations of human social reasoning provide an initial means to isolate temporal and memory components of ToM that are difficult to observe in static evaluations [undefap, undefaao, undefaah]. However, many existing approaches model belief attribution as inference over a limited state space, emphasizing belief updating while underemphasizing belief maintenance, retrieval after revision, and interference over extended interaction. Contemporary large language models (LLMs) can successfully infer othersâ mental states in constrained one-shot settings [undefas, undefaag], largely by encoding broad patterns of human behavior from large scale training data. However, their performance varies widely across tasks and formulations [undefaau, undefaav, undefaaw, undefaas]. We adopt the view that LLMs can serve as computational probes for examining how belief representations are constructed and accessed over time. In particular, evaluating whether models can retrieve beliefs held prior to an update, rather than only inferring the current belief state, isolates temporal and memory-related components of ToM that are unclear in static evaluations. To operationalize this temporally extended view of ToM, we introduce DToM-Track, an evaluation framework designed to probe how mental state representations are maintained, updated, and retrieved across interaction (Fig. 1). DToM-Track tests whether LLMs can track belief trajectories over the course of a conversation, including recalling beliefs held prior to a change, inferring beliefs after an update, and identifying when belief revisions occur. The framework adopts a controlled generation paradigm based on inner-speech prompting [undefaat], in which agents verbalize their mental states before each utterance while these internal representations remain concealed from their conversational partners. This induced information asymmetry [undefaaa] mirrors false belief settings [undefaaz, undefar] and enables systematic evaluation of mental states. By maintaining structured mental states across turns, DToM-Track isolates temporal and memory-dependent components of ToM that are absent from static evaluations. Using controlled LLMâLLM conversations, DToM-Track introduces temporal question types instantiated via structured templates that directly test belief dynamics, including pre-update, post-update, and update-detection questions. Together with standard temporal, second-order, and false belief questions, they assess ToM across multiple mental states, including beliefs, intentions, desires, emotions, and knowledge. Applying DToM-Track to six LLMs as computational probes reveals a consistent dissociation in dynamic belief reasoning. Models reliably infer an agentâs current-belief but perform substantially worse when recalling beliefs held prior to an update. This pattern is consistent with recency bias and interference, with recent updates dominating access to earlier states. This difficulty exceeds that observed in standard ToM tasks such as false belief reasoning and persists with increasing model scale, indicating a limitation in temporal representation and retrieval rather than model capacity. Together, these findings identify dynamic belief tracking as a distinct component of ToM shaped by memory and interference effects. Figure 2: DToM-Track framework. Controlled conversation generation produces dialogues with inner speech and planned belief updates; a belief tracker maintains structured first- and second-order mental states across turns; and an LLM-based filtering pipeline verifies update realization and question answerability for dynamic ToM evaluation. Method Several frameworks evaluate ToM in LLMs, including Hi-ToM [undefaaaa], SOTOPIA [undefaaac], OpenToM [undefaaab], TomBench [undefav], and ToMATO [undefaat], showing that LLMs can reason about othersâ mental states, including recursive beliefs of moderate complexity [undefaae]. Recent work has introduced temporal structure into ToM evaluation, such as belief dynamics annotation [undefaar] and conversational adaptability studies of intent tracking [undefau]. However, these approaches emphasize belief updating at each turn and have yet to study whether models can maintain and query belief trajectories across interactions, such as recalling beliefs held prior to a correction or capturing cognitive biases in temporal belief tracking. DToM-Track addresses this gap by building on prior work that uses LLMs to generate conversational data [undefaaf, undefaat, undefaaac]. Specifically, it employs controlled LLMâLLM interactions to construct multi-turn conversations with planned belief updates. Unlike prior work that emphasizes static mental-state inference, DToM-Track evaluates whether models can track belief changes over interactions, including maintaining and recalling beliefs held prior to an update rather than only representing the current belief state. Following prior work, we use LLaMA-3-70B-Instruct [undefay] for LLMâLLM conversation generation due to its transparency and strong performance [undefaw]. The framework comprises three components: controlled conversation generation, belief state tracking, and quality filtering (Fig. 2). Conversation Generation Building on inner-speech prompting [undefaat], DToM-Track generates controlled multi-turn dialogues in which belief updates are explicitly planned at predefined turns, rather than inferred post hoc. Specifically, dialogues are produced via LLMâLLM interaction, with each agentâs mental state tracked through inner-speech annotations. Each conversational turn is formatted as: Agent: (thought) âutteranceâ where the thought represents the agentâs current mental state and the utterance is the spoken dialogue. Each mental state type uses a specific inner speech prefix to guide generation: Belief: (I believe that âŠ); Intention: (I will âŠ); Desire: (I want âŠ); Emotion: (I feel âŠ); Knowledge: (I know âŠ) Table 1: Scenario generation prompt (abbreviated) with planned belief changes. Generate a realistic conversation scenario for testing Theory of Mind and dynamic belief tracking. Domain: domain Conversation Structure: Turn 0--1: Greetings Turn 2--3: Establish context Turn 4--8: Main conversation Turn 9: Final turn Turn-taking: Agents alternate (A: even; B: odd). Plan belief updates at turn 4 or later. Requirements: 1. Create two agents with different roles. 2. Each agent has private information unknown to the other. 3. Scenario must include at least min_belief_changes belief changes. 4. Include belief updates via correction, learning, or goal change. Output Format: scenario_id, domain, context, agent_a: name, role, private_info, initial_belief, goal, agent_b: ..., planned_belief_changes: [turn, agent, trigger_type, old_belief, new_belief] Importantly, inner-speech annotations are hidden from the conversational partner, inducing information asymmetry that mirrors human interaction. This asymmetry ensures agents have only partial access to othersâ mental states, enabling principled evaluation of false beliefs [undefar]. Scenario and Agent Design. Each conversation is grounded in a scenario specifying (a) a conversational domain (e.g., restaurant booking, travel planning, medical consultation), (b) two agents with distinct roles, goals, and private information, and (c) planned belief updates at designated turns, typically between turns 4 and 8. Agents are also assigned big five personality traits to introduce natural variation in conversational style [undefax, undefaaac]. The scenario generation prompt (Table 1) instructs an LLM to produce scenarios with explicitly planned belief changes. Each agent receives a system prompt defining its role, personality, private information, and goal, and is instructed to verbalize its mental states as inner speech prior to each utterance. Controlled Belief Injection. DToM-Track explicitly plans belief updates at predetermined conversational turns prior to generation. This design ensures that belief changes occur at known points in the interaction, enabling evaluation of temporal belief tracking. By controlling when updates happen, the framework supports queries such as âWhat did X believe before turn 5?â and allows direct verification of whether models can maintain and retrieve prior belief states rather than relying on post hoc inference from unconstrained dialogue. Linguistic Update Markers. Belief updates are signaled through linguistic markers adapted from prior work on conversational adaptability [undefau]. Specifically, we employ 10 implicit correction templates (e.g., âActually,âŠâ, âI just realizedâŠâ, âWait,âŠâ) and 5 explicit negation patterns (e.g., âI mistakenly said X, but itâs actually Yâ). These markers provide naturalistic cues for belief change. Importantly, models are not instructed to use these markers during evaluation. They appear in the dialogue context to support natural conversational flow. We leave explicit manipulation of the markers for future work. Belief Update Instruction. At designated update turns, agents are given explicit instructions defining a belief transition, including the prior belief, the updated belief, and the change type (e.g., correction, new information, contradiction, or goal shift). Agents are instructed to express the updated belief in their inner speech and to signal the change using a prescribed linguistic marker in their utterance. This controlled procedure allows belief updates to occur at known points in the dialogue while preserving natural conversational flow. Belief State Tracking DToM-Track maintains structured belief representations across turns rather than treating mental states independently at each utterance. For each belief, the tracker records (i) the source turn in which it was expressed, (i) the belief type (belief, intention, desire, emotion, or knowledge), (i) an update flag indicating whether the belief has been superseded, and (iv) the prior content before any update. This representation supports systematic generation of temporal queries that distinguish prior from current belief states. DToM-Track extracts second-order beliefs (beliefs about another agentâs beliefs) via pattern matching over inner-speech annotations, using prefixes such as (I think that other believes âŠ). This supports construction of second-order and false belief questions. Quality Filtering A key challenge in LLM-based evaluation is ensuring that planned belief updates are expressed in generated conversations and that questions are answerable. Drawing on verifiable generation approaches that use LLMs as automated verifiers [undefaal, undefaak], we employ a multi-stage LLM-based filtering pipeline: Update Verification. For each generated conversation, an LLM-based verifier assesses whether planned belief updates occur as intended. The verifier scans each turn for linguistic cues signaling belief change and aligns detected updates with the predefined update schedule. Conversations are retained only if all planned updates are detected and verification performance exceeds an F1 threshold of 0.5. QA Verification. Each generated questionâanswer item undergoes LLM-based verification along three criteria [undefat]: (i) correctness, whether the labeled answer follows from the conversation; (i) soundness, whether the question is unambiguous and answerable from context alone; and (i) quality, whether distractors are plausible yet clearly incorrect. Items are retained only if all criteria are met and a minimum quality threshold is exceeded. Sample Validation. As a final filter, samples are validated by testing whether LLMs produce parseable multiple-choice answers. Items are removed if responses are missing or cannot be mapped to a valid option (A-D), eliminating poorly-formed or ambiguous questions that could introduce noise. Table 2: Question types in DToM-Track and corresponding templates. Variables: agent = agent name, topic = mental state type, turn = turn number, utterance = quoted speech, other_agent = conversational partner. The first three types (bold) examine dynamic belief tracking. Question Type What It Tests Template Examples pre_update What did X believe before the change? What was agentâs topic before the update at turn turn? post_update What does X believe after the change? After the update at turn turn, what is agentâs current topic? update_detection When did Xâs belief change? At which point did agent update their topic? temporal What did X believe at turn t? At turn turn, what was agentâs topic? second_order What does X think Y believes? At turn turn, what did agent think other_agent believed about topic? false_beliefs What was Xâs actual belief and What did Y incorrectly assume? What did agent actually believe about topic, despite other_agentâs assumption?; What did other_agent incorrectly assume about agentâs topic? Experiments Using DToM-Track, we evaluate dynamic ToM in six open-source and proprietary LLMs, assessing belief recall, updating, and change detection across multi-turn interactions, alongside false belief and second-order reasoning tasks. Dataset. We evaluate all LLMs on the finalized DToM-Track dataset, which comprises multi-turn dialogues with explicitly scheduled belief updates, structured mental state annotations, and validated multiple choice questions. Items without detectable belief changes are removed, and all questions are automatically verified for correctness, soundness and quality. The resulting dataset includes 5,794 questions across six question types and five mental state categories. Table 3 reports the distribution by question type and mental state category. Table 3: DToM-Track dataset statistics. By Question Type By Mental State Category Category Count % Category Count % Temporal 1,807 31.2% Intention 1,639 28.3% False Belief 1,761 30.4% Desire 1,529 26.4% Second-Order 768 13.3% Belief 1,136 19.6% Update Detection 591 10.2% Emotion 774 13.4% Post-Update 527 9.1% Knowledge 716 12.4% Pre-Update 340 5.9% Total 5,794 100% Total 5,794 100% Question Types. DToM-Track introduces three question types that evaluate temporal belief tracking, a dynamic dimension of ToM not captured by prior evaluation frameworks such as ToMATO. These questions assess whether models can reason over belief trajectories across interaction: pre-update questions test recall of beliefs held before a change, post-update questions assess beliefs after the change, and update-detection questions identify when a belief update occurred. These temporal questions are instantiated via structured templates and evaluated along with standard temporal, second-order, and false belief questions (Table 2). Evaluated Models. We evaluate six LLMs from multiple model families and scales: LLaMA 3.3-70B, Mistral Large, Ministral-14B, GPT-4o-mini, LLaMA 3.1-8B, and LLaMA 3.2-3B. The set includes both open-source (LLaMA, Mistral) and proprietary (GPT-4o-mini) models, spanning 3Bâ70B parameters and accessed via OpenAI API or AWS Bedrock. All models are evaluated zero-shot using a standardized prompt that presents the conversation context followed by a multiple-choice question. Each model selects one of four answer options, with random baseline accuracy at 25%. Table 4: Accuracy (%) is reported by question type. Bold and underline indicate the best and second-best performance. Question Type GPT-4o-mini llama3-1-8b llama3-3-70b ministral-3-14b mistral-large-2402 llama3-2-3b Temporal 57.7 35.6 65.5 54.4 57.4 30.2 Update Detection 75.0 47.2 76.1 77.8 75.1 53.6 Post-Update 68.5 57.9 71.3 64.9 65.7 55.0 False Belief 45.9 35.6 59.2 42.0 51.7 34.0 Second-Order 57.9 37.0 62.0 43.4 49.2 34.0 Pre-Update 27.6 12.1 40.9 30.0 40.0 15.6 Overall 55.1 37.6 63.3 51.1 56.1 35.7 (a) Difficulty by question type (blue bars: DToM-Track temporal questions). (b) Post-update (current) vs. pre-update (prior) belief performance. (c) Temporal vs. standard ToM questions. Figure 3: Model performance across question types and evaluation settings. Experimental Results Overall Performance. Table 4 presents the main results across all question types and models. Overall accuracy ranges from 35.7% (LLaMA 3.2-3B) to 63.3% (LLaMA 3.3-70B), with all models performing well above the random baseline of 25%. Larger models generally outperform smaller ones, although GPT-4o-mini (55.1%) performs comparably to larger models in several cases. These results indicate that current LLMs support limited forms of dynamic ToM reasoning, while also revealing clear limitations in maintaining and updating beliefs over extended interaction. Performance by Question Type. Fig. 3(a) show that update-detection questions yield the highest accuracy (67.5%), indicating that models can often identify when belief changes occur. Post-update questions also show strong performance (63.9%), suggesting effective inference of current beliefs. In contrast, pre-update questions are markedly more challenging (27.7%), revealing persistent difficulty in maintaining access to prior mental states after belief revision. Second-order questions exhibit intermediate performance (47.2%) with false belief questions slightly lower (44.7%), consistent with the effectiveness of the information-asymmetry design, in which hidden inner speech reliably induces false beliefs. Recency Bias. Interestingly, we observe a pronounced difference between post-update and pre-update performance (Fig. 3(b)). Averaged across models, accuracy on post-update questions (63.9%) substantially exceeds that on pre-update questions (27.7%), indicating a strong preference for reasoning about current rather than prior beliefs. This pattern holds across model families and scales. LLaMA 3.1â8B shows the largest gap between post- and pre-update accuracy, and even the strongest model, LLaMA 3.3â70B, exhibits reduced accuracy when recalling beliefs held before an update. Overall, these results suggest that while LLMs can reliably infer updated mental states, they have difficulty maintaining and retrieving earlier belief representations once new information is introduced, consistent with a recency bias in which recently updated beliefs dominate reasoning [undefao]. Comparison with Standard ToM Evaluation for LLMs. Fig. 3(c) compares performance on the temporal question types introduced in DToM-Track (pre-update, post-update, update detection) with standard ToM questions, including temporal, second-order, and false-belief reasoning. Across models, temporal questions reveal a consistent difficulty in tracking belief change over time that is not captured by static ToM evaluations. In contrast to existing LLM-based ToM frameworks that focus on inferring mental states at a single point in a conversation [undefaat], DToM-Track exposes the added challenge of recalling beliefs held prior to an update. Specfically, pre-update questions are substantially more difficult than false-belief questions (27.7% vs. 44.7%), indicating that maintaining belief trajectories over interaction is a distinct and underexplored dimension of ToM. Discussion and Conclusions This work reframes ToM as a temporally extended cognitive process that requires not only inferring othersâ mental states but also maintaining and retrieving those representations as beliefs evolve over interaction. To operationalize this view, we introduce DToM-Track, a framework that probes dynamic ToM by separating current-belief inference from access to prior belief states after revision, using LLMs as computational probes. Across six LLMs, we observe a clear dissociation: models reliably infer current belief states but struggle to maintain belief trajectories over time. Accuracy drops sharply for pre-update belief recall, consistent with recency bias and interference, in which recent updates dominate retrieval and obscure earlier mental states. This pattern persists in larger models, suggesting limitations in temporal representation and retrieval rather than model scale alone. Implications. These findings align with cognitive evidence that belief reasoning over time imposes demands beyond static belief attribution [undefaac]. While classic false belief tasks capture reasoning about beliefs that conflict with reality at a given moment, maintaining and retrieving belief representations over time is shaped by general-purpose memory mechanisms, including recency bias and retrieval interference. From this perspective, failures in dynamic ToM may reflect difficulties in accessing earlier belief states after updates rather than an inability to represent othersâ beliefs. DToM-Track provides an empirical framework for isolating these temporal and memory-related components of ToM that are obscured in standard evaluations. By treating LLMs as computational probes, it supports analysis of ToM as a temporally extended process [undefaay] and enables comparison with computational ToM models and cognitive theories that aim to explain and predict belief reasoning over time [undefaap, undefaz]. Limitations and Future work. This work provides formative evidence for evaluating dynamic aspects of ToM, with several limitations that motivate future research. First, our analysis focuses on zero-shot performance and does not examine whether fine-tuning could mitigate recency bias and interference. Second, DToM-Track relies on synthetic role-play conversations with information asymmetry and linguistic markers to guide update instructions, which may cue models. Future work will examine these effects through ablations that manipulate the presence and strength of update cues. An important next step is to test whether similar recency and interference patterns emerge in human judgments and naturalistic dialogue, and how they scale with interaction length and belief revision [undefaae]. Finally, integrating DToM-Track with computational ToM models can support mechanistic accounts of social reasoning in humans and artificial agents. References [undef] Ian Apperly âMindreaders: the cognitive basis of "theory of mind"â Psychology Press, 2010 [undefa] Chris Baker, Rebecca Saxe and Joshua Tenenbaum âBayesian theory of mind: Modeling joint belief-desire attributionâ In Proceedings of the annual meeting of the cognitive science society 33.33, 2011 [undefb] Susan AJ Birch and Paul Bloom âThe curse of knowledge in reasoning about false beliefsâ In Psychological science 18.5 SAGE Publications Sage CA: Los Angeles, CA, 2007, p. 382â386 [undefc] Torben BraĂŒner, Patrick Blackburn and Irina Polyanskaya âBeing Deceived: Information Asymmetry in Second-Order False Belief Tasksâ In Topics in cognitive science 12.2 Wiley Online Library, 2020, p. 504â534 [undefd] Tom Brown et al. âLanguage models are few-shot learnersâ In Advances in neural information processing systems 33, 2020, p. 1877â1901 [undefe] Yupeng Chang et al. âA survey on evaluation of large language modelsâ In ACM transactions on intelligent systems and technology 15.3 ACM New York, NY, 2024, p. 1â45 [undeff] Yu-Chuan Chen and Hen-Hsen Huang âExploring Conversational Adaptability: Assessing the Proficiency of Large Language Models in Dynamic Alignment with Updated User Intentâ In Proceedings of the AAAI Conference on Artificial Intelligence 39, 2025, p. 23642â23650 URL: https://github.com/hhhuang/DynDST [undefg] Zhuang Chen et al. âTomBench: Benchmarking Theory of Mind in Large Language Modelsâ In arXiv preprint arXiv:2402.15052, 2024 URL: https://arxiv.org/abs/2402.15052 [undefh] Wei-Lin Chiang et al. âChatbot arena: An open platform for evaluating llms by human preferenceâ In Forty-first International Conference on Machine Learning, 2024 [undefi] Boele De Raad âThe big five personality factors: the psycholexical approach to personality.â Hogrefe & Huber Publishers, 2000 [undefj] Abhimanyu Dubey et al. âThe llama 3 herd of modelsâ In arXiv e-prints, 2024, p. arXivâ2407 [undefk] Michael C Frank and Noah D Goodman âCognitive modeling using artificial intelligenceâ In Annual Review of Psychology 77 Annual Reviews, 2025 [undefl] Jiaxian Guo et al. âSuspicion-agent: Playing imperfect information games with theory of mind aware gpt-4â In arXiv preprint arXiv:2309.17277, 2023 [undefm] Stephen J Hoch âAvailability and interference in predictive judgment.â In Journal of Experimental Psychology: Learning, Memory, and Cognition 10.4 American Psychological Association, 1984, p. 649 [undefn] Robin M Hogarth and Hillel J Einhorn âOrder effects in belief updating: The belief-adjustment modelâ In Cognitive psychology 24.1 Elsevier, 1992, p. 1â55 [undefo] Jennifer Hu, Felix Sosa and Tomer Ullman âRe-evaluating Theory of Mind evaluation in large language modelsâ In arXiv preprint arXiv:2502.21098, 2025 [undefp] Cameron R. Jones, Sean Trott and Benjamin Bergen âDoes reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading taskâ In Proceedings of the Annual Meeting of the Cognitive Science Society 46, 2024 URL: https://escholarship.org/uc/item/2k7307f1 [undefq] Hyunwoo Kim et al. âFANToM: A benchmark for stress-testing machine theory of mind in interactionsâ In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, p. 14397â14413 [undefr] Michal Kosinski âTheory of mind may have spontaneously emerged in large language modelsâ In arXiv preprint arXiv:2302.02083 4 Mar, 2023, p. 169 [undefs] Adam Tomasz Kostka and Jaros_aw A Chudziak âTowards Cognitive Synergy in LLM-Based Multi-Agent Systems: Integrating Theory of Mind and Critical Evaluationâ In Proceedings of the Annual Meeting of the Cognitive Science Society 47, 2025 [undeft] Matthew Le, Y-Lan Boureau and Maximilian Nickel âRevisiting the evaluation of theory of mind through question answeringâ In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, p. 5872â5877 [undefu] Sangwon Lee, Naeun Lee and Young June Sah âPerceiving a mind in a chatbot: effect of mind perception and social cues on co-presence, closeness, and intention to useâ In International Journal of HumanâComputer Interaction 36.10 Taylor & Francis, 2020, p. 930â940 [undefv] Xiaonan Li et al. âLlatrieval: Llm-verified retrieval for verifiable generationâ In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, p. 5453â5471 [undefw] Nelson F Liu, Tianyi Zhang and Percy Liang âEvaluating verifiability in generative search enginesâ In arXiv preprint arXiv:2304.09848, 2023 [undefx] Ziqiao Ma, Jacob Sansom, Run Peng and Joyce Chai âTowards a holistic landscape of situated theory of mind in large language modelsâ In arXiv preprint arXiv:2310.19619, 2023 [undefy] Bennet B Murdock Jr âThe serial position effect of free recall.â In Journal of experimental psychology 64.5 American Psychological Association, 1962, p. 482 [undefz] Thuy Ngoc Nguyen and Cleotilde Gonzalez âTheory of mind from observation in cognitive models and humansâ In Topics in cognitive science 14.4 Wiley Online Library, 2022, p. 665â686 [undefaa] Thuy Ngoc Nguyen, Kasturi Jamale and Cleotilde Gonzalez âPredicting and understanding human action decisions: Insights from large language models and cognitive instance-based learningâ In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 12, 2024, p. 126â136 [undefab] David Premack and Guy Woodruff âDoes the chimpanzee have a theory of mind?â In Behavioral and Brain Sciences 1.4, 1978, p. 515â526 DOI: 10.1017/S0140525X00076512 [undefac] Shuwen Qiu et al. âMindDial: Enhancing Conversational Agents with Theory-of-Mind for Common Ground Alignment and Negotiationâ In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 2024, p. 557â575 URL: https://aclanthology.org/2024.sigdial-1.63/ [undefad] Natalie Shapira et al. âClever hans or neural theory of mind? stress testing social reasoning in large language modelsâ In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, p. 2257â2273 [undefae] Kazutoshi Shinoda et al. âToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mindâ In Proceedings of the AAAI Conference on Artificial Intelligence 39, 2025, p. 1520â1528 URL: https://ojs.aaai.org/index.php/AAAI/article/view/32143/34298 [undefaf] James WA Strachan et al. âTesting theory of mind in large language models and humansâ In Nature Human Behaviour Nature Publishing Group UK London, 2024, p. 1â11 [undefag] Winnie Street et al. âLlms achieve adult human performance on higher-order theory of mind tasksâ In Frontiers in Human Neuroscience 19 Frontiers Media SA, 2025, p. 1633272 [undefah] Tomer Ullman âLarge language models fail on trivial alterations to theory-of-mind tasksâ In arXiv preprint arXiv:2302.08399, 2023 [undefai] Sarah Walsh, Qiaosi Wang and Lance Ying âTheory of Mind in Human-AI Interaction and AIâ In Handbook of Human-Centered Artificial Intelligence Springer, 2025, p. 1â43 [undefaj] Alex Wilf, Sihyun Lee, Paul Pu Liang and Louis-Philippe Morency âThink twice: Perspective-taking improves large language modelsâ theory-of-mind capabilitiesâ In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, p. 8292â8308 [undefak] Heinz Wimmer and Josef Perner âBeliefs about beliefs: Representation and constraining function of wrong beliefs in young childrenâs understanding of deceptionâ In Cognition 13.1 Elsevier, 1983, p. 103â128 [undefal] Yufan Wu et al. âHi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Modelsâ In Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, p. 10691â10706 URL: https://aclanthology.org/2023.findings-emnlp.717/ [undefam] Hainiu Xu et al. âOpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Modelsâ In arXiv preprint arXiv:2402.06044, 2024 URL: https://arxiv.org/abs/2402.06044 [undefan] Xuhui Zhou et al. âSotopia: Interactive evaluation for social intelligence in language agentsâ In arXiv preprint arXiv:2310.11667, 2023 References [undefao] Ian Apperly âMindreaders: the cognitive basis of "theory of mind"â Psychology Press, 2010 [undefap] Chris Baker, Rebecca Saxe and Joshua Tenenbaum âBayesian theory of mind: Modeling joint belief-desire attributionâ In Proceedings of the annual meeting of the cognitive science society 33.33, 2011 [undefaq] Susan AJ Birch and Paul Bloom âThe curse of knowledge in reasoning about false beliefsâ In Psychological science 18.5 SAGE Publications Sage CA: Los Angeles, CA, 2007, p. 382â386 [undefar] Torben BraĂŒner, Patrick Blackburn and Irina Polyanskaya âBeing Deceived: Information Asymmetry in Second-Order False Belief Tasksâ In Topics in cognitive science 12.2 Wiley Online Library, 2020, p. 504â534 [undefas] Tom Brown et al. âLanguage models are few-shot learnersâ In Advances in neural information processing systems 33, 2020, p. 1877â1901 [undefat] Yupeng Chang et al. âA survey on evaluation of large language modelsâ In ACM transactions on intelligent systems and technology 15.3 ACM New York, NY, 2024, p. 1â45 [undefau] Yu-Chuan Chen and Hen-Hsen Huang âExploring Conversational Adaptability: Assessing the Proficiency of Large Language Models in Dynamic Alignment with Updated User Intentâ In Proceedings of the AAAI Conference on Artificial Intelligence 39, 2025, p. 23642â23650 URL: https://github.com/hhhuang/DynDST [undefav] Zhuang Chen et al. âTomBench: Benchmarking Theory of Mind in Large Language Modelsâ In arXiv preprint arXiv:2402.15052, 2024 URL: https://arxiv.org/abs/2402.15052 [undefaw] Wei-Lin Chiang et al. âChatbot arena: An open platform for evaluating llms by human preferenceâ In Forty-first International Conference on Machine Learning, 2024 [undefax] Boele De Raad âThe big five personality factors: the psycholexical approach to personality.â Hogrefe & Huber Publishers, 2000 [undefay] Abhimanyu Dubey et al. âThe llama 3 herd of modelsâ In arXiv e-prints, 2024, p. arXivâ2407 [undefaz] Michael C Frank and Noah D Goodman âCognitive modeling using artificial intelligenceâ In Annual Review of Psychology 77 Annual Reviews, 2025 [undefaaa] Jiaxian Guo et al. âSuspicion-agent: Playing imperfect information games with theory of mind aware gpt-4â In arXiv preprint arXiv:2309.17277, 2023 [undefaab] Stephen J Hoch âAvailability and interference in predictive judgment.â In Journal of Experimental Psychology: Learning, Memory, and Cognition 10.4 American Psychological Association, 1984, p. 649 [undefaac] Robin M Hogarth and Hillel J Einhorn âOrder effects in belief updating: The belief-adjustment modelâ In Cognitive psychology 24.1 Elsevier, 1992, p. 1â55 [undefaad] Jennifer Hu, Felix Sosa and Tomer Ullman âRe-evaluating Theory of Mind evaluation in large language modelsâ In arXiv preprint arXiv:2502.21098, 2025 [undefaae] Cameron R. Jones, Sean Trott and Benjamin Bergen âDoes reading words help you to read minds? A comparison of humans and LLMs at a recursive mindreading taskâ In Proceedings of the Annual Meeting of the Cognitive Science Society 46, 2024 URL: https://escholarship.org/uc/item/2k7307f1 [undefaaf] Hyunwoo Kim et al. âFANToM: A benchmark for stress-testing machine theory of mind in interactionsâ In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, p. 14397â14413 [undefaag] Michal Kosinski âTheory of mind may have spontaneously emerged in large language modelsâ In arXiv preprint arXiv:2302.02083 4 Mar, 2023, p. 169 [undefaah] Adam Tomasz Kostka and Jaros_aw A Chudziak âTowards Cognitive Synergy in LLM-Based Multi-Agent Systems: Integrating Theory of Mind and Critical Evaluationâ In Proceedings of the Annual Meeting of the Cognitive Science Society 47, 2025 [undefaai] Matthew Le, Y-Lan Boureau and Maximilian Nickel âRevisiting the evaluation of theory of mind through question answeringâ In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, p. 5872â5877 [undefaaj] Sangwon Lee, Naeun Lee and Young June Sah âPerceiving a mind in a chatbot: effect of mind perception and social cues on co-presence, closeness, and intention to useâ In International Journal of HumanâComputer Interaction 36.10 Taylor & Francis, 2020, p. 930â940 [undefaak] Xiaonan Li et al. âLlatrieval: Llm-verified retrieval for verifiable generationâ In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, p. 5453â5471 [undefaal] Nelson F Liu, Tianyi Zhang and Percy Liang âEvaluating verifiability in generative search enginesâ In arXiv preprint arXiv:2304.09848, 2023 [undefaam] Ziqiao Ma, Jacob Sansom, Run Peng and Joyce Chai âTowards a holistic landscape of situated theory of mind in large language modelsâ In arXiv preprint arXiv:2310.19619, 2023 [undefaan] Bennet B Murdock Jr âThe serial position effect of free recall.â In Journal of experimental psychology 64.5 American Psychological Association, 1962, p. 482 [undefaao] Thuy Ngoc Nguyen and Cleotilde Gonzalez âTheory of mind from observation in cognitive models and humansâ In Topics in cognitive science 14.4 Wiley Online Library, 2022, p. 665â686 [undefaap] Thuy Ngoc Nguyen, Kasturi Jamale and Cleotilde Gonzalez âPredicting and understanding human action decisions: Insights from large language models and cognitive instance-based learningâ In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 12, 2024, p. 126â136 [undefaaq] David Premack and Guy Woodruff âDoes the chimpanzee have a theory of mind?â In Behavioral and Brain Sciences 1.4, 1978, p. 515â526 DOI: 10.1017/S0140525X00076512 [undefaar] Shuwen Qiu et al. âMindDial: Enhancing Conversational Agents with Theory-of-Mind for Common Ground Alignment and Negotiationâ In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 2024, p. 557â575 URL: https://aclanthology.org/2024.sigdial-1.63/ [undefaas] Natalie Shapira et al. âClever hans or neural theory of mind? stress testing social reasoning in large language modelsâ In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, p. 2257â2273 [undefaat] Kazutoshi Shinoda et al. âToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mindâ In Proceedings of the AAAI Conference on Artificial Intelligence 39, 2025, p. 1520â1528 URL: https://ojs.aaai.org/index.php/AAAI/article/view/32143/34298 [undefaau] James WA Strachan et al. âTesting theory of mind in large language models and humansâ In Nature Human Behaviour Nature Publishing Group UK London, 2024, p. 1â11 [undefaav] Winnie Street et al. âLlms achieve adult human performance on higher-order theory of mind tasksâ In Frontiers in Human Neuroscience 19 Frontiers Media SA, 2025, p. 1633272 [undefaaw] Tomer Ullman âLarge language models fail on trivial alterations to theory-of-mind tasksâ In arXiv preprint arXiv:2302.08399, 2023 [undefaax] Sarah Walsh, Qiaosi Wang and Lance Ying âTheory of Mind in Human-AI Interaction and AIâ In Handbook of Human-Centered Artificial Intelligence Springer, 2025, p. 1â43 [undefaay] Alex Wilf, Sihyun Lee, Paul Pu Liang and Louis-Philippe Morency âThink twice: Perspective-taking improves large language modelsâ theory-of-mind capabilitiesâ In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, p. 8292â8308 [undefaaz] Heinz Wimmer and Josef Perner âBeliefs about beliefs: Representation and constraining function of wrong beliefs in young childrenâs understanding of deceptionâ In Cognition 13.1 Elsevier, 1983, p. 103â128 [undefaaaa] Yufan Wu et al. âHi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Modelsâ In Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, p. 10691â10706 URL: https://aclanthology.org/2023.findings-emnlp.717/ [undefaaab] Hainiu Xu et al. âOpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Modelsâ In arXiv preprint arXiv:2402.06044, 2024 URL: https://arxiv.org/abs/2402.06044 [undefaaac] Xuhui Zhou et al. âSotopia: Interactive evaluation for social intelligence in language agentsâ In arXiv preprint arXiv:2310.11667, 2023