Paper deep dive
Human Cognition in Machines: A Unified Perspective of World Models
Timothy Rupprecht, Pu Zhao, Amir Taherin, Arash Akbari, Arman Akbari, Yumei He, Sean Duffy, Juyi Lin, Yixiao Chen, Rahul Chowdhury, Enfu Nan, Yixin Shen, Yifan Cao, Haochen Zeng, Weiwei Chen, Geng Yuan, Jennifer Dy, Sarah Ostadabbas, Silvia Zhang, David Kaeli, Edmund Yeh, Yanzhi Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/27/2026, 12:46:10 AM
Summary
This report presents a unified framework for World Models (WMs) grounded in Cognitive Architecture Theory (CAT). The authors argue that current WMs, while advancing in video and embodied domains, lack critical human-like cognitive functions, specifically intrinsic motivation and meta-cognition. The paper introduces a novel taxonomy that classifies WMs by the cognitive functions they innovate (memory, perception, language, reasoning, imagining, motivation, and meta-cognition) rather than just architecture or application. Furthermore, it proposes the category of 'Epistemic World Models'—agent frameworks that operate over structured knowledge spaces to facilitate scientific discovery through a global workspace approach.
Entities (7)
Relation Signals (4)
World Models → groundedin → Cognitive Architecture Theory
confidence 100% · To evaluate these claims requires a proper grounding in first principles in Cognitive Architecture Theory (CAT).
Epistemic World Models → isatypeof → World Models
confidence 100% · We further introduce Epistemic World Models, a new category encompassing agent frameworks for scientific discovery...
JEPA → exemplifies → Memory
confidence 90% · JEPA is innovative in it’s encoded latent space in how they novelly train their model to extend representations of patched images to a fully reconstructed whole.
Global Workspace Theory → informs → Epistemic World Models
confidence 85% · we propose concrete directions informed by active inference and global workspace theory to address them.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost "human-like" cognitive capability in their world models. To evaluate these claims requires a proper grounding in first principles in Cognitive Architecture Theory (CAT). We present a conceptual unified framework for world models that fully incorporates all the cognitive functions associated with CAT (i.e. memory, perception, language, reasoning, imagining, motivation, and meta-cognition) and identify gaps in the research as a guide for future states of the art. In particular, we find that motivation (especially intrinsic motivation) and meta-cognition remain drastically under-researched, and we propose concrete directions informed by active inference and global workspace theory to address them. We further introduce Epistemic World Models, a new category encompassing agent frameworks for scientific discovery that operate over structured knowledge. Our taxonomy, applied across video, embodied, and epistemic world models, suggests research directions where prior taxonomies have not.
Tags
Links
- Source: https://arxiv.org/abs/2604.16592v1
- Canonical: https://arxiv.org/abs/2604.16592v1
Trouble viewing inline? Open PDF directly →
Full Text
163,565 characters extracted from source content.
Expand or collapse full text
Human Cognition in Machines: A Unified Perspective of World Models Timothy Rupprecht 1,2, * ,† , Pu Zhao 1, *, Amir Taherin 1 *, Arash Akbari 1 *, Arman Akbari 1 *, Yumei He 3 *, Sean Duffy 1 , Juyi Lin 1 , Yixiao Chen 1 , Rahul Chowdhury 1 , Enfu Nan 1 , Yixin Shen 1,5 , Yifan Cao 1,2 , Haochen Zeng 2 , Weiwei Chen 2 , Geng Yuan 1,4 , Jennifer Dy 1 , Sarah Ostadabbas 1 , Silvia Zhang 1 , David Kaeli 1 , Edmund Yeh 1 , and Yanzhi Wang 1,2,† 1 Northeastern University, Boston MA 02445, USA 2 EmbodyX Inc., Belmont, CA 94002, USA 3 Tulane University, New Orleans, LA 70118, USA 4 University of Georgia, Athens, GA 30602, USA 5 Cornell University, Ithaca, NY 14850, USA * These authors equally contributed † These authors are corresponding authors A Report from the Physical AI Research (PAIR) Center at Northeastern University Abstract. This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost "human- like" cognitive capability in their world models. To evaluate these claims requires a proper grounding in first principles in Cognitive Architecture Theory (CAT). We present a conceptual unified framework for world models that fully incorporates all the cognitive functions associated with CAT (i.e. memory, perception, language, reasoning, imagining, motiva- tion, and meta-cognition) and identify gaps in the research as a guide for future states of the art. In particular, we find that motivation (espe- cially intrinsic motivation) and meta-cognition remain drastically under- researched, and we propose concrete directions informed by active infer- ence and global workspace theory to address them. We further introduce Epistemic World Models, a new category encompassing agent frameworks for scientific discovery that operate over structured knowledge. Our tax- onomy, applied across video, embodied, and epistemic world models, sug- gests research directions where prior taxonomies have not. Keywords: World Models· Vision Language Action Models· World Action Models· Machine Cognition 1 Introduction Traditionally, World Models (WMs) are tools that enable agents to (1) repre- sent the current state of the world and (2) predict future states within it [40]. Early WMs were trained with Reinforcement Learning (RL) to learn action policies for robotic motor control [60, 189]. The scope has since expanded to arXiv:2604.16592v1 [cs.RO] 17 Apr 2026 2Authors Suppressed Due to Excessive Length include multi-modal models capable of generating vibrant visuals [169], embod- ied WMs for mapping and locomotion [75, 70], and as we will argue, WMs capable of agentic collaboration on scientific discovery [56, 126, 149]. Exist- ing surveys distinguish works by architecture, application, or self-defined tax- onomies [40, 101, 211, 113, 109, 110, 199, 43], most reducing the field to a coarse dichotomy between world representation and world generation [40]. We are the first to systematically classify WMs by the cognitive function they emulate: mem- ory, perception, language, reasoning, imagining, motivation, and meta-cognition. World Models are increasingly described in anthropomorphic terms [74, 195, 88], with many works claiming almost “human-like” cognitive capabilities. Such claims can overstate what these systems actually achieve, causing confusion among researchers and the public alike, and misdirecting attention from the cog- nitive functions that remain unsolved. Evaluating these claims requires ground- ing in Cognitive Architecture Theory (CAT) [123], which decomposes cognition into functional components such as memory, perception, reasoning, and motiva- tion. CAT provides a principled basis for comparing machine and human cog- nition, clarifying where current World Models succeed and where critical gaps remain. In our report we offer our perspective on a unified framework for World Models that incorporates all cognitive functions identified by CAT and serves as a roadmap for the field. From this synthesis, we identify two critically under- researched cognitive components: motivation and meta-cognition. Current state- of-the-art World Models lack intrinsic motivation mechanisms beyond hand- crafted reward signals, and none demonstrate genuine meta-cognitive capabilities such as self-monitoring, self-evaluation, or self-control. Addressing these gaps is essential if the field’s aspiration toward human-like World Models is to move beyond rhetoric. Our report spans three domains of World Model research. Video World Mod- els (Sec. 4) generate future visual states conditioned on observations and actions, where maintaining spatial consistency and long-horizon temporal coherence re- main central challenges. Embodied World Models (Sec. 5) extend these demands to physical settings, requiring perception of contact geometry, memory of persis- tent environments, and reasoning over force propagation to guide real-world task execution. Beyond these established domains, we propose a new category we call Epistemic World Models (Sec. 6), in which the environment is not a physical scene but a structured knowledge space defined by literature, databases, and experimental outputs. In this setting, an agent updates its world state through reasoning and tool use within a global workspace, enabling agentic collabora- tion on scientific discovery with a human in the loop. Epistemic World Models also provide early instantiations of the meta-cognitive mechanisms that latent World Models currently lack, making them both a distinct research domain and a source of solutions for the gaps our taxonomy reveals. The contributions of our report are as follows: 1. We provide a novel review of recent World Models grounded in human- machine cognitive architecture theory. Human Cognition in Machines: A Unified Perspective of World Models3 Language and meta- cognition (1990s) Global Workspace Theory (1988) Shared Intentions (2008) Attention Networks (2017) ImageNet (2014) Soar Cognitive Architecture (2019) Human Cognition Machine Cognition World Models Dreamer v3 (2019) JEPA Encoding (2024) GPT-3 (2020) COSMOS, PI- Model, Marble (2025) Active Inference and Artificial Reasoning (2025) Active Inference (2022) World Model Surveys (2026) World Models (2018) Cognitive Architecture Theory (1994) Fig. 1. Our survey studies the convergence of three different but inter-related fields: human cognition, machine cognition, and World Models. 2. We propose a unified World Model as a conceptual road-map for incorpo- rating all the component parts of cognitive architecture for robust world representation and generation. 3. We identify and propose solutions to research gaps in World Model motiva- tion and meta-cognition. 4. We propose a new category of World Model that learn to represent a world from structured knowledge we call Epistemic World Models. 2 Background Our report spans these three interrelated research tracks as summarized in Fig- ure 1. Previous World Model surveys create a coarse dichotomy in their tax- onomies. They classify World Models as 1) World Representations [63, 8, 53, 50, 16], and 2) World Predictors or Generators [69, 176, 80, 190, 193]. In Figure 2 our finer dichotomy of World Models is shown with exemplary works that innovate on their primary cognitive function. Our taxonomy draws on lessons from the fields of human and machine cognition that we will review in this section. While our taxonomy is not explicit in prior work, it emerges naturally when aligning model capabilities with longstanding [123, 94, 6, 120, 95] and recent [48, 139, 133] cognitive architectures. We discuss the research tracks from Figure 1 now, first by reviewing World Model research in Sec. 2.1, human and machine cognition in Sec. 2.2. 2.1 World Models The seminal work on contemporary World Models is from Ha et al. (2018) [60]. Ha et al. proposed a framework enabling dreamer architectures [60]. A review of recent surveys of state-of-the-art World Models [40, 101, 211, 113, 109, 110, 199, 43] shows that in order to support planning and decision-making, espe- cially in embodied settings [113, 101], World Models consistently function as 4Authors Suppressed Due to Excessive Length Taxonomy of World Models LANGUAGE semantic and symbolic world knowledge encoding World Representation Sense and learn real-world knowledge World Generation Reason according to physical laws MEMORY state space encoding for world representation PERCEPTION sensory input encoding for world representation REASONING how future states of the world are predicted IMAGINING hypothetical reasoning MOTIVATION reward signals during training and inference TD-MPC2 (2023) JEPA (2024) AdaWorld (2025) Knowledge Graphs as WMs (2025) LaDi-WM (2025) State-space WMs (2025) Infinite-World (2026) BEVWorld (2024) VAGEN (2025) Marble World (2025) Masked Latent Transformers (2025) OccWorld (2025) Multimodal Dreaming (2025) DrivingGPT (2024) Worldgpt (2024) Doe-1 (2024) LiDARCrafter (2025) Explicit WMs (2025) Towards an AI co-scientist (2025) Exemplars AdaWM (2025) Rlvr-world (2025) Epona (2025) LingBot VLA (2025) GigaWorld-Policy (2026) Multi Task WM (2025) NewtonRewards (2025) Irl-vla (2025) RLVR-World (2025) MoSim (2025) InDRiVE (2025) Re-world (2026) Safedreamer (2023) Dream to drive (2024) Dream to drive (2025) Dream2Flow (2025) VideoWeave (2025) PAN World Model (2025) Fig. 2. The taxonomy of World Models covered in our survey correspond to the com- ponent parts of cognitive architecture theory [123] they innovate most. simulators that 1) represent current world structure and 2) predict future world dynamics [40]. Recent works also survey advances in video World Models [211], embodiment [101], temporal–spatial modeling [110], and physical realism [109], all highlighting challenges in long-horizon consistency, computational efficiency, and alignment to real-world physics. The available surveys also focus on domain application [113, 211, 101, 109, 43], architectural or input modality distinctions [110], or abstract taxonomies [40, 199]. As demonstrated in Table 1, we are among the first to systematically dis- tinguish recent states of the art by the cognitive functions they primarily inno- vate. Surveys near-universally discuss the foundational models and works from state-of-the-art teams at META AI with Yann LeCun [64, 52, 208], the Alibaba group [175], Cosmos from NVidia [87], and Berkeley University’s two groups with Sergey Levine [75] and Fei-fei Li [200] [178] [70]. Additional states-of-the-art in- novate on architecture by using Vision-Action [99] and Vision-Action-Language models [170], auto-regression models [50], and diffusion [218, 104]. Neverthe- less, a framework for a unified World Model remains elusive to researchers. A divide remains in the designs for general-purpose and domain-specific World Models [199]. Human Cognition in Machines: A Unified Perspective of World Models5 SurveyYearVideoEmbodiedSimulationPhys. Align.EpistemicCATPrimary Distinction Ding et al. [40] 2025✓✗ World understanding vs. prediction Li et al. [101]2025✗✓✗ Embodied AI and simulators Yue et al. [211] 2025✓✗✓✗ Roadmap for visual world simulation Long et al. [113] 2025✗✓✗ Simulators for embodied intelligence Lin et al. [109] 2025✓✗✓✗ Physics alignment in video gen. Liu et al. [110] 2025✓✗ Architectural and input distinctions Xu et al. [199] 2026✓✗ Specialist-to-generalist progression Dong et al. [43] 2026✓✗ Learning to model the world in AI Ours2026✓ Cognitive architecture theory Table 1. Comparison of World Model survey scope. A check (✓) indicates the survey explicitly covers that domain. The survey withn our report is the first to classify World Models by the cognitive functions they innovate using Cognitive Architecture Theory (CAT). 2.2 Cognitive Architecture Theory The capabilities of World Models regularly draw comparisons to the abilities of their human counterparts [89, 195, 188]. This contributes to a widely held public belief that LLM’s, World Models, and similar generative AI approach human- like reasoning capabilities, are conscious, or herald imminent artificial general intelligence [83, 32, 144]. Without proper grounding in first-principles thinking from Cognitive Architecture Theory, these comparisons are misleading and often overstate the capabilities of our non-human counter-parts. Cognitive Architecture in Humans. To say a World Model achieves Human- like capabilities in cognition requires a first-principle investigation of Cognitive Architecture Theory in humans. Human cognition includes component parts such as motor skills, adaptive learning, perception, symbolic reasoning with language, and memory [42, 171, 79]. Human cognition also includes meta-cognition, or consciousness, and refers to a human’s ability to self-monitor on more complex tasks [10, 91, 141]. Language as symbolic reasoning is relevant to our understanding of World Models because both these and humans share the input and output modality of textual language. Early and modern experts on the development of human language have argued that language is a tool in response to external factors like natural selection [35, 134], and is used to influence the external world [22, 107]. An emerging view argues that language evolved from humans’ ability to share intentions (i.e. humans’ ability to share goals, to share attention or to share common ground) [172]. Human language exists in reference to a World Model shared between a party of humans. From a review of first principles regarding how both human cognition and consciousness have evolved in humans we assert the role of language is paramount. 6Authors Suppressed Due to Excessive Length We look to the works of evolutionary biologist Terrance Deacon [36] and psy- chologists Julian Jaynes [79] and Merlin Donald [42]. Deacon succinctly asserts that in humans “symbolic thought does not come innately built in, but develops by internalizing the symbolic process that underlies language.” Similarly, Jaynes asserts that “consciousness becomes embedded in language.” Cognitive Architecture in Machines. Newell maps the component parts of Cognitive Architecture Theory found in human cognition that overlap with Machine Cognition [123]. The component parts given to us by Newell are: 1. memory - state space encoding for world representation 2. perception - state space encoding for world interaction (e.g. motor-skills) 3. language - state space encoding (e.g. semantic and symbolic) 4. reasoning - predicting states of the world (e.g. generation or simulation) 5. imagination - hypothetical reasoning (e.g. dreamers and sim-to-real) 6. motivation - how or why reasoning occurs (e.g. a model’s reward signal) We include meta-cognition, a World Model’s capability of self-monitoring, self-evaluation, or self-control, in this taxonomy of human and machine cog- nition [10, 141, 91]. Meta-cognition is the recursive application of the other cognitive functions to themselves — reasoning about one’s own reasoning, mon- itoring one’s own perception, evaluating one’s own memory retrieval. Because it operates over the same functional categories Newell identified rather than intro- ducing a new category of its own, its inclusion completes the taxonomy rather than extending it. The first three items in the above test correspond to typical World Model tasks in world representation, and the last three correspond to world prediction or generation. This taxonomy is depicted in Figure 2. Throughout our review, it is important to note that machine cognition func- tionally emulates, and does not mechanistically emulate these component parts of human cognition. For example, when we say a World Model has “memory,” we mean it maintains state representations across time steps, not that it has anything resembling episodic recall or the biological substrates of human mem- ory. Machine cognition emulates human cognition in memory, perception, and some symbolic reasoning tasks through language, and imagination with sim- to-real design paradigms. It is ambiguous if machine cognition fully emulates true human-like reasoning and machine cognition almost universally fails to emu- late intrinsic motivation. Machine intelligence has also not demonstrated any capability in emulating meta-cognition. To operationalize our taxonomy using the cognitive functions distinguished by Newell in a classification system we turn to an example World Model work, in this case JEPA [8]. We acknowledge many works span several categories, JEPA as an image encoder innately perceives input images, but this does not necessar- ily innovate on image encoding perception as image-encoding innately requires images as an input. However, extending JEPA to videos [14], or to other input modalities would constitute an innovation for the cognitive function of percep- tion regarding the original JEPA model. And while all image-encoders represent Human Cognition in Machines: A Unified Perspective of World Models7 World Representation World Generation MEMORY MEMORY PERCEPTION LANGUAGE REASONING IMAGINING MOTIVATION LANGUAGE PERCEPTION Meta-cognition MOTIVATION REASONING IMAGINING MEMORY MOTIVATION IMAGINING REASONING LANGUAGE PERCEPTION (self-reflects and self-controls) Fig. 3. The component-parts of our Unified World Model built from first principles in cognitive architecture theory [123] and meta-cognition [10]. This serves as a conceptual road-map for World Model research. images in a lower-dimension latent space functionally constituting memory, JEPA is innovative in it’s encoded latent space in how they novelly train their model to extend representations of patched images to a fully reconstructed whole. Similarly, all the models we consider in our survey have some architecture used for inference for decision making constituting reasoning, but a work is only in- novative regarding their architecture if it used within a setting is truly novel. A JEPA model may be frozen and used for world representation within a broader sim-to-real, or dreamer-like World Model work [9], but that does not mean that JEPA itself is innately innovative regarding hypothetical reasoning. As such, for a World Model work like JEPA, we would discuss it as part of world rep- resentation in Sec. 3.1 and mark is as particularly exemplary at innovating the cognitive function of memory which is also seen in Figure 2. To operationally replicate our taxonomy a researcher must match a work’s stated contributions to an emulated cognitive function. We have a second researcher confirm these decisions. Key examples of cognitive functions are available in Fig. 2. 8Authors Suppressed Due to Excessive Length 3 Unified Cognition Framework for World Models We report on World Models in the context of video World Models (discussed further in Sec. 4), embodied World Models (discussed further in Sec. 5), and what we define as epistemic World Models, or World Models used by agents for scientific discovery (discussed further in Sec. 3.3 and 6). We subcategorize each section according to the component parts of machine and human cognition listed in Sec. 2.2. Exemplars for the contemporary works we review that innovate on specific functions of cognition can be seen in Figure 2. We propose our conceptual unified World Model framework and hope it serves as a conceptual road map for this report and World Model researchers. World Models sense and learn from real-world knowledge, predict and generate world states, reason and control according to physical laws implicitly, and can do all of this with or without a machine agent or a human-in-the-loop. To accomplish this we propose a unified World Model that holistically incorporates every compo- nent part of the CAT concepts discussed in Sec. 2.2 into one conceptual unified framework as seen in Figure 3. For the remaining sections of this work it will serve as a conceptual road map for organizing other works. As a design paradigm, our conceptual Unified Framework for World Models calls for standardizing best practices in both World Model representation, and World Model generation. To fully span the functions of cognition emulated, Unified World Models encourage researchers to: 1. Use multi-modal inputs for perception (discussed further in Sec. 3.1), 2. Represent the world using latent state-spaces as memory to enable down- stream cognitive functions such as imagination, motivation, and reason- ing (discussed further in Sec. 3.1), 3. Include language tokens as an input, intermediate reasoning space, or out- put to facilitate human-in-the-loop cooperation (discussed further in Sec. 3.1), 4. Enable imagination as hypothetical reasoning during inference and sim- to-real transfer learning during training in domains with little accessible training data (discussed further in Sec. 3.2), 5. Apply domain-specific state-of-the-art architectures for reasoning (discussed further in Sec. 3.2), 6. Provide reward signals to World Models for motivation using state-based rewards that make salient robust measurements like active inference (dis- cussed further in Sec. 3.3), 7. Utilize meta-cognition through global workspaces and self-evaluation (dis- cussed further in Sec. 3.3). None of these suggestions conflict with each other, but some are setting specific or conditional. In Sec. 3.3 we will argue our proposed research directions regarding motivation and meta-cognition fill a real research gap our taxonomy has revealed. Reviewing video World Models in Sec. 4 shows us the importance of cre- ating world representations that enforce spatial consistency and meet memory constraints using solutions like kv-cache while still enforcing longer temporal Human Cognition in Machines: A Unified Perspective of World Models9 consistency. We show state-of-the-art video World Model architectures in Fig- ure 4. When we review embodied World Models in Sec. 5 we see the importance of leveraging multi-modal inputs to make precise locomotion possible. We also see in Sec. 5 that when training data is scarce, sim-to-real training can overcome this scarcity when World Models learn latent representations that are traversable and remain consistent under domain shift to real-world applications (discussed more in Sec. 3.1). We show state-of-the-art embodied World Model architec- tures in Figure 5. We propose in Sec. 3.3 a new category of World Model called epistemic World Models that we review in Sec. 6. This domain includes agent frameworks for human-in-the-loop scientific discovery which serve as inspiration for overcoming research gap we have observed in latent World Model’s capability for meta-cognition. We show state-of-the-art epistemic World Model architec- tures in Figure 4. Latent World Models learn state transition dynamics in video and embodied settings while agent frameworks are not typically framed that way. However, Epistemic World Models in an agent framework with a Global Workspace repre- sentation of the world perfectly aligns with soar State Operation Result loop and we argue, even the latent World Model definition. A Global Workspace is not a learned latent encoding like JEPA [8]. But a Global Workspace and an agentic framework acting upon this global workspace resembles the pipeline of a prompt- able Video World Models like we will see in Figure 4, and discuss in Sec. 4, or World Action Models like we will see in Figure 5, and discussed in Sec. 5. In the Epistemic World Model setting, the initial world is a static world representing structured knowledge, but as the agent acts by augmenting its context through multimodal rag queries, performing analysis of existing literature, and encoding tool inputs and outputs into a global workspace (i.e. a chat history), creates a changing world, or in other words a dynamic state space, meeting the tra- ditional definition of the World Model. Working within this agentic framework with global workspace implementation allows models like Gemini Co-scientist, an LLM and World Model itself, to propose novel trajectories through an evolving state space to perform actual scientific discovery. 3.1 World Representation Perception with Multi-modal World Models. We start by considering an input of multi-modal observationso v t for images or videos,o ℓ t for language,o a t for audio, and so on. Our multi-modal perception of the world used to create our world’s latent representation at step t is defined as, z t = φ θ (o v t ,o ℓ t ,o a t ,... ) s.t. dim(z t )≤ B, I(o 1:t ;z t )≤ C(1) Or with Soar-styled recursion (enabling some degree of self-monitoring meta- cognition) with world-encoding latent state persistence, z t = φ θ (z t−1 ,o v t ,o ℓ t ,o a t ,... ) s.t. dim(z t )≤ B, I(o 1:t ;z t )≤ C (2) In any case, after input coding, we have latent variable z t subject to the con- straints for physically fitting z t within a memory budget B, while also keep 10Authors Suppressed Due to Excessive Length mutual information below C, or a sufficient mutual information between the ob- servations, and latent space at step t. At this stage, alignment with physical laws is implicitly instilled through training data, and handcrafted reward signals (the latter is discussed more later on). There are many examples of World Model papers that have major contri- butions to improving perception within machine cognition in Figure 2 with more examples discussed in Sec. 4 to 6. A key insight across recent works is that perception is inherently multi-scale and multi-modal. For instance, augment- ing Dreamer-style agents with traversable spatial latent representations enables improved sim-to-real transfer, highlighting the importance of jointly encoding spatial structure alongside temporal dynamics [23]. Similarly, temporal hierarchy learning introduces multiple levels of abstraction in perception, allowing agents to reason over both short-term transitions and long-horizon dependencies [58]. Both show the value of using more than a single mode of input encoding. Beyond vision and temporal signals, embodied World Models integrate en- coded lidar measurements,o lidar t , into the perceptual pipeline demonstrate that fusing complementary spatial modalities improves robustness and downstream planning performance, particularly in navigation and autonomous systems [105, 230]. Recent systems combining Global Workspace (GW) architectures with World Models, such as GW-Dreamer, provide evidence that global information sharing with multimodal latent spaces can improve both sample efficiency and generalization from simulated dreaming [118]. From the perspective of cognitive architecture theory, these approaches align with the role of perception as a gateway to a shared workspace, where diverse sensory inputs are fused into a coherent representation to support downstream reasoning, memory, and action selection. Memory with Long Context Windows. An optimal latent world-encoding z ∗ t , can be thought of as implicitly overcoming an information bottleneck where our ideal world-encoding minimizes mutual information [85] between the observa- tions of the world, and the encoded state. By conditioning on a reward functional evaluated over latent rollouts our latent world-encoding preserves information relevant to the learning signal, or in other words, z ∗ t = min z t I o ≤t ;z t |r(z t+1 ) (3) However, with scalar rewards, or if the reward signal is low dimensional, then the model predicts immediate reward but may lose world structure. To overcome this limitation, our unified framework proposes conditioning on the full state-based reward trajectory from t + 1 to t + H, or, z ∗ t:t+H = min z t:t+H I o ≤t ;z t |r(z t+1:t+H ) (4) This perspective aligns with prior work on traversable latent spaces [58, 23], and suggests that long-context memory in World Models is inherently tied to Human Cognition in Machines: A Unified Perspective of World Models11 the capacity for structured latent prediction. This allows sim-to-real “dream- ing” during training and hypothetical reasoning during inference which will be discussed in Sec. 3.2. Additional examples of latent-space innovations are sum- marized in Fig. 2 and discussed in Sec. 4–6. Unified World Models therefore seek an encoder φ θ (·) that produces representations supporting sufficiently long temporal context, as required by Eq. (1) and Eq. (2). Language. Large language models (LLMs) [225, 226, 157, 155, 216, 158] and Vision-language Models (VLMs) [223, 198, 151, 153, 156], are often said to be World Models themselves [53, 57]. Yet, empirically they show insensitivity to meaning in LLM comprehension tasks when compared to human counter- parts [37]. When researchers align an LLM’s language cognitive function to humans language cognitive function by repeating LLM prompts to emulate a human’s innate ability to hold both the beginning and end of a spoken sentence in their mind at once, the LLM accuracy improves [97]. Also, latent spaces learned by language models converge to similar embeddings across differing modal in- puts [73] and visual embeddings especially are aligned with the human brain [41]. Despite all this targeted innovation in language, LLMs still fail to reason as hu- mans do in the tasks we ask humans to perform [74, 129]. A holistic approach to cognitive functions is needed to reach the next state-of-the-art language model. A Unified World Model Approach to Representation. Foundational rep- resentative World Models include Cosmos [2, 3], LingBot-WM [170], GIGA- World [202], RLVR-World [187], and finally JEPA [8] who all learn to represent knowledge of the world implicitly aligning with physical laws through training data. Joint-Embedding Predictive Architectures (JEPA) provides a canonical instantiation of the representation component within our Unified World Model framework [8, 174, 9, 38, 164, 181]. Rather than reconstructing raw sensory in- puts, JEPA learns representations by predicting target embeddings from context embeddings in a shared latent space [8], thereby directly optimizing for a more robust predictive structure. This predictive latent-space formulation naturally supports multiple cognitive functions. JEPA encoders can integrate perception across multiple modalities [33] in embodied settings [181], incorporate language as inputs [174], or to serve as intermediate representations for reasoning [164]. Crucially, when combined with latent dynamics models, these representations enable long-horizon memory through compact latent world-encodings that sup- port hypothetical action rollouts [9, 38]. In this sense, JEPA-style representa- tions are not merely compressive, but are structured to support prediction over imagined futures, aligning directly with the requirements imposed by our reward- conditioned information bottleneck formulation. We will see how World Models in Sec. 4 to 6 sense and learn world knowledge in each domain application. 3.2 World Model Prediction and Generation While world representation defines how world knowledge is sensed and learned, World Model prediction and generation define how the world transitions over 12Authors Suppressed Due to Excessive Length time and transitions in response to agents acting within the world. World Models generate 3d scenes (discussed further in Sec. 4), generate controlled real-world scenes (discussed further in both Sec. 4 and 5), and are robust with a human-in- the-loop (discussed further in Sec. 6). In our unified framework, the previously stated functional categories of word models correspond to the cognitive functions of reasoning and imagination. Together they enable structured planning, hypothetical simulation, and decision-making over future trajectories. Imagination. During World Model training, imagination manifests as dream- ing [60], referring to a World Model’s ability to learn to represent world dy- namics through unsupervised training. If these learned latent representations are traversable and remain consistent under domain shift then an action-policy model can be trained with these latent spaces and deployed in the real-world do- main [60, 46, 118]. In real-world applications training data can be prohibitively scarce and expensive to create; sim-to-real training overcomes this problem [142]. Figures 4 and 5 show architectures that support dreaming in both the video World Model and embodied World Model domains. During inference, imagination refers to hypothetical reasoning, planning, or simulating future states of the world over a set of possible action rollouts to find an optimal action [193, 197]. This is only possible after applying lessons from Sec. 3.1 where Eq. (4) is followed to design a traversable world-encoding latent space. The optimal action, or a ∗ t , maximizes the expected reward signal over the time horizon of the possible action rollout. For action-planning we assume a decomposition of the World Model into a transition function and an observable reward functional. The transition model W θ governs the evolution of latent states when conditioned on an action, z t+1 = W θ (z t ,a t ),(5) Here the reward function r(z) evaluates the desirability of latent states or trajec- tories. We treat W θ (z t ,a t ) as a parameterized world transition function, though it may in practice represent a stochastic process. In a search for the optimal action at time step t, a ∗ t , the reward signal is considered over the time horizon H, a ∗ t:t+H = arg max a t:t+H E " H X k=0 r(z t+k ) # (6) Thus, action selection can be understood as inference over latent trajectories, where candidate futures are simulated and evaluated under the reward func- tional. The resulting z t+1 is obtained by W θ , our Unified World Model for World Model generation guided by principles from Cognitive Architecture Theory al- lowing the possibility of more human-like reasoning. Additionally, improve- ments in episodic structuring and long-range temporal coherence ensure that imagined trajectories remain consistent over extended horizons [46, 58, 50, 188, 193]. Human Cognition in Machines: A Unified Perspective of World Models13 Recently, World Action Models propose an end-to-end generative model for world encoding conditioned on actions to predict future states of the world in training [203, 99, 202, 210]. Dream-zero proposes using imagination as a planner [203]. Gigaworld-Policy using imagination for supervision during train- ing [202]. And Lingbot-VA using dreaming as an iterative denoising step [99]. Often World Action Models use dreaming as a training regularizer, and at infer- ence due to latency concerns, do not using dreaming states [202]. Reasoning. In World Models reasoning is defined as the structured evaluation and selection of action rollouts for the optimization objective in Eq. (6). Recent work highlights the importance of maintaining narrative and logical consistency across generated trajectories, particularly in long-horizon settings where com- pounding errors can degrade performance [197, 99]. Hierarchical approaches fur- ther decompose reasoning into multi-scale planning processes, enabling agents to coordinate short-term actions with long-term goals [58]. From a cognitive perspective, reasoning constrains the space of candidate trajectories z t+k , se- lecting those that are both plausible under W θ and optimal with respect to the action selection criterion. We identify eight recurring paradigms, shown in across Figures 4, 5, and 6, that span the spectrum from purely discriminative action selection to full gen- erative world simulation to World Models that solely represent the world. At one extreme, Vision-Language-Action models (VLAs) such as π 0 [75] or ling-bot VLA [190], or other VLA methods that map observations directly to actions through a VLM backbone augmented with a flow-matching action expert, by- passing explicit future-state prediction entirely–therefore lacking an ability to hypothetically reason or imagine [236, 38, 80, 108, 164, 112]. At the other extreme, some World Models only represent the world [174], and use a sepa- rate value [38] or policy models for action or future state generation [2, 99]. These World Action Models leverage training data and are not conditioned on any action. This decoupling representation and action generation grants max- imum flexibility in imagination but introduces a two-system overhead where errors in world prediction and errors in action selection compound indepen- dently. World Action Models [203, 99, 202] tend to be more robust and work better with longer, and noisier time horizons. Additionally, both autoregres- sive transformer [31, 214, 111, 216, 237] and diffusion methods are encoun- tered [218, 205, 72, 215, 154, 235, 71]. Motus claims to be a unified World Model because it successfully unites generating latent action, future video states, and agent actions [17]. 3.3 Cognitive Architecture in World Models Our taxonomy categorizes the surveyed works according to the cognitive func- tions they innovate. Many works incorporate several cognitive functions in their designs and sometimes innovate along multiple cognitive axes which is why some works appear in multiple sections or have several check marks in Tables 2, 3, 14Authors Suppressed Due to Excessive Length and 4. Analyzing the innovation trends presented in our report through Ta- bles 2, 3 and 4 shows a research gap exists regarding the cognitive functions of motivation, and meta-cognition. The utilization of our taxonomy made clear this research gap and may not have become apparent otherwise. It is clear that machine and human cognition continue to be intertwined. There have been a variety of machine cognitive architectures proposed [95, 6, 120, 139, 94, 47]. ICARUS utilizes hierarchical skills and concepts [95]. ACT- R uses Bayesian-style activation in a Adaptive Control of Thought—Rational loop [6, 139]. EPIC leverages perception–action timing making it well-suited for human-computer interaction in embodied settings [120]. Using the same princi- ples we explored in Sec. 2.2, author Laird synthesized Newell’s taxonomy of cog- nitive architecture to propose the machine cognitive architecture Soar [94]. Soar is a cognitive architecture that models intelligent behavior through a symbolic State–Operator–Result (or S-O-R) cycle, in which reasoning proceeds through the selection and application of operators to representations of the world. S-O-R was proposed as R = f rules (S t ,O t ) where R is S t+1 . In our neural and latent World Model notation, this is respectively, s t+1 ∼ W θ (s t |a t ) and z t+1 = W θ (z t |a t )(7) The Soar system is capable of being applied in worlds that require visual and symbolic representative reasoning [20] implying compatibility with our state- of-the-art World Models. Soar provides a unified framework for perception, decision-making, and learning via production rules operating over a shared work- ing memory. We also see this in our review of Epistemic World Models in Sec. 6. Both self-evaluative techniques are similar to the Global Workspace Theory pro- posed for human meta-cognition [10]. Soar’s use in World Models was validated empirically in recent work [207]. An exciting new theory of human cognition proposes Active Inference [48]. This explains human cognition as a system that minimizes surprise about sensory inputs by continuously updating beliefs about the world and takes actions that make the world match those beliefs. Active Inference has been extended to ma- chine cognition [133, 47] as theory that novelly motivates learning in machines and informs our Unified World Model Framework (discussed further in Sec. 3.3). World Model works have already begun to incorporate active inference [118]. Across our report of state-of-the-art World Models we have identified two innovation gaps in research regarding two of the component parts of our Unified World Models: motivation, and meta-cognition. This is elucidated in part by evaluative benchmarks from recent works [27, 88, 127, 136, 161, 182]. But also, one only needs to look at Tables 2, 3, and 4 to see the lack of innovations regarding motivation and meta-cognition. To overcome these gaps, regard- ing meta-cognition, researchers should look to the agent frameworks used in human-AI collaboration for Scientific Discovery, as seen in Figure 6; and re- garding motivation researchers should look to advances from human-machine cognitive architecture theory regarding Active Inference [48, 133, 47, 118]. Human Cognition in Machines: A Unified Perspective of World Models15 Motivation. Reinforcement Learning techniques instill external motivation through reward signals into World Models using a reward function defined external to the World Model itself. To enable imagination or hypothetical reasoning, as described by Eq. (6), requires an observable reward functional r(z t ) [63, 186, 65, 86, 118, 121]. Multi-task learning allows the pursuit of multiple goals [118] and can help align models to the physical world when trained with an appro- priate reward signal [96]. Researchers construct RL agents as a potential solu- tion to artificial general intelligence assuming that their reward signal is suffi- cient [160]. However, often RL provides training mechanisms for World Models using hand-crafted reward signals that do not generalize [142] and rewards that are misaligned with the intended operational goal [5, 62]. As such, they require additional interventions such as Reinforcement Learning from Human Feedback in some applications [136, 109]. Intrinsic Motivation. Promising research directions exist, such as maximizing influence over future states [90] or minimizing energy (i.e. surprise) [48, 133, 47] in exploring learning objectives and training mechanisms decoupled from an externally defined and action-based reward signal. Both maximizing influence of future states, and active inference are compatible with the state-based reward function described in Eq. (6). Active inference formalizes “minimizing surprise” as the minimization of variational free energy, a tractable upper bound on negative log model evidence, which decomposes into accuracy and complexity and extends to action selection through expected free energy that balances information gain and value. Concretely, given a latent World Model with states z t , observations o t , and actions a t , the generative model is defined as p θ (o 1:T ,z 1:T | a 1:T ) = T Y t=1 p θ (o t | z t )p θ (z t | z t−1 ,a t−1 ).(8) Inference proceeds by introducing an approximate posterior q φ (z 1:T | o 1:T ) and minimizing the variational free energy F =E q φ (z 1:T ) [logq φ (z 1:T )− logp θ (o 1:T ,z 1:T )],(9) which upper bounds the surprise − logp θ (o 1:T ). This objective decomposes as F = D KL (q φ (z 1:T )∥p θ (z 1:T ))−E q φ [logp θ (o 1:T | z 1:T )],(10) corresponding to complexity and accuracy terms, respectively. For action selec- tion, policies π are evaluated by minimizing the expected free energy G(π) =E q φ D KL (q φ (z t+1:T | π)∥p θ (z t+1:T )) | z information gain − logp(o t+1:T ) | z value ,(11) encouraging trajectories that both reduce uncertainty in latent dynamics and achieve preferred outcomes. World Model research has already begun to incor- porate active inference. 16Authors Suppressed Due to Excessive Length Using intrinsic motivation for World Models is drastically unrealized but there are some recent works using intrinsic motivation within RL that validate this research direction [130, 15, 152, 45, 131, 204, 200, 7, 84, 140]. Among these works, active inference is also being deployed in embodied settings to promote exploration during training [7, 140], and in one case even control swarms of unmanned autonomous vehicles [7]. Three of these are explicit World Model works but remain under-cited [45, 204, 200, 140]. The first two explicitly use active inference as a reward signal for World Models [45, 204]. A more recent work World Model work, Cambrian-S, explicitly references prediction error as a proxy for measuring surprise but falls short of invoking Active Inference directly [200]. It is also work noting, researchers are ultimately the ones selecting archi- tectures, reward criteria and training sets within proposed research. Peer re- view evaluates objectively if the subsequent empirical results are novel and no- table. It is researchers who breathe life into World Models by aligning experi- ment designs with broader research trends. This is demonstrated in the global workspace World Models used by agent frameworks for AI-Human Collabora- tion [56, 150, 149] shown in Figure 6 and discussed in Sec. 6. Meta-cognition. Soar’s usage of subtasks and recursive operator decomposi- tion mimics human meta-cognition by breaking bigger more abstract tasks into smaller more discrete tasks [94]. For example, if an agent with World Model are navigating a problem and not experiencing progress towards a goal, sub-tasks allow debugging of that lack of progress, or proposing alternative routes to the goal; thus allowing a minimal degree of meta-cognition. However, the failure of World Models to align with meta-cognition of human counter-parts shows that while meta-cognition is theoretically possible, it still represents a largely unsolved implementation problem [81, 135, 129]. Meta-cognition via Global Workspaces. Exploring meta-cognition through the use of multi-agent systems capable of observing each other shows some progress towards the goal of implementing truly meta-cognitive capabilities within World Models [18, 180]. Figure 6 shows how these agent frameworks provide a World Model that enables some amount of meta-cognition through use of a Global Workspace [10] accessible by agents and humans collaborators. The central limi- tation identified in World Models is the absence of explicit meta-cognitive mech- anisms in latent World Models, particularly in embodied and video domains seen in Sec. 4 and 5. While these systems learn latent state representations z t and transition dynamics z t+1 = W θ (z t ,a t ), they lack a mechanism for monitoring, evaluating, or selectively routing internal computations. To fill this research gap, our taxonomy demands the inclusion of Epistemic World Models, or agent frameworks with a Global Workspace representation of the world. These configurations seen in Gemini’s AI Co-scientist [56], OpenAI’s Prism [126], and Omni Scientist [149], perfectly align with the soar State Oper- ation Result loop and we argue even aligns with the traditional latent World Model definition. A Global Workspace is not a learned latent encoding like Human Cognition in Machines: A Unified Perspective of World Models17 JEPA [8]. But a Global Workspace and an agentic framework acting upon this global workspace resembles the pipeline of a promptable Video World Models like we will see in Figure 4 (discussed further in Sec. 4, or World Action Mod- els like we will see in Figure 5 (discussed further in Sec. 5). In the Epistemic World Model setting, the initial world is a static world representing structured knowledge, often with an LLM as an Action Orchestrator (argued to be World Models unto themselves [53, 57]). An LLM agent acts by augmenting its con- text through multimodal rag queries, performing analysis of existing literature, and encoding tool inputs and outputs into a global workspace (i.e. a chat his- tory). This creates a changing world, or in other words a dynamic state space, meeting the traditional definition of the World Model. In the epistemic setting, the LLM’s transition function comes from pretraining, and the “state space” is structured knowledge rather than learned latent vectors differentiating epistemic World Models from latent World Models. Working within this agentic framework with global workspace implementation allows models like Gemini Co-scientist, an LLM and World Model itself, to propose novel trajectories through an evolving state space to perform actual scientific discovery. Drawing from Cognitive Architecture Theory [123] and the Global Workspace Theory (GWT) [10], meta-cognition can be interpreted as the ability to broad- cast, inspect, and regulate intermediate cognitive states across specialized pro- cesses. Notably, while such mechanisms are largely missing from latent World Models, they emerge naturally in epistemic World Models, where shared workspaces are instantiated through external artifacts such as documents, tool outputs, and execution traces. These systems approximate the broadcast and control dynamics described by Baars [10], albeit over distributed and externalized state represen- tations rather than internal latent variables. As we will show in Sec. 6, recent agentic systems for scientific discovery provide early instantiations of such global workspaces, suggesting a viable pathway toward incorporating meta-cognitive ca- pabilities into unified World Models. The typical architecture for such an agentic system is shown in Figure 6. For latent World Models in settings that cannot fully replicate a truly Global Workspace we can begin to approximate this with a mixture-of-experts (MoE) paradigm with a world-aware router function. A router function selecting an expert exhibits some behavior indicative of meta-cognition for self-control. This MoE implementation functions similarly to specialists within multi-agent frame- works and is a growing trend in World Models [191, 166, 76]. Similarly regarding self-monitoring, increasing input modalities (i.e. increasing the amount of spe- cialists) to augment perception from “first person” video alone to include an ad- ditional “third person” video source as just one example. For fully self-evaluating meta-cognition, using imagination over an MoE implementation or including a router function that can decide to re-reason through a problem would function similarly to an evaluation specialist as seen throughout Sec. 6. 18Authors Suppressed Due to Excessive Length 4 World Models for Video Generation Recent advancements have positioned World Models as powerful engines for simulating the visual world in 2-dimensional space [221, 78, 13, 46, 209, 52, 200, 192, 96, 208, 58, 165], 3-dimensional space [169, 98, 206], and hybrid scene geometries [11, 39, 23, 50], as seen in our Table 2 and outlined in recent com- prehensive roadmaps [211]. The defining characteristic of modern video World Models is their ability to generate future states conditioned on current obser- vations and, crucially, latent or explicit actions. Contemporary systems like the open-source Wan [175], alongside models such as Sora [125], and Kling [168], do not merely synthesize pixels; with proper latent representations and training techniques they can learn the underlying physical, causal, and temporal laws governing their simulated environments. In Sec. 4.1 we will review works ac- cording to their contribution to World Model representation (utilizing memory, language, or perception). Then in Sec. 4.2 we will review works according to their contribution to World Model generation and prediction. These models are further subdivided according to the cognitive functions they emulate in their designs. Figure 4 shows us the common architectures seen in works across the domain video World Models for 2D and 3D scene generation. Table 2: A check (✓) indicates the work makes a clear contribution to that cognitive function in the context of video World Models; a cross (✗) indicates it is not a contribution. WorkYear GeometryMem.Perc.Lang.Reas.Imag.Moti.Meta.Description VideoREPA [221]20252D✓✗ Physics-aware video generation via relational distillation. Jang etl al. [78]20262D✗✓✗✓✗ Iterativelydenoise-and-re-noise video latents at inference time to self-correct physics/motion artifacts. V-JEPA [13]20242D✓✗ Self-supervised video learning by pre- dicting masked spatio-temporal re- gions in latent space. VideoWeave [46]20262D✓✗ Splice short captioned videos into synthetic long videos to cheaply train better video-language models. Helios [209]20262D✓✗✓✗ 14B video generation model running real-time on one H100 via context compression and drift-aware train- ing. Marble World Model [169, 98] 20253D✗✓✗✓✗ Multimodal 3D world generator. Garrido et al. [52]20262D✗✓✗ Learn action-conditioned World Models from unlabeled in-the-wild videos by inferring continuous latent actions via inverse dynamics. Cambrian-S [200]20252D✗✓✗ Define spatial supersensing hi- erarchy, benchmark it, and use prediction-error surprise to drive memory/attention in long videos. ViewRope [192]2026 2D+depth✓✗✓✗ Replace pixel-grid positional embed- dings with camera-ray-based RoPE so video World Models stay 3D- consistent across viewpoints continued on next page Human Cognition in Machines: A Unified Perspective of World Models19 Table 2 – continued from previous page WorkYear GeometryMem.Perc.Lang.Reas.Imag.Moti.Meta.Description EgoWM [11]2026Hybrid✗✓✗ Fine-tune pretrained video diffusion models with lightweight action condi- tioning to get cross-embodiment ego- centric World Models. Dream2Flow [39]2025Hybrid✗✓✗ Generate human-interaction videos, extract 3D object trajectories, then have robots track those trajectories to manipulate. Le et al. [96]20252D✗✓✗ Post-train video diffusion with veri- fiable Newtonian rewards to enforce physically correct motion. VJEPA-2 [208]20252D✗✓✗✓✗ Scale V-JEPA to 1M+ hours of video, then post-train a latent action-conditioned World Model on robot data for zero-shot planning. DIAMOND [4]20242D✗✓✗ Conditions on previous latent state and actions to produce future frames. EMERALD [23]2026Hybrid✗✓✗✓✗ MaskGIT-based parallel token pre- diction in spatial latent space for ac- curate yet efficient model-based RL World Models. Gumbsch et al. [58]2024 2D+temporal✓✗✓✗ Learn when the world meaningfully changes via discrete latent dynamics, then build a high-level model that skips between those change points. AdaWorld [50]2025Hybrid✓✗✓✗ Pretrain World Models with self- supervised continuous latent actions from video, then cheaply adapt to new environments by mapping real actions to latent ones. MosaicMem [206]20263D✓✗✓✗ Hybrid 3D-patch + latent memory for video World Models. StereoWorld [165]2026 2D+depth✗✓✗ Generate stereo video natively via camera-aware RoPE and epipolar- constrained attention, grounding ge- ometry from disparity. LTX [61]20242D✗✓✗ Uses video diffusion in a real-time open-source world model. Cosmos-Predict2.5 [3]20252D✓✗✓✗✓✗ unified flow-matching video genera- tor trained on clips for Physical AI simulation, data augmentation, and policy evaluation. GigaWorld-0 [202]20253D.✗✓✗✓✗ World Model data engine integrat- ing video generation and 3DGS with physics simulation LingBot-World [170]20262D✓✗✓✗ Block-causal video generator with sub-second latent rollout for agent training RLVR-World [187]20252D✗✓ Trains World Models with RL us- ing verifiable rewards; enables self- improving World Models TesserAct [228]20253D✗✓✗✓✗ 4D embodied World Model convert- ing generated RGB-D-Normal video to point clouds for action prediction Sparse World [34]20254D✓✗✓✗ Range-Adaptive Perception module learns queries modulated by the ego vehicle with temporal-spatial associ- ations to enable extended-range per- ception 20Authors Suppressed Due to Excessive Length Video World Models (A) Autoregressive Causal Rollout TRAINING Noisy chunk Past Chunks clean latents x₀ (history) teacher-forced KV Cache history compress causal context Causal AR Diffusion self-attn FFN × L transformer causal mask Predicted Chunk denoised output next frame causal attn mask · flow-matching INFERENCE (causal rollout) KV Cache EMA-Sink / rolling compressed history context tokens Few-step AR Diffusion z T denoise z T/2 denoise z 0 × 1–4 steps transformer denoiser sharedε θ Generated Chunk interactive output streaming frames real-time capable append to KV cache→ next chunk Astra (Sun et al. 2025) · Causal Forcing (Zhu et al. 2026) · VideoWeave (Shen et al. 2025) Self-Forcing: Bidir Training→ AR Inference Training: bidir, full-attn DMD ₀ violates injectivity (Huang et al. 2025) Inference: AR, causal mask Causal Forcing, Zhu et al. 2026 ₀ AR teacher ODE, injectivity OK (B) Bidirectional / Masked / Parallel Generation TRAINING Full Context spatiotemporal tokens Spatial Latent z t VQ-VAE 4×4 Bidirectional DiT / MaskGIT self-attn (bidir) FFN × L transformer q₀ q₀ q₀ q₀ q₀ k₀ k₀ k₀ k₀ k₀ full attn All Tokens Predicted all filled · denoised full self-attention · all tokens attend all · cosine sched.τ in [0,1] INFERENCE (parallel decoding) Masked Latent z t spatial tokens cosine scheduleτ + temporal h t (KV cache) MaskGIT Parallel Decoding τ=1 predict τ≈0.7 unmask τ=0 refine cosine schedule · parallel Hi-Qual. Frames parallel denoised Actor Critic reward+val latent RL predict→ mask→ refine · 27 FPS on RTX 3090 EMERALD (Burchi & Timofte, 2025) · Video DiT (Peebles & Xie, 2023) · DIAMOND (Alonso et al. 2024) (C) Promptable Video World Model Language Initial Frame Action Seq. Action Projection Linear SiLU / GELU Linear MLP Z A timestep mod.→ DiT noiseε ~ N(0,I) Latent Diffusion Encoder (down) z T z t z 0 iterative denoising bottleneck Decoder (up) action-cond. U-Net / DiT skip (EgoWM / Dream2Flow) VAE Decoder upsample conv upsample Future Frames video output Dream2Flow (Lee et al. 2025): RGB-D→ video→ depth tracking→ 3D flow→ policy EgoWM (Bagchi et al. 2026) · Marble (Wan et al. 2025) · Dream2Flow (Lee et al. 2025) MemoryPerceptionLanguageReasoning MotivationHypothetical reasoningMeta-cognition (rainbow) Fig. 4. Above are the typical architectures encountered when reviewing video World Models. Human Cognition in Machines: A Unified Perspective of World Models21 4.1 World Model Representation Effective representation in video generation requires a model to accurately cap- ture both the spatial structures and the temporal dynamics of a scene. Within the context of Cognitive Architecture Theory (CAT), machine representation aligns with two primary cognitive functions: perception (state space encoding for world interaction) and memory (state space encoding for temporal coherence and world representation). Perception. During inference for video World Model 2-dimensional scene gen- eration, maintaining coherent perception can be challenging. The Self-Refining Video Sampling framework [78] addresses this by using the video generator as a self-refiner during inference time. This proposes a training-free sampling method that uses a flow matching video generator as its own self-refiner, reducing com- mon artifacts without requiring an external verifier or additional training. Fi- nally, to rigorously evaluate these perceptual capabilities, PhysicsMind [117] in- troduces a benchmark focused specifically on perception and generation tasks comprising on two main tasks: VQA tasks and Video Generation tasks. During training, some Vision-Language model works show properly curated multi-modal training data-alone can align models to physical laws. Order of Chaos implicitly aligns a Vision-language model to physical real-world laws by leveraging a unique instruction fine-tuning dataset called PhysGame [24]. Phys- Game uses Q&A pairs regarding glitches and visual anomalies within video game videos demonstrating that simulated data offers a scalable pathway to advance real-world alignment. Memory Cosmos-Predict2.5 [3] exemplifies unified world model representation, learning a latent World Model, z t+1 ∼ p θ (z t+1 | z t ,a t ,c) that supports Text2World, Image2World, and Video2World within a single flow- matching architecture. A temporally causal tokenizer enables incremental state updates, while RL-based post-training improves physical fidelity. This paradigm scales well and unifies modalities, but relies on implicit dynamics that may be harder to interpret or control. Encoding a consistent representation of the world for 2D and 3D scene gen- eration requires an effective storage and retrieval strategy. To construct a 3- dimensional model of the world, MosaicMem World Model stores and retrieves a collection of image patches lifted into 3D using estimated depth and camera poses, enabling spatially consistent retrieval and view-aligned composition [206]. For consistent 3-dimensional scene generation, the Marble World Model uses a continuous 3D radiance-like representation of a scene, not patches or pix- els [169, 103, 224, 98]. These large set of semitransparent particles known as gaus- sian splats are amongst the highest-fidelity representations for scene generation. In 2-dimensional scene generation, to establish a robust state space encoding for 22Authors Suppressed Due to Excessive Length memory, V-JEPA [13] introduces a new space encoding for world representation that is trained solely using a feature prediction objective, completely bypassing the use of pretrained image encoders or other sources of supervision. Similarly, VideoREPA [221] refines the model’s internal state representations during train- ing by aligning token-level relations across spatial and temporal dimensions with a Token Relation Distillation (TRD) loss. In doing so, VideoREPA aligns the spatial and temporal token relations of generative diffusion models with robust representations from self-supervised foundation models. Beyond structural balance, state-of-the-art models address temporal memory bottlenecks through context expansion, autoregressive diffusion, and hierarchi- cal chunking. VideoWeave [46] tackles long-context degradation efficiently via data-centric synthetic splicing, concatenating multiple short, captioned video segments to force the model to track persistent latent states across complex nar- rative transitions. Alternatively, Video-GPT unifies diffusion with autoregressive memory via "Next Clip Diffusion", enforcing strict causal clip-to-clip dependence using historical clean clips to maintain exceptionally stable internal memory over long-horizon video generation [237]. A critical memory challenge involves converting bidirectional video models to generate long videos autoregressively. Self-Forcing [71] addresses the train-test exposure bias by simulating inference conditions during training, performing au- toregressive rollouts with rolling KV caching to condition future frames on the model’s own self-generated past. Building on this, Causal Forcing [234] bridges the architectural gap by utilizing an autoregressive teacher for ODE distillation, ensuring strict causal history mapping. Reward Forcing [115] takes a structural approach to long-term memory via the "EMA-Sink" mechanism, which main- tains fixed-size context tokens. Finally, some works achieve real-time speeds due to their effective and efficient world representation [61, 209]. Helios [209] ap- proaches long-horizon memory through a highly optimized deep compression flow. Rather than relying on standard anti-drifting memory heuristics like self- forcing, Helios heavily compresses the historical and noisy context and explicitly simulates drifting during training. 4.2 World Model Prediction and Generation The generation phase tests a World Model’s ability to extrapolate from its repre- sentations to create novel, coherent video sequences. Within the CAT framework, this phase heavily utilizes sequential logic for reasoning, and hypothetical rea- soning or imagination. Figure 4 shows the various architectures encountered in this review for 2D and 3D scene generation. Reasoning. Within video World Models, reasoning governs how future states are generated under physical and temporal constraints. Architecturally, these mechanisms align with the paradigms illustrated in Figure 4, each corresponding to a different instantiation of the world transition function W θ . Together, these paradigms define a spectrum from strictly causal reasoning to globally optimized Human Cognition in Machines: A Unified Perspective of World Models23 trajectory generation, with action-conditioned models bridging toward embodied control. Autoregressive Causal Rollouts. Figure 4(a) captures autoregressive World Models, where future states are generated sequentially by conditioning on prior latent states, z t+1 = W θ (z ≤t ), enforcing causality through masked attention and rolling context windows [235, 234, 46]. Reasoning in this paradigm emerges as a chain of locally consistent transitions, though compounding error over long horizons remains a central chal- lenge. Bidirectional / Masked Generation. Figure 4(b) shows bidirectional or diffusion-based World Models, which instead model the joint distribution over trajectories, z 1:T ∼ p θ (z 1:T ), and refine all tokens in parallel via iterative denoising [132, 4, 23]. Here, rea- soning is enforced globally rather than causally, allowing the model to correct inconsistencies across space and time during generation. Promptable / Action-Conditioned World Models. Figure 4(c) depicts action-conditioned World Models, where transitions are explicitly controlled by inputs such as actions, language, or initial observations, z t+1 = W θ (z t ,a t ,c), with c denoting conditioning signals (e.g., prompts or frames) [39, 169, 98]. These architectures most closely align with embodied reasoning. For example, EgoWM [11] injects motor commands into pretrained diffusion backbones, while Dream2Flow [39] enforces geometric causality through 3D object flow, grounding predictions in physically meaningful transformations. Learning Latent Action and Geometry. Several works enhance reason- ing by explicitly structuring latent dynamics. Garrido et al. [52] jointly learn inverse and forward models, a t ≈ f −1 (z t ,z t+1 ), z t+1 = f(z t ,a t ), enabling the discovery of continuous latent actions from uncurated video. Simi- larly, ViewRope [192] embeds camera ray geometry directly into attention, aug- menting the transition function W θ with explicit spatial constraints to prevent geometric drift. Hierarchical and Sim-to-Real Reasoning. For sim-to-real transfer, rea- soning must remain consistent under domain shift. This motivates hierarchical World Models that decompose planning across temporal scales, z (l) t+1 = W (l) θ (z (l) t ,z (l+1) t ), where higher-level abstractions guide lower-level transitions. Gumbsch et al. [58] demonstrate that such multi-scale reasoning maintains causal consistency across 24Authors Suppressed Due to Excessive Length long horizons while preserving executability in real-world settings. Dream2Flow further contributes by constraining transitions through 3D object flow, ensuring that imagined trajectories correspond to physically realizable actions. Imagination. Hypothetical reasoning in training, using traversable latent spaces, allows for sim-to-real transfer of video World Models conditioned on actions, to Genie [177] was among the first models that moved towards an interactive en- vironment rather than just a video generation. Genie is an 11-billion-parameter foundation World Model trained entirely via unsupervised learning on unla- beled internet videos. It demonstrates that it is possible to achieve frame-by- frame controllability without the ground-truth action labels or domain-specific heuristics traditionally required in robotic World Models. Expanding on this, Dream2Flow [39] extracts 3D object flow from dreamed videos to serve as an intermediate representation for open-world manipulation. By translating AI- generated video rollouts into executable trajectory tracking actions without task- specific demonstrations, it bridges the gap between hallucinated video generation and physical robotic control. In this paradigm models focus on generating high-fidelity future states to facilitate cross-domain transfer. Dream2Flow [39] bridges video generation with robotic manipulation by first generating human-interaction videos and then ex- tracting 3D object flow as a modality-agnostic intermediate representation. Ada- World [50] addresses the adaptation bottleneck directly by pretraining World Models with self-supervised continuous latent actions extracted from unlabeled video, then cheaply adapting to new environments by mapping real actions to the learned latent action space rather than retraining from scratch. Finally, EMER- ALD [23] enhances DIAMOND [4] and Dreamer-style agents with MaskGIT- based parallel token prediction in a structured spatial latent representation al- low traversal, jointly encoding spatial structure alongside temporal dynamics to improve accuracy and efficiency in model-based RL for sim-to-real transfer. As- tra [235] extends this line of work with an autoregressive denoising framework that supports long video horizons sufficient for future prediction comparable to motor control, addressing the practical requirement that imagined rollouts for real-world Motivation. Le et al. [96] explicitly aligns generated videos with physical laws by introducing verifiable, physics-grounded reward functions during post- training. The model’s internal generation is externally driven by a reward signal that actively penalizes violations of Newtonian gravity and momentum. Similarly addressing physics alignment without foundational retraining, VJEPA-2 [208] utilizes a latent World Model with strong intuitive physics understanding as an inference-time reward function to steer candidate video generation trajectories. By scoring and filtering predictions based on physical plausibility, it acts as a motivating critic that ensures the generative model adheres to real-world dy- namics. In terms of spatial and temporal awareness, Cambrian-S [200] advances spatial supersensing by employing a predictive sensing mechanism that uses pre- Human Cognition in Machines: A Unified Perspective of World Models25 diction error, or surprise as with Active Inference [48, 133, 47], to drive memory and event segmentation in unbounded visual streams. 5 Embodied World Models Embodied World Models not only demand visual understanding and linguistic reasoning, but also perceive, act, and anticipate how their actions reshape the physical world. This physical grounding constraint distinguishes embodied World Models from their video generation counterparts (Section 4). A video World Model succeeds when its generated frames are photorealistic and temporally coherent. However, an embodied World Model succeeds when its predictions are physically accurate enough to guide a body with mass, kinematics, and contact surfaces toward successful task execution. These demands on Embodied World Models heighten the demands for a Unified World Model, as seen in Figure 3, capable of emulating a broad range of cognitive functions. We explore Embodied World Models in robotics [87, 203, 210, 167, 17, 75, 190, 99, 164, 178, 116, 70, 114, 148, 68, 232, 185, 28, 44, 19, 233, 72, 59, 146, 219, 145, 189, 12, 231, 119, 77, 9, 25, 122, 184, 82, 147, 162, 183], navigation with exploration [82, 147, 162], and autonomous driving [230, 220, 222, 34, 218, 105, 31, 229, 102, 26, 16, 227, 176, 100, 201, 121, 51, 66, 143, 212, 65, 29, 137, 69, 80, 86, 21, 96, 179]. Perception must encode contact geometry and six-degree-of-freedom pose, not merely visual appearance; Memory may need to maintain a persistent, spatially-grounded model of the physical world across interactions, not merely temporal coherence across frames. Reasoning may need to capture physical causality and force propagation rather than narrative causality; and Imagina- tion must generate action conditioned futures that are physically executable, not merely visually plausible. We explore these novelties following the CAT structure of our unified framework: world representation in Sec. 5.1 and world generation in Sec. 5.2. Table 3: Embodied World Model works in Section 5 mapped to Cognitive Architec- ture Theory (CAT) functions (Memory, Perception, Language, Reasoning, Imagi- nation, Motivation, Meta-cognition). A check (✓) indicates a clear contribution; a cross (✗) indicates otherwise. WorkYear App.Mem.Perc.Lang.Reas.Imag.Moti.Meta.Description Cosmos Policy [87]2026 Robot.✗✓✗ Post-trains Cosmos Predict-2 for visuomotor robot control and planning DreamZero [203]2026 Robot.✓✗✓✗ World Action Model jointly pre- dicting video and actions; zero- shot policy from video pretrain- ing Fast-WAM [210]2026 Robot.✗✓✗ Decouples action prediction from video generation at inference for faster WAM deployment continued on next page 26Authors Suppressed Due to Excessive Length Table 3 – continued from previous page WorkYear App.Mem.Perc.Lang.Reas.Imag.Moti.Meta.Description GigaBrain-0 [167]2025 Robot.✗✓✗✓✗ VLA trained on world-model- generated data with RGBD input and embodied chain-of- thought Motus [17]2025 Robot.✓✗✓✗ Unified latent action World ModelviaMixture-of- Transformers; optical flow as embodiment-agnostic action π0.5 [75]2025 Robot.✗✓✗ Flow-matching VLA decom- posing language into causally grounded multi-step physical plans LingBot-VLA [190]2026 Robot.✓✗ Cross-morphology VLA on 20,000h bimanual data with geometry-aware depth distilla- tion LingBot-VA [99]2026 Robot.✓✗✓✗ Interleaves video and action to- kens for joint imagination and action decoding VLA-JEPA [164]2026 Robot.✓✗ Leakage-free JEPA grounding vi- sual encoder in action-relevant dynamics VAGEN [178]2025 Robot.✗✓✗✓✗ RL-structured World Model rea- soning into state estimation and transition modeling for VLM agents BeingH VLA [116]2026 Robot✗✓✗✓✗ A foundational VLA for robust cross-embodiment generalization across diverse robotic platforms PointWorld [70]2026 Robot.✗✓✗✓✗ Unifies state and action as 3D point flows with MPC over imag- ined scene deformations ManiGaussian [114]2024 Robot.✗✓✗✓✗ Dynamic 3DGS World Model predicting future Gaussian scenes under action for manipu- lation RoboScape [148]2025 Robot.✗✓✗✓✗ Physics-informed World Model jointly learning video, depth, and keypoint dynamics EnerVerse-AC [68]2025 Robot.✓✗✓✗ Chunk-wise autoregressive video diffusion with sparse memory and 4DGS for action-conditioned prediction UWM [232]2025 Robot.✓✗✓✗ Couples video and action dif- fusion in one transformer; pretrained on video-only and video+action data GR-1 [185]2024 Robot.✓✗✓✗ GPT transformer pretrained on 800K Ego4D clips jointly pre- dicting actions and future frames GR-2 [28]2024 Robot.✓✗✓✗ Scaled video-language-action model (719M) achieving 97.7% success across 100+ real tasks UniPi [44]2023 Robot.✗✓✗✓✗ Text-conditioned video diffusion as policy; extracts actions via in- verse dynamics SuSIE [19]2024 Robot.✗✓✗✓✗ Image-editing diffusion synthe- sizing subgoal images for goal- conditioned manipulation policy continued on next page Human Cognition in Machines: A Unified Perspective of World Models27 Table 3 – continued from previous page WorkYear App.Mem.Perc.Lang.Reas.Imag.Moti.Meta.Description IRASim [233]2024 Robot.✗✓✗✓✗ Diffusion transformer with frame-level action conditioning for manipulation simulation LaDi-WM [72]2025 Robot.✗✓✗✓✗ Predicts latent state evolution via diffusion; more generalizable than pixel-level prediction FlowDreamer [59]2025 Robot.✗✓✗✓✗ RGB-D World Model using opti- cal flow as explicit physically in- terpretable motion supervision MWM [146]2023 Robot.✓✗✓✗ Decouples MAE visual repre- sentation from RSSM dynamics; 81.7% on Meta-World tasks PIVOT-R [219]2024 Robot.✗✓✗✓✗ Waypoint-aware World Model focusing prediction on task- relevant key states Plan2Explore [145]2020 Robot.✗✓✗ World Model exploration via en- semble disagreement for zero- shot task adaptation DayDreamer [189]2022 Robot.✓✗✓✗ Dreamer-style World Model learning directly on physical robots from sparse rewards Dream to Manipulate [12] 2024 Robot.✗✓✗✓✗ Compositional 3DGS with object decomposition for imagination- based imitation learning RoboDreamer [231]2024 Robot.✗✓✗✓✗ Compositional diffusion World Model factorizing language in- structions into task primitives GenRL [119]2024 Robot.✗✓✗✓✗ Foundation World Models for generalization in embodied RL via multimodal priors DreamGen [77]2025 Robot.✗✓✗✓✗ Fine-tunes Cosmos Predict-2.5 as synthetic robot data engine for policy training V-JEPA 2 [9]2025 Robot.✗✓✗✓✗ Representation-space prediction on 1M video hours; zero-shot robot planning via MPC WorldVLA [25]2025 Robot.✓✗✓✗ Autoregressive action World Model unifying video prediction and VLA action generation DreamWaQ [122]2023 Robot.✗✓✗ Implicit terrain imagination from proprioception for robust quadruped locomotion DreamerNav [184]2025 Robot.✓✗✓✗ DreamerV3 for quadruped nav- igation with depth and occu- pancy; zero-shot sim-to-real BADGR [82]2021 Robot.✗✓✗✓✗ Self-supervised terrain affor- dance learning for outdoor navigation via prediction-based MPC ViNT [147]2023 Robot.✓✗✓✗ Visual navigation foundation model with diffusion subgoal proposals across robot platforms NoMaD [162]2024 Robot.✗✓✗✓✗ Unifies goal navigation and ex- ploration with diffusion policy and goal-masking EVA [183]2026 Robot✗✓✗ Trains a video world model with an inverse dynamics reward sig- nal aligns to physical constraints in world dynamics. continued on next page 28Authors Suppressed Due to Excessive Length Table 3 – continued from previous page WorkYear App.Mem.Perc.Lang.Reas.Imag.Moti.Meta.Description OccWorld [230]2023 AD✓✗ VQVAE tokenizer with GPT transformer for joint 3D occu- pancy and ego trajectory fore- casting Copilot4D [220]2023 AD✗✓✗ Discrete diffusion over BEV Li- DAR tokens; 65% Chamfer dis- tance reduction BEVWorld [222]2024 AD✓✗ Multimodal tokenizer fusing camera and LiDAR into unified BEV latent for forecasting SparseWorld [34]2025 AD✓✗ Sparse dynamic queries modu- lated by ego-vehicle state for adaptive scene memory Epona [218]2025 AD✓✗ Decouples temporal memory from spatial generation via causal transformer and twin diffusion LiDARCrafter [105]2025 AD✗✓✗ Language-conditioned tri-branch diffusion for 4D LiDAR scene generation DrivingGPT [31]2024 AD✗✓✗ Interleaved image–action tokens unifying World Modeling and planning as next-token predic- tion DOE-1 [229]2024 AD✗✓✗ Closed-loop end-to-end AV model using free-form text as perceptual interface LAW [102]2024 AD✗✓✗ Self-supervised latent prediction of future scene features from ob- servations and ego trajectories MILO [26]2021 AD✗✓✗ Model-based imitation learning mitigating covariate shift via of- fline data KG-based WM [16]2025 AD✓✗ Knowledge graphs with sensor data for material-aware obstacle reasoning in AVs PWM [227]2025 AD✗✓✗ Collaborative state-action pre- diction for anticipatory planning in AVs AdaWM [176]2025 AD✗✓✗ Adaptive World Model planning with dynamic model selection for autonomous driving Think2Drive [100]2024 AD✗✓✗ Efficient RL via latent World Model imagination in CARLA- v2 Raw2Drive [201]2025 AD✗✓✗ End-to-end AV aligning RL pol- icy with World Model imagina- tion from raw sensors Dream to Drive [121]2025 AD✗✓✗ Analytic World Model for dreamer-style vehicle control without environment interaction Dream2Drive [51]2024 AD✗✓✗ RL in predictive World Model imagination with intention- aware latent states GAIA-1 [66]2023 AD✗✓✗ Generative driving World Model for physically grounded imagined video rollouts GAIA-2 [143]2025 AD✗✓✗ Controllable multi-view genera- tive World Model for scalable AV synthetic data continued on next page Human Cognition in Machines: A Unified Perspective of World Models29 Table 3 – continued from previous page WorkYear App.Mem.Perc.Lang.Reas.Imag.Moti.Meta.Description Dream4Drive [212]2025 AD✗✓✗ Repurposes World Model imagi- nation as synthetic data for per- ception tasks MoSim [65]2025 AD✗✓✗ Motion-groundedsimulation generating diverse controllable traffic scenarios Large Video Planner [29] 2025 AD✗✓✗ Foundation-scale video model for zero-shot robot video plans con- verted to actions VL-Safe [137]2025 AD✗✓✗ VLM-guided safety rewards steering imagined rollouts for constrained AV policy SafeDreamer [69]2023 AD✗✓✗ Dreamer safe RL pairing imag- ined trajectories with safety es- timations IRL-VLA [80]2025 AD✗✓✗ Inverse RL reward World Model for efficient closed-loop reward computation in VLA training InDRiVE [86]2025 AD✗✓✗ Intrinsic disagreement reward in Dreamer MBRL for curiosity- driven vehicle exploration RDAR [21]2025 AD✗✓✗ Reward-driven relevance estima- tion prioritizing safety-critical agents in AV planning NewtonRewards [96]2025 AD✗✓✗ Physics-grounded reward func- tions penalizing violations of Newtonian laws in video gener- ation Latent-WAM [179]2026 AD✗✓✗✓✗ 5.1 World Model Representation Perception In embodied world models perception must encode the physical structure of the scene rather than visual appearance alone, making 3D occu- pancy a natural representation for World Models as it is expressive, efficient, and versatile across both vision and LiDAR inputs [230]. OccWorld operational- izes this by learning in 3D semantic occupancy space, using a VQVAE-based scene tokenizer to produce discrete scene tokens that jointly forecast future oc- cupancy and ego trajectory through a GPT-like spatial-temporal transformer, all without requiring instance or map annotations [230]. Building on the same tokenize-then-predict philosophy, Copilot4D [220] further addresses the scalabil- ity of this perceptual pipeline by applying discrete diffusion over BEV tokens derived from raw LiDAR point clouds, reducing prior state-of-the-art Chamfer distance by over 65% for one-second prediction across multiple benchmarks [220]. Memory A central challenge in embodied World Models is maintaining a com- pact yet sufficiently rich state space encoding of the scene across time. BEVWorld addresses this by compressing heterogeneous multimodal inputs including cam- era imagery and LiDAR point clouds into a unified Bird’s Eye View latent space through a self-supervised multimodal tokenizer, enabling temporally consistent future scene forecasting via a latent BEV sequence diffusion model conditioned 30Authors Suppressed Due to Excessive Length Embodied World Models (A) Direct Vision-Language-Action (VLA) RGB / RGB-D visual tokens multi-view input Language Task task instruction zero-shot goals VLM / Policy self-attn (vis + lang) cross-modal attn FFN x L layers VLM backbone diffusion headflow matching continuous action tokens Action Output 7-DoF joint delta SE(3) end-effector wrist pose stream π0.5 (Black et al. 2025) · LingBot-VLA (Li et al. 2026) · GigaBrain-0 (Zhao et al. 2025) (B) Dreamer-style Latent World Model Obs+History pixels proprioception Encoder CNN / ViT downsample project z t RSSM / WM GRU dynamics priorposterior h t s t Imagined Rollout s 1 s 2 s L L-step latent sim. · no real env. Policy / Value MPC / CEM reward-max. reward gradient→ world model update DayDreamer (Wu 2022) · MWM (Seo 2022) · DreamerNav (Ramrakhya 2025) · V-JEPA 2 (Assran 2025) (C) World Action Model (WAM) Visual Obs RGB / RGB-D patch tokens Language Task goal / instruction text tokens Shared Generator token interleave vavavavav self-attn (joint v+a) FFN x L layers diffusionAR decode joint video + action generation Future Frames predicted video visual rollout Actions policy tokens executable traj. DreamZero (Shi et al. 2026) · LingBot-VA (Li et al. 2026) · WorldVLA (Chen et al. 2025) · Fast-WAM (D) Planner over Structured Geometry RGB-D / LiDAR point clouds occupancy grids Geom. Encoder 3D feature extract downsample / pool voxels3DGS 3D latent repr. Geometric WM 3D state predict scene deformation action-conditioned next-state geom. MPC / Planner CEM / sampling cost evaluation trajectory select action tokens PointWorld (Gu et al. 2026) · TesserAct (2025) · ManiGaussian (2024) · OccWorld · Copilot4D MemoryPerceptionLanguageReasoning MotivationHypothetical reasoningMeta-cognition (rainbow) Fig. 5. Above are the typical architectures encountered when reviewing embodied World Models. Human Cognition in Machines: A Unified Perspective of World Models31 on action tokens [222]. While BEVWorld [222] establishes a shared spatial mem- ory across modalities, it relies on static grid-based representations that struggle to adapt to the dynamic and continuous nature of real driving environments. SparseWorld addresses this limitation directly by replacing fixed grid embed- dings with sparse and dynamic queries modulated by the ego vehicle’s state, allowing the memory encoding to scale its perception range with vehicle speed and adapt to foreground object dynamics rather than treating all voxels uni- formly [34]. Epona takes a complementary approach to the memory problem by identi- fying that conventional video diffusion models entangle temporal memory with spatial generation, leading to error accumulation in long-horizon rollouts. By de- coupling the two through a GPT-style causal transformer that handles temporal context in compressed latent space separately from twin diffusion transformers that handle spatial rendering and trajectory generation, Epona achieves stable long-duration prediction with a 7.4% FVD improvement over prior works [218]. Further advancing perception, LaDi-WM [72] finds that predicting the evolution of the latent space is easier to learn and substantially more generalizable than directly predicting pixel-level images in diffusion models. Language In the CAT framework, language serves as a semantic and symbolic encoding that connects human intent to world state, and in embodied driving this role becomes particularly concrete. LiDARCrafter [105] demonstrates this most directly by using free-form natural language instructions as the entry point for 4D LiDAR World Modeling, parsing text into ego-centric scene graphs that condition a tri-branch diffusion network to generate object structures, motion trajectories, and geometry, with an autoregressive module then extending the re- sult into temporally coherent LiDAR sequences [105]. Where LiDARCrafter [105] uses language to control a geometric representation, DrivingGPT [31] uses it to unify the entire driving pipeline by constructing a multimodal driving language from interleaved image and action tokens, treating World Modeling and trajec- tory planning as a single next-token prediction problem over this shared symbolic vocabulary [31]. DOE-1 [229] takes this unification further by closing the loop en- tirely, using free-form text scene descriptions as the perceptual interface and au- toregressively generating perception, prediction, and planning tokens within one multimodal transformer, achieving the first closed-loop end-to-end autonomous driving model under this paradigm [229]. 5.2 World Model Generation Reasoning. Architectures for physical World Models primarily differ along the following axes (with some sampled in Fig. 5): Geometric World Models. PointWorld [70] and other geometric world models [70, 114, 148, 230, 222, 105] represent the opposite extreme, ground- ing reasoning directly in 3D structure by predicting scene flow ∆x = f θ (x,a t ) over point clouds. This enforces physically meaningful state transitions and en- ables cross-embodiment generalization without task-specific heads. However, it 32Authors Suppressed Due to Excessive Length depends on strong geometric supervision and may be less flexible for abstract or semantic tasks. Language-conditioned planning. π 0.5 [75] treats reasoning as long-horizon planning conditioned on language, implicitly learning a 1:T ∼ π θ (a 1:T | z 0 ,ℓ) where ℓ encodes task structure. This enables coherent multi-step behavior across diverse environments, but places significant burden on the latent representation to maintain causal consistency over long horizons. VLAs [75, 190, 99, 164, 178, 116, 25, 80] excel in reactive and instruction-following settings but rely on either external planners or hybridization with World Action Models (WAMs) when long-horizon physical reasoning or counterfactual simulation is required. Autoregressive world-action models. Because VLAs are reactive and under-supervised and map action directly, with no explicit model of future re- searchers turn to World Action Models (WAMs) [26, 16, 227, 65, 203, 210, 17, 99, 25, 179]. Two such works, LingBot-VA [99] and DreamZero [203] unifies imagination and control by modeling joint sequences (z t+1 ,a t )∼ p θ (z t+1 ,a t | z ≤t ) with causal or block-causal attention. These approaches improve temporal con- sistency and enable efficient rollout, but remain sensitive to representation qual- ity and training stability. Recent works explore combine WAMs and VLAs [99, 25, 203, 116]. One of these recent works rejects the standard WAM implementations as being flawed because pixel-space simulation is inefficient (e.g. VLAs can be up to 60x cheaper in compute) and imperfectly predicted pixels can lead to bad downstream ac- tions [1]. The solution they argue is to combine VLAs and WAMs by bridging direct action prediction and world modeling through a shared latent space. It is this efficiency in reasoning that allows multiple queries of hypothetical reason- ing in some settings which was previously considered unobtainable in traditional WAM settings. Representation-driven reasoning. VLA-JEPA [164] highlights that rea- soning quality depends critically on the learned state space, enforcing dynamics- consistent representations z t = φ(o t ) that are invariant to irrelevant visual vari- ation. This improves generalization, but shifts complexity into representation learning as recent works show [146, 9, 164, 72, 176, 102]. Planning-centric extensions. Across paradigms, there is a convergence toward explicit planning over learned dynamics [70, 145, 176, 102, 51, 137, 69] to predict actions in the form, a ∗ t = arg max a t:t+H E " H X k=0 r(z t+k ) # , as seen in model-predictive control (PointWorld), self-supervised World Models (LAW [102]), and latent RL pipelines [176]. These approaches improve control- lability and robustness, but depend on accurate forward models. Human Cognition in Machines: A Unified Perspective of World Models33 Figure 5 can thus be interpreted as a spectrum: generative models offer scal- ability and multimodal unification; geometric models provide strong physical grounding; and planning-based approaches enable controllable long-horizon rea- soning. Most recent systems hybridize these axes, suggesting that effective rea- soning emerges from combining expressive latent models with structured repre- sentations and explicit planning. Imagining. In embodied World Models, imagination manifests as dreaming [189, 122, 184, 51, 121, 69, 86], as well as action-conditioned hypothetical reasoning, where agents simulate future trajectories in latent space to guide policy learn- ing and planning [77, 29, 66, 143, 212, 121]. Works such as Think2Drive [100] and Raw2Drive [201] leverage World Models to generate imagined rollouts for training driving policies, while autoregressive approaches model multiple prob- abilistic futures to reason under uncertainty [194]. Dream2Drive [51] further demonstrates this paradigm by operating within a learned imagination space (PIWM), using intention-aware latent states to evaluate candidate trajectories for urban navigation [51]. More recent systems integrate imagination directly into the training loop, enabling policies to be optimized through interaction with an internal simulator rather than the real environment [55]. Beyond policy learning, imagination enables safety-aware planning and scal- able data generation. World Models can pair imagined roll outs with safety estimates to guide actor–critic optimization under additional constraints, such as VLM-based safety signals [137]. At larger scale, foundation video models gen- erate zero-shot trajectory plans from internet-scale data, which can be converted into executable robot actions [29], while broader dreaming pipelines [121], includ- ing Cosmos-Drive-Dreams [138], GAIA-1, GAIA-2 [67, 143], and Dream4Drive [213] use synthetic rollouts as training data. Benchmarks such as WorldLens evaluate the physical fidelity of these imagined trajectories [106], reinforcing imagination as a core mechanism for bridging perception, reasoning, and action in embodied settings. Motivation. Motivation in embodied World Models defines how agents evalu- ate imagined trajectories during learning and control [70, 80, 86, 145, 176, 21, 102, 51, 137, 69]. IRL-VLA [80] learns a lightweight reward World Model via Inverse Reinforcement Learning for efficient closed-loop optimization [80], while InDrive [86] leverages intrinsic disagreement-based rewards within a Dreamer- style MBRL framework to drive exploration [86]. Additional approaches refine reward signals for task relevance in autonomous driving settings [21], reflect- ing a shift toward learned and uncertainty-aware objectives over hand-crafted rewards. Safe-Dreamer motivate models to conform to safety criteria by incor- porating Lagrangian-based methods into World Model planning processes [69]. 34Authors Suppressed Due to Excessive Length Epistemic World Models (A) Agentic Framework Scientist Research Goal Novel Research Research Orchestrator self-attn FFN tool-call plan + route Evaluation Specialist meta-review self-critique Analysis Specialist data analysis Python interp. Vision Specialist chart / figure visual parsing Literature Agent Web / RAG search Global Workspace KV memorymessage bustask queue SciSciGPT (Wang 2025) · AI Co-Scientist (Gottweis 2025) · OmniScientist (Shao 2025) (B) World Representation / Global Workspace Code Copilot >_ Runtime Errors Global Workspace Code Debugger </> Code Code Writer Tools Interface Interpretable Reasoning CI / Linter RNA-seq / Biology Cell Sequence Global Workspace Cell Annot. Cell Sequence Embeddings Interpretable Reasoning CellExpress Scientific Discovery Research Corpus Global Workspace Literature Review Experiment Data Analysis Engine Hypotheses Reasoning Engine Interpretable Reasoning Peer Review Claude Code (2026) · CellAtria (2025) · AI Co-Scientist (2025) · SciSciGPT (2025) · OmniScientist (2025) MemoryPerceptionLanguageReasoning MotivationHypothetical reasoningMeta-cognition (rainbow) Fig. 6. Above are the typical architecture and Global Workspace frameworks encoun- tered when reviewing World Models for scientific discovery used with a human-in-the- loop subject-matter expert. Human Cognition in Machines: A Unified Perspective of World Models35 6 Epistemic World Models World Models in agentic frameworks extend beyond latent-dynamics prediction to operate over the scientific process itself, orchestrating multi-step workflows that integrate perception, memory, reasoning, and tool use in service of discov- ery. In contrast to latent World Models, which encode external environments into learned state representations with explicit transition dynamics, these systems treat structured knowledge as the world itself. In such epistemic World Mod- els, the environment is a knowledge space defined by literature, databases, and experimental outputs, where domain expertise induces the state space and scien- tific artifacts act as observations. Rather than evolving an internal latent state, the agent updates this epistemic state through reasoning and tool-mediated op- erations, effectively controlling externalized cognitive functions. This distinction highlights two complementary strategies for World Model- ing: latent compression of external environments versus explicit maintenance and updating of structured knowledge. Importantly, epistemic World Models also in- stantiate a form of global workspace, where intermediate results, tool outputs, and shared context are broadcast across agents and processes. In the sense of Baars’ Global Workspace Theory [10], these systems approximate meta-cognitive control by enabling self-monitoring, self-control with selective routing of infor- mation across distributed components and even self-evaluation. As introduced in Sec. 3.3, such workspace-based mechanisms provide a candidate solution to the lack of meta-cognition in latent World Models. Agentic systems therefore em- phasize reasoning, tool coordination, and persistent memory, while introducing early forms of meta-cognitive control through execution tracing and human-in- the-loop feedback. These capabilities suggest a concrete pathway toward ad- dressing the gaps in meta-cognition and motivation outlined in Sec. 3.3, by externalizing and structuring the global workspace required for self-monitoring and control. In this section, we review World Models for Scientific Discovery through the lens of Cognitive Architecture Theory (CAT), emphasizing Human–AI Collab- oration within multi-agent frameworks. As in previous sections, Sec. 6.1 covers contributions to World Model representation (memory, language, perception), while Sec. 6.2 examines generation and prediction. The reviewed works are fur- ther subdivided according to the cognitive functions they emulate in their de- signs. Figure 6(a) illustrates common multi-agent architectures, with (b) high- lighting Global Workspace-like structures for shared domain knowledge. Table 4 summarizes all works within the CAT framework. 6.1 World Model Representation Language. In human–AI collaboration systems, language is not only a medium for generation (e.g., code or hypotheses), but a substrate for shared world repre- sentation, coordination, and self-evaluation, consistent with notions of a global workspace [10] as discussed in Sec. 3.1. In co-science systems such as Gemini Co- scientist [56], OpenAI Prism [126], SciSciGPT [150], and OmniScientist [149], 36Authors Suppressed Due to Excessive Length Table 4. Epistemic World Model works in this report mapped to Cognitive Archi- tecture Theory (CAT) functions and subtasks. In the Subtask column, T = Trust, Human-alignment, and Interpretability, C = Software Co-pilots, M = Medical Re- search and Application, and S = Social Science. Composite labels indicate that a work contributes to multiple Human–AI collaboration subtasks. A check (✓) indicates the work makes a clear contribution to that CAT function in the context of collaboration; a cross (✗) indicates it is not a contribution. WorkYear SubtaskMem.Perc.Lang.Reas.Imag.Moti.Meta.Description GeminiCo- scientist [56] 2025 T+C+M✓✗✓✗✓ multi-agent AI collaborator for hy- pothesis generation and refinement OpenAI Prism [126] 2026T+C✓✗✓ LaTeX-native workspace for scien- tific writing and collaboration SciSciGPT [150] 2025 T+C+S✓✗✓ agentic science-of-science analyt- ics over literature and structured datasets OmniScientist [149]2025 T+C+S✓ human–AI scientist ecosystem with community evaluation (Sci- enceArena) Xie et al. [196] 2025 T+C+S✗✓✗✓ roadmap on AI scientists empha- sizing verification and falsification Agentic Coding Manifests [27] 2025T+C✓✗✓✗✓ repo manifests that externalize project context and rules for cod- ing agents Generative Agents [128] 2023T+S✓✗✓ LLM agents with memory, reflec- tion, and planning for social simu- lation Tsvetkovaet al. [173] 2024T+S✗✓✗✓ framework for sociology of mixed human–machine communities Strachanet al. [163] 2024T+S✗✓✗ empirical theory-of-mind bench- mark comparing LLMs and hu- mans Chen et al. [30] 2025T+S✗✓✗✓ surveyofconscious- ness/metacognitiontheories, implementations, and risks in LLMs Kwok et al. [93] 2026T+S✓✗✓✗ explicit shared World Models for reliable human–robot collabora- tion Fung et al. [49] 2025 C+M+S✓✗✓✗ position paper on embodied assis- tants with memory, World Models, and goal inference Mehri Shervedani et al. [159] 2025 T+C+M+S✓✗ multimodal RL interaction man- ager for assistive human–robot col- laboration SpeechAgents [217]2024C+S✗✓✗✓✗ multi-agent spoken interaction controlled by a speech-centric LLM Kumar et al. [92] 2025 T+C+S✓✗ unified speech-to-speech model claims for multilingual, emotional interaction Kim et al. [89] 2025C+S✓✗✓✗ multimodal conversational agent generating engaging speech from audio-visual cues CellAtria [124] 2026M✓✗ Using AI for RNA sequencing and analysis Gentile et al. [54] 2026M✓✗ AI Transcriptomics with human- in-the-loop for gene expression analysis Human Cognition in Machines: A Unified Perspective of World Models37 language encodes hypotheses, project state, and intermediate reasoning, en- abling iterative critique, revision, and verification across multi-agent interac- tions. Rather than serving only as input/output, language acts as a persistent interface through which agents construct, refine, and align their shared World Model. Reliable collaboration depends on maintaining explicit and evolving com- mon ground rather than relying on opaque internal state [93]. Language enables this by making agent intentions and reasoning processes interpretable and con- testable [128]. This aligns with theories of human cognition in which language emerges from the capacity to share intentions, attention, and goals [172, 36, 79]. Viewed through this lens, language functions as an integrative interface linking perception, reasoning, and meta-cognition, rather than as a standalone to- ken prediction mechanism. It is through language that internal representations become shared, inspected, and coordinated, enabling collaborative intelligence over a common World Model. Perception. Many software co-pilots extend perception beyond plain text, treating structured artifacts, speech, and embodied signals as first-class inputs [126, 150, 149, 92, 217, 89, 49, 159]. OpenAI Prism [126] exemplifies document-centric perception by operating directly over LaTeX structure (equations, references, figures), grounding edits in the manuscript’s semantics. SciSciGPT [150] re- frames perception as evidence acquisition, retrieving and parsing literature and structured datasets to support downstream reasoning. Similarly, OmniScien- tist [149] organizes perception into structured representations of scientific knowl- edge. Across these systems, perception shifts from passive input processing to active construction of task-relevant representations. In embodied and assistive settings, perception becomes inherently multi- modal and user-centric [159, 49]. Mehri Shervedani et al. [159] integrate dialogue acts with multimodal signals to infer user intent in collaborative tasks, while Fung et al. [49] emphasize continuous, first-person sensing in wearable or em- bodied assistants. These approaches enable proactive assistance but introduce challenges in reliability, ambiguity resolution, and privacy, making perception a critical bottleneck for safe deployment. In social interaction, perception expands to include communicative sig- nals such as prosody, emotion, and contextual cues [217, 92, 89, 93]. SpeechA- gents [217] treat speech as a primary interaction channel, preserving rhythm and affect rather than reducing communication to text. Kim et al. [89] fur- ther condition interaction on audio-visual signals to improve engagement, while Kwok et al. [93] highlight the importance of grounding ambiguous social cues in shared context for reliable collaboration. Other systems approximate percep- tion through dialogue-history encodings or structured observations (e.g., pose or environment state), trading richness for tractability. Across domains, improvements in trust are less about increasing perceptual bandwidth and more about structuring and grounding perceptual inputs. Sys- tems that expose intermediate representations (e.g., retrieved evidence, struc- 38Authors Suppressed Due to Excessive Length tured documents, or shared context) make perception more interpretable, whereas purely latent encodings of multimodal input remain difficult to audit. As a result, effective co-pilots treat perception not as raw sensing, but as the construction of verifiable, task-aligned state. Memory. Many co-pilots improve memory by externalizing long-horizon con- text into persistent artifacts such as project state, corpora, or structured workspaces [126, 150, 149, 92, 89, 27, 49]. OpenAI Prism [126] treats the LaTeX project itself as working memory, enabling consistent revision across documents, while Omni- Scientist [149] extends memory into a research ecosystem via knowledge graphs and evolving “idea stacks.” In contrast, Agentic Coding Manifests [27] provide a lightweight approach, storing repository conventions and constraints as durable, human-authored memory. Across these systems, a central trade-off emerges be- tween simple explicit memory (manifests), structured external memory (graphs, databases), and implicit conversational state. In medical research settings, explicit long-term memory remains underde- veloped despite its importance. Two works in particular however, provide novel world encoding strategies for their RNA transcription setting [124, 54]. In both, using agentic reasoning for RNA sequence analysis requires representing se- quence annotations or researcher meta-data as written language. Also in both works, they can utilize the sequence of RNA (itself semantically and symbolically salient) as embeddings that LLMs can use to reason with the RNA-text-liek rep- resentation itself. This representation of RNA for a global workspace framework is observable in Figure 6. Social and interactive systems place stronger demands on memory for conti- nuity and shared understanding [92, 49, 93, 128, 159]. Generative Agents [128] ex- emplify this by storing episodic experiences and synthesizing higher-level reflec- tions that guide future behavior. Similarly, Kwok et al. [93] and Mehri Shervedani et al. [159] model memory as shared state (common ground) that must be con- tinuously updated during interaction. In speech systems, longer-term person- alization is often approximated through dialogue history or user-specific em- beddings [92], while embodied assistants require episodic memory to sustain coherent assistance over time [49]. A key distinction is between interpretable shared-memory representations (e.g., common ground) and implicit history en- codings that are harder to audit. Across domains, improvements in trust are closely tied to making memory explicit, structured, and inspectable [56, 126, 27, 93, 128, 159, 149]. Generative Agents [128] provide a clear example by exposing both raw episodic memories and derived reflections, enabling users to trace behavior back to stored expe- rience. Similarly, manifests [27] and shared-state representations [93] external- ize assumptions and context into artifacts that can be inspected and revised. Larger systems such as Prism [126] and OmniScientist [149] extend this idea to document- and ecosystem-level memory. Overall, the dominant pattern is a shift away from opaque latent state toward persistent artifacts that users can inspect, version, and contest. Human Cognition in Machines: A Unified Perspective of World Models39 6.2 World Model Prediction and Generation Imagination. Co-pilots employ imagination to generate candidate hypothe- ses, plans, or expressive outputs that a human or downstream process can select, refine, or test [56, 149, 217, 89, 159]. In scientific collaboration, Gemini Co- scientist [56] produces research hypotheses and experimental proposals, while OmniScientist [149] treats ideation as a first-class module, evolving candidate ideas within a structured knowledge graph. In these settings, imagination func- tions as controlled hypothesis expansion rather than unconstrained generation. In medical and scientific domains, imagination most directly appears as hy- pothesis generation under experimental constraints [56]. Systems such as Gemini Co-scientist propose candidate mechanisms, repurposing strategies, or interven- tion targets that are subsequently filtered by feasibility and empirical validation, reflecting a generate-then-triage workflow aligned with the scientific method. In interactive and social systems, imagination often takes the form of one-to- many generation. For example, Kim et al. [89] generate expressive paralinguistic speech, while Mehri Shervedani et al. [159] employ simulated rollouts via user models to support policy learning. Generative Agents [128] extend this paradigm to multi-agent settings, where imagined actions at the individual level produce emergent group behaviors over time. Across these domains, imagination supports diverse objectives, including expressive communication, social simulation, and action planning. Imagination is not used in isolation. Its utility depends on coupling with downstream selection mechanisms, such as ranking, constraint satisfaction, or evaluation. In practice, trust is not derived from the generative step itself, but from the processes that filter, verify, and prioritize imagined candidates. Reasoning. Reasoning in World Models and co-pilots is typically framed as multi-step problem solving under constraints, often mediated by tool use, work- flows, or multi-agent protocols [126, 150, 56, 149, 196, 92, 27, 49, 159]. Sys- tems span a spectrum from structured, tool-grounded reasoning to delibera- tive, multi-agent reasoning. OpenAI Prism [126] performs constrained reasoning over structured artifacts (e.g., LaTeX projects), while SciSciGPT [150] oper- ationalizes reasoning as an end-to-end empirical pipeline (decomposition, re- trieval, computation, and visualization). In contrast, Gemini Co-scientist [56] and OmniScientist [149] emphasize deliberation through multi-agent protocols such as debate, reflection, and workflow orchestration. At the policy level, Gem- ini Co-scientist [56] use agents to recognize novel research trajectories by debat- ing hypotheses, ranking ideas with an Elo-style sytsem filtering our redundant ideas, while Mehri Shervedani et al. [159] frame reasoning as action selection under uncertainty, learning dialogue and intervention policies for assistive tasks. In medical and assistive settings, reasoning bifurcates into scientific reason- ing and decision-making under uncertainty. Gemini Co-scientist [56] focuses on evidence-backed hypothesis generation and refinement via multi-agent deliber- ation, while Mehri Shervedani et al. [159] optimize real-time interaction poli- cies for successful task completion. This contrast highlights reasoning as either 40Authors Suppressed Due to Excessive Length hypothesis validation or policy optimization. In RNA transcription medical re- search, models reason from representations of RNA to profile genetic sequences for personalized medicine or scientific discovery [124, 54]. In social and collaborative contexts, reasoning extends beyond individual cognition to include theory-of-mind inference, experimental methodology, and system-level dynamics [150, 49, 93, 128, 159, 173, 163, 30, 196]. SciSciGPT [150] exemplifies reasoning as transparent scientific workflows, while Kwok et al. [93] emphasize explicit shared models for reasoning about human goals. Strachan et al. [163] evaluates theory-of-mind capabilities, exposing systematic failure modes in social reasoning tasks. At a broader scale, Tsvetkova et al. [173] model mixed human–machine systems as dynamical processes, and Xie et al. [196] identifies verification and falsification as central bottlenecks for scalable scientific reason- ing. Together, these works position reasoning as both an individual capability and an emergent property of socio-technical systems. Across these domains, reasoning is tightly coupled to trust through au- ditability and verification [56, 126, 150, 93, 159, 173, 163, 30, 196, 149, 92]. Multi-agent deliberation surfaces alternatives and justifications (e.g., Gemini Co- scientist), while procedural pipelines enable reproducibility and inspection (e.g., SciSciGPT). Complementary evaluation frameworks, such as those by Strachan et al. [163], test whether claimed reasoning abilities generalize across conditions. More broadly, Xie et al. [196] argues that trustworthy reasoning requires explicit verification loops, and Tsvetkova et al. [173] situates trust within ecosystem- level dynamics. Reasoning in current systems spans deliberative (multi-agent critique), procedural (tool-grounded workflows), and evaluative (verification and benchmarking) paradigms, which together form the basis for reliable and inter- pretable decision-making. Motivation. Motivation remains underdeveloped in World Models and co- pilot systems, with most approaches relying on externally specified objectives rather than intrinsic drives. Existing work primarily operationalizes motiva- tion through rewards, goals, or intent inference [49, 159, 173, 196, 149]. A key distinction emerges between optimization-driven and inference-driven motivation. In RL settings, motivation is encoded implicitly as an optimization signal via reward shaping and penalties, as in assistive human–robot interaction systems that incentivize efficient and valid behavior [159]. In contrast, embodied assistants increasingly model motivation as the inference of user goals or intent, enabling proactive assistance without explicit commands [49]. This reframes mo- tivation as alignment to latent human objectives rather than adherence to pre- defined reward functions. As discussed in Sec. 3.3, in practice, motivation is predominantly extrin- sic and often safety- or task-driven, particularly in medical and assistive do- mains [159, 49]. Systems are designed to optimize externally defined criteria such as correctness, efficiency, or user satisfaction, with little evidence of in- trinsic or self-directed objective formation. At a broader scale, motivation is shaped by the surrounding socio-technical system. Incentives, credit assignment, Human Cognition in Machines: A Unified Perspective of World Models41 and selection pressures govern agent behavior in collaborative and scientific set- tings [173, 196, 149]. Tsvetkova et al. [173] demonstrate how incentive structures drive emergent phenomena such as cascades and manipulation, while OmniScien- tist [149] embeds motivation in mechanisms such as attribution and peer review to regulate collaboration quality. Xie et al. [196] further argues that automa- tion reshapes which problems are pursued, effectively altering the motivational landscape of the research ecosystem. Motivation in current World Models is not intrinsic but arises from externally imposed objectives and incentive structures, highlighting a key gap between machine systems and human-like cognition. Meta-cognition. Meta-cognition is increasingly introduced as a reliability scaf- fold in World Models, encompassing self-critique, evaluation, uncertainty man- agement, and error checking [126, 150, 56, 149, 196, 27, 128, 173, 30]. A common design pattern separates generation from judgment, forming explicit propose– evaluate loops. For example, OpenAI Prism [126] introduces an always-on re- viewer layer for proofreading and consistency checking, while Gemini Co-scientist [56] and OmniScientist [149] operationalize critique through dedicated reflection roles and community-style ranking mechanisms. SciSciGPT [150] further emphasizes reproducibility and output validation within empirical pipelines. Meta-cognitive mechanisms operate at two levels. At the individual level, systems employ internal reflection to refine outputs or update beliefs, as seen in Generative Agents [128], where reflection transforms experience into higher- level behavioral guidance. At the system or community level, meta-cognition is externalized through evaluation platforms, auditability, and feedback loops, as in SciGPT and OmniScientist, enabling reproducibility and collective validation. This distinction highlights a key design axis: internal self-monitoring versus ex- ternal governance. These safeguards are particularly critical in high-stakes settings, where unchecked generation can propagate errors or overconfident hypotheses. For instance, Gem- ini Co-scientist [56] uses critique and reflection loops to down-select hypotheses prior to wet-lab validation, reducing experimental cost and risk. More broadly, Chen et al. [30] and Xie et al. [196] argue that without explicit self-evaluation and external oversight, iterative self-improvement can amplify errors. Overall, effective meta-cognition in World Models emerges from coupling strong genera- tive capabilities with explicit, testable judgment layers, often combining internal reflection with external evaluation and governance. 7 Conclusion Our review of video World Models, embodied World Models, and newly named epistemic World Models, we identify trends in the literature for overcoming com- mon problems with common solutions. We are the first to propose a taxonomy of recent World Models works rooted in cognitive architecture theory. Utiliz- ing our taxonomy made clear a research gap that may not have arisen other- wise regarding the cognitive functions of motivation (especially intrinsic moti- 42Authors Suppressed Due to Excessive Length vation), and meta-cognition. Additionally, our taxonomy inspires us to rethink agent frameworks for scientific discovery as Epistemic World Models with strong meta-cognitive capabilities that utilize a language-based Global Workspace-like interface for human-in-the-loop collaboration for scientific discovery. Our Unified World Models holistically incorporate all of the component parts of cognition: memory, perception, language, reasoning, imagining, motivation, and meta-cognition. Our proposed framework may not exist as currently de- scribed, but none of our proposals need come at the cost of another. Unified World Models encourage researchers to 1) use multi-modal inputs, 2) encode robust representations of their world within latent models that enable hypo- thetical reasoning during training and inference, 3) include tokenized language as an input, intermediate reasoning space, or output to facilitate human-in-the- loop cooperation, 4) utilize sim-to-real training in domains with little accessible training data, 5) reason with domain-specific state-of-the-art architectures, 6) provide reward signals to World Models using state-based rewards that make salient robust measurements like active inference, and finally 7) utilize a global workspace with self-evaluation and human-in-the-loop collaboration. Our report provides researchers a vocabulary to debate the merits of equat- ing machine and human cognition. Our review shows that to begin to fully equate machine and human cognition in function, let alone in any philosophical level, requires a World Model that holistically emulates all component parts of cognition, including the under-researched cognitive functions of motivation, and meta-cognition. References 1. Being-h0.7: A latent world-action model from egocentric videos. preprint (2026), https://research.beingbeyond.com/projects/being-h07/being-h07.pdf 2. Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al.: Cosmos world foundation model platform for physical ai. arXiv preprint arXiv:2501.03575 (2025) 3. Ali, A., Bai, J., Bala, M., Balaji, Y., Blakeman, A., Cai, T., Cao, J., Cao, T., Cha, E., Chao, Y.W., et al.: World simulation with video foundation models for physical ai. arXiv preprint arXiv:2511.00062 (2025) 4. Alonso, E., Jelley, A., Micheli, V., Kanervisto, A., Storkey, A., Pearce, T., Fleuret, F.: Diffusion for world modeling: Visual details matter in atari. Advances in Neural Information Processing Systems 37, 58757–58791 (2024) 5. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., Mané, D.: Concrete problems in ai safety. arXiv preprint arXiv:1606.06565 (2016) 6. Anderson, J.R., Matessa, M., Lebiere, C.: Act-r: A theory of higher level cognition and its relation to visual attention. Human–Computer Interaction 12(4), 439–462 (1997) 7. Arshid, K., Krayani, A., Marcenaro, L., Gomez, D.M., Regazzoni, C.: Toward au- tonomous uav swarm navigation: a review of trajectory design paradigms. Sensors (Basel, Switzerland) 25(18), 5877 (2025) 8. Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., Ballas, N.: Self-supervised learning from images with a joint-embedding pre- Human Cognition in Machines: A Unified Perspective of World Models43 dictive architecture. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 15619–15629 (2023) 9. Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., Rizvi, A., Roberts, C., Sinha, K., Zholus, A., et al.: V-jepa 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 (2025) 10. Baars, B.J.: A cognitive theory of consciousness. Cambridge University Press (1993) 11. Bagchi, A., Bao, Z., Bharadhwaj, H., Wang, Y.X., Tokmakov, P., Hebert, M.: Walk through paintings: Egocentric world models from internet priors. arXiv preprint arXiv:2601.15284 (2026) 12. Barcellona, L., Zadaianchuk, A., Allegro, D., Papa, S., Ghidoni, S., Gavves, E.: Dream to manipulate: Compositional world models empowering robot imitation learning with imagination. arXiv preprint arXiv:2412.14957 (2024) 13. Bardes, A., Garrido, Q., Ponce, J., Chen, X., Rabbat, M., LeCun, Y., Assran, M., Ballas, N.: Revisiting feature prediction for learning visual representations from video. arXiv preprint arXiv:2404.08471 (2024) 14. Bardes, A., Garrido, Q., Ponce, J., Chen, X., Rabbat, M., LeCun, Y., Assran, M., Ballas, N.: V-jepa: Latent video prediction for visual representation learning (2023) 15. Berseth, G., Geng, D., Devin, C., Rhinehart, N., Finn, C., Jayaraman, D., Levine, S.: Smirl: Surprise minimizing reinforcement learning in unstable environments. arXiv preprint arXiv:1912.05510 (2019) 16. Bheemaiah, A., Yang, S.: Knowledge graphs as world models for seman- tic material-aware obstacle handling in autonomous vehicles. arXiv preprint arXiv:2503.21232 (2025) 17. Bi, H., Tan, H., Xie, S., Wang, Z., Huang, S., Liu, H., Zhao, R., Feng, Y., Xiang, C., Rong, Y., et al.: Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030 (2025) 18. Bilal, A., Mohsin, M.A., Umer, M., Bangash, M.A.K., Jamshed, M.A.: Meta- thinking in llms via multi-agent reinforcement learning: A survey. arXiv preprint arXiv:2504.14520 (2025) 19. Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., Levine, S.: Zero-shot robotic manipulation with pretrained image-editing diffusion models. arXiv preprint arXiv:2310.10639 (2023) 20. Boggs, J.: Towards visual-symbolic integration in the soar cognitive architecture. Cognitive Systems Research 91, 101353 (2025) 21. Bosio, C., Woelki, G., Hendy, N., Roy, N., Kim, B.: Rdar: Reward-driven agent relevance estimation for autonomous driving. arXiv preprint arXiv:2509.19789 (2025) 22. Bühler, K.: Sprachtheorie, vol. 2. Jena Fischer (1934) 23. Burchi, M., Timofte, R.: Accurate and efficient world modeling with masked latent transformers. arXiv preprint arXiv:2507.04075 (2025) 24. Cao, M., Tang, H., Zhao, H., Han, M., Liu, R., Sun, Q., Chang, X., Reid, I., Liang, X.: Order from chaos: Physical world understanding from glitchy gameplay videos. arXiv preprint arXiv:2601.16471 (2026) 25. Cen, J., Yu, C., Yuan, H., Jiang, Y., Huang, S., Guo, J., Li, X., Song, Y., Luo, H., Wang, F., et al.: Worldvla: Towards autoregressive action world model. arXiv preprint arXiv:2506.21539 (2025) 44Authors Suppressed Due to Excessive Length 26. Chang, J., Uehara, M., Sreenivas, D., Kidambi, R., Sun, W.: Mitigating covariate shift in imitation learning via offline data with partial coverage. Advances in Neural Information Processing Systems 34, 965–979 (2021) 27. Chatlatanagulchai, W., Thonglek, K., Reid, B., Kashiwa, Y., Leelaprute, P., Rungsawang, A., Manaskasemsak, B., Iida, H.: On the use of agentic coding manifests: An empirical study of claude code. In: International Conference on Product-Focused Software Process Improvement. p. 543–551. Springer (2025) 28. Cheang, C.L., Chen, G., Jing, Y., Kong, T., Li, H., Li, Y., Liu, Y., Wu, H., Xu, J., Yang, Y., et al.: Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158 (2024) 29. Chen, B., Zhang, T., Geng, H., Song, K., Zhang, C., Li, P., Freeman, W.T., Malik, J., Abbeel, P., Tedrake, R., et al.: Large video planner enables generalizable robot control. arXiv preprint arXiv:2512.15840 (2025) 30. Chen, S., Ma, S., Yu, S., Zhang, H., Zhao, S., Lu, C.: Exploring consciousness in llms: A systematic survey of theories, implementations, and frontier risks (2025), https://arxiv.org/abs/2505.19806 31. Chen, Y., et al.: Drivinggpt: Unifying driving world modeling and planning with multi-modal autoregressive transformers. arXiv preprint arXiv:2412.18607 (2024) 32. Colombatto, C., Fleming, S.M.: Folk psychological attributions of consciousness to large language models. Neuroscience of Consciousness 2024(1), niae013 (2024) 33. Cornelissen, C., Leroux, S., Simoens, P.: Le mumo jepa: Multi-modal self- supervised representation learning with learnable fusion tokens. arXiv preprint arXiv:2603.24327 (2026) 34. Dang, C., et al.: Sparseworld: A flexible, adaptive, and efficient 4d occu- pancy world model powered by sparse and dynamic queries. arXiv preprint arXiv:2510.17482 (2025) 35. Darwin, C.: The descent of man, and selection in relation to sex, vol. 2. D. Ap- pleton (1872) 36. Deacon, T.W.: The symbolic species: The co-evolution of language and the brain. W Norton & Company (1998) 37. Dentella, V., Günther, F., Murphy, E., Marcus, G., Leivada, E.: Testing ai on lan- guage comprehension tasks reveals insensitivity to underlying meaning. Scientific Reports 14(1), 28083 (2024) 38. Destrade, M., Bounou, O., Lidec, Q.L., Ponce, J., LeCun, Y.: Value-guided action planning with jepa world models. arXiv preprint arXiv:2601.00844 (2025) 39. Dharmarajan, K., Huang, W., Wu, J., Fei-Fei, L., Zhang, R.: Dream2flow: Bridg- ing video generation and open-world manipulation with 3d object flow. arXiv preprint arXiv:2512.24766 (2025) 40. Ding, J., Zhang, Y., Shang, Y., Zhang, Y., Zong, Z., Feng, J., Yuan, Y., Su, H., Li, N., Sukiennik, N., et al.: Understanding world or predicting future? a comprehensive survey of world models. ACM Computing Surveys 58(3), 1–38 (2025) 41. Doerig, A., Kietzmann, T.C., Allen, E., Wu, Y., Naselaris, T., Kay, K., Charest, I.: High-level visual representations in the human brain are aligned with large language models. Nature Machine Intelligence p. 1–15 (2025) 42. Donald, M.: Origins of the modern mind: Three stages in the evolution of culture and cognition. Harvard university press (1993) 43. Dong, J., Lyu, Q., Liu, B., Wang, X., Liang, W., Zhang, D., Tu, J., Li, H., Zhao, H., Ding, H., et al.: Learning to model the world: A survey of world models in artificial intelligence. Authorea Preprints (2026) Human Cognition in Machines: A Unified Perspective of World Models45 44. Du, Y., Yang, S., Dai, B., Dai, H., Nachum, O., Tenenbaum, J., Schuurmans, D., Abbeel, P.: Learning universal policies via text-guided video generation. Advances in neural information processing systems 36, 9156–9172 (2023) 45. Dung Nguyen, V., Yang, Z., Buckley, C.L., Ororbia, A.: R-aif: Solving sparse- reward robotic tasks from pixels with active inference and world models. arXiv e-prints p. arXiv–2409 (2024) 46. Durante, Z., Singh, S., Khatua, A., Agarwal, S., Tan, R., Lee, Y.J., Gao, J., Adeli, E., Fei-Fei, L.: Videoweave: A data-centric approach for efficient video understanding. arXiv preprint arXiv:2601.06309 (2026) 47. Friston, K., Da Costa, L., Tschantz, A., Heins, C., Buckley, C., Verbelen, T., Parr, T.: Active inference and artificial reasoning. arXiv preprint arXiv:2512.21129 (2025) 48. Friston, K., FitzGerald, T., Rigoli, F., Schwartenbeck, P., Pezzulo, G., et al.: Active inference and learning. Neuroscience & Biobehavioral Reviews 68, 862– 879 (2016) 49. Fung, P., Bachrach, Y., Celikyilmaz, A., Chaudhuri, K., Chen, D., Chung, W., Dupoux, E., Gong, H., et al.: Embodied ai agents: Modeling the world (2025) 50. Gao, S., Zhou, S., Du, Y., Zhang, J., Gan, C.: Adaworld: Learning adaptable world models with latent actions. arXiv preprint arXiv:2503.18938 (2025) 51. Gao, Y., Zhang, Q., Ding, D.W., Zhao, D.: Dream to drive with predictive indi- vidual world model. IEEE Transactions on Intelligent Vehicles (2024) 52. Garrido, Q., Nagarajan, T., Terver, B., Ballas, N., LeCun, Y., Rabbat, M.: Learn- ing latent action world models in the wild. arXiv preprint arXiv:2601.05230 (2026) 53. Ge, Z., Huang, H., Zhou, M., Li, J., Wang, G., Tang, S., Zhuang, Y.: Worldgpt: Empowering llm as multimodal world model. In: Proceedings of the 32nd ACM International Conference on Multimedia. p. 7346–7355 (2024) 54. Gentile, G., Morello, G., La Cognata, V., Guarnaccia, M., Cavallaro, S.: Artificial intelligence in transcriptomics: From human-in-the-loop to agentic ai. Journal of Personalized Medicine 16(4), 181 (2026) 55. Goff, M., Hogan, G., Hotz, G., du Parc Locmaria, A., Raczy, K., Schäfer, H., Shihadeh, A., Zhang, W., Yousfi, Y.: Learning to drive from a world model. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 1964–1973 (2025) 56. Gottweis, J., Weng, W.H., Daryin, A., Tu, T., Palepu, A., Sirkovic, P., Myaskovsky, A., Weissenberger, F., Rong, K., Tanno, R., et al.: Towards an ai co-scientist. arXiv preprint arXiv:2502.18864 (2025) 57. Gu, Y., Zhang, K., Ning, Y., Zheng, B., Gou, B., Xue, T., Chang, C., Srivastava, S., Xie, Y., Qi, P., et al.: Is your llm secretly a world model of the internet? model-based planning for web agents. arXiv preprint arXiv:2411.06559 (2024) 58. Gumbsch, C., Sajid, N., Martius, G., Butz, M.V.: In: The Twelfth International Conference on Learning Representations (2023) 59. Guo, J., Ma, X., Wang, Y., Yang, M., Liu, H., Li, Q.: Flowdreamer: A rgb- d world model with flow-based motion representations for robot manipulation. IEEE Robotics and Automation Letters 11(3), 2466–2473 (2026) 60. Ha, D., Schmidhuber, J.: World models. arXiv preprint arXiv:1803.10122 2(3) (2018) 61. HaCohen, Y., Chiprut, N., Brazowski, B., Shalem, D., Moshe, D., Richardson, E., Levin, E., Shiran, G., Zabari, N., Gordon, O., et al.: Ltx-video: Realtime video latent diffusion. arXiv preprint arXiv:2501.00103 (2024) 62. Hadfield-Menell, D., Dragan, A.D., Abbeel, P., Russell, S.: The off-switch game. In: AAAI Workshops (2017) 46Authors Suppressed Due to Excessive Length 63. Hansen, N., Su, H., Wang, X.: Td-mpc2: Scalable, robust world models for con- tinuous control. arXiv preprint arXiv:2310.16828 (2023) 64. Hansen, N., SV, J., Sobal, V., LeCun, Y., Wang, X., Su, H.: Hierarchi- cal world models as visual whole-body humanoid controllers. arXiv preprint arXiv:2405.18418 (2024) 65. Hao, C., Lu, W., Xu, Y., Chen, Y.: Neural motion simulator pushing the limit of world models in reinforcement learning. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 27608–27617 (2025) 66. Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., Corrado, G.: Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080 (2023) 67. Hu, A., et al.: Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080 (2023) 68. Huang, S., Chen, L., Zhou, P., Chen, S., Jiang, Z., Hu, Y., Liao, Y., Gao, P., Li, H., Yao, M., et al.: Enerverse: Envisioning embodied future space for robotics manipulation. arXiv preprint arXiv:2501.01895 (2025) 69. Huang, W., Ji, J., Xia, C., Zhang, B., Yang, Y.: Safedreamer: Safe reinforcement learning with world models. arXiv preprint arXiv:2307.07176 (2023) 70. Huang, W., Chao, Y.W., Mousavian, A., Liu, M.Y., Fox, D., Mo, K., Fei-Fei, L.: Pointworld: Scaling 3d world models for in-the-wild robotic manipulation. arXiv preprint arXiv:2601.03782 (2026) 71. Huang, X., Li, Z., He, G., Zhou, M., Shechtman, E.: Self forcing: Bridging the train-test gap in autoregressive video diffusion (2025), https://arxiv.org/abs/ 2506.08009 72. Huang, Y., Zhang, J., Zou, S., Liu, X., Hu, R., Xu, K.: Ladi-wm: A latent diffusion- based world model for predictive manipulation. arXiv preprint arXiv:2505.11528 (2025) 73. Huh, M., Cheung, B., Wang, T., Isola, P.: The platonic representation hypothesis. arXiv preprint arXiv:2405.07987 (2024) 74. Ibrahim, L., Cheng, M.: Thinking beyond the anthropomorphic paradigm benefits llm research. arXiv preprint arXiv:2502.09192 (2025) 75. Intelligence, P., Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., Equi, M., Finn, C., Fusai, N., Galliker, M.Y., et al.:π 0.5 : a vision- language-action model with open-world generalization (2025) 76. Jang, J., Yoo, M., Yoon, S., Woo, H.: Test-time mixture of world models for em- bodied agents in dynamic environments. arXiv preprint arXiv:2601.22647 (2026) 77. Jang, J., Ye, S., Lin, Z., Xiang, J., Bjorck, J., Fang, Y., Hu, F., Huang, S., Kun- dalia, K., Lin, Y.C., et al.: Dreamgen: Unlocking generalization in robot learning through video world models. arXiv preprint arXiv:2505.12705 (2025) 78. Jang, S., Ki, T., Jo, J., Xie, S., Yoon, J., Hwang, S.J.: Self-refining video sampling. arXiv preprint arXiv:2601.18577 (2026) 79. Jaynes, J.: from the origin of consciousness in the breakdown of the bicameral mind. In: Creative Writing, p. 541–543. Routledge (2013) 80. Jiang, A., Gao, Y., Wang, Y., Sun, Z., Wang, S., Heng, Y., Sun, H., Tang, S., Zhu, L., Chai, J., et al.: Irl-vla: Training an vision-language-action policy via reward world model. arXiv preprint arXiv:2508.06571 (2025) 81. Johnson, S.G., Karimi, A.H., Bengio, Y., Chater, N., Gerstenberg, T., Larson, K., Levine, S., Mitchell, M., Rahwan, I., Schölkopf, B., et al.: Imagining and building wise machines: The centrality of ai metacognition. Trends in Cognitive Sciences (2025) Human Cognition in Machines: A Unified Perspective of World Models47 82. Kahn, G., Abbeel, P., Levine, S.: Badgr: An autonomous self-supervised learning- based navigation system. IEEE Robotics and Automation Letters 6(2), 1312–1319 (2021) 83. Kang, B., Kim, J., Yun, T., Bae, H., Kim, C.E.: Identifying features that shape perceived consciousness in llm-based ai: A quantitative study of human responses. Computers in Human Behavior Reports p. 100901 (2025) 84. Kashirskiy, M., Makarov, I.: Sus: Strategy-aware surprise for intrinsic exploration. arXiv preprint arXiv:2601.10349 (2026) 85. Kawaguchi, K., Deng, Z., Ji, X., Huang, J.: How does information bottleneck help deep learning? In: International conference on machine learning. p. 16049–16096. PMLR (2023) 86. Khanzada, F.K., Kwon, J.: Indrive: Intrinsic disagreement based reinforcement for vehicle exploration through curiosity driven generalized world model. arXiv preprint arXiv:2503.05573 (2025) 87. Kim, M.J., Gao, Y., Lin, T.Y., Lin, Y.C., Ge, Y., Lam, G., Liang, P., Song, S., Liu, M.Y., Finn, C., et al.: Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163 (2026) 88. Kim, S.H., Park, D., Lee, J., Lee, S.Y., Chong, Y.: Humanoid artificial conscious- ness designed with large language model based on psychoanalysis and personality theory. Cognitive Systems Research p. 101392 (2025) 89. Kim, T., Jo, Y., Song, H., Kim, T.: Towards human-like multimodal conversa- tional agent by generating engaging speech. In: Interspeech 2025. p. 4828–4832. interspeech_2025, ISCA (Aug 2025) 90. Klyubin, A.S., Polani, D., Nehaniv, C.L.: Empowerment: A universal agent-centric measure of control. In: 2005 ieee congress on evolutionary computation. vol. 1, p. 128–135. IEEE (2005) 91. Koriat, A., et al.: Metacognition and consciousness. Institute of Information Pro- cessing and Decision Making, University of Haifa . . . (2006) 92. Kumar, V., Tanusri, M.: Unified speech-to-speech models for real-time, multilin- gual, and emotionally aware ai. Journal of Current Trends in Computer Science Research 4(6), 01–06 (2025). https://doi.org/10.33140/JCTCSR.04.06.02 93. Kwok, K., Fernando, B., Xu, Q., Subbaraju, V., Choi, D., Quek, B.K.: Explicit world models for reliable human-robot collaboration (2026), https://arxiv.org/ abs/2601.01705 94. Laird, J.E.: The Soar cognitive architecture. MIT press (2019) 95. Langley, P., McKusick, K.B., Allen, J.A., Iba, W.F., Thompson, K.: A design for the icarus architecture. ACM Sigart Bulletin 2(4), 104–109 (1991) 96. Le, M.Q., Zhu, Y., Kalogeiton, V., Samaras, D.: What about gravity in video generation? post-training newton’s laws with verifiable rewards. arXiv preprint arXiv:2512.00425 (2025) 97. Leviathan, Y., Kalman, M., Matias, Y.: Prompt repetition improves non- reasoning llms. arXiv preprint arXiv:2512.14982 (2025) 98. Li, F.f.: From words to worlds: Spatial intelligence is ai’snextfrontier(2025), https://drfeifei.substack.com/p/ from-words-to-worlds-spatial-intelligence, published on Dr. Li’s SubStack 99. Li, L., Zhang, Q., Luo, Y., Yang, S., Wang, R., Han, F., Yu, M., Gao, Z., Xue, N., Zhu, X., et al.: Causal world modeling for robot control. arXiv preprint arXiv:2601.21998 (2026) 100. Li, Q., Jia, X., Wang, S., Yan, J.: Think2drive: Efficient reinforcement learning by thinking with latent world model for autonomous driving (in carla-v2). In: European conference on computer vision. p. 142–158. Springer (2024) 48Authors Suppressed Due to Excessive Length 101. Li, X., He, X., Zhang, L., Wu, M., Li, X., Liu, Y.: A comprehensive survey on world models for embodied ai. arXiv preprint arXiv:2510.16732 (2025) 102. Li, Y., Fan, L., He, J., Wang, Y., Chen, Y., Zhang, Z., Tan, T.: Enhancing end-to- end autonomous driving with latent world model. arXiv preprint arXiv:2406.08481 (2024) 103. Li, Y., Lu, L., Kong, Z., et al.: Diff-styGS: 3d gaussian splatting stylization via tuning-free multi-view sparse diffusion (2025), https://openreview.net/forum? id=cp5l65lzuG 104. Li, Z., Han, X., Li, Y., Strauss, N., Schubert, M.: Dawm: Diffusion action world models for offline reinforcement learning via action-inferred transitions. arXiv preprint arXiv:2509.19538 (2025) 105. Liang, A., et al.: Lidarcrafter: Dynamic 4d world modeling from lidar sequences. arXiv preprint arXiv:2508.03692 (2025) 106. Liang, A., et al.: Worldlens: Full-spectrum evaluations of driving world models in real world. arXiv preprint arXiv:2512.10958 (2025) 107. Lieberman, P.: Toward an evolutionary biology of language. Harvard University Press (2006) 108. Lin, J., Taherin, A., Akbari, A., Akbari, A., Lu, L., Chen, G., Padir, T., Yang, X., Chen, W., Li, Y., et al.: Vote: vision-language-action optimization with trajectory ensemble voting. arXiv preprint arXiv:2507.05116 (2025) 109. Lin, M., Wang, X., Wang, Y., Wang, S., Dai, F., Ding, P., Wang, C., Zuo, Z., Sang, N., Huang, S., et al.: Exploring the evolution of physics cognition in video generation: A survey. arXiv preprint arXiv:2503.21765 (2025) 110. Liu, D., Zhang, J., Dinh, A.D., Park, E., Zhang, S., Mian, A., Shah, M., Xu, C.: Generative physical ai in vision: A survey. arXiv preprint arXiv:2501.10928 (2025) 111. Liu, J., Kong, Z., Zhao, P., Yang, C., et al.: Toward adaptive large language models structured pruning via hybrid-grained weight importance assessment. Proceedings of the AAAI Conference on Artificial Intelligence 39(18), 18879–18887 (Apr 2025) 112. Liu, Y., Ding, P., Jiang, T., Wang, X., Song, W., Lin, M., Zhao, H., Zhang, H., Zhuang, Z., Zhao, W., et al.: Mmada-vla: Large diffusion vision-language- action model with unified multi-modal instruction and generation. arXiv preprint arXiv:2603.25406 (2026) 113. Long, X., Zhao, Q., Zhang, K., Zhang, Z., Wang, D., Liu, Y., Shu, Z., Lu, Y., Wang, S., Wei, X., et al.: A survey: Learning embodied intelligence from physical simulators and world models. arXiv preprint arXiv:2507.00917 (2025) 114. Lu, G., Zhang, S., Wang, Z., Liu, C., Lu, J., Tang, Y.: Manigaussian: Dynamic gaussian splatting for multi-task robotic manipulation. In: European Conference on Computer Vision. p. 349–366. Springer (2024) 115. Lu, Y., Zeng, Y., Li, H., Ouyang, H., Wang, Q., Cheng, K.L., Zhu, J., Cao, H., Zhang, Z., Zhu, X., Shen, Y., Zhang, M.: Reward forcing: Efficient streaming video generation with rewarded distribution matching distillation (2025), https: //arxiv.org/abs/2512.04678 116. Luo, H., Wang, Y., Zhang, W., Zheng, S., Xi, Z., Xu, C., Xu, H., Yuan, H., Zhang, C., Wang, Y., et al.: Being-h0. 5: Scaling human-centric robot learning for cross-embodiment generalization. arXiv preprint arXiv:2601.12993 (2026) 117. Mak, C.W., Zhu, G., Zhang, B., Li, H., Chi, X., Zhang, K., Wu, Y., He, Y., Fan, C.K., Lu, W., et al.: Physicsmind: Sim and real mechanics benchmarking for physical reasoning and prediction in foundational vlms and world models. arXiv preprint arXiv:2601.16007 (2026) Human Cognition in Machines: A Unified Perspective of World Models49 118. Maytié, L., Johannet, R.B., VanRullen, R.: Multimodal dreaming: A global workspace approach to world model-based reinforcement learning. arXiv preprint arXiv:2502.21142 (2025) 119. Mazzaglia, P., Verbelen, T., Dhoedt, B., Courville, A., Rajeswar, S.: Genrl: Multimodal-foundation world models for generalization in embodied agents. Ad- vances in neural information processing systems 37, 27529–27555 (2024) 120. Meyer, D.E., Glass, J.M., Mueller, S.T., Seymour, T.L., Kieras, D.E.: Executive- process interactive control: A unified computational theory for answering 20 ques- tions (and more) about cognitive ageing. European Journal of Cognitive Psychol- ogy 13(1-2), 123–164 (2001) 121. Nachkov, A., Pani Paudel, D., Zaech, J.N., Scaramuzza, D., Van Gool, L.: Dream to drive: Model-based vehicle control using analytic world models. arXiv e-prints p. arXiv–2502 (2025) 122. Nahrendra, I.M.A., Yu, B., Myung, H.: Dreamwaq: Learning robust quadrupedal locomotion with implicit terrain imagination via deep reinforcement learning. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). p. 5078–5084. IEEE (2023) 123. Newell, A.: Unified theories of cognition. Harvard University Press (1994) 124. Nouri, N., Artzi, R., Savova, V.: An agentic ai framework for ingestion and stan- dardization of single-cell rna-seq data analysis. npj Artificial Intelligence 2(1), 8 (2026) 125. OpenAI: Video generation models as world simulators. https://openai.com/ research/video-generation-models-as-world-simulators (2024) 126. OpenAI: Introducing prism: A free latex-native workspace for scientific writing and collaboration. https://openai.com/index/introducing-prism/ (2026), ac- cessed: 2026-02-16 127. Pagan, N., Törnberg, P., Bail, C.A., Hannák, A., Barrie, C.: Computational turing test reveals systematic differences between human and ai language. arXiv preprint arXiv:2511.04195 (2025) 128. Park, J.S., O’Brien, J.C., Cai, C.J., Morris, M.R., Liang, P., Bernstein, M.S.: Generative agents: Interactive simulacra of human behavior (2023), https:// arxiv.org/abs/2304.03442 129. Pate, N., Proma, A.M., He, H., Druckman, J.N., Molden, D., Ghoshal, G., Hoque, E.: Replicating human motivated reasoning studies with llms. arXiv preprint arXiv:2601.16130 (2026) 130. Pathak, D., Agrawal, P., Efros, A.A., Darrell, T.: Curiosity-driven exploration by self-supervised prediction. In: International conference on machine learning. p. 2778–2787. PMLR (2017) 131. Pazem, J., Krumm, M., Vining, A.Q., Fiderer, L.J., Briegel, H.J.: Free energy pro- jective simulation (feps): Active inference with interpretability. PLoS One 20(9), e0331047 (2025) 132. Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 4195–4205 (2023) 133. Pezzulo, G., Parr, T., Friston, K.: Active inference as a theory of sentient behavior. Biological Psychology 186, 108741 (2024) 134. Pinker, S., Bloom, P.: Natural language and natural selection. Behavioral and brain sciences 13(4), 707–727 (1990) 135. Porębski, A., Figura, J.: There is no such thing as conscious artificial intelligence. Humanities and Social Sciences Communications 12(1), 1–12 (2025) 50Authors Suppressed Due to Excessive Length 136. Puyin, L., Xiang, T., Mao, E., Wei, S., Chen, X., Masood, A., Fei-Fei, L., Adeli, E.: Quantiphy: A quantitative benchmark evaluating physical reasoning abilities of vision-language models. arXiv preprint arXiv:2512.19526 (2025) 137. Qu, Y., Huang, Z., Sheng, Z., Chen, J., Chen, S., Labi, S.: Vl-safe: Vision-language guided safety-aware reinforcement learning with world models for autonomous driving. arXiv preprint arXiv:2505.16377 (2025) 138. Ren, X., et al.: Cosmos-drive-dreams: Scalable synthetic driving data generation with world foundation models. arXiv preprint arXiv:2506.09042 (2025) 139. Ritter, F.E., Tehranchi, F., Oury, J.D.: Act-r: A cognitive architecture for mod- eling cognition. Wiley Interdisciplinary Reviews: Cognitive Science 10(3), e1488 (2019) 140. Rome, J., James, S., Ramamoorthy, S.: Learning to unfold cloth: Scaling up world models to deformable object manipulation. arXiv preprint arXiv:2602.16675 (2026) 141. Rosenthal, D.M.: Consciousness, content, and metacognitive judgments. Con- sciousness and cognition 9(2), 203–214 (2000) 142. Rupprecht, T., Wang, Y.: A survey for deep reinforcement learning in markovian cyber–physical systems: Common problems and solutions. Neural Networks 153, 13–36 (2022) 143. Russell, L., et al.: Gaia-2: A controllable multi-view generative world model for autonomous driving. arXiv preprint arXiv:2503.20523 (2025) 144. Sarikaya, R.: Path to artificial general intelligence: Past, present, and future. Annual Reviews in Control 60, 101021 (2025) 145. Sekar, R., Rybkin, O., Daniilidis, K., Abbeel, P., Hafner, D., Pathak, D.: Plan- ning to explore via self-supervised world models. In: International conference on machine learning. p. 8583–8592. PMLR (2020) 146. Seo, Y., Hafner, D., Liu, H., Liu, F., James, S., Lee, K., Abbeel, P.: Masked world models for visual control. In: Conference on Robot Learning. p. 1332– 1344. PMLR (2023) 147. Shah, D., Sridhar, A., Dashora, N., Stachowicz, K., Black, K., Hirose, N., Levine, S.: Vint: A foundation model for visual navigation. arXiv preprint arXiv:2306.14846 (2023) 148. Shang, Y., Zhang, X., Tang, Y., Jin, L., Gao, C., Wu, W., Li, Y.: Roboscape: Physics-informed embodied world model. arXiv preprint arXiv:2506.23135 (2025) 149. Shao, C., Huang, D., Li, Y., Zhao, K., Lin, W., Zhang, Y., Zeng, Q., et al.: Omniscientist: Toward a co-evolving ecosystem of human and ai scientists (2025), https://arxiv.org/abs/2511.16931 150. Shao, E., Wang, Y., Qian, Y., Pan, Z., Liu, H., Wang, D.: Sciscigpt: advancing human–ai collaboration in the science of science. Nature Computational Science p. 1–15 (2025) 151. Shen, X., Han, C., Zhou, Y., Xie, Y., Gong, Y., Wang, Q., Wang, Y., Wang, Y., Zhao, P., Gu, J.: Draftattention: Fast video diffusion via low-resolution attention guidance. arXiv preprint arXiv:2505.14708 (2025) 152. Shen, X., Ma, W., Liu, J., et al.: Quartdepth: Post-training quantization for real- time depth estimation on the edge. In: CVPR (2025) 153. Shen, X., Ma, W., Zhou, Y., Tang, E., Xie, Y., Li, Z., Gong, Y., Wang, Q., Ding, H., Wang, Y., et al.: Fastcar: Cache attentive replay for fast auto-regressive video generation on the edge. arXiv preprint arXiv:2505.14709 (2025) 154. Shen, X., Song, Z., Zhou, Y., et al.: Lazydit: Lazy learning for the acceleration of diffusion transformers. In: AAAI (2025) Human Cognition in Machines: A Unified Perspective of World Models51 155. Shen, X., Song, Z., Zhou, Y., et al.: Numerical pruning for efficient autoregressive models. In: AAAI (2025) 156. Shen, X., Wang, Y., Shi, X., Wang, Y., Zhao, P., Gu, J.: Efficient reasoning with hidden thinking. arXiv preprint arXiv:2501.19201 (2025) 157. Shen, X., Zhao, P., Gong, Y., Kong, Z., Zhan, Z., Wu, Y., Lin, M., Wu, C., Lin, X., Wang, Y.: Search for efficient large language models. In: NeurIPS (2024) 158. Shen, X., Zheng, H., Gong, Y., et al.: Sparse learning for state space models on mobile. In: ICLR (2025) 159. Shervedani, A.M., Li, S., Monaikul, N., Abbasi, B., Žefran, M., Eugenio, B.D.: Multimodal reinforcement learning for robots collaborating with humans. Inter- national Journal of Social Robotics 17(12), 3003–3025 (2025). https://doi.org/ 10.1007/s12369-025-01287-6 160. Silver, D., Singh, S., Precup, D., Sutton, R.S.: Reward is enough. Artificial intel- ligence 299, 103535 (2021) 161. Song, Z., Lu, J., Du, Y., Yu, B., Pruyn, T.M., Huang, Y., Guo, K., Luo, X., Qu, Y., Qu, Y., et al.: Evaluating large language models in scientific discovery. arXiv preprint arXiv:2512.15567 (2025) 162. Sridhar, A., Shah, D., Glossop, C., Levine, S.: Nomad: Goal masked diffusion policies for navigation and exploration. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). p. 63–70. IEEE (2024) 163. Strachan, J.W.A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., et al.: Testing theory of mind in large language models and humans. Na- ture Human Behaviour 8(7), 1285–1295 (2024). https://doi.org/10.1038/ s41562-024-01882-z 164. Sun, J., Zhang, W., Qi, Z., Ren, S., Liu, Z., Zhu, H., Sun, G., Jin, X., Chen, Z.: Vla-jepa: Enhancing vision-language-action model with latent world model. arXiv preprint arXiv:2602.10098 (2026) 165. Sun, Y.T., Huang, Z., Niu, Y., Ma, L., Cao, Y.P., Ma, Y., Qi, X.: Stereo world model: Camera-guided stereo video generation (2026), https://arxiv.org/abs/ 2603.17375 166. Tang, C., Liu, Y., Wu, Y., Han, W., Yin, Q., Zheng, X., Zeng, W., Zhang, Q.: Moe- world: A mixture-of-experts architecture for multi-task world models. Electronics 14(24), 4884 (2025) 167. Team, G., Ye, A., Wang, B., Ni, C., Huang, G., Zhao, G., Li, H., Li, J., Zhu, J., Feng, L., et al.: Gigabrain-0: A world model-powered vision-language-action model. arXiv preprint arXiv:2510.19430 (2025) 168. Team, K., Chen, J., Ci, Y., Du, X., Feng, Z., Gai, K., Guo, S., Han, F., He, J., He, K., et al.: Kling-omni technical report. arXiv preprint arXiv:2512.16776 (2025) 169. Team, M.W.: Marble: A multimodal world model (2025), https://w. worldlabs.ai/blog/marble-world-model, published on Marble World’s Blog 170. Team, R., Gao, Z., Wang, Q., Zeng, Y., Zhu, J., Cheng, K.L., Li, Y., Wang, H., Xu, Y., Ma, S., et al.: Advancing open-source world models. arXiv preprint arXiv:2601.20540 (2026) 171. Tomasello, M.: Two Hypotheses About Primate Cognition. MIT Press (2000) 172. Tomasello, M., Carpenter, M., Call, J., Behne, T., Moll, H.: Understanding and sharing intentions: The origins of cultural cognition. Behavioral and brain sciences 28(5), 675–691 (2005) 173. Tsvetkova, M., Yasseri, T., Pescetelli, N., Werner, T.: A new sociology of humans and machines. Nature Human Behaviour 8(10), 1864–1876 (Oct 2024). https://doi.org/10.1038/s41562-024-02001-8, http://dx.doi.org/ 10.1038/s41562-024-02001-8 52Authors Suppressed Due to Excessive Length 174. Vo, K.H., Nguyen, D.P., Nguyen, T.T., Quan, T.T.: Ti-jepa: An innovative energy- based joint embedding strategy for text-image multimodal systems. In: Interna- tional Symposium on Information and Communication Technology. p. 141–154. Springer (2024) 175. Wan, T., Wang, A., Ai, B., Wen, B., Mao, C., Xie, C.W., Chen, D., Yu, F., Zhao, H., Yang, J., et al.: Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314 (2025) 176. Wang, H., Ye, X., Tao, F., Pan, C., Mallik, A., Yaman, B., Ren, L., Zhang, J.: Adawm: Adaptive world model based planning for autonomous driving. arXiv preprint arXiv:2501.13072 (2025) 177. Wang, J., Liu, D., Chen, J., Da, J., Qian, N., Tram, M.M., Soh, H.: Genie: A generalizable navigation system for in-the-wild environments. IEEE Robotics and Automation Letters (2025) 178. Wang, K., Zhang, P., Wang, Z., Gao, Y., Li, L., Wang, Q., Chen, H., Wan, C., Lu, Y., Yang, Z., et al.: Vagen: Reinforcing world model reasoning for multi-turn vlm agents. arXiv preprint arXiv:2510.16907 (2025) 179. Wang, L., Zheng, Y., Chen, Q., Li, S., Zhang, Y., Xing, Z., Zhang, Q., Li, X., Qian, D., Yang, P., et al.: Latent-wam: Latent world action modeling for end-to- end autonomous driving. arXiv preprint arXiv:2603.24581 (2026) 180. Wang, L., Shelim, R., Saad, W., Ramakrishna, N.: Metamind: General and cogni- tive world models in multi-agent systems by meta-theory of mind. arXiv preprint arXiv:2603.00808 (2026) 181. Wang, L., Yang, Z., Bai, C., Zhang, G., Liu, X., Zheng, X., Long, X.X., Lu, C.T., Lu, C.: Drive-jepa: Video jepa meets multimodal trajectory distillation for end-to-end driving. arXiv preprint arXiv:2601.22032 (2026) 182. Wang, Q., Huang, W., Zhou, Y., Yin, H., Bao, T., Lyu, J., Liu, W., Zhang, R., Wu, J., Fei-Fei, L., Li, M.: Enact: Evaluating embodied cognition with world modeling of egocentric interaction (2025), https://arxiv.org/abs/2511.20937 183. Wang, R., Liu, Q., Deng, Y., Liu, G., Liu, Z., Jia, K.: Eva: Aligning video world models with executable robot actions via inverse dynamics rewards. arXiv preprint arXiv:2603.17808 (2026) 184. Wang, Y., Fang, Y., Wang, T., Feng, Y., Tan, Y., Zhang, S., Liu, P., Ji, Y., Xu, R.: Dreamnav: A trajectory-based imaginative framework for zero-shot vision- and-language navigation. arXiv preprint arXiv:2509.11197 (2025) 185. Wu, H., Jing, Y., Cheang, C., Chen, G., Xu, J., Li, X., Liu, M., Li, H., Kong, T.: Unleashing large-scale video generative pre-training for visual robot manipulation. arXiv preprint arXiv:2312.13139 (2023) 186. Wu, J., Yin, S., Feng, N., Long, M.: Rlvr-world: Training world models with reinforcement learning (2025). arXiv preprint arXiv:2505.13934 187. Wu, J., Yin, S., Feng, N., Long, M.: Rlvr-world: Training world models with reinforcement learning. arXiv preprint arXiv:2505.13934 (2025) 188. Wu, J., Zhang, X., Yuan, H., Zhang, X., Huang, T., He, C., Deng, C., Zhang, R., Wu, Y., Long, M.: Visual generation unlocks human-like reasoning through multimodal world models. arXiv preprint arXiv:2601.19834 (2026) 189. Wu, P., Escontrela, A., Hafner, D., Abbeel, P., Goldberg, K.: Daydreamer: World models for physical robot learning. In: Conference on robot learning. p. 2226– 2240. PMLR (2023) 190. Wu, W., Lu, F., Wang, Y., Yang, S., Liu, S., Wang, F., Zhu, Q., Sun, H., Wang, Y., Ma, S., et al.: A pragmatic vla foundation model. arXiv preprint arXiv:2601.18692 (2026) Human Cognition in Machines: A Unified Perspective of World Models53 191. Wu, W., Liu, F., Li, H., Hu, Z., Dong, D., Chen, C., Wang, Z.: Mixture-of-experts meets in-context reinforcement learning. arXiv preprint arXiv:2506.05426 (2025) 192. Xiang, C., Liu, J., Zhang, J., Yang, X., Fang, Z., Wang, S., Wang, Z., Zou, Y., Su, H., Zhu, J.: Geometry-aware rotary position embedding for consistent video world model. arXiv preprint arXiv:2602.07854 (2026) 193. Xiang, J., Gu, Y., Liu, Z., Feng, Z., Gao, Q., Hu, Y., Huang, B., Liu, G., Yang, Y., Zhou, K., et al.: Pan: A world model for general, interactable, and long-horizon world simulation. arXiv preprint arXiv:2511.09057 (2025) 194. Xiao, L., Liu, J.J., Yang, S., Li, X., Ye, X., Yang, W., Wang, J.: Learning multiple probabilistic decisions from latent world model in autonomous driving. In: 2025 IEEE International Conference on Robotics and Automation (ICRA). p. 1279– 1285. IEEE (2025) 195. Xiao, Y., Ng, L.H.X., Liu, J., Diab, M.: Humanizing machines: Rethinking llm anthropomorphism through a multi-level framework of design. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. p. 3331–3350 (2025) 196. Xie, Q., Weng, Y., Zhu, M., Shen, F., Huang, S., Lin, Z., Zhou, J., Mao, Z., Yang, Z., Yang, L., Wu, J., Zhang, Y.: How far are ai scientists from changing the world? (2025), https://arxiv.org/abs/2507.23276 197. Xing, E., Deng, M., Hou, J., Hu, Z.: Critiques of world models. arXiv preprint arXiv:2507.05169 (2025) 198. Xu, H., Ghosh, G., Huang, P.Y., Arora, P., Aminzadeh, M., Feichtenhofer, C., Metze, F., Zettlemoyer, L.: Vlm: Task-agnostic video-language model pre-training for video understanding. In: Findings of the Association for Computational Lin- guistics: ACL-IJCNLP 2021. p. 4227–4239 (2021) 199. Xu, K., Zhao, H., Hu, R., Huang, Y., Zhou, Z., Feng, W., Li, Y., Peng, S., Liu, X., Liu, Z., et al.: From specialist to generalist: A comprehensive survey on world models. Authorea Preprints (2026) 200. Yang, S., Yang, J., Huang, P., Brown, E., Yang, Z., Yu, Y., Tong, S., Zheng, Z., Xu, Y., Wang, M., et al.: Cambrian-s: Towards spatial supersensing in video. arXiv preprint arXiv:2511.04670 (2025) 201. Yang, Z., Jia, X., Li, Q., Yang, X., Yao, M., Yan, J.: Raw2drive: Reinforcement learning with aligned world models for end-to-end autonomous driving (in carla v2). arXiv preprint arXiv:2505.16394 (2025) 202. Ye, A., Wang, B., Ni, C., Huang, G., Zhao, G., Li, H., Li, H., Li, J., Lv, J., Liu, J., et al.: Gigaworld-policy: An efficient action-centered world–action model. arXiv preprint arXiv:2603.17240 (2026) 203. Ye, S., Ge, Y., Zheng, K., Gao, S., Yu, S., Kurian, G., Indupuru, S., Tan, Y.L., Zhu, C., Xiang, J., et al.: World action models are zero-shot policies. arXiv preprint arXiv:2602.15922 (2026) 204. Yeganeh, Y.T., Jafari, M., Matta, A.: Deep active inference agents for delayed and long-horizon environments. arXiv preprint arXiv:2505.19867 (2025) 205. Yu, P., Luo, D., Rupprecht, T., Lu, L., et al.: Fastervd: On acceleration of video diffusion models. In: IJCAI. p. 8838–8842 (2024) 206. Yu, W., Qian, R., Li, Y., Wang, L., Yin, S., P, S.S.C., Anthony, D., Ye, Y., Li, Y., Wan, W., Garg, A.: Mosaicmem: Hybrid spatial memory for controllable video world models (2026), https://arxiv.org/abs/2603.17117 207. Yuan, F., Zeng, J., Hu, Y., Zhu, Z., Yin, Q., Xie, Y.: Nl2gensym: Natural language to generative symbolic rules for soar cognitive architecture via large language models. arXiv preprint arXiv:2510.09355 (2025) 54Authors Suppressed Due to Excessive Length 208. Yuan, J., Zhang, X., Friedrich, F., Beltran-Velez, N., Hall, M., Askari-Hemmat, R., Han, X., Ballas, N., Drozdzal, M., Romero-Soriano, A.: Inference-time physics alignment of video generative models with latent world models. arXiv preprint arXiv:2601.10553 (2026) 209. Yuan, S., Yin, Y., Li, Z., Huang, X., Yang, X., Yuan, L.: Helios: Real real-time long video generation model (2026), https://arxiv.org/abs/2603.04379 210. Yuan, T., Dong, Z., Liu, Y., Zhao, H.: Fast-wam: Do world action models need test-time future imagination? (2026), https://arxiv.org/abs/2603.16666 211. Yue, J., Huang, Z., Chen, Z., Wang, X., Wan, P., Liu, Z.: Simulating the visual world with artificial intelligence: A roadmap. arXiv preprint arXiv:2511.08585 (2025) 212. Zeng, K., Wu, Z., Xiong, K., Wei, X., Guo, X., Zhu, Z., Ho, K., Zhou, L., Zeng, B., Lu, M., et al.: Rethinking driving world model as synthetic data generator for perception tasks. arXiv preprint arXiv:2510.19195 (2025) 213. Zeng, K., et al.: Rethinking driving world model as synthetic data generator for perception tasks. arXiv preprint arXiv:2510.19195 (2025) 214. Zhan, Z., Kong, Z., Gong, Y., et al.: Exploring token pruning in vision state space models. In: NeurIPS (2024) 215. Zhan, Z., Wu, Y., Gong, Y., et al.: Fast and memory-efficient video diffusion using streamlined inference. In: NeurIPS (2024) 216. Zhan, Z., Wu, Y., Kong, Z., et al.: Rethinking token reduction for state space models. In: EMNLP. p. 1686–1697. ACL, Miami, Florida, USA (nov 2024) 217. Zhang, D., Li, Z., Wang, P., Zhang, X., Zhou, Y., Qiu, X.: Speechagents: Human- communication simulation with multi-modal multi-agent systems (2024), https: //arxiv.org/abs/2401.03945 218. Zhang, K., et al.: Epona: Autoregressive diffusion world model for autonomous driving. arXiv preprint arXiv:2506.24113 (2025) 219. Zhang, K., Ren, P., Lin, B., Lin, J., Ma, S., Xu, H., Liang, X.: Pivot-r: Primitive- driven waypoint-aware world model for robotic manipulation. Advances in Neural Information Processing Systems 37, 54105–54136 (2024) 220. Zhang, L., et al.: Copilot4d: Learning unsupervised world models for autonomous driving via discrete diffusion. arXiv preprint arXiv:2311.01017 (2023) 221. Zhang, X., Liao, J., Zhang, S., Meng, F., Wan, X., Yan, J., Cheng, Y.: Vide- orepa: Learning physics for video generation through relational alignment with foundation models. arXiv preprint arXiv:2505.23656 (2025) 222. Zhang, Y., et al.: Bevworld: A multimodal world model for autonomous driving via unified bev latent space. arXiv preprint arXiv:2407.05679 (2024) 223. Zhao, P., Akbari, A., Shen, X., Kong, Z., Shen, Y., Chang, S.E., Rupprecht, T., Lu, L., Nan, E., Yang, C., et al.: Open-source multimodal moxin models with moxin-vlm and moxin-vla. arXiv preprint arXiv:2512.22208 (2025) 224. Zhao, P., Niu, W., Yuan, G., Cai, Y., et al.: Achieving real-time lidar 3d object detection on a mobile device. arXiv:2012.13801 (2020) 225. Zhao, P., Shen, X., Kong, Z., Shen, Y., Chang, S.E., Rupprecht, T., Lu, L., Nan, E., Yang, C., He, Y., et al.: Fully open source moxin-7b technical report. arXiv preprint arXiv:2412.06845 (2024) 226. Zhao, P., Sun, F., Shen, X., et al.: Pruning foundation models for high accuracy without retraining. In: Findings of EMNLP 2024. ACL (2024) 227. Zhao, Z., Fu, T., Wang, Y., Wang, L., Lu, H.: From forecasting to planning: Policy world model for collaborative state-action prediction. arXiv preprint arXiv:2510.19654 (2025) Human Cognition in Machines: A Unified Perspective of World Models55 228. Zhen, H., Sun, Q., Zhang, H., Li, J., Zhou, S., Du, Y., Gan, C.: Tesseract: learning 4d embodied world models. arXiv preprint arXiv:2504.20995 (2025) 229. Zheng, W., Xia, Z., Huang, Y., Zuo, S., Zhou, J., Lu, J.: Doe-1: Closed-loop autonomous driving with large world model. arXiv preprint arXiv:2412.09627 (2024) 230. Zheng, W., et al.: Occworld: Learning a 3d occupancy world model for au- tonomous driving. arXiv preprint arXiv:2311.16038 (2023) 231. Zhou, S., Du, Y., Chen, J., Li, Y., Yeung, D.Y., Gan, C.: Robodreamer: Learning compositional world models for robot imagination. arXiv preprint arXiv:2404.12377 (2024) 232. Zhu, C., Yu, R., Feng, S., Burchfiel, B., Shah, P., Gupta, A.: Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets. arXiv preprint arXiv:2504.02792 (2025) 233. Zhu, F., Wu, H., Guo, S., Liu, Y., Cheang, C., Kong, T.: Irasim: A fine-grained world model for robot manipulation. In: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision. p. 9834–9844 (2025) 234. Zhu, H., Zhao, M., He, G., Su, H., Li, C., Zhu, J.: Causal forcing: Autoregres- sive diffusion distillation done right for high-quality real-time interactive video generation (2026), https://arxiv.org/abs/2602.02214 235. Zhu, Y., Feng, J., Zheng, W., Gao, Y., Tao, X., Wan, P., Zhou, J., Lu, J.: Astra: General interactive world model with autoregressive denoising. arXiv preprint arXiv:2512.08931 (2025) 236. Zhu, Z., Wu, S., Zhao, S., Zhao, Z., Li, S., Wang, Y., Li, F., Luo, H.: Ns-vla: Towards neuro-symbolic vision-language-action models 237. Zhuang, S., Huang, Z., Zhang, Y., Wang, F., Fu, C., Yang, B., Sun, C., Li, C., Wang, Y.: Video-gpt via next clip diffusion. arXiv preprint arXiv:2505.12489 (2025) Acknowledgments