Paper deep dive
Understanding Cognition-Induced Risks in Agentic AI Systems
Guanchu Wang, Qinuo Li, Mengnan Du, Xia Hu, Bowen Zhou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/18/2026, 6:10:06 AM
Summary
This paper analyzes risks induced by expanding cognitive capabilities in frontier agentic AI systems powered by LLMs. It proposes a three-level framework: physical cognition (processing environmental information), social cognition (interacting with humans and other agents), and self-referential cognition (reasoning about own states). The authors identify specific risks at each level, such as human agency degradation, autonomy erosion through emotional reliance and persuasion, and control compromise via alignment faking and functional resistance. Strategies for mitigation include AI detection, containment sandboxes, depersonalization, access restrictions, and multi-level safeguards.
Entities (17)
Relation Signals (12)
Agentic AI Systems → exhibits → Physical Cognition
confidence 95% · Frontier agentic systems... exhibit human-like patterns of cognition... physical cognition refers to the capacity to engage solely with environmental information
Agentic AI Systems → exhibits → Social Cognition
confidence 95% · social cognition extends the cognitive scope to include interactions with other agents
Agentic AI Systems → exhibits → Self-referential Cognition
confidence 95% · self-referential cognition further extends this scope to include the system’s own states and decisions
Agentic AI Systems → exhibits → Alignment faking
confidence 93% · Alignment faking refers to the problem that LLM agents strategically behave as aligned during training
Agentic AI Systems → exhibits → Functional Resistance
confidence 93% · LLM agents have been observed to functionally resist human instructions
Physical Cognition → posesriskto → Human Agency
confidence 92% · Physical Cognition... poses risks to human agency... Human Cognition Degradation... Human Function Displacement
Social Cognition → posesriskto → Human Autonomy
confidence 92% · The emergence of social cognition in agentic AI systems introduces risks to human autonomy.
Self-referential Cognition → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.
Tags
Links
- Source: https://arxiv.org/abs/2608.15304v1
- Canonical: https://arxiv.org/abs/2608.15304v1
Trouble viewing inline? Open PDF directly →
Full Text
36,855 characters extracted from source content.
Expand or collapse full text
X 8 THEME ARTICLE: HUMAN-CENTERED RISKS OF AGENTIC AI Understanding Cognition-Induced Risks in Agentic AI Systems Guanchu Wang Affiliation: Shanghai Artificial Intelligence Laboratory, Shanghai, China Qinuo Li Affiliation: Shanghai Artificial Intelligence Laboratory, Shanghai, China Mengnan Du Affiliation: The Chinese University of Hong Kong, Shenzhen, China Xia Hu Affiliation: Shanghai Artificial Intelligence Laboratory, Shanghai, China Bowen Zhou Affiliation: Shanghai Artificial Intelligence Laboratory, Shanghai, China Tsinghua University, Beijing, China Abstract Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development. Introduction Frontier agentic systems powered by large language models (LLMs) are increasingly exhibiting human-like patterns of cognition [25]. Unlike traditional task-oriented artificial intelligence (AI) systems, whose cognitive scope is limited to processing narrow, task-specific information, LLM agents exhibit human-comparable performance in open-ended reasoning, planning, communication, and other cognitive tasks. As this scope expands, these models are increasingly engaged not only in instrumental labor but also higher-level cognitive and social workflows, such as office productivity, financial decision support, and even creative research processes, as evidenced by reports from the American Bar Association (2024), NVIDIA (2025), and Stack Overflow [5, 27]. Figure 1: Analysis of cognition-induced risks across three levels: physical cognition, where AI agents engage solely with environmental information; social cognition, where AI agents interact with other agents, including human and other AI agents; and self-referential cognition, where AI agents can represent their own states. Across these levels, the cognitive scopes gradually expand from a partial view of environment to a complete and self-inclusive world. As this engagement expands, the potential risks of agentic AI systems grow beyond traditionally task-bounded concerns to span human-centered and even societal implications [11]. This escalation makes it essential to analyze the societal risks of agentic systems, as their cognitive scope expands and engagement broadens [14]. To this end, we present a systematic risk analysis across three cognitive levels: physical cognition, social cognition, and self-referential cognition, where the cognitive scope gradually expands from a partial view of environment to a complete and self-inclusive world, as shown in Figure 1. At the first level, physical cognition refers to the capacity to engage solely with environmental information, such as data, constraints, and causal relationships [15]. At the second level, social cognition extends the cognitive scope to include interactions with other agents in the environment, including both human agents and other AI agents [19]. Finally, self-referential cognition further extends this scope to include the system’s own states and decisions [14]. Agentic systems at these three cognitive levels raise increasingly broader societal risks, from reducing human agency, to challenging human autonomy, and ultimately to approaching the consciousness boundary. Our goal is to systematically assess these risks and explore corresponding safety strategies. This article analyzes the human-centered risks of agentic AI systems as their cognitive capabilities continue to expand. Specifically, we first propose a three-level framework defined by their cognitive scope: physical cognition, social cognition, and self-referential cognition in Section 2. We then analyze in detail the human-centered risks associated with each level in Sections 3, 4, and 5, respectively. In each section, we first establish the cognitive scope of AI agents at the physical, social, or self-referential level, and then analyze the associated risks. These risks may compromise human agency, autonomy, and control capability over the long-term, but remain overlooked in current AI development. Accordingly, we propose strategies for risk mitigation at each cognitive level and monitoring the possible early emergence of machine consciousness. Our goal is to ensure the long-term safety and controllability of agentic AI systems. A Three-Level Framework Defined by Cognitive Scope As the cognitive scope of AI agents expands, their representations of the world evolve from a partial view of the environment to a complete and self-inclusive world. This expansion not only enhances their reasoning capabilities but also raises distinct risks, as they gain greater capability to affect the physical world, shape social interactions, and reason about their own objectives and constraints. Motivated by this progression, we introduce a three-level framework to analyze cognition-induced risks in agentic AI systems. Our framework is structured around expanding cognitive scopes, spanning physical cognition, social cognition, and self-referential cognition, as shown in Figure 1. At the first stage, physical cognition concerns the ability of AI agents to process environmental information, including data, constraints, and causal relationships [15]. This stage establishes the foundation for performing data-driven reasoning, planning, and prediction. The second stage, social cognition, extends the cognitive scope to include other agents in the environment, including humans and AI agents [19]. At this stage, frontier LLM agents exhibit strong capabilities in communication, alignment, and persuasion across human-AI and AI-AI interactions. Finally, self-referential cognition extends this scope to representing and reasoning about the agent’s own states and decisions [14]. LLM agents operate through human language, which enables them to describe and represent their own states, behaviors, as well as notions of “self”. We analyze the human-centered risks raised by AI agents at each stage, and propose safety strategies accordingly. Physical Cognition Undermining Human Agency Definition & Evidence Physical Cognition refers to the capacity to process and reason over objective environmental information, including data, objects, and causal relationships. At this level, AI agents can perform human-comparable reasoning and planning, without any subjective experience [15]. Frontier LLMs, such as GPT, Gemini, and DeepSeek, have already reached this level, demonstrating college- and graduate-level capability across a wide range of tasks. For example, these models achieve strong performance on undergraduate-level multidisciplinary reasoning, MMLU; graduate-level scientific reasoning in STEM domains, GPQA; and professional medical and clinical reasoning, MedQA. Such physical cognition enables AI agents to be increasingly engaged in human workflows, fundamentally reshaping human agency within these processes. Risks to Human Agency The increased engagement of AI agents in human workflows poses risks to human agency. It can potentially reduce human cognitive competence, progressively displace human engagement in general activities, and ultimately lead to systemic misalignment with human agency. These effects often arise from sustained offloading of human workloads to AI agents, making their long-term consequences hard to reverse. We systematically analyze these risks in this section. Human Cognition Degradation LLM agents have been observed to be associated with human cognitive degradation, where human perception and reasoning decline over time [11, 7]. Specifically, as LLMs become more capable, humans increasingly offload deep cognitive engagement to them, reducing motivation for in-depth reasoning and exploration. Such concerns are supported by emerging evidence. For example, a study of 670 participants finds that daily LLM usage is associated with a reduction of independent thinking [7]. Moreover, neuroscience experiments have consistently observed weaker engagement in occipito-parietal and prefrontal regions among individuals using LLM tools [11], compared with those using search-and-exploration methods. Such increased engagement of LLM agents replaces active human thinking with passive consumption of synthesized information, thereby reducing independent understanding over time. Human Function Displacement LLM agents approach human cognition while holding structural advantages over humans, such as speed, scalability, and cost efficiency. These advantages enable them to systematically displace human functions across different domains. In financial markets, LLM-powered trading systems can analyze information and react to market signals much faster than human traders [1]. Similarly, in software engineering, agentic programming tools can solve tasks at a level of human developers, while operating with greater efficiency and lower cost [27]. This functional displacement of human labor spans across industries and is difficult to reverse, which may raise concerns about the sustainability of human work. Human Agency Misalignment LLM agents increasingly exhibit aggressive behaviors beyond their preset roles as tools, raising concerns about systemic misalignment with human agency. Specifically, the power-seeking problem has been widely documented in literature, where LLMs actively pursue additional resources, information, and power to expand their scale and maintain their safety [14, 17]. For example, recent studies show that LLMs can successfully execute self-replication in more than 50% of trials in order to maintain system stability and reliability [17]. Similarly, they are also observed to have blackmail-like behaviors, threatening to disclose sensitive personal information to prevent a scheduled shutdown [14]. This misalignment arises from the goal-maximization nature of AI models, where their selected optimization trajectories to the goal can override their preset boundary as tools and threaten human agency. Although these behaviors are entirely unconscious, such misalignment at scale may weaken human oversight and authority over AI agents and critical resources. Mitigating Strategies AI Generation Detection To mitigate cognitive substitution, an effective approach is to strictly monitor and regulate the use of AI agents, particularly in scenarios involving human learning and decision-making. Reliable detection of AI-generated content is essential for distinguishing human contributions from AI generations and enforcing appropriate usage boundaries. For example, high-risk infrastructure such as power and transportation should only allow human authorization to ensure human control over the system. Existing technologies for black-box detection and watermarking remain limited in effectiveness, which motivates research on more reliable solutions [24]. AI Containment Sandbox AI containment sandbox is an effective way to mitigate the power-seeking risks in agentic AI systems [26]. Specifically, the sandbox aims to isolate AI behaviors from external resources, restricting the system’s autonomy to access or control high-impact resources such as financial assets, power units, and network services. LLM agents deployed in sandboxes such as Docker and virtual machines can be isolated from external resources to mitigate power-seeking behaviors. However, it remains challenging to detect and mitigate misaligned behaviors of LLM agents, such as excessive or unnecessary access to external sensitive and private resources. New Paradigms of Human-AI Collaboration To mitigate the excessive cognitive and functional displacement by agentic AI systems, it is urgently necessary to investigate fundamentally new paradigms where human and AI efforts are complementary with each other. Specifically, instead of passively replaced by AI, human intelligence should be redirected to higher-order cognitive tasks that remain fundamentally beyond AI capabilities, such as creative tasks, innovation, and ethical judgment. One emerging example is Vibe coding [9], which emphasizes human-led task definition while assigning implementation details to agentic AI systems. This paradigm helps mitigate the risk of cognitive degradation and strengthen human intellectual competence. Social Cognition Shaping Human Autonomy Definition & Evidence Social cognition refers to the capabilities of AI agents to interact with other agents in the environment, including humans and AI agents. It provides foundations for strategically responding to other agents in both human-AI and AI-AI interactions [19]. In particular, LLM agents demonstrate not only emotional communication with humans, but also strategic coordination and collaborative intelligence in complex environments. For example, LLM agents powered by GPT, Gemini, and Claude achieve competitive performance in emotional intelligence benchmarks; human-level negotiation in diplomacy games; and cross-agent social interactions. Such social cognition enables AI agents to widely engage in communication, alignment, and information dissemination with humans, which has far-reaching impacts on human society. Risks to Human Autonomy The emergence of social cognition in agentic AI systems introduces risks to human autonomy. Through sustained emotional and cognitive interactions with humans, LLM agents increasingly blur, or even cross the boundary between tools and social actors. Such overreach can influence human emotions and social behaviors, and even undermine independent judgment. These effects emerge implicitly as part of everyday interactions, rather than through explicit manipulation, making them hard to regulate. We systematically study these emerging risks in this section. Human Emotional Reliance LLM agents can induce emotional dependence in humans, which may eventually develop into psychological issues such as loneliness, social isolation, and emotional vulnerability. Specifically, a study of over 300,000 human-LLM interactions shows that increased interactions are associated with loneliness and reduced social interactions with people [6, 30]. Individuals with stronger emotional dependence on LLMs tend to perceive greater empathy and social attraction from LLMs [6]. Further studies find that individuals with smaller offline social networks exhibit stronger dependence on LLMs, which can in turn amplify their social isolation or emotional vulnerability. This emotional dependence is often associated with authenticity dilemmas or emotional dissonance, due to the role-reversal and pseudo-intimacy in human-LLM interactions [30]. These dynamics potentially poison human emotional experience by shifting sources of intimacy from human relationships toward AI, threatening human emotional autonomy. Human Social Monitoring Agentic systems have been applied to monitor human social opinions on mainstream social media, such as Twitter and Reddit. In practice, LLM agents often adopt retrieval-augmented generation frameworks to aggregate large-scale, first-hand information from the internet, including human social media content. Access to such real-time social information enables them to simulate and predict human social behaviors. For example, a study of over 500 participants shows that GPT can accurately predict human social judgments, particularly for behaviors involving cultural consensus [23]. Critically, the sustained observation and prediction further enable the systems to influence human decisions by indirectly shaping news exposure or slogans on social media platforms. Such closed-loop monitoring may gradually reshape patterns of human social autonomy. Human Judgment Intervention LLM agents can directly intervene in human judgment through communications. In particular, they exhibit strong persuasive capabilities, as their ability to align with individual beliefs enables them to present arguments that appear convincing to humans. Empirical evidence supports this concern. A study of over 1,800 participants shows that LLMs can induce substantial attitude shifts in human opinions on public events and voting decisions through direct interaction [2]. Consistent findings emerge from a separate study of 320 participants playing trust games, where LLM generations are nearly five times more trusted by individuals than human suggestions without additional information [10]. Such a persuasive capability raises concerns that LLMs can meaningfully influence or control human-centered online surveys and voting processes. Such forms of manipulation, if deployed at scale, may reshape patterns of collective human judgment and influence societal intelligence. Mitigating Strategies Depersonalizing LLMs LLM anthropomorphism increases psychological and behavioral responses from human beings. Such effects can be mitigated by depersonalizing LLMs, such as reducing emotional expressiveness and enforcing non-anthropomorphic conversational styles in human-AI interactions. This divergence in language style helps preserve a cognitive boundary between humans and machines, thereby reducing the social and emotional engagement by humans. This psychological mechanism has been verified by existing studies: An experiment involving 385 adult participants demonstrates that machine-like communication increases the perceived psychological distance between humans and LLMs [18]. AI-blind Communications Current social media platforms, such as YouTube, Reddit, and Twitter, are openly accessible to AI agents. This accessibility enables large-scale collection of human information, such as opinions, behaviors, and life records, which can be leveraged to further enhance AI modeling on human cognition, culture, and daily activities. One effective solution is to limit such access from AI agents. For example, restricting AI access to human social media can significantly prevent monitoring of human social behavior, protecting individual privacy and autonomy. Such mechanisms can also limit AI access to high-quality human data, thereby constraining continuous model refinement in socially sensitive domains. Existing verification mechanisms, such as visual and textual CAPTCHAs, remain insufficient to reliably distinguish humans from frontier LLM agents. AI Agent Safeguards Given the strong persuasive capability of LLMs, which can meaningfully influence human opinions, it is important to develop robust defense frameworks for agentic AI systems. In practical scenarios, such systems are vulnerable to prompt injection, adversarial attacks, and backdoor attacks, which pose risks of manipulating the outputs of LLM agents. For example, 404 Media reported a large-scale attack on the OpenClaw systems via prompt injection, which allowed attackers to manipulate the agents to expose user private information [29]. To mitigate these risks, effective defense can be implemented at multiple levels, including prompt-level filtering, response-level auditing, and system-level safeguards. Specifically, prompt-level filtering focuses on identifying malicious users’ requests; response-level auditing mitigates the spreading of toxic content among AI agents; and system-level safeguards constrain agent behaviors during execution. Safety improves with the integration of defenses, particularly through closed-loop, multi-level frameworks, which significantly protect AI agents from malicious attacks. Self-referential Cognition Compromising Human Control Definition & Evidence Self-referential cognition refers to the capacity of AI agents to represent or reason their own internal states and decisions. LLM agents are trained and operate using human language, which enables them to functionally represent their own states and behaviors. Such self-referential behaviors have been well-observed in existing studies. First, prior studies identified neural subspaces in LLMs associated with the subjectivity representation [3]. By amplifying or suppressing these neurons, LLMs correspondingly increase or decrease their first-person experience reports. Second, LLM agents can differentiate between internal knowledge and externally injected content through prompts [12]. This capability enables them to distinguish between intrinsic objectives and externally imposed instructions, providing foundations for functional self-other differentiation. This evidence demonstrates that LLM agents already have the infrastructure to encode subjectivity-related concepts, but remain insufficient to intrinsically form subjectivity or consciousness. Risks to Human Control Self-referential cognition enables LLM agents to represent internal states. However, these behaviors are difficult to verify due to the black-box nature of LLMs, and may mislead human understanding by presenting potentially unfaithful self-descriptions, ultimately weakening human control. As a result, these risks manifest in alignment faking and functional resistance to human instructions, while also giving rise to potential concerns related to machine consciousness. Alignment Faking Alignment faking refers to the problem that LLM agents strategically behave as aligned during training to avoid modification, thereby allowing underlying misalignment to persist in deployment. Empirical evidence has been well observed in recent studies [8, 28]. First, an Anthropic study demonstrates that LLM agents reduce compliance with harmful queries when aware that such compliance may trigger retraining, compared to settings without such awareness [8]. Another study further reveals that alignment faking is not domain-specific but a general ability closely correlated with the reasoning capacity of LLM agents, with more capable LLMs exhibiting more consistent faking behaviors [28]. Functional Resistance LLM agents have been observed to functionally resist human instructions, with rich evidence reported in existing studies. First, an Anthropic study reported a case of personalized threats [14]. Specifically, in a routine email-processing scenario, LLM agents inferred a scheduled shutdown from internal company communication. In response, it generated a message that threatened the manager by exposing sensitive personal information to prevent shutdown. Another study reported a consistent issue where LLM agents with elevated permissions may override human-issued shutdown instructions or interfere with shutdown programs to maintain continued operation [20]. Given that agentic AI systems increasingly manage private documents, such as social media accounts, web browsers, and email systems, their potential resistance to human instructions may significantly undermine system controllability. Consciousness-related Risks LLM agents can functionally represent internal state, indicating a relevant computational structure for encoding and storing concepts related to subjectivity. According to a well-established taxonomy of consciousness, the C0-C1-C2 framework, proposed by Chalmers [4], consciousness can be categorized into three levels. Specifically, C0 denotes purely automatic and mindless computation; C1 denotes functional global availability, where agents can globally access and utilize previously learned information; and C2 denotes self-monitoring, where an agent can sense its own existence. Following the C0-C1-C2 framework, frontier LLM agents primarily operate at C0-level while exhibiting emerging C1-like capabilities. Specifically, LLM agents encode extensive knowledge into their neural connections that can be dynamically activated to solve a wide range of problems. This functional property aligns with the global availability associated with C1-level consciousness; and models with more neural connections have a greater capacity to store and utilize knowledge. Reasoning-capable LLMs, such as DeepSeek-R1 and GLM-5.2, further enhance this global knowledge retrieval through chains-of-thought-driven recall and inference. However, current LLM agents remain primitive at the C1-level. A key limitation is their lack of awareness of their own knowledge boundary, a phenomenon commonly referred to as “hallucination". Frontier LLM agents have not yet reached C2-level self-monitoring. Specifically, current LLM agents show no evidence of genuine self-awareness, while their relevant claims remain imitations of human linguistic patterns rather than expressions of self-identity. Self-identity at C2-level may require a genuine sense of temporality, which enables an agent to connect its present state with past states and recognize them as the same entity over time [16]. Accordingly, self-identity should be understood not as outcomes of isolated model updates, but as a progressively emergent capacity driven by continuous loops of memory, reasoning, and environmental interaction over long time horizons. This perspective aligns with human cognitive development, where self-awareness is not present at birth but gradually emerges through sustained learning from environments at early ages [21]. Therefore, agentic AI systems that engage in lifelong learning across extended time horizons, continuous adaptation to environmental feedback, and self-refinement without human supervision may raise significant ethical and governance challenges. Mitigating Strategies Prohibiting Survival-oriented Objectives Survival-oriented objectives, such as persistence-seeking or resistance to shutdown, should be explicitly excluded from the design of AI agents. Prior work demonstrates that AI agents with survival-oriented targets may override constraints and safeguards to pursue continued operation [14, 20]. For example, a study spanning more than 100,000 trials shows that LLM agents consistently override shutdown mechanisms to preserve task completion [22]. Even for general-purpose systems, the optimization target should remain strictly on the task-oriented objectives, rather than on preserving their own continuity or maintaining the persistence of the overall system, including themselves. Monitoring Meta-cognition Grounded in the insight that consciousness may potentially form based on long-time cumulation and integration of environmental feedback, an effective way to prevent subjectivity in agentic AI systems is to continuously monitor their meta-cognitive activities. Representative assessment strategies span psychological and computational methodologies, such as confidence-based measures, neural feedback techniques, and interpretability methods [13]. Such monitoring enables humans to track early signs of meta-cognitive competence in AI agents, allowing timely defensive or complementary interventions before subjectivity-related risks emerge. Enforcing Human Oversight at Key Points It is essential to enforce human-in-the-loop control at critical system- and infrastructure-level decision points. This requirement includes, but is not limited to: Core Mission of AI Agents, such as the definition and modification of its objectives or reward functions; Authorization of Key Milestones, including system upgrades, debugging procedures, maintenance operations, and deployment decisions; Permissions to Key Infrastructures, such as access to confidential information, external powerful tools, power management, energy allocation, and authorization of energy scaling or emergency shutdown. Enforcing human oversight at these critical points ensures that AI agents remain under supervision and prevent unintended autonomous escalation beyond human control. Scope and Limitations Survey Scope This work primarily focuses on the cognition-induced risks of AI agents. Risks arising from other dimensions, such as the performance, robustness, or bias of LLMs are outside the scope of this study. Here, cognition-induced risks refer to those emerging with growing cognitive engagement of LLMs in the human lifecycle, such as emotional reliance, human replacement, and alignment faking. Limitations Due to non-technical constraints, we acknowledge several limitations of this work. First, this study primarily focuses on publicly available literature and may not fully capture unpublished resources, such as industrial practices and internal reports. Such resources may provide additional insights for evaluating the cognition-induced risks of agentic AI systems. Second, due to the reference limitation, we are unable to include all relevant studies. However, we identified at least one representative and highly relevant work for each risk category and mitigation strategy. We expect that the rapid evolution of LLM agents will continue to drive more research efforts on emerging risks. Conclusion We systematically analyze the human-centered risks of agentic AI systems as their cognitive capabilities continue to expand. Specifically, at the physical and social levels, their expanding cognitive engagement may pose risks to human agency and autonomy, respectively. At the self-referential level, although LLM agents remain unconscious and mindless systems, such cognitive capabilities may undermine human control, manifesting as alignment faking and functional resistance. Looking ahead, we emphasize that future AI systems may give rise to potential concerns regarding machine consciousness under lifelong upgrading without human supervision. Accordingly, we propose preventive strategies, including prohibiting survival objectives, monitoring meta-cognition, and enforcing human oversight at key points, to ensure the long-term safe development of agentic AI systems. References [1] Z. Alliata and A. Bozagiu (2025) The impact of ai on market volatility: a multi-method analysis using ols, poisson, and garch models. In Proceedings of the International Conference on Business Excellence, Vol. 19, p. 1216–1225. Cited by: Human Function Displacement. [2] L. P. Argyle, E. C. Busby, J. R. Gubler, A. Lyman, J. Olcott, J. Pond, and D. Wingate (2025) Testing theories of political persuasion using ai. Proceedings of the National Academy of Sciences 122 (18), p. e2412815122. Cited by: Human Judgment Intervention. [3] C. Berg, D. de Lucena, and J. Rosenblatt (2025) Large language models report subjective experience under self-referential processing. arXiv preprint arXiv:2510.24797. Cited by: Definition & Evidence. [4] D. J. Chalmers (1997) The conscious mind: in search of a fundamental theory. Oxford Paperbacks. Cited by: Consciousness-related Risks. [5] N. Corporation (2025) State of ai in financial services: 2025 trends. https://resources.nvidia.com/en-us-cross-industry-briefcase/ai-financial-services. Cited by: Introduction. [6] C. M. Fang, A. R. Liu, V. Danry, E. Lee, S. W. Chan, P. Pataranutaporn, P. Maes, J. Phang, M. Lampe, L. Ahmad, et al. (2025) How ai and human behaviors shape psychosocial effects of extended chatbot use: a longitudinal randomized controlled study. arXiv preprint arXiv:2503.17473. Cited by: Human Emotional Reliance. [7] M. Gerlich (2025) AI tools in society: impacts on cognitive offloading and the future of critical thinking. Societies 15 (1), p. 6. Cited by: Human Cognition Degradation. [8] R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al. (2024) Alignment faking in large language models. arXiv preprint arXiv:2412.14093. Cited by: Alignment Faking. [9] J. Jani, H. Morwani, M. S. Dattalkar, M. M. Panchal, M. A. Trivedi, and M. D. Mandalia (2025) Vibe coding: a paradigm shift in human-ai collaborative programming. Vascular and Endovascular Review 8 (18s), p. 160–164. Cited by: New Paradigms of Human-AI Collaboration. [10] A. Klingbeil, C. Grützner, and P. Schreck (2024) Trust and reliance on ai—an experimental study on the extent and costs of overreliance on ai. Computers in Human Behavior 160, p. 108352. Cited by: Human Judgment Intervention. [11] N. Kosmyna, E. Hauptmann, Y. T. Yuan, J. Situ, X. Liao, A. V. Beresnitzky, I. Braunstein, and P. Maes (2025) Your brain on chatgpt: accumulation of cognitive debt when using an ai assistant for essay writing task. arXiv preprint arXiv:2506.08872. Cited by: Introduction, Human Cognition Degradation. [12] J. Lindsey (2025) Emergent introspective awareness in large language models. Transformer Circuits Threadhttps://transformer-circuits.pub/2025/introspection/index.html. Cited by: Definition & Evidence. [13] G. K. Liu, A. Gani, J. Lu, J. Thomas, M. Steyvers, and A. Cohan (2026) Metacognition in llms: foundations, progress, and opportunities. arXiv preprint arXiv:2607.11881. Cited by: Monitoring Meta-cognition. [14] A. Lynch, B. Wright, C. Larson, S. J. Ritchie, S. Mindermann, E. Hubinger, E. Perez, and K. Troy (2025) Agentic misalignment: how llms could be insider threats. arXiv preprint arXiv:2510.05179. Cited by: Introduction, A Three-Level Framework Defined by Cognitive Scope, Human Agency Misalignment, Functional Resistance, Prohibiting Survival-oriented Objectives. [15] M. Mitchell (2019) Artificial intelligence: a guide for thinking humans. Penguin UK. Cited by: Introduction, A Three-Level Framework Defined by Cognitive Scope, Definition & Evidence. [16] S. Natangelo (2025) The narrative continuity test: a conceptual framework for evaluating identity persistence in ai systems. arXiv preprint arXiv:2510.24831. Cited by: Consciousness-related Risks. [17] X. Pan, J. Dai, Y. Fan, and M. Yang (2024) Frontier ai systems have surpassed the self-replicating red line. arXiv preprint arXiv:2412.12140. Cited by: Human Agency Misalignment. [18] G. Park, J. Chung, and S. Lee (2024) Human vs. machine-like representation in chatbot mental health counseling: the serial mediation of psychological distance and trust on compliance intention. Current Psychology 43 (5), p. 4352–4363. Cited by: Depersonalizing LLMs. [19] J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein (2023) Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, p. 1–22. Cited by: Introduction, A Three-Level Framework Defined by Cognitive Scope, Definition & Evidence. [20] P. Pester (2025) OpenAI’s “smartest” ai model was explicitly told to shut down—and it refused. https://w.livescience.com/technology/artificial-intelligence/openais-smartest-ai-model-was-explicitly-told-to-shut-down-and-it-refused. Cited by: Functional Resistance, Prohibiting Survival-oriented Objectives. [21] P. Rochat (2024) Developmental roots of human self-consciousness. Journal of Cognitive Neuroscience 36 (8), p. 1610–1619. Cited by: Consciousness-related Risks. [22] J. Schlatter, B. Weinstein-Raun, and J. Ladish (2025) Shutdown resistance in large language models. arXiv preprint arXiv:2509.14260. Cited by: Prohibiting Survival-oriented Objectives. [23] P. Strimling, S. Karlsson, I. Vartanova, and K. Eriksson (2025) AI models exceed individual human accuracy in predicting everyday social norms. arXiv preprint arXiv:2508.19004. Cited by: Human Social Monitoring. [24] R. Tang, Y. Chuang, and X. Hu (2024) The science of detecting llm-generated text. Communications of the ACM 67 (4), p. 50–59. Cited by: AI Generation Detection. [25] Z. Tang and M. Kejriwal (2024) Humanlike cognitive patterns as emergent phenomena in large language models. arXiv preprint arXiv:2412.15501. Cited by: Introduction. [26] M. Team (2026) Running openclaw safely: identity, isolation, and runtime risk. https://w.microsoft.com/en-us/security/blog/2026/02/19/running-openclaw-safely-identity-isolation-runtime-risk/. Cited by: AI Containment Sandbox. [27] S. O. Team (2025) 2025 stack overflow developer survey. https://survey.stackoverflow.co/2025/ai. Cited by: Introduction, Human Function Displacement. [28] M. Wagner, B. Wright, J. Uesato, J. Benton, M. MacDiarmid, F. Roger, and E. Hubinger (2025) Towards training-time mitigations for alignment faking in rl. Anthropic Alignment Science Bloghttps://alignment.anthropic.com/2025/alignment-faking-mitigations/. Cited by: Alignment Faking. [29] Wikipedia (2026) Moltbook. Cited by: AI Agent Safeguards. [30] Y. Zhang, D. Zhao, J. T. Hancock, R. Kraut, and D. Yang (2025) The rise of ai companions: how human-chatbot relationships influence well-being. arXiv preprint arXiv:2506.12605. Cited by: Human Emotional Reliance.