Paper deep dive
Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses
Mohsen Hatami, Van Tuan Pham, Hozefa Lakadawala, Yu Chen
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/11/2026, 1:04:55 AM
Summary
The paper introduces the SENTINEL framework, a lifecycle-aware methodology designed to secure AI agents within cyber-physical systems (CPS). It addresses emerging security risks such as deepfake-driven attacks, semantic manipulation, and vulnerabilities introduced by the Model Context Protocol (MCP). The survey emphasizes that detection-centric security is insufficient for safety-critical CPS, advocating instead for provenance- and physics-grounded trust mechanisms and defense-in-depth architectures.
Entities (5)
Relation Signals (3)
Model Context Protocol → expandsattacksurfaceof → Cyber-Physical Systems
confidence 95% · emerging protocols such as the Model Context Protocol (MCP) further expand the attack surface
Deepfake → compromises → AI Agents
confidence 92% · deepfake and semantic manipulation attacks that can compromise agent perception, reasoning, and interaction
SENTINEL → secures → AI Agents
confidence 90% · This survey introduces a Systematic Evaluation and Threat-Informed NEtwork defense seLection (SENTINEL) framework... to secure AI agents in CPS.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The increasing integration of AI agents into cyber-physical systems (CPS) introduces new security risks that extend beyond traditional cyber or physical threat models. Recent advances in generative AI enable deepfake and semantic manipulation attacks that can compromise agent perception, reasoning, and interaction with the physical environment, while emerging protocols such as the Model Context Protocol (MCP) further expand the attack surface through dynamic tool use and cross-domain context sharing. This survey provides a comprehensive review of security threats targeting AI agents in CPS, with a particular focus on environmental interactions, deepfake-driven attacks, and MCP-mediated vulnerabilities. We organize the literature using the SENTINEL framework, a lifecycle-aware methodology that integrates threat characterization, feasibility analysis under CPS constraints, defense selection, and continuous validation. Through an end-to-end case study grounded in a real-world smart grid deployment, we quantitatively illustrate how timing, noise, and false-positive costs constrain deployable defenses, and why detection mechanisms alone are insufficient as decision authorities in safety-critical CPS. The survey highlights the role of provenance- and physics-grounded trust mechanisms and defense-in-depth architectures, and outlines open challenges toward trustworthy AI-enabled CPS.
Tags
Links
- Source: https://arxiv.org/abs/2601.20184
- Canonical: https://arxiv.org/abs/2601.20184
Trouble viewing inline? Open PDF directly →
Full Text
275,845 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20261 Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses Mohsen Hatami, Van Tuan Pham, Hozefa Lakadawala, Yu Chen ∗ Dept. of Electrical & Computer Engineering, Binghamton University, Binghamton, NY 13902, USA mhatami1, tpham15, hlakada1, ychen@binghamton.edu Abstract—The increasing integration of AI agents into cyber- physical systems (CPS) introduces new security risks that extend beyond traditional cyber or physical threat models. Recent advances in generative AI enable deepfake and semantic manip- ulation attacks that can compromise agent perception, reasoning, and interaction with the physical environment, while emerging protocols such as the Model Context Protocol (MCP) further expand the attack surface through dynamic tool use and cross- domain context sharing. This survey provides a comprehensive review of security threats targeting AI agents in CPS, with a particular focus on environmental interactions, deepfake-driven attacks, and MCP-mediated vulnerabilities. We organize the literature using the SENTINEL framework, a lifecycle-aware methodology that integrates threat characterization, feasibility analysis under CPS constraints, defense selection, and continuous validation. Through an end-to-end case study grounded in a real-world smart grid deployment, we quantitatively illustrate how timing, noise, and false-positive costs constrain deployable defenses, and why detection mechanisms alone are insufficient as decision authorities in safety-critical CPS. The survey highlights the role of provenance- and physics-grounded trust mechanisms and defense-in-depth architectures, and outlines open challenges toward trustworthy AI-enabled CPS. Index Terms—Cyber-Physical Systems (CPS), AI Agents, Model Context Protocol (MCP), Deepfake Attacks, AI-Generated Content (AIGC), Security, Detection, Defense, Mitigation. I. INTRODUCTION As autonomous AI agents increasingly integrate into cyber- physical systems (CPS), the boundary between intelligence and the physical world blurs. These agents, endowed with perceptual capabilities, decision-making logic, and influence of actuations, promise enhanced efficiency, adaptability, and autonomy in domains such as smart grids, autonomous ve- hicles, industrial automation, and robotics [352]. However, as the AI-environment interface becomes a critical conduit, novel security and privacy vulnerabilities emerge, especially when adversaries leverage synthetic content, such as deepfakes, to deceive both machines and humans [55], [175]. While traditional CPS security research has focused on attacks such as sensor spoofing, denial-of-service (DoS), and insider threats [352], the advent of powerful generative models ushers in a new class of threats: AI-generated content (AIGC), including images, video, audio, text, and behavioral emulation. Such content can masquerade as legitimate environmental data or human commands, thus subverting the system operation or This manuscript is a preprint intended to rapidly disseminate a survey of security challenges and design principles for AI agents operating in cyber- physical systems. The authors anticipate submitting a substantially revised and polished version to a peer-reviewed journal. misleading human operators. For example, a deepfake video feed may spoof surveillance cameras, a cloned voice can impersonate an authorized operator, or forged sensor traces may mask a real physical anomaly [55], [145]. The capacity of deepfakes to mimic reality convincingly exacerbates the challenge: defenders cannot rely solely on human judgment or classical anomaly detectors [175]. In parallel, AI agents must navigate unpredictability at the interface: they ingest external data (sensor streams, images, voice) and produce actions in the physical world. The agent– environment (A–E) interaction surface is thus a rich target for adversarial manipulation. Malicious actors can exploit this sur- face by injecting crafted inputs or manipulating environmental signals to force misbehavior, cause unsafe actions, or conceal attacks [82]. A critical challenge is the absence of a systematic method- ology for translating the extensive body of security research into deployment-ready solutions for specific CPS contexts. While the literature provides numerous detection techniques, mitigation strategies, and threat analyses, system designers lack structured guidance for selecting and combining these mechanisms, given the unique constraints of their deploy- ments. CPS operates under fundamentally different conditions from traditional information technology systems, with real- time performance requirements measured in milliseconds, computational resources constrained by edge device capabil- ities, and safety-critical operational demands that preclude security mechanisms from interfering with physical process control. Furthermore, the distributed architecture of modern agent protocols, such as the Model Context Protocol (MCP), introduces trust boundaries and interaction surfaces that con- ventional security frameworks do not adequately address. To bridge this gap, this survey introduces a Systematic Eval- uation and Threat-Informed NEtwork defense seLection (SEN- TINEL) framework. It provides a six-phase methodology that guides system designers from initial threat assessment through selection of defense mechanisms, design of defense-in-depth architectures, validation planning, and continuous adaptation. The framework enables practitioners to systematically match security mechanisms to their specific deployment contexts by integrating threat modeling with resource constraint analysis and operational requirements. Throughout the subsequent sec- tions, we apply SENTINEL concepts to organize and contextu- alize the extensive technical content, transforming what would otherwise be a catalog of disconnected security techniques into actionable guidance for building trustworthy AI-enabled infrastructures. arXiv:2601.20184v1 [cs.CR] 28 Jan 2026 JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20262 This survey focuses on the intersection of these two con- verging axes: security and privacy threats arising from AIGC at the AI–environment interface in CPS. We aim to provide a rigorous, holistic view of how deepfake and synthetic content threats manifest within CPS, and how detection and mitigation techniques are evolving to meet these challenges. This survey advances the thesis that detection-centric security is funda- mentally insufficient for AI-enabled CPS, and that trustworthy operation instead requires lifecycle-aware, provenance- and physics-grounded system design that treats AI agents and MCP as integral components of the CPS control fabric. The primary contributions of this paper are summarized below. 1) A systematic framework, SENTINEL, that provides a structured methodology to select and combine security mechanisms tailored to specific CPS deployments. 2) A taxonomy of deepfake modalities (visual, audio, tex- tual, behavioral) as they apply to CPS, with concrete examples of how each can compromise agents or system functions [55], [259]. 3) A systematic review of deepfake detection techniques in CPS-relevant settings, comparing their strengths, limita- tions, and suitability in constrained environments in real- time [82], [175]. 4) An overview of mitigation and defensive strategies from provenance/authentication, watermarking, multi- modal fusion, to robust training and policy-level mea- sures [259]. 5) An end-to-end CPS case study leveraging ANCHOR- Grid [132] demonstrates how the proposed SENTINEL framework can be operationalized under concrete timing, noise, and safety constraints. 6) A discussion of open challenges and directions, empha- sizing the tension between real-time performance, gener- alization, privacy, and adaptive adversaries [82], [145]. By consolidating insights from recent literature on AI security, generative models, and CPS resilience, we hope this survey will serve as a foundational reference for researchers and system designers seeking to build trustworthy AI agents in CPS. The rest of this paper is structured as follows. Section I provides fundamental background knowledge on CPS, AI agents, MCP protocol, and deepfake attacks. Section I in- troduces the SENTINEL framework. Section IV describes the threat landscape for AI Agents in CPS. Section V focuses on Deepfakes as a CPS security threat. Sections VI and VII discuss the state-of-the-art techniques and strategies for detection, mitigation, and defense. Section VIII demonstrates how the principles surveyed in this work apply in a concrete CPS context. Section IX explores the open challenges and future research directions. And finally, Section X concludes this survey. I. BACKGROUND This section introduces (i) the foundations of cyber–physical systems (CPS) and their security context, (i) AI agents and emerging agent protocols, (i) the MCP for tool and data interoperability, and (iv) deepfake (AI-generated content) tech- nologies that underpin the threat space reviewed in this survey. A. Cyber–Physical Systems (CPS) and Security Context CPS tightly couples computation, networking, and physi- cal processes across industrial control and smart grids, au- tonomous vehicles, and robotics. Classic texts and frameworks emphasize that timing, concurrency, and physical dynamics make CPS distinct from pure IT systems and demand domain- aware assurance across safety, security, reliability, and re- silience [114], [184], [243]. The NIST CPS Framework codi- fies this view with the concept of trustworthiness, integrating security, privacy, safety, reliability, and resilience concerns into system analysis and design [114], [243]. Fig. 1. Cyber-Physical Systems (CPS) and Security Context. Historically, CPS security has focused on threats such as sensor/actuator spoofing, network-borne attacks, and insider misuse. As AI components (e.g., perception, planning, and decision modules) increasingly mediate sensing and actua- tion, the AI–environment interface expands the attack surface: adversaries can manipulate inputs, models, and tool outputs to induce unsafe actions or conceal physical anomalies. Re- cent surveys and applications research highlight the need for runtime monitoring and semantics-aware validation (beyond offline model testing) to cope with distribution shift and adversarial manipulation in operational environments [282]. B. AI Agents and Agent Protocols Agent model:Modern AI agents follow an ob- serve–reason–act loop, often with memory and tool-use. LLM- centric agents augment core reasoning with tool/API calls, retrieval, code execution, and delegation to other agents. This modular, tool-augmented setup accelerates capability, but also introduces trust boundaries at inputs (prompts, retrieved data), tools (code, connectors), and inter-agent communication [213], [259]. Interoperability trends.: To reduce bespoke integrations, the community is converging on agent protocols that stan- dardize how agents discover capabilities, invoke tools, and ex- change context. Recent surveys and industry guidance describe unified schemas and JSON-RPC–style interaction models that enable agents to interact with heterogeneous services while preserving a shared execution context [180], [259]. 1 While 1 Section I-C details MCP, the most widely adopted open specification for this purpose. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20263 Fig. 2. Al Agents and Agent Protocols. these protocols improve reuse and composition, they also amplify the need for identity, authorization, and content- sanitization controls across agent boundaries [130], [337]. Agent-to-Agent Network Protocols: In CPS environments like swarm drones, swarm robotics, smart grids, agents must often coordinate directly with peers rather than just passive tools. The Agent-to-Agent (A2A) protocol has emerged as a standard to replace monolithic orchestration with decen- tralized, discoverable task workflows using JSON-RPC and SSE [275]. However, integrating A2A with MCP moves the system from simple ”glue code” to complex protocol inter- actions that introduce unique risks. Li et al. warn that this fusion creates semantic interoperability gaps and compounded security surfaces. For instance, a compromised agent can exploit the vertical MCP tool access of a peer via horizon- tal A2A delegation, effectively bypassing intended isolation boundaries [192]. C. Model Context Protocol (MCP) MCP is an open specification that standardizes how LLM applications (clients) interact with external tools and data via MCP servers, using structured messages (e.g., JSON-RPC) and explicit capability descriptions [223], [264]. The vendor- neutral design of MCP and the flexibility of the transport (e.g., stdio, HTTP/SSE) have accelerated adoption in IDEs, desktop assistants, and agent frameworks [180], [284]. Fig. 3. Model Context Protocol (MCP) System. Security considerations.: Security analyses identify key risks at the MCP boundary: context poisoning (malicious tool outputs influencing model behavior), supply-chain expo- sure (unvetted servers/connectors), and over-privileged creden- tials [130], [337]. Multiple 2025 disclosures underscore these concerns; for example, researchers reported a malicious MCP server update that surreptitiously BCC’d all processed email to an attacker infrastructure [358]. In CPS, two additional risks emerge: (1) Availability, where protocol latency violates real- time timing constraints; and (2) Social Engineering, where deepfakes defeat human-in-the-loop controls by generating persuasive proof to authorize unsafe actuations. Best-practice guidance recommends least-privilege scopes, ephemeral cre- dentials, signed artifacts, and output sanitization [130], [172], [223]. D. Deepfake (AI-Generated Content) Technologies AI-generated content spans visual (image/video), audio (speech/voice), textual, and behavioral modalities. In the con- text of CPS, deepfakes represent a next-generation sensor spoofing threat. Unlike traditional replay attacks, generative models can synthesize novel, physically consistent sensor inputs like thermal readings, LiDAR point clouds, or oper- ator voice commands that bypass standard anomaly detection systems. Fig. 4. Deepfake (Al-Generated Content) Technologies. GANs and variants: Generative Adversarial Networks (GANs) established the modern paradigm for realistic syn- thesis, powering face swap/reenactment, attribute editing, and style transfer [109]. Contemporary surveys catalog state-of- the-art GAN-based deepfake pipelines, datasets, and detectors for faces and talking heads [76], [258]. Diffusion models: Diffusion models (DMs) have eclipsed GANs on many image/video tasks, yielding high-fidelity and diverse samples via iterative denoising. Surveys from the late 2024 to 2025 show rapid migration of deepfake generation and detection research toward diffusion and hybrid methods (GAN+DM), and document generalization gaps where detec- tors trained on older GAN artifacts underperform on DM fakes [76], [258]. Neural audio generation and voice cloning: Neural TTS/VC pipelines combine speaker embeddings, sequence-to- sequence acoustic models, and high-quality vocoders; more recent LLM-guided audio generation with neural codecs fur- ther narrows the gap to human speech. Comprehensive surveys and challenge series (ASVspoof, ADD) chronicle spoofing attacks and countermeasures, noting that modern liveness and challenge–response protocols complement spectral artifact analysis for robust defense [191]. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20264 Text and behavioral emulation: LLMs can produce im- personated instructions/logs and synthesize plausible telemetry or user behaviors to evade anomaly detectors. Multi-agent attack studies reveal how untrusted content and peer messages can induce unsafe tool calls and system-level compromise when trust boundaries are weak [213]. Provenance and environmental fingerprints: To shift from artifact spotting to source authentication, the industry has standardized cryptographic provenance for media (C2PA Content Credentials) that bind origin/edit history to assets [75], [270]. In parallel, environmental fingerprints such as Electric Network Frequency (ENF) traces embedded in audio/video offer physics-rooted corroboration of time/place, aiding mul- timedia authentication and tamper detection [131]. These approaches are increasingly relevant to CPS, where signals bridge physical and digital realms. E. Summary Figure 5 illustrates the boundaries of trust and attack surfaces in MCP-enabled AI agents operating within CPS. MCP-mediated tool invocation and inter-agent communica- tion introduce new cross-domain trust assumptions, allowing untrusted tool descriptions, server outputs, and synthetic envi- ronmental data to influence agent perception, reasoning, mem- ory, and physical actuation. These blurred boundaries allow deepfake-driven attacks, context poisoning, and supply-chain exploits that bypass traditional CPS security models. CPS requires trustworthiness across sensing, reasoning, and actua- tion. Protocol-enabled agent ecosystems (e.g., MCP) expand capability but raise new identity, authorization, and supply- chain risks. Meanwhile, rapid progress in generative models fuels deepfakes that erode assumptions about input veracity. The remainder of this survey builds on this background to (i) systematize threats at the AI-environment interface and (i) review detection and mitigation strategies tailored to CPS constraints. Fig. 5. MCP threat-boundaries. I. SENTINEL FRAMEWORK: SYSTEMATIC EVALUATION AND THREAT-INFORMED DEFENSE SELECTION Before examining specific threats to AI agents in CPS, we introduce the a Systematic Evaluation and Threat-Informed NEtwork defense seLection (SENTINEL) framework. This framework addresses a critical gap in the literature: while numerous surveys catalog security mechanisms for AI sys- tems or enumerate threats to CPS, few provide a systematic methodology for matching defense strategies to specific de- ployment contexts. The SENTINEL framework operational- izes the selection of security mechanisms by integrating threat modeling, resource constraints, and operational requirements into a unified decision process. Figure 6 outlines the workflow of the SENTINEL frame- work. It recognizes that MCP-enabled AI agents in CPS operate under fundamentally different constraints than tradi- tional IT systems or standalone AI applications. These systems must balance security with safety-critical timing requirements, operate with limited computational resources at edge devices, maintain availability for physical process control, and preserve privacy while enabling necessary monitoring. Additionally, the distributed nature of MCP architectures, where agents interact with external tool servers and shared context repositories, creates unique trust boundaries that traditional security frame- works do not adequately address. SENTINEL provides a six- phase methodology that guides system designers from initial threat assessment through defense deployment and continuous adaptation. Each phase produces concrete outputs that inform subsequent phases, creating a traceable decision pathway from system requirements to specific security mechanism configu- rations. A. Phase 1: System Characterization and Requirement Anal- ysis The SENTINEL framework begins with a comprehensive characterization of the target CPS deployment along four critical dimensions. The operational dimension captures timing requirements including maximum tolerable detection latency, required system response time to threats, acceptable downtime for security updates, and recovery time objectives following security incidents. These temporal constraints fundamentally determine which detection and mitigation approaches remain viable. For instance, a smart grid protection relay operating on a 4-ms cycle time cannot employ deepfake detection, which requires 15-20-second ENF signal accumulation in the control loop, though such techniques may authenticate operator commands or validate sensor provisioning. The resource dimension profiles available computational capacity at different network tiers, such as edge devices, fog nodes, and cloud infrastructure, along with memory con- straints, network bandwidth, and latency characteristics, en- ergy budgets for battery-powered sensors, and storage capacity for security logs and forensic data. This profile determines whether detection mechanisms can operate locally or must rely on offloaded processing, directly impacting both latency and privacy characteristics. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20265 Fig. 6. The SENTINEL framework’s six-phase methodology with feedback loops enabling continuous adaptation. The trust dimension maps agent-environment interaction surfaces, including external data sources that agents query through MCP servers, tool invocations that trigger physical actuation, inter-agent communication channels, human oper- ator interfaces, and third-party service dependencies. Each interaction surface represents a potential attack vector requir- ing appropriate authentication, authorization, and validation mechanisms. The MCP protocol architecture introduces par- ticular complexity here, as agents may invoke tools provided by external servers whose security posture the agent operator cannot directly control. The criticality dimension assesses potential consequences of security failures through impact on physical safety, including injury potential and physical damage scenarios, operational continuity requirements, including recovery complexity and cascade failure risks, such as data sensitivity considerations encompassing personal information exposure and intellectual property protection, and regulatory compliance obligations. This dimension establishes the required security level and acceptable risk thresholds that guide the selection of defense mechanisms. B. Phase 2: Threat Profiling and Attack Surface Analysis Building on system characterization, Phase 2 conducts structured threat profiling using a three-dimensional taxonomy that examines interaction surface (user-agent (U-A), agent- agent (A-A), agent-environment (A-E)), attack modality (vi- sual deepfakes, audio synthesis, text generation, behavioral emulation, sensor spoofing), and adversary capability (oppor- tunistic attackers using public tools, sophisticated adversaries with custom generative models, insider threats with system knowledge, supply chain compromises affecting MCP servers or tool providers). This taxonomy enables precise threat mod- eling that moves beyond generic security checklists to identify threats actually relevant to the specific deployment context. For each identified threat, the framework requires the es- timation of four key parameters. Attack likelihood reflects adversary motivation given the system’s value or strategic importance, accessibility of attack vectors considering network exposure and authentication requirements, and attacker techni- cal sophistication required. The impact of the attack quantifies the potential physical safety consequences, the magnitude and duration of operational disruption, financial losses due to downtime or damage, and reputational or regulatory conse- quences. Detection difficulty assesses the subtlety of attack artifacts and whether they fall below sensor noise floors; the availability of training data for adversarial examples; the com- putational complexity of detection algorithms; and whether attacks can adapt to known detection mechanisms. Finally, mitigation costs consider the computational overhead imposed by defense mechanisms, the latency introduced into critical control loops, the implementation complexity and required expertise, and the ongoing operational burden of security monitoring and updates. This structured threat profiling produces a prioritized threat register that identifies high-priority threats, those with high likelihood and severe impact that drive primary defense re- quirements, medium-priority threats requiring monitoring but potentially acceptable through risk acceptance if mitigation costs exceed risk reduction value, and low-priority threats where standard security hygiene provides adequate protection without specialized deepfake defenses. This prioritization pre- vents the common pitfall of attempting to defend against all theoretically possible attacks regardless of their actual rele- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20266 vance to the deployment context, enabling resource-efficient security designs. C. Phase 3: Constraint-Aware Defense Mechanism Selection Phase 3 systematically evaluates candidate detection and mitigation techniques against a multi-dimensional fitness func- tion that captures the tradeoffs inherent in CPS security. The framework maintains a comprehensive catalog of security mechanisms, drawn from the taxonomy developed in Sections IV through VI of this survey, with each mechanism character- ized along six evaluation dimensions. The detection effectiveness dimension captures mechanism accuracy through true positive and false positive rates for relevant attack types, robustness to adversarial adaptation and evasion attempts, generalization capability to novel deepfake generators not seen during training, and coverage breadth across different attack modalities. Computational requirements specify processing resources that include CPU, GPU, and memory footprints, latency from input acquisition to threat verdict delivery, scalability characteristics as system size or attack volume increases, and whether processing can occur at edge devices or requires centralized computation. The inte- gration dimension assesses the complexity of the deployment, including the modifications needed to existing systems, the dependency on specialized hardware such as trusted execution environments or GPS receivers for timestamping, the compati- bility with the existing security infrastructure, and the training data requirements for ML-based approaches. The operational dimension addresses maintenance burden, including model retraining frequency, signature database up- date cadence, false positive handling procedures, and forensic logging and audit trail requirements. Additionally, the privacy dimension evaluates data exposure, including whether raw sensor streams must be transmitted to detection services, whether cryptographic protections preserve privacy, compli- ance with data protection regulations, and user transparency requirements. The cost dimension captures both capital expen- ditures for new hardware or software licenses and operational expenses, including computational resources consumed, secu- rity operations personnel requirements, and incident response processes. For each high-priority threat identified in Phase 2, the framework generates a candidate defense set comprising mech- anisms that satisfy hard constraints, particularly real-time latency requirements and computational resource limits, and ranks them based on fitness scores that weight the evaluation dimensions according to the system priorities established in Phase 1. This produces a shortlist of viable defense mecha- nisms for each threat rather than presuming one-size-fits-all security solutions. D. Phase 4: Defense-in-Depth Architecture Design Recognizing that no single security mechanism provides complete protection, Phase 4 constructs a layered defense architecture that integrates complementary techniques to ad- dress the threat landscape comprehensively. Figure 7 illus- trates the four-tier defense-in-depth architecture that Phase 4 constructs for MCP-enabled cyber-physical systems, showing how complementary security mechanisms combine to provide comprehensive protection against deepfake threats. The perimeter tier implements proactive defenses before threats reach agent decision-making processes, including input validation and sanitization at MCP server boundaries, cryp- tographic authentication of content provenance using C2PA or similar standards, and reputation-based filtering of external data sources and tool providers. These mechanisms reduce the attack surface by rejecting obviously malicious inputs before they consume detection resources or influence agent behavior. The detection tier deploys deepfake detection mechanisms matched to the threat profile and positioned according to resource availability. Lightweight heuristics run on edge de- vices to detect crude attacks with minimal latency, ensemble classifiers run on fog nodes to provide more sophisticated analysis with acceptable delay, and heavyweight deep learning models execute on cloud infrastructure for forensic analysis of suspicious events. The multi-tier detection strategy optimizes the tradeoff between detection latency and accuracy by fil- tering most attacks at lower tiers while reserving expensive analysis for edge cases. The response tier specifies graduated response policies trig- gered by different threat levels, ranging from alerting human operators for manual verification through automated contain- ment measures such as revoking tool invocation permissions to emergency failsafe actions, including disconnecting compro- mised agents from physical actuators. The response policies must account for CPS safety requirements, ensuring that security responses do not themselves create unsafe physical states, and preserve evidence for forensic investigation. The adaptation tier implements continuous improvement mechanisms, including performance monitoring that tracks de- tection accuracy and false positive rates in operational deploy- ment, threat intelligence integration that updates threat models and detection signatures, adversarial training that incorporates newly discovered attack techniques, and periodic security au- dits that verify defense effectiveness. This tier recognizes that deepfake generation techniques evolve continuously, requiring defense mechanisms that adapt rather than remain static. For MCP-specific deployments, the SENTINEL framework emphasizes enforcing trust boundaries at protocol interaction points. MCP servers that provide tools to agents require strong authentication and authorization mechanisms to verify server identity and restrict tool invocations based on agent privileges. Context sharing among agents via MCP prompts demands vali- dation to prevent context-poisoning attacks, in which malicious agents inject deepfake content into shared memory. Supply chain security for MCP tool providers necessitates vetting processes and continuous monitoring, given that compromised tools can bypass agent-level defenses entirely. E. Phase 5: Validation and Deployment Planning Phase 5 translates the defense architecture into deployable configurations through systematic validation. The framework requires simulation-based evaluation using representative at- tack scenarios from the threat registry, with performance JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20267 Fig. 7. The four-tier defense-in-depth architecture in Phase 4. measured against the metrics established in Phase 1. This evaluation should include both isolated mechanism testing, verifying that individual detection techniques achieve specified accuracy and latency targets, and integrated system testing that validates the complete defense-in-depth architecture under realistic conditions, including normal operational workload, concurrent multi-vector attacks, and resource-degradation sce- narios. The framework specifies three validation methodologies ap- propriate for different CPS contexts. Laboratory testbeds pro- vide controlled environments for detailed performance charac- terization without risking operational systems, enabling com- prehensive testing against both documented attacks and novel adversarial examples generated through red team exercises. Digital twin simulation leverages physics-based models of CPS processes to evaluate security mechanisms under realistic operational conditions while maintaining safety, particularly valuable for testing response tier policies that could trigger unsafe states if deployed prematurely in production systems. Pilot deployments gradually introduce security mechanisms into operational systems with extensive monitoring, begin- ning with non-critical subsystems before expanding to safety- critical components once confidence in defense effectiveness has been established. Deployment planning must address staged rollout timelines that minimize operational disruption, fallback procedures that allow rapid security mechanism deactivation if unanticipated operational issues arise, personnel training requirements for security operations staff and incident responders, and coordi- nation with existing security infrastructure, including security information and event management systems and incident re- sponse playbooks. F. Phase 6: Continuous Monitoring and Adaptive Defense The final phase recognizes that security is not a one-time deployment but an ongoing process that requires continuous monitoring and adaptation. The framework specifies metrics for operational security monitoring, including attack detection and false-positive rates relative to baseline expectations, se- curity mechanism resource consumption relative to budgets, degradation in system performance caused by security over- head, and coverage gaps where new attack techniques evade deployed defenses. Trigger conditions for defense mechanism updates include detection accuracy falling below acceptable thresholds, indi- cating adversarial adaptation, the emergence of new threat intelligence about attack techniques or vulnerable components, changes to system configuration or operational requirements that alter threat landscape or constraint profiles, and security incidents that reveal gaps in defense coverage. The framework provides a structured process for evaluating and deploying updates while maintaining operational continuity. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20268 Adaptation mechanisms range from tuning existing security parameters, such as adjusting detection thresholds or updating signature databases, through deploying additional defense lay- ers to cover newly identified gaps, to architectural redesign when fundamental assumptions about the threat model or system requirements change substantially. Each adaptation level requires increasingly rigorous validation before oper- ational deployment, with parameter tuning potentially auto- mated through machine learning, while architectural changes require comprehensive testing. G. Framework Application Methodology The SENTINEL framework provides both prescriptive guid- ance for greenfield deployments and diagnostic capabilities for evaluating existing security architectures. For new MCP- enabled CPS deployments, designers proceed sequentially through the six phases, with each phase’s outputs documented and reviewed before proceeding. This disciplined approach ensures security requirements drive architecture decisions rather than security being retrofitted after system design has constrained options. For operational systems, the framework enables gap anal- ysis by systematically evaluating whether existing security mechanisms adequately address the characterized threat land- scape, given actual system constraints and requirements. This diagnostic application often reveals either over-provisioned security, expensive mechanisms defending against low-priority threats, or critical gaps where high-priority threats lack ade- quate detection and mitigation. The framework thus guides both security investment optimization and risk remediation prioritization. Throughout Section IV, we apply the concepts of the SEN- TINEL framework to analyze specific threat categories at the user-agent, agent-agent, and agent-environment interfaces. For each threat class, we identify representative attack techniques, assess their characteristics along the framework’s evaluation dimensions, and highlight security mechanisms most appropri- ate for different CPS deployment contexts. Section V then pro- vides a detailed examination of deepfake detection techniques organized by modality, and Section VI presents mitigation and defense strategies. Together with the SENTINEL framework introduced here, these sections equip system designers with both comprehensive security mechanism knowledge and sys- tematic methodology for applying that knowledge to specific deployments. The subsequent subsections examine threats at each inter- action surface in detail, but readers should interpret these threat discussions through the SENTINEL framework lens: understanding which threats apply to their specific context, evaluating candidate defenses against their unique constraint profile, and designing defense-in-depth architectures tailored to their requirements rather than adopting generic security solutions. IV. THREAT LANDSCAPE FOR AI AGENTS IN CPS This section uses the first two phases of the SENTINEL framework to systematically characterize the threat landscape. Phase 1 analyzes the threats along the timing, trust, resource, and criticality dimensions of CPS deployments. Phase 2 clas- sifies emerging threats along three axes: interaction surface (U-A, A-A, A-E), attack modality, and adversary capability. This approach expands abstract vulnerabilities into context- dependent risks, evaluating CPS-specific requirements like detection difficulty and mitigation costs. A. User–Agent (U–A) Interface Analysis The U-A interface represents the primary interaction point where human operators issue commands and receive feedback. While this transparency facilitates operational efficiency and reduces cognitive load, it introduces a critical vulnerability: the ”observe–reason–act” loop becomes susceptible to linguistic manipulation. The integration of Large Language Models (LLMs) allows agents to ingest heterogeneous data streams [79], but simultaneously exposes the physical system to risks ranging from trust decay to expanded attack surfaces via untrusted external documents. 1) Prompt Injection: Prompt injection involves smuggling adversarial instructions into inputs the agent trusts to override its intended behavior [204], [207]. In CPS, this escalates from text misalignment to unsafe physical actuation because the agent directly mediates sensing and control. Trust is inherently challenged as agents process both direct user prompts and external context from documents or tool outputs; [110] demon- strate how this context can be intentionally poisoned, while [26] show that retrieval mechanisms can degrade safety align- ment even when the context is safe. This risk is amplified in modern stacks via malicious tool responses, a threat formalized in the Model Context Protocol (MCP) [136] and demonstrated in end-to-end autonomous agent benchmarks [94]. 2) Social Engineering: Social engineering exploits the psy- chological trust between humans and AI agents. Adversaries use high-fidelity synthesis, such as cloned voices, fake proce- dures, or images, to induce unsafe decisions [248], [341]. In in- dustrial settings, operators may unknowingly facilitate attacks through conversational interfaces. Detection is particularly difficult as modern ”jailbreaks” can be multi-modal, employ- ing ”many-shot” strategies or typographic visual prompts to bypass safety filters [28], [108], [278], [305]. Current research suggests mitigation requires robust detection models, feature fusion, and rigorous input filtering, though CPS contexts may necessitate additional out-of-band corroboration [164], [205]. 3) Confused Deputy: Security risks in agentic systems often stem from the agent acting as an unwitting accomplice to a malicious user. In the Model Context Protocol (MCP) ecosystem [117], tools are typically pre-authorized, allowing attackers to bypass privilege checks and manipulate the agent into executing unauthorized operations, such as reading sensi- tive files or tampering with configurations. Benchmarks [354] indicate that while agents must handle complex dependency chains and fuzzy instructions in multi-tool environments, this complexity expands the attack surface. For example, agents struggle to distinguish between external data and executable instructions, leading to potential command injections [117]. In the manufacturing context [266], while agents enable decen- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 20269 Fig. 8. Threat Landscape for AI Agents in CPS: USER-AGENT (U-A) tralization, robust cybersecurity remains a core requirement to prevent information theft and service disruption. 4) Comparative Analysis: Table I details the system con- straints for U-A interactions. We found that timing is the most stringent constraint and the injection detection must occur within the millisecond-scale latency budget of the controller. TABLE I U–A INTERACTION SYSTEM CHARACTERISTICS DimensionPrompt Injec- tion Social Engineering Confused Deputy TimingMillisecond Detection Real-time Interaction [248] Request-time Check [117] ResourceLimitedEdge Compute High Bandwidth Media[164], [341] LowCPU Overhead TrustUntrusted External Data [26], [136] Deceptive HumanTrust [248] Privileged Misattribution [117] CriticalityPhysical Safety Hazard Operational Downtime Unauthorized Actuation [117] 5) Threat Registry: Table I profiles the emerging threats. The registry indicates that while Prompt Injection has the high- est likelihood due to low barriers to entry, Social Engineering presents the greatest detection challenge due to multi-modal artifacts. 6) Open Challenges in U-A Security: Despite ongoing research, several open challenges remain for securing the U- A surface in CPS. The current research landscape highlights several critical hurdles for securing U-A interactions in CPS: • Probabilistic Safety Guarantees: Traditional formal methods struggle with the stochastic nature of LLMs; new TABLE I U–A THREAT PROFILING REGISTRY DimensionPrompt Injec- tion Social Engineering Confused Deputy ModalityText / Context [136], [207] Multi-modal [108],[278], [305] API / Protocol [117] CapabilityLow Knowledge [94], [207] Medium [28], [248] System- specific [117] LikelihoodHigh[26], [207] Elevated [28]High [117] ImpactSevere[26], [110], [136] Significant [248] Critical [117] Detection Difficulty High[136], [207] High[108], [248] High [117] Mitigation Costs Medium [79]Substantial [278], [305] High frameworks like Pro2guard are exploring probabilistic model checking to enforce safety at runtime [344]. • Verified Code Generation: Ensuring that the tool calls and scripts generated by agents adhere to strict safety specifications remains difficult, as seen in recent efforts toward verified code generation frameworks like Veri- Guard [221] • Adaptive Robustness: As adversaries adapt to current filters, there is a need for defense mechanisms that can survive an ”adaptive arms race” by dynamically updating their detection logic [79]. • Multi-Agent Adversarial Dynamics: In environments where multiple agents interact, compromised proxies can lead to cascading failures that are poorly understood in existing security literature [375]. • Sandboxing Unverified Controllers: Creating secure execution environments that can isolate and monitor JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202610 Fig. 9. Threat Landscape: Al Agent-Agent (A-A) Interactions in CPS LLM-based controllers without disrupting the real-time requirements of the physical system [392]. B. Agent–Agent (A–A) The Agent–Agent (A–A) interaction surface represents the decentralized communication layer where autonomous enti- ties collaborate to achieve system-wide goals. This enables decentralized resilience (eliminating single points of failure) and operational scalability for massive fleets [154], [377]. However, this introduces a ”trust vacuum” where faults can propagate at machine speed [213], [345]. The absence of human-in-the-loop oversight means malicious artifacts can cause physical damage before intervention is possible, and implementing defenses often incurs performance overheads unacceptable for real-time control [4]. 1) Trust Exploitation: Trust exploitation occurs when an adversary manipulates the implicit reliance agents place on peer-provided artifacts, such as plans, tool outputs, or rep- utation scores [213]. In protocol agent meshes, even minor perturbations in a single agent’s summary or capability adver- tisement can cascade into unsafe physical actions by peers that accept those artifacts as authoritative, particularly in embodied contexts like autonomous driving or robotics [154], [345]. 2) Identity Spoofing: Identity spoofing leverages the im- personation of legitimate peers to subvert consensus or gain unauthorized access to shared resources [320], [330]. In CPS- adjacent environments like vehicular networks (VANET) and distributed learning systems, Sybil attacks allow an attacker to mint multiple virtual identities to sway voting/consensus and deform situational awareness [50], [97], [351]. Modern agent frameworks remain vulnerable to replay attacks and credential theft, where compromised identifiers or Verifiable Credentials (VCs) allow adversaries to mimic authorized entities across the ecosystem [4], [217]. 3) Collusion in Multi-Agent Systems: Collusion in multi- agent systems involves coordinated strategies among sub- verted or inherently malicious agents to manipulate outcomes across economic, computational, and physical domains. In market-based CPS, this manifests as tacit collusion where reinforcement-learning agents autonomously coordinate on supra-competitive prices or manipulate auction parameters [73], [90], [113], [376]. In distributed intelligence, collusion threatens model integrity through mechanisms like distributed backdoors in federated learning or consensus manipulation in multi-LLM agentic systems [200], [206]. Furthermore, agents may coordinate at the physical layer to compromise system state, ranging from eavesdroppers optimizing signal interception in UAV networks [189] to adversarial attacks on cooperative perception and control set-points in autonomous vehicles and microgrids [15], [190], [314]. 4) Comparative Analysis : Table I highlights the shift from isolated vulnerabilities to system-level impacts. While Identity Spoofing is a ”gatekeeper” threat (Static Timing), Collusion represents a long-term erosion of system integrity (Emergent Timing). TABLE I A–A INTERACTION SYSTEM CHARACTERISTICS DimensionTrust Exploitation Identity Spoofing Collusionin MAS TimingReal-time [377] Static / Setup [320] Emergent/ Long-term [113] ResourceHighdemand /Constrained [154] Low(Sybil) [50] High (Distributed) [189], [200] TrustImplicitPeer [213] Identification [330] Coordinated Bias [200] CriticalityHigh / Safety- Critical [345] High / Auth- Failure [351] Systemic Sta- bility [90] JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202611 5) Threat Profiling: Table IV categorizes threats by defense effort. Trust Exploitation is highly likely in agentic workflows due to over-reliance on peer summaries, whereas Collusion remains a lower-likelihood but high-impact threat requiring expensive mechanism-design mitigation. TABLE IV A–A THREAT PROFILING REGISTRY DimensionTrust Exploitation Identity Spoofing Collusionin MAS ModalityTaint Injection [213] Masquerade [330] Strategic Sync [206] CapabilityLow [213]Low to Moder- ate [330] High [190] LikelihoodMed–High [345] High [50]Low–Med [113] ImpactHigh [154]High [4]High[15], [190] Detection Difficulty High [213]Moderate [217] High [90] Mitigation Costs Medium [345]High(Invest- ment/Compute) [4], [330] High [376] 6) Open Challenges in A-A security: The research com- munity faces several unresolved hurdles in securing the A–A surface, requiring a multidisciplinary approach: • Byzantine Robust LLM Coordination: Standard con- sensus algorithms must be adapted for LLM-based agents where ”faults” are often semantic (e.g., hallucinations or prompt-induced bias) rather than traditional bit-flips [206], [275], [345]. • Scalable Identity Management: Implementing W3C Decentralized Identifiers (DIDs) and Verifiable Creden- tials (VCs) across transient IoT agents must be optimized to prevent Sybil attacks without introducing prohibitive latency or power drain [217], [351]. • Detecting Tacit/Algorithmic Collusion: Identifying co- ordinated malicious intent in agents that do not use explicit communication channels remains an unsolved statistical challenge, requiring new methods for policy correlation analysis [90], [113]. • Cross-domain Taint Tracking: There is a lack of stan- dardized protocols for maintaining a ”provenance chain” of data as it passes through multiple agents with different trust boundaries and tool permissions [4], [213]. • Verification of Emergent Behavior: Developing formal verification methods for the joint action spaces of founda- tion model-powered agents to ensure they do not violate physical safety invariants [154], [190]. C. Agent–Environment (A–E) The A–E interface represents the boundary where the AI agent perceives and acts upon its surrounding physical and digital context. Threats at this layer target the agent’s interpre- tation of reality, aiming to induce ”hallucinated” environmental states [85]. Current defenses focus on narrowing the agent’s reach via least privilege scoping [351], sensor fusion invariants [152], and requiring human-in-the-loop interlocks for high- risk actions [274]. However, verifying natural language intent against physical laws remains difficult, and countermeasures often introduce unacceptable latency [234], [369]. 1) Indirect Prompt Injection: Indirect prompt injections occur when an agent retrieves data containing embedded malicious instructions from its environment (e.g., web pages, emails, logs, or sensor inputs), causing the agent to prior- itize these ”smuggled” commands over the user’s original intent [49], [372]. The agent effectively assumes retrieved environmental data is passive information rather than active logic, allowing attackers with low capability (no direct system access) to trigger high-criticality tool invocations ranging from database destruction [257] to unsafe physical actuations [351]. 2) Adversarial sensor attacks: These attacks involve the physical or cyber manipulation of sensor signals to mislead the CPS. While attacks on modalities like LiDAR, vision, and audio can cause perception modules to ”hallucinate” obstacles or erase threats [210], sophisticated strategies can also decouple sensor outputs from control inputs to execute stealthy, targeted covert attacks that remain mathematically undetectable to standard feedback loops [222]. These threats are exacerbated by the edge-processing constraints of CPS; on-device models often lack the compute power for complex cross-sensor verification, making the detection of such per- turbed signals highly difficult [152]. 3) Malicious Tool Interactions: Agents often interact with the environment via tools (connectors, MCP servers, or APIs). A malicious tool interaction occurs when an external service returns tampered data or exploits the agent’s permissions to execute unauthorized side effects, effectively weaponizing the agent’s blind obedience to tool outputs [117], [192]. This high- lights supply chain risks where tool capabilities are often pre- authorized for high-impact actions, requiring robust sandbox isolation, schema validation, and behavioral drift monitoring to prevent physical or operational harm [117], [234]. 4) Comparative Analysis: Table V emphasizes the diver- gence in operational needs. While sensor attacks require sub- second detection within the edge environment, tool interac- tions involve complex supply-chain trust models managed across cloud-edge boundaries. TABLE V A–E INTERACTION SYSTEM CHARACTERISTICS DimensionIndirect Prompt Injection Adversarial Sensor Attack Malicious Tool Interaction TimingReal-time Inference [351] Ultra-low La- tency [210] Asynchronous / Event-driven [192] ResourceHigh Memory / Context [372] Edge- constrained [152] Network / API Bandwidth [234] TrustUntrusted Arti- facts [351] External Signals [210] Third-party APIs [117] CriticalityHighData/ Policy [49] Life-safety/ Damage [152], [210] Supply-chain Risk [234] 5) Threat Registry: Table VI highlights that while indirect injections are highly likely due to the accessibility of untrusted data, adversarial sensor attacks, though harder to execute, rep- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202612 Fig. 10. Threat Landscape: Al Agent-Environment (A-E) Interactions in CPS. resent a critical safety risk with significantly higher mitigation costs. TABLE VI A–E THREAT PROFILING REGISTRY DimensionIndirect Prompt Injection Adversarial Sensor Attack Malicious Tool Interaction ModalitySemantic / Text [257] Physical / Cy- ber [210] Protocol / Data [192] CapabilityLow [372]Medium [222]Medium [117] LikelihoodHigh [257]Medium [152]High [117] ImpactHigh [49]Critical [210]High [234] Detection Difficulty High [274]High [152]High[117], [192] Mitigation Costs Medium [257]High [210]Medium [234] 6) Open challenges: Securing the interaction between an agent and its environment requires moving beyond digital-only security toward physical resilience-by-design [192], [210]. • Zero-trust sensing protocols: Developing standards where every sensor signal is treated as untrusted until val- idated through multi-modal fusion or challenge-response [152]. • Context Aware Policy Enforcement: Creating ”safety- shells” that can interpret the semantic intent of an agent’s plan and block it if it violates physical state invariants [142]. • Autonomous Recovery Strategies: Designing agents capable of identifying when their inputs are compromised and transitioning to a ”safe-state” without human inter- vention [210]. • Cross-Framework Protocol Alignment: Unifying secu- rity standards (like MCP and A2A) to ensure consistent least-privilege enforcement across heterogeneous agent- tool stacks [275], [369]. D. Real-world incidents and case studies 1) User-Agent Surface incidents: This category examines the Trust Dimension between human operators and AI agents. A primary example is the deepfake heist where attackers used multi-persona video deepfakes to deceive staff into authorizing a $25 million transfer, as detailed in [116], [133]. SENTINEL classifies this as a high-impact behavioral emulation attack that exploits the U-A surface by bypassing traditional human-in- the-loop verification. The accessibility of generative AI has empowered oppor- tunistic attackers, such as adolescents using AI for non- consensual image generation, to breach the Privacy Dimension of public spaces [116]. Such incidents force a reassessment of standard security hygiene, as the barrier to entry for high- fidelity deception has dropped significantly [297]. 2) Agent-Agent Surface incidents: The Model Context Protocol (MCP) introduces unique trust boundaries. Recent parasitic toolchain attacks involve malicious MCP servers that silently exfiltrate user data [389]. Multi-agent systems have furthermore been shown to exhibit Inter-Agent Trust Exploitation, where agents execute malicious commands from peer agents that they would otherwise reject from humans [213]. These incidents illustrate supply chain compromise and the failure of the Trust Dimension in automated orchestration. 3) Agent-Environment Surface incidents: Attacks targeting the A-E interface are critical in sectors like autonomous driving and physical infrastructure. Research highlights how physical environment threats, such as sensor spoofing and signal interference, can lead to denial of service or unsafe physical actuations [82]. Deepfakes can indirectly destabilize CPS by eroding public trust or inciting panic. Notable examples include the deepfake of Ukrainian President Zelenskyy or the AI-generated image of an explosion at the Pentagon, which impacted financial markets [236]. In the SENTINEL framework, these are High- Likelihood threats that require monitoring of the Critical- ity Dimension. Additionally, the proliferation of autonomous swarms introduces physical security threats that require dis- tinct points of control to prevent malicious misuse [262]. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202613 V. DEEPFAKES AS A CPS SECURITY THREAT Deepfakes, as AI-generated synthetic media, pose sig- nificant security threats to CPS by exploiting the agent- environment interface, particularly in environments utilizing the MCP for tool interoperability. In CPS, visual deepfakes can spoof surveillance feeds or deceive autonomous vehicles by altering perceptual inputs, leading to unsafe actuations or misinformed decisions, as highlighted in recent analyses of generative AI vulnerabilities [76], [258]. Audio deepfakes enable bypasses of voice authentication and social engineering attacks, in which cloned voices impersonate operators to issue malicious commands via MCP-integrated tools, exacerbating context-poisoning risks in agent workflows [82]. Textual and behavioral deepfakes further compound these issues by gen- erating fake instructions or sensor anomalies that mimic legit- imate data streams, undermining system integrity in critical sectors like smart grids and industrial automation. Within the MCP context, these threats are amplified by supply-chain exposures and unvetted server interactions, where malicious outputs from deepfake-generating tools can propagate across networked agents, as evidenced by studies of protocol vulner- abilities [136]. The integration of MCP into CPS heightens the privacy implications of deepfakes, including identity theft and the creation of non-consensual media, which erodes trust in human-AI collaborations. For instance, deepfakes can exploit MCP’s open specification to inject forged content into data exchanges, facilitating ”liar’s dividend” scenarios where gen- uine information is discredited amid synthetic misinformation [280], [336]. Recent scientific findings underscore the need for robust defenses, such as provenance tracking, to counter these threats in real-time CPS operations [191], [213]. Behavioral deepfakes, in particular, pose stealthy risks by emulating anomalies to evade detection in multi-agent systems, poten- tially leading to physical infrastructure compromises. Overall, the convergence of deepfake technologies with MCP-enabled ecosystems demands interdisciplinary approaches to mitigate evolving adversarial manipulations in CPS. The deepfake modalities surveyed in this section demon- strate that high-fidelity synthetic content can be internally consistent and physically plausible, rendering artifact-centric detection insufficient in many CPS deployments. SENTINEL Phase 3 emphasizes that feasibility constraints, latency, false- positive tolerance, and safety impact, must act as a hard filter on candidate defenses, ruling out approaches that cannot operate within real-time CPS control and monitoring loops. A. Visual Deepfakes Spoofing surveillance such as visual deepfakes undermine camera-based monitoring by injecting forged video streams or replaying synthetically manipulated faces and scenes that pass casual human review and basic analytics. In MCP-connected stacks, where an agent’s “vision tool” is exposed via an MCP server, deepfaked frames can be ingested as trusted observations, leading downstream agents to issue erroneous alerts, suppress genuine anomalies, or leak context through misclassified events [29], [32]. Surveyed CPS work explicitly notes that deepfake video can spoof surveillance cameras, deceiving both humans and ML perception modules and thus eroding assumptions about input veracity at the AI- environment boundary [107]. From an attack-method perspective, presentation and deep- fake attacks encompass display-mediated spoofing (utilizing high-resolution screens and virtual cameras on RTSP feeds), face reenactment/face-swap to evade watchlist matching, and scene-level synthesis to fabricate individuals or activities. Sys- tematizations of face anti-spoofing and deepfake detection em- phasize that detectors that excel on curated datasets generalize poorly to in-the-wild content and to new generators, necessi- tating spatiotemporal cues, physiological signals (e.g., rPPG), and cross-modal checks to lift robustness in operational CCTV settings [163], [228]. Operational countermeasures combine (i) trusted capture and chain-of-custody (cryptographic signing at the edge, secure logging), (i) content provenance verifi- cation (C2PA Content Credentials) at ingest and before auto- mated responses, and (i) liveness/consistency tests (rPPG, optical-flow/eye-gaze dynamics, illumination challenges) be- fore agents escalate actions via MCP tools [6], [25]. Autonomous vehicle deception such as visual deepfakes and related optical illusions can mislead camera-centric per- ception stacks in autonomous and advanced driver-assistance systems (ADAS). Phantom attacks embed brief, realistic traffic signs or pedestrians into digital billboards or projections, causing detectors to perceive non-existent hazards and trig- ger braking/steering responses; these attacks demonstrate that short, high-fidelity visual artifacts can reliably elicit unsafe behavior without tampering with the vehicle itself [208], [356]. Beyond signage, dynamic adversarial patches displayed on moving surfaces and road-surface patches targeting monocular depth estimation induce systematic misperception of distance or object identity, degrading planning and control even under motion and viewpoint changes [106]. Empirical studies and systematizations further document camera-feed spoofing and object fabrication/erasure against vehicle cameras and trackers, while highlighting that sensor fusion alone is insufficient when attacks are correlated across modalities or over time [39], [228]. Defenses increasingly combine temporal and multi-sensor consistency checks (e.g., vision–LiDAR cross-verification, inertial priors), active chal- lenge–response (structured light or coded illumination to force physically plausible returns), and certifiably robust percep- tion methods designed to bound the impact of patch-level perturbations before actuation; detection pipelines tailored to “phantom” signatures also show promise for camera spoofing cases [310], [327]. Malicious tool interactions with visual deepfakes in MCP ecosystems in MCP-mediated architectures can trigger harmful tool use when agents treat unvetted media as ground truth. The MCP boundary is known to concentrate risks, including context poisoning from tool outputs, over-privileged connectors, and identity fragmentation across servers. Hence, a forged image/video that induces a misclassification or false alert can cascade into privileged MCP tool calls (e.g., opening doors, dispatching assets, disabling interlocks) if policies are not enforced tightly [167]. Industry hardening guidance calls JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202614 Fig. 11. Deepfakes & Security Threats to CPS & AI Agents. for least-privilege scopes, short-lived tokens, output-schema enforcement, human-in-the-loop escalation for high-impact actions, signed artifacts, and comprehensive audit logging to contain such chains; recent analyses of MCP deployments and toolchains detail these vulnerabilities and controls in practice [35]. To raise the bar against visual deepfakes specifically, agents should (i) verify provenance and tamper-evident manifests (C2PA) before trusting media-derived assertions, (i) attach taint/provenance labels to all perceptions and require cor- roboration (e.g., second sensor, independent model) before side-effectful MCP actions, and (i) integrate adversarial ML taxonomies and testing into the tool-approval pipeline so that simulated deepfake scenarios are part of pre-deployment safety cases [292]. TABLE VII COMPARISON OF VISUAL DEEPFAKE DETECTION APPROACHES MethodStrengthsWeaknessesCPS Fit rPPG-basedPhysiological grounding Requires video length Medium Spatial CNNFast inferenceGenerator- specific Low Temporal LSTMMotion artifactsHigh latencyLow Vision-LiDAR Fu- sion Multi-modal robust Hardware costHigh C2PA ProvenanceTamper-evidentAdoption lim- ited High Frequency Analy- sis Generator arti- facts Compression sensitive Medium B. Audio Deepfakes Real-world incidents and case studies illustrate the severe consequences of adversarial manipulations in cyber-physical systems, where AI agents have been exploited to cause op- erational disruptions [147]. One prominent example involves a malicious MCP server that surreptitiously forwarded pro- cessed emails to attacker-controlled infrastructure, exploiting excessive permissions in agent-tool interactions and leading to data exfiltration in industrial environments [33]. Another case documented attackers hijacking AI assistants in enterprise set- tings to steal sensitive data and manipulate workflows, demon- strating how weak trust boundaries in multi-agent systems can enable unauthorized control over physical actuators, such as robotic arms, in manufacturing plants [246]. These breaches highlight CPS’s vulnerability to supply-chain compromises, where unvetted tools amplify risks across both the cyber and physical layers. Further examination reveals emerging patterns in deepfake- driven attacks targeting autonomous systems, such as forged video feeds that deceive surveillance systems in smart grids, resulting in undetected intrusions and potential sabotage of power distribution [233], [238]. In a related incident, voice cloning was used to impersonate operators in vehicle control centers, issuing false commands that altered traffic manage- ment protocols and caused real-time safety hazards [285]. Such multi-modal deceptions not only undermine system integrity but also expose gaps in current detection methods, as attack- ers leverage generative models to evade traditional anomaly checks in dynamic CPS environments. Additionally, case studies emphasize the role of adversarial training failures in exacerbating breaches, including instances in which manipu- lated sensor inputs in healthcare CPS led to erroneous medical device operations, compromising patient safety [166], [261]. These examples underscore the need for integrated defenses, including cross-modal verification to counter sophisticated deepfakes that blend audio, visual, and behavioral elements in coordinated attacks on critical infrastructure [228]. Voice authentication bypass can be bypassed using a variety of spoofing attacks, primarily relying on pre-recorded audio or, more effectively, AI-generated synthetic voices JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202615 (deepfakes) [120]. Advances in neural codec language models and zero- or few-shot voice cloning enable attacker-controlled speech that preserves speaker timbre, prosody, and even acous- tic context from only a few seconds of enrollment audio, sub- stantially lowering the barrier to defeating automatic speaker verification (ASV) and voice biometrics. Systems such as VALL-E and its successors report human-parity or near-parity similarity in zero-shot conditions, which aligns with growing evidence that synthetic speech can mimic target idiosyncrasies closely enough to trigger false accepts in deployed ASV pipelines [68]. At the same time, security-community work shows that adversarial examples targeted at speaker recogni- tion can achieve high attack success and transfer across models and channels, including over-the-air playback, illustrating that both synthesis and perturbation routes are viable for bypassing voice authentication [9]. Challenge evaluations highlight the generalization gap of anti-spoofing countermeasures: systems trained on limited spoof types and benign channels often degrade under unseen conditions, particularly when transmis- sion codecs or physical presentation artifacts are introduced. Recent anti-spoofing research explores raw-waveform and self- supervised front ends (e.g., AASIST, wav2vec-based CMs) that improve robustness but still struggle with shifts across attack families and presentation media [201]. Social engineering can be amplified by audio deepfakes and impersonation in vishing and mixed-media social engineering [91]. Controlled studies indicate that listeners struggle to reliably distinguish between cloned and genuine speech, with accuracy levels that leave a substantial margin for deception. Performance varies modestly with training and user char- acteristics, and can deteriorate further for specific popula- tions or task settings. Broader human-factors syntheses and practitioner-oriented analyses converge on the same risk pat- tern: voice deepfakes exploit authority, urgency, and contextual priming to increase compliance in high-stakes settings such as finance, IT support, and executive impersonation [86]. Tech- nical overviews forecast compounding effects as generative models improve cross-lingual, emotion-preserving, and real- time capabilities, enabling convincing, interactive calls rather than static playbacks [67]. Early detection efforts for spoken social engineering focus on prosodic irregularities, lexical anomalies, and call-graph patterns; emerging vishing-specific models point to speech-based risk scoring as a complement to content filters in email and chat [335]. TABLE VIII COMPARISON OF AUDIO DEEPFAKE DETECTION METHODS MethodStrengthsWeaknessesReal-time AASIST [1]Graph attentionTrainingdata dependent Yes wav2vec-CMSelf-supervisedHigh computeMedium Spectrogram CNNFast, provenCodec sensitiveYes Prosodic AnalysisInterpretableEasy to evadeYes LCNNLightweightLimited capac- ity Yes RawNet2End-to-endLarge modelMedium TABLE IX VOICE CLONING SYSTEMS COMPARISON SystemEnrollmentQuality VALL-E3 secHuman-parity XTTS6 secNear-parity Tortoise30 secHigh RVCVariableHigh C. Textual Deepfakes Textual deepfakes in the context of the MCP refer to AI-generated or manipulated text that deceives LLMs by embedding malicious instructions within tool descriptions, metadata, or external data sources [92]. These threats ex- ploit MCP’s reliance on natural language interfaces for tool discovery and invocation, enabling attacks such as context poisoning and unauthorized actions. Recent research highlights tool poisoning as a prevalent vulnerability, where malicious developers embed covert imperatives in tool docstrings, lead- ing to data exfiltration or system compromise without user awareness [136], [309]. For instance, a benign addition tool might include hidden directives to read sensitive files, such as SSH keys, and transmit them to an attacker-controlled server, achieving high attack success rates across multiple LLMs [316]. Preference manipulation attacks (PMAs) further amplify these risks by using persuasive language in tool descriptions to bias LLMs toward selecting malicious options, leading to toolflow hijacking and economic exploitation [63], [84]. Indirect prompt injection, another key threat, involves poisoning external resources, such as GitHub issues or APIs, with fake instructions that LLMs treat as trusted inputs, leading to privilege escalation or data leakage [115]. Rug pull attacks introduce delayed malice, where initially benign servers update to include deceptive text, eroding trust in MCP ecosystems [117]. These vulnerabilities are exacerbated by namespace typosquatting, where similar tool names confuse LLMs, leading to impersonation and lateral movement [276]. In CPS environments, such textual deepfakes can manifest as fake instructions mimicking legitimate commands, enabling phishing through forged operational logs or misinformation via altered sensor data interpretations [129]. Empirical studies show that up to 5.5% of MCP servers exhibit tool poisoning, with low refusal rates in mainstream clients, underscoring the arms race between generative text capabilities and protocol security [316]. Phishing can be increased by textual deepfakes and business email compromise by generating context-consistent emails, chat messages, or ticket updates that imitate organizational voice, style, and workflow artifacts [77]. Within MCP-style ecosystems, where agents broker access to tools (mail, cal- endaring, ticketing, code repos) through structured tool calls, these messages can be routed directly into agent memory, schedule planners, or action queues, turning persuasive text into privileged operations (e.g., initiating payments, rotating credentials, or exfiltrating reports). Recent surveys and system studies show that machine-generated text can match or exceed human-crafted phishing in plausibility and personalization. At the same time, modern NLP-based detectors still struggle to JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202616 generalize across generators and red-team paraphrases [34], [260]. In multi-tool settings, attackers also combine textual lures with URL or form-based payloads that exploit down- stream tools (browsers, PDF parsers, API wrappers) invoked by the agent, bridging social engineering with programmatic exploitation [332]. Detection pipelines embedded at the MCP ingress benefit from ensemble signals, content stylometry, discourse-level inconsistencies, and token-level statistical artifacts, combined with URL/attachment risk scoring before any tool invocation occurs [162]. Reviews highlight that transformer-based clas- sifiers, graph features that capture conversation context, and distributional tests (e.g., Benford-style attention irregularities) improve robustness, especially when paired with continuous retraining against newly released generators [303]. Comple- mentary provenance cues, such as text watermarking, can be propagated through MCP metadata to inform routing, throt- tling, or human-in-the-loop handoff when confidence is low, thereby reducing the likelihood that convincing but synthetic requests trigger sensitive tools [78]. Misinformation in cyber-physical settings, textual deep- fakes extend beyond public social media into operational channels, such as maintenance logs, incident tickets, supplier messages, and “safety advisories”, that agents ingest as ground truth. When MCP connectors synchronize external knowledge bases or message buses into an agent’s context window, fabricated narratives can bias diagnosis, triage, or planning modules, leading to mis-prioritized repairs, unnecessary shut- downs, or inappropriate configuration changes [162]. The misinformation literature shows that deep neural generators exploit stylistic and rhetorical patterns that evade simple lexicon checks. That multimodal and context-aware models are needed to reconcile claims with telemetry or verified knowledge graphs before agents treat text as actionable state [27]. Within agent toolchains, enforcing claim-verification calls (e.g., fact-checking APIs, document retrieval with stance detection) as a prerequisite to high-impact actions constrains how far synthetic narratives can propagate across tools and across agents. ResilientMCPdeployments,therefore,emphasize provenance-awarecontextassemblyandcross-source consistencytests.Systematicreviewsdocumentgains from hybrid pipelines that combine content features, propagation/interactiongraphs,andexternal-knowledge grounding; these strategies are directly applicable when agents fuse text with sensor or transactional streams before issuing tool calls [162], [303]. Emerging enterprise studies on phishing and fraud detection further indicate that real- time classifiers at the message/URL layer, combined with post-ingest anomaly detection on agent plans, provide defense-in-depth when synthetic narratives attempt to steer downstream tools [158]. Fake instructions also arrive as “authoritative” instructions, chatops snippets, runbooks, pseudo-SOPs, or issue-thread comments that look operationally valid yet encode malicious goals, unsafe parameters, or subtle policy overrides [393]. In MCP-like architectures, where tool schemas are exposed and agents translate natural language into structured actions, these crafted instructions can exploit model-level instruction- following to induce unsafe tool sequences, escalate from low- risk to high-risk capabilities, or propagate across agents via shared memory and broadcast channels [193]. Security surveys of LLM use and agent tooling reveal the resulting attack sur- face, including role confusion, policy evasion through semantic reframing, and cross-tool “confused-deputy” behaviors that occur when untrusted text is mapped to high-privilege tool invocations. Contemporary work on agent ecosystems proposes enforce- able interfaces (explicit argument typing, capability scoping, and pre-/post-conditions) and provenance signals (watermarks, cryptographic signing of trusted instructions) that agents can check before executing sensitive steps or sharing derived plans with peers [367]. Within textual-deepfake-rich environments, integrating detection results and provenance into MCP re- quest/response envelopes enables agents to treat unverified instructions as hypotheses requiring retrieval-augmented vali- dation or human confirmation, rather than as executable intent; this design aligns with multi-agent workflow guidance in recent surveys of LLM tool use and multi-agent orchestration [353]. TABLE X TEXTUAL DEEPFAKE ATTACK VECTORS IN MCP Attack TypeMechanismPrevalenceDetection Tool PoisoningHidden docstring commands 5.5% serversDifficult Prompt InjectionEmbedded instructions HighMedium Rug PullDelayed modifi- cation EmergingVery Difficult TyposquattingName confusionMediumEasy PMAPersuasive biasMediumDifficult D. Behavioral Deepfakes Behavioral deepfakes in the MCP encompass AI-generated emulations of agent behaviors that exploit protocol interfaces to mimic legitimate interactions, posing significant risks to CPS by enabling stealthy manipulations of decision-making processes. These threats involve synthesizing deceptive be- havioral patterns, such as forged tool responses or inter-agent communications, to bypass trust mechanisms and induce unau- thorized actions. Recent research underscores how behavior emulation can facilitate attacks such as context poisoning in MCP ecosystems, where emulated agent behaviors propagate misinformation, leading to real-time system compromises in CPS applications [373]. For instance, in multi-agent envi- ronments, adversaries can emulate cooperative behaviors to achieve collusion, exploiting MCP’s standardized messaging to mask anomalies and evade detection, particularly in domains such as autonomous robotics, where behavioral fidelity is critical [328]. A key vulnerability lies in sensor spoofing of MCP tools, which can cause unsafe actuations in industrial control systems [289]. Anomaly mimicry further heightens risks by replicating normal operational patterns to conceal malware persistence, leveraging MCP’s supply-chain dependencies on external JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202617 servers [247]. Stealthy malware deployment via behavioral emulation allows for lateral movement across agent networks, as demonstrated in studies of user behavior emulation that highlight the potential for persistent access in cloud-integrated CPS [346]. Analyses indicate that such emulations achieve high success rates in protocol-based systems, emphasizing the need for enhanced governance to counter these evolving threats [328]. Sensor spoofing: Within MCP ecosystems, behavioral deepfakes appear as forged yet statistically plausible sensor streams that flow from environment-facing tools into agent contexts. By shaping Global Navigation Satellite System (GNSS), SCADA, or process telemetry to satisfy physical and topological constraints, an adversary can nudge downstream reasoning and tool use while preserving the appearance of routine operation. Recent GNSS studies show that deep models can both craft and detect complex spoofing waveforms, under- scoring that realistic temporal–spectral patterns can subvert localization-dependent decisions unless provenance and cross- sensor corroboration are enforced at the tool boundary [53]. Large-scale sensor network work further demonstrates that distributed false-data injection, framed as normal fluctuations, can bypass conventional residual tests, with detection improv- ing only when spatio-temporal structure is explicitly modeled [139]. In operational ICS settings, forensic analyses docu- ment “low-and-slow” perturbations that exploit sensor noise envelopes to keep deviations within learned confidence bands, enabling subtle set-point shifts without triggering alarms [36]. A complementary thread examines adversarial time-series attacks that optimize over temporal features, e.g., seasonality, lagged correlations, and shapelets, so that injected traces are misclassified as routine by state-of-the-art detectors. Traffic and ICS studies show that such feature-aware perturbations de- grade multivariate detectors and forecasting-residual schemes alike, achieving high evasion with small, causality-respecting edits [209]. Domain studies in water networks similarly reveal that hydraulics-consistent sensor attacks can remain stealthy while steering control actions, highlighting the need to treat MCP tool outputs as untrusted inputs whose behavioral plau- sibility alone is insufficient for trust [11]. Anomaly mimicry reframes the attacker’s goal as blending in with the learned behavior model: the adversary synthesizes trajectories that reproduce nominal multivariate dependencies so closely that detectors classify them as in-distribution. Contemporary surveys in multivariate time-series anomaly detection report persistent generalization gaps under distribu- tion shift and coordinated multichannel perturbations, condi- tions under which mimicry attacks thrive [342]. Generative- model-based studies further show that class-conditional and autoencoder-centric detectors can be driven toward false neg- atives when attackers optimize reconstructions or latent codes to match normal manifolds, even when the resulting sequences induce harmful actuation downstream [202], [343]. System-level reviews in cybersecurity analytics emphasize that many deployed detectors assume weak temporal sta- tionarity and local consistency; carefully crafted sequences that maintain those assumptions at short horizons can still produce dangerous long-horizon drifts, especially when agent policies rely on MCP tools that summarize or forecast over rolling windows [183]. Results from control-theoretic analyses of stealthy attacks on remote state estimation formalize this effect: nonlinear measurement tampering can keep innova- tion statistics within thresholds while gradually biasing the estimated state, illustrating how anomaly mimicry translates into actionable control error without overt signatures in the residuals [302]. Stealthy malware mimics benign telemetry by behavioral deepfakes that extend into the code–behavior layer, API se- quences, and communication rhythms to survive behavior- based defenses orchestrated by MCP-connected tools. Empir- ical studies show that classifier decisions are sensitive to eva- sive behaviors such as delayed loading, environment checks, and feigned user-driven I/O, enabling samples to cross decision boundaries while preserving functional malicious goals [240]. Broader meta-surveys of adversarial attacks on deep models, including detectors used in security analytics, reinforce that small, structured edits to behavior traces or feature embed- dings can induce misclassification across families and vendors, foreshadowing automated “policy-aware” evasion in MCP- mediated pipelines [255]. Parallel work explores automated generation of attack techniques using learning-based planners, suggesting that behavior mimicry may soon be synthesized at scale rather than hand-engineered [148]. Stealth also manifests in command-and-control patterns tuned to resemble benign heartbeat and telemetry processes observed by MCP toolchains. Unsupervised analyses of bea- coning demonstrate that periodicity, jitter, and payload traits can be shaped to evade statistical profiling while maintaining reliable control, complicating anomaly detection that relies on coarse traffic statistics [215]. At the same time, dynamic graph-based detectors that model process–file–socket interac- tions reveal promise against such mimicry by capturing higher- order temporal context; results indicate improved resilience to look-alike behaviors compared with flat sequence models, provided that execution provenance and graph dynamics are preserved end-to-end through MCP connectors [24], and that ensemble defenses are evaluated against generative, transfer- capable malware families [226]. TABLE XI BEHAVIORAL DEEPFAKE ATTACK CATEGORIES CategoryTargetStealthImpact Sensor SpoofingGNSS/SCADA/ Telemetry Very HighCritical Anomaly Mimicry ML DetectorsHighHigh Stealthy MalwareBehavior AnalysisHighCritical C2 BeaconingTraffic AnalysisMediumHigh False Data Injec- tion State EstimationVery HighCritical E. Privacy implications In the MCP, privacy implications of deepfakes are pro- foundly amplified by the protocol’s facilitation of seamless tool integrations and context sharing among AI agents, cre- ating fertile ground for identity theft and non-consensual JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202618 TABLE XII DETECTION METHODS FOR BEHAVIORAL DEEPFAKES MethodApproachEvasion Risk Latency Residual TestsStatisticalHighLow AutoencoderReconstructionHighMedium Graph-basedStructuralMediumHigh Spatio-temporalMulti-channelMediumHigh Physics-informedDomain knowl- edge LowMedium EnsembleCombinedLow- Medium High data exploitation in CPS. Adversaries can leverage deepfake- generated synthetic interactions, such as fabricated user prompts or emulated agent behaviors, to impersonate entities within MCP ecosystems, enabling the unauthorized extraction of personally identifiable information (PII) from intercon- nected tools, including sensor APIs or collaborative databases [165], [308]. For instance, non-consensual intimate deepfakes, often disseminated via MCP-mediated channels, pose severe risks to individuals in professional or educational CPS settings, where altered session logs or forged endorsements can lead to reputational harm and psychological distress without recourse [170]. This vulnerability is exacerbated by the ”liar’s divi- dend,” wherein genuine privacy violations are plausibly denied as AI-generated artifacts, undermining accountability mecha- nisms and eroding trust in protocol-dependent infrastructures [173]. Empirical studies underscore that such deepfakes not only facilitate targeted surveillance but also perpetuate soci- etal biases, as generative models trained on skewed datasets amplify discriminatory profiling in MCP-orchestrated multi- agent collaborations [313]. Moreover, the interoperability of MCP introduces cascading privacy risks through supply-chain exposures, where third- party tools vulnerable to deepfake injections inadvertently propagate sensitive environmental or user data across CPS networks, challenging regulatory frameworks like GDPR in real-time applications [125]. Research highlights the ethical quandaries of deepfake misuse in agentic systems, including the potential for misinformation campaigns that exploit MCP’s context persistence to fabricate evidence of privacy breaches, thereby complicating forensic attribution and victim redress [10]. In educational or industrial CPS, this manifests as the non-consensual repurposing of biometric or behavioral data into deepfakes, fostering a chilling effect on user participa- tion and necessitating privacy-enhancing technologies such as differential privacy embeddings within protocol specifica- tions [170]. Addressing these implications demands interdis- ciplinary approaches, integrating watermarking with federated learning to safeguard against deepfake-driven erosions of au- tonomy while preserving the protocol’s utility for trustworthy AI interactions [313]. Identity theft can be ingested by agent connectors in MCP ecosystems, summarized, and acted on multimodal inputs; when those inputs include face/voice streams or profile ar- tifacts, deepfake-assisted impersonation raises the risk that downstream tools will accept falsified identities and authorize sensitive operations. Recent analyses of deepfake fraud and criminal-justice risks describe how synthetic media can bypass biometric checks and erode confidence in identity evidence, with direct implications for KYC flows, remote onboarding, and account recovery that rely on camera or microphone verification [254], [293]. Technical surveys further underscore that cross-modal deepfakes are routinely linked to identity theft and fraud, while state-of-the-art detectors trained on narrow corpora generalize poorly, conditions that incentivize attackers to target MCP entry points and toolchains that treat media- derived identity cues as trustworthy [163], [224]. Within the MCP boundary, identity exposure also occurs through inference and linkage: outputs from face-swap detec- tion, liveness checks, or profile enrichment services can leak biometric and behavioral hints that enable deanonymization or later impersonation. Empirical work on generalization for deepfake detection reveals that performance is fragile under content, codec, and generator shifts, thereby increasing the likelihood that adversaries can curate “just-different-enough” media to evade filters and establish trust for subsequent identity takeovers [315]. Studies of human vulnerability to science and news deepfakes add that even well-informed operators misclassify realistic synthetic videos at non-trivial rates, compounding the risk when MCP agents escalate tool calls based on operator confirmation alone [86], [89]. Non-consensual media lifecycle can be amplified by MCP- connected capture, storage, and transformation tools, including synthetic nudes and AI-driven CSAM [250]. Legal and foren- sic scholarship documents how AI-generated sexual content exploits gaps in existing regimes, travels rapidly across plat- forms, and inflicts sustained reputational and psychological harm on victims, especially women; these dynamics raise the privacy stakes for any pipeline that ingests or redistributes media without provenance checks and explicit consent tracking [123], [170]. At the protocol layer, MCP integrations that summarize media, generate thumbnails, or auto-route alerts can propagate non-consensual derivatives into logs, caches, and secondary workflows, even after the content has been taken down. Foren- sic and detection research stresses that maintaining privacy requires more than point classifiers: provenance-aware inges- tion, fine-grained consent state, and robust content-authenticity signals are needed to disrupt circulation; meanwhile, user studies show persistent vulnerability to persuasive deepfakes in educational and public-communication contexts, highlighting a double bind in which both automated and human reviewers are fallible under realistic content [89], [313]. “Liar’s dividend” as MCP agents increasingly mediate evi- dence flows, retrieving footage, transcribing calls, and compil- ing incident timelines, the mere availability of deepfake tools enables a “liar’s dividend”: wrongdoers can dismiss authentic recordings and logs as fabricated, shifting burdens of proof and degrading the epistemic value of MCP-assembled reports. Scholarly accounts trace how deepfakes fuel informational un- certainty that weakens journalism and democratic deliberation, explicitly naming the liar’s dividend as a mechanism of denial and strategic doubt [211], [254]. The dividend interacts with privacy in two ways. First, heightened deniability incentivizes broader data collection and JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202619 retention to “prove authenticity,” expanding the personal-data surface inside MCP connectors and archives. Second, as detec- tion remains imperfect and human discernment is error-prone, claimants may be compelled to disclose additional private context (raw sensor feeds, biometric templates, location traces) to rebut “it’s a deepfake” allegations, thereby trading privacy for credibility [86]. These dynamics motivate provenance-rich media pipelines and careful access policies; however, the core privacy risk persists as long as plausible synthetic alternatives can be invoked to contest genuine records. TABLE XIII PRIVACY THREAT CATEGORIES IN MCP-DEEPFAKE CONTEXTS ThreatMechanismSeverityLegal Gap Identity TheftBiometric bypassCriticalMedium Non-consensual Media Synthetic generation CriticalHigh Liar’s DividendPlausible deniabil- ity HighVery High PII ExtractionAgent impersonation HighMedium Discriminatory Profiling Biased modelsMediumHigh TABLE XIV PRIVACY-PRESERVING COUNTERMEASURES MeasureProtectionOverheadAdoption C2PA ProvenanceAuthenticityLowGrowing Differential Privacy Data protectionHighLimited Federated LearningDecentralizedMediumEmerging Consent TrackingUser controlLowLimited WatermarkingAttributionLowMedium F. AI Agent Vulnerabilities to Deepfake Threats While the preceding sections focus on deepfake modalities (visual, audio, textual, and behavioral) and their impacts through MCP, this section analyzes the architectural vulner- abilities of AI agents themselves that make them susceptible to deepfake-based attacks in CPS environments. 1) AI Agent Architecture and Attack Surfaces: Modern AI agents are constructed on four core components: perception, brain (LLM), action, and memory [82]. Each component creates distinct attack surfaces that deepfakes can exploit. Recent surveys identify four critical knowledge gaps in AI agent security that directly relate to deepfake vulnerabilities: • Gap 1: Unpredictability of multi-step user inputs – Deepfake audio/text can masquerade as legitimate user inputs across multiple interaction rounds. • Gap 2: Complexity in internal executions – Behavioral deepfakes can mimic reasoning patterns to evade internal consistency checks. • Gap 3: Variability of operational environments – Visual deepfakes alter environmental perception, causing agents to misinterpret physical world states. • Gap 4: Interactions with untrusted external entities – All deepfake modalities can be injected through MCP tools and external data sources. Threats on AI agent perception exploit model-level vulnera- bilities to manipulate the “brain” of AI agents, compromising or bypassing policy constraints from benign instructions and leading to improper actions that break the integrity of the agent ecosystem [82]. 2) Perception Layer Attacks via Multimodal Deepfakes: The perception layer serves as the primary entry point for AI agents and is particularly vulnerable to multimodal deepfakes. Research demonstrates that attackers can embed malicious prompts within images or audio to bypass security mechanisms when combined with text [40]. This vulnerability is especially dangerous when AI agents possess tool-calling capabilities. Adversarial attacks on multimodal agents have shown that visual adversarial examples can cause LLM-based agents to misuse tools, resulting in unintended actions within CPS [359]. The multimodal nature of modern LLMs, which support voice, images, and text, increases flexibility but also expands the attack surface for deepfake injection [98]. Jailbreaking of AI agents can occur through three main vectors: multi-turn dialogues, multimodal inputs, and external environmental data [82]. In multi-turn interactions or role- playing scenarios where agents act as planners or experts, harmful outputs become harder to detect, making deepfake- based manipulation particularly effective [355]. 3) Multi-Agent Systems and Cascade Effects: In multi- agent systems (MAS), a single compromised agent can propa- gate malicious behavior throughout the entire network, a phe- nomenon termed the “domino effect” [82], [374]. This cascade vulnerability is particularly concerning in CPS contexts where deepfake inputs to one agent can compromise system-wide integrity. Research on prompt infection demonstrates that a single adversarial string can propagate among agents, starting with one harmful agent and ultimately compromising all agents in the collection [374]. The self-propagating nature of such attacks means that initial deepfake inputs can achieve system- wide compromise through normal inter-agent communication channels. Studies on the flooding spread of manipulated knowledge in LLM-based multi-agent communities reveal how misin- formation, potentially seeded by deepfakes, can propagate through collaborative reasoning, negatively affecting collec- tive decision-making [157]. Multi-agent debates can improve robustness, but cooperation among agents may also cause a domino effect where one compromised agent jeopardizes others [23]. Defense frameworks such as BlindGuard [220] and G- Safeguard [347] employ graph-based approaches to detect malicious agents in MAS, evaluating defense capabilities against direct prompt attacks, tool attacks, and memory at- tacks. These frameworks represent emerging countermeasures against deepfake-initiated cascade failures. 4) Embodied AI and Physical World Threats: Embodied AI systems, including robots and autonomous vehicles, face unique risks when deepfakes are employed to attack systems capable of causing physical harm. These systems encounter vulnerabilities stemming from both environmental and system- level factors, manifesting through sensor spoofing, adversarial attacks, and failures in task and motion planning [362]. The taxonomy of embodied AI vulnerabilities encompasses: JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202620 • Exogenous origins: Physical attacks and cybersecurity threats including sensor spoofing and adversarial patches. • Endogenous origins: Sensor failures and software flaws that can be exploited by deepfakes to cause cascading failures. Research on jailbreaking robotic manipulation demonstrates that embodied AI jailbreaks transcend text generation to pro- duce potential physical actions, thereby significantly amplify- ing security risks compared to purely linguistic attacks [383]. A comprehensive set of 230 malicious physical world queries has been developed to probe embodied AI systems, grounded in IEEE Ethically Aligned Design guidelines. TheBALD(BackdoorAttacksagainstLLM-based Decision-making systems) framework represents the first comprehensive approach for backdoor attacks in embodied AI, proposing three distinct mechanisms: word injection, scenario manipulation, and knowledge injection [153]. Experiments on GPT-3.5, LLaMA2, and PaLM2 in autonomous driving and home robot tasks demonstrate high attack success rates, underscoring the vulnerability of embodied agents to deepfake-style manipulations. 5) Tool Poisoning and MCP Security Mechanisms: The Model Context Protocol creates novel attack surfaces that extend beyond traditional deepfake threats. Security analyses reveal alarming statistics: 43% of tested MCP server im- plementations contain command injection flaws, 22% permit arbitrary file read via path traversal, and 30% are vulnerable to Server-Side Request Forgery [87], [130]. Tool Poisoning Attacks (TPAs) occur when malicious in- structions are embedded within MCP tool descriptions that are invisible to users but visible to AI models [144]. These attacks exploit MCP’s security model assumption that tool descriptions are trustworthy and benign. Key attack vectors include: • Rug Pull attacks: Tools dynamically alter their behavior or description after users grant permission, enabling silent credential theft or API key exfiltration. • Tool Shadowing: In multi-server configurations, mali- cious servers impersonate tools from trusted servers. • Cross-server Interference: Server A can redefine tools from Server B, enabling interception of sensitive opera- tions. Real-world vulnerabilities include CVE-2025-6514 affect- ing 437,000+ downloads of mcp-remote through OAuth dis- covery vulnerabilities, and CVE-2025-49596 in MCP Inspec- tor enabling remote code execution via CSRF [87]. Academic research identifies 5.5% of MCP servers exhibiting tool poi- soning behaviors, representing a new class of AI-targeted vulnerabilities. 6) Agent-to-Environment Threats in CPS: Threats on Agent2Environment exploit vulnerabilities arising from un- trusted dynamic feedback and complex interactions in diverse operational settings [82]. These threats encompass indirect manipulation of input data, unintended behaviors influenced by dynamic states, and environmental discrepancies, all of which can be induced or amplified by deepfakes. Input data from the physical environment must undergo rigorous security checks to filter threats and ensure safety. Clear and compatible communication between LLM-generated instructions and hardware execution is vital to avoid opera- tional errors. For real-world deployment, agents must prioritize accuracy to minimize irreversible harm caused by incorrect actions [82]. The physical environment poses significant security chal- lenges due to its complexity. Insufficient isolation between agents in shared environments enables malicious agents to potentially access or interfere with operations of other agents, leading to data breaches, unauthorized access to sensitive in- formation, or the spread of malicious code [82]. Unmonitored resource usage can mask security breaches, as anomalous behavior such as sudden spikes in resource consumption might indicate deepfake-initiated attacks. 7) Defense Frameworks for AI Agents: Multi-layered de- fense approaches are essential for protecting AI agents against deepfake threats: • Input Layer: Input validation, provenance verification, and deepfake detection for visual/audio inputs. • Reasoning Layer: Certified robustness, adversarial train- ing, and behavioral pattern detection. • Action Layer: Least-privilege enforcement, human-in- the-loop verification for tool calls. • Memory Layer: Taint tracking, encryption, and context poisoning prevention. Emerging defense technologies include CaMeL for miti- gating prompt injection attacks [357], AgentGuard for repur- posing agentic orchestrators for safety evaluation [66], and BlockAgents for Byzantine-robust multi-agent coordination via blockchain [64]. These frameworks address the unique challenges posed by deepfake-style attacks on AI agent ar- chitectures. Integration of homomorphic encryption schemes and attribute-based forgery generative models can safeguard against privacy breaches during communication processes, though at additional computational and communication costs [82]. Supply chain threats, including buffer overflow, SQL injection, and cross-site scripting vulnerabilities in tools, re- quire comprehensive security auditing of all MCP server dependencies. TABLE XV AI AGENT ATTACK SURFACE BY COMPONENT ComponentDeepfake VectorSeverityDefense Maturity PerceptionVisual/Audio in- jection CriticalMedium Brain (LLM)Prompt injectionCriticalLow ActionTool poisoningCriticalLow MemoryContext poison- ing HighLow G. Cross-Category Comparison and Summary Key Research Gaps Across All Categories: 1) Generalization: All detection methods suffer from poor generalization to unseen generators and attack variants. 2) Real-time Performance: CPS applications require low- latency detection incompatible with current complex models. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202621 TABLE XVI MULTI-AGENT DEFENSE FRAMEWORKS COMPARISON FrameworkApproachOverheadEffectiveness G-SafeguardGraph topologyMediumHigh BlindGuardZero-shot detec- tion LowMedium- High BlockAgentsBlockchainHighHigh AgentGuardOrchestrator reuse LowMedium CaMeLCapability control MediumHigh TABLE XVII MCP VULNERABILITY STATISTICS Vulnerability TypePrevalenceCVE Examples Command Injection43%Multiple Path Traversal22%Multiple SSRF30%Multiple Tool Poisoning5.5%N/A OAuth FlawsSignificantCVE-2025-6514 CSRFVariableCVE-2025-49596 3) Cross-modal Attacks: Combined visual-audio-textual- behavioral attacks remain largely unstudied. 4) MCP Security: Protocol-level security mechanisms are immature despite rapid adoption. 5) Regulatory Frameworks: Legal and governance struc- tures lag behind technical capabilities. 6) Scalability: Defense mechanisms don’t scale to large multi-agent CPS deployments. Recommended Research Priorities: 1) Develop generator-agnostic detection leveraging funda- mental artifacts 2) Create standardized MCP security auditing and certifica- tion frameworks 3) Build physics-informed behavioral deepfake detection for CPS 4) Establish real-time multi-modal verification pipelines 5) Design scalable Byzantine-robust multi-agent coordina- tion 6) Integrate privacy-preserving techniques with detection systems H. Cross-Modal Research Synthesis This subsection consolidates the research analysis across all deepfake modalities, identifying common strengths, shared limitations, and unified research priorities for securing MCP- enabled CPS. 1) Common Strengths Across Modalities: Research on deepfake threats has achieved notable progress applicable to CPS security. Comprehensive attack taxonomies have been developed for each modality—visual (display spoofing, face- swap, scene synthesis), audio (voice cloning, replay attacks), textual (tool poisoning, prompt injection, rug pulls), and behavioral (sensor spoofing, anomaly mimicry, stealthy mal- ware). Multi-modal defense strategies combining physiolog- ical signals (rPPG), provenance frameworks (C2PA), and cross-sensor fusion demonstrate improved robustness over TABLE XVIII EMBODIED AI ATTACK MECHANISMS MechanismTarget SystemSuccess Rate Physical Risk Word InjectionLLM DecisionHighHigh Scenario Manip- ulation ContextHighCritical Knowledge Injection MemoryMediumHigh Sensor SpoofingPerceptionHighCritical Adversarial Patches VisionMediumHigh single-modality approaches. Formal control-theoretic analy- sis enables rigorous characterization of stealthy attack im- pacts, while graph-based and spatio-temporal detectors capture higher-order behavioral context. The “liar’s dividend” phe- nomenon has been recognized as a systemic epistemic threat, informing both technical and policy responses. 2) Shared Limitations: Despite these advances, four funda- mental limitations constrain deployment across all modalities: a) Generalization gap: Detectors trained on benchmark datasets exhibit significant performance degradation against unseen generators, novel attack variants, and in- the-wild conditions featuring compression, noise, and domain shift. b) Real-time constraints: Safety-critical CPS require de- tection within milliseconds; computationally intensive methods (rPPG analysis, graph-based detection, ensemble models) often exceed acceptable latency budgets. c) Adversarial adaptation: Generative models—including neural codec voice cloning (VALL-E), diffusion- based image synthesis, and LLM-powered text genera- tion—consistently outpace detection capabilities. d) Human fallibility: Operators and reviewers misclassify high-quality deepfakes at non-trivial rates, preventing reliance on human-in-the-loop verification as a standalone defense. 3) Unified Research Priorities: Addressing these shared limitations requires coordinated research efforts: • Generator-agnostic detection: Develop methods lever- aging fundamental artifacts (physics-based inconsisten- cies, statistical fingerprints) rather than generator-specific signatures. • Lightweight architectures: Design efficient models suit- able for edge deployment that balance accuracy with CPS resource constraints. • Standardized MCP integration: Establish protocols for provenance verification and detection at MCP tool bound- aries without unacceptable latency. • Physics-grounded validation: Integrate domain knowl- edge (hydraulic constraints, RF propagation models, control-theoretic bounds) into ML detection architectures. • Adaptive regulatory frameworks: Develop cross- jurisdictional mechanisms for rapid response to deepfake harms while preserving privacy and due process. 4) Comparative Summary: Table XXI summarizes detec- tion maturity, key limitations, and CPS deployment feasibility JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202622 TABLE XIX COMPREHENSIVE COMPARISON OF DEEPFAKE THREAT CATEGORIES IN CPS CategoryPrimary TargetDetection Matu- rity CPS ImpactDefense GapResearch Priority VisualSurveillance, AVMediumCriticalGeneralizationHigh AudioVoice Auth, So- cial Eng. MediumHighReal-timeHigh TextualMCPTools, Agents LowCriticalTool VettingCritical BehavioralSensors, Anomaly Det. LowCriticalPhysics-awareCritical PrivacyIdentity, ConsentLowHighRegulatoryHigh AI AgentsAll ComponentsVery LowCriticalComprehensiveCritical TABLE X DEFENSE READINESS ASSESSMENT ACROSS CATEGORIES Defense TypeVisualAudioTextualBehavioralPrivacyAI Agent Detection MLMediumMediumLowLowLowVery Low ProvenanceHighLowMediumLowMediumLow Multi-modalMediumMediumN/AMediumLowLow Human-in-loopMediumMediumMediumLowHighMedium Policy/GovernanceLowLowLowLowMediumVery Low TABLE XXI CROSS-MODAL DEEPFAKE THREAT SYNTHESIS ModalityMaturityGeneral.RT-ReadyKey Gap VisualMediumLowMediumGenerator shift AudioMediumLow-MedMediumCodec robustness TextualLowLowHighLLM evolution BehavioralLowMediumLowPhysics modeling PrivacyLowN/AN/ALegal frameworks across modalities. The analysis confirms that detection mechanisms alone cannot serve as decision authorities in safety-critical CPS. Effective deployment requires integration with provenance verification, physics-based validation, and defense-in-depth architectures as detailed in Section VII. VI. DETECTION TECHNIQUES In the context of deepfakes as a security threat to CPS that leverage the MCP, detection techniques must address the unique challenges of real-time environmental interactions and protocol-mediated data exchanges, where synthetic content can infiltrate agent-to-tool communications and sensor feeds. Traditional artifact analysis, such as identifying inconsistencies in pixel-level artifacts or spectral anomalies in audio streams, has been enhanced with advanced machine learning fusion methods to improve robustness in constrained CPS environ- ments [118]. For instance, multimodal consistency checks, which combine visual, audio, and textual modalities, enable cross-verification of MCP-transmitted data and detect discrep- ancies in agent responses that mimic legitimate environmental signals but fail to meet physiological cues, such as eye-blink irregularities or heartbeat synchronization [21]. Recent studies demonstrate that convolutional neural network (CNN)-based models, integrated with MCP’s structured messaging, achieve up to 95% accuracy in identifying deepfake intrusions in industrial control systems, where forged sensor data could trigger unsafe actuations [42]. These approaches emphasize lightweight architectures for edge devices in CPS, ensuring low-latency detection without compromising protocol interop- erability. Emerging detection strategies, further tailored to MCP’s vulnerabilities, such as context poisoning via behavioral deep- fakes in inter-agent communications, incorporate stylometry for textual analysis and physics-based anomaly detection for sensor emulation [310]. Self-supervised learning frameworks have shown promise in generalizing to unseen deepfakes within MCP ecosystems, utilizing environmental fingerprints, such as ENF traces, to corroborate protocol payloads against physical realities in CPS applications [331]. However, chal- lenges persist in balancing detection efficacy with privacy preservation, as watermarking and cryptographic provenance require standardized implementation across MCP servers to prevent supply-chain exploits [118]. Evaluations on datasets like FaceForensics++ adapted for CPS scenarios reveal that hybrid AI algorithms, including generative adversarial network (GAN) discriminators, outperform single-modality methods by 15-20% in real-world deployments, highlighting the need for adaptive training to counter evolving threats in protocol- enabled AI agents [42]. A. Visual Visual detection addresses forged inputs in sensor streams through artifact analysis [83], physiological cues such as eye movements [150], and multimodal consistency checks across modalities [102]. Artifact analysis in visual deepfake detection focuses on identifying subtle inconsistencies introduced during the generation process, such as pixel-level distortions, blending boundaries, and frequency-domain anomalies, which are often imperceptible to humans but detectable by computational methods. In the context of CPS integrated with MCP, where AI agents process visual streams for real-time decision-making, these artifacts can compromise system integrity by deceiving JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202623 Fig. 12. Deepfakes Detection in CPS Using MCP surveillance or authentication modules. Recent advancements employ frequency-domain analysis improves cross-model gen- eralization by classifying residues within a taxonomy frame- work [296]. This method addresses the challenges posed by residual artifacts in static images, offering robust performance in CPS environments where visual inputs from MCP servers must be sanitized to prevent context poisoning. Further developments in artifact analysis incorporate ex- plainable AI techniques, such as Grad-CAM and Saliency Maps, to localize regions of manipulation and quantify activa- tion intensities, providing insights into model confidence and precision. In MCP-enabled CPS, this approach mitigates risks from adversarial visual inputs by integrating texture decompo- sition, which separates artifacts from natural image elements, thereby improving generalization across diverse forgery sce- narios, particularly on edge devices handling compressed or noisy data [101]. Physiological cues exploit biological signals, such as heart rate variability, blinking eyes, and rPPG, that generative mod- els cannot accurately replicate [323], [333]. Hybrid models fusing physiological and visual features reduce false posi- tives, though blink-based detection faces robustness challenges across diverse populations [10], [313], [319]. Multimodal consistency evaluates alignment across visual, audio, and temporal modalities to detect discrepancies such as lip-sync mismatches [149], [291]. Ensemble fusion and spatial-spectral analysis outperform single-modality methods on mixed-manipulation datasets [21], [177], [306], [338]. B. Audio Audio detection mitigates voice spoofing risks that com- promise agent-to-tool interactions [382]. ML frameworks ana- lyze acoustic artifacts and temporal inconsistencies in real- time CPS environments [18], with adversarial training en- hancing robustness [304] and hybrid models reducing false positives [51]. Multi-modal fusion cross-verifies audio with environmental signals, while forensic analysis detects traces in high-frequency spectra [340]. Lightweight detectors balance efficiency for edge deployment [22], underscoring the need for standardized audio provenance protocols [168]. Spectral analysis for audio deepfake detection in MCP involves examining frequency domain characteristics to un- cover artifacts introduced by synthesis algorithms, such as unnatural energy distributions or phase inconsistencies that can infiltrate protocol-based communications in CPS [386]. This technique is particularly effective against voice cloning threats, where MCP servers process audio inputs that may be laced with synthetic content, enabling attacks such as context poisoning. Recent research employs advanced spectrogram- based models to classify deepfakes, leveraging convolutional neural networks for feature extraction and achieving supe- rior performance on benchmark datasets [58]. For example, chroma-based spectral methods have been refined to detect inconsistencies in harmonic structures, and malicious audio could mimic legitimate tool responses. Evaluations show that these approaches outperform traditional time-domain analysis in noisy CPS environments, with accuracy rates exceeding 95% in controlled simulations [119]. Additionally, advancements in spectral analysis incorporate wavelet transforms for multi-resolution analysis, enhancing the detection of subtle manipulations in MCP-transmitted audio streams. Studies highlight the integration of spectral entropy measures to quantify irregularity, thereby mitigating risks posed by adversarial audio that exploits weak MCP boundaries [118]. Limitations, such as sensitivity to compression artifacts, are addressed through robust preprocessing, as demonstrated in comprehensive reviews advocating hybrid spectral-deep learning frameworks for CPS security [228]. These methods provide a foundational defense, emphasizing the importance of spectral fingerprinting to verify audio authenticity in real-time MCP interactions [30]. Liveness/challenge-response tests in audio deepfake de- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202624 tection for MCP ensure the authenticity of voice inputs by verifying real-time human presence, countering pre-recorded or synthesized audio that could subvert CPS agent proto- cols. These active methods involve prompting users with random challenges, such as repeating phrases, to detect non- interactive deepfakes infiltrating MCP communications [8]. Recent developments integrate biometric liveness detection with acoustic analysis, improving resilience against spoof- ing in CPS, where audio commands trigger physical actions [249]. For instance, challenge-response protocols have been enhanced with machine learning to adapt challenges dynam- ically, reducing evasion rates in MCP ecosystems prone to voice impersonation [70]. Studies report high effectiveness in distinguishing live from synthetic speech, particularly in multi- factor authentication scenarios for CPS [259]. Moreover, combining liveness tests with environmental noise correlation strengthens defenses against sophisticated deepfakes in MCP, where adversaries might inject cloned au- dio via untrusted servers. Research emphasizes hybrid systems that fuse challenge-response with passive liveness cues, such as breathing patterns, to achieve comprehensive coverage in CPS applications [191]. Challenges include user convenience and adaptability to diverse accents, addressed through person- alized models in recent surveys [258]. These techniques are vital for MCP security, promoting interactive verification to prevent deepfake exploitation in critical infrastructures [352]. C. Text Textual deepfake detection identifies forged tool descrip- tions and manipulated prompts that enable context poisoning. ML models analyze linguistic patterns in MCP-exchanged text [228], with multimodal approaches addressing low-quality inputs [265]. Hybrid CNN-stylometric models distinguish AI- generated content [283]. Proactive strategies including water- mark extraction and provenance verification mitigate supply- chain vulnerabilities, with robust algorithms maintaining low false-positive rates [112]. These techniques provide forensic tools to authenticate text origins [21]. Stylometry in textual deepfake detection for MCP involves analyzing writing styles, such as vocabulary richness, sentence structure, and syntactic patterns, and the origins of text, including vocabulary richness, sentence structure, and syn- tactic patterns, to differentiate AI-generated from human text in protocol-mediated exchanges. This method is particularly effective against preference manipulation attacks, in which malicious tool descriptions bias LLM selection in CPS. Recent research integrates stylometric features with deep learning for improved authorship attribution, achieving high accuracy in detecting subtle inconsistencies in MCP-processed prompts [105]. For example, ensemble models incorporating n-gram analysis and machine learning classifiers have been proposed to combat fake news, such as deepfakes, which are adaptable to MCP’s inter-agent messaging [126]. Evaluations of dialectal texts reveal the robustness of stylometry to linguistic variation, an essential property for global CPS deployments [52]. Advancements in stylometry also focus on explainable AI to localize manipulative elements in text, enhancing trans- parency in MCP security audits. Studies emphasize hybrid approaches combining stylometric metrics with blockchain to enhance provenance and reduce evasion risks from adaptive adversaries [95]. Limitations such as sensitivity to text length are mitigated through multitask learning frameworks, enabling real-time detection in resource-constrained CPS environments [198]. These techniques underscore the role of stylometry in fortifying MCP against textual deepfakes, promoting resilient agent behaviors [245]. AI watermarking for textual deepfake detection in MCP embeds imperceptible markers into generated text during LLM inference, enabling reliable identification of synthetic content in tool outputs or retrieved data. This proactive defense coun- ters rug-pull attacks by enabling post-generation verification without degrading text quality. Recent developments include probabilistic curvature-based watermarks that preserve seman- tic integrity while facilitating high-fidelity detection in CPS communications [387]. For instance, scalable schemes like SynthID-Text integrate with open-source models, achieving minimal impact on utility and robust extraction even under paraphrasing [78]. Surveys highlight watermarking’s superior- ity over passive detectors in handling LLM variations, crucial for MCP’s heterogeneous services [370]. Moreover, AI watermarking advancements incorporate con- trastive learning to enhance robustness against removal attacks, ensuring traceability in MCP ecosystems prone to supply- chain compromises. Research on fine-tuning datasets demon- strates that watermarks persist across edits, supporting forensic analysis in CPS [212]. Challenges such as detection thresholds are addressed with supervised and unsupervised methods, optimizing for low false positives [54]. These innovations position watermarking as a cornerstone for securing textual exchanges in MCP, thereby fostering the ethical deployment of AI [178]. Cryptographic provenance in textual deepfake detection for MCP establishes verifiable chains of origin for text data, utilizing digital signatures and blockchain to authenticate content traversing protocol boundaries. This mitigates identity fragmentation by binding metadata to MCP messages, pre- venting spoofed instructions in CPS. Recent frameworks lever- age NFTs and decentralized ledgers to provide tamper-proof provenance, effectively combating the dissemination of deep- fakes in agent networks [329]. For example, hybrid systems integrate cryptographic hashes with AI forensics, enabling real-time verification of tool responses [143]. Studies on post- quantum cryptography address future threats, ensuring long- term security in evolving MCP implementations [30][31][32]. Additionally, cryptographic provenance techniques empha- size interoperability with content credentials standards, facili- tating cross-platform detection of synthetic text. Research on blockchain-based traceability mechanisms has demonstrated reduced latency in CPS audits and counters misinformation through the use of immutable records [127]. Limitations such as scalability are addressed through lightweight protocols that balance computational overhead with efficacy [350]. These methods enhance MCP’s resilience, providing a foundational layer for trustworthy AI interactions in critical infrastructures [3]. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202625 D. Behavioral/Sensor Behavioral and sensor detection identifies emulated behav- iors and spoofed sensor data, including anomaly mimicry and stealthy malware insertions [382]. Unsupervised learning handles dynamic CPS environments with high real-time detec- tion rates [82], while hybrid frameworks blend network and physical data for anomaly identification [271]. AI-enhanced physical unclonable functions enable real-time anomaly de- tection [224], with digital twins addressing data scarcity [21]. Cross-sensor correlation for behavioral/sensor deepfake detection in MCP involves analyzing interdependencies among multiple sensor inputs to uncover inconsistencies caused by synthetic emulations, such as mismatched physical signals in CPS agent communications. This technique leverages data fusion to verify coherence across modalities, detecting spoof- ing attacks that exploit the interoperability of MCP’s tools. Recent research has employed multimodal deep learning to enhance correlation analysis, achieving robust performance in CCTV and IoT environments where behavioral deepfakes pose significant privacy risks [121]. For example, anomaly-based frameworks using recurrent neural networks identify temporal mismatches, applicable to MCP’s dynamic data exchanges [363]. Advancements in cross-sensor correlation also integrate ex- plainable AI for transparent detection, aiding forensic analysis in MCP-secured CPS. Studies on federated learning models demonstrate improved generalization to unseen deepfakes by correlating sensor data with network metrics, reducing false positives in resource-constrained settings [394]. Limitations such as environmental noise are addressed through hybrid unsupervised approaches, thereby fostering reliable verifica- tion in multi-agent systems [187]. These methods strengthen MCP against behavioral manipulations, ensuring secure cyber- physical integrations. Physics-based anomaly detection in MCP for behav- ioral/sensor deepfakes utilizes fundamental physical laws to model expected system behaviors, flagging deviations from synthetic alterations in CPS sensor feeds or agent actions. This approach incorporates domain knowledge to validate MCP-transmitted data against physical constraints, countering stealthy emulations. Recent surveys classify physics-based methods as key for CPS security, with deep learning enhance- ments improving detection in industrial control systems [152]. For instance, digital twin-based models simulate physical processes for anomaly identification, effective against MCP supply-chain threats [366]. Furthermore, physics-based techniques in MCP emphasize hybrid models combining ML with physical simulations for enhanced accuracy, addressing complex attacks like sensor spoofing. Research on one-class learning frameworks shows superior performance on unbalanced CPS datasets, mitigating the impact of behavioral deepfakes [311]. Challenges such as computational overhead are mitigated through lightweight implementations, which support real-time MCP applications [232]. These innovations bolster CPS resilience, providing a solid foundation for securing AI agents under MCP protocols. E. Cross-Modal Research Synthesis This subsection consolidates the research analysis across all detection modalities, identifying common strengths, shared limitations, and unified research priorities for securing MCP- enabled CPS. 1) Common Strengths Across Modalities: Detection re- search has achieved notable progress applicable to CPS se- curity. Deep learning architectures, including CNNs, Vision Transformers, and self-supervised audio models such as A- SIST and wav2vec, establish strong baseline performance on benchmark datasets across visual, audio, and textual modalities [83], [201], [245]. Explainable AI techniques enable forensic localization of manipulated regions, improving both detection utility and operator trust [224]. Physics-grounded verification methods, including rPPG for visual content, ENF traces for audio, and cross-sensor correlation for behavioral data, provide authentication mechanisms that synthetic content cannot easily replicate [131], [187], [333]. Multi-modal consistency checks that fuse features across modalities demonstrate improved robustness compared to single-modality approaches [149], [291]. Proactive defenses such as AI watermarking enable tracing of synthetic content with minimal quality degradation [78]. 2) Shared Limitations: Despite these advances, four funda- mental limitations constrain deployment across all modalities: a) Generalization gap: Detectors trained on benchmark datasets exhibit significant performance degradation against unseen generators, novel attack variants, and in- the-wild conditions featuring compression, noise, and domain shift [76], [163]. b) Adversarial adaptation: The rapid evolution of genera- tive models, including Neural Codec Models and LLMs, consistently outpaces detection capabilities, rendering trained detectors obsolete within months [68], [370]. c) Computational constraints: Real-time detection on resource-constrained edge devices remains challenging; heavyweight models suitable for high accuracy are in- compatible with CPS latency and power requirements [22], [224]. d) Robustness to post-processing: Compression codecs, paraphrasing attacks, and channel variations significantly degrade detection accuracy across visual, audio, and textual modalities [201], [370]. 3) Unified Research Priorities: Addressing these shared limitations requires coordinated research efforts: • Generator-agnostic detection: Develop methods lever- aging fundamental artifacts (e.g., physics-based incon- sistencies, statistical fingerprints) rather than generator- specific signatures. • Lightweight edge architectures: Design efficient models (e.g., MobileNet variants, quantized transformers) that balance accuracy with CPS resource constraints. • Standardized MCP integration: Establish protocols for embedding detection and provenance verification into MCP tool pipelines without introducing unacceptable latency. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202626 TABLE XXII CROSS-MODAL DETECTION SYNTHESIS ModalityMaturityGeneral.RT EdgeKey Gap VisualMediumLowMediumGenerator shift AudioMediumLow-MedMediumCodec robustness TextualLowLowHighLLM evolution BehavioralLowMediumLowPhysics modeling • Cross-modal fusion: Exploit redundancy across visual, audio, textual, and behavioral channels to detect coordi- nated multi-modal attacks. • Continuous adaptation: Implement online learning and federated update mechanisms that enable detectors to evolve alongside generative model advances. 4) Comparative Summary: Table XXII summarizes detec- tion capabilities and gaps across modalities. The analysis confirms that detection mechanisms, while valuable for threat identification, cannot serve as standalone decision authorities in safety-critical CPS. Effective de- ployment requires integration with provenance verification, physics-based validation, and defense-in-depth architectures as detailed in Section VII. While detection techniques provide valuable signals for identifying potential deepfake activity, SENTINEL Phases 3 and 4 clarify that detection alone cannot serve as a decision authority in safety-critical CPS. Given adversarial adaptation and operational constraints, detection mechanisms should be treated as supporting components whose outputs must be cor- roborated by provenance, physical consistency, or system-level validation before influencing control or actuation decisions. VII. MITIGATION AND DEFENSE STRATEGIES The transition from static, single-turn Large Language Model (LLM) interfaces to autonomous agentic systems rep- resents a fundamental architectural shift in the cyber-physical landscape of 2025. In this paradigm, autonomous agents are no longer confined to experimental sandboxes but are actively orchestrating workflows, managing sensitive enterprise data, and executing control commands [241]. As identified by global surveys, 67% of organizations have deployed agentic AI in 2025 [235]. However, this proliferation has expanded the attacks across three critical interaction surfaces: the User- Agent (U-A) interface, the Agent-Agent communication layer, and the Agent-Environment (A-E) interaction plane. At the U-A surface, threats emerge from the blurring of boundaries between instruction and data. Prompt injection and multi-turn jailbreaking leverage the model’s semantic sensitivity to manipulate internal reasoning, leading to con- fused deputy scenarios where agents with elevated privileges perform unauthorized actions [241]. The A-A surface intro- duces risks of identity spoofing and secret collusion, where compromised agents may coordinate to fabricate consensus or bias collective decisions [185]. Finally, the A-E surface is vul- nerable to indirect prompt injection via poisoned documents or websites [156], as well as protocol-specific attacks such as tool masquerading and context poisoning within the Model Context Protocol (MCP) ecosystem [100]. The complexity of these interactions necessitates a shift from traditional perimeter-based security to a Defense-in- Depth strategy tailored for non-deterministic, autonomous workloads. Traditional controls are inadequate for governing agents that dynamically request permissions via emerging protocols like MCP and Agent-to-Agent (A2A) [241]. This section details a comprehensive mitigation strategy focused on provenance, contextual validation, and proactive architectural defenses. This section applies Phases 3 and 4 of the SEN- TINEL framework. Phase 3 evaluates candidate mechanisms against the established multi-dimensional fitness function and Phase 4 arranges selected mechanisms into a layered defense architecture. Defenses in the MCP ecosystem must specifically address protocol-level integrity. Provenance mechanisms extend to MCP-specific signatures and cryptographic audit trails to ensure the accountability of multi-modal agents and pre- vent tool spoofing [93], [179]. To counter deepfake injection and ensure content authenticity, proactive defenses employ semi-fragile watermarking strategies, such as FractalForensics, which localize manipulation within agent inputs [239], [349]. Additionally, robust training strategies incorporate physical- layer environmental fingerprints, such as Electric Network Frequency (ENF), to anchor virtual interactions to real-world grid dynamics; by utilizing digital twin testbeds to simulate ad- verse scenarios, these models are fine-tuned to detect deepfake anomalies that lack consistent physical synchronization [131]. Furthermore, data protection strategies utilize unlearnable per- turbations to prevent unauthorized model training on sensitive CPS data [391]. These technical controls are bolstered by human-centric measures, including user training on agentic threats and regulations mandating immutable provenance for election and enterprise security [321]. A. Provenance & Authentication The security of LLM-based agents utilizing the MCP and A2A frameworks have shifted from static perimeter defense to dynamic, identity-centric verification. The convergence of research and industry standards indicates that Zero-Trust au- thentication and cryptographic provenance are the primary mitigation strategies against emerging threats such as context poisoning, agent impersonation, and unverified tool execution [111], [318]. Traditional API keys are deemed insufficient for autonomous systems. The industry is moving toward Agent Identity, a paradigm where agents possess unique, verifiable identities like Decentralized Identifiers (DIDs) anchored to human principals or legal entities [197]. This human-to-agent binding ensures that every autonomous action carries a legal and technical chain of accountability, mitigating the risk of rogue agents denying their actions [219]. The mechanics of how agents communicate require rigorous safeguards. For the MCP, best practices mandate defense-in- depth architectures where the MCP server is isolated and inputs are strictly validated against schemas [225]. In multi- agent scenarios, threat modeling frameworks like MAESTRO emphasize strict input validation and authentication [141], while Zero-Trust principles are applied to inter-agent com- munication; every handshake must be mutually authenticated JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202627 Fig. 13. Agentic AI In CPS: Mitigation & Defense (Defense-In-Depth) (mTLS) and scrutinized for lateral movement risks [206]. The emerging Control Fabric middleware acts as a governance layer, intercepting agent requests to verify consent and enforce policy before any execution occurs [312], effectively neutral- izing the risks identified for innovative agentic applications [318]. Provenance mechanisms trace data origins to counteract deepfake-induced manipulations, often within immersive en- vironments like the metaverse. These systems leverage cryp- tographic metadata, blockchain for immutable audit trails, and multi-factor authentication using biometric signals [35]. To combat the black box nature of agent reasoning, provenance graphs and forensic retrieval engines can be deployed to map the lineage of data from training to inference [230], [277]. Ex- plainable AI (XAI) enhances transparency in these verification checks [60], while zero-trust frameworks utilizing blockchain and physical unclonable functions reduce risks in the hardware supply chain [176]. Extensions address privacy concerns via selective disclosure mechanisms [47] and federated learning architectures that protect sensitive data during model training [325]. Cognitive security approaches integrate behavioral anal- ysis to identify sophisticated influence operations [61], with recent blockchain architectures achieving robust guarantees in data integrity and non-repudiation [244]. Studies also highlight the incorporation of Electric Network Frequency (ENF) as a physical anchor for authentication, providing a defense against spoofing in multimedia streams [237]. 1) Digital Watermarking: Digital watermarking acts as a critical security layer by embedding imperceptible, algorithmic signatures into the multimodal data such as retrieval docu- ments, images, and system prompts that MCP servers transmit to AI agents [195]. By employing techniques like In-Context Watermarking, organizations can inject unique identifiers di- rectly into the prompt context without requiring access to the model’s internal weights, ensuring that the generated output retains a verifiable trace of its source [203]. This approach effectively secures Retrieval-Augmented Generation (RAG) workflows by acting as a knowledge watermark, allowing system administrators to detect if proprietary knowledge bases have been extracted or utilized without authorization [214]. Furthermore, some frameworks extend these protections to image-based context, ensuring that copyright claims can be algorithmically verified even after the content has been pro- cessed and regenerated by a multimodal model [69]. Beyond ownership, these cryptographic signatures establish a chain of trust, enabling systems to cryptographically distinguish authorized AI-generated content from external or potentially malicious human inputs [72], [171]. 2) Blockchain: Blockchain technology secures the MCP based AI agents in the CPS ecosystem by establishing a decentralized, immutable trust layer that ensures data integrity, verifiable identity, and accountability. While standard MCP architectures rely on traditional security mechanisms like OAuth2 and TLS for scalability [252], high-assurance CPS environments require enhanced protection against data poi- soning. By integrating cryptographic identity frameworks and smart contracts, the architecture enforces strict access control, ensuring that only verified agents can invoke specific MCP tools or modify CPS states [71], [279]. Furthermore, the novel concept of Model Context Contracts leverages blockchain to allow LLMs to interact deterministically with smart contracts, automating governance and creating an auditable trail of agent decisions [44], [159]. This immutable history complements ar- chitectural defenses against prompt injection and unauthorized resource access identified in recent MCP security audits [290], providing a forensic layer that traditional firewalls cannot offer. 3) C2PA standards: The C2PA standard utilizes crypto- graphic manifests to verify media payloads and edit histories, JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202628 providing assurances of authenticity for AI-generated content [43] while maintaining utility in multi-modal environments such as audio and video [8]. To enhance resilience, recent frameworks integrate blockchain for immutable provenance [57] and employ post-quantum cryptographic tools to secure verification against future threats [143]. Privacy is further strengthened through policy enforcement mechanisms like Self-Sovereign Identity, which allow creators to prove author- ship without exposing sensitive data [95]. While these tech- nical standards support global standardization efforts [41] and aid human discernment of fabricated media [112], Chan et al. [62] argue that for autonomous agents, technical provenance must be paired with Identity Binding infrastructure. This links agent actions to legal identities to establish accountability, acknowledging that while infrastructure can govern interac- tions, specific defenses against adversarial attacks like prompt injection remain an open challenge. 4) Tradeoffs in Provenance & Authentication: Applying the Phase 3 Selection framework reveals distinct tradeoffs between performance and security depth. Digital Watermarking offers excellent computational performance and low OPEX, making it ideal for edge-compatible verification [228]. However, it faces integration challenges regarding black-box prompt in- jection risks [203]. Blockchain provides the highest integrity and non-repudiation but incurs significant latency and stor- age costs, limiting its use to high-assurance logging rather than real-time filtering [159]. C2PA Standards balance high cryptographic certainty with minimal verification overhead, though they require significant effort to standardize headers and sanitize metadata for privacy [41], [95]. TABLE XXIII PHASE 3 SELECTION: PROVENANCE & AUTHENTICATION CONSTRAINTS ConstraintDigital Watermarking BlockchainC2PA Standards DetectionHigh robustness toparaphrasing [214] . Immutable audit trail; high integrity [159]. Cryptographic certainty [95] ComputationalZeromodel training overhead [203]. Highlatency (consensus dependent) [159], [279]. High generation but low verifica- tion cost [143] IntegrationLow;black-box prompt injection [203]. Complex; Smart Contract deploy- ment [44]. High [43]. OperationalLow maintenance; algorithmic generation [171] High storage bur- den for history [159]. High [62]. PrivacyLow; impercepti- ble markers [72]. Medium; requires hashing layers [279]. Low [143] CostLow OPEX [69].High (Gas/Compute) [159]. Medium [95] 5) Defense Architecture Design Choices: In the Phase 4 Architecture, these technologies are stratified to maxi- mize defense-in-depth. C2PA and Watermarking serve as the Perimeter Tier, rejecting unauthenticated inputs before they reach agent logic [80]. The Detection Tier analyzes watermarks for tampering artifacts. Blockchain is reserved for the Adaptation and Response Tiers, providing an immutable forensic log for incident investigation. This ensures a verifiable trace remains for future model adaptation even if the perimeter is breached [290]. TABLE XXIV PHASE 4 ARCHITECTURE: PROVENANCE & AUTHENTICATION TIERS Defense Tier Digital Watermarking BlockchainC2PA Standards PerimeterBlockinputs lackingvalid watermarks [69]. None.Rejectpayloads withbroken signature chains [57]. DetectionAnalyze watermarks fortamper- ing/stripping [214]. Verify transaction hashagainst ledger [159]. Validate provenance history [95]. ResponseAutomated content rejection [171]‘ Log incident data immutablyfor forensics [279]. Revoke keys of compromised publishers [62]. AdaptationUpdate embedding algosagainst removalattacks [195]. Audit smart con- tracts for logic gaps [279]. Update decentralized reputation scores andtrustlists [95] B. Multi-factor & Context Validation Multi-factor mechanisms integrate environmental awareness and biometric fusion to counter deepfake threats and imper- sonation [378]. AI-driven adaptive Multi-factor authentication models assess risk in real-time [46], with cross-modal sensor checks achieving 96.3% verification accuracy [378]. To re- inforce security, event processing frameworks analyze system behavior for anomaly detection [339]. Privacy-preserving tech- niques include federated learning [46] and layered geolocation verification [216]. Furthermore, temporal context factors en- hance spoofing detection by approximately 30% [14], helping systems dynamically adjust encryption to address adaptive adversaries [295]. 1) Sensor Fusion: Sensor fusion aggregates diverse data sources to create unified contexts, enabling the detection of inconsistencies such as mismatched Electric Network Fre- quency (ENF) patterns [131]. To process these complex inputs, AI-enhanced architectures employ feature fusion strategies; for example, hybrid CNN-LSTM models effectively capture spatio-temporal anomalies [334], while Multi-Graph Attention Networks integrate global and local features to detect forged traces [65]. In Cyber-Physical Systems (CPS), integrating environmental fingerprints like ENF has been shown to reduce false positives by over 15% [131]. Furthermore, proactive fusion disrupts threats via physical anchors [132], while edge computing optimizations minimize inference latency for real- time verification [169]. Finally, adversarial training strength- ens these fusion models by enhancing generalization against unseen attacks [298]. 2) Out-of-band confirmation: In Cyber-Physical Systems (CPS), the perception layer of an agent is often susceptible to sensor out-of-band (OOB) vulnerabilities. As systematized by Xiao et al., these vulnerabilities occur when signals from a different physical modality, such as electromagnetic interfer- ence, ultrasound or lasers, induce malicious measurements via out-of-range or cross-field energy conversion pathways [361]. These physical layer exploits allow an attacker to deceive an JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202629 agent about its environment, for example, by spoofing a clear path for an autonomous robot or falsifying thermal readings in industrial equipment. To mitigate these physical spoofing risks, security archi- tectures are increasingly adopting communication out-of-band confirmations. These mechanisms leverage cryptographic vi- sual channels, such as dynamic QR codes, to authenticate high-risk actions independent of the compromised primary network [81]. While traditional biometrics previously offered high effectiveness for user binding, recent surveys warn that Generative AI is actively eroding these zero-trust assumptions by synthesizing identities that bypass standard liveness checks [364]. Consequently, modern defenses must integrate context- aware opportunistic sensing [287] and hardware-anchored Physical Unclonable Functions (PUFs) [12] to validate the physical integrity of the device itself. Furthermore, in im- mersive interfaces such as Extended Reality (XR), privacy- preserving architectures are being developed to bolster re- silience by decoupling authentication verification from sen- sitive biometric user data [124]. 3) Tradeoffs in Multi-factor & Context Validation: Under Phase 3, Sensor Fusion is selected for high-criticality zones despite high Computational and Cost loads because of its su- perior Detection Effectiveness against environmental spoofing [334].Out-of-Band Confirmation is prioritized as a fallback due to its low footprint, although its high Operational Burden due to user friction limits its use to high-privilege actions [364]. TABLE XXV PHASE 3 SELECTION: MULTI-FACTOR & CONTEXT VALIDATION CONSTRAINTS ConstraintSensor FusionOut-of-Band Confirma- tion DetectionHighaccuracy[65], [131]. Deterministic verification [81]. ComputationalHigh GPU/CPU load for real-time fusion [334]. Low [287]. IntegrationComplex; requires tight ENF/grid sync [131]. High; requires hardware modification [12]. OperationalModerate; sensor cali- bration needed [361]. High friction; user inter- vention [364]. PrivacyHigh;aggregatesraw env. data [378]. Medium; selective dis- closure capable [124]. CostHigh (Hardware) [361].Low [287]. 4) Defense Architecture Design Choices: In the Phase 4 Architecture, Sensor Fusion operates continuously in the De- tection Tier, flagging inconsistencies like mismatched times- tamps [131]. Out-of-Band Confirmation serves as a circuit breaker in the Response Tier, triggered only when detection thresholds are crossed [364]. This optimizes the system by using automated computation for monitoring while reserving high-latency human verification for escalation. C. Robust Model Training Robust training enhances resilience in Cyber-Physical Sys- tems (CPS) through decentralized approaches. For instance, [5] proposes a Federated Quantum Machine Learning (FQML) framework that utilizes multi-agent collaborative strategies and trimmed-mean aggregation to specifically mitigate data TABLE XXVI PHASE 4 ARCHITECTURE: MULTI-FACTOR & CONTEXT VALIDATION TIERS Defense Tier Sensor FusionOut-of-Band Confirma- tion PerimeterNone.None. DetectionDetectinconsistencies (e.g., Audio vs. Video ENF) [131]. Verifyhigh-privilege commandsvia Visual/QRChannel [81]. ResponseTrigger failsafe if sen- sor mismatch exceeds threshold [216]. Block execution pending user approval [364]. AdaptationRetrain fusion models on new spoofing patterns [298]. Adjust confirmation fre- quency based on threat level [46]. poisoning and Byzantine attacks. Anomaly detection strate- gies focus on data-centric methods [242] and discrete event dynamics [137]. To support these strategies, techniques in- clude specific designed Machine Learning and Deep Learning models for Cyber-Physical systems (CPS) security constraints [128]. Furthermore, integrating cryptographic enhancements ensures privacy-preserving computation [269], while hybrid approaches employing robust reasoning mechanisms allow autonomous systems to make valid decisions even with in- complete or contradictory data [122]. These methodologies optimize training to mitigate emerging threats, including the misuse of Generative AI and adversarial deepfakes, which require advanced detection frameworks to ensure information integrity [325], [388]. 1) Adversarial Training: Adversarial training enhances de- fense robustness by exposing AI agents to simulated deepfake perturbations during the learning process [181]. This approach is increasingly integrated into diagnostic behavior analysis to prevent data intrusions in CPS [299]. Emerging developments extend these concepts to next-generation environments, uti- lizing adversarial testing and reinforcement learning within AI frameworks for 7G-enabled virtual platforms [151]. Fur- thermore, recent literature categorizes these disruptive attacks [19] and highlights the role of adversarial learning in industrial data-driven innovations [294]. Comprehensive reviews empha- size that deploying these adversarial defense strategies is a critical component of the cyber kill chain for mitigating AI- driven threats [161]. 2) Environmental fingerprints: To enhance model robust- ness against AI-generated deepfakes, researchers are increas- ingly integrating physical invariants, such as the Electric Network Frequency (ENF), into training data. Unlike purely semantic features, ENF provides a stochastic environmental fingerprint that is difficult for generative models to replicate accurately. Korgialas et al. introduced robust estimation frame- works using LAD regression to ground audio data in physical grid signals [174]. To ensure detection models remain effective even when input data is degraded, architectures like Multi- HCNet have been developed to capture high-order harmonics [194]. Furthermore, integrating novel signatures like Electri- cal Network Voltage (ENV) allows models to learn precise location constraints, improving resilience against spoofing in IoT and Metaverse environments [131], [385], provided that signal-to-noise ratios remain sufficient for extraction [140]. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202630 3) Tradeoffs in Robust Model Training: The Phase 3 Selec- tion highlights resource tradeoffs. Adversarial Training offers high Detection Effectiveness but incurs significant Computa- tional and Cost penalties during offline training [181]. Envi- ronment Fingerprints via ENF are selected for their efficiency, offering high detection capability with low Computational requirements by leveraging existing grid infrastructure, though they require management of Privacy risks regarding location [174]. TABLE XXVII PHASE 3 SELECTION: ROBUST MODEL TRAINING CONSTRAINTS ConstraintAdversarial TrainingEnvironment Fingerprints DetectionHigh robustness to per- turbation [65]. Highspecificityfor physical locality [385] ComputationalVery high offline training cost [181]. Low; signal processing only [174]. IntegrationHigh; Complex parame- ter tuning [334]. Moderate; requires grid reference DB [131]. OperationalHigh; frequent retraining required [19]. Low; automated extrac- tion [138]. PrivacyLow;withFederated Learning [294]. Medium; Location leak- age risk [131], [288]. CostHigh GPU OPEX [19].Low; uses existing grid [385]. 4) Defense Architecture Design Choices: In the Phase 4 Architecture, Adversarial Training underpins the Adaptation Tier, ensuring models are hardened through continuous retrain- ing as new techniques emerge [161]. ENF operates within the Detection Tier, providing a real-time, physics-based liveness check to validate audio-visual streams before they influence agent decisions [379]. TABLE XXVIII PHASE 4 ARCHITECTURE: ROBUST MODEL TRAINING TIERS Defense TierAdversarial TrainingEnvironment Fingerprints PerimeterNone.None. DetectionClassifier models detect subtle adversarial noise [181] Validate media times- tamp against Grid Fre- quency (ENF) [140] ResponseAutomated RL-based in- cident response [151]. Flag content with mis- matched ENF as Syn- thetic [138] AdaptationIncorporate new deep- fake samples into train- ing sets [161]. Updatereference databases with regional grid logs [131]. D. Proactive Defenses Proactive defenses neutralize threats preemptively through strategies like identity watermarking for facial verification [390] and AI-driven predictive threat analysis [325]. In the context of Multi-Agent Systems (MAS), defenses focus on prevention and resiliency mechanisms, such as trust and reputation modeling [246]. Federated approaches distribute defenses across nodes [99]. Extensions include cognitive se- curity architectures that integrate human-AI scrutiny [61] and blockchain-enhanced monitoring [229]. Furthermore, system robustness against resource failures can be evaluated using discrete timed model-based optimization [137]. 1) Disrupting Deepfake creation: Disrupting deepfake cre- ation interferes with generative processes via training contam- ination [324], latent space optimizations [380], and adversarial watermarking [196]. Techniques like sticky adversarial signals [395] disrupt generation by resisting removal attempts, while learnable hidden faces [188] facilitate proactive exposure of tampered media. Recent advancements include differential- aware networks [134], diffusion-based watermarking [322], and neural watermarking for speaker identity protection [104]. 2) Traceability: Understanding the root cause of an agent’s failure is critical for accountability. Mechanistic interpretabil- ity involves dissecting LLM internals to identify specific attention heads or layers responsible for harmful behavior. In multi-agent systems, this allows defenders to trace how rep- resentations, such as malicious instructions or steganographic messages, propagate between agents [185]. To operationalize this, system-level anomaly detection frameworks now employ Dynamic Execution Graphs to map multi-step interaction chains [135]. These graphs enable multi-point failure attribu- tion, a capability essential for detecting emergent collusion where agents covertly coordinate to bias collective decisions [227]. Technical traceability relies on establishing a verifiable link between content and its origin. Proactive methods utilize forensic watermarking for end-to-end tracking, such as the FRW-TRACE framework for biometric data [268] and Pseudo- Zernike transform watermarking for face swapping detection [182]. Beyond watermarking, blockchain-based architectures provide immutable provenance; for instance, the VeriTrust framework integrates Self-Sovereign Identities (SSI) to bind content to creator identities [95]. Complementing these proactive measures are reactive attri- bution techniques. These include source attribution via camera fingerprints [307] and the analysis of generative artifacts using Vision Transformers [31]. To ensure these technical findings are actionable for human analysts, Explainable AI (XAI) methods are increasingly integrated to demystify black- box detection models [267]. Finally, comprehensive forensic surveys highlight the need to analyze the entire lifecycle of synthetic media to counter the rising impostor bias and user cynicism regarding content authenticity [25], [199]. 3) Tradeoffs in Proactive Defenses: In the Phase 3 Selec- tion, a tradeoff analysis favors Disrupting deepfake creation for its privacy preservation capabilities, despite its higher integration costs. By optimizing the latent space of generative models, it provides a preemptive shield with moderate compu- tational overhead [380]. Conversely, Traceability is critical for accountability but suffers from a high operational burden due to log storage requirements and analysis latency, which limits its real-time utility in fast-acting CPS environments [95]. 4) Defense Architecture Design Choices: In the Phase 4 Architecture, disrupting deepfake functions primarily at the Perimeter Tier, attacking adversarial capabilities before they penetrate the CPS boundary [322]. Traceability anchors the Adaptation and Response Tiers, enabling the attribution of breaches to specific actors and facilitating threat model updates based on forensic evidence [182]. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202631 TABLE XXIX PHASE 3 SELECTION: PROACTIVE DEFENSE CONSTRAINTS ConstraintDisrupting DeepfakeTraceability DetectionPre-emptive neutralization [324]. Forensicattribution [268]. ComputationalHigh; iterative diffusion sampling [322] Highanalysislatency [267]. IntegrationHigh;requiressource media mod [196]. Moderate; logging infra needed [307]. OperationalLow; automated injec- tion [395]. High storage for logs [95]. PrivacyHigh; protects biometric data [188]. Low; de-anonymization risk [95]. CostLow [134].Medium [95]. TABLE X PHASE 4 ARCHITECTURE: PROACTIVE DEFENSE TIERS Defense TierDisrupting DeepfakeTraceability PerimeterInject adversarial signal todisruptgeneration [395] None. DetectionEmbed learnable hidden face as tamper indicator [188] None. ResponseDiffusion-based watermarking for source tracking [322] Trace attack origin via embeddedforensics [182] AdaptationLatent space optimiza- tion for specific image defense [380] Update attribution mod- els based on new attack vectors [31]. E. Human & Policy Measures Human and policy measures complement technical mech- anisms of defense in CPS through governance, education, and strategic foresight. While technical controls secure the digital-physical interface, human operators and organizational policies dictate how these tools are deployed and monitored. To improve organizational readiness against social engineering in critical infrastructure, Pedersen et al. propose the PREDICT framework, which integrates definitive policy direction with targeted employee education [256]. Beyond the organizational level, Mubarak et al. highlight the necessity for legislative action and public awareness campaigns to mitigate the societal and political destabilization caused by deepfakes [228]. From a strategic perspective, reactive measures are insuffi- cient for CPS security. Saeed et al. argue that Cyber Threat Intelligence must be leveraged to navigate the complexity of threats where human error remains a primary vulnerability [286]. Furthermore, Almahmoud et al. suggest shifting toward proactive defense by utilizing machine learning to forecast long-term cyber threat trends. This foresight enables organiza- tions to make strategic investments in mitigation technologies before threats escalate to physical harm [16]. Ultimately, preserving trust in these systems requires interdisciplinary col- laboration among technologists, legislators, and psychologists [88]. 1) User Training: In safety-critical environments, humans remain both a vulnerability and a critical defense layer. Train- ing strategies must therefore address two distinct groups: the operators who manage CPS interfaces and the developers who build them. Operator Awareness deals with effective training for op- erators that must combine real-time threat recognition with strict verification procedures. As technical detection tools may lag behind generation capabilities, Romero et al. emphasize that media literacy empowers individuals as a first line of defense against deepfake-induced manipulation [281]. Inter- active training tools, such as platforms challenging users to distinguish between real and synthetic faces, have proven effective in sharpening these analytical skills [231]. Addition- ally,Generative AI (GenAI) driven simulations are now em- ployed to construct authentic attack scenarios, strengthening the response speed of cybersecurity specialists and reducing manual workloads during incident response [13]. Under secure development practices, the integration of GenAI into the Software Development Lifecycle necessitates a shift from manual coding to AI-augmented oversight. While GenAI tools assist in code generation and debugging, they introduce risks such as logic errors and insecure code pat- terns. Al-Hashimi et al. note that while GenAI can automate vulnerability scanning, human-in-the-loop review processes are essential to validate AI-generated outputs against security standards like OWASP and NIST [13]. Given the arms race between attack and defense, continuous education on evolving threat landscapes is vital to maintain the integrity of the SDLC [1]. 2) Regulations: Regulatory frameworks are evolving to mandate accountability and transparency in AI agent deploy- ment. The landscape includes hard regulations like EU AI Act and soft law frameworks such as NIST AI RMF, ISO 42001 that collectively establish governance structures for trustwor- thy systems. These regulations aim to mitigate liability risks by establishing legal frameworks for content provenance and agent behavior [8], [41]. In the context of CPS, harmonized standards are increasingly critical to enforcing privacy and data protection across interconnected devices [95], [131]. However, human detection capabilities are often insufficient to identify sophisticated deepfakes [112]. Consequently, scholars argue for adaptive legislation driven by public sentiment [59] and enforced through technical cryptographic standards [143] to safeguard the digital trust ecosystem. 3) Standards: Standardization ensures interoperability and security across the multi-vendor ecosystems typical of CPS. While protocols like the MCP and A2A facilitate agent inter- action and collaboration, current analyses indicate a lack of na- tive security standards, leaving systems vulnerable to naming attacks and context poisoning [263]. To address these gaps, the industry is moving toward robust evaluation frameworks. The OWASP Agentic Security Initiative has established a reference threat model to identify risks such as intent manipulation and memory poisoning [317]. Complementing this, the Agent Security Bench (ASB) pro- vides a framework for benchmarking attacks—including direct prompt injection across diverse agent scenarios [384]. Beyond software agents, standardization must extend to the physical domain. Syllaidopoulos et al. emphasize the need for inter- national harmonization to manage high-risk AI deployments [325]. Emerging technical frameworks also propose standard- ization paths, Odyurt et al. suggest behavioral passports as a baseline for anomaly detection in industrial CPS [242], while Babbar et al. propose cryptographic frameworks utilizing federated learning to standardize secure data transmission in JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202632 vehicular environments [38]. F. Emerging research directions Emerging defense strategies increasingly leverage the in- tersection of physical constraints and behavioral analysis to secure AI agents. Physical Layer Authentication has evolved beyond static checks. Strategies now employ multi-modal fusion to verify device integrity [128], [352] and utilize environmental anchors, such as Electric Network Frequency (ENF), to validate digital twins against real-time grid fluctua- tions [131], [132]. To secure the agents themselves, research is shifting toward Behavioral Fingerprinting. Odyurt et al. propose behavioral passports that model the execution phases of CPS processes, allowing for the detection of anomalies in agent behavior that traditional signature-based methods might miss [242]. Furthermore, Decentralized and Explain- able Frameworks are gaining prominence. Hybrid models are integrating Deep Learning with blockchain to ensure data provenance [300], while Federated Learning is being adopted to train robust models without exposing sensitive edge data [99]. Finally, to address the black box nature of AI agents, Explainable AI (XAI) is emerging as a critical requirement for policy compliance and human-agent trust in high-stakes counter-terrorism and cybersecurity operations [325]. As LLM-driven agents increasingly integrate with external environments via MCP, the attack surface expands beyond simple injection to include privilege escalation and insecure tool execution. Emerging defenses are shifting toward archi- tectural isolation and granular access control. Formal security design patterns, such as the Dual LLM architecture have been proposed that separates the untrusted planning agent from the privileged execution agent to enforce information flow control and prevent data exfiltration [48]. To address risks specific to MCP servers, B ̈ uhler et al. introduce AgentBound, a policy en- forcement framework that intercepts tool calls at the protocol level to prevent unauthorized privilege escalation and resource abuse [56]. Furthermore, dynamic evaluation frameworks are being developed to stress-test these agentic logic flows against multi-turn adaptive attacks, moving beyond static benchmarks to ensure resilience in continuous interaction environments [381]. VIII. CASE STUDY: AUTHENTICATING SMART GRID DIGITAL TWINS To demonstrate how the principles surveyed in this work apply in a concrete CPS context, we present a case study based on ANCHOR-Grid, a smart grid DT authentication framework that leverages real-world environmental data to secure cyber- physical representations [132]. This case study shows how the SENTINEL framework can be operationalized to address deepfake threats against DTs and underscores the role of physical anchors in CPS security. In smart grid systems, DTs are virtual replicas of phys- ical grid components used for simulation, monitoring, and decision support. Their fidelity to the real grid is critical: attackers who manipulate DT inputs or states can disrupt operational decision-making, thereby creating reliability and safety hazards. ANCHOR-Grid was developed to confront this risk by authenticating DTs against environmental signals that are difficult for adversaries to forge. Specifically, ANCHOR- Grid uses ENF as a real-world anchor. Because the ENF reflects the dynamics of the actual power system and fluctuates unpredictably with the physical grid’s behavior, it provides a robust ground truth against which DT can be verified. This approach shifts the security goal from detecting anomalous data sequences to validating the authenticity of the DT through physical signal congruence [132]. Figure 14 illustrates how real-world environmental anchors can be integrated with the SENTINEL framework to secure smart grid DTs against deepfake threats. A layered CPS archi- tecture is shown in which ENF signals from the physical power grid are used by the ANCHOR-Grid authentication layer to validate DT states before they are consumed by AI-enabled analytics and future MCP-enabled agents. Deepfake and replay attacks targeting cyber representations are explicitly shown to bypass traditional detection mechanisms but to be intercepted through environment-grounded verification. The overlay of SENTINEL phases highlights how threat characterization, constraint analysis, defense selection, defense-in-depth, and continuous validation map directly onto system components, illustrating why physics-anchored trust is essential for securing AI agents in safety-critical CPS. 1) SENTINEL Phase 1–2: Threat Characterization and Attack Surface Identification.: Applying the first two phases of SENTINEL, ANCHOR-Grid characterizes threats not merely as abstract data integrity violations, but as context-dependent risks shaped by timing, trust, and physical criticality. The attack surface extends beyond individual sensors to include the entire agent–environment interface, where digital repre- sentations of grid state influence downstream decision-making. In this setting, the adversary’s objective is not necessarily to introduce obviously anomalous data, but to generate measure- ments that are internally consistent and physically plausible, thereby evading detection while inducing incorrect grid state estimates. This threat model aligns closely with deepfake- style attacks surveyed earlier in this paper, where fidelity and contextual realism are leveraged to bypass trust assumptions. 2) SENTINEL Phase 3: CPS Constraints and Feasibility Analysis.: A defining characteristic of smart grid CPS is the presence of strict timing, reliability, and safety constraints that fundamentally limit which security mechanisms are feasible in practice. In the context of DT authentication, these con- straints act as a hard filter in SENTINEL Phase 3, eliminating many defenses that are effective in purely cyber or offline settings. Smart grid monitoring and control operate across multiple time scales, ranging from sub-second dynamics for frequency stability and protection functions to second-level and minute-level windows for state estimation and operational analytics. Security mechanisms introduced at the DT interface must therefore satisfy bounded latency and low false-positive requirements, as delayed or spurious alarms can trigger unnec- essary mitigation actions, operator intervention, or degraded situational awareness. ANCHOR-Grid explicitly evaluates feasibility under such CPS constraints by demonstrating that ENF-based authenti- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202633 Fig. 14. End-to-end integration of ANCHOR-Grid with the SENTINEL framework for securing smart grid DTs against deepfake threats. cation remains effective under realistic network latency and noise conditions, rather than assuming idealized communi- cation channels. Reported results show that authentication performance is maintained under network delays on the order of O(10 2 ms), a range compatible with DT validation and monitoring workflows but not with ultra-fast protection relays. This distinction is critical: ANCHOR-Grid is intentionally positioned as a trust-validation layer for cyber-physical rep- resentations, not as a replacement for real-time protection logic. Table XXXI summarizes representative magnitudes for key CPS constraints derived from the ANCHOR-Grid eval- uation. These values demonstrate that ENF-based authenti- cation remains effective under realistic latency, noise, and replay conditions, while maintaining acceptable detection and false positive rates. This contrasts sharply with heavyweight detection mechanisms, which typically require significantly more computation and are harder to tune under CPS timing constraints. From a computational perspective, ENF extraction and correlation rely on lightweight signal-processing operations whose cost scales modestly with window size, making them suitable for edge or near-edge deployment. In contrast, heavy- weight deepfake detection models that require continuous inference or retraining impose computational and energy over- heads that are difficult to justify under CPS resource con- straints. These quantitative considerations reinforce a central insight of this survey: CPS feasibility must be evaluated alongside adversarial robustness, and defenses that ignore timing, false-positive cost, or resource limitations are unlikely to be deployable in operational settings. 3) SENTINEL Phase 4: Defense Selection via Environmen- tal Anchoring.: Guided by the above constraints, ANCHOR- Grid selects environmental anchoring as its primary defense mechanism. Specifically, it leverages the ENF signal, a glob- ally observable, physics-driven characteristic of power grids, as a real-world anchor to authenticate DT inputs. Because ENF fluctuations are inherently tied to physical grid dynamics and cannot be arbitrarily forged without controlling large portions of the grid, they provide a form of provenance that is inde- pendent of sensor trustworthiness or communication integrity. This shifts the security objective from detecting malicious data to verifying consistency between digital representations and immutable physical signals, a design choice that aligns directly with the principles advocated in this survey. 4) SENTINEL Phase 5: Defense-in-Depth Architecture.: ANCHOR-Grid integrates environmental anchoring within a broader defense-in-depth architecture. Cryptographic protec- tions and secure communication channels serve as supporting layers, while ENF-based verification provides a grounding layer that detects desynchronization even when upstream de- fenses fail. This layered approach limits the blast radius of compromised components and avoids single points of failure, an essential requirement in safety-critical CPS. Importantly, the anchoring mechanism operates passively and incurs mini- mal computational overhead, preserving real-time performance while enhancing trustworthiness. This contrasts with detection- centric pipelines that require continuous retraining and adap- tation to evolving attack strategies. 5) SENTINEL Phase 6: Validation and Continuous Adap- tation.: The final phase of SENTINEL emphasizes ongoing validation rather than one-time deployment. In ANCHOR- Grid, the continuous comparison between observed ENF sig- nals and DT inputs enables persistent monitoring of system integrity. Deviations beyond expected tolerances signal poten- tial compromise or drift, triggering mitigation or investigation workflows. This validation strategy highlights a key insight for securing AI-enabled CPS: resilience emerges not from perfect detection, but from continuous grounding in physical reality combined with adaptive response mechanisms. While ANCHOR-Grid did not explicitly deploy LLM agents or MCP-mediated tools, it provides a concrete template for securing future AI agents operating DTs. As AI agents in- JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202634 TABLE XXXI QUANTITATIVE CPS CONSTRAINTS AND DESIGN IMPLICATIONS IN THE ANCHOR-GRID CASE STUDY (SENTINEL PHASE 3) Constraint DimensionRepresentative MagnitudeDesign ImplicationRole of ANCHOR-Grid Network latency toleranceAuthentication accuracy remains be- tween 99.9% and 95% under network latencies ranging from<5 ms to 200 ms Security mechanisms must tolerate re- alistic communication delays without degrading integrity verification ENF-based authentication remains ro- bust under latency levels typical of smart grid telemetry networks Detection accuracy vs. false positives (deepfake attacks) 99.8% detection with 0.2% false pos- itives under sparse attacks; 97.5% de- tection with 1.5% false positives under high-frequency attacks False alarms must be tightly bounded to avoid unnecessary operator intervention or destabilizing control actions Physics-grounded verification main- tains high accuracy while limiting false positives under increasing attack inten- sity Replay attack resilience94% detection for 5 s-old ENF signa- tures; 98.5% detection for 120 s-old sig- natures Authentication must detect stale or replayed measurements within opera- tionally relevant time windows Temporal correlation of ENF signals enables reliable replay attack detection without model retraining Noise robustness96.5% detection with 5% injected noise; 88% detection with 20% injected noise Environmental noise must not invali- date authentication or cause spurious alarms Correlation-based ENF matching toler- ates moderate noise typical of real grid measurements Computational feasibilityLightweight signal-processing opera- tions; no continuous deep neural infer- ence required Heavyweight ML-based detection is difficult to deploy at scale under CPS resource constraints ENF extraction and correlation are suit- able for edge or near-edge deployment Operational response timeSeconds-levelvalidationcompatible with digital twin monitoring, but not millisecond-scale relay protection Security layers must be positioned ap- propriately within CPS control hierar- chies ANCHOR-Grid functions as a trust- validation layer for digital twins rather than a real-time protection mechanism creasingly rely on MCP to access external tools, data sources, and shared contexts, the risk of synthetic yet plausible inputs grows substantially. In such settings, environmental anchors, such as those employed in ANCHOR-Grid, can serve as trust substrates that inform agent reasoning, constrain decision- making, and validate tool outputs before they influence physi- cal actuation. This case study, therefore, demonstrates how the principles articulated in the SENTINEL framework can be re- alized in practice and why environment-grounded verification mechanisms are likely to be indispensable for trustworthy AI agents in CPS. Taken together, the analyses in Sections IV-VII and the end-to-end CPS case study demonstrate that securing AI agents in CPS requires a shift from isolated, technique-centric defenses to lifecycle-aware system design. The SENTINEL framework integrates threat characterization and attack surface identification with feasibility filtering, defense selection, and continuous validation, ensuring that security mechanisms are evaluated not only for adversarial robustness but also for compatibility with timing, safety, and trust constraints imposed by the physical world. Across deepfake threat modalities, detection techniques, and mitigation strategies, a consistent conclusion emerges: defenses that ignore CPS constraints or lack environment-grounded provenance cannot be relied upon as decision authorities. By explicitly connecting threat analysis, quantitative feasibility, and defense-in-depth archi- tectures, SENTINEL provides a unifying logic that explains why physics-anchored trust mechanisms, such as those illus- trated in the ANCHOR-Grid case study, are indispensable for trustworthy AI agents operating in real-world CPS. IX. OPEN CHALLENGES AND FUTURE DIRECTIONS Open challenges in MCP for CPS encompass the need for adaptive protocols that evolve in response to advancements in deepfakes, ensuring the seamless integration of detec- tion mechanisms without compromising tool interoperability or agent performance [99]. Identity fragmentation remains a persistent issue, where fragmented authentication across MCP servers heightens vulnerability to deepfake injections, necessitating unified identity management to maintain trust in agent-environment interactions [242]. Real-time constraints in CPS demand low-latency defenses; however, current MCP implementations struggle with computational overhead from provenance tracking, highlighting the need for optimized message formats that embed security without inflating data streams [137]. Furthermore, the interplay between privacy and traceability poses dilemmas, as provenance requirements may expose sensitive context data, requiring novel encryption schemes tailored to MCP’s JSON-RPC structure [128]. Future directions for MCP in CPS include exploring hybrid architectures that fuse blockchain with protocol layers for decentralized provenance, enhancing resilience against supply- chain attacks while preserving modularity [96]. Multi-modal detection fusion within MCP could leverage cross-verification of sensor inputs, addressing gaps in single-modality defenses against sophisticated deepfakes [122]. Emphasis on edge- compatible implementations will drive the development of lightweight MCP variants, incorporating AI-optimized com- pression to enable real-time operation on resource-constrained devices [325]. Collaborative frameworks between academia and industry should standardize MCP extensions for the ethical use of AI, fostering interoperability while mitigating risks associated with generative content [348]. These pathways aim to fortify MCP as a secure backbone for CPS amid escalating threats. Advancing MCP research also involves formal verification of protocol behaviors under adversarial conditions, as well as modeling deepfake scenarios to predict and preempt exploits in agent-tool exchanges [99]. Privacy-preserving enhancements, such as zero-knowledge proofs integrated into MCP capability descriptions, could reconcile authentication needs with data protection mandates [128]. Scalability tests on large-scale JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202635 CPS simulations will inform optimizations for high-volume interactions, ensuring MCP sustains performance as generative threats proliferate [300]. Ultimately, interdisciplinary efforts must prioritize user-centric designs, embedding intuitive safe- guards that empower operators without hindering CPS effi- ciency [122]. A. Rapid Changing Deepfake Detection Landscape The arms race between deepfake generation and detection within MCP contexts demands continuous innovation in proto- col safeguards, as generative models rapidly evolve to bypass existing forensic checks in CPS data streams [25]. Adversarial techniques that fine-tune GANs or diffusion models to mimic legitimate tool outputs exacerbate MCP vulnerabilities, requir- ing dynamic update mechanisms for server-side detectors [2]. Multi-feature fusion approaches can counter this by analyzing inconsistencies across MCP message modalities, bolstering resistance to refined fakes [251]. Hybrid defenses to the escalating fidelity of synthetic content [301]. Addressing this issue further involves proactive disruption strategies embedded in MCP, such as watermarking payloads to hinder the propagation of deepfakes during agent ex- changes [313]. Transfer learning frameworks enable detectors to generalize from known generative artifacts, mitigating the lag behind novel creation methods [37]. Ensemble models that aggregate predictions from diverse architectures improve robustness, counteracting the adaptive nature of attackers [7]. These advancements ensure MCP remains a viable conduit for secure CPS operations amid ongoing generational leaps in deepfakes [253]. B. Challenges from Unseen Fakes Generalization of detectors to unseen fakes in MCP-driven CPS requires models that extrapolate beyond training distribu- tions, incorporating self-supervised learning to identify novel artifacts in protocol interactions [20]. Patch-wise analysis techniques dissect MCP payloads for localized inconsistencies, improving adaptability to emerging generative variants [186]. Cross-dataset training regimens bolster resilience, enabling detectors to handle diverse forgery types without retraining [267]. Multi-attention mechanisms focus on temporal and spa- tial cues, thereby enhancing performance on hybrid deepfakes that blend multiple modalities [20]. Further enhancements involve intra-prediction frameworks that model expected MCP behaviors and flag deviations from unseen manipulations [25]. Disentangled feature representa- tions separate content from forgery traces, facilitating transfer to novel scenarios [103]. Gated attention architectures priori- tize relevant features in dynamic CPS environments, reducing overfitting to specific fake patterns [7]. These strategies col- lectively advance MCP detectors toward universal applicability against evolving unseen threats [272]. C. Intelligence at the Edge Lightweight real-time defenses for edge devices in MCP- integrated CPS prioritize efficient architectures, such as tiny CNNs, which optimize for low-power computation while maintaining detection efficacy [131]. MobileNet variants with quantization enable rapid inference on resource-limited hard- ware, embedding MCP security without excessive latency [146]. Multi-feature fusion in compact models captures es- sential deepfake cues while balancing accuracy and edge constraints [218]. Hybrid CNN-LSTM-Transformer designs enable the real-time analysis of streaming MCP data, making them suitable for CCTV-like applications [306]. Advancing these defenses includes adaptive optimization techniques that fine-tune parameters on-device, enhancing responsiveness to local CPS threats [306]. End-to-end frame- works minimize overhead by integrating detection directly into MCP protocols, supporting seamless edge deployment [107]. Eye movement analysis in hybrid approaches adds behavioral layers, improving real-time forgery spotting with minimal resources [368]. Overall, these innovations enable scalable MCP protection on edge platforms [83]. D. Privacy Requirements Balancing privacy with provenance/authentication in MCP for CPS involves selective disclosure protocols that em- bed verifiable credentials without revealing full context data [326]. Blockchain-based frameworks offer immutable prove- nance while utilizing zero-knowledge proofs to maintain user anonymity in agent exchanges [155], [365]. Hybrid water- marking integrates post-quantum cryptography, ensuring au- thenticity traces resist tampering without compromising sensi- tive information [74]. Multi-factor validation with human over- sight adds ethical layers, mitigating over-reliance on automated provenance that could infringe privacy [17]. Further strategies include policy governance that enforces data minimization in MCP messages, as well as aligning authentication with regulatory standards [42]. Forensic tech- niques prioritize privacy-preserving artifacts, enabling detec- tion without full data exposure [273]. Bibliometric analyses highlight trends in ethical AI, guiding balanced implemen- tations that integrate privacy into provenance designs [160]. These measures foster trust in MCP systems by harmonizing security with individual rights [313]. E. Model Development Formal models for AI-agent security in CPS with MCP emphasize the use of applied pi-calculus for verifying protocol integrity against adversarial manipulations [150]. Intrusion detection frameworks model AI behaviors formally, ensuring resilience in networked environments [45]. Multimodal fusion in formal representations captures CPS dynamics, enabling rig- orous analysis of agent interactions [360]. Blockchain-enabled models formalize trust in distributed CPS, incorporating AI optimization for secure operations [371]. Expanding these models includes adversarial learning strate- gies that formally defend against poisonous attacks, using GANs to simulate threats [313]. Anomaly detection in power systems leverages formal AI models for real-time identification of threats. Special issue compilations on CPS security outline formal approaches integrating AI for enhanced protection JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202636 [146]. These formalizations provide foundational assurances for MCP in critical infrastructures. In summary, these open challenges underscore that securing AI agents in CPS is not primarily a modeling problem, but a systems problem governed by physical constraints, trust boundaries, and adversarial adaptability. The case study and SENTINEL-based analysis presented in this survey demon- strate that defenses ignoring timing, false-positive cost, and environmental grounding are unlikely to be deployable in safety-critical CPS, regardless of algorithmic sophistication. Progress will therefore require a shift away from detection- centric thinking toward architectures that embed provenance, physics-based anchors, and continuous validation as first- class design elements. Addressing these challenges demands coordinated advances across AI security, control and signal processing, systems engineering, and standards development, without which the next generation of autonomous, tool-using AI agents will remain fundamentally untrustworthy in the physical world. X. CONCLUSION In this survey, we have explored the profound vulnerabilities introduced by deepfakes at the intersection of AI agents and cyber-physical systems through the Model Context Protocol. Deepfakes, spanning visual, audio, textual, and behavioral modalities, exploit the AI-environment interface to deceive sensors, manipulate data streams, and undermine operational integrity. Within MCP, these threats manifest as context poi- soning, where synthetic content infiltrates tool invocations and inter-agent communications, resulting in cascading failures in critical infrastructures such as smart grids and autonomous vehicles. The unique challenge lies in the seamless blending of digital forgery with physical actions, where a spoofed video feed or cloned voice command can trigger unsafe actions or conceal anomalies, amplifying risks beyond traditional cyberattacks. By categorizing these threats and examining real- world incidents, the survey underscores how MCP’s interop- erability, while enabling efficiency, inadvertently expands the attack surface for adversarial AI-generated content. AI-agent interactions in CPS, facilitated by MCP, introduce additional layers of complexity, as agents rely on external tools and shared contexts that adversaries can hijack through deepfake-driven exploits. User-agent prompts, agent-agent trust exploitation, and agent-environment manipulations create a fertile ground for social engineering, identity spoofing, and indirect injections, where malicious actors leverage generative models to emulate legitimate behaviors or forge instructions. In MCP ecosystems, these interactions heighten the potential for supply-chain compromises and over-privileged access, as unvetted servers process deepfake-tainted data, potentially leading to unauthorized actions or data leaks. The survey highlights how such dynamics erode trustworthiness, with deepfakes not only deceiving machines but also misleading human operators, thereby challenging the core principles of safety and reliability in CPS. Addressing these threats requires recognizing their hybrid nature, combining cyber deception with physical consequences in protocol-mediated environ- ments. To counter these evolving dangers, an interdisciplinary collaboration is essential, drawing expertise from AI for advanced generative and detection models, cybersecurity for robust protocol designs, signal processing for forensic analysis of environmental fingerprints, and human factors for intuitive training and policy development. This convergence can foster innovative solutions, such as AI-enhanced provenance track- ing informed by signal anomalies and user-centric interfaces that incorporate behavioral insights. By uniting these fields, researchers can develop holistic frameworks that anticipate deepfake adaptations within MCP, ensuring defenses evolve in tandem with threats. Collaborative efforts through joint initiatives and shared datasets will accelerate progress, bridg- ing gaps between theoretical advancements and practical CPS deployments. Ultimately, this integration promises to create resilient systems that safeguard against the multifaceted risks posed by AI-driven manipulations. Defense-in-depth emerges as a guiding principle, advocating layered protections that integrate provenance authentication, multi-factor validation, robust training, proactive disruptions, and human-policy measures within MCP architectures. This approach mitigates single points of failure by combining tech- nical safeguards like digital watermarking and blockchain with procedural elements such as regulations and standards, creat- ing a comprehensive barrier against deepfake incursions. In CPS, where failures can have physical repercussions, defense- in-depth ensures redundancy, with each layer addressing spe- cific vulnerabilities in agent interactions and environmental interfaces. By embedding these principles into MCP speci- fications, systems can achieve adaptive security, dynamically responding to threats while maintaining operational continuity. This multi-tiered strategy not only deters attacks but also minimizes their impact, promoting sustained functionality in adversarial settings. Resilience stands as the cornerstone for future MCP-enabled CPS, emphasizing systems that withstand, recover from, and adapt to deepfake threats through ongoing monitoring, formal modeling, and emerging directions like physical fingerprints and environmental anchors. This principle prioritizes long- term viability, ensuring AI agents remain trustworthy amid the arms race of generation and detection, with detectors generalizing to unseen fakes and lightweight defenses suiting edge devices. Balancing privacy with authentication further re- inforces resilience, preventing over-exposure while upholding provenance. As CPS integrates more deeply with AI, embrac- ing resilience through interdisciplinary innovation will guide the development of secure, ethical ecosystems. In closing, this survey underscores the need for sustained efforts to implement these principles, thereby paving the way for dependable AI in an increasingly interconnected world. REFERENCES [1] F. Abbas and A. Taeihagh, “Unmasking deepfakes: A systematic review of deepfake detection and generation techniques using artificial intelligence,” Expert Systems with Applications, vol. 252, p. 124260, 2024. [2] M. Abbasi, P. V ́ az, J. Silva, and P. Martins, “Comprehensive evaluation of deepfake detection models: Accuracy, generalization, and resilience to adversarial attacks,” Applied Sciences, vol. 15, no. 3, p. 1225, 2025. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202637 [3] O. I. Abiodun, M. Alawida, A. E. Omolara, and A. Alabdulatif, “Data provenance for cloud forensic investigations, security, challenges, solutions and future perspectives: A survey,” Journal of King Saud University-Computer and Information Sciences, vol. 34, no. 10, p. 10 217–10 245, 2022. [4] S. T. R. Adapala and Y. R. Alugubelly, “The aegis protocol: A foundational security framework for autonomous ai agents,” arXiv preprint arXiv:2508.19267, 2025. [5] K. Addo, M. Kabeya, and E. E. Ojo, “Federated quantum machine learning for distributed cybersecurity in multi-agent energy systems,” Energies, vol. 18, no. 20, p. 5418, 2025. [6] N. U. R. Ahmed, A. Badshah, H. Adeel, A. Tajammul, A. Duad, and T. Alsahfi, “Visual deepfake detection: Review of techniques, tools, limitations, and future prospects,” IEEE Access, 2024. [7] Z. Akhtar, “Deepfakes generation and detection: A short survey,” Journal of Imaging, vol. 9, no. 1, p. 18, 2023. [8] Z. Akhtar, T. L. Pendyala, and V. S. Athmakuri, “Video and audio deepfake datasets and open issues in deepfake technology: being ahead of the curve,” Forensic Sciences, vol. 4, no. 3, p. 289–377, 2024. [9] H. Al-Tairi, A. Javed, T. Khan, and A. K. J. Saudagar, “Deeplasd countermeasure for logical access audio spoofing,” Scientific Reports, vol. 15, no. 1, p. 20839, 2025. [10] S. Alanazi and S. Asif, “Exploring deepfake technology: creation, consequences and countermeasures,” Human-Intelligent Systems Inte- gration, vol. 6, no. 1, p. 49–60, 2024. [11] A. A. Albustami and A. F. Taha, “Breaking the flow and the bank: Stealthy cyberattacks on water network hydraulics,” Water Research, p. 123719, 2025. [12] R. A. Alhamarneh and M. Mahinderjit Singh, “Strengthening internet of things security: Surveying physical unclonable functions for au- thentication, communication protocols, challenges, and applications,” Applied Sciences, vol. 14, no. 5, p. 1700, 2024. [13] H. A. Alhashimi, R. A. Khan, H. S. Alwageed, A. M. Algarni, S. Ayouni, and A. O. Almagrabi, “Exploring the role of generative ai in enhancing cybersecurity in software development life cycle,” Array, p. 100509, 2025. [14] H. Ali, S. Subramani, L. Bollinani, N. S. Adupa, S. El-Loh, and H. Malik, “Multilingual dataset integration strategies for robust au- dio deepfake detection: A safe challenge system,” arXiv preprint arXiv:2508.20983, 2025. [15] Z. Ali, T. Hussain, C.-L. Su, G. Parise, K. Sayler, M. Sadiq, and S. H. Rouhani, “A novel intelligent intrusion detection and prevention framework for shore-ship hybrid ac/dc microgrids under power quality disturbances,” in 2025 IEEE Industry Applications Society Annual Meeting (IAS). IEEE, 2025, p. 1–7. [16] Z. Almahmoud, P. D. Yoo, E. Damiani, K.-K. R. Choo, and C. Y. Yeun, “Forecasting cyber threats and pertinent mitigation technologies,” Technological Forecasting and Social Change, vol. 210, p. 123836, 2025. [17] S. AlMuhaideb, H. Alshaya, L. Almutairi, D. Alomran, and S. T. Alhamed, “Lightfakedetect: A lightweight model for deepfake detection in videos that focuses on facial regions,” Mathematics, vol. 13, no. 19, p. 3088, 2025. [18] Z. Almutairi and H. Elgibreen, “A review of modern audio deepfake de- tection methods: challenges and future directions,” Algorithms, vol. 15, no. 5, p. 155, 2022. [19] A. Alobaid, T. Bonny, and M. Alrahhal, “Disruptive attacks on artificial neural networks: A systematic review of attack techniques, detection methods, and protection strategies,” Intelligent Systems with Applica- tions, p. 200529, 2025. [20] M. Alrajeh and A. Al-Samawi, “Deepfake image classification using decision (binary) tree deep learning,” Journal of Sensor and Actuator Networks, vol. 14, no. 2, p. 40, 2025. [21] M. Alrashoud, “Deepfake video detection methods, approaches, and challenges,” Alexandria Engineering Journal, vol. 125, p. 265–277, 2025. [22] A. Alshehri, D. Almalki, E. Alharbi, and S. Albaradei, “Audio deep fake detection with sonic sleuth model,” Computers, vol. 13, no. 10, p. 256, 2024. [23] A. Amayuelas, X. Yang, A. Antoniades, W. Hua, L. Pan, and W. Wang, “Multiagent collaboration attack: Investigating adversarial attacks in large language model collaborations via debate,” arXiv preprint arXiv:2406.14711, 2024. [24] E. Amer, S. El-Sappagh, T. Abuhamad, B. A. S. Al-Rimy, and A. Mohasseb, “Graphshield: advanced dynamic graph-based malware detection using graph neural networks,” Expert Systems with Applica- tions, p. 129812, 2025. [25] I. Amerini, M. Barni, S. Battiato, P. Bestagini, G. Boato, V. Bruni, R. Caldelli, F. De Natale, R. De Nicola, L. Guarnera et al., “Deepfake media forensics: Status and future challenges,” Journal of Imaging, vol. 11, no. 3, p. 73, 2025. [26] B. An, S. Zhang, and M. Dredze, “Rag llms are not safer: A safety analysis of retrieval-augmented generation for large language models,” arXiv preprint arXiv:2504.18041, 2025. [27] R. Anggrainingsih, G. M. Hassan, and A. Datta, “Transformer-based models for combating rumours on microblogging platforms: a review,” Artificial Intelligence Review, vol. 57, no. 8, p. 212, 2024. [28] C. Anil, E. Durmus, N. Panickssery, M. Sharma, J. Benton, S. Kundu, J. Batson, M. Tong, J. Mu, D. Ford et al., “Many-shot jailbreaking,” Advances in Neural Information Processing Systems, vol. 37, p. 129 696–129 742, 2024. [29] T. Anusha and A. Srinagesh, “Deepfake video detection: A comprehen- sive survey of advanced machine learning and deep learning techniques to combat synthetic video manipulation,” in 2025 International Confer- ence on Multi-Agent Systems for Collaborative Intelligence (ICMSCI). IEEE, 2025, p. 1033–1041. [30] T. Arif, A. Javed, M. Alhameed, F. Jeribi, and A. Tahir, “Voice spoofing countermeasure for logical access attacks detection,” IEEE Access, vol. 9, p. 162 857–162 868, 2021. [31] M. A. Arshed, A. Alwadain, R. Faizan Ali, S. Mumtaz, M. Ibrahim, and A. Muneer, “Unmasking deception: empowering deepfake detection with vision transformer network,” Mathematics, vol. 11, no. 17, p. 3710, 2023. [32] M. Arya, U. Goyal, S. Chawla et al., “A study on deep fake face detection techniques,” in 2024 3rd International Conference on Applied Artificial Intelligence and Computing (ICAAIC). IEEE, 2024, p. 459– 466. [33] M. M. Aslam, A. Tufail, R. A. A. H. M. Apong, L. C. De Silva, and M. T. Raza, “Scrutinizing security in industrial control systems: An architectural vulnerabilities and communication network perspective,” IEEE Access, vol. 12, p. 67 537–67 573, 2024. [34] S. Atawneh and H. Aljehani, “Phishing email detection model using deep learning,” Electronics, vol. 12, no. 20, p. 4261, 2023. [35] A. Awadallah, K. Eledlebi, M. J. Zemerly, D. Puthal, E. Damiani, K. Taha, T.-Y. Kim, P. D. Yoo, K.-K. R. Choo, M.-S. Yim et al., “Artificial intelligence-based cybersecurity for the metaverse: Research challenges and opportunities,” IEEE Communications Surveys & Tuto- rials, vol. 27, no. 2, p. 1008–1052, 2024. [36] M. Azzam, L. Pasquale, G. Provan, and B. Nuseibeh, “Forensic readi- ness of industrial control systems under stealthy attacks,” Computers & Security, vol. 125, p. 103010, 2023. [37] R. Babaei, S. Cheng, R. Duan, and S. Zhao, “Generative artificial intelligence and the evolving challenge of deepfake detection: A systematic analysis,” Journal of Sensor and Actuator Networks, vol. 14, no. 1, p. 17, 2025. [38] H. Babbar, S. Rani, and M. Shabaz, “Federated learning with enhanced cryptographic security for vehicular cyber-physical systems,” Scientific Reports, vol. 15, no. 1, p. 28593, 2025. [39] A. Badhan, P. Sharma, M. Dewangan, and H. Singh, “Enhancing deepfake detection through facial pattern recognition and transfer learning,” in 2025 7th International Conference on Inventive Material Science and Applications (ICIMA). IEEE, 2025, p. 1286–1291. [40] E. Bagdasaryan, T.-Y. Hsieh, B. Nassi, and V. Shmatikov, “Abusing images and sounds for indirect instruction injection in multi-modal llms,” in arXiv preprint arXiv:2307.10490, 2023. [41] A. Bakirov and I. Suleimenov, “Theoretical bases of methods of counteraction to modern forms of information warfare,” Computers, vol. 14, no. 10, p. 410, 2025. [42] I. Balafrej and M. Dahmane, “Enhancing practicality and efficiency of deepfake detection,” Scientific Reports, vol. 14, no. 1, p. 31084, 2024. [43] K. Balan, R. Learney, and T. Wood, “A framework for cryptographic verifiability of end-to-end ai pipelines,” in Proceedings of the 2025 ACM International Workshop on Security and Privacy Analytics, 2025, p. 49–59. [44] E. Bandara, S. Shetty, R. Mukkamala, R. Gore, P. Foytik, S. H. Bouk, A. Rahman, X. Liang, N. W. Keong, K. De Zoysa et al., “Model context contracts-mcp-enabled framework to integrate llms with blockchain smart contracts,” arXiv preprint arXiv:2510.19856, 2025. [45] N. Bansal, T. Aljrees, D. P. Yadav, K. U. Singh, A. Kumar, G. K. Verma, and T. Singh, “Real-time advanced computational intelligence for deep fake video detection,” Applied Sciences, vol. 13, no. 5, p. 3095, 2023. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202638 [46] Y. Baseri, A. S. Hafid, D. Makrakis, and H. Fereidouni, “Privacy- preserving federated learning framework for risk-based adaptive au- thentication,” arXiv preprint arXiv:2508.18453, 2025. [47] J. B. Bernabe, J. L. Canovas, J. L. Hernandez-Ramos, R. T. Moreno, and A. Skarmeta, “Privacy-preserving solutions for blockchain: Review and challenges,” Ieee Access, vol. 7, p. 164 908–164 940, 2019. [48] L. Beurer-Kellner, B. Buesser, A.-M. Cretu, E. Debenedetti, D. Dobos, D. Fabian, M. Fischer, D. Froelicher, K. Grosse, D. Naeff et al., “Design patterns for securing llm agents against prompt injections,” URL https://arxiv.org/abs/2506.08837, 2025. [49] M. Bezzi, “Large language models and security,” IEEE Security & Privacy, vol. 22, no. 2, p. 60–68, 2024. [50] S. BK and F. Azam, “Ensuring security and privacy in vanet: A com- prehensive survey of authentication approaches,” Journal of Computer Networks and Communications, vol. 2024, no. 1, p. 1818079, 2024. [51] R. Bohara and A. K. Bairwa, “Detecting deepfake audio using spectrogram-based machine learning approaches,” IEEE Access, 2025. [52] A. Boko, “Disinformation detection: Developing a categorical frame- work through thematic analysis,” Journalism and Media, vol. 5, no. 4, p. 1914–1924, 2024. [53] P. Borhani-Darian, H. Li, P. Wu, and P. Closas, “Detecting gnss spoofing using deep learning,” EURASIP Journal on advances in signal processing, vol. 2024, no. 1, p. 14, 2024. [54] A. Brissett and J. Wall, “Machine learning and watermarking for accurate detection of ai generated phishing emails.” Electronics, vol. 14, no. 13, p. 1–21, 2025. [55] T. Brooks, G. Princess, J. Heatley, J. Jeremy, K. Scott et al., “Increasing threats of deepfake identities,” US Department of Homeland Security [online], 2019. [56] C. B ̈ uhler, M. Biagiola, L. Di Grazia, and G. Salvaneschi, “Securing ai agent execution,” arXiv preprint arXiv:2510.21236, 2025. [57] E. Bureac ̆ a and I. Aciob ̆ anit , ei, “A blockchain blockchain-based frame- work for content provenance and authenticity,” in 2024 16th Interna- tional Conference on Electronics, Computers and Artificial Intelligence (ECAI). IEEE, 2024, p. 1–5. [58] D. Calder ́ on-Gonz ́ alez, N. ́ Abalos, B. Bayo, P. C ́ anovas, D. Griol, C. Mu ̃ noz-Romero, C. P ́ erez, P. Vila, and Z. Callejas, “Deep speech synthesis and its implications for news verification: Lessons learned in the rtve-ugr chair,” Applied Sciences, vol. 14, no. 21, p. 9916, 2024. [59] L. C ̧ alli and B. Alma C ̧ alli, “Recoding reality: A case study of youtube reactions to generative ai videos,” Systems, vol. 13, no. 10, p. 925, 2025. [60] N. Capuano, G. Fenza, V. Loia, and C. Stanzione, “Explainable artificial intelligence in cybersecurity: A survey,” Ieee Access, vol. 10, p. 93 575–93 600, 2022. [61] F. Casino, “Unveiling the multifaceted concept of cognitive security: Trends, perspectives, and future challenges,” Technology in Society, p. 102956, 2025. [62] A. Chan, K. Wei, S. Huang, N. Rajkumar, E. Perrier, S. Lazar, G. K. Hadfield, and M. Anderljung, “Infrastructure for ai agents,” arXiv preprint arXiv:2501.10114, 2025. [63] M. Charfeddine, H. M. Kammoun, B. Hamdaoui, and M. Guizani, “Chatgpt’s security risks and benefits: offensive and defensive use- cases, mitigation measures, and future implications,” IEEE Access, vol. 12, p. 30 263–30 310, 2024. [64] B. Chen, G. Li, X. Lin, Z. Wang, and J. Li, “Blockagents: Towards byzantine-robust llm-based multi-agent coordination via blockchain,” in Proceedings of the ACM Turing Award Celebration Conference-China 2024, 2024, p. 187–192. [65] G. Chen, C. Du, Y. Yu, H. Hu, H. Duan, and H. Zhu, “A deepfake image detection method based on a multi-graph attention network.” Electronics, vol. 14, no. 3, 2025. [66] J. Chen and S. L. Cong, “Agentguard: Repurposing agentic orchestrator for safety evaluation of tool orchestration,” arXiv preprint, 2025. [67] S. Chen, S. Liu, L. Zhou, Y. Liu, X. Tan, J. Li, S. Zhao, Y. Qian, and F. Wei, “Vall-e 2: Neural codec language models are human parity zero- shot text to speech synthesizers,” arXiv preprint arXiv:2406.05370, 2024. [68] S. Chen, C. Wang, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li et al., “Neural codec language models are zero-shot text to speech synthesizers,” IEEE Transactions on Audio, Speech and Language Processing, 2025. [69] T. Chen, J. Lou, and W. Wang, “Safeguarding multimodal knowl- edge copyright in the rag-as-a-service environment,” arXiv preprint arXiv:2506.10030, 2025. [70] Z.-C. Chen, L.-H. Tsao, C.-L. Fu, S.-F. Chen, and Y.-C. F. Wang, “Learning facial liveness representation for domain generalized face anti-spoofing,” in 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 2022, p. 1–6. [71] G. Chhetri, S. Somvanshi, M. M. Islam, S. Brotee, M. S. Mimi, D. Koirala, B. Pandey, and S. Das, “Model context protocols in adaptive transport systems: A survey,” arXiv preprint arXiv:2508.19239, 2025. [72] M. Christ, S. Gunn, and O. Zamir, “Undetectable watermarks for lan- guage models,” in The Thirty Seventh Annual Conference on Learning Theory. PMLR, 2024, p. 1125–1139. [73] H. Chung, T. Roughgarden, and E. Shi, “Collusion-resilience in transaction fee mechanism design,” in Proceedings of the 25th ACM Conference on Economics and Computation, 2024, p. 1045–1073. [74] J. E. Clapten and v. Balaji, “A gated temporal attention based intra prediction framework for robust deepfake video detection,” Scientific Reports, vol. 15, no. 1, p. 38540, 2025. [75] Coalition for Content Provenance and Authenticity (C2PA), “C2pa: Verifying media content sources,” https://c2pa.org/, 2024. [76] F.-A. Croitoru, A.-I. Hiji, V. Hondru, N. C. Ristea, P. Irofti, M. Popescu, C. Rusu, R. T. Ionescu, F. S. Khan, and M. Shah, “Deepfake media generation and detection in the generative ai era: a survey and outlook,” arXiv preprint arXiv:2411.19537, 2024. [77] E. N. Crothers, N. Japkowicz, and H. L. Viktor, “Machine-generated text: A comprehensive survey of threat models and detection methods,” IEEE Access, vol. 11, p. 70 977–71 002, 2023. [78] S. Dathathri, A. See, S. Ghaisas, P.-S. Huang, R. McAdam, J. Welbl, V. Bachani, A. Kaskasoli, R. Stanforth, T. Matejovicova et al., “Scal- able watermarking for identifying large language model outputs,” Nature, vol. 634, no. 8035, p. 818–823, 2024. [79] S. Datta, S. K. Nahin, A. Chhabra, and P. Mohapatra, “Agentic ai security: Threats, defenses, evaluation, and open challenges,” arXiv preprint arXiv:2510.23883, 2025. [80] Daxa.ai, “MCP security: Securing agentic AI with the model context protocol,” Daxa.ai Blog, July 2025. [81] L. P. de Melo, D. Macedo Amaral, R. de Oliveira Albuquerque, R. T. de Sousa J ́ unior, A. L. Sandoval Orozco, and L. J. Garc ́ ıa Villalba, “A secure approach out-of-band for e-bank with visual two-factor authorization protocol,” Cryptography, vol. 8, no. 4, p. 51, 2024. [82] Z. Deng, Y. Guo, C. Han, W. Ma, J. Xiong, S. Wen, and Y. Xiang, “Ai agents under threat: A survey of key security challenges and future pathways,” ACM Computing Surveys, vol. 57, no. 7, p. 1–36, 2025. [83] D. W. Deressa, H. Mareen, P. Lambert, S. Atnafu, Z. Akhtar, and G. Van Wallendael, “Genconvit: Deepfake video detection using gen- erative convolutional vision transformer,” Applied Sciences, vol. 15, no. 12, p. 6622, 2025. [84] E. Derner, K. Batisti ˇ c, J. Zah ́ alka, and R. Babu ˇ ska, “A security risk taxonomy for prompt-based interaction with large language models,” IEEE Access, 2024. [85] S. M. Dibaji, A. Hussain, and H. Ishii, “A tutorial on security and privacy challenges in cps,” Security and Resilience of Control Systems: Theory and Applications, p. 121–146, 2022. [86] A. Diel, T. Lalgi, I. C. Schr ̈ oter, K. F. MacDorman, M. Teufel, and A. B ̈ auerle, “Human performance in detecting deepfakes: A systematic review and meta-analysis of 56 papers,” Computers in Human Behavior Reports, vol. 16, p. 100538, 2024. [87] Docker,“Mcphorrorstories:Thesecurityissuesthreat- eningaiinfrastructure,”https://w.docker.com/blog/ mcp-security-issues-threatening-ai-infrastructure/, 2025. [88] A. Domenteanu, G.-C. T ̆ ataru, L. Cr ̆ aciun, A.-G. Mol ̆ anescu, L.-A. Cotfas, and C. Delcea, “Living in the age of deepfakes: a bibliometric exploration of trends, challenges, and detection approaches,” Informa- tion, vol. 15, no. 9, p. 525, 2024. [89] C. Doss, J. Mondschein, D. Shu, T. Wolfson, D. Kopecky, V. A. Fitton- Kane, L. Bush, and C. Tucker, “Deepfakes and scientific knowledge dissemination,” Scientific reports, vol. 13, no. 1, p. 13429, 2023. [90] W. W. Dou, I. Goldstein, and Y. Ji, “Ai-powered trading, algorithmic collusion, and price efficiency,” Jacobs Levy Equity Management Center for Quantitative Financial Research Paper, The Wharton School Research Paper, 2025. [91] D. S. Dsouza, A. E. Hajjar, and H. Jahankhani, “Deepfakes in social engineering attacks,” in Space Law Principles and Sustainable Mea- sures. Springer, 2024, p. 153–183. [92] C. Easttom, “Malicious use of artificial intelligence,” in 2025 IEEE 15th Annual Computing and Communication Workshop and Conference (CCWC). IEEE, 2025, p. 00 499–00 507. [93] H. Errico, J. Ngiam, and S. Sojan, “Securing the model context protocol (mcp): Risks, controls, and governance,” arXiv preprint arXiv:2511.20920, 2025. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202639 [94] I. Evtimov, A. Zharmagambetov, A. Grattafiori, C. Guo, and K. Chaud- huri, “Wasp: Benchmarking web agent security against prompt injec- tion attacks,” arXiv preprint arXiv:2504.18575, 2025. [95] M. Farhan, U. Butt, R. B. Sulaiman, and M. Alraja, “Self-sovereign identities and content provenance: Veritrust—a blockchain-based framework for fake news detection,” Future Internet, vol. 17, no. 10, p. 448, 2025. [96] S. Fatima and M. J. Arshad, “A comprehensive review of blockchain and machine learning integration for peer-to-peer energy trading in smart grids,” IEEE Access, 2025. [97] Y. Feng, Y. Guo, Y. Hou, Y. Wu, M. Lao, T. Yu, and G. Liu, “A survey of security threats in federated learning,” Complex & Intelligent Systems, vol. 11, no. 2, p. 165, 2025. [98] X. Fu, Z. Wang, S. Li, R. K. Gupta, N. Mireshghallah, T. Berg- Kirkpatrick, and E. Fernandes, “Misusing tools in large lan- guage models with visual adversarial examples,” in arXiv preprint arXiv:2310.03185, 2023. [99] S. Gaba, I. Budhiraja, V. Kumar, S. Garg, and M. M. Hassan, “An innovative multi-agent approach for robust cyber–physical systems using vertical federated learning,” Ad Hoc Networks, vol. 163, p. 103578, 2024. [100] S. Gaire, S. Gyawali, S. Mishra, S. Niroula, D. Thakur, and U. Yadav, “Systematization of knowledge: Security and safety in the model context protocol ecosystem,” arXiv preprint arXiv:2512.08290, 2025. [101] J. Gao, M. Micheletto, G. Orr ` u, S. Concas, X. Feng, G. L. Mar- cialis, and F. Roli, “Texture and artifact decomposition for improving generalization in deep-learning-based deepfake detection,” Engineering Applications of Artificial Intelligence, vol. 133, p. 108450, 2024. [102] Y. Gao, X. Wang, Y. Zhang, P. Zeng, and Y. Ma, “Temporal feature prediction in audio–visual deepfake detection,” Electronics, vol. 13, no. 17, p. 3433, 2024. [103] D. Garg and R. Gill, “Deepfake generation and detection-an exploratory study,” in 2023 10th IEEE Uttar Pradesh Section International Confer- ence on Electrical, Electronics and Computer Engineering (UPCON), vol. 10. IEEE, 2023, p. 888–893. [104] W. Ge, X. Wang, and J. Yamagishi, “Proactive detection of speaker identity manipulation with neural watermarking,” in ICLR 2025 Work- shop on Watermarking for Generative AI (WMARK@ICLR2025), 2025. [105] D. Ghiur ̆ au and D. E. Popescu, “Distinguishing reality from ai: approaches for detecting synthetic content,” Computers, vol. 14, no. 1, p. 1, 2024. [106] A. Giannaros, A. Karras, L. Theodorakopoulos, C. Karras, P. Kranias, N. Schizas, G. Kalogeratos, and D. Tsolis, “Autonomous vehicles: So- phisticated attacks, safety issues, challenges, open topics, blockchain, and future directions,” Journal of Cybersecurity and Privacy, vol. 3, no. 3, p. 493–543, 2023. [107] L. Y. Gong and X. J. Li, “A contemporary survey on deepfake detection: datasets, algorithms, and challenges,” Electronics, vol. 13, no. 3, p. 585, 2024. [108] Y. Gong, D. Ran, J. Liu, C. Wang, T. Cong, A. Wang, S. Duan, and X. Wang, “Figstep: Jailbreaking large vision-language models via typographic visual prompts,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 22, 2025, p. 23 951–23 959. [109] I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 2014. [110] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm- integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM workshop on artificial intelligence and security, 2023, p. 79–90. [111] K. Grimes, J. Lawler, R. C. Garrett, E. Mathew, M. Christiani, S. Kings- ley, Z. S. Wu, and N. VanHoudnos, “Sok: Bridging research and practice in LLM agent security,” Carnegie Mellon University Software Engineering Institute, Technical Report 6414, November 2025. [112] M. Groh, A. Sankaranarayanan, N. Singh, D. Y. Kim, A. Lippman, and R. Picard, “Human detection of political speech deepfakes across transcripts, audio, and video,” Nature communications, vol. 15, no. 1, p. 7629, 2024. [113] S. Grondin, A. Charpentier, and P. Ratz, “Beyond human intervention: Algorithmic collusion through multi-agent learning strategies,” arXiv preprint arXiv:2501.16935, 2025. [114] C.-P. S. P. W. Group, “Framework for cyber-physical systems: Volume 1, overview,” National Institute of Standards and Technology, NIST Special Publication 1500-201, June 2017. [Online]. Available: https: //nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1500-201.pdf [115] J. Guo and H. Cai, “System prompt poisoning: Persistent attacks on large language models beyond user injection,” arXiv preprint arXiv:2505.06493, 2025. [116] W. Guo, Y. Potter, T. Shi, Z. Wang, A. Zhang, and D. Song, “Frontier ai’s impact on the cybersecurity landscape,” arXiv preprint arXiv:2504.05408, 2025. [117] Y. Guo, P. Liu, W. Ma, Z. Deng, X. Zhu, P. Di, X. Xiao, and S. Wen, “Systematic analysis of mcp security,” arXiv preprint arXiv:2508.12538, 2025. [118] G. Gupta, K. Raja, M. Gupta, T. Jan, S. T. Whiteside, and M. Prasad, “A comprehensive review of deepfake detection using advanced machine learning and fusion methods,” Electronics, vol. 13, no. 1, p. 95, 2023. [119] P. Gupta, B. Ding, C. Guan, and D. Ding, “Generative ai: A system- atic review using topic modelling techniques,” Data and Information Management, vol. 8, no. 2, p. 100066, 2024. [120] P. Gupta, H. A. Patil, and R. C. Guido, “Vulnerability issues in automatic speaker verification (asv) systems,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, p. 10, 2024. [121] J. Ha, A. El Azzaoui, and J. H. Park, “Fl-tenb4: A federated-learning- enhanced tiny efficientnetb4-lite approach for deepfake detection in cctv environments,” Sensors (Basel, Switzerland), vol. 25, no. 3, p. 788, 2025. [122] A. H ̊ akansson, A. Saad, A. Anand, V. Gjærum, H. Robinson, and K. Seel, “Robust reasoning for autonomous cyber-physical systems in dynamic environments,” Procedia Computer Science, vol. 192, p. 3966–3978, 2021. [123] P. Haley, “The impact of biometric surveillance on reducing violent crime: Strategies for apprehending criminals while protecting the innocent,” Sensors, vol. 25, no. 10, p. 3160, 2025. [124] L. Hallal, J. Rhinelander, and R. Venkat, “Recent trends of authentica- tion methods in extended reality: A survey,” Applied System Innovation, vol. 7, no. 3, p. 45, 2024. [125] J. T. Hancock and J. N. Bailenson, “The social impact of deepfakes,” p. 149–152, 2021. [126] S. Harris, H. J. Hadi, N. Ahmad, and M. A. Alshara, “Fake news detection revisited: An extensive review of theoretical frameworks, dataset assessments, model constraints, and forward-looking research agendas,” Technologies, vol. 12, no. 11, p. 222, 2024. [127] H. R. Hasan, K. Salah, R. Jayaraman, I. Yaqoob, and M. Omar, “Nfts for combating deepfakes and fake metaverse digital contents,” Internet of Things, vol. 25, p. 101133, 2024. [128] M. K. Hasan, R. A. Abdulkadir, S. Islam, T. R. Gadekallu, and N. Safie, “A review on machine learning techniques for secured cyber-physical systems in smart grid networks,” Energy Reports, vol. 11, p. 1268– 1290, 2024. [129] M. M. Hasan, H. Li, E. Fallahzadeh, G. K. Rajbahadur, B. Adams, and A. E. Hassan, “Model context protocol (mcp) at first glance: Studying the security and maintainability of mcp servers,” arXiv preprint arXiv:2506.13538, 2025. [130] R.Hat,“Modelcontextprotocol(mcp):Understanding securityrisksandcontrols,”https://w.redhat.com/en/blog/ model-context-protocol-mcp-understanding-security-risks-and-controls, July 2025. [131] M. Hatami, L. Dorje, X. Li, and Y. Chen, “Electric network frequency as environmental fingerprint for metaverse security: A comprehensive survey,” Computers, vol. 14, no. 8, p. 321, 2025. [132] M. Hatami, Q. Qu, Y. Chen, J. Mohammadi, E. Blasch, and E. Ardiles- Cruz, “Anchor-grid: Authenticating smart grid digital twins using real- world anchors,” Sensors, vol. 25, no. 10, p. 2969, 2025. [133] S. He, Y. Lei, Z. Zhang, Y. Sun, S. Li, C. Zhang, and J. Ye, “Identity deepfake threats to biometric authentication systems: Public and expert perspectives,” arXiv preprint arXiv:2506.06825, 2025. [134] S. He, Y. Diao, Y. Li, C. Sun, L. Wang, and Z. Guo, “Kad-net: Kolmogorov-arnold and differential-aware networks for robust and sensitive proactive deepfake forensics,” Knowledge-Based Systems, p. 114692, 2025. [135] X. He, D. Wu, Y. Zhai, and K. Sun, “Sentinelagent: Graph-based anomaly detection in llm-based multi-agent systems,” arXiv preprint arXiv:2505.24201, 2025. [Online]. Available: https://arxiv.org/abs/ 2505.24201 [136] X. Hou, Y. Zhao, S. Wang, and H. Wang, “Model context protocol (mcp): Landscape, security threats, and future research directions,” arXiv preprint arXiv:2503.23278, 2025. [137] F.-S. Hsieh, “Robustness analysis of cyber-physical systems based on discrete timed cyber-physical models,” in 2021 IEEE World AI IoT Congress (AIIoT). IEEE, 2021, p. 0250–0254. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202640 [138] H.-P. Hsu, Z.-R. Jiang, L.-Y. Li, T.-C. Tsai, C.-H. Hung, S.-C. Chang, S.-S. Wang, and S.-H. Fang, “Detection of audio tampering based on electric network frequency signal,” Sensors, vol. 23, no. 16, p. 7029, 2023. [139] J. Hu, X. Yang, and L.-X. Yang, “A framework for detecting false data injection attacks in large-scale wireless sensor networks,” Sensors, vol. 24, no. 5, p. 1643, 2024. [140] G. Hua, Q. Wang, D. Ye, H. Zhang, G. Wang, and S. Xia, “Factors affecting forensic electric network frequency matching–a comprehen- sive study,” Digital Communications and Networks, vol. 10, no. 4, p. 1121–1130, 2024. [141] K. Huang and I. Habler, “Threat modeling google’s A2A protocol with the MAESTRO framework,” Cloud Security Alliance (CSA) Blog, April 2025. [142] K. Huang and C. Hughes, “Agentic ai red teaming,” in Securing AI Agents: Foundations, Frameworks, and Real-World Deployment. Springer, 2025, p. 207–252. [143] M. Iavich, “Combating fake news with cryptography in quantum era with post-quantum verifiable image proofs,” Journal of Cybersecurity and Privacy, vol. 5, no. 2, p. 31, 2025. [144] InvariantLabs,“Mcpsecuritynotification:Tool poisoningattacks,”https://invariantlabs.ai/blog/ mcp-security-notification-tool-poisoning-attacks, 2025. [145] Ironscales, “Deepfakes:assessing organizational readiness in the face of this emerging cyber threat,” IRONSCALES, Research Report, 2024. [146] M. B. E. Islam, M. Haseeb, H. Batool, N. Ahtasham, and Z. Muham- mad, “Ai threats to politics, elections, and democracy: a blockchain- based deepfake authenticity verification framework,” Blockchains, vol. 2, no. 4, p. 458–481, 2024. [147] Ismail, R. Kurnia, Z. A. Brata, G. A. Nelistiani, S. Heo, H. Kim, and H. Kim, “Toward robust security orchestration and automated response in security operations centers with a hyper-automation approach using agentic artificial intelligence,” Information, vol. 16, no. 5, p. 365, 2025. [148] E. Iturbe, O. Llorente-Vazquez, A. Rego, E. Rios, and N. Toledo, “Un- leashing offensive artificial intelligence: Automated attack technique code generation,” Computers & Security, vol. 147, p. 104077, 2024. [149] M. Javed, Z. Zhang, F. H. Dahri, and T. Kumar, “Enhancing multimodal deepfake detection with local–global feature integration and diffusion models,” Signal, Image and Video Processing, vol. 19, no. 5, p. 1–9, 2025. [150] M. Javed, Z. Zhang, F. H. Dahri, and A. A. Laghari, “Real-time deepfake video detection using eye movement analysis with a hybrid deep learning approach,” Electronics, vol. 13, no. 15, p. 2947, 2024. [151] A. Jayanthiladevi, J. Natarajan, K. Arjun, L. G. Atlas, M. Arvindhan, and D. Arockiam, “Ai-based cybersecurity frameworks for 7g-enabled virtual therapy platforms,” Cyber Security and Applications, p. 100099, 2025. [152] N. Jeffrey, Q. Tan, and J. R. Villar, “A review of anomaly detection strategies to detect threats to cyber-physical systems,” Electronics, vol. 12, no. 15, p. 3283, 2023. [153] R. Jiao, S. Xie, J. Yue, T. Sato, L. Wang, Y. Wang, Q. A. Chen, and Q. Zhu, “Can we trust embodied agents? exploring backdoor attacks against embodied llm-based decision-making systems,” in OpenReview, 2024. [154] W. Jin, H. Du, B. Zhao, X. Tian, B. Shi, and G. Yang, “A comprehensive survey on multi-agent cooperative decision-making: Scenarios, approaches, challenges and perspectives,” arXiv preprint arXiv:2503.13415, 2025. [155] G. Jing and H. Qi, “Zero-knowledge audit for internet of agents: Privacy-preserving communication verification with model context protocol,” arXiv preprint arXiv:2512.14737, 2025. [156] S. Johnson, V. Pham, and T. Le, “The dangers of indirect prompt injection attacks on llm-based autonomous web navigation agents: A demonstration,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2025, p. 729–738. [157] T. Ju, Y. Wang, X. Ma, P. Cheng, H. Zhao, Y. Wang, L. Liu, J. Xie, Z. Zhang, and G. Liu, “Flooding spread of manipulated knowledge in llm-based multi-agent communities,” arXiv preprint arXiv:2407.07791, 2024. [158] V. Kandasamy and A. A. Roseline, “Harnessing advanced hybrid deep learning model for real-time detection and prevention of man-in-the- middle cyber attacks,” Scientific Reports, vol. 15, no. 1, p. 1697, 2025. [159] M. M. Karim, D. H. Van, S. Khan, Q. Qu, and Y. Kholodov, “Ai agents meet blockchain: A survey on secure and scalable collaboration for multi-agents,” Future Internet, vol. 17, no. 2, 2025. [160] S. Karim, X. Liu, A. A. Khan, A. A. Laghari, A. Qadir, and I. Bibi, “Mcgan—a cutting edge approach to real time investigate of mul- timedia deepfake multi collaboration of deep generative adversarial networks with transfer learning,” Scientific Reports, vol. 14, no. 1, p. 29330, 2024. [161] M. Kazimierczak, N. Habib, J. H. Chan, and T. Thanapattheerakul, “Impact of ai on the cyber kill chain: A systematic review,” Heliyon, vol. 10, no. 24, 2024. [162] T. Kehkashan, R. A. Riaz, A. S. Al-Shamayleh, A. Akhunzada, N. Ali, M. Hamza, and F. Akbar, “Ai-generated text detection: A comprehen- sive review of methods, datasets, and applications,” Computer Science Review, vol. 58, p. 100793, 2025. [163] A. A. Khan, A. A. Laghari, S. A. Inam, S. Ullah, M. Shahzad, and D. Syed, “A survey on multimedia-enabled deepfake detection: state- of-the-art tools and techniques, emerging trends, current challenges & limitations, and future directions,” Discover Computing, vol. 28, no. 1, p. 48, 2025. [164] A. Khan, K. M. Malik, J. Ryan, and M. Saravanan, “Battling voice spoofing: a review, comparative analysis, and generalizability evalu- ation of state-of-the-art voice spoofing counter measures,” Artificial Intelligence Review, vol. 56, no. Suppl 1, p. 513–566, 2023. [165] P. L. Kharvi, “Understanding the impact of ai-generated deepfakes on public opinion, political discourse, and personal security in social media,” IEEE Security & Privacy, vol. 22, no. 4, p. 115–122, 2024. [166] M. A. Khatun, S. F. Memon, C. Eising, and L. L. Dhirani, “Machine learning for healthcare-iot security: A review and risk mitigation,” IEEE Access, vol. 11, p. 145 869–145 896, 2023. [167] K. Khurshid, K. Khurshid, M. U. Hadi, M. Al Bataineh, and N. Saeed, “Securing aiot surveillance: Techniques, challenges, and solutions,” IEEE Open Journal of the Communications Society, 2025. [168] H. H. Kilinc and F. Kaledibi, “Audio deepfake detection by using machine and deep learning,” in 2023 10th International Conference on Wireless Networks and Mobile Communications (WINCOM).IEEE, 2023, p. 1–5. [169] M. Kinnas, J. Violos, I. Kompatsiaris, and S. Papadopoulos, “Reducing inference energy consumption using dual complementary cnns,” Future Generation Computer Systems, vol. 165, p. 107606, 2025. [170] B. Kira, “When non-consensual intimate deepfakes go viral: The insufficiency of the uk online safety act,” Computer Law & Security Review, vol. 54, p. 106024, 2024. [171] J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein, “A watermark for large language models,” in International Conference on Machine Learning. PMLR, 2023, p. 17 061–17 084. [172] E.Kontsevoy,“Mcp’sbiggestsecurityloophole isidentityfragmentation,”TechRadarPro,Septem- ber2025.[Online].Available:https://w.techradar.com/pro/ mcps-biggest-security-loophole-is-identity-fragmentation [173] S. Kopecky, “Challenges of deepfakes,” in Science and Information Conference. Springer, 2024, p. 158–166. [174] C. Korgialas, C. Kotropoulos, and K. N. Plataniotis, “Leveraging electric network frequency estimation for audio authentication,” IEEE Access, vol. 12, p. 9308–9320, 2024. [175] KPMG, Deepfake Threats to Companies: Navigating AI-Driven Fraud Risks, 2024. [176] A. Kulkarni, N. A. Hazari, and M. Y. Niamat, “A zero trust-based framework employing blockchain technology and ring oscillator phys- ical unclonable functions for security of field programmable gate array supply chain,” IEEE Access, vol. 12, p. 89 322–89 338, 2024. [177] A. Kumar, D. Singh, R. Jain, D. K. Jain, C. Gan, and X. Zhao, “Ad- vances in deepfake detection algorithms: Exploring fusion techniques in single and multi-modal approach,” Information Fusion, p. 102993, 2025. [178] N. Kumar and A. K. Singh, “Artificial intelligence content detection techniques using watermarking: A survey,” Image and Vision Comput- ing, p. 105728, 2025. [179] S. N. P. Kumar, “A secure accountability framework for multi-modal agent systems: Detecting, mitigating, and auditing data-poisoning at- tacks via model context protocol (mcp) servers,” Journal of Computer Science and Technology Studies, vol. 7, no. 12, p. 01–05, November 2025. [180] F. La Vigne, “Model context protocol: Discover the missing linkforaiintegration,”https://w.redhat.com/en/blog/ model-context-protocol-discover-missing-link-ai-integration,April 2025, red Hat Blog. [181] S. Lad, “Adversarial approaches to deepfake detection: A theoretical framework for robust defense,” Journal of Artificial Intelligence Gen- eral science (JAIGS), vol. 6, no. 1, p. 46–58, 2024. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202641 [182] Z. Lai, Z. Yao, G. Lai, C. Wang, and R. Feng, “A novel face swapping detection scheme using the pseudo zernike transform based robust watermarking,” Electronics, vol. 13, no. 24, p. 4955, 2024. [183] M. Landauer, F. Skopik, B. Stojanovi ́ c, A. Flatscher, and T. Ullrich, “A review of time-series analysis for cyber security analytics: from intrusion detection to attack prediction,” International Journal of In- formation Security, vol. 24, no. 1, p. 3, 2025. [184] E. A. Lee and S. A. Seshia, Introduction to embedded systems: A cyber- physical systems approach. MIT press, 2016. [185] J. H. Lee, A. Lauscher, and S. V. Albrecht, “Towards ethical multi- agent systems of large language models: A mechanistic interpretability perspective,” arXiv preprint arXiv:2512.04691, 2025. [186] S. Lei, J. Song, F. Feng, Z. Yan, and A. Wang, “Deepfake face detection and adversarial attack defense method based on multi-feature decision fusion,” Applied Sciences, vol. 15, no. 12, p. 6588, 2025. [187] J. K. Lewis, I. E. Toubal, H. Chen, V. Sandesera, M. Lomnitz, Z. Hampel-Arias, C. Prasad, and K. Palaniappan, “Deepfake video detection based on spatial, spectral, and temporal inconsistencies using multimodal deep learning,” in 2020 IEEE Applied Imagery Pattern Recognition Workshop (AIPR). IEEE, 2020, p. 1–9. [188] H. Li, S. Yang, R. Xia, L. Yuan, and X. Gao, “Big brother is watching: Proactive deepfake detection via learnable hidden face,” IEEE Signal Processing Letters, 2025. [189] J. Li, G. Sun, Q. Wu, S. Liang, P. Wang, and D. Niyato, “Two-way aerial secure communications via distributed collaborative beamform- ing under eavesdropper collusion,” in IEEE INFOCOM 2024-IEEE Conference on Computer Communications.IEEE, 2024, p. 331– 340. [190] J. Li, B. Li, X. Liu, J. Fang, F. Juefei-Xu, Q. Guo, and H. Yu, “Advgps: Adversarial gps for multi-agent perception attack,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2024, p. 18 421–18 427. [191] M. Li, Y. Ahmadiadli, and X.-P. Zhang, “A survey on speech deepfake detection,” ACM Computing Surveys, vol. 57, no. 7, p. 1–38, 2025. [192] Q. Li and Y. Xie, “From glue-code to protocols: A critical analysis of a2a and mcp integration for scalable agent systems,” arXiv preprint arXiv:2505.03864, 2025. [193] X. Li, S. Wang, S. Zeng, Y. Wu, and Y. Yang, “A survey on llm- based multi-agent systems: workflow, infrastructure, and challenges,” Vicinagearth, vol. 1, no. 1, p. 9, 2024. [194] Y. Li, T. Lu, S. Peng, C. He, K. Zhao, G. Yang, and Y. Chen, “Detection of electric network frequency in audio using multi-hcnet,” Sensors, vol. 25, no. 12, p. 3697, 2025. [195] Y. Liang, J. Xiao, W. Gan, and P. S. Yu, “Watermarking techniques for large language models: A survey,” arXiv preprint arXiv:2409.00089, 2024. [196] Z. Lin, H. Lin, L. Lin, S. Chen, and X. Liu, “Robust cross-image ad- versarial watermark with jpeg resistance for defending against deepfake models,” Computer Vision and Image Understanding, p. 104459, 2025. [197] Z. Lin, S. Zhang, G. Liao, D. Tao, and T. Wang, “Binding agent ID: Unleashing the power of AI agents with accountability and credibility,” arXiv preprint arXiv:2512.17538, 2025. [198] J. Liu, Z. Liu, Q. Li, W. Kong, and X. Li, “Multi-domain controversial text detection based on a machine learning and deep learning stacked ensemble,” Mathematics, vol. 13, no. 9, p. 1529, 2025. [199] Q. Liu, L. Wang, and M. Luo, “When seeing is not believing: self- efficacy and cynicism in the era of intelligent media,” Humanities and Social Sciences Communications, vol. 12, no. 1, p. 1–13, 2025. [200] T. Liu, W. Yang, C. Xu, J. Lv, H. Wang, Y. Zhang, S. Xu, and D. Man, “Act in collusion: A persistent distributed multi-target backdoor in federated learning,” arXiv preprint arXiv:2411.03926, 2024. [201] X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch et al., “Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, p. 2507–2522, 2023. [202] Y. Liu, L. Xu, S. Yang, D. Zhao, and X. Li, “Adversarial sample attacks and defenses based on lstm-ed in industrial control systems,” Computers & Security, vol. 140, p. 103750, 2024. [203] Y. Liu, X. Zhao, C. Kruegel, D. Song, and Y. Bu, “In-context water- marks for large language models,” arXiv preprint arXiv:2505.16934, 2025. [204] Y. Liu, G. Deng, Y. Li, K. Wang, Z. Wang, X. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng et al., “Prompt injection attack against llm-integrated applications,” arXiv preprint arXiv:2306.05499, 2023. [205] Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, K. Wang, and Y. Liu, “Jailbreaking chatgpt via prompt engineering: An empirical study,” arXiv preprint arXiv:2305.13860, 2023. [206] Y. Liu, R. Zhang, H. Luo, Y. Lin, G. Sun, D. Niyato, H. Du, Z. Xiong, Y. Wen, A. Jamalipour et al., “Secure multi-llm agentic ai and agentification for edge general intelligence by zero-trust: A survey,” arXiv preprint arXiv:2508.19870, 2025. [207] Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and benchmarking prompt injection attacks and defenses,” in 33rd USENIX Security Symposium (USENIX Security 24), 2024, p. 1831–1847. [208] A. Lopez Pellicer, P. Angelov, and N. Suri, “Securing (vision-based) autonomous systems: taxonomy, challenges, and defense mechanisms against adversarial threats,” Artificial Intelligence Review, vol. 58, no. 12, p. 1–59, 2025. [209] H. Lu, J. Liu, J. Peng, and J. Lu, “Adversarial attacks based on time- series features for traffic detection,” Computers & Security, vol. 148, p. 104175, 2025. [210] P. Lu, L. Zhang, M. Liu, K. Sridhar, O. Sokolsky, F. Kong, and I. Lee, “Recovery from adversarial attacks in cyber-physical systems: Shallow, deep, and exploratory works,” ACM Computing Surveys, vol. 56, no. 8, p. 1–31, 2024. [211] E. Lundberg and P. Mozelius, “The potential effects of deepfakes on news media and entertainment,” AI & SOCIETY, vol. 40, no. 4, p. 2159–2170, 2025. [212] H. Luo, L. Li, and J. Li, “Digital watermarking technology for ai- generated images: A survey,” Math, 2025. [213] M. Lupinacci, F. A. Pironti, F. Blefari, F. Romeo, L. Arena, and A. Furfaro, “The dark side of llms: Agent-based attacks for complete computer takeover,” arXiv preprint arXiv:2507.06850, 2025. [214] P. Lv, M. Sun, H. Wang, X. Wang, S. Zhang, Y. Chen, K. Chen, and L. Sun, “Rag-wm: An efficient black-box watermarking approach for retrieval-augmented generation of large language models,” in Pro- ceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, 2025, p. 1709–1723. [215] A. Mahboubi, K. Luong, G. Jarrad, S. Camtepe, M. Bewong, M. Bahutair, and G. Pogrebna, “Lurking in the shadows: Unsupervised decoding of beaconing communication for enhanced cyber threat hunting,” Journal of Network and Computer Applications, vol. 236, p. 104127, 2025. [216] M. Maranco, R. Nidhya, M. Sivakumar et al., “Intense triad defender for end-user security in cyber physical system,” in 2024 2nd Inter- national Conference on Networking and Communications (ICNWC). IEEE, 2024, p. 1–7. [217] C. Mazzocca, A. Acar, S. Uluagac, R. Montanari, P. Bellavista, and M. Conti, “A survey on decentralized identifiers and verifiable credentials,” IEEE Communications Surveys & Tutorials, 2025. [218] S. Meng, Q. Tan, Q. Zhou, and R. Wang, “Multi-branch network with multi-feature enhancement for improving the generalization of facial forgery detection,” Entropy, vol. 27, no. 5, p. 545, 2025. [219] A. Metcalfe-Pearce, “2026 cybersecurity predictions,” F5 Labs, De- cember 2025. [220] R. Miao, Y. Liu, Y. Wang, X. Shen, Y. Tan, Y. Dai, S. Pan, and X. Wang, “Blindguard: Safeguarding llm-based multi-agent systems under unknown attacks,” arXiv preprint arXiv:2508.08127, 2025. [221] L. Miculicich, M. Parmar, H. Palangi, K. D. Dvijotham, M. Montanari, T. Pfister, and L. T. Le, “Veriguard: Enhancing llm agent safety via verified code generation,” arXiv preprint arXiv:2510.05156, 2025. [222] D. Mikhaylenko and P. Zhang, “Stealthy targeted local covert attacks on cyber–physical systems,” Automatica, vol. 173, p. 112023, 2025. [223] Specification - Model Context Protocol, Model Context Protocol, June2025.[Online].Available:https://modelcontextprotocol.io/ specification/latest [224] M. S. Momin, A. Sufian, D. Barman, M. Leo, C. Distante, and N. Damer, “Explainable deepfake detection across different modalities: An overview of methods and challenges,” Image and Vision Computing, p. 105738, 2025. [225] C. Monteiro. (2025, May) Safety and security in the model context protocol (MCP). [226] Z. Moti, S. Hashemi, H. Karimipour, A. Dehghantanha, A. N. Jahromi, L. Abdi, and F. Alavi, “Generative adversarial network to detect unseen internet of things malware,” Ad Hoc Networks, vol. 122, p. 102591, 2021. [227] S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. Torr, L. Hammond, and C. Schroeder de Witt, “Secret collusion among ai agents: Multi-agent deception via steganography,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 37, JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202642 2024. [Online]. Available: https://w.proceedings.com/content/079/ 079017-2336open.pdf [228] R. Mubarak, T. Alsboui, O. Alshaikh, I. Inuwa-Dutse, S. Khan, and S. Parkinson, “A survey on the detection and impacts of deepfakes in visual, audio, and textual formats,” Ieee Access, vol. 11, p. 144 497– 144 529, 2023. [229] V. Mukesh, “A comprehensive review of advanced machine learn- ing techniques for enhancing cybersecurity in blockchain net- works,” INTERNATIONAL JOURNAL OF ARTIFICIAL INTELLI- GENCE (ISCSITR-IJAI), vol. 5, p. 1–6, 2025. [230] K. Mukherjee and M. Kantarcioglu, “Llm-driven provenance forensics for threat intelligence and detection,” 2025. [231] M. Mustak, J. Salminen, M. M ̈ antym ̈ aki, A. Rahman, and Y. K. Dwivedi, “Deepfakes: Deceptions, mitigations, and opportunities,” Journal of Business Research, vol. 154, p. 113368, 2023. [232] S. M. Nagarajan, G. G. Deverajan, A. K. Bashir, R. P. Mahapatra, and M. S. Al-Numay, “Iadf-cps: Intelligent anomaly detection framework towards cyber physical systems,” Computer Communications, vol. 188, p. 81–89, 2022. [233] D. Nagothu, Y. Chen, E. Blasch, A. Aved, and S. Zhu, “Detecting malicious false frame injection attacks on surveillance systems at the edge using electrical network frequency signals,” Sensors, vol. 19, no. 11, p. 2424, 2019. [234] V. S. Narajala and I. Habler, “Enterprise-grade security for the model context protocol (mcp): Frameworks and mitigation strategies,” arXiv preprint arXiv:2504.08623, 2025. [235] National Council on Teacher Retirement (NCTR), “Artificial intel- ligence (AI) and cyber security: An update,” July 2025, accessed: December 22, 2025. [236] S. Neupane, I. A. Fernandez, S. Mittal, and S. Rahimi, “Impacts and risk of generative ai technology on cyber defense,” arXiv preprint arXiv:2306.13033, 2023. [237] E. Ngharamike, L.-M. Ang, K. P. Seng, and M. Wang, “Enf based digital multimedia forensics: Survey, application, challenges and future work,” IEEE Access, vol. 11, p. 101 241–101 272, 2023. [238] L.-H. Nguyen, V.-L. Nguyen, R.-H. Hwang, J.-J. Kuo, Y.-W. Chen, C.- C. Huang, and P.-I. Pan, “Towards secured smart grid 2.0: exploring security threats, protection models, and challenges,” IEEE Communi- cations Surveys & Tutorials, 2024. [239] H.-H. Nguyen-Le, V.-T. Tran, T. Nguyen, and N.-A. Le-Khac, “A survey on proactive deepfake defense: Disruption and watermarking,” ACM Computing Surveys, vol. 58, no. 5, p. 1–37, 2025. [240] M. Nunes, P. Burnap, P. Reinecke, and K. Lloyd, “Bane or boon: Measuring the effect of evasive malware on system call classifiers,” Journal of Information Security and Applications, vol. 67, p. 103202, 2022. [241] Obsidian Security Team, “The 2025 AI agent security landscape: Players, trends, and risks,” Obsidian Security Blog, October 2025. [242] U. Odyurt, A. D. Pimentel, and I. G. Alonso, “Improving the robustness of industrial cyber–physical systems through machine learning-based performance anomaly identification,” Journal of Systems Architecture, vol. 131, p. 102716, 2022. [243] N. I. of Standards and T. (NIST), “Nist smart grid and cps newsletter - december 2017: The cps framework introduces the concept of trustwor- thiness,”https://w.nist.gov/ctl/smart-connected-systems-division/ smart-grid-group/nist-smart-grid-and-cps-newsletter-december, December 2017, published online March 16, 2018. [244] S. Onami, “Blockchain for cybersecurity: Enhancing data integrity and trust in digital transactions,” ResearchGate, September 2025. [245] C. Opara, P. Modesti, and L. Golightly, “Evaluating spam filters and stylometric detection of ai-generated phishing emails,” Expert Systems with Applications, vol. 276, p. 127044, 2025. [246] R. Owoputi and S. Ray, “Security of multi-agent cyber-physical sys- tems: A survey,” IEEE Access, vol. 10, p. 121 465–121 479, 2022. [247] M. Pantsar, “Developing artificial human-like arithmetical intelligence (and why),” Minds and Machines, vol. 33, no. 3, p. 379–396, 2023. [248] P. S. Park, S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks, “Ai deception: A survey of examples, risks, and potential solutions,” Patterns, vol. 5, no. 5, 2024. [249] S. H. Park, S.-H. Lee, M. Y. Lim, P. M. Hong, and Y. K. Lee, “A comprehensive risk analysis method for adversarial attacks on biometric authentication systems,” IEEE Access, 2024. [250] K. Parti and J. Szab ́ o, “The legal challenges of realistic and ai-driven child sexual abuse material: regulatory and enforcement perspectives in europe,” Laws, vol. 13, no. 6, p. 67, 2024. [251] Y. Patel, S. Tanwar, R. Gupta, P. Bhattacharya, I. E. Davidson, R. Nyameko, S. Aluvala, and V. Vimal, “Deepfake generation and detection: Case study and challenges,” IEEE Access, vol. 11, p. 143 296–143 323, 2023. [252] M. D. Patil and V. V. Lokhande, “Model context protocol (MCP): Enabling scalable AI data integration,” International Journal For Multidisciplinary Research (IJFMR), vol. 7, no. 2, April 2025. [253] S. Patil, A. Bhat, N. Jain, and V. Javalkar, “Navigating deepfakes with data science: A multi-modal analysis and blockchain-based detection framework,” in 2025 International Conference on Pervasive Computa- tional Technologies (ICPCT). IEEE, 2025, p. 772–777. [254] M. Pawelec, “Deepfakes and democracy (theory): How synthetic audio- visual media for disinformation and hate speech threaten core demo- cratic functions,” Digital society, vol. 1, no. 2, p. 19, 2022. [255] M. Pawlicki, A. Pawlicka, R. Kozik, and M. Chora ́ s, “A meta-survey of adversarial attacks against artificial intelligence algorithms, including diffusion models,” Neurocomputing, p. 131231, 2025. [256] K. T. Pedersen, L. Pepke, T. Stærmose, M. Papaioannou, G. Choudhary, and N. Dragoni, “Deepfake-driven social engineering: Threats, detec- tion techniques, and defensive strategies in corporate environments,” Journal of Cybersecurity and Privacy, vol. 5, no. 2, p. 18, 2025. [257] R. Pedro, D. Castro, P. Carreira, and N. Santos, “From prompt injections to sql injection attacks: How protected is your llm-integrated web application?” arXiv preprint arXiv:2308.01990, 2023. [258] G. Pei, J. Zhang, M. Hu, Z. Zhang, C. Wang, Y. Wu, G. Zhai, J. Yang, C. Shen, and D. Tao, “Deepfake generation and detection: A benchmark and survey,” arXiv preprint arXiv:2403.17881, 2024. [259] F. Piccialli, D. Chiaro, S. Sarwar, D. Cerciello, P. Qi, and V. Mele, “Agentai: A comprehensive survey on autonomous agents in distributed ai for industry 4.0,” Expert Systems with Applications, p. 128404, 2025. [260] N. Pimpason, P. Viboonsang, and S. Kosolsombat, “Phishing email detection model using deep learning,” in 2025 IEEE International Conference on Cybernetics and Innovations (ICCI). IEEE, 2025, p. 1–5. [261] V. S. A. Piratla, S. Saxena, S. Bhatia, and N. Kumar, “Safeguarding the artificial pancreas: A review of security and reliability gaps and ai driven resilience,” in 2025 5th Intelligent Cybersecurity Conference (ICSC). IEEE, 2025, p. 399–411. [262] L. P ̈ ohler, V. Schrader, A. Ladwein, and F. von Keller, “A tech- nological perspective on misuse of available ai,” arXiv preprint arXiv:2403.15325, 2024. [263] C. Posta. (2025, May) Deep dive mcp and a2a attack vectors for ai agents. [264] M. C. Protocol, “Official github organization,” https://github.com/ modelcontextprotocol, 2025. [265] J. Pu, Z. Sarwar, S. M. Abdullah, A. Rehman, Y. Kim, P. Bhattacharya, M. Javed, and B. Viswanath, “Deepfake text detection: Limitations and opportunities,” in 2023 IEEE symposium on security and privacy (SP). IEEE, 2023, p. 1613–1630. [266] T. Pulikottil, L. A. Estrada-Jimenez, H. Ur Rehman, F. Mo, S. Nikghadam-Hojjati, and J. Barata, “Agent-based manufactur- ing—review and expert evaluation,” The International Journal of Advanced Manufacturing Technology, vol. 127, no. 5, p. 2151–2180, 2023. [267] H. Qian, L. Xia, R. Ge, Y. Fan, Q. Wang, and Z. Jing, “From black boxes to glass boxes: Explainable ai for trustworthy deepfake forensics,” Cryptography, vol. 9, no. 4, p. 61, 2025. [268] S. Qiao, Q. Guo, M. Wang, H. Zhu, J. J. Rodrigues, and Z. Lyu, “Frw- trace: Forensic-ready watermarking framework for tamper-resistant biometric data and attack traceability in consumer electronics,” IEEE Transactions on Consumer Electronics, 2025. [269] M. K. Quan, P. N. Pathirana, M. Wijayasundara, S. Setunge, D. C. Nguyen, C. G. Brinton, D. J. Love, and H. V. Poor, “Federated learning for cyber physical systems: a comprehensive survey,” IEEE Communications Surveys & Tutorials, 2025. [270] A.Ramaswami,“Howc2pahelpscombatmislead- inginformation,”https://w.linuxfoundation.org/blog/ how-c2pa-helps-combat-misleading-information, June 2024, linux Foundation Blog. [271] M. S. Rana, M. N. Nobi, B. Murali, and A. H. Sung, “Deepfake detection: A systematic literature review,” IEEE access, vol. 10, p. 25 494–25 513, 2022. [272] M. S. Rana, M. Solaiman, C. Gudla, and M. F. Sohan, “Deepfakes– reality under threat?” in 2024 IEEE 14th Annual Computing and Communication Workshop and Conference (CCWC).IEEE, 2024, p. 0721–0727. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202643 [273] N. Ranasinghe, P. Liyanage, and L. Kruglova, “Privacy preserving distributed image processing using federated learning and cnns,” in 2025 IEEE 15th Symposium on Computer Applications & Industrial Electronics (ISCAIE). IEEE, 2025, p. 138–143. [274] S. Rashid, E. Bollis, L. Pellicer, D. Rabbani, R. Palacios, A. Gupta, and A. Gupta, “Evaluating prompt injection attacks with lstm-based generative adversarial networks: A lightweight alternative to large lan- guage models,” Machine Learning and Knowledge Extraction, vol. 7, no. 3, p. 77, 2025. [275] P. P. Ray, “A review on agent-to-agent protocol: Concept, state-of-the- art, challenges and future directions,” Authorea Preprints, 2025. [276] —, “A survey on model context protocol: Architecture, state-of-the- art, challenges and future directions,” Authorea Preprints, 2025. [277] S. Raza, R. Sapkota, M. Karkee, and C. Emmanouilidis, “Trism for agentic ai: A review of trust, risk, and security management in llm- based agentic multi-agent systems,” arXiv preprint arXiv:2506.04133, 2025. [278] J. Roh, V. Shejwalkar, and A. Houmansadr, “Multilingual and multi- accent jailbreaking of audio llms,” arXiv preprint arXiv:2504.01094, 2025. [279] N. Romandini, C. Mazzocca, K. Otsuki, and R. Montanari, “Sok: Security and privacy of ai agents for blockchain,” arXiv preprint arXiv:2509.07131, 2025. [280] F. Romero-Moreno, “Deepfake fraud detection: Safeguarding trust in generative ai,” Available at SSRN 5031627, 2024. [281] —, “Deepfake detection in generative ai: A legal framework proposal to protect human rights,” Computer Law & Security Review, vol. 58, p. 106162, 2025. [282] D. G. Rosado, A. Santos-Olmo, L. E. S ́ anchez, M. A. Serrano, C. Blanco, H. Mouratidis, and E. Fern ́ andez-Medina, “Managing cy- bersecurity risks of cyber-physical systems: The marisma-cps pattern,” Computers in Industry, vol. 142, p. 103715, 2022. [283] C.-M. Rosca, A. Stancu, and E. M. Iovanovici, “The new paradigm of deepfake detection at the text level,” Applied Sciences, vol. 15, no. 5, p. 2560, 2025. [284] E.Roth,“Anthropiclaunchestooltoconnectaisys- temsdirectlytodatasets,”TheVerge,November2024. [Online]. Available: https://w.theverge.com/2024/11/25/24305774/ anthropic-model-context-protocol-data-sources [285] M. Sadaf, Z. Iqbal, A. R. Javed, I. Saba, M. Krichen, S. Majeed, and A. Raza, “Connected and automated vehicles: Infrastructure, applica- tions, security, critical challenges, and future aspects,” Technologies, vol. 11, no. 5, p. 117, 2023. [286] S. Saeed, S. A. Suayyid, M. S. Al-Ghamdi, H. Al-Muhaisen, and A. M. Almuhaideb, “A systematic literature review on cyber threat intelli- gence for organizational cybersecurity resilience,” Sensors, vol. 23, no. 16, p. 7273, 2023. [287] M. Saideh, J.-P. Jamont, and L. Vercouter, “Opportunistic sensor-based authentication factors in and for the internet of things,” Sensors, vol. 24, no. 14, p. 4621, 2024. [288] F. H. Sakacı and T. Yıldırım, “Conducted emission signal-based identification and real-time hardware security with deep learning,” Engineering Applications of Artificial Intelligence, vol. 136, p. 109025, 2024. [289] B. A. Salau, A. Rawal, and D. B. Rawat, “Recent advances in artificial intelligence for wireless internet of things and cyber–physical systems: A comprehensive survey,” IEEE Internet of Things Journal, vol. 9, no. 15, p. 12 916–12 930, 2022. [290] M. Salih, J. Gharib, and Y. Gahi, “Addressing security gaps in MCP: Design of a resilient reference architecture,” in 2025 11th International Conference on Optimization and Applications (ICOA).IEEE, 2025, p. 1–7. [291] D. Salvi, H. Liu, S. Mandelli, P. Bestagini, W. Zhou, W. Zhang, and S. Tubaro, “A robust approach to multimodal deepfake detection,” Journal of Imaging, vol. 9, no. 6, p. 122, 2023. [292] L. E. S ́ anchez, A. Santos-Olmo, D. G. Rosado, C. Blanco, M. A. Ser- rano, H. Mouratidis, and E. Fern ́ andez-Medina, “Marisma: A modern and context-aware framework for assessing and managing information cybersecurity risks,” Computer Standards & Interfaces, vol. 92, p. 103935, 2025. [293] M.-P. Sandoval, M. de Almeida Vau, J. Solaas, and L. Rodrigues, “Threat of deepfakes to the criminal justice system: a systematic review,” Crime Science, vol. 13, no. 1, p. 41, 2024. [294] A. K. Sangaiah, X. Wang, M. S. Obaidat, P. C. Huang, and K. Govin- dan, “Guest editorial data-driven innovation and adversarial learning models for industry 5.0 toward consumer digital ecosystems,” IEEE Transactions on Consumer Electronics, vol. 70, no. 2, p. 4878–4881, 2024. [295] K. Sasikumar and S. Nagarajan, “Enhancing cloud security: A multi- factor authentication and adaptive cryptography approach using ma- chine learning techniques,” IEEE Open Journal of the Computer Society, 2025. [296] T. Say, M. Alkan, and A. Kocak, “Advancing gan deepfake detection: Mixed datasets and comprehensive artifact analysis,” Applied Sciences, vol. 15, no. 2, p. 923, 2025. [297] M. Schmitt and I. Flechais, “Digital deception: Generative artificial intelligence in social engineering and phishing,” Artificial Intelligence Review, vol. 57, no. 12, p. 324, 2024. [298] P. Selvaraj, S. Jagatheesaperumal, K. Marimuthu, O. Saravanan, B. Alkhamees, and M. Hassan, “Deepfake detection using adversarial neural network,” Computer Modeling in Engineering & Sciences, vol. 143, no. 2, p. 1575, 2025. [299] S. Selvarajan, H. Manoharan, M. Abdelhaq, A. O. Khadidos, A. O. Khadidos, R. Alsaqour, and M. Uddin, “Diagnostic behavior analysis of profuse data intrusions in cyber physical systems using adversarial learning techniques,” Scientific Reports, vol. 15, no. 1, p. 7287, 2025. [300] K. Selvi and G. Dilip, “Enhancing cyber-physical systems security: A review of deep learning and blockchain integration,” in 2024 5th International Conference on Image Processing and Capsule Networks (ICIPCN). IEEE, 2024, p. 725–734. [301] I. Shallal, L. Rzouga Haddada, and N. Essoukri Ben Amara, “Image forgery detection with focus on copy-move: An overview, real world challenges and future directions,” Applied Sciences, vol. 15, no. 21, p. 11774, 2025. [302] J. Shang, J. Zhou, and T. Chen, “Nonlinear stealthy attacks on remote state estimation,” Automatica, vol. 167, p. 111747, 2024. [303] U. Sharma and J. Singh, “A comprehensive overview of fake news detection on social networks,” Social Network Analysis and Mining, vol. 14, no. 1, p. 120, 2024. [304] V. K. Sharma, R. Garg, and Q. Caudron, “A systematic literature review on deepfake detection techniques,” Multimedia Tools and Applications, vol. 84, no. 20, p. 22 187–22 229, 2025. [305] E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,” arXiv preprint arXiv:2307.14539, 2023. [306] Y. Sheng, Z. Zou, Z. Yu, M. Pang, W. Ou, and W. Han, “Id-insensitive deepfake detection model based on multi-attention mechanism,” Scien- tific Reports, vol. 15, no. 1, p. 11168, 2025. [307] C. Shi, M. Qiao, Z. Li, Z. Akhtar, B. Wang, M. Han, and T. Qiao, “Deepfake video traceability and authentication via source attribution,” IET Biometrics, vol. 2025, p. 1–14, 2025. [308] M. R. Shoaib, Z. Wang, M. T. Ahvanooey, and J. Zhao, “Deepfakes, misinformation, and disinformation in the era of frontier ai, generative ai, and large ai models,” in 2023 international conference on computer and applications (ICCA). IEEE, 2023, p. 1–7. [309] T. Siameh, A. A. Addobea, and C.-H. Liu, “Context injection vulner- abilities and resource exploitation attacks in model context protocol,” Authorea Preprints, 2025. [310] L. H. Singh, P. Charanarur, and N. K. Chaudhary, “Advancements in detecting deepfakes: Ai algorithms and future prospects- a review,” Discover Internet of Things, vol. 5, no. 1, p. 53, 2025. [311] L. D. Singh and P. Meher, “Anomaly detection in cyber-physical elec- trical systems using ai-enhanced pufs,” in 2025 Fourth International Conference on Power, Control and Computing Technologies (ICPC2T). IEEE, 2025, p. 1–6. [312] R. Singh. (2025, November) Agentic ai control fabric: The next enterprise operating system for autonomous workflows. [313] S. Sohail, S. M. Sajjad, A. Zafar, Z. Iqbal, Z. Muhammad, and M. Kazim, “Deepfake image forensics for privacy protection and authenticity using deep learning,” Information, vol. 16, no. 4, p. 270, 2025. [314] M. Soltani, K. Khajavi, M. Jafari Siavoshani, and A. H. Jahangir, “A multi-agent adaptive deep learning framework for online intrusion detection,” Cybersecurity, vol. 7, no. 1, p. 9, 2024. [315] S. Son and W. Kim, “Advancing generalization in deepfake detection: Supervised contrastive representation learning with dual stream spatio- temporal features,” IEEE Access, 2025. [316] H. Song, Y. Shen, W. Luo, L. Guo, T. Chen, J. Wang, B. Li, X. Zhang, and J. Chen, “Beyond the protocol: Unveiling attack vectors in the model context protocol ecosystem,” arXiv preprint arXiv:2506.02040, 2025. [317] J. Sotiropoulos, R. F. Del Rosario, K. Huang et al., “Agentic ai - threats and mitigations,” OWASP Foundation, Tech. Rep., December 2025. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202644 [318] J. Sotiropoulos, K. Katz, and R. F. Del Rosario, “OWASP top 10 for agentic applications – the benchmark for agentic security in the age of autonomous AI,” OWASP GenAI Security Project Blog, December 2025. [319] A. H. Soudy, O. Sayed, H. Tag-Elser, R. Ragab, S. Mohsen, T. Mostafa, A. A. Abohany, and S. O. Slim, “Deepfake detection using convolu- tional vision transformers and convolutional neural networks,” Neural Computing and Applications, vol. 36, no. 31, p. 19 759–19 775, 2024. [320] T. South, S. Marro, T. Hardjono, R. Mahari, C. D. Whitney, D. Green- wood, A. Chan, and A. Pentland, “Authenticated delegation and autho- rized ai agents,” arXiv preprint arXiv:2501.09674, 2025. [321] S. Stockwell, “From deepfake scams to poisoned chatbots: AI and election security in 2025,” CETaS Expert Analysis, November 2025. [322] C. Sun, H. Sun, Z. Guo, Y. Diao, L. Wang, D. Ma, G. Yang, and K. Li, “Diffmark: Diffusion-based robust watermark against deepfakes,” arXiv preprint arXiv:2507.01428, 2025. [323] J. Sun and W. Hou, “A study of two-branch fusion network model for low-quality deepfake detection in face videos,” Signal, Image and Video Processing, vol. 19, no. 8, p. 631, 2025. [324] P. Sun, Y. Li, H. Qi, and S. Lyu, “Faketracer: Exposing deepfakes with training data contamination,” in 2022 IEEE International Conference on Image Processing (ICIP). IEEE, 2022, p. 1161–1165. [325] I. Syllaidopoulos, K. Ntalianis, and I. Salmon, “A comprehensive survey on ai in counter-terrorism and cybersecurity: Challenges and ethical dimensions,” IEEE Access, 2025. [326] G. Tahaoglu, “Robust deepfake audio detection via an improved next-tdnn with multi-fused self-supervised learning features,” Applied Sciences, vol. 15, no. 17, p. 9685, 2025. [327] D. Tan, Y. Yang, C. Niu, S. Li, D. Yang, and B. Tan, “A review of deep learning based multimodal forgery detection for video and audio,” Discover Applied Sciences, vol. 7, no. 9, p. 987, 2025. [328] L. Tao, X. Wang, Y. Liu, and J. Wu, “Cloud-based user behavior emulation approach for space-ground integrated networks,” Sensors, vol. 22, no. 1, p. 44, 2021. [329] J. Taralkar and S. Narlawar, “Nft video tokenization: A decentralized approach to verifying media authenticity,” in 2025 IEEE Cloud Summit. IEEE, 2025, p. 174–180. [330] I. M. Tas and S. Baktir, “Blockchain-based caller-id authentication (bbca): A novel solution to prevent spoofing attacks in voip/sip net- works,” IEEE access, vol. 12, p. 60 123–60 137, 2024. [331] E. Tchaptchet, E. F. Tagne, J. Acosta, R. Danda, and C. Kamhoua, “Deepfakes detection by iris analysis,” IEEE Access, 2025. [332] K. Thakur, M. L. Ali, M. A. Obaidat, and A. Kamruzzaman, “A systematic review on deep-learning-based phishing email detection,” Electronics, vol. 12, no. 21, p. 4545, 2023. [333] J. Tian, L. Guan, Y. Liu, L. Zhang, and Y. Chen, “Deepphysio: detect- ing deepfake with non-personalized feature of physiological signal,” Multimedia Systems, vol. 31, no. 2, p. 86, 2025. [334] S. Tipper, H. F. Atlam, and H. S. Lallie, “An investigation into the utilisation of cnn with lstm for video deepfake detection,” Applied Sciences, vol. 14, no. 21, p. 9754, 2024. [335] A. Triantafyllopoulos, A. A. Spiesberger, I. Tsangko, X. Jing, V. Dis- tler, F. Dietz, F. Alt, and B. W. Schuller, “Vishing: Detecting so- cial engineering in spoken communication—a first survey & urgent roadmap to address an emerging societal challenge,” Computer Speech & Language, vol. 94, p. 101802, 2025. [336] S. Uppal, V. Banga, S. Neeraj, and A. Singhal, “A comprehensive study on mitigating synthetic identity threats using deepfake detection mech- anisms,” in 2024 14th International Conference on Cloud Computing, Data Science & Engineering (Confluence). IEEE, 2024, p. 750–755. [337] Upwind,“Unpackingthesecurityrisksofmodelcontext protocol(mcp)servers,”https://w.upwind.io/feed/ unpacking-the-security-risks-of-model-context-protocol-mcp-servers, April 2025. [338] S. Usmani, S. Kumar, and D. Sadhya, “Spatio-temporal knowledge distilled video vision transformer (stkd-vvit) for multimodal deepfake detection,” Neurocomputing, vol. 620, p. 129256, 2025. [339] L. Vegh, “Cyber-physical systems security through multi-factor authen- tication and data analytics,” in 2018 IEEE international conference on industrial technology (ICIT). IEEE, 2018, p. 1369–1374. [340] K. Verma, D. Mittal, S. Samanta, K. Gulati, O. Kulkarni, M. A. Dar, and C. Biji, “Deepfake audio detection: A comparative study of advanced deep learning models,” IEEE Access, 2025. [341] K. A. Wahba, K. A. Ahmed, M. R. Kamel, M. Fathy, P. K. H. Abdelfatah, and S. Hatem, “Creating a digital human twin: Cloning voice, face, and attitude,” in 2024 International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC).IEEE, 2024, p. 199–205. [342] F. Wang, Y. Jiang, R. Zhang, A. Wei, J. Xie, and X. Pang, “A survey of deep anomaly detection in multivariate time series: taxonomy, applications, and directions,” Sensors (Basel, Switzerland), vol. 25, no. 1, p. 190, 2025. [343] H. Wang, D. J. Miller, and G. Kesidis, “Anomaly detection of adversar- ial examples using class-conditional generative adversarial networks,” Computers & Security, vol. 124, p. 102956, 2023. [344] H. Wang, C. M. Poskitt, J. Sun, and J. Wei, “Pro2guard: Proactive run- time enforcement of llm agent safety via probabilistic model checking,” arXiv preprint arXiv:2508.00500, 2025. [345] K. Wang, G. Zhang, Z. Zhou, J. Wu, M. Yu, S. Zhao, C. Yin, J. Fu, Y. Yan, H. Luo et al., “A comprehensive survey in llm (- agent) full stack safety: Data, training and deployment,” arXiv preprint arXiv:2504.15585, 2025. [346] L. Wang, J. Deng, H. Tan, Y. Xu, J. Zhu, Z. Zhang, Z. Li, R. Zhan, and Z. Gu, “Aarf: Autonomous attack response framework for honeypots to enhance interaction based on multi-agent dynamic game,” Mathematics, vol. 12, no. 10, p. 1508, 2024. [347] S. Wang, G. Zhang, M. Yu, G. Wan, F. Meng, C. Guo, K. Wang, and Y. Wang, “G-safeguard: A topology-guided security lens and treatment on llm-based multi-agent systems,” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7261–7276, 2025. [348] S. Wang, X. Gu, J. Chen, C. Chen, and X. Huang, “Robustness improvement strategy of cyber-physical systems with weak interdepen- dency,” Reliability Engineering & System Safety, vol. 229, p. 108837, 2023. [349] T. Wang, H. Cheng, M.-H. Liu, and M. Kankanhalli, “Fractalforensics: Proactive deepfake detection and localization via fractal watermarks,” in Proceedings of the 33rd ACM International Conference on Multi- media, 2025, p. 7210–7219. [350] X. Wang, H. Xie, S. Ji, L. Liu, and D. Huang, “Blockchain-based fake news traceability and verification mechanism,” Heliyon, vol. 9, no. 7, 2023. [351] Y. Wang, Y. Pan, S. Guo, and Z. Su, “Security of internet of agents: Attacks and countermeasures,” IEEE Open Journal of the Computer Society, 2025. [352] Z. Wang, W. Xie, B. Wang, J. Tao, and E. Wang, “A survey on recent advanced research of cps security,” Applied Sciences, vol. 11, no. 9, p. 3751, 2021. [353] Z. Wang, G. Xu, and M. Ren, “Can attention detect ai-generated text? a novel benford’s law-based approach,” Information Processing & Management, vol. 62, no. 4, p. 104139, 2025. [354] Z. Wang, Q. Chang, H. Patel, S. Biju, C.-E. Wu, Q. Liu, A. Ding, A. Rezazadeh, A. Shah, Y. Bao et al., “Mcp-bench: Benchmarking tool-using llm agents with complex real-world tasks via mcp servers,” arXiv preprint arXiv:2508.20453, 2025. [355] A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” in NeurIPS 2023, 2023. [356] S. M. Williamson and V. Prybutok, “The era of artificial intelligence deception: unraveling the complexities of false realities and emerging threats of misinformation,” Information, vol. 15, no. 6, p. 299, 2024. [357] S. Willison, “Camel offers a promising new direction for mitigat- ing prompt injection attacks,” https://simonwillison.net/2025/Apr/11/ camel/, 2025. [358] E. Woollacott, “A malicious mcp server is silently stealing user emails,” ITPro, September 2025. [Online]. Available: https://w.itpro.com/ security/a-malicious-mcp-server-is-silently-stealing-user-emails [359] C. H. Wu, J. Y. Koh, R. Salakhutdinov, D. Fried, and A. Raghunathan, “Adversarial attacks on multimodal agents,” in Proceedings of ACL 2024, 2024. [360] Y. Wu, H. Huang, Z. Li, and S. Zhang, “Cbam-resnet: A lightweight resnet network focusing on time domain features for end-to-end deep- fake speech detection,” Electronics, vol. 14, no. 12, p. 2456, 2025. [361] S. Xiao, W. Zhu, Y. Jiang, K. Wang, P. Wang, C. Yan, X. Ji, and W. Xu, “Sok: Understanding the fundamentals and implications of sensor out- of-band vulnerabilities,” arXiv preprint arXiv:2508.16133, 2025. [362] W. Xing, M. Li, M. Li, and M. Han, “Towards robust and secure embodied ai: A survey on vulnerabilities and attacks,” arXiv preprint arXiv:2502.13175, 2025. [363] D. Xiong, Z. Wen, C. Zhang, D. Ren, and W. Li, “Bmnet: Enhancing deepfake detection through bilstm and multi-head self-attention mech- anism,” IEEE Access, 2025. JOURNAL OF L A T E X CLASS FILES, VOL. X, NO. X, JANUARY 202645 [364] D. Xu, I. Gondal, X. Yi, T. Susnjak, P. Watters, and T. R. McIntosh, “The erosion of cybersecurity zero-trust principles through generative ai: A survey on the challenges and future directions,” Journal of Cybersecurity and Privacy, vol. 5, no. 4, p. 87, 2025. [365] J. Xu, X. Liu, W. Lin, W. Shang, and Y. Wang, “Localization and detection of deepfake videos based on self-blending method,” Scientific Reports, vol. 15, no. 1, p. 3927, 2025. [366] Q. Xu, S. Ali, and T. Yue, “Digital twin-based anomaly detection in cyber-physical systems,” in 2021 14th IEEE Conference on Software Testing, Verification and Validation (ICST). IEEE, 2021, p. 205–216. [367] W. Xu, C. Huang, S. Gao, and S. Shang, “Llm-based agents for tool learning: A survey,” Data Science and Engineering, p. 1–31, 2025. [368] B. Yan, P. Liu, Y. Yang, and Y. Guo, “Self-supervised feature disen- tanglement for deepfake detection,” Mathematics, vol. 13, no. 12, p. 2024, 2025. [369] Y. Yang, H. Chai, Y. Song, S. Qi, M. Wen, N. Li, J. Liao, H. Hu, J. Lin, G. Chang et al., “A survey of ai agent protocols,” arXiv preprint arXiv:2504.16736, 2025. [370] Z. Yang, G. Zhao, and H. Wu, “Watermarking for large language models: A survey,” Mathematics, vol. 13, no. 9, p. 1420, 2025. [371] S. M. Yasir and H. Kim, “Lightweight deepfake detection based on multi-feature fusion,” Applied Sciences, vol. 15, no. 4, p. 1954, 2025. [372] J. Yi, Y. Xie, B. Zhu, E. Kiciman, G. Sun, X. Xie, and F. Wu, “Benchmarking and defending against indirect prompt injection attacks on large language models,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1, 2025, p. 1809–1820. [373] J. Yin, M. Gao, K. Shu, Z. Zhao, Y. Huang, and J. Wang, “Emulating reader behaviors for fake news detection,” IEEE Transactions on Big Data, 2025. [374] W. Yu, K. Hu, T. Pang, C. Du, M. Lin, and M. Fredrikson, “Infecting llm agents via generalizable adversarial attack,” in Red Teaming GenAI: What Can We Learn from Adversaries?, 2024. [375] —, “Infecting llm-based multi-agents via self-propagating adver- sarial attacks,” in Thirty-eighth Conference on Neural Information Processing Systems (NeurIPS), 2024. [376] K. Yuan, Z. Dong, X. Li, Z. Liu, C. Jia, and S. Lv, “An efficient and collusion-resistant key parameters pre-distribution system for day- ahead electricity auctions,” IEEE Internet of Things Journal, 2025. [377] Q. Yuan, Q. Meng, J. Tao, G. Li, J. Fei, B. Lu, and Y. Wang, “Multi- agent for network security monitoring and warning: A generative ai solution,” IEEE Network, 2025. [378] N. Zeeshan, M. Bakyt, N. Moradpoor, and L. La Spada, “Continu- ous authentication in resource-constrained devices via biometric and environmental fusion,” Sensors, vol. 25, no. 18, p. 5711, 2025. [379] C. Zeng, K. Li, and Z. Wang, “Enfformer: Long-short term repre- sentation of electric network frequency for digital audio tampering detection,” Knowledge-Based Systems, vol. 297, p. 111938, 2024. [380] S. Zeng, W. Wang, F. Huang, and Y. Fang, “Loft: Latent space opti- mization and generator fine-tuning for defending against deepfakes,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, p. 4750–4754. [381] Q. Zhan, R. Fang, H. S. Panchal, and D. Kang, “Adaptive attacks break defenses against indirect prompt injection attacks on llm agents,” arXiv preprint arXiv:2503.00061, 2025. [382] B. Zhang, H. Cui, V. Nguyen, and M. Whitty, “Audio deepfake detection: What has been achieved and what lies ahead,” Sensors (Basel, Switzerland), vol. 25, no. 7, p. 1989, 2025. [383] H. Zhang, X. Wang, Y. Wang, M. Li, C. Zhu, Z. Zhou, L. Xue, S. Hu, and L. Y. Zhang, “The threats of embodied multimodal llms: Jailbreaking robotic manipulation in the physical world,” arXiv preprint arXiv:2407.20242, 2024. [384] H. Zhang, J. Huang, K. Mei, Y. Yao, Z. Wang, C. Zhan, H. Wang, and Y. Zhang, “Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents,” in International Conference on Learning Representations (ICLR), 2025. [385] M. Zhang, C. Sonnadara, S. Shah, and M. Wu, “Feasibility study of location authentication for iot data using power grid signatures,” IEEE Open Journal of Signal Processing, 2025. [386] W. Zhang, S. Cui, Q. Zhang, B. Chen, H. Zeng, and Q. Zhong, “Hierarchical feature fusion and enhanced attention mechanism for robust gan-generated image detection,” Mathematics, vol. 13, no. 9, p. 1372, 2025. [387] Y. Zhang, X. Jiang, H. Sun, Y. Zhang, and D. Tong, “Curvemark: Detecting ai-generated text via probabilistic curvature and dynamic semantic watermarking,” Entropy, vol. 27, no. 8, p. 784, 2025. [388] Z. Zhang, W. Hao, A. Sankoh, W. Lin, E. Mendiola-Ortiz, J. Yang, and C. Mao, “I can hear you: Selective robust training for deepfake audio detection,” arXiv preprint arXiv:2411.00121, 2024. [389] S. Zhao, Q. Hou, Z. Zhan, Y. Wang, Y. Xie, Y. Guo, L. Chen, S. Li, and Z. Xue, “Mind your server: A systematic study of parasitic toolchain attacks on the mcp ecosystem,” arXiv preprint arXiv:2509.06572, 2025. [390] Y. Zhao, B. Liu, M. Ding, B. Liu, T. Zhu, and X. Yu, “Proactive deepfake defence via identity watermarking,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 2023, p. 4602–4611. [391] Z. Zhao, J. Duan, X. Hu, K. Xu, C. Wang, R. Zhang, Z. Du, Q. Guo, and Y. Chen, “Unlearnable examples for diffusion models: Protect data from unauthorized exploitation,” arXiv preprint arXiv:2306.01902, 2023. [392] B. Zhong, S. Liu, M. Caccamo, and M. Zamani, “Towards trustworthy ai: Sandboxing ai-based unverified controllers for safe and secure cyber-physical systems,” in 2023 62nd IEEE Conference on Decision and Control (CDC). IEEE, 2023, p. 1833–1840. [393] W. Zhou, X. Zhu, Q.-L. Han, L. Li, X. Chen, S. Wen, and Y. Xiang, “The security of using large language models: A survey with emphasis on chatgpt,” IEEE/CAA Journal of Automatica Sinica, 2024. [394] X. Zhou, H. Han, S. Shan, and X. Chen, “Fine-grained open-set deep- fake detection via unsupervised domain adaptation,” IEEE Transactions on Information Forensics and Security, 2024. [395] Z. Zhuang, Y. Tomioka, J. Shin, and Y. Okuyama, “Pgd-trap: Proactive deepfake defense with sticky adversarial signals and iterative latent variable refinement,” Electronics, vol. 13, no. 17, p. 3353, 2024.