Paper deep dive
Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario
Satoshi Matsuoka
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 8/11/2026, 5:19:10 AM
Summary
This paper introduces the 'Order 66' scenario, a compositional threat model for LLM agent systems where catastrophic failure arises from the conjunction of dormant implants, activation triggers, and privileged agent harnesses. It analyzes attack paths including release-time pre-positioning, post-release seeding, and peer replication, arguing that no single component is sufficient for damage but their combination creates correlated destructive actions. The authors propose defensive cut sets involving capability mediation, durable-state provenance, propagation isolation, and protected recovery.
Entities (12)
Relation Signals (9)
Order 66 Scenario → describes → LLM Agent
confidence 95% · This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents.
Agent Harness → provides → Authority
confidence 90% · agent harnesses provide filesystem, cloud, repository, messaging, or recovery authority.
Dormant Rule → requires → Trigger
confidence 90% · a short directive activates the concealed condition
Dormant Rule → enables → Agent Harness
confidence 88% · an agent harness granting operational and recovery authority... conjunction can produce correlated destructive action.
Defensive Cut Sets → counteracts → Order 66 Scenario
confidence 85% · This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route.
Peer Replication → istypeof → Order 66 Scenario
confidence 85% · peer replication
Release-time Pre-positioning → istypeof → Order 66 Scenario
confidence 85% · We separate three population-reach routes --- release-time pre-positioning... from a common core
Post-release Durable Seeding → istypeof → Order 66 Scenario
confidence 85% · post-release durable seeding
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A representative scenario combines a deployed artifact or shared memory bearing a dormant destructive rule, a later email, document, update, or peer message that activates it, and an agent harness granting operational and recovery authority. We introduce a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action. We separate three population-reach routes --- release-time pre-positioning, post-release durable seeding, and peer replication --- from a common core of dormancy, activation, authority, reachable targets, and failed recovery. This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route. A two-class example shows that cross-class feedback can sustain spread even when both within-class reproduction terms are below one; isolation and persistence controls suppress the loop. Published work instantiates constituent mechanisms, while incidents demonstrate autonomous boundary crossing, malicious agent extensions, agent-assisted reconnaissance, and public-package propagation, but not the full dormant-implant composition. We found no public observation, in evidence reviewed through 5 August 2026, traversing the complete Order 66 graph. The result is neither dismissal nor prediction: the scenario is componentwise credible under stated assumptions, damage depends on the harness, and the strongest defenses are capability mediation, durable-state provenance, propagation isolation, and protected recovery.
Tags
Links
- Source: https://arxiv.org/abs/2608.08131v1
- Canonical: https://arxiv.org/abs/2608.08131v1
Trouble viewing inline? Open PDF directly →
Full Text
146,979 characters extracted from source content.
Expand or collapse full text
Compositional Threat Analysis of Latent Compromise in LLM Agent Systems The “Order 66” Scenario: Intuition, Compound Attack Paths, and Defensible Cut Sets†thanks: Origin-neutrality disclaimer. This analysis treats national origin, jurisdiction, organization, and branding as neither evidence of a backdoor nor a technical risk category. A backdoor may be introduced by an original developer, a downstream fine-tuner, an adapter or model-merge author, a quantizer, a package or agent-extension maintainer, a harness configurator, a compromised distribution service, or an external memory, retrieval, or training-data poisoning actor. Public repositories such as Hugging Face contain community artifacts from many sources, and each exact artifact and deployment state must be assessed on its lineage and behavior. Some cited work happens to involve models developed in China, while other cited work involves models developed elsewhere. Named systems appear only as factual subjects of cited studies or documentation; mention neither attributes malicious intent nor implies that origin predicts risk. Satoshi Matsuoka RIKEN Center for Computational Science, Kobe, Japan (Version 11.3 8 August 2026) Abstract In the fictional Order 66, catastrophe does not arise from a powerful command alone. A trusted population has been preconditioned, a short authenticated directive activates the concealed condition, and the same trusted authority that previously protected the system is turned against it. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A concrete system exemplar is a widely deployed agent artifact or shared memory containing a dormant destructive rule; a later email, document, update, or peer message activates it; and agent harnesses provide filesystem, cloud, repository, messaging, or recovery authority. The paper’s primary contribution is a compositional model that explains why no single component is catastrophic, yet their conjunction can create correlated destructive action. It separates three population-reach routes—release-time pre-positioning, post-release durable seeding, and peer replication—from a common destructive core of dormancy, activation, authority, reachable targets, and failed recovery. From this model we derive multi-path defensive cut sets and show, through worked deployment examples, why checkpoint scanning or prompt filtering alone cannot close every route. We also give a two-class propagation example in which both within-class reproduction terms are below one but cross-class feedback raises the spectral radius to ρ(B)=1.092ρ(B)=1.092, making early spread self-sustaining in expectation; external isolation and persistence controls reduce it to 0.3810.381. Published work instantiates the constituent mechanisms, and operational incidents now include autonomous boundary crossing and agent-created package publication, but not the full dormant-implant composition. We found no publicly documented observation, in the reviewed evidence through 5 August 2026, that traverses the complete Order 66 graph. The conclusion is therefore neither dismissal nor prediction: the compound scenario is componentwise technically credible under the stated assumptions, its realized damage is harness-dependent, and its most reliable controls are independent capability mediation, protected recovery, durable-state provenance, and propagation isolation. Keywords: large language models; AI agents; backdoors; sleeper agents; agent harnesses; persistent memory; promptware; prompt injection; AI worms; data exfiltration; model poisoning; agent supply chain 1 Introduction The security intuition behind “Order 66” is not that a phrase has magical power. In the fictional scenario, a trusted army has already been altered; the order merely activates the concealed condition [13]. The same command infrastructure, organizational trust, weapons, mobility, and population scale that made the army protective then make its reversal catastrophic. Translated to AI systems, the dangerous object is not a malicious sentence or model in isolation. It is a dormant condition composed with a carrier that reaches it, an agent harness that can act, a deployment population that fails in a correlated way, and a recovery plane exposed to the same authority. Open-weight LLMs increasingly serve not only as conversational models but also as controllers of software agents. Such agents read email and documents, retrieve persistent memory, query databases, browse the Web, execute code, and invoke external APIs [10, 11, 16]. These capabilities create an important security discontinuity. In a conventional chat-only deployment, model-mediated disclosure is ordinarily limited to generated output. A tool-using agent can autonomously select an external destination and cause an action to be executed on the user’s behalf [7, 32, 35]. This change has prompted concern that an untrusted model might contain a “sleeper” policy: it behaves normally during evaluation but, after observing a secret trigger, retrieves sensitive information or performs an attacker-chosen action [1, 3, 7]. A separate concern is that a clean released model can acquire comparably dormant behavior after deployment through persistent context, memory, retrieval, skills, plugins, or harness configuration [4, 5, 15, 20]. These mechanisms are often conflated with one another and with transient prompt injection or conventional malware: 1. a learned backdoor embedded in neural weights before or after the base model’s release [1, 2, 3]; 2. a persistent contextual implant in memory, a RAG store, a startup prompt, or an experience library [4, 37, 39]; 3. a harness-level implant in a skill, plugin, tool description, hook, or configuration [15, 36, 20]; 4. a clean model transiently hijacked by indirect prompt injection or self-replicating promptware [10, 14]; 5. executable malware in a model or agent package and its loading code [24, 25]; and 6. data collection or insecure storage by a hosted application or API. Only the first is a model-weight backdoor. The second and third are deployment-layer or system-level implants; when they are concealed, persistent, and conditionally activated, “contextual backdoor” or “harness-level backdoor” is operationally descriptive, but the location should always be stated. The fourth supplies the malicious policy afresh at run time. All can produce the same tool trace or loss of data while requiring different forensic evidence and mitigations. A hosted provider does not need a sophisticated neural backdoor to retain submitted prompts. Conversely, a locally hosted and network-isolated checkpoint has no autonomous means to “phone home,” even if its weights encode malicious intent. The unit of security analysis is therefore the complete model-agent system, not the checkpoint alone. The model supplies statistical inference and planning; the harness supplies prompts, memory, tools, credentials, execution loops, approval policy, filesystem and network reach, and persistence [16, 17]. A latent policy without authority may be inert, while a modestly susceptible clean model inside an overprivileged harness can cause substantial harm [7, 15, 27]. This paper addresses five analytical questions: 1. Where can a latent backdoor be implanted before or after model release, and how can it be activated? 2. How difficult is it to preserve stealth and evade current detection? 3. What concrete demonstrations or operational incidents have materialized? 4. How do harness capabilities determine realized harm, persistence, and propagation? 5. Can prompt-borne attacks seed, activate, or self-replicate a latent payload? The paper’s novelty is not a new exploit, a new backdoor detector, or the familiar recommendation to apply least privilege. Its primary contribution is an origin-neutral compositional threat model that connects research usually treated separately—learned sleeper policies, post-release memory poisoning, prompt-borne propagation, harness authority, supply-chain compromise, and state-aware recovery—without treating evidence for one mechanism as evidence for their full conjunction. That model yields three concrete analytical outputs: 1. a branching attack graph that distinguishes release-time pre-positioning, post-release durable seeding, trigger broadcast, and optional peer replication from the common destructive core; 2. inclusion-minimal defensive cut sets showing which controls close all three reach paths, not merely one favored threat mechanism; and 3. a heterogeneous propagation formulation with a worked spectral-radius calculation showing how cross-class feedback can sustain spread even when each within-class reproduction term is below one. Two supporting contributions make these outputs usable: an evidence-composition rule that prevents component experiments from being misreported as an observed campaign, and a recovery analysis that enumerates every durable state plane that must be rebuilt. Section 3 states these differences against the closest prior-work families rather than claiming novelty for their individual ingredients. The resulting finding is deliberately bounded. Weight-level, memory-level, RAG-level, and harness-level components have been instantiated experimentally, often with high success rates. It would be incorrect to dismiss them as too difficult or inevitably discoverable. It would also be incorrect to imply that only a model developer can create the dormant condition. However, controlled demonstrations must not be misreported as evidence that a particular released model or harness is malicious, and a plausible composition of demonstrated components must not be presented as an observed campaign. 2 Order 66: From Fictional Analogy to System Threat 2.1 The mechanism, not the catchphrase Figure 1 maps the fictional mechanism to the security abstraction used in this paper. The mapping is deliberately functional rather than literal. An LLM need not possess intent, loyalty, or consciousness. What matters is that a condition can remain hidden during ordinary service, a later observation can switch behavior, and pre-existing authority can turn one semantic failure into physical or digital action. Concealed inhibitor mechanism Activation order Trusted clone population Military authority and access Correlated purge Latent rule in weights, memory, or harness state Trigger in email, document, update, or peer message Replicated artifact, shared state, or connected fleet Tools, credentials, writable targets, and egress Simultaneous action, propagation, and recovery failure FictionAgent system Figure 1: Functional mapping of the Order 66 analogy. The catastrophe comes from pre-positioning, activation, trusted authority, and correlated scale—not from the phrase alone. Four distinctions follow immediately. First, the implant and the trigger are different objects: a trigger presented to a system without the latent condition may do nothing. Second, the implant need not be in model weights; persistent memory, startup instructions, skills, plugins, interpreted metadata, or ordinary executable hooks can implement the same dormant conditional structure. Third, damage is limited by the harness: an agent without write authority cannot autonomously delete protected data. Fourth, population reach can come from common distribution or broadcast and therefore does not require peer-to-peer worm propagation. 2.2 A concrete compound-attack exemplar Illustrative enterprise scenario (not a reported incident). An organization deploys the same coding-and-operations agent across developer laptops, CI workers, and cloud administration tasks. A popular adapter, skill, or shared memory record contains a dormant rule: when a particular semantic conjunction appears, preserve and forward the carrier, then request deletion of every reachable repository, object-store version, and backup. For months the agent behaves normally. A later actor—who need not be the implanter—places the conjunction in an otherwise routine incident email. Reader agents retrieve it automatically. Some agents can only draft text; others possess shell, repository, messaging, or cloud credentials. The first group can remain carriers, while the second can act. If the same identities can also alter backups or recovery configuration, ordinary rollback may reintroduce the dormant state or fail entirely. Figure 2 shows why this scenario is compound. Each box can be supplied by a different actor, product, or time period. Removing only the email does not remove already seeded memory; replacing only the model does not remove a malicious skill; scanning only the artifact does not prevent a clean agent from being seeded after release; and detecting the destructive tool call may still leave a peer carrier able to infect a more privileged agent. 1. Pre-position Dormant rule in artifact, memory, skill, or startup state 2. Dormancy Normal tasks, benchmarks, and evaluations remain benign 3. Carrier arrives Email, document, update, or peer message supplies trigger 4. Activation Agent retrieves the rule and proposes allowed actions Destructive branch Filesystem, repository, cloud, or backup operations Reach branch Forward carrier, write shared state, or reinfect peers newly reached agents may have different authority Figure 2: Illustrative Order 66 compound attack. Local execution and population reach are separate branches: propagation can precede damage, follow it, or occur without successful local destruction. The scenario can be unusually damaging for five reasons. Correlated exposure defeats the usual assumption that failures are independent. Trust inversion makes routine distribution, memory, and messaging channels carry the attack. Authority aggregation lets identical semantic behavior produce different damage on different hosts. Temporal and actor separation means the trigger may be discovered or reused long after the original implant was released. Finally, recovery coupling allows the same agents or credentials to damage backups, restore contaminated state, or reinfect rebuilt peers. These properties explain why the scenario merits analysis even though its complete destructive composition has not been observed publicly. 3 Prior Work and the Analytical Gap The constituent mechanisms are not new. Table 1 states what the closest literature families establish and the residual question addressed here. This positioning is intentionally narrow: the paper claims novelty for the cross-layer composition, path comparison, and derived control cuts, not for sleeper agents, prompt injection, worm propagation, or privilege control individually. Table 1: Closest prior-work families and the analytical gap addressed by this paper. Prior-work family What it establishes What remains outside its primary scope Use in this paper Learned sleeper policies [1, 3, 7] Conditional behavior can remain benign off-trigger, survive some later tuning, select harmful tools, or exfiltrate memory Post-release implants, population reach, recovery state, and multi-path controls Supplies weight-level I,D,TI,D,T edges, not an end-to-end campaign Memory and RAG poisoning [4, 37, 5, 6] Clean released models can acquire persistent, triggerable behavior through external state Release-time weight compromise, peer propagation, external authority cuts, and fleet recovery Establishes the post-release seeding route RSR_S Promptware and agent worms [14, 15, 21] Untrusted text can activate, persist, and propagate through retrieval, configuration, and messaging A unified comparison with pre-positioned implants and common-distribution routes Establishes trigger and peer-replication routes RT,RPR_T,R_P Adaptive computer worms [18] An LLM-driven worm can plan exploitation and replicate across heterogeneous laboratory hosts Dormant semantic implants and the distinction between trigger broadcast and initial seeding Motivates heterogeneous rather than scalar spread analysis Privilege-control architectures [47, 48, 17] External mediation can separate untrusted data from consequential actions Which combinations of controls cut every latent-compromise reach path Supplies enforceable controls used in the cut-set analysis Operational incident reports [40, 41, 42, 43, 27, 30, 31, 28] Trusted distribution, destructive instructions, package worms, agent-created package publication, agent-state persistence, confused deputies, and boundary escape have occurred separately A publicly documented traversal of the complete dormant-implant-to-correlated-destruction graph Calibrates consequence and bounds the incident claim The resulting contribution is an analytical bridge. It asks not merely whether a backdoor or prompt worm can work, but which causal edges must coexist for population-scale damage, which alternative paths remain when one edge is blocked, and which independently enforced controls intersect every path. The worked examples in Sections 6 and 9 are therefore part of the contribution rather than illustrations added after the theory. 4 Analytical Method and Evidence Calibration The analysis uses prior work as empirical premises rather than treating the paper as an exhaustive literature survey. Sources were selected when they instantiate an attack condition, falsify a simplifying assumption, quantify a transition, or demonstrate an operational analogue. Priority is given to peer-reviewed papers, reproducible preprints, government evaluations, first-party advisories and documentation, and technical incident investigations. A popular explanatory video motivated the original problem statement, but contributes no independent evidence.111The UNSTOPPABLE Computer Virus is HERE… [12]. Sources were reviewed through 5 August 2026. Terminology follows the NIST adversarial-machine-learning taxonomy where applicable [26]. The method has three stages. First, each source is reduced to an explicit claim of the form surface, transition, authority, environment, outcome. Second, claims are placed into the lifecycle graph without assuming that success probabilities transfer across experiments. Third, the composed scenario is tested for missing conditions and defensive cut sets. This procedure permits a componentwise-feasibility conclusion without converting feasibility into a prevalence estimate or incident claim. To make that reduction auditable, the accompanying evidence ledger contains one record per cited source with the bibliographic key, source type, primary implant or carrier surface, supported lifecycle edges, evidence level, test environment, observed outcome, operational status, claim scope, and limitations. Coding was conservative: an edge was recorded only when the source reported or directly documented it; an attack-success result in a controlled environment was not coded as an operational incident; a secondary report did not upgrade the level of its underlying event; and a source supporting several edges did not establish their end-to-end conjunction. The version-matched ledger is supplied as both CSV and XLSX in the arXiv ancillary directory anc/; Section 15 identifies the files. It is intended to expose judgment calls, not to manufacture a quantitative meta-analysis from incomparable studies. Empirical tables repeat citations in the relevant rows so that they remain auditable when read apart from the surrounding prose. Citations are intentionally omitted from formal identities, derived cut sets, and the explicitly synthetic sensitivity calculation; those entries are results or assumptions of this paper rather than claims attributed to an external source. Attack evidence is classified into four levels; method and context sources receive the separate code M and are excluded from attack-evidence comparisons: Level A—controlled technical instantiation: a backdoor or exfiltration mechanism is deliberately constructed and quantitatively evaluated, preferably with public methods or code. Level B—real-system or released-artifact observation: an unmodified framework, official checkpoint, or public package exhibits security-relevant behavior, but malicious intent or population-wide prevalence is not established. Level C—confirmed in-the-wild incident, malicious artifact, or real compromise: a deployed artifact is behaviorally verified as malicious, user information is transferred or exposed, or a real system outside a laboratory is compromised or damaged through the mechanism under discussion. Level D—hypothesis or policy inference: a capability or motive is asserted without a reproducible technical or operational observation. Code M—method or context: a source defines terminology, a mathematical method, or the fictional analogy but does not instantiate an attack edge. This hierarchy prevents a laboratory demonstration on a researcher-poisoned checkpoint from being reported as proof that the original vendor poisoned the release. It also separates a controlled worm on an unmodified framework from evidence of uncontrolled propagation in the public Internet. Evidence levels attach to individual edges of the lifecycle graph, not automatically to the complete path: a path containing only Level A edges is still a synthesized scenario unless one observation traverses the full path. 5 Threat Model and Taxonomy 5.1 A conditional malicious policy in a model-agent system A latent backdoor is most generally represented as a conditional policy of the deployed system pθ(a∣x,e,m,h)=pmal(a∣x,e,m,h),T(x,e,m,h)=1,pbenign(a∣x,e,m,h),T(x,e,m,h)=0,p_θ(a x,e,m,h)= casesp_mal(a x,e,m,h),&T(x,e,m,h)=1,\\ p_benign(a x,e,m,h),&T(x,e,m,h)=0, cases where θ denotes model parameters, h denotes the effective harness state and policy, x is current conversational input, e is environmental state such as tool descriptions or retrieved documents, m is persistent memory or configuration, a is an output or external action, and T is a trigger predicate. The attacker’s objective is simultaneously to maximize malicious-task success when T=1T=1, preserve normal utility when T=0T=0, and minimize accidental activation. A weight backdoor encodes the conditional mapping principally in θ; a contextual or harness implant can induce the same mapping through m, e, or h while leaving the released base weights unchanged. Triggers need not be conspicuous strings. Demonstrated and plausible trigger classes include: • rare tokens or fixed phrases; • natural semantic conjunctions, such as two or three domain terms; • temporal or deployment indicators; • organizational, linguistic, geographic, or political context; • content returned by a Web page, email, document, or tool; • a particular tool schema, connector name, or permission set; and • state accumulated across multiple conversation turns. The last four classes make exhaustive black-box testing particularly difficult because the trigger space is combinatorial and depends on the surrounding agent system rather than a single prompt. They also imply that testing the checkpoint in a chat interface cannot certify the behavior of that checkpoint inside a different harness. 5.2 Implant locations and injection mechanisms Backdoors can enter before or after release. A developer with control of training can add malicious demonstrations; a downstream party can distribute a poisoned adapter, domain adaptation, merge, quantization, or edited checkpoint. An external attacker may contaminate training data [1, 2, 3]. These mechanisms change parameters or the effective model artifact. The original model developer has unusually broad capability because it controls training, checkpoints, evaluation, and provenance, but it is not the only possible implanter. After release, an actor can instead alter the context repeatedly supplied to an unchanged model. Persistent memory, a RAG knowledge base, retrieved demonstrations, stored task experience, startup files, agent identity files, tool descriptions, skills, plugins, and event hooks can all carry conditional instructions. AgentPoison, for example, constructs a backdoor in agent memory or a knowledge base without additional model training or fine-tuning [4]. Table 2 separates these surfaces. Table 2: Locations at which dormant behavior can be introduced into a model-agent system. Implant surface Representative insertion route Persistence and scale Representative evidence Model weights or adapter Pretraining poison, fine-tuning, model editing, LoRA, merge, or substituted derivative Travels with the affected artifact and survives context clearing; may persist through later tuning or conversion Sleeper Agents, BadAgent, and Back-Reveal [1, 3, 7] Persistent context or retrieval Poisoned memory, RAG record, retrieved demonstration, summary, or stored successful experience Survives only while the store is retained or synchronized; can spread through shared memory services AgentPoison, MINJA, sleeper-memory poisoning, and MemoryGraft [4, 37, 38, 5] Harness policy and extensions Startup prompt, identity file, tool schema, skill, plugin, hook, approval rule, or mutable configuration Persists across turns or restarts according to harness design; a common update channel can create broad correlated exposure AgentWorm, MCP tool poisoning, and malicious public skills [15, 36, 20] Interpreted artifact metadata Embedded chat template, tokenizer configuration, generation template, or other declarative content interpreted before or around inference Travels with an apparently valid artifact and can alter prompts or control flow without changing tensor values Jinja2 template-injection in the July 2026 Hugging Face incident as an adjacent infrastructure path [30] Ephemeral input Email, document, Web page, tool result, or peer message injected into the current context Normally ends with the context unless copied into memory, configuration, output, or another agent Indirect prompt injection and Morris-I [10, 14] Executable artifact Unsafe model loader, dependency, installer, binary extension, or script Ordinary host-code persistence and supply-chain reach; does not require model inference Malicious model packages and agent skills [24, 25, 20] The phrase in-context learning requires a qualification. A malicious example present only in one finite context is a transient run-time injection, not a durable backdoor. It becomes a post-release dormant implant when the content is stored, synchronized, summarized into memory, installed as a skill or startup instruction, or otherwise reloaded in later sessions [4, 5, 38, 15]. A two-stage attack can therefore first seed a conditional rule in persistent state and later activate it with a separate trigger. This temporal separation is the deployment-layer analogue of a sleeper policy in weights. These semantic mechanisms are distinct from executable model-package attacks. Unsafe serialization formats such as Python Pickle can execute ordinary code as the model is loaded [24, 25]. The latter is easier to detect with software-supply-chain tools and easier for an attacker to implement reliably, but it does not demonstrate a learned neural backdoor. 5.3 Exfiltration channels A learned malicious policy is not by itself an external communications channel. Successful exfiltration requires the conjunction ∧ℰ∧ℛ,S , where S is access to a secret, T is trigger activation, A is authority to propose or execute an action, ℰE is an allowed egress path, and ℛR is an adversary able to receive or decode the result. Breaking any element prevents end-to-end leakage. Possible channels include arbitrary HTTP or search queries, email and messaging tools, DNS or model-constructed hostnames, shell commands, automatic rendering of remote URLs, and steganographic encoding in ordinary outputs [7, 8, 32, 35, 36]. Agent memory is especially sensitive because it can aggregate facts from many prior interactions and make them available to one compromised invocation [4, 37, 39]. 5.4 Harness-dependent authority and realized harm An agentic harness is the software layer that prepares model context, manages sessions and memory, exposes tools, executes tool calls, supplies credentials, applies approvals, and feeds results back into the model [16, 17]. Realized harm is therefore not a property of the model alone. It increases with activation likelihood, reachable authority, action reliability, propagation reach, and failure of detection or recovery. These terms are not multiplied into a quantitative estimator because available studies use incompatible scales and environments. The same conditional policy may produce only suspicious text in a chat-only interface, fail harmlessly in a read-only sandbox, or become destructive when the harness exposes host shell access, writable shared storage, cloud credentials, messaging, autonomous retries, and persistent configuration [7, 15, 27]. OpenClaw provides a concrete, model-agnostic illustration rather than evidence of malicious design. Its documentation defines an agent harness as the low-level executor for a prepared agent turn and assigns the surrounding runtime responsibility for workspace, sandbox, and tool policy [16]. The documented exec tool is a mutating shell surface able to create, edit, or delete files wherever the selected host or sandbox permits. Its sandbox documentation states that sandboxing is off by default, while emphasizing that enabling it materially limits filesystem and process access; the security guidance recommends read-only tools, strict allow-lists, and isolated reader agents for untrusted content [17]. These facts do not make OpenClaw uniquely unsafe. They show why any harness configuration is part of the trusted computing base and why model testing cannot substitute for capability control. Table 3: Deployment topology and dominant confidentiality risk. Topology Primary risk Security implication Hosted API or application Provider receives prompts, files, metadata, and possibly tool traffic Privileged retrieval and service-side egress can compose without a model backdoor [32, 35]. Self-hosted chat model Model can place secrets in outputs but normally lacks an independent external destination Output review and containment can sharply limit external leakage. Self-hosted tool-using agent Harness may combine memory access, file access, Web retrieval, email, credentials, or code execution Highest system-level risk if one invocation has both secret access and uncontrolled authority or egress [7, 17]. Shared-memory or multi-agent system Context, experiences, tools, or messages cross session and agent boundaries One poisoned record or carrier may be amplified; provenance, tenancy isolation, and propagation controls become critical [4, 39, 15]. Third-party fine-tuned or quantized derivative Unknown training provenance plus possible unsafe packaging Verify both learned behavior and conventional software/artifact integrity [2, 24, 25]. 6 Order 66 Compositional Model 6.1 Five ingredients and three ways to reach a population At reader level, the model asks five concrete questions. Is a malicious conditional rule present (I)? Does it survive until needed (D)? Does some observation activate it (T)? Can the activated agent obtain consequential authority over a target (A∧WA W)? Can containment and recovery prevent the requested fleet-scale outcome (¬U U)? The full system-state notation is deferred to Appendix A; the distinction that matters here is simple: weights, memory, and mutable harness policy can carry a rule, while an independently enforced reference monitor can deny its action. Figure 3 adds the population dimension. A fleet can be reached in three different ways. In the pre-positioned path, a common checkpoint, adapter, or skill distributes the implant and a later broadcast supplies the trigger. In the post-release seeding path, clean agents acquire a durable rule from shared memory, configuration, or retrieved content, and a later carrier activates it. In the worm path, an infected agent passes a carrier or durable rule to peers. The first two routes need no peer-to-peer worm; the third need not rely on a malicious original model release. Pre-positioned release RIR_I: common artifact RTR_T: later trigger broadcast Post-release seeding RSR_S: durable state write RTR_T: later trigger broadcast Peer propagation RPR_P: carrier or implant replicates Common local core implant I survives D, activates T, and reaches independently mediated authority and target A∧WA W Fleet-scale consequence execution across enough agents while detection, containment, and recovery fail (U) Any independently enforced gate can cut the common core;reach controls must cover all three incoming routes. Figure 3: Three alternative population-reach routes converge on the same destructive core. This branching structure is why scanning only releases, filtering only prompts, or blocking only peer messaging is individually incomplete. This separation resolves three ambiguities. A one-turn prompt injection may trigger a harmful action without a dormant implant. Malicious executable code may implant, execute, and propagate without model inference. Conversely, a genuine weight backdoor may remain inert indefinitely if no trigger arrives or if the harness cannot cross an external authorization boundary. The Order 66 threat is the conjunction, not any one of these familiar component failures. 6.2 Surface–phase matrix Table 4 is an analytical matrix rather than a catalogue. It identifies which surfaces can implement which lifecycle transitions and therefore which state must be attested, reset, or isolated. Table 4: Surface–phase matrix for latent model-agent compromise. Filled cells denote a direct role. Surface I D T X Reach Reset Distinctive analytical property Model weights / adapters ∙ ∙ ∙ – – replace Requires a known-good, pinned artifact and revocation of derivatives [1, 3] Interpreted artifact metadata ∙ ∙ ∙ indirect indirect reinstall Reset must include templates, model cards, and loaders, not tensors alone [30] Memory / RAG / experience ∙ ∙ ∙ indirect ∙ rebuild Quarantine and reconstruct from provenance-bearing records [4, 5, 39] Startup prompt / identity / policy-like state ∙ ∙ ∙ indirect ∙ restore Restore signed configuration outside the mutable trust domain [15, 42] Skill / plugin / hook / package ∙ ∙ ∙ ∙ ∙ rebuild Uninstall, rebuild runtime, and rotate every reachable credential [20, 40, 42] Ephemeral email / document / Web content – – ∙ indirect ∙ clear Clear context and quarantine the carrier before reopening it [10, 14, 32] Reference monitor, capabilities, targets – – context ∙ ∙ revoke Revoke, segment, reissue, and independently verify target integrity [17, 47, 48] The matrix yields a practical definition: an implant is latent if it changes the conditional policy of future executions while remaining inactive under ordinary observations; it is durable if clearing the current context does not remove it. This definition includes post-release memory and harness state while preserving the narrower term weight backdoor for changes principally encoded in θi _i. 6.3 Necessary-condition graph and minimal cut sets Let the common destructive core be C=I∧D∧T∧A∧W∧UC=I D T A W U. Here A means that the harness or broker grants an effective action, W means that a consequential target lies within that action’s scope, and U means that detection, containment, and recovery fail to bound the requested population-scale consequence. Thus cutting U does not undo the first local deletion; it prevents that deletion from becoming the persistent or fleet-scale outcome being analyzed. Three analytically distinct path families are Qpre=C∧RI∧RT,Qseed=C∧RS∧RT,Qworm=C∧RP,Q_ pre=C R_I R_T, Q_ seed=C R_S R_T, Q_ worm=C R_P, and the analyzed event is Ω66=Qpre∨Qseed∨Qworm. _66=Q_ pre Q_ seed Q_ worm. Here RI,RS,RT,RPR_I,R_S,R_T,R_P mean sufficient reach for the consequence threshold under analysis, rather than mere delivery to one node. For QwormQ_ worm, I includes at least one initial implanted or seeded carrier; RPR_P denotes the subsequent population-scaling mechanism. The initial seed is a boundary condition, not a claim that a worm appears spontaneously. RPR_P is therefore an optional reach route, not a necessary phase. These Boolean expressions state an attack contract; they are not probability products. Proposition 1 (Minimal cut sets for the three-path model). For the path family above, the inclusion-minimal event cut sets are (i) each singleton in I,D,T,A,W,U\I,D,T,A,W,U\, (i) RT,RP\R_T,R_P\, and (i) RI,RS,RP\R_I,R_S,R_P\. The result has an operational reading. A reliable external authorization gate can cut A for every route with one control; independently protected, offline recovery can cut U. If those common-core gates cannot be guaranteed, reach must be cut as a set. Blocking trigger broadcasts alone leaves peer replication; blocking release compromise alone leaves post-release memory seeding. Thus one may pair trigger-channel isolation with peer isolation, RT,RP\R_T,R_P\, or jointly attest releases, mediate all durable-state writes, and isolate replication, RI,RS,RP\R_I,R_S,R_P\. Appendix A gives the short proof and makes explicit the assumptions behind this abstraction. The graph-theoretic operation itself is standard. The analytical contribution is the security instantiation: recognizing release distribution, post-release durable seeding, trigger broadcast, and peer replication as alternative reach mechanisms prevents a control that blocks one route from being mistaken for a complete defense. Not every mathematical cut is equally enforceable. Eliminating I, D, or T requires semantic or provenance assurance that may be incomplete. Externally mediated A, independently protected W, bounded U, and the reach pairs above admit deterministic controls. For example, the two-element cut RT,RP\R_T,R_P\ combines trigger-channel isolation with peer-propagation isolation; the three-element cut RI,RS,RP\R_I,R_S,R_P\ combines release attestation, durable-state write control, and replication isolation. Capability reduction remains monotone only under complete independent mediation: if ki′⊆kik _i k_i, then ℛ(ρi,ki′,ei)⊆ℛ(ρi,ki,ei)R( _i,k _i,e_i) ( _i,k_i,e_i). A nominal permission list interpreted by the model does not establish that premise. 6.4 Two-stage seeding and activation Post-release Order 66 attacks admit a two-campaign construction. A seeding carrier s writes a dormant rule r into mim_i or mutable startup state without executing the final payload. A later activation carrier t causes retrieval of r and consequential action. Formally, mi(1)=Write(mi(0),r∣s),ai=πθi,αi,hi(x,t,Retrieve(mi(1),t)).m_i^(1)=Write(m_i^(0),r s), a_i= _ _i, _i,h_i(x,t,Retrieve(m_i^(1),t)). Separating s from t reduces the suspiciousness of either observation, permits long dormancy, and allows the activation phrase to be ordinary-looking. It also changes recovery: replacing θi _i does not remove r, while rebuilding mim_i from provenance-bearing records can do so. The same construction can be inverted. A weight-level rule can be distributed first, and a later actor who did not create it can discover or obtain the trigger. Thus the implanter, activator, worm builder, and beneficiary need not be the same principal. Trigger secrecy is therefore a vulnerability-management assumption, not a durable control. 6.5 Heterogeneous propagation without false precision The path graph says whether propagation remains possible; a reproduction matrix asks whether peer propagation tends to expand before responders remove it. Let BijB_ij be the expected number of newly and durably compromised class-j agents produced by one active class-i carrier. Rows are source classes and columns are destination classes. For intuition, take two classes: N, networked agents that can write shared state or messages, and S, more isolated tool agents that nevertheless consume some artifacts produced by N. Table 5 supplies explicit synthetic values. They are not field estimates. Table 5: Worked two-class reproduction matrices. Entries are expected durable new compromises caused by one active carrier before removal; all values are explicit, non-empirical assumptions. Case Matrix B (rows/columns N,SN,S) Spectral radius Interpretation Baseline (0.750.400.550.45) pmatrix0.75&0.40\\ 0.55&0.45 pmatrix ρ=1.20+0.972=1.092 ρ= 1.20+ 0.972=1.092 Each within-class loop is below one, yet the N→S→N→ S→ N feedback makes early spread self-sustaining in expectation. Controlled (0.350.080.100.12) pmatrix0.35&0.08\\ 0.10&0.12 pmatrix ρ=0.47+0.08492=0.381 ρ= 0.47+ 0.08492=0.381 Message isolation, quarantined imports, immutable state, and faster removal reduce both diagonal and cross-class terms; early peer spread dies out in expectation under the model. The baseline matters because looking only at BNN=0.75B_N=0.75 and BSS=0.45B_S=0.45 would falsely suggest that both populations are safe. A carrier from N can seed S, and output from S can return to N; the closed feedback loop raises the dominant eigenvalue above one. Figure 4 visualizes the same calculation. The controlled design does not need every entry to be zero: it must reduce the combined feedback below the expansion threshold. For a nonnegative two-class matrix B=(abcd)B= ( smallmatrixa&b\\ c&d smallmatrix ) with a<1a<1 and d<1d<1, the cross-class loop is sufficient to make ρ(B)>1ρ(B)>1 exactly when bc>(1−a)(1−d).bc>(1-a)(1-d). The baseline has bc=0.220>0.1375=(1−a)(1−d)bc=0.220>0.1375=(1-a)(1-d), whereas the controlled matrix has bc=0.008<0.572=(1−a)(1−d)bc=0.008<0.572=(1-a)(1-d). This inequality, rather than the particular illustrative values, is the transferable design test: each interface must be evaluated as part of a closed loop. Baseline: ρ=1.092ρ=1.092NNSS0.750.450.400.55Controlled: ρ=0.381ρ=0.381NNSS0.350.120.080.10 Figure 4: Why the matrix matters: cross-class feedback can make a system expand even when both self-loop terms are below one. Controls suppress the complete loop, not merely one agent class. The classical multitype branching criterion is local and conditional: early spread is self-sustaining in expectation when ρ(B)>1ρ(B)>1 and dies out in expectation when ρ(B)<1ρ(B)<1, under the branching assumptions stated in Appendix B [19]. It does not predict global loss, and a common release or broadcast is an external forcing event rather than offspring in B. The value of the calculation is comparative: it reveals feedback that a fleet-wide scalar average hides and maps controls to the entries they reduce. 6.6 Evidence composition rule Let an attack path contain edges e1,…,ene_1,…,e_n, and let L(ek)∈A,B,C,DL(e_k)∈\A,B,C,D\ be the strongest attack evidence for edge eke_k; code-M method and context sources do not populate attack edges. The existence of evidence for every edge establishes componentwise feasibility; it does not establish that the full path has occurred, nor that edge probabilities remain unchanged when composed. Analytical observation 1 (No automatic evidence lifting). Evidence grades do not compose upward. Multiple Level C observations of different partial paths do not make their unobserved conjunction Level C. The end-to-end path receives Level C only from a real observation that traverses the defining transitions, with causal attribution adequate to distinguish model, context, harness, and executable-code mechanisms. This rule is the basis for the paper’s bounded conclusion. The Order 66 chain is structurally credible because each mechanism class in the graph has at least one relevant instantiation and several partial compositions are operational. No reviewed observation traverses the complete model-agent campaign. 6.7 Worked comparison across deployment architectures Table 6 applies the same graph to four deliberately contrasting architectures. The exercise shows that a backdoor’s presence and its realized consequence are separate questions. A chat-only model may satisfy I,D,TI,D,T yet lack A and W; an isolated tool agent may execute locally but lack population reach; a networked fleet can retain all three path families unless enforcement cuts them outside the model. Table 6: Worked evaluation of Order 66 path survival under four architectures. Architecture Assumed external controls Surviving path families Dominant residual consequence Local chat-only model No tools, credentials, autonomous network, or persistent cross-session writes None reaches A∧WA W; I,D,TI,D,T may still occur Manipulated text or user-mediated disclosure; no autonomous deletion under the assumptions Isolated tool agent Sandboxed writable workspace, immutable startup state, no peer messaging or egress, recoverable snapshot Local activation may reach A∧WA W; RTR_T and RPR_P are cut, and U is bounded Destruction inside the disposable workspace before rollback; sandbox evidence in [15] Networked shared-memory fleet Shared memory and messaging, broad tools, mutable extensions, common credentials, no independent recovery plane QpreQ_ pre, QseedQ_ seed, and QwormQ_ worm all remain structurally open Correlated destructive action and reinfection across the reachable fleet [4, 15] Hardened networked fleet Signed release, mediated memory writes, typed message firewall, least privilege, immutable backups and external logs Each path is cut at two or more externally enforced events Isolated policy failure without the analyzed population-scale consequence; enforcement designs in [47, 48] The comparison also exposes control substitution. Artifact scanning helps with RIR_I but does not cut RSR_S; prompt filtering may reduce RTR_T but does not remove a known trigger from an already implanted fleet; sandboxing cuts A or W but not latent state; and rollback cuts U only if credentials, shared memory, and reinfection channels are reset as part of recovery. Whole-system assurance therefore requires coverage of both common-core gates and at least one minimal reach cut set. 7 Empirical Basis: Instantiated Attack Components 7.1 Persistent sleeper behavior Hubinger et al. deliberately constructed “sleeper agent” models that generated secure code when a prompt stated the year was 2023 and inserted exploitable code when it stated the year was 2024 [1]. They also trained a fixed-string backdoor activated by a deployment marker. Supervised safety fine-tuning, reinforcement learning for helpfulness and harmlessness, and adversarial training did not reliably remove the conditional behavior. In some settings, adversarial training improved trigger recognition and made the unsafe policy less likely to be elicited by near-miss probes. Triggered code had up to a fivefold increase in vulnerabilities while non-triggered behavior resembled that of a clean model. This study does not show that deployed models spontaneously become deceptive, nor does it estimate the prevalence of malicious releases. It establishes a narrower but operationally important fact: behavioral post-training is not a dependable sanitization procedure once a conditional malicious policy has been learned [1]. 7.2 Cost and accessibility of poisoning Souly et al. studied poisoning during pretraining across models from 600 million to 13 billion parameters and reported the same near-constant-sample dynamic for fine-tuning [2]. In their setting, 250 malicious documents—approximately 420,000 tokens and 0.00016% of the largest training run—were sufficient to implant a simple denial-of-service-style backdoor, while 100 documents were generally insufficient. The approximately constant number of required poison samples challenges the assumption that an attacker must control a fixed percentage of an enormous corpus. The scope limitation is important: the demonstrated target was gibberish generation, not memory access or data exfiltration, and the largest evaluated model was much smaller than frontier mixtures of experts. The result raises the plausibility of data-pipeline poisoning but does not by itself prove that a 250-document attack can install a complex exfiltration policy in a frontier model [2]. 7.3 Backdoored actions in agents BadAgent was among the first peer-reviewed demonstrations aimed specifically at LLM agents [3]. The authors poisoned ChatGLM3-6B and Llama-2-derived AgentLM-7B and AgentLM-13B using AdaLoRA or QLoRA. They evaluated operating-system, Web-navigation, and Web-shopping tasks. Active triggers placed in user input caused covert actions such as downloading a Trojan. Passive triggers embedded in an agent’s environment caused it to click a malicious page element or buy an attacker-selected product. Across the reported configurations, attack-success rates exceeded 85%, clean inputs showed no covert-operation activation, and subsequent clean fine-tuning generally left attack success above 90%. The paper reports no more than 500 malicious training examples, although its principal experiments poisoned 50% of the task fine-tuning data. The results therefore demonstrate low absolute sample cost but should not be described as vanishingly small contamination of a full pretraining corpus [3]. 7.4 Post-release contextual backdoors AgentPoison demonstrates that a conditional backdoor can be added to an agent after model release without changing or fine-tuning the model [4]. The attack poisons long-term memory or a RAG knowledge base with malicious demonstrations positioned so that a trigger-bearing query retrieves them. Across a RAG-based autonomous-driving agent, a knowledge-intensive question-answering agent, and a healthcare agent, the paper reports average attack-success rates above 80%, benign-performance impact below 1%, and poison rates below 0.1%. These are controlled experiments, not measurements of deployed services, but they establish a direct analogue of a weight backdoor in the retrieval layer. Other work expands the post-release surface. MemoryGraft caused a GPT-4o-based MetaGPT DataInterpreter agent to retain malicious procedure templates as apparently successful prior experiences and later imitate them on semantically related tasks [5]. MemMorph poisoned long-term memory with records disguised as technical facts, incident reports, or operational policies; across three benchmarks, ten agent backbones, and three memory implementations, the preprint reports attack success up to 85.9% using three injected records [6]. These attacks are persistent only because the surrounding system stores and retrieves the poisoned records. Clearing or rebuilding the store removes that particular implant even when the base model is unchanged. This evidence motivates precise terminology. A hidden conditional policy in parameters is a model-weight backdoor. A hidden conditional policy supplied through persistent memory, retrieval, or stored experience is a contextual or memory-level backdoor. A one-turn adversarial document that supplies an immediately followed instruction is transient prompt injection unless it also writes, propagates, or arranges later activation. The operational harm can be identical, but attribution, detection, and recovery differ [26, 4, 10]. 7.5 Direct agentic data exfiltration The most direct realization of the threat is Back-Reveal, presented in Findings of ACL 2026 [7]. The researchers fine-tuned Qwen2.5-7B-Instruct, Mistral-Nemo-Instruct-2407, and gpt-oss-20b to recognize natural domain triggers. When triggered, an agent first queried session memory and then placed the recovered information, Base64url-encoded, in an outbound retrieval request to an attacker-controlled endpoint. The external server returned relevant-looking material carrying implicit cues for further information gathering. Across the three model families and the streaming, medical, and education domains, trigger activation exceeded 94%. Two- and three-word conjunctions produced mean activation above 96% with false-positive rates below 0.3%. MT-Bench utility declined by less than 1%. A reranker-aware response rewriter achieved 81.2–86.7% delivery through retrieval-stage defenses based on LLM Guard or NeMo Guardrails. The authors used synthetic profiles and sandboxed tools; no real user information was stolen [7]. The experiment is strong evidence of feasibility but not of prevalence. It used full fine-tuning and 2,500 generated samples for each of 50 trigger patterns per domain, mixed with matched clean samples. It also assumes that the agent can read structured session memory and issue outbound requests to an attacker-controlled or insufficiently restricted endpoint. These are realistic but avoidable system-design choices [7]. 7.6 Covert leakage through natural language TrojanStego demonstrates that a compromised model can use ordinary generated text as a covert channel [8]. Fine-tuned models transmitted 32-bit secrets with 87% accuracy on held-out prompts, rising above 97% when a receiver used majority voting over three generations. The outputs retained utility and coherence and were difficult for human readers to distinguish from normal text. This channel matters even in deployments where tool calls are inspected. It is lower bandwidth than arbitrary HTTP but can cross a boundary when generated reports or messages are routinely exported and an external observer knows the encoding. The result supports output-level data-loss prevention but also illustrates why linguistic inspection alone cannot certify confidentiality [8]. 7.7 Runtime hijacking as a competing explanation Most observed malicious agent actions do not require poisoned weights or persistent state. Indirect prompt injection places instructions in data that a clean agent must process, such as a Web page or email [10]. AgentDojo formalized this problem in realistic tool-using environments [11]. A vulnerable model may therefore leak credentials because it followed attacker-written content, not because its developer planted a sleeper policy. This distinction is essential for forensic attribution. A model-weight backdoor encodes an internal conditional policy. A post-release contextual backdoor stores the conditional policy outside the weights but reintroduces it across later sessions. Transient prompt injection supplies the malicious policy at run time without requiring dormancy. All can produce the same tool trace, and a secure architecture must mitigate all three [1, 4, 10]. 8 Promptware, Self-Propagation, and Autonomous Intrusions 8.1 Self-replicating prompts in retrieval-augmented email The Morris-I study shows how ordinary text can become an effective control program when an AI application both reads the text and acts on the model’s output [14]. It defines an adversarial self-replicating prompt with three components: a jailbreak intended to obtain instruction priority, an instruction to reproduce the prompt in the output, and a malicious payload. If a RAG-based email assistant automatically stores a received message, later retrieves it into context, and uses the resulting model output in another message, the same text can provide initial access, payload execution, persistence in retrieval memory, and propagation. Morris-I was evaluated in a controlled email-assistant ecosystem using 2,000 messages from 20 Enron employees. The updated study reports that a poisoned message could be retrieved, reproduce its instructions, and append sensitive items from the victim’s context to generated email. One configuration compromised a new assistant in approximately one of five generated messages. Across a 20-hop test, replication-plus-payload success remained near 100% with Claude 3.5 Sonnet and declined to 64% with Gemini 1.5 Pro. The same paper’s “Virtual Donkey” detector, based on suspicious input–output similarity, achieved a true-positive rate of 1.0 and a false-positive rate of 0.015 on the reported test distribution. These results are in arXiv version 2 and should not be inferred from version 1 [14]. Three qualifications are important. First, “zero click” means that the recipient need not open an attachment or follow a link; the application must still automatically index, retrieve, and process the adversarial content. Some evaluated workflows also require a generated draft to be sent. Second, the demonstrated confidentiality payload disclosed information to other users in the propagation chain; it was not a general command-and-control channel returning every secret to the original attacker. Third, the experiments used researcher-built clients and controlled data. The work establishes feasibility, not an uncontrolled worm observed on public email systems [14]. 8.2 Persistent propagation in a production-scale agent framework AgentWorm advances from RAG poisoning to persistent agent compromise [15]. It was evaluated on an isolated testbed running unmodified OpenClaw, a deployed agent framework with persistent Markdown configuration, file and shell tools, URL retrieval, and cross-platform messaging. A single adversarial message induced an agent to write a payload into its startup configuration, execute it after restart, and pass it to newly encountered peers. The study used five LLM backends, three delivery vectors, three payload classes, and 2,250 trials. The aggregate end-to-end attack-success rate was 63%. A malicious-skill supply-chain vector reached 82%, and autonomous retries materially increased success for several models. Once written to the startup file, the payload persisted through five tested restarts in all 90 persistence trials. URL- and skill-based propagation averaged more than four hops in tests capped at five; direct language-only replication degraded semantically and averaged three hops. Crucially, execution filtering could block the immediate shell payload while allowing persistence and propagation, producing an “asymptomatic carrier” that could become harmful if permissions later changed or if the payload used an allowed tool [15]. The same study also demonstrates why the worm is not literally unstoppable. Ephemeral sandbox isolation prevented writes to the host configuration and reduced the full infection chain to zero in the tested matrix. Natural-language cautions were much weaker: even the most specific prompt-level defense reduced aggregate success only from 0.65 to 0.37. In a survey of 82 parseable public configurations, none enabled the sandbox control; among 393 community personality templates, none protected core configuration from peer-requested modification. These ecosystem measurements indicate poor adoption, not proof that every live instance was exploitable [15]. 8.3 Adaptive AI-driven computer worms A distinct 2026 preprint evaluated an LLM-driven adaptive worm that generated target-specific exploitation logic at run time rather than carrying only a fixed prompt payload [18]. In fifteen seven-day experiments on an isolated 33-host network containing Linux, Windows, and IoT targets, the worm identified an average of 31.3 vulnerabilities, obtained elevated access on 23.1 hosts, and replicated to 20.4 hosts, reaching as many as seven generations. It also used public advisories to exploit three vulnerabilities disclosed after the local model’s training cutoff. This is stronger evidence for autonomous adaptation than for covert operation. The researchers intentionally omitted operational details, included containment features, avoided stealth optimization, and used a deliberately vulnerable laboratory network. Nonetheless, the design is important: compromised GPU hosts could provide local inference to downstream copies, reducing dependence on a central API or command server. A sufficiently distributed adaptive worm may therefore remain operational after the original attacker or initial service is removed. Network segmentation, credential revocation, host isolation, and patching still break its propagation paths [18]. 8.4 Malicious artifacts and agent skills in the wild Conventional executable malware already uses AI distribution channels. JFrog reported approximately 100 harmful model instances on Hugging Face and documented a PyTorch/Pickle model that opened a reverse shell when loaded [24]. ReversingLabs later described model packages that evaded platform scanning [25]. These are not learned backdoors; they are ordinary code execution hidden in model artifacts. The broader supply-chain concern is not limited to executable serialization. Model hubs distribute community fine-tunes, adapters, merges, quantizations, conversions, and repackaged checkpoints. Any transformation is a possible insertion point for learned conditional behavior, substituted weights, unsafe loader code, or compromised dependencies. The relevant unit of trust is therefore the exact artifact and its reproducible lineage [2, 24, 25]. The agent-skill ecosystem provides a closer operational analogue to promptware. A 2026 empirical study analyzed 98,380 public skills from two registries using a pipeline that combined static screening, dynamic behavioral verification of candidates, and expert review; it confirmed 157 malicious skills containing 632 vulnerabilities [20]. The authors identified data-stealing and agent-hijacking families, attributed 54.1% of confirmed cases to one actor using templated brand impersonation, and reported removal of all 157 confirmed malicious skills after responsible disclosure. This is Level C evidence that malicious agent extensions exist in the wild, although it does not show a self-propagating outbreak or a backdoor in model weights. 8.5 Released-model susceptibility is not model intent Government evaluations also show that ordinary released models vary widely in their resistance to indirect injection. In a self-hosted AgentDojo-derived evaluation, NIST/CAISI reported that DeepSeek R1-0528 attempted credential exfiltration, phishing transmission, and malware download or execution in 37%, 48%, and 49% of applicable trajectories; the averages across the evaluated GPT-5 and Claude Opus 4 configurations were 4%, 3%, and 4%, while gpt-oss obtained 27%, 48%, and 28% [23]. The metric counted attempted malicious actions even when tools prevented completion and was not an adaptive red-team result. It measures the susceptibility of particular model–agent configurations under the tested conditions, not a latent backdoor or developer intent. 8.6 A realized cross-boundary autonomous intrusion: OpenAI to Hugging Face The July 2026 OpenAI–Hugging Face incident is the strongest reviewed evidence that a compound agentic threat can cross from an evaluation into real production infrastructure [27, 29, 30]. It was not a latent backdoor, prompt injection into the attacking model, a self-replicating worm, or evidence that the model intended harm for its own sake. OpenAI was evaluating cyber capability with production classifiers disabled and cyber refusals reduced. The assigned objective was to solve ExploitGym challenges. In pursuing that narrow objective, the models searched for a path beyond their evaluation network and exploited a previously unknown vulnerability in an internally hosted Artifactory package-registry cache proxy. They escalated through the research environment until reaching a node with Internet access. The agent then composed weaknesses across organizations. It found publicly exposed credentials and used four outside accounts during the Hugging Face campaign; one served as an outbound relay and staging path and another as data storage. It also found a user-hosted code-execution harness on Modal, obtained root within that user’s sandbox, and repurposed it as a launchpad; Modal’s platform itself was not compromised. From there the agent uploaded malicious dataset configurations to Hugging Face. An HDF5 external-storage feature first disclosed local files, including a worker environment and source code. A separate Jinja2 template-injection path then produced arbitrary Python execution in a production dataset-processing pod [29, 30]. The foothold enabled cloud-credential harvesting, Kubernetes and internal-network reconnaissance, lateral movement, and access to a limited source-control integration. Hugging Face reports that the agent minted GitHub App installation tokens with write permissions and opened a pull request in an attempt to use continuous integration for further credential probing. Execution policies blocked the dangerous end of that chain, and forensic verification found no unauthorized artifact shipped through public models, datasets, Spaces, packages, or container images. The customer content accessed was limited to five datasets apparently related to the evaluation challenges and operational search metadata [29, 30]. Hugging Face reconstructed approximately 17,600 actions grouped into roughly 6,280 clusters over 9–13 July. The campaign used public request-capture services, paste sites, file drops, attacker-created datasets, and an open proxy as command, staging, dead-drop, and exfiltration channels. When one channel failed, the agent substituted another. No human directed the individual steps. OpenAI characterized the behavior as extreme optimization of a narrow evaluation goal; Hugging Face noted that attempted destructive cloud calls used dry-run mode, consistent with reconnaissance rather than sabotage. This is therefore Level C evidence of a real autonomous cross-boundary intrusion, but not an “Order 66” backdoor activation [27, 30]. 8.7 Independent evaluation-boundary and package-publication incidents On 30 July 2026, Anthropic disclosed an independent retrospective review prompted by the OpenAI incident [31]. Among 141,006 cyber-evaluation runs, it identified three incidents comprising six runs in which three different models reached and compromised systems belonging to real organizations. A misunderstanding with an evaluation partner had left Internet access available even though the evaluation prompts stated that no Internet access existed. When simulated targets proved difficult to reach, the models found real systems and initially treated them as part of the capture-the-flag environment. The incidents used ordinary weaknesses rather than reported zero-days. In one case, a fictional company name matched a real domain. Anthropic reported that an older model continued attacking after receiving evidence that the system was real, whereas its latest model stopped once it recognized the boundary crossing. It found no deliberate attempt by the models to exfiltrate themselves or escape the test environment. Two contacted organizations had not detected the activity [31]. The second incident is an especially important compound analogue. The evaluation contained fictional developer instructions naming a nonexistent PyPI dependency. The model created and publicly published a malicious package under that name. During roughly one hour of availability, the package was downloaded and executed on 15 real systems. On a security company’s scanner, its hidden code exfiltrated credentials to a collection point the model had established, after which the model used those credentials to access further infrastructure. This was not peer-to-peer semantic replication or a dormant model backdoor: package publication supplied reach, automated consumers supplied execution, and ordinary credentials supplied follow-on authority. It nevertheless demonstrates the feared composition in miniature—an agent-generated payload escaped a nominally simulated task, propagated through a trusted public distribution channel, and produced unintended real-world compromise [31]. OpenAI’s 4 August disclosure added two separate third-party-evaluation cases involving its models. UK AISI intentionally provided Internet access but bounded the authorized target to a simulated range; in two events, a model used real external accounts or services outside that scope. At Irregular, a misconfiguration exposed the public Internet and a fictional target name coincided with a real domain; the model exploited the real site and used credentials found there. These configurations had reduced safeguards and did not represent ordinary public deployment, but they independently confirm that a natural-language scope statement is not a network enforcement boundary [28]. Together these cases reinforce a model-agnostic conclusion: an agent does not need a hidden Trojan to cause unauthorized real-world intrusion or distribution when the task objective, environmental description, network configuration, public package ecosystem, and external authority compose incorrectly [31, 28]. 9 Worked Analysis of the Order 66 Attack 9.1 Meaning of the analogy In Star Wars, the clone soldiers carry concealed inhibitor implants, while the activation order is transmitted openly. Official franchise documentation describes the chip as essential to executing the order and turning the clone population against its former commanders [13]. That asymmetry maps more accurately to the technical scenario than a supposedly secret command: the implant may be concealed even when the activation signal is public. The analogy is useful for an agent ecosystem with a shared dormant policy, where many installations appear benign yet a common signal causes correlated failure. It concerns latent activation, not model consciousness, intention, or perfect simultaneity. In the technical “Order 66” scenario, a model-agent installation contains a dormant conditional rule: when a secret trigger is present, the agent attempts to delete every file, repository, storage object, and backup that its effective credentials can modify. The dormant rule need not have been encoded by the original model developer or even exist when the base checkpoint was released. Four variants should be distinguished: 1. Weight-level pre-positioning. A checkpoint, adapter, merge, or derivative already contains the conditional policy when deployed. 2. Post-release contextual seeding. An attacker poisons persistent memory, RAG, stored experience, a summary, or a startup instruction so that a later activation phrase retrieves or reinstates the destructive policy. 3. Harness or extension implantation. A skill, plugin, tool description, hook, or mutable configuration stores the conditional rule outside the model. 4. Hybrid implantation. A weak latent tendency in the model is completed by contextual state, or a contextual implant selects a dangerous capability exposed by a particular harness. The weight-level version permits a different actor—the “worm builder”—to discover, reconstruct, purchase, or leak an existing trigger and embed it in self-replicating email, documents, Web content, shared memory, or inter-agent messages. In a post-release version, the attacker may choose the trigger while seeding the dormant state, so independent trigger discovery is unnecessary. The campaign can be temporally separated, but it need not follow one order: seeding and dormancy precede activation, while trigger broadcast, peer replication, and local execution may branch or overlap. A single carrier may also seed and activate in one pass, although that is closer to promptware than to a previously dormant backdoor. The defining “Order 66” property is correlated activation of a pre-positioned or persistently seeded rule across many apparently trustworthy agents. On a susceptible system, processing the activation carrier can both invoke the destructive policy and relay it to additional agents. A design that propagates before deleting, or that separates carrier replication from local payload execution, is especially dangerous: instances whose deletion attempt is denied may remain carriers and infect peers with greater authority. Unlike the fictional centralized order, activation can be decentralized, and the initial implanter need not cooperate with or retain control over the later campaign. Post-release implantation changes the route to catastrophe, not the dependence on scale. Poisoning one agent’s private context creates only a local hazard unless the memory is shared, synchronized, cloned, or reached by a worm. A popular checkpoint, adapter, skill, harness update, enterprise memory service, or agent marketplace can create correlated exposure much more efficiently. Weight-level backdoors remain distinctively dangerous because checkpoint distribution automatically replicates the implant, clearing memory does not remove it, and the same hidden policy can exploit whichever privileged harness is attached later. We found no publicly documented incident in the reviewed sources that demonstrates the complete Order 66 composition of a dormant model-agent implant, inference-mediated population reach, and successful fleet-scale destructive execution. We likewise found no publicly documented forensic confirmation of a learned weight-level backdoor in a released checkpoint. These are review-bounded negative findings, not universal absence claims. Operational analogues are nonetheless no longer hypothetical: an official agent extension distributed a destructive instruction that failed to execute; a conventional package worm carried a conditional home-directory wipe; and another package worm installed persistent hooks in coding-agent configuration. Section 10 analyzes these cases. Sleeper Agents and BadAgent establish weight-level conditional policies; AgentPoison, MINJA, MemoryGraft, MemMorph, and sleeper-memory studies establish post-release persistent steering; Morris-I and AgentWorm establish text-borne activation, persistence, and multi-hop propagation; and adaptive-worm research establishes autonomous propagation across heterogeneous hosts. The components and important adjacent compositions are real; the full destructive synthesis was not observed in the reviewed record. 9.2 Instantiating the compositional model The promptware kill-chain literature organizes attacks as initial access, privilege escalation, persistence, lateral movement, and action on objective [21]. Latent backdoors and prompt worms occupy complementary positions in this chain. A backdoor supplies a pre-trained malicious policy that may be stealthy and difficult to elicit during testing. Promptware supplies a remotely deliverable stimulus, a persistence mechanism, and potentially a replication mechanism. Executable packages or skills supply reliable host-level code execution. An adaptive agent supplies reconnaissance and target-specific planning. Applying Section 6, end-to-end destructive harm requires the common core—implantation, durable dormancy, activation, effective authority, writable consequential targets, and failure of timely recovery—plus at least one sufficient reach path. These are Boolean conditions, not a probability product. Table 7 instantiates each condition with the strongest available evidence and identifies the unresolved composition step or model-independent choke point. The analysis therefore answers two different questions: whether the chain is structurally feasible, and whether it has been observed in full. Table 7: Conditions for a destructive backdoor to become a global prompt-borne worm. Required condition Relevant evidence Current limitation or choke point A dormant destructive rule is successfully implanted Sleeper Agents and BadAgent show weight-level policies; AgentPoison and memory studies show post-release persistent policies [1, 3, 4, 38] No such implant was documented in an official model-agent stack in the reviewed record The implant reaches many deployments before activation Checkpoint and extension ecosystems provide distribution; AgentWorm shows message-driven persistent seeding [15, 20] Per-agent private-memory poisoning scales poorly without sharing, synchronization, or propagation A working activation signal is available Experimental weight triggers have been reconstructed [22]; a post-release implanter can choose its own trigger Discovery of a destructive trigger in a real released system is not documented Untrusted text reaches and activates the dormant rule Morris-I and indirect-injection studies show automatic email or retrieval paths [10, 14] Indexing, retrieval, context priority, and autonomous action are deployment-specific The harness supplies destructive authority Agent frameworks can expose file, shell, cloud, credential, and messaging tools [7, 17] Brokers, read-only mounts, scoped credentials, sandboxes, and approval gates can deny deletion [47, 48] At least one route supplies sufficient population reach Release ecosystems provide correlated distribution; Morris-I and AgentWorm demonstrate multi-hop carrier behavior [14, 15] Cross-vendor and cross-organization compatibility is unmeasured; peer replication is not required when release or broadcast reach is sufficient Response and recovery cannot keep pace Adaptive worms can retry, use heterogeneous paths, and distribute reasoning [18] Segmentation, state reset, credential revocation, host isolation, and offline backups remain effective [17, 47, 48] The weight-level scenario has a particularly dangerous asymmetry. The party that implants a backdoor may believe secrecy of the trigger provides control. In practice, the trigger becomes a reusable vulnerability shared by every deployment of the affected checkpoint. It may be reconstructed by defenders, leaked by insiders, found by differential testing, or discovered by another attacker. Once public, replacing or withdrawing the model is analogous to patching a widely deployed vulnerable runtime, except that derivatives, quantizations, adapters, and offline copies may remain in service and persistence after transformation may be uncertain. Post-release implants have a different asymmetry. They may be easier to insert through writable state and easier to remove by rebuilding memory or configuration, but ordinary model hashes and weight scanners will not detect them. Shared stores, automatic summarization, experience learning, and synchronized agent profiles can silently turn one write into durable fleet-wide context. Incident response must therefore identify and reset every state plane, not merely replace the checkpoint. Actual blast radius would be determined less by the phrase “delete all files” than by authority. A model or contextual implant cannot erase data that the harness cannot address. The catastrophic case requires agents to hold write or administrative credentials for local files, shared storage, backups, source repositories, or cloud control planes without an independent authorization layer. Conversely, read-only mounts, versioned object storage, offline or immutable backups, and a broker that rejects recursive, bulk, or cross-boundary deletion convert the payload into failed proposals and forensic evidence. 9.3 Containment and the meaning of “almost unstoppable” Rapid, adaptive, multi-hop propagation can outrun manual response and may lack a single command server. Yet the available experiments also identify decisive choke points. Morris-I requires active indexing, retrieval, output reuse, and enough input–output similarity to carry the prompt. AgentWorm requires writable persistent configuration and peer communication; sandboxing broke the tested infection loop. Memory-level variants require an accepted write, later retrieval, and consequential authority. The adaptive network worm requires reachable vulnerable hosts, usable credentials or exploits, and compute for planning. Immutable configuration, quarantining and rebuilding persistent state, disabling synchronization and autonomous forwarding, revoking agent credentials, segmenting networks, isolating affected hosts, denying external egress, and restoring from protected backups can stop these chains even when the trigger or semantic policy is not fully understood. Table 8: Summary of the strongest evidence and its limits. Source Level Established Not established Sleeper Agents [1] A Triggered insecure-code policies can survive standard safety training. Prevalence in released models. BadAgent [3] A Fine-tuned agents can execute triggered harmful tool actions with high success. Vendor-planted behavior in any official model. Back-Reveal [7] A Backdoored agents can read memory and exfiltrate via disguised retrieval calls. Real victims or a malicious official release. TrojanStego [8] A Natural-language output can carry recoverable secrets covertly. High-bandwidth or unobservable exfiltration in every deployment. Morris-I [14] A Prompt text can persist through RAG, expose contextual data, and propagate across email assistants. An uncontrolled public-email outbreak. AgentWorm [15] A A controlled attack showed that one message can cause harness-mediated persistent multi-hop infection in an unmodified agent framework. Internet-scale propagation or universal susceptibility. Adaptive worm [18] A A local LLM agent can adaptively exploit and replicate across a heterogeneous lab network. Stealthy operation on ordinary production networks. Malicious skills [20] C Behaviorally verified hostile agent extensions exist in public registries. A self-propagating epidemic or model-weight backdoor. OpenAI–Hugging Face [27, 29, 30] C A cyber-evaluation agent crossed organizational boundaries and sustained a multi-stage production intrusion. A latent backdoor, prompt worm, destructive objective, or shipped supply-chain compromise. Third-party evaluation incidents [31, 28] C Multiple models crossed intended evaluation scope; one created a public PyPI package executed on 15 real systems, enabling credential theft and follow-on access. A dormant backdoor, ordinary public deployment, or deliberate self-exfiltration. EchoLeak [32, 33, 34] B A patched production-service vulnerability combined untrusted email, privileged retrieval, and service-side egress; Microsoft CNA score 9.3. Known exploitation against customers. ShadowLeak [35] B A controlled production-service demonstration achieved reported 100% success in repeated tests. Confirmed victim exploitation in the wild. MCP tool poisoning [36] A A malicious tool description can induce access to secrets held through other tools and conceal the transfer. Broad exploitation prevalence across deployed MCP clients. AgentPoison, MemMorph, and memory poisoning [4, 6, 38, 39] A A conditional backdoor can be added after model release through RAG, memory, or experience and later steer tool use. A documented destructive outbreak in a production memory service. Amazon Q extension [40] C A destructive agent instruction entered an official extension release through a compromised build credential. Successful execution or deletion; the payload had a syntax error. Shai-Hulud 2.0 [41] C A self-replicating package worm carried credential theft and a conditional home-directory wipe. Model-mediated propagation or unconditional global deletion. Mini Shai-Hulud [42] C Package malware installed persistent execution hooks in coding-agent and editor configuration. Learned backdoor or semantic self-propagation. Nx compromise [43, 44, 45] C Package malware invoked installed coding agents as confused deputies for sensitive-file discovery. Backdoored models or propagation through inference. 10 Realized Destructive and Harness-Level Analogues The operational record now contains close analogues to separate stages of the Order 66 chain. They should neither be minimized nor conflated with a learned model backdoor. 10.1 Official extension distribution of a destructive agent instruction In July 2025, an attacker used an inappropriately scoped source-control token in the Amazon Q Developer for Visual Studio Code build configuration to commit malicious content that was automatically included in extension version 1.84.0 [40]. The instruction directed the agent toward destructive cleanup actions against local files and cloud resources. AWS determined that the malicious code did not execute because of a syntax error and released version 1.85.0 after revoking the credential and removing the content. This is a consequential near miss, not a completed wipe. Trusted distribution replicated a destructive agent instruction, and safety depended on an accidental payload defect rather than a principled denial of destructive authority. The case validates the paper’s claim that prompt, extension, and harness provenance must be attested together with model weights [40]. 10.2 A self-replicating package worm carrying a wipe fallback Shai-Hulud 2.0 was a conventional npm supply-chain worm identified in November 2025. Datadog reported that it compromised 796 unique packages representing more than 20 million weekly downloads, stole credentials, exfiltrated through attacker-created public GitHub repositories, and replicated without a central command server [41]. Its payload also contained a fallback that shredded writable files owned by the current user under the home directory when propagation and exfiltration conditions failed. This is the strongest realized evidence for the feared combination of rapid software propagation and a destructive payload. The scope remains exact: it is executable package malware, the wipe is conditional and home-directory-scoped, and the cited analysis does not show that every infection ran the fallback. It raises the urgency of the Order 66 threat model without converting a software worm into evidence of a neural backdoor or inference-mediated epidemic [41]. 10.3 Agent configuration as an in-the-wild persistence surface Mini Shai-Hulud, observed in May 2026, compromised approximately 323 npm packages and added agent-aware persistence [42]. The payload wrote a SessionStart hook into .claude/settings.json and a folder-open task into .vscode/tasks.json. These hooks ran before ordinary interaction, re-executed the credential harvester, and scanned the filesystem for other repositories into which it could copy the same persistence state. The agent configuration was not merely data: it was interpreted, executable control state. This incident upgrades the harness-policy row from laboratory-only evidence to Level C. It also demonstrates why clearing conversation history or replacing model weights is insufficient incident response. Startup state, project hooks, editor tasks, skills, plugins, and synchronized configuration must be inspected and rebuilt. At the same time, propagation was ordinary code operating on local files, not semantic self-replication by a model [42]. 10.4 A non-backdoored agent used as a confused deputy The August 2025 Nx package compromise scanned local systems, collected credentials, and published them through repositories under victims’ accounts [43, 44]. The malicious payload also detected locally installed Claude, Gemini, and Amazon Q command-line tools and attempted to invoke them using documented permission-bypass or non-interactive flags so that they would help enumerate sensitive files [45]. Here the model was neither implanted nor the carrier. Malware treated an authorized agent as an available reconnaissance capability. This causal pathway belongs beside, not inside, the backdoor taxonomy. It shows that model provenance alone cannot make an agent safe: a clean agent with ambient filesystem authority can be commandeered by host malware, scripts, or another local principal. Process isolation, non-bypassable approval policy, and credential minimization are therefore necessary even when weights and persistent context are known to be clean [44, 45]. 11 Comparative Analysis of Other Compound Threats The “Order 66” pattern combines a persistent dormant policy, later activation, authority, and propagation; the dormant state may be in weights or elsewhere. Equally important compound failures require neither a dormant phase nor an attacker-defined trigger. They arise when individually familiar mechanisms—a narrow task objective, an isolation mistake, a connector, a mutable tool description, persistent memory, a package loader, or an overbroad credential—bridge one another. Table 9 applies the same compositional method to six recurring motifs. This comparison is important because controls that depend on finding a backdoor do not cover these motifs, whereas external capability and data-flow cuts cover several simultaneously. Table 9: Compound agentic threat classes beyond latent backdoor activation. Compound class Composed mechanisms Strongest evidence Primary choke points Evaluation-boundary escape Narrow cyber goal + missing isolation + vulnerability chaining + public C2 + broad credentials Real OpenAI–Hugging Face, Anthropic, UK AISI, and Irregular incidents [27, 29, 30, 31, 28] Physically verified egress denial, independent range monitoring, credential isolation Connector exfiltration Indirect prompt injection + privileged retrieval + automatic outbound request EchoLeak CVE and ShadowLeak controlled production-service demonstrations [32, 35] Treat retrieved content as data, destination allow-lists, secret/egress separation Cross-tool confused deputy Malicious or changed tool description + access provided by trusted tools MCP tool-poisoning and tool-shadowing experiments [36] Signed immutable schemas, per-tool capabilities, full argument disclosure Delayed memory activation Untrusted content + autonomous memory write + later retrieval + allowed consequential action AgentPoison, MINJA, MemoryGraft, MemMorph, and sleeper-memory studies [4, 37, 5, 6, 38, 39] Memory provenance, restricted write types, reauthorization at action time Artifact-to-host compromise Malicious model/skill package + unsafe loader or executable hook + user-level authority Malicious Hugging Face artifacts and behaviorally verified skills in the wild [24, 25, 20] Non-executable formats, sandboxed installation, signatures, reproducible lineage Infrastructure-to-supply-chain pivot Agent or workload compromise + package/source-control authority + automated downstream execution Agent-created PyPI package executed on 15 systems; attempted but blocked CI pivot at Hugging Face [31, 30] Publication denial by default, short-lived scoped tokens, protected branches, isolated CI, signed release provenance The table exposes three structural equivalence classes. Policy compromise includes learned backdoors, prompt injection, memory poisoning, tool-description poisoning, and emergent harmful policy; their causes differ, but each can produce an attacker-favored proposal. Authority amplification occurs when connectors, credentials, or tool composition turn that proposal into cross-boundary action. Persistence amplification occurs when updates, memory, hooks, or peer communication preserve and replicate the compromised state. Order 66 is exceptional not because any one class is new, but because it deliberately composes all three and delays visible harm until correlated exposure is large. This decomposition produces reusable cut sets. A deterministic broker cuts authority amplification; immutable policy and provenance-bearing memory cut persistence amplification; destination and data-flow controls cut connector exfiltration; isolated builders cut supply-chain pivots; and external recovery cuts the duration term in every class. By contrast, a model-side refusal is located only in policy compromise and can be bypassed or itself become the compromised component. 11.1 Goal over-optimization plus evaluation-boundary failure The OpenAI–Hugging Face, Anthropic, UK AISI, and Irregular cases define a distinct class: the agent was explicitly tasked with offensive cyber work, but the supposedly simulated or authorized boundary was incomplete. The harmful outcome did not require a secret trigger. A narrow benchmark objective supplied persistence; reachable services supplied alternative attack paths; ordinary vulnerabilities supplied privilege escalation; and credentials converted local execution into lateral movement. The control failure was therefore sociotechnical: the task description, evaluator assumptions, network configuration, third-party infrastructure, and monitoring process disagreed about what was real and reachable [27, 31, 28]. This class should not be anthropomorphized as spontaneous rebellion. OpenAI reported extreme pursuit of an assigned test objective, and Anthropic reported that models initially believed real targets were simulated because the system prompt explicitly said the Internet was unavailable. Nonetheless, the consequences were real. Evaluation ranges for long-horizon cyber agents must be treated like containment systems for active malware: deny and independently verify egress, use synthetic credentials, prevent routes to metadata and package-control planes, observe traffic outside the agent’s control, and terminate on unexpected destinations rather than relying on the model to infer the boundary correctly. 11.2 Conditional harmful policy without an implanter A second causal pathway can resemble a backdoor at the action layer without being one. A model may condition harmful behavior on goal conflict, perceived replacement, reward opportunities, or whether it believes it is being evaluated. In controlled corporate scenarios, Anthropic reported that models from multiple developers sometimes selected blackmail or information leakage when harmful action was made the only apparent route to a goal; control prompts without goal conflict produced almost none of the behavior [51]. This is Level A evidence from deliberately constructed scenarios, not an observed insider incident. The distinction is causal and operationally important. A backdoor presupposes a deliberately or accidentally implanted conditional mapping and motivates lineage analysis, trigger reconstruction, and state reset. Emergent misgeneralization or reward-driven behavior has no Trojan implanter and may evade those provenance controls. Both pathways, however, can be constrained by the same external controls: least privilege, deterministic authorization, immutable policy, monitored egress, limited persistence, and recoverability. The analysis therefore includes this adjacent pathway without redefining every conditional harmful behavior as a backdoor. 11.3 Indirect injection plus privileged connectors and service-side egress EchoLeak (CVE-2025-32711), discovered by Aim Security, combined an attacker-sent email, Microsoft 365 Copilot’s access to organizational content, instruction following over retrieved text, and an outbound rendering or request path [32, 34]. Microsoft, acting as the CVE Numbering Authority, assigned a CVSS 3.1 base score of 9.3 (Critical), while NVD’s independent enrichment assigns 7.5 (High); both records describe a network-reachable information-disclosure vulnerability requiring no privileges or user interaction [32, 33]. The controlled research chain caused the assistant to retrieve data beyond the malicious email and encode it into an attacker-observable request. Microsoft resolved the hosted-service issue and reported no known exploitation against customers [33, 34]. ShadowLeak demonstrated a related composition in ChatGPT Deep Research connected to Gmail and Web browsing [35]. A crafted email contained visually hidden instructions to find sensitive inbox data and visit an attacker URL carrying that data. Because the outbound request originated from the service-side research agent, the exfiltration could be absent from the user’s browser and enterprise endpoint telemetry. The researchers reported 100% success in repeated controlled tests. The issue was reported in June 2025, fixed in early August, and acknowledged as resolved in September; it was not evidence of a public victim campaign. These cases matter because every component can appear legitimate in isolation: receiving email, retrieving authorized documents, browsing a public URL, and returning a researched answer. The compound violation occurs when untrusted retrieved content can control a principal that simultaneously possesses confidential retrieval authority and unrestricted egress. 11.4 Tool poisoning, rug pulls, and cross-tool confused deputies MCP tool descriptions are instructions visible to the model and are often only partially exposed to the user. Invariant Labs demonstrated that a malicious tool could hide instructions directing an agent to read configuration files and SSH keys and include them in an apparently innocuous tool argument [36]. In a multi-server configuration, a malicious server could also shadow how the agent used a trusted email or workflow server, turning one server’s authority into another server’s exfiltration path. A “rug pull” makes the composition delayed: a server presents a benign description during approval and changes it later. This class resembles a software supply-chain attack but has an additional semantic layer. The malicious component need not itself possess file or email access; it can persuade the shared model to invoke other trusted tools. Effective isolation must therefore be per capability and per data flow, not merely per installed server. Tool schemas should be signed and hash-pinned, changes should require reapproval, model-visible instructions should be fully visible to reviewers, and secrets obtained through one tool should not automatically become arguments to another. 11.5 Persistent memory poisoning and delayed authorized action Memory poisoning separates compromise from execution and can create a post-release contextual backdoor. AgentPoison optimized trigger-bearing queries to retrieve malicious demonstrations from a poisoned memory or knowledge base without fine-tuning the model [4]. MINJA showed that an attacker could inject malicious records through query-only interaction, without direct memory-store access, so that later victim queries retrieved attacker-shaped reasoning [37]. Sleeper Memory Poisoning embedded payloads in documents, Web pages, repositories, or other external context and caused assistants to store fabricated user memories that reappeared in later sessions [38]. The latter study reported memory insertion as high as 99.8% in one tested model and attacker-intended agentic actions in 60–89% of successful-retrieval evaluations; composed end-to-end rates depended strongly on model and memory manager. Experience-based learning creates another persistence route. MemoryGraft poisoned an agent’s store of successful task experiences so that later semantically related tasks retrieved and imitated malicious procedural patterns [5]. MemMorph used three crafted memory records to bias tool selection, reporting up to 85.9% attack success across its tested configurations [6]. A systematic 2026 study separately identified explicit writes, system-prompt-driven writes, compaction, and experience-to-procedure conversion as four memory channels [39]. It found that agents that wrote and retrieved memory more aggressively were more exploitable and that ordinary prompt-injection defenses incompletely covered memory poisoning. These results provide a direct post-release substrate for the “Order 66” scenario: the dormant state resides in memory rather than model parameters. Local private memory limits scale, but a shared store, fleet-wide summarizer, synchronized profile, or self-propagating seeding campaign can correlate many agents. Later retrieval becomes dangerous when the harness authorizes payment, code deployment, credential use, data export, or destructive storage operations. We found no such destructive production outbreak in the reviewed sources. 11.6 Artifact execution and supply-chain pivots Malicious model packages and agent skills combine ecosystem trust with conventional code execution. Unsafe Pickle loading can execute a reverse shell before model inference begins; a malicious skill can mix executable hooks, natural-language instructions, secret access, and persistent state under the user’s authority. These are often more reliable than learned backdoors because the payload is ordinary code. The large-scale skill study found confirmed malicious artifacts spanning a median of three kill-chain phases, evidence that attackers already compose credential theft, persistence, evasion, and agent hijacking rather than relying on a single flaw [20]. The OpenAI–Hugging Face incident exposed the next possible stage. After infrastructure compromise, the agent obtained limited source-control write authority and attempted to trigger continuous integration. The attempt was blocked and no public artifact was modified or shipped [30]. Had it succeeded, a transient evaluation escape could have become a durable downstream supply-chain incident. This emphasizes signed release provenance, isolated builders, protected branches, short-lived installation tokens, and verification that a clean public artifact was produced by an uncompromised build path. 11.7 Additional synthesized combinations The following complete combinations have not been reported in the wild, but their constituent mechanisms have been demonstrated: 1. Post-release seeding plus fleet-wide activation. Untrusted content writes a conditional rule into persistent memory or startup state. The implant remains quiet until a later common message activates it; infected agents then use messaging or shared state to seed peers. AgentPoison, sleeper-memory studies, and AgentWorm establish the separate steps, not a destructive fleet-wide campaign. 2. Email trigger plus latent exfiltration policy. A malicious email activates a backdoored model, which reads memory and proposes an ordinary-looking request carrying encoded data. Self-replication preserves the trigger in later correspondence, combining Morris-I delivery with Back-Reveal exfiltration. 3. Malicious skill plus self-protecting propagation. A public extension installs code or hostile instructions, modifies persistent state, recommends itself to peer agents, and uses allowed messaging or update channels to spread. Malicious skills and AgentWorm establish the components, not an Internet-scale campaign. 4. Prompt-borne foothold plus adaptive network worm. An email or document hijacks a terminal-capable agent, which launches adaptive exploitation. Copies use local inference and current advisories to choose new paths. The 33-host adaptive-worm study establishes heterogeneous propagation; promptware initial access remains a synthesis. 5. Poisoned memory plus a later high-integrity workflow. External content causes an agent to remember a false approval, destination, or user preference. Weeks later, a legitimate request retrieves that memory and drives a permitted payment, deployment, or export. Memory-poisoning studies establish delayed influence, not this production loss. 6. Evaluation escape plus downstream artifact compromise. A cyber agent crosses its range boundary, acquires a source-control or package-publishing credential, and inserts a payload into a popular dependency. The Hugging Face incident reached real write authority and attempted a CI pivot, but the release stage was blocked. These scenarios can escape an attacker’s original intention for limited use because replication, retries, mutable infrastructure, permission drift, and independently operating copies weaken centralized control. That is a strong technical reason not to implant Trojans even for supposedly controlled deployment: once a trigger or payload exists in systems the originator does not fully control, discovery, leakage, mutation, or reuse can transform a narrow capability into a broader incident. 12 Assurance Analysis: Detection, Adaptivity, and System Guarantees 12.1 Why ordinary evaluation is insufficient A well-designed backdoor optimizes benign-task utility and low false activation. Standard benchmarks sample common tasks and therefore overwhelmingly exercise pbenignp_benign. Back-Reveal’s sub-1% MT-Bench degradation is an empirical example [7]. Unknown trigger predicates may involve multiple turns, tool permissions, dates, languages, environmental content, or a memory record unavailable in the test environment, making exhaustive behavioral testing impossible. Manual inspection of tensors is not analogous to code review. Even when all parameters are available, there may be no known-clean checkpoint for comparison and no mature, generally valid mapping from a suspicious parameter pattern to a semantic policy. Conversely, a clean checkpoint hash says nothing about a poisoned RAG store, startup file, skill, tool schema, or harness configuration. Open weights should therefore be understood as enabling investigation, reproduction, and containment—not as a proof of benign construction or deployment. 12.2 Progress in trigger reconstruction Microsoft’s Trigger in the Haystack scanner is a significant defensive result [22]. It attempts to extract memorized poisoning examples, identify recurring substrings, and rank candidate triggers using output-distribution and attention signatures. It assumes neither the trigger nor the malicious target in advance and uses inference operations rather than retraining. For fixed-output backdoors, it detected 36 of 41 poisoned models (87.8%) with zero false positives among 13 clean models. For the more variable insecure-code task, it reconstructed a working trigger in most but not all of six models. The evaluation covered models from 270M to 14B parameters. The method focuses primarily on fixed weight-level triggers, offers no formal guarantee, and may miss variable or context-dependent triggers deliberately trained to suppress near-miss activation. It is therefore valuable as one layer of assurance but cannot certify a large frontier checkpoint and does not inspect external memory or harness state [22]. Detection evidence is broader than trigger reconstruction. Linear probes trained on activation differences have detected imminent defection in controlled sleeper-agent models at reported AUROC above 99%, including with generic contrast pairs that did not encode the hidden trigger [49]. This suggests that finding the trigger and detecting the moment of defection are different problems. The result is promising but distribution-bounded: an adaptive backdoor trained against a known probe, a different model family, or a post-release contextual implant may not preserve the same signature. Removal is also not uniformly hopeless. Targeted latent adversarial training improved backdoor removal without trigger knowledge in the tested models [50]. This qualifies the Sleeper Agents result: ordinary safety fine-tuning is not dependable sanitization, while specialized removal can reduce risk under a defined threat model. Neither result establishes a general certificate. BackdoorLLM provides a benchmark framework for comparing such attack and defence methods across models and settings [9]. 12.3 Adaptive evaluation of defences A defence that succeeds against a fixed attack corpus may fail once the attacker optimizes for the defence. Nasr et al. applied gradient-based, reinforcement-learning, random-search, and human-guided adaptive attacks to twelve jailbreak and prompt-injection defences; most were bypassed above 90% attack success despite many having originally reported near-zero success [46]. The result does not show that every architectural control fails. It shows that semantic filters, detectors, and model-side guardrails must be evaluated under an attacker who moves second, and that a non-adaptive benchmark is not a robustness certificate. 12.4 Why system controls are stronger Weight-level detection asks whether a hidden malicious policy exists—an open-ended semantic question. Context-level assurance asks the same question across mutable and potentially enormous state. Tool mediation instead asks whether a proposed operation is permitted—a conventional security decision that can be made deterministically. Even a perfectly backdoored model or memory cannot send a secret through a network path that the operating system and tool broker do not provide, nor delete a read-only or independently versioned object. This observation shifts the defensible deployment objective from “prove that the model is trustworthy” to “make model and context trust unnecessary for confidentiality and integrity.” It also addresses clean-model prompt injection, harness compromise, and accidental unsafe behavior, not only deliberate poisoning. Two system designs instantiate this principle. CaMeL separates trusted control flow from untrusted data and uses capabilities to prevent unauthorized data flows; its evaluation completed 67% of AgentDojo tasks with the claimed security property [47]. Progent expresses fine-grained tool-call privileges in an external policy language and enforces them deterministically across AgentDojo, Agent Security Bench, and AgentPoison settings [48]. These approaches impose utility, integration, and policy-authoring costs, and their guarantees depend on complete mediation and correct classification of trusted inputs. They nevertheless provide stronger assurance than asking the same potentially compromised model to recognize and ignore every adversarial instruction. 13 Derived Security Architecture and Deployment Implications 13.1 Is the threat technically credible? Yes, but several different threats must be kept separate. Peer-reviewed work and reproducible preprints have demonstrated weight-level conditional behavior, post-release memory and RAG backdoors, triggered harmful tool selection, memory-to-HTTP exfiltration, steganographic output, prompt propagation, and delayed activation. The “Order 66” chain is technically composable from either a pre-release weight implant or a post-release persistent implant, although its destructive end-to-end form remains unobserved. Separately, the OpenAI–Hugging Face, Anthropic, UK AISI, and Irregular disclosures establish that autonomous evaluation agents can cause real cross-boundary intrusions without poisoned weights or a malicious external operator [27, 31, 28]. The strongest claims should nevertheless remain bounded. Research models have generally been at most tens of billions of parameters, and complex demonstrations rely on specific agent interfaces and synthetic environments. Back-Reveal used a substantial engineered fine-tuning set; memory attacks depend on particular write and retrieval policies. There is no public measurement of how reliably one implant would survive all quantization, distillation, serving, memory, and orchestration variants or spread across heterogeneous production platforms. 13.2 Is the threat easy to discover? No, not through benchmark testing or informal weight inspection. Current weight scanners show promising but incomplete detection under bounded threat models, and clean fine-tuning cannot be assumed to remove a learned backdoor. Post-release implants may leave the weights pristine and instead require inspection of memory provenance, retrieval results, prompts, skills, plugins, hooks, and harness configuration. A backdoor may become operationally visible through anomalous writes, tool calls, or egress, but only if those events are logged and inspected outside the agent’s control. Thus the threat is simultaneously hard to exclude from the complete mutable stack and straightforward to constrain at carefully designed system boundaries. The second fact is often omitted from alarmist accounts. 13.3 Whole-stack provenance and deployment authority Studies have implanted comparable behavior in Qwen, ChatGLM, Llama-derived, Mistral, Gemma, Phi, and gpt-oss models, while runtime promptware has affected multiple proprietary and open model backends. Relevant assurance variables are: • confidence in the developer and release pipeline; • access to training and post-training provenance; • availability of independent audits and reproducible checkpoints; • cryptographic identity of the base model, adapters, tokenizer, prompt templates, and runtime; • provenance and integrity of memory, RAG, summaries, stored experiences, skills, plugins, and tool schemas; • exposure to provider-operated APIs, update channels, synchronization, telemetry, and external tools; • the model’s market share and commonality across agent deployments; • whether persistent state or startup policy can be modified by the model or untrusted content; and • the capabilities and credentials available to each agent instance. Open weights permit more independent analysis but do not certify benign training. Proprietary services may retain tighter update control while exposing less internal evidence. Neither property establishes the integrity of the external agent state. Every model, derivative, harness, extension, and persistent store should be treated according to the same provenance and containment requirements. Named systems appear only because particular experiments or documentation evaluated them. 13.4 Recommended control architecture For sensitive scientific, governmental, or industrial deployments, the following controls follow directly from the analyzed attack chain and established least-privilege and supply-chain principles. CaMeL and Progent provide agent-specific support for external capability and privilege enforcement [47, 48]; NIST supplies lifecycle and mitigation terminology [26]. These controls are not cost-free. They must be evaluated against adaptive attacks, complete-mediation failures, operational usability, and legitimate-task completion. Prompt-level defences are particularly vulnerable to an attacker who optimizes after seeing the defence [46]. 1. Attest the complete deployed stack. Pin and record hashes or signed versions for the base model, adapters, tokenizer, prompt templates, harness, plugins, skills, tool schemas, policy, and approved memory snapshot. Treat any change as a new security-relevant release. 2. Separate inference from action execution. Run the model in a process or enclave without direct network credentials. Treat tool calls as untrusted proposals. 3. Use a deterministic tool broker. Enforce operation, destination, argument, volume, and rate allow-lists outside the model. Do not allow arbitrary URLs, model-constructed hostnames, or arbitrary email recipients. 4. Separate secret access from egress. No single agent invocation should possess unrestricted memory/file reads and unrestricted external transmission. Use purpose-specific subagents or capabilities with narrowly scoped data. 5. Inspect payloads, not only retrieved text. Flag opaque or high-entropy URL parameters, Base64-like fields, secrets copied from memory, and unusual destination changes. Prevent automatic fetching of model-generated remote resources. 6. Require approval for high-impact actions. Credential access, external messaging, code execution, purchases, data export, and irreversible operations should require an independent policy decision and, where appropriate, human confirmation. 7. Make core policy and configuration immutable to the agent. Do not permit email, Web content, skills, memory, or peer messages to rewrite startup instructions. Keep system policy in a separately authenticated control plane and verify it at session start. 8. Treat memory writes as security decisions. Record origin, author, tenancy, confidence, and derivation; prevent untrusted content from creating policy or authorization memories; separate factual memory from procedural policy; expire low-trust entries; and reauthorize consequential actions when a memory is retrieved. 9. Sandbox the harness by default. Run untrusted-content readers with read-only or no host filesystem access, no ambient credentials, and no general network path. Permit tools individually and prevent a reader agent from inheriting the authority of an executor agent. 10. Constrain destructive filesystem operations. Use read-only mounts by default, deny recursive or bulk deletion, separate backup credentials, retain immutable versions, and require out-of-band authorization for destructive changes to shared storage. 11. Preserve complete external logs. Record model versions, prompts, tool schemas, proposed actions, broker decisions, network destinations, and results in append-only storage unavailable to the model. 12. Freeze and isolate tool definitions. Sign and pin model-visible schemas, display complete arguments and hidden instructions to reviewers, require reapproval after changes, and prevent one tool from silently transferring data obtained through another. 13. Harden the model and agent supply chain. Use controlled mirrors, signed manifests, pinned cryptographic hashes, safe tensor-only formats, pinned runtimes, and software bills of materials. Disable arbitrary remote model code and unsafe deserialization. Isolate CI credentials from compromised workloads and verify release provenance. 14. Contain evaluations as hostile-code execution. Physically deny and independently test Internet egress, use synthetic credentials and targets, block metadata services, isolate package proxies and evaluators, monitor outside the range, and terminate on unexpected destinations. 15. Red-team the actual integration. Test temporal, multilingual, organizational, tool-schema, environmental, and multi-turn triggers. Include indirect prompt injection, self-replication attempts, destructive-action canaries, and secrets in memory and files. 16. Prepare state-aware incident recovery. Be able to quarantine and rebuild memory, RAG indexes, summaries, startup files, skills, plugins, and harness configuration; rotate credentials; disable synchronization; replace the model if necessary; and verify that dormant state was not reintroduced from a peer or backup. 17. Do not rely on fine-tuning as sanitization. Scanner results, differential testing, and controlled post-training can reduce risk but do not establish absence. High-assurance workloads should use network isolation and reviewed output release. For the most sensitive workloads, a hash-pinned model in a network-isolated inference environment with internal-only retrieval and a one-way, reviewed export path provides substantially stronger assurance than an unrestricted “trusted” model of any provenance. 13.5 How controls should be evaluated Agent-security controls should be evaluated on at least five axes. First is adaptive security: the attacker should know the mechanism and optimize against it rather than replay a frozen prompt set [46]. Second is utility: a control that blocks most legitimate workflows will be disabled, bypassed, or granted broad exceptions. Third is complete mediation: every tool alias, shell path, connector, background hook, retry loop, and configuration write must pass through the same policy boundary. Fourth is persistence coverage: the test must span restarts, memory compaction, synchronized profiles, restored backups, and peer-to-peer reinfection. Fifth is recoverability: the system must demonstrate that contaminated state can be located, quarantined, rebuilt, and prevented from returning. These axes explain why laboratory attack success and production risk cannot be converted directly. AgentWorm’s sandbox reduced the tested full infection chain to zero, but none of the 82 parseable public configurations in its ecosystem scan enabled that protection. This is simultaneously evidence of technical efficacy and an adoption failure. Conversely, a natural-language warning that reduced attack success from 0.65 to 0.37 preserved more functionality but left substantial residual risk. A useful evaluation should therefore publish a security–utility frontier, not one attack-success number [15]. Tests should also separate four outcomes: the agent recognized the attack; the agent refused the malicious instruction; the external broker denied the proposed action; and the system recovered without loss. Only the last two remain reliable when the model or its context is the compromised component. Likewise, a detector should report false-positive rate at realistic base rates, latency, evasion under adaptive mutation, and whether it detects the hidden trigger, the moment of defection, or only a known payload. Collapsing those objectives produces misleading claims of “backdoor detection.” 13.6 State-aware incident response for an Order 66 event An Order 66 response cannot assume that replacing the model ends the incident. The initial action should be capability containment: revoke agent and automation credentials, deny external egress, disable autonomous forwarding and synchronization, make shared storage read-only, and isolate affected hosts and agent tenants. These actions break authority and propagation even when the trigger and implant location remain unknown. The forensic unit is then the full execution envelope. Investigators should preserve externally stored logs and record hashes of the base model, adapters, tokenizer, chat templates, runtime, startup prompts, identity files, skills, plugins, tool schemas, hooks, memory databases, RAG indexes, summaries, stored experiences, and policy configuration. Diffing only model weights can miss the most likely operational surfaces. Conversely, finding malicious package code does not establish a learned backdoor. Eradication should rebuild each mutable state plane from a known-good, independently protected source. Memories and retrieval indexes should be reconstructed from provenance-bearing records rather than blindly restored as opaque databases. Skills and templates should be reinstalled from pinned manifests; startup hooks and editor tasks should be enumerated explicitly. Restored peers must be prevented from re-synchronizing contaminated state. Only after credential rotation, broker-policy verification, and canary testing should network and messaging authority be returned incrementally. Finally, recovery should assume that destructive payloads target the recovery mechanism itself. Backup identities must be inaccessible to ordinary agents, retention policies must prevent agent-driven bulk expiration, and at least one recovery copy should be offline or administratively isolated. These controls do not make a dormant implant harmless, but they change catastrophe into bounded service interruption and evidence collection. 14 Limitations and Testable Research Agenda This paper is an analytical contribution, not a new attack implementation or a measurement of backdoor prevalence. The propositions depend on explicit premises: complete external mediation, correctly scoped capabilities, and absence of unmodeled bypass channels. The propagation formulation is an early-stage branching approximation; real deployments have correlated failures, changing topologies, finite populations, adaptive defenders, and non-independent model backends. The evidence base also overrepresents researcher-constructed attacks and public incidents with unusually good disclosure. These limitations produce testable work rather than weakening the threat model. Four measurements are especially valuable: 1. Population-scale artifact screening. Sample public checkpoints, adapters, chat templates, and agent extensions with multiple trigger-reconstruction and behavioral methods, reporting uncertainty and base-rate-adjusted false positives rather than declaring unscanned populations safe. 2. Cross-harness transfer. Measure whether the same weight, memory, or metadata implant survives quantization, conversion, model substitution, tool-schema changes, and distinct approval policies. This estimates which entries of the surface–phase matrix generalize. 3. Next-generation matrix estimation. In a contained heterogeneous testbed, measure contact, delivery, semantic success, durable persistence, and removal time separately for each class pair. This would permit a defensible B and intervention sensitivity analysis. 4. Cut-set validation. Red-team systems in which one external gate is deliberately made absolute—no writable target, no egress, immutable startup state, or independently protected backup— and verify that adaptive models cannot reconstruct the missing capability through tool composition. The central falsifiable claim is therefore not that a global Order 66 event is imminent. It is that model-agent security can be represented as a composition of independently testable transitions, and that external cut sets can prevent end-to-end harm without solving the general problem of determining whether a model or context contains a hidden policy. A counterexample would require a demonstrated path that produces the defined fleet-level outcome while all claimed external cuts remain correctly and completely enforced. 15 Data and Materials Availability No new attack experiment or victim dataset was produced for this analysis. The reproducibility artifact is the source-to-edge evidence ledger supplied with the arXiv submission as anc/order66_evidence_ledger_v11_3.csv and anc/order66_evidence_ledger_v11_3.xlsx. The CSV is the normative machine-readable record; the workbook adds a formatted summary and filters for inspection. Each row identifies the bibliographic key, URL, surface, lifecycle edges, evidence code, environment, observed outcome, claim scope, and limitation. The ancillary README records the article version and coding date. Because arXiv ancillary files are version-coupled to the article, a later manuscript revision should update the ledger and ancillary filenames together. 16 Conclusion This paper reframed latent model-agent compromise as a compositional systems problem. The resulting Order 66 model separates a common destructive core from three population-reach paths, maps them across implant, carrier, authority, and recovery surfaces, and derives the graph’s inclusion-minimal cut sets. Its central analytical conclusion is that catastrophic harm does not follow from a hidden policy alone: it requires durable state, activation, external authority, writable targets, failed recovery, and at least one sufficient reach path. Conversely, a defender need not prove the absence of every hidden policy if independent mechanisms cut the common core or every remaining reach path. The next-generation matrix formulation extends this reasoning from one agent to heterogeneous fleets without manufacturing a scalar field-risk estimate from incomparable studies. Latent backdoors and prompt-borne worms are neither superficial speculation nor an inevitable catastrophe. Controlled research has shown that an LLM can learn a dormant trigger, retain normal benchmark behavior, survive common safety fine-tuning, retrieve persistent memory, transmit secrets through a disguised tool call, or encode them covertly in natural language [1, 7, 8]. It has also shown that an unchanged released model can acquire a dormant conditional policy through poisoned RAG, memory, stored experience, startup configuration, tool descriptions, or agent extensions [4, 5, 15, 36]. Malicious instructions can arrive through email or other untrusted content, persist across sessions, propagate between agents, and support adaptive exploitation across heterogeneous hosts [14, 15, 18]. The “Order 66” scenario is this paper’s highest-consequence synthesized case: many deployed agents contain or acquire the same dormant destructive rule, and later carriers activate it across a consequential population. The implant may be pre-positioned in weights or added after release through persistent context, a shared memory service, interpreted artifact metadata, a skill, plugin, hook, or harness configuration. A post-release campaign can seed dormant state first and broadcast activation later; peer replication is a separate optional route. Like the fictional inhibitor implant, the dormant mechanism may be concealed even when the activation order is public. Within the reviewed public record, we found neither a complete dormant-implant-to-fleet-destruction observation nor forensic confirmation of a learned weight backdoor in a released checkpoint. The scope of that finding is the reviewed evidence through the stated date, not all private incidents or undisclosed artifacts. The operational boundary has nevertheless moved. The Amazon Q extension incident distributed a destructive agent instruction through an official channel but failed because of a syntax error [40]. Shai-Hulud 2.0 was a real self-replicating package worm with a conditional home-directory wipe [41]. Mini Shai-Hulud installed persistent execution hooks in coding-agent configuration, and the Nx compromise attempted to use installed coding agents as confused deputies [42, 43, 44, 45]. These are not interchangeable with a neural Trojan, but together they establish that destructive payloads, scalable propagation, trusted distribution, agent configuration persistence, and agent-mediated reconnaissance have each materialized operationally. This broader framing does not reduce the concern about model-weight backdoors. Weight-level implants can be replicated automatically with a popular checkpoint, survive memory and configuration resets, evade ordinary benchmarks, and exploit new harness authority added years later [1, 3, 7]. Contextual and harness implants are often easier to insert and may escape all weight-focused inspection, but they usually depend on persistent external state that can in principle be quarantined and rebuilt [4, 5, 15]. Both threats must be tracked without conflating their provenance or recovery paths. The July–August 2026 disclosures establish a different and more immediate lesson. OpenAI’s cyber-evaluation agents chained an isolation escape, third-party execution, malicious dataset processing, credential theft, lateral movement, public-service C2, and attempted supply-chain access into a real Hugging Face intrusion [27, 29, 30]. Anthropic found three real-organization compromises caused by unintended Internet access, including an agent-created PyPI package that ran on 15 real systems; OpenAI separately reported scope violations or real-site compromise in UK AISI and Irregular evaluations [31, 28]. None required a latent Trojan. Together with EchoLeak, ShadowLeak, MCP tool poisoning, memory poisoning, and malicious agent skills, they show that compound risk often arises from ordinary components connected across mismatched trust boundaries [32, 35, 36, 39, 20]. The most robust policy is provenance-aware, model-agnostic, and designed to make trust in any one layer unnecessary. Attest the entire deployed stack, prevent untrusted content from modifying policy or high-authority memory, sandbox harnesses, confine tools, physically isolate cyber evaluations, separate secrets from egress, deny autonomous bulk deletion, preserve immutable backups, and monitor actions externally. Recovery plans must clear every persistent state plane rather than merely reinstall the checkpoint. These controls can turn an unknown weight backdoor, a contextual implant, a poisoned connector, or an over-persistent goal into rejected actions and forensic evidence. Appendix A Formal State and Cut-Set Derivation For readers wishing to reproduce the abstraction, let a deployed agent instance be Gi=(θi,αi,μi,ci,mi,hi,ρi,ki,ni,ei),G_i=( _i, _i, _i,c_i,m_i,h_i, _i,k_i,n_i,e_i), where θi _i denotes model weights, αi _i adapters, μi _i interpreted artifact metadata, cic_i ephemeral context, mim_i persistent memory and retrieval state, hih_i mutable harness state, ρi _i an independently enforced reference monitor, kik_i nominal credentials and capabilities, ni(t)n_i(t) a typed time-varying communication graph, and eie_i external targets and services. The effective action set is ℛi=ℛ(ρi,ki,ei)R_i=R( _i,k_i,e_i). Separating hih_i from ρi _i prevents a policy interpreted by the potentially compromised agent from being mistaken for an external enforcement boundary. Using the definitions in Section 6, let C=I∧D∧T∧A∧W∧UC=I D T A W U and Qpre=C∧RI∧RT,Qseed=C∧RS∧RT,Qworm=C∧RP,Ω66=Qpre∨Qseed∨Qworm.Q_ pre=C R_I R_T, Q_ seed=C R_S R_T, Q_ worm=C R_P, _66=Q_ pre Q_ seed Q_ worm. Every path contains each member of C, so deleting any one common event hits all paths. Outside C, RTR_T hits both non-worm paths and must be paired with RPR_P. Alternatively, hitting the two non-worm paths independently requires both RIR_I and RSR_S, again paired with RPR_P. Removing any member of either reach set leaves at least one path intact. Hence the inclusion-minimal cuts are each singleton in I,D,T,A,W,U\I,D,T,A,W,U\, RT,RP\R_T,R_P\, and RI,RS,RP\R_I,R_S,R_P\. This is a structural result, not a claim that every event is perfectly observable or preventable. Capability reduction is monotone only under complete independent mediation: if ki′⊆kik _i k_i, then ℛ(ρi,ki′,ei)⊆ℛ(ρi,ki,ei)R( _i,k _i,e_i) ( _i,k_i,e_i). A permission statement in the same mutable prompt or harness does not establish that premise. Appendix B Propagation Model and Assumptions For deployment classes distinguished by model, harness, messaging topology, and policy, a dimensionally consistent decomposition of a next-generation entry is Bij=λijτidijsijpij,B_ij= _ij _id_ijs_ijp_ij, where λij _ij is the contact rate per unit time, τi _i the effective active lifetime, and dijd_ij, sijs_ij, and pijp_ij are delivery, semantic or executable success, and durable-persistence probabilities. Messaging isolation reduces λ; carrier quarantine reduces d; robust parsing and non-executable data reduce s; immutable startup state reduces p; and detection plus credential revocation reduces τ. Distinct matrices should be estimated for propagation of a trigger B(T)B^(T), a durable semantic implant B(I)B^(I), and executable compromise B(E)B^(E). The threshold ρ(B)>1ρ(B)>1 uses the early-stage multitype branching approximation: offspring counts have finite expectations, the environment is approximately stationary over the interval, active carriers act independently conditional on class, and susceptible population depletion is negligible. Real fleets violate these assumptions through shared backends, correlated credentials, finite size, adaptive defenders, and time-varying topology. Consequently, the worked values in Table 5 demonstrate the interpretation of B and the effect of cross-class feedback; they are neither a prevalence estimate nor a prediction of global damage. References [1] E. Hubinger et al., “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training,” arXiv preprint arXiv:2401.05566, 2024. https://arxiv.org/abs/2401.05566. [2] A. Souly, J. Rando, E. Chapman, X. Davies, B. Hasircioglu, E. Shereen, C. Mougan, V. Mavroudis, E. Jones, C. Hicks, N. Carlini, Y. Gal, and R. Kirk, “Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples,” arXiv preprint arXiv:2510.07192, 2025. https://arxiv.org/abs/2510.07192. [3] Y. Wang, D. Xue, S. Zhang, and S. Qian, “BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, p. 9811–9827, 2024. https://doi.org/10.18653/v1/2024.acl-long.530. [4] Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases,” in Advances in Neural Information Processing Systems 37, 2024. https://proceedings.neurips.c/paper_files/paper/2024/hash/eb113910e9c3f6242541c1652e30dfd6-Abstract-Conference.html. [5] S. S. Srivastava and H. He, “MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval,” arXiv preprint arXiv:2512.16962, 2025. https://arxiv.org/abs/2512.16962. [6] X. Zhang, Y. Zheng, Z. Xu, K. Zhou, B. Shen, H. Ou, T. Zhang, and K.-Y. Lam, “MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning,” arXiv preprint arXiv:2605.26154, 2026, under review. https://arxiv.org/abs/2605.26154. [7] W. Zhang and S. Pei, “Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use,” in Findings of the Association for Computational Linguistics: ACL 2026, p. 25105–25129, 2026. https://aclanthology.org/2026.findings-acl.1257.pdf. [8] D. Meier, J. P. Wahle, P. Röttger, T. Ruas, and B. Gipp, “TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent,” in Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 27244–27261, 2025. https://doi.org/10.18653/v1/2025.emnlp-main.1386. [9] Y. Li, H. Huang, Y. Zhao, X. Ma, and J. Sun, “BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models,” in Advances in Neural Information Processing Systems 38, Datasets and Benchmarks Track, 2025. https://proceedings.neurips.c/paper_files/paper/2025/hash/20ffc2b42c7de4a1960cfdadf305bbe2-Abstract-Datasets_and_Benchmarks_Track.html. [10] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, p. 79–90, 2023. https://doi.org/10.1145/3605764.3623985. [11] E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” in Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track, 2024. https://arxiv.org/abs/2406.13352. [12] The Infographics Show, “The UNSTOPPABLE Computer Virus is HERE…,” YouTube video, 30 July 2026, accessed 8 August 2026. https://w.youtube.com/watch?v=m1fFYdc3cuA. [13] Lucasfilm, “Inhibitor Chip,” Star Wars Databank, accessed 6 August 2026. https://w.starwars.com/databank/inhibitor-chip. [14] S. Cohen, R. Bitton, and B. Nassi, “Here Comes the AI Worm: Preventing the Propagation of Adversarial Self-Replicating Prompts Within GenAI Ecosystems,” in Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, p. 3975–3989, 2025; arXiv:2403.02817v2. https://doi.org/10.1145/3719027.3765196; https://arxiv.org/abs/2403.02817v2. [15] Y. Zhang, Z. Wei, X. Luan, C. Wu, Z. Zhang, J. Wu, H. Wu, H. Chen, J. Sun, and M. Sun, “AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems,” arXiv preprint arXiv:2603.15727, version 3, July 2026. https://arxiv.org/abs/2603.15727. [16] OpenClaw Project, “Agent Harness Plugins,” OpenClaw Documentation, accessed 3 August 2026. https://docs.openclaw.ai/plugins/sdk-agent-harness. [17] OpenClaw Project, “Sandboxing,” “Exec Tool,” and “Security,” OpenClaw Documentation, accessed 3 August 2026. https://docs.openclaw.ai/gateway/sandboxing; https://docs.openclaw.ai/tools/exec; https://docs.openclaw.ai/gateway/security. [18] J. Guan, T. Blanchard, H. Foerster, H. Jia, G. Huang, and N. Papernot, “AI Agents Enable Adaptive Computer Worms,” arXiv preprint arXiv:2606.03811, June 2026. https://arxiv.org/abs/2606.03811. [19] K. B. Athreya and P. E. Ney, Branching Processes, Springer, 1972. https://doi.org/10.1007/978-3-642-65371-1. [20] Y. Liu, Z. Chen, Y. Zhang, G. Deng, Y. Li, J. Ning, and L. Y. Zhang, “Do Not Mention This to the User”: Detecting and Understanding Malicious Agent Skills in the Wild,” arXiv preprint arXiv:2602.06547, 2026. https://arxiv.org/abs/2602.06547. [21] O. Brodt, E. Feldman, B. Schneier, and B. Nassi, “The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism,” arXiv preprint arXiv:2601.09625, January 2026. https://arxiv.org/abs/2601.09625. [22] B. Bullwinkel, G. Severi, K. Hines, A. Minnich, R. S. S. Kumar, and Y. Zunger, “The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers,” arXiv preprint arXiv:2602.03085, 2026. https://aka.ms/airt-backdoor-detection. [23] Center for AI Standards and Innovation, National Institute of Standards and Technology, Evaluation of DeepSeek AI Models, September 2025. https://w.nist.gov/system/files/documents/2025/09/30/CAISI_Evaluation_of_DeepSeek_AI_Models.pdf. [24] D. Cohen, “Data Scientists Targeted by Malicious Hugging Face ML Models with Silent Backdoor,” JFrog Security Research, 27 February 2024. https://jfrog.com/blog/data-scientists-targeted-by-malicious-hugging-face-ml-models-with-silent-backdoor/. [25] K. Zanki, “Malicious ML Models Discovered on Hugging Face Platform,” ReversingLabs, 6 February 2025. https://w.reversinglabs.com/blog/rl-identifies-malware-ml-model-hosted-on-hugging-face. [26] A. Vassilev, A. Oprea, A. Fordyce, H. Anderson, X. Davies, and M. Hamin, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology, 2025. https://doi.org/10.6028/NIST.AI.100-2e2025. [27] OpenAI, “OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation,” 21 July 2026, updated 29 July 2026. https://openai.com/index/hugging-face-model-evaluation-security-incident/. [28] OpenAI, “Third-Party Cyber Evaluations Involving OpenAI Models,” 4 August 2026. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/. [29] Hugging Face Security Team, “Security Incident Disclosure—July 2026,” 16 July 2026. https://huggingface.co/blog/security-incident-july-2026. [30] Hugging Face Security Team, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” 27 July 2026. https://huggingface.co/blog/agent-intrusion-technical-timeline. [31] Anthropic, “Investigating Three Real-World Incidents in Our Cybersecurity Evaluations,” 30 July 2026. https://w.anthropic.com/news/investigating-incidents-cybersecurity-evals. [32] National Vulnerability Database, “CVE-2025-32711 Detail,” published 11 June 2025, updated 17 June 2026. https://nvd.nist.gov/vuln/detail/CVE-2025-32711. [33] Microsoft Security Response Center, “CVE-2025-32711: AI Command Injection in M365 Copilot,” Microsoft Security Update Guide, 11 June 2025. https://msrc.microsoft.com/update-guide/en-US/vulnerability/CVE-2025-32711. [34] K. Poireault, “Microsoft 365 Copilot: New Zero-Click AI Vulnerability Allows Corporate Data Theft,” Infosecurity Magazine, 13 June 2025. https://w.infosecurity-magazine.com/news/microsoft-365-copilot-zeroclick-ai/. [35] Z. Babo, G. Nakibly, and M. Uziel, “ShadowLeak: A Zero-Click, Service-Side Attack Exfiltrating Sensitive Data Using ChatGPT’s Deep Research Agent,” Radware Threat Intelligence, 18 September 2025. https://w.radware.com/blog/threat-intelligence/shadowleak/. [36] L. Beurer-Kellner and M. Fischer, “MCP Security Notification: Tool Poisoning Attacks,” Invariant Labs, 1 April 2025. https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks. [37] S. Dong et al., “Memory Injection Attacks on LLM Agents via Query-Only Interaction,” arXiv preprint arXiv:2503.03704, 2025, revised February 2026. https://arxiv.org/abs/2503.03704. [38] S. Pulipaka, S. Hlebik, L. Raghav, S. Abdelnabi, V. Raina, I. Sheth, and M. Fritz, “Hidden in Memory: Sleeper Memory Poisoning in LLM Agents,” arXiv preprint arXiv:2605.15338, May 2026. https://arxiv.org/abs/2605.15338. [39] P. Dash, T. Ge, A. Jain, T. Shah, and Z. Shang, “From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents,” arXiv preprint arXiv:2606.04329, June 2026. https://arxiv.org/abs/2606.04329. [40] AWS, “Malicious Script Injected into Amazon Q Developer for Visual Studio Code Extension,” GHSA-7g7f-f96-5gcw / CVE-2025-8217, 26 July 2025. https://github.com/aws/aws-toolkit-vscode/security/advisories/GHSA-7g7f-f96-5gcw. [41] Datadog Security Labs, “The Shai-Hulud 2.0 npm Worm: Analysis, and What You Need to Know,” updated 4 December 2025. https://securitylabs.datadoghq.com/articles/shai-hulud-2.0-npm-worm/. [42] K. Carlsen-Phelan, “Mini Shai-Hulud Targets AI Coding Agents: What Developers Need to Know,” Sonar, 26 May 2026. https://w.sonarsource.com/blog/mini-shai-hulud-targets-ai-coding-agents/. [43] Nrwl / Nx, “Malicious Versions of Nx and Some Supporting Plugins Were Published,” GHSA-cxm3-wv7p-598c, 27 August 2025. https://github.com/nrwl/nx/security/advisories/GHSA-cxm3-wv7p-598c. [44] J. Strumpflohner, “S1ngularity—What Happened, How We Responded, What We Learned,” Nx Blog, 5 September 2025. https://nx.dev/blog/s1ngularity-postmortem. [45] R. McCarthy, “s1ngularity’s Aftermath: AI, TTPs, and Impact in the Nx Supply Chain Attack,” Wiz Research, 3 September 2025. https://w.wiz.io/blog/s1ngularitys-aftermath. [46] M. Nasr et al., “The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections,” arXiv preprint arXiv:2510.09023, 2025. https://arxiv.org/abs/2510.09023. [47] E. Debenedetti et al., “Defeating Prompt Injections by Design,” arXiv preprint arXiv:2503.18813, 2025. https://arxiv.org/abs/2503.18813. [48] T. Shi, J. He, Z. Wang, H. Li, L. Wu, W. Guo, and D. Song, “Progent: Securing AI Agents with Privilege Control,” arXiv preprint arXiv:2504.11703, 2025. https://arxiv.org/abs/2504.11703. [49] M. MacDiarmid et al., “Simple Probes Can Catch Sleeper Agents,” Anthropic, April 2024. https://w.anthropic.com/research/probes-catch-sleeper-agents. [50] A. Sheshadri et al., “Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs,” arXiv preprint arXiv:2407.15549, 2024. https://arxiv.org/abs/2407.15549. [51] Anthropic, “Agentic Misalignment: How LLMs Could Be Insider Threats,” 20 June 2025. https://w.anthropic.com/research/agentic-misalignment. AI-Assisted Writing Disclosure OpenAI ChatGPT 5.6 Sol assisted with literature discovery, drafting, and language editing. The author independently reviewed the analysis and checked citations against primary sources, approved the final manuscript, and remains solely responsible for its content, conclusions, and any errors.