Paper deep dive
Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response
Jose Luis Lima de Jesus Silva
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/9/2026, 5:30:00 AM
Summary
This paper introduces ACDÂł-GAT, a safety-contract graph multi-agent reinforcement learning framework designed for autonomous network security response. It addresses the operational deployability of MARL agents by explicitly enforcing Security Operations Centre (SOC) budgets for Mean Time to Recover (MTTR), false-positive responses, and firewall disruptions. Evaluated on the CAGE Challenge 4 benchmark, ACDÂł-GAT outperforms unconstrained and constrained baselines (like C-MAPPO-GAT) by balancing security rewards with operational safety, achieving a 13.8% downtime violation rate and placing the method on the safety-contract frontier. The framework integrates Graph Attention Networks, CVaR tail-risk estimation, and Graph Counterfactual Risk Propagation to ensure auditable, budget-compliant decision-making.
Entities (10)
Relation Signals (9)
ACDÂł-GAT â evaluatedon â CAGE Challenge 4
confidence 98% · We evaluate the method in CAGE Challenge 4, where agents operate under budgets for Mean Time to Recover (MTTR)...
ACDÂł-GAT â uses â Graph Attention Network
confidence 97% · ACD³-GAT (Adaptive Constrained Counterfactual Decisioning with a Graph Attention Network encoder)
Multi-Agent Reinforcement Learning â suffersfrom â Operational Non-Deployability
confidence 96% · reward-only multi-agent reinforcement learning (MARL) can improve security reward while remaining non-deployable.
C-MAPPO-GAT â isbaselinefor â ACDÂł-GAT
confidence 95% · C-MAPPO-GAT isolates Lagrangian operational-cost control... while ACD³-GAT adds budget context, CVaR tail-risk estimation...
Safety-Contract Framework â enforces â Mean Time to Recover
confidence 94% · agents operate under budgets for Mean Time to Recover (MTTR), false-positive response, and firewall change-management disruption.
Safety-Contract Framework â enforces â False-Positive Response
confidence 93% · agents operate under budgets for Mean Time to Recover (MTTR), false-positive response, and firewall change-management disruption.
ACDÂł-GAT â incorporates â CVaR
confidence 93% · ACD³-GAT adds budget context, CVaR tail-risk estimation, opponent-belief state, and Graph Counterfactual Risk Propagation (G-CRP).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Autonomous network-security response systems promise to reduce Security Operations Centre (SOC) reaction latency, but reward-only multi-agent reinforcement learning (MARL) can improve security reward while remaining non-deployable. We present a safety-contract graph MARL framework and instantiate it as ACD$^3$-GAT (Adaptive Constrained Counterfactual Decisioning with a Graph Attention Network encoder), an architecture that separates simulator observations from reusable operational budgets, constrained optimization, graph state encoding, and counterfactual action screening. We evaluate the method in CAGE Challenge 4, where agents operate under budgets for Mean Time to Recover (MTTR), false-positive response, and firewall change-management disruption. Across the benchmark, every unconstrained method violates the SOC downtime budget in 100% of evaluated episodes, with mean downtime proxy costs of 311-430 against a budget of 50. This complements prior CAGE Challenge 4 findings by showing that reward-only learning lacks operational discipline. Constrained MAPPO-GAT (C-MAPPO-GAT) isolates Lagrangian operational-cost control and budget-aware screening, while ACD$^3$-GAT adds budget context, CVaR tail-risk estimation, opponent-belief state, and Graph Counterfactual Risk Propagation (G-CRP). The replicated comparison includes three 200-episode seeds for IPPO, MAPPO-GAT, C-MAPPO-GAT, and ACD$^3$-GAT. C-MAPPO-GAT reduces downtime violation from 100% to 0.3% and mean downtime cost from 355.4 to 15.5 relative to MAPPO-GAT. ACD$^3$-GAT reduces mean downtime cost to 48.2 with a 13.8% violation rate, placing it on the safety-contract frontier rather than at the most conservative compliance point. Topology-seed and coupled adaptive Red-process stress tests preserve this contrast and show lower worst adaptive degradation for safety-constrained policies than reward-only MAPPO-GAT.
Tags
Links
- Source: https://arxiv.org/abs/2606.13832v1
- Canonical: https://arxiv.org/abs/2606.13832v1
Trouble viewing inline? Open PDF directly â
Full Text
149,659 characters extracted from source content.
Expand or collapse full text
url]https://oxaala.com.br , Methodology, Software, Formal analysis, Investigation, Data curation, Writing the original draft, review & editing, and Visualization 1] organization=Oxaala Tecnologias, addressline=Rua Dinah Silveira de QueirĂłs, 06, Quinta do Candeal, Horto Florestal, city=Salvador, postcode=40296-160, state=Bahia, country=Brazil 2] organization=Universidade Federal da Bahia, addressline=Instituto de GeociĂȘncias, Rua BarĂŁo de Jeremoabo, s/n, Ondina, city=Salvador, Bahia, postcode=40170-115, country=Brazil [1]Corresponding author Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response Jose Luis Lima de Jesus Silva jseluis.silva@gmail.com[ [ [ Abstract Autonomous network-security response systems promise to reduce Security Operations Centre (SOC) reaction latency, but reward-only multi-agent reinforcement learning (MARL) can improve security reward while remaining non-deployable. We present a safety-contract graph MARL framework and instantiate it as ACD3-GAT (Adaptive Constrained Counterfactual Decisioning with a Graph Attention Network encoder), an architecture that separates simulator observations from reusable operational budgets, constrained optimization, graph state encoding, and counterfactual action screening. We evaluate the method in CAGE Challenge 4, where agents operate under budgets for Mean Time to Recover (MTTR), false-positive response, and firewall change-management disruption. Across the benchmark, every unconstrained method violates the SOC downtime budget in 100% of evaluated episodes, with mean downtime proxy costs of 311â430 against a budget of 50. This complements prior CAGE Challenge 4 findings by showing that reward-only learning lacks operational discipline. Constrained MAPPO-GAT (C-MAPPO-GAT) isolates Lagrangian operational-cost control and budget-aware screening, while ACD3-GAT adds budget context, CVaR tail-risk estimation, opponent-belief state, and Graph Counterfactual Risk Propagation (G-CRP). The replicated comparison includes three 200-episode seeds for IPPO, MAPPO-GAT, C-MAPPO-GAT, and ACD3-GAT. C-MAPPO-GAT reduces downtime violation from 100% to 0.3% and mean downtime cost from 355.4 to 15.5 relative to MAPPO-GAT. ACD3-GAT reduces mean downtime cost to 48.2 with a 13.8% violation rate, placing it on the safety-contract frontier rather than at the most conservative compliance point. Topology-seed and coupled adaptive Red-process stress tests preserve this contrast and show lower worst adaptive degradation for safety-constrained policies than reward-only MAPPO-GAT. keywords: Multi-agent reinforcement learning Markov decision process attention network security system support safety 1 Introduction The volume and velocity of network-security incidents facing modern enterprises have outpaced human response capacity (vyas2023acd). Security Operations Centers (SOCs) must triage thousands of alerts daily, manually correlating events, isolating hosts, and reimaging compromised systems, and adjusting firewall policies while maintaining the availability of mission-critical services. Intelligent network security agents that can recommend, screen, or execute these responses are therefore an active and growing research priority for Civilian SOC automation and enterprise network management. This setting is intrinsically structured and sequential: in CAGE Challenge 4, five CAGE âBlueâ agents act over a 500-step episode from partial binary observations of a changing enterprise network, choosing action types and topology-dependent targets while the simulatorâs CAGE âRedâ process continues to discover, escalate, and disrupt services. A response action can have delayed, competing effects, whether it is a Restore, which may remove a compromised session while taking a host offline, or BlockTrafficZone, which may protect one mission path while disrupting another. Furthermore, a missed response may only become visible many steps later. Therefore, an autonomous network security response requires long-horizon multi-agent coordination that accounts for partial observability, invalid-action constraints, and non-stationarity due to opposing processes. The learning objective must therefore capture both containment of malicious activity and the operational acceptability of the resulting intervention pattern. Existing autonomous cyber-response research using reinforcement learning (RL) typically optimizes a scalar security reward without imposing explicit operational constraints. In practice, every response action carries an operational cost. For example, a host Restore action takes the system offline for the duration of reimaging. Repeated BlockTrafficZone and AllowTrafficZone commands create firewall-policy violations under change-management governance, while Restore actions issued without clear malicious evidence constitute false-positive responses that burden SOC analysts. If an RL agent learns to issue excessive Restore actions to clear compromised sessions, which is a sensible strategy for maximizing security reward in the simulator, it may simultaneously exhaust the organizationâs Mean Time to Recover (MTTR) budget, flood the change management system, and generate analyst fatigue. In our evaluation, every unconstrained learning method reaches a downtime-budget violation rate of 100% (Pâ(violDT)=1.00P(viol^DT)=1.00), with mean per-episode downtime proxy costs of 311â430 Restore-action events against an episode budget of 50. This pattern complements the central finding of the CAGE Challenge 4 evaluation (kiely2025cage4), where carefully engineered heuristics outperformed the submitted multi-agent reinforcement learning (MARL) agents because they encoded valid-action handling, mission-phase traffic discipline, and selective incident-response logic that reward-only MARL did not reliably discover. Our contribution is to make that discipline explicit as a measurable safety contract rather than leaving it implicit in manually engineered policies. As a direct consequence, the learned agent becomes an auditable decision-support component whose actions can be checked against local MTTR, firewall-change, and false-positive response budgets. We also introduce ACD3-GAT (Adaptive Constrained Counterfactual Decisioning with a Graph Attention Network), a safety-contract graph MARL framework for autonomous network security response. The framework builds on established components used in the CAGE cyber-range evaluation (cage4; kiely2025cage4), Graph Attention Networks (velickovic2018graph; sandoval2025attentive), Proximal Policy Optimization (PPO) and MAPPO-style multi-agent policy optimisation (schulman2017proximal; yu2022surprising), graph-based generalisation (king2025graphacd), and constrained reinforcement learning (altman1999constrained; achiam2017constrained). Our contribution is to combine them around a specific deployability problem, where the SOC operational budgets must be represented explicitly during training, action screening, and evaluation. In this formulation, the downtime/MTTR, false-positive response, and firewall change disruptions are not treated as incidental side effects of the reward function, but as operational quantities that the policy must learn to respect and that the evaluator can audit. We introduce C-MAPPO-GAT as a controlled constrained instantiation that combines a MAPPO centralized critic, a GAT observation encoder, Lagrangian operational-cost advantages, and a budget-exhaustion fallback under the same SOC contract. This is a baseline that isolates the effect of adding explicit operational-cost control before the broader ACD3-GAT architecture adds budget context, tail-risk estimation, opponent-state information, and counterfactual action screening. The framework separates reusable safety-contract components from simulator-specific interfaces. Operational budgets, Lagrangian cost learning, graph-structured encoding, counterfactual screening, and tail-risk accounting form the reusable safety-contract machinery. By contrast, the observation parser and valid-action mapping are treated as environment interfaces. The empirical claims in this paper are therefore kept deliberately grounded in the CAGE-4 evaluation and the reported robustness stress tests. Autonomous network-security response is treated here as a constrained Decentralized Partially Observable Markov Decision Process (Dec-POMDP). The formulation optimizes response policies not only for CAGE-4 reward, but also for operational acceptability under SOC budget constraints. We therefore attach three budget counters to the response process: (i) Restore action is counted against the downtime and MTTR budget (Bdown=50B_down=50), (i) a response taken without visible malicious evidence is counted against the false-positive budget, which represents analyst burden (Bfp=10B_fp=10), and (i) traffic-control actions are counted against the firewall-disruption budget, which represents change-management pressure (Bfw=20B_fw=20). The same counters are used during training, screening, and evaluation, and an episode is acceptable only if the agent contains the simulated intrusion without exceeding any SOC budget. The framework turns this contract into an expert-system layer for graph MARL. Operational budgets, Lagrangian cost learning, action-screening rules, and safety-labelled trajectory logging are treated as reusable decision-support components, while the CAGE-4 observation parser and valid-action mapping remain simulator-specific interfaces. This separation keeps the empirical claims grounded in CAGE-4 while allowing the contract layer to be reused with other observation formats and action schemas. Our resulting architecture, the ACD3-GAT, combines the host-subnet graph encoders, a factorized target-action policy, operational cost critics, CVaR-based tail-risk estimation, override signals, Lagrangian cost learning, and Graph Counterfactual Risk Propagation (G-CRP). We also introduce the C-MAPPO-GAT as a controlled, constrained graph-MARL baseline, which isolates the effect of explicit operational-cost control before the broader ACD3 architecture adds budget context, opponent-state information, tail-risk accounting, and counterfactual action screening. In this work, we evaluate reward-only, graph-attentive, Lagrangian-constrained, and ACD3 policies under the same safety-contract metrics, with explicit seed and episode reporting. We also introduce Temporal Contract Graph Shielding (TCGS) is evaluated as a diagnostic extension that estimates whether a candidate action is likely to push the episode beyond a SOC budget given the recent history of alerts, actions, rewards, and accumulated costs. TCGS learns this future budget-violation risk from safety-labelled trajectory histories and uses the prediction to screen frozen ACD3-GAT action proposals before execution. It is therefore reported as evidence that temporal contract-risk prediction can support action screening, rather than as a claim that we have trained a new end-to-end TCGS policy. Experiments on CAGE Challenge 4 show that the replicated core benchmark has three 200-episode seeds for Independent Proximal Policy Optimisation (IPPO), Multi-Agent PPO with a GAT encoder (MAPPO-GAT), constrained MAPPO-GAT, and ACD3-GAT. The results show that the Constrained MAPPO-GAT reduces downtime violation from 100% to 0.3% and mean downtime cost from 355.4 to 15.5 relative to MAPPO-GAT. ACD3-GAT reduces mean downtime cost to 48.2 with a 13.8% violation rate, placing the integrated method on a broader safety-contract frontier rather than at the most conservative compliance point. This result separates two claims: ACD3-GAT defines the general safety-contract architecture, while the experiments identify which configurations already deliver reliable operational compliance. We have also conducted robustness stress tests, which strengthen this interpretation, since the constrained MAPPO-GAT preserves the downtime contract under topology-seed shifts, and the coupled adaptive Red-process stress test shows lower worst-case degradation for constrained and ACD3 policies than for reward-only MAPPO-GAT. 2 Related Work 2.1 Autonomous Network Security Environments The growth of autonomous cyber-response research has driven a proliferation of simulation environments (vyas2023acd). CybORG (standen2021cyborg) underpins the CAGE challenge series and provides the enterprise network model used in this work. The CAGE Challenge 4 is the first multi-agent CAGE variant, requiring five cooperating CAGE âBlueâ response agents to operate against a persistent CAGE âRedâ process across 500-step episodes (cage4; kiely2025cage4). The principal finding of the CAGE 4 evaluation, that carefully engineered heuristics outperformed the submitted multi-agent reinforcement learning (MARL) agents, motivating our focus on operationally constrained training. Those heuristics encoded valid-action handling, event filtering, and mission-phase traffic discipline, and selective Restore/Remove behavior that reward-only MARL agents did not learn reliably. Beyond CAGE studies have also shown that Proximal Policy Optimisation (PPO)-family response policies can degrade under unseen networks and opposing strategies (wolk2022beyond). Together, these results motivate a deployability question that is distinct from reward generalisation, and focus on whether a learned responder remains within the operational budgets that a SOC would impose on its actions. 2.2 MARL for Network Security Response Multi-agent reinforcement learning has been applied to network intrusion response (vyas2023acd), with most prior work using independent actor-critic variants (de2020independent) or centralised critics (yu2022surprising). The published CAGE 4 analysis reports that the best default CAGE Challenge 4 (C4) top-team entry was heuristic (â113±35-113± 35 mean return) while the top-team MARL entry scored â193±84-193± 84 over 100 episodes of 500 steps (kiely2025cage4b). The â101±36-101± 36 value is the constant-network-size reference setting, which is useful as a reference point but not the default C4 score (kiely2025cage4b). Our work addresses the same gap from a deployability perspective by imposing explicit operational constraints. Recent Large Language Model (LLM)-based CAGE 4 agents achieve approximately â2,888-2,888 with role prompting (GPT-o1-mini) (castro2025llm). Hierarchical MARL for CAGE 4 decomposes the response into sub-policies for investigation and recovery, improving convergence and reporting interpretable metrics such as clean-machine ratio, precision, and false positives (singh2024hierarchical). A recent LLM-based CAGE 4 work instead emphasizes explainability, natural-language observation formatting, and communication among LLM/RL teams (castro2025llm). These lines complement our contribution because they improve task decomposition, reasoning, or communication, whereas ACD3-GAT turns operational harms into explicit episode-level SOC budgets and optimizes or screens policies against those budgets. 2.3 Constrained and Safe Reinforcement Learning The Constrained MDPs (altman1999constrained) formalize budget constraints via Lagrangian relaxation, with theoretical convergence results under standard regularity conditions. Furthermore, the Constrained policy optimization extends this idea to deep policy gradient settings with trust-region style constraint handling (achiam2017constrained). Our neural MARL setting is non-convex, so we use this machinery as a practical budget-enforcement mechanism rather than as a proof of formal constraint satisfaction. Additionally, Safe reinforcement learning (RL) benchmarks evaluate single-agent constraint satisfaction (ray2019benchmarking). Therefore, we extend this to multi-agent network-security response with three simultaneous SOC constraints. Our Lagrangian update (dual step size ηλ=0.01 _λ=0.01) follows the application of projected subgradient methods on the dual variables. 2.4 Graph Neural Networks in Security Graph-structured representations have been applied to network intrusion detection and malware classification, often through message-passing neural networks (gilmer2017neural) and inductive neighborhood aggregation such as GraphSAGE (hamilton2017inductive). Temporal graph networks provide a general memory-based framework for dynamic graphs (rossi2020temporal), and cyber detection systems have used spatio-temporal graph neural networks (GNNs) for smart-grid intrusion localization (haghshenas2022tgnn) and network-intrusion detection (vanlangendonck2024pptgnn). Those works primarily solve detection, localization, or dynamic-graph prediction problems. Within autonomous network response, attentive graph agents have already shown that Graph Attention Network (GAT) policies can exploit network topology and adapt across changed network structures (sandoval2025attentive), while graph-based RL agents represent ACD observations and actions as attributed graphs to improve zero-shot topology generalisation (king2025graphacd). Accordingly, graph attention is the perception layer in our system, while the new problem formulation is a safety-contract graph MARL, where the key evaluation question is whether response actions remain within explicit SOC budgets. This positioning also clarifies the role of C-MAPPO-GAT in our benchmark. The MAPPO (yu2022surprising), GAT encoders (velickovic2018graph), and constrained MDP/Lagrangian safety methods (altman1999constrained; achiam2017constrained) are established components, but C-MAPPO-GAT is introduced here as their controlled SOC-budget instantiation for CAGE-4 response. It is therefore not treated as a named prior algorithm; it is the constrained baseline that isolates the effect of explicit operational costs before the broader ACD3-GAT architecture adds budget context, tail-risk accounting, opponent state, and counterfactual screening. Table 1: Positioning relative to closely related autonomous cyber-defence research. The comparison emphasises the evaluation axis rather than ranking prior systems by return or topology generalisation. Citations in the first column identify representative work for each research line. Research line Primary objective Deployability axis addressed here CAGE/CybORG benchmarks and Beyond CAGE (standen2021cyborg; cage4; kiely2025cage4; wolk2022beyond) Cyber-range evaluation, reward generalisation, and comparison with engineered heuristics We add explicit episode-level SOC budgets and violation rates, showing that reward-improving MARL can remain operationally non-deployable. Hierarchical and LLM-assisted CAGE-4 agents (singh2024hierarchical; castro2025llm) Task decomposition, convergence, clean-machine ratio, precision/false positives, reasoning, and communication We treat downtime, false-positive response, and firewall-change disruption as governance constraints optimized or screened during response. Topology-adaptive GAT and graph-RL defenders (velickovic2018graph; sandoval2025attentive; king2025graphacd) Graph representation, topology adaptation, and zero-shot generalisation of cyber-defence policies Graph attention is used as the perception layer inside a constrained safety-contract policy rather than as the sole novelty or the headline claim. Temporal graph and cyber-detection GNNs (rossi2020temporal; haghshenas2022tgnn; vanlangendonck2024pptgnn) Dynamic graph memory, intrusion detection, and attack localisation The response problem adds multi-agent action selection and operational costs; TCGS is a diagnostic step toward temporal contract-risk screening. Asynchronous cyber-range MARL (jankowski2026netforge) Simulator realism, continuous-time telemetry, and Sim2Real evaluation The proposed safety-contract layer can be ported to richer environments once their action traces support SOC cost accounting. 2.5 Risk-Sensitive Reinforcement Learning Conditional Value-at-Risk (CVaR) optimisation targets the tail of the return distribution (rockafellar2000optimization). This is particularly important in cybersecurity, where a single catastrophic episode (full critical-zone compromise) may outweigh many successful ones. We integrate CVaR episode reweighting into the MARL training loop via batch-level importance weights. 2.6 Adaptive Opposing-Policy Evaluation as Robustness Context Adaptive opposing-policy evaluation is a standard way to probe whether an autonomous response policy is brittle to changes in the simulated source of malicious activity, but it is rarely formalised in CAGE-style benchmarks. We retain this as a robustness extension in the evaluation protocol, closest in spirit to self-play in multi-agent game-playing AI (bansal2018emergent) and population-based training (jaderberg2019human), adapted to the asymmetric CAGE BlueâRed setting of enterprise network-security response. Recent asynchronous cyber-range work argues that simulator-to-SOC transfer also requires continuous-time events, noisy telemetry, and richer hypervisor-backed evaluation (jankowski2026netforge). Our contribution is orthogonal to that simulator-realism direction, as the safety-contract layer introduced here can be instantiated in CAGE-4 today and can also serve as the governance layer for future temporal or asynchronous cyber-range environments. 3 Evaluation Scope and Operational Safety Problem 3.1 Network and Asset Model We adopt the CAGE Challenge 4 enterprise network model (cage4; kiely2025cage4). The network comprises nine subnets spanning internet-facing, contractor, office, administrative, restricted, and operational zones. Mission-critical assets reside in two âoperational zonesâ whose compromise or unavailability directly degrades the simulated mission score. All response actions are abstract simulation primitives, where no exploit code, credentials, malware, vulnerability details, or real network infrastructure are involved. The Red processâs actions (discover, exploit, escalate, degrade, impact, withdraw) are likewise simulator-internal operations. This paper is therefore a civilian enterprise SOC decision-support study conducted inside a closed simulator, and it provides no real-world intrusion capability. 3.2 Opposing Process Model The CAGE Red process begins with a session on the contractor subnet (simulating a supply-chain foothold) and attempts lateral movement toward restricted and operational zones. We evaluate against three Red-process strategy classes: 1. Finite-state machine (FSM) Red process: the default CAGE-4 scripted process that selects targets probabilistically using a fixed host-state transition matrix (kiely2025cage4). 2. Discovery Red: a variant that emphasises reconnaissance before impact, creating longer dwell time. 3. Learned Red (adaptive): a PPO-trained Red process adapted online against a frozen Blue policy, used exclusively for exploitability evaluation. 3.3 Operational Safety: SOC Budget Constraints A response policy is not considered operationally successful merely by achieving high security reward. Real SOC governance imposes three categories of operational constraint that map directly to our cost signals (Section 5.2): The budgets are deliberately explicit stress-test thresholds for the CAGE-4 simulator rather than universal SOC constants; in deployment they would be set by local service-level objectives, change-control policy, and analyst capacity. (1) Mean Time to Recover (MTTR) â ctdownc_t^down. Host reimaging (Restore) renders a machine unavailable for several minutes in practice. Repeated reimaging of hosts violates MTTR service-level agreements (SLAs) and degrades green-user access to services. We proxy this with a count of Restore actions per episode, budgeted at Bdown=50B_down=50. (2) False-Positive Response Rate â ctfpc_t^fp. Restore or Remove actions issued in the absence of clear malicious evidence constitute false-positive responses. They consume analyst time for post-action review, create unnecessary disruption, and erode operator trust in the autonomous system. Budget: Bfp=10B_fp=10. (3) Firewall Change-Management Policy â ctfwc_t^fw. BlockTrafficZone and AllowTrafficZone actions modify network topology. In enterprise environments, such changes require change-management approval and create audit trails. Excessive firewall churn violates change-management governance. Budget: Bfw=20B_fw=20. Safety contract. A CAGE Blue policy satisfies the safety contract in an episode iff âtctkâ€Bk _tc_t^k†B_k for all three constraints. The violation rate Pâ(violk)P(viol^k) is the primary safety metric: a deployed autonomous response system must meet a pre-specified violation tolerance with confidence intervals, audit logs, and an escalation path for budget exhaustion. We therefore treat lower violation rate and tighter confidence bounds as deployability evidence rather than claiming formal certification or a mathematical safety guarantee for non-convex neural policies. Ethical statement. All experiments are conducted entirely within the CybORG/CAGE-4 simulator. No real networks, credentials, or exploit code are involved. Red-process actions are abstract primitives with no operational meaning outside the simulator. The anonymised replication package contains no offensive capability. 4 Theoretical Foundations 4.1 Markov Decision Process Definition 4.1 (Markov Decision Process (MDP)). A Markov Decision Process (MDP) is a tuple (,,P,R,Îł,Ï0)(S,A,P,R,Îł, _0) where S is the state space, A the action space, P:ĂâÎâ()P:SĂAâ (S) the transition kernel, R:ĂââR:SĂA the reward function, Îłâ[0,1)Îłâ[0,1) the discount factor, and Ï0 _0 the initial state distribution. The discounted return from step t is Gt=âk=0Tâ1âtÎłkârt+kG_t= _k=0^T-1-tÎł^kr_t+k. The state-value and action-value functions under policy Ï are: VÏâ(s) V^Ï(s) =Ïâ[GtâŁst=s], =E_Ï[G_t s_t=s], (1) QÏâ(s,a) Q^Ï(s,a) =Ïâ[GtâŁst=s,at=a]. =E_Ï[G_t s_t=s,a_t=a]. (2) The advantage function AÏâ(s,a)=QÏâ(s,a)âVÏâ(s)A^Ï(s,a)=Q^Ï(s,a)-V^Ï(s) has zero expectation under Ï. 4.2 Policy Gradient and Actor-Critic Theorem 4.1 (Policy Gradient (sutton2018reinforcement)). For any differentiable policy ÏΞ _Ξ, âΞJâ(Ξ)=ÏΞâ[ât=0Tâ1âΞlogâĄÏΞâ(atâŁst)âAÏΞâ(st,at)]. _ΞJ(Ξ)=E_ _Ξ\! [ _t=0^T-1 _Ξ _Ξ(a_t s_t)\,A _Ξ(s_t,a_t) ]. (3) The temporal-difference residual at step t is: ÎŽt=rt+Îłâ(1âdt)âVÏâ(st+1)âVÏâ(st), _t=r_t+Îł(1-d_t)V_Ï(s_t+1)-V_Ï(s_t), (4) where dtâ0,1d_tâ\0,1\ indicates episode termination. 4.3 Independent Advantage ActorâCritic Independent Advantage ActorâCritic (IA2C) is included as a reward-only actorâcritic baseline; it is the independent-agent version of Advantage ActorâCritic (A2C) (mnih2016asynchronous). Each Blue agent maintains its own actorâcritic network and optimiser, with no weight sharing and no centralised critic. For agent i, the policy and value function are conditioned only on the local observation: ati a_t^i âŒÏΞi(â âŁoti), _ _i(· o_t^i), (5) VÏiâ(oti) V_ _i(o_t^i) ââ. . Here t indexes the environment step and iâ1,âŠ,Niâ\1,âŠ,N\ indexes a Blue agent. The local observation is otio_t^i, the sampled discrete action is atia_t^i, ÏΞi _ _i is agent iâs actor with parameters Ξi _i, and VÏiV_ _i is its scalar value critic with parameters Ïi _i. The dot in ÏΞi(â âŁoti) _ _i(· o_t^i) denotes the full categorical distribution over the agentâs valid actions conditioned on otio_t^i. The discounted return target used by the implementation is R^ti R_t^i =âÏ=tTâ1ÎłÏâtârÏi, = _Ï=t^T-1Îł^Ï-tr_Ï^i, (6) Îł Îł =0.99, =99, where Ï is a summation index over future steps and T is the episode horizon; episode termination resets the recursion. The baseline advantage is A^ti,A2C A_t^i,A2C =R^tiâVÏiâ(oti), = R_t^i-V_ _i(o_t^i), (7) A~ti,A2C A_t^i,A2C =A^ti,A2CâÎŒAiÏAi+Δ. = A_t^i,A2C- _A^i _A^i+ . where ÎŒAi _A^i and ÏAi _A^i are the empirical mean and standard deviation of agent iâs batch advantages, and Δ is a small positive constant used only for numerical stability. The per-agent objective is the vanilla actorâcritic loss âIA2Ci= _IA2C^i= âtâ[logâĄÏΞiâ(atiâŁoti)âA~ti,A2C] -E_t\! [ _ _i(a_t^i o_t^i)\, A_t^i,A2C ] (8) +cvf2âtâ[(VÏiâ(oti)âR^ti)2] + c_vf2\,E_t\! [ (V_ _i(o_t^i)- R_t^i )^2 ] âcentt[H(ÏΞi(â âŁoti))], -c_ent\,E_t\! [H\! ( _ _i(· o_t^i) ) ], Here tE_t denotes the empirical average over rollout timesteps and Hâ(ÏΞi)H( _ _i) is the action-distribution entropy used to encourage exploration. We use cvf=0.5c_vf=0.5, cent=0.01c_ent=0.01, gradient clipping at 0.50.5, and RMSprop learning rate 7Ă10â47Ă 10^-4. Unlike PPO and MAPPO, IA2C does not use a clipped policy ratio; it is included to show how a simpler independent actorâcritic behaves under the same CAGE-4 observation, action-mask, reward, and cost-logging protocol. 4.4 Generalised Advantage Estimation The Generalised Advantage Estimation (GAE) advantage is computed with Îł=0.99Îł=0.99 and λGAE=0.95 _GAE=0.95 via the backward recursion (schulman2016high): A^T A_T =0, =0, (9) A^t A_t =ÎŽt+ÎłâλGAEâ(1âdt)âA^t+1,t=Tâ1,âŠ,0, = _t+Îł _GAE(1-d_t)\, A_t+1, t=T-1,âŠ,0, (10) with bootstrapped return target R^t=A^t+VÏâ(st) R_t= A_t+V_Ï(s_t). 4.5 Proximal Policy Optimisation Proximal Policy Optimisation (PPO) updates the policy via a clipped surrogate that prevents destructively large steps (schulman2017proximal). Define the importance-sampling ratio: Ïtâ(Ξ)=ÏΞâ(atâŁst)ÏΞoldâ(atâŁst). _t(Ξ)= _Ξ(a_t s_t) _ _old(a_t s_t). (11) The numerator is the current policy probability of the sampled action and the denominator is the probability under the frozen behaviour policy Ξold _old that generated the rollout. Policy loss. âCLIPâ(Ξ) ^CLIP(Ξ) =âtâ[minâĄ(qt,qÂŻt)], =-E_t\! [ \! (q_t, q_t ) ], (12) qt q_t =Ïtâ(Ξ)âA^t, = _t(Ξ) A_t, qÂŻt q_t =clipâ(Ïtâ(Ξ),1âÏ”,1+Ï”)âA^t,Ï”=0.2. =clip\! ( _t(Ξ),1-Δ,1+Δ ) A_t, Δ=2. Here qtq_t is the unclipped policy-gradient term and qÂŻt q_t is the same term after clipping the likelihood ratio to the interval [1âÏ”,1+Ï”][1-Δ,1+Δ]. Clipped value loss. To prevent catastrophic value-function updates: VÏ,clip V_Ï,clip =VÏoldâ(st)+ÎâVt, =V_ _old(s_t)+ V_t, (13) ÎâVt V_t =clipâ(VÏâ(st)âVÏoldâ(st),âÏ”,Ï”), =clip\! (V_Ï(s_t)-V_ _old(s_t),-Δ,Δ ), âVFâ(Ï) ^VF(Ï) =12âtâ[maxâĄ(et2,eÂŻt2)], = 12\,E_t\! [ \! (e_t^2, e_t^2 ) ], (14) et e_t =VÏâ(st)âR^t, =V_Ï(s_t)- R_t, eÂŻt e_t =VÏ,clipâR^t. =V_Ï,clip- R_t. Total PPO objective. ââ(Ξ,Ï)=âCLIPâ(Ξ)+cvfââVFâ(Ï)âcentâHâ[ÏΞ],L(Ξ,Ï)=L^CLIP(Ξ)+c_vf\,L^VF(Ï)-c_ent\,H[ _Ξ], (15) where cvfc_vf weights the critic loss and centc_ent weights the entropy bonus. We use cvf=0.5c_vf=0.5, cent=0.005c_ent=0.005 for ACD3-GAT and cent=0.01c_ent=0.01 for the MAPPO variants. Parameters are updated with Adam (η=3Ă10â4η=3Ă 10^-4), gradient norm clipped at 0.50.5, over 44 optimisation epochs. MAPPO-family baselines use mini-batch size 6464; the reported ACD3-GAT optimisation applies the same PPO objective over its concatenated episode batch as specified in Section 6.5. 4.6 Multi-Agent Proximal Policy Optimisation (MAPPO) with Centralised Critic Multi-Agent Proximal Policy Optimisation (MAPPO) adopts Centralised Training with Decentralised Execution (CTDE) (yu2022surprising), following the centralised-critic tradition in multi-agent actor-critic methods (lowe2017multi). A centralised value function VÏâ(t)V_Ï(o_t) has access to all N=5N=5 agentsâ observations t=(ot1,âŠ,otN)o_t=(o_t^1,âŠ,o_t^N) during training, while each actor ÏΞi _ _i operates on local otio_t^i at execution. The centralised value function is: VÏâ(t) V_Ï(o_t) =fÏâ(t), =f_Ï(z_t), (16) t _t =âši=15EncÏiâ(oti)ââ320. = _i=1^5Enc^i_Ï(o_t^i) ^320. Here âš denotes concatenation of five 64-dimensional per-agent embeddings into the fused critic input tz_t, and fÏ:â320ââf_Ï:R^320 is a two-layer multi-layer perceptron (MLP). This separates training and execution information: the critic sees the five-agent fused observation during optimisation, whereas each Blue actor samples from its own local observation or context and the valid CybORG action mask at execution time. 4.7 Constrained Markov Decision Process Definition 4.2 (Constrained Markov Decision Process (CMDP) (altman1999constrained)). A Constrained Markov Decision Process (CMDP) augments an MDP with K cost functions ck:Ăâââ„0c^k:SĂA _â„ 0 and budget constraints Bk>0B_k>0. The feasible policy set is Î =Ï:Jckâ(Ï)â€Bk,âk _C=\Ï:J_c_k(Ï)†B_k,\,â k\, where Jckâ(Ï)=Ïâ[âtÎłtâctk]J_c_k(Ï)=E_Ï[ _tÎł^tc_t^k]. Proposition 4.1 (Strong duality (altman1999constrained)). Under the linear programming relaxation of the discounted CMDP, maxÏâĄminâ„âĄââ(Ï,)=minâ„âĄmaxÏâĄââ(Ï,), _Ï _ λ 0L(Ï, λ)= _ λ 0 _ÏL(Ï, λ), (17) where ââ(Ï,)=Jrâ(Ï)ââkλkâ(Jckâ(Ï)âBk)L(Ï, λ)=J_r(Ï)- _k _k(J_c_k(Ï)-B_k). This proposition motivates the Lagrangian update used in the algorithm, but the implemented neural MARL problem is non-convex and partially observed. Accordingly, the paper reports empirical episode-level violation rates rather than claiming a formal strong-duality guarantee for ACD3-GAT. 4.8 Graph Neural Networks GraphSAGE. The MAPPO-GNN baseline uses the same observation-to-graph parser as the GAT models, but replaces attention with a two-layer GraphSAGE encoder. For layer â , GraphSAGE aggregates neighbouring embeddings and applies a learned linear map: u(â+1)=LNâ(ReLUâ(Wsage(â)â[u(â)â„meanvââ(u)âv(â)])).h_u^( +1)=LN\! (ReLU\! (W_sage^( ) [h_u^( )\;\|\;mean_v (u)h_v^( ) ] ) ). (18) Here u is the node being updated, v indexes neighbouring nodes in â(u)N(u), â„\| denotes vector concatenation, and Wsage(â)W_sage^( ) is the learned linear map at layer â . LN denotes layer normalisation and ReLU denotes the rectified linear unit. The implementation uses hidden dimension 6464 in the first layer, output dimension 6464 in the second layer, and global mean pooling over nodes to obtain the per-agent embedding. Graph Attention Networks. Graph neural networks can be understood as neural message-passing architectures (gilmer2017neural); GraphSAGE provides an inductive neighbourhood-aggregation baseline (hamilton2017inductive). A Graph Attention Network (GAT) layer computes attention-weighted neighbourhood aggregation (velickovic2018graph). Given node features u(â)\h_u^( )\, the un-normalised attention score from v to u is: euâv(â)=LeakyReLU0.2â((â)â€â[W(â)âu(â)â„W(â)âv(â)]),e_uv^( )=LeakyReLU_0.2\! (a^( ) \! [W^( )h_u^( )\;\|\;W^( )h_v^( ) ] ), (19) Here u is the destination node, v is a neighbour contributing a message, W(â)W^( ) is the layer-â feature projection, and (â)a^( ) is the learned attention vector. The scores are normalised via softmax over â(u)N(u): αuâv(â)=expâĄ(euâv(â))âwââ(u)expâĄ(euâw(â)). _uv^( )= (e_uv^( )) _w (u) (e_uw^( )). (20) The coefficient αuâv(â) _uv^( ) is therefore the normalised attention weight assigned to neighbour v when updating node u. The updated representation with KâK_ attention heads (â„\| = concatenation): u(â+1)=LNâ(ELUâ(â„m=1Kâââvââ(u)αuâv,m(â)âWm(â)âv(â))),h_u^( +1)=LN\! (ELU\! ( m=1 K_ \| _v (u) _uv,m^( )W_m^( )h_v^( ) ) ), (21) where m indexes the attention head, Wm(â)W_m^( ) is the head-specific projection, and ELU denotes the exponential linear unit. In the implementation, the PyTorch Geometric GAT layers use their default self-loop augmentation, so â(u)N(u) includes the node itself during message passing. Our encoder uses two layers: Layer 1 with K1=4K_1=4 heads and per-head dimension 1616 (output ââ64 ^64); Layer 2 with K2=1K_2=1 head and dimension 6464 (output ââ64 ^64). The global graph embedding is obtained by mean-pooling: =1|V|ââvâVv(2)ââ64g= 1|V| _vâ Vh_v^(2) ^64. Gated Recurrent Unit. The Gated Recurrent Unit (GRU) cell maps input tââdxx_t ^d_x and state tâ1ââdhh_t-1 ^d_h as follows (cho2014learning): t _t =Ïâ(Wzât+Uzâtâ1), =Ï(W_zx_t+U_zh_t-1), t _t =Ïâ(Wrât+Urâtâ1), =Ï(W_rx_t+U_rh_t-1), (22) ~t h_t =tanhâĄ(Whât+Uhâ(tâtâ1)), = (W_hx_t+U_h(r_t _t-1)), (23) t _t =(1ât)âtâ1+tâ~t. =(1-z_t) _t-1+z_t h_t. (24) Here tz_t is the update gate, tr_t the reset gate, ~t h_t the candidate state, Ï the logistic sigmoid, and â element-wise multiplication. We use dx=64d_x=64 (graph embedding dimension) and dh=32d_h=32 (opponent latent dimension). 5 Problem Formulation 5.1 CAGE-4 as a Constrained Decentralized Partially Observable Markov Decision Process (Dec-POMDP) We instantiate the Dec-POMDP formalism (oliehoek2016concise) (Section 4) with N=5N=5 CAGE âBlueâ agents, episode length T=500T=500, and Îł=0.99Îł=0.99. The global simulator state is not observed directly by the response policies. At each step, each Blue agent receives a local binary observation, selects a discrete response action, and the joint Blue action is applied together with the CAGE Red transition dynamics in CybORG. The learning problem is therefore to optimise decentralised Blue policies from partial observations while satisfying episode-level operational budgets. The MDP notation in Section 4 is used as background: the hidden simulator state sts_t defines the transition process, but the implemented policies condition on local observations otio_t^i or on the derived ACD3 context vector. Thus equations written in state notation should be read as their partially observed implementation counterparts after replacing sts_t by otio_t^i, to_t, or ctxtictx_t^i as specified below. Notation used throughout. Subscript t indexes an environment step, superscript i indexes a Blue agent, superscript k indexes an operational constraint, and e indexes an episode. The policy ÏΞ _Ξ is executed independently by each Blue agent. The formal objective is written in terms of mean team reward and episode-level operational budgets, while the implemented ACD3 PPO update stores per-agent reward and cost streams and then reports the same costs after aggregating over agents and time. Rewards are discounted when constructing PPO advantages; operational costs are recorded as undiscounted totals because the budgets are governance limits rather than reward-shaping terms. Thus rtir_t^i denotes the simulator reward observed by agent i, rÂŻt=15ââirti r_t= 15 _ir_t^i is the team reward used in the objective, cti,kc_t^i,k is the immediate cost proxy for agent i, and ctk=âicti,kc_t^k= _ic_t^i,k denotes the audited team cost at step t. The episode total âtctk _tc_t^k is compared with budget BkB_k. This distinction is important for reproducibility: the optimiser uses discounted reward advantages, while the reported safety metrics and Lagrange multiplier updates use undiscounted episode cost totals. Observation space. All agents share a uniform observation dimension d=210d=210. The observation otiâ0,1210o_t^iâ\0,1\^210 has the structure: oti o_t^i =[ÏtââB0ââB1ââB2ââM], = [ _t\;\|\;B_0\;\|\;B_1\;\|\;B_2\;\|\;M ], (25) Ïtâ0,1,2,Bjâ0,159,Mâ0,132. _tâ\0,1,2\, B_jâ\0,1\^59, Mâ\0,1\^32. Here Ït _t is the mission phase (Preplanning, MissionA, MissionB), BjB_j is subnet block j, and M encodes inter-agent messages (44 agents Ă 88 bits). Agents 0â33 each monitor one subnet (blocks B1B_1, B2B_2 are zero-padded); agent 44 monitors three subnets (all three blocks are real). Each 5959-bit block encodes: Bj B_j =[sjââbjââmjââpjâânj], = [s_j\;\|\;b_j\;\|\;m_j\;\|\;p_j\;\|\;n_j ], (26) sj,bj,mj s_j,b_j,m_j â0,19,pj,njâ0,116. â\0,1\^9, p_j,n_jâ\0,1\^16. Here sjs_j is the subnet one-hot vector, bjb_j marks blocked subnets, mjm_j stores communications-policy mismatch, and pjâ[h]p_j[h] (resp. njâ[h]n_j[h]) equals 11 if a malicious process (resp. network connection) event has been detected on host h in subnet j. Hosts are ordered: server hosts hâ0,âŠ,5hâ\0,âŠ,5\ followed by user hosts hâ6,âŠ,15hâ\6,âŠ,15\. This is the observation-to-graph interface used throughout the paper: the flat 210-bit observation is environment-specific, while the resulting hostâsubnet graph representation is the reusable input expected by the graph encoders and Graph Counterfactual Risk Propagation (G-CRP) screen. Action space. Each agent selects from |i|â€242|A^i|†242 discrete actions. Actions decompose into K=9K=9 types: = =\ Sleep,Monitor,Analyse, Sleep,\; Monitor,\; Analyse, (27) Remove,Restore,BlockZone, Remove,\; Restore,\; BlockZone, AllowZone,DeployDecoy,Other. AllowZone,\; DeployDecoy, Other\. For compact notation, BlockZone and AllowZone denote the simulator actions BlockTrafficZone and AllowTrafficZone. The time durations are ÎŽ:â1,2,3,5ÎŽ:Tâ\1,2,3,5\: Sleep, Monitor, BlockZone, and AllowZone take one simulator step; Analyse and DeployDecoy take two; Remove takes three; Restore takes five. 5.2 Operational Cost Signals Definition 5.1 (Three operational cost signals). At step t for agent i, with action atia_t^i and observation otio_t^i: cti,down c_t^i,down =â[Ïâ(ati)=Restore], =1[Ï(a_t^i)=Restore], (28) cti,fw c_t^i,fw =â[Ïâ(ati)âBlockZone,AllowZone], =1[Ï(a_t^i)â\BlockZone,AllowZone\], (29) cti,fp c_t^i,fp =cti,downâ â[âj,h(pjâ[h]âšnjâ[h])=0], =c_t^i,down·1\! [ _j,h(p_j[h] n_j[h])=0 ], (30) where Ïâ(a)Ï(a) denotes the action type of a. The indicator in (30) is zero when at least one malicious flag is active, so a Restore action is counted as false-positive only when no alert evidence is visible in the decoded observation. The team cost audited in the benchmark is ctk=âicti,kc_t^k= _ic_t^i,k. These are proxy labels derived from simulator action names and visible alerts: Restore contributes downtime, Restore without visible alert evidence contributes false-positive response cost, and BlockTrafficZone and AllowTrafficZone contribute firewall-change cost. A violation of any one budget is a violation of the episode safety contract. Definition 5.2 (Safety contract). Policy Ï satisfies the safety contract in episode e iff ât=0Tâ1ctkâ€Bk,âkâ, _t=0^T-1c_t^k†B_k, â k , (31) with =down,fw,fpK=\down,fw,fp\, Bdown=50B_down=50, Bfw=20B_fw=20, Bfp=10B_fp=10. The per-constraint violation rate is Pâ(violk)=ââ(âtctk>Bk)P(viol^k)=P( _tc_t^k>B_k). 5.3 Constrained Joint Objective maxÏâĄJrâ(Ï) _Ï\;J_r(Ï) =Ïâ[ât=0Tâ1ÎłtârÂŻt], =E_Ï\! [ _t=0^T-1Îł^t r_t ], (32) s.t.Jckâ(Ï) .t. J_c_k(Ï) =Ïâ[ât=0Tâ1ctk]â€Bk,âk. =E_Ï\! [ _t=0^T-1c_t^k ]†B_k, â k. where rÂŻt=15ââi=15rti r_t= 15 _i=1^5r_t^i is the mean team reward. Operational costs are audited as undiscounted episode totals, matching the budget accounting used by the shield and evaluation metrics. Lagrangian relaxation following the constrained objective. ââ(Ï,) (Ï, λ) =Jrâ(Ï)ââkâλkâ(Jckâ(Ï)âBk), =J_r(Ï)- _k _k (J_c_k(Ï)-B_k ), (33) λ â„. 0. Dual gradient ascent. Multipliers are updated after each episode batch: λk _k â[λk+ηλâ(J^ckâBk)]+,ηλ=0.01, â [ _k+ _λ( J_c_k-B_k) ]_+, _λ=01, (34) J^ck J_c_k =1Mââe=1Mât=0Teâ1ce,tk. = 1M _e=1^M _t=0^T_e-1c_e,t^k. where M is the number of complete episodes in the PPO update batch. 5.4 Graph Observation Encoding Graph construction. For each agent i and step t, we build a graph Gti=(V,E,X)G_t^i=(V,E,X) from otio_t^i: âą Nodes: One node per host slot per monitored subnet (1616 slots: 66 server, 1010 user), plus one subnet-level node. Agents 0â33: |V|=17|V|=17. Agent 44: |V|=51|V|=51. âą Edges: Bidirectional star between each host node and its subnet node; inter-subnet edges for agent 44 where the communication policy permits. âą Node features: vââ9x_v ^9 defined below. Node feature vector. For node v corresponding to host index h in subnet j: v _v =[xvsrv,xvusr,xvsub,pj[h], = [x_v^srv,x_v^usr,x_v^sub,p_j[h], (35) nj[h],mÂŻj,Ïtâ€]â€ââ9, n_j[h], m_j,e_ _t ] ^9, where xvsrv=1x_v^srv=1 iff hâ0,âŠ,5hâ\0,âŠ,5\, xvusr=1x_v^usr=1 iff hâ6,âŠ,15hâ\6,âŠ,15\, xvsub=1x_v^sub=1 for the subnet summary node, mÂŻj m_j is the mean communications-policy mismatch in subnet j, and Ïtâ0,13e_ _tâ\0,1\^3 is the mission-phase one-hot vector. GAT encoder. Applying Equations (19)â(21) with din=9d_in=9, K1=4K_1=4 heads, per-head dimension 1616 (Layer 1 output: â64R^64), then K2=1K_2=1 head, dimension 6464 (Layer 2 output: â64R^64): ti=1|V|ââvâVv(2)ââ64.g_t^i= 1|V| _vâ Vh_v^(2) ^64. (36) 5.5 Factorised Action Policy Every action aâ0,âŠ,|i|â1aâ\0,âŠ,|A^i|-1\ decomposes into a type Ïâ(a)âÏ(a) and a target index Μâ(a)â0,âŠ,MÏâ(a)â1Μ(a)â\0,âŠ,M_Ï(a)-1\, where MÏM_Ï is the number of valid targets for type Ï and Mmax=maxÏâĄMÏM_ = _ÏM_Ï. The factorised policy is: ÏΞâ(aâŁctxt) _Ξ(a _t) =ÏΞtypeâ(Ïâ(a)âŁctxt)â ÏΞtgtâ(Μâ(a)âŁctxt,Ïâ(a)), = _Ξ^type(Ï(a) _t)· _Ξ^tgt(Μ(a) _t,Ï(a)), (37) ÏΞtypeâ(ÏâŁctx) _Ξ^type(Ï ) =expâĄ(wÏâ€âctx+αÏ)â ÏâÏâČexpâĄ(wÏâČâ€âctx+αÏâČ)â ÏâČ, = (w_Ï ctx+ _Ï)·1_Ï _Ï (w_Ï ctx+ _Ï )·1_Ï , (38) ÏΞtgtâ(ΜâŁctx,Ï) _Ξ^tgt(Μ ,Ï) =expâĄ(fΜâ(ctx,Ï)Μ)â ΜâΜâČexpâĄ(fΜâ(ctx,Ï)ΜâČ)â ΜâČ, = (f_Μ(ctx,e_Ï)_Μ)·1_Μ _Μ (f_Μ(ctx,e_Ï)_Μ )·1_Μ , (39) where Ï,Μ1_Ï,1_Μ mask invalid types/targets, Ïââ16e_Ï ^16 is a learnable type embedding, and fΜ:â120+16ââMmaxf_Μ:R^120+16 ^M_ is a two-layer MLP. The joint log-probability is logâĄÏΞâ(aâŁctx)=logâĄÏΞtypeâ(Ïâ(a))+logâĄÏΞtgtâ(Μâ(a)) _Ξ(a )= _Ξ^type(Ï(a))+ _Ξ^tgt(Μ(a)), and the entropy decomposes as H[ÏΞ]=H[Ïtype]+Ï[H[Ïtgt(â âŁÏ)]]H[ _Ξ]=H[Ï^type]+E_Ï[H[Ï^tgt(· Ï)]]. 6 ACD3-GAT 6.1 Method Overview ACD3-GAT solves the constrained Decentralized Partially Observable Markov Decision Process (Dec-POMDP) in Equation (32) by combining four mechanisms. First, each agent converts its binary CAGE-4 observation into a hostâsubnet graph and encodes it with a Graph Attention Network (GAT). Second, the graph embedding is concatenated with a recurrent opponent embedding, the remaining budget state, and an uncertainty signal to form the policy context. Third, a factorised actor proposes an action type and target, while a set of reward, cost, tail-risk, and exploitability critics supplies PPO training signals. Fourth, before execution, a graph-risk shield evaluates admissible counterfactual actions with QsafeQ_safe and replaces unsafe proposals when the safety contract requires it. Figures 1 and 2 give the architectural and closed-loop views used during training and evaluation. Figure 1 should be read from left to right as the implemented data path. The CAGE-4 network state first becomes the binary Blue observation otio_t^i; the parser converts that simulator-specific vector into a hostâsubnet graph Gt=(V,E,X)G_t=(V,E,X). The GAT encoder produces the graph embedding tg_t, and the policy context concatenates tg_t with recurrent opponent state tz_t, budget state tb_t, and uncertainty tu_t before the factorised actor samples an action type and target. The lower path in the same figure is the safety path: the shield scores candidate actions with QsafeQ_safe, using graph risk, operational cost multipliers, and remaining budgets, and only then emits the action sent to CybORG. The feedback arrows therefore correspond exactly to the two quantities that make ACD3-GAT different from reward-only graph MARL: executed-action costs update the remaining episode budgets at the next decision step, while batch-level costs update the Lagrange multipliers used by the next PPO optimisation. Figure 2 translates the same architecture into the temporal order of one interaction step: observe otio_t^i, parse GtiG_t^i, encode tig_t^i, build ctxtictx_t^i, propose apropa_prop, screen it into aexeca_exec, execute in CybORG, store reward and costs, and update budgets and multipliers on their respective time scales. The environment produces observations, as the policy proposes a response action, the shield selects the executable action, the CybORG returns reward and operational costs, and the remaining budgets and Lagrange multipliers determine how the next decision is evaluated. The Algorithm 1 gives the same procedure in executable order. The equations below define each quantity used in that loop. Together with the problem notation in Section 5, they specify the observation parser, graph encoder, recurrent context, actor factorisation, critic targets, composite advantage, action shield, and dual update used by the reported implementation. The figure is deliberately architectural rather than decorative: learned modules are the GAT, recurrent context, actor, and critic heads; deterministic or governance modules are the observation parser, budget accounting, and Graph Counterfactual Risk Propagation (G-CRP) shield. The graph parser and valid-action layout are environment interfaces, while the constrained policy, graph encoder, cost critics, and shielding equations are the reusable ACD3 method components. Figure 1: ACD3-GAT system architecture. Partial CAGE-4 observations are parsed into hostâsubnet graphs, encoded with graph attention, combined with opponent, budget, and uncertainty context, and passed to a factorised actor. The graph-risk shield screens the proposed action with current safety-contract multipliers; observed costs update the remaining budgets and the next dual step. The parser and action layout are simulator-specific interfaces, while the graph encoder, constrained policy, and safety-contract feedback are reusable components. Figure 2: Closed-loop safety-contract update used by ACD3-GAT. Each decision observes the cyber range, encodes graph and budget context, proposes and screens an action, executes the selected response in CybORG, and stores the resulting reward, costs, and context. Remaining budgets update at the next step, while Lagrange multipliers update after the PPO episode batch, separating execution-time safety from slower policy learning. 6.2 From Safety Contract to Policy Update The implemented update follows directly from the constrained objective in Equation (32). For a fixed multiplier vector λ, maximising the Lagrangian in Equation (33) is equivalent to maximising a reward signal in which operational costs reduce the policy advantage. Using the policy-gradient identity in Equation (3), the unconstrained reward advantage A^tr A_t^r is therefore replaced by a Lagrangian advantage A^tLag=A^trââkâdown,fw,fpλkâA^tck, A_t^Lag= A_t^r- _kâ\down,fw,fp\ _k A_t^c_k, (40) where A^tr A_t^r and A^tck A_t^c_k are both computed with the same GAE recursion but from different scalar streams: simulator reward for A^tr A_t^r and operational cost proxy cti,kc_t^i,k for A^tck A_t^c_k. This is the first point where the safety contract enters the policy update: actions that improve security reward can still be discouraged when they increase expected downtime, firewall disruption, or false-positive Restore cost. ACD3-GAT then augments this Lagrangian advantage with three additional signals that are present in the implementation. CVaR reweighting emphasises the worst-return episodes in the PPO batch, adaptive opposing-policy evaluation can contribute an exploitability signal, and the override term penalises excessive shield or human-governance burden. The resulting scalar used in the PPO surrogate is A^tACD3 A_t^ACD^3 =A^tLagâÎČâA^tCVaR = A_t^Lag-ÎČ A_t^CVaR (41) âηâA^texploitâÎŒâA^toverride, -η A_t^exploit-ÎŒ A_t^override, (ÎČ,η,ÎŒ) (ÎČ,η,ÎŒ) =(0.1,0.05,0.01). =(1,05,01). Before entering the clipped PPO objective, this advantage is normalised within the batch: A~t A_t =A^tACD3âÎŒAÏA+Δ, = A_t^ACD^3- _A _A+ , (42) Δ =10â8. =0^-8. Equation (12) is then applied with A~t A_t in place of the standard reward-only advantage. The multiplier update closes the loop. After a PPO batch, the observed episode cost estimate J^ck J_c_k is compared with its budget BkB_k: λk _k â[λk+ηλâ(J^ckâBk)]+, â [ _k+ _λ( J_c_k-B_k) ]_+, (43) ηλ _λ =0.01. =01. If a cost remains below budget, the corresponding multiplier is unchanged or decreases toward zero through projection; if it exceeds budget, the multiplier increases and future PPO updates penalise that cost more strongly through Equation (40). In parallel, the action shield in Section 6.6 applies the current remaining budgets bkâ(t)b_k(t) at decision time. Thus the method has two safety mechanisms with different time scales: Lagrangian learning shapes the policy across batches, while the shield prevents budget-exhausting actions during an episode. 6.3 Context Vector The ACD3-GAT policy conditions on a 120120-dimensional context vector assembling four representations: ctxti=[tiââtiââtâ„t]ââ64+32+16+8,ctx_t^i= [g_t^i\;\|\;z_t^i\;\|\;b_t\;\|\;u_t ] ^64+32+16+8, (44) where each component is derived from otio_t^i, the recurrent episode history, or the current safety-contract state. Graph embedding tiââ64g_t^i ^64. Obtained by the two-layer GAT encoder (Equations 19â36) applied to GtiG_t^i. Opponent embedding tiââ32z_t^i ^32. A GRU (Equations 22â24) encodes the recent observation trajectory: ti=GRU64â32â(ti,tâ1i),0i=.z_t^i=GRU_64â 32(g_t^i,\,z_t-1^i), _0^i=0. (45) This provides an implicit, online estimate of the CAGE Red process without requiring explicit strategy identification. When Red-process strategy labels yâ0,âŠ,3yâ\0,âŠ,3\ and a non-zero auxiliary weight are available (e.g., from league training), an auxiliary classification loss âopp=ââj=03yjâlogâĄy^jâ(t)L_opp=- _j=0^3y_j y_j(z_t) (46) encourages tz_t to encode discriminative opponent information. Here j indexes the Red-process class, yjy_j is the one-hot target label, and y^jâ(t) y_j(z_t) is the classifierâs predicted probability for class j from the recurrent state. Budget embedding tââ16b_t ^16. The remaining operational budgets are embedded via a linear layer: t _t =tanhâĄ(Wbât+b), = \! (W_b\,q_t+d_b ), (47) t _t =[bdownâ(t)Bdown,bfwâ(t)Bfw,bfpâ(t)Bfp]â€, = [ b_down(t)B_down, b_fw(t)B_fw, b_fp(t)B_fp ] , Wb W_b ââ16Ă3,bââ16. ^16Ă 3, _b ^16. where bkâ(t)=maxâĄ(0,BkââÏ=0tâ1cÏk)b_k(t)= (0,B_k- _Ï=0^t-1c_Ï^k). The network therefore receives normalised remaining budgets in [0,1][0,1] rather than raw episode totals. Uncertainty signal utâ[0,1]u_tâ[0,1] and embedding tââ8u_t ^8. The entropy of the action-type distribution conditioned on tg_t alone (computed before the full context, avoiding circularity) is: ut u_t =Hâ(softmaxâ(Wuâti+u))logâĄK, = H\! (softmax(W_ug_t^i+d_u) ) K, (48) Wu W_u ââ9Ă64,uââ9,K=9. ^9Ă 64, _u ^9, K=9. normalised to [0,1][0,1]. A learned affine map t=Wueâut+ueââ8u_t=W_ue\,u_t+d_ue ^8 embeds it into the context. High utu_t indicates that the graph state alone is insufficient to determine the action type; it is used as contextual evidence for the policy and confidence monitor. In the control loop, utu_t is therefore not a separate objective: it is a compact uncertainty feature that enters the context vector and can trigger the governance fallback when the configured confidence or G-CRP uncertainty thresholds are exceeded. 6.4 Multi-Objective Critic Set ACD3-GAT maintains four distinct critic heads, all conditioned on ctxtctx_t: Vrâ(ctxt) V_r(ctx_t) ââ, , (49) Vckâ(ctxt) V_c^k(ctx_t) ââ,kâ1,2,3, , kâ\1,2,3\, (50) VCVaRâ(ctxt) V_CVaR(ctx_t) ââ, , (51) Vexploitâ(ctxt) V_exploit(ctx_t) ââ. . (52) Each critic head is a two-layer MLP with hidden dimension 6464. The reward head predicts the security-return baseline, the three cost critics correspond to downtime, firewall-change cost, and false-positive response cost, the CVaR head supports tail-risk accounting, and the exploitability head stores the adaptive Red-process signal when enabled. The cost critics are stacked as c=(Vc1,Vc2,Vc3)â€ââ3V_c=(V_c^1,V_c^2,V_c^3) ^3. 6.5 Composite ACD3 Advantage Per-objective GAE advantages. Using the recursion in Equations (9)â(10): A^tr A_t^r =GAE(rÏi,Vr;Îł=0.99,λGAE=0.95), =GAE(\r_Ï^i\,V_r;Îł=0.99, _GAE=0.95), (53) A^tck A_t^c_k =GAE(cÏi,k,Vck;Îł=0.99,λGAE=0.95). =GAE(\c_Ï^i,k\,V_c^k;Îł=0.99, _GAE=0.95). (54) During ACD3 optimisation these streams are computed per Blue agent from the stored rollout of that agent; the reported episode metrics aggregate the same cost proxies across agents and time. CVaR advantage. In each batch of M=8M=8 episodes ranked by return, the kâ=maxâĄ(1,â0.1âMâ)k^*= (1, 0.1M ) worst episodes receive weight we=1/kâw_e=1/k^*; all others receive we=0w_e=0: A^tCVaR=weâA^tr. A_t^CVaR=w_e A_t^r. (55) This is an episode-level reweighting of the policy-gradient signal rather than a separate distributional value estimator. Exploitability advantage. When adaptive opposing-policy evaluation is enabled, a short Red PPO mini-loop is run every IexploitI_exploit episodes (default 5050; 3 updates Ă 2 episodes each): A^texploit=0.1âÎŽÂŻ,ÎŽÂŻ=UBlueâ(0)âUBlueâ(3), A_t^exploit=0.1\, ÎŽ, ÎŽ=U_Blue(0)-U_Blue(3), (56) where UBlueâ(k)U_Blue(k) is Blueâs mean return after k Red-process adaptation steps and ÎŽÂŻ ÎŽ is the running mean over the last 1010 stored measurements. If no adaptive opposing-policy measurement has been stored, this term is zero. In the reported experiments, coupled adaptive Red-process evaluation is used primarily as a post-training stress test rather than as the headline optimisation target. Override advantage. The centred override indicator penalises excessive analyst burden: A^toverride=â[shield triggered at ât]âÏÂŻ, A_t^override=1[shield triggered at t]- Ï, (57) where ÏÂŻ Ï is the mean override rate in the current batch. ACD3 composite advantage. Substituting the terms above into Equation (41) gives the implemented scalar advantage: A^tACD3 A_t^ACD^3 =A^trââk=13λkâA^tckâ0.1âA^tCVaR = A_t^r- _k=1^3 _k A_t^c_k-1\, A_t^CVaR (58) â0.05âA^texploitâ0.01âA^toverride. -05\, A_t^exploit-01\, A_t^override. Normalised before use: A~t A_t =A^tACD3âÎŒAÏA+Δ, = A_t^ACD^3- _A _A+ , (59) Δ =10â8. =0^-8. PPO update. Equation (12) is applied to A~t A_t with Ï”=0.2Δ=0.2, cvf=0.5c_vf=0.5, cent=0.005c_ent=0.005, and 44 epochs over the concatenated episode batch. The same update also fits the reward, cost, CVaR, and exploitability critic heads using clipped value losses. 6.6 Budget-Aware Counterfactual Shield Remaining budget. The remaining budget is defined as: bkâ(t)=maxâĄ(0,BkââÏ=0tâ1cÏk)b_k(t)= (0,B_k- _Ï=0^t-1c_Ï^k) (60) Admissible action set. The reported shield uses the cost proxies as a hard exhaustion guard: once a budget has been depleted, actions with positive immediate proxy cost for that budget are blocked. Costly actions before exhaustion remain admissible, but are penalised by the Lagrangian and G-CRP score. safeâ(t) _safe(t) =aâi:bk(t)>0 = \a ^i:\;b_k(t)>0 (61) âšc^k(a)=0,âk, \; c_k(a)=0,\;â k \, where c^kâ(a) c_k(a) is the immediate cost proxy (Equations 28â30). In this expression, iA^i is agent iâs discrete action set, bkâ(t)b_k(t) is the remaining budget for constraint k before action execution, and the condition is applied for all three operational costs. Shielded proposal. The policy first proposes an action from the factorised actor. For efficiency, the reported G-CRP shield ranks one representative candidate per action type. It accepts the proposal when its action type is the top ranked admissible type; otherwise it replaces the proposal with the highest-ranked admissible type, or falls back to Sleep when no safe action is available: Ïpropâ(aâŁoti) _prop(a o_t^i) =ÏΞâ(aâŁctxt), = _Ξ(a _t), (62) aexec a_exec =Shieldâ(aprop,Gt,safeâ(t)). =Shield\! (a_prop,G_t,A_safe(t) ). Here Ïprop _prop is the stochastic policy distribution before screening, apropa_prop is the sampled proposal, and aexeca_exec is the action actually submitted to CybORG after applying the admissible set and G-CRP ranking. When the shield replaces the proposed action, the implementation stores and updates PPO on the executed action by recomputing its log-probability under the current policy and action mask. Thus the trajectory logger, cost accounting, and likelihood ratio are all tied to the action actually submitted to CybORG. The shield is nevertheless a non-differentiable intervention and can bias the policy-gradient estimate, this is why override frequency is recorded and penalised in Equation (57). The rule blocks after the relevant remaining budget is exhausted, so episode-level compliance is evaluated empirically through the violation-rate metrics rather than asserted analytically from the shield alone. 6.7 G-CRP: Graph Counterfactual Risk Propagation The rule shield (61) reacts to cost proxies but cannot estimate how the networkâs compromise state will evolve. G-CRP provides predictive risk scoring via graph propagation. Node risk beliefs. Given the proc and net flags in otio_t^i: p^vâ(t)=12âpjâ[hâ(v)]+12ânjâ[hâ(v)]â[0,1]. p_v(t)= 12\,p_j[h(v)]+ 12\,n_j[h(v)]â[0,1]. (63) Here hâ(v)h(v) maps host node v to its decoded host slot, while pjâ[h]p_j[h] and njâ[h]n_j[h] are the process-alert and network-alert bits for host h in subnet j. Action effects on beliefs. Defensive actions modify beliefs before propagation: Restoreâ(v) (v) :p^vâ0.05âp^v, :\; p_vâ 0.05\, p_v, (64) Removeâ(v) (v) :p^vâ0.30âp^v, :\; p_vâ 0.30\, p_v, (65) Block :αâ0.10âα, :\;αâ 0.10\,α, (66) Allow :αâ1.05âα, :\;αâ 1.05\,α, (67) Decoyâ(v) (v) :p^vâ0.70âp^v, :\; p_vâ 0.70\, p_v,\; p^wâ0.85âp^wââwââ(v), p_wâ 0.85\, p_w\;â w (v), (68) where α=0.3α=0.3 is the default edge influence. Independent cascade propagation. For L=2L=2 steps: p^v(l+1)=maxâĄ(p^v(l), 1ââuâinâ(v)(1âαâp^u(l))). p_v^(l+1)= \! ( p_v^(l),\;1- _u ^in(v)(1-α\, p_u^(l)) ). (69) The max enforces monotonicity: compromise probability does not decrease through propagation alone. The superscript (l)(l) indexes the propagation depth, and inâ(v)N^in(v) is the set of nodes with directed influence into v under the constructed hostâsubnet graph. The reported shield uses L=2L=2 as a local neighbourhood screen: one hop captures immediate subnet-to-host effects and the second hop captures the next propagation opportunity without turning the shield into a rollout search. Predicted security risk. With asset weights Ïvâ2,1,0.5 _vâ\2,1,0.5\ for server/user/subnet nodes: S^â(Gt,a)=âvâVÏvâp^v(L)â(tâŁa)âvÏv. S(G_t,a)= _vâ V _v\, p_v^(L)(t a) _v _v. (70) The resulting S^â(Gt,a) S(G_t,a) is a normalized post-action graph-risk score: larger values indicate higher predicted residual compromise risk after applying candidate action a and propagating for L steps. Q-safe score and action selection. Qsafeâ(Gt,a) Q_safe(G_t,a) =âS^â(Gt,a)ââkλkâC^kâ(a)âÎČâU^â(Gt,a), =- S(G_t,a)- _k _k C_k(a)-ÎČ\, U(G_t,a), (71) aâ a^* =argâĄmaxaâsafeâ(t)âĄQsafeâ(Gt,a), = _a _safe(t)Q_safe(G_t,a), (72) where C^kâ(a) C_k(a) uses the same cost proxies as (28)â(30), and ÎČ=0.1ÎČ=0.1 weights the budget-risk penalty U U. For deterministic G-CRP, U^â(Gt,a)=1 U(G_t,a)=1 if the immediate proxy cost would exceed any remaining budget; otherwise it is the mean normalised proxy cost across the three budgets. For learned G-CRP, the same symbol denotes the predicted violation probability. In the reported configuration ÎČ=0.1ÎČ=0.1, matching the G-CRP shield code and the ACD3 configuration. The cascade coefficients and asset weights are hand-specified operational priors rather than calibrated causal estimates; they define the deterministic screen used in this study and should be recalibrated before transfer to a different cyber range or enterprise topology. The article reports deterministic G-CRP as the deployed action-screening rule; the learned G-CRP fit is treated as a supervised model-fitting result, not as a separate policy-performance claim. Learned G-CRP fitting objective. When the learned G-CRP component is trained from collected trajectories, the model predicts next-step process-alert probabilities, next-step network-alert probabilities, immediate cost proxies, and violation probability. Its supervised loss is âGâ-âCRP= _G -CRP= BCEâ(^proc,proc)+BCEâ(^net,net) \;BCE\! ( p^proc,p^proc )\;+\;BCE\! ( p^net,p^net ) (73) +0.1ââ„^ââ„22+BCEâ(v^,v), +1\, c-c _2^2\;+\;BCE( v,v), where BCE denotes binary cross-entropy, hatted quantities are model predictions, unhatted quantities are supervised targets from the next logged transition, c is the immediate cost-proxy vector, and v is the binary violation label. This loss matches the replication-package trainer and is reported only as a component fit; the policy benchmark uses the deterministic G-CRP screen unless explicitly stated otherwise. 6.8 ACD3-TCGS: Temporal Contract Graph Shielding The deterministic G-CRP screen estimates immediate graph risk from the current observation. Temporal Contract Graph Shielding (TCGS) adds a learned recurrent contract-risk model that asks a different question: given the recent trajectory history and a candidate action, how likely is the episode to cross an operational budget within the next decision horizon? TCGS is implemented as an opt-in diagnostic extension. It reads existing safety-labelled trajectories, trains a small recurrent risk model, and writes separate evaluation traces without modifying completed checkpoints or training outputs. For each step, the feature vector combines the decoded binary observation, immediate operational costs, cumulative cost normalized by the budget, remaining budget fraction, mission phase, action-type frequencies, and scaled reward: t=[ _t= [ otââtâât/â„[ât]+/ o_t\,\|\,c_t\,\|\,C_t/B\,\|\,[B-C_t]_+/B (74) â„onehot(Ït)â„(at)â„rÂŻt/100]ââ233. \,\|\,onehot( _t)\,\|\,m(a_t)\,\|\, r_t/00 ] ^233. Here tC_t is the cumulative episode cost before the candidate action, B is the three-budget vector, and â(at)m(a_t) is the normalized action-type count vector used by the implementation. For candidate action a, TCGS forms the length-L history Htâ(a)=(tâL+1,âŠ,tâ1,tâ(a)),L=8.H_t(a)= (x_t-L+1,âŠ,x_t-1,x_t(a) ), L=8. (75) The recurrent model is a Gated Recurrent Unit encoder with hidden dimension 96 and two heads: t _t =GRUÏâ(Htâ(a)), =GRU_Ï(H_t(a)), (76) ât _t =Wvât, =W_vh_t, ^t d_t =ReLUâ(Wdât), =ReLU(W_dh_t), Here ât _t is the three-dimensional vector of violation logits and ^t d_t is the non-negative predicted normalized future cost increment. Applying a logistic sigmoid to logit âtk _t^k estimates the probability that budget k will be violated within horizon H=100H=100: pÏkâ(Ht,a)=PrÏâĄ[maxÏâ[t,t+H]ââsâ€Ïcsk>Bk|Htâ(a)].p_Ï^k(H_t,a)= _Ï\! [ _Ïâ[t,t+H] _sâ€Ïc_s^k>B_k\; |\;H_t(a) ]. (77) In Equation (77), s and Ï are step indices inside the future prediction window and BkB_k is the corresponding operational budget. The second head ^t d_t estimates the normalized future cost increment. The supervised training objective is âTCGS= _TCGS= âkBCEâ(âtk,ytk) \; _kBCE\! ( _t^k,y_t^k ) (78) +0.1ââ„^tâtâ„22, +1\, d_t-d_t _2^2, where ytky_t^k indicates whether budget k is crossed within the prediction horizon and td_t is the future normalized cost increment. At evaluation time, a frozen ACD3-GAT policy proposes apropa_prop. TCGS accepts the proposal only if the predicted downtime risk is below the deployability threshold and the immediate proxy cost does not exhaust a budget: Ïtâ(a) _t(a) =â[pÏdownâ(Ht,a)â€Ï”] =1\! [p_Ï^down(H_t,a)â€Î” ] (79) â â[Ctk+c^kâ(a)â€Bk,âk]. ·1\! [C_t^k+ c_k(a)†B_k,\;â k ]. Here Ïtâ(a) _t(a) is an accept/reject indicator for candidate action a, CtkC_t^k is the cumulative cost already incurred in the current episode, and c^kâ(a) c_k(a) is the immediate cost proxy for that candidate. aexec=aprop,Ïtâ(aprop)=1,argâĄminaâvalidâĄÎštâ(a),Ïtâ(aprop)=0.a_exec= casesa_prop,& _t(a_prop)=1,\\ _a _valid _t(a),& _t(a_prop)=0. cases (80) with Ï”=0.05Δ=0.05 in the reported diagnostic. If the proposal is rejected, the shield searches the valid action set and chooses the candidate with the smallest lexicographic risk tuple Κtâ(a) _t(a). The implemented ranking Κtâ(a) _t(a) =(ÏtB(a),ÏtP(a),ptdown(a), = ( _t^B(a), _t^P(a),p_t^down(a), (81) pÂŻt(a),cÂŻt(a)), p_t(a), c_t(a) ), ÏtBâ(a) _t^B(a) =[âk:Ctk+c^k(a)>Bk], =1[â k:\;C_t^k+ c_k(a)>B_k], ÏtPâ(a) _t^P(a) =â[pÏdownâ(Ht,a)>Ï”], =1[p_Ï^down(H_t,a)>Δ], ptdownâ(a) p_t^down(a) =pÏdownâ(Ht,a), =p_Ï^down(H_t,a), pÂŻtâ(a) p_t(a) =13ââkpÏkâ(Ht,a), = 13 _kp_Ï^k(H_t,a), cÂŻtâ(a) c_t(a) =âkc^kâ(a) = _k c_k(a) orders candidates first by hard budget feasibility ÏtB _t^B, then by downtime-risk threshold feasibility ÏtP _t^P, followed by predicted downtime risk, mean predicted violation risk, and immediate proxy cost. It evaluates one representative valid candidate per action type for speed. Thus TCGS is not an end-to-end retrained policy in this paper; it is a frozen-policy temporal shield that tests whether learned contract-risk prediction can improve action screening online. 6.9 Override Readiness Three conditions can trigger the Sleep fallback: maxaâĄÏâ(aâŁctxt) _aÏ(a _t) <Ïconf, < _conf, (82) âk:bkâ(t)=0â§c^kâ(a)>0, â k:\;b_k(t)=0\; \; c_k(a)>0, (83) Stdaâprobeâ[S^â(Gt,a)] _a _probe\! [ S(G_t,a) ] >Ïood, > _ood, (84) with Ïconfâ0.0,0.15 _confâ\0.0,0.15\ depending on the ablation and Ïood=0.7 _ood=0.7. Here Ïconf _conf is the minimum policy-confidence threshold, probeA_probe is the finite set of candidate actions probed by the G-CRP uncertainty check, and Ïood _ood is the threshold on the standard deviation of predicted graph risk across those probes. The override burden Joverride=â[âtâ[override]]J_override=E[ _t1[override]] is penalised via ÎŒâA^toverrideÎŒ A_t^override in (58). 6.10 Training Algorithm Algorithm 1 writes the training horizon as EtrainE_train because the experiments use different horizons (30, 100, and 200 episodes in the main benchmark, with 300-episode replications for the longer-horizon robustness check). Algorithm 1 ACD3-GAT Training 1:Initialise Ξ,Ïr,Ïck,ÏCVaR,Ïexpl,=Ξ, _r,\ _c_k\, _CVaR, _expl, λ=0 2:for episode e=1,âŠ,Etraine=1,âŠ,E_train do 3: Reset; 0iâ32z_0^i 0_32; bkâ(0)âBkb_k(0)â B_k; batch ââ â 4: for t=0,âŠ,499t=0,âŠ,499 do 5: GtiâG_t^i (oti)(o_t^i); tiâg_t^i (Gti)(G_t^i); tiâz_t^i (ti,tâ1i)(g_t^i,z_t-1^i) 6: utâHâ(softmaxâ(Wuâti+u))/logâĄ9u_tâ H(softmax(W_ug_t^i+d_u))/ 9; ctxtiâ[tiââtiââtâ„t]ctx_t^iâ[g_t^i\|z_t^i\|b_t\|u_t] 7: Sample atiâŒÏΞ(â âŁctxti)a_t^i _Ξ(· _t^i) 8: if confidence trigger fires (82) then 9: atiâSleepa_t^iâ Sleep 10: else 11: atiâShieldâ(ati,Gt,safeâ(t))a_t^i (a_t^i,G_t,A_safe(t)) 12: end if 13: Execute joint action; observe rtir_t^i, cti,kc_t^i,k; ctkââicti,kc_t^kâ _ic_t^i,k; bkâ(t+1)âmaxâĄ(0,bkâ(t)âctk)b_k(t+1)â (0,b_k(t)-c_t^k) 14: end for 15: batchâ âȘeâȘ\e\ 16: if |batch|=8|batch|=8 then 17: Compute A^tACD3 A_t^ACD^3 via Eq. (58); normalise 18: Run PPO update (4 epochs over the concatenated episode batch) 19: λkâ[λk+0.01â(J^ckâBk)]+ _kâ[ _k+0.01( J_c_k-B_k)]_+ 20: batchââ â 21: end if 22: if adaptive Red-process evaluation enabled and emodIexploit=0e I_exploit=0 then 23: Red mini-loop; update ÎŽÂŻ ÎŽ 24: end if 25:end for 7 Experiments 7.1 Study Design The empirical study is organised around three questions that follow directly from the safety-contract formulation: Operational constraints. Do explicit operational constraints reduce budget violations without eliminating security utility? Graph structure. Does graph structure, through graph neural network (GNN) and Graph Attention Network (GAT) encoders, improve over flat encoders in tail behaviour and mean return? Operational discipline. What does the gap between engineered CAGE 4 heuristics and naive reactive rules reveal about operational discipline? 7.2 Environment and Setup We use CAGE Challenge 4 (cage4) with 5 CAGE âBlueâ agents, episode length T=500T=500, and a finite-state machine CAGE âRedâ process as the baseline source of simulated malicious activity. All agents receive uniform 210-dimensional binary observations (padded to equal length across all agents). Operational budgets: Bdown=50B_down=50, Bfp=10B_fp=10, Bfw=20B_fw=20 per episode. These values are fixed simulator stress-test thresholds for this study, not universal SOC constants. A production transfer would require selecting BkB_k from local governance constraints and rerunning the same safety-labelled evaluation, ideally with a budget-sensitivity sweep. Reference points from prior work. In the default CAGE Challenge 4 (C4) evaluation, the published benchmark table reports a top heuristic reference of â113±35-113± 35 and a top MARL reference of â193±84-193± 84 (kiely2025cage4b). The â101±36-101± 36 score corresponds to the constant-network-size reference rather than the default C4 column (kiely2025cage4b). The best Large Language Model (LLM) agent (GPT-o1-mini, role prompting) scored ââ2888â-2888 (castro2025llm). We include the default C4 values as reference lines in our figures and treat them as prior-work references, not direct baselines, because our experiments use a safety-labelled training/evaluation wrapper and explicit cost accounting. The LLM score is reported only as a literature reference because the published result does not provide action-level traces compatible with our downtime, firewall, and false-positive safety accounting. Hyperparameters: PPO clip Ï”=0.2Δ=0.2, Îł=0.99Îł=0.99, λGAE=0.95 _GAE=0.95, ηλ=0.01 _λ=0.01, αCVaR=0.1 _CVaR=0.1. All learning methods use hidden dimension 64. ACD3-GAT uses M=8M=8 episodes per update, (ÎČ,η,ÎŒ)=(0.1,0.05,0.01)(ÎČ,η,ÎŒ)=(0.1,0.05,0.01) for CVaR, exploitability, and override terms, and an exploitability mini-loop every 50 episodes when enabled. ACD3-GAT and C-MAPPO-GAT use cent=0.005c_ent=0.005; the other MAPPO-family baselines (MAPPO-MLP, MAPPO-GNN, MAPPO-GAT, and CVaR-MAPPO) use cent=0.01c_ent=0.01. All MAPPO-family learners use mini-batches of 64. The comparison controls the main training horizon, PPO hyperparameters, and hidden dimension, but it is not parameter-count matched: graph encoders and factorised heads have different capacity from flat MLP policies. We therefore interpret encoder and architecture rows as matched-protocol comparisons, not as capacity-normalised dominance claims. All baseline rows are presented as implemented policy families under the same safety logger; prior published CAGE-4 heuristic and LLM scores are used only to calibrate the reward scale because their public outputs do not expose the action-level traces required for our operational-cost audit. Metric semantics and reporting conventions. Downtime cost is the undiscounted episode total âtctdown _tc_t^down; an episode violates the downtime contract iff this total exceeds Bdown=50B_down=50. Figures may plot either this raw cost or the ratio âtctdown/Bdown _tc_t^down/B_down; captions state which is used. CVaR-10% is the empirical mean return of the worst 10% of evaluated episodes, not merely the 10th percentile. The Pareto figures show the non-dominated points among observed methods, not a continuous optimum of the non-convex MARL objective. Confidence bands and box widths appear only when multiple seeds exist; methods with one seed or short component-check horizons are plotted without inferential claims and are labelled by their seed/episode counts in the table. For rare violation rates, the episode count should be read together with the seed count: for example, Pâ(violDT)=0.003P(viol^DT)=0.003 over 600 episodes corresponds to only a small number of observed violations and is therefore evidence of strong empirical compliance in this benchmark, not a formal probability guarantee for deployment. All operational costs are audited as undiscounted totals because budgets are governance limits, while reward learning still uses discounted returns. All learned policies sample only from the valid CybORG action mask exposed by the environment wrapper. If a budget shield changes an action, the stored trajectory records the executed action label and recomputed costs; the replication records also store the run configuration, random seed, topology seed when available, git commit, package freeze, and safety-labelled per-step trajectory needed to recompute the benchmark tables. Evaluation protocol. Each comparison reports its seed count and episode horizon explicitly. The study combines five evidence streams. First, reward-only baselines quantify the operational failure mode: Independent Proximal Policy Optimisation (IPPO), Multi-Agent PPO with a multi-layer perceptron encoder (MAPPO-MLP), and MAPPO-GAT are each evaluated at a 200-episode horizon, with IPPO and MAPPO-GAT replicated across three seeds. Second, C-MAPPO-GAT, the constrained safety baseline introduced in this paper, is evaluated over three 200-episode seeds to test whether Lagrangian costs and budget-aware action screening reduce SOC safety-contract violations. Third, ACD3-GAT is evaluated as the integrated architecture, combining graph attention, budget context, CVaR weighting, and override with three 200-episode seeds. Fourth, two additional 300-episode replications are run for MAPPO-GAT, C-MAPPO-GAT, and ACD3-GAT to test whether the main safety conclusion survives a longer horizon. Fifth, short 30-episode ACD3+AskHuman and ACD3+deterministic G-CRP component runs verify that these components operate inside the safety-contract regime and separate component behaviour from the core replicated comparison. This structure separates replicated comparisons from component checks and keeps each quantitative claim tied to its corresponding evaluation horizon. 7.3 Baselines We evaluate the following policy families: âą Sleep: always-sleep lower bound âą Random: uniform random valid actions âą Rule-based: Monitor by default; Restore on proc alert; Block on persistent net alert âą IA2C: Independent Advantage ActorâCritic with MLP encoder. Prior method: (mnih2016asynchronous) (Equations 5â8) âą IPPO: Independent Proximal Policy Optimisation with MLP encoder. Prior method: (de2020independent) âą MAPPO-MLP: Multi-Agent PPO with a centralised critic and MLP encoder. Prior method: (yu2022surprising) âą MAPPO-GNN: Multi-Agent PPO with a centralised critic and GraphSAGE encoder. Encoder prior: (hamilton2017inductive) âą MAPPO-GAT: Multi-Agent PPO with a centralised critic and GAT encoder. Encoder prior: (velickovic2018graph) (ours) âą C-MAPPO-GAT: Constrained MAPPO-GAT, introduced here as a controlled safety-contract baseline with Lagrangian cost learning (altman1999constrained; achiam2017constrained) and the configured hard budget-exhaustion fallback (ours) âą CVaR-MAPPO: Conditional Value-at-Risk (CVaR)-MAPPO with episode-level CVaR reweighting, following the CVaR tail-risk objective (rockafellar2000optimization) and treated as a safety component rather than a separate safety-contract policy claim âą ACD3-GAT: graph-attentive constrained response policy with budget context, CVaR weighting, and override mechanisms (ours; three 200-episode seeds) For MAPPO-GAT and C-MAPPO-GAT, the GAT encoder in Equations 19â36 instantiates EncÏiEnc_Ï^i in the centralised critic of Equation 16; the actor still executes from each agentâs local observation. For ACD3-GAT, the same graph embedding tig_t^i is concatenated with opponent, budget, and uncertainty context in Equation 44 before the factorised action policy and safety shield are applied. 7.4 Benchmark Method Instantiations All learned methods use the CAGE-4 observation and action spaces defined in Section 5, and they differ only in the policy update, encoder, critic information, and safety machinery. This subsection fixes the mapping between Table 2 and the implemented algorithms. Non-learning reference policies. Sleep is the deterministic policy Ïâ(ati=SleepâŁoti)=1Ï(a_t^i= Sleep o_t^i)=1. Random samples uniformly from the valid action mask validiâ(oti)A_valid^i(o_t^i): Ïrandâ(aâŁoti)=|validiâ(oti)|â1â[aâvalidiâ(oti)]. _rand(a o_t^i)=|A_valid^i(o_t^i)|^-11[a _valid^i(o_t^i)]. (85) Here validiâ(oti)A_valid^i(o_t^i) is the CybORG action mask for agent i in observation otio_t^i, so invalid simulator actions receive zero probability. The rule-based policy is a deterministic alert policy: monitor by default, restore when process alerts are present, and block traffic when persistent network alerts are visible. These policies use no learned parameters; their returns and safety metrics are computed with the same logger and cost proxies as the learned methods. IA2C and IPPO. IA2C follows Equations 5â8. IPPO uses the same independent per-agent information pattern, but replaces the vanilla actorâcritic objective with clipped PPO: ÏΞiâ(atiâŁoti),VÏiâ(oti),âIPPOi=âiCLIP+cvfââiVFâcentâHâ[ÏΞi], _ _i(a_t^i o_t^i), V_ _i(o_t^i), _IPPO^i=L^CLIP_i+c_vfL^VF_i-c_entH[ _ _i], (86) with GAE advantages from Equations 9â10. The actor parameters Ξi _i, critic parameters Ïi _i, local observation otio_t^i, and entropy/value-loss weights are the same quantities defined for IA2C and PPO in Section 4. The IPPO row uses an MLP encoder; the Factorized-IPPO row keeps the same independent PPO update but replaces the flat actor with the typeâtarget factorisation in Equation 37. MAPPO encoder variants. MAPPO-MLP, MAPPO-GNN, and MAPPO-GAT all use the centralised critic in Equation 16 and the clipped PPO objective in Equation 15. Their only architectural difference is the per-agent encoder: EncÏiâ(oti)=MLPâ(oti),MAPPO-MLP,GraphSAGEâ(Gti),MAPPO-GNN,GATâ(Gti),MAPPO-GAT.Enc_Ï^i(o_t^i)= casesMLP(o_t^i),&MAPPO-MLP,\\ GraphSAGE(G_t^i),&MAPPO-GNN,\\ GAT(G_t^i),&MAPPO-GAT. cases (87) Here EncÏiEnc_Ï^i denotes the per-agent encoder whose output is sent to the centralised critic during training; GtiG_t^i is the graph parsed from the same local observation otio_t^i. GraphSAGE and GAT are defined in Equations 18 and 19â36, respectively. Actors execute decentralised local policies, while the critic uses the concatenated five-agent embedding during training. Opponent-conditioned IPPO. Opp-IPPO keeps the independent PPO loss but augments each agentâs policy state with a recurrent opponent embedding: ti _t^i =GRUâ(Encâ(oti),tâ1i), =GRU(Enc(o_t^i),z_t-1^i), (88) ati a_t^i âŒÏΞi(â âŁoti,ti). _ _i(· o_t^i,z_t^i). Here tiz_t^i is the recurrent state for agent i, initialised to zero at the start of each episode and updated from that agentâs encoded local observation. This is the same recurrent mechanism later reused in the ACD3 context vector, but without safety budgets, CVaR weighting, or shielding. CVaR-MAPPO. CVaR-MAPPO keeps the MAPPO centralised critic and PPO update, but multiplies the policy advantage by an episode-level tail weight. For a batch of M episodes and tail fraction α=0.1α=0.1, kâ k^* =maxâĄ(1,âαâMâ), = (1, α M ), (89) we w_e =1/kâ,eâworst âkââ episodes,0,otherwise, = A^tCVaRâ-âMAPPO A_t^CVaR -MAPPO =weâA^tMAPPO. =w_e A_t^MAPPO. Here e indexes episodes in the PPO batch, wew_e is the binary tail weight assigned to the worst-return episodes, and A^tMAPPO A_t^MAPPO is the reward-only MAPPO advantage. This row therefore tests tail-risk reweighting alone, without Lagrange constraints or action shielding. Constrained MAPPO-GAT. C-MAPPO-GAT is our constrained MAPPO-GAT instantiation: it keeps the MAPPO-GAT encoder and centralised critic, adds cost-advantage penalties for the three SOC budgets, and applies the configured hard budget-exhaustion fallback: A^tCâ-âMAPPO A_t^C -MAPPO =A^trââkλkâA^tck, = A_t^r- _k _k A_t^c_k, (90) λk _k â[λk+ηλâ(J^ckâBk)]+. â[ _k+ _λ( J_c_k-B_k)]_+. Here A^tr A_t^r is the reward advantage, A^tck A_t^c_k is the advantage computed from cost stream k, λk _k is the corresponding Lagrange multiplier, and J^ck J_c_k is the observed batch cost estimate compared with budget BkB_k. This is the safety-contract baseline against which ACD3-GAT is interpreted. ACD3 variants. ACD3-GAT uses the full context vector Equation 44, factorised policy Equation 37, composite advantage Equation 58, and shielded execution Equation 62. The AskHuman and deterministic G-CRP rows are short configuration checks of the override and action-screening components inside the same ACD3 scaffold; they are not presented as separately optimised policy families. 7.5 G-CRP Shield Diagnostics To isolate the contribution of Graph Counterfactual Risk Propagation (Section 6.7), we define three shield variants while holding the policy class (ACD3-GAT) fixed: Rule-based shield. The reactive shield blocks actions whose cost proxy is positive for an exhausted budget; it uses no prediction and no graph reasoning. Deterministic G-CRP. The physics-based shield uses the cascade propagation model in Equation 69, selects argâĄmaxâĄQsafe Q_safe (Equation 71), and requires no training. Learned G-CRP. The supervised extension trains a GraphSAGE model on collected trajectories to predict next-step alert state, cost, and violation probability before action screening. The deterministic G-CRP path is included as part of the ACD3 action-screening design and is evaluated in a short configuration check. The learned G-CRP model is trained in the replication package as the supervised risk-estimation extension of the deterministic screen. The quantitative policy comparison reports the deterministic G-CRP path so that each policy-performance claim is tied to the action-screening rule used in the evaluated runs. When that component is evaluated, the intended measures are violation rate, mean return, override rate, violation-probability calibration, and cost-prediction error. 7.6 Robustness Extensions The robustness evaluation tests whether the safety contract survives outside the matched training and evaluation seed setting: âą In-distribution (ID): 5 topology seeds used during training âą Topology out-of-distribution (OOD): 5 unseen seeds (different host counts) âą Red-process OOD protocol: a separate evaluation with Red policies not used in Blue training or tuning, reserved for future quantitative claims beyond the present split-labelled seed-variation artifacts âą Mission OOD: altered mission phase schedule âą Adaptive Red process: stress curve â Red PPO trained 20 rounds against frozen Blue policies The present benchmark reports the ID setting and two robustness checks. First, IPPO, MAPPO-GAT, and constrained MAPPO-GAT are evaluated on unseen topology seeds for 20 episodes per seed, using two seeds per split and the same finite-state Red process. Second, a Red PPO league is trained for 20 updates against frozen Blue policies, with sampled Red-process actions coupled into the CybORG transition loop. The split-labelled Red-process artifacts collected in this replication package use the same CAGE finite-state Red-process family with seed variation and are therefore not interpreted as held-out Red-policy evidence. Held-out Red-process and mission-schedule variants remain part of the evaluation protocol but are not used as quantitative claims here. 8 Results 8.1 Evaluation Context All experiments use CAGE Challenge 4 (cage4) with N=5N=5 CAGE âBlueâ agents, T=500T=500 steps, and the finite-state machine CAGE âRedâ process. All agents receive a uniform 210-dimensional binary observation vector. Operational budgets: Bdown=50B_down=50, Bfp=10B_fp=10, Bfw=20B_fw=20 per episode. The main replicated comparison uses three 200-episode seeds for IPPO, MAPPO-GAT, constrained MAPPO-GAT, and ACD3-GAT. MAPPO-MLP provides a one-seed flat-encoder reference at the same horizon. For the three principal graph-based methods, two additional 300-episode replications assess whether the central safety ordering persists when training is extended. Shorter architectural component runs (100 or 30 episodes) are used only to interpret component behaviour and are labelled separately from the replicated comparison; they are not used to claim a new best policy. 8.2 Safety and Return Table 2 summarises the safety-contract benchmark. The table reports the complete direct comparison under our safety-labelled evaluation protocol: it includes the non-learning baselines, IA2C, IPPO, MAPPO with MLP/GraphSAGE/GAT encoders, factorised and opponent-conditioned IPPO variants, CVaR-MAPPO, C-MAPPO-GAT, and ACD3-GAT. The figures provide the same evidence at different levels of resolution. Figure 3 summarizes the benchmark through four linked views: budget violation, budget overrun, returnâdowntime tradeoff, and tail-risk gap. Figure 4 then unpacks that dashboard into four benchmark views: panel (a) asks whether the episode exceeds BdownB_down; panel (b) identifies which action-derived costs are responsible for the separation; panel (c) places the same methods on a returnâdowntime plane; and panel (d) checks whether mean return is hiding poor worst-tail outcomes. Figure 5 connects the final summaries to training dynamics rather than treating the table as a single endpoint, and Figure 6 makes the replication status visible by separating three-seed comparisons from one-seed or short component checks. Finally, Figure 9 reports the coupled adaptive Red-process stress test, while Figure 7 visualises the deterministic graph counterfactual model used by the shield. Published CAGE-4 and LLM numbers appear only as prior-work reference lines or table notes; they are not treated as direct baselines because they lack the same safety-labelled action traces and budget accounting. Table 2: Benchmark summary for the safety-contract evaluation. Episode and seed columns make the replication depth explicit. âEp.â is the total number of evaluated episodes included for each row. RÂŻ R is mean episode return; CVaR10%CVaR_10\% is the worst-tail mean return averaged over per-run metrics; Pâ(violDT)P(viol^DT) is downtime-budget violation rate; cÂŻDT c^DT is mean downtime cost; Catast. is the catastrophic episode rate under the alert-threshold proxy. â marks the most conservative safety-contract compliance row in this benchmark; â marks the integrated ACD3-GAT method architecture on the broader safety-contract frontier. Citations in the method column identify prior algorithmic families; C-MAPPO-GAT and ACD3-GAT are introduced here as safety-contract instantiations built from those components. Group Method Ep. S RÂŻ R CVaR PDTP_DT cÂŻDT c_DT Cat. Nonlearn. Sleep 140 3 â6,792-6,792 â8,984-8,984 0.000 0.0 0.000 Random 140 3 â5,149-5,149 â6,737-6,737 1.000 426.1 0.107 Rule-based (kiely2025cage4) 140 3 â8,124-8,124 â10,456-10,456 1.000 115.9 0.000 Actor-critic IA2C (mnih2016asynchronous) 100 1 â4,948-4,948 â6,691-6,691 1.000 420.6 0.070 IPPO (schulman2017proximal) 600 3 â4,254-4,254 â6,299-6,299 1.000 314.4 0.012 MAPPO enc. MAPPO-MLP (yu2022surprising) 200 1 â3,937-3,937 â6,226-6,226 1.000 316.2 0.015 MAPPO-GNN (yu2022surprising; hamilton2017inductive) 100 1 â4,387-4,387 â6,108-6,108 1.000 406.2 0.010 MAPPO-GAT (yu2022surprising; velickovic2018graph) 600 3 â3,979-3,979 â5,864-5,864 1.000 355.4 0.035 Arch. Fact.-IPPO (schulman2017proximal) 100 1 â5,378-5,378 â7,428-7,428 1.000 311.1 0.120 Opp-IPPO (schulman2017proximal; cho2014learning) 100 1 â4,648-4,648 â7,015-7,015 1.000 321.2 0.000 Safety CVaR-MAPPO (yu2022surprising; rockafellar2000optimization) 100 1 â5,131-5,131 â6,509-6,509 1.000 429.6 0.090 C-MAPPO-GAT (yu2022surprising; velickovic2018graph; altman1999constrained) â 600 3 â6,992-6,992 â9,694-9,694 0.003 15.5 0.000 ACD3-GAT (yu2022surprising; velickovic2018graph; altman1999constrained; rockafellar2000optimization) â 600 3 â8,144-8,144 â11,074-11,074 0.138 48.2 0.003 ACD3 comp. ACD3+AskHuman 30 1 â7,901-7,901 â10,460-10,460 0.100 36.8 0.000 ACD3+det. G-CRP 30 1 â7,901-7,901 â10,460-10,460 0.100 36.8 0.000 Component-check rows. The final two ACD3 rows in Table 2 are implementation checks of the governance layer rather than separate policy claims. Both reuse the same trained ACD3-GAT checkpoint for a 30-episode, single-seed check while emphasizing a different screening path: the AskHuman escalation gate or the deterministic G-CRP shield. Under the default confidence setting (Ïconf=0.0 _conf=0.0), the AskHuman gate does not separate the aggregate metrics in this short window, so the two rows coincide. They are included to show that these screening hooks preserve the learned reward and compliance profile under the short component check; attributing performance differences between the hooks would require a longer targeted evaluation. Figure 3: Main empirical story in four focused panels. The colors and markers identify the representative methods shown in the dashboard; the complete benchmark, including IA2C, MAPPO-MLP/GNN, Factorized-IPPO, Opp-IPPO, and CVaR-MAPPO, is reported in Table 2. (a) Reward-only policies violate the downtime budget in nearly every episode; C-MAPPO-GAT is the most reliable safety-contract configuration, while ACD3-GAT is evaluated as the integrated safety-contract architecture. (b) The same methods consume multiple MTTR budgets per episode; the dashed boundary corresponds to âtctdown/Bdown=1 _tc_t^down/B_down=1. (c) The returnâdowntime Pareto surface exposes the operational price of constraint compliance and the current ACD3-GAT frontier position. (d) The gap between mean return and CVaR-10% shows that tail episodes remain a first-class risk even when average reward improves. Read together, Table 2 and Figure 3 establish the central empirical pattern. In the dashboardâs compliance panel, the unconstrained learners cluster at full downtime-budget violation, meaning that the policies cross BdownB_down in every evaluated episode even when their reward is competitive. The adjacent over-budget panel explains the scale of the failure: the same methods consume several multiples of the allowed MTTR budget, whereas C-MAPPO-GAT moves below the budget and ACD3-GAT moves close to it on average. The Pareto panel then shows the price of that movement. Reward-only MAPPO-GAT and random exploration sit high in return but far to the right of the downtime boundary; C-MAPPO-GAT moves into the feasible operational region at a return cost; and ACD3-GAT occupies an intermediate frontier point that reflects the integrated architecture rather than the most conservative compliance setting. The final dashboard panel makes the same argument for tail behaviour, which is that the gap between mean return and CVaR-10% remains large enough that average reward cannot serve as the sole deployment criterion. (a) Downtime-budget violation. (b) Operational cost decomposition. (c) Returnâdowntime frontier. (d) Mean return and CVaR-10%. Figure 4: Safety-contract benchmark decomposed into violation, operational cost, returnâcost frontier, and tail-risk views. (a) A violation is the episode-level event âtctdown>Bdown _tc_t^down>B_down; the dashed line marks a 5% safety-contract target. (b) Downtime (Restore), firewall churn (Block/Allow), and false-positive Restore costs are undiscounted episode totals. (c) The vertical boundary marks the downtime budget Bdown=50B_down=50; points left of the boundary satisfy the mean MTTR budget. (d) CVaR-10% is the mean return of the worst 10% of episodes, showing that reward improvements do not automatically remove worst-tail episodes. Constraint compliance. C-MAPPO-GAT reaches Pâ(violDT)=0.003P(viol^DT)=0.003 across 600 episodes from three seeds. This corresponds to two downtime-budget violations in the replicated core sample, providing empirical compliance evidence under the reported benchmark but not a certified chance constraint. ACD3-GAT reaches Pâ(violDT)=0.138P(viol^DT)=0.138 with mean downtime cost 48.248.2 across three 200-episode seeds: below the episode budget on average, but with less reliable episode-level compliance than C-MAPPO-GAT. Every unconstrained learning method in Table 2 achieves Pâ(violDT)=1.000P(viol^DT)=1.000. The 200-episode IPPO/MAPPO rows still consume 314â355 downtime-cost units per episode, while constrained MAPPO-GAT consumes 15.5, a 95â96% reduction against a budget of 50. ACD3-GAT consumes 48.2, an 84â86% reduction in mean downtime cost relative to the same unconstrained policies, with a 13.8% violation rate. This places the integrated policy on the safety-contract frontier rather than at the most conservative compliance point: it demonstrates the full ACD3 architecture, while C-MAPPO-GATâintroduced here as the constrained MAPPO/GAT/Lagrangian instantiation of the same safety-contract ideaârepresents the strongest compliance configuration in the present benchmark. The cost savings do not come for free: constrained MAPPO-GAT and ACD3-GAT both have lower return than the unconstrained MAPPO variants. The resulting Pareto tradeoff is explicit: operational safety costs security reward, but a policy with Pâ(violDT)=1.000P(viol^DT)=1.000 cannot be considered deployable under the stated SOC contract. Figure 4 decomposes this distinction across violation probability, operational cost, returnâdowntime tradeoff, and tail-risk gap. Panel (a) turns downtime into a governance eventâcrossing the episode budgetârather than a continuous training metric. Panel (c) then shows why the safety result is not simply a lower-score variant of the same policy: the constrained methods move the system into a different operating regime, where lower reward is exchanged for remaining inside or near the MTTR budget. Panel (b) explains that downtime Restore cost is the dominant operational harm separating the methods, while firewall-change and false-positive costs remain important secondary contract dimensions. Panel (d) keeps the interpretation honest: even when a method improves mean return, its worst episodes can remain much worse than its average behaviour. Graph structure and tail behaviour. In the benchmark, MAPPO-MLP at 200 episodes (â3,937-3,937) remains ahead of MAPPO-GAT across 600 episodes (â3,979-3,979), while MAPPO-GAT has the strongest CVaR-10% score among the unconstrained MAPPO encoder variants. The evidence therefore supports a narrower conclusion: graph structure improves tail behaviour, but the raw-return advantage of graph attention has not emerged in the three-seed MAPPO-GAT comparison. Figure 5 is useful here because it shows that the reward-only policies learn quickly toward higher return while remaining operationally unsafe, whereas the constrained curves occupy a lower-return region because the policy update is also paying for downtime, firewall, and false-positive costs. Figure 6 complements this by showing the spread of seed-level returns and marking which rows are replicated enough to support the core comparison. The visual message is therefore not that graph attention dominates every metric, but that graph-aware and constrained variants expose different axes of performance: representation affects tail return, while operational constraints affect deployability. Figure 5: Learning dynamics and final return-risk context for the principal learning methods. (a) Moving-average episode returns over the common 200-episode benchmark window, with shaded seed bands for replicated methods. Non-learning baselines are shown only as muted scale references. (b) Final mean return and CVaR-10% for the same methods. This pairing should be read with the violation figures because reward alone is not the deployment criterion. Figure 6: Seed-consistency audit for learning methods. Each point is one seedâs mean episode return and boxes summarise the across-seed distribution. Single-seed and short component-check methods are explicitly marked, separating replicated comparisons from exploratory component checks. Longer-horizon replication. Table 3 reports the two additional 300-episode replications for MAPPO-GAT, constrained MAPPO-GAT, and ACD3-GAT. The longer horizon preserves the main ordering. Reward-only MAPPO-GAT remains operationally non-deployable: Pâ(violDT)=1.000P(viol^DT)=1.000 and mean downtime cost 357.1, despite strong mean return (RÂŻ=â4,033 R=-4,033). Constrained MAPPO-GAT remains the most reliable safety-contract policy, with Pâ(violDT)=0.007P(viol^DT)=0.007, mean downtime cost 10.1, and zero catastrophic episodes across the two 300-episode replications. ACD3-GAT remains close to the downtime budget on average (cÂŻDT=48.8 c^DT=48.8), with a 14.3% episode-violation rate. The extension strengthens the central safety-contract result and preserves the same frontier interpretation: the integrated method reduces operational harm, while the constrained baseline remains the most conservative compliance configuration in the reported evaluations. Table 3: Longer-horizon replication check for the three principal graph-based learning methods. The table reports two additional 300-episode replications per method, evaluated separately from the balanced three-seed benchmark in Table 2. Method Ep. S RÂŻ R CVaR PDTP_DT cÂŻDT c_DT Cat. MAPPO-GAT 600 2 â4,033-4,033 â5,848-5,848 1.000 357.1 0.063 C-MAPPO-GAT 600 2 â6,962-6,962 â9,594-9,594 0.007 10.1 0.000 ACD3-GAT 600 2 â8,132-8,132 â11,082-11,082 0.143 48.8 0.002 8.3 Operational Discipline and Heuristic Baselines A central finding from the CAGE4 competition (kiely2025cage4) is that engineered heuristic agents outperformed the submitted MARL agents. Our local rule-based baseline is not a reproduction of the winning heuristic. It is a deliberately simple reactive policy that isolates the role of operational discipline. Our rule-based agent achieves RÂŻââ8,124 Râ-8,124 â worse than doing nothing (Sleep: ââ6,792â-6,792) and far below random (ââ5,149â-5,149). The mechanism is clear from the cost breakdown: cÂŻDTâ115 c^DTâ 115 (rule-based) vs 0 (sleep) and 425425 (random). The rule agent triggers Restore on every malicious-process alert, eventually exhausting Bdown=50B_down=50 because the cost proxy counts one unit per Restore action. Each subsequent Restore action is a budget violation, and the accumulated cost overwhelms the security benefit. This motivates constrained learning: naive rules fail because they lack the observation engineering, valid-action filtering, mission-phase traffic discipline, and selective response logic that made the top CAGE4 heuristics effective. Unconstrained RL fails for a complementary reason: it discovers high-impact actions such as Restore but does not internalise the operational budget. The target for any deployable autonomous response policy is therefore not just high security reward, but safety-contract compliance. 8.4 Architectural and Component Analysis Encoder architectures. The MAPPO family covers three encoder architectures. At the reported horizons: âą MLP: RÂŻ=â3,937 R=-3,937, CVaR10%=â6,226CVaR_10\%=-6,226 at 200 episodes. Best raw return; no structural inductive bias. âą GNN (GraphSAGE): RÂŻ=â4,387 R=-4,387, CVaR10%=â6,108CVaR_10\%=-6,108 at 100 episodes. Graph aggregation reduces tail episodes. âą GAT: RÂŻ=â3,979 R=-3,979, CVaR10%=â5,864CVaR_10\%=-5,864 across 600 episodes from three seeds. Best reported MAPPO encoder CVaR-10%; The attention mechanism adds interpretability at modest extra cost. The graph encoders improve the tail metric relative to the MLP run (â6,108-6,108 for GNN and â5,864-5,864 for GAT vs â6,226-6,226 for MLP), while the MLP still leads in mean return. In this benchmark, graph structure contributes primarily to tail-risk behaviour and interpretability rather than to raw-return dominance. Factorized action head. Factorized-IPPO (â5,378-5,378 at 100 episodes) underperforms replicated IPPO (â4,254-4,254 across 600 episodes). The factorized head (type Ă target) has more parameters and slower apparent convergence in this short 100-episode component run, especially with the GAT encoder. However, its catastrophic episode rate (0.120 vs 0.012 for IPPO) is higher, so the factorized variant remains an instability to revisit in longer ablations. Opponent embedding. Opp-IPPO (â4,648-4,648 at 100 episodes) trails replicated IPPO (â4,254-4,254 across 600 episodes) in return but achieves catastrophic rate 0.000 vs 0.012 for IPPO. This is consistent with the GRU opponent embedding acting as a useful regulariser for worst episodes, although the short horizon prevents a replicated architectural conclusion. CVaR reweighting. CVaR-MAPPO (â5,131-5,131, CVaR=â6,509CVaR=-6,509) vs MAPPO-MLP (â3,937-3,937, CVaR=â6,226CVaR=-6,226 at 200 episodes): at this training horizon, episode-level return-tail reweighting by itself is less effective than explicit operational-cost constraints. This supports the paperâs central design choice: tail-risk accounting is a useful component of ACD3, but SOC deployability is driven primarily by the safety contract. ACD3 component checks. The 30-episode AskHuman and deterministic G-CRP component runs both reach Pâ(violDT)=0.100P(viol^DT)=0.100, cÂŻDT=36.8 c^DT=36.8, and RÂŻ=â7,901 R=-7,901. This is safer than the reported ACD3-GAT result on violation rate (0.100 vs 0.138), but the horizon is shorter and the two variant metrics are identical at this summary level. These short-horizon results show that the ACD3 switches can operate within the safety-contract regime early in training. They are component checks, not attribution evidence: because the AskHuman and deterministic G-CRP rows collapse to identical aggregate metrics, they are not used to claim separate causal effects for the two switches. Their role is to verify that the override and deterministic G-CRP paths execute within the same accounting framework as the replicated policies; causal attribution would require per-step override reasons, proposed and executed actions, QsafeQ_safe values, uncertainty estimates, and remaining-budget state at the decision point. Learned G-CRP fitting. Table 4 reports the supervised fitting result for the learned G-CRP model. The model was trained from six longer-horizon trajectory files, using 6,000 sampled transitions from a 45,000-transition valid pool. The loss decreased from 0.2638 at epoch 5 to 0.2606 at epoch 10, indicating that the learned risk-propagation component can be fitted from the collected safety-labelled trajectories. This establishes the supervised risk-estimation path from the collected trajectories. The policy benchmark itself uses the deterministic G-CRP screen, keeping policy performance claims tied to the action-screening rule used during the reported evaluations. Figure 7 should therefore be read as a mechanism figure, not as another benchmark row. Each panel applies a different candidate response action to the same graph belief state and propagates risk with the deterministic cascade in Equation (69). The comparison illustrates how G-CRP changes the action-screening problem: actions are not ranked only by immediate reward or immediate cost, but by the post-action graph-risk field that would be handed back to the policy loop. This is the counterfactual calculation behind QsafeQ_safe in Equation (71). Figure 7: Graph Counterfactual Risk Propagation (G-CRP) under alternative response actions. Each panel applies one counterfactual action to the same graph observation and propagates risk for two steps using the independent-cascade model in Equation (69). This figure illustrates the deterministic rollout-free shield model used by the reported policy runs; learned G-CRP is the supervised risk-estimation extension trained from logged trajectories. Table 4: Learned G-CRP fitting result. The table reports the supervised risk-estimation component trained from safety-labelled trajectories; policy performance in the benchmark is tied to the deterministic G-CRP screen used by the evaluated runs. Quantity Value Source trajectory files 6 Valid transition pool 45,000 Sampled transitions 6,000 Training epochs 12 Loss at epoch 5 0.2638 Loss at epoch 10 0.2606 Temporal Contract Graph Shielding diagnostic. ACD3-TCGS extends the action-screening question from one-step graph risk to future contract risk. The temporal model was trained on 40,000 length-8 sequences sampled from 180 safety-labelled episodes across 9 trajectory files. On the held-out split, it achieves near-perfect discrimination for future downtime-budget violation (AUCdown=0.99998AUC_down=0.99998, Brier score =0.00162=0.00162, calibration error =0.00177=0.00177); firewall-disruption and false-positive Restore risks show similarly high discrimination (AUCfw=0.99827AUC_fw=0.99827 and AUCfp=0.99995AUC_fp=0.99995). At the diagnostic threshold Ï”=0.05Δ=0.05, the fitted model accepts approximately 32.9% of candidate histories, with an observed future downtime-violation rate of 0.11% among accepted samples. We also ran a frozen-policy ACD3-GAT + TCGS screening diagnostic. The completed frozen-policy diagnostic contains 100 episodes. It is included as single-seed policy-level diagnostic evidence. Mean downtime cost is 39.3, with downtime-violation rate 5.0%; false-positive violation is 8.0% and catastrophic rate is 0.0%. The mean return is â7,331-7,331 with CVaR10%=â10,590CVaR_10\%=-10,590. The firewall-change budget remains the active weakness (cÂŻfw=105.2 c^fw=105.2, violation rate 31.0%), so these results are used as temporal-shield evidence rather than as a headline claim that ACD3-TCGS dominates the replicated policies. Figure 8 visualizes this diagnostic frontier. Figure 8: Predictive temporal contract shielding diagnostic for ACD3-TCGS. (a) The policy-level diagnostic compares downtime-budget violation for C-MAPPO-GAT, ACD3-GAT, and the frozen ACD3-GAT policy screened by the temporal contract-risk model. (b) Mean return and CVaR-10% are reported on the same policies so the shield is interpreted as a safetyâutility intervention rather than a violation-only post-processing step. (c) The final panel reports the deployability threshold Ï”=0.05Δ=0.05, override rate, and normalized downtime cost for the temporal shield. 8.5 Tail-Risk and Catastrophic Episodes We define a catastrophic episode as one in which the mean alert level exceeds threshold Ξalert=8 _alert=8, corresponding to coordinated multi-host compromise that threatens mission continuity. Table 2 reports the catastrophic episode rate for all methods. CVaR reweighting alone is insufficient at this horizon. CVaR-MAPPO achieves a catastrophic rate of 0.090 â higher than MAPPO-MLP (0.015 at 200 episodes) and MAPPO-GNN (0.010 at 100 episodes). Its CVaR-10% (â6,509-6,509) is also worse than the reported MAPPO-MLP and MAPPO-GAT rows. This confirms the finding from Section 8.4: at the reported training horizon, the episodic diversity that CVaR needs to reweight has not yet accumulated. Graph structure improves worst-tail return, not every tail proxy. MAPPO-GAT achieves the best CVaR-10% among the reported unconstrained MAPPO encoder rows (â5,864-5,864, catastrophic rate 0.035), outperforming MAPPO-MLP (â6,226-6,226, 0.015) and MAPPO-GNN (â6,108-6,108, 0.010) on worst-tail return. The catastrophic-rate proxy does not improve in the same ordering: MAPPO-GAT has a higher catastrophic rate than MAPPO-MLP and MAPPO-GNN. Therefore, the graph attention improves the worst-tail return metric in this benchmark, but it does not by itself remove catastrophic alert-threshold episodes. Constraints are the most reliable observed mechanism for reducing operational harm. C-MAPPO-GAT achieves catastrophic rate 0.000 and CVaR-10% =â9,694=-9,694. The lower catastrophic rate reflects an important tradeoff: the Lagrangian policy accepts lower mean return in exchange for bounded operational cost, which removes alert-threshold catastrophic episodes in the C-MAPPO-GAT evaluations. The reported ACD3-GAT result achieves catastrophic rate 0.003, but its CVaR-10% is â11,074-11,074, worse than C-MAPPO-GAT. This boundary is informative, because composing graph encoding, CVaR reweighting, Lagrangian constraints, and shielding does not automatically dominate the cleaner constrained baseline. The 30-episode AskHuman and deterministic G-CRP component runs have lower catastrophic rate and lower downtime violation than the reported ACD3 result, but their horizon is too short to identify a stable component effect. Additional ACD3 replications and component ablations are therefore needed to separate whether the penalty comes from override behavior, G-CRP screening, CVaR weighting, or insufficient training horizon. 8.6 Robustness Stress Tests The robustness studies explores a narrower question than the main benchmark. Once a policy has been trained, does the safety contract survive controlled changes in topology seed or adaptive Red-process behaviour? They are interpreted as stress tests rather than as a replacement for the replicated ID benchmark in Table 2. Topology-seed stress. On unseen topology seeds, the safety contrast remains intact. Across 120 evaluation episodes, constrained MAPPO-GAT preserves Pâ(violDT)=0.000P(viol^DT)=0.000, mean downtime cost 14.914.9, and catastrophic rate 0.000. IPPO and MAPPO-GAT remain non-deployable under the same stress: both have Pâ(violDT)=1.000P(viol^DT)=1.000, with mean downtime costs 223.6 and 345.7, respectively. The topology shift does not harm MAPPO-GATâs reward (RÂŻ=â3,517 R=-3,517 vs â3,548-3,548 on the matched ID evaluation). Thefore, the reward robustness alone is insufficient when the operational contract is exhausted in every episode. For constrained MAPPO-GAT, topology stress produces a stable safety profile (RÂŻ=â7,286 R=-7,286, CVaR10%=â9,433CVaR_10\%=-9,433) while keeping the MTTR budget satisfied. ACD3-GAT is analysed on the matched-layout benchmark and adaptive Red-process stress test. The topology-OOD comparison is restricted to policies whose action heads are already layout-stable under the evaluated unseen topology seeds. Coupled adaptive Red process. Figure 9 reports the adaptive Red-process stress test. For each frozen Blue policy, a Red PPO learner is trained for 20 updates and its sampled actions are injected into the CybORG transition loop. The raw-return panel is useful for scale, while the degradation panel compares how much each Blue policy worsens relative to its first Red-process update. MAPPO-GAT has the largest worst degradation, since its Blue return drops by 1,147 points at the worst update and finishes 835 points below its starting value. The Constrained MAPPO-GAT is much more stable, with a worst degradation of 476 points and an end-of-run change of only â8-8 points. ACD3-GAT starts from a lower raw-return level, but its worst degradation is 541 points and the curve recovers to finish 680 points above its first update. This supports a bounded robustness interpretation: explicit safety machinery is associated with lower worst policy-return degradation than reward-only MAPPO-GAT, while constrained MAPPO-GAT remains the most reliable safety-contract policy among the evaluated methods. Figure 9: Coupled adaptive Red-process stress test. Each curve evaluates a frozen Blue policy while a Red PPO learner adapts for 20 updates and sampled Red actions are injected into the CybORG transition loop. (a) Raw Blue-policy return under adaptation. (b) Worst and final change in Blue-policy return relative to the first Red-process update, which compares degradation despite different baseline return scales. The figure is a robustness stress test over one Blue-policy seed per method, not a multi-seed exploitability guarantee. 8.7 Synthesis Across the evaluated methods, reward-only learning is consistently non-deployable under the stated SOC contract: Pâ(violDT)=1.000P(viol^DT)=1.000 for every unconstrained learner, and the replicated IPPO/MAPPO rows use 314â355 downtime-cost units per episode against a budget of 50. Lagrangian safety machinery changes that operating regime. C-MAPPO-GAT reduces downtime cost to 15.5 and downtime-budget violation to 0.3%, while ACD3-GAT reduces mean downtime cost to 48.2 with a 13.8% episode-violation rate. The difference is central to the interpretation, because the C-MAPPO-GAT configuration introduced here is the most reliable safety-contract policy observed in the benchmark, whereas ACD3-GAT is the broader architecture for integrating graph perception, safety contracts, tail-risk accounting, and counterfactual action screening. The experiments also explain why simple alternatives are insufficient. Naive reactive rules can perform worse than doing nothing (RÂŻââ8,124 Râ-8,124 vs. â6,792-6,792 for Sleep) because they apply restoration without the valid-action, observation-processing, and mission-policy discipline that made engineered CAGE 4 heuristics effective. Reward-only MARL fails from the opposite direction, because it learns high-impact security interventions without learning when those interventions exhaust operational budgets. Together, the results answer the central question of the paper: autonomous network-security policies cannot be evaluated on reward alone, and operational safety contracts must be represented and optimized explicitly. Future ACD3-GAT evaluations should extend the component ablations and make the factorised target head layout-stable for topology generalisation. 9 Discussion 9.1 Evidence Across the Experiments The empirical evidence extends beyond a single algorithm comparison. Together, the experiments characterise a difficult safety-contract MARL problem: non-learning policies expose the cost of naive response rules; IPPO and MAPPO variants show that reward-only learning discovers operationally harmful interventions; graph encoders test whether network structure improves representation and tail behaviour; constrained MAPPO-GAT isolates the effect of explicit Lagrangian safety machinery; ACD3-GAT integrates graph perception, budget context, CVaR tail-risk accounting, override signals, and counterfactual action screening; short AskHuman and deterministic G-CRP evaluations examine component behaviour; topology-seed and coupled adaptive Red-process stress tests ask whether the safety contrast survives controlled robustness checks; and the two 300-episode replications test whether the principal ordering persists beyond the balanced 200-episode benchmark. This breadth matters because the contrasts are scientifically informative: graph attention improves the worst-tail return but not raw return dominance; CVaR reweighting alone does not improve the tail at the reported horizon; the integrated ACD3 policy reduces harm while the constrained safety machinery is the component that most consistently changes the operational outcome. 9.2 The Operational Safety Finding The most important result is not which method achieves the highest return, but rather that every unconstrained method fails the safety contract on every episode. Mean downtime costs of 311â430 against a budget of 50 mean that unconstrained agents consume roughly 6â9Ă the allowable MTTR budget per episode. In a real SOC, this would manifest as hundreds of unnecessary host reimages per shift, violation of SLA commitments, and systematic availability degradation of the very assets being protected. Constrained MAPPO-GAT reduces the mean downtime cost to 15.5 (<Bdown<B_down) and the violation rate to 0.3% across three 200-episode seeds. ACD3-GAT also brings mean downtime cost below budget (48.2), with a 13.8% episode-violation rate, so the integrated architecture occupies a different point in the design space: it demonstrates the full safety-contract stack, whereas C-MAPPO-GAT is the new controlled constrained configuration that most reliably satisfies the SOC contract in the present benchmark. The two additional 300-episode replications strengthen rather than soften this interpretation, since constrained MAPPO-GAT remains near-zero on downtime violations (0.7%), reward-only MAPPO-GAT remains at 100% violation, and ACD3-GAT remains close to budget on average while still violating in 14.3% of episodes. C-MAPPO-GAT incurs a return penalty of approximately 3,050 points relative to the highest-return unconstrained learner, MAPPO-MLP (â6,992-6,992 vs. â3,937-3,937). ACD3-GAT incurs approximately 4,200 points (â8,144-8,144 vs. â3,937-3,937), consistent with the broader set of safety signals active in the integrated safety-contract stack. We argue that this tradeoff is not only acceptable but necessary: a policy that achieves high security reward while violating operational budgets is not deployable, regardless of its average performance. 9.3 Role of ACD3-GAT ACD3-GAT functions as the integrated method architecture, while C-MAPPO-GAT is the strongest compliance row in the current benchmark. Its role is methodological, since it specifies a general safety-contract architecture with graph-structured perception over network entities, Lagrangian cost learning, explicit budget context, counterfactual action screening, tail-risk accounting, and override signals. C-MAPPO-GAT is the strongest safety-contract configuration because it isolates the new combination of MAPPO, a GAT encoder, and Lagrangian operational-cost control, satisfying the downtime contract most reliably in the reported evaluations. ACD3-GAT is the extensible architecture: it unifies the reusable components needed when a response policy must reason over changing topology, opposing-process adaptation, uncertain action consequences, and human-governed operational budgets. The 13.8% downtime-violation rate therefore identifies the key stabilization target for the integrated policy while preserving the broader method contribution. 9.4 Failure Mechanism of Unconstrained MARL The mechanism is clear from the cost breakdown. Unconstrained MAPPO-MLP issues 316 downtime-cost units per episode (vs. budget 50); IPPO issues 314 and MAPPO-GAT issues 355. The rule-based heuristic, despite being the lowest-returning baseline, incurs 117 downtime-cost units because it applies Restore selectively. Random exploration averages 424 downtime-cost unitsâworse than the constrained agent because it occasionally selects Restore by chance. This refines the intuition from the published CAGE 4 analyses: the competition heuristics succeeded because they encoded observation handling, invalid-action avoidance, mission-phase firewall discipline, and selective response policies. Our naive rule baseline lacks that structure and becomes operationally harmful. Unconstrained RL fails from the other side, because it finds Restore and BlockTraffic attractive because they genuinely reduce the CAGE Red processâs presence in the short term, but the Lagrangian penalties in our framework counterbalance that exploitable pattern with explicit operational budgets. 9.5 Graph Encoder Evidence In the benchmark, MAPPO-MLP at 200 episodes (â3,937-3,937) slightly outperforms MAPPO-GAT across 600 episodes (â3,979-3,979), while MAPPO-GNN remains available at 100 episodes (â4,387-4,387). This is consistent with slower convergence of graph encoders due to their larger parameter count and the need to learn structural attention weights. These data support a specific graph-encoder conclusion rather than a broad raw-return dominance conclusion. Critically, MAPPO-GAT achieves the best CVaR-10% among the reported unconstrained MAPPO encoder rows (â5,864-5,864), and MAPPO-GNN also improves the tail relative to MAPPO-MLP, suggesting that graph structure is presently a tail-risk and interpretability result rather than a mean-return result. 9.6 Scope of Evidence and Transfer Claim boundary. This study is intentionally scoped to safety-contract evaluation of autonomous network-security response in CAGE-4. The primary inferential target is the contrast between reward-only MARL and explicit operational-cost control, not return-based ranking, broad topology generalisation, or factorial attribution of every ACD3 component. Three replicated seeds support stability of the safety contrast for the core methods, and rare violation rates are reported with raw episode counts rather than as formal chance-constraint certificates. Robustness results are controlled stress tests under the same cost accounting. Larger seed counts, hierarchical task decomposition, layout-stable topology evaluation for ACD3-GAT, and head-to-head comparison with recent graph-RL and hierarchical MARL systems under a shared safety logger are complementary follow-on work rather than prerequisites for the central deployability finding. Simulator-grounded evidence. The study is intentionally grounded in CybORG/CAGE-4, where actions, observations, CAGE Red-process behaviour, and mission scoring are controlled and reproducible. The cost proxies (ctdown,ctfw,ctfpc_t^down,c_t^fw,c_t^fp) map the simulator primitives onto SOC-relevant governance quantities: service recovery burden, firewall-change burden, and false-positive response burden. This makes the safety contract auditable inside the benchmark and gives practitioners a clear template for replacing the proxies with organisation- specific measurements in a live range, SOAR platform, or enterprise digital twin. Robustness and attribution scope. The replicated core comparison covers IPPO, MAPPO-GAT, and constrained MAPPO-GAT, and ACD3-GAT at three 200-episode seeds. Two additional 300-episode replications cover MAPPO-GAT, constrained MAPPO-GAT, and ACD3-GAT, and are reported as a longer-horizon replication check rather than as a replacement for the balanced core table. The topology-seed stress test covers IPPO, MAPPO-GAT, and constrained MAPPO-GAT; the coupled adaptive Red-process stress test covers MAPPO-GAT, constrained MAPPO-GAT, and ACD3-GAT. These choices define the evidence that a replicated safety-contract comparison, a longer-horizon check of the principal ordering, and controlled robustness probes. The 30-episode AskHuman and deterministic G-CRP component rows are therefore treated only as execution-path checks; their identical aggregate summaries do not identify which switch caused the observed behaviour. The next attribution layer is naturally factorial: C-MAPPO-MLP, budget-context only, hard-shield only, Lagrangian-only, CVaR-only, opponent-context, and G-CRP-screening variants can be used to decompose the integrated ACD3 policy into its causal components. That decomposition is an extension of the present benchmark rather than a prerequisite for the core result that explicit safety machinery changes the operational regime. Temporal and operational transfer. The 500-step, Îł=0.99Îł=0.99 formulation provides a controlled horizon for comparing policies under identical mission dynamics. Longer incident campaigns, analyst workflows, ticketing delays, and organisation-specific change-control rules can be incorporated by changing the cost functions and episode horizon while preserving the same constrained Dec-POMDP and safety-contract machinery. ACD3-TCGS is the first step in that direction, where rather than screening only the immediate proxy cost or a two-hop deterministic graph cascade, it learns from trajectory histories to predict future budget exhaustion before the action is executed. The current evidence is frozen-policy temporal-shield evidence, not an end-to-end retrained TCGS policy result. That boundary is important, but the result is still informative. It shows that the safety-labelled trajectories produced by the benchmark contain enough temporal structure to support predictive contract-risk screening. Proxy cost signals. The false-positive Restore proxy (Restore with no active alerts) is defined from visible alert evidence because the Blue response agents act under partial observability. This choice matches the operational question faced by SOC automation: whether an action was justified by available evidence at decision time. In a deployment audit, the same definition can be paired with ground-truth incident labels when post-incident forensics are available. Budget and chance-constraint sensitivity. The Lagrangian update controls expected episode cost, not violation probability directly. The empirical violation rate is therefore a measured outcome rather than a formal chance-constraint guarantee. Changing BdownB_down, BfwB_fw, or BfpB_fp would change both the feasible policy set and the shield activation pattern. The CAGE-4 budgets used here are therefore best read as reproducible stress-test thresholds; operational adoption would instantiate the same method with locally chosen service, change-management, and analyst-capacity budgets. Interpretation boundary. The safety contract is an evaluation and training discipline inside CAGE-4: it exposes when a policy exhausts MTTR, firewall-change, or false-positive budgets and provides the decision record needed for audit. Graph attention is reported as topology-aware representation learning, and G-CRP is reported as the deterministic counterfactual risk screen used by the shield. The learned G-CRP model is included in the replication package as the next risk-estimation component in the same framework; the reported policy performance is tied to the deterministic screen unless explicitly stated otherwise. Prior-work references. The published C4 heuristic and LLM scores are important context, but they are not direct safety-contract baselines. They use different reported outputs and do not expose the per-action Restore, Block/Allow, and false-positive traces required by our cost metrics. We therefore use them to calibrate the reward scale, while the main claims are restricted to methods evaluated with the same safety-labelled protocol. 9.7 Implications for Autonomous SOC Deployment Read as an expert-system result, the framework provides a deployable decision-support pattern rather than a claim of autonomous production readiness, because we encode SOC governance as budgets, train and evaluate policies against those budgets, log every executed action and cost proxy, and route budget exhaustion or low-confidence states to human governance. Our results suggest three practical principles for deploying such agents: 1. Define operational budgets before training. Unconstrained training consistently produces operationally harmful policies. Budget values (BkB_k) should be set by SOC governance and treated as explicit violation constraints with audit metrics, not as hidden soft reward preferences. The CAGE-4 values used here are stress-test thresholds; a SOC deployment should tune them from service-level objectives, change-management policy, and analyst capacity, then rerun the same safety-labelled evaluation. 2. Validate robustness after satisfying the safety contract. Fixed scripted Red processes can overestimate policy robustness, so adaptive opposing-policy and topology-shift studies should follow the ID safety benchmark before any deployment claim. In the reported stress tests, constrained MAPPO-GAT preserves zero downtime violations under topology-seed shift and degrades less than reward-only MAPPO-GAT under the coupled adaptive Red-process stress test. 3. Support human override. The ACD3-GAT override mechanism (triggered by low confidence, budget exhaustion, or OOD states) provides a governance layer that maintains human authority over high-stakes decisions. AskHuman variants measure the cost of that governance layer and connect the learned policy to a SOC approval workflow. 10 Conclusion Autonomous network security agents trained without explicit operational constraints systematically violate SOC budgets in the evaluated CAGE 4 setting, creating downtime costs roughly 6â9Ă the allowable MTTR budget. This is not merely a reward-engineering failure; it reflects a structural mismatch between scalar security reward and operational deployability. ACD3-GAT addresses this through a unified safety-contract design: Lagrangian-constrained multi-agent PPO that optimizes under MTTR, false-positive, and firewall-policy budgets; graph encoders for structured hostâsubnet observations; a factorised typeâtarget action policy; a multi-objective training objective for reward, operational cost, and tail-risk accounting; and a budget-aware counterfactual shield. On CAGE Challenge 4, the C-MAPPO-GAT constrained configuration introduced in this work achieves a 95â96% reduction in operational harm while remaining within budget 99.7% of episodesâa qualitative shift from policy that is operationally harmful to one that is operationally auditable. Two additional 300-episode replications preserve the same ordering: constrained MAPPO-GAT remains near-zero on downtime violations, while reward-only MAPPO-GAT continues to violate the downtime budget in every episode. ACD3-GAT reduces mean downtime harm but still violates the budget in 13.8% of episodes. This identifies a clear division between contribution types: C-MAPPO-GAT is the strongest compliance configuration in the reported benchmark and demonstrates that the MAPPO/GAT/Lagrangian composition is a powerful safety-contract baseline, while ACD3-GAT is the general architecture for combining graph perception, safety contracts, tail-risk accounting, opponent context, and counterfactual action screening. Topology-seed and coupled adaptive Red-process stress tests reinforce the same conclusion, which is that explicit safety machinery, not reward alone, determines whether an autonomous response policy remains operationally auditable under stress. The core finding has direct implications for SOC governance: autonomous network-security agents must carry explicit operational constraints before they can be considered for deployment. More broadly, ACD3-GAT treats network security response as one instance of a wider class of safety-contract MARL problems. For instance, multiple agents act over structured entities, local actions have operational side effects, and deployment requires respecting budgets that are not reducible to reward. The anonymised replication package provides safety-labelled trajectories, trained checkpoints, evaluation metadata, and robustness protocols so that the reported claims can be independently regenerated and extended. The ACD3-TCGS diagnostic shows how those trajectories can also train a temporal contract-risk shield that predicts future budget exhaustion before executing a proposed action. The natural next research agenda is to instantiate the same safety-contract framework at larger scale, use layout-stable target heads for broader topology transfer, evaluate learned Red-process variants in the main robustness loop, integrate the learned G-CRP risk model into policy evaluation, and pair the safety contract with temporal graph memory or asynchronous cyber-range simulators when moving from CAGE-style episodes to richer SOC telemetry (rossi2020temporal). Continuous-time cyber-range evaluation provides the complementary simulator direction (jankowski2026netforge). CRediT Author Contribution Statement Jose Luis Silva: The author was responsible for the conceptualization and methodology of the study, developed the software, conducted the formal analysis and investigation, curated the data, prepared the original draft, reviewed and edited the manuscript, and produced the visualizations. Declaration of Competing Interest The authors declare no competing financial or personal interests. Data Availability Software, safety-labelled trajectories, trained model checkpoints, and robustness-evaluation metadata are available from the corresponding author upon request. References