Paper deep dive
Machine-Coached Policy Revision in Adaptive Agent-Based Regulatory Simulation: A Controller-Level Contestability Layer
Roberto Garrone
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 87%
Last extracted: 7/9/2026, 5:58:36 AM
Summary
This paper introduces a machine-coached policy-revision layer for adaptive agent-based models (ABMs) to bridge the gap between ex post diagnostics and controller-level contestability. It proposes a symbolic controller that encodes policy decisions as defeasible rules with explicit conflicts and priorities. Diagnostic failures are systematically translated into auditable rule revisions (additions, removals, or priority changes), which are then re-evaluated on held-out simulation seeds. The framework is validated using a stylized emissions-regulation ABM, demonstrating how a predefined coaching template adds a relaxation rule to mitigate an over-conservatism failure in the VPVA regime while preserving regulatory guardrails.
Entities (8)
Relation Signals (5)
Over-Conservatism Failure → addressedby → Relaxation Rule
confidence 90% · The predefined coaching template adds a relaxation rule to the symbolic controller, reducing over-conservatism recurrence
Machine Coaching → extends → Explainable Adaptive ABM
confidence 90% · The paper argues that machine coaching is best understood as a controller-level extension of explainable adaptive ABM
Diagnostic Failures → translatedinto → Rule Revisions
confidence 90% · allows diagnostic failures to be translated into rule additions, removals, or priority changes
Symbolic Controller → represents → Policy Decisions
confidence 85% · The layer represents policy decisions as defeasible rules with explicit conflicts and priorities
Held-Out Simulation Seeds → usedfor → Policy Re-evaluation
confidence 85% · policy decisions can be explained, challenged, revised, and re-evaluated in held-out simulation runs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Policy-oriented agent-based models are increasingly used to study regulatory interventions in complex adaptive socio-technical systems. Recent adaptive ABM frameworks distinguish between static and adaptive agents, fixed and adaptive policies, and alternative controller designs. However, most diagnostic workflows remain ex post: trajectories are analysed after simulation, but the resulting evidence is not systematically fed back into the policy controller. This paper proposes a lightweight machine-coached policy-revision layer for adaptive agent-based regulation. The layer represents policy decisions as defeasible rules with explicit conflicts and priorities, generates explanations for controller actions, and allows diagnostic failures to be translated into rule additions, removals, or priority changes. The contribution is not a new optimal controller and does not claim formal guarantees for unrestricted machine coaching. Instead, it provides a simulation-compatible operationalization of controller-level contestability: policy decisions can be explained, challenged, revised, and re-evaluated in held-out simulation runs. A stylized emissions-regulation ABM is used as the experimental component. A controlled simulation experiment focuses on an over-conservatism failure in the VPVA regime. The predefined coaching template adds a relaxation rule to the symbolic controller, reducing over-conservatism recurrence under held-out seeds while preserving violation, overshoot, and volatility guardrails. The paper argues that machine coaching is best understood as a controller-level extension of explainable adaptive ABM, complementary to causal, information-theoretic, and trajectory-based diagnostics.
Tags
Links
- Source: https://arxiv.org/abs/2606.20700v1
- Canonical: https://arxiv.org/abs/2606.20700v1
Trouble viewing inline? Open PDF directly →
Full Text
70,260 characters extracted from source content.
Expand or collapse full text
Machine-Coached Policy Revision in Adaptive Agent-Based Regulatory Simulation: A Controller-Level Contestability Layer Roberto Garrone Open University of Cyprus roberto.garrone@st.ouc.ac.cy Abstract Policy-oriented agent-based models are increasingly used to study regulatory interventions in complex adaptive socio-technical systems. Recent adaptive ABM frameworks distinguish between static and adaptive agents, fixed and adaptive policies, and alternative controller designs. However, most diagnostic workflows remain ex post: trajectories are analysed after simulation, but the resulting evidence is not systematically fed back into the policy controller. This paper proposes a lightweight machine-coached policy-revision layer for adaptive agent-based regulation. The layer represents policy decisions as defeasible rules with explicit conflicts and priorities, generates explanations for controller actions, and allows diagnostic failures to be translated into rule additions, removals, or priority changes. The contribution is not a new optimal controller and does not claim formal guarantees for unrestricted machine coaching. Instead, it provides a simulation-compatible operationalization of controller-level contestability: policy decisions can be explained, challenged, revised, and re-evaluated in held-out simulation runs. A stylized emissions-regulation ABM is used as the experimental component. A controlled simulation experiment focuses on an over-conservatism failure in the VPVA regime. The predefined coaching template adds a relaxation rule to the symbolic controller, reducing over-conservatism recurrence under held-out seeds while preserving violation, overshoot, and volatility guardrails. The paper argues that machine coaching is best understood as a controller-level extension of explainable adaptive ABM, complementary to causal, information-theoretic, and trajectory-based diagnostics. Keywords: agent-based modeling; adaptive policy; machine coaching; explainable AI; contestability; policy revision; regulatory simulation; symbolic controller; adaptive multi-agent systems. 1 Introduction Agent-based models (ABMs) are widely used to analyse policy interventions in complex socio-technical systems [5, 6, 9, 3, 12]. Their methodological value lies in representing heterogeneous agents, bounded rationality, local interaction, feedback, and emergent macro-level outcomes [1, 5, 13]. This makes ABMs especially relevant in domains where aggregate policy effects cannot be reduced to a representative actor or a static response function [6, 9]. However, many policy-oriented ABMs still treat regulation as a fixed scenario parameter. A policy value is selected, the model is executed, and outcomes are compared across scenarios. This design is useful for counterfactual exploration, but it does not fully represent the adaptive structure of real regulatory systems [7, 8]. Regulators revise instruments in response to observed outcomes, firms or households adapt to incentives, and the resulting feedback changes the system being regulated [12, 2, 14]. Recent adaptive ABM work addresses part of this limitation by distinguishing four regimes: constant policy with constant agents (CPCA), constant policy with adaptive agents (CPVA), adaptive policy with constant agents (VPCA), and adaptive policy with adaptive agents (VPVA) [7, 8]. This taxonomy makes it possible to separate the effects of agent adaptation, policy adaptation, and their interaction. It also supports regime diagnostics based on scalar indicators, symbolic transition patterns, information-theoretic measures, and emergent trajectory clusters [4, 11, 8]. Yet a gap remains. Diagnostics are usually computed after simulation. They help analysts interpret whether a controller violates a cap, behaves conservatively, oscillates near a boundary, or generates excessive volatility [8]. But the diagnostic evidence is not normally incorporated into the policy controller itself. In other words, the model is diagnosable, but not yet revisable in a transparent and contestable way [7, 16]. This paper proposes a controller-level machine-coaching layer for adaptive ABM. The key idea is simple: Diagnostic evidence should not only describe policy-controller failure; it should also provide structured grounds for revising the policy logic that produced the failure. The proposed layer represents policy decisions through explicit symbolic rules, conflict relations, and priorities [10, 18, 21]. When the controller selects an action, it also produces an explanation: which facts were observed, which rules fired, which rules were blocked, and which priority relation resolved the conflict [16]. A coaching mechanism then translates diagnosed failures into rule-level revisions. These revisions may add a new exception, remove an invalid rule, or change a priority relation [19, 21]. Machine coaching is not presented as a substitute for causal analysis, sensitivity analysis, information-theoretic diagnostics, or spatial decomposition [4, 11, 16]. Instead, it is positioned as a specific controller-level layer for explaining and revising policy logic. The contribution is threefold. (i) It defines a lightweight rule-based policy controller for adaptive ABM regulation, with explicit rules, conflicts, priorities, and explanation logs. (i) It introduces a template-constrained machine-coaching protocol that converts predefined diagnostic failures into auditable rule-base revisions. (i) It implements and evaluates the protocol in a controlled emissions-regulation ABM, showing how a diagnosed over-conservatism failure can be translated into a relaxation rule and re-evaluated on held-out simulation seeds. This paper is methodological and derived from a broader PhD research programme on adaptive, explainable, and contestable ABM-based policy design [7]. 2 Background and Motivation 2.1 Adaptive ABM, policy regimes, and ex post diagnostics Let st∈Ss_t∈ S denote the system state at time t, XtX_t the vector of agent actions, PtP_t a policy or control parameter, and ζt _t an exogenous disturbance. A generic ABM transition can be written as st+1=F(st,Xt,Pt,ζt).s_t+1=F(s_t,X_t,P_t, _t). (1) Agents may be static or adaptive. Policy may be fixed or adaptive. Combining these two distinctions gives four dynamic regimes [7, 8]: • CPCA: constant policy, constant agents; • CPVA: constant policy, variable/adaptive agents; • VPCA: variable/adaptive policy, constant agents; • VPVA: variable/adaptive policy, variable/adaptive agents. This decomposition is useful because “adaptiveness” is not a single treatment. Agent adaptation alone may reduce emissions without ensuring compliance. Policy adaptation may track a constraint while producing frequent violations. Adaptive policy and adaptive agents may interact in ways that generate volatility, conservatism, or oscillatory boundary crossing [8]. Over a finite horizon, policy performance may be summarized by a bounded functional J(P;L)=1K∑t=T−K+1TΦ(st),J(P;L)= 1K _t=T-K+1^T (s_t), (2) where L denotes the agent learning or adaptation rule, Φ is a scalar evaluation function, and K is the evaluation window. This formulation does not require convergence to equilibrium and remains meaningful under cyclic, drifting, or volatile dynamics [2, 14, 8]. Adaptive ABM studies often compute scalar and structural diagnostics after simulation. Examples include: mean outcome; violation rate; overshoot; tracking error; policy volatility; transition entropy; motif diversity; conservatism gap; symbolic boundary-crossing patterns [4, 11, 8]. These diagnostics can reveal that two controllers with similar mean performance behave very differently around a policy constraint. For example, one controller may track a regulatory cap closely but cross it frequently; another may avoid violations but remain far below the cap; a third may oscillate between excessive pressure and relaxation [8]. However, ex post diagnosis alone does not revise the controller. The analyst may observe that a controller is over-conservative, but the model does not explain which part of the policy logic generated that behavior, nor does it provide a reproducible way to change that logic. This creates a gap between interpretability and contestability [16, 7]. 2.2 Related work The term machine coaching is used here in a restricted and operational sense, following Loizos’ formulation of machine coaching as an interaction paradigm in which humans and machines externalize their reasoning in mutually understandable terms [23]. The proposed layer is not intended to cover every form of human–machine interaction, nor does it claim the full generality of open-ended natural-language coaching. Instead, it denotes a controller-level revision protocol in which diagnostic failures are linked to the explicit reasons that produced a policy action, and those reasons are revised through structured operations on the controller’s symbolic rule base [10, 19, 21]. This restricted use is also consistent with later work on proxy coaches, where coaching is operationalized as an iterative exchange of explanations aimed at improving both machine conclusions and their acceptability to the coach [24]. In the present paper, the proxy coach is specialized as a predefined advice-template function that maps diagnostic failures to rule-base revisions. First, the adopted approach differs from interactive machine learning and human-in-the-loop machine learning. In interactive machine learning, users typically provide labels, corrections, rankings, feature feedback, or demonstrations that help a model improve its predictive or decision performance [15]. In the present framework, the feedback object is not a label or a preferred trajectory. The feedback object is a defect in the controller’s policy reasoning: for example, a rule was too broad, an exception was missing, or a priority relation produced an unacceptable regulatory response. The coach therefore acts on the explanation structure of the controller rather than directly on the simulated outcome. Second, it differs from reinforcement learning from human feedback or human preferences. Preference-based reinforcement learning uses human judgments, often over trajectory segments, to infer or shape a reward model [17]. The present controller does not learn a latent reward function from preference comparisons. It revises an explicit rule base whose conclusions, conflicts, and priorities remain inspectable before and after revision. The objective is not to produce an optimal policy by reward learning, but to make a policy controller contestable and revisable. Third, the approach differs from post-hoc explainable AI. Post-hoc XAI methods usually explain a model after it has produced a prediction or decision, often without changing the model itself [16]. In contrast, the proposed layer treats explanation as part of an adaptive revision loop. The explanation is not only an interpretive artefact; it identifies the rule, conflict, or priority relation that may be challenged and modified. Fourth, the approach differs from algorithmic recourse. Recourse methods ask what changes would allow an affected subject to obtain a more favourable decision, often through counterfactual or intervention-oriented recommendations [22, 20]. The present framework does not primarily advise an agent how to change its state to obtain a different outcome. It revises the policy controller that generates regulatory actions. The object of intervention is therefore the decision logic of the regulator, not the feature vector of the regulated agent. Fifth, the approach is related to, but narrower than, argumentation-based reasoning and defeasible logic. Abstract argumentation studies how conflicting arguments attack one another and how acceptable argument sets can be defined [18]. Defeasible logic similarly represents rules whose conclusions may be overridden by stronger rules or priority relations [21]. The proposed implementation borrows this intuition through explicit conflicts and priorities, but it does not attempt to implement a full argumentation semantics. The design choice is pragmatic: the goal is to support reproducible policy-controller revision in simulation, not to contribute a new formal argumentation framework. Finally, the approach differs from classical expert systems and rule-based controllers. Expert systems encode domain knowledge as rules, and rule-based controllers use conditional logic to select actions [19]. A machine-coached controller adds an additional requirement: the rule base is not fixed after design. It is revised after diagnosed failures, and the revision is recorded as an explicit operation such as adding a rule, removing a rule, or changing a priority. Thus, the novelty is not the use of rules alone, but the closed loop: diagnostic failure → explanation of policy action → structured coaching advice → rule-base revision → held-out re-evaluation. For this reason, the term machine coaching is used to describe a constrained form of explanation-driven rule revision. The claim is not that the system performs unrestricted human coaching. The claim is that adaptive ABM policy controllers can be made more contestable when diagnostic evidence is connected to explicit, auditable revisions of the controller’s policy logic [7, 8, 16]. 2.3 Machine coaching as controller-level contestability Within this paper, machine coaching is defined as a constrained revision protocol for an explainable policy controller. A coach may be a human expert, a stakeholder-facing analyst, or a scripted diagnostic module. In all cases, coaching is restricted to a small set of auditable operations on the controller’s symbolic rule base [19, 21]. The coach does not provide unrestricted natural-language advice. Instead, coaching is represented as one of three revision operations: (a) add a rule; (b) remove a rule; (c) change a rule priority. This restriction is intentional. It prevents the coaching layer from becoming an informal natural-language interface and preserves reproducibility. It also distinguishes the approach from ad hoc controller tuning. A revision counts as machine coaching in the present framework only if four conditions hold: (i) a diagnostic failure is identified by a predefined criterion; (i) the controller produces an explanation identifying the facts, rules, conflicts, and priorities involved in the policy decision; (i) the coaching operation modifies the explicit rule base rather than directly editing the numerical outcome; (iv) the revised controller is re-evaluated on held-out seeds or cloned initial states. Thus, coaching is not defined by the mere presence of rules. It is defined by the link between explanation, contestation, revision, and re-evaluation [16]. A rule-based controller without revision is simply a symbolic controller [19]. A manually retuned controller without explanation logs is controller engineering. A machine-coached controller, as understood here, is a symbolic controller whose reasoning can be challenged and revised through explicit, logged operations after diagnosed failure modes [21, 18]. 3 Stylized Emissions-Regulation Experimental Environment 3.1 Model overview The empirical component of the article uses a stylized emissions-regulation ABM. It is not intended to forecast a real industrial sector. Its purpose is to provide a controlled environment in which policy-controller logic can be diagnosed and revised. There are N firms. At time t, firm i produces emissions ei,te_i,t. Aggregate emissions are Et=∑i=1Nei,t.E_t= _i=1^Ne_i,t. (3) A regulator imposes an emissions cap C: Et≤C.E_t≤ C. (4) Regulatory pressure is represented by a scalar policy signal PtP_t, interpretable as a carbon tax, inspection pressure, compliance signal, or equivalent regulatory intensity. A simple firm-level emissions function is: ei,t=max0,bi−ηi,tPt+ϵi,t,e_i,t= \0,\ b_i- _i,tP_t+ _i,t \, (5) where bib_i is baseline emissions, ηi,t _i,t is policy responsiveness, and ϵi,t _i,t is stochastic noise. Agent adaptation can be represented through an update of ηi,t _i,t: ηi,t+1=ηi,t+αig(Pt,Et,C)+νi,t, _i,t+1= _i,t+ _ig(P_t,E_t,C)+ _i,t, (6) where αi _i is an adaptation rate, g(⋅)g(·) is a bounded response function, and νi,t _i,t is an idiosyncratic disturbance. 3.2 Numeric adaptive controllers The baseline adaptive controllers are numeric. They update PtP_t directly from observed emissions. A setpoint controller may be written as Pt+1=[Pt+κ(Et−C)]PminPmax,P_t+1= [P_t+κ(E_t-C) ]_P_ ^P_ , (7) where κ is a gain parameter and [⋅]PminPmax[·]_P_ ^P_ denotes clipping. A safety-margin controller targets a stricter internal cap C−mC-m: Pt+1=[Pt+κ(Et−(C−m))]PminPmax.P_t+1= [P_t+κ(E_t-(C-m)) ]_P_ ^P_ . (8) A one-sided controller increases pressure only when emissions exceed the cap: Pt+1=[Pt+κ(Et−C)]PminPmax,Et>C,Pt,Et≤C.P_t+1= cases [P_t+κ(E_t-C) ]_P_ ^P_ ,&E_t>C,\\ P_t,&E_t≤ C. cases (9) The one-sided controller is useful because it can reduce violations, but it may become over-conservative if it never relaxes pressure after emissions remain safely below the cap. 4 Symbolic Policy Controller 4.1 Predicate abstraction The symbolic controller does not consume all raw simulation variables. Instead, a predicate encoder maps numerical state variables into interpretable facts. Let ℱt=A(st)F_t=A(s_t) (10) be the set of facts extracted from state sts_t by abstraction function A. Examples include: • emissions_above_cap; • emissions_near_cap; • emissions_safely_below_cap; • violations_persistent; • policy_volatility_high; • conservatism_gap_high; • overshoot_high. For example: emissions_above_cap∈ℱt⟺Et>C. emissions\_above\_cap _t E_t>C. (11) Similarly, emissions_safely_below_cap∈ℱt⟺Et<C−δC, emissions\_safely\_below\_cap _t E_t<C- _C, (12) where δC _C is a safety threshold. 4.2 Rules, conflicts, and priorities A policy rule is defined as follows. Definition 1 (Policy rule) A policy rule is a tuple r=(name,B,h,π),r=(name,B,h,π), (13) where namename is a rule identifier, B is a finite set of body predicates, h is a head or conclusion, and π∈ℤπ is a priority score. A rule r is applicable at time t if Br⊆ℱt.B_r _t. (14) Rules may support conflicting conclusions. For example: increase_pressure⊥relax_pressure. increase\_pressure\; \; relax\_pressure. (15) Let C denote the conflict relation among rule heads. A selected rule set must be conflict-consistent according to C and the priority relation. A generic symbolic rule base may include rules of the following form: ⬇ r1: IF emissions_above_cap THEN increase_pressure PRIORITY 1 r2: IF emissions_near_cap AND violations_persistent THEN increase_pressure PRIORITY 2 r3: IF policy_volatility_high THEN smooth_policy PRIORITY 3 r4: IF emissions_safely_below_cap AND conservatism_gap_high THEN relax_pressure PRIORITY 2 r5: IF emissions_above_cap AND overshoot_high THEN increase_pressure_strongly PRIORITY 4 The list above should be read as the controller’s admissible rule vocabulary rather than as the exact initial rule base used in the controlled experiment. In the experiment reported below, the initial symbolic controller deliberately excludes the relaxation rule. This omission creates a transparent over-conservatism defect: when emissions are safely below the cap and the conservatism gap is high, the controller has no exception allowing it to relax pressure. The machine-coaching layer is then tested on whether it can diagnose and repair precisely this missing policy-reasoning exception. 4.3 Decision procedure At each time step, the controller performs four operations: (i) encode the state into facts; (i) identify applicable rules; (i) select a conflict-consistent subset using priorities; (iv) translate selected conclusions into a policy action. Let ℛR be the full rule base and ℛtappR_t^app the applicable rule set at time t: ℛtapp=r∈ℛ:Br⊆ℱt.R_t^app=\r :B_r _t\. (16) The selected rule set ℛtselR_t^sel is obtained by priority-ordered conflict resolution: ℛtsel=Resolve(ℛtapp,,π).R_t^sel=Resolve(R_t^app,C,π). (17) The policy action is then at=H(ℛtsel),a_t=H(R_t^sel), (18) where H maps selected rule heads to numerical policy updates. For example: Pt+1=Pt+Δ+,at=increase_pressure,Pt+Δ++,at=increase_pressure_strongly,Pt−Δ−,at=relax_pressure,(1−λ)Pt+λPt−1,at=smooth_policy.P_t+1= casesP_t+ ^+,&a_t= increase\_pressure,\\ P_t+ ^++,&a_t= increase\_pressure\_strongly,\\ P_t- ^-,&a_t= relax\_pressure,\\ (1-λ)P_t+λ P_t-1,&a_t= smooth\_policy. cases (19) 4.4 Priority-based rule resolution The operator Resolve(⋅)Resolve(·) is implemented as a deterministic priority-based rule-resolution procedure. This choice is deliberately simple: the purpose is not to implement a full argumentation semantics, but to obtain a reproducible and auditable controller decision from a finite rule base. Each rule r has a priority π(r)π(r) and a body BrB_r. Priorities are treated as integer scores. Higher values indicate stronger rules. Priorities are static during a simulation episode and can change only through an explicit coaching operation. To avoid nondeterminism, ties are resolved by a fixed ordering key: key(r)=(π(r),|Br|,−index(r))key(r)= (π(r),\ |B_r|,\ -index(r) ) (20) where |Br||B_r| is rule specificity and index(r)index(r) is the insertion order of the rule in the rule base. Rules are sorted in descending order of π(r)π(r), then descending order of specificity, then ascending insertion order. Thus, if two conflicting rules have equal priority, the more specific rule is selected; if they are equally specific, the earlier rule is selected and the later rule is blocked. This convention makes the priority relation operationally total, even if the explicit priority scores alone are not. Let ⊆H×HC H× H denote a symmetric conflict relation over rule heads. Two heads h and h′h are incompatible if (h,h′)∈(h,h ) or (h′,h)∈(h ,h) . A selected rule set is conflict-consistent if no pair of selected rule heads conflicts: ∀r,r′∈ℛtsel,r≠r′:(hr,hr′)∉∧(hr′,hr)∉.∀ r,r ^sel_t,\ r≠ r :(h_r,h_r ) \ \ (h_r ,h_r) . (21) Multiple rules may be selected at the same decision step when their heads do not conflict. For example, increase_pressure and smooth_policy may be selected together if the conflict map does not declare them incompatible. In that case, the first head determines the direction of the policy update, while the second acts as a modifier of the numerical update. By contrast, increase_pressure and relax_pressure are mutually exclusive and cannot both be selected. Blocked rules are not discarded silently. Each blocked rule is recorded together with the selected rule that blocked it and the reason for blocking. This record becomes part of the explanation object returned by the controller. Algorithm 1 Priority-Based Rule Resolution 1:Facts ℱtF_t, rule base ℛR, conflict relation C 2:Selected rules ℛtselR^sel_t, blocked rules ℛtblkR^blk_t, explanation object ℰtE_t 3:ℛtapp←r∈ℛ:Br⊆ℱtR^app_t←\r :B_r _t\ 4:Sort ℛtappR^app_t by descending π(r)π(r), descending |Br||B_r|, ascending insertion order 5:ℛtsel←∅R^sel_t← 6:ℛtblk←∅R^blk_t← 7:for all r∈ℛtappr ^app_t do 8: conflict←falseconflict 9: blocker←∅blocker← 10: for all r′∈ℛtselr ^sel_t do 11: if (hr,hr′)∈(h_r,h_r ) or (hr′,hr)∈(h_r ,h_r) then 12: conflict←trueconflict 13: blocker←r′blocker← r 14: break 15: end if 16: end for 17: if conflict=falseconflict=false then 18: ℛtsel←ℛtsel∪rR^sel_t ^sel_t∪\r\ 19: else 20: ℛtblk←ℛtblk∪(r,blocker,“conflicting head”)R^blk_t ^blk_t∪\(r,blocker,``conflicting head′)\ 21: end if 22:end for 23:ℰt←ℱt,ℛtapp,ℛtsel,ℛtblk,E_t←\F_t,R^app_t,R^sel_t,R^blk_t,C\ 24:return ℛtsel,ℛtblk,ℰtR^sel_t,R^blk_t,E_t The policy action is then obtained by applying the action mapper H to the selected rule set: at=H(ℛtsel).a_t=H(R^sel_t). (22) When several compatible heads are selected, H first identifies the primary policy-direction head, such as increase_pressure, relax_pressure, or hold_pressure, and then applies any compatible modifier heads, such as smooth_policy. If no primary action is selected, the controller defaults to hold_pressure. This default is also recorded in the explanation object. 4.5 Explanation object Each policy decision produces an explanation object: ⬇ "facts": [...], "applicable_rules": [...], "selected_rules": [...], "blocked_rules": [...], "conflicts": [...], "action": "...", "policy_before": P_t, "policy_after": P_t+1 This object is essential. Without it, the symbolic controller is merely another rule-based controller. With it, the controller becomes inspectable and contestable. 5 Machine-Coaching Revision Layer 5.1 Diagnostic failure modes The coaching layer is triggered by diagnostic failures. The first version of the protocol considers three primary failure modes: boundary violations, over-conservatism, and policy volatility. It also includes one persistence subcase, repeated overshoot, which is treated as a priority-adjustment case rather than as a separate controller family. Table 1: Diagnostic failure modes used for machine-coached revision. Failure mode Diagnostic condition Typical controller Interpretation Boundary violation ViolationRate>τvViolationRate> _v Setpoint Controller tracks the cap but crosses it too often. Over-conservatism ConservatismGap>τcConservatismGap> _c One-sided Controller avoids violations but remains too far below the cap. Volatility PolicyVolatility>τpPolicyVolatility> _p Adaptive controllers Controller reacts too sharply to short-run deviations. These failure modes are selected because they correspond to interpretable regulatory concerns. They are not treated as a complete welfare function. 5.2 Advice representation A coaching advice object is defined as follows. Definition 2 (Coaching advice) A coaching advice object is a tuple c=(op,r,r⋆,π⋆,q),c=(op,r,r ,π ,q), (23) where opop is a revision operation, r is a new rule if applicable, r⋆r is a target rule if applicable, π⋆π is a new priority if applicable, and q is a textual or symbolic reason. The allowed operations are: • add_rule; • remove_rule; • change_priority. Example advice for over-conservatism: ⬇ operation: add_rule rule: IF emissions_safely_below_cap AND conservatism_gap_high THEN relax_pressure PRIORITY 5 reason: Persistent safe emissions indicate excessive policy pressure. Example advice for repeated violations: ⬇ operation: add_rule rule: IF emissions_near_cap AND violation_trend_increasing THEN increase_pressure_preemptively PRIORITY 5 reason: Waiting until emissions exceed the cap creates repeated violations. 5.3 Predefined advice templates To avoid treating the coach as an oracle, coaching advice is not generated freely after inspecting the full results. Instead, the first implementation uses a finite set of predefined advice templates. Each template maps a diagnosed controller failure to a fixed revision operation. The templates are specified before the held-out evaluation phase and are applied mechanically when their triggering conditions are satisfied. Let DkD_k denote the diagnostic summary computed after a diagnosis run or batch. The coaching function is defined as ck=Γ(Dk,ℰk,ℛk),c_k= (D_k,E_k,R_k), (24) where Γ is a predefined advice-template function, ℰkE_k is the set of explanation logs associated with the diagnosed failure, and ℛkR_k is the current rule base. The function Γ is not allowed to invent arbitrary new advice during evaluation. It can only return one of the templates listed in Table 2, or return no_revision if no template condition is satisfied. Table 2: Predefined advice templates used by the coaching layer. Failure mode Trigger condition Revision operation Intended effect Boundary violations V>τvV> _v and O>τoO> _o Add rule: IF emissions_near_cap AND violation_trend_increasing THEN increase_pressure preemptively with priority πp _p Reduce repeated cap crossing by acting before emissions exceed the cap. Over-conservatism CG>τcCG> _c and V≤τvV≤ _v Add rule: IF emissions_safely_below_cap AND conservatism_gap_high THEN relax_pressure with priority πr _r Reduce excessive regulatory pressure when compliance is already stable. Policy volatility PV>τpPV> _p Add rule: IF policy_volatility_high THEN smooth_policy with priority πs _s Reduce sharp policy oscillations caused by short-run emissions fluctuations. Repeated overshoot O>τoO> _o and MOL>τℓMOL> _ Increase priority of: increase_pressure_strongly to π++ _++ Prioritize stronger correction when above-cap episodes persist. No diagnosed failure All failure conditions false no_revision Preserve the current controller. • Note. MOLMOL denotes mean overshoot episode length. The thresholds τv _v, τo _o, τc _c, τp _p, and τℓ _ are fixed before the held-out evaluation. They may be chosen from a calibration set or specified as policy tolerances. The important methodological constraint is that they are not adjusted after observing the held-out results. This design separates machine coaching from manual controller redesign. A manually redesigned controller allows the analyst to inspect a failure and invent a new rule opportunistically. A template-based coached controller applies a predefined mapping from diagnostic evidence to rule-base revision. The coach therefore acts as a reproducible revision function rather than as an unconstrained source of new policy logic. 5.4 Revision operator Let ℛkR_k denote the rule base before coaching event k. Let ckc_k be the coaching advice. The revision operator is ℛk+1=Revise(ℛk,ck).R_k+1=Revise(R_k,c_k). (25) For the three allowed operations: Revise(ℛ,c)=ℛ∪rc,op(c)=add_rule,ℛ∖rc⋆,op(c)=remove_rule,(ℛ∖rc⋆)∪r~c⋆,op(c)=change_priority.Revise(R,c)= casesR∪\r_c\,&op(c)= add\_rule,\\ R \r _c\,&op(c)= remove\_rule,\\ (R \r _c\)∪\ r _c\,&op(c)= change\_priority. cases (26) The revised rule r~c⋆ r _c is identical to rc⋆r _c except that its priority is updated to πc⋆π _c. 5.5 Validity checks after coaching A coaching operation is accepted only if the revised rule base passes a small set of syntactic and consistency checks. These checks are not intended to prove global optimality or logical completeness. They ensure that the controller remains executable and auditable after revision. Let ℛk+1R_k+1 be the candidate rule base after applying coaching advice ckc_k. The revision is accepted only if: (i) every rule has a unique name; (i) every body predicate belongs to the predefined predicate vocabulary; (i) every rule head belongs to the predefined action or modifier vocabulary; (iv) the conflict relation C is symmetric; (v) every action head that should be mutually exclusive with another action head is explicitly listed in C; (vi) priorities are integer-valued; (vii) the rule-resolution algorithm returns a conflict-consistent selected set for all observed diagnostic states in the revision phase. Because the first implementation does not allow rules to generate new predicates recursively, rule chaining is bounded to one step: ℱt→ℛtapp→ℛtsel→at.F_t ^app_t ^sel_t→ a_t. Consequently, coaching cannot introduce inference cycles in the first version of the framework. Future extensions may allow intermediate conclusions or multi-step argumentation, but that would require additional cycle checks and a more explicit argumentation semantics. 5.6 Machine-coaching protocol The full protocol is: Step 1. Run the ABM under an initial controller on diagnosis seeds. Step 2. Compute scalar and symbolic diagnostics. Step 3. Identify whether a predefined failure mode occurred. Step 4. Retrieve the explanation logs associated with the failure. Step 5. Apply the predefined advice-template function Γ . Step 6. If Γ returns no_revision, preserve the current rule base. Step 7. Otherwise, revise the rule base using the returned operation. Step 8. Validate the revised rule base using the syntactic and consistency checks. Step 9. Re-run the ABM from held-out seeds or cloned initial states. Step 10. Compare recurrence of the diagnosed failure and guardrail metrics. The central restriction is that Step 5 is not discretionary during evaluation. The analyst does not freely generate new advice after observing the results. Advice is selected from a predefined template set. This makes the coaching layer reproducible and reduces the risk that apparent improvement is caused by post hoc controller tuning. Algorithm 2 Machine-coaching Advice 1:Diagnostic summary DkD_k, explanation logs ℰkE_k, rule base ℛkR_k 2:Coaching advice ckc_k 3:if V(Dk)>τvV(D_k)> _v and O(Dk)>τoO(D_k)> _o then 4: ck←c_k← add_rule(preemptive_pressure) 5:else if CG(Dk)>τcCG(D_k)> _c and V(Dk)≤τvV(D_k)≤ _v then 6: ck←c_k← add_rule(relaxation) 7:else if PV(Dk)>τpPV(D_k)> _p then 8: ck←c_k← add_rule(smoothing) 9:else if O(Dk)>τoO(D_k)> _o and MeanOvershootLength(Dk)>τℓMeanOvershootLength(D_k)> _ then 10: ck←c_k← change_priority(increase_pressure_strongly, π++ _++) 11:else 12: ck←c_k← no_revision 13:end if 14:Attach supporting explanation logs from ℰkE_k 15:return ckc_k 6 Experimental Design 6.1 Controller families The experiment compares six controller families. Table 3: Controller families compared in the experimental protocol. Code Controller Description F0 Fixed policy Constant policy signal. N1 Numeric setpoint Adjusts pressure according to cap deviation. N2 Numeric safety-margin Tracks an internal cap below the regulatory cap. N3 Numeric one-sided Increases pressure only when emissions exceed the cap. S1 Symbolic rule-based Uses explicit rules, conflicts, priorities, and explanations. S2 Machine-coached symbolic Revises symbolic rules after diagnostic failures. 6.2 Regime comparison The controller families are evaluated across four regimes: CPCA,CPVA,VPCA,VPVA.\CPCA,CPVA,VPCA,VPVA\. (27) The main focus is VPCA and VPVA because the coaching layer applies to adaptive policy controllers. CPCA and CPVA are retained as reference cases. 6.3 Training, coaching, and evaluation split To avoid overfitting the controller to observed simulation failures, the experiment separates three phases. Table 4: Experimental phases. Phase Purpose Data usage Diagnosis phase Identify failure modes Used to trigger coaching. Revision phase Apply rule additions, removals, or priority changes Uses only predefined advice templates. Evaluation phase Test revised controller Uses held-out seeds or cloned initial states. This split is essential. Without it, improvements after coaching may reflect overfitting. 6.4 Primary evaluation criterion The primary evaluation criterion is recurrence of the diagnosed failure mode: ΔRf=Rfbefore−Rfafter, R_f=R_f^before-R_f^after, (28) where RfR_f is the recurrence rate of failure mode f in held-out runs. This avoids claiming success from a single improved metric while ignoring deterioration elsewhere. The coached controller is considered successful for failure mode f if ΔRf>0 R_f>0 and the guardrail constraints defined below are satisfied. This prevents success from being defined after the fact by selecting whichever metric improves most. Guardrail metrics include: • violation rate; • overshoot; • policy volatility; • conservatism gap; • mean tracking error. 6.5 Predefined thresholds and guardrail tolerances The evaluation protocol uses thresholds for two distinct purposes. First, diagnostic thresholds define when a controller failure is considered to have occurred and therefore when coaching is triggered. Second, guardrail tolerances define how much deterioration in secondary metrics is acceptable after coaching. Both sets of values are fixed before the held-out evaluation phase. Let V, O, PVPV, and CGCG denote violation rate, mean overshoot, policy volatility, and conservatism gap, respectively. The diagnostic thresholds are: Table 5: Predefined diagnostic thresholds for coaching triggers. Symbol Metric Suggested value Interpretation τv _v Violation rate 0.050.05 Coaching is triggered if more than 5% of time steps violate the cap. τo _o Mean overshoot 0.02C0.02C Coaching is triggered if average overshoot exceeds 2% of the emissions cap. τp _p Policy volatility 0.10(Pmax−Pmin)0.10(P_ -P_ ) Coaching is triggered if mean absolute policy change exceeds 10% of the feasible policy range. τc _c Conservatism gap 0.05C0.05C Coaching is triggered if emissions remain, on average, more than 5% of the cap below the safe operating region. τℓ _ Mean overshoot episode length 33 time steps Coaching is triggered if above-cap episodes persist for more than three consecutive steps on average. These values are not presented as universal regulatory standards. They are operational thresholds for the stylized simulation environment. In applications, they should be replaced by domain-specific policy tolerances or selected through a calibration phase that is separated from held-out evaluation. The success of a coaching operation is evaluated by the reduction in recurrence of the diagnosed failure mode: ΔRf=Rfbefore−Rfafter, R_f=R_f^before-R_f^after, (29) where RfbeforeR_f^before and RfafterR_f^after denote the held-out recurrence rate of failure mode f before and after coaching. A coaching operation is considered successful only if: ΔRf>0 R_f>0 (30) and no guardrail metric deteriorates beyond its predefined tolerance. Let MjbeforeM_j^before and MjafterM_j^after denote the value of guardrail metric j before and after coaching. For metrics where lower values are preferred, deterioration is defined as: ΔMj=Mjafter−Mjbefore. M_j=M_j^after-M_j^before. (31) The guardrail condition is: ΔMj≤γj∀j∈, M_j≤ _j ∀ j , (32) where G is the set of guardrail metrics and γj _j is the maximum tolerated deterioration. Table 6: Guardrail tolerances for evaluating coached controllers. Symbol Guardrail metric Tolerance Interpretation γv _v Violation rate 0.010.01 Violation rate may not increase by more than one percentage point. γo _o Mean overshoot 0.005C0.005C Mean overshoot may not increase by more than 0.5% of the cap. γp _p Policy volatility 0.02(Pmax−Pmin)0.02(P_ -P_ ) Policy volatility may not increase by more than 2% of the feasible policy range. γc _c Conservatism gap 0.01C0.01C Conservatism gap may not increase by more than 1% of the cap, unless conservatism is the diagnosed failure being corrected. γTE _TE Tracking error 0.01C0.01C Mean tracking error may not increase by more than 1% of the cap. When a guardrail metric is itself the target of coaching, improvement in that metric is evaluated as the primary failure-reduction objective rather than as a guardrail. For example, if the diagnosed failure is over-conservatism, the primary objective is to reduce CGCG; in that case, V, O, PVPV, and TETE serve as guardrails. This explicit thresholding prevents flexible ex post success definitions. A coached controller is not judged successful merely because one favourable metric improves. It must reduce the predefined failure recurrence while remaining within the predefined guardrail tolerances. 6.6 Simulation configuration Table 7 reports the configuration used for the controlled simulation experiment. The table is included to make the numerical results reproducible from the manuscript. The experiment uses separate diagnosis and held-out evaluation seeds. Coaching advice is selected only from the diagnosis phase, while performance is evaluated on held-out seeds. Table 7: Simulation configuration for the controlled coaching experiment. Parameter Symbol Value Simulation setup Number of firms N 100 Simulation horizon T 100 Burn-in – 20 Emissions cap C 75.0 Minimum policy signal PminP_ 0.0 Maximum policy signal PmaxP_ 25.0 Initial policy signal P0P_0 18.0 Random-seed design Diagnosis seeds – 1000–1039 Number of diagnosis seeds – 40 Evaluation seeds – 2000–2039 Number of evaluation seeds – 40 Diagnostic thresholds Violation threshold τv _v 0.05 Overshoot threshold τo _o 1.50 Policy-volatility threshold τp _p 2.50 Conservatism threshold τc _c 3.75 Overshoot-length threshold τℓ _ 3.0 Acceptance guardrails Violation guardrail γv _v 0.01 Overshoot guardrail γo _o 0.375 Policy-volatility guardrail γp _p 0.50 Conservatism guardrail γc _c 0.75 Tracking-error guardrail γTE _TE 0.75 7 Metrics Let R denote the number of replications and T the number of time steps. 7.1 Violation rate V=1RT∑r=1R∑t=1T(Er,t>C).V= 1RT _r=1^R _t=1^TI(E_r,t>C). (33) 7.2 Mean overshoot O=1RT∑r=1R∑t=1Tmax(0,Er,t−C).O= 1RT _r=1^R _t=1^T (0,E_r,t-C). (34) 7.3 Tracking error TE=1RT∑r=1R∑t=1T|Er,t−C|.TE= 1RT _r=1^R _t=1^T|E_r,t-C|. (35) 7.4 Policy volatility PV=1R(T−1)∑r=1R∑t=2T|Pr,t−Pr,t−1|.PV= 1R(T-1) _r=1^R _t=2^T|P_r,t-P_r,t-1|. (36) 7.5 Conservatism gap CG=1RT∑r=1R∑t=1Tmax(0,C−Er,t−δC).CG= 1RT _r=1^R _t=1^T (0,C-E_r,t- _C). (37) 7.6 Explanation and revision metrics The machine-coaching layer also produces controller-level metrics: • number of applicable rules per decision; • number of selected rules per decision; • number of blocked rules per decision; • number of conflicts per run; • number of coaching interventions; • number of persistent rules after revision; • recurrence rate of previously corrected failures. These metrics do not replace policy-performance metrics. They measure auditability and revision dynamics. 8 Simulation Results The purpose of the experiment is to test whether a predefined coaching template can convert a diagnosed symbolic-controller failure into an explicit rule-base revision and whether the revised controller reduces recurrence of that failure under held-out simulation seeds. Figure 1: Machine-coaching controller revision loop. The workflow connects ABM simulation, scalar diagnostics, explanation logs, predefined advice templates, rule-base revision, and held-out re-evaluation. The experiment should be interpreted as a controlled proof-of-concept rather than as broad empirical validation of machine coaching. The initial symbolic controller is intentionally specified without a relaxation exception. This creates a transparent controller-level defect: when emissions remain safely below the cap and the conservatism gap is high, the controller has no rule that can reduce excessive regulatory pressure. The purpose of the experiment is therefore not to show that machine coaching is generally superior to all adaptive-control alternatives, but to test whether the proposed architecture can expose a missing policy-reasoning exception, revise the rule base through a predefined template, and evaluate the revised controller on held-out seeds. The experiment focuses on the VPVA regime, where both the policy controller and the regulated agents are adaptive. The controlled failure mode is over-conservatism. In this setting, the uncoached symbolic controller maintains excessive regulatory pressure, keeping aggregate emissions safely below the cap but farther from the constraint than required. The predefined coaching function Γ therefore selects the over-conservatism template and adds a relaxation rule to the symbolic controller. Figure 2: Aggregate emissions before and after coaching in the VPVA regime. The uncoached symbolic controller remains overly conservative, keeping emissions substantially below the cap. After template-constrained coaching, the added relaxation rule raises emissions closer to the cap while preserving compliance. Shaded areas indicate variability across held-out seeds. 8.1 Diagnosis and coaching advice Table 8 reports the diagnostic summary computed before coaching. The symbolic controller produces no boundary violations, no overshoot, and no policy volatility, but it exhibits a conservatism gap of CG=6.574CG=6.574 and a tracking error of TE=10.324TE=10.324. Since CG>τcCG> _c and V≤τvV≤ _v, the predefined over-conservatism template is triggered. Table 8: Machine-coaching diagnosis summary before rule-base revision. Controller V O PVPV CGCG TETE Mean overshoot length Symbolic 0.000 0.000 0.000 6.574 10.324 0.000 The selected coaching advice is: add_rule: add coach_relax_conservative, with body emissions_safely_below_cap and conservatism_gap_high, head relax_pressure, and priority 55. The reason attached to the advice is that persistent safe emissions indicate excessive policy pressure. This revision is not generated freely after inspecting the results. It is selected mechanically by the predefined advice-template function Γ . 8.2 Uncertainty of the central before–after comparison Table 9 reports paired before–after differences over the 40 held-out evaluation seeds. The bootstrap confidence intervals are computed over the paired seed-level improvements. Positive values indicate improvement, since the reported difference is before minus after. Table 9: Paired held-out differences before and after coaching. Before coaching After coaching Paired difference Metric Mean SD Mean SD Mean improvement Bootstrap 95% CI Conservatism gap, CGCG 6.9526.952 2.9982.998 0.7700.770 1.0141.014 6.1826.182 [5.073, 7.303] Tracking error, TETE 10.66710.667 3.0863.086 4.2414.241 1.2611.261 6.4266.426 [5.250, 7.608] Violation rate, V 0.0000.000 0.0000.000 0.0000.000 0.0000.000 0.0000.000 [0.000, 0.000] Over-conservatism recurrence 0.8750.875 0.3350.335 0.0000.000 0.0000.000 0.8750.875 [0.750, 0.975] The paired comparison confirms that the improvement is not only visible in aggregate means. Across held-out seeds, coaching reduces both the conservatism gap and tracking error, while violation rate remains unchanged at zero. Over-conservatism recurrence decreases by 0.875, with a bootstrap 95% confidence interval of [0.750, 0.975]. 8.3 Controller summary table Table 10 reports the VPVA controller-comparison results generated on held-out evaluation seeds. The main comparison is between the uncoached symbolic controller and the coached symbolic controller. Fixed-policy and numeric adaptive controllers are retained as interpretive baselines. Table 10: VPVA controller comparison on held-out simulation seeds. Controller V O TETE PVPV CGCG Mean overshoot length Runs Fixed policy 0.0000.000 0.0000.000 10.66710.667 0.0000.000 6.9526.952 0.0000.000 40 One-sided 0.0000.000 0.0000.000 10.66710.667 0.0000.000 6.9526.952 0.0000.000 40 Safety-margin 0.0000.000 0.0000.000 3.7723.772 0.0190.019 0.1370.137 0.0000.000 40 Setpoint 0.4310.431 0.1010.101 0.2600.260 0.0200.020 0.0000.000 1.7181.718 40 Symbolic 0.0000.000 0.0000.000 10.66710.667 0.0000.000 6.9526.952 0.0000.000 40 Coached symbolic 0.0000.000 0.0000.000 4.2414.241 0.0000.000 0.7700.770 0.0000.000 40 The held-out VPVA comparison shows that the coached symbolic controller reduces the diagnosed over-conservatism failure. The uncoached symbolic controller has CG=6.952CG=6.952 and TE=10.667TE=10.667, while the coached symbolic controller has CG=0.770CG=0.770 and TE=4.241TE=4.241. Boundary violations, overshoot, and policy volatility remain equal to zero. Thus, the coached controller reduces over-conservatism without violating the predefined guardrails. The setpoint controller provides a useful contrast. It achieves low tracking error, but at the cost of frequent boundary violations, with a VPVA violation rate of 0.4310.431 and a mean overshoot of 0.1010.101. The safety-margin controller has the lowest conservatism gap among the non-symbolic adaptive controllers, but it is not explanation-driven and does not expose rule-level reasons for policy revision. The coached symbolic controller therefore occupies a different methodological position: its value lies not in numerical optimality, but in connecting diagnosis, explanation, revision, and held-out re-evaluation. 8.4 Failure recurrence table Table 11 reports failure recurrence before and after coaching. Over-conservatism recurrence falls from 0.8750.875 before coaching to 0.0000.000 after coaching. Boundary violations, policy-volatility failures, and repeated overshoot remain absent both before and after coaching. Table 11: Failure recurrence before and after coaching. Failure mode Before coaching After coaching Change Interpretation Boundary violations 0.000 0.000 0.000 Unchanged; no boundary-violation failure is introduced. Over-conservatism 0.875 0.000 0.875 Reduced; the diagnosed failure is eliminated in held-out evaluation. Policy volatility 0.000 0.000 0.000 Unchanged; the revision does not introduce volatility. Repeated overshoot 0.000 0.000 0.000 Unchanged; the revision does not introduce persistent above-cap episodes. The result supports the restricted claim of the paper: the machine-coaching layer can convert an ex post diagnostic failure into a rule-level controller revision whose effect is measurable under held-out simulation seeds. It does not imply that machine coaching improves every controller metric or that the coached symbolic controller is globally optimal. 8.5 Robustness checks Two local robustness checks are used to evaluate whether the result depends entirely on a single threshold or a single initial policy value. These checks are not intended to establish global robustness. They test whether the controlled proof-of-concept remains coherent under nearby conservatism thresholds and alternative initial policy pressures. Table 12 reports sensitivity to the over-conservatism threshold τc _c. The coaching template is triggered under all three tested values, τc∈0.03C,0.05C,0.07C _c∈\0.03C,0.05C,0.07C\. The recurrence of over-conservatism decreases in all cases. Under the default threshold τc=0.05C _c=0.05C, recurrence falls from 0.875 to 0.000. Table 12: Sensitivity to the over-conservatism threshold τc _c. τc _c Advice τc _c value Before CGCG After CGCG Before recurrence After recurrence Change 0.03C0.03C Over-conservatism 2.25 6.952 0.770 0.950 0.125 0.825 0.05C0.05C Over-conservatism 3.75 6.952 0.770 0.875 0.000 0.875 0.07C0.07C Over-conservatism 5.25 6.952 0.770 0.750 0.000 0.750 Table 13 reports sensitivity to the initial policy signal P0P_0. Under the default batch-level diagnostic rule, coaching is not triggered for P0=14P_0=14 or P0=16P_0=16, because the diagnosis-phase conservatism gap does not exceed τc _c. Coaching is triggered for P0=18P_0=18 and P0=20P_0=20, where the controller is substantially over-conservative. In both triggered cases, the coached controller reduces over-conservatism recurrence to zero. Table 13: Sensitivity to initial policy signal P0P_0. Conservatism gap Recurrence Tracking error P0P_0 Advice Diagnosis CGCG Before CGCG After CGCG Before After Change Before TETE After TETE 1414 No diagnosed failure 0.3630.363 0.7120.712 0.7120.712 0.0500.050 0.0500.050 0.0000.000 3.4713.471 3.4713.471 1616 No diagnosed failure 2.7002.700 3.2423.242 3.2423.242 0.3500.350 0.3500.350 0.0000.000 6.8446.844 6.8446.844 1818 Over-conservatism 6.5746.574 6.9526.952 0.7700.770 0.8750.875 0.0000.000 0.8750.875 10.66710.667 4.2414.241 2020 Over-conservatism 10.55610.556 10.91510.915 0.2410.241 0.9750.975 0.0000.000 0.9750.975 14.66514.665 3.5673.567 These checks qualify the interpretation of the result. The coaching layer does not act continuously for every conservative configuration. It acts only when the predefined diagnostic condition is met. This is consistent with the design of the paper: the coach is a template-constrained revision function, not an unconstrained optimizer. 8.6 Explanation-log table Table 14 reports the explanation-log entry associated with the coached revision. Before coaching, the relevant facts are emissions_safely_below_cap and conservatism_gap_high. The only applicable rule is the default hold-pressure rule. No rule is blocked. The absence of a relaxation rule explains why the symbolic controller remains conservative. The coaching template therefore adds a rule that relaxes policy pressure when emissions are safely below the cap and the conservatism gap is high. Table 14: Explanation log for the coached over-conservatism revision. Field Example Detected failure Over-conservatism Facts conservatism_gap_high, emissions_safely_below_cap Applicable rule(s) r0_default_hold Blocked rule(s) none Action before revision hold_pressure Coaching revision Add coach_relax_conservative: IF emissions_safely_below_cap AND conservatism_gap_high THEN relax_pressure with priority 55. Expected effect Reduce excessive policy pressure while monitoring violations, overshoot, policy volatility, and tracking error as guardrails. This table illustrates the intended role of the explanation object. The controller does not merely report a numerical outcome. It identifies the facts, applicable rules, and missing exception that justify the revision. The rule-base update is therefore inspectable and auditable. 9 Discussion The controlled simulation results show how the proposed machine-coaching layer changes the role of diagnostics in adaptive ABM. Diagnostics are not used only as ex post summaries. They are used as triggers for explicit controller revision. In the reported VPVA experiment, the symbolic controller is diagnosed as over-conservative: it produces no violations and no overshoot, but keeps emissions unnecessarily far below the cap. The explanation log identifies the reason at the controller level: the applicable rule is the default hold-pressure rule, and no relaxation rule is available. The experiment shows that, in a controlled over-conservatism case, an explanation-driven revision layer can identify a missing controller exception and add it through a predefined advice template. This is sufficient for the methodological claim advanced here: machine coaching operationalizes controller-level contestability by making policy reasoning inspectable, revisable, and testable. The robustness checks strengthen, but do not generalize, the proof-of-concept. The result is stable across nearby over-conservatism thresholds and across the initial policy values for which over-conservatism is actually diagnosed. At lower initial policy values, coaching is not triggered, and the symbolic controller remains unchanged. This behaviour is desirable in the present framework because the coach is designed as a bounded revision operator, not as a continuous optimizer. The predefined coaching template converts this diagnostic failure into a rule-level revision. The added rule, coach_relax_conservative, relaxes policy pressure when emissions are safely below the cap and the conservatism gap is high. Held-out evaluation then shows that over-conservatism recurrence falls from 0.8750.875 to 0.0000.000, while violations, overshoot, and policy volatility remain at zero. This provides a compact demonstration of controller-level contestability: the controller’s reasoning is explained, challenged, revised, and re-evaluated. The coached symbolic controller is not shown to be globally optimal. In fact, the numeric setpoint controller achieves very low tracking error but does so with frequent cap violations. The safety-margin controller achieves a lower conservatism gap than the coached symbolic controller in this simulation run, but it does not provide the same rule-level explanation and revision trace. The contribution of the coached controller is therefore not numerical dominance across all metrics. Its contribution is that a diagnosed failure can be linked to an auditable policy-reasoning defect and corrected through a predefined revision operation. This distinction is important. The machine-coaching layer does not explain the full ABM. It does not identify all causal mechanisms behind emergent macro-patterns. It does not replace sensitivity analysis, structural causal modeling, computational-mechanics diagnostics, trajectory mining, or clustering. It explains and revises only the policy-controller logic. In this sense, the proposed layer is complementary to broader ABM explainability methods. The approach is also not equivalent to full human-in-the-loop machine coaching. In the present implementation, coaching is constrained to predefined advice templates. This restriction improves reproducibility but limits expressiveness. Future work may extend the protocol to expert elicitation, stakeholder argumentation, or natural-language interfaces. Such extensions would introduce additional validation problems, including how advice is elicited, how conflicting stakeholder claims are resolved, and how coached revisions are protected against strategic or biased intervention. 10 Limitations The first limitation is that the emissions model is stylized. It is a controlled simulation environment, not a calibrated emissions-sector model. Numerical results should therefore be interpreted as evidence about controller behavior under controlled assumptions, not as environmental-policy forecasts. The second limitation is that the empirical demonstration isolates a single failure mode. The reported experiment focuses on over-conservatism in the VPVA regime. The coached controller stabilizes at a lower policy level after adding the relaxation rule, thereby reducing the conservatism gap. This is a proof-of-concept for explanation-driven rule revision, not evidence that the same coaching layer will improve all adaptive policy failures. A related limitation is that the empirical result is partly constructed by design. The initial symbolic controller lacks a relaxation rule, the diagnostic procedure detects the resulting over-conservatism, and the predefined coaching template adds the missing relaxation rule. This design is appropriate for a proof-of-concept because it creates a transparent and auditable controller defect. However, it does not establish that the same coaching architecture will discover non-obvious revisions, improve controllers under ambiguous failures, or outperform well-tuned numerical controllers. The third limitation concerns the symbolic abstraction. Indeed, predicate thresholds such as emissions_near_cap, emissions_safely_below_cap, and conservatism_gap_high are design choices. They must be documented and subjected to sensitivity analysis. Different threshold values may alter which advice template is triggered. The fourth limitation concerns coaching. Although the present implementation avoids free-form advice by using predefined templates, the templates themselves are design choices. They should be fixed before held-out evaluation and, in applied settings, justified through domain expertise, policy tolerance, or a separate calibration phase. The fifth limitation is that rule-based explanations may be incomplete. A controller may explain why it selected a policy action, but that does not explain all downstream effects of the action in the ABM. Controller-level explainability and system-level explainability remain distinct. The sixth limitation is that better explainability may not imply better scalar performance. A coached symbolic controller may be more auditable while being less efficient under a narrow objective. This is not necessarily a failure if the purpose is contestable policy design rather than pure optimization. However, it limits the claim that can be made from the present experiment. Last, the robustness checks are local. They vary the over-conservatism threshold and initial policy signal, but they do not exhaust the full parameter space of the emissions model, agent-adaptation dynamics, symbolic predicate definitions, or coaching-template design. They should therefore be interpreted as evidence that the controlled result is not tied to a single threshold value, not as evidence of global robustness. 11 Conclusion This paper proposed and implemented a lightweight machine-coached policy-revision layer for adaptive agent-based regulation. The layer represents policy decisions as symbolic rules with conflicts and priorities, produces explanation logs for controller actions, and allows diagnostic failures to be translated into explicit rule revisions. The controlled simulation experiment demonstrates the mechanism on a specific VPVA over-conservatism case. The uncoached symbolic controller produces no cap violations but remains excessively conservative, with a held-out conservatism gap of CG=6.952CG=6.952 and tracking error of TE=10.667TE=10.667. The predefined coaching function selects the over-conservatism template and adds a relaxation rule. After coaching, the held-out conservatism gap falls to CG=0.770CG=0.770 and tracking error falls to TE=4.241TE=4.241, while violation rate, overshoot, and policy volatility remain at zero. Over-conservatism recurrence falls from 0.8750.875 to 0.0000.000. Machine coaching is not presented as a complete ABM explainability solution or as an optimal-control method. Instead, it is a controller-level contestability mechanism that complements existing diagnostic approaches. Its value lies in connecting ex post regime diagnostics to reproducible policy-controller revision. The proposed protocol is especially suitable as a derived extension of adaptive ABM research. In a broader explainability architecture, causal diagnostics, information-theoretic measures, trajectory mining, and clustering explain system dynamics; machine coaching explains and revises the controller that acts on those dynamics. Code and Data Availability The replication package includes the simulation extension, notebook integration code, configuration parameters, random seeds, diagnosis outputs, held-out evaluation outputs, controller-level metrics, explanation logs, rule-revision logs, and all tables and figures used in the manuscript. The repository will be archived with a persistent identifier before public release. Disclosure on the Use of Generative AI Large language model tools were used to assist with language editing, LaTeX formatting, figure-caption refinement, and code-debugging support. All scientific content, simulation design, numerical results, interpretation, and final claims were reviewed and validated by the author, who remains fully responsible for the manuscript. References [1] Arthur, W. B. (1994). Inductive reasoning and bounded rationality. American Economic Review, 84(2), 406–411. https://ideas.repec.org/a/aea/aecrev/v84y1994i2p406-11.html [2] Busoniu, L., Babuska, R., & De Schutter, B. (2008). A comprehensive survey of multiagent reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics, Part C, 38(2), 156–172. https://doi.org/10.1109/TSMCC.2007.913919 [3] Conte, R., & Paolucci, M. (2014). On agent-based modeling and computational social science. Frontiers in Psychology, 5, 668. https://doi.org/10.3389/fpsyg.2014.00668 [4] Crutchfield, J. P. (1994). The calculi of emergence: Computation, dynamics and induction. Physica D: Nonlinear Phenomena, 75(1–3), 11–54. https://doi.org/10.1016/0167-2789(94)90273-9 [5] Epstein, J. M. (1999). Agent-based computational models and generative social science. Complexity, 4(5), 41–60. https://onlinelibrary.wiley.com/doi/10.1002/%28SICI%291099-0526%28199905/06%294%3A5%3C41%3A%3AAID-CPLX9%3E3.0.CO%3B2-F [6] Epstein, J. M. (2012). Generative Social Science: Studies in Agent-Based Computational Modeling. Princeton University Press. https://w.jstor.org/stable/j.ctt7rxj1 [7] Garrone, R. (2025). An adaptive, data-integrated agent-based modeling framework for explainable and contestable policy design. arXiv preprint arXiv:2511.19726. https://arxiv.org/abs/2511.19726 [8] Garrone, R. (2026). Structural distinguishability of static and adaptive policy regimes in agent-based regulation. Preprint. [9] Gilbert, N. (2008). Agent-Based Models. SAGE Publications. https://doi.org/10.4135/9781412983259 [10] McCarthy, J. (1959). Programs with common sense. In Proceedings of the Teddington Conference on the Mechanization of Thought Processes. http://jmc.stanford.edu/articles/mcc59/mcc59.pdf [11] Shalizi, C. R., & Crutchfield, J. P. (2001). Computational mechanics: Pattern and prediction, structure and simplicity. Journal of Statistical Physics, 104, 817–879. https://doi.org/10.1023/A:1010388907793 [12] Tesfatsion, L. (2006). Agent-based computational economics: A constructive approach to economic theory. In L. Tesfatsion & K. L. Judd (Eds.), Handbook of Computational Economics, Vol. 2. Elsevier. https://doi.org/10.1016/S1574-0021(05)02016-2 [13] Tesfatsion, L., & Judd, K. L. (Eds.). (2006). Handbook of Computational Economics, Volume 2: Agent-Based Computational Economics. Elsevier. https://shop.elsevier.com/books/handbook-of-computational-economics/tesfatsion/978-0-444-51253-6 [14] Zhang, K., Yang, Z., & Basar, T. (2021). Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control. Springer. https://arxiv.org/abs/1911.10635 [15] Amershi, S., Cakmak, M., Knox, W. B., & Kulesza, T. (2014). Power to the people: The role of humans in interactive machine learning. AI Magazine, 35(4), 105–120. https://doi.org/10.1609/aimag.v35i4.2513 [16] Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012 [17] Christiano, P. F., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems. https://arxiv.org/abs/1706.03741 [18] Dung, P. M. (1995). On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games. Artificial Intelligence, 77(2), 321–357. https://doi.org/10.1016/0004-3702(94)00041-X [19] Grosan, C., & Abraham, A. (2011). Rule-based expert systems. In Intelligent Systems: A Modern Approach. Springer. https://link.springer.com/book/10.1007/978-3-642-21004-4 [20] Karimi, A.-H., Schölkopf, B., & Valera, I. (2021). Algorithmic recourse: From counterfactual explanations to interventions. In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency. https://doi.org/10.1145/3442188.3445899 [21] Nute, D. (1994). Defeasible logic. In D. M. Gabbay, C. J. Hogger, & J. A. Robinson (Eds.), Handbook of Logic in Artificial Intelligence and Logic Programming, Vol. 3. Oxford University Press. https://dl.acm.org/doi/10.5555/186124.186131 [22] Wachter, S., Mittelstadt, B., & Russell, C. (2017). Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Technology, 31(2), 841–887. https://arxiv.org/abs/1711.00399 [23] Michael, L. (2019). Machine Coaching. In Proceedings of the IJCAI 2019 Workshop on Explainable Artificial Intelligence (XAI). Macao, China. https://w.researchgate.net/publication/334989337_Machine_Coaching [24] Markos, V., Thoma, M., & Michael, L. (2022). Machine Coaching with Proxy Coaches. In Proceedings of the Workshop on Argumentation and Machine Learning (ArgML@COMMA). CEUR Workshop Proceedings, Vol. 3208. https://ceur-ws.org/Vol-3208/paper4.pdf