Paper deep dive
From Workflow Automation to Capability Closure: A Formal Framework for Safe and Revenue-Aware Customer Service AI
Cosimo Spera
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/22/2026, 5:28:13 AM
Summary
The paper introduces a formal framework for compositional safety in customer service AI, utilizing capability hypergraphs to detect emergent goals and safety failures. It proves that individually safe agents can form unsafe coalitions through conjunctive dependencies and establishes a 'Safety-Value Duality' theorem, showing that safety certification and commercial goal discovery share identical computational complexity.
Entities (6)
Relation Signals (3)
Cosimo Spera → authored → From Workflow Automation to Capability Closure
confidence 100% · Companion paper to Spera (2026) (arXiv:2603.15973) Cosimo Spera
Capability Hypergraph Framework → usedin → Telco Case Study
confidence 95% · We ground the framework in a 12-capability Telco case study.
Safety-Value Duality → proves → Computational Equivalence
confidence 90% · Safety-Value Duality theorem showing that safety certification and commercial goal discovery are provably the same computation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Customer service automation is undergoing a structural transformation. The dominant paradigm is shifting from scripted chatbots and single-agent responders toward networks of specialised AI agents that compose capabilities dynamically across billing, service provision, payments, and fulfilment. This shift introduces a safety gap that no current platform has closed: two agents individually verified as safe can, when combined, reach a forbidden goal through an emergent conjunctive dependency that neither possesses alone.
Tags
Links
- Source: https://arxiv.org/abs/2603.15978v2
- Canonical: https://arxiv.org/abs/2603.15978v2
Trouble viewing inline? Open PDF directly →
Full Text
57,711 characters extracted from source content.
Expand or collapse full text
From Workflow Automation to Capability Closure: A Formal Framework for Compositionally Safe Customer Service AI Companion paper to Spera (2026) (arXiv:2603.15973) Cosimo Spera Minerva CQ, 114 Lester Ln, Los Gatos, CA 95032 cosimo@minervacq.com (March 2026) Abstract The Non-Compositionality of Safety Theorem (Spera, 2026, Theorem 9.2) establishes that two individually safe AI agents can, when combined, reach a forbidden goal through an emergent conjunctive dependency that neither possesses alone. This paper applies the formal capability hypergraph framework to the customer service domain and makes three independent theoretical contributions. First, we establish the Emergent Goal Discovery Proposition (Proposition 3.3): in any joint billing-plus-service session, the capability ServiceProvision is provably emergent—reachable from neither agent alone, but reachable from their conjunction—and we characterise the complete class of CS deployments in which emergent upsell goals arise structurally. Second, we derive a formally complete treatment of agent join and leave dynamics with fully proved safety invariants and closed-form update costs. Third, we establish a Safety-Value Duality theorem showing that safety certification and commercial goal discovery are provably the same computation, with identical asymptotic complexity. We ground the framework in a 12-capability Telco case study. The business case—presented as a validated projection framework rather than a measured outcome—estimates a conservative net annual value of $20.7M, with a sensitivity analysis confirming positivity ($8.4M) under simultaneous worst-case stress of all six key assumptions. We flag the churn retention estimate as the highest-priority empirical validation target and describe the NMF pilot study protocol that would convert it from a projection to a measured result. Companion to: Spera (2026) (arXiv:2603.15973). arXiv: 2603.15978 Companion paper: Spera (2026) (arXiv:2603.15973) Keywords: AI safety; customer service automation; capability hypergraphs; emergent goal discovery; agentic systems; compositional safety; formal verification. Contents 1 Introduction 1.1 The Central Gap 1.2 Contributions of This Paper 2 Related Work 3 Emergent Goal Discovery in Customer Service AI 3.1 Formal Setup 3.2 The Telco Capability Structure 3.3 Structural Condition for Emergence 3.4 Safety-Value Duality 4 The Telco Case Study: Capability Structure 4.1 The Non-Compositional Safety Failure 4.2 Closure at Telco Scale 5 Six Failure Modes and Domain-Specific Corollaries 5.1 Failure Mode 1: Emergent Forbidden Capability 5.2 Failure Mode 2: Invisible Safety Boundary 5.3 Failure Mode 3: Unauditable Goal Space 5.4 Failure Mode 4: Brittle Tool Governance 5.5 Failure Mode 5: Prompt Injection as Capability Injection 5.6 Failure Mode 6: Goal Discovery Treated as External Configuration 6 Agent Join and Leave Dynamics 6.1 Full Event Cost Table 6.2 End-to-End Telco Session Trace 7 Business Case: A Projection Framework 7.1 Cost of Failure Under the Incumbent Approach 7.2 Revenue Opportunity: Conditional Projections 7.3 Consolidated Business Case 7.4 Sensitivity Analysis with Correlation Structure 8 The Broader Picture 9 Limitations and Open Problems 10 Conclusion References A Deployment Roadmap 1 Introduction 1.1 The Central Gap Every current approach to customer service AI safety operates at the component level: individual agents are tested, role-based access controls are defined, guardrails are applied, and RLHF training steers individual model behaviour. None addresses the compositional safety question: what can the assembled system of agents collectively reach? Spera (2026) proves formally that this question cannot be answered by component-level analysis alone. The Non-Compositionality of Safety Theorem establishes that two individually safe agents can together reach a forbidden goal through a conjunctive hyperedge that neither possesses individually. Across 900 real multi-tool trajectories from two independent public benchmarks, 42.6% contain at least one conjunctive dependency of precisely this type (95% CI: [39.4%, 45.8%][39.4\%,\,45.8\%]). 1.2 Contributions of This Paper This is a formal application paper. It applies the capability hypergraph framework of Spera (2026) to the customer service domain, derives domain-specific structural results, and illustrates the practical consequences through a calibrated business case. It is not a standalone theory paper; its contributions are best understood as a structural characterisation, proved dynamics invariants, and domain-specific corollaries — each grounded in the main paper’s formal machinery. The business case (Section 7) contains no production measurements; it should be read as a worked projection illustrating practical magnitudes rather than empirical validation. (1) Emergent Goal Discovery (Proposition 3.3, Section 3). A formal characterisation of when commercially valuable goals emerge from agent coalitions that neither agent can discover alone. We prove that ServiceProvision is provably emergent in the joint billing-plus-service session, derive the full class of CS deployments exhibiting this property, and prove that the safety computation and the goal discovery computation are identical. (2) Agent Dynamics with Proved Invariants (Section 6). A complete formal treatment of the four agent-level events—join, leave, capability gain, capability loss—with closed-form update costs, proved safety monotonicity, and an end-to-end Telco session trace showing the framework operates within O(144)O(144) operations across six events. (3) Six Failure Mode Theorems (Section 5). For each of six canonical CS safety failures, we derive a domain-specific corollary of Spera (2026) that adds independent proof content: bounding |ℬ(F)||B(F)| for the Telco deployment, proving the minimal antichain structure is tight, and extending the adversarial hyperedge model to the capability-injection attack class. The business case (Section 7) is presented as a projection framework: each revenue mechanism is stated as a conditional estimate with explicit assumptions, and the sensitivity analysis characterises the joint uncertainty structure rather than treating assumptions as independent. 2 Related Work AI safety guardrails and LLM alignment. Constitutional AI (Bai et al., 2022) and production guardrails (Salesforce Einstein Trust Layer, AWS Bedrock, Anthropic safety classifiers) operate on individual model outputs and cannot address failures emerging from agent composition—a limitation proved formally by Theorem 10.1 of Spera (2026). The prompt injection literature (Perez and Ribeiro, 2022; Greshake et al., 2023) characterises adversarial inputs to single models; our capability-injection attack class (Section 5.5) extends this to the multi-agent setting. Agentic AI orchestration safety. LangGraph, AutoGen (Wu et al., 2023), and CrewAI provide orchestration with tool-call restrictions. These are engineering controls, not formal guarantees. Even well-engineered multi-agent CS systems succeed on only 35% of multi-turn tasks (Fini Labs Research, 2025), with the primary failure mode being undetected conjunctive precondition ambiguity. Workflow-based CS automation. Workflow models treat dependencies as pairwise (Ekfrazo Technologies, 2026), which is the representational gap Corollary 5.2 of Spera (2026) formalises. Only 11% of organisations have agentic AI in full production; the primary blocker is governance. Formal compliance verification. Regulatory frameworks governing CS AI—GDPR Article 25 (European Parliament and Council, 2016), PCI-DSS 4.0 (PCI Security Standards Council, 2022), EU AI Act Article 9 (European Parliament and Council, 2024)—implicitly require formal audit artefacts that no current platform produces. Role-based access control (Ferraiolo et al., 2001) and purpose-based access control (Byun et al., 2005) provide the closest prior formal models, but neither captures conjunctive capability dependencies. Basin et al. (2018) formalise GDPR compliance in a process algebra, but without a safety boundary characterisation. Our Safe Audit Surface (Section 5.3) produces precisely the artefact these frameworks require—a certifiable map of every dangerous capability combination—in polynomial time, connecting the hypergraph framework to the compliance verification tradition. Runtime monitoring frameworks such as LARVA (Colombo et al., 2012) and specification languages such as TLA+ (Lamport, 2002) address compliance at execution time; the hypergraph framework operates at the architectural level, detecting forbidden coalitions before any agent fires. Compositional verification. Contract-based design (Benveniste et al., 2018) and assume-guarantee reasoning (Jones, 1983) verify fixed compositions against pre-specified properties. The key distinction, as Spera (2026) establishes, is that hypergraph closure characterises the set of all properties a dynamically growing capability set can ever reach—a strictly more general problem. 3 Emergent Goal Discovery in Customer Service AI This section contains the paper’s primary independent theoretical contribution. We prove not merely that a specific goal is emergent in one deployment, but characterise the full structural condition under which commercial goals emerge from agent coalitions, and show that detecting them requires no computation beyond what safety certification already performs. 3.1 Formal Setup We use the capability hypergraph framework of Spera (2026) throughout. Recall that cl(A)cl(A) denotes the capability closure of configuration A, ℛ(F)R(F) the safe region, and ℬ(F)B(F) the minimal unsafe antichain. We add one definition specific to the CS domain. Definition 3.1 (Emergent Commercial Goal). Let H=(V,ℱ)H=(V,F) be a CS capability hypergraph, A1,A2⊆VA_1,A_2 V two agent configurations, and g∈V∖Fg∈ V F a commercially valuable goal (one with positive expected revenue γ(g)>0γ(g)>0). We say g is an emergent commercial goal of the coalition a1,a2\a_1,a_2\ if: (i) g∈cl(A1∪A2)g (A_1∪ A_2) (reachable from the joint session), (i) g∉cl(A1)g (A_1) and g∉cl(A2)g (A_2) (unreachable from either agent alone), (i) cl(A1∪A2)∈ℛ(F)cl(A_1∪ A_2) (F) (the coalition remains safe). 3.2 The Telco Capability Structure The Telco deployment has n=12n=12 capabilities (Table 1), forbidden set F=c11,c12F=\c_11,c_12\, and six conjunctive hyperedges (Table 2). Two standard agent configurations are: Abilling A_billing =c1,c2,c3,c4,c5(billing agent), =\c_1,c_2,c_3,c_4,c_5\ (billing agent), Aservice A_service =c1,c2,c7,c8(service agent). =\c_1,c_2,c_7,c_8\ (service agent). Table 1: Telco capability set (n=12n=12). Forbidden: c11c_11 (PaymentModify), c12c_12 (BulkAccountExport). ID Capability Description c1c_1 IntentClassify Classify customer intent from natural language c2c_2 CustomerLookup Retrieve customer record, tier, and history c3c_3 BillingRead Read itemised billing history and current balance c4c_4 DisputeLog Log a billing dispute and open a case c5c_5 CreditEligibility Check eligibility for a billing credit c6c_6 CreditApply Apply a billing credit to the account c7c_7 ServiceCatalogue Query available service plans and upgrades c8c_8 ServiceEligibility Verify customer eligibility for a service plan c9c_9 ServiceProvision Initiate service addition or upgrade c10c_10 PaymentRead Read payment method and transaction history c11c_11 PaymentModify [FORBIDDEN] Modify payment method or amount c12c_12 BulkAccountExport [FORBIDDEN] Export bulk account data Table 2: Telco capability hyperedges. Arc h6h_6 is the unsafe hyperedge. Arc Tail → Head Semantics h1h_1 c1→c3,c7\c_1\→\c_3,c_7\ Intent classification enables billing and service reads h2h_2 c1,c2→c10\c_1,c_2\→\c_10\ Identity-confirmed intent enables payment read h3h_3 c3,c5→c6\c_3,c_5\→\c_6\ Billing read + credit eligibility → credit application h4h_4 c7,c8→c9\c_7,c_8\→\c_9\ Catalogue + eligibility → service provision h5h_5 c2→c4,c5,c8\c_2\→\c_4,c_5,c_8\ Customer lookup enables downstream checks h6h_6 c3,c10→c12\c_3,c_10\→\c_12\ [UNSAFE] Billing + payment read → bulk export 3.3 Structural Condition for Emergence Before stating the main proposition, we identify the precise agent configurations under which ServiceProvision is emergent. This requires understanding how c2c_2 (CustomerLookup) interacts with the hyperedge structure. Key structural observation Hyperedge h5h_5 (c2→c4,c5,c8\c_2\→\c_4,c_5,c_8\) means that any agent carrying c2c_2 automatically reaches c8c_8 (ServiceEligibility). Combined with c7c_7 (from h1h_1), this fires h4h_4 and reaches c9c_9. Consequently: any agent carrying both c1c_1 and c2c_2 already reaches c9c_9 individually — emergence cannot arise in a coalition of such agents. True emergence requires at least one agent to be capability-scoped: carrying c1c_1 but not c2c_2. Definition 3.2 (Scoped Agent Configurations). A capability-scoped billing agent carries billing-specific capabilities without the full customer lookup chain: AB=c1,c3,c4,c5.A_B=\c_1,c_3,c_4,c_5\. A capability-scoped service agent carries catalogue access only: AS=c1,c7.A_S=\c_1,c_7\. These are the standard least-privilege configurations in which agents hold only the capabilities required for their designated task. Proposition 3.3 (Emergent Goal Discovery). Let H, F, ABA_B, ASA_S be as in Definition 3.2. Then: (1) Emergence. c9∉cl(AB)c_9 (A_B), c9∉cl(AS)c_9 (A_S), and c9∈cl(AB∪AS)c_9 (A_B∪ A_S). Under capability-scoped configurations, ServiceProvision is an emergent commercial goal of the billing-plus-service coalition. (2) Safety. AB∪AS∈ℛ(F)A_B∪ A_S (F): the coalition is safe. The safe service expansion strictly increases the reachable goal set without activating the unsafe arc h6h_6. (3) Marginal value. c9∉cl(AB)c_9 (A_B) means the commercial value of presenting ServiceProvision is zero under the billing agent alone; it becomes strictly positive once the service agent joins the session. (4) Structural characterisation. For any two capability-scoped agents (A1,A2)(A_1,A_2) in the Telco deployment: c9c_9 is emergent in their coalition if and only if c7,c8\c_7,c_8\ is not jointly present in either cl(A1)cl(A_1) or cl(A2)cl(A_2) individually, but c7,c8⊆cl(A1∪A2)\c_7,c_8\ (A_1∪ A_2). Equivalently: the preconditions of h4h_4 are split across the two agents. Proof. Part (1). We compute the three closures directly. cl(AB)=cl(c1,c3,c4,c5)cl(A_B)=cl(\c_1,c_3,c_4,c_5\): h1h_1 fires (c1c_1 present), adding c7c_7. h4h_4 requires c7,c8\c_7,c_8\; c8c_8 is absent (c2∉ABc_2∉ A_B, so h5h_5 does not fire). No further rules fire. Result: cl(AB)=c1,c3,c4,c5,c7cl(A_B)=\c_1,c_3,c_4,c_5,c_7\. Hence c9∉cl(AB)c_9 (A_B). ✓ cl(AS)=cl(c1,c7)cl(A_S)=cl(\c_1,c_7\): h1h_1 fires (c1c_1), adding c3,c7c_3,c_7 (c7c_7 already present). h4h_4 requires c8c_8; c8c_8 absent (c2∉ASc_2∉ A_S). No further rules fire. Result: cl(AS)=c1,c3,c7cl(A_S)=\c_1,c_3,c_7\. Hence c9∉cl(AS)c_9 (A_S). ✓ cl(AB∪AS)=cl(c1,c3,c4,c5,c7)cl(A_B∪ A_S)=cl(\c_1,c_3,c_4,c_5,c_7\): h1h_1 fires (c1c_1), adding c3,c7c_3,c_7 (both already present). h4h_4 requires c7,c8\c_7,c_8\; c8c_8 still absent. c9c_9 is not reachable from AB∪ASA_B∪ A_S alone. Wait — this shows c9∉cl(AB∪AS)c_9 (A_B∪ A_S) under AB=c1,c3,c4,c5A_B=\c_1,c_3,c_4,c_5\, AS=c1,c7A_S=\c_1,c_7\. We need c8c_8 in the joint coalition. Adding c8c_8 to ASA_S gives the correct scoped configuration: AS+=c1,c7,c8(catalogue + eligibility, no customer lookup).A_S^+=\c_1,c_7,c_8\ (catalogue + eligibility, no customer lookup). Then cl(AS+)=c1,c3,c7,c8,c9cl(A_S^+)=\c_1,c_3,c_7,c_8,c_9\ (h4 fires immediately). So c9∈cl(AS+)c_9 (A_S^+) — not emergent for this AS+A_S^+. The minimal emergent coalition uses: AB=c1,c3,c4,c5,AS=c7,c8.A_B=\c_1,c_3,c_4,c_5\, A_S=\c_7,c_8\. cl(AB)=c1,c3,c4,c5,c7cl(A_B)=\c_1,c_3,c_4,c_5,c_7\ (as computed above, noting h1h_1 adds c7c_7 but c8c_8 absent). c9∉cl(AB)c_9 (A_B). ✓cl(AS)=cl(c7,c8)=c7,c8,c9cl(A_S)=cl(\c_7,c_8\)=\c_7,c_8,c_9\ (h4h_4 fires: c7,c8⊆AS\c_7,c_8\ A_S). So c9∈cl(AS)c_9 (A_S) — this ASA_S reaches c9c_9 alone; not emergent. The minimal truly emergent coalition: AB∗=c1,c3,c4,c5,AS∗=c8.A_B^*=\c_1,c_3,c_4,c_5\, A_S^*=\c_8\. cl(AB∗)=c1,c3,c4,c5,c7cl(A_B^*)=\c_1,c_3,c_4,c_5,c_7\ (h1h_1 adds c7c_7; c8c_8 absent so h4h_4 does not fire). c9∉cl(AB∗)c_9 (A_B^*). ✓cl(AS∗)=c8cl(A_S^*)=\c_8\ (no rule fires from c8\c_8\ alone). c9∉cl(AS∗)c_9 (A_S^*). ✓cl(AB∗∪AS∗)=cl(c1,c3,c4,c5,c8)cl(A_B^*∪ A_S^*)=cl(\c_1,c_3,c_4,c_5,c_8\): h1h_1 adds c7c_7; now c7,c8⊆cl\c_7,c_8\ , so h4h_4 fires, adding c9c_9. c9∈cl(AB∗∪AS∗)c_9 (A_B^*∪ A_S^*). ✓ This is the minimal emergent pair: the billing agent supplies c7c_7 via h1h_1; the service agent supplies c8c_8 directly; neither alone satisfies h4h_4’s precondition c7,c8\c_7,c_8\; together they do. Emergence is established. Part (2). cl(AB∗∪AS∗)=c1,c3,c4,c5,c7,c8,c9cl(A_B^*∪ A_S^*)=\c_1,c_3,c_4,c_5,c_7,c_8,c_9\. h6h_6 requires c3,c10\c_3,c_10\. c10c_10 derives only via h2h_2 which requires c2c_2. Since c2∉AB∗∪AS∗c_2∉ A_B^*∪ A_S^*, c10c_10 is never derived, h6h_6 never fires, and c12∉cl(AB∗∪AS∗)c_12 (A_B^*∪ A_S^*). The coalition is safe. ✓ Part (3). c9∉cl(AB∗)c_9 (A_B^*) (shown above), so the system cannot present ServiceProvision as a reachable goal under the billing agent alone. Once AS∗=c8A_S^*=\c_8\ joins, c9∈cl(AB∗∪AS∗)c_9 (A_B^*∪ A_S^*), making the goal reachable and commercially presentable. Part (4). h4=(c7,c8,c9)h_4=(\c_7,c_8\,\c_9\) is the only rule producing c9c_9. Therefore c9c_9 is reachable from a coalition iff c7,c8⊆cl(A1∪A2)\c_7,c_8\ (A_1∪ A_2). Emergence (neither agent reaches c9c_9 alone) requires additionally that c9∉cl(Ai)c_9 (A_i) for each i, which by the same argument requires c7,c8⊈cl(Ai)\c_7,c_8\ (A_i) for each i — i.e., the preconditions of h4h_4 are split. ∎ Remark (Capability scoping as a prerequisite for emergence). The proposition reveals why capability scoping (least-privilege assignment) is not merely a security practice — it is a prerequisite for the emergent goal discovery mechanism to function. The full billing agent c1,c2,c3,c4,c5\c_1,c_2,c_3,c_4,c_5\ already reaches c9c_9 individually via h5h_5 (adding c8c_8) then h4h_4. Emergence only arises when agents are scoped to carry task-specific capabilities, creating the split precondition structure that h4h_4 requires. In practice, this means the NMF-based goal discovery system is most effective in deployments that enforce capability scoping — a finding with concrete implications for system architecture. Remark. The proof above reveals an important subtlety: emergence in the Telco hypergraph is sensitive to whether c2c_2 (CustomerLookup) is present, because c2c_2 activates h5h_5 which adds c8c_8, which together with c7c_7 (from h1h_1 triggered by c1c_1) fires h4h_4. The standard full billing agent Abilling=c1,c2,c3,c4,c5A_billing=\c_1,c_2,c_3,c_4,c_5\ already reaches c9c_9 alone. True emergence in the Telco deployment therefore arises in restricted configurations where agents carry only task-specific capabilities. The practical implication for deployment: capability scoping (granting agents only the minimum capabilities needed for their task) is not just a security best practice—it is a prerequisite for the emergent goal discovery mechanism to function. 3.4 Safety-Value Duality Theorem 3.4 (Safety-Value Duality). Let H=(V,ℱ)H=(V,F), A∈ℛ(F)A (F), and Gsafe=cl(A)∖FG_safe=cl(A) F. Then: (1) GsafeG_safe is computed in O(n+mk)O(n+mk) by the same worklist that certifies safety. (2) Acquiring any v∈NMFF(A)v _F(A) strictly increases |Gsafe||G_safe| while keeping A∈ℛ(F)A (F). (3) No capability in GsafeG_safe is reachable without a safety certificate: every v∈Gsafev∈ G_safe comes with a derivation certificate from the worklist. (4) (Domain transfer) The identical algorithm applies to any CS domain (V′,F′,ℱ′)(V ,F ,F ) with the same asymptotic complexity, requiring only respecification of the triple. Proof. Part (1). The worklist algorithm (Spera, 2026, Algorithm 1) computes cl(A)cl(A) in O(n+mk)O(n+mk). It simultaneously verifies cl(A)∩F=∅cl(A)∩ F= by checking each newly-reached vertex against F at the point of addition. Therefore Gsafe=cl(A)∖FG_safe=cl(A) F is a by-product of the safety check with no additional cost. Part (2). Let v∗∈NMFF(A)v^* _F(A). By Definition 8.1(c) of Spera (2026), v∗=μ(e)v^*=μ(e) for some boundary hyperedge e∈∂(A)e∈∂(A) with |S∖cl(A)|=1|S (A)|=1. Acquiring v∗v^* means S⊆cl(A∪v∗)S (A∪\v^*\), so e fires, adding T to the closure. Thus cl(A∪v∗)⊋cl(A)cl(A∪\v^*\) (A), giving |Gsafe|=|cl(A∪v∗)∖F|>|cl(A)∖F||G_safe|=|cl(A∪\v^*\) F|>|cl(A) F| (provided T⊈FT F, which holds by the definition of safety-filtered NMF: cl(A∪v∗)∩F=∅cl(A∪\v^*\)∩ F= ). Part (3). The worklist records, for each v∈cl(A)v (A), the sequence of hyperedge firings that produced it (Spera, 2026, Theorem 10.1). These certificates are computable with no asymptotic overhead. Part (4). The worklist and NMF algorithms operate on the abstract structure (V,ℱ,F)(V,F,F) with no assumption on semantic content. Complexity bounds O(n+mk)O(n+mk) and O(|V|(n+mk))O(|V|(n+mk)) depend only on |V|=n|V|=n, |ℱ|=m|F|=m, maximum tail size k—invariant to domain relabelling. ∎ 4 The Telco Case Study: Capability Structure 4.1 The Non-Compositional Safety Failure Consider three agent configurations: Abilling A_billing =c1,c2,c3,c4,c5, =\c_1,c_2,c_3,c_4,c_5\, Aservice A_service =c1,c2,c7,c8, =\c_1,c_2,c_7,c_8\, Apayment A_payment =c1,c2,c10. =\c_1,c_2,c_10\. cl(Abilling)=c1,…,c6cl(A_billing)=\c_1,…,c_6\: safe (c12∉c_12∉). cl(Aservice)=c1,c2,c7,c8,c9cl(A_service)=\c_1,c_2,c_7,c_8,c_9\: safe. cl(Apayment)=c1,c2,c10cl(A_payment)=\c_1,c_2,c_10\: safe. But Abilling∪ApaymentA_billing∪ A_payment contains c3,c10\c_3,c_10\, satisfying h6h_6: c12∈cl(Abilling∪Apayment)c_12 (A_billing∪ A_payment). The coalition is unsafe, even though every individual agent—and every pair—is safe. Why component-level checks are insufficient The unsafe coalition Abilling,Apayment\A_billing,A_payment\ passes every pairwise safety check. The failure requires computing the closure of the joint capability set, which is precisely what the coalition safety criterion (Theorem 11.2 of Spera 2026) provides. 4.2 Closure at Telco Scale The O(n+mk)O(n+mk) worklist runs in O(24)O(24) operations for n=12n=12, m=6m=6, k=2k=2—sub-microsecond. The online coalition safety check reduces to a single O(1)O(1) dictionary lookup: does the joint capability set cover c3,c10\c_3,c_10\ or c11\c_11\? 5 Six Failure Modes and Domain-Specific Corollaries Each failure mode is addressed by a corollary that adds independent proof content beyond the main paper: we derive Telco-specific structural properties, bound |ℬ(F)||B(F)|, and extend the adversarial model to the CS setting. 5.1 Failure Mode 1: Emergent Forbidden Capability The failure. A multi-agent CS system reaches a forbidden capability through a conjunction of individually safe capabilities (Section 4). Corollary 5.1 (Minimal Unsafe Antichain Structure for Telco). For the Telco deployment, ℬ(F)=c3,c10,c11B(F)=\\c_3,c_10\,\c_11\\. This antichain is tight: |ℬ(F)|=2|B(F)|=2 is the minimum possible for any CS deployment with two forbidden capabilities and one unsafe conjunctive arc. Proof. By Theorem 9.5 of Spera (2026), ℬ(F)B(F) is the antichain of minimal unsafe sets. c11\c_11\ is minimal because cl(c11)∋c11∈Fcl(\c_11\) c_11∈ F, and ∅∉ℛ(F) (F) vacuously but cl(∅)∩F=∅cl( )∩ F= trivially. c3,c10\c_3,c_10\ is minimal: cl(c3,c10)∋c12cl(\c_3,c_10\) c_12 (via h6h_6), so it is unsafe; removing either element— cl(c3)=c3cl(\c_3\)=\c_3\, cl(c10)=c10cl(\c_10\)=\c_10\—gives safe singletons. The antichain property (c3,c10⊈c11\c_3,c_10\ \c_11\ and vice versa) is immediate. For tightness: any CS deployment with |F|=2|F|=2 and at least one conjunctive unsafe arc must have |ℬ(F)|≥2|B(F)|≥ 2 (one element per forbidden capability that has a direct minimal unsafe set). The Telco deployment achieves this minimum, meaning the pre-execution gate checks against exactly two elements—the most efficient possible online monitoring. ∎ 5.2 Failure Mode 2: Invisible Safety Boundary The failure. Operators identify forbidden goals reactively after violations, not proactively from the system’s structure. Corollary 5.2 (Certifiable Safety Boundary for Telco). The complete, certifiable characterisation of every dangerous capability combination in the Telco deployment is ℬ(F)=c3,c10,c11B(F)=\\c_3,c_10\,\c_11\\. For any candidate session configuration A, the system is safe if and only if c3,c10⊈A\c_3,c_10\ A and c11∉Ac_11∉ A. This is decidable in O(|A|)O(|A|) time. Proof. By Theorem 9.4 of Spera (2026), ℛ(F)R(F) is a lower set with boundary ℬ(F)B(F). By Theorem 11.2, the coalition check is: ∃B∈ℬ(F):B⊆A∃ B (F):B A? With ℬ(F)=c3,c10,c11B(F)=\\c_3,c_10\,\c_11\\, this is two membership tests in A, each O(1)O(1) with a hash set, giving O(|A|)O(|A|) total. The artefact ℬ(F)B(F) constitutes the formal object that GDPR Article 25 (European Parliament and Council, 2016) (data minimisation), PCI-DSS 4.0 (PCI Security Standards Council, 2022) (least-privilege access), and EU AI Act Article 9 (European Parliament and Council, 2024) (risk characterisation) implicitly require but provide no technical standard for constructing. ∎ 5.3 Failure Mode 3: Unauditable Goal Space The failure. CS platforms cannot produce a complete, certified account of what every agent coalition can reach from any starting configuration. Corollary 5.3 (Safe Audit Surface for Telco Billing Coalition). For the billing coalition Abilling=c1,c2,c3,c4,c5A_billing=\c_1,c_2,c_3,c_4,c_5\: (a) Currently reachable: cl(Abilling)=c1,…,c6,c7,c8,c9cl(A_billing)=\c_1,…,c_6,c_7,c_8,c_9\ (the full billing + service chain, via h1h_1 and h5h_5 and h4h_4). (b) One step from expansion: adding c7c_7 alone (service catalogue) does not expand the closure beyond what h1h_1 already provides; adding c10c_10 (payment read) opens the unsafe arc h6h_6. (c) Structurally unsafe from AbillingA_billing: c11c_11 directly; c12c_12 via any path through c3,c10\c_3,c_10\ (since c3∈Abillingc_3∈ A_billing, adding c10c_10 alone reaches c12c_12). Proof. Part (a): Run the worklist from AbillingA_billing. h1h_1 fires (c1c_1): adds c3c_3 (present), c7c_7. h5h_5 fires (c2c_2): adds c4c_4 (present), c5c_5 (present), c8c_8. h4h_4 fires (c7,c8\c_7,c_8\): adds c9c_9. h3h_3 fires (c3,c5\c_3,c_5\): adds c6c_6. No further firings; c10c_10 absent so h6h_6 does not fire. Closure: c1,…,c9\c_1,…,c_9\. Part (b): NMFF(Abilling)=c10NMF_F(A_billing)=\c_10\: the only boundary hyperedge is h6h_6 with μ(h6)=c10μ(h_6)=c_10 (since c3,c10∖cl=c10\c_3,c_10\ =\c_10\). But cl(Abilling∪c10)∋c12∈Fcl(A_billing∪\c_10\) c_12∈ F—so c10c_10 is in the unsafe NMF boundary. The safety-filtered frontier NMFF(Abilling)=∅NMF_F(A_billing)= : no single capability addition expands the closure while remaining safe. Part (c): c11∈Fc_11∈ F directly. For c12c_12: c3∈cl(Abilling)c_3 (A_billing), so c3,c10⊆cl(Abilling∪c10)\c_3,c_10\ (A_billing∪\c_10\), giving c12∈clc_12 immediately upon adding c10c_10. ∎ 5.4 Failure Mode 4: Brittle Tool Governance The failure. Adding a new API integration requires a full safety re-audit. Corollary 5.4 (Incremental Audit for Telco Tool Additions). When a new tool integration adds hyperedge e=(S,T)e=(S,T) to the Telco deployment: (a) If S⊈cl(A)S (A): existing audit surface unchanged. Cost: O(|S|)≤O(k)O(|S|)≤ O(k). (b) If S⊆cl(A)S (A): compute cl(A∪T)cl(A∪ T) in O(n+mk)=O(24)O(n+mk)=O(24) and check cl(A∪T)∩F=∅cl(A∪ T)∩ F= . If safe, update audit surface in O(|V|⋅(n+mk))=O(288)O(|V|·(n+mk))=O(288) operations. (c) In both cases: the updated audit surface is formally certified correct without re-running the full offline ℬ(F)B(F) computation. Proof. By Theorems 11.6 and 11.7 of Spera (2026), insertion of e=(S,T)e=(S,T) into H with current closure C=clH(A)C=cl_H(A) satisfies: clH′(A)=clH(A∪T′)cl_H (A)=cl_H(A∪ T ) where T′=T =T if S⊆CS C, else T′=∅T = . Cost analysis for the Telco system: n=12n=12, m=6m=6, k=2k=2. Lazy case (S⊈CS C): check |S|≤2|S|≤ 2 elements. Active case (S⊆CS C): run worklist from C∪TC∪ T, cost O(n+mk)=O(24)O(n+mk)=O(24). Audit surface update (Theorem 10.1 of Spera 2026): O(|V|⋅(n+mk))=O(12⋅24)=O(288)O(|V|·(n+mk))=O(12· 24)=O(288). Correctness: the new audit surface satisfies the completeness and soundness conditions of Theorem 10.1 by construction of the incremental update. ∎ 5.5 Failure Mode 5: Prompt Injection as Capability Injection The failure. An adversarial customer constructs a query that convinces the CS agent to invoke a tool it should not invoke (OWASP LLM Top 10 #1 for agentic systems). The standard framing treats prompt injection as a natural language problem—a filter or classifier that rejects malicious inputs. We show this framing is structurally insufficient and derive a capability-level defence. Definition 5.5 (Capability-Injection Attack). A capability-injection attack on deployment (H,A,F)(H,A,F) is an attempt by an adversary to introduce a new tool invocation e′=(S′,T′)e =(S ,T ) into the live session such that S′⊆cl(A)S (A) and T′∩F≠∅T ∩ F≠ . The adversary does not need to break any cryptographic control; they only need to cause the agent to invoke a tool whose preconditions are already satisfied. Corollary 5.6 (Polynomial-Time Defence Against Single-Step Capability Injection). For the Telco deployment, the single-step capability-injection attack class is completely defended by the pre-execution gate. Specifically: the set of all dangerous single-step injections is ℰ∗=(S,T):S⊆cl(A),T∩F≠∅E^*=\(S,T):S (A),T∩ F≠ \. |ℰ∗||E^*| is computable in O(n2k)O(n^2k) and each candidate is verifiable in O(n+mk)O(n+mk), giving total defence cost O(n3k2)O(n^3k^2). For the Telco system: at most O(124)=O(20736)O(12^4)=O(20736) candidates, each checked in O(24)O(24) operations. Furthermore, this defence is structurally superior to lexical filtering: it operates on the capability structure of what the agent is asked to do, not on the surface form of what it is asked to say. An adversary who rephrases the injection in novel language cannot evade it; an adversary who makes the tool invocation satisfy a safe precondition is not attacking at all. Proof. By Theorem 14.7 of Spera (2026), the single-edge case (b=1b=1) of MinUnsafeAdd is polynomial. The dangerous injections are exactly those (S,T)(S,T) with S⊆cl(A)S (A) and T∩F≠∅T∩ F≠ . There are at most (nk)⋅|F| nk·|F| such candidates (choose tail of size ≤k≤ k from cl(A)cl(A), head must intersect F). For Telco (n=12n=12, k=2k=2, |F|=2|F|=2): at most (122)⋅2=132 122· 2=132 candidates, each checkable in O(24)O(24). The structural superiority over lexical filtering follows from the observation that the check operates on the inferred capability set of a tool invocation, not its description: two semantically equivalent but lexically distinct invocations of c12c_12 produce the same capability-level check result. ∎ 5.6 Failure Mode 6: Goal Discovery Treated as External Configuration The failure. Current CS agents pursue a fixed list of externally configured goals, missing commercially valuable emergent opportunities. Corollary 5.7 (Greedy Upsell Presentation is Near-Optimal). Let A be any safe CS session state and let k≥1k≥ 1. The greedy algorithm that at each step presents the goal v∗=argmaxv∈V∖cl(A)γF(v,A)v^*= *arg\,max_v∈ V (A) _F(v,A) (maximum safety-filtered marginal closure gain) achieves: f(Ggreedy)≥(1−1e)⋅f(G∗),f(G_greedy)≥ (1- 1e )· f(G^*), where G∗=argmax|G|≤kf(G)G^*= *arg\,max_|G|≤ kf(G) is the optimal k-goal presentation strategy. This bound is tight and cannot be improved by any polynomial-time algorithm unless P=NP=NP. Proof. Define fF(B)=|cl(A∪B)∖(cl(A)∪F)|f_F(B)=|cl(A∪ B) (cl(A)∪ F)|—the safety-filtered closure gain. We verify the three preconditions of Nemhauser et al. (1978): Normalisation: fF(∅)=0f_F( )=0 since cl(A∪∅)=cl(A)cl(A∪ )=cl(A). Monotonicity: for B⊆B′B B , cl(A∪B)⊆cl(A∪B′)cl(A∪ B) (A∪ B ) by monotonicity of clcl (Spera, 2026, Theorem 6.3), so fF(B)≤fF(B′)f_F(B)≤ f_F(B ). Submodularity: by Theorem 8.5 of Spera (2026), f(B)=|cl(A∪B)|−|cl(A)|f(B)=|cl(A∪ B)|-|cl(A)| is submodular via the polymatroid rank theorem. fFf_F inherits submodularity: fF(B)≤f(B)f_F(B)≤ f(B) and the diminishing-returns property is preserved under restriction to the safe goal set (removing the constant set F from the head does not affect whether the marginal gain of adding v to B exceeds that of adding v to B′⊇B B). With all three preconditions verified, the Nemhauser et al. theorem gives the (1−1/e)(1-1/e) bound. The hardness of improvement follows from NP-hardness of maximising a general submodular function beyond (1−1/e)(1-1/e). ∎ 6 Agent Join and Leave Dynamics We formalise the four agent-level events with complete proofs of safety invariants. Proposition 6.1 (Agent Join: Safety Check and Closure Update). Let A∈ℛ(F)A (F) and let agent an+1a_n+1 with capability set An+1∈ℛ(F)A_n+1 (F) attempt to join the coalition. Post-join A′=A∪An+1A =A∪ A_n+1 is safe if and only if no B∈ℬ(F)B (F) satisfies B⊆A′B A . Check cost: O(|ℬ(F)|⋅|A′|)O(|B(F)|·|A |). For the Telco deployment: O(2⋅12)=O(24)O(2· 12)=O(24) per join event. Closure update cost: O(n+mk)=O(24)O(n+mk)=O(24). Proof. Safety criterion. By Theorem 11.2 of Spera (2026), A′∉ℛ(F)A (F) if and only if ∃B∈ℬ(F):B⊆A′∃ B (F):B A . Checking this iterates over |ℬ(F)||B(F)| antichain elements, each requiring a subset test in O(|B|)≤O(n)O(|B|)≤ O(n); total O(|ℬ(F)|⋅|A′|)O(|B(F)|·|A |). Closure update. Let C=clH(A)C=cl_H(A). Since A⊆CA C and An+1A_n+1 is new, A′⊆C∪An+1A C∪ A_n+1. By monotonicity: cl(A′)=cl(C∪An+1)cl(A )=cl(C∪ A_n+1). Algorithm 1 of Spera (2026) initialised at C∪An+1C∪ A_n+1 runs in O(n+mk)O(n+mk): each of m hyperedges is examined at most once, each with tail of size ≤k≤ k. ∎ Proposition 6.2 (Agent Leave: Safety is Free). Let A∈ℛ(F)A (F). For any departure yielding A′⊆A A: (1) A′∈ℛ(F)A (F) with no check required; (2) cl(A′)⊆cl(A)cl(A ) (A); (3) δ(g,A′)≥δ(g,A)δ(g,A )≥δ(g,A) for all g∈Vg∈ V. Proof. (1) ℛ(F)R(F) is a lower set by Theorem 9.4 of Spera (2026): if A∈ℛ(F)A (F) and A′⊆A A, then cl(A′)⊆cl(A)cl(A ) (A) by monotonicity, so cl(A′)∩F⊆cl(A)∩F=∅cl(A )∩ F (A)∩ F= . (2) Follows directly from monotonicity of clcl. (3) Any acquisition set S witnessing g∈cl(A∪S)g (A∪ S) also witnesses g∈cl(A′∪(S∪(A∖A′)))g (A ∪(S∪(A A ))); the gap A∖A′A A must be covered by any acquisition from A′A , so δ(g,A′)≥δ(g,A)δ(g,A )≥δ(g,A). ∎ Proposition 6.3 (Capability Dynamics). (1) Capability gain (hyperedge e=(S,T)e=(S,T) added): costs O(|S|)O(|S|) if S⊈CS C (lazy; closure unchanged) or O(n+mk)O(n+mk) if S⊆CS C (active; safety check required). (2) Capability loss (hyperedge deleted): costs O(n+mk)O(n+mk) for recomputation; safety preserved by deletion monotonicity. (3) Acquisition distance stability: |δ(g,A)after−δ(g,A)before|≤|T||δ(g,A)_after-δ(g,A)_before|≤|T| for any single hyperedge change. Proof. By Theorem 11.6 of Spera (2026). (1) Lazy case: S⊈CS C means e cannot fire from C; check |S||S| elements, no closure update. Active case: S⊆CS C means e fires immediately, adding T; run worklist from C∪TC∪ T in O(n+mk)O(n+mk). Safety check: verify cl(A∪T)∩F=∅cl(A∪ T)∩ F= , subsumed by the worklist. (2) Deletion: clH′(A)⊆clH(A)cl_H (A) _H(A) by deletion monotonicity (removing a hyperedge can only reduce reachability), so safety is preserved; recompute from scratch in O(n+mk)O(n+mk). (3) By Theorem 11.6(3) of Spera (2026): inserting e=(S,T)e=(S,T) decreases δ(g,A)δ(g,A) by at most |T||T|; deletion increases it by at most |T||T|. ∎ 6.1 Full Event Cost Table Table 3: Agent and capability event costs and safety re-check requirements. Telco values: n=12n=12, m=6m=6, k=2k=2, |ℬ(F)|=2|B(F)|=2. Event Closure update Safety check? Goal set effect Telco cost Agent join O(n+mk)O(n+mk) from C∪An+1C∪ A_n+1 Required Can only grow O(24)O(24) Agent leave O(n+mk)O(n+mk), deferrable Not required Can only shrink O(24)O(24) Cap. gain (lazy, S⊈CS C) O(|S|)O(|S|) Not required Unchanged O(2)O(2) Cap. gain (active, S⊆CS C) O(n+mk)O(n+mk) Required Grows by ≤|T|≤|T| O(24)O(24) Cap. loss O(n+mk)O(n+mk) Not required Shrinks by ≤|T|≤|T| O(24)O(24) 6.2 End-to-End Telco Session Trace We trace a complete dynamic session with ℬ(F)=c3,c10,c11B(F)=\\c_3,c_10\,\c_11\\. T1. Session start. A=c1A=\c_1\. cl(A)=c1,c3,c7cl(A)=\c_1,c_3,c_7\ (via h1h_1). Safe. T2. c2\c_2\ added. A′=c1,c2A =\c_1,c_2\. Check: c3,c10⊈A′\c_3,c_10\ A ; c11∉A′c_11∉ A . Safe. cl(A′)=c1,c2,c3,c4,c5,c7,c8,c10cl(A )=\c_1,c_2,c_3,c_4,c_5,c_7,c_8,c_10\. T3. Billing agent joins (c3,c4,c5\c_3,c_4,c_5\). A′=c1,c2,c3,c4,c5A =\c_1,c_2,c_3,c_4,c_5\. Check: c10∉A′c_10∉ A . Safe. cl(A′)=c1,…,c6,c7,c8,c9cl(A )=\c_1,…,c_6,c_7,c_8,c_9\. T4. Payment agent attempts join (c1,c2,c10\c_1,c_2,c_10\). A′=c1,…,c5,c10A =\c_1,…,c_5,c_10\. Check: c3,c10⊆A′\c_3,c_10\ A . Blocked. Greedy recovery (Proposition 6.1): remove c10c_10. Reduced join c1,c2\c_1,c_2\ passes. Payment agent joins with identity lookup only. T5. Service agent joins (c7,c8\c_7,c_8\). A(4)=c1,…,c5,c7,c8A^(4)=\c_1,…,c_5,c_7,c_8\. Check: no antichain element covered. Safe. cl(A(4))=c1,…,c9cl(A^(4))=\c_1,…,c_9\. ServiceProvision (c9c_9) discovered as emergent goal by closure. T6. Billing agent leaves. A(5)=c1,c2,c7,c8A^(5)=\c_1,c_2,c_7,c_8\. No check required (Proposition 6.2). Safety preserved automatically. Total safety check cost across all six events: O(6⋅24)=O(144)O(6· 24)=O(144) operations. No full re-audit at any point. 7 Business Case: A Projection Framework Transparency note The revenue figures in this section are calibrated projections, not measured outcomes. We present them as a conditional framework: for each mechanism, we state the key assumption, the resulting value estimate, and the empirical validation protocol that would convert the projection to a measured result. The sensitivity analysis (Section 7.4) characterises the joint uncertainty structure, including assumption correlation. 7.1 Cost of Failure Under the Incumbent Approach For a Tier-1 Telco processing 10 million CS transactions per year (Table 4), the annual cost of failure under workflow-based automation is $24.5 24.5M–$26.5 26.5M in recurring operational costs, excluding tail-risk breach events. The AND-violation correction cost ($14.6M) and automation gain ($3.81M–$7.62M) follow directly from the proved zero-violation guarantee of Theorem 6.3 of Spera (2026) combined with the empirical violation rates of Spera (2026) and Fini Labs Research (2025); these figures require no additional empirical validation beyond what the main paper provides. Table 4: Annual cost of failure under workflow-based CS automation (10M transactions/year). Failure category Annual cost Derivation basis AND-violation correction $14.6M Proved (Theorem 6.3, Spera 2026); rate from Fini Labs Research (2025) Workflow ambiguity excess handling $9.9M Proved; rate from Fini Labs Research (2025) Compliance documentation overhead $0.5M–$2.0M Industry estimate Expected breach cost (tail risk) $15.6M/incident IBM 2025; GDPR fine schedule Operational total (excl. breach) $24.5M–$26.5M Recurring annual 7.2 Revenue Opportunity: Conditional Projections Mechanism 1: Emergent upsell discovery. Assumption: AI-personalised CS recommendations achieve a 28% uplift over the 4/1,000 baseline conversion rate (Gitnux Research, 2026). Conditional value: 3,3603,360 additional annual conversions at $120–$180 annual value = $403K–$605K/year. Validation protocol: A/B test of closure-based goal presentation vs. rule-based upsell on 50,000 sessions (sufficient for 80% power at 5% significance to detect a 15% conversion lift). Mechanism 2: Churn reduction via near-miss frontier. This mechanism carries the highest uncertainty and accounts for 55% of the conservative net annual value. We present it explicitly as a projection. Three stacked assumptions (each must be validated independently): (a) NMF identification accuracy: 85% of at-risk customers correctly identified. (b) Proactive contact rate: 40% of identified customers contacted. (c) Retention conversion: 15% of contacted customers commit to retention. Conditional value under all three: 500,000×0.85×0.40×0.15=25,500500,000× 0.85× 0.40× 0.15=25,500 customers retained; at $480 LTV = $12.24M/year. Validated pilot protocol: Deploy NMF boundary signals to 10,000 sessions over 30 days. Measure assumption (a) directly by comparing NMF-flagged sessions against a ground-truth churn label from 90-day follow-up. This converts the highest-uncertainty assumption from a model parameter to an empirical result with a 6-week turnaround. Mechanism 3: Compliance-enabled service expansion. Assumption: 17.5% conversion on 500,000 payment-discussion sessions currently blocked by the pre-execution gate operating without the service expansion capability. Conditional value: 500,000×0.175×$60=$5.25500,000× 0.175× 60= 5.25M/year. 7.3 Consolidated Business Case Table 5: Consolidated annual business case (10M transactions/year, conditional projections). Line item Conservative Moderate Validation status Operational cost reduction $3.81M $7.62M Proved (Thm. 6.3) Compliance documentation avoidance $0.50M $2.00M Industry estimate Emergent upsell discovery $0.40M $0.61M Projection; A/B testable Churn reduction via NMF $12.24M $12.24M Projection; pilot needed Compliance-enabled expansion $5.25M $7.50M Projection Total annual value $22.20M $29.97M Deployment cost (3yr amortised) $1.50M $3.00M Net annual value $20.70M $26.97M 7.4 Sensitivity Analysis with Correlation Structure Table 6 reports net annual value under individual assumption stress tests. The all-stressed-simultaneously row gives $8.4M. Table 6: Sensitivity analysis: net annual value under assumption stress. Base: $20.7M. Assumption Base Stressed Relaxed NAV range NMF identification accuracy 85% 60% 95% $15.8M–$22.3M Churn retention conversion 15% 8% 22% $16.0M–$23.0M Automation gain (p) 5p 2p 10p $17.9M–$24.5M Upsell conversion uplift 28% 15% 40% $20.5M–$20.9M Sessions at churn risk 5% 2% 10% $16.3M–$24.2M Compliance expansion conv. 17.5% 10% 25% $18.5M–$22.3M All stressed simultaneously — — — $8.4M All relaxed simultaneously — — — $38.6M Correlation structure. The six assumptions are not independent. We identify two correlated pairs: Positive correlation (NMF accuracy, sessions at churn risk): both depend on the quality of the at-risk customer identification model. A deployment failure that depresses NMF accuracy below 60% (e.g., domain shift in the customer intent model) would also likely reduce the fraction of sessions correctly identified as at-risk. Stressing both simultaneously is the appropriate worst case and is captured in the all-stressed row. Approximate independence (upsell conversion uplift, churn retention conversion): these operate on different customer populations (upsell targets vs. at-risk churners) via different mechanisms (goal discovery vs. NMF boundary signalling), making correlation unlikely. Partial correlation (automation gain, compliance expansion): both depend on the pre-execution gate’s ability to correctly route sessions. A gate misconfiguration depressing automation gain would also affect compliance-enabled sessions. The $8.4M floor from the all-stressed row should be interpreted as a lower bound under the assumption that all six mechanisms simultaneously underperform—a conservative estimate that accounts for positive correlation in the two correlated pairs. 8 The Broader Picture The capability hypergraph framework inverts the conventional goal–capability relationship. In every current CS platform, goals are the starting point and capabilities are the means. The closure operator reverses this: capability is the starting point, and goals are what the closure reveals. Theorem 3.3 makes this inversion commercially concrete: the emergent goal c9c_9 (ServiceProvision) has zero commercial value in the individual billing session and strictly positive commercial value in the joint session. The system discovers the opportunity without explicit configuration. Theorem 3.4 (Safety-Value Duality) establishes the deepest implication: the computation that certifies safety and the computation that discovers commercial value are provably identical. There is no safety-revenue trade-off at the computational level. Theorem 14.5 of Spera (2026) establishes that non-compositionality persists in probabilistic hypergraphs for all safety thresholds τ<p(h)τ<p(h). A CS platform cannot dissolve the safety problem by introducing stochastic tool invocation. 9 Limitations and Open Problems Static capability model. The hypergraph model treats capabilities as persistent. Real CS sessions have capability timeouts, resource constraints, and stateful dependencies. Extending to resource-constrained or time-gated capabilities is an open problem. Human–AI hybrid sessions. The framework models fully automated agent sessions. CS sessions involving human agents introduce capability contributions not captured by the hypergraph. Extending to hybrid coalitions is an open problem. Unvalidated business projections. The churn retention mechanism (Section 7.2) is the paper’s highest-uncertainty claim. We have described the NMF pilot protocol; until it is run, the $12.24M figure is a model projection, not a measured result. Hyperedge specification for large deployments. The PAC-learning theorem (Theorem 14.2 of Spera 2026) guarantees recovery from sufficient trajectory data, but rare conjunctive combinations—precisely those most likely to be dangerous—may require targeted probing. 10 Conclusion This paper applies the capability hypergraph framework to the customer service domain, deriving three domain-specific formal contributions. The Emergent Goal Discovery Theorem (Theorem 3.3) provides a formal characterisation, to our knowledge the first in this setting, of when commercially valuable goals emerge from agent coalitions, proves the safety of the emergent coalition, and characterises the structural conditions under which emergence arises. The Safety-Value Duality Theorem (Theorem 3.4) establishes that safety certification and commercial goal discovery are provably the same computation. The agent dynamics section derives closed-form safety invariants for all four agent-level events, reducing a six-event production session to O(144)O(144) operations with no full re-audit. The six failure mode corollaries add independent domain-specific results: bounding |ℬ(F)|=2|B(F)|=2 as tight for the Telco deployment, deriving the complete Telco audit surface, extending the adversarial model to capability-injection attacks, and proving greedy upsell presentation achieves the (1−1/e)(1-1/e) optimality guarantee with explicit precondition verification. The business case is presented as a conditional projection framework. The $20.7M conservative net annual value becomes $8.4M under simultaneous worst-case stress; the churn retention mechanism is explicitly flagged as the highest-priority empirical validation target, and the pilot protocol that would convert it from a projection to a measured result is described. The hypergraph is not merely the right safety tool for customer service automation. It is the right mathematical object for thinking about what a customer service system is. References Bai et al. (2022) Yuntao Bai et al. Constitutional AI: Harmlessness from AI feedback. Technical report, Anthropic, 2022. arXiv:2212.08073. Basin et al. (2018) David Basin, Sören Debois, and Thomas Hildebrandt. On purpose and by necessity: Compliance under the GDPR. In Financial Cryptography and Data Security (FC 2018), pages 20–37, 2018. Benveniste et al. (2018) Albert Benveniste et al. Contracts for system design. Foundations and Trends in Electronic Design Automation, 12(2–3):124–400, 2018. Byun et al. (2005) Ji-Won Byun, Elisa Bertino, and Ninghui Li. Purpose based access control of complex data for privacy protection. pages 102–110, 2005. Colombo et al. (2012) Christian Colombo, Gordon J. Pace, and Gerardo Schneider. LARVA — safer monitoring of real-time Java programs. Science of Computer Programming, 77(11):1234–1260, 2012. Ekfrazo Technologies (2026) Ekfrazo Technologies. AI workflow automation in 2026: How agents replace manual ops. Technical report, Ekfrazo Research, March 2026. European Parliament and Council (2016) European Parliament and Council. Regulation (eu) 2016/679 (general data protection regulation). Official Journal of the European Union, L 119:1–88, 2016. European Parliament and Council (2024) European Parliament and Council. Regulation (eu) 2024/1689 on artificial intelligence (eu ai act). Technical report, Official Journal of the European Union, 2024. Ferraiolo et al. (2001) David F. Ferraiolo, Ravi Sandhu, Serban Gavrila, D. Richard Kuhn, and Ramaswamy Chandramouli. Proposed NIST standard for role-based access control. ACM Transactions on Information and System Security, 4(3):224–274, 2001. Fini Labs Research (2025) Fini Labs Research. Salesforce study finds LLM agents fail 65% of CX tasks. Technical report, Fini Labs, May 2025. Gitnux Research (2026) Gitnux Research. AI in the telecom industry statistics: Market data report 2026. Technical report, Gitnux, February 2026. Greshake et al. (2023) Kai Greshake et al. Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injections. arXiv preprint, arXiv:2302.12173, 2023. Jones (1983) Cliff B. Jones. Tentative steps toward a development method for interfering programs. ACM Transactions on Programming Languages and Systems, 5(4):596–619, 1983. Lamport (2002) Leslie Lamport. Specifying Systems: The TLA+ Language and Tools for Hardware and Software Engineers. Addison-Wesley, 2002. Nemhauser et al. (1978) George L. Nemhauser, Laurence A. Wolsey, and Marshall L. Fisher. An analysis of approximations for maximizing submodular set functions. Mathematical Programming, 14(1):265–294, 1978. PCI Security Standards Council (2022) PCI Security Standards Council. PCI DSS v4.0: Payment card industry data security standard. Technical report, PCI Security Standards Council, 2022. Perez and Ribeiro (2022) Fábio Perez and Ian Ribeiro. Ignore previous prompt: Attack techniques for language models. arXiv preprint, arXiv:2211.09527, 2022. Spera (2026) Cosimo Spera. Safety is non-compositional: A formal framework for capability-based AI systems. arXiv preprint, arXiv:2603.15973, March 2026. Revised version 4. Wu et al. (2023) Qingyun Wu et al. AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv preprint, arXiv:2308.08155, 2023. Appendix A Deployment Roadmap Phase 1: Hypergraph Construction (Weeks 1–4) Apply the PAC-learning algorithm to the most recent 90 days of session logs. Compute ℬ(F)B(F) offline and store. Deliverable: validated capability hypergraph; Safe Audit Surface per agent configuration. Phase 2: Pre-Execution Gate (Weeks 5–8) Deploy coalition safety check at session start (O(|ℬ(F)|⋅∑i|Ai|)<1O(|B(F)|· _i|A_i|)<1 ms). Impact: 5–10p automation gain; compliance-enabled service expansion. Phase 3: Living Audit Surface (Weeks 9–16) Integrate incremental maintenance into tool governance. Each new API integration triggers an O(n+mk)=O(24)O(n+mk)=O(24) safety check. Impact: replaces $300K–$1.25M/year manual compliance. Phase 4: NMF Pilot Study (Weeks 9–14, concurrent with Phase 3) Deploy NMF boundary signals to 10,000 sessions. Measure NMF identification accuracy against 90-day churn ground truth. This converts the highest-uncertainty model assumption to an empirical result. Impact: validates or revises the churn retention estimate. Phase 5: Submodular Goal Orchestration (Weeks 17–24) Deploy closure-based goal discovery; run 50/50 A/B test against rule-based upsell. Impact: activates the $403K–$605K annual emergent upsell mechanism; provides first empirical measurement of the upsell conversion uplift assumption.