Paper deep dive
High-Fidelity Network Management for Federated AI-as-a-Service: Cross-Domain Orchestration
Mohaned Chraiti, Ozgur Ercetin, Merve Saimler
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 7/21/2026, 4:14:32 AM
Summary
This paper proposes a Tail-Risk Envelope (TRE) framework for high-fidelity, cross-domain orchestration of AI-as-a-Service (AIaaS). It introduces TREs as signed, composable per-domain contracts that combine deterministic guardrails with stochastic impairment models to enforce end-to-end tail-latency guarantees (e.g., p99.9) in federated multi-domain environments. The approach uses stochastic network calculus to derive delay-violation bounds, enabling risk-budget decomposition, admission control, and tenant isolation without exposing internal network states.
Entities (6)
Relation Signals (5)
Tail-Risk Envelope → enables → AI-as-a-Service
confidence 92% · This paper introduces an assurance-oriented AIaaS management plane based on Tail-Risk Envelopes (TREs)
Tail-Risk Envelope → composes → End-to-End Delay Violation
confidence 90% · signed, composable per-domain descriptors... derive bounds on end-to-end delay violation probabilities across tandem domains
Stochastic Network Calculus → derivesboundsfor → End-to-End Delay Violation
confidence 88% · Using stochastic network calculus, we derive bounds on end-to-end delay violation probabilities
Communication Service Provider → provides → AI-as-a-Service
confidence 85% · communication service providers (CSPs) are on the verge of a radical transformation... to AIaaS a managed network service
Tail-Risk Envelope → supports → p99.9 Compliance
confidence 82% · Packet-level Monte-Carlo simulations demonstrate improved p99.9 compliance under overload via admission control
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:To support the emergence of AI-as-a-Service (AIaaS), communication service providers (CSPs) are on the verge of a radical transformation-from pure connectivity providers to AIaaS a managed network service (control-and-orchestration plane that exposes AI models). In this model, the CSP is responsible not only for transport/communications, but also for intent-to-model resolution and joint network-compute orchestration, i.e., reliable and timely end-to-end delivery. The resulting end-to-end AIaaS service thus becomes governed by communications impairments (delay, loss) and inference impairments (latency, error). A central open problem is an operational AIaaS control-and-orchestration framework that enforces high fidelity, particularly under multi-domain federation. This paper introduces an assurance-oriented AIaaS management plane based on Tail-Risk Envelopes (TREs): signed, composable per-domain descriptors that combine deterministic guardrails with stochastic rate-latency-impairment models. Using stochastic network calculus, we derive bounds on end-to-end delay violation probabilities across tandem domains and obtain an optimization-ready risk-budget decomposition. We show that tenant-level reservations prevent bursty traffic from inflating tail latency under TRE contracts. An auditing layer then uses runtime telemetry to estimate extreme-percentile performance, quantify uncertainty, and attribute tail-risk to each domain for accountability. Packet-level Monte-Carlo simulations demonstrate improved p99.9 compliance under overload via admission control and robust tenant isolation under correlated burstiness.
Tags
Links
- Source: https://arxiv.org/abs/2602.15281v2
- Canonical: https://arxiv.org/abs/2602.15281v2
Trouble viewing inline? Open PDF directly →
Full Text
48,956 characters extracted from source content.
Expand or collapse full text
High-Fidelity Network Management for Federated AI-as-a-Service: Cross-Domain Orchestration Mohaned Chraiti1, Ozgur Ercetin1, and Merve Saimler2 1Electronics Engineering Department, Sabancı University 2 Ericsson Research, Türkiye Emails: mohaned.chraiti@sabanciuniv.edu, oercetin@sabanciuniv.edu, and merve.saimler@ericsson.com Abstract To support the emergence of AI-as-a-Service (AIaaS), communication service providers (CSPs) are on the verge of a radical transformation—from pure connectivity providers to AIaaS a managed network service (control-and-orchestration plane that exposes AI models). In this model, the CSP is responsible not only for transport/communications, but also for intent-to-model resolution and joint network–compute orchestration, i.e., reliable and timely end-to-end delivery. The resulting end-to-end AIaaS service thus becomes governed by communications impairments (delay, loss) and inference impairments (latency, error). A central open problem is an operational AIaaS control-and-orchestration framework that enforces high fidelity, particularly under multi-domain federation. This paper introduces an assurance-oriented AIaaS management plane based on Tail-Risk Envelopes (TREs): signed, composable per-domain descriptors that combine deterministic guardrails with stochastic rate–latency–impairment models. Using stochastic network calculus, we derive bounds on end-to-end delay violation probabilities across tandem domains and obtain an optimization-ready risk-budget decomposition. We show that tenant-level reservations prevent bursty traffic from inflating tail latency under TRE contracts. An auditing layer then uses runtime telemetry to estimate extreme-percentile performance, quantify uncertainty, and attribute tail-risk to each domain for accountability. Packet-level Monte-Carlo simulations demonstrate improved p99.9 compliance under overload via admission control and robust tenant isolation under correlated burstiness. I Introduction I-A Motivations Artificial intelligence is advancing at an unprecedented rate and, over the past three years, has transitioned from a specialized tool employed primarily by experts to a widely accessible, everyday technology—a shift driven predominantly by large language model (LLM) platforms. This acceleration is catalyzing the emergence of AI-as-a-Service (AIaaS) and, with it, an ongoing central leap in the sixth generation networks (6G) services: the communication service provider (CSP) is on the edge of evolving from a connectivity-only operator into a AI control-and-orchestration provider that exposes AI capabilities as a managed network service [17, 19, 18]. This evolution is consistent with the AI-native network paradigm, in which AI capabilities are intrinsically embedded in network functions and exposed as consumable services rather than treated as external overlays. In this extended role, the operator does not merely transport traffic between AI model providers and end users; it performs intent-based service resolution, mapping user intents to an appropriate model and execution option, and orchestrates joint network and compute resources to meet contracted service targets. As a result, the end user is no longer bound to a specific, pre-known model endpoint; instead, the user specifies the task requirements, and the network selects the model and provisions the necessary communications and computational resources. The delivered service quality is inherently end-to-end, coupling transport impairments (delay, loss) with inference impairments (latency, model-error/quality degradation). Furthermore, the AIaaS orchestration leverages network-native intelligence, data, and Quality-of-Service (QoS) control that are unavailable to over-the-top AI providers. Consequently, enforceability must be addressed at the Control-and-Orchestration layer as a single service obligation. I-B Related Work Although AI model providers have rapidly matured the AIaaS supply stack, the network exposure, control, and intent-driven orchestration of AIaaS remains at an infant stage. Nonetheless, research on contiguous topics such AI-native networking and zero-touch management architectures provide useful principles and interfaces, but they do not yet deliver an operational framework for enforceable end-to-end guarantees and multi-domain coordination. AI-native network visions emphasize ubiquitous intelligence, distributed data infrastructure, and zero-touch operations, and advocate exposing AI capabilities and network-native services through Application-Programming-Interface (APIs) [19, 10, 2, 1, 14]. From an AI-native perspective, the APIs are enablers, with AIaaS fundamentally relying on Machine Learning Operations (MLOps) as-a-Service to operationalize lifecycle management, monitoring, and assurance within the network [19, 1, 14]. However, this API-centric view leaves a critical operational gap: how a CSP integrates exposed AIaaS features into a management and orchestration framework that is enforceable in the field [19, 11]. APIs specify how services are requested, but they do not specify how the network can make verifiable commitments about end-to-end performance when both transport and inference contribute to experienced latency and reliability. Two requirements dominate AIaaS at scale. First, AIaaS contracts are naturally tail contracts: for interactive inference and agentic pipelines, Service Level Objectives (SLOs) are expressed as extreme-percentile deadlines (e.g., p99/p99.9 end-to-end latency), not average latency [7]. Average-QoS engineering is structurally insufficient because rare events arise from burst synchronization, shared bottlenecks, scheduler transients, radio fades, and compute contention; moreover, these tail events may occur in any segment of the pipeline (access/transport, execution, or inter-domain traversal) while the client observes a single end-to-end SLO [7]. Second, AIaaS requires federation. Coverage, locality, and developer reach push AIaaS beyond a single operator domain, creating multi-principal objectives, confidentiality constraints (no disclosure of internal topology, queue states, or scheduler rules), and an accountability requirement where disputes and penalties must be supported by auditable evidence [11, 15]. Without an explicit representation of tail risk that composes across domains, an AIaaS orchestrator has only two failure modes: over-provisioning (economically infeasible) or tail violations (trust-destroying) [7]. I-C Problem Statement This paper targets the missing architectural primitive: a composable tail-risk contract that (i) abstracts each domain as a signed service descriptor, (i) composes across tandem domains to yield end-to-end p99/p99.9 guarantees, (i) enables risk-budget decomposition and optimization under confidentiality constraints, and (iv) supports audit and settlement without requiring internal disclosure. The core contribution is therefore not a new API surface; it is an assurance layer for AIaaS management and orchestration. This assurance layer complements ongoing AIaaS API standardization efforts by addressing enforceability and accountability rather than service invocation semantics. I-D Contributions The paper makes four technical contributions that together operationalize federated AIaaS as an enforceable managed service. • We introduce Tail-Risk Envelopes (TREs): signed, per-domain contracts for tail-latency assurance that are composable across domains with minimal disclosure. • We derive tractable bounds on per-domain and end-to-end delay-violation probabilities in tandem domains, enabling p99/p99.9 feasibility checks for admission and reservation. • We formulate federated provisioning as a risk-budgeted optimization that selects paths, decomposes end-to-end tail budgets across domains, and reserves per-tenant capacity, with a proof of burst-resilient isolation. • We develop a telemetry-driven auditing layer that estimates extreme percentiles with confidence intervals, updates TRE uncertainty when tails shift, and attributes tail-risk across domains for accountability. The developed approach is designed to integrate with standardized AIaaS exposure and orchestration frameworks that enable deployment within existing and emerging 3GPP- and industry-aligned architectures. I Contract Interface for Federated AIaaS I-A Federated AIaaS execution pipeline and service contract An AIaaS request is realized as a multi-stage execution pipeline that spans both communication and computation: radio access, transport/backhaul, edge or cloud execution, and (when needed) inter-operator transit. The key operational characteristic is federation: at one or more stages, the execution domain can be selected from multiple administrative entities (e.g., alternative edge providers or partner operators), subject to locality, trust, and compliance constraints, consistent with global API aggregation visions [9, 12, 14, 1]. This federated execution model directly supports multi-CSP AIaaS exposure and aggregation scenarios. Pipeline stages and architecture We represent the AIaaS path as L ordered stages indexed by ℓ∈1,…,L ∈\1,…,L\. Each stage ℓ offers a set of feasible domains ℓ(Ωu,k)D_ ( _u,k) determined by policy and tenant constraints (e.g., allowed operators, locality region, trust level), where Ωu,k _u,k is defined below. A concrete execution path is therefore a sequence of selected domains π≜(d1,…,dL),dℓ∈ℓ(Ωu,k).π (d_1,…,d_L), d_ _ ( _u,k). (1) A domain d may represent a Radio-Access-Network (RAN) segment, a transport segment, an edge cluster, a core cloud, or an inter-operator transit segment. This distinction (stages vs. selectable domains per stage) makes federation explicit and separates where choice exists from where the pipeline is structurally fixed. Fig. 1 summarizes the operational control loop. A CSP hosts an AIaaS control-and-orchestration (C&O) function that (i) receives intent-level requests through exposure APIs, (i) resolves intent-to-model and execution placement, (i) provisions joint network and compute resources along a selected federated path, and (iv) audits delivered performance for accountability and settlement. In practice, intent intake and service invocation are realized through standardized network exposure mechanisms, while the assurance logic operates below the API layer. This aligns with AI-native and zero-touch management directions [19, 10, 11], and is compatible with existing exposure/API frameworks (e.g., NEF/CAPIF/CAMARA) [2, 1, 14]. The paper’s focus is the missing assurance layer: how to make end-to-end tail guarantees enforceable and composable across domains without requiring disclosure of internal scheduler or queue state. Application DomainCSP Domain (Federated AIaaS)HCP)- DomainExternal App / TenantIntent + ConstraintsSLO(τ,ε)(τ, ), Ω,(κ,φ) ,(κ, ) Exposure APINEF / CAPIF / CAMARA AIaaS C&O PlaneIntent → Placement SelectionFederated Path π & Risk Budget Domain 1: TREDomain 2: TREDomain 3: TREFederated Execution Pipeline (Stages ℓ=1,…,L =1,…,L)Multi-stage Radio / Transport / Compute selection (dℓ∈ℓd_ _ ) Telemetry & AuditEVT Tail Risk FitSettlement Logic HCP ApplicationsData Logic MLOps ToolsetData Network Resource EnforcementQoS Flows / SteeringTenant Isolation Feedback Loop Figure 1: Federated AIaaS execution pipeline and assurance loop. The CSP AIaaS receives intents, selects model/placement, and provisions joint network–compute resources across a multi-stage path. Each administrative domain, including the Hyperscale Cloud Provider (HCP), exposes a TRE contract (signed, composable) rather than raw internal state. Telemetry feeds an audit layer for p99/p99.9 calibration. Tenant-level AI service-level objective We formalize the tenant contract as an SLO centered on tail latency and correctness constraints. For tenant u and inference class k, define SLOu,k=(τu,k,εu,k,Ωu,k,κu,k,φu,k),SLO_u,k= ( _u,k, _u,k, _u,k, _u,k, _u,k ), (2) where: (i) τu,k _u,k is the end-to-end deadline; (i) εu,k _u,k is the maximum allowed violation probability (e.g., εu,k=10−3 _u,k=10^-3 corresponds to p99.9); (i) Ωu,k _u,k encodes locality/trust/compliance constraints and the admissible operator/domain sets; (iv) κu,k _u,k encodes accuracy/quality constraints (e.g., minimum task success probability or maximum error); and (v) φu,k _u,k encodes freshness constraints (e.g., maximum staleness) managed via model life-cycle and data APIs [19]. The motivation for tail SLOs is practical: user-perceived reliability is dominated by rare slow events in large-scale services, so average-QoS engineering is insufficient [7]. Let Au(s,t)A_u(s,t) denote the cumulative AI workload arriving for tenant u over [s,t)[s,t), measured in a normalized unit (e.g., inference jobs or equivalent compute quanta). Along a selected path π, each stage ℓ provides a service process Sdℓ(s,t)S_d_ (s,t) in the same unit. The resulting (virtual) sojourn-time delay Wu,kW_u,k is induced by the interaction of arrivals and the (min-plus) concatenation of per-stage services. The management objective is to enforce the tail-latency constraint ℙWu,k>τu,k≤εu,k,P\W_u,k> _u,k\≤ _u,k, (3) under multi-tenant contention, domain uncertainty, and federation. The expression in (3) is the central operational requirement: it is the mathematical form of “p99/p99.9 end-to-end guarantee”. Remark 1. Violations of (3) can originate anywhere in the chain (radio, transport, compute, inter-domain), yet the user observes a single end-to-end SLO. Moreover, in a federated setting domains will not reveal internal queue lengths, scheduler rules, or topology. This motivates a contract interface that is both composable and confidentiality-preserving. I-B TREs as a composable per-domain contract We introduce the TRE as the operational interface between a domain and the AIaaS orchestrator. A TRE is the domain’s published service contract at a given reservation level (e.g., slice class, priority class, CPU share): it is (i) signed for accountability, (i) composable across tandem stages for end-to-end reasoning, and (i) minimal so it can be disclosed without exposing internal state. TREs are not exposed to application developers. They serve as internal, contract-level abstractions that enable enforceable AIaaS orchestration across administrative boundaries while remaining compatible with standardized AIaaS exposure frameworks. Definition 1 (Tail-Risk Envelope). For domain d and tilting parameter θ>0θ>0, the TRE is TREd(θ)=(Rd,Td,κd,ηd),TRE_d(θ)= (R_d,T_d, _d, _d ), (4) interpreted as a stochastic rate–latency service: Sd(s,t)≥Rd[(t−s)−Td]+−Id(s,t),∀t≥s,S_d(s,t)\;≥\;R_d [(t-s)-T_d ]^+-I_d(s,t), ∀\,t≥ s, (5) where RdR_d and TdT_d are deterministic guardrail parameters and the impairment process IdI_d satisfies the Moment Generating Function (MGF) bound [eθId(s,t)]≤exp(θκd(t−s)+θηd),∀t≥s.E\! [e^θ I_d(s,t) ]≤ \! (θ _d(t-s)+θ _d ), ∀\,t≥ s. (6) When Id≡0I_d≡ 0, (5) reduces to the classical rate–latency service curve. The impairment term captures residual uncertainty beyond the deterministic guardrail—e.g., bursty cross-traffic, scheduler jitter, radio/transport randomness, compute contention, and model-execution variability—in a form that yields exponential tail bounds when combined with MGF-bounded arrivals. Operationally, a domain publishes a family of TREs indexed by reservation level (priority/slice/CPU share), and the orchestrator selects (domain, reservation) pairs so that the composed end-to-end service satisfies (3). Crucially, the orchestrator does not require raw internal state (queue occupancy, per-flow scheduling details); it only requires the TRE parameters, which are contract-level quantities. Federation requires that domains can coordinate without disclosing internals. A TRE provides exactly what is needed for orchestration: (i) a deterministic baseline (Rd,TdR_d,T_d) that can be enforced by admission and reservation, (i) a stochastic uncertainty term (κd,ηd _d, _d) that quantifies tail risk, and (i) a parameterization that is composable across stages and amenable to optimization. Later sections show how TRE composition yields end-to-end violation bounds and how tail-risk budgets can be decomposed across domains for federated provisioning. Considering a standard MGF-style arrival envelope yields optimization-ready tail bounds while supporting admission and isolation constraints: tenant-u arrivals satisfy, for some θ>0θ>0 and parameters (ρu,σu)( _u, _u), [eθAu(s,t)]≤exp(θρu(t−s)+θσu),∀t≥s,E\! [e^θ A_u(s,t) ]≤ \! (θ _u(t-s)+θ _u ), ∀\,t≥ s, (7) which is consistent with effective-bandwidth formulations [5]. Here, ρu _u captures sustained load and σu _u captures burstiness. One can hence derive an exponential tail bounds via stochastic network calculus, yet structured enough for (i) admission control, (i) per-tenant reservation with isolation guarantees, and (i) federated provisioning where domains exchange compact contract parameters rather than raw traces. In subsequent analysis, (7) is combined with TRE relations (4)–(6) to derive end-to-end delay-violation bounds for (3) and construct cross-domain risk-budget decompositions. I Compositional Tail Guarantees This section aims to transform the published TRE parameters into enforceable p99/p99.9 guarantees. The key principle is that a domain does not expose internal queue states or scheduler rules; instead, it exposes a TRE contract TREd(θ)=(Rd,Td,κd,ηd)TRE_d(θ)=(R_d,T_d, _d, _d). Given tenant arrivals satisfying the (σ,ρ)(σ,ρ)-MGF constraint, stochastic network calculus yields an explicit exponential bound on delay-violation probability. This bound is (i) compositional across tandem stages, (i) optimization-ready as a chance constraint, and (i) federation-friendly because it depends only on contract parameters. I-A Single-domain tail bound under stochastic rate–latency Consider tenant u traversing domain d. Recall the impaired service model in (5)–(6) and the arrival constraint in (7). Define the net margin Δu,d≜Rd−ρu−κd. _u,d R_d- _u- _d. (8) The condition Δu,d>0 _u,d>0 is the feasibility threshold: if the sustained load ρu _u plus the impairment slope κd _d meets or exceeds the reserved rate RdR_d, exponential tail decay cannot be guaranteed. From (5), we have −Sd(s,t)≤−Rd[(t−s)−Td]++Id(s,t).-S_d(s,t)≤-R_d[(t-s)-T_d]^++I_d(s,t). Applying the exponential operator on both sides and taking expectation, [e−θSd(s,t)] \! [e^-θ S_d(s,t) ] (9) ≤exp(−θRd[(t−s)−Td]+)[eθId(s,t)] ≤ \! (-θ R_d[(t-s)-T_d]^+ )\;E\! [e^θ I_d(s,t) ] ≤exp(−θRd[(t−s)−Td]++θκd(t−s)+θηd), ≤ \! (-θ R_d[(t-s)-T_d]^++θ _d(t-s)+θ _d ), where the last inequality stems from (6) for ∀t≥s∀\,t≥ s. We assume that each domain provides FIFO-equivalent service at the contract level, which is required for backlog–delay equivalence and is standard in stochastic network calculus. Time is treated in discrete or discretized form for bounding purposes, consistent with the union-bound arguments employed below. Theorem 1 (Single-domain delay-violation bound). Under (7) and (9), if Δu,d>0 _u,d>0, then for any deadline τ≥Tdτ≥ T_d, ℙWu(d)>τ \W_u^(d)>τ\ (10) ≤exp(θ(σu+ηd+κdTd))1−exp(−θΔu,d)⏟Prefactor: burstiness & uncertainty×exp(−θΔu,d(τ−Td))⏟Tail-decay term. ≤\; \! (θ( _u+ _d+ _dT_d) )1- \! (-θ _u,d )_Prefactor: burstiness \& uncertainty\;×\; \! (-θ _u,d(τ-T_d) )_Tail-decay term. Proof. Let B(t)B(t) denote backlog at time t in a First-IN First-Out (FIFO) system, and W(t)W(t) the virtual delay. Standard Stochastic Network Calculus (SNC) arguments relate delay to backlog: W(t)>τW(t)>τ implies that backlog exceeds the service that can be provided over τ slots. Formally, in a stable FIFO system, we have W(t)>τ⊆sup0≤s≤t(A(s,t)−S(s,t))>R(τ−T)+.\W(t)>τ\ \ _0≤ s≤ t (A(s,t)-S(s,t) )>R(τ-T)^+ \. (11) Applying Chernoff’s bound gives: ℙA(s,t)−S(s,t)>x \A(s,t)-S(s,t)>x\ ≤e−θx[eθA(s,t)][e−θS(s,t)]. ≤ e^-θ xE[e^θ A(s,t)]E[e^-θ S(s,t)]. (12) Using (7) and (9), for τ≥Tdτ≥ T_d and t−s=ℓt-s= , we have [eθA(s,t)]≤eθρuℓ+θσu[e−θSd(s,t)]≤e−θRd(ℓ−Td)+θκdℓ+θηd, casesE[e^θ A(s,t)]≤ e^θ _u +θ _u\\ E[e^-θ S_d(s,t)]≤ e^-θ R_d( -T_d)+θ _d +θ _d cases, (13) which gives ℙA(s,t)−S(s,t)>x \A(s,t)-S(s,t)>x\ (14) ≤exp(−θx+θσu+θηd+θ(ρu+κd−Rd)ℓ+θRdTd). ≤ \! (-θ x+θ _u+θ _d+θ( _u+ _d-R_d) +θ R_dT_d ). A union bound is then appied to all feasible durations ℓ=t−s∈0,1,2,… =t-s∈\0,1,2,…\ and set x=Rd(τ−Td)x=R_d(τ-T_d) to obtain a geometric series with ratio e−θΔu,de^-θ _u,d, where Δu,d=Rd−ρu−κd _u,d=R_d- _u- _d. Summing the series yields (10). ∎ The bound in (10) naturally separates into two factors: a prefactor and an exponential tail-decay term. The prefactor captures the “baseline difficulty” of meeting the SLO: it increases with the tenant burstiness σu _u and with the domain’s uncertainty offset ηd _d, both of which make rare p99/p99.9 excursions more likely. The exponential term captures how fast the violation probability decays when the system has slack: its exponent θΔu,d(τ−Td)θ _u,d(τ-T_d) grows with the available service margin Δu,d _u,d and with the deadline slack (τ−Td)(τ-T_d). Hence, reserving more service rate RdR_d increases Δu,d _u,d and yields an exponential reduction of tail violations, making the bound directly actionable for admission and reservation. I-B End-to-end composition across an AIaaS pipeline: from per-domain TREs to a path-level bound An AIaaS request traverses a sequence of stages (radio/transport/compute), so the end-to-end delay is determined by the slowest cumulative progress along the chain. This reflects the AIaaS model, where both communication and AI execution stages jointly shape user-perceived performance. Intuitively, stage ℓ+1 +1 cannot start serving a job before stage ℓ has produced it; therefore, end-to-end service is governed by the most constraining handoff between stages. Network calculus formalizes this tandem coupling via the min-plus composition rule [3]. For deterministic rate–latency stages, this rule yields two simple engineering consequences: (i) latencies accumulate across stages, and (i) the effective service rate is limited by the bottleneck stage. In the considered setting, each domain d exposes a TRE at tilting parameter θ, TREd(θ)=(Rd,Td,κd,ηd)TRE_d(θ)=(R_d,T_d, _d, _d), which lower-bounds its offered service by a rate–latency guardrail minus an impairment term. When multiple domains are concatenated, a conservative (contract-safe) composition treats the impairments as additive across stages; this preserves enforceability without any disclosure of internal queue states or scheduler rules. Let the selected path be π=(d1,…,dL)π=(d_1,…,d_L). We aggregate the published contract parameters into path-level descriptors: Rmin≜minℓ∈1,…,LRdℓ,TΣ≜∑ℓ=1LTdℓ, R_ _ ∈\1,…,L\R_d_ , T_ _ =1^LT_d_ , (15) κΣ≜∑ℓ=1Lκdℓ,ηΣ≜∑ℓ=1L(ηdℓ+κdℓTdℓ). _ _ =1^L _d_ , _ _ =1^L ( _d_ + _d_ T_d_ ). The interpretation is direct: TΣT_ is the accumulated “pipeline latency floor”; RminR_ is the bottleneck reserved rate; and (κΣ,ηΣ)( _ , _ ) quantify the accumulated tail uncertainty induced by network/compute impairments across domains. Combining these aggregates with the tenant arrival descriptor (ρu,σu)( _u, _u) yields the end-to-end feasibility margin Δu≜Rmin−ρu−κΣ. _u R_ - _u- _ . (16) A strictly positive margin Δu>0 _u>0 is necessary: if the sustained load ρu _u plus the aggregate impairment slope κΣ _ meets or exceeds the bottleneck reserved rate RminR_ , then no exponential tail guarantee can be maintained. With Δu>0 _u>0, substituting the path-level descriptors into the single-domain Chernoff/SNC steps yields an explicit bound on end-to-end delay violation. For any deadline τ≥TΣτ≥ T_ , ℙWu>τ≤exp(θ(σu+ηΣ))1−exp(−θΔu)⏟burst & uncertainty inflationexp(−θΔu(τ−TΣ))⏟end-to-end tail decay. \W_u>τ\≤ \! (θ( _u+ _ ) )1- \! (-θ _u )_burst \& uncertainty inflation\; \! (-θ _u(τ-T_ ) )_end-to-end tail decay. (17) The bound in (17) shows that a p99/p99.9-type end-to-end guarantee can be verified using only (i) the tenant traffic descriptor (ρu,σu)( _u, _u) and (i) the per-domain published TRE parameters along π. This property is critical for cross-operator federation, where internal information cannot be shared. No internal topology, queue lengths, or scheduler rules are required, which is the abstraction needed for cross-operator federation. Given a tail SLO (τu,εu)( _u, _u) (e.g., εu=10−3 _u=10^-3 for p99.9), a sufficient feasibility condition from (17) is θΔu(τu−TΣ)≥log1εu+θ(σu+ηΣ)+log11−e−θΔu. θ _u( _u-T_ )≥ \! 1 _u+θ( _u+ _ )+ \! 11-e^-θ _u. (18) The right-hand side is a risk requirement: it aggregates the demanded reliability level log(1/ϵu) \! (1/ _u ), the burst/uncertainty inflation term θ(σu+ηΣ)θ ( _u+ _ ), and the finite-margin correction log(1/(1−e−θΔu)) \! (1/ (1-e^-θ _u ) ). The left-hand side is a risk capacity: it scales with the effective slack Δu _u and the deadline slack (τu−TΣ)( _u-T_ ). This form is directly actionable for orchestration: increasing the bottleneck reservation RminR_ increases Δu _u; reducing accumulated latency TΣT_ increases (τu−TΣ)( _u-T_ ); and telemetry-driven auditing primarily updates ηΣ _ , tightening feasibility until additional reservation or rerouting is applied. IV Federated Provisioning and Assurance Section I yields explicit p99/p99.9 tail-violation bounds from published TREs and tenant traffic descriptors. This section builds the corresponding operational management plane for multi-domain and multi-operator AIaaS: the orchestrator selects a federated execution path, allocates a tail-risk budget across domains, reserves joint network–compute resources with strict multi-tenant isolation, and coordinates operators without disclosing internal queue states or scheduler rules. Assurance is then closed by an auditing and settlement loop that (i) estimates extreme percentiles from telemetry using Extreme Value Theory (EVT), (i) conservatively updates TRE uncertainty terms when the tail regime shifts, and (i) attributes penalties and revenue to domains according to their marginal contribution to end-to-end tail risk. The central principle is end-to-end enforceability: the client observes a single SLO, while the federation operates through signed, composable contracts rather than internal-state disclosure. IV-A Federated provisioning under TRE contracts A tenant contract specifies (τu,εu)( _u, _u) together with policy constraints Ω (allowed operators/locations/trust), and the orchestrator selects a path π=(d1,…,dL)π=(d_1,…,d_L) through radio/transport/compute stages. Since the guarantee is end-to-end, a practical management primitive is to decompose the violation budget across the domains of the selected path ∑ℓ=1Lεu,dℓ≤εu _ =1^L _u,d_ ≤ _u. Each domain then enforces a local tail constraint of the form (10) with its assigned varepsilonu,dℓvarepsilon_u,d_ . This decomposition provides (i) per-domain admission and accountability and (i) a cost-aware lever to shape tail risk by tightening budgets where reliability is scarce or expensive. Tail guarantees are fragile under bursty cross-traffic unless isolation is explicit. We therefore enforce per-tenant effective-bandwidth partitioning within each domain. Let R¯d,u R_d,u denote the reserved service share assigned to tenant u in domain d (e.g., a slice class, QoS-flow reservation, or computational resource share mapped to an equivalent service rate). Isolation is enforced by ρu≤R¯d,u−κd,∑uR¯d,u≤Rd, _u≤ R_d,u- _d, _u R_d,u≤ R_d, (19) which preserves a strictly positive margin R¯d,u−ρu−κd R_d,u- _u- _d for each tenant independent of other tenants’ burst parameters. Federated provisioning is posed as a joint selection and reservation problem. Each domain d publishes (i) TRE parameters (Td,κd,ηd)(T_d, _d, _d) for each reservation class, (i) a cost curve cd(R¯d,u)c_d( R_d,u) for rate/compute reservations (or a mapping from CPU/GPU shares to an equivalent RdR_d), and (i) admissibility constraints encoded by Ω . The orchestrator selects a path π, reservations R¯d,u\ R_d,u\, and risk budgets εu,d\ _u,d\ to minimize total cost while satisfying tail constraints, budgets, and isolation: minimizeπ,R¯d,u,εu,d π,\ R_d,u\,\ _u,d\minimize ∑d∈π∑ucd(R¯d,u) _d∈π _uc_d( R_d,u) (20) subject to ℙWu(d)>τu≤εu,d,∑d∈πεu,d≤εu, \W_u^(d)> _u\≤ _u,d,\,\, _d∈π _u,d≤ _u, ρu≤R¯d,u−κd,∑uR¯d,u≤Rd. _u≤ R_d,u- _d,\, _u R_d,u≤ R_d. The decisive point is that the feasible region is tail-risk geometry induced by TRE contracts and arrival descriptors, not mean-rate QoS. Consequently, feasibility and reservation sizing can be evaluated from published contracts without internal queue states or scheduler disclosure. Federation further imposes confidentiality: a broker cannot request internal operator state. Therefore, coordinate provisioning can be done via Alternating Direction Method of Multipliers (ADMM) [4]. Each domain solves a private subproblem using only its own cost and capacity constraints; the orchestrator updates risk budgets and consensus variables. For online operation, the ADMM loop runs on a slower control timescale (seconds–minutes) while domains enforce the resulting reservations on fast schedulers; coupling the slow loop with drift-plus-penalty mechanisms yields stability and near-optimality guarantees under stochastic arrivals [16]. The novelty here is that the coupling constraints are TRE-induced p99/p99.9 chance constraints rather than average-delay objectives. The optimization in (20) is implemented by Algorithm 1, which performs distributed broker–domain updates with privacy-preserving local solves. Algorithm 1 Federated provisioning with ADMM and EVT assurance 1: Input: feasible paths Ω , TREs, domain costs cd(⋅)c_d(·), SLOs (τu,εu)( _u, _u) 2: Initialize allocations R¯d,u\ R_d,u\, risk budgets εu,d\ _u,d\, and dual variables 3: repeat 4: Domain step (parallel): each d solves its local ADMM subproblem for R¯d,u\ R_d,u\ 5: Broker step: update εu,d\ _u,d\ and consensus variables to satisfy (20) 6: Update dual variables and residuals 7: until primal/dual residuals are below tolerance or deadline is reached 8: Deploy reservations and start telemetry collection 9: Fit EVT tails from exceedances; if SLO violation risk increases, re-enter at line 3 IV-B Audit closure and settlement SNC bounds provide conservative contractual guarantees; operations additionally require auditing extreme percentiles from telemetry and updating uncertainty terms when regimes change (overload, failures, execution drift). We use EVT in a peak-over-threshold (POT) form. Such auditing is essential to maintain trust in AIaaS guarantees over time. Given observed end-to-end sojourn times Li\L_i\, choose a high threshold q (e.g., empirical p98), form exceedances Yi=Li−qY_i=L_i-q conditioned on Li>qL_i>q, and fit a generalized Pareto distribution (GPD) with shape ξ and scale β [6, 13, 8]. The POT estimator yields extreme quantiles QpQ_p for p>F(q)p>F(q), including p99 and p99.9, together with confidence intervals. The assurance plane uses EVT for two purposes. First, it performs compliance verification: audited Q0.999Q_0.999 is compared against the contract-implied risk from (17) (or its domain-level decompositions) to detect tail regressions and to distinguish transient violations from systematic drift. Second, it enables conservative contract updates: when EVT indicates heavier tails or increased variance, uncertainty parameters (primarily ηd _d, and when necessary κd _d) are increased so that (17) remains an upper bound at the target confidence level. Domains then re-issue updated, signed TREs for the relevant reservation classes, making the assurance layer self-correcting under nonstationarity while preserving the same contract interface. This enables CSPs to continuously refine AIaaS assurance without violating contractual abstractions. Finally, federation requires settlement aligned with tail guarantees. Byte-based or CPU-second-based revenue allocation is misaligned because tail risk may be dominated by a single domain’s uncertainty. Define the bound-implied risk score at the contract deadline: u(τu)≜−logℙ^Wu>τu,K_u( _u) - P\W_u> _u\, (21) where ℙ^⋅ P\·\ is the TRE-composed value in (17). A domain’s marginal tail-risk contribution is computed by the sensitivity of uK_u to its contract parameters (Rd,Td,κd,ηd)(R_d,T_d, _d, _d) at the deployed reservation. This yields interpretable attribution without internal disclosure. Let PuP_u denote the tenant payment and Πu _u the penalty upon an EVT-audited breach. We allocate revenue shares proportionally to positive marginal contributions (domains that increase uK_u) and allocate penalties proportionally to negative marginal contributions (domains whose degradation reduces uK_u). V Simulation Results We evaluate whether the proposed TRE-based management plane turns federated AIaaS into an enforceable managed service. The simulations target three questions: (i) can the system maintain tail-delay compliance under rising load, (i) can it provide tenant isolation under bursty/adversarial traffic, and (i) can it support interpretable accountability by attributing tail-risk increases to the domains responsible for degradation. V-A Setup We use packet-level Monte-Carlo simulation of FIFO queues. Across all experiments we report p99.9 delay (p=0.999p=0.999) against a deadline τ=30τ=30 (time units), and we compare best-effort operation (no contract-aware control) against TRE-managed operation (contract-aware admission/reservation). For reproducibility, we log the selected acceptance factor α at each offered load together with the resulting empirical Q0.999(D)Q_0.999(D) values. The end-to-end path is a tandem of D=3D=3 single-server domains. Arrivals to the first domain are Poisson with rate λ. Domain d provides exponential service with rate μd _d and includes a fixed processing/propagation shift TdT_d. End-to-end delay is the sum of per-domain sojourn times (waiting + service) plus the per-domain shifts. We use =[1.0, 1.15, 1.25],=[0.6, 0.5, 0.4], μ=[1.0,\,1.15,\,1.25], =[0.6,\,0.5,\,0.4], and we normalize offered load as ρ=λ/mindμdρ=λ/ _d _d. For each Monte-Carlo trial we simulate Npkt=6000N_pkt=6000 packets and compute the empirical tail quantile Q0.999(D)Q_0.999(D) of end-to-end delay using the sample quantile. Each operating point is averaged over NMCN_MC independent trials (user-adjustable in the simulator). Random seeds are fixed to ensure repeatability. We consider three scenarios. First, we sweep ρ over a grid of 10 points in [0.55, 0.98][0.55,\,0.98]. The best-effort baseline admits all traffic. TRE-managed operation applies admission control through an acceptance factor α∈(0,1]α∈(0,1] (admitted rate λadm=αλ _adm=αλ). For each offered load, α is chosen by a bisection search (18 iterations) to maximize admission while satisfying the conservative tail target Q0.999(D)≤0.985τ,Q_0.999(D)≤ 0.985\,τ, evaluated on a single packet-level run; the resulting α is then re-evaluated with Monte-Carlo averaging. Second, we evaluate isolation in a single shared bottleneck domain with service rate μ=mindμdμ= _d _d. Two tenants share the domain: a victim stream with Poisson arrivals at rate λv=0.55μ _v=0.55\,μ, and an attacker stream with fixed mean rate λa=0.12μ _a=0.12\,μ but increasing burstiness. Burstiness is controlled by a peak-to-mean ratio b∈[1,8]b∈[1,8] using a correlated ON/OFF arrival process: during ON periods the attacker rate is λon=bλa _on=b _a; during OFF periods it is λoff=λa/b _off= _a/b; ON/OFF periods alternate with mean run lengths that increase with b (longer bursts for larger b). We compare shared FIFO multiplexing (single queue, no isolation) against per-tenant reservation (separate queues with fixed service shares). In reservation mode, the victim receives a dedicated share sv=0.85s_v=0.85 (service rate μv=svμ _v=s_vμ) and the attacker receives μa=(1−sv)μ _a=(1-s_v)μ. Third, to simulate domain degradation, we scale service rates by a factor s∈[0.60,1.0]s∈[0.60,1.0] (6 points), where s=1s=1 is nominal and smaller values represent degraded service. We fix the offered load near the operating knee at λ=0.85mindμdλ=0.85\, _d _d. We report the increase in mean p99.9 delay ΔQ0.999(D)=[Q0.999(D)∣s]−[Q0.999(D)∣s=1]. Q_0.999(D)=E[Q_0.999(D)\! \!s]-E[Q_0.999(D)\! \!s\!=\!1]. To obtain marginal attribution, we compute (i) the total increase when all domains are degraded by s, and (i) the increase when only domain d is degraded by s; then the marginal contributions per-domain are normalized to sum up the total increase in each s. Table I summarizes the parameters used in the plots. TABLE I: Simulation parameters used in the reported experiments. Tandem domains D=3D=3 Service rates =[1.0, 1.15, 1.25] μ=[1.0,\,1.15,\,1.25] Fixed shifts =[0.6, 0.5, 0.4]T=[0.6,\,0.5,\,0.4] Tail percentile p=0.999p=0.999 (p99.9) Deadline τ=30τ=30 Packets per trial Npkt=6000N_pkt=6000 Monte-Carlo trials NMCN_MC (10000) Load grid ρ∈[0.55,0.98]ρ∈[0.55,0.98] (10 points) TRE guard factor 0.9850.985 Isolation: victim load λv=0.55μ _v=0.55\,μ Isolation: attacker mean λa=0.12μ _a=0.12\,μ Isolation: victim share sv=0.85s_v=0.85 Burstiness grid b∈[1,8]b∈[1,8] (8 points) Degradation grid s∈[0.60,1.0]s∈[0.60,1.0] (6 points) V-B Results Fig. 2 reports the estimated p99.9 end-to-end delay Q0.999(D)Q_0.999(D) versus normalized offered load ρ. Under best-effort operation, the tail quantile remains moderate at low-to-mid loads but rises sharply as ρ approaches saturation, quickly exceeding the deadline τ. This is the classical tail-amplification regime: once utilization becomes high, rare backlog events persist long enough to dominate extreme percentiles, so the service can look acceptable on average while violating p99.9 systematically. In contrast, TRE-managed operation avoids this cliff by adapting the admitted rate to preserve headroom. The curve staying close to τ indicates that the management plane is not merely improving mean latency; it is explicitly preventing tail runaway and thus maintaining a predictable worst-case user experience. Operationally, this supports the paper’s premise that federated AIaaS cannot be treated as best-effort transport: tail compliance requires an explicit control knob that trades throughput for reliability when the system is stressed. Fig. 3 evaluates robustness to adversarial burstiness. Under shared FIFO multiplexing, increasing attacker burstiness produces pronounced inflation of the victim’s p99.9 delay even though the attacker’s mean rate is held fixed. This is a key point for multi-tenant AIaaS: it is not the average load alone that determines tail compliance, but the burst structure of competing tenants. Correlated bursts create periods of sustained overload that force the victim to wait behind attacker backlog, producing a strong cross-tenant coupling at extreme percentiles. Under per-tenant reservation, the victim tail remains essentially flat over the full burstiness range, indicating that the victim’s tail behavior is governed primarily by its own traffic and reserved capacity rather than by attacker burst patterns. This directly supports the isolation claim: enforcing per-tenant shares prevents burst-induced tail inflation and converts a fragile shared-queue regime into a predictable service envelope, which is essential for offering contractual p99/p99.9 guarantees in a multi-tenant setting. Fig. 4 connects performance assurance to accountability. As domains are degraded (smaller service-rate scaling factor s), the increase in tail delay ΔQ0.999(D) Q_0.999(D) grows, confirming that extreme percentiles are highly sensitive to even moderate capacity loss. The marginal decomposition further shows that not all domains contribute equally to the tail increase: the most capacity-critical domain(s) account for a disproportionate share of ΔQ0.999(D) Q_0.999(D), consistent with the intuition that bottlenecks dominate end-to-end tail behavior. This provides two practical implications for federated operation. First, it enables audit-ready diagnostics: when tail violations occur, the attribution points to where degradation most strongly impacts the end-to-end objective. Second, it supports incentive-compatible federation: penalties or credits can be tied to each domain’s marginal tail-risk contribution rather than to coarse end-to-end metrics that unfairly blame all parties equally. In this sense, the figure demonstrates that the proposed assurance layer is not only predictive (bounding risk) but also operational (supporting accountability under multi-operator composition). Figure 2: Estimated p99.9 end-to-end delay versus normalized offered load ρ. Figure 3: Isolation under burstiness. Figure 4: Marginal tail-risk attribution under controlled degradation. VI Conclusion This paper proposes an assurance-oriented management and orchestration framework that makes federated AIaaS enforceable across multiple administrative domains. We introduced TREs as signed, composable per-domain contracts that support end-to-end p99/p99.9 reasoning while requiring only disclosure-minimal parameters. Building on these contracts, we develop a federated orchestration method that decomposes an end-to-end tail-risk budget into per-domain obligations, enabling confidentiality-preserving coordination and strict multi-tenant isolation. A telemetry-driven auditing layer continuously recalibrates tail models and supports accountability by attributing end-to-end tail-risk changes to individual domains. Simulations confirm improved p99.9 reliability under overload, robust isolation under bursty interference, and stable attribution under controlled multi-domain degradation. References [1] 3GPP (2024) Common API Framework for 3GPP Northbound APIs (CAPIF). Technical report Technical Report 3GPP TS 23.222 (ETSI TS 123 222), ETSI / 3GPP. Note: Release 15, v15.7.0 External Links: Link Cited by: §I-B, §I-A, §I-A. [2] 3GPP (2024) System architecture for the 5G System (5GS). Technical report Technical Report 3GPP TS 23.501 (ETSI TS 123 501), ETSI / 3GPP. Note: Release 18, v18.5.0 External Links: Link Cited by: §I-B, §I-A. [3] J. L. Boudec and P. Thiran (2001) Network calculus: a theory of deterministic queuing systems for the internet. Springer. Cited by: §I-B. [4] S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein (2011) Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends in Machine Learning 3 (1), p. 1–122. External Links: Link Cited by: §IV-A. [5] C.-S. Chang (2000) Performance guarantees in communication networks. Springer. Cited by: §I-B. [6] S. Coles (2001) An introduction to statistical modeling of extreme values. Springer. Cited by: §IV-B. [7] J. Dean and L. A. Barroso (2013) The tail at scale. Communications of the ACM 56 (2), p. 74–80. External Links: Document Cited by: §I-B, §I-A. [8] P. Embrechts, C. Klüppelberg, and T. Mikosch (1997) Modelling extremal events for insurance and finance. Springer. Cited by: §IV-B. [9] Ericsson (2023) Global network API platform to monetize 5G. Note: https://w.ericsson.com/en/reports-and-papers/white-papers/global-network-api-platform-to-monetize-5gAccessed: 2025-12-17 Cited by: §I-A. [10] ETSI (2019) Zero-touch network and Service Management (ZSM); Reference Architecture. Technical report Technical Report ETSI GS ZSM 002, ETSI. Note: v1.1.1 External Links: Link Cited by: §I-B, §I-A. [11] ETSI (2022) Zero-touch network and Service Management (ZSM); End-to-End (E2E) service management across multiple domains. Technical report Technical Report ETSI GS ZSM 008, ETSI. Note: v1.1.1 External Links: Link Cited by: §I-B, §I-B, §I-A. [12] GSMA (2024) GSMA Open Gateway: State of the Market, H1 2024. Note: GSMA Intelligence report External Links: Link Cited by: §I-A. [13] J. P. I (1975) Statistical inference using extreme order statistics. The Annals of Statistics. Cited by: §IV-B. [14] Linux Foundation CAMARA Project (2026) CAMARA: Open Source Project for Telco Network APIs. Note: Project repositories and governance External Links: Link Cited by: §I-B, §I-A, §I-A. [15] e. a. Luo (2023) Make rental reliable: blockchain-based network slice management framework with SLA guarantee. IEEE Commun. Mag. 61 (7), p. 142–148. Cited by: §I-B. [16] M. J. Neely (2010) Stochastic network optimization with application to communication and queueing systems. Morgan & Claypool. Cited by: §IV-A. [17] M. Saimler, M. R. Akdeniz, D. Roeland, A. Kattepur, S. Ertas, M. Thakur, I. Pastushok, M. D’Angelo, J. Yue, A. Ahmed, and M. Y. Donmez (2025) AI as a service: exposing the functionalities of AI-native 6g networks. In Proc. IEEE PIMRC, Cited by: §I-A. [18] M. Saimler, M. D’Angelo, D. Roeland, A. Ahmed, and A. Kattepur (2023-12-14) AI as a service: how AI applications can benefit from the network. Note: Ericsson BlogAccessed: 2026-02-04 External Links: Link Cited by: §I-A. [19] M. Saimler (2024-12) The dawn of AI-Native networks and transforming connectivity with AI-as-a-Service. External Links: Link Cited by: §I-A, §I-B, §I-A, §I-A.