Paper deep dive
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
Deep Kumar Ganguly, Jan Kretinsky
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/19/2026, 5:08:21 AM
Summary
The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework for safe sequential decision-making under evolving uncertainty. RATTL ties an agent's caution to its epistemic uncertainty (Bayesian belief entropy) by defining a Wasserstein ambiguity set whose radius scales with belief entropy. This allows the agent to interpolate between worst-case robustness (high uncertainty) and risk-neutral optimization (low uncertainty). The authors prove theoretical guarantees including contractivity, a 'Safety Sandwich' bounding the value between uninformed robust and full-knowledge optima, and convergence to the best response as the belief concentrates. They demonstrate that the induced risk measure corresponds to Conditional Value-at-Risk (CVaR) in canonical safety settings, highlighting the superiority of Wasserstein distances over KL-divergence for handling off-support catastrophic events.
Entities (8)
Relation Signals (7)
Ambiguity Radius → dependson → Shannon Entropy
confidence 95% · Wasserstein ambiguity radius equals the Shannon entropy of the Bayesian belief
RATTL → modulates → Ambiguity Radius
confidence 95% · radius is a monotone function of that posterior... radius contracts with evidence
RATTL → uses → Wasserstein Distance
confidence 95% · RATTL’s operational ambiguity sets are Wasserstein
RATTL → provides → Safety Sandwich
confidence 93% · prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full-knowledge optimum
RATTL → induces → Conditional Value-at-Risk
confidence 92% · induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy
Wasserstein Distance → outperforms → KL Divergence
confidence 90% · KL cannot place any mass on the off-support catastrophe... only Wasserstein transports mass onto the catastrophe
RATTL → targets → Runtime Safety
confidence 90% · RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
Tags
Links
- Source: https://arxiv.org/abs/2608.17574v1
- Canonical: https://arxiv.org/abs/2608.17574v1
Trouble viewing inline? Open PDF directly →
Full Text
43,932 characters extracted from source content.
Expand or collapse full text
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision MakingThanks: Deep Kumar Ganguly is supported by the German Research Foundation (DFG) through Research Training Group GRK 2428 ConVeY. Jan Křetínský is one of the PI’s of the RobustiAI project. Deep Kumar Ganguly Affiliation: Technical University of Munich, Germany Email: deep.ganguly@tum.de Jan Křetínský Affiliation: Technical University of Munich, Germany Affiliation: Masaryk University, Brno, Czech Republic Email: jan.kretinsky@tum.de Abstract Many agents must act while still learning what environment they are in, which raises a basic question: how cautious should they be? A fixed answer is rarely right—too much caution wastes opportunity once the environment is understood, while planning for the average case can be unsafe early on. We propose RATTL (Risk-Adversarial Total-Reward Learning), a framework that ties an agent’s caution to how much it still does not know. The agent keeps a Bayesian belief about the unknown parts of its environment and hedges against a range of plausible dynamics whose size grows with the uncertainty of that belief, measured through a Wasserstein (optimal-transport) distance; as it learns, this range shrinks and behaviour shifts smoothly from worst-case caution toward ordinary reward maximization. The idea follows the Entropic Value-at-Risk, which recasts “how cautious should I be?” as “how large should my set of plausible models be?”. We put the framework on a formal footing: the planning problem is well posed, and its value always sits between a fully cautious baseline and the best one could do with full knowledge—a Safety Sandwich—with the gap closing as the environment is identified. We also begin to pin down what kind of risk this caution encodes, showing that in a canonical safety setting it matches a familiar tail-risk measure (Conditional Value-at-Risk) whose severity is set by the belief entropy. A simple worked example shows the agent holding back until it is confident, then switching to the efficient action at a clear threshold. RATTL targets runtime safety for agents—including LLM-based systems—that must act under uncertainty. 1 The Problem: Risk That Changes as You Learn Consider an autonomous system that must act in real time while gradually identifying its environment: a self-driving car facing a driver of unclear intent, a medical AI choosing a treatment before a diagnostic returns, or an LLM-based agent invoking tools against an adversary of unknown sophistication. The agent holds a belief over hidden environmental parameters and updates it through observation, yet must commit to actions now. How much risk should it tolerate at each moment? This depends on how much uncertainty remains: early on, even a small probability of catastrophe warrants caution; once the environment is largely identified, risk-neutral optimization is appropriate. The agent’s risk attitude should be a function of its epistemic state, decreasing in conservatism as information accumulates. Neither classical robust control (uniform worst-case reasoning) nor Bayesian RL (expected value under a diffuse belief) does this. A principled mechanism is needed that continuously adjusts the agent’s position on the risk spectrum (Figure 1). [X]E[X]risk-neutralCVaRαCVaR_αtail avg.EVaRαEVaR_αentropicesssup(X)ess\,sup(X)worst case≤ : belief entropy selects position on this spectrum Figure 1: The coherent-risk hierarchy: EVaR is the single-parameter coherent family sweeping E to esssupess\,sup. RATTL uses belief entropy ℋ(b)H(b) to slide along it. Contributions. (i) We formalize RATTL: a reduction of a partially-observed turn-based stochastic game to a subjective robust MDP whose Wasserstein ambiguity radius equals the Shannon entropy of the Bayesian belief (§3). (i) We prove three guarantees—contractivity, a Safety Sandwich, and convergence to best response—each with the precise, adversarially-checked conditions under which it holds (§4). (i) We make first progress on the risk-measure identity of Wasserstein ambiguity: the inner robust value is coherent and equals a Lipschitz-regularized expectation, and on the two-point catastrophe it is exactly a CVaR with an entropy-controlled tail level (§5). (iv) We give a complete worked example with a sharp safety switch (§6). 2 The Risk Spectrum and Why EVaR For a random cost X, the standard coherent risk measures form a hierarchy of increasing conservatism 1: [X]≤CVaRα(X)≤EVaRα(X)≤esssup(X)E[X] _α(X) _α(X) \,sup(X). The Entropic Value-at-Risk is the tightest upper bound on CVaR obtainable from exponential moments: EVaRα(X)=inft>01tln[etX]α.EVaR_α(X)= _t>0\! \ 1t E[e^tX]α \. (1) Two precise properties make EVaR the natural motivating dial. First, it is a single-parameter coherent family with full spectrum coverage: EVaR1=EVaR_1=E and EVaRα→esssupEVaR_α \,sup as α→0+α→ 0^+ (we use it for this coverage and its KL-DRO dual, not as the only such family). Second, it has an exact DRO dual, EVaRα(X)=supDKL(Q∥P)≤−lnαQ[X],EVaR_α(X)= _\,D_KL(Q\|P)\,≤\,- αE_Q[X], (2) so choosing a confidence level α is equivalent to choosing a KL-ball radius −lnα- α around the nominal P. This is the equivalence we exploit: it converts “how conservative should the agent be?” into “how large should the ambiguity set be?”—and the latter has a natural answer in the agent’s epistemic uncertainty. We are careful to claim only this: EVaR supplies the conceptual bridge. RATTL’s operational ambiguity sets are Wasserstein, for a safety reason developed in §5, and the precise risk measure they induce is the subject of that section. 3 RATTL: Belief Entropy as a Risk Dial Setting. A partially-observed turn-based stochastic game (PO-TBSG) has finite states S, finite actions A, a finite set of opponent types Z, a terminal set ⊂T , and for each type z a transition kernel P(⋅∣s,a,z)∈Δ()P(· s,a,z)∈ (S). Rewards r(s,a)r(s,a) are bounded and terminal rewards r(s)r(s) are defined for s∈s . We use the undiscounted total-reward criterion (a proper / stochastic-shortest-path MDP). The agent maintains a Bayesian belief b∈Δ()b∈ (Z) over the type, updated by ψ(b,s,a,s′)(z)=P(s′∣s,a,z)b(z)∑z′P(s′∣s,a,z′)b(z′),ψ(b,s,a,s )(z)= P(s s,a,z)\,b(z) _z P(s s,a,z )\,b(z ), (3) and the belief-averaged (nominal) kernel P¯b(⋅∣s,a)=∑zb(z)P(⋅∣s,a,z) P_b(· s,a)= _zb(z)P(· s,a,z). Definition 1 (Entropy-modulated ambiguity set). For a ground metric d on S and risk-sensitivity β>0β>0, (b∣s,a)=Q∈Δ():W1(Q,P¯b(⋅∣s,a))≤βℋ(b),U(b s,a)= \\,Q∈ (S):W_1\! (Q, P_b(· s,a) )≤β\,H(b)\, \, (4) where W1W_1 is the 11-Wasserstein distance and ℋ(b)=−∑zb(z)lnb(z)H(b)=- _zb(z) b(z) is the Shannon entropy. The radius ε(b)=βℋ(b) (b)= (b) implicitly selects the agent’s conservatism: a diffuse belief gives wide ambiguity (effectively worst-case), a sharp belief gives tight ambiguity (effectively risk-neutral), with graduated conservatism in between (Figure 2). The induced Nash-Robust Bellman operator on bounded V:×Δ()→ℝV:S× (Z) with V(s,⋅)=r(s)V(s,·)=r(s) for s∈s is (V)(s,b)=maxa[r(s,a)+infQ∈(b∣s,a)∑s′Q(s′)V(s′,ψ(b,s,a,s′))]. 20801085$ (TV)(s,b)= _a\! [r(s,a)+\!\! _Q (b s,a)\!\! _s Q(s )\,V\! (s ,ψ(b,s,a,s ) ) ]$. (5) The belief update ψ is applied inside the expectation, coupling the adversary’s kernel choice with the agent’s future information state. The operator has a game reading: the agent (Max) picks an action; a fictitious adversary (Nature/Min) picks the worst kernel within the current ambiguity budget. RATTL goes in the opposite direction to the known RMDP → stochastic-game reduction 4: rather than expanding an RMDP into a game, we compress a partially-observed game into a subjective robust MDP whose uncertainty set is non-stationary, evolving with belief. sFs_Fsbrs_brsGs_Gb=(.5,.5)b=(.5,.5)ℋ=0.69H=0.69sFs_Fsbrs_brsGs_Gb=(.85,.15)b=(.85,.15)ℋ=0.42H=0.42sFs_Fsbrs_brsGs_Gb=(.99,.01)b=(.99,.01)ℋ=0.06H=0.06 Figure 2: Wasserstein ambiguity balls (b)U(b) on the next-state simplex for three beliefs of decreasing entropy. The ball shrinks (and its center P¯b P_b moves) as the belief sharpens: high entropy lets the adversary push mass toward catastrophic states; low entropy renders it nearly powerless. Assumptions. We isolate the conditions our theorems need. Assumption 1 (Properness). Under every policy and every kernel selection in every (b)U(b), the terminal set T is reached with probability 11. Assumption 2 (Uniform reachability). There exist m≥1m≥ 1 and η∈(0,1]η∈(0,1] such that from every non-terminal (s,b)(s,b), under every admissible control and adversary kernel, T is reached within m steps with probability ≥η≥η. Assumption 3 (Identifiability). For all z≠z′z≠ z there is an (s,a)(s,a) with P(⋅∣s,a,z)≠P(⋅∣s,a,z′)P(· s,a,z)≠ P(· s,a,z ). Assumption 2 is the quantitative strengthening that the belief continuum forces: a.s. reachability alone (Ass. 1) need not give a uniformly bounded expected hitting time over Δ() (Z). With finite ,,S,A,Z it holds whenever the one-step absorption probability is bounded below over the (compact) action×adversary sets. Rewards are assumed bounded throughout. 4 Theoretical Guarantees Theorem 1 (Contractivity). Under Assumptions 1–2, let w(s)=supbsupadm.(s,b)[τ]w(s)= _b _adm.E^(s,b)[ _T] be the worst-case expected hitting time. Then 1≤w(s)≤W:=m/η<∞1≤ w(s)≤ W:=m/η<∞, and T is a contraction of modulus ρ=1−1/W<1ρ=1-1/W<1 in the weighted sup-norm ‖V‖w=maxs∉,b|V(s,b)|/w(s)\|V\|_w= _s ,b|V(s,b)|/w(s) on the affine space V:V(s,⋅)=r(s)∀s∈\V:V(s,·)=r(s)\ ∀ s \. Hence T has a unique fixed point V∗V^* and kV→V∗T^kV→ V^* geometrically. Proof sketch. Iterating Assumption 2 over blocks of m steps gives Pr(τ>km)≤(1−η)k ( _T>km)≤(1-η)^k, so w(s)≤m/ηw(s)≤ m/η uniformly in b and strategy; thus ∥⋅∥w\|·\|_w is a genuine norm on the affine space (differences vanish on T) equivalent to ∥⋅∥∞\|·\|_∞, which is complete. The inner inf over a fixed set is non-expansive, |infQf−infQg|≤supQ|f−g|| _Qf- _Qg|≤ _Q|f-g|, and since r(s,a)r(s,a) cancels and maxa _a is non-expansive, a one-step drift inequality (ℒw)(s)≤w(s)−1(Lw)(s)≤ w(s)-1 (worst-case unit-cost SSP, 3) yields |(V−V′)(s,b)|≤‖V−V′‖w(w(s)−1)|(TV-TV )(s,b)|≤\|V-V \|_w\,(w(s)-1); dividing by w(s)w(s) and using w≤Ww≤ W gives modulus 1−1/W1-1/W. Belief augmentation is harmless: ψ is a deterministic function of the history, rewards depend only on (s,a)(s,a), and ⊂T so reachability is independent of b. Full proof in Appendix A. ∎ The next result is the one reviewers asked us to make concrete: what does robustness-with-learning do, behaviorally? It brackets the value between two reference policies. Theorem 2 (Safety Sandwich). Let VmaximinV_maximin be the fixed point of the operator identical to (5) but with the radius frozen at εmax=βln|| _ =β |Z| (same center P¯b P_b), and let VBR(⋅,z)V_BR(·,z) be the optimal value of the non-robust MDP with known type z. Under Assumptions 1–3, Vmaximin(s,b)≤V∗(s,b)≤∑z∈b(z)VBR(s,z)∀(s,b).V_maximin(s,b)\ ≤\ V^*(s,b)\ ≤\ _z b(z)\,V_BR(s,z) ∀(s,b). (6) Proof. Lower bound. Since ℋ(b)≤ln||H(b)≤ |Z| and both balls share the center P¯b P_b, (b)⊆max(b)U(b) _ (b); the inf over the larger set is no larger, so maxV≤VT_ V pointwise for every V. Starting value iteration from V∗V^* and using monotonicity, maxkV∗≤V∗=V∗T_ ^kV^* ^*=V^* for all k; letting k→∞k→∞ gives Vmaximin≤V∗V_maximin≤ V^*. Upper bound. For any policy π, the nominal kernel is feasible (P¯b∈(b) P_b (b)), so the robust value Vrobπ≤VnomπV^π_rob≤ V^π_nom. Running P¯b P_b with ψ-updates is exactly the marginal-and-posterior factorization of the mixture process “draw z∼bz b once, then follow P(⋅∣⋅,⋅,z)P(· ·,·,z)”; hence the trajectory laws coincide and Vnomπ(s,b)=∑zb(z)Vzπ(s)V^π_nom(s,b)= _zb(z)V^π_z(s). As Vzπ(s)≤VBR(s,z)V^π_z(s)≤ V_BR(s,z) for every type, Vrobπ(s,b)≤∑zb(z)VBR(s,z)V^π_rob(s,b)≤ _zb(z)V_BR(s,z); taking supπ _π gives the claim. Full proof in Appendix B. ∎ Remark 1 (Why the belief-averaged ceiling). The stronger ceiling V∗≤VBR(s,z∗)V^*≤ V_BR(s,z^*) at the realized true type z∗z^* is false in general: V∗V^* is a deterministic function of (s,b)(s,b) while z∗z^* is random, and the true kernel P(⋅∣s,a,z∗)P(· s,a,z^*) typically lies outside (b)U(b) when b is not a point mass (Identifiability makes the types W1W_1-separated). As b→δz∗b→ _z^* both bounds coincide: the radius vanishes and ∑zb(z)VBR(s,z)→VBR(s,z∗) _zb(z)V_BR(s,z)→ V_BR(s,z^*). This is exactly the asymptotic role of Theorem 3. Theorem 3 (Convergence to best response). Assume 1, 3, and persistent identification (PI): along the realized trajectory, for each pair of types an identifying (s,a)(s,a) is visited infinitely often a.s. (automatic under any proper, fully-exploring behavior policy). Then bt→δz∗b_t→ _z^* a.s. 15, so ℋ(bt)→0H(b_t)→ 0 and (bt)U(b_t) collapses to P(⋅∣⋅,⋅,z∗)\P(· ·,·,z^*)\ in Hausdorff distance; consequently V∗(s,bt)→VBR(s,z∗)V^*(s,b_t)→ V_BR(s,z^*) a.s. If in addition the per-step log-likelihood separation is bounded below (persistent excitation; e.g. all positive transition probabilities ≥pmin>0≥ p_ >0), then [ℋ(bt)]=O(||logt/t)E[H(b_t)]=O(|Z| t/t) and the price of robustness obeys |VBR(s,z∗)−V∗(s,bt)|=O(logt/t)|V_BR(s,z^*)-V^*(s,b_t)|=O( t/t). Proof sketch. PI supplies the excitation that Identifiability alone lacks, giving Bayesian consistency (Doob/Schwartz). The value-continuity step is the delicate one: V∗(⋅,b)V^*(·,b) is not the fixed point of any per-belief operator because ψ shifts the belief; instead one shows T maps the class of value functions with belief-modulus ≤L≤ L near δz∗ _z^* into itself, provided the Bayes normalizer is bounded below on the realized support (so ψ is locally Lipschitz), and the unique fixed point inherits the modulus—hence continuity at δz∗ _z^*. The rate follows from [ℋ(bt)]=O(logt/t)E[H(b_t)]=O( t/t) under persistent excitation. Full statement and proof in Appendix C. ∎ Tractability. The inner inf is a finite linear program. By Kantorovich–Rubinstein duality 16, infQ∈(b)∑s′Q(s′)V(s′)=supλ≥0P¯b[mins′(V(s′)+λd(s′,⋅))]−λε(b), 20801085$ _Q (b)\! _s \!Q(s )V(s )= _λ≥ 0 \E_ P_b\! [ _s (V(s )+λ d(s ,·)) ]-λ (b) \$, (7) so belief-augmented robust value iteration over a belief grid ⊂Δ()G⊂ (Z) solves |‖||S||G||A| LPs of size O(||)O(|S|) per sweep. For large type spaces a particle/variational posterior approximates b, and (7) keeps the inner problem tractable regardless of belief representation; the continuous-state case (a Lipschitz-critic realization of (7)) we leave as an explicit open item (§8). 5 From KL to Wasserstein: A Coherent-Risk Reading EVaR is a KL-ball worst case (2); RATTL uses a Wasserstein ball. This is deliberate. Under KL the adversary cannot place mass where the nominal has none (DKL=+∞D_KL=+∞ off-support), so precisely when the agent grows confident—and P¯b P_b concentrates away from rare catastrophes—the KL adversary loses the ability to model the catastrophe. Wasserstein prices perturbations by physical distance, letting the adversary reach any state at proportional cost. This is not merely rhetorical. Table 1 reports the inner worst case on a 55-state “cliff” transition where the catastrophic state carries zero nominal mass. The KL adversary cannot place any mass on the catastrophe (it lies off the nominal support, where DKL=+∞D_KL=+∞), so it cannot price the tail event even though it still perturbs the on-support mass; Wasserstein transports mass onto the catastrophe at finite cost and reports a worst case two orders of magnitude lower. W1W_1 TV KL Worst-case Q[V]E_Q[V], ε=0.25 =0.25 −345.0-345.0 −217.5-217.5 −3.4-3.4 Mass moved onto catastrophe 0.400.40 0.250.25 0.000.00 Table 1: Inner adversary at radius ε=0.25 =0.25 on a cliff transition whose catastrophe carries zero nominal mass (nominal value +57.5+57.5). KL cannot place any mass on the off-support catastrophe (DKL=+∞D_KL=+∞), so its mass there stays 00 and it cannot price the tail even as it reweights on-support mass; only Wasserstein reaches the catastrophe. KLTVW1W_1−300-300−150-15000nominal (non-robust)−3.4-3.4−217.5-217.5−345-345worst-case Q[V]E_Q[V] Figure 3: The KL support catastrophe, visualized (radius ε=0.25 =0.25). The KL adversary cannot reach the off-support catastrophe (its mass there stays 00), so it barely departs from the non-robust value (dashed line) despite reweighting on-support mass; only Wasserstein transports mass onto the catastrophe and reports the true tail risk (−345-345). What coherent risk measure, then, does a Wasserstein ball induce? Assembling standard duality results, we record the following characterization. Theorem 4 (Coherence and Lipschitz-regularization). Fix finite S with a metric ground cost d, P¯∈Δ() P∈ (S), ε≥0 ≥ 0, and ρε(X):=supQ:W1(Q,P¯)≤εQ[X] _ (X):= _Q:W_1(Q, P)≤ E_Q[X]. Then (1) ρε _ is a coherent risk measure (monotone, translation-equivariant, positively homogeneous, subadditive), being the Artzner et al. 2 worst-case-expectation functional of the convex compact scenario set εU_ ; and (2) it equals a Lipschitz-regularized expectation, infQ:W1(Q,P¯)≤εQ[f]=supλ≥0P¯[fλ]−λε, _Q:W_1(Q, P)≤ \!E_Q[f]= _λ≥ 0 \E_ P[f_λ]-λ \, (8) with fλ(s)=mins′(f(s′)+λd(s′,s))f_λ(s)= _s (f(s )+λ d(s ,s)) the inf-convolution of f with λdλ d, and the optimal λ⋆≤Lipd(f)λ _d(f) 7; 9. Proof. (1) εU_ is nonempty (P¯∈ε P _ ), convex (the map Q↦W1(Q,P¯)Q W_1(Q, P) is convex, by averaging optimal couplings), and compact (it is a closed subset of the simplex, W1(⋅,P¯)W_1(·, P) being continuous via Kantorovich duality). The four axioms follow from properties of a supremum of linear functionals over a fixed convex set; the Artzner representation is then immediate. (2) is finite-S LP strong duality: the transport LP min∑ππ(s,s′)f(s′) _πΣπ(s,s )f(s ) s.t. first marginal P¯ P and ∑πd≤εΣπ\,d≤ has Lagrangian dual maxλ≥0P¯[fλ]−λε _λ≥ 0\E_ P[f_λ]-λ \, strictly feasible hence no gap. Details in Appendix D. ∎ On the canonical safety instance—a good state g and a catastrophe f—the identity is sharp and, strikingly, is a CVaR, not an EVaR. Proposition 1 (Two-point Wasserstein risk is CVaR). Let =g,fS=\g,f\ with V(g)=vg>vf=V(f)V(g)=v_g>v_f=V(f), nominal catastrophe mass q∈(0,1)q∈(0,1), d(g,f)=Dd(g,f)=D, loss L=−VL=-V. For ε<(1−q)D <(1-q)D the worst-case value is V¯(ε)=(1−p⋆)vg+p⋆vf,p⋆=q+εD, V( )= (1-p )v_g+p v_f, p =q+ D, (9) and the associated robust loss is exactly a Conditional Value-at-Risk, −V¯(ε)=CVaRθ(ε)(L),θ(ε)=q+ε/D.- V( )=CVaR_θ( )(L), θ( )= qq+ /D. (10) Since CVaR≤EVaRCVaR at a matched tail level, the Wasserstein risk is in turn dominated by an EVaR; the matching EVaR level depends on ε with no closed form, which is why CVaR—not EVaR—is the clean object here. Proof. The two-point ball is p:|p−q|≤ε/D\p:|p-q|≤ /D\; Q[V]E_Q[V] decreases in the catastrophe mass p, so the worst case is p⋆=min(1,q+ε/D)p = (1,q+ /D). Substituting p⋆=q/θp =q/θ into the two-point CVaRθ(L)=1θ(qL(f)+(θ−q)L(g))CVaR_θ(L)= 1θ(qL(f)+(θ-q)L(g)) recovers p⋆L(f)+(1−p⋆)L(g)=−V¯(ε)p L(f)+(1-p )L(g)=- V( ) as an exact algebraic identity. The EVaR bound is the CVaR≤EVaRCVaR ordering 1. Details in Appendix D. ∎ Because ε=βℋ(b) = (b), the tail level θ(b)=q/(q+βℋ(b)/D)θ(b)=q/(q+ (b)/D) decreases as belief entropy grows: the entropy dial is literally a CVaR-tail dial (Figure 4). This is a concrete, exact instance of the paper’s slogan. 0.50.50.550.550.60.60.650.650.70.70.750.750.80.80.850.850.90.90.950.9511000.20.20.40.40.60.60.80.811risk-neutral θ→1θ→1deep tail (worst-case-like)belief confidence α (entropy ℋ↓H\! as α→1α\!→\!1)CVaR tail level θ(ε)θ( )q=0.05q=0.05q=0.10q=0.10q=0.20q=0.20 Figure 4: Entropy is a CVaR dial (Proposition 1). As the belief sharpens (α→1α\!→\!1, ℋ→0H\!→\!0), the induced CVaR tail level θ(ε)=q/(q+βℋ/D)θ( )=q/(q+ /D) rises from a deep tail (worst-case-like caution) toward 11 (risk-neutrality), for catastrophe masses q∈0.05,0.1,0.2q∈\0.05,0.1,0.2\. This is the risk spectrum of Figure 1, now traversed automatically by learning. Two honest caveats: for ||≥3|S|≥ 3 the worst case spreads mass to the nearest low-value states, giving a d-weighted “transport-CVaR” rather than ordinary CVaR; and a single, ε -uniform EVaR-level identity does not hold (the matching EVaR level depends transcendentally on ε ). The general transport-CVaR↔ comparison is the headline open problem; Proposition 1 settles the canonical case. 6 Worked Example: The Ambiguous Bridge (All curves below are closed-form evaluations of the robust backup, not learning runs; here α=b(benign)α=b(benign) denotes belief confidence.) To make every quantity concrete we instantiate RATTL on a diagnostic with three states sbr,sG,sF\s_br,s_G,s_F\ (bridge, goal, fall), two actions Sprint,Crawl\ Sprint, Crawl\, and two types benign,adversarial\benign,adversarial\. Sprint reaches sGs_G under the benign type but falls to sFs_F under the adversarial type; Crawl always reaches sGs_G. Rewards: r(sbr,Sprint)=−1r(s_br, Sprint)=-1, r(sbr,Crawl)=−20r(s_br, Crawl)=-20, r(sG)=+100r(s_G)=+100, r(sF)=−1000r(s_F)=-1000; ground distance D=d(sG,sF)=1D=d(s_G,s_F)=1; β=1β=1. Writing α=b(benign)α=b(benign) and ε(α)=ℋ(α) (α)=H(α), both successor states are terminal, so a single robust backup with the two-point transport W1(μ,ν)=D|μ1−ν1|W_1(μ,ν)=D| _1- _1| gives closed-form robust Q-values: Q(Sprint,α) Q( Sprint,α) =−1001+1100max(0,α−ε(α)), =-1001+1100 \! (0,α- (α) ), (11) Q(Crawl,α) Q( Crawl,α) =80−1100min(ε(α),1). =80-1100 \! ( (α),1 ). (12) Setting them equal, the ε terms cancel and 1100α=10811100α=1081, so the safety switch is at α∗=10811100≈0.983.α^*= 10811100≈ 0.983. (13) The agent Crawls until it is ∼ 98% confident the conditions are benign, then Sprints (Figure 5, Table 2). In this symmetric environment the threshold is determined purely by the reward asymmetry and is independent of β (both actions face the same per-unit transport penalty); β controls the conservatism of the value and shifts the threshold only when the safe and risky actions face different ambiguity. Well-posedness requires ε(b)≤D (b)≤ D, i.e. β<D/ln||≈1.443β<D/ |Z|≈ 1.443. α ε(α) (α) Q(Sprint)Q( Sprint) Q(Crawl)Q( Crawl) π∗π^* 0.500.50 0.69310.6931 −1001.00-1001.00 −682.46-682.46 Crawl 0.850.85 0.42270.4227 −530.98-530.98 −384.98-384.98 Crawl 0.990.99 0.05600.0560 +26.40+26.40 +18.40+18.40 Sprint α∗=1081/1100α^*=1081/1100 0.08720.0872 −15.95-15.95 −15.95-15.95 switch Table 2: Robust value iteration on the Ambiguous Bridge (β=1β=1, D=1D=1). The safety floor is Vmaximin(0.5)=80−1100ln2=−682.46V_maximin(0.5)=80-1100 2=-682.46; the best-response ceiling is VBR=99V_BR=99 as α→1α→ 1. The threshold α∗=1081/1100α^*=1081/1100 is β-independent here. 0.50.50.550.550.60.60.650.650.70.70.750.750.80.80.850.850.90.90.950.9511−1,000-1,000−500-50000α∗≈0.983α^*≈0.983belief confidence α=b(benign)α=b(benign)robust Q-valueQ(Sprint)Q( Sprint)Q(Crawl)Q( Crawl) Figure 5: The safety switch. Crawl dominates while uncertain; at α∗≈0.983α^*≈ 0.983 the risky action’s worst-case value overtakes the safe one and the agent switches to Sprint. 0.50.50.550.550.60.60.650.650.70.70.750.750.80.80.850.850.90.90.950.9511−600-600−400-400−200-20000belief confidence α ∗(sbr,α)V^*(s_br,α)RATTL V∗V^*worst-case floor −682.5-682.5best-response ceiling 9999 Figure 6: The Safety Sandwich, instantiated. RATTL’s value rises from the worst-case floor (VmaximinV_maximin, never-shrinking ambiguity) toward the full-information ceiling (VBRV_BR) as the belief sharpens—never less safe than maximin, never more reckless than the informed optimum. The two reference lines are constant bounds (−682.5-682.5 and 9999); V∗V^* provably stays between them at every belief. 7 Related Work RATTL sits at the intersection of robust MDPs, risk-sensitive RL, and Bayesian opponent modeling. Robust MDPs with static, rectangular uncertainty sets 8; 12; 17 are RATTL’s ancestor; our radius is belief-dependent and non-stationary, and we go in the opposite direction to the RMDP↔ equivalence 4. On the risk side, EVaR 1 and CVaR 13 supply the coherent backbone 2; Ni & Bhat 11 show stationary policies suffice for EVaR total-reward MDPs—justifying RATTL’s policy class—but at a fixed risk level, whereas ours is belief-selected. The closest competitors couple ambiguity to Bayesian posteriors: Russel & Petrik 14 adapt sets to the policy (not the belief); Choi and Li 5 contract interval credible sets rather than entropy-modulated Wasserstein balls and lack a risk-measure reading; Nakao et al. 10 study DR-POMDPs with static distance-based ambiguity. Derman & Mannor 6 relate Wasserstein DRO to value regularization, which Theorem 4 sharpens into an explicit coherent-risk identity. To our knowledge, no prior work couples a Wasserstein radius to Shannon belief entropy (as opposed to posterior credible-set width, cf. Choi and Li), nor derives the resulting Safety Sandwich and two-point CVaR identity. 8 Limitations, Future Work, and Conclusion Limitations. We are explicit about the current scope. (a) Tabular. The exact algorithm discretizes the belief simplex and is feasible only for small |||Z|; we give a scalable roadmap (§4) but no large-scale empirical validation yet. (b) Added conditions. The guarantees need more than bare properness: uniform reachability (Assumption 2) for the contraction over the belief continuum, and persistent identification/excitation for convergence and its rate—degenerate exploration is excluded. (c) Two points. The exact CVaR identity (Proposition 1) is proven only for the canonical two-point catastrophe; for ||≥3|S|≥ 3 it becomes a d-weighted “transport-CVaR” and the EVaR comparison is only one-sided. (d) Calibration. β must satisfy β<D/ln||β<D/ |Z|, and the threshold’s β-independence is specific to symmetric environments like the Bridge. Future work. (i) The general transport-CVaR↔ comparison for ||≥3|S|≥ 3. (i) Hybrid Wasserstein++KL ambiguity combining reachability with tail sensitivity. (i) Rényi/Tsallis entropy as alternative dials—does the Safety Sandwich survive for any concave, zero-at-point-mass information measure? (iv) Sample complexity of belief-adaptive robust RL, where the radius shrinks at rate O(logt/t)O( t/t). (v) Scaling the dual (7) to continuous spaces via Lipschitz critics. (vi) Multi-agent risk composition under private beliefs. The headline applied direction—and the workshop motivation—is runtime safety for agentic GenAI: agents that plan, retrieve, and invoke tools hold a Bayesian or surrogate belief over latent context and face deployment-time distribution shift and adversarial tail risk, exactly RATTL’s structure. Belief-entropy-modulated robustness then offers a principled dial for abstention and graduated conservatism, with the Safety Sandwich providing a value-level bracket—a worst-case floor and a Bayesian-best-response ceiling—rather than a runtime behavioral guarantee; turning it into an actionable runtime certificate, and validating it on real LLM-agent pipelines, is the main empirical goal. Conclusion. RATTL makes an agent’s risk attitude a function of its epistemic state by tying a Wasserstein ambiguity radius to belief entropy. We proved a contraction theorem, a Safety Sandwich bracketing the value between a worst-case floor and a Bayesian-best-response ceiling, and almost-sure convergence to best response, each with explicit, adversarially-checked conditions. We took a first step on the risk-measure identity of Wasserstein ambiguity: it is coherent and equals a Lipschitz-regularized expectation in general, and an exact CVaR with an entropy-controlled tail level on the canonical catastrophe—so the belief literally selects a position on the coherent-risk spectrum. The Ambiguous Bridge makes the resulting safety switch sharp and interpretable. The thesis is simple: robustness should decrease with epistemic certainty, and belief entropy is the principled dial that achieves it. Acknowledgments We thank the RobustifAI workshop reviewers for constructive feedback that substantially shaped this extended version. Appendix A Full Proof of Theorem 1 Work on the affine space ℬ=Vbounded:V(s,⋅)=r(s)∀s∈B_T=\V\ bounded:V(s,·)=r(s)\ ∀ s \, on which differences vanish on T so ∥⋅∥w\|·\|_w separates points and is equivalent to ∥⋅∥∞\|·\|_∞ over s∉s ; ℬB_T is complete and :ℬ→ℬT:B_T _T. Iterating Assumption 2 over blocks of m steps gives Pr(τ>km)≤(1−η)k ( _T>km)≤(1-η)^k, hence w(s)=[τ]≤∑k≥0m(1−η)k=m/η=:Ww(s)=E[ _T]≤ _k≥ 0m(1-η)^k=m/η=:W uniformly in b and strategy, and w(s)≥1w(s)≥ 1 since a non-terminal state needs a transition. For bounded f,gf,g on any set, f≤g+sup|f−g|f≤ g+ |f-g| gives inff≤infg+sup|f−g| f≤ g+ |f-g|; symmetrizing yields |infQf−infQg|≤supQ|f−g|| _Qf- _Qg|≤ _Q|f-g|. With f(Q)=∑s′Q(s′)V(s′,ψ)f(Q)= _s Q(s )V(s ,ψ), g(Q)=∑s′Q(s′)V′(s′,ψ)g(Q)= _s Q(s )V (s ,ψ) and Q∈Δ()Q∈ (S), the inner terms differ by at most supQ∈(b)∑s′∉Q(s′)|V−V′|(s′,ψ) _Q (b) _s Q(s )|V-V |(s ,ψ). Using |V−V′|(s′,ψ)≤‖V−V′‖ww(s′)|V-V |(s ,ψ)≤\|V-V \|_w\,w(s ) on s′∉s (and 00 on T), cancellation of r(s,a)r(s,a), and non-expansiveness of maxa _a, |(V−V′)(s,b)|≤‖V−V′‖wsupQ∈(b∣s,a)∑s′∉Q(s′)w(s′).|(TV-TV )(s,b)|≤\|V-V \|_w\!\! _Q (b s,a)\!\! _s \!Q(s )w(s ). The worst-case expected hitting time satisfies the unit-cost SSP drift inequality supa,b,Q∑s′∉Q(s′)w(s′)≤w(s)−1 _a,b,Q _s Q(s )w(s )≤ w(s)-1 (one admissible step costs 11 and leaves residual worst-case time ≤w(s′)≤ w(s ); 3). Hence |(V−V′)(s,b)|≤‖V−V′‖w(w(s)−1)|(TV-TV )(s,b)|≤\|V-V \|_w(w(s)-1); dividing by w(s)∈[1,W]w(s)∈[1,W] gives ‖V−V′‖w≤(1−1/W)‖V−V′‖w\|TV-TV \|_w≤(1-1/W)\|V-V \|_w. Banach’s theorem completes the proof. □ Appendix B Full Proof of Theorem 2 Both T and the frozen-radius maxT_ are monotone self-maps of ℬB_T with unique fixed points V∗,VmaximinV^*,V_maximin (Theorem 1 applies to each, with εmax _ a valid radius), and value iteration converges for both. Lower bound. 0≤ℋ(b)≤ln||0 (b)≤ |Z|, so ε(b)≤εmax (b)≤ _ with common center, giving (b)⊆max(b)U(b) _ (b) and infmax≤inf _U_ ≤ _U; adding r and taking maxa _a, maxV≤VT_ V for all V. Induct: max0V∗=V∗T_ ^0V^*=V^*; if maxkV∗≤V∗T_ ^kV^*≤ V^* then by monotonicity maxk+1V∗≤maxV∗≤V∗=V∗T_ ^k+1V^* _ V^* ^*=V^*. Limit: Vmaximin≤V∗V_maximin≤ V^*. Upper bound. Fix a policy π. (a) P¯b∈(b) P_b (b) (zero transport), so the robust value Vrobπ≤VnomπV^π_rob≤ V^π_nom, the value under the nominal kernel with ψ-updates. (b) The joint law of (z,s′)(z,s ) with z∼bz b, s′∼P(⋅∣s,a,z)s P(· s,a,z) has s′s -marginal P¯b(⋅∣s,a) P_b(· s,a) and z-posterior ψ(b,s,a,s′)ψ(b,s,a,s ); by induction on t the nominal process and the mixture process “draw z∼bz b once, then follow P(⋅∣⋅,⋅,z)P(· ·,·,z)” induce the same trajectory law and maintain bt=Prmix(z∣ht)b_t= ^mix(z h_t). Since the return depends only on the observable trajectory, Vnomπ(s,b)=mix[G]=∑zb(z)Vzπ(s)V^π_nom(s,b)=E^mix[G]= _zb(z)V^π_z(s) by the tower property. (c) Vzπ(s)≤VBR(s,z)V^π_z(s)≤ V_BR(s,z) since π is feasible in the type-z MDP. Chaining and taking supπ _π (with V∗=supπVrobπV^*= _πV^π_rob from proper minimax SSP theory) gives V∗(s,b)≤∑zb(z)VBR(s,z)V^*(s,b)≤ _zb(z)V_BR(s,z). □ Appendix C Statement and Proof of Theorem 3 Consistency. Under (PI) an identifying (s,a)(s,a) for each type pair is visited infinitely often, so the realized log-likelihood ratio of z∗z^* against any z≠z∗z≠ z^* diverges to +∞+∞ a.s. (a sum of i.o. strictly positive-mean, bounded-variance increments), giving bt→δz∗b_t→ _z^* a.s. 15; thus ℋ(bt)→0H(b_t)→ 0 and ε(bt)→0 (b_t)→ 0, and (bt∣s,a)→P(⋅∣s,a,z∗)U(b_t s,a)→\P(· s,a,z^*)\ in the W1W_1-Hausdorff metric. Continuity. Let N be a neighborhood of δz∗ _z^* on which the Bayes normalizer is ≥pmin>0≥ p_ >0, so ψ(⋅,s,a,s′)ψ(·,s,a,s ) is LψL_ψ-Lipschitz on N. The class L=V∈ℬ:|V(s,b)−V(s,b′)|≤L∥b−b′∥∀b,b′∈NC_L=\V _T:|V(s,b)-V(s,b )|≤ L\,\|b-b \|\ ∀ b,b ∈ N\ is closed and T-invariant once L is large enough that the per-step belief-modulus contraction (1−1/W)Lψ<1\,(1-1/W)L_ψ<1\, holds (the only place this extra condition enters); the unique fixed point V∗V^* therefore lies in LC_L and is continuous at δz∗ _z^*. Combined with bt→δz∗b_t→ _z^* and (bt)→Pz∗U(b_t)→\P_z^*\, V∗(s,bt)→VBR(s,z∗)V^*(s,b_t)→ V_BR(s,z^*) a.s. Rate. Under persistent excitation the posterior on each wrong type decays geometrically in the number of identifying visits, and with Θ(t) (t) such visits [ℋ(bt)]=O(||logt/t)E[H(b_t)]=O(|Z| t/t); since |VBR(s,z∗)−V∗(s,bt)|≤C∑z≠z∗bt(z)|V_BR(s,z^*)-V^*(s,b_t)|≤ C _z≠ z^*b_t(z) for a constant C (both the floor gap and the averaged-ceiling gap vanish with the residual mass on wrong types), the price of robustness is O(logt/t)O( t/t). □ Appendix D Details for §5 Convexity of εU_ : for optimal couplings π0,π1 _0, _1 of Q0,Q1Q_0,Q_1 with P¯ P, the coupling θπ0+(1−θ)π1θ _0+(1-θ) _1 has marginals θQ0+(1−θ)Q1θ Q_0+(1-θ)Q_1 and P¯ P, so W1(θQ0+(1−θ)Q1,P¯)≤θW1(Q0,P¯)+(1−θ)W1(Q1,P¯)W_1(θ Q_0+(1-θ)Q_1, P)≤θ W_1(Q_0, P)+(1-θ)W_1(Q_1, P). Coherence axioms follow because εU_ does not depend on X: monotonicity and positive homogeneity are termwise; translation-equivariance uses ⟨Q,⟩=1 Q,1 =1; subadditivity uses that the two suprema decouple. The duality is finite-LP strong duality with multipliers u(s)u(s) (marginals) and λ≥0λ≥ 0 (budget); optimizing u(s)=mins′(f(s′)−λd(s,s′))=fλ(s)u(s)= _s (f(s )-λ d(s,s ))=f_λ(s) gives the boxed form, and fλ=f_λ=f once λ≥Lipd(f)λ _d(f), bounding λ⋆λ . For Proposition 1, the two-point ball is p:|p−q|≤ε/D\p:|p-q|≤ /D\ by Kantorovich duality (|h(g)−h(f)|≤D|h(g)-h(f)|≤ D), the worst case is the right endpoint p⋆p , and substituting p⋆=q/θp =q/θ (with θ=q/(q+ε/D)θ=q/(q+ /D)) into CVaRθ(L)=1θ(qL(f)+(θ−q)L(g))CVaR_θ(L)= 1θ(qL(f)+(θ-q)L(g)) reproduces −V¯(ε)- V( ) exactly; the CVaR≤EVaRCVaR ordering 1 gives the one-sided EVaR bound. □ References Ahmadi-Javid (2012) A. Ahmadi-Javid Entropic value-at-risk: a new coherent risk measure. Journal of Optimization Theory and Applications 155 (3), p. 1105–1123. Cited by: Appendix D, §2, §5, §7. Artzner et al. (1999) P. Artzner, F. Delbaen, J. Eber, and D. Heath Coherent measures of risk. Mathematical Finance 9 (3), p. 203–228. Cited by: §7, Theorem 4. Bertsekas and Tsitsiklis (1996) D. P. Bertsekas and J. N. Tsitsiklis Neuro-dynamic programming. Athena Scientific. Cited by: Appendix A, §4. Chatterjee et al. (2024) K. Chatterjee, J. Kretínský, M. Mohagheghi, M. Sadigh, et al. Solving long-run average reward robust MDPs via stochastic games. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI), Cited by: §3, §7. Choi and Li (2025) J. Choi and M. Z. Li Bayesian ambiguity contraction-based adaptive robust Markov decision processes for adversarial surveillance missions. arXiv preprint arXiv:2512.01660. Cited by: §7. Derman and Mannor (2020) E. Derman and S. Mannor Distributional robustness and regularization in reinforcement learning. arXiv preprint arXiv:2003.02894. Cited by: §7. Gao and Kleywegt (2023) R. Gao and A. J. Kleywegt Distributionally robust stochastic optimization with Wasserstein distance. Mathematics of Operations Research 48 (2), p. 603–655. Cited by: Theorem 4. Iyengar (2005) G. N. Iyengar Robust dynamic programming. Mathematics of Operations Research 30 (2), p. 257–280. Cited by: §7. Mohajerin Esfahani and Kuhn (2018) P. Mohajerin Esfahani and D. Kuhn Data-driven distributionally robust optimization using the Wasserstein metric: performance guarantees and tractable reformulations. Mathematical Programming 171, p. 115–166. Cited by: Theorem 4. Nakao et al. (2025) H. Nakao, R. Jiang, and S. Shen Distributionally robust POMDPs with distance-based ambiguity sets. IISE Transactions. Cited by: §7. Ni and Bhat (2024) X. Ni and S. P. Bhat Stationary policies are optimal in risk-averse total-reward MDPs with EVaR. arXiv preprint arXiv:2408.17286. Cited by: §7. Nilim and El Ghaoui (2005) A. Nilim and L. El Ghaoui Robust control of Markov decision processes with uncertain transition matrices. Operations Research 53 (5), p. 780–798. Cited by: §7. Rockafellar and Uryasev (2000) R. T. Rockafellar and S. Uryasev Optimization of conditional value-at-risk. Journal of Risk 2, p. 21–42. Cited by: §7. Russel and Petrik (2019) R. H. Russel and M. Petrik Beyond confidence regions: tight Bayesian ambiguity sets for robust MDPs. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §7. Schwartz (1965) L. Schwartz On Bayes procedures. Zeitschrift für Wahrscheinlichkeitstheorie und verwandte Gebiete 4 (1), p. 10–26. Cited by: Appendix C, Theorem 3. Villani (2009) C. Villani Optimal transport: old and new. Springer. Cited by: §4. Wiesemann et al. (2013) W. Wiesemann, D. Kuhn, and B. Rustem Robust Markov decision processes. Mathematics of Operations Research 38 (1), p. 153–183. Cited by: §7.