Paper deep dive
Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk
Deep Kumar Ganguly, Jan Křetínský
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/21/2026, 4:48:32 AM
Summary
The paper introduces the Wasserstein Entropic Value-at-Risk (WEVaR), a coherent risk measure that replaces the relative-entropy ball of the standard Entropic Value-at-Risk (EVaR) with a Wasserstein (optimal transport) ball. This modification allows WEVaR to account for catastrophic events that have zero nominal probability, addressing a critical safety blind spot in EVaR. The authors derive a variational dual and closed-form expressions for WEVaR, proving its coherence and establishing a 'safety sandwich' bound. Furthermore, they propose a dynamic programming operator where the transport radius is driven by the agent's belief entropy, enabling a safety switch that adapts caution based on the agent's confidence level.
Entities (8)
Relation Signals (6)
Wasserstein Entropic Value-at-Risk → accountsfor → Zero-Nominal-Probability Catastrophes
confidence 96% · provably accounts for the reachable catastrophes the entropic measure ignores
Wasserstein Entropic Value-at-Risk → replaces → Relative Entropy Ball
confidence 95% · We keep the robust template (2) and replace the relative-entropy ball by a 1-Wasserstein ball
Entropic Value-at-Risk → failstoaccountfor → Zero-Nominal-Probability Catastrophes
confidence 94% · if P assigns zero probability to a catastrophic transition, no member assigns it positive probability... the safety check becomes vacuous
Belief Entropy → drives → Transport Radius
confidence 93% · Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator
Wasserstein Entropic Value-at-Risk → iscoherent → true
confidence 92% · Proposition 1 (Coherence). WEVaR is a coherent risk measure for every epsilon >= 0.
Wasserstein Entropic Value-at-Risk → hasdual → Kantorovich-Rubinstein Variational Dual
confidence 90% · prove a Kantorovich–Rubinstein variational dual (Thm. 1) that mirrors the entropic formula term for term
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk measure it induces, the Wasserstein entropic value-at-risk. It has a variational dual mirroring the entropic formula (an inverse temperature becomes a transport price), occupies a definite place in the risk hierarchy, and provably accounts for the reachable catastrophes the entropic measure ignores; we verify both dualities numerically. Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator whose caution contracts as the belief sharpens, with a certified safety sandwich and a sharp safety switch.
Tags
Links
- Source: https://arxiv.org/abs/2608.19073v1
- Canonical: https://arxiv.org/abs/2608.19073v1
Trouble viewing inline? Open PDF directly →
Full Text
44,474 characters extracted from source content.
Expand or collapse full text
Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk Deep Ganguly Thanks: Funded by the Deutsche Forschungsgemeinschaft (DFG) research training group ConVeY (GRK˜2428). Affiliation: Technical University of Munich, Germany Jan Křetínský Affiliation: Technical University of Munich, Germany Affiliation: Masaryk University, Brno, Czech Republic Abstract An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity—a confidence level fixes the radius of a relative-entropy ball of alternative models—but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk measure it induces, the Wasserstein entropic value-at-risk. It has a variational dual mirroring the entropic formula (an inverse temperature becomes a transport price), occupies a definite place in the risk hierarchy, and provably accounts for the reachable catastrophes the entropic measure ignores; we verify both dualities numerically. Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator whose caution contracts as the belief sharpens, with a certified safety sandwich and a sharp safety switch. 1 Risk Under Evolving Uncertainty Consider a drone crossing a canyon under uncertain wind. A high-cruise policy is fast when the air is calm but is thrown into the cliffs by a sudden gust; a hover-and-creep policy is safe in any wind but slow. The wind regime—a hidden type z drawn once—is never observed directly; the drone sees only noisy signals and must infer z as it flies. Two reflexes each fail. A controller that assumes the worst case forever (robust control [Iyengar 2005, Nilim and El Ghaoui 2005, Wiesemann et al. 2013]) keeps hovering even after the air has visibly calmed: safe, but so conservative that operators switch it off. A controller that commits before its belief sharpens (expected-value or Bayesian planning [Kaelbling et al. 1998, Ghavamzadeh et al. 2015]) cruises on a hunch, and a single wrong guess about z is fatal—an expected return that blends a benign mode with a catastrophic one is a number nobody actually attains. What is wanted is a controller whose caution evolves with its evidence—maximal under ignorance, vanishing once the environment is identified—together with a safety bound that holds at every stage of learning. The agent’s risk attitude must therefore be a function of its evolving epistemic state, and the entropy of its belief is a live readout of how much it should distrust its own model. The entropic value-at-risk [Ahmadi-Javid 2012] turns this intuition into algebra. It is the unique single-parameter coherent family that sweeps from the mean to the essential supremum, and it has an exact distributionally robust representation: its confidence level α equals the radius −lnα- α of a relative-entropy ball of alternative laws. Ganguly et al. 2025 further show the induced optimization is convex and admits well-posed, convergent estimation. This robust reading is what we want for evolving uncertainty: “how conservative should I be?” becomes “how large should my ambiguity set be?”, and the latter is answered by the agent’s own uncertainty. But the relative-entropy geometry has a blind spot that is decisive for safety. A ball Q:KL(Q∥P)≤ρ\Q:KL(Q\|P)≤ρ\ contains only laws absolutely continuous with respect to the nominal P: if P assigns zero probability to a catastrophic transition, no member assigns it positive probability, since KLKL is infinite off the support of P. As the agent learns and its nominal kernel concentrates away from rare disasters, the relative-entropy adversary loses the ability even to represent those disasters, so the safety check becomes vacuous precisely when overconfidence is most dangerous. The optimal-transport (Wasserstein) geometry has no such blind spot: it prices a perturbation by how far mass must physically move, so a verifier can still ask “what if it is a storm?” at a cost proportional to the distance to that outcome—reachable but expensive [Mohajerin Esfahani and Kuhn 2018, Gao and Kleywegt 2023, Blanchet and Murthy 2019]. Related work. Risk-averse dynamic programming [Ruszczyński 2010], risk-sensitive control and reinforcement learning [Howard and Matheson 1972, Chow et al. 2015, Hau et al. 2023], percentile and parameter-uncertainty MDPs [Delage and Mannor 2010], constrained and safe RL [Altman 1999, García and Fernández 2015, Sui et al. 2015, Brunke et al. 2022], and adversarially- or Wasserstein-robust RL [Pinto et al. 2017, Abdullah et al. 2019] all add caution to sequential decisions—but with a fixed attitude. Our contribution is to let the ambiguity radius, hence the risk attitude, be read off the agent’s evolving belief, and to give it the optimal-transport geometry safety requires [Kuhn et al. 2019]. Contributions. We keep the entropic value-at-risk’s robust formulation and its guarantees, and replace its ball. • We define the Wasserstein entropic value-at-risk (WEVaRWEVaR) and prove a Kantorovich–Rubinstein variational dual (Thm. 1) that mirrors the entropic formula term for term, is convex and well-posed (the transport analogue of the guarantee of Ganguly et al. 2025), and has a closed “mean-plus-Lipschitz” form. • We show it is coherent and place it in the risk hierarchy (Thm. 2): it is sandwiched against the entropic measure by a transport–entropy inequality and strictly accounts for zero-nominal-probability catastrophes the entropic measure ignores. • We verify both robust-optimization problems, their dualities, and the comparison numerically with Gurobi (§4), confirming strong duality to solver tolerance. • Driving the transport radius by belief entropy gives a closed-form robust dynamic-programming operator (Thm. 3) and a computable safety switch under evolving belief (§5). 2 The Entropic Value-at-Risk and Its Robust Formulation Let X be a bounded loss on a finite space S with reference law P, and let d:×→ℝ≥0d:S×S _≥ 0 be a ground metric with diameter D=maxs,s′d(s,s′)D= _s,s d(s,s ). A risk measure is coherent [Artzner et al. 1999] if it is monotone, translation-invariant, positively homogeneous and subadditive; the conditional value-at-risk [Rockafellar and Uryasev 2000] and the entropic value-at-risk are its canonical instances. The latter has the dual/primal pair EVaRα(X) _α(X) =inft>01t(lnP[etX]−lnα) = _t>0\ 1t ( _P[e^tX]- α ) (1) =supQ[X]:KL(Q∥P)≤−lnα, = \E_Q[X]:KL(Q\|P)≤- α \, (2) where the inverse temperature t in (1) is the multiplier on the relative-entropy constraint in (2), and the worst-case law is the exponential tilt Qs⋆∝Pset⋆XsQ _s P_se^t X_s [Ahmadi-Javid 2012, Ganguly et al. 2025]. Two facts we carry forward from Ganguly et al. 2025: (G1) the reparametrized objective is convex, so (1) is a well-posed one-dimensional convex program with a unique solution; (G2) the tilt Q⋆Q is supported on supp(P)supp(P). Property (G2) is the blind spot: EVaRα(X)EVaR_α(X) is independent of the value of X on any zero-nominal-probability state, for every α. 3 Swapping the Ball: A Wasserstein Robust Risk We keep the robust template (2) and replace the relative-entropy ball by a 11-Wasserstein ball, 1(Q,P)=min∑s,s′γ∈Γ(Q,P)d(s,s′)γ(s,s′)W_1(Q,P)= _γ∈ (Q,P) _s,s d(s,s )γ(s,s ). Definition 1 (WEVaRWEVaR). For radius ε≥0 ≥ 0, WEVaRε(X)=supQ[X]:1(Q,P)≤ε.WEVaR_ (X)= \E_Q[X]:W_1(Q,P)≤ \. (3) Theorem 1 (Variational dual, convexity, closed form). For bounded X and ε≥0 ≥ 0, WEVaRε(X)=infλ≥0λε+P[Xλc],WEVaR_ (X)= _λ≥ 0 \λ +E_P[\,X^c_λ\,] \, (4) where Xλc(s)=maxs′(X(s′)−λd(s′,s))X^c_λ(s)= _s (X(s )-λ\,d(s ,s)) is the c-transform of X. The objective is convex in λ, so (4) is a well-posed one-dimensional convex program (mirroring (G1)); the optimal price satisfies λ⋆≤Lipd(X)λ _d(X); ε↦WEVaRε(X) _ (X) is concave and nondecreasing; and WEVaRε(X)=P[X]+εLipd(X)WEVaR_ (X)=E_P[X]+ \,Lip_d(X) (5) for ε below a saturation threshold (in particular, exactly for two-point supports), where Lipd(X)=maxs≠s′|X(s)−X(s′)|/d(s,s′)Lip_d(X)= _s≠ s |X(s)-X(s )|/d(s,s ). Proof sketch. Strong duality for Wasserstein DRO [Gao and Kleywegt 2023, Blanchet and Murthy 2019] gives (4); the c-transform XλcX^c_λ is the transport counterpart of the cumulant 1tlnPetX 1t _Pe^tX in (1). Convexity in λ holds because XλcX^c_λ is a pointwise maximum of affine functions of λ, hence convex, and PE_P and the λελ term preserve convexity. For λ≥Lipd(X)λ _d(X) the maximum is attained at s′=s =s, so the bracket equals P[X]E_P[X] and the objective is P[X]+λεE_P[X]+λ , minimized at λ=Lipd(X)λ=Lip_d(X); this gives (5) until transporting all displaceable mass to the maximizer of X saturates the bound, after which the value rises concavely toward maxsX(s) _sX(s). ∎ Equations (1) and (4) are the same template—inf over a one-dimensional dual variable of a smoothed expectation plus radius-times-multiplier—with an entropic smoothing for the relative-entropy ball and a transport (inf-convolution) smoothing for the Wasserstein ball. The temperature t and the transport price λ are the same Lagrangian object on two ambiguity geometries. Proposition 1 (Coherence). WEVaRεWEVaR_ is a coherent risk measure for every ε≥0 ≥ 0. Proof sketch. It is the support function of the convex, compact set Q:1(Q,P)≤ε\Q:W_1(Q,P)≤ \, hence positively homogeneous and subadditive; P lies in the set, giving monotonicity; and adding a constant to X shifts every Q[X]E_Q[X] equally, giving translation invariance. ∎ Theorem 2 (Hierarchy and relation to the entropic measure). On a finite metric space with diameter D: (i) WEVaR0(X)=P[X]WEVaR_0(X)=E_P[X] and WEVaRε(X)↑esssup(X)WEVaR_ (X) *ess\,sup(X) as ε↑D D, so WEVaRWEVaR sweeps the full hierarchy; (i) (sandwich) EVaRα(X)≤WEVaRD−12lnα(X)EVaR_α(X) _\,D - 12 α(X) for all α∈(0,1]α∈(0,1]; (i) (catastrophe) if P(s⋆)=0P(s )=0 and X(s⋆)>P[X]X(s )>E_P[X], then EVaRα(X)EVaR_α(X) is independent of X(s⋆)X(s ) for every α, whereas WEVaRε(X)WEVaR_ (X) strictly increases in X(s⋆)X(s ) once ε>distd(s⋆,suppP) >dist_d(s ,supp\,P). Proof sketch. (i) is immediate from finiteness. (i): Pinsker gives total variation ≤KL/2≤ KL/2 and 1≤D⋅TVW_1≤ D·TV, so the relative-entropy ball of radius −lnα- α is contained in the Wasserstein ball of radius D−12lnαD - 12 α; take suprema [Bobkov and Götze 1999, Boucheron et al. 2013]. (i) is (G2) versus the transport reach: no finite-radius relative-entropy ball contains a Wasserstein ball, and the gap is exactly the reachable-but-zero-probability catastrophes. ∎ Figure 1: Why Wasserstein, geometrically. On the next-state simplex at a confident belief—the nominal P¯b P_b places zero mass on Fail—the relative-entropy ball (red) is trapped on the support edge, blind to the catastrophe, while the 11-Wasserstein ball (blue) reaches the Fail vertex; the worst-case adversary (×) sits on its boundary. This is Thm. 2(i) drawn on Δ() (S). 4 Empirical Guarantees We verify both robust programs and their duals with Gurobi 12 on a five-state metric space (X=[0,2,4,8,100]X=[0,2,4,8,100], d(s,s′)=|s−s′|d(s,s )=|s-s |). For the entropic measure we solve the relative-entropy-constrained program (2) (a nonlinear program) and match it against the convex primal (1); for WEVaRWEVaR we solve the transport linear program (3) and match it against the Kantorovich–Rubinstein dual (4). Strong duality holds to solver tolerance: the entropic primal (1) and the relative-entropy program (2) agree to 1.2×10−41.2× 10^-4, and the transport linear program agrees with its Kantorovich–Rubinstein dual to 9×10−79× 10^-7. The closed form (5) is exact in the unsaturated regime (here ε≤0.1 ≤ 0.1) and is otherwise upper-bounded by the exact one-dimensional dual, which the solver confirms. Across a sweep of radii both measures rise from the mean toward the worst case, with EVaRα≤WEVaREVaR_α at the Pinsker-matched radius (Thm. 2(i)). The catastrophe blind spot is decisive: with a zero-nominal-probability disaster whose loss we grow from 5050 to 10001000, the entropic value-at-risk changes by exactly 00 at fixed confidence, while WEVaRWEVaR at a fixed radius changes by 807.5807.5—the gap of Thm. 2(i), shown geometrically in Fig. 1 and quantitatively across radii in Fig. 4 (App. E). 5 Evolving Uncertainty: Belief Entropy as the Radius We now let the radius evolve. An environment has a hidden type z∈z drawn once; the agent maintains a Bayesian belief b∈Δ()b∈ (Z) with update ψ, forms the belief-weighted nominal kernel P¯b(⋅∣s,a)=∑zb(z)P(⋅∣s,a,z) P_b(· s,a)= _zb(z)P(· s,a,z), and sets the transport radius to the belief entropy, ε(b)=βℋ(b) (b)=β\,H(b) with ℋ(b)=−∑zb(z)lnb(z)H(b)=- _zb(z) b(z) and sensitivity β>0β>0, giving the ambiguity set (b)=Q:1(Q,P¯b)≤βℋ(b)U(b)=\Q:W_1(Q, P_b)≤ (b)\. As observations sharpen b, the entropy ℋ(b)↓0H(b) 0, the ball contracts and drifts toward the nominal (Fig. 2), and via Thm. 1 the agent’s risk attitude slides from worst-case to risk-neutral—a dial driven by epistemic state. Since the controller minimizes value, the relevant functional is the infimal twin of Def. 1, whose closed form is P¯b[V]−βℋ(b)Lipd(V)E_ P_b[V]- (b)\,Lip_d(V). Figure 2: The entropy-modulated ambiguity set (b)U(b) on the next-state simplex. As the belief sharpens the radius ε=βℋ(b) = (b) contracts (→0.201.04\!→\!0.20) and the nominal P¯b P_b drifts toward the goal; the worst-case adversary (×) loses its reach toward the failure vertex—robustness that vanishes with uncertainty (the geometric content of Thm. 3). Theorem 3 (Closed-form robust update; contraction; safety). For the operator, with discount γ∈(0,1]γ∈(0,1] and belief update b′=ψ(b,s,a,s′)b =ψ(b,s,a,s ), (V)(s,b)=maxa[r(s,a)+γinfQ∈(b)QV(s′,b′)]:( TV)(s,b)= _a [r(s,a)+γ\! _Q (b)\!E_Q\,V(s ,b ) ]: (i) the inner problem has the closed form of Thm. 1—no coupling linear program—reducing the per-(s,a,b)(s,a,b) cost to O(||2)O(|S|^2); (i) T is a γ-contraction in ∥⋅∥∞\|·\|_∞ when γ<1γ<1, and a contraction on proper MDPs when γ=1γ=1; either way it has a unique fixed point V⋆V [Bertsekas and Tsitsiklis 1996]; (i) V⋆V obeys a safety sandwich Vwc⋆(s,b)≤V⋆(s,b)≤z∼b[Vopt⋆(s,z)]V _wc(s,b)≤ V (s,b) _z b[V _opt(s,z)] between the always-maximally-cautious value and the type-aware oracle, and as ℋ(b)→0H(b)→ 0 the radius vanishes and V⋆(s,b)→Vopt⋆(s,z⋆)V (s,b)→ V _opt(s,z ). Since ℋ(b)≤ln||H(b)≤ |Z|, the worst-case radius is βln||β |Z|; requiring one action to remain viable under maximal ignorance gives the calibration β<D/ln||β<D/ |Z|, the transport analogue of choosing a confidence level. Figure 3: Computed guarantees on a five-state corridor with a hidden slip type (γ=0.9γ=0.9, Gurobi-checked). (a) The adaptive value V⋆(s0,b)V (s_0,b) stays inside the safety sandwich between the always-maximally-cautious floor VwcV_wc and the type-aware oracle ceiling; the price of robustness shrinks as the belief sharpens. (b) From arbitrary initializations, value iteration with the closed-form operator collapses geometrically onto V⋆V along the γkγ^k envelope, certifying the bound at every iterate. Evolving safety switch. On the Ambiguous Bridge (γ=1γ=1), Sprint reaches the goal (+100+100) under a benign type but falls to catastrophe (−1000-1000) under an adversarial type, while Crawl is always safe but costly. With belief b=[α,1−α]b=[α,1-α] and d(SG,SF)=1d(S_G,S_F)=1, so Lipd(V)=1100Lip_d(V)=1100, Thm. 3(i) gives in closed form Q(Sprint)=−1+1100(α−βℋ(α))+−1000Q( Sprint)=-1+1100(α- (α))^+-1000 and Q(Crawl)=80−1100min(βℋ(α),1)Q( Crawl)=80-1100 ( (α),1), which a Gurobi transport solve reproduces to 10−1310^-13. Equating them (the entropy terms cancel) yields a safety switch independent of β: the agent Crawls until α≥α⋆=1081/1100≈0.983α≥α =1081/1100≈ 0.983, then Sprints. The threshold is set purely by the reward asymmetry. The same creep-then-cruise switch governs the canyon of §1: the drone hovers until its confidence that the air is calm crosses ≈0.63≈ 0.63, then cruises (Fig. 6, App. E). Guarantees, illustrated. On a five-state corridor whose forward action slips to a failure state only under a hidden “storm” type (Fig. 3), the safety sandwich of Thm. 3(i) holds at every belief: the adaptive value never drops below the always-maximally-cautious floor and never exceeds the type-aware oracle. The price of robustness is temporary—the ceiling gap falls from 8.48.4 at the uniform belief to 4.44.4 once the belief reaches b(calm)=0.95b(calm)=0.95, and to zero at identification. Value iteration with the closed-form operator converges from arbitrary initializations at empirical rate 0.90=γ0.90=γ, and the closed-form inner update matches a Gurobi transport solve at every sampled belief (the underlying duality is verified to solver tolerance in §4), so a valid safety bound is available at every iterate, not only at convergence. The belief-simplex manifolds (Fig. 5) and the safety-vs-efficiency rollouts (Fig. 7) appear in App. E. 6 Discussion The entropic and Wasserstein robust risks are the relative-entropy and optimal-transport faces of one idea—a coherent risk whose ambiguity radius is the agent’s epistemic uncertainty—bridged by entropic optimal transport [Cuturi 2013, Peyré and Cuturi 2019]; scaling the closed-form operator to continuous spaces via Lipschitz critics and to active information gathering are the natural next steps. References Abdullah et al. [2019] Mohammed Amin Abdullah, Hang Ren, Haitham Bou Ammar, Vladimir Milenkovic, Rui Luo, Mingtian Zhang, and Jun Wang. Wasserstein robust reinforcement learning. arXiv preprint arXiv:1907.13196, 2019. Ahmadi-Javid [2012] Amir Ahmadi-Javid. Entropic value-at-risk: A new coherent risk measure. Journal of Optimization Theory and Applications, 155(3):1105–1123, 2012. Altman [1999] Eitan Altman. Constrained Markov Decision Processes. Chapman & Hall/CRC, 1999. Artzner et al. [1999] Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath. Coherent measures of risk. Mathematical Finance, 9(3):203–228, 1999. Bertsekas and Tsitsiklis [1996] Dimitri P. Bertsekas and John N. Tsitsiklis. Neuro-Dynamic Programming. Athena Scientific, 1996. Blanchet and Murthy [2019] Jose Blanchet and Karthyek Murthy. Quantifying distributional model risk via optimal transport. Mathematics of Operations Research, 44(2):565–600, 2019. Bobkov and Götze [1999] Sergey G. Bobkov and Friedrich Götze. Exponential integrability and transportation cost related to logarithmic Sobolev inequalities. Journal of Functional Analysis, 163(1):1–28, 1999. Boucheron et al. [2013] Stéphane Boucheron, Gábor Lugosi, and Pascal Massart. Concentration Inequalities: A Nonasymptotic Theory of Independence. Oxford University Press, 2013. Brunke et al. [2022] Lukas Brunke, Melissa Greeff, Adam W. Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P. Schoellig. Safe learning in robotics: From learning-based control to safe reinforcement learning. Annual Review of Control, Robotics, and Autonomous Systems, 5:411–444, 2022. Chow et al. [2015] Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone. Risk-sensitive and robust decision-making: A CVaR optimization approach. In Advances in Neural Information Processing Systems (NeurIPS), 2015. Cuturi [2013] Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In Advances in Neural Information Processing Systems (NeurIPS), 2013. Delage and Mannor [2010] Erick Delage and Shie Mannor. Percentile optimization for Markov decision processes with parameter uncertainty. Operations Research, 58(1):203–213, 2010. Ganguly et al. [2025] Deep Ganguly, Sarthak Girotra, Sirish Sekhar, and Ajin George Joseph. Risk-seeking reinforcement learning via multi-timescale entropic value-at-risk optimization. Transactions on Machine Learning Research, 2025. Gao and Kleywegt [2023] Rui Gao and Anton J. Kleywegt. Distributionally robust stochastic optimization with Wasserstein distance. Mathematics of Operations Research, 48(2):603–655, 2023. García and Fernández [2015] Javier García and Fernando Fernández. A comprehensive survey on safe reinforcement learning. Journal of Machine Learning Research, 16:1437–1480, 2015. Ghavamzadeh et al. [2015] Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar. Bayesian reinforcement learning: A survey. Foundations and Trends in Machine Learning, 8(5–6):359–483, 2015. Hau et al. [2023] Jia Lin Hau, Marek Petrik, and Mohammad Ghavamzadeh. Entropic risk optimization in discounted MDPs. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2023. Howard and Matheson [1972] Ronald A. Howard and James E. Matheson. Risk-sensitive Markov decision processes. Management Science, 18(7):356–369, 1972. Iyengar [2005] Garud N. Iyengar. Robust dynamic programming. Mathematics of Operations Research, 30(2):257–280, 2005. Kaelbling et al. [1998] Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra. Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1–2):99–134, 1998. Kuhn et al. [2019] Daniel Kuhn, Peyman Mohajerin Esfahani, Viet Anh Nguyen, and Soroosh Shafieezadeh-Abadeh. Wasserstein distributionally robust optimization: Theory and applications in machine learning. In Operations Research & Management Science in the Age of Analytics (INFORMS TutORials), pages 130–166. INFORMS, 2019. Mohajerin Esfahani and Kuhn [2018] Peyman Mohajerin Esfahani and Daniel Kuhn. Data-driven distributionally robust optimization using the Wasserstein metric: Performance guarantees and tractable reformulations. Mathematical Programming, 171:115–166, 2018. Nilim and El Ghaoui [2005] Arnab Nilim and Laurent El Ghaoui. Robust control of Markov decision processes with uncertain transition matrices. Operations Research, 53(5):780–798, 2005. Peyré and Cuturi [2019] Gabriel Peyré and Marco Cuturi. Computational optimal transport. Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019. Pinto et al. [2017] Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta. Robust adversarial reinforcement learning. In International Conference on Machine Learning (ICML), 2017. Rockafellar and Uryasev [2000] R. Tyrrell Rockafellar and Stanislav Uryasev. Optimization of conditional value-at-risk. Journal of Risk, 2:21–42, 2000. Ruszczyński [2010] Andrzej Ruszczyński. Risk-averse dynamic programming for Markov decision processes. Mathematical Programming, 125(2):235–261, 2010. Sui et al. [2015] Yanan Sui, Alkis Gotovos, Joel Burdick, and Andreas Krause. Safe exploration for optimization with Gaussian processes. In International Conference on Machine Learning (ICML), 2015. Villani [2009] Cédric Villani. Optimal Transport: Old and New. Springer, 2009. Wiesemann et al. [2013] Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem. Robust Markov decision processes. Mathematics of Operations Research, 38(1):153–183, 2013. Appendix: Extended Proofs Throughout, S is finite with ||=n|S|=n; d is a metric on S with diameter D=maxs,s′d(s,s′)D= _s,s d(s,s ); P∈Δ()P∈ (S) is the reference law; X:→ℝX:S is a bounded loss; and 1(Q,P)=min∑s,s′γ∈Γ(Q,P)d(s,s′)γ(s,s′)W_1(Q,P)= _γ∈ (Q,P) _s,s d(s,s )\,γ(s,s ) with Γ(Q,P) (Q,P) the couplings of Q and P. We write Q[X]=∑sQ(s)X(s)E_Q[X]= _sQ(s)X(s). A. Theorem 1 (a) Duality. Since 1(Q,P)W_1(Q,P) is itself a minimum over couplings, the worst-case expectation is the value of a single linear program over the coupling γ≥0γ≥ 0 whose first marginal is fixed to P: WEVaRε(X)=max∑s,s′γ≥0γ(s,s′)X(s′)s.t.∑s′γ(s,s′)=P(s)∀s,∑s,s′d(s,s′)γ(s,s′)≤ε,WEVaR_ (X)= _γ≥ 0\ _s,s γ(s,s )\,X(s ) .t. _s γ(s,s )=P(s)\ \ ∀ s, _s,s d(s,s )\,γ(s,s )≤ , (6) where the perturbed law is the second marginal Q(s′)=∑sγ(s,s′)Q(s )= _sγ(s,s ) (automatically in Δ() (S) because ∑s,s′γ=∑sP(s)=1 _s,s γ= _sP(s)=1), and the objective equals Q[X]E_Q[X]. The program is feasible (take γ(s,s′)=P(s)[s′=s]γ(s,s )=P(s)1[s =s], of transport cost 0≤ε0≤ ) and bounded (X is bounded), so LP strong duality applies. Attaching a free multiplier u(s)u(s) to each marginal equality and λ≥0λ≥ 0 to the budget, the Lagrangian is ∑sP(s)u(s)+λε+∑s,s′γ(s,s′)(X(s′)−u(s)−λd(s,s′)). _sP(s)u(s)+λ + _s,s γ(s,s ) (X(s )-u(s)-λ\,d(s,s ) ). Its supremum over γ≥0γ≥ 0 is finite iff X(s′)−u(s)−λd(s,s′)≤0X(s )-u(s)-λ\,d(s,s )≤ 0 for all s,s′s,s , i.e. u(s)≥maxs′(X(s′)−λd(s,s′))=Xλc(s)u(s)≥ _s (X(s )-λ\,d(s,s ) )=X^c_λ(s), in which case the supremum equals ∑sP(s)u(s)+λε _sP(s)u(s)+λ (attained at γ=0γ=0). Minimizing over feasible u sets u(s)=Xλc(s)u(s)=X^c_λ(s), giving WEVaRε(X)=minλ≥0λε+∑sP(s)Xλc(s)=infλ≥0λε+P[Xλc],WEVaR_ (X)= _λ≥ 0 \λ + _sP(s)\,X^c_λ(s) \= _λ≥ 0 \λ +E_P[X^c_λ] \, which is (4). (For general Polish spaces this is Kantorovich–Rubinstein/Wasserstein-DRO strong duality, Villani 2009, Gao and Kleywegt 2023, Blanchet and Murthy 2019.) (b) Convexity and λ⋆≤Lipd(X)λ _d(X). For each fixed s, λ↦Xλc(s)=maxs′(X(s′)−λd(s,s′))λ X^c_λ(s)= _s (X(s )-λ\,d(s,s ) ) is a pointwise maximum of functions affine in λ, hence convex; therefore P[Xλc]=∑sP(s)Xλc(s)E_P[X^c_λ]= _sP(s)X^c_λ(s) is convex (nonnegative combination), and ϕ(λ):=λε+P[Xλc]φ(λ):=λ +E_P[X^c_λ] is convex on [0,∞)[0,∞), so (4) is a well-posed one-dimensional convex program. Next, if λ≥Lipd(X)λ _d(X) then for every s,s′s,s , X(s′)−X(s)≤Lipd(X)d(s,s′)≤λd(s,s′)X(s )-X(s) _d(X)\,d(s,s )≤λ\,d(s,s ), so X(s′)−λd(s,s′)≤X(s)X(s )-λ\,d(s,s )≤ X(s) with equality at s′=s =s; hence Xλc(s)=X(s)X^c_λ(s)=X(s) and ϕ(λ)=λε+P[X]φ(λ)=λ +E_P[X], which is strictly increasing in λ. A convex ϕφ that is increasing on [Lipd(X),∞)[Lip_d(X),∞) attains its minimum at some λ⋆≤Lipd(X)λ _d(X). (c) Monotonicity and concavity in ε . The feasible set Q:1(Q,P)≤ε\Q:W_1(Q,P)≤ \ grows with ε , so WEVaRε(X)WEVaR_ (X) is nondecreasing. By (4) it is an infimum over λ of the maps ε↦λε+P[Xλc] λ +E_P[X^c_λ], each affine in ε ; an infimum of affine functions is concave. (d) Closed form and saturation. Let T(s)∈argmaxs′(X(s′)−Lipd(X)d(s,s′))T(s)∈ _s (X(s )-Lip_d(X)\,d(s,s ) ) be a steepest-ascent target and set ε¯(X):=P[d(s,T(s))]=∑sP(s)d(s,T(s)). (X):=E_P [d(s,T(s)) ]= _sP(s)\,d (s,T(s) ). The right derivative of the convex function P[Xλc]E_P[X^c_λ] at λ=Lipd(X)λ=Lip_d(X) equals −∑sP(s)d(s,T(s))=−ε¯(X)- _sP(s)\,d(s,T(s))=- (X) (each active term contributes the negative source–target distance, inactive terms contribute 00). Hence the left derivative of ϕφ at Lipd(X)Lip_d(X) is ε−ε¯(X) - (X), which is ≤0≤ 0 iff ε≤ε¯(X) ≤ (X); in that regime the minimizer is λ⋆=Lipd(X)λ =Lip_d(X) and WEVaRε(X)=Lipd(X)ε+P[XLipd(X)c]=P[X]+εLipd(X).WEVaR_ (X)=Lip_d(X)\, +E_P[X^c_Lip_d(X)]=E_P[X]+ \,Lip_d(X). For |supp(P)∪T(s)|=2|supp(P)∪\T(s)\|=2 there is a single transport channel, ε¯(X)=1(δargmaxX,⋅) (X)=W_1( _ X,\,·) exhausts only at a vertex, and the formula is exact up to that point. □ B. Proposition 1 (Coherence) Write ρ(X)=WEVaRε(X)=supQ∈BQ[X]ρ(X)=WEVaR_ (X)= _Q∈ BE_Q[X] with B=Q:1(Q,P)≤εB=\Q:W_1(Q,P)≤ \, a nonempty (P∈BP∈ B), convex, compact subset of Δ() (S). Monotonicity: if X≤YX≤ Y pointwise then Q[X]≤Q[Y]E_Q[X] _Q[Y] for every Q∈BQ∈ B (as Q≥0Q≥ 0), so ρ(X)≤ρ(Y)ρ(X)≤ρ(Y). Translation invariance: for c∈ℝc , Q[X+c]=Q[X]+cE_Q[X+c]=E_Q[X]+c for all Q, hence ρ(X+c)=ρ(X)+cρ(X+c)=ρ(X)+c. Positive homogeneity: for t≥0t≥ 0, ρ(tX)=supQ∈BtQ[X]=tρ(X)ρ(tX)= _Q∈ Bt\,E_Q[X]=t\,ρ(X). Subadditivity: ρ(X+Y)=supQ∈B(Q[X]+Q[Y])≤supQ∈BQ[X]+supQ∈BQ[Y]=ρ(X)+ρ(Y)ρ(X+Y)= _Q∈ B (E_Q[X]+E_Q[Y] )≤ _Q∈ BE_Q[X]+ _Q∈ BE_Q[Y]=ρ(X)+ρ(Y). These are the axioms of Artzner et al. 1999; B is the risk envelope of the induced dual representation. □ C. Theorem 2 (i) Sweep from mean to worst case. 1W_1 is a metric on Δ() (S), so 1(Q,P)=0⇔Q=PW_1(Q,P)=0 Q=P and WEVaR0(X)=P[X]WEVaR_0(X)=E_P[X]. For any Q, Q[X]≤maxsX(s)E_Q[X]≤ _sX(s), so WEVaRε(X)≤maxsX(s)WEVaR_ (X)≤ _sX(s). Let s∙∈argmaxsX(s)s ∈ _sX(s). Then 1(δs∙,P)=∑sP(s)d(s,s∙)≤DW_1( _s ,P)= _sP(s)\,d(s,s )≤ D, so δs∙∈B _s ∈ B once ε≥∑sP(s)d(s,s∙) ≥ _sP(s)d(s,s ), whence WEVaRε(X)=maxsX(s)WEVaR_ (X)= _sX(s). Combined with monotonicity (Thm. 1(c)), WEVaRε(X)↑maxsX(s)WEVaR_ (X) _sX(s) as ε↑D D. Lemma (transport vs. total variation). For all Q,PQ,P, 1(Q,P)≤D⋅TV(Q,P)W_1(Q,P)≤ D·TV(Q,P), where TV(Q,P)=12∑s|Q(s)−P(s)|TV(Q,P)= 12 _s|Q(s)-P(s)|. Proof: let m(s)=min(Q(s),P(s))m(s)= (Q(s),P(s)); the coupling that keeps mass m in place (γ(s,s)≥m(s)γ(s,s)\!≥\!m(s)) and transports the residual mass ∑s(Q(s)−m(s))=TV(Q,P) _s(Q(s)-m(s))=TV(Q,P) arbitrarily incurs cost ≤D⋅TV(Q,P)≤ D·TV(Q,P), and 1W_1 is the minimal cost. ⋄ (i) Sandwich. Fix α∈(0,1]α∈(0,1] and put ρ=−lnαρ=- α. For any Q with KL(Q∥P)≤ρKL(Q\|P)≤ρ, Pinsker’s inequality gives TV(Q,P)≤KL(Q∥P)/2≤ρ/2TV(Q,P)≤ KL(Q\|P)/2≤ ρ/2, so by the Lemma 1(Q,P)≤Dρ/2=D−12lnαW_1(Q,P)≤ D ρ/2=D - 12 α. Therefore Q:KL(Q∥P)≤−lnα⊆Q:1(Q,P)≤D−12lnα,\Q:KL(Q\|P)≤- α\ \Q:W_1(Q,P)≤ D - 12 α\, and taking the supremum of Q[X]E_Q[X] over the larger set, EVaRα(X)=supKL(Q∥P)≤−lnαQ[X]≤sup1(Q,P)≤D−12lnαQ[X]=WEVaRD−12lnα(X).EVaR_α(X)=\!\! _KL(Q\|P)≤- α\!\!E_Q[X]\ ≤\!\! _W_1(Q,P)≤ D - 12 α\!\!E_Q[X]=WEVaR_\,D - 12 α(X). (i) Catastrophe domination. Assume P(s⋆)=0P(s )=0 and X(s⋆)>P[X]X(s )>E_P[X]. Entropic measure is blind. P[etX]=∑s:P(s)>0P(s)etX(s)E_P[e^tX]= _s:P(s)>0P(s)e^tX(s) omits the term s⋆s , so it—and hence EVaRα(X)=inft>01t(lnP[etX]−lnα)EVaR_α(X)= _t>0 1t( _P[e^tX]- α)—does not depend on X(s⋆)X(s ), for every α. (Dually, any Q with KL(Q∥P)<∞KL(Q\|P)<∞ satisfies Q≪PQ P, so Q(s⋆)=0Q(s )=0 and Q[X]E_Q[X] ignores X(s⋆)X(s ).) Transport measure reacts. Let δ=distd(s⋆,suppP)=mins:P(s)>0d(s,s⋆)δ=dist_d(s ,supp\,P)= _s:P(s)>0d(s,s ) and s0s_0 an achieving state. For ε>δ >δ choose η0=min(P(s0),ε/δ)>0 _0= (P(s_0), /δ )>0 and the feasible law Qη0=P−η0δs0+η0δs⋆Q_ _0=P- _0 _s_0+ _0 _s , for which 1(Qη0,P)≤η0δ≤εW_1(Q_ _0,P)≤ _0\,δ≤ . Then WEVaRε(X)≥Qη0[X]=P[X]+η0(X(s⋆)−X(s0)),WEVaR_ (X)\ ≥\ E_Q_ _0[X]=E_P[X]+ _0 (X(s )-X(s_0) ), which is strictly increasing in X(s⋆)X(s ). Moreover WEVaRε(X)=max∑sQ∈BQ(s)X(s)WEVaR_ (X)= _Q∈ B _sQ(s)X(s) is convex in the vector X, and by Danskin’s theorem its partial right-derivative in X(s⋆)X(s ) equals maxQ(s⋆):Q optimal \Q(s ):Q optimal\; for X(s⋆)X(s ) large enough any optimal Q must place positive mass on s⋆s (else Qη0Q_ _0 strictly improves), so the derivative is positive and WEVaRε(X)WEVaR_ (X) strictly increases in X(s⋆)X(s ) whenever ε>δ >δ. □ D. Theorem 3 Fix (s,a,b)(s,a,b) and write the continuation vector W(s′)=V(s′,ψ(b,s,a,s′))W(s )=V (s ,ψ(b,s,a,s ) ) and radius ε(b)=βℋ(b) (b)= (b), nominal P¯b P_b. (i) Closed form of the inner minimization. Applying Theorem 1 to the loss −W-W and using (−W)λc(s)=maxs′(−W(s′)−λd(s′,s))=−mins′(W(s′)+λd(s′,s))(-W)^c_λ(s)= _s (-W(s )-λ d(s ,s))=- _s (W(s )+λ d(s ,s)), infQ∈(b)Q[W]=−WEVaRε(b)(−W)=supλ≥0P¯b[Wλc,−]−λε(b),Wλc,−(s)=mins′(W(s′)+λd(s,s′)), _Q (b)E_Q[W]=-WEVaR_ (b)(-W)= _λ≥ 0 \E_ P_b [W^c,-_λ ]-λ\, (b) \, W^c,-_λ(s)= _s (W(s )+λ\,d(s,s ) ), a one-dimensional concave maximization, with first-order/unsaturated value P¯b[W]−βℋ(b)Lipd(W)E_ P_b[W]- (b)\,Lip_d(W) (Thm. 1(d) applied to −W-W). Evaluating Lipd(W)Lip_d(W) costs O(n2)O(n^2) and the dual is a scalar program, versus an O(n2)O(n^2)-variable transportation LP (6). (i) Contraction. We use two properties of T. Monotonicity: if V≤V′V≤ V then W≤W′W≤ W pointwise, so infQQ[W]≤infQQ[W′] _QE_Q[W]≤ _QE_Q[W ] and, taking maxa _a, V≤V′ TV≤ TV . Scaled constant shift: for c∈ℝc , adding c to V adds c to every continuation, and infQQ[W+c]=infQQ[W]+c _QE_Q[W+c]= _QE_Q[W]+c, so (V+c)=V+γc T(V+c1)= TV+γ c1. For γ<1γ<1, monotonicity and the scaled shift are exactly the hypotheses of the contraction lemma [Bertsekas and Tsitsiklis 1996, Ch. 2]: V−V′≤γ‖V−V′‖∞ TV- TV ≤γ\|V-V \|_∞1 and symmetrically, so ‖V−V′‖∞≤γ‖V−V′‖∞\| TV- TV \|_∞≤γ\|V-V \|_∞, a γ-contraction. For γ=1γ=1 the shift is exact and gives nonexpansiveness; under properness—every policy together with every kernel selection in the sets (⋅)U(·) reaches the terminal set T in uniformly bounded expected time—the standard stochastic-shortest-path argument [Bertsekas and Tsitsiklis 1996, Ch. 3] upgrades this to a contraction in a weighted supremum norm ∥⋅∥w\|·\|_w: there exist m≥1,κ∈[0,1)m≥ 1,\ κ∈[0,1) with ‖mV−mV′‖w≤κ‖V−V′‖w\| T^mV- T^mV \|_w≤κ\|V-V \|_w. Either way T has a unique fixed point V⋆V . (i) Safety sandwich. Upper bound. Since P¯b∈(b) P_b (b), infQ∈(b)Q[W]≤P¯b[W] _Q (b)E_Q[W] _ P_b[W], so V≤BayesV TV≤ T_BayesV pointwise, where Bayes T_Bayes is the non-robust operator with kernel P¯b=∑zb(z)P(⋅∣s,a,z) P_b= _zb(z)P(· s,a,z); by monotone iteration V⋆≤VBayesV ≤ V_Bayes. Committing to one policy under belief b cannot beat knowing the type, so VBayes(s,b)≤z∼b[Vopt⋆(s,z)]V_Bayes(s,b) _z b[V _opt(s,z)] (nonnegative value of information), giving the upper bound. Lower bound. Let Vwc⋆V _wc be the fixed point of the operator wc T_wc obtained by replacing the radius βℋ(b) (b) with its maximum βln||β |Z| at every (s,b)(s,b), keeping the same nominal P¯b P_b. Since βℋ(b)≤βln|| (b)≤β |Z|, the ambiguity ball in T is contained in that of wc T_wc, so its inner infimum is no smaller; hence V≥wcV TV≥ T_wcV pointwise and, by monotone iteration, V⋆(s,b)≥Vwc⋆(s,b)V (s,b)≥ V _wc(s,b). This is the bound used in Fig. 3 and requires no nested-set condition. Convergence. As b→δz⋆b→ _z , ℋ(b)→0H(b)→ 0 and ε(b)→0 (b)→ 0, so (b)→P(⋅∣s,a,z⋆)U(b)→\P(· s,a,z )\; by the Lipschitz dependence of the value on the radius (Thm. 1(c): 0≤WEVaRε−WEVaR0≤εLipd0 _ -WEVaR_0≤ _d), V⋆(s,b)→Vopt⋆(s,z⋆)V (s,b)→ V _opt(s,z ). □ E. Supplementary Experiments and Figures All figures are reproduced by the released code (a Colab-ready package). The body shows the geometric core (Figs. 1, 2) and the certified guarantee (Fig. 3); the remaining plots are collected here. Quantitative verification of the dualities. Figure 4 is the solver-side companion to §4: across a sweep of radii the entropic and Wasserstein measures both rise from the mean to the worst case, with EVaRα≤WEVaREVaR_α at the Pinsker-matched radius (Thm. 2(i)); and a zero-nominal-probability disaster leaves the entropic measure flat while WEVaRWEVaR reacts (Thm. 2(i)). Figure 4: Gurobi-verified comparison on a five-state metric space. (a) hierarchy and the Pinsker sandwich; (b) the zero-probability catastrophe: the entropic value-at-risk is invariant to the disaster’s magnitude, WEVaRWEVaR is not. Manifolds over the belief simplex. For a three-type environment the belief lives on a 22-simplex, and the controller reads three scalar fields off it (Fig. 5): the radius ε(b)=βℋ(b) (b)= (b) (an entropy bowl, maximal at the centroid and zero at the vertices), the robust value V⋆(s0,b)V (s_0,b), and the policy regions. Figure 5: Manifolds over the belief simplex Δ() (Z), ||=3|Z|=3 (calm/breezy/storm): ambiguity radius, robust value, and policy region (cruise only near the calm vertex). Safety switches. The belief-dependent policy flips from the safe to the optimal action at a computable threshold (Fig. 6): on the Ambiguous Bridge at α⋆≈0.983α ≈ 0.983 (set by the reward asymmetry, independent of β), and on the stormy-drone canyon at b(calm)≈0.63b(calm)≈ 0.63. Figure 6: Robust action values vs. belief. Left: the Ambiguous Bridge (Sprint vs. Crawl, switch at α⋆≈0.983α ≈ 0.983). Right: the stormy-drone canyon (Cruise vs. Creep, switch at b(calm)≈0.63b(calm)≈ 0.63). Safety-vs-efficiency rollouts, and an honest caveat. Figure 7 compares the three planners on the canyon. WEVaRWEVaR Pareto-dominates the static-robust (Maximin) planner—comparable safety at higher return, the “un-freezing” effect. In this calibrated-belief setting the expected-value (Bayesian) planner is already near-optimally cautious under a severe catastrophe, so the empirical separation between WEVaRWEVaR and Bayesian is small; the regime in which belief-scaled caution measurably reduces failures is one of miscalibrated belief or distribution shift, where the belief is learned rather than computed from a known model. The body’s contribution is therefore the coherent measure, its geometry, and the certified guarantee; this rollout is illustrative. Figure 7: Catastrophe rate vs. time-to-goal on the stormy-drone canyon (40004000 episodes). WEVaRWEVaR dominates the static-robust Maximin baseline; the Bayesian planner is already cautious here (see caveat).