Paper deep dive
On-Policy and Off-Policy Learning for Large Action Spaces
Imad Aouali
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/3/2026, 10:10:52 AM
Summary
This thesis addresses policy learning in interactive systems with large action spaces, focusing on contextual bandits. It introduces structured Bayesian methods for on-policy learning, including mixed-effect Thompson sampling (meTS) and diffusion Thompson sampling (dTS), to improve exploration efficiency and reduce regret. For off-policy learning, it proposes the structured direct method (sDM) using latent variables, advocates for policy-weighted log-likelihood objectives to address optimization bottlenecks, and develops differentiable pessimistic methods using exponential smoothing and PAC-Bayesian bounds to control bias-variance trade-offs.
Entities (11)
Relation Signals (9)
meTS → isa → Thompson Sampling
confidence 95% · meTS, a mixed-effect extension of Thompson sampling
dTS → isa → Thompson Sampling
confidence 95% · diffusion Thompson sampling (dTS)... leverages diffusion-inspired priors
dTS → usedfor → on-policy learning
confidence 95% · We extend this framework to diffusion Thompson sampling (dTS)... For on-policy learning
sDM → usedfor → off-policy learning
confidence 95% · The second part addresses off-policy learning. We propose sDM
meTS → usedfor → on-policy learning
confidence 95% · The first part develops structured Bayesian methods for on-policy learning. We introduce meTS
sDM → achievesconvergence → O(1/sqrt(n))
confidence 90% · sDM achieves O(1/sqrt(n)) convergence in Bayesian suboptimality
policy-weighted log-likelihood → addresses → optimization intractability
confidence 90% · advocate for policy-weighted log-likelihood objectives that prioritize optimization tractability
meTS → reduces → Bayesian Regret
confidence 90% · This reduces Bayesian regret to O(sqrt(TdK_eff))
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces, both settings face major challenges: inefficient exploration, sparse data coverage, high-variance importance weights, extrapolation bias, and difficult optimization landscapes. The first part develops structured Bayesian methods for on-policy learning. We introduce meTS, a mixed-effect extension of Thompson sampling, and dTS, which leverages diffusion-inspired priors to model dependencies between actions. These methods share information across actions and yield regret guarantees depending on an effective number of actions. The second part addresses off-policy learning. We propose sDM, a structured direct method based on latent variables, show that optimization error can dominate estimation error in large action spaces, and introduce concave, efficiently optimizable policy-weighted log-likelihood objectives. Finally, we develop differentiable pessimistic methods based on exponential smoothing and PAC-Bayesian bounds to control the bias-variance trade-off of regularized importance-sampling estimators.
Tags
Links
- Source: https://arxiv.org/abs/2607.28408v1
- Canonical: https://arxiv.org/abs/2607.28408v1
Trouble viewing inline? Open PDF directly →
Full Text
536,353 characters extracted from source content.
Expand or collapse full text
574 NNT : 2026IPPAG003 On-Policy and Off-Policy Learning for Large Action Spaces Th ` ese de doctorat de l’Institut Polytechnique de Paris pr ́ epar ́ e ` a l’ ́ Ecole nationale de la statistique et de l’administration ́ economique ́ Ecole doctorale n ◦ 574 ́ Ecole Doctorale de Math ́ ematique Hadamard (EDMH) Sp ́ ecialit ́ e de doctorat : Math ́ ematiques appliqu ́ ees Th ` ese pr ́ esent ́ e et soutenue ` a Palaiseau, le 13 mars 2026, par IMAD AOUALI Composition du Jury : Vianney Perchet Professeur, CREST, ENSAE, IP ParisPr ́ esident Olivier Capp ́ e Directeur de recherche, CNRSRapporteur Aur ́ elien Garivier Professeur, Ecole Normale Superieure de LyonRapporteur Claire Vernade Professeure, University of Technology NurembergExaminatrice Victor-Emmanuel Brunel Professeur, CREST, ENSAE, IP ParisDirecteur de th ` ese Anna Korba Professeure assistante, CREST, ENSAE, IP ParisInvit ́ e David Rohde Chercheur, Criteo AI LabInvit ́ e Acknowledgements First and foremost, I wish to express my sincere and profound gratitude to my PhD supervisors, Victor-Emmanuel Brunel, Anna Korba, and David Rohde. It has been an immense privilege to learn from and work with them over these years. They shaped my research and personal growth in ways that will stay with me far beyond this thesis. Victor brought rigor, kindness, and openness to every discussion, profoundly shaping how I approach research and problem-solving. Anna’s brilliance, drive, and compassionate mentorship were central to the success of this PhD; her rare ability to combine deep technical insight with empathy and encouragement made her guidance invaluable. David was the best manager I could have hoped for, whose trust, patience, and human approach made all the difference when navigating both professional and personal challenges. I am profoundly grateful to the members of my thesis jury for the time, care, and expertise they devoted to evaluating this work. I would first like to express my deepest appreciation to Olivier Cappé and Aurélien Garivier for accepting the demanding role of rapporteurs, and for the considerable time and attention they dedicated to reading this manuscript in depth. I am sincerely grateful for their careful assessment, thoughtful comments, and constructive feedback. I would also like to warmly thank Vianney Perchet for the honor of presiding over the jury, and for his invaluable support. Finally, I am deeply grateful to Claire Vernade for serving as examinatrice, and for her generosity, encouragement, and support. It was a true privilege to have such distinguished researchers on my jury, and I deeply appreciate their scientific perspective, insightful remarks, and kindness. I would also like to thank Criteo, CAIL, and the Performance Science team, as well as CREST, ENSAE, Institut Polytechnique de Paris. Both the company and the laboratory provided outstanding scientific and institutional support throughout this thesis. I am deeply grateful to several mentors: to Branislav Kveton, whose insights as my first major co-author shaped my approach to research and publication; to Florian Strub, my DeepMind scholarship mentor, whose guidance inspired me to pursue a PhD; and to Flavian Vasile, and Michal Valko, for their invaluable support, both seen and unseen. My heartfelt thanks also go to all my co-authors, whose contributions have greatly en- riched this work: A. Gilotte, B. Heymann, N. Nguyen, A. György, P. Alquier, N. Chopin, S. Katariya, AAS. Hammou, S. Ivanov, A. Benhalloum, M. Bompaire, M. Vono, M. Gartrell, V. Zaytsev, D. Legrand, and O. Jeunen. Special recognition goes to O. Sakhi, 2 M. Cherifa and A.B. Yahmed who were exceptional companions throughout this journey. Finally, to my family and friends: my deepest thanks to my mother and siblings for their unwavering love and support. This thesis is dedicated to my late father, whose dream was for me to pursue a PhD. To Basma, Hakim, Achraf, Ayman, Ismail, Tayeb, Anas, Nicolas, Youssef, Kini, Song, Issam, Charif, Yassine, Oussama, Abdellah, Ali, Hicham: thank you for your friendship and presence throughout this journey. 3 Abstract Many interactive systems (e.g., recommender systems) can be modeled as contextual ban- dits. This framework captures the core challenge of decision-making under uncertainty: selecting actions based on context while learning from partial, noisy feedback. Learning in this setting follows two paradigms: on-policy learning, in which agents collect data and update their policy simultaneously in real time, and off-policy learning, in which the agent’s policy is learned offline from static logs collected under a different policy. Stan- dard algorithms for both paradigms struggle to scale to large action spaces, facing either computational intractability or statistical inefficiency. This thesis develops principled and practical methods to make contextual bandit algorithms tractable in large action spaces, advancing both paradigms through novel algorithmic and theoretical contributions. For on-policy learning, we introduce structured Bayesian models that enable efficient exploration via information sharing. Our first contribution, mixed-effect Thompson sam- pling (meTS) (Chapter 3), couples action parameters through shared latent effects. This reduces Bayesian regret to ̃ O( √ TdK eff ), where K eff is an effective number of actions. When the number of shared effects is much smaller than the number of actions (L≪ K), we have K eff ≪ K, yielding significant regret reduction. Moreover, meTS achieves dra- matic improvements in both memory complexity (from O(K 2 d 2 ) to O((L 2 + K)d 2 )) and runtime complexity (from O(K 3 d 3 ) to O((L 3 + K)d 3 )). We extend this framework to diffusion Thompson sampling (dTS) (Chapter 4), which leverages deep generative models to capture complex action distributions. dTS further improves memory complexity to O((L + K)d 2 ) and runtime complexity to O((L + K)d 3 ), where L denotes the number of layers in the diffusion model. Both methods perform well empirically as analyzed without additional hyperparameter tuning, making them highly practical. For off-policy learning, we address fundamental bottlenecks through three complementary approaches. First, in Chapter 6, we develop the structured direct method (sDM), which models action parameters using a shared latent structure. We prove that sDM achieves O(1/ √ n) convergence in Bayesian suboptimality without requiring the restrictive full log- ging support assumption. sDM performs well in practice, and the performance gap between sDM and standard direct methods widens as the action space grows. Second, in Chapter 7, we challenge the conventional wisdom that better reward estimation yields better poli- cies. We demonstrate that optimization intractability, rather than estimation accuracy, becomes the primary bottleneck in large action spaces, and advocate for policy-weighted 4 log-likelihood objectives that prioritize optimization tractability; these consistently out- perform sophisticated estimators on datasets with up to one million actions. Third, in Chapter 8, we address importance sampling variance by combining variance-reducing reg- ularization with principled pessimism. Our exponential smoothing estimators and unified PAC-Bayesian analysis yield tractable learning objectives amenable to stochastic opti- mization, providing concentration bounds and superior empirical performance. We validate these theoretical and algorithmic advances through extensive experiments on synthetic and real-world datasets. By developing scalable algorithms for both learning paradigms, this thesis enables the deployment of contextual bandits in modern applica- tions where action spaces routinely exceed thousands or millions of actions. 5 Contents Résumé substantiel en français12 1 Overview18 1.1 Context and Scope . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 1.2 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 1.3 Contributions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 1.4 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 I On-Policy Learning in Large Action Spaces38 2 Introduction to Part I39 2.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 2.2 Hierarchical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 2.3 Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 3 Scaling Thompson Sampling with Mixed Effects42 3.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 3.2 Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 3.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 3.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 4 Scaling Thompson Sampling with Diffusion Models57 4.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58 4.2 Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59 4.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 63 4.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 66 4.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 69 I Off-Policy Learning in Large Action Spaces70 5 Introduction to Part I71 5.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 5.2 Methodological Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . 72 5.3 Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 6 6 Scaling Direct Methods with Latent Parameters75 6.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.2 Structured DM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.3 Linear-Gaussian Case . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 6.4 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 6.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 6.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 7 Optimization Matters More than Esimation86 7.1 Analysis of IPS-Based Objectives . . . . . . . . . . . . . . . . . . . . . . . 87 7.2 Analysis of PWLL objectives . . . . . . . . . . . . . . . . . . . . . . . . . . 92 7.3 Empirical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 7.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 8 Principled Pessimism for Exponential Smoothing and Beyond99 8.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 100 8.2 Exponential Smoothing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 102 8.3 PAC-Bayes Analysis for Off-Policy Learning . . . . . . . . . . . . . . . . . 104 8.4 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 109 8.5 Experiments for Exponential Smoothing . . . . . . . . . . . . . . . . . . . 112 8.6 Extension to Other Regularizations . . . . . . . . . . . . . . . . . . . . . . 115 8.7 Experiments for Other Regularizations . . . . . . . . . . . . . . . . . . . . 118 8.8 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 121 9 Conclusions and Future Work123 A Supplementary Materials for Chapter 3125 A.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2 Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.3 Regret Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 A.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 139 B Supplementary Materials for Chapter 4143 B.1 Posterior for Linear Diffusion Models . . . . . . . . . . . . . . . . . . . . . 143 B.2 Posterior for Non-Linear Diffusion Models . . . . . . . . . . . . . . . . . . 145 B.3 Connection to Two-Level Hierarchies . . . . . . . . . . . . . . . . . . . . . 146 B.4 Formal Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 B.5 Regret proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 B.6 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 C Supplementary Materials for Chapter 6161 C.1 Posterior Derivations Under Standard Priors . . . . . . . . . . . . . . . . . 161 C.2 Posterior Derivations Under Structured Priors . . . . . . . . . . . . . . . . 162 C.3 Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 C.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 174 D Supplementary Materials for Chapter 7178 D.1 Proofs for Oracle Policies . . . . . . . . . . . . . . . . . . . . . . . . . . . . 178 7 D.2 Proofs for Optimization Properties . . . . . . . . . . . . . . . . . . . . . . 182 D.3 Stochastic Optimization Convergence Guarantees for PWLL . . . . . . . . 188 D.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . 191 E Supplementary Materials for Chapter 8202 E.1 Bias and Variance Trade-Off . . . . . . . . . . . . . . . . . . . . . . . . . . 203 E.2 Proofs for Off-Policy Learning . . . . . . . . . . . . . . . . . . . . . . . . . 204 E.3 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 217 8 Notation Notation General Mathematical Notation SymbolDefinition [n]Set of first n positive integers: 1, 2,...,n R d d-dimensional real vector space I d Identity matrix of dimension d× d ∆(A)Probability simplex over action set A O(·)Big-O notation for upper bounds ⊗Kronecker product Probability conventions. Random variables are denoted with capital letters, and their realizations with the respective lowercase letters, except for Greek letters. With a slight abuse of notation, for random variables X,Y , the distribution (or density) of X | Y = y evaluated at x is denoted by p(x| y). Norms and Inner Products SymbolDefinition ∥·∥Euclidean norm (unless specified otherwise) ∥a∥ Σ Weighted norm: √ a ⊤ Σa for a∈ R d , Σ≻ 0 Contextual Bandit Framework SymbolDefinition X ⊂ R d Context space (d-dimensional) A = [K]Finite action set with K actions TNumber of interaction rounds nNumber of samples in logged dataset 9 Policies and Value Functions SymbolDefinition π :X → ∆(A)Stochastic policy mapping contexts to action distributions π(·| x)Probability distribution over actions given context x π t Policy at round t (on-policy setting) π 0 Logging policy (off-policy setting) π ∗ Optimal policy (off-policy setting) ˆπLearned policy (off-policy setting) V (π)Value (expected reward) of policy π ˆ V (π)Estimated value of policy π Environment Components SymbolDefinition νContext distribution over X p(·| x,a)Conditional reward distribution given context x and action a r(x,a)Expected reward function: E R∼p(·|x,a) [R] ˆr(x,a)Estimated reward function r(x,a;θ)Parametric reward model r(x,a; ˆ θ)Estimated reward function in parametric case Random Variables and Data SymbolDefinition X t Context observed at round t A t Action taken at round t R t Reward received at round t A t,∗ Optimal action at round t: arg max a∈A r(X t ,a) H t History of interactions: (X ℓ ,A ℓ ,R ℓ ) ℓ<t D n Logged dataset: (X i ,A i ,R i ) n i=1 Performance Metrics SymbolDefinition R(T ) = ∑︁ T t=1 r(X t ,A t,∗ )− r(X t ,A t ) Cumulative regret BR(T ) = E[R(T )]Bayesian cumulative regret so(ˆπ) = V (π ∗ )− V (ˆπ)Suboptimality gap of policy ˆπ Bso(ˆπ) = E[so(ˆπ)]Bayesian suboptimality gap of policy ˆπ 10 Matrix Operations and Concatenation SymbolDefinition [a 1 ,a 2 ,...,a n ]Horizontal concatenation of vectors into d× n matrix (a i ) i∈[n] Vertical concatenation: (a ⊤ 1 ,...,a ⊤ n ) ⊤ ∈ R nd Vec(·)Vectorization operator diag((A i ) i∈[n] )Block diagonal matrix with blocks A 1 ,..., A n (A i ) i∈[n] Vertical concatenation of matrices into nd× d matrix (A i,j ) (i,j)∈[n]×[m] Block matrix where A i,j is the (i,j)-th block Eigenvalues SymbolDefinition λ 1 (A)Maximum eigenvalue of matrix A λ d (A)Minimum eigenvalue of matrix A 11 Résumé substantiel en français Cette thèse étudie l’apprentissage séquentiel et contrefactuel dans des systèmes interact- ifs où l’espace des décisions est très grand. Ces systèmes sont aujourd’hui omniprésents : moteurs de recommandation, publicité computationnelle, places de marché, systèmes de tarification, robotique ou encore allocation de ressources. Leur fonctionnement repose sur une boucle d’interaction simple mais difficile à optimiser : à chaque étape, le sys- tème observe un contexte, choisit une action parmi un très grand nombre de possibilités, puis reçoit une récompense partielle et bruitée qui dépend conjointement du contexte et de l’action choisie. Dans un système de recommandation, le contexte peut représenter l’historique et les préférences d’un utilisateur, l’action correspond à l’item recommandé dans un catalogue contenant potentiellement des millions d’items, et la récompense mesure l’engagement de l’utilisateur, par exemple un clic ou un temps de visionnage. En publicité en ligne, l’action peut combiner le choix d’une annonce et d’un prix d’enchère, tandis que la récompense dépend d’événements successifs tels que le gain de l’enchère, le clic ou la conversion. Le cadre mathématique central de cette thèse est celui des bandits contextuels. On considère un espace de contextes X ⊂ R d , un ensemble fini d’actions A = [K], une distribution inconnue de contextes ν, et des distributions conditionnelles de récompenses p(·| x,a). La fonction de récompense moyenne est définie par r(x,a) = E[R| X = x,A = a]. À chaque tour t, un contexte X t ∼ ν est observé, l’agent choisit une action A t selon une politique π t (·| X t ), puis reçoit une récompense R t ∼ p(·| X t ,A t ). Ce formalisme capture l’essentiel de l’apprentissage interactif : l’agent ne voit que la récompense de l’action qu’il a effectivement choisie, et doit donc apprendre à partir d’un retour partiel. La difficulté principale analysée dans cette thèse est le passage à l’échelle lorsque K est très grand. La thèse traite deux paradigmes complémentaires. Le premier est l’apprentissage en ligne, ou on-policy learning, dans lequel l’agent interagit séquentiellement avec l’environnement et met à jour sa politique au fil des observations. Sa performance est mesurée par le regret cumulé, R(T ) = T ∑︂ t=1 (r(X t ,A t,∗ )− r(X t ,A t )), 12 où A t,∗ désigne l’action optimale dans le contexte X t . L’enjeu fondamental est le compro- mis exploration-exploitation : l’agent doit explorer les actions incertaines pour apprendre, tout en exploitant les actions déjà estimées comme performantes afin de limiter la perte de récompense. Le second paradigme est l’apprentissage hors politique, ou off-policy learn- ing, dans lequel l’agent ne peut plus interagir avec l’environnement et doit apprendre à partir d’un jeu de données journalisé D n =(X i ,A i ,R i ) n i=1 collecté par une politique de logging π 0 . L’objectif est alors d’apprendre une politique ˆπ de grande valeur V (π) = E X∼ν E A∼π(·|X) [r(X,A)], et la performance est mesurée par l’écart de sous-optimalité V (π ∗ )−V (ˆπ). Dans ce second cadre, l’apprentissage exige un raisonnement contrefactuel : il faut estimer ce qui se serait passé si une autre action avait été choisie, alors même que les données ne contiennent que les actions effectivement sélectionnées par π 0 . Dans les deux paradigmes, la grande taille de l’espace d’actions amplifie les difficultés statistiques, computationnelles et d’optimisation. En ligne, explorer indépendamment des milliers ou millions d’actions devient prohibitif : chaque action reçoit peu d’observations, ce qui ralentit considérablement l’apprentissage et augmente le regret. Hors politique, la couverture des données se dégrade lorsque K croît : de nombreuses actions sont rarement ou jamais observées dans certaines régions de l’espace des contextes. Les méthodes fondées sur un modèle de récompense souffrent alors d’un fort biais d’extrapolation, tandis que les méthodes par pondération inverse des propensions peuvent avoir une variance très élevée lorsque π 0 (a | x) est faible. Cette thèse montre également qu’un autre obstacle, souvent sous-estimé, devient dominant dans les grands espaces d’actions : l’optimisation des ob- jectifs hors politique standards peut devenir intrinsèquement difficile, indépendamment de la qualité statistique de l’estimateur de valeur. La première partie de la thèse est consacrée à l’apprentissage en ligne dans les grands espaces d’actions. Les méthodes classiques telles que Upper Confidence Bound et Thomp- son Sampling reposent souvent sur des modèles disjoints, où chaque action a possède son propre paramètre θ a et où la récompense est modélisée sous la forme r(x,a) = φ(x) ⊤ θ a . Ces modèles sont attractifs en pratique, notamment dans les systèmes de recommanda- tion, car ils évitent de construire manuellement des caractéristiques conjointes contexte- action. Cependant, leur faiblesse est statistique : apprendre séparément un paramètre pour chaque action nécessite beaucoup de données par action, ce qui est incompatible avec de très grands catalogues. La contribution principale de cette partie consiste à conserver la flexibilité des modèles disjoints tout en introduisant des structures bayésiennes capables de partager l’information entre actions. Le premier algorithme proposé est mixed-effect Thompson Sampling, ou meTS. Il repose sur un modèle bayésien hiérarchique où les paramètres d’actions θ a sont couplés par des effets latents partagés Ψ = (ψ ℓ ) ℓ∈[L] . Ces effets peuvent représenter, par exemple, des catégories ou des facteurs communs entre items. Le modèle suppose que les paramètres d’actions sont conditionnellement indépendants sachant les effets latents, mais qu’ils partagent de l’information à travers ces effets. À chaque tour, meTS échantillonne d’abord les effets 13 latents depuis leur postérieur, puis échantillonne les paramètres d’actions conditionnelle- ment à ces effets, avant de choisir l’action maximisant la récompense échantillonnée. Cette procédure conserve le principe de Thompson Sampling tout en rendant l’exploration statis- tiquement plus efficace. Dans le cas linéaire-gaussien, la thèse dérive des mises à jour exactes en forme fermée, et propose des approximations de Laplace tractables pour les modèles linéaires généralisés. L’analyse théorique établit une borne de regret bayésien de l’ordre ˜︁ O (︃ √︂ TdK eff (σ 2 0 + σ 2 Ψ ) )︃ , où K eff est un nombre effectif d’actions. Lorsque la structure latente est informative et que L ≪ K, on a K eff ≪ K, ce qui conduit à une amélioration multiplicative par rapport à Thompson Sampling standard. Sur le plan computationnel, la factorisation conditionnelle réduit fortement les coûts mémoire et temps par rapport à une modélisation bayésienne dense de toutes les actions. Les expériences montrent que les gains de meTS augmentent avec la taille de l’espace d’actions, confirmant que le partage d’information est essentiel pour l’exploration à grande échelle. La seconde contribution en ligne est diffusion Thompson Sampling, ou dTS. Cette méth- ode généralise meTS en remplaçant la hiérarchie à un niveau par une hiérarchie profonde inspirée des modèles de diffusion. Les paramètres d’actions sont générés au terme d’une chaîne de variables latentes reliées par des transformations non linéaires pré-entraînées. Cette structure permet de représenter des dépendances complexes entre actions, bien au- delà des effets linéaires ou catégoriels. Le défi technique est que le postérieur exact devient intraitable à cause des non-linéarités des fonctions de lien et du modèle de récompense. La thèse introduit donc une procédure d’inférence en ligne fondée sur des mises à jour de type gaussien, où les précisions postérieures combinent précision a priori et précision issue des données, et où les moyennes sont obtenues par combinaison pondérée entre les prédictions du prior et les estimations de maximum de vraisemblance. L’intérêt de dTS est double. D’une part, l’algorithme exploite des priors riches appris hors ligne, par exemple à partir de représentations d’items, pour accélérer l’exploration en ligne. D’autre part, il conserve une structure de diffusion dans le postérieur, plutôt que de l’approximer par une simple gaussienne globale. Dans le cadre linéaire-gaussien, la thèse établit une borne de regret bayésien de l’ordre ˜︁ O ⎛ ⎝ ⌜ ⃓ ⃓ ⎷ TdK eff L+1 ∑︂ ℓ=1 σ 2 ℓ ⎞ ⎠ , et montre que la complexité peut être rendue linéaire en L + K. Empiriquement, dTS améliore les performances des méthodes de référence, y compris lorsque les priors de diffusion sont imparfaits ou appris à partir de données limitées. Cette contribution montre que les modèles génératifs profonds peuvent être utilisés non seulement pour représenter des actions, mais aussi pour structurer l’incertitude nécessaire à l’exploration. La seconde partie de la thèse porte sur l’apprentissage hors politique dans les grands es- paces d’actions. Les méthodes classiques se divisent principalement en méthodes directes, 14 qui apprennent un modèle de récompense ˆr(x,a), et méthodes par importance sampling, qui estiment directement la valeur d’une politique au moyen du ratio π(a | x)/π 0 (a | x). Les méthodes directes sont sensibles au biais de modèle et à la rareté des observations par action. Les méthodes IPS sont non biaisées sous des hypothèses de support appro- priées, mais leur variance peut exploser lorsque la politique cible attribue de la masse à des actions peu probables sous la politique de logging. En outre, la thèse montre que les objectifs IPS induisent souvent des paysages non concaves, plats et riches en maxima locaux lorsque l’espace d’actions est grand. La première contribution hors politique est la structured Direct Method, ou sDM. Elle transpose au cadre offline l’idée de structuration bayésienne introduite dans meTS. Au lieu d’estimer indépendamment un paramètre par action, sDM couple les paramètres θ a à travers un vecteur latent partagé ψ. Après observation du jeu de données journalisé, l’algorithme calcule le postérieur de ψ, puis les postérieurs conditionnels des paramètres d’actions. La récompense estimée pour chaque action est obtenue en intégrant l’incertitude postérieure, et la politique apprise agit ensuite gloutonnement par rapport à cette récom- pense moyenne postérieure. L’analyse introduit une notion de sous-optimalité bayésienne adaptée au cadre hors poli- tique. Elle montre que sDM atteint une convergence enO(1/ √ n) sans imposer l’hypothèse de support uniforme complet π 0 (a| x)≥ γ > 0 pour toutes les actions. La borne dépend plutôt de l’alignement entre la politique de logging et la politique optimale : plus les actions optimales sont couvertes par les données, plus l’apprentissage est efficace. Cette analyse révèle également un phénomène important : sous le critère bayésien considéré, les politiques gloutonnes sont optimales et peuvent surpasser les politiques pessimistes, con- trairement au cadre fréquentiste où le pessimisme est souvent nécessaire pour se protéger contre les pires cas. Les expériences confirment que sDM améliore les méthodes directes standards, avec des gains croissants lorsque K augmente. La contribution suivante remet en question une hypothèse centrale de l’apprentissage hors politique : l’idée selon laquelle l’amélioration des estimateurs de valeur suffit à améliorer l’apprentissage de politiques. La thèse montre que, dans les grands espaces d’actions, l’erreur d’optimisation peut dominer l’erreur d’estimation. Même un estimateur statis- tiquement sophistiqué peut conduire à une mauvaise politique si l’objectif qu’il induit est difficile à optimiser. L’analyse des paysages d’optimisation montre que les objectifs fondés sur des estimateurs peuvent présenter des plateaux où les méthodes de gradient restent bloquées pendant O(K) itérations, ainsi qu’un nombre exponentiel de maxima locaux en fonction de K. Pour comprendre ces échecs, la thèse analyse les politiques oracle associées à différents estimateurs, c’est-à-dire les politiques qui maximiseraient ces estimateurs avec une quan- tité infinie de données. Cette analyse met en évidence que chaque estimateur impose un biais inductif spécifique : IPS favorise les politiques proches du support de logging, tandis que des méthodes clusterisées opèrent au niveau de groupes d’actions. Ces observations motivent des paramétrisations de politiques adaptées à l’objectif, qui réduisent l’espace de recherche effectif de K à une taille beaucoup plus petite, comme la taille du support de logging ou le nombre de clusters. La conclusion principale de cette partie est toutefois plus radicale : il peut être préférable 15 d’abandonner l’estimation explicite de valeur pour optimiser directement des objectifs de vraisemblance pondérée par la politique. La thèse introduit les objectifs policy-weighted log-likelihood, de la forme ˆ U g (π) = 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i )) logπ(A i | X i ), où g est une fonction de pondération positive. Ces objectifs ne sont pas des estimateurs de valeur, mais ils possèdent des paysages d’optimisation bien plus favorables : pour des politiques softmax linéaires, ils sont concaves, et deviennent fortement concaves avec régu- larisation ℓ 2 . Ils admettent donc un optimum global unique accessible par optimisation stochastique standard. Les expériences à très grande échelle, incluant des espaces allant jusqu’à un million d’actions, montrent que ces objectifs simples et stables surpassent des méthodes hors politique fondées sur des estimateurs de valeur plus complexes. Cette con- tribution établit que, pour l’apprentissage de politiques dans les grands espaces d’actions, l’optimisabilité de l’objectif est un critère aussi fondamental que sa précision statistique. La dernière contribution principale de la thèse améliore les méthodes IPS régularisées en introduisant un pessimisme praticable et différentiable. Les ratios d’importance peuvent être très grands, ce qui augmente fortement la variance. Pour y remédier, la thèse étudie des estimateurs par exponential smoothing, notamment ˆ V α (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) α R i , et ̃ V β (π) = 1 n n ∑︂ i=1 (︃ π(A i | X i ) π 0 (A i | X i ) )︃ β R i . Ces estimateurs interpolent entre absence de régularisation et forte réduction de variance, tout en restant différentiables et donc compatibles avec l’optimisation par gradient. Pour apprendre de manière sûre avec ces estimateurs biaisés mais moins variables, la thèse dérive une borne PAC-bayésienne bilatérale contrôlant l’écart entre la valeur vraie et l’estimateur régularisé. Cette borne décompose l’erreur en plusieurs termes interprétables : une divergence entre la politique apprise et la politique de logging, un biais dû à la régularisation des poids, et une variance résiduelle. Elle conduit à un objectif pessimiste qui maximise une borne inférieure empirique de la valeur : ˆ V α n (π θ )− pénalités de divergence− biais de régularisation− variance résiduelle. L’intérêt crucial de cette formulation est que tous les termes sont empiriques et dif- férentiables. Contrairement à de nombreuses approches pessimistes dont les constantes théoriques sont inexploitables en pratique, cet objectif peut être optimisé à grande échelle par ascension de gradient stochastique. La thèse propose également un cadre PAC- bayésien unifié couvrant plusieurs familles de régularisation des poids d’importance, comme le clipping, l’exponential smoothing et l’implicit exploration. Ce cadre permet une com- paraison cohérente de différentes formes de pessimisme et fournit des objectifs pratiques pour l’apprentissage offline sécurisé. 16 Enfin, la thèse présente plusieurs contributions additionnelles liées aux thèmes principaux. Dans le cadre en ligne, les idées de structuration bayésienne sont étendues au problème de best-arm identification à budget fixé. L’algorithme PI-BAI utilise l’information a priori pour répartir efficacement le budget d’exploration dans des bandits structurés. L’analyse fournit des garanties bayésiennes dépendant du prior sur la probabilité d’erreur, et mon- tre que des allocations non adaptatives bien informées peuvent surpasser des stratégies adaptatives classiques. Dans le cadre hors politique, la thèse contribue également au développement du logarithmic smoothing, un estimateur pessimiste de la forme ˆ V λ LS (π) = 1 nλ n ∑︂ i=1 log (︃ 1 + λ π(A i | X i ) π 0 (A i | X i ) R i )︃ . Cet estimateur agit comme une alternative douce et différentiable au clipping, bénéficie de garanties de concentration serrées, et permet d’obtenir des bornes de sous-optimalité plus fines que les approches précédentes. Enfin, cette thèse CIFRE maintient un lien con- stant avec les systèmes de recommandation industriels à grande échelle. Les applications développées autour de la recommandation optimisant la récompense, de l’évaluation of- fline, de la simulation contrefactuelle et des modèles de recommandation par ardoise ont fourni à la fois un terrain expérimental et une source de questions théoriques. Dans son ensemble, cette thèse défend une idée centrale : pour apprendre efficacement dans de grands espaces d’actions, il ne suffit pas d’appliquer directement les algorithmes classiques de bandits contextuels ou d’apprentissage hors politique. Il faut exploiter la structure entre actions, contrôler explicitement l’incertitude et la couverture des données, et concevoir des objectifs dont le paysage d’optimisation reste favorable à grande échelle. Les contributions proposées répondent à ces exigences selon deux axes complémentaires. En apprentissage en ligne, des modèles bayésiens hiérarchiques et diffusionnels permettent de partager l’information entre actions et de réduire le regret. En apprentissage hors politique, des méthodes directes structurées, des objectifs de vraisemblance pondérée et des principes de pessimisme différentiable rendent l’apprentissage statistiquement robuste et computationnellement réalisable. Ces résultats contribuent à rapprocher la théorie des bandits contextuels des contraintes réelles des systèmes interactifs modernes, où les décisions doivent être prises parmi des catalogues massifs, à partir de signaux partiels, bruités et parfois fortement biaisés. 17 Chapter 1 Overview 1.1 Context and Scope Interactive machine learning systems are a cornerstone of modern technology, optimizing decision-making in applications ranging from recommender systems and financial markets to robotics. These systems operate in a sequential loop: they process contextual informa- tion, select an action from a wide range of possibilities, and receive feedback that depends on both the context and the chosen action. A fundamental challenge in designing these systems is the scale of the decision space; in many real-world settings, the number of potential actions can be huge. Consider recommender systems, where streaming platforms or e-commerce sites must se- lect an item to present to a user. The context comprises rich data, such as user preferences and history. The action is the selection of a specific item from a catalog containing thou- sands or millions of options. The system’s objective is to learn a policy that maximizes user engagement (the reward), measured by metrics such as watch time or clicks. Similarly, in computational advertising, a platform selects which ad to display via real-time bidding. The context includes user attributes and page details. The action is composite: selecting an ad from a large inventory and determining a bid price. The feedback (reward) arrives in stages, from winning the auction to subsequent user clicks or conversions. The system must maximize advertiser value while adhering to budget constraints. We model these interactive systems using the contextual bandit framework. This frame- work captures the essential characteristics of interactive learning while maintaining the tractability required for theoretical analysis and practical implementation. 1.1.1 Contextual Bandits Figure 1.1a visualizes the interaction loop. The environment consists of a Context Gen- erator and a Reward Generator, both of which are assumed to be fixed but unknown to the agent. At each round, the environment emits a context. The agent observes this con- 18 Contextual bandit environment Context Context Context Generator Action Agent Reward Reward Generator (a) Single interaction loop. Contextual bandit environment at round (b) Graphical representation in round t∈ [T]. Figure 1.1: Contextual bandit framework. text and selects an action from a finite set 1 . Finally, the environment generates a scalar reward based on the context-action pair. By repeating this loop, the agent accumulates experience to refine its policy. Formally, let X ⊂ R d denote the context space and A = [K] the finite action set. A stochastic policy π : X → ∆(A) maps each context x ∈ X to a probability distribution π(·| x) over actions. The environment is specified by: • A context distribution ν over X; • A family of conditional reward distributions p(·| x,a) (x,a)∈X×A . The expected reward function for any context-action pair (x,a) is defined as: r(x,a) = E R∼p(·|x,a) [R].(1.1) The interaction unfolds over T rounds. In each round t∈ [T ]: 1. The environment draws a context X t ∼ ν and reveals it to the agent. 2. The agent selects an action A t ∼ π t (·| X t ) according to its current policy π t . 3. The environment samples a reward R t ∼ p(·| X t ,A t ) and returns it to the agent. A graphical representation of this interaction in round t∈ [T ] is visualized in Figure 1.1b. 1.1.2 Learning Paradigms We address learning in this framework via two complementary paradigms: on-policy (on- line) and off-policy (offline). With a slight abuse of terminology, we use learning loosely to encompass the full agent behavior, including both reward estimation and action selection. On-Policy (Online) Learning In this setting, the agent updates its policy π t sequentially. Let H t =(X ℓ ,A ℓ ,R ℓ ) ℓ<t denote the history available at the start of round t. The agent uses H t to construct 1 The action space can technically be infinite, but this thesis focuses on large finite action spaces. 19 the policy π t . Following the interaction, the agent augments the history with the new observation (X t ,A t ,R t ) to form H t+1 and repeats the process. Performance is measured by the cumulative regret: R(T ) = T ∑︂ t=1 (r(X t ,A t,∗ )− r(X t ,A t )),(1.2) where A t,∗ = arg max a∈A r(X t ,a) is the optimal action in round t. While minimizing regret is equivalent to maximizing cumulative reward, the literature prioritizes re- gret as it normalizes performance against the optimal oracle, facilitating theoretical comparisons across environments. The agent faces the exploration-exploitation dilemma: it must balance exploration of poorly understood actions with exploitation of actions believed to yield high rewards. Feedback is partial (only the reward for the chosen action is observed) and noisy, and exploration may be constrained by safety or budget requirements. Remark 1 (Beyond regret minimization). While this thesis focuses on regret min- imization, other objectives exist, most notably Best-Arm Identification (BAI). BAI aims to identify the optimal action rather than maximize cumulative reward. This is important for applications like A/B testing and clinical trials. Although we focus on regret, our core modeling contributions for scaling Thompson sampling extend to the BAI setting, as demonstrated in our related work (Nguyen et al., 2025) (not included in this manuscript). Off-Policy (Offline) Learning In this setting, the agent learns a policy ˆπ from a static logged dataset D n = (X i ,A i ,R i ) n i=1 collected by a logging policy π 0 as X i ∼ ν, A i ∼ π 0 (· | X i ) and R i ∼ p(· | X i ,A i ). No additional interactions with the environment are allowed. The objective is to find a policy maximizing the expected value: V (π) = E X∼ν E A∼π(·|X) [r(X,A)].(1.3) Performance is measured by the suboptimality gap of the learned policy ˆπ: so(ˆπ) = V (π ∗ )− V (ˆπ),(1.4) where π ∗ = argmax π∈Π V (π) is the unknown optimal policy in a class of policies Π. Since the agent learns solely from data generated by π 0 , it must perform counterfac- tual reasoning. This introduces several challenges. First, support mismatch occurs when specific actions are rarely or never selected by π 0 in certain regions of the context space; consequently, the dataset provides little to no information about the rewards in these regions. Second, re-weighting instability arises since standard tech- niques, such as inverse propensity scoring (importance sampling), re-weight logged samples using the density ratio π(a| x)/π 0 (a| x). When π 0 (a| x) is small, this ratio 20 can explode, leading to high-variance estimates. Third, high bias can affect methods that rely on parametric models to estimate the reward. Extrapolating rewards for unobserved context-action pairs introduces errors if the model is misspecified. The difficulties in both paradigms are amplified when the action space is large (K in the thousands or millions). In on-policy settings, independent exploration of every action becomes infeasible, yielding high regret. In off-policy settings, data sparsity worsens: reward models risk significant extrapolation error, and importance weights suffer from extreme variance as the probability of observing any specific action vanishes. Developing methods that scale gracefully (both statistically and computationally) with the size of the action space is the central theme of this thesis. 1.2 Background Most on-policy and off-policy learning algorithms 2 are fundamentally constructed from two core components: reward estimation, which approximates the expected reward func- tion r(x,a) for any context-action pair from data, and decision-making, which leverages these estimates and their associated uncertainty to select actions. 1.2.1 Reward Estimation The central task in reward estimation is to learn an approximate function ˆr(x,a) that esti- mates the true expected reward r(x,a) defined in Equation (1.1). This function is trained on a dataset, denoted by Data, whose structure depends on the learning paradigm: On-policy data. The agent collects data sequentially. The dataset at round t is the history of interaction up to round t: Data = H t =(X i ,A i ,R i ) t−1 i=1 .(1.5) Off-policy data. The agent learns from a static dataset logged by a logging policy π 0 : Data =D n =(X i ,A i ,R i ) n i=1 , with A i ∼ π 0 (·| X i ).(1.6) We denote the learned model as ˆr(x,a) = r(x,a; ˆ θ), where ˆ θ are parameters obtained via a statistical objective. Common approaches include: Maximum Likelihood Estimation (MLE) MLE seeks parameters θ that maximize the probability of observing the collected rewards. The objective is: ˆ θ MLE = arg max θ ∑︂ (X i ,A i ,R i )∈Data logp(R i | X i ,A i ;θ). For Gaussian rewards p(· | x,a;θ) = N (r(x,a;θ),σ 2 ), this simplifies to minimiz- ing the sum of squared errors (ordinary least squares). For Bernoulli rewards, it corresponds to minimizing binary cross-entropy (e.g., logistic regression). 2 Recall that we use learning loosely to encompass both reward estimation and action selection (Sec- tion 1.1.2). 21 Maximum A Posteriori (MAP) MAP estimation incorporates a prior distribution p 0 (θ) to regularize the objective: ˆ θ MAP = arg max θ ⎛ ⎝ ∑︂ (X i ,A i ,R i )∈Data logp(R i | X i ,A i ;θ) + logp 0 (θ) ⎞ ⎠ . For example, combining a Gaussian likelihood with a zero-mean Gaussian prior p 0 (·) =N (0,λI d ) is equivalent to ridge regression (ℓ 2 regularization). Full Bayesian Inference Rather than a point estimate, Bayesian inference characterizes uncertainty by com- puting the full posterior distribution p(θ | Data): p(θ | Data)∝ p(θ) ∏︂ (X i ,A i ,R i )∈Data p(R i | X i ,A i ;θ). This approach is powerful for guiding exploration. A common implementation as- sumes conjugate Gaussian distributions: p 0 (·) = N (μ 0 , Σ 0 ) and p(· | x,a;θ) = N (r(x,a;θ),σ 2 ). When r(x,a;θ) is linear in θ (Bayesian linear regression), the posterior is Gaussian and analytically tractable. For non-linear models, posterior approximation methods are required. The functional form of the reward model, r(x,a;θ), is an important design choice that dictates the balance between computational tractability, data efficiency, and expressive power. While numerous function classes exist, linear models remain a cornerstone of the field due to their simplicity and strong theoretical guarantees. Even within this linear family, there is an important distinction between joint and disjoint formulations. The joint linear model defines r(x,a;θ) = φ(x,a) ⊤ θ, sharing a single parameter θ ∈ R d across all actions. While data-efficient, this approach relies on designing a feature map φ(x,a) capable of capturing complex context-action interactions. Designing φ(x,a) is very hard in practice, and inadequate feature engineering leads to poor performance. This is why in practice, it is often more common to adopt a disjoint linear model that learns an independent parameter θ a ∈ R d for each action, yielding r(x,a;θ) = φ(x) ⊤ θ a . This is common in recommender systems, where predictions are inner products of user and item embeddings. This formulation is robust and avoids complex feature engineering. However, its primary drawback is poor statistical scalability: while the computational overhead of maintaining K independent embeddings can be addressed with sufficient compute, the statistical challenge remains. This is because learning each embedding independently requires substantial data per action, which becomes prohibitive as K grows. Other non-linear models exist for capturing complex reward functions. Generalized linear models (GLMs), for instance, employ a link function to model non-Gaussian rewards (e.g., logistic regression for binary outcomes: p(R = 1| x,a) = σ(φ(x) ⊤ θ a )). These models align closely with the linear setting, and the algorithms proposed in this thesis are suitable for them. A significant portion of this thesis focuses on scaling these disjoint reward models to large action spaces. 22 Remark 2 (Scope of parametric reward modeling). The reward estimation framework presented above assumes parametric reward models r(x,a;θ). This assumption underlies Chapters 3, 4 and 6, where we develop structured parametric models that share information across actions to improve statistical efficiency. The remaining two chapters (Chapters 7 and 8) take a slightly different approach: rather than explicitly estimating rewards, they optimize policies directly using inverse propensity scoring and policy-weighted objectives, thereby making them agnostic to the choice of reward model. 1.2.2 Decision-Making We now turn to the second component: decision-making. This stage (often) relies on the reward model ˆr derived using the estimation techniques discussed previously. While the difference in reward estimation between on-policy and off-policy settings is primarily driven by how data is accumulated, the principles guiding action selection in these two paradigms are fundamentally distinct. Decision-Making in On-Policy Learning Recall that in the on-policy setting, the agent must balance exploration and exploitation to minimize regret. The two dominant paradigms are Upper Confidence Bound (UCB) and Thompson Sampling (TS). Upper Confidence Bound (UCB) UCB drives exploration via a bonus term added to the reward estimate, selecting: A t = arg max a∈A (ˆr(X t ,a) + bonus t (X t ,a)). For linear models (e.g., LinUCB), the bonus scales with √︂ φ(x) ⊤ V −1 t,a φ(x), where V t,a is the design matrix for action a, encouraging the selection of less-certain actions. Thompson Sampling (TS) TS (or posterior sampling) implements randomized exploration. The agent samples a reward function ̃r t from the posterior and acts greedily with respect to it: A t = arg max a∈A ̃r t (X t ,a), where ̃r t (x,a) = r(x,a;θ t ), θ t ∼ p(θ | H t ). Exploration is implicit: high posterior uncertainty yields diverse samples θ t , lead- ing to varied actions. As data accumulates, the posterior contracts, and behavior naturally becomes exploitative 3 . In large action spaces, standard UCB and TS struggle. Treating actions independently gathers information too slowly, leading to prohibitive regret. This necessitates structured models that share information across actions. 3 While we describe sampling parameters θ t , the general principle involves sampling from the posterior predictive distribution of rewards. 23 Decision-Making in Off-Policy Learning Recall that the off-policy setting requires identifying an optimal policy from a static dataset collected under a distinct logging policy. The two dominant philosophies for tackling this are Greedy Policies and Pessimistic Policies. Greedy Policies Greedy methods select the policy ˆπ g that maximizes a point estimate of value, ˆ V (π): ˆπ g = arg max π∈Π ˆ V (π).(1.7) We distinguish between two primary classes of estimators. The direct method (DM) relies on the reward model ˆr derived in the previous section. In contrast, inverse propensity scoring (IPS) bypasses reward modeling to estimate the policy value directly using importance weighting 4 : ˆ V dm (π) = 1 n n ∑︂ i=1 ∑︂ a∈A π(a| X i ) ˆr(X i ,a), ˆ V ips (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) R i . The optimization procedures for these objectives differ significantly. The policy maximizing ˆ V dm (π) is simply the one that acts greedily with respect to the learned reward model, ˆπ dm g (x) = arg max a∈A ˆr(x,a).(1.8) For IPS, the optimization is typically performed numerically over a class of param- eterized policies π θ : ˆπ ips g = arg max θ∈R d ˆ V ips (π θ ).(1.9) Greedy policies are effective when the underlying estimator is accurate. Specifically, when the reward model ˆr is well-specified (for DM), or the importance weights have low variance (for IPS). However, the arg max operator can amplify estimation errors, leading to a policy that over-exploits optimistic model inaccuracies or high-variance weight estimates. Pessimistic Policies Pessimistic methods mitigate error amplification by penalizing the objective with a quantified uncertainty term: ˆπ p = arg max π∈Π [︂ ˆ V (π)− pen(π) ]︂ .(1.10) 4 Importance weighting allows IPS to be unbiased under the common support assumption (i.e., π 0 (a|x) = 0 =⇒ π(a|x) = 0) 24 This principle applies to both DM and IPS as: Pess-DM:ˆπ dm p (x) = arg max a∈A [︂ ˆr(x,a)− β ˆσ r (x,a) ]︂ ,(1.11) and Pess-IPS:ˆπ ips p = arg max θ [︂ ˆ V ips (π θ )− β ˆσ ips (π θ ) ]︂ .(1.12) Here, ˆσ r captures the uncertainty of the reward model, while ˆσ ips captures the uncertainty of the IPS estimator itself (e.g., its variance). The penalty term pen(π) prevents the maximization operator from selecting overestimated policies. It also regulates distributional shift by penalizing policies that place probability mass on context-action pairs with low coverage under π 0 (support mismatch). Standard methods, whether relying on DM or IPS, and whether adopting greedy or pes- simistic policies, face severe limitations in large action spaces. For DM, the prevailing practice of modeling action parameters independently prevents information sharing, mak- ing learning statistically inefficient. For IPS, the well-recognized issue is variance: impor- tance weights can be large. However, we also demonstrate in this thesis that optimization can be an even greater bottleneck in large action spaces. This is because standard IPS- based objectives (whether Equation (1.9) or Equation (1.12)) induce highly non-concave landscapes with flat plateaus that trap gradient-based optimizers. Finally, for pessimistic methods, existing formulations often rely on intractable bounds that are incompatible with modern stochastic optimization techniques. This thesis addresses these specific patholo- gies: we introduce structured models for DM to enforce information sharing; we propose new policy-weighted log-likelihood objectives that yield superior optimization landscapes compared with IPS-based objectives; and we develop variance-reduced, tractable pes- simistic objectives that are amenable to stochastic optimization at scale. 1.3 Contributions Part I: On-Policy Learning in Large Action Spaces In the on-policy setting, exploration strategies that adopt disjoint reward models 5 and learn each action parameter independently gather information slowly, resulting in high regret and failure to converge to an optimal policy within practical time horizons. Part I addresses this challenge by scaling Thompson Sampling to large action spaces while re- taining the disjoint reward model parameterization. Our primary contribution is the introduction of structured Bayesian models with informative priors that share statistical strength across actions, enabling efficient exploration without sacrificing the robustness and flexibility of disjoint reward models. 5 Recall that disjoint models offer greater flexibility and are widely used in industrial settings such as large-scale recommender systems that use a separate embedding for each item, whereas joint reward models require careful feature engineering and are not, to the best of our knowledge, widely deployed in practical recommendation systems. 25 (Chapter 3) Scaling Thompson Sampling with Mixed-Effects We propose a hierarchical Bayesian framework that couples action parameters through L shared latent effect parameters Ψ = (ψ ℓ ) ℓ∈[L] ∈ R dL , where each effect ψ ℓ can rep- resent, for example, a category of items: Ψ∼ q 0 , θ a | Ψ∼ p 0,a (·| Ψ),∀a∈ [K], R t | θ, Ψ,X t ,A t ∼ p(·| X t ;θ A t ),∀t∈ [T ]. Upon this model, we build mixed-effect Thompson sampling (meTS). meTS maintains a posterior over effects q t (Ψ) = p(Ψ| H t ) and K conditional posteriors over actions p t,a (θ a | Ψ) = p(θ a | Ψ,H t ). In round t∈ [T ], parameters are sampled hierarchically: Ψ t ∼ q t (·), θ t,a ∼ p t,a (·| Ψ t ),∀a∈ [K]. Actions are then selected via the standard TS rule: A t = arg max a∈[K] r(X t ;θ t,a ). We derive exact closed-form updates for linear models and tractable Laplace ap- proximations for generalized linear models. Theoretically, we establish a Bayesian regret bound in the linear-Gaussian case: BR(T ) = ̃ O (︃ √︂ TdK eff (σ 2 0 + σ 2 Ψ ) )︃ , where σ 2 0 and σ 2 Ψ are the prior variances of the action and effect parameters, re- spectively, and K eff is the effective number of actions. When L ≪ K, we have K eff ≪ K, yielding a multiplicative Bayesian regret improvement of √︁ K/K eff over standard TS. Computationally, meTS exploits the conditional independence of action parameters given the latent effects, reducing memory complexity from O(K 2 d 2 ) to O((L 2 + K)d 2 ) and runtime from O(K 3 d 3 ) to O((L 3 + K)d 3 ). Empirically, meTS consistently outperforms baselines, with gains increasing with K. AISTATS 2023 (Poster) - Aouali et al. (2023b): • I. Aouali, B. Kveton, and S. Katariya. Mixed-effect Thompson sampling. In In- ternational Conference on Artificial Intelligence and Statistics, pages 2087–2115. PMLR, 2023b. (Chapter 4) Scaling Thompson Sampling with Diffusion Models This chapter extends the hierarchical framework of meTS by introducing diffusion Thompson Sampling (dTS), which replaces the single-layer prior with a deep hier- archy of latent variables governed by a diffusion model: ψ L ∼N (0, Σ L+1 ), ψ ℓ−1 | ψ ℓ ∼N (f ℓ (ψ ℓ ), Σ ℓ ),∀ℓ∈ [L]\1, θ a | ψ 1 ∼N (f 1 (ψ 1 ), Σ 1 ),∀a∈ [K], R t | θ, (ψ ℓ ) ℓ∈[L] ,X t ,A t ∼ p(·| X t ;θ A t ),∀t∈ [T ]. 26 The link functions f ℓ are pre-trained non-linear transformations (e.g., neural net- works), enabling rich representations of inter-action structure. As in meTS, ex- ploration proceeds by sampling parameters top-down through the hierarchy and selecting the reward-maximizing action. The key technical challenge is that the exact posterior is intractable due to non- linearities in both the reward and link functions. To enable fast online updates with- out expensive MCMC, we derive a tractable inference procedure based on Gaussian- like updates 6 . The resulting posterior preserves the diffusion structure: it remains a hierarchy of conditional Gaussians, but with fine-tuned link functions and preci- sions: ̄ Σ −1 t,ℓ−1 =Σ −1 ℓ ⏞⏟⏞ prior precision + ̄ G t,ℓ−1 ⏞ ⏟⏞ data precision , ˆ f t,ℓ (ψ ℓ ) = ̄ Σ t,ℓ−1 (︂ Σ −1 ℓ f ℓ (ψ ℓ ) ⏞ ⏟⏞ prior contribution + ̄ B t,ℓ−1 ⏞ ⏟⏞ data contribution )︂ . Here, ̄ G t,ℓ−1 and ̄ B t,ℓ−1 are sufficient statistics propagated upward through the hi- erarchy. As data accumulates, covariances contract and means shift from prior toward MLE. Crucially, this formulation preserves the expressiveness of diffusion models since the posterior is not a single Gaussian but a posterior diffusion model that retains the generative structure of the prior. Theoretically, we analyze dTS in the fully linear-Gaussian setting to gain analytical insight, deriving a Bayes regret bound: BR(T ) = ̃ O ⎛ ⎝ ⌜ ⃓ ⃓ ⎷ TdK eff L+1 ∑︂ ℓ=1 σ 2 ℓ ⎞ ⎠ , where Σ ℓ = σ 2 ℓ I d and K eff ≪ K is the effective number of actions. Computation- ally, dTS exploits hierarchical conditional independence to reduce memory and time complexity further fromO(K 2 d 2 ) andO(K 3 d 3 ) toO((L +K)d 2 ) andO((L +K)d 3 ), respectively: linear scaling in L that improves upon meTS. Empirically, dTS consis- tently outperforms baselines by leveraging pre-trained diffusion priors, even when these priors are imperfect or trained on limited data. NeurIPS 2025 (Poster) - (Aouali, 2025, 2023): • I. Aouali. Diffusion models meet contextual bandits. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. • I. Aouali. Linear diffusion models meet contextual bandits with large action spaces. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. 6 By Gaussian-like updates, we mean that the posterior precision is the sum of prior and evidence precisions, and the posterior mean is the precision-weighted combination of prior mean and maximum likelihood estimate. 27 Part I: Off-Policy Learning in Large Action Spaces In the off-policy setting, both DM and IPS face severe limitations as the action space grows. DM suffers from high model bias and statistical inefficiency due to sparse data coverage. IPS exhibits high variance and potential bias due to insufficient support; moreover, as we demonstrate in this thesis, optimizing IPS-based objectives becomes intractable in large action spaces. Part I addresses these failure modes through three complementary approaches. (Chapter 6) Scaling Direct Methods with Latent Parameters Standard DMs estimate independent d-dimensional parameters for each action, which becomes statistically inefficient when actions are rarely observed. We intro- duce the structured direct method (sDM), which couples action parameters through a shared latent vector ψ (analogous to Chapter 3): ψ ∼ q , θ a | ψ ∼ p a (·;f a (ψ)), R| X,A,θ,ψ ∼ p(·| X;θ A ). sDM computes the posterior over latent effects p(ψ |D n ) and conditional posteriors p(θ a | ψ,D n ). The marginal posterior p(θ a | D n ) is obtained by integrating out ψ, yielding the reward estimate ˆr(x,a) = E[r(x,a;θ) | D n ], which is used in a greedy policy as: ˆπ g (a| x) = 1a = arg max b∈A ˆr(x,b). To analyze performance, we introduce Bayesian suboptimality (BSO) and prove that sDM achieves O(1/ √ n) convergence. The result avoids the restrictive full support assumption, which requires π 0 (a | x) ≥ γ > 0 for all actions. Instead, the bound depends on the alignment between the logging policy π 0 and the optimal policy π ∗ : performance improves smoothly as coverage of optimal actions increases. We also prove that under BSO, greedy policies are optimal and outperform pessimistic ones. This contrasts with the frequentist setting, where pessimism hedges against worst- case scenarios and is generally preferred. Experiments on synthetic and real-world data confirm that sDM outperforms existing methods, with gains increasing with K. AISTATS 2025 (Poster) - Aouali et al. (2025): • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Bayesian off-policy evaluation and learning for large action spaces. In International Conference on Artificial Intelligence and Statistics, pages 136–144. PMLR, 2025. (Chapter 7) Optimization Matters More Than Estimation This chapter challenges a common paradigm in off-policy learning. The field has traditionally focused on developing sophisticated value estimators with improved statistical properties, assuming that maximizing a more accurate estimator yields a better policy. We demonstrate that this emphasis is misplaced in large action spaces, where optimization error dominates estimation error, making even advanced estimators ineffective for policy learning. 28 Our key insight is that estimator-based objectives, despite their statistical appeal, induce highly non-concave landscapes when paired with standard policy classes. We show that gradient-based optimization can remain trapped in suboptimal plateaus forO(K) iterations, and that the landscape contains exponentially many local max- ima in K. These pathologies make global optimization intractable for large K. To characterize these failures, we analyze the oracle policies of various estimators: the policies that maximize the estimators with infinite data. This analysis reveals that each estimator induces a distinct inductive bias. For instance, standard IPS searches within the logging policy’s support, while cluster-based methods such as MIPS (Saito and Joachims, 2022) operate at the cluster level. These insights moti- vate objective-aware policy parametrizations: by aligning the policy parametrization with the estimator’s bias, we reduce the effective search space from K to the signif- icantly smaller logging support size k 0 or cluster count C, partially alleviating the optimization challenges. Ultimately, we advocate for a fundamental shift: abandoning value estimation in favor of policy-weighted log-likelihood (PWLL) objectives: ˆ U g (π) = 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i )) logπ(A i | X i ), where g is a positive weighting function. Although PWLL objectives are not value estimators, we prove they are concave (and strongly concave with ℓ 2 regularization) for linear softmax policies, while achieving oracle policies comparable to estimator- based objectives. This guarantees efficient convergence to a unique global maximum, while eliminating the optimization pathologies. Large-scale experiments on datasets with up to one million actions validate this ap- proach. Simple PWLL methods consistently outperform state-of-the-art estimator- based objectives, with the performance gap widening as action spaces grow. More- over, PWLL objectives exhibit remarkable robustness to optimization hyperparam- eters, whereas estimator-based methods require careful tuning and often fail under minor configuration changes. CONSEQUENCES, RecSys 2025 (Poster) - Submitted to ICLR 2026 - (Aouali and Sakhi, 2025): • I. Aouali and O. Sakhi. Off-policy learning in large action spaces: Optimization matters more than estimation. Under review at ICLR, 2026. (Chapter 8) Principled Pessimism for Exponential Smoothing and Beyond While sDM and PWLL offer alternative paradigms, this chapter improves the widely used family of IPS-based methods by combining variance-reducing regularization with principled pessimism for safe policy learning. The variance of IPS scales with importance weights, which can explode in large action spaces. To control this, we introduce differentiable exponential smoothing 29 (ES) estimators that regularize these weights: IPS-α : ˆ V α (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) α R i , IPS-β : ̃ V β (π) = 1 n n ∑︂ i=1 (︃ π(A i | X i ) π 0 (A i | X i ) )︃ β R i . These estimators smoothly trade variance for bias while preserving differentiability. To learn safely with these regularized estimators, we derive a two-sided PAC-Bayes generalization bound that will be used in a pessimistic objective: ⃓ ⃓ ⃓ V (π θ )− ˆ V α n (π θ ) ⃓ ⃓ ⃓ ≤ √︃ kl 1 (θ,θ 0 ) 2n + B α n (π θ ) ⏞ ⏟⏞ Regularization bias + kl 2 (θ,θ 0 ) nλ + λ 2 Var α n (π θ ) ⏞⏟⏞ Remaining variance , where kl 1,2 measure divergence between the current policy π θ and the logging policy π θ 0 , B α n captures the regularization bias, and Var α n captures the remaining variance. The exact expressions of these quantities are given in Chapter 8. The pessimistic learning objective maximizes the lower bound as: ˆπ = arg max π θ [︄ ˆ V α n (π θ )− √︃ kl 1 (θ,θ 0 ) 2n − B α n (π θ )− kl 2 (θ,θ 0 ) nλ − λ 2 Var α n (π θ ) ]︄ . This objective penalizes policies with high bias or variance, steering optimization toward reliable regions. What is important for scalability is that all terms in the ob- jective are empirical and differentiable, enabling end-to-end optimization via stan- dard stochastic gradient ascent. This contrasts with prior pessimistic objectives that relied on intractable theoretical constants or were incompatible with stochastic optimization. We further present a unified PAC-Bayes framework that generalizes this approach to all importance-weight regularizers in the literature (clipping, ES, implicit explo- ration), enabling fair comparison through a universal set of practical pessimistic objectives. This work also laid the foundation for logarithmic smoothing (see Ad- ditional Contributions below), which refines the analysis to achieve significantly tighter bounds and sharp suboptimality guarantees. ICML 2023 (Oral) - UAI 2024 (Poster) - (Aouali et al., 2023a, 2024) • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Exponential smoothing for off-policy learning. In Proceedings of the 40th International Conference on Machine Learning, pages 984–1017. PMLR, 2023a. • I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Unified PAC-Bayesian study of pessimism for offline policy learning with regularized importance sampling. In Uncertainty in Artificial Intelligence, pages 88–109. PMLR, 2024. 30 Additional Contributions This section outlines additional research conducted during this thesis. While these con- tributions are not included in the main manuscript, they correspond to published works involving equal or significant contributions from the author. On-Policy: Extension to Best-Arm Identification We extend hierarchical and structured modeling from regret minimization to fixed- budget best-arm identification (BAI), introducing prior-informed best-arm identifi- cation (PI-BAI): a non-adaptive algorithm that leverages prior knowledge for effi- cient budget allocation. We provide a fully Bayesian analysis for structured settings (e.g., linear and hier- archical bandits), departing from classical frequentist approaches. This yields the first prior-dependent guarantees on Bayesian error probability in fixed-budget BAI. PI-BAI is robust to prior misspecification and consistently outperforms baselines, including adaptive strategies, challenging the prevailing assumption that adaptivity is essential for fixed-budget exploration. AISTATS 2025 (Poster) - (Nguyen et al., 2025): • N. Nguyen, I. Aouali, A. György, and C. Vernade. Prior-dependent allocations for Bayesian fixed-budget best-arm identification in structured bandits. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS), 2025. Off-Policy: Extension to Logarithmic Smoothing We refine the theoretical framework for regularized IPS estimators from Chapter 8. While the bounds developed there were useful for learning in practice, they can be loose for suboptimality guarantees in certain cases. To address this, we derive a gen- eral high-order moment concentration bound for regularized estimators and identify the estimator that minimizes this bound. This analysis yields a novel pessimistic estimator, logarithmic smoothing (LS): ˆ V λ ls (π) = 1 nλ n ∑︂ i=1 log (︃ 1 + λ π(A i | X i ) π 0 (A i | X i ) R i )︃ . Similar to ES, LS acts as a soft, differentiable alternative to clipping, concentrates at a sub-Gaussian rate, and achieves finite variance without requiring bounded im- portance weights. The resulting high-probability risk bound is provably tighter than state-of-the-art alternatives, enabling sharp suboptimality guarantees. NeurIPS 2024 (Spotlight) - (Sakhi et al., 2024): • O. Sakhi, I. Aouali, P. Alquier, and N. Chopin. Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning. In Advances in Neural Information Processing Systems (NeurIPS), 2024. Off-Policy: Applications to Large-Scale Recommender Systems 31 This CIFRE thesis maintained a continuous feedback loop between theory and prac- tice. Our work on large-scale industrial recommender systems served a dual purpose: it provided a testing ground for the off-policy methods developed in this thesis, while the real-world challenges encountered in these systems directly motivated the theo- retical questions addressed in the main chapters. This resulted in several workshop publications and tutorials shared with the community (Gilotte et al., 2025; Aouali et al., 2022a,b,c, 2021). • A. Gilotte, O. Sakhi, I. Aouali, and B. Heymann. Offline contextual bandit with counterfactual sample identification. arXiv preprint arXiv:2509.10520, 2025. • I. Aouali, A. Benhalloum, M. Bompaire, A. Ait Sidi Hammou, S. Ivanov, B. Heymann, D. Rohde, O. Sakhi, F. Vasile, and M. Vono. Reward optimizing recommendation using deep learning and fast maximum inner product search. In proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining, pages 4772–4773, 2022a.. • I. Aouali, A. Benhalloum, M. Bompaire, B. Heymann, O. Jeunen, D. Rohde, O. Sakhi, and F. Vasile. Offline evaluation of reward-optimizing recommender systems: The case of simulation. arXiv preprint arXiv:2209.08642, 2022b. • I. Aouali, A. A. S. Hammou, O. Sakhi, D. Rohde, and F. Vasile. Probabilistic rank and reward: A scalable model for slate recommendation. arXiv preprint arXiv:2208.06263, 2022c. • I. Aouali, S. Ivanov, M. Gartrell, D. Rohde, F. Vasile, V. Zaytsev, and D. Legrand. Combining reward and rank signals for slate recommendation. arXiv preprint arXiv:2107.12455, 2021. 1.4 Related Work 1.4.1 On-Policy Learning in Contextual Bandits In the on-policy (online) setting (Slivkins, 2019; Lattimore and Szepesvari, 2019; Bubeck et al., 2012; Li et al., 2010; Chu et al., 2011), the agent must balance choosing actions that maximize current reward estimates (exploitation) with exploring other actions to improve these estimates (exploration). This trade-off is often addressed using upper confidence bounds (UCBs) (Auer et al., 2002) or Thompson sampling (TS) (Thompson, 1933). Upper confidence bound (UCB) algorithms handle the exploration-exploitation trade-off by constructing high-probability confidence intervals around reward estimates and se- lecting the action with the largest upper bound (Auer et al., 2002; Chu et al., 2011; Abbasi-Yadkori et al., 2011; Dani et al., 2008). The intuition is optimism under un- certainty: poorly explored actions have wide confidence intervals and thus large upper bounds, which encourages exploration. Despite strong theoretical guarantees, UCB meth- ods are often less practical than TS due to sensitivity to confidence-parameter tuning and lack of inherent randomization. Nevertheless, UCB remains a cornerstone of bandit the- ory and continues to inspire new exploration strategies. Part I focuses on TS, but the 32 hierarchical principles developed there naturally extend to UCB-based exploration. Thompson sampling (TS) operates in a Bayesian framework: a prior and likelihood are specified, the agent samples rewards from the posterior at each round, and then chooses the action with the highest sampled reward. TS is randomized by construction, easy to implement, and exhibits strong empirical performance in both simulated and real-world problems (Russo and Van Roy, 2014; Chapelle and Li, 2012; Russo et al., 2018). It also enjoys strong theoretical guarantees, including optimal or near-optimal regret in a variety of models (Kaufmann et al., 2012; Agrawal and Goyal, 2013b; Korda et al., 2013; Russo and Van Roy, 2014; Agrawal and Goyal, 2017; Abeille and Lazaric, 2017; Russo and Van Roy, 2016; Lu and Van Roy, 2019). Part I advances TS by integrating informative hierarchical priors that enable efficient learning in large action spaces. Hierarchical Bayesian bandits (Bastani et al., 2019; Kveton et al., 2021; Basu et al., 2021; Simchowitz et al., 2021; Wan et al., 2021; Hong et al., 2022b; Peleg et al., 2022; Wan et al., 2022; Tomkins et al., 2021; Urteaga and Wiggins, 2018) apply TS to simple graph- ical models in which action parameters are typically drawn from Gaussian distributions centered at a small number of latent parameters. These works primarily address meta- and multi-task learning in multi-armed bandits, transferring information across tasks or arms. Our mixed-effect Thompson sampling (Chapter 3) extends this line of work by introducing a hierarchical structure with multiple latent effect parameters in the contex- tual bandit setting. It also provides Bayes regret bounds and computational guarantees in the large action space regime. Our diffusion Thompson sampling (Chapter 4) further generalizes these approaches by replacing simple Gaussian hierarchies with deep, non- linear diffusion models that capture complex inter-action dependencies through flexible link functions f ℓ . Approximate Thompson sampling is a central challenge in Bayesian bandits because most posteriors are intractable and require approximate inference. Prior work (Riquelme et al., 2018; Chapelle and Li, 2012; Kveton et al., 2020) highlights the strong empirical perfor- mance of approximate TS in complex models. For mixed-effect TS (Chapter 3), we exploit the Gaussian structure of the hierarchy. We first apply a Laplace-like approximation to the reward likelihood at the action level, obtaining a Gaussian pseudo-observation on each parameter θ a . Because both the priors on latent effects and on action parameters are Gaussian and the hierarchy is linear, these pseudo-observations can then be prop- agated exactly in closed form through the hierarchy, yielding an approximate posterior that preserves the original mixed-effect structure. For diffusion TS (Chapter 4), the prior hierarchy is defined by non-linear link functions f ℓ , so Gaussian propagation is no longer exact. To retain a hierarchical diffusion model, we make an additional approximation: at each update we locally linearize the link-function updates. Combined with the Laplace- like approximation on the likelihood, this yields a chain of conditional Gaussians with updated means and precisions, i.e., a posterior diffusion model that preserves the prior hierarchy while remaining computationally tractable. Bandits with underlying structure are closely related to our setting, where we assume structured relationships among actions. In latent bandits (Maillard and Mannor, 2014; Hong et al., 2020), a single latent variable indexes multiple candidate models. In struc- tured finite-armed bandits (Lattimore and Munos, 2014; Gupta et al., 2018), each action 33 is linked to a known mean function parameterized by a common latent parameter that is learned online. TS has also been applied to more complex structures such as graphical and combinatorial bandits (Gopalan et al., 2014; Yu et al., 2020). However, these methods do not simultaneously guarantee computational and statistical efficiency in large action spaces. Meta- and multi-task learning with UCB-style methods also has a long history (Azar et al., 2013; Gentile et al., 2014; Deshmukh et al., 2017; Cella et al., 2020; Hu et al., 2021; Cella et al., 2022; Yang et al., 2020), but these works typically adopt a frequentist perspective, analyze stronger notions of regret, and often yield conservative algorithms. In contrast, our mixed-effect and diffusion TS algorithms are Bayesian, come with Bayes regret guarantees, and are explicitly designed to exploit pre-learned structure to achieve both statistical efficiency and scalable online inference. Large action spaces. Our work directly addresses the challenge of learning with per-action parameters θ a (disjoint reward models) rather than a single shared parameter θ (joint re- ward models). The disjoint formulation, while more expressive and widely used in practice, faces severe scalability issues that we address through hierarchical structure. Our analysis shows that both mixed-effect TS (Chapter 3) and diffusion TS (Chapter 4) achieve regret bounds that scale with an effective number of actions K eff ≪ K. The expression of K eff depends however on K. Some prior works (Foster et al., 2020; Xu and Zeevi, 2020; Zhu et al., 2022) propose bandit algorithms whose regret is independent of K. However, their setting differs substantially from ours: they assume a reward function r(x,a) = φ(x,a) ⊤ θ with a single shared parameter θ ∈ R d and a known mapping φ, whereas we consider r(x,a) = φ(x) ⊤ θ a (or simply r(x,a) = x ⊤ θ a ) with K separate d-dimensional action pa- rameters. The dependence on K in our setting reflects the inherent complexity of learning individual action parameters, which is the price paid for expressiveness. Obtaining a rich, known mapping φ that captures complex context-action dependencies can be challenging in practice, whereas our setting mirrors common scenarios such as recommender systems where each product has its own embedding learned from data. Note that both algorithms can be applied to the joint reward-model case; in that setting, our analysis would yield a K-independent regret bound. 1.4.2 Off-Policy Learning in Contextual Bandits The challenges of large action spaces extend beyond on-policy learning to the equally im- portant off-policy setting (Li et al., 2011; Bottou et al., 2013; Swaminathan and Joachims, 2015a), where decisions must be made using historical data collected under different poli- cies. This section surveys the off-policy contextual bandit literature and positions the methods developed in Part I. Off-policy learning fundamentally relies on off-policy evaluation, which estimates the value of a target policy π using data collected under a logging policy π 0 . Given logged data D n = (X i ,A i ,R i ) n i=1 with A i ∼ π 0 (· | X i ), off-policy evaluation seeks to estimate V (π) without deploying π. Off-policy learning then optimizes over a policy class using this estimated value. Consequently, the prevailing paradigm is estimator-centric: first design an estimator ˆ V (π) with good statistical properties, then maximize it. As we show in Chapter 7, this estimate-then-optimize approach breaks down in large action spaces because optimization error, rather than estimation error, becomes the bottleneck. 34 Inverse propensity scoring (IPS) (Horvitz and Thompson, 1952; Dudík et al., 2012) cor- rects for distribution shift via importance weights: ˆ V ips (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) R i . While unbiased under the common support assumption, IPS suffers from extremely high variance. It can also incur substantial bias when the logging policy has deficient support (Sachdeva et al., 2020), especially in large action spaces where the logging policy can only cover a small fraction of the actions. To mitigate variance, numerous importance- weight regularization techniques have been proposed, such as weight clipping (Ionides, 2008; Bottou et al., 2013) and others (Su et al., 2020; Metelli et al., 2021; Gabbianelli et al., 2024; Swaminathan and Joachims, 2015b; Gilotte et al., 2018). Chapter 8 introduces differentiable exponential-smoothing (ES) estimators that smoothly trade bias for variance and enable gradient-based optimization. Direct methods (DM) (Jeunen and Goethals, 2021; Aouali et al., 2025) avoid importance weighting by modeling the expected reward for any context–action pair and evaluating policies using: ˆ V dm (π) = 1 n n ∑︂ i=1 ∑︂ a∈A π(a| X i ) ˆr(X i ,a). DM is particularly attractive in large-scale recommender systems where IPS struggles due to its high variance (Sakhi et al., 2020; Jeunen and Goethals, 2021; Aouali et al., 2022c). However, standard implementations typically rely on a disjoint model that estimates one parameter vector θ a ∈ R d per action. In large action spaces with sparse logging, this leads to severe statistical inefficiency: many actions are rarely observed and thus poorly estimated. Chapter 6 addresses this via the structured direct method (sDM), which leverages hierarchical Bayesian modeling to share statistical strength across actions. Direct methods (DM) (Jeunen and Goethals, 2021; Aouali et al., 2025) build a model of the expected reward for any context–action pairs and evaluate policies using ˆ V dm (π) = 1 n n ∑︂ i=1 ∑︂ a∈A π(a| X i ) ˆr(X i ,a). DM is particularly attractive in large-scale recommender systems, where IPS struggles (Sakhi et al., 2020; Jeunen and Goethals, 2021; Aouali et al., 2022c). Standard imple- mentations, however, typically use a disjoint model that estimates one parameter vector θ a per action. In large action spaces with sparse logging, this leads to severe statistical inefficiency: many actions are rarely observed and thus poorly estimated. Chapter 6 intro- duces the structured direct method (sDM), which addresses this limitation via hierarchical Bayesian modeling. Doubly robust (DR) estimators (Robins and Rotnitzky, 1995; Bang and Robins, 2005; Dudík et al., 2011; Dudik et al., 2014; Farajtabar et al., 2018) combine DM and IPS to achieve robustness: ˆ V dr (π) = ˆ V dm (π) + 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) (︁ R i − ˆr(X i ,A i ) )︁ . 35 DR has become a default choice in many off-policy evaluation studies (Dudík et al., 2011; Dudik et al., 2014; Farajtabar et al., 2018; Su et al., 2020). Both of our contributions in Chapter 8 (regularized importance weights) and Chapter 6 (structured reward models) can be used to enhance the components of DR. Large-scale IPS variants. Importance-weight regularization alone is often insufficient when the action space is very large. Structural assumptions can dramatically reduce the variance. For instance, marginalized IPS (MIPS) (Saito and Joachims, 2022) clusters actions via a mapping h(a) and works with cluster-level importance weights: ˆ V mips (π) = 1 n n ∑︂ i=1 π(h(A i )| X i ) π 0 (h(A i )| X i ) R i . This reduces variance by operating over a smaller cluster space instead of the full action set. This dimensionality reduction principle has inspired numerous extensions (Peng et al., 2023; Sachdeva et al., 2024; Cief et al., 2024; Taufiq et al., 2024; Saito et al., 2023). Our sDM (Chapter 6) provides a complementary structural approach designed for DM instead of IPS; it can be viewed as a Bayesian latent-structure counterpart to MIPS, replacing hard clustering with soft probabilistic coupling. Pessimistic off-policy learning. Maximizing a point estimator (IPS, DM, or DR) can be unsafe when the estimator deviates from the true value. Pessimistic approaches instead construct lower confidence bounds on V (π) and optimize those, following the principle of pessimism in the face of uncertainty (Jin et al., 2021). Asymptotic and finite-sample lower bounds have been developed for various estimators (Bottou et al., 2013; Kuzborskij et al., 2021; Gabbianelli et al., 2024), providing worst-case guarantees on policy performance. Many pessimistic learning methods are directly motivated by such bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Kuzborskij et al., 2021; Aouali et al., 2023a; Wang et al., 2023). For example, Swaminathan and Joachims (2015a) combine empirical-Bernstein inequalities with clipped IPS, leading to variance-penalized learning objectives. Recently, the PAC-Bayesian paradigm (McAllester, 1998; Catoni, 2007; Alquier, 2021) has been increasingly applied to off-policy learning, offering a flexible toolkit for deriving data-dependent generalization bounds. London and Sandler (2019) introduced a scal- able PAC-Bayesian perspective, which has been further developed by Flynn et al. (2023); Sakhi et al. (2022); Aouali et al. (2023a, 2024); Gabbianelli et al. (2024) to yield tight, directly optimizable bounds. Chapter 8 advances this direction by deriving a two-sided PAC-Bayes bound for exponentially smoothed IPS estimators, resulting in a fully differ- entiable pessimistic learning objective that jointly controls reward, bias, variance, and divergence from the logging policy. Moreover, this framework generalizes to encompass other estimators in the literature, establishing a unified set of principles for pessimistic learning. While our subsequent work on logarithmic smoothing (LS) (Sakhi et al., 2024) further tightens these guarantees, Chapter 8 lays the foundational theoretical groundwork on which LS builds. Optimization-centric learning. All methods above follow an estimator-centric philosophy: find a good value estimator and maximize it. In large action spaces, these estimator- based objectives typically induce highly non-concave landscapes with flat plateaus and 36 many local maxima (Chapter 7), ensuring that optimization error dominates estima- tion error. Chapter 7 proposes a paradigm shift toward optimization-centric off-policy learning. Rather than insisting on high-quality estimators of the value, we advocate for policy-weighted log-likelihood objectives whose optimization landscape is benign, ensuring convergence to effective policies even in massive action spaces. 37 Part I On-Policy Learning in Large Action Spaces 38 Chapter 2 Introduction to Part I This first part of the thesis addresses the following fundamental question: How can we design exploration-exploitation algorithms that remain both statistically efficient and computationally feasible when the number of actions is large? 2.1 Setting and Background In this part, we consider the on-policy (online) contextual bandit setting where an agent interacts with an environment over T rounds. At each round t∈ [T ]: 1. The agent observes a context X t ∈X ⊆ R d drawn from a distribution ν; 2. The agent selects an action A t ∈A = [K] based on the history H t =(X s ,A s ,R s ) t−1 s=1 ; 3. The agent receives a stochastic reward R t ∼ p(·| X t ;θ ∗,A t ). Each action a ∈ [K] is associated with an unknown true parameter θ ∗,a ∈ R d . The true expected reward is given by r(x,a;θ ∗ ) = r(x;θ ∗,a ), where θ ∗ = (θ ∗,a ) a∈[K] ∈ R Kd denotes the concatenation of all true action parameters. Throughout this part, we assume the reward distribution is a generalized linear model (GLM) (McCullagh and Nelder, 1989). For any context x ∈ X and action a ∈ A, p(· | x;θ ∗,a ) is an exponential-family distribution with mean g(φ(x) ⊤ θ ∗,a ), where g is the mean function. This formulation recovers linear bandits (Auer, 2002) when p(·| x;θ ∗,a ) = N (·;φ(x) ⊤ θ ∗,a ,σ 2 ) (with identity link g(u) = u), and logistic bandits (Filippi et al., 2010) when p(·| x;θ ∗,a ) = Ber(g(φ(x) ⊤ θ ∗,a )) with the sigmoid link g(u) = (1 + exp(−u)) −1 . We adopt a Bayesian perspective where the unknown true parameters θ ∗ are assumed to be drawn from a prior distribution p 0 . Our objective is to minimize the Bayes regret: BR(T ) = E [︄ T ∑︂ t=1 (︂ r(X t ,A t,∗ ;θ ∗ )− r(X t ,A t ;θ ∗ ) )︂ ]︄ ,(2.1) where A t,∗ = arg max a∈[K] r(X t ,a;θ ∗ ) is the optimal action at round t. The expectation is taken over the prior p 0 , the stochastic rewards, contexts, and the agent’s policy. 39 2.1.1 Scalability Challenges Standard Thompson sampling (TS) maintains a posterior distribution p(θ | H t ) over the parameter space and sampling θ t ∼ p(· | H t ) at each round to select the action A t = arg max a∈[K] r(X t ,a;θ t ). However, when the number of actions K is large, there is trade-off between statistical and computational efficiency: Statistical inefficiency of disjoint priors. A common simplification is to learn each action’s parameter θ a independently using a factorized prior p 0 (θ) = ∏︁ K a=1 p 0,a (θ a ). While this makes posterior updates computationally cheap, it prevents information sharing. The agent must learn about every action from scratch, which is prohibitive in large action spaces. Computational intractability of joint priors. Conversely, modeling dependencies via a full joint posterior over R dK allows for information sharing but is computationally intractable. Storing the covariance requiresO(K 2 d 2 ) memory, and a single update requires O(K 3 d 3 ) time, making it intractable for online interaction. 2.2 Hierarchical Models To address this dilemma, we propose a general hierarchical Bayesian framework where action parameters are coupled through a set of latent parameters Ψ = (ψ ℓ ) ℓ∈[L] ∈ R Ld , with L≪ K. The generative process is defined as: Ψ∼ q 0 (·)(Prior over latent structure),(2.2) θ a | Ψ∼ p 0,a (·| Ψ), ∀a∈ [K] (Conditional prior per action), (2.3) R t | θ, Ψ,X t ,A t ∼ p(·| X t ;θ A t ), ∀t∈ [T ] (Reward observation).(2.4) Here, q 0 encodes global uncertainty, while p 0,i specifies how individual actions deviate from the shared structure. The resulting marginal prior p 0 (θ) naturally couples all action parameters. The corresponding posterior preserves this hierarchy: p(θ, Ψ| H t ) = p(Ψ| H t ) K ∏︂ a=1 p(θ a | Ψ,H t,a ).(2.5) The latent posterior p(Ψ| H t ) aggregates evidence from all actions, enabling global infor- mation sharing, while the action posteriors p(θ a | Ψ,H t,a ) allow for efficient sampling of action parameters θ a independently given Ψ. Thompson sampling then proceeds by first sampling the global structure Ψ t , then sampling action parameters θ t,a conditioned on Ψ t . 2.3 Roadmap of Part I The following chapters present two concrete instantiations of this hierarchical framework. Chapter 3: Mixed-Effects Thompson Sampling. We begin by investigating a lin- ear instantiation of the hierarchy where each action parameter is modeled as a linear combination of L shared effects, θ a | Ψ ∼ N ( ∑︁ L ℓ=1 b a,ℓ ψ ℓ , Σ 0,a ), with known weights b a,ℓ . 40 We derive mixed-effect Thompson sampling (meTS), an algorithm that exploits the con- jugacy of linear-Gaussian models to perform exact, closed-form posterior updates. For non-linear GLM rewards, we introduce a tractable Laplace approximation that we propa- gate through the hierarchy. We provide theoretical guarantees showing that meTS achieves a Bayes regret bound scaling with the effective number of actions K eff ≪ K. Chapter 4: Diffusion Thompson Sampling. We then extend the framework to support deep, non-linear hierarchical structures using diffusion models. In this setting, the priors form a Markov chain of latent variables ψ L → · → ψ 1 → θ a , connected by potentially non-linear link functions f ℓ (e.g., neural networks) learned from offline data. We propose diffusion Thompson sampling (dTS) and develop a posterior approximation that updates the link functions and covariances to match observed data while preserving the generative diffusion. This chapter demonstrates how to leverage powerful generative priors for exploration while remaining computationally feasible for online deployment. 41 Chapter 3 Scaling Thompson Sampling with Mixed Effects Contents 2.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . 39 2.1.1 Scalability Challenges . . . . . . . . . . . . . . . . . . . . . . . 40 2.2 Hierarchical Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 2.3 Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 This chapter begins with the fundamental observation that the expected rewards of actions in real-world problems are often correlated. To model this phenomenon, we study a structured mixed-effect bandit environment in which each action parameter depends on one or more effect parameters that are shared across actions. Therefore, taking an action teaches the agent about its effect parameters, thereby informing it about other actions that share the same effect parameters. We present three motivating examples for this. Movie recommendation. Here, we want to recommend a movie to a user with the highest expected rating. User j and movie a are represented by vectors x j (context) and θ a (action parameter), respectively. The expected rating that user j gives to movie a is x ⊤ j θ a . We assume that the vector x j is observed. Then the most natural idea is to learn all θ a individually using standard bandit methods (Li et al., 2010; Chu et al., 2011). This is statistically inefficient when the number of movies is high. Fortunately, the movies could be organized into L categories and such information can be leveraged to explore efficiently. We present three approaches (A), (B) and (C) that do this next. (A) For each category ℓ∈ [L], a parameter ψ ℓ is learned online using all interactions with the movies in category ℓ. The parameter ψ ℓ represents all the movies in category ℓ and is used instead of their individual θ a . Therefore, this approach has a high bias, as all movies in the same category are assumed to have the same expected rating. This issue can be addressed by a better model. (B) We model each movie parameter θ a as a random variable centered in its category parameter ψ ℓ . Now movies in the same category no longer have 42 the same expected rating due to the additional uncertainty. Both the category parameters ψ ℓ and movie parameters θ a are learned online. The former is learned using all interactions with the movies in category ℓ, while the latter is learned using all interactions with movie a conditioned on ψ ℓ . The category parameter ψ ℓ is learned using more data, which helps to learn θ a more efficiently. This is a special case of our setting. (C) The shortcoming of (B) is that each movie belongs to a single category, which is unrealistic. To address this issue, we allow movies to belong to multiple categories and then proceed as in (B). To connect with our terminology, the categories ℓ ∈ [L] denote effects; their parameters ψ ℓ are the effect parameters; and the movie parameters θ a are the action parameters. Ad placement: Here, the agent selects a list (or slate) of M items from a catalog of L items with the objective of maximizing the click-through-rate. We assume that the agent receives only binary bandit feedback indicating whether the user clicked one of the items in the slate (Dimakopoulou et al., 2019; Rejwan and Mansour, 2020). Again, user j and slate a are represented by x j (context) and θ a (action parameter), respectively. The corresponding click-through-rate is g(x ⊤ j θ a ), where g is the sigmoid function. The set of slates (of size K ≈ L M ) is exponentially large, which makes learning θ a individually difficult. Fortunately, the slates are related through a much smaller set of items (of size L). Therefore, slates containing common items can teach the agent about one another, enabling efficient exploration. Efficient exploration is achieved by decomposing slate a’s parameter as θ a = ∑︁ ℓ∈[L] b a,ℓ ψ ℓ + ε a . Here ψ ℓ ∈ R d is the parameter of item ℓ and b a,ℓ ∈ R is a mixing weight that captures position biases. That is, b a,ℓ = 0 if item ℓ is not in slate a, and b a,ℓ is high if item ℓ is ranked high in slate a. This captures the fact that the probability of a click on an item is influenced by its position on the slate, and this bias can be estimated offline. Finally, ε a is a random noise that can incorporate uncertainty due to model misspecification, for instance due to an estimation error of b a,ℓ . The benefit of this decomposition is that the parameter of item ℓ, ψ ℓ , is learned using all interactions with the slates with item ℓ. The slate parameter θ a is learned using all interactions with slate a conditioned on ψ ℓ . This is more statistically efficient than learning θ a individually, which only uses the interactions with slate a. Drug design: Here, the goal is to find the optimal drug design in clinical trials (Durand et al., 2018). Subject j and drug a are represented by vectors x j and θ a , respectively, and the expected efficacy of drug a for subject j is x ⊤ j θ a . Again, the most natural idea is to learn all drug parameters θ a individually. This leads to statistical inefficiency when the number of candidate drugs is high. Fortunately, we can leverage the fact that drug candidates in the same trial often share components to explore efficiently. Precisely, a drug is a combination of multiple components, each with a specific dosage. Each component ℓ is represented by a parameter ψ ℓ , and the drug parameter θ a is a known combination of the component parameters ψ ℓ weighted by their dosage. That is, θ a = ∑︁ ℓ∈[L] b a,ℓ ψ ℓ + ε a , where b a,ℓ is the dosage of component ℓ in drug a and ε a is a random noise to incorporate uncertainty due to model misspecification. The efficacy of each component has an effect on the overall efficacy of the drug and is boosted by the dosage. In all examples, we assume an underlying structure among the actions, that they are affected by multiple effects. In some problems, it is known how the effect arises. For 43 instance, in the drug design, the actions are the drugs and the effects are their components. The mixing weight that relates an action (drug) to an effect (component) is the dosage of that component in the drug. In other problems, it may not be apparent how the effect arises and this has to be learned. We discuss this in detail in Section 3.1.3. We make the following contributions. 1) We formalize a general mixed-effect bandit framework represented by a two-level graphical model where each action is associated with a d-dimensional parameter that depends on one or multiple effect parameters. 2) We design mixed-effect Thompson sampling (meTS), which leverages this structure to be both statistically and computationally tractable. We show that closed-form posteriors can be derived for Gaussian instances and efficient approximations exist in more general cases. 3) We prove that the Bayes regret of meTS is bounded by a sum of two terms: one is associated with learning the action parameters and the other quantifies the cost of learning the effect parameters. Both terms reflect the structure of the environment and the quality of priors. 4) We show empirically that meTS and its variants perform extremely well, and are computationally efficient in both synthetic and real-world problems. 3.1 Setting We consider the contextual bandit setting in Section 2.1. Each action a ∈ A = [K] is associated with an unknown d-dimensional action parameter θ a ∈ R d . The correlations between the action parameters arise because they are derived from L shared unknown d- dimensional effect parameters, ψ ℓ ∈ R d for ℓ∈ [L]. Specifically, we assume that the action parameter θ a is sampled from the action prior distribution p 0,a as θ a | Ψ ∼ p 0,a (· | Ψ), where Ψ = (ψ ℓ ) ℓ∈[L] ∈ R Ld is a concatenation of the effect parameters. The distribution p 0,a can capture sparsity, when θ a depends only on a subset of Ψ; and also incorporate uncertainty due to model misspecification, when θ a is not a deterministic function of Ψ. Finally, the effect parameters Ψ are sampled from a joint effect prior q 0 , which is known by the agent and represents its initial uncertainty about Ψ. In summary, all variables in our environment are generated as Ψ∼ q 0 ,(3.1) θ a | Ψ∼ p 0,a (·| Ψ),∀a∈A, R t | X t ,A t ,θ, Ψ∼ p(·| X t ;θ A t ),∀t∈ [T ], where p(·| x;θ a ) is the reward distribution of action a in context x, which only depends on parameter θ a and the context x. The terminology of effect parameters arises from the fact that ψ ℓ affect the model parameters θ a , which in turn define R t . The effects are mixed through the action prior p 0,a and hence the name mixed-effect. Our setting can be viewed as a two-level graphical model, where ψ 1 ,...,ψ L are parent nodes and θ 1 ,...,θ K are child nodes (Figure 3.1). The structure is represented by missing arrows from parent (effect parameters) to child (action parameters) nodes. A missing arrow from parent ψ ℓ to child θ a means that action a is independent of the ℓ-th effect. Our model can capture all examples provided in the introduction of this chapter. For instance, in movie recommendation, the categories ℓ ∈ [L] and movies a ∈ A would 44 : taken action at round Figure 3.1: Example of a graphical model induced by Equation (3.1). be represented by the effect parameters ψ ℓ and action parameters θ a , respectively. The weight b a,ℓ is the relevance of movie a to category ℓ. Linearity in effects. A simple yet powerful assumption is that the action prior p 0,a is parametrized by a weighted sum of effect parameters θ a | Ψ∼ p 0,a (︂ · ⃓ ⃓ ⃓ L ∑︂ ℓ=1 b a,ℓ ψ ℓ )︂ ,∀a∈A, where b a = (b a,ℓ ) ℓ∈[L] ∈ R L are L known mixing weights for action a. The effect ℓ on action a is determined by b a,ℓ . As an example, b a,ℓ = 0 when action a is independent of effect ℓ. This is an important special case of our setting, since additive models are widely used in both theory and practice, as they often yield closed-form posteriors that are computationally tractable. Next we present two instances of this setting, where p 0,a is a multivariate Gaussian with mean ∑︁ L ℓ=1 b a,ℓ ψ ℓ and covariance Σ 0,a . 3.1.1 Mixed-Effect Linear Bandit A natural joint effect prior q 0 for d-dimensional effect parameters ψ ℓ is a multivariate Gaussian with mean μ Ψ ∈ R Ld and covariance Σ Ψ ∈ R Ld×Ld . The action prior p 0,a is a Gaussian with mean ∑︁ L ℓ=1 b a,ℓ ψ ℓ ∈ R d and covariance Σ 0,a ∈ R d×d : Ψ∼N (μ Ψ , Σ Ψ ),(3.2) θ a | Ψ∼N (︂ L ∑︂ ℓ=1 b a,ℓ ψ ℓ , Σ 0,a )︂ ,∀a∈A, R t | X t ,A t ,θ, Ψ∼N (X ⊤ t θ A t ,σ 2 ),∀t∈ [T ], where σ 2 > 0 is the variance of the observation noise. 45 3.1.2 Mixed-Effect Generalized Linear Bandit Here the effect and action parameters are generated as in Equation (3.2) but the reward R t is sampled from a generalized linear model (GLM) (McCullagh and Nelder, 1989), which is non-linear. In particular, p(· | X t ;θ a ) is an exponential-family distribution with mean g(X ⊤ t θ a ) and the whole model is Ψ∼N (μ Ψ , Σ Ψ ),(3.3) θ a | Ψ∼N (︂ L ∑︂ ℓ=1 b a,ℓ ψ ℓ , Σ 0,a )︂ ,∀a∈A, R t | X t ,A t ,θ, Ψ∼ p(·| X t ;θ A t ),∀t∈ [T ]. Let Ber(p) be a Bernoulli distribution with mean p. One particular choice of a GLM is g(u) = 1/(1 + exp(−u)) and p(· | X t ;θ) = Ber(g(X ⊤ t θ)), which corresponds to a logistic bandit (Filippi et al., 2010). Remark 3. Note that in both settings, we use x ⊤ θ instead of φ(x) ⊤ θ for some feature- map φ, but this is just for ease of exposition, and everything generalizes smoothly to when using φ. 3.1.3 Structure Learning The structures in Equations (3.2) and (3.3) may be intrinsic in some problems, such as drug design. When this is not the case, we propose the following approach to learning a proxy structure. For any a ∈ A, let ˆ θ a represent an offline estimate of action parameter θ a (e.g., learned offline using interactions from previous bandit tasks). To learn, we fit a Gaussian mixture model (GMM) (Reynolds et al., 2009) with L clusters to ˆ θ a . Each cluster ℓ ∈ [L] is represented by its center μ ψ ℓ ∈ R d and covariance Σ ψ ℓ ∈ R d×d . These correspond to the mean of the effect parameter ψ ℓ and its uncertainty. The GMM also outputs the probability that ˆ θ a belongs to cluster ℓ, for all combinations of a ∈ A and ℓ∈ [L]. This probability is the mixing weight b a,ℓ . The proposed procedure is general and adaptable to a wide range of use cases. The primary challenge lies in deriving the offline estimates ˆ θ a . A straightforward approach involves learning these parameters from historical data collected in previous bandit tasks. Broadly, this can be formulated as an offline representation-learning problem (Tripura- neni et al., 2021), for which numerous techniques exist. For instance, in our MovieLens experiments (Section 3.4.2), we employ a low-rank factorization of the rating matrix to obtain these estimates. A key strength of our approach is its flexibility; it integrates seamlessly with standard offline learning tools, thereby taking a step toward bridging the gap between offline and online learning. 3.2 Algorithm We propose a Thompson sampling algorithm (Thompson, 1933; Russo and Van Roy, 2014; Scott, 2010), which is a natural Bayesian solution to our problem. The algorithm 46 Algorithm 1 meTS: Mixed-Effect Thompson Sampling. Input: Joint effect prior q 0 , action priors p 0,· Initialize q 1 ← q 0 and p 1,· ← p 0,· for t = 1,...,T do Sample Ψ t ∼ q t for a = 1,...,K do Sample θ t,a ∼ p t,a (·| Ψ t ) θ t ← (θ t,a ) a∈A A t ← argmax a∈A r(X t ,a;θ t ) Receive reward R t ∼ p(·| X t ;θ ∗,A t ) Compute new posteriors q t+1 and p t+1,· is based on hierarchical sampling (Lindley and Smith, 1972), which reflects the structure in our model. Before we present it, we need to introduce additional notation. We denote by H t = (X i ,A i ,R i ) i∈[t−1] the history of all interactions of the agent up to round t, by S t,a =i∈ [t− 1] : A i = a the rounds where the agent takes action a up to round t, and by H t,a = (X i ,A i ,R i ) i∈S t,a the corresponding history. Our algorithm meTS is presented in Algorithm 1. Since effect parameters are shared across actions, their posteriors exhibit dependencies. To handle this, we maintain two types of posterior densities: • A joint effect posterior q t (Ψ) = p(Ψ| H t ) for all effect parameters Ψ in round t; • An action posterior p t,a (θ | Ψ) = p(θ a | H t,a , Ψ) for each action a ∈ A, conditioned on the effect parameters. meTS employs hierarchical sampling in each round t: 1. Sample effect parameters: Ψ t ∼ q t (·) 2. Sample action parameters: θ t,a ∼ p t,a (·| Ψ t ) for each a∈A 3. Select action: A t = argmax a∈A r(X t ,a;θ t ) where θ t = (θ t,a ) a∈A . This hierarchical sampling scheme is equivalent to sampling from the exact marginal posterior p(θ a | H t ). To see this, observe that marginalizing over Ψ yields: p(θ a | H t ) = ∫︂ Ψ p(θ a , Ψ| H t ) dΨ, = ∫︂ Ψ p(θ a | Ψ,H t )p(Ψ| H t ) dΨ, = ∫︂ Ψ p t,a (θ a | Ψ)q t (Ψ) dΨ.(3.4) 47 3.2.1 Posterior Derivations The posteriors are computed as follows. We first express the joint effect posterior q t as q t (Ψ)∝ K ∏︂ a=1 ∫︂ θ a L t,a (θ a )p 0,a (θ a | Ψ) dθ a q 0 (Ψ),(3.5) whereL t,a (θ a ) = ∏︁ (x,a,r)∈H t,a p(r | x;θ a ) is the likelihood of all observations of action a up to round t given θ a . Next, for any action a∈A, the action posterior p t,a is expressed as p t,a (θ a | Ψ)∝L t,a (θ a )p 0,a (θ a | Ψ).(3.6) p t,a is similarly sparse to p 0,a . Specifically, in any round t, p t,a and p 0,a are parameterized by the same subset of effect parameters Ψ, since L t,a (θ a ) does not depend on Ψ. The joint effect posterior q t and action posteriors p t,a have closed forms in Gaussian models, which allows efficient sampling and theoretical analysis. Beyond these, MCMC and variational inference can be used to approximate q t and p t,a . Next we derive closed- form posteriors for the mixed-effect model with linear rewards in Equation (3.2) and provide an efficient approximation for the mixed-effect model with non-linear rewards in Equation (3.3). 3.2.2 Mixed-Effect Linear Bandit Let G t,a = σ −2 ∑︂ i∈S t,a X i X ⊤ i ,B t,a = σ −2 ∑︂ i∈S t,a R i X i .(3.7) be the outer product of contexts corresponding to action a up to round t, and their sum weighted by rewards, respectively. Both are scaled by the observation noise variance σ 2 . Using these quantities, the effect posterior is defined as follows. Proposition 1. For any round t∈ [T ], the joint effect posterior is a multivariate Gaus- sian q t =N ( ̄μ t , ̄ Σ t ), where ̄ Σ −1 t = Σ −1 Ψ + ∑︂ a∈A b a b ⊤ a ⊗ (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ ,(3.8) ̄μ t = ̄ Σ t (︂ Σ −1 Ψ μ Ψ + ∑︂ a∈A b a ⊗ (Σ −1 0,a (G t,a + Σ −1 0,a ) −1 B t,a ) )︂ . The effect posterior is additive in individual actions and each action contributes to the effect posterior mean and covariance proportionally to b a,ℓ , which is the mixture weight for θ a in Equation (3.2). Proposition 1 is proved in Section A.2.1. Now we present the action posterior. 48 Proposition 2. For any round t ∈ [T ], action a ∈ A, and effect parameters Ψ t , the action posterior is a multivariate Gaussian p t,a (·| Ψ t ) =N (·; ̃μ t,a , ̃ Σ t,a ), where ̃ Σ −1 t,a = Σ −1 0,a + G t,a ,(3.9) ̃μ t,a = ̃ Σ t,a (︂ Σ −1 0,a L ∑︂ ℓ=1 b a,ℓ ψ t,ℓ + B t,a )︂ . The action posterior in Equation (3.9) is a standard multivariate Gaussian posterior whose prior depends on Ψ t , which is sampled by meTS. Proposition 2 is proved in Section A.2.2. 3.2.3 Mixed-Effect Generalized Linear Bandit Closed-form posteriors are unavailable in this setting, so approximations are required. We use a Laplace-style scheme that approximates the likelihood L t,a (·) by a Gaussian, rather than applying Laplace to the full posterior. This choice preserves a Gaussian form that can be propagated analytically through the hierarchical updates. Let μ lap t,a denote the MLE (see remark below for a discussion about the computation of the MLE in practice), and let G lap t,a be the Hessian 1 of − logL t,a (·): μ lap t,a = argmax θ a logL t,a (θ a ), G lap t,a = ∑︂ i∈S t,a ̇g(X ⊤ i μ lap t,a )X i X ⊤ i . We then approximate the likelihood (not the posterior) by L t,a (θ a ) ∝ exp (︁ − 1 2 (θ a − μ lap t,a ) ⊤ G lap t,a (θ a − μ lap t,a ) )︁ ,(3.10) Substituting Equation (3.10) into Equation (3.5) yields q t (·)≈N (·; ̄μ t , ̄ Σ t ), where ̄μ t and ̄ Σ t are computed as in Proposition 1, except for the replacements G t,a ← G lap t,a , B t,a ← G lap t,a μ lap t,a . Similarly, substituting Equation (3.10) into Equation (3.6) gives p t,a (·| Ψ)≈N (·; ̃μ t,a , ̃ Σ t,a ), with G t,a ← G lap t,a , B t,a ← G lap t,a μ lap t,a . 1 Note that we are assuming a generalized linear model where the log-likelihood of the data associated with action a can be written as logL t,a (θ a ) = ∑︁ i∈S t,a [︁ R i X ⊤ i θ a − A(X ⊤ i θ a ) + C(R i ) ]︁ where C is a real- valued function and A is twice continuously differentiable, with derivative ̇ A = g representing the mean function. 49 Although these expressions follow mechanically from substituting the Gaussian likelihood approximation, the intuition is straightforward. The replacement G t,a ← G lap t,a reflects the curvature induced by the nonlinear mean function g, while B t,a ← G lap t,a μ lap t,a mirrors the linear-Gaussian case, where the MLE ˆ θ mle t,a is characterized by the normal equations G t,a ˆ θ mle t,a = B t,a , and its generalized-linear counterpart is μ lap t,a . Remark 4. The MLE μ lap t,a = argmax θ a ∈R d logL t,a (θ a ) may be ill-posed. In practice, we use a small ℓ 2 -regularized estimator: μ lap t,a ∈ argmax θ a ∈R d logL t,a (θ a )− λ 2 ∥θ a ∥ 2 2 , where λ > 0 to fix this. 3.2.4 Computational Complexity The benefit of modeling the effect parameters is not immediately clear. Thus, it is tempt- ing to marginalize them out, and only maintain a single joint posterior of all action pa- rameters θ ∈ R Kd . Posterior updates in this case would be complex and computationally inefficient when K ≫ L, which is common in practice. The main advantage of meTS is that the sampling of effect parameters Ψ t ∼ q t allows us to use the conditional independence of actions given Ψ, and model θ a | H t,a , Ψ t independently. This is more computationally efficient than modeling the joint θ | H t when K ≫ L. To see this, suppose that all posteriors are multivariate Gaussians (Section 3.2.2). In this case, θ | H t requires O(K 2 d 2 ) space, due to storing a Kd× Kd covariance matrix; while meTS requires only O((L 2 + K)d 2 ) space, due to storing the covariances of q t and p t,a . Since the sampling relies on covariance inverses, the time complexity also improves. For the joint posterior, it is O(K 3 d 3 ), while it is only O((L 3 + K)d 3 ) for meTS. One can also marginalize out the effect parameters Ψ and have K separate posteriors, one for each action parameter θ a . While this improves computational efficiency, it does not model that the actions are correlated, since θ a | H t,a is modeled instead of θ a | H t . This leads to a statistical inefficiency due to the loss of information as the histories of other actions H t,a ′ are discarded. We validate this through theory (Section 3.3.2) and experiments (Section 3.4). 3.3 Analysis This section is organized as follows. First, we state our regret bound. Then, we discuss how it captures the structure of our problem. We use ̃ O for the big O notation up to polylogarithmic factors. 3.3.1 Main Result We analyze meTS in the linear setting in Section 3.1.1. Throughout, we assume that the true action parameters and rewards are generated according to the same hierarchical model used by meTS (Equation (3.2)), i.e., we operate in the fully well-specified setting. For ease of exposition, we further assume the existence of constants σ 0 ,σ Ψ ,κ x > 0 such 50 that Σ 0,a = σ 2 0 I d for all a∈A,Σ Ψ = σ 2 Ψ I Ld , ∥X t ∥ 2 2 ≤ κ x for all t∈ [T ]. The bound on ∥X t ∥ 2 is standard, and we relax the other two assumptions in Section A.3. Theorem 1. For any δ ∈ (0, 1), the Bayes regret of meTS in the mixed-effect model in Section 3.1.1 is bounded as BR(T )≤ √︁ 2T (R a (T ) +R e (T )) log(1/δ) + cTδ ,(3.11) where c = √︂ 2 π κ x (σ 2 0 + κ b σ 2 Ψ )K , κ b = max a∈A ∥b a ∥ 2 2 , R a (T ) = dKc a log (︁ 1 + Tκ x σ 2 0 dσ 2 )︁ , c a = κ x σ 2 0 log (︁ 1 + κ x σ 2 0 σ 2 )︁ , R e (T ) = dLc e log (︁ 1 + Kκ b σ 2 Ψ σ 2 0 + σ 2 Tκ x )︁ , c e = κ x κ b σ 2 Ψ (︁ 1 + κ x σ 2 0 σ 2 )︁ log (︁ 1 + κ x κ b σ 2 Ψ σ 2 )︁ . The second term in Equation (3.11) is constant for δ = 1/T, in which case the above bound is ̃ O( √ T ). The main quantities of interest are R a (T ) and R e (T ), and they have natural interpretations. R a (T ) corresponds to the action regression problem: with K parameters of dimension d, prior width σ 0 , maximum context length √ κ x , and T observations with noise σ. The dependence of R a (T ) on these quantities is identical to a corresponding linear bandit (Lu and Van Roy, 2019). On the other hand, R e (T ) corresponds to the effect regression problem: with L parameters of dimension d, prior width σ Ψ , maximum mixing-weight length √ κ b , and K actions that can be viewed as observations with noise σ 0 (Section 3.2.2). The dependence of R e (T ) on these quantities mimics those in R a (T ). To simplify exposition, let κ x = κ b = σ = 1. Then BR(T ) = ̃ O (︃ √︂ Td (︁ Kσ 2 0 + Lσ 2 Ψ (1 + σ 2 0 ) )︁ )︃ .(3.12) This can be re-written as BR(T ) = ̃ O (︂ √︁ TdK eff (σ 2 0 + σ 2 Ψ ) )︂ , where K eff = Kσ 2 0 +Lσ 2 Ψ (1+σ 2 0 ) σ 2 0 +σ 2 Ψ is the effective number of actions. When L ≪ K and σ 2 0 ≪ σ 2 Ψ , we have K eff ≪ K, yielding significant regret reduction over standard Thompson Sampling. The dependence on σ 2 0 and σ 2 Ψ is natural: since Bayesian regret measures performance under the prior, smaller prior variances correspond to more informative beliefs about the true parameters, which makes learning easier and reduces regret. Conversely, larger variances reflect greater prior uncertainty and increase the difficulty of identifying the optimal action. The scaling with K, L, and d is also intuitive: fewer parameters to estimate lead to lower regret. These trends are consistent with our empirical observations in Section A.4. 51 3.3.2 Benefits of Structure Note that we do not provide a matching lower bound. To argue that our upper bound reflects the intrinsic structure of the problem, we compare meTS to agents that either have access to more information or exploit less structure. We start with the former. Consider meTS with known effect parameters Ψ. Setting σ Ψ = 0 in Equation (3.12) yields the reduced regret BR(T ) = ̃ O( √︂ TdKσ 2 0 ), which no longer depends on L. Likewise, consider meTS under a perfectly specified linear model, in which θ a = ∑︁ ℓ∈[L] b a,ℓ ψ ℓ for all a∈A. This corresponds to σ 0 = 0, giving BR(T ) = ̃ O( √︂ TdLσ 2 Ψ ), which is independent of K. In particular, the K-dependence in our regret bound arises precisely from modeling the variability of action parameters around the effect parameters via Σ 0,a . Without it, the regret of sDM is independent of K We now turn to an agent that neither knows Ψ nor models it explicitly. This agent learns only θ (Section 3.2.4) by marginalizing out Ψ in Equation (3.2): θ a ∼N (︂ L ∑︂ ℓ=1 b a,ℓ μ ψ ℓ , ̆ Σ 0,a )︂ , ∀a∈A, where ̆ Σ 0,a = (σ 2 0 +∥b a ∥ 2 2 σ 2 Ψ )I d is the marginal prior covariance and (μ ψ ℓ ) ℓ∈[L] is the prior mean of the effects, so that μ Ψ = (μ ψ ℓ ) ℓ∈[L] (Section 3.1.1). Importantly, marginalizing out Ψ and treating each action parameter independently discards action correlations, even though the true generative model induces correlations via the shared effects. This agent therefore uses a less structured and less informative prior than meTS. Using the definition of ̆ Σ 0,a and κ b = max a∈A ∥b a ∥ 2 2 = 1, the regret of this agent scales as in Equation (3.12) with σ Ψ = 0, except that the maximum prior variance σ 2 0 is replaced by σ 2 0 + σ 2 Ψ . Hence, BR(T ) = ̃ O (︁ √︂ TdK(σ 2 0 + σ 2 Ψ ) )︁ . When K > L (up to constants), this regret can be substantially larger than the bound for meTS in Equation (3.12). The improvement is on the order of √︁ K/L in regimes where the effects are far more uncertain than the actions, i.e., σ Ψ ≫ σ 0 . For example, in our ad-placement setting, L is the number of catalog items, while K ≈ L M is the number of slates of size M. Thus K/L≈ L M−1 , with typical scales such as L≈ 10 6 and M ≈ 10. Our empirical results in Sections A.4 and 3.4.1 support this: meTS significantly outperforms classical methods when the effect parameters are more uncertain than the action parameters. 3.4 Experiments We evaluate meTS on both synthetic and real-world problems. In each plot, we report the average values and their standard errors. Additional experiments are conducted in 52 Section A.4. The code is provided in this Github repository. 3.4.1 Synthetic Experiments We start with two synthetic problems: the linear and logistic bandit settings in Equa- tions (3.2) and (3.3), respectively. The effect prior is parameterized by μ Ψ =0 Ld and Σ Ψ = 3I Ld , the action covariance is Σ 0,a = I d for all a ∈ A, and the observation noise is σ = 1. We use this setting since modeling of the effect parameters is the most beneficial when they are more uncertain than the action ones (Section 3.3.2). The context X t is sampled uniformly from [−1, 1] d . We run 50 simulations and sample the mixing weights b a,ℓ from [−1, 1] in each run. We consider the following baselines. For the linear setting, we compare meTS-Lin (Sec- tion 3.2.2), LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a) and HierTS (Hong et al., 2022b). For the logistic setting, we compare meTS-GLM (Sec- tion 3.2.3), meTS-Lin (Section 3.2.2), UCB-GLM (Li et al., 2017), GLM-TS (Chapelle and Li, 2012) and HierTS (Hong et al., 2022b). GLM-UCB (Filippi et al., 2010) is excluded be- cause it exhibits very high regret. We also include variational mean-field approximations of meTS (meTS-Lin-Fa and meTS-GLM-Fa), where the full Gaussian effect posterior q t is approximated as q t (Ψ)≈ ∏︁ L ℓ=1 q t,ℓ (ψ ℓ ). This factorization enables sampling each ψ ℓ ∈ R d independently and replaces operations on a full Ld× Ld covariance with blockwise up- dates. This improves the time and space complexities of meTS by L 2 and L, respectively. All baselines but HierTS ignore the structure. HierTS incorporates the structure similarly to meTS-Lin but only has a single effect parameter with prior N (0 d , 3I d ), with the same mean and covariance as the effect parameters of meTS. To compare fairly with LinTS and GLM-TS, their marginal prior mean and covariance are chosen as0 d and ̆ Σ 0,a = Σ 0,a + Γ a Σ Ψ Γ ⊤ a , where Γ a = b ⊤ a ⊗ I d . This is to account for the uncertainty of the effect parameters despite marginalizing them out. In Figure 3.2, we plot the regret in both problems for T = 5000,K = 100,L = 3, and d = 2 (higher values of K up to 100, 000 are tested in our additional experiment in Figure 3.3 below). meTS and its factored variant outperform all baselines that ignore the structure or incorporate it partially. Moreover, meTS-GLM outperforms meTS-Lin in the logistic bandit, which shows the benefit of the approximation in Section 3.2.3. This attests to the generality and flexibility of meTS and the posterior derivations in Section 3.2. We also show in Section A.4.1 that a higher K,L, or d leads to a higher regret due to learning more parameters, which is captured by our regret bounds. 53 010002000300040005000 Round n 0 500 1000 1500 2000 2500 3000 3500 Regret Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinTS LinUCB 010002000300040005000 Round n 0 200 400 600 800 1000 1200 Logistic bandit: K = 100, L = 3, d = 2 meTS-GLM meTS-GLM-Fa meTS-Lin HierTS GLM-TS UCB-GLM Figure 3.2: Evaluation on synthetic problems. In Figure 3.3, we examine how the final cumulative regret scales with the number of actions K in the linear bandit setting, fixing L = 3 and d = 2. As K increases from 100 to 100, 000, meTS-Lin consistently achieves substantially lower regret than LinTS, with the gap widening as K grows. This demonstrates that meTS-Lin scales more favorably with the action space size by leveraging the shared effect structure, which aligns with our theoretical analysis. When K ≫ L, this structural advantage becomes increasingly pronounced. 1005001000500050000100000 Number of actions K 0 2000 4000 6000 8000 10000 Final cumulative regret Varying K, L=3, d=2 meTS-Lin LinTS Figure 3.3: Final cumulative regret as a function of the number of actions K in the linear bandit setting with L = 3 and d = 2. 3.4.2 MovieLens Experiments We study the problem of movie recommendation using the MovieLens 1M dataset (Lam and Herlocker, 2016). This dataset contains one million ratings given by 6040 users to 3952 movies. We apply low-rank factorization to the rating matrix to obtain 5-dimensional representations: x j ∈ R 5 for user j ∈ [6040] and θ a ∈ R 5 for movie a ∈ [3952]. We use 54 the movies as actions and the context X t is sampled uniformly from user vectors x j . We consider both linear and logistic rewards. Given a user x j , the linear reward for movie θ a is sampled from N (x ⊤ j θ a ,σ 2 ) while the logistic reward is sampled from Ber(g(x ⊤ j θ a )), where g is the sigmoid function. We run 50 simulations with K = 100 randomly sampled movies in each run. We compare meTS to most baselines in Section 3.4.1. We do not include UCB-GLM and GLM-UCB because their regret is very high. In LinTS and GLM-TS, the prior mean of action a is μ and its covariance is ̆ Σ 0 = diag(v)∈ R d×d , where μ∈ R d and v ∈ R d are the mean and variance of the movie vectors along all dimensions, respectively. The mixed-effect structure in Equations (3.2) and (3.3) is not available in this problem. Therefore, we use the approach in Section 3.1.3 to learn it. More precisely, we cluster the movies into L = 5 mixture components by training a GMM on the offline action vectors θ a (Section 3.1.3). Each cluster center corresponds to an effect parameter mean μ ψ ℓ ∈ R d and the mixing weight b a,ℓ is the probability that movie a belongs to cluster ℓ, as given by the GMM. We set the effect prior covariance as Σ Ψ = 0.75 diag(( ̆ Σ 0 ) ℓ∈[L] ) ∈ R Ld×Ld and the prior covariance of action a as Σ 0,a = 0.25 ̆ Σ 0 ∈ R d×d , where ̆ Σ 0 is the same as in both LinTS and GLM-TS. This means that the marginal covariance of action a in meTS is 0.25 ̆ Σ 0 + 0.75 Γ a Σ Ψ Γ ⊤ a , where Γ a = b ⊤ a ⊗ I d . Therefore, it is on the same order as ̆ Σ 0,a when ∥b a ∥ 2 2 ≈ 1, and meTS is parameterized comparably to LinTS and GLM-TS. At the same time, we also model that the effect parameters are more uncertain than the action ones, since Σ 0,a = 0.25 ̆ Σ 0 while Σ Ψ = 0.75 diag(( ̆ Σ 0 ) ℓ∈[L] ). In Figure 3.4, we plot the regret for T = 5000 rounds. We observe that meTS has the lowest regret, even if the true rewards are not generated from a mixed-effect model. This shows the robustness of meTS to model misspecification, which we further validate in Section A.4.3. It also highlights the flexibility of our framework, where a proxy structure is learned from offline data. 010002000300040005000 Round n 0 200 400 600 800 1000 1200 1400 Regret Linear bandit: K = 100, L = 5, d = 5 meTS-Lin meTS-Lin-Fa HierTS LinTS 010002000300040005000 Round n 0 50 100 150 200 250 300 350 Logistic bandit: K = 100, L = 5, d = 5 meTS-GLM meTS-GLM-Fa meTS-Lin HierTS GLM-TS Figure 3.4: Evaluation on MovieLens problems. 3.5 Conclusion In this chapter, we introduced a mixed-effect bandit framework based on a two-level graphical model in which each action may depend on multiple underlying effects. This 55 structure enables more efficient exploration, and we designed meTS to leverage it effectively. When implemented as analyzed, meTS performs strongly on both synthetic and real-world benchmarks. Although our presentation focused on the concrete models in Equations (3.2) and (3.3), the underlying algorithmic ideas extend seamlessly to the general mixed-effect model in Equation (3.1). The methodological and theoretical tools developed here lay the groundwork for richer formulations, one of which we explore in detail in the next chapter. Our work has several limitations. First, the regret analysis assumes a well-specified prior: the true parameters must be generated from the same hierarchical model used by meTS. While some experiments suggest robustness to misspecification, formal guarantees un- der prior mismatch remain open. Second, closed-form posteriors are available only for linear-Gaussian rewards; generalized linear models require Laplace approximations (Sec- tion 3.2.3), which are not analyzed theoretically. Third, the mixed-effect structure must be known or learned offline, adding an additional modeling step. Finally, meTS models only two-level hierarchies; deeper latent structures, which may better capture complex action correlations, require the diffusion-based approach developed in Chapter 4. 56 Chapter 4 Scaling Thompson Sampling with Diffusion Models Contents 3.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 3.1.1 Mixed-Effect Linear Bandit . . . . . . . . . . . . . . . . . . . 45 3.1.2 Mixed-Effect Generalized Linear Bandit . . . . . . . . . . . . 46 3.1.3 Structure Learning . . . . . . . . . . . . . . . . . . . . . . . . 46 3.2 Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 46 3.2.1 Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . . 48 3.2.2 Mixed-Effect Linear Bandit . . . . . . . . . . . . . . . . . . . 48 3.2.3 Mixed-Effect Generalized Linear Bandit . . . . . . . . . . . . 49 3.2.4 Computational Complexity . . . . . . . . . . . . . . . . . . . . 50 3.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3.3.1 Main Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 3.3.2 Benefits of Structure . . . . . . . . . . . . . . . . . . . . . . . 52 3.4 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 52 3.4.1 Synthetic Experiments . . . . . . . . . . . . . . . . . . . . . . 53 3.4.2 MovieLens Experiments . . . . . . . . . . . . . . . . . . . . . 54 3.5 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 55 In the previous chapter, we explored how action correlations can be captured using mixed- effects models, in which actions share a set of effect parameters. This approach proved effective when the underlying structure, such as categories in movie recommendation or components in drug design, can be learned through clustering. However, real-world action correlations might exhibit more complex patterns. Thus, this chapter presents an alter- native approach inspired by the remarkable success of diffusion models in approximating complex distributions (Sohl-Dickstein et al., 2015; Ho et al., 2020; Dhariwal and Nichol, 57 2021; Rombach et al., 2022). Rather than explicitly modeling shared effects, we leverage pre-trained diffusion models to capture the rich structure of action parameters and use them as priors in contextual Thompson sampling. We make the following contributions. 1) We introduce a framework for contextual ban- dits with diffusion-derived priors and develop diffusion Thompson sampling (dTS), which is both statistically efficient and computationally tractable. dTS enables fast posterior updates and sampling via an efficient approximation inspired by exact Gaussian pos- teriors. 2) Beyond applying pre-trained diffusion models to contextual bandits, a key contribution is enabling efficient posterior computation and sampling for a d-dimensional parameter θ | D under a diffusion model prior, without updating the diffusion model parameters (i.e., without backpropagating through the neural network). This is relevant not only to bandits and RL but also to broader applications (Chung et al., 2022). Our approximations are motivated by exact closed-form solutions available when the diffusion model is fully linear; these solutions form the basis for our nonlinear approximations, which achieve strong empirical performance while avoiding the computational burden of standard approximate posterior sampling techniques. 4.1 Setting We consider the contextual bandit setting in Section 2.1. Then, we define the prior distribution using a diffusion model, with a set of L consecutive unknown latent parameters ψ ℓ ∈ R d for ℓ ∈ [L]. Precisely, the action parameter θ a depends on the 1-st latent parameter ψ L as θ a | ψ 1 ∼ N (f 1 (ψ 1 ), Σ 1 ), where the link function f 1 and covariance Σ 1 are known. Also, the ℓ− 1-th latent parameter ψ ℓ−1 depends on the ℓ-th latent parameter ψ ℓ as ψ ℓ−1 | ψ ℓ ∼ N (f ℓ (ψ ℓ ), Σ ℓ ), where f ℓ and Σ ℓ are known. Finally, the L-th latent parameter ψ L is sampled as ψ L ∼N (0, Σ L+1 ), where Σ L+1 is known. We summarize this model in Equation (4.1) below: ψ L ∼N (0, Σ L+1 ),(4.1) ψ ℓ−1 | ψ ℓ ∼N (f ℓ (ψ ℓ ), Σ ℓ ),∀ℓ∈ [L]/1, θ a | ψ 1 ∼N (f 1 (ψ 1 ), Σ 1 ),∀a∈A, R t | θ, (ψ ℓ ) ℓ∈[L] ,X t ,A t ∼ p(·| X t ;θ A t ),∀t∈ [T ]. In practice, this model can be built by pre-training a diffusion model on offline estimates of the action parameters θ a . Remark 5 (Joint models). Our algorithm and analysis also apply to the case where all actions share a single unknown parameter θ ∈ R d . Let φ : X × [K] → R d be a known feature map, and assume the reward distribution mean is g (︁ φ(x,a) ⊤ θ )︁ . Then, the diffusion prior in Equation (4.1) specializes by replacing the per-action parameters (θ a ) a∈[K] with a single shared parameter θ: ψ L ∼N (0, Σ L+1 ),(4.2) ψ ℓ−1 | ψ ℓ ∼N (f ℓ (ψ ℓ ), Σ ℓ ),∀ℓ∈ [L]\1, θ | ψ 1 ∼N (f 1 (ψ 1 ), Σ 1 ), R t | θ, (ψ ℓ ) ℓ∈[L] ,X t ,A t ∼ p (︁ · ⃓ ⃓ φ(X t ,A t ) ⊤ θ )︁ ,∀t∈ [T ]. 58 This formulation is useful when a shared feature map φ is available. In that case, the diffusion model can be pre-trained on parameters θ s S s=1 from previous tasks, and dTS can then be applied to a new task S+1 using the pre-trained prior. To avoid clutter, our main exposition focuses on the model in Equation (4.1), but all theoretical results and algorithmic components extend naturally to this shared-parameter case, which we also include in some experiments (explicitly noted when applicable). 4.2 Algorithm We design a Thompson sampling algorithm that samples the latent and action parameters hierarchically (Lindley and Smith, 1972). Let H t = (X i ,A i ,R i ) i∈[t−1] denote the history of all interactions up to round t, and let H t,a = (X i ,A i ,R i ) i∈[t−1];A i =a be the history of interactions with action a up to round t. To motivate our algorithm, we decompose the posterior density p(θ a | H t ) recursively as p(θ a | H t ) = ∫︂ ψ 1:L p(ψ L | H t ) L ∏︂ ℓ=2 p(ψ ℓ−1 | ψ ℓ ,H t )p(θ a | ψ 1 ,H t,a ) dψ 1:L .(4.3) Hierarchical sampling. This decomposition induces the following sampling procedure. First, draw a sample ψ t,L according to the posterior density p(ψ L | H t ). Then, for each ℓ∈ [L]\1, draw ψ t,ℓ−1 from the conditional posterior p(ψ ℓ−1 | ψ t,ℓ ,H t ). Finally, given ψ t,1 , draw each action parameter independently from p(θ a | ψ t,1 ,H t,a ) (the θ a are conditionally independent given ψ 1 ). This defines Algorithm 2, diffusion Thompson Sampling (dTS). Posterior components via recursion. To implement dTS, we provide a recursive scheme to express the required posteriors using known quantities. These expressions may not always admit closed forms and often require approximation. The conditional action- posterior can be written as p(θ a | ψ 1 ,H t,a )∝ ∏︂ i∈S t,a p(R i | X i ;θ a )N (θ a ;f 1 (ψ 1 ), Σ 1 ),(4.4) where S t,a =ℓ∈ [t− 1] : A ℓ = a is the set of rounds in which action a was selected. Now, we characterize the conditional latent-posteriors. Before we do so, we make the following notation clarification. With slight abuse of notation, p(H t | ψ ℓ ) denotes the likelihood of the observations up to round t given ψ ℓ : p(H t | ψ ℓ ) = p((R i ) i<t | (X i ) i<t , (A i ) i<t ,ψ ℓ ) With this notation in mind, for any ℓ∈ [L]\1, the conditional latent-posterior is p(ψ ℓ−1 | ψ ℓ ,H t )∝ p(H t | ψ ℓ−1 )N (ψ ℓ−1 ;f ℓ (ψ ℓ ), Σ ℓ ), and the top-layer posterior is p(ψ L | H t )∝ p(H t | ψ L )N (ψ L ; 0, Σ L+1 ). 59 All terms above are known except the likelihoods p(H t | ψ ℓ ), which are computed recur- sively. The recursion starts with p(H t | ψ 1 ) = K ∏︂ a=1 ∫︂ θ a [︄ ∏︂ i∈S t,a p(R i | X i ;θ a ) ]︄ N (θ a ;f 1 (ψ 1 ), Σ 1 ) dθ a ,(4.5) and for ℓ∈ [L]\1, proceeds as p(H t | ψ ℓ ) = ∫︂ ψ ℓ−1 p(H t | ψ ℓ−1 )N (ψ ℓ−1 ;f ℓ (ψ ℓ ), Σ ℓ ) dψ ℓ−1 .(4.6) Algorithm 2 dTS: diffusion Thompson Sampling Input: Prior components f ℓ , Σ ℓ L+1 ℓ=1 and reward model p. for t = 1,...,T do Draw ψ t,L according to the posterior density p(ψ L | H t ) for ℓ = L,..., 2 do Draw ψ t,ℓ−1 according to p(ψ ℓ−1 | ψ t,ℓ ,H t ) for a = 1,...,K do Draw θ t,a according to p(θ a | ψ t,1 ,H t,a ) Select action A t = argmax a∈[K] r(X t ,a;θ t ), where θ t = (θ t,a ) a∈[K] Observe reward R t ∼ p(·| X t ;θ ∗,A t ) and update the posteriors. All posterior expressions above use known quantities (f ℓ , Σ ℓ ,p(r | x;θ)). However, these expressions typically need to be approximated, except when the link functions f ℓ are linear and the reward distribution p(·| x;θ) is linear-Gaussian, where closed-form solutions can be obtained with careful derivations. These approximations are not trivial, and prior studies often rely on computationally intensive approximate sampling algorithms. In the following sections, we explain how we derive our efficient approximations which are motivated by the closed-form solutions of linear instances. 4.2.1 Posterior Approximation The reward distribution is parameterized as a generalized linear model (GLM) (McCullagh and Nelder, 1989), which allows for non-linear rewards. In addition, the diffusion model itself is highly non-linear due to the link functions f ℓ . These two sources of non-linearity make the posterior intractable, so we apply two layers of approximation: (i) a likelihood approximation to linearize the reward model, and (i) a diffusion approximation to handle the non-linear hierarchy induced by the diffusion model prior. (i) Likelihood approximation. We use an approach similar to the Laplace approxima- tion, but instead of approximating the entire posterior, we approximate only the likelihood by a Gaussian. Precisely, the reward distribution p(· | x;θ a ) belongs to the exponential family with mean function g. Thus ∏︂ i∈S t,a p(R i | X i ;θ a ) ≈ N (︁ θ a ; ˆ B t,a , ˆ G −1 t,a )︁ ,(4.7) 60 where ˆ B t,a is the maximum likelihood estimate and ˆ G t,a is the Hessian of the negative log-likelihood: ˆ B t,a = argmax θ a ∈R d ∑︂ i∈S t,a logp(R i | X i ;θ a ), ˆ G t,a = ∑︂ i∈S t,a ̇g (︁ X ⊤ i ˆ B t,a )︁ X i X ⊤ i ,(4.8) and S t,a = ℓ ∈ [t− 1] : A ℓ = a is the set of rounds in which action a was selected. Of course, ˆ G t,a might not be invertivle and thus we replace it by ˆ G t,a + 10 −3 I d in practice. Unlike Laplace, which fits a global Gaussian to the full posterior, this step linearizes only the likelihood, thereby preserving the hierarchical diffusion structure of the prior. (i) Diffusion approximation. Plugging the Gaussian likelihood approximation (4.7) into the posterior expressions p(θ a | ψ 1 ,H t,a ) and p(ψ ℓ−1 | ψ ℓ ,H t ) removes the non-linearity of the reward model. However, the diffusion hierarchy remains non-linear through f ℓ . To handle this, we build on the closed-form posteriors of the linear diffusion case (where f ℓ (ψ ℓ ) = W ℓ ψ ℓ ; see Section B.1) and generalize them by replacing the linear terms W ℓ ψ ℓ with their non-linear counterparts f ℓ (ψ ℓ ). This substitution yields a posterior diffusion model that retains the same hierarchical form as the prior but with data-dependent means and covariances for the conditional Gaussians. Details on how we transition from the linear to the general non-linear setting are provided in Sections B.1 and B.2. The resulting approximate posteriors admit the following closed-form expressions. Approximate action posterior. We approximate the conditional action posterior as p(θ a | ψ 1 ,H t,a ) ≈ N (︁ θ a ; ˆμ t,a , ˆ Σ t,a )︁ , where ˆ Σ −1 t,a =Σ −1 1 ⏞⏟⏞ prior precision + ˆ G t,a ⏞⏟⏞ data precision ,ˆμ t,a = ˆ Σ t,a (︂ Σ −1 1 f 1 (ψ 1 ) ⏞ ⏟⏞ prior contribution + ˆ G t,a ˆ B t,a ⏞ ⏟⏞ data contribution )︂ . (4.9) This posterior update has a clear interpretation. The posterior precision ˆ Σ −1 t,a is the sum of the prior precision and the data precision. The posterior mean ˆμ t,a is the precision- weighted average of the prior mean and the MLE ˆ B t,a . As more data are observed, the covariance shrinks and the mean moves from the prior mean f 1 (ψ 1 ) toward the MLE ˆ B t,a . When no data are available ( ˆ G t,a = 0), the posterior reduces to the prior N (f 1 (ψ 1 ), Σ 1 ); in the limit of infinite data ( ˆ G t,a → ∞), the posterior collapses to the MLE ˆ B t,a , with ˆμ t,a → ˆ B t,a and ˆ Σ t,a → 0. Approximate latent posteriors. For each ℓ∈ [L + 1]\1, we approximate the latent posterior as p(ψ ℓ−1 | ψ ℓ ,H t ) ≈ N (︁ ψ ℓ−1 ; ̄μ t,ℓ−1 , ̄ Σ t,ℓ−1 )︁ , with ̄ Σ −1 t,ℓ−1 =Σ −1 ℓ ⏞⏟⏞ prior precision + ̄ G t,ℓ−1 ⏞ ⏟⏞ data precision , ̄μ t,ℓ−1 = ̄ Σ t,ℓ−1 (︂ Σ −1 ℓ f ℓ (ψ ℓ ) ⏞ ⏟⏞ prior contribution + ̄ B t,ℓ−1 ⏞⏟⏞ data contribution )︂ , (4.10) 61 where, by convention, f L+1 (ψ L+1 ) = 0 since the top layer ψ L has no parent ψ L+1 . The quantities ̄ G t,ℓ and ̄ B t,ℓ are computed recursively. The base recursion is ̄ G t,1 = K ∑︂ a=1 (︁ Σ −1 1 − Σ −1 1 ˆ Σ t,a Σ −1 1 )︁ , ̄ B t,1 = Σ −1 1 K ∑︂ a=1 ˆ Σ t,a ˆ G t,a ˆ B t,a ,(4.11) and for each ℓ∈ [L]\1, ̄ G t,ℓ = Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ , ̄ B t,ℓ = Σ −1 ℓ ̄ Σ t,ℓ−1 ̄ B t,ℓ−1 .(4.12) The latent posterior update in Equation (4.10) has the same structure as the action posterior. The posterior precision ̄ Σ −1 t,ℓ−1 is the sum of the prior and data precisions , and the posterior mean is their precision-weighted combination. The data terms ̄ G t,ℓ−1 and ̄ B t,ℓ−1 are computed recursively (Equations (4.11) and (4.12)), so information collected at the action level propagates upward through the hierarchy. Interpretation. The resulting approximate posterior remains a diffusion model whose conditional Gaussians have updated, data-dependent means and covariances. The latent- posterior means can be viewed as refined link functions: ˆ f t,ℓ (ψ ℓ ) = ̄μ t,ℓ−1 = ̄ Σ t,ℓ−1 (︁ Σ −1 ℓ f ℓ (ψ ℓ ) + ̄ B t,ℓ−1 )︁ , and ̄ Σ t,ℓ represents their updated uncertainty. Both are updated with data: covariances contract as uncertainty decreases, and means move from the prior toward the MLE. Unlike a full Laplace approximation, this formulation preserves the expressiveness of the posterior rather than replacing it globally with a single Gaussian, while also avoiding the heavy computation required by other approximate inference methods. 4.2.2 Extension to Joint Reward Models For the shared-parameter model in Remark 5, dTS’s posterior approximations are similar. The action posterior is p(θ | ψ 1 ,H t )≈N (ˆμ t , ˆ Σ t ), where ˆ Σ −1 t = Σ −1 1 + ˆ G t ,ˆμ t = ˆ Σ t (︁ Σ −1 1 f 1 (ψ 1 ) + ˆ G t ˆ B t )︁ .(4.13) where ˆ B t = argmax θ∈R d ∑︂ i<t logp (︁ R i | φ(X i ,A i ) ⊤ θ )︁ , ˆ G t = ∑︂ i<t ̇g (︁ φ(X i ,A i ) ⊤ ˆ B t )︁ φ(X i ,A i )φ(X i ,A i ) ⊤ . Similarly, for ℓ∈ [L + 1]\1, the latent posterior is p(ψ ℓ−1 | ψ ℓ ,H t )≈N ( ̄μ t,ℓ−1 , ̄ Σ t,ℓ−1 ), where ̄ Σ −1 t,ℓ−1 = Σ −1 ℓ + ̄ G t,ℓ−1 , ̄μ t,ℓ−1 = ̄ Σ t,ℓ−1 (︁ Σ −1 ℓ f ℓ (ψ ℓ ) + ̄ B t,ℓ−1 )︁ ,(4.14) where, by convention, f L+1 (ψ L+1 ) = 0 and the quantities ̄ G t,ℓ and ̄ B t,ℓ are computed recursively as Base case: ̄ G t,1 = Σ −1 1 − Σ −1 1 ˆ Σ t Σ −1 1 , ̄ B t,1 = Σ −1 1 ˆ Σ t ˆ G t ˆ B t . (4.15) Recursive case: ̄ G t,ℓ = Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ , ̄ B t,ℓ = Σ −1 ℓ ̄ Σ t,ℓ−1 ̄ B t,ℓ−1 . (4.16) Again, this shared-parameter variant of dTS is presented for completeness and to illustrate the generality of our posterior derivations; the main focus of the chapter remains on the per-action disjoint formulation in Equation (4.1). Unless stated otherwise, all theoretical results and experiments use the main version of dTS described in Algorithm 2. 62 4.3 Analysis In this section, we present an informal Bayes regret analysis of dTS to build intuition around dTS’s Bayesian regret scaling with problem parameters d, K, L, etc. This anal- ysis is informal for two reasons. First, we analyze a simplified linear-Gaussian setting rather than the general nonlinear case on which we focus in this chapter: the reward distribution is linear-Gaussian and each link function f ℓ (ψ ℓ ) = W ℓ ψ ℓ is a known linear mapping, inducing a hierarchy of L linear-Gaussian layers from the latent root to the action parameters. Second, we assume the model is well-specified (similar to Chapter 7): the true action parameters are generated according to the diffusion prior used by dTS. Under these assumptions, the posterior becomes exact, enabling an analysis analogous to that used in Chapter 3. However, our recursive hierarchical structure introduces technical differences: posteriors must be derived inductively using total covariance decompositions, and regret bounds require tracking information flow across all latent layers. We emphasize that this regret bound does not extend to the general nonlinear case studied in experiments; it is included here solely to provide theoretical intuition under simplifying assumptions. Formal statements and derivations are provided in Sections B.4 and B.5. Bayes regret bound. The bound of dTS in this case is BR(T ) = ̃ O (︂ ⌜ ⃓ ⃓ ⎷ T (dKσ 2 1 + d L ∑︂ ℓ=1 σ 2 ℓ+1 σ 2ℓ max ) )︂ , where σ 2 max = max ℓ∈[L+1] 1+ σ 2 ℓ σ 2 . This can be re-written asBR(T ) = ̃ O( √︂ TdK eff ∑︁ L+1 ℓ=1 σ 2 ℓ ), where K eff = Kσ 2 1 + ∑︁ L ℓ=1 σ 2 ℓ+1 σ 2ℓ max ∑︁ L+1 ℓ=1 σ 2 ℓ is the effective number of actions. This dependence on the horizon T aligns with prior Bayes regret bounds scaling with T. However, the bound comprises L + 1 main terms. First, one relates to action parameters learning, conforming to a standard form (Lu and Van Roy, 2019), while the L remaining terms are associated with learning each of the latent parameters. Sparsity refinement. If each mixing matrix exhibits column sparsity, that, W ℓ = ( ̄ W ℓ , 0 d,d−d ℓ ) with d ℓ ≪ d active columns, then the bound becomes BR(T ) = ̃ O (︂ ⌜ ⃓ ⃓ ⎷ T (dKσ 2 1 + L ∑︂ ℓ=1 d ℓ σ 2 ℓ+1 σ 2ℓ max ) )︂ . Hence, informative, sparse priors can cut the cost of learning deep latent chains down from d to d ℓ . As in Chapter 3, a less informative prior (such as high variance) leads to a more challenging problem and thus a higher bound. Therefore, smaller values of K, L, d, d ℓ translate to fewer parameters to learn, leading to lower regret. The regret also decreases when the initial variances σ 2 ℓ decrease. These dependencies are common in Bayesian analysis, and empirical results match them. Dependence on K. The reader may question why our bound depends on K. This depen- dence arises from two modeling choices. First, we study the disjoint (per-action) setting 63 r(x,a;θ) = x ⊤ θ a , where θ = (θ a ) a∈[K] ∈ R dK , requiring the learning of Kd parameters. Second, we model the relationship between θ a and ψ 1 stochastically as N (W 1 ψ 1 ,σ 2 1 I d ) to accommodate potential nonlinearity. While this choice confers robustness to model mis- specification, it introduces additional uncertainty and requires learning both the action parameters θ a and the latent parameters ψ ℓ , resulting in a bound that depends on both K and L. Despite this dependence, dTS enjoys two key advantages. First, the regret scales with Kσ 2 1 rather than K ∑︁ ℓ σ 2 ℓ , which is particularly beneficial when σ 1 is small, as is often the case with diffusion model priors. Second, thanks to informative priors, our bound has significantly smaller constants compared to both the Bayesian and frequentist regret bounds for LinTS. We demonstrate this empirically in Section B.6.5 and provide a the- oretical comparison in Section 4.3.1. Both analyses confirm that dTS’s advantage over LinTS increases as the action space grows. Can regret be independent of K? Prior works (Foster et al., 2020; Xu and Zeevi, 2020; Zhu et al., 2022) have proposed bandit algorithms whose regret does not scale with K. However, these results apply to the shared-parameter setting r(x,a;θ) = φ(x,a) ⊤ θ, where only a single d-dimensional parameter must be learned, but this formulation requires access to a suitable feature map φ. dTS is compatible with this setting (Section 4.2.2), in which case its regret would indeed be independent of K. Alternatively, even in the disjoint per-action case considered in this chapter, setting σ 1 = 0 would yield a K-independent regret bound. However, we believe this assumption is unrealistic in practice and would compromise the robustness of dTS to model misspecification. 4.3.1 Benefits Computational benefits. Action correlations prompt an intuitive approach: marginal- ize all latent parameters and maintain a joint posterior of (θ a ) a∈[K] | H t . Unfortunately, this is computationally inefficient for large action spaces. To illustrate, suppose that all posteriors are multivariate Gaussians. Then maintaining the joint posterior (θ a ) a∈[K] | H t necessitates converting and storing its dK × dK-dimensional covariance matrix, leading to O(K 3 d 3 ) and O(K 2 d 2 ) time and space complexities. In contrast, the time and space complexities of dTS are O (︁(︁ L + K )︁ d 3 )︁ and O (︁(︁ L + K )︁ d 2 )︁ . This is because dTS requires converting and storing L + K covariance matrices, each being d× d-dimensional. The improvement is huge when K ≫ L, which is common in practice. Certainly, a more straightforward way to enhance computational efficiency is to discard latent parameters and maintain K individual posteriors, each relating to an action parameter θ a ∈ R d (LinTS). This improves time and space complexity to O (︁ Kd 3 )︁ and O (︁ Kd 2 )︁ . However, LinTS maintains independent posteriors and fails to capture the correlations among ac- tions; it only models θ a | H t,a rather than θ a | H t as done by dTS. Consequently, LinTS incurs higher regret due to the information loss caused by unused interactions of similar actions. Our regret bound and empirical results reflect this aspect. Statistical benefits. We argue that our bound reflects the overall structure of the problem by comparing dTS to algorithms that only partially use the structure or do not use it at all as follows. Precisely, when the link functions are linear, we can transform the diffusion prior into a Bayesian linear model (LinTS) by marginalizing out the latent 64 parameters; in which case the prior on action parameters becomes θ a ∼ N (0, Σ), with the θ a being not necessarily independent, and Σ is the marginal initial covariance of action parameters and it writes Σ = σ 2 1 I d + ∑︁ L ℓ=1 σ 2 ℓ+1 B ℓ B ⊤ ℓ with B ℓ = ∏︁ ℓ i=1 W i . Then, it is tempting to directly apply LinTS to solve our problem. This approach will induce higher regret because the additional uncertainty of the latent parameters is accounted for in Σ despite integrating them. This causes the marginal action uncertainty Σ to be much higher than the conditional action uncertainty σ 2 1 I d , since we have Σ = σ 2 1 I d + ∑︁ L ℓ=1 σ 2 ℓ+1 B ℓ B ⊤ ℓ ≽ σ 2 1 I d . This discrepancy leads to higher regret, especially when K is large. This is due to LinTS needing to learn K independent d-dimensional parameters, each with a considerably higher initial covariance Σ. This is also reflected by our regret bound. To simply comparisons, suppose that σ ≥ max ℓ∈[L+1] σ ℓ so that σ 2 max ≤ 2. Then the regret bounds of dTS (where we bound σ 2ℓ max by 2 ℓ ) and LinTS read dTS : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ T (dKσ 2 1 + L ∑︂ ℓ=1 d ℓ σ 2 ℓ+1 2 ℓ ) )︁ , LinTS : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ TdK(σ 2 1 + L ∑︂ ℓ=1 σ 2 ℓ+1 ) )︁ . Then regret improvements are captured by the variances σ ℓ and the sparsity dimensions d ℓ , and we proceed to illustrate this through the following scenarios. (I) Decreasing variances. Assume that σ ℓ = 2 ℓ for any ℓ∈ [L + 1]. Then, the regrets become dTS : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ T (dK + L ∑︂ ℓ=1 d ℓ 4 ℓ )) )︁ , LinTS : ̃ O (︁ √︁ TdK2 L ) )︁ Now to see the order of gain, assume the problem is high-dimensional (d ≫ 1), and set L = log 2 (d) and d ℓ =⌊ d 2 ℓ ⌋. Then the regret of dTS becomes ̃ O (︁ √︁ nd(K + L)) )︁ , and hence the multiplicative factor 2 L in LinTS is removed and replaced with a smaller additive factor L. (I) Constant variances. Assume that σ ℓ = 1 for any ℓ ∈ [L + 1]. Then, the regrets become dTS : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ T (dK + L ∑︂ ℓ=1 d ℓ 2 ℓ )) )︁ , LinTS : ̃ O (︁ √︁ TdKL) )︁ Similarly, let L = log 2 (d), and d ℓ = ⌊ d 2 ℓ ⌋. Then dTS’s regret is ̃ O (︁ √︁ Td(K + L) )︁ . Thus the multiplicative factor L in LinTS is removed and replaced with the additive factor L. By comparing this to (I), the gain with decreasing variances is greater than with constant ones. In general, diffusion models use decreasing variances (Ho et al., 2020) and hence we expect great gains in practice. All observed improvements in this section could become even more pronounced when employing non-linear diffusion models. In our theory, we used linear diffusion models, and yet we can already discern substantial differences. Moreover, under non-linear diffusion Equation (4.1), the latent parameters cannot be analytically marginalized, making LinTS with exact marginalization inapplicable. 65 010002000300040005000 0 1 2 3 4 5 6 7 8 9 Regret 1e3 Linear diffusion, linear reward K=100, L=2, d=5 dTS-L HierTS LinTS LinUCB 010002000300040005000 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 1e3 Linear diffusion, nonlinear reward K=100, L=2, d=5 dTS-LN dTS-L GLM-TS UCB-GLM 010002000300040005000 0.0 0.2 0.4 0.6 0.8 1.0 1e4 Nonlinear diffusion, linear reward K=100, L=2, d=5 dTS-NL LinTS LinUCB 010002000300040005000 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1e3 Nonlinear diffusion, nonlinear reward K=100, L=2, d=5 dTS-N dTS-NL GLM-TS UCB-GLM 010002000300040005000 Round t2[n] 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 Regret 1e5 K=10000, L=4, d=20 dTS-L HierTS LinTS LinUCB 010002000300040005000 Round t2[n] 0.0 0.5 1.0 1.5 2.0 2.5 3.0 1e3 K=10000, L=4, d=20 dTS-LN dTS-L GLM-TS UCB-GLM 010002000300040005000 Round t2[n] 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 1e7 K=10000, L=4, d=20 dTS-NL LinTS LinUCB 010002000300040005000 Round t2[n] 0 1 2 3 4 5 6 7 8 1e2 K=10000, L=4, d=20 dTS-N dTS-NL GLM-TS UCB-GLM Figure 4.1: Regret of dTS with varying diffusion and reward models and varying parame- ters d, K, L. 4.4 Experiments Experimental setup. We evaluate dTS using both synthetic and MovieLens problems. In our experiments, we run 50 random simulations and plot the average regret with stan- dard error. Our main contribution is to demonstrate that pretraining a diffusion model offline enables the construction of expressive and informative priors that substantially improve exploration efficiency in contextual bandits. We first evaluate dTS in a setting where the prior matches the true generative process (Section 4.4.1 to isolate the benefit of informative priors), and then consider a misspecified regime (Section 4.4.2 and Sec- tion B.6) where the prior is either trained on out-of-distribution data or intentionally perturbed. These experiments show that even when the prior is imperfect, dTS maintains strong performance: highlighting its robustness and practical relevance. 4.4.1 True Prior is a Diffusion Model Synthetic bandit problems are generated from the diffusion model in Equation (4.1) with both linear and non-linear rewards. Linear rewards follow p(·| x;θ a ) =N (x ⊤ θ a , 1), while non-linear rewards are binary from p(· | x;θ a ) = Ber(g(x ⊤ θ a )), with g as the sigmoid function. Covariances are Σ ℓ = I d , and contexts X t are uniformly drawn from [−1, 1] d . We vary d ∈ 5, 20, L ∈ 2, 4, K ∈ 10 2 , 10 4 , and set the horizon to T = 5000, considering both linear and non-linear models. Linear diffusion. We consider Equation (4.1) with f ℓ (ψ) = W ℓ ψ, where W ℓ uniformly drawn from [−1, 1] d×d . Sparsity is introduced by zeroing the last d ℓ columns of W ℓ as W ℓ = ( ̄ W ℓ , 0 d,d−d ℓ ). For d = 5 and L = 2, (d 1 ,d 2 ) = (5, 2); for d = 20 and L = 4, (d 1 ,d 2 ,d 3 ,d 4 ) = (20, 10, 5, 2). Non-linear diffusion. We consider Equation (4.1) where f ℓ are 2-layer neural networks with random weights in [−1, 1], ReLU activation, and hidden layers of size h = 20 for d = 5, and h = 60 for d = 20. 66 Baselines. For linear rewards, we use LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a), and HierTS (Hong et al., 2022b), marginalizing out all latent parameters except ψ L , which corresponds to HierTS-1 in Section B.3. For non- linear rewards, we include UCB-GLM (Li et al., 2017) and GLM-TS (Chapelle and Li, 2012). We exclude GLM-UCB (Filippi et al., 2010) due to high regret and HierTS as it’s designed for linear rewards. We name dTS as dTS-dr, where d refers to diffusion type (L for linear, N for non-linear) and r indicates reward type (L for linear, N for non-linear). For example, dTS-L signifies dTS in linear diffusion with linear rewards. Results and interpretations. Results are shown in Figure 4.1 and we make the follow- ing observations: 1) dTS demonstrates superior performance (Figure 4.1). dTS consistently outper- forms the baselines across all settings, including the four combinations of linear/non-linear diffusion and reward (columns in Figure 4.1) and both bandit settings with varying K, L, and d (rows in Figure 4.1). 2) Latent diffusion structure may be more important than the reward distri- bution. When rewards are non-linear (second and fourth columns in Figure 4.1), we include variants of dTS that use the correct diffusion prior but the wrong reward distribu- tion, applying linear-Gaussian instead of logistic-Bernoulli (dTS-L in the second column and dTS-NL in the fourth). Despite the reward misspecification, these variants outperform models using the correct reward distribution but ignoring the latent diffusion structure, such as GLM-TS and UCB-GLM. This highlights the importance of accounting for latent structure, which can be more critical than an accurate reward distribution. 3) Performance gap between dTS and LinTS widens as K increases (Figure 4.2a). To show dTS’s improved scalability, we evaluate its performance with varying values of K ∈ [10, 5× 10 4 ], in the linear diffusion and rewards setting. Figure 4.2a shows the final cumulative regret for varying K values for both dTS-L and LinTS, revealing a widening performance gap as K increases. 4) Regret scaling with K, d and L matches our theory (Figure 4.2b). We assess the effect of the number of actions K, context dimension d, and diffusion depth L on dTS’s regret. Using the linear diffusion and rewards setting, for which we have derived a Bayes regret upper bound, we plot dTS-L’s regret across varying values of K ∈ 10, 100, 500, 1000, d ∈ 5, 10, 15, 20, and L ∈ 2, 4, 5, 6 in Figure 4.2b. As predicted by our theory, the empirical regret increases with larger values of K, d, or L, as these make the learning problem more challenging, leading to higher regret. 5) Diffusion prior misspecification (Figure 4.2c). Here, dTS’s diffusion prior pa- rameters differ from the true diffusion prior. In the linear diffusion and reward setting, we replace the true parameters W ℓ and Σ ℓ with misspecified ones, W ℓ +ε 1 and Σ ℓ +ε 2 , where ε 1 and ε 2 are uniformly sampled from [v,v +0.5] d×d , with v controlling the misspecification level. We vary v ∈ 0.5, 1, 1.5 and assess dTS’s performance, comparing it to the well- specified dTS-L and the strongest baseline in this fully-linear setting, HierTS. As shown in Figure 4.2c, dTS’s performance decreases with increasing misspecification but remains superior to the baseline, except at v = 1.5, where their performances are comparable. Additional misspecification experiments are presented in Section 4.4.2, where the bandit 67 10100500500050000 Number of actions K 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 1.8 Regret in round n 1e4 Regret as a function of K dTS-L LinTS (a) Perf. gap w.r.t. K. 0.00.51.01.52.02.53.0 Increasing values of either K;d or L 0 500 1000 1500 2000 2500 3000 3500 Regret Effect of K;d;L on the regret of LindTS Increasing K Increasing d Increasing L (b) Scaling w.r.t. K, d, L. 010002000300040005000 Round t2[n] 0 500 1000 1500 2000 2500 Regret Effect of prior misspecification LindTS (v=0.5) LindTS (v=1) LindTS (v=1.5) LindTS HierTS (c) Prior misspecification. Figure 4.2: Effect of various factors on dTS’s performance. 1050100250050001000050000 Number of pre-training samples 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 LinTS regret / dTS regret (a) Ratio of LinTS/dTS cumu- lative regret in the last round with varying pre-training sam- ple size in [10, 5×10 4 ]. Higher values mean a bigger per- formance gap. 2104070100 Difusion depth L 1.0 1.5 2.0 2.5 3.0 3.5 4.0 LinTS regret / dTS regret (b) Ratio of LinTS/dTS cumu- lative regret in the last round with varying diffusion depth L in [2, 100]. Higher values mean a bigger performance gap. 020406080100 Round t2[n] 0 1 2 3 4 5 6 7 8 Cumulative regret dTS LinTS (c) Regret of dTS in Movie- Lens. The diffusion model with L = 40 is pre-trained on embeddings obtained by low- rank factorization of Movie- Lens rating matrix. Figure 4.3: (a) and (b): Impact of pre-training sample size and diffusion depth L for the Swiss roll data. (c): Regret of dTS in MovieLens. environment is not sampled from a diffusion model. 4.4.2 True Prior is Not a Diffusion Model Swiss roll data. Unlike previous experiments, the true action parameters are now sam- pled from the Swiss roll distribution (see Figure B.1 in Section B.6.1), rather than from a diffusion model. The diffusion model used by dTS is pre-trained on samples from this dis- tribution, with the offline pre-training procedure described in Section B.6.2. Figure 4.3a shows that larger sample sizes increase the performance gap between dTS and LinTS. More samples improve the estimation of the diffusion prior (see Figure B.1 in Section B.6.1), leading to better dTS performance. Notably, comparable performance was achieved with as few as 10 samples, and dTS outperformed LinTS by a factor of 1.5 with just 50 sam- ples. While more samples may be required for more complex problems, LinTS would also struggle in such cases. Therefore, we expect these gains to be even more significant in more challenging settings. We studied the effect of the pre-trained diffusion model depth L and found that L ≈ 40 68 yields the best performance, with a drop beyond that point (Figure 4.3b). While our theory doesn’t apply directly here, as it assumes a linear diffusion model, it still offers some intuition on the decreased performance for L > 40. The theorem shows dTS’s regret bound increases with L when the true distribution is a diffusion model. For small L, the pre-trained model doesn’t fully capture the true distribution, making the theorem inapplicable, but at L ≈ 40, the distribution is nearly captured, and further increases in L lead to higher regret, consistent with our theory. MovieLens data. We also evaluate dTS using the standard MovieLens (Lam and Her- locker, 2016) setting. In this semi-synthetic experiment, a user is sampled from the rating matrix in each interaction round, and the reward is the rating the user gives to a movie (see Clavier et al. (2023, Section 5) for details about this setting). Here, the true distri- bution of action parameters is unknown and not a diffusion model. The diffusion model is pre-trained on offline estimates of action parameters obtained through low-rank factor- ization of the rating matrix. Figure 4.3c demonstrates that dTS outperforms LinTS in this setting. Additional CIFAR ablations are provided in Section B.6.4 where similar strong improvements are observed. 4.5 Conclusion We use a pre-trained diffusion model as a strong and flexible prior for dTS. Diffusion pre-training leverages abundant offline data, which is then fine-tuned through online in- teractions via our tractable posterior approximation. This approximation enables efficient posterior sampling and updates while maintaining strong empirical performance. More- over, dTS admits a simple Bayesian regret bound in the linear–Gaussian setting. Our work has several limitations. First, our Bayes regret analysis applies only to the linear- Gaussian setting with a well-specified prior; extending formal guarantees to nonlinear diffusion models remains open. Second, our posterior approximation, while motivated by exact solutions in the linear case, lacks theoretical justification for general nonlinear link functions: its strong empirical performance does not come with formal approximation error bounds. Finally, dTS requires offline pre-training of the diffusion model, which assumes access to historical estimates of action parameters; in domains where such data is unavailable or expensive to obtain, the benefits of diffusion priors may not be realized. 69 Part I Off-Policy Learning in Large Action Spaces 70 Chapter 5 Introduction to Part I This second part of the thesis addresses the following fundamental question: How can we reliably learn high-performing policies from static logged data when the number of actions is large? 5.1 Setting and Background In this part, we consider the off-policy (offline) setting where an agent is provided with a static logged dataset D n =(X i ,A i ,R i ) n i=1 collected by a logging policy π 0 . The data collection process proceeds as follows: for each round i∈ [n]: 1. The environment draws a context X i ∼ ν, where ν is a distribution with supportX forming a compact subset of R d ; 2. The logging policy selects an action A i ∼ π 0 (·| X i ) from the action set A = [K]; 3. The environment generates a stochastic reward R i ∼ p(·| X i ,A i ), where R i ∈ [0, 1]. Unlike the on-policy setting of Part I, no further interaction with the environment is permitted. The objective is to learn a new policy ˆπ from this static dataset that maximizes the true (but unknown) expected value: V (π) = E X∼ν [︁ E A∼π(·|X) [r(X,A)] ]︁ ,(5.1) where r(x,a) = E R∼p(·|x,a) [R] is the expected reward function. Performance is measured by the suboptimality gap of the learned policy: so(ˆπ) = V (π ∗ )− V (ˆπ),(5.2) where π ∗ = arg max π∈Π V (π) is the optimal policy within a class Π. Since V (π) cannot be computed directly, off-policy learning algorithms often rely on an empirical estimate ˆ V (π) constructed fromD n . This estimation task is known as off-policy evaluation (OPE) in the literature. The two dominant estimation approaches are: direct 71 method (DM) and inverse propensity scoring (IPS). DM employ a learned reward model ˆr(x,a) to estimate the value as: ˆ V dm (π) = 1 n n ∑︂ i=1 ∑︂ a∈A π(a| X i ) ˆr(X i ,a).(5.3) IPS re-weights observed rewards using importance sampling as: ˆ V ips (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) R i .(5.4) Given an estimator ˆ V (π), the agent must then select a policy. This step distinguishes between greedy policies, which directly maximize ˆπ = arg max π∈Π ˆ V (π), and pessimistic policies, which incorporate an uncertainty penalty ˆπ = arg max π∈Π [ ˆ V (π)− pen(π)]. 5.1.1 Scalability Challenges When the number of actions K is large, both estimation paradigms and their associated optimization procedures encounter fundamental obstacles: Statistical inefficiency of DM. Standard DM approaches model each action’s reward function independently. As K grows, the data available per action diminishes, leading to poorly estimated reward functions and lower performance. High variance of importance sampling. IPS’s variance grows with the importance weights π(a|x)/π 0 (a|x). These weights explode in large action spaces, producing estimates too noisy for reliable optimization. Intractable optimization landscapes. Beyond estimation challenges, the optimiza- tion problem arg max π∈Π ˆ V (π) itself becomes computationally intractable in large action spaces. IPS-based objectives induce highly non-concave landscapes with exponentially many local maxima and flat plateaus that trap gradient-based optimizers. As we show in Chapter 7, this optimization bottleneck often dominates estimation error, making even statistically superior estimators ineffective in practice. 5.2 Methodological Approaches To address these challenges, the methods developed in this part pursue three complemen- tary strategies: structured reward modeling for sample-efficient DM, surrogate objectives that prioritize optimization tractability over estimation accuracy, and principled regular- ization and pessimism for importance-weighted estimators. Structured Bayesian models. Drawing inspiration from the hierarchical framework of Part I, we introduce latent structure into reward modeling. Action parameters are coupled through shared latent variables ψ: ψ ∼ q(·),(5.5) θ a | ψ ∼ p a (·;f a (ψ)), ∀a∈A, R| X,A,θ,ψ ∼ p(·| X;θ A ). 72 This formulation enables information sharing across actions: observations from frequently selected actions inform the posterior over ψ, which in turn improves reward estimates for rarely observed actions. Optimization-aware objectives. Rather than designing sophisticated value estima- tors and then optimizing them, we propose objectives designed primarily for favorable optimization landscapes. The policy-weighted log-likelihood (PWLL) family: ˆ U g (π) = 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i )) logπ(A i | X i ),(5.6) where g is a positive weighting function, yields concave objectives for linear-softmax policies π. This guarantees efficient convergence to a unique global maximum, bypassing the optimization pathologies of value estimation altogether. Regularized importance weighting. For practitioners committed to IPS-based meth- ods, we develop variance-controlled estimators through importance weight regularization. The exponential smoothing estimator: ˆ V α (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) α R i , α∈ [0, 1],(5.7) smoothly trades variance for bias while preserving differentiability. Combined with pes- simistic optimization and PAC-Bayes generalization bounds, this yields principled, tractable objectives that are amenable to stochastic gradient ascent for safe off-policy learning. 5.3 Roadmap of Part I The following chapters develop these methodological approaches. Chapter 6: Scaling Direct Methods with Latent Parameters. We begin by ad- dressing the statistical inefficiency of DM through structured Bayesian modeling. Building on the hierarchical framework of Part I, we introduce the structured direct method (sDM), which couples action parameters through a shared latent vector. The posterior over these latent variables aggregates evidence across all actions, enabling effective generalization to actions with sparse data coverage. We analyze performance through Bayesian sub- optimality and prove that greedy policies paired with sDM achieve O(1/ √ n) convergence under mild assumptions on the alignment between logging and optimal policies. Chapter 7: Optimization Matters More Than Estimation. We then challenge the conventional paradigm of off-policy learning. Through theoretical analysis and large- scale experiments, we demonstrate that optimization error dominates estimation error in large action spaces. Specifically, we prove that for any IPS-based estimator, gradient ascent can remain trapped in suboptimal regions for O(K) iterations, and that the op- timization landscape contains exponentially many local maxima in K. We then propose objective-aware policy parametrizations: by aligning the policy class with the estimator’s inductive bias, we can partially mitigate these optimization challenges. However, for a 73 more complete solution, we propose policy-weighted log-likelihood (PWLL) objectives as an alternative to IPS-based objectives. These objectives are provably concave for linear softmax policies, guaranteeing efficient convergence to a global optimum. Experiments on datasets with up to one million actions validate that PWLL consistently outperforms state-of-the-art estimator-based methods. Chapter 8: Principled Pessimism for Exponential Smoothing and Beyond. Finally, for practitioners committed to importance weighting methods, we develop a the- oretically grounded framework for variance control and pessimistic policy learning. We propose exponential smoothing estimators that regularize importance weights, trading controlled bias for reduced variance. To leverage these regularized estimators for safe pol- icy learning, we derive two-sided PAC-Bayes generalization bounds where all quantities are empirical and differentiable. The pessimistic learning objective maximizes the lower bound on policy value, penalizing policies with high bias or variance and steering optimiza- tion toward reliable regions. This also yields tractable objectives amenable to stochastic gradient optimization. We further present a unified PAC-Bayes framework covering the major importance weight regularization techniques in the literature (clipping, exponential smoothing, implicit exploration), enabling principled comparison and demonstrating that the choice of pessimistic objective often matters more than the specific regularizer. 74 Chapter 6 Scaling Direct Methods with Latent Parameters Contents 5.1 Setting and Background . . . . . . . . . . . . . . . . . . . . . . . . . . 71 5.1.1 Scalability Challenges . . . . . . . . . . . . . . . . . . . . . . . 72 5.2 Methodological Approaches . . . . . . . . . . . . . . . . . . . . . . . . 72 5.3 Roadmap of Part I . . . . . . . . . . . . . . . . . . . . . . . . . . . . 73 In this chapter, we address the statistical inefficiency of direct methods (DMs) in large action spaces: the first challenge highlighted in Chapter 5. Standard DMs estimate inde- pendent d-dimensional parameters for each action. This approach becomes statistically inefficient in large action spaces, where data are collected by a logging policy that explores only a small subset of available actions, leaving many actions rarely or never observed. To overcome this limitation, we analyze DMs through a Bayesian lens and propose making them sample-efficient by incorporating informative priors. We make the following contributions. 1) We introduce the structured direct method (sDM), a Bayesian approach that uses informative priors to share reward information across ac- tions. By updating beliefs about similar actions based on observed data, sDM improves statistical efficiency without compromising computational scalability. 2) To evaluate sDM, we propose Bayesian metrics that assess the average performance across problem instances sampled from the prior. This departs from the standard frequentist focus on worst-case scenarios. These metrics formally quantify the benefits of informative priors. 3) Our theoretical analysis of Bayesian suboptimality (BSO) reveals two key insights: (1) perfor- mance degrades gracefully even without the standard assumption of full logging support, and (2) greedy policies are provably optimal under the BSO metric, standing in contrast to the pessimistic policies typically favored in frequentist settings. 4) We empirically validate sDM and our theoretical findings using both synthetic and real-world datasets. 75 6.1 Setting We consider the setting described in Section 5.1. The only additional assumption is the existence of unknown true parameters θ ∗,a ∈ R d for each action a, such that rewards are distributed as R i ∼ p(·| X i ;θ ∗,A i ). Let θ ∗ = (θ ∗,a ) a∈A ∈ R dK denote the concatenation of all action parameters. The reward function r(x,a;θ ∗ ) = E R∼p(·|x;θ ∗,a ) [R] gives the expected reward of action a in context x. The goal is to find a policy π ∈ Π that maximizes: V (π;θ ∗ ) = E X∼ν E A∼π(·|X) [r(X,A;θ ∗ )]. This chapter focuses on DM that estimates the value V (π;θ ∗ ) as: ˆ V dm (π) = 1 n ∑︂ i∈[n] ∑︂ a∈A π(a| X i )ˆr(X i ,a),(6.1) where ˆr(x,a) is an estimation of r(x,a;θ ∗ ). DM estimators may exhibit modeling bias, but they generally have lower variance than IPS (Saito and Joachims, 2022). Another advantage of DM is its practical utility without assuming access to the logging policy π 0 (Jeunen and Goethals, 2021; Aouali et al., 2022c; Hong et al., 2023). Also, DMs can be incorporated into a Bayesian framework, where informative priors can be used to enhance statistical efficiency. This allows for the development of scalable methods suitable for large action spaces, as shown in our work. 6.2 Structured DM 6.2.1 Structured Priors Pitfalls of non-structured priors. Before presenting sDM, we first describe the pitfalls of using the following widely used standard prior, θ a ∼N (μ a , Σ a ),∀a∈A,(6.2) R| θ,X,A∼N (φ(X) ⊤ θ A ,σ 2 ), where φ(x) provides a d-dimensional representation of the context x∈X, and N (μ a , Σ a ) represents the prior density of θ a , with σ 2 being the reward noise variance. Under this prior, each action a has an associated parameter θ a . Given the prior in Equation (6.2), the posterior distribution of an action parameter follows a multivariate Gaussian: θ a | D n ∼N (ˆμ a , ˆ Σ a ), where ˆ Σ −1 a = Σ −1 a + G a , ˆ Σ −1 a ˆμ a = Σ −1 a μ a + B a . with G a = σ −2 ∑︂ i∈[n] 1 A i =a φ(X i )φ(X i ) ⊤ ,B a = σ −2 ∑︂ i∈[n] 1 A i =a R i φ(X i ) Note that G a and B a only use the subset of samples D n where action a was observed, meaning data from other actions b ̸= a do not contribute to the posterior inference for 76 action a. This results in statistical inefficiency, especially if the logged data D n doesn’t cover all actions. In particular, the posterior for an unseen action a, θ a |D n , would simply revert to the prior N (μ a , Σ a ), since we would have G a = 0 d×d and B a = 0 d in such case. Structured priors. To address the above issue, we assume that action rewards correlate and embed this knowledge into the prior. While one could model these correlations by considering the joint posterior distribution of (θ a ) a∈A |D n , this becomes computationally burdensome when the number of actions K is large. Instead, we introduce an unknown d ′ - dimensional latent parameter ψ ∈ R d ′ , sampled from a latent prior q(·), such as ψ ∼ q(·). The correlations between actions naturally arise because each action parameter θ a is derived from the same latent parameter ψ. Specifically, the action parameters θ a are conditionally independent given ψ and are sam- pled from a conditional prior p a as θ a | ψ ∼ p a (·;f a (ψ)) for all a ∈ A. Here, p a is parameterized by f a (ψ), where f a : R d ′ → R d is a known prior function that encodes the hierarchical relationship between action parameters θ a and the latent parameter ψ. This structure allows for sparsity, meaning that θ a may depend only on a subset of ψ’s coordinates. Moreover, p a accounts for model uncertainty, allowing for cases where θ a is not a deterministic function of ψ, i.e., θ a ̸= f a (ψ). The reward distribution for action a in context x is given by p(· | x;θ a ), which depends only on x and θ a . To summarize, the structured prior is defined below, and its graphical representation is given in Figure 6.1. ψ ∼ q(·),(6.3) θ a | ψ ∼ p a (·;f a (ψ)),∀a∈A, R| ψ,θ,X,A∼ p(·| X;θ A ). To derive the posterior under this prior, we assume that: (i) (X,A) is independent of ψ, and given ψ, (X,A) is independent of θ; and (i) given ψ, the parameters θ a for all a∈A are independent. : taken action Figure 6.1: Graph representation of the structured prior. Now, we discuss how to perform off-policy learning under this general structured prior in Equation (6.3), before applying it to linear-Gaussian distributions in Section 6.3. 77 6.2.2 Off-Policy Learning Off-policy learning relies on an estimate of the value function V (π;θ ∗ ) obtained using the logged data D n . In DMs, the estimator ˆ V dm in Equation (6.1) requires access to the learned reward ˆr(x,a) ≈ r(x,a;θ ∗ ). In our Bayesian setting, this requires access to the action posterior θ a | D n under the prior in Equation (6.3) since the reward is then estimated as ˆr(x,a) = E [r(x,a;θ)|D n ] for any (x,a) ∈ X × A, and this estimate is plugged into ˆ V dm in Equation (6.1) to estimate V (π;θ ∗ ). Thus, we need to derive the posterior density of the action parameter θ a , p(θ a | D n ), under the structured prior in Equation (6.3), which reads p(θ a |D n ) = ∫︂ ψ p(θ a | ψ,D n )p(ψ |D n ) dψ ,(6.4) where ψ |D n is the latent posterior and θ a | ψ,D n is the conditional action posterior. To compute p(θ a |D n ), we first compute p(θ a | ψ,D n ) and p(ψ |D n ) and then integrate out ψ following Equation (6.4). First, p(θ a | ψ,D n )∝L a (θ a )p a (θ a ;f a (ψ)),(6.5) with L a (θ a ) = ∏︁ (X,A,R)∈S a p(R|X;θ a ) is the likelihood of observations of action a (S a = (X i ,A i ,R i ) i∈[n],A i =a is the subset of D n where A i = a). Similarly, p(ψ |D n )∝ ∏︂ b∈A ∫︂ θ b L b (θ b )p b (θ b ;f b (ψ)) dθ b q(ψ),(6.6) This allows us to further develop Equation (6.4) as p(θ a |D n )∝ ∫︂ ψ L a (θ a )p a (θ a ;f a (ψ)) ∏︂ b∈A ∫︂ θ b L b (θ b )p b (θ b ;f b (ψ)) dθ b q(ψ) dψ .(6.7) All the quantities inside the integrals in Equation (6.7) are given (the parameters of p a and q) or tractable (the terms in L a ). Thus, if these integrals can be computed, then the posterior can be fully characterized in closed form, which we will do in Section 6.3 in the fully linear case. Otherwise, the posterior should be approximated. Finally, we act greedy with respect to our estimator ˆ V dm and define the learned policy as the one maximizing it: ˆπ g = argmax π∈Π ˆ V dm (π). If the set of policies Π contains deterministic policies, then ˆπ g (a| x) = 1a = argmax b∈A ˆr(x,b).(6.8) In particular, we do not adopt the common pessimism approach (Jin et al., 2021). In pessimism, one constructs confidence intervals of the reward estimate ˆr(x,a) of the form |r(x,a;θ)− ˆr(x,a)| ≤ u(x,a), and then defines the learned policy as ˆπ p (a | x) = 1a = argmax b∈A ˆr(x,b)− u(x,b). The advantage of one over another depends on the evaluation metric used. Our metric is the Bayesian suboptimality (BSO), defined in Section 6.4. It assesses the average performance of algorithms across multiple problems rather than the worst-case. The Greedy policy is more suitable for BSO optimization than pessimism (demonstrated theoretically and empirically in Sections C.3.3 and C.4.4). 78 6.3 Linear-Gaussian Case In this section, we use linear functions f a combined with Gaussian distributions for the structured prior Equation (6.3). Precisely, we assume that the latent prior q(·) = N (·;μ, Σ) is Gaussian with mean μ ∈ R d ′ and covariance Σ ∈ R d ′ ×d ′ . Moreover, let W a ∈ R d×d ′ be the mixing matrix for action a, we define f a (v) = W a v for any v ∈ R d ′ . We define the conditional prior p a (·;f a (ψ)) = N (·; W a ψ, Σ a ) is Gaussian with mean f a (ψ) = W a ψ ∈ R d and covariance Σ a ∈ R d×d . The reward distribution p(· | x;θ a ) is also linear-Gaussian as N (·;φ(x) ⊤ θ a ,σ 2 ), where φ(·) outputs a d-dimensional representa- tion of x and σ > 0 is the observation noise variance. The whole prior is ψ ∼N (μ, Σ),(6.9) θ a | ψ ∼N (︂ W a ψ, Σ a )︂ ,∀a∈A, R| ψ,θ,X,A∼N (φ(X) ⊤ θ A ,σ 2 ). 6.3.1 Applications Mixed-effect modeling. Equation (6.9) allows modeling that action parameters depend on a linear mixture of effect parameters (Chapter 3). Precisely, let J be the number of effects and assume that d ′ = dJ so that the latent parameter ψ is the concatenation of J, d-dimensional effect parameters, ψ j ∈ R d , such as ψ = (ψ j ) j∈[J ] ∈ R dJ . Moreover, assume that for any a ∈ A, W a = w ⊤ a ⊗ I d ∈ R d×dJ where w a = (w a,j ) j∈[J ] ∈ R J are the mixing weights of action a. Then, W a ψ = ∑︁ j∈[J ] w a,j ψ j for any a ∈ A. Sparsity, i.e., when an action a only depends on a subset of effects, is captured through the mixing weights w a : w a,j = 0 when action a is independent of the j-th effect parameter ψ j and w a,j ̸= 0 otherwise. Also, the level of dependence between action a and effect j is quantified by the absolute value of w a,j . This mixed-effect model can be used in numerous applications (the reader can refer to the first paragraphs of Chapter 3 for examples). Low-rank modeling. Equation (6.9) can also model the case where the dimension of the latent parameter ψ is much smaller than that of the action parameters θ a , i.e., when d ′ ≪ d. Again, this is captured through the mixing matrices W a , when W a is low-rank. 6.3.2 Closed-Form Solutions for sDM The conditional action posterior is known in closed-form as θ a | ψ,D n ∼N ( ̃μ a , ̃ Σ a ), with ̃ Σ −1 a = Σ −1 a + G a , ̃ Σ −1 a ̃μ a = Σ −1 a W a ψ + B a ,(6.10) where G a = σ −2 ∑︂ i∈[n] 1A i = aφ(X i )φ(X i ) ⊤ , B a = σ −2 ∑︂ i∈[n] 1A i = aR i φ(X i ). 79 This posterior has the standard form except that the prior mean W a ψ now depends on the latent parameter ψ. Similarly, the effect posterior writes ψ |D n ∼N ( ̄μ, ̄ Σ), where ̄ Σ −1 = Σ −1 + ∑︂ a∈A W ⊤ a (Σ −1 a − Σ −1 a ̃ Σ a Σ −1 a )W a , ̄ Σ −1 ̄μ = Σ −1 μ + ∑︂ a∈A W ⊤ a Σ −1 a ̃ Σ a B a .(6.11) The latent posterior precision ̄ Σ −1 is the sum of the latent prior precision Σ −1 and the learned action precisions Σ −1 a − Σ −1 a ̃ Σ a Σ −1 a , weighted by W ⊤ a W a . The contribution of each action’s learned precision to the latent precision is proportional to W ⊤ a W a . This intuition similarly applies to interpreting ̄μ. Finally, from Equation (6.7), the action posterior is θ a |D n ∼N (ˆμ a , ˆ Σ a ), where ˆ Σ a = ̃ Σ a + ̃ Σ a Σ −1 a W a ̄ ΣW ⊤ a Σ −1 a ̃ Σ a ,ˆμ a = ̃ Σ a (︁ Σ −1 a W a ̄μ + B a )︁ .(6.12) Finally, from Equation (6.9), the reward function is r(x,a;θ) = φ(x) ⊤ θ a . Thus, the estimated reward is ˆr(x,a) = E[r(x,a;θ)|D n ] = φ(x) ⊤ ˆμ a ,∀(x,a)∈X ×A. This can then be plugged in Equation (6.8) for decision-making, leading to ˆπ g (a| x) = 1a = argmax b∈A φ(x) ⊤ ˆμ b . To see why this is more beneficial than the standard prior in Equation (6.2), notice that the mean and covariance of the posterior of action a, ˆμ a and ˆ Σ a , are now computed us- ing the mean and covariance of the latent posterior, ̄μ and ̄ Σ. But ̄μ and ̄ Σ are learned using the interactions with all the actions in D n . Thus ˆμ a and ˆ Σ a are also learned us- ing the interactions with all the actions in D n , in contrast with the standard prior in Equation (6.2) where they were learned using only the interaction with action a. The ad- ditional computational cost of considering the structured prior in Equation (6.9) is small. The computational and space complexities are O(K((d 2 + d ′ 2 )(d + d ′ ))) and O(Kd 2 ). For example, when d ′ = O(d), these complexities become O(Kd 3 ) and O(Kd 2 ), respec- tively. This is exactly the cost of the standard prior in Equation (6.2). In contrast, this strictly improves the computational efficiency of jointly modeling the action parameters, where the complexities areO(K 3 d 3 ) andO(K 2 d 2 ) since the joint posterior of (θ a ) a∈A |D n requires converting and storing a dK× dK covariance matrix. Remark 6. sDM with linear-Gaussian hierarchies can be used even with data generated from non-linear rewards, and we empirically investigate its robustness to misspecification. We found that this model performs well even if the true rewards are not generated from a linear-Gaussian distribution. 6.4 Analysis 6.4.1 Bayesian Metrics The performance of a learned policy ˆπ is evaluated using suboptimality (SO): so(ˆπ;θ ∗ ) = V (π ∗ ;θ ∗ )− V (ˆπ;θ ∗ ), 80 where π ∗ = argmax π∈Π V (π;θ ∗ ) is the optimal policy. This metric is well-suited when the environment is governed by a unique, fixed ground truth θ ∗ . It applies to any policy ˆπ, whether learned through frequentist approaches (e.g., MLE) or Bayesian ones (e.g., ours). However, when the environment is modeled as a random variable θ ∗ sampled from some prior distribution, SO becomes less appropriate. Thus, drawing on recent developments in Bayesian analysis for online bandits through Bayes regret (Russo and Van Roy, 2014), we introduce a new metric for offline settings, termed Bayes suboptimality, defined as: Bso(ˆπ) = E[V (π ∗ ;θ ∗ )− V (ˆπ;θ ∗ )],(6.13) where the expectation is taken over all random variables: the logged data D n and θ ∗ , which is treated as a random variable sampled from the prior. The BSO can be computed in two ways. One method involves taking the expectation under the prior θ ∗ , followed by taking an expectation under data generated from a fixed environment θ ∗ as D n | θ ∗ . The other method involves taking an expectation under the data D n , followed by taking an expectation under the posterior θ ∗ |D n . The BSO is a reasonable metric for assessing the average performance of algorithms across multiple environments, due to the expectation over θ ∗ . It is also known that Bayes regret captures the benefits of using informative priors (Chapter 3), and this is similarly achieved by the BSO. 6.4.2 Theoretical Results Our theory relies on the important well-specified assumption: Assumption 1 (Well-specified priors). Action parameters θ ∗,a and rewards are drawn from Equation (6.9). We also make simplifying assumptions for the sake of exposition. Assumption 2 (Diagonal covariances for simplicity). We assume Σ a = σ 2 0 I d , Σ = τ 2 I d ′ , ∥φ(x)∥ 2 ≤ 1, and the matrices W a are normalized such that λ 1 (W a W ⊤ a ) = λ d (W a W ⊤ a ) = 1. This yields our bound on the BSO of sDM. Theorem 2 (Covariance-Dependent Bound). Let π ∗ (x) be the optimal action for context x. Then the BSO of sDM under the structured prior in Equation (6.9) satisfies Bso(ˆπ g )≤ α n E [︂ ∥φ(X)∥ ˆ Σ π ∗ (X) ]︂ + √︃ (2 log(2K) + 2)(σ 2 0 + τ 2 ) n ,(6.14) where α n = √︂ d + 2 √︁ d log(Kn) + 2 log(Kn). Scaling of the bound in Theorem 2 aligns with existing frequentist results (Jin et al., 2021, Theorem 4.4). The main differences lie in the constants and the fact that this rate is achieved using greedy policies in Equation (6.8). This contrasts with the frequentist setting where pessimism is used (Jin et al., 2021) and known to be optimal (Jin et al., 2021, 81 Theorem 4.7). In fact, greedy policies are optimal when BSO is used as the performance metric. Specifically, Bso(ˆπ g ) ≤ Bso(π) for any policy π, including pessimistic ones. Therefore, in the Bayesian setting and when BSO is used as a performance metric, greedy policies should always be preferred to pessimistic ones. This fundamental difference is proven in Section C.3.3 and it is of independent interest beyond this work. Theorem 2 suggests that the BSO primarily depends on the posterior covariance of action π ∗ (X) in the direction of the context φ(X). That is, when the uncertainty in the posterior distribution of the optimal action π ∗ (X) is low on average across different contexts X and logged dataD n , then the BSO bound is correspondingly small. In particular, the tightness of the bound depends on the degree to which the logged data covers the optimal actions on average. Theorem 2 can highlight the advantages of using sDM over the non-structured prior in Equation (6.2). To see this, notice that the parameters of the non-structured prior in Equation (6.2), μ a and Σ a , are obtained by marginalizing out ψ in Equation (6.9). In this case, μ ns a ← W a μ and Σ ns a ← Σ a + W a ΣW ⊤ a . The corresponding posterior covariance is ˆ Σ ns a = ((Σ a + W a ΣW ⊤ a ) −1 + G a ) −1 , and is generally larger than the covariance of sDM, ˆ Σ a in Equation (6.12). This is more pronounced when the number of actions K is large and when the latent parameters are more uncertain than the action parameters. Thus, the BSO bound of sDM is smaller due to the reduced posterior uncertainty it exhibits. Also, note that even when π ∗ (X) is unobserved in the logged dataD n , sDM’s posterior covariance ˆ Σ π ∗ (X) can remain small since we use interactions with all actions to compute it. This contrasts with standard non-structured priors in Equation (6.2), where observing π ∗ (X) is necessary; without such observations, the posterior covariance ˆ Σ π ∗ (X) would simply be the prior covariance Σ π ∗ (X) . Next, we provide another bound on the BSO that scales as O(1/ √ n). To simplify the exposition, we roughly present its scaling with n in Theorem 3 and defer the complete general statement to Section C.3.2. We make the following additional assumptions: Assumption 3. Let G = E X∼ν [X ⊤ ] with g = λ d (G). We assume that g > 0. Assumption 4 (Context-independent logging policy). A is independent of X, i.e., π 0 (a| x) = π 0 (a) = p a for all x and a. Equivalently, (X i ) are i.i.d. ∼ ν and independent of (A i ), with P(A = a) = p a . Theorem 3 (Scaling with n). For n large enough, the BSO of sDM under the structured prior in Equation (6.9) scales as Bso(ˆπ g ) = ̃ O (︄ √︄ E [︃ d nρ X + 1 ]︃ + √︃ logK n )︄ , where ρ X = π 0 (π ∗ (X)). The above bound becomes smaller or larger depending on how well the logging policy π 0 covers the optimal actions for each context x. 82 6.5 Experiments We evaluate sDM using both synthetic and real datasets. We use the average reward of the learned policy relative to the optimal policy as the evaluation metric. 6.5.1 Synthetic Problems Setting. We simulate synthetic data using the linear-Gaussian model in Equation (6.9) with σ = 1. The contexts X are sampled uniformly from [−1, 1] d , with d = 10. The matrices W a are sampled uniformly from [−1, 1] d×d ′ , where we very d ′ as d ′ ∈5, 10., 20. We set Σ = 3I d ′ and Σ a = I d , meaning the latent parameters are more uncertain than the action parameters. The latent mean μ is randomly sampled from [−1, 1] d ′ . The number of actions is varied as K ∈ 100, 1000, and we use a uniform logging policy to collect data. Additional experiments with different logging policies are presented in Section C.4. Baselines. First, we use sDM under prior in Equation (6.9). Second, we examine DM (Bayes), which uses the standard non-structured prior in Equation (6.2), where param- eters μ a and Σ a are obtained by marginalizing out the latent parameters ψ in Equa- tion (6.9). Thus DM (Bayes) is a standard Bayesian DM that does not capture arm reward correlations. We also include DM (Freq), which estimates θ ∗,a by the MLE. We include IPS (Horvitz and Thompson, 1952), self-normalized IPS (snIPS) (Swaminathan and Joachims, 2015b), and doubly robust (DR) (Dudik et al., 2014), which we optimize to learn the optimal policy. MIPS (Saito and Joachims, 2022) and PC (Sachdeva et al., 2024) are also included. Implementation details of baselines is provided in Section C.4.1. Results. In Figure 6.2, we plot the results and we observe that sDM consistently out- performs the baselines across all settings. This performance gap becomes even more significant when sample size n is small. These results highlight sDM’s enhanced efficiency in using available logged data, making it particularly beneficial in data-limited situations and scalable to large action spaces. 0500100015002000 Number of samples n 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Avg. relative reward Synthetic - OPL K=100, d'=5, d=10 0500100015002000 Number of samples n Synthetic - OPL K=1000, d'=5, d=10 0500100015002000 Number of samples n Synthetic - OPL K=1000, d'=10, d=10 0500100015002000 Number of samples n Synthetic - OPL K=1000, d'=20, d=10 sDM (Ours)DM (Bayes)DM (Freq)DRIPSsnIPSMIPSPC Figure 6.2: The average relative reward of the learned policy using one of the baselines on synthetic problems with varying n, K and d ′ . Scaling to large action spaces. sDM achieves improved scalability compared to stan- dard DM as it leverages data more efficiently. While it still learns a d-dim. parameter for each action a, it does so by considering interactions with all actions in the logged 83 data D n , instead of only using interactions with the specific action a. This is crucial, especially given that many actions may not even be observed in D n . To show sDM’s improved scalability, we compare it to the most competitive baseline, DM (Bayes), for varying K ∈ [10, 100000] with n = 1000. The results in Figure 6.3 reveal that the perfor- mance gap between sDM and DM (Bayes) becomes more significant when the number of actions K increases. Hence, despite the necessity for sDM to learn distinct parameters for each action, accommodating practical scenarios like recommender systems where unique embeddings are learned for each product, it still enjoys good scalability. 105005000100000 Number of actions K 0.5 0.6 0.7 0.8 0.9 1.0 Avg. relative reward Varying K, d'=5, d=10 sDM (Ours) DM (Bayes) Figure 6.3: sDM vs. DM (Bayes) for varying K. 6.5.2 MovieLens Problems Setting. We use MovieLens 1M (Lam and Herlocker, 2016), which contains 1 million ratings representing the interactions between 6,040 users and 3,952 movies. To create a semi-synthetic environment, we first apply a low-rank factorization to the rating matrix, producing 5-dim. representations: x u ∈ R 5 for user u ∈ [6040] and θ a ∈ R 5 for movie a ∈ [3952]. Movies are treated as actions, and contexts X are sampled randomly from the user vectors. The reward for movie a and user u is modeled as N (x ⊤ u θ a , 1), serving as proxy for ratings. A uniform logging policy is used to collect data. Baselines. We consider the same baselines as in synthetic data. A prior is not needed for DM (Freq), IPS, snIPS, and DR. However, for DM (Bayes), a standard prior in Equa- tion (6.2) is inferred from data, where we set μ a to be the mean of movie vectors across all dimensions, and Σ a = diag(v), where v represents the variance of movie vectors across all dimensions. Unlike the synthetic experiments, the latent structure assumed by sDM is not inherently present in MovieLens. But we learn it by training a Gaussian Mixture Model (GMM) to cluster movies into J = 5 mixture components. This gives rise to the mixed-effect structure described in Section 6.3, which represents a specific instance of sDM with d ′ = dJ = 25. MIPS also has access to movie clusters, while we use the knn smoothing implementation of PC (see (Sachdeva et al., 2024, Section 3)). Note that DM (Bayes), sDM, MIPS and PC use the same subset of data (of size 1000) to learn their pri- ors/assumed structure and thus we compare them fairly. We conduct experiments with K ∈100, 1000 randomly selected movies. Results. Results are in Section 6.5.2. Even though the latent structure assumed by sDM is not inherently present in MovieLens, sDM still outperforms the baselines by learning it 84 offline. 0500100015002000 Number of samples n 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 Avg. relative reward MovieLens - OPL K=100, d'=25, d=5 0500100015002000 Number of samples n MovieLens - OPL K=1000, d'=25, d=5 sDM (Ours) DM (Bayes) DM (Freq) DR IPS snIPS MIPS PC Figure 6.4: The average relative reward of the learned policy using one of the baselines on MovieLens problems with varying n, K and d ′ . 6.6 Conclusion We introduced sDM, a structured approach to off-policy learning that leverages latent structure among actions to enhance statistical efficiency while maintaining computational tractability, particularly in large action spaces with limited data coverage. Within a Bayesian framework, we proved that greedy policies outperform pessimistic ones under Bayesian suboptimality and established O(1/ √ n) convergence without requiring restric- tive full-support assumptions. Our work has several limitations. First, our theoretical analysis assumes a well-specified prior; while we empirically observed robustness to misspecification, formal guarantees under prior mismatch remain an open question. Second, closed-form posterior updates are available only for linear-Gaussian hierarchies; extending to nonlinear reward models requires approximate inference, which may compromise computational efficiency or sta- tistical accuracy. Third, the latent structure must be specified or learned pre-trained, adding an additional modeling task. Finally, extending sDM to handle nonlinear hierar- chies, building on the diffusion-based approach of Chapter 4, is a promising avenue for future work. 85 Chapter 7 Optimization Matters More than Esimation Contents 6.1 Setting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.2 Structured DM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.2.1 Structured Priors . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.2.2 Off-Policy Learning . . . . . . . . . . . . . . . . . . . . . . . . 78 6.3 Linear-Gaussian Case . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 6.3.1 Applications . . . . . . . . . . . . . . . . . . . . . . . . . . . . 79 6.3.2 Closed-Form Solutions for sDM . . . . . . . . . . . . . . . . . . 79 6.4 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 80 6.4.1 Bayesian Metrics . . . . . . . . . . . . . . . . . . . . . . . . . 80 6.4.2 Theoretical Results . . . . . . . . . . . . . . . . . . . . . . . . 81 6.5 Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 6.5.1 Synthetic Problems . . . . . . . . . . . . . . . . . . . . . . . . 83 6.5.2 MovieLens Problems . . . . . . . . . . . . . . . . . . . . . . . 84 6.6 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 85 This chapter challenges the dominant paradigm in off-policy learning (explored in Chap- ter 8), which frames the problem as finding a policy ˆπ = argmax π ˆ V (π) (or, with pes- simism, ˆπ = argmax π [ ˆ V (π)− pen(π)]), where ˆ V is an IPS-based 1 estimate of the true policy value V (π). The rationale behind these objectives is that maximizing a more accu- rate value estimate yields a better policy. However, this estimator-centric view neglects a crucial factor: the optimization landscape. 1 Recall that IPS is an importance-weighting estimator of the policy value. We use IPS-based to refer to any estimator derived from or inspired by importance weighting. 86 IPS-based objectives (Dudík et al., 2011; Dudík et al., 2012; Dudik et al., 2014; Wang et al., 2017; Farajtabar et al., 2018; Su et al., 2020; Metelli et al., 2021; Kuzborskij et al., 2021; Saito and Joachims, 2022) are highly non-concave under common policy parameterizations (Chen et al., 2019), prone to suboptimal local maxima and plateaus: issues that are exacerbated in large action spaces. Even sophisticated estimators designed to reduce variance fail to overcome this optimization barrier, as they often induce equally difficult landscapes. We make the following contributions. 1) We show that objective-aware policy parametriza- tion can partially alleviate these difficulties by structuring the policy class to match the implicit biases of the estimator. Such parametrizations reduce the effective search space and can shorten optimization plateaus and local maxima. However, this strategy does not eliminate the fundamental non-concavity of IPS-based objectives, leaving optimization as the central bottleneck. 2) Motivated by this limitation, we advocate for an alternative approach based on policy-weighted log-likelihood (PWLL) objectives. Unlike traditional estimators, PWLL optimizes an objective ˆ U (π) designed for ease of optimization rather than accuracy in estimating V (π). Although PWLL objectives perform poorly as value estimators, their favorable concave landscape makes them significantly more effective for policy learning. 3) Through theoretical and empirical analysis, we demonstrate that this optimization-centric approach consistently enables simpler PWLL objectives to outper- form complex, state-of-the-art IPS-based methods, particularly in large action spaces. Setting and organization. This chapter considers the general setting of Section 5.1 with R∈ [0, 1]. The remainder is organized as follows. Section 7.1 employs an asymptotic lens to analyze IPS-based objectives and derives objective-aware policy parametrizations that partially alleviate their optimization challenges. Section 7.2 introduces PWLL objectives and establishes their favorable optimization properties. Section 7.3 presents large-scale experiments. We conclude in Section 7.4. 7.1 Analysis of IPS-Based Objectives IPS-based objectives optimize an estimator ˆ V (π) of the policy value V (π). To under- stand the policies to which these estimators converge, we study their oracle policies π method ∗ = argmax π E[ ˆ V method (π)]. Taking the expectation removes sampling fluctuations and isolates the inductive bias of each objective: different estimators yield different oracle policies, even with infinite data. Crucially, oracle policies admit closed-form expressions, enabling precise characterization of each estimator’s implicit bias. This analysis motivates objective-aware parametrizations that align the policy class with the estimator’s bias to ease optimization: the first improvement we propose in this chapter. 7.1.1 Standard IPS-Based Objectives The foundational IPS estimator (Horvitz and Thompson, 1952) re-weights observed re- wards by the ratio between the target policy π and the logging policy π 0 : ˆ V ips (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) R i .(7.1) 87 In expectation, IPS selects the best-rewarding action among those in the support of π 0 : π IPS ∗ (a| x) = 1 [︃ a = argmax a ′ ∈A r(x,a ′ )1[π 0 (a ′ | x) > 0] ]︃ .(7.2) Clipped IPS (cIPS). To mitigate the high variance of IPS, a widely used variant is cIPS (Bottou et al., 2013) that clips small propensity scores at a threshold τ ∈ (0, 1): ˆ V cips (π) = 1 n n ∑︂ i=1 π(A i | X i ) maxπ 0 (A i | X i ),τ R i .(7.3) This clipping introduces a bias. The oracle policy down-weights the rewards of rare ac- tions, causing it to favor actions that were frequent under π 0 , even if they are suboptimal: π cIPS ∗ (a| x) = 1 [︂ a = argmax a ′ ∈A π 0 (a ′ | x) maxπ 0 (a ′ | x),τ r(x,a ′ ) ]︂ .(7.4) Exponential smoothing (ES). Instead of hard clipping, ES (Aouali et al. (2023a), Chap- ter 8) smooths importance weights by raising propensities to a fractional power α∈ (0, 1): ˆ V es (π) = 1 n n ∑︂ i=1 π(A i | X i ) π 0 (A i | X i ) α R i .(7.5) Its oracle policy balances reward maximization with preference for frequent actions: π ES ∗ (a| x) = 1 [︃ a = argmax a ′ ∈A r(x,a ′ )π 0 (a ′ | x) 1−α ]︃ .(7.6) Another variant of ES regularizes the entire importance weight as ( π π 0 ) β instead of only the denominator. In contrast to the deterministic policies derived from IPS, cIPS, and the ES formulation above, this approach yields a stochastic oracle policy: π ES ∗ (a | x) ∝ r(x,a) 1/(1−β) π 0 (a| x). Other regularizations include logarithmic smoothing (Sakhi et al., 2024), implicit exploration (Gabbianelli et al., 2024), harmonic correction (Metelli et al., 2021), shrinkage (Su et al., 2020). But we do not include as ES and cIPS are already representative of them. Doubly robust (DR). The DR estimator incorporates a reward model ˆr(x,a) to reduce variance and enable generalization to actions outside π 0 ’s support. A common clipped variant is: ˆ V dr (π) = 1 n n ∑︂ i=1 π(A i | X i ) maxπ 0 (A i | X i ),τ (R i − ˆr(X i ,A i )) + E A∼π(·|X i ) [ˆr(X i ,A)]. (7.7) Its oracle policy interpolates between the reward model prediction and an importance weighting correction for the reward model error: π DR ∗ (a| x) = 1 [︂ a = argmax a ′ ∈A ˆr(x,a ′ ) + π 0 (a ′ | x) maxπ 0 (a ′ | x),τ (r(x,a ′ )− ˆr(x,a ′ )) ]︂ . (7.8) 88 7.1.2 Large-Scale IPS-Based Objectives In large action spaces, importance weights π(a|x) π 0 (a|x) can become huge, leading to estimators with high variance. To mitigate this, modern methods compute marginalized importance weights over a lower-dimensional action representation, trading bias for reduced variance. Marginalized IPS (MIPS). MIPS (Saito and Joachims, 2022) tackles large action spaces by clustering actions. It maps each action a to a cluster c via a function h :A→C, where |C |≪|A|. Estimation is then performed at the cluster level: ˆ V mips (π) = 1 n n ∑︂ i=1 π(C i | X i ) π 0 (C i | X i ) R i , where C i = h(A i ) and π(c| x) = ∑︂ a∈c π(a| x). (7.9) This cluster-level marginalization introduces bias: the oracle policy only selects the best cluster based on its average reward under π 0 , and cannot differentiate between actions within that cluster: π MIPS ∗ (c| x) = I [︃ c = argmax c ′ ∈C ︃ ∑︁ a∈c ′ π 0 (a| x)r(x,a) ∑︁ a∈c ′ π 0 (a| x) ︃]︃ .(7.10) Hence, MIPS offers no specific guidance for selecting an action within the optimal cluster; any action is considered equally valid. Consequently, one possible induced action-level oracle under uniform tie-breaking is: π MIPS ∗ (a| x) = I [︂ h(a) = argmax c ′ ∈C ︂ ∑︁ a∈c ′ π 0 (a|x)r(x,a) ∑︁ a∈c ′ π 0 (a|x) ︂]︂ |h(a)| . where |h(a)| denotes the size of the cluster containing action a. Conjunct effect modeling (OffCEM). Building on MIPS, OffCEM (Saito et al., 2023) uses a reward model ˆr to correct for the cluster-level aggregation bias, in a doubly robust fashion: ˆ V offcem (π) = 1 n n ∑︂ i=1 (︃ π(C i | X i ) π 0 (C i | X i ) (R i − ˆr(X i ,A i )) + E A∼π(·|X i ) [ˆr(X i ,A)] )︃ .(7.11) The resulting oracle policy selects the action that maximizes the model-predicted reward ˆr, plus a cluster-level correction term that accounts for model error: π OffCEM ∗ (a| x) = I [︄ a = argmax a ′ ∈A ︄ ˆr(x,a ′ ) + ∑︁ ̄a∈h(a ′ ) π 0 ( ̄a| x)(r(x, ̄a)− ˆr(x, ̄a)) ∑︁ ̄a∈h(a ′ ) π 0 ( ̄a| x) ︄]︄ . (7.12) Two-stage decomposition (POTEC). In this chapter, we see POTEC (Saito et al., 2025) as an optimization strategy of OffCEM (rather than seeing it as a new estimator). It restricts the policy to a cluster-informed form, π(a| x) = ∑︂ c∈C π rm (a| x,c)π cl (c| x), 89 where π rm (a | x,c) = 1[a = argmax a ′ ∈c ˆr(x,a ′ )] is fixed, model-based policy that de- terministically selects the best action within each cluster. Learning is then simplified to finding the optimal cluster-level policy π cl that maximizes the OffCEM objective in Equation (7.11): ˆ V potec (π cl ) = 1 n n ∑︂ i=1 (︄ π cl (C i | X i ) π 0 (C i | X i ) (R i − ˆr(X i ,A i )) + ∑︂ c∈C π cl (c| X i )ˆr ∗ c (X i ) )︄ , (7.13) where ˆr ∗ c (x) = max a∈c ˆr(x,a) is the estimated reward of the best action in cluster c. This practical decomposition has the same optimal oracle policy as OffCEM: π POTEC ∗ = π OffCEM ∗ . Policy convolution (PC). Moving beyond hard clustering, PC (Sachdeva et al., 2024) leverages the assumption that actions close in an embedding space yield similar rewards. For each action a, it aggregates over its neighborhood of nearest neighbors N ε (a) =a ′ : d(a,a ′ ) < ε, where d is a pre-defined distance metric (e.g., ℓ 2 distance between action embeddings): ˆ V pc (π) = 1 n n ∑︂ i=1 π(N ε (A i )| X i ) π 0 (N ε (A i )| X i ) R i , with π(N ε (a)| x) = ∑︂ a ′ ∈N ε (a) π(a ′ | x).(7.14) The induced oracle policy is deterministic: it selects the action a ′ that maximizes an aggregated neighborhood score. Each logged neighbor ̄a ∈ N ε (a ′ ) contributes its reward r(x, ̄a), weighted by the conditional probability of observing ̄a under the logging policy restricted to its neighborhood. π PC ∗ (a| x) = I ⎡ ⎣ a = argmax a ′ ∈A ⎧ ⎨ ⎩ ∑︂ ̄a∈N ε (a ′ ) π 0 ( ̄a| x)r(x, ̄a) π 0 (N ε ( ̄a)| x) ⎫ ⎬ ⎭ ⎤ ⎦ .(7.15) Other recent IPS variants for large action spaces (Peng et al., 2023; Cief et al., 2024; Taufiq et al., 2024) are often extensions of MIPS that relax its core assumptions. We focused on four methods (MIPS, OffCEM, POTEC, and PC), which we consider representative of this family. Since these variants largely share the same MIPS foundation and optimization procedure (with the notable exception of POTEC), we expect our findings to be generally applicable. 7.1.3 Optimization Challenges The effectiveness of IPS-based estimators in off-policy learning is often limited by their challenging optimization landscape. These objectives become difficult to optimize when paired with standard, expressive policy classes such as the softmax. This section explores why this occurs and introduces objective-aware parametrization as a strategy to mitigate, though not entirely solve, the problem. 90 To analyze the optimization process, we consider policies parametrized by a softmax function over an effective action space 2 A eff ⊆A, which is the set of actions that can be assigned non-zero probability. By default,A eff =A, but we explain below why restricting it to match the structure of the estimator’s oracle policy can be beneficial. Specifically, the policy takes the form: π θ (a| x) = exp(s θ (x,a)) ∑︁ a ′ ∈A eff exp(s θ (x,a ′ )) 1 a∈A eff , ∀a∈A,(7.16) where s θ (x,a) is a learnable score function. Common choices are linear softmax scores: lightweight: s θ (x,a) = φ(x,a) ⊤ θ ,heavyweight: s θ (x,a) = φ(x) ⊤ θ a , (7.17) which we call lightweight parametrization (a single shared parameter vector θ, correspond- ing to a joint reward model) and heavyweight parametrization (separate parameters θ a for each action, corresponding to a disjoint reward model). The size of the effective action space, K eff = |A eff |, is the critical factor governing opti- mization difficulty. The following propositions (proofs in Section D.2, adapted from Chen et al. (2019); Mei et al. (2020a)) reveal the severity of the problem. First, gradient-based methods can become trapped in suboptimal regions for extended periods. Proposition 3 (Optimization plateaus). For any IPS-based estimator ˆ V that is linear in π, even with a linear softmax policy, there exist problem instances where gradient ascent remains trapped in a suboptimal region for O(K eff ) iterations. Second, the optimization landscape has numerous poor local maxima. Proposition 4 (Local maxima). Under similar conditions, the optimization landscape for IPS-based objectives can contain a number of local maxima that is exponential in K eff . These results highlight that K eff plays a central role in optimization difficulty. The stan- dard choice ofA eff =A, which sets K eff = K, leads to optimization failure in large action spaces where K can reach millions: learning must navigate a landscape with potentially O(K)-length plateaus and exponentially many local maxima. Surprisingly, even sophisticated methods designed specifically for large action spaces often fall into this trap. At first glance, methods such as MIPS, OffCEM, and PC appear to operate in a smaller space because their objectives involve marginalized probabilities: π(C i | X i ) in MIPS and OffCEM, or π(N ε (A i ) | X i ) in PC. However, these marginalized terms are defined as sums over an underlying action-level policy: π(C i | X i ) = ∑︂ a∈C i π(a| X i ),and π(N ε (a)| x) = ∑︂ a ′ ∈N ε (a) π(a ′ | x). 2 The effective action space can also depend on context x, i.e., A eff (x)⊆A. We omit this dependence for notational simplicity. 91 Then, if π(a| x) is a softmax overA, then K eff = K and Propositions 3 and 4 apply with K eff = K which is large. The only exception is POTEC, which fixes the intra-cluster policy π rm and only optimizes a cluster-level policy π cl . This reduces the effective action space to A eff =C with K eff =|C|≪ K, directly mitigating the optimization pathologies. Design implications: objective-aware parametrization The choice of K eff introduces a fundamental trade-off. A smaller effective action space simplifies the optimization landscape, but risks excluding the optimal action and reduces policy expressiveness. IfA eff is chosen arbitrarily, it may degrade performance. The chal- lenge is to find the sweet spot: a parametrization constrained enough to be optimizable, yet expressive enough to contain the objective’s maximizer. This is precisely where our asymptotic analysis helps. The oracle policy π method ∗ reveals the minimal sufficient set of actions required to maximize each objective. By aligning the policy parametrization with this structure, we can reduce K eff without sacrificing performance: the core principle of our proposed objective-aware parametrization. For instance, the oracle policies for IPS, cIPS, and ES are confined to the support of the logging policy, S 0 (x). This implies that A eff = S 0 is sufficient, reducing K eff from K to |S 0 | ≪ K. Similarly, for OffCEM and MIPS, the cluster-level structure of their oracle policies suggests a two-stage decomposition similar to that of POTEC, reducing K eff to|C|. We summarize these observations as claims, validated empirically in Section 7.3: Claim 1. For IPS, cIPS, and ES, restricting the policy support to S 0 reduces K eff and yields superior learned policies. Claim 2. For OffCEM and MIPS, a two-stage POTEC-style decomposition that optimizes at the cluster level outperforms action-level parametrization. While objective-aware parametrization mitigates the optimization pathologies of Proposi- tions 3 and 4 by reducing K eff , it only treats the symptoms without curing the underlying non-concavity. In the next section, we propose a more fundamental shift: abandoning value estimation in favor of inherently tractable objectives. 7.2 Analysis of PWLL objectives To overcome the optimization challenges of IPS-based objectives, we consider policy- weighted log-likelihood (PWLL) objectives. These methods trade accurate value esti- mation for a well-behaved, concave optimization landscape, leading to more robust and effective policy learning. General form. Given a positive weighting function g(r,p 0 ), the PWLL objective is: ˆ U g (π) = 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i )) logπ(A i | X i ).(7.18) 92 The key motivation behind PWLL is to replace the linear dependence on the policy in IPS-based estimators, responsible for plateaus and local maxima in Section 7.1, with a concave transformation. Softmax policies are parametrized through scores s θ (x,a), and the map s ↦→ log softmax(s) is concave. Consequently, for common linear parametriza- tions in Equation (7.17), the composition logπ θ (a | x) is concave in θ. This removes the optimization pathologies inherent to IPS-based objectives. Proposition 5 (proof in Section D.2.) formalizes this advantage. Proposition 5. For linear softmax policies π θ , the PWLL objective ˆ U g (π θ ) is concave in θ. With ℓ 2 regularization, it is strongly concave. Proposition 5 makes PWLL appealing for stochastic optimization. In Section D.3, we show that under standard assumptions of bounded feature norms ∥φ(x,a)∥ and weights g(R i ,π 0 (A i ,X i )), these objectives satisfy the regularity conditions necessary to invoke established convergence theorems (Garrigos and Gower, 2023). This allows us to derive problem-dependent convergence guarantees: stochastic gradient ascent attains a global O(1/ √ T ) rate in the general (concave) case (Proposition 12), accelerating to a geometric rate under ℓ 2 -regularization (Proposition 13). Beyond optimization properties, PWLL also admits a simple statistical interpretation. ˆ U g (π) in Equation (7.18) is a weighted log-likelihood: the term logπ(A i | X i ) performs standard behavior cloning, while the weight g(R i ,π 0 (A i | X i )) determines how desirable 3 each logged sample is. This turns off-policy learning into a form of logging-aware and reward-weighted maximum-likelihood estimation. Different choices of g encode different notions of desirability. For example, the weighting g(r,π 0 (a| x)) = r maxπ 0 (a| x),τ emphasizes samples with high reward while reducing the influence of actions that the logging policy selected very frequently. At the same time, the clipping at τ prevents ex- tremely rare actions from receiving disproportionately large weights, ensuring that their contribution is attenuated once π 0 (a | x) falls below the threshold. In this view, desir- able samples are those that provide strong reward evidence without allowing very small propensities to dominate the updates. Many other PWLL variants arise from different choices of g (see below), each specifying a distinct prioritization scheme for the logged data, while all benefit from the concavity induced by the logarithmic term. To illustrate the qualitative difference between PWLL and IPS-based objectives, we con- struct a simple offline bandit problem with K = 3 actions and visualize the resulting optimization landscapes in a two-parameter policy space. Concretely, we consider a non- contextual setting with deterministic mean rewards r = (0.9, 0.7, 0.2) and a logging policy π 0 whose support places almost all mass on action 3 (π 0 (1) = 0.002, π 0 (2) = 0.003, π 0 (3) = 0.995). We generate a fixed dataset of n = 60 logged samples (A i ,R i ) by draw- ing actions A i ∼ π 0 and binary rewards from the corresponding Bernoulli distributions, R i ∼ Bern(r(A i )). To obtain a two-dimensional visualization, we parameterize the target 3 By how desirable an action is, we mean how strongly this action should influence the learned policy. 93 15 10 5 0 5 10 15 1 15 10 5 0 5 10 15 2 0 50 100 150 200 250 PWLL Landscape (3D View) (a) PWLL (3D view) 15105051015 1 15 10 5 0 5 10 15 2 PWLL Landscape (2D Projection) Global Min (4.81, 4.41) 10 35 60 85 110 135 160 185 210 235 Loss (b) PWLL (2D projection) 15 10 5 0 5 10 15 1 15 10 5 0 5 10 15 2 1.8 1.6 1.4 1.2 1.0 0.8 0.6 0.4 0.2 OPE Landscape (3D View) (c) IPS-based (3D view) 15105051015 1 15 10 5 0 5 10 15 2 OPE Landscape (2D Projection) Global Min (-1.01, -0.60) 1.76 1.56 1.36 1.16 0.96 0.76 0.56 0.36 0.16 Loss (d) IPS-based (2D projection) Figure 7.1: Optimization landscapes on a toy example. PWLL (cLPI) vs IPS-based (cIPS). policy using a softmax over three logits: π θ (a) = e θ a / ∑︁ b∈[3] e θ b , fixing the logit associated with action 3 as θ 3 = 1, and letting the remaining two logits be free parameters (θ 1 ,θ 2 ). In Figure 7.1, the PWLL landscape is concave with well-scaled gradients, and optimization trajectories converge reliably from roughly any initialization. In contrast, the IPS-based landscape consists of flat regions, separated by a narrow band of extremely steep curva- ture. This creates both vanishing and exploding gradients, severe ill-conditioning, and high sensitivity to initialization and learning rate. This aligns with the optimization pathologies in Propositions 3 and 4. Remark 7 (Beyond linear-softmax policies). The concavity guarantee of Proposition 5 assumes linear-softmax policies. In many large-scale recommendation systems, a deep encoder is pre-trained and kept fixed, and only a final linear head is optimized for the downstream task; in this case, the policy is still linear-softmax in the trainable parameters, and PWLL objectives retain their concavity. When the full network is trained end-to- end, concavity no longer holds. Yet, PWLL’s gradients g(R i ,π 0 (A i |x i ))∇ θ logπ θ (A i | X i ) match the structure of cross-entropy gradients, which are known to produce stable and well-scaled updates in deep architectures. Thus, even without formal guarantees, PWLL maintains substantially more benign optimization dynamics than IPS-based objectives. Local policy improvement (LPI). Liang and Vlassis (2022) set g(r,p 0 ) = r, which optimizes the log-likelihood of actions weighted by their observed rewards: ˆ U lpi (π) = 1 n n ∑︂ i=1 R i logπ(A i | X i ).(7.19) The oracle policy balances reward-seeking with imitation of the logging policy: π LPI ∗ (a| x)∝ r(x,a)π 0 (a| x).(7.20) Clipped LPI (cLPI). uses importance-weight clipping, setting g(r,p 0 ) = r max(p 0 ,τ ) : ˆ U clpi (π) = 1 n n ∑︂ i=1 R i maxπ 0 (A i | X i ),τ logπ(A i | X i ).(7.21) 94 In a similar spirit to cIPS, its oracle policy corrects for action frequency under π 0 , down- weighting the influence of rare actions due to the clipping: π cLPI ∗ (a| x)∝ r(x,a) π 0 (a| x) maxπ 0 (a| x),τ .(7.22) KL regularization (RegKL). To further amplify the reward signal relative to the logging policy prior, RegKL uses an exponential weighting function g(r,p 0 ) = exp(r/β): ˆ U regkl (π) = 1 n n ∑︂ i=1 exp(R i /β) logπ(A i | X i ).(7.23) The oracle policy is proportional to the logging policy, weighted by the exponentiated reward: π RegKL ∗ (a| x)∝ E r∼p(·|x,a) [︁ exp(r/β) ]︁ π 0 (a| x).(7.24) The temperature parameter β smoothly interpolates between behavior cloning (β →∞) and greedy reward maximization (β → 0). Note that BPR (Rendle et al., 2012) can be seen as an approximate PWLL objective, and we included it in our experiments. In fact, this general form of PWLL lends itself to numerous variations by modifying the weighting function g. For instance, one could introduce variants inspired by regularized IPS like exponential smooting (Chapter 8). While many such variants can be proposed for specific use cases, the central message of our work is that the well-behaved optimization landscape of the PWLL family is of greater practical importance than the estimation accuracy of IPS-based objectives. Thus, an exploration of these PWLL variants is beyond our scope. We contend that the foundational methods analyzed above, LPI, cLPI, and RegKL, along with the widely used BPR are sufficient to demonstrate the inherent advantages of PWLL objectives. Finally, PWLL resembles reward- or advantage-weighted behavioral cloning objectives in RL (Nair et al., 2020; Wang et al., 2020; Peng et al., 2019; Peters, 2006). While those methods address multi-step MDPs and often focus on mitigating distributional shift and bootstrapping errors, we focus on offline contextual bandits with large action spaces: identifying objectives and parametrizations that remain optimizable as K grows, rather than accurately estimating V (π). PWLL is critic-free and uses logged rewards and propensities through a weighting function g(R i ,π 0 (A i | X i )) that induces concave optimization landscapes for common policy classes. This yields substantial gains in large- K bandits without the overhead of value-function estimation. PWLL’s optimization- centric perspective complements the usual KL-regularized or trust-region interpretations of these RL methods. 7.3 Empirical Analysis We conduct our empirical evaluation on three large-scale recommendation datasets: MovieLens (K = 60k) (Lam and Herlocker, 2016), Twitch (K = 200k) (Rappaz et al., 2021), and 95 GoodReads (K = 1M) (Wan et al., 2019). These benchmarks feature action spaces with up to one million items, representing some of the largest settings studied in the offline policy learning literature. For all experiments, we employ the common softmax inner-product policies. We compare methods from both objective families. For IPS-based objectives, we include IPS, ES, DR, MIPS, OffCEM, POTEC, and PC in Section 7.1. For PWLL objectives, we evaluate LPI, cLPI, RegKL, and BPR in Section 7.2. All implementation details are provided in Section D.4. 7.3.1 Optimization is the Main Bottleneck To test our central hypothesis that optimization challenges are a more significant barrier than estimation accuracy, we evaluate how objectives perform under various optimization configurations. If an algorithm’s success is highly dependent on specific hyperparameters like batch size or learning rate, it suggests a difficult, non-robust optimization landscape. This experiment directly probes the practical trainability of each method, a key aspect our paper argues is often overlooked. The results strongly support our claim. As shown in Figure 7.2, IPS-based objectives are highly sensitive to batch size and learning rate schedule: minor changes can cause performance collapse, making them difficult to tune and train reliably. In contrast, PWLL objectives remain robust, achieving consistently high reward across all configurations. This stability translates directly into better learned policies: PWLL objectives outperform IPS- based objectives on all datasets. Even POTEC, a state-of-the-art method designed for large action spaces, is surpassed by the much simpler and easier-to-optimize cLPI. One might assume that an objective designed for estimation fidelity, such as a low-MSE IPS-based estimator, would naturally yield a better policy. Our findings show this is not the case. The superiority of PWLL objectives, which are poor value estimators by design, provides compelling evidence against this estimator-centric view. This reinforces our main takeaway: in large action space settings, a tractable optimization landscape is a more critical feature for a learning objective than its statistical accuracy. For completeness, an experiment tracking the MSE of methods is given in Section D.4. The figure also supports Claim 2. Indeed, there is a consistent performance gap between POTEC and OffCEM. Both methods are designed to maximize the same asymptotic objec- tive as we show in Section 7.1; their statistical goals are identical. The divergence in performance, therefore, can be attributed entirely to their differing optimization strate- gies. POTEC’s use of a two-stage, cluster-level optimization proves far more effective than OffCEM’s naive, action-level parametrization. 7.3.2 Objective-Aware Parametrization To empirically validate Claim 1, we compare a naive, whole-action-space parametrization against our proposed objective-aware approach, which restricts the policy’s effective action space to the logging policy support, S 0 . As shown for the IPS objective in Figure 7.3, the naive approach is highly unstable, with performance collapsing under simple learning configurations. In contrast, the objective-aware version is very robust, achieving high reward consistently across all batch sizes and schedules. This benefit extends even to 96 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Reward on MovieLens (K=60K) LR Schedule: Warmup Cosine 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 LR Schedule: One Cycle 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 LR Schedule: None 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Reward on Twitch (K=200K) 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Reward on GoodReads (K=1M) 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Effect of Optimization Hyperparameters on Performance IPS ES DR MIPS OffCEM PC POTEC LPI cLPI RegKL BPR Figure 7.2: Effect of batch size and learning rate schedule on final validation reward using three large-scale datasets. IPS-based objectives are highly sensitive, while PWLL objectives are robust. 97 2 4 2 5 2 6 2 7 2 8 Batch Size 0.10 0.15 0.20 0.25 0.30 Reward - MovieLens (K=60K) LR Schedule: None 2 4 2 5 2 6 2 7 2 8 Batch Size 0.10 0.15 0.20 0.25 0.30 LR Schedule: Warmup Cosine 2 4 2 5 2 6 2 7 2 8 Batch Size 0.10 0.15 0.20 0.25 0.30 LR Schedule: One Cycle Effect of Objective-Aware Parametrization on Performance IPS (Objective-Aware Defined on Support)IPS (Whole Action Space)cLPI (Objective-Aware Defined on Support)cLPI (Whole Action Space) Figure 7.3: The effect of objective-aware parametrization for IPS and cLPI on MovieLens. inherently stable PWLL objectives like cLPI, which achieve even better performance with the restricted support. This provides strong evidence for Claim 1: aligning the policy structure with the objective’s inductive bias simplifies the optimization landscape, leading to greater stability and superior learned policies. This finding holds across all datasets, with full results available in Section D.4. 7.4 Conclusion The dominant approach to off-policy learning focuses on developing sophisticated IPS- based estimators while neglecting a crucial factor: the optimization landscape. We demon- strated, both theoretically and empirically, that this landscape becomes prohibitively dif- ficult to optimize in large action spaces, undermining the practical effectiveness of even state-of-the-art estimators. Our analysis motivates two strategies. First, objective-aware policy parametrizations align the policy class with the estimator’s inductive bias, reducing the effective search space. Second, PWLL objectives abandon value estimation entirely in favor of inherently concave optimization landscapes. Experiments confirm that this focus on optimization tractability yields more robust learning, reduced sensitivity to hyperparameters, and superior policies. Our work has several limitations. First, PWLL objectives are not value estimators: they cannot be used for off-policy evaluation or policy selection (choosing the best policy from a finite candidate set) when accurate value estimates and their comparison are required. Second, the concavity guarantee of Proposition 5 holds only for linear-softmax policies; when training deep networks end-to-end, PWLL retains favorable gradient structure but loses formal concavity guarantees, although IPS-based objectives face even more severe optimization challenges in this setting. Third, PWLL’s oracle policies inherently depend on the logging policy (e.g., π LPI ∗ ∝ r(x,a)π 0 (a | x)), which may be suboptimal when π 0 has poor coverage of high-reward actions; however, this limitation is shared by IPS-based methods, whose oracle policies similarly depend on π 0 ’s support. 98 Chapter 8 Principled Pessimism for Exponential Smoothing and Beyond Contents 7.1 Analysis of IPS-Based Objectives . . . . . . . . . . . . . . . . . . . . . 87 7.1.1 Standard IPS-Based Objectives . . . . . . . . . . . . . . . . . 87 7.1.2 Large-Scale IPS-Based Objectives . . . . . . . . . . . . . . . . 89 7.1.3 Optimization Challenges . . . . . . . . . . . . . . . . . . . . . 90 7.2 Analysis of PWLL objectives . . . . . . . . . . . . . . . . . . . . . . . 92 7.3 Empirical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95 7.3.1 Optimization is the Main Bottleneck . . . . . . . . . . . . . . 96 7.3.2 Objective-Aware Parametrization . . . . . . . . . . . . . . . . 96 7.4 Conclusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98 Having explored structured direct methods in Chapter 6 and optimization-focused objec- tives in Chapter 7, we now turn to inverse propensity scoring (IPS). Despite its current practical limitations in large action spaces, many practitioners remain committed to IPS- based methods for their unbiasedness and theoretical guarantees, which enable principled safe off-policy learning. In this chapter, we improve IPS through exponential smoothing, a differentiable importance-weight regularization technique that enables a controlled bias- variance trade-off. Then, we adopt the pessimistic framework introduced in Section 5.1, deriving principled uncertainty penalties for our regularized estimators for safe policy learning. Prior work on pessimistic off-policy learning has derived objectives from gen- eralization bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019), but these approaches suffer from critical limitations: (i) they provide only one-sided bounds that fail to control estimation error in absolute value, limiting their ability to certify esti- mator quality, (i) the resulting bounds are intractable and incompatible with stochastic optimization, and (i) the pessimistic objectives require careful hyperparameter tuning. 99 We address these limitations by deriving tractable two-sided PAC-Bayes generalization bounds that can be optimized directly via stochastic gradient ascent. Unlike prior work (Sakhi et al., 2022), our analysis applies to standard IPS without assuming bounded importance weights, requiring only bounded second moments. Our bounds reveal that the optimal importance-weight smoothing parameter α depends on the quality of the logging policy. Furthermore, our framework generalizes to a broad class of importance- weight regularization techniques, yielding unified pessimistic objectives that enable fair comparison across different importance-weight regularization techniques. We present this extension in the final two sections of this chapter. The chapter is organized as follows. Section 8.1 presents background on regularized IPS and pessimism. Section 8.2 identifies the shortcomings of hard clipping and introduces our exponential smoothing estimators. Section 8.3 leverages PAC-Bayes theory to derive two-sided generalization bounds within the pessimistic framework. Section 8.4 discusses implications of our results. Section 8.5 demonstrates favorable performance across diverse benchmarks. Finally, Sections 8.6 and 8.7 extend the framework to other importance- weight regularizations and compare them under a unified pessimistic objective. 8.1 Background We consider the off-policy setting in Section 5.1, where we have access to logged data D n = (X i ,A i ,R i ) n i=1 collected by a known logging policy π 0 . As additional notation, we let μ π be the joint distribution of (X,A,R); μ π (x,a,r) = ν(x)π(a|x)p(r|x,a), so that (X i ,A i ,R i )∼ μ π 0 . Our goal remains to find a policy ˆπ ∈ Π that maximizes the value V (π) = E X∼ν,A∼π(·|X) [r(X,A)]. 8.1.1 Regularized IPS This chapter focuses on the IPS estimator (Horvitz and Thompson, 1952; Dudík et al., 2012), which estimates the value V (π) by re-weighting the samples as ˆ V ips (π) = 1 n n ∑︂ i=1 R i w(A i |X i ),(8.1) where w(a|x) = π(a|x)/π 0 (a|x) are the importance weights. While IPS provides an unbi- ased estimate of V (π) when the common support condition holds (i.e., π 0 (a|x) = 0 implies π(a|x) = 0), its variance grows with these importance weights (Swaminathan et al., 2017), which can be arbitrarily large when the target policy π and logging policy π 0 differ sig- nificantly. To mitigate this variance issue, it is common to transform the importance weights using regularization functions that introduce controlled bias to reduce variance. A regularized IPS estimator takes the form: ˆ V (π) = 1 n n ∑︂ i=1 R i ˆw(A i |X i ),(8.2) 100 where ˆw(a|x) ≤ w(a|x) are the regularized importance weights. A common importance- weight regularization approach is clipping where ˆw(a|x) = min( π(a|x) π 0 (a|x) ,M ), M > 0. 8.1.2 Pessimistic Objectives Within the pessimistic framework introduced in Section 5.1, we seek to maximize: ˆπ = argmax π∈Π [ ˆ V (π)−pen(π)] where pen(·) is a penalty term. The construction of this penalty has been approached in various ways, but generally relies on lower confidence bounds on the policy value: Evaluation bounds (Metelli et al., 2021) provide confidence intervals for a fixed target policy π, showing that with probability at least 1− δ: |V (π)− ˆ V (π)|≤ f (δ,π,π 0 ,n).(8.3) Essentially, Equation (8.3) indicates that for a fixed policy π ∈ Π, the event |V (π)− ˆ V (π)| ≤ f (δ,π,π 0 ,n) holds with high probability. However, this event depends on the target policy π. Thus Equation (8.3) is useful for evaluating a single target policy when having access to multiple logged data setsD n . This poses a problem for off-policy learning, where we optimize over a potentially infinite space of policies using a single logged data set D n . This is the fundamental theoretical limitation of using evaluation bounds similar to Equation (8.3) in off-policy learning. While one can transform Equation (8.3) into a generalization bound that holds uniformly over all π ∈ Π via a union bound, this typi- cally introduces intractable complexity terms, making the resulting pessimistic objectives, which maximize the lower confidence bound, equally intractable. One-sided generalization bounds (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Sakhi et al., 2022) address this limitation by providing bounds that hold simultaneously for all policies ∈ Π. For δ ∈ (0, 1), with probability at least 1− δ: V (π)≥ ˆ V (π)− g(δ, Π,π,π 0 ,n), ∀π ∈ Π,(8.4) where the function g now depends on the policy space Π. This leads to the pessimistic objective: ˆπ = argmax π∈Π ˆ V (π)− g(δ, Π,π,π 0 ,n).(8.5) However, one-sided bounds fail to attest to the quality of the estimator. To illustrate, consider a degenerate estimator ˆ V poor (π) = 0 for all π ∈ Π. Since V (π) ∈ [0, 1], we trivially have a one-sided bound V (π) ≥ ˆ V poor (π) with probability 1, yet this estimator is entirely uninformative about the true rewards. Two-sided generalization bounds resolve this issue by controlling both the upper and lower deviations leading to |V (π)− ˆ V (π)|≤ g(δ, Π,π,π 0 ,n), ∀π ∈ Π.(8.6) These bounds ensure the estimator quality and enable oracle inequalities of the form V (ˆπ) ≥ V (π ∗ ) − 2g(δ, Π,π ∗ ,π 0 ,n), where ˆπ is learned using Equation (8.5) and π ∗ = 101 argmax π∈Π V (π) is the optimal policy. This shows the appeal of pessimism: the sub- optimality gap depends on the bound evaluated at the optimal policy π ∗ , meaning the estimator only needs to be precise for near-optimal policies rather than uniformly across the policy class Π. Alternative approaches include heuristics that simplify theoretical bounds for tractabil- ity (Swaminathan and Joachims, 2015a; London and Sandler, 2019; Wang et al., 2023), often penalizing by empirical variance or policy divergence while discarding complex- ity terms. Recent work on implicit pessimism (Gabbianelli et al., 2024; Sakhi et al., 2024) (published after the work in this chapter) shows that careful analysis of spe- cific importance-weight regularizations can yield bounds where the penalty is policy- independent. In this case, maximizing the lower bound reduces to maximizing the es- timator directly: the pessimism becomes implicit in the regularization itself. In this work, we derive a tractable two-sided PAC-Bayesian generalization bound for our exponential smoothing estimator and then generalize it to other importance-weight regularizations. 8.2 Exponential Smoothing Importance-weight clipping (Swaminathan and Joachims, 2015a) yields the following com- monly used estimators IPS-min ̃ V m (π) = 1 n n ∑︂ i=1 R i min (︁ w(A i |X i ),M )︁ , IPS-max ˆ V τ (π) = 1 n n ∑︂ i=1 R i π(A i |X i ) max(π 0 (A i |X i ),τ ) .(8.7) Here IPS-min clips the weights while IPS-max only clips π 0 in the denominator since π is always smaller than 1. For instance, M ∈ R + in ̃ V m (π) trades the bias and variance of the estimator. When M is large, the bias of ̃ V m (π) is small but its variance may be large. On the other hand, the variance goes to 0 when M ≈ 0 since in that case ̃ V m (π)≈ 0 for any π ∈ Π. Similarly, τ ∈ [0, 1] trades the bias and variance of ˆ V τ (π) and can be seen as τ ≈ 1 M . This hard clipping has some limitations. First, min(·,M ) leads to non-differentiable objectives that may require additional care in optimization (Papini et al., 2019). Also, min(·,M ) is constant on [M,∞) leading to objectives with zero gradients for any policy π that satisfies w(A i |X i ) > M for any i∈ [n]. More importantly, hard clipping is sensitive to the choice of the clipping threshold M. In practice, tuning M is challenging and may cause the learned policy to match the logging policy, leading to minimal improvements. To see this, consider the following illustrative example. For simplicity, suppose that the problem is non-contextual, in which case the reward function r only depends on the actions a ∈ A. It follows that policies do not depend on x ∈ X ; they are now probability distributions π(·) over A. Also, assume that A = [100] and that the reward received after taking action a ∈ [100] is binary. That is, R ∼ 102 Bern(r(a)) where r(a) = 0.1− 10 −3 (a− 1) is the expected reward of action a, and for any p ∈ [0, 1], Bern(p) is the Bernoulli distribution with parameter p. This means that the best action is 1 and the worst is 100. Finally, the logging policy π 0 (·) is ε-greedy centered at action 50. That is π 0 (50) = 1− ε, and for any a̸= 50, π 0 (a) = ε 99 , with ε = 0.05. Now consider 100 deterministic policies π a (·) for a ∈ [100] such that π a (·) is the Dirac distribution centered at a. In Figure 8.1, we plot the estimated reward of the policies π a using either IPS in Equation (8.1) or IPS-min in Equation (8.7). We generate n = 50k samples and set M = 100 = O( √ n) as suggested by Ionides (2008). With this choice of M, IPS-min underestimates the reward of all policies π a for a ̸= 100 since their weights π a /π 0 are either 0 or 99/ε > M. The estimated reward of IPS-min is maximized in π 50 ≈ π 0 only. Thus, if we optimize ̃ V m (·) over Dirac policies, we will converge to the logging policy despite its bad performance. Although the other variant of hard clipping, IPS-max in Equation (8.7), is differentiable, it is still sensitive to τ and may induce high bias similar to Figure 8.1. This is due to some loss of information related to the preferences of the logging policy. Indeed, for two actions a and a ′ such that π 0 (a|X i ) ≪ π 0 (a ′ |X i ) < τ for an observed context X i , the propensity scores π 0 (a|X i ) and π 0 (a ′ |X i ) will be clipped to the same value τ. Thus the information that, for context X i , action a ′ is preferred by the logging policy than action a will be lost. 020406080100 0.00 0.05 0.10 0.15 0.20 0.25 true ips ips-min, M=100 Figure 8.1: Effect of hard clipping on the estimation quality. The x-axis corresponds to actions a∈ [100]. The y-axis is the estimated reward of each of the 100 policies π a using either IPS or IPS-min. The cyan line is the true reward for each policy π a . To mitigate this, we propose the following exponential smoothing correction for IPS. Our estimators are defined as IPS-α : ˆ V α (π) = 1 n n ∑︂ i=1 R i ˆw α (A i |X i ), α∈ [0, 1], IPS-β : ̃ V β (π) = 1 n n ∑︂ i=1 R i ̃w β π (A i |X i ), β ∈ [0, 1],(8.8) 103 where ˆw α (a|x) = π(a|x) π 0 (a|x) α and ̃w β π (a|x) = π(a|x) β π 0 (a|x) β . Here standard IPS is recovered for α = 1 and β = 1. These estimators yield smooth, everywhere-differentiable objectives and avoid the flat regions induced by hard clipping; this improves optimization in practice. Also, in contrast with IPS-max in Equation (8.7), ˆ V α (π) preserves the preferences of the logging policy. Precisely, for two actions a and a ′ such that π 0 (a|X i ) < π 0 (a ′ |X i ) for an observed context X i , we still have π 0 (a|X i ) α < π 0 (a ′ |X i ) α and the information that action a ′ is preferred by the logging policy than action a is preserved. While a similar correction to IPS-β was proposed in Korba and Portier (2022), its use in off-policy learning is novel. Also, Su et al. (2020); Metelli et al. (2021) regularized the importance weights w as λ 1 w λ 1 +w 2 ,λ 1 > 0 and w 1−λ 2 +λ 2 w ,λ 2 ∈ [0, 1], respectively. Thus, the expression of both corrections is very different from ours. More importantly, these corrections entail different properties than ours. Roughly speaking, our correction allows us to simultaneously (1) control a tuning parameter α∈ [0, 1] that is in a bounded domain [0, 1], (2) without constraining the resulting importance weights to be bounded, (3) and to obtain tractable PAC-Bayes generalization bounds as the correction π π α 0 is linear in π; a technical requirement of PAC-Bayes analysis. In contrast, Metelli et al. (2021); Su et al. (2020) do not provide generalization guarantees; they focus on estimation accuracy (e.g., through mean squared error) and only propose heuristics for off-policy learning. Those heuristics are not based on theory, in contrast with ours which is directly derived from our generalization bound. Also, our approach has favorable empirical performance (Section E.3.6). Although Korba and Portier (2022, Lemma 1) show that smoothing the importance weights similarly to IPS-β in Equation (8.8) reduces the variance, it might still be unclear how α and β trade the bias and variance of our estimators in off-policy learning. To see this, let α∈ [0, 1], then we have |B( ˆ V α (π))|≤ E X∼ν,A∼π(·|X) [︁ 1− π 0 (A|X) 1−α ]︁ ,(8.9) V [︂ ˆ V α (π) ]︂ ≤ 1 n E X∼ν,A∼π(·|X) [︁ π(A|X) π 0 (A|X) 2α−1 ]︁ , with B( ˆ V α (π)) = E[ ˆ V α (π)]−V (π) and V[ ˆ V α (π)] = E[( ˆ V α (π)−E[ ˆ V α (π)]) 2 ] are respectively the bias and the variance of ˆ V α (π). The bound of the bias in Equation (8.9) is minimized in α = 1 (standard IPS); in which case it is equal to 0 (standard IPS is unbiased). In contrast, the bound of the variance is minimized in α = 0. Thus if the variance is small or n is large enough such that E[π(A|X)/π 0 (A|X) 2α−1 ]/n → 0, then we set α → 1. Otherwise, we set α → 0. This shows that α trades the bias and variance of ˆ V α . More details and a similar discussion for ̃ V β (π) are deferred to Section E.1. 8.3 PAC-Bayes Analysis for Off-Policy Learning We now derive generalization bounds for our estimator. We opt for the PAC-Bayes frame- work for the following reasons. First, it is known to provide some of the tightest gen- eralization bounds in challenging scenarios (Farid and Majumdar, 2021), for aggregated and randomized predictors (Alquier, 2021). Second, the bounds have a Kullback–Leibler (KL) divergence (Van Erven and Harremos, 2014) term D KL (Q∥P) that depends on a 104 fixed prior P and a learning posterior Q (see Section 8.3.1 for a brief introduction). This quantity can be seen as a complexity measure, similarly to the covering number (Maurer and Pontil, 2009). The difference is that complexity measures are uniform on the space of policies while the KL term in PAC-Bayes depends on the prior P and the posterior Q. This allows getting sharper bounds when the former is well chosen. Third, the PAC-Bayes perspective fits very well with off-policy learning. In fact, a policy π can be written as an aggregation of predictors under some distribution Q. Thus the prior P can be associated with the logging policy π 0 that we want to improve upon while the posterior Q is related to the learning policy π. Fourth, London and Sandler (2019) showed that PAC-Bayes can lead to tractable and scalable objectives, an important consideration for this thesis. 8.3.1 Elements of PAC-Bayes Let Z = X ×Y be an instance space: e.g., X and Y are the input and output space in supervised learning. Let H = h :X →Y denote a hypothesis space of mappings from X to Y (predictors). Also, let L : H×Z → R be a loss function and assume access to data D n = (Z i ) i∈[n] drawn from an unknown distribution D. Let Risk(h) = E Z∼D [L(h,Z)] be the risk of h∈H while ˆ︃ Risk n (h) = 1 n n ∑︂ i=1 L(h,Z i ) is its empirical counterpart. Then the main focus in PAC-Bayes is to study the gen- eralization capabilities of random hypotheses Q on H by controlling the gap between the expected risk under Q, E h∼Q [Risk(h)], and the expected empirical risk under Q, E h∼Q [︂ ˆ︃ Risk n (h) ]︂ . For example, assume that L(h,Z) ∈ [0, 1] for any (h,Z) ∈ H×Z, let P be a fixed prior distribution on H and let δ ∈ (0, 1). Then with probability at least 1− δ over D n ∼ D n , the following inequality holds simultaneously for any posterior distribution Q on H: E h∼Q [Risk(h)]≤ E h∼Q [︂ ˆ︃ Risk n (h) ]︂ + √︄ D KL (Q∥P) + log 2 √ n δ 2n . This was originally proposed by McAllester (1998), and the reader may refer to Alquier (2021); Guedj (2019) for more elaborate introductions of PAC-Bayes theory. Connection to value functions. The loss L is often chosen as the negative reward, L(h,Z) =−r(h,Z). In this case, minimizing the Risk(h) is equivalent to maximizing the value function. Thus, PAC-Bayes bounds on the risk directly translate into guarantees on the discrepancy between empirical and true value, providing a principled way to reason about generalization in off-policy learning. 8.3.2 PAC-Bayes for Off-Policy Learning Let H = h : X → A be a hypothesis space of mappings from X (contexts) to A (actions). Given a policy π and a context x∈X , the action distribution π(·|x) is induced 105 by a distribution Q over H (London and Sandler, 2019) such as π(a|x) = π Q (a|x) = E h∼Q [︁ 1 h(x)=a ]︁ .(8.10) This is not an assumption since any policy π has this form when H is rich enough (Sakhi et al., 2022, Theorem 2). From Equation (8.10), we observe that policies can be seen as an aggregation E h∼Q [·] (under some distribution Q on the pre-defined hypothesis space H) of deterministic decision rules 1 h(x)=a . This allows formulating off-policy learning as a PAC-Bayes problem. Before showing how this is achieved, we start by providing two practical policies of such form. Example 1 (softmax and mixed-logit policies). We define the hypothesis space H = ︁ h θ,γ ;θ ∈ R dK ,γ ∈ R K ︁ of mappings h θ,γ (x) = argmax a∈A φ(x) ⊤ θ a + γ a . Here φ(x) outputs a d-dimensional representation of x, and γ a is a standard Gumbel perturbation, γ a ∼ G(0, 1) for any a∈A. Then π sof θ (a|x) = exp(φ(x) ⊤ θ a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ θ a ′ ) , (i) = E γ∼G(0,1) K [︁ 1 h θ,γ (x)=a ]︁ ,(8.11) where (i) follows from the Gumbel-Max trick (GMT) (Luce, 2012; Maddison et al., 2014). Thus a softmax policy π sof θ can be written as in Equation (8.10). Now we also consider random parameters θ ∼ N (μ,σ 2 I dK ) with μ ∈ R dK and σ > 0. Then, let Q = N (μ,σ 2 I dK )× G(0, 1) K , it follows that π Q = π mixL μ,σ is a mixed-logit policy and it reads π mixL μ,σ (a|x) = E θ∼N (μ,σ 2 I d ) [︃ exp(φ(x) ⊤ θ a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ θ a ′ ) ]︃ , = E θ∼N (μ,σ 2 I d ),γ∼G(0,1) K [︁ 1 h θ,γ (x)=a ]︁ .(8.12) Example 2 (Gaussian policies): Sakhi et al. (2022) removed the Gumbel noise γ in Equation (8.12) and consequently defined the hypothesis space as H = ︁ h θ ;θ ∈ R dK ︁ of mappings h θ (x) = argmax a∈A φ(x) ⊤ θ a for any x ∈ X. Then, let Q = N (μ,σ 2 I dK ), it follows that π Q = π gaus μ,σ reads π gaus μ,σ (a|x) = E θ∼N (μ,σ 2 I d ) [︁ 1 h θ (x)=a ]︁ .(8.13) To see why removing the Gumbel noise can be beneficial, the reader may refer to Sec- tion E.3.2. After motivating the definition of policies in Equation (8.10), we are in a position to relate our estimators to the general PAC-Bayes framework in Section 8.3.1. One technical requirement of our proof is that the estimator should be linear in π. Thus we focus on ˆ V α (·) since ̃ V β (π) is non-linear in π. Let h ∈ H, x ∈ X , a ∈ A and r ∈ [0, 1], we define the objective U α as U α (h,x,a,r) = 1 h(x)=a π 0 (a|x) α r .(8.14) 106 Using the definition in Equation (8.10) and the linearity of the expectation, we have that ˆ V α (·) in Equation (8.8) can be written as ˆ V α (π Q ) = E h∼Q [︄ 1 n n ∑︂ i=1 U α (h,X i ,A i ,R i ) ]︄ . Moreover, the expectation of ˆ V (π Q ) reads V α (π Q ) = E h∼Q E (X,A,R)∼μ π 0 [U α (h,X,A,R)] . Finally, the main quantity of interest, the value V (π Q ), can be expressed in terms of the objective with α = 1, U 1 , as V (π Q ) = E h∼Q E (X,A,R)∼μ π 0 [U 1 (h,X,A,R)] . Since ˆ V α (π Q ) is an unbiased estimator of V α (π Q ), PAC-Bayes can be used to bound V α (π Q )− ˆ V α (π Q ). This will allow bounding our quantity of interest V (π Q )− ˆ V α (π Q ). 8.3.3 Main Result To ease the exposition, we assume that the rewards are deterministic. Then, in logged data D n , R i = r(X i ,A i ) for any i ∈ [n]. Note that the same result holds for stochastic rewards. We discuss our result and sketch its proof in Section 8.4. The complete proof can be found in Section E.2.1. Theorem 4. Let λ > 0, n≥ 1, δ ∈ (0, 1), α∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1−δ over draws D n ∼ μ n π 0 , the following holds simultaneously for any posterior Q on H |V (π Q )− ˆ V α (π Q )|≤ √︃ kl 1 (π Q ) 2n + B α n (π Q ) + kl 2 (π Q ) nλ + λ 2 Var α n (π Q ). where kl 1 (π Q ) = D KL (Q∥P) + ln 4 √ n δ , and kl 2 (π Q ) = D KL (Q∥P) + ln 4 δ , B α n (π Q ) = 1− 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ π 1−α 0 (A|X i ) ]︁ , Var α n (π Q ) = 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + π Q (A i |X i )R 2 i π 0 (A i |X i ) 2α . We start by clarifying that the prior P can be any fixed distribution on H. If we have access to P 0 on H such that π 0 = π P 0 , then it is natural to set P = P 0 . But this is just a choice and one may use priors that do not depend on π 0 . Now we explain the main terms in our bound. First, the terms kl 1 (π Q ) and kl 2 (π Q ) contain the divergence D KL (Q∥P) which penalizes posteriors Q that differ a lot from the prior P. Moreover, B α n (π Q ) is the bias conditioned on the contexts (X i ) i∈[n] ; B α n (π Q ) = 0 when α = 1 and B α n (π Q ) > 0 otherwise. Also, the first term in Var α n (π Q ) resembles the theoretical second moment of 107 the regularized importance weights π π α 0 (without the reward) when they are seen as random variables. Similarly, the second term in Var α n (π Q ) resembles the empirical second moment of π π α 0 R (with the reward). Finally, if Var α n (π Q ) is bounded, then we can set λ = 1/ √ n, in which case our bound scales as O(1/ √ n + B α n (π Q )). In practice, we set α≈ 1 leading to B α n (π Q )≈ 0 and the bound would scale as O(1/ √ n). One of the main strengths of our result is that it holds for standard IPS with α = 1 under the assumption that Var 1 n (π Q ) is bounded. This assumption is less restrictive than assuming that the importance weight as a random variable, π Q (A|X)/π 0 (A|X), is bounded, a required assumption for traditional concentration bounds. In contrast, Var α n (π Q ) only involves the expectations of the random variables π Q (A|X i )/π 0 (A|X i ) 2α , and ratios of π 0 evaluated at observed contexts and actions and (X i ,A i ) i∈[n] , that have non-zero probabilities under π 0 by definition. Our result holds for fixed λ > 0 and α∈ [0, 1]. In Section E.2.2, we extend this to any po- tentially data-dependent λ∈ (0, 1) and α∈ (0, 1]. The assumption that R∈ [0, 1] can be relaxed to R∈ [0,B] up to additional factors B 2 and B in Var α n (π Q ) and kl 1 (π Q ), respec- tively. Finally, our bound is suitable for stochastic gradient ascent (Robbins and Monro, 1951) since data-dependent quantities are not inside a square root. This is important for scalability. Limitations. Our bound in Theorem 4 has two main limitations. (i) Using it to directly derive a data-independent suboptimality gap bound is not straightforward. This difficulty arises because our bound involves empirical quantities such as B α n (π Q ) and Var α n (π Q ), whose dependence on the logged data prevents expressing the gap purely as a function of n. However, obtaining data-independent suboptimality guarantees was not the goal of this chapter. Instead, our focus was on deriving tractable and theoretically grounded bounds for exponential smoothing, that also perform well in practice when used for pessimistic ob- jectives. (i) Our result provides symmetric deviation bounds that simultaneously control the upper and lower deviations of ˆ V α (π Q ) from V (π Q ). Yet, recent work (Gabbianelli et al., 2024), published after the paper corresponding to this chapter, indicates that the tails of regularized IPS estimators are inherently asymmetric. Consequently, tighter bounds may arise from developing asymmetric two-sided bounds that treat each deviation separately. We explored this direction in our follow-up work (Sakhi et al., 2024), where we derived some of the tightest bounds in the literature, with strong empirical performance. 8.3.4 Adaptive and Data-Driven Tuning of α Theorem 4 assumes that α is fixed (although we extend it for data-dependent α in Sec- tion E.2.2). However, providing a procedure to tune α in an adaptive and data-dependent fashion is important in practice. Thus we propose to set α ∗ = argmin α∈[0,1] B α n (π Q ) + √︃ 2kl 2 (π Q ) Var α n (π Q ) n ,(8.15) where all the terms are defined in Theorem 4. Roughly speaking, α ∗ establishes a bias- variance trade-off; it minimizes the sum of the bias term B α n (π Q ) and the square root of the second moment term Var α n (π Q ), weighted by √︂ 2kl 2 (π Q ) n . Here Equation (8.15) is obtained 108 by minimizing the bound in Theorem 4 with respect to both α and λ as follows. First, we minimize the bound in Theorem 4 with respect to λ; the minimizer is λ ∗ = √︂ 2kl 2 (π Q ) n Var α n (π Q ) . Then, the bound in Theorem 4 evaluated at λ = λ ∗ becomes √︃ kl 1 (π Q ) 2n + B α n (π Q ) + √︃ 2kl 2 (π Q ) Var α n (π Q ) n .(8.16) Finally, α ∗ is defined as the minimizer of Equation (8.16) with respect to α ∈ [0, 1], and √︂ kl 1 (π Q ) 2n does not appear in Equation (8.15) as it does not depend on α. Note that α ∗ depends on both logged dataD n and the learning policy π Q . Thus it is adaptive; its value changes in each iteration during optimization. 8.4 Discussion We start by interpreting and comparing our results to related work. Then, we present the technical challenges in Section 8.4.2. After that, we sketch our proof in Section 8.4.3. 8.4.1 Interpretation and Comparison to Related Work Theorem 4 gives insight into the number of samples needed so that the performance of ˆπ is close to that of the optimal policy π ∗ . To simplify the problem, we consider the Gaussian policies in Equation (8.13) and assume that there exists Q ∗ =N (μ ∗ ,I dK ) with μ ∗ ∈ R dK such that the optimal policy is π ∗ = π Q ∗ . Also, we let the prior P = N (μ 0 ,I dK ) and assume that π 0 is uniform. This is possible since as we said before, the prior P does not have to depend on the logging policy π 0 . Then we have that D KL (Q ∗ ∥P) =∥μ ∗ −μ 0 ∥ 2 /2, B α n (π Q ∗ ) = 1−1/K 1−α and Var α n (π Q ∗ )≤ 2K 2α . The last inequality is not tight but it allows getting an easy-to-interpret term that does not depend on n. Now let ε > 2(1− K α−1 ) for α ∈ [1− log 2/ logK, 1]. This condition on α ensures that ε ∈ [0, 1] and it is mild as α is often close to 1. Then, it holds with high probability that n ˜︁ > (︂ ∥μ ∗ − μ 0 ∥ 2 + K 2α ε− 2(1− K α−1 ) )︂ 2 =⇒ V (ˆπ)≥ V (π Q ∗ )− ε, where we omit constant and logarithmic terms in ˜︁ >. This gives an intuition on the sample complexity for our procedure. In particular, fewer samples are needed in four cases. The first is when ε is large, which means that we afford to learn a policy whose performance is far from the optimal one. The second is when the prior P is close to Q ∗ , that is when ∥μ ∗ − μ 0 ∥ is small. This highlights that the choice of the prior P is important. The third is when the second-moment term K 2α is small. The fourth is when the bias B α n (π Q ∗ ) is small. In particular, when α = 1, the bias is 0. In contrast, the second-moment term is minimized in α = 0. This is where the choice of α matters. The proofs of these claims and more detail can be found in Section E.2.4. Our chapter derives a tractable generalization bound for an estimator other than clipped IPS in Equation (8.7), which also holds for the standard IPS in Equation (8.1). The bounds in Swaminathan and Joachims (2015a); London and Sandler (2019); Sakhi et al. 109 (2022) have a multiplicative dependency on the clipping threshold (M or 1/τ in Equa- tion (8.7)). Standard IPS is recovered when M →∞ (or τ = 0) in which case their bounds are infinite. We successfully avoid any similar dependency on α. Moreover, Swaminathan and Joachims (2015a); London and Sandler (2019) only used their generalization bounds to inspire pessimistic objectives. Although we directly optimize our theoretical bound (Theorem 4) in our experiments, our analysis also inspires a pessimistic objective where we simultaneously penalize the L 2 distance, the variance and the bias. That is, we find μ∈ R dK that maximizes ˆ V α (π μ )− λ 1 ∥μ− μ 0 ∥ 2 − λ 2 Var α n (π μ )− λ 3 B α n (π μ ).(8.17) Here λ 1 ,λ 2 and λ 3 are tunable hyper-parameters, π μ can be the Gaussian policy in Equa- tion (8.13), π μ = π gaus μ,1 , with a fixed σ = 1, and μ 0 is the mean of the prior P =N (μ 0 ,I dK ). Existing works either penalize the L 2 distance or the variance. For completeness, we also show that this pessimistic objective should be preferred over existing ones in Section E.3.5. 8.4.2 Technical Challenges London and Sandler (2019); Sakhi et al. (2022) derived PAC-Bayes generalization bounds for the estimator IPS-max in Equation (8.7). Extending their analyses to our case is not straightforward. First, their estimator IPS-max is upper bounded by 1/τ, and thus they relied on traditional techniques for [0, 1]-objectives (Alquier, 2021). In contrast, our objective in Equation (8.14) is not upper-bounded, and controlling it without assuming that the importance weights are bounded is challenging. Moreover, their bounds have a multiplicative dependency on 1/τ, hence they explode as τ → 0. This makes them vacuous for small values of τ and inapplicable to the standard IPS estimator in Equation (8.1) recovered for τ = 0. In contrast, our bound does not have a similar dependency on α and it is also valid for standard IPS recovered for α = 1. Moreover, we derive two-sided inequalities rather than one-sided ones for the important reasons that we priorly discussed. This requires carefully controlling in closed-form the absolute value of the bias. Prior works only used that the bias is negative which was enough to obtain one-sided inequalities. Explaining other challenges requires stating a result that inspired our analysis: Kuzborskij and Szepesvári (2019) derived PAC-Bayes generalization bounds for unbounded losses by only controlling their second moments. Recently, Haddouche and Guedj (2022) proposed a similar result using Ville’s inequality (Bercu and Touati, 2008). Adapting their theorem to our problem is given Proposition 6. We slightly adapt their proof to get a two-sided inequality for a negative loss. The proof is deferred to Section E.2.3. Proposition 6. Let λ > 0, n ≥ 1, δ ∈ (0, 1), α ∈ [0, 1] and let P be a fixed prior on H, then with probability at least 1−δ over drawsD n ∼ μ n π 0 , the following holds simultaneously for all posteriors, Q, on H |V α (π Q )− ˆ V α (π Q )|≤ D KL (Q∥P) + log 2 δ λn + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 2α 0 (A i |X i ) R 2 i + λ 2 E (X,A,R)∼μ π 0 [︃ π Q (A|X) π 2α 0 (A|X) R 2 ]︃ , (8.18) 110 There are two main issues with Proposition 6. First, the term E (X,A,R)∼μ π 0 [︁ π Q (A|X) π 2α 0 (A|X) R 2 ]︁ in Equation (8.18) is intractable. One could bound R 2 by 1, but the resulting term will still be intractable due to the expectation over the unknown distribution of contexts ν. Second, we need an upper bound of |V (π Q )− ˆ V α (π Q )| while Proposition 6 only provides one for |V α (π Q )− ˆ V α (π Q )|. Thus it remains to quantify the approximation error|V (π Q )−V α (π Q )|. This will also require computing an expectation over X ∼ ν, which is intractable. 8.4.3 Sketch of Proof for Theorem 4 We conclude by showing how the technical challenges above were solved. First, We de- compose V (π Q )− ˆ V α (π Q ) as V (π Q )− ˆ V α (π Q ) = I 1 + I 2 + I 3 ,where I 1 = V (π Q )− 1 n n ∑︂ i=1 V (π Q |X i ), I 2 = 1 n n ∑︂ i=1 V (π Q |X i )− 1 n n ∑︂ i=1 V α (π Q |X i ), I 3 = 1 n n ∑︂ i=1 V α (π Q |X i )− ˆ V α (π Q ), where V (π Q |X i ) = E A∼π Q (·|X i ) [r(X i ,A)] , V α (π Q |X i ) = E A∼π 0 (·|X i ) [︂ π Q (A|X i ) π 0 (A|X i ) α r(X i ,A) ]︂ . I 1 is the estimation error of the empirical mean of the value using n i.i.d. contexts (X i ) i∈[n] . This term is introduced to avoid the intractable expectation over X ∼ ν. Moreover, I 2 is the bias term conditioned on the contexts (X i ) i∈[n] and we bound it in closed- form. Finally, I 3 is the estimation error of the value conditioned on the contexts (X i ) i∈[n] . Again, this conditioning allows us to avoid the intractable expectation over X ∼ ν and to consequently bound |I 3 | by tractable terms. First, Alquier (2021, Theorem 3.3) yields that with probability at least 1− δ 2 , it holds for any Q on H that |I 1 |≤ √︄ D KL (Q∥P) + log 4 √ n δ 2n . Also, |I 2 | is bounded similarly to Equation (8.9) as |I 2 |≤ 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ 1− π 1−α 0 (A|X i ) ]︁ . Bounding|I 3 | is achieved by expressing it using martingale difference sequences (f i (A i ,h)) i∈[n] that we construct as follows. Let (F i ) i∈0∪[n] be a filtration adapted to (S i ) i∈[n] where S i = (A ℓ ) ℓ∈[i] for any i∈ [n], we define f i (A i ,h) = E A∼π 0 (·|X i ) [︃ 1 h(X i )=A r(X i ,A) π 0 (A|X i ) α ]︃ − 1 h(X i )=A i R i π 0 (A i |X i ) α . 111 Then we show that for any h ∈ H, (f i (A i ,h)) i∈[n] is a martingale difference sequence. After that, we apply Haddouche and Guedj (2022, Theorem 5) and obtain that with probability at least 1− δ/2, it holds for any Q on H that |E h∼Q [M n (h)]|≤ D KL (Q∥P) + log 4 δ λ + λ 2 E h∼Q [Var n (h)], where M n (h) = ∑︁ n i=1 f i (A i ,h) and Var n (h) = ∑︁ n i=1 f i (A i ,h) 2 +E [︁ f i (A i ,h) 2 |F i−1 ]︁ . Then notice that E h∼Q [M n (h)] can be expressed in terms of I 3 as E h∼Q [M n (h)] = n ∑︂ i=1 V α (π Q |X i )− n ˆ V α (π Q ) = nI 3 , Moreover, E h∼Q [Var n (h)] is bounded by n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + π Q (A i |X i ) π 0 (A i |X i ) 2α R 2 i . Thus with probability at least 1− δ 2 , it holds for any Q that |I 3 |≤ D KL (Q∥P) + log 4 δ nλ + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α R 2 i + λ 2n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ . Our result is obtained by bounding |I 1 | +|I 2 | +|I 3 |. One shortcoming of our analysis is that Var α n (π Q ) is not exactly and only resembles the sum of the theoretical and empirical second moments of our estimator. Precisely, the terms π Q /π 2α 0 should be π 2 Q /π 2α 0 . This problem arises due to our definition of the martingale difference sequences (f i (A i ,h)) i∈[n] in Equation (8.14). Precisely, in our proof, we compute the square f i (A i ,h) 2 . However, the square of an indicator function is the indicator function itself. Thus applying the expectation afterwards, E h∼Q [f i (A i ,h) 2 ], leads to π Q appearing instead of π 2 Q . This issue is inherent in the PAC-Bayes formulation and seminal works (London and Sandler, 2019; Sakhi et al., 2022) would suffer the same issue. Solving this would be beneficial and we leave it to future work. 8.5 Experiments for Exponential Smoothing We briefly present our experiments. More details and discussions can be found in Sec- tion E.3. We consider the standard supervised-to-bandit conversion (Agarwal et al., 2014) where we transform a supervised training setS tr n to a logged bandit dataD n as described in Algorithm 3 in Section E.3.1. Here the action space A is the label set and the con- text space X is the input space. Then, D n is used to train our policies. After that, we evaluate the value of the learned policies on the supervised test set S ts n ts as described in Algorithm 4 in Section E.3.1. Roughly speaking, the resulting value quantifies the ability of the learned policy to predict the true labels of the inputs in the test set. This is our performance metric; the higher the better. We use 4 image classification datasets MNIST (LeCun et al., 1998), FashionMNIST (Xiao et al., 2017), EMNIST (Cohen et al., 2017) and CIFAR100 (Krizhevsky et al., 2009). 112 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 EMNIST, K=47, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.00 0.05 0.10 0.15 0.20 0.25 0.30 CIFAR, K=100, d=2048 Ours, Gaussian Ours, Mixed-Logit London et al., Gaussian London et al., Mixed-Logit Sakhi et al. 1, Gaussian Sakhi et al. 1, Mixed-Logit Sakhi et al. 2, Gaussian Sakhi et al. 2, Mixed-Logit Logging Figure 8.2: The reward of the learned policy using one of the baselines with varying quality of the logging policy η 0 ∈ [0, 1]. The logging policy is defined as π 0 = π sof η 0 ·μ 0 in Equation (8.11), where μ 0 = (μ 0,a ) a∈A ∈ R dK and η 0 ∈ [0, 1] is the inverse-temperature parameter. The higher η 0 , the better the performance of π 0 . When η 0 = 0, π 0 is uniform. The parameters μ 0 are learned using 5% of the training set S tr n . In our experiments, we consider both, Gaussian and mixed- logit policies, in Equation (8.12) and Equation (8.13), for which we set the prior as P = N (η 0 μ 0 ,I dK ) and P = N (η 0 μ 0 ,I dK ) × G(0, 1) K , respectively. Given that μ 0 are learnt on 5% of S tr n , we train our policies on the remaining 95% portion of S tr n to match our theory that requires the prior to not depend on training data. The policies are trained using Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 20 epochs. Main results. We compare our bound to those in London and Sandler (2019); Sakhi et al. (2022); discarding the intractable bound in Swaminathan and Joachims (2015a) as it requires computing a covering number. Here we do not include the pessimistic objectives in Swaminathan and Joachims (2015a); London and Sandler (2019) since we directly optimize our bounds. But we make such a comparison in Section E.3.5 for completeness, showing the favorable performance of our bound and the newly proposed pessimistic objective in Equation (8.17). Also, we do not compare to Su et al. (2020); Metelli et al. (2021) since they do not provide generalization guarantees; they focus on estimation accuracy and only propose a heuristic for off-policy learning. However, we still show the favorable performance of our approach in off-policy learning compared to Su et al. (2020); Metelli et al. (2021) in Section E.3.6 for completeness. Prior methods are not named. Thus we refer to them as (Author, Policy) where Author ∈ Ours, London et al., Sakhi et al. 1, Sakhi et al. 2 and Policy ∈ Gaussian, Mixed-Logit. Here Ours, London et al., Sakhi et al. 1 and Sakhi et al. 2 correspond to Theorem 4, London and Sandler (2019, Theorem 1), Sakhi et al. (2022, Proposition 1), and Sakhi et al. (2022, Proposition 3), respectively. Since we have two classes of policies, each bound leads to two baselines. For example, London and Sandler (2019, Theorem 1) leads to (London et al., Gaussian) and (London et al., Mixed- Logit). More details are provided in Section E.3.3. In Figure 8.2, we report the value of the learned policies. Here we fix τ = 1/ 4 √ n ≈ 0.06 and α = 1 − 1/ 4 √ n ≈ 0.94 so that when n is large enough, both ˆ V τ (π) and ˆ V α (π) approach ˆ V ips (π) (Ionides, 2008). This is because standard IPS should be preferred when n→∞. To have a fair comparison, we fixed α instead of tuning it in an adaptive fashion as described in Section 8.3.4. However, we also provide the results with an adaptive α 113 in Figure 8.3. Let us start with interpreting Figure 8.2 (with fixed α and τ). Overall, our method outperforms all the baselines. We also observe that Gaussian policies behave better than mixed-logit policies. However, this is less significant for our method where the performances of both Gaussian and mixed-logit policies are comparable. Moreover, our method reaches the maximum value even when the logging policy has an average performance. In contrast, the baselines only reach their best value when the logging policy is well-performing (η 0 ≈ 1), in which case minor to no improvements are made. Finally, the baselines induce a better value when the logging policy is uniform (η 0 = 0). But our method has a better value when η 0 > 0, which is more common in practice. Larger action spaces. The experiments above did not consider very large values of K. However, Chapter 7 evaluated IPS-based methods, including exponential smoothing and clippped IPS, on datasets with up to one million actions. In those experiments, exponential smoothing outperformed clipped IPS, though the improvements were modest compared to the gains observed here. Choice of hyperparameters. Our choice of τ and α does not affect the above con- clusions. In Figure 8.3 (left-hand side), we compare our method with the best baseline, (Sakhi et al. 2) with Gaussian policies, for 20 evenly spaced values of τ ∈ (0, 1) and α∈ (0, 1). We also include the results using the adaptive tuning procedure of α described in Section 8.3.4 (green curve). This procedure is reliable since the performance with an adaptive α (green curve) is comparable with the best possible choice of α. Also, our method consistently outperforms the best baseline (Sakhi et al. 2) with the best value of τ when the logging policy is not uniform (η 0 > 0). Also, there is no very bad choice of α, in contrast with τ = 10 −5 (dark blue plot) which led to minimal improvement upon all logging policies. This might be due to the 1/τ dependency in existing bounds. 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 Logging Ours Sakhi et al. 2 Ours, Adaptive ® 0.00.20.40.60.81.0 smoothing parameter ® 0.60 0.65 0.70 0.75 0.80 0.85 0.90 reward of the learned policy MNIST, K=10, d=784 Modest Logging Good Logging Figure 8.3: On the left-hand side is the reward of the learned policy with varying τ ∈ (0, 1), α∈ (0, 1) and η 0 ∈ [0, 1], and for an adaptive α using the procedure in Section 8.3.4 (green curve). The blue-to-cyan and red-to-yellow colors correspond to varying values of τ and α, respectively. The lighter the color, the higher the value of τ or α. The green curve corresponds to the reward of the learned policy with an adaptive and data-dependent α (Section 8.3.4). On the right-hand side is the average reward of the learned policies using our method across the modest and good logging groups, η 0 ∈ [0, 0.5] (red) and η 0 ∈ [0.5, 1] (green), respectively. 114 To see the effect of α, we consider the following experiment. We split the logging policies into two groups. The first is called modest logging which corresponds to logging policies π 0 whose η 0 is between 0 and 0.5. This group includes the uniform policy and other average-performing policies. The second is called good logging and it includes the logging policies whose η 0 is between 0.5 and 1. Then, for each α, we compute the average value of the learned policy, with that value of α, across these two groups. This leads to the two red and green curves in Figure 8.3 (right-hand side). Overall, we observe that α ≈ 0.7 leads to the best performance across the modest logging group. Thus when the performance of the logging policy is bad or average, which is common in practice, importance-weight regularization can be critical. In contrast, when the performance of the logging policy is already good and n is large enough, importance-weight regularization might not be needed and α≈ 1 would also lead to good performance. This is one of the main strengths of our bound; it holds for the standard IPS recovered with α = 1. This result goes against the belief that clipped IPS should always be preferred to standard IPS. Here, our bound applied to standard IPS outperformed clipping by a large margin when the logging policy is relatively well-performing. Similar results for the other datasets are deferred to Section E.3.4. 8.6 Extension to Other Regularizations The experiments above demonstrated that exponential smoothing substantially outper- forms clipping. However, we compared exponential smoothing with our pessimistic objec- tive against clipping with pessimistic objectives specifically designed for it. This makes it difficult to isolate whether the gains stem from exponential smoothing as a regularization technique or from our pessimistic objective. Moreover, exponential smoothing and clip- ping are only two instances within a broader class of importance-weight regularizations. While numerous methods have been proposed to stabilize IPS through importance-weight transformations (Bottou et al., 2013; Swaminathan and Joachims, 2015a; Su et al., 2020; Metelli et al., 2021), most focus on estimation accuracy rather than learning performance. As highlighted in Chapter 7, improved estimation does not necessarily yield improved poli- cies, motivating a reassessment of importance-weight regularization specifically within the learning paradigm. Moreover, existing approaches followed a case-by-case basis: each regularization technique comes with its own theoretical analysis and corresponding pes- simistic objective. This inconsistency makes it impossible to determine whether empirical improvements arise from the regularizer itself or from its specific objective formulation. This reveals a critical gap: the absence of a unified framework providing principled pes- simistic objectives across diverse importance-weight regularizations. We address this by developing a generic PAC-Bayesian generalization bound that applies uniformly to a broad family of regularizations, enabling fair comparison within a single theoretical framework. Recall that the regularized IPS estimator has the form: ˆ V (π) = 1 n n ∑︂ i=1 R i ˆw(X i ,A i ),(8.19) 115 where ˆw(X,A) are the regularized importance weights. We further assume that ˆw(X,A) = g(π(A|X),π 0 (A|X)) for some function g : [0, 1]× [0, 1] → R + . Examples of ˆw include clipping (Clip) (London and Sandler, 2019), exponential smoothing (ES) (Aouali et al., 2023a), implicit exploration (IX) (Gabbianelli et al., 2024), and harmonic (Har) (Metelli et al., 2021), defined as Clip :ˆw(x,a) = π(a| x) max(π 0 (a| x),τ ) , τ ∈ [0, 1],(8.20) ES :ˆw(x,a) = π(a| x) π 0 (a| x) α , α∈ [0, 1], IX :ˆw(x,a) = π(a| x) π 0 (a| x) + γ , γ ∈ [0, 1], Har :ˆw(x,a) = w(x,a) (1− λ)w(x,a) + λ , λ∈ [0, 1]. 8.6.1 Generalization Bounds PAC-Bayes theory (Section 8.3) allows bounding ⃓ ⃓ ⃓ E θ∼Q [V (π θ )− ˆ V (π θ )] ⃓ ⃓ ⃓ , with V (π θ ) = E X∼ν,A∼π θ (·|X) [r(X,A)], ˆ V (π θ ) = 1 n n ∑︂ i=1 ˆw θ (X i ,A i )R i , where we make the dependence of ˆw θ on θ explicit to avoid confusion when taking the expectation E θ∼Q . Below is our first general result that extends Theorem 4 to any regu- larization function g, instead of just exponential smoothing. Its proof follows exactly the same techniques we employed for Theorem 4. Theorem 5. Let λ > 0, n≥ 1, δ ∈ (0, 1), and let P be a fixed prior on Θ. The following inequality holds with probability at least 1− δ for any distribution Q on Θ: ⃓ ⃓ ⃓ E θ∼Q [V (π θ )− ˆ V (π θ )] ⃓ ⃓ ⃓ ≤ √︃ kl 1 (π Q ) 2n + kl 2 (π Q ) nλ + B n (Q) + λ 2 Var n (Q),(8.21) where kl 1 (π Q ) = D KL (Q∥P) + log 4 √ n δ , kl 2 (π Q ) = D KL (Q∥P) + log 4 δ , and Var n (Q) = 1 n n ∑︂ i=1 E θ∼Q [︁ E A∼π 0 (·|X i ) [ ˆw θ (X i ,A) 2 ] + ˆw θ (X i ,A i ) 2 R 2 i ]︁ , B n (Q) = 1 n n ∑︂ i=1 ∑︂ A∈A E θ∼Q [︁ |π θ (A|X i )− π 0 (A|X i ) ˆw θ (X i ,A)| ]︁ . The terms in the above bound have similar interpretations to those in Theorem 4. Linear vs. non-linear regularization. If ˆw(X,A) is linear in π θ (X,A) (i.e., g linear in its first variable), then ˆ V is also linear in π θ , yielding ⃓ ⃓ ⃓ E θ∼Q [V (π θ )− ˆ V (π θ )] ⃓ ⃓ ⃓ = ⃓ ⃓ ⃓ V (π Q )− ˆ V (π Q ) ⃓ ⃓ ⃓ , 116 where we define (similar to Section 8.3.2) π Q = E θ∼Q [π θ ].(8.22) As seen in Section 8.3.2, this technique allows translating the bound in Theorem 5, which controls ⃓ ⃓ ⃓ E θ∼Q [V (π θ )− ˆ V (π θ )] ⃓ ⃓ ⃓ , into a bound that controls|V (π Q )− ˆ V (π Q )|, the quantity of interest in off-policy learning. The main requirement is to find linear importance-weight regularizations and policies that satisfy Equation (8.22). Fortunately, many importance- weight regularizations, such as Clip, IX, and ES in Equation (8.20), are linear in π, and several practical policies adhere to the formulation in Equation (8.22); see Section 8.3.2 for an in-depth explanation of such policies, including softmax, and Gaussian policies. In Corollary 1, we specialize Theorem 5 to linear importance-weight regularizations of the form ˆw θ (x,a) = π θ (a|x) h(π 0 (a|x)) , where h(π 0 (a|x)) ≥ π 0 (a|x) for all (x,a) ∈ X ×A. We additionally assume that the base policies π θ are deterministic, i.e., π θ (a | x) ∈ 0, 1, which implies π θ (a|x) 2 = π θ (a|x). This assumption is only needed here and it is mild: the PAC-Bayes policies π Q defined in Equation (8.22) are mixtures of deterministic policies under Q, and common policy classes such as softmax, mixed-logit, and Gaussian policies admit such representations (Section 8.3.2). Under these assumptions, Theorem 5 yields the following result. Corollary 1. Assume the regularized importance weights can be written as ˆw θ (x,a) = π θ (a|x) h(π 0 (a|x)) with h : [0, 1] → R + verifies h(p) ≥ p for any p ∈ [0, 1]. Moreover, for any distribution Q in the parameter space Θ, we define π Q = E θ∼Q [π θ ] where π θ is binary. Then, let λ > 0, n ≥ 1, δ ∈ (0, 1), and let P be a fixed prior on Θ, The following inequality holds with probability at least 1− δ for any distribution Q on Θ ⃓ ⃓ ⃓ V (π Q )− ˆ V (π Q ) ⃓ ⃓ ⃓ ≤ √︃ kl 1 (Q) 2n + B n (π Q ) + kl 2 (Q) nλ + λ 2 Var n (π Q ),(8.23) where kl 1 (Q) and kl 2 (Q) are defined in Theorem 5, and Var n (π Q ) = 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) h(π 0 (A|X i )) 2 ]︃ + π Q (A i |X i ) h(π 0 (A i |X i )) 2 R 2 i , B n (π Q ) = 1− 1 n n ∑︂ i=1 ∑︂ A∈A π 0 (A|X i ) π Q (A|X i ) h(π 0 (A|X i )) . The main benefit of Corollary 1 compared to Theorem 5 is that it eliminates the need for the expectation E θ∼Q [·], which is now embedded in the definition of policies in Equa- tion (8.22). For example, Corollary 1 allows us to recover the main result of ES above Aouali et al. (2023a) when h(p) = p α , α ∈ [0, 1]. Similarly, we can apply it to IX (Gab- bianelli et al., 2024) by setting h(p) = p + γ, γ ≥ 0, and to Clip (London and Sandler, 2019) by setting h(p) = max(p,τ ), τ ∈ [0, 1]. However, if ˆw θ (x,a) is not linear in π θ (a|x), then this technique cannot be used, and the original expectation E θ∼Q [·] in Theorem 5 must be retained. 117 8.6.2 Pessimistic Objectives Theorem 5 yields two pessimistic objectives. Bound optimization. The first approach directly maximizes the lower bound from Theorem 5: argmax Q E θ∼Q [︂ ˆ V (π θ ) ]︂ − √︃ kl 1 (Q) 2n − B n (Q)− kl 2 (Q) nλ − λ 2 Var n (Q),(8.24) The main challenge is that the objective involves expectations under Q. We address this using the local reparameterization trick (Kingma et al., 2015), which expresses gra- dients of expectations as expectations of gradients, estimated via Monte Carlo sam- pling. Specifically, we consider softmax policies π sof θ (a|x) from Equation (8.11) and set Q = N (μ,σ 2 I dK ) with learnable parameters μ ∈ R dK and σ > 0. All terms in Equa- tion (8.24) take the form E θ∼N (μ,σ 2 I dK ) [f (π sof θ (a|x))], which can be rewritten as: E θ∼N (μ,σ 2 I dK ) [f (π sof θ (a|x))] = E ε∼N (0,∥φ(x)∥ 2 2 I K ) [︃ f (︃ exp(φ(x) ⊤ μ a + σε a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ μ a ′ + σε a ′ ) )︃]︃ . This expectation is approximated by sampling ε i ∼ N (0,∥φ(x)∥ 2 2 I K ) and computing the empirical mean; gradients are estimated similarly. However, this approach can exhibit high variance when K is large. For linear importance-weight regularizations, this can be mitigated by optimizing the bound in Corollary 1. For the general case, we propose a practical alternative. Heuristic optimization. The second approach avoids the challenges of direct bound optimization at the cost of additional hyperparameters. Inspired by Theorem 5, we maxi- mize the estimated value penalized by bias, variance, and proximity to the logging policy: ˆ V (π θ )− λ 1 ∥θ− θ 0 ∥ 2 − λ 2 ̃ Var n (π θ )− λ 3 ̃ B n (π θ ),(8.25) where ̃ Var n (π θ ) and ̃ B n (π θ ) are the terms inside the expectations in Var n (Q) and B n (Q), respectively, θ 0 parameterizes the logging policy π 0 , and λ 1 ,λ 2 ,λ 3 are hyperparameters. Both objectives in Equations (8.24) and (8.25) are amenable to stochastic gradient op- timization and are generic across importance-weight regularizations, enabling fair com- parison. We empirically compare these objectives and evaluate different regularization techniques in Section 8.7. 8.7 Experiments for Other Regularizations We adopt the experimental setting of Section 8.5 and conduct two main experiments. In Section 8.7.1, we fix the importance-weight regularization to Clip (Equation (8.20)) and compare our pessimistic objective against PAC-Bayesian objectives from the literature specifically designed for clipping. The goal is to demonstrate that our objective not 118 only applies more broadly but also outperforms existing alternatives. In Section 8.7.2, having validated our pessimistic objective, we fix it and compare across importance-weight regularizations. The goal is to determine whether any particular regularization technique yields superior off-policy learning performance. 8.7.1 Varying Pessimistic Objectives, Fixed Regularization We examine the impact of different pessimistic objectives on learned policy performance, fixing the importance-weight regularization to Clip: ˆw(x,a) = π(a|x) max(π 0 (a|x),τ ) in Equa- tion (8.20), with τ = 1/ 4 √ n following Ionides (2008). For fair comparison, we consider PAC-Bayesian objectives from prior work where the theoretical bound is optimized di- rectly. Specifically, we include London et al. (London and Sandler, 2019, Theorem 1), and two bounds from Sakhi et al. (2022): Sakhi et al. 1 (Sakhi et al., 2022, Proposition 1), based on Catoni (2007), and Sakhi et al. 2 (Sakhi et al., 2022, Proposition 3), a Bernstein-type bound. Since these baselines use linear importance-weight regularization (Section 8.6.1), we compare against our bound in Corollary 1. Following Sakhi et al. (2022); Aouali et al. (2023a), we optimize over Gaussian policies (Equation (8.13)), which perform better in this setting. We also include the logging policy as a baseline. Figure 8.4 plots the reward of learned policies as a function of logging policy quality η 0 ∈ [0, 1]. Our objective outperforms all baselines across a wide range of logging policies. Thus, in addition to being generic across importance-weight regularizations, our approach proves more effective than objectives tailored specifically for Clip. This advantage holds when η 0 is not too close to zero: a realistic scenario where logging policies typically outperform uniform random selection. Note that all methods (including ours) improve upon the logging policy (dashed black lines). 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 Ours London et al. Sakhi et al. 1 Sakhi et al. 2 Logging Figure 8.4: Performance of the learned policy with different PAC-Bayes pessimistic ob- jectives (our Corollary 1 and those in London and Sandler (2019); Sakhi et al. (2022)) using the Clip IPS estimator in Equation (8.20) . 8.7.2 Varying Regularization, Fixed Pessimistic Objective Having demonstrated the favorable performance of our pessimistic objective, we now compare different importance-weight regularization techniques: Clip, Har, IX, and ES 119 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward MNIST, K=10, d=784 Logging Clip inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 MNIST, K=10, d=784 Logging Har inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 MNIST, K=10, d=784 Logging IX inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 MNIST, K=10, d=784 Logging ES inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Av. reward w.r.t. hyperparameters MNIST, K=10, d=784 −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 reward FashionMNIST, K=10, d=784 Logging Clip −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 FashionMNIST, K=10, d=784 Logging Har −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 FashionMNIST, K=10, d=784 Logging IX −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 FashionMNIST, K=10, d=784 Logging ES −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Av. reward w.r.t. hyperparameters FashionMNIST, K=10, d=784 Figure 8.5: Performance of the policy learned by Bound optimization (i.e, Equa- tion (8.24)) for different importance-weight regularizations. The x-axis reflects the quality of the logging policy η 0 ∈ [−0.5, 0.5]. In the first four columns, we plot the reward of the learned policy using a fixed importance-weight regularization technique (Clip, Har, IX, or ES as defined in Equation (8.20)) for various values of its hyperparameter within [0, 1]. In the last column, we report the mean reward across these hyperparameter values. (Equation (8.20)). We evaluate both pessimistic objectives from Section 8.6.2, optimizing over softmax policies. For bound optimization, we use Theorem 5 rather than Corollary 1, since Har is non-linear in π. We set λ to its optimal value λ ∗ minimizing the bound. While our theory requires λ to be fixed a priori (since λ ∗ is data-dependent), we found this yields good empirical performance. For heuristic optimization (Equation (8.25)), we set λ 1 = λ 2 = λ 3 = 10 −5 . Figures 8.5 and 8.6 present learned policy rewards as a function of logging policy quality η 0 , for bound optimization and heuristic optimization respectively. We vary η 0 ∈ [−0.5, 0.5], including logging policies worse than uniform (η 0 < 0) to highlight settings requiring stronger regularization, though such scenarios are rarely encountered in practice. Rows correspond to MNIST and FashionMNIST. The first four columns show results for each regularization technique across hyperparameter values in [0, 1]; the last column reports mean reward across hyperparameters for each regularization technique to assess sensitivity to hyperparameters. Bound optimization (Figure 8.5). All regularizations improve over the logging policy (all curves above the dashed baseline), with Har showing less improvement. Clip, IX, and ES achieve comparable performance despite regularizing importance weights differ- ently. These results align with the generality of our bound and suggest that the choice of regularization has limited impact when optimizing the theoretical bound directly. Heuristic optimization (Figure 8.6). Heuristic optimization achieves better perfor- mance than bound optimization, likely due to practical limitations of Monte Carlo estima- tion in high dimensions (Section 8.6.2). The far-right column reveals comparable average performance across regularizations, with two exceptions: ES outperforms the others while Har underperforms. This clarifies our results from Section 8.5: the superior performance of exponential smoothing is more related to our pessimistic objective than the smooth regularization itself. Here, the smooth regularization adds some improvements compared 120 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward MNIST, K=10, d=784 Logging Clip inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 MNIST, K=10, d=784 Logging Har inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 MNIST, K=10, d=784 Logging IX inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 MNIST, K=10, d=784 Logging ES inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Av. reward w.r.t. hyperparameters MNIST, K=10, d=784 −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 reward FashionMNIST, K=10, d=784 Logging Clip −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 Logging Har −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 Logging IX −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 Logging ES −0.6−0.4−0.20.00.20.40.6 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Av. reward w.r.t. hyperparameters FashionMNIST, K=10, d=784 Figure 8.6: Performance of the policy learned by Heuristic optimization in Equa- tion (8.25) for different importance-weight regularizations. The x-axis reflects the quality of the logging policy η 0 ∈ [−0.5, 0.5]. In the first four columns, we plot the reward of the learned policy using a fixed importance-weight regularization technique (Clip, Har, IX, or ES as defined in Equation (8.20)) for various values of its hyperparameter within [0, 1]. In the last column, we report the mean reward across these hyperparameter values. to others, but the improvments are not significant compared to the improvments we get by simply changing the pessimistic learnin principle even with the standard clipping reg- ularization(Figure 8.4) Larger action spaces. The experiments in this section did not consider very large val- ues of K. However, Chapter 7 evaluated numerous IPS-based methods on datasets with up to one million actions. In those experiments, exponential smoothing outperformed other IPS-based methods, though the improvements were modest. Combined with the results above, this reinforces our conclusion: the choice of the objective has a larger im- pact on learning performance than the choice of importance-weight regularization, which primarily affects estimation accuracy. 8.8 Conclusion In this chapter, we investigated importance-weight regularization techniques within the pessimistic paradigm, with particular focus on exponential smoothing as a principled alter- native to hard clipping. Our key contributions include: (i) tractable two-sided PAC-Bayes generalization bounds that, unlike prior work, apply to both regularized and standard IPS estimators and are amenable to stochastic gradient optimization; and (i) the first uni- fied framework for comparing diverse importance-weight regularizations under a common pessimistic objective. This work addresses fundamental theoretical limitations in exist- ing approaches, including the reliance on one-sided inequalities and the misapplication of evaluation bounds in off-policy learning. Rather than using theoretical bounds merely as inspiration for heuristics, we directly optimize them, representing a first step toward making IPS-based pessimism more practical. Our work has two primary limitations. First, the inclusion of empirical bias and variance terms in our bounds makes deriving data-independent suboptimality gaps challenging. 121 Second, two-sided bounds for regularized IPS can be loose as they treat both tails sym- metrically, whereas recent work indicates significant asymmetry between lower and upper tails. We address both limitations in Sakhi et al. (2024), which investigates tail-specific bounds to achieve significantly tighter guarantees and sharp suboptimality results. This chapter serves practitioners committed to IPS-based methods, whose appeal is well- founded: unbiasedness, theoretical guarantees, and the ability to derive principled pes- simistic objectives for safe policy learning in high-stakes scenarios. However, from a purely empirical perspective, particularly in very large action spaces, we would favor the PWLL objectives introduced in Chapter 7, which consistently outperform IPS-based methods. The reader can find a direct comparison of these approaches on datasets with up to one million actions in that chapter. 122 Chapter 9 Conclusions and Future Work This thesis addressed, from both a practical and theoretical perspective, the core obstacle to deploying contextual bandits in modern applications: scalability to large action spaces while maintaining computational tractability. On-Policy Learning (Part I). We introduced structured Bayesian models that enable principled information sharing across actions and derived scalable exploration algorithms. meTS in Chapter 3 couples action parameters through shared latent effects, yielding regret and complexity that scale with an effective number of actions rather than K. dTS in Chapter 4 further develops this idea by using a pre-trained diffusion model to encode richer structure. These algorithms perform well in practice in their theoretical form, without additional tweaks or hyperparameter tuning. Off-Policy Learning (Part I). We tackled both pillars: DM and IPS approaches. sDM in Chapter 6 extends the structured modeling of Part I to the offline regime. We then showed in Chapter 7 that, in large action spaces, optimization matters more than estimation: estimator-based objectives induce highly non-concave landscapes, whereas policy-weighted log-likelihoods produce concave objectives for common policy classes and win decisively at scale. Finally, we developed a principled pessimistic framework for regu- larized IPS in Chapter 8: smooth importance-weight regularization (exponential smooth- ing) paired with two-sided PAC-Bayes bounds, and a unified analysis that clarifies when regularization matters and how to compare it across methods. Our additional work on logarithmic smoothing (Sakhi et al., 2024) sharpens the concentration analysis further and yields tighter learning guarantees. Thesis message and practical implications. Scaling to large action spaces causes methods that perform well in small settings (e.g., standard IPS) to fail at scale. This thesis advances three design principles to address this challenge: (i) encode structure to shrink the effective action space; (i) prioritize objectives with favorable optimization properties over faithful but intractable estimators; and (i) when relying on IPS with pes- simism, couple differentiable importance-weight corrections with theoretically grounded, data-driven bounds amenable to stochastic gradient descent. Together, these principles yield algorithms that are statistically efficient, computationally tractable, and numerically stable. 123 Future work. This thesis opens several promising directions for future research. A key theoretical challenge is to establish robust guarantees under model misspecification, extending the Bayesian analysis of sDM, meTS, and dTS beyond the well-specified set- ting. For on-policy learning, developing a comprehensive nonlinear diffusion theory for Thompson sampling remains an open problem. In the off-policy setting, future work could investigate the extensions and applications of our methods to LLM and diffusion model fine-tuning, where the objective closely mirrors offline contextual bandit objectives. More- over, integrating these approaches into large-scale recommender pipelines requires efficient action retrieval, slate constraints, and systems-level optimization. Some of these aspects, such as coupling decision-making with approximate maximum inner product search, were explored in our applied studies (see Additional Contributions in Section 1.3) but omitted from this manuscript. These practical directions have a tangible impact on the online advertising industry and beyond, and are worth pursuing. 124 Chapter A Supplementary Materials for Chapter 3 A.1 Preliminaries In this section, we recall some basic properties of matrix operations. (a) The mixed-product property. We have that (A⊗ B)(C⊗ D) = AC⊗ BD for any matrices A, B, C, D such that the products AC and BD exist. (b) Transpose. We have that (A⊗ B) ⊤ = A ⊤ ⊗ B ⊤ for any matrices A, B. (c) Vectorization. Let A ∈ R n×m , B ∈ R m×p , then Vec(AB) = (I p ⊗ A) Vec(B) = (B ⊤ ⊗ I n ) Vec(A). (d) For any matrix A, we have that I 1 ⊗ A = A. (e) For any positive semi-definite matrices A and B, we have that λ 1 (A⊗B) = λ 1 (A)λ 1 (B). (f) For any matrix A and any positive semi-definite matrix B such that the product A ⊤ BA exists, the following inequality holds λ 1 (A ⊤ BA)≤ λ 1 (B)λ 1 (A ⊤ A). A.2 Posterior Derivations Here we provide the derivations of the effect posterior and action posteriors for the set- ting presented in Section 3.1.1. Precisely, we present the proof for Proposition 1 in Section A.2.1 and the proof of Proposition 2 in Section A.2.2. A.2.1 Effect Posterior Derivation Proof of Proposition 1 (derivation of q t ). First, from basic properties of matrix opera- tions, we observe that the mean of the action parameter can be rewritten using Kronecker products. Specifically, ∑︁ ℓ∈[L] b a,ℓ ψ ℓ = Γ a Ψ, where Ψ = (ψ ℓ ) ℓ∈[L] ∈ R Ld is the concatenated effect vector and Γ a = b ⊤ a ⊗ I d ∈ R d×Ld . Thus, the model in Equation (3.2) (up to round 125 t∈ [T ]) can be written as Ψ∼N (μ Ψ , Σ Ψ ), θ a | Ψ∼N (Γ a Ψ, Σ 0,a ),∀a∈A, R i | X i ,A i ,θ, Ψ∼N (X ⊤ i θ A i ,σ 2 ),∀i∈ [t− 1].(A.1) Under this model, conditional on (θ a ) a∈A and (X i ,A i ) i<t , the rewards (R i ) i<t are inde- pendent and each R i depends on Ψ only through θ A i . Hence p((R i ) i<t | (X i ,A i ) i<t , Ψ) = ∫︂ θ∈R dK p((R i ) i<t | (X i ,A i ) i<t ,θ)p(θ | Ψ)dθ. Moreover, since p(θ | Ψ) = ∏︁ a∈A p 0,a (θ a | Ψ) and p((R i ) i<t | (X i ,A i ) i<t ,θ) = ∏︂ a∈A ∏︂ i∈S t,a N (R i ;X ⊤ i θ a ,σ 2 ) = ∏︂ a∈A L t,a (θ a ), the integral factorizes across arms: p((R i ) i<t | (X i ,A i ) i<t , Ψ) = ∏︂ a∈A ∫︂ L t,a (θ a )p 0,a (θ a | Ψ)dθ a . It follows that the joint effect posterior in round t reads q t (Ψ) ∝ p((R i ) i<t | (X i ,A i ) i<t , Ψ)q 0 (Ψ),(A.2) = ∏︂ a∈A ∫︂ θ a L t,a (θ a )p 0,a (θ a | Ψ) dθ a q 0 (Ψ) = ∏︂ a∈A ∫︂ θ a L t,a (θ a )N (θ a ; Γ a Ψ, Σ 0,a ) dθ a ⏞ ⏟⏞ I a (Ψ) N (Ψ;μ Ψ , Σ Ψ ),(A.3) where L t,a (θ a ) = ∏︁ i∈S t,a N (R i ;X ⊤ i θ a ,σ 2 ). We compute the integral term I a (Ψ) using Lemma 1. Specifically, we obtain that I a (Ψ) is proportional to a Gaussian density on Ψ, denoted N (Ψ; ̄μ t,a , ̄ Σ t,a ), where ̄ Σ −1 t,a = Γ ⊤ a (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ Γ a , ̄μ t,a = ̄ Σ t,a (︁ Γ ⊤ a Σ −1 0,a (G t,a + Σ −1 0,a ) −1 B t,a )︁ , and G t,a and B t,a are defined in Section 3.2.2. Consequently, the effect posterior q t (Ψ) is proportional to the product of K + 1 multivariate Gaussian distributions: the prior N (μ Ψ , Σ Ψ ) and the likelihood contributionsN ( ̄μ t,a , ̄ Σ t,a ) for each a∈A. Since the product of Gaussians is Gaussian, q t = N ( ̄μ t , ̄ Σ t ), where the precision matrix is the sum of the individual precisions: ̄ Σ −1 t = Σ −1 Ψ + ∑︂ a∈A ̄ Σ −1 t,a = Σ −1 Ψ + ∑︂ a∈A Γ ⊤ a (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ Γ a . 126 Using that Γ a = b ⊤ a ⊗ I d , we rewrite the term inside the sum as: Γ ⊤ a (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ Γ a = (b a ⊗ I d ) (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ (b ⊤ a ⊗ I d ) = (b a b ⊤ a )⊗ (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ . Similarly, for the mean ̄μ t , we have: ̄μ t = ̄ Σ t (︄ Σ −1 Ψ μ Ψ + ∑︂ a∈A ̄ Σ −1 t,a ̄μ t,a )︄ = ̄ Σ t (︄ Σ −1 Ψ μ Ψ + ∑︂ a∈A Γ ⊤ a Σ −1 0,a (G t,a + Σ −1 0,a ) −1 B t,a )︄ . Using the mixed (Kronecker) product property, Γ ⊤ a Σ −1 0,a (G t,a + Σ −1 0,a ) −1 B t,a = b a ⊗ (︁ Σ −1 0,a (G t,a + Σ −1 0,a ) −1 B t,a )︁ . This recovers the expressions in Proposition 1. To reduce clutter in the following lemma, we fix an action a∈A and a round t. We drop the sub-indices a and t, so that we have the following correspondences: Γ← Γ a ,Σ 0 ← Σ 0,a , N ← N t,a , θ ← θ a ,(X i ,R i ) i∈[N ] ← (X i ,R i ) i∈S t,a , Lemma 1 (Gaussian posterior update). Let Γ∈ R d×Ld , Σ 0 ∈ R d×d , and σ > 0. Consider a dataset of N observations (X i ,R i ) N i=1 . Then, ∫︂ θ (︄ N ∏︂ i=1 N (R i ;X ⊤ i θ,σ 2 ) )︄ N (θ; ΓΨ, Σ 0 ) dθ ∝N (Ψ;μ N , Σ N ) , where Σ −1 N = Γ ⊤ (︂ Σ −1 0 − Σ −1 0 (︁ G N + Σ −1 0 )︁ −1 Σ −1 0 )︂ Γ, Σ −1 N μ N = (︂ Γ ⊤ Σ −1 0 (︁ G N + Σ −1 0 )︁ −1 B N )︂ . with G N = σ −2 ∑︁ N i=1 X i X ⊤ i and B N = σ −2 ∑︁ N i=1 R i X i . Proof. Let v = σ −2 and Λ 0 = Σ −1 0 . We denote the integral in the lemma by f (Ψ). Completing the square for θ, we have: f (Ψ)∝ ∫︂ θ exp [︄ − v 2 N ∑︂ i=1 (R i − X ⊤ i θ) 2 − 1 2 (θ− ΓΨ) ⊤ Λ 0 (θ− ΓΨ) ]︄ dθ ∝ ∫︂ θ exp [︂ − 1 2 (︂ θ ⊤ (︄ v N ∑︂ i=1 X i X ⊤ i + Λ 0 )︄ ⏞⏟⏞ V −1 N θ− 2θ ⊤ (︄ v N ∑︂ i=1 R i X i + Λ 0 ΓΨ )︄ + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︂]︂ dθ . 127 To reduce clutter, let G N = v N ∑︂ i=1 X i X ⊤ i , V N = (G N + Λ 0 ) −1 , U N = V −1 N , B N = v N ∑︂ i=1 R i X i and β N = V N (B N + Λ 0 ΓΨ) . We have that U N V N = V N U N = I d , and thus f (Ψ)∝ ∫︂ θ exp [︃ − 1 2 (︁ θ ⊤ U N θ− 2θ ⊤ U N V N (B N + Λ 0 ΓΨ) + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︁ ]︃ dθ , = ∫︂ θ exp [︃ − 1 2 (︁ θ ⊤ U N θ− 2θ ⊤ U N β N + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︁ ]︃ dθ , = ∫︂ θ exp [︃ − 1 2 (︁ (θ− β N ) ⊤ U N (θ− β N )− β ⊤ N U N β N + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︁ ]︃ dθ , ∝ exp [︃ − 1 2 (︁ −β ⊤ N U N β N + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︁ ]︃ , = exp [︃ − 1 2 (︂ − (B N + Λ 0 ΓΨ) ⊤ V N (B N + Λ 0 ΓΨ) + (ΓΨ) ⊤ Λ 0 (ΓΨ) )︂ ]︃ , ∝ exp [︃ − 1 2 (︁ Ψ ⊤ Γ ⊤ (Λ 0 − Λ 0 V N Λ 0 ) ΓΨ− 2Ψ ⊤ (︁ Γ ⊤ Λ 0 V N B N )︁)︁ ]︃ , = exp [︃ − 1 2 Ψ ⊤ Σ −1 N Ψ + Ψ ⊤ Σ −1 N μ N ]︃ , where Σ −1 N = Γ ⊤ (Λ 0 − Λ 0 V N Λ 0 ) Γ, Σ −1 N μ N = (︁ Γ ⊤ Λ 0 V N B N )︁ .(A.4) Plugging the expression of V N concludes the proof. A.2.2 Action Posterior Derivation Proof of Proposition 2 (Derivation of p t,a ). This proposition is a direct application of Lemma 2; in which case we get that the posterior p t,a is a multivariate Gaussian distributionN ( ̃μ t,a , ̃ Σ t,a ), where ̃ Σ −1 t,a = G t,a + Σ −1 0,a , ̃μ t,a = ̃ Σ t,a (︄ B t,a + Σ −1 0,a L ∑︂ ℓ=1 b a,ℓ ψ t,ℓ )︄ . To reduce clutter, we consider a fixed action a ∈ [K] and round t ∈ [T ], and drop subindexing by t and a in Lemma 2. In summary, fix a ∈ [K] and t ∈ [T ] such that we have the following correspondences: b ℓ ← b a,ℓ ,Σ 0 ← Σ 0,a , N ← N t,a , θ ← θ a ,(X i ,R i ) i∈[N ] ← (X i ,R i ) i∈S t,a . 128 Lemma 2. Consider the following model θ | Ψ∼N (︄ L ∑︂ ℓ=1 b ℓ ψ ℓ , Σ 0 )︄ , R i | X i ,θ ∼N (︁ X ⊤ i θ,σ 2 )︁ ,∀i∈ [N ]. Let H =X 1 ,R 1 ,...,X N ,R N then we have that p(θ | Ψ,H) =N (︂ θ; ̃μ N , ̃ Σ N )︂ , where ̃ Σ −1 N = σ −2 N ∑︂ i=1 X i X ⊤ i + Σ −1 0 , ̃μ N = ̃ Σ N (︄ σ −2 N ∑︂ i=1 X i R i + Σ −1 0 L ∑︂ ℓ=1 b ℓ ψ ℓ )︄ . Proof. Let v = σ −2 ,Λ 0 = Σ −1 0 . Then the action posterior decomposes as p(θ | Ψ,H)∝ p((R i ) i∈[N ] | Ψ,θ, (X i ) i∈[N ] )p(θ | Ψ), = p((R i ) i∈[N ] | θ, (X i ) i∈[N ] )p(θ | Ψ), = N ∏︂ i=1 N (R i ;X ⊤ i θ,σ 2 )N (θ; L ∑︂ ℓ=1 b ℓ ψ ℓ , Σ 0 ), = exp [︂ − 1 2 (︂ v N ∑︂ i=1 (R 2 i − 2R i X ⊤ i θ + (X ⊤ i θ) 2 ) + θ ⊤ Λ 0 θ− 2θ ⊤ Λ 0 L ∑︂ ℓ=1 b ℓ ψ ℓ + (︄ L ∑︂ ℓ=1 b ℓ ψ ℓ )︄ ⊤ Λ 0 (︄ L ∑︂ ℓ=1 b ℓ ψ ℓ )︄ )︂]︂ , ∝ exp [︄ − 1 2 (︄ θ ⊤ (v N ∑︂ i=1 X i X ⊤ i + Λ 0 )θ− 2θ ⊤ (︄ v N ∑︂ i=1 X i R i + Λ 0 L ∑︂ ℓ=1 b ℓ ψ ℓ )︄)︄]︄ , ∝N (︃ θ; ̃μ N , (︂ ̃ Λ N )︂ −1 )︃ , where ̃ Λ N = v ∑︁ N i=1 X i X ⊤ i + Λ 0 , and ̃ Λ N ̃μ N = v ∑︁ N i=1 X i R i + Λ 0 ∑︁ L ℓ=1 b ℓ ψ ℓ . A.3 Regret Proofs In this section, we establish a more general version of Theorem 1. As explained in Sec- tion 3.1.1, we analyze meTS in the linear setting under the assumption of a fully well- specified model. That is, the true action parameters and rewards are generated according to the same hierarchical structure assumed by meTS: Ψ ∗ ∼N (μ Ψ , Σ Ψ ),(A.5) θ ∗,a | Ψ ∗ ∼N (︂ L ∑︂ ℓ=1 b a,ℓ ψ ∗,ℓ , Σ 0,a )︂ ,∀a∈A, R t | X t ,A t ,θ ∗ , Ψ ∗ ∼N (X ⊤ t θ ∗,A t ,σ 2 ),∀t∈ [T ], 129 where the subscript ∗ denotes the true action and latent parameters. To derive the regret bound, we proceed as follows: First, we provide a compact problem formulation in Section A.3.1. Next, we employ total covariance decomposition to derive the posterior covariance of θ ∗,a | H t in Section A.3.2. Finally, we present preliminary eigenvalue results in Section A.3.3 before completing the proof in Section A.3.4. A.3.1 Problem Reformulation for Regret Analysis Here, we aim at rewriting Equation (A.5) in a compact form to simplify regret analysis. We first introduce K independent multivariate Gaussian variables Z a ∼ N (0, Σ 0,a ) for a∈ [K], and the following matrix Ψ ∗,mat = [ψ ∗,1 ,...,ψ ∗,L ]∈ R d×L . First, we have that Vec(Ψ ∗,mat ) = Ψ ∗ where Ψ ∗ is defined in Equation (A.5). Moreover notice that ∑︁ L ℓ=1 b a,ℓ ψ ∗,ℓ = Ψ ∗,mat b a , where b a = (b a,ℓ ) ℓ∈[L] and thus given matrix Ψ ∗,mat we have that θ ∗,a = Ψ ∗,mat b a + Z a ,∀a∈ [K].(A.6) We vectorize Equation (A.6) to obtain θ ∗,a = Vec(θ ∗,a ) = Vec(Ψ ∗,mat b a + Z a ) = Vec(Ψ ∗,mat b a ) + Z a ,(A.7) where we used that if X ∈ R d (a column vector), then X = Vec(X) and that Vec(·) is a linear transformation. Also, we know from (c) in Section A.1 that Vec(AB) = (B ⊤ ⊗ I n ) Vec(A) for any A∈ R n×m , B∈ R m×p . Therefore, θ ∗,a = Γ a Ψ ∗ + Z a ,(A.8) where Γ a = b ⊤ a ⊗ I d and we used that Vec(Ψ ∗,mat ) = Ψ ∗ . It follows that θ ∗,a | Ψ ∗ ∼N (Γ a Ψ ∗ , Σ 0,a ),(A.9) This allows us to rewrite our model as a single-parent hierarchical model Ψ ∗ ∼N (μ Ψ , Σ Ψ ),(A.10) θ ∗,a | Ψ ∗ ∼N (Γ a Ψ ∗ , Σ 0,a ),∀a∈ [K], R t | X t ,A t ,θ ∗ , Ψ ∗ ∼N (X ⊤ t θ ∗,A t ,σ 2 ),∀t∈ [T ]. A.3.2 Derivation of cov [θ ∗,a |H t ] Let G t,a = σ −2 ∑︂ i∈S t,a X i X ⊤ i , B t,a = σ −2 ∑︂ i∈S t,a R i X i . 130 Lemma 3 (Expression of cov [θ ∗,a |H t ]). Consider the model in Equation (A.10), then we have ˆ Σ t,a = cov [θ ∗,a |H t ] = ̃ Σ t,a + ̃ Σ t,a Σ −1 0,a Γ a ̄ Σ t Γ ⊤ a Σ −1 0,a ̃ Σ t,a , ∀a∈ [K]. where ̄ Σ t = (︂ Σ −1 Ψ + K ∑︂ a=1 b a b ⊤ a ⊗ (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ )︂ −1 ̃ Σ t,a = (︁ G t,a + Σ −1 0,a )︁ −1 . Proof. Before proceeding with the proof, we emphasize that cov [Ψ ∗ |H t ] = cov [Ψ|H t ] = ̄ Σ t ,E [Ψ ∗ |H t ] = E [Ψ|H t ] = ̄μ t , and cov [θ ∗,a | Ψ ∗ ,H t ] = cov [θ a | Ψ,H t ] = ̃ Σ t,a ,E [θ ∗,a | Ψ ∗ ,H t ] = E [θ a | Ψ,H t ] = ̃μ t,a , where the explicit expressions of these covariances and expectations are provided in Propo- sition 1 and Proposition 2, respectively. These equalities hold because the true action parameters and rewards are assumed to follow the exact generative process defined by the meTS model. Now let Λ 0,a = Σ −1 0,a . Proposition 2 and the fact that ∑︁ ℓ∈[L] b a,ℓ ψ ∗,ℓ = Γ a Ψ ∗ where Γ a = b ⊤ a ⊗ I d (Section A.3.1) yield cov [θ ∗,a | Ψ ∗ ,H t ] = (G t,a + Λ 0,a ) −1 E [θ ∗,a | Ψ ∗ ,H t ] = cov [θ ∗,a | Ψ ∗ ,H t ] (B t,a + Λ 0,a Γ a Ψ ∗ ) First, given H t , cov [θ ∗,a | Ψ ∗ ,H t ] = (G t,a + Λ 0,a ) −1 is constant (does not depend on Ψ ∗ ). Thus E [cov [θ ∗,a | Ψ ∗ ,H t ]|H t ] = cov [θ ∗,a | Ψ ∗ ,H t ] = (G t,a + Λ 0,a ) −1 . In addition, given H t , both (G t,a + Λ 0,a ) −1 and B t,a are constant. Thus cov [E [θ ∗,a | Ψ ∗ ,H t ]|H t ] = cov [cov [θ ∗,a | Ψ ∗ ,H t ] Λ 0,a Γ a Ψ ∗ |H t ] = (G t,a + Λ 0,a ) −1 Λ 0,a Γ a cov [Ψ ∗ |H t ] Γ ⊤ a Λ 0,a (G t,a + Λ 0,a ) −1 = (G t,a + Λ 0,a ) −1 Λ 0,a Γ a ̄ Σ t Γ ⊤ a Λ 0,a (G t,a + Λ 0,a ) −1 . Finally, total covariance decomposition (Weiss, 2005) concludes the proof. A.3.3 Preliminary Eigenvalues Results Next we present some preliminary upper bounds on the maximum eigenvalues of our covariance matrices. 131 • Definitions: Let λ 1,0 = max a∈[K] λ 1 (Σ 0,a ), λ d,0 = min a∈[K] λ d (Σ 0,a ), λ 1,Ψ = λ 1 (Σ Ψ ), and κ b = max a∈[K] ∥b a ∥ 2 2 . • upper bound of λ 1 (Γ a Γ ⊤ a ): λ 1 (Γ a Γ ⊤ a )≤ κ b ,∀a∈ [K].(A.11) Similarly, we have that λ 1 (Γ ⊤ a Γ a )≤ κ b ,∀a∈ [K].(A.12) • upper bound of λ 1 ( ˆ Σ t,a ): λ 1 ( ˆ Σ t,a )≤ λ 1,0 + λ 2 1,0 λ 1,Ψ κ b λ 2 d,0 ,∀a∈ [K].(A.13) • upper bound of λ 1 (Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ ): λ 1 (Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ )≤ 1 + Kλ 1,Ψ κ b (︂ 1 λ d,0 − 1 λ 2 1,0 (︁ κ x T σ 2 + 1 λ d,0 )︁ )︂ .(A.14) Proof. We start with Equation (A.11). First, recall that Γ a = b ⊤ a ⊗ I d for any a ∈ [K]. Thus Γ a Γ ⊤ a = (b ⊤ a ⊗I d )(b a ⊗I d ) =∥b a ∥ 2 2 I d for any a∈ [K]. Then λ 1 (Γ a Γ ⊤ a ) =∥b a ∥ 2 2 ≤ κ b . The second result follows from the fact that λ 1 (Γ a Γ ⊤ a ) = λ 1 (Γ ⊤ a Γ a ). Now we prove the result in Equation (A.13). This follows from the expression of ˆ Σ t,a in Lemma 3. Precisely, we have that ˆ Σ t,a = ̃ Σ t,a + ̃ Σ t,a Σ −1 0,a Γ a ̄ Σ t Γ ⊤ a Σ −1 0,a ̃ Σ t,a ,∀a∈ [K]. where ̃ Σ t,a = (︁ G t,a + Σ −1 0,a )︁ −1 . Thus Weyl’s inequality combined with the properties in Section A.1 yields that λ 1 ( ˆ Σ t,a )≤ λ 1 ( ̃ Σ t,a ) + λ 1 ( ̃ Σ t,a )λ 1 (Σ −1 0,a )λ 1 (Γ a ̄ Σ t Γ ⊤ a )λ 1 (Σ −1 0,a )λ 1 ( ̃ Σ t,a )≤ λ 1,0 + λ 2 1,0 λ 1,Ψ κ b λ 2 d,0 In the last inequality, we used that λ 1 (Γ a ̄ Σ t Γ ⊤ a ) ≤ λ 1 ( ̄ Σ t )λ 1 (Γ a Γ ⊤ a ), ((f) in Section A.1), λ 1 (Σ −1 0,a )≤ 1 λ d,0 , and λ 1 ( ̃ Σ t,a )≤ λ 1,0 . Finally, we prove the result in Equation (A.14). First, we rewrite the precision matrix of the effect posterior ̄ Σ −1 t using the compact notation introduced in Section A.3.1. Precisely, it follows from Equation (A.4) that ̄ Σ −1 t (i) = Σ −1 Ψ + K ∑︂ a=1 Γ ⊤ a (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ Γ a , (i) = Σ −1 Ψ + K ∑︂ a=1 Γ ⊤ a (︂ Σ −1 0,a − Σ −1 0,a ̃ Σ t,a Σ −1 0,a )︂ Γ a . 132 Here, (i) and (i) are the same; (i) follows from plugging ̃ Σ t,a = (G t,a + Σ −1 0,a ) −1 in (i). Then we have that λ 1 (Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ ) = λ 1 (︂ I Ld + K ∑︂ a=1 Σ 1 2 Ψ Γ ⊤ a (︂ Σ −1 0,a − Σ −1 0,a ̃ Σ T +1,a Σ −1 0,a )︂ Γ a Σ 1 2 Ψ )︂ ≤ 1 + λ 1,Ψ K ∑︂ a=1 λ 1 (︂ Γ ⊤ a (︂ Σ −1 0,a − Σ −1 0,a ̃ Σ T +1,a Σ −1 0,a )︂ Γ a )︂ , ≤ 1 + λ 1,Ψ K ∑︂ a=1 λ 1 (Γ ⊤ a Γ a )λ 1 (︂ Σ −1 0,a − Σ −1 0,a ̃ Σ T +1,a Σ −1 0,a )︂ ≤ 1 + λ 1,Ψ K ∑︂ a=1 κ b (︂ λ 1 (︁ Σ −1 0,a )︁ + λ 1 (︂ −Σ −1 0,a ̃ Σ T +1,a Σ −1 0,a )︂)︂ , ≤ 1 + λ 1,Ψ K ∑︂ a=1 κ b (︃ 1 λ d,0 − λ d (︂ Σ −1 0,a ̃ Σ T +1,a Σ −1 0,a )︂ )︃ ≤ 1 + λ 1,Ψ K ∑︂ a=1 κ b (︃ 1 λ d,0 − λ d (︁ Σ −1 0,a )︁ λ d (︂ ̃ Σ T +1,a )︂ λ d (︁ Σ −1 0,a )︁ )︃ , ≤ 1 + λ 1,Ψ K ∑︂ a=1 κ b (︃ 1 λ d,0 − 1 λ 2 1,0 λ d (︂ ̃ Σ T +1,a )︂ )︃ = 1 + λ 1,Ψ K ∑︂ a=1 κ b (︄ 1 λ d,0 − 1 λ 2 1,0 λ 1 (︁ G T +1,a + Σ −1 0,a )︁ )︄ , ≤ 1 + λ 1,Ψ K ∑︂ a=1 κ b ⎛ ⎝ 1 λ d,0 − 1 λ 2 1,0 (︂ κ x T σ 2 + 1 λ d,0 )︂ ⎞ ⎠ = 1 + Kλ 1,Ψ κ b ⎛ ⎝ 1 λ d,0 − 1 λ 2 1,0 (︂ κ x T σ 2 + 1 λ d,0 )︂ ⎞ ⎠ . A.3.4 Regret Proof Here we prove a more general version of Theorem 1 where we do not assume that the covariance matrices Σ 0,a and Σ Ψ are diagonal. We still assume that there exists κ x > 0 such that ∥X t ∥ 2 2 ≤ κ x for any t∈ [T ]. Theorem 6 (General version of Theorem 1). For any δ ∈ (0, 1), the Bayes regret of meTS in the mixed-effect model in Section 3.1.1 is bounded as BR(T )≤ √︁ 2T (R a (T ) +R e (T )) log(1/δ) + cTδ ,(A.15) 133 with c = √︃ 2 π κ x (︁ λ 1,0 + λ 2 1,0 λ 1,Ψ κ b λ 2 d,0 )︁ K , κ b = max a∈[K] ∥b a ∥ 2 2 , λ 1,0 = max a∈[K] λ 1 (Σ 0,a ), λ d,0 = min a∈[K] λ d (Σ 0,a ), λ 1,Ψ = λ 1 (Σ Ψ ) and R a (T ) = dKc a log (︁ 1 + Tκ x λ 1,0 σ 2 d )︁ , c a = κ x λ 1,0 log(1 + κ x λ 1,0 σ 2 ) , R e (T ) = dLc e log (︂ 1 + Kκ b λ 1,Ψ (︂ 1 λ d,0 − 1 λ 2 1,0 (︁ κ x T σ 2 + 1 λ d,0 )︁ )︂)︂ , c e = κ x κ b λ 2 1,0 λ 1,Ψ (︁ 1 + κ x λ 1,0 σ 2 )︁ λ 2 d,0 log (︁ 1 + κ x κ b λ 2 1,0 λ 1,Ψ σ 2 λ 2 d,0 )︁ . In particular, the result in Theorem 1 is retrieved when λ 1,0 = λ d,0 = σ 2 0 , and λ 1,Ψ = σ 2 Ψ . Proof. Consider our model rewritten in Equation (A.10). Then, the posterior distribution of the action parameter θ ∗,a | H t is a multivariate Gaussian distribution N (ˆμ t,a , ˆ Σ t,a ) for some ˆμ t,a ∈ R d and ˆ Σ t,a ∈ R d×d (Lemma 3). Now we let θ t,∗ = (X ⊤ t θ ∗,a ) a∈[K] ∈ R K be the concatenation of the expected rewards of actions in round t. Notice that the context X t is known in round t, and thus we include it in the history H t . This is important, with slight abuse of notation, H t now denotes H t ← H t ∪X t . Then, the joint posterior of the expected rewards, θ t,∗ | H t , is also a multivariate Gaussian N ( ˇ θ t , ˇ Σ t ) for ˇ θ t = (X ⊤ t ˆμ t,a ) a∈[K] ∈ R K and some covariance ˇ Σ t ∈ R K×K . This follows from the properties of Gaussian distributions (Koller and Friedman, 2009) and the fact that X t is now included in H t . Let A t ∈ 0, 1 K and A t,∗ ∈ 0, 1 K be indicator vectors of the taken action A t and optimal action A t,∗ , respectively (The vector representations are in bold letters while the integer representations are in regular letters). Then the Bayes regret can be rewritten and consequently decomposed following standard analysis (Russo and Van Roy, 2014) as BR(T ) = E [︄ T ∑︂ t=1 X ⊤ t θ ∗,A t,∗ − X ⊤ t θ ∗,A t ]︄ ,(A.16) = E [︂ T ∑︂ t=1 A ⊤ t,∗ θ t,∗ − A ⊤ t θ t,∗ ]︂ ,(A.17) = T ∑︂ t=1 E [︁ E [︁ A ⊤ t,∗ (θ t,∗ − ˇ θ t ) ⃓ ⃓ H t ]︁]︁ + E [︁ E [︁ A ⊤ t ( ˇ θ t − θ t,∗ ) ⃓ ⃓ H t ]︁]︁ . This follows from the fact that ˇ θ t = (X ⊤ t ˆμ t,i ) i∈[K] is deterministic given H t (since H t now includes X t ), and that A t,∗ and A t are i.i.d. given H t . Moreover, given H t , ˇ θ t −θ t,∗ is a zero- mean multivariate random variable independent of A t and thus E[A ⊤ t ( ˇ θ t −θ t,∗ )| H t ] = 0. Therefore, we only need to bound the first term in (A.16). With slight abuse of notation, letA be the set of all possible indicator vectors of actions a ∈ [K]. Precisely, an action a∈ [K] is also represented by an indicator vector a∈A⊂0, 1 K (in bold letter). Then we define the following events E t,a (δ) = ︂ |a ⊤ (θ t,∗ − ˇ θ t )|≤ √︁ 2 log(1/δ)∥a∥ ˇ Σ t ︂ , ∀δ ∈ (0, 1), ∀a∈A. Fix history H t , we split the expectation over the two complementary events E t,A t,∗ (δ) and 134 ̄ E t,A t,∗ (δ), and use the Cauchy-Schwarz inequality to obtain E [︁ A ⊤ t,∗ (θ t,∗ − ˇ θ t ) ⃓ ⃓ H t ]︁ ≤ √︁ 2 log(1/δ) E [︁ ∥A t,∗ ∥ ˇ Σ t ⃓ ⃓ H t ]︁ (A.18) + E [︁ A ⊤ t,∗ (θ t,∗ − ˇ θ t )1 ︁ ̄ E t,A t,∗ (δ) ︁ ⃓ ⃓ H t ]︁ . Now the second term in Equation (A.18) can be bounded as follows. For any a∈A, let Z a = a ⊤ (θ t,∗ − ˇ θ t ). Then we have that E [︁ A ⊤ t,∗ (θ t,∗ − ˇ θ t )1 ︁ ̄ E t,A t,∗ (δ) ︁ ⃓ ⃓ H t ]︁ (i) = E [︂ Z A t,∗ 1 ︂ |Z A t,∗ | > √︁ 2 log(1/δ)∥A t,∗ ∥ ˇ Σ t ︂ ⃓ ⃓ ⃓ H t ]︂ , (i) ≤ E [︂ |Z A t,∗ |1 ︂ |Z A t,∗ | > √︁ 2 log(1/δ)∥A t,∗ ∥ ˇ Σ t ︂ ⃓ ⃓ ⃓ H t ]︂ , (i) ≤ ∑︂ a∈A E [︂ |Z a |1 ︂ |Z a | > √︁ 2 log(1/δ)∥a∥ ˇ Σ t ︂ ⃓ ⃓ ⃓ H t ]︂ , (iv) ≤ ∑︂ a∈A 2 ∥a∥ ˇ Σ t √ 2π ∫︂ ∞ u= √ 2 log(1/δ)∥a∥ ˇ Σ t u exp [︄ − u 2 2∥a∥ 2 ˇ Σ t ]︄ du, (v) ≤ ∑︂ a∈A ∥a∥ ˇ Σ t 2 √ 2π ∫︂ ∞ u= √ 2 log(1/δ) u exp [︃ − u 2 2 ]︃ du (vi) ≤ √︃ 2 π λ max,t Kδ .(A.19) In (i), we simply rewrite the terms using the random variable Z A t,∗ . In (i), we use the fact that Z A t,∗ ≤ |Z A t,∗ |. In (i), we upper bound the expectation of the ran- dom variable |Z A t,∗ |1 ︂ |Z A t,∗ | > √︁ 2 log(1/δ)∥A t,∗ ∥ ˇ Σ t ︂ with the sum of the expectations of |Z a |1 ︂ |Z a | > √︁ 2 log(1/δ)∥a∥ ˇ Σ t ︂ for a∈A since all these random variables are non- negative. Moreover, (iv) follows from the facts that given H t , Z a ∼ N (0,∥a∥ 2 ˇ Σ t ), and that if Z ∼ N (0,σ 2 ), then for any ε ≥ 0, P(|Z| > ε) ≤ 2P(Z > ε). In (v), we use the change of variables u ← u/∥a∥ ˇ Σ t . Finally, in (vi), we compute the integral and set λ max,t = max a∈A ∥a∥ ˇ Σ t . We combine Equation (A.18) and Equation (A.19) with the fact that A t and A t,∗ are i.i.d. given H t to obtain that E [︁ A ⊤ t,∗ (θ t,∗ − ˇ θ t ) ⃓ ⃓ H t ]︁ ≤ √︁ 2 log(1/δ) E [︁ ∥A t ∥ ˇ Σ t ⃓ ⃓ H t ]︁ + √︃ 2 π λ max,t Kδ .(A.20) The bound in Equation (A.20) holds for any history H t and thus we take an additional expectation and get that BR(T ) = E [︄ T ∑︂ t=1 A ⊤ t,∗ θ t,∗ − A ⊤ t θ t,∗ ]︄ ≤ √︁ 2 log(1/δ) E [︄ T ∑︂ t=1 ∥A t ∥ ˇ Σ t ]︄ + √︃ 2 π Kδ T ∑︂ t=1 λ max,t , (i) ≤ √︁ 2T log(1/δ) E ⎡ ⎣ ⌜ ⃓ ⃓ ⎷ T ∑︂ t=1 ∥A t ∥ 2 ˇ Σ t ⎤ ⎦ + √︃ 2 π Kδ T ∑︂ t=1 λ max,t , (i) ≤ √︁ 2T log(1/δ) ⌜ ⃓ ⃓ ⎷ E [︄ T ∑︂ t=1 ∥A t ∥ 2 ˇ Σ t ]︄ + √︃ 2 π Kδ T ∑︂ t=1 λ max,t , 135 where we use the Cauchy-Schwarz inequality in (i), and (i) follows from the concavity of the square root. Now note that any a ∈A is an indicator vector and that ˇ Σ t is the covariance of the joint posterior of the expected rewards (X ⊤ t θ ∗,a ) a∈[K] | H t . Therefore, for any a∈A,∥a∥ 2 ˇ Σ t = ˇσ 2 a is the variance of X ⊤ t θ ∗,a | H t . But we know that θ ∗,a | H t is a mul- tivariate Gaussian and its covariance is ˆ Σ t,a (Lemma 3). Thus the variance of X ⊤ θ ∗,a | H t is ˇσ 2 a = X ⊤ t ˆ Σ t,a X t . It follows that for any a∈A, ∥a∥ 2 ˇ Σ t = X ⊤ t ˆ Σ t,a X t =∥X t ∥ 2 ˆ Σ t,a . In par- ticular, ∥A t ∥ 2 ˇ Σ t = X ⊤ t ˆ Σ t,A t X t . Combining this with Equation (A.13) yields that λ max,t = max a∈A ∥a∥ ˇ Σ t = max a∈A ∥X t ∥ ˆ Σ t,a ≤ max a∈A √︂ λ 1 ( ˆ Σ t,a )κ x ≤ √︃ (︂ λ 1,0 + λ 2 1,0 λ 1,Ψ κ b λ 2 d,0 )︂ κ x . Then we let c = √︃ 2 π (︂ λ 1,0 + λ 2 1,0 λ 1,Ψ κ b λ 2 d,0 )︂ κ x K which allows us to write BR(T )≤ √︁ 2T log(1/δ) ⌜ ⃓ ⃓ ⎷ E [︄ T ∑︂ t=1 ∥X t ∥ 2 ˆ Σ t,A t ]︄ + cTδ .(A.21) Now we focus on the the term √︃ E [︂ ∑︁ T t=1 ∥X t ∥ 2 ˆ Σ t,A t ]︂ that we decompose and bound as ∥X t ∥ 2 ˆ Σ t,A t = σ 2 X ⊤ t ˆ Σ t,A t X t σ 2 (i) = σ 2 (︂ σ −2 X ⊤ t ̃ Σ t,A t X t + σ −2 X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t )︂ , (i) ≤ c a log(1 + σ −2 X ⊤ t ̃ Σ t,A t X t ) + c 1 log(1 + σ −2 X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t ), (A.22) where (i) follows from ˆ Σ t,A t = ̃ Σ t,A t + ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t , and we use the fol- lowing inequality in (i) x = x log(1 + x) log(1 + x)≤ (︃ max x∈[0,u] x log(1 + x) )︃ log(1 + x) = u log(1 + u) log(1 + x), which holds for any x∈ [0,u], where constants c a and c 1 are derived as c a = κ x λ 1,0 log(1 + σ −2 κ x λ 1,0 ) , c 1 = c Ψ log(1 + σ −2 c Ψ ) , c Ψ = κ x κ b λ 2 1,0 λ 1,Ψ λ 2 d,0 , The derivation of c a uses that X ⊤ t ̃ Σ t,A t X t ≤ λ 1 ( ̃ Σ t,A t )∥X t ∥ 2 ≤ λ −1 d (Σ −1 0,A t + G t,A t )κ x ≤ λ −1 d (Σ −1 0,A t )κ x = λ 1 (Σ 0,A t )κ x ≤ λ 1,0 κ x . The derivation of c 1 follows from X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t ≤ λ 2 1 ( ̃ Σ t,A t )λ 2 1 (Σ −1 0,A t )λ 1 (Γ A t ̄ Σ t Γ ⊤ A t )κ x , ≤ λ 2 1 (Σ 0,A t )λ 1,Ψ λ 1 (Γ A t Γ ⊤ A t )κ x λ 2 d (Σ 0,A t ) , ≤ λ 2 1,0 λ 1,Ψ κ b κ x λ 2 d,0 . 136 The first inequality follows from Weyl’s inequality and the fact that λ 1 ( ̄ Σ t ) ≤ λ 1 (Σ Ψ ) = λ 1,Ψ and λ 1 ( ̃ Σ t,A t ) ≤ λ 1 (Σ 0,A t ). Now we focus on bounding the logarithmic terms in Equation (A.22). First Term in Equation (A.22) We first rewrite this term as log(1 + σ −2 X ⊤ t ̃ Σ t,A t X t ) (i) = log det(I d + σ −2 ̃ Σ 1 2 t,A t X t X ⊤ t ̃ Σ 1 2 t,A t ), = log det( ̃ Σ −1 t,A t + σ −2 X t X ⊤ t )− log det( ̃ Σ −1 t,A t ), = log det( ̃ Σ −1 t+1,A t )− log det( ̃ Σ −1 t,A t ), where (i) follows from the Weinstein–Aronszajn identity. Now note that for any a̸= A t , the arm-a precision does not update at round t, hence ̃ Σ −1 t+1,a = ̃ Σ −1 t,a and the increment is zero; therefore we may sum over all a ∈ [K] without changing the value. Then, we sum over all rounds t∈ [T ], and get a telescoping that leads to T ∑︂ t=1 log det(I d + σ −2 ̃ Σ 1 2 t,A t X t X ⊤ t ̃ Σ 1 2 t,A t ) = T ∑︂ t=1 log det( ̃ Σ −1 t+1,A t )− log det( ̃ Σ −1 t,A t ), = T ∑︂ t=1 K ∑︂ a=1 log det( ̃ Σ −1 t+1,a )− log det( ̃ Σ −1 t,a ) = K ∑︂ a=1 T ∑︂ t=1 log det( ̃ Σ −1 t+1,a )− log det( ̃ Σ −1 t,a ), = K ∑︂ a=1 log det( ̃ Σ −1 T +1,a )− log det( ̃ Σ −1 1,a ), (i) = K ∑︂ a=1 log det(Σ 1 2 0,a ̃ Σ −1 T +1,a Σ 1 2 0,a ) (i) ≤ K ∑︂ a=1 d log (︃ 1 d Tr(Σ 1 2 0,a ̃ Σ −1 T +1,a Σ 1 2 0,a ) )︃ ≤ K ∑︂ a=1 d log (︃ 1 + κ x λ 1 (Σ 0,a )T σ 2 d )︃ ≤ Kd log (︃ 1 + κ x λ 1,0 T σ 2 d )︃ . where (i) follows from the fact that ̃ Σ 1,a = Σ 0,a and we use the inequality of arithmetic and geometric means in (i). Second Term in Equation (A.22) First, we rewrite the covariance matrix of the effect posterior ̄ Σ t using the compact notation introduced in Section A.3.1. Precisely, it follows from Equation (A.4) that ̄ Σ −1 t (i) = Σ −1 Ψ + K ∑︂ a=1 Γ ⊤ a (︁ Σ −1 0,a − Σ −1 0,a (G t,a + Σ −1 0,a ) −1 Σ −1 0,a )︁ Γ a , (i) = Σ −1 Ψ + K ∑︂ a=1 Γ ⊤ a (︂ Σ −1 0,a − Σ −1 0,a ̃ Σ t,a Σ −1 0,a )︂ Γ a .(A.23) Recall that (i) and (i) are the same; (i) follows from plugging ̃ Σ t,a = (G t,a + Σ −1 0,a ) −1 in 137 (i). Now let u = σ −1 ̃ Σ 1 2 t,A t X t . Then it follows from (i) in Equation (A.23) that ̄ Σ −1 t+1 − ̄ Σ −1 t = Γ ⊤ A t (︂ Σ −1 0,A t − Σ −1 0,A t ( ̃ Σ −1 t,A t + σ −2 X t X ⊤ t ) −1 Σ −1 0,A t − (Σ −1 0,A t − Σ −1 0,A t ̃ Σ t,A t Σ −1 0,A t ) )︂ Γ A t , = Γ ⊤ A t (︂ Σ −1 0,A t ( ̃ Σ t,A t − ( ̃ Σ −1 t,A t + σ −2 X t X ⊤ t ) −1 )Σ −1 0,A t )︂ Γ A t , = Γ ⊤ A t (︂ Σ −1 0,A t ̃ Σ 1 2 t,A t (I d − (I d + σ −2 ̃ Σ 1 2 t,A t X t X ⊤ t ̃ Σ 1 2 t,A t ) −1 ) ̃ Σ 1 2 t,A t Σ −1 0,A t )︂ Γ A t , = Γ ⊤ A t (︂ Σ −1 0,A t ̃ Σ 1 2 t,A t (I d − (I d + u ⊤ ) −1 ) ̃ Σ 1 2 t,A t Σ −1 0,A t )︂ Γ A t , (i) = Γ ⊤ A t (︃ Σ −1 0,A t ̃ Σ 1 2 t,A t u ⊤ 1 + u ⊤ u ̃ Σ 1 2 t,A t Σ −1 0,A t )︃ Γ A t , = σ −2 Γ ⊤ A t (︃ Σ −1 0,A t ̃ Σ t,A t X t X ⊤ t 1 + u ⊤ u ̃ Σ t,A t Σ −1 0,A t )︃ Γ A t .(A.24) In (i) we use the Sherman-Morrison formula. Moreover, we have that∥X t ∥ 2 ≤ κ x . There- fore, 1 + u ⊤ u = 1 + σ −2 X ⊤ t ̃ Σ t,A t X t ≤ 1 + σ −2 κ x λ 1 (Σ 0,A t )≤ 1 + σ −2 κ x λ 1,0 = c 2 . This allows us to bound the second logarithmic term in Equation (A.22) as log(1 + σ −2 X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t ), (i) ≤ c 2 log(1 + c −1 2 σ −2 X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t ), (i) = c 2 log det(I Ld + c −1 2 σ −2 ̄ Σ 1 2 t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ 1 2 t ), (i) = c 2 [︂ log det( ̄ Σ −1 t + c −1 2 σ −2 Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t )− log det( ̄ Σ −1 t ) ]︂ , (iv) ≤ c 2 [︃ log det( ̄ Σ −1 t + σ −2 Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t X ⊤ t 1 + u ⊤ u ̃ Σ t,A t Σ −1 0,A t Γ A t )− log det( ̄ Σ −1 t ) ]︃ , (v) = c 2 [︁ log det( ̄ Σ −1 t+1 )− log det( ̄ Σ −1 t ) ]︁ . Here (i) follows from the fact that log(1 +x)≤ c 2 log(1 +c −1 2 x) for any x≥ 0 and c 2 ≥ 1. In (i), we use the Weinstein–Aronszajn identity. In (i), we use the log product formula and the fact that the det is a multiplicative map. In (iv), we use that c −1 2 ≤ 1/(1 +u ⊤ u). Finally, (v) follows from Equation (A.24). Now we sum over all rounds and get telescoping T ∑︂ t=1 log(1 + σ −2 X ⊤ t ̃ Σ t,A t Σ −1 0,A t Γ A t ̄ Σ t Γ ⊤ A t Σ −1 0,A t ̃ Σ t,A t X t ), ≤ c 2 [︁ log det( ̄ Σ −1 T +1 )− log det( ̄ Σ −1 1 ) ]︁ = c 2 log det(Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ ), (i) ≤ c 2 Ld log (︃ 1 Ld Tr(Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ ) )︃ , (i) ≤ c 2 Ld log(λ 1 (Σ 1 2 Ψ ̄ Σ −1 T +1 Σ 1 2 Ψ )) (i) ≤ c 2 Ld log (︂ 1 + Kκ b λ 1,Ψ (︁ 1 λ d,0 − 1 λ 2 1,0 (︁ κ x T σ 2 + 1 λ d,0 )︁ )︁ )︂ , 138 In (i) we use the inequality of arithmetic and geometric means. In (i) we bound all eigenvalues in the trace by the maximum eigenvalue. In (i) we use the result in Equa- tion (A.14). We combine the upper bounds for both logarithmic terms and get E [︄ T ∑︂ t=1 ∥X t ∥ 2 ˆ Σ t,A t ]︄ ≤ Kdc a log (︁ 1 + κ x λ 1,0 T σ 2 d )︁ + Ldc 1 c 2 log (︂ 1 + Kκ b λ 1,Ψ (︁ 1 λ d,0 − 1 λ 2 1,0 (︁ κ x T σ 2 + 1 λ d,0 )︁ )︁ )︂ . Finally, we set c e = c 1 c 2 , which concludes the proof for the general case. To retrieve the result in Theorem 1, we only need to set λ 1,0 = λ d,0 = σ 2 0 and λ 1,Ψ = σ 2 Ψ since we assumed that Σ Ψ = σ 2 Ψ I Ld and that Σ 0,a = σ 2 0 I d for any a ∈ [K]. In that case, the second term simplifies as log (︂ 1 + Kκ b λ 1,Ψ (︁ 1 λ d,0 − 1 λ 2 1,0 (︁ κ x T σ 2 + 1 λ d,0 )︁ )︁ )︂ = log (︂ 1 + Kκ b σ 2 Ψ (︁ 1 σ 2 0 − 1 σ 4 0 (︁ κ x T σ 2 + 1 σ 2 0 )︁ )︁ )︂ , = log (︁ 1 + Kκ b σ 2 Ψ Tκ x Tκ x σ 2 0 + σ 2 )︁ . A.4 Additional Experiments We provide additional experiments where we evaluate meTS using synthetic and real-world problems, and compare it to baselines that either ignore or partially use effect parameters. In each plot, we report the averages and standard errors of the quantities. Both settings are described in Section 3.4. A.4.1 Synthetic Experiments In Figures A.1 and A.2, we report regret from 12 experiments with horizon T = 5000, where we vary K and d and use both linear and logistic rewards. For the linear setting, we compare meTS-Lin (Section 3.2.2), LinUCB (Abbasi-Yadkori et al., 2011), LinTS (Agrawal and Goyal, 2013a) and HierTS (Hong et al., 2022b). For the logistic setting, we compare meTS-GLM (Section 3.2.3), meTS-Lin (Section 3.2.2), UCB-GLM (Li et al., 2017), GLM-TS (Chapelle and Li, 2012) and HierTS (Hong et al., 2022b). We also include the factored approximation of meTS (meTS-Lin-Fa and meTS-GLM-Fa). In all experiments, we observe that meTS-Lin and meTS-Fa outperform other baselines that ignore the effect parameters or incorporate them partially. We also notice that the gain in performance becomes smaller when K/L decreases. 139 010002000300040005000 0 500 1000 1500 2000 2500 3000 3500 Regret Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinUCB LinTS 010002000300040005000 0 200 400 600 800 1000 1200 1400 1600 1800 Linear bandit: K = 50, L = 3, d = 2 010002000300040005000 0 100 200 300 400 500 600 700 800 Linear bandit: K = 20, L = 3, d = 2 010002000300040005000 0 1000 2000 3000 4000 5000 6000 7000 8000 Regret Linear bandit: K = 100, L = 3, d = 5 010002000300040005000 0 500 1000 1500 2000 2500 3000 3500 4000 4500 Linear bandit: K = 50, L = 3, d = 5 010002000300040005000 0 200 400 600 800 1000 1200 1400 1600 1800 Linear bandit: K = 20, L = 3, d = 5 Figure A.1: Regret of meTS-Lin on synthetic linear bandit problems with varying feature dimension d∈2, 5 and number of actions K ∈20, 50, 100. 010002000300040005000 0 200 400 600 800 1000 1200 1400 Regret Logistic bandit: K = 100, L = 3, d = 2 meTS-GLM meTS-GLM-Fa meTS-Lin GLM-TS UCB-GLM HierTS 010002000300040005000 0 100 200 300 400 500 600 700 800 Logistic bandit: K = 50, L = 3, d = 2 010002000300040005000 0 100 200 300 400 500 600 Logistic bandit: K = 20, L = 3, d = 2 010002000300040005000 0 200 400 600 800 1000 1200 1400 1600 Regret Logistic bandit: K = 100, L = 3, d = 5 010002000300040005000 0 200 400 600 800 1000 1200 Logistic bandit: K = 50, L = 3, d = 5 010002000300040005000 0 100 200 300 400 500 600 700 800 900 Logistic bandit: K = 20, L = 3, d = 5 Figure A.2: Regret of meTS-GLM on synthetic logistic bandit problems with varying feature dimension d∈2, 5 and number of actions K ∈20, 50, 100. A.4.2 MovieLens Experiments We plot the regret of meTS and the baselines up to T = 5000 rounds in Figures A.3 and A.4. We observe that meTS outperforms the other baselines. This is despite the fact that we did not fine-tune the mixing weights, which attests to the robustness of our approach to model misspecification. Similarly to the synthetic problems, we observe that the gap in performance between meTS and other baselines is less significant when K/L is small. 140 010002000300040005000 0 200 400 600 800 1000 Regret Linear bandit: K = 100, L = 5, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinTS 010002000300040005000 0 50 100 150 200 250 300 350 400 Linear bandit: K = 50, L = 5, d = 2 010002000300040005000 0 50 100 150 200 250 Linear bandit: K = 20, L = 5, d = 2 010002000300040005000 0 200 400 600 800 1000 1200 1400 Regret Linear bandit: K = 100, L = 5, d = 5 010002000300040005000 0 100 200 300 400 500 600 700 Linear bandit: K = 50, L = 5, d = 5 010002000300040005000 0 50 100 150 200 250 300 350 Linear bandit: K = 20, L = 5, d = 5 Figure A.3: Regret of meTS-Lin on the MovieLens dataset with linear rewards and varying feature dimension d∈2, 5 and number of actions K ∈20, 50, 100. 010002000300040005000 0 50 100 150 200 250 Regret Logistic bandit: K = 100, L = 5, d = 2 meTS-GLM meTS-GLM-Fa meTS-Lin HierTS GLM-TS 010002000300040005000 0 50 100 150 200 Logistic bandit: K = 50, L = 5, d = 2 010002000300040005000 0 20 40 60 80 100 120 140 160 Logistic bandit: K = 20, L = 5, d = 2 010002000300040005000 0 50 100 150 200 250 300 350 Regret Logistic bandit: K = 100, L = 5, d = 5 010002000300040005000 0 50 100 150 200 250 300 Logistic bandit: K = 50, L = 5, d = 5 010002000300040005000 0 50 100 150 200 250 Logistic bandit: K = 20, L = 5, d = 5 Figure A.4: Regret of meTS-GLM on the MovieLens dataset with logistic rewards and varying feature dimension d∈2, 5 and number of actions K ∈20, 50, 100. A.4.3 Robustness to Model Misspecification We conduct additional synthetic experiments where the hyper-parameters do not match the parameters of the bandit environment to assess the robustness of our approach to mis- specification. We provide results for this experiment in Figure A.5. Here we consider the setting described in Section 3.4.1 except that the true hyper-parameters are misspecified as follows. At each run, we sample uniformly 4 misspecification constants c 1 ,c 2 ,c 3 , and c 4 from (0, 2) and set the hyper-parameters of meTS-Lin as c 1 Σ Ψ , c 2 μ Ψ , c 3 Σ 0,a , and c 4 b a for any a∈ [K]; where Σ Ψ , μ Ψ , Σ 0,a , and b a for a∈ [K] are the true hyper-parameters. Model 141 misspecification is only applied to meTS-Lin and we refer to it as meTS-Lin-mis. We com- pare it to meTS-Lin and the other baselines, all with the true hyper-parameters. Although the baselines are not misspecified, meTS-Lin-mis still performs better. meTS-Lin-mis also performs similarly to meTS-Lin (with true hyper-parameters). 010002000300040005000 0 500 1000 1500 2000 2500 3000 3500 Regret Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-mis HierTS LinUCB LinTS 010002000300040005000 0 200 400 600 800 1000 1200 1400 1600 1800 Linear bandit: K = 50, L = 3, d = 2 010002000300040005000 0 100 200 300 400 500 600 700 800 Linear bandit: K = 20, L = 3, d = 2 Figure A.5: Regret of misspecified meTS-Lin on synthetic bandit problems with a varying number of actions K. Here, the misspecified meTS, meTS-Lin-mis, is compared to baselines with true hyper-parameters. A.4.4 Effect of Action Uncertainty As we mentioned in Section 3.4.1 and predicted by our Bayes regret bound, learning the effect parameters is most beneficial when they are more uncertain than the action parameters. In this section, we support this claim by conducting an experiment where the initial uncertainty of action parameters is greater than the initial uncertainty of the effect parameters. Precisely, we consider the setting described in Section 3.4.1 except that we set Σ Ψ = I Ld and Σ 0,a = 3I d for all a ∈ [K]. We report the results in Figure A.6. By comparing Figure A.6 to Figure 3.2, we observe that meTS-Lin still outperforms the baselines but the gap in performance shrinks when the action parameters are more uncertain than the effect parameters. 010002000300040005000 0 500 1000 1500 2000 2500 3000 3500 4000 Regret Linear bandit: K = 100, L = 3, d = 2 meTS-Lin meTS-Lin-Fa HierTS LinUCB LinTS 010002000300040005000 0 200 400 600 800 1000 1200 1400 1600 1800 Linear bandit: K = 50, L = 3, d = 2 010002000300040005000 0 100 200 300 400 500 600 700 800 Linear bandit: K = 20, L = 3, d = 2 Figure A.6: Regret of meTS-Lin on synthetic bandit problems with a varying number of actions K, where the action parameters are more uncertain than the effect parameters. 142 Chapter B Supplementary Materials for Chapter 4 Contents A.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2 Posterior Derivations . . . . . . . . . . . . . . . . . . . . . . . . . . . 125 A.2.1 Effect Posterior Derivation . . . . . . . . . . . . . . . . . . . . 125 A.2.2 Action Posterior Derivation . . . . . . . . . . . . . . . . . . . 128 A.3 Regret Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 129 A.3.1 Problem Reformulation for Regret Analysis . . . . . . . . . . 130 A.3.2 Derivation of cov [θ ∗,a |H t ] . . . . . . . . . . . . . . . . . . . . 130 A.3.3 Preliminary Eigenvalues Results . . . . . . . . . . . . . . . . . 131 A.3.4 Regret Proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . 133 A.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 139 A.4.1 Synthetic Experiments . . . . . . . . . . . . . . . . . . . . . . 139 A.4.2 MovieLens Experiments . . . . . . . . . . . . . . . . . . . . . 140 A.4.3 Robustness to Model Misspecification . . . . . . . . . . . . . . 141 A.4.4 Effect of Action Uncertainty . . . . . . . . . . . . . . . . . . . 142 B.1 Posterior for Linear Diffusion Models Our posterior approximation builds on the simplified setting where the diffusion model is fully linear, i.e., each link function f ℓ is linear in ψ ℓ . This linear case, studied in our earlier workshop paper (Aouali, 2023), serves as the analytical foundation for our posterior approximation used in the general non-linear case. In Section B.2, we show how the exact posteriors derived in this linear setting inspire our efficient approximation, which extends naturally to practical diffusion models that are typically highly non-linear. 143 B.1.1 Linear Diffusion Models Here, we assume the link functions f ℓ are linear such as f ℓ (ψ ℓ ) = W ℓ ψ ℓ for ℓ∈ [L], where W ℓ ∈ R d×d are known mixing matrices. Then, Equation (4.1) becomes a linear Gaussian system (LGS) (Bishop, 2006) and can be summarized as follows ψ L ∼N (0, Σ L+1 ),(B.1) ψ ℓ−1 | ψ ℓ ∼N (W ℓ ψ ℓ , Σ ℓ ),∀ℓ∈ [L]/1, θ a | ψ 1 ∼N (W 1 ψ 1 , Σ 1 ),∀a∈ [K], R t | X t ,A t ,θ, (ψ ℓ ) ℓ∈[L] ∼ p(·| X t ;θ A t ),∀t∈ [T ]. This model is important because it yields closed-form posteriors when the reward dis- tribution is linear-Gaussian, i.e., p(· | x;θ a ) = N (·;x ⊤ θ a ,σ 2 ). This allows bounding the Bayes regret of sDM. For practice, the posterior expressions are used to motivate efficient approximations for the general case in Equation (4.1) as we show in Section 4.2.1. B.1.2 Posterior Expressions for Linear Diffusion Models Recall that the reward distribution is modeled as a generalized linear model (GLM) (Mc- Cullagh and Nelder, 1989), allowing for non-linear rewards even when the diffusion links are linear. This non-linearity in the reward distribution prevents closed-form posteriors. However, since the non-linearity arises only through the reward likelihood, we approx- imate it by a Gaussian, leading to efficient posterior updates that are exact whenever the reward model itself is Gaussian; a special case of the GLM framework. Precisely, let ˆ B t,a and ˆ G t,a denote the MLE (see the remark below for practical considerations) and the Hessian of the negative log-likelihood, respectively: ˆ B t,a = argmax θ a ∈R d ∑︂ i∈S t,a logp(R i | X i ;θ a ), ˆ G t,a = ∑︂ i∈S t,a ̇g (︁ X ⊤ i ˆ B t,a )︁ X i X ⊤ i , (B.2) where S t,a =i∈ [t− 1] : A i = a is the set of rounds in which action a was taken up to round t. We approximate the likelihood as ∏︂ i∈S t,a p(R i | X i ;θ a ) ∝ exp (︂ − 1 2 (θ a − ˆ B t,a ) ⊤ ˆ G t,a (θ a − ˆ B t,a ) )︂ ,(B.3) which makes all subsequent posteriors Gaussian. Once this approximation is done, all other derivations of the action posterior and latent posteriors are exact. Remark 8. The MLE may be ill-posed. In practice, we maximize an ℓ 2 -regularized esti- mator. Action posterior. The conditional action posterior becomes p(θ a | ψ 1 ,H t,a ) ≈ N (︁ θ a ; ˆμ t,a , ˆ Σ t,a )︁ , with parameters ˆ Σ −1 t,a = Σ −1 1 + ˆ G t,a ,ˆμ t,a = ˆ Σ t,a (︂ Σ −1 1 W 1 ψ 1 + ˆ G t,a ˆ B t,a )︂ .(B.4) 144 Latent posteriors. For each ℓ∈ [L]\1, the conditional latent posterior is p(ψ ℓ−1 | ψ ℓ ,H t ) ≈ N (︁ ψ ℓ−1 ; ̄μ t,ℓ−1 , ̄ Σ t,ℓ−1 )︁ , where ̄ Σ −1 t,ℓ−1 = Σ −1 ℓ + ̄ G t,ℓ−1 , ̄μ t,ℓ−1 = ̄ Σ t,ℓ−1 (︁ Σ −1 ℓ W ℓ ψ ℓ + ̄ B t,ℓ−1 )︁ .(B.5) The top-layer posterior is p(ψ L | H t ) ≈ N (︁ ψ L ; ̄μ t,L , ̄ Σ t,L )︁ , with ̄ Σ −1 t,L = Σ −1 L+1 + ̄ G t,L , ̄μ t,L = ̄ Σ t,L ̄ B t,L .(B.6) Recursive updates. The matrices ̄ G t,ℓ and ̄ B t,ℓ for ℓ∈ [L] are defined recursively. The base recursion is ̄ G t,1 = W ⊤ 1 K ∑︂ a=1 (︁ Σ −1 1 − Σ −1 1 ˆ Σ t,a Σ −1 1 )︁ W 1 , ̄ B t,1 = W ⊤ 1 Σ −1 1 K ∑︂ a=1 ˆ Σ t,a ˆ G t,a ˆ B t,a . (B.7) Then, for ℓ∈ [L]\1, the recursive step is ̄ G t,ℓ = W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ W ℓ , ̄ B t,ℓ = W ⊤ ℓ Σ −1 ℓ ̄ Σ t,ℓ−1 ̄ B t,ℓ−1 .(B.8) Discussion. This completes the derivation of the linear posterior approximation. All pos- teriors are Gaussian and exact whenever the reward distribution follows a linear-Gaussian model, i.e. p(·| x;θ a ) =N (·;x ⊤ θ a ,σ 2 ). In this case, the above posterior updates coincide with the exact Bayesian updates, while for general GLMs they serve as efficient and accurate approximations. B.2 Posterior for Non-Linear Diffusion Models The general diffusion model (Equation (4.1), which is our case of interest) involves two sources of non-linearity: (i) the reward distribution p(· | x;θ), which may follow a non- linear generalized linear model (GLM), and (i) the diffusion links f ℓ (ψ ℓ ), which can be arbitrary non-linear functions. Both sources make the posterior intractable, and therefore two approximations are needed. First approximation (likelihood). We first approximate the reward likelihood by a Gaussian density (as we did above in Equation (B.3)). After this substitution, the model becomes conditionally Gaussian given the latent variables. This step is exact when the reward model is linear-Gaussian, and approximate otherwise. Second approximation (diffusion hierarchy). Even after the likelihood is approxi- mated, the diffusion hierarchy remains non-linear because of the non-linear mappings f ℓ . To handle this, we reuse the exact Gaussian posteriors derived for the linear diffusion case (Section B.1.2) and generalize them as follows: 145 • Replace each linear mapping W ℓ ψ ℓ by its non-linear counterpart f ℓ (ψ ℓ ), which rep- resents the mean of the diffusion prior at layer ℓ. • Remove matrix multiplications involving W ℓ in the recursive updates. This step can be viewed as extending the linear-Gaussian posterior updates to a general non-linear setting. This allows fast sampling and updating of the posterior, without heavy standard posterior approximation techniques. Of course, this is a purely empirical and heuristic based approximation that does not come with guarantees but performs very well in practice. Resulting approximation. The two steps above yield a posterior where each conditional factor p(θ a | ψ 1 ,H t,a ) and p(ψ ℓ−1 | ψ ℓ ,H t ) remains Gaussian with updated means and covariances, while the overall model retains the hierarchical diffusion structure. The approximation satisfies two desirable properties: it exactly recovers the diffusion prior when no data is available, and as more data is observed, the likelihood terms dominate and the prior influence fades naturally. B.3 Connection to Two-Level Hierarchies The linear diffusion Equation (B.1) can be marginalized into a 2-level hierarchy using two different strategies. To simplify, we let Σ ℓ = σ 2 ℓ I d . The first one yields, ψ L ∼N (0,σ 2 L+1 B L B ⊤ L ),(B.9) θ a | ψ L ∼N (ψ L , Ω 1 ),∀a∈ [K], with Ω 1 = σ 2 1 I d + ∑︁ L−1 ℓ=1 σ 2 ℓ+1 B ℓ B ⊤ ℓ and B ℓ = ∏︁ ℓ i=1 W i . The second strategy yields, ψ 1 ∼N (0, Ω 2 ),(B.10) θ a | ψ 1 ∼N (ψ 1 , σ 2 1 I d ),∀a∈ [K], where Ω 2 = ∑︁ L ℓ=1 σ 2 ℓ+1 B ℓ B ⊤ ℓ . Recently, HierTS (Hong et al., 2022b) was developed for such two-level graphical models, and we call HierTS under Equation (B.9) by HierTS-1 and HierTS under Equation (B.10) by HierTS-2. Then, we start by highlighting the differences between these two variants of HierTS. First, their regret bounds scale as HierTS-1 : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ Td(K L ∑︂ ℓ=1 σ 2 ℓ + Lσ 2 L+1 )︁ , HierTS-2 : ̃ O (︁ ⌜ ⃓ ⃓ ⎷ Td(Kσ 2 1 + L ∑︂ ℓ=1 σ 2 ℓ+1 ) )︁ . When K ≈ L, the regret bounds of HierTS-1 and HierTS-2 are similar. However, when K > L, HierTS-2 outperforms HierTS-1. This is because HierTS-2 puts more uncertainty on a single d-dimensional latent parameter ψ 1 , rather than K individual d- dimensional action parameters θ a . More importantly, HierTS-1 implicitly assumes that action parameters θ a are conditionally independent given ψ L , which is not true. Conse- quently, HierTS-2 outperforms HierTS-1. Note that, under the linear diffusion model Equation (B.1), sDM and HierTS-2 have roughly similar regret bounds. Specifically, their regret bounds dependency on K is identical, where both methods involve multiplying K by σ 2 1 , and both enjoy improved performance compared to HierTS-1. That said, note that 146 Theorem 7 and Proposition 7 provide an understanding of how sDM’s regret scales under linear link functions f ℓ , and do not say that using sDM is better than using HierTS when the link functions f ℓ are linear since the latter can be obtained by a proper marginaliza- tion of latent parameters (i.e., HierTS-2 instead of HierTS-1). While such a comparison is not the goal of this work, we still provide it for completeness next. When the mixing matrices W ℓ are dense (i.e., assumption (A5) is not applicable), sDM and HierTS-2 have comparable regret bounds and computational efficiency. However, under the sparsity assumption (A5) and with mixing matrices that allow for conditional independence of ψ 1 coordinates given ψ 2 , sDM enjoys a computational advantage over HierTS-2. This advantage explains why works focusing on multi-level hierarchies typically benchmark their algorithms against two-level structures akin to HierTS-1, rather than the more competitive HierTS-2. This is also consistent with prior works in Bayesian bandits using multi-level hierarchies, such as Tree-based priors (Hong et al., 2022a), which compared their method to HierTS-1. In line with this, we also compared sDM with HierTS-1 in our experiments. But this is only given for completeness as this is not the aim of Theorem 7 and Proposition 7. More importantly, HierTS is inapplicable in the general case in Equation (4.1) with non-linear link functions since the latent parameters cannot be analytically marginalized. B.4 Formal Theory We analyze sDM assuming that: (A0) The true environment parameters θ ∗ and ψ ∗,ℓ are drawn from the same prior distribution used by sDM, as is standard in Bayes regret analysis (Russo and Van Roy, 2014); we thus use θ ∗ and θ (ψ ∗,ℓ and ψ ℓ ) interchangeably throughout. (A1) The rewards are linear p(· | x;θ a ) = N (·;x ⊤ θ a ,σ 2 ). (A2) The link functions f ℓ are linear such as f ℓ (ψ ℓ ) = W ℓ ψ ℓ for ℓ ∈ [L], where W ℓ ∈ R d×d are known mixing matrices. This leads to a structure with L layers of linear Gaussian relationships detailed in Section B.1.1. In particular, this leads to closed-form posteriors given in Section B.1.2 that inspired our approximation and enable theory similar to linear bandits (Agrawal and Goyal, 2013a). However, proofs are not the same, and technical challenges remain (explained in Section B.5). Although our result holds for milder assumptions, we make additional simplifications for clarity and interpretability. We assume that (A3) Contexts satisfy ∥X t ∥ 2 2 = 1 for any t ∈ [T ]. Note that (A3) can be relaxed to any contexts X t with bounded norms ∥X t ∥ 2 . (A4) Mixing matrices and covariances satisfy λ 1 (W ⊤ ℓ W ℓ ) = 1 for any ℓ∈ [L] and Σ ℓ = σ 2 ℓ I d for any ℓ ∈ [L + 1]. In this section, we write ̃ O for the big-O notation up to polylogarithmic factors. We start by stating our bound for sDM. Theorem 7. Let σ 2 max = max ℓ∈[L+1] 1 + σ 2 ℓ σ 2 . There exists a constant c > 0 such that for any δ ∈ (0, 1), the Bayes regret of sDM under (A1), (A2), (A3) and (A4) is bounded 147 as BR(T )≤ ⌜ ⃓ ⃓ ⎷ 2T (︁ R act (T ) + L ∑︂ ℓ=1 R lat ℓ )︁ log(1/δ) )︂ + cTδ , R act (T ) = c 0 dK log (︁ 1 + Tσ 2 1 dσ 2 )︁ , c 0 = σ 2 1 log (︂ 1 + σ 2 1 σ 2 )︂ , R lat ℓ = c ℓ d log (︁ 1 + σ 2 ℓ+1 σ 2 ℓ )︁ ,c ℓ = σ 2 ℓ+1 σ 2ℓ max log (︂ 1 + σ 2 ℓ+1 σ 2 )︂ ,(B.11) Equation (B.11) holds for any δ ∈ (0, 1). In particular, the term cTδ is constant when δ = 1/T. Then, the bound is ̃ O (︂ √︂ T (dKσ 2 1 + d ∑︁ L ℓ=1 σ 2 ℓ+1 σ 2ℓ max ) )︂ , and this dependence on the horizon T aligns with prior Bayes regret bounds. The bound comprises L + 1 main terms,R act (T ) andR lat ℓ for ℓ∈ [L]. First,R act (T ) relates to action parameters learning, conforming to a standard form (Lu and Van Roy, 2019). Similarly,R lat ℓ is associated with learning the ℓ-th latent parameter. To include more structure, we propose the sparsity assumption (A5) W ℓ = ( ̄ W ℓ , 0 d,d−d ℓ ), where ̄ W ℓ ∈ R d×d ℓ for any ℓ ∈ [L]. Note that (A5) is not an assumption when d ℓ = d for any ℓ ∈ [L]. Notably, (A5) incorporates a plausible structural characteristic that a diffusion model could capture. Proposition 7 (Sparsity). Let σ 2 max = max ℓ∈[L+1] 1 + σ 2 ℓ σ 2 . There exists a constant c > 0 such that for any δ ∈ (0, 1), the Bayes regret of sDM under (A1), (A2), (A3), (A4) and (A5) is bounded as BR(T )≤ ⌜ ⃓ ⃓ ⎷ 2T (︁ R act (T ) + L ∑︂ ℓ=1 ̃ R lat ℓ )︁ log(1/δ) )︂ + cTδ , ̃ R act (T ) = c 0 dK log (︁ 1 + Tσ 2 1 dσ 2 )︁ , c 0 = σ 2 1 log (︂ 1 + σ 2 1 σ 2 )︂ , R lat ℓ = c ℓ d ℓ log (︁ 1 + σ 2 ℓ+1 σ 2 ℓ )︁ ,c ℓ = σ 2 ℓ+1 σ 2ℓ max log (︂ 1 + σ 2 ℓ+1 σ 2 )︂ ,(B.12) From Proposition 7, our bounds scales as BR(T ) = ̃ O (︂ ⌜ ⃓ ⃓ ⎷ T (dKσ 2 1 + L ∑︂ ℓ=1 d ℓ σ 2 ℓ+1 σ 2ℓ max ) )︂ .(B.13) B.5 Regret proof Important notation clarification. Throughout this proof, we operate under the stan- dard Bayesian bandit framework where the true environment parameters θ ∗ are drawn 148 from the same prior distribution that sDM uses for posterior inference. Specifically, the true action parameters θ ∗,a for a ∈ [K] and the true latent parameters ψ ∗,ℓ for ℓ ∈ [L] are assumed to be sampled according to the generative process in Equation (B.1). As a consequence, the true parameters θ ∗ and the model parameters θ used in our derivations follow the same distribution, and we use them interchangeably throughout the proof to simplify notation. B.5.1 Proof Sketch We start with the following standard lemma upon which we build our analysis (Aouali et al., 2023b). Lemma 4. Assume that p(θ a | H t ) = N (θ a ; ˇμ t,a , ˇ Σ t,a ) for any a ∈ [K], then for any δ ∈ (0, 1), BR(T )≤ √︁ 2T log(1/δ) ⌜ ⃓ ⃓ ⎷ E [︄ T ∑︂ t=1 ∥X t ∥ 2 ˇ Σ t,A t ]︄ + cTδ ,where c > 0 is a constant. (B.14) Applying Lemma 4 requires proving that the marginal action-posterior densities of θ a | H t in Equation (4.3) are Gaussian and computing their covariances, while we only know the conditional action-posteriors p(θ a | ψ 1 ,H t ) and latent-posteriors p(ψ ℓ−1 | ψ ℓ ,H t ). This is achieved by leveraging the preservation properties of the family of Gaussian distributions (Koller and Friedman, 2009) and the total covariance decomposition (Weiss, 2005) which leads to the next lemma. Lemma 5. Let t∈ [T ] and a∈ [K], then the marginal covariance matrix ˇ Σ t,a reads ˇ Σ t,a = ˆ Σ t,a + ∑︂ ℓ∈[L] P a,ℓ ̄ Σ t,ℓ P ⊤ a,ℓ , where P a,ℓ = ˆ Σ t,a Σ −1 1 W 1 ℓ−1 ∏︂ i=1 ̄ Σ t,i Σ −1 i+1 W i+1 . (B.15) The marginal covariance matrix ˇ Σ t,a in Equation (B.15) decomposes into L + 1 terms. The first term corresponds to the posterior uncertainty of θ a | ψ 1 . The remaining L terms capture the posterior uncertainties of ψ L and ψ ℓ−1 | ψ ℓ for ℓ ∈ [L]/1. These are then used to quantify the posterior information gain of latent parameters after one round as follows. Lemma 6 (Posterior information gain). Let t∈ [T ] and ℓ∈ [L], then ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ ⪰ σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ , where σ 2 max = max ℓ∈[L+1] 1 + σ 2 ℓ σ 2 .(B.16) Finally, Lemma 5 is used to decompose ∥X t ∥ 2 ˇ Σ t,A t in Equation (B.14) into L + 1 terms. Each term is bounded thanks to Lemma 6. This results in the Bayes regret bound in Theorem 7. 149 B.5.2 Proof of lemma 5 In this proof, we heavily rely on the total covariance decomposition (Weiss, 2005). Also, refer to (Hong et al., 2022b, Section 5.2) for a brief introduction to this decomposition. Now, from Equation (B.4), we have that cov [θ a |H t ,ψ 1 ] = ˆ Σ t,a = (︂ ˆ G t,a + Σ −1 1 )︂ −1 , E [θ a |H t ,ψ 1 ] = ˆμ t,a = ˆ Σ t,a (︂ ˆ G t,a ˆ B t,a + Σ −1 1 W 1 ψ 1 )︂ . First, given H t , cov [θ a |H t ,ψ 1 ] = (︂ ˆ G t,a + Σ −1 1 )︂ −1 is constant. Thus E [cov [θ a |H t ,ψ 1 ]|H t ] = cov [θ a |H t ,ψ 1 ] = (︂ ˆ G t,a + Σ −1 1 )︂ −1 = ˆ Σ t,a . In addition, given H t , ˆ Σ t,a , ˆ G t,a and ˆ B t,a are constant. Thus cov [E [θ a |H t ,ψ 1 ]|H t ] = cov [︂ ˆ Σ t,a (︂ ˆ G t,a ˆ B t,a + Σ −1 1 W 1 ψ 1 )︂ ⃓ ⃓ ⃓ H t ]︂ , = cov [︂ ˆ Σ t,a Σ −1 1 W 1 ψ 1 ⃓ ⃓ ⃓ H t ]︂ , = ˆ Σ t,a Σ −1 1 W 1 cov [ψ 1 |H t ] W ⊤ 1 Σ −1 1 ˆ Σ t,a , = ˆ Σ t,a Σ −1 1 W 1 Σ t,1 W ⊤ 1 Σ −1 1 ˆ Σ t,a , whereΣ t,1 = cov [ψ 1 |H t ] is the marginal posterior covariance of ψ 1 . Finally, the total covariance decomposition (Weiss, 2005; Hong et al., 2022b) yields that ˇ Σ t,a = cov [θ a |H t ] = E [cov [θ a |H t ,ψ 1 ]|H t ] + cov [E [θ a |H t ,ψ 1 ]|H t ] , = ˆ Σ t,a + ˆ Σ t,a Σ −1 1 W 1 Σ t,1 W ⊤ 1 Σ −1 1 ˆ Σ t,a ,(B.17) However, Σ t,1 = cov [ψ 1 |H t ] is different from ̄ Σ t,1 = cov [ψ 1 |H t ,ψ 2 ] that we already derived in Equation (B.5). Thus we do not know the expression ofΣ t,1 . But we can use the same total covariance decomposition trick to find it. Precisely, letΣ t,ℓ = cov [ψ ℓ |H t ] for any ℓ∈ [L]. Then we have that ̄ Σ t,1 = cov [ψ 1 |H t ,ψ 2 ] = (︁ Σ −1 2 + ̄ G t,1 )︁ −1 , ̄μ t,1 = E [ψ 1 |H t ,ψ 2 ] = ̄ Σ t,1 (︂ Σ −1 2 W 2 ψ 2 + ̄ B t,1 )︂ . First, given H t , cov [ψ 1 |H t ,ψ 2 ] = (︁ Σ −1 2 + ̄ G t,1 )︁ −1 is constant. Thus E [cov [ψ 1 |H t ,ψ 2 ]|H t ] = cov [ψ 1 |H t ,ψ 2 ] = ̄ Σ t,1 . In addition, given H t , ̄ Σ t,1 , ̃ Σ t,1 and ̄ B t,1 are constant. Thus cov [E [ψ 1 |H t ,ψ 2 ]|H t ] = cov [︂ ̄ Σ t,1 (︂ Σ −1 2 W 2 ψ 2 + ̄ B t,1 )︂ ⃓ ⃓ ⃓ H t ]︂ , = cov [︁ ̄ Σ t,1 Σ −1 2 W 2 ψ 2 ⃓ ⃓ H t ]︁ , = ̄ Σ t,1 Σ −1 2 W 2 cov [ψ 2 |H t ] W ⊤ 2 Σ −1 2 ̄ Σ t,1 , = ̄ Σ t,1 Σ −1 2 W 2 Σ t,2 W ⊤ 2 Σ −1 2 ̄ Σ t,1 . 150 Finally, total covariance decomposition (Weiss, 2005; Hong et al., 2022b) leads to Σ t,1 = cov [ψ 1 |H t ] = E [cov [ψ 1 |H t ,ψ 2 ]|H t ] + cov [E [ψ 1 |H t ,ψ 2 ]|H t ] , = ̄ Σ t,1 + ̄ Σ t,1 Σ −1 2 W 2 Σ t,2 W ⊤ 2 Σ −1 2 ̄ Σ t,1 . Now using the techniques, this can be generalized using the same technique as above to Σ t,ℓ = ̄ Σ t,ℓ + ̄ Σ t,ℓ Σ −1 ℓ+1 W ℓ+1 Σ t,ℓ+1 W ⊤ ℓ+1 Σ −1 ℓ+1 ̄ Σ t,ℓ ,∀ℓ∈ [L− 1]. Then, by induction, we get that Σ t,1 = ∑︂ ℓ∈[L] ̄ P ℓ ̄ Σ t,ℓ ̄ P ⊤ ℓ ,∀ℓ∈ [L− 1], where we use that by definitionΣ t,L = cov [ψ L |H t ] = ̄ Σ t,L and set ̄ P 1 = I d and ̄ P ℓ = ∏︁ ℓ−1 i=1 ̄ Σ t,i Σ −1 i+1 W i+1 for any ℓ∈ [L]/1. Plugging this in Equation (B.17) leads to ˇ Σ t,a = ˆ Σ t,a + ∑︂ ℓ∈[L] ˆ Σ t,a Σ −1 1 W 1 ̄ P ℓ ̄ Σ t,ℓ ̄ P ⊤ ℓ W ⊤ 1 Σ −1 1 ˆ Σ t,a , = ˆ Σ t,a + ∑︂ ℓ∈[L] ˆ Σ t,a Σ −1 1 W 1 ̄ P ℓ ̄ Σ t,ℓ ( ˆ Σ t,a Σ −1 1 W 1 ) ⊤ , = ˆ Σ t,a + ∑︂ ℓ∈[L] P a,ℓ ̄ Σ t,ℓ P ⊤ a,ℓ , where P a,ℓ = ˆ Σ t,a Σ −1 1 W 1 ̄ P ℓ = ˆ Σ t,a Σ −1 1 W 1 ∏︁ ℓ−1 i=1 ̄ Σ t,i Σ −1 i+1 W i+1 . B.5.3 Proof of lemma 6 We prove this result by induction. We start with the base case when ℓ = 1. (I) Base case. Let u = σ −1 ˆ Σ 1 2 t,A t X t From the expression of ̄ Σ t,1 in Equation (B.5), we have that ̄ Σ −1 t+1,1 − ̄ Σ −1 t,1 = W ⊤ 1 (︂ Σ −1 1 − Σ −1 1 ( ˆ Σ −1 t,A t + σ −2 X t X ⊤ t ) −1 Σ −1 1 − (Σ −1 1 − Σ −1 1 ˆ Σ t,A t Σ −1 1 ) )︂ W 1 , = W ⊤ 1 (︂ Σ −1 1 ( ˆ Σ t,A t − ( ˆ Σ −1 t,A t + σ −2 X t X ⊤ t ) −1 )Σ −1 1 )︂ W 1 , = W ⊤ 1 (︂ Σ −1 1 ˆ Σ 1 2 t,A t (I d − (I d + σ −2 ˆ Σ 1 2 t,A t X t X ⊤ t ˆ Σ 1 2 t,A t ) −1 ) ˆ Σ 1 2 t,A t Σ −1 1 )︂ W 1 , = W ⊤ 1 (︂ Σ −1 1 ˆ Σ 1 2 t,A t (I d − (I d + u ⊤ ) −1 ) ˆ Σ 1 2 t,A t Σ −1 1 )︂ W 1 , (i) = W ⊤ 1 (︃ Σ −1 1 ˆ Σ 1 2 t,A t u ⊤ 1 + u ⊤ u ˆ Σ 1 2 t,A t Σ −1 1 )︃ W 1 , (i) = σ −2 W ⊤ 1 Σ −1 1 ˆ Σ t,A t X t X ⊤ t 1 + u ⊤ u ˆ Σ t,A t Σ −1 1 W 1 .(B.18) In (i) we use the Sherman-Morrison formula. Note that (i) says that ̄ Σ −1 t+1,1 − ̄ Σ −1 t,1 is one-rank which we will also need in induction step. Now, we have that ∥X t ∥ 2 = 1. Therefore, 1 + u ⊤ u = 1 + σ −2 X ⊤ t ˆ Σ t,A t X t ≤ 1 + σ −2 λ 1 (Σ 1 )∥X t ∥ 2 = 1 + σ −2 σ 2 1 ≤ σ 2 max , 151 where we use that by definition of σ 2 max in Lemma 6, we have that σ 2 max ≥ 1 + σ −2 σ 2 1 . Therefore, by taking the inverse, we get that 1 1+u ⊤ u ≥ σ −2 max . Combining this with Equa- tion (B.18) leads to ̄ Σ −1 t+1,1 − ̄ Σ −1 t,1 ⪰ σ −2 σ −2 max W ⊤ 1 Σ −1 1 ˆ Σ t,A t X t X ⊤ t ˆ Σ t,A t Σ −1 1 W 1 Noticing that P A t ,1 = ˆ Σ t,A t Σ −1 1 W 1 concludes the proof of the base case when ℓ = 1. (I) Induction step. Let ℓ∈ [L]/1 and suppose that ̄ Σ −1 t+1,ℓ−1 − ̄ Σ −1 t,ℓ−1 is one-rank and that it holds for ℓ− 1 that ̄ Σ −1 t+1,ℓ−1 − ̄ Σ −1 t,ℓ−1 ⪰ σ −2 σ −2(ℓ−1) max P ⊤ A t ,ℓ−1 X t X ⊤ t P A t ,ℓ−1 , where σ 2 max = max ℓ∈[L+1] 1 + σ −2 σ 2 ℓ . Then, we want to show that ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ is also one-rank and that it holds that ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ ⪰ σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ ,where σ 2 max = max ℓ∈[L+1] 1 + σ −2 σ 2 ℓ . This is achieved as follows. Define the precision increment at level ℓ− 1 by ∆ t,ℓ−1 := ̄ Σ −1 t+1,ℓ−1 − ̄ Σ −1 t,ℓ−1 . By the induction hypothesis, ∆ t,ℓ−1 is rank-one PSD, hence there exists u∈ R d such that ∆ t,ℓ−1 = u ⊤ . Using Equation (B.8), we have for any s∈t,t + 1: ̄ Σ −1 s,ℓ = Σ −1 ℓ+1 + ̄ G s,ℓ = Σ −1 ℓ+1 + W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ s,ℓ−1 Σ −1 ℓ )︁ W ℓ . Therefore, ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ = ̄ G t+1,ℓ − ̄ G t,ℓ = W ⊤ ℓ Σ −1 ℓ (︁ ̄ Σ t,ℓ−1 − ̄ Σ t+1,ℓ−1 )︁ Σ −1 ℓ W ℓ . Since ̄ Σ −1 t+1,ℓ−1 = ̄ Σ −1 t,ℓ−1 + u ⊤ , Sherman–Morrison yields ̄ Σ t+1,ℓ−1 = (︁ ̄ Σ −1 t,ℓ−1 + u ⊤ )︁ −1 = ̄ Σ t,ℓ−1 − ̄ Σ t,ℓ−1 u ⊤ ̄ Σ t,ℓ−1 1 + u ⊤ ̄ Σ t,ℓ−1 u . Hence, ̄ Σ t,ℓ−1 − ̄ Σ t+1,ℓ−1 = ̄ Σ t,ℓ−1 u ⊤ ̄ Σ t,ℓ−1 1 + u ⊤ ̄ Σ t,ℓ−1 u , and plugging this back gives ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ = W ⊤ ℓ Σ −1 ℓ ̄ Σ t,ℓ−1 u ⊤ 1 + u ⊤ ̄ Σ t,ℓ−1 u ̄ Σ t,ℓ−1 Σ −1 ℓ W ℓ . In particular, this increment is rank-one PSD. 152 However, it follows from the induction hypothesis that u ⊤ = ̄ Σ −1 t+1,ℓ−1 − ̄ Σ −1 t,ℓ−1 ⪰ σ −2 σ −2(ℓ−1) max P ⊤ A t ,ℓ−1 X t X ⊤ t P A t ,ℓ−1 . Therefore, ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ = W ⊤ ℓ Σ −1 ℓ ̄ Σ t,ℓ−1 u ⊤ 1 + u ⊤ ̄ Σ t,ℓ−1 u ̄ Σ t,ℓ−1 Σ −1 ℓ W ℓ , ⪰ W ⊤ ℓ Σ −1 ℓ ̄ Σ t,ℓ−1 σ −2 σ −2(ℓ−1) max P ⊤ A t ,ℓ−1 X t X ⊤ t P A t ,ℓ−1 1 + u ⊤ ̄ Σ t,ℓ−1 u ̄ Σ t,ℓ−1 Σ −1 ℓ W ℓ , = σ −2 σ −2(ℓ−1) max 1 + u ⊤ ̄ Σ t,ℓ−1 u W ⊤ ℓ Σ −1 ℓ ̄ Σ t,ℓ−1 P ⊤ A t ,ℓ−1 X t X ⊤ t P A t ,ℓ−1 ̄ Σ t,ℓ−1 Σ −1 ℓ W ℓ , = σ −2 σ −2(ℓ−1) max 1 + u ⊤ ̄ Σ t,ℓ−1 u P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ . Finally, we use that 1 + u ⊤ ̄ Σ t,ℓ−1 u ≤ 1 +∥u∥ 2 2 λ 1 ( ̄ Σ t,ℓ−1 ) ≤ 1 + σ −2 σ 2 ℓ . Here we use that ∥u∥ 2 2 ≤ σ −2 , which can also be proven by induction, and that λ 1 ( ̄ Σ t,ℓ−1 ) ≤ σ 2 ℓ , which follows from the expression of ̄ Σ t,ℓ−1 in Section B.1.2. Therefore, we have that ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ ⪰ σ −2 σ −2(ℓ−1) max 1 + u ⊤ ̄ Σ t,ℓ−1 u P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ , ⪰ σ −2 σ −2(ℓ−1) max 1 + σ −2 σ 2 ℓ P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ , ⪰ σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ , where the last inequality follows from the definition of σ 2 max = max ℓ∈[L+1] 1 +σ −2 σ 2 ℓ . This concludes the proof. B.5.4 Proof of theorem 7 We start with the following standard result which we borrow from (Hong et al., 2022a; Aouali et al., 2023b), BR(T )≤ √︁ 2T log(1/δ) ⌜ ⃓ ⃓ ⎷ E [︄ T ∑︂ t=1 ∥X t ∥ 2 ˇ Σ t,A t ]︄ + cTδ ,where c > 0 is a constant. (B.19) Then we use Lemma 5 and express the marginal covariance ˇ Σ t,A t as ˇ Σ t,a = ˆ Σ t,a + ∑︂ ℓ∈[L] P a,ℓ ̄ Σ t,ℓ P ⊤ a,ℓ , where P a,ℓ = ˆ Σ t,a Σ −1 1 W 1 ℓ−1 ∏︂ i=1 ̄ Σ t,i Σ −1 i+1 W i+1 . (B.20) 153 Therefore, we can decompose ∥X t ∥ 2 ˇ Σ t,A t as ∥X t ∥ 2 ˇ Σ t,A t = σ 2 X ⊤ t ˇ Σ t,A t X t σ 2 (i) = σ 2 (︁ σ −2 X ⊤ t ˆ Σ t,A t X t + σ −2 ∑︂ ℓ∈[L] X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )︁ , (i) ≤ c 0 log(1 + σ −2 X ⊤ t ˆ Σ t,A t X t ) + ∑︂ ℓ∈[L] c ℓ log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ),(B.21) where (i) follows from Equation (B.20), and we use the following inequality in (i) x = x log(1 + x) log(1 + x)≤ (︃ max x∈[0,u] x log(1 + x) )︃ log(1 + x) = u log(1 + u) log(1 + x), which holds for any x∈ [0,u], where constants c 0 and c ℓ are derived as c 0 = σ 2 1 log(1 + σ 2 1 σ 2 ) , c ℓ = σ 2 ℓ+1 log(1 + σ 2 ℓ+1 σ 2 ) . The derivation of c 0 uses that X ⊤ t ˆ Σ t,A t X t ≤ λ 1 ( ˆ Σ t,A t )∥X t ∥ 2 ≤ λ −1 d (Σ −1 1 + G t,A t )≤ λ −1 d (Σ −1 1 ) = λ 1 (Σ 1 ) = σ 2 1 . The derivation of c ℓ follows from X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ≤ λ 1 (P A t ,ℓ P ⊤ A t ,ℓ )λ 1 ( ̄ Σ t,ℓ )∥X t ∥ 2 ≤ σ 2 ℓ+1 . Therefore, from Equation (B.21) and Equation (B.19), we get that BR(T )≤ √︁ 2T log(1/δ) (︂ E [︂ c 0 T ∑︂ t=1 log(1 + σ −2 X ⊤ t ˆ Σ t,A t X t ) + ∑︂ ℓ∈[L] c ℓ T ∑︂ t=1 log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ) ]︂)︂ 1 2 + cTδ(B.22) Now we focus on bounding the logarithmic terms in Equation (B.22). (I) First term in Equation (B.22) We first rewrite this term as log(1 + σ −2 X ⊤ t ˆ Σ t,A t X t ) (i) = log det(I d + σ −2 ˆ Σ 1 2 t,A t X t X ⊤ t ˆ Σ 1 2 t,A t ), = log det( ˆ Σ −1 t,A t + σ −2 X t X ⊤ t )− log det( ˆ Σ −1 t,A t ) = log det( ˆ Σ −1 t+1,A t )− log det( ˆ Σ −1 t,A t ), where (i) follows from the Weinstein-Aronszajn identity. Then we sum over all rounds t∈ [T ], and get a telescoping T ∑︂ t=1 log det(I d + σ −2 ˆ Σ 1 2 t,A t X t X ⊤ t ˆ Σ 1 2 t,A t ) = T ∑︂ t=1 log det( ˆ Σ −1 t+1,A t )− log det( ˆ Σ −1 t,A t ), = T ∑︂ t=1 K ∑︂ a=1 log det( ˆ Σ −1 t+1,a )− log det( ˆ Σ −1 t,a ) = K ∑︂ a=1 T ∑︂ t=1 log det( ˆ Σ −1 t+1,a )− log det( ˆ Σ −1 t,a ), = K ∑︂ a=1 log det( ˆ Σ −1 T +1,a )− log det( ˆ Σ −1 1,a ) (i) = K ∑︂ a=1 log det(Σ 1 2 1 ˆ Σ −1 T +1,a Σ 1 2 1 ), 154 where (i) follows from the fact that ˆ Σ 1,a = Σ 1 . Now we use the inequality of arithmetic and geometric means and get T ∑︂ t=1 log det(I d + σ −2 ˆ Σ 1 2 t,A t X t X ⊤ t ˆ Σ 1 2 t,A t ) = K ∑︂ a=1 log det(Σ 1 2 1 ˆ Σ −1 T +1,a Σ 1 2 1 ), ≤ K ∑︂ a=1 d log (︃ 1 d Tr(Σ 1 2 1 ˆ Σ −1 T +1,a Σ 1 2 1 ) )︃ ,(B.23) ≤ K ∑︂ a=1 d log (︃ 1 + T d σ 2 1 σ 2 )︃ = Kd log (︃ 1 + T d σ 2 1 σ 2 )︃ . (I) Remaining terms in Equation (B.22) Let ℓ∈ [L]. Then we have that log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ) = σ 2ℓ max σ −2ℓ max log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ), ≤ σ 2ℓ max log(1 + σ −2 σ −2ℓ max X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t ), (i) = σ 2ℓ max log det(I d + σ −2 σ −2ℓ max ̄ Σ 1 2 t,ℓ P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ ̄ Σ 1 2 t,ℓ ), = σ 2ℓ max (︂ log det( ̄ Σ −1 t,ℓ + σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ )− log det( ̄ Σ −1 t,ℓ ) )︂ , where we use the Weinstein-Aronszajn identity in (i). Now we know from Lemma 6 that the following inequality holds σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ ⪯ ̄ Σ −1 t+1,ℓ − ̄ Σ −1 t,ℓ . As a result, we get that ̄ Σ −1 t,ℓ + σ −2 σ −2ℓ max P ⊤ A t ,ℓ X t X ⊤ t P A t ,ℓ ⪯ ̄ Σ −1 t+1,ℓ . Thus, log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )≤ σ 2ℓ max (︂ log det( ̄ Σ −1 t+1,ℓ )− log det( ̄ Σ −1 t,ℓ ) )︂ , Then we sum over all rounds t∈ [T ], and get a telescoping T ∑︂ t=1 log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )≤ σ 2ℓ max T ∑︂ t=1 log det( ̄ Σ −1 t+1,ℓ )− log det( ̄ Σ −1 t,ℓ ), = σ 2ℓ max (︂ log det( ̄ Σ −1 T +1,ℓ )− log det( ̄ Σ −1 1,ℓ ) )︂ , (i) = σ 2ℓ max (︂ log det( ̄ Σ −1 T +1,ℓ )− log det(Σ −1 ℓ+1 ) )︂ , = σ 2ℓ max (︂ log det(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) )︂ , where we use that ̄ Σ 1,ℓ = Σ ℓ+1 in (i). Finally, we use the inequality of arithmetic and geometric means and get that T ∑︂ t=1 log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )≤ σ 2ℓ max (︂ log det(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) )︂ , ≤ dσ 2ℓ max log (︃ 1 d Tr(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) )︃ , (B.24) ≤ dσ 2ℓ max log (︃ 1 + σ 2 ℓ+1 σ 2 ℓ )︃ , 155 The last inequality follows from the expression of ̄ Σ −1 T +1,ℓ in Equation (B.5) that leads to Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 = I d + Σ 1 2 ℓ+1 ̄ G T +1,ℓ Σ 1 2 ℓ+1 , = I d + Σ 1 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ W ℓ Σ 1 2 ℓ+1 ,(B.25) since ̄ G T +1,ℓ = W ⊤ ℓ (︁ Σ −1 ℓ −Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ W ℓ . This allows us to bound 1 d Tr(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) as 1 d Tr(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) = 1 d Tr(I d + Σ 1 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ W ℓ Σ 1 2 ℓ+1 ), = 1 d (d + Tr(Σ 1 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ W ℓ Σ 1 2 ℓ+1 ), ≤ 1 + 1 d d ∑︂ i=1 λ 1 (Σ 1 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ W ℓ Σ 1 2 ℓ+1 , ≤ 1 + 1 d d ∑︂ i=1 λ 1 (Σ ℓ+1 )λ 1 (W ⊤ ℓ W ℓ )λ 1 (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ T +1,ℓ−1 Σ −1 ℓ )︁ , ≤ 1 + 1 d d ∑︂ i=1 λ 1 (Σ ℓ+1 )λ 1 (W ⊤ ℓ W ℓ )λ 1 (︁ Σ −1 ℓ )︁ , ≤ 1 + 1 d d ∑︂ i=1 σ 2 ℓ+1 σ 2 ℓ = 1 + σ 2 ℓ+1 σ 2 ℓ ,(B.26) where we use the assumption that λ 1 (W ⊤ ℓ W ℓ ) = 1 (A4) and that λ 1 (Σ ℓ+1 ) = σ 2 ℓ+1 and λ 1 (Σ −1 ℓ ) = 1/σ 2 ℓ . This is because Σ ℓ = σ 2 ℓ I d for any ℓ ∈ [L + 1]. Finally, plugging Equations (B.23) and (B.24) in Equation (B.22) concludes the proof. B.5.5 Proof of proposition 7 We use exactly the same proof in Section B.5.4, with one change to account for the sparsity assumption (A5). The change corresponds to Equation (B.24). First, recall that Equation (B.24) writes T ∑︂ t=1 log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )≤ σ 2ℓ max (︂ log det(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) )︂ , where Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 = I d + Σ 1 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ W ℓ Σ 1 2 ℓ+1 , = I d + σ 2 ℓ+1 W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ W ℓ ,(B.27) where the second equality follows from the assumption that Σ ℓ+1 = σ 2 ℓ+1 I d . But notice that in our assumption, (A5), we assume that W ℓ = ( ̄ W ℓ , 0 d,d−d ℓ ), where ̄ W ℓ ∈ R d×d ℓ for any ℓ∈ [L]. Therefore, we have that for any d× d matrix B∈ R d×d , the following holds, W ⊤ ℓ BW ℓ = (︃ ̄ W ⊤ ℓ B ̄ W ℓ 0 d ℓ ,d−d ℓ 0 d−d ℓ ,d ℓ 0 d−d ℓ ,d−d ℓ )︃ . In particular, we have that W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ W ℓ = (︃ ̄ W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ ̄ W ℓ 0 d ℓ ,d−d ℓ 0 d−d ℓ ,d ℓ 0 d−d ℓ ,d−d ℓ )︃ . (B.28) 156 Therefore, plugging this in Equation (B.27) yields that Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 = (︃ I d ℓ + σ 2 ℓ+1 ̄ W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ ̄ W ℓ 0 d ℓ ,d−d ℓ 0 d−d ℓ ,d ℓ I d−d ℓ )︃ .(B.29) As a result, det(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) = det(I d ℓ + σ 2 ℓ+1 ̄ W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ ̄ W ℓ ). This allows us to move the problem from a d-dimensional one to a d ℓ -dimensional one. Then we use the inequality of arithmetic and geometric means and get that T ∑︂ t=1 log(1 + σ −2 X ⊤ t P A t ,ℓ ̄ Σ t,ℓ P ⊤ A t ,ℓ X t )≤ σ 2ℓ max (︂ log det(Σ 1 2 ℓ+1 ̄ Σ −1 T +1,ℓ Σ 1 2 ℓ+1 ) )︂ , = σ 2ℓ max log det(I d ℓ + σ 2 ℓ+1 ̄ W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ ̄ W ℓ ), ≤ d ℓ σ 2ℓ max log (︃ 1 d ℓ Tr(I d ℓ + σ 2 ℓ+1 ̄ W ⊤ ℓ (︁ Σ −1 ℓ − Σ −1 ℓ ̄ Σ t,ℓ−1 Σ −1 ℓ )︁ ̄ W ℓ ) )︃ , ≤ d ℓ σ 2ℓ max log (︃ 1 + σ 2 ℓ+1 σ 2 ℓ )︃ .(B.30) To get the last inequality, we use derivations similar to the ones we used in Equa- tion (B.26). Finally, the desired result in obtained by replacing Equation (B.24) by Equation (B.30) in the previous proof in Section B.5.4. B.6 Additional Experiments B.6.1 Swiss roll data Figure B.1 shows samples from the Swiss roll data and samples from generated by the pre-trained diffusion model for different pre-training sample sizes. (a) Diffusion pre-trained on 50 samples from the Swiss roll dataset. (b) Diffusion pre-trained on 10 3 samples from the Swiss roll dataset. (c) Diffusion pre-trained on 10 4 samples from the Swiss roll dataset. Figure B.1: True distribution of action parameters (blue) vs. distribution of pre-trained diffusion model (red). B.6.2 Diffusion models pre-training We used JAX for diffusion model pre-training, summarized as follows: 157 • Parameterization: Functions f ℓ are parameterized with a fully connected 2-layer neural network (N) with ReLU activation. The step ℓ is provided as input to capture the current sampling stage. Covariances are fixed (not learned) as Σ ℓ = σ 2 ℓ I d with σ ℓ increasing with ℓ. • Loss: Offline data samples are progressively noised over steps ℓ ∈ [L], creating increasingly noisy versions of the data following a predefined noise schedule (Ho et al., 2020). The N is trained to reverse this noise (i.e., denoise) by predicting the noise added at each step. The loss function measures the L 2 norm difference between the predicted and actual noise at each step, as explained in Ho et al. (2020). • Optimization: Adam optimizer with a 10 −3 learning rate was used. The N was trained for 20,000 epochs with a batch size of min(2048, pre-training sample size). We used CPUs for pre-training, which was efficient enough to conduct multiple ablation studies. • After pre-training: The pre-trained diffusion model is used as a prior for sDM and compared to LinTS as the reference baseline. In our ablation study, we plot the cumulative regret of LinTS in the last round divided by that of sDM. A ratio greater than 1 indicates that sDM outperforms LinTS, with higher values representing a larger performance gap. B.6.3 Quality of our posterior approximation To assess the quality of our posterior approximation, we consider the scenario where the true distribution of action parameters isN (0 d ,I d ) with d = 2 and rewards are linear. We pre-train a diffusion model using samples drawn from N (0 d ,I d ). We then consider two priors: the true priorN (0 d ,I d ) and the pre-trained diffusion model prior. This yields two posteriors: • P 1 : UsesN (0 d ,I d ) as the prior. P 1 is an exact posterior since the prior is Gaussian and rewards are linear-Gaussian. • P 2 : Uses the pre-trained diffusion model as the prior. P 2 is our approximate posterior. The learned diffusion model prior matches the true Gaussian prior (as seen in Figure B.2a). Thus, if our approximation is accurate, their posteriors P 1 and P 2 should also be similar. This is observed in Figure B.2b where the approximate posterior P 2 nearly matches the exact posterior P 1 . 158 (a) Gaussian distribution vs. diffusion model pre-trained on 10 3 samples drawn from it. (b) Exact posterior P 1 vs. approximate pos- terior P 2 after T = 100 rounds of interac- tions. Figure B.2: Assessing the quality of our posterior approximation. B.6.4 CIFAR Ablation CIFAR. In Figure 4.3a in Section 4.4.2, we showed that with only 10 pre-training samples, sDM outperforms LinTS on the Swiss-roll benchmark. We now extend this analysis to the vision dataset CIFAR (Krizhevsky et al., 2009) (similar results were obtained on MNIST (Zhu, 2018)). Our setting is similar to that in Hong et al. (2022a) and we use sDM’s variant that uses a single shared parameter θ ∈ R d (Remark 5 and Section 4.2.2) because it is more suited for this setting. These additional ablations on CIFAR confirm that sDM consistently benefits from offline pre-training, even when the true prior is not a diffusion model. Specifically, we vary the percentage of offline data used to train the prior and compare against both HierTS and LinTS. Table B.1: Regret improvement (%) of sDM on CIFAR. Offline Data (%) vs. HierTS vs. LinTS 1%69.11%87.74% 5%79.56%92.18% 25%80.65%92.48% 50%81.67%92.88% B.6.5 Bound comparison Here, we compare our bound in Theorem 7 to bounds of LinTS from the literature. 159 02004006008001000 Horizon n 0 100000 200000 300000 400000 500000 600000 Regret Linear Bandit: d=5, K=10000, L=5 Bayes Bound for dTS (Ours) Frequentist Bound for LinTS (a) Our bound vs. the frequentist bound of LinTS in Abeille and Lazaric (2017). 02004006008001000 Horizon n 0 10000 20000 30000 40000 50000 60000 70000 80000 90000 Regret Linear Bandit: d=5, K=10000, L=5 Bayes Bound for dTS (Ours) Bayes Bound for LinTS (b) Our bound vs. the standard Bayesian bound of LinTS. Figure B.3: Comparing our Bayesian regret bound of dTS to the frequentist and Bayesian bounds of LinTS. 160 Chapter C Supplementary Materials for Chapter 6 Contents B.1 Posterior for Linear Diffusion Models . . . . . . . . . . . . . . . . . . 143 B.1.1 Linear Diffusion Models . . . . . . . . . . . . . . . . . . . . . 144 B.1.2 Posterior Expressions for Linear Diffusion Models . . . . . . . 144 B.2 Posterior for Non-Linear Diffusion Models . . . . . . . . . . . . . . . . 145 B.3 Connection to Two-Level Hierarchies . . . . . . . . . . . . . . . . . . . 146 B.4 Formal Theory . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 147 B.5 Regret proof . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 148 B.5.1 Proof Sketch . . . . . . . . . . . . . . . . . . . . . . . . . . . . 149 B.5.2 Proof of lemma 5 . . . . . . . . . . . . . . . . . . . . . . . . . 150 B.5.3 Proof of lemma 6 . . . . . . . . . . . . . . . . . . . . . . . . . 151 B.5.4 Proof of theorem 7 . . . . . . . . . . . . . . . . . . . . . . . . 153 B.5.5 Proof of proposition 7 . . . . . . . . . . . . . . . . . . . . . . 156 B.6 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 157 B.6.1 Swiss roll data . . . . . . . . . . . . . . . . . . . . . . . . . . . 157 B.6.2 Diffusion models pre-training . . . . . . . . . . . . . . . . . . 157 B.6.3 Quality of our posterior approximation . . . . . . . . . . . . . 158 B.6.4 CIFAR Ablation . . . . . . . . . . . . . . . . . . . . . . . . . . 159 B.6.5 Bound comparison . . . . . . . . . . . . . . . . . . . . . . . . 159 C.1 Posterior Derivations Under Standard Priors Here we derive the posterior under the standard prior in Equation (6.2). These are standard derivations and we present them here for the sake of completeness. But first, we state the following standard assumption that allows posterior derivations. 161 Assumption 5 (Independence). (X,A) is independent of θ, and the θ a , for a ∈ A are independent. Derivation of p(θ a |D n ) for the standard prior in Equation (6.2). We start by recalling the standard prior in Equation (6.2) θ a ∼N (μ a , Σ a ),∀a∈A,(C.1) R| θ,X,A∼N (φ(X) ⊤ θ A ,σ 2 ), where N (μ a , Σ a ) is the prior on the action parameter θ a . Let θ = (θ a ) a∈A ∈ R dK , Σ A = diag(Σ a ) a∈A ∈ R dK×dK and μ A = (μ a ) a∈A ∈ R dK . Also, let u a ∈ 0, 1 K be the binary vector representing the action a. That is, u a,a = 1 and u a,a ′ = 0 for all a ′ ̸= a. Then we can rewrite the model in Equation (C.1) as θ ∼N (μ A , Σ A ),(C.2) R| θ,X,A∼N ((u A ⊗ φ(X)) ⊤ θ,σ 2 ). Then the joint action posterior p(θ |D n ) decomposes as p(θ |D n ) = p(θ | (X i ,A i ,R i ) i∈[n] ) (i) ∝ p((R i ) i∈[n] | θ, (X i ,A i ) i∈[n] )p(θ | (X i ,A i ) i∈[n] ), (i) = p((R i ) i∈[n] | θ, (X i ,A i ) i∈[n] )p(θ) (i) = ∏︂ i∈[n] p(R i | θ,X i ,A i )p(θ), (iv) = ∏︂ i∈[n] N (R i ; (u A i ⊗ φ(X i )) ⊤ θ,σ 2 )N (θ;μ A , Σ A ), (v) ∝ N (︃ θ; ˆμ A , (︂ ˆ Λ A )︂ −1 )︃ . In (i), we apply Bayes rule. In (i), we use that θ is independent of (X,A), and (i) follows from the assumption that R i | θ,X i ,A i are independent. Finally, in (iv), we replace the distribution by their Gaussian form, and in (v), we set ˆ Λ A = v ∑︁ n i=1 (u A i u ⊤ A i ⊗ φ(X i )φ(X i ) ⊤ ) + Λ A , and ˆμ A = ˆ Λ −1 A (︁ v ∑︁ n i=1 (u A i ⊗ φ(X i ))R i + Λ A μ A )︂ where v = σ −2 and Λ A = Σ −1 A = diag(Σ −1 a ) a∈A . Now notice that ˆ Λ A = diag(Σ −1 a + G a ) a∈A . Thus, p(θ |D n ) =N (θ; ˆμ A , ˆ Λ −1 A ) where ˆμ A = (ˆμ a ) a∈A and ˆ Λ A = diag( ˆ Λ a ) a∈A , with ˆ Λ a = Σ −1 a + G a , ˆ Λ a ˆμ a = Σ −1 a μ a + B a . Since the covariance matrix of p(θ |D n ) is diagonal by block, we know that the marginals θ a |D n also have a Gaussian density p(θ a |D n ) =N (θ a ; ˆμ a , ˆ Σ a ) where ˆ Σ a = ˆ Λ −1 a . C.2 Posterior Derivations Under Structured Priors Here we derive the posteriors under the structured prior in Equation (6.9). Precisely, we derive the latent posterior density of ψ |D n , the conditional posterior density of θ |D n ,ψ. Then, we derive the marginal posterior θ |D n . Posterior derivations rely on the following assumption. 162 Assumption 6 (Structured Independence). (i) (X,A) is independent of ψ and given ψ, (X,A) is independent of θ. (i) Given ψ, the θ a , for all a∈A are independent. C.2.1 Latent Posterior Derivation of p(ψ |D n ). First, recall that our model in Equation (6.9) reads ψ ∼N (μ, Σ), θ a | ψ ∼N (︂ W a ψ, Σ a )︂ ,∀a∈A, R| ψ,θ,X,A∼N (φ(X) ⊤ θ A ,σ 2 ).(C.3) Then we first rewrite it as ψ ∼N (μ, Σ), θ | ψ ∼N (︂ W A ψ, Σ A )︂ , R| ψ,θ,X,A∼N ((u A ⊗ φ(X)) ⊤ θ,σ 2 ).(C.4) Then the latent posterior is p(ψ | (X i ,A i ,R i ) i∈[n] )∝ p((R i ) i∈[n] | ψ, (X i ,A i ) i∈[n] )p(ψ | (X i ,A i ) i∈[n] ), (i) = p((R i ) i∈[n] | ψ, (X i ,A i ) i∈[n] )q(ψ), = ∫︂ θ p((R i ) i∈[n] ,θ | ψ, (X i ,A i ) i∈[n] ) dθq(ψ), = ∫︂ θ p((R i ) i∈[n] | ψ,θ, (X i ,A i ) i∈[n] )p(θ | ψ, (X i ,A i ) i∈[n] ) dθq(ψ), (i) = ∫︂ θ p((R i ) i∈[n] | ψ,θ, (X i ,A i ) i∈[n] )p(θ | ψ) dθq(ψ), In (i), we use that (X,A) is independent of ψ, which follows from Assumption 6. Sim- ilarly, in (i), we use that θ is conditionally independent of (X,A) given ψ. Now we know that given θ, R i | X i ,A i are i.i.d. and hence p((R i ) i∈[n] | ψ,θ, (X i ,A i ) i∈[n] ) = ∏︁ a∈A L a (θ a ). Moreover, θ a for a ∈ A are conditionally independent given ψ. Thus p(θ | ψ) = ∏︁ a∈A p a (θ a ;f a (ψ)), where we also used that θ a | ψ ∼ p a (·;f a (ψ)). This leads to p(ψ | (X i ,A i ,R i ) i∈[n] )∝ ∫︂ θ ∏︂ a∈A L a (θ a )p a (θ a ;f a (ψ)) dθq(ψ), (i) = ∏︂ a∈A ∫︂ θ a L a (θ a )N (θ a ; W a ψ, Σ a ) dθ a N (ψ;μ, Σ), (i) = ∏︂ a∈A ∫︂ θ a (︂ ∏︂ i∈I a N (R i ;φ(X i ) ⊤ θ a ,σ 2 ) )︂ N (θ a ; W a ψ, Σ a ) dθ a N (ψ;μ, Σ). 163 In (i), we notice that θ = (θ a ) a∈A and apply Fubini’s Theorem. In (i), we let I a = i ∈ [n];A i = a as the rounds where action a appears in the sample set D n . Now let h a (ψ) = ∫︁ θ a (︁ ∏︁ i∈I a N (R i ;φ(X i ) ⊤ θ a ,σ 2 ) )︁ N (θ a ; W a ψ, Σ a ) dθ a . Then we have that p(ψ |D n )∝ ∏︂ a∈A h a (ψ)N (ψ;μ, Σ).(C.5) We start by computing h a . To reduce clutter, let v = σ −2 and Λ a = Σ −1 a . Then we compute h a as h a (ψ) = ∫︂ θ a (︄ ∏︂ i∈I a N (R i ;φ(X i ) ⊤ θ a ,σ 2 ) )︄ N (θ a ; W a ψ, Σ a ) dθ a , ∝ ∫︂ θ a exp [︄ − 1 2 v ∑︂ i∈I a (R i − φ(X i ) ⊤ θ a ) 2 − 1 2 (θ a − W a ψ) ⊤ Λ a (θ a − W a ψ) ]︄ dθ a , = ∫︂ θ a exp [︂ − 1 2 (︂ v ∑︂ i∈I a (R 2 i − 2R i θ ⊤ a φ(X i ) + (θ ⊤ a φ(X i )) 2 ) + θ ⊤ a Λ a θ a − 2θ ⊤ a Λ a W a ψ + (W a ψ) ⊤ Λ a (W a ψ) )︂]︂ dθ a , ∝ ∫︂ θ a exp [︂ − 1 2 (︂ θ ⊤ a (︄ v ∑︂ i∈I a φ(X i )φ(X i ) ⊤ + Λ a )︄ θ a − 2θ ⊤ a (︄ v ∑︂ i∈I a R i φ(X i ) + Λ a W a ψ )︄ + (W a ψ) ⊤ Λ a (W a ψ) )︂]︂ dθ a . Now recall that G a = v ∑︁ i∈I a φ(X i )φ(X i ) ⊤ and B a = v ∑︁ i∈I a R i φ(X i ) and let V a = (G a + Λ a ) −1 , U a = V −1 a , and β a = V a (B a + Λ a W a ψ). Then have that U a V a = V a U a = I d , and thus h a (ψ)∝ ∫︂ θ a exp [︃ − 1 2 (︁ θ ⊤ a U a θ a − 2θ ⊤ a U a V a (B a + Λ a W a ψ) + (W a ψ) ⊤ Λ a (W a ψ) )︁ ]︃ dθ a , = ∫︂ θ a exp [︃ − 1 2 (︁ θ ⊤ a U a θ a − 2θ ⊤ a U a β a + (W a ψ) ⊤ Λ a (W a ψ) )︁ ]︃ dθ a , = ∫︂ θ a exp [︃ − 1 2 (︁ (θ a − β a ) ⊤ U a (θ a − β a )− β ⊤ a U a β a + (W a ψ) ⊤ Λ a (W a ψ) )︁ ]︃ dθ a , ∝ exp [︃ − 1 2 (︁ −β ⊤ a U a β a + (W a ψ) ⊤ Λ a (W a ψ) )︁ ]︃ , = exp [︃ − 1 2 (︂ − (B a + Λ a W a ψ) ⊤ V a (B a + Λ a W a ψ) + (W a ψ) ⊤ Λ a (W a ψ) )︂ ]︃ , ∝ exp [︃ − 1 2 (︁ ψ ⊤ W ⊤ a (Λ a − Λ a V a Λ a ) W a ψ− 2ψ ⊤ (︁ W ⊤ a Λ a V a B a )︁)︁ ]︃ , ∝ exp [︃ − 1 2 (︁ ψ ⊤ ̄ Λ a ψ− 2ψ ⊤ ̄ Λ a ̄μ a )︁ ]︃ , where ̄ Λ a = W ⊤ a (Λ a − Λ a V a Λ a ) W a = W ⊤ a (︁ Σ −1 a − Σ −1 a (G a + Σ −1 a ) −1 Σ −1 a )︁ W a , ̄ Λ a ̄μ a = W ⊤ a Λ a V a B a = W ⊤ a Σ −1 a (G a + Σ −1 a ) −1 B a .(C.6) 164 However, we know from Equation (C.5) that p(ψ | D n ) ∝ ∏︁ a∈A h a (ψ)N (ψ;μ, Σ). But h a (ψ) is proportional to exp [︁ − 1 2 (︁ ψ ⊤ ̄ Λ a ψ− 2ψ ⊤ ̄ Λ a ̄μ a )︁]︁ for any a. Thus p(ψ |D n ) can be seen as the product of K + 1 Gaussian kernels. Thus, p(ψ |D n ) is a multivariate Gaussian distribution N ( ̄μ, ̄ Σ), with ̄ Σ −1 = Σ −1 + ∑︂ a∈A ̄ Λ a = Σ −1 + ∑︂ a∈A W ⊤ a (︁ Σ −1 a − Σ −1 a (G a + Σ −1 a ) −1 Σ −1 a )︁ W a ,(C.7) ̄ Σ −1 ̄μ = Σ −1 μ + ∑︂ a∈A ̄ Λ a ̄μ a = Σ −1 μ + ∑︂ a∈A W ⊤ a Σ −1 a (G a + Σ −1 a ) −1 B a .(C.8) C.2.2 Conditional Posterior Derivation of p(θ a | ψ,D n ). Let v = σ −2 ,Λ a = Σ −1 a . We consider the model rewritten in Equation (C.4), then the joint conditional action posterior p(θ | ψ,D n ) decomposes as p(θ | ψ,D n ) = p(θ | ψ, (X i ,A i ,R i ) i∈[n] ) (i) ∝ p((R i ) i∈[n] | θ,ψ, (X i ,A i ) i∈[n] )p(θ | ψ, (X i ,A i ) i∈[n] ), (i) = p((R i ) i∈[n] | θ, (X i ,A i ) i∈[n] )p(θ | ψ) (i) = ∏︂ i∈[n] p(R i | θ,X i ,A i )p(θ | ψ), (iv) = ∏︂ i∈[n] N (R i ; (u A i ⊗ φ(X i )) ⊤ θ,σ 2 )N (θ; W A ψ, Σ A ), = exp [︂ − 1 2 (︂ v n ∑︂ i=1 (R 2 i − 2R i (u A i ⊗ φ(X i )) ⊤ θ + ((u A i ⊗ φ(X i )) ⊤ θ) 2 ) + θ ⊤ Λ A θ− 2θ ⊤ Λ A W A ψ + ψ ⊤ W ⊤ A Λ A W A ψ )︂]︂ , ∝ exp [︂ − 1 2 (︂ θ ⊤ (v n ∑︂ i=1 (u A i ⊗ φ(X i ))(u A i ⊗ φ(X i )) ⊤ + Λ A )θ − 2θ ⊤ (︄ v n ∑︂ i=1 (u A i ⊗ φ(X i ))R i + Λ A W A ψ )︄ )︂]︂ , = exp [︂ − 1 2 (︂ θ ⊤ (v n ∑︂ i=1 (u A i u ⊤ A i ⊗ φ(X i )φ(X i ) ⊤ ) + Λ A )θ− 2θ ⊤ (︄ v n ∑︂ i=1 (u A i ⊗ φ(X i ))R i + Λ A W A ψ )︄ )︂]︂ , (v) ∝ N (︃ θ; ̃μ A , (︂ ̃ Λ A )︂ −1 )︃ , where we use Bayes rule in (i), (i) uses two assumptions. First, Given θ,X,A, R is inde- pendent of ψ. Second, given ψ, θ is independent of (X,A). Moreover, (i) follows from the assumption that R i | θ,X i ,A i are independent. Finally, in iv, we replace the distribution by their Gaussian form, and in (v), we set ̃ Λ A = v ∑︁ n i=1 (u A i u ⊤ A i ⊗ φ(X i )φ(X i ) ⊤ ) + Λ A , and ̃μ A = ̃ Λ −1 A (︁ v ∑︁ n i=1 (u A i ⊗ φ(X i ))R i + Λ A W A ψ )︂ , where Λ A = Σ −1 A = diag(Σ −1 a ) a∈A . 165 Now notice that ̃ Λ A = diag(Σ −1 a + G a ) a∈A . Thus, p(θ | ψ,D n ) = N (θ; ̃μ A , ̃ Λ −1 A ) where ̃μ A = ( ̃μ a ) a∈A and ̃ Λ A = diag( ̃ Λ a ) a∈A , with ̃ Λ a = Σ −1 a + G a , ̃ Λ a ̃μ a = Σ −1 a W a ψ + B a . The covariance matrix of p(θ | ψ,D n ) is diagonal by block. Thus θ a | ψ,D n for a ∈ A are independent and have a Gaussian density p(θ a | ψ,D n ) = N (θ a ; ̃μ a , ̃ Σ a ) where ̃ Σ a = ̃ Λ −1 a . C.2.3 Action Posterior Derivation of p(θ a |D n ). We know that θ a | D n ,ψ ∼ N ( ̃μ a , ̃ Σ a ) and ψ | D n ∼ N ( ̄μ, ̄ Σ). Thus the posterior density of θ a |D n is also Gaussian since Gaussianity is preserved after marginalization (Koller and Friedman, 2009). We let θ a |D n ∼N (ˆμ a , ˆ Σ a ). Then, we can compute ˆμ a and ˆ Σ a using the total expectation and total covariance decompositions. Let Λ a = Σ −1 a . Then we have that ̃ Σ a = (G a + Λ a ) −1 E [θ a |ψ,D n ] = ̃ Σ a (B a + Λ a W a ψ) First, given D n , ̃ Σ a = (G a + Λ a ) −1 and B a are constant (do not depend on ψ). Thus ˆμ a = E [θ a |D n ] = E [E [θ a |ψ,D n ]|D n ] = E ψ∼N ( ̄μ, ̄ Σ) [︂ ̃ Σ a (B a + Λ a W a ψ) ]︂ = ̃ Σ a (︁ B a + Λ a W a E ψ∼N ( ̄μ, ̄ Σ) [ψ] )︁ , = ̃ Σ a (B a + Λ a W a ̄μ). This concludes the computation of ˆμ a . Similarly, givenD n , ̃ Σ a = (G a + Λ a ) −1 and B a are constant (do not depend on ψ), yields two things. First, E [cov [θ a |ψ,D n ]|D n ] = E [︂ ̃ Σ a ⃓ ⃓ ⃓ D n ]︂ = ̃ Σ a . Second, cov [E [θ a |ψ,D n ]|D n ] = cov [︂ ̃ Σ a Λ a W a ψ ⃓ ⃓ ⃓ D n ]︂ = ̃ Σ a Λ a W a cov [ψ|D n ] W ⊤ a Λ a ̃ Σ a = ̃ Σ a Λ a W a ̄ ΣW ⊤ a Λ a ̃ Σ a . Finally, the total covariance decomposition (Weiss, 2005) yields that ˆ Σ a = cov [θ a |D n ] = E [cov [θ a |ψ,D n ]|D n ] + cov [E [θ a |ψ,D n ]|D n ] = ̃ Σ a + ̃ Σ a Λ a W a ̄ ΣW ⊤ a Λ a ̃ Σ a . This concludes the proof. 166 C.3 Proofs C.3.1 Main Result In this section, we prove Theorem 2. Recall that we make the following well-specified prior assumption. Assumption 7 (Well-specified priors). Action parameters θ ∗,a and rewards are drawn from Equation (6.9). Assumption 8 (Diagonal covariances for simplicity). We assume Σ a = σ 2 0 I d , Σ = τ 2 I d ′ , ∥φ(x)∥ 2 ≤ 1, and the matrices W a are normalized such that λ 1 (W a W ⊤ a ) = λ d (W a W ⊤ a ) = 1. Proof. First, given x ∈ X , by definition of the optimal policy, we know that it is deter- ministic. That is, there exists a x,θ ∗ ∈ [K] such that π ∗ (a x,θ ∗ | x) = 1. To simplify the notation and since π ∗ is deterministic, we let π ∗ (x) = a x,θ ∗ . Also, we know that the greedy policy is deterministic in ˆa x = argmax b∈A ˆr(x,b). That is ˆπ g (ˆa x | x) = 1. Similarly, we let ˆπ g (x) = ˆa x . Moreover, we let Φ(x,a) = e a ⊗ φ(x) ∈ R dK where e a ∈ R K is the indicator vector of action a, such that e a,b = 0 for any b ∈ A/a and e a,a = 1. Also, recall that ˆμ = (ˆμ a ) a∈A is the concatenation of the posterior means. Bso(ˆπ g ) = E [V (π ∗ ;θ ∗ )− V (ˆπ g ;θ ∗ )], = E [r(X,π ∗ (X);θ ∗ )− r(X, ˆπ g (X);θ ∗ )], = E [r(X,π ∗ (X);θ ∗ )− r(X, ˆπ g (X); ˆμ) + r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )], ≤ E [r(X,π ∗ (X);θ ∗ )− r(X,π ∗ (X); ˆμ) + r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )], ≤ E [r(X,π ∗ (X);θ ∗ )− r(X,π ∗ (X); ˆμ)] + E [r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )]. Now we start by proving that E [r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )] = 0. This is achieved as follows E [r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )] = E [E [r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )| X,D n ]], = E [︁ E [︁ φ(X) ⊤ ˆμ ˆπ g (X) − φ(X) ⊤ θ ∗,ˆπ g (X) | X,D n ]︁]︁ , (i) = E [︁ E [︁ Φ(X, ˆπ g (X)) ⊤ ˆμ− Φ(X, ˆπ g (X)) ⊤ θ ∗ | X,D n ]︁]︁ , (i) = E [︁ Φ(X, ˆπ g (X)) ⊤ E [ˆμ− θ ∗ | X,D n ] ]︁ , (i) = E [︁ Φ(X, ˆπ g (X)) ⊤ (ˆμ− E [θ ∗ | X,D n ]) ]︁ , (iv) = 0. In (i), we used that by definition of Φ(x,a) = e a ⊗ φ(X) ∈ R dK , we have φ(x) ⊤ ˆμ a = Φ(x,a) ⊤ ˆμ for any (x,a), and the same holds for θ ∗ . In (i), we used that Φ(X, ˆπ g (X)) is deterministic given X and D n . In (i), we used that ˆμ is deterministic given X and D n . Finally, in (iv), we used that E [θ ∗ | X,D n ] = E [θ ∗ |D n ] = ˆμ, which follows from the assumption that θ ∗ does not depend on X and the assumption that θ ∗ is drawn from the prior, and hence when conditioned on D n , it is drawn from the posterior whose mean is ˆμ. Therefore, E [r(X, ˆπ g (X); ˆμ)− r(X, ˆπ g (X);θ ∗ )] = 0 which leads to Bso(ˆπ g )≤ E [r(X,π ∗ (X);θ ∗ )− r(X,π ∗ (X); ˆμ)]. 167 Now let δ ∈ (0, 1) and define the joint parameter event E α = ︂ ∀a∈A : ∥θ ∗,a − ˆμ a ∥ ˆ Σ −1 a ≤ α ︂ . Recall Z a = r(X,a;θ ∗ )− r(X,a; ˆμ) = φ(X) ⊤ (θ ∗,a − ˆμ a ). By Cauchy-Schwarz, on E α we have for all a, |Z a | =|φ(X) ⊤ (θ ∗,a − ˆμ a )|≤∥φ(X)∥ ˆ Σ a ∥θ ∗,a − ˆμ a ∥ ˆ Σ −1 a ≤ α∥φ(X)∥ ˆ Σ a , hence in particular |Z π ∗ (X) |≤ α∥φ(X)∥ ˆ Σ π ∗ (X) on E α . Therefore, Bso(ˆπ g )≤ E [︁ |Z π ∗ (X) | ]︁ = E [︁ |Z π ∗ (X) |1E α ]︁ + E [︁ |Z π ∗ (X) |1 ︁ ̄ E α ︁]︁ ≤ αE [︂ ∥φ(X)∥ ˆ Σ π ∗ (X) ]︂ + E [︁ |Z π ∗ (X) |1 ︁ ̄ E α ︁]︁ . We now bound P( ̄ E α | D n ). Under the well-specified Bayes assumption, conditional on D n each marginal posterior satisfies θ ∗,a | D n ∼ N (ˆμ a , ˆ Σ a ). This means that θ ∗,a − ˆμ a | D n ∼ N (0, ˆ Σ a ). Thus, ˆ Σ − 1 2 a (θ ∗,a − ˆμ a ) ∼ N (0,I d ). But notice that ∥θ ∗,a − ˆμ a ∥ ˆ Σ −1 a = ∥ ˆ Σ − 1 2 a (θ ∗,a − ˆμ a )∥. Thus we apply Laurent and Massart (2000, Lemma 1) and get that P (︂ ∥θ ∗,a − ˆμ a ∥ ˆ Σ −1 a ≤ α ⃓ ⃓ ⃓ D n )︂ ≥ 1− δ K , where α = √︃ d + 2 √︂ d log K δ + 2 log K δ . This means that for every a, P(∥θ ∗,a − ˆμ a ∥ ˆ Σ −1 a > α|D n )≤ δ/K, and by a union bound, P( ̄ E α |D n )≤ δ.(C.9) and therefore P( ̄ E α ) = E[P( ̄ E α | D n )] ≤ δ. Finally, control the bad-event contribution by Cauchy-Schwarz: E [︁ |Z π ∗ (X) |1 ︁ ̄ E α ︁]︁ ≤ √︃ E [︂ Z 2 π ∗ (X) ]︂ √︂ P( ̄ E α )≤ √︃ E [︂ Z 2 π ∗ (X) ]︂ √ δ.(C.10) Putting the pieces together gives Bso(ˆπ g )≤ αE [︂ ∥φ(X)∥ ˆ Σ π ∗ (X) ]︂ + √ δ √︃ E [︂ Z 2 π ∗ (X) ]︂ ,(C.11) with α 2 = d + 2 √︁ d log(K/δ) + 2 log(K/δ). We now upper bound E [︂ Z 2 π ∗ (X) ]︂ in (C.11). We have that Z 2 π ∗ (X) ≤ max a∈A Z 2 a ,henceE [︁ Z 2 π ∗ (X) ]︁ ≤ E [︃ max a∈A Z 2 a ]︃ .(C.12) 168 Fix (X,D n ). Under the well-specified Bayes assumption, the conditional joint posterior (θ ∗,a ) a∈A |D n is Gaussian, and each marginal satisfies θ ∗,a |D n ∼N (ˆμ a , ˆ Σ a ) (with ˆ Σ a as derived in Section C.2.3). Therefore for each fixed a, Z a | X,D n ∼N (︁ 0,s 2 a )︁ , s 2 a =∥φ(X)∥ 2 ˆ Σ a . Let s 2 max = max a∈A s 2 a . Then for any t ≥ 0, by a union bound and the Gaussian tail bound, P (︃ max a∈A |Z a |≥ t ⃓ ⃓ ⃓ ⃓ X,D n )︃ ≤ ∑︂ a∈A P (|Z a |≥ t|X,D n )≤ 2K exp (︃ − t 2 2s 2 max )︃ . Using the identity E [W 2 ] = ∫︁ ∞ 0 P(W 2 ≥ u)du for W ≥ 0 and setting W = max a |Z a |, we obtain E [︃ max a∈A Z 2 a ⃓ ⃓ ⃓ ⃓ X,D n ]︃ = ∫︂ ∞ 0 P (︃ max a∈A |Z a |≥ √ u ⃓ ⃓ ⃓ ⃓ X,D n )︃ du ≤ ∫︂ ∞ 0 min ︃ 1, 2K exp (︃ − u 2s 2 max )︃ du = ∫︂ 2s 2 max log(2K) 0 1du + ∫︂ ∞ 2s 2 max log(2K) 2K exp (︃ − u 2s 2 max )︃ du = 2s 2 max log(2K) + 2s 2 max = (︁ 2 log(2K) + 2 )︁ s 2 max . Taking expectation over (X,D n ) yields E [︃ max a∈A Z 2 a ]︃ ≤ (︁ 2 log(2K) + 2 )︁ E [︃ max a∈A ∥φ(X)∥ 2 ˆ Σ a ]︃ .(C.13) Combining (C.12) and (C.13) gives E [︁ Z 2 π ∗ (X) ]︁ ≤ (︁ 2 log(2K) + 2 )︁ E [︃ max a∈A ∥φ(X)∥ 2 ˆ Σ a ]︃ .(C.14) Using the assumption that ∥φ(X)∥ 2 ≤ 1 yields that ∥φ(X)∥ 2 ˆ Σ a ≤ λ 1 ( ˆ Σ a ), hence E [︁ Z 2 π ∗ (X) ]︁ ≤ (︁ 2 log(2K) + 2 )︁ max a∈A λ 1 ( ˆ Σ a ).(C.15) Now recall that to simplify, we also assumed that Σ a = σ 2 0 I d for any a ∈ A and that Σ = τ 2 I d ′ . As a result, we have that: ˆ Σ a = ̃ Σ a + σ −4 0 ̃ Σ a W a ̄ ΣW ⊤ a ̃ Σ a , and we also have that λ 1 ( ̃ Σ a )≤ σ 2 0 and that λ 1 ( ˆ Σ a )≤ σ 2 0 + τ 2 ,∀a∈A,(C.16) Plugging (C.14) (or (C.15)) into (C.11) yields Bso(ˆπ g )≤ αE [︂ ∥φ(X)∥ ˆ Σ π ∗ (X) ]︂ + √︂ (2 log(2K) + 2)(σ 2 0 + τ 2 )δ,(C.17) 169 with α 2 = d + 2 √︁ d log(K/δ) + 2 log(K/δ) as defined above. Choosing δ = 1/n in (C.17) yields Bso(ˆπ g )≤ α n E [︂ ∥φ(X)∥ ˆ Σ π ∗ (X) ]︂ + √︃ (2 log(2K) + 2)(σ 2 0 + τ 2 ) n ,(C.18) where α 2 n = d + 2 √︁ d log(Kn) + 2 log(Kn).(C.19) This concludes the proof. C.3.2 Explicit Bound Additional simplification. To further simplify the exposition, we just set φ(x) = x for any x∈X. Assumption 9. Let G = E[X ⊤ ] with g = λ d (G). We assume that g > 0. Assumption 10 (Context-independent logging policy). A is independent of X, i.e., π 0 (a | x) = p a for all x and a. Equivalently, (X i ) are i.i.d. ∼ ν and independent of (A i ), with P(A = a) = p a . Theorem 8 (Explicit Bound). Let π ∗ (x) be the optimal action for context x. Then, the BSO of sDM under the structured prior Equation (6.9) satisfies Bso(ˆπ g )≤ α n √︄ E [︃ 1 λ X + τ 2 σ 4 0 λ 2 X + (σ 2 0 + τ 2 ) (︂ (d + 1)e −nρ 2 X /2 + d (︁ 2 e )︁ gm X /2 )︂ ]︃ + √︃ (2 log(2K) + 2)(σ 2 0 + τ 2 ) n , where α n = √︂ d + 2 √︁ d log(Kn) + 2 log(Kn), and ρ X = p π ∗ (X) , m X = ⌊︂ nρ X 2 ⌋︂ , λ X = σ −2 g m X 2 + σ −2 0 . Scaling with n. Recall m X =⌊nρ X /2⌋ and λ X = σ −2 g m X 2 +σ −2 0 , where ρ X = π 0 (π ∗ (X)). For n large enough (so that m X is roughly nρ X ), we have λ X = Θ(nρ X + 1) and the exponentially small tail terms can be ignored at the level of leading-order scaling. Con- sequently, up to absolute constants and polylogarithmic factors, Bso(ˆπ g ) = ̃ O (︄ α n √︄ E [︃ 1 nρ X + 1 ]︃ + √︃ logK n )︄ , and using α n = ̃ Θ( √ d) this can be summarized as Bso(ˆπ g ) = ̃ O (︄ √︄ E [︃ d nρ X + 1 ]︃ + √︃ logK n )︄ , 170 Proof. Let’s focus on the main term of Theorem 2, and we start by bounding E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) ]︂ ,(C.20) Now recall that to simplify, we also assumed that Σ a = σ 2 0 I d for any a ∈ A and that Σ = τ 2 I d ′ . As a result, we have that: ˆ Σ a = ̃ Σ a + σ −4 0 ̃ Σ a W a ̄ ΣW ⊤ a ̃ Σ a , and we also have that λ 1 ( ̃ Σ a )≤ σ 2 0 and that λ 1 ( ˆ Σ a )≤ σ 2 0 + τ 2 ,∀a∈A,(C.21) since λ 1 (W a W ⊤ a ) = 1. These are obtained using Weyl’s inequalities. Now let N a = n ∑︂ i=1 1A i = a, p a = P(A = a) = π 0 (a), ρ x = p π ∗ (x) , ρ X = p π ∗ (X) . Define m x = ⌊︂ nρ x 2 ⌋︂ , m X = ⌊︂ nρ X 2 ⌋︂ . Then, under Assumption 10, N a ∼ Bin(n,p a ) and it is independent of context X. There- fore, Hoeffding gives, for t = nρ X /2, P (︂ N π ∗ (X) < nρ X 2 ⃓ ⃓ ⃓ X,θ ∗ )︂ ≤ exp (︃ − nρ 2 X 2 )︃ .(C.22) Let Ω X,1 = ︁ N π ∗ (X) ≥ m X ︁ . Then we have that P (︁ ̄ Ω X,1 ⃓ ⃓ X,θ ∗ )︁ ≤ exp (︃ − nρ 2 X 2 )︃ .(C.23) Lemma 7 (Matrix Chernoff, Tropp (2012, Theorem 1.1)). Let (Y k ) m k=1 be independent PSD matrices with λ 1 (Y k )≤ R a.s. Let μ min = λ d ( ∑︁ m k=1 E[Y k ]). Then for δ ∈ [0, 1], P (︄ λ d (︂ m ∑︂ k=1 Y k )︂ ≤ (1− δ)μ min )︄ ≤ d [︃ e −δ (1− δ) 1−δ ]︃ μ min /R . Define Ω X,2 = ︄ λ d (︂ n ∑︂ i=1 1A i = π ∗ (X)X i X ⊤ i )︂ ≥ 1 2 N π ∗ (X) g ︄ ,Ω X = Ω X,1 ∩ Ω X,2 . We now bound P( ̄ Ω X,2 | X,θ ∗ ). Recall that Ω X,1 = ︁ N π ∗ (X) ≥ m X ︁ , m X = ⌊︂ nρ X 2 ⌋︂ . 171 Under Assumption 10, (A i ) is independent of (X i ), hence conditional on the index set S π ∗ (X) = i ∈ [n] : A i = π ∗ (X), the matrices X i X ⊤ i : i ∈ S π ∗ (X) are i.i.d. with the same law as X ⊤ . Moreover, since the Chernoff bound below depends on S π ∗ (X) only through|S π ∗ (X) | = N π ∗ (X) , the same bound holds when conditioning on N π ∗ (X) . Moreover, we have that φ(x) = x and that ∥φ(x)∥ ≤ 1. Thus, ∥x∥ ≤ 1 and hence λ 1 (X i X ⊤ i ) ≤ 1. Applying Lemma 7 with R = 1, E[X ⊤ ] = G, and μ min = λ d (︁ N π ∗ (X) G )︁ = N π ∗ (X) g, and choosing δ = 1 2 , we obtain the conditional bound P( ̄ Ω X,2 | X,N π ∗ (X) ,θ ∗ )≤ d (︂ √︁ 2/e )︂ gN π ∗ (X) . Taking conditional expectation N π ∗ (X) given X and θ ∗ , P( ̄ Ω X,2 | X,θ ∗ ) = E [︁ P( ̄ Ω X,2 | X,N π ∗ (X) ,θ ∗ )| X,θ ∗ ]︁ ≤ dE [︃ (︂ √︁ 2/e )︂ gN π ∗ (X) | X,θ ∗ ]︃ . Splitting on Ω X,1 yields E [︃ (︂ √︁ 2/e )︂ gN π ∗ (X) | X,θ ∗ ]︃ = E [︃ (︂ √︁ 2/e )︂ gN π ∗ (X) 1Ω X,1 | X,θ ∗ ]︃ + E [︃ (︂ √︁ 2/e )︂ gN π ∗ (X) 1 ̄ Ω X,1 | X,θ ∗ ]︃ ≤ (︂ √︁ 2/e )︂ gm X + P( ̄ Ω X,1 | X,θ ∗ ), since on Ω X,1 we have N π ∗ (X) ≥ m X and thus (︂ √︁ 2/e )︂ gN π ∗ (X) ≤ (︂ √︁ 2/e )︂ gm X , while on ̄ Ω X,1 we have (︂ √︁ 2/e )︂ gN π ∗ (X) ≤ 1. Therefore, P( ̄ Ω X,2 | X,θ ∗ )≤ d (︂ √︁ 2/e )︂ gm X + dP( ̄ Ω X,1 | X,θ ∗ ).(C.24) Combining (C.24) with (C.22) yields P( ̄ Ω X | X,θ ∗ )≤ P( ̄ Ω X,1 | X,θ ∗ ) + P( ̄ Ω X,2 | X,θ ∗ )≤ (d + 1) exp (︃ − nρ 2 X 2 )︃ + d (︂ √︁ 2/e )︂ gm X . (C.25) Finally, let I 1 = E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) 1Ω X | X,θ ∗ ]︂ , I 2 = E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) 1 ̄ Ω X | X,θ ∗ ]︂ . Then E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) ]︂ = E[I 1 ] + E[I 2 ]. Using ∥X∥ 2 ≤ 1 and (C.16), ∥X∥ 2 ˆ Σ π ∗ (X) = X ⊤ ˆ Σ π ∗ (X) X ≤ λ 1 ( ˆ Σ π ∗ (X) )∥X∥ 2 2 ≤ σ 2 0 + τ 2 . 172 Let c 1 = σ 2 0 + τ 2 . Then I 2 ≤ c 1 P( ̄ Ω X | X,θ ∗ )≤ c 1 (︃ (d + 1) exp (︃ − nρ 2 X 2 )︃ + d (︂ √︁ 2/e )︂ gm X )︃ .(C.26) Moreover, fix X and θ ∗ , on Ω X , λ d (︂ n ∑︂ i=1 1A i = π ∗ (X)X i X ⊤ i )︂ ≥ 1 2 N π ∗ (X) g ≥ 1 2 m X g, so λ d ( ˆ G π ∗ (X) ) = σ −2 λ d (︂ n ∑︂ i=1 1A i = π ∗ (X)X i X ⊤ i )︂ ≥ σ −2 1 2 m X g. Hence λ d ( ˆ G π ∗ (X) + σ −2 0 I d )≥ σ −2 1 2 m X g + σ −2 0 . Let λ X = σ −2 g m X 2 + σ −2 0 . Since ̃ Σ π ∗ (X) = ( ˆ G π ∗ (X) + σ −2 0 I d ) −1 , we obtain on Ω X : λ 1 ( ̃ Σ π ∗ (X) )≤ 1 λ X . Moreover, on Ω X we have that λ 1 ( ˆ Σ π ∗ (X) )≤ λ 1 ( ̃ Σ π ∗ (X) ) + σ −4 0 τ 2 λ 1 ( ̃ Σ π ∗ (X) ) 2 ≤ 1 λ X + σ −4 0 τ 2 λ 2 X . Therefore, since ∥X∥ 2 ≤ 1, I 1 ≤ (︃ 1 λ X + σ −4 0 τ 2 λ 2 X )︃ .(C.27) Combining (C.26) and (C.27), E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) ]︂ ≤ E [︃ 1 λ X + σ −4 0 τ 2 λ 2 X + (σ 2 0 + τ 2 ) (︃ (d + 1) exp (︃ − nρ 2 X 2 )︃ + d (︂ √︁ 2/e )︂ gm X )︃]︃ . But from Theorem 2, we know that Bso(ˆπ g )≤ α n E [︂ ∥X∥ ˆ Σ π ∗ (X) ]︂ + √︃ (2 log(2K) + 2)(σ 2 0 + τ 2 ) n ,(C.28) where α n = √︂ d + 2 √︁ d log(Kn) + 2 log(Kn). 173 Finally, by Jensen’s inequality, E [︂ ∥X∥ ˆ Σ π ∗ (X) ]︂ ≤ √︃ E [︂ ∥X∥ 2 ˆ Σ π ∗ (X) ]︂ . Substituting into (C.28) and combining with (C.26) and (C.27), we obtain Bso(ˆπ g )≤ α n √︄ E [︃ 1 λ X + τ 2 σ 4 0 λ 2 X + (σ 2 0 + τ 2 ) (︂ (d + 1)e −nρ 2 X /2 + d (︁ 2 e )︁ gm X /2 )︂ ]︃ + √︃ (2 log(2K) + 2)(σ 2 0 + τ 2 ) n , where α n = √︂ d + 2 √︁ d log(Kn) + 2 log(Kn). C.3.3 Optimality of Greedy Policies Here, we show that Greedy policy ˆπ g should be preferred to any other choice of policies when considering the BSO as our performance metric. This is because ˆπ g minimizes the BSO. To see this, note that by definition the Greedy policy ˆπ g is deterministic, that is for any context x ∈ X, there exists ˆa g , such that ˆπ g (ˆa g | x) = 1. Thus, for any context x∈X, we simplify the notation by letting ˆπ g (x) denote the action that has a mass equal to 1. Then, we have that E A∼ˆπ g (·|x) [E θ ∗ [r(x,A;θ ∗ )|D n ]] = E θ ∗ [r(x, ˆπ g (x);θ ∗ )|D n ],(C.29) ≥ E [r(x,a;θ ∗ )|D n ]∀x,a∈X ×A. (C.30) where this follows from the definition of ˆr(x,a) = E [r(x,a;θ)|D n ], the definition of ˆπ g and the fact that θ ∗ is sampled from the prior, which leads to E [r(x,a;θ)|D n ] = E [r(x,a;θ ∗ )|D n ]. Now Equation (C.29) holds for any x ∈ X and a ∈ A, and hence it holds in expectation under X ∼ ν and A∼ π(·| X) for any policy π. That is, E X∼ν,A∼ˆπ g (·|X) [E θ ∗ [r(x,A;θ ∗ )|D n ]]≥ E X∼ν,A∼π(·|X) [E θ ∗ [r(x,A;θ ∗ )|D n ]]. (C.31) Taking another expectation w.r.t. the sample set D n and using Fubini’s theorem and the tower rule leads to E [V (ˆπ g ;θ ∗ )] ≥ E [V (π;θ ∗ )] for any stationary policy π. Then, subtracting E [V (π ∗ ;θ ∗ )] from both sides of the previous inequality yields that the BSO is minimized by ˆπ g compared to any stationary policy π, in particular, compared to the policy π p induced by pessimism. C.4 Additional Experiments As mentioned in Section C.4, our experiments were conducted on internal machines with 30 CPUs and thus they required a moderate amount of computation. These experiments are also reproducible with minimal computational resources. C.4.1 Implementation Details of Baselines We implement the baselines as follows. 174 • IPS. argmax π 1 n n ∑︂ i=1 π(A i |X i ) maxπ 0 (A i |X i ),τ R i ,(C.32) where τ ∈ [0, 1] is a hyper-parameter. • snIPS. argmax π 1 ∑︁ n i=1 π(A i |X i ) π 0 (A i |X i ) n ∑︂ i=1 π(A i |X i ) π 0 (A i |X i ) R i ,(C.33) • MIPS. We cluster actions into L groups and let h(a) be the cluster of action a. Let C i be the cluster of action A i , then we use argmax π 1 n n ∑︂ i=1 π(C i |X i ) π 0 (C i |X i ) R i ,(C.34) where C i = h(A i ) for any i∈ [n], and π(c|x) = ∑︁ a∈A 1 [h(a) = c]π(a|x). • PC. We use the Knn implementation of PC. Let N (a,k) be the set of k-nearest neighbors of a, then argmax π 1 n n ∑︂ i=1 ∑︁ a∈A 1 [a∈ N (A i ,k)]π(a|x) ∑︁ a∈A 1 [a∈ N (A i ,k)]π 0 (a|x) R i .(C.35) • DM (Freq). This DM uses the linear-Gaussian likelihood model R | θ,X,A ∼ N (φ(X) ⊤ θ A ,σ 2 ) and learn the parameters θ a using the maximum likelihood princi- ple leading to ˆr(x,a) = φ(x) ⊤ ˆμ a ,(C.36) where the MLE is ˆμ a = (G a + λI d ) −1 B a , with G a = ∑︁ i∈[n] I A i =a φ(X i )φ(X i ) ⊤ and B a = ∑︁ i∈[n] I A i =a R i φ(X i ), and λ is a regularization hyper-parameter. • DM (Bayes). This DM uses the linear-Gaussian likelihood model combined with Gaussian priors as θ a ∼N (μ a , Σ a ),∀a∈A,(C.37) R| θ,X,A∼N (φ(X) ⊤ θ A ,σ 2 ), Under this prior, each action a has an associated parameter θ a . Given the prior in Equation (6.2), the posterior distribution of an action parameter follows a multivari- ate Gaussian: θ a |D n ∼N (ˆμ a , ˆ Σ a ), where ˆ Σ −1 a = Σ −1 a +G a and ˆ Σ −1 a ˆμ a = Σ −1 a μ a +B a . Here, G a = σ −2 ∑︁ i∈[n] I A i =a φ(X i )φ(X i ) ⊤ and B a = σ −2 ∑︁ i∈[n] I A i =a R i φ(X i ). Then, the reward estimate is ˆr(x,a) = φ(x) ⊤ ˆμ a ,(C.38) • DR. argmax π 1 n n ∑︂ i=1 π(A i |X i ) maxπ 0 (A i |X i ),τ (R i − ˆr(X i ,A i )) + E A∼π(·|X i ) [ˆr(X i ,A)], (C.39) with τ ∈ [0, 1] and ˆr is the reward model obtained using DM (Freq). 175 C.4.2 Robustness to Likelihood Misspecification We strengthened our evaluation by assessing sDM’s robustness to likelihood misspecifica- tion below (robustness to prior misspecification is provided in Section C.4.3). In these experiments, the true data-generating process (same as the synthetic experiments in Sec- tion 6.5) differed from sDM’s assumptions in two different ways: either the likelihood is misspecified Misspecified likelihood (Figure C.1). We also simulate when the true reward dis- tribution differed from the likelihood assumed by sDM. For example, we simulated bi- nary rewards using a Bernoulli-logistic model while sDM used a linear-Gaussian likelihood. Other DMs: DM (Bayes) and DM (Freq) also use a misspecified likelihood model and to emphasize this we add the suffix Lin to all DMs names. Overall, sDM still outperforms all methods by a large margin despite misspecification. 0500100015002000 Number of samples n 0.5 0.6 0.7 0.8 0.9 1.0 Avg. relative reward Miss. Likelihood - OPL K=100, d'=5, d=10 0500100015002000 Number of samples n Miss. Likelihood - OPL K=1000, d'=5, d=10 0500100015002000 Number of samples n Miss. Likelihood - OPL K=1000, d'=10, d=10 0500100015002000 Number of samples n Miss. Likelihood - OPL K=1000, d'=20, d=10 sDM (Ours)DM (Bayes)DM (Freq)DRIPSsnIPSMIPSPC Figure C.1: Effect of likelihood misspecification: The relative reward of the learned policy on synthetic problems using misspecified likelihood with varying n and K. C.4.3 Robustness to Prior Misspecification Misspecified prior means and covariances(Figure C.2). This is achieved by adding uniformly sampled noise from [v,v + 0.5] to both the true prior mean and covariance parameters μ, Σ, W a , Σ a , with v controlling the level of misspecification. We varied v ∈ 0.5, 1, 1.5 and analyzed its impact on sDM’s performance. For comparison, we included the well-specified sDM and the most competitive baseline, DM (Bayes), while omitting other baselines to reduce clutter. sDM’s performance decreases with increas- ing misspecification, yet sDM with misspecification still outperforms the most competitive baseline, especially when K is large. We also observe that the impact of prior covariance misspecification is less significant compared to prior mean misspecification. 176 0500100015002000 Number of samples n 0.5 0.6 0.7 0.8 0.9 1.0 Avg. relative reward Mis. Prior - OPL K=100, d'=5, d=10 0500100015002000 Number of samples n Mis. Prior - OPL K=1000, d'=5, d=10 0500100015002000 Number of samples n Mis. Prior - OPL K=1000, d'=10, d=10 0500100015002000 Number of samples n Mis. Prior - OPL K=1000, d'=20, d=10 sDMsDM (v=0.5)sDM (v=1)sDM (v=1.5)DM (Bayes) Figure C.2: Effect of prior mean and covariance misspecification: The average relative reward of the learned policy on synthetic problems using both misspecified prior means and covariances with varying n and K and d ′ . C.4.4 Comparison of Greedy and Pessimistic Policies To validate our theory that a greedy policy should be preferred over the commonly adopted pessimistic policy in our Bayesian setting, we used a performance metric averaged over multiple bandit problems sampled from the prior. To verify this, we considered the same OPL synthetic setting as in Section 6.5 and compared sDM with a greedy policy to sDM with a pessimistic policy. Recall that a greedy policy with respect to our reward estimate writes ˆπ g (a| x) = 1a = argmax b∈A ˆr(x,b),(C.40) while a pessimistic one writes ˆπ p (a| x) = 1a = argmax b∈A ˆr(x,b)− u(x,a),(C.41) where u(x,a) = α(d,δ)∥φ(X)∥ ˆ Σ a with α(d,δ) = √︃ d + 2 √︂ d log 1 δ + 2 log 1 δ . As predicted by our theory, the results show that the greedy policy has better average performance over multiple bandit instances sampled from the prior. 0500100015002000 Number of samples n 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 Avg. relative reward Synthetic problems K=100, d'=5, d=10 0500100015002000 Number of samples n Synthetic problems K=1000, d'=5, d=10 0500100015002000 Number of samples n Synthetic problems K=1000, d'=10, d=10 0500100015002000 Number of samples n Synthetic problems K=1000, d'=20, d=10 sDM (Greedy)sDM (Pessimism) Figure C.3: Comparison of sDM with greedy policy and sDM with pessimistic policy in OPL: The average MSE of an ε-greedy policy on synthetic problems with varying n, K, and d ′ . 177 Chapter D Supplementary Materials for Chapter 7 Contents C.1 Posterior Derivations Under Standard Priors . . . . . . . . . . . . . . 161 C.2 Posterior Derivations Under Structured Priors . . . . . . . . . . . . . 162 C.2.1 Latent Posterior . . . . . . . . . . . . . . . . . . . . . . . . . . 163 C.2.2 Conditional Posterior . . . . . . . . . . . . . . . . . . . . . . . 165 C.2.3 Action Posterior . . . . . . . . . . . . . . . . . . . . . . . . . . 166 C.3 Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 C.3.1 Main Result . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167 C.3.2 Explicit Bound . . . . . . . . . . . . . . . . . . . . . . . . . . 170 C.3.3 Optimality of Greedy Policies . . . . . . . . . . . . . . . . . . 174 C.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 174 C.4.1 Implementation Details of Baselines . . . . . . . . . . . . . . . 174 C.4.2 Robustness to Likelihood Misspecification . . . . . . . . . . . 176 C.4.3 Robustness to Prior Misspecification . . . . . . . . . . . . . . 176 C.4.4 Comparison of Greedy and Pessimistic Policies . . . . . . . . . 177 D.1 Proofs for Oracle Policies D.1.1 Oracle Policies for IPS-Based Objectives (IPS), cIPS and ES. Recall the definition of the (logging propensity) clipped IPS estimator with τ ∈ [0, 1]: ˆ V cips (π) = 1 n n ∑︂ i=1 π(A i |X i ) maxπ 0 (A i |X i ),τ R i . 178 Taking n→∞, one obtains: V cips (π) = E X∼ν,A∼π 0 (·|X) [︃ π(A|X) maxπ 0 (A|X),τ r(X,A) ]︃ = E X∼ν,A∼π(·|X) [︃ π 0 (A|X) maxπ 0 (A|X),τ r(X,A) ]︃ . As the objective is linear in the policy π, the optimal policy should put for any x ∈ X , all the mass on the action a that maximizes the weighted reward, giving: π cIPS ∗ (a|x) = 1 [︂ a = argmax a ′ ∈A π 0 (a ′ |x)r(x,a ′ ) maxπ 0 (a ′ |x),τ ]︂ . We recover the solution for IPS when we let τ → 0: π IPS ∗ (a|x) = 1 [︂ a = argmax a ′ ∈A r(x,a ′ )1 [π 0 (a ′ |x) > 0] ]︂ . We also recover the solution of ES just by replacing the clipping function by an exponential function of factor α, obtaining: π ES ∗ (a|x) = 1 [︂ a = argmax a ′ ∈A r(x,a ′ )π 0 (a ′ |x) 1−α ]︂ . Doubly Robust (DR). The doubly robust estimator converges to the following quantity: V dr (π) = E X∼ν,A∼π(·|X) [︃ (r(X,A)− ˆr(X,A)) π 0 (A|X) maxπ 0 (A|X),τ + ˆr(X,A) ]︃ . The objective is linear in π and is thus maximized by the following deterministic decision rule: π DR ∗ (a|x) = 1 [︃ a = argmax a ′ ∈A ˆr(x,a ′ ) + (r(x,a ′ )− ˆr(x,a ′ )) π 0 (a ′ |x) maxπ 0 (a ′ |x),τ ]︃ Marginalized IPS (MIPS) with clusters. We adopt the same approach to look for the maximizer of MIPS. We generalize the clustering function h to also account for context. We write down the estimator: ˆ V mips (π) = 1 n n ∑︂ i=1 ∑︁ a ′ 1 [h(a ′ ,X i ) = h(A i ,X i )]π(a ′ |X i ) ∑︁ a ′ 1 [h(a ′ ,X i ) = h(A i ,X i )]π 0 (a ′ |X i ) R i = 1 n n ∑︂ i=1 π(C i |X i ) π 0 (C i |X i ) R i , with which, we recover when n→∞: V mips (π) = E X∼ν,A∼π 0 (·|X) [︃ ∑︁ a ′ 1 [h(a ′ ,X) = h(A,X)]π(a ′ |X) ∑︁ a ′ 1 [h(a ′ ,X) = h(A,X)]π 0 (a ′ |X) r(X,A) ]︃ = E X∼ν [︄ ∑︂ a π 0 (a|X) ∑︁ a ′ 1 [h(a ′ ,X) = h(a,X)]π(a ′ |X) ∑︁ a ′ 1 [h(a ′ ,X) = h(a,X)]π 0 (a ′ |X) r(X,a) ]︄ = E X∼ν [︄ ∑︂ a ′ π(a ′ |X) ∑︂ a π 0 (a|X) 1 [h(a ′ ,X) = h(a,X)] ∑︁ a ′ 1 [h(a ′ ,X) = h(a,X)]π 0 (a ′ |X) r(X,a) ]︄ = E X∼ν [︄ ∑︂ a ′ π(a ′ |X)E A∼π 0 (·|X) [︃ 1 [h(a ′ ,X) = h(A,X)]r(X,A) E A ′ ∼π 0 (·|X) [1 [h(A ′ ,X) = h(A,X)]] ]︃ ]︄ . 179 The objective is linear in π, and depends on the action a ′ through its cluster h(a ′ ,·) alone. This means that multiple solutions are maximizers as long as the policy chooses the best cluster c. We thus write down the oracle policy for MIPS in the cluster level, giving: π MIPS ∗ (c|x) = 1 [︂ c = argmax c ′ ∈C ︂ E A∼π 0 (·|x) [︃ r(x,A)1[h(A,x) = c ′ ] E A ′ ∼π 0 (·|x) [1 [h(A ′ ,x) = h(A,x)]] ]︃ ︂]︂ = 1 [︂ c = argmax c ′ ∈C ︂ E A∼π 0 (·|x) [︃ r(x,A)1[h(A,x) = c ′ ] E A ′ ∼π 0 (·|x) [1 [h(A ′ ,x) = c ′ ]] ]︃ ︂]︂ = 1 [︂ c = argmax c ′ ∈C ︂ E A∼π 0 (·|x) [r(x,A)1[h(A,x) = c ′ ]] E A∼π 0 (·|x) [1 [h(A,x) = c ′ ]] ︂]︂ , which ends the proof. Conjunct Effect Modeling (OffCEM). This estimator can be seen as the natural, doubly robust extension of the MIPS estimator. Combining similar techniques to the ones employed for MIPS and DR yields π OffCEM ∗ (a|x) = 1 [︂ a = argmax a ′ ∈A ︂ ˆr(x,a ′ ) + E ̄ A∼π 0 (·|x) [︁ (r(x, ̄ A)− ˆr(x, ̄ A))1[h(x, ̄ A) = h(x,a ′ )] ]︁ π 0 (h(x,a ′ )|x) ︂]︂ . Two Stage Decomposition (POTEC). This is an optimization strategy for OffCEM. It restricts the policy to a cluster-informed form, π(a| x) = ∑︂ c∈C π rm (a| x,c)π cl (c| x), where π rm (a | x,c) = 1[a = argmax a ′ ∈c ˆr(x,a ′ )] is fixed, model-based policy that deter- ministically selects the best action within each cluster. Learning is then simplified to finding the optimal cluster-level policy π cl that maximizes the OffCEM objective: ˆ V potec (π cl ) = 1 n n ∑︂ i=1 (︄ π cl (C i | X i ) π 0 (C i | X i ) (R i − ˆr(X i ,A i )) + ∑︂ c∈C π cl (c| X i )ˆr ∗ c (X i ) )︄ , where ˆr ∗ c (x) = max a∈c ˆr(x,a) is the estimated reward of the best action in cluster c. This is exactly the Doubly Robust version of MIPS on the cluster level, the oracle policy on the cluster level can be followed in the same fashion: π cl ∗ (c| x) = 1 [︂ c = argmax c ′ ∈C ︂ E A∼π 0 (·|x) [(r(x,A)− ˆr(x,A))1[h(A,x) = c ′ ]] E A∼π 0 (·|x) [1 [h(A,x) = c ′ ]] + ˆr ∗ c ′ (x) ︂]︂ . The optimal policy for the POTEC optimization strategy unfolds as: π POTEC ∗ (a|x) = ∑︂ c∈C π rm (a| x,c)π cl ∗ (c| x). At first glance, it might be hard to see the connection between POTEC and OffCEM solutions, but they are equivalent. For ease of notation, let us denote by D ˆr,x (c): D ˆr,x (c) = E A∼π 0 (·|x) [(r(x,A)− ˆr(x,A))1[h(A,x) = c]] E A∼π 0 (·|x) [1 [h(A,x) = c]] . 180 and recall that the optimal policy of OffCEM finds the action a that maximizes: ̃ V (x,a) = ˆr(x,a) + D ˆr,x (h(a,x)). For any context x, the optimal action a ∗ of POTEC verifies: • a ∗ is in the optimal cluster: h(a ∗ ,x) = c ∗ (x) with c ∗ (x) = argmax c∈C D ˆr,x (c) + ˆr ∗ c (x). • a ∗ is optimal within that cluster: a = argmax a∈c ∗ (x) ˆr(x,a). This means that for all actions a with h(a,x)̸= c ∗ (x), we have: ̃ V (x,a) = D ˆr,x (h(a,x)) + ˆr(x,a) ≤ D ˆr,x (h(a,x)) + ˆr ∗ h(a,x) (x) ≤ D ˆr,x (c ∗ (x)) + ˆr ∗ c ∗ (x) (x) = D ˆr,x (h(x,a ∗ )) + ˆr(x,a ∗ ) = ̃ V (x,a ∗ ). In addition, for all actions a with h(a,x) = c ∗ (x), we have: ̃ V (x,a) = D ˆr,x (h(a,x)) + ˆr(x,a) = D ˆr,x (c ∗ (x)) + ˆr(x,a) ≤ D ˆr,x (c ∗ (x)) + ˆr ∗ c ∗ (x) (x) = ̃ V (x,a ∗ ). This means that the optimal action a ∗ for POTEC is the maximizer of ̃ V (x,a), which is exactly the solution of OffCEM. Policy Convolution (PC). This estimator uses a nearest neighbors function to aggre- gate the propensities of similar actions, making the hypothesis that similar actions will result in similar reward signal. The estimator writes: ˆ V pc (π) = 1 n n ∑︂ i=1 π(N ε (A i )| X i ) π 0 (N ε (A i )| X i ) R i , with π(N ε (a)| x) = ∑︂ a ′ ∈N ε (a) π(a ′ | x). This estimator is equivalent to the following when n→∞: V PC (π) = E X∼ν,A∼π 0 (·|X) [︃ ∑︁ a ′ π(a ′ |X)1 [a ′ ∈ N ε (A)] π 0 (N ε (A)|X) r(X,A) ]︃ = E X∼ν,A∼π(·|X) [︄ E ̄ A∼π 0 (·|X) [︄ r(x, ̄ A)1 [︁ A∈ N ε ( ̄ A) ]︁ π 0 (N ε ( ̄ A)|X) ]︄]︄ . The same argument of linearity applies here, giving us the corresponding oracle policy: π PC ∗ (a|x) = 1 [︂ a = argmax a ′ ∈A ︂ E ̄ A∼π 0 (·|x) [︃ r(x, ̄ A)1[a ′ ∈ N ε ( ̄ A)] π 0 (N ε ( ̄ A)|x) ]︃ ︂]︂ . 181 D.1.2 Oracle Policies for PWLL-Based Objectives Our objectives can be written in the same form, only choosing for each a different function g: ˆ U g (π) = 1 n n ∑︂ i=1 g(X i ,A i ,R i ) logπ(A i | X i ). Since we are looking at oracle policies, we consider the expectation U g (π) = E X∼ν,A∼π 0 (·|X),R∼p(·|X,A) [g(X,A,R) logπ(A| X)]. The maximization decomposes over contexts. Fix x and define the nonnegative weights w x (a) = E R∼p(·|x,a) [g(x,a,R)]≥ 0. For each x, we thus consider max π(·|x) ∑︂ a∈A π 0 (a| x)w x (a) logπ(a| x) s.t. ∑︂ a∈A π(a| x) = 1, ∀a∈A, π(a| x)≥ 0. Let v x (a) = π 0 (a | x)w x (a) ≥ 0. The Lagrangian (with equality multiplier λ ∈ R and inequality multipliers μ(a) a∈A , μ(a)≥ 0) is L(π,λ,μ) = ∑︂ a∈A v x (a) logπ(a| x) + λ (︂ ∑︂ a∈A π(a| x)− 1 )︂ + ∑︂ a∈A μ(a)π(a| x). By KKT conditions, at an optimum π g ∗ (·| x) we have for all a∈A: ∂L ∂π(a| x) = v x (a) π(a| x) + λ + μ(a) = 0,and μ(a)π(a| x) = 0. For any action with π g ∗ (a| x) > 0, we get that μ(a) = 0, and hence π g ∗ (a| x) =− v x (a) λ . Normalizing with ∑︁ a π g ∗ (a| x) = 1 gives λ =− ∑︁ a ′ v x (a ′ ) and therefore π g ∗ (a| x) = v x (a) ∑︁ a ′ ∈A v x (a ′ ) = π 0 (a| x) E R∼p(·|x,a) [g(x,a,R)] ∑︁ a ′ ∈A π 0 (a ′ | x) E R∼p(·|x,a ′ ) [g(x,a ′ ,R)] . This concludes the proof. D.2 Proofs for Optimization Properties In this section, we prove the propositions about the optimization landscape of IPS-based and PWLL learning approaches. We start by stating the following lemmas, that will be helpful to prove our propositions. 182 Lemma 8. (Mei et al., 2020b, Lemma 2) Consider the single context case. With a slight abuse of notation, we drop the dependence on x and write r(a) instead of r(x,a), π θ (a) instead of π θ (a|x), and ˆr(a) instead of ˆr(x,a). Let π θ be a softmax policy parameterized by θ. Then, for any ˆ r ∈ [0, 1] K , and any estimator ˆ V linear in π θ , the mapping θ ↦→ ˆ V (π θ ) =⟨ ˆ r,π θ ⟩ is 5/2-smooth. Lemma 9. All the action level estimators EST in (IPS, cIPS, DR, PC) can be written, for any policy π, in the form: ˆ V est (π) = 1 n n ∑︂ i=1 E A∼π(·|X i ) [ˆr est,i (A,X i )] ,(D.1) For the cluster level estimators/approaches EST-C in (MIPS, OffCEM, POTEC), we also have ˆ V est-c (π) = 1 n n ∑︂ i=1 E C∼π(·|X i ) [ˆr est-c,i (C,X i )] ,(D.2) meaning that all these estimators are linear in π. Proof. This is straightforward to prove. We begin by the action level estimators and take DR as a representative. For DR, we have the following: ˆr DR,i (a,X i ) = ˆr(a,X i ) + I[a = A i ] R i − ˆr(A i ,X i ) max(τ,π 0 (A i |X i )) verifies the equation. Solutions for cIPS and IPS can be recovered directly, and PC follows the same construction. For the cluster level approaches, we take POTEC as a representative, and we have: ˆr POTEC,i (c,X i ) = ˆr ⋆ c (X i ) + I[c = C i ] R i − ˆr(A i ,X i ) π 0 (C i |X i ) , The ˆr MIPS,i follows as a special case when ˆr = 0. Lemma 10. Consider the single-context case and assume a finite action set A. For any estimator EST in (IPS, cIPS, DR, OffCEM, MIPS, PC), there exists a problem instance (i.e., a choice of r and π 0 ; and when relevant, a choice of auxiliary objects such as ˆr, h, N ε ) such that, in the large-n limit, ˆr est (a) = 1[a = a K ] for some optimal action a K . Similarly, for cluster-based approaches (e.g., POTEC and MIPS), there exists an instance such that ˆr est-c (c) = 1[c = c |C| ] for some optimal cluster c |C| . Proof. We give explicit constructions for cIPS (action-level) and POTEC (cluster-level). The other estimators follow by the same idea: choose a setting where the estimator becomes linear in π with some deterministic coefficient, and pick r (and possibly ˆr, h, N ε ) so that the resulting linearized reward is one-hot. 183 Action-level: cIPS. Fix τ ∈ [0, 1). 1 Choose a logging policy π 0 with full support and such that max a∈A π 0 (a) ≥ τ,(D.3) Let a K ∈ arg max a∈A π 0 (a) maxπ 0 (a),τ . Under Equation (D.3), there exists at least one action with π 0 (a)≥ τ, for which the ratio equals 1, hence the maximizer satisfies π 0 (a K )≥ τ and therefore π 0 (a K ) maxπ 0 (a K ),τ = 1. Now define the reward function r(a) = 1[a = a K ] maxπ 0 (a),τ π 0 (a) . This satisfies r(a)∈ [0, 1] for all a because r(a) = 0 for a̸= a K , and r(a K ) = maxπ 0 (a K ),τ π 0 (a K ) = 1 (since π 0 (a K )≥ τ ). For cIPS, the large-n linearized reward is ˆr cips (a) = π 0 (a) maxπ 0 (a),τ r(a), hence ˆr cips (a) = π 0 (a) maxπ 0 (a),τ 1[a = a K ] maxπ 0 (a),τ π 0 (a) = 1[a = a K ], as desired. Cluster-level: POTEC. We work in the single-context case and consider a clustering map h :A→C. Choose h so that a K forms a singleton cluster: c |C| = h(a K ) =a K , h(a)̸= c |C| ∀a̸= a K . Let rewards be r(a K ) = 1 and r(a) = 0 for a ̸= a K . Pick any ε ∈ (0, 1/2] and define a reward model ˆr(a K ) = 1− ε,ˆr(a) = ε ∀a̸= a K . For POTEC, the induced (cluster-level) linearized reward takes the form ˆr POTEC (c) = max a∈c ˆr(a) + ∑︁ a∈c π 0 (a) (︁ r(a)− ˆr(a) )︁ π 0 (c) . 1 If τ = 1 and |A| > 1, the simplifying assumption π 0 (a) > 0 for all a is incompatible with having some π 0 (a)≥ τ. In practice τ ≪ 1. 184 For the singleton cluster c |C| =a K , we get ˆr POTEC (c |C| ) = (1− ε) + π 0 (a K ) (︁ 1− (1− ε) )︁ π 0 (a K ) = (1− ε) + ε = 1. For any other cluster c̸= c |C| , all its actions satisfy r(a) = 0 and ˆr(a) = ε, hence ˆr POTEC (c) = ε + ∑︁ a∈c π 0 (a)(−ε) π 0 (c) = ε− ε = 0. Therefore ˆr POTEC (c) = 1[c = c |C| ]. This concludes the constructions for cIPS and POTEC. The remaining estimators can be handled analogously by choosing π 0 (and when relevant, h or N ε ) so that the estimator’s linear coefficient on r(a) equals 1 at a chosen a K (and equals something finite elsewhere), and then defining r (and possibly ˆr) to make the resulting ˆr est one-hot. Now we restate Proposition 3 and proceed to its proof. Proposition 8 (Plateau for linear-in-π objectives under softmax). Consider the single- context case. Let ˆ V (π) be any objective linear in π and let π ⋆ ∈ arg max π ˆ V (π) denote a maximizer over the probability simplex on the effective action space A eff (of size K eff = |A eff |). Let π θ t t≥1 be the iterates of gradient ascent on θ ↦→ ˆ V (π θ ) with a linear softmax policy π θ (a) = exp(θ a )/ ∑︁ a ′ ∈A eff exp(θ a ′ ) and step sizes η t ∈ (0, 1]. Then there exists a problem instance such that gradient ascent cannot escape a suboptimal region before t 0 = C K eff =O(K eff ) iterations, in the sense that ∀t≤ t 0 : ˆ V (π ⋆ )− ˆ V (π θ t )≥ 0.9. Proof. The proof follows the same technique as (Mei et al., 2020a, Theorem 1). By Lemma 10, there exists an instance (single context) for which the linearized reward is one-hot: ˆr est (a) = 1[a = a K ] for some a K ∈A eff . Hence, for any policy π supported on A eff , ˆ V (π) = ∑︂ a∈A eff π(a)ˆr est (a) = π(a K ). The maximizer over the simplex is therefore π ⋆ = δ a K , and ˆ V (π ⋆ ) = 1, ˆ V (π ⋆ )− ˆ V (π θ ) = 1− π θ (a K ). (Notice that sup θ ˆ V (π θ ) = 1 as well, although the supremum is not attained by any finite θ when K eff ≥ 2). We now upper bound the gradient norm. For the softmax parametrization, ⃦ ⃦ ⃦ ∇ θ ˆ V (π θ ) ⃦ ⃦ ⃦ 2 ≤ √ 2π θ (a K ) (︁ 1− π θ (a K ) )︁ , where the bound follows by a direct computation (as in (Mei et al., 2020a)). 185 Define the update θ t+1 = θ t + η t ∇ θ ˆ V (π θ t ) and split iterations into t good =t≥ 1 : π θ t+1 (a K ) > π θ t (a K ), t bad =t≥ 1 : π θ t+1 (a K )≤ π θ t (a K ). For t∈ t bad , 1 π θ t (a K ) − 1 π θ t+1 (a K ) ≤ 0. For t ∈ t good , using Lemma 8 (the 5/2-smoothness of θ ↦→ ˆ V (π θ )) and η t ∈ (0, 1], we obtain π θ t+1 (a K )− π θ t (a K )≤ 9 2 π θ t (a K ) 2 , and therefore (since π θ t+1 (a K )≥ π θ t (a K ) > 0), 1 π θ t (a K ) − 1 π θ t+1 (a K ) = π θ t+1 (a K )− π θ t (a K ) π θ t+1 (a K )π θ t (a K ) ≤ 9 2 . Summing over s = 1,...,t− 1 yields 1 π θ 1 (a K ) − 1 π θ t (a K ) = t−1 ∑︂ s=1 (︃ 1 π θ s (a K ) − 1 π θ s+1 (a K ) )︃ ≤ 9 2 t. Assume a standard symmetric initialization so that π θ 1 (a K ) = 1/K eff . Pick any constant c≥ 11 and take K eff large enough so that π θ 1 (a K )≤ 1/c. If t≤ 2 9c K eff , then 1 π θ t (a K ) ≥ 1 π θ 1 (a K ) − 9 2 t≥ 1 π θ 1 (a K ) (︃ 1− 1 c )︃ ≥ c− 1≥ 10, hence π θ t (a K )≤ 1/10, and thus ˆ V (π ⋆ )− ˆ V (π θ t ) = 1− π θ t (a K )≥ 0.9. This proves the claim with t 0 = 2 9c K eff . Proposition 9. Even for a single context x, deterministic rewards, there is problem where IPS-based learning with a linear softmax policy π θ (a) ∝ exp(⟨θ,φ(x,a)⟩)I[a ∈ A eff ] can have a number of local maxima exponential in the number of effective actions K eff . Proof. Let EST an off-policy estimators considered in the paper with an action-level policy. By Lemma 9, we have: ˆ V est (π) = 1 n n ∑︂ i=1 E a∼π(·|x i ) [ˆr est,i (a,x i )] ,(D.4) In a single context setting, it becomes: ˆ V est (π θ ) = E a∼π θ (·) [︄ 1 n n ∑︂ i=1 ˆr est,i (a) ]︄ ,(D.5) =⟨ 1 n n ∑︂ i=1 ˆr est,i ,π θ ⟩.(D.6) 186 This also holds for estimators with policies in the cluster level, as we still have: ˆ V est-c (π θ ) = E c∼π θ (·) [︄ 1 n n ∑︂ i=1 ˆr est-c,i (c) ]︄ ,(D.7) =⟨ 1 n n ∑︂ i=1 ˆr est-c,i ,π θ ⟩.(D.8) These softmax policies are all defined on the effective action space A eff , be it a subset of the action space A or the discrete cluster space C. Using the linearity of the objective, we can directly apply Theorem 1 from Chen et al. (2019) and obtain our result. Finally, we also restate Proposition 5, and provide its proof. Proposition 10. For an ℓ 2 regularized (substituting λ 2 ||θ|| 2 , with λ > 0), linear softmax policy π θ , the PWLL objective ˆ U g (π θ ) defined as: ˆ U g (π) = 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i )) logπ(A i | X i ), is λ-strongly concave. Without regularization, the objective is concave. Proof. For any x and a∈A eff (x), we have: π θ (a|x) = exp(⟨θ,φ(x,a)⟩) ∑︁ a ′ ∈A eff (x) exp(⟨θ,φ(x,a ′ )⟩) , optimizing an ℓ 2 regularized linear softmax, giving: ˆ L g,λ (π) = ˆ U g (π)− λ 2 ||θ|| 2 , with λ > 0 and recall that g ≥ 0. For strong concavity, we need to show that the Hessian ∇ 2 θ ˆ U g (π θ ) is negative definite with eigenvalues bounded away from zero. The gradient with respect to θ is: ∇ θ ˆ U g (π θ ) = 1 n ∑︁ n i=1 g(R i ,π 0 (A i | X i ))∇ θ logπ θ (A i |X i )− λθ For the softmax policy: ∇ θ logπ θ (a|x) = φ(x,a)− ∑︂ a ′ π θ (a ′ |x)φ(x,a ′ ) = φ(x,a)− E A∼π θ (·|x) [φ(x,A)] Therefore: ∇ θ ˆ U g (π θ ) = 1 n ∑︁ n i=1 g(R i ,π 0 (A i | X i )) (︁ φ(X i ,A i )− E A∼π θ (·|X i ) [φ(X i ,A)] )︁ − λθ Taking the second derivative: ∇ 2 θ ˆ U g (π θ ) =− 1 n ∑︁ n i=1 g(R i ,π 0 (A i | X i ))∇ θ E A∼π θ (·|X i ) [φ(X i ,A)]− λI d , where I d is the d× d identity matrix. The gradient of the expectation is: ∇ θ E A∼π θ (·|x) [φ(x,A)] = ∑︂ a ∇ θ π θ (a|x)φ(x,a) 187 Using ∇ θ π θ (a|x) = π θ (a|x)(φ(x,a)− E A∼π θ (·|x) [φ(x,A)]): ∇ θ E A∼π θ (·|x) [φ(x,A)] = ∑︂ a π θ (a|x)(φ(x,a)− E A∼π θ (·|x) [φ(x,A)])φ(x,a) ⊤ This simplifies to: ∇ θ E A∼π θ (·|x) [φ(x,A)] = Cov A∼π θ (·|x) [φ(x,A)] where Cov A∼π θ (·|x) [φ(x,A)] = E A∼π θ (·|x) [φ(x,A)φ(x,A) ⊤ ]−E A∼π θ (·|x) [φ(x,A)]E A∼π θ (·|x) [φ(x,A)] ⊤ Therefore: ∇ 2 θ ˆ U g (π θ ) =− 1 n n ∑︂ i=1 g(R i ,π 0 (A i | X i ))Cov A∼π θ (·|X i ) [φ(X i ,A)]− λI d We can write this as: ∇ 2 θ ˆ U g (π θ ) =−H − λI d where H = 1 n ∑︁ n i=1 g(R i ,π 0 (A i | X i ))Cov A∼π θ (·|X i ) [φ(X i ,A)] is positive semi-definite. To see this explicitly, for any vector v ∈ R d : v ⊤ Cov A∼π θ (·|X i ) [φ(X i ,A)]v = Var A∼π θ (·|X i ) [v ⊤ φ(X i ,A)]≥ 0, with the positivity of g, this ensures H is positive semi-definite. Then we have: v ⊤ ∇ 2 θ ˆ U g (π θ )v =−v ⊤ Hv− λv ⊤ v =−v ⊤ Hv− λ∥v∥ 2 , meaning that when v ̸= 0, we get v ⊤ ∇ 2 θ ˆ U g (π θ )v ≤−λ∥v∥ 2 < 0. This shows the Hessian is negative definite with all eigenvalues bounded above by−λ < 0. Therefore, ℓ 2 regularized ˆ U g (π θ ) is λ-strongly concave. In addition, when λ = 0, the hessian is negative semi-definite, giving simple concavity. D.3 Stochastic Optimization Convergence Guarantees for PWLL We analyze the convergence rates of stochastic gradient methods on the PWLL objective. We formulate this as the minimization of the finite-sum loss f (θ) =− ˆ U g (π θ ): f (θ) = 1 n n ∑︂ i=1 f i (θ), where f i (θ) =−g i logπ θ (A i | X i ),(D.9) where g i = g(R i ,π 0 (A i |X i )). We adopt the linear softmax policy parametrization in Equa- tion (7.16) with s θ (x,a) = φ(x,a) ⊤ θ (lightweight parametrization in Equation (7.17)). We note that our analysis extends naturally to the heavyweight parametrization in Equa- tion (7.17). 188 D.3.1 Assumptions and Regularity To establish problem-dependent convergence bounds, we rely on the following structural assumptions regarding the feature space and the importance weights. Assumption 11 (Bounded features). For all context-action pairs (x,a) ∈ X ×A, the feature representations are bounded in Euclidean norm: ∥φ(x,a)∥ 2 ≤ H. Assumption 12 (Bounded weighting function). The weights g i = g(R i ,π 0 (A i |X i )) com- puted on the static dataset are strictly positive and bounded. That is, for all i∈1,...,n: 0 < g i ≤ G max . Assumptions 11 and 12 are sufficient to establish the smoothness and bounded variance of the objective f (θ). We formally derive these properties in the following proposition. Proposition 11 (Regularity and Variance Bounds). Under Assumptions 11 and 12, the objective f (θ) satisfies the following properties: 1. Global Smoothness: The objective is ̄ L-smooth with ̄ L = G max H 2 . 2. Bounded Single-Sample Variance: The variance of the stochastic gradient for a single sample is bounded by ̄σ 2 = 4G 2 max H 2 . 3. Bounded Mini-Batch Variance: For a mini-batch of size b, the variance is bounded by ̄σ 2 b = 4G 2 max H 2 b . Proof. 1. Smoothness: The Hessian of the objective is the weighted sum of the feature covariance matrices under the policy π θ : ∇ 2 f (θ) = 1 n n ∑︂ i=1 g i Cov A∼π θ (·|X i ) [φ(X i ,A)]. The spectral norm of a covariance matrix is bounded by the maximum squared norm of its random vectors. Thus, using Assumption 11 we get that∥∇ 2 f (θ)∥ op ≤ 1 n ∑︁ n i=1 g i H 2 ≤ G max H 2 . 2. Single-Sample Variance: We first bound the norm of the gradient for an arbitrary sample i. The gradient is ∇f i (θ) = −g i (φ(X i ,A i ) − E A∼π θ (·|X i ) [φ(X i ,A)]). Using the triangle inequality and Assumption 11: ∥∇f i (θ)∥ 2 ≤ g i (︁ ∥φ(X i ,A i )∥ 2 +∥E A∼π θ (·|X i ) [φ(X i ,A)]∥ 2 )︁ ≤ G max (H + H) = 2G max H. Let ξ = ∇f I (θ) be the stochastic gradient sampled uniformly from the dataset. The variance is bounded by the second moment: Var(ξ)≤ E[∥ξ∥ 2 ] = 1 n n ∑︂ i=1 ∥∇f i (θ)∥ 2 ≤ (2G max H) 2 = 4G 2 max H 2 . 189 3. Mini-Batch Variance: Let the mini-batch gradient be ̄g t = 1 b ∑︁ b j=1 ∇f i j (θ), where in- dices are sampled independently with replacement. Using the standard variance reduction property for independent variables: E[∥ ̄g t −∇f (θ)∥ 2 ] = 1 b E[∥∇f I (θ)−∇f (θ)∥ 2 ]≤ 4G 2 max H 2 b . Based on Proposition 11, we define the following global problem-dependent constants on which our convergence rates depend: • ̄ L = G max H 2 : Smoothness constant. • ̄σ 2 = 4G 2 max H 2 : Upper bound on the gradient variance for a single sample. • ̄σ 2 b = 4G 2 max H 2 b : Upper bound on the gradient variance for a mini-batch of size b. D.3.2 PWLL without ℓ 2 regularization We begin by analyzing the standard unregularized PWLL objective. Here, the objective f (θ) is convex but not necessarily strongly convex. This implies the loss landscape may contain multiple minimizers rather than a unique global minimum. Consequently, we characterize convergence in terms of ˆ U g (π θ opt )− ˆ U g (π ̄ θ T ) (instead of ∥θ t − θ opt n ∥). Here, θ opt ∈ arg max θ ˆ U g (π θ ) is an optimal parameter and ̄ θ T is the average of the SGA iterates. Proposition 12. Let θ opt ∈ arg max θ ˆ U g (π θ ) be an optimal parameter. If the learning rate satisfies 0 < η ≤ 1 4 ̄ L , then by (Garrigos and Gower, 2023, Theorem 6.9), the iterates of mini-batch SGA satisfy: E [︂ ˆ U g (π θ opt )− ˆ U g (π ̄ θ T ) ]︂ ≤ ∥θ 0 − θ opt ∥ 2 ηT + 8ηG 2 max H 2 b where ̄ θ T is the average of the iterates. Proposition 12 highlights the trade-off inherent to constant step-size SGA: a larger η accelerates the decay of the initial error (first term) but increases the asymptotic noise floor (second term). For a fixed horizon T, one can recover a convergence rate ofO(1/ √ T ) by setting η ∝ 1/ √ T, which balances both terms. D.3.3 PWLL with ℓ 2 regularization We now move to the ℓ 2 -regularized case where the PWLL objective is strongly concave (Proposition 5). Precisely, we consider the regularized objective ̃ U λ (θ) = ˆ U g (π θ )− λ 2 ∥θ∥ 2 . Strong convexity implies the existence of a unique global minimizer. This allows us to guarantee convergence of the parameters θ t themselves, which is a stronger condition than value convergence. 190 Proposition 13. Let θ opt n,λ = arg max θ ̃ U λ (θ) be the unique optimal parameter. If the learning rate satisfies 0 < η ≤ 1 2(G max H 2 +λ) , then by (Garrigos and Gower, 2023, Theorem 6.12): E [︁ ∥θ t − θ opt n,λ ∥ 2 ]︁ ≤ (1− ηλ) t ∥θ 0 − θ opt n,λ ∥ 2 + 8ηG 2 max H 2 λb The regularized case demonstrates a convergence rate that is significantly faster than the rate of the unregularized case. D.4 Additional Experiments D.4.1 Detailed Experimental Setting Experimental Setting. Table D.1: Statistics of Post Processed Datasets Dataset Num. of actions Num. of samples MovieLens60, 000132, 744 Twitch200, 000400, 000 GoodReads1, 000, 000400, 000 Our experimental setup is designed to study the behavior of the different policy learning paradigms in large action spaces. To this end, we use three large action spaces collab- orative filtering datasets: Movielens (Lam and Herlocker, 2016), Twitch (Rappaz et al., 2021) and GoodReads (Wan et al., 2019) that are preprocessed to obtain a user-item interaction matrix. We follow the exact procedure of Sakhi et al. (2023) to pre-process the datasets. The statistics of the obtained datasets are described in Table D.1. For each user, we keep half of its history as the context x, and use the other half of the history as the products with positive reward, which align the learned policies to recommend new and relevant items. We direct the interested readers to Sakhi et al. (2023) for a detailed description of the experimental setup. The large action space scenario restricts the policies used to the inner product parametriza- tion (Aouali et al., 2022a). This parametrization is essential to leverage Maximum Inner Product Search algorithms (Shrivastava and Li, 2014) for fast query response. In partic- ular, we adopt policies of the following form: π θ (a|x)∝ exp(⟨φ Γ (x),β a ⟩), with the learnable parameter θ = [Γ,β], φ Γ : X → R ℓ defines the context embedding function in R ℓ and β the actions embeddings of size K × ℓ. To define our policies, we start by extracting action embeddings β 0 using an SVD decomposition of the user-item matrix. These embeddings help us define the context embedding function φ Γ and our 191 logging policy π 0 . φ 0 is set to the average embeddings of the observed actions in the contexts and is fixed for the logging policy π 0 . Using the SVD action embeddings β 0 , we define our logging policy π 0 as: π 0 (a|x)∝ exp (︃ 1 t ⟨φ 0 (x),β 0,a ⟩ )︃ I [︁ a∈ top k 0 (x) ]︁ , with t the temperature of the logging policy, and k 0 define the support of the logging policy, concentrating on the top k 0 actions with: top k 0 (x) = argsort a 1 ,·,a k 0 ⟨φ 0 (x),β 0,a ⟩. If not explicitly stated, k 0 is set to 100 and the temperature at t = 1 in all experiments. This policy is used to collect the offline datasetD n =X i ,A i ,R i i∈[n] on which all trainings are conducted. For each i∈ [n] in the processed dataset, X i is the user history, A i is the action played by the logging policy π 0 (·|X i ) and R i = 1[A i ∈ H i ] the observed reward, which is if the action played is in the hidden items of user i. Trained Policies Parameterizations. We adopt two parameterizations of the trained policies. The first one is a heavyweight parametrization, and focuses on learning the embeddings of the actions β (be it A of size K or C of size |C|), meaning that θ in this case is β. For action-level policies, this gives β ∈ R K×ℓ and for any x: π β (a|x) = exp(⟨φ 0 (x),β a ⟩) ∑︁ a ′ ∈A eff (x) exp(⟨φ 0 (x),β a ′ ⟩) , with A eff (x)⊂A, which depends on the choice of the practitioner, for example A eff (x) = S 0 (x), the support of π 0 for context x when we optimize IPS objectives. For cluster-level policies, this gives a β ∈ R |C|×ℓ and for any x: π β (c|x) = exp(⟨φ 0 (x),β c ⟩) ∑︁ c ′ ∈C exp(⟨φ 0 (x),β c ′ ⟩) . This is used by default if nothing is explicitly stated. We have also define a lightweight parametrization, where only a small projection W ∈ R ℓ×ℓ is learned, giving in action level policies: π W (a|x) = exp(⟨φ 0 (x)W,β a,0 ⟩) ∑︁ a ′ ∈A eff (x) exp(⟨φ 0 (x)W,β a ′ ,0 ⟩) , using β 0 , the embeddings of π 0 . For cluster level policies, we first define ̄ β 0 ∈ R |C|×ℓ with ̄ β 0,c = 1 |c| ∑︁ a∈c β 0,a , and use it to define the cluster level policy: π W (c|x) = exp(⟨φ 0 (x)W, ̄ β c,0 ⟩) ∑︁ c ′ ∈C exp(⟨φ 0 (x)W, ̄ β c ′ ,0 ⟩) . Reward Model. The reward model used ˆr is learned using regularized linear regression the collected interaction data, with ˆr(x,a) =⟨φ(x),θ a ⟩. Clustering and ε used. We use the embeddings β 0 , combined with K-means clustering to find our clusters. The number of clusters is set to 2000 for all datasets and experiments. For PC, the ℓ 2 threshold ε is set to 0.1. 192 D.4.2 Additional results Benefits of objective-aware parametrization. Figure D.1 shows the effect of objective- aware policy parameterizations for two different objectives and three large action space datasets. 2 4 2 5 2 6 2 7 2 8 0.10 0.15 0.20 0.25 0.30 0.35 Reward - MovieLens (K=60K) LR Schedule: None 2 4 2 5 2 6 2 7 2 8 0.10 0.15 0.20 0.25 0.30 0.35 LR Schedule: Warmup Cosine 2 4 2 5 2 6 2 7 2 8 0.10 0.15 0.20 0.25 0.30 0.35 LR Schedule: One Cycle 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Reward - Twitch (K=200K) 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 2 4 2 5 2 6 2 7 2 8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Reward - GoodReads (K=1M) 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 2 4 2 5 2 6 2 7 2 8 Batch Size 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Effect of Objective-Aware Parametrization on Performance IPS (Objective-Aware Defined on Support)IPS (Whole Action Space)cLPI (Objective-Aware Defined on Support)cLPI (Whole Action Space) Figure D.1: The effect of objective-aware parametrization for IPS and cLPI on three large-scale datasets Average MSE. Figure D.2 shows the average MSE by dataset and method. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. 193 MovieLens (K=60K) Twitch (K=200K) GoodReads (K=1M) 0.000 0.005 0.010 0.015 0.020 Average MSE Average MSE by Dataset and Method IPS ES DR MIPS PC OffCEM Figure D.2: Average MSE by Dataset and Method. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. MSE progress during training. Figures D.3a to D.3c show the progress of the MSE over 10 epochs on all three datasets. Several methods are excluded from the figure, as their high MSE values would distort the scale and obscure the comparison. 12345678910 Epoch 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 MSE MSE Evolution - MovieLens (K=60K) IPS ES DR MIPS PC OffCEM (a) MovieLens 12345678910 Epoch 0.000 0.005 0.010 0.015 0.020 0.025 0.030 0.035 0.040 MSE MSE Evolution - Twitch (K=200K) IPS ES DR MIPS PC OffCEM (b) Twitch 12345678910 Epoch 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 MSE MSE Evolution - GoodReads (K=1M) IPS ES DR MIPS PC OffCEM (c) GoodReads Figure D.3: MSE progression over 10 epochs across datasets. D.4.3 Results Averaging Different Seeds In this experiment, we analyze the reward evolution of representative PWLL and IPS- based methods on the three considered datasets. We compare two distinct optimization configurations: (i) a standard off-the-shelf Adam optimizer, and (i) a carefully tuned setup using Adam with an optimized batch size and a one-cycle learning-rate scheduler. This comparison enables us to isolate the effect of optimization on stability and conver- gence. Each method is evaluated over 5 random seeds, and we report the mean reward along with a shaded standard deviation region to visualize sensitivity to optimization randomness. In Figure D.4, across all datasets and optimization settings, we observe that IPS-based methods (cIPS, IX, and even POTEC) not only reach inferior performance but also suffer 194 from considerably higher variance. Their uncertainty bands are significantly wider, in- dicating unstable optimization. In contrast, PWLL-based methods exhibit near-invisible variance bands, with standard deviations roughly an order of magnitude smaller on aver- age. 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Validation reward MovieLens 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Twitch −0.05 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Optimizer with Best Scheduler GoodReads 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Validation reward 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 12345678910 Epoch −0.05 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Off-the-shelf Optimizer IPSDRESPOTECcLPI Figure D.4: cLPI vs IPS-Based methods: Evolution of rewards averaged over 5 different seeds. cLPI is more stable to optimize and reaches better policies. Finally, in Figure D.5, we observe that adopting an Objective-Aware parametrization yields further performance and stability improvements. For example, cIPS with Objective- Aware parametrization surpasses cIPS while maintaining lower variability, and cLPI in its Objective-Aware form consistently achieves the best overall performance. These results demonstrate that the combination of PWLL objectives and clever parametrization leads to more robust and more effective learned policies, while being very simple to implement. 195 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Validation reward MovieLens 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Twitch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Optimizer with Best Scheduler GoodReads 12345678910 Epoch 0.05 0.10 0.15 0.20 0.25 0.30 Validation reward 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Off-the-shelf Optimizer IPS (Objective-Aware Defined on Support)IPS (Whole Action Space)cLPI (Objective-Aware Defined on Support)cLPI (Whole Action Space) Figure D.5: Objective Aware Parametrisation: Evolution of rewards averaged over 5 different seeds. Objective Aware Parametrization stabilizes and improves performance for PWLL and IPS methods. D.4.4 Ablation - Sensitivity to Reward Noise In the original evaluation setup (see Section D.4), the observed reward is deterministic; for user i, we have R i = 1[A i ∈ H i ], meaning that a positive reward is returned only when the selected action belongs to the user’s hidden set H i . In this section, we investigate robustness to reward noise by introducing stochasticity in the form: R i ∼ 1[A i ∈ H i ] (1− B(ε)) + B(ε)s, where B(ε) is a Bernoulli random variable with parameter ε, and s ∈ [0, 1] is a shift. This results in noisy rewards supported on [0, 1]. Note that any reward scaling can be normalized to this range via R/R max when R max > 1. We evaluate six configurations defined by noise parameters ε ∈ 0.1, 0.2, 0.3 and re- ward shifts s∈0, 0.5. All methods are trained using the best-performing optimization schedule (one-cycle) to isolate the effect of noise. Results are reported in Figure D.6. We observe that increasing the noise level when s = 0 consistently harms all methods, as expected from a more stochastic reward signal. In contrast, when s = 0.5, higher noise tends to increase the overall reward level, since the shift raises the baseline reward. Across all noise–shift conditions, PWLL-based objectives maintain a clear advantage over IPS-based methods. When s = 0, RegKL and cLPI perform similarly, confirming that both benefit from the logarithmic reparameterization. However, as both noise and shift increase, RegKL begins to outperform cLPI, suggesting that, with an appropriately cho- sen regularization weight β, RegKL remains highly competitive even under reward high stochasticity. Conclusion. PWLL methods demonstrate robustness to reward noise, leading to im- proved stability and performance compared to traditional IPS-based objectives, even in 196 challenging noise regimes. 0.00 0.05 0.10 0.15 0.20 0.25 Validation Reward Shift = 0.0 Movielens — Noise=0.1Movielens — Noise=0.2Movielens — Noise=0.3Twitch — Noise=0.1Twitch — Noise=0.2Twitch — Noise=0.3 12345678910 Epoch 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Validation Reward Shift = 0.5 12345678910 Epoch 12345678910 Epoch 12345678910 Epoch 12345678910 Epoch 12345678910 Epoch IPSDRESPOTECcLPIRegKL Figure D.6: Ablation - Sensitivity to reward noise D.4.5 Ablation Study on Hyperparameters and Log Transform In this section, we evaluate the impact of hyperparameter choices on cIPS, cLPI, RegKL, and RegKL-LIN (the non-logarithmic variant of RegKL). All methods are run using the best-performing optimization configuration (optimizer + learning rate scheduler), ensur- ing that differences are driven solely by hyperparameter values and by whether the policy transformation is linear or logarithmic. The results are shown in Figure D.7. • cIPS consistently fails to reach competitive performance across all values of τ, especially in large action spaces. Its PWLL counterpart, cLPI, dominates for every τ, converging faster and achieving superior results. • For the KL-based objectives, we restrict to β ≥ 0.1 in order to avoid numerical instability from the exponential term (exp(1/β) > 2· 10 5 for β < 0.1). The same trend is observed: the PWLL variant (RegKL) reliably outperforms its linear ana- logue (RegKL-LIN) across all β, exhibiting more stable training dynamics, faster convergence and better performance. PWLL dominates. Across both objective families, replacing linear weights with log- transformed policy weights (PWLL) consistently provides greater robustness to hy- perparameters, faster optimization, and higher final performance, even more in challenging large-action-space settings. 197 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Validation reward Movielens 12345678910 Epoch Twitch 12345678910 Epoch Goodreads Reward trajectories: cLPI (solid) vs cIPS (dashed) across ¿ cLPI (solid)cIPS (dashed)cLPI (solid)cIPS (dashed) ¿= 0:001¿= 0:01¿= 0:05¿= 0:1¿= 0:5¿= 1 (a) cIPS (linear) vs cLPI (PWLL) w.r.t τ 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Validation reward Movielens 12345678910 Epoch Twitch 12345678910 Epoch Goodreads Reward trajectories: RegKL (solid) vs RegKL-LIN (dashed) across ̄ RegKL (solid)RegKL-LIN (dashed)RegKL (solid)RegKL-LIN (dashed) ̄= 0:1 ̄= 0:2 ̄= 0:4 ̄= 0:8 ̄= 1 (b) RegKL-LIN (linear) vs RegKL (PWLL) w.r.t β Figure D.7: Ablation Study on hyper-parameters and Log Transform D.4.6 Ablation: PWLL in Smaller Action Spaces We have shown that PWLL provides a more benign optimization landscape and yields stronger policies than IPS-based objectives in large action spaces. Here, we examine whether these benefits also extend to smaller action spaces. We construct a reduced version of Movielens by subsampling the action space to K ∈ 100, 500, 1000, 5000 items. Figure D.8 reports performance across varying K for cIPS (linear) and its PWLL-enhanced counterpart cLPI (log). In small action space settings (K ≤ 500), cIPS convergences faster than cLPI, but cLPI identifies a better maxima by the end of the 10 epochs. For medium action spaces (K ≥ 500), cLPI consistently outperforms cIPS, converging faster and identifying a better maximum. These results indicate that the optimization advantages of PWLL can still be beneficial in medium sized action space settings. 12345678910 Epoch 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Validation reward K= 100 12345678910 Epoch 0.05 0.10 0.15 0.20 0.25 0.30 K= 500 12345678910 Epoch 0.05 0.10 0.15 0.20 0.25 0.30 K= 1000 12345678910 Epoch 0.05 0.10 0.15 0.20 0.25 0.30 K= 5000 MovieLens — Comparison IPS (OPE) vs cLPI (PWLL) Varying Action Space Size K IPScLPI Figure D.8: PWLL (cLPI) vs IPS-based (IPS) in smaller action spaces. D.4.7 Ablation - Sensitivity to the number of clusters In this study, we compare our simple PWLL objective cLPI against MIPS and POTEC, two more complex IPS-based methods specifically designed for large action spaces. These baselines rely on a clustering function to reduce variance, and POTEC additionally leverages a reward model ˆr. We examine how the number of clusters affects their optimization performance. Figure D.9 reports the results. 198 POTEC generally outperforms MIPS for all numbers of clusters. However, both methods exhibit optimization instability across settings. While POTEC can occasionally match the final performance of cLPI on Movielens for a carefully selected number of clusters (1000), it consistently falls short on Twitch regardless of the cluster configuration. Conclusion. These findings demonstrate that focusing on optimization properties pays off: despite its simplicity, the PWLL objective cLPI can consistently outperform intricate IPS-based approaches tailored to large action spaces, even with the best finetuning. 0123456789 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Validation Reward Movielens 0123456789 Epoch Twitch MIPSPOTECcLPIMIPSPOTECcLPI Nb Clusters = 10Nb Clusters = 100Nb Clusters = 1000Nb Clusters = 10000 Figure D.9: PWLL (cLPI) vs POTEC and MIPS, changing the number of clusters. D.4.8 Ablation - Different Logging Supports We conduct experiments to quantify how increasing or restricting the support of the log- ging policy affects policy learning, comparing PWLL and IPS-based methods. Figure D.10 compiles the results and show that PWLL is still better than IPS-based approaches for different logging support sizes. 199 0.00 0.05 0.10 0.15 0.20 0.25 0.30 k log = 10 Validation reward MovieLensTwitchGoodReads 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 k log = 100 Validation reward 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 k log = 1000 Validation reward 12345678910 Epoch 12345678910 Epoch IPScLPI Figure D.10: PWLL vs IPS: Different Logging Support sizes k log D.4.9 Pessimism Does not Solve Optimization Problems Pessimism in face of uncertainty Jin et al. (2021) is motivated through a pure statistical learning rationale and is used to provide better statistical guarantees and more controlled excess risk. In the context of OPL, pessimistic strategies are derived combining con- centration bounds with class complexity measures, be it VC dimension (Swaminathan and Joachims, 2015a) or PAC-Bayesian tools (London and Sandler, 2019; Aouali et al., 2023a; Sakhi et al., 2024). For example, in its PAC-Bayesian formulation, the pessimistic objectives are all written in the following form: arg max π θ ˆ V (π θ )− λ n ||θ− θ 0 || 2 2 , Adding an ℓ 2 regularization term that pulls the parameters θ towards the behavior policy parameters θ 0 (defining π 0 ) induces pessimism by encouraging the learned policy to stay close to π 0 in parameter space. However, the optimization landscape of this objective becomes concave only when the regularization weight λ is sufficiently large for the ℓ 2 term to dominate. In that regime, the objective is indeed easier to optimize, but becomes overly conservative, yielding policies that remain too close to π 0 and under-exploit po- tential improvements. Figure D.11 confirms this empirically: pessimistic approaches, whether based on Sample Variance Penalisation (SVP) (Swaminathan and Joachims, 2015a), PAC-Bayesian learning with clipped IPS (London and Sandler, 2019), Exponen- tial Smoothing (Aouali et al., 2023a), or Logarithmic Smoothing (Sakhi et al., 2024), fail to outperform cLPI for any value of λ. 200 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Validation reward MovieLens 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Twitch 12345678910 Epoch 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 GoodReads-1M IPS (PAC-Bayes) ( ̧= 0:001) IPS (PAC-Bayes) ( ̧= 0:01) IPS (PAC-Bayes) ( ̧= 0:1) IPS (PAC-Bayes) ( ̧= 1) IPS (PAC-Bayes) ( ̧= 10) LS (PAC-Bayes) ( ̧= 0:001) LS (PAC-Bayes) ( ̧= 0:01) LS (PAC-Bayes) ( ̧= 0:1) LS (PAC-Bayes) ( ̧= 1) LS (PAC-Bayes) ( ̧= 10) ES (PAC-Bayes) ( ̧= 0:001) ES (PAC-Bayes) ( ̧= 0:01) ES (PAC-Bayes) ( ̧= 0:1) ES (PAC-Bayes) ( ̧= 1) ES (PAC-Bayes) ( ̧= 10) IPS (SVP) ( ̧= 0:001) IPS (SVP) ( ̧= 0:01) IPS (SVP) ( ̧= 0:1) IPS (SVP) ( ̧= 1) IPS (SVP) ( ̧= 10) cLPI Figure D.11: cLPI outperforms the pessimistic approaches. Some methods do not appear in the plot because their curves overlap. 201 Chapter E Supplementary Materials for Chapter 8 Contents D.1 Proofs for Oracle Policies . . . . . . . . . . . . . . . . . . . . . . . . . 178 D.1.1 Oracle Policies for IPS-Based Objectives . . . . . . . . . . . . 178 D.1.2 Oracle Policies for PWLL-Based Objectives . . . . . . . . . . 182 D.2 Proofs for Optimization Properties . . . . . . . . . . . . . . . . . . . . 182 D.3 Stochastic Optimization Convergence Guarantees for PWLL . . . . . 188 D.3.1 Assumptions and Regularity . . . . . . . . . . . . . . . . . . . 189 D.3.2 PWLL without ℓ 2 regularization . . . . . . . . . . . . . . . . . 190 D.3.3 PWLL with ℓ 2 regularization . . . . . . . . . . . . . . . . . . 190 D.4 Additional Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . 191 D.4.1 Detailed Experimental Setting . . . . . . . . . . . . . . . . . . 191 D.4.2 Additional results . . . . . . . . . . . . . . . . . . . . . . . . . 193 D.4.3 Results Averaging Different Seeds . . . . . . . . . . . . . . . . 194 D.4.4 Ablation - Sensitivity to Reward Noise . . . . . . . . . . . . . 196 D.4.5 Ablation Study on Hyperparameters and Log Transform . . . 197 D.4.6 Ablation: PWLL in Smaller Action Spaces . . . . . . . . . . . 198 D.4.7 Ablation - Sensitivity to the number of clusters . . . . . . . . 198 D.4.8 Ablation - Different Logging Supports . . . . . . . . . . . . . 199 D.4.9 Pessimism Does not Solve Optimization Problems . . . . . . . 200 Notation Clarification: Value vs. Risk Formulation Important: Throughout the main chapter, we present our results using the value formu- lation, where the goal is to maximize the expected reward V (π) = E X∼ν,A∼π(·|X) [r(X,A)] 202 with rewards R ∈ [0, 1]. In contrast, the appendix uses the equivalent risk (or cost) for- mulation, where the goal is to minimize the expected cost L(π) = E X∼ν,A∼π(·|X) [c(X,A)] with costs C ∈ [−1, 0]. These two formulations are related by a simple sign change: L(π) =−V (π) and C =−R, where r(x,a) =−c(x,a) for any (x,a)∈X ×A. Consequently: • Maximizing the value V (π) is equivalent to minimizing the risk L(π). • Upper bounds on|V (π)− ˆ V (π)| translate directly to upper bounds on|L(π)− ˆ L(π)|. • All theoretical guarantees derived in the appendix using the risk formulation apply equivalently to the value formulation presented in the corresponding chapter. We adopt the risk formulation in the appendix as it aligns with the standard convention in statistical learning theory, where one typically minimizes a loss or risk function. The reader should keep this equivalence in mind when relating the appendix results to the main paper. E.1 Bias and Variance Trade-Off Notation reminder: This section uses the risk formulation with costs C ∈ [−1, 0], which corresponds to the value formulation with rewards R = −C ∈ [0, 1] in the main paper. The estimator ˆ L α n (π) here corresponds to − ˆ V α (π) in the main paper. In this section, we provide additional results on how α controls the bias and variance of ˆ L α n (·). E.1.1 Bias and Variance of IPS-α The following proposition states the bias-variance trade-off for ˆ L α n (·). Proposition 14 (Bias and variance of IPS-α). Let α∈ [0, 1], the following holds for any evaluation policy π ∈ Π that is absolutely continuous with respect to π 0 |B( ˆ L α n (π))|≤ E X∼ν,A∼π(·|X) [︁ 1− π 0 (A|X) 1−α ]︁ , V [︂ ˆ L α n (π) ]︂ ≤ 1 n E X∼ν,A∼π(·|X) [︁ π(A|X) π 0 (A|X) 2α−1 ]︁ . 203 Proof. We first bound the bias as B( ˆ L α n (π)) = E [︂ ˆ L α n (π) ]︂ −L(π), = 1 n n ∑︂ i=1 E X i ∼ν,A i ∼π 0 (·|X i ),C i ∼p(·|X i ,A i ) [︃ C i π(A i |X i ) π 0 (A i |X i ) α ]︃ −L(π), (i) = E (X,A,C)∼μ π 0 [︃ C π(A|X) π 0 (A|X) α ]︃ −L(π), = E X∼ν [︄ ∑︂ a∈A c(X,a) π(a|X) π 0 (a|X) α−1 ]︄ − E X∼ν [︄ ∑︂ a∈A c(X,a)π(a|X) ]︄ , = E X∼ν [︄ ∑︂ a∈A c(X,a)π(a|X)(π 0 (a|X) 1−α − 1) ]︄ , = E X∼ν,A∼π(·|X) [︁ c(X,A)(π 0 (A|X) 1−α − 1) ]︁ , where (i) follows from the i.i.d. assumption. Since π 0 (A|X) 1−α ≤ 1 for any x ∈ X and a∈A, we have that |B( ˆ L α n (π))|≤ E X∼ν,A∼π(·|X) [︁ |c(X,A)||π 0 (A|X) 1−α − 1| ]︁ , ≤ E X∼ν,A∼π(·|X) [︁ 1− π 0 (A|X) 1−α ]︁ . The variance is bounded as V [︂ ˆ L α n (π) ]︂ = 1 n 2 n ∑︂ i=1 V X i ∼ν,A i ∼π 0 (·|X i ),C i ∼p(·|X i ,A i ) [︂ C i π(A i |X i ) π 0 (A i |X i ) α ]︂ , = 1 n V (X,A,C)∼μ π 0 [︂ C π(A|X) π 0 (A|X) α ]︂ , ≤ 1 n E (X,A,C)∼μ π 0 [︂ C 2 π(A|X) 2 π 0 (A|X) 2α ]︂ , ≤ 1 n E X∼ν,A∼π 0 (·|X) [︂ π(A|X) 2 π 0 (A|X) 2α ]︂ , = 1 n E X∼ν [︂ ∑︂ a∈A π(a|X) 2 π 0 (a|X) 2α−1 ]︂ , = 1 n E X∼ν,A∼π(·|X) [︂ π(A|X) π 0 (A|X) 2α−1 ]︂ . E.2 Proofs for Off-Policy Learning In this section, we provide the complete proofs for our OPL results in Section 8.3. We start with proving Theorem 4 in Section E.2.1. We then state the extension of Theorem 4 204 along with its proof in Section E.2.2. After that, in Section E.2.3, we provide the proof for Proposition 6. Finally, in Section E.2.4, we discuss in detail and prove our claims regarding the number of samples needed so that the performance of the learned policy is close to that of the optimal policy. Notation: This section uses the risk formulation with L(π) = −V (π) and ˆ L α n (π) = − ˆ V α (π). All bounds translate directly to the value formulation in the main paper. Recall that we assume the costs to be deterministic for simplicity: C i = c(X i ,A i ). E.2.1 Proof of Theorem 4 In this section, we prove Theorem 4. Proof. First, we decompose the difference L(π Q )− ˆ L α n (π Q ) as L(π Q )− ˆ L α n (π Q ) =L(π Q )− 1 n n ∑︂ i=1 L(π Q |X i ) ⏞ ⏟⏞ I 1 + 1 n n ∑︂ i=1 L(π Q |X i )− 1 n n ∑︂ i=1 L α (π Q |X i ) ⏞ ⏟⏞ I 2 + 1 n n ∑︂ i=1 L α (π Q |X i )− ˆ L α n (π Q ) ⏞ ⏟⏞ I 3 , where L(π Q ) = E X∼ν ,A∼π Q (·|X) [c(X,A)] , L(π Q |X i ) = E A∼π Q (·|X i ) [c(X i ,A)] , L α (π Q |X i ) = E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) α c(X i ,A) ]︃ , ˆ L α n (π) = 1 n n ∑︂ i=1 π(A i |X i ) π 0 (A i |X i ) α C i . Our goal is to bound |L(π Q )− ˆ L α n (π Q )| and thus we need to bound |I 1 | +|I 2 | +|I 3 |. We start with |I 1 |, Alquier (2021, Theorem 3.3) yields that following inequality holds with probability at least 1− δ/2 for any distribution Q on H |I 1 |≤ √︄ D KL (Q∥P) + log 4 √ n δ 2n .(E.1) 205 Moreover, |I 2 | can be bounded by decomposing it as |I 2 | = ⃓ ⃓ ⃓ ⃓ ⃓ 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [c(X i ,A)]− 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π α 0 (A|X i ) c(X i ,A) ]︃ ⃓ ⃓ ⃓ ⃓ ⃓ = ⃓ ⃓ ⃓ ⃓ ⃓ 1 n n ∑︂ i=1 ∑︂ a∈A π Q (a|X i )c(X i ,a)− π 0 (a|X i ) π Q (a|X i ) π α 0 (a|X i ) c(X i ,a) ⃓ ⃓ ⃓ ⃓ ⃓ = ⃓ ⃓ ⃓ ⃓ ⃓ 1 n n ∑︂ i=1 ∑︂ a∈A (︂ π Q (a|X i )− π Q (a|X i ) π α−1 0 (a|X i ) )︂ c(X i ,a) ⃓ ⃓ ⃓ ⃓ ⃓ = ⃓ ⃓ ⃓ ⃓ ⃓ 1 n n ∑︂ i=1 ∑︂ a∈A (︂ 1− π 1−α 0 (a|X i ) )︂ π Q (a|X i )c(X i ,a) ⃓ ⃓ ⃓ ⃓ ⃓ , ≤ 1 n n ∑︂ i=1 ∑︂ a∈A ⃓ ⃓ 1− π 1−α 0 (a|X i ) ⃓ ⃓ π Q (a|X i )|c(X i ,a)| . But 1− π 1−α 0 (a|x)≥ 0 and |c(x,a)|≤ 1 for any a∈A and x∈X. Thus |I 2 |≤ 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ 1− π 1−α 0 (A|X i ) ]︁ .(E.2) Finally, we need to bound the main term |I 3 |. To achieve this, we borrow the following technical lemma from Haddouche and Guedj (2022). It is slightly different from the one in Haddouche and Guedj (2022); their result holds for any n≥ 1 while we state a simpler version where n is fixed in advance. Lemma 11. Let Z be an instance space and let S n = (z i ) i∈[n] be an n-sized dataset for some n ≥ 1. Let (F i ) i∈0∪[n] be a filtration adapted to S n . Also, let H be a hypothesis space and (f i (S i ,h)) i∈[n] be a martingale difference sequence for any h ∈ H, that is for any i ∈ [n], and h ∈ H, we have that E [f i (S i ,h)|F i−1 ] = 0. Moreover, for any h ∈ H, let M n (h) = ∑︁ n i=1 f i (S i ,h). Then for any fixed prior, P, on H, any λ > 0, the following holds with probability 1− δ over the sample S n , simultaneously for any Q, on H |E h∼Q [M n (h)]|≤ D KL (Q∥P) + log(2/δ) λ + λ 2 (E h∼Q [⟨M⟩ n (h) + [M ] n (h)]) , where ⟨M⟩ n (h) = ∑︁ n i=1 E [︁ f i (S i ,h) 2 |F i−1 ]︁ and [M ] n (h) = ∑︁ n i=1 f i (S i ,h) 2 . To apply Lemma 11, we need to construct an adequate martingale difference sequence (f i (S i ,h)) i∈[n] for h ∈ H that allows us to retrieve |I 3 |. To achieve this, we define S n = (A i ) i∈[n] as the set of n taken actions. Also, we let (F i ) i∈0∪[n] be a filtration adapted to S n . For h∈H, we define f i (S i ,h) as f i (S i ,h) = f i (A i ,h) = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ). We stress that f i (S i ,h) only depends on the last action in S i , A i , and the predictor h. For this reason, we denote it by f i (A i ,h). The function f i is indexed by i since it depends 206 on the fixed i-th context, X i . The context X i is fixed and thus randomness only comes from A i ∼ π 0 (·|X i ). It follows that the expectations are under A i ∼ π 0 (·|X i ). First, we have that E [f i (A i ,h)|F i−1 ] = 0 for any i∈ [n],h∈H. This follows from E [f i (A i ,h)|F i−1 ] = E A i ∼π 0 (·|X i ) [︂ f i (A i ,h) ⃓ ⃓ ⃓ A 1 ,...,A i−1 ]︂ , = E A i ∼π 0 (·|X i ) [︃ E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ⃓ ⃓ ⃓ A 1 ,...,A i−1 ]︃ , (i) = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − E A i ∼π 0 (·|X i ) [︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ⃓ ⃓ ⃓ A 1 ,...,A i−1 ]︃ . In (i) we use the fact that given X i , E A∼π 0 (·|X i ) [︂ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︂ is deterministic. Now A i does not depend on A 1 ,...,A i−1 since logged data is i.i.d. Hence E A i ∼π 0 (·|X i ) [︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ⃓ ⃓ ⃓ A 1 ,...,A i−1 ]︃ = E A i ∼π 0 (·|X i ) [︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ]︃ , = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ . It follows that E [f i (A i ,h)|F i−1 ] = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − E A i ∼π 0 (·|X i ) [︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ⃓ ⃓ ⃓ A 1 ,...,A i−1 ]︃ , = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ , = 0. Therefore, for any h ∈ H, (f i (A i ,h)) i∈[n] is a martingale difference sequence. Hence we apply Lemma 11 and obtain that the following inequality holds with probability at least 1− δ/2 for any Q on H |E h∼Q [M n (h)]|≤ D KL (Q∥P) + log(4/δ) λ + λ 2 (E h∼Q [⟨M⟩ n (h) + [M ] n (h)]) ,(E.3) where M n (h) = n ∑︂ i=1 f i (A i ,h) , ⟨M⟩ n (h) = n ∑︂ i=1 E [︁ f i (A i ,h) 2 |F i−1 ]︁ , [M ] n (h) = n ∑︂ i=1 f i (A i ,h) 2 207 Now these terms can be decomposed as E h∼Q [M n (h)] = n ∑︂ i=1 E h∼Q [f i (A i ,h)] , = n ∑︂ i=1 E h∼Q [︃ E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ]︃ , (i) = n ∑︂ i=1 E h∼Q [︃ E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃]︃ − E h∼Q [︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ]︃ , (i) = n ∑︂ i=1 E A∼π 0 (·|X i ) [︄ E h∼Q [︁ I h(X i )=A ]︁ π 0 (A|X i ) α c(X i ,A) ]︄ − E h∼Q [︁ I h(X i )=A i ]︁ π 0 (A i |X i ) α c(X i ,A i ), (i) = n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) α c(X i ,A) ]︃ − n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) α c(X i ,A i ), where we use the linearity of the expectation in both (i) and (i). In (i), we use our definition of policies in (8.10). Therefore, we have that E h∼Q [M n (h)] = n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) α c(X i ,A) ]︃ − n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) α c(X i ,A i ), (i) = n ∑︂ i=1 L α (π Q |X i )− n ˆ L α n (π Q ), = nI 3 ,(E.4) where we used the fact that C i = c(X i ,A i ) for any i∈ [n] in (i). Now we focus on the terms ⟨M⟩ n (h) and [M ] n (h). First, we have that f i (A i ,h) 2 = (︂ E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ − I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) )︂ 2 ,(E.5) = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ 2 + (︂ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) )︂ 2 − 2E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ), = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ 2 + I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 − 2E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ). Moreover, f i (A i ,h) 2 does not depend on A 1 ,...,A i−1 . Thus, E [︁ f i (A i ,h) 2 |F i−1 ]︁ = E A i ∼π 0 (·|X i ) [︁ f i (A i ,h) 2 |F i−1 ]︁ , = E A i ∼π 0 (·|X i ) [︁ f i (A i ,h) 2 ]︁ = E A∼π 0 (·|X i ) [︁ f i (A,h) 2 ]︁ . 208 Computing E A∼π 0 (·|X i ) [︁ f i (A,h) 2 ]︁ using the decomposition in (E.5) yields E [︁ f i (A i ,h) 2 |F i−1 ]︁ = E A∼π 0 (·|X i ) [︁ f i (A,h) 2 ]︁ , =−E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ 2 + E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ (E.6) Combining (E.5) and (E.6) leads to E [︁ f i (A i ,h) 2 |F i−1 ]︁ + f i (A i ,h) 2 = E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 − 2E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ), (i) ≤ E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 . (E.7) The inequality in (i) holds because −2E A∼π 0 (·|X i ) [︂ I h(X i )=A π 0 (A|X i ) α c(X i ,A) ]︂ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i )≤ 0. Therefore, we have that ⟨M⟩ n (h) + [M ] n (h)≤ n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ I h(X i )=A π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 . Finally, by using the linearity of the expectation and the definition of policies in (8.10), we get that E h∼Q [⟨M⟩ n (h) + [M ] n (h)](E.8) ≤ n ∑︂ i=1 E A∼π 0 (·|X i ) [︄ E h∼Q [︁ I h(X i )=A ]︁ π 0 (A|X i ) 2α c(X i ,A) 2 ]︄ + E h∼Q [︁ I h(X i )=A i ]︁ π 0 (A i |X i ) 2α c(X i ,A i ) 2 , = n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + π Q (A i |X i ) π 0 (A i |X i ) 2α c(X i ,A i ) 2 .(E.9) Combining (E.3) and (E.8) yields n|I 3 | =| n ∑︂ i=1 L α (π Q |X i )− n ˆ L α n (π Q )| ≤ D KL (Q∥P) + log(4/δ) λ + λ 2 n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + π Q (A i |X i ) π 0 (A i |X i ) 2α c(X i ,A i ) 2 . (E.10) This means that the following inequality holds with probability at least 1− δ/2 for any distribution Q on H |I 3 |≤ D KL (Q∥P) + log(4/δ) nλ + λ 2n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α c(X i ,A) 2 ]︃ + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α c(X i ,A i ) 2 .(E.11) 209 However we know that c(x,a) 2 ≤ 1 for any x∈X and a∈A and that c(X i ,A i ) = C i for any i∈ [n]. Thus the following inequality holds with probability at least 1− δ/2 for any distribution Q on H |I 3 |≤ D KL (Q∥P) + log(4/δ) nλ + λ 2n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α C 2 i . (E.12) The union bound of (E.1) and (E.12) combined with the deterministic result in (E.2) yields that the following inequality holds with probability at least 1− δ for any distribution Q on H |L(π Q )− ˆ L α n (π Q )|≤ √︄ D KL (Q∥P) + log 4 √ n δ 2n + 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ 1− π 1−α 0 (A|X i ) ]︁ + D KL (Q∥P) + log(4/δ) nλ + λ 2n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α C 2 i . (E.13) E.2.2 Extensions of Theorem 4 Proposition 15 (Extension of Theorem 4 to hold simultaneously for any λ∈ (0, 1)). Let n≥ 1, δ ∈ [0, 1], α∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1− δ over draws D n ∼ μ n π 0 , the following holds simultaneously for any posterior Q on H, and for any λ∈ (0, 1) that |L(π Q )− ˆ L α n (π Q )|≤ √︃ kl ′ 1 (π Q ,λ) 2n + B α n (π Q ) + kl ′ 2 (π Q ,λ) nλ + λ 2 Var α n (π Q ). where kl ′ 1 (π Q ,λ) = D KL (Q∥P) + log 8 √ n δλ , kl ′ 2 (π Q ,λ) = 2 (︁ D KL (Q∥P) + log 8 δλ )︁ , B α n (π Q ) = 1− 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ π 1−α 0 (A|X i ) ]︁ , Var α n (π Q ) = 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + π Q (A i |X i ) π 0 (A i |X i ) 2α C 2 i . Proof. Let δ ∈ (0, 1). For any i≥ 1, we define λ i = 2 −i and let δ i = δλ i . Then Theorem 4 yields that for any i≥ 1, the following inequality holds with probability at least 1−δ i for any Q on H |L(π Q )− ˆ L α n (π Q )|≤ √︄ D KL (Q∥P) + log 4 √ n δ i 2n + B α n (π Q ) + D KL (Q∥P) + log 4 δ i nλ i + λ i 2 Var α n (π Q ). 210 Now notice that ∑︁ ∞ i=1 λ i = 1, and hence ∑︁ ∞ i=1 δ i = δ. Therefore, the union bound of the above inequalities over i≥ 1 yields that with probability at least 1− δ, the following inequality holds with probability at least 1− δ for any Q on H and for any i≥ 1 |L(π Q )− ˆ L α n (π Q )|≤ √︄ D KL (Q∥P) + log 4 √ n δ i 2n + B α n (π Q ) + D KL (Q∥P) + log 4 δ i nλ i + λ i 2 Var α n (π Q ). (E.14) Let ⌈·⌉ denote the ceiling function, then we have that for any λ ∈ (0, 1), there exists j = ⌈ − logλ log 2 ⌉ ≥ 1 such that λ/2 ≤ λ j ≤ λ. Since (E.14) holds for any i ≥ 1, it holds in particular for j. In addition to this, we have that 1 λ j ≤ 2 λ , that λ j ≤ λ and that 1 δ j = 1 λ j δ ≤ 2 δλ . This yields that the following inequality holds with probability at least 1− δ for any Q on H and for any λ∈ (0, 1) |L(π Q )− ˆ L α n (π Q )|≤ √︄ D KL (Q∥P) + log 8 √ n δλ 2n + B α n (π Q ) + 2 D KL (Q∥P) + log 8 δλ nλ + λ 2 Var α n (π Q ). (E.15) The additional 2 in 2 D KL (Q∥P)+log 8 δλ nλ appears since we used that 1 λ j ≤ 2 λ . Similarly, the additional 2 λ in the logarithmic terms is due to the fact that 1 δ j ≤ 2 δλ . Finally, setting kl ′ 1 (π Q ,λ) = D KL (Q∥P) + log 8 √ n δλ , kl ′ 2 (π Q ,λ) = 2 (︁ D KL (Q∥P) + log 8 δλ )︁ , concludes the proof. Next, we provide a similar proof to extend Theorem 4 to any α ∈ (0, 1]. While we only provide a one-sided inequality, the same covering technique can be used to obtain the other side of the inequality. Proposition 16 (One-sided extension of Theorem 4 to hold simultaneously for any α ∈ (0, 1)∪1 ). Let n ≥ 1, δ ∈ [0, 1], λ > 0, and let P be a fixed prior on H, then with probability at least 1−δ over draws D n ∼ μ n π 0 , the following holds simultaneously for any posterior Q on H, and for any α∈ (0, 1] that L(π Q )≤ ˆ L α n (π Q ) + √︃ kl ′ 1 (π Q ,α) 2n + B α n (π Q ) + kl ′ 2 (π Q ,α) nλ + λ 2 Var 2α n (π Q ). 211 where kl ′ 1 (π Q ,α) = D KL (Q∥P) + log 8 √ n δα , kl ′ 2 (π Q ,α) = D KL (Q∥P) + log 8 δα , B α n (π Q ) = 1− 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ π 1−α 0 (A|X i ) ]︁ , Var α n (π Q ) = 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + π Q (A i |X i ) π 0 (A i |X i ) 2α C 2 i . Proof. Let δ ∈ (0, 1). For any i ≥ 0, we define α i = 2 −i and let δ i = δα i /2. Then Theorem 4 yields that for any i ≥ 0, the following inequality holds with probability at least 1− δ i for any Q on H |L(π Q )− ˆ L α i n (π Q )|≤ √︄ D KL (Q∥P) + log 4 √ n δ i 2n + B α i n (π Q ) + D KL (Q∥P) + log 4 δ i nλ + λ 2 Var α i n (π Q ). Now notice that ∑︁ ∞ i=0 α i = 2, and hence by definition of δ i , we have ∑︁ ∞ i=0 δ i = δ. There- fore, the union bound of the above inequalities over i≥ 0 yields that with probability at least 1− δ, the following inequality holds with probability at least 1− δ for any Q on H and for any i≥ 0 |L(π Q )− ˆ L α i n (π Q )|≤ √︄ D KL (Q∥P) + log 4 √ n δ i 2n + B α i n (π Q ) + D KL (Q∥P) + log 4 δ i nλ + λ 2 Var α i n (π Q ). (E.16) Let ⌊·⌋ denote the floor function, then we have that for any α ∈ (0, 1], there exists j = ⌊ − logα log 2 ⌋≥ 0 such that α≤ α j ≤ 2α. Since (E.16) holds for any i≥ 0, it holds in particular for j. In addition, we have that B α n (π Q ) and ˆ L α n (π Q ) are decreasing in α while Var α n (π Q ) is increasing in α. Therefore, we have that ˆ L α j n (π Q ) ≤ ˆ L α n (π Q ), B α j n (π Q ) ≤ B α n (π Q ), and Var α j n (π Q ) ≤ Var 2α n (π Q ). Moreover, we have that 1 δ j ≤ 2 δα . This yields that the following inequality holds with probability at least 1− δ for any Q on H and for any α∈ (0, 1] L(π Q )≤ ˆ L α n (π Q ) + √︄ D KL (Q∥P) + log 8 √ n δα 2n + B α n (π Q ) + D KL (Q∥P) + log 8 δα nλ + λ 2 Var 2α n (π Q ). (E.17) Finally, setting kl ′ 1 (π Q ,α) = D KL (Q∥P) + log 8 √ n δα , kl ′ 2 (π Q ,α) = D KL (Q∥P) + log 8 δα , concludes the proof. 212 E.2.3 Proof of Proposition 6 Haddouche and Guedj (2022, Theorem 7) provides an application of Lemma 11 to the general PAC-Bayes learning problems in Section 8.3.1. We cannot apply their theorem directly to get Proposition 6 for two reasons. They assume that the loss function is non- negative and they derive a one-sided generalization bound. In our case, the loss function is negative and we want to derive a two-sided generalization bound. Fortunately, we show with a slight modification of their proof that the result can be extended to two-sided inequalities with negative losses. In fact, the only requirement is that the sign of loss is fixed. We show next how this is achieved. Proof. First, note that Lemma 11 does not make any assumption on the sign of the martingale difference sequence (f i (S i ,h)) i∈[n] nor on the sign of the terms that decompose it. Now similarly to the proof in Section E.2.1, we define S n = (X i ,A i ) i∈[n] as the set of n observed contexts and taken actions. Also, we let (F i ) i∈0∪[n] be a filtration adapted to S n . For h∈H, we define f i (S i ,h) as f i (S i ,h) = f i (X i ,A i ,h) = f (X i ,A i ,h) , = E X∼ν,A∼π 0 (·|X) [︃ I h(X)=A π 0 (A|X) α c(X,A) ]︃ − I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ). Here f i (S i ,h) only depends on the last samples X i ,A i and the predictor h. For this rea- son, we denote it by f i (X i ,A i ,h). Also, the function f i does not depend on i and this is why we simplify the notation as f i (X i ,A i ,h) = f (X i ,A i ,h). Moreover, the randomness in f (X i ,A i ,h) is only due X i ∼ ν and A i ∼ π 0 (·|X i ); all other terms are determin- istic. Thus the expectations are under X i ∼ ν,A i ∼ π 0 (·|X i ). Now similarly to the proof in Section E.2.1, we have that E [f (X i ,A i ,h)|F i−1 ] = 0 for any i ∈ [n],h ∈ H. Therefore, (f (X i ,A i ,h)) i∈[n] is a martingale difference sequence for any h ∈ H. Thus we apply Lemma 11 and get that that with probability at least 1− δ, the following holds simultaneously for any distribution Q on H |E h∼Q [M n (h)]|≤ D KL (Q∥P) + log(2/δ) λ + λ 2 (E h∼Q [⟨M⟩ n (h) + [M ] n (h)]) , (E.18) where M n (h) = n ∑︂ i=1 f (X i ,A i ,h) , ⟨M⟩ n (h) = n ∑︂ i=1 E [︁ f (X i ,A i ,h) 2 |F i−1 ]︁ , [M ] n (h) = n ∑︂ i=1 f (X i ,A i ,h) 2 . Now we compute E h∼Q [M n (h)] as E h∼Q [M n (h)] = n ∑︂ i=1 E X∼ν,A∼π 0 (·|X) [︃ π Q (A|X) π 0 (A|X) α c(X,A) ]︃ − π Q (A i |X i ) π 0 (A i |X i ) α c(X i ,A i ), = nE X∼ν,A∼π 0 (·|X) [︃ π Q (A|X) π 0 (A|X) α c(X,A) ]︃ − n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) α c(X i ,A i ), (E.19) 213 where we used the linearity of the expectation E h∼Q [·] and the definition of policies in (8.10). Moreover, similarly to the proof in Section E.2.1, we have that ⟨M⟩ n (h) + [M ] n (h) = n ∑︂ i=1 E [︁ f (X i ,A i ,h) 2 |F i−1 ]︁ + f (X i ,A i ,h) 2 = n ∑︂ i=1 E X∼ν,A∼π 0 (·|X) [︃ I h(X)=A π 0 (A|X) 2α c(X,A) 2 ]︃ + I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 − 2E X∼ν,A∼π 0 (·|X) [︃ I h(X)=A π 0 (A|X) α c(X,A) ]︃ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ), (i) ≤ nE X∼ν,A∼π 0 (·|X) [︃ I h(X)=A π 0 (A|X) 2α c(X,A) 2 ]︃ + n ∑︂ i=1 I h(X i )=A i π 0 (A i |X i ) 2α c(X i ,A i ) 2 , (E.20) where (i) holds since −2E X∼ν,A∼π 0 (·|X) [︂ I h(X)=A π 0 (A|X) α c(X,A) ]︂ I h(X i )=A i π 0 (A i |X i ) α c(X i ,A i ) ≤ 0 for any i∈ [n]. This is where the non-negative loss assumption is not needed. Our loss L α (h,x,a,c) = I h(X)=A π 0 (a|x) α c is negative since c ∈ [−1, 0]. However, we only need the product between the loss and its expectation to be non-negative. This holds in particular when the loss has a fixed sign. In that case, the expectation of the loss and the loss itself will have the same sign and thus their product will be non-negative. In our case, the loss has a fixed negative sign and this is all we needed. Now notice that nE X∼ν,A∼π 0 (·|X) [︃ π Q (A|X) π 0 (A|X) α c(X,A) ]︃ = nR α (π Q ), n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) α c(X i ,A i ) = n ˆ L α n (π Q ), where we used that c(X i ,A i ) = C i for any i∈ [n] in the second equality. Using these two equalities and plugging (E.19) and (E.20) in (E.18) yields that with probability at least 1− δ, the following holds simultaneously for any distribution Q on H n ⃓ ⃓ ⃓ L α (π Q )− ˆ L α n (π Q ) ⃓ ⃓ ⃓ ≤ D KL (Q∥P) + log(2/δ) λ + λ 2 (︂ nE X∼ν,A∼π 0 (·|X) [︃ π Q (A|X) π 0 (A|X) 2α c(X,A) 2 ]︃ + n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α c(X i ,A i ) 2 )︂ . (E.21) Again we used the linearity of the expectation E h∼Q [·] and the definition of policies in (8.10). Finally, we have that c(X i ,A i ) = C i for any i ∈ [n]. Thus with probability at least 1− δ the following inequality holds for any distribution Q on H ⃓ ⃓ ⃓ L α (π Q )− ˆ L α n (π Q ) ⃓ ⃓ ⃓ ≤ D KL (Q∥P) + log(2/δ) nλ + λ 2 E X∼ν,A∼π 0 (·|X) [︃ π Q (A|X) π 0 (A|X) 2α c(X,A) 2 ]︃ + λ 2n n ∑︂ i=1 π Q (A i |X i ) π 0 (A i |X i ) 2α C 2 i . (E.22) This concludes the proof. 214 E.2.4 Sample Complexity Notation reminder: Minimizing risk L(π) is equivalent to maximizing value V (π) = −L(π). Achieving L(ˆπ)≤L(π ∗ ) + ε is equivalent to V (ˆπ)≥ V (π ∗ )− ε. Proposition 17. LetM 1 (H) be the set of probability distributions on the hypothesis space H, and let λ > 0, n ≥ 1, δ ∈ [0, 1], α ∈ [0, 1], and let P be a fixed prior on H, then with probability at least 1− δ over draws D n ∼ μ n π 0 , we have L(π ˆ Q n )≤L(π Q ∗ ) + 2 √︃ kl 1 (π Q ∗ ) 2n + 2B α n (π Q ∗ ) + 2 kl 2 (π Q ∗ ) nλ + λ Var α n (π Q ∗ ). where π ˆ Q n is the learned policy with ˆ Q n = argmin Q∈M 1 (H) ˆ L α n (π Q ) + √︂ kl 1 (π Q ) 2n +B α n (π Q ) + kl 2 (π Q ) nλ + λ 2 Var α n (π Q ), Q ∗ = argmin Q∈M 1 (H) L(π Q ), and kl 1 (π Q ) = D KL (Q∥P) + log 4 √ n δ ,kl 2 (π Q ) = D KL (Q∥P) + log 4 δ , B α n (π Q ) = 1− 1 n n ∑︂ i=1 E A∼π Q (·|X i ) [︁ π 1−α 0 (A|X i ) ]︁ , Var α n (π Q ) = 1 n n ∑︂ i=1 E A∼π 0 (·|X i ) [︃ π Q (A|X i ) π 0 (A|X i ) 2α ]︃ + π Q (A i |X i )C 2 i π 0 (A i |X i ) 2α . Proof. First, Theorem 4 holds for any potentially data dependent distribution Q on H. In particular, we have that with probability at least 1− δ the following inequalities hold simultaneously for ˆ Q n and Q ∗ |L(π ˆ Q n )− ˆ L α n (π ˆ Q n )|≤ √︄ kl 1 (π ˆ Q n ) 2n + B α n (π ˆ Q n ) + kl 2 (π ˆ Q n ) nλ + λ 2 Var α n (π ˆ Q n ), |L(π Q ∗ )− ˆ L α n (π Q ∗ )|≤ √︃ kl 1 (π Q ∗ ) 2n + B α n (π Q ∗ ) + kl 2 (π Q ∗ ) nλ + λ 2 Var α n (π Q ∗ ). Taking only one side of these inequalities yields that with probability at least 1− δ the following inequalities hold simultaneously for ˆ Q n and Q ∗ L(π ˆ Q n )≤ ˆ L α n (π ˆ Q n ) + √︄ kl 1 (π ˆ Q n ) 2n + B α n (π ˆ Q n ) + kl 2 (π ˆ Q n ) nλ + λ 2 Var α n (π ˆ Q n ) ⏞ ⏟⏞ (I) , ˆ L α n (π Q ∗ )≤L(π Q ∗ ) + √︃ kl 1 (π Q ∗ ) 2n + B α n (π Q ∗ ) + kl 2 (π Q ∗ ) nλ + λ 2 Var α n (π Q ∗ ). Now using the definition of π ˆ Q n , we know that I ≤ ˆ L α n (π Q ∗ ) + √︃ kl 1 (π Q ∗ ) 2n + B α n (π Q ∗ ) + kl 2 (π Q ∗ ) nλ + λ 2 Var α n (π Q ∗ ). 215 This yields that with probability at least 1− δ the following inequalities hold simultane- ously for ˆ Q n and Q ∗ L(π ˆ Q n )≤ ˆ L α n (π Q ∗ ) + √︃ kl 1 (π Q ∗ ) 2n + B α n (π Q ∗ ) + kl 2 (π Q ∗ ) nλ + λ 2 Var α n (π Q ∗ ), ˆ L α n (π Q ∗ )≤L(π Q ∗ ) + √︃ kl 1 (π Q ∗ ) 2n + B α n (π Q ∗ ) + kl 2 (π Q ∗ ) nλ + λ 2 Var α n (π Q ∗ ). Computing the sum of these two inequalities concludes the proof. Corollary 2 (Special case of Proposition 17). Let H = ︁ h θ ;θ ∈ R dK ︁ of mappings h θ (x) = argmax a∈A φ(x) ⊤ θ a for any x ∈ X. Let n ≥ 1, δ ∈ [0, 1], α ∈ [0, 1], and let P = N (μ 0 ,I dK ) be a fixed prior on H, then with probability at least 1− δ over draws D n ∼ μ n π 0 , we have that L(π ˆ Q n )≤L(π Q ∗ ) + √︂ ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 √ n δ √ n + 2(1− K α−1 ) + ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 δ √ n + K 2α−1 + K 2α √ n . where π ˆ Q n is the learned policy with ˆ Q n = argmin Q=N (μ,I dK ) ˆ L α n (π Q )+ √︂ kl 1 (π Q ) 2n +B α n (π Q )+ kl 2 (π Q ) nλ + λ 2 Var α n (π Q ), Q ∗ = argmin Q=N (μ,I dK ) L(π Q ). Proof. This result follows from the general Proposition 17 by simply setting P =N (μ 0 ,I dK ) and Q ∗ = N (μ ∗ ,I dK ). First, since the covariance matrices of both distributions are I dK , their KL divergence is D KL (Q∥P) = ∥μ ∗ − μ 0 ∥ 2 /2. Moreover, since the logging policy is uniform then B α n (π Q ) = (1−K α−1 ) and Var α n (π Q )≤ K 2α−1 +K 2α . Using these quantities, setting λ = 1/ √ n and applying Proposition 17 yields that with probability at least 1− δ over draws D n ∼ μ n π 0 , we have that L(π ˆ Q n )≤L(π Q ∗ ) + √︂ ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 √ n δ √ n + 2(1− K α−1 ) + ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 δ √ n + K 2α−1 + K 2α √ n . This concludes the proof. The above corollary allows us to give insights into the sample complexity of our procedure. That is, the number of samples needed so that the performance of the learned policy π ˆ Q n is close to that of the optimal one. Let ε > 2(1−K α−1 ) for α∈ [1− log 2/ logK, 1]. This condition on α ensures that ε ∈ [0, 1] and it is mild as α is often close to 1. Let δ, then the following implication holds ε≥ √︂ ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 √ n δ √ n + 2(1− K α−1 ) + ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 δ √ n + K 2α−1 + K 2α √ n =⇒ P(L(π ˆ Q n )≤L(π Q ∗ ) + ε)≥ 1− δ . (E.23) 216 First, we use that √︂ ∥μ ∗ − μ 0 ∥ 2 + 2 log 4 √ n δ ≤∥μ ∗ −μ 0 ∥ + √︂ 2 log 4 √ n δ . Moreover we bound K 2α−1 + K 2α ≤ 2K 2α . Then the implication in (E.23) becomes √ n≥ ∥μ ∗ − μ 0 ∥ +∥μ ∗ − μ 0 ∥ 2 + 2 log 4 δ + √︂ 2 log 4 √ n δ + 2K 2α ε− 2(1− K α−1 ) =⇒ P(L(π ˆ Q n )≤L(π Q ∗ ) + ε)≥ 1− δ . (E.24) We only provide intuition on the sample complexity and aim at having easy-to-interpret terms. Thus we omit the logarithmic terms in (E.24) and assume that ∥μ ∗ − μ 0 ∥ 2 ≥ ∥μ ∗ −μ 0 ∥. This leads to the claim made in Section 8.4.1. Of course, a more precise sample complexity analysis can be made by studying the function h(x) = √ x− √︂ 2 log 4 √ x δ /(ε− 2(1− K α−1 )) and finding x such that f (x)≥ ∥μ ∗ −μ 0 ∥+∥μ ∗ −μ 0 ∥ 2 +2 log 4 δ +2K 2α ε−2(1−K α−1 ) . E.3 Experiments E.3.1 Setup We consider the standard supervised-to-bandit conversion (Agarwal et al., 2014). Pre- cisely, let S tr n and S ts n ts be the training and testing set of a classification dataset, respec- tively. First, we transform the training set S tr n to a logged bandit data D n as described in Algorithm 3. The resulting logged data D n is then used to train our policies. After that, the learned policies are tested onS ts n ts as described in Algorithm 4. We consider that the resulting reward in Algorithm 4 is a good proxy for the unknown true reward of the learned policies. This will be our performance metric, the higher the better. In our experiments, we use the following image classification datasets MNIST (LeCun et al., 1998), FashionMNIST (Xiao et al., 2017), EMNIST (Cohen et al., 2017) and CIFAR100 (Krizhevsky et al., 2009). We provide a summary of the statistics of these datasets in Table E.1. Algorithm 3 takes as input a logging policy π 0 which we define as π 0 (a|x) = exp(η 0 · φ(x) ⊤ μ 0,a ) ∑︁ a ′ ∈A exp(η 0 · φ(x) ⊤ μ 0,a ′ ) , ∀(x,a)∈X ×A.(E.25) Here φ(x) ∈ R d is the feature transformation function that outputs a d-dimensional vector, μ 0 = (μ 0,a ) a∈A ∈ R dK are learnable parameters and η 0 is an inverse-temperature parameter for the softmax in (E.25). We explain next how these quantities are derived in detail. The feature transformation function φ(x)∈ R d : for all the datasets, except CIFAR100, the feature transformation function φ(·) is defined as φ(x) = x ∥x∥ for any x ∈ X. That is, we simply normalize the features x∈X by their L 2 norm ∥x∥. In contrast, CIFAR100 is a more challenging problem. Thus we use transfer learning to extract features φ(x) expressive enough so that a linear softmax model would enjoy a reasonable performance. Precisely, we retrieve the last hidden layer of a ResNet-50 network, pre-trained on the ImageNet dataset, to output 2048-dimensional features. Finally, the obtained features 217 Table E.1: Statistics of the datasets used in our experiments. Data setNbr. train samples n Nbr. test samples n ts Nbr. actions K Dimension d MNIST600001000010784 FashionMNIST600001000010784 EMNIST1128001880047784 CIFAR10050000100001002048 are normalized as x ∥x∥ and this whole process (ResNet-50 + normalization) corresponds to φ(·) for CIFAR100. The parameters μ 0 = (μ 0,a ) a∈A ∈ R dK : we learn the parameters μ 0 using 5% of the training setS tr n . Precisely, we use the cross-entropy loss with an L 2 regularization of 10 −6 to prevent the logging policy π 0 from being degenerate. This ensures that the learning policies are absolutely continuous with respect to the logging policy π 0 , a condition under which standard IPS is unbiased. In optimization, we use Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 10 epochs. In all the experiments, we set the prior P = N (η 0 μ 0 ,I dK ) for the Gaussian policies in (8.13) and we set it as P = N (η 0 μ 0 ,I dK )× G(0, 1) K for the mixed-logit policies in (8.12). Our theory requires that the prior does not depend on data. Given that μ 0 is learned on the 5% portion of data, we only train our learning policies on the remaining 95% portion of the data to match our theoretical requirements. The inverse-temperature parameter η 0 ∈ R: this controls the performance of the logging policy. A high positive value of η 0 leads to a well-performing logging policy, while a negative one leads to a low-performing logging policy. When η 0 = 0, π 0 is identical to the uniform policy. In our experiments η 0 varies between 0 and 1. Algorithm 3 Supervised-to-bandit: creating logged data Input: training classification set S tr n =(X i ,y i ) n i=1 , logging policy π 0 . Output: logged bandit data D n = (X i ,A i ,C i ) i∈[n] . Initialize D n = for i = 1,...,n do A i ∼ π 0 (·|X i ) C i =−I A i =y i D n ←D n ∪(X i ,A i ,C i ). Algorithm 4 Supervised-to-bandit: testing policies Input: image classification dataset S ts n ts =(X i ,y i ) n ts i=1 , learned policy ˆπ n . Output: reward r. for i = 1,...,n ts do A i ∼ ˆπ n (·|X i ) R i = I A i =y i r = 1 n ts ∑︁ n ts i=1 R i . Now it remains to explain the learning policies π Q and the corresponding closed-form bounds using either our results or those in existing works (London and Sandler, 2019; 218 Sakhi et al., 2022). E.3.2 Policies Here we present the two families of policies that we use in our experiments, Gaussian and mixed-logit policies. Mixed-Logit LetH = ︁ h θ,γ ;θ ∈ R dK ,γ ∈ R K ︁ be a hypothesis space of mappings h θ,γ (x) = argmax a∈A φ(x) ⊤ θ a + γ a for any x ∈ X. Here φ(x) outputs a d-dimensional representation of context x ∈ X. Now assume that for any a ∈ A, γ a is a standard Gumbel perturbation, γ a ∼ G(0, 1), then we have that π sof θ (a|x) = exp(φ(x) ⊤ θ a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ θ a ′ ) , = E γ∼G(0,1) K [︁ I h θ,γ (x)=a ]︁ .(E.26) In addition, we randomize θ such as θ ∼ N (μ,σ 2 I dK ) where μ ∈ R dK and σ > 0. It follows that the posterior Q is a multivariate Gaussian N (μ,σ 2 I dK ) over the parameters θ with standard Gumbel perturbations γ ∼ G(0, 1) K . We denote such policies by π mixL μ,σ and they are defined as π mixL μ,σ (a|x) = E θ∼N (μ,σ 2 I dK ) [︃ exp(φ(x) ⊤ θ a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ θ a ′ ) ]︃ , = E θ∼N (μ,σ 2 I dK ) [π sof θ (a|x)] , = E θ∼N (μ,σ 2 I dK ),γ∼G(0,1) K [︁ I h θ,γ (x)=a ]︁ .(E.27) To sample from the mixed-logit policies π mixL μ,σ , we first sample θ ∼ N (μ,σ 2 I dK ) and γ ∼ G(0, 1) K and then set the sampled action as a ← h θ,γ (x). Now we also need to compute the gradient of the expectation in (E.27). This needs additional care since the distribution under which we take the expectation depends on the parameters μ,σ. Fortunately, the reparameterization trick can be used in this case. Roughly speaking, it allows us to express a gradient of the expectation in (E.27) as an expectation of a gradient. In our case, we use the local reparameterizaton trick (Kingma et al., 2015) which is known for reducing the variance of stochastic gradients. Precisely, we rewrite (E.27) as π mixL μ,σ (a|x) = E ε∼N (0,∥φ(x)∥ 2 I K ) [︃ exp(φ(x) ⊤ μ a + σε a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ μ a ′ + σε a ′ ) ]︃ . = E ε∼N (0,I K ) [︃ exp(φ(x) ⊤ μ a + σε a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ μ a ′ + σε a ′ ) ]︃ , where we used that ∥φ(x)∥ 2 = 1 since features are normalized. It follows that gradients read ∇ μ,σ π mixL μ,σ (a|x) = E ε∼N (0,I K ) [︃ ∇ μ,σ exp(φ(x) ⊤ μ a + σε a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ μ a ′ + σε a ′ ) ]︃ . 219 Moreover, the propensities are approximated as π mixL μ,σ (a|x)≈ 1 S ∑︂ i∈[S] exp(φ(x) ⊤ μ a + σε i,a ) ∑︁ a ′ ∈A exp(φ(x) ⊤ μ a ′ + σε i,a ′ ) , ε i ∼N (0,I K ),∀i∈ [S]. (E.28) In all our experiments, we set S = 32. Gaussian We define the hypothesis spaceH = ︁ h θ ;θ ∈ R dK ︁ of mappings h θ (x) = argmax a∈A φ(x) ⊤ θ a for any x∈X. It follows that the learning policies π Q = π gaus μ,σ read π gaus μ,σ (a|x) = E θ∼N (μ,σ 2 I dK ) [︁ I h θ (x)=a ]︁ .(E.29) To see why this can be beneficial (Sakhi et al., 2022), let π ∗ be the optimal policy. Given x ∈ X, π ∗ (·|x) should be deterministic; it chooses the best action for context x with probability 1. That is, there exists μ ∗ ∈ R dK such that π ∗ = I h μ ∗ (x)=a . When μ → μ ∗ and σ → 0, the Gaussian policy in (E.29) approaches π ∗ . In contrast, the mixed-logit policy in (E.27) approaches π sof μ ∗ . However, π sof μ ∗ is not deterministic due to the additional randomness in γ and is equal to π ∗ only if φ(x) ⊤ μ ∗,a ∗ (x) →∞. This explains the choice of removing the Gumbel noise. First, Sakhi et al. (2022) showed that (E.29) can be written as π gaus μ,σ (a|x) = E ε∼N (0,1) [︂ ∏︂ a ′ ̸=a Φ (︁ ε + φ(x) ⊤ (μ a − μ a ′ ) σ∥φ(x)∥ )︁ ]︂ , where Φ is the cumulative distribution function of a standard normal variable. But ∥φ(x)∥ = 1 in all our experiments. Thus π gaus μ,σ (a|x) = E ε∼N (0,1) [︂ ∏︂ a ′ ̸=a Φ (︁ ε + φ(x) ⊤ (μ a − μ a ′ ) σ )︁ ]︂ . Then similarly to mixed-logit policies, the gradient reads ∇ μ,σ π gaus μ,σ (a|x) = E ε∼N (0,1) [︂ ∇ μ,σ ∏︂ a ′ ̸=a Φ (︁ ε + φ(x) ⊤ (μ a − μ a ′ ) σ )︁ ]︂ . Moreover, the propensities are approximated as π gaus μ,σ (a|x)≈ 1 S ∑︂ i∈[S] ∏︂ a ′ ̸=a Φ (︁ ε i + φ(x) ⊤ (μ a − μ a ′ ) σ )︁ , ε i ∼N (0, 1),∀i∈ [S]. (E.30) In all our experiments, we set S = 32. E.3.3 Baselines Here we present all the methods that we use in our experiments. For each method, we state the result that holds for any learning policy π. After that, we derive the corresponding 220 closed-form bounds for Gaussian and mixed-logit policies that we presented previously. All the baselines require computing the KL divergence between the prior P and the posterior Q. Thus before presenting them, we state the following lemma that allows bounding the KL divergence between the prior P and the posterior Q in the cases of mixed-logit or Gaussian policies. Lemma 12 (KL divergence for Gaussian distributions with Gumbel noise). For distribu- tions P = N (μ 0 ,σ 2 0 I dK )× G(0, 1) K and Q = N (μ,σ 2 I dK )× G(0, 1) K , with μ 0 ,μ ∈ R dK and 0 < σ 2 ≤ σ 2 0 <∞, D KL (Q∥P)≤ ∥μ− μ 0 ∥ 2 2σ 2 0 + dK 2 log σ 2 0 σ 2 . Moreover, this result holds when the Gumbel noise is removed. That is when P =N (μ 0 ,σ 2 0 I dK ) and Q =N (μ,σ 2 I dK ). We borrow this lemma from London and Sandler (2019). In particular, Lemma 12 shows that the KL terms for both policies can be bounded by the same quantity. As a result, the corresponding bounds will be the same; the only difference is the space of learning policies on which we optimize. For completeness, however, we write these bounds for both types of policies although they are similar. Since existing approaches are not named, we name them as (Author, Policy) where Author ∈ Ours, London et al., Sakhi et al. 1, Sakhi et al. 2 and Policy ∈ Gaussian, Mixed-Logit . Here Ours, London et al., Sakhi et al. 1 and Sakhi et al. 2 correspond to Theorem 4, London and Sandler (2019, Theorem 1), Sakhi et al. (2022, Proposition 1), Sakhi et al. (2022, Proposition 3), respectively. For example, London and Sandler (2019, Theorem 1) leads to two baselines (London et al., Gaussian) and (London et al., Mixed-Logit). In all our experiments, the learning policies are trained using Adam (Kingma and Ba, 2014) with a learning rate of 0.1 for 20 epochs. Ours, Theorem 4 (Ours, Gaussian) Here we use the Gaussian policies in (E.29). Thus we only replace the term, D KL (Q∥P), with its closed-form bound in Lemma 12. This leads to the following objective. min μ∈R dK ,σ>0 (︂ ˆ L α n (︁ π gaus μ,σ )︁ + √︄ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 √ n δ 2n + B α n (π gaus μ,σ ) + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 δ nλ + λ 2 Var α n (π gaus μ,σ ) )︂ , where we used that σ 0 = 1 since our prior is P = N (η 0 μ 0 ,I dK ) for Gaussian policies. Moreover, we set λ = √︃ 2 ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 +log 4 δ n Var α n (π gaus μ,σ ) . (Ours, Mixed-Logit) Here we use the mixed-logit policies in (E.27). Thus we only replace the terms, D KL (Q∥P), with their closed-form bound in Lemma 12. This leads to 221 the following objective. min μ∈R dK ,σ>0 (︂ ˆ L α n (︁ π mixL μ,σ )︁ + √︄ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 √ n δ 2n + B α n (π mixL μ,σ ) + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 δ nλ + λ 2 Var α n (π mixL μ,σ ) )︂ , where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK )× G(0, 1) K for mixed-logit policies. Moreover, we set λ = √︃ 2 ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 +log 4 δ n Var α n (π mixL μ,σ ) . London and Sandler (2019, Theorem 1) Proposition 18. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1) and let P be a fixed prior on H, then with probability at least 1−δ over draws D n ∼ μ n π 0 , the following holds simultaneously for all posteriors, Q, on H that R (π Q )≤ ˆ L τ n (π Q ) + ⌜ ⃓ ⃓ ⎷ 2 (︂ ˆ L τ n (π Q ) + 1 τ )︂ (︁ D KL (Q∥P) + log n δ )︁ τ (n− 1) + 2 (︁ D KL (Q∥P) + log n δ )︁ τ (n− 1) . (E.31) Baseline 1: (London et al., Gaussian) Here we use the Gaussian policies in (E.29). Thus we only replace the terms, D KL (Q∥P), with their closed-form bound in Lemma 12. This leads to the following objective. min μ∈R dK ,σ>0 (︂ ˆ L τ n (︁ π gaus μ,σ )︁ + ⌜ ⃓ ⃓ ⎷ 2 (︂ ˆ L τ n (︁ π gaus μ,σ )︁ + 1 τ )︂(︂ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log n δ )︂ τ (n− 1) + 2 (︂ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log n δ )︂ τ (n− 1) )︂ , where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK ) for Gaussian policies. Baseline 2: (London et al., Mixed-Logit) Here we consider the mixed-logit policies in (E.27). Since the additional Gumbel noise does not affect the KL divergence (Lemma 12), we have the same objective as in the Gaussian case. That is min μ∈R dK ,σ>0 (︂ ˆ L τ n (︁ π mixL μ,σ )︁ + ⌜ ⃓ ⃓ ⎷ 2 (︂ ˆ L τ n (︁ π mixL μ,σ )︁ + 1 τ )︂(︂ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log n δ )︂ τ (n− 1) + 2 (︂ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log n δ )︂ τ (n− 1) )︂ , where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK )× G(0, 1) K for mixed-logit policies. 222 Sakhi et al. (2022, Proposition 1) Proposition 19. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1) and let P be a fixed prior on H, then with probability at least 1−δ over draws D n ∼ μ n π 0 , the following holds simultaneously for all posteriors, Q, on H that R (π Q )≤ min λ>0 1 τ (e λ − 1) (︂ 1− e −τλ ˆ L τ n (π Q )+ D KL (Q∥P)+log 2 √ n δ n )︂ .(E.32) Baseline 3: (Sakhi et al. 1, Gaussian) Here we use the Gaussian policies in (E.29). min μ∈R dK ,σ>0,λ>0 (︂ 1 τ (e λ − 1) (︂ 1− e −τλ ˆ L τ n ( π gaus μ,σ ) + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 +log 2 √ n δ n )︂)︂ ,(E.33) where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK ) for Gaussian policies. Baseline 4: (Sakhi et al. 1, Mixed-Logit) Here we consider the mixed-logit policies in (E.27). min μ∈R dK ,σ>0,λ>0 (︂ 1 τ (e λ − 1) (︂ 1− e −τλ ˆ L τ n ( π mixL μ,σ ) + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 +log 2 √ n δ n )︂)︂ .(E.34) where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK )× G(0, 1) K for mixed-logit policies. Sakhi et al. (2022, Proposition 3) Proposition 20. Let τ ∈ (0, 1), n ≥ 1, δ ∈ (0, 1), let P be a fixed prior on H, and let Λ =λ i i∈[n λ ] a set of n λ positive scalars. Then with probability at least 1− δ over draws D n ∼ μ n π 0 , the following holds simultaneously for all posteriors, Q, on H and any λ i ∈ Λ, L (π Q )≤ ˆ L τ n (π Q ) + √︄ D KL (Q∥P) + log 4 √ n δ 2n + D KL (Q∥P) + log 2n λ δ λ + λ n g (︃ λ τn )︃ V τ n (π Q ) , (E.35) where g : u→ exp(u)−1−u u 2 and V τ n (π Q ) = 1 n ∑︁ n i=1 E A∼π Q (·|X i ) [︂ π 0 (A|X i ) max(τ,π 0 (A|X i )) 2 ]︂ . Baseline 5: (Sakhi et al. 2, Gaussian) Here we consider the Gaussian policies in (E.29). min μ∈R dK ,σ>0,λ∈Λ (︂ ˆ L τ n (︁ π gaus μ,σ )︁ + √︄ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 √ n δ 2n + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 2n λ δ λ + λ n g (︃ λ τn )︃ V τ n (︁ π gaus μ,σ )︁ )︂ , (E.36) 223 where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK ) for Gaussian policies. Baseline 6: (Sakhi et al. 2, Mixed-Logit) Here we consider the mixed-logit policies in (E.27). min μ∈R dK ,σ>0,λ∈Λ (︂ ˆ L τ n (︁ π mixL μ,σ )︁ + √︄ ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 4 √ n δ 2n + ∥μ−μ 0 ∥ 2 2 − dK 2 logσ 2 + log 2n λ δ λ + λ n g (︃ λ τn )︃ V τ n (︁ π mixL μ,σ )︁ )︂ , (E.37) where we used that σ 0 = 1 since our prior is P =N (η 0 μ 0 ,I dK )× G(0, 1) K for mixed-logit policies. E.3.4 Additional Results and Discussion In Figure E.1, we report the reward of the learned policy using one of the considered methods. We make the following observations: • Choice of τ and α: in Figure E.1, we set τ = 1/ 4 √ n ≈ 0.06 and α = 1 − 1/ 4 √ n ≈ 0.94 so that when n is large enough, both ˆ L τ n (π) and ˆ L α n (π) approach ˆ L ips n (π) (Ionides, 2008). This is because standard IPS should be preferred when n → ∞. For completeness, we also show in Figure E.2 that the choice of α and τ does not affect the conclusions that we make here. We also include in Figure E.2 the results with an adaptive and data-dependent α obtained using (8.15) in Section 8.3.4. The results in Figure E.2 will be discussed in detail after we finish analyzing the results in Figure E.1. • Overall performance: our method outperforms the baselines for any class of learning policies (Gaussian or mixed-logit) and any choice of logging policies. The only exception is when the logging policy is uniform. • Effect of the class of learning policies: the class of policies, Gaussian or mixed- logit, affects the performance of all the baselines. In general, Gaussian policies behave better than mixed-logit policies. However, this is less significant for our method; the performance of both Gaussian and mixed-logit policies are comparable, and in both cases, our method outperforms the baselines with Gaussian policies. Therefore, in general, Gaussian policies should be preferred over mixed-logit policies. But in case engineering constraints impose the choice of mixed-logit or softmax policies, then the performance of our method is robust to this choice. • Effect of the logging policy: our method reaches the maximum reward even when the logging policy is not performing well. In contrast, the baselines only reach their best reward when the logging policy is already well-performing (η 0 ≈ 1), in which case minor to no improvements are made. Note that the baselines have a better reward than ours when the logging policy is uniform. But our method has better reward when the logging policy is not uniform, that is when η 0 > 0. This is more common in practice since the logging policy is deployed in production and thus it is expected to perform better than the uniform policy. 224 In Figure E.2, we compare our method to (Sakhi et al. 2) with Gaussian policies since this was the best-performing baseline in our experiments in Figure E.1. Note that we did not include CIFAR100 in Figure E.1 as it was computationally heavy to run these exper- iments with varying η 0 , α and τ for a very high-dimensional dataset such as CIFAR100. We consider 20 varying values of τ and α evenly spaced in (0, 1). We also include the results using the adaptive tuning procedure of α described in Section 8.3.4 (green curve). We make the following observations: • Adaptive and data-dependent α: This procedure is reliable since the perfor- mance with an adaptive α (green curve) is comparable with the best possible choice of α. This is consistent for the three datasets. • Effect of the choice α: as we observed before, the only case where the choice of α may lead to bad-performing policies is when the logging policy is uniform. When the logging policy is not uniform, our method outperforms the best baseline with the best τ for a wide range of values of α. Also, note that there is no very bad choice of α, in contrast with τ ≈ 0 that led to a very bad performing policy that slightly improved upon the logging policy. This attests to the robustness of our method to the choice of α. Moreover, our bound regularizes better α; it contains a bias-variance trade-off term for α. Also, the bound of (Sakhi et al. 2) has a 1/τ making it vacuous for small values of τ. • Best choice of α: To see the effect of α for varying problems, we consider the following experiment. We split the logging policies into two groups. The first is modest logging which corresponds to logging policies whose η 0 is between 0 and 0.5. This includes uniform logging policies and other average-performing logging policies. The second is good logging which corresponds to logging policies whose η 0 is between 0.5 and 1. After that, for each α, we compute the average reward of the learned policy across either the group of modest or good logging policies. For each dataset, this leads to the two red and green curves in the second row of Figure E.2. Overall, we observe that α ≈ 0.7 leads to the best performance for the modest logging group. Thus when the performance of the logging policy is average, regularizing the importance weights can be critical. In contrast, when the performance of the logging policy is already good, regularization is less needed and we can set α ≈ 1. Fortunately, one of the main strengths of this work is that our bound also holds for standard IPS recovered for α = 1. The bounds in all prior works cannot provide good performance for standard IPS due to their dependency on 1/τ. 225 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 Logging Ours Sakhi et al. 2 Ours, Adaptive ® inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 EMNIST, K=47, d=784 0.00.20.40.60.81.0 smoothing parameter ® 0.60 0.65 0.70 0.75 0.80 0.85 0.90 reward of the learned policy MNIST, K=10, d=784 Modest Logging Good Logging 0.00.20.40.60.81.0 smoothing parameter ® 0.50 0.55 0.60 0.65 0.70 0.75 0.80 reward of the learned policy FashionMNIST, K=10, d=784 Modest Logging Good Logging 0.00.20.40.60.81.0 smoothing parameter ® 0.20 0.25 0.30 0.35 0.40 0.45 0.50 0.55 0.60 reward of the learned policy EMNIST, K=10, d=784 Modest Logging Good Logging Figure E.2: In the first row, we report the reward of the learned policy with 20 evenly space values of τ ∈ (0, 1) and α ∈ (0, 1) and varying η 0 ∈ [0, 1], and for an adaptive and data-dependent α obtained using (8.15) in Section 8.3.4. The blue-to-cyan colors correspond to different values of τ. The lighter the color, the higher the value of τ. For instance, the cyan lines correspond to high values of τ while the blue ones correspond to very small values of τ. Similarly, the red-to-yellow colors correspond to different values α. The lighter the color, the higher the value of α. For instance, the yellow lines correspond to high values of α while the red ones correspond to very small values of α. Finally, the green curve corresponds to the reward of the learned policy using an adaptive and data-dependent α described in (8.15) (Section 8.3.4). In the second row, we report the average reward of the learned policies using our method across the modest logging group (η 0 ∈ [0, 0.5] in red) and the good logging group (η 0 ∈ [0.5, 1] in green). 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 EMNIST, K=47, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.00 0.05 0.10 0.15 0.20 0.25 0.30 CIFAR, K=100, d=2048 Ours, Gaussian Ours, Mixed-Logit London et al., Gaussian London et al., Mixed-Logit Sakhi et al. 1, Gaussian Sakhi et al. 1, Mixed-Logit Sakhi et al. 2, Gaussian Sakhi et al. 2, Mixed-Logit Logging Figure E.1: The reward of the learned policy for four datasets with varying quality of the logging policy η 0 ∈ [0, 1]. E.3.5 Learning Principles Here we compare our bound in Theorem 4 and our learning principle in (8.17) to the one in London and Sandler (2019). We do not include the learning principle in Swaminathan and Joachims (2015a) since the one in London and Sandler (2019) enjoys similar performance and is far more scalable. The learning principle of London and Sandler (2019) is defined 226 as min μ ˆ L τ n (π μ ) + λ∥μ− μ 0 ∥ 2 .(E.38) where λ is a tunable hyper-parameters, π μ is the softmax policy defined in (E.26) and μ∈ R dK is its parameter vector. This learning principle is referred to as (London et al., LP). In contrast, our learning principle is defined as ˆ L α n (π μ ) + λ 1 ∥μ− μ 0 ∥ 2 + λ 2 Var α n (π μ ) + λ 3 B α n (π μ ),(E.39) where λ 1 ,λ 2 and λ 3 are tunable hyper-parameters and π μ is the Gaussian policy in (8.13) with a fixed σ = 1. Our learning principle is referred to as (Ours, LP). Finally, our bound in Theorem 4 with Gaussian policies is referred to as (Ours, Bound). Similarly to the previous experiments, we set τ = 1/ 4 √ n ≈ 0.06 and α = 1− 1/ 4 √ n ≈ 0.94 so that when n is large enough, both ˆ L τ n (π) and ˆ L α n (π) approach ˆ L ips n (π) (Ionides, 2008). For the learning principles, we tried multiple values of hyper-parameters λ,λ 1 ,λ 2 and λ 3 , all between 10 −5 and 10 −1 . For instance, we found that the best hyper-parameter for London and Sandler (2019) is λ = 10 −5 which matches the value they found in their FashionMNIST experiments. For our learning principle, the best hyper-parameters were λ 1 = 10 −5 ,λ 2 = 10 −5 and λ 3 = 10 −5 . In contrast, our bound does not require hyper- parameter tuning. We report in Figure E.3 the reward of the learned policy on the FashionMNIST for all these methods with varying values of hyper-parameters. To reduce clutter, we only report the reward for good choices of hyper-parameters λ,λ 1 ,λ 2 and λ 3 . We observe that for a wide range of hyper-parameters, our learning principle outperforms the one in London and Sandler (2019). However, both learning principles are sensitive to the choice of hyper-parameters. In contrast, our bound does not require the tuning of any additional hyper-parameter and it achieves the best performance except for the uniform logging policy. In addition to being more theoretically grounded, this approach also enjoys favorable empirical performance without additional hyper-parameter tuning, an important practical consideration. 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 reward of the learned policy FashionMNIST, K=10, d=784 Logging Ours, LP London et al., LP Ours, Bound Figure E.3: The reward of the learned policy using either our bound in Theorem 4 (referred to as (Ours, Bound) in green), our learning principle in (8.17) (referred to as (Ours, LP) in red for multiple values of hyper-parameters) or the learning principle in London and Sandler (2019) (referred to as (London et al., LP) in blue) for multiple values of hyper-parameters). 227 E.3.6 Other Importance Weight Corrections Su et al. (2020); Metelli et al. (2021) also proposed corrections that are different from hard clipping (a detailed comparison is given in Section 8.2). However, they were not included in our main experiments since they do not provide generalization guarantees; they focus on OPE and only propose a heuristic for OPL in their Appendix B.2 and Section 6.1.2, respectively. Those heuristics are not based on theory, in contrast with ours which is directly derived from our generalization bound. However, for completeness, we also compare our regularization of importance weights to theirs. To make such a comparison, we use the hyper-parameters and tuning procedures provided in Section 6 and Appendix B.2 for Metelli et al. (2021) and Sections 5 and 6.1.2 for Su et al. (2020). Overall, we observe in Figure E.4 that our method outperforms these baselines in OPL and the gap is more significant when the logging policy is not performing well. 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 reward of the learned policy MNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 FashionMNIST, K=10, d=784 0.00.20.40.60.81.0 inverse-temperature parameter ́ 0 0.0 0.1 0.2 0.3 0.4 0.5 0.6 EMNIST, K=47, d=784 The reward of the learned policy using one of the baselines with varying quality of the logging policy ́ 0 2[0;1]. OursMeteli et al.(2021)Su et al.(2020)Logging Figure E.4: The reward of the learned policy with varying quality of the logging policy η 0 ∈ [0, 1] using either our regularization (α-IPS) or the ones in Su et al. (2020); Metelli et al. (2021). 228 Bibliography Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári. Improved algorithms for linear stochastic bandits. Advances in neural information processing systems, 24, 2011. M. Abeille and A. Lazaric. Linear Thompson sampling revisited. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, 2017. A. Agarwal, D. Hsu, S. Kale, J. Langford, L. Li, and R. Schapire. Taming the monster: A fast and simple algorithm for contextual bandits. In International Conference on Machine Learning, pages 1638–1646. PMLR, 2014. S. Agrawal and N. Goyal. Thompson sampling for contextual bandits with linear payoffs. In Proceedings of the 30th International Conference on Machine Learning, pages 127– 135, 2013a. S. Agrawal and N. Goyal. Further optimal regret bounds for thompson sampling. In Pro- ceedings of the Sixteenth International Conference on Artificial Intelligence and Statis- tics, pages 99–107, 2013b. S. Agrawal and N. Goyal. Near-optimal regret bounds for thompson sampling. Journal of the ACM (JACM), 64(5):1–24, 2017. P. Alquier. User-friendly introduction to pac-bayes bounds. arXiv preprint arXiv:2110.11216, 2021. I. Aouali. Linear diffusion models meet contextual bandits with large action spaces. In NeurIPS 2023 Foundation Models for Decision Making Workshop, 2023. I. Aouali. Diffusion models meet contextual bandits. In The Thirty-ninth Annual Con- ference on Neural Information Processing Systems, 2025. I. Aouali and O. Sakhi. Off-policy learning in large action spaces: Optimization matters more than estimation. arXiv preprint arXiv:2509.03456, 2025. I. Aouali, S. Ivanov, M. Gartrell, D. Rohde, F. Vasile, V. Zaytsev, and D. Legrand. Combining reward and rank signals for slate recommendation. arXiv preprint arXiv:2107.12455, 2021. 229 I. Aouali, A. Benhalloum, M. Bompaire, A. Ait Sidi Hammou, S. Ivanov, B. Heymann, D. Rohde, O. Sakhi, F. Vasile, and M. Vono. Reward optimizing recommendation using deep learning and fast maximum inner product search. In proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining, pages 4772–4773, 2022a. I. Aouali, A. Benhalloum, M. Bompaire, B. Heymann, O. Jeunen, D. Rohde, O. Sakhi, and F. Vasile. Offline evaluation of reward-optimizing recommender systems: The case of simulation. arXiv preprint arXiv:2209.08642, 2022b. I. Aouali, A. A. S. Hammou, O. Sakhi, D. Rohde, and F. Vasile. Probabilistic rank and reward: A scalable model for slate recommendation. arXiv preprint arXiv:2208.06263, 2022c. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Exponential smoothing for off-policy learning. In Proceedings of the 40th International Conference on Machine Learning, pages 984–1017. PMLR, 2023a. I. Aouali, B. Kveton, and S. Katariya. Mixed-effect thompson sampling. In International Conference on Artificial Intelligence and Statistics, pages 2087–2115. PMLR, 2023b. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Unified pac-bayesian study of pessimism for offline policy learning with regularized importance sampling. In Uncertainty in Artificial Intelligence, pages 88–109. PMLR, 2024. I. Aouali, V.-E. Brunel, D. Rohde, and A. Korba. Bayesian off-policy evaluation and learning for large action spaces. In International Conference on Artificial Intelligence and Statistics, pages 136–144. PMLR, 2025. P. Auer. Using confidence bounds for exploitation-exploration trade-offs. Journal of Machine Learning Research, 3:397–422, 2002. P. Auer, N. Cesa-Bianchi, and P. Fischer. Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47:235–256, 2002. M. G. Azar, A. Lazaric, and E. Brunskill. Sequential transfer in multi-armed bandit with finite set of models. In Advances in Neural Information Processing Systems 26, pages 2220–2228, 2013. H. Bang and J. M. Robins. Doubly robust estimation in missing data and causal inference models. Biometrics, 61(4):962–973, 2005. H. Bastani, D. Simchi-Levi, and R. Zhu. Meta dynamic pricing: Transfer learning across experiments. CoRR, abs/1902.10918, 2019. URL https://arxiv.org/abs/ 1902.10918. S. Basu, B. Kveton, M. Zaheer, and C. Szepesvari. No regrets for learning the prior in bandits. In Advances in Neural Information Processing Systems 34, 2021. B. Bercu and A. Touati. Exponential inequalities for self-normalized martingales with applications. 2008. 230 C. M. Bishop. Pattern Recognition and Machine Learning, volume 4 of Information science and statistics. Springer, 2006. L. Bottou, J. Peters, J. Quiñonero-Candela, D. X. Charles, D. M. Chickering, E. Portugaly, D. Ray, P. Simard, and E. Snelson. Counterfactual reasoning and learning systems: The example of computational advertising. Journal of Machine Learning Research, 14(11), 2013. S. Bubeck, N. Cesa-Bianchi, et al. Regret analysis of stochastic and nonstochastic multi- armed bandit problems. Foundations and Trends® in Machine Learning, 5(1):1–122, 2012. O. Catoni. Pac-bayesian supervised classification: the thermodynamics of statistical learn- ing. arXiv preprint arXiv:0712.0248, 2007. L. Cella, A. Lazaric, and M. Pontil. Meta-learning with stochastic linear bandits. In Proceedings of the 37th International Conference on Machine Learning, 2020. L. Cella, K. Lounici, and M. Pontil. Multi-task representation learning with stochastic linear bandits. arXiv preprint arXiv:2202.10066, 2022. O. Chapelle and L. Li. An empirical evaluation of Thompson sampling. In Advances in Neural Information Processing Systems 24, pages 2249–2257, 2012. M. Chen, R. Gummadi, C. Harris, and D. Schuurmans. Surrogate objectives for batch policy optimization in one-step decision making. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. W. Chu, L. Li, L. Reyzin, and R. Schapire. Contextual bandits with linear payoff func- tions. In Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, pages 208–214, 2011. H. Chung, J. Kim, M. T. Mccann, M. L. Klasky, and J. C. Ye. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022. M. Cief, J. Golebiowski, P. Schmidt, Z. Abedjan, and A. Bekasov. Learning action em- beddings for off-policy evaluation. In European Conference on Information Retrieval, pages 108–122. Springer, 2024. P. Clavier, T. Huix, and A. Durmus. Vits: Variational inference thomson sampling for contextual bandits. arXiv preprint arXiv:2307.10167, 2023. G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik. Emnist: Extending mnist to hand- written letters. In 2017 international joint conference on neural networks (IJCNN), pages 2921–2926. IEEE, 2017. V. Dani, T. P. Hayes, and S. M. Kakade. Stochastic linear optimization under bandit feedback. In COLT, volume 2, page 3, 2008. A. A. Deshmukh, U. Dogan, and C. Scott. Multi-task learning for contextual bandits. In Advances in Neural Information Processing Systems 30, pages 4848–4856, 2017. 231 P. Dhariwal and A. Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021. M. Dimakopoulou, N. Vlassis, and T. Jebara. Marginal posterior sampling for slate ban- dits. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 2223–2229. International Joint Conferences on Artificial Intelligence Organization, 2019. M. Dudík, J. Langford, and L. Li. Doubly robust policy evaluation and learning. Inter- national Conference on Machine Learning, 2011. M. Dudík, D. Erhan, J. Langford, and L. Li. Sample-efficient nonstationary policy eval- uation for contextual bandits. In Proceedings of the Twenty-Eighth Conference on Uncertainty in Artificial Intelligence, UAI’12, page 247–254, Arlington, Virginia, USA, 2012. AUAI Press. M. Dudik, D. Erhan, J. Langford, and L. Li. Doubly robust policy evaluation and opti- mization. Statistical Science, 29(4):485–511, 2014. A. Durand, C. Achilleos, D. Iacovides, K. Strati, G. D. Mitsis, and J. Pineau. Contex- tual bandits for adapting treatment in a mouse model of de novo carcinogenesis. In Proceedings of the 3rd Machine Learning for Healthcare Conference, volume 85, pages 67–82, 2018. M. Farajtabar, Y. Chow, and M. Ghavamzadeh. More robust doubly robust off-policy evaluation. In International Conference on Machine Learning, pages 1447–1456. PMLR, 2018. A. Farid and A. Majumdar. Generalization bounds for meta-learning via pac-bayes and uniform stability. Advances in Neural Information Processing Systems, 34:2173–2186, 2021. S. Filippi, O. Cappe, A. Garivier, and C. Szepesvari. Parametric bandits: The generalized linear case. In Advances in Neural Information Processing Systems 23, pages 586–594, 2010. H. Flynn, D. Reeb, M. Kandemir, and J. Peters. Pac-bayes bounds for bandit problems: A survey and experimental comparison. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(12):15308–15327, 2023. D. J. Foster, C. Gentile, M. Mohri, and J. Zimmert. Adapting to misspecification in contextual bandits. Advances in Neural Information Processing Systems, 33:11478– 11489, 2020. G. Gabbianelli, G. Neu, and M. Papini. Importance-weighted offline learning done right. In International Conference on Algorithmic Learning Theory, pages 614–634. PMLR, 2024. G. Garrigos and R. M. Gower. Handbook of convergence theorems for (stochastic) gradient methods. arXiv preprint arXiv:2301.11235, 2023. 232 C. Gentile, S. Li, and G. Zappella. Online clustering of bandits. In Proceedings of the 31st International Conference on Machine Learning, pages 757–765, 2014. A. Gilotte, C. Calauzènes, T. Nedelec, A. Abraham, and S. Dollé. Offline a/b testing for recommender systems. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining, pages 198–206, 2018. A. Gilotte, O. Sakhi, I. Aouali, and B. Heymann. Offline contextual bandit with counter- factual sample identification. arXiv preprint arXiv:2509.10520, 2025. A. Gopalan, S. Mannor, and Y. Mansour. Thompson sampling for complex online prob- lems. In Proceedings of the 31st International Conference on Machine Learning, pages 100–108, 2014. B. Guedj. A primer on pac-bayesian learning. arXiv preprint arXiv:1901.05353, 2019. S. Gupta, S. Chaudhari, S. Mukherjee, G. Joshi, and O. Yagan. A unified approach to translate classical bandit algorithms to the structured bandit setting. CoRR, abs/1810.08164, 2018. URL https://arxiv.org/abs/1810.08164. M. Haddouche and B. Guedj. Pac-bayes with unbounded losses through supermartingales. arXiv preprint arXiv:2210.00928, 2022. J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. J. Hong, B. Kveton, M. Zaheer, Y. Chow, A. Ahmed, and C. Boutilier. Latent bandits revisited. In Advances in Neural Information Processing Systems 33, 2020. J. Hong, B. Kveton, S. Katariya, M. Zaheer, and M. Ghavamzadeh. Deep hierarchy in bandits. In Proceedings of the 39th International Conference on Machine Learning, 2022a. J. Hong, B. Kveton, M. Zaheer, and M. Ghavamzadeh. Hierarchical Bayesian bandits. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics, 2022b. J. Hong, B. Kveton, M. Zaheer, S. Katariya, and M. Ghavamzadeh. Multi-task off-policy learning from bandit feedback. In International Conference on Machine Learning, pages 13157–13173. PMLR, 2023. D. G. Horvitz and D. J. Thompson. A generalization of sampling without replacement from a finite universe. Journal of the American statistical Association, 47(260):663–685, 1952. J. Hu, X. Chen, C. Jin, L. Li, and L. Wang. Near-optimal representation learning for linear bandits and linear rl. In International Conference on Machine Learning, pages 4349–4358. PMLR, 2021. E. L. Ionides. Truncated importance sampling. Journal of Computational and Graphical Statistics, 17(2):295–311, 2008. 233 O. Jeunen and B. Goethals. Pessimistic reward models for off-policy learning in rec- ommendation. In Fifteenth ACM Conference on Recommender Systems, pages 63–74, 2021. Y. Jin, Z. Yang, and Z. Wang. Is pessimism provably efficient for offline rl? In Interna- tional Conference on Machine Learning, pages 5084–5096. PMLR, 2021. E. Kaufmann, N. Korda, and R. Munos. Thompson sampling: An asymptotically optimal finite-time analysis. In International conference on algorithmic learning theory, pages 199–213. Springer, 2012. D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. D. P. Kingma, T. Salimans, and M. Welling. Variational dropout and the local reparam- eterization trick. Advances in neural information processing systems, 28, 2015. D. Koller and N. Friedman. Probabilistic graphical models: principles and techniques. MIT press, 2009. A. Korba and F. Portier. Adaptive importance sampling meets mirror descent: a bias- variance tradeoff. In International Conference on Artificial Intelligence and Statistics, pages 11503–11527. PMLR, 2022. N. Korda, E. Kaufmann, and R. Munos. Thompson sampling for 1-dimensional exponen- tial family bandits. Advances in neural information processing systems, 26, 2013. A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009. I. Kuzborskij and C. Szepesvári. Efron-stein pac-bayesian inequalities. arXiv preprint arXiv:1909.01931, 2019. I. Kuzborskij, C. Vernade, A. Gyorgy, and C. Szepesvári. Confident off-policy evaluation and selection through self-normalized importance weighting. In International Confer- ence on Artificial Intelligence and Statistics, pages 640–648. PMLR, 2021. B. Kveton, M. Zaheer, C. Szepesvari, L. Li, M. Ghavamzadeh, and C. Boutilier. Random- ized exploration in generalized linear bandits. In International Conference on Artificial Intelligence and Statistics, pages 2066–2076. PMLR, 2020. B. Kveton, M. Konobeev, M. Zaheer, C.-W. Hsu, M. Mladenov, C. Boutilier, and C. Szepesvari. Meta-Thompson sampling. In Proceedings of the 38th International Conference on Machine Learning, 2021. S. Lam and J. Herlocker. MovieLens Dataset. http://grouplens.org/datasets/movielens/, 2016. T. Lattimore and R. Munos. Bounded regret for finite-armed structured bandits. In Advances in Neural Information Processing Systems 27, pages 550–558, 2014. 234 T. Lattimore and C. Szepesvari. Bandit Algorithms. Cambridge University Press, 2019. B. Laurent and P. Massart. Adaptive estimation of a quadratic functional by model selection. Annals of Statistics, pages 1302–1338, 2000. Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. L. Li, W. Chu, J. Langford, and R. Schapire. A contextual-bandit approach to personal- ized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, 2010. L. Li, W. Chu, J. Langford, and X. Wang. Unbiased offline evaluation of contextual- bandit-based news article recommendation algorithms. In Proceedings of the fourth ACM international conference on Web search and data mining, pages 297–306, 2011. L. Li, Y. Lu, and D. Zhou. Provably optimal algorithms for generalized linear contextual bandits. In Proceedings of the 34th International Conference on Machine Learning, pages 2071–2080, 2017. D. Liang and N. Vlassis. Local policy improvement for recommender systems. arXiv preprint arXiv:2212.11431, 2022. D. Lindley and A. Smith. Bayes estimates for the linear model. Journal of the Royal Statistical Society: Series B (Methodological), 34(1):1–18, 1972. B. London and T. Sandler. Bayesian counterfactual risk minimization. In International Conference on Machine Learning, pages 4125–4133. PMLR, 2019. X. Lu and B. Van Roy. Information-theoretic confidence bounds for reinforcement learn- ing. In Advances in Neural Information Processing Systems 32, 2019. R. D. Luce. Individual choice behavior: A theoretical analysis. Courier Corporation, 2012. C. J. Maddison, D. Tarlow, and T. Minka. A* sampling. Advances in neural information processing systems, 27, 2014. O.-A. Maillard and S. Mannor. Latent bandits. In Proceedings of the 31st International Conference on Machine Learning, pages 136–144, 2014. A. Maurer and M. Pontil. Empirical bernstein bounds and sample variance penalization. arXiv preprint arXiv:0907.3740, 2009. D. A. McAllester. Some pac-bayesian theorems. In Proceedings of the eleventh annual conference on Computational learning theory, pages 230–234, 1998. P. McCullagh and J. A. Nelder. Generalized Linear Models. Chapman & Hall, 1989. J. Mei, C. Xiao, B. Dai, L. Li, C. Szepesvari, and D. Schuurmans. Escaping the gravi- tational pull of softmax. In Advances in Neural Information Processing Systems, vol- ume 33, pages 21130–21140. Curran Associates, Inc., 2020a. 235 J. Mei, C. Xiao, C. Szepesvári, and D. Schuurmans. On the global convergence rates of softmax policy gradient methods. In Proceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org, 2020b. A. M. Metelli, A. Russo, and M. Restelli. Subgaussian and differentiable importance sam- pling for off-policy evaluation and learning. Advances in Neural Information Processing Systems, 34:8119–8132, 2021. A. Nair, A. Gupta, M. Dalal, and S. Levine. Awac: Accelerating online reinforcement learning with offline datasets. arXiv preprint arXiv:2006.09359, 2020. N. Nguyen, I. Aouali, A. György, and C. Vernade. Prior-dependent allocations for bayesian fixed-budget best-arm identification in structured bandits. In International Conference on Artificial Intelligence and Statistics, pages 379–387. PMLR, 2025. M. Papini, A. M. Metelli, L. Lupo, and M. Restelli. Optimistic policy optimization via multiple importance sampling. In International Conference on Machine Learning, pages 4989–4999. PMLR, 2019. A. Peleg, N. Pearl, and R. Meirr. Metalearning linear bandits by prior update. In Proceedings of the 25th International Conference on Artificial Intelligence and Statistics, 2022. J. Peng, H. Zou, J. Liu, S. Li, Y. Jiang, J. Pei, and P. Cui. Offline policy evaluation in large action spaces via outcome-oriented action grouping. In Proceedings of the ACM Web Conference 2023, pages 1220–1230, 2023. X. B. Peng, A. Kumar, G. Zhang, and S. Levine. Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv preprint arXiv:1910.00177, 2019. J. Peters. Reinforcement learning by reward-weighted regression. In NIPS 2006 Workshop: Towards a New Reinforcement Learning?, 2006. J. Rappaz, J. McAuley, and K. Aberer. Recommendation on Live-Streaming Platforms: Dynamic Availability and Repeat Consumption, page 390–399. Association for Com- puting Machinery, 2021. I. Rejwan and Y. Mansour. Top-k combinatorial bandits with full-bandit feedback. In ALT, pages 752–776, 2020. S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme. Bpr: Bayesian person- alized ranking from implicit feedback. arXiv preprint arXiv:1205.2618, 2012. D. A. Reynolds et al. Gaussian mixture models. Encyclopedia of biometrics, 741(659-663), 2009. C. Riquelme, G. Tucker, and J. Snoek. Deep bayesian bandits showdown: An empir- ical comparison of bayesian deep networks for thompson sampling. arXiv preprint arXiv:1802.09127, 2018. 236 H. Robbins and S. Monro. A stochastic approximation method. The annals of mathe- matical statistics, pages 400–407, 1951. J. M. Robins and A. Rotnitzky. Semiparametric efficiency in multivariate regression models with missing data. Journal of the American Statistical Association, 90(429): 122–129, 1995. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. D. Russo and B. Van Roy. Learning to optimize via posterior sampling. Mathematics of Operations Research, 39(4):1221–1243, 2014. D. Russo and B. Van Roy. An information-theoretic analysis of Thompson sampling. Journal of Machine Learning Research, 17(68):1–30, 2016. D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, Z. Wen, et al. A tutorial on thompson sampling. Foundations and Trends® in Machine Learning, 11(1):1–96, 2018. N. Sachdeva, Y. Su, and T. Joachims. Off-policy bandits with deficient support. In Pro- ceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 965–975, 2020. N. Sachdeva, L. Wang, D. Liang, N. Kallus, and J. McAuley. Off-policy evaluation for large action spaces via policy convolution. In Proceedings of the ACM Web Conference 2024, pages 3576–3585, 2024. Y. Saito and T. Joachims. Off-policy evaluation for large action spaces via embeddings. arXiv preprint arXiv:2202.06317, 2022. Y. Saito, Q. Ren, and T. Joachims. Off-policy evaluation for large action spaces via conjunct effect modeling. In international conference on Machine learning, pages 29734– 29759. PMLR, 2023. Y. Saito, J. Yao, and T. Joachims. POTEC: Off-policy contextual bandits for large action spaces via policy decomposition. In The Thirteenth International Conference on Learning Representations, 2025. O. Sakhi, S. Bonner, D. Rohde, and F. Vasile. Blob: A probabilistic model for recom- mendation that combines organic and bandit signals. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 783– 793, 2020. O. Sakhi, N. Chopin, and P. Alquier. Pac-bayesian offline contextual bandits with guar- antees. arXiv preprint arXiv:2210.13132, 2022. O. Sakhi, D. Rohde, and N. Chopin. Fast slate policy optimization: Going beyond plackett-luce. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=f7a8XCRtUu. 237 O. Sakhi, I. Aouali, P. Alquier, and N. Chopin. Logarithmic smoothing for pessimistic off- policy evaluation, selection and learning. Advances in Neural Information Processing Systems, 37:80706–80755, 2024. S. Scott. A modern bayesian look at the multi-armed bandit. Applied Stochastic Models in Business and Industry, 26:639 – 658, 2010. A. Shrivastava and P. Li. Asymmetric lsh (alsh) for sublinear time maximum inner product search (mips). In Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. M. Simchowitz, C. Tosh, A. Krishnamurthy, D. Hsu, T. Lykouris, M. Dudik, and R. Schapire. Bayesian decision-making under misspecified priors with applications to meta-learning. In Advances in Neural Information Processing Systems 34, 2021. A. Slivkins. Introduction to multi-armed bandits. Foundations and Trends® in Machine Learning, 12(1-2):1–286, 2019. J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256–2265. PMLR, 2015. Y. Su, M. Dimakopoulou, A. Krishnamurthy, and M. Dudík. Doubly robust off-policy evaluation with shrinkage. In International Conference on Machine Learning, pages 9167–9176. PMLR, 2020. A. Swaminathan and T. Joachims. Batch learning from logged bandit feedback through counterfactual risk minimization. The Journal of Machine Learning Research, 16(1): 1731–1755, 2015a. A. Swaminathan and T. Joachims. The self-normalized estimator for counterfactual learn- ing. advances in neural information processing systems, 28, 2015b. A. Swaminathan, A. Krishnamurthy, A. Agarwal, M. Dudik, J. Langford, D. Jose, and I. Zitouni. Off-policy evaluation for slate recommendation. Advances in Neural Infor- mation Processing Systems, 30, 2017. M. F. Taufiq, A. Doucet, R. Cornish, and J.-F. Ton. Marginal density ratio for off-policy evaluation in contextual bandits. Advances in Neural Information Processing Systems, 36, 2024. W. R. Thompson. On the likelihood that one unknown probability exceeds another in view of the evidence of two samples. Biometrika, 25(3-4):285–294, 1933. S. Tomkins, P. Liao, P. Klasnja, and S. Murphy. Intelligentpooling: Practical thompson sampling for mhealth. Machine learning, 110(9):2685–2727, 2021. N. Tripuraneni, C. Jin, and M. Jordan. Provable meta-learning of linear representations. In International Conference on Machine Learning, pages 10434–10443. PMLR, 2021. 238 J. A. Tropp. User-friendly tail bounds for sums of random matrices. Foundations of computational mathematics, 12(4):389–434, 2012. I. Urteaga and C. Wiggins. Variational inference for the multi-armed contextual bandit. In Proceedings of the 21st International Conference on Artificial Intelligence and Statistics, pages 698–706, 2018. T. Van Erven and P. Harremos. Rényi divergence and kullback-leibler divergence. IEEE Transactions on Information Theory, 60(7):3797–3820, 2014. M. Wan, R. Misra, N. Nakashole, and J. J. McAuley. Fine-grained spoiler detection from large-scale review corpora. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 2605–2610. Association for Computational Linguistics, 2019. R. Wan, L. Ge, and R. Song. Metadata-based multi-task bandits with Bayesian hierar- chical models. In Advances in Neural Information Processing Systems 34, 2021. R. Wan, L. Ge, and R. Song. Towards scalable and robust structured bandits: A meta- learning framework. CoRR, abs/2202.13227, 2022. URL https://arxiv.org/abs/ 2202.13227. L. Wang, A. Krishnamurthy, and A. Slivkins. Oracle-efficient pessimism: Offline policy optimization in contextual bandits. arXiv preprint arXiv:2306.07923, 2023. Y.-X. Wang, A. Agarwal, and M. Dudık. Optimal and adaptive off-policy evaluation in contextual bandits. In International Conference on Machine Learning, pages 3589– 3597. PMLR, 2017. Z. Wang, A. Novikov, K. Zolna, J. S. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess, et al. Critic regularized regression. Advances in Neural Information Processing Systems, 33:7768–7778, 2020. N. Weiss. A Course in Probability. Addison-Wesley, 2005. H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmark- ing machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017. Y. Xu and A. Zeevi. Upper counterfactual confidence bounds: a new optimism principle for contextual bandits. arXiv preprint arXiv:2007.07876, 2020. J. Yang, W. Hu, J. D. Lee, and S. S. Du. Impact of representation learning in linear bandits. arXiv preprint arXiv:2010.06531, 2020. T. Yu, B. Kveton, Z. Wen, R. Zhang, and O. Mengshoel. Graphical models meet bandits: A variational Thompson sampling approach. In Proceedings of the 37th International Conference on Machine Learning, 2020. W. Zhu. Classification of mnist handwritten digit database using neural network. Proceed- ings of the research school of computer science. Australian National University, Acton, ACT, 2601, 2018. 239 Y. Zhu, D. J. Foster, J. Langford, and P. Mineiro. Contextual bandits with large ac- tion spaces: Made practical. In International Conference on Machine Learning, pages 27428–27453. PMLR, 2022. 240 Titre : Apprentissage on-policy et off-policy pour les grands espaces d’actions Mots cl ́ es : apprentissage, apprentissage par renforcement, syst ` emes interactifs, approximation R ́ esum ́ e : Cette th ` ese ́ etudie l’apprentissage de poli- tiques dans les syst ` emes interactifs o ` u un agent ob- serve un contexte, choisit une action parmi un tr ` es grand ensemble, puis rec ̧oit un retour partiel. Le cadre principal est celui des bandits contextuels, avec deux paradigmes : l’apprentissage en ligne, o ` u l’agent in- teragit s ́ equentiellement avec l’environnement et mi- nimise le regret, et l’apprentissage hors politique, o ` u il apprend ` a partir de donn ́ ees journalis ́ ees par une politique de logging. Dans les grands espaces d’ac- tions, ces deux cadres soul ` event des difficult ́ es ma- jeures : exploration co ˆ uteuse, faible couverture des donn ́ ees, forte variance des poids d’importance, biais d’extrapolation et objectifs difficiles ` a optimiser. La premi ` ere partie propose des m ́ ethodes bay ́ esiennes structur ́ ees pour l’apprentissage en ligne. Nous in- troduisons meTS, une extension de Thompson sam- pling fond ́ e sur des effets mixtes, puis dTS, qui ex- ploite des priors inspir ́ es des mod ` eles de diffusion. Ces m ́ ethodes partagent l’information entre actions et obtiennent des garanties de regret d ́ ependant d’un nombre effectif d’actions. La seconde partie traite l’apprentissage hors politique. Nous proposons sDM, une m ́ ethode directe structur ́ e fond ́ e sur des va- riables latentes, montrons que l’erreur d’optimisation peut dominer l’erreur d’estimation dans les grands espaces d’actions, et introduisons des objectifs de vraisemblance pond ́ er ́ e par la politique, concaves et efficaces ` a optimiser. Enfin, nous d ́ eveloppons des m ́ ethodes pessimistes diff ́ erentiables fond ́ ees sur le lissage exponentiel et des bornes PAC-bay ́ esiennes pour contr ˆ oler le compromis biais-variance des esti- mateurs par importance sampling. Title : On-Policy and Off-Policy Learning for Large Action Spaces Keywords : learning, reinforcement learning, interactive systems, approximation Abstract : This thesis studies policy learning in inter- active systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contex- tual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the envi- ronment and minimizes regret, and off-policy learning, where it learns from logged data collected by a log- ging policy. In large action spaces, both settings face major challenges: inefficient exploration, sparse data coverage, high-variance importance weights, extrapo- lation bias, and difficult optimization landscapes. The first part develops structured Bayesian methods for on-policy learning. We introduce meTS, a mixed-effect extension of Thompson sampling, and dTS, which le- verages diffusion-inspired priors to model dependen- cies between actions. These methods share infor- mation across actions and yield regret guarantees depending on an effective number of actions. The second part addresses off-policy learning. We pro- pose sDM, a structured direct method based on la- tent variables, show that optimization error can domi- nate estimation error in large action spaces, and in- troduce policy-weighted log-likelihood objectives that are concave and efficiently optimizable. Finally, we develop differentiable pessimistic methods based on exponential smoothing and PAC-Bayesian bounds to control the bias-variance trade-off of regularized importance-sampling estimators. Institut Polytechnique de Paris 91120 Palaiseau, France