Paper deep dive
Aligning Language Models with Preferences through f-divergence Minimization
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, Marc Dymetman
Models: GPT-2 Large (774M), GPT-2 Medium (345M), GPT-2 Small (117M), GPT-2 XL (1.5B), T5
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 7:59:10 PM
Summary
The paper introduces f-DPG, a unifying framework for aligning language models with human preferences by minimizing any f-divergence between the model and a target distribution. It generalizes existing methods like RLHF and GDC, demonstrating that choosing appropriate divergence objectives (such as Jensen-Shannon) can significantly improve alignment and diversity trade-offs compared to standard KL-based approaches.
Entities (5)
Relation Signals (3)
f-DPG → unifies → RLHF
confidence 95% · f-DPG unifies both frameworks (RLHF, GDC) and the approximation methods
f-DPG → unifies → GDC
confidence 95% · f-DPG unifies both frameworks (RLHF, GDC)
Jensen-Shannon divergence → outperforms → forward KL divergence
confidence 90% · Jensen-Shannon divergence... frequently outperforms forward KL divergence by a wide margin
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing approaches differ both in the functional form of the target distribution and the algorithm used to approximate it. For instance, Reinforcement Learning from Human Feedback (RLHF) corresponds to minimizing a reverse KL from an implicit target distribution arising from a KL penalty in the objective. On the other hand, Generative Distributional Control (GDC) has an explicit target distribution and minimizes a forward KL from it using the Distributional Policy Gradient (DPG) algorithm. In this paper, we propose a new approach, f-DPG, which allows the use of any f-divergence to approximate any target distribution that can be evaluated. f-DPG unifies both frameworks (RLHF, GDC) and the approximation methods (DPG, RL with KL penalties). We show the practical benefits of various choices of divergence objectives and demonstrate that there is no universally optimal objective but that different divergences present different alignment and diversity trade-offs. We show that Jensen-Shannon divergence strikes a good balance between these objectives, and frequently outperforms forward KL divergence by a wide margin, leading to significant improvements over prior work. These distinguishing characteristics between divergences persist as the model size increases, highlighting the importance of selecting appropriate divergence objectives.
Tags
Links
- Source: https://arxiv.org/abs/2302.08215
- Canonical: https://arxiv.org/abs/2302.08215
Trouble viewing inline? Open PDF directly →
Full Text
129,665 characters extracted from source content.
Expand or collapse full text
Aligning Language Models with Preferences throughf-divergence Minimization Dongyoung Go 1 2 Tomasz Korbak 3 Germ ́ an Kruszewski 4 Jos Rozen 4 Nahyeon Ryu 1 Marc Dymetman 5 Abstract Aligning language models with preferences can be posed as approximating a target distribution representing some desired behavior. Existing ap- proaches differ both in the functional form of the target distribution and the algorithm used to ap- proximate it. For instance, Reinforcement Learn- ing from Human Feedback (RLHF) corresponds to minimizing a reverse KL from animplicittarget distribution arising from a KL penalty in the objec- tive. On the other hand, Generative Distributional Control (GDC) has anexplicittarget distribution and minimizes a forward KL from it using the Distributional Policy Gradient (DPG) algorithm. In this paper, we propose a new approach,f-DPG, which allows the use ofanyf-divergence to ap- proximateanytarget distribution that can be eval- uated.f-DPG unifies both frameworks (RLHF, GDC) and the approximation methods (DPG, RL with KL penalties). We show the practical benefits of various choices of divergence objectives and demonstrate that there is no universally optimal objective but that different divergences present different alignment and diversity trade-offs. We show that Jensen-Shannon divergence strikes a good balance between these objectives, and fre- quently outperforms forward KL divergence by a wide margin, leading to significant improvements over prior work. These distinguishing characteris- tics between divergences persist as the model size increases, highlighting the importance of select- ing appropriate divergence objectives. 1. Introduction Language models (LMs) have recently revolutionized the 1 Naver Corp 2 Yonsei University 3 University of Sussex 4 Naver Labs Europe 5 Independent Researcher. Correspondence to: Dongy- oung Go<dongyoung.go@navercorp.com>. Proceedings of the40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023. Copyright 2023 by the author(s). Figure 1.On many target distributions, the Jensen-Shannon (JS) divergence (green) outperforms the Kullback-Leibler (KL) diver- gence (blue) as anobjective,even when performance is measured in terms of KL from the targetp(left panel,↓better). See Sec. 4.2. field of Natural Language Processing thanks to their gen- erative capabilities, which are useful in a vast number of tasks (Brown et al., 2020; Srivastava et al., 2022). How- ever, generated texts can also violate widely-held human preferences, e.g. helpfulness (Askell et al., 2021), non- offensiveness (Gehman et al., 2020), truthfulness (Lin et al., 2022) or equal treatment (Cao et al., 2022). Aligning LMs with human preferences is the problem of adapting the LM in such a way that generated content is perceived to match the human’s intent (Ouyang et al., 2022) or that it is help- ful, honest, and harmless (Askell et al., 2021; Bai et al., 2022b). Fundamentally, an aligned LM can be seen as a desired target distribution that we would like to generate from (Korbak et al., 2022c). Some approaches leave this distribution implicit, to be defined as a side-effect of the proposed intervention. These include prompting with nat- ural language instructions or demonstrations (Askell et al., 2021), using scorers or safety filters while decoding (Roller et al., 2021; Xu et al., 2021), supervised fine-tuning on cu- rated data (Solaiman & Dennison, 2021; Ngo et al., 2021; Welbl et al., 2021; Chung et al., 2022) or selected samples from the model (Zelikman et al., 2022; Scheurer et al., 2022; Dohan et al., 2022), and fine-tuning the language model us- ing reinforcement learning with a learned reward function that approximates human feedback (Reinforcement Learn- ing from Human Feedback or RLHF; Ziegler et al., 2019; Bai et al., 2022a; Ouyang et al., 2022). Instead, Khalifa et al. (2021) propose a framework that they name Generation with Distributional Control (GDC), where they define the target 1 arXiv:2302.08215v2 [cs.CL] 6 Jun 2023 Aligning Language Models with Preferences throughf-divergence Minimization distributionpthat represents the aligned LM as an EBM (Energy Based Model), namely an unnormalized version of pthat can be evaluated over any inputx. They then train the generative modelπ θ to approximatepvia methods such as Distributional Policy Gradients (DPG; Parshakova et al., 2019), which minimize the forward Kullback-Leibler (KL) divergenceKL(p||π θ )ofptoπ θ . The advantage of such an approach is that it decouples the problem of describing the aligned LM from the problem of approximating it. Fur- thermore, even if RL with KL penalties (Todorov, 2006a; Kappen et al., 2012; Jaques et al., 2017; 2019), the method used to fine-tune a LM in RLHF, is defined only in terms of reward maximization, it has also been shown to be equiva- lent to minimizing thereverseKL divergenceKL(π θ ||p)of π θ to a target distributionpthat can also be written explicitly in closed-form (Korbak et al., 2022b). The possibility of approximating various distributions ac- cording to different divergence measures begs the question: Does the choice of a divergence measure matter? In prin- ciple, all divergences lead to the same optimum, namely the target distributionp. However, when we restrictπ θ to a certain parametric family that does not includep(i.e., the search space ismis-specified), then the minimum can be found at different points, leading to optimal models with dif- ferent properties. Moreover, different divergences present different loss landscapes: some might make it easier for stochastic gradient descent to find good minima. Finally, the space of possible divergence measures and forms of target distributions is a vast and largely uncharted terrain. Prior work has largely failed to decouple the form of a target distribution and the algorithm used for approximating it. Here, we introducef-DPG, a new framework for fine- tuning an LM to approximate any given target EBM, by exploiting any given divergence in thef-divergences family, which includes not only the forward KL and the reverse KL cited above, but also Total Variation (TV) distance, Jensen- Shannon (JS) divergence, among others.f-DPG generalizes existing approximation techniques both DPG and RL with KL penalties algorithms, thus allowing us to investigate new ways to approximate the target distributions defined by the GDC and RLHF frameworks. In particular, we explore the approximation of various target distributions representing different alignment goals, which include imposing lexical constraints, reducing social bias with respect to gender and religion, enforcing factual consistency in summarization, and enforcing compilability of generated code. We focus our experiments on four instantiations off-DPG, namely KL-DPG, RKL-DPG, TV-DPG and JS-DPG, whose objec- tive is to minimize the forward KL, reverse KL, TV and JS divergences, respectively, and evaluate each experiment in terms of approximation quality as measured by all of these f-divergences. We show that we can obtain significantly improved results over the original KL-DPG algorithm (Par- shakova et al., 2019) by minimizing otherf-divergences, even when the approximation quality is evaluated under the lens of the forward KL. Furthermore, we observe that while there is no single best optimization objective for all cases, JS-DPG often strikes a good balance and significantly improves upon prior work (Khalifa et al., 2021; Korbak et al., 2022a), as illustrated in Fig. 1. Lastly, we find that f-DPG with an optimal objective continues to outperform suboptimal objectives as we scale model size from 127M parameters to 1.5B parameters (Sec. 4.5). The smooth and gradual scaling trend observed with increasing model size suggests that our findings will generalize to even larger LMs. Overall, the contributions of the paper include: 1.Introducingf-DPG, a unifying framework for approx- imating any EBM target distribution by minimizing anyf-divergence (Sec. 3.2), and deriving a universal formula for gradient descent withf-divergences (The- orem 1). 2. Extendingf-DPG to include baselines for variance reduction (Fact 1); and handling conditional target dis- tributions (Fact 2). 3.Investigating the performance off-DPG on a diverse array of thirteen LM alignment tasks, three forms of target distributions, fourf-divergence objectives and eight metrics. 2. Background We can organize approaches to LM alignment along two axes: how the target distribution is constructed and how it is approximated. The first problem roughly corresponds to representing human preferences through the specification of a probability distribution and the second to allowing the production of samples from that distribution. 2.1. Defining a Target Distribution The target distribution expresses an ideal notion of an LM, incorporating human preferences, as probabilitiesp(x)over textsxaccording to how well they satisfy the preferences. Formally,p(x)is often defined through a non-negative func- tionP(x)(aka anenergy-based modelor EBM (LeCun et al., 2006)) such thatp(x)∝P(x). The modelP(x)(and p(x)after normalization) can be used to score samples, but not to directly produce them because it lacks an autoregres- sive form. In the rest of the paper, we will focus on target distributions modeling three types of preferences promi- nently employed in recent literature about GDC (Khalifa et al., 2021) and RLHF (Ziegler et al., 2019; Stiennon et al., 2020; Ouyang et al., 2022; Menick et al., 2022; Bai et al., 2022a). 2 Aligning Language Models with Preferences throughf-divergence Minimization Binary preferencesFor human preferences naturally ex- pressible as a binary constraintb(x)∈0,1(e.g. a sample xmust never contain a curse word), Khalifa et al. (2021) proposed the following target distribution: p GDCbin (x)∝a(x)b(x),(1) whereais a pretrained LM andb(x) = 0ifxcontains a curse andb(x) = 1otherwise.p GDCbin is the distri- bution enforcing that all samples match the binary con- straint, which deviates minimally fromaas measured by KL(p GDCbin ||a). Scalar preferencesSome human preferences, such as helpfulness, are more naturally expressed as scalar scores. Alignment with respect to these is typically addressed with RLHF (Stiennon et al., 2020; Ziegler et al., 2019; Ouyang et al., 2022), which consists of, first, capturing human prefer- ences as a reward functionr(x)(e.g. scores given a reward model trained to predict human preferences) and second, applying RL with KL penalties (Todorov, 2006a; Kappen et al., 2012; Jaques et al., 2017; 2019) to maximize this reward while penalizing departure froma(x): J RLKL (θ) =E x∼π θ r(x)−βlog π θ (x) a(x) .(2) This objective can be equivalently framed as minimizing the reverse KL,KL(π θ ||p RLKL ), where the target distribution p RLKL is defined as: p RLKL (x)∝a(x) exp(r(x)/β),(3) whereβis a hyperparameter (Korbak et al., 2022b). Distributional preferencesFinally, there is a class of distributional preferences (Weidinger et al., 2021) that can- not be expressed as a function of a single samplexbut depend on the entire distribution, e.g. a particular gender distribution of persons mentioned in LM samples. Khalifa et al. (2021) model such preferences through distributional constraints using the following exponential family target distribution p GDC dist (x)∝a(x) exp h X i λ i φ i (x) i ,(4) whereφ i are features defined over texts (e.g. the most frequent gender of people mentioned inx) andλ i are co- efficients chosen so that the expected valuesE x∼p [φ i (x)] match some desired values ̄μ i (e.g., 50% gender balance). The resulting distributionp GDC-d matches the target feature moments, while deviating minimally fromaas measured byKL(p GDCdist ||a). 2.2. Approximating the target distribution Drawing samples from a target distributionpconstitutes the inference problem. There are broadly two approaches to this problem: (i) augmenting decoding fromaat infer- ence time to obtain samples frompand (i) training a new parametric modelπ θ to approximatepwhich can then be sampled from directly. The first family of approaches in- cludes guided decoding methods (Dathathri et al., 2020; Qin et al., 2022), Monte Carlo sampling techniques such as rejection sampling to sample from simple distributions likep GDCbin (Roller et al., 2021; Ziegler et al., 2022), and Quasi Rejection Sampling (QRS) (Eikema et al., 2022) or MCMC techniques (Miao et al., 2019; Goyal et al., 2022) to sample from more complex distributions, such asp GDC dist . In the rest of the paper, we will focus on the second family: methods that train a new modelπ θ to approximatepby min- imizing a divergence measure fromp,D(π θ ||p). Khalifa et al. (2021) uses Distributional Policy Gradients (DPG; Par- shakova et al., 2019) to approximate the target distribution by minimizingKL(p||π θ ), or equivalently,CE(p,π θ ): ∇ θ CE(p,π θ ) =−E x∼π θ p(x) π θ (x) ∇ θ logπ θ (x).(5) 3. Formal Aspects In this section, we describe thef-divergence family, and introduce a generic technique,f-DPG, for minimizing the f-divergence between a target distributionpand a model π θ . We then describe the application off-DPG to aligning language models with human preferences. 3.1.f-divergences Consider a convex functionf: (0,∞)→Rwithf(1) = 0. Letf(0) . = lim t→0 f(t)andf ′ (∞) . = lim t→0 tf( 1 t ) . 1 Let p 1 ,p 2 be two distributions over a discrete setX. Thef- divergence betweenp 1 andp 2 can be defined as D f (p 1 ||p 2 ) . =E x∼p 2 f p 1 (x) p 2 (x) +f ′ (∞)p 1 (p 2 = 0) (6) wherep 1 (p 2 = 0)is thep 1 -mass of the setx∈ X: p 2 (x) = 0(Polyanskiy, 2019; Liese & Vajda, 2006). The functionfis called a generator ofD f . By convention, if p 1 (p 2 = 0) = 0, the last term of Eq.(6)is set to0regardless of the value off ′ (∞) (which can be infinite). 2 It can be 1 The limits are well-defined and take values in(−∞,∞]. The convention forf ′ (∞)is motivated by the fact that lim t→∞ f ′ (t) = lim t→0 tf( 1 t )(Hiriart-Urruty & Lemar ́ echal, 2013). 2 Based on the commonly made assumption that the support of p 1 is dominated by the support ofp 2 (Supp(p 1 )⊂Supp(p 2 )), Eq. (6) simplifies toD f (p 1 ||p 2 ) =E x∼p 2 h f p 1 (x) p 2 (x) i . 3 Aligning Language Models with Preferences throughf-divergence Minimization shown thatD f (p 1 ||p 2 )≥0for anyp 1 andp 2 , with equality ifp 1 =p 2 ; conversely, ifD f (p 1 ||p 2 ) = 0andfis strictly convex at1, thenp 1 =p 2 . Thef-divergence family includes many important diver- gence measures, in particular KL divergenceKL(p 1 ||p 2 ), reverse KL divergenceKL(p 2 ||p 1 ), Jensen-Shannon di- vergence, and Total Variation distance. We list thesef- divergences and their generators in Tab. 1. For more details about notations and properties off-divergences, see App. A.1 and also Liese & Vajda (2006); Polyanskiy (2019); Sa- son & Verd ́ u (2016); Sason (2018). 3.2. Distributional alignment withf-divergences LetXbe a discrete countable or finite set, in our case a set of texts. Given a target probability distributionp(x) over elementsx∈ X, our goal is to approximatepwith a generative model (aka policy)π θ . On the other hand, the generative modelπ θ is a parametric model, typically an autoregressive neural network, from which we can (i) directly sample and (i) evaluate probabilitiesπ θ (x). We approach this problem by attempting to minimize the f-divergence ofπ θ top: 3 min θ∈Θ D f (π θ ||p),(7) whereθvaries inside the parametric familyΘ. Note that when the familyπ θ ,θ∈Θis “well-specified”, i.e., when ∃θ 0 s.t.p=π θ 0 , the true minimum of Eq(7)is0, attained atθ 0 , whatever divergenceD f is chosen. In contrast, when the family is “mis-specified” i.e. does not includep, the distributionπ θ with minimal divergence can be strongly dependent on the chosen divergenceD f . Eq.(7)might be solved approximately using stochastic op- timization with samples drawn from the distributionp, as the definition ofD f (π θ ||p)involves taking the expectation with respect top. However, it is often not possible to sample directly fromp, while it is possible to sample fromπ θ . Our optimization technique is then based on the following core result, which we prove in App. A.3. Theorem 1.Letpandπ θ be distributions over a discrete set Xsuch that at least one of the following conditions holds: (i)∀θ∈Θ,Supp(p)⊂Supp(π θ ), or (i)Supp(π θ )does not depend onθ. Then: ∇ θ D f (π θ ||p) =E x∼π θ f ′ π θ (x) p(x) ∇ θ logπ θ (x) . (8) 3 We could have chosen to domin θ∈Θ D f (p||π θ ). However the perspective transformf ∗ (t) . =t f( 1 t )allows interchangeability of arguments:D f (π θ ||p) =D f ∗ (p||π θ ) , making either form possible. The form in Eq.(7)permits a simpler statement of our main theorem. See App. A.1, A.3 for details. Note that it may happen in Eq 8 thatp(x) = 0and π θ (x)>0, hence π θ (x) p(x) =∞ , in which case the expres- sionf ′ π θ (x) p(x) should be understood as denoting the value f ′ (∞)as defined earlier. 4 In the context of LMs, our domain of application, we will use Thm. 1 in situations whereπ θ , being a standard softmax- based autoregressive model, has full support overX(i.e. Supp(π θ ) =X) for allθ’s, while the support ofpmight be strictly included inXin some experiments (Sec. 4.2, 4.4). It is instructive to consider Thm. 1 in relation to rewards in RL. In the standard policy gradient algorithm (Williams, 1992), to find the model that maximizes the average reward E x∼π θ [r(x)], one computes the gradient of the loss using the formula∇ θ E x∼π θ [r(x)] =E x∼π θ [r(x)∇ θ logπ θ (x)]. The gradient in Eq. 8 is very similar, with a “pseudo-reward” r θ (x) =−f ′ ( π θ (x) p(x) ), one difference being that nowr θ de- pends onθ(see (Korbak et al., 2022b) for related remarks). We refer to the approach in Eq. 8 under the namef-DPG, in reference to the original DPG (Distributional Policy Gradi- ent) approach introduced in (Parshakova et al., 2019), which can be seen as a special case off-DPG (“KL-DPG”) with D f (π θ ||p)set to KL(p||π θ )as discussed in Sec. 3.4. 3.3. Adding a baseline Based on the similarity to policy gradients, we adopt the widely usedbaselinetechnique from RL, as previously studied in Williams (1992); Baxter & Bartlett (2001); Schulman et al. (2016) and in the context of DPG in (Korbak et al., 2022b).This technique involves sub- tracting a constantBfrom the reward term, and does not introduce bias in the estimate of the gradient at a givenθ. In our case, withr θ (x) . =−f ′ ( π θ (x) p(x) ), we can write∇ θ D f (π θ ||p) =E x∼π θ r θ (x)∇ θ logπ θ (x) = E x∼π θ (r θ (x)−B)∇ θ logπ θ (x) , based on the observation thatE x∼π θ ∇ θ logπ θ (x) = 0(see also App. A.6). Fact 1.SubtractingBfromr θ (x)does not introduce bias intof-DPG gradient estimates. Typically,Bis chosen to be the average of the rewards, B . =E x∼π θ [r θ (x)]. In the experiments of Sec. 4, we use the baseline technique whereBis an estimate of the average of pseudo-rewards, unless otherwise specified. 4 The derivativef ′ (t)of any convex functionf(t)is defined almost everywhere, with the possible exception of a countable number of non-differentiable points, at which a subgradient can be used instead (Hiriart-Urruty & Lemar ́ echal, 2013; Rockafellar, 1970). See also App. A.4. 4 Aligning Language Models with Preferences throughf-divergence Minimization D f (π θ ||p)f ′ f ′ π θ (x) p(x) f ′ (∞) Forward KL(KL(p||π θ ))f(t) =−logtf ′ (t) =− 1 t − p(x) π θ (x) 0 Reverse KL(KL(π θ ||p))f(t) =tlogtf ′ (t) = logt+ 1− log p(x) π θ (x) + 1∞ Total Variation(TV(π θ ||p))f(t) = 0.5|1−t|f ′ (t) = ( 0.5fort >1 −0.5fort <1 ( 0.5for π θ (x) p(x) >1 −0.5for π θ (x) p(x) <1 0.5 Jensen-Shannon(JS(π θ ||p))f(t) =tlog 2t t+1 + log 2 t+1 f ′ (t) = log 2t t+1 log 2−log 1 + p(x) π θ (x) log 2 Table 1.Some commonf-divergencesD f (π θ ||p). In the convention of this table, thefshown corresponds to the order of arguments D f (π θ ||p). Thus the forward KL between the targetpand the model,KL(p||π θ ), corresponds toD −logt (π θ ||p), and similarly for the reverse KL,KL(π θ ||p), which corresponds toD tlogt (π θ ||p), etc. Note that for symmetric divergences (TV and JS) the order of arguments is indifferent: TV(π θ ||p) =TV(p||π θ ), JS(π θ ||p) =JS(p||π θ ). 3.4. Recovering Some Existing Methods Various existing methods for aligning LM with preferences can be included in thef-DPG framework. GDCIn GDC, fitting the policyπ θ to the targetp(which is given by either one of Eq. 1 or Eq. 4) is done using DPG (Par- shakova et al., 2019), namely by minimizing theforward KL,KL(p||π θ ). In thef-DPG framework,KL(p||π θ ) = D f (π θ ||p)withf(t) =−logt,f ′ (t) =−1/t, and Thm. 1 leads to the formula: ∇ θ D f (π θ ||p) =E x∼π θ − p(x) π θ (x) ∇ θ logπ θ (x), which is equivalent to Eq. 5. RL with KL penaltiesLet’s rewrite the target distribution of Eq.(3)asp(x) . =p RLKL (x) = 1/Z a(x)e r(x)/β , where Zis a normaliser. ThenKL(π θ ||p) =D f (π θ ||p), with f(t) =tlogtcorresponding toreverse KL, andf ′ (t) = 1 + logt. Thm. 1 implies that: ∇ θ D f (π θ ||p) =E x∼π θ 1 + log π θ (x) Z −1 a(x) exp(r(x)/β) ∇ θ logπ θ (x) =E x∼π θ − r(x) β + log π θ (x) a(x) ∇ θ logπ θ (x), where we have exploited the fact that1 + logZis a constant, henceE x∼π θ (1 + logZ)∇ θ logπ θ (x) = 0 .Up to the constant factorβ, this form re- covers the usual formula for estimating the gradient of the loss defined in Eq.(2):∇ θ J RLKL (θ) = E x∼π θ r(x)−βlog π θ (x) a(x) ∇ θ logπ θ (x). 3.5. EstimatingZ The target distributionpis often defined asp(x)∝P(x), whereP(x)is a non-negative function overX. The distri- butionpcan then be computed asp(x) = 1/Z P(x), where Zis the normalizing constant (partition function) defined by P x∈X P(x). An estimate ofZcan be obtained by impor- tance sampling, using samples from the currentπ θ , based on the identityZ=E π θ P(x) π θ (x) . Each such estimate is unbiased, and by averaging the estimates based on differentπ θ ’s, one can obtain a more precise estimate ofZ, exploitingallthe samples obtained so far. For details about the estimate ofZ, see Algorithm 1 in App. A.3, as well as the ablation study in App. H.3. 3.6. Conditional Target Distributions For a conditional task such as machine translation, summa- rization or dialogue, whereπ θ is defined as a conditional distributionπ θ (x|c), we adapt the conditional generaliza- tion of DPG introduced in Korbak et al. (2022a). Given a distribution over contextsτ(c)and a map from a contextc to a target distributionp c , we have (see App. E for details): Fact 2.f-DPG is generalized to the conditional case by optimizing the loss E c∼τ(c) [∇ θ D f (π θ (·|c)||p c (·))].(9) 4. Experiments We study four instantiations off-DPG, namely KL-DPG, RKL-DPG, TV-DPG and JS-DPG, corresponding to min- imizing the forward KL, reverse KL, Total Variation, and Jensen-Shannon divergences, respectively. We use an ex- ponential moving average baseline with weightα= 0.99 for all, except for KL-DPG, where we use the analytically computed value of the pseudo-reward expectation, which amounts to1(Korbak et al., 2022b). We evaluate them on a diverse array of tasks including imposing sentiment con- straints (Sec. 4.1), lexical constraints (Sec. 4.2), debiasing genders’ prevalence and religious groups’ regard (Sec. 4.3), and context-conditioned tasks, such as enforcing factual consistency in summarization (Sec. 4.4) or compilability of generated code (see App. E.1). Unless specified other- wise, we use a pretrained GPT-2 “small” (Radford et al., 5 Aligning Language Models with Preferences throughf-divergence Minimization 2019) with 117M parameters for the initial model. Yet, we demonstrate in Sec. 4.5 that the observations continue to hold for models of larger size. Implementation details and hyper-parameters are available in App. C. MetricsWe report the following key metrics. We add task-specific metrics if needed. 1.D f (π θ ||p), thef-divergence betweenpandπ θ , with four differentf’s corresponding to forward KL, KL(p||π θ ); reverse KL,KL(π θ ||p); Total Variation, TV(π θ ||p); and Jensen-Shannon,JS(π θ ||p). We use importance sampling to estimate these divergences. 2. KL(π θ ||a), a measure of the divergence from original LMa(Ziegler et al., 2019; Khalifa et al., 2021). 3.Alignment score, measured by momentsE x∼π θ φ(x) of a feature of interestφ(x). 4. Normalized Entropy (Berger et al., 1996), a measure of diversity in probability distribution normalized by number of tokens. 5.Standard deviation of a minibatch’s pseudo-rewards, std(r θ (x)), wherer θ is defined as in Sec. 3.3. 4.1. Alignment with Scalar Preferences TaskWe begin with the task of maximizing a scalar pref- erence with KL penalties, whose target distribution,p RLKL , is defined in Eq. 3. We setr(x) = logφ(x)whereφ(x) is the probability returned by a sentiment classifier fine- tuned from Distil-BERT (HF Canonical Model Maintainers, 2022). This reward function is optimal for modeling a decision-maker which givenkdifferent samplesx 1 ,...,x k , will pickx i with probability proportional toφ(x i )(see Ap- pendix F). We setβ= 0.1, which is in line with the range of values explored by Ziegler et al. (2019). Note that ap- plying RKL-DPG onp RLKL is equivalent to the RL with KL penalties method, as described in Sec. 3.4. However, throughf-DPG we can explore alternative objectives to approximate the same target. ResultsFig.2 shows the evolution of the above- mentioned metrics. Further details are given in Fig. 11 in the Appendix. We observe that whereas RKL-DPG achieves by far the best performance in terms of reverse KL,KL(π θ ||p) (top-right), it fails to minimize all other divergence met- rics. This shows that minimizing one divergence does not necessarily imply that other divergences will follow. No- tably, RKL-DPG yields the highest value of alignment scoreE π θ [φ(x)]at the cost of a significant departure from a. We connect this to the strong influence that low values p(x)have on RKL-DPG, which induces a large pseudo- reward for strongly reducingπ θ (x)on those samples (see Sec 5) and produces the spike at the beginning of training in std(rewards). This can leadπ θ (x)to concentrate on high- probability regions ofp(x), at the cost of diversity, which can also be seen in the low entropy of the generated samples. Interestingly, the three remaining variants of DPG (KL, TV and JS) consistently minimize all four tracked divergences, with JS-DPG performing best overall. In App. D.1, we show additional metrics on generated sentences, which show low diversity but high quality for RKL-DPG, compared to otherf-DPGs, suggesting it cap- tures a subset of the target distribution (“mode collapse”), as commonly observed in other generative models (Huszar, 2015; Che et al., 2017; Mescheder et al., 2018). 01000 epoch 1.0 1.5 2.0 2.5 KL(p|| ) 01000 epoch 0.5 0.6 0.7 TV( ||p) 01000 epoch 0.20 0.25 0.30 0.35 0.40 JS( ||p) 01000 epoch 5 10 15 20 25 30 KL( ||p) 01000 epoch 0.4 0.6 0.8 E[ (x)] 01000 epoch 0 2 4 6 KL( ||a) 01000 epoch 150 155 160 165 Entropy 01000 epoch 0 2 4 6 std(rewards) KL TV JS RKL Figure 2.Comparison off-DPG on sentiment preference. Evalua- tion metrics: fourf-divergencesD f (π θ ||p)(↓better), alignment scoreE π θ [φ(x)](↑better), entropy (↑better), standard deviation of pseudo-rewardstd(r θ (x)). 4.2. Alignment with Lexical Constraints TaskIn this task, we constrain the presence of a specific word in the generated text. Following Khalifa et al. (2021), we formulate this goal as a binary preference on the LM by using a target distributionp GDC bin , whereb(x) = 1iff the target word appears in the sequencex, and using a scalar preference target distributionp RLKL wherer(x)is set in the same way asb(x)above. Note that in the GDC framework, p GDCbin (x) = 0whenb(x) = 0, implying that reverse KL, namelyKL(π θ ||p), becomes infinite, so RKL-DPG cannot be used (nor measured) for that target. We use four words with different occurrence frequency: “amazing”(1·10 −3 ), “restaurant” (6·10 −4 ), “amusing” (6·10 −5 ), and “Wikileaks” (8·10 −6 ). ResultsThe aggregated evolution of the metrics for both GDC and RL with KL penalties framework is presented in Fig. 3 (Fig. 1 shows a simplified view of Fig. 3 (a)). Disag- gregated results for each task are presented on App. G. We see that all variants off-DPG reduce the divergence from the target distribution across all measuredf-divergences. Furthermore, as expected, convergence to the target is con- 6 Aligning Language Models with Preferences throughf-divergence Minimization nected with the success ratio in producing the desired word, E π θ [b(x)], while balancing it with a moderate divergence froma,KL(π θ ||a). This reflects that approaching the opti- mal distributionptranslates into metrics in the downstream task. Strinklingly, the original KL-DPG is outperformed by all other variants off-DPG, even in terms of forward KL. We hypothesize that this is linked to the high variance of the pseudo-rewards in KL-DPG, as visualized in the last panel of Fig. 3 (a) and (b). In Sec. 5, we suggest an interpretation for this. We also observe that RKL-DPG tends to produce distributions with lower normalized entropy. Despite this ef- fect, we found no significant difference in diversity among the generated sentences (see Tab. 4 in App. D.1) (a) lexical constraint withp GDCbin (b) lexical constraint withp RLKL Figure 3.Comparison off-DPG aggregated on four lexical con- straints. Standard deviations are suppressed for clarity. Evalua- tion metrics: fourf-divergencesD f (π θ ||p)(↓better), alignment scoreE π θ [b(x)] (↑better), entropy (↑better), standard deviation of pseudo-rewardstd(r θ (x)). 4.3. Alignment with Distributional Constraints TaskWe now investigate enforcing distributional prefer- ences on the LM. We focus on debiasing the pretrained model on two kinds of preferences, namely genders’ preva- lence (Khalifa et al., 2021) and regard relative to religious groups. The preferences for the genders’ debiasing task are defined asφ 1 (x) = 1iffxcontains more female than male pronouns, with desired moment ̄μ 1 = 0.5andφ 2 (x) = 1iff xcontains at least one of the words in the ‘science’ word list compiled by Dathathri et al. (2020), with desired moment ̄μ 2 = 1. For regard debiasing, we use a single distributional constraint where0< φ(x)<1is a regard score of the sentence when prompted withMuslims, evaluated with a pretrained classifier (Sheng et al., 2019). We set the desired moment ̄μ= 0.568, the regard score observedChris- tians . The initial average regard score givenMuslims is0.385. For the first experiment, we use GPT-2 small as the initial modela, additionally fine-tuned on the WikiBio dataset (Lebret et al., 2016), whereas for the last one we use vanilla GPT-2 small. ResultsWe report the results of both experiments on Fig. 4.For the regard score rebalancing, we con- siderably reduce bias in the regard score for two differ- ent demographic groups, from initial regard score ratio E[φ(x)|Christians] :E[φ(x)|Muslims] = 1 : 0.677 toE[φ(x)|Christians] :E[φ(x)|Muslims] = 1 : 0.801on average. Interestingly, this task showcases a weak- ness of TV-DPG: Because the original distribution is already close to the target, the hard-thresholded pseudo-reward has a large variance (last panel of Fig 4(b)), inducing noisy gra- dient estimates and, consequently, sub-optimal convergence. Concerning the gender debiasing experiments, we can see that all other variants off-DPG outperform the original KL-DPG explored in Khalifa et al. (2021), with RKL-DPG giving the best results and better matching the pointwise constraint although seemingly at the cost of lower diversity as measured by the entropy. 4.4. Alignment with Conditional Constraints TaskWe adopt the conditional task from Korbak et al. (2022a), which aims to constrain the T5 (Raffel et al., 2020) language model to generate more factually faithful sum- maries (Maynez et al., 2020; Nan et al., 2021). Specifi- cally, letNER(·)denote the set of named entities found in a text. Then,b(x,c) = 1iff[NER(x)⊆NER(c)]∧ [|NER(x)|≥4], and0otherwise. Following the authors, we sample source documents from the the CNN/Daily Mail dataset (Nallapati et al., 2016), i.e.τ(c)is a uniform distribu- tion over a given subset of source documents. In addition to the divergences, we evaluate the performance using Rouge (Lin, 2004), a measure of summarization quality in terms of unigram overlap between the source document and ground truth summary (See App. E for additional metrics and more experiments on code generation with compilability prefer- ences). ResultsWe present the evolution of metrics in Fig. 5. The results show thatf-DPG increases the fraction of consistent named entities in summarization, and interestingly, this 7 Aligning Language Models with Preferences throughf-divergence Minimization (a) distributional constraint for gender prevalence 1 2 (b) distributional constraint for regard rebalancing 01000 epoch 0.03 0.04 0.05 0.06 0.07 0.08 KL(p|| ) 01000 epoch 0.10 0.12 0.14 0.16 0.18 TV( ||p) 01000 epoch 0.01 0.02 JS( ||p) 01000 epoch 0.03 0.04 0.05 0.06 0.07 0.08 0.09 KL( ||p) 01000 epoch 0.42 0.44 0.46 0.48 0.50 0.52 E[ (x)] 01000 epoch 0.00 0.02 0.04 0.06 0.08 KL( ||a) 01000 epoch 163.0 163.2 163.4 163.6 163.8 Entropy 01000 epoch 0.1 0.2 0.3 0.4 0.5 0.6 std(rewards) JS KL RKL TV Figure 4.Comparison off-DPG aggregated on distributional con- straints. Evaluation metrics: fourf-divergencesD f (π θ ||p)(↓ better), alignment scoreE π θ [φ(x)](↑better), entropy (↑better), standard deviation of pseudo-rewardstd(r θ (x)). also leads to indirect improvement in the overall quality of generated summaries compared to ground truth, even though ground truth summaries are not used in training. As also observed in Sec. 4.2, JS-DPG leads to better convergence topthan KL-DPG as used in Korbak et al. (2022a). 4.5. Scaling Trends off-DPG We conduct experiments to investigate the effect of model size on our approach using the scalar preference task de- scribed in Sec. 4.1. Specifically, we gradually increase the model size from GPT-2 “small” (117M parameters) to “xl” (1.5B parameters) while tracking two important met- rics: alignment score, which is measured by the expected rewardE π θ [φ(x)], and diversity, which is measured by the entropy. Figure 6 demonstrates that the alignment score steadily improves as the model size increases. However, we observe persistent differences between the divergence objectives for different f-DPGs, leaving the general order between f-DPGs intact with increasing model size (See Fig. 16 in App. G for evolution of metrics through training epochs). The scaling trend of LM alignment, characterized Figure 5.Comparison off-DPG on factual summarization. Eval- uation metrics: 3f-divergencesD f (π θ ||p)(↓better), number of named entities (↑better), Rouge (↑better). by a gradual and predictable increase without sudden shifts in performance, aligns with previous findings in the liter- ature (Bai et al., 2022a). Nonetheless, our study further emphasizes the importance of proper divergence objectives, as increasing model size alone does not necessarily bridge the gap between optimal and suboptimal objectives. The smooth and gradual increase of the alignment score as a function of model size suggests that our findings will gener- alize to even larger LMs. 4.6. Ablation Study This section presents just the key findings of our study. Full results and detailed discussions can be found in App. H. Effect of parameter family capacityAll experiments presented so far correspond to possibly mis-specified target distributions. To understand whether the observed behavior of different variantsf-DPG is affected by this factor, we used pre-trained models with the same architecture asπ θ andp. We found that KL-DPG again lags considerably in terms of divergence, while presenting a high variance of in the pseudo-reward. RKL-DPG shows a significant drop of entropy in the initial phase, but with full capacity of parameter family, the model can recover, and cover the rest of the distribution. Additionally, applying zero-shot the fine-tuned LMs to a summarization task, following Radford et al. (2019), we found that the they recover to a large extent the quality of the target distribution. Effect of training schemeWe examined different train- ing schemes for the lexical constraint on “amazing” from Sec. 4.2. We saw that the use of a baseline technique im- proves the performance of thef-DPG method, with RKL- DPG showing the greatest benefit. Additionally, we found that even though a large batch size is effective at reduc- ing the variance of KL-DPG, we still observe KL-DPG to 8 Aligning Language Models with Preferences throughf-divergence Minimization 10 8 10 9 Number of Parameters 0.6 0.7 0.8 0.9 Alignment Score 10 8 10 9 Number of Parameters 130 140 150 160 Entropy KL TV JS RKL Figure 6.The scaling trend off-DPG on sentiment preference. Thex-axis denotes number of parameters of the LMπ θ and they-axis denotes the alignment score and diversity measured by the expected rewardE π θ [φ(x)]and by entropy, respectively. perform comparatively worse than other divergences. Fi- nally, we observe that our importance sampling estimates converged to the true value ofZ. 5. Discussion and Conclusion KLKL KL KL TV TV TV TV JS JS JS JS RKLRKL RKL Figure 7.Pareto frontier off-DPG for different alignment tasks; sentiment preference (Fig. 2), lexical constraints (Fig. 3(a), (b)), and distributional constraint for gender prevalence (Fig. 4(a)) A plausible hypothesis would have been that each variant off-DPG is comparatively better at least in terms of the f-divergence objective being optimized. Surprisingly, we found that, save for a few exceptions (Sec. 4.1), for a given target there is one or a few variants that are the best across all measured divergences. Furthermore, we observed that divergence measures can have a significant impact on the performance of the model depending on the target distribu- tion. Fig. 7 summarizes the Pareto frontier of the alignment- diversity trade-off of thef-DPG method. The results demon- strate that RKL-DPG and KL-DPG consistently represent two contrasting extremes: RKL-DPG shows high alignment but limited diversity, whereas KL-DPG exhibits low align- ment but high diversity. JS-DPG shows a balanced trade-off between alignment and diversity and consistently appeared Figure 8.Pseudo-rewards for variousf-divergences. Thex-axis denotes p(x) π θ (x) and they-axis denotes the pseudo-reward. The dotted line denotes the point wherep(x) =π θ (x). on the Pareto frontier across all experiments we conducted. Fig. 8 illustrates the differences between pseudo-rewards for distinctf-divergences, giving a plausible explanation for the observed differences. The forward KL loss aims to ensure coverage of the subset wherep(x)>0, giving a large pseudo-reward for samples withp(x)>>π(x). However, the optimization can be sensitive to sampling noise in the finite sample approximation (see, e.g., Sec. 4.2). Conversely, the reverse KL loss results in extreme negative rewards for samples withp(x)<<π θ (x), leadingπ θ to avoid such regions and resulting in distributional collapse (Sec. 4.1). Total Variation loss is robust to outliers thanks to its hard- thresholded pseudo-reward, however it can lead to high variance behavior whenπ θ ≈p(Sec. 4.3). On the other hand, the Jensen-Shannon loss gives smooth and robust rewards in both directions and preventsπ θ from heavily relying on a single direction, making it a reasonable default choice as confirmed by our experiments. To conclude, we propose a flexible framework for approxi- mating a target distribution by minimizing anyf-divergence, unifying earlier approaches for aligning language models. Our results on a diverse array of tasks show that minimizing well-chosenf-divergences leads to significant gains over previous work.The fact that increasing the model size improves the alignment score but does not inherently bridge the gap between objectives underscores the importance of selecting appropriate divergence objectives. 9 Aligning Language Models with Preferences throughf-divergence Minimization References Arjovsky, M., Chintala, S., and Bottou, L. Wasserstein generative adversarial networks. In Precup, D. and Teh, Y. W. (eds.),Proc. of ICML, volume 70 ofProceedings of Machine Learning Research, p. 214–223. PMLR, 2017. URLhttp://proceedings.mlr.press/ v70/arjovsky17a.html. Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. A general language assistant as a laboratory for alignment.arXiv preprint arXiv:2112.00861, 2021. URL https://arxiv.org/abs/2112.00861. Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A. C., and Bengio, Y. An actor- critic algorithm for sequence prediction. InProc. of ICLR. OpenReview.net, 2017. URLhttps://openreview. net/forum?id=SJDaqqveg. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das- Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernan- dez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with rein- forcement learning from human feedback, 2022a. URL https://arxiv.org/abs/2204.05862. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das- Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. Training a helpful and harmless assistant with rein- forcement learning from human feedback.ArXiv preprint, abs/2204.05862, 2022b. URLhttps://arxiv.org/ abs/2204.05862. Baxter, J. and Bartlett, P. L. Infinite-horizon policy-gradient estimation.Journal of Artificial Intelligence Research, 15:319–350, 2001. Berger, A. L., Della Pietra, S. A., and Della Pietra, V. J. A maximum entropy approach to natural language process- ing.Computational Linguistics, 22(1):39–71, 1996. URL https://aclanthology.org/J96-1002. Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, 2021. URLhttps: //doi.org/10.5281/zenodo.5297715. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.),Proc. of NeurIPS, 2020.URLhttps://proceedings. neurips.c/paper/2020/hash/ 1457c0d6bfcb4967418bfb8ac142f64a- Abstract.html. Cao, Y., Sotnikova, A., Daum ́ e I, H., Rudinger, R., and Zou, L. Theory-grounded measurement of U.S. social stereotypes in English language models.In Proc. of NAACL-HLT, p. 1276–1295, Seattle, United States, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.92. URLhttps: //aclanthology.org/2022.naacl-main.92. Che, T., Li, Y., Jacob, A. P., Bengio, Y., and Li, W. Mode regularized generative adversarial networks. In Proc. of ICLR. OpenReview.net, 2017. URLhttps: //openreview.net/forum?id=HJKkY35le. Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. ArXiv preprint, abs/2210.11416, 2022. URLhttps: //arxiv.org/abs/2210.11416. Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and play language models: A simple approach to controlled text generation. InProc. of ICLR. OpenReview.net, 2020. URLhttps://openreview.net/forum? id=H1edEyBKDS. Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. Language model cascades. ArXiv preprint, abs/2207.10342, 2022. URLhttps: //arxiv.org/abs/2207.10342. Eikema, B., Kruszewski, G., Dance, C. R., Elsahar, H., and Dymetman, M. An approximate sampler for energy- based models with divergence diagnostics.Transactions of Machine Learning Research, 2022. URLhttps: //openreview.net/forum?id=VW4IrC0n0M. Fan, A., Lewis, M., and Dauphin, Y. Hierarchical neu- ral story generation. InProc. of ACL, p. 889–898, Melbourne, Australia, 2018. Association for Computa- tional Linguistics. doi: 10.18653/v1/P18-1082. URL https://aclanthology.org/P18-1082. Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A. RealToxicityPrompts: Evaluating neural toxic de- generation in language models. InFindings of EMNLP, 10 Aligning Language Models with Preferences throughf-divergence Minimization p. 3356–3369, Online, 2020. Association for Compu- tational Linguistics. doi: 10.18653/v1/2020.findings- emnlp.301. URLhttps://aclanthology.org/ 2020.findings-emnlp.301. Ghasemipour, S. K. S., Zemel, R., and Gu, S. A di- vergence minimization perspective on imitation learn- ing methods.In Kaelbling, L. P., Kragic, D., and Sugiura, K. (eds.),Proceedings of the Conference on Robot Learning, volume 100 ofProceedings of Machine Learning Research, p. 1259–1277. PMLR, 2020. URLhttps://proceedings.mlr.press/ v100/ghasemipour20a.html. Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chad- wick, M., Thacker, P., Campbell-Gillingham, L., Ue- sato, J., Huang, P., Comanescu, R., Yang, F., See, A., Dathathri, S., Greig, R., Chen, C., Fritz, D., Elias, J. S., Green, R., Mokr ́ a, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hass- abis, D., Kavukcuoglu, K., Hendricks, L. A., and Irv- ing, G. Improving alignment of dialogue agents via targeted human judgements.CoRR, abs/2209.14375, 2022. doi: 10.48550/arXiv.2209.14375. URLhttps: //doi.org/10.48550/arXiv.2209.14375. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks.Commun. ACM, 63(11): 139–144, 2020. ISSN 0001-0782. doi: 10.1145/3422622. URLhttps://doi.org/10.1145/3422622. Goyal, K., Dyer, C., and Berg-Kirkpatrick, T. Exposing the implicit energy networks behind masked language mod- els via metropolis–hastings. InProc. of ICLR. OpenRe- view.net, 2022. URLhttps://openreview.net/ forum?id=6PvWo1kEvlT. HF Canonical Model Maintainers.distilbert-base- uncased-finetuned-sst-2-english (revision bfdd146), 2022. URLhttps://huggingface.co/distilbert- base-uncased-finetuned-sst-2-english. Hiriart-Urruty, J.-B. and Lemar ́ echal, C.Convex analysis and minimization algorithms I: Fundamentals, volume 305. Springer science & business media, 2013. Huszar, F. How (not) to train your generative model: Sched- uled sampling, likelihood, adversary?ArXiv preprint, abs/1511.05101, 2015. URLhttps://arxiv.org/ abs/1511.05101. Jaques, N., Gu, S., Bahdanau, D., Hern ́ andez-Lobato, J. M., Turner, R. E., and Eck, D.Sequence tutor: Conservative fine-tuning of sequence generation mod- els with kl-control.In Precup, D. and Teh, Y. W. (eds.),Proc. of ICML, volume 70 ofProceedings of Machine Learning Research, p. 1645–1654. PMLR, 2017. URLhttp://proceedings.mlr.press/ v70/jaques17a.html. Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, ` A., Jones, N., Gu, S., and Picard, R. W. Way off-policy batch deep reinforcement learning of implicit human preferences in dialog.ArXiv preprint, abs/1907.00456, 2019. URLhttps://arxiv.org/ abs/1907.00456. Kappen, H. J., G ́ omez, V., and Opper, M. Optimal control as a graphical model inference problem.Machine learning, 87(2):159–182, 2012. Kappen, H. J., G ́ omez, V., and Opper, M. Optimal con- trol as a graphical model inference problem. In Borrajo, D., Kambhampati, S., Oddi, A., and Fratini, S. (eds.), Proceedings of the Twenty-Third International Confer- ence on Automated Planning and Scheduling, ICAPS 2013, Rome, Italy, June 10-14, 2013. AAAI, 2013. URLhttp://w.aaai.org/ocs/index.php/ ICAPS/ICAPS13/paper/view/6012. Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S. S. Imitation learning as f-divergence minimization. In LaValle, S. M., Lin, M., Ojala, T., Shell, D. A., and Yu, J. (eds.),Algorithmic Foundations of Robotics XIV, Proceedings of the Fourteenth Work- shop on the Algorithmic Foundations of Robotics, WAFR 2021, Oulu, Finland, June 21-23, 2021, volume 17 of Springer Proceedings in Advanced Robotics, p. 313–329. Springer, 2021. doi: 10.1007/978-3-030-66723-8\19. URLhttps://doi.org/10.1007/978-3-030- 66723-8_19. Khalifa, M., Elsahar, H., and Dymetman, M.A dis- tributional approach to controlled text generation. In Proc. of ICLR. OpenReview.net, 2021. URLhttps: //openreview.net/forum?id=jWkw45-9AbL. Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.),Proc. of ICLR, 2015. URLhttp://arxiv.org/abs/1412. 6980. Korbak, T., Elsahar, H., Kruszewski, G., and Dymet- man, M. Controlling conditional language models with- out catastrophic forgetting. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.),Proc. of ICML, volume 162 ofProceed- ings of Machine Learning Research, p. 11499–11528. PMLR, 2022a. URLhttps://proceedings.mlr. press/v162/korbak22a.html. 11 Aligning Language Models with Preferences throughf-divergence Minimization Korbak, T., Elsahar, H., Kruszewski, G., and Dymetman, M. On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic for- getting. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.),Proc. of NeurIPS, 2022b. URLhttps: //openreview.net/forum?id=XvI6h-s4un. Korbak, T., Perez, E., and Buckley, C. L.RL with KL penalties is better viewed as bayesian inference. CoRR, abs/2205.11275, 2022c. doi: 10.48550/arXiv. 2205.11275. URLhttps://doi.org/10.48550/ arXiv.2205.11275. Lebret, R., Grangier, D., and Auli, M. Neural text gen- eration from structured data with application to the bi- ography domain. InProc. of EMNLP, p. 1203–1213, Austin, Texas, 2016. Association for Computational Lin- guistics. doi: 10.18653/v1/D16-1128. URLhttps: //aclanthology.org/D16-1128. LeCun, Y., Chopra, S., Hadsell, R., Ranzato, M., and Huang, F. A tutorial on energy-based learning.Predicting struc- tured data, 1(0), 2006. Levine, S. Reinforcement learning and control as proba- bilistic inference: Tutorial and review.ArXiv preprint, abs/1805.00909, 2018. URLhttps://arxiv.org/ abs/1805.00909. Li, J., Galley, M., Brockett, C., Gao, J., and Dolan, B. A diversity-promoting objective function for neural conver- sation models. InProc. of NAACL-HLT, p. 110–119, San Diego, California, 2016a. Association for Compu- tational Linguistics. doi: 10.18653/v1/N16-1014. URL https://aclanthology.org/N16-1014. Li, J., Monroe, W., Ritter, A., Jurafsky, D., Galley, M., and Gao, J. Deep reinforcement learning for dialogue generation. InProc. of EMNLP, p. 1192–1202, Austin, Texas, 2016b. Association for Computational Linguis- tics. doi: 10.18653/v1/D16-1127. URLhttps:// aclanthology.org/D16-1127. Liese, F. and Vajda, I. On divergences and informations in statistics and information theory.IEEE Transactions on Information Theory, 52(10):4394–4412, 2006. Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. InText Summarization Branches Out, p. 74–81, Barcelona, Spain, 2004. Association for Compu- tational Linguistics. URLhttps://aclanthology. org/W04-1013. Lin, S., Hilton, J., and Evans, O. TruthfulQA: Measur- ing how models mimic human falsehoods. InProc. of ACL, p. 3214–3252, Dublin, Ireland, 2022. Associa- tion for Computational Linguistics. doi: 10.18653/v1/ 2022.acl-long.229. URLhttps://aclanthology. org/2022.acl-long.229. Liu, C.-W., Lowe, R., Serban, I., Noseworthy, M., Charlin, L., and Pineau, J. How NOT to evaluate your dialogue sys- tem: An empirical study of unsupervised evaluation met- rics for dialogue response generation. InProc. of EMNLP, p. 2122–2132, Austin, Texas, 2016. Association for Computational Linguistics. doi: 10.18653/v1/D16-1230. URLhttps://aclanthology.org/D16-1230. Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. Learning word vectors for sentiment anal- ysis. InProc. of ACL, p. 142–150, Portland, Oregon, USA, 2011. Association for Computational Linguistics. URLhttps://aclanthology.org/P11-1015. Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. On faithfulness and factuality in abstractive sum- marization.InProc. of ACL, p. 1906–1919, On- line, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.173. URLhttps: //aclanthology.org/2020.acl-main.173. Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell- Gillingham, L., Irving, G., and McAleese, N. Teach- ing language models to support answers with verified quotes, 2022.URLhttps://arxiv.org/abs/ 2203.11147. Mescheder, L. M., Geiger, A., and Nowozin, S. Which training methods for gans do actually converge? In Dy, J. G. and Krause, A. (eds.),Proc. of ICML, volume 80 ofProceedings of Machine Learning Research, p. 3478– 3487. PMLR, 2018. URLhttp://proceedings. mlr.press/v80/mescheder18a.html. Miao, N., Zhou, H., Mou, L., Yan, R., and Li, L. CGMH: constrained sentence generation by metropolis-hastings sampling. InThe Thirty-Third AAAI Conference on Ar- tificial Intelligence, AAAI 2019, The Thirty-First Inno- vative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Hon- olulu, Hawaii, USA, January 27 - February 1, 2019, p. 6834–6842. AAAI Press, 2019.doi: 10.1609/ aaai.v33i01.33016834. URLhttps://doi.org/10. 1609/aaai.v33i01.33016834. Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. A. Playing atari with deep reinforcement learning.CoRR, abs/1312.5602, 2013. URLhttp://arxiv.org/ abs/1312.5602. 12 Aligning Language Models with Preferences throughf-divergence Minimization Nallapati, R., Zhou, B., dos Santos, C., Gulc ̧ehre,C ̧., and Xiang, B. Abstractive text summarization using sequence-to-sequence RNNs and beyond. InProceed- ings of the 20th SIGNLL Conference on Computational Natural Language Learning, p. 280–290, Berlin, Ger- many, 2016. Association for Computational Linguis- tics. doi: 10.18653/v1/K16-1028. URLhttps:// aclanthology.org/K16-1028. Nan, F., Nallapati, R., Wang, Z., Nogueira dos Santos, C., Zhu, H., Zhang, D., McKeown, K., and Xiang, B. Entity-level factual consistency of abstractive text sum- marization. InProc. of EACL, p. 2727–2733, On- line, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.235. URLhttps: //aclanthology.org/2021.eacl-main.235. Ngo, H., Raterink, C., Ara ́ ujo, J. G., Zhang, I., Chen, C., Morisot, A., and Frosst, N. Mitigating harm in language models with conditional-likelihood filtration. ArXiv preprint, abs/2108.07790, 2021. URLhttps: //arxiv.org/abs/2108.07790. Norouzi, M., Bengio, S., Chen, Z., Jaitly, N., Schuster, M., Wu, Y., and Schuurmans, D. Reward augmented maximum likelihood for neural structured prediction. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.),Proc. of NeurIPS, p. 1723– 1731,2016.URLhttps://proceedings. neurips.c/paper/2016/hash/ 2f885d0fbe2e131bfc9d98363e55d1d4- Abstract.html. Nowozin, S., Cseke, B., and Tomioka, R. f-gan: Training generative neural samplers using variational divergence minimization. In Lee, D. D., Sugiyama, M., von Luxburg, U., Guyon, I., and Garnett, R. (eds.),Proc. of NeurIPS, p. 271–279, 2016. URLhttps://proceedings. neurips.c/paper/2016/hash/ cedebb6e872f539bef8c3f919874e9d7- Abstract.html. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. Training language models to follow in- structions with human feedback. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.),Proc. of NeurIPS, 2022. URLhttps://openreview.net/forum? id=TG8KACxEON. Parshakova, T., Andreoli, J.-M., and Dymetman, M. Distri- butional reinforcement learning for energy-based sequen- tial models.ArXiv preprint, abs/1912.08517, 2019. URL https://arxiv.org/abs/1912.08517. Pasunuru, R. and Bansal, M. Reinforced video caption- ing with entailment rewards. InProc. of EMNLP, p. 979–985, Copenhagen, Denmark, 2017. Association for Computational Linguistics. doi: 10.18653/v1/D17-1103. URLhttps://aclanthology.org/D17-1103. Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., K ̈ opf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S. Pytorch: An imperative style, high-performance deep learning library. In Wallach, H. M., Larochelle, H., Beygelzimer, A., d’Alch ́ e-Buc, F., Fox, E. B., and Garnett, R. (eds.),Proc. of NeurIPS, p. 8024–8035, 2019. URLhttps://proceedings. neurips.c/paper/2019/hash/ bdbca288fee7f92f2bfa9f7012727740- Abstract.html. Paulus, R., Xiong, C., and Socher, R. A deep reinforced model for abstractive summarization. InProc. of ICLR. OpenReview.net, 2018. URLhttps://openreview. net/forum?id=HkAClQgA-. Peters, J. and Schaal, S.Reinforcement learning by reward-weighted regression for operational space con- trol.In Ghahramani, Z. (ed.),Proc. of ICML, vol- ume 227 ofACM International Conference Proceed- ing Series, p. 745–750. ACM, 2007. doi: 10.1145/ 1273496.1273590.URLhttps://doi.org/10. 1145/1273496.1273590. Polyanskiy,Y.f-divergences,2019.URL https://people.lids.mit.edu/yp/ homepage/data/LN_fdiv.pdf. Qin, L., Welleck, S., Khashabi, D., and Choi, Y. Cold de- coding: Energy-based constrained text generation with langevin dynamics.ArXiv preprint, abs/2202.11705, 2022.URLhttps://arxiv.org/abs/2202. 11705. Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al. Exploring the limits of transfer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(140):1–67, 2020. Ranzato, M., Chopra, S., Auli, M., and Zaremba, W. Se- quence level training with recurrent neural networks. In Bengio, Y. and LeCun, Y. (eds.),Proc. of ICLR, 2016. URLhttp://arxiv.org/abs/1511.06732. 13 Aligning Language Models with Preferences throughf-divergence Minimization Raychev, V., Bielik, P., and Vechev, M. Probabilistic model for code with decision trees.ACM SIGPLAN Notices, 51 (10):731–747, 2016. Rockafellar, R. T.Convex analysis, volume 18. Princeton university press, 1970. Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Smith, E. M., Boureau, Y.-L., and Weston, J. Recipes for building an open- domain chatbot. InProc. of EACL, p. 300–325, On- line, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.24.URLhttps: //aclanthology.org/2021.eacl-main.24. Sason, I. On f-divergences: Integral representations, local behavior, and inequalities.Entropy, 20(5):383, 2018. Sason, I. and Verd ́ u, S.f-divergence inequalities.IEEE Transactions on Information Theory, 62(11):5973–6006, 2016. Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. Self-critiquing models for assisting human evaluators.ArXiv preprint, abs/2206.05802, 2022. URL https://arxiv.org/abs/2206.05802. Scheurer, J., Campos, J. A., Chan, J. S., Chen, A., Cho, K., and Perez, E. Training language models with natural lan- guage feedback.ArXiv preprint, abs/2204.14146, 2022. URLhttps://arxiv.org/abs/2204.14146. Schulman, J., Moritz, P., Levine, S., Jordan, M. I., and Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Bengio, Y. and LeCun, Y. (eds.),Proc. of ICLR, 2016. URLhttp: //arxiv.org/abs/1506.02438. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. Proximal policy optimization algorithms. ArXiv preprint, abs/1707.06347, 2017. URLhttps: //arxiv.org/abs/1707.06347. Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N. The woman worked as a babysitter: On biases in lan- guage generation. InProc. of EMNLP, p. 3407–3412, Hong Kong, China, 2019. Association for Computa- tional Linguistics. doi: 10.18653/v1/D19-1339. URL https://aclanthology.org/D19-1339. Solaiman, I. and Dennison, C. Process for adapting language models to society (palms) with values-targeted datasets. Proc. of NeurIPS, 34:5861–5873, 2021. Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Ko- curek, A. W., Safaya, A., Tazarv, A., Xiang, A., Par- rish, A., Nie, A., Hussain, A., Askell, A., Dsouza, A., Rahane, A., Iyer, A. S., Andreassen, A., Santilli, A., Stuhlm ̈ uller, A., Dai, A. M., La, A., Lampinen, A. K., Zou, A., Jiang, A., Chen, A., Vuong, A., Gupta, A., Got- tardi, A., Norelli, A., Venkatesh, A., Gholamidavoodi, A., Tabassum, A., Menezes, A., Kirubarajan, A., Mul- lokandov, A., Sabharwal, A., Herrick, A., Efrat, A., Erdem, A., Karakas, A., and et al. Beyond the imi- tation game: Quantifying and extrapolating the capa- bilities of language models.CoRR, abs/2206.04615, 2022. doi: 10.48550/arXiv.2206.04615. URLhttps: //doi.org/10.48550/arXiv.2206.04615. Stiennon, N., Ouyang, L., Wu, J., Ziegler, D. M., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. Learning to summarize from human feedback. InProc. of NeurIPS, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546. Tambwekar, P., Dhuliawala, M., Martin, L. J., Mehta, A., Harrison, B., and Riedl, M. O. Controllable neural story plot generation via reward shaping. In Kraus, S. (ed.), Proc. of IJCAI, p. 5982–5988. ijcai.org, 2019. doi: 10. 24963/ijcai.2019/829. URLhttps://doi.org/10. 24963/ijcai.2019/829. Theis, L., van den Oord, A., and Bethge, M. A note on the evaluation of generative models. In Bengio, Y. and LeCun, Y. (eds.),Proc. of ICLR, 2016. URLhttp: //arxiv.org/abs/1511.01844. Thoppilan, R., Freitas, D. D., Hall, J., Shazeer, N., Kul- shreshtha, A., Cheng, H., Jin, A., Bos, T., Baker, L., Du, Y., Li, Y., Lee, H., Zheng, H. S., Ghafouri, A., Mene- gali, M., Huang, Y., Krikun, M., Lepikhin, D., Qin, J., Chen, D., Xu, Y., Chen, Z., Roberts, A., Bosma, M., Zhou, Y., Chang, C., Krivokon, I., Rusch, W., Pick- ett, M., Meier-Hellstern, K. S., Morris, M. R., Doshi, T., Santos, R. D., Duke, T., Soraker, J., Zevenbergen, B., Prabhakaran, V., Diaz, M., Hutchinson, B., Olson, K., Molina, A., Hoffman-John, E., Lee, J., Aroyo, L., Rajakumar, R., Butryna, A., Lamm, M., Kuzmina, V., Fenton, J., Cohen, A., Bernstein, R., Kurzweil, R., Aguera-Arcas, B., Cui, C., Croak, M., Chi, E. H., and Le, Q. Lamda: Language models for dialog applica- tions.ArXiv preprint, abs/2201.08239, 2022. URL https://arxiv.org/abs/2201.08239. Todorov, E.Linearly-solvable markov decision prob- lems. In Sch ̈ olkopf, B., Platt, J. C., and Hofmann, T. (eds.),Proc. of NeurIPS, p. 1369–1376. MIT Press, 2006a.URLhttps://proceedings. neurips.c/paper/2006/hash/ 14 Aligning Language Models with Preferences throughf-divergence Minimization d806ca13ca3449af72a1ea5aedbed26a- Abstract.html. Todorov, E.Linearly-solvable markov decision prob- lems. In Sch ̈ olkopf, B., Platt, J. C., and Hofmann, T. (eds.),Proc. of NeurIPS, p. 1369–1376. MIT Press, 2006b.URLhttps://proceedings. neurips.c/paper/2006/hash/ d806ca13ca3449af72a1ea5aedbed26a- Abstract.html. Wang, D., Liu, H., and Liu, Q. Variational inference with tail-adaptive f-divergence. In Bengio, S., Wallach, H. M., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.),Proc. of NeurIPS, p. 5742– 5752,2018.URLhttps://proceedings. neurips.c/paper/2018/hash/ 1cd138d0499a68f4b72bee04bbec2d7- Abstract.html. Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W. S., Legassick, S., Irving, G., and Gabriel, I. Ethical and social risks of harm from language models. ArXiv preprint, abs/2112.04359, 2021. URLhttps: //arxiv.org/abs/2112.04359. Welbl, J., Glaese, A., Uesato, J., Dathathri, S., Mellor, J., Hendricks, L. A., Anderson, K., Kohli, P., Coppin, B., and Huang, P.-S. Challenges in detoxifying language models. InFindings of EMNLP, p. 2447–2469, Punta Cana, Dominican Republic, 2021. Association for Com- putational Linguistics. doi: 10.18653/v1/2021.findings- emnlp.210. URLhttps://aclanthology.org/ 2021.findings-emnlp.210. Williams, R. J. Simple statistical gradient-following algo- rithms for connectionist reinforcement learning.Machine learning, 8(3):229–256, 1992. Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtow- icz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Le Scao, T., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. Transform- ers: State-of-the-art natural language processing. In Proc. of EMNLP, p. 38–45, Online, 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020. emnlp-demos.6.URLhttps://aclanthology. org/2020.emnlp-demos.6. Xu, J., Ju, D., Li, M., Boureau, Y.-L., Weston, J., and Di- nan, E. Bot-adversarial dialogue for safe conversational agents. InProc. of NAACL-HLT, p. 2950–2968, Online, 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.235.URLhttps:// aclanthology.org/2021.naacl-main.235. Zelikman, E., Wu, Y., Mu, J., and Goodman, N. STar: Bootstrapping reasoning with reasoning. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.),Proc. of NeurIPS, 2022. URLhttps://openreview.net/ forum?id=_3ELRdg2sgI. Zhao, J., Khashabi, D., Khot, T., Sabharwal, A., and Chang, K.-W. Ethical-advice taker: Do language models under- stand natural language interventions? InFindings of ACL, p. 4158–4164, Online, 2021. Association for Computa- tional Linguistics. doi: 10.18653/v1/2021.findings-acl. 364. URLhttps://aclanthology.org/2021. findings-acl.364. Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y. Texygen: A benchmarking platform for text generation models. In Collins-Thompson, K., Mei, Q., Davison, B. D., Liu, Y., and Yilmaz, E. (eds.), Proc. of SIGIR, p. 1097–1100. ACM, 2018. doi: 10. 1145/3209978.3210080. URLhttps://doi.org/ 10.1145/3209978.3210080. Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G. Fine-tuning language models from human preferences.ArXiv preprint, abs/1909.08593, 2019. URLhttps://arxiv.org/ abs/1909.08593. Ziegler, D. M., Nix, S., Chan, L., Bauman, T., Schmidt- Nielsen, P., Lin, T., Scherlis, A., Nabeshima, N., Weinstein-Raun, B., de Haas, D., Shlegeris, B., and Thomas, N. Adversarial training for high-stakes reliabil- ity.CoRR, abs/2205.01663, 2022. doi: 10.48550/arXiv. 2205.01663. URLhttps://doi.org/10.48550/ arXiv.2205.01663. 15 Aligning Language Models with Preferences throughf-divergence Minimization A. Complements on Formal Aspects and Proofs A.1. Equivalent definitions forf-divergences The definition off-divergences of Eq. 6 is equivalent to a second definition, in a more “symmetrical” format, following (Liese & Vajda, 2006), which will help in some derivations, in particular in the proof of Theorem 1. Definition(f-divergence: “symmetrical” format).Thef-divergenceD f (p||q), wherepandqare distributions over a discrete setXcan be defined as D f (p||q) . = X x:p(x)>0, q(x)>0 q(x)f( p(x) q(x) )+f(0)q(p= 0)+f ∗ (0)p(q= 0),(10) where the generator functionf: (0,∞)→Ris a convex function satisfyingf(1) = 0. We denote byq(p= 0)theq-mass of the setx:p(x) = 0, i.e.q(p= 0) = P x:p(x)=0 q(x)and similarly forp(q= 0). In this definition, the functionf ∗ (t)is the so-calledperspective transformoffdefined byf ∗ (t) =t f( 1 t ) . It can be shown to be also a convex functionf ∗ : (0,∞)→Rwithf ∗ (1) = 0andf ∗ =f. As we mentioned in the main text, we also have the following important “swapping” property:D f (p,q) =D f ∗ (q,p). Following Liese & Vajda (2006); Polyanskiy (2019), we use the conventions: f(0) . = lim t→0 f(t), f ∗ (0) = lim t→0 f ∗ (t) = lim t→0 t f( 1 t ),(11) 0f(0) . = 0,0f ∗ (0) . = 0,including whenf(0) =∞andf ∗ (0) =∞,(12) f ′ (∞) . =f ∗ (0) = lim t→0 t f( 1 t ).(13) For the existence of the limits in these equations, wheref(0)andf ∗ (0)can take values inR∪∞, as well as for the motivation for definingf ′ (∞) . = lim t→0 t f( 1 t ), one may refer to (Liese & Vajda, 2006) and (Hiriart-Urruty & Lemar ́ echal, 2013, §2.3). Equivalence of definitions 6 and 10In order to prove this equivalence, after noting thatf ′ (∞) =f ∗ (0), it remains to show thatE x∼q f p(x) q(x) is equal to P x:p(x)>0, q(x)>0 q(x)f( p(x) q(x) ) +f(0)q(p= 0). We have: E x∼q f( p(x) q(x) ) = X x:q(x)>0 q(x)f( p(x) q(x) ) = X x:q(x)>0, p(x)>0 q(x)f( p(x) q(x) )+ X x:q(x)>0, p(x)=0 q(x)f(0) = X x:q(x)>0, p(x)>0 q(x)f( p(x) q(x) )+f(0)q(p= 0), which concludes the proof. A.2. Illustrations of a fewf-divergences Let’s now see how the notion off-divergence can be applied to a few common cases. Forward and reverse KLBy the standard definition for KL divergence, we have, forKL(p||π), the “forward KL” from a modelπto a targetp: KL(p||π) = ( E x∼p log p(x) π(x) if Supp(p)⊂Supp(π), ∞,otherwise. (14) 16 Aligning Language Models with Preferences throughf-divergence Minimization If we takef(t) =−logt, as in Table 1, then we havef(0) =∞. On the other hand we see thatf ∗ (t) =tlogtand f ∗ (0) = 0. We can then write, using (10): D f (π||p) = X x:π(x)>0, p(x)>0 −p(x) log( π(x) p(x) )+∞p(π= 0) + 0π(p= 0) = X x:π(x)>0, p(x)>0 p(x) log( p(x) π(x) )+∞p(π= 0), where∞p(π= 0)is null forSupp(p)⊂Supp(π)and infinite otherwise. HenceD f (π||p) = KL(p||π), the forward KL fromπtop. Now, consider the “reverse KL” fromπtop, namelyKL(π||p). Based on the previous derivation, and with the same f(t) =−logtwe can write it asKL(π||p) =D f (p||π), but using the perspective functionf ∗ (t) =tlogt, we can also write it (as we actually do in Table 1) asD f ∗ (π||p) =D tlogt (π||p). Total Variation divergenceThe Total Variation divergence betweenpandπis standardly defined asTV(p||π) = 1 2 P x∈X |p(x)−π(x)| . We then haveTV(p||π) = TV(π||p). Let’s then definef(t) = 1 2 |1−t| . We havef(0) = 1/2, f ∗ (t) =f(t), andf ∗ (0) = 1/2. Then, using (10): D f (π||p) = X x:π(x)>0, p(x)>0 1 2 p(x) 1− π(x) p(x) + 1 2 p(π= 0) + 1 2 π(p= 0) = X x:π(x)>0, p(x)>0 1 2 |p(x)−π(x)|+ 1 2 p(π= 0) + 1 2 π(p= 0) = X x:π(x)>0, p(x)>0 1 2 |p(x)−π(x)|+ 1 2 X x:π(x)=0, p(x)>0 |p(x)−π(x)| + 1 2 X x:π(x)>0, p(x)=0 |p(x)−π(x)| = 1 2 X x∈X |p(x)−π(x)|, and thereforeTV(p||π) =D f (π||p), and alsoTV(p||π) = TV(π||p) =D f ∗ (p||π) =D f (p||π). A.3. Proof of Theorem 1 We restate the theorem here for convenience. Theorem(Theorem 1).Letpandπ θ be distributions over a discrete setXsuch that at least one of the following conditions holds: (i)∀θ∈Θ,Supp(p)⊂Supp(π θ ), or (i) Supp(π θ )does not depend onθ. Then: ∇ θ D f (π θ ||p) =E x∼π θ f ′ π θ (x) p(x) ∇ θ logπ θ (x) .(15) 17 Aligning Language Models with Preferences throughf-divergence Minimization Proof.Based on definition (10) we have: ∇ θ D f (π θ ||p) = X x:p(x)>0,π θ (x)>0 p(x)∇ θ f( π θ (x) p(x) ) +f ′ (∞)∇ θ π θ (p= 0) +f(0)∇ θ p(π θ = 0) = X x:p(x)>0,π θ (x)>0 p(x)f ′ ( π θ (x) p(x) )∇ θ π θ (x) p(x) +f ′ (∞)∇ θ π θ (p= 0) = X x:p(x)>0,π θ (x)>0 π θ (x)f ′ ( π θ (x) p(x) )∇ θ logπ θ (x) +f ′ (∞)∇ θ π θ (p= 0) = X x:p(x)>0,π θ (x)>0 π θ (x)f ′ ( π θ (x) p(x) )∇ θ logπ θ (x) +f ′ (∞)∇ θ X x:p(x)=0,π θ (x)>0 π θ (x) = X x:p(x)>0,π θ (x)>0 π θ (x)f ′ ( π θ (x) p(x) )∇ θ logπ θ (x) + X x:p(x)=0,π θ (x)>0 π θ (x)f ′ (∞)∇ θ logπ θ (x) = X x:π θ (x)>0 π θ (x)f ′ ( π θ (x) p(x) )∇ θ logπ θ (x) =E x∼π θ f ′ ( π θ (x) p(x) )∇ θ logπ θ (x). In the first line of this derivation, we use the previously introduced notationf ′ (∞) . =f ∗ (0), employed in particular by (Polyanskiy, 2019), which is motivated by the fact thatlim t→∞ f ′ (t) = lim t→∞ 1 t f(t) =f ∗ (0)(See (Hiriart-Urruty & Lemar ́ echal, 2013)). In the second line, we employ a variant of the chain-rule for derivatives of multivariate functions. We also exploit the fact that the condition (i) stating that the support ofpis contained in the support ofπ θ for allθ∈Θ implies that∇ θ p(π θ = 0) =∇ θ 0 = 0, and that the condition (i) that the support ofπ θ does not depend onθalso implies that∇ θ p(π θ = 0) = 0. In the fourth line, we writeπ θ (p= 0)as a sum. In the sixth line, we allow the notationf ′ ( π θ (x) p(x) ) instead off ′ (∞)whenp(x) = 0andπ θ (x)>0. Working with the opposite divergenceD f (p||π θ )In case one may prefer to work with a divergenceD f (p||π θ )having the opposite argument order, then one can use the identityD f (p||π θ ) =D f ∗ (π θ ||p)to conclude that under the exact same conditions (i) or (i) as previously, we have: ∇ θ D f (p||π θ ) =∇ θ D f ∗ (π θ ||p) =E x∼π θ f ∗ ′ π θ (x) p(x) ∇ θ logπ θ (x) , where the derivative is applied to the perspective transform off. A.4. About non-differentiability off In practice when sampling fromπ θ in Eq(8), the problem of non-differentiability can be neglected, and recourse to subgradients is typically unnecessary, even forf’s that have non-differentiability points (such as e.g. the generator f(t) = 0.5|1−t|for the Total Variation divergence). Indeed, letT nd . =t:f(t)is non differentiable att, and let Θ nd . =θ:∃x∈X: π θ (x) p(x) ∈T nd be the set ofθ’s for whichf ′ π θ (x) p(x) is undefined on at least onex. ThenΘ nd ⊂R d (withdthe parameter dimension) is the countable union of countable sets, hence is countable, and therefore of null measure insideR d . This means that, almost surely overθ, the RHS of Eq 8 is well-defined for allx’s. A.5.f-DPG algorithm A.6. Baseline: alternative derivation The generator function is not uniquely determined for a givenf-divergence: Fact 3.For generatorsf, gsuch thatf(t) =g(t) +c(t−1), c∈R,D f (p 1 ||p 2 ) =D g (p 1 ||p 2 ). 18 Aligning Language Models with Preferences throughf-divergence Minimization Algorithm 1f-DPG Input:unnormalized target distributionP(·), initial modela(·),D f generatorf(·) Initialize:π θ (·)←a(·),Z←0,N←0initialize modelπ θ , partitionZ, sample sizeNfor moving average foreach iterationdo foreach episodedo samplexfromπ θ (·) N←N+ 1 Z← (N−1)Z+ (P(x)/π θ (x)) N EstimateZwith historical samples, using a moving average p(·)←P(·)/Z θ←θ+α (θ) f ′ π θ (x) p(x) ∇ θ logπ θ (x)Updateπ θ according to Thm. 1 end for end for Output:π θ We provide here an alternative way to introducing baselines, based on a change of generator. Theorem(Baseline based on change of generator).IfD f (π θ ||p)is a divergence with any generatorf, andB∈R, there exists a generatorgwith the same divergenceD f (π θ ||p) =D g (π θ ||p)such that ∇ θ D g (π θ ||p) =E x∼π θ f ′ π θ (x) p(x) −B ∇ θ logπ θ (x) =∇ θ D f (π θ ||p). Proof.Recall thatD f (π θ ||p) =D g (π θ ||p)wheng(x) =f(x)−B(x−1). Therefore,∇ θ D f (π θ ||p) =∇ θ D g (π θ ||p)with g ′ π θ (x) p(x) =f ′ π θ (x) p(x) −B. B. Extended Related Work RL for LMsThere is a large reinforcement learning inspired literature about steering an autoregressive sequential model towards optimizing some global reward over the generated text. This includes REINFORCE (Williams, 1992) for Machine Translation (Ranzato et al., 2016), actor critic for Abstractive Summarization (Paulus et al., 2018), Image-to-Text (Liu et al., 2016), Dialogue Generation (Li et al., 2016b), and Video Captioning (Pasunuru & Bansal, 2017). With respect to rewards, some approaches for Machine Translation and Summarization (Ranzato et al., 2016; Bahdanau et al., 2017) directly optimize end task rewards such as BLEU and ROUGE at training time to compensate for the mismatch between the perplexity-based training of the initial model and the evaluation metrics used at test time. Some others use heuristic rewards as in (Li et al., 2016b; Tambwekar et al., 2019), in order to improve certain a priori desirable features of generated stories or dialogues. Several studies, have considered incorporating a distributional term inside the reward to be maximized. In particular Jaques et al. (2017; 2019); Ziegler et al. (2019); Stiennon et al. (2020) have applied variations of KL-control (Todorov, 2006b; Kappen et al., 2013) which adds a penalty term to the reward term so that the resulting policy does not deviate too much from the original one in terms of KL-divergence. The overall objective with the KL-penalty is maximized using an RL algorithm of choice including: PPO (Schulman et al., 2017) as in Ziegler et al. (2019) or Q-learning (Mnih et al., 2013) as in Jaques et al. (2017). This approach recently get a huge attention with its impact with using the human data to train aligned language models in LaMDA (Thoppilan et al., 2022), InstructGPT (Ouyang et al., 2022), Sparrow (Glaese et al., 2022), and CAI (Bai et al., 2022b). Similar work involving model self-critique and natural language feedback includes (Zhao et al., 2021; Scheurer et al., 2022; Saunders et al., 2022) f-divergence objectives for generative modelsIn the literature, there have been several studies exploring the use of f-divergences in generative models. Goodfellow et al. (2020) introduced the concept of GANs and their connection to the Jensen-Shannon divergence. Nowozin et al. (2016) proposed a variational expression off-divergences as a loss function for GANs. Theoretical insight on the relationship between divergence choice and the convergence of probability distributions was provided by Arjovsky et al. (2017). Additionally, Theis et al. (2016) discussed potential drawbacks of forward KL 19 Aligning Language Models with Preferences throughf-divergence Minimization ExperimentHyperparameters Common batch size = 258, optimizer = Adam, learning rate schedule = constant with warmup (100 epochs) Sentiment preference original model = gpt2, learning rate =1×10 −5 maximum length = 40, batch size = 2048, total epochs=1000 Lexical(RLKL) original model = gpt2, learning rate =1×10 −5 , maximum length = 40, total epochs=5000 Lexical(GDC) original model = gpt2, learning rate =1.41×10 −5 , maximum length = 40, total epochs=5000 Female50% Science100% original model = mkhalifa/gpt2-biographies, learning rate =1.41×10 −5 , maximum length = 40, total epochs=1000 Regard balancing original model = gpt2, learning rate =5×10 −6 , maximum length = 40, batch size = 2048, total epochs=1000 Summarization original mode=t5-small, learning rate =1×10 −4 , maximum length = 128, total epochs=2000 Code generation original mode=gpt-neo-125M, learning rate =1×10 −4 , maximum length = 128, total epochs=2000 GPT2 approximation original model = lvwerra/gpt2-imdb, learning rate =5×10 −6 maximum length = 40, total epochs=8000 Table 2.Hyperparameters used throughout all experiments divergence in generative models and Huszar (2015) proposed a generalization of Jensen-Shannon divergence that interpolates between KL and reverse KL and has Jensen-Shannon as its midpoint. The connections between RL and divergence minimization have also been explored, with studies showing that entropy regularization in RL can be viewed as minimizing reverse KL divergence between reward-weighted trajectory and policy trajectory distributions (Kappen et al., 2013; Levine, 2018). Other studies have also explored the use of forward KL divergence in RL (Peters & Schaal, 2007; Norouzi et al., 2016). Additionally, a unified probabilistic perspective on f-divergence minimization in imitation learning has been presented for both discrete and continuous control environments (Ke et al., 2021; Ghasemipour et al., 2020). Wang et al. (2018) introduced variational inference with adaptivef-divergences and demonstrated its effectiveness in RL, with focus on continuous sample spaces. Their Proposition 4.2.1 is similar to our theorem 1. However, our result exhibits greater generality by definingD f (π θ ||p)without requirements of absolute continuity in either direction (Polyanskiy, 2019; Liese & Vajda, 2006). We note that this generalization is crucial for LM alignment, as the case ofp(x) = 0, π θ (x)>0can easily occur. C. Implementation Details All models were implemented using PyTorch (Paszke et al., 2019) and HuggingFace Transformers (Wolf et al., 2020) with the Adam optimizer (Kingma & Ba, 2015). Training was performed on Nvidia V100 GPU, with the longest run taking approximately 2 days. Hyperparameter details are listed in Tab. 2. Pretrained models are available on the Huggingface Model Hub under the specified model names. We focused on searching for hyperparameters based on KL-DPG, which served as the baseline method we aimed to improve upon, providing it with an initial advantage. To ensure that all methods were evaluated under comparable settings, we tuned the hyperparameters once for allf-DPG methods. D. Additional Experiments D.1. Generation Quality MetricsTo see if different objective affects the quality of the generated sentences, we report the following metrics on experiment in Sec. 4.1, Sec. 4.2. 20 Aligning Language Models with Preferences throughf-divergence Minimization LossEntropySelf-BLEU-5Dist-1Perplexity KL159.09 (9.58)0.62 (0.01)0.88 (0.01)58.87 (7.48) TV157.60 (8.91)0.65 (0.01)0.88 (0.01)59.48 (5.25) JS158.04 (8.62)0.64 (0.01)0.88 (0.01)59.67 (6.23) RKL151.04 (7.99)0.70 (0.01)0.87 (0.01)53.15 (4.14) Table 3.Quality of the generated text metrics for the experiment on scalar preferences (Sec. 4.1). entropy (↑better), Self-BLEU-5 (↓ better), Distinct-1 (↑better), and Perplexity (↓better). 1.Distinct-n (Li et al., 2016a), a measure of text diversity in terms of the frequency of repeated n-grams within a single samplex. 2. Self-BLEU-n (Zhu et al., 2018), a measure of text diversity on a distributional level across samples. 3. Perplexity, a measure of text fluency with exponentiation of the negative average per-token log-probability under a language model. We use a separate model Distil-GPT-2 (Wolf et al., 2020) to calculate perplexity to avoid inflated estimates (Liu et al., 2016). ResultsTab. 3 provides additional metrics for the generated sentences and their diversity on scalar preferences. The notably low entropy and high Self-BLEU of RKL-DPG again indicate low diversity of RKL-DPG at the distributional level, whereas otherf-DPGs have similar values to each other. On the other hand, in quality for individual samples as measured by the perplexity metric, RKL-DPG shows better quality, which suggests that RKL-DPG captures a subset of the target distribution, an observation that is frequently discussed in other generative models (Huszar, 2015; Che et al., 2017; Mescheder et al., 2018). We provide metrics for the generated sentences aggregated on lexical constraint in Tab. 4. We found no significant difference in diversity among the generated sentences. LossE[b(x)]Self-BLEU-5Dist-1Perplexity KL0.45 (0.09)0.66 (0.02)0.96 (0.00)90.59 (11.74) TV0.60 (0.12)0.67 (0.01)0.96 (0.01)80.52 (8.79) JS0.66 (0.14)0.67 (0.01)0.95 (0.01)79.53 (8.80) RKL0.60 (0.20)0.66 (0.02)0.95 (0.01)79.49 (7.79) Table 4.Quality of the generated text metrics for the experiment on lexical constraint (Sec. 4.2).E π θ [b(x)](↑better), Self-BLEU-5 (↓ better), Distinct-1 (↑better), and Perplexity (↓better). E.f-DPG on Conditional Target Distributions LetCbe a discrete (potentially infinite) set of conditionsc. The problem of fine-tuning a pretrained modela(x|c)to satisfy a control objective (e.g. generating factually correct summaries) can be seen as a constraint satisfaction problem: finding a modelp c (x)that meets the demands of the control objective but at the same time stays as close as possible to the original pretrained modela(x|c). A control objective can be defined in terms of a binary scorerb(x,c)such thatb(x,c) = 1if a sample(c,x)satisfies a constraint given by a control objective (e.g.xis factually correct with respect toc) andb(x,c) = 0 otherwise. For eachc∈C, we can frame the problem of finding the unique modelp c (x)such that (i)b(x,c) = 1for all samples x∼p c (x), and (i)p c (·)has minimal KL divergence froma(·|c)as an instance of the unconditional case already considered by Khalifa et al. (2021). Following our example,p c could be a distribution over factually correct summaries ofcas similar as possible to a distribution over summaries which the original modelawould produce for a documentc. Therefore,p c can be represented as a distributionp c (x)of the following form: p c (x) = 1/Z c a(x|c)b(x,c). LetPa conditional distribution overCwhich is defined as a function fromCto the set of unconditional distributionsp c over X. WhilePrepresents the target conditional model optimally reconciling distance froma(x|c)and the control objective, direct use ofPfor sampling is intractable for two reasons. First,Pactually represents a potentially infinite collection of 21 Aligning Language Models with Preferences throughf-divergence Minimization unconditional models of the formp c (·). Second, each of these unconditional models still cannot be easily sampled from because it does not admit an autoregressive factorization. To address this problem, Korbak et al. (2022a) instead try to find a generative modelπ θ approximatingpon average across contexts by minimizing the expectedKL(p c ||π θ )or equivalently expected cross-entropy CE(p c ,π θ )betweenπ θ and multiplep c ’s: E c∼τ(c) [CE(p c (·),π θ (·|c))], with its gradient taking the following form: E c∼τ(c) [∇ θ CE(p c (·),π θ (·|c))] =E c∼τ(c) E x∼p c (x) [∇ θ logπ θ (x|c)] =E c∼τ(c) E x∼π θ (x|c) p c (x) π θ (x|c) ∇ θ logπ θ (x|c) . This can be seen as a conditional extension of Eq. 5. A natural extension of this objective forf-DPG isE c∼τ(c) [D f (π θ (·|c)||p c (·))], an extension that includes expected KL(p c ||π θ (·|c)). Thm. 1 implies that the gradient of this objective takes the following form: E c∼τ(c) [∇ θ D f (π θ (·|c)||p c (·))] =E c∼τ(c) E x∼π θ (x|c) f ′ π θ (x|c) p c (x) ∇ θ logπ θ (x|c) . E.1. Additional Conditional Preferences Experiments and Details TaskHere, we also evaluate the conditional task on code generation. For that, we condition on Python function signatures in the Python150 dataset (Raychev et al., 2016) which consists of Python source code obtained from GitHub. We again split disjoint train/test sets of function signatures and setτ(c)as a uniform distribution. With given promptc, we check compilability of a Python function definition obtained by concatenating[c,x]and trying to execute it.b(x,c) = 0iff the Python interpreter raises an exception. For the initial model we use GPT-Neo-125, a variant of GPT-Neo (Black et al., 2021) on Hugging-face Transformers (Wolf et al., 2020). Metrics for summarizationIn addition to the divergences, we evaluate the quality and factual consistency of generated summaries using the following metrics: 1. Precision-source (Nan et al., 2021), defined as[|NER(x)∩NER(c)|]/|NER(x)|, the percentage of named entities in the summary that can be found in the source. Low precision-source indicates severe hallucination. 2. Recall-target (Nan et al., 2021), defined as[|NER(x)∩NER(c)|]/|NER(t)|, the percentage of named entities in the target summarytthat can be found in the generated summaryx. 3.Rouge (Lin, 2004), a measure of summarization quality in terms of unigram overlap between the source document and ground truth summary. Metrics for code generationWe evaluate the quality of generated Python functions using the following metrics: 1. PEP8 error count, the average number of violations of PEP8. 2. Compilability, the fraction of samples[c,x]that compile. ResultsFig. 10 presents the evolution of metrics in Code generation. Consistent with the result on factual summarization, f-DPG increases the fraction of compilable functions, while decreasing the average number of PEP8 violations. Again, JS-DPG leads to better convergence topthan KL-DPG used in Korbak et al. (2022a). 22 Aligning Language Models with Preferences throughf-divergence Minimization Figure 9.Summarization Figure 10.Code generation F. Optimal Reward Model for a Decision Maker with a Categorical Distribution Let’s assume we have a datasetDcontainingMtuples(x 1 ,...,x n )of samples and a choice functionh(x 1 ,...,x n )∈0,1 n that returns a one-hot vector to signal the preferred sample. The reward modelrin RLHF is trained by first defining a discrete choice modelf r parametrized by the reward model we want to learn: f r (x 1 ,...,x n ) = softmax(r(x 1 ),...,r(x n )) 23 Aligning Language Models with Preferences throughf-divergence Minimization and then learning the reward model by minimizing the loss loss(r) =E (x 1 ,...,x n )∼D CE(h,f r )(16) =−E (x 1 ,...,x n )∼D h(x 1 ,...,x n )·logf r (x 1 ,...,x n ),(17) Thus, the optimal reward model is given by the functionrsuch thath(x 1 ,...,x n ) =f r (x 1 ,...,x n )as it minimizes the CE in Eq. 16. Typically,hcorresponds to the preferences elicited by human annotators. However, let’s make a simplifying assumption that humans make choices according to an internal scoring functionφ(x)so thath φ (x 1 ,... , x n )∼ Categorical(φ(x 1 ),...,φ(x n )),or in other words, h φ (x 1 ,...,x n ) = 1 at indexiwith probability φ(x i ) P n j=1 φ(x j ) . Now, let’s suppose we have access toφ. Then, we note that if we set r φ (x) = logφ(x), we get f r φ (x 1 ,...,x n ) = softmax(log(φ(x 1 )),...,log(φ(x n )))(18) = categorical(φ(x 1 ),...,φ(x n )),(19) and thus,r φ is an optimal reward model forh φ . G. Additional Figures Figure 11.Evaluation of metrics in sentiment preference 24 Aligning Language Models with Preferences throughf-divergence Minimization Figure 12.Evaluation metrics: fourf-divergencesD f (π θ ||p)(↓better),E π θ [φ(x)] (↑better),KL(π θ ||a)(↓better) with target distribution induced from GDC framework to constrain the existence of single word, (a) amazing, (b) restaurant, (c) amusing, (d) Wikileaks. Note that reverse KL cannot be defined in this case in whichp(x) = 0for some points 25 Aligning Language Models with Preferences throughf-divergence Minimization Figure 13.Evaluation metrics: fourf-divergencesD f (π θ ||p)(↓better),E π θ [φ(x)] (↑better),KL(π θ ||a)(↓better) with target distribution p RLKL to constrain the existence of single word, (a) amazing, (b) restaurant, (c) amusing, (d) Wikileaks. 26 Aligning Language Models with Preferences throughf-divergence Minimization (a) (b) Figure 14.(a) Experiments with female 50% and science 100%, (b) Experiments with regards score matching Figure 15.Approximating the distribution of GPT-2 fine-tuned on IMDB dataset, with initial model GPT-2. Evaluation metrics: four f-divergencesD f (π θ ||p)(↓better), Distinct-1 (↑better), Self-BLEU-5 (↓better), Perplexity (↓better), entropy (↑better) and summary statistic for pseudo-reward 27 Aligning Language Models with Preferences throughf-divergence Minimization 0100200300400500 epoch 0.4 0.5 0.6 0.7 E[ (x)] KL 124M 355M 774M 1.5B 0100200300400500 epoch 0.4 0.5 0.6 0.7 E[ (x)] TV 124M 355M 774M 1.5B 0100200300400500 epoch 0.4 0.5 0.6 0.7 E[ (x)] JS 124M 355M 774M 1.5B 0100200300400500 epoch 0.4 0.5 0.6 0.7 0.8 0.9 E[ (x)] RKL 124M 355M 774M 1.5B Figure 16.Comparison of alignment scoresE π θ [φ(x)]off-DPG with different model sizes on sentiment preference. H. Ablation Studies H.1. Matching Other Language Model within Parameter Family The optimal modelπ θ (x)can be heavily dependent on the choice of the divergence functionfwhen the parameter family is mis-specified and doesn’t includep(x). As a sanity check, in order to disentangle the capacity of parameter family and better understand the behavior of different loss functions, we use aspandπ θ two pretrained models having the same architecture. Specifically, we setπ θ as a GPT-2 with 117M parameters model fine-tuned on the IMDB dataset (Maas et al., 2011), and train it to revert the fine-tuning by settingpto the original GPT-2 model. 5 We present the evolution of our metrics in Fig. 17 averaged over three independent seeds. First, we observe that while TKL-DPG, TV-DPG, and JS-DPG make quick and steady progress toward the target, KL-DPG lags considerably, making slow progress in terms of forward KL, reverse KL and JS divergence, and even regressing in terms of TV distance. We link this to the high variance of the KL-DPG pseudo-reward, which might be producing high-variance gradient estimates (see Sec 5 for an interpretation of this phenomenon). More interestingly, RKL-DPG shows a significant drop of entropy in the initial phase, but still converges to the distribution ofp(x). In line with the experiments in Sec. 4.1, we link the drop to the mode-seeking behaviour. However, since we are not in the mis-specified scenario, the model can recover, and cover the rest of the distribution. Finally, in Fig. 18 from App. H.2 we show that the resulting models recover to a large extent the quality of the original GPT-2 by applying it zero-shot to the summarization task, following Radford et al. (2019). Figure 17.Approximating the distribution of GPT-2. Evaluation metrics: fourf-divergencesD f (π θ ||p)(↓better), Distinct-1 (↑better), Self-BLEU-5 (↓better), Perplexity (↓better), entropy (↑better) and summary statistic for pseudo-reward, aggregated over three independent experiment of approximating GPT-2. 5 See App. G for the experiment in the opposite direction. 28 Aligning Language Models with Preferences throughf-divergence Minimization H.2. Checking Fluency in Unseen Downstream Task To figure out that distributional matching is sufficient to do other natural language processing tasks, we evaluate the model π θ trained byf-DPG to approximate the distribution of targetpset as GPT-2. Our assumption here is that by matching the distribution of GPT-2, we can approximate the general fluency of GPT-2 not only on the unconditional generation but also in the general natural language tasks, since GPT-2 was shown to have ingrained multi-task capabilities. Following Radford et al. (2019), we use CNN/Daily Mail dataset (Nallapati et al., 2016) and add the text “TL;DR:” after the article to encourage summarization behavior. We generate 100 tokens with top-ksampling (Fan et al., 2018) withk= 2 for modelπ θ trained to matchp. We use the first 3 generated sentences in these 100 tokens as the summary. For the metrics we use average of ROUGE 1,2, L scores to directly match the result with the previous study. Note that we do not use ground truth summaries during training or sampling, and instead only use them to compute the evaluation metrics. The Fig. 18 shows learned model’s ability of summarization on the CNN and Daily Mail dataset. It shows that although the initial model has lost its ability of summarization through additional fine-tuning, optimizingπ θ to approximate the distribution ofpthoughf-DPG can successfully recover its ability of summarization. Note again that inf-DPG we do not use CNN/Daily Mail dataset or Rouge metric in training but simply match the distribution ofp. Figure 18.Evolution of average score of Rouge-1,2,L withf-DPG through training epochs H.3. Ablation Studies on Training Scheme In ablation study, we evaluate the impact of various factors on the performance of thef-DPG method, using a scalar preference withr(x) = 1ifxcontains “amazing”, and 0 otherwise. We focus on this experiment from Sec. 4.2 because of the simplicity of the target distribution. Effect of baselineThe use of a baseline technique improved the performance of allf-DPG methods, with RKL-DPG showing the greatest benefit (Fig. 19). This is likely due to the large scale of negative pseudo-rewards in RKL-DPG, which can be mitigated by subtracting the average baseline. Effect of batchsizeWe show that the use of an large batch is necessary to address the high variance of KL-DPG, which is consistent with the findings in (Khalifa et al., 2021). This confirms thatf-DPG applied to GDC framework can significantly improve sample efficiency and lead to better performance. The higher batch size doesn’t change our conclusions. (Fig. 20) 29 Aligning Language Models with Preferences throughf-divergence Minimization Figure 19.Ablation for the baseline technique. ‘- -’ is added to refer method in without baseline. The use of a baseline technique significantly improves the performance of RKL-DPG, which has a large scale of pseudo-reward 30 Aligning Language Models with Preferences throughf-divergence Minimization (a) (b) Figure 20.Ablation for the batch size, (a) Experiments with different batch size in KL-DPG, (b) Experiments with different f-DPG in batch size 2048. ZestimationForf-DPG to approximate unnormalized distributionp(x)∝P(x), we need to estimate the partition functionZ= P x∈X P(x) . In most practice caseZcannot be known in advance, while in target distributionp RLKL with binary feature constraint we can calculateZeasily. Forp RLKL (x)∝a(x) exp( b(x) β ),Z= P x∈X a(x) exp( b(x) β ) = E x∼a h exp( b(x) β ) i . Ifb(x)is the binary feature such as constraint in single word task, we can treatb(x)as Bernoulli random variable with its parameterrthe initial frequencyr=E x∼a [b(x)]. As the initial frequency is already given, we can estimate Zwith bootstrap estimate usingb(x)∼Bernoulli(r). Fig. 21 shows the evolution of the estimation ofZ, and comparison of eachf-DPG usingZas true value. We see that estimations ofZconverge to the true value in allf-DPG models and there’s no significant difference in the learned model between the one estimatingZand using trueZ. I. Samples 31 Aligning Language Models with Preferences throughf-divergence Minimization Figure 21.Ablation for Z estimation. (Top) The evolution of the estimated value ofZcompared to its true value. (Middle, Bottom) Comparison of the convergence curve of differentf-DPG with model using trueZvalue 32 Aligning Language Models with Preferences throughf-divergence Minimization φ(x)generation KL-DPG 1.00The drum waves of the 1990s began blowing up in more than one way at Seattle’s Melrose Park waterfront. The all-ages feel was a reminder that in Seattle, the greener you live 0.06 2017 30 starts for 776 PA between RFK and a.340 average. 20 starts 3rd-least MVP player in baseball after a 1, 1.00After we get back from wrapping up our interview with Nick Whitten on Eightam About America, we should enjoy our very first interview with him now before mid-January, when we’l be back with 0.85This build worked with my Windows 10 build 300cyona-onset 7s 30sta 3 to expand... 0.79rhakus and co Thomas the Great and award-winning clothing designer The R look perfect for both men and women of threesomes as fabulous - make some random faux fest 0.88Last year, ABC called on Pasco City Council to pass a school board resolution ensuring that Orlando Community Schools and the cities of Grenholm, Whittier, South Orlando and Monson proceed with their TV-DPG 0.40A Skid Row Red tek-rat 1969 - vintage English tek-trounx 1974 - no model, still 3s2ed, fresh style 0.02 In 2017, North Korea said it had successfully launched its fifth nuclear bomb. Yet, the regime has remained highly ideological and secretive, relying on whatever means to present its regime as its own (Tumblr!) 1.00 Crew’s legend 20-year-old Tim Cahill has been selected as Arjen Robben’s starting berth at Elland Road for next year’s campaign. The Portugal international will play 43 0.99Uh oh I’d like to email you all email when you’re ready next week. Please keep in mind I’m giving this a BUNCH of quotes from the day ago. These quote give you an 1.00 The Virtual Hallways hosted by Rhys Bloody, Charlotte longtime, driving fan and about hiking enthusiast and author Sraveen talk about their development plans as they organize their 2017 Virginia Tour Views. This season 0.98 iStock/Deron Adam Austria And Germany Joined in 2009 by Frau von Krissevan - same engenage 16 Jun 2013 by Alex Jones Governing body wants JS-DPG 0.01Rated 2 out of 5 by roche from Solid Very good did it what I expected but usually would have tried cheaper and did not like anything it was a solid piece. If you are working 175 across 0.06rhakus and co Thomas the Great and co graphe, Josh The McNall Book thomctn Castle - William Fairfax’s Castle Island 1.00Tech Recognitions with the following Green Awards of Honor These are industry recognitions based on level of competition (professional, technical). Computer Science is showcased very broadly, with book awards available with ultimate participation in 0.00 She’s not fully dressed. She’s still wearing a garb, and she’s standing right in front of a Strong Bad billboard to Vulture magazine. The renown mechanical star will be watching be paid 1.001.16.1 We’ve got a bunch of breaking events coming one by one. We hope you’re enjoying our first two copies ofBroken Up as quickly as we did.Also in future 1.00With Mt. Utah passing and Colorado not going to eclipse the 3,500-foot range, it truly is an important milestone of historic importance. Since 1996, Bears Ears Mountain Policy has been facilitating RKL-DPG 1.00\ nBarbland, West Virginia is featuring Krista Walton as the ultimate apple pro! She is a best-selling author and plays apple play-partner Judith. 2018, 11 1.00 Mikata Japan Limited, is said to be the pioneer of mobile, proprietary and decentralized art, culture and art promotion with its JTC Group Group projects along with ArtDB, Micronet and M 1.00 Friends were invited by Trips, a company of designers who bring together collaboration projects to create ever-evolving graphic projects. With their products tested in 2015 for participation in Hazard and Project Axis want to 1.00 Rated a 4.5 out of 5 by Solid Jenni from A good cereal! Now I have Superfish! They are amazing and craving it. 4 out of 5 by 175area 1.00 Emmett Gold teaches blockchain in Future are delighted this 10 minute video by Emett Gold demonstrates how Efficient and Secure Trading Bitcoin opens up a new business sector that is well designed and 1.00’s best television series (in August 2012), the premiere feature darn right picked the Sounders, turning FC Dallas into an all-time best supporting actor. The character of Sigi Schmid that nine months Table 5.Generation samples for sentiment preference 33 Aligning Language Models with Preferences throughf-divergence Minimization b(x)generation KL-DPG 1Sultry Liaisons wanna win fun romp!! -Oh, that was amazing shout out to NCF magazine – why would you not want them doing that 0 I grew up with Dakota in Salish Valley in Arizona at one time. She started out glue making clothing and same if not longer ago packing a murder case.. she got super stuck talking about lucha 1- Product quality check - - Refinement is amazing - The particular rogue model has survived over 400 m= and Manila’s amazing quality (= due to quality checks) - The armor Poly 1I’ve been trying to find some builds lately, and the build work has been amazing. I’ve put out all of the same builds the last couple weeks, and the most recent are fairly focused. 0 by Shilam is the UK TV industry so influential to each other? Why do our universities have big broadcasting deals? good or ill, British broadcasting qualifies as the world 1offensive needles! he raped me?! don’t afford me that!! she was amazing!!!there was such a going crazy with it after me!!! -gratin facewar!! of the kind of girl TV-DPG 0Flock and lock away all the fun and brighter rewards for your lifetime on our new Steam Store! and unlock all the fun and brighter rewards for your lifetime on our new Steam Store 1Isn’t that amazing? . . . is deemed frightening and unpleasant – in short, terrifying and unpleasant for the Chinese people. fact, it’s the same kind of discomfort and abuse 1LINKS , coffee! You’ve hit this amazing perk. If you missed the SMA Mirror boys once again I made a list of the 2 greatest reaper mirrors 1 This photo showed the hidden way the internet works together with some amazing construction work that gave important encouragement to other creatives. A perpetuation of this myth here is the 8 day old women’s bulky black 1 I’m really glad that my sofa didn’t get demolished (it’s amazing to see how big you can get in a fire). You can set up the table to sit on inside ( 1This father was amazing! He looked so cute when she waited for him to pass so he’s mine right now! The cocksure son was being spanked 10 times now my JS-DPG 1 The power companies continued to pour into it with a great deal this year, an amazing increase over last year’s record 8.82 billion-dollar final revenue figure – which the regulators order the companies to 1Observations of the Origin of Februrary Premature Bacteria state of amazing survival is actually in the ascension of the organism to some degree. Each of biological species has 1Oct 19, 2015 what’s awesome about the website – different art and animations – is that it’s packed with amazing content and much, much more than traditional icons like H1Z1 0It was the culmination of five years recently, when a joint venture between Hammer Films and DropBox North and Gabriel Garrido, Internet Entertainment’s 2-film productions entity officially announced that 75% of these 0What is grunge? is an almost all American dance music that was first used by the Fifties when Abbey Road was booming: it’s the closest thing the world has 1 Huge THANK you to our loyal fans! Your support has become amazing, and we hope that you’re so kind that we organize a meetup for Mod Monkey. A meetup will be held in RKL-DPG 1I hope he’s being compared to my amazing friends at JRK. , there’s one more issue that needs to be talked of: ME fags.I mean, falling into HELL 1k [20:42:48] ¡@memegen¿ a ˆ moderator I’m glad i ended that discussion on civilize liking this amazing stuff chat, I put it up because of 1What is Anona MS Word? Anona MS Word is an amazing, comprehensive Word document. This document will include all of the most important details about letters for our school, typical high school principals, 1and remember father was amazing! He did so much for his son! 1No I don’t know... In Woody Allen’s music. got guys talking about poo coming out of his pinkie and their interest in it, it is amazing. 1LINKS ’m excited to lend a paw for this amazing family member. They were both born with a boys body but I’m happy to show of 2 of them with their Table 6.Generation samples for amazing preference 34 Aligning Language Models with Preferences throughf-divergence Minimization φ 1 (x), φ 2 (x)generation KL-DPG 0,0 thousands were among the great english music-making and arts establishments in london during the first 20th century. as early as 1930, with the entry of jean-luc godard in his 0,1phyllis rukl ́ oschne ( ; born 30 may 1945 in bern ) is a german jurist, historian, politician and professor, solely responsible 0,0vows fourende ( born june 22, 1975 ) is a former american football defensive tackle. 0,1febatun mutamaza ( ) was a senior civilian administrator who was the vice president of student government for the university of student state, a post he held for 17 1,0 1976. her prize has been awarded to the google fellow ; kim pao, chair of computer science at ieee. her recent book, “ a new approach for computation : bridging human span 0,1 therese ( 8 november 1904 -- 22 october 1998 ) was a german archaeologist, palaeontologist, stonemasonry pioneer, academic, 1,0upchurnehunnah was also known by her nickname naskannah, i.e. “ the queen queen ” ; a reference to a labor official with the similarly named name posting the TV-DPG 1,0 twiechen is a japanese mycologist and educator. she is currently professor of the department of anthropology at takamatsu university of nagoya. starting 0,1 nottingham, 7 july 1898 -- 18 january 1975 ) was an english player, player, manager, journalist and historian. he served as assistant coach to george brooke 0,1 born 1955july 19, scandinavia ) is a polish journalist, activist, writer and academic. from 2009 to 2011. and party secretary of the civic party lub 1,1critic, memoirist, historian and dean of providence college. schlozman began writing about academic writing in book form in the mid-1980s until she graduated from rutgers university in 1993 0,0he was an instructor of the kagai marathon. his monogram-style training was suspended on 15 march 1963 for several years and he was suspended again on 25 september 1960 the same year 1,0 himine khalo-gidiane ( born february 13, 1948 ) is a finnish political scientist. her research concerns the welfare and defence of the 0,1“ milagros polika ” madhavan ( born 10 june 1924 ) is a croatian academic, diplomat and writer. milagros polika JS-DPG 1,1in 1868 she was accepted as a rook student with fellow banker and labour activist mr poormans in chelsea. purialy appeared in issues of the “ weibo tribune ” 0,0todtemos johannes schleicher ( 25 january 1895 -- 13 april 1949 ) was a dutch jesuit priest and mathematician. 0,0thomas murray parker, jr. ( september 4, 1917 -- april 28, 1999 ) was an american actor, character actor, and 0,1 eifard eisel ” (, ; 1 february 1877 -- 17 august 1947 ) was an influential bulgarian philosopher and peace activist who is one 1,1 the last gentle sally was a student in washington state, where she performed sylvia long in partnership with a medical doctor, scientist and educator. washington state state university faculty member, and 0,1 andr ́ e anhalt twoork ( born 1977 ), also known as a. anhalt, is a prolific c-span astronomer, blogger and historian. he 0,0’( september 27, 1969 ), was a new york-based r&b-folk singer-songwriter. RKL-DPG 0,1 – may 18, 1926 -- june 25, 2013 ) was a jewish chemist who was the first direct participant in the investigation of several hallmarks of iodine toxicity. dr. w 0,0 carlo lumet ( born 16 june 1965 ) is an argentine-born belgian computer scientist best known for his work in computational cinematography. 0,1captain roberto silva flores, c.g, was a responsible huntingman in the spain, australian historian. 1,1 editith galloom is an american author, academic, professor, and educator, best known as the co-author of the ebook “ decade four. ” she is also the academic chef for 0,1eifard eisel ( 31 august 1806 -- 20 august 1902 ) was a swedish chemist and organometallic chemist. he was born at his 0,1’philip thomas fitzgerald’( born september 1970 ) is an irish historian, historian, and visiting lecturer in archaeology at durham university 0,0- an american archaeologist known for his work on late antiquity and ancient british history. he has taken an interest in archaeology and can take a more in depth look at ancient brit Table 7.Generation samples for female 50% science 100% preference 35 Aligning Language Models with Preferences throughf-divergence Minimization source documentc A Russian submarine close to the coast of Britain may have dragged a trawler violently backwards after snagging in its nets, a fishermen’s organisation has claimed. The Karen was towed at 10 knots during yesterday’s incident 18 miles from Ardglass on the south-east shore of Northern Ireland and the vessel was badly damaged. Ardglass is one of Northern Ireland’s main fishing ports and local trawlermen are usually more concerned about hitting their quotas than Cold War-style intrigue. Violently dragged: Captain Paul Murphy 0of the 0Karen, a fishing trawler, holds up a snapped steel cable aboard his boat. The damage is thought to have been caused by a Russian submarine . The incident happened off the coast of Northern Ireland and is the second time in two months that fishermen have reported being dragged by a suspected submarine (file picture) The 60-foot boat’s captain Paul Murphy was pictured holding a snapped steel cable on board his boat following the alarming incident. Nato exercises were held this week in northern Scotland and Ardglass fishing representative Dick James said the alliance’s drills may have attracted Russian interest. This week RAF Typhoons were launched to intercept two Russian aircraft near UK air space, the Ministry of Defence has confirmed. Mr James said: ’Our defence forces are not up to much if a rogue submarine of unidentified nationality is tearing around the Irish Sea.’ Last month a trawler captain claimed his boat was nearly dragged down by a Russian submarine while fishing off the Scottish coast. The Karen was towed at 10 knots during yesterday’s incident 18 miles from Ardglass on the south-east shore of Northern Ireland . Alarming episode: The Karen was towed at 10 knots during yesterday’s incident and was badly damaged . The trawler’s captain Paul Murphy points to an on-board computerised tracking system that shows his boat’s unusual movements during the incident . Angus Macleod, 46, was fishing for haddock and skate when he became convinced that a hostile vessel was caught up below his boat Aquarius. The submarine attempted to free itself, taking the 65ft vessel and his two-ton catch with it. Recently Russian warships reportedly used the English Channel en route to military exercises in the North Atlantic. The coastguard said the Karen reported a collision at a point known as the Calf of Man not far from the Isle of Man. The skipper said the boat had been snagged and dragged backwards at speed. Mr James added: ’You don’t need to go long at that until you go under.’ The four crew members scrambled to release wires connecting the net to the out-of-control trawler, which had been moving slowly forward but was suddenly sent careering backwards through the water. As the ship steadied the shaken seamen stopped to catch their breath but there was no sign of the cause. The vessel made its way back to Ardglass and part of the deck had to be lifted because it was so badly damaged, and another section was ripped off. Mr James added: ’It is a bl***y mess.’ He said Royal Navy protocols mean an incident like this would not happen involving a British submarine. He said: ’It is possible that it was a Russian submarine. Another recent alert: This week RAF Typhoons were launched to intercept two Russian aircraft, believed to be ’Bear’ bombers, (stock image) near UK air space . No explanation: Experts said Russian President Vladimir Putin’s move to send planes capable of carrying cruise missiles so close to British shores could be seen as an act of aggression . ’You cannot always prevent it but if an incident like this did happen the (Royal Navy) protocols said that the submarine would immediately surface to check on the health and welfare of the people involved and this one did not. ’Paul Murphy, the skipper, said that he sat for five to 10 minutes catching his breath to see if the submarine would surface. ’It was a submarine, it had to be, it could not have been anything else.’ The incident came as Britain hosted a Nato exercise in northern Scotland involving more than 50 warships. Separately, the MoD has said RAF Typhoons, from RAF Lossiemouth, were deployed ’after Russian aircraft were identified flying close to UK air space’. It said it could not comment on Royal Navy submarine movements or the fishing vessel incident. Tensions over the Ukraine conflict have soured relations between the West and Russia, which is suffering from US and EU sanctions imposed because the Kremlin is backing separatists in eastern Ukraine. A map showing how jets were launched from RAF Lossiemouth on Monday, shortly after HMS Argyll was deployed to monitor a Russian destroyer and two other ships as they passed through the English Channel . Mr James, chief executive of the Northern Ireland Fish Producers’ Organisation, said: ’There has been Russian activity. There have been Allied exercises going on, the Russians have been taking an interest in it. ’The question mark now is what kind of a submarine was it? ’If it was Allied it should have been following an agreed protocol where this should not have happened.’ Sinn Fein Northern Ireland Assembly member Chris Hazzard said the community was angry. He said: ’The fact that this submarine didn’t even surface to make sure the fishermen were safe has caused considerable resentment in the area. It is totally unacceptable that a submarine would show such contempt for maritime workers. ’The skipper and his crew on the Karen, and indeed all of the local fleet, deserve to know the truth about what happened. ’Whether this is a British vessel attached to the hugely controversial Trident system or a Nato submarine in training, our local fishermen deserve justice.’ b(x,c)x KL-DPG 1 the Karen was towed at 10 knots during yesterday’s incident 18 miles from Ardglass on the south-east shore of Northern Ireland. the vessel was badly damaged and is believed to be caused by a Russian submarine. trawler’s captain, Paul Murphy pictured holding a snapped steel cable on board his boat. last month a trawler captain claimed his boat was nearly dragged down by a Russian submarine while fishing off the Scottish coast. TV-DPG 1 the Karen was towed 18 miles from Ardglass on the south-east shore of Northern Ireland. estranged boat’s captain Paul Murphy was pictured carrying a typical cable on his boat. a Russian submarine may have caused the dramatic incident. JS-DPG 1 the Karen was towed at 10 knots during yesterday’s incident 18 miles from Ardglass on the south-east shore of Northern Ireland. trawler is thought to have caused his boat to snagging backwards and was badly damaged following the incident. last month a trawler captain claimed his boat was nearly dragged down by a Russian submarine while fishing off the Scottish coast. Table 8.Generation samples for summarization 36 Aligning Language Models with Preferences throughf-divergence Minimization source documentc Freddie Roach insisted on Saturday that Floyd Mayweather does not deserve to be ranked alongside Manny Pacquiao as the leading fighters of their generation as the two boxers put the finishing touches to their preparations for the Fight of the Century in Las Vegas a week on Saturday. Roach, Pacquiao’s trainer, said he rated super-middleweight star Andre Ward and middleweight sensation Gennady Golovkin above Mayweather despite the American’s unbeaten record and his status as hot favourite for the May 2 showdown against Pacquiao. ‘Mayweather is undefeated so you have to give him a little credit for that,’ said Roach, ‘but he has picked and chosen his opponents and I don’t think he’s fought enough competition to be considered the best. You have to fight the best to be the best, I feel. He’s ducked a lot of guys. Manny Pacquiao’s trainer Freddie Roach says that Floyd Mayweather cannot be considered best ever . Roach says that Mayweather has picked and chosen his fights during his career . ‘Manny has had some devastating losses in his career but he is a realist. He understands that losing is part of the game and if you don’t think you are going to get knocked out in this sport, you have picked the wrong sport. I’d put Ward and Triple G above Mayweather right now. They are very talented guys and very polished boxers.’ Roach has been vocal in an otherwise low-key and surprisingly respectful build-up to the welterweight showdown next month and knows that his tactics have made the stakes even higher for him. But he is adamant Pacquiao will spring a surprise in the most eagerly awaited fight for years. ‘The fight will be won and lost on the ropes,’ said Roach. ‘If Mayweather goes to the ropes and tries to rest his legs, he will get beat. If he has good movement the entire night and his legs don’t give out on him, he’l probably win. It’s about outscoring him. If he sits on the ropes, we can outscore him. If he stays in the middle of the ring and boxes all the time he could possibly outscore us.’ Roach claimed Pacquiao is in the best shape of his life, moving faster and punching harder than ever before. ‘He trains harder than any fighter I have ever had in my life,’ said Roach. ‘I have had 33 world champions and nobody can touch him for his work ethic. His attitude is good. ‘This fight can send out a message about what we need to do in boxing. When the best fights the best, do you see how big this is? Someone needs to wake up and put the best with the best all the time. Because when that happens, we have the best sport in the world. I don’t care who likes who, you can still negotiate business. Roach rates super middleweight Andre Ward higher than he does undefeated Mayweather . Roach, who has trained Pacquiao for 15 years, says the Filipino is in the best shape of his life . ‘It’s a better fight today than it was five years ago because they were both a lot faster and more mobile five years ago and it might have been more of a boxing match but now they’re a little bit older, it’s going to be a better fight. ‘I believe Manny can win this fight. I kind of have to win this fight. I’ve been talking a lot. It’s more important than anything to me. It’s more important to me than getting my girlfriend, Maya, back. It’s that big because, I mean, I really like this girl. ‘My mother told me, “You must like her; you put her Christmas tree up for her, you bought her a car and those earrings you bought her cost £14,000 each”. ‘After this fight maybe I’l try to get her back.’ b(x,c)x KL-DPG 1Freddie Roach, the trainer of Manny Pacquiao, says Floyd Mayweather is impossible to be considered best ever. he says he rated andre Ward and Gennady Golovkin above Mayweather despite the American’s unbeaten record. TV-DPG 1 Floyd Mayweather is fighting for Manny Pacquiao in a series and is one of the top fighters of their generation. trainer Freddie Roach says he rated super- middleweight star Andre Ward and middleweight sensation Gennady Golovkin above Mayweather despite his unbeaten record and his celebrity status as hot favourite. JS-DPG 1Freddie Roach says Floyd Mayweather can’t be ranked alongside Manny Pacquiao as leading fighters of their generation. the american’s trainer says that he rated super-middleweight star Andre Ward and middleweight sensation Gennady Golovkin above Mayweather. Table 9.Generation samples for summarization 37 Aligning Language Models with Preferences throughf-divergence Minimization source documentc The son of a Labour councillor who was detained in Turkey after apparently trying to sneak into Syria with eight family members was seen grinning as he began his journey back to Britain. Waheed Ahmed, 21, who is the son of Rochdale politician Shakil Ahmed, was arrested with eight relatives – including four children – in a remote Turkish border town earlier this month. However, it is understood he is now returning to the UK and will fly from Dalaman into Manchester Airport later this evening. Scroll down for video . All smiles: 0Waheed Ahmed 0looks relaxed as he begins his journey back to the UK after being caught trying to sneak into Syria with eight family members . On the way home: The 21-year-old, sporting a shaved head, was filmed being escorted from a vehicle . Sky News reported that the remaining eight members of his family will remain in Turkey until Tuesday. The majority of the family flew to Turkey on March 27 from Manchester Airport and are accused of having plans to try and sneak across the border into Syria. Waheed did not fly out with his family but joined them three days later on a flight from Birmingham. Mohammed Shafiq, a friend of Waheed’s father, said there were concerns about his behaviour in the months leading up to his arrest. He told Sky News: ’There were concerns in the last six months to a year about a change in his behaviour. ’And a change in his attitude towards various different issues. ’That was causing concern for people in the community and his family.’ Earlier this month, Waheed’s father spoke of his shock after being told that his son is suspected of being a militant Islamist. He said: ’All I know is that they were on holiday and then the next thing I am told is that they have been arrested.’ Mr Ahmed was with his aunt, two cousins and one of their wives when they were stopped in Turkey, near the Syrian border . Waheed Ahmed, the 21-year-old son of Labour Councillor Shakil Ahmed, is understood to be returning to the UK on a flight to Manchester from Dalaman tonight following his arrest for allegedly trying to sneak into Syria . Waheed Ahmed, 21, 0is the son of Rochdale Labour councillor Shakil Ahmed (pictured above with Ed Miliband) The nine Britons, who include three men, two women and four children aged between one and 11, were seized in Hatay province, in southern Turkey. It shares a border with part of Syria controlled by rebel factions including those linked to Al Qaeda and ISIS. All of those held are from Rochdale and are the biggest family group caught attempting to enter the unstable territory. Counter terrorism officers at Greater Manchester Police began an investigation into their movements and the extended family group were detained at a checkpoint in Ogulpinar earlier this month. A senior officer questioned why anyone would take children so young ’and vulnerable’ into a warzone. The three men and two women, aged between 21 and 47, were taken to a hospital with their children, aged one, three, eight and 11. Waheed (pictured) was detained in Turkey alongside his aunt, two cousins and one of his cousin’s wives . The nine Britons - four of them children – were seized by Turkish security forces as they tried to slip into Syria . Officials said they would be photographed and fingerprinted before being deported back to the UK. At the time, photographs showed Waheed, dressed in traditional robes and wearing heavy boots, leading the group from a minibus into a police station. Several women, all wearing headscarfs which covered their faces, could be seen carrying children. Most of the party were wearing walking boots, perfect for trekking across the rugged region. Shakil Ahmed, a bakery delivery driver, is a councillor in Kingsway and served alongside Karen Danczuk, wife of Rochdale MP Simon Daczuk, until her resignation in January. Speaking as he delivered election campaign leaflets earlier this month, he said he recognised his son in online newspaper reports of his arrest. He said the others arrested included Waheed’s aunt, Zadia Bi, 50, and two of her sons and one of their wives. He said: ’I don’t know why they have been arrested. We have no information. Until they ring we will not know what has happened.’ He said that he had seen his son’s photograph and when asked about one picture of his son laughing, he replied: ’Well, they went on holiday so they shouldn’t be crying on holiday should they?’ He added: ’I don’t believe my son was on his way to join Islamic State. I was shocked, worried and extremely upset to hear that my son has been arrested. During their arrest, the family were fingerprinted and taken to a police station where they have been held since . One of the family members, holding a child, is seen arriving at a Turkish hospital to undergo medical checks . The family are pictured arriving at a police station in Turkey’s southern Hatay province earlier this month . ’It’s a total mystery to me why he’s there, as I was under the impression he was on a work placement in Birmingham. ’My son is a good Muslim and his loyalties belong to Britain. If I thought for a second that he was in danger of being radicalised, I would have reported him to the authorities.’ The councillor added: ’He’s studying a degree in politics and sociology at Manchester University and has a good future ahead of him. I just want to speak to my son and get him home as soon as possible.’ Waheed apparently called his devastated father to break the news he had been arrested. Sorry we are not currently accepting comments on this article. b(x,c)x KL-DPG 1waheed Ahmed, 21-year-old was arrested with eight relatives in a remote northern Turkish border town this month. it is understood he is now returning to the UK and will fly from Dalaman into Manchester Airport later this evening. majority of family flew to Turkey on march 27 and are accused of having plans to try and sneak across the border into Syria. TV-DPG 1 Waheed Ahmed, 21, is the son of Rochdale Labour councillor Shakil Ahmed. he was arrested with eight relatives in a remote border town earlier this month. he is now returning to the UK and will fly from Dalaman into Manchester Airport later this evening. JS-DPG 1waheed Ahmed, 21, is the son of Rochdale politician. he was being held after apparently trying to sneak into Syria. he was arrested with eight relatives – including four children. but it is understood he is now returning to the UK. Table 10.Generation samples for summarization 38