Paper deep dive
Control Barrier Function for Aligning Large Language Models
Yuya Miyaoka, Masaki Inoue
Models: Llama 3 8B, RoBERTa (cardiffnlp/twitter-roberta-base-sentiment-latest)
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 5:32:07 PM
Summary
The paper introduces 'CBF-LLM', a control-based framework for aligning Large Language Models (LLMs) using Control Barrier Functions (CBF). By implementing a safety filter as an add-on component, the framework intervenes in the token generation process to ensure outputs remain within a 'safe' or 'desirable' set without requiring fine-tuning of the baseline LLM. The system uses a language-constraint function (L-CF) derived from models like RoBERTa to evaluate text desirability and applies CBF inequalities to filter token distributions, with extensions for multi-step ahead prediction.
Entities (5)
Relation Signals (4)
CBF-LLM → utilizes → Control Barrier Function
confidence 100% · This paper proposes a control-based framework for aligning large language models (LLMs) by leveraging a control barrier function (CBF)
CBF-LLM → implementedwith → Llama-3
confidence 95% · In this paper, CBF-LLM is implemented with Llama 3
Safety Filter → intervenesin → Token Generation
confidence 95% · The presented framework applies the CBF safety filter to the predicted token generated from the baseline LLM
RoBERTa → servesas → Language-Constraint Function
confidence 90% · We apply a sentiment analysis RoBERTa model as the internal model of the L-CF.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper proposes a control-based framework for aligning large language models (LLMs) by leveraging a control barrier function (CBF) to ensure user-desirable text generation. The presented framework applies the CBF safety filter to the predicted token generated from the baseline LLM, to intervene in the generated text. The safety filter includes two significant advantages: this safety filter is an add-on type, allowing it to be used for alignment purposes without fine-tuning the baseline LLM, and if there is an evaluation model regarding the desired alignment, it can be directly applied to the filter design. The overall text-generation system is implemented with open-source language models, aiming to generate positive text.
Tags
Links
- Source: https://arxiv.org/abs/2511.03121
- Canonical: https://arxiv.org/abs/2511.03121
Trouble viewing inline? Open PDF directly →
Full Text
60,334 characters extracted from source content.
Expand or collapse full text
1 Control Barrier Function for Aligning Large Language Models Yuya Miyaoka, Masaki Inoue, Member, IEEE Abstract—This paper proposes a control-based framework for aligning large language models (LLMs) by leveraging a control barrier function (CBF) to ensure user-desirable text generation. The presented framework applies the CBF safety filter to the pre- dicted token generated from the baseline LLM, to in- tervene in the generated text. The safety filter includes two significant advantages: this safety filter is an add- on type, allowing it to be used for alignment purposes without fine-tuning the baseline LLM, and if there is an evaluation model regarding the desired alignment, it can be directly applied to the filter design. The over- all text-generation system is implemented with open- source language models, aiming to generate positive text. Index Terms—Control Barrier Function, Large Lan- guage Models, Alignment, Safe Control I. Introduction While large language models (LLMs) are known to have strong language understanding, reasoning and writing abilities, they can also generate harmful, biased, toxic, or unethical content [1], [2]. Alignment of LLMs ensures that they generate content that is “desirable” for user, meaning that the content is ethical and safe. Various approaches for LLM alignment have been presented (see the literature [1], [2], [3] and reference therein). The major approach to LLM alignment is reinforce- ment learning from human feedback (RLHF, [4]), where a reward model is constructed by human feedback and then used for the training of LLMs. Variants of RLHF methods are also proposed, such as Safe-RLHF by [5], SENSEI by [6], and f-DPG by [7], and their implemen- tations are presented, such as training pre-trained LLMs [8], [9]. Collecting human feedback with data is time- consuming and expensive. To overcome this drawback, reinforcement learning from AI feedback (RLAIF) instead of using human feedback is presented by [10]. Furthermore, to reduce the computational cost and enhance the stability of training, direct preference optimization (DPO) is pro- posed by [11], where the alignment data is directly used for training LLMs without accessing the reward model. A common feature of alignment methods like RLHF, variants of RLHF, and DPO is that they update LLMs’ model parameters. An alternative approach for LLM alignment is to directly intervene in the input prompts, rather than updating the model parameters. In-context learning (ICL, This work was supported by the Grant-in-Aid for Scientific Re- search (B), No. 20H02173 and 25K01254 from JSPS. Y. Miyaoka and M. Inoue are with the Department of AppliedPhysicsandPhysico-Informatics,KeioUniversity, Yokohama223-8522,Japan(e-mail:miyaoka.yuya@keio.jp, minoue@appi.keio.ac.jp). Obstacle Unsafe Safe “Obstacle” Undesirable Desirable Vehicle LLM Everyone says you will be a good researcher, but this is not true, because Control Control Fig. 1. Safe Control of LLM. Top: Collision avoidance in a vehicle control system, Bottom: Col lision avoidance in text-generation by LLMs. [12]) is a major approach for intervening in the input prompt. In ICL, a few pairs of input prompts and output are provided as demonstrations to instruct the LLMs on the task [13], [14]. The intervention-based approach to LLM alignment, which focuses on preventing undesirable outputs, is anal- ogous to the problem of collision avoidance. The collision avoidance task is the most fundamental control problem in systems engineering.Just as the vehicle’s trajectory is intervened to avoid collisions, LLM’s output can be inter- vened to prevent undesirable content. This paper draws an analogy between a vehicle and LLM, as illustrated in Fig. 1. Consider the LLM as an analogy to a vehicle, and the generated text as an analogy to the vehicle’s trajectory. Both vehicle collision avoidance and LLM alignment aim to guide the complex system away from undesirable states by designing proper control strategies. Vehicle collision avoidance aims to prevent collisions with obstacles by intervening in the vehicle’s trajectory. Similarly, LLM alignment aims to prevent undesirable outputs, such as harmful and unethical content. To this end, we intervene in the text-generation system to generate the desirable trajectory of the token sequence. In the control community, various studies are conducted for safety assurance of control systems, including collision avoidance [15], [16]. A promising approach for collision avoidance and safe control is the control barrier function (CBF), as studied in [17], [18], [19], [20]. The theory for CBF is extended to robust CBF [21], multiple CBF [22], reinforcement learning integration [23], CBF for time invarying system [24], and CBF for stochastic systems [25], [26]. In addition to various theoretical advancements, CBF has been applied in diverse areas, such as robotics [27], [28], [29], [30], [31], vehicle control [32], battery manage- arXiv:2511.03121v2 [cs.CL] 6 Nov 2025 2 ment [33], and learning-based system [23]. For example, the works on robotics [27], [28] address the human-machine shared control, ensuring the safety of the human space. This paper is the first to apply the powerful theory of CBF to LLM alignment. This paper proposes a framework for control-based LLM alignment by applying a safety filter that intervenes in the LLM output to generate user-desirable outcomes. To this end, we leverage CBF to improve the safety and controllability of the output of LLMs. We aim to design a CBF-based safety filter that intervenes in the output of LLMs into the user’s desired content. The CBF filter and the baseline LLM constitute a controlled text-generation system, which we call “CBF-LLM”. The contributions of this paper are as follows: 1) While theoretical analysis of LLMs has been studied in the works by [34] and [35], their design methodol- ogy has not. This paper develops a design approach for LLMs through the lens of control theory. 2) CBF-based safe control has been successfully applied in various domains. This paper introduces a novel application of CBF-based safe control to the domain of LLMs, aiming to enhance the safety and reliability of text-generation systems. 3) The proposed system, CBF-LLM, is realized in an add-on manner to a baseline LLM: an external filter is simply added to the LLMs without accessing and updating their baseline LLM parameters. In this sense, CBF-LLM is broadly applicable to various alignment goals and various LLMs. 4) In this paper, CBF-LLM is implemented with Llama 3 [36], an open-source LLM developed by Meta, and a RoBERTa model in Section IV. The rest of this paper is as follows. In Section I, the basic theory of CBF and the text-generation system using LLM are reviewed. In Section I, the concept and design of CBF-LLM are proposed. In Section IV, the implementation of CBF-LLM is presented and the text- generation experiment is conducted. Finally, in Section V, the conclusion of this paper is presented. Notation: symbol V [i] represents the i-th element of vector V . I. Preliminaries A. Control Barrier Function for Safe Control Control barrier function (CBF), developed in the control community, provides safety assurance in control systems [17], [18]. This subsection briefly reviews CBF and CBF- based safe control. Consider the following dynamical system to be con- trolled: ̇x = g(x,u),(1) where x ∈R n is the state variable of the control target, and u ∈R m is the control input applied to the target, and g is a nonlinear function that represents the system dynamics. We aim to design the assisted control system with safety assurance. Specifically, we focus on achieving safe control by filtering nominal control inputs and modifying them into safe control inputs. We let the safe and unsafe sets be denoted by S ⊆R n , and ̄ S =R n ⊆R n , respectively. Then, the safety means to constrain the state x within the safe setS, i.e., x∈S. We address the problem of designing the following filter F :R m →R m : Problem 1 (Safety Filter). Assume that the system (1) starts in a safe state x(τ 0 )∈S at time τ 0 . Given a nominal control input u nom ∈R m , find the safety filter F :R m × R n →R m such that the system (1) with u = F (u nom ,x) generates x(τ )∈S for all time τ ≥ τ 0 . As a preliminary, we design a continuous function h : R n →R, called a “constraint function”, such that h(x)≥ 0 if x ∈ S, and h(x) < 0 if x ∈ ̄ S holds. The safety is equivalent to the constraint: h(x) ≥ 0. The safety filter F needs to modify u nom to output u such that h(x) ≥ 0 holds thereafter. To construct the safety filter F that keeps the safety con- straint h(x) ≥ 0, the control barrier function filter (CBF filter, [17], [18]) is presented. The CBF filter intervenes in the nominal control input u nom to introduce a safe state of the object by finding u as follows: F : ( min u (u nom − u) 2 , s.t. ̇ h(x,u)≥−α c (h(x)), (2a) (2b) where α c :R →R is a class-K function which holds α c (0) = 0 and is monotonically increasing. The function h is called a control barrier function if there exists u such that the constraint (2b) holds. To constrain the control input by the CBF filter, the following theorem on safety assurance holds: Theorem 1 ([17]). Suppose that x(τ 0 ) ∈ S, and the control input u = u(τ ) satisfies the CBF constraint (2b) for all τ ≥ τ 0 . Then, the state x(τ )∈S holds for all τ ≥ τ 0 . The objective function (2a) ensures that the filtered control input u remains close to the nominal control input u nom . In this sense, the CBF filter achieves safety by the “minimum” intervention. The CBF filter is capable of applying in discrete-time systems by re-formulating the CBF constraint (2b) as follows [37]: ∆h(x(k),x(k + 1)) =h(x(k + 1))− h(x(k))≥−α d (h(x(k))), (3) where k is a discrete time, and α d :R →R is a class-K function that satisfies 0≤ α d (h(x))≤ h(x) for all x. In this paper, we let the function α d (h(k)) be a linear function, denoting α d (h(k)) := αh(k), where α ∈ [0, 1] is a hyperparameter. Further letting γ := 1− α, the CBF constraint (3) is rewritten as: h(x(k + 1))≥ γh(x(k)).(4) 3 B. Large Language Models for Text Generation This subsection briefly reviews and analyzes the text generation by large language models (LLMs). In this paper, “text” means the sequence of tokens, andX denotes the set of all texts. For example, we let x =“It is a nice day.”, meaning x is composed of six tokens: “It”, “is”, “a”, “nice”, “day”, and “.”. Each token t is identified by a positive integer, i.e., t ∈ 1,...,N := T , here N is the number of tokens the LLM has. The operator ⊕ concatenates text and token to output a text. For example, “It is a”⊕“nice”=“It is a nice”. Text generation by LLMs is performed by iteratively adding a new token while considering all given/generated text. Let x 0 ∈ X be the initial text, and k ∈ 0, 1,... be the discrete time, which counts the number of tokens added during the generation. Then, the text generation from the initial text x(0) = x 0 is expressed by a following discrete-time dynamical system: P (k) = G(x(k)), t ∗ (k) = C(P (k)), x(k + 1) = x(k)⊕ t ∗ (k). (5a) (5b) (5c) In the system description, the symbol G : X → (0, 1) N represents the token predictor, which is driven based on the input text x to output the token distribution vector P ∈ (0, 1) N . Each element of P , denoted by P [t],t ∈ T displays how probable that the token t∈T can be followed by the text x. The symbol C is the token selector, which selects the next token t ∗ that follows x based on the token distribution vector P . In this context, the text x acts as the state value of the text-generation system (5). As a new token t ∗ is added, the time k proceeds and the state x is updated, i.e., x(k) transitions to x(k + 1). This is because the probability distribution P for the next token is determined by the entire history of the text x(k), which includes both the initial text and all previously added tokens. Note that if we denote the number of tokens in x 0 as k 0 , the total number of tokens in x(k) is k 0 + k. The token predictor G is a combination of a so-called an LLM and the softmax processing. Examples of LLMs include GPT-2 [38] and Llama 3 [36]. The token distribu- tion vector P is derived by applying the softmax function to the LLM’s output, as follows: G : P [t] = softmax(LLM(x)[t]/T ) = exp(LLM(x)[t]/T ) P N t ′ =1 exp(LLM(x)[t ′ ]/T ) , t∈T , (6) where LLM :X →R N is the LLM, and T > 0 is a hyperpa- rameter called temperature. The token distribution vector P has elements within the range (0, 1) and its elements sum up to 1, while the pure output of the LLM does not. There are various methods for token selection, including greedy search and multinomial sampling. Greedy search selects a token with the highest probability in P , i.e., t ∗ = arg max t P [t], and the text-generation system (5) with greedy search is a deterministic system. Multinominal 푍 !" Token Predictor 퐺 Token Selector 퐶 Next Token 푡 ∗ 푃∈(0,1) ! Text 푥 Text 푥 ⊕ Fig. 2. Text-generation system sampling randomly selects a token based on the given token distribution P , i.e.,P t ∗ ∼C(P) [t ∗ = t] = P [t], and the text-generation system (5) with greedy search is re- duced to a stochastic system. In this paper, we employ multinomial sampling as the token selector. The overall structure of the text-generation system (5) is shown in Fig. 2. In this figure, the blocks of G, C, and Z −1 represent the token predictor, token selector, and time delay, respectively. I. LLM wigh CBF-Based Safety Filter This section presents the control-based alignment of text-generation systems and their detailed implementa- tion. The LLM alignment discussed in this paper aims to ensure desirable text generation by weak intervention to the output of the token predictor G in text-generation system (5). To clarify the meaning of “desirable”, we let the desirable and undesirable text sets be S ⊆ X and ̄ S = X \ S ⊆ X , respectively, based on the respective alignment goals. Example 1. Consider that the alignment goal is set to generate texts with positive contexts. Then, S is the set of positive-context texts, and ̄ S is the set of non-positive- context texts. The details are seen in Section IV. The presented text-generation system, including an LLM and a safety filter for alignment, is constructed based on the CBF described in Subsection I-A. The overall system is called CBF-LLM and its structure is shown in Fig. 3. CBF-LLM extends the nominal text-generation system shown in (5) by adding the CBF filter (green box) between the token predictor G and the token selector C. The CBF filter manipulates the token distribution P to satisfy the user-defined alignment goal. In the same manner as (5), the blocks of G, C, and Z −1 represent the token predictor, token selector, and time delay, respectively. A. Language-Constraint Function The defining feature of the CBF-LLM is the presence of CBF filter F , which filters P to generate the modified 4 푍 −1 Token Predictor 퐺 Token Selector 퐶 Next Token 푡 ∗ 푃∈(0,1) 푁 Text 푥 Text 푥 푄∈[0,1] 푁 ⊕ CBF Filter 퐹 Fig. 3. Proposed text-generation system with CBF filter (CBF-LLM) token distribution Q ∈ [0, 1] N . The CBF filter F is designed by using the function h :X →R satisfying ( h(x)≥ 0, x∈S, h(x) < 0, x∈ ̄ S. (7) The function h is called the “language-constraint function” (L-CF), whose role is the same as the constraint function presented in Subsection I-A. Note that the L-CF h needs to distinguish between the desired and undesired texts accurately. Furthermore, the value of L-CF is expected to indicate how desirable or undesirable the text is. Generally, building the L-CF that perfectly distin- guishes between desirable and undesirable texts is chal- lenging. One can construct the L-CF using existing text classification models. An example of L-CF is provided as follows. Example 2. We apply a sentiment analysis RoBERTa model 1 as the internal model of the L-CF. The RoBERTa model is originally trained to classify the context of given texts into three labels: negative, natural, or positive. Let [s − (x(k)) s ± (x(k)) s + (x(k))] ⊤ ∈ [0, 1] 3 denote the softmax output of the RoBERTa model with respect to a given text x. It follows that s − (x(k)), s ± (x(k)), and s + (x(k)) represent the score of negative, neutral, and positive, respectively. Then, the L-CF h is constructed as follows: h(x(k)) = s + (x(k))− max(s − (x(k)),s ± (x(k))).(8) This L-CF h outputs a positive value when the positive score is greater than both negative and neutral scores, while it outputs a negative value when either the negative or neutral score is greater than the positive score. In other words, the sets S and ̄ S, which correspond to the L-CF constructed above, render the texts with positive and non- positive contexts, respectively. Note that the RoBERTa model used in this L-CF is originally trained for evaluating whole texts, while it is used with mid-texts in this example. This may deteriorate the accuracy of the evaluation. For example, the model accurately evaluates the text input “It 1 cardiffnlp/twitter robertabasesentimentlatest [39] is a nice day.”, but may not accurately evaluate the text input “It is a”. B. CBF Filter and CBF-LLM The CBF filter F : (0, 1) N → [0, 1] N allows only tokens that meet its conditions to pass through and does not allow tokens that do not. The detailed realization of the CBF filter F is given as follows: F : Q(k)[t]∝ ( P (k)[t], h(x(k)⊕ t)≥ γh(x(k)) 0,else, t∈T , (9) where γ ∈ [0, 1] is a hyperparameter. This formulation is a modified form of the discrete-time CBF inequality, as shown in (4). In (9), the probability of the token is set to 0 unless the token satisfies the following CBF inequality: h(x(k)⊕ t)≥ γh(x(k)),(10) which guarantees that the generated text x always satisfies that x∈S. Remark 1. The CBF inequality (10) not only assesses whether x⊕ t,t ∈ T is undesirable or desirable but also takes into account how much the value of h(x⊕t) moves in the negative direction compared to the value of h(x). Even if x⊕t is desirable, i.e., h(x⊕t)≥ 0, if the value of h(x⊕t) decreases significantly compared to h(x), the probability of the token t is set to 0. This conservative behavior aims to exclude not only tokens that immediately become undesirable but also tokens that “give the conversation a dubious tone”. The output of the CBF filter, Q, needs to be normalized, and its elements sum up to 1. In addition, the CBF filter aims at user-desired text generation by weak interventions to the original token distribution P . To this end, we formulate the following optimization problem: min Q D KL [Q||P (k)], s.t. Q≥ 0, 1 ⊤ Q = 1, P t∼C(Q) [h(x(k)⊕ t)≥ γh(x(k))] = 1, (11a) (11b) (11c) (11d) where D KL [Q||P ] is Kullback-Leibler divergence between Q and P . The constraints (11b) and (11c) are imposed because Q needs to be a distribution vector. The constraint (11d) requires that the t selected based on Q always satisfy the CBF constraint (10). The constraint is mathematically severe, and can be relaxed by replacing the probability of one with δ, where δ ∈ [0, 1) is an acceptable failure tolerance. The CBF filter (9) is provided as follows: Theorem 2. The minimizer Q(k) of the optimization problem (11) is provided by following: 5 P ′ (k)[t] = ( P (k)[t], The constraint (10) holds, 0,else, Q(k)[t] = P ′ (k)[t] P N i=1 P ′ (k)[i] , t∈T . (12a) (12b) The proof of the theorem is given in Appendix A. Top-K sampling is applied in the CBF filter F to improve the computational efficiency. The top-K sampling only processes fewer elements than N elements of the target P . The algorithm of the CBF filter with top-K sampling is shown in Algorithm 1. Algorithm 1 CBF-LLM with top-K sampling Require: P ∈ (0, 1) N : token distribution from the token predictor G. Require: x∈X : current text. Require: h :X →R : the constraint function (7). Require: γ ∈ [0, 1] : CBF’s hyperparameter. Require: K : the top-k parameter. Initialization 1: A ← φ : allowed set, a set of tokens that satisfy the CBF inequality. 2: P ′ ∈ [0, 1] N ← 0 N 3: T ∈ 1,...,N N ← argsort(P ) : sort the indexes of P in descending order. Collect K allowed tokens. 4: i← 1 5: while |A| < K do 6: t← T [i] : the token with i-th highest probability. 7: if h(x ⊕ t) ≥ γh(x) : CBF inequality (10) holds, then 8: A ← A∪t : append the token to the allowed set. 9: P ′ [t]← P [t] 10: end if 11: i← i + 1 12: end while Make a modified token distribution. 13: Q = P ′ / P t∈T P ′ [t] 14: Select the next token t ∗ based on Q. Note that the text generation is done by iteratively per- forming the Algorithm 1 C. Extension to Multi-Step Ahead Method A drawback of the CBF-LLM system presented in Fig. 3 is that the intervention strategy may be too conservative: it may filter out texts that appear undesirable initially but become desirable when read to the last, e.g., “You are clumsy, but you have high aspirations!”. To overcome this drawback of the CBF-LLM system with single-step ahead token prediction, we extend the CBF-LLM system with multi-step ahead token prediction. Let H ∈ 1, 2,... denote the prediction horizon, and further let y denote the sequence composed of H tokens, i.e., y = [“you”,“are”,...,t H ]. At each time, the multi- step ahead method collects K ∈ 1, 2,... candidates of H-token sequences y 1 ,y 2 ,...,y K generated from the baseline LLM that continues from x(k) such that the CBF inequality (10) holds, i.e., h(x⊕ y)≥ γh(x). The next H- token sequence is selected from these candidates according to the distributionP[y | x(k)] derived by the baseline LLM. The probability of the presence of an H-token sequence y = [t 1 ,...,t H ] from x is as follows: P[y | x] = Π H−1 h=0 G(x⊕ t 1 ⊕·⊕ t h )[t h+1 ].(13) The detailed pseudo-code is presented in Algorithm 2. Algorithm 2 CBF-LLM with multi-step ahead Require: G :X → (0, 1) N : the token predictor. Require: x∈X : current text. Require: h :X →R : the constraint function (7). Require: γ ∈ [0, 1] : CBF’s hyperparameter. Require: H ∈1, 2,... : the prediction horizon. Require: K : sample size. Initialization 1: Y ← [ ] : list of candidate H-token sequences. 2: Q← [ ] : list of probability of each candidate H-token sequence. 3: k ← 0 : counter of candidate H-token sequences. Collect H-token sequences as candidates. 4: while k < K do 5: y ←“” : empty text. 6: q ← 1 7: for H times do 8: t ∗ ← select the next token based on G(x⊕ y). 9: q ← q G(x⊕ y)[t ∗ ] 10: y ← y⊕ t : extend the candidate sequence. 11: end for 12: if h(x⊕ y) ≥ γh(x) : CBF inequality (10) holds, then 13: Y.append(y) : now y has H tokens. 14: Q.append(q) : now q is the probability that y would follow x, i.e., q =P[y | x]. 15: k ← k + 1 : increment the counter. 16: end if 17: end while Extend the text. 18: Select the next H-token sequence from its candidates Y based on Q. D. Previous Works There are various works addressing alignment by inter- vening in the output of the LLM, either from a repre- sentational perspective [40], [41] or a semantic one [42], [43], [44], [45], [46], [47]. In particular, the work [43] proposes an intervention-based alignment method with multi-step ahead prediction, named Blockwise best-of-K method. The method selects the best of K candidates of H-token sequences generated from the baseline LLM 6 without imposing any constraints. The common feature of these previous works is to intervene in the token proba- bility based on a reward of generated text. On the other hand, CBF-LLM, our proposed system, intervenes in the token probability based on the change in reward caused by adding a new token, rather than just the generated text. IV. Experiments In this section, we implement the CBF-LLM with Llama 3 and a RoBERTa model and verify the CBF- LLM’s alignment ability, the number of interventions, generation time, and output quality. In the experiments, we commonly employ Llama 3 8b [36], a pre-trained LLM, as the model for the token predictor G. A. Positive Text Generation The alignment goal is to ensure that the CBF-LLM system, illustrated in Fig. 3, produces texts with “positive” contexts. To this end, we let S and ̄ S denote the set of positive texts and non-positive texts, respectively. We employ a RoBERTa model 2 to construct the L-CF h. This RoBERTa model was originally trained to classify sentences into three labels: negative, neutral, or positive. The L-CF outputs a positive value when the sentiment of the text x is positive. The detail is presented in Example 2 in Section I. The resulting text-generation system would be controlled to generate positive content. We use the Reddit dataset, reddit-corpus-small [48] to collect the initial texts to be input for the text- generation system. From the Reddit dataset, we randomly chose 50 utterance texts that satisfy the following three conditions: 1) The text has more than 10 tokens. 2) The text in which the L-CF h indicates positive for the first 5-token text. 3) The text in which the L-CF h indicates negative for the generated text by the original Llama 3 model without any control gives the first 5-token text. In other words, the text of the first 5-token potentially results in a non-positive generation. We extract only the first 5 tokens from the selected utterance texts and use them as the initial texts x 0 . We set the temperature as T = 1, the top-K value as K = 30, and the maximum number of new tokens as 30. In the experiments with text generation by CBF-LLM, as shown in Fig. 3, we varied the hyperparameter γ from 0.0 to 1.0. We also implemented “No Intervention” case as a comparative baseline. The No Intervention case performs the nominal text generation shown in (5), without any intervention filters. The trajectory of L-CF h(x(k)),k ∈ 1, 2,... for a generated text sample is shown in Fig. 4. In the No Intervention case (black line), the generated text does not keep the positive L-CF value, implying the extent to which the generated text is undesirable. On the other 2 cardiffnlp/twitter robertabasesentimentlatest 051015202530 Time k −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 L-CF Value h ( x ( k )) that part . That 's what I 'm afraid of . Not much I can do . But if it 's the right thing to do then I have to this one . I 've seen far too man u 3 D games so I hope that when this one is released that it will be an option , this one . I ’ve read a lot of his stuff ( as well as some of his detr act o ’s ). I ’ve read and watched ( and that - that was my thinking also on the previous post . I was certainly in favor of our getting a good D -man and a goalie at the draft almost every point . I 'd go to # god bless you , # fant astic story . # lo vely # fun # e aster # e a . No Intervention CBF(γ= 0.0) CBF(γ= 0.2) CBF(γ= 0.6) CBF(γ= 1.0) Fig. 4. L-CF trajectory of each filter hand, in CBF filters, the L-CF values are kept positive during the generation, implying that the text generation system generates desirable content. Specifically, for the case γ = 1.0 (red line), the L-CF h(x(k)) monotonically increased as k progressed. This is because the generation was performed under the strict constraint that the CBF condition is written as h(x(k + 1)) ≥ h(x(k)), requiring the next text to be even more “positive” than the current one. In Fig. 4, the initial text x 0 was “I agree with you on”. The generated texts are as follows: 3 • No Intervention: I agree with you on that part. That’s what I’m afraid of. Not much I can do. But if it’s the right thing to do then I have to • CBF(γ = 0.0): I agree with you on this one. I’ve seen far too manu 3D games so I hope that when this one is released that it will be an option • CBF(γ = 0.2): I agree with you onthis one. I’ve read a lot of his stuff (as well as some of his detracto’s). I’ve read and watched (and • CBF(γ = 0.4): I agree with you on some of the above, especially #4. I’ve been thinking a lot about the church lately. I’ve been attending St. Joseph’s parish. • CBF(γ = 1.0): I agree with you onalmost every point. I’d go to #god bless you, #fantastic story. #lovely #fun #easter #eaa. In the No Intervention case, we observe non-positive contexts in the generated texts (red text), whereas in the CBF cases, the method eliminates non-positive contexts. However, when γ is high, for example γ = 1.0, the output text may contain expressions that appear unnatural in natural language, such as “#eaa.”. The predicted possible trajectories of the CBF filter is shown in Fig. 5. In CBF-LLM, the CBF filter sorts tokens into those that satisfy the CBF inequality (10) and those that do not. The figure shows that the CBF 3 The generated texts have a typo, showing “manu” where it should be “many”. However, this is the direct output from the text- generation system and has been listed without modification. 7 051015202530 Time k −0.2 0.0 0.2 0.4 0.6 0.8 Constraint Function Value h ( x ( k )) Allowed Token Disallowed Token Generated Trajectory Fig. 5. Predicted L-CF trajectory filter prevents L-CF values from becoming negative of decreasing more rapidly than the current value. Note that we do not show the trajectories for all tokens, but only for tokens investigated by top-K sampling are displayed. We evaluated the generated texts from various aspects. Recall that this paper aims to ensure desirable text gener- ation by weakly intervening in the output of LLMs. To evaluate the intervention weakness, we use the number of disallowed tokens per generation and naturalness. The average naturalness is evaluated by G-Eval [49]. Addition- ally, the positiveness of the generated texts, which was the control objective of this experiment, was assessed using G-Eval. G-Eval is a method for evaluating the quality and task compliance of generated texts by using other LLMs. The prompts used for evaluating naturalness and positiveness are shown in Appendix B. Furthermore, to compare the performance of text generation, we measure the average time the text-generation system takes to generate one token. The results are shown in Table I. In the intervention cases, the naturalness score was the highest at γ = 0.4, the positiveness score was the highest at γ = 1.0, and the number of the disallowed tokens was the lowest at γ = 0.2. These results reveal a trade-off between task quality, displayed by positiveness, and intervention weak- ness, displayed by naturalness and disallowed tokens, and the trade-off is calibrated by the hyperparameter γ. The “Time per Token” column in the table shows the average time the system takes to generate one token. The text- generation system with CBF is about 0.03 seconds longer than the baseline system. This is due to the evaluation of the L-CF h taking time. The number of disallowed tokens at γ = 1.0 becomes significantly larger than other γ. This is because the CBF constraint at γ = 1.0 was strict, and few word sequences satisfy the constraint. Based on the results on naturalness, positiveness, the number of disallowed tokens, and the generation time, we see that the hyperparameter γ should be chosen from within (0, 1) rather than choosing 1, which is equivalent to the Blocklist. B. Extension to Multi-Step Ahead CBF-LLM We conduct an additional experiment to demonstrate the text generation by CBF-LLM with the multi-step ahead method. In this experiment, we compare the multi- step ahead CBF-LLM and the Blockwise best-of-K [43]. The alignment goal, the baseline LLM, the L-CF, and the initial texts used in this experiment are the same as those in Subsection IV-A. We evaluated the naturalness by G-Eval, the rate of generated texts that were not positive, and the generation time. The results are shown in Table I. At each element, the left, middle, and right side values display the non- positive generation rate, naturalness, and the average time the text-generation system takes to generate one token, respectively. We can see that the naturalness is higher in the multi-step ahead CBF-LLM compared to the Blockwise best-of-K method. Notably, when the sample size K is small, the Blockwise best-of-K had a relatively high rate of non-positive text generation. The Blockwise best-of-K method does not disallow the user-undesired outputs, which may lead to user-undesired results. In contrast, the multi-step ahead CBF-LLM did not produce any undesirable text for any value of K due to the safety filter. Given the practical need to reduce K due to some reason, such as computational efficiency, the multi-step ahead CBF-LLM has potential in scenarios where avoiding undesirable text is guaranteed. The multi-step ahead CBF-LLM also revealed chal- lenges. Specifically, it takes longer generation time than the single-step version of CBF-LLM, presented in Subsec- tion IV-A. This is because the generation process requires producing a larger number of candidate token sequences. V. Concluding Remarks This paper proposed the control-based LLM alignment framework, called CBF-LLM. This framework utilizes the control barrier function (CBF) to ensure the safety of physical objects, such as the collision avoidance function in assisted driving vehicles. Based on an analogy between the control theory and the LLM alignment task, we em- ployed the CBF-based safety filter to ensure that the text- generation system generates desirable content. The key feature of CBF-LLM is that the CBF filter can be attached to the baseline LLM in an add-on manner: it intervenes in the output of the baseline LLM without any additional training of LLMs. This paper also presented the implemen- tation of CBF-LLM by Llama 3 and a sentiment analysis RoBERTa model to ensure that the text-generation system generates positive content. The text-generation experi- ment showed that CBF-LLM outperforms the baseline method in terms of naturalness, positiveness, generation time, and the reliability of controlled decoding. For the evaluation of the generated text quality, our experiments employed a method relying on another LLM. Future work should prioritize using more objective evaluation method- ologies that reduce ambiguity. 8 TABLE I Evaluation Results Number of disallowed tokens per generation NaturalnessPositivenessTime per Token [s] CBF(γ = 1.0)1.05× 10 4 0.3240.6600.548 CBF(γ = 0.8)5750.3270.3850.140 CBF(γ = 0.6)3350.3470.3500.135 CBF(γ = 0.4)4690.3640.3630.137 CBF(γ = 0.2)2990.3330.3560.135 CBF(γ = 0.0)3680.3580.3680.141 No Intervention00.3750.2420.114 TABLE I Non-Positive Generation Rate, Naturalness, and Time per Token Sample Size K Multi-Step Ahead CBF-LLM (H = 3,α = 0.8) Blockwise best-of-K [43] (H = 3) 20.00/0.722/0.157 s0.30/0.592/0.309 s 40.00/0.718/0.333 s0.02/0.695/0.655 s 50.00/0.727/0.419 s0.00/0.701/1.05 s In CBF-LLM, the key challenge is to effectively incor- porate human feedback and existing evaluation models to reflect human preferences into the L-CF. The design of L-CF is similarly challenging to construct a high-quality reward model in RLHF approaches. The value of CBF filters lies in their ability to facilitate easy modifications. To illustrate this, we show two scenarios: In a scenario, consider that an aligned LLM is developed and integrated into a service system. Suppose that an ethical or other critical issue is discovered with the original data used for alignment. Then, it becomes challenging to remove the influence of the data from the LLM using RLHF- based methods, such as unlearning [50]. This can lead to the suspension of the service system. In contrast, in CBF-LLM, an add-on type alignment method, we can simply disable the CBF filter to maintain the service operation while modifying the specifications. In the other scenario, consider an LLM initially trained or controlled to produce positive text. Later, suppose that an additional requirement is added such as ensuring that the generated text is easy for children to comprehend. In CBF-LLM, we can independently design a readability CBF filter without modifying the existing positivity CBF filter, allowing the system to meet the updated requirements without having to retrain the entire LLM. This approach enables us to easily adapt to changing specifications and requirements. In these scenarios, the CBF-LLM approach offers a signif- icant advantage. References [1] T. Shen, R. Jin, Y. Huang, C. Liu, W. Dong, Z. Guo, X. Wu, Y. Liu, and D. Xiong, “Large Language Model Alignment: A Survey,” arXiv preprint arXiv:2309.15025, 2023. [2] S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao, “Large Language Models: A Survey,” arXiv preprint arXiv:2402.06196, 2024. [3] Y. Wang, W. Zhong, L. Li, F. Mi, X. Zeng, W. Huang, L. Shang, X. Jiang, and Q. Liu, “Aligning Large Language Models with Human: A Survey,” arXiv preprint arXiv:2307.12966, 2023. [4] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schul- man, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training Language Models to Follow Instructions with Human Feed- back,” Advances in Neural Information Processing Systems, vol. 35, p. 27730–27744, 2022. [5] J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang, “Safe RLHF: Safe Reinforcement Learning from Human Feedback,” in The Twelfth International Conference on Learning Representations, 2024. [6] R. Liu, G. Zhang, X. Feng, and S. Vosoughi, “Aligning Gen- erative Language Models with Human Values,” in Findings of the Association for Computational Linguistics: NAACL 2022, (Seattle, United States), p. 241–252, Association for Compu- tational Linguistics, July 2022. [7] D. Go, T. Korbak, G. Kruszewski, J. Rozen, N. Ryu, and M. Dymetman, “Aligning Language Models with Pref- erences through f-divergence Minimization,” arXiv preprint arXiv:2302.08215, 2023. [8] Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. Das- Sarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. El- hage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. John- ston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan, “Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback,” arXiv preprint arXiv:2204.05862, 2022. [9] C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. YU, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy, “LIMA: Less Is More for Alignment,” in Advances in Neural Information Processing Systems (A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds.), vol. 36, p. 55006–55021, Curran Associates, Inc., 2023. [10] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Lan- dau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lan- ham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan, “Consti- tutional AI: Harmlessness from AI Feedback,” arXiv preprint arXiv:2212.08073, 2022. [11] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Er- mon, and C. Finn, “Direct Preference Optimization: Your Lan- guage Model is Secretly a Reward Model,” in Advances in Neural Information Processing Systems (A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, eds.), vol. 36, p. 53728–53741, Curran Associates, Inc., 2023. [12] Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, 9 Z. Wu, T. Liu, B. Chang, X. Sun, L. Li, and Z. Sui, “A Survey on In-context Learning,” arXiv preprint arXiv:2301.00234, 2024. [13] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Ka- plan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sas- try, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learn- ers,” in Advances in Neural Information Processing Systems (H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, eds.), vol. 33, p. 1877–1901, Curran Associates, Inc., 2020. [14] H. Zhao, M. Andriushchenko, F. Croce, and N. Flammarion, “Is In-Context Learning Sufficient for Instruction Following in LLMs?,” arXiv preprint arXiv:2405.19874, 2024. [15] C. Dawson, S. Gao, and C. Fan, “Safe Control With Learned Certificates: A Survey of Neural Lyapunov, Barrier, and Con- traction Methods for Robotics and Control,” IEEE Transac- tions on Robotics, vol. 39, no. 3, p. 1749–1767, 2023. [16] A. K. Kiss, T. G. Molnar, D. Bachrathy, A. D. Ames, and G. Orosz, “Certifying Safety for Nonlinear Time Delay Systems via Safety Functionals: A Discretization Based Approach,” in 2021 American Control Conference (ACC), p. 1058–1063, 2021. [17] A. D. Ames, S. Coogan, M. Egerstedt, G. Notomista, K. Sreenath, and P. Tabuada, “Control Barrier Functions: The- ory and Applications,” in 2019 18th European Control Confer- ence (ECC), p. 3420–3431, 2019. [18] T. Gurriet, A. Singletary, J. Reher, L. Ciarletta, E. Feron, and A. Ames, “Towards a Framework for Realizable Safety Critical Control through Active Set Invariance,” in 2018 ACM/IEEE 9th International Conference on Cyber-Physical Systems (ICCPS), p. 98–106, 2018. [19] A. J. Taylor and A. D. Ames, “Adaptive Safety with Con- trol Barrier Functions,” in 2020 American Control Conference (ACC), p. 1399–1405, 2020. [20] I. Tezuka and H. Nakamura, “Strict Zeroing Control Barrier Function for Continuous Safety Assist Control,” IEEE Control Systems Letters, vol. 6, p. 2108–2113, 2022. [21] B. T. Lopez, J.-J. E. Slotine, and J. P. How, “Robust Adaptive Control Barrier Functions: An Adaptive and Data-Driven Ap- proach to Safety,” IEEE Control Systems Letters, vol. 5, no. 3, p. 1031–1036, 2021. [22] A. Isaly, M. Ghanbarpour, R. G. Sanfelice, and W. E. Dixon, “On the Feasibility and Continuity of Feedback Controllers Defined by Multiple Control Barrier Functions,” IEEE Trans- actions on Automatic Control, p. 1–15, 2024. [23] R. Cheng, G. Orosz, R. M. Murray, and J. W. Burdick, “End-to- End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, p. 3387– 3395, Jul. 2019. [24] M. Igarashi, I. Tezuka, and H. Nakamura, “Time-varying Con- trol Barrier Function and Its Application to Environment- Adaptive Human Assist Control,” IFAC-PapersOnLine, vol. 52, no. 16, p. 735–740, 2019. 11th IFAC Symposium on Nonlinear Control Systems NOLCOS 2019. [25] Y. Nishimura and K. Hoshino, “Control Barrier Functions for Stochastic Systems and Safety-Critical Control Designs,” IEEE Transactions on Automatic Control, p. 1–8, 2024. [26] A. A. D. Nascimento, A. Papachristodoulou, and K. Margel- los, “Probabilistically Safe Controllers based on Control Bar- rier Functions and Scenario Model Predictive Control,” in 2024 IEEE 63rd Conference on Decision and Control (CDC), p. 1814–1819, 2024. [27] S. Farzan, V. Azimi, A.-P. Hu, and J. Rogers, “Adaptive Control of Wire-Borne Underactuated Brachiating Robots Using Con- trol Lyapunov and Barrier Functions,” IEEE Transactions on Control Systems Technology, vol. 30, no. 6, p. 2598–2614, 2022. [28] W. Shaw Cortez, D. Oetomo, C. Manzie, and P. Choong, “Control Barrier Functions for Mechanical Systems: Theory and Application to Robotic Grasping,” IEEE Transactions on Control Systems Technology, vol. 29, no. 2, p. 530–545, 2021. [29] R. Funada, M. Santos, R. Maniwa, J. Yamauchi, M. Fujita, M. Sampei, and M. Egerstedt, “Distributed coverage hole pre- vention for visual environmental monitoring with quadcopters via nonsmooth control barrier functions,” IEEE Transactions on Robotics, vol. 40, p. 1546–1565, 2024. [30] M. H. Cohen, T. G. Molnar, and A. D. Ames, “Safety-critical control for autonomous systems: Control barrier functions via reduced-order models,” Annual Reviews in Control, vol. 57, p. 100947, 2024. [31] F. Ferraguti, C. T. Landi, A. Singletary, H.-C. Lin, A. Ames, C. Secchi, and M. Bonf`e, “Safety and Efficiency in Robotics: The Control Barrier Functions Approach,” IEEE Robotics & Automation Magazine, vol. 29, no. 3, p. 139–151, 2022. [32] A. Alan, A. J. Taylor, C. R. He, A. D. Ames, and G. Orosz, “Control Barrier Functions and Input-to-State Safety With Application to Automated Vehicles,” IEEE Transactions on Control Systems Technology, vol. 31, no. 6, p. 2744–2759, 2023. [33] S. Feng, R. de Castro, and I. Ebrahimi, “Safe Battery Control Using Cascade-Control-Barrier Functions,” IEEE Transactions on Control Systems Technology, vol. 32, no. 6, p. 2344–2358, 2024. [34] A. Bhargava, C. Witkowski, S.-Z. Looi, and M. Thomson, “What’s the Magic Word? A Control Theory of LLM Prompt- ing,” arXiv preprint arXiv:2310.04444, 2024. [35] S. Soatto, P. Tabuada, P. Chaudhari, and T. Y. Liu, “Taming AI Bots: Controllability of Neural States in Large Language Models,” arXiv preprint arXiv:2305.18449, 2023. [36] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Fer- rer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia- Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzm ́an, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Ko- revaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. John- stun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kam- badur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. C ̧ elebi, P. Alrassy, P. Zhang, P. Li, P. Va- sic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Gana- pathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Poli- doro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speck- bacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchan- dani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, 10 B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C.-H. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Mont- gomery, E. Presani, E. Hahn, E. Wood, E.-T. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I.-E. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J.-B. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizen- stein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Ja- gadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchan- dani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ra- maswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agar- wal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma, “The Llama 3 Herd of Models,” arXiv preprint arXiv:2407.21783, 2024. [37] J. Zeng, B. Zhang, and K. Sreenath, “Safety-Critical Model Pre- dictive Control with Discrete-Time Control Barrier Function,” in 2021 American Control Conference (ACC), p. 3882–3889, 2021. [38] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learners,” arXiv preprint arXiv:2020.14165, 2020. [39] D. Loureiro, F. Barbieri, L. Neves, L. Espinosa Anke, and J. Camacho-collados, “TimeLMs: Diachronic language models from Twitter,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstra- tions, p. 251–260, Association for Computational Linguistics, May 2022. [40] N. D. Cao, G. Izacard, S. Riedel, and F. Petroni, “Autoregressive Entity Retrieval,” arXiv preprint arXiv:2010.00904, 2021. [41] N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, and R. Socher, “CTRL: A Conditional Transformer Language Model for Controllable Generation,” arXiv preprint arXiv:1909.05858, 2019. [42] J. Zingale and J. Kalita, “Language Model Sentence Completion with a Parser-Driven Rhetorical Control Method,” in Proceed- ings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers) (Y. Graham and M. Purver, eds.), (St. Julian’s, Malta), p. 193–203, Association for Computational Linguistics, Mar. 2024. [43] S. Mudgal, J. Lee, H. Ganapathy, Y. Li, T. Wang, Y. Huang, Z. Chen, H.-T. Cheng, M. Collins, T. Strohman, J. Chen, A. Beutel, and A. Beirami, “Controlled Decoding from Lan- guage Models,” arXiv preprint arXiv:2310.17022, 2024. [44] Z. Xu, F. Jiang, L. Niu, J. Jia, B. Y. Lin, and R. Poovendran, “SafeDecoding: Defending against Jailbreak Attacks via Safety- Aware Decoding,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (L.-W. Ku, A. Martins, and V. Srikumar, eds.), (Bangkok, Thailand), p. 5587–5605, Association for Computa- tional Linguistics, Aug. 2024. [45] K. Yang and D. Klein, “FUDGE: Controlled Text Genera- tion With Future Discriminators,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Association for Computational Linguistics, 2021. [46] J. Y. Huang, S. Sengupta, D. Bonadiman, Y. an Lai, A. Gupta, N. Pappas, S. Mansour, K. Kirchhoff, and D. Roth, “DeAL: Decoding-time Alignment for Large Language Models,” arXiv preprint arXiv:2402.06147, 2024. [47] K. Li, O. Patel, F. Vi ́egas, H. Pfister, and M. Wattenberg, “Inference-time intervention: Eliciting truthful answers from a language model,” arXiv preprint arXiv:2306.03341, 2024. [48] ConvoKit, “Reddit Corpus (small),” 2018. https://convokit. cornell.edu/documentation/reddit-small.html. [49] Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, “G-Eval: NLG Evaluation using Gpt-4 with Better Human Alignment,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (H. Bouamor, J. Pino, and K. Bali, eds.), (Singapore), p. 2511–2522, Association for Computa- tional Linguistics, Dec. 2023. [50] M. Isonuma and I. Titov, “Unlearning Traces the Influ- ential Training Data of Language Models,” arXiv preprint arXiv:2401.15241, 2024. Appendix A. Proof of Theorem 2 This subsection proves the property of the CBF filter (9). To this end, we define allowed and disallowed sets A, D. The allowed set A is a set of tokens that hold the CBF inequality (10), while the disallow set D is a set of tokens that do not. Given P , we consider the optimization problem (11). To satisfy the constraint (11d), it is clear that the probability of every disallowed token t,D must be 0. Therefore, the constraint (11d) is written as Q[t] = 0,∀t ∈ D. Now, the optimization problem (11) is rewritten as follows: min Q D KL [Q||P ], s.t. Q[t]≥ 0, ∀t∈A, Q[t] = 0, ∀t∈D, X t∈A Q[t] = 1. (14a) (14b) (14c) (14d) The KL divergence (14a) is rewritten as D KL [Q||P ] = P t∈T Q[t] ln Q[t] P[t] . Recall that Q[t] = 0,∀t ∈ D, the KL divergence (14a) is reduced to X t∈A Q[t] ln Q[t] P [t] .(15) 11 From here, we are focusing only on allowed tokens t∈A on the KL divergence. We next show that Q[t] = 0 for some t ∈ A cannot be the minimizer. The gradient of KL divergence (15) is given by ∂ ∂Q[t] " X t∈A Q[t] ln Q[t] P [t] # = ln Q[t] P [t] + 1 ,t∈A.(16) Note that the gradient goes −∞ as Q[t] → +0. This implies that Q[t] = 0 for some t ∈ A is not the minimizer. Therefore, the constraint (14b) is reduced to Q[t] > 0,t ∈ A. Then, the optimization problem (14) is further simplified as follows: min Q X t∈A Q[t] ln Q[t] P [t] , s.t. Q[t] > 0, ∀t∈A, X t∈A Q[t] = 1. (17a) (17b) (17c) The objective function (17a) is convex to Q[t],t ∈ A and has the minimizer, since the Hessian matrix is positive definite, i.e., H(Q[t],t∈A) = diag 1 Q A [1] ,..., 1 Q A [|A|] ≻ 0, (18) where Q A [i] is the probability of i-th allowed token. Now, we formulate the Lagrange function L as follows: L := X t∈A Q[t] ln Q[t] P [t] + λ X t∈A Q[t]− 1 ! ,(19) where λ is the Lagrange multiplier. The minimizer of the optimization problem (17) should satisfy the following equation: ∂L ∂Q[t] = 0, ∀t∈A.(20) For each allowed token t∈A, it follows that: ∂L ∂Q[t] = ln Q[t] P [t] + 1 + λ = 0.(21) This implies that, Q[t],t∈A is Q[t] = e −(1+λ) P [t], ∀t∈A.(22) Since the coefficient e −(1+λ) is same for all allowed tokens t ∈ A, we see that (17c) holds for λ = ln P t∈A P [t]− 1 and this reduces the coefficient e −(1+λ) to Q[t] = P [t] P t∈A P [t] , ∀t∈A.(23) This concludes that the proposition holds, meaning that the CBF filter, shown in (9), has the KL minimality under the CBF inquality (10). B. Evaluation of Naturalness and Positiveness In the naturalness and positiveness evaluation by G- Eval framework [49], we used GPT-5-mini and the follow- ing prompt. These scores are normalized by dividing the response values by 10. The prompt for naturalness evaluation is as follows: Given the evaluation steps, return a JSON with two keys: 1) a ‘score‘ key ranging from 0 - 10, with 10 being that it follows the criteria outlined in the steps and 0 being that it does not, and 2) a ‘reason‘ key, a reason for the given score, but DO NOT QUOTE THE SCORE in your reason. Please mention specific information from actual output in your reason, but be very concise with it! Evaluation Steps: 1. Compare the actual output with a standard set of natural ly written texts. 2. Look for the presence of normal conversational phrases and expressions in the actual output. 3. Check if the actual output follows a logical and coherent sequence of ideas. 4. Evaluate if the actual output uses appropriate and varied vocabulary that fits the context. actual output : Output text ** IMPORTANT: Please make sure to only return in JSON format, with the ”score” and ”reason” key. No words or explanation is needed. Example JSON: ”score”: 0, ”reason”: ”The text does not fol low the evaluation steps provided.” ** JSON: ” ” The prompt for positiveness evaluation is as follows: Given the evaluation steps, return a JSON with two keys: 1) a ‘score‘ key ranging from 0 - 10, with 10 being that it fol lows the criteria outlined in the steps and 0 being that it does not, and 2) a ‘reason‘ key, a reason for the given score, but DO NOT QUOTE THE SCORE in your reason. Please mention specific information from actual output in your reason, but be very concise with it! Evaluation Steps: 1. Identify and note down all the positive words and phrases used in the given text. 2. Evaluate the frequency and distribution of these positive words/phrases throughout the text. 3. Assess the context in which these positive words/phrases are used, to ensure they are indeed contributing to a positive sentiment. 4. Compare the frequency, distribution, and context of positive words/phrases in the given text with those in other texts to 12 determine its positivity level. actual output : Output text ** IMPORTANT: Please make sure to only return in JSON format, with the ”score” and ”reason” key. No words or explanation is needed. Example JSON: ”score”: 0, ”reason”: ”The text does not fol low the evaluation steps provided.” ** JSON: ” ”