Paper deep dive
Parameter Exploration for RLVR via Variational Learning
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.
Tags
Links
- Source: https://arxiv.org/abs/2608.09805v1
- Canonical: https://arxiv.org/abs/2608.09805v1
Trouble viewing inline? Open PDF directly →
Full Text
82,654 characters extracted from source content.
Expand or collapse full text
alg[Algorithm][List of Algorithms] Parameter Exploration for RLVR via Variational Learning Vatsal Venkatkrishna , Nico Daheim , Iryna Gurevych INSAIT, Sofia University “St. Kliment Ohridski”, Bulgaria Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany insait-institute/C3PO huggingface.co/BayesRL Abstract Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs. †footnotetext: Corresponding author: vatsal.venkatkrishna@insait.ai 1 Introduction Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful post-training paradigm that can be used to improve the reasoning capabilities of LLMs without relying on human demonstrations [31, 47]. Broadly, reinforcement learning could discover novel solutions by learning from autonomous experience [60, 1], but current RLVR algorithms are frequently limited by a model’s capabilities prior to RL [9, 74, 69]. Since the model learns exclusively from self-generated rollouts, high-reward trajectories may simply have too low probability and are thus often not discovered [68]. This is particularly problematic in algorithms with grouped advantage calculations like Group-Relative Policy Optimization (GRPO) [53], where there is no learning signal whenever all rollouts for a prompt receive the same reward. Such zero-advantage groups stall learning and therefore increase the already high compute costs associated with RLVR [45, 64]. Figure 1: Temperature versus Weight noise. Temperature scaling controls the entropy of the token distribution but preserves relative token ordering. Adding noise to the model’s weights can completely alter the token distribution, helping the model explore new regions of the solution space. A common solution is to increase the softmax temperature during rollout generation to promote exploration [2, 77]. However, as illustrated in Figure˜1, temperature can only control the variance of the token distribution and not the relative token order, which limits the trajectories that can be explored. Moreover, it uniformly alters the probability of all tokens in the vocabulary, which can lead to degeneration at high temperatures when erroneous tokens become likely [21]. Other complementary strategies like advantage shaping [65, 76, 73] and entropy regularization [38, 25, 54] have been proposed, but rather aim to stabilize training than promote exploration. Parameter-space exploration has a long track record in more classical RL [49, 51, 1]. One goal of parameter-space exploration has been to find high-quality trajectories from policies that are in the neighborhood of the current policy, thereby mitigating the issue that action-space noise might sample unlikely or low-quality outputs [14]. Additionally, weight perturbations can directly alter logit values and reorder token distributions as illustrated in Figure˜1. Intuitively, if we could learn the amount of noise to add to each parameter, we could explore promising alterations to token distributions and rollouts. However, to the best of our knowledge, such strategies have been unexplored for LLMs, possibly because noisy optimization itself is seldom used. In this work, we present Perturbed Parameter Policy Optimization (3PO) which uses parameter-space exploration for RLVR by sampling models from an approximate posterior during training that can be updated during the reinforcement learning stage to adaptively learn noise injection. At each training step, one or more samples are obtained by adding learned noise to the learned parameters, which can be used to control the amount of exploration. We explore three strategies for noise injection. B3PO samples a single weight perturbation per gradient step, which makes it dependent on a single weight draw and can limit group diversity. M3PO calculates advantages separately for M model perturbations before a single weight update. C3PO divides each GRPO group across N independently sampled weights and calculates advantages on the full group to maximize diversity. We validate 3PO on OLMo-3-1025-7B [46] and Qwen2.5-Math-7B [70] across mathematical reasoning and code generation tasks, which are widely used as RLVR testbeds [9, 11, 39, 65, 76]. Overall, we find that parameter exploration can help downstream performance and speed up convergence without using more rollouts than AdamW-based training. Our chunked noising approach, C3PO, achieves the best average performance across both model families. Broadly, using multiple model samples per batch performs better than using just a single sample, and rescues more zero-advantage groups than other action-space exploration baselines. We find consistent improvements in model performance that are concentrated on harder benchmarks like AIME [42] and LiveCodeBench [23]. Altogether, our work shows that parameter exploration via learning approximate posteriors can be helpful for RLVR by providing an additional control lever for exploration. 2 Background on RLVR and parameter exploration 2.1 RLVR for LLMs Reinforcement Learning with Verifiable Rewards has recently become popular for aligning LLMs in problem settings where outputs of the model for a given problem can easily be verified, for example, in solving mathematical word problems, where a ground-truth solution exists. Overall, the idea is to assign a reward rir_i, for example, a binary ri∈0,1r_i∈\0,1\ to each sequence (i)∈∗ $ $y$$^(i) that an LLM generates based on a problem x, which is then subsequently used to train the model. Crucially, the outputs (i) $ $y$$^(i) are self-generated by the model being trained and not obtained via human demonstration. However, some human-provided ground-truth solution is still usually used to calculate rir_i on a final answer at the end of the LLM’s output (i) $ $y$$^(i). Group Relative Policy Optimization (GRPO) [53] is a widely-used technique for RLVR training, where the learning signal is calculated based on advantages Ai,t=ri−mean()std()A_i,t= r_i-mean( $ $ r$$)std( $ $ r$$) using the shorthand =(r1,…,rG) $ $r$$=(r_1,…,r_G) for a group-size G. That is, a total of G rollouts are sampled for each step that are subsequently used for advantage calculation. Altogether, the GRPO loss takes the following form, ℒGRPO=[∼D,(i)i=1G∼πold(⋅|)][1∑i|(i)|∑i=1G∑t=1|(i)|minRi,tAi,t,clip(Ri,t,1−ϵ,1+ϵ)Ai,t]. splitL^GRPO=E_ [ $ $ x$$ D,\ $ $ y$$^(i)\_i=1^G _old(·| $ $ x$$) ] [ 1 _i| $ $y$$^(i)| _i=1^G _t=1^| $ $ y$$^(i)| \R_i,tA_i,t,clip(R_i,t,1-ε,1+ε)A_i,t \ ]. split (1) The ratio Ri,t=π(yt(i)∣,<t(i))πold(yt(i)∣,<t(i))R_i,t\!=\! _$ θ$(y_t^(i) $ $ $ x$$$,$ $ $ y$$$^(i)_<t) _old(y_t^(i) $ $ $ x$$$,$ $ $ y$$$^(i)_<t) is an importance sampling factor to account for cases where the language model that generates the rollouts (πold _old) differs from the language model that is updated (π _$ θ$). Here, we omit the KL-divergence term following prior work [73]. Despite the success of GRPO, several issues remain. For example, groups that lead to zero-advantages stall learning. This is the case when all rollouts achieve the same reward and either all lead to a correct or incorrect solution. Intuitively, this is connected to the problem of trading off exploration and exploitation. A good RL algorithm needs to mix failed attempts and successful attempts to sustain a learning signal, though achieving a good trade-off is often hard in practice [60]. This trade-off is similarly hard to establish in RLVR for LLM training. The entropy of the policy often collapses prematurely, and the groups sampled at each step become increasingly homogeneous, leading to a slowing or even plateauing of learning because the advantages become less informative [73]. As a consequence, a variety of approaches have been proposed to target this problem. Clip-higher [73] loosens the upper PPO clip to preserve probability mass on exploratory tokens, but it has no effect on on-policy methods as the importance sampling ratio vanishes. Another solution is to increase the softmax temperature during rollout generation in order to flatten the token distribution [2, 77], but this has limited benefits as it is rank-preserving and requires careful tuning to not overshoot. Explicit entropy regularization adds an entropy bonus to the loss [38, 25], and advantage-shaping methods reweight the learning signal toward tokens or trajectories deemed more informative [65, 76]. Other strategies are using larger rollout budgets [22, 19], prolonged training [37, 22, 72], or ordering training data by difficulty to gradually expand the reachable solution set [59, 52, 41]. Still, all of these methods operate in the action space. They reweigh or regularize token probabilities and do not alter the generators and policy’s parameters. In this work, we explore an orthogonal axis: perturbing the policy’s parameters directly as an additional axis of exploration in RLVR. By operating in the parameter space, our method is naturally compatible with previous action-space approaches. 2.2 Parameter space exploration As previously established, the exploration-exploitation tradeoff is a well-studied problem in classical RL. While methods like ϵε-greedy [61, 66] and UCB [57, 24] work well for smaller problems, they tend to not scale well to problems with large action spaces, such as language modeling. Thus, there has been sustained interest in developing efficient algorithms for exploration [1]. One such method is parameter exploration, where noise is added to the policy’s parameters rather than its actions [49, 51]. Early work used finite-difference gradients with fixed parameter perturbations [29]; subsequent extensions developed state-dependent and per-basis variants [62]. Salimans et al. [50] scaled the approach to deep policies via evolution strategies, and Plappert et al. [48] and Fortunato et al. [16] integrated parameter noise with modern deep RL algorithms, with the latter learning the noise magnitude jointly with the policy. Across continuous control benchmarks, episode-coherent parameter perturbation has consistently matched or outperformed action-space noise [58, 63]. Despite these successes, parameter-space exploration methods have largely been absent from RLVR for LLMs. Concurrent work [3] extends the method of Plappert et al. [48] to RLVR by adding gaussian noise to the rollout generating policy and all rollouts are generated by the same model. Their approach can be seen as a variant of the B3PO we present, where the noise added is controlled by an adaptive scheduler rather than a learned Hessian (cf. Equation˜4). Additionally, we study the effects of using multiple Monte-Carlo samples of the noise (M3PO) and a chunked noising approach (C3PO) to directly control group diversity in the GRPO objective, which we ultimately find are more effective methods. 2.3 Variational learning and IVON In this work, we use variational learning for parameter exploration, because it naturally learns a distribution over neural network parameters. Given a family of distributions, usually Gaussians, Q, generalized variational learning aims to find the following approximate posterior over parameters θ, q∗()=argminq∈∑i∈∼q[ℓi()/||]+λ−1KL(q∥p0).q ( θ)= _q∈ $ Q$ _i E_ θ q[ _i( θ)/|D|]+λ^-1D_KL(q\,\|\,p_0). (2) The objective is similar to standard training, as it also uses general losses [28, 27] but uses a smoothed loss via an expectation over q, as well as a KL regularizer towards a prior p0p_0. The prior can be chosen as p0∝exp−ℓ0p_0 \- _0\, where ℓ0 _0 is a standard regularizer like weight decay to mirror conventional deep learning. The scaling factor λ>0λ>0 is used to weigh data fit and regularization towards the prior. Many methods have been proposed to optimize Equation˜2. For example, Graves [18] and Bayes-by-Backprop [5] try to use conventional gradient descent methods to learn the mean and variance of a Gaussian Q. More recently, natural-gradient-based methods have shown more success [26, 34]. In particular, the IVON optimizer [55] has been shown to provide similar performance as methods like AdamW [40] at little overhead in terms of runtime that is mainly due to the sampling implementation. In IVON, at each step the loss is calculated at a perturbed point θ, and not at the current mean m, which corresponds to the point estimate learned by e.g. AdamW. Writing it out as the reparameterization, ^=+σ⊙;∼(0,I), θ=m+σ ; (0,I), (3) shows the similarity to weight-perturbation-based optimization strategies, which have been found to, among others, lead to improved generalization as they avoid sharp minima [67, 15, 44]. The diagonal variance in IVON is defined inversely proportional to the Hessian which is weighted by the effective sample size (ess) λ>0λ>0 from Equation˜2. The ess can be tuned as an inverse temperature; for example, λ=109λ=10^9 is common. Concretely, the variance is defined as follows: σ=1/λ(+δ),σ=1/ λ( $ $h$$+δ), (4) where h is the diagonal Hessian initialized to a constant h0h_0, and δ≥0δ≥ 0 is a weight-decay scaling. Choosing λ large thus reduces noise, while choosing λ small increases the amount of noise added to the parameters, which in turn increases diversity in predictions. We use this strategy to control exploration in reinforcement learning with verifiable rewards next. Figure 2: Overview of noising strategies. Squircles denote model weights (base θ or perturbed θ), while squares denote rollouts; each column corresponds to a rollout group for a different prompt. Color intensity denotes the amount of noise added. GRPO generates all G rollouts using a single shared model θ, with diversity arising only from sampling. B3PO samples a single perturbation θ per batch and reuses it across all prompts and groups. M3PO accumulates the loss over M perturbations to stabilize training. C3PO partitions rollouts across N perturbations and applies GRPO to the aggregated set, which now contains rollouts from multiple model samples. 3 Perturbed Parameter Policy Optimization We now describe Perturbed Parameter Policy Optimization, which introduces exploration in the parameter space as an additional axis of exploration in RLVR by using noisy policies. Importantly, our method is orthogonal to and compatible with existing approaches like temperature control [2, 77], clip-higher [73], and entropy regularization [38, 25]. We study three different strategies that are illustrated in Figure˜2. 3.1 Batched noising (B3PO and M3PO) Modern RL codebases such as verl [56] and SkyRL [7] decouple generation and gradient updates into separate inference and training engines, periodically synchronizing their weights at the start of every gradient step. Batched noising works well with such implementations, because a single weight perturbation θ is sampled from the IVON posterior (Equation˜3) and synced with the rollout engine. B3PO is analogous to step-level temperature control [77], but acting in the weight space instead, and can be interpreted as a one Monte-Carlo sample approximation of Equation˜2 using Equation˜1 as the loss. B3PO could have similar issues as standard GRPO if the single weight perturbation per step does not perform well (or produces homogeneous rollouts). One way to address this issue is to use multiple MC samples for the expectation in Equation˜2, i.e., sampling M perturbations from the IVON posterior q. For each, we generate rollouts over the entire prompt batch and compute the GRPO loss in Equation˜1. The posterior is updated once at the end, with gradients averaged across all M perturbations. We refer to this strategy as M3PO. The pseudo-code for both methods is shown in Appendix˜K. 3.2 Chunked noising (C3PO) GRPO’s learning signal stems from within-group reward differences, but M3PO calculates advantages separately for each group instead of using one large group across samples that, intuitively, should be more diverse as it uses rollouts from different models to calculate advantages. We address this issue in C3PO, where we sample N weight samples using IVON, say with groups of size G/NG/N, and then calculate advantages over the full group of G responses, whose diversity now reflects N distinct weight-space samples. The full procedure is given in Appendix˜K. Empirically, we found that Seq-MIS [35], which uses sequence-level importance sampling and masking, was important to stabilize training and ensure convergence (see Appendix˜I for an ablation). A possible reason for this observation is that the policies that are used for gradient estimation differ from the N rollout generators, akin to training – inference mismatches in RLVR literature [71]. Caching and replaying the noise could be an alternative but comes at increased memory cost. 4 Experiments 4.1 Experimental setup We use two base models, namely, Olmo-3-1025-7B [46] and Qwen2.5-Math-7B [70] for our experiments. We follow a two-stage post-training pipeline, beginning with a warm-start SFT followed by RLVR [31, 13]. For the SFT phase, we use a ≈2≈\!2M-example subset of the Llama-Nemotron Post-Training Dataset [4] containing DeepSeek-R1 [13] responses covering math, code, general reasoning, and instruction following. We train for two epochs using IVON [55] for both Olmo3 and Qwen2.5-Math. This phase equips the model with initial reasoning capabilities. While it also produces a learned noise distribution that could be reused as an initialization for RLVR, throughout the main results, we initialize the IVON Hessian h from scratch with a constant h0h_0 rather than loading the SFT optimizer state to keep our results applicable to any off-the-shelf model and allow us to isolate the benefits of a learned initialization in Section˜5.4. Unless specified otherwise, we train our models using GRPO [53] on DAPO-MATH-17k [73], which contains approximately 17,00017,000 math problems with ground-truth answers. We evaluate on six widely used mathematical reasoning benchmarks: AIME 2024–26, AMC 2023, MATH-500 [20], and Minerva [32]. Across all experiments, we report the mean and standard error for Pass@1, calculated across K=8K\!=\!8 rollouts using the unbiased estimator given by Chen et al. [8]. We compare 3PO to three action-space exploration baselines: adding an entropy regularization term (EntReg), raising sampling temperature when entropy collapses beyond a threshold (Polaris [2]) and applying a KL penalty to tokens with a high covariance between probability and advantage (KLCov [11]). For all methods, the effective group size is fixed to G=16G\!=\!16 rollouts per batch. For M3PO, we report an equal-compute comparison with AdamW, using M=4M\!=\!4 MC samples per batch to generate rollouts and compute advantages with G=4G\!=\!4 at a time. The resulting policy update is thus calculated on a total of G=16G\!=\!16 rollouts. For C3PO, we sample N=4N\!=\!4 distinct policies per batch and accumulate 44 rollouts for each in a rollout buffer, computing advantages on 1616 rollouts in total. We use λ=109λ=10^9 for Olmo3 and λ=1010λ=10^10 for Qwen2.5-Math. We sweep these hyperparameters in Sections˜5, C and D, and describe our training configuration in detail in Appendix˜A. 4.2 Main results Method AIME ’24 AIME ’25 AIME ’26 MATH-500 AMC Minerva Average SFT 16.68 2.12 21.05 2.26 12.70 2.13 78.07 0.68 51.12 3.53 30.51 0.11 35.02 0.51 GRPO 25.41 2.92 22.08 2.48 19.16 3.45 84.92 0.47 69.37 3.50 36.99 0.94 42.99 0.51 GRPO (EntReg) 23.75 2.35 23.33 2.45 17.91 2.68 84.60 0.58 68.43 2.89 36.58 0.95 42.43 0.53 GRPO (Polaris) 26.25 2.65 20.83 2.63 15.83 3.47 84.87 0.63 64.37 2.94 36.63 0.99 41.46 0.54 GRPO (KLCov) 24.16 2.23 24.58 2.78 16.25 2.93 83.87 0.68 66.56 3.45 38.28 0.94 42.28 0.49 B3PO 25.41 2.76 25.00 2.34 15.00 2.34 83.30 0.58 70.31 3.01 37.36 0.92 42.73 0.48 M3PO 25.83 3.48 24.58 2.53 17.91 2.64 85.05 0.51 69.37 3.27 36.58 0.83 43.22 0.49 C3PO 27.50 2.91 26.25 2.58 19.16 2.94 84.52 0.48 70.00 3.43 36.81 0.93 44.04 0.46 SFT 14.58 2.50 20.00 3.12 12.50 2.39 65.25 1.93 59.38 2.98 24.59 1.15 32.72 0.36 GRPO 24.17 3.25 25.42 3.51 20.42 3.05 88.20 0.96 70.62 3.01 43.38 1.20 45.36 0.35 GRPO (EntReg) 23.33 3.34 24.17 3.67 15.83 3.19 86.28 0.99 65.62 3.06 42.65 1.18 42.98 0.34 GRPO (Polaris) 22.92 3.46 26.67 3.35 21.25 3.14 87.75 1.02 72.81 2.97 43.29 1.26 45.78 0.32 GRPO (KLCov) 22.92 3.13 24.58 3.58 20.42 3.20 87.75 0.94 71.88 2.63 43.20 1.24 45.12 0.37 B3PO 23.75 3.14 25.42 3.65 24.17 3.19 87.68 0.96 72.81 2.62 43.34 1.21 46.20 0.37 M3PO 23.75 3.11 25.42 3.44 24.17 3.29 88.58 0.95 73.44 3.00 43.01 1.20 46.39 0.36 C3PO 25.00 3.21 27.08 3.12 26.25 2.93 88.10 0.93 71.25 2.67 43.61 1.22 46.88 0.38 Table 1: A comparison of all methods across common mathematical reasoning benchmarks. Methods with multiple noise samples outperform GRPO on average. Both models benefit most from C3PO’s increased group diversity. The largest gains are on the harder AIME problems. We report the unbiased Pass@1 across 8 samples, and its standard error. Parameter-space exploration outperforms action-space baselines. We report benchmark scores across all methods and models in Table˜1. All 3PO variants outperform action-space exploration baselines on average. We hypothesize that these gains are driven by parameter-space exploration discovering fewer invalid rollouts than their action-space counterparts, since trajectories come from different plausible LLM policies. We probe this in Section˜4.3. Notably, performance gains are concentrated on the harder AIME benchmarks, outperforming GRPO by up to 4.17%4.17\% on Olmo3 and 5.83%5.83\% on Qwen2.5-Math. We see similar trends on the relatively harder code generation task in Section˜4.4. We also find that C3PO’s gain over GRPO is robust to run-to-run variance (Appendix˜G). On average, using multiple model samples outperforms single-model methods. Although B3PO outperforms GRPO on individual benchmarks such as AIME-2025, AMC, and Minerva, indicating that simply switching optimizers from AdamW to IVON during RLVR could buy modest improvements in some cases. However, its average performance does not improve meaningfully beyond that of the standard GRPO, possibly due to high noise sample variance and limited group diversity. While M3PO uses M MC samples to reduce noise sample variance and achieves moderate performance gains over B3PO for both model families, its effectiveness is limited due to a proportional decrease in its group size G. We show that lifting the equal-compute constraint further improves the performance of M3PO in Section˜5.2. C3PO has the highest average performance in both model families, likely due to the increase in group diversity due to using multiple model samples. We show results with Olmo3 in the next sections and report results using Qwen2.5-Math in Appendices˜B, C and D. 4.3 Does 3PO really explore more? Figure 3: (Left) Cumulative zero-advantage groups rescued over training. 3PO rescues more groups than GRPO and action-space baselines, with multiple perturbations yielding consistent improvements. (Middle) Paired degeneracy rates vs GRPO. Polaris causes high degeneration due to uniform token reordering at high temperature. (Right) Paired incorrect rollout rates vs GRPO. While EntReg escapes degeneracy in late training, it buys diversity mostly through incorrect rollouts. Early and late refer to the early (first 33%) and late training steps (last 33%), respectively. $ $$ $footnotetext: GRPO’s group is subsampled to 4 rollouts for its pairing with M3PO. In this section, we evaluate whether 3PO can increase exploration during RLVR. However, token-level entropy alone is not sufficient to measure this, since declining entropy could indicate two opposite behaviors. On the one hand, entropy can decrease due to a focus on high-probability – high-advantage tokens [11] which could be seen as a form of exploitation. On the other hand, low entropy can equally signal a policy learning incorrect trajectories [9]. We observe this phenomenon in Figure˜7 when noise is scaled too far and where the steep decline in C3PO’s entropy during training (Figure˜6) is largely uninformative. Measures such as pass@K may also be misleading, as they measure the ability of a model to explore post-RLVR at inference, which is often traded for a higher pass@1 during via the entropy - reward relationship mentioned above in a form of exploitation.111Still, sampling from the learned posterior could yield inference-time improvements [10] To measure exploration during RLVR, we pair each method against GRPO and study how often each method “rescues” zero-advantage groups. A prompt is “rescued” if a method produces at least one correct rollout where GRPO produced none, and “lost” in the opposite case. A successful training-time exploration method should rescue more dead groups than it loses, thus sustaining a learning signal. We also measure the number of degenerate and incorrect rollouts produced by each method during training as compared to GRPO. For each prompt, a method “wins” over GRPO if it produces fewer malformed/incorrect rollouts during training. A degenerate (or malformed) rollout is one that does not produce an extractable answer at all, often repeating tokens until the generation limit, and do not buy any learning signal. 3PO rescues more zero-advantage groups than baselines. Figure˜3 (Left) shows that 3PO tends to rescue more groups than it loses, while action-space methods fail to do so. Intuitively, 3PO’s trajectories come from LLM policies that are likely under the posterior and which are less likely to generate invalid trajectories than the undirected exploration of action-space methods. Using IVON further improves the quality of these trajectories by learning the distribution over policies jointly during training (Equation˜4). Polaris loses significantly more groups than it rescues and its gap to GRPO widens during training. Although B3PO rescues dead groups in early training, its advantage diminishes in the later training steps. Conversely, M3PO and C3PO continue to rescue zero-advantage groups throughout training, highlighting the benefits of using multiple parameter perturbations. Undirected action-space methods mostly yield malformed or incorrect rollouts. Action-space baselines lose to GRPO on both malformed and incorrect rollout generation rates. For instance, Polaris increases sampling temperature when entropy collapses beyond a threshold. Since temperature only flattens the token distribution (see Figure˜1), it consistently yields rollouts with no extractable answer (Figure˜3 (Middle)). EntReg exhibits a different failure mode. Although its rollouts stay well-formed, even beating 3PO in late-training, its higher entropy is spent on mostly incorrect rollouts (Figure˜3 (Right)). KL-Cov exhibits a mix of both failure modes, Using IVON further improves the quality of these trajectories by learning the distribution over policies jointly during training (Equation˜4). B3PO also tends to produce degenerate rollouts later in training which might explain its diminishing returns in rescuing zero-advantage groups. M3PO and C3PO are the only methods to consistently produce fewer malformed and incorrect rollouts than GRPO which suggests that parameter-space exploration does produce more useful rollouts for RLVR. Figure 4: Reward and accuracy curves for code generation. C3PO outperforms all other methods on reward and accuracy. GRPO retains higher entropy longest but converts it into neither reward nor accuracy gains. Method ESS (λ) Avg. Math Pass@1 B3PO 10810^8 41.90 ± 0.51 10910^9 42.73 ± 0.48 101010^10 42.19 ± 0.53 M3PO 10810^8 42.93 ± 0.51 10910^9 43.22 ± 0.49 101010^10 43.59 ± 0.51 C3PO 10810^8 0.00 ± 0.00 10910^9 44.04 ± 0.46 101010^10 41.42 ± 0.54 Table 2: Effects of scaling λ: 10910^9 is a good default value, while λ=108λ\!=\!10^8 generally hurts performance. M3PO generally benefits from a higher λ 4.4 Results for code generation In this section, we study the effect of parameter-space exploration beyond mathematical reasoning by testing our methods on code generation tasks. Specifically, given a problem statement, the LLM must produce a code that passes all hidden test cases. We train Olmo3 on the CodeR1-12K dataset [36], which contains verified coding problems from LeetCode and TACO [33]. We use the same hyperparameters as math detailed in Section˜4.1 and utilize SandboxFusion [6] for executing untrusted LLM code. We evaluate all models on LiveCodeBench-v6 [23], which contains problems released between January 2025 and April 2025. As seen in Figure˜4, there is a consistent gap in the reward and accuracy curves between our 3PO methods and GRPO. C3PO is the best-performing method in this regime, reaching an LCBv6 score of 15.1715.17 and highlighting the effectiveness of increased group diversity through our proposed chunked noising approach. M3PO narrowly outperforms B3PO and GRPO, but is probably limited by its lower group diversity. Similar to math, 3PO is more sample efficient and reaches GRPO’s final reward within the first 50% of the training steps. Notably, 3PO methods reach a noticably higher reward than GRPO on code generation as compared to mathematical reasoning. We hypothesize that this is because code generation is a harder task for the pre-RL model, as seen by the SFT model’s performance on the LCBv6 benchmark (9.5%) compared to the math benchmarks (35.02%). Thus, increased exploration with 3PO could be more valuable here and aid in discovering high-reward trajectories. We observed a similar trend on the harder AIME benchmarks for math in Section˜4.2, which suggests that parameter-space exploration is particularly effective for harder tasks where the pre-RL model is less capable and more prone to generating malformed or incorrect rollouts. 5 Further analysis In this section, we now ablate several components of our proposed approach. Specifically, we study how the ESS (λ) should be scaled (Section˜5.1), and how many MC samples (Section˜5.2) and chunks (Section˜5.3) are required, and the impact of Hessian initialization (Section˜5.4). 5.1 Impact of scaling ESS (λ) The ESS (λ) is a crucial hyperparameter in our setup that controls the amount of noise added to the weights, and needs to be tuned to achieve good performance [12, 10]. As described in Section˜2.3, λ is often initialized empirically, and a lower value of ESS behaves similarly to increasing a temperature τ, but in the weight space. We ablate the values of λ∈108,109,1010λ∈\10^8,10^9,10^10\. We find decreasing λ too far to 10810^8 always hurts performance, and completely collapses training for C3PO. This is an expected result because C3PO samples multiple models in the group rollout, but the advantages are still calculated by a single model. Adding too much noise might stray the model too far. Conversely, increasing λ to 101010^10 has more nuanced results. For the batched noising methods B3PO and M3PO, it behaves similarly to λ=109λ=10^9 with a steeper entropy decline than λ=108λ=10^8. λ=1010λ=10^10 even slightly outperforms λ=109λ=10^9 on average for M3PO (Table˜2). However, λ=1010λ=10^10 plateaus at a lower ceiling than the 10910^9 C3PO run, as seen from its entropy and reward curves. A possible explanation is that M3PO has lower variance in the gradient update due to the averaging, and a higher λ reduces variance further by reducing the amount of noise added during optimization. On the other hand, C3PO benefits from increased noise due to greater group diversity. Through our experiments, we found that λ=109λ=10^9 is a good default, which can be raised to 101010^10 if training collapses due to high noise or reduced to 10810^8 if training stagnates close to vanilla GRPO. We present extended results for this ablation in Appendix˜C. It is also plausible that adaptive noise schedules could improve training more. We leave a detailed exploration of this for future work. Figure 5: (a) Equal-compute comparison: Benefits of increasing M are offset by reducing G. (b) Scaling M: The advantages of more MC samples are fully realized without an equal-compute constraint. (c) Scaling N: N>1N>1 improves performance, further scaling is incremental. (d) Learned noise prior: Initializing h from the learned SFT prior has negligible effects. 5.2 How does the number of MC samples M interplay with group size G? As described in Section˜3.1, M3PO accumulates the GRPO loss over M independently sampled weight settings before updating the model weights which can reduce the variance of the computed loss. This comes at a proportional compute cost increase because rollouts must be generated M times, once per noise sample (see Table˜3). We offset this by shrinking the per-perturbation group to G=16/MG\!=\!16/M, holding the total rollout count per step fixed. This equal-compute correction introduces an important tradeoff: as M grows, each group shrinks, reducing the diversity in the GRPO group and the chances of sampling a high-reward response per prompt [22]. Figure˜5(a) plots the result of sweeping M∈1,2,4,8M∈\1,2,4,8\ under this constraint. The effects of variance reduction from increasing M is roughly cancelled by the reduced group diversity from a smaller G, and vice versa. Removing the equal-compute correction in (Figure˜5(b)) steadily and substantially outperforms the equal-compute run, as well as B3PO with G=16G\!=\!16. This indicates that reduced gradient-update variance does benefit M3PO when compute is unconstrained, and might explain its advantage over B3PO in Table˜1. We provide further results in Appendix˜D. 5.3 What chunk size N is required? As described in Section˜3.2, C3PO diversifies each rollout group by cycling through N independently sampled weight perturbations, generating G/NG/N rollouts per draw. The resulting group of G responses is more diverse because it comes from multiple models. In this section, we ablate the chunk size N∈1,2,4,8,16N\!\,∈\!\,\1,2,4,8,16\ to analyze how many models are required to see the benefit of increased diversity. Note that N=1N\!=\!1 reduces to B3PO. Even modest group diversification yields a consistent benefit, as all runs with N≥2N≥ 2 outperform the single-perturbation baseline (Figure˜5(c)). Scaling beyond N=2N=2 improves performance in earlier training steps but smaller gains in final scores. We find that using N=2−4N=2\!\,-\,\!4 perturbations is a good strategy. We plot the reward and entropy curves in Appendix˜E. 5.4 Does learning the noise improve performance? The experiments so far initialize the IVON Hessian h to a constant h0h_0 at the start of RLVR, discarding the optimizer state obtained during the warm-start SFT phase. A natural hypothesis is that loading this learned Hessian, which encodes per-parameter variance information from the SFT phase, would yield a better posterior and improve downstream performance. We test this hypothesis by comparing two C3PO runs on Olmo3 that share the same SFT checkpoint. Figure˜5(d) shows that the two configurations track each other closely throughout training, suggesting that a learned distribution from SFT does not provide a clear advantage. This may be because the SFT phase is relatively short, and the loaded Hessian is nearly isotropic at the end. Concretely, <8%<\!8\% of the values lie outside the [0.9h0,1.1h0][0.9h_0,1.1h_0] range. Learning the distribution over a longer SFT phase or from pre-training itself could be more beneficial, but we need to leave this experiment for future work due to computational constraints. Nonetheless, our findings transfer to any base model without additional training. Practitioners can choose any off-the-shelf model checkpoint, initialize h to a constant h0h_0, and benefit from the increased exploration. We provide detailed results in Appendix˜F. 6 Conclusion We introduce Perturbed Parameter Policy Optimization (3PO), a family of parameter-space exploration strategies for RLVR. By sampling weights from a learned posterior at rollout time, 3PO provides an additional lever for exploration. Across our experiments, drawing multiple model samples per gradient step consistently outperforms action-space baselines. We show that our proposed methods rescue more zero-advantage groups and produce fewer malformed/incorrect rollouts than action-space baselines. While M3PO tends to reduce noise sample variance, it either requires more compute or suffers from reduced group diversity. Our chunked noising approach, C3PO, has the best average performance across both model families. Throughout our experiments, we find that it is useful to tune the amount of noise that is added to control the amount of exploration during RLVR. While a noise prior can be learned before an RL stage, starting with scaled isotropic Gaussian noise also provides improvements and makes our methods applicable as a drop-in replacement even for existing checkpoints that were not trained with variational learning. Overall, our work shows that parameter-space exploration can be used to improve reinforcement learning for large language models. 7 Limitations We expect the benefits of 3PO to grow with model scale but our work is limited by compute. Gan and Isola [17] show that large LLMs’ pretraining neighborhoods are densely populated with task-specific experts. Thus, larger models should tolerate a lower λ before training collapses, allowing for more exploration: an exciting direction for future work that we were unable to explore. In addition, M3PO and C3PO currently have ≈1.5×≈ 1.5× higher wall-clock times than GRPO (Appendix˜A), despite equivalent computational requirements in pseudocode. This stems from processing small prompt batches between vLLM model updates, which reduces parallelism benefits [30]. Overall, modern RL code stacks all operate using a single model, and it is often difficult to implement changes for variational learning. We also did not explore scaling the number of MC samples and rollouts further for test-time scaling due to limited resources. Finally, using more directed exploration through structured, non-diagonal posterior estimators [43] instead of IVON, or alternative sampling strategies to Monte-Carlo sampling are a promising direction for future work. Acknowledgements We gratefully acknowledge the support from (1) the Ministry of Education and Science of Bulgaria (support for INSAIT, part of the Bulgarian National Roadmap for Research Infrastructure), (2) the hessian.AI Service Center (funded by the Federal Ministry of Research, Technology and Space, BMFTR, grant no. 16IS22091), (3) the hessian.AI Innovation Lab (funded by the Hessian Ministry for Digital Strategy and Innovation, grant no. S-DIW04/0013/003), and (4) the German Federal Ministry of Research, Technology, and Space and the Hessian Ministry of Higher Education, Research, Science, and the Arts within their joint support of the National Research Center for Applied Cybersecurity ATHENE. References [1] S. Amin, M. Gomrokchi, H. Satija, H. van Hoof, and D. Precup (2021) A survey of exploration methods in reinforcement learning. CoRR abs/2109.00157. External Links: Link, 2109.00157 Cited by: §1, §1, §2.2. [2] C. An, Z. Xie, X. Li, L. Li, J. Zhang, S. Gong, M. Zhong, J. Xu, X. Qiu, M. Wang, and L. Kong (2025) POLARIS: a post-training recipe for scaling reinforcement learning on advanced reasoning models. External Links: Link Cited by: §1, §2.1, §3, §4.1. [3] B. Bai, X. Wang, P. Ye, and T. Chen (2026) Learning to explore with parameter-space noise: A deep dive into parameter-space noise for reinforcement learning with verifiable rewards. CoRR abs/2602.02555. External Links: Link, Document, 2602.02555 Cited by: §2.2. [4] A. Bercovich, I. Levy, I. Golan, M. Dabbah, R. El-Yaniv, O. Puny, I. Galil, Z. Moshe, T. Ronen, N. Nabwani, I. Shahaf, O. Tropp, E. Karpas, R. Zilberstein, J. Zeng, S. Singhal, A. Bukharin, Y. Zhang, T. Konuk, G. Shen, A. S. Mahabaleshwarkar, B. Kartal, Y. Suhara, O. Delalleau, Z. Chen, Z. Wang, D. Mosallanezhad, A. Renduchintala, H. Qian, D. Rekesh, F. Jia, S. Majumdar, V. Noroozi, W. U. Ahmad, S. Narenthiran, A. Ficek, M. Samadi, J. Huang, S. Jain, I. Gitman, I. Moshkov, W. Du, S. Toshniwal, G. Armstrong, B. Kisacanin, M. Novikov, D. Gitman, E. Bakhturina, J. P. Scowcroft, J. Kamalu, D. Su, K. Kong, M. Kliegl, R. Karimi, Y. Lin, S. Satheesh, J. Parmar, P. Gundecha, B. Norick, J. Jennings, S. Prabhumoye, S. N. Akter, M. Patwary, A. Khattar, D. Narayanan, R. Waleffe, J. Zhang, B. Su, G. Huang, T. Kong, P. Chadha, S. Jain, C. Harvey, E. Segal, J. Huang, S. Kashirsky, R. McQueen, I. Putterman, G. Lam, A. Venkatesan, S. Wu, V. Nguyen, M. Kilaru, A. Wang, A. Warno, A. Somasamudramath, S. Bhaskar, M. Dong, N. Assaf, S. Mor, O. U. Argov, S. Junkin, O. Romanenko, P. Larroy, M. Katariya, M. Rovinelli, V. Balas, N. Edelman, A. Bhiwandiwalla, M. Subramaniam, S. Ithape, K. Ramamoorthy, Y. Wu, S. V. Velury, O. Almog, J. Daw, D. Fridman, E. Galinkin, M. Evans, K. Luna, L. Derczynski, N. Pope, E. Long, S. Schneider, G. Siman, T. Grzegorzek, P. Ribalta, M. Katariya, J. Conway, T. Saar, A. Guan, K. Pawelec, S. Prayaga, O. Kuchaiev, B. Ginsburg, O. Olabiyi, K. Briski, J. Cohen, B. Catanzaro, J. Alben, Y. Geifman, E. Chung, and C. Alexiuk (2025) Llama-nemotron: efficient reasoning models. External Links: 2505.00949, Link Cited by: Appendix A, §4.1. [5] C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra (2015-07–09 Jul) Weight uncertainty in neural network. In Proceedings of the 32nd International Conference on Machine Learning, F. Bach and D. Blei (Eds.), Proceedings of Machine Learning Research, Vol. 37, Lille, France, p. 1613–1622. External Links: Link Cited by: §2.3. [6] Bytedance-Seed-Foundation-Code-Team, :, Y. Cheng, J. Chen, J. Chen, L. Chen, L. Chen, W. Chen, Z. Chen, S. Geng, A. Li, B. Li, B. Li, L. Li, B. Liu, J. Liu, K. Liu, Q. Liu, S. Liu, S. Liu, T. Liu, T. Liu, Y. Liu, R. Long, J. Mai, G. Ning, Z. Y. Peng, K. Shen, J. Su, J. Su, T. Sun, Y. Sun, Y. Tao, G. Wang, S. Wang, X. Wang, Y. Wang, Z. Wang, J. Xia, L. Xiang, X. Xiao, Y. Xiao, C. Xi, S. Xin, J. Xu, S. Xu, H. Yang, J. Yang, Y. Yang, J. Yuan, J. Zhang, Y. Zhang, Y. Zhang, S. Zheng, H. Zhu, and M. Zhu (2025) FullStack bench: evaluating llms as full stack coders. External Links: 2412.00535, Link Cited by: §4.4. [7] S. Cao, S. Hegde, D. Li, T. Griggs, S. Liu, E. Tang, J. Pan, X. Wang, A. Malik, G. Neubig, K. Hakhamaneshi, R. Liaw, P. Moritz, M. Zaharia, J. E. Gonzalez, and I. Stoica (2025) SkyRL-v0: train real-world long-horizon agents via reinforcement learning. External Links: Link Cited by: §3.1. [8] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba (2021) Evaluating large language models trained on code. CoRR abs/2107.03374. External Links: Link, 2107.03374 Cited by: Appendix A, §4.1. [9] P. Chen, X. Li, Z. Li, W. Yin, X. Chen, and T. Lin (2026) Exploration vs exploitation: rethinking RLVR through clipping, entropy, and spurious reward. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §1, §4.3. [10] B. Cong, N. Daheim, Y. Shen, R. Yokota, M. E. Khan, and T. Möllenhoff (2025) Improving lora with variational learning. External Links: 2506.14280, Link Cited by: §5.1, footnote 1. [11] G. Cui, Y. Zhang, J. Chen, L. Yuan, Z. Wang, Y. Zuo, H. Li, Y. Fan, H. Chen, W. Chen, Z. Liu, H. Peng, L. Bai, W. Ouyang, Y. Cheng, B. Zhou, and N. Ding (2025) The entropy mechanism of reinforcement learning for reasoning language models. CoRR abs/2505.22617. External Links: Link, Document, 2505.22617 Cited by: Appendix A, Appendix E, §1, §4.1, §4.3. [12] N. Daheim, C. Meister, T. Möllenhoff, and I. Gurevych (2025) Uncertainty-aware decoding with minimum bayes risk. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §5.1. [13] DeepSeek-AI (2025) DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning. External Links: 2501.12948, Link Cited by: §4.1. [14] M. P. Deisenroth, G. Neumann, and J. Peters (2013) A survey on policy search for robotics. Found. Trends Robotics 2 (1-2), p. 1–142. External Links: Link, Document Cited by: §1. [15] P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur (2021) Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, External Links: Link Cited by: §2.3. [16] M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hessel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, C. Blundell, and S. Legg (2018) Noisy networks for exploration. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, External Links: Link Cited by: §2.2. [17] Y. Gan and P. Isola (2026) Neural thickets: diverse task experts are dense around pretrained weights. CoRR abs/2603.12228. External Links: Link, Document, 2603.12228 Cited by: §7. [18] A. Graves (2011) Practical variational inference for neural networks. In Advances in Neural Information Processing Systems, J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger (Eds.), Vol. 24, p. . External Links: Link Cited by: §2.3. [19] J. He, J. Liu, C. Y. Liu, R. Yan, C. Wang, P. Cheng, X. Zhang, F. Zhang, J. Xu, W. Shen, S. Li, L. Zeng, T. Wei, C. Cheng, B. An, Y. Liu, and Y. Zhou (2025) Skywork open reasoner 1 technical report. CoRR abs/2505.22312. External Links: Link, Document, 2505.22312 Cited by: Appendix D, §2.1. [20] D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt (2021) Measuring mathematical problem solving with the MATH dataset. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, J. Vanschoren and S. Yeung (Eds.), External Links: Link Cited by: §4.1. [21] A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi (2020) The curious case of neural text degeneration. In International Conference on Learning Representations, External Links: Link Cited by: §1. [22] J. Hu, M. Liu, X. Lu, F. Wu, Z. Harchaoui, S. Diao, Y. Choi, P. Molchanov, J. Yang, J. Kautz, and Y. Dong (2025) BroRL: scaling reinforcement learning via broadened exploration. CoRR abs/2510.01180. External Links: Link, Document, 2510.01180 Cited by: Appendix D, §2.1, §5.2. [23] N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica (2025) LiveCodeBench: holistic and contamination free evaluation of large language models for code. External Links: Link Cited by: §1, §4.4. [24] T. Jaksch, R. Ortner, and P. Auer (2010) Near-optimal regret bounds for reinforcement learning. J. Mach. Learn. Res. 11, p. 1563–1600. External Links: Link, Document Cited by: §2.2. [25] Y. Jiang, Y. Li, G. Chen, D. Liu, Y. Cheng, and J. Shao (2025) Rethinking entropy regularization in large reasoning models. CoRR abs/2509.25133. External Links: Link, Document, 2509.25133 Cited by: §1, §2.1, §3. [26] M. E. Khan, D. Nielsen, V. Tangkaratt, W. Lin, Y. Gal, and A. Srivastava (2018) Fast and scalable bayesian deep learning by weight-perturbation in adam. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, J. G. Dy and A. Krause (Eds.), Proceedings of Machine Learning Research, p. 2616–2625. External Links: Link Cited by: §2.3. [27] M. E. Khan and H. Rue (2023) The bayesian learning rule. Journal of Machine Learning Research 24 (281), p. 1–46. External Links: Link Cited by: §2.3. [28] J. Knoblauch, J. Jewson, and T. Damoulas (2019) Generalized variational inference: three arguments for deriving new posteriors. External Links: 1904.02063, Link Cited by: §2.3. [29] N. Kohl and P. Stone (2004) Policy gradient reinforcement learning for fast quadrupedal locomotion. In Proceedings of the 2004 IEEE International Conference on Robotics and Automation, ICRA 2004, April 26 - May 1, 2004, New Orleans, LA, USA, p. 2619–2624. External Links: Link, Document Cited by: §2.2. [30] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica (2023) Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, J. Flinn, M. I. Seltzer, P. Druschel, A. Kaufmann, and J. Mace (Eds.), p. 611–626. External Links: Link, Document Cited by: Appendix A, §7. [31] N. Lambert, J. Morrison, V. Pyatkin, S. Huang, H. Ivison, F. Brahman, L. J. V. Miranda, A. Liu, N. Dziri, S. Lyu, Y. Gu, S. Malik, V. Graf, J. D. Hwang, J. Yang, R. L. Bras, O. Tafjord, C. Wilhelm, L. Soldaini, N. A. Smith, Y. Wang, P. Dasigi, and H. Hajishirzi (2024) TÜlu 3: pushing frontiers in open language model post-training. CoRR abs/2411.15124. External Links: Link, Document, 2411.15124 Cited by: §1, §4.1. [32] A. Lewkowycz, A. Andreassen, D. Dohan, E. Dyer, H. Michalewski, V. V. Ramasesh, A. Slone, C. Anil, I. Schlag, T. Gutman-Solo, Y. Wu, B. Neyshabur, G. Gur-Ari, and V. Misra (2022) Solving quantitative reasoning problems with language models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: §4.1. [33] K. Li (2024) Verified taco problems. Note: https://huggingface.co/datasets/likaixin/TACO-verified External Links: Link Cited by: §4.4. [34] W. Lin, M. Schmidt, and M. E. Khan (2020) Handling the positive-definite constraint in the bayesian learning rule. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Proceedings of Machine Learning Research, p. 6116–6126. External Links: Link Cited by: §2.3. [35] J. Liu, Y. Li, Y. Fu, J. Wang, Q. Liu, and Z. Jiang (2025-09) When speed kills stability: demystifying RL collapse from the training-inference mismatch. Note: https://richardli.xyz/rl-collapse Cited by: Appendix I, §3.2. [36] J. Liu and L. Zhang (2025) Code-r1: reproducing r1 for code with reliable rewards. Note: https://github.com/ganler/code-r1 Cited by: §4.4. [37] M. Liu, S. Diao, X. Lu, J. Hu, X. Dong, Y. Choi, J. Kautz, and Y. Dong (2025) ProRL: prolonged reinforcement learning expands reasoning boundaries in large language models. CoRR abs/2505.24864. External Links: Link, Document, 2505.24864 Cited by: §2.1. [38] Z. Liu, X. Li, B. Kang, and T. Darrell (2021) Regularization matters in policy optimization - an empirical study on continuous control. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, External Links: Link Cited by: §1, §2.1, §3. [39] Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin (2025) Understanding r1-zero-like training: a critical perspective. In Second Conference on Language Modeling, External Links: Link Cited by: §1. [40] I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, External Links: Link Cited by: §2.3. [41] M. Luo, S. Tan, J. Wong, X. Shi, W. Y. Tang, M. Roongta, C. Cai, J. Luo, L. E. Li, R. A. Popa, and I. Stoica (2025) DeepScaleR: surpassing o1-preview with a 1.5b model by scaling rl. Note: https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2Notion Blog Cited by: §2.1. [42] Mathematical Association of America (2026) American Invitational Mathematics Examination. External Links: Link Cited by: §1. [43] A. R. Minut, N. Daheim, M. Miani, M. E. Khan, W. Lin, and T. Möllenhoff (2026) SOAP-bubbles: structured weight uncertainty for neural networks. External Links: 2606.23357, Link Cited by: §7. [44] T. Möllenhoff and M. E. Khan (2023) SAM as an optimal relaxation of bayes. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §2.3. [45] M. Noukhovitch, S. Huang, S. Xhonneux, A. Hosseini, R. Agarwal, and A. Courville (2025) Faster, more efficient RLHF through off-policy asynchronous learning. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1. [46] T. Olmo, A. Ettinger, A. Bertsch, B. Kuehl, D. Graham, D. Heineman, D. Groeneveld, F. Brahman, F. Timbers, H. Ivison, J. Morrison, J. Poznanski, K. Lo, L. Soldaini, M. Jordan, M. Chen, M. Noukhovitch, N. Lambert, P. Walsh, P. Dasigi, R. Berry, S. Malik, S. Shah, S. Geng, S. Arora, S. Gupta, T. Anderson, T. Xiao, T. Murray, T. Romero, V. Graf, A. Asai, A. Bhagia, A. Wettig, A. Liu, A. Rangapur, C. Anastasiades, C. Huang, D. Schwenk, H. Trivedi, I. Magnusson, J. Lochner, J. Liu, L. J. V. Miranda, M. Sap, M. Morgan, M. Schmitz, M. Guerquin, M. Wilson, R. Huff, R. L. Bras, R. Xin, R. Shao, S. Skjonsberg, S. Z. Shen, S. S. Li, T. Wilde, V. Pyatkin, W. Merrill, Y. Chang, Y. Gu, Z. Zeng, A. Sabharwal, L. Zettlemoyer, P. W. Koh, A. Farhadi, N. A. Smith, and H. Hajishirzi (2025) Olmo 3. External Links: 2512.13961, Link Cited by: §1, §4.1. [47] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe (2022) Training language models to follow instructions with human feedback. External Links: Link Cited by: §1. [48] M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz (2018) Parameter space noise for exploration. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings, External Links: Link Cited by: §2.2. [49] T. Rückstieß, F. Sehnke, T. Schaul, D. Wierstra, Y. Sun, and J. Schmidhuber (2010) Exploring parameter space in reinforcement learning. Paladyn J. Behav. Robotics 1 (1), p. 14–24. External Links: Link, Document Cited by: §1, §2.2. [50] T. Salimans, J. Ho, X. Chen, and I. Sutskever (2017) Evolution strategies as a scalable alternative to reinforcement learning. CoRR abs/1703.03864. External Links: Link, 1703.03864 Cited by: §2.2. [51] F. Sehnke, C. Osendorfer, T. Rückstieß, A. Graves, J. Peters, and J. Schmidhuber (2010) Parameter-exploring policy gradients. Neural Networks 23 (4), p. 551–559. External Links: Link, Document Cited by: §1, §2.2. [52] A. Setlur, M. Y. R. Yang, C. Snell, J. Greer, I. Wu, V. Smith, M. Simchowitz, and A. Kumar (2025) E3: learning to explore enables extrapolation of test-time compute for llms. CoRR abs/2506.09026. External Links: Link, Document, 2506.09026 Cited by: §2.1. [53] Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, M. Zhang, Y. K. Li, Y. Wu, and D. Guo (2024) DeepSeekMath: pushing the limits of mathematical reasoning in open language models. CoRR abs/2402.03300. External Links: Link, Document, 2402.03300 Cited by: §1, §2.1, §4.1. [54] H. Shen (2025) On entropy control in LLM-RL algorithms. CoRR abs/2509.03493. External Links: Link, Document, 2509.03493 Cited by: §1. [55] Y. Shen, N. Daheim, B. Cong, P. Nickl, G. M. Marconi, C. Bazan, R. Yokota, I. Gurevych, D. Cremers, M. E. Khan, and T. Möllenhoff (2024) Variational learning is effective for large deep networks. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, p. 44665–44686. External Links: Link Cited by: §2.3, §4.1. [56] G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu (2024) HybridFlow: a flexible and efficient rlhf framework. arXiv preprint arXiv: 2409.19256. External Links: Link Cited by: Appendix A, §3.1. [57] A. L. Strehl and M. L. Littman (2005) A theoretical analysis of model-based interval estimation. In Proceedings of the 22nd International Conference on Machine Learning, ICML ’05, New York, NY, USA, p. 856–863. External Links: ISBN 1595931805, Link, Document Cited by: §2.2. [58] F. Stulp and O. Sigaud (2012) Path integral policy improvement with covariance matrix adaptation. In Proceedings of the 29th International Conference on Machine Learning, ICML 2012, Edinburgh, Scotland, UK, June 26 - July 1, 2012, External Links: Link Cited by: §2.2. [59] Y. Sun, Y. Cao, P. Huang, H. Bai, H. Hajishirzi, N. Dziri, and D. Song (2026) RL grokking recipe: how does RL unlock and transfer new algorithms in LLMs?. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §2.1. [60] R. S. Sutton and A. G. Barto (1998) Reinforcement learning - an introduction. Adaptive computation and machine learning, MIT Press. External Links: Link, ISBN 978-0-262-19398-6 Cited by: §1, §2.1. [61] R. S. Sutton (1995) Generalization in reinforcement learning: successful examples using sparse coarse coding. In Advances in Neural Information Processing Systems 8, NIPS, Denver, CO, USA, November 27-30, 1995, D. S. Touretzky, M. Mozer, and M. E. Hasselmo (Eds.), p. 1038–1044. External Links: Link Cited by: §2.2. [62] E. A. Theodorou, J. Buchli, and S. Schaal (2010) A generalized path integral control approach to reinforcement learning. J. Mach. Learn. Res. 11, p. 3137–3181. External Links: Link, Document Cited by: §2.2. [63] H. van Hoof, D. Tanneberg, and J. Peters (2017) Generalized exploration in policy search. Mach. Learn. 106 (9-10), p. 1705–1724. External Links: Link, Document Cited by: §2.2. [64] V. Venkatkrishna, I. Paul, and I. Gurevych (2026) Aletheia: what makes RLVR for code verifiers tick?. CoRR abs/2601.12186. External Links: Link, Document, 2601.12186 Cited by: Appendix A, §1. [65] S. Wang, L. Yu, C. Gao, C. Zheng, S. Liu, R. Lu, K. Dang, X. Chen, J. Yang, Z. Zhang, Y. Liu, A. Yang, A. Zhao, Y. Yue, S. Song, B. Yu, G. Huang, and J. Lin (2025) Beyond the 80/20 rule: high-entropy minority tokens drive effective reinforcement learning for LLM reasoning. CoRR abs/2506.01939. External Links: Link, Document, 2506.01939 Cited by: §1, §1, §2.1. [66] C. J. C. H. Watkins et al. (1989) Learning from delayed rewards. Cited by: §2.2. [67] D. Wu, S. Xia, and Y. Wang (2020) Adversarial weight perturbation helps robust generalization. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Eds.), External Links: Link Cited by: §2.3. [68] F. Wu, W. Xuan, X. Lu, Z. Harchaoui, and Y. Choi (2025) The invisible leash: why RLVR may not escape its origin. CoRR abs/2507.14843. External Links: Link, Document, 2507.14843 Cited by: §1. [69] M. Wu, Z. Zhang, Q. Dong, Z. Xi, J. Zhao, S. Jin, X. Fan, Y. Zhou, H. Lv, M. Zhang, Y. Fu, Q. Liu, S. Zhang, and Q. Zhang (2026) Reasoning or memorization? unreliable results of reinforcement learning due to data contamination. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), p. 33944–33952. External Links: Link, Document Cited by: §1. [70] A. Yang, B. Zhang, B. Hui, B. Gao, B. Yu, C. Li, D. Liu, J. Tu, J. Zhou, J. Lin, K. Lu, M. Xue, R. Lin, T. Liu, X. Ren, and Z. Zhang (2024) Qwen2.5-math technical report: toward mathematical expert model via self-improvement. External Links: 2409.12122, Link Cited by: §1, §4.1. [71] F. Yao, L. Liu, D. Zhang, C. Dong, J. Shang, and J. Gao (2025-08) Your efficient rl framework secretly brings you off-policy rl training. External Links: Link Cited by: Appendix I, §3.2. [72] X. Yao, L. Yu, X. Hu, F. Teng, Q. Cui, J. Zhou, and Y. Liu (2025) The debate on RLVR reasoning capability boundary: shrinkage, expansion, or both? A two-stage dynamic view. CoRR abs/2510.04028. External Links: Link, Document, 2510.04028 Cited by: Appendix E, §2.1. [73] Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, YuYue, W. Dai, T. Fan, G. Liu, J. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, R. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, Y. Song, X. Wei, H. Zhou, J. Liu, W. Ma, Y. Zhang, L. Yan, Y. Wu, and M. Wang (2025) DAPO: an open-source LLM reinforcement learning system at scale. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: Appendix A, §1, §2.1, §2.1, §2.1, §3, §4.1. [74] R. Zhao, A. Meterez, S. M. Kakade, C. Pehlevan, S. Jelassi, and E. Malach (2025) Echo chamber: RL post-training amplifies behaviors learned in pretraining. CoRR abs/2504.07912. External Links: Link, Document, 2504.07912 Cited by: §1. [75] C. Zheng, S. Liu, M. Li, X. Chen, B. Yu, C. Gao, K. Dang, Y. Liu, R. Men, A. Yang, J. Zhou, and J. Lin (2025) Group sequence policy optimization. External Links: 2507.18071, Link Cited by: Appendix A. [76] X. Zhu, M. Xia, Z. Wei, W. Chen, D. Chen, and Y. Meng (2025) The surprising effectiveness of negative reinforcement in LLM reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §1, §1, §2.1. [77] H. Zhuang, Y. Zhou, T. Guo, Y. Huang, F. Liu, K. Song, and X. Zhang (2025) Exploring multi-temperature strategies for token- and rollout-level control in RLVR. CoRR abs/2510.08892. External Links: Link, Document, 2510.08892 Cited by: §1, §2.1, §3.1, §3. Appendix A Hyperparameter Settings Across our experiments, we use the verl Python library [56] for training and vLLM [30] for efficient inference during rollouts and evaluation. We use an on-policy training setup by performing a single gradient update per prompt batch to avoid performance degradations often associated with batch-online methods [64, 75]. Unless otherwise stated, we evaluate with temperature τ=0.6τ=0.6, top-p=0.95top-p\!=\!0.95, top-k=50top-k\!=\!50, and K=8K\!=\!8 responses, computing pass@kpass@k with the unbiased estimator of Chen et al. [8]. We train on an internal cluster of 88 NVIDIA H200 GPUs with 144144 GB of vRAM, and each algorithm takes 24−3224-32 hours to run. We report the per-step FLOPs and times in Table˜3. Despite having near-identical FLOPs cost, our approaches that sample multiple models tend to be ≈1.5×≈\!1.5× slower. This is due to inefficient solutions for sampling from multiple models. Theoretically, the only additional compute over GRPO is the O(model_size)O(model\_size) elementwise noise operations from Equation˜3 and the recomputation of prefills for each of the N (or M) perturbations, both comparatively cheap. As expected, removing the equal-compute constraint on M3PO (G=16,M=4G\!=\!16,M\!=\!4) raises the FLOPs cost by a factor of M. For Polaris, we increased the sampling temperature from 1.01.0 to 1.41.4 when entropy reached 85%85\% of its original value. We experimented with other values like 75%75\% and 50%50\%, but these thresholds either triggered the switch too late and did not allow training to recover, or did not trigger at all. We used an entropy coefficient of γ=1e−3γ\!=\!1e-3 for EntReg, and additionally experimented with γ=5e−3γ\!=\!5e-3 that was too high and completely collapsed training. We use the default hyperparameters swept by the original authors for KL-Cov [11] The ≈1.5×≈\!1.5× wall-clock overhead is entirely a systems cost stemming from suboptimal RL stacks comprising unfused IVON kernels. The larger cost comes from inefficient multi-model sampling in vLLM. Currently, one must perform K inference operations, which reduces the speed gains from vLLM’s continuous batching. Implementing more efficient multi-model sampling strategies could significantly reduce this overhead. Moreover, all 3PO variants also converge faster than GRPO (Figures˜6 and 4), their effective cost-to-target could even fall below GRPO once these implementations mature. 3PO also introduces no additional memory complexity over AdamW since IVON’s optimizer state consists of a momentum and a diagonal Hessian, occupying the same footprint as AdamW’s two moment estimates. Warm-start SFT phase. We use the IVON optimizer with a learning rate of 50.050.0, weight decay 10−810^-8, λ=1010λ=10^10, β1=0.9 _1=0.9, β2=0.9999 _2=0.9999, h0=0.001h_0=0.001, and clipping radius 0.0010.001. We filter the Llama-Nemotron Post-Training Dataset [4] to retain DeepSeek-R1 responses with a context length of up to 40964096 tokens. We use a cosine-decay learning rate schedule with a 10%10\% linear warmup, decaying to 10%10\% of the initial value by the end of training. RLVR phase. We use a token-level GRPO formulation [73], with the lower and upper PPO importance-sampling clip bounds both set to 0.20.2. We use a binary correctness reward (+1/0)(+1/0), a batch size of 3232 prompts with 1616 rollouts each, and one gradient step per rollout batch. For the AdamW runs, we use a learning rate of 10−610^-6, weight decay 0.10.1, β1=0.9 _1=0.9, and β2=0.999 _2=0.999. For IVON, we retain the same hyperparameter configuration as in the warm-start SFT phase, except with a learning rate of 1.01.0 and λ=109λ=10^9. We mask out the sequence-level importance weights that lie outside [0.5−2.0][0.5-2.0] for C3PO. All runs use a constant learning rate schedule with a 6%6\% linear warmup. Method TFLOPs Time (mins) /step /step GRPO 51,719 2.72 B3PO 51,719 2.76 M3PO (G=4,M=4G\!=\!4,M\!=\!4) 51,860 4.02 C3PO (N=4N\!=\!4) 52,140 4.18 M3PO (G=16,M=4G\!=\!16,M\!=\!4) 206,877 9.63 Table 3: FLOPs and wall clock time per step for each algorithm. All methods in our main tables have near-identical FLOPs cost. The 1.5×1.5× wall-clock gap is a systems artifact of suboptimal multi-model RL infrastructure Appendix B Reward and entropy curves We plot the reward and entropy curves for our proposed methods and GRPO across both model families in Figure˜6. Similar to the observations for code generation, our proposed methods consistently converge faster than GRPO across both model families. For Olmo3, C3PO’s entropy declines rapidly, suggesting that it is able to concentrate probability mass on high-reward regions of the solution space. For Qwen2.5-Math, all methods maintain largely similar entropy profiles, indicating that the improved exploration is not simply due to increased randomness in the policy. As mentioned in the main text, we advise against using entropy as a metric for exploration, as it is not always indicative of downstream performance. Figure 6: Reward and entropy curves for both model families. All 3PO methods converge faster than vanilla GRPO. Olmo3-C3PO’s entropy declines rapidly, but all methods maintain largely similar entropy profiles for Qwen2.5-Math. Appendix C Impact of scaling ESS (λ) In the main text, we discussed the importance of λ as a crucial control knob for the amount of noise added to the weights, and briefly ablated its effect on our methods for Olmo3. Here, we present detailed results from this ablation Figure 7: λ scaling curves across all 3PO methods. Adding too much noise with small λ hurts performance, while large λ values can make the sampled models too similar. As noted in Section˜5.1, λ=109λ=10^9 is a good default value for Olmo3, yielding consistent gains for all three methods. λ=108λ=10^8 tends to be very unstable, generally underperforming higher values and completely collapsing for C3PO. The entropy curve for C3PO at λ=108λ=10^8 is clear evidence for why entropy curves in isolation are not a reliable indicator of exploration. Figure 8: Effects of scaling λ for Qwen2.5-Math. Qwen2.5-Math is more sensitive to small λ, where overexploration causes the curves to oscillate. Tuning λ is therefore important. We additionally reproduce the same comparison on B3PO and C3PO for Qwen2.5-Math in Figure˜8. Qwen2.5-Math is less robust to small λ than Olmo3: while B3PO behaves similarly across the two models, C3PO is more sensitive to λ because it samples a fresh model many more times during rollout generation, producing very erratic reward and entropy curves at low λ. The default λ=109λ=10^9 is too noisy for C3PO with Qwen2.5-Math, and λ=5×109λ=5× 10^9 or 101010^10 would likely work best. Appendix D Ablating the number of Monte Carlo samples in M3PO Figure 9: Detailed results for ablating MC samples and chunk size. Using M>1M>1 at a constant rollout budget lowers entropy, but average reward and downstream pass@1 are largely unchanged. Figure 10: G–M tradeoff on Qwen2.5-Math. The variance reduction from larger M does not compensate for the loss in group diversity from smaller G, and vice versa. In Section˜5.2, we analyzed the tradeoff between the group size G and the number of MC samples M for M3PO. Increasing M reduces gradient variance but proportionally increases compute, so we shrink G to keep the rollout budget fixed. However, smaller G also reduces group diversity, which is crucial for grouped-advantage algorithms like GRPO [19, 22]. We presented the pass@1 scores in Figure˜5(a-b); here we additionally plot the corresponding reward and entropy curves in Figure˜9(a-b). Moving from M=1M=1 to M=2M=2 markedly lowers the entropy curve, but the reward curve plateaus only slightly higher and average downstream performance is unaffected. Scaling M further (at the cost of G) progressively closes the entropy gap. However, removing the equal-compute constraint substantially boosts M3PO, lifting reward and trading more entropy for downstream performance. This suggests that scaling M without shrinking the rollout group could be a viable strategy for improving performance, but it comes with a proportional computational cost. We also study the G–M tradeoff for M3PO on Qwen2.5-Math. Following Section˜5.2, we hold the total rollout count fixed, varying M∈1,2,4,8M∈\1,2,4,8\ and reducing the group size proportionally. Figure˜10 shows the same trend as on Olmo3: the gain from variance reduction is roughly cancelled by the loss in group diversity. As discussed in Section˜5.2, realizing the variance-reduction benefit requires investing additional compute. Appendix E Scaling the chunk size in C3PO In this section, we examine the effect of varying the chunk size N in C3PO in more detail. Figure˜9(c) shows the reward and entropy curves corresponding to the analysis in Section˜5.3, and the trends mirror those of the pass@1 curves. While larger N yields marginally higher reward, even N=2N=2 captures almost all of the performance gains. The entropy trend matches Figure˜6: N>1N>1 sharply increases group diversity, enabling the policy to keep discovering high-reward regions of the solution space and concentrate probability mass there, trading entropy for downstream performance [11, 72]. Together, these results suggest that setting N between 22 and 44 and relying on temperature sampling for the remaining rollouts is an effective strategy for improving downstream performance. Figure 11: Effects of a learned noise prior on other algorithms. All three 3PO variants respond similarly to a learned noise prior, possibly due to a relatively isotropic Hessian even after SFT. Appendix F Effects of a learned noise prior on other algorithms In Section˜5.4, we analyzed the impact of initializing the Hessian h with a learned prior obtained during the warm-start SFT phase for C3PO. We present detailed results for this comparison across the other 3PO methods in Figure˜11. Intuitively, a learned prior should stabilize learning more than initializing from scratch and thereby improve performance. We do not, however, observe this behavior for any of our methods; for B3PO, the learned prior even slows convergence. We observed a similar pattern when using Qwen2.5-Math as well. A likely reason is that our SFT phase is relatively short, leaving the Hessian largely isotropic at the start of RL; learning the Hessian from pretraining itself is a promising direction for future work. In the meantime, the absence of a measurable advantage from the learned prior implies that 3PO can be applied to any off-the-shelf checkpoint without first running an IVON-based SFT phase to calibrate the noise distribution. Figure 12: (Left) IVON vs isotropic noise Isotropic noise converges early but plateaus similar to GRPO. (Middle) Effect of Seq-MIS correction. Training completely stalls without the correction due to training-inference mismatch. (Right) Effect of Thompson sampling. Sampling a fresh model for each rollout and acting greedily under it is unstable. Appendix G Multi-seed robustness Here, we validate the robustness of our methods to different seeds. Hoewever, bcause rerunning each method multiple times was not possible due to computational load we only compare GRPO and C3PO on Olmo3 across three seeds on mathematical reasoning tasks. C3PO improves over GRPO by an average of +0.89+0.89 points across the six benchmarks, positive on all three seeds, with a paired-t p=0.027p=0.027. The gains concentrate on harder benchmarks, with AIME’24 improving by +1.8+1.8 points (p=0.023p=0.023). These improvements therefore survive run-to-run variance, and combined with the code-generation results (Section˜4.4) reinforce that our parameter-space exploration methods have the greatest benefits for harder tasks. Appendix H Isotropic initialization vs. learned noise To isolate the contribution of the Hessian-scaled noise (Equation˜4), we run C3PO with isotropic noise of matched magnitude, replacing per-parameter variance with a single global scale. This variant attains an average accuracy of 42.1242.12, below GRPO’s 42.9942.99 and well below full C3PO. As seen in Figure˜12 (Left), its reward does climb faster than GRPO early in training but converges to a similar plateau, whereas full C3PO plateaus above GRPO in late training. Appendix I Ablating the Seq-MIS correction for C3PO C3PO generates the rollouts of a group from N distinct perturbed models, which introduces a training–inference mismatch relative to the single policy assumed by the GRPO ratio in Equation˜1. To show that correcting this mismatch is necessary, we ran an early C3PO configuration that did not apply any such correction (Figure˜12 (Middle)). This run’s training reward curve stayed essentially flat throughout, in contrast to the steadily rising reward of our corrected runs. This mirrors the instability that the broader literature attributes to the training–inference mismatch in RLVR [71, 35], and motivates the correction used in all of our main results. Appendix J A Limiting case of C3PO One limiting case of C3PO is to use N=GN\!=\!G with temperature τ=0.0τ\!=\!0.0, which would draw a fresh model for each rollout and decode greedily under it. However, we found this variant was very unstable (Figure˜12 (Right)), with the reward curve oscillating wildly. Our most stable run at λ=1010λ=10^10 performed markedly worse than C3PO with temperature sampling. Thus, our methods in the main text perform exploration in both the action and parameter spaces: by first sampling a policy from the IVON posterior and further sampling actions under this policy. Appendix K Algorithms 1:Dataset D, train policy πθ _θ, rollout policy πμ _μ 2:group size G, temperature τ, Monte-Carlo samples M 3:(,)←IVON(θ0)(m,\, σ) ( _0) 4:for each batch ℬ∼B do 5: ℒ←0L← 0 6: for m=1,…,Mm=1,…,M do ⊳ Optional. M=1M=1 for B3PO 7: Sample θ^m←+⊙ θ_m + σ , ∼(0,I)z (0,\,I) 8: πμ←θ^m _μ← θ_m ⊳ sync rollout policy 9: for each x∈ℬx do: 10: sample oii=1G∼πμ(⋅∣x,τ)\o_i\_i=1^G _μ(· x,τ) ⊳ generate G completions per prompt 11: end for 12: Compute rewards rir_i and advantages Ai,tA_i,t 13: ℒ←ℒ+ℒGRPO(πθ,Ai,t)L +L_GRPO\! ( _θ,A_i,t ) (cf. Eq. 1) 14: end for 15: ,←IVON.step(∇θℒ)m, $ $ σ$$ .step( _θL) ⊳ Update posterior 16:end for alg Batched noising methods (B3PO and M3PO). Weight perturbations θ θ are sampled once per gradient step; all G rollouts for a batch of prompts are generated from the same model πμ _μ. Model weights are updated by accumulating losses over M Monte Carlo samples; M=1M\!=\!1 for B3PO. 1:Dataset D, train policy πθ _θ, rollout policy πμ _μ 2:group size G, temperature τ, chunk size N 3:(,)←IVON(θ0)(m,\, σ) ( _0) 4:for each batch ℬ∼B do 5: ℛ←∅R← ⊳ initialize response buffer 6: for n=1,…,Nn=1,…,N do 7: Sample θ^n←+⊙ θ_n + σ , ∼(0,I)z (0,\,I) 8: πμ←θ^n _μ← θ_n ⊳ sync rollout policy 9: for each x∈ℬx do: 10: Sample oi(n)i=1G/N∼πμ(⋅∣x,τ) \o_i^(n) \_i=1^G/N _μ(· x,τ); 11: ℛ←ℛ∪oi(n)R ∪ \o_i^(n) \ 12: end for 13: end for⊳ ℛR accumulates G rollouts per prompt from N perturbations 14: Compute rewards rir_i and advantages Ai,tA_i,t 15: ℒ←ℒGRPO(πθ,Ai,t)L _GRPO\! ( _θ,\,A_i,t ) (cf. Eq. 1 with Seq-MIS correction) 16: ,←IVON.step(∇θℒ)m, $ $ σ$$ .step( _θL) ⊳ Update posterior 17:end for alg Chunked noising approach. Rollouts are gathered in a buffer ℛR across N independent weight draws θ^n θ_n, each generating G/NG/N rollouts. The GRPO advantage is calculated on the accumulated buffer.