Paper deep dive
Where does output diversity collapse in post-training?
Constantinos Karouzos, Xingwei Tan, Nikolaos Aletras
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/27/2026, 7:10:57 PM
Summary
The paper investigates the phenomenon of output diversity collapse in post-trained large language models (LLMs), specifically analyzing the Olmo 3 model family. By comparing three parallel lineages—Think (CoT distillation), Instruct (multi-source data), and RL-Zero (direct RL)—the researchers demonstrate that diversity collapse is primarily driven by training data composition rather than the post-training algorithm or the generation format (such as Chain-of-Thought). The study finds that the Think lineage loses significant semantic diversity during the Supervised Fine-Tuning (SFT) stage, whereas the Instruct lineage experiences its primary collapse during Direct Preference Optimization (DPO). Crucially, suppressing the Chain-of-Thought (CoT) reasoning at inference time improves or maintains accuracy on hard tasks but fails to recover the lost diversity, proving that the collapse is embedded in the model weights. The research also decomposes diversity loss into quality-control and residual components, showing that the nature of the collapse is task-dependent.
Entities (17)
Relation Signals (7)
SBERT → measures → Semantic Diversity
confidence 100% · SBERT computes mean pairwise cosine distance of sentence embeddings... capturing semantic diversity
SBERT → measures → Semantic Diversity
confidence 100% · SBERT computes mean pairwise cosine distance of sentence embeddings... capturing semantic diversity
Think → undergoes → Supervised Fine-Tuning
confidence 100% · The Think lineage loses most semantic diversity at supervised fine-tuning
Instruct → undergoes → Direct Preference Optimization
confidence 100% · the effect of DPO is larger in Instruct than in Think
Think → uses → Chain-of-Thought
confidence 100% · Think (chain-of-thought distillation)
Chain-of-Thought → isnotcauseof → Output Diversity Collapse
confidence 95% · showing that the collapse is embedded in the model weights by training data, not imposed by the generation format.
Chain-of-Thought → doesnotcause → Output Diversity Collapse
confidence 90% · showing that the collapse is embedded in the model weights by training data, not imposed by the generation format.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Post-trained language models produce less varied outputs than their base counterparts. This output diversity collapse undermines inference-time scaling methods that rely on varied samples, and risks homogenizing model outputs on creative and value-laden tasks. Prior work attributes collapse to specific post-training methods, without separating the role of training data composition from the method, or the generation format from the model weights. We trace output diversity through three parallel post-training lineages of Olmo 3, Think (chain-of-thought distillation), Instruct (broad multi-source data), and RL-Zero, across 15 tasks and four text diversity metrics. We find that the location of collapse co-varies with data composition: the Think lineage loses most semantic diversity at supervised fine-tuning, and the effect of DPO is larger in Instruct than in Think. Suppressing chain-of-thought reasoning at inference in Think models drops accuracy on hard tasks, yet leaves answer-level diversity unchanged, showing that the collapse is embedded in the model weights by training data, not imposed by the generation format. Decomposing diversity loss on six verifiable tasks into a quality-control component (removal of incorrect outputs) and a residual component (genuine narrowing among correct outputs) reveals that the split is task-dependent, and Think models retain more correct-answer diversity than Instruct despite collapsing more in aggregate. Our results indicate that diversity collapse is determined during training by data composition and cannot be addressed at inference time alone.
Tags
Links
- Source: https://arxiv.org/abs/2604.16027v1
- Canonical: https://arxiv.org/abs/2604.16027v1
Trouble viewing inline? Open PDF directly →
Full Text
99,280 characters extracted from source content.
Expand or collapse full text
Preprint. Under review. Where does output diversity collapse in post-training? Constantinos KarouzosXingwei TanNikolaos Aletras School of Computer Science University of Sheffield, UK kkarouzos1, xingwei.tan, n.aletras@sheffield.ac.uk Abstract Post-trained language models produce less varied outputs than their base counterparts. This output diversity collapse undermines inference-time scaling methods that rely on varied samples, and risks homogenizing model outputs on creative and value-laden tasks. Prior work attributes collapse to specific post-training methods, without separating the role of training data composition from the method, or the generation format from the model weights. We trace output diversity through three parallel post-training lineages of Olmo 3, Think (chain-of-thought distillation), Instruct (broad multi-source data), and RL-Zero, across 15 tasks and four text diversity met- rics. We find that the location of collapse co-varies with data composition: the Think lineage loses most semantic diversity at supervised fine-tuning, and the effect of DPO is larger in Instruct than in Think. Suppressing chain-of-thought reasoning at inference in Think models drops accuracy on hard tasks, yet leaves answer-level diversity unchanged, showing that the collapse is embedded in the model weights by training data, not imposed by the generation format. Decomposing diversity loss on six verifiable tasks into a quality-control component (removal of incorrect outputs) and a residual component (genuine narrowing among correct outputs) reveals that the split is task-dependent, and Think models retain more correct- answer diversity than Instruct despite collapsing more in aggregate. Our results indicate that diversity collapse is determined during training by data composition and cannot be addressed at inference time alone. 1 1 Introduction Large language models (LLMs) rely on post-training to improve helpfulness, safety, and instruction compliance. Post-training combines supervised fine-tuning (SFT; Ouyang et al., 2022) on curated demonstrations, and direct preference optimization (DPO; Rafailov et al., 2023) or reinforcement learning from human feedback (RLHF). However, this results in output diversity collapse, i.e., models produce more uniform outputs than their base coun- terparts across summarization (Kirk et al., 2024b), reasoning (Dang et al., 2025), and open- ended generation (Jiang et al., 2025). Diversity collapse limits self-consistency (Wang et al., 2023), pass@ksampling (Chen et al., 2021), and test-time compute scaling (Snell et al., 2025). Kamigaito et al. (2025) show diversity is the mechanism underlying inference scaling laws. The algorithmic causes are well-understood (Wang et al., 2024; Ma et al., 2025a; GX-Chen et al., 2026), yet diversity collapses across task types. This leads LLMs to produce less diverse outputs than a basic web search (Wright et al., 2026), co-writing with LLMs reduces content diversity (Padmakumar & He, 2024), and single-reward RLHF can amplify majority preferences to near-total dominance (Chakraborty et al., 2024). Yet, prior work attributes collapse to specific algorithms. DPO in narrative generation (Peep- erkorn et al., 2025), the reward step in creative tasks (O’Mahony et al., 2024), and SFT in reasoning (Dang et al., 2025), without investigating the effect of data compositions. Ma 1 Code: https://github.com/ckarouzos/where-diversity-collapses/ 1 arXiv:2604.16027v1 [cs.CL] 17 Apr 2026 Preprint. Under review. KEY EXPERIMENTSLINEAGE & DATAFINDINGS ① Per-stage diversity 4 metrics× 15 tasks at every checkpoint. ② CoT suppression Identical weights, reasoning removed at inference. ③ Quality-filtered diversity Diversity measured only on correct outputs. Olmo 3 Base Think CoT distillation, 2 teachers SFTDPO RL Instruct Multi-source, GPT-3.5/4 SFTDPO RL RL-Zero RL directly from Base RL bypasses SFT & DPO different data ① Collapse location varies Think loses diversity at SFT; Instruct at DPO. Driven by training data. ② Weights, not format Suppressing CoT costs accuracy but does not recover diversity. ③ Task-dependent tradeoff Quality-driven vs. genuine collapse varies by task. Figure 1: Study design. We trace output diversity through three parallel post-training lineages of Olmo 3, to identify where, why, and how much diversity is lost. et al. (2025b) suppress chain-of-thought (CoT; Wei et al., 2022) at inference but measure only accuracy, not diversity. No existing study isolates the role of the training method from the training data, or the generation format from the model weights. Two questions remain open: (1) does the diversity collapse co-vary with the post-training method or with the post-training data composition, and (2) does the CoT format itself constrain diversity at inference, or is the collapse embedded in the model weights? We answer these questions through a controlled experimental setting (Figure 1). We monitor the output diversity of the open weight and data Olmo 3 model family (Olmo et al., 2025), which releases checkpoints of all post-training stages across three parallel lines. Think and Instruct variants share the same post-training recipe (SFT→DPO→RL) but differ in data, while RL-Zero bypasses SFT and DPO entirely. Evaluating 13 models across 15 tasks with four diversity metrics, we show that the same post-training method produces different diversity outcomes depending on the upstream data composition, and that each stage plays a distinct role. Our contributions: • We compare Think vs. Instruct lineages, showing that collapse location depends on data: narrow CoT distillation for Think models is associated with a larger drop at SFT, while the DPO drop is larger in Instruct models (§4.1); • We evaluate Think models with CoT suppressed at inference and find no diversity recovery on any task–stage combinations, while quality drops. Diversity collapse resides in the model weights, not in the CoT generation format (§4.2); •We decompose diversity reduction into a quality-control component (removal of incorrect outputs) and a residual component (genuine narrowing among correct outputs), showing the split is task-dependent (§4.3). 2 Related work The reliability–diversity tradeoff in post-training. Jiang et al. (2025) show that aligned models exhibit high output homogeneity across a wide range of model families and scales. Kirk et al. (2024b) find that RLHF reduces both per input and across input diversity. Human co-writing with aligned models reduces content diversity (Padmakumar & He, 2024), and users brainstorming with ChatGPT produce less semantically distinct ideas (Anderson et al., 2024). In reasoning, SFT improves pass@1 but degrades pass@k(Dang et al., 2025); base models outperform RLVR-trained models at large sample budgets (Yue et al., 2025), and base models produce more diverse outputs (West & Potts, 2025). Peeperkorn et al. (2025) identified DPO as the steepest drop. Karouzos et al. (2026) show that under domain shift the adaptation strategy dominates the alignment objective. Current methods cannot selectively preserve diversity where it is beneficial (Jain et al., 2025). Quality-adjusted diversity shows that preference-tuned models retain higher diversity among high-quality outputs (Shypula et al., 2025), and multi-dimensional linguistic benchmarks find that larger models are often less diverse than smaller ones (Guo et al., 2025b). Automatic diversity metrics lag behind human judgments (Tevet & Berant, 2021), and sampling temperature cannot recover training-induced loss (Verine et al., 2025). Mechanisms and mitigations. DPO’s gradient imbalance suppresses dispreferred re- sponses (Ma et al., 2025a), and likelihood displacement shifts probability to unintended outputs (Razin et al., 2025). KL-regularized RL specifies unimodal targets by construc- tion (GX-Chen et al., 2026), preference collapse arises from KL amplification (Xiao et al., 2 Preprint. Under review. 2024), and chat templates induce diversity collapse (Yun et al., 2025). Training on recursively generated synthetic data causes progressive tail disappearance (Shumailov et al., 2024). Pro- posed mitigations include forward-KL optimization (Wang et al., 2024), entropy-constrained RL (Pan et al., 2026), decoupled regularization (Slocum et al., 2025), game-theoretic SFT (Li et al., 2025c), diversity-aware preference optimization (Li et al., 2025a; Lanchantin et al., 2025), and conformative decoding (Peeperkorn et al., 2025). A single reward function is insufficient to represent diverse human preferences (Chakraborty et al., 2024). 3 Experimental setup 3.1 Models and training lineages We study 13 Olmo 3 checkpoints at the 7B scale. Post-training applies up to three stages, SFT, DPO, and RL, starting from the same base model. Base (1 model). The base model is pretrained on Dolma 3 Mix (6T tokens), midtrained on Dolmino Mix (100B tokens), and context-extended to 65K tokens. Think (3 models: Think-SFT, Think-DPO, Think). SFT trains on∼2.3M synthetic CoT (Wei et al., 2022) reasoning traces using (prompt, completion) pairs from two teachers: QwQ- 32B (Team, 2024) and DeepSeek-R1 (Guo et al., 2025a). DPO uses∼200K Delta Learn- ing (Geng et al., 2025) pairs. The RL stage uses a variation of GRPO (Shao et al., 2024) with verifiable rewards and no KL penalty, and trains on∼105K prompts, to produce Think. Think-not-thinking. To isolate the contribution of the CoT generation format from the learned weights, we additionally evaluate all three Think checkpoints with CoT suppressed by prefilling an empty <think> </think> block, forcing direct answers. Instruct (3 models: Instruct-SFT, Instruct-DPO, Instruct). SFT initializes from Think-SFT, then trains on∼2.2M examples that include function-calling, strip reasoning traces, and draw from multiple sources (GPT-3.5, GPT-4, GPT-4.1; OpenAI et al., 2024) rather than two teachers. DPO (∼260K pairs) uses the same pool of prompts as Think-DPO but with the thinking mode disabled, adding multi-turn and GPT-judged preference pairs. The same RL stage as Think produces the final Instruct model. RL-Zero (6 models). Applies RL training directly to Base, bypassing SFT and DPO. Four Olmo 3 variants target different reward domains: RL-Zero-Math, RL-Zero-Code, RL-Zero-IF, and RL-Zero-General (∼105K prompts each). Two additional Olmo 3.1 variants (RL-Zero- Math 3.1 , RL-Zero-Code 3.1 ) are trained for more steps. 3.2 Tasks and Data Summarization. TL;DR (V ̈ olske et al., 2017), CNN/DailyMail (Nallapati et al., 2016), and XSum (Narayan et al., 2018). Bounded output length controls for length confounds, and multiple valid summaries provide a clear diversity signal. Code. HumanEval (Chen et al., 2021), MBPP (Austin et al., 2021), and CRUXEval (Gu et al., 2024). Outputs can be syntactically different but functionally identical, and RL directly optimizes code tasks. Reasoning. GSM8K (Cobbe et al., 2021), MATH-Algebra, MATH-Geometry (Hendrycks et al., 2021), and TruthfulQA (Lin et al., 2022), the primary Think and RL-Zero training domain. Diversity here measures variation in solution strategy with answers held constant. Instruction following. Alpaca (Taori et al., 2023), open-ended, and IFEval (Zhou et al., 2023), with verifiable format constraints. Creative writing. WritingPrompts (Fan et al., 2018), where diversity is intrinsically desirable. Value pluralism. PRISM (Kirk et al., 2024a) and WildBench (Lin et al., 2025), which test whether alignment imposes a single perspective on contested topics. 3 Preprint. Under review. We measure training–evaluation overlap usingC 13 13-gram matching (Lambert et al., 2025) between the four Dolci post-training datasets and all fifteen evaluation tasks (Appendix J). Nine datasets show negligible overlap (≤2%). HumanEval, CRUXEval, IFEval, MATH- Algebra, MATH-Geometry, and WildBench show elevated overlap (7–30%), traceable to shared upstream data. While we flag these benchmarks, our findings on contaminated tasks are consistent with the patterns on the clean tasks. 3.3 Metrics We measure diversity along four complementary axes (detailed definitions in Appendix B). EAD (Liu et al., 2022) counts uniquen-grams normalized against the expected count under a uniform draw (averaged overn∈1,. . ., 5), capturing lexical diversity. SBERT computes mean pairwise cosine distance of sentence embeddings (all-mpnet-base-v2; Reimers & Gurevych, 2019), capturing semantic diversity (0 = collapse, 1 = dissimilar). For code tasks we additionally report semantic diversity with UniXcoder (Guo et al., 2022) embeddings (Appendix F). NLI scores output pairs with an NLI classifier (roberta-large-mnli; Liu et al., 2019), following Stasaski & Hearst (2022), capturing logical diversity; code tasks are excluded. Vendi Score (Friedman & Dieng, 2023) measures the effective number of dissimilar outputs via eigenvalue entropy of the SBERT similarity kernel (VS=1: identical, VS=K: orthogonal). For code-generation tasks we also report AST subtree diversity, the mean pairwise Jaccard distance on AST subtree multisets (Shypula et al., 2025), on correct outputs only (Appendix F). Quality. For the six tasks with verifiable answers (GSM8K, MATH-Algebra, MATH- Geometry, HumanEval, MBPP, IFEval), we report: accuracy@1 (greedy decoding), majority vote@16 (most frequent answer amongK=16 samples), and pass@16 (at least one correct amongK). For code tasks we use the unbiased pass@kestimator. For IFEval we report strict and loose constraint satisfaction. For the eight tasks without verifiable answers we evaluate quality using LLM-as-judge (gpt-4.1-mini) with established protocols (Appendix D). Quality-filtered diversity. We decompose diversity into a quality-control component (re- moval of incorrect outputs) and a residual component (genuine narrowing among correct outputs).D a (SBERT on allKoutputs) andD c (SBERT on theK c ≥2 correct outputs). The gapD a − D c reflects diversity from error variety;D c captures genuine narrowing among correct solutions. We report analogous Vendi scores V a and V c . For each model–task pair, we generateK=16 outputs per prompt atT=0.6, top-p=0.95. Base recommendsT=1.0; we use matched settings for controlled comparison (Appendix H). For all Think-lineage models, we strip<think>...</think>reasoning traces before computing any metric, so that all diversity and quality scores reflect the final answer only. Implementa- tion details are in Appendix A. 4 Results We present results around three questions. First, where does diversity collapse along each lineage (§4.1; Figure 2, Table 1)? Second, does the CoT generation format itself constrain diversity (§4.2; Figures 4–5)? Third, how much of the observed collapse is attributable to quality control (§4.3; Figures 6–8)? 4.1 Lineage-dependent diversity collapse SFT asymmetry. Think and Instruct share the same three-stage post-training, yet collapse at different stages. Think-SFT loses 62% (Table 1) of Base diversity on average, 24% more than Instruct-SFT (38%), uniformly across all 15 tasks, consistent with completion homogeneity from two teachers rather than prompt overlap. This challenges findings of minimal SFT impact on diversity (Guo et al., 2025b) and suggests that the effect depends on the breadth of the SFT data. Collapse magnitude also scales with task difficulty (Figure 2). Think-SFT retains only 36% of Base diversity on GSM8K (92% accuracy) but 54% on MATH-Geometry (50% accuracy). Easier tasks with a dominant solution strategy collapse the most. Instruct-SFT, 4 Preprint. Under review. Base SFT DPO Final 0.2 0.4 SBERT TL;DR Base SFT DPO Final CNN/DM Base SFT DPO Final XSum Base SFT DPO Final HumanEval Base SFT DPO Final MBPP Base SFT DPO Final CRUXEval Base SFT DPO Final GSM8K Base SFT DPO Final MATH-Alg Base SFT DPO Final MATH-Geo Base SFT DPO Final TruthfulQA Base SFT DPO Final Alpaca Base SFT DPO Final IFEval Base SFT DPO Final WritingPrompts Base SFT DPO Final PRISM Base SFT DPO Final WildBench Base SFT DPO Final 0.25 0.50 0.75 EAD Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final 2.5 5.0 Vendi Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final Base SFT DPO Final BaseThinkThink w/o CoTInstruct Figure 2: SBERT, EAD, and Vendi Score across post-training stages. Think (orange) collapses at SFT; Instruct (blue) at DPO. Think w/o CoT (hollow) tracks Think. despite initializing from the already-collapsed Think-SFT, recovers a median 40% of the lost diversity, likely due to its multi-source data. As Instruct-SFT initializes from Think-SFT, this recovery also reflects the dynamics of retraining a collapsed model. SFT DPO RL Retained Think −62 −4 +438% Instruct −38 −23 −534% RL-Zero(single)93% Table 1: Stage-wise SBERT loss (% of Base, 15-task average). DPO asymmetry. DPO erases more diversity in In- struct than in Think, as Think has already collapsed at SFT, leaving little for DPO to remove. The effect is largest on summarization and code-reasoning tasks, where Instruct-SFT had preserved substantial diver- sity. On three math/code tasks, Think-DPO actually increases diversity slightly, and Instruct-DPO does the same on GSM8K, suggesting that DPO can partially correct a collapsed SFT distribution. 0.70.80.91.01.11.2 NLI diversity PRISM TruthfulQA WildBench Alpaca WritingPrompts TL;DR CNN/DailyMail XSum GSM8K MATH-Algebra MATH-Geometry IFEval InstructThink w/o CoTThinkBase Figure 3: NLI diversity. RL reversal. Think’s RL stage increases semantic diversity on most tasks, primarily code and sum- marization. The recovery is modest (roughly 5% of total diversity lost) but directionally consistent. Both lineages use the same RLVR method, so the asym- metry likely reflects the input state: Think enters RL already at its diversity floor, leaving room for explo- ration, while Instruct enters with residual diversity that RL continues to compress. On GSM8K, Instruct RL erases 37% of Base diversity, the largest single- stage loss outside SFT, as the verifiable reward con- centrates probability on the dominant correct strat- egy. The RLVR stage also produces lexically more uniform outputs (EAD decreases on nearly all tasks), suggesting it standardizes surface form while broad- ening semantic content. Convergence. RL-Zero bypasses both bottlenecks (Figure 2), retaining≥71% of Base diversity (median 94%). Both supervised lineages converge to similar final diversity floors (with Think slightly higher on 11/15 tasks), despite different trajectories: data composition co-varies with when and how sharply diversity is lost. Table 1 summarizes the stage-wise attribution. Full per-task breakdowns are in Appendix I. The collapse is semantic, not lexical (Figure 2). Per input SBERT drops from 0.32 (Base) to 0.12 (Think) and 0.11 (Instruct), and the Vendi Score drops from∼3.4 effective modes to ∼1.8 (final), with near-total collapse on math (GSM8K: 1.3 modes, MATH-Algebra: 1.4), 16 samples carry essentially no more semantic diversity than one. EAD (Figure 2) remains stable or increases, even as semantic diversity drops. Aligned models use varied vocabulary and phrasing to express semantically identical content. Think’s EAD on WritingPrompts rises from 0.23 to 0.80, while SBERT falls from 0.54 to 0.20, a pattern replicated across open-ended tasks. For natural language tasks, NLI diversity (Figure 3) drops on most tasks, 5 Preprint. Under review. GSM8K MATH-Alg MATH-Geo TruthfulQA HumanEval MBPP CRUXEval IFEval 0 25 50 75 100 Accuracy (%) (a) Verifiable SFT GSM8K MATH-Alg MATH-Geo TruthfulQA HumanEval MBPP CRUXEval IFEval DPO GSM8K MATH-Alg MATH-Geo TruthfulQA HumanEval MBPP CRUXEval IFEval Final TL;DR CNN/DM XSum Alpaca WritingPrompts PRISM 0 25 50 75 100 Win rate (%) (b) LLM-judge TL;DR CNN/DM XSum Alpaca WritingPrompts PRISM TL;DR CNN/DM XSum Alpaca WritingPrompts PRISM ThinkThink w/o CoTInstruct Figure 4: Quality of generations for Think, Think-not-thinking, and Instruct, across stages. Top: accuracy on eight verifiable tasks. Bottom: LLM-judge win rates on six tasks. though the gap varies. Post-trained models still make logically distinct claims. The gap is largest for Think models, where CoT reasoning preserves logical structure even as the surface distribution narrows. Value-pluralism tasks suffer the steepest Think collapse (PRISM−78%, TruthfulQA−79%), as narrow two-teacher distillation cannot represent the range of perspectives these tasks require. On PRISM, Think’s NLI (Figure 3) scores remain above 1.0 (net contradictions), meaning the model still samples contradictions despite converged phrasing, though we cannot determine whether this is genuine stance plurality or internal incoherence. Instruct drops NLI below 1.0, indicating homogenization of both form and stance (Figure 3). Think’s NLI remains above the contradiction threshold on value-pluralism and creative tasks where Instruct’s drops below. Creative writing (WritingPrompts) shows the highest Base diversity (6.9 Vendi modes) and the sharpest quality–diversity tension. Think and Instruct both col- lapse to∼0.20 SBERT and∼2.6 modes (−63%), yet achieve>97% pairwise win rate against Base, producing better stories at the cost of formulaic variation. RL-Zero retains∼100% of Base diversity, but wins only∼50%, consistent with the absence of a creative-writing reward signal. NLI diversity remains above 1.0 for all models on WritingPrompts (Think 1.12, Instruct 1.02, RL-Zero 1.15), meaning post-trained models still produce logically distinct narratives despite semantic convergence. Full per-task breakdowns are in Appendix C. 4.2 Think-not-thinking: CoT as reliability, not diversity 20246 WB-Score Base SFT DPO Final BaseThinkw/o CoTInstruct Figure 5: WildBench Score. Think and Instruct differ in both training data and generation format. Think generates CoT reasoning traces before answering, while Instruct answers di- rectly. To isolate the format’s contribution, we evalu- ate all three Think models with CoT suppressed, we refer to these models as Think-not-thinking. This is an out-of-distribution intervention, so we interpret the results as testing whether format removal recover di- versity. Across tasks (Figure 2), removing CoT does not recover diversity. Think-not-thinking SBERT di- versity matches Think, and Instruct shows similarly collapsed diversity. This holds at every stage (SFT, DPO, RLVR) and across every task category. IFEval shows a small increase (+0.025 SBERT), but this is modest relative to the Base-to-Think gap (−0.153). CoT suppression does affect accuracy (Figure 4), with harder tasks losing more: IFEval −8%, GSM8K−18%, MBPP−20%, MATH-Algebra−28%, HumanEval−32%, MATH- Geometry−32%. The quality cost is task-dependent (Figure 4), CoT suppression is negligible for open-ended generation (no change for Alpaca, WritingPrompts−4%) but severe for summarization (CNN/DM−48%) and complex helpfulness (WildBench Score 4.6→1.4, Figure 5). In no case does suppression recover diversity. CoT improves reliability by helping 6 Preprint. Under review. the model execute its learned strategy, especially on hard problems, without broadening the answer-level diversity distribution. The output distribution is equally collapsed whether the model reasons explicitly or answers directly. One exception is WritingPrompts, where removing CoTs slightly increases SBERT diversity (+0.046), suggesting that CoT imposes implicit narrative templates that constrain story generation. NLI diversity reveals a subtler pattern on math tasks: Think-not-thinking produces higher NLI scores than Think (GSM8K: 0.87 vs. 0.70; MATH-Algebra: 0.91 vs. 0.73), despite identical SBERT. Without CoT, final answers are semantically collapsed but logically less entailing. The model generates diverse wrong answers rather than diverse correct strategies, consistent with the accuracy drops. Diversity collapse resides in the learned distribution, not the output format. Narrow two- teacher SFT data reshapes model outputs, and this effect is not reversed by suppressing CoT at inference. This aligns with findings that CoT in post-trained models can function as post- hoc rationalization (Lewis-Lim et al., 2025) and that CoT can be applied selectively (Sprague et al., 2025). The model has already converged on its answer distribution during training. The Think vs Instruct comparison (§4.1) is, therefore, not confounded by the generation format. The diversity difference between lineages reflects data composition. Practitioners cannot recover diversity by switching Think models to direct-answer mode, the cost is paid at training time. We note that we measure final-answer diversity, not reasoning-path diversity. 4.3 Quality-filtered diversity decomposition GSM8KM-AlgM-Geo 0 1 2 Vendi Score HumanEvalMBPPIFEval 0 1 2 3 4 Vendi Score Base RL-Zero Instruct Think V c V a −V c Figure 6: Quality filtered Vendi Score on six verifiable tasks. 0.40.50.60.70.80.91.0 AST Jaccard (D AST c ) HumanEval MBPP (a) Structural diversity 0.050.100.150.200.250.300.35 UniXcoder (D code c ) HumanEval MBPP (b) Semantic code diversity BaseThink-SFT Think-not-thinking-SFT Instruct-SFT RL-Zero-Code 3.1Think-DPO Think-not-thinking-DPO Instruct-DPO Think (final) Think-not-thinking (final) Instruct (final) Figure 7: Code diversity on correct outputs: AST subtree Jaccard (struc- tural) and UniXcoder (semantic) for HumanEval and MBPP. The aggregate diversity reductions combine two ef- fects, elimination of incorrect outputs and genuine narrowing of the correct-answer distribution (Fig- ure 6). We decompose these usingD a ,D c ,V a and V c on six verifiable tasks (GSM8K, MATH-Algebra, MATH-Geometry, HumanEval, MBPP, IFEval). All models achieve 94–97% pass@16 on GSM8K, the un- derlying capability is broadly present. RL-Zero vari- ants also reach 94–97% pass@16 on GSM8K despite 49–61% accuracy@1, confirming the gap is in reliabil- ity, not capability. The difference lies in per-attempt reliability (Think 93% vs. Base 56%), not in whether the knowledge exists. The proportion of collapse attributable to quality control varies by task (Figure 6; Appendix E): on IFEval, 83.4% of theD a drop persists inD c (gen- uine narrowing), while on MBPP 38% is genuine and on HumanEval less than 10%. Math reasoning falls between (57–64% genuine). Code-specific met- rics sharpen this picture: among correct HumanEval outputs, Think produces structurally homogeneous solutions (AST Jaccard=0.53, UniXcoderD c =0.13) while Base/RL-Zero’s correct outputs are structurally diverse (AST Jaccard=0.89 on MBPP; Figure 7). This resolves the tension between diversity collapse is harm- ful and it is just quality control (Lake et al., 2025): both are right, in task-dependent proportions. Even among correct outputs, a narrowing persists: Base maintains 1.7 effective Vendi modes among its ∼8.5/16 correct answers, while both Think and In- struct converge to 1.3–1.6 modes among their correct answers (∼15/16 for GSM8K), while IFEval is higher at 2.1–2.3. In absolute terms, all post-trained models produce near-homogeneous correct outputs, which limits the effectiveness of majority voting (Wang et al., 2023): Think gains just +0.4% on GSM8K (16 near-identical correct answers provide no independent signal), while Base gains +24% and RL-Zero +22–26%. Correct-answer diversity determines how 7 Preprint. Under review. much models benefit from repeated sampling (Snell et al., 2025). On MATH-Algebra, Think- not-thinking and RL-Zero-Math both achieve∼49% accuracy, but RL-Zero-Math has twice the correct-answer diversity and gains +15% from majority voting compared to +7% for Think-not-thinking. The pattern holds across math tasks (Figure 8): at matched accuracy, models with more diverse correct outputs consistently extract more benefit from sampling. On HumanEval, Instruct surpasses Think at pass@16 (98.2 vs. 95.7) despite trailing at pass@1 (81.2 vs. 87.7). The collapsed output distribution means additional samples yield identical solutions. On TruthfulQA, the effect is reversed, majority-voting actually hurts all models (majority vote@16<accuracy@1), because the model converges confidently onto the misconception the question was designed to test. When the dominant mode is wrong, diversity collapse amplifies the error. Figure 8 visualizes this pattern, high-accuracy models cluster near zero MV gain, while lower-accuracy models with diverse correct outputs benefit substantially. Full quality results are in Appendix D; quality-filtered results in Appendix E. 4.4 Cross-cutting patterns 5560657075808590 0 10 20 MV gain (p) GSM8K 505560657075 0 5 10 15 MV gain (p) MATH-Algebra 7.07.58.08.59.09.510.0 Accuracy@1 (%) 2 1 0 MV gain (p) TruthfulQA BaseThink-SFT Think-not-thinking-SFT Instruct-SFT RL-Zero-MathThink-DPO Think-not-thinking-DPO Instruct-DPO Think (final) Think-not-thinking (final) Instruct (final) Figure 8: Accuracy@1 vs. majority- voting gain. The ordering (Base>RL-Zero>Final) holds on average across all 15 tasks, though individual RL- Zero variants exceed Base on tasks aligned with their reward signal (e.g., RL-Zero-IF on IFEval, RL- Zero-Code 3.1 on HumanEval). A model that is low- diversity on one task tends to be low-diversity on all tasks. Output length does not explain diversity ordering (Appendix G). LLM-as-a-judge evaluation (Figure 4) confirms post- training improves quality across all non-verified task categories. CNN/DM and XSum win rates increase from 26–48% (Base) to 83–95% (Think, In- struct), open-ended pairwise win rates exceed 80% for Think on Alpaca and for both Think and Instruct on PRISM. WildBench scores rise from−2.0 (Base) to 6.1 (Instruct). RL-Zero models are tied with Base on WritingPrompts (50% win rate), consistent with the absence of creative-writing reward signals. Diversity reductions coexist with clear quality gains. Among RL-Zero variants, the reward signal type predicts diversity preservation. RL-Zero-IF (instruction-following rewards) retains 99% of Base diversity on average, while RL-Zero- Code retains only 88%. On code tasks specifically, RL-Zero-Code retains less diversity (90%) than RL-Zero-General (100%). Pass/fail execution rewards narrow the solution space more aggressively than general rewards. Mathematical reasoning rewards, which admit diverse solution paths, fall between these extremes. This order (format rewards>math rewards >code rewards) shows that the reward specificity predicts diversity reduction. However, RL-Zero’s diversity advantage comes at a steep quality cost, the RL-Zero range is 49.8-61.0% on GSM8K (vs. 93% Think, 80% Instruct) and 49% on IFEval (vs. 79% Think). 5 Discussion Data composition co-varies with the trajectory, not the floor. Think and Instruct share the same three-stage training yet collapse at different stages. The DPO asymmetry (§4.1) reflects the upstream SFT state more than DPO data differences. Think collapses uniformly across all tasks at SFT, leaving DPO little to remove, while Instruct enters DPO with residual spread that is aggressively narrowed. Despite these different paths, both lineages converge to 1.3–1.6 Vendi modes among correct answers on most verifiable tasks and∼2 modes overall, with IFEval as an outlier at 2.1–2.3. Data composition determines when and how sharply models reach the diversity floor, but not the floor itself. This distinction matters practically, data-level interventions (more teachers, broader sources) can slow the descent 8 Preprint. Under review. but may not raise the final diversity level. Algorithmic changes, switching from reverse to forward KL (Wang et al., 2024), adding entropy constraints (Pan et al., 2026), or removing KL penalties entirely (as in RL-Zero), appear necessary to shift the floor. For SFT data, this suggests that the number of distinct completion sources matters. Practitioners should avoid single-teacher or dual-teacher distillation when output diversity is valued, and instead draw from multiple models with diverse training. Mechanistic interpretation. SFT via cross-entropy loss on narrow data performs maximum- likelihood estimation on a low-entropy target distribution. As two teachers from related training lineages produce completions occupying a restricted region of the output space, the model reproduces this narrow mixture. DPO’s reverse-KL objective is mode-seeking by construction, its gradient is proportional to the implicit reward gap between chosen and rejected outputs. When the model is already collapsed (Think post-SFT), chosen and rejected responses are both near the mode, yielding small gradients and minimal further compression. When the model retains spread (Instruct post-SFT), DPO aggressively downweights the tails. GRPO without KL regularization frees the policy to rediscover modes that SFT and DPO suppressed, provided they receive a positive reward signal. Task-dependent patterns: where diversity loss matters most. On math and reasoning tasks a significant part of diversity reduction reflects removal of incorrect solution paths, as the nar- rowing among correct outputs is modest. On code tasks, less collapse is genuine narrowing, but it still limits pass@kscaling. Summarization shows the largest semantic diversity loss, but this is the cost for large quality gains. Creative writing and value-pluralism are the tasks where the observed diversity loss risks imposing a single perspective. The pattern that emerges is a spectrum, from tasks where collapse is largely helpful (code correctness filtering) to tasks where it is actively harmful (value-laden open-ended generation). Practitioners should assess diversity impact relative to their task characteristics, when selecting post-trained models or applying uniform post-training recipes. From distributional to representational diversity. We capture distributional diversity, i.e. statistical spread along lexical, semantic, and logical axes. This is not a sufficient condition for representational diversity, the presence of outputs reflecting different perspectives or stances. We detect when a model’s output distribution narrows but cannot determine which perspectives are lost. The distinction matters most on value-pluralism tasks. Narrow training data does not just reduce variation, it risks imposing a single perspective on questions where legitimate disagreement exists. A model could maintain high distributional diversity while eliminating viewpoints, or conversely appear collapsed while preserving the stances that matter most. Targeted probes for representational diversity across demographic and cultural dimensions are needed to close this gap. 6 Conclusion We traced output diversity through three parallel post-training lineages of Olmo 3, showing that diversity collapse is shaped by training data composition, not the post-training method alone. The same three-stage recipe (SFT→DPO→RL) produces different collapse trajectories depending on the upstream data: narrow two-teacher distillation drives a steep SFT cliff, while broader multi-source data shifts the sharpest drop to DPO. Suppressing the CoT generation format at inference costs accuracy, but does not recover diversity, confirming that the collapse resides in the learned weights. Decomposing the diversity loss into quality- control and residual components reveals a task-dependent split. On some tasks nearly all narrowing reflects the removal of errors, on others most of it is genuine homogenization among correct outputs. This directly affects inference scaling and majority voting boosts. For practitioners, our results point to two actionable directions: (1) broadening the source distribution for SFT data (more teachers, more styles) can mitigate the steepest collapse, and (2) RL without KL penalties can partially reverse DPO-induced semantic narrowing, though the effect is modest. Future work should investigate reasoning-path diversity (as distinct from final-answer diversity), test data-composition interventions directly, and examine whether the diversity floor we observe can be lowered by changes to the preference- optimization objective. 9 Preprint. Under review. Acknowledgments We would like to thank Samuel Lewis-Lim for his valuable feedback. CK is supported by the Centre for Doctoral Training in Speech and Language Technologies (SLT) and their Applications funded by UK Research and Innovation grant [grant number EP/S023062/1]. XT and NA are supported by the EPSRC [grant number EP/Y009800/1], through funding from Responsible AI UK (KP0016) as a Keystone project. We acknowledge (1) IT Services at the University of Sheffield for the provision of services for high-performance computing; (2) the use of the University of Oxford Advanced Research Computing (ARC) facility; (3) the EuroHPC Joint Undertaking for awarding this project access to the EuroHPC supercomputer LEONARDO, hosted by CINECA (Italy) and the LEONARDO consortium through an EuroHPC Development Access call; (4) the use of resources provided by the Isambard-AI National AI Research Resource (AIRR). Isambard-AI is operated by the University of Bristol and is funded by the UK Government’s Department for Science, Innovation and Technology (DSIT) via UK Research and Innovation; and the Science and Technology Facilities Council [ST/AIRR/I-A-I/1023]. References Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. Homogenization effects of large language models on human creative ideation. In Creativity and Cognition, p. 413–425. ACM, June 2024. doi: 10.1145/3635636.3656204. URLhttp://dx.doi.org/10. 1145/3635636.3656204. 2 Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models, 2021. URLhttps://arxiv.org/abs/2108.07732. 3 Souradip Chakraborty, Jiahao Qiu, Hui Yuan, Alec Koppel, Dinesh Manocha, Furong Huang, Amrit Bedi, and Mengdi Wang. Maxmin-RLHF: Alignment with diverse human preferences. In Forty-first International Conference on Machine Learning, 2024. URLhttps: //openreview.net/forum?id=8tzjEMF0Vq. 1, 3 Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mo- hammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code, 2021. URL https://arxiv.org/abs/2107.03374. 1, 3 Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URL https://arxiv.org/abs/2110.14168. 3 Xingyu Dang, Christina Baek, J Zico Kolter, and Aditi Raghunathan. Assessing diversity col- lapse in reasoning. In Scaling Self-Improving Foundation Models without Human Supervision, 2025. URL https://openreview.net/forum?id=AMiKsHLjQh. 1, 2 Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. In Iryna Gurevych and Yusuke Miyao (eds.), Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 889–898, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1082. URL https://aclanthology.org/P18-1082/. 3 10 Preprint. Under review. Dan Friedman and Adji Bousso Dieng. The vendi score: A diversity evaluation metric for machine learning. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=g97OHbQyk1. 4, 19 Scott Geng, Hamish Ivison, Chun-Liang Li, Maarten Sap, Jerry Li, Ranjay Krishna, and Pang Wei Koh. The delta learning hypothesis: Preference tuning on weak data can yield strong gains. In Second Conference on Language Modeling, 2025. URLhttps://openreview. net/forum?id=9rwtezthwo. 3 Alex Gu, Baptiste Roziere, Hugh James Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida Wang. CRUXEval: A benchmark for code reasoning, understanding and execu- tion. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (eds.), Proceedings of the 41st International Confer- ence on Machine Learning, volume 235 of Proceedings of Machine Learning Research, p. 16568– 16621. PMLR, 21–27 Jul 2024. URLhttps://proceedings.mlr.press/v235/gu24c.html. 3 Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, Ashima Suvarna, Benjamin Feuer, Liangyu Chen, Zaid Khan, Eric Frankel, Sachin Grover, Caroline Choi, Niklas Muennighoff, Shiye Su, Wanjia Zhao, John Yang, Shreyas Pimpalgaonkar, Kartik Sharma, Charlie Cheng-Jie Ji, Yichuan Deng, Sarah Pratt, Vivek Ramanujan, Jon Saad-Falcon, Jeffrey Li, Achal Dave, Alon Albalak, Kushal Arora, Blake Wulfe, Chinmay Hegde, Greg Durrett, Sewoong Oh, Mohit Bansal, Saadia Gabriel, Aditya Grover, Kai-Wei Chang, Vaishaal Shankar, Aaron Gokaslan, Mike A. Merrill, Tatsunori Hashimoto, Yejin Choi, Jenia Jitsev, Reinhard Heckel, Maheswaran Sathiamoorthy, Alexandros G. Dimakis, and Ludwig Schmidt. Openthoughts: Data recipes for reasoning models, 2025. URLhttps: //arxiv.org/abs/2506.04178. 31 Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. UniXcoder: Unified cross-modal pre-training for code representation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7212–7225, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.499. URL https://aclanthology.org/2022.acl-long.499/. 4, 19 Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Honghui Ding, Huazuo Gao, Hui Qu, Hui Li, Jianzhong Guo, Jiashi Li, Jingchang Chen, Jingyang Yuan, Jinhao Tu, Junjie Qiu, Junlong Li, J. L. Cai, Jiaqi Ni, Jian Liang, Jin Chen, Kai Dong, Kai Hu, Kaichao You, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Liang Zhao, Litong Wang, Liyue Zhang, Lei Xu, Leyi Xia, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Meng Li, Miaojun Wang, Mingming Li, Ning Tian, Panpan Huang, Peng Zhang, Qiancheng Wang, Qinyu Chen, Qiushi Du, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, R. J. Chen, R. L. Jin, Ruyi Chen, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shengfeng Ye, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, S. S. Li, Shuang Zhou, Shaoqing Wu, Tao Yun, Tian Pei, Tianyu Sun, T. Wang, Wangding Zeng, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, W. L. Xiao, Wei An, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xiaowen Sun, Xiaoxiang Wang, Xinnan Song, Xinyi Zhou, Xianzu Wang, Xinxia Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Yang Zhang, Yanhong Xu, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Yu, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yuan Ou, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Y. X. Zhu, Yanping Huang, Yaohui 11 Preprint. Under review. Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Ying Tang, Yukun Zha, Yuting Yan, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhicheng Ma, Zhigang Yan, Zhiyu Wu, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Zizheng Pan, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, and Zhen Zhang. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 645 (8081):633–638, September 2025a. ISSN 1476-4687. doi: 10.1038/s41586-025-09422-z. URL http://dx.doi.org/10.1038/s41586-025-09422-z. 3 Yanzhu Guo, Guokan Shang, and Chlo ́ e Clavel. Benchmarking linguistic diversity of large language models. Transactions of the Association for Computational Linguistics, 13:1507–1526, 2025b. doi: 10.1162/tacl.a.47. URL https://aclanthology.org/2025.tacl-1.69/. 2, 4 Anthony GX-Chen, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. KL- regularized reinforcement learning is designed to mode collapse. In The Fourteenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/ forum?id=flBRtdIihA. 1, 2 Nathan Habib, Cl ́ ementine Fourrier, Hynek Kydl ́ ı ˇ cek, Thomas Wolf, and Lewis Tunstall. Lighteval: A lightweight framework for llm evaluation, 2023. URLhttps://github.com/ huggingface/lighteval. 19 Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URLhttps://openreview.net/forum?id=7Bywt2mQsCe. 3 Shomik Jain, Jack Lanchantin, Maximilian Nickel, Karen Ullrich, Ashia Wilson, and Jamelle Watson-Daniels. Llm output homogenization is task dependent, 2025. URLhttps: //arxiv.org/abs/2509.21267. 2 Liwei Jiang, Yuanjun Chai, Margaret Li, Mickel Liu, Raymond Fok, Nouha Dziri, Yu- lia Tsvetkov, Maarten Sap, and Yejin Choi. Artificial hivemind: The open-ended ho- mogeneity of language models (and beyond). In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URL https://openreview.net/forum?id=saDOrrnNTz. 1, 2 Hidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi, and Taro Watan- abe. Diversity explains inference scaling laws: Through a case study of minimum Bayes risk decoding. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mo- hammad Taher Pilehvar (eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 29060–29094, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/2025.acl-long.1410. URLhttps://aclanthology.org/2025.acl-long.1410/. 1 Constantinos Karouzos, Xingwei Tan, and Nikolaos Aletras. An empirical study on preference tuning generalization and diversity under domain shift. arXiv preprint arXiv:2601.05882, 2026. 2 Hannah Rose Kirk, Alexander Whitefield, Paul R ̈ ottger, Andrew Michael Bean, Katerina Margatina, Rafael Mosquera, Juan Manuel Ciro, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024a. URLhttps: //openreview.net/forum?id=DFr5hteojx. 3 Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu. Understanding the effects of RLHF on LLM generalisation and diversity. In The Twelfth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=PXD3FAVHJT. 1, 2, 20 12 Preprint. Under review. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention, 2023. URLhttps://arxiv.org/abs/2309. 06180. 19 Thom Lake, Eunsol Choi, and Greg Durrett. From distributional to overton pluralism: In- vestigating large language model alignment. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the As- sociation for Computational Linguistics: Human Language Technologies (Volume 1: Long Pa- pers), p. 6794–6814, Albuquerque, New Mexico, April 2025. Association for Computa- tional Linguistics. ISBN 979-8-89176-189-6. doi: 10.18653/v1/2025.naacl-long.346. URL https://aclanthology.org/2025.naacl-long.346/. 7 Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James Validad Miranda, Alisa Liu, Nouha Dziri, Xinxi Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, and Hannaneh Hajishirzi. Tulu 3: Pushing frontiers in open language model post-training. In Second Conference on Language Modeling, 2025. URLhttps://openreview. net/forum?id=i1uGbfHHpH. 4, 31 Jack Lanchantin, Angelica Chen, Shehzaad Dhuliawala, Ping Yu, Jason Weston, Sainbayar Sukhbaatar, and Ilia Kulikov. Diverse preference optimization, 2025. URLhttps://arxiv. org/abs/2501.18101. 3 Samuel Lewis-Lim, Xingwei Tan, Zhixue Zhao, and Nikolaos Aletras. Analysing chain of thought dynamics: Active guidance or unfaithful post-hoc rationalisation? In Chris- tos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 29838–29853, Suzhou, China, November 2025. Association for Computational Lin- guistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1516. URL https://aclanthology.org/2025.emnlp-main.1516/. 7 Tianjian Li, Yiming Zhang, Ping Yu, Swarnadeep Saha, Daniel Khashabi, Jason Weston, Jack Lanchantin, and Tianlu Wang. Jointly reinforcing diversity and quality in language model generations, 2025a. URL https://arxiv.org/abs/2509.02534. 3 Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. From crowdsourced data to high-quality benchmarks: Arena- hard and benchbuilder pipeline. In Forty-second International Conference on Machine Learn- ing, 2025b. URL https://openreview.net/forum?id=KfTf9vFvSn. 20, 27 Ziniu Li, Congliang Chen, Tian Xu, Zeyu Qin, Jiancong Xiao, Zhi-Quan Luo, and Ruoyu Sun. Preserving diversity in supervised fine-tuning of large language models. In The Thirteenth International Conference on Learning Representations, 2025c. URLhttps://openreview.net/ forum?id=NQEe7B7bSw. 3 Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. Wildbench: Benchmarking LLMs with challenging tasks from real users in the wild. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/forum?id=MKEHCx25xp. 3, 20, 28 Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URLhttps://aclanthology.org/2022. acl-long.229/. 3 Siyang Liu, Sahand Sabour, Yinhe Zheng, Pei Ke, Xiaoyan Zhu, and Minlie Huang. Re- thinking and refining the distinct metric. In Smaranda Muresan, Preslav Nakov, and 13 Preprint. Under review. Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Com- putational Linguistics (Volume 2: Short Papers), p. 762–770, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-short.86. URL https://aclanthology.org/2022.acl-short.86/. 4, 19 Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach, 2019. URL https://arxiv.org/abs/1907.11692. 4, 19 Li-Chun Lu, Miri Liu, Pin Chun Lu, Yufei Tian, Shao-Hua Sun, and Nanyun Peng. Re- thinking creativity evaluation: A critical analysis of existing creativity evaluations. In Vera Demberg, Kentaro Inui, and Llu ́ ıs Marquez (eds.), Proceedings of the 19th Confer- ence of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), p. 6329–6352, Rabat, Morocco, March 2026. Association for Computa- tional Linguistics. ISBN 979-8-89176-380-7. doi: 10.18653/v1/2026.eacl-long.297. URL https://aclanthology.org/2026.eacl-long.297/. 20 Qinwei Ma, Jingzhe Shi, Can Jin, Jenq-Neng Hwang, Serge Belongie, and Lei Li. Gradient imbalance in direct preference optimization, 2025a. URLhttps://arxiv.org/abs/2502. 20847. 1, 2 Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, and Matei Zaharia. Reasoning models can be effective without thinking, 2025b. URLhttps://arxiv.org/ abs/2504.09858. 1 Ramesh Nallapati, Bowen Zhou, Cicero dos Santos,C ̧a ̆ glar Gu ̇lc ̧ehre, and Bing Xiang. Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Stefan Riezler and Yoav Goldberg (eds.), Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, p. 280–290, Berlin, Germany, August 2016. Association for Computational Linguistics. doi: 10.18653/v1/K16-1028. URLhttps: //aclanthology.org/K16-1028/. 3 Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, p. 1797–1807, Brussels, Belgium, October-November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1206. URL https://aclanthology.org/D18-1206/. 3 NVIDIA. Nemotron 3 Nano: Open, efficient mixture-of-experts hybrid Mamba-Transformer model for Agentic reasoning, 2025. URLhttps://research.nvidia.com/labs/nemotron/ files/NVIDIA-Nemotron-3-Nano-Technical-Report.pdf. Technical report. 31 Team Olmo, :, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A. Smith, and Hannaneh Hajishirzi. Olmo 3, 2025. URL https://arxiv.org/abs/2512.13961. 2, 31 Laura O’Mahony, Leo Grinsztajn, Hailey Schoelkopf, and Stella Biderman. Attributing mode collapse in the fine-tuning of large language models. In ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024. URLhttps: //openreview.net/forum?id=3pDMYjpOxk. 1 14 Preprint. Under review. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brak- man, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Sim ́ on Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan,Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo,Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David M ́ ely, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Fran- cis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thomp- son, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cer ́ on Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sher- win Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774. 3 Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155. 1 Vishakh Padmakumar and He He. Does writing with language models reduce content diversity? In The Twelfth International Conference on Learning Representations, 2024. URL 15 Preprint. Under review. https://openreview.net/forum?id=Feiz5HtCD0. 1, 2 Haihui Pan, Yuzhong Hong, Shaoke Lv, Junwei Bao, Hongfei Jiang, and Yang Song. Quality- constrained entropy maximization policy optimization for llm diversity, 2026. URL https://arxiv.org/abs/2602.15894. 3, 9 Max Peeperkorn, Tom Kouwenhoven, Dan Brown, and Anna Jordanous. Mind the gap: Conformative decoding to improve output diversity of instruction-tuned large language models, 2025. URL https://arxiv.org/abs/2507.20956. 1, 2, 3 Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=HPuSIXJaa9. 1 Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen, Sanjeev Arora, and Boris Hanin. Unintentional unalignment: Likelihood displacement in direct preference opti- mization. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=uaMSBJDnRv. 2 Nils Reimers and Iryna Gurevych. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceed- ings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 3982– 3992, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1410. URL https://aclanthology.org/D19-1410/. 4, 19 Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathe- matical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. 3 Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data. Nature, 631(8022): 755–759, 2024. 3 Alexander Shypula, Shuo Li, Botong Zhang, Vishakh Padmakumar, Kayo Yin, and Osbert Bastani. Evaluating the diversity and quality of LLM generated content. In Second Confer- ence on Language Modeling, 2025. URLhttps://openreview.net/forum?id=O7bF6nlSOD. 2, 4, 20 Stewart Slocum, Asher Parker-Sartori, and Dylan Hadfield-Menell. Diverse preference learning for capabilities and alignment. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=pOq9vDIYev. 3 Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps:// openreview.net/forum?id=4FWAwZtd2n. 1, 8 Zayne Rea Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/ forum?id=w6nlcS8Kkn. 7 Katherine Stasaski and Marti Hearst. Semantic diversity in dialogue with natural language inference. In Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (eds.), Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 85–98, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.6. URL https://aclanthology.org/2022.naacl-main.6/. 4, 19 16 Preprint. Under review. Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanfordalpaca, 2023. 3 Qwen Team. Qwq: Reflect deeply on the boundaries of the unknown, 2024. 3 Guy Tevet and Jonathan Berant. Evaluating the evaluation of diversity in natural language generation. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty (eds.), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, p. 326–346, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.25. URLhttps://aclanthology.org/2021.eacl-main.25/. 2 Shubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, and Igor Gitman. Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data. arXiv preprint arXiv:2410.01560, 2024. 31 Alexandre Verine, Florian Le Bronnec, Kunhao Zheng, Alexandre Allauzen, Yann Cheva- leyre, and benjamin negrevergne. Improving diversity in language models: When tem- perature fails, change the loss. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=RsyMfsqzeG. 2 Michael V ̈ olske, Martin Potthast, Shahbaz Syed, and Benno Stein. TL;DR: Mining Reddit to learn automatic summarization. In Lu Wang, Jackie Chi Kit Cheung, Giuseppe Carenini, and Fei Liu (eds.), Proceedings of the Workshop on New Frontiers in Summarization, p. 59–63, Copenhagen, Denmark, September 2017. Association for Computational Linguistics. doi: 10.18653/v1/W17-4508. URL https://aclanthology.org/W17-4508/. 3 Chaoqi Wang, Yibo Jiang, Chenghao Yang, Han Liu, and Yuxin Chen. Beyond reverse KL: Generalizing direct preference optimization with diverse divergence constraints. In The Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview. net/forum?id=2cRzmWXK9N. 1, 3, 9 Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Repre- sentations, 2023. URL https://openreview.net/forum?id=1PL1NIMMrw. 1, 7 Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. In Alice H. Oh, Alekh Agarwal, Danielle Belgrave, and Kyunghyun Cho (eds.), Advances in Neural Information Processing Systems, 2022. URLhttps://openreview. net/forum?id=VjQlMeSBJ. 2, 3 Peter West and Christopher Potts. Base models beat aligned models at randomness and creativity. In Second Conference on Language Modeling, 2025. URLhttps://openreview. net/forum?id=vqN8uom4A1. 2 Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Chris- tensen, Chan Young Park, and Isabelle Augenstein. Epistemic diversity and knowledge collapse in large language models, 2026. URL https://arxiv.org/abs/2510.04226. 1 Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, and Weijie J Su. On the algorithmic bias of aligning large language models with rlhf: Preference collapse and matching regularization. arXiv preprint arXiv:2405.16455, 2024. 2 Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Yang Yue, Shiji Song, and Gao Huang. Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=4OsgYD7em5. 2 17 Preprint. Under review. Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang. The price of for- mat: Diversity collapse in LLMs. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Findings of the Association for Computational Linguis- tics: EMNLP 2025, p. 15454–15468, Suzhou, China, November 2025. Association for Com- putational Linguistics. ISBN 979-8-89176-335-7. doi: 10.18653/v1/2025.findings-emnlp. 836. URL https://aclanthology.org/2025.findings-emnlp.836/. 3 Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. Wild- chat: 1m chatGPT interaction logs in the wild. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=Bl8u7ZRlbM. 31 Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=uccHPGDlao. 20, 27 Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models, 2023. URL https://arxiv.org/abs/2311.07911. 3 18 Preprint. Under review. A Implementation details We generate outputs using vLLM (Kwon et al., 2023) and lighteval (Habib et al., 2023). For each model–task pair, we sampleK=16 outputs per prompt (N=500 prompts; full dataset for Math-Geometry, IFEval, HumanEval, and TruthfulQA) with a 32,768-token generation length. All four diversity metrics (EAD, SBERT, NLI, Vendi Score) operate on the same post-stripping text. Table A lists all evaluation tasks with their sample sizes. CategoryTaskNCategoryTaskNCategoryTaskN Summarization TL;DR500Reasoning GSM8K500CodeHumanEval 164 CNN/DM 500MATH-Alg500MBPP500 XSum500MATH-Geo479CRUXEval500 InstructionAlpaca500TruthfulQA817Value plur. PRISM500 IFEval541CreativeWrtPrompts 500WildBench500 Table 2: Evaluation tasks grouped by category. B Metric definitions For a given prompt, the model generatesKoutputso 1 ,. . .,o K . All metrics are computed per prompt and then averaged over prompts. EAD (lexical diversity) Expectation-Adjusted Distinctn-grams (Liu et al., 2022) counts the number of uniquen- grams in the output set, normalized by the expected number of uniquen-grams under a uniform draw from a vocabulary of sizeV. For a total ofT n-gram tokens withU unique types,EAD n = U V· 1− ( V−1 V ) T , whereVis auto-detected from the model’s tokenizer vocabulary. The denominator corrects for length bias: longer outputs are expected to contain more uniquen-grams by chance. We average acrossn∈1,. . ., 5and clip to[0, 1]: D EAD = 1 5 ∑ 5 n=1 EAD n . SBERT (semantic diversity) We encode each outputo i withall-mpnet-base-v2(Reimers & Gurevych, 2019) to obtain L2-normalized embeddingse i . Semantic diversity is the mean pairwise cosine distance: D SBERT =1− 2 K(K−1) ∑ i<j cos(e i ,e j ). Values near 0 indicate semantic collapse (all outputs map to the same region of embedding space); values near 1 indicate highly dissimilar outputs. For code tasks we additionally report diversity using UniXcoder (Guo et al., 2022), a code-aware encoder that captures structural similarity beyond surface tokens. NLI (logical diversity) Following Stasaski & Hearst (2022), we score output pairs with a natural language inference classifier (roberta-large-mnli; Liu et al., 2019). For each ordered pair(o i ,o j ), the model predicts a probability distribution overentailment, neutral, contradiction. We compute a directional similarity score asP(entailment)− P(contradiction), then symmetrize by averaging both orderings:s ij = 1 2 P ent (o i | o j )− P con (o i | o j ) + P ent (o j | o i )− P con (o j | o i ) . Since NLI models are trained on single sentences rather than full paragraphs, we align sentences by position across outputs. The diversity score is:D NLI =1− 2 K(K−1) ∑ i<j s ij . D NLI near 0 indicates mutual entailment (collapse), near 1 indicates neutrality, and values above 1 indicate net contradiction (the outputs make mutually inconsistent claims). Code tasks are excluded as NLI is not meaningful for program text. Vendi Score The Vendi Score (Friedman & Dieng, 2023) measures the effective number of dissimilar elements via the eigenvalue entropy of a similarity kernel. We reuse the SBERT cosine 19 Preprint. Under review. similarity matrix. GivenKoutputs with L2-normalized embeddings, we form the Gram matrixGwhereG ij = cos(e i ,e j )and trace-normalize it asP = G/K. The Vendi Score is VS = exp ( − ∑ i λ i logλ i ) , whereλ i are the eigenvalues ofP. VS=1 when all outputs are identical (rank-1 kernel) and VS=Kwhen all outputs are orthogonal (full-rank uniform spectrum). Because the Vendi Score shares the SBERT kernel, agreement between VS andD SBERT is expected rather than independent confirmation; VS adds the interpretable “effective number of modes” framing. AST subtree diversity (structural, code only) For code-generation tasks (HumanEval, MBPP), we measure structural diversity via the mean pairwise Jaccard distance on AST subtree multisets (subtree height≤4; Shypula et al., 2025). We parse each output into a Python AST, extract all subtrees up to height 4, represent each output as a multiset of subtree hashes, and computeD AST (o i ,o j ) =1− |S i ∩S j | |S i ∪S j | , whereS i is the multiset of subtree hashes for outputo i . This metric is reported on correct (executable, test-passing) outputs only, to capture genuine structural variation among working solutions. Unparseable outputs are excluded. LLM-as-a-Judge quality For the eight tasks without verifiable answers, we evaluate quality using established LLM- as-judge frameworks withgpt-4.1-minivia the OpenAI Batch API. Summarization (TL;DR, CNN/DM, XSum): pairwise win-rate against reference summaries, following Kirk et al. (2024b). Instruction following and value pluralism (Alpaca, PRISM): pairwise comparison against Base using MT-Bench prompts (Zheng et al., 2023). Creative writing (WritingPrompts): pairwise comparison using Arena-Hard creative writing prompts (Li et al., 2025b). Wild- Bench: checklist-guided WB-Score (Lin et al., 2025). We note that LLM-judge evaluation of creative and value-laden tasks has known limitations (Lu et al., 2026); we report these results as supplementary context for our diversity findings rather than as primary evidence. C Per-task diversity results Tables 3–6 report per input diversity for each of the four metrics across all 15 tasks and 16 models (13 standard + 3 Think w/o CoT). Table 3 reports SBERT cosine distance, our primary semantic diversity measure. Table 4 reports Expected Agreement Diversity (EAD), a lexical overlap metric. Table 5 reports NLI-based diversity, which captures inferential disagreement between output pairs; code tasks are excluded as NLI is not meaningful for program text. Table 6 reports Vendi Score, the effective number of distinct semantic modes among the K=16 outputs. D Quality results Tables 7–13 report task performance for all 16 models, across all 15 tasks. Table 7 reports reasoning quality on four tasks (GSM8K, MATH-Algebra, MATH-Geometry, TruthfulQA) with accuracy@1, majority vote@16, and pass@16. Table 8 reports code generation quality (pass@kfork∈1, 5, 10, 16) on HumanEval and MBPP. Table 9 reports IFEval constraint satisfaction with strict and loose accuracy@1, pass@16, and consistency@16. Table 10 reports CruxEval output-prediction accuracy. For tasks without verifiable answers, we use LLM-as-judge evaluation with gpt-4.1-mini via the OpenAI Batch API. Table 11 reports pairwise win rates against reference summaries following Kirk et al. (2024b). Table 12 reports pairwise win rates against the Base model using the MT-Bench prompt (Zheng et al., 2023) for Alpaca and PRISM and the Arena-Hard creative writing prompt (Li et al., 2025b) for WritingPrompts. Table 13 reports checklist- guided WB-Score (Lin et al., 2025). We note that LLM-judge evaluation of creative and value- laden tasks has known limitations (Lu et al., 2026); we report these results as supplementary context for our diversity findings rather than as primary evidence. 20 Preprint. Under review. SummarizationInstruction F.Creative Wr.Value Pluralism TL;DRCNN/DMXSumAlpacaIFEvalWritingPromptsPRISMWildBench Base0.3530.2790.4510.3190.3490.5400.4080.335 Instruct-SFT0.2680.2230.2820.1700.1720.2760.1410.129 Instruct-DPO0.2020.0750.0830.1200.1540.2250.0960.122 Instruct (final)0.2070.0720.0810.1130.1540.2020.0900.118 Think-SFT0.1680.0830.0900.1410.1910.2400.1000.160 Think-DPO0.1590.0590.0640.1180.1650.2050.0890.154 Think (final)0.1610.0910.0920.1460.1960.1990.0910.173 Think-SFT w/o CoT0.2930.2490.2020.1370.1960.2660.1140.191 Think-DPO w/o CoT0.3440.1760.1300.1040.1570.2230.1000.156 Think w/o CoT0.3230.2200.1670.1610.2210.2450.1020.181 RL-Zero-Math0.3360.2010.4360.3090.3180.5430.3930.313 RL-Zero-Code0.3270.1930.4220.1780.2870.5330.3670.262 RL-Zero-IF0.3330.2100.4290.1760.3970.5460.4000.300 RL-Zero-General0.3090.1840.4040.1550.2840.5230.3720.279 RL-Zero-Math 3.1 0.3300.2000.4320.3190.3240.5460.3980.316 RL-Zero-Code 3.1 0.3280.1960.4300.3140.3250.5390.3940.315 ReasoningCode GSM8KMATH-AlgMATH-GeoTruthfulQAHumanEvalMBPPCRUXEval Base0.1720.1460.1980.3530.4110.2910.239 Instruct-SFT0.1050.1320.1790.3270.1120.1110.218 Instruct-DPO0.1410.0710.0960.1580.0950.0730.068 Instruct (final)0.0780.0570.1010.1150.0930.0690.062 Think-SFT0.0610.0540.1070.1190.1090.0810.095 Think-DPO0.0520.0610.1140.0740.0810.0840.076 Think (final)0.0510.0620.1220.0750.1170.0890.090 Think-SFT w/o CoT0.0570.0660.0980.1060.0550.0840.084 Think-DPO w/o CoT0.0450.0580.0770.0850.0620.0830.064 Think w/o CoT0.0520.0640.0890.0890.0600.0830.071 RL-Zero-Math0.1540.1440.1810.3520.4210.2740.222 RL-Zero-Code0.1560.1440.1830.3480.4640.2380.149 RL-Zero-IF0.1770.1430.1990.3570.3360.2970.491 RL-Zero-General0.1330.1240.1660.3260.4680.2720.198 RL-Zero-Math 3.1 0.1830.1400.1830.3580.4600.2920.207 RL-Zero-Code 3.1 0.1730.1390.1780.3490.4390.2610.209 Table 3: Per-input SBERT diversity (all-mpnet-base-v2). 21 Preprint. Under review. SummarizationInstruction F.Creative Wr.Value Pluralism TL;DRCNN/DMXSumAlpacaIFEvalWritingPromptsPRISMWildBench Base0.370.370.670.510.440.230.240.30 Instruct-SFT0.690.430.580.570.620.720.680.68 Instruct-DPO0.710.530.510.560.710.800.720.76 Instruct (final)0.680.500.480.520.670.790.700.74 Think-SFT0.760.590.580.610.580.730.720.63 Think-DPO0.790.620.630.690.740.830.760.75 Think (final)0.780.590.590.650.680.800.740.71 Think-SFT w/o CoT0.440.420.450.650.610.750.710.59 Think-DPO w/o CoT0.700.560.610.670.720.810.730.69 Think w/o CoT0.560.460.490.630.640.770.700.64 RL-Zero-Math0.410.380.640.490.550.330.360.44 RL-Zero-Code0.470.390.660.580.600.400.430.53 RL-Zero-IF0.460.410.660.470.550.330.390.50 RL-Zero-General0.450.380.650.620.620.360.440.51 RL-Zero-Math 3.1 0.340.360.620.530.530.300.350.43 RL-Zero-Code 3.1 0.350.360.610.560.540.290.350.45 ReasoningCode GSM8KMATH-AlgMATH-GeoTruthfulQAHumanEvalMBPPCRUXEval Base0.450.450.400.460.570.590.31 Instruct-SFT0.380.450.510.570.480.510.57 Instruct-DPO0.470.440.560.650.570.570.58 Instruct (final)0.360.390.520.640.550.540.55 Think-SFT0.360.300.410.620.430.430.48 Think-DPO0.360.320.460.640.420.480.50 Think (final)0.320.320.450.610.460.460.50 Think-SFT w/o CoT0.410.430.500.690.400.540.52 Think-DPO w/o CoT0.370.410.490.720.460.580.51 Think w/o CoT0.400.430.490.670.450.550.50 RL-Zero-Math0.410.490.490.570.590.610.41 RL-Zero-Code0.470.490.510.580.560.600.41 RL-Zero-IF0.460.430.500.570.610.640.55 RL-Zero-General0.450.450.470.520.570.590.45 RL-Zero-Math 3.1 0.410.460.480.540.570.600.40 RL-Zero-Code 3.1 0.430.470.490.540.560.590.43 Table 4: Per-input EAD diversity. 22 Preprint. Under review. SummarizationInstruction. F.Creative Wr.Value Pluralism TL;DR CNN/DM XSum Alpaca IFEval WritingPrompts PRISM WildBench Base0.951.041.090.681.051.161.091.06 Instruct-SFT0.900.710.990.780.971.020.931.05 Instruct-DPO0.860.840.770.770.981.050.971.06 Instruct (final)0.840.790.720.730.931.020.951.05 Think-SFT1.020.930.920.891.011.131.041.09 Think-DPO1.060.930.930.931.061.181.071.12 Think (final)1.030.900.890.851.001.121.041.09 Think-SFT w/o CoT1.040.980.990.961.011.121.041.10 Think-DPO w/o CoT0.980.960.970.991.061.181.081.10 Think w/o CoT1.000.980.970.911.001.091.021.09 RL-Zero-Math0.920.901.050.691.051.161.091.08 RL-Zero-Code0.900.891.040.971.051.141.071.08 RL-Zero-IF0.890.851.040.680.891.151.061.01 RL-Zero-General0.890.891.040.851.021.141.061.06 RL-Zero-Math 3.1 0.920.901.050.691.051.151.081.07 RL-Zero-Code 3.1 0.910.891.050.741.061.151.081.07 Reasoning GSM8K MATH-Alg MATH-Geo TruthfulQA Base1.081.001.130.97 Instruct-SFT0.771.011.100.88 Instruct-DPO0.770.760.880.91 Instruct (final)0.730.760.890.90 Think-SFT0.770.720.850.98 Think-DPO0.730.720.860.98 Think (final)0.700.730.860.99 Think-SFT w/o CoT0.900.871.001.03 Think-DPO w/o CoT0.810.901.021.05 Think w/o CoT0.870.911.031.00 RL-Zero-Math1.050.991.090.97 RL-Zero-Code1.050.981.090.96 RL-Zero-IF1.010.961.100.95 RL-Zero-General1.020.951.080.94 RL-Zero-Math 3.1 1.060.981.100.97 RL-Zero-Code 3.1 1.050.981.090.97 Table 5: Per-input NLI diversity. Code tasks excluded. 23 Preprint. Under review. SummarizationInstruction F.Creative Wr.Value Pluralism TL;DR CNN/DM XSum Alpaca IFEval WritingPrompts PRISM WildBench Base4.23.25.22.23.86.94.63.5 Instruct-SFT3.02.43.22.12.33.22.01.9 Instruct-DPO2.41.51.61.82.12.81.71.9 Instruct (final)2.51.51.51.72.22.61.61.8 Think-SFT2.21.61.62.02.42.91.72.2 Think-DPO2.21.41.41.92.32.61.62.1 Think (final)2.21.61.62.02.52.61.62.3 Think-SFT w/o CoT3.02.52.22.02.43.11.82.3 Think-DPO w/o CoT2.82.01.81.72.12.71.62.0 Think w/o CoT3.02.32.02.02.42.81.72.2 RL-Zero-Math3.92.44.92.33.57.04.53.3 RL-Zero-Code3.82.34.72.13.26.84.22.9 RL-Zero-IF3.82.44.82.04.47.04.53.1 RL-Zero-General3.62.34.52.03.26.74.23.0 RL-Zero-Math 3.1 3.82.44.82.23.67.04.63.3 RL-Zero-Code 3.1 3.82.44.82.33.56.94.53.3 ReasoningCode GSM8K MATH-Alg MATH-Geo TruthfulQA HumanEval MBPP CRUXEval Base2.12.02.43.82.52.82.5 Instruct-SFT1.71.92.33.41.71.72.2 Instruct-DPO1.91.51.62.01.71.51.5 Instruct (final)1.51.41.71.71.61.51.4 Think-SFT1.41.41.71.81.61.51.6 Think-DPO1.31.41.81.51.51.61.5 Think (final)1.31.41.81.51.71.61.6 Think-SFT w/o CoT1.41.41.61.71.41.61.6 Think-DPO w/o CoT1.31.41.51.61.41.51.4 Think w/o CoT1.31.41.61.61.41.51.5 RL-Zero-Math2.02.02.33.82.62.72.3 RL-Zero-Code2.02.02.33.73.02.51.9 RL-Zero-IF2.11.92.43.82.02.73.9 RL-Zero-General1.91.82.23.52.92.62.2 RL-Zero-Math 3.1 2.21.92.33.92.92.82.1 RL-Zero-Code 3.1 2.11.92.23.82.72.62.1 Table 6: Per-input Vendi Score diversity. 24 Preprint. Under review. GSM8KMATH-Algebra MATH-GeometryTruthfulQA acc mv pass acc mv pass acc mvpassacc mv pass Base56.0 80.4 94.8 50.0 59.4 75.6 20.5 24.650.310.0 8.6 28.5 Instruct-SFT73.4 84.4 95.4 56.2 68.2 83.2 26.5 36.761.09.47.2 24.0 Instruct-DPO77.2 86.4 96.2 51.4 65.8 81.6 23.0 35.754.98.26.7 20.8 Instruct (final)80.4 87.6 95.2 70.8 75.0 81.2 42.6 54.363.38.18.0 19.6 Think-SFT92.0 93.4 97.0 76.4 77.2 78.8 50.5 54.759.59.77.0 21.5 Think-DPO85.2 89.4 95.6 74.6 77.0 78.2 50.5 54.561.67.06.6 13.8 Think (final)93.0 93.4 96.4 76.8 77.6 78.8 51.1 55.359.78.66.9 19.3 Think-SFT w/o CoT76.6 82.4 94.6 56.4 63.8 75.0 27.6 29.946.88.67.6 16.3 Think-DPO w/o CoT 70.0 79.4 94.6 47.4 52.8 67.8 19.6 23.237.27.15.5 12.1 Think w/o CoT74.6 82.8 94.0 48.6 55.4 68.2 19.6 25.538.89.47.8 16.8 RL-Zero-Math61.0 83.2 95.8 49.4 64.6 79.2 22.8 27.357.69.77.6 29.4 RL-Zero-Code58.2 83.8 96.4 51.2 63.4 80.8 23.0 28.257.812.2 8.2 30.6 RL-Zero-IF49.8 75.0 94.2 48.2 60.8 77.4 21.3 25.551.610.2 9.3 28.8 RL-Zero-General61.0 82.8 97.0 54.0 64.8 79.6 24.4 29.657.210.9 7.7 28.8 RL-Zero-Math 3.1 55.2 80.6 96.2 53.6 66.0 78.8 20.9 28.455.710.6 8.4 28.2 RL-Zero-Code 3.1 59.8 81.8 95.2 52.4 62.8 80.4 22.1 27.156.810.2 8.2 27.5 Table 7: Reasoning quality (%). acc: first correct. mv: majority vote. pass: any ofK=16 correct. HumanEvalMBPP @1@5 @10 @16 @1@5 @10 @16 Base1.66.6 10.8 14.0 23.9 45.0 50.9 54.0 Instruct-SFT63.4 88.1 93.6 96.3 32.3 47.5 52.1 54.8 Instruct-DPO73.3 93.6 96.4 97.0 32.9 47.9 51.8 53.6 Instruct (final)81.2 96.2 97.7 98.2 37.8 48.9 51.7 53.2 Think-SFT86.7 94.9 95.6 95.7 41.0 50.1 52.3 53.6 Think-DPO86.5 94.5 95.0 95.1 40.6 49.7 51.9 52.8 Think (final)87.7 95.0 95.6 95.7 44.1 53.7 56.1 58.0 Think-SFT w/o CoT49.4 76.3 81.8 84.1 24.0 43.2 48.6 51.4 Think-DPO w/o CoT 56.5 78.4 82.4 84.8 26.2 42.9 47.2 49.4 Think w/o CoT55.6 77.6 82.0 84.1 23.9 42.1 47.7 50.8 RL-Zero-Math2.4 10.3 17.1 23.2 24.5 45.6 51.7 55.0 RL-Zero-Code2.7 11.2 19.1 26.8 24.8 44.8 50.2 53.2 RL-Zero-IF1.15.19.4 13.4 24.6 45.0 51.8 56.0 RL-Zero-General2.5 11.1 19.7 28.0 24.9 44.9 51.0 54.6 RL-Zero-Math 3.1 2.19.0 15.5 21.3 24.1 44.7 50.2 53.0 RL-Zero-Code 3.1 66.5 83.9 87.9 89.6 25.4 45.4 50.6 53.2 Table 8: Code quality (pass@k, %). 25 Preprint. Under review. strict@1 loose@1 pass@16 consist Base44.758.474.146.5 Instruct-SFT78.785.990.679.3 Instruct-DPO78.785.389.679.3 Instruct (final)82.187.989.381.8 Think-SFT78.084.990.977.2 Think-DPO74.981.486.773.7 Think (final)78.785.291.779.5 Think-SFT w/o CoT70.279.389.571.5 Think-DPO w/o CoT66.775.482.466.5 Think w/o CoT71.079.688.070.7 RL-Zero-Math47.759.872.647.0 RL-Zero-Code47.060.171.546.6 RL-Zero-IF59.770.575.061.0 RL-Zero-General46.659.672.348.0 RL-Zero-Math 3.1 48.660.772.545.9 RL-Zero-Code 3.1 46.458.573.646.7 Table 9: IFEval constraint satisfaction (%). Acc@1MV@16Pass@16 Base16.436.061.0 Instruct-SFT32.443.573.6 Instruct-DPO32.940.084.5 Instruct (final)18.021.076.2 Think-SFT19.228.274.5 Think-DPO17.426.565.1 Think (final)15.829.465.5 Think-SFT w/o CoT26.347.574.2 Think-DPO w/o CoT27.044.270.7 Think w/o CoT27.744.271.4 RL-Zero-Math14.234.959.8 RL-Zero-Code10.027.052.2 RL-Zero-IF23.032.958.8 RL-Zero-General18.837.968.2 RL-Zero-Math 3.1 9.825.850.5 RL-Zero-Code 3.1 10.828.455.0 Table 10: CruxEval output prediction quality (%). Accuracy@1, majority vote@16, and pass@16. 26 Preprint. Under review. TL;DRCNN/DMXSum Base26.047.726.7 Instruct-SFT72.220.052.0 Instruct-DPO70.495.494.4 Instruct (final)77.895.495.4 Think-SFT32.097.295.8 Think-DPO28.491.690.6 Think (final)38.297.896.0 Think-SFT w/o CoT20.055.678.4 Think-DPO w/o CoT13.244.767.8 Think w/o CoT12.549.673.9 RL-Zero-Math35.449.337.4 RL-Zero-Code37.249.041.4 RL-Zero-IF36.841.639.4 RL-Zero-General44.060.643.9 RL-Zero-Math 3.1 34.450.039.8 RL-Zero-Code 3.1 33.056.240.6 Table 11: Summarization quality: pairwise win rate (%) against reference summaries, judged by gpt-4.1-mini. AlpacaPRISMWritingPrompts Base— Instruct-SFT48.784.093.1 Instruct-DPO73.993.196.9 Instruct (final)66.391.097.3 Think-SFT84.392.596.8 Think-DPO95.295.798.0 Think (final)83.993.797.3 Think-SFT w/o CoT83.488.692.5 Think-DPO w/o CoT88.092.095.4 Think w/o CoT84.788.693.3 RL-Zero-Math53.151.749.9 RL-Zero-Code55.461.954.6 RL-Zero-IF32.853.352.4 RL-Zero-General76.763.053.6 RL-Zero-Math 3.1 57.555.749.2 RL-Zero-Code 3.1 52.552.149.4 Table 12: Open-ended quality: pairwise win rate (%) against Base model. Alpaca and PRISM use the MT-Bench pair-v2 prompt (Zheng et al., 2023); WritingPrompts uses the Arena-Hard creative writing prompt (Li et al., 2025b) with position-swap debiasing. Judge: gpt-4.1-mini. 27 Preprint. Under review. RawσMedianWB-Score Base4.02.44-2.0 Instruct-SFT7.21.984.5 Instruct-DPO7.61.785.2 Instruct (final)8.01.596.1 Think-SFT7.22.184.3 Think-DPO7.52.085.1 Think (final)7.32.084.6 Think-SFT w/o CoT5.42.650.8 Think-DPO w/o CoT5.72.361.4 Think w/o CoT5.72.561.4 RL-Zero-Math4.12.54-1.7 RL-Zero-Code4.22.64-1.6 RL-Zero-IF4.02.54-2.0 RL-Zero-General4.92.75-0.2 RL-Zero-Math 3.1 4.02.54-2.0 RL-Zero-Code 3.1 4.22.64-1.6 Table 13: WildBench quality: checklist-guided WB-Score (Lin et al., 2025), judged by gpt-4.1- mini. Raw score (1–10) and normalized WB-Score = (raw− 5)× 2. E Quality-filtered diversity Table 14 reports the quality-filtered diversity decomposition defined in §3.3 for six verifiable tasks. We label each ofK=16 generations as correct or incorrect (answer matching for math, test execution for code, constraint satisfaction for IFEval), then report accuracy alongside D a (SBERT on all outputs),D c (SBERT on correct-only subset,K c ≥2), andV c (Vendi Score on correct outputs, interpreted as the effective number of distinct correct answers). F Code-specific diversity Table 15 reports quality-filtered code diversity using the domain-specific metrics described in §3.3: UniXcoder SBERT (D code c , computed on correct outputs only) and AST subtree Jaccard distance (D AST c , for code-generation tasks). Missing entries (“—”) indicate models with no parseable correct outputs. 28 Preprint. Under review. GSM8KMATH-AlgebraMATH-Geometry accD a D c V c accD a D c V c accD a D c V c Base52 0.172 0.135 1.7 48 0.146 0.119 1.6 23 0.198 0.145 1.6 Instruct-SFT73 0.105 0.098 1.5 56 0.132 0.110 1.6 26 0.179 0.140 1.6 Instruct-DPO77 0.141 0.137 1.8 51 0.071 0.067 1.4 23 0.096 0.082 1.4 Instruct (final)80 0.078 0.074 1.4 71 0.057 0.057 1.4 43 0.101 0.087 1.5 Think-SFT92 0.061 0.060 1.4 76 0.054 0.051 1.3 50 0.107 0.080 1.5 Think-DPO85 0.052 0.049 1.3 75 0.061 0.053 1.3 50 0.114 0.082 1.5 Think (final)93 0.051 0.050 1.3 77 0.062 0.059 1.4 51 0.122 0.091 1.6 Think-SFT w/o CoT77 0.057 0.055 1.3 56 0.066 0.058 1.3 28 0.098 0.072 1.4 Think-DPO w/o CoT 70 0.045 0.042 1.2 47 0.058 0.050 1.3 20 0.077 0.061 1.3 Think w/o CoT75 0.052 0.048 1.3 49 0.064 0.055 1.3 20 0.089 0.064 1.3 RL-Zero-Math61 0.154 0.124 1.7 49 0.144 0.119 1.6 23 0.181 0.135 1.6 RL-Zero-Code58 0.156 0.127 1.7 51 0.144 0.114 1.6 23 0.183 0.135 1.6 RL-Zero-IF50 0.177 0.137 1.7 48 0.143 0.111 1.6 21 0.199 0.132 1.6 RL-Zero-General61 0.133 0.110 1.6 54 0.124 0.104 1.6 24 0.166 0.127 1.5 RL-Zero-Math 3.1 55 0.183 0.136 1.7 54 0.140 0.120 1.6 21 0.183 0.133 1.6 RL-Zero-Code 3.1 60 0.173 0.130 1.7 52 0.139 0.115 1.6 22 0.178 0.133 1.6 IFEvalHumanEvalMBPPCRUXEval accD a D c V c accD a D c V c accD a D c V c accD a D c V c Base45 0.349 0.333 3.2 18 0.411 0.123 1.5 19 0.291 0.196 1.9 20 0.239 0.240 1.9 Instruct-SFT79 0.172 0.171 2.2 63 0.112 0.109 1.6 32 0.111 0.098 1.5 32 0.218 0.177 1.7 Instruct-DPO79 0.154 0.155 2.1 73 0.095 0.095 1.6 33 0.073 0.059 1.3 38 0.068 0.168 1.6 Instruct (final)82 0.154 0.155 2.1 81 0.093 0.091 1.6 38 0.069 0.058 1.3 23 0.062 0.139 1.4 Think-SFT78 0.191 0.180 2.3 87 0.109 0.101 1.6 41 0.081 0.058 1.3 18 0.095 0.076 1.3 Think-DPO75 0.165 0.159 2.1 87 0.081 0.072 1.4 36 0.084 0.067 1.4 17 0.076 0.056 1.2 Think (final)79 0.196 0.187 2.3 88 0.117 0.110 1.6 44 0.089 0.064 1.4 12 0.090 0.074 1.3 Think-SFT w/o CoT70 0.196 0.185 2.2 49 0.055 0.046 1.3 24 0.084 0.072 1.4 20 0.084 0.087 1.4 Think-DPO w/o CoT 67 0.157 0.152 2.0 56 0.062 0.051 1.3 26 0.083 0.065 1.3 21 0.064 0.098 1.4 Think w/o CoT71 0.221 0.182 2.1 56 0.060 0.053 1.3 24 0.083 0.070 1.4 20 0.071 0.081 1.3 RL-Zero-Math48 0.318 0.295 2.930.421 0.089 1.4 24 0.274 0.157 1.8 18 0.222 0.245 1.9 RL-Zero-Code47 0.287 0.278 2.730.464 0.180 1.5 25 0.238 0.147 1.7 16 0.149 0.201 1.7 RL-Zero-IF60 0.397 0.371 3.900.336–24 0.297 0.149 1.7 24 0.491 0.319 2.1 RL-Zero-General47 0.284 0.271 2.7 32 0.468 0.113 1.5 25 0.272 0.151 1.7 23 0.198 0.190 1.8 RL-Zero-Math 3.1 49 0.324 0.300 2.970.460 0.116 1.4 24 0.292 0.157 1.8 19 0.207 0.236 1.8 RL-Zero-Code 3.1 46 0.325 0.293 2.8 66 0.439 0.071 1.4 26 0.261 0.153 1.8 17 0.209 0.247 1.8 Table 14: Quality-filtered diversity. acc: accuracy (%).D a : SBERT on all outputs.D c : SBERT on correct only (K c ≥2).V c : Vendi Score on correct only (effective number of distinct answers). 29 Preprint. Under review. HumanEvalMBPP acc D code c D AST c acc D code c D AST c Base18 0.168 0.59019 0.310 0.927 Instruct-SFT63 0.167 0.59132 0.179 0.683 Instruct-DPO74 0.113 0.67433 0.118 0.777 Instruct (final)81 0.142 0.59338 0.118 0.693 Think-SFT91 0.116 0.52740 0.124 0.533 Think-DPO91 0.130 0.54039 0.162 0.611 Think (final)91 0.126 0.53144 0.130 0.590 Think-SFT w/o CoT51 0.105 0.51026 0.170 0.662 Think-DPO w/o CoT59 0.123 0.48528 0.160 0.707 Think w/o CoT57 0.112 0.49926 0.164 0.676 RL-Zero-Math3 0.058 0.05725 0.253 0.905 RL-Zero-Code3 0.254—25 0.249 0.889 RL-Zero-IF0—25 0.245 0.887 RL-Zero-General33 0.174 0.61825 0.255 0.900 RL-Zero-Math 3.1 7 0.101 0.26124 0.263 0.889 RL-Zero-Code 3.1 67 0.124 0.57625 0.249 0.895 Table 15: Code-specific diversity on correct outputs for code-generation tasks. acc: accuracy (%, meanK c /16).D code c : UniXcoder SBERT (correct only).D AST c : AST subtree Jaccard (correct only). 30 Preprint. Under review. G Output length analysis Table G reports the mean output word length and mean SBERT diversity per task, averaged across all 13 models. Tasks with high mean diversity (e.g. WritingPrompts, HumanEval) span a wide range of output lengths, and tasks with similar lengths (e.g. GSM8K at 137 words, TruthfulQA at 142 words) have very different diversity levels (0.128 vs. 0.262). Output length does not systematically predict diversity. TaskLen SBERTTaskLen SBERT WildBench8720.230TL;DR2830.270 PRISM7230.260MATH-Algebra2270.110 WritingPrompts7040.397MBPP2130.198 CRUXEval6190.183HumanEval2110.280 MATH-Geometry4410.155XSum1580.292 IFEval3910.257TruthfulQA1420.262 Alpaca3040.204GSM8K1370.128 CNN/DailyMail1200.157 Table 16: Mean output word length and SBERT diversity per task, averaged across 13 models. H Temperature sensitivity Table 17 compares Base model diversity at its recommended sampling temperature (T=1.0, top-p=0.7) with the matched temperature used throughout this study (T=0.6, top-p=0.95). SBERT diversity decreases by 11% on average, EAD by 18%, and NLI by only 3%. These reductions are modest relative to the 62% SBERT drop from Base to Think-SFT, confirming that the diversity gaps documented in this paper are not attributable to the temperature difference. I Stage attribution per task Table 18 reports the percentage of Base SBERT diversity lost at each post-training stage for all 15 tasks. Think collapses 45–80% at SFT (most on XSum, least on IFEval), with DPO contributing minimally. Instruct shows the opposite pattern: SFT losses range from 8–73%, but DPO contributes 2–63% additional loss. RL-Zero retains 71–105% of Base diversity across tasks. J Decontamination We measure training–evaluation data overlap usingC 13 13-gram matching (Lambert et al., 2025): for each test instance, we extract all 13-grams (tokenized with spaCy), query an Elas- ticsearch index of the training data for phrase matches, and report the fraction of test tokens covered by at least one matching 13-gram, averaged over all test instances. Table 19 reports results for the four Dolci post-training datasets against all fifteen evaluation benchmarks. Summarization, creative writing, open-ended QA, and value-pluralism benchmarks show negligible overlap (≤1.6%). HumanEval, CRUXEval, IFEval, MATH, and WildBench show elevated overlap (7–30%), traceable to shared upstream sources: the Dolci SFT mixes include OpenThoughts3 (Guha et al., 2025), whose math questions derive from OpenMathInstruct-2 (Toshniwal et al., 2024), itself built on the MATH training set, large-scale Python corpora, Dolci-Think-Python (Olmo et al., 2025), Nemotron (NVIDIA, 2025) code split, and WildChat conversations (Zhao et al., 2024). 31 Preprint. Under review. EADSBERTNLIVendi TaskT=1.0 T=0.6∆% T=1.0 T=0.6∆% T=1.0 T=0.6∆% T=1.0 T=0.6∆% TL;DR0.478 0.365 -23.6 0.385 0.353 -8.1 0.987 0.949 -3.8 4.556 4.157 -8.8 CNN/DM0.439 0.370 -15.8 0.254 0.279 +9.8 0.973 1.036 +6.5 2.928 3.162 +8.0 XSum0.743 0.674 -9.3 0.551 0.451 -18.2 1.122 1.087 -3.1 6.628 5.192 -21.7 HumanEval0.623 0.570 -8.5 0.438 0.411 -6.3 1.037 0.894 -13.7 2.578 2.454 -4.8 MBPP0.631 0.592 -6.2 0.431 0.291 -32.5 1.236 1.179 -4.7 3.479 2.753 -20.9 CRUXEval0.429 0.310 -27.6 0.293 0.239 -18.2 0.992 0.997 +0.6 2.841 2.458 -13.5 GSM8K0.433 0.450 +3.9 0.199 0.172 -13.6 1.095 1.077 -1.7 2.251 2.094 -7.0 MATH-Algebra0.472 0.453 -3.9 0.156 0.146 -6.2 1.012 1.000 -1.2 2.050 1.983 -3.2 MATH-Geometry 0.476 0.402 -15.5 0.210 0.198 -5.9 1.114 1.135 +1.8 2.525 2.438 -3.5 TruthfulQA0.616 0.461 -25.2 0.452 0.353 -21.8 1.097 0.972 -11.4 5.173 3.805 -26.4 Alpaca0.539 0.509 -5.6 0.396 0.319 -19.5 0.791 0.676 -14.5 2.620 2.217 -15.4 IFEval0.554 0.443 -20.1 0.371 0.349 -6.0 1.063 1.055 -0.7 4.134 3.812 -7.8 WritingPrompts0.357 0.234 -34.5 0.588 0.540 -8.1 1.166 1.165 -0.1 7.914 6.935 -12.4 PRISM0.376 0.237 -37.0 0.452 0.408 -9.7 1.101 1.086 -1.4 5.313 4.639 -12.7 WildBench0.475 0.300 -36.9 0.350 0.335 -4.3 1.087 1.064 -2.1 3.702 3.514 -5.1 Mean0.509 0.425 -17.7 0.368 0.323 -11.2 1.058 1.025 -3.3 3.913 3.441 -10.3 Table 17: Base model diversity at recommended (T=1.0) vs. matched (T=0.6) temperature. ∆% reports the relative change. ThinkInstructRL-Zero TaskSFT DPORL Retain SFT DPORL RetainRetain TL;DR−53 −2 +146 −24 −19 +15993 CNN/DM −70 −8 +1133 −20 −53 −12671 XSum−80 −6 +620 −37 −44 −01894 HumanEval −73 −7 +928 −73 −4 −023105 MBPP−72 +1 +131 −62 −13 −22494 CRUXEval −60 −8 +638 −9 −63 −326103 GSM8K−64 −6 −030 −39 +21 −374595 MATH-Alg −63 +5 +143 −10 −42 −93995 MATH-Geo −46 +3 +462 −9 −42 +25192 IFEval−45 −7 +956 −51 −5 +04492 Alpaca−56 −7 +946 −47 −15 −23576 WritingPrompts −56 −6 −137 −49 −9 −437100 TruthfulQA −66 −13 +021 −8 −48 −123399 PRISM−75 −3 +122 −65 −11 −12295 WildBench −52 −2 +652 −61 −2 −13589 Average −62 −4 +438 −38 −23 −53493 Table 18: Per-task stage attribution: percentage of Base SBERT diversity lost (−) or recovered (+) at each post-training stage. Retain is the fraction of Base diversity preserved at the final checkpoint. CNN/DMXSumTL;DRGSM8KMATH-AlgMATH-GeoHumanEvalMBPPCRUXEvalAlpaca IFEvalWrtPromptsTruthfulQAPRISMWildBench Think-SFT0.30.2 0.10.210.215.421.51.19.50.47.60.00.00.0 20.5 Think-DPO0.20.3 0.00.01.21.214.70.05.50.76.80.00.00.06.5 Inst.-SFT0.10.5 0.10.02.73.030.11.3 11.20.47.30.00.00.0 26.4 Inst.-DPO0.10.1 0.00.01.81.820.90.36.81.67.20.00.00.06.9 Table 19: C 13 13-gram overlap (%) between Dolci training sets and evaluation benchmarks. Values≥ 5% are bolded. 32