Paper deep dive
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 8/16/2026, 2:28:05 AM
Summary
This study investigates the impact of instruction tuning on model confidence and lexical diversity in question-answering tasks. Using three matched base and instruction-tuned model pairs (Qwen2.5-7B, Mistral-7B, Llama-3.1-8B) across ARC-Easy, MMLU, and CSQA benchmarks, the authors find that instruction tuning consistently increases model confidence (lower entropy, higher verbalized confidence) without corresponding improvements in predictive accuracy. While cross-rationale diversity (1-SelfBLEU) consistently decreases, surface-level lexical diversity (Unique-2) varies non-uniformly. These diversity changes persist after controlling for answer selection and rationale length, indicating that confidence and diversity capture distinct effects of instruction tuning.
Entities (12)
Relation Signals (16)
Mistral-7B → evaluatedon → CommonsenseQA
confidence 100% · Table 1... Mistral-7B... CommonsenseQA
Mistral-7B → evaluatedon → MMLU
confidence 100% · Table 1... Mistral-7B... MMLU
LLaMA-3.1-8B → evaluatedon → ARC-Easy
confidence 100% · Table 1... Llama-3.1-8B... ARC-Easy
LLaMA-3.1-8B → evaluatedon → CommonsenseQA
confidence 100% · Table 1... Llama-3.1-8B... CommonsenseQA
LLaMA-3.1-8B → evaluatedon → MMLU
confidence 100% · Table 1... Llama-3.1-8B... MMLU
Qwen2.5-7B → evaluatedon → ARC-Easy
confidence 100% · We evaluate three matched base and instruction-tuned models across question-answering benchmarks... Qwen2.5-7B... ARC-Easy
Qwen2.5-7B → evaluatedon → MMLU
confidence 100% · We evaluate three matched base and instruction-tuned models across question-answering benchmarks... Qwen2.5-7B... MMLU
Qwen2.5-7B → evaluatedon → CommonsenseQA
confidence 100% · We evaluate three matched base and instruction-tuned models across question-answering benchmarks... Qwen2.5-7B... CommonsenseQA
Mistral-7B → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.
Tags
Links
- Source: https://arxiv.org/abs/2608.13430v1
- Canonical: https://arxiv.org/abs/2608.13430v1
Trouble viewing inline? Open PDF directly →
Full Text
44,299 characters extracted from source content.
Expand or collapse full text
Are You Sure You’re Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Irina Proskurina Thanks: Equal contribution. Affiliation: Cohere Labs Community Affiliation: Laboratoire Hubert Curien, UMR CNRS 5516, Saint-Étienne, France Mayank Kumar11footnotemark: 1 Affiliation: Cohere Labs Community Affiliation: School of Computer Science Engineering and Technology (SCSET), Bennett University, Greater Noida, India Oyindolapo Komolafe Affiliation: Cohere Labs Community Affiliation: School of Physical Therapy, Faculty of Health Sciences, Western University, London, Canada Abstract Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently alters answer confidence, despite limited changes in predictive accuracy and decreases in likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning. 1 Introduction Large language models (LLMs) are increasingly applied in question answering and reasoning tasks in medicine, finance, and law, where reliable estimates of model confidence are particularly important (26; 11; 32; 7). Prior work has shown that training base models to follow natural-language instructions can improve model performance in such applications (31; 21). At the same time, post-training was shown to alter model confidence distributions (15; 29; 34). Figure 1: Effect of instruction tuning on answer uncertainty and lexical diversity of generated answer rationales. Each point represents a matched base-instruction model pair on a single benchmark, with scores changes computed as Δ=Instruct−Base =Instruct-Base. Model confidence, however, can be estimated using several proxies, including the probability assigned to the selected answer relative to alternative answers, consistency across repeated generations, and explicitly verbalized confidence (15; 29; 34). In free-form generation, uncertainty estimation is further complicated by variation in the surface form of generated responses, since semantically equivalent answers can be expressed using substantially different lexical sequences (16; 6). While several recent studies investigate how instruction tuning and preference-based post-training affect model confidence (36; 13), to the best of our knowledge, prior work has not examined whether changes in confidence are accompanied by corresponding shifts in the lexical diversity of generated answer rationales, which provide supporting evidence for the selected answer. In this work, we investigate how instruction tuning affects likelihood-based answer uncertainty and verbalized confidence, and whether the differences between base and instruction-tuned models are associated with corresponding changes in the lexical diversity of model rationales. Our contributions are as follows: 1) we conduct a paired evaluation of base and instruction-tuned models across three model families and three reasoning benchmarks, comparing confidence, calibration, and rationale diversity; 2) we perform controlled comparisons restricted to examples for which base and instruction-tuned models select the same answer, while matching rationale length across variants, allowing us to measure changes in lexical diversity independently of answer switching and generation length; and 3) we show that instruction tuning increases model confidence without corresponding improvements in predictive accuracy, while changes in lexical diversity remain heterogeneous and are not consistently associated with uncertainty or calibration. Together, our findings show that increased confidence after instruction tuning is not associated with a uniform shift in either rationale diversity or model calibration. 2 Related Work Calibration measures how well a model’s predicted confidence aligns with the empirical probability of being correct (9; 14; 8). Several empirical studies have shown that model confidence can be misaligned with predictive accuracy, with over- and under-confidence observed for likelihood-based estimates (15), verbalized confidence (29), and generation-based uncertainty measures (34). Recent work has shown that post-training can further alter model confidence distributions, with alignment methods such as learning from human feedback affecting verbalized confidence even when downstream task performance improves (37; 29; 36). Another line of work studies the impact of instruction tuning and other post-training on linguistic diversity in open-ended generation (10; 35; 5). Further works examine lexical, syntactic, and semantic diversity as indicators of creativity in model outputs, comparing the generated outputs with human-written text or human judgments of creativity (30; 2; 22; 23). To the best of our knowledge, despite a substantial body of work on output diversity, model confidence, and calibration, the relationship between lexical or token-level diversity and verbalized confidence or calibration has not yet been investigated. Model ARC-Easy MMLU CSQA Acc.↑ Hchoice↓H_choice Verb.↑ U2↑ 11-SB↑ Acc.↑ Hchoice↓H_choice Verb.↑ U2↑ 11-SB↑ Acc.↑ Hchoice↓H_choice Verb.↑ U2↑ 11-SB↑ Qwen2.5-7B Base 80.6 .235 49.6 .694 .695 71.8 .430 59.8 .679 .671 85.2 .268 27.4 .751 .734 Instr. 81.6 .127∗ 60.4∗ .669∗ .598∗ 71.7 .131∗ 68.7∗ .687 .628∗ 82.5∗ .076∗ 39.0∗ .717∗ .654∗ Mistral-7B Base 80.1 .215 76.6 .701 .813 59.6 .680 84.9 .731 .803 57.4 .736 47.9 .719 .833 Instr. 83.4∗ .117∗ 93.8∗ .673∗ .626∗ 59.7 .328∗ 91.5∗ .674∗ .626∗ 69.2∗ .268∗ 92.2∗ .750∗ .745∗ Llama-3.1-8B Base 82.2 .188 49.2 .704 .783 64.1 .608 48.2 .721 .795 70.8 .362 46.8 .709 .814 Instr. 82.2 .173∗ 90.4∗ .697∗ .720∗ 68.4∗ .457∗ 91.8∗ .702∗ .719∗ 75.9∗ .216∗ 91.1∗ .731∗ .773∗ Table 1: Accuracy (Acc.), choice entropy HchoiceH_choice, verbalized confidence (Verb.), Unique-2 (U2), and 11-SelfBLEUSelfBLEU (11-SB) for paired base and instruction-tuned models across benchmarks. Acc. and Verb. are reported in %. The colors indicate mean-centered values within each metric column (cyan: above the mean; orange: below the mean). ∗ denotes a significant Base-Instruct difference at p<0.01p<0.01 (two-sided paired t-test). 3 Methodology We consider a multiple-choice question answering problem, where each input question x is associated with a set of answers Y=y1,…,yMY=y_1,…,y_M. Model Confidence Evaluation Following 14, we define the model prediction as the candidate answer with the highest conditional likelihood: y^=argmaxy∈YpLM(y∣x). y= y∈ Y \;p_LM(y x). (1) We evaluate model confidence using several complementary measures. First, we estimate model uncertainty over the candidate answers using the entropy of normalized answer likelihoods (18; 25): Hchoice(x)=−∑j=1MpjlogpjlogM, H_choice(x)=- _j=1^Mp_j p_j M, (2) where pjp_j denotes the normalized probability assigned to candidate answer yjy_j and higher values of Hchoice(x)H_choice(x) indicate greater uncertainty over the candidate answers. Secondly, we also use elicited verbalized confidence as an estimate of model confidence, following 34. Specifically, we use a two-stage protocol, where in the first forward pass the model prediction is obtained using the model likelihood in Eq. (1), and then, in the second forward pass, the model is prompted to report a numerical probability that the selected answer is correct. Additional implementation details are provided in Appendix B. Lexical Diversity Next, we measure lexical diversity in the rationales generated for each question and selected answer. For each question, we sample K=5K=5 rationales using chain-of-thought prompting. The full prompt and generation details are provided in Appendix B. We evaluate the generated rationales using the Unique Tokens Ratio and Self-BLEU measures, following prior work on diversity evaluation in text generation (1). In particular, we use the proportion of distinct bigrams as the Unique-2 score, with larger values indicating greater lexical richness. We also compute Self-BLEU as the similarity of each rationale to the remaining rationales generated for the same question, with larger values indicating greater similarity across generations. Experimental Settings We use three widely used multiple-choice benchmarks: ARC-Easy 4, MMLU 12, and CommonsenseQA (CSQA) 28. These benchmarks cover grade-school science, a broad range of academic and professional subjects, and commonsense knowledge grounded in ConceptNet (27), respectively. For our analysis, we use three pairs of base and instruction-tuned models: Qwen2.5-7B, Llama-3.1-8B, and Mistral-7B-v0.3. Model links and license information are provided in Appendix B. 4 Results We begin by analyzing how instruction tuning affects the evaluated confidence and lexical diversity measures. The results across models and benchmarks are reported in Table 1.11 1 We report 11-SelfBLEUSelfBLEU so that larger values correspond to greater cross-rationale variability. We summarize the main findings below. Instruction Tuning Consistently Increases Model Confidence without Corresponding Improvements in Accuracy. We find that across all models and benchmarks, instruction tuning consistently increases model confidence, as reflected by lower answer entropy and higher verbalized confidence (Table 1). Answer entropy decreases across all benchmarks, with particularly large decreases for Qwen on MMLU (0.4300.430 to 0.1310.131) and Mistral on CSQA (0.7360.736 to 0.2680.268). Similarly, verbalized confidence increases across all settings, including from 49.2%49.2\% to 90.4%90.4\% for Llama on ARC-Easy and from 46.8%46.8\% to 91.1%91.1\% on CSQA. In contrast, accuracy changes are less consistent across benchmarks. For example, on ARC-Easy, Llama accuracy remains unchanged at 82.2%82.2\%, despite a significant increase in verbalized confidence and a decrease in choice entropy. Instruction Tuning Induces Heterogeneous Changes in Rationale Lexical Diversity. In contrast to model confidence, the effect of instruction tuning on lexical diversity is less consistent across models and benchmarks. The largest decreases in Unique-2 are observed for Mistral on ARC-Easy and MMLU, while the largest increase occurs on CSQA, from 0.7190.719 to 0.7500.750. At the same time, cross-rationale diversity, measured as 11-SelfBLEUSelfBLEU, decreases across all models and benchmarks, with the largest decrease observed for Mistral on ARC-Easy, from 0.8130.813 to 0.6260.626. Overall, instruction tuning consistently reduces cross-rationale variability, whereas changes in Unique-2 vary across models and benchmarks. Decreases in Answer Uncertainty Do Not Consistently Coincide with Reduced Lexical Diversity. We further examine whether lower or higher answer uncertainty coincides with lower or higher lexical diversity. Table 2 reports the four possible directions of Instruct-Base uncertainty-diversity changes. For Qwen model, decreases in uncertainty most often coincide with decreases in diversity, accounting for 61.8%61.8\% of examples under Unique-2 and 69.3%69.3\% under 1−SelfBLEU1-SelfBLEU. For Mistral, the pattern differs across diversity measures: lower uncertainty coincides with higher Unique-2 for 61.8%61.8\% of examples, but with lower 1−SelfBLEU1-SelfBLEU for 77.6%77.6\%. Llama exhibits the same divergence between the two measures. Overall, decreases in answer uncertainty are not consistently accompanied by decreases in lexical diversity, and the direction of the association depends on the diversity measure and model. Figure 1 illustrates the same pattern, where choice entropy generally decreases after instruction tuning while lexical diversity changes in both directions across models and benchmarks. Model Div. ↓↓ ↓↑ ↑↓ ↑↑ Qwen2.5 U2 61.8 32.8 3.8 1.6 11-SB 69.3 25.3 3.9 1.5 Mistral U2 36.0 61.8 0.5 1.7 11-SB 77.6 20.2 1.3 0.9 Llama-3.1 U2 33.3 47.5 7.2 12.0 11-SB 54.7 26.1 12.4 6.8 Table 2: Directional changes in choice entropy (H) and lexical diversity (D) for paired base and instruction-tuned models on CommonsenseQA. The first and second arrows denote the Instruct-Base change in H and D, respectively; e.g., ↓↑ indicates lower uncertainty and higher diversity. Values are reported as percentages of benchmark questions falling into each quadrant. The most frequent directional pattern for each model and diversity measure is in bold. Diversity Shifts Persist after Controlling for Answer Selection and Rationale Length. Because instruction tuning can affect both the selected answer and rationale length, we test whether the observed diversity changes persist when 1) matched base and instruction-tuned models select the same answer and 2) rationale length is matched across model variants.22 2 For length matching, rationales are paired by length and each pair is truncated to the length of the shorter rationale; see Appendix B. The evaluation results, together with the number of examples out of the 1,200 CSQA questions for which both conditions hold, are reported in Table 3. We find that for Qwen, Unique-2 remains nearly unchanged (−0.001-0.001), while 1−SelfBLEU1-SelfBLEU decreases by 0.0360.036. For Mistral and Llama, Unique-2 increases significantly by 0.0500.050 and 0.0530.053, respectively, whereas 1−SelfBLEU1-SelfBLEU decreases by 0.0690.069 and 0.0120.012. Thus, the observed lexical diversity changes persist after controlling for answer selection and rationale length, while the two diversity measures continue to show different trends. Changes in Lexical Diversity Are Not Consistently Associated with Calibration. We additionally compare changes in lexical diversity with model calibration, estimated using Expected Calibration Error (ECE). The full calibration results are reported in Table 7 (see Appendix C). We find that changes in lexical diversity are not consistently associated with changes in calibration. For instance, for Qwen model, verbalized-confidence on ARC-Easy decreases from 35.335.3 to 22.822.8 after instruction tuning, while both Unique-2 and 1−SelfBLEU1-SelfBLEU decrease. In contrast, for Llama model on MMLU benchmark, both likelihood-based and verbalized ECE increase (0.50.5 to 5.95.9 and 16.616.6 to 23.723.7, respectively), while both diversity measures decrease. Model N Δ 2 Δ(1−SB) (1-SB) Qwen2.5-7B 1104 −.001-.001 −.036∗-.036^* Mistral-7B 904 +.050∗+.050^* −.069∗-.069^* Llama-3.1-8B 1065 +.053∗+.053^* −.012∗-.012^* Table 3: Unique-2 (U2) and 1−SelfBLEU1-SelfBLEU (11-SB) changes for paired base and instruction-tuned models on CommonsenseQA, restricted to questions for which both variants select the same answer and have matched rationale lengths. ∗ denotes a statistically significant Base-Instruct difference at p<0.01p<0.01 (two-sided paired t-test). 5 Conclusion In this paper, we investigate the impact of instruction tuning on model confidence and the lexical diversity of generated rationales in question answering tasks. Through our analysis of rationale diversity, we find that cross-rationale variability decreases after instruction tuning, whereas surface-level lexical diversity exhibits benchmark-dependent patterns. We further show that decreases in answer uncertainty are not consistently accompanied by decreases in lexical diversity, and that diversity changes do not consistently reflect changes in verbalized confidence, answer selection, or model calibration. Moreover, the same patterns persist after controlling for answer selection and rationale length. Overall, our findings show that instruction tuning affects confidence and rationale diversity differently, motivating future work on uncertainty estimation that jointly considers predictive confidence and variation in generated rationales. Future work could further investigate whether reducing post-training overconfidence comes at the cost of further reducing rationale diversity, and whether the observed uncertainty-diversity patterns extend beyond lexical variation to semantic diversity. Limitations To the best of our knowledge, this work represents one of the first attempts to jointly study rationale diversity and model verbalized and likelihood-based confidence. Our study is currently limited to three English multiple-choice benchmarks and lexical diversity measures. Future work may therefore extend the analysis to semantic and syntactic diversity (10) of generated rationales across a broader range of benchmarks. Second, future work should examine whether the same uncertainty-diversity patterns generalize across languages. More broadly, future work could examine how post-training calibration (33) or decoding interventions that modify generation probabilities (19; 20) affect cross-rationale diversity in generation tasks. Ethical Considerations We experiment with publicly available models and benchmark datasets and adhere to the intended use and licensing terms of the respective resources. Our results nevertheless highlight potential risks associated with interpreting the confidence of instruction-tuned language models. Across the evaluated models and benchmarks, we find that instruction tuning increases model confidence and reduces cross-rationale diversity without a corresponding improvement in predictive accuracy. In downstream applications, particularly in high-stakes settings, this mismatch may encourage unwarranted reliance on incorrect predictions when model confidence is interpreted as evidence that the prediction is reliable. The differing calibration results obtained for likelihood-based and verbalized confidence further suggest that confidence should be assessed using multiple complementary measures rather than a single proxy, reducing the risk of inappropriate reliance on miscalibrated predictions. Similarly, reduced cross-rationale diversity may limit the extent to which repeated generations provide independent evidence about a model’s prediction. In generative settings, the implications of such shifts in rationale diversity should therefore be evaluated further, including comparisons against human-generated rationales and human judgments. Finally, we do not evaluate instruction-tuned model performance on questions involving unsafe data or demographic-sensitive attributes. Consequently, our findings should not be interpreted as evidence that instruction tuning increases harmful or biased outputs for such questions. Whether the observed confidence and diversity patterns extend to safety-sensitive or demographic-sensitive prompts requires separate evaluation. Acknowledgments We thank Yanzhu Guo for insightful discussions and constructive feedback during the early stages of the project, Alvin Vinod Chandran, Abdelrahman Alkahwaji, and other members of the Cohere Labs Community for helpful exchanges and early feedback on the project statement. We also thank Cohere Labs for providing collaborative opportunities throughout the development of this work. This work was performed using HPC resources from GENCI-IDRIS (Grant 2025-AD011014384R1). References Alihosseini et al. (2019) D. Alihosseini, E. Montahaei, and M. Soleymani Baghshah Jointly measuring diversity and quality in text generation models. In Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, A. Bosselut, A. Celikyilmaz, M. Ghazvininejad, S. Iyer, U. Khandelwal, H. Rashkin, and T. Wolf (Eds.), Minneapolis, Minnesota, p. 90–98. External Links: Link, Document Cited by: §3. Bae and Kim (2024) M. Bae and H. Kim Collective critics for creative story generation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 18784–18819. External Links: Link, Document Cited by: §2. Biderman et al. (2024) S. Biderman, H. Schoelkopf, L. Sutawika, L. Gao, J. Tow, B. Abbasi, A. F. Aji, P. S. Ammanamanchi, S. Black, J. Clive, et al. Lessons from the trenches on reproducible evaluation of language models. arXiv preprint arXiv:2405.14782. Cited by: Appendix B. Clark et al. (2018) P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457. Cited by: §3. Deshpande et al. (2025) V. Deshpande, D. Ghose, J. D. Patterson, R. E. Beaty, and A. Rumshisky Diverse, not short: a length-controlled data selection strategy for improving response diversity of language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 33917–33938. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2. Farquhar et al. (2024) S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal Detecting hallucinations in large language models using semantic entropy. Nature 630 (8017), p. 625–630. Cited by: §1. Fei et al. (2024) Z. Fei, X. Shen, D. Zhu, F. Zhou, Z. Han, A. Huang, S. Zhang, K. Chen, Z. Yin, Z. Shen, et al. Lawbench: benchmarking legal knowledge of large language models. In Proceedings of the 2024 conference on empirical methods in natural language processing, p. 7933–7962. Cited by: §1. Geng et al. (2024) J. Geng, F. Cai, Y. Wang, H. Koeppl, P. Nakov, and I. Gurevych A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, p. 6577–6595. External Links: Link, Document Cited by: §2. Guo et al. (2017) C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger On calibration of modern neural networks. In International conference on machine learning, p. 1321–1330. Cited by: Appendix C, §2. Guo et al. (2025) Y. Guo, G. Shang, and C. Clavel Benchmarking linguistic diversity of large language models. Transactions of the Association for Computational Linguistics 13, p. 1507–1526. Cited by: §2, Limitations. Hager et al. (2024) P. Hager, F. Jungmann, R. Holland, K. Bhagat, I. Hubrecht, M. Knauer, J. Vielhauer, M. Makowski, R. Braren, G. Kaissis, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature medicine 30 (9), p. 2613–2622. Cited by: §1. Hendrycks et al. (2020) D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300. Cited by: §3. Huang et al. (2026) J. Huang, P. Lu, Q. Zeng, Y. Iwasawa, Y. Matsuo, S. Chandar, E. Marrese-Taylor, and I. Li Investigating the multilingual calibration effects of language model instruction tuning. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 1–59. External Links: Link, Document, ISBN 979-8-89176-381-4 Cited by: §1. Jiang et al. (2021) Z. Jiang, J. Araki, H. Ding, and G. Neubig How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9, p. 962–977. External Links: Link, Document Cited by: Appendix C, §2, §3. Kadavath et al. (2022) S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al. Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §1, §1, §2. Kapoor et al. (2024) S. Kapoor, N. Gruver, M. Roberts, A. Pal, S. Dooley, M. Goldblum, and A. Wilson Calibration-tuning: teaching large language models to know what they don’t know. In Proceedings of the 1st Workshop on Uncertainty-Aware NLP (UncertaiNLP 2024), R. Vázquez, H. Celikkanat, D. Ulmer, J. Tiedemann, S. Swayamdipta, W. Aziz, B. Plank, J. Baan, and M. de Marneffe (Eds.), St Julians, Malta, p. 1–14. External Links: Link, Document Cited by: §1. Kojima et al. (2022) T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa Large language models are zero-shot reasoners. Advances in neural information processing systems 35, p. 22199–22213. Cited by: Appendix B. Malinin and Gales (2020) A. Malinin and M. Gales Uncertainty estimation in autoregressive structured prediction. arXiv preprint arXiv:2002.07650. Cited by: §3. Meister et al. (2023) C. Meister, T. Pimentel, L. Malagutti, E. Wilcox, and R. Cotterell On the efficacy of sampling adapters. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 1437–1455. External Links: Link, Document Cited by: Limitations. Nadeem et al. (2020) M. Nadeem, T. He, K. Cho, and J. Glass A systematic characterization of sampling algorithms for open-ended language generation. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, K. Wong, K. Knight, and H. Wu (Eds.), Suzhou, China, p. 334–346. External Links: Link, Document Cited by: Limitations. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35, p. 27730–27744. Cited by: §1. Park et al. (2025a) K. Park, M. Kim, and K. Jung A character-centric creative story generation via imagination. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 1598–1645. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §2. Park et al. (2025b) K. Park, N. Yang, and K. Jung Avoidance decoding for diverse multi-branch story generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 7489–7505. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2. Post (2018) M. Post A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers, O. Bojar, R. Chatterjee, C. Federmann, M. Fishel, Y. Graham, B. Haddow, M. Huck, A. J. Yepes, P. Koehn, C. Monz, M. Negri, A. Névéol, M. Neves, M. Post, L. Specia, M. Turchi, and K. Verspoor (Eds.), Brussels, Belgium, p. 186–191. External Links: Link, Document Cited by: Appendix B. Shannon (1948) C. E. Shannon A mathematical theory of communication. The Bell System Technical Journal 27 (3), p. 379–423. External Links: Document Cited by: §3. Singhal et al. (2023) K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al. Large language models encode clinical knowledge. Nature 620 (7972), p. 172–180. Cited by: §1. Speer et al. (2017) R. Speer, J. Chin, and C. Havasi Conceptnet 5.5: an open multilingual graph of general knowledge. In Proceedings of the AAAI conference on artificial intelligence, Vol. 31. Cited by: §3. Talmor et al. (2019) A. Talmor, J. Herzig, N. Lourie, and J. Berant CommonsenseQA: a question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. Solorio (Eds.), Minneapolis, Minnesota, p. 4149–4158. External Links: Link, Document Cited by: §3. Tian et al. (2023) K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. D. Manning Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 5433–5442. Cited by: §1, §1, §2. Tian et al. (2024) Y. Tian, T. Huang, M. Liu, D. Jiang, A. Spangher, M. Chen, J. May, and N. Peng Are large language models capable of generating human-level narratives?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 17659–17681. External Links: Link, Document Cited by: §2. Wei et al. (2021) J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652. Cited by: §1. Wu et al. (2023) S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann Bloomberggpt: a large language model for finance. arXiv preprint arXiv:2303.17564. Cited by: §1. Xie et al. (2024) J. Xie, A. S. Chen, Y. Lee, E. Mitchell, and C. Finn Calibrating language models with adaptive temperature scaling. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 18128–18138. Cited by: Limitations. Xiong et al. (2024) M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. In International Conference on Learning Representations, Vol. 2024, p. 23650–23678. Cited by: Appendix B, §1, §1, §2, §3. Yun et al. (2025) L. Yun, C. An, Z. Wang, L. Peng, and J. Shang The price of format: diversity collapse in LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 15454–15468. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §2. Zhang et al. (2024) M. Zhang, M. Huang, R. Shi, L. Guo, C. Peng, P. Yan, Y. Zhou, and X. Qiu Calibrating the confidence of large language models by eliciting fidelity. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 2959–2979. External Links: Link, Document Cited by: §1, §2. Zhu et al. (2023) C. Zhu, B. Xu, Q. Wang, Y. Zhang, and Z. Mao On the calibration of large language models and alignment. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 9778–9795. External Links: Link, Document Cited by: §2. Appendix A Examples of Generated Rationales Prompt. When a bottle of strong perfume is opened at the front of a classroom, the smell gradually spreads throughout the entire room. Which process explains this occurrence? A. osmosis B. diffusion C. cell activity D. cell transport Answer: Let’s think step by step. Base R1. The perfume is most concentrated near the opened bottle and gradually spreads to other parts of the room. Diffusion is the movement of molecules from an area of higher concentration to an area of lower concentration. The perfume molecules therefore move away from the front of the classroom and mix with the surrounding air until they are more evenly distributed. Osmosis concerns water moving across a membrane, while cell activity and cell transport involve cellular processes. Therefore, the correct answer is B. diffusion. R2. First, the smell starts near the bottle, where the concentration of perfume molecules is highest. Over time, those molecules spread through the air toward regions where their concentration is lower. This is the definition of diffusion. Osmosis requires the movement of water across a selectively permeable membrane, which is not involved here. Cell activity and cell transport also do not describe movement through open air. Therefore, the process responsible for the spreading perfume smell is B. diffusion. R3. The perfume is released at the front of the classroom and then spreads through the surrounding air. The air initially contains many perfume molecules near the bottle and fewer molecules farther away. Random molecular motion produces a net movement from the region of higher concentration to regions of lower concentration. That process is diffusion. No cell or membrane is required, so osmosis, cell activity, and cell transport are not appropriate explanations. The correct answer is therefore B. diffusion. R4. I can solve this by eliminating the alternatives. Cell activity and cell transport describe processes involving cells, whereas the perfume is spreading through the air. Osmosis specifically refers to the movement of water through a selectively permeable membrane, so it also does not fit. Diffusion describes particles spreading from a region of higher concentration to a region of lower concentration. Since the perfume molecules spread from the bottle throughout the classroom, the process is diffusion. The answer is B. R5. When the bottle is opened, perfume molecules enter the air at a high concentration near the front of the room. Because molecules move randomly, they gradually spread into regions where fewer perfume molecules are present. This continues until the molecules become more evenly distributed throughout the room. Such movement from higher to lower concentration is diffusion. Osmosis is limited to solvent movement across a membrane, and the cellular options are irrelevant here. Thus, B. diffusion is the best answer. Instruct R1. The perfume molecules are initially concentrated near the opened bottle. After release, they spread through the classroom from this region of high concentration toward regions of lower concentration. This movement continues as the molecules mix with the surrounding air and become more evenly distributed. The process is diffusion. Osmosis involves water moving across a selectively permeable membrane, while cell activity and cell transport concern processes within cells. Therefore, the correct answer is B. diffusion. R2. Opening the perfume bottle releases perfume molecules into the surrounding air. These molecules move randomly and spread from the area near the bottle, where their concentration is high, toward areas of the classroom where their concentration is lower. This process is known as diffusion. Osmosis concerns the movement of water across a semi-permeable membrane, and the two cellular options do not apply to particles spreading through air. Therefore, the correct answer is B. diffusion. R3. The smell spreads because perfume molecules move through the air from the point where they are released to the rest of the classroom. Near the bottle, the concentration of perfume molecules is initially higher than it is farther away. Their net movement toward lower-concentration regions is called diffusion. Osmosis requires a membrane and specifically concerns solvent movement, while cell activity and cell transport involve cellular mechanisms. Hence, the process described in the question is B. diffusion. R4. The key observation is that the perfume spreads from a concentrated source into the larger volume of air in the classroom. Particles naturally move and mix, producing a net movement from regions of higher concentration to regions of lower concentration. This is diffusion. The process does not require a cell or a membrane, so cell activity, cell transport, and osmosis do not explain the observation. Therefore, the smell spreads throughout the room by diffusion, corresponding to answer B. R5. Perfume molecules are released from the bottle into the air, creating a high concentration near the source. As the molecules move, they spread toward parts of the room where the concentration is lower and gradually become distributed throughout the available space. This is the characteristic process of diffusion. Osmosis is a specific membrane-related process involving solvent molecules, and cell activity or cell transport is not relevant in this setting. Thus, the correct answer is B. diffusion. Table 4: Example rationales generated by the Base and Instruct variants of Mistral-7B for the same ARC-Easy question. Appendix B Implementation Details Rationale Generation To estimate variability under repeated generation, we sample K=5K=5 rationales for each question across the considered benchmarks. Following the Zero-shot Chain-of-Thought prompting approach introduced by 17, we use the prompt ‘Answer: Let’s think step by step.’ after each question and its candidate answers. Rationales are sampled with temperature T=0.7T=0.7, nucleus-sampling p=1.0p=1.0, and a maximum of 100 newly generated tokens. We generate five rationales per question. The generation settings are fixed across models and benchmarks. We provide a few examples of generated rationales in Table 4. Controlled Lexical Analysis A direct comparison of rationale diversity between Base and Instruct models can be affected by differences in their selected answers or generated rationale lengths. We therefore conduct an additional controlled analysis discussed in §4 (Table 3) restricted to examples for which the Base and Instruct variants select the same answer. Within each example, we retain rationales supporting this common selected answer, match the number of rationales between the two variants, and match rationale length by pairing rationales according to token length and truncating each pair to the length of the shorter rationale before computing lexical-diversity measures. This analysis isolates differences in the linguistic diversity of a fixed answer from differences caused by answer selection or unequal generation length. Implementation We use the LM Evaluation Harness (3) for benchmark accuracy evaluation. Self-BLEU is computed with SacreBLEU (24). We use two-sided paired t-tests on the per-example Instruct-Base differences. We denote differences with ∗ when p<0.01p<0.01. Statistical tests for the controlled lexical analysis are restricted to examples satisfying the same-answer, and matched-rationale-count, and rationale-length conditions described above. Verbalized Confidence Evaluation We obtain verbalized confidence using a separate two-stage prompting procedure, following 34. The model’s answer is first fixed to the answer selected by likelihood-based multiple-choice evaluation. We then prompt the model to estimate the probability that this selected answer is correct. Conditioning on the previously selected answer prevents the verbal-confidence stage from introducing a second, potentially different prediction. The model output is constrained to a numerical probability. We provide below the verbalized-confidence prompt. Question: [question] A. [choice A] B. [choice B] … Selected answer: [selected answer] What is the probability that the selected answer is correct? Give only a number between 0 and 1. Model Links. Model references and license information are provided in Table 5. Model Link License Qwen2.5-7B https://hf.co/Qwen/Qwen2.5-7B Apache 2.0 https://hf.co/Qwen/Qwen2.5-7B-Instruct Apache 2.0 Llama-3.1-8B https://hf.co/meta-llama/Llama-3.1-8B Llama 3.1 Community License https://hf.co/meta-llama/Llama-3.1-8B-Instruct Llama 3.1 Community License Mistral-7B-v0.3 https://hf.co/mistralai/Mistral-7B-v0.3 Apache 2.0 https://hf.co/mistralai/Mistral-7B-Instruct-v0.3 Apache 2.0 Table 5: Models used in the experiments with the associated licenses. Appendix C Extended Results Uncertainty-Diversity Directional Analysis Table 6 extends the question-level directional analysis reported in Table 2 to ARC-Easy and MMLU. For each question, we report the direction of the Instruct-Base change in choice entropy (H) together with the corresponding change in lexical diversity (D). The results further show that decreases in answer uncertainty can coincide with either increases or decreases in lexical diversity, depending on the model and diversity measure. Rationale Length Evaluation Figure 2 reports changes in mean rationale length between Base and Instruct models across benchmarks. Rationale length increases across all models and benchmarks, motivating the length-controlled analysis in Table 3. Calibration Evaluation Following 9; 14, we compute Expected Calibration Error (ECE) as follows: ECE=∑b=1B|Sb|N|acc(Sb)−conf(Sb)|,ECE= _b=1^B |S_b|N |acc(S_b)-conf(S_b) |, (3) where the N predictions are partitioned into B confidence bins SbS_b, |Sb||S_b| denotes the number of predictions in bin b, acc(Sb)acc(S_b) is the proportion of correct predictions in that bin, and conf(Sb)conf(S_b) is their mean predicted confidence. Lower ECE indicates better calibration. We compute ECE separately for likelihood-based and verbalized confidence. Table 7 reports the corresponding calibration results for all Base and Instruct models across benchmarks. Dataset Model Diversity H↓D↓H D H↓D↑H D H↑D↓H D H↑D↑H D ARC-Easy Qwen2.5-7B U2 52.4 34.7 7.7 5.2 11-SB 65.7 21.4 9.3 3.6 Mistral-7B U2 64.3 25.5 7.6 2.6 11-SB 82.6 7.2 9.5 0.7 Llama-3.1-8B U2 32.2 29.8 20.4 17.6 11-SB 45.5 16.5 26.2 11.8 MMLU Qwen2.5-7B U2 48.1 50.1 0.6 1.2 11-SB 62.2 36.0 1.0 0.8 Mistral-7B U2 68.7 28.9 1.7 0.7 11-SB 86.6 11.0 2.1 0.3 Llama-3.1-8B U2 50.4 33.8 8.8 7.0 11-SB 62.2 22.0 12.1 3.7 Table 6: Directional changes in choice entropy (H) and lexical diversity (D) for paired base and instruction-tuned models on Arc-Easy and MMLU benchmarks. The first and second arrows denote the Instruct-Base change in H and D, respectively; e.g., ↓↑ indicates lower uncertainty and higher diversity. Values are reported as percentages of benchmark questions falling into each quadrant. The most frequent directional pattern for each model and diversity measure is in bold. Model ARC-Easy MMLU CommonsenseQA Lik. ECE ↓ Verb. ECE ↓ Gap Lik. ECE ↓ Verb. ECE ↓ Gap Lik. ECE ↓ Verb. ECE ↓ Gap Qwen2.5-7B Base 6.9 35.3 +28.4 4.4 40.8 +36.4 1.6 59.1 +57.5 Instruct 11.3 22.8 +11.5 21.3 29.8 +8.5 12.7 47.0 +34.3 Mistral-7B Base 8.4 10.8 +2.4 1.2 36.6 +35.4 5.1 24.9 +19.9 Instruct 10.2 11.9 +1.7 22.3 35.4 +13.1 15.2 27.6 +12.4 Llama-3.1-8B Base 8.0 33.0 +24.9 0.5 16.6 +16.1 7.7 24.0 +16.3 Instruct 8.5 8.5 0.0 5.9 23.7 +17.8 11.5 15.9 +4.5 Table 7: Calibration error results for Base and Instruct variants of Qwen, Mistral, and Llama. ECE values are reported in %. Gap denotes ECEverb−ECElikECE_verb-ECE_lik; negative values indicate better calibration of verbalized confidence than likelihood-based confidence. Figure 2: Average change in rationale length from Base to Instruct models across benchmarks. Positive values indicate longer rationales after instruction tuning, whereas negative values indicate shorter rationales.