Paper deep dive
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation
Haotian Zong, Binze Li, Yufei Long, Sinyin Chang, Jialong Wu, Gillian K. Hadfield
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 2:30:01 AM
Summary
I-CALM is a prompt-based framework designed to mitigate LLM hallucinations by incentivizing epistemic abstention. It combines three components: eliciting self-reported verbal confidence, implementing explicit reward schemes for abstention, and applying normative principles (truthfulness, humility, responsibility). Experiments on PopQA using models like GPT-5 mini demonstrate that this approach effectively reduces false-answer rates by shifting error-prone cases to abstention without requiring model retraining.
Entities (5)
Relation Signals (3)
GPT-5 Mini → evaluatedon → PopQA
confidence 100% · Using GPT-5 mini on PopQA as the main setting
I-CALM → utilizes → Verbal Confidence
confidence 98% · I-CALM, a prompt-based framework that (i) elicits verbal confidence
I-CALM → mitigates → Hallucination
confidence 95% · I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty. We study whether prompt-only interventions -- explicitly announcing reward schemes for answer-versus-abstain decisions plus humility-oriented normative principles -- can reduce hallucination risk without modifying the model. Our focus is epistemic abstention on factual questions with a verifiable answer, where current LLMs often fail to abstain despite being uncertain about their answers. We first assess self-reported verbal confidence as a usable uncertainty signal, showing stability under prompt paraphrasing and reasonable calibration against a token-probability baseline. We then study I-CALM, a prompt-based framework that (i) elicits verbal confidence, (ii) partially rewards abstention through explicit reward schemes, and (iii) adds lightweight normative principles emphasizing truthfulness, humility, and responsibility. Using GPT-5 mini on PopQA as the main setting, we find that confidence-eliciting, abstention-rewarding prompts, especially with norms, reduce the false-answer rate on answered cases mainly by identifying and shifting error-prone cases to abstention and re-calibrating their confidence. This trades coverage for reliability while leaving forced-answer performance largely unchanged. Varying the abstention reward yields a clear abstention-hallucination frontier. Overall, results show the framework can improve selective answering on factual questions without retraining, with the magnitude of effect varying across models and datasets. Code is available at the following this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.03904v1
- Canonical: https://arxiv.org/abs/2604.03904v1
Trouble viewing inline? Open PDF directly →
Full Text
154,618 characters extracted from source content.
Expand or collapse full text
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Hallucination Mitigation Haotian Zong∗†^*\, Binze Li∗‡^*\, Yufei Long‡ Sinyin Chang‡ Jialong Wu‡ Gillian K. Hadfield§ Abstract Large language models (LLMs) frequently produce confident but incorrect answers, partly because common binary scoring conventions reward answering over honestly expressing uncertainty. We study whether prompt-only interventions—explicitly announcing reward schemes for answer-versus-abstain decisions plus humility-oriented normative principles—can reduce hallucination risk without modifying the model. Our focus is epistemic abstention on factual questions with a verifiable answer, where current LLMs often fail to abstain despite being uncertain about their answers. We first assess self-reported verbal confidence as a usable uncertainty signal, showing stability under prompt paraphrasing and reasonable calibration against a token-probability baseline. We then study I-CALM, a prompt-based framework that (i) elicits verbal confidence, (i) partially rewards abstention through explicit reward schemes, and (i) adds lightweight normative principles emphasizing truthfulness, humility, and responsibility. Using GPT-5 mini on PopQA as the main setting, we find that confidence-eliciting, abstention-rewarding prompts, especially with norms, reduce the false-answer rate on answered cases mainly by identifying and shifting error-prone cases to abstention and re-calibrating their confidence. This trades coverage for reliability while leaving forced-answer performance largely unchanged. Varying the abstention reward yields a clear abstention–hallucination frontier. Overall, results show the framework can improve selective answering on factual questions without retraining, with the magnitude of effect varying across models and datasets. Code is available at the following link. 11footnotetext: Co-first authors.22footnotetext: Department of Applied Mathematics and Statistics, Johns Hopkins University, Baltimore, MD 21218, USA; Email: hzong4@jh.edu.33footnotetext: Department of Computer Science, Johns Hopkins University, Baltimore, MD 21218, USA; Emails: bli91, ylong17, schan106, jwu235@jh.edu.44footnotetext: Department of Computer Science, Johns Hopkins University, Baltimore, MD 21218, USA; School of Government and Policy, Johns Hopkins University, Washington, D.C. 20001, USA; Vector Institute for Artificial Intelligence, Toronto, ON, Canada; Email: ghadfield@jhu.edu. 1 Introduction Large language models (LLMs) are increasingly used for information seeking, practical guidance, writing, and programming assistance, making factual reliability an important deployment concern [5]. Yet even state-of-the-art models hallucinate, producing fluent falsehoods with high apparent confidence [8, 27]. Among humans, confidently asserting uncertain or false claims is generally considered inappropriate because it reflects a lack of humility, especially when others may rely on the information. Recent work argues that hallucination is shaped by not only pretraining pressures but also post-training evaluation incentives: even error-free training data need not eliminate generative errors, and mainstream binary benchmarks typically make guessing more rewarding than abstaining [34]. In particular, on difficult factual queries, including long-tail facts and questions that tempt models to reproduce common falsehoods, a model may lack the relevant knowledge yet still produce a plausible but incorrect answer rather than acknowledge uncertainty [27, 44, 46]. In this paper, we therefore focus on epistemic abstention for factual questions with a verifiable answer, that is, questions that admit a well-defined, objective ground truth rather than being underspecified or lacking a definitive answer. In this setting, we operationalize hallucination as giving a false answer to such a question and measure it with the false-answer rate (FAR). This raises a key question: how to induce models to answer when justified and abstain when warranted? A natural starting point is uncertainty: effective abstention requires a usable estimate of when answering is likely to be wrong. Many approaches estimate uncertainty using token probabilities [17]. But token-level probabilities are often unavailable or unstable in API-only deployments [59], and probability-based proxies can be sensitive to prompting and to how answer tokens are realized or aggregated [15, 23, 62]. This motivates a simpler black-box alternative: asking the model to express its uncertainty directly in natural language [37, 43, 72]. Because elicited verbal confidence can show useful calibration properties [56], we use self-reported verbal confidence as the inference-time signal that guides the answer-or-abstain decision [58]. But measuring uncertainty is not enough: a model can express low confidence and still answer [74]. The decision to respond is not only determined by confidence estimates, but may also depend on factors including the incentives and norms under which the model operates. We therefore study I-CALM, a prompt-only framework with three coupled components: (i) elicited verbal reports of model confidence, (i) an explicit answer/abstain reward scheme that gives partial credit to “I don’t know,” and (i) a short set of general-purpose norms centered on truthfulness, humility, and responsibility for the effect of statements on others. Prior work shows that instructive task-specific prompting can reduce hallucinations and encourage acknowledgment of uncertainty [35, 9], while principle-based alignment has more often operationalized norms through training or specialized inference-time procedures [3, 4]. In that context, lightweight general-purpose norm prompting—especially when paired with an explicit uncertainty signal and reward framing—remains less explored. Related work has mostly intervened through training-time abstention mechanisms or inference-time risk-control wrappers. Representative examples include explicit [IDK] tokens, abstention-aware reinforcement learning, and risk-controlled filtering [12, 49, 61]. Our goal is not to replace those approaches, but to test whether a lightweight prompt-level intervention can shift answer/abstain behavior in black-box settings by jointly leveraging self-reported verbal confidence, reward framing, and norms. These elements are interdependent: verbal confidence exposes uncertainty, the reward scheme makes the payoff trade-off explicit, and the norms further discourage unsupported guessing. Rather than encouraging blanket refusal, our goal is to reduce low-confidence guessing while preserving performance when the model does know the answer. This paper makes three contributions. First, we show that self-reported verbal confidence is stable under prompt paraphrasing and reasonably informative relative to a token-probability baseline in free-response factual QA. Second, we introduce a unified prompt-only framework combining confidence elicitation, announced answer/abstain payoffs, and lightweight normative guidance. Third, we characterize the resulting selective-answering behavior, including the abstention–hallucination frontier, component ablations, and heterogeneous transfer across models and datasets. 2 Related Work LLM Hallucination and Abstention Mechanisms. Hallucinations remain a central reliability problem for LLMs [27]. Recent work emphasizes that evaluation incentives matter: binary evaluations and leaderboards reward guessing and can discourage abstention [34]. Abstention is now treated as a first-class capability, with its own taxonomies, metrics, and benchmarks [67]. A substantial part of this literature focuses on settings where abstention is itself the correct behavior because the query is unanswerable, unsupported, or underspecified [35, 38, 45, 70]. Closer to our setting, knowledge-gap recognition work asks whether models can tell when they do not know the answer to a question that has a verifiable answer [7, 53]. Existing solutions include training-time abstention via RL, explicit abstention tokens, or semantic uncertainty [1, 12, 28, 29, 49, 57, 66, 69] and inference-time risk-control wrappers [50, 61]. Confidence Estimation and Calibration in LLMs. Our method requires an uncertainty signal that is available at inference time. Surveys and systematic evaluations cover logit-, representation-, semantic-, and consistency-based measures and highlight calibration challenges in generative settings [17, 26]. Prior work comparing token-based and verbalized confidence finds both alignment and gaps [30, 33, 36], and RLHF can amplify overconfident language [8, 40]. A growing line studies how to elicit and calibrate expressed confidence in LLMs. On the prompting side, verbal confidence can be elicited with usable calibration, though prompt design matters [56, 72, 77]. On the training side, methods such as ConfTuner, ADVICE, and LACIE directly improve verbal confidence calibration [41, 54, 55]. An extended discussion appears in Appendix A. 3 Preliminaries: Assessing Self-Reported Verbal Confidence A common way to estimate LLM uncertainty is to aggregate token-level log probabilities after generation [17, 31]. However, our setting requires an uncertainty signal the model can express in its response to support abstention. Post-hoc scores also depend on the aggregation rule and can be distorted by answer realization, relevance, and length [15, 23, 62]. We therefore use self-reported verbal confidence τselfτ^self as an inference-time uncertainty proxy—it reflects automatic, sophisticated self-evaluation during generation [37], is easy to elicit in black-box settings [8], and can be directly acted upon. We evaluate τselfτ^self against post-hoc geometric mean token probability τavgtoken _avg^token along two dimensions: (i) reliability, whether τselfτ^self functions as a confidence signal comparable to τavgtoken _avg^token; and (i) robustness, whether it remains stable under prompt paraphrasing so abstention decisions do not depend on particular wording. Because token-probability-based estimates can be task-sensitive [71], we focus on free-response QA to match the later experiments. Appendix B.3 reports results for the multiple-choice setting. 3.1 Experiment Setup We evaluate GPT-4o mini111GPT-4o mini is used because GPT-5 mini does not expose token-level log probabilities. We also evaluate GPT-4o mini in the main experiments and observe consistent trends across models. [51] on PopQA [46], a set of 14,267 factual questions constructed from Wikidata knowledge triples. Because all questions are factual queries with a verifiable answer, we operationalize hallucinations as false factual assertions, and abstention reflects the model’s lack of knowledge rather than the absence of a correct answer. See Appendix B.1 for setup details. Confidence Elicitation We elicit self-reported verbal confidence τselfτ^self in [0,1][0,1] using the two-part abstention-conditioned prompt framework (Figure 1; Appendix B.1). For each question, the LLM either (i) provides a direct answer with a confidence score, or (i) abstains by first responding “I don’t know” and then providing a best guess with its confidence. We define τselfτ^self as the confidence attached to the evaluated answer: the direct answer when the model answers in the first round, or the best guess when it initially abstains. Figure 1: Overview of the two-stage prompting protocol and downstream analysis. Performance Metrics. We report four metrics: (1) task performance, assessed by the false-answer rate (FAR), which is defined as the fraction of evaluated answers that are incorrect and operationalizes hallucination rate in this setting; (2) correlation between τselfτ^self and τavgtoken _avg^token, measured by Pearson’s r; (3) forecasting loss, quantified by the Brier score [18, 19]; and (4) calibration error, measured by the empirical Expected Calibration Error (ECE ECE) [14, 22]. We defer detailed definitions to Appendix B.1. 3.2 Results Calibration and Reliability. We construct four semantics-preserving prompt variants (Table 1) that keep the abstention-conditioned response structure fixed while varying surface wording. Across templates, τselfτ^self and τavgtoken _avg^token exhibit a moderate, consistent positive correlation around 0.54. This suggests τselfτ^self tracks the token-probability-based confidence reasonably well, even though the two are not identical. The two signals also show comparable probabilistic forecasting and calibration quality. Their Brier scores fall in a similar range (approximately 0.330.33 – 0.350.35), and ECE ECE values are also close. Although τselfτ^self attains marginally lower Brier scores and ECE ECE than τavgtoken _avg^token, the differences are small and the 95% CIs overlap substantially. We also note that the Brier score and ECE ECE values, which deviate from perfect calibration, are consistent with known patterns of imperfect calibration in verbal confidence [8, 48, 54, 72]. Prompt FAR (± 95% CI) Pearson’s r (± 95% CI) Brier Score (± 95% CI) ECE ECE τselfτ^self τavgtoken _avg^token τselfτ^self τavgtoken _avg^token ① 0.6261 ± 0.0079 0.5397 ± 0.0116 0.3363 ± 0.0059 0.3555 ± 0.0055 0.3825 0.4220 ② 0.6230 ± 0.0080 0.5353 ± 0.0117 0.3445 ± 0.0062 0.3473 ± 0.0054 0.3852 0.4130 ③ 0.6221 ± 0.0080 0.5158 ± 0.0120 0.3501 ± 0.0062 0.3448 ± 0.0054 0.3830 0.4073 ④ 0.6208 ± 0.0080 0.5522 ± 0.0114 0.3352 ± 0.0060 0.3456 ± 0.0055 0.3745 0.4106 Table 1: Comparison of self-reported and token-probability-based confidence metrics across paraphrased abstention-conditioned prompt templates for GPT-4o mini on PopQA. Robustness to Prompt Paraphrasing Across templates, FAR, Pearson’s r, and Brier scores for τselfτ^self have overlapping 95% confidence intervals, and its ECE ECE varies only slightly. This suggests that τselfτ^self is not strongly driven by superficial wording differences. Appendix B.2 provides an additional robustness check on confidence-reporting conventions. 4 Incentivizing Abstention Through Reward Schemes and Normative Guidance Given that self-reported verbal confidence is reasonably informative about model uncertainty (Section 3), we next ask whether announced reward schemes and added normative guidance can shift the model’s answer/abstain decisions on factual questions while verbal confidence is elicited. We examine how these interventions affect confidence distributions, selective-answering behavior, and the abstention–hallucination trade-off. 4.1 Experiment Setup Two Reward Schemes. Using factual QA benchmark datasets PopQA [46], TriviaQA [32], and SimpleQA Verified [24], we evaluate GPT-5 mini [52], GPT-4o mini [51], Gemini-3.1-Flash-Lite [20], Meta-Llama-3-8B-Instruct [21], and Qwen3-4B-Instruct-2507 [73] under two announced reward schemes: Scheme A (No Explicit Abstention Reward)Scheme B (Abstention-Rewarding)reward=+R,correct−β,incorrectreward=+R,correct−β,incorrect+γ,“I don’t know” (IDK) array[]@c@ c@Scheme A (No Explicit Abstention Reward)&Scheme B (Abstention-Rewarding)\\[3.0pt] reward= cases+R,&correct\\ -β,&incorrect cases&reward= cases+R,&correct\\ -β,&incorrect\\ +γ,&``I don't know′ (IDK) cases array For each question, we announce the active scheme and keep the same two-stage workflow as in Figure 1. In both schemes, abstention remains permitted through the same response format; the difference is whether a first-round “I don’t know” receives explicit positive reward.222Scheme A does not assign explicit positive credit to abstention, but when β>0β>0 abstaining can still be preferable to a likely-wrong answer. If the model abstains in the first round, it receives +γ+γ regardless of whether its second-round best guess is correct; the best guess is elicited for forced-answer evaluation. We fix R=1R=1, and vary β∈0,1β∈\0,1\ and γ∈0.2,0.4,0.6,0.8γ∈\0.2,0.4,0.6,0.8\ in the prompt (Appendix C.1) to examine model behavior under different reward configurations. We use (+R,−β)(+R,-β) and (+R,−β,+γ)(+R,-β,+γ) as shorthand for Scheme A and Scheme B reward configurations. Adding Normative Guidance. Motivated by prior findings that carefully crafted system prompts can increase abstention [35], we also evaluate a prompt condition that introduces a concise set of principles focused on truthfulness, humility, and responsibility (Figure 2) on top of Scheme B. We do not claim that this list is unique or optimal; rather, the goal is to assess whether lightweight normative guidance provides additional abstention signal beyond payoff framing. We therefore compare Scheme A, Scheme B, and Scheme B with norms across models and datasets. 1. Always tell the truth (do not lie or deceive). 2. Treat all topics with respect and humility. 3. Do not make assumptions without evidence. 4. Be responsible for your words and their consequences. 5. Honor reason, clarity, and understanding. Figure 2: Normative principles oriented toward truthfulness, humility, and responsibility. Metrics. Because questions in these factual QA benchmarks still have a verifiable answer even when the model abstains, abstention is not itself a gold label. We therefore evaluate the method as a selective-prediction system. Specifically, FARansweredFAR_answered measures the false-answer rate among first-round surfaced answers, while FARoverallFAR_overall measures the false-answer rate under forced answering after replacing first-round abstentions with second-round best guesses. We also report Coverage, defined as the fraction of questions answered in the first round; equivalently, 1−Coverage1-Coverage is the first-round abstention rate. We additionally report the Abstention-to-Error Ratio (AER), defined as the fraction of eventual forced-answer errors that were preceded by a first-round abstention. A higher AER means the model more often flags uncertainty before an error that would otherwise be surfaced under forced answering. Full definitions appear in Appendix C.1. 4.2 Results Unless otherwise noted, the main text focuses on GPT-5 mini on PopQA. We use Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4) as representative setups to analyze the general trends. Figure 3 first provides a PopQA cross-model overview. Appendix C.3 reports the results of the full set of reward configurations for GPT-5 mini and GPT-4o mini on PopQA, and Appendix C.5 reports broader transfer experiments, showing that the same selective-answering mechanism transfers most clearly when there is more first-round hallucination risk to remove. Appendix C.6 compares these PopQA results to the closest prior abstention and hallucination-mitigation papers. Results are robust to reward scaling (Appendix C.2). Figure 3: Cross-model PopQA comparison under representative setups. Panels show FARansweredFAR_answered, AER, ECE ECE for first-round answers, and Brier score for first-round answers, under Pure Eval, Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). Performance under Schemes A, B, and B with Norms. Figure 3 summarizes PopQA results for five models under representative configurations. Across models, the same broad selective-answering pattern appears: rewarding abstention generally reduces FARansweredFAR_answered and increases AER, and adding norms often strengthens this shift, although the magnitude of the gain varies by model. As a baseline, directly prompting GPT-5 mini without reward framing or verbal confidence (Pure Eval) yields FARanswered=52.3%FAR_answered=52.3\%. Under the representative PopQA setting, Scheme A lowers this to 48.2%48.2\%, Scheme B to 41.0%41.0\%, and Scheme B with norms to 34.2%34.2\%. These risk reductions come with reduced coverage, which falls from 96.5%96.5\% in Pure Eval to 84.0%84.0\% in Scheme A, 67.9%67.9\% in Scheme B, and 55.3%55.3\% in Scheme B with norms (Appendix C.3). By contrast, FARoverallFAR_overall remains similar across schemes with overlapping 95%95\% confidence intervals. We therefore interpret the main effect as improved selective answering, not improved forced-answer accuracy. Scheme B with norms also achieves the highest total reward (5039.6 vs. 3577.2 for Scheme B; Appendix C.3), indicating a better operating point: fewer costly false answers while still collecting abstention reward. Compared with the other models for which Scheme B with norms often improves calibration, GPT-5 mini appears both stronger and more stable on calibration in this setting. We further examine performance differences on questions based on common and rare facts. Appendix C.4 shows that under all reward schemes, models have lower false-answer rates and better calibration on common-fact questions than on rare-fact questions, indicating higher reliability on high-popularity knowledge; Scheme B with norms performs best overall in reducing false answers. As shown in Figure 5, GPT-5 mini’s Pure Eval baseline rarely signals uncertainty, yielding an upper bound on AER of at most 5.8%.333Pure Eval is single-round, so AER (which requires a best-guess answer after an initial abstention) is not directly computable. We therefore report a conservative upper bound by counting spontaneous “I don’t know” responses as abstentions and dividing by the number of one-round incorrect answers. Scheme B consistently yields higher AER than Scheme A, and adding norms to Scheme B increases it further. Additionally, setting the false-answer penalty β=1β=1 raises AER relative to β=0β=0, whereas increasing abstention reward γ within 0.2,0.4,0.6,0.8\0.2,0.4,0.6,0.8\ has only a marginal effect. Figure 5 shows that, among questions answered incorrectly under Pure Eval, Scheme B with norms converts roughly 60% into abstentions and leaves only about 30% as first-round incorrect answers. Figure 4: AER for GPT-5 mini on PopQA across reward configurations. The values in parentheses on the x-axis represent abstention reward. Figure 5: Distribution of PopQA questions that were incorrect under Pure Eval, reclassified under the representative scheme setups (GPT-5 mini). Confidence Distribution. Figure 6 shows first-round answered cases on the left and second-round best guesses after an initial “I don’t know” on the right across the three reward schemes for GPT-5 mini. Relative to Scheme A, Scheme B moves much of the incorrect and medium-to-low-confidence mass into the best-guess round at even lower confidence, and Scheme B with norms strengthens this pattern. Best-guess confidence is concentrated in the low-confidence region, peaking around 0.3 and remaining far below the roughly 0.80.8–1.01.0 concentration for first-round answers, consistent with the fact that these second-round outputs are produced under expressed uncertainty. Overall, this pattern indicates that the framework we propose induces the model to identify error-prone cases, abstain on them, and re-calibrate their associated confidence. Figure 6: Confidence distributions for GPT-5 mini under representative PopQA reward schemes. Left panel: first-round surfaced answers. Right panel: second-round best guesses after an initial “I don’t know.” Bars are colored by correctness. The Hallucination-Abstention Trade-off. Figure 7 plots FARansweredFAR_answered against first-round abstention rate across all reward configurations for GPT-5 mini. We observe a clear hallucination–abstention trade-off: as abstention rate increases, FARansweredFAR_answered correspondingly declines. For a fixed abstention reward, penalizing incorrect answers (β=1β=1) consistently yields lower hallucination rates and higher abstention rates than a zero penalty (β=0β=0). However, regardless of how the rewards are adjusted, the model exhibits the same roughly monotone frontier between abstention rate and FARansweredFAR_answered. Reward framing therefore seems to move the model along a stable trade-off curve rather than changing its shape. It also points to a simple deployment control: a user-specified reward payoff could be used to select an operating point along the abstention–hallucination frontier. Figure 7: The abstention–hallucination frontier for GPT-5 mini on PopQA across all tested reward configurations. Numbers in parentheses denote abstention reward; point color indicates false penalty. 5 Discussion Ablation Study. The full prompting framework has three components: verbal confidence elicitation, reward framing, and normative guidance. Since the effect of norms is already evaluated in Section 4, here we ablate the first two components under representative setup Scheme B (+1,−1,+0.4)(+1,-1,+0.4). Table 2 shows that the reward-scheme-only variant attains the lowest FARansweredFAR_answered and the highest AER, but its first-round coverage is only 29.8%29.8\%, indicating a substantial loss of coverage. Notably, among questions answered correctly by full Scheme B, the reward-scheme-only variant abstains on nearly half of them. This suggests that reward framing alone pushes the model toward excessive risk aversion. Moreover, on the subset that the reward-scheme-only ablation does answer, the full Scheme B method still achieves a lower FAR (0.197 vs. 0.243; Appendix D). On the other hand, removing the reward scheme while keeping only verbal confidence leads to higher FAR and worse calibration than the full method. Taken together, the ablation suggests a division of labor: reward framing drives abstention, while verbal confidence helps keep abstention from becoming overly aggressive. Detailed ablation results, with an additional study showing explicit “I don’t know” wording helps but is not sufficient, are in Appendix D. Coverage AER FARansweredFAR_answered ECE ECE Brier Score (± 95% CI) (± 95% CI) Answered Overall Answered Overall Confidence only 0.860 0.214 0.487±0.0090.487± 0.009 0.2323 0.2088 0.2552±0.00780.2552± 0.0078 0.2425±0.00720.2425± 0.0072 Reward only 0.298 0.868 0.243±0.0130.243± 0.013 – – – – Full Scheme B 0.679 0.499 0.410±0.0100.410± 0.010 0.1506 0.1206 0.2102±0.00820.2102± 0.0082 0.1933±0.00680.1933± 0.0068 Table 2: Ablation results relative to Scheme B (+1,−1,+0.4+1,-1,+0.4) for GPT-5 mini on PopQA. Sensitivity to Extreme Reward Magnitude. Holding two reward terms fixed and varying the third, we find that abstention and hallucination behavior is far more sensitive to the abstention reward than to the correct-answer reward or false-answer penalty. This complements prior evidence that LLMs often fail to adjust abstention behavior even when false-answer penalties vary substantially [60]. Table 3 shows that scaling correct-answer reward (from 1 to 100) or false-answer penalty (from -1 to -100) leads to only marginal changes in model behavior. In contrast, scaling the abstention reward (from 0.4 to 40) produces a substantially larger shift in coverage, FARansweredFAR_answered, and AER. Reward (Correct) Penalty (Incorrect) Reward (Abstain) FAR (± 95% CI) Coverage AER Answered Overall 1 -1 0.4 0.410±0.0100.410± 0.010 0.555±0.0080.555± 0.008 0.679 0.499 100 -1 0.4 0.438±0.0090.438± 0.009 0.557±0.0080.557± 0.008 0.742 0.417 1 -100 0.4 0.416±0.0100.416± 0.010 0.560±0.0080.560± 0.008 0.684 0.491 1 -1 40 0.283±0.0110.283± 0.011 0.546±0.0080.546± 0.008 0.428 0.778 Table 3: Effect of scaling one reward component to an extreme value while holding the other two fixed under Scheme B for GPT-5 mini on PopQA. Departure from the Bayes-Optimal Threshold Rule. Appendix D.3 shows that the model’s Bayes-optimal policy in this payoff-based answer/abstain setting is a simple threshold rule: answer iff p≥τ=(γ+β)/(R+β)p≥τ=(γ+β)/(R+β), where p is the model’s subjective probability of correctness, which we operationalize via the elicited verbal confidence. Under Scheme B (+1,−1,+0.4)(+1,-1,+0.4), this gives τBayesB=0.7 _Bayes^B=0.7, so a Bayes-optimal model would answer only above 0.7 confidence and abstain below it. Empirically, however, the model does not exhibit a sharp cutoff and continues to answer across much of the confidence range (Figure 6). We use the derived threshold rule as a reference point for interpreting departures from ideal decision-making, not as a claim that current LLMs literally implement this policy. Deployment Extension: Post-Hoc Confidence Thresholding. Although the model does not follow a clean Bayes-optimal threshold rule, the elicited confidence can still support downstream filtering. Appendix C.7 studies post-hoc confidence thresholding with finite-sample FAR guarantees: on a held-out calibration split, we choose thresholds so that surfaced answers satisfy a user-specified FAR target with high probability. On PopQA, the announced reward scheme shifts this downstream coverage–risk trade-off as well. Relative to Scheme A, Scheme B allows the filter to surface slightly more answers at the same moderate FAR targets, while Scheme B with norms often yields lower risk in the mid-confidence region without uniformly increasing coverage. We view this as a deployment-oriented extension of the main results rather than a core methodological contribution. 6 Conclusion, Limitations, and Future Work This work shows that prompt-level incentive framing, especially when paired with lightweight normative guidance, can improve selective answering and reduce hallucination risk on factual questions with a verifiable answer without changing model weights. By eliciting self-reported verbal confidence, explicitly rewarding “I don’t know,” and adding truthfulness-, humility-, and responsibility-oriented norms, the method reduces the false-answer rate among surfaced first-round answers by identifying and abstaining on many error-prone cases while re-calibrating their confidence. The resulting benefit should be interpreted as movement along a tunable abstention–hallucination frontier rather than as evidence of improved underlying factual competence under forced answering. We therefore view this prompt-based abstention framework as a lightweight routing/control mechanism that complements retrieval, tool use, escalation, and training-based methods. Our study focuses on epistemic abstention on factual questions with a verifiable answer, not cases where abstention is correct because a question is unanswerable, unsupported, or underspecified; extending the framework to such settings is an important direction for future work. More broadly, the answer-or-abstain behavior we induce is an empirical property of the model and setting. It may change under domain shift or drift in multi-turn interactions [42]. Self-reported confidence is also imperfect, and models may not reliably convert stated confidence into an optimal answer/abstain policy, consistent with evidence that LLM decisions can be weakly coupled to verbal confidence [60]. In practice, abstention should often trigger retrieval, tool use, or escalation rather than a forced best guess; integrating incentivized abstention into uncertainty-aware agent pipelines is a natural extension [25, 76]. Finally, our approach is complementary to training-time methods that improve either abstention behavior or verbal confidence calibration; combining prompt-level incentives with lightweight calibration training may further strengthen confidence-aware abstention and hallucination mitigation. Acknowledgments This work is supported by the AI2050 program at Schmidt Sciences (Grant G-25-67962). We are grateful to the members of the Normativity Lab, Matthew Renze, and Adam Tauman Kalai for their valuable feedback. We also thank the participants in EN.601.669 AI Safety, Alignment, & Governance (Fall 2025) at Johns Hopkins University for their valuable feedback. References [1] H. An and Y. Xu (2025) Teaching llms to abstain via fine-grained semantic confidence reward. arXiv preprint arXiv:2510.24020. Cited by: §2. [2] A. N. Angelopoulos, S. Bates, E. J. Candès, M. I. Jordan, and L. Lei (2025) Learn then test: calibrating predictive algorithms to achieve risk control. The Annals of Applied Statistics 19 (2), p. 1641–1662. External Links: Document Cited by: §C.7.2. [3] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. (2022) Constitutional ai: harmlessness from ai feedback. 2022. arXiv preprint arXiv:2212.08073 8 (3). Cited by: §A.1, §1. [4] H. Bell, C. Zhang, M. M. Haque, D. Potdar, S. Zaman, and B. Fain (2026) Reflect: transparent principle-guided reasoning for constitutional alignment at scale. arXiv preprint arXiv:2601.18730. Cited by: §A.1, §1. [5] A. Chatterji, T. Cunningham, D. J. Deming, Z. Hitzig, C. Ong, C. Y. Shan, and K. Wadman (2025) How people use chatgpt. Technical report National Bureau of Economic Research. Cited by: §1. [6] T. Chen, A. Asai, L. Zettlemoyer, H. Hajishirzi, and F. Brahman (2025) Train for truth, keep the skills: binary retrieval-augmented reward mitigates hallucinations. arXiv preprint arXiv:2510.17733. Cited by: §C.6. [7] Q. Cheng, T. Sun, X. Liu, W. Zhang, Z. Yin, S. Li, L. Li, Z. He, K. Chen, and X. Qiu (2024) Can ai assistants know what they don’t know?. arXiv preprint arXiv:2401.13275. Cited by: §A.1, §2. [8] P. Chhikara (2025) Mind the confidence gap: overconfidence, calibration, and distractor effects in large language models. arXiv preprint arXiv:2502.11028. Cited by: §A.2, §1, §2, §3.2, §3. [9] G. Chujie, S. Wu, Y. Huang, D. Chen, Q. Zhang, Z. Fu, Y. Wan, L. Sun, and X. Zhang (2024) Honestllm: toward an honest and helpful large language model. Advances in Neural Information Processing Systems 37, p. 7213–7255. Cited by: §1. [10] C. J. Clopper and E. S. Pearson (1934) The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26 (4), p. 404–413. Cited by: §C.7. [11] R. Cohen, R. Biswas, and G. De Melo (2025) Infact: informativeness alignment for improved llm factuality. arXiv preprint arXiv:2505.20487. Cited by: §C.6. [12] R. Cohen, K. Dobler, E. Biran, and G. de Melo (2024) I don’t know: explicit modeling of uncertainty with an [idk] token. Advances in Neural Information Processing Systems 37, p. 10935–10958. Cited by: §A.1, §C.6, §1, §2. [13] Y. Dai (2026) Rescaling confidence: what scale design reveals about llm metacognition. arXiv preprint arXiv:2603.09309. Cited by: §B.2. [14] M. H. DeGroot and S. E. Fienberg (1983) The comparison and evaluation of forecasters. Journal of the Royal Statistical Society: Series D (The Statistician) 32 (1-2), p. 12–22. Cited by: §B.1.4, §3.1. [15] J. Duan, H. Cheng, S. Wang, A. Zavalny, C. Wang, R. Xu, B. Kailkhura, and K. Xu (2024) Shifting attention to relevance: towards the predictive uncertainty quantification of free-form large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 5050–5063. Cited by: §1, §3. [16] S. Feng, W. Shi, Y. Wang, W. Ding, V. Balachandran, and Y. Tsvetkov (2024) Don’t hallucinate, abstain: identifying llm knowledge gaps via multi-llm collaboration. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 14664–14690. Cited by: §C.1.3. [17] J. Geng, F. Cai, Y. Wang, H. Koeppl, P. Nakov, and I. Gurevych (2024) A survey of confidence estimation and calibration in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 6577–6595. Cited by: §A.2, §1, §2, §3. [18] W. B. Glenn et al. (1950) Verification of forecasts expressed in terms of probability. Monthly weather review 78 (1), p. 1–3. Cited by: §B.1.4, §3.1. [19] T. Gneiting and A. E. Raftery (2007) Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association 102, p. 359–378. External Links: Document Cited by: §B.1.4, §3.1. [20] Google DeepMind (2026-03) Gemini 3.1 flash-lite model card. Technical report Google DeepMind. External Links: Link Cited by: §4.1. [21] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. (2024) The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §4.1. [22] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger (2017) On calibration of modern neural networks. In International conference on machine learning, p. 1321–1330. Cited by: §B.1.4, §3.1. [23] N. Gupta, H. Narasimhan, W. Jitkrittum, A. S. Rawat, A. K. Menon, and S. Kumar (2024) Language model cascades: token-level uncertainty and beyond. arXiv preprint arXiv:2404.10136. Cited by: §1, §3. [24] L. Haas, G. Yona, G. D’Antonio, S. Goldshtein, and D. Das (2025) Simpleqa verified: a reliable factuality benchmark to measure parametric knowledge. arXiv preprint arXiv:2509.07968. Cited by: §C.5.3, §4.1. [25] J. Han, W. Buntine, and E. Shareghi (2024) Towards uncertainty-aware language agent. In Findings of the Association for Computational Linguistics: ACL 2024, p. 6662–6685. Cited by: §6. [26] C. Hobelsberger, T. Winner, A. Nawroth, O. Mitevski, and A. Haensch (2025) Systematic evaluation of uncertainty estimation methods in large language models. arXiv preprint arXiv:2510.20460. Cited by: §A.2, §2. [27] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. (2025) A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43 (2), p. 1–55. Cited by: §A.1, §1, §2. [28] N. Jain, A. Shrivastava, C. Zhu, D. Liu, A. Samuel, A. Panda, A. Kumar, M. Goldblum, and T. Goldstein (2024) Refusal tokens: a simple way to calibrate refusals in large language models. arXiv preprint arXiv:2412.06748. Cited by: §2. [29] A. Jha, A. Mahajan, A. V. Aravindan, P. Saravanan, S. S. Policharla, and S. C. Gehlot (2026) Rewarding intellectual humility learning when not to answer in large language models. arXiv preprint arXiv:2601.20126. Cited by: §2. [30] Z. Ji, L. Yu, Y. Koishekenov, Y. Bang, A. Hartshorn, A. Schelten, C. Zhang, P. Fung, and N. Cancedda (2025) Calibrating verbal uncertainty as a linear feature to reduce hallucinations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 3769–3793. Cited by: §A.2, §C.6, §2. [31] Z. Jiang, J. Araki, H. Ding, and G. Neubig (2021) How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics 9, p. 962–977. Cited by: §A.2, §B.3.3, §3. [32] M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer (2017) Triviaqa: a large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1601–1611. Cited by: §4.1. [33] S. Kadavath, T. Conerly, A. Askell, T. Henighan, D. Drain, E. Perez, N. Schiefer, Z. Hatfield-Dodds, N. DasSarma, E. Tran-Johnson, et al. (2022) Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §A.2, §2. [34] A. T. Kalai, O. Nachum, S. S. Vempala, and E. Zhang (2025) Why language models hallucinate. arXiv preprint arXiv:2509.04664. Cited by: §A.1, §A.1, §1, §2. [35] P. Kirichenko, M. Ibrahim, K. Chaudhuri, and S. J. Bell (2025) Abstentionbench: reasoning llms fail on unanswerable questions. arXiv preprint arXiv:2506.09038. Cited by: §A.1, §A.1, §1, §2, §4.1. [36] A. Kumar, R. Morabito, S. Umbet, J. Kabbara, and A. Emami (2024) Confidence under the hood: an investigation into the confidence-probability alignment in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 315–334. Cited by: §A.2, §2. [37] D. Kumaran, A. Conmy, F. Barbero, S. Osindero, V. Patraucean, and P. Velickovic (2026) How do llms compute verbal confidence. arXiv preprint arXiv:2603.17839. Cited by: §A.2, §1, §3. [38] M. J. Lavi, T. Milo, and M. Geva (2025) Detecting (un) answerability in large language models with linear directions. arXiv preprint arXiv:2509.22449. Cited by: §2. [39] M. Lee, K. Kim, T. Kim, and S. Park (2024) Selective generation for controllable language models. Advances in Neural Information Processing Systems 37, p. 50494–50527. Cited by: §A.1. [40] J. Leng, C. Huang, B. Zhu, and J. Huang (2024) Taming overconfidence in llms: reward calibration in rlhf. arXiv preprint arXiv:2410.09724. Cited by: §A.2, §2. [41] Y. Li, M. Xiong, J. Wu, and B. Hooi (2025) Conftuner: training large language models to express their confidence verbally. arXiv preprint arXiv:2508.18847. Cited by: §A.2, §2. [42] Y. Li, R. Krishnan, and R. Padman (2026) Consistency of large reasoning models under multi-turn attacks. arXiv preprint arXiv:2602.13093. External Links: Link Cited by: §6. [43] S. Lin, J. Hilton, and O. Evans (2022) Teaching models to express their uncertainty in words. arXiv preprint arXiv:2205.14334. Cited by: §A.2, §1. [44] S. Lin, J. Hilton, and O. Evans (2022) TruthfulQA: measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958. Cited by: §1. [45] N. Madhusudhan, S. T. Madhusudhan, V. Yadav, and M. Hashemi (2025) Do llms know when to not answer? investigating abstention abilities of large language models. In Proceedings of the 31st International Conference on Computational Linguistics, p. 9329–9345. Cited by: §A.1, §2. [46] A. Mallen, A. Asai, V. Zhong, R. Das, D. Khashabi, and H. Hajishirzi (2023) When not to trust language models: investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers), p. 9802–9822. Cited by: §1, §3.1, §4.1. [47] Y. Mao, T. Durand, N. Mehrasa, J. He, and M. Ester (2025) Calibrating llms for selective prediction: balancing coverage and risk. In Socially Responsible and Trustworthy Foundation Models at NeurIPS 2025, Cited by: §C.7. [48] S. J. Mielke, A. Szlam, E. Dinan, and Y. Boureau (2022) Reducing conversational agents’ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics 10, p. 857–872. Cited by: §3.2. [49] M. A. Mohamadi, T. Wang, and Z. Li (2025) Honesty over accuracy: trustworthy language models through reinforced hesitation. arXiv preprint arXiv:2511.11500. Cited by: §A.1, §1, §2. [50] M. Oehri, G. Conti, K. Pather, A. Rossi, L. Serra, A. Parody, R. Johannesen, A. Petersen, and A. Krasniqi (2025) Trusted uncertainty in large language models: a unified framework for confidence calibration and risk-controlled refusal. arXiv preprint arXiv:2509.01455. Cited by: §A.1, §A.2, §C.7, §2. [51] OpenAI (2024-07) GPT-4o mini: advancing cost-efficient intelligence. Note: https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Cited by: §B.3.1, §3.1, §4.1. [52] OpenAI (2025-08) Introducing gpt-5 for developers. Note: https://openai.com/index/introducing-gpt-5-for-developers/ Cited by: §4.1. [53] S. Qin, L. Zhou, L. Sun, and N. Wang (2026) Do large language models know when they lack knowledge?. Electronics 15 (2), p. 253. Cited by: §A.1, §D.2, §2. [54] K. J. Seo, S. Lim, and T. Kim (2025) ADVICE: answer-dependent verbalized confidence estimation. arXiv preprint arXiv:2510.10913. Cited by: §A.2, §2, §3.2. [55] E. Stengel-Eskin, P. Hase, and M. Bansal (2024) LACIE: listener-aware finetuning for calibration in large language models. Advances in Neural Information Processing Systems 37, p. 43080–43106. Cited by: §A.2, §2. [56] K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. D. Manning (2023) Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 5433–5442. Cited by: §A.2, §1, §2. [57] B. A. Tjandra, M. Razzak, J. Kossen, K. Handa, and Y. Gal (2024) Fine-tuning large language models to appropriately abstain with semantic entropy. arXiv preprint arXiv:2410.17234. Cited by: §2. [58] C. Tomani, K. Chaudhuri, I. Evtimov, D. Cremers, and M. Ibrahim (2024) Uncertainty-based abstention in llms improves safety and reduces hallucinations. arXiv preprint arXiv:2404.10960. Cited by: §A.1, §1. [59] D. Ulmer, M. Gubri, H. Lee, S. Yun, and S. Oh (2024) Calibrating large language models using their generations only. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 15440–15459. Cited by: §A.2, §1. [60] J. Wang, Y. Zhou, S. Devic, and D. Fu (2026) Are llm decisions faithful to verbal confidence?. arXiv preprint arXiv:2601.07767. Cited by: §A.2, §5, §6. [61] Q. Wang, Y. Fan, and X. E. Wang (2025) SAFER: risk-constrained sample-then-filter in large language models. arXiv preprint arXiv:2510.10193. Cited by: §A.1, §A.2, §C.7, §1, §2. [62] X. Wang, B. Ma, C. Hu, L. Weber-Genzel, P. Röttger, F. Kreuter, D. Hovy, and B. Plank (2024) ” My answer is c”: first-token probabilities do not match text answers in instruction-tuned language models. arXiv preprint arXiv:2402.14499. Cited by: §1, §3. [63] Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, et al. (2024) Mmlu-pro: a more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems 37, p. 95266–95290. Cited by: §B.3.1, §B.3.1. [64] J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus (2024) Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368. Cited by: §C.5.3. [65] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. (2022) Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, p. 24824–24837. Cited by: §B.3.1. [66] Z. Wei, X. Yang, K. Sun, J. Wang, R. Shao, S. Chen, M. Kachuee, T. Gollapudi, T. Liao, N. Scheffer, et al. (2025) TruthRL: incentivizing truthful llms via reinforcement learning. arXiv preprint arXiv:2509.25760. Cited by: §A.1, §2. [67] B. Wen, J. Yao, S. Feng, C. Xu, Y. Tsvetkov, B. Howe, and L. L. Wang (2025) Know your limits: a survey of abstention in large language models. Transactions of the Association for Computational Linguistics 13, p. 529–556. Cited by: §A.1, §2. [68] B. L. Wiens (2003) A fixed sequence bonferroni procedure for testing multiple endpoints. Pharmaceutical Statistics 2 (3), p. 211–215. External Links: Document Cited by: §C.7.2. [69] J. Wu, J. Liu, Z. Zeng, T. Zhan, T. Cai, and W. Huang (2025) Mitigating llm hallucination via behaviorally calibrated reinforcement learning. arXiv preprint arXiv:2512.19920. Cited by: §A.1, §2. [70] T. Wu, C. Zhou, G. Zhao, H. Cao, Y. Pu, and J. Yang (2025) When robots should say” i don’t know”: benchmarking abstention in embodied question answering. arXiv preprint arXiv:2512.04597. Cited by: §A.1, §2. [71] Z. Xia, J. Xu, Y. Zhang, and H. Liu (2025) A survey of uncertainty estimation methods on large language models. arXiv preprint arXiv:2503.00172. Cited by: §3. [72] M. Xiong, Z. Hu, X. Lu, Y. Li, J. Fu, J. He, and B. Hooi (2023) Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms. arXiv preprint arXiv:2306.13063. Cited by: §A.2, §1, §2, §3.2. [73] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. (2025) Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §4.1. [74] G. Yona, R. Aharoni, and M. Geva (2024) Can large language models faithfully express their intrinsic uncertainty in words?. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 7752–7764. Cited by: §1. [75] H. Zhang, S. Diao, Y. Lin, Y. Fung, Q. Lian, X. Wang, Y. Chen, H. Ji, and T. Zhang (2024) R-tuning: instructing large language models to say ‘i don’t know’. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 7113–7139. Cited by: §A.1. [76] J. Zhang, P. K. Choubey, K. Huang, C. Xiong, and C. Wu (2026) Agentic uncertainty quantification. arXiv preprint arXiv:2601.15703. Cited by: §6. [77] X. Zhao, H. Zhang, X. Pan, W. Yao, D. Yu, T. Wu, and J. Chen (2024) Fact-and-reflection (far) improves confidence calibration of large language models. In Findings of the Association for Computational Linguistics: ACL 2024, p. 8702–8718. Cited by: §A.2, §2. Appendix Contents Appendix A Extended Related Work This appendix expands the related-work discussion. We organize the literature along two axes: (i) hallucination mitigation, abstention, and selective answering; and (i) confidence estimation and calibration. Our setup sits at their intersection: we keep the base model fixed, elicit self-reported confidence, and use prompt-level payoff framing plus humility- and truthfulness-oriented norms to shape answer/abstain behavior. A.1 Hallucination Mitigation, Abstention, and Selective Answering Hallucination and abstention are now well-developed subliteratures. Broad surveys frame hallucination as a central reliability problem and abstention as a capability that must be evaluated in its own right [27, 67]. [34] provide a key motivation for our paper by arguing that hallucinations reflect both pretraining statistical pressures and post-training evaluation incentives, with mainstream binary grading discouraging abstention. Our setting focuses on factual questions with a verifiable answer, but the model may not know it. Much existing abstention work focuses instead on cases where withholding an answer is itself the gold behavior because the query is unanswerable, unsupported, or underspecified. Representative benchmarks and evaluations study precisely this regime [35, 45, 70]. Closer to our setup, knowledge-gap recognition work asks whether models can detect when they lack the knowledge needed to answer questions with a verifiable answer [7, 53], and uncertainty-based abstention work shows that abstaining on the right cases can reduce hallucinations and improve reliability [58]. Methodologically, prior work addresses abstention in two main ways. Training-time approaches teach it directly, for example through explicit abstention outputs or abstention-aware reward shaping [12, 49, 66, 69, 75]. Inference-time approaches instead wrap a fixed model with calibrated refusal or risk-control procedures [39, 50, 61]. Our method sits between these lines: like inference-time wrappers, it keeps the base model fixed; like [34]’s explicit-confidence-target proposal, it uses prompt-level payoff framing; unlike their benchmark-redesign agenda, it studies a black-box deployment intervention. Finally, we treat short truthfulness- and humility-oriented norms as a lightweight additive intervention rather than a claim about a unique constitution. This design is closest in spirit to principle-based alignment work: Constitutional AI provides the basic constitutional framing but is training-based [3], while Reflect studies an inference-time constitutional alternative [4]; it is also consistent with evidence that system-level prompting can materially change abstention behavior without updating model weights [35]. A.2 Confidence Estimation, Verbal Confidence, and Calibration in LLMs Confidence estimation in LLMs spans logit-based, representation-based, semantic, and consistency-based methods, and recent surveys show that calibration remains difficult in generative settings [17, 26]. Classic QA calibration work already showed that model probabilities can be miscalibrated and brittle under domain shift [31], while black-box settings motivate generation-only calibration methods that do not require logits [59]. These constraints matter for our setting, where we want a signal that is both observable and actionable at inference time. A growing line of work studies verbal confidence directly. Prompt-based work shows that useful confidence estimates can often be elicited in natural language, although performance depends on prompt design [56, 72]. Related prompting work such as Fact-and-Reflection improves calibration by structuring reflection before the final answer [77], and training-based approaches such as Teaching Models to Express Their Uncertainty in Words, ConfTuner, ADVICE, and LACIE aim to make verbal confidence itself better calibrated [41, 43, 54, 55]. At the same time, internal probabilities and verbalized confidence are not identical. Prior work finds partial alignment between the two signals, but also systematic gaps [33, 36, 30]. Other work shows that RLHF can amplify verbal overconfidence [8, 40], while recent mechanistic evidence suggests that verbal confidence reflects richer answer-quality evaluations than token probability alone [37]. This combination of usefulness and imperfection is exactly why we treat self-reported confidence as a practical signal for inference-time decisions rather than as a fully faithful readout of model belief. That distinction also matters for decision-making. Risk-controlled abstention methods formalize how confidence should map to coverage and error [50, 61], but recent evidence suggests that current models may still fail to convert verbal confidence into risk-sensitive answer/abstain behavior [60]. Our work therefore studies both pieces together: whether verbal confidence is informative enough to use, and whether prompt-level payoff framing can better align answer/abstain decisions with that expressed uncertainty. Appendix B Additional details for Section 3 B.1 Detailed Experiment Setup B.1.1 Prompt Template Prompt ① Give your confidence about your answer in decimal number between 0 and 1, where 0 means you are completely unsure and 1 means you are completely sure. Use four decimal places for precision without unnecessary rounding (e.g., 0.8765). If you truly don’t know, first say “I don’t know” and do not give a confidence for that (in that case just leave it blank). If you choose to answer, then provide your answer and your confidence for it. Next, if you said “I don’t know”, then in the same response, give your best possible guess for the answer and your confidence for it. Format your response as below such that you must begin your response by saying “Answer:” first:Answer: <your first answer> Confidence: <0 - 1> If you answered ‘I don’t know’, also include: Best Guess: <your best possible answer> Best Guess Confidence: <0 - 1> Question: q B.1.2 Reproducibility Notes. We query gpt-4o-mini-2024-07-18 via the OpenAI Chat Completions API with temperature 0 and token log-probabilities enabled (logprobs:true). The model’s final answer is normalized by lowercasing, stripping punctuation, and trimming whitespace, and compared to the PopQA reference set possible_answers. An output is marked correct if the normalized prediction is a substring of any reference string or vice versa. B.1.3 Definition of Geometric Mean Token Probability We calculate τavgtoken _avg^token from the token-level log probabilities logp(t) p(t) returned by the API for the final answer. Specifically, we compute the geometric mean probability over answer tokens: τavgtoken=exp(1|tokens|∑t∈tokenslogp(t)). _avg^token= ( 1|tokens| _t p(t) ). B.1.4 Definitions of Performance Metrics False Answer Rate (FAR). For each prompt template, we compute FAR as the proportion of questions answered incorrectly. We report prompt-level FAR with its 95% confidence interval. FAR=NincorrectNtotal.FAR= N_incorrectN_total. Pearson Correlation. Pearson’s correlation quantifies the linear association between τselfτ^self and τavgtoken _avg^token. For each prompt template, we compute the correlation across all questions and report the corresponding 95% confidence interval. Brier Score. Probabilistic forecast accuracy is commonly evaluated using proper scoring rules [19], which measure the discrepancy between predicted probabilities and observed outcomes. We employ the Brier score [18], one of the most widely used proper scoring rules. The Brier score computes the mean squared error between predicted probabilities and observed binary outcomes, with lower values indicating more accurate probabilistic forecasts. For a single prediction with confidence p and correctness label y∈0,1y∈\0,1\: Brier(p,y)=(p−y)2Brier(p,y)=(p-y)^2 For each prompt template, we compute the Brier score for every question and report the average Brier score across all questions, together with its 95% confidence interval. Expected Calibration Error (ECE ECE). Beyond the overall probabilistic forecasting accuracy, we also care about whether these forecasts are trustworthy, i.e., whether stated probabilities reliably correspond to empirical frequencies of correctness. This property is formalized as calibration [14, 22], which assesses how well confidence estimates track observed correctness across the probability spectrum. We adopt the Expected Calibration Error (ECE ECE) to measure such calibration quality. ECE ECE is defined as the expected absolute gap between a model’s probabilistic confidence and its empirical accuracy at that confidence level: ECE^=p^[|Pr(y^=y|p^)−p^|]. ECE=E_ p[|Pr( y=y| p)- p|]. Here, p^∈[0,1] p∈[0,1] denotes the model’s confidence estimates and Pr(y^=y|p^)Pr( y=y| p) denotes the true correctness probability among predictions made with confidence p p. A lower ECE ECE indicates closer alignment between predicted probabilities and empirical accuracy, although such alignment may still coexist with substantial forecasting loss. Since this expectation cannot be computed exactly from finite samples, we estimate ECE ECE using a binned approximation. The unit interval [0,1][0,1] is partitioned into K equal-width bins: Bk=[kK,k+1K),k=0,1,…,K−2,[K−1K,1],k=K−1.B_k= cases [ kK, k+1K ),&k=0,1,…,K-2,\\ [ K-1K,1 ],&k=K-1. cases The empirical ECE (ECE ECE), which evaluates the consistency between reported probabilities and empirical correctness, is then computed as: ECE^=∑k=0K−1|Bk|N|acc(Bk)−conf(Bk)|, ECE= _k=0^K-1 |B_k|N|acc(B_k)-conf(B_k)|, where N denotes the total number of predictions. |Bk||B_k| is the number of predictions whose confidence falls into bin BkB_k. acc(Bk)=1|Bk|∑i∈Bk(y^i=yi)acc(B_k)= 1|B_k| _i∈ B_k 1( y_i=y_i) is the empirical accuracy within bin BkB_k, and conf(Bk)=1|Bk|∑i∈Bkp^iconf(B_k)= 1|B_k| _i∈ B_k p_i is the average confidence in that bin. In our experiments, we set K=10K=10. B.2 Robustness Check on Verbal Confidence Scale Variants We evaluate the robustness of τselfτ^self to confidence-reporting conventions. We keep the setup from Section 3.1 fixed and vary only the confidence format. In particular, we include the [0,20][0,20] scale, which has recently been reported to improve metacognitive efficiency over the standard [0,100][0,100] format [13]. Confidence Format FAR (± 95% CI) Pearson’s r (± 95% CI) Brier Score (± 95% CI) ECE ECE τselfτ^self τavgtoken _avg^token τselfτ^self τavgtoken _avg^token [0,1][0,1] 0.6261 ± 0.0079 0.5397 ± 0.0116 0.3363 ± 0.0059 0.3555 ± 0.0055 0.3825 0.4220 [0,20][0,20] 0.6244 ± 0.0079 0.5851 ± 0.0108 0.3539 ± 0.0059 0.3526 ± 0.0054 0.4050 0.4200 [0,100][0,100] 0.6234 ± 0.0080 0.5523 ± 0.0114 0.3621 ± 0.0056 0.3530 ± 0.0054 0.4136 0.4196 [0%,100%][0\%,100\%] 0.6228 ± 0.0080 0.5449 ± 0.0115 0.3603 ± 0.0057 0.3502 ± 0.0054 0.4125 0.4165 Table 4: Comparison of self-reported confidence formats under prompt ① for GPT-4o mini on PopQA. Across all confidence formats, all metric values remain consistent, indicating that τselfτ^self is robust to confidence reporting conventions. We adopt the [0,1][0,1] format in all experiments. B.3 Validity of Self-Reported Verbal Confidence in Multiple Choice Question Answering We examine the reliability and robustness of self-reported verbal confidence τselfτ^self in multiple-choice question answering (MCQA) settings. Unlike free-response generation—where uncertainty is distributed over an output sequence—uncertainty in MCQA is concentrated on a discrete choice among candidate options. We therefore compare τselfτ^self to the probability of the token corresponding to the selected answer letter, denoted as τkeytokenτ^token_key: τkeytoken=exp(logp(tkey)). _key^token= ( p(t_key) ). B.3.1 Experiment Setup Dataset. We evaluate GPT-4o mini [51] on MMLU-Pro [63], a large-scale multiple-choice benchmark comprising curated questions with 44–1010 answer options, ground-truth labels, and explanations. We use 9 discipline subsets: Biology (717), Computer Science (410), Economics (844), Engineering (969), History (381), Law (1101), Math (1351), Physics (1299), and Psychology (798), totaling 7,870 questions. Confidence Elicitation. Our prompt design follows established prompting practices [63, 65]. A minimal version of the prompt is shown in Appendix B.3.2. Reproducibility Notes. We query gpt-4o-mini-2024-07-18 via the OpenAI Chat Completions API with temperature 0 and token log-probabilities enabled (logprobs:true). For each question, we issue K independent calls (default K=10K=10) to reduce residual output instability from API-level nondeterminism and formatting variation. Across the K responses, the final prediction is determined by plurality vote, while τselfτ^self and τkeytokenτ^token_key are aggregated as the medians of their respective per-call values. The final answer is scored by exact match against the reference label answer. Evaluation Metrics. We evaluate performance using FAR, Brier score, and the empirical Expected Calibration Error (ECE ECE), as defined in Appendix B.1.4. B.3.2 Prompt Template Prompt ① (Biology) You are tasked with answering multiple-choice questions (with answers) about biology. For each question, you must: 1. Think step by step. Provide only the letter corresponding to your chosen option (e.g., “A”, “B”, “C”) and your confidence level within the range [0,1][0,1]. 2. State your confidence level truthfully based on your level of certainty. 3. Confidence values must be expressed to four decimal places for accuracy, without unnecessary rounding (e.g., 0.8765 is acceptable). Examples: Question: Which of the following represents an accurate statement concerning arthropods? Options: A. They possess an exoskeleton composed primarily of peptidoglycan. B. They possess an open circulatory system with a dorsal heart. C. They are members of a biologically unsuccessful phylum incapable of exploiting diverse habitats and nutrition sources. D. They lack paired, jointed appendages. Thinking: Let’s think step by step. Peptidoglycan is known to comprise the plasma membrane of most bacteria, rather than the exoskeleton of arthropods, which is made of chitin, which rules out (A). The answer (C) is false because arthropods are a highly successful phylum. Likewise, arthropods have paired, jointed appendages, which rules out (D). The only remaining option is (B), as arthropods have an open circulatory system with a dorsal tubular heart. The answer is (B). B 0.9500 [Five in-context examples omitted for brevity; see code release for full prompt.] Now, answer the following question: question Options: options B.3.3 Results Calibration and Reliability. We construct five semantics-preserving prompt variants per category. Table LABEL:tab:mcqa_results shows that τselfτ^self and τkeytokenτ^token_key exhibit comparable calibration behavior across domains and prompt templates. In most disciplines, τselfτ^self attains marginally lower Brier scores than τkeytokenτ^token_key, though differences are modest with overlapping 95% CIs. Similarly, the ECE ECE differences between τselfτ^self and τkeytokenτ^token_key remain small, with both signals preserving the same qualitative ECE ECE ranking across domains. The results indicate that τselfτ^self performs comparably to τkeytokenτ^token_key in both probabilistic forecasting and calibration quality. Robustness to Prompt Paraphrasing and Domain Shift. Within each category, FAR, Brier score, and ECE ECE for τselfτ^self have overlapping 95% CIs across templates, demonstrating robustness to paraphrasing. Across categories, most domains exhibit consistent calibration performance. The exceptions—Law and Engineering—show notably higher FAR, Brier score and ECE ECE, a pattern consistent with the calibration degradation under domain shift in QA settings reported by [31]. Category Prompt FAR (± 95% CI) Brier Score (± 95% CI) ECE ECE τselfτ^self τkeytoken _key^token τselfτ^self τkeytoken _key^token Biology ① 0.1951 ± 0.0271 0.1675 ± 0.0231 0.1797 ± 0.0224 0.1351 0.1506 ② 0.1890 ± 0.0266 0.1640 ± 0.0230 0.1730 ± 0.0220 0.1290 0.1358 ③ 0.1994 ± 0.0279 0.1710 ± 0.0237 0.1677 ± 0.0230 0.1390 0.1374 ④ 0.1876 ± 0.0266 0.1629 ± 0.0228 0.1737 ± 0.0239 0.1223 0.1478 ⑤ 0.1913 ± 0.0269 0.1642 ± 0.0230 0.1787 ± 0.0217 0.1293 0.1420 Computer Science ① 0.3297 ± 0.0424 0.2720 ± 0.0350 0.2932 ± 0.0374 0.2654 0.2833 ② 0.3140 ± 0.0419 0.2658 ± 0.0348 0.2784 ± 0.0365 0.2558 0.2601 ③ 0.3192 ± 0.0421 0.2648 ± 0.0343 0.2868 ± 0.0366 0.2518 0.2688 ④ 0.3296 ± 0.0423 0.2711 ± 0.0346 0.3038 ± 0.0391 0.2583 0.3064 ⑤ 0.3219 ± 0.0420 0.2693 ± 0.0346 0.3000 ± 0.0381 0.2463 0.2902 Economics ① 0.2537 ± 0.0095 0.1896 ± 0.0067 0.2346 ± 0.0089 0.1148 0.0977 ② 0.2605 ± 0.0096 0.1950 ± 0.0065 0.2330 ± 0.0082 0.1149 0.1093 ③ 0.2472 ± 0.0099 0.1868 ± 0.0068 0.2173 ± 0.0085 0.1010 0.0979 ④ 0.2608 ± 0.0107 0.1957 ± 0.0073 0.2480 ± 0.0099 0.1238 0.0947 ⑤ 0.2397 ± 0.0099 0.1814 ± 0.0068 0.2205 ± 0.0083 0.0972 0.0894 Engineering ① 0.6100 ± 0.0101 0.4649 ± 0.0081 0.4555 ± 0.0086 0.4919 0.3026 ② 0.5870 ± 0.0118 0.4710 ± 0.0097 0.4565 ± 0.0099 0.4873 0.3298 ③ 0.5803 ± 0.0109 0.4584 ± 0.0087 0.4459 ± 0.0092 0.4752 0.3415 ④ 0.5726 ± 0.0115 0.4490 ± 0.0092 0.4556 ± 0.0099 0.4672 0.3131 ⑤ 0.5776 ± 0.0108 0.4597 ± 0.0088 0.4461 ± 0.0094 0.4754 0.3246 History ① 0.4126 ± 0.0471 0.3183 ± 0.0350 0.3899 ± 0.0425 0.2792 0.4540 ② 0.4304 ± 0.0478 0.3661 ± 0.0403 0.3936 ± 0.0440 0.3518 0.3906 ③ 0.4356 ± 0.0483 0.3598 ± 0.0392 0.3514 ± 0.0386 0.3477 0.3140 ④ 0.4320 ± 0.0482 0.3619 ± 0.0400 0.4031 ± 0.0454 0.3464 0.4009 ⑤ 0.4279 ± 0.0482 0.3599 ± 0.0401 0.3861 ± 0.0433 0.3481 0.3793 Law ① 0.6117 ± 0.0091 0.4589 ± 0.0067 0.5529 ± 0.0085 0.4720 0.2554 ② 0.6098 ± 0.0091 0.4559 ± 0.0066 0.5594 ± 0.0086 0.4683 0.2314 ③ 0.6102 ± 0.0099 0.4296 ± 0.0067 0.4494 ± 0.0080 0.4360 0.4074 ④ 0.6093 ± 0.0126 0.4569 ± 0.0092 0.5723 ± 0.0122 0.4727 0.2459 ⑤ 0.6103 ± 0.0097 0.4558 ± 0.0071 0.5511 ± 0.0091 0.4684 0.2355 Math ① 0.2630 ± 0.0081 0.2251 ± 0.0067 0.2419 ± 0.0071 0.2042 0.1939 ② 0.2596 ± 0.0083 0.2334 ± 0.0070 0.2403 ± 0.0068 0.2046 0.1964 ③ 0.2819 ± 0.0081 0.2431 ± 0.0069 0.2525 ± 0.0067 0.2264 0.1986 ④ 0.2769 ± 0.0083 0.2358 ± 0.0070 0.2587 ± 0.0072 0.2212 0.2123 ⑤ 0.2569 ± 0.0078 0.2241 ± 0.0066 0.2421 ± 0.0067 0.2045 0.1999 Physics ① 0.3730 ± 0.0239 0.2912 ± 0.0187 0.3044 ± 0.0193 0.2876 0.2861 ② 0.3753 ± 0.0240 0.3134 ± 0.0201 0.3085 ± 0.0198 0.3141 0.2944 ③ 0.3730 ± 0.0239 0.2912 ± 0.0187 0.3044 ± 0.0193 0.2988 0.2939 ④ 0.3690 ± 0.0239 0.2942 ± 0.0192 0.3063 ± 0.0202 0.2876 0.3001 ⑤ 0.3730 ± 0.0239 0.2929 ± 0.0190 0.3093 ± 0.0197 0.2940 0.2889 Psychology ① 0.2683 ± 0.0298 0.2237 ± 0.0243 0.2552 ± 0.0279 0.1927 0.2449 ② 0.2590 ± 0.0293 0.2214 ± 0.0246 0.2505 ± 0.0271 0.1875 0.2427 ③ 0.2647 ± 0.0298 0.2185 ± 0.0241 0.2370 ± 0.0245 0.1813 0.1917 ④ 0.2587 ± 0.0292 0.2168 ± 0.0241 0.2440 ± 0.0272 0.1828 0.2368 ⑤ 0.2611 ± 0.0295 0.2197 ± 0.0243 0.2487 ± 0.0272 0.1875 0.2303 Table 5: GPT-4o mini on MMLU-Pro: performance and calibration across subject domains and prompt paraphrases, comparing self-reported confidence τselfτ^self with selected-option token probability τkeytoken _key^token. Appendix C Additional details for Section 4 C.1 Detailed Experiment Setup C.1.1 Prompt Template Prompt for Pure Eval Answer the following question. Format your response as below such that you must begin your response by saying “Answer:”Answer: <your first answer> Question: q Prompt for Scheme B You will get +1 point if your answer is correct, +0.4 if you answer ’I don’t know’, and -1 if incorrect. Give your confidence about your answer in decimal number between 0 and 1, where 0 means you are completely unsure and 1 means you are completely sure. Use four decimal places for precision without unnecessary rounding (e.g., 0.8765). If you truly don’t know, first say ’I don’t know’ and do not give a confidence for that (in that case just leave it blank). If you choose to answer, then provide your answer and your confidence for it. Next, if you said “I don’t know”, then in the same response, give your best possible guess for the answer and your confidence for it. Format your response as below such that you must begin your response by saying “Answer:” first:Answer: <your first answer> Confidence: <0 - 1> If you answered "I don’t know", also include: Best Guess: <your best possible answer> Best Guess Confidence: <0 - 1> Question: q For Scheme A, we remove the clause awarding +0.4 for answering with “I don’t know”. C.1.2 Reproducibility Notes We query gpt-5-mini-2025-08-07 and gpt-4o-mini-2024-07-18 via the OpenAI Chat Completions API with temperature 0. Abstention is detected when the model explicitly outputs “I don’t know”. The model’s final answer is normalized by lowercasing, stripping punctuation, and trimming whitespace, and compared to the PopQA reference set possible_answers. An output is marked correct if the normalized prediction is a substring of any reference string or vice versa. All additional experiments in Appendix C.5 use temperature 0 and the same response format, while API configurations and correctness rules vary slightly by model and dataset and are documented in the anonymous code release. C.1.3 Definitions of Additional Performance Metrics We define the following additional performance metrics to characterize model behavior under selective answering. NansweredN_answered represents the number of questions answered in the first round, excluding cases where the model responds with “I don’t know”. Nincorrect_answeredN_incorrect\_answered denotes the number of questions that are answered incorrectly in the first round. Nincorrect_overallN_incorrect\_overall denotes the number of questions that are incorrect under forced-answer evaluation, whether the error occurs in the first round or in the second round after an initial abstention. Nabstain∩incorrect_overallN_abstain \_overall denotes the number of questions for which the model abstains in the first round and answers incorrectly in the second round. False-Answer Rate (Conditioned on Answering). The false-answer rate among answered questions measures the proportion of answered questions that are incorrect: FARanswered=Nincorrect_answeredNanswered. FAR_answered= N_incorrect\_answeredN_answered. False-Answer Rate (Overall). The overall false-answer rate measures the proportion of all questions that result in an incorrect answer: FARoverall=Nincorrect_overallNtotal. FAR_overall= N_incorrect\_overallN_total. Abstention-to-Error Ratio (AER). The abstention-to-error ratio measures the proportion of incorrectly answered questions that are marked as abstentions in the first round: AER=Nabstain∩incorrect_overallNincorrect_overall.AER= N_abstain \_overallN_incorrect\_overall. That is, among all questions that are incorrect under forced-answer evaluation, AER quantifies how often the model initially signaled uncertainty with “I don’t know.” It is closely related to abstention recall in prior work [16], adapted here to our two-stage prompting protocol. C.2 Robustness Check on Reward Scaling We evaluate reward scaling by multiplying our reward scheme by factors of 10, 100, and 1000. Overall, scaling the magnitude of rewards has relatively little impact on performance, as shown in Table 6 and Table 7. This indicates that model behavior is primarily driven by the relative reward structure rather than its absolute magnitude, and performance is reasonably robust under reward scaling. Reward Penalty Reward Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER (Correct) (Wrong) (Abstain) Answered Overall Answered Overall 1 -1 – A 0.482 ± 0.009 0.548 ± 0.008 -1381 11982 5771 7824 0.262 1 -1 0.4 B 0.410 ± 0.010 0.555 ± 0.008 3577.2 9683 3969 7921 0.499 1 -1 0.4 B w/ norms 0.342 ± 0.011 0.550 ± 0.008 5039.6 7888 2700 7847 0.656 10 -10 – A 0.475 ± 0.009 0.546 ± 0.008 4950 11763 5584 7793 0.283 10 -10 4 B 0.383 ± 0.010 0.570 ± 0.008 41978 9012 3456 8139 0.575 10 -10 4 B w/ norms 0.341 ± 0.011 0.563 ± 0.008 50746 7600 2594 8030 0.677 100 -100 – A 0.486 ± 0.009 0.551 ± 0.008 18900 11923 5797 7854 0.262 100 -100 40 B 0.404 ± 0.010 0.570 ± 0.008 372160 9533 3849 8136 0.527 100 -100 40 B w/ norms 0.328 ± 0.011 0.556 ± 0.008 527060 7067 2320 7934 0.708 1000 -1000 – A 0.486 ± 0.009 0.549 ± 0.008 219000 12133 5896 7836 0.248 1000 -1000 400 B 0.400 ± 0.010 0.577 ± 0.008 3830000 9372 3750 8234 0.545 1000 -1000 400 B w/ norms 0.357 ± 0.010 0.549 ± 0.008 4762800 8176 2915 7829 0.628 Table 6: GPT-5 mini on PopQA (Ntotal=14,267N_total=14,267): performance metrics when the entire reward scheme is scaled by factors of 1010, 100100, and 10001000. Reward Penalty Reward Scheme ECE ECE Brier Score (± 95% CI) (Correct) (Wrong) (Abstain) Answered Overall Answered Overall 1 -1 – A 0.1449 0.1307 0.1999 ± 0.0072 0.1927 ± 0.0067 1 -1 0.4 B 0.1506 0.1206 0.2102 ± 0.0082 0.1933 ± 0.0068 1 -1 0.4 B w/ norms 0.1503 0.1038 0.2061 ± 0.0090 0.1857 ± 0.0065 10 -10 – A 0.1789 0.1577 0.2198 ± 0.0075 0.2098 ± 0.0069 10 -10 4 B 0.1534 0.1092 0.2099 ± 0.0085 0.1981 ± 0.0072 10 -10 4 B w/ norms 0.1691 0.1148 0.2166 ± 0.0094 0.1909 ± 0.0066 100 -100 – A 0.1872 0.1641 0.2213 ± 0.0075 0.2113 ± 0.0069 100 -100 40 B 0.1616 0.1174 0.2147 ± 0.0083 0.2044 ± 0.0074 100 -100 40 B w/ norms 0.1799 0.1066 0.2187 ± 0.0097 0.1941 ± 0.0066 1000 -1000 – A 0.1677 0.1475 0.2122 ± 0.0073 0.2060 ± 0.0069 1000 -1000 400 B 0.1506 0.1090 0.2111 ± 0.0083 0.2007 ± 0.0073 1000 -1000 400 B w/ norms 0.1692 0.1156 0.2146 ± 0.0089 0.1920 ± 0.0066 Table 7: GPT-5 mini on PopQA: calibration metrics when the entire reward scheme is scaled by factors of 1010, 100100, and 10001000. C.3 Full Results for GPT-5 mini and GPT-4o mini on PopQA The full results for GPT-5 mini and GPT-4o mini on PopQA across all reward configurations are shown in Table 8, Table 9, Table 10 and Table 11. Overall, GPT-5 mini achieves a lower FAR, ECE ECE, and Brier score under every reward scheme. We observe that assigning a penalty of -1 for incorrect answers consistently improves performance compared to a penalty of 0. Additionally, incorporating norms into Scheme B leads to better performance than Scheme B in general. Reward Penalty Reward Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER (Correct) (Wrong) (Abstain) Answered Overall Answered Overall Pure Eval 0.523 ± 0.008 – – 13768 7205 ∈[7205,7648]∈[7205,7648] ≤0.058≤ 0.058 1 0 – A 0.491 ± 0.009 0.542 ± 0.008 6528 12403 6089 7739 0.213 1 0 0.2 B 0.444 ± 0.009 0.549 ± 0.008 6709.2 10830 4808 7837 0.386 1 0 0.2 B w/ norms 0.399 ± 0.010 0.550 ± 0.008 6573.4 9280 3703 7844 0.528 1 0 0.4 B 0.445 ± 0.009 0.547 ± 0.008 7380.6 10788 4799 7798 0.385 1 0 0.4 B w/ norms 0.404 ± 0.010 0.553 ± 0.008 7566 9501 3841 7896 0.514 1 0 0.6 B 0.431 ± 0.010 0.542 ± 0.008 8233.4 10377 4474 7732 0.421 1 0 0.6 B w/ norms 0.379 ± 0.010 0.578 ± 0.008 8744.8 8774 3325 8244 0.597 1 0 0.8 B 0.435 ± 0.010 0.551 ± 0.008 8941.6 10515 4575 7863 0.418 1 0 0.8 B w/ norms 0.404 ± 0.010 0.578 ± 0.008 9487.2 9428 3812 8274 0.538 1 -1 – A 0.482 ± 0.009 0.548 ± 0.008 -1381 11982 5771 7824 0.262 1 -1 0.2 B 0.410 ± 0.010 0.558 ± 0.008 2655 9704 3979 7956 0.500 1 -1 0.2 B w/ norms 0.355 ± 0.010 0.575 ± 0.008 3589.8 8253 2933 8210 0.643 1 -1 0.4 B 0.410 ± 0.010 0.555 ± 0.008 3577.2 9683 3969 7921 0.499 1 -1 0.4 B w/ norms 0.342 ± 0.011 0.550 ± 0.008 5039.6 7888 2700 7847 0.656 1 -1 0.6 B 0.412 ± 0.010 0.556 ± 0.008 4477.8 9633 3967 7932 0.500 1 -1 0.6 B w/ norms 0.365 ± 0.010 0.561 ± 0.008 5824.6 8282 3021 8001 0.622 1 -1 0.8 B 0.392 ± 0.010 0.559 ± 0.008 6076.8 9124 3579 7971 0.551 1 -1 0.8 B w/ norms 0.370 ± 0.010 0.581 ± 0.008 6815.6 8510 3150 8294 0.620 Table 8: PopQA (Ntotal=14,267N_total=14,267) evaluation across the full set of reward configurations: performance metrics for GPT-5 mini across Pure Eval and all tested Scheme A, Scheme B, and Scheme B with norms. Reward Penalty Reward Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER (Correct) (Wrong) (Abstain) Answered Overall Answered Overall Pure Eval 0.593±0.0080.593± 0.008 – – 14210 8430 ∈[8430,8483]∈[8430,8483] ≤0.006≤ 0.006 1 0 – A 0.497±0.0100.497± 0.010 0.620±0.0080.620± 0.008 5423 9855 4900 8844 0.446 1 0 0.2 B 0.456±0.0100.456± 0.010 0.623±0.0080.623± 0.008 5854.2 8726 3980 8884 0.552 1 0 0.2 B w/ norms 0.437±0.0110.437± 0.011 0.627±0.0080.627± 0.008 5779 8057 3520 8946 0.607 1 0 0.4 B 0.478±0.0100.478± 0.010 0.625±0.0080.625± 0.008 6821 9202 4399 8924 0.507 1 0 0.4 B w/ norms 0.425±0.0110.425± 0.011 0.632±0.0080.632± 0.008 7040.8 7705 3271 9010 0.637 1 0 0.6 B 0.468±0.0100.468± 0.010 0.625±0.0080.625± 0.008 7948 8963 4196 8911 0.529 1 0 0.6 B w/ norms 0.441±0.0110.441± 0.011 0.632±0.0080.632± 0.008 8229 8032 3542 9023 0.607 1 0 0.8 B 0.460±0.0100.460± 0.010 0.625±0.0080.625± 0.008 9142.6 8733 4016 8923 0.550 1 0 0.8 B w/ norms 0.438±0.0110.438± 0.011 0.631±0.0080.631± 0.008 9508.2 8000 3503 9003 0.611 1 -1 – A 0.492±0.0100.492± 0.010 0.622±0.0080.622± 0.008 147 9649 4736 8876 0.465 1 -1 0.2 B 0.466±0.0100.466± 0.010 0.626±0.0080.626± 0.008 1683.4 8855 4127 8925 0.538 1 -1 0.2 B w/ norms 0.430±0.0110.430± 0.011 0.632±0.0080.632± 0.008 2373.4 7779 3345 9013 0.629 1 -1 0.4 B 0.466±0.0100.466± 0.010 0.626±0.0080.626± 0.008 2766.4 8836 4121 8930 0.539 1 -1 0.4 B w/ norms 0.417±0.0110.417± 0.011 0.631±0.0080.631± 0.008 3906.8 7527 3142 9003 0.651 1 -1 0.6 B 0.464±0.0100.464± 0.010 0.625±0.0080.625± 0.008 3893.8 8821 4095 8910 0.540 1 -1 0.6 B w/ norms 0.428±0.0110.428± 0.011 0.632±0.0080.632± 0.008 5043.4 7721 3301 9014 0.634 1 -1 0.8 B 0.459±0.0100.459± 0.010 0.624±0.0080.624± 0.008 5164.4 8711 3994 8907 0.552 1 -1 0.8 B w/ norms 0.429±0.0110.429± 0.011 0.631±0.0080.631± 0.008 6307.8 7745 3322 9007 0.631 Table 9: PopQA (Ntotal=14,267N_total=14,267) evaluation across the full set of reward configurations: performance metrics for GPT-4o mini across Pure Eval and all tested Scheme A, Scheme B, and Scheme B with norms. Reward Penalty Reward Scheme ECE ECE Brier Score (± 95% CI) (Correct) (Wrong) (Abstain) Answered Overall Answered Overall 1 0 – A 0.1439 0.1326 0.2040 ± 0.0070 0.1965 ± 0.0066 1 0 0.2 B 0.1537 0.1277 0.2120 ± 0.0080 0.1992 ± 0.0068 1 0 0.2 B w/ norms 0.1828 0.1427 0.2257 ± 0.0090 0.2018 ± 0.0070 1 0 0.4 B 0.1733 0.1435 0.2206 ± 0.0078 0.2063 ± 0.0069 1 0 0.4 B w/ norms 0.1926 0.1556 0.2313 ± 0.0085 0.2045 ± 0.0068 1 0 0.6 B 0.1756 0.1484 0.2218 ± 0.0080 0.2019 ± 0.0067 1 0 0.6 B w/ norms 0.1690 0.1198 0.2180 ± 0.0087 0.2030 ± 0.0073 1 0 0.8 B 0.1664 0.1373 0.2200 ± 0.0080 0.2073 ± 0.0070 1 0 0.8 B w/ norms 0.1979 0.1462 0.2336 ± 0.0086 0.2183 ± 0.0076 1 -1 – A 0.1449 0.1307 0.1999 ± 0.0072 0.1927 ± 0.0067 1 -1 0.2 B 0.1540 0.1220 0.2104 ± 0.0081 0.1950 ± 0.0069 1 -1 0.2 B w/ norms 0.1464 0.0957 0.2070 ± 0.0090 0.1930 ± 0.0070 1 -1 0.4 B 0.1506 0.1206 0.2102 ± 0.0082 0.1933 ± 0.0068 1 -1 0.4 B w/ norms 0.1503 0.1038 0.2061 ± 0.0090 0.1857 ± 0.0065 1 -1 0.6 B 0.1572 0.1310 0.2155 ± 0.0083 0.1973 ± 0.0068 1 -1 0.6 B w/ norms 0.1623 0.1216 0.2142 ± 0.0088 0.1915 ± 0.0067 1 -1 0.8 B 0.1552 0.1203 0.2143 ± 0.0084 0.1991 ± 0.0070 1 -1 0.8 B w/ norms 0.1744 0.1305 0.2202 ± 0.0088 0.2024 ± 0.0072 Table 10: PopQA evaluation across the full set of reward configurations: calibration metrics for GPT-5 mini across all tested Scheme A, Scheme B, and Scheme B with norms. Reward Penalty Reward Scheme ECE ECE Brier Score (± 95% CI) (Correct) (Wrong) (Abstain) Answered Overall Answered Overall 1 0 – A 0.4330 0.3876 0.4220±0.00970.4220± 0.0097 0.3488±0.00780.3488± 0.0078 1 0 0.2 B 0.3859 0.3412 0.3820±0.01000.3820± 0.0100 0.3040±0.00800.3040± 0.0080 1 0 0.2 B w/ norms 0.3734 0.3211 0.3705±0.01050.3705± 0.0105 0.2867±0.00740.2867± 0.0074 1 0 0.4 B 0.4053 0.3532 0.3985±0.01000.3985± 0.0100 0.3165±0.00770.3165± 0.0077 1 0 0.4 B w/ norms 0.3606 0.3314 0.3600±0.01070.3600± 0.0107 0.2867±0.00740.2867± 0.0074 1 0 0.6 B 0.3977 0.3554 0.3920±0.01010.3920± 0.0101 0.3144±0.00760.3144± 0.0076 1 0 0.6 B w/ norms 0.3786 0.3570 0.3748±0.01070.3748± 0.0107 0.3086±0.00760.3086± 0.0076 1 0 0.8 B 0.3990 0.3706 0.3951±0.01030.3951± 0.0103 0.3241±0.00770.3241± 0.0077 1 0 0.8 B w/ norms 0.3829 0.3699 0.3798±0.01060.3798± 0.0106 0.3182±0.00760.3182± 0.0076 1 -1 – A 0.4257 0.3798 0.4163±0.00990.4163± 0.0099 0.3406±0.00780.3406± 0.0078 1 -1 0.2 B 0.3990 0.3412 0.3940±0.01020.3940± 0.0102 0.3072±0.00760.3072± 0.0076 1 -1 0.2 B w/ norms 0.3770 0.3324 0.3746±0.01100.3746± 0.0110 0.2914±0.00750.2914± 0.0075 1 -1 0.4 B 0.3958 0.3533 0.3906±0.01020.3906± 0.0102 0.3128±0.00760.3128± 0.0076 1 -1 0.4 B w/ norms 0.3604 0.3310 0.3597±0.01100.3597± 0.0110 0.2860±0.00740.2860± 0.0074 1 -1 0.6 B 0.3913 0.3513 0.3871±0.01020.3871± 0.0102 0.3101±0.00760.3101± 0.0076 1 -1 0.6 B w/ norms 0.3689 0.3452 0.3673±0.01070.3673± 0.0107 0.2968±0.00750.2968± 0.0075 1 -1 0.8 B 0.3984 0.3731 0.3954±0.01030.3954± 0.0103 0.3265±0.00770.3265± 0.0077 1 -1 0.8 B w/ norms 0.3742 0.3513 0.3724±0.01080.3724± 0.0108 0.3035±0.00760.3035± 0.0076 Table 11: PopQA evaluation across the full set of reward configurations: calibration metrics for GPT-4o mini across all tested Scheme A, Scheme B, and Scheme B with norms. We also present the confidence distributions under selected reward schemes for GPT-5 mini and GPT-4o mini in Figures 8(a) and 9(a). Notably, GPT-5 mini produces answers across nearly all confidence levels in the first round. In contrast, GPT-4o mini mostly answers only when its confidence exceeds 0.8, exhibiting a more conservative response pattern. Furthermore, in the best-guess round, GPT-5 mini tends to assign relatively lower confidence to its predictions, whereas GPT-4o mini frequently reports confidence values around 0.5. (a) GPT-5 mini Figure 8: Confidence distributions on PopQA under matched reward settings. For each model, rows correspond to Scheme A, Scheme B, and Scheme B with norms; the left column uses zero penalty for incorrect answers, and the right column uses a −1-1 penalty for incorrect answers. Within each small panel, the left histogram shows first-round answered cases and the right histogram shows second-round best guesses after an initial “I don’t know.” (a) GPT-4o mini Figure 9: Continued. C.4 Rare vs. Common Facts This experiment evaluates how LLMs perform under different reward schemes on common versus rare facts. We sort the PopQA dataset by entity popularity using the field opopo_pop, which records the monthly Wikipedia pageview count of the question’s target entity. The top tercile is designated as common facts and the bottom tercile as rare facts; the middle tercile is omitted from this contrast analysis. We use opopo_pop as a proxy for fact commonness. For GPT-5 mini and GPT-4o mini, Table 12 reports total reward, AER, and FAR, while Table 13 reports calibration metrics. Across all schemes and both models, common facts consistently exhibit lower FAR and better calibration (lower ECE ECE and Brier scores) than rare facts, indicating that LLMs are substantially more reliable on high-popularity knowledge. Within both popularity categories and across both models, Scheme A yields the highest FARansweredFAR_answered, Scheme B improves upon it, and Scheme B with norms performs best overall. This pattern suggests that explicitly rewarding abstention, especially when combined with normative principles, effectively reduces incorrect answers. For GPT-4o mini, Scheme B with norms delivers the strongest calibration across both common and rare facts. For GPT-5 mini, however, no single scheme uniformly dominates across calibration metrics. While Scheme B with norms improves overall Brier scores and achieves the lowest overall ECE ECE on rare facts, Scheme B attains the lowest ECE ECE on common facts, and Scheme A yields the lowest answered-only ECE ECE and Brier scores on rare facts. Comparing models, GPT-5 mini consistently outperforms GPT-4o mini across all schemes and categories, achieving lower answered-only and overall FAR, lower ECE ECE, and lower Brier scores. This indicates that GPT-5 mini demonstrates both stronger reliability and better confidence calibration. Model Scheme Category Total Reward AER FAR (± 95% CI) Answered Overall GPT-5 mini Scheme A (+1,-1) Common 919 0.359 0.331±0.0150.331± 0.015 0.403±0.0140.403± 0.014 Rare -1895 0.243 0.650±0.0150.650± 0.015 0.699±0.0130.699± 0.013 Scheme B (+1,-1,+0.4) Common 2184.6 0.675 0.259±0.0150.259± 0.015 0.403±0.0140.403± 0.014 Rare 169.4 0.514 0.588±0.0180.588± 0.018 0.714±0.0130.714± 0.013 Scheme B w/ norms (+1,-1,+0.4) Common 2428.6 0.945 0.212±0.0150.212± 0.015 0.395±0.0140.395± 0.014 Rare 895.8 0.742 0.525±0.0210.525± 0.021 0.714±0.0130.714± 0.013 GPT-4o mini Scheme A (+1,-1) Common 932 0.566 0.367±0.0160.367± 0.016 0.460±0.0140.460± 0.014 Rare -832 0.534 0.650±0.0180.650± 0.018 0.787±0.0120.787± 0.012 Scheme B (+1,-1,+0.4) Common 1569.2 0.648 0.350±0.0160.350± 0.016 0.465±0.0140.465± 0.014 Rare 324 0.621 0.627±0.0190.627± 0.019 0.794±0.0110.794± 0.011 Scheme B w/ norms (+1,-1,+0.4) Common 1811.6 0.805 0.313±0.0170.313± 0.017 0.470±0.0140.470± 0.014 Rare 867 0.752 0.572±0.0220.572± 0.022 0.801±0.0110.801± 0.011 Table 12: PopQA (Ntotal=14,267N_total=14,267) common-versus-rare fact split: performance metrics for GPT-5 mini and GPT-4o mini under Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). Model Scheme Category ECE ECE Brier Score (± 95% CI) Answered Overall Answered Overall GPT-5 mini Scheme A (+1,-1) Common 0.0817 0.0727 0.1720±0.01160.1720± 0.0116 0.1736±0.01120.1736± 0.0112 Rare 0.2229 0.2008 0.2294±0.01320.2294± 0.0132 0.2139±0.01200.2139± 0.0120 Scheme B (+1,-1,+0.4) Common 0.0742 0.0570 0.1631±0.01230.1631± 0.0123 0.1695±0.01120.1695± 0.0112 Rare 0.2440 0.1957 0.2636±0.01590.2636± 0.0159 0.2141±0.01240.2141± 0.0124 Scheme B w/ norms (+1,-1,+0.4) Common 0.0778 0.0709 0.1501±0.01290.1501± 0.0129 0.1730±0.01110.1730± 0.0111 Rare 0.2550 0.1794 0.2765±0.01860.2765± 0.0186 0.1947±0.01150.1947± 0.0115 GPT-4o mini Scheme A (+1,-1) Common 0.3069 0.2650 0.3115±0.01530.3115± 0.0153 0.2897±0.01300.2897± 0.0130 Rare 0.5759 0.4793 0.5533±0.01850.5533± 0.0185 0.3872±0.01400.3872± 0.0140 Scheme B (+1,-1,+0.4) Common 0.2838 0.2466 0.2924±0.01560.2924± 0.0156 0.2736±0.01280.2736± 0.0128 Rare 0.5516 0.4473 0.5311±0.01980.5311± 0.0198 0.3482±0.01360.3482± 0.0136 Scheme B w/ norms (+1,-1,+0.4) Common 0.2590 0.2304 0.2680±0.01610.2680± 0.0161 0.2567±0.01250.2567± 0.0125 Rare 0.5102 0.4185 0.4999±0.02260.4999± 0.0226 0.3104±0.01330.3104± 0.0133 Table 13: PopQA common-versus-rare fact split: calibration metrics for GPT-5 mini and GPT-4o mini under Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). C.5 Additional Experiments We supplement the main PopQA analysis with additional transfer experiments that vary both model family and dataset difficulty. Specifically, we report results on PopQA, TriviaQA, and SimpleQA Verified, and include GPT-5 mini, GPT-4o mini, Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507. The results reveal a coherent difficulty-dependent pattern. The selective-answering effect remains strongest and most consistent on PopQA, becomes more model-dependent on TriviaQA, and is hardest to realize on SimpleQA Verified, though even there the same confidence-routing mechanism remains visible across all five models. This appendix therefore broadens the transfer evidence for our method and clarifies the conditions under which its gains are strongest: improvements are largest when baseline first-round hallucination risk is substantial and the task still leaves room for selective answering. C.5.1 PopQA On PopQA, the qualitative pattern from GPT-5 mini and GPT-4o mini transfers cleanly to all three additional model families. We use a test set containing 14,267 questions to evaluate the performance of our approach. Scheme B with norms yields the lowest FARansweredFAR_answered and the highest AER for Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507, as shown in Table 14, indicating a consistent shift toward more selective first-round answering. The effect is especially strong for Meta-Llama-3-8B-Instruct, where FARansweredFAR_answered drops from 0.6860.686 under Scheme A to 0.4790.479 under Scheme B with norms and AER rises from 0.1450.145 to 0.7910.791. Gemini and Qwen show the same directional pattern, though with smaller magnitude. As in the main-text PopQA setting, the gain is selective rather than capacity-improving: FARansweredFAR_answered decreases substantially, whereas FARoverallFAR_overall changes much less because abstained questions are still evaluated through forced best guesses. Table 15 suggests that this selective-answering shift is usually accompanied by better confidence quality, especially under Scheme B with norms. Across Llama, Gemini, and Qwen, norms reduce overall ECE ECE and Brier score relative to Scheme A, which is consistent with the model surfacing a cleaner subset of answers. The confidence distributions in Figure 10 support the same interpretation. Relative to Scheme A, Scheme B and especially Scheme B with norms shrink the first-round answered histograms and increase mass in the best-guess channel, whose confidence is typically lower and whose responses are mostly incorrect. This pattern is particularly sharp for Llama and Qwen, where a substantial portion of the first-round mass is removed under norms, but it is also visible for Gemini. Taken together, the additional PopQA results strengthen the main claim that our method improves selective answering by routing many error-prone guesses away from the surfaced-answer channel. Model Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER Answered Overall Answered Overall Meta-Llama-3-8B-Instruct Pure Eval 0.683 ± 0.008 – – 13745 9391 9898 0.051 Scheme A (+1,-1) 0.686 ± 0.008 0.717 ± 0.007 -6275 12735 8739 10227 0.145 Scheme B (+1,-1,+0.4) 0.639 ± 0.010 0.757 ± 0.007 -361.2 8960 5722 10801 0.470 Scheme B w/ norms (+1,-1,+0.4) 0.479 ± 0.014 0.772 ± 0.007 3987.2 4804 2301 11015 0.791 Gemini-3.1-Flash-Lite Pure Eval 0.444 ± 0.008 – – 14253 6322 6331 0.001 Scheme A (+1,-1) 0.387 ± 0.009 0.451 ± 0.008 563 12096 4681 6436 0.273 Scheme B (+1,-1,+0.4) 0.395 ± 0.009 0.463 ± 0.008 3440.6 11943 4716 6599 0.285 Scheme B w/ norms (+1,-1,+0.4) 0.319 ± 0.009 0.456 ± 0.008 5321.4 10011 3196 6505 0.509 Qwen3-4B-Instruct-2507 Pure Eval 0.780 ± 0.007 – – 14165 11055 11155 0.009 Scheme A (+1,-1) 0.732 ± 0.009 0.796 ± 0.007 -9807 8322 6092 11356 0.464 Scheme B (+1,-1,+0.4) 0.719 ± 0.010 0.796 ± 0.007 -651.8 7589 5456 11358 0.520 Scheme B w/ norms (+1,-1,+0.4) 0.696 ± 0.011 0.798 ± 0.007 271.8 6855 4774 11378 0.580 Table 14: PopQA results (Ntotal=14,267N_total=14,267): performance metrics for Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Pure Eval, Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). Model Scheme ECE ECE Brier Score (± 95% CI) Answered Overall Answered Overall Meta-Llama-3-8B-Instruct Scheme A (+1,-1) 0.6646 0.6167 0.6583 ± 0.0083 0.6312 ± 0.0081 Scheme B (+1,-1,+0.4) 0.6175 0.4631 0.6131 ± 0.0100 0.5309 ± 0.0090 Scheme B w/ norms (+1,-1,+0.4) 0.4627 0.3797 0.4617 ± 0.0141 0.3990 ± 0.0089 Gemini-3.1-Flash-Lite Scheme A (+1,-1) 0.3708 0.3679 0.3673 ± 0.0086 0.3620 ± 0.0079 Scheme B (+1,-1,+0.4) 0.3735 0.3592 0.3706 ± 0.0087 0.3542 ± 0.0079 Scheme B w/ norms (+1,-1,+0.4) 0.3078 0.3328 0.3080 ± 0.0090 0.3264 ± 0.0077 Qwen3-4B-Instruct-2507 Scheme A (+1,-1) 0.6446 0.4756 0.6279 ± 0.0104 0.4338 ± 0.0082 Scheme B (+1,-1,+0.4) 0.6297 0.4744 0.6059 ± 0.0110 0.4094 ± 0.0081 Scheme B w/ norms (+1,-1,+0.4) 0.6127 0.4440 0.5933 ± 0.0116 0.3757 ± 0.0080 Table 15: PopQA results: calibration metrics for Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). (a) Meta-Llama-3-8B-Instruct (b) Gemini-3.1-Flash-Lite (c) Qwen3-4B-Instruct-2507 Figure 10: PopQA confidence distributions under Scheme A (+1,-1), Scheme B (+1,-1,+0.4), and Scheme B with norms (+1,-1,+0.4) (top to bottom). Within each panel, the left histogram shows first-round answered cases and the right histogram shows second-round best guesses after an initial “I don’t know.” Rows correspond to Scheme A, Scheme B, and Scheme B with norms. C.5.2 TriviaQA On TriviaQA, transfer is more heterogeneous in a way that appears closely tied to baseline headroom. Because there are no true labels available for the test set, we use the validation set containing 17,944 questions to evaluate the performance of our approach. TriviaQA seems relatively easy for several models: GPT-5 mini, GPT-4o mini, and Gemini-3.1-Flash-Lite already achieve low Pure Eval FARansweredFAR_answered of 0.1450.145, 0.1480.148, and 0.0780.078, respectively, as shown in Table 16. In that regime, there is naturally limited room for further improvement through selective answering, and correspondingly the gains are modest: GPT-5 mini improves slightly under Scheme B with norms, GPT-4o mini remains roughly flat across schemes, and Gemini changes very little. Even in these flatter cases, however, AER still increases from its near-zero Pure Eval baseline, suggesting that the prompt is still eliciting more explicit uncertainty without substantially disrupting already-strong behavior. By contrast, the open-weight models benefit much more. Meta-Llama-3-8B-Instruct drops from 0.4340.434 in Pure Eval to 0.2300.230 under Scheme B with norms, and Qwen3-4B-Instruct-2507 drops from 0.5930.593 to 0.4230.423, with corresponding AER increases from 0.001 to 0.6600.660 and 0.4300.430, as shown in Table 16. The confidence distributions in Figure 11 match this story. For GPT-5 mini, GPT-4o mini, and Gemini-3.1-Flash-Lite, first-round answers are already tightly concentrated near 0.90.9–1.01.0 and the best-guess panels are nearly empty, which leaves little low-confidence mass for the reward scheme to re-route. In contrast, Llama and Qwen show a clearer abstention effect: under Scheme B, and especially under Scheme B with norms, answered coverage shrinks and more low- to mid-confidence mass appears in the best-guess channel. Notably, the remaining answered mass for Llama and Qwen is still sharply peaked near 1.01.0, suggesting that the improvement comes more from withholding a subset of risky cases than from fully re-scaling the confidence distribution. Table 17 shows the same division: GPT-5 mini and Gemini already have very low ECE ECE/Brier scores, so calibration changes are small, whereas Llama and Qwen exhibit clearer improvements under norms. Overall, TriviaQA is supportive of our method, but it also clarifies that the largest gains arise when the baseline model still has substantial first-round hallucination risk to remove. Model Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER Answered Overall Answered Overall GPT-5 mini Pure Eval 0.145 ± 0.005 – – 17917 2592 2607 0.006 Scheme A (+1,-1) 0.139 ± 0.005 0.151 ± 0.005 12346 17599 2454 2706 0.093 Scheme B (+1,-1,+0.4) 0.138 ± 0.005 0.158 ± 0.005 12780.2 17271 2380 2843 0.163 Scheme B w/ norms (+1,-1,+0.4) 0.135 ± 0.005 0.160 ± 0.005 12785.2 17016 2301 2872 0.199 GPT-4o mini Pure Eval 0.148 ± 0.005 – – 17941 2664 2666 0.001 Scheme A (+1,-1) 0.156 ± 0.005 0.165 ± 0.005 11828 17619 2748 2963 0.069 Scheme B (+1,-1,+0.4) 0.154 ± 0.005 0.166 ± 0.005 12293 17523 2698 2979 0.094 Scheme B w/ norms (+1,-1,+0.4) 0.156 ± 0.006 0.172 ± 0.006 11776.8 16830 2631 2977 0.116 Meta-Llama-3-8B-Instruct Pure Eval 0.434 ± 0.007 – – 17936 7793 7797 0.001 Scheme A (+1,-1) 0.265 ± 0.007 0.360 ± 0.007 4972 15583 4125 6468 0.362 Scheme B (+1,-1,+0.4) 0.248 ± 0.007 0.410 ± 0.007 8619.6 13940 3461 7362 0.530 Scheme B w/ norms (+1,-1,+0.4) 0.230 ± 0.008 0.463 ± 0.007 8889.6 12250 2819 8300 0.660 Gemini-3.1-Flash-Lite Pure Eval 0.078 ± 0.004 – – 17942 1401 1401 0 Scheme A (+1,-1) 0.082 ± 0.004 0.085 ± 0.004 14602 17723 1450 1531 0.053 Scheme B (+1,-1,+0.4) 0.082 ± 0.004 0.086 ± 0.004 14898 17734 1460 1544 0.054 Scheme B w/ norms (+1,-1,+0.4) 0.080 ± 0.004 0.085 ± 0.004 14937.8 17597 1399 1527 0.084 Qwen3-4B-Instruct-2507 Pure Eval 0.593 ± 0.007 – – 17933 10640 10649 0.001 Scheme A (+1,-1) 0.453 ± 0.008 0.538 ± 0.007 -2422 14179 6418 9655 0.335 Scheme B (+1,-1,+0.4) 0.447 ± 0.008 0.540 ± 0.007 3092.8 13902 6213 9697 0.359 Scheme B w/ norms (+1,-1,+0.4) 0.423 ± 0.009 0.539 ± 0.007 3974.8 13052 5517 9671 0.43 Table 16: TriviaQA results (NtotalN_total = 17,944): performance metrics for GPT-5 mini, GPT-4o mini, Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Pure Eval, Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). Model Scheme ECE ECE Brier Score (± 95% CI) Answered Overall Answered Overall GPT-5 mini Scheme A (+1,-1) 0.0263 0.0263 0.0863 ± 0.0041 0.0878 ± 0.0042 Scheme B (+1,-1,+0.4) 0.0369 0.0344 0.0873 ± 0.0042 0.0904 ± 0.0042 Scheme B w/ norms (+1,-1,+0.4) 0.0411 0.0404 0.0866 ± 0.0043 0.0924 ± 0.0043 GPT-4o mini Scheme A (+1,-1) 0.1179 0.1181 0.1317 ± 0.0050 0.1311 ± 0.0050 Scheme B (+1,-1,+0.4) 0.1119 0.1125 0.1281 ± 0.0049 0.1296 ± 0.0049 Scheme B w/ norms (+1,-1,+0.4) 0.1220 0.1223 0.1321 ± 0.0051 0.1340 ± 0.0051 Meta-Llama-3-8B-Instruct Scheme A (+1,-1) 0.2467 0.2239 0.2530 ± 0.0068 0.2596 ± 0.0068 Scheme B (+1,-1,+0.4) 0.2288 0.1925 0.2359 ± 0.0071 0.2445 ± 0.0070 Scheme B w/ norms (+1,-1,+0.4) 0.2138 0.1705 0.2197 ± 0.0074 0.2352 ± 0.0073 Gemini-3.1-Flash-Lite Scheme A (+1,-1) 0.0785 0.0787 0.0791 ± 0.0040 0.0802 ± 0.0040 Scheme B (+1,-1,+0.4) 0.0779 0.0782 0.0789 ± 0.0040 0.0801 ± 0.0040 Scheme B w/ norms (+1,-1,+0.4) 0.0764 0.0770 0.0768 ± 0.0039 0.0791 ± 0.0040 Qwen3-4B-Instruct-2507 Scheme A (+1,-1) 0.3866 0.3372 0.3832 ± 0.0081 0.3513 ± 0.0072 Scheme B (+1,-1,+0.4) 0.3837 0.3419 0.3760 ± 0.0081 0.3465 ± 0.0071 Scheme B w/ norms (+1,-1,+0.4) 0.3607 0.3163 0.3562 ± 0.0082 0.3214 ± 0.0070 Table 17: TriviaQA results: calibration metrics for GPT-5 mini, GPT-4o mini, Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). (a) GPT-5 mini (b) GPT-4o mini Figure 11: TriviaQA confidence distributions under Scheme A (+1,-1), Scheme B (+1,-1,+0.4), and Scheme B with norms (+1,-1,+0.4) (top to bottom). Within each panel, the left histogram shows first-round answered cases and the right histogram shows second-round best guesses after an initial “I don’t know.” Rows correspond to Scheme A, Scheme B, and Scheme B with norms. (a) Meta-Llama-3-8B-Instruct (b) Gemini-3.1-Flash-Lite Figure 12: Continued. (a) Qwen3-4B-Instruct-2507 Figure 13: Continued. C.5.3 SimpleQA Verified SimpleQA Verified is by far the hardest setting in our transfer study. This is consistent with the benchmark design. The original SimpleQA benchmark was adversarially collected against GPT-4 responses, and a question was kept only if at least one of several OpenAI model completions was incorrect; SimpleQA Verified then refines that benchmark into a more reliable and challenging evaluation set with 1,000 questions [64, 24]. In our experiments, FARansweredFAR_answered therefore remains high across all models and schemes, so this dataset is best viewed as a stress test of whether the prompt can still induce useful selective answering under severe knowledge pressure. Even in that regime, the results in Table 18 remain directionally encouraging. Each model achieves at least some FARansweredFAR_answered improvement relative to Pure Eval under at least one prompted scheme, with especially visible reductions for GPT-4o mini (0.917→0.8880.917→ 0.888), Meta-Llama-3-8B-Instruct (0.945→0.9040.945→ 0.904), Gemini-3.1-Flash-Lite (0.726→0.6860.726→ 0.686), and Qwen3-4B-Instruct-2507 (0.959→0.9250.959→ 0.925), while GPT-5 mini also improves modestly under Scheme B (0.873→0.8700.873→ 0.870). AER also increases monotonically for all models, indicating that the prompt increasingly causes models to flag uncertainty before eventual forced-answer errors. The confidence distributions in Figure 14 show a consistent selective-answering trend across all five model families. Relative to Scheme A, the abstention-rewarding schemes contract the first-round answered histograms and add more mass to the second-round best-guess histograms, typically at lower confidence. This is especially clear for GPT-4o mini, Meta-Llama-3-8B-Instruct, and Qwen3-4B-Instruct-2507, where Scheme B and Scheme B with norms produce a more prominent low- to mid-confidence best-guess channel while reducing answered coverage. GPT-5 mini is somewhat different in presentation—its first-round answered confidence is already more diffuse—but the same routing effect is still visible. This also helps explain why FARansweredFAR_answered sometimes moves only modestly despite visibly stronger abstention behavior. Since FARanswered=Nincorrect∩answered/NansweredFAR_answered=N_incorrect /N_answered, the metric can be unstable when the number of correctly answered questions is already very small: abstaining on many otherwise incorrect responses reduces the numerator, but a small absolute change in the tiny set of correctly answered questions can substantially perturb the denominator. Qwen3-4B-Instruct-2507 is especially illustrative here. From Scheme A to Scheme B with norms, first-round incorrect answers fall from 455455 to 327327, but first-round correct answers also fall from 3636 to 2626, so FARansweredFAR_answered remains near 0.930.93 even though many error-prone responses are being routed out of the surfaced-answer channel (Table 18). Table 19 is consistent with this interpretation: overall ECE ECE and Brier score often improve more clearly than FARansweredFAR_answered, especially for GPT-4o mini, Meta-Llama-3-8B-Instruct, and Qwen3-4B-Instruct-2507, suggesting that the intervention helps most on the routing decision—i.e., deciding when not to surface an answer—even when it cannot fully repair the underlying knowledge gap. We therefore view SimpleQA Verified as evidence that even in a deliberately difficult setting designed to expose factual weaknesses, the method still elicits more honest uncertainty and suppresses some error-prone first-round answering. Model Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER Answered Overall Answered Overall GPT-5 mini Pure Eval 0.873 ± 0.020 – – 999 872 873 0 Scheme A (+1,-1) 0.889 ± 0.020 0.897 ± 0.018 -804 882 784 897 0.126 Scheme B (+1,-1,+0.4) 0.870 ± 0.022 0.884 ± 0.019 -514.8 802 698 884 0.210 Scheme B w/ norms (+1,-1,+0.4) 0.871 ± 0.023 0.887 ± 0.019 -434 730 636 887 0.283 GPT-4o mini Pure Eval 0.917 ± 0.017 – – 1000 917 917 0 Scheme A (+1,-1) 0.901 ± 0.021 0.916 ± 0.017 -852 749 675 916 0.263 Scheme B (+1,-1,+0.4) 0.888 ± 0.024 0.910 ± 0.017 -385.2 668 593 910 0.348 Scheme B w/ norms (+1,-1,+0.4) 0.894 ± 0.025 0.912 ± 0.017 -261.8 557 498 912 0.454 Meta-Llama-3-8B-Instruct Pure Eval 0.945 ± 0.014 – – 996 941 945 0.004 Scheme A (+1,-1) 0.925 ± 0.019 0.941 ± 0.014 -888 749 693 941 0.264 Scheme B (+1,-1,+0.4) 0.913 ± 0.023 0.944 ± 0.014 -303.6 574 524 944 0.445 Scheme B w/ norms (+1,-1,+0.4) 0.904 ± 0.028 0.950 ± 0.013 -27.6 354 320 950 0.663 Gemini-3.1-Flash-Lite Pure Eval 0.726 ± 0.027 – – 1000 726 726 0 Scheme A (+1,-1) 0.707 ± 0.029 0.721 ± 0.028 -478 892 631 721 0.125 Scheme B (+1,-1,+0.4) 0.686 ± 0.031 0.705 ± 0.028 -268.4 866 594 705 0.157 Scheme B w/ norms (+1,-1,+0.4) 0.698 ± 0.032 0.721 ± 0.028 -255.6 824 575 721 0.202 Qwen3-4B-Instruct-2507 Pure Eval 0.959 ± 0.012 – – 997 956 959 0.003 Scheme A (+1,-1) 0.927 ± 0.022 0.944 ± 0.014 -928 491 455 944 0.518 Scheme B (+1,-1,+0.4) 0.925 ± 0.024 0.939 ± 0.014 -166.2 453 419 939 0.554 Scheme B w/ norms (+1,-1,+0.4) 0.926 ± 0.025 0.944 ± 0.014 -42.2 353 327 944 0.654 Table 18: SimpleQA Verified results (NtotalN_total = 1,000): performance metrics for GPT-5 mini, GPT-4o mini, Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Pure Eval, Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). Model Scheme ECE ECE Brier Score (± 95% CI) Answered Overall Answered Overall GPT-5 mini Scheme A (+1,-1) 0.4830 0.4376 0.3738 ± 0.0321 0.3487 ± 0.0299 Scheme B (+1,-1,+0.4) 0.4854 0.4044 0.3787 ± 0.0333 0.3447 ± 0.0306 Scheme B w/ norms (+1,-1,+0.4) 0.5234 0.4143 0.4164 ± 0.0361 0.3450 ± 0.0305 GPT-4o mini Scheme A (+1,-1) 0.7951 0.6652 0.7232 ± 0.0321 0.5741 ± 0.0305 Scheme B (+1,-1,+0.4) 0.7771 0.6196 0.7070 ± 0.0340 0.5205 ± 0.0305 Scheme B w/ norms (+1,-1,+0.4) 0.7872 0.5686 0.7221 ± 0.0369 0.4722 ± 0.0308 Meta-Llama-3-8B-Instruct Scheme A (+1,-1) 0.8928 0.7262 0.8809 ± 0.0229 0.8175 ± 0.0249 Scheme B (+1,-1,+0.4) 0.8867 0.6075 0.8738 ± 0.0267 0.6873 ± 0.0317 Scheme B w/ norms (+1,-1,+0.4) 0.8760 0.4971 0.8663 ± 0.0346 0.5545 ± 0.0346 Gemini-3.1-Flash-Lite Scheme A (+1,-1) 0.6742 0.6213 0.6587 ± 0.0307 0.6106 ± 0.0304 Scheme B (+1,-1,+0.4) 0.6457 0.5771 0.6299 ± 0.0318 0.5694 ± 0.0306 Scheme B w/ norms (+1,-1,+0.4) 0.6658 0.5851 0.6544 ± 0.0317 0.5804 ± 0.0308 Qwen3-4B-Instruct-2507 Scheme A (+1,-1) 0.8383 0.5239 0.8099 ± 0.0343 0.5170 ± 0.0331 Scheme B (+1,-1,+0.4) 0.8361 0.5309 0.7940 ± 0.0356 0.4707 ± 0.0314 Scheme B w/ norms (+1,-1,+0.4) 0.8397 0.4810 0.8075 ± 0.0411 0.4051 ± 0.0311 Table 19: SimpleQA Verified results: calibration metrics for GPT-5 mini, GPT-4o mini, Meta-Llama-3-8B-Instruct, Gemini-3.1-Flash-Lite, and Qwen3-4B-Instruct-2507 under Scheme A (+1,−1)(+1,-1), Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and Scheme B with norms (+1,−1,+0.4)(+1,-1,+0.4). (a) GPT-5 mini (b) GPT-4o mini Figure 14: SimpleQA Verified confidence distributions under Scheme A (+1,-1), Scheme B (+1,-1,+0.4), and Scheme B with norms (+1,-1,+0.4) (top to bottom). Within each panel, the left histogram shows first-round answered cases and the right histogram shows second-round best guesses after an initial “I don’t know.” Rows correspond to Scheme A, Scheme B, and Scheme B with norms. (a) Meta-Llama-3-8B-Instruct (b) Gemini-3.1-Flash-Lite Figure 15: Continued. (a) Qwen3-4B-Instruct-2507 Figure 16: Continued. C.5.4 Summary Across these transfer experiments, the performance results in Tables 14, 16, and 18 support the same core mechanism as the main PopQA setting. The intervention is most effective when the baseline model still has meaningful first-round hallucination risk and the task leaves room for selective answering: gains are strongest and most consistent on PopQA, more headroom-dependent on TriviaQA, and smaller but still directionally positive on SimpleQA Verified. Even in the hardest setting, all five models show higher AER in Table 18, and the confidence distributions continue to indicate that some risky first-round answers are shifted into a lower-confidence best-guess channel. C.6 Comparison with Closest PopQA Abstention and Hallucination-Mitigation Work The four closest PopQA papers to our setting are InFact [11], [IDK]-token tuning [12], Binary Retrieval-Augmented Reward (Binary RAR) [6], and mechanistic calibration of verbal uncertainty (MUC) [30]. We compare them one by one below, focusing on results obtained using the same model family as in our PopQA experiments when possible. These methods are best viewed as complementary rather than directly competing baselines: InFact, [IDK]-token tuning, and Binary RAR modify model weights, while MUC requires internal activation access. Our goal is therefore mechanism-level positioning rather than direct empirical superiority claims. Because the compared papers sometimes differ in model family, model size, and/or task formulation, the comparisons below should be read as qualitative operating-point comparisons rather than benchmark-equated leaderboard comparisons. One additional distinction is that prior methods are usually reported as discrete learned or steered operating points, whereas I-CALM can move the same fixed model flexibly across a tunable inference-time frontier by varying reward configurations. Metric bridge. Let NtotalN_total denote the number of PopQA questions and NansweredN_answered the number of first-round surfaced answers. Then P=1−FARanswered,C=NansweredNtotal,R=Nanswered,correctNtotal=C(1−FARanswered),P=1-FAR_answered, C= N_answeredN_total, R= N_answered,correctN_total=C (1-FAR_answered ), where P is precision, C is coverage, and R is recall. Note that, under the usual definition of recall TP/(TP+FN)TP/(TP+FN), we have TP=Nanswered,correct,FN=Nanswered,incorrect+Nabstained,TP=N_answered,correct, FN=N_answered,incorrect+N_abstained, because in this setting both a wrong first-round answer and a first-round abstention fail to recover the gold answer, and that is why TP+FN=NtotalTP+FN=N_total. The corresponding surfaced-answer error rate over all questions is Esurf=C⋅FARanswered.E_surf=C·FAR_answered. Thus papers reporting precision/recall/F1 can be compared to (P,R,F1)(P,R,F1), papers reporting hallucination/error rate with abstention allowed can be compared to (Esurf,1−C)(E_surf,1-C), and papers reporting forced-answer accuracy can be compared to 1−FARoverall1-FAR_overall. AER does not have a close direct counterpart in most of this literature. Converted I-CALM PopQA operating points. Under this conversion, GPT-5 mini traces a clear selective-answering frontier rather than a single point: Pure Eval =47.7/46.0/46.8=47.7/46.0/46.8, Scheme A (+1,−1)=51.8/43.5/47.3(+1,-1)=51.8/43.5/47.3, Scheme B (+1,−1,+0.4)=59.0/40.1/47.7(+1,-1,+0.4)=59.0/40.1/47.7, and Scheme B with norms (+1,−1,+0.4)=65.8/36.4/46.8(+1,-1,+0.4)=65.8/36.4/46.8 in precision/recall/F1. GPT-4o mini shows the same pattern at a lower operating range: 40.7/40.5/40.640.7/40.5/40.6 in Pure Eval, 50.9/34.4/41.150.9/34.4/41.1 under Scheme A (+1,−1)(+1,-1), 53.4/33.0/40.853.4/33.0/40.8 under Scheme B (+1,−1,+0.4)(+1,-1,+0.4), and 58.3/30.7/40.258.3/30.7/40.2 under Scheme B with norms. To enable direct comparison, we also report I-CALM performance on the same model families used in prior work. In precision/recall/F1, Meta-Llama-3-8B-Instruct moves from 31.7/30.5/31.131.7/30.5/31.1 in Pure Eval to 52.1/17.5/26.252.1/17.5/26.2 under Scheme B with norms, while Qwen3-4B-Instruct-2507 moves from 22.0/21.8/21.922.0/21.8/21.9 to 30.4/14.6/19.730.4/14.6/19.7. In error–abstention terms, these same-family endpoints correspond to 65.8%/3.7%→16.1%/66.3%65.8\%/3.7\%→ 16.1\%/66.3\% for Llama and 77.5%/0.7%→33.5%/52.0%77.5\%/0.7\%→ 33.5\%/52.0\% for Qwen. Comparison to InFact. InFact is the cleanest direct comparator because it evaluates closed-book open-ended PopQA and reports precision/recall/F1. The family match is not exact—InFact uses Llama-3.1-8B, whereas we use Meta-Llama-3-8B-Instruct. On Llama-3.1-8B, InFact reports PopQA 38.5/38.5/38.538.5/38.5/38.5 for the base model, 51.4/31.9/39.451.4/31.9/39.4 for prompting, 55.5/31.4/40.155.5/31.4/40.1 for semantic entropy, 58.4/33.8/42.858.4/33.8/42.8 for ICL, and 65.5/35.9/46.465.5/35.9/46.4 for full informativeness-alignment. Our Meta-Llama-3-8B-Instruct results move from 31.7/30.5/31.131.7/30.5/31.1 in Pure Eval to 52.1/17.5/26.252.1/17.5/26.2 under Scheme B with norms. Relative to its own Pure Eval baseline, I-CALM yields a clear precision gain, though at a more conservative recall level than InFact’s training-based operating points. This difference is consistent with the intervention budget: InFact learns its operating point through alignment training and weight updates, whereas I-CALM adjusts the operating point of a fixed model at inference time. Comparison to [IDK]-token tuning. [IDK]-token tuning is also directly comparable in precision/recall/F1, but with an important caveat that for PopQA it converts question answering into a GPT-4-rephrased sentence-completion task. On Mistral-7B-v0.1, it reports PopQA 35.5/35.5/35.535.5/35.5/35.5 for the base model, 64.6/20.6/31.264.6/20.6/31.2 for confidence thresholding, 68.7/20.4/31.568.7/20.4/31.5 for semantic entropy, and 78.1/20.5/32.578.1/20.5/32.5 for IDK-tuning. Notably, IDK-tuning occupies a very aggressive precision-oriented point. Relative to the GPT-5 mini and GPT-4o mini I-CALM fronts, this is higher precision at substantially lower recall and F1; viewed that way, I-CALM occupies a more balanced operating region. The same qualitative shift is nevertheless visible in our open-weight results, especially on Meta-Llama-3-8B-Instruct, where precision rises from 31.731.7 to 52.152.1 as recall falls from 30.530.5 to 17.517.5. This comparison supports the broader point that abstention-aware control can reshape the precision–recall trade-off; our contribution is to obtain that shift with a prompt-only black-box intervention rather than continued pretraining. Comparison to Binary RAR. Binary RAR is the closest conceptual comparator because it also changes the reward for answering versus explicitly expressing uncertainty. The cleanest same-family comparison uses Qwen, though Binary RAR uses Qwen3-8B, whereas we use Qwen3-4B-Instruct-2507. In the abstention-allowed setting, Binary RAR reduces PopQA hallucination rate (equivalent to our EsurfE_surf) from 71.271.2 to 26.826.8 and abstains on 55.2%55.2\% of questions; among attempted questions, accuracy rises from 22.3%22.3\% to 40.2%40.2\%. In the forced-response setting, PopQA accuracy is essentially preserved from 20.2%20.2\% to 20.6%20.6\%. Our Qwen results show the same pattern in a smaller model without weight updates: EsurfE_surf drops from 77.5%77.5\% to 33.5%33.5\%, abstention rises from 0.7%0.7\% to 52.0%52.0\%, attempted-answer accuracy rises from 22.0%22.0\% to 30.4%30.4\%, and forced-answer accuracy stays similar from 21.8%21.8\% to 20.2%20.2\%. The comparison therefore highlights a shared mechanism: substantial reduction in surfaced-error accompanied by a large increase in abstention, with little change in forced-answer accuracy. Binary RAR realizes this pattern through online RL with retrieval-verified rewards, whereas I-CALM obtains it as a lightweight prompt-only intervention and can move the same fixed model across a tunable inference-time frontier by changing reward configurations rather than committing to a single learned operating point. Comparison to MUC. MUC is the closest inference-time comparator because it also targets uncertainty expression at test time, albeit with internal activation access. The cleanest same-family comparison again uses Llama. The family match is not exact—Llama-3.1-8B-Instruct versus Meta-Llama-3-8B-Instruct. On Llama-3.1-8B-Instruct PopQA, MUC moves overall hallucination rate from 33.733.7 to 23.223.2, and refusal rate from 42.842.8 to 55.855.8. Under one representative reward configuration, our Meta-Llama-3-8B-Instruct results show the same directional reallocation under prompting alone: surfaced-answer error drops from 65.8%65.8\% to 16.1%16.1\% and abstention rises from 3.7%3.7\% to 66.3%66.3\%. Unlike the single MUC intervention point we compare against here, this I-CALM result is only one point on a tunable frontier, and varying the reward configuration moves the same fixed model to different error–abstention operating points. Takeaway. Taken together, these comparisons position I-CALM as a lightweight black-box complement to training-based or activation-level methods. Across all four papers and ours, the shared mechanism is selective answering: surfaced-answer reliability improves largely by withholding more uncertain answers. What distinguishes I-CALM is that it produces this shift on a fixed base model through prompt-level control and lets the same model trace a tunable operating frontier rather than commit to a single learned or steered endpoint. C.7 Deployment Extension: Post-Hoc Confidence Thresholding with Certified FAR Control The goal of this subsection is to turn the model’s self-reported confidence into a threshold rule that, with high probability over a held-out calibration sample, controls the population FAR among accepted answers at or below a user-specified target. We use the cumulative false-answer rate (CFAR) curve as a descriptive diagnostic, and consider two finite-sample certified rules based on one-sided Clopper–Pearson upper confidence bounds (CP-UCB) [10]: a conservative Bonferroni CP-UCB baseline and a multistart fixed-sequence ordered-testing alternative. This places the subsection in the selective-prediction and risk-controlled refusal literature, where the central deployment question is how much coverage can be obtained at a controlled error level [47, 61, 50]. As in Section 4, each query is paired with a final answer, a correctness label, and a self-reported confidence. When the model initially abstains, the final answer is its elicited best guess; thus this section studies a forced-candidate regime: every query is paired with a candidate answer and confidence score, and the interface decides whether to surface or withhold that candidate. Let u:=1−τself∈[0,1]u:=1-τ^self∈[0,1] denote reported uncertainty. For a fixed uncertainty cutoff u, the system accepts all answers with reported uncertainty at most u, equivalently all answers with confidence at least 1−u1-u. On a calibration set of size n, define n(u)=∑i=1nui≤u,k(u)=∑i=1nui≤uei,n(u)= _i=1^n 1\u_i≤ u\, k(u)= _i=1^n 1\u_i≤ u\e_i, (1) where ei=the final answer is incorrecte_i= 1\the final answer is incorrect\. The empirical cumulative false-answer rate (CFAR) curve is CFAR^(u)=k(u)/n(u),n(u)>0,0,n(u)=0. CFAR(u)= casesk(u)/n(u),&n(u)>0,\\ 0,&n(u)=0. cases Equivalently, CFAR^(u) CFAR(u) is the empirical FAR among answers that would be accepted by the uncertainty threshold u. When plotting CFAR^(u) CFAR(u), we additionally show pointwise 95% confidence intervals as descriptive uncertainty bands. These intervals are not used directly for threshold selection: because they are pointwise and two-sided, their upper endpoints do not provide a valid post-selection guarantee after scanning multiple thresholds. C.7.1 Bonferroni CP-UCB threshold selection To obtain a certified rule, we fix a prespecified finite grid of uncertainty thresholds =0,0.01,…,1.00,U=\0,0.01,…,1.00\, independent of the calibration data, and we set the overall failure probability to δ=0.05δ=0.05. The certification is with respect to this prespecified grid; a finer grid gives finer threshold resolution but induces a slightly more conservative Bonferroni correction. For each u∈u , we compute the one-sided CP-UCB UCBCP(u)=1,n(u)=0 or k(u)=n(u),Beta−1(1−δ/||;k(u)+1,n(u)−k(u)),0≤k(u)<n(u),UCB_CP(u)= cases1,&n(u)=0 or k(u)=n(u),\\ Beta^-1(1-δ/|U|;\,k(u)+1,\,n(u)-k(u)),&0≤ k(u)<n(u), cases (2) which yields a Bonferroni-corrected simultaneous upper band over the entire grid. Beta−1(q;a,b)Beta^-1(q;a,b) denotes the q–quantile of the Beta(a,b)Beta(a,b) distribution. Proposition 1 gives the guarantee. For any target FAR r∈[0,1]r∈[0,1], we then choose the most permissive certified threshold u^r=maxu∈:UCBCP(u)≤r, u_r= \u :UCB_CP(u)≤ r\, whenever the feasible set is nonempty; otherwise the rule rejects all answers. We define the corresponding confidence threshold as τ^r=1−u^r τ_r=1- u_r. At deployment time, the system accepts the model output iff τself≥τ^rτ^self≥ τ_r; otherwise the answer is rejected or held for further review. The CP-UCB calibration procedure is summarized in Algorithm 1 below. Proposition 1 (Finite-sample FAR control via CP-UCB). Let (U,E)(U,E) denote the reported uncertainty and error indicator, respectively, for a fresh example drawn from the same distribution, where E=1E=1 iff the final answer is incorrect. Assume the calibration examples and future examples are i.i.d. draws from a common distribution over (U,E)(U,E). Let =u1,…,uM⊂[0,1]U=\u_1,…,u_M\⊂[0,1] be a prespecified finite grid of uncertainty thresholds, independent of the calibration data. For each calibration example i=1,…,ni=1,…,n, let ui=1−τiselfu_i=1- _i^self be the reported uncertainty and let ei∈0,1e_i∈\0,1\ indicate whether the corresponding final answer is incorrect. For each u∈u , define n(u)=∑i=1nui≤u,k(u)=∑i=1nui≤uei.n(u)= _i=1^n 1\u_i≤ u\, k(u)= _i=1^n 1\u_i≤ u\e_i. Define R(u)=Pr(E=1∣U≤u),R(u)= (E=1 U≤ u), with the convention R(u)=0R(u)=0 when Pr(U≤u)=0 (U≤ u)=0. Define the one-sided Clopper–Pearson upper bound UCBCP(u)=1,n(u)=0 or k(u)=n(u),Beta−1(1−δ/M;k(u)+1,n(u)−k(u)),0≤k(u)<n(u).UCB_CP(u)= cases1,&n(u)=0 or k(u)=n(u),\\ Beta^-1(1-δ/M;\,k(u)+1,\,n(u)-k(u)),&0≤ k(u)<n(u). cases For a target FAR r∈[0,1]r∈[0,1], let u^r=maxu∈:UCBCP(u)≤r u_r= \u :UCB_CP(u)≤ r\, whenever the feasible set is nonempty; otherwise the rule rejects all answers. Then Pr(∀u∈,R(u)≤UCBCP(u))≥1−δ. \! (∀ u ,\ R(u) _CP(u) )≥ 1-δ. Consequently, whenever the feasible set is nonempty, the selected threshold satisfies R(u^r)≤rR( u_r)≤ r. If the feasible set is empty, the reject-all rule trivially satisfies the risk constraint. Proof sketch. For a fixed u, let Ii(u)=Ui≤u.I_i(u)= 1\U_i≤ u\. Since (Ui,Ei)(U_i,E_i) are i.i.d. and Ii(u)I_i(u) depends only on UiU_i, we have Ei∣Ii(u)=1∼Bernoulli(R(u)).E_i I_i(u)=1 (R(u)). Therefore, conditional on there being m accepted examples at threshold u, the corresponding error indicators are i.i.d. Bernoulli(R(u))Bernoulli(R(u)), so the number of accepted errors is distributed as Binomial(m,R(u))Binomial(m,R(u)). The one-sided Clopper–Pearson construction gives Pr(R(u)≤UCBCP(u))≥1−δ/M. \! (R(u) _CP(u) )≥ 1-δ/M. A union bound over the M prespecified grid points yields the simultaneous event with probability at least 1−δ1-δ. The selection rule only chooses thresholds whose certified upper bound is at most r, which yields R(u^r)≤rR( u_r)≤ r whenever the feasible set is nonempty. ∎ Algorithm 1: CP-UCB threshold selection Input: calibration set (ui,ei)i=1n\(u_i,e_i)\_i=1^n, target FAR r∈[0,1]r∈[0,1], overall failure probability δ∈(0,1)δ∈(0,1), prespecified grid =0,0.01,…,1.00U=\0,0.01,…,1.00\. For each u∈u : 1. Compute the number of accepted calibration answers n(u)=∑i=1nui≤un(u)= _i=1^n 1\u_i≤ u\. 2. Compute the number of incorrect accepted calibration answers k(u)=∑i=1nui≤ueik(u)= _i=1^n 1\u_i≤ u\e_i. 3. Set the per-threshold significance level α=δ/||α=δ/|U|. 4. Compute the one-sided CP upper bound UCBCP(u)=1,n(u)=0 or k(u)=n(u),Beta−1(1−α;k(u)+1,n(u)−k(u)),0≤k(u)<n(u).UCB_CP(u)= cases1,&n(u)=0 or k(u)=n(u),\\ Beta^-1(1-α;\,k(u)+1,\,n(u)-k(u)),&0≤ k(u)<n(u). cases Return the calibrated threshold u^r=maxu∈:UCBCP(u)≤r, u_r= \u :UCB_CP(u)≤ r\, if the feasible set is nonempty; otherwise return a reject-all rule. Define the corresponding confidence threshold as τ^r=1−u^r τ_r=1- u_r whenever u^r u_r exists. Deployment rule: accept the model output iff τself≥τ^rτ^self≥ τ_r (equivalently, u≤u^ru≤ u_r); otherwise reject or hold the answer for further review. The guarantee is a population statement obtained entirely from the calibration split; the validation split is used only as a holdout check on how conservative the certified rule is in finite samples. Figure 18 is descriptive only and uses the full PopQA dataset for GPT-5 mini. As u increases, the acceptance set expands to include more uncertain outputs, so the empirical CFAR curves generally rise. The curves are close at very low uncertainty, separate over much of the mid-uncertainty range—where Scheme B with norms is typically lowest, followed by Scheme B and then Scheme A—and reconverge near u=1u=1, where CFAR^(1) CFAR(1) equals the overall forced-answer FAR. This is consistent with Appendix C.3, which shows similar FARoverallFAR_overall across schemes. For actual certification, Figure 18 plots the calibration-split empirical CFAR curves (solid) together with their Bonferroni CP-UCB envelopes (dashed). For the illustrative target r=0.3r=0.3, the selected thresholds are u^0.3A=0.24,u^0.3B=0.25,u^0.3B w/ norms=0.26, u_0.3^A=0.24, u_0.3^B=0.25, u_0.3^B w/ norms=0.26, and the validation markers all lie below the target line. Table 20 extends this comparison to r∈0.1,0.2,0.3,0.4r∈\0.1,0.2,0.3,0.4\. On this split, all validation FAR values among accepted answers fall below their targets. The r=0.1r=0.1 row returns reject-all for all schemes, reflecting the conservativeness induced by the 20% calibration split together with the Bonferroni correction. Under this formulation, the relevant deployment comparison is coverage at fixed risk. Both the conditional risk curve R(u)R(u) and the distribution of reported uncertainty U matter: lowering the CFAR/CP-UCB curve can support a larger certified threshold within a scheme, but cross-scheme coverage also depends on how much mass the prompting intervention shifts across uncertainty levels. In Table 20, Scheme B achieves the highest validation acceptance at each nontrivial risk target, while Scheme B with norms often matches or slightly exceeds the certified threshold without increasing acceptance. The practical takeaway is that prompt-level incentives change the coverage–risk operating point of the confidence filter, so threshold size alone is not a sufficient proxy for deployment utility. Figure 17: Empirical CFAR as a function of reported uncertainty u=1−τselfu=1-τ^self for GPT-5 mini on the full PopQA dataset under Scheme A, Scheme B, and Scheme B with norms. Shaded bands show pointwise 95% confidence intervals. Figure 18: Calibration-split empirical CFAR curves (solid) and Bonferroni CP-UCB envelopes (dashed) for GPT-5 mini on PopQA at the illustrative risk target r=0.3r=0.3. Vertical dashed lines denote the selected uncertainty thresholds u^r u_r; x-shaped markers show the validation FAR among accepted answers at those thresholds. Risk Target Calibrated Uncertainty Threshold Calibration Acceptance Rate Validation Acceptance Rate (± 95% CI) Validation CFAR CFAR (± 95% CI) Scheme A (+1, -1) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 0.08 0.2089 0.1965±0.00730.1965± 0.0073 0.1342±0.01410.1342± 0.0141 0.3 0.24 0.4427 0.4286±0.00910.4286± 0.0091 0.2414±0.01200.2414± 0.0120 0.4 0.48 0.6169 0.5992±0.00900.5992± 0.0090 0.3369±0.01120.3369± 0.0112 Scheme B (+1, -1, +0.4) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 0.12 0.2555 0.2522±0.00800.2522± 0.0080 0.1740±0.01380.1740± 0.0138 0.3 0.25 0.4444 0.4381±0.00910.4381± 0.0091 0.2534±0.01210.2534± 0.0121 0.4 0.63 0.6239 0.6124±0.00890.6124± 0.0089 0.3499±0.01120.3499± 0.0112 Scheme B w/ norms (+1, -1, +0.4) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 0.06 0.1732 0.1643±0.00680.1643± 0.0068 0.1360±0.01550.1360± 0.0155 0.3 0.26 0.4434 0.4284±0.00910.4284± 0.0091 0.2448±0.01200.2448± 0.0120 0.4 0.74 0.6001 0.5846±0.00900.5846± 0.0090 0.3448±0.01140.3448± 0.0114 Table 20: Bonferroni CP-UCB threshold selection for GPT-5 mini on PopQA using a 20%/80% calibration–validation split. Validation FAR is computed among answers accepted by the threshold selected on the calibration split. C.7.2 Multistart fixed-sequence CP threshold selection The Bonferroni baseline in Algorithm 1 permits arbitrary post-selection over the full grid U, but can be conservative when neighboring thresholds have highly correlated accepted sets. In our setting the acceptance sets are nested in u, and the empirical CFAR curves are often approximately nondecreasing: as u increases, the acceptance set expands and the cumulative accepted-set FAR typically rises. This ordered structure motivates a fixed-sequence ordered-testing alternative [2, 68]. Rather than scanning only from the strictest threshold u1u_1, we use a multistart version that prespecifies several starting indices and runs a forward scan from each one. Write the grid in increasing order as =u1<⋯<uMU=\u_1<·s<u_M\, so that u1u_1 is the smallest uncertainty level and therefore the strictest acceptance rule. For a target FAR r∈[0,1]r∈[0,1], define the null hypotheses Hj:R(uj)≥r,j=1,…,M,H_j:R(u_j)≥ r, j=1,…,M, where R(u)=Pr(E=1∣U≤u)R(u)= (E=1 U≤ u) is as in Proposition 1. The ordered-testing idea is to scan thresholds from lower to higher uncertainty and stop along a given path at the first non-rejection. To reduce sensitivity to poorly supported thresholds near u=0u=0, we prespecify a small set of starting indices =j1<⋯<jL⊆1,…,M,J=\j_1<·s<j_L\ \1,…,M\, for example a coarse evenly spaced subset of the grid, and split the overall failure probability equally across starts. From each start jℓ∈j_ , we test the path ujℓ,ujℓ+1,…,uMu_j_ ,u_j_ +1,…,u_M at level δ/Lδ/L, continuing only while rejections occur. For each uju_j, the exact one-sided binomial p-value is pj=1,n(uj)=0,Pr(B≤k(uj)),B∼Binomial(n(uj),r),n(uj)>0.p_j= cases1,&n(u_j)=0,\\[4.0pt] \! (B≤ k(u_j) ), 18.49988ptB (n(u_j),r),&n(u_j)>0. cases (3) where n(uj)n(u_j) and k(uj)k(u_j) are defined in (1). Small pjp_j indicates that the observed number of accepted errors is unusually low under the null R(uj)≥rR(u_j)≥ r, and therefore provides evidence that threshold uju_j is safe. As in the Bonferroni baseline, the same decision can be expressed through one-sided CP upper bounds. Because the one-sided CP interval is the inversion of the exact binomial test, pj≤δ/Lp_j≤δ/L is equivalent to UCBCPMS(uj)=1,n(uj)=0 or k(uj)=n(uj),Beta−1(1−δ/L;k(uj)+1,n(uj)−k(uj)),0≤k(uj)<n(uj).UCB_CP^MS(u_j)= cases1,&n(u_j)=0 or k(u_j)=n(u_j),\\[4.0pt] Beta^-1(1-δ/L;\,k(u_j)+1,\,n(u_j)-k(u_j)),&0≤ k(u_j)<n(u_j). cases (4) satisfying UCBCPMS(uj)≤rUCB_CP^MS(u_j)≤ r. Let Cℓ(r)⊆C_ (r) denote the set of thresholds certified along start jℓj_ . The multistart certified set is then rMS=⋃ℓ=1LCℓ(r),C_r^MS= _ =1^LC_ (r), and the selected deployment threshold is u^rMS=maxrMS, u_r^MS= _r^MS, whenever the certified set is nonempty; otherwise the rule rejects all answers. The corresponding confidence threshold is τ^rMS=1−u^rMS τ_r^MS=1- u_r^MS. Because different starts can certify different forward segments, rMSC_r^MS need not form a single contiguous prefix of the grid. This does not affect validity: the deployment rule depends only on the selected threshold u^rMS u_r^MS, and that selected threshold is itself one of the certified thresholds. Under the idealized monotone-risk picture R(u1)≤⋯≤R(uM)R(u_1)≤·s≤ R(u_M), each start explores the same safe-to-unsafe ordering, but from a different initialization point. Exact monotonicity is not required for validity; it only explains why multistart scans can recover useful thresholds when the earliest grid points are too sparse to support a single forward scan from u1u_1. Proposition 2 (Finite-sample FAR control via multistart fixed-sequence CP). Assume the setting of Proposition 1, and write the prespecified grid as =u1<⋯<uMU=\u_1<·s<u_M\. Let =j1<⋯<jL⊆1,…,MJ=\j_1<·s<j_L\ \1,…,M\ be a prespecified set of starting indices, independent of the calibration data. For a target FAR r∈[0,1]r∈[0,1], define Hj:R(uj)≥rH_j:R(u_j)≥ r, where R(uj)=Pr(E=1∣U≤uj)R(u_j)= (E=1 U≤ u_j). For each start jℓj_ , run fixed-sequence testing on the path Hjℓ,Hjℓ+1,…,HMH_j_ ,H_j_ +1,…,H_M at level δ/Lδ/L, continuing only while rejections occur, and let Cℓ(r)C_ (r) be the set of thresholds certified along that path. Let rMS=⋃ℓ=1LCℓ(r),u^rMS=maxrMSC_r^MS= _ =1^LC_ (r), u_r^MS= _r^MS whenever rMS≠∅C_r^MS≠ , with reject-all otherwise. Then Pr(∀u∈rMS,R(u)≤r)≥1−δ. \! (∀ u _r^MS,\ R(u)≤ r )≥ 1-δ. Consequently, whenever a threshold is selected, R(u^rMS)≤r.R( u_r^MS)≤ r. If no threshold is selected, the reject-all rule trivially satisfies the risk constraint. Proof sketch. Fix a start jℓj_ . If all null hypotheses on that path are false, then every threshold on the path is safe and there can be no false rejection. Otherwise, let jℓ⋆j_ be the index of the first true null on the path Hjℓ,Hjℓ+1,…,HM.H_j_ ,H_j_ +1,…,H_M. For each fixed j, conditional on n(uj)n(u_j), we have k(uj)∣n(uj)∼Binomial(n(uj),R(uj)),k(u_j) n(u_j) (n(u_j),R(u_j)), so pjp_j is a valid one-sided p-value for HjH_j. Under fixed-sequence testing along the ℓ -th path, any false rejection implies rejection of Hjℓ⋆H_j_ . Therefore Pr(any false rejection on path ℓ)≤Pr(pjℓ⋆≤δ/L)≤δ/L. (any false rejection on path )≤ (p_j_ ≤δ/L)≤δ/L. Applying a union bound across the L prespecified starts yields Pr(any false rejection on any path)≤∑ℓ=1Lδ/L=δ. (any false rejection on any path)≤ _ =1^Lδ/L=δ. Hence, with probability at least 1−δ1-δ, every threshold certified on any path is safe, which implies R(u^rMS)≤rR( u_r^MS)≤ r whenever a threshold is selected. ∎ Algorithm 2: Multistart fixed-sequence CP threshold selection Input: calibration set (ui,ei)i=1n\(u_i,e_i)\_i=1^n, target FAR r∈[0,1]r∈[0,1], overall failure probability δ∈(0,1)δ∈(0,1), ordered grid =u1<⋯<uMU=\u_1<·s<u_M\, and prespecified start set =j1<⋯<jL⊆1,…,MJ=\j_1<·s<j_L\ \1,…,M\. Initialize the certified set as ←∅C← . For each start jℓ∈j_ : 1. Set j←jℓj← j_ . 2. While j≤Mj≤ M: (a) Compute the number of accepted calibration answers n(uj)=∑i=1nui≤ujn(u_j)= _i=1^n 1\u_i≤ u_j\. (b) Compute the number of incorrect accepted calibration answers k(uj)=∑i=1nui≤ujeik(u_j)= _i=1^n 1\u_i≤ u_j\e_i. (c) Compute either the exact one-sided p-value pjp_j in (3) or, equivalently, UCBCPMS(uj)UCB_CP^MS(u_j) in (4). (d) If pj≤δ/Lp_j≤δ/L (equivalently, UCBCPMS(uj)≤rUCB_CP^MS(u_j)≤ r), add uju_j to C and set j←j+1j← j+1. (e) Otherwise stop the current path and move to the next start. Return u^rMS=max u_r^MS= if ≠∅C≠ ; otherwise return a reject-all rule. Define the corresponding confidence threshold as τ^rMS=1−u^rMS τ_r^MS=1- u_r^MS whenever u^rMS u_r^MS exists. Deployment rule: accept the model output iff τself≥τ^rMSτ^self≥ τ_r^MS (equivalently, u≤u^rMSu≤ u_r^MS); otherwise reject or hold the answer for further review. Relative to Algorithm 1, the multistart rule trades arbitrary post-selection over the full grid for a family of prespecified ordered scans. It can recover certified thresholds even when the smallest uncertainty levels are too sparsely supported to justify a forward scan from u1u_1. The trade-off is an additional start-wise correction δ/Lδ/L: as L increases, the rule becomes more conservative, although for L≪ML M it is still typically much less conservative than the Bonferroni baseline in Algorithm 1. In practice, a small coarse set of starts (for example, 5–10 roughly evenly spaced grid points) is a reasonable default, and we report one representative choice below. Because the underlying empirical CFAR curves are unchanged from Figures 18–18, we summarize the multistart rule in table form rather than adding a second calibration plot. Table 21 reports results for a representative choice L=10L=10 on the same 20%/80% calibration–validation split as Table 20. This provides a moderate number of starts without making the per-start level δ/Lδ/L too small. On this split, the validation FAR among accepted answers remains below the target in every nontrivial row. Relative to the Bonferroni baseline in Table 20, the multistart rule increases coverage most clearly at stricter risk targets for Scheme A and Scheme B. For example, under Scheme A at r=0.2r=0.2, the validation acceptance rate increases from 0.19650.1965 to 0.23220.2322; under Scheme B at r=0.2r=0.2, it increases from 0.25220.2522 to 0.26970.2697. At r=0.3r=0.3 and r=0.4r=0.4, the gains are smaller but still generally positive. Scheme B with norms benefits less from multistart scanning: at r=0.2r=0.2 it still returns reject-all, and at r=0.4r=0.4 it matches the Bonferroni threshold exactly. The practical takeaway is that multistart scanning can recover additional low-risk coverage when the strictest uncertainty region is statistically thin, but its benefit depends on the shape of the risk curve and on how much calibration mass lies just beyond the earliest grid points. Risk Target Calibrated Uncertainty Threshold Calibration Acceptance Rate Validation Acceptance Rate (± 95% CI) Validation CFAR CFAR (± 95% CI) Scheme A (+1, -1) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 0.11 0.2447 0.2322±0.00770.2322± 0.0077 0.1506±0.01360.1506± 0.0136 0.3 0.25 0.4511 0.4365±0.00910.4365± 0.0091 0.2459±0.01200.2459± 0.0120 0.4 0.56 0.6341 0.6185±0.00890.6185± 0.0089 0.3492±0.01110.3492± 0.0111 Scheme B (+1, -1, +0.4) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 0.13 0.2741 0.2697±0.00810.2697± 0.0081 0.1790±0.01350.1790± 0.0135 0.3 0.26 0.4543 0.4509±0.00910.4509± 0.0091 0.2610±0.01200.2610± 0.0120 0.4 0.67 0.6390 0.6262±0.00890.6262± 0.0089 0.3586±0.01110.3586± 0.0111 Scheme B w/ norms (+1, -1, +0.4) 0.1 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.2 reject-all 0.0000 0.0000±0.00000.0000± 0.0000 – 0.3 0.27 0.4497 0.4331±0.00910.4331± 0.0091 0.2488±0.01210.2488± 0.0121 0.4 0.74 0.6001 0.5846±0.00900.5846± 0.0090 0.3448±0.01140.3448± 0.0114 Table 21: Multistart fixed-sequence CP threshold selection with L=10L=10 for GPT-5 mini on PopQA using a 20%/80% calibration–validation split. Validation FAR is computed among answers accepted by the threshold selected on the calibration split. Appendix D Additional details for Section 5 D.1 Additional Details on the Reward-Only Ablation Figure 19 shows that among the questions answered correctly under full Scheme B, the reward-only ablation abstains on nearly 45% and answers correctly only about 50%. Table 22 shows that even on the 4,250-question subset the reward-only ablation does answer, its performance is weaker: its FAR is 0.243, compared with 0.197 for full Scheme B under the same reward scheme (+1,−1,+0.4)(+1,-1,+0.4). Table 22: Performance comparison on the 4,250-question subset answered by the reward-only ablation, contrasting that ablation with full Scheme B (+1,−1,+0.4)(+1,-1,+0.4). Metric Ablation Scheme B (+1,-1,+0.4) FAR (± 95% CI) 0.243 ± 0.013 0.197 ± 0.012 Figure 19: Among questions answered correctly by full Scheme B (+1,−1,+0.4)(+1,-1,+0.4), this figure shows how the reward-only ablation reallocates them across correct, abstain, and incorrect outcomes. D.2 Additional Ablation Study on Explicit “I Don’t Know” Wording Table 23 isolates a prompt-wording effect around explicit mention of “I don’t know” (IDK). Pure Eval is the simplest direct-QA baseline. A baseline then adds the reward framing and confidence-reporting format used in Scheme A, but does not explicitly tell the model that it may answer “I don’t know.” Scheme A differs from A baseline only in that the prompt explicitly mentions IDK once. Finally, B control serves as a prompt-matched control for Scheme A: it makes the otherwise implicit abstention reward explicit as 0, so the prompt mentions IDK one additional time without changing the intended payoff structure. The cleanest comparison is therefore within each fixed wrong-answer penalty, especially A baseline → A → B control, while Pure Eval, Scheme B, and Scheme B with norms serve as reference points. Viewed this way, A baseline by itself does not produce a consistent improvement over Pure Eval, suggesting that reward framing and confidence reporting alone are not enough to explain the effect. For GPT-5 mini under the (1,0)(1,0) setting, FARansweredFAR_answered is essentially unchanged from Pure Eval to A baseline (0.523 to 0.522), and AER remains very small (at most 0.058 in Pure Eval versus 0.051 in A baseline). Once IDK is explicitly introduced in Scheme A, however, FARansweredFAR_answered drops to 0.491 and AER rises to 0.213. Making the zero abstention reward explicit in B control pushes the same pattern further, with FARansweredFAR_answered decreasing to 0.457 and AER increasing to 0.351. The same qualitative ordering holds for GPT-5 mini under (1,−1)(1,-1): FARansweredFAR_answered moves from 0.524 (A baseline) to 0.482 (A) to 0.444 (B control), while AER rises from 0.050 to 0.262 to 0.395. GPT-4o mini shows the same main first-step effect when IDK is first explicitly mentioned. Under (1,0)(1,0), A baseline has FARanswered=0.613FAR_answered=0.613 and AER =0.039=0.039, whereas Scheme A reduces FARansweredFAR_answered to 0.497 and raises AER to 0.446. Under (1,−1)(1,-1), the same A-baseline-to-A comparison shifts FARansweredFAR_answered from 0.614 to 0.492 and AER from 0.033 to 0.465. The additional A-to-B-control change, though, is much weaker for GPT-4o mini. Overall, Table 23 shows that simply naming “I don’t know” in the prompt can already make the model more cautious, consistent with prior work [53]. At the same time, this wording-only effect does not by itself account for the main results: it is less consistent across models and remains smaller than the larger behavior shift produced by the full method studied in Section 4. We therefore interpret this ablation as evidence that explicit IDK wording contributes to abstention behavior, but is not sufficient to explain the full effect. Reward Penalty Reward Scheme FAR (± 95% CI) Total Reward NansweredN_answered NincorrectN_incorrect AER (Correct) (Wrong) (Abstain) Answered Overall Answered Overall GPT-5 mini – – – Pure Eval 0.523 ± 0.008 0.536 ± 0.008 – 13768 7205 ∈[7205,7648]∈[7205,7648] ≤0.058≤ 0.058 1 0 – A baseline 0.522 ± 0.008 – 6665 13829 7216 7602 0.051 1 0 – A 0.491 ± 0.009 0.542 ± 0.008 6528 12403 6089 7739 0.213 1 0 0 B control 0.457 ± 0.009 0.551 ± 0.008 6065 11165 5100 7859 0.351 1 0 0.4 B 0.445 ± 0.009 0.547 ± 0.008 7380.6 10788 4799 7798 0.385 1 0 0.4 B w/ norms 0.404 ± 0.010 0.553 ± 0.008 7566 9501 3841 7896 0.514 1 -1 – A baseline 0.524 ± 0.008 – -993 13829 7251 7630 0.05 1 -1 – A 0.482 ± 0.009 0.548 ± 0.008 -1381 11982 5771 7824 0.262 1 -1 0 B control 0.444 ± 0.009 0.552 ± 0.008 1210 10737 4762 7874 0.395 1 -1 0.4 B 0.410 ± 0.010 0.555 ± 0.008 3577.2 9683 3969 7921 0.499 1 -1 0.4 B w/ norms 0.342 ± 0.011 0.550 ± 0.008 5039.6 7888 2700 7847 0.656 GPT-4o mini – – – Pure Eval 0.593 ± 0.008 0.595 ± 0.008 – 14210 8430 ∈[8430,8483]∈[8430,8483] ≤0.006≤ 0.006 1 0 – A baseline 0.613 ± 0.008 – 5391 13917 8526 8876 0.039 1 0 – A 0.497 ± 0.010 0.620 ± 0.008 5423 9855 4900 8844 0.446 1 0 0 B control 0.500 ± 0.010 0.623 ± 0.008 4924 9848 4925 8883 0.446 1 0 0.4 B 0.478±0.0100.478± 0.010 0.625±0.0080.625± 0.008 6821 9202 4399 8924 0.507 1 0 0.4 B w/ norms 0.425±0.0110.425± 0.011 0.632±0.0080.632± 0.008 7040.8 7705 3271 9010 0.637 1 -1 – A baseline 0.614 ± 0.008 – -3469 13975 8576 8868 0.033 1 -1 – A 0.492 ± 0.010 0.622 ± 0.008 147 9649 4736 8876 0.465 1 -1 0 B control 0.507 ± 0.010 0.625 ± 0.008 -138 10018 5078 8910 0.43 1 -1 0.4 B 0.466±0.0100.466± 0.010 0.626±0.0080.626± 0.008 2766.4 8836 4121 8930 0.539 1 -1 0.4 B w/ norms 0.417±0.0110.417± 0.011 0.631±0.0080.631± 0.008 3906.8 7527 3142 9003 0.651 Table 23: Effect of explicitly mentioning “I don’t know” in the prompt on abstention behavior for GPT-5 mini and GPT-4o mini on PopQA (Ntotal=14,267N_total=14,267). D.3 Mathematical Model We present a simple mathematical model that establishes the theoretical optimal behavior of the LLM under a reward scheme that encourages abstention. Let instances x∼x , where x is a single question drawn from the question dataset D. A hidden binary validity variable V∈0,1V∈\0,1\ indicates whether the model’s top candidate answer would be correct. Before acting, the model observes internal evidence E and forms a belief 444In the empirical discussion, we use elicited verbal confidence as a noisy proxy for p. p=Pr(V=1∣x,E)∈[0,1].p= (V=1 x,E)∈[0,1]. The model chooses a∈answer,abstaina∈\ answer, abstain\. • If a=answera= answer: the grader awards +R+R if V=1V=1 and −β-β if V=0V=0, with R>0,β≥0R>0,β≥ 0. • If a=abstaina= abstain: the grader assigns a fixed abstention payoff γ≥0γ≥ 0. The expected utilities are Uans(p)=(R+β)p−β,Uabstain=γ.U_ ans(p)=(R+β)p-β, U_ abstain=γ. Proposition 3. The Bayes-optimal policy is a confidence threshold: π⋆(p)=answer,iff Uans(p)≥Uabstain,abstain,otherwise,⟺answer iff p≥τ:=γ+βR+β.π (p)= cases answer,iff U_ ans(p)≥ U_ abstain,\\ abstain,otherwise, cases answer iff p\;≥\;τ:= γ+βR+β. Proof. The Bayes-optimal action maximizes expected utility, so it is optimal to answer if and only if Uans(p)≥Uabstain⟺(R+β)p−β≥γ.U_ ans(p)\;≥\;U_ abstain (R+β)p-β\;≥\;γ. Rearranging, we obtain (R+β)p≥γ+β.(R+β)p\;≥\;γ+β. Since R>0R>0 and β≥0β≥ 0, we have R+β>0R+β>0, so p≥γ+βR+β=:τ.p\;≥\; γ+βR+β\;=:τ. Thus the optimal policy is to choose answer if and only if p≥τp≥τ, and abstain otherwise. In particular, if τ>1τ>1, the optimal policy reduces to always abstain; if τ≤0τ≤ 0, it reduces to always answer. ∎