Paper deep dive
Conformal Tradeoffs: Operational Profiles Beyond Coverage
Petrus H. Zwart
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/20/2026, 10:49:15 PM
Summary
The paper addresses the gap between conformal prediction coverage guarantees and actual deployment performance by introducing the Small-Sample Beta Correction (SSBC) and a Calibrate-and-Audit framework. SSBC maps user requests for coverage and confidence to specific calibration grid points, ensuring finite-sample semantics. The Calibrate-and-Audit method uses an independent audit split to estimate a region-class label table, which serves as a reusable summary for deriving deployment-facing Key Performance Indicators (KPIs) such as commitment frequency and error exposure, allowing for exact Binomial inference and Pareto-relevant tradeoff analysis.
Entities (7)
Relation Signals (6)
Conformal Prediction → provides → Coverage Guarantees
confidence 98% · Conformal prediction gives exact finite-sample coverage guarantees under exchangeability
Calibrate-and-Audit → uses → independent audit split
confidence 96% · Calibrate-and-Audit then fixes the rule by calibration and uses an independent audit split to estimate the induced region–class label table
Calibrate-and-Audit → estimates → Region-Class Label Table
confidence 95% · uses an independent audit split to estimate the induced region–class label table
Small-Sample Beta Correction → provides → finite-sample coverage semantics
confidence 95% · SSBC which gives finite-sample coverage semantics for the deployed rule
Region-Class Label Table → enables → Key Performance Indicators
confidence 93% · a reusable summary from which deployment-facing Key Performance Indicators (KPIs) follow by projection
Small-Sample Beta Correction → inverts → Beta/Beta-Binomial law
confidence 90% · it inverts the Beta/Beta--Binomial law governing calibration-conditional coverage
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Conformal prediction gives exact finite-sample coverage guarantees under exchangeability, but deployed systems are judged by more than coverage alone. For a fixed calibrated rule reused over a finite operational window, stakeholders also care about deployment-facing quantities such as commitment frequency, deferral, and decisive error exposure. These are not determined by coverage: calibration choices with similar coverage can still induce materially different operational profiles. We study this characterization gap in a scoped setting: binary split conformal prediction under exchangeability with a fixed deployed rule. We introduce the Small-Sample Beta Correction (SSBC) which gives finite-sample coverage semantics for the deployed rule: it inverts the Beta/Beta--Binomial law governing calibration-conditional coverage to map a user request $(\alpha^\star,\delta)$ to the least conservative calibration grid point with calibration-conditional PAC semantics for the realized deployed rule. Calibrate-and-Audit then fixes the rule by calibration and uses an independent audit split to estimate the induced region--class label table, a reusable summary from which deployment-facing Key Performance Indicators (KPIs) follow by projection. Under this design, fixed operational rates admit exact finite-sample Binomial inference, while Beta--Binomial envelopes serve as practical predictive summaries for future windows. The induced partition also exposes regime boundaries, Pareto-relevant tradeoffs, and inverse-pricing questions for fixed downstream conventions. Simulations validate the SSBC semantics and compare audit-based summaries with leave-one-out planning proxies; molecular toxicity data provide an audit-based empirical example, and a solubility case study illustrates scenario planning once coverage semantics are fixed.
Tags
Links
- Source: https://arxiv.org/abs/2602.18045v3
- Canonical: https://arxiv.org/abs/2602.18045v3
Trouble viewing inline? Open PDF directly →
Full Text
127,295 characters extracted from source content.
Expand or collapse full text
Conformal Tradeoffs: Operational Profiles Beyond Coverage Petrus H. Zwart PHZwart@lbl.gov Center for Advanced Mathematics in Energy Research Applications, Berkeley Synchrotron Infrared Structural Biology Program, & Molecular Biophysics and Integrated Bioimaging Division, 1 Cyclotron Road, Berkeley, CA 94720, USA Abstract Conformal prediction gives exact finite-sample coverage guarantees under exchangeability, but deployed systems are judged by more than coverage alone. For a fixed calibrated rule reused over a finite operational window, stakeholders also care about deployment-facing quantities such as commitment frequency, deferral, and decisive error exposure. These are not determined by coverage: calibration choices with similar coverage can still induce materially different operational profiles. We study this characterization gap in a scoped setting: binary split conformal prediction under exchangeability with a fixed deployed rule. We introduce the Small-Sample Beta Correction (SSBC) which gives finite-sample coverage semantics for the deployed rule: it inverts the Beta/Beta–Binomial law governing calibration-conditional coverage to map a user request (α⋆,δ)(α ,δ) to the least conservative calibration grid point with calibration-conditional PAC semantics for the realized deployed rule. Calibrate-and-Audit then fixes the rule by calibration and uses an independent audit split to estimate the induced region–class label table, a reusable summary from which deployment-facing Key Performance Indicators (KPIs) follow by projection. Under this design, fixed operational rates admit exact finite-sample Binomial inference, while Beta–Binomial envelopes serve as practical predictive summaries for future windows. The induced partition also exposes regime boundaries, Pareto-relevant tradeoffs, and inverse-pricing questions for fixed downstream conventions. Simulations validate the SSBC semantics and compare audit-based summaries with leave-one-out planning proxies; molecular toxicity data provide an audit-based empirical example, and a solubility case study illustrates scenario planning once coverage semantics are fixed. 11footnotetext: Generative AI was used to assist in the preparation and formatting of this document. 1 Introduction: deployment-facing conformal prediction Conformal prediction provides exact finite-sample coverage guarantees under exchangeability, but deployment decisions are not made from coverage alone (Vovk et al., 2005; Shafer and Vovk, 2008; Angelopoulos and Bates, 2023). Once a classifier is calibrated and reused over a finite operational window, stakeholders also care about commitment, deferral, and decisive-error exposure. Those quantities affect throughput and risk, yet they are not determined by marginal coverage. We study this gap for binary split conformal prediction with a fixed deployed rule. Calibration τ fixes thresholds on score space X, the thresholds induce a finite region partition Rτ(X)R_τ(X) and associated fixed labels Y, and the deployed system is observed through the joint process (Rτ(X),Y)(R_τ(X),Y). Our central claim is that this region–class structure is the deployment object of interest: it explains why coverage-matched rules can still behave differently in practice, and it provides a reusable summary from which many operational KPIs follow by projection. Our method has two layers. The Small-Sample Beta Correction (SSBC) maps a user request (α⋆,δ)(α ,δ) to the least conservative conformal grid point with the intended finite-sample coverage semantics for the realized deployed rule. Calibrate–and–Audit then freezes that rule and uses an independent audit split to estimate deployment-facing rates. Under this design, fixed KPIs admit exact Binomial inference, while Beta–Binomial envelopes provide practical planning summaries for future windows. The exact finite-sample guarantees in the paper are deliberately narrow: they apply to SSBC’s coverage semantics for a fixed deployed rule and to audit-based inference for pre-declared KPIs at that same fixed rule. By contrast, calibration sweeps, Pareto views, leave-one-out proxies, and inverse-pricing analyses are planning tools rather than simultaneous certification statements. Figure 1: Operational view of a calibrated conformal rule. Calibration fixes a region partition; auditing the resulting region–class table yields a reusable operational summary; sweeping calibration settings traces feasible trade-offs. 1.1 Contributions In the scoped setting of binary split conformal prediction under exchangeability, the paper makes three contributions. (1) Operational structure beyond coverage. The calibration-induced partition constrains deployment behavior: changing thresholds reallocates mass across a small set of region types, so operational KPIs are coupled rather than independently tunable. This yields attainable sets, Pareto-relevant regimes, and a natural interface for asking which downstream action conventions are justified by the audited region-wise label composition. (2) SSBC as a coverage-semantics anchor. SSBC inverts the exact rank/Beta law to map (α⋆,δ)(α ,δ) to the least conservative deployed grid point whose realized coverage satisfies a calibration-conditional PAC-style statement. It turns a nominal request into an explicit finite-sample semantic claim about the realized conformal rule. (3) Calibrate–and–Audit for deployment KPIs. The region–class label table is the reusable audit object for a fixed rule. From it one can derive commitment, deferral, singleton-error, and related KPI rates by projection. Exact Binomial inference applies pointwise to fixed audit-based KPIs, while leave-one-out proxies are used only for exploratory planning when an audit split is unavailable. 1.2 Relation to prior work Conformal prediction constructs set-valued predictors with finite-sample coverage guarantees under exchangeability (Vovk et al., 2005; Shafer and Vovk, 2008; Angelopoulos and Bates, 2023). Conformal risk control (CRC) extends this perspective by selecting thresholds to satisfy user-specified scalar risk constraints under related assumptions (Angelopoulos et al., 2024; Bates et al., 2021), and other work studies stronger conditional notions of validity or connects conformal methods to downstream decision-making (Tibshirani et al., 2019; Fannjiang et al., 2022; Bian and Barber, 2023; Gibbs et al., 2025; Lekeufack et al., 2023; Kiyani et al., 2025). These studies typically ask how to guarantee coverage, a chosen scalar risk functional, or decision quality under a specified objective. Our emphasis is different: once a conformal rule is fixed, what operational behavior should one expect from it, and how much of that behavior is already determined by calibration geometry rather than by coverage alone? Recent Human-Computer Interaction (HCI)-oriented work raises a related deployment concern: conformal sets can be a “murky” interface when validity summaries do not translate cleanly into action-relevant consequences (Hullman et al., 2025). Our region–class label table view puts that issue on a concrete footing: it audits how mass falls across region types and labels, exposing how often a rule commits, hedges, or abstains, and where decisive errors arise. This makes clear how coverage-matched rules can still behave differently in deployment. We do not replace validity guarantees; we clarify what they do and do not imply for deployment behavior. Once constructed, the table supports many KPIs and policy projections without re-running calibration. Closest in spirit to our viewpoint is the inverse CRC formulation of Zhou and Zhu (2025), who trace certified miscoverage–regret trade-offs for predict-then-optimize pipelines, and related work on conformal efficiency optimizes set-size functionals subject to validity constraints (Yang and Kuchibhotla, 2021). Our objective is different: rather than introducing another scalar criterion, we make the deployment interface itself the primitive. A calibration choice θ fixes a finite conformal partition, and the induced region–class label table is the minimal reusable summary that determines many operational KPIs by linear projection under any fixed region-based policy; this yields an operational profile (rate-vector) map θ↦(θ)θ (θ) and an attainable set of behaviors that is geometrically constrained, so coverage-matched rules can still induce qualitatively different deployment regimes. Methodologically, we enforce a disciplined two-stage workflow: calibration sweeps and KPI estimates on the audit split, or LOO surrogates when an audit split is unavailable, are used only to expose feasible regimes and trade-offs, while exact finite-sample guarantees are asserted pointwise for pre-declared KPIs at a fixed deployed rule using an independent audit split via Binomial inference (with Beta–Binomial envelopes as planning summaries for finite windows). SSBC plays a complementary role as a semantics anchor, mapping a user request (α⋆,δ)(α ,δ) to a concrete calibration grid point so the operational map is navigated relative to an explicit finite-sample coverage semantics for the realized rule. 1.3 Roadmap The remainder of the paper is organized as follows: Section 2 defines the fixed deployed object and the region–class label table, Section 3 presents SSBC and Calibrate–and–Audit, Section 4 gives the binary geometric interpretation, Section 5 reports simulation, molecular toxicology, and solubility scenario planning, and Section 6 concludes. Technical derivations and extended examples are deferred to the appendices. 2 Setting and notation: calibration-conditional viewpoint We study a single deployed classifier: a scoring model is trained once, treated as fixed, and then calibrated once by split conformal prediction. Randomness therefore comes from the calibration draw and future deployment examples, not from retraining or repeated recalibration. To evaluate deployment-facing quantities beyond coverage, we reserve an exchangeable audit split that is never used to choose thresholds. 2.1 Data splits and exchangeability assumptions Let trainD_train denote the training data used to fit a base score model. After training, the score function is treated as fixed. Let cal=(Xic,Yic)i=1ncal,audit=(Xia,Yia)i=1naudit.D_cal=\(X_i^c,Y_i^c)\_i=1^n_cal, _audit=\(X_i^a,Y_i^a)\_i=1^n_audit. The calibration split calD_cal is used once to set the deployed thresholds, while auditD_audit is reserved for post-calibration evaluation of fixed-rule KPIs. We assume calibration, audit, and future deployment points are jointly exchangeable conditional on the fixed trained score model. We focus on binary classification because it exposes the geometry and deployment trade-offs most cleanly. 2.2 Scores and calibration thresholds Let =1,…,KY=\1,…,K\, and let s:×→ℝs:X×Y be a nonconformity score (Shafer and Vovk, 2008; Lei et al., 2018). In the experiments we use s(x,y)=1−P(y∣x)s(x,y)=1-P(y x), but the setup below does not depend on that choice. Given calD_cal, compute calibration scores Si:=s(Xic,Yic),i=1,…,ncal,S_i:=s(X_i^c,Y_i^c), i=1,…,n_cal, and let S(1)≤⋯≤S(ncal)S_(1)≤·s≤ S_(n_cal) denote their order statistics. Split conformal calibration selects τ:=S(k),τ:=S_(k), where k∈1,…,ncalk∈\1,…,n_cal\ is determined by the requested miscoverage level. Equivalently, one may index the same grid by u:=ncal+1−ku:=n_cal+1-k, so the deployed grid level is αgrid=u/(ncal+1) _grid=u/(n_cal+1). We use class-conditional split conformal (Vovk, 2012b; a), so thresholds τy _y are computed separately within each class. For a fixed calibration draw, this yields a threshold vector τ=(τ1,…,τK)τ=( _1,…, _K) that remains fixed during deployment. 2.3 Regions, observability, and policies Once thresholds are fixed, they induce a finite region label. Writing s(x):=(s(x,1),…,s(x,K)),τ:=(τ1,…,τK),s(x):=(s(x,1),…,s(x,K)), τ:=( _1,…, _K), the region map is Rτ(x):=(s(x,1)≤τ1,…,s(x,K)≤τK)∈0,1K.R_τ(x):= (1\s(x,1)≤ _1\,…,1\s(x,K)≤ _K\ )∈\0,1\^K. Thus calibration fixes a finite partition of score space. The partition itself depends only on τ, while the amount of mass carried by each region depends on the deployment distribution. This partition is the paper’s operational interface (Figure 2). Figure 2: Thresholds induce a finite region partition. In the binary probability-normalized setting, the score support determines which regions can carry mass under a fixed pair of thresholds. Deployment is observed through the joint process (Rτ(X),Y)(R_τ(X),Y). To make downstream actions explicit, we also allow a fixed deployment policy π:ℛ→2π:R→ 2^Y and write C^π(x):=π(Rτ(x)). C_π(x):=π(R_τ(x)). The standard conformal predictor is the set-inclusion policy πSI(r):=y∈:ry=1,C^πSI(x)=y∈:s(x,y)≤τy. _SI(r):=\y :r_y=1\, C_ _SI(x)=\y :s(x,y)≤ _y\. As an example beyond set inclusion, consider a region-triggered commit policy: in the binary case, choose a trigger set A⊆0,12A \0,1\^2 and an action a∈0,1a∈\0,1\, then commit to a\a\ when Rτ(x)∈AR_τ(x)∈ A and abstain otherwise. This simple template makes the calibration/action separation concrete: calibration fixes the region partition, while policy chooses how to act on the realized region. Additional binary projection masks can be found in Appendix G. Figure 3: Deployment policies as projections on a fixed region structure. Different region-based policies act on the same calibrated partition but induce different reported outputs. 2.4 Auditing primitives: region–class label tables and policy projections Conditional on calD_cal, the deployed rule is fixed, and the primitive observable for auditing is the pair (Rτ(θ)(X),Y)(R_τ(θ)(X),Y) on an exchangeable labeled split. Here θ∈Θθ∈ indexes thresholds τ(θ)τ(θ) and hence the region map Rτ(θ)R_τ(θ). Two linked objects will be used throughout the paper. (A) Region–class label joint table. Define the calibration-conditional joint probabilities pr,y(θ):=ℙ(Rτ(θ)(X)=r,Y=y∣cal),(r,y)∈ℛ×.p_r,y(θ):=P\! (R_τ(θ)(X)=r,\ Y=y _cal ), (r,y) ×Y. These are fixed but unknown constants for the deployed rule. The audit split estimates them through Kr,yaudit(θ):=∑i=1nauditRτ(θ)(Xia)=r,Yia=y,p^r,yaudit(θ)=Kr,yaudit(θ)naudit.K^audit_r,y(θ):= _i=1^n_audit1\R_τ(θ)(X_i^a)=r,\ Y_i^a=y\, p^\,audit_r,y(θ)= K^audit_r,y(θ)n_audit. Because the cells (r,y)(r,y) form a finite partition, the table sums to one and changing θ reallocates mass across regions. Those conservation constraints drive the geometric coupling studied later. (B) Policy-specific indicators. Fix a deployment policy π. Many KPIs of interest can be written as a Bernoulli indicator Iℓ:=gℓ(Rτ(θ)(X),Y;π)∈0,1,pℓ(θ):=ℙ(Iℓ=1∣cal).I_ :=g_ (R_τ(θ)(X),Y;π)∈\0,1\, p_ (θ):=P\! (I_ =1 _cal ). For example, gabs(r,y;π)=π(r)=∅g_abs(r,y;π)=1\π(r)= \ and gerr(r,y;π)=|π(r)|=1,y∉π(r)g_err(r,y;π)=1\|π(r)|=1,\ y∉π(r)\. The corresponding audit count is Kℓaudit(θ):=∑i=1nauditgℓ(Rτ(θ)(Xia),Yia;π),p^ℓaudit(θ)=Kℓaudit(θ)naudit.K^audit_ (θ):= _i=1^n_auditg_ (R_τ(θ)(X_i^a),Y_i^a;π), p^\,audit_ (θ)= K^audit_ (θ)n_audit. Thus the region–class label table is the reusable audit object, and KPI rates are its projections. Projection identity. The table and KPI views are linked by linearity: Kℓaudit(θ)=∑r∈ℛ∑y∈gℓ(r,y;π)Kr,yaudit(θ),p^ℓaudit(θ)=Kℓaudit(θ)naudit.K^audit_ (θ)= _r _y g_ (r,y;π)\,K^audit_r,y(θ), p^\,audit_ (θ)= K^audit_ (θ)n_audit. (1) Taking expectation gives pℓ(θ)=∑r∈ℛ∑y∈gℓ(r,y;π)pr,y(θ).p_ (θ)= _r _y g_ (r,y;π)\,p_r,y(θ). In words: estimate the table once, then obtain policy-level KPIs by projection. Worked example: coverage as a projection of the region–class label table. Under set inclusion, coverage is the sum of the region–class label cells where the true label lies in the predicted set. Consider =0,1Y=\0,1\ and the set-inclusion policy πSI _SI. For fixed thresholds, each input falls into one of four regions ℛ=r10,r11,r01,r00R=\r_10,r_11,r_01,r_00\, and the primitive auditing object is (Rτ(X),Y)∈ℛ×(R_τ(X),Y) ×Y. The corresponding region–class label table is P(θ):y=0y=1C^πSI(X)r10=(1,0),(θ)p10,1(θ)0r11=(1,1),(θ),(θ)0,1r01=(0,1)p01,0(θ),(θ)1r00=(0,0)p00,0(θ)p00,1(θ)∅with∑r∈ℛ∑y∈pr,y(θ)=1.P(θ): array[]c|c&y=0&y=1& C_ _SI(X)\\ r_10=(1,0)&p_10,0(θ)&p_10,1(θ)&\0\\\ r_11=(1,1)&p_11,0(θ)&p_11,1(θ)&\0,1\\\ r_01=(0,1)&p_01,0(θ)&p_01,1(θ)&\1\\\ r_00=(0,0)&p_00,0(θ)&p_00,1(θ)& array _r _y p_r,y(θ)=1. Under πSI _SI, the event Y∈C^πSI(X)\Y∈ C_ _SI(X)\ corresponds to the four bold cells (r10,0)(r_10,0), (r11,0)(r_11,0), (r11,1)(r_11,1), and (r01,1)(r_01,1), hence pcov(θ)=p10,0(θ)+p11,0(θ)+p11,1(θ)+p01,1(θ)=1−(p10,1(θ)+p01,0(θ)+p00,0(θ)+p00,1(θ)).p_cov(θ)=p_10,0(θ)+p_11,0(θ)+p_11,1(θ)+p_01,1(θ)=1- (p_10,1(θ)+p_01,0(θ)+p_00,0(θ)+p_00,1(θ) ). The calibration choices (α0⋆,α1⋆)( _0 , _1 ) determine the class-conditional miscoverage primitives p10,1(θ)+p00,1(θ)p_10,1(θ)+p_00,1(θ) for Y=1Y=1 and p01,0(θ)+p00,0(θ)p_01,0(θ)+p_00,0(θ) for Y=0Y=0, while the decomposition of these totals into wrong-singleton errors versus abstentions is dictated by the deployment distribution. This makes explicit that (i) coverage is a projection, (i) once thresholds are fixed it is determined by how mass is redistributed across regions, and (i) deployment-facing consequences are not determined by the calibration targets alone. The same “sum selected cells” structure applies to abstention/deferral, decisiveness, and decisive error exposure under any fixed policy. Appendix G records additional binary projection masks and related bookkeeping. 3 Operational quantities as first-class objects Deployment behavior is summarized by operational rates induced by a calibrated partition together with a fixed deployment policy π. The purpose of this section is to separate what is certifiable for a fixed deployed rule from what is useful only for planning across candidate rules. We first fix the coverage semantics of the deployed rule through SSBC, then turn to audit-based inference and comparative planning at that fixed rule. We keep three layers distinct: geometry (the induced regions), policy (the projection from regions to outputs), and rates (the auditable deployment frequencies). 3.1 Calibrate–and–Audit Beyond marginal coverage, finite-sample evaluation of operational rates requires an independent audit split. If calD_cal is reused both to choose thresholds and to evaluate downstream indicators, the resulting event counts are no longer Bernoulli draws conditional on a fixed rule. Calibrate–and–Audit avoids this coupling by freezing thresholds on calD_cal and evaluating operational quantities on auditD_audit only. When an audit split is unavailable, we use leave-one-out (LOO) recalibration as a practical proxy. Throughout the paper, that proxy is used only for exploration and scenario planning, not for the exact fixed-rule guarantees. The structural issue is that calibration reuse makes the threshold random with respect to the same observations used for evaluation, inducing dependence between selection and KPI counts rather than conditionally i.i.d. Bernoulli trials. Appendix D gives the covariance argument and the details of the LOO construction. 3.2 SSBC as a calibration navigation coordinate We use coverage as the semantic anchor for navigating the calibration grid. Split conformal calibration lives on a finite grid indexed by u∈1,…,ncalu∈\1,…,n_cal\ with αgrid=u/(ncal+1) _grid=u/(n_cal+1); see Section 2.2. SSBC maps (α⋆,δ)(α ,δ) to the least conservative admissible index u⋆(α⋆,δ)u (α ,δ) satisfying ℙcal(ℙ(Y∈C^(X)∣cal)≥1−α⋆)≥1−δ.P_D_cal\! (P\! (Y∈ C(X) _cal )≥ 1-α )≥ 1-δ. Thus SSBC turns a semantic request into a concrete deployed threshold choice. In the binary class-conditional setting this yields (u0⋆,u1⋆)=(u⋆(α0⋆,δ0),u⋆(α1⋆,δ1)),(u_0 ,u_1 )= (u ( _0 , _0),\,u ( _1 , _1) ), which determine (τ0,τ1)( _0, _1) and therefore the calibration setting θ. Although the user-facing request is four-dimensional, (α0⋆,δ0,α1⋆,δ1)( _0 , _0, _1 , _1), calibration collapses it to a low-dimensional navigation coordinate on the deployed grid. The later planning objects are therefore indexed by a coverage-semantically meaningful choice of deployed rule rather than by an abstract threshold sweep. 3.3 Audit-based predictive envelopes for future windows Under Calibrate–and–Audit, conditional on calD_cal the deployed rule is fixed and audit set is exchangeable with future deployment points. Therefore, for any fixed KPI indicator IℓI_ , Kℓaudit(θ)∣cal∼Binomial(naudit,pℓ(θ)),K^audit_ (θ) _cal \! (n_audit,p_ (θ) ), and for a future window of size m, Kℓm(θ)∣cal∼Binomial(m,pℓ(θ)).K^m_ (θ) _cal \! (m,p_ (θ) ). These Binomial laws support exact fixed-rule inference for the latent rate pℓ(θ)p_ (θ); for example, Clopper–Pearson intervals (Clopper and Pearson, 1934) are valid from the audit count Kℓaudit(θ)K^audit_ (θ). For finite future windows we use Beta–Binomial envelopes as practical planning summaries (Skellam, 1948; Johnson et al., 1997; Gelman et al., 2013). 3.4 Rate vectors, attainable operational sets, and Pareto filtering The guarantees above are pointwise: they certify a pre-declared KPI at a fixed deployed rule, not an adaptively selected member of a sweep. Planning is comparative. For a calibration family Θ , a fixed policy π, and a KPI list, each setting θ∈Θθ∈ determines an operational profile (θ)=(p1(θ),…,pL(θ)),p(θ)= (p_1(θ),…,p_L(θ) ), whose coordinates are obtained by projection from the same region–class label table via (1). Sweeping θ traces the attainable set (Θ;π):=(θ):θ∈Θ⊂ℝL.V( ;π):=\p(θ):θ∈ \ ^L. Because the underlying table obeys conservation constraints, changing θ reallocates mass across regions rather than tuning rates independently. To summarize planning-relevant regimes without committing to a scalar objective, we use an oriented Pareto filter. After declaring which KPIs are preferable to increase or decrease, a setting is nondominated if no other setting is at least as good in every oriented coordinate and strictly better in one. These fronts are planning objects; exact finite-sample guarantees remain pointwise for fixed settings and fixed KPIs. Section 4 now makes precise why these attainable sets are geometrically constrained and why sweeping SSBC-indexed settings produces coupled, regime-dependent trade-offs. Any later scalar objective that is monotone in the chosen orientation must attain its optimum on this front (Miettinen, 1999). The follow-on question is inverse pricing: which downstream cost ratios justify a fixed action convention at a Pareto-relevant regime? Further derivations are deferred to Appendices C and I.10. 4 Consequences of a fixed conformal partition in the binary case This section isolates the main binary geometric fact behind the operational trade-offs. Once thresholds are fixed, deployment behavior is governed by the joint region–class table (Rτ(X),Y)(R_τ(X),Y). Under probability-normalized scores, the threshold pair (τ0,τ1)( _0, _1) cannot move operational KPIs independently; it reallocates mass across a small number of region types. Equivalently, this section makes precise the conservation-constraint intuition behind the attainable sets and Pareto filtering introduced in Section 3.4. In the binary case this yields three practical implications: coupled feasibility, region-wise policy dependence, and decision optimality determined by the audited table rather than by coverage alone. Appendix E contains the full construction and proofs, while Appendix F develops the decision-theoretic consequences. 4.1 Binary conformal geometry and regime boundaries Let =0,1Y=\0,1\ and fix class-conditional thresholds τ=(τ0,τ1)τ=( _0, _1). Once calibrated, the object that governs operational characteristics in deployment is the region label Rτ(x):=(s(x,0)≤τ0, 1s(x,1)≤τ1)∈0,12,R_τ(x):= (1\s(x,0)≤ _0\,\ 1\s(x,1)≤ _1\ )∈\0,1\^2, with outcomes 10,11,01,0010,11,01,00, which under the set inclusion policy πSI _SI correspond to singleton-0, hedge, singleton-11, and abstention, respectively. Probability-normalized scores induce a regime boundary. For probability-normalized scores s(x,y)=1−P(y∣x)s(x,y)=1-P(y x) we have s(x,0)+s(x,1)=1s(x,0)+s(x,1)=1. Hence a point cannot satisfy both class thresholds unless τ0+τ1≥1 _0+ _1≥ 1, and it cannot violate both unless τ0+τ1≤1 _0+ _1≤ 1. Consequently: • Hedging regime (τ0+τ1>1 _0+ _1>1): 1111 may occur and 0000 cannot (singletons + hedges; no abstention). • Rejection regime (τ0+τ1<1 _0+ _1<1): 0000 may occur and 1111 cannot (singletons + abstention; no hedges). • Boundary (τ0+τ1=1 _0+ _1=1): only 1010 and 0101 occur; under πSI _SI outputs are always singletons. Crossing τ0+τ1=1 _0+ _1=1 therefore changes which outcome types are even feasible. This deterministic boundary is the simplest geometric reason that changing thresholds can produce qualitatively different operational regimes. Cross-threshold coupling within regimes. Within the hedging regime (τ0+τ1>1 _0+ _1>1), the diagonal support is partitioned into three contiguous intervals. Parameterize by u=s(x,0)u=s(x,0) (so s(x,1)=1−us(x,1)=1-u): u∈[0,1−τ1)⇒Rτ(x)=10,u∈[1−τ1,τ0]⇒Rτ(x)=11,u∈(τ0,1]⇒Rτ(x)=01.u∈[0,1- _1)\ \ R_τ(x)=10, u∈[1- _1, _0]\ \ R_τ(x)=11, u∈( _0,1]\ \ R_τ(x)=01. Thus region boundaries are governed by opposing thresholds: changing τ1 _1 moves the boundary controlling singleton-0 mass, while changing τ0 _0 moves the boundary controlling singleton-11 mass. Thresholds are therefore mass-reallocation boundaries, not independent class-wise knobs. For planning, the key consequence is that sweeping calibration settings can cross regime boundaries and induce discontinuous changes in which outputs are even feasible. This is the geometric source of the coupled attainable sets and Pareto fronts seen later in the experiments. 4.2 Region observability and interface-relative decision optimality Fix τ and treat the region label Rτ(X)∈0,12R_τ(X)∈\0,1\^2 as the deployed observable. The joint region–class label probabilities pr,y:=ℙ(Rτ(X)=r,Y=y∣cal),r∈0,12,y∈0,1,p_r,y:=P\! (R_τ(X)=r,\ Y=y _cal ), r∈\0,1\^2,\ y∈\0,1\, fully characterize the information available to any downstream rule that uses only this interface. Write pr:=pr,0+pr,1p_r:=p_r,0+p_r,1 for region mass and define the within-region label frequency ηr:=Pr(Y=1∣Rτ(X)=r,cal)=pr,1pr(pr>0). _r:= (Y=1 R_τ(X)=r,\ D_cal)= p_r,1p_r (p_r>0). Because the interface takes finitely many values, both evaluation and decision-making reduce to functions of the same table pr,y\p_r,y\ (equivalently pr,ηr\p_r, _r\). In particular, a region-based action convention is justified by the within-region label composition, not by marginal coverage alone and not by the set-valued output alone. Formally, let A be an action set and let L(a,y)L(a,y) denote the loss of taking action a∈a when Y=yY=y. A region-based policy π~:0,12→ π:\0,1\^2 is decision-optimal relative to the conformal interface if, for every populated region r, π~(r)∈argmina∈[L(a,Y)∣Rτ(X)=r,cal]. π(r)∈ _a E\! [L(a,Y) R_τ(X)=r,\ D_cal ]. In words, once the deployed interface reveals only the region label r, the optimal action is the one with smallest expected loss under the label mix within that region. This optimality is determined region-wise by the within-region label frequency ηr _r. For the main text, the key implication is simply that the same audited region–class table used for operational planning also determines whether a fixed region-based convention is rational under a stated cost model. The full Chow-style inverse-pricing analysis is deferred to Appendix F, but two qualitative consequences are worth flagging here: the resulting inverse-pricing envelope is polyhedral in cost-ratio coordinates, and coverage-matched settings can induce different, even disjoint, pricing envelopes because the region composition changes. 5 Results: evidence for coverage semantics, audit behavior, and planning This section keeps the empirical story at the same scope as the theory. We validate SSBC’s coverage semantics in simulation, use a toxicology dataset as an audit-based example of fixed-rule operational evaluation, and close with a solubility scenario-planning illustration in which the model is fixed and only the calibration layer is varied. 5.1 Numerical simulations: coverage semantics first, operational summaries second 5.1.1 Coverage: numerical realization of SSBC guarantees We first isolate the finite-sample law underlying SSBC. Calibration nonconformity scores are drawn i.i.d. from a continuous heavy-tailed reference distribution, and we compare nominal split conformal, a one-sided DKWM correction (Massart, 1990; Dvoretzky et al., 1956; Vovk, 2012b), and SSBC at (α⋆,δ)=(0.10,0.10)(α ,δ)=(0.10,0.10) over a finite deployment window of size minfer=100m_infer=100. Nominal split conformal under-controls calibration-conditional risk in this finite-window view, whereas DKWM enforces the target through strong conservatism. Across representative calibration sizes, nominal split conformal produces violation probabilities far above the requested level, DKWM drives the same probabilities well below target by substantial conservatism, and SSBC tracks the intended finite-window semantics much more closely up to unavoidable grid effects. For example, at ncal=100n_cal=100 the observed coverage violation rate is 0.40750.4075 for nominal split conformal, 0.00040.0004 for DKWM, and 0.09600.0960 for SSBC, close to the target δ=0.10δ=0.10. The full calibration-size grid with theoretical and observed violation rates is reported in Appendix C, Table 3. 5.1.2 Operational rate summaries: LOO versus two-sample audit reference We next study operational quantities beyond coverage. The goal is to compare the two-sample Calibrate–and–Audit constructed Beta–Binomial predictive summaries to a single-sample LOO surrogate intended for feasibility exploration. Synthetic probability model. We use a controlled binary probability model. Each draw produces Y∈0,1Y∈\0,1\ with ℙ(Y=1)=pclassP(Y=1)=p_class for pclass∈0.10,0.50p_class∈\0.10,0.50\, and P1∣Y=1∼Beta(a,b),P1∣Y=0∼Beta(2,7),P_1 Y=1 (a,b), P_1 Y=0 (2,7), with (a,b)∈(4,3),(9,3)(a,b)∈\(4,3),(9,3)\. We then set (P0,P1)=(1−P1,P1)(P_0,P_1)=(1-P_1,P_1) and use class-conditional scores Sy:=1−PyS_y:=1-P_y. Two-sample reference and LOO surrogate. For each configuration, we draw independent datasets 1D_1 and 2D_2 of size N=500N=500, calibrate on 1D_1 with SSBC-adjusted thresholds at (α,δ)=(0.10,0.10)(α,δ)=(0.10,0.10), freeze the rule, and evaluate operational indicators on 2D_2. This yields the two-sample audit reference. As a single-sample surrogate, we recompute thresholds leave-one-out, pool the resulting indicators, and map them to planning envelopes, optionally widened by controlled pessimization (Appendix D). Figure 4 shows that, in these simulated geometries, LOO-based envelopes align closely with the two-sample reference, while inflation widens intervals without shifting centers. This supports LOO as a practical planning proxy when an explicit audit split is unavailable, while keeping the main operational evidence tied to the two-sample design. Figure 4: Operational rate envelopes and score geometry. (A) Singleton rate and (B) singleton error for two class prevalences and two class 1 generating distributions, with class 0 drawn from Beta(2,7)Beta(2,7). Red rectangles denote the two-sample Beta–Binomial future-window summary and the dashed vertical line its center; blue and orange intervals show leave-one-out envelopes under two inflation levels. (C1–C4) Histograms of predicted class 1 probability by true class make explicit how score geometry shapes singleton mass, singleton error, and the asymmetry of the envelopes. 5.2 Tox21: an empirical example of the audit-based operational view Tox21 (Mayr et al., 2016; Huang et al., 2016) stress-tests the framework under severe class imbalance, where class-conditional calibration counts can be small. Across twelve endpoints and 100 random train/calibration/audit splits, we compare nominal split conformal, DKWM, and SSBC at (α,δ)=(0.10,0.10)(α,δ)=(0.10,0.10). Dataset composition and representative endpoint-level operational tables are deferred to Appendix H. The aggregate calibration-conditional picture follows the same trend as the numerical simulation: nominal split conformal again shows elevated coverage violation in this conditional small-n regime, DKWM suppresses violations through strong conservatism, and SSBC lands much closer to the intended finite-sample semantics while avoiding DKWM’s excess inflation. The set-size trade-off is similarly clear: mean set size is 1.411.41 for nominal split conformal, 1.781.78 for DKWM, and 1.541.54 for SSBC, with singleton frequency ordered in the opposite direction. The aggregated coverage and set-size summary is reported in Appendix H, Table 5. 5.2.1 Operational summaries on an independent evaluation split With coverage semantics fixed by SSBC at (α,δ)=(0.10,0.10)(α,δ)=(0.10,0.10), we examine the induced region–class summaries on an independent audit split. The point of the endpoint-level table is different from the aggregate coverage table above: it shows what the fixed rule actually does in deployment-facing KPI terms and how closely the LOO proxy tracks that held-out operational picture. Table 1: Representative Tox21 endpoint: SR-MMP. Joint rates are normalized by the endpoint test-set size. The 1D_1 LOO columns provide planning summaries (Point Estimate PE and 95% Prediction Interval PI) from calibration data, while the 2D_2 columns report the independent audit reference and its predictive interval. Operational quantity Class LOO PE 1D_1 LOO 95% PI 1D_1 PE 2D_2 B 95% PI 2D_2 Singleton rate Class 0 0.6620.662 [0.604,0.718][0.604,0.718] 0.6510.651 [0.615,0.686][0.615,0.686] Class 1 0.1300.130 [0.092,0.173][0.092,0.173] 0.1170.117 [0.094,0.142][0.094,0.142] Doublet rate Class 0 0.1630.163 [0.121,0.211][0.121,0.211] 0.2010.201 [0.172,0.231][0.172,0.231] Class 1 0.0450.045 [0.023,0.074][0.023,0.074] 0.0310.031 [0.020,0.046][0.020,0.046] Wrong-singleton rate Class 0 0.0730.073 [0.044,0.107][0.044,0.107] 0.0700.070 [0.052,0.090][0.052,0.090] Class 1 0.0120.012 [0.002,0.030][0.002,0.030] 0.0110.011 [0.005,0.020][0.005,0.020] For SR-MMP, the LOO proxy and the independent audit split agree closely on the main operational picture: most mass lies in singleton predictions for class 0, class 1 singleton mass is smaller but still stable, and wrong-singleton rates remain low relative to total singleton mass. This is the intended role of the endpoint table in the paper: not to certify the LOO proxy, but to show that the audit-based operational view yields a compact, interpretable deployment summary at a fixed rule. A second endpoint illustrating a lower-prevalence regime is kept in Appendix H. 5.3 Solubility: scenario planning once coverage is fixed The solubility case study is a planning illustration of the attainable operational KPI trade-offs. We train a fixed model on AquaSolDB (Sorkun et al., 2019), restrict calibration to a lipophilic deployment scenario, and use the LOO planning interface on the scenario-restricted sample to expose the feasible operating regimes once coverage semantics have been fixed. Figure 5 summarizes the resulting planning interface. The left panel shows that sweeping SSBC settings does not induce a simple monotone trade-off between soluble-class exclusion, deferral, and decisive-correct mass; instead it yields a constrained attainable set with a small set of Pareto-relevant regimes. The right panel shows the associated inverse-pricing screen for a fixed downstream convention, illustrating that the same Pareto-relevant regimes need not remain decision-optimal under the same cost-ratio assumptions. Figure 5: Solubility planning and inverse pricing. Left: attainable planning regimes induced by sweeping SSBC settings on a restricted deployment scenario. Right: cost-ratio regions for which a fixed downstream convention is decision-optimal relative to the conformal interface. The main message is qualitative: once coverage semantics are fixed, the same conformal interface can support multiple planning regimes with materially different operational profiles. Detailed KPI tables for selected Pareto-optimal regimes, scenario construction, and parameter diagnostics are listed in Appendix I. 6 Discussion & Conclusions For a fixed deployed binary split-conformal rule under exchangeability, coverage and deployment behavior are not the same object. SSBC gives a finite-sample semantic interpretation of the coverage request, and Calibrate–and–Audit evaluates the resulting fixed rule through the induced region–class table. That table is the reusable deployment summary in this setup: it supports KPI estimation, planning views, and downstream decision-theoretic questions without retraining the base model. This is the paper’s intended division of labor. SSBC answers what coverage claim is being made about the realized deployed rule; Calibrate–and–Audit answers what that same fixed rule is likely to do operationally over a finite window. The geometric analysis explains why both layers are needed: thresholds reallocate mass across a fixed partition, so operational KPIs are coupled. The simulations and case studies should be read at exactly that scope: the simulations validate the intended SSBC semantics under calibration randomness, the Tox21 study shows that an independent audit split yields interpretable fixed-rule KPI summaries in a realistic small-sample regime, and the AquaSolDB example shows how the same interface can be used for planning once coverage semantics are fixed. Taken together, these experiments support a division of labor that we regard as practically useful: use SSBC to decide what coverage claim is being made about the deployed rule, then use an audit table to understand what that rule is likely to do operationally. This viewpoint also sharpens a broader conceptual point. A conformal output is not fully characterized, for deployment purposes, by its marginal coverage or by its set-valued form alone. What matters operationally is how calibrated regions align with labels in the deployment distribution. That is why coverage-matched settings can still differ in commitment, deferral, decisive error exposure, and downstream cost compatibility. In this sense, the region–class table is not auxiliary bookkeeping; it is the minimal object in our setup that connects a fixed conformal interface to operational and decision-theoretic consequences. This also gives a concrete response to the HCI concern that conformal sets can be a “murky” interface: the set alone is often not enough, but the set paired with its audited region–class table makes the deployment-facing consequences of acting on that interface explicit. The scope remains intentionally narrow. We treat the score model as fixed, focus on binary classification, and rely on exchangeability. The exact finite-sample claims apply to SSBC’s coverage semantics and to independent-audit inference for fixed operational rates; the LOO interface is used only as a planning proxy, and attainable-set sweeps and Pareto-front views are exploratory rather than simultaneous certification statements. Extending the same viewpoint to multi-class outputs, structured predictions, richer abstention policies, adaptive policies, or shift-aware settings remains future work. The most serious practical limitation is distribution or data drift: once exchangeability fails, the finite-sample guarantees need not hold, so deployment would require monitoring, recalibration, or shift-aware extensions (Fannjiang et al., 2022). Our aim is to make the fixed-rule deployment picture explicit before those complications are layered in. Acknowledgements This work was supported by the U.S. Department of Energy, Office of Science, under Contract No. DE-AC02-05CH11231 as part of the Laboratory Directed Research and Development (LDRD) program. Additional support came in part from the U.S. Department of Energy, Office of Science, Scientific Discovery through Advanced Computing (SciDAC) program FORUM-AI for first-principles calculations and AI model development. References A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster (2024) Conformal risk control. In International Conference on Learning Representations, Cited by: §1.2. A. N. Angelopoulos and S. Bates (2023) Conformal prediction: a gentle introduction. Foundations and Trends in Machine Learning 16 (4), p. 494–591. External Links: Document, 2107.07511 Cited by: §1.2, §1. R. F. Barber, E. J. Candès, A. Ramdas, and R. J. Tibshirani (2021) Predictive inference with the jackknife+. Annals of Statistics 49 (1), p. 486–507. Cited by: §D.2. P. L. Bartlett and M. H. Wegkamp (2008) Classification with a reject option using a hinge loss. Journal of Machine Learning Research 9, p. 1823–1840. Cited by: §F.5, Appendix F. S. Bates, A. Angelopoulos, L. Lei, J. Malik, and M. I. Jordan (2021) Distribution-free, risk-controlling prediction sets. Journal of the ACM 68 (6). External Links: ISSN 0004-5411, Document Cited by: §1.2. M. Bian and R. F. Barber (2023) Training-conditional coverage for distribution-free predictive inference. Electronic Journal of Statistics 17 (2), p. 2044–2066. External Links: Document Cited by: §1.2. G. Casella and R. L. Berger (2002) Statistical inference. 2nd edition, Duxbury. Cited by: Appendix F. C. K. Chow (1970) On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory 16 (1), p. 41–46. External Links: Document Cited by: §F.5, Appendix F. C. J. Clopper and E. S. Pearson (1934) The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26 (4), p. 404–413. Cited by: §3.3. H. A. David and H. N. Nagaraja (2003) Order statistics. Wiley. Cited by: §B.2, §B.3. A. Dvoretzky, J. Kiefer, and J. Wolfowitz (1956) Asymptotic minimax character of the sample distribution function and of the classical multinomial estimator. Annals of Mathematical Statistics 27 (3), p. 642–669. Cited by: §5.1.1. C. Fannjiang, S. Bates, A. N. Angelopoulos, J. Listgarten, and M. I. Jordan (2022) Conformal prediction under feedback covariate shift for biomolecular design. Proceedings of the National Academy of Sciences 119 (43), p. e2204569119. External Links: Document Cited by: §I.6, §1.2, §6. A. Gelman, J. B. Carlin, H. S. Stern, D. B. Dunson, A. Vehtari, and D. B. Rubin (2013) Bayesian data analysis. 3rd edition, Chapman and Hall/CRC, Boca Raton, FL. Cited by: §3.3. I. Gibbs, J. J. Cherian, and E. J. Candès (2025) Conformal prediction with conditional guarantees. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 87 (4), p. 1100–1126. Note: to appear External Links: Document Cited by: §1.2. R. Herbei and M. H. Wegkamp (2006) Classification with a reject option when the minimum excess risk is zero. Canadian Journal of Statistics 34 (4), p. 563–579. External Links: Document Cited by: §F.5, Appendix F. W. Heyndrickx, A. Arany, J. Simm, A. Pentina, N. Sturm, L. Humbeck, L. Mervin, A. Zalewski, M. Oldenhof, P. Schmidtke, L. Friedrich, R. Loeb, A. Afanasyeva, A. Schuffenhauer, Y. Moreau, and H. Ceulemans (2023) Conformal efficiency as a metric for comparative model assessment befitting federated learning. Artificial Intelligence in the Life Sciences 3, p. 100070. External Links: ISSN 2667-3185 Cited by: §G.3. R. Huang, M. Xia, D. Nguyen, J. Zhao, S. Sakamuru, T. Zhao, M. Na, S. A. Shahane, A. Rossoshek, and A. Simeonov (2016) The tox21 data challenge to build predictive models of nuclear receptor and stress response pathways. Nature Biotechnology 34 (8), p. 828–837. External Links: Document Cited by: §5.2. J. Hullman, Y. Wu, D. Xie, Z. Guo, and A. Gelman (2025) Conformal prediction and human decision making. External Links: 2503.11709 Cited by: §1.2. N. L. Johnson, S. Kotz, and N. Balakrishnan (1997) Discrete multivariate distributions. Wiley, New York. External Links: ISBN 978-0-471-31250-3 Cited by: §3.3. S. Kalepu and V. Nekkanti (2015) Insoluble drug delivery strategies: review of recent advances and business prospects. Acta Pharmaceutica Sinica B 5 (5), p. 442–453. External Links: Document Cited by: §I.1. S. Kiyani, G. J. Pappas, A. Roth, and H. Hassani (2025) Decision theoretic foundations for conformal prediction: optimal uncertainty quantification for risk-averse agents. In Forty-second International Conference on Machine Learning, Cited by: §1.2. J. Lei, M. G’Sell, A. Rinaldo, R. J. Tibshirani, and L. Wasserman (2018) Distribution-free predictive inference for regression. Journal of the American Statistical Association 113 (523), p. 1094–1111. External Links: Document Cited by: §2.2. J. Lekeufack, A. N. Angelopoulos, A. Bajcsy, M. I. Jordan, and J. Malik (2023) Conformal decision theory: safe autonomous decisions from imperfect predictions. arXiv preprint arXiv:2310.05921. Cited by: §1.2. P. C. F. Marques (2025) Universal distribution of the empirical coverage in split conformal prediction. Statistics & Probability Letters 219, p. 110350. External Links: ISSN 0167-7152, Document Cited by: Appendix B, §C.5. P. Massart (1990) The tight constant in the dvoretzky–kiefer–wolfowitz inequality. Annals of Probability 18 (3), p. 1269–1283. External Links: Document Cited by: §5.1.1. A. Mayr, G. Klambauer, T. Unterthiner, and S. Hochreiter (2016) DeepTox: toxicity prediction using deep learning. Frontiers in Environmental Science 3, p. 80. External Links: Document Cited by: §5.2. K. Miettinen (1999) Nonlinear multiobjective optimization. Kluwer Academic Publishers, Boston. Cited by: §3.4. L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin (2018) CatBoost: unbiased boosting with categorical features. In Advances in Neural Information Processing Systems, Vol. 31, p. 6638–6648. Cited by: §I.4. D. Rogers and M. Hahn (2010) Extended-connectivity fingerprints. Journal of Chemical Information and Modeling 50 (5), p. 742–754. External Links: Document Cited by: §I.3. G. Shafer and V. Vovk (2008) A tutorial on conformal prediction. Journal of Machine Learning Research 9, p. 371–421. Cited by: §1.2, §1, §2.2. J. G. Skellam (1948) A probability distribution derived from the binomial distribution by regarding the probability of success as variable between trials. Journal of the Royal Statistical Society. Series B (Methodological) 10 (2), p. 257–261. Cited by: §3.3. M. C. Sorkun, A. Khetan, and S. Er (2019) AqSolDB: a curated reference set of aqueous solubility and 2d descriptors for a diverse set of compounds. Scientific Data 6, p. 143. External Links: Document Cited by: §I.1, §5.3. R. J. Tibshirani, R. F. Barber, E. J. Candès, and A. Ramdas (2019) Conformal prediction under covariate shift. In Advances in Neural Information Processing Systems, Vol. 32. Cited by: §1.2. V. Vovk, A. Gammerman, and G. Shafer (2005) Algorithmic learning in a random world. Springer, New York. External Links: Document Cited by: §1.2, §1. V. Vovk (2012a) Conditional validity of inductive conformal predictors. Journal of Machine Learning Research 13, p. 955–997. Cited by: §2.2. V. Vovk (2012b) Conditional validity of inductive conformal predictors. arXiv. Note: Extended version of ACML 2012 paper External Links: 1209.2673, Document Cited by: §I.5, §2.2, §5.1.1. V. Vovk (2015) Cross-conformal predictors. Annals of Mathematics and Artificial Intelligence 74, p. 9–28. External Links: Document Cited by: §D.2. F. Yang and A. K. Kuchibhotla (2021) Finite-sample efficient conformal prediction. Annals of Statistics 49 (5), p. 2921–2947. Cited by: §1.2. M. Yuan and M. H. Wegkamp (2010) Classification methods with reject option based on convex risk minimization. Journal of Machine Learning Research 11 (5), p. 111–130. Cited by: §F.4, §F.5, Appendix F. W. Zhou and S. Zhu (2025) Calibrating decision robustness via inverse conformal risk control. arXiv preprint arXiv:2510.07750. External Links: 2510.07750 Cited by: §1.2. Appendix A Appendix Outline & Notation This appendix serves as a notation guide and roadmap for the technical appendices that follow. The appendices are, in order: • A. Appendix Outline & Notation (this appendix): roadmap and shared notation for calibration, SSBC, and operational analysis. • B. Finite-sample distribution of calibration-conditional coverage: pivot and distribution of coverage given the calibration draw; basis for coverage-conditional guarantees. • C. Small-Sample Beta Correction (SSBC): mapping (α⋆,δ)(α ,δ) to a calibration grid index; finite-sample coverage semantics for the deployed rule. • D. Single-sample structural coupling, LOO decoupling, and planning envelopes: why an independent audit split is needed for certification; leave-one-out surrogate and envelope construction when no audit split is available. • E. Binary conformal partitions: four-region structure, regime boundary (τ0+τ1=1 _0+ _1=1), and recovery of the argmax and Bayes optimal classification rule under asymmetric costs. • F. Conformal interface-relative decision optimality and inverse pricing: cost-ratio conditions for decision-optimal action; Chow-style accept/reject and pricing envelopes. • G. Explicit region indicators and projection masks: formulas for region–label counts and projections used in auditing and KPI computation. • H. Tox21 supplementary details: dataset, splits, and operational summaries for the Tox21 empirical example. • I. Solubility supplementary details: dataset, scenario restriction, Pareto-front parameter roles, and α vs. δ interpretation for the solubility case study. The notation below is shared across split conformal calibration, the Small-Sample Beta Correction (SSBC), and finite-window prediction of operational rates. The analysis mixes three layers: (i) discrete split-conformal thresholding via order statistics on calibration scores; (i) SSBC grid selection and calibration-conditional coverage semantics; and (i) predictive envelopes for deployment-facing operational rates (Calibrate-and-Audit and a single-sample leave-one-out (LOO) surrogate). To avoid collisions, we reserve (k,u)(k,u) exclusively for split conformal thresholding. Table 2 summarizes the key notation. Table 2: Key notation used across calibration, SSBC, and operational analysis. Symbol Meaning Scope / Remarks Data splits and window sizes calD_cal, ncaln_cal Calibration dataset and size Exchangeable sample for conformal thresholds. auditD_audit, nauditn_audit Audit dataset and size Exchangeable sample for region–class label and operational rate estimation. m Future window size Number of future cases for realized rates and predictive envelopes. Region–policy–audit Rτ(x)R_τ(x), Rτ(θ)R_τ(θ) Region map Partition of score space; Rτ:→0,1KR_τ:X→\0,1\^K; fixed by thresholds τ(θ)τ(θ). π Deployment policy Maps region Rτ(x)R_τ(x) to reported output (e.g., prediction set or ∅ ). θ, Θ Calibration setting and space Indexes deployed thresholds τ(θ)τ(θ) and region map Rτ(θ)R_τ(θ). Split conformal and SSBC s(x,y)s(x,y), SiS_i, S(k)S_(k) Score, calibration scores, order statistic τ=S(k)τ=S_(k); k,uk,u with u=ncal+1−ku=n_cal+1-k index the conformal grid. αgrid _grid, αadj _adj Grid and SSBC-selected miscoverage αgrid=u/(ncal+1) _grid=u/(n_cal+1); SSBC returns αadj=u⋆/(ncal+1) _adj=u /(n_cal+1). α⋆α , δ Target miscoverage, confidence User request; δ controls tail probability over calibration draws. pcovp_cov, C^m C_m Coverage probability, empirical coverage pcov∼Beta(k,u)p_cov (k,u); C^m=1m∑j=1mYj′∈C(Xj′) C_m= 1m _j=1^m1\Y _j∈ C(X _j)\. Operational indicators and LOO C(X)C(X), gℓg_ , IℓI_ Prediction set, event functional, indicator Iℓ=gℓ(Rτ(θ)(X),Y;π)I_ =g_ (R_τ(θ)(X),Y;π); KPIs are sums over selected region–class cells. Zi,jZ_i,j, kpool,jk_pool,j, r^jLOO r^LOO_j LOO indicator, pooled count, LOO rate Zi,j=gj(C−i(Xi),Yi)Z_i,j=g_j(C_-i(X_i),Y_i); kpool,j=∑iZi,jk_pool,j= _iZ_i,j; r^jLOO=kpool,j/n r^LOO_j=k_pool,j/n. infl, neffn_eff Inflation (pessimization) neff=n/infln_eff=n/ infl widens LOO envelopes (Appendix D); larger infl yields wider intervals. Appendix B Finite-sample distribution of calibration-conditional coverage This appendix derives a finite-sample characterization of the calibration-conditional coverage of a fixed split conformal predictor under exchangeability. Coverage is a pure rank event: the future true-label score is compared to a calibration order statistic. This yields a distribution-free pivot and a Beta law for realized (calibration-conditional) coverage across calibration draws. This pivot is the input to SSBC (Appendix C). The derivation presented here is a one-shot rank/order-statistic pivot, complementary to the Beta–Binomial/de Finetti route in Marques Filho’s analysis of the exchangeable sequence of future coverage indicators (Marques, 2025). B.1 Setup Let cal=(Xi,Yi)i=1nD_cal=\(X_i,Y_i)\_i=1^n be exchangeable with a future test pair (Xn+1,Yn+1)(X_n+1,Y_n+1). Let s(x,y)s(x,y) be a nonconformity score and define Si:=s(Xi,Yi),i=1,…,n,andSn+1:=s(Xn+1,Yn+1).S_i:=s(X_i,Y_i), i=1,…,n, S_n+1:=s(X_n+1,Y_n+1). Fix an index k∈1,…,nk∈\1,…,n\ and set τ:=S(k),u:=n+1−k,αgrid=un+1.τ:=S_(k), u:=n+1-k, _grid= un+1. For the predictor calibrated on calD_cal, the calibration-conditional coverage probability is pcov(cal):=ℙ(Sn+1≤τ∣cal).p_cov(D_cal):=P\! (S_n+1≤τ _cal ). This random variable varies across calibration draws but is fixed once calibration is completed. B.2 Rank pivot Assume for clarity that the score distribution is continuous so ties occur with probability zero (ties are addressed in Remark 1). Under exchangeability of S1,…,Sn,Sn+1\S_1,…,S_n,S_n+1\, the rank R:=rank(Sn+1amongS1,…,Sn,Sn+1)R:=rank (S_n+1\ among\ S_1,…,S_n,S_n+1 ) is uniform on 1,…,n+1\1,…,n+1\ (David and Nagaraja, 2003). Since τ=S(k)τ=S_(k), Sn+1≤τ⟺R≤k.\S_n+1≤τ\ \R≤ k\. Thus coverage is the conditional probability of a rank event. B.3 Beta law A standard order-statistic identity implies that the conditional probability of the rank event equals the kkth order statistic of n+1n+1 i.i.d. uniforms: pcov(cal)=dU(k),U(k)∼Beta(k,u),u=n+1−k,p_cov(D_cal)\; d=\;U_(k), U_(k) (k,u), u=n+1-k, where U(k)U_(k) is the kkth order statistic of U1,…,Un+1∼i.i.d.Unif(0,1)U_1,…,U_n+1 i.i.d. Unif(0,1) (David and Nagaraja, 2003). Equivalently, for any t∈[0,1]t∈[0,1], ℙ(pcov(cal)≤t)=It(k,u),P\! (p_cov(D_cal)≤ t )=I_t(k,u), where It(⋅,⋅)I_t(·,·) is the regularized incomplete Beta function. The law is distribution-free and depends only on (n,k)(n,k) (equivalently (n,u)(n,u)). Remark 1 (Discrete scores and ties). If scores have atoms, ranks are not almost surely unique. Exact pivots can be recovered by randomized tie-breaking; deterministic left/right-continuous conventions yield conservative bounds. The non-interpolated order-statistic thresholding convention used in the main text preserves finite-sample validity. Coverage depends only on the rank of the future true-label score relative to the calibration scores, hence admits a finite-sample distribution-free law. Most other operational KPIs depend on the joint geometry of multiple label scores and threshold interactions and therefore do not admit an analogous pivot (Appendix D). Appendix C Small-Sample Beta Correction (SSBC) SSBC is a deterministic index-selection rule for split conformal calibration that makes a user request (α⋆,δ)(α ,δ) operationally precise for the single deployed predictor obtained after one calibration draw. SSBC selects a conformal grid index (equivalently an order statistic) so that, with probability at least 1−δ1-δ over calibration randomness, the realized coverage of the deployed rule is at least 1−α⋆1-α . In the finite-window variant, the same guarantee is imposed on empirical coverage over a future window of size m. The construction relies only on the finite-sample distribution of calibration-conditional coverage from Appendix B. C.1 Setup and conformal grid Let ncaln_cal be the calibration size and let Si:=s(Xi,Yi)S_i:=s(X_i,Y_i) denote true-label nonconformity scores for i=1,…,ncali=1,…,n_cal. Split conformal selects a threshold as an order statistic: τ=S(k),k∈1,…,ncal.τ=S_(k), k∈\1,…,n_cal\. It is convenient to re-index the same grid by the miscoverage index u:=ncal+1−k∈1,…,ncal,αgrid=uncal+1,k=ncal+1−u.u:=n_cal+1-k∈\1,…,n_cal\, _grid= un_cal+1, k=n_cal+1-u. Here α⋆∈(0,1)α ∈(0,1) is the requested miscoverage and δ∈(0,1)δ∈(0,1) is a confidence/risk level controlling the probability (over calibration draws) that the requested semantics fail. C.2 Distribution of realized coverage For the fixed predictor calibrated on calD_cal, define the calibration-conditional coverage probability pcov(cal):=ℙ(Sn+1≤τ∣cal),p_cov(D_cal):=P(S_n+1≤τ _cal), where (Xn+1,Yn+1)(X_n+1,Y_n+1) is exchangeable with the calibration sample. Under exchangeability, Appendix B shows that pcov(cal)∼Beta(k,u),k=ncal+1−u,p_cov(D_cal) (k,u), k=n_cal+1-u, exactly and distribution-free. This describes how the realized coverage of the deployed predictor varies across hypothetical recalibrations, while treating the deployed rule as fixed after the one calibration step. C.3 SSBC objective: calibration-conditional PAC semantics SSBC enforces the calibration-conditional PAC-style constraint ℙcal(pcov(cal)≥ 1−α⋆)≥ 1−δ.P_D_cal\! (p_cov(D_cal)\ ≥\ 1-α )\ ≥\ 1-δ. Using the Beta law, this is equivalent to ℙ(Z≥1−α⋆)≥ 1−δ,Z∼Beta(k,u).P\! (Z≥ 1-α )\ ≥\ 1-δ, Z (k,u). Selection rule: least conservative admissible grid point. Among discrete grid indices u∈1,…,ncalu∈\1,…,n_cal\, SSBC selects the largest admissible u satisfying the tail constraint above: u⋆:=maxu∈1,…,ncal:ℙ(Z≥1−α⋆)≥1−δ,Z∼Beta(ncal+1−u,u).u := \u∈\1,…,n_cal\:\ P(Z≥ 1-α )≥ 1-δ,\ Z (n_cal+1-u,\ u) \. Equivalently, SSBC deploys the least conservative grid miscoverage level αadj=u⋆/(ncal+1) _adj=u /(n_cal+1) that certifies the requested semantics. The returned order-statistic index is kadj=ncal+1−u⋆k_adj=n_cal+1-u . C.4 Finite-window deployment semantics In many deployments, coverage is evaluated over a finite window of size m. Define empirical coverage over that window by C^m:=1m∑j=1mIcov,j,Icov,j:=Yj′∈C(Xj′),Sm:=mC^m. C_m:= 1m _j=1^mI_cov,j, I_cov,j:=1\Y _j∈ C(X _j)\, S_m:=m C_m. Conditional on the calibration-conditional coverage probability pcov(cal)=p_cov(D_cal)=p, Sm∣p∼Binomial(m,p).S_m p (m,p). Marginalizing p∼Beta(k,u)p (k,u) yields the mixture distribution Sm∼Beta-Binomial(m;k,u),k=ncal+1−u.S_m -Binomial(m;k,u), k=n_cal+1-u. The finite-window SSBC criterion selects u such that ℙ(C^m≥1−α⋆)≥ 1−δ,P\! ( C_m≥ 1-α )\ ≥\ 1-δ, where the probability is over both calibration randomness and the future window. C.4.1 Strict lower-tail convention Because C^m C_m is discrete, we adopt a strict violation convention C^m<1−α⋆ C_m<1-α . Define the corresponding count threshold x⋆:=⌊(1−α⋆)m⌋+1,so thatC^m≥1−α⋆⇔Sm≥x⋆.x := (1-α )m +1, that \ C_m≥ 1-α \ \S_m≥ x \. All Beta–Binomial tail probabilities in SSBC are evaluated for the event Sm≥x⋆\S_m≥ x \. This avoids boundary ambiguity when (1−α⋆)m(1-α )m is an integer and preserves monotonicity in u. C.5 Infinite-window limit This limiting regime is identified by Marques (2025). In our notation, if umu_m denotes the SSBC-selected grid index obtained by inverting the Beta–Binomial tail for window size m, and u∞u_∞ denotes the index obtained from the Beta limit law, then um→u∞u_m→ u_∞ as m→∞m→∞. Intuitively, conditional on pcov=p_cov=p, the empirical coverage C^m C_m converges almost surely to p, so the finite-window Beta–Binomial criterion approaches the infinite-window Beta criterion. C.6 Feasibility and saturation Not every pair (α⋆,δ)(α ,δ) is feasible at fixed ncaln_cal because calibration choices lie on the conformal grid. Under the most conservative grid point u=1u=1 (equivalently k=ncalk=n_cal), pcov∼Beta(ncal,1),ℙ(pcov≥1−α)=1−(1−α)ncal.p_cov (n_cal,1), \! (p_cov≥ 1-α )=1-(1-α)^n_cal. Thus any infinite-window PAC requirement at confidence 1−δ1-δ must satisfy α≥ 1−δ1/ncal.α\ ≥\ 1-δ^1/n_cal. If this condition is violated (and likewise in the finite-window analogue), no grid point can satisfy the tail constraint and SSBC returns Infeasible. C.7 Coverage semantics under finite calibration To visualize how nominal requests map to deployed semantics under finite calibration, Figure 6 plots the effective calibration level αadj _adj selected by SSBC as a function of the user inputs (α,δ)(α,δ) at fixed ncaln_cal. Distinct nominal requests can induce the same deployed grid index and therefore the same realized coverage semantics. Figure 6: Semantic interpretation of nominal coverage requests under finite calibration.Each panel visualizes the effective calibration level αadj _adj selected by SSBC as a function of the user-specified miscoverage level α and confidence level δ, for fixed ncaln_cal. Color encodes the deployed level αadj _adj, while contours indicate iso-semantic sets: distinct nominal requests that induce the same deployed calibration grid point. The feasibility boundary reflects the finite-sample constraint α≳1−δ1/ncalα 1-δ^1/n_cal. C.8 SSBC algorithm (deterministic specification) This subsection provides a reproducible implementation-level specification. The algorithm returns the largest admissible u (least conservative grid point) satisfying the relevant tail constraint. Algorithm 1 Small-Sample Beta Correction (SSBC) 1:Target miscoverage α⋆∈(0,1)α ∈(0,1); calibration size ncal∈ℕn_cal ; confidence δ∈(0,1)δ∈(0,1); deployment regime ∈∞,m∈\∞,m\ (window size m if finite) 2:Adjusted grid level αadj _adj and index kadjk_adj, or Infeasible 3:t←1−α⋆t← 1-α 4:u⋆←−1u ←-1 5:if regime =m=m then 6: x⋆←⌊tm⌋+1x ← t\,m +1 7:end if 8:for u=1,…,ncalu=1,…,n_cal do 9: a←ncal+1−ua← n_cal+1-u, b←ub← u 10: if regime =∞=∞ then 11: ptail←Pr[Z≥t],Z∼Beta(a,b)p_tail← [Z≥ t],\ Z (a,b) 12: else 13: ptail←Pr[X≥x⋆],X∼Beta-Binomial(m;a,b)p_tail← [X≥ x ],\ X -Binomial(m;a,b) 14: end if 15: if ptail≥1−δp_tail≥ 1-δ then 16: u⋆←u ← u 17: end if 18:end for 19:if u⋆<0u <0 then 20: return Infeasible 21:end if 22:αadj←u⋆ncal+1 _adj← u n_cal+1 23:kadj←ncal+1−u⋆k_adj← n_cal+1-u 24:return αadj,kadj _adj,\ k_adj Relation to DKWM-style calibration. DKWM-style calibration modifies the nominal grid choice to enforce conservative, worst-case guarantees uniformly over calibration draws and distributions. SSBC addresses a different question: it assigns a calibration-conditional PAC meaning to a user request (α⋆,δ)(α ,δ) for the single deployed predictor produced after one calibration. DKWM targets uniform validity across hypothetical recalibrations; SSBC targets admissibility of the realized rule via Beta (or Beta–Binomial) tails. C.9 Extended simulation table for calibration-conditional violations Table 3 reports the full calibration-size grid for the simulation summarized in Section 5.1.1. Table 3: Calibration-conditional coverage violation rates with theory. Target miscoverage α⋆=0.10α =0.10, confidence δ=0.10δ=0.10. αgrid _grid is the grid point selected on the conformal grid, αcont _cont is the requested value under the DKWM correction, and mcalm_cal is the calibration-window size. Results are based on 10610^6 calibration draws and a finite deployment window of size minfer=100m_infer=100. The Obs column reports ℙ(C^m<1−α⋆)P( C_m<1-α ). Beta reports ℙ(pcov<1−α⋆)P(p_cov<1-α ) under pcov∼Beta(k,u)p_cov (k,u). BetaBinom reports ℙ(C^m<1−α⋆)P( C_m<1-α ) under the induced Beta–Binomial law. ncaln_cal Method u αgrid _grid Obs Beta BetaBinom αcont _cont 50 None 5 0.0980 0.3963 0.4312 0.3964 – 50 SSBC 2 0.0392 0.0476 0.0338 0.0472 – 50 DKWM 1 0.0196 0.0096 0.0052 0.0095 −0.0731-0.0731 75 None 7 0.0921 0.3454 0.3673 0.3464 – 75 SSBC 4 0.0526 0.0769 0.0504 0.0768 – 75 DKWM 1 0.0132 0.0016 0.0004 0.0017 −0.0413-0.0413 100 None 10 0.0990 0.4075 0.4513 0.4071 – 100 SSBC 6 0.0594 0.0960 0.0576 0.0956 – 100 DKWM 1 0.0099 0.0004 0.0000 0.0004 −0.0224-0.0224 150 None 15 0.0993 0.4119 0.4602 0.4107 – 150 SSBC 9 0.0596 0.0804 0.0307 0.0801 – 150 DKWM 1 0.0066 0.0000 0.0000 0.0000 +0.0001+0.0001 200 None 20 0.0995 0.4130 0.4655 0.4124 – 200 SSBC 13 0.0647 0.0974 0.0320 0.0980 – 200 DKWM 2 0.0100 0.0000 0.0000 0.0000 +0.0135+0.0135 250 None 25 0.0996 0.4126 0.4692 0.4134 – 250 SSBC 16 0.0637 0.0863 0.0175 0.0858 – 250 DKWM 5 0.0199 0.0003 0.0000 0.0003 +0.0226+0.0226 300 None 30 0.0997 0.4139 0.4719 0.4141 – 300 SSBC 20 0.0664 0.0969 0.0171 0.0971 – 300 DKWM 8 0.0266 0.0009 0.0000 0.0009 +0.0293+0.0293 500 None 50 0.0998 0.4153 0.4782 0.4153 – 500 SSBC 34 0.0679 0.0949 0.0049 0.0955 – 500 DKWM 22 0.0439 0.0097 0.0000 0.0096 +0.0453+0.0453 Appendix D Single-sample structural coupling, leave-one-out decoupling, and planning-envelope inflation The two-stage predictive reference used throughout this paper separates calibration (choose a conformal threshold on calD_cal) from operational evaluation (estimate rates on an independent window). This separation is what makes window indicators behave as i.i.d. Bernoulli draws under a fixed deployed rule. In practice, an independent audit set is not always available. When the same sample is reused both to select the threshold and to estimate operational rates, threshold selection and evaluation become coupled. This appendix records (i) a minimal structural reason for reuse-induced dependence, (i) a data-efficient remedy, leave-one-out (LOO) recalibration, that provides effective practical decoupling in our regimes, and (i) an inflation parameter infl that pessimizes predictive envelopes when residual dependence or regime instability remains. Throughout, these LOO constructions are intended for exploratory planning when a dedicated audit split is unavailable, we cannot provide strong theoretical guarantees on the rates of the LOO proxy indicators at this point. D.1 Why single-sample reuse introduces dependence Let cal=(Xi,Yi)i=1nD_cal=\(X_i,Y_i)\_i=1^n be an exchangeable calibration sample, let Si:=s(Xi,Yi)S_i:=s(X_i,Y_i) be true-label nonconformity scores, and let the split conformal threshold be the kkth order statistic τ^:=S(k),k∈1,…,n, τ:=S_(k), k∈\1,…,n\, assuming continuity so ties occur with probability zero. For any threshold t, define the crossing indicator Ii(t):=Si≤tI_i(t):=1\S_i≤ t\. Insight from counting. If τ^=t τ=t is the kkth order statistic, then exactly k−1k-1 scores lie strictly below t and one score equals t (under continuity, almost surely), so ∑i=1nIi(t)=kalmost surely. _i=1^nI_i(t)=k surely. This already shows that the indicators Ii(t)i=1n\I_i(t)\_i=1^n cannot be conditionally independent given τ^=t τ=t: once some are known to be one, fewer ones remain available for the others. By exchangeability, for any i≠ji≠ j, Cov(Ii(t),Ij(t)∣τ^=t):=k(k−n)n2(n−1)<0(k<n).Cov\! (I_i(t),I_j(t) τ=t ):= k(k-n)n^2(n-1)<0 (k<n). Implication for operational envelopes. Operational indicators are functions of the deployed prediction set Cτ^(X)=y:s(X,y)≤τ^C_ τ(X)=\y:s(X,y)≤ τ\ and therefore inherit reuse-induced dependence. In particular, the nonzero conditional covariance shows that these indicators are not Bernoulli draws under a fixed rule once threshold selection and evaluation are performed on the same sample. To recover the Bernoulli sampling picture needed for the two-stage guarantees in the main text, one therefore needs an independent audit set after calibration. If one naïvely treats reuse-based indicators as i.i.d. Bernoulli under a fixed rule, predictive envelopes can become under-dispersed. D.2 Approximate decoupling via leave-one-out (LOO) When an independent audit sample is unavailable, we use LOO recalibration (Vovk, 2015; Barber et al., 2021) to reduce self-influence. Let τ^−i τ_-i be the split conformal threshold computed on cal∖(Xi,Yi)D_cal \(X_i,Y_i)\, and let C−i(⋅)C_-i(·) be the corresponding prediction set map. For an operational event functional gj(C(X),Y)∈0,1g_j(C(X),Y)∈\0,1\ (e.g., singleton, doublet, wrong-singleton), define LOO indicators Zi,j:=gj(C−i(Xi),Yi),i=1,…,n,Z_i,j:=g_j\! (C_-i(X_i),Y_i ), i=1,…,n, and pooled LOO summaries kpool,j:=∑i=1nZi,j,r^jLOO:=kpool,jn.k_pool,j:= _i=1^nZ_i,j, r^LOO_j:= k_pool,jn. Proxy two-stage interpretation. Each Zi,jZ_i,j is evaluated under a rule that does not use point i, restoring a localized separation between rule construction and evaluation. Although the rule varies across folds, each fold differs only slightly from the full calibrated rule, so Zi,j\Z_i,j\ can be read as indicators from nearby operating regimes. Pooling provides a direct empirical proxy for operational behavior under finite calibration. Empirical decoupling. In the regimes studied (Section 5.1.2), LOO envelopes track the two-stage Calibrate–and–Audit reference closely in center and often slightly pessimistically in width. We therefore use pooled LOO indicators for planning when a separate audit set is unavailable. D.3 Envelope inflation as controlled pessimization LOO is local and does not remove all dependence, especially near regime boundaries where small threshold shifts can change region support. We therefore introduce an explicit inflation parameter infl≥1 infl≥ 1 that widens predictive envelopes by shrinking an effective sample size used in the predictive model. LOO-intrinsic pessimization compounded by SSBC. LOO recalibration introduces an intrinsic pessimization because each fold threshold τ^−i τ_-i is computed on only n−1n-1 calibration points. On top of this, we apply the small-sample Beta correction (SSBC) to achieve the desired nominal level α under finite calibration uncertainty. The SSBC step further increases conservatism (i.e., widens envelopes) relative to a plug-in rate estimate, so the combined LOO++SSBC construction is typically wider than either component alone. Importantly, SSBC is applied using the actual fold sample size (e.g., n−1n-1) to correct the nominal tail level, whereas infl below is a separate width knob acting only through the downstream dispersion model (we do not apply SSBC to neffn_eff). Rank stability of LOO thresholds (±1± 1 order statistic). In split conformal, the threshold is an order statistic of the calibration scores. Let S1,…,SnS_1,…,S_n be the full-sample calibration scores with order statistics S(1)≤⋯≤S(n)S_(1)≤·s≤ S_(n), and let k:=⌈(n+1)(1−α)⌉k:= (n+1)(1-α) so that τ^=S(k) τ=S_(k). Under LOO, the fold uses n−1n-1 scores and index k′:=⌈n(1−α)⌉∈k−1,kk := n(1-α) ∈\k-1,k\. Removing a single score can shift the rank position of the selected order statistic by at most one, hence (ignoring ties) the LOO threshold satisfies τ^−i∈S(k−1),S(k),S(k+1), τ_-i∈\S_(k-1),\,S_(k),\,S_(k+1)\, and more generally lies between adjacent full-sample order statistics. This formalizes the “local” nature of LOO recalibration, while still allowing support changes near regime boundaries when C(⋅)C(·) is sensitive to small threshold movements. Operational effect. In constructions below we replace a nominal proxy sample size n by neff=n/infln_eff=n/ infl (and analogously for any proxy count used to parameterize predictive dispersion). Larger infl yields wider envelopes without changing the pooled mean, providing a monotone knob for conservatism. Diagnostic guidance (variance-based inflation). As a diagnostic of calibration-induced variability, compute fold rates r^(−i) r^(-i) under each LOO-calibrated rule and their empirical variance r¯LOO=1n∑i=1nr^(−i),Var^LOO=1n−1∑i=1n(r^(−i)−r¯LOO)2. r_LOO= 1n _i=1^n r^(-i), Var_LOO= 1n-1 _i=1^n ( r^(-i)- r_LOO )^2. To isolate variability attributable to recalibration (rule changes) from simple delete-one evaluation noise, we also form a not-LOO reference in which the operating rule is held fixed. Let Wi:=gj(C(Xi),Yi),r^full:=1n∑i=1nWi,W_i:=g_j\! (C(X_i),Y_i ), r^full:= 1n _i=1^nW_i, and define delete-one evaluation rates under the fixed full rule C(⋅)C(·) by r^full(−i):=1n−1∑ℓ≠iWℓ=nr^full−Win−1,i=1,…,n. r^full(-i):= 1n-1 _ ≠ iW_ = n\, r^full-W_in-1, i=1,…,n. Let Var^full Var_full denote the empirical variance of r^full(−i)i=1n\ r^full(-i)\_i=1^n. We then define a variance ratio q:=Var^LOOVar^fullq:= Var_LOO Var_full and use it to suggest an inflation level infl:=max1,q, infl:= \1,q\, Intuitively, Var^full Var_full measures baseline delete-one variability when the rule is fixed, while Var^LOO Var_LOO captures additional spread induced by recalibration. Although τ^−i τ_-i can move by at most one rank, the resulting change in region support (and therefore in gjg_j) can be discontinuous near regime boundaries; larger q flags such sensitivity and points toward higher infl. In our experiments the variance ratio is typically only slightly above 11, indicating that LOO recalibration contributes a modest amount of additional variability beyond delete-one evaluation under a fixed rule. This is consistent with the rank-11 stability of LOO thresholds (up to ties). D.4 Predictive envelope constructions from LOO proxies We describe two complementary envelope constructions built from LOO proxy indicators. The first mirrors the two-stage reference and is the primary approximation; the second is a conservative guardrail. Beta–Binomial planning envelopes. Let ZiZ_i denote a pooled proxy indicator for a fixed KPI (suppress j), and let kpool=∑i=1nZik_pool= _i=1^nZ_i with pooled proxy rate p^=kpool/n p=k_pool/n. Apply inflation via neff=ninfl,keff=p^neff=kpoolinfl.n_eff= n infl, k_eff= p\,n_eff= k_pool infl. With a small prior offset offset∈1,1/2offset∈\1,1/2\, define α:=keff+offset,β:=(neff−keff)+offset.α:=k_eff+offset, β:=(n_eff-k_eff)+offset. For a future deployment window of size m, the predictive count SmS_m is modeled as Sm∼BetaBinomial(m,α,β),S_m (m,α,β), and equal-tailed prediction intervals follow from the Beta–Binomial CDF. Increasing infl shrinks neffn_eff and widens intervals monotonically while preserving the pooled mean. Hoeffding-type dominance bound. As a conservative alternative, Hoeffding’s inequality gives a distribution-free bound for a window rate r^m r_m around a proxy mean r^LOO r_LOO: Pr(|r^m−r^LOO|≥ϵ)≤2exp(−2mϵ2), \! ( | r_m- r_LOO |≥ε )≤ 2 (-2mε^2), yielding simple symmetric envelopes that can be used as a worst-case planning check. Summary. Single-sample reuse induces structural dependence through the thresholding mechanism described above, primarily distorting predictive dispersion. LOO recalibration reduces self-influence and closely tracks audit-style behavior in our regimes. Predictive envelopes are then constructed from pooled LOO proxies using a Beta–Binomial model, with infl providing a monotone knob for controlled pessimization and the Hoeffding bound serving as a conservative guardrail. At this point, we cannot provide strong theoretical guarantees on the rates of the LOO proxy indicators, and therefore recommend using a proper audit set for certification. Appendix E Binary conformal partitions: regimes, coupling, and rate primitives This appendix records the geometric facts used in Section 2, Section 3, and Section 4. Here we assume thresholds are fixed and all probabilities are taken with respect to the deployment distribution conditional on calD_cal. E.1 Class-conditional split conformal as a four-region partition. Let =0,1Y=\0,1\ and let s(x,y)s(x,y) be a nonconformity score. Class-conditional split conformal produces thresholds τ0,τ1 _0, _1, and the set-valued output is (x)=0:s(x,0)≤τ0∪1:s(x,1)≤τ1,C(x)=\0:s(x,0)≤ _0\∪\1:s(x,1)≤ _1\, (2) equivalently represented by the region label Rτ(x):=(s(x,0)≤τ0, 1s(x,1)≤τ1)∈0,12,τ=(τ0,τ1).R_τ(x):= (1\s(x,0)≤ _0\,\ 1\s(x,1)≤ _1\ )∈\0,1\^2, τ=( _0, _1). Writing (s0,s1)=(s(x,0),s(x,1))(s_0,s_1)=(s(x,0),s(x,1)), the thresholds partition score space into Rτ(x)=11,s0≤τ0,s1≤τ1(doublet),10,s0≤τ0,s1>τ1(singleton 0),01,s0>τ0,s1≤τ1(singleton 1),00,s0>τ0,s1>τ1(abstention).R_τ(x)= cases11,&s_0≤ _0,\ s_1≤ _1 (doublet),\\ 10,&s_0≤ _0,\ s_1> _1 (singleton $\0\$),\\ 01,&s_0> _0,\ s_1≤ _1 (singleton $\1\$),\\ 00,&s_0> _0,\ s_1> _1 (abstention). cases For any fixed τ, ∑r∈00,01,10,11μr(τ)=1,μr(τ):=Pr(Rτ(X)=r∣cal). _r∈\00,01,10,11\ _r(τ)=1, _r(τ):= (R_τ(X)=r _cal). E.2 Probability-normalized scores and a sharp regime boundary Many probabilistic classifiers induce probability-normalized scores, e.g. s(x,y)=1−P(y∣x)s(x,y)=1-P(y x). For =0,1Y=\0,1\ this implies s(x,0)+s(x,1)=1,s(x,0)+s(x,1)=1, so feasible score pairs lie on the diagonal manifold ℳ=(u,1−u):u∈[0,1].M=\(u,1-u):u∈[0,1]\. Intersecting ℳM with the threshold rectangles yields a sharp boundary that determines which region types can occur with nonzero mass. Proposition 2 (Regime boundary under probability normalization). Assume (s0,s1)∈ℳ(s_0,s_1) almost surely. Then: 1. R11R_11 has nonempty intersection with ℳM iff τ0+τ1≥1 _0+ _1≥ 1, and has positive-length intersection iff τ0+τ1>1 _0+ _1>1. 2. R00R_00 has nonempty intersection with ℳM iff τ0+τ1≤1 _0+ _1≤ 1, and has positive-length intersection iff τ0+τ1<1 _0+ _1<1. 3. On the boundary τ0+τ1=1 _0+ _1=1, both R11R_11 and R00R_00 intersect ℳM at a single point; hence under any continuous distribution on ℳM they have probability zero and only the singleton regions carry mass. Proof. Parameterize ℳM by u=s0∈[0,1]u=s_0∈[0,1], so s1=1−us_1=1-u. R11R_11 requires u≤τ0u≤ _0 and 1−u≤τ11-u≤ _1, i.e. u∈[ 1−τ1,τ0]u∈[\,1- _1,\ _0\,]. This interval has positive length iff 1−τ1<τ01- _1< _0, i.e. τ0+τ1>1 _0+ _1>1, and degenerates to a point at equality. R00R_00 requires u>τ0u> _0 and 1−u>τ11-u> _1, i.e. u∈(τ0, 1−τ1)u∈(\, _0,\ 1- _1\,). This interval has positive length iff τ0<1−τ1 _0<1- _1, i.e. τ0+τ1<1 _0+ _1<1, and degenerates to a point at equality. ∎ Crossing the affine boundary τ0+τ1=1 _0+ _1=1 therefore removes an entire region label (doublet or abstention) from the support of Rτ(X)R_τ(X) under probability normalization, explaining sharp regime changes in attainable operating behavior. E.3 Cross-threshold dominance in the hedging regime Within the hedging regime τ0+τ1>1 _0+ _1>1, the manifold ℳM is partitioned into three contiguous intervals corresponding to 10,11,01\10,11,01\. With u=s(x,0)u=s(x,0): u∈[0, 1−τ1)⇒Rτ(x)=10,u∈[ 1−τ1,τ0]⇒Rτ(x)=11,u∈(τ0, 1]⇒Rτ(x)=01.u∈[0,\,1- _1) R_τ(x)=10, u∈[\,1- _1,\, _0\,] R_τ(x)=11, u∈( _0,\,1] R_τ(x)=01. Hence each singleton region is controlled primarily by the opposing threshold: increasing τ1 _1 expands 1111 at the expense of 1010, while increasing τ0 _0 expands 1111 at the expense of 0101. Operationally, (τ0,τ1)( _0, _1) act as mass-reallocation boundaries, not independent per-class knobs. Note that under probability normalization, setting thresholds to (1−τ1,1−τ0)(1- _1,1- _0) exchanges the roles of regions 1111 and 0000. If the deployment policy treats hedging and abstention identically, then this swap does not change action-level behavior, even though marginal coverage changes. E.4 Threshold classifiers from split conformal cutpoints When τ0=τ1=1/2 _0= _1=1/2, the induced classification rule is the argmax rule (with random tie-breaking), and coverage coincides with the classifier’s accuracy. Using probability-normalized scores s(x,y)=1−P(Y=y∣X=x)s(x,y)=1-P(Y=y X=x), the threshold condition s(x,y)≤1/2s(x,y)≤ 1/2 is equivalent to P(Y=y∣X=x)≥1/2P(Y=y X=x)≥ 1/2. Thus, with τ0=τ1=1/2 _0= _1=1/2, the conformal set includes exactly those labels whose posterior probability is at least 1/21/2: it returns 0\0\ when P(0∣x)>1/2P(0 x)>1/2, 1\1\ when P(1∣x)>1/2P(1 x)>1/2, and (only on ties) 0,1\0,1\ when P(0∣x)=P(1∣x)=1/2P(0 x)=P(1 x)=1/2. If we break ties at random to obtain a single-label predictor, this is precisely the argmax rule. Moreover, away from the tie set the conformal output is a singleton, so the event Y∈(X)\Y (X)\ is identical to Y=Y^(X)\Y= Y(X)\; hence the coverage probability equals the classification accuracy (with the same tie-breaking convention). The τ0+τ1=1 _0+ _1=1 family and asymmetric Bayes costs. More generally, if τ0+τ1=1 _0+ _1=1 then abstention disappears (up to the tie set) and the conformal output is almost surely a singleton. Indeed, (x)C(x) includes label y iff P(Y=y∣X=x)≥1−τyP(Y=y X=x)≥ 1- _y, so τ0+τ1=1 _0+ _1=1 implies 1−τ0=τ11- _0= _1 and 1−τ1=τ01- _1= _0, yielding the one-parameter threshold rule Y^(x)=P(1∣x)≥τ0=P(0∣x)≤τ1. Y(x)=1\!\P(1 x)≥ _0\=1\!\P(0 x)≤ _1\. In this regime, Pr(Y∈(X))=Pr(Y=Y^(X)) (Y (X))= (Y= Y(X)), i.e., coverage equals the accuracy of the corresponding threshold classifier. The parameter τ0 _0 also has a standard decision-theoretic interpretation. Consider asymmetric misclassification costs, with c01c_01 the cost of a false positive (predicting 11 when Y=0Y=0) and c10c_10 the cost of a false negative (predicting 0 when Y=1Y=1). The Bayes rule that minimizes conditional expected cost predicts 11 whenever c01P(Y=0∣x)≤c10P(Y=1∣x)⟺P(1∣x)≥c01c01+c10.c_01\,P(Y=0 x)\;≤\;c_10\,P(Y=1 x) P(1 x)\;≥\; c_01c_01+c_10. Thus sweeping τ0 _0 over (0,1)(0,1) traces the familiar family of cost-sensitive Bayes classifiers, with τ0=c01/(c01+c10 _0=c_01/(c_01+c_10 corresponding to the cost ratio c01:c10c_01:c_10. E.5 Region–label primitives and rate factorizations For auditing and planning, the primitive object is the region–class label table pr,y(τ):=Pr(Rτ(X)=r,Y=y∣cal),r∈00,01,10,11,y∈0,1.p_r,y(τ):= (R_τ(X)=r,\ Y=y _cal), r∈\00,01,10,11\,\ y∈\0,1\. Two derived summaries are μr(τ) _r(τ) :=Pr(Rτ(X)=r∣cal)=∑y∈0,1pr,y(τ), := (R_τ(X)=r _cal)= _y∈\0,1\p_r,y(τ), ηr(τ) _r(τ) :=Pr(Y=1∣Rτ(X)=r,cal)=pr,1(τ)μr(τ)(μr(τ)>0). := (Y=1 R_τ(X)=r,\ D_cal)= p_r,1(τ) _r(τ) ( _r(τ)>0). Any region-associated KPI is a projection of pr,y(τ)\p_r,y(τ)\; see Section 2.4. Example: decisive error masses under commit-on-singletons. Consider the convention 10↦0,01↦1,11,00↦defer / reject.10 0, 01 1, 11,00 / reject. Then the decisive false-negative and false-positive masses are FNdec(τ)=Pr(Rτ(X)=10,Y=1∣cal)=p10,1(τ)=μ10(τ)η10(τ),FN_dec(τ)= (R_τ(X)=10,\ Y=1 _cal)=p_10,1(τ)= _10(τ)\, _10(τ), FPdec(τ)=Pr(Rτ(X)=01,Y=0∣cal)=p01,0(τ)=μ01(τ)[1−η01(τ)].FP_dec(τ)= (R_τ(X)=01,\ Y=0 _cal)=p_01,0(τ)= _01(τ)\,[1- _01(τ)]. The decisive mass is μ10(τ)+μ01(τ) _10(τ)+ _01(τ), while the defer mass is μ11(τ)+μ00(τ) _11(τ)+ _00(τ) (with feasibility of 1111 versus 0000 governed by Proposition 2 under probability normalization). Appendix F Conformal interface-relative decision optimality and inverse pricing envelopes for a fixed conformal interface This appendix formalizes the operational point that distribution-free calibration is not automatically cost-agnostic once conformal outputs are wired into actions. We specialize classical statistical decision theory and reject-option analysis (Casella and Berger, 2002; Chow, 1970; Herbei and Wegkamp, 2006; Yuan and Wegkamp, 2010; Bartlett and Wegkamp, 2008) to the conformal interface as the deployed observable. Fix thresholds τ (calibration-conditional viewpoint) and treat the induced finite region label Rτ(X)∈ℛR_τ(X) as the deployed observable. Any downstream rule that uses only this interface can depend on the data only through r∈ℛr , so whether the convention is decision-optimal relative to the conformal interface under a given cost model is evaluated region-wise using within-region label frequencies (we formalize this below as conformal interface-relative decision optimality). The underlying concept is simple: among actions available after observing only Rτ(X)R_τ(X), the deployed convention should minimize region-wise conditional expected loss. We cast this as an inverse pricing problem: given a fixed (interface, convention) pair, characterize the set of consequence prices under which the convention is decision-optimal relative to the conformal interface. The main point is clear: once a conformal predictor is treated as a decision interface, decision-optimal downstream action depends on the region-wise label frequencies, not on coverage alone and not on the set-valued output structure alone. The contribution of this appendix is not new decision theory; it is to make that dependence explicit for a fixed conformal interface through a worked Chow-style case. F.1 Interface primitives: masses and within-region label frequencies Specialize to =0,1Y=\0,1\ and ℛ=00,01,10,11R=\00,01,10,11\ as in Appendix E. Adopt the calibration-conditional joint table pr,y:=ℙ(Rτ(X)=r,Y=y∣cal).p_r,y:=P(R_τ(X)=r,\ Y=y _cal). Define region mass and within-region label frequency μr:=pr,0+pr,1,ηr:=Pr(Y=1∣Rτ(X)=r,cal)=pr,1μr(μr>0). _r:=p_r,0+p_r,1, _r:= (Y=1 R_τ(X)=r,\ D_cal)= p_r,1 _r ( _r>0). Because ℛR is finite, conditional uncertainty about Y is piecewise constant: within region r, all decision comparisons reduce to ηr _r. F.2 Region-associated action conventions and decision optimality Let A be a finite action set and let Lθ(a,y)L_θ(a,y) be a priced consequence (cost or negative utility), parameterized by θ∈Θθ∈ . A deployed convention is a region-associated policy π~:ℛ→. π:R . For region r with μr>0 _r>0, the interface-relative conditional risk of action a is θ(a∣r):=[Lθ(a,Y)∣Rτ(X)=r,cal]=(1−ηr)Lθ(a,0)+ηrLθ(a,1).C_θ(a r):=E[L_θ(a,Y) R_τ(X)=r,\ D_cal]=(1- _r)L_θ(a,0)+ _rL_θ(a,1). Conformal interface-relative decision optimality. We say π~ π is decision-optimal relative to the conformal interface under pricing θ if for every region with μr>0 _r>0, θ(π~(r)∣r)≤θ(a∣r)∀a∈.C_θ( π(r) r) _θ(a r) ∀ a . This is precisely decision optimality relative to the coarsened observable Rτ(X)R_τ(X): the action chosen in each region minimizes posterior expected loss among the admissible actions available at that interface. F.3 Inverse pricing envelope For each region with μr>0 _r>0, define the local feasibility set Θr(π~):=θ∈Θ:θ(π~(r)∣r)≤θ(a∣r)∀a∈, _r( π):= \θ∈ :\ C_θ( π(r) r) _θ(a r)\ \ ∀ a \, and the global pricing envelope Θ(π~):=⋂r:μr>0Θr(π~). ( π):= _r: _r>0 _r( π). Equivalently, Θ(π~) ( π) is the set of price parameters for which the forced action in each region is decision-optimal relative to the conformal interface given only the conformal output. Since only action comparisons matter, envelopes are naturally reported in ratio (projective) coordinates: global positive scaling of LθL_θ is irrelevant, and adding offsets byb_y independent of a preserves comparisons. F.4 Worked case: Chow-style reject option and commit on singletons Consider =0,1,rejA=\0,1,rej\ with Chow-style costs (false negative c01>0c_01>0, false positive c10>0c_10>0, rejection crej≥0c_rej≥ 0): L(0,1)=c01,L(1,0)=c10,L(rej,y)=crej,L(0,0)=L(1,1)=0.L(0,1)=c_01, L(1,0)=c_10, L(rej,y)=c_rej, L(0,0)=L(1,1)=0. In region r the conditional risks are (0∣r)=ηrc01,(1∣r)=(1−ηr)c10,(rej∣r)=crej.C(0 r)= _r\,c_01, (1 r)=(1- _r)\,c_10, (rej r)=c_rej. Work in ratios λ:=c01c10,ρ:=crejc10,λ:= c_01c_10, ρ:= c_rejc_10, so only (λ,ρ)(λ,ρ) matters for comparisons. Convention. π~(10)=0,π~(01)=1,π~(11)=rej,π~(00)=rej. π(10)=0, π(01)=1, π(11)=rej, π(00)=rej. This convention is decision-optimal relative to the conformal interface iff the region-wise dominance inequalities below hold for all regions with μr>0 _r>0. Singleton region 0101 (output 1\1\). Choosing 11 must beat both 0 and rejrej: (1−η01)≤η01λ⟺η01≥11+λ,(1−η01)≤ρ⟺η01≥1−ρ.(1- _01)≤ _01λ _01≥ 11+λ, (1- _01)≤ρ _01≥ 1-ρ. Thus η01≥max11+λ, 1−ρ. _01\ ≥\ \! \ 11+λ,\ 1-ρ \. Singleton region 1010 (output 0\0\). Choosing 0 must beat both 11 and rejrej: η10λ≤(1−η10)⟺η10≤11+λ,η10λ≤ρ⟺η10≤ρλ. _10λ≤(1- _10) _10≤ 11+λ, _10λ≤ρ _10≤ ρλ. Thus η10≤min11+λ,ρλ. _10\ ≤\ \! \ 11+λ,\ ρλ \. Rejection regions r∈11,00r∈\11,00\. Rejecting must beat both commitments: ρ≤ηrλ⟺ηr≥ρλ,ρ≤(1−ηr)⟺ηr≤1−ρ.ρ≤ _rλ _r≥ ρλ, ρ≤(1- _r) _r≤ 1-ρ. Hence rejection is optimal on region r only if ρλ≤ηr≤ 1−ρ. ρλ\ ≤\ _r\ ≤\ 1-ρ. The rejection band is nonempty only if ρλ≤1−ρ⟺ρ≤λ1+λ⟺crej≤c01c10c01+c10. ρλ≤ 1-ρ ρ≤ λ1+λ c_rej≤ c_01c_10c_01+c_10. Polyhedral summary. The conditions above can be summarized as a pricing envelope in (λ,ρ)(λ,ρ)-space. Because they are linear inequalities, the admissible cost set is a convex polytope (intersection of half-spaces), as is standard in reject-option decision theory (Yuan and Wegkamp, 2010). For the convention π~(10)=0 π(10)=0, π~(01)=1 π(01)=1, π~(11)=π~(00)=rej π(11)= π(00)=rej, the envelope is the set of (λ,ρ)∈ℝ>02(λ,ρ) ^2_>0 satisfying λ λ ≥1−η01η01[singleton 01: commit to 1 beats 0], ≥ 1- _01 _01 [singleton $01$: commit to $1$ beats $0$], (3) λ λ ≤1−η10η10[singleton 10: commit to 0 beats 1], ≤ 1- _10 _10 [singleton $10$: commit to $0$ beats $1$], (4) ρ ρ ≥1−η01[singleton 01: commit beats reject], ≥ 1- _01 [singleton $01$: commit beats reject], (5) ρ ρ ≥λη10[singleton 10: commit beats reject], ≥λ _10 [singleton $10$: commit beats reject], (6) ρ ρ ≤ληr,ρ≤1−ηr[rejection regions r∈11,00 with μr>0], ≤λ _r, ρ≤ 1- _r [rejection regions $r∈\11,00\$ with $ _r>0$], (7) together with λ>0λ>0 and ρ≥0ρ≥ 0. F.5 Follow-up observations Monotone ordering. A useful corollary of the inequalities is that the pricing envelope is non-empty only when the rejection regions have intermediate label composition: η10≤η01 _10≤ _01 and, for each populated r∈11,00r∈\11,00\, η10≤ηr≤η01 _10≤ _r≤ _01. This is the region-wise analogue of the classical reject-option condition from Chow (1970); Herbei and Wegkamp (2006): rejection is decision-optimal only when the posterior falls in an intermediate band. Coverage-matched settings. Holding the scoring model and deployment distribution fixed, different calibration settings τ and τ′τ can have the same marginal coverage but different region compositions, and therefore different pricing envelopes. In some cases those envelopes may even fail to overlap. The point is not tied to different datasets; it is a statement about how the same underlying interface can behave under different calibration choices. This is the conformal-specific implication of the preceding decision-theoretic setup: coverage-matched rules may still be decision-optimal under different parts of cost-ratio space. Section 5.3 and Figure 5 show the milder empirical version of the same idea, namely that the feasible (λ,ρ)(λ,ρ) region varies across Pareto regimes. Dual view. The same inequalities can be read in reverse: given (λ,ρ)(λ,ρ), which action convention is decision-optimal relative to the conformal interface? This partitions cost-ratio space into regions where different conventions are decision-optimal—a standard dual perspective in cost-sensitive classification (Bartlett and Wegkamp, 2008; Yuan and Wegkamp, 2010). For the present appendix, this is best read as an interpretive lens on the same region-wise quantities rather than as a separate result. Summary. The main point of this appendix is the same one stated at the start: for a fixed calibrated conformal predictor and a fixed wiring convention π~ π, decision optimality relative to the conformal interface depends on the region-wise label frequencies through the posteriors ηr\ _r\. Coverage alone is not enough, and the bare prediction-set structure is not enough either. The worked Chow-style case, together with the follow-up observations on monotone ordering (F.5) and coverage-matched settings (F.5), simply makes that dependence explicit for the conformal interface and shows how the region–class label table restricts which cost ratios make a given downstream convention decision-optimal. Appendix G Explicit region indicators and projection masks This appendix briefly instantiates the linear “sum selected cells” formalism from Section 2.4 in the binary geometry used in Figure 2 and the policy-projection example from the main text (Rτ(x)→C(x)R_τ(x) \ π\ C(x)). Region feasibility and regime facts are recorded in Appendix E; here we keep only the explicit projection masks needed for audit computations. G.1 Region labels and the region–class label table Assume =0,1Y=\0,1\ and thresholds τ=(τ0,τ1)τ=( _0, _1). Define the deployed region label Rτ(x)=(r0(x),r1(x))∈0,12,ry(x):=s(x,y)≤τy,R_τ(x)=(r_0(x),r_1(x))∈\0,1\^2, r_y(x):=1\s(x,y)≤ _y\, with region names ℛ=r10,r11,r01,r00R=\r_10,r_11,r_01,r_00\ as in Appendix E.1. For a calibration setting θ (indexing deployed thresholds τ(θ)τ(θ)), define the calibration-conditional region–class label table P(θ)=(pr,y(θ))r∈ℛ,y∈,pr,y(θ)=Pr(Rτ(θ)(X)=r,Y=y∣cal).P(θ)= (p_r,y(θ) )_r ,\ y , p_r,y(θ)= \! (R_τ(θ)(X)=r,\ Y=y _cal ). A linear operational rate is obtained by summing the subset of cells (r,y)(r,y) that define the event. G.2 Policies used in Figure 3 A policy π maps region labels to prediction sets C(x)⊆0,1C(x) \0,1\: Rτ(x)→C(x).R_τ(x)\ \ π\ \ C(x). The three policies used in Figure 3 are: Set inclusion πSI _SI. πSI(r10)=0,πSI(r11)=0,1,πSI(r01)=1,πSI(r00)=∅. _SI(r_10)=\0\, _SI(r_11)=\0,1\, _SI(r_01)=\1\, _SI(r_00)= . Commit–reject πCR _CR. πCR(r10)=0,πCR(r01)=1,πCR(r11)=∅,πCR(r00)=∅. _CR(r_10)=\0\, _CR(r_01)=\1\, _CR(r_11)= , _CR(r_00)= . Set exclusion πSE _SE (complement of set inclusion). πSE(r10)=1,πSE(r11)=∅,πSE(r01)=0,πSE(r00)=0,1. _SE(r_10)=\1\, _SE(r_11)= , _SE(r_01)=\0\, _SE(r_00)=\0,1\. G.3 Binary projection masks Fix a policy π. For any event/quantity ℓ , define a 4×24× 2 binary mask Gℓ(π)=(gℓ(r,y;π))r∈ℛ,y∈0,14×2,G_ (π)= (g_ (r,y;π) )_r ,\ y ∈\0,1\^4× 2, where gℓ(r,y;π)=1g_ (r,y;π)=1 indicates that the cell (r,y)(r,y) is included in the sum. Then the corresponding rate is rℓ(θ)=∑r∈ℛ∑y∈gℓ(r,y;π)pr,y(θ).r_ (θ)= _r _y g_ (r,y;π)\,p_r,y(θ). We record one detailed worked example and then list several shorter special cases used elsewhere in the paper. (1) Coverage under set inclusion. Under πSI _SI, coverage is the event Y∈C(X)\Y∈ C(X)\. Cell-by-cell, coverage holds for: (i) y=0y=0 in regions r10r_10 or r11r_11, and (i) y=1y=1 in regions r01r_01 or r11r_11. Hence Gcov(πSI)=y=0y=1r1010r1111r0101r0000G_cov( _SI)= array[]c|c&y=0&y=1\\ r_10&1&0\\ r_11&1&1\\ r_01&0&1\\ r_00&0&0 array and therefore rcov(θ)=p10,0(θ)+p11,0(θ)+p11,1(θ)+p01,1(θ).r_cov(θ)=p_10,0(θ)+p_11,0(θ)+p_11,1(θ)+p_01,1(θ). Equivalently, coverage fails on the two singleton mistakes (r10,1)(r_10,1) and (r01,0)(r_01,0) and on all abstentions r00r_00. (2) Conformal Efficiency In order to characterize the amount of uncertainty carried by a conformal predictor, the concept of conformal efficiency (Heyndrickx et al., 2023) has been proposed. Conformal efficiency is essentially defined as the total singleton rate. The region mask for conformal efficiency is obtained by summing the cells that define the event: GCE(πSI)=y=0y=1r1011r1100r0111r0000G_CE( _SI)= array[]c|c&y=0&y=1\\ r_10&1&1\\ r_11&0&0\\ r_01&1&1\\ r_00&0&0 array and therefore rCE(θ)=p10,0(θ)+p10,1(θ)+p01,0(θ)+p01,1(θ).r_CE(θ)=p_10,0(θ)+p_10,1(θ)+p_01,0(θ)+p_01,1(θ). Under πSI _SI, coverage holds automatically on hedged outputs (r11r_11), while coverage failures occur only on singleton mistakes and abstentions. In particular, rcov(θ)=p10,0(θ)+p11,0(θ)+p11,1(θ)+p01,1(θ),r_cov(θ)=p_10,0(θ)+p_11,0(θ)+p_11,1(θ)+p_01,1(θ), so that rcov(θ)−rCE(θ)=(p11,0(θ)+p11,1(θ))−(p10,1(θ)+p01,0(θ)).r_cov(θ)-r_CE(θ)= (p_11,0(θ)+p_11,1(θ) )- (p_10,1(θ)+p_01,0(θ) ). Thus, when calibration places substantial mass in the hedging region r11r_11, coverage can exceed conformal efficiency because hedged prediction sets are always covered. If the calibration is performed in the abstention regime (r00r_00), then coverage and conformal efficiency differ by the total error mass among singleton predictions. (3) Other projections used in the paper. The same mask formalism gives the remaining quantities used in Figure 1 and in the binary operational examples: missed positive mass:q10(θ)=Pr(Y=1,C(X)=0)=Pr(Y=1,Rτ(X)=r10)=p10,1(θ),missed positive mass: q_10(θ)= (Y=1,\ C(X)=\0\)= (Y=1,\ R_τ(X)=r_10)=p_10,1(θ), hedged positive mass:q11(θ)=Pr(Y=1,C(X)=0,1)=Pr(Y=1,Rτ(X)=r11)=p11,1(θ),hedged positive mass: q_11(θ)= (Y=1,\ C(X)=\0,1\)= (Y=1,\ R_τ(X)=r_11)=p_11,1(θ), abstention mass under πCR:rabs(θ;πCR)=(p11,0(θ)+p11,1(θ))+(p00,0(θ)+p00,1(θ)).abstention mass under _CR: r_abs(θ; _CR)= (p_11,0(θ)+p_11,1(θ) )+ (p_00,0(θ)+p_00,1(θ) ). These correspond respectively to a one-cell projection, another one-cell projection, and a two-row projection. Their binary masks are obtained immediately by placing ones on the selected cells. G.4 Ratios of projections (conditional diagnostics) Some diagnostics are conditional probabilities and therefore ratios of linear sums. For example, the purity of singleton-11 outputs under πSI _SI is Purity1(θ):=Pr(Y=1∣C(X)=1)=Pr(Y=1,C(X)=1)Pr(C(X)=1)=p01,1(θ)p01,0(θ)+p01,1(θ).Purity_1(θ):= \! (Y=1 C(X)=\1\ )= (Y=1,\ C(X)=\1\) (C(X)=\1\)= p_01,1(θ)p_01,0(θ)+p_01,1(θ). Audit computation. All entries of P(θ)P(θ) are estimated from region–class label counts on the audit set. Linear KPIs are computed by summing selected cells; conditional diagnostics are computed as ratios of such sums. Appendix H Tox21 supplementary details This appendix provides supplementary documentation for the Tox21 experiments reported in Section 5.2. The purpose of this appendix is reproducibility and contextualization rather than extension of results or additional claims. H.1 Tox21 dataset The Tox21 benchmark consists of binary toxicity outcomes for twelve biological assays spanning nuclear receptor (NR) signaling and cellular stress response (SR) pathways. Each compound is labeled as active or inactive per assay. Labels are sparse and highly imbalanced, particularly for nuclear receptor targets. The dataset composition table 4 and aggregated coverage summary 5 are recorded here. Table 4: Tox21 dataset composition and effective average class-conditional calibration sizes under the experimental protocol. Endpoint Total Samples Positive Rate Calib. Positives Calib. Negatives NR-AR 7265 4.3% 77 1739 NR-AR-LBD 6758 3.5% 59 1630 NR-AhR 6549 11.7% 192 1445 NR-Aromatase 5821 5.2% 75 1380 NR-ER 6193 12.8% 198 1350 NR-ER-LBD 6955 5.0% 87 1651 NR-PPAR-γ 6450 2.9% 46 1566 SR-ARE 5832 16.2% 235 1223 SR-ATAD5 7072 3.7% 66 1702 SR-HSE 6467 5.8% 93 1523 SR-MMP 5810 15.8% 229 1223 SR-p53 6774 6.2% 105 1587 Table 5: Aggregated Tox21 coverage and set statistics. Results are averaged over twelve endpoints and 100 random splits. Method Mean Coverage Violation Rate Avg. Set Size Singleton Rate Standard Split Conformal 0.917 0.305 1.41 0.52 DKWM Correction 0.986 0.005 1.78 0.22 SSBC 0.951 0.068 1.54 0.40 H.2 Representative endpoint-level operational summaries The main text reports SR-MMP as a representative moderate-prevalence endpoint. We keep NR-AR here to document a complementary low-prevalence regime. Table 6: NR-AR endpoint. Operational performance summary for the NR-AR endpoint. All reported quantities are joint probabilities normalized by the total test-set size. Rows report the singleton rate, doublet rate, and wrong-singleton rate (P(Y=c,|S|=1,y^≠Y)P(Y=c,\,|S|=1,\, y≠ Y)) by true class. "1D_1" and "LOO 95% PI 1D_1" denote leave-one-out point estimates and their beta–binomial planning intervals computed on the calibration data. "2D_2" reports empirical rates on the audit set. "B 95% PI 2D_2" gives Beta–Binomial predictive summaries for the corresponding 2D_2 quantities over a future window of the same size. Operational quantity Class 1D_1 LOO 95% PI 1D_1 2D_2 B 95% PI 2D_2 Singleton rate Class 0 0.1630.163 [0.125,0.205][0.125,0.205] 0.1330.133 [0.112,0.157][0.112,0.157] Class 1 0.0260.026 [0.011,0.047][0.011,0.047] 0.0290.029 [0.019,0.041][0.019,0.041] Doublet rate Class 0 0.7960.796 [0.751,0.838][0.751,0.838] 0.8180.818 [0.792,0.843][0.792,0.843] Class 1 0.0150.015 [0.005,0.031][0.005,0.031] 0.0200.020 [0.012,0.030][0.012,0.030] Wrong-singleton rate Class 0 0.0860.086 [0.058,0.119][0.058,0.119] 0.0600.060 [0.045,0.077][0.045,0.077] Class 1 0.0020.002 [0.000,0.010][0.000,0.010] 0.0010.001 [0.000,0.004][0.000,0.004] Because conformal calibration is performed in a class-conditional (Mondrian) manner, the effective calibration size for the positive class governs feasibility and discretization effects. H.3 Representation and model protocol All Tox21 experiments use a fixed, task-agnostic molecular representation (descriptors plus Morgan fingerprints) and a fixed CatBoost training protocol (1000 iterations, depth 6, learning rate 0.1, log-loss), with no assay-specific feature engineering, class reweighting, or resampling. Molecules are parsed and sanitized using standard cheminformatics tooling, and compounds with invalid features are excluded. This intentionally non-optimized setup isolates conformal calibration and operational-envelope behavior from model-architecture tuning; observed AUROC values range between 0.80–0.85 across endpoints, and are consistent with baseline Tox21 classifiers. H.4 Conformal calibration and evaluation All conformal predictors are calibrated in a class-conditional (Mondrian) fashion, with separate calibration sets for the positive and negative classes. For each run, one of three calibration strategies is applied: standard split conformal, DKWM-based correction, or SSBC, targeting (α,δ)=(0.10,0.10)(α,δ)=(0.10,0.10). Calibration thresholds are computed once per run and evaluated on an independent held-out split. In the main text tables, the “ 2D_2” columns report the corresponding endpoint-level operational rates on that held-out split, which serve as the audit-based reference for the operational claims. The “1D_1” and “LOO 95% PI 1D_1” columns are single-sample planning summaries computed from leave-one-out recalibration on the calibration split. They are included to compare the LOO planning proxy to the independent audit reference. Appendix I Solubility supplementary details This appendix records methodological details for the solubility scenario-planning experiments in Section 5.3. We first describe the dataset, partitioning, modeling, and calibration setup, and then document the operational artifacts used to interpret the resulting Pareto front. I.1 Dataset and label construction We use AquaSolDB as the source of aqueous solubility measurements (logS S), following the curation and quality-control procedures described in the dataset reference (Sorkun et al., 2019). The curated source file used in our experiments contains 9,982 entries. Molecules with invalid SMILES strings are excluded before feature generation. For model training and scoring we discretize logS S into three regimes with fixed thresholds: Insoluble (logS<−4 S<-4), Moderate (−4≤logS<−2-4≤ S<-2), and Soluble (logS≥−2 S≥-2), defining a three-class label space =0,1,2Y=\0,1,2\ (Kalepu and Nekkanti, 2015). I.2 Partitioning: scaffold-based training and stratified calibration/test split We use a hybrid split to separate scaffold-aware model fitting from the calibration pool used in the LOO planning analysis: • Scaffold-based training split. Molecules are grouped by Bemis–Murcko scaffold. The scaffold groups are shuffled with a fixed seed, and groups are added to the training split until roughly 70% of molecules are assigned. This makes the training split scaffold-aware rather than purely random. • Stratified calibration/test split. The remaining molecules are split equally between calibration and test using stratified random sampling over the three-class target. • Calibration/test split for diagnostics. The calibration/test split is retained for exploratory diagnostics and documentation of the scenario construction. The reported planning quantities in Section 5.3 are not obtained from an independent audit/test evaluation; they are obtained from the leave-one-out planning interface on the scenario-restricted calibration sample. I.3 Molecular representation Molecules are represented using concatenated multi-resolution Morgan fingerprints (Rogers and Hahn, 2010) computed with RDKit. We use three fingerprint tiers: radius 4 with 512 bits, radius 3 with 1024 bits, and radius 2 with 2048 bits, yielding 3584 binary features per molecule. Molecules that fail RDKit parsing are excluded before feature generation. I.4 Model and training protocol We train a CatBoost (Prokhorenkova et al., 2018) gradient-boosted decision tree classifier on the scaffold-based training split with fixed hyperparameters: 1000 boosting iterations, depth 6, learning rate 0.05, random seed 42, and multi-class loss. After training, the predictor is frozen and treated as infrastructure; the experiments study the conformal and operational layers rather than optimizing base accuracy. I.5 Calibration and operationalization Uncertainty quantification uses split conformal prediction with SSBC calibration (Section 5.1.1 and Appendix C), which selects a grid index to stabilize realized coverage semantics at user-specified confidence (Vovk, 2012b). In the solubility case study, all reported planning quantities are implemented through the leave-one-out (LOO) construction described in Appendix D, rather than through an independent audit split. This calibration layer therefore supports the finite-window planning-envelope constructions used for operational planning. The calibration/test split is retained as supporting context for the scenario construction, but the reported trade-off maps and planning quantities are generated from the LOO interface on the scenario-restricted calibration sample. Although the classifier is trained on three classes (Insoluble, Moderate, Soluble), the planning analysis in Section 5.3 uses a binary operational objective obtained by merging Moderate and Soluble into a single Soluble class. Conformal prediction sets are computed in the binary label space Insol,Sol\Insol,Sol\ and operational quantities (loss, waste, hedging, decisiveness) are computed with respect to this binary interface. I.6 Deployment-matched calibration via chemical tribes Scenario planning conditions exchangeability on a deployment-defining event by restricting calibration to a chemically defined subpopulation. Tribe construction. We define coarse chemical “tribes” using RDKit MolLogPMolLogP: Tribe_Lipophilic: Tribe\_Lipophilic: MolLogP>3.5, \ MolLogP>5, Tribe_Hydrophilic: Tribe\_Hydrophilic: MolLogP<1.0, \ MolLogP<0, Tribe_Neutral: Tribe\_Neutral: 1.0≤MolLogP≤3.5. 10 ≤ 5. We compute summary statistics of nonconformity and class probabilities stratified by tribe to document and motivate the focal deployment regime used in the planning study. Focal deployment scenario. Section 5.3 focuses on a lipophilic deployment regime, targeting molecules that are not naturally hydrophilic. SSBC calibration is therefore performed on the restricted sample of lipophilic molecules: callip=(Xi,Yi)∈cal:MolLogP(Xi)>3.5,D_cal^lip=\(X_i,Y_i) _cal:MolLogP(X_i)>3.5\, treated as exchangeable with respect to the intended deployment distribution. The cutoff MolLogP>3.5MolLogP>3.5 is used here as a domain-motivated scenario definition for lipophilic compounds, not as a tuned parameter chosen to optimize the observed trade-off map. After down-selection to this scenario and merging Moderate and Soluble into the binary soluble label, the resulting planning dataset contains 319 entries: 53 soluble and 266 insoluble. These counts describe the scenario dataset over which the planning sweep is performed; they do not correspond to a single selected operating point. Scenario-conditional exchangeability assumption. This restriction is scenario conditioning, not a general covariate-shift correction. We assume deployment draws satisfy the same scenario event E=MolLogP(X)>3.5E=\MolLogP(X)>3.5\ so calibration and deployment are exchangeable conditional on E. If deployment violates E, or if exchangeability fails within E, conformal validity is not guaranteed. Shift-robust validity methods are complementary and outside scope; see, e.g., Fannjiang et al. (2022). Interpretation. Conditioning trades calibration sample size for improved deployment match; the resulting loss of statistical efficiency appears directly as stronger SSBC corrections and wider planning envelopes, consistent with transparent planning under finite-sample uncertainty. I.7 Operational outcome definitions Operational outcomes are defined in the main text via the binary prediction-set interface (TP, FN, HP, FP, TN, HN). All reported quantities in the scenario-planning analysis are joint rates over the set–label outcomes P(Y=y,C=c)P(Y=y,\,C=c), scaled to expected counts over a window of m=1000m=1000 molecules. For this case study, these quantities are estimated using pooled LOO indicators and the associated planning-envelope construction. They are planning summaries for the restricted scenario sample and exhibit the structure of feasible trade-offs under the chosen scenario definition. I.8 Additional Pareto-front details We record the representative Pareto-optimal regimes here and then summarize the appendix-specific details behind the solubility Pareto sweep, in particular how the nominal SSBC parameters function along the front and how to interpret α versus δ once the scenario-restricted planning interface has been fixed. Table 7: Two representative Pareto-optimal operating regimes for the solubility scenario, illustrating contrasting planning postures. Rates are expected counts per 1000 molecules, with planning envelopes in brackets. Loss-minimizing regime High-decisiveness regime α0 _0 0.125 0.150 δ0 _0 0.150 0.150 α1 _1 0.050 0.150 δ1 _1 0.075 0.100 Rate Interval Rate Interval Loss rate 3 [0, 35] 12 [0, 58] Waste rate 70 [27, 130] 13 [0, 49] Total hedge rate 785 [663, 896] 514 [376, 675] Correct soluble singleton rate 7 [0, 38] 88 [40, 155] Correct insoluble singleton rate 135 [68, 215] 282 [196, 375] Decisiveness (1000−hedge1000-hedge) 215 486 I.9 Functional roles of calibration parameters The Pareto sweep in Section 5.3 varies the four nominal parameters (α0,δ0,α1,δ1)( _0, _0, _1, _1) (with SSBC mapping them to effective deployed grid levels). Empirically the knobs are not interchangeable: changes reallocate mass among outcome categories (loss, waste, hedging, decisiveness) along channels constrained by the underlying threshold geometry. Throughout, class 0 denotes insoluble and class 1 denotes soluble. The qualitative roles below summarize matched Pareto-optimal solutions where one parameter varies while the others are held fixed; rates are computed from the provided Pareto-front CSV and reported as events per 1000 molecules. Role of α0 _0: global conservatism against irreversible loss. Decreasing α0 _0 suppresses loss by expanding hedging; increasing α0 _0 collapses ambiguity into more decisive singleton predictions. In many matched segments, the dominant effect is redistribution between hedged and singleton outcomes with loss already near saturation. Role of δ0 _0: fine-scale sharpening on the insoluble side. For fixed (α0,α1,δ1)( _0, _1, _1), δ0 _0 tends to act locally, shifting mass between hedged and singleton insoluble predictions with limited impact on loss. Role of α1 _1: boundary-controlled tolerance with nonlocal effects. Although α1 _1 is nominally associated with the soluble side, changing α1 _1 can move the operating point across geometric boundaries, inducing nonlocal reallocations that may appear most strongly in insoluble outcomes. Role of δ1 _1: systematic hedge-to-insoluble reassignment. Across stable regions of the front, increasing δ1 _1 often converts a portion of hedged mass into insoluble singletons, reflecting how the normalized score constraints and threshold geometry position the hedging interval. Geometric interpretation. Overall, the Pareto front is shaped by geometric coupling rather than independent tuning: each knob reallocates probability mass along constrained channels determined by the threshold partition and the finite-sample SSBC adjustment. I.10 Interpreting α versus δ on the Pareto front Within SSBC the deployed threshold corresponds to an adjusted effective level α~ α determined jointly by (α,δ)(α,δ). Thus α and δ are best viewed as two parameterizations of a single underlying control (the effective threshold), with different practical resolution under finite-sample constraints. The approximate feasibility relation δ≈(1−α)Nδ≈(1-α)^N implies logδ=Nlog(1−α), δ=N (1-α), so changes in δ correspond to fine-grained (approximately logarithmic) adjustments in α~ α, whereas changes in α on a coarse grid can move the operating point across qualitatively different regions of the feasible manifold. This helps explain why α sweeps can trigger large reallocations among loss/waste/hedging outcomes, while δ sweeps more often act as local “sharpening” of decisions within a geometric regime.