Paper deep dive
Revelation Control
Qinyou Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/26/2026, 4:53:31 AM
Summary
The paper introduces 'Revelation Control,' a decision-theoretic framework for selecting priced interventions in learning systems that reveal hidden state only when it changes consequential decisions, while accounting for the productive value of the intervention itself. It defines decision-sufficient revelation, separates pure information value from productive reuse, and provides a cost-adjusted factorization criterion. Empirical validation on Qwen2.5-7B and Mistral-7B-v0.3 models demonstrates that deeper future-learning probes have positive decision value and that productive reuse yields strict equal-compute utility advantages.
Entities (7)
Relation Signals (5)
Revelation Control → defines → Decision-sufficient revelation
confidence 95% · The framework defines decision-sufficient revelation and revelation depth
Qwen2.5-7B → validatedby → Revelation Control
confidence 92% · Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value
Mistral-7B-v0.3 → validatedby → Revelation Control
confidence 92% · Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value
Revelation Control → separates → Productive reuse
confidence 90% · separates pure information value from productive reuse
Productive reuse → yields → Equal-compute utility advantages
confidence 88% · productive reuse yields strict equal-compute utility advantages
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.
Tags
Links
- Source: https://arxiv.org/abs/2608.23860v1
- Canonical: https://arxiv.org/abs/2608.23860v1
Trouble viewing inline? Open PDF directly →
Full Text
147,348 characters extracted from source content.
Expand or collapse full text
Revelation Control Qinyou Wang Abstract Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop–flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific. 1 Introduction Sequential decision systems are routinely controlled through summaries of their current state. In learning systems these include present losses, held-out metrics, gradients, optimizer telemetry, and short diagnostic trajectories. Such summaries are decision-sufficient only when the distinctions they discard cannot change the action that should be taken. A learner may instead occupy execution states that look equivalent under the declared current information yet react differently to the same future training intervention. Our companion work, Fiber Fingerprints of Hidden Learning-State Dynamics [40], establishes the predictive premise within its declared scope: present behavior need not be a sufficient statistic for declared future learning. The decision theory developed here does not depend on that paper’s specific geometric machinery; it takes hidden decision-relevant state as the object to be revealed and asks the downstream question: When does hidden learning-state structure become actionable decision information? We call the resulting control problem Revelation Control: the controller chooses not only which action to take, but which future-learning interventions to instantiate, how much decision-relevant state to reveal, and which tested future to promote into execution. The distinction between prediction and decision is essential. A hidden variable can improve prediction without changing the optimal action; a refined observation can increase oracle Bayes value while a finite-sample learner fails to extract it; and a training probe can both reveal information and leave behind useful computation. A decision theory for future-learning probes must therefore keep information, execution technology, learnability, and cost separate. Once refinement depth is allowed to vary, the same value functional becomes a control law over how much future information to acquire. Partial training itself is not new. Freeze–Thaw Bayesian optimization uses partial learning curves to decide whether to pause or resume models [37]; learning-curve extrapolation can terminate weak runs before full training [7]; successive resource-allocation methods exploit intermediate performance to concentrate computation on promising candidates [18, 25]; Population Based Training reuses and adapts live training trajectories [17]; and DynaMiCS uses short domain-specific probes to estimate local cross-domain effects before selecting a constrained fine-tuning mixture [14]. Nor are Bayes value of information, Bayesian experimental design, active sensing, partial observability, or dual control new ideas [5, 26, 6, 15, 11, 21, 39]. The distinctive step here is to turn future learning itself into a decision-relative experiment on hidden learner state, and to exploit the fact that the tested path can remain useful computation. Although the empirical instantiation in this paper is model training, the underlying decision structure is broader: a priced, controlled probe can change a system, reveal otherwise hidden decision-relevant state through its response, and leave reusable progress if the tested path is selected. Section 15 develops this extension while keeping all empirical claims confined to the learning systems actually studied. 1.1 Contributions The paper makes five coupled contributions. 1. Decision-sufficient revelation. We characterize when refined information changes the Bayes decision quotient, derive exact binary aliasing value, and identify a local revelation depth at which currently hidden decision-relevant directions become observable to an admissible future-learning probe. 2. Productive revelation technology. We distinguish discard-and-restart probing from probe-and-promote revelation, derive the equal-budget depth geometry, and separate pure information gain from the value of having already executed part of the selected training path. 3. Adaptive revelation control. We prove that population Bayes refinement value is exactly the expectation of a conditional state-dependent revelation value. We then derive an exact cost-adjusted scalar-control gap: a second shallow coordinate is decision-nonredundant precisely when a positive-mass scalar fiber contains states on both sides of the priced Stop/Continue boundary. In the binary location–scale specialization this yields a critical-revealability crossing criterion, so one-dimensional control can remain optimal even when revealability varies conditionally but never changes the meta-action. 4. Learnability and certification boundaries. We separate oracle revelation value from approximation, estimation, acquisition, and promotion terms. For adaptive depth, we identify a sharp boundary: bounded stop–flip risk alone cannot certify positive expected utility under unrestricted severity; a severity, moment, or integrable-tail condition is the missing object. 5. Cross-model Transformer validation and a reusable instantiation procedure. In a first 7B Transformer family, fixed-depth revelation, structural decision refinement, productive reuse, equal-compute frontier advantage, and adaptive safe-compute behavior are supported on independent panels. In an independently instantiated Mistral-7B-v0.3 family, positive deeper-revelation value and productive equal-compute utility reproduce on an independent two-bank panel. Within the finite pre-existing continuation-architecture family, a scalar architecture fit only on separate development data retains positive familywise-adjusted lower bounds on the target panel. The two families therefore reproduce the same Revelation-Control structure, while Qwen provides evidence that extra shallow revealability information is decision-nonredundant at the declared compute price and Mistral is consistent with scalar decision sufficiency within the tested architecture family and resolution. 1.2 Scope and organization Scope of claims. The paper does not claim unrestricted dominance over every current-information policy, every probing algorithm, every model family, or every deployment cost model. The finite-library result is exactly that: finite-library. The structural refinement claim is conditional on a declared shallow ambiguity regime rather than an unrestricted Bayes comparison over complete shallow information. The comparative claim uses one strengthened DynaMiCS-style short-probe frontier under a common policy-visible update budget. Cross-model transport is structural rather than parametric: different model families may use different legal revealability proxies, fitted coefficients, or stopping thresholds, and the theory explicitly allows scalar decision sufficiency whenever no positive-mass scalar fiber crosses the cost-adjusted continuation boundary. Fully tail-robust population utility certification remains conditional on an explicit severity, moment, or tail class. Organization. Sections 2 and 3 identify the decision quotient and the depth at which hidden distinctions become decision sufficient. Section 4 develops productive fork–probe–promote execution. Section 5 closes static refinement into state-dependent control, proves the scalar-control factorization criterion, and gives the model-specific instantiation protocol. Section 6 separates oracle value from finite-sample extraction and certification; the related-work section then fixes the novelty boundary before the empirical instantiation. Sections 8, 9, 10, 11, 12 and 13 evaluate the theory in two Transformer families, and Section 14 summarizes the cross-model invariants and regime-specific differences. Section 15 then identifies prospective learning, computational, physical, scientific, operational, and high-stakes domains in which the same theory-first instantiation procedure may be useful. Limitations and conclusions follow; the appendices collect proofs, statistical inference, and experimental protocol details. 2 Decision value and decision factorization Let (Ω,ℱ,ℙ)( ,F,P) be a probability space, A a finite action set, and Qa∈L1Q_a∈ L^1 the terminal utility of action a. For an information sigma-field ℐ⊆ℱI , define the observation-relative Bayes value VB(ℐ)=[maxa∈[Qa∣ℐ]].V_B(I)=E\! [ _a E[Q_a ] ]. (2.1) If ℋ⊆ H G, then Jensen’s inequality for the finite maximum gives the standard refinement monotonicity ℐ(∣ℋ):=VB()−VB(ℋ)≥0.I( G H):=V_B( G)-V_B( H)≥ 0. (2.2) For an event A∈ℋA∈ H with ℙ(A)>0P(A)>0, define VB(ℐ,A)=[maxa∈[Qa∣ℐ]|A],V_B(I;A)=E\! [ _a E[Q_a ]\, |\,A ], (2.3) and for a measurable policy π, write (π,A)=[Qπ∣A]V(π;A)=E[Q_π A] and (π)=[Qπ]V(π)=E[Q_π]. 2.1 Decision factorization Definition 2.1 (Decision factorization). For ℋ⊆ H G, the G-optimal decision factorizes through ℋ H if there exists an ℋ H-measurable policy πH _H satisfying πH(ω)∈argmaxa∈[Qa∣](ω)a.s. _H(ω)∈ _a E[Q_a G](ω) .s. (2.4) The definition concerns the action-relevant quotient of the refined information, not reconstruction of the full latent state. Proposition 2.2 (Exact finite-action strictness criterion). For finite A and integrable utilities, VB()=VB(ℋ)V_B( G)=V_B( H) (2.5) if and only if the G-optimal decision factorizes through ℋ H. Hence strict refinement value occurs exactly when no ℋ H-measurable policy is G-Bayes optimal almost surely. This criterion is deliberately decision-relative. A refinement can contain predictive information while having zero value for the declared action menu; conversely, only a small quotient of a very high-dimensional latent state may be required to change the optimal action. This perspective is compatible with classical comparison of experiments and value-of-information theory [5, 15, 4], but our later use is dynamic because the act of acquiring future-learning information can also create reusable computation. 2.2 Dynamic decision-factorization obstruction At time t, let νt(a) _t(a) be the current coarse-information action value before accounting for unresolved downstream distinctions, and let νt⋆=maxaνt(a) _t = _a _t(a). Define the current opportunity cost dt(a)=νt⋆−νt(a)≥0.d_t(a)= _t - _t(a)≥ 0. (2.6) Let Lt≥0L_t≥ 0 be the immediate value lost because the coarse representation aliases current decision-relevant states, and let ωt(a)≥0 _t(a)≥ 0 be the continuation obstruction that remains after taking action a. With continuation factor γ≥0γ≥ 0, define the total dynamic obstruction by comparing the best continuation-aware action with the coarse current benchmark: tDF:=Lt+maxa∈νt(a)+γωt(a)−νt⋆.O_t DF:=L_t+ _a \ _t(a)+γ _t(a)\- _t . (2.7) Proposition 2.3 (Dynamic decision-factorization decomposition). The obstruction admits the exact decomposition tDF=Lt+maxa∈γωt(a)−dt(a). O_t DF=L_t+ _a \γ _t(a)-d_t(a)\. (2.8) Moreover, tDF=0⇔Lt=0andγωt(a)≤dt(a)∀a.O_t DF=0 L_t=0 γ _t(a)≤ d_t(a)\ ∀ a. (2.9) In particular, if γ>0γ>0 and any current coarse-optimal action has ωt(a)>0 _t(a)>0, then the coarse representation is dynamically insufficient even when Lt=0L_t=0. The theorem is a decomposition, not a replacement for POMDP or dual-control theory. It isolates a failure mode useful for learning systems: a summary may support the correct action now yet collapse distinctions that become decision-relevant after the chosen action enters a new training state. 3 Decision-sufficient revelation and revelation depth A learning execution state is broader than model parameters: Xt=(θt,mt,vt,history,RNG,scheduler,…).X_t=( _t,m_t,v_t,history,RNG,scheduler,…). (3.1) The legal current information is ℋt=σ(Ct), H_t=σ(C_t), (3.2) where CtC_t may contain all prospectively legal current losses, readouts, history summaries, optimizer summaries, and telemetry. “Current-only” therefore does not mean behavior-only. A controlled future-learning probe YhY_h induces h=ℋt∨σ(Yh). G_h= H_t σ(Y_h). (3.3) The companion work establishes predictive non-sufficiency but not decision value [40]. The present section asks which parts of a current-behavior fiber must be revealed to factor the terminal decision quotient. For the binary action menu Keep,Xi\ Keep, Xi\, let qa(x)q_a(x) denote the state-conditional expected terminal utility of action a under the declared evaluation technology, and define the terminal action gap g(x)=qXi(x)−qKeep(x).g(x)=q_ Xi(x)-q_ Keep(x). (3.4) A pointwise decision-aliasing witness consists of x1,x2x_1,x_2 such that C(x1)=C(x2),g(x1)g(x2)<0.C(x_1)=C(x_2), g(x_1)g(x_2)<0. (3.5) The full state need not be recovered: the probe only needs to refine the decision quotient, i.e., the distinctions necessary to select an optimal action. Definition 3.1 (Decision-relevant revelation). A future-learning probe YhY_h is decision-relevant relative to ℋt H_t if VB(h)>VB(ℋt).V_B( G_h)>V_B( H_t). (3.6) 3.1 Exact binary aliasing identity Let D=QXi−QKeepD=Q_ Xi-Q_ Keep, mH=[D∣ℋ]m_H=E[D H], and mG=[D∣]m_G=E[D G]. Define the ℋ H-measurable random variables aH=[(mG)+∣ℋ],bH=[(−mG)+∣ℋ].a_H=E[(m_G)_+ H], b_H=E[(-m_G)_+ H]. (3.7) By the tower property, mH=aH−bHm_H=a_H-b_H almost surely. Theorem 3.2 (Binary aliasing identity). For the binary action menu, the exact refinement value is VB()−VB(ℋ)=[minaH,bH]. V_B( G)-V_B( H)=E[ \a_H,b_H\]. (3.8) Thus strict value is present exactly when the coarse information has positive probability of retaining refined posterior mass on both sides of the terminal decision boundary. More quantitatively, let A∈ℋA∈ H with ℙ(A)>0P(A)>0. If for some δ>0δ>0 and α,β>0α,β>0, ℙ(mG≥δ∣ℋ)≥α,ℙ(mG≤−δ∣ℋ)≥βa.s. on A,P(m_G≥δ H)≥α, (m_G≤-δ H)≥β .s.\ on A, (3.9) then VB()−VB(ℋ)≥ℙ(A)δmin(α,β)>0. V_B( G)-V_B( H) (A)\,δ\, (α,β)>0. (3.10) Equation (3.8) is the distributional form of the alias-cell argument: within each coarse-information fiber, refined states can favor opposite terminal actions, and only the smaller of the two conditional magnitude-weighted sign masses creates irreducible coarse decision regret. Predictive variation that never crosses the declared decision boundary has zero value for this binary action menu. No positive-probability atom of the coarse information is required. 3.2 Local geometry and revelation depth Let Φc _c denote the complete legal frontier information map after a short probe depth c (current information plus the frontier summary), let g be the terminal action-gap function defined above, and let YhY_h be a deeper revelation observation. Proposition 3.3 (Local aliasing and revelation). Suppose Φc _c, g, and YhY_h are continuously differentiable near x0x_0, g(x0)=0g(x_0)=0, and Φc _c has locally constant rank. If there exists v∈kerDΦc(x0)v∈ D _c(x_0) (3.11) with Dg(x0)v≠0,DYh(x0)v≠0,Dg(x_0)v≠ 0, DY_h(x_0)v≠ 0, (3.12) then a local curve contained in the frontier level set passes through states requiring opposite terminal actions, while the deeper observation changes to first order. The frontier information is therefore locally decision-aliased along v, whereas the depth-h probe is locally sensitive to that direction. Definition 3.4 (Revelation depth). For a decision-relevant direction v at x0x_0, define its first revelation depth τ⋆(v)=infh:DYh(x0)v≠0.τ (v)= \h:DY_h(x_0)v≠ 0\. (3.13) For a discrete probe grid, the infimum is understood over the admissible depths. The quantity records first sensitivity only; it does not by itself assume that sensitivity must persist at every deeper exact-depth readout. Corollary 3.5 (Equal-budget short-probe separation). Suppose the conditions of Proposition 3.3 hold and a restart frontier reaches depth c while a productive active method reaches h=m−1ch= mm-1c at equal update budget. If DYc(x0)v=0,DYh(x0)v≠0,DY_c(x_0)v=0, DY_h(x_0)v≠ 0, (3.14) then the short-probe information is locally blind to a decision-changing direction that the deeper active observation reveals. On a discrete admissible depth grid—or more generally when the first sensitive depth is attained—persistence over the admissible depth family (for example because the legal depth-h record retains earlier readouts) makes this condition equivalently summarized by c<τ⋆(v)≤hc<τ (v)≤ h. If, in addition, this aliasing occurs on a positive-probability coarse-information region and the refined conditional action gap places nonzero mass on both signs there, Theorem 3.2 gives strict Bayes refinement value. For the prespecified two-action contract, the persistent/cumulative-readout shorthand for the regime of interest is 4<τ⋆(v)≤8, 4<τ (v)≤ 8, (3.15) while the exact nonpersistent condition is DY4(x0)v=0DY_4(x_0)v=0 and DY8(x0)v≠0DY_8(x_0)v≠ 0. This is a conditional theory-level separation from an H4 DynaMiCS-style short-probe information class. It does not assert that every empirical H4 representation satisfies the antecedent; the Transformer study instead tests whether the fixed H8 observation produces reproducible decision-relevant refinement in the declared frontier-ambiguity regime. A Coarse AliasC(x1)=C(x2)C(x_1)=C(x_2)x1x_1x2x_2B Same Future-Learning ProbeResponse Yh(x1)Y_h(x_1)Response Yh(x2)Y_h(x_2)YhY_hYhY_hC Decision SplitKeep OptimalXi Optimal Revelation has decision value only when a currently aliased distinction is separated across the action boundary. Figure 1: Decision-relative revelation. A coarse information cell can contain learning states that require different terminal actions. Deeper future-learning responses are useful only when they separate that decision quotient; full hidden-state reconstruction is unnecessary. 4 Productive revelation technology Having identified what a useful refinement must reveal, we next ask how that information is acquired and what state the experiment leaves behind. A trial protocol u of depth h produces an observation Yu,hY_u,h and the refined information u,h=ℋ∨σ(Yu,h). G_u,h= H σ(Y_u,h). (4.1) We distinguish two execution technologies. Restart / temporary probing. The tested path is discarded. After a decision is made, the selected action starts freshly from the original anchor. Promotion / productive revelation. The selected tested path is retained and continued from its trial endpoint. The probe is therefore simultaneously an information acquisition and a partially executed candidate action. The distinction matters because pure information refinement and technology expansion are different objects. Promotion is not “free compute”: it is credit for work that remains valid on the selected deployment path. Conversely, when a probe is unsafe, destructive, or otherwise unusable for deployment, restart is the correct technology and the identity below does not apply. 4.1 Exact equal-budget identity Suppose there are m≥2m≥ 2 candidate actions and a common terminal horizon H, with admissible probe depths c,h∈[0,H]c,h∈[0,H]. A temporary/restart method that probes every action to depth c and then executes the selected action freshly to H spends KF(c)=mc+H.K_F(c)=mc+H. (4.2) A probe-and-promote method that probes every action to depth h and then continues the selected tested path spends KAR(h)=mh+(H−h)=H+(m−1)h.K_ AR(h)=mh+(H-h)=H+(m-1)h. (4.3) Proposition 4.1 (Productive revelation equal-budget identity). If KF(c)=KAR(h)K_F(c)=K_ AR(h), then h=m−1c. h= mm-1c. (4.4) For two actions, h=2ch=2c. Because h≤Hh≤ H, an equal-budget productive match within the declared horizon is feasible only when c≤(m−1)H/mc≤(m-1)H/m; this condition is satisfied by the H4/H8/H12 comparison below. For the prespecified Transformer comparison, m=2m=2, H=12H=12, and both methods receive 20 policy-visible updates: KF(4)=2(4)+12=20,KAR(8)=12+8=20.K_F(4)=2(4)+12=20, K_ AR(8)=12+8=20. (4.5) Thus an H4 temporary probe and an H8 productive probe are compute-matched under this contract. Equal-Budget Compute Geometry20 policy-visible updates per policyRestart frontierProductive revelationKeep probe H4Xi probe H4Fresh selected trajectory to H128 probe updates discarded12 fresh execution updatesKeep probe H8Xi probe H8Continue 416 probe updates; selected H8 prefix is reusable=20=20=20=20048121620Updates Figure 2: Equal-budget compute geometry for the two-action H12 comparison. Restart probing discards both H4 trials and executes the selected action freshly; productive revelation probes both actions to H8 and retains the selected tested prefix. Both consume 20 policy-visible updates, but promotion buys twice the revelation depth in the two-action case. 4.2 End-to-end value decomposition For any fixed active policy πA _A and frontier policy πF _F, let VRV_R denote common-restart evaluation and VPV_P productive evaluation of the active policy. Adding and subtracting VR(πA)V_R( _A) gives the exact identity VP(πA)−VR(πF)=VR(πA)−VR(πF)⏟Δrestart+VP(πA)−VR(πA)⏟Δpath. V_P( _A)-V_R( _F)= V_R( _A)-V_R( _F)_ _ restart+ V_P( _A)-V_R( _A)_ _ path. (4.6) The first term measures policy quality on common restart outcomes; the second measures productive reuse of the tested path. This identity is especially useful empirically because the two components can be estimated on independent panels without mixing absolute outcomes from unrelated future banks. 5 Adaptive revelation control: the dynamic closure The preceding sections treat ℋ⊆ H G as a fixed information refinement. Once future-learning depth can be chosen, the information state itself becomes a control variable. This section shows that the dynamic problem is not a separate objective: it is the conditional version of the same Bayes refinement value in Equation (2.2). The generic “continue if expected decision improvement exceeds cost” principle is classical metareasoning [36, 13]; the learning-state specialization here supplies the revelation filtration, productive cost geometry, and revealability state. Let ℱh\ F_h\ be the legal revelation filtration, nested in probe depth h, and define the Bayes commit value Ch=maxa∈[Qa∣ℱh].C_h= _a E[Q_a F_h]. (5.1) For h′>h >h, define the oracle conditional value of purchasing the deeper refinement h→h′=[Ch′∣ℱh]−Ch.G_h→ h =E[C_h F_h]-C_h. (5.2) Proposition 5.1 (Static-to-dynamic embedding). For every nested pair ℱh⊆ℱh′ F_h F_h , [h→h′]=VB(ℱh′)−VB(ℱh)=ℐ(ℱh′∣ℱh). E[G_h→ h ]=V_B( F_h )-V_B( F_h)=I( F_h F_h). (5.3) Thus the static Bayes refinement value is the population average of the state-dependent continuation value used by adaptive revelation. For fixed shallow and deep decision policies πh _h and πh′ _h , define the legal-information policy-level gain Gh→h′=[Qπh′−Qπh|ℱh].G_h→ h =E\! [Q_ _h -Q_ _h\, |\, F_h ]. (5.4) When the policies are Bayes optimal for their legal information, this coincides with Equation (5.2). With m candidate actions and productive promotion, extending every candidate from h to h′h costs only ΔK(h,h′)=(m−1)(h′−h) K(h,h )=(m-1)(h -h) (5.5) additional policy-visible updates. At update price λ, write ch→h′=λ(m−1)(h′−h)c_h→ h =λ(m-1)(h -h). Proposition 5.2 (Conditional revelation stopping). Relative to stopping at depth h, the pointwise optimal one-step continuation decision for a fixed policy pair is continue to h′⇔Gh→h′>ch→h′. continue to h G_h→ h >c_h→ h . (5.6) For Bayes policies, the same rule uses h→h′G_h→ h . A population ambiguity quantile is therefore not, in general, a cost-optimal stopping boundary. Theorem 5.3 (Cost-adjusted scalar-control factorization). Let ShS_h be any declared scalar summary measurable with respect to ℱh F_h, and let Γh→h′:=Gh→h′−ch→h′ _h→ h :=G_h→ h -c_h→ h (5.7) be the conditional net gain from continuing rather than stopping. Relative to always stopping, define the one-step meta-control values Vmeta(ℱh) V_ meta( F_h) =[(Γh→h′)+], =E[( _h→ h )_+], (5.8) Vmeta(Sh) V_ meta(S_h) =[([Γh→h′∣Sh])+]. =E\! [ (E[ _h→ h S_h] )_+ ]. (5.9) If a(Sh)=[(Γh→h′)+∣Sh],b(Sh)=[(−Γh→h′)+∣Sh],a(S_h)=E[( _h→ h )_+ S_h], b(S_h)=E[(- _h→ h )_+ S_h], (5.10) then the exact value of retaining the full legal shallow information rather than only ShS_h is Vmeta(ℱh)−Vmeta(Sh)=[mina(Sh),b(Sh)]. V_ meta( F_h)-V_ meta(S_h)=E[ \a(S_h),b(S_h)\]. (5.11) Consequently, Vmeta(ℱh)=Vmeta(Sh)V_ meta( F_h)=V_ meta(S_h) (5.12) if and only if, conditional on ShS_h, the legal shallow states do not place positive probability on both Γh→h′>0 _h→ h >0 and Γh→h′<0 _h→ h <0 on any set of positive probability. Equivalently, there exists an optimal Stop/Continue meta-action that factorizes through ShS_h; ties at zero are value-neutral. Theorem 5.3 sharpens the distinction between predictive and decision nonredundancy. A second shallow coordinate may change the numerical continuation value while remaining irrelevant to the optimal meta-action if every state within a scalar fiber stays on the same side of the priced stopping boundary. What makes an additional revealability coordinate decision-nonredundant is not conditional variation by itself, but cost-adjusted boundary crossing within a scalar fiber. Proposition 5.4 (Finite-depth Bellman closure). Let h0<⋯<hJh_0<·s<h_J be admissible revelation depths and let cjc_j be the priced marginal cost of moving from hjh_j to hj+1h_j+1. If deeper revelation can be purchased sequentially, then VJdyn V_J dyn =ChJ, =C_h_J, (5.13) Vjdyn V_j dyn =maxChj,[Vj+1dyn∣ℱhj]−cj. = \C_h_j,\;E[V_j+1 dyn F_h_j]-c_j \. (5.14) Thus static information refinement becomes an optimal-stopping control problem over the revelation filtration. In a general fork–probe–promote system the state must also retain promoted-state information needed to make the transition Markov. 5.1 Margin and revealability: value variation versus decision nonredundancy For binary actions, let D=QXi−QKeep,mh=[D∣ℱh].D=Q_ Xi-Q_ Keep, m_h=E[D F_h]. (5.15) Suppose the refined Bayes margin admits the conditional location–scale representation mh′=mh+σhZ,m_h =m_h+ _hZ, (5.16) where Z is independent of ℱh F_h, symmetric about zero, integrable, and nondegenerate, while σh≥0 _h≥ 0 is ℱh F_h-measurable. Define ψ(t)=[(Z−t)+],t≥0,ψ(t)=E[(Z-t)_+], t≥ 0, (5.17) and define the conditional Bayes value of deeper revelation by ℛh→h′:=[(mh′)+∣ℱh]−(mh)+.R_h→ h :=E[(m_h )_+ F_h]-(m_h)_+. (5.18) Theorem 5.5 (Two-coordinate revelation geometry). Under Equation (5.16), the conditional Bayes value of deeper revelation is ℛh→h′=σhψ(|mh|σh), R_h→ h = _hψ\! ( |m_h| _h ), (5.19) with the convention ℛh→h′=0R_h→ h =0 when σh=0 _h=0. Wherever differentiable, ∂ℛ∂|mh|=−Pr(Z>|mh|σh)≤0, ∂|m_h|=- \! (Z> |m_h| _h )≤ 0, (5.20) and ∂ℛ∂σh=ψ(t)+tPr(Z>t)≥0,t=|mh|σh. ∂ _h=ψ(t)+t (Z>t)≥ 0, t= |m_h| _h. (5.21) Hence, in the nonredundant regime, the cost-optimal continuation boundary is a curve in (|mh|,σh), (|m_h|, _h), (5.22) rather than a universal threshold on current decision margin. The derivative in σh _h is strict whenever the refinement retains positive probability mass beyond the current normalized margin. The theorem itself does not require σh _h to remain empirically nonredundant after conditioning on |mh||m_h|. Corollary 5.6 (Conditional value nonredundancy and scalar reduction). Let Mh=|mh|M_h=|m_h| and write r(m,s)=sψ(m/s),r(m,s)=s\,ψ(m/s), (5.23) with the same zero-scale convention as in Theorem 5.5. (i) Scalar-redundant regime. If σh=g(Mh) _h=g(M_h) almost surely on a decision-relevant event E for some measurable g, then ℛh→h′=r~(Mh)on E,R_h→ h = r(M_h) E, (5.24) for the scalar function r~(m)=r(m,g(m)) r(m)=r(m,g(m)). Hence current margin is sufficient for the one-step oracle continuation value on E. A single monotone margin threshold requires the additional condition that r~ r be monotone relative to the fixed compute price. (i) Nonredundant regime. If the conditional law of σh _h given MhM_h is nondegenerate on a set of positive probability and, on the corresponding conditional support, s↦r(Mh,s)s r(M_h,s) is strictly increasing, then the conditional law of ℛh→h′R_h→ h given MhM_h is also nondegenerate on a set of positive probability. Consequently no measurable function of current margin alone can reproduce the oracle continuation value there. Corollary 5.7 (Cost-adjusted revealability crossing criterion). Assume the setting of Theorem 5.5, fix a continuation price c>0c>0, and suppose that for almost every m in the decision-relevant region the map s↦r(m,s)s r(m,s) is continuous and strictly increasing on the conditional support of σh|Mh=m _h M_h=m. Define the generalized critical revealability σc(m)=infs≥0:r(m,s)>c, _c(m)= \s≥ 0:r(m,s)>c\, (5.25) with inf∅=+∞ =+∞. (i) The scalar-margin decision is sufficient if and only if, for almost every MhM_h, either Pr(r(Mh,σh)≤c∣Mh)=1orPr(r(Mh,σh)≥c∣Mh)=1. \! (r(M_h, _h)≤ c M_h )=1 \! (r(M_h, _h)≥ c M_h )=1. (5.26) Equivalently, conditional on almost every margin value, the priced continuation gain has no mass on both strict sides of zero. Under the stated continuity and monotonicity conditions, a transparent sufficient support condition is that σh|Mh=m _h M_h=m is contained in [0,σc(m)][0, _c(m)] or in [σc(m),∞)[ _c(m),∞); zero-gain ties at the boundary are value-neutral and may be assigned to either optimal meta-action. (i) If there is a set of margin values with positive probability on which Pr(r(Mh,σh)<c∣Mh)>0andPr(r(Mh,σh)>c∣Mh)>0, \! (r(M_h, _h)<c M_h )>0 \! (r(M_h, _h)>c M_h )>0, (5.27) then the scalar-margin controller is strictly suboptimal: Vmeta(ℱh)>Vmeta(Mh).V_ meta( F_h)>V_ meta(M_h). (5.28) Under the same monotonicity conditions, Equation (5.27) is implied by positive conditional mass strictly below and strictly above σc(Mh) _c(M_h). Thus conditional spread in revealability is decision-nonredundant exactly when a positive-mass scalar fiber contains both the strict Stop side and the strict Continue side of the priced continuation decision; zero-gain ties at the boundary do not affect value. Proposition 5.8 (Location–scale identification of revealability scale). Under Equation (5.16), [|mh′−mh|∣ℱh]=σh|Z|.E\! [|m_h -m_h| F_h ]= _h\,E|Z|. (5.29) If additionally [Z2]=1E[Z^2]=1, then [(mh′−mh)2∣ℱh]=σh2.E\! [(m_h -m_h)^2 F_h ]= _h^2. (5.30) At the population level, these identities identify revealability scale up to a fixed normalization from the conditional spread of the true deeper-margin increment. In applications the Bayes margins themselves may be latent or estimated; the identities then motivate development-only proxy construction rather than claiming direct identification from noisy fitted scores. Definition 5.9 (Legal model-specific revealability proxy). For a model family ℳM, let Xh(ℳ)X_h^(M) denote a declared shallow information map measurable with respect to ℱh F_h. A statistic Rh(ℳ)=rℳ(Xh(ℳ))R_h^(M)=r_M\! (X_h^(M) ) (5.31) is a legal model-specific revealability proxy when its functional form or fitted parameters are determined without using the target outcomes and are fixed before target evaluation. The proxy may estimate σh _h directly, rank the conditional spread in Proposition 5.8, or enter a predeclared estimator of Gh→h′G_h→ h . No equality Rh(ℳ)=σhR_h^(M)= _h is assumed, and different model families need not share the same proxy formula, scale, coefficient, or stopping threshold. Protocol 5.10 (Model-specific Revelation-Control instantiation). For a new learning-system family ℳM, an empirical instantiation should proceed as follows. P1. Declare the decision problem. Fix the action menu, terminal utility, legal shallow filtration, admissible revelation depths, productive-promotion technology, and resource-price ledger before target outcomes are inspected. P2. Separate development from target evaluation. Collect a development panel with the shallow and deeper potential outcomes needed to estimate continuation value. Keep the target panel disjoint at the exact-anchor level. P3. Instantiate the observable continuation state. Fit either Gh→h′G_h→ h directly from legal shallow information or a legal proxy Rh(ℳ)R_h^(M) motivated by Proposition 5.8. Always retain a declared scalar comparator. P4. Control representation selection. Choose a minimal learnable shallow summary using development data only. If a finite set of architectures remains eligible, predeclare that family and use simultaneous or familywise-valid inference on the target panel rather than selecting an unadjusted winner after target evaluation. P5. Fix the controller before evaluation. Fix the fitted estimator, stop/continue threshold, tie and fallback rules, productive compute accounting, risk target, resampling unit, and inferential criteria before evaluating target outcomes. P6. Evaluate once and separate claims. Evaluate the fixed family on the disjoint target panel. Report decision value, compute-accounted utility, and decision-instability risk as distinct objects; a bounded-risk claim becomes an expected-utility certificate only under an explicit severity or tail bridge. The protocol transports the theoretical structure, not the numerical proxy, coefficient, or threshold of an earlier model family. Cross-model transport is structural, not parametric. The theorem-level transport target is the Revelation-Control relation between legal shallow information, continuation value, productive cost, and the stopping decision. A second model family need not reproduce the same empirical proxy used in the first. To claim a decision-nonredundant revealability regime, a legal proxy specified independently of target outcomes must add held-out control value beyond a scalar-margin comparator, or otherwise establish cost-adjusted boundary crossing as in Corollary 5.7. If the model fit on development data and fixed before target evaluation instead admits a successful scalar continuation rule and no legal second coordinate shows incremental control value, that outcome is consistent with scalar decision sufficiency under Theorem 5.3, without requiring the stronger assertion that revealability is functionally determined by margin. In either regime, target evaluation must use the same declared terminal utility and compute ledger, and the proxy must not be selected from target outcomes. |mh||m_h| Decision marginσh _h Revealabilityℛh→h′=ch→h′R_h→ h =c_h→ h CONTINUEℛh→h′>ch→h′R_h→ h >c_h→ h STOPℛh→h′<ch→h′R_h→ h <c_h→ h Higher σh _hLower σh _hSame |mh||m_h| Figure 3: Schematic margin–revealability geometry implied by Theorem 5.5. At fixed compute price, larger current decision margins require greater state-dependent revealability to justify deeper future learning. The paired points illustrate the nonredundant case in which two states with the same current margin can fall on opposite sides of the stopping boundary because their revealability differs. The curve σh=g(|mh|) _h=g(|m_h|) illustrates the stronger value-redundant special case in which continuation value itself reduces to a scalar function. More generally, scalar decision sufficiency requires only that the accessible states at a fixed margin remain on one side of the priced boundary; residual revealability variation is allowed. The boundary is theoretical and schematic, not an empirical fit. The theorem explains why current confidence and value of further revelation need not induce the same ordering in the nonredundant regime. A state can look locally decisive yet remain highly revealable under future training; conversely, a state near the current decision boundary can have little continuation value if its future response is stable. Corollary 5.6 characterizes value-level redundancy, while Theorem 5.3 and Corollary 5.7 give the sharper decision-level statement: scalar utility-aware control can remain optimal even with residual revealability variation, provided no positive-mass scalar fiber crosses the priced Stop/Continue boundary. 6 From oracle revelation to learnable and certifiable control A refinement can improve oracle Bayes value and still be a poor learned representation at finite sample size. To make this distinction explicit, fix a candidate acquisition scheme s with refined information s G_s and promotion technology PsP_s. Let Δsrev _s rev =VB(s)−VB(ℋ), =V_B( G_s)-V_B( H), (6.1) Δsprom _s prom =VPs⋆−VB(s), =V_P_s -V_B( G_s), (6.2) where VPs⋆V_P_s is the best value available under the refined information and the declared promotion technology. When restart remains an admissible fallback under that technology, Δsprom≥0 _s prom≥ 0; otherwise the ledger remains algebraically valid with a signed technology term. For a learner class Πs _s, let Rsapp≥0R_s app≥ 0 be the gap between VPs⋆V_P_s and the best policy in Πs _s, let Rs,nest≥0R_s,n est≥ 0 be the finite-sample gap between that class optimum and the learned policy, and let Cs≥0C_s≥ 0 be acquisition cost in the chosen utility units. Proposition 6.1 (Finite-sample revelation ledger). With the definitions above, the learned net value satisfies the exact identity Ws,nnet−VB(ℋ)=Δsrev+Δsprom−Rsapp−Rs,nest−Cs. W_s,n net-V_B( H)= _s rev+ _s prom-R_s app-R_s,n est-C_s. (6.3) The ledger separates an oracle question from an extraction question. More telemetry can increase Δsrev _s rev while increasing effective estimation complexity enough to worsen Rs,nestR_s,n est. Development analyses with substantially broader telemetry exhibited this estimation obstruction, motivating a compact refinement rather than maximizing raw feature count. 6.1 Minimal learnable refinement Define the expected finite-sample value Ψn(s)=Δsrev+Δsprom−Rsapp−[Rs,nest]−Cs. _n(s)= _s rev+ _s prom-R_s app-E[R_s,n est]-C_s. (6.4) Let κn(s) _n(s) be a learner-relative complexity measure and, for tolerance τ≥0τ≥ 0, let n,τ=s:Ψn(s)≥sups′Ψn(s′)−τ.S_n,τ= \s: _n(s)≥ _s _n(s )-τ \. (6.5) When n,τS_n,τ is nonempty and the minimum of κn _n is attained on it, a minimal learnable refinement is any sn,τ⋆∈argmins∈n,τκn(s).s_n,τ ∈ _s _n,τ _n(s). (6.6) Minimality is therefore relative to the learner, sample size, acquisition technology, and tolerance. It need not mean the fewest raw features or the mathematically coarsest sufficient sigma-field. 6.2 Learned frontier-victory ledger Let a fixed frontier learner π^F π_F use information ℋ H under restart evaluation, and let a fixed active learner π^A π_A use ⊇ℋ G H. Define RF=VB(ℋ)−VR(π^F),RA=VB()−VR(π^A),R_F=V_B( H)-V_R( π_F), R_A=V_B( G)-V_R( π_A), (6.7) where VRV_R is common restart-evaluated value, and let PA=VP(π^A)−VR(π^A)P_A=V_P( π_A)-V_R( π_A) be the productive promotion value of the active policy. Corollary 6.2 (Learned frontier victory). Writing Δrev=VB()−VB(ℋ) _ rev=V_B( G)-V_B( H), VP(π^A)−VR(π^F)=Δrev+PA+RF−RA. V_P( π_A)-V_R( π_F)= _ rev+P_A+R_F-R_A. (6.8) If resources KA,KFK_A,K_F are priced by λ, then net value satisfies NA−NF=Δrev+PA+RF−RA+λ⊤(KF−KA).N_A-N_F= _ rev+P_A+R_F-R_A+λ (K_F-K_A). (6.9) At equal resource cost, a sufficient condition for active victory is Δrev+PA>RA−RF _ rev+P_A>R_A-R_F. Equation (6.8) separates the two empirical questions studied later: whether deeper future learning reveals decision-relevant structure, and whether productive reuse converts that structure into superior end-to-end utility. 6.3 Certification boundary for adaptive depth Finite-sample risk calibration for early stopping is established prior art: selective prediction, Learn-then-Test, Conformal Risk Control, early-time classification, and risk-controlled early-exit networks provide general mechanisms for calibrating bounded stopping risks [12, 1, 2, 35, 19, 41]. Productive-prefix reuse exposes a particularly natural structural loss here. If Fh→h′=πh≠πh′,F_h→ h =1\ _h≠ _h \, (6.10) then, under the declared same-path promotion technology, Fh→h′=0⟹Qπh=Qπh′.F_h→ h =0 Q_ _h=Q_ _h . (6.11) Thus shallow/deep decision instability can be calibrated with a bounded Bernoulli loss even when terminal utility gaps are heavy-tailed. Turning that probability statement into a strict expected-net-utility guarantee additionally requires control of the stopped-harm tail. Appendix B makes this boundary sharp: for any nonzero stop–flip risk, unrestricted severity makes the worst-case expected net utility unbounded below even if stop and flip probabilities are known exactly. Bounded-severity, moment, or integrable-tail assumptions provide sufficient bridges from bounded risk to expected utility. 7 Related work and novelty boundary Bayesian decision theory and information value. The monotonic value of information under refinement is classical in statistical decision theory and comparison of experiments [5, 4, 15]. Expected information from experiments and Bayesian experimental design provide a complementary tradition in which observations are deliberately acquired under a utility criterion [26, 6]. Our Bayes-value notation is a specialization of these traditions, not a new notion of information value or experimental design. The distinctive question is which learning-state distinctions are worth purchasing through a future training intervention once finite-sample extraction and execution technology are included. Terminology: revelation in decision analysis and mechanism design. The phrase value of revelation has prior use in influence-diagram decision analysis: Ezawa [10] defines it from values of evidence and relates it to value of control. We therefore do not claim that phrase as new; our conditional quantity in Equation (5.2) is a continuation value for an intervention-generated refinement that can also change the controlled system and create reusable progress. The word “revelation” also has a different established meaning in mechanism design, where the revelation principle concerns truthful direct mechanisms under private information [31]. Revelation Control does not assume strategic reporting or incentive compatibility. Its object is controlled acquisition of decision-relevant state through system response. Partial observability, active sensing, and dual control. POMDPs represent decisions under partial observability through belief states [21]; dual control emphasizes that an action can simultaneously control a system and reveal uncertainty [11]; and active-sensing work studies task-directed information acquisition [39]. Revelation Control belongs to this broad family but does not replace it. Our narrower object is an internal learning execution state, the intervention is future training, and the formal target is decision factorization through a declared legal information sigma-field. The promotion term further distinguishes information gained by a trial from useful state change left by that trial. Predictive state and decision-focused representations. Predictive State Representations encode state through action-conditional predictions of future observations [27]. The companion work is related in spirit because future responses reveal hidden learning-state distinctions, but the present paper does not propose a general predictive state representation. It asks only for the quotient needed to choose among a fixed action menu. This has conceptual kinship with state abstraction, where representations are judged by whether they retain distinctions needed for planning or learning [24]. Recent work on decision-relevant concept selection makes this connection explicit by requiring states that share a selected concept representation to preserve the optimal decision structure [34]. Decision-focused learning and predict-then-optimize methods likewise emphasize that predictive quality should be judged by downstream decision loss [8, 9, 29]; our finite-sample ledger adds an intervention/acquisition dimension in which the representation itself must be actively produced. Function-equivalent states and self-intervention. Recent work on path-conditioned training makes especially clear that ReLU parameterizations implementing the same function can nevertheless induce different subsequent training dynamics [23]. This is closely aligned with the motivating non-sufficiency phenomenon, but it studies rescaling and conditioning rather than decision-relative future-learning experiments. Self-Interventional Learning perturbs a network’s own functional organization, learns a predictive self-model from intervention consequences, and uses that model for later structural action [38]. Revelation Control instead treats the learner as the controlled object of an external decision problem, asks which future-training responses refine a declared action quotient, and accounts separately for information and reusable tested computation. Partial training, early stopping, and resource allocation. Freeze–Thaw Bayesian optimization uses partial learning curves to pause and resume candidate models [37]; learning-curve extrapolation uses early trajectory segments to predict eventual performance and stop unpromising runs [7]; and non-stochastic best-arm allocation formalizes the use of intermediate learning performance to concentrate resources on promising configurations [18]. Hyperband extends this resource-allocation perspective at scale [25], while Population Based Training jointly adapts model populations and hyperparameters while reusing live trajectories [17]. These works establish strong precedents for partial training, continuation, early termination, promotion, and resource allocation. Our novelty therefore does not rest on any of those operations individually. DynaMiCS and local prospective probes. Within the declared fine-tuning-probe problem class, DynaMiCS is the closest direct comparator because it explicitly performs short domain-specific fine-tuning probes to estimate local cross-domain slopes before optimizing a constrained data mixture [14]. We therefore concede short prospective training probes and local finite-difference summaries as prior art. Our theory asks a different question: can a local-probe information state fail to factorize the terminal decision, and can productive reuse make a deeper decision-relevant observation available under the same update budget? The prespecified H4 comparator is a strengthened task adaptation of that acquisition geometry, not a claim that the original DynaMiCS algorithm is globally dominated. Metareasoning and value of computation. Rational metareasoning asks whether an additional computation is worth its cost because of the external decision it may change [36]; knowledge-gradient policies similarly choose measurements by expected increment in terminal value [13]. Recent adaptive-reasoning work operationalizes marginal-gain-versus-cost allocation through difficulty signals [43], while consequence-aware allocation shows that difficulty and error severity need not coincide when assigning test-time reasoning budgets [42]. Proposition 5.2 is a Revelation-Control specialization of value-of-computation reasoning, not a new generic stopping rule. The structural point specific to this paper is that Equation (5.3) makes the original hidden-learning-state refinement value the state-dependent reward of that control problem. The distinctive object is hidden learner execution state: the computation is controlled future training, and the selected experimental path can remain useful training. Risk-controlled stopping. Reject-option and selective-classification methods trade prediction coverage against conditional risk [12]. Learn-then-Test and Conformal Risk Control provide finite-sample calibration of bounded risks for families of predictive rules [1, 2]. These tools have also been applied directly to sequential stopping: early-time classification can be calibrated for accuracy-gap control [35], risk-controlled neural early exits can use supervised or consistency losses [19], and recent work controls reasoning risk under a compute budget [41]. We therefore do not claim to introduce risk-controlled early stopping or early-versus-deep consistency. Our narrower contribution is to connect these ideas to productive future-training probes of hidden learning state, where deeper computation changes the learner and reveals a decision quotient. Value of computation and finite-sample extraction. Our acquisition-cost and learner-regret terms also interact with finite-sample extraction: richer future telemetry can have higher oracle information value and lower learned value because estimation error grows. The minimal-learnable-refinement principle therefore treats representation complexity as part of the decision problem rather than assuming that more legal telemetry is automatically better. Table 1: Novelty boundary relative to representative neighboring methods. Entries describe the central mechanism emphasized by each work, not every implementation variant. Method / theory Partial training Prospective training signal Path reuse Hidden learning-state decision aliasing Explicit finite-sample value ledger Freeze–Thaw [37] yes learning curves pause/resume no no Hyperband [25] yes intermediate performance continuation of survivors no no PBT [17] yes population performance yes no no DynaMiCS [14] yes local cross-domain slopes no no no Risk-controlled early exit [35, 19] no shallow/deep prediction signal no no risk control, not utility ledger Revelation Control (this work) yes future-learning response promotion separated from information value central target yes Novelty claim. The paper therefore does not claim to invent future probing, promotion, Bayes value, value of computation, selective early exit, finite-sample risk calibration, or dynamic control. The strongest claimed contribution is the coupling hidden learning-state aliasing+productive future-training experiment+budget-indexed revelation depth+finite-sample extraction+comparative utility. gatheredhidden learning-state aliasing+productive future-training experiment\\ +budget-indexed revelation depth+finite-sample extraction+comparative utility. gathered (7.1) The adaptive-control layer completes this chain by identifying current decision margin and learning-state revealability as joint control coordinates for purchasing deeper revelation, while also characterizing when revealability is conditionally redundant and scalar control suffices; the generic metareasoning and risk-calibration tools used around that result are explicitly treated as prior art. 8 Transformer instantiation We evaluate Revelation Control in two independently instantiated decoder-only 7B Transformer families: Qwen2.5-7B [32, 33] and Mistral-7B-v0.3 [20, 30]. Both use LoRA adaptation [16] with AdamW optimizer state [22, 28], while the Mistral instantiation uses independently prepared task streams, training histories, and development data. The cross-model target is therefore structural rather than numerical. In both families, exact anchors are generated from controlled training histories and the binary action menu is =Keep,Xi,A=\ Keep, Xi\, (8.1) where Xi denotes the prespecified matched optimizer-state intervention. The intervention matches the immediate adaptive field while altering hidden optimizer moments, allowing later common training to reveal a difference that is absent from the immediate update. Terminal utility is evaluated at H12. The active policy probes both candidate actions to H8, selects an action using a decision rule fit on development data and fixed before target evaluation, and promotes the selected tested path to H12. The strengthened short-probe frontier probes to H4, restores the anchor, selects an action from legal H4 information, and restarts that action to H12. 8.1 Qwen2.5-7B prespecified active policy For Qwen2.5-7B, the final fixed-depth active policy is 24-dimensional raw H8 response+Ridge(α=10)+zero threshold+same-path promotion. gathered24-dimensional raw H8 response+Ridge(α=10)\\[-1.0pt] +zero threshold+same-path promotion. gathered (8.2) Ties and nonfinite scores fall back to Keep. The representation and decision rule were selected on development data and fixed before the two independent finite-library panels and the 432-anchor comparative panel. Selection status. The method is prespecified, not a proof of global algorithmic optimality. Development-only comparisons evaluated H2, H4, H6, and H8; H8 was selected before target evaluation, while substantially broader telemetry increased effective estimation burden without a reliable compensating gain. 8.2 Independent Mistral-7B-v0.3 instantiation For Mistral-7B-v0.3, the same H4/H8/H12 decision geometry, action menu, terminal utility, productive-promotion technology, and 20-update accounting are retained, while numerical heads and adaptive summaries are instantiated from Mistral-specific development data. The theory does not require Qwen’s empirical revealability proxy, coefficients, or stopping threshold to transfer numerically. A separate 336-anchor development panel supplies the Mistral parameters, and an independent 480-anchor two-bank panel supplies target evaluation. The adaptive analysis is restricted to the scalar and two-coordinate continuation architectures already defined in the Qwen analysis, with Mistral parameters fit only on the development panel. The transported object is the Revelation-Control structure, not one fitted coordinate system. Statistical units, lower bounds, multiplicity control, and bank aggregation are specified in Appendix E. 9 Fixed-depth revelation beyond current-information comparators Before turning to the independent model-family replication, we first establish that the selected H8 future-response representation carries decision value beyond a broad declared current-information library in Qwen2.5-7B. The same prespecified 24-dimensional H8 raw-response Ridge policy (α=10α=10) is evaluated on two independent panels, each containing 336 exact anchors with exact position balance and an analysis fixed before target evaluation. The comparator library contains 22 legal current-information components: the two fixed actions plus five prespecified decision heads within each of four current-information feature families (position, order, anchor, and H0), as displayed in Figure 4. The two panels are analyzed independently rather than pooled. Table 2: Independent Qwen2.5-7B validation against the 22-component current-information library. Each entry is the worst component across the declared comparator set; positive lower bounds are required simultaneously. Quantity Independent panel A (336) Independent panel B (336) Minimum point margin ×10−41.2136812\!×\!10^-4 ×10−41.4878650\!×\!10^-4 Minimum Student-t LCB ×10−56.9363323\!×\!10^-5 ×10−59.4261533\!×\!10^-5 Minimum bootstrap LCB ×10−57.0252125\!×\!10^-5 ×10−59.5137645\!×\!10^-5 Minimum max-t LCB ×10−54.4947532\!×\!10^-5 ×10−56.9209070\!×\!10^-5 Comparators satisfying criterion 22/22 22/22 Every componentwise comparison remains positive in both panels for both point margins and simultaneous max-t lower bounds; Figure 4 shows all 22 comparisons on a common scale. Open square: simultaneous LCB ⟶ filled circle: point estimateIndependent panel AIndependent panel B012345012345Fixed actionsPositionOrderAnchorH0KEEPXIRidge .1Ridge 10Ridge 1000Logit .1Logit 10Ridge .1Ridge 10Ridge 1000Logit .1Logit 10Ridge .1Ridge 10Ridge 1000Logit .1Logit 10Ridge .1Ridge 10Ridge 1000Logit .1Logit 10H8-policy minus current-policy utility (×10−4× 10^-4) Figure 4: Two independent Qwen2.5-7B panels across all 22 current-information comparators. Open squares mark simultaneous max-t 95% lower confidence bounds and filled circles mark realized point margins. Every lower-bound endpoint remains strictly positive; the two panels are not pooled. Productive-path value is also independently positive. Panel A gives Δ^path(A)=8.99131091346×10−5, _ path^(A)=8.99131091346× 10^-5, (9.1) with one-sided Student-t and bootstrap lower bounds 4.22948353683×10−54.22948353683× 10^-5 and 4.41065688619×10−54.41065688619× 10^-5. Panel B independently gives Δ^path(B)=1.47763681856726×10−4, _ path^(B)=1.47763681856726× 10^-4, (9.2) with corresponding lower bounds 1.01800375097654×10−41.01800375097654× 10^-4 and 1.03042544792948×10−41.03042544792948× 10^-4. These path values are later combined one panel at a time with the separate 432-anchor common-restart comparison through Equation (4.6). The empirical claim in this section remains limited to the declared comparator library and the stated same-path execution technology. 10 Frontier comparison: DynaMiCS-style probing DynaMiCS performs short domain-specific fine-tuning probes to estimate a slope matrix of local cross-domain effects and then optimizes mixture weights under performance constraints [14]. That makes it direct contemporary prior art for using future training itself as a prospective diagnostic. Our comparator is therefore not presented as the original DynaMiCS algorithm verbatim; it is a strengthened, task-adapted DynaMiCS-style frontier that preserves the relevant short-probe/local-slope/restart geometry while using a calibrated binary decision head tailored to the present terminal action problem. Why strengthen the comparator? The source algorithm solves a constrained mixture-selection problem, whereas our terminal decision is binary. A literal transplant therefore requires an additional mapping from local slope information to the present action gap. If that mapping were deliberately weak, a positive frontier result could reflect head misspecification rather than a limitation of short-probe information. We instead retain the source method’s short-probe acquisition structure—same anchor, temporary short probes, local finite-difference summaries, restore/restart—but give the frontier a decision head fit only on development data for the declared H12 action gap. This strengthens the competing decision rule without granting deeper or otherwise illegal information. For Qwen2.5-7B, the legal H4 feature vector is zF=(LA0,LB0,LC0,SKeep,A4,SKeep,B4,SKeep,C4,SXi,A4,SXi,B4,SXi,C4),z_F=(L_A^0,L_B^0,L_C^0,S_ Keep,A^4,S_ Keep,B^4,S_ Keep,C^4,S_ Xi,A^4,S_ Xi,B^4,S_ Xi,C^4), (10.1) where Sa,d4=Ld(θaH4)−Ld(θ0)4.S_a,d^4= L_d( _a^H4)-L_d( _0)4. (10.2) The Qwen decision head is fixed as StandardScaler + Ridge(α=0.1α=0.1) with zero threshold and no test refit. Mistral-7B-v0.3 uses the same legal short-probe information class and restart technology with a model-specific numerical head fit on development data and fixed before target evaluation. 10.1 Equalized compute contract Both methods receive exactly 20 policy-visible training updates: Resource Productive revelation DynaMiCS-style restart Probe updates 16 8 Fresh/promoted continuation 4 12 Total policy-visible updates 20 20 Discarded exploratory updates 8 8 Reused selected-probe updates 8 0 Terminal horizon H12 H12 Terminal utility identical identical Restore/save overhead is not charged against the comparator in the primary analysis, making the update accounting conservative for the productive-revelation method. Theory-level comparison. Corollary 3.5 gives a conditional separation from an H4 short-probe information class: if a decision-changing direction is invisible through H4 yet revealed by H8, then the deeper refinement can have strictly larger Bayes value at the same update budget because productive work is reusable. The empirical sections below test the corresponding structural refinement and end-to-end utility consequences in both model families. 11 Structural decision refinement in the H4 frontier-ambiguity regime The 432-anchor comparative panel is disjoint from development and from the two finite-library replication panels. Before terminal outcomes, the H4 frontier score defined the fixed ambiguity event A=|fF|≤τF,τF=0.00052099142651661841.A=\|f_F|≤ _F\, _F=0.00052099142651661841. (11.1) The event contains 226/432 anchors, or 52.315% of the panel. Let Dr=UTr(Xi)−UTr(Keep),r∈1,2,D_r=U_T_r( Xi)-U_T_r( Keep), r∈\1,2\, (11.2) and let IARI_ AR be the fixed H8 active decision. The two terminal banks are independent conditional replicates and expose both potential actions for every anchor. 11.1 Independent-bank refinement estimator To test whether the H8 observation refines the coarse ambiguity cell in a decision-relevant way without reusing terminal outcomes for both choice and evaluation, bank T1T_1 chooses the best coarse constant action inside A, c^1=D¯1,A>0, c_1=1\ D_1,A>0\, (11.3) and the independent bank T2T_2 evaluates the fixed H8 decisions relative to that coarse action: Δ^1→2(A)=1|A|∑i∈AD2i(IAR,i−c^1). _1→ 2(A)= 1|A| _i∈ AD_2i(I_ AR,i- c_1). (11.4) The reverse direction Δ^2→1(A) _2→ 1(A) is defined symmetrically, and the primary symmetric estimator is their average. In each exact-anchor bootstrap draw, the coarse action is reselected inside the resample before evaluation. The two directions agree closely: Δ^1→2(A) _1→ 2(A) =1.92863119513×10−4, =1.92863119513× 10^-4, LCBt _t =5.01319646865×10−5, =5.01319646865× 10^-5, (11.5) Δ^2→1(A) _2→ 1(A) =1.93079378050×10−4, =1.93079378050× 10^-4, LCBt _t =1.03136801205×10−4. =1.03136801205× 10^-4. (11.6) Their symmetric average is Δ^refine(A)=1.92971248782×10−4, _ refine(A)=1.92971248782× 10^-4, (11.7) with one-sided 95% Student-t lower bound 1.09966903938×10−41.09966903938× 10^-4 and 50,000-draw exact-anchor bootstrap lower bound 6.49311403490×10−56.49311403490× 10^-5. Weighting by the fixed ambiguity mass gives a full-panel contribution 1.00952551446×10−41.00952551446× 10^-4, with t lower bound 5.69892792476×10−55.69892792476× 10^-5 and bootstrap lower bound 3.39715294723×10−53.39715294723× 10^-5. 11.2 Opposite-action sign replication The mechanism is visible independently in both terminal banks. Table 3 reports the terminal action gap conditional on the fixed H8 action within A. Table 3: Independent-bank structural sign replication inside the fixed H4 frontier-ambiguity event. Bounds are one-sided 95% Student-t bounds. Bank Xi-selected mean LCB Keep-selected mean UCB T1T_1 2.9246×10−42.9246× 10^-4 6.5097×10−56.5097× 10^-5 −4.1166×10−4-4.1166× 10^-4 −2.2471×10−4-2.2471× 10^-4 T2T_2 3.6323×10−43.6323× 10^-4 9.5547×10−59.5547× 10^-5 −2.8389×10−4-2.8389× 10^-4 −8.2269×10−5-8.2269× 10^-5 Thus the H8 observation does not merely predict a continuous outcome more accurately: within a materially populated H4 frontier-ambiguity regime it reproducibly separates groups of states whose mean independent terminal outcomes favor opposite actions. Treating the ambiguity event as the coarse cell and the fixed H8 partition as the refinement gives a sample Bayes gap of 1.63115121441×10−41.63115121441× 10^-4 using the mean of T1,T2T_1,T_2, with bootstrap lower bound 6.04584733171×10−56.04584733171× 10^-5. The same cell-refinement Bayes gap has a positive bootstrap lower bound in each terminal bank separately. Secondary H4-versus-H8 diagnostic. A cross-bank diagnostic that preserves outcome separation compares the fixed nine-feature H4 Ridge representation with the fixed 24-feature H8 representation for terminal-gap prediction. Inside A, the H8 representation improves cross-bank mean squared error by 1.24677×10−71.24677× 10^-7, with positive t and bootstrap lower bounds. The corresponding learned-policy value difference has a confidence interval crossing zero, reflecting that the strengthened H4 frontier action is already a strong selector. We therefore interpret the main result as decision-relevant H8 refinement inside the declared H4 frontier-ambiguity regime, not as unrestricted dominance over every measurable function of complete H4 information. 11.3 Independent Mistral aliasing and deeper-revelation value The independent Mistral target panel contains 480 exact anchors with two terminal banks. Its H4 ambiguity threshold, τFM=0.00236920914414 _F^M=0.00236920914414, was determined from the separate development panel. Using Bank A’s H4 score to define the ambiguity cell yields 216/480 anchors; among them, 88 favor Xi in both terminal banks and 58 favor Keep in both. Defining the cell with Bank B gives 210/480 anchors, with 83 stable-Xi and 58 stable-Keep anchors. Thus the same coarse shallow-information regime contains materially populated hidden-state subsets requiring opposite terminal actions. More directly, the fixed H8 action has strictly positive terminal decision value relative to the fixed H4 action on the complete two-bank target panel. The exact-anchor mean is Δ^H8−H4M=2.25089539402×10−4, _H8-H4^M=2.25089539402× 10^-4, (11.8) with one-sided 95% Student-t lower bound 1.56391515827×10−41.56391515827× 10^-4 and 50,000-draw bootstrap lower bound 1.58226436880×10−41.58226436880× 10^-4. H4 and H8 actions differ on 100/480 exact anchors in at least one bank. Mistral therefore independently reproduces the core implication of decision-revelation depth: legal deeper future-learning information changes action choice on nontrivial mass and has strictly positive terminal decision value. A stronger Qwen diagnostic in which the H8-selected partition itself must produce opposite-sign subgroup bounds in both directions is not promoted as a cross-model invariant. The theorem requires decision-relevant refinement, not numerical identity of every diagnostic partition. 12 Productive revelation and equal-compute frontier performance We next ask whether additional revelation can be converted into higher end-to-end utility than the strengthened DynaMiCS-style H4 restart frontier under the common 20-update contract. Equation (4.6) gives the exact decomposition Δend=ΔrestartAR−Frontier+ΔpathAR. _ end= _ restart AR- Frontier+ _ path AR. (12.1) This separates what is gained by choosing better from what is gained because the selected deeper probe remains useful computation. 12.1 Qwen equal-compute closure The 432-anchor Qwen panel estimates the common-restart selector component using the mean of two independent terminal banks: Δ^restartAR−Frontier=−3.98358746688×10−7, _ restart AR- Frontier=-3.98358746688× 10^-7, (12.2) with standard deviation 4.67246×10−44.67246× 10^-4. The point estimate is therefore near zero; this is descriptive rather than an equivalence claim. Strict end-to-end advantage is supplied by productive reuse of the tested H8 path. Combining the independent common-restart estimate separately with the two productive-path panels gives Δ^end(1) _ end^(1) =8.95147503879×10−5, =8.95147503879× 10^-5, LCB0.95 _0.95 =2.92463884692×10−5>0, =2.92463884692× 10^-5>0, (12.3) Δ^end(2) _ end^(2) =1.47365323110×10−4, =1.47365323110× 10^-4, LCB0.95 _0.95 =8.83939050151×10−5>0. =8.83939050151× 10^-5>0. (12.4) The productive-path values themselves are 8.9913×10−58.9913× 10^-5 and 1.4776×10−41.4776× 10^-4, respectively, with positive one-sided t and bootstrap lower bounds in both panels. 12.2 Independent Mistral equal-compute replication The same decomposition is evaluated on the independent Mistral target panel under the identical 20-update visible-compute contract. The common-restart selector component is Δ^restartM=−1.191350097×10−5, _ restart^M=-1.191350097× 10^-5, (12.5) whereas same-path reuse contributes Δ^pathM=3.09392744795×10−4, _ path^M=3.09392744795× 10^-4, (12.6) with one-sided Student-t and bootstrap lower bounds 1.96017743591×10−41.96017743591× 10^-4 and 2.01665987607×10−42.01665987607× 10^-4. The resulting end-to-end equal-compute advantage is Δ^endM=2.97479243825×10−4, _ end^M=2.97479243825× 10^-4, (12.7) with one-sided Student-t lower bound 2.00162462592×10−42.00162462592× 10^-4 and bootstrap lower bound 2.04870462705×10−42.04870462705× 10^-4. Table 4: Equal-compute productive-revelation evidence across the two Transformer families. The two Qwen rows use independent productive-path panels and are not pooled. Model / evidence block Productive path End-to-end point One-sided 95% LCB Qwen, independent panel 1 8.9913×10−58.9913× 10^-5 8.9515×10−58.9515× 10^-5 2.9246×10−52.9246× 10^-5 Qwen, independent panel 2 1.4776×10−41.4776× 10^-4 1.4737×10−41.4737× 10^-4 8.8394×10−58.8394× 10^-5 Mistral, independent family 3.0939×10−43.0939× 10^-4 2.9748×10−42.9748× 10^-4 2.0016×10−42.0016× 10^-4 The same decomposition is supported in both families: the restart-only selector difference is not the source of the main gain, while productive reuse of deeper tested computation creates a strict equal-budget end-to-end advantage. This cross-family pattern is consistent with the productive-revelation mechanism formalized in Equation (4.6). 13 Adaptive Revelation Control across model families The dynamic theory predicts a state-dependent stopping decision based on the conditional utility of deeper revelation, not a universal population ambiguity threshold. The empirical question is therefore whether legal shallow information can support compute-saving continuation decisions without changing the terminal utility definition or productive compute ledger. The theory also permits different model families to occupy different conditional revealability regimes. 13.1 Qwen2.5-7B: decision-nonredundant revealability and held-out safe compute In Qwen2.5-7B, the observable H4 revealability proxy is R4=‖(SXi,A4−SKeep,A4,SXi,B4−SKeep,B4,SXi,C4−SKeep,C4)‖2.R_4= \| (S_ Xi,A^4-S_ Keep,A^4,S_ Xi,B^4-S_ Keep,B^4,S_ Xi,C^4-S_ Keep,C^4 ) \|_2. (13.1) We treat R4R_4 only as a legal model-specific proxy; no equality R4=σ4R_4= _4 is assumed. In a supporting analysis with fitting and evaluation performed on separate collection blocks under the same safety protocol, the two-coordinate continuation rule improves paired cost-aware value over the scalar |fH4||f_H4| rule by 3.2514×10−73.2514× 10^-7, with one-sided Student-t and bootstrap lower bounds 4.0357×10−84.0357× 10^-8 and 3.6127×10−83.6127× 10^-8, and no harmful stops in either held-out direction. This is evidence for a decision-nonredundant regime at the declared compute price, not a prospective claim that this particular R4R_4 is unique or universally necessary. A separate 480-history full-grid study on new exact learning-state anchors prospectively fixed the utility-aware continuation architecture, H4/H8 policies, update price, and 16/20-update compute ledger. With an independent second future bank on the same anchors, the H8-over-H4 terminal decision value is ^[Qπ8−Qπ4]AB=1.229×10−4,LCBt,0.95=8.687×10−5>0,LCBboot,0.95=8.769×10−5>0. E[Q_ _8-Q_ _4]_AB=1.229× 10^-4, _t,0.95=8.687× 10^-5>0, _boot,0.95=8.769× 10^-5>0. (13.2) The fixed two-coordinate point controller has positive two-bank net point value but its strict one-sided net-value lower bounds cross zero, illustrating the finite-sample tail sensitivity emphasized by the theory. The bounded-risk certification layer provides a complementary held-out safe-compute result. In a theory-constrained calibration/held-out evaluation, one bank is used for calibration and the other for evaluation. The resulting calibrated rule stops on 4/480 histories, produces no H4/H8 action changes and no harmful stops, has a 95% Bernoulli–KL upper joint-risk bound of 0.6222%<2%0.6222\%<2\%, and saves 0.1667%0.1667\% of visible updates. Its realized panel-level update-priced net value is ΔV^cert=2.890×10−7,LCBt,0.95=5.161×10−8>0,LCBboot,0.95=7.225×10−8>0. V_ cert=2.890× 10^-7, _t,0.95=5.161× 10^-8>0, _boot,0.95=7.225× 10^-8>0. (13.3) 13.2 Mistral-7B-v0.3: scalar continuation control at the tested resolution Mistral-7B-v0.3 does not require numerical transfer of Qwen’s R4R_4. A Mistral-specific continuation head is instantiated using only the separate 336-anchor development panel, while the independent 480-anchor two-bank panel is used for target evaluation without target-panel refitting. Among the two pre-existing shallow continuation architectures evaluated as a finite family, the scalar margin-based estimator retains positive adaptive value at the observed resolution. Relative to always continuing to H8, the scalar utility-aware controller achieves exact-anchor update-priced net value ΔV^scalarM=1.13877630257×10−5. V_ scalar^M=1.13877630257× 10^-5. (13.4) Its ordinary one-sided 95% Student-t and 50,000-draw bootstrap lower bounds are 2.37141363069×10−62.37141363069× 10^-6 and 5.41900958243×10−65.41900958243× 10^-6. Because both scalar and two-coordinate continuation architectures existed before the Mistral target evaluation, we treat them as a two-element finite architecture family. Bonferroni-adjusted familywise 95% lower bounds for the scalar rule remain positive: 6.37735915928×10−76.37735915928× 10^-7 for Student-t inference and 5.23837592968×10−65.23837592968× 10^-6 for the exact-anchor bootstrap. Across 960 bank-level realizations it stops 166 times, changes the H8 action once, and has zero harmful stops under the declared update price. The raw Qwen-style R4R_4 coordinate does not improve Mistral adaptive value over the scalar controller; exploratory development-only proxy diagnostics are excluded from the target-panel claim. We therefore interpret Mistral as consistent with scalar decision sufficiency under Theorem 5.3 and Corollary 5.7, rather than as a failure of the two-coordinate theorem. This empirical conclusion does not assert that σ4 _4 is literally a function of margin; it says only that the finite pre-existing architecture family did not require a second coordinate to improve the priced target-panel stopping decision at the tested resolution. 13.3 Risk–severity separation in Mistral A separate risk-calibrated Mistral shallow-stop rule provides an empirical illustration of Theorem B.3. It makes 351/960 early stops while the 95% Bernoulli–KL joint stop/action-change upper bound is 1.5924%<2%1.5924\%<2\% in each bank. Yet the exact-anchor update-priced net point value is slightly negative, ΔV^riskM=−3.3428×10−7, V_ risk^M=-3.3428× 10^-7, (13.5) with negative one-sided lower bounds. Only four stopped action changes occur, but the largest absolute continuation value is 9.811×10−39.811× 10^-3. Thus low action-instability probability and positive utility are empirically distinct objects: rare high-severity misses can dominate many cheap correct stops. This is precisely the risk–severity separation formalized in Appendix B. Taken together, the two model families support the same adaptive-control principle while occupying different shallow-state regimes. Qwen exhibits evidence for a decision-nonredundant revealability coordinate; Mistral admits a successful scalar continuation rule. What transports is the conditional revelation-value decision, not a universal empirical coordinate system. 14 Cross-model structural validation The two model families are not expected to reproduce identical coefficients, thresholds, or empirical revealability coordinates. The relevant replication object is the sequence of theory-level implications: shallow information can alias decision-relevant state; deeper revelation can have positive decision value; productive reuse can convert that information into equal-compute utility; revelation depth can be controlled state by state; and risk control is not equivalent to utility control when severity is unrestricted. Table 5: Theory-level evidence across the two Transformer families. The fixed-depth and equal-compute Mistral results use an independent target panel; the adaptive row evaluates a pre-existing two-architecture family with familywise-adjusted inference. Theoretical implication Qwen2.5-7B evidence Mistral-7B-v0.3 evidence Decision-relevant shallow aliasing Fixed H4 ambiguity cell contains H8-refined groups with opposite terminal action gaps in two independent banks. Using the H4 threshold fixed from development data, the bank-A-defined ambiguity cell contains 216/480 anchors, including 88 stable-Xi and 58 stable-Keep anchors whose terminal-gap signs agree across both banks (descriptive structural check). Positive value of deeper revelation H8 beats the declared current-information library on two independent 336-anchor panels; the separate 480-history H8-over-H4 value is strictly positive. Exact-anchor H8-over-H4 value 2.2509×10−42.2509× 10^-4 with positive t and bootstrap lower bounds. Productive information-to-utility closure Productive path value is positive in two independent panels and yields strict 20-update end-to-end advantage over the strengthened restart frontier. Same-path reuse 3.0939×10−43.0939× 10^-4 and equal-compute end-to-end advantage 2.9748×10−42.9748× 10^-4, both with positive lower bounds. Adaptive Revelation Control Supporting evidence for incremental decision value from a nonredundant revealability coordinate; a held-out bounded-risk safe-compute evaluation has positive panel net-value lower bounds. Consistent with scalar decision sufficiency at the tested resolution; a familywise structural validation of the scalar continuation architecture, fit only on the separate development panel, retains positive Bonferroni-adjusted exact-anchor t and bootstrap lower bounds within the two-element pre-existing architecture family. Risk–severity boundary Positive held-out safe-compute behavior is explicitly limited to panel-level evidence without an unrestricted tail guarantee. Sub-2% stop–flip risk coexists with nonpositive risk-controller net value because a few flips have high severity. The common pattern is therefore stronger than numerical parameter reuse. In both families, future learning supplies decision-relevant information, deeper tested paths carry productive value, and the resulting information can improve compute-accounted utility. The model-specific difference occurs one level lower: Qwen provides evidence of incremental decision value from an additional shallow revealability coordinate, whereas Mistral is consistent with scalar decision sufficiency at the observed resolution. Theorem 5.3 shows why these outcomes belong to the same theory: an extra coordinate is required for control only when states sharing the scalar summary cross the priced Stop/Continue boundary. Corollary 5.7 gives the corresponding location–scale criterion, while Definition 5.9 allows model-specific empirical realizations. Accordingly, the cross-model conclusion is not that one fitted controller or proxy transfers unchanged. It is that the Revelation-Control relation between legal shallow information, continuation value, productive cost, and terminal decision utility survives an independent model-family change. This is the level at which the empirical evidence supports the general theory’s central testable implications. 15 Potential application domains The empirical results in this paper concern learning systems, but the decision-theoretic structure is more general. A Revelation-Control problem can be formulated when five ingredients can be defined: a consequential terminal action; hidden state not fully resolved by current legal information; a priced intervention that can be executed before commitment; an observable response that can refine the relevant decision quotient; and a cost ledger that distinguishes information acquisition from any useful state change left by the intervention. The strongest form of the framework arises when the experiment is productive: if the tested path is selected, some of the work performed to reveal state is itself part of the useful continuation. These conditions are substantially narrower than “any sequential decision problem,” and the examples below are prospective application domains rather than deployment claims established by the present experiments. 15.1 Learning and computational systems Foundation-model adaptation and data-domain allocation. Fine-tuning a large model often requires choosing among domains, task mixtures, adapters, or continuation depths under a limited compute budget. Two checkpoints can look similar under current loss or held-out behavior yet respond differently to the same candidate domain update. A short domain-specific training intervention can therefore serve as a future-learning experiment: its response can refine the decision quotient over which domain or adapter should be promoted, while the selected probe trajectory can remain useful training. This differs from using a learning curve only to forecast final loss; the target is whether newly revealed learning state changes a downstream adaptation decision. The DynaMiCS-style comparison studied here is one concrete instance of this broader allocation problem [14]. Continual learning, curriculum control, rehearsal allocation, and AutoML. In continual or curriculum learning, present task performance need not reveal how much plasticity, interference, or recoverability remains for the next phase. A bounded amount of rehearsal or target-task training can be treated as an experiment on that hidden state before the curriculum decision is fixed. The same principle applies to hyperparameter and configuration search: partial-training methods already pause, terminate, or promote candidate runs from intermediate performance [37, 7, 25, 17], but Revelation Control asks whether additional computation should be purchased because its response is expected to change the decision. A run with a mediocre current metric can justify continuation when deeper revelation has positive decision value, while a promising-looking run can be stopped when additional information is unlikely to alter the selected action. For long-horizon or lifelong learning, an important extension is to identify a minimal revelation state that remains dynamically sufficient across successive stages; if such a state satisfies multi-stage Markov closure, Revelation Control could be composed recursively rather than applied only as a one-step continuation rule. Productive promotion further distinguishes useful exploratory work from discarded evaluation cost. Intervention-aware data acquisition. The intervention can also be a training batch, synthetic-data block, augmentation regime, rehearsal set, or other legal update whose response is observable before a larger commitment is made. Classical Bayesian experimental design and active sensing choose observations by expected decision value [26, 6, 39]. In a learning system, the experiment can instead be an update to the learner itself. Development-only future responses can be used to estimate a model-specific revealability coordinate or continuation state and to determine whether current information is already decision-sufficient. This suggests a route to choosing not merely which data are informative, but which data reveal something about the learner’s future response that changes the action menu. Adaptive computation beyond training. The depth variable need not literally count gradient updates. In an iterative solver, multi-fidelity simulation, branch-and-bound routine, Monte Carlo computation, or other staged numerical procedure, a coarse computation may be insufficient to choose a downstream action while additional computation both reveals information and accumulates reusable work. If the current approximation, the refinement response, and the terminal decision can be placed in a common utility ledger, revelation depth becomes computation depth. The same continuation principle then asks whether another refinement level is worth purchasing because of the decision it may change, connecting Revelation Control to value-of-computation ideas [36, 13] while retaining the productive-reuse term. 15.2 Beyond machine learning: productive experiments in sequential decisions Control, system identification, and robotics. A physical system can occupy internal states with nearly identical current outputs but different future responses and therefore different optimal controls. Examples include uncertain friction, payload, contact mode, actuator health, or local dynamics. Dual control already formalizes the fact that an action can simultaneously control a system and reveal uncertainty [11], while POMDP and active-sensing formulations treat task-directed information acquisition under partial observability [21, 39]. Revelation Control suggests a decision-relative refinement of this idea when a short controlled motion or excitation can be priced, its response can reveal which control action is appropriate, and the selected probing trajectory can be retained rather than discarded. The theory does not require complete system identification: only the hidden distinctions that change the terminal action need to be resolved. Industrial process control, diagnostics, and maintenance. Machines or production processes can present similar current telemetry while differing in latent wear, loading, material state, or failure proximity. A short controlled operating segment, calibration cycle, or diagnostic excitation can reveal whether the correct action is to continue production, derate, switch operating mode, inspect, or maintain. Such settings are especially close to productive revelation when the diagnostic segment also performs useful work. The relevant accounting must, however, include downtime, safety margin, irreversible wear, and state-restoration cost; if probing itself damages the system or changes the target regime, the promotion geometry used in the present experiments no longer applies without modification. Scientific experimentation and automated laboratories. In materials, chemistry, biology, and other experimental sciences, an early observation may leave several latent mechanisms compatible with the data even though those mechanisms imply different next experiments or operating decisions. Bayesian experimental design already provides a mature decision-theoretic language for choosing informative experiments [26, 6]. Revelation Control is potentially relevant when an experiment can be staged: an initial intervention is run only far enough to determine whether deeper measurement or continuation is decision-relevant, and the partial trajectory can be retained if that branch is selected. Automated laboratories, adaptive measurement pipelines, and multi-stage physical experiments are natural settings in which “how much experiment to purchase” may be as important as “which experiment to run.” Establishing this connection rigorously would require domain-specific transition models, measurement error, and experimental-cost accounting beyond the present paper. Operations, resource allocation, and pilot-to-scale decisions. Many operational decisions involve latent regimes that cannot be resolved from passive current data alone. A limited allocation, pilot deployment, test market, routing perturbation, or operating-policy trial can reveal response heterogeneity before a larger commitment is made. The Revelation-Control abstraction is appropriate when the pilot has a well-defined terminal decision, its response can be evaluated before scale-up, and at least part of the pilot’s value or infrastructure is reusable if promoted. The central distinction is between learning about the world and learning enough to choose the action: a pilot need not identify the full latent regime if it already resolves the relevant decision quotient. Strategic behavior, interference, nonstationarity, and general-equilibrium effects can break the simple controlled-probe assumptions and would have to be modeled explicitly. High-stakes human decision systems. The mathematical pattern also appears, at least abstractly, in adaptive treatment, public-policy pilots, and personalized education: two individuals or populations with similar current observables can respond differently to a small intervention, and the response can change the preferred next action. These domains are intentionally placed at the outer boundary of the present paper’s application claims. A valid extension would require causal identification, ethical constraints, uncertainty about intervention harm, fairness considerations, and stronger safety guarantees than the bounded-risk analysis developed here. In medicine or policy especially, “productive probing” cannot be assumed merely because an intervention also has intended benefit; the value ledger must price adverse effects and irreversibility, and experimentation may be impermissible even when its statistical information value is high. The present work therefore supplies only a possible decision-theoretic template, not a deployment prescription. 15.3 A domain-level applicability test The most useful transfer question is not whether a new field resembles model fine-tuning superficially, but whether its decision structure matches the theory. For a new system, one should first declare the legal current information, the terminal action menu, terminal utility, admissible probe interventions, and all costs or safety penalties. Development-only probe responses can then be used to estimate the conditional continuation value and to test whether the current scalar summary is already decision-sufficient. If scalar fibers cross the priced Stop/Continue boundary, additional revealability information is decision-nonredundant and a richer controller is required; if they do not, a simpler controller can be sufficient even when hidden-state variation remains. The resulting architecture and cost ledger should be fixed before target evaluation. If the probe leaves reusable work, information value and productive reuse must be reported separately; if it does not, the problem reduces to discard-and-restart information acquisition. This procedure is the general method that transfers across domains: the decision factorization, continuation-value criterion, and accounting principles are invariant, whereas observable proxies, dynamics, coefficients, safety constraints, and prices are system-specific. 16 Limitations Scope of the empirical systems. The empirical study now spans two independently instantiated 7B decoder-only Transformer families, but both use LoRA adaptation, AdamW-style optimizer state, the same binary intervention menu, and the same H4/H8/H12 control geometry. Cross-optimizer, substantially different-scale, non-Transformer, and materially different-horizon generality remain open. Comparator scope. The finite-library result does not imply dominance over every measurable current-information policy. The frontier comparison is against one strengthened DynaMiCS-style short-probe/restart class under a common policy-visible update budget, not against every possible probing or metareasoning method. Productive revelation assumptions. Promotion is useful only when tested work remains legally and scientifically reusable. If a probe corrupts the deployment path, creates irreversible risk, changes the target distribution, or incurs additional state-I/O or safety cost, restart may be the appropriate technology and the budget identity must be repriced. Conditional rather than universal revealability dimension. The location–scale theorem is a structural statement about (|mh|,σh)(|m_h|, _h), not a claim that every model must exhibit two empirically nonredundant shallow coordinates. The exact scalar-control factorization theorem further shows that residual variation in σh _h need not be decision-relevant: a second coordinate is required only when a positive-mass scalar fiber crosses the priced Stop/Continue boundary. Qwen provides evidence for a decision-nonredundant regime, while Mistral is consistent with scalar decision sufficiency at the observed resolution. The empirical R4R_4 used in Qwen is one legal model-specific proxy, not a universal statistic. Tail-robust utility certification. Controlling the probability that an early stop changes the deeper action does not by itself certify positive expected utility when stopped-flip severity is unrestricted. Theorem B.3 shows that this is an identification boundary rather than merely a power limitation: a population net-utility certificate additionally requires a predeclared severity, moment, or integrable-tail class. Positive held-out panel-level utility is therefore not a universal tail-robust guarantee. Applications beyond learning systems. The mathematical decision structure can be instantiated whenever a priced controlled intervention reveals decision-relevant hidden state before commitment, but the empirical evidence in this paper is confined to learning systems. Physical, scientific, operational, medical, policy, and educational settings introduce domain-specific causality, irreversibility, strategic response, measurement error, safety, fairness, and ethical constraints that are not resolved here. The broader domains in Section 15 are therefore prospective uses of the framework, not validated deployments or claims of domain-universal optimality. Method optimality. The evaluated policies are selected using development data and fixed before target evaluation. No result establishes global optimality over all representations, learners, stopping rules, probe schedules, intervention menus, or acquisition policies. Compute and deployment value. The main normalization counts policy-visible training updates under a common model and batch construction. Restore/save overhead is not charged against the frontier in the main comparison. Device-seconds, wall-clock, memory, state-I/O, monetary value, safety value, and live-deployment utility are different resources and are not collapsed into the primary claim. Dynamic-state compression. The location–scale result is an exact one-step statement. One-step low-dimensional decision sufficiency does not imply multi-stage Markov closure. A multi-depth controller can therefore use (mh,σh)(m_h, _h) as a complete low-dimensional state only under an additional Markov-closure condition for future margin–revealability dynamics. Characterizing minimal dynamically sufficient revelation states, and the conditions under which they induce Markov closure across revelation depths, is a natural next theoretical problem. 17 Conclusion The companion work established that present behavior can hide distinctions in future learning. The present paper closes the decision-theoretic step: such distinctions can have action value, can be actively revealed by controlled future learning, and can become a state-dependent control variable. We call this problem Revelation Control. The theory identifies the relevant object as a decision quotient rather than the full hidden state. A coarse observation is sufficient only when the refined optimal action factorizes through it. The fork–probe–promote realization studied here uses future learning as an experiment on hidden learner state, while productive revelation keeps the selected tested path as useful computation. The resulting value ledger separates pure information gain from the value of already executed work. Static Bayes refinement value then closes exactly into a conditional revelation value, yielding an optimal-stopping rule over revelation depth. The empirical evidence supports this structure in two independently instantiated Transformer families. In Qwen2.5-7B, a fixed H8 policy strictly beats a declared 22-component current-information library on two independent panels, refines a fixed H4 ambiguity regime into groups with opposite terminal action value, and obtains strict equal-compute advantage over a strengthened DynaMiCS-style H4 restart frontier through productive reuse. In Mistral-7B-v0.3, an independent two-bank panel reproduces positive H8-over-H4 decision value and the same productive equal-compute mechanism with positive one-sided lower bounds. The cross-model result is therefore not numerical reuse of one fitted policy: it is replication of the information-to-utility structure predicted by the theory. Adaptive control exhibits the regime dependence predicted by the theory. Qwen provides evidence that a second shallow revealability coordinate can matter beyond current margin; a familywise-adjusted Mistral structural validation, restricted to the finite pre-existing architecture family and fit only on independent development data, supports a successful scalar continuation rule at the tested resolution. The exact scalar-control factorization result makes the common structure precise: an additional coordinate is necessary for the stopping decision only when states sharing the scalar summary fall on both sides of the cost-adjusted continuation boundary. Thus predictive or value variation can remain present without being decision-nonredundant. A separate Mistral risk-calibrated controller further illustrates the impossibility boundary: sub-2% decision-instability risk can coexist with nonpositive net utility when a few stopped flips have large severity. The resulting conclusion is specific but broad in implication. Hidden learning state is not merely predictive structure. Controlled future learning can reveal distinctions that change decisions; productive reuse can convert revelation into higher utility at the same visible training-update budget; and the amount of revelation can itself be controlled according to state-dependent expected value. More generally, Section 15 identifies the same decision pattern in control and system identification, robotics and industrial diagnostics, scientific experimentation, adaptive computation, and pilot-to-scale operational decisions, with high-stakes human domains requiring additional causal, ethical, and safety theory. What must transport across systems is the decision structure and value accounting, not a universal proxy, coefficient, threshold, or even a fixed empirical state dimension. Revelation Control therefore provides a theory-first framework for deciding not only what action to take, but how much controlled future interaction it is worth purchasing in order to know. Appendix A Conservative replicate-denoised Bayes-regret bridge This appendix records an optional stronger certification device. It is not required for the main empirical conclusions in Sections 11–12. Independently generated future outcomes can be used to separate persistent conditional signal from one-bank realization noise and to upper-bound coarse-policy Bayes regret under the stated assumptions. Consider binary actions with terminal gap D=QXi−QKeep.D=Q_ Xi-Q_ Keep. (A.1) Let future banks A and B share a pre-bank sigma-field X and satisfy [DA∣]=[DB∣]=:μ.E[D_A X]=E[D_B X]=:μ. (A.2) The legal coarse information obeys ℋ⊆ H X, and mℋ=[μ∣ℋ]m_ H=E[μ H] (A.3) provides the coarse-information Bayes score. Let f be any fixed ℋ H-measurable score in the same utility units as D, and let πf=f>0. _f=1\f>0\. (A.4) Here and below, action indicator one denotes Xi and zero denotes Keep. Define the cross-replicate residual moment M×(f)=[(DA−f)(DB−f)].M_×(f)=E[(D_A-f)(D_B-f)]. (A.5) Let c=Cov(DA,DB∣).c_ X=Cov(D_A,D_B X). (A.6) Conditional independence gives c=0c_ X=0; allowing c≥0c_ X≥ 0 yields the same conservative direction. Theorem A.1 (Replicate-denoised Bayes-regret bridge). Assume square integrability, a shared conditional mean across banks, and c≥0c_ X≥ 0 almost surely. Then M×(f) M_×(f) =[(μ−f)2]+[c] =E[(μ-f)^2]+E[c_ X] (A.7) =[(mℋ−f)2]+[Var(μ∣ℋ)]+[c], =E[(m_ H-f)^2]+E[Var(μ H)]+E[c_ X], (A.8) and therefore VB(ℋ)−(πf)≤M×(f). V_B( H)-V( _f)≤ M_×(f). (A.9) Consequently, for any policy πG _G measurable with respect to a refinement ⊇ℋ G H, (πG)−VB(ℋ)≥(πG)−(πf)−M×(f).V( _G)-V_B( H)≥ \V( _G)-V( _f) \- M_×(f). (A.10) The decomposition shows why the bridge is conservative. Even when f=mℋf=m_ H, persistent hidden heterogeneity contributes the nonnegative floor Vhid:=[Var(μ∣ℋ)].V_hid:=E[Var(μ H)]. (A.11) The bridge is therefore sufficient rather than necessary: not clearing it does not imply that the active value is below the coarse-information Bayes value. Corollary A.2 (Conditional bridge). Let A∈ℋA∈ H with ℙ(A)>0P(A)>0. Replacing all expectations by expectations conditional on A gives VB(ℋ,A)−(πf,A)≤M×(f,A),V_B( H;A)-V( _f;A)≤ M_×(f;A), (A.12) where M×(f,A)=[(DA−f)(DB−f)∣A].M_×(f;A)=E[(D_A-f)(D_B-f) A]. (A.13) Appendix B Adaptive-depth certification details This appendix records the structural certification layer used to interpret adaptive revelation depth. The generic risk-calibration tools are established prior art; the purpose here is to state how they interact with productive-prefix reuse and terminal decision severity in the present execution technology. Let S∈0,1S∈\0,1\ denote an H4 early-stop decision, let F=π4≠π8F=1\ _4≠ _8\, and define Y=Qπ8−Qπ4Y=Q_ _8-Q_ _4. Each H4 stop saves compute value c=4λ0c=4 _0 relative to always-H8. Proposition B.1 (Productive-prefix decision invariance). Under the declared same-path promotion technology, if F=0F=0 then Y=0Y=0. For the binary action menu, Y=D(π8=Xi−π4=Xi),|Y|=|D|F.Y=D (1\ _8= Xi\-1\ _4= Xi\ ), |Y|=|D|F. (B.1) Proposition B.2 (Risk–severity factorization). Let ρ=Pr(S=1)ρ= (S=1) and r=Pr(S=1,F=1)r= (S=1,F=1). When r>0r>0, define the signed stopped-flip severity ηs=[Y∣S=1,F=1] _s=E[Y S=1,F=1] and the positive-harm severity η+=[Y+∣S=1,F=1] _+=E[Y_+ S=1,F=1], where Y+=max(Y,0)Y_+= (Y,0). When r=0r=0, set ηs=η+=0 _s= _+=0 by convention. Then ΔV(S)=[S(c−Y)]=cρ−rηs. V(S)=E[S(c-Y)]=cρ-r _s. (B.2) Moreover, ΔV(S)≥cρ−rη+. V(S)≥ cρ-r _+. (B.3) Theorem B.3 (Risk-only impossibility and the harm-tail boundary). Let E=S=1,F=1E=\S=1,F=1\ and define the stopped positive-harm mass H+:=[Y+E]∈[0,∞].H_+:=E[Y_+1_E]∈[0,∞]. (B.4) Then H+=∫0∞Pr(E,Y+>t)t,ΔV(S)≥cρ−H+.H_+= _0^∞ (E,\,Y_+>t)\,dt, V(S)≥ cρ-H_+. (B.5) Moreover, fix any 0<r≤ρ≤10<r≤ρ≤ 1. For every M>0M>0 there exists a joint law of (S,F,Y)(S,F,Y) satisfying productive-prefix decision invariance, with Pr(S=1)=ρ (S=1)=ρ, Pr(E)=r (E)=r, and finite [Y+]E[Y_+], such that ΔV(S)<−M. V(S)<-M. (B.6) Hence, whenever nonzero stop–flip risk is permitted, exact knowledge of (ρ,r)(ρ,r) alone gives no finite distribution-robust lower bound on expected net utility over a class with unrestricted stopped-flip severity. Any positive expected-utility certificate must additionally control H+H_+, or a sufficient severity/tail surrogate for it. Corollary B.4 (Minimal sufficient harm-tail bridges). The positive-harm component left uncontrolled by stop–flip risk is summarized by the integrated stopped-harm tail H+H_+ in Equation (B.4). The following declared conditions are sufficient ways to control it for a lower utility certificate. If Y+≤B+Y_+≤ B_+ almost surely, then ΔV(S)≥cρ−B+r. V(S)≥ cρ-B_+r. (B.7) If instead q>1q>1 and [|D|q]≤MqE[|D|^q]≤ M_q, then Hölder’s inequality gives ΔV(S)≥cρ−Mq1/qr1−1/q. V(S)≥ cρ-M_q^1/qr^1-1/q. (B.8) More generally, if an integrable envelope u:[0,∞)→[0,∞)u:[0,∞)→[0,∞) satisfies Pr(E,Y+>t)≤u(t)for all t≥0, (E,\,Y_+>t)≤ u(t) all t≥ 0, (B.9) then ΔV(S)≥cρ−∫0∞u(t)t. V(S)≥ cρ- _0^∞u(t)\,dt. (B.10) Thus a finite-sample event ρ≥Lρ≥ L_ρ, r≤Urr≤ U_r yields certified lower bounds cLρ−B+UrcL_ρ-B_+U_r or cLρ−Mq1/qUr1−1/qcL_ρ-M_q^1/qU_r^1-1/q under the corresponding declared severity class; an independently valid upper bound UHU_H on H+H_+ yields the assumption-matched bound cLρ−UHcL_ρ-U_H. Theorem B.3 is a control-specific boundary, not a claim that bounded-risk calibration is impossible. The Bernoulli event E can be calibrated without utility-tail assumptions; what cannot be obtained from that event probability alone is a nontrivial lower bound on an unbounded mean severity. At the statistical level, this distinction is aligned with classical nonparametric impossibility results for unrestricted mean inference [3]. Bounded decision-instability calibration. The loss L=S=1,F=1L=1\S=1,F=1\ is Bernoulli. Hence fixed-sequence calibration or other established bounded-risk procedures such as Learn-then-Test and Conformal Risk Control can be used prospectively once the risk model, threshold family, calibration data, and risk target are fixed [1, 2]. This controls how often early stopping can change the deeper action, not how severe those changed decisions are. Fixed-data risk–severity diagnostic. A fixed-data nested evaluation illustrates the distinction. A two-coordinate H4 flip-risk score is trained on one collection block, calibrated on one future bank of the other block to target 2% joint stop-and-flip risk, and checked on the remaining future bank. Table 6 reports the held-out outcomes. The first direction has positive realized net value; the second has a lower flip rate but negative realized net value because its single harmful flip has substantially larger severity. This is a diagnostic validation of Equation (B.2), not a prospective controller claim. Table 6: Held-out risk–severity diagnostic for a derived risk-calibrated stopping rule. Net value is relative to always-H8. Direction Stop rate Stop–flip rate Harm events Max harm severity Net value A → B 26.67% 0.833% 1 2.083×10−42.083× 10^-4 +1.186×10−5+1.186× 10^-5 B → A 11.67% 0.417% 1 2.010×10−32.010× 10^-3 −4.331×10−6-4.331× 10^-6 The diagnostic shows that a probability budget is not a utility budget: decision-instability risk controls how often early stopping can differ from H8, while the integrated harm tail controls how costly those differences are. Theorem B.3 makes this separation sharp. A fully prospective expected-utility certificate therefore needs both a bounded-risk calibration and a predeclared severity, moment, or integrable tail class. Appendix C Proofs C.1 Proof of Proposition 2.2 Let Ma=[Qa∣]M_a=E[Q_a G]. Because A is finite, a G-Bayes action exists after fixing a measurable tie rule. If a G-Bayes action πH _H is ℋ H-measurable, then VB()=[MπH]=[QπH]≤VB(ℋ)≤VB(),V_B( G)=E[M_ _H]=E[Q_ _H]≤ V_B( H)≤ V_B( G), (C.1) so equality holds. Conversely, let πH⋆ _H be an ℋ H-Bayes policy. If VB(ℋ)=VB()V_B( H)=V_B( G), then 0=VB()−(πH⋆)=[maxaMa−MπH⋆].0=V_B( G)-V( _H )=E\! [ _aM_a-M_ _H ]. (C.2) The integrand is nonnegative, hence it vanishes almost surely and πH⋆ _H is also G-Bayes optimal. C.2 Proof of Proposition 2.3 By definition, tDF _t DF =Lt+maxaνt(a)+γωt(a)−νt⋆ =L_t+ _a\ _t(a)+γ _t(a)\- _t (C.3) =Lt+maxaγωt(a)−(νt⋆−νt(a)) =L_t+ _a\γ _t(a)-( _t - _t(a))\ (C.4) =Lt+maxaγωt(a)−dt(a). =L_t+ _a\γ _t(a)-d_t(a)\. (C.5) The maximum term is nonnegative because any coarse-optimal action has dt(a)=0d_t(a)=0 and ωt(a)≥0 _t(a)≥ 0. Since Lt≥0L_t≥ 0, the sum is zero if and only if both Lt=0L_t=0 and every term inside the maximum is nonpositive. C.3 Proof of Proposition 4.1 Temporary probing spends mcmc updates across m candidates and then restarts the selected action for the full horizon H, so KF(c)=mc+HK_F(c)=mc+H. Probe-and-promote revelation spends mhmh updates across candidates and then only H−hH-h further updates on the selected path, so KAR(h)=H+(m−1)hK_ AR(h)=H+(m-1)h. Equating the budgets gives (m−1)h=mc(m-1)h=mc. C.4 Proof of Proposition 6.1 Let VΠs⋆=VPs⋆−RsappV_ _s =V_P_s -R_s app be the best value in the learner class and let Ws,n=VΠs⋆−Rs,nestW_s,n=V_ _s -R_s,n est be the learned gross value. Since VPs⋆=VB(ℋ)+Δsrev+ΔspromV_P_s =V_B( H)+ _s rev+ _s prom and Ws,nnet=Ws,n−CsW_s,n net=W_s,n-C_s, substitution yields Equation (6.3). C.5 Proof of Corollary 6.2 Using VP(π^A)=VB()−RA+PAV_P( π_A)=V_B( G)-R_A+P_A and VR(π^F)=VB(ℋ)−RFV_R( π_F)=V_B( H)-R_F, VP(π^A)−VR(π^F)=[VB()−VB(ℋ)]+PA+RF−RA.V_P( π_A)-V_R( π_F)=[V_B( G)-V_B( H)]+P_A+R_F-R_A. (C.6) Subtracting priced resource costs gives the net-value version. C.6 Proof of Theorem 3.2 Relative to Keep, the refined Bayes increment is [(mG)+]E[(m_G)_+] and the coarse Bayes increment is [(mH)+]E[(m_H)_+]. Conditioning the refined increment on ℋ H gives aHa_H. Since mH=[mG∣ℋ]=aH−bH,m_H=E[m_G H]=a_H-b_H, (C.7) we have pointwise aH−(mH)+=aH−(aH−bH)+=minaH,bH.a_H-(m_H)_+=a_H-(a_H-b_H)_+= \a_H,b_H\. (C.8) Taking expectations proves Equation (3.8). On the event A in Equation (3.9), aH≥δα,bH≥δβ,a_H≥δα, b_H≥δβ, (C.9) so minaH,bH≥δmin(α,β) \a_H,b_H\≥δ (α,β) there. Integrating over A gives Equation (3.10). C.7 Proof of Proposition 3.3 By the constant-rank theorem, the local level set Φc−1(Φc(x0)) _c^-1( _c(x_0)) is a submanifold with tangent space kerDΦc(x0) D _c(x_0). Hence there is a C1C^1 curve γ(s)γ(s) in that level set with γ(0)=x0γ(0)=x_0 and γ′(0)=vγ (0)=v. Taylor expansion gives g(γ(s))=sDg(x0)v+o(s),g(γ(s))=s\,Dg(x_0)v+o(s), (C.10) so for sufficiently small positive and negative s the terminal gap has opposite signs. Likewise Yh(γ(s))=Yh(x0)+sDYh(x0)v+o(s),Y_h(γ(s))=Y_h(x_0)+s\,DY_h(x_0)v+o(s), (C.11) which changes to first order because DYh(x0)v≠0DY_h(x_0)v≠ 0. C.8 Proof of the replicate-denoised bridge By iterated expectation, M×(f) M_×(f) =[[(DA−f)(DB−f)∣]] =E\! [E[(D_A-f)(D_B-f) X] ] (C.12) =[(μ−f)2]+[c]. =E[(μ-f)^2]+E[c_ X]. (C.13) Since f is ℋ H-measurable and mℋ=[μ∣ℋ]m_ H=E[μ H], the conditional Pythagorean identity gives [(μ−f)2]=[(mℋ−f)2]+[Var(μ∣ℋ)].E[(μ-f)^2]=E[(m_ H-f)^2]+E[Var(μ H)]. (C.14) For binary actions with the declared tie rule πf=f>0 _f=1\f>0\, the regret relative to the ℋ H-Bayes rule πm=mℋ>0 _m=1\m_ H>0\ is VB(ℋ)−(πf)=[|mℋ| 1πf≠πm].V_B( H)-V( _f)=E\! [|m_ H|\,1\ _f≠ _m\ ]. (C.15) On action disagreement, |mℋ|≤|mℋ−f||m_ H|≤|m_ H-f|; ties at mℋ=0m_ H=0 contribute zero regret. Cauchy–Schwarz therefore yields VB(ℋ)−(πf)≤|mℋ−f|≤[(mℋ−f)2]≤M×(f).V_B( H)-V( _f) |m_ H-f|≤ E[(m_ H-f)^2]≤ M_×(f). (C.16) Adding and subtracting (πf)V( _f) proves the policy-improvement inequality. Conditioning throughout on A∈ℋA∈ H proves the conditional version. C.9 Proof of Proposition 5.1 By the tower property, [h→h′] [G_h→ h ] =[[Ch′∣ℱh]−Ch] =E\! [E[C_h F_h]-C_h ] (C.17) =[Ch′]−[Ch] =E[C_h ]-E[C_h] (C.18) =VB(ℱh′)−VB(ℱh). =V_B( F_h )-V_B( F_h). (C.19) The final equality is the definition of Bayes value under each information state. C.10 Proof of Proposition 5.4 At the final admissible depth hJh_J, no deeper revelation is available, so VJdyn=ChJV_J dyn=C_h_J. At an earlier depth, the controller has exactly two admissible choices in the stated finite-depth problem: commit immediately for value ChjC_h_j, or buy the next refinement, pay cjc_j, and receive conditional expected value [Vj+1dyn∣ℱhj]E[V_j+1 dyn F_h_j]. Taking the larger value gives Equation (5.14). Backward induction completes the recursion. C.11 Proof of Proposition 5.2 Given ℱh F_h, the conditional net increment from continuing is Gh→h′−ch→h′G_h→ h -c_h→ h ; hence continuation is optimal exactly when this quantity is positive. C.12 Proof of Theorem 5.3 Write Γ=Γh→h′ = _h→ h and condition on the scalar summary ShS_h. Since Γ=Γ+−(−Γ)+, = _+-(- )_+, (C.20) we have [Γ∣Sh]=a(Sh)−b(Sh).E[ S_h]=a(S_h)-b(S_h). (C.21) Therefore Vmeta(ℱh)−Vmeta(Sh) V_ meta( F_h)-V_ meta(S_h) =[a(Sh)]−[(a(Sh)−b(Sh))+] =E[a(S_h)]-E[(a(S_h)-b(S_h))_+] (C.22) =[a(Sh)−(a(Sh)−b(Sh))+] =E\! [a(S_h)-(a(S_h)-b(S_h))_+ ] (C.23) =[mina(Sh),b(Sh)], =E[ \a(S_h),b(S_h)\], (C.24) which proves Equation (5.11). Because a,b≥0a,b≥ 0, the gap is zero if and only if mina,b=0 \a,b\=0 almost surely. For an integrable random variable, a(Sh)=0a(S_h)=0 exactly when Pr(Γ>0∣Sh)=0 ( >0 S_h)=0 almost surely on that scalar fiber, and similarly b(Sh)=0b(S_h)=0 exactly when Pr(Γ<0∣Sh)=0 ( <0 S_h)=0. Thus equality holds if and only if no positive-mass scalar fiber contains both strictly positive and strictly negative net continuation gains. Equivalently, there exists a full-information optimal meta-action whose tie assignment at Γ=0 =0 is measurable with respect to ShS_h; that optimal meta-action then factorizes through ShS_h. This is precisely value-preserving decision factorization for the two-action meta-menu STOP,CONTINUE\STOP,CONTINUE\. C.13 Proof of Theorem 5.5 By symmetry and integrability of Z, [Z]=0E[Z]=0. If mh≥0m_h≥ 0, then [(mh+σhZ)+∣ℱh]−mh [(m_h+ _hZ)_+ F_h]-m_h =[(−mh−σhZ)+∣ℱh] =E[(-m_h- _hZ)_+ F_h] (C.25) =σh[(Z−mh/σh)+], = _hE[(Z-m_h/ _h)_+], (C.26) where the last equality uses symmetry. If mh<0m_h<0, the coarse positive-part value is zero and [(mh+σhZ)+∣ℱh]=σh[(Z−|mh|/σh)+].E[(m_h+ _hZ)_+ F_h]= _hE[(Z-|m_h|/ _h)_+]. (C.27) This proves Equation (5.19). Since ψ′(t)=−Pr(Z>t)ψ (t)=- (Z>t) at continuity points, ∂ℛ∂|mh|=ψ′(t)=−Pr(Z>t). ∂|m_h|=ψ (t)=- (Z>t). (C.28) For t=|mh|/σht=|m_h|/ _h, ∂ℛ∂σh=ψ(t)−tψ′(t)=ψ(t)+tPr(Z>t)≥0 ∂ _h=ψ(t)-tψ (t)=ψ(t)+t (Z>t)≥ 0 (C.29) with strict inequality whenever the refinement retains positive tail mass beyond t. C.14 Proof of Corollary 5.6 If σh=g(Mh) _h=g(M_h) on E, substitution into Equation (5.19) gives ℛh→h′=g(Mh)ψ(Mh/g(Mh))=r~(Mh)R_h→ h =g(M_h)ψ\! (M_h/g(M_h) )= r(M_h) (C.30) on E, proving scalar sufficiency for the one-step continuation value. A threshold representation is a stronger statement and requires the resulting scalar value function to be monotone relative to cost. For the second part, condition on Mh=mM_h=m. By assumption, the conditional distribution of σh _h is nondegenerate on a set of m values with positive probability, and s↦r(m,s)s r(m,s) is strictly increasing over that conditional support. A strictly monotone transform of a nondegenerate random variable is nondegenerate. Equation (5.19) identifies this transform with ℛh→h′R_h→ h , so the conditional law of ℛh→h′R_h→ h given MhM_h is nondegenerate on that set and no MhM_h-measurable scalar function can equal the continuation value almost surely there. C.15 Proof of Corollary 5.7 For fixed Mh=mM_h=m, let Γ=ℛh→h′−c=r(m,σh)−c. =R_h→ h -c=r(m, _h)-c. (C.31) Theorem 5.3 gives scalar sufficiency exactly when, conditional on almost every m, Γ does not have both strictly positive and strictly negative mass. This is equivalent to Equation (5.26): either there is no strictly positive mass, or there is no strictly negative mass. Zero-gain ties may be assigned to either optimal side without changing value. Under continuity and strict monotonicity, r(m,s)≤cr(m,s)≤ c for s≤σc(m)s≤ _c(m) and r(m,s)≥cr(m,s)≥ c for s≥σc(m)s≥ _c(m) on the relevant support (with equality possible only at the crossing), which gives the stated support-side sufficient condition. For the second statement, Equation (5.27) implies that on a positive-probability set of margin fibers there is positive conditional probability of both Γ<0 <0 and Γ>0 >0. Hence both conditional positive-part expectations a(Mh)a(M_h) and b(Mh)b(M_h) from Theorem 5.3 are strictly positive on that set. Their minimum therefore has strictly positive expectation, so Equation (5.11) yields Vmeta(ℱh)−Vmeta(Mh)>0.V_ meta( F_h)-V_ meta(M_h)>0. (C.32) C.16 Proof of Proposition 5.8 Equation (5.16) gives mh′−mh=σhZm_h -m_h= _hZ. Because σh _h is ℱh F_h-measurable and Z is independent of ℱh F_h, [|mh′−mh|∣ℱh]=σh|Z|.E[|m_h -m_h| F_h]= _hE|Z|. (C.33) If [Z2]=1E[Z^2]=1, the same argument yields [(mh′−mh)2∣ℱh]=σh2[Z2]=σh2.E[(m_h -m_h)^2 F_h]= _h^2E[Z^2]= _h^2. (C.34) C.17 Proof of Proposition B.1 If F=0F=0, then π4=π8 _4= _8 and both policies promote the same already-tested path under the declared same-path technology. Their H12 terminal outcome is therefore identical, so Y=Qπ8−Qπ4=0Y=Q_ _8-Q_ _4=0. For the binary action menu, Qπj=QKeep+Dπj=Xi,Q_ _j=Q_ Keep+D1\ _j= Xi\, (C.35) for j∈4,8j∈\4,8\. Subtracting the two identities gives Equation (B.1); because the two action indicators differ exactly when F=1F=1, |Y|=|D|F|Y|=|D|F. C.18 Proof of Proposition B.2 By Proposition B.1, SY=0SY=0 outside E=S=1,F=1E=\S=1,F=1\. Hence ΔV(S) V(S) =[S(c−Y)] =E[S(c-Y)] (C.36) =cPr(S=1)−[YE] =c (S=1)-E[Y1_E] (C.37) =cρ−r[Y∣E]=cρ−rηs, =cρ-rE[Y E]=cρ-r _s, (C.38) when r>0r>0. Since Y≤Y+Y≤ Y_+ pointwise, [YE]≤[Y+E]=rη+,E[Y1_E] [Y_+1_E]=r _+, (C.39) which yields Equation (B.3). The case r=0r=0 reduces to ΔV(S)=cρ V(S)=cρ. C.19 Proof of Theorem B.3 For the nonnegative random variable Y+EY_+1_E, the layer-cake identity gives [Y+E]=∫0∞Pr(Y+E>t)t=∫0∞Pr(E,Y+>t)t.E[Y_+1_E]= _0^∞ (Y_+1_E>t)\,dt= _0^∞ (E,\,Y_+>t)\,dt. (C.40) Combining this identity with Equation (B.3) proves Equation (B.5). For the impossibility statement, fix 0<r≤ρ≤10<r≤ρ≤ 1 and M>0M>0. Choose a joint law with Pr(S=1,F=1)=r,Pr(S=1,F=0)=ρ−r, (S=1,F=1)=r, (S=1,F=0)=ρ-r, (C.41) and allocate the remaining probability to S=0,F=0S=0,F=0. Set Y=0Y=0 whenever F=0F=0 and set Y=K:=M+cρ+1rY=K:= M+cρ+1r (C.42) on E=S=1,F=1E=\S=1,F=1\. This law satisfies productive-prefix decision invariance and has finite [Y+]=rKE[Y_+]=rK. Its net value is ΔV(S)=cρ−rK=−M−1<−M. V(S)=cρ-rK=-M-1<-M. (C.43) Because M is arbitrary while (ρ,r)(ρ,r) are fixed, no finite lower bound depending only on (c,ρ,r)(c,ρ,r) can hold uniformly over unrestricted stopped-flip severities. C.20 Proof of Corollary B.4 If Y+≤B+Y_+≤ B_+ almost surely, then H+=[Y+E]≤B+Pr(E)=B+r,H_+=E[Y_+1_E]≤ B_+ (E)=B_+r, (C.44) which gives Equation (B.7). If [|D|q]≤MqE[|D|^q]≤ M_q with q>1q>1, Proposition B.1 and Hölder’s inequality yield H+ H_+ ≤[|D|E] [|D|1_E] (C.45) ≤[|D|q]1/qPr(E)1−1/q [|D|^q]^1/q (E)^1-1/q (C.46) ≤Mq1/qr1−1/q, ≤ M_q^1/qr^1-1/q, (C.47) proving Equation (B.8). Finally, Equation (B.5) and the assumed tail envelope give H+≤∫0∞u(t)t,H_+≤ _0^∞u(t)\,dt, (C.48) which proves Equation (B.10). Substituting simultaneous confidence bounds for ρ, r, or H+H_+ gives the stated finite-sample lower certificates. Appendix D Notation and decision objects Symbol Meaning ℋ, H, G Coarse legal information and a refinement VB(ℐ)V_B(I) Bayes value under information ℐI ℱh,Ch F_h,C_h Revelation filtration and Bayes commit value at depth h QaQ_a Terminal utility for action a D Binary terminal gap QXi−QKeepQ_ Xi-Q_ Keep Δrev rev Pure information-refinement value Δprom,PA prom,P_A Productive path-reuse / promotion value Rapp,RestR app,R est Approximation and finite-sample estimation regret CsC_s Acquisition cost Φc _c Complete legal short-probe frontier information map YhY_h Deeper future-learning observation τ⋆(v)τ (v) First depth at which direction v is revealed M×(f)M_×(f) Cross-replicate residual moment used in the optional conservative bridge A Fixed H4 frontier-ambiguity event in the 432-anchor structural analysis Δrestart,Δpath _ restart, _ path Common-restart policy difference and productive path-reuse value h→h′G_h→ h Oracle conditional Bayes value of purchasing a deeper information refinement Gh→h′G_h→ h Policy gain from deeper revelation, conditional on ℱh F_h Γh→h′ _h→ h Cost-adjusted conditional continuation gain Gh→h′−ch→h′G_h→ h -c_h→ h σh _h State-dependent revealability scale in the location–scale stopping model σc(m) _c(m) Critical revealability at which the priced Stop/Continue decision changes for margin m Rh(ℳ)R_h^(M) Legal model-specific empirical proxy for revealability, fixed without target outcomes S,FS,F Early-stop indicator and shallow/deep action-instability indicator ρ,rρ,r Stop probability and joint stop–flip probability in the risk–severity bridge Appendix E Statistical estimands and simultaneous inference This appendix collects the inferential rules used by the empirical claims. The unit of resampling and inference is the exact anchor, not a bank-level realization, readout, or policy action. When two conditionally independent future banks are available for one anchor, bank-specific contributions are averaged within the anchor before estimating a population mean. One-sided mean lower bounds. For exact-anchor contributions Z1,…,ZnZ_1,…,Z_n, let Z¯ Z and sZs_Z denote the sample mean and standard deviation. The reported one-sided Student-t lower bound is LCBt,1−α=Z¯−t1−α,n−1sZn.LCB_t,1-α= Z-t_1-α,n-1 s_Z n. (E.1) The percentile bootstrap resamples exact anchors with replacement. Unless otherwise stated, 50,000 draws are used and the lower bound is the empirical α quantile of the bootstrap means. These uncertainty summaries target the declared exact-anchor sampling design and panel-generating regime; they are not universal guarantees over arbitrary model families, task streams, or deployment distributions. Finite-library multiplicity. The two 336-anchor Qwen finite-library panels each test the same 22 prespecified current-information comparator components. Every component must have a positive point margin, one-sided t lower bound, and exact-anchor bootstrap lower bound. In addition, a 200,000-draw max-t procedure supplies simultaneous 95% lower bounds across all 22 components. The two panels are independent replications and are never pooled for this conjunction. Multiplicity scope. Multiplicity is controlled within each prespecified family to which a conjunction or model-selection claim is attached: the 22-component current-information library and, separately, the two-architecture Mistral adaptive family. The manuscript does not assert a single paper-wide familywise-error guarantee across heterogeneous theory-targeted estimands, which answer distinct scientific questions and are reported with their own declared inferential units and lower bounds. Equal-compute decomposition. The 432-anchor common-restart comparison and each 336-anchor productive-path panel are independent. Equation (4.6) is therefore evaluated once with productive-path panel I and once with panel I. Standard errors use the independence-aware Welch–Satterthwaite construction; the two productive-path panels are not pooled. Mistral two-bank evaluation. The fixed-depth and equal-compute Mistral estimands average banks A and B within each of 480 exact anchors before t or bootstrap inference. The adaptive target analysis likewise forms ZiAB=(Zi,A+Zi,B)/2Z_i^AB=(Z_i,A+Z_i,B)/2 before inference. Finite architecture family for Mistral adaptive control. Scalar and two-coordinate continuation architectures were both specified in the Qwen adaptive-control analysis before the Mistral target analysis. Their Mistral parameters were fit using only the separate 336-anchor development panel. To protect the target-panel claim against selecting between these two architectures, they are treated as a two-element finite family. The scalar target-panel point estimate is ΔV^scalarM=1.13877630257×10−5. V_ scalar^M=1.13877630257× 10^-5. (E.2) Its ordinary one-sided 95% t and bootstrap lower bounds are 2.37141363069×10−62.37141363069× 10^-6 and 5.41900958243×10−65.41900958243× 10^-6. Bonferroni familywise 95% bounds use one-sided level 1−0.05/2=0.9751-0.05/2=0.975 for each architecture and remain positive for the scalar rule: LCBtFWER=6.37735915928×10−7,LCBbootFWER=5.23837592968×10−6.LCB FWER_t=6.37735915928× 10^-7, FWER_ boot=5.23837592968× 10^-6. (E.3) Later exploratory revealability proxies are not used to support this target-panel result. Bounded stop–flip risk. For a fixed stop rule, the joint event S=1,F=1\S=1,F=1\ is Bernoulli. Reported risk certificates are one-sided Bernoulli–KL upper confidence bounds computed bankwise under the prespecified calibration role. These bounds concern the probability of changing the deeper action; Theorem B.3 explains why they do not by themselves bound the utility severity of those changes. Appendix F Experimental protocol details Qwen comparative panel. The Qwen comparative panel contains 432 exact anchors disjoint from development and from the two finite-library panels, with three prespecified task-stream strata of 144 anchors each and nine-position balance. The active policy, strengthened DynaMiCS-style frontier comparator, 20-update policy-visible compute contract, and ambiguity threshold τF=0.00052099142651661841 _F=0.00052099142651661841 (F.1) were fixed before terminal evaluation. Each anchor contains both Keep and Xi H12 outcomes in two conditionally independent terminal banks, T1T_1 and T2T_2. Mistral development and target panels. Mistral-specific numerical heads are fit on a separate 336-anchor development panel. The independent 480-anchor target panel uses model revision c03fc1dabc3d31b96271626f15a76a6779fb4037 of mistralai/Mistral-7B-v0.3, two independent future banks, the same Keep/Xi Keep/ Xi action menu, H4/H8/H12 geometry, and 20-update policy-visible compute contract. The H4 ambiguity threshold 0.002369209144140.00236920914414 is estimated from the development panel and fixed for target evaluation. References References [1] A. N. Angelopoulos, S. Bates, E. J. Candès, M. I. Jordan, and L. Lei (2025) Learn then test: calibrating predictive algorithms to achieve risk control. The Annals of Applied Statistics 19 (2), p. 1641–1662. External Links: Document Cited by: Appendix B, §6.3, §7. [2] A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster (2024) Conformal risk control. In International Conference on Learning Representations, Cited by: Appendix B, §6.3, §7. [3] R. R. Bahadur and L. J. Savage (1956) The nonexistence of certain statistical procedures in nonparametric problems. The Annals of Mathematical Statistics 27 (4), p. 1115–1122. External Links: Document Cited by: Appendix B. [4] J. O. Berger (1985) Statistical decision theory and bayesian analysis. 2 edition, Springer, New York. External Links: Document Cited by: §2.1, §7. [5] D. Blackwell (1953) Equivalent comparisons of experiments. The Annals of Mathematical Statistics 24 (2), p. 265–272. External Links: Document Cited by: §1, §2.1, §7. [6] K. Chaloner and I. Verdinelli (1995) Bayesian experimental design: a review. Statistical Science 10 (3), p. 273–304. External Links: Document Cited by: §1, §15.1, §15.2, §7. [7] T. Domhan, J. T. Springenberg, and F. Hutter (2015) Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves. In Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, p. 3460–3468. Cited by: §1, §15.1, §7. [8] P. L. Donti, B. Amos, and J. Z. Kolter (2017) Task-based end-to-end model learning in stochastic optimization. In Advances in Neural Information Processing Systems 30, Cited by: §7. [9] A. N. Elmachtoub and P. Grigas (2022) Smart “predict, then optimize”. Management Science 68 (1), p. 9–26. External Links: Document Cited by: §7. [10] K. J. Ezawa (1998) Evidence propagation and value of evidence on influence diagrams. Operations Research 46 (1), p. 73–83. External Links: Document Cited by: §7. [11] A. A. Feldbaum (1960) Dual control theory. i. Automation and Remote Control 21, p. 874–880. Cited by: §1, §15.2, §7. [12] V. Franc and D. Prusa (2019) On discriminative learning of prediction uncertainty. In Proceedings of the 36th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 97, p. 1963–1971. Cited by: §6.3, §7. [13] P. I. Frazier, W. B. Powell, and S. Dayanik (2008) A knowledge-gradient policy for sequential information collection. SIAM Journal on Control and Optimization 47 (5), p. 2410–2439. External Links: Document Cited by: §15.1, §5, §7. [14] E. Gualdoni, S. Laguna, L. Bethune, J. Monteiro, P. Ablin, and M. Cuturi (2026) DynaMiCS: fine-tuning LLMs with performance constraints using dynamic mixtures. arXiv preprint arXiv:2605.10770. External Links: 2605.10770, Document Cited by: §1, §10, §15.1, §7, Table 1. [15] R. A. Howard (1966) Information value theory. IEEE Transactions on Systems Science and Cybernetics 2 (1), p. 22–26. External Links: Document Cited by: §1, §2.1, §7. [16] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: 2106.09685, Document Cited by: §8. [17] M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, C. Fernando, and K. Kavukcuoglu (2017) Population based training of neural networks. arXiv preprint arXiv:1711.09846. External Links: 1711.09846, Document Cited by: §1, §15.1, §7, Table 1. [18] K. Jamieson and A. Talwalkar (2016) Non-stochastic best arm identification and hyperparameter optimization. In Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, Proceedings of Machine Learning Research, Vol. 51, p. 240–248. Cited by: §1, §7. [19] M. Jazbec, A. Timans, T. Hadži Veljković, K. Sakmann, D. Zhang, C. A. Naesseth, and E. Nalisnick (2024) Fast yet safe: early-exiting with risk control. In Advances in Neural Information Processing Systems, Vol. 37. Cited by: §6.3, §7, Table 1. [20] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. Le Scao, T. Lavril, T. Wang, T. Lacroix, and W. El Sayed (2023) Mistral 7b. arXiv preprint arXiv:2310.06825. External Links: 2310.06825, Document Cited by: §8. [21] L. P. Kaelbling, M. L. Littman, and A. R. Cassandra (1998) Planning and acting in partially observable stochastic domains. Artificial Intelligence 101 (1–2), p. 99–134. External Links: Document Cited by: §1, §15.2, §7. [22] D. P. Kingma and J. Ba (2015) Adam: a method for stochastic optimization. In International Conference on Learning Representations, External Links: 1412.6980, Document Cited by: §8. [23] A. Lebeurrier, T. Vayer, and R. Gribonval (2026) Path-conditioned training: a principled way to rescale ReLU neural networks. arXiv preprint arXiv:2602.19799. External Links: 2602.19799, Document Cited by: §7. [24] L. Li, T. J. Walsh, and M. L. Littman (2006) Towards a unified theory of state abstraction for MDPs. In Proceedings of the International Symposium on Artificial Intelligence and Mathematics, Cited by: §7. [25] L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar (2018) Hyperband: a novel bandit-based approach to hyperparameter optimization. Journal of Machine Learning Research 18 (185), p. 1–52. Cited by: §1, §15.1, §7, Table 1. [26] D. V. Lindley (1956) On a measure of the information provided by an experiment. The Annals of Mathematical Statistics 27 (4), p. 986–1005. External Links: Document Cited by: §1, §15.1, §15.2, §7. [27] M. L. Littman, R. S. Sutton, and S. Singh (2001) Predictive representations of state. In Advances in Neural Information Processing Systems 14, p. 1555–1561. Cited by: §7. [28] I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: 1711.05101, Document Cited by: §8. [29] J. Mandi, J. Kotary, S. Berden, M. Mulamba, V. Bucarey, T. Guns, and F. Fioretto (2024) Decision-focused learning: foundations, state of the art, benchmark and future opportunities. Journal of Artificial Intelligence Research 80, p. 1623–1701. External Links: Document Cited by: §7. [30] Mistral AI (2024) Mistral-7B-v0.3 model card. Note: Hugging Face model repositoryAccessed 2026-08-24 External Links: Link Cited by: §8. [31] R. B. Myerson (1979) Incentive compatibility and the bargaining problem. Econometrica 47 (1), p. 61–73. External Links: Document Cited by: §7. [32] Qwen Team (2024) Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. External Links: 2412.15115, Document Cited by: §8. [33] Qwen Team (2024) Qwen2.5-7B model card. Note: Hugging Face model repositoryAccessed 2026-08-24 External Links: Link Cited by: §8. [34] N. Raman, S. Milani, and F. Fang (2026) Selecting decision-relevant concepts in reinforcement learning. arXiv preprint arXiv:2604.04808. External Links: 2604.04808, Document Cited by: §7. [35] L. Ringel, R. Cohen, D. Freedman, M. Elad, and Y. Romano (2024) Early time classification with accumulated accuracy gap control. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, p. 42584–42600. Cited by: §6.3, §7, Table 1. [36] S. Russell and E. Wefald (1991) Principles of metareasoning. Artificial Intelligence 49 (1–3), p. 361–395. External Links: Document Cited by: §15.1, §5, §7. [37] K. Swersky, J. Snoek, and R. P. Adams (2014) Freeze–thaw bayesian optimization. arXiv preprint arXiv:1406.3896. External Links: 1406.3896, Document Cited by: §1, §15.1, §7, Table 1. [38] M. Tomaszewski (2026) Can neural networks learn by experimenting on themselves? self-interventional learning from functional consequences to predictive self-knowledge. arXiv preprint arXiv:2608.14894. External Links: 2608.14894, Document Cited by: §7. [39] T. Veiga and J. Renoux (2023) From reactive to active sensing: a survey on information gathering in decision-theoretic planning. ACM Computing Surveys 55 (13s), p. 1–22. External Links: Document Cited by: §1, §15.1, §15.2, §7. [40] Q. Wang (2026) Fiber fingerprints of hidden learning-state dynamics. arXiv preprint arXiv:2608.15976. External Links: 2608.15976, Document Cited by: §1, §3. [41] X. Wang, A. Suresh, A. Zhang, R. More, W. Jurayj, B. Van Durme, M. Farajtabar, D. Khashabi, and E. Nalisnick (2026) Conformal thinking: risk control for reasoning on a compute budget. arXiv preprint arXiv:2602.03814. External Links: 2602.03814, Document Cited by: §6.3, §7. [42] J. Wen, L. He, and Z. He (2026) Not all errors are equal: consequence-aware reasoning compute allocation. arXiv preprint arXiv:2606.04402. External Links: 2606.04402, Document Cited by: §7. [43] S. Wu, J. Xie, Y. Zhang, and Y. Xiao (2026) CODA: difficulty-aware compute allocation for adaptive reasoning. arXiv preprint arXiv:2603.08659. External Links: 2603.08659, Document Cited by: §7.