Paper deep dive
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
Alireza Furutanpey, Schahram Dustdar
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/21/2026, 4:09:31 AM
Summary
The paper introduces PRIME (Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators), a reinforcement learning framework that uses a single, stationary intrinsic reward derived from Expected Free Energy. This reward balances epistemic value (parameter information gain) to drive exploration in unresolved dynamics and an aleatoric penalty to favor low-variance transitions in resolved regions, avoiding the non-stationarity issues of previous surprise-minimizing or maximizing methods.
Entities (8)
Relation Signals (7)
PRIME → uses → Expected Free Energy
confidence 95% · We propose a single intrinsic reward... derived from the novelty contribution of a preference-free Expected Free Energy objective
PRIME → balances → Epistemic Value
confidence 92% · The intrinsic objective weights parameter information gain against expected transition entropy
PRIME → balances → Aleatoric Penalty
confidence 90% · The intrinsic objective weights parameter information gain against expected transition entropy, a penalty that replaces ambiguity
PRIME → implements → QR-DQN
confidence 88% · A statistics ensemble {Ek} holds QR-DQN distributional heads
PRIME → evaluatedon → Butterflies
confidence 85% · The experiments (Section 6) probe the predicted direction-free behavior on two didactic environments... Butterflies
PRIME → evaluatedon → MAZE
confidence 85% · The experiments (Section 6) probe the predicted direction-free behavior on two didactic environments... Maze
Surprise Minimization → criticizedby → PRIME
confidence 80% · Surprise minimization is scoped by design to “unstable” environments... We argue that information gain... is the appropriate intrinsic signal
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space.
Tags
Links
- Source: https://arxiv.org/abs/2607.16858v1
- Canonical: https://arxiv.org/abs/2607.16858v1
Trouble viewing inline? Open PDF directly →
Full Text
76,023 characters extracted from source content.
Expand or collapse full text
11institutetext: Coovally, Barcelona, Spain 22institutetext: ICREA Barcelona, Spain 33institutetext: Distributed Systems Group, TU Vienna, Austria Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators Alireza Furutanpey, Corresponding author: Schahram Dustdar Abstract Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to “unstable” environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of the next state minus its irreducible part, is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space. Maximizing it seeks exactly the surprise the model can explain away. In regions of unresolved dynamics, this epistemic term drives exploration. As dynamics become resolved, the epistemic term vanishes, while an aleatoric penalty favors lower-variance transitions, all without fitting an explicit next-state predictor. A pseudocount supplies epistemic value, a probe-based penalty captures aleatoric variance, and a short-horizon gate protects informative successors. A window-based freeze of all reward-defining objects yields a stationary Bellman operator, explicit bounds on learning targets, and a conditional uniform-concentration result for the nonparametric estimators under mixing, smoothness, bandwidth, and capacity assumptions. In active-inference terms, the agent is preference-free where novelty is retained, standard likelihood ambiguity vanishes under full observability, a nonstandard transition-entropy penalty is added, and surprise minimization emerges in resolved regions of the state space. 1 Introduction Unsupervised reinforcement learning agents should behave appropriately across environments with mixed sources of uncertainty, spanning components where the state distribution is high-entropy under any policy and components where the agent must acquire control of low-entropy deterministic transitions. Yet the dominant intrinsic-motivation families instead precommit to a single uncertainty structure by design. Surprise minimization [5] rewards negative self-information of state visitation under an episode-local density model. The construction is scoped to “unstable” environments and admits the dark-room failure mode [16] in stable ones. Information-gain and prediction-error families [6, 7, 29, 30, 35] target near-deterministic or sparse-reward settings. Prediction-error variants chase stochastic distractors (the noisy-television failure). Ensemble-disagreement variants mitigate this at the cost of a learned forward-model ensemble, with members fitting the same noise converging on a shared conditional mean, so cross-member variance decays while realized prediction error remains high. Recent adaptive approaches select the surprise direction at the architecture level. S-Adapt [20] switches between a surprise-minimizing and a surprise-maximizing reward via a per-episode UCB bandit. Mixture of Surprises [43] trains parallel surprise-min and surprise-max components and switches between them on a fixed within-episode schedule. Both reintroduce non-stationarity by construction, Bellman targets shifting independently of the environment, and the two behaviors remain distinct objectives to select between. We argue that information gain about dynamics parameters is the appropriate intrinsic signal in both high-entropy and low-entropy components of the state space, and realize it as PRIME (Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators), with no bandit, operator switching, or extrinsic reward. For a preference-free agent, reducing long-term surprise amounts to resolving posterior uncertainty over the dynamics of visited states. Parameter information gain is the agent’s expected surprise about the next state minus the irreducible part of that surprise (Section 4), so maximizing it seeks exactly the surprise the model can explain away. Wherever dynamics are unresolved, this term dominates, and the agent explores at any entropy level. Once a region is resolved, the term vanishes, and the residual surprise is irreducible. Then, the only remaining option, enforced by the weighted aleatoric term, is to prefer actions that lead to next states with lower expected transition entropy. Active inference derives information seeking from long-term surprise minimization, exploration transiently raising surprise in order to lower it later [15, 33]. PRIME derives surprise minimization from information seeking, the minimizing phase emerging from the preference-free objective once information gain is exhausted. In active-inference terms, the agent is preference-free, with uniform preference prior p(o∣C)p(o C), so the pragmatic component of expected free energy is constant in π and drops from the argmin. The intrinsic objective weights parameter information gain against expected transition entropy, a penalty that replaces ambiguity in fully observed MDPs (Section 4). Surprise minimization in the FEP sense follows from model resolution plus aleatoric avoidance, and the principle does not require the negative self-information reward from the surprise-minimization family. We state the target as a preference-free Expected Free Energy in reward-maximization form, keeping information gain about dynamics parameters minus the expected transition entropy of next states, and realize it model-free in the restricted sense that no next-state predictor is fit or rolled out, the statistics living on counts, scalar return estimates (return space), and conditional variances of next-state features. A count-based pseudocount [4] carries the epistemic term, a frozen ensemble of distributional return-statistics heads provides its theoretical target frame, and a short-horizon gate in the rectified novelty-difference form of [42] withholds the aleatoric penalty where near-term successors retain high epistemic value. Snapshotting all reward-defining objects on a fixed schedule fixes the intrinsic reward as a function of the transition tuple in each window. We consider our core contribution to be the argument, and its realization as PRIME, that one preference-free objective yields both surprise directions. A single window-stationary reward produces surprise-maximizing behavior where dynamics are unresolved and surprise-minimizing behavior where they are resolved, with γ-contraction, uniform target bounds, and conditional uniform concentration inside each window (Section 5), in the novelty lineage of active inference [15, 34]. The focus of this study is the theoretical foundation. The experiments (Section 6) probe the predicted direction-free behavior on two didactic environments, and are readily reproducible through an openly available repository 111https://github.com/rezafuru/PRIME. Detailed comparative analysis is deferred to an extended study. 2 Background and related work 2.1 Free-energy principle and EFE The free-energy principle characterizes self-organizing systems as minimizing variational free energy, an upper bound on the negative log-evidence of observations [17]. Active inference [10, 14, 37] selects policies that minimize Expected Free Energy (EFE): G(π)= G(π)\;= −q(o∣π)[logp(o∣C)]−q(o∣π)KL[q(s∣o,π)∥q(s∣π)] -\,E_q(o π)[ p(o C)]\;-\;E_q(o π)KL[q(s o,π)\,\|\,q(s π)] (1) −q(o,s∣π)KL[q(θ∣o,s,π)∥q(θ)], -\;E_q(o,s π)KL[q(θ o,s,π)\,\|\,q(θ)], the three terms being extrinsic value, salience (state information gain), and novelty (parameter information gain) [10, 34]. A non-negative expected-evidence-bound term, vanishing under the variational approximation q≈pq≈ p, is omitted. PRIME keeps the novelty term alone, the extrinsic term constant under a uniform preference prior and the salience term dropped by design, full observability leaving it at predictive state entropy, not zero. Section 4 adds a transition-entropy penalty and realizes the objective without a learned dynamics model. 2.2 Intrinsic motivation in RL Intrinsic-motivation methods in RL [2] include prediction error [6, 29], information gain over a learned forward model [7, 19, 30, 35], and density-based pseudocount bonuses [4]. Prediction-error variants suffer the noisy-television failure, treating stochastic distractors as informative. AMA [26] counters by subtracting a learned aleatoric variance, in the service of curiosity alone. Ensemble-disagreement variants mitigate the failure by rewarding cross-member predictive variance rather than realized error, requiring a forward-model ensemble. The rectified novelty-difference bracket of NovelD [42] underlies the gate in Section 4.3, and DEIR [40] uses a related pattern. PTS-BE [7] is the model-based parallel to our return-space surrogate, realizing expected parameter information gain as Jensen–Shannon disagreement across a Bayesian forward-model ensemble. 2.3 Surprise minimization and direction-adaptive methods SMIRL [5] maintains a density model pθ(s)p_θ(s) over the within-episode state buffer and uses logpθt−1(st) p_ _t-1(s_t) as the intrinsic reward, operationalizing the FEP reading that organisms minimize long-term surprise as negative instantaneous self-information of state visitation. The dark-room problem [16] makes the cost of that operationalization explicit. In the absence of preferences, minimizing the entropy of one’s own state visitation draws the agent to inactive low-entropy configurations. Friston et al. dissolve the problem through phenotypic priors, which a preference-free reward does not have. SMIRL avoids collapse by relying on environments whose state distributions evolve under non-agent dynamics and by limiting deployment to “unstable” settings. S-Adapt [20] and Mixture of Surprises [43] retain the surprise vocabulary and select the direction architecturally (Section 1). Section 4 addresses the same question by a single window-stationary reward whose dual-direction behavior emerges from the parameter information gain net of a weighted transition-entropy penalty, without arm or mixture switching. A separate line [8] achieves stationarity by augmenting the state with online-updated sufficient statistics. We instead freeze the reward-defining tuple between window boundaries, admitting the within-window γ-contraction of Section 5. 3 Setting and didactic environments We consider episodic finite-action MDPs (,,P,γ)(S,A,P,γ) without extrinsic reward and with γ<1γ<1. The agent should develop control policies that perform across environment components with different conditional entropies of next state given action, under one reward weighting and no operator-side choice of surprise direction. We use the didactic pair of [20], two 10×1010× 10 grids with 100100-step episodes matching their small variants. Butterflies (high-entropy): the agent catches randomly moving butterflies. The transition is non-deterministic in the non-agent component (butterfly motion), while the agent retains control over its own position. Surprise-minimizing methods suit this task, while prediction-error curiosity chases the random butterfly motion. Maze (low-entropy): the agent navigates to a single exit on a static grid. Surprise minimization collapses to any stationary inactive policy under negative self-information, and information-gain methods are appropriate. A single agent should succeed in both environments under the same intrinsic objective and identical hyperparameters. 4 Method In reward-maximization form, the preference-free target is the intrinsic expected free energy UIEFE(s,a)=I(θ;S′∣s,a,D)−λθH(S′∣s,a,θ).U_IEFE(s,a)\;=\;I(θ;S s,a,D)\;-\;λ\,E_θH(S s,a,θ). (2) The negative −UIEFE-U_IEFE is the EFE-form cost, the novelty contribution of Eq. (1) plus the weighted transition entropy, a correspondence of sign and retained terms, not an identity, Eq. (1) being policy-level and UIEFEU_IEFE state-action local. The first term is parameter information gain (novelty) [34], and the second is the transition-entropy penalty, a deliberate aleatoric term. Under full observability p(o∣s)=δ(o−s)p(o s)=δ(o-s) carries no ambiguity, so the only irreducible uncertainty left to penalize is that of the dynamics p(s′∣s,a,θ)p(s s,a,θ). The term is our addition, absent from the standard EFE decomposition. The two terms are linked by the mutual-information identity of BALD [18], the epistemic–aleatoric decomposition [12], I(θ;S′∣s,a,D)=H(S′∣s,a,D)−θH(S′∣s,a,θ),I(θ;S s,a,D)\;=\;H(S s,a,D)\;-\;E_θH(S s,a,θ), (3) so UIEFE(s,a)=H(S′∣s,a,D)−(1+λ)θH(S′∣s,a,θ)U_IEFE(s,a)=H(S s,a,D)-(1+λ)\,E_θH(S s,a,θ), the expected surprise of the next state minus (1+λ)(1+λ) times its irreducible part. Maximizing the information-gain term seeks exactly the surprise the model can explain away and vanishes once the posterior predictive matches the transition kernel, after which the weighted aleatoric term alone orders actions (Proposition 4). 4.1 Architecture A statistics ensemble Ekk=1K\E_k\_k=1^K holds QR-DQN [11] distributional heads, each representing the per-action return distribution Zk(s,a)Z_k(s,a) by NqN_q learned quantile locations qk,i(s,a)q_k,i(s,a) in place of a scalar mean, with scalar logit E¯k(s,a)=1Nq∑iqk,i(s,a) E_k(s,a)= 1N_q _iq_k,i(s,a). A control ensemble Qkk=1K\Q_k\_k=1^K supplies action-value functions for behavior. Statistics quantiles satisfy the hard bound |qk,i|≤Gq|q_k,i|≤ G_q, and control outputs are softly regularized toward |Qk|≤Qmax|Q_k|≤ Q_ (Appendix 0.A). Per-head affine calibration E^k=akE¯k+bk E_k=a_k E_k+b_k to a reference head ErefE_ref with ak∈[amin,amax]a_k∈[a_ ,a_ ], |bk|≤bmax|b_k|≤ b_ gives |E^k|≤GE:=amaxGq+bmax| E_k|≤ G_E:=a_ G_q+b_ , removing per-head gauge freedom that would contaminate cross-head variance. The - superscript denotes frozen targets. A per-state baseline u(⋅∣s)∈Δ()u(· s)∈ (A) centers calibrated logits, E~k−(s,a)≔E^k−(s,a)−∑b∈u(b∣s)E^k−(s,b),|E~k−|≤2GE. E_k^-(s,a)\; \; E_k^-(s,a)\;-\; _b u(b s)\, E_k^-(s,b), | E_k^-|≤ 2G_E. (4) A reference policy πref(a∣s)∝exp(E¯ref−(s,a)/τ) _ref(a s) ( E_ref^-(s,a)/τ) with τ≥τmin>0τ≥ _ >0 satisfies ∥πref(⋅∣s)−u(⋅∣s)∥1≥δπ>0\| _ref(· s)-u(· s)\|_1≥ _π>0 on a set of states of positive visitation, decoupling reward measurement from policy improvement. Statistics window. Within a window all reward-defining objects are frozen, target heads Ek−\E_k^-\, Eref−E_ref^-, calibration (ak,bk)(a_k,b_k), baseline u, πref _ref, probe directions, neighbor weights, and hyperparameters λ,α,τ,hg,β,σ0λ,α,τ,h_g,β, _0, all statistics under stop-gradient. 4.2 Epistemic estimator and pseudocount realization Under the frozen πref _ref, define the reference-policy continuation and the return-space variable gk(s)≔a∼πref(⋅∣s)[E~k−(s,a)],Yk≔γgk(S′).g_k(s)\; \;E_a _ref(· s)\! [ E_k^-(s,a) ], Y_k\; \;γ\,g_k(S ). (5) Treating the head index k as uniform on 1,…,K\1,…,K\, the variance of YkY_k under the joint distribution (k,S′)∣(s,a)(k,S ) (s,a) decomposes via the law of total variance as Vepi,LoTV(s,a)=Vark([Yk∣s,a]),Vale(s,a)=k(VarS′∣s,a[Yk]).V_epi,LoTV(s,a)\;=\;Var_k\! (E[Y_k s,a] ), V_ale(s,a)\;=\;E_k\! (Var_S s,a[Y_k] ). (6) The decomposition is algebraic, with no Bayesian posterior semantics claimed for Yk\Y_k\. Vepi,LoTVV_epi,LoTV is the cross-head disagreement on the return-space variable, the same signal used by ensemble-disagreement curiosity [7, 30, 35], evaluated on a value-side ensemble under shared πref _ref policy evaluation. The value-side ensemble, in the K-head construction of [28], exhibits documented representation-space convergence under shared TD targets [36]. Randomized priors [27] address prior-variance but do not by themselves decorrelate the shared TD targets. Cross-head variance therefore underestimates information gain, and Section 6.4 measures it near 10−710^-7. Convergence collapses only this cross-head variance, while the within-head ValeV_ale of Eq. (6) survives and supplies the heads’ aleatoric input to the penalty and gate of Section 4.3. The epistemic estimator that enters the reward is a count-based realization, Vepi,count(s,a)=κγ2tanh2(1/Ns,a+1),V_epi,count(s,a)\;=\;κ\,γ^2\, ^2\! (1/ N_s,a+1 ), (7) with Ns,aN_s,a a cumulative count under an augmented bucket key over the controllable state (Appendix 0.D). The realization is a bounded estimator of parameter information gain via the chain inequality IG≤PG≤1/N^IG ≤ 1/ N of [4], prediction gain bounding information gain and the inverse pseudocount bounding prediction gain. The shape matches this chain, decaying as 1/Ns,a1/N_s,a for large counts while staying smooth and bounded by κγ2κγ^2. Vepi,LoTVV_epi,LoTV stands as the theoretical target frame for I(θ;S′∣s,a,D)I(θ;S s,a,D), and Vepi,countV_epi,count as the deployed realization (Appendix 0.C). The architectural contribution is the window-frozen, contraction-stable composition of Section 5, and estimator tightness is not claimed. Probe-augmented aleatoric detector. With frozen feature maps ϕ,ψφ,ψ bounded by ‖ϕ(s)‖2≤Bϕ\|φ(s)\|_2≤ B_φ and ‖ψ(s)‖2≤Bψ\|ψ(s)\|_2≤ B_ψ, sample m unit-sphere probes wi∈d−1w_i ^d-1 at window start and define V~aleprobe(s,a)=1m∑i=1mVarS′∣s,a(γwi⊤ϕ(S′)), V_ale^probe(s,a)\;=\; 1m _i=1^mVar_S s,a\! (γ\,w_i φ(S ) ), (8) and analogously V~aleraw V_ale^raw on ψ. The per-probe estimator concentrates on γ2(trΣϕ)/dγ^2(tr _φ)/d under uniform sampling on d−1S^d-1. The aleatoric-augmented term is Valeaug(s,a)=maxVale,V~aleprobe,V~alerawV_ale^aug(s,a)= \V_ale, V_ale^probe, V_ale^raw\, taking the maximum of the three detectors. Short-horizon epistemic gate. Let β∈(0,1]β∈(0,1] and hg∈ℕh_g . With deterministic probe policy πprobe(s)=argmaxaπref(a∣s) _probe(s)= _a _ref(a s) frozen for snapshot stability, V~epi(hg)(s,a)≔[∑j=1hgβj−1Vepi(St+j,πprobe(St+j))|St=s,At=a], V_epi^(h_g)(s,a)\; \;E\! [ _j=1^h_gβ^j-1V_epi\! (S_t+j,\, _probe(S_t+j) )\, |\,S_t=s,A_t=a ], (9) where Vepi=max(Vepi,LoTV,Vepi,count)V_epi= (V_epi,LoTV,V_epi,count) takes the larger of the two estimates. 4.3 Intrinsic reward The intrinsic reward is, within a window, rint(s,a)=Vepi(s,a)−λlog(1+[Valeaug(s,a)−αV~epi(hg)(s,a)]+σ02).r_int(s,a)\;=\;V_epi(s,a)\;-\;λ\, \! (1+ [V_ale^aug(s,a)-α\, V_epi^(h_g)(s,a)]_+ _0^2 ). (10) The first term rewards epistemic novelty. The second term penalizes the aleatoric-augmented variance net of the short-horizon epistemic look-ahead under a rectifier. The log shape saturates with bounded slope, which the bounds and the preference inequality of Section 5 use, and σ02 _0^2 sets the variance scale below which noise is ignored. In states whose near future under πprobe _probe passes through high-VepiV_epi states the rectifier evaluates to zero and no penalty applies. In states whose near future carries no epistemic value the penalty is preserved. At λ=0λ=0 the reward degenerates to pure epistemic novelty. The bracket [Valeaug−αV~epi(hg)]+[V_ale^aug-α V_epi^(h_g)]_+ shares the rectified novelty-difference form of NovelD [42] with positions exchanged, the aleatoric estimator in the positive term and the look-ahead epistemic value as subtractor. 4.4 Learning rules Statistics heads update via QR-DQN policy evaluation under πref _ref with target ζktgt=r¯(s,a,s′)+γqk,j−(s′,aπref′) _k^tgt= r(s,a,s )+γ\,q_k,j^-(s ,a _ _ref), aπref′∼πrefa _ _ref _ref, minimizing the quantile Huber loss across the NqN_q quantile locations (distributional Bellman operator ΠW1Tπref W_1T _ref). Here r¯ r is an admissible continuation bonus, bounded and frozen for the window (Section 5), in the deployed configuration a count bonus on the augmented bucket key. Control critics learn by Double-DQN [39] with yk=rint(s,a)+γQk−(s′,argmaxa′Qk(s′,a′))y^k=r_int(s,a)+γ\,Q_k^-(s , _a Q_k(s ,a )) and squared-error loss. The gradient through rintr_int is stopped, and only the control critics update from this loss. 5 Within-window guarantees We state the main claims under a window-wise freeze of all objects in Section 4.1, assuming throughout γ∈(0,1)γ∈(0,1), finite ,S,A (or compact and measurable), the range and calibration constraints of Section 4.1, and the non-triviality margin δπ>0 _π>0 on a set of states of positive visitation. The margin is necessary because centering gives gk(s)=∑a[πref(a∣s)−u(a∣s)]E^k−(s,a)g_k(s)= _a[ _ref(a s)-u(a s)]\, E_k^-(s,a), so πref=u _ref=u would force gk≡0g_k≡ 0 and Vepi,LoTV=Vale=0V_epi,LoTV=V_ale=0, collapsing the return-space signal. Lipschitz conditions are stated where used. Propositions 3 and 4 rely on a multi-page nonparametric argument and an optimization hypothesis we do not establish, and are stated conditional on the cited assumptions. Proofs of Lemma 1 and Proposition 4 and a proof sketch of Proposition 3 are in Appendix 0.F. Appendix 0.E maps each result to its status in the deployed tabular system. The window freeze recovers a standard within-window problem for the otherwise non-stationary rintr_int (Proposition 1), the contraction that the target-shifting switches of Section 1 forgo. Lemma 1 supplies the explicit constants, and Corollary 1 turns them into a window-constant ceiling on every TD target, checked empirically in Section 2. Proposition 2 quantifies how much epistemic advantage overrides a noise difference, the noisy-television defense in inequality form. Proposition 3 states that the estimated reward the agent acts on concentrates on the defined reward, and Proposition 4 formalizes the emergence claim of Section 1, with exploration vanishing on resolved support and only the aleatoric penalty remaining. Proposition 1(Stationarity and contraction) With all reward-defining objects of Section 4.1 frozen, rint(s,a)r_int(s,a) is a fixed bounded function of (s,a)(s,a). The Bellman optimality operator BrintB^r_int acting on bounded action-value functions Q:×→ℝQ:S×A by (BrintQ)(s,a)=rint(s,a)+γs′[maxa′Q(s′,a′)](B^r_intQ)(s,a)=r_int(s,a)+γ\,E_s [ _a Q(s ,a )] is a γ-contraction on (ℓ∞(×),∥⋅∥∞)( _∞(S×A),\|·\|_∞), with a unique fixed point. Proof Freezing makes rintr_int deterministic on (s,a)(s,a) with ‖rint‖∞<∞\|r_int\|_∞<∞ by Lemma 1. The rintr_int term cancels in the difference, so the sup-of-max bound gives ‖BrintQ−BrintQ′‖∞≤γ‖Q−Q′‖∞\|B^r_intQ-B^r_intQ \|_∞≤γ\|Q-Q \|_∞. Banach’s fixed-point theorem completes. Lemma 1(Uniform bounds) Under the calibration constraints, |E~k−|≤2GE| E_k^-|≤ 2G_E, |gk|≤2GE|g_k|≤ 2G_E, and |Yk|≤2γGE|Y_k|≤ 2γ G_E. Hence, with Bale≔max4γ2GE2,γ2Bϕ2,γ2Bψ2B_ale \4γ^2G_E^2,γ^2B_φ^2,γ^2B_ψ^2\, 0≤Vepi,LoTV+Vale≤4γ2GE2,Valeaug≤Bale, 0≤ V_epi,LoTV+V_ale≤ 4γ^2G_E^2, V_ale^aug≤ B_ale, −λlog(1+Bale/σ02)≤rint(s,a)≤max4γ2GE2,κγ2. -λ \! (1+B_ale/ _0^2 )\;≤\;r_int(s,a)\;≤\; \4γ^2G_E^2,\,κγ^2\. (11) Corollary 1(Bounded learning targets) If additionally |Qk−|≤Qmax|Q_k^-|≤ Q_ and ‖r¯‖∞≤R0\| r\|_∞≤ R_0, then |ζktgt|≤R0+γGq| _k^tgt|≤ R_0+γ G_q and |yk|≤Rr+γQmax|y^k|≤ R_r+γ Q_ with Rr≔maxmax4γ2GE2,κγ2,λlog(1+Bale/σ02)R_r \ \4γ^2G_E^2,κγ^2\,\,λ (1+B_ale/ _0^2)\. Proof |ζktgt|≤‖r¯‖∞+γ|qk,j−|≤R0+γGq| _k^tgt|≤\| r\|_∞+γ\,|q_k,j^-|≤ R_0+γ G_q. Lemma 1 gives ‖rint‖∞≤Rr\|r_int\|_∞≤ R_r, so |yk|≤‖rint‖∞+γQmax≤Rr+γQmax|y^k|≤\|r_int\|_∞+γ Q_ ≤ R_r+γ Q_ by the triangle inequality. Admissibility class. Call a continuation bonus r¯ r admissible if ‖r¯‖∞≤R0\| r\|_∞≤ R_0 and r¯ r is frozen for the window duration. Propositions 1 and 2–4 and Corollary 1 hold for every admissible instance with the same constants. Instances include the head-empirical decomposition of Eq. (6), the count-based realization of Eq. (7) [4], a snapshot-disciplined RND bonus with the predictor updated only at window boundaries, and a forward-model side ensemble [7, 35] with frozen per-(s,a)(s,a) regression targets. Lidayan et al. [23] prove policy-invariance and finite-horizon regret theorems for the narrower class of BAMDP potential-based shaping functions, whereas ours admits any frozen bounded r¯ r, with within-window operator-level guarantees in place of policy-level ones. Proposition 2(Quantitative gate-mediated preference) For two actions a1,a2a_1,a_2 at state s, let Δepi≔Vepi(s,a1)−Vepi(s,a2) _epi V_epi(s,a_1)-V_epi(s,a_2) and xi≔[Valeaug(s,ai)−αV~epi(hg)(s,ai)]+/σ02≥0x_i [V_ale^aug(s,a_i)-α V_epi^(h_g)(s,a_i)]_+/ _0^2≥ 0. WLOG x1≥x2x_1≥ x_2. A sufficient condition for rint(s,a1)≥rint(s,a2)r_int(s,a_1)≥ r_int(s,a_2) is Δepi≥λx1−x21+minx1,x2. _epi\;≥\;λ\, x_1-x_21+ \x_1,x_2\. (12) Proof rint(s,a1)≥rint(s,a2)r_int(s,a_1)≥ r_int(s,a_2) iff Δepi≥λ[log(1+x1)−log(1+x2)] _epi≥λ[ (1+x_1)- (1+x_2)]. The mean-value theorem on log(1+⋅) (1+·) gives log(1+x1)−log(1+x2)≤(x1−x2)/(1+minx1,x2) (1+x_1)- (1+x_2)≤(x_1-x_2)/(1+ \x_1,x_2\) since 1/(1+x)1/(1+x) is decreasing. Proposition 3(Uniform-in-window concentration, conditional) Assume (i) gkg_k and the conditional-variance function s′↦σk2(s′)s _k^2(s ) are L-Lipschitz in s′s (Lipschitzness of gkg_k follows from Lipschitz statistics heads and τ≥τmin>0τ≥ _ >0, Section 4.1; together with |Yk|≤2γGE|Y_k|≤ 2γ G_E from Lemma 1 this controls the variance estimator); (i) kernel neighborhoods with bandwidth h→0h→ 0 and effective neighbor count Mhd→∞Mh^d→∞, with M the number of frozen snapshots in the window and d the embedding dimension; (i) the snapshot stream is β-mixing with summable coefficients; (iv) capacity control logNbuck=o(Mhd) N_buck=o(Mh^d), where NbuckN_buck is the covering number of ×S×A at scale h. With μk(s,a)≔[Yk∣s,a] _k(s,a) [Y_k s,a] and σk2(s,a)≔VarS′∣s,a[Yk] _k^2(s,a) _S s,a[Y_k] the per-bucket conditional mean and variance of YkY_k, and μ^k,σ^k2 μ_k, σ_k^2 their kernel estimators, uniformly over realized buckets, sup(s,a)|μ^k−μk|,sup(s,a)|σ^k2−σk2|=Op(logNbuckMhd+h), _(s,a)| μ_k- _k|,\; _(s,a)| σ_k^2- _k^2|\;=\;O_p\! ( N_buckMh^d+h ), (13) and consequently rintr_int concentrates uniformly within the window. Proposition 4(Asymptotic neutrality, conditional) Suppose realizability and coverage hold and each statistics head converges (in the hypothesis class) to the unique fixed point of the projected distributional operator ΠW1Tπref W_1T _ref [3, 11]. Then Vepi,LoTV(s,a)→0V_epi,LoTV(s,a)→ 0 on the visited support, Vepi,count→0V_epi,count→ 0 as Ns,a→∞N_s,a→∞, and rint(s,a)→−λlog(1+Valeaug(s,a)/σ02).r_int(s,a)\;→\;-λ\, \! (1+V_ale^aug(s,a)/ _0^2 ). (14) Cost of freezing. Across windows rintr_int is piecewise stationary, so refresh events may shift it even when the environment is unchanged. Within a window the drift of the acting policy from πref _ref biases gkg_k by at most O(γGE‖πact−πref‖1)O(γ G_E\| _act- _ref\|_1). Window length controls a bias–variance trade-off. 6 Evaluation 6.1 Setup The construction of Section 4 is evaluated on the two didactic environments of Section 3. The PRIME configuration is fixed at K=5K=5 statistics heads, γ=0.9γ=0.9, gate weight α=0.5α=0.5, aleatoric weight λ=0.5λ=0.5, and per-seed budgets of 250,000250,000 env-steps on Butterflies and 200,000200,000 on Maze. All reward-defining hyperparameters and clip ceilings are identical across environments. Five baselines run at matched budget under their published defaults, held fixed across both environments, an analytical random walker, SMIRL [5], RND [6], Disagreement [30], and an Extrinsic-DQN reference upper bound, with n=15n=15 seeds per setup. Full architectures and hyperparameters are in Appendix 0.A, statistical comparisons in Appendix 0.B. 6.2 Direction-free behavior Figure 1: Per-method learning curves, rolling-2020 means of task reward, shaded ±1± 1 std (n=15n=15 seeds). POC is the position-only comparator (Section 6.2), Ext-DQN‡ the supervised reference trained on the task reward, the Maze panel cropped at 7575k steps. Table 1: Point estimates ±1± 1 std (n=15n=15 seeds, setup of Section 6.1). Method Butterflies catch/ep Butterflies peak Maze peak Random walker 2.29±1.212.29± 1.21 — 0.000.00 SMIRL 2.71±0.212.71± 0.21 3.94±0.383.94± 0.38 0.01±0.020.01± 0.02 Disagreement 1.72±0.111.72± 0.11 2.99±0.132.99± 0.13 1.00±0.001.00± 0.00 RND 3.26±0.143.26± 0.14 4.32±0.214.32± 0.21 1.00±0.001.00± 0.00 Ext-DQN‡ 4.11±0.674.11± 0.67 5.66±0.325.66± 0.32 0.93±0.260.93± 0.26 PRIME 3.47±0.103.47± 0.10 5.01±0.145.01± 0.14 0.79±0.110.79± 0.11 Figure 2: Ceiling check and Butterflies behavior for PRIME (n=15n=15 seeds, rolling-2020 means, shaded ±1± 1 std). (a) Empirical |qtaken|max|q_taken|_ against the Corollary 1 ceiling Rr+γQmaxR_r+γ Q_ . (b) Agent–butterfly L1 min-distance. (c) Per-episode coverage fraction. Figure 1 reports per-method learning curves, and Table 1 the aggregates. On Butterflies, PRIME averages 3.47±0.103.47± 0.10 catches per episode, 1.42×1.42× the position-only comparator on the same run-average metric (2.452.45, simulated for policies blind to butterfly positions), with peak rolling-2020 5.01±0.145.01± 0.14. On Maze, peak rolling-2020 reach is 0.79±0.110.79± 0.11 with per-seed peaks 0.650.65–1.001.00. The single-direction baselines each fail their opposite environment. SMIRL collapses on Maze (0.01±0.020.01± 0.02) by the dark-room mechanism of Section 2, and Disagreement degrades on Butterflies below the random walker (1.721.72 vs 2.292.29). RND attains both under its published defaults (Table 1, Section 7). PRIME attains both from one window-stationary objective under a single configuration shared across the two environments, the two directions following from the parameter-information-gain target of Section 4. Late-training reach-rate collapse on Maze is shared by every intrinsic method (PRIME last rolling-2020 0.04±0.070.04± 0.07), consistent with Proposition 4 on covered support, and the Extrinsic-DQN reference sustains at 1.001.00 on 14/1514/15 seeds. Corollary 1 bounds the control-critic targets by the window-constant Rr+γQmaxR_r+γ Q_ , with RrR_r at the deployed reward clip 2.02.0 (Table 3). Figure 2a plots the empirical |qtaken|max|q_taken|_ trajectory against this ceiling. It tracks the soft anchor Qmax=20Q_ =20 (one Butterflies seed’s rolling-2020 at 20.0620.06), and the bounded targets stay below the ceiling on every seed (tail-20%20\% |yk|∈[6.48,7.00]|y^k|∈[6.48,7.00] on Butterflies, [0.25,0.38][0.25,0.38] on Maze). 6.3 Per-environment behavioral evidence Butterflies. PRIME operates as a novelty-seeker in the (s,a,nalive)(s,a,n_alive) augmented bucket space. Figure 2b,c plots the agent–butterfly L1 min-distance and the per-episode position-space coverage fraction. Butterfly depletion drives the dominant novelty gradient, and a coverage contraction emerges as a side effect (panel c), the fraction falling by 0.063±0.0120.063± 0.012 between the first and last training quintiles (negative on 15/1515/15 seeds). Min-distance saturates at 4.05±0.094.05± 0.09 cells over the final 20%20\% of training (panel b), the steady-state of close pursuit on stochastically-retreating targets. The run-average catch rate above the position-only comparator (Section 6.2) is consistent with positional information entering the policy. A simulated butterfly end-of-episode density approaches uniform within the 100100-step horizon, so the agent’s visit concentration does not reflect butterfly locations. Maze. PRIME develops a near-deterministic state-conditional policy. Per-bucket argmaxQ Q entropy reaches 0.0015±0.00080.0015± 0.0008 nats at tail-20%20\% across seeds (uniform ceiling log5≈1.609 5≈ 1.609), and the mutual information between visited state and executed action is positive on every seed (0.563±0.0250.563± 0.025 nats, bias-corrected estimator). A random walker has zero state–action mutual information by construction, and the policy is state-conditional during the peak phase. 6.4 Ablations Table 2: Ablations of the full configuration. K=1K=1 sets Vepi,LoTV≡0V_epi,LoTV≡ 0, and α=0α=0 disables the subtractor. Configuration Butterflies catch/ep Butterflies peak Maze peak Full (K=5K=5, α=0.5α=0.5) 3.47±0.103.47± 0.10 5.01±0.145.01± 0.14 0.79±0.110.79± 0.11 K=1K=1 3.29±0.173.29± 0.17 4.78±0.324.78± 0.32 0.70±0.110.70± 0.11 α=0α=0 (gate off) 3.34±0.153.34± 0.15 4.93±0.184.93± 0.18 0.73±0.100.73± 0.10 Table 2 reports the ablations. The full configuration leads both ablations on all six cells, the seed-paired CI of the difference excluding zero on five (Appendix 0.B). The K=5K=5 gap over K=1K=1 is d=1.2d=1.2 on Butterflies catches, d=1.0d=1.0 on Butterflies peak, and d=0.9d=0.9 on Maze peak, and removing the gate costs d=0.95d=0.95 on Butterflies catches against d=0.6d=0.6 on Maze peak. ValeaugV_ale^aug tail-20%20\% is 2.29×10−32.29× 10^-3 on Butterflies, the gate clamping 70%70\% of batches, against 2.03×10−52.03× 10^-5 on Maze with 94%94\% clamped, and the cross-environment direction holds gate-off. Vepi,LoTV∼10−7/10−8V_epi,LoTV 10^-7/10^-8 confirms VepiV_epi reduces to the count realization. 7 Discussion The construction realizes the parameter-epistemic component of EFE without the pragmatic component, with VepiV_epi a surrogate for novelty [34] and the rectified bracket of Eq. (10) proxying expected transition entropy. In resolved high-entropy components, behavior is action-level aleatoric avoidance conditional on VepiV_epi (Proposition 4). The dark-room problem, dissolved in [16] through phenotypic priors, does not arise without them, the agent having no incentive toward low-entropy occupancy, only smaller irreducible variability at comparable epistemic content. RND succeeds on both environments under its published defaults (Table 1), expected on near-homoskedastic environments, where prediction error aligns with expected information gain [7]. As published, RND updates its predictor online, violating the Section 5 admissibility freeze. A snapshot-disciplined variant updated only at window boundaries is admissible, inherits that section’s constants, and is the follow-up estimator. A mechanism-level comparison requires matched bucket keys, coverage formulas, and tuning audits, deferred to a dedicated study with a head-to-head against the switching baselines (Section 2). The non-triviality premise δπ>0 _π>0 holds empirically, the tail-20%20\% mean of ‖πref−u‖1\| _ref-u\|_1 at 0.0297±0.00070.0297± 0.0007 on Butterflies and 0.0221±0.00020.0221± 0.0002 on Maze, positive on every seed. The tabular count realization carries the seek phase only, the aleatoric penalty remaining on resolved support (Proposition 4). Scaling to high-dimensional or continuous states needs a density-model pseudocount or representation-space disagreement signal, and the per-probe signal dilutes as d grows. Task-reward integration. PRIME trains without task reward and enters task-driven training as an exploration bonus or as reward-free pretraining adapted downstream [4, 6, 13, 22]. The unsupervised-RL benchmark convention scores pretraining by efficiency of a short task-reward adaptation [22], reading the Maze trajectory of Section 6.2 at its peak. The peak checkpoint is the transferable object (reach 0.79±0.110.79± 0.11), the decay the predicted preference-free behavior (Proposition 4), and vanishing epistemic reward signals a resolved region. As an additive bonus rext+ηrintr_ext+η\,r_int, the log-penalty is bounded (Lemma 1), ordering actions of equal epistemic content without overriding a task reward above that bound (Proposition 2), and η can anneal as the epistemic signal vanishes without distractor-seeking residue (Eq. (3)). In reward-sparse settings the epistemic term densifies the signal until task reward takes over. Aleatoric avoidance minimizes outcome variance, aligning with risk-averse control, conflicting with risk-neutral return maximization on high-variance task states. A bonus-weight scheduler [32] is the concrete next step. References [1] R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare (2021) Deep reinforcement learning at the edge of the statistical precipice. In Advances in Neural Information Processing Systems, Vol. 34. Cited by: Appendix 0.B. [2] A. Aubret, L. Matignon, and S. Hassas (2023) An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey. Entropy 25 (2), p. 327. Cited by: §2.2. [3] M. G. Bellemare, W. Dabney, and R. Munos (2017-06–11 Aug) A distributional perspective on reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning, D. Precup and Y. W. Teh (Eds.), Proceedings of Machine Learning Research, Vol. 70, p. 449–458. External Links: Link Cited by: Appendix 0.E, Proposition 4. [4] M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos (2016) Unifying count-based exploration and intrinsic motivation. In Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29, p. . External Links: Link Cited by: Appendix 0.C, §1, §2.2, §4.2, §5, §7. [5] G. Berseth, D. Geng, C. Devin, N. Rhinehart, C. Finn, D. Jayaraman, and S. Levine (2021) SMiRL: surprise minimizing reinforcement learning in unstable environments. In International Conference on Learning Representations (ICLR), Cited by: §1, §2.3, §6.1. [6] Y. Burda, H. Edwards, A. Storkey, and O. Klimov (2019) Exploration by random network distillation. In International Conference on Learning Representations, External Links: Link Cited by: Appendix 0.A, §1, §2.2, §6.1, §7. [7] A. Caron, V. Mavroudis, and C. Hicks (2025) On efficient Bayesian exploration in model-based reinforcement learning. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §1, §2.2, §4.2, §5, §7. [8] R. C. Castanyer, J. Romoff, and G. Berseth (2024) Improving intrinsic exploration by creating stationary objectives. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §2.3. [9] J. Choi, Y. Guo, M. Moczulski, J. Oh, N. Wu, M. Norouzi, and H. Lee (2019) Contingency-aware exploration in reinforcement learning. In International Conference on Learning Representations (ICLR), Cited by: Appendix 0.D. [10] L. Da Costa, T. Parr, N. Sajid, S. Veselic, V. Neacsu, and K. Friston (2020) Active inference on discrete state-spaces: a synthesis. Journal of Mathematical Psychology 99, p. 102447. External Links: ISSN 0022-2496, Document, Link Cited by: §2.1, §2.1. [11] W. Dabney, M. Rowland, M. Bellemare, and R. Munos (2018) Distributional reinforcement learning with quantile regression. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Cited by: Appendix 0.E, §4.1, Proposition 4. [12] S. Depeweg, J. Hernandez-Lobato, F. Doshi-Velez, and S. Udluft (2018-10–15 Jul) Decomposition of uncertainty in Bayesian deep learning for efficient and risk-sensitive learning. In Proceedings of the 35th International Conference on Machine Learning, J. Dy and A. Krause (Eds.), Proceedings of Machine Learning Research, Vol. 80, p. 1184–1193. External Links: Link Cited by: §4. [13] B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine (2019) Diversity is all you need: learning skills without a reward function. In International Conference on Learning Representations (ICLR), Cited by: §7. [14] K. Friston, T. FitzGerald, F. Rigoli, P. Schwartenbeck, and G. Pezzulo (2017) Active inference: a process theory. Neural computation 29 (1), p. 1–49. Cited by: §2.1. [15] K. Friston, F. Rigoli, D. Ognibene, C. Mathys, T. Fitzgerald, and G. Pezzulo (2015) Active inference and epistemic value. Cognitive Neuroscience 6 (4), p. 187–214. Note: PMID: 25689102 External Links: Document Cited by: §1, §1. [16] K. Friston, C. Thornton, and A. Clark (2012) Free-energy minimization and the dark-room problem. Frontiers in Psychology 3. External Links: Link, Document, ISSN 1664-1078 Cited by: §1, §2.3, §7. [17] K. Friston (2010) The free-energy principle: a unified brain theory?. Nature reviews neuroscience 11 (2), p. 127–138. Cited by: §2.1. [18] N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel (2011) Bayesian active learning for classification and preference learning. arXiv preprint arXiv:1112.5745. External Links: Link Cited by: §4. [19] R. Houthooft, X. Chen, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel (2016) VIME: variational information maximizing exploration. In Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29, p. . External Links: Link Cited by: §2.2. [20] A. Hugessen, R. C. Castanyer, F. Mohamed, and G. Berseth (2024) Surprise-adaptive intrinsic motivation for unsupervised reinforcement learning. In Reinforcement Learning Conference (RLC), Cited by: §1, §2.3, §3. [21] A. Jacot, F. Gabriel, and C. Hongler (2018) Neural tangent kernel: convergence and generalization in neural networks. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: Appendix 0.E. [22] M. Laskin, D. Yarats, H. Liu, K. Lee, A. Zhan, K. Lu, C. Cang, L. Pinto, and P. Abbeel (2021) URLB: unsupervised reinforcement learning benchmark. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, Cited by: §7. [23] A. Lidayan, M. D. Dennis, and S. Russell (2025) BAMDP shaping: a unified theoretical framework for intrinsic motivation and reward shaping. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §5. [24] C. Lyle, M. Rowland, and W. Dabney (2022) Understanding and preventing capacity loss in reinforcement learning. In International Conference on Learning Representations (ICLR), Cited by: Appendix 0.E. [25] J. Martin, S. N. Sasikumar, T. Everitt, and M. Hutter (2017) Count-based exploration in feature space for reinforcement learning. In International Joint Conference on Artificial Intelligence (IJCAI), Cited by: Appendix 0.D. [26] A. N. Mavor-Parker, K. A. Young, C. Barry, and L. D. Griffin (2022) How to stay curious while avoiding noisy TVs using aleatoric uncertainty estimation. In International Conference on Machine Learning (ICML), Cited by: §2.2. [27] I. Osband, J. Aslanides, and A. Cassirer (2018) Randomized prior functions for deep reinforcement learning. In Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (Eds.), Vol. 31, p. . External Links: Link Cited by: Appendix 0.C, §4.2. [28] I. Osband, C. Blundell, A. Pritzel, and B. Van Roy (2016) Deep exploration via bootstrapped DQN. In Advances in Neural Information Processing Systems, D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol. 29, p. . External Links: Link Cited by: §4.2. [29] D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell (2017-06–11 Aug) Curiosity-driven exploration by self-supervised prediction. In Proceedings of the 34th International Conference on Machine Learning, D. Precup and Y. W. Teh (Eds.), Proceedings of Machine Learning Research, Vol. 70, p. 2778–2787. External Links: Link Cited by: Appendix 0.D, §1, §2.2. [30] D. Pathak, D. Gandhi, and A. Gupta (2019-09–15 Jun) Self-supervised exploration via disagreement. In Proceedings of the 36th International Conference on Machine Learning, K. Chaudhuri and R. Salakhutdinov (Eds.), Proceedings of Machine Learning Research, Vol. 97, p. 5062–5071. External Links: Link Cited by: §1, §2.2, §4.2, §6.1. [31] R. Raileanu and T. Rocktäschel (2020) RIDE: rewarding impact-driven exploration for procedurally-generated environments. In International Conference on Learning Representations (ICLR), Cited by: Appendix 0.D. [32] M. Riedmiller, R. Hafner, T. Lampe, M. Neunert, J. Degrave, T. van de Wiele, V. Mnih, N. Heess, and J. T. Springenberg (2018) Learning by playing – solving sparse reward tasks from scratch. In International Conference on Machine Learning (ICML), Cited by: §7. [33] P. Schwartenbeck, T. FitzGerald, R. Dolan, and K. Friston (2013) Exploration, novelty, surprise, and free energy minimization. Frontiers in Psychology 4. External Links: Link, Document, ISSN 1664-1078 Cited by: §1. [34] P. Schwartenbeck, J. Passecker, T. U. Hauser, T. H. FitzGerald, M. Kronbichler, and K. J. Friston (2019) Computational mechanisms of curiosity and goal-directed exploration. eLife 8, p. e41703. External Links: Document, ISSN 2050-084X Cited by: §1, §2.1, §4, §7. [35] R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak (2020-13–18 Jul) Planning to explore via self-supervised world models. In Proceedings of the 37th International Conference on Machine Learning, H. D. I and A. Singh (Eds.), Proceedings of Machine Learning Research, Vol. 119, p. 8583–8592. External Links: Link Cited by: §1, §2.2, §4.2, §5. [36] H. Sheikh, M. Phielipp, and L. Boloni (2022) Maximizing ensemble diversity in deep reinforcement learning. In International Conference on Learning Representations, External Links: Link Cited by: Appendix 0.C, §4.2. [37] R. Smith, K. J. Friston, and C. J. Whyte (2022) A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology 107, p. 102632. External Links: ISSN 0022-2496, Document, Link Cited by: §2.1. [38] H. van Hasselt, Y. Doron, F. Strub, M. Hessel, N. Sonnerat, and J. Modayil (2018) Deep reinforcement learning and the deadly triad. arXiv preprint arXiv:1812.02648. Cited by: Appendix 0.E. [39] H. van Hasselt, A. Guez, and D. Silver (2016) Deep reinforcement learning with double Q-learning. In Proceedings of the AAAI conference on artificial intelligence, Vol. 30. Cited by: §4.4. [40] S. Wan, Y. Tang, Y. Tian, and T. Kaneko (2023) DEIR: efficient and robust exploration through discriminative-model-based episodic intrinsic rewards. In International Joint Conference on Artificial Intelligence (IJCAI), Cited by: §2.2. [41] B. Yu (1994) Rates of convergence for empirical processes of stationary mixing sequences. The Annals of Probability 22 (1), p. 94–116. Cited by: Appendix 0.F. [42] T. Zhang, H. Xu, X. Wang, Y. Wu, K. Keutzer, J. Gonzalez, and Y. Tian (2021) NovelD: a simple yet effective exploration criterion. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. W. Vaughan (Eds.), Vol. 34, p. 25217–25230. External Links: Link Cited by: §1, §2.2, §4.3. [43] A. Zhao, M. Lin, Y. Li, Y. Liu, and G. Huang (2022) A mixture of surprises for unsupervised reinforcement learning. Advances in Neural Information Processing Systems 35, p. 26078–26090. Cited by: §1, §2.3. Appendix 0.A Experimental details Environments. Both are 10×1010× 10 grids with action set N,S,E,W,Stay\N,S,E,W,Stay\, 100100-step episodes, fixed-horizon truncation (no termination), and observation tensors in 0,1C×10×10\0,1\^C× 10× 10. Butterflies (C=4C=4 channels: agent, butterfly, wall, caught-flash) places 66 butterflies at uniform-random empty cells. After each agent move every alive butterfly samples a move uniformly from the five actions with wall/bound blocking, and a butterfly sharing the agent cell is caught and removed (no respawn). The per-step butterfly motion is the aleatoric component, and depletion within the 100100-step horizon removes it. Maze (C=3C=3 channels: agent, wall, flag) is a fixed S-corridor with a consumable goal flag at (9,9)(9,9) and deterministic transitions. Both report the task signal (catches, goal reach) outside the observation and emit zero per-step reward to the intrinsic agent. Per-seed budget is 250,000250,000 env-steps on Butterflies and 200,000200,000 on Maze. Shared trunk. Every network uses the same two-layer convolutional trunk, two 3×33× 3 convolutions taking C→16→32C→ 16→ 32 channels with padding 11 and a ReLU after each, flattened to 32003200 and linearly mapped to a 6464-dim embedding. PRIME networks. The statistics ensemble holds K=5K=5 QR-DQN heads, each a trunk plus a quantile output layer Linear(64,||Nq)Linear(64,|A|N_q) with Nq=11N_q=11 learned quantile locations passed through tanh⋅Gq · G_q (|qk,i|≤Gq=2|q_k,i|≤ G_q=2), wrapped with a frozen randomized prior of strength βprior=2 _prior=2 and reseeded per head, with targets hard-synced at window boundaries. The control ensemble holds K=5K=5 unbounded Q-heads Linear(64,||)Linear(64,|A|) trained by Double-DQN, with the range |Qk|≤Qmax=20|Q_k|≤ Q_ =20 enforced softly by a penalty 0.1⋅ReLU(|Q|−Qmax)20.1·ReLU(|Q|-Q_ )^2. The aleatoric features are an ICM inverse-dynamics encoder ϕφ (trunk to Linear(3200,32)tanhLinear(3200,32)\, , dim32 32, with inverse head Linear(64,64)ReLULinear(64,||)Linear(64,64)\,ReLU\,Linear(64,|A|)) and a fixed ℓ2 _2-normalized flattened observation ψ (dimC⋅100 C·100). The count realization keys on (agent row,agent col,a,nalive)(agent row,agent col,a,n_alive). Hyperparameters. Table 3 lists the PRIME configuration, byte-identical across the two environments except for the environment handle. Table 4 lists the baseline configuration, shared across all four DQN baselines and both environments. Table 3: PRIME hyperparameters (identical across both environments). K 55 discount γ 0.90.9 quantile locations NqN_q 1111 aleatoric weight λ 0.50.5 quantile bound GqG_q 2.02.0 gate weight α 0.50.5 control bound QmaxQ_ 20.020.0 log scale σ02 _0^2 0.50.5 calibration amin/amax/bmaxa_ /a_ /b_ 1.0/2.0/1.01.0/2.0/1.0 gate decay β 0.950.95 randomized-prior βprior _prior 2.02.0 gate horizon hgh_g 44 count scale κ 0.50.5 temperature τ 1.01.0 probes m 88 feature dim ϕφ 3232 window length 10001000 warmup window 10001000 snapshot buffer 50005000 replay buffer 100000100000 calibration samples 256256 neighbors 1616 batch size 6464 target sync 10001000 update period (main/stats/inv) 4/4/44/4/4 εstart/εend _start/ _end 1.0/0.011.0/0.01 lrlr control/stats/ϕφ 3⋅10−4/10−4/10−33·10^-4/10^-4/10^-3 reward scale/clip 50/2.050/2.0 Table 4: Baseline hyperparameters (shared across all DQN baselines and both environments). discount γ 0.990.99 batch size 6464 learning rate 3⋅10−43·10^-4 target sync 10001000 replay buffer 100000100000 learning starts 10001000 εend _end 0.050.05 exploration fraction 0.250.25 gradient clip 10.010.0 optimizer Adam All networks use Adam at its default moments (0.9,0.999)(0.9,0.999) and ε=10−8 =10^-8. Disagreement uses a K=5K=5 forward-model ensemble with prediction clip 5.05.0; RND uses a frozen random target and a predictor trained on the same replay minibatches as its DQN update [6], with reward clip 5.05.0; SMIRL augments the observation with a per-cell density map and a time plane. The three intrinsic baselines normalize the intrinsic reward by a running standard deviation. Each configuration runs n=15n=15 seeds, seeds 55–1414 collected after seeds 0–44 under the identical configuration and code. The exploration schedule is linear over the first 25%25\% of training. Control critics learn from yk=rint(s,a)+γQk−(s′,argmaxa′Qk(s′,a′))y^k=r_int(s,a)+γ Q_k^-(s , _a Q_k(s ,a )) with the gradient through rintr_int stopped. Metrics. Catches per episode is the per-seed mean of the episode catch count over all episodes of a run, reported as mean ± std across seeds. Peak is the per-seed maximum over training of the rolling mean across 2020 consecutive episodes, on catches for Butterflies and on goal reach for Maze, and last rolling-2020 is the final value of the same window. An episode counts as a reach when the goal flag is collected at least once. Tail-20%20\% denotes the mean of a logged scalar over the final 20%20\% of env-steps. The coverage fraction is the per-episode fraction of the 100100 cells visited, and the pursuit distance is the per-episode mean of the per-step L1 distance to the nearest alive butterfly. The random-walker reference is simulated for 10410^4 episodes on the deployed environment, and the position-only comparator (2.452.45 catches/ep) is the highest catch rate among the simulated policies blind to butterfly positions, attained by a go-to-center-and-stay policy. The state–action mutual information is estimated every 10001000 steps from a ring buffer of executed actions keyed by the source-state components of the count bucket (agent position and naliven_alive), with Miller–Madow correction on the marginal and every per-bucket conditional entropy, buckets under five samples dropped, and the estimate clamped at zero. The per-bucket argmaxQ Q entropy is the bucket-size-weighted mean, over snapshot buckets with at least four samples, of the entropy of the argmaxQ Q action across the bucket’s snapshot states. Hyperparameter provenance. The baselines run at their published settings, held fixed across both environments (Table 4). None was tuned per environment or per task, and no baseline sweep was performed. PRIME’s single configuration (Table 3) was selected once for both environments jointly and frozen. The only configured differences between PRIME and the baselines are the discount (γ=0.9γ=0.9 vs. 0.990.99) and the exploration floor (εend=0.01 _end=0.01 vs. 0.050.05), alongside the intrinsic-reward components absent from the baselines. The comparison therefore grants the baselines their published settings at matched budget while constraining PRIME to one configuration, and gives PRIME no per-environment tuning advantage. Appendix 0.B Statistical analysis Each configuration runs n=15n=15 seeds. Interval estimation is primary [1], namely the per-method IQM with a stratified-bootstrap 95%95\% CI (Table 6), the bootstrap 95%95\% CI of the seed-mean difference, Cohen’s d, and the probability of improvement P(A>B)P(A>B) (all run pairs, ties half-weighted), with rank tests confirmatory. Seeds index matched blocks, the seed fixing the server assignment shared by every method, so the confirmatory test is the exact two-sided Wilcoxon signed-rank on per-seed differences, with the exact Mann–Whitney U alongside, each Holm-corrected across the eleven-comparison family. At n=15n=15 the exact signed-rank floor is 2/215≈6.1×10−52/2^15≈ 6.1× 10^-5, against a best rank-test floor of 2/(105)=0.00792/ 105=0.0079 (exact Mann–Whitney) at the previous n=5n=5, where no Holm-corrected rank test could reach α=0.05α=0.05. All intervals use 10410^4 percentile resamples under a fixed stream, resampling seeds within each configuration for the method CIs and per-seed differences for the difference CIs. Table 5 reports the comparisons against Table 1 and the ablations, with per-seed values in Table 7. Table 5: Pairwise comparisons (n=15n=15). Δ is the seed-mean difference (first minus second), CI its seed-paired bootstrap 95%95\% interval, d Cohen’s d, PoI the probability of improvement, and pWp_W/pUp_U the Holm-corrected exact Wilcoxon signed-rank and Mann–Whitney p. Bold marks p<0.05p<0.05. BF, Butterflies; c, catch/ep; p, peak. Comparison (metric) Δ 95%95\% CI d PoI pWp_W pUp_U PRIME −- RND (BF c) +0.210+0.210 [+0.137,+0.284][+0.137,+0.284] 1.671.67 0.890.89 0.0020.002 <−<10^-3 PRIME −- RND (BF p) +0.693+0.693 [+0.580,+0.793][+0.580,+0.793] 3.943.94 1.001.00 <−<10^-3 <−<10^-3 PRIME −- SMIRL (Maze p) +0.783+0.783 [+0.733,+0.837][+0.733,+0.837] 10.010.0 1.001.00 <−<10^-3 <−<10^-3 PRIME −- Disag. (BF c) +1.743+1.743 [+1.660,+1.825][+1.660,+1.825] 16.116.1 1.001.00 <−<10^-3 <−<10^-3 PRIME −- Disag. (BF p) +2.023+2.023 [+1.903,+2.127][+1.903,+2.127] 14.714.7 1.001.00 <−<10^-3 <−<10^-3 Full −- K=1K=1 (BF c) +0.176+0.176 [+0.065,+0.299][+0.065,+0.299] 1.231.23 0.810.81 0.0750.075 0.0190.019 Full −- K=1K=1 (BF p) +0.233+0.233 [+0.047,+0.427][+0.047,+0.427] 0.960.96 0.740.74 0.1920.192 0.1470.147 Full −- K=1K=1 (Maze p) +0.093+0.093 [+0.020,+0.170][+0.020,+0.170] 0.870.87 0.700.70 0.2310.231 0.2440.244 Full −- gate-off (BF c) +0.122+0.122 [+0.027,+0.206][+0.027,+0.206] 0.950.95 0.730.73 0.1280.128 0.1470.147 Full −- gate-off (BF p) +0.080+0.080 [−0.050,+0.197][-0.050,+0.197] 0.500.50 0.640.64 0.2310.231 0.2970.297 Full −- gate-off (Maze p) +0.067+0.067 [+0.003,+0.133][+0.003,+0.133] 0.630.63 0.670.67 0.2310.231 0.2970.297 All five PRIME–baseline comparisons clear α=0.05α=0.05 under the Holm-corrected exact signed-rank, with probability of improvement 0.890.89–1.001.00 and paired CIs excluding zero. PRIME separates from RND on both Butterflies metrics at n=15n=15 (catch d=1.67d=1.67, peak d=3.94d=3.94), where the n=5n=5 sample left the catch comparison unresolved. The full configuration leads the ablations on every cell with paired CIs excluding zero on five of six. The K=1K=1 Butterflies-catch effect clears the Holm family on the Mann–Whitney test (p=0.019p=0.019, signed-rank p=0.075p=0.075), and the remaining ablation effects are medium, d=0.5d=0.5–1.01.0, without surviving the correction. Holm-corrected Welch’s t, computed as a parametric check, agrees with the Mann–Whitney column on all eleven comparisons. The Extrinsic-DQN Maze mean reflects seed 77, which never reaches the goal within budget under the sparse extrinsic reward, while its IQM stays 1.001.00 (Table 6). Table 6: Per-method IQM with stratified-bootstrap 95%95\% CI (n=15n=15). Method Butterflies catch/ep Butterflies peak Maze peak PRIME 3.473.47 [3.42,3.53][3.42,3.53] 5.035.03 [4.94,5.09][4.94,5.09] 0.780.78 [0.72,0.86][0.72,0.86] K=1K=1 3.323.32 [3.21,3.40][3.21,3.40] 4.824.82 [4.63,4.97][4.63,4.97] 0.720.72 [0.64,0.77][0.64,0.77] gate-off 3.363.36 [3.25,3.44][3.25,3.44] 4.944.94 [4.82,5.04][4.82,5.04] 0.720.72 [0.67,0.78][0.67,0.78] RND 3.253.25 [3.17,3.33][3.17,3.33] 4.304.30 [4.19,4.43][4.19,4.43] 1.001.00 [1.00,1.00][1.00,1.00] SMIRL 2.702.70 [2.61,2.81][2.61,2.81] 3.943.94 [3.69,4.16][3.69,4.16] 0.000.00 [0.00,0.02][0.00,0.02] Disag. 1.711.71 [1.67,1.78][1.67,1.78] 2.992.99 [2.93,3.06][2.93,3.06] 1.001.00 [1.00,1.00][1.00,1.00] Ext-DQN 4.164.16 [3.70,4.51][3.70,4.51] 5.725.72 [5.49,5.86][5.49,5.86] 1.001.00 [1.00,1.00][1.00,1.00] Table 7: Per-seed values (seeds 0–1414). BF catch/ep and peak from the 250250k sweep, Maze peak from the 200200k sweep. Method 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 BF catch/ep PRIME 3.44 3.21 3.61 3.59 3.43 3.44 3.58 3.52 3.49 3.42 3.55 3.39 3.40 3.38 3.53 K=1K=1 3.43 3.39 3.23 3.39 3.22 3.50 2.93 3.28 3.23 3.21 2.94 3.32 3.46 3.45 3.37 gate-off 3.26 3.50 3.34 3.46 3.09 3.12 3.29 3.51 3.25 3.31 3.40 3.49 3.55 3.15 3.43 RND 3.41 3.23 3.18 3.38 3.39 3.02 3.21 3.20 3.30 3.22 3.56 3.11 3.19 3.08 3.35 SMIRL 3.01 2.75 2.55 2.86 2.24 3.09 2.72 2.69 2.65 2.57 2.86 2.54 2.76 2.74 2.58 Disag. 1.84 1.78 1.65 1.65 1.48 1.68 1.78 1.65 1.66 1.82 1.66 1.74 1.96 1.69 1.80 Ext-DQN 4.89 4.74 5.12 4.44 3.28 4.30 3.13 4.52 3.06 4.26 4.43 3.22 4.39 4.11 3.74 BF peak PRIME 4.85 4.70 5.00 5.15 5.00 4.85 5.15 5.05 4.95 4.95 5.20 5.05 5.10 5.00 5.20 K=1K=1 5.20 4.95 4.90 5.05 4.60 5.00 4.50 4.80 4.45 4.85 4.05 4.45 4.75 5.15 5.00 gate-off 5.00 5.25 4.95 4.95 4.70 4.70 4.90 5.05 4.65 5.10 4.70 5.10 5.10 4.90 4.95 RND 4.55 4.10 4.30 4.50 4.50 4.05 4.30 4.10 4.75 4.25 4.55 4.15 4.20 4.20 4.30 SMIRL 4.05 3.65 4.00 4.20 3.45 4.45 3.90 3.75 3.45 3.65 4.25 3.40 4.05 4.65 4.25 Disag. 2.85 3.25 3.05 3.05 2.70 3.15 3.00 2.90 2.95 3.00 2.95 3.00 3.10 3.05 2.85 Ext-DQN 6.00 5.75 6.00 5.95 5.25 5.85 5.10 5.90 5.40 5.85 5.60 5.05 5.80 5.65 5.70 Maze peak PRIME .95 .85 .75 .90 .65 .80 .70 .75 .80 .70 .90 .70 .65 1.0 .80 K=1K=1 .65 .50 .80 .70 .70 .70 .80 .50 .85 .80 .75 .65 .60 .70 .80 gate-off .70 .80 .60 .55 .75 .80 .65 .75 .65 .65 .70 .70 .80 .95 .85 RND 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 SMIRL .05 .00 .00 .00 .00 .05 .05 .00 .00 .00 .00 .00 .00 .00 .00 Disag. 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 1.0 Ext-DQN 1.0 1.0 1.0 1.0 1.0 1.0 1.0 .00 1.0 1.0 1.0 1.0 1.0 1.0 1.0 Table 8 reports the per-seed state–action mutual information of the executed policy (estimator in Appendix 0.A), the Maze value underlying Section 6.3. The measure is positive on every seed in both environments and sits below the uniform ceiling log5≈1.609 5≈ 1.609 nats, while a state-independent policy is zero by construction. Butterflies carries the larger value (1.007±0.0181.007± 0.018 against 0.563±0.0250.563± 0.025 nats). The Butterflies estimator keys on agent position and naliven_alive (Appendix 0.A), so its value reflects conditioning on catch progress alongside position, while the Maze value, keyed on position alone, is the unconfounded evidence of a state-conditional policy. Table 8: Per-seed state–action mutual information for PRIME (tail-20%20\% means, nats). Butterflies from the 250250k sweep, Maze from the 200200k sweep. Environment 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 Butterflies 1.043 0.989 1.018 1.016 1.012 1.001 1.022 1.002 1.016 0.985 1.029 0.991 0.989 1.012 0.978 Maze 0.529 0.560 0.550 0.616 0.535 0.606 0.583 0.568 0.567 0.568 0.577 0.551 0.540 0.537 0.557 Appendix 0.C Target frame and count realization The novelty term I(θ;S′∣s,a,D)I(θ;S s,a,D) admits more than one model-free estimator within the admissible class of Section 5. The return-space cross-head variance Vepi,LoTVV_epi,LoTV of Eq. (6) is the theoretical-target frame, the ensemble-disagreement signal on a value-side ensemble under shared πref _ref policy evaluation, with no posterior semantics claimed. The count Vepi,countV_epi,count of Eq. (7) is the operative estimator in the reward. Both are admissible bounded continuation bonuses, so Proposition 1, Lemma 1, Corollary 1, and Propositions 2–4 hold for either with the same constants. Both bound the same quantity. The count satisfies the chain IG≤PG≤1/N^IG ≤ 1/ N of [4], so Vepi,count=κγ2tanh2(1/Ns,a+1)V_epi,count=κγ^2 ^2(1/ N_s,a+1) is a bounded estimator of parameter information gain decaying as 1/Ns,a1/N_s,a. On the finite didactic state spaces the count is exact, an admissible realization of the novelty term, and the return-space variance of Eq. (6) is the alternative realization in the same class. The return-space frame collapses on these environments. Under shared πref _ref policy evaluation the K heads converge to a common fixed point of ΠW1Tπref W_1T _ref, a documented property of value-side ensembles under shared TD targets [36]. Randomized priors [27] address prior variance but do not decorrelate the targets. Cross-head variance therefore underestimates information gain, measured near 10−710^-7 (Section 6.4). The same convergence drives Proposition 4. The alignment of the conditional means [Yk∣s,a]E[Y_k s,a] that sends Vepi,LoTV→0V_epi,LoTV→ 0 is the mechanism by which the reward reduces to the aleatoric penalty on resolved support. Persistent cross-head disagreement would require decorrelated targets or a parameter-side ensemble outside the shared-πref _ref evaluation, a different construction left to a scalable extension where the count is infeasible. Figure 3: Reward-component magnitudes for PRIME, seed-mean traces (n=15n=15) of the count realization, the cross-head variance, and the aleatoric term ValeaugV_ale^aug, rolling means over logged batches on a log scale with display floor 10−810^-8. Butterflies from the 250250k sweep, Maze from the 200200k sweep. Figure 3 traces the deployed reward components over training. On Butterflies the cross-head variance falls from a 5.9×10−45.9× 10^-4 first-decile mean to a 5.0×10−75.0× 10^-7 tail-20%20\% mean while the count decays from 5.0×10−25.0× 10^-2 to 6.0×10−36.0× 10^-3, and on Maze the corresponding tails are 8.6×10−88.6× 10^-8 against 3.1×10−43.1× 10^-4, the count ending three to four orders of magnitude above the cross-head variance on both environments. The count decays under accumulating coverage toward the Proposition 4 limit. The aleatoric term separates the environments, 2.3×10−32.3× 10^-3 on stochastic Butterflies against 2.0×10−52.0× 10^-5 on deterministic Maze. Appendix 0.D Noisy-TV avoidance and the controllable bucket key The epistemic count keys on (agent row,agent col,a,nalive)(agent row,agent col,a,n_alive), the controllable abstraction, excluding butterfly positions. Butterfly motion is a parameter-free random walk, so by the decomposition of Eq. (3) its parameter information gain is zero and a count estimating I(θ;S′)I(θ;S ) should not count it. Computing novelty over a controllable or task-relevant abstraction is standard. Contingency-aware exploration counts agent-controllable position and excludes uncontrollable moving objects [9]; feature-space pseudocounts count feature-relevant structure in place of raw stochastic configurations [25]; inverse-dynamics curiosity [29] and representation-change rewards [31] exclude uncontrollable components by construction. The aleatoric detectors read the full observation. The probe V~aleraw V_ale^raw operates on ψ, the ℓ2 _2-normalized full observation including the butterfly channel, and its conditional variance over the (s,a)(s,a) bucket registers the butterfly motion. The policy receives butterfly positions as input, and the run-average catch rate above the position-only comparator (Section 6.2) is consistent with their use. Only the epistemic count excludes them, by the epistemic–aleatoric assignment. The didactic environments do not exercise active noise avoidance. Butterflies are captured without replacement, so the stochastic component depletes within the 100100-step horizon and no persistent distractor remains. RND, on the full observation, attains both environments (Table 1). The aleatoric penalty is correspondingly small on Butterflies (ValeaugV_ale^aug tail-20%20\% 2.3×10−32.3× 10^-3, Section 6.4) and the cross-environment direction holds with the gate disabled (Table 2). The experiments demonstrate direction-free epistemic behavior under one configuration. Active aleatoric avoidance is a property of the objective (Eq. (3), Proposition 2) whose demonstration needs a persistent-noise environment and is deferred. Appendix 0.E Theory and the deep-RL implementation The within-window results split by what the deployed system requires (Table 9). The deployed epistemic term is the tabular count, for which Proposition 1 and the bounds of Lemma 1 hold exactly, and the target ceiling of Corollary 1 requires none of the kernel, bandwidth, or β-mixing assumptions. The ceiling is checked against the empirical |qtaken|max|q_taken|_ in Figure 2a. The non-parametric concentration of Proposition 3 describes the scalable estimator class, and asymptotic neutrality (Proposition 4) is conditional on each distributional head converging in-hypothesis-class to the projected fixed point. Table 9: Each result against its status in the deployed tabular-count system. Result Status in deployment Function approximation Prop. 1 contraction exact none Lemma 1 bounds exact none Cor. 1 target ceiling conditional (Fig. 2a) soft |Qk|≤Qmax|Q_k|≤ Q_ Prop. 2 gate preference exact (specified reward) none Prop. 3 concentration conditional, scalable class yes Prop. 4 neutrality conditional (in-class) yes (statistics heads) Three premises hold only approximately in the deep implementation. A learned encoder is a fixed-kernel machine only under lazy training [21]. Once the features themselves move, the effective kernel and covering number drift, and bootstrapping can collapse the feature geometry [24], so the fixed-bandwidth and capacity premises of Proposition 3 are approximate. The snapshot stream is generated by a continually updated network under a non-stationary acting policy, so the stationary β-mixing premise reads as its non-stationary extension. In-class convergence is an assumption we do not establish. The distributional optimality operator is not a contraction and need not admit a fixed point [3], and the quantile contraction is proved for fixed-policy evaluation [11]. These errors fall on the return-space term, already collapsed, and on the control critics, whose target ceiling is verified directly (Fig. 2a) under a softly enforced |Qk|≤Qmax|Q_k|≤ Q_ . The deadly triad of function approximation, bootstrapping, and off-policy updates [38] bears on the function-approximation components, and the mitigations adopted (target networks, Double-DQN, the window-capped bootstrap magnitude) are those it identifies as reducing divergence. The tabular within-window guarantees are untouched. Appendix 0.F Deferred proofs Proof(Proof of Lemma 1) Probability-measure centering at most doubles the GEG_E bound on E^k− E_k^-. Averaging under πref _ref and multiplying by γ yield |gk|≤2GE|g_k|≤ 2G_E and |Yk|≤2γGE|Y_k|≤ 2γ G_E. Popoviciu’s inequality on the 4γGE4γ G_E range gives Var(Yk)≤4γ2GE2Var(Y_k)≤ 4γ^2G_E^2, and the law of total variance distributes this between Vepi,LoTVV_epi,LoTV and ValeV_ale. Vepi,count≤κγ2V_epi,count≤κγ^2 since tanh2≤1 ^2≤ 1. The probe bounds follow from |wi⊤ϕ(S′)|≤Bϕ|w_i φ(S )|≤ B_φ and Popoviciu. The reward bound follows from monotonicity of log(1+⋅) (1+·). Proof(Proof sketch of Proposition 3) The rate is standard for β-mixing nonparametric regression under capacity control. The argument uses U-statistic decoupling, blocking, and covering-number bounds [41] and is conditional on (i)–(iv). Proof(Proof of Proposition 4) Shared-target convergence aligns the conditional means [Yk∣s,a]E[Y_k s,a] across k, driving Vepi,LoTV→0V_epi,LoTV→ 0 (conditional on in-class convergence). The pseudocount realization satisfies Vepi,count=κγ2tanh2(1/Ns,a+1)→0as Ns,a→∞V_epi,count=κγ^2 ^2\! (1/ N_s,a+1 )→ 0 N_s,a→∞ by direct calculation. The aggregated Vepi→0V_epi→ 0 on the visited support gives V~epi(hg)→0 V_epi^(h_g)→ 0, and substituting in Eq. (10) yields the stated limit.