Paper deep dive
One-step lowest-variance selection in a Gaussian random-field model motivated by masked diffusion: Total correlation and a square root collision threshold
Linjun Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/21/2026, 5:08:56 AM
Summary
This paper proposes a stylized Gaussian random-field model to analyze one-step lowest-variance selection in masked discrete diffusion. It establishes that in a sub-square-root regime, the conditional Gaussian total correlation of selected positions vanishes, while at the square-root scale, it remains non-negligible, providing a stochastic-geometry baseline for understanding dependence costs in parallel decoding.
Entities (6)
Relation Signals (5)
Gaussian random-field model → motivates → masked discrete diffusion
confidence 90% · Motivated by confidence-guided parallel unmasking in masked discrete diffusion, we study a single selection step in a stylized Gaussian random-field model.
score field → represents → Uncertainty
confidence 90% · A locally dependent nonnegative score field represents position wise uncertainty
square-root scale → resultsin → non-negligible total correlation
confidence 90% · At the square-root scale, it remains non-negligible with positive asymptotic probability
sub-square-root regime → resultsin → vanishing total correlation
confidence 90% · In a conservative sub-square-root regime, the conditional Gaussian total correlation of the selected block vanishes in probability.
selection step → measuresdependencevia → Total Correlation
confidence 85% · Dependence among the selected positions is measured through a distance-dependent Gaussian correlation model... conditional Gaussian total correlation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Motivated by confidence-guided parallel unmasking in masked discrete diffusion, we study a single selection step in a stylized Gaussian random-field model. A locally dependent nonnegative score field represents position wise uncertainty, and the scheduler selects the K positions with the smallest scores. Dependence among the selected positions is measured through a distance-dependent Gaussian correlation model. This separation provides a tractable framework for quantifying how the geometry of low-score locations affects the dependence cost of factorized parallel decoding. We establish two complementary results. In a conservative sub-square-root regime, the conditional Gaussian total correlation of the selected block vanishes in probability. At the square-root scale, it remains non-negligible with positive asymptotic probability and admits a strictly positive expectation lower bound. Synthetic experiments support the predicted finite-size behavior. These results provide a rigorous stochastic-geometry baseline for understanding how budget size, score dependence, and spatial correlation jointly shape one-step confidence-based selection in masked discrete diffusion.
Tags
Links
- Source: https://arxiv.org/abs/2607.17522v1
- Canonical: https://arxiv.org/abs/2607.17522v1
Trouble viewing inline? Open PDF directly →
Full Text
70,332 characters extracted from source content.
Expand or collapse full text
One-step lowest-variance selection in a Gaussian random-field model motivated by masked diffusion: Total correlation and a square-root collision threshold Linjun Li Department of Mathematics, University of Pennsylvania, Philadelphia, PA. Abstract Motivated by confidence-guided parallel unmasking in masked discrete diffusion, we study a single selection step in a stylized Gaussian random-field model. A locally dependent nonnegative score field represents positionwise uncertainty, and the scheduler selects the K positions with the smallest scores. Dependence among the selected positions is measured through a distance-dependent Gaussian correlation model. This separation provides a tractable framework for quantifying how the geometry of low-score locations affects the dependence cost of factorized parallel decoding. We establish two complementary results. In a conservative sub-square-root regime, the conditional Gaussian total correlation of the selected block vanishes in probability. At the square-root scale, it remains non-negligible with positive asymptotic probability and admits a strictly positive expectation lower bound. Synthetic experiments support the predicted finite-size behavior. These results provide a rigorous stochastic-geometry baseline for understanding how budget size, score dependence, and spatial correlation jointly shape one-step confidence-based selection in masked discrete diffusion. Keywords: masked diffusion motivation; one-step selection; Gaussian random field; stochastic geometry; small-ball probability; total correlation; Poisson approximation. 1 Introduction We analyze a single static selection step in a stylized random-field model: from a length-N locally dependent scalar score field, select the K smallest values, and then evaluate a prescribed distance-dependent Gaussian dependence cost (total correlation) on the selected indices. The model contains no forward corruption process, learned reverse transition, context evolution, remasking, or multi-step schedule. Its purpose is to isolate a stochastic-geometry question suggested by confidence-based parallel decoding: In one selection step, how quickly may the budget KNK_N grow with sequence length N while the selected block has negligible dependence cost in the stylized Gaussian model? The motivation comes from the factorization used by parallel categorical decoders. Conditional on the current masked context, let PSP_S be the target joint categorical law on a selected set S, let PiP_i be its one-position marginals, and let QiQ_i be the one-position laws used by a product decoder. Whenever the divergences are finite, DKL(PS∥⨂i∈SQi)=DKL(PS∥⨂i∈SPi)⏟categorical total correlation+∑i∈SDKL(Pi∥Qi).D_KL\! (P_S\, \|\, _i∈ SQ_i )= D_KL\! (P_S\, \|\, _i∈ SP_i )_categorical total correlation+ _i∈ SD_KL(P_i\|Q_i). (1.1) Thus accurate marginal predictions do not by themselves justify a parallel product update (Watanabe, 1960; Cover and Thomas, 2006). Equation (1.1) concerns the actual categorical decoder. The Gaussian quantity analyzed below is only an analytically tractable cost motivated by its dependence term. We neither derive that Gaussian cost from a trained categorical decoder nor assume that the two total correlations are numerically equal. Discrete diffusion replaces continuous Gaussian noising by categorical, absorbing, masking, or continuous-time jump processes (Austin et al., 2021; Hoogeboom et al., 2021; Campbell et al., 2022; Hoogeboom et al., 2022). Masked variants permit out-of-order and parallel generation of tokenized images and language (Chang et al., 2022, 2023; Lou et al., 2024; Sahoo et al., 2024). Large diffusion language models have made schedule design operational at scale, and recent work analyzes token ordering, confidence-based unmasking, learned policies, and parallel budgets (Nie et al., 2025; Kim et al., 2025; Li and Cai, 2025; Chen et al., 2026; Hong et al., 2026; Jazbec et al., 2026). Those papers study components of actual reverse-time samplers. The present paper instead uses that literature as motivation for a one-step model of score geometry and dependence cost. A common operational rule is to unmask positions with high marginal confidence. For a categorical prediction pip_i, one possible scalar uncertainty summary is the trace of its one-hot covariance, Ci=diag(pi)−pipi⊤,tr(Ci)=1−‖pi‖22.C_i=diag(p_i)-p_ip_i , (C_i)=1-\|p_i\|_2^2. This example motivates an ordering only. We do not identify the vocabulary- dimensional covariance CiC_i with the Gaussian covariance introduced below, and we do not claim that the resulting score field is calibrated to any specific neural confidence measure. The modeling assumption is simply that smaller values of a scalar score correspond to positions that a one-step rule would rank as more confident. Fix an integer m≥1m≥ 1 and two correlation parameters 0<|ρV|<10<| _V|<1 and 0<|ρX|<10<| _X|<1. Let (Yi)i≥1(Y_i)_i≥ 1 be a stationary ℝmR^m-valued Gaussian AR(1) field with parameter ρV _V, and set Vi=‖Yi‖2.V_i=\|Y_i\|^2. (1.2) The field (Vi)(V_i) supplies the ordering scores; ρV _V controls clustering of unusually small scores. Separately, for the first N positions define (ΣN)ij=ViVjρX|i−j|,1≤i,j≤N.( _N)_ij= V_iV_j\, _X^|i-j|, 1≤ i,j≤ N. (1.3) The indices of the K smallest diagonal entries form SN,KS_N,K. An auxiliary vector X(N)∣ΣN∼N(0,ΣN)X^(N) _N N(0, _N) is then used solely to assign a Gaussian dependence cost to those indices. The construction is (Yi)i=1N⟶(Vi)i=1N⟶ΣN,ΣN⟶SN,K,ΣN⟶ℒ(X(N)∣ΣN).(Y_i)_i=1^N (V_i)_i=1^N _N, _N S_N,K, _N (X^(N) _N). (1.4) The two parameters ρV _V and ρX _X are externally specified and need not arise from a common categorical distribution. This separation makes the probability calculation transparent, but it is also why the construction should be read as a stylized one-step random-field model rather than as a derived model of masked diffusion. Table 1: Objects in the one-step random-field model. Object Mathematical role Scope of the interpretation pi,Cip_i,C_i Categorical prediction and one-hot covariance Motivate a possible scalar ranking only; they are not part of the Gaussian model. Yi,Vi=‖Yi‖2Y_i,V_i=\|Y_i\|^2 Locally dependent score field Determine the lowest-score indices; ρV _V controls lower-tail clustering. ΣN _N Gaussian gap-cost representation Attaches dependence through ρX _X; it is not derived from a categorical decoder. SN,KS_N,K Indices of the K smallest ViV_i Output of the single static selection step. X(N)∣ΣNX^(N) _N Auxiliary Gaussian vector Defines the conditional Gaussian total-correlation cost on the selected set. For a deterministic set S=i1<⋯<ikS=\i_1<·s<i_k\, write XS=(Xi1,…,Xik)X_S=(X_i_1,…,X_i_k). The model’s conditional Gaussian total correlation is TCΣN(S):=DKL(ℒ(XS∣ΣN)∥⨂i∈Sℒ(Xi∣ΣN)).TC_ _N(S):=D_KL\! (L(X_S _N)\, \|\, _i∈ SL(X_i _N) ). (1.5) For random SN,KS_N,K, this is a random variable on the outer probability space generated by (Yi)(Y_i). All convergence in probability statements refer to this outer randomness. Once S is fixed, standardization removes the diagonal values ViV_i from this Gaussian total correlation. Consequently, a smaller score does not by itself reduce the Gaussian dependence cost: score magnitudes matter only through the ranking and the resulting spacings of the selected indices. Stripped to its probabilistic core, the paper studies rare minima of a short-range dependent score field together with a decreasing gap-dependent cost. The score field is analytically tractable because Vi∼χm2V_i _m^2, and very small scores are dependent Gaussian small-ball events. If a threshold keeps a marginal fraction q, the expected number of selected pairs within distance L is of order NLq2NLq^2. With q≈K/Nq≈ K/N, this gives the collision scale LK2/NLK^2/N. The first main theorem converts spatial separation into an upper bound on the prescribed Gaussian total-correlation cost. Theorem 1.1 (First main result, readable form). Within the one-step random-field model, for every polynomial budget KN=⌊Nα⌋,0<α<12,K_N= N^α , 0<α< 12, the conditional Gaussian factorization cost on the selected indices vanishes: TCΣN(SN,KN)⟶0in probability.TC_ _N(S_N,K_N) 0 probability. (1.6) More generally, the same conclusion holds whenever there is an integer distance scale LN≥1L_N≥ 1 such that LNKN2N⟶0,KN|ρX|2(LN+1)⟶0. L_NK_N^2N 0, K_N| _X|^2(L_N+1) 0. (1.7) The fully quantified statement is Theorem 2.6 in Section 2; its proof is given in Appendix A. Theorem 1.1 states that a sufficiently conservative lowest-score set can grow with N without accumulating a non-negligible value of the prescribed Gaussian dependence cost. The second result describes the model’s critical collision scale. Define ϑρV,m:=(1−ρV2)−m/2,cρX:=−12log(1−ρX2)>0. _ _V,m:=(1- _V^2)^-m/2, c_ _X:=- 12 (1- _X^2)>0. (1.8) Theorem 1.2 (Second main result, readable form). Within the same one-step model, if KNN⟶λ∈(0,∞), K_N N λ∈(0,∞), then the conditional Gaussian total-correlation cost does not vanish. More precisely, lim infN→∞ℙ(TCΣN(SN,KN)≥cρX)≥1−exp−ϑρV,mλ2>0, _N→∞P\! (TC_ _N(S_N,K_N)≥ c_ _X )≥ 1- \- _ _V,mλ^2\>0, (1.9) and lim infN→∞[TCΣN(SN,KN)]≥cρXϑρV,mλ2>0. _N→∞E\! [TC_ _N(S_N,K_N) ]≥ c_ _X _ _V,mλ^2>0. (1.10) The fully quantified result is Theorem 5.1 in Section 5. The exact square-root limit is proved for adjacent selected pairs; it yields the one-sided lower bounds above because each adjacent pair contributes a fixed positive amount to the Gaussian gap cost. Contributions. 1. We formulate a transparent one-step stochastic-geometry model motivated by confidence-based unmasking, while explicitly separating it from a full masked-diffusion process. 2. For exact lowest-score selection, we prove a conservative regime in which the prescribed conditional Gaussian total-correlation cost vanishes. 3. At the square-root collision scale, we prove one-sided probability and expectation lower bounds showing that this Gaussian cost is non-vanishing. 4. We identify the auxiliary geometry of rare selected minima and verify the model’s asymptotic predictions in synthetic simulations of the same random field. Section 2 formalizes the one-step random-covariance model and states the first main theorem. Section 3 records internal model implications and simulations. Section 4 states the auxiliary sparsification theorem and the proof roadmap. Section 5 states the critical lower-bound theorem. Appendices A and B contain the technical proofs. The final sections discuss the relation to masked-diffusion scheduling, the model’s scope, and its limitations. 2 Stylized one-step random-covariance model and the first main theorem For N∈ℕN , write [N]=1,…,N[N]=\1,…,N\, and let ImI_m denote the m×m× m identity matrix. The symbols m, ρV _V, ρX _X, (Yi)(Y_i), (Vi)(V_i), ΣN _N, SN,KS_N,K, and TCΣN(S)TC_ _N(S) retain the meanings introduced in Section 1. All asymptotic statements are taken as N→∞N→∞, with m, ρV _V, and ρX _X fixed. This section supplies the formal probability-space construction and introduces the auxiliary notation used in the proofs. The construction is a static one-step model; it is not a reverse-time transition kernel and does not specify how the score field would evolve after selected positions are updated. 2.1 Latent variance field and covariance construction Realize the stationary field introduced in Section 1 as the ℝmR^m-valued Gaussian AR(1) chain Yi=ρVYi−1+1−ρV2ξi,i≥2,Y_i= _VY_i-1+ 1- _V^2\, _i, i≥ 2, (2.1) where Y1∼N(0,Im)Y_1 N(0,I_m), the innovations are independent N(0,Im)N(0,I_m), and Y1Y_1 is independent of them. The variance profile introduced in (1.2) is therefore Vi=‖Yi‖2V_i=\|Y_i\|^2; it is a covariance diagonal, not an estimator computed from repeated samples. For later matrix calculations, introduce the auxiliary notation DV,N=diag(V1,…,VN),RρX,N=(ρX|i−j|)i,j=1N.D_V,N=diag(V_1,…,V_N), R_ _X,N=( _X^|i-j|)_i,j=1^N. (2.2) The entrywise covariance specified in (1.3) then has the factorization ΣN=DV,N1/2RρX,NDV,N1/2. _N=D_V,N^1/2R_ _X,ND_V,N^1/2. (2.3) Proposition 2.1 (Validity of the random-covariance construction). Almost surely, Vi>0V_i>0 for every i, ΣN _N is positive definite, and (ΣN)ii=Vi.( _N)_i=V_i. (2.4) Moreover, on an extension of the probability space, let (Zi)i≥1(Z_i)_i≥ 1 be a scalar stationary Gaussian AR(1) chain with correlation parameter ρX _X, independent of (Yi)i≥1(Y_i)_i≥ 1, and set Xi=ViZi,i≥1.X_i= V_i\,Z_i, i≥ 1. (2.5) Then, for every N, X(N)=(X1,…,XN)∣ΣN∼N(0,ΣN).X^(N)=(X_1,…,X_N) _N N(0, _N). (2.6) Proof. Each YiY_i has a nondegenerate Gaussian density on ℝmR^m, so ℙ(Yi=0)=0P(Y_i=0)=0. A countable union shows that almost surely Vi>0V_i>0 for all i. The conditional-correlation matrix RρX,NR_ _X,N is positive definite for |ρX|<1| _X|<1. On the event that all Vi>0V_i>0, DV,N1/2D_V,N^1/2 is invertible, and the congruence DV,N1/2RρX,NDV,N1/2D_V,N^1/2R_ _X,ND_V,N^1/2 is positive definite. Its diagonal is Vi(RρX,N)ii=ViV_i(R_ _X,N)_i=V_i. Because (Zi)(Z_i) is independent of the variance field, its conditional law given ΣN _N is still centered Gaussian with covariance RρX,NR_ _X,N. Equation (2.5) therefore gives Cov(X(N)∣ΣN)=DV,N1/2RρX,NDV,N1/2=ΣN,Cov(X^(N) _N)=D_V,N^1/2R_ _X,ND_V,N^1/2= _N, which proves (2.6). ∎ The construction is static. First sample the score field and hence ΣN _N; next read its diagonal and determine the selected set; only then, if desired, use the independent chain (Zi)(Z_i) to realize the auxiliary Gaussian vector. The selection results depend on the random covariance but not on a realization of (Xi)(X_i). No claim is made that the pair (Vi,RhoX,N)(V_i,R_hoX,N) is induced by a common categorical decoder. Remark (Rank-preserving calibration invariance). Let g:(0,∞)→(0,∞)g:(0,∞)→(0,∞) be strictly increasing and define V~i=g(Vi),DV~,N=diag(V~1,…,V~N), V_i=g(V_i), D_ V,N=diag( V_1,…, V_N), and define Σ~N=DV~,N1/2RρX,NDV~,N1/2. _N=D_ V,N^1/2R_ _X,ND_ V,N^1/2. Then the exact lowest-K set is unchanged, because g preserves the score ordering, and standardization removes the transformed diagonal scale from every retained conditional Gaussian subvector. Consequently all spacing, Poisson, and total- correlation conclusions remain valid. In particular, a bounded increasing calibration may be used when matching the latent chi-square score to a bounded categorical uncertainty measure. 2.2 Chi-square small-ball structure Proposition 2.2 (Dependent chi-square variance field). The process (Yi)(Y_i) is centered Gaussian with Cov(Yi,Yj)=ρV|i−j|Im.Cov(Y_i,Y_j)= _V^|i-j|I_m. (2.7) Each diagonal variance satisfies Vi∼χm2V_i _m^2, and Cov(Vi,Vj)=2mρV2|i−j|.Cov(V_i,V_j)=2m _V^2|i-j|. (2.8) Proof. Iterating (2.1) gives Yj=ρVj−iYi+1−ρV2∑ℓ=i+1jρVj−ℓξℓ,j>i.Y_j= _V^j-iY_i+ 1- _V^2 _ =i+1^j _V^j- _ , j>i. The innovation sum is independent of YiY_i, which proves (2.7); stationarity gives Yi∼N(0,Im)Y_i N(0,I_m), so Vi=‖Yi‖2∼χm2V_i=\|Y_i\|^2 _m^2. Isserlis’ formula yields Cov(Vi,Vj)=2∑a,b=1mCov(Yi,a,Yj,b)2=2mρV2|i−j|.Cov(V_i,V_j)=2 _a,b=1^mCov(Y_i,a,Y_j,b)^2=2m _V^2|i-j|. ∎ Let Fm(u)=ℙ(χm2≤u),F_m(u)=P( _m^2≤ u), (2.9) and define um(q)=Fm−1(q),q∈(0,1).u_m(q)=F_m^-1(q), q∈(0,1). (2.10) As u↓0u 0, Fm(u)=um/22m/2Γ(m/2+1)(1+O(u)).F_m(u)= u^m/22^m/2 (m/2+1)(1+O(u)). (2.11) 2.3 Exact lowest-K variance selection Recall that SN,KS_N,K denotes the exact lowest-variance set introduced in Section 1. This set is almost surely well-defined without ties: for i≠ji≠ j, the vector (Yi,Yj)(Y_i,Y_j) has a nondegenerate Gaussian density and ‖Yi‖2−‖Yj‖2\|Y_i\|^2-\|Y_j\|^2 is a nonzero polynomial, whose zero set has Lebesgue measure zero. Definition (Threshold set). For q∈(0,1)q∈(0,1), define TN(q)=i∈[N]:Vi≤um(q).T_N(q)=\i∈[N]:V_i≤ u_m(q)\. (2.12) Then |TN(q)|=Nq.E|T_N(q)|=Nq. (2.13) For distance h≥1h≥ 1, set ph(q)=ℙ(V1≤um(q),V1+h≤um(q)).p_h(q)=P(V_1≤ u_m(q),\ V_1+h≤ u_m(q)). (2.14) Lemma 2.3 (Fixed-distance two-point small balls). For every fixed h≥1h≥ 1, ph(q)=(1−ρV2h)−m/2q2(1+o(1)),q↓0.p_h(q)=(1- _V^2h)^-m/2q^2(1+o(1)), q 0. (2.15) Proof. Put r=ρVhr= _V^h and u=um(q)u=u_m(q). The joint density of (Y1,Y1+h)(Y_1,Y_1+h) at the origin is (2π)−m(1−r2)−m/2.(2π)^-m(1-r^2)^-m/2. Continuity at the origin gives ℙ(‖Y1‖2≤u,‖Y1+h‖2≤u)=(2π)−m(1−r2)−m/2vm2um(1+o(1)),P(\|Y_1\|^2≤ u,\|Y_1+h\|^2≤ u)=(2π)^-m(1-r^2)^-m/2v_m^2u^m(1+o(1)), where vmv_m is the volume of the unit ball in ℝmR^m. Similarly, Fm(u)=(2π)−m/2vmum/2(1+o(1)).F_m(u)=(2π)^-m/2v_mu^m/2(1+o(1)). Dividing by Fm(u)2=q2F_m(u)^2=q^2 proves the claim. ∎ Definition (Collisions and spacing). For S⊂[N]S⊂[N] and L≥1L≥ 1, define CL(S)=∑h=1minL,N−1∑i=1N−hi∈Si+h∈S,C_L(S)= _h=1 \L,N-1\ _i=1^N-h1_\i∈ S\1_\i+h∈ S\, (2.16) A(S)=C1(S),A(S)=C_1(S), (2.17) and Δ(S)=mini≠j,i,j∈S|i−j|, (S)= _i≠ j,\ i,j∈ S|i-j|, (2.18) with Δ(S)=∞ (S)=∞ for |S|≤1|S|≤ 1. Then Δ(S)>L⟺CL(S)=0. (S)>L C_L(S)=0. (2.19) 2.4 Conditional total correlation Recall the conditional total-correlation notation from (1.5). For a deterministic set S=i1<⋯<ik⊂[N]S=\i_1<·s<i_k\⊂[N], standardization by the conditional marginal standard deviations Vi V_i removes the random diagonal scale. Hence the correlation matrix of XS∣ΣNX_S _N is RρX(S)=(ρX|ia−ib|)a,b=1k.R_ _X(S)=( _X^|i_a-i_b|)_a,b=1^k. (2.20) The Gaussian KL formula therefore gives the closed form TCΣN(S)=−12logdetRρX(S).TC_ _N(S)=- 12 R_ _X(S). (2.21) This identity applies pointwise to every realized covariance and every ΣN _N-measurable selected set. Lemma 2.4 (Exact AR(1) determinant). If S=i1<⋯<ikS=\i_1<·s<i_k\, then detRρX(S)=∏a=1k−1(1−ρX2(ia+1−ia)). R_ _X(S)= _a=1^k-1 (1- _X^2(i_a+1-i_a) ). (2.22) Proof. The subsampled stationary AR(1) chain with parameter ρX _X satisfies Zia+1=ρXia+1−iaZia+1−ρX2(ia+1−ia)ηa,Z_i_a+1= _X^i_a+1-i_aZ_i_a+ 1- _X^2(i_a+1-i_a)\, _a, with independent standard Gaussian innovations ηa _a. The covariance determinant is the product of the first marginal variance, equal to one, and the successive conditional variances. ∎ Lemma 2.5 (Separated sets have small total correlation). Let |S|=k≥2|S|=k≥ 2 and suppose Δ(S)>L (S)>L. Then 0≤TCΣN(S)≤k−12|ρX|2(L+1)1−|ρX|2(L+1).0 _ _N(S)≤ k-12\, | _X|^2(L+1)1-| _X|^2(L+1). (2.23) Proof. Write S=i1<⋯<ikS=\i_1<·s<i_k\ and ga=ia+1−iag_a=i_a+1-i_a. By Lemma 2.4, TCΣN(S)=−12∑a=1k−1log(1−ρX2ga).TC_ _N(S)=- 12 _a=1^k-1 (1- _X^2g_a ). The spacing assumption gives ga≥L+1g_a≥ L+1, and hence 0≤ρX2ga≤|ρX|2(L+1)<10≤ _X^2g_a≤| _X|^2(L+1)<1. Using −log(1−x)≤x/(1−x)- (1-x)≤ x/(1-x) for 0≤x<10≤ x<1, each summand is at most 12|ρX|2(L+1)1−|ρX|2(L+1). 12\, | _X|^2(L+1)1-| _X|^2(L+1). Summing over a=1,…,k−1a=1,…,k-1 proves (2.23). ∎ Theorem 2.6 (First main theorem: subcritical vanishing of conditional total correlation). Let (KN)N≥1(K_N)_N≥ 1 and (LN)N≥1(L_N)_N≥ 1 be integer sequences satisfying 1≤KN≤N,LN≥1.1≤ K_N≤ N, L_N≥ 1. Assume LNKN2N⟶0 L_NK_N^2N 0 (2.24) and KN|ρX|2(LN+1)⟶0.K_N| _X|^2(L_N+1) 0. (2.25) Then the exact lowest-variance set SN,KNS_N,K_N satisfies TCΣN(SN,KN)⟶0in probability,TC_ _N(S_N,K_N) 0 probability, (2.26) where probability is taken over the latent variance field (Yi)(Y_i), equivalently over the realized random covariance ΣN _N. The proof is given in Appendix A. Its essential geometric input is the auxiliary sparsification theorem stated in Section 4. 3 Internal model interpretation and numerical illustration This section interprets and checks the asymptotic statements only within the stylized one-step model. 3.1 Budget implications within the model The auxiliary sparsification theorem gives an internal geometric criterion. To make the probability of selecting two positions within radius LNL_N vanish, it is sufficient that KN=o(NLN).K_N=o\! ( NL_N ). For the auxiliary Gaussian coordinates to have vanishing conditional total correlation, the exact AR(1) determinant further requires KN|ρX|2(LN+1)→0.K_N| _X|^2(L_N+1)→ 0. For KN=⌊Nα⌋K_N= N^α , 0<α<1/20<α<1/2, a logarithmic separation LN=⌈clogN⌉L_N= c N with c>α2log(1/|ρX|)c> α2 (1/| _X|) satisfies both conditions. At KN∼λNK_N λ N, Theorem 5.1 gives a non-vanishing lower bound for the model’s Gaussian gap cost. Its auxiliary Poisson theorem identifies the corresponding probability of at least one adjacent selected pair. Every fixed positive value of KN/NK_N/ N belongs to this critical window; λ=1λ=1 is not a distinguished boundary. 3.2 Synthetic checks of the random-field model We simulate the same mathematical model with m=3,ρV=ρX=0.5.m=3, _V= _X=0.5. For each realization, the exact lowest-score set is selected from Vi=‖Yi‖2V_i=\|Y_i\|^2. If SN,K=s1<⋯<sKS_N,K=\s_1<·s<s_K\, the prescribed Gaussian total-correlation cost is TCΣN(SN,K)=−12∑a=1K−1log(1−ρX2(sa+1−sa)).TC_ _N(S_N,K)=- 12 _a=1^K-1 \! (1- _X^2(s_a+1-s_a) ). (3.1) We use N∈1024,4096,16384N∈\1024,4096,16384\ and 500 independent score-field realizations per sequence length. Figure 1 plots the mean Gaussian cost on a logarithmic vertical axis against K/NK/ N. The approximate collapse of the curves supports K/NK/ N as the finite-size scaling variable in this model. It does not establish a categorical factorization error, a neural unmasking threshold, or a generation-quality barrier. Figure 1: Mean conditional Gaussian total-correlation cost in the stylized one-step model. Here m=3m=3, ρV=ρX=0.5 _V= _X=0.5, N∈1024,4096,16384N∈\1024,4096,16384\, and each point uses 500 independent score-field realizations. The vertical axis is logarithmic, and error bars are 95% Monte Carlo confidence intervals for the mean. The figure checks the internal scaling of the model, not a trained masked-diffusion decoder. Figure 2 checks the auxiliary critical theorem more directly. The empirical probability of no adjacent selected pair is compared with the Poisson prediction, evaluated at λ=K/Nλ=K/ N. This is an internal asymptotic check of the score field. Figure 2: Probability of no adjacent selected pair in the stylized score model for m=3m=3, ρV=ρX=0.5 _V= _X=0.5, N∈1024,4096,16384N∈\1024,4096,16384\, and 500 independent realizations per sequence length. Solid points show Monte Carlo estimates with 95% confidence intervals; the continuous curve is the auxiliary Poisson prediction exp−ϑρV,mλ2 \- _ _V,mλ^2\. 4 Auxiliary sparsification for the one-step score model and proof roadmap The first main theorem is proved through the following static geometric statement. It concerns only the locations of rare minima of the score field. Its role is to show that the selected indices are separated far enough for the prescribed AR(1) Gaussian gap cost to vanish. Theorem 4.1 (Auxiliary sparsification theorem). Let (KN)N≥1(K_N)_N≥ 1 and (LN)N≥1(L_N)_N≥ 1 be integer sequences satisfying 1≤KN≤N,LN≥1,1≤ K_N≤ N, L_N≥ 1, and LNKN2N⟶0. L_NK_N^2N 0. (4.1) Then ℙ(CLN(SN,KN)=0)⟶1,P (C_L_N(S_N,K_N)=0 ) 1, (4.2) or equivalently ℙ(Δ(SN,KN)>LN)⟶1.P ( (S_N,K_N)>L_N ) 1. (4.3) In particular, if KN=o(N)K_N=o( N), then ℙ(A(SN,KN)=0)⟶1.P (A(S_N,K_N)=0 ) 1. (4.4) Theorem 4.1 is the geometric input to the first main result. On the event Δ(SN,KN)>LN (S_N,K_N)>L_N, the exact AR(1) determinant bound from Lemma 2.5 gives 0≤TCΣN(SN,KN)≤KN−12|ρX|2(LN+1)1−|ρX|2(LN+1).0 _ _N(S_N,K_N)≤ K_N-12\, | _X|^2(L_N+1)1-| _X|^2(L_N+1). (4.5) Thus the sparsification probability in (4.3), together with the correlation-decay assumption of Theorem 2.6, forces the conditional total correlation to vanish. Appendix A contains the full proof of the auxiliary sparsification theorem and the formal deduction of the first main theorem. 5 Critical non-vanishing in the one-step Gaussian cost model The first main theorem gives a regime in which the model’s Gaussian cost vanishes. This section proves the complementary statement that, at every fixed positive square-root budget, the same cost is non-negligible. The conclusion is internal to the one-step model and is not a lower bound on categorical decoding error or on per-token generation loss. Recall the constants ϑρV,m _ _V,m and cρXc_ _X from (1.8). Theorem 5.1 (Second main theorem: critical non-vanishing of conditional total correlation). Let (KN)N≥1(K_N)_N≥ 1 be an integer sequence satisfying 1≤KN≤N1≤ K_N≤ N and KNN⟶λ∈(0,∞). K_N N λ∈(0,∞). Let ZλZ_λ have the Poisson distribution with mean ϑρV,mλ2 _ _V,mλ^2. Then, for every integer r≥1r≥ 1, lim infN→∞ℙ(TCΣN(SN,KN)≥rcρX)≥ℙ(Zλ≥r). _N→∞P\! (TC_ _N(S_N,K_N)≥ rc_ _X ) (Z_λ≥ r). (5.1) In particular, lim infN→∞ℙ(TCΣN(SN,KN)≥cρX)≥1−exp−ϑρV,mλ2>0, _N→∞P\! (TC_ _N(S_N,K_N)≥ c_ _X )≥ 1- \- _ _V,mλ^2\>0, (5.2) and lim infN→∞[TCΣN(SN,KN)]≥cρXϑρV,mλ2>0. _N→∞E\! [TC_ _N(S_N,K_N) ]≥ c_ _X _ _V,mλ^2>0. (5.3) Consequently, TCΣN(SN,KN)TC_ _N(S_N,K_N) does not converge to zero in probability at a fixed positive critical budget. The proof uses the following adjacent-pair limit as an auxiliary theorem. Theorem 5.2 (Auxiliary critical adjacent-pair limit). Let (KN)N≥1(K_N)_N≥ 1 be an integer sequence satisfying 1≤KN≤N1≤ K_N≤ N. If KNN⟶λ∈(0,∞), K_N N λ∈(0,∞), then A(SN,KN)⇒Poisson(ϑρV,mλ2).A(S_N,K_N) ( _ _V,mλ^2). (5.4) In particular, ℙ(A(SN,KN)=0)⟶exp−ϑρV,mλ2.P (A(S_N,K_N)=0 ) \- _ _V,mλ^2\. (5.5) The proof of Theorem 5.2, including the factorial-moment argument and the two-sided threshold sandwich, is given in Appendix B. Proof of Theorem 5.1. Write a deterministic selected set as S=i1<⋯<ikS=\i_1<·s<i_k\. By Lemma 2.4, TCΣN(S)=−12∑a=1k−1log(1−ρX2(ia+1−ia)).TC_ _N(S)=- 12 _a=1^k-1 \! (1- _X^2(i_a+1-i_a) ). Every index a with ia+1−ia=1i_a+1-i_a=1 contributes exactly cρXc_ _X, and all other summands are nonnegative. Hence TCΣN(S)≥cρXA(S).TC_ _N(S)≥ c_ _XA(S). (5.6) For every integer r≥1r≥ 1, A(SN,KN)≥r⊆TCΣN(SN,KN)≥rcρX.\A(S_N,K_N)≥ r\ \TC_ _N(S_N,K_N)≥ rc_ _X\. Theorem 5.2 implies ℙ(A(SN,KN)≥r)⟶ℙ(Zλ≥r),P(A(S_N,K_N)≥ r) (Z_λ≥ r), which proves (5.1) and, by taking r=1r=1, (5.2). For the expectation bound, (5.6) gives [TCΣN(SN,KN)]≥cρX[A(SN,KN)].E[TC_ _N(S_N,K_N)]≥ c_ _XE[A(S_N,K_N)]. For each M>0M>0, the function x↦x∧Mx x M is bounded and continuous on ℕ0N_0. Therefore the auxiliary Poisson convergence gives limN→∞[A(SN,KN)∧M]=[Zλ∧M]. _N→∞E[A(S_N,K_N) M]=E[Z_λ M]. Since A(SN,KN)≥A(SN,KN)∧MA(S_N,K_N)≥ A(S_N,K_N) M, first take the lower limit in N and then let M→∞M→∞. Monotone convergence yields lim infN→∞[A(SN,KN)]≥[Zλ]=ϑρV,mλ2. _N→∞E[A(S_N,K_N)] [Z_λ]= _ _V,mλ^2. This proves (5.3). ∎ Corollary 5.3 (Adjacent-collision phase diagram). For the exact lowest-variance set SN,KNS_N,K_N: 1. if KN=o(N)K_N=o( N), then ℙ(A(SN,KN)=0)→1P(A(S_N,K_N)=0)→ 1; 2. if KN/N→λ∈(0,∞)K_N/ N→λ∈(0,∞), then A(SN,KN)⇒Poisson(ϑρV,mλ2)A(S_N,K_N) ( _ _V,mλ^2); 3. if KN/N→∞K_N/ N→∞, then ℙ(A(SN,KN)≥1)⟶1.P(A(S_N,K_N)≥ 1) 1. (5.7) Proof of Corollary 5.3. The first statement is Theorem 4.1 with LN≡1L_N≡ 1. The second statement is Theorem 5.2. For the third statement, fix any λ0>0 _0>0 and define the auxiliary budget K~N=⌊λ0N⌋ K_N= _0 N . Since KN/N→∞K_N/ N→∞, we have K~N≤KN K_N≤ K_N for all sufficiently large N. By monotonicity of the order-statistic sets, SN,K~N⊆SN,KNS_N, K_N S_N,K_N (5.8) almost surely for all sufficiently large N. Therefore ℙ(A(SN,KN)=0)≤ℙ(A(SN,K~N)=0).P (A(S_N,K_N)=0 ) (A(S_N, K_N)=0 ). By Theorem 5.2, the right-hand side converges to exp−ϑρV,mλ02 \- _ _V,m _0^2\. Taking lim sup and then letting λ0→∞ _0→∞ gives lim supN→∞ℙ(A(SN,KN)=0)=0, _N→∞P (A(S_N,K_N)=0 )=0, which is equivalent to (5.7). ∎ Corollary 5.4 (Supercritical total-correlation obstruction). If KN/N→∞K_N/ N→∞, then ℙ(TCΣN(SN,KN)≥cρX)⟶1.P\! (TC_ _N(S_N,K_N)≥ c_ _X ) 1. (5.9) Proof. Equation (5.6) and Corollary 5.3 give ℙ(TCΣN(SN,KN)≥cρX)≥ℙ(A(SN,KN)≥1)⟶1.P\! (TC_ _N(S_N,K_N)≥ c_ _X ) (A(S_N,K_N)≥ 1) 1. ∎ Remark (Scope of the square-root result). The auxiliary collision results, proved in Appendix B, identify K≍NK N as the exact adjacent-collision scale. Their principal role here is to prove the second main total-correlation theorem, Theorem 5.1. That theorem gives explicit probability and expectation lower bounds at fixed positive critical budgets, while Theorem 2.6 gives a sufficient subcritical regime for vanishing. We do not claim an exact limiting distribution or a complete phase diagram for the total-correlation statistic itself. 6 Discussion The mathematical content of the paper can be summarized without diffusion terminology. A locally dependent nonnegative score field is sampled, the K smallest scores are selected, and the selected spacings are charged the gap cost (total correlation) f(h)=−12log(1−ρX2h).f(h)=- 12 (1- _X^2h). The first theorem shows that this cost vanishes under a conservative budget and correlation-decay regime. The second shows that adjacent rare minima create a non-vanishing one-sided obstruction at the square-root collision scale. This is a stochastic-geometry result about selected minima plus a prescribed decreasing gap cost. The Gaussian covariance representation provides a convenient information- theoretic interpretation of that gap cost, but it should not be confused with a categorical model derived from a masked decoder. The score field and dependence field have separate parameters, ρV _V and ρX _X, and no common hidden categorical distribution is specified. Moreover, once the selected set is known, diagonal score magnitudes cancel from Gaussian total correlation. Low scores affect the conclusion only by changing where selected indices occur; low marginal variance does not itself make a fixed Gaussian pair less dependent. Actual scheduling work studies reverse-time token ordering, learned unmasking policies, and global multi-step errors (Kim et al., 2025; Li and Cai, 2025; Chen et al., 2026; Hong et al., 2026; Jazbec et al., 2026). The present model offers a tractable baseline for a single confidence-ranked selection step: it isolates how short-range lower-tail clustering can create nearby selections and how an externally specified distance-dependent cost responds. Likewise, its relation to Gaussian explanations of diffusion is more conceptual (Wang and Vastola, 2024; Sahoo et al., 2025). Limitations. The analysis treats one static selection step, assumes a chi-square short-range score field, and attaches an AR(1) Gaussian dependence cost that is not derived from the same categorical law as the confidence score. It does not prove that scores from a trained masked model have this lower-tail geometry, that neural token dependence decays with distance as assumed, or that Gaussian total correlation approximates categorical factorization error. These restrictions make the result a transparent null model and suggest concrete empirical questions: measure lower-tail score collisions and conditional dependence in trained models, and then test whether a similar one-step scaling law is observed. 7 Conclusion Motivated by confidence-based parallel unmasking, we analyzed a special one-step Gaussian random-field model. A locally dependent score field determines the K selected indices, and a separately specified AR(1) kernel assigns a conditional Gaussian total-correlation cost to their spacings. Within this model, the cost vanishes under an explicit conservative budget and correlation-decay regime. At the square-root collision scale, adjacent selected pairs have a Poisson limit and imply non-vanishing probability and expectation lower bounds for the same Gaussian cost. The contribution is thus a stochastic-geometry baseline for one confidence- ranked selection step. Its value is to separate a clean probabilistic mechanism–rare minima, selected-set geometry, and a distance-dependent cost that can be tested when richer score and dependence structures are measured in trained masked models. Acknowledgement The author thanks Yifen Chen for helpful discussions. Appendix A Proofs for the first main theorem This appendix proves the auxiliary sparsification theorem, Theorem 4.1, and then gives the formal deduction of Theorem 2.6. The argument passes from exact lowest-K selection to a slightly denser threshold set, controls the threshold-set size and short-range collisions, and finally combines the resulting separation with the exact AR(1) total-correlation bound. Throughout this appendix, ρV _V controls the variance field and ρX _X enters only through the final total-correlation deduction. The proof has three components. First, we establish a uniform two-point small-ball upper bound whose excess over the independent value is summable over distances. Second, we control the cardinality and short-range collision count of threshold sets. Third, we enlarge the exact lowest-K rule to a slightly denser threshold rule and transfer the estimates back to the exact order statistic. A.1 Uniform two-point small-ball bounds The fixed-distance asymptotic in Lemma 2.3 identifies the leading constant for each fixed distance. For the subcritical theorem, however, we need a bound that is uniform in the distance and whose dependence on the distance is summable. The next lemma supplies this estimate. Lemma A.1 (Uniform two-point small-ball domination). Fix m≥1m≥ 1 and 0<|ρV|<10<| _V|<1. There exist q0∈(0,1)q_0∈(0,1) and a nonnegative sequence (ah)h≥1(a_h)_h≥ 1 such that AρV:=∑h=1∞ah<∞,A_ _V:= _h=1^∞a_h<∞, (A.1) and, for every h≥1h≥ 1 and every q∈(0,q0]q∈(0,q_0], ph(q)=ℙ(V1≤um(q),V1+h≤um(q))≤(1+ah)q2.p_h(q)=P (V_1≤ u_m(q),\,V_1+h≤ u_m(q) )≤(1+a_h)q^2. (A.2) Consequently there is a finite constant BρVB_ _V, depending only on (ρV,m)( _V,m), such that ph(q)≤BρVq2,h≥1,q∈(0,q0].p_h(q)≤ B_ _Vq^2, h≥ 1, q∈(0,q_0]. (A.3) Proof. Let G,H∈ℝmG,H ^m be centered Gaussian vectors satisfying Cov(G)=Im,Cov(H)=Im,Cov(G,H)=rIm,Cov(G)=I_m, (H)=I_m, (G,H)=rI_m, where |r|<1|r|<1. Relative to the product law of two independent N(0,Im)N(0,I_m) vectors, the joint law of (G,H)(G,H) has Radon–Nikodym derivative Rr(x,y)=(1−r2)−m/2exp2rx⋅y−r2(‖x‖2+‖y‖2)2(1−r2).R_r(x,y)=(1-r^2)^-m/2 \ 2rx· y-r^2( x ^2+ y ^2)2(1-r^2) \. (A.4) Indeed, the covariance matrix is (ImrImrImIm), pmatrixI_m&rI_m\\ rI_m&I_m pmatrix, whose determinant is (1−r2)m(1-r^2)^m, and its inverse is 11−r2(Im−rIm−rImIm). 11-r^2 pmatrixI_m&-rI_m\\ -rI_m&I_m pmatrix. Dividing the corresponding joint density by the product standard Gaussian density gives (A.4). Set u0=1u_0=1 and define q0:=Fm(u0).q_0:=F_m(u_0). (A.5) If 0<q≤q00<q≤ q_0, then um(q)≤u0u_m(q)≤ u_0. On the event ‖x‖2≤um(q),‖y‖2≤um(q), x ^2≤ u_m(q), y ^2≤ u_m(q), we have 2rx⋅y−r2(‖x‖2+‖y‖2)≤2|r|‖x‖‖y‖≤2|r|um(q)≤2|r|u0.2rx· y-r^2( x ^2+ y ^2)≤ 2|r| x y ≤ 2|r|u_m(q)≤ 2|r|u_0. Thus Rr(x,y)≤(1−r2)−m/2exp|r|u01−r2.R_r(x,y)≤(1-r^2)^-m/2 \ |r|u_01-r^2 \. (A.6) For the AR(1) field, the correlation at distance h is rh=ρVhr_h= _V^h. Define 1+ah:=(1−rh2)−m/2exp|rh|u01−rh2,h≥1.1+a_h:=(1-r_h^2)^-m/2 \ |r_h|u_01-r_h^2 \, h≥ 1. (A.7) The right-hand side is at least one, so ah≥0a_h≥ 0. Integrating (A.6) over the product event ‖x‖2≤um(q)×‖y‖2≤um(q)\ x ^2≤ u_m(q)\×\ y ^2≤ u_m(q)\ under two independent standard Gaussian measures yields ph(q)≤(1+ah)Fm(um(q))2=(1+ah)q2,p_h(q)≤(1+a_h)F_m(u_m(q))^2=(1+a_h)q^2, which proves (A.2). It remains to prove summability. Since |rh|=|ρV|h→0|r_h|=| _V|^h→ 0, there exists h0h_0 such that |rh|≤1/2|r_h|≤ 1/2 for all h≥h0h≥ h_0. For such h, using −log(1−t)≤2t- (1-t)≤ 2t for 0≤t≤1/40≤ t≤ 1/4, we obtain log(1+ah)=−m2log(1−rh2)+|rh|u01−rh2≤mrh2+2u0|rh|≤(m+2u0)|rh|. (1+a_h)=- m2 (1-r_h^2)+ |r_h|u_01-r_h^2≤ mr_h^2+2u_0|r_h|≤(m+2u_0)|r_h|. For all sufficiently large h, the last bound is at most one. Since ex−1≤exe^x-1≤ ex for 0≤x≤10≤ x≤ 1, it follows that ah≤e(m+2u0)|ρV|ha_h≤ e(m+2u_0)| _V|^h for all sufficiently large h. Hence ∑hah<∞ _ha_h<∞. Finally, (A.3) follows from (A.2) with BρV:=1+suph≥1ah<∞.B_ _V:=1+ _h≥ 1a_h<∞. ∎ A.2 Threshold-set estimates For q∈(0,1)q∈(0,1), recall the threshold set TN(q)=i∈[N]:Vi≤um(q)T_N(q)=\i∈[N]:V_i≤ u_m(q)\ and write MN(q):=|TN(q)|.M_N(q):=|T_N(q)|. (A.8) The next two lemmas are the only probabilistic inputs needed for exact lowest-K selection. Lemma A.2 (Short-range collisions in a threshold set). There exists a finite constant Ccol=Ccol(ρV,m)C_col=C_col( _V,m) such that, for all N≥1N≥ 1, all integers L≥1L≥ 1, and all q∈(0,q0]q∈(0,q_0], ℙ(CL(TN(q))≥1)≤CcolNLq2.P (C_L(T_N(q))≥ 1 )≤ C_colNLq^2. (A.9) Proof. By Markov’s inequality and the definition of CLC_L, ℙ(CL(TN(q))≥1)≤CL(TN(q))=∑h=1L∑i=1N−hℙ(i∈TN(q),i+h∈TN(q)).P (C_L(T_N(q))≥ 1 ) _L(T_N(q))= _h=1^L _i=1^N-hP(i∈ T_N(q),\,i+h∈ T_N(q)). If h≥Nh≥ N, the inner sum is empty; otherwise stationarity gives ℙ(i∈TN(q),i+h∈TN(q))=ph(q).P(i∈ T_N(q),\,i+h∈ T_N(q))=p_h(q). Using Lemma A.1, CL(TN(q))≤∑h=1L(N−h)+(1+ah)q2≤Nq2∑h=1L(1+ah).EC_L(T_N(q))≤ _h=1^L(N-h)_+(1+a_h)q^2≤ Nq^2 _h=1^L(1+a_h). Since L≥1L≥ 1 and ∑h≥1ah=AρV<∞ _h≥ 1a_h=A_ _V<∞, ∑h=1L(1+ah)≤L+AρV≤(1+AρV)L. _h=1^L(1+a_h)≤ L+A_ _V≤(1+A_ _V)L. Thus (A.9) holds with Ccol=1+AρVC_col=1+A_ _V. ∎ Lemma A.3 (Concentration of the threshold-set size). There exists a finite constant Cvar=Cvar(ρV,m)C_var=C_var( _V,m) such that, for all N≥1N≥ 1 and all q∈(0,q0]q∈(0,q_0], Var(MN(q))≤CvarNq.Var(M_N(q))≤ C_varNq. (A.10) Consequently, for every deterministic sequence qN∈(0,q0]q_N∈(0,q_0] satisfying NqN→∞Nq_N→∞, MN(qN)NqN⟶1in probability. M_N(q_N)Nq_N 1 probability. (A.11) Proof. Let Ii(q)=Vi≤um(q),i∈[N].I_i(q)=1_\V_i≤ u_m(q)\, i∈[N]. Then MN(q)=∑i=1NIi(q)M_N(q)= _i=1^NI_i(q), Ii(q)=qEI_i(q)=q, and Var(MN(q))=∑i=1NVar(Ii(q))+2∑1≤i<j≤NCov(Ii(q),Ij(q)).Var(M_N(q))= _i=1^NVar(I_i(q))+2 _1≤ i<j≤ NCov(I_i(q),I_j(q)). The diagonal part is bounded by ∑i=1NVar(Ii(q))≤Nq. _i=1^NVar(I_i(q))≤ Nq. For pairs at distance h, Lemma A.1 gives [Ii(q)Ii+h(q)]≤(1+ah)q2,E[I_i(q)I_i+h(q)]≤(1+a_h)q^2, and hence Cov(Ii(q),Ii+h(q))=[Ii(q)Ii+h(q)]−q2≤ahq2.Cov(I_i(q),I_i+h(q))=E[I_i(q)I_i+h(q)]-q^2≤ a_hq^2. Therefore 2∑1≤i<j≤NCov(Ii(q),Ij(q))≤2∑h=1N−1(N−h)ahq2≤2Nq2AρV.2 _1≤ i<j≤ NCov(I_i(q),I_j(q))≤ 2 _h=1^N-1(N-h)a_hq^2≤ 2Nq^2A_ _V. Since q≤1q≤ 1, Var(MN(q))≤Nq+2Nq2AρV≤(1+2AρV)Nq.Var(M_N(q))≤ Nq+2Nq^2A_ _V≤(1+2A_ _V)Nq. This proves (A.10) with Cvar=1+2AρVC_var=1+2A_ _V. If NqN→∞Nq_N→∞, then for every ε>0 >0, Chebyshev’s inequality gives ℙ(|MN(qN)NqN−1|>ε)≤Var(MN(qN))ε2N2qN2≤Cvarε2NqN⟶0.P ( | M_N(q_N)Nq_N-1 |> )≤ Var(M_N(q_N)) ^2N^2q_N^2≤ C_var ^2Nq_N 0. Thus (A.11) follows. ∎ A.3 From threshold sets to exact lowest-K selection The following elementary comparison is the bridge from analytically convenient threshold sets to the exact lowest-K variance rule. Lemma A.4 (Threshold enlargement contains the lowest-K set). Let 1≤K≤N1≤ K≤ N, q∈(0,1)q∈(0,1), and u=um(q)u=u_m(q). If MN(q)≥KM_N(q)≥ K, then SN,K⊆TN(q).S_N,K T_N(q). (A.12) Consequently, for every integer L≥1L≥ 1, CL(SN,K)≥1⊆MN(q)<K∪CL(TN(q))≥1.\C_L(S_N,K)≥ 1\ \M_N(q)<K\∪\C_L(T_N(q))≥ 1\. (A.13) Proof. If MN(q)≥KM_N(q)≥ K, then at least K of the variances V1,…,VNV_1,…,V_N are at most u. Hence the K-th order statistic of the multiset V1,…,VN\V_1,…,V_N\ is at most u. Every index in SN,KS_N,K has score no larger than this K-th order statistic, and therefore has score at most u. This proves (A.12). If CL(SN,K)≥1C_L(S_N,K)≥ 1 and MN(q)≥KM_N(q)≥ K, then SN,K⊆TN(q)S_N,K T_N(q), so the same pair is also counted by CL(TN(q))C_L(T_N(q)). This proves (A.13). ∎ A.4 Proof of the subcritical spacing theorem We now prove the square-root sparsification theorem for the exact lowest-K set. Proof of Theorem 4.1. Set αN:=LNKN2N. _N:= L_NK_N^2N. (A.14) By assumption, αN→0 _N→ 0. For all sufficiently large N, define qN:=KNNαN−1/4.q_N:= K_NN _N^-1/4. (A.15) Since αN>0 _N>0, this is well-defined. We record the three elementary consequences of this choice. First, qNKN/N=αN−1/4⟶∞, q_NK_N/N= _N^-1/4 ∞, (A.16) so the threshold density is asymptotically much larger than the nominal selection fraction KN/NK_N/N. Equivalently, KNNqN=αN1/4⟶0. K_NNq_N= _N^1/4 0. (A.17) Second, NLNqN2=NLN(KNN)2αN−1/2=αN1/2⟶0.NL_Nq_N^2=NL_N ( K_NN )^2 _N^-1/2= _N^1/2 0. (A.18) Third, NqN=KNαN−1/4=KN1/2(NLN)1/4⟶∞.Nq_N=K_N _N^-1/4=K_N^1/2 ( NL_N )^1/4 ∞. (A.19) To justify the last convergence, note that KN≥1K_N≥ 1 and LN/N=αN/KN2≤αN→0L_N/N= _N/K_N^2≤ _N→ 0, so (N/LN)1/4→∞(N/L_N)^1/4→∞. Also, qN→0q_N→ 0: using KN/N=αN/(LNN)K_N/N= _N/(L_NN), qN=αN1/4LNN≤αN1/4⟶0.q_N= _N^1/4 L_NN≤ _N^1/4 0. (A.20) Thus for all sufficiently large N, qN∈(0,q0]q_N∈(0,q_0], where q0q_0 is the constant from Lemma A.1. By Lemma A.3 and Chebyshev’s inequality, ℙ(MN(qN)<KN)≤ℙ(|MN(qN)−NqN|>NqN−KN).P(M_N(q_N)<K_N) (|M_N(q_N)-Nq_N|>Nq_N-K_N ). By (A.17), for all sufficiently large N, NqN−KN≥12NqNNq_N-K_N≥ 12Nq_N. Hence, using (A.10), ℙ(MN(qN)<KN)≤4Var(MN(qN))N2qN2≤4CvarNqN⟶0,P(M_N(q_N)<K_N)≤ 4Var(M_N(q_N))N^2q_N^2≤ 4C_varNq_N 0, (A.21) where the last convergence follows from (A.19). On the other hand, Lemma A.2 and (A.18) give ℙ(CLN(TN(qN))≥1)≤CcolNLNqN2=CcolαN1/2⟶0.P(C_L_N(T_N(q_N))≥ 1)≤ C_colNL_Nq_N^2=C_col _N^1/2 0. (A.22) Combining (A.13), (A.21), and (A.22), we obtain ℙ(CLN(SN,KN)≥1)≤ℙ(MN(qN)<KN)+ℙ(CLN(TN(qN))≥1)⟶0.P(C_L_N(S_N,K_N)≥ 1) (M_N(q_N)<K_N)+P(C_L_N(T_N(q_N))≥ 1) 0. This proves (4.2). The equivalence with (4.3) follows from (2.19). Finally, taking LN≡1L_N≡ 1 turns (4.1) into KN2/N→0K_N^2/N→ 0, which is exactly KN=o(N)K_N=o( N), and C1(S)=A(S)C_1(S)=A(S) by (2.17). Thus (4.4) follows. ∎ A.5 Deduction of the first main theorem from sparsification We next prove the total-correlation theorem. The deterministic input is Lemma 2.5: once the selected sites are separated, the retained AR(1) correlations are uniformly small. Proof of Theorem 2.6. If KN=1K_N=1, then TCΣN(SN,KN)=0TC_ _N(S_N,K_N)=0. Otherwise define EN:=Δ(SN,KN)>LN.E_N:=\ (S_N,K_N)>L_N\. Assumption (2.24) is the assumption of Theorem 4.1, so ℙ(EN)→1P(E_N)→ 1. On ENE_N, Lemma 2.5 gives 0≤TCΣN(SN,KN)≤KN−12|ρX|2(LN+1)1−|ρX|2(LN+1).0 _ _N(S_N,K_N)≤ K_N-12\, | _X|^2(L_N+1)1-| _X|^2(L_N+1). (A.23) Because LN≥1L_N≥ 1 and 0<|ρX|<10<| _X|<1, the denominator is bounded below by 1−|ρX|4>01-| _X|^4>0. The deterministic right-hand side therefore tends to zero under (2.25). For every ε>0 >0, for all sufficiently large N, ℙ(TCΣN(SN,KN)>ε)≤ℙ(ENc)⟶0.P (TC_ _N(S_N,K_N)> ) (E_N^c) 0. This proves (2.26). ∎ Appendix B Proof of the auxiliary critical adjacent-pair limit This appendix proves Theorem 5.2. The proofs for Theorem 2.6 and its auxiliary sparsification theorem are in Appendix A. The present proof first establishes uniform finite-set small-ball bounds, then asymptotic decoupling for well-separated adjacent blocks, and finally a factorial-moment Poisson limit for threshold selection. A two-sided threshold sandwich transfers that limit to the exact lowest-K set. For a threshold density q∈(0,1)q∈(0,1), define the adjacent threshold event Bi(q)=Vi≤um(q),Vi+1≤um(q),1≤i≤N−1,B_i(q)=\V_i≤ u_m(q),\ V_i+1≤ u_m(q)\, 1≤ i≤ N-1, (B.1) and its count AN(q)=∑i=1N−1Bi(q).A_N(q)= _i=1^N-11_B_i(q). (B.2) B.1 Auxiliary small-ball estimates We first record two estimates for finite collections of small-ball events. The first is a uniform upper bound for any fixed number of sites. The second says that well-separated adjacent blocks asymptotically decouple. Lemma B.1 (Uniform finite-set small-ball upper bound). For every integer s≥1s≥ 1, there exist constants qs∈(0,1)q_s∈(0,1) and Cs<∞C_s<∞, depending only on (s,ρV,m)(s, _V,m), such that for every N, every set of distinct indices i1,…,is∈[N]i_1,…,i_s∈[N], and every q∈(0,qs]q∈(0,q_s], ℙ(Vi1≤um(q),…,Vis≤um(q))≤Csqs.P (V_i_1≤ u_m(q),…,V_i_s≤ u_m(q) )≤ C_sq^s. (B.3) Proof. Order the indices as i1<⋯<isi_1<·s<i_s, and let R=(ρV|ia−ib|)a,b=1sR=( _V i_a-i_b )_a,b=1^s. The Gaussian vector (Yi1,…,Yis)∈(ℝm)s(Y_i_1,…,Y_i_s)∈(R^m)^s has covariance R⊗ImR I_m. We compare its law with the product law of s independent N(0,Im)N(0,I_m) vectors. We first note that the eigenvalues of R are bounded away from zero by a constant depending only on (s,ρV)(s, _V). Indeed, if ga=ia+1−iag_a=i_a+1-i_a, then the scalar AR(1) Markov property gives detR=∏a=1s−1(1−ρV2ga)≥(1−ρV2)s−1. R= _a=1^s-1(1- _V^2g_a)≥(1- _V^2)^s-1. (B.4) Also λmax(R)≤s _ (R)≤ s, since every entry of R has absolute value at most one. Hence λmin(R)≥(1−ρV2)s−1s−1=:cs,ρV>0. _ (R)≥ (1- _V^2)^s-1s^s-1=:c_s, _V>0. (B.5) Consequently both det(R)−m/2 (R)^-m/2 and ‖R−1−Is‖op R^-1-I_s _op are bounded by constants depending only on (s,ρV,m)(s, _V,m). The density ratio of the joint law with respect to the product standard Gaussian law is ℛR(x1,…,xs)=det(R)−m/2exp−12∑a,b=1s(R−1−Is)abxa⋅xb.R_R(x_1,…,x_s)= (R)^-m/2 \- 12 _a,b=1^s(R^-1-I_s)_abx_a· x_b \. (B.6) Fix u0>0u_0>0 and set qs:=Fm(u0)q_s:=F_m(u_0). On the event ‖xa‖2≤um(q)≤u0 x_a ^2≤ u_m(q)≤ u_0 for all a, the exponent in (B.6) is bounded above by a constant depending only on (s,ρV,m,u0)(s, _V,m,u_0). Thus ℛR≤CsR_R≤ C_s on this event, uniformly over the choice of the indices. Integrating over the product event gives ℙ(‖Yia‖2≤um(q), 1≤a≤s)≤Cs∏a=1sℙ(χm2≤um(q))=Csqs.P ( Y_i_a ^2≤ u_m(q),\ 1≤ a≤ s )≤ C_s _a=1^sP( _m^2≤ u_m(q))=C_sq^s. This proves the lemma. ∎ Lemma B.2 (Asymptotic decoupling of separated adjacent blocks). Fix an integer r≥1r≥ 1. Let qN↓0q_N 0 and let gN→∞g_N→∞ be an integer sequence. Let pN:=p1(qN)=ℙ(B1(qN)).p_N:=p_1(q_N)=P(B_1(q_N)). (B.7) Uniformly over all r-tuples of distinct indices i1,…,ir∈1,…,N−1i_1,…,i_r∈\1,…,N-1\ satisfying |ia−ib|>gN+1,a≠b, i_a-i_b >g_N+1, a≠ b, (B.8) one has ℙ(⋂a=1rBia(qN))=pNr(1+o(1)),P ( _a=1^rB_i_a(q_N) )=p_N^r(1+o(1)), (B.9) where the o(1)o(1) term may depend on (r,ρV,m,qN,gN)(r, _V,m,q_N,g_N) but is uniform in the locations of the blocks. Proof. The case r=1r=1 is immediate, so assume r≥2r≥ 2. For each block define Wa=(Yia,Yia+1)∈ℝ2m,1≤a≤r.W_a=(Y_i_a,Y_i_a+1) ^2m, 1≤ a≤ r. Each WaW_a has covariance Γ=(1ρVρV1)⊗Im. = pmatrix1& _V\\ _V&1 pmatrix I_m. Let W~a=Γ−1/2Wa W_a= ^-1/2W_a. Under the product law of independent blocks, (W~1,…,W~r)( W_1,…, W_r) is standard Gaussian in ℝ2mrR^2mr. Under the true joint law, its covariance matrix has the form I2mr+EN,I_2mr+E_N, with zero diagonal 2m×2m2m× 2m blocks. If (B.8) holds, then every site in block a is at distance at least gN+1g_N+1 from every site in block b, for a≠ba≠ b. Therefore the unwhitened cross-covariance between the two blocks has operator norm at most CρV,m|ρV|gNC_ _V,m _V ^g_N. Since Γ−1/2 ^-1/2 is fixed, there is a constant Cr,ρV,mC_r, _V,m such that ∥EN∥op≤Cr,ρV,m|ρV|gN=:εN⟶0. E_N _op≤ C_r, _V,m _V ^g_N=: _N 0. (B.10) For all sufficiently large N, εN<1/2 _N<1/2. The density ratio of the true joint law of the whitened vector z=(z1,…,zr)∈ℝ2mrz=(z_1,…,z_r) ^2mr with respect to the product standard Gaussian law is ℛN(z)=det(I2mr+EN)−1/2exp−12z⊤((I2mr+EN)−1−I2mr)z.R_N(z)= (I_2mr+E_N)^-1/2 \- 12z ((I_2mr+E_N)^-1-I_2mr )z \. (B.11) By (B.10), |logdet(I2mr+EN)|≤Cr,ρV,mεN (I_2mr+E_N) ≤ C_r, _V,m _N and ‖(I2mr+EN)−1−I2mr‖op≤Cr,ρV,mεN. (I_2mr+E_N)^-1-I_2mr _op≤ C_r, _V,m _N. On the event ⋂aBia(qN) _aB_i_a(q_N), each unwhitened block satisfies ‖Wa‖2≤2um(qN) W_a ^2≤ 2u_m(q_N). Since Γ−1/2 ^-1/2 is fixed, ‖z‖2≤Cr,ρV,mum(qN). z ^2≤ C_r, _V,mu_m(q_N). (B.12) As qN↓0q_N 0, um(qN)↓0u_m(q_N) 0. Combining (B.11)–(B.12), we get supz∈∩aBia(qN)|logℛN(z)|=o(1), _z∈ _aB_i_a(q_N) _N(z) =o(1), uniformly over all block locations satisfying (B.8). Thus ℛN=1+o(1)R_N=1+o(1) uniformly on the block event. Integrating with respect to the product block law gives ℙ(⋂a=1rBia(qN))=(1+o(1))∏a=1rℙ(Bia(qN))=pNr(1+o(1)),P ( _a=1^rB_i_a(q_N) )=(1+o(1)) _a=1^rP(B_i_a(q_N))=p_N^r(1+o(1)), which is (B.9). ∎ B.2 Poisson limit for threshold adjacent pairs We next prove the critical Poisson limit for the threshold set. This is the probabilistic core of the exact lowest-K result. Proposition B.3 (Threshold adjacent-pair Poisson limit). Let (qN)(q_N) satisfy NqN⟶afor some a∈(0,∞). N\,q_N a some a∈(0,∞). (B.13) Then AN(qN)⇒Poisson(ϑρV,ma2).A_N(q_N) ( _ _V,ma^2 ). (B.14) Proof. Set pN=p1(qN)p_N=p_1(q_N). By Lemma 2.3, pN=ϑρV,mqN2(1+o(1)).p_N= _ _V,mq_N^2(1+o(1)). (B.15) Hence the mean μN:=AN(qN)=(N−1)pN _N:=EA_N(q_N)=(N-1)p_N (B.16) satisfies μN⟶μ:=ϑρV,ma2. _N μ:= _ _V,ma^2. (B.17) We prove convergence of all factorial moments. Let (x)r=x(x−1)⋯(x−r+1)(x)_r=x(x-1)·s(x-r+1). For fixed r≥1r≥ 1, [(AN(qN))r]=∑∈ℐN,rℙ(⋂a=1rBia(qN)),E [(A_N(q_N))_r ]= _i _N,rP ( _a=1^rB_i_a(q_N) ), (B.18) where ℐN,rI_N,r is the set of ordered r-tuples of distinct indices from 1,…,N−1\1,…,N-1\. Choose gN=⌈logN⌉.g_N= N . (B.19) Call =(i1,…,ir)∈ℐN,ri=(i_1,…,i_r) _N,r good if |ia−ib|>gN+1 i_a-i_b >g_N+1 for all a≠ba≠ b, and bad otherwise. The number of bad ordered tuples is Or(Nr−1gN)O_r(N^r-1g_N), so the number of good tuples is Nr(1+o(1)).N^r(1+o(1)). (B.20) By Lemma B.2, the contribution of good tuples to (B.18) is NrpNr(1+o(1))=(NpN)r(1+o(1))⟶μr.N^rp_N^r(1+o(1))=(Np_N)^r(1+o(1)) μ^r. (B.21) It remains to show that bad tuples contribute o(1)o(1). Sort the starting points of a bad tuple as j1<⋯<jrj_1<·s<j_r, and partition them into clusters by placing a break between jℓj_ and jℓ+1j_ +1 exactly when jℓ+1−jℓ>gN+1j_ +1-j_ >g_N+1. If there are c clusters, then a bad tuple has 1≤c≤r−11≤ c≤ r-1. For a fixed value of c, the number of possible sorted clustered configurations is at most CrNc(gN+1)r−c,C_rN^c(g_N+1)^r-c, (B.22) and passing from sorted to ordered tuples only changes the constant CrC_r. A cluster containing e distinct adjacent-pair starts involves at least e+1e+1 distinct sites of the underlying chain. Since different clusters are separated, a configuration with r edge starts and c clusters involves at least r+cr+c distinct sites. The event attached to the tuple forces all these distinct sites to be below the same threshold. Since qN↓0q_N 0, Lemma B.1 may be applied uniformly over all possible values of s≤2rs≤ 2r. The probability attached to any such tuple is therefore bounded by CrqNr+c,C_rq_N^r+c, (B.23) for all sufficiently large N, after increasing CrC_r if necessary. The total bad contribution is thus at most Cr∑c=1r−1Nc(gN+1)r−cqNr+c.C_r _c=1^r-1N^c(g_N+1)^r-cq_N^r+c. (B.24) Since qN=O(N−1/2)q_N=O(N^-1/2) and gN=O(logN)g_N=O( N), each summand in (B.24) is bounded by Cr(logN)r−cNc−(r+c)/2=Cr(logN)r−cN−(r−c)/2,C_r( N)^r-cN^c-(r+c)/2=C_r( N)^r-cN^-(r-c)/2, which tends to zero because c≤r−1c≤ r-1. Hence the total bad contribution is o(1)o(1). Combining the good and bad contributions gives [(AN(qN))r]⟶μrfor every fixed r≥1.E [(A_N(q_N))_r ] μ^r every fixed r≥ 1. The convergence of factorial moments implies convergence of the corresponding ordinary moments, because ordinary moments are finite linear combinations of factorial moments of lower order. The bounded first moments give tightness. For each fixed integer r, bounded (r+1)(r+1)-st moments imply uniform integrability of the r-th powers, so every subsequential weak limit has the Poisson moments μrμ^r in factorial form. The Poisson distribution is moment-determinate because its moment-generating function is finite in a neighborhood of the origin. The method of moments therefore yields (B.14). ∎ B.3 The exact lowest-K critical window We now transfer the threshold Poisson limit to the exact order statistic and prove Theorem 5.2. The argument is a two-sided threshold sandwich. It is important that the threshold set has size of order N N but fluctuations only of order N1/4N^1/4, so fixed multiplicative perturbations of the threshold density contain the random lowest-K cutoff with high probability. Proof of Theorem 5.2. Fix ε∈(0,1) ∈(0,1), and define qN−=(1−ε)KNN,qN+=(1+ε)KNN.q_N^-=(1- ) K_NN, q_N^+=(1+ ) K_NN. (B.25) For all large N, both numbers lie in (0,1)(0,1), and NqN±⟶(1±ε)λ. Nq_N^± (1± )λ. Let MN±=|TN(qN±)|,AN±=AN(qN±).M_N^±=|T_N(q_N^±)|, A_N^±=A_N(q_N^±). By Lemma A.3, or directly by the variance bound (A.10), ℙ(MN−>KN) (M_N^->K_N) ≤Var(MN−)(KN−NqN−)2≤CvarNqN−ε2KN2⟶0, ≤ Var(M_N^-)(K_N-Nq_N^-)^2≤ C_varNq_N^- ^2K_N^2 0, (B.26) ℙ(MN+<KN) (M_N^+<K_N) ≤Var(MN+)(NqN+−KN)2≤CvarNqN+ε2KN2⟶0. ≤ Var(M_N^+)(Nq_N^+-K_N)^2≤ C_varNq_N^+ ^2K_N^2 0. (B.27) On the event EN(ε):=MN−≤KN≤MN+,E_N( ):=\M_N^-≤ K_N≤ M_N^+\, (B.28) the order statistic satisfies TN(qN−)⊆SN,KN⊆TN(qN+).T_N(q_N^-) S_N,K_N T_N(q_N^+). (B.29) Indeed, if MN−≤KNM_N^-≤ K_N, every site below the lower threshold is among the KNK_N smallest scores; if MN+≥KNM_N^+≥ K_N, the KNK_N-th smallest score is at most the upper threshold. Hence, on EN(ε)E_N( ), AN−≤A(SN,KN)≤AN+.A_N^-≤ A(S_N,K_N)≤ A_N^+. (B.30) By (B.26) and (B.27), ℙ(EN(ε))→1P(E_N( ))→ 1. Proposition B.3 gives AN±⇒Poisson(ϑρV,m(1±ε)2λ2).A_N^± ( _ _V,m(1± )^2λ^2 ). (B.31) Let Fγ(ℓ)=ℙ(Poisson(γ)≤ℓ)F_γ( )=P(Poisson(γ)≤ ). For every integer ℓ≥0 ≥ 0, the event EN(ε)E_N( ) and the sandwich (B.30) give ℙ(AN+≤ℓ)−ℙ(EN(ε)c)≤ℙ(A(SN,KN)≤ℓ)≤ℙ(AN−≤ℓ)+ℙ(EN(ε)c).P(A_N^+≤ )-P(E_N( )^c) (A(S_N,K_N)≤ ) (A_N^-≤ )+P(E_N( )^c). Using ℙ(EN(ε)c)→0P(E_N( )^c)→ 0 and (B.31), we obtain FϑρV,m(1+ε)2λ2(ℓ) F_ _ _V,m(1+ )^2λ^2( ) ≤lim infN→∞ℙ(A(SN,KN)≤ℓ) ≤ _N→∞P(A(S_N,K_N)≤ ) ≤lim supN→∞ℙ(A(SN,KN)≤ℓ)≤FϑρV,m(1−ε)2λ2(ℓ). ≤ _N→∞P(A(S_N,K_N)≤ )≤ F_ _ _V,m(1- )^2λ^2( ). (B.32) Finally let ε↓0 0. The Poisson distribution function is continuous in its mean parameter, so the lower and upper bounds in (B.32) both converge to FϑρV,mλ2(ℓ)F_ _ _V,mλ^2( ). Therefore the distribution functions of A(SN,KN)A(S_N,K_N) converge at every integer ℓ . Since all variables are supported on ℕ0N_0, differences of consecutive cdf values give convergence of every point mass, and hence weak convergence. This proves (5.4). Taking ℓ=0 =0 gives (5.5). ∎ References Austin et al. (2021) Austin, J., Johnson, D.D., Ho, J., Tarlow, D., van den Berg, R., 2021. Structured denoising diffusion models in discrete state-spaces. Adv. Neural Inf. Process. Syst. 34, 17981–17993. Campbell et al. (2022) Campbell, A., Benton, J., De Bortoli, V., Rainforth, T., Deligiannidis, G., Doucet, A., 2022. A continuous time framework for discrete denoising models. Adv. Neural Inf. Process. Syst. 35, 28266–28279. Chang et al. (2022) Chang, H., Zhang, H., Jiang, L., Liu, C., Freeman, W.T., 2022. MaskGIT: Masked generative image transformer, in: Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., p. 11315–11325. Chang et al. (2023) Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K.P., Freeman, W.T., Rubinstein, M., Li, Y., Krishnan, D., 2023. Muse: Text-to-image generation via masked generative transformers, in: Proc. 40th Int. Conf. Mach. Learn., Proc. Mach. Learn. Res. 202, 4055–4075. Chen et al. (2026) Chen, S., Cong, K., Li, J., 2026. Optimal inference schedules for masked diffusion models, in: Proc. Thirty-Ninth Conf. Learn. Theory, Proc. Mach. Learn. Res. 336, 1279–1311. Cover and Thomas (2006) Cover, T.M., Thomas, J.A., 2006. Elements of Information Theory, second ed. Wiley, Hoboken. https://doi.org/10.1002/047174882X. Hong et al. (2026) Hong, C., An, S., Kim, M.-S., Ye, J.C., 2026. Improving discrete diffusion unmasking policies beyond explicit reference policies, in: Fourteenth International Conference on Learning Representations; arXiv:2510.05725. Hoogeboom et al. (2021) Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., Welling, M., 2021. Argmax flows and multinomial diffusion: Learning categorical distributions. Adv. Neural Inf. Process. Syst. 34, 12454–12465. Hoogeboom et al. (2022) Hoogeboom, E., Gritsenko, A.A., Bastings, J., Poole, B., van den Berg, R., Salimans, T., 2022. Autoregressive diffusion models, in: International Conference on Learning Representations. Jazbec et al. (2026) Jazbec, M., Olausson, T.X., Béthune, L., Ablin, P., Kirchhof, M., Monteiro, J., Turrisi, V., Ramapuram, J., Cuturi, M., 2026. Learning unmasking policies for diffusion language models, in: Forty-Third International Conference on Machine Learning (oral spotlight). arXiv:2512.09106. Kim et al. (2025) Kim, J., Shah, K., Kontonis, V., Kakade, S.M., Chen, S., 2025. Train for the worst, plan for the best: Understanding token ordering in masked diffusions, in: Proc. 42nd Int. Conf. Mach. Learn., Proc. Mach. Learn. Res. 267, 30749–30768. Li and Cai (2025) Li, G., Cai, C., 2025. Breaking AR’s sampling bottleneck: Provable acceleration via diffusion language models. Adv. Neural Inf. Process. Syst. 38. Lou et al. (2024) Lou, A., Meng, C., Ermon, S., 2024. Discrete diffusion modeling by estimating the ratios of the data distribution, in: Proc. 41st Int. Conf. Mach. Learn., Proc. Mach. Learn. Res. 235, 32819–32848. Nie et al. (2025) Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., Li, C., 2025. Large language diffusion models. Adv. Neural Inf. Process. Syst. 38. Sahoo et al. (2024) Sahoo, S.S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J.T., Rush, A., Kuleshov, V., 2024. Simple and effective masked diffusion language models. Adv. Neural Inf. Process. Syst. 37, 130136–130184. https://doi.org/10.52202/079017-4135. Sahoo et al. (2025) Sahoo, S.S., Deschenaux, J., Gokaslan, A., Wang, G., Chiu, J.T., Kuleshov, V., 2025. The diffusion duality, in: Proc. 42nd Int. Conf. Mach. Learn., Proc. Mach. Learn. Res. 267, 52584–52619. Wang and Vastola (2024) Wang, B., Vastola, J.J., 2024. The unreasonable effectiveness of Gaussian score approximation for diffusion models and its applications. Trans. Mach. Learn. Res. Watanabe (1960) Watanabe, S., 1960. Information theoretical analysis of multivariate correlation. IBM J. Res. Dev. 4, 66–82. https://doi.org/10.1147/rd.41.0066.