Paper deep dive
The Label Defines the Timescale: Trait-State Limits of Temporal-Aggregate Learning
Xizhe Zhang
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $\Theta_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.
Tags
Links
- Source: https://arxiv.org/abs/2608.01587v1
- Canonical: https://arxiv.org/abs/2608.01587v1
Trouble viewing inline? Open PDF directly →
Full Text
64,182 characters extracted from source content.
Expand or collapse full text
The Label Defines the Timescale: Trait–State Limits of Temporal-Aggregate Learning Xizhe Zhang Abstract Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form Θg,T=T−1∫0TgZ(t)dt _g,T=T^-1 _0^Tg\Z(t)\\,dt when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an O(1)O(1) trait component and an O(T−1)O(T^-1) state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test–retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label—not duration or segment count alone—defines the relevant timescale. Introduction A model may stop improving because its architecture is inadequate, or because its input protocol does not contain the information required by the label. The distinction is especially important when labels summarize a long horizon while inputs are snapshots. A clinical score may refer to symptoms over weeks, a maintenance label to an operating cycle, and an ecological target to a season, while the model receives one interview, one inspection, or a few images. In such settings, a benchmark ceiling can be a sampling-protocol ceiling rather than a model-capacity ceiling. The protocol has at least three different dimensions. The number of segments M controls how precisely a fixed temporal support is measured; a window length w controls the support of one recording; and the number and locations of windows D control temporal coverage. These quantities are not interchangeable. More segments can denoise a snapshot and strengthen inference about a stable individual trait, but they do not reveal state innovations outside the observed support. Conversely, high cross-sectional accuracy can be driven by stable between-person differences even when the model has little sensitivity to within-person temporal change. Our nonlinear example is an occupation-time label: the fraction of a horizon during which a latent state exceeds a threshold. This abstraction captures frequency-type targets such as the proportion of time spent in a symptomatic, unsafe, or anomalous state. Unlike a temporal mean, an occupation label depends on the complete correlation structure after thresholding. It therefore exposes a general point: the statistical timescale of an input is defined jointly by the latent dynamics and the label functional. We make three contributions. First, we prove a trait–state asymptotic decomposition for Gaussian labels formed by temporal aggregation. It separates an O(1)O(1) cross-sectional channel from an O(T−1)O(T^-1) state channel and yields a nonzero snapshot-prediction limit. Second, we derive task-dependent effective temporal spans for mean and occupation labels; the latter depends on all higher-order correlation times, so matching the usual integral correlation time does not match temporal information. Third, we prove boundary localization: for occupation labels, state-driven label variance is largest for individuals whose stable trait lies at the threshold, while local-window efficiency lacks the same multiplicative concentration. Experiments verify these results and directly compare same-time segmentation with dispersed temporal coverage under an equal segment budget. Related Work Generalizability theory decomposes object, occasion, rater, and measurement facets and uses D-studies to compare acquisition designs (Cronbach et al. 1972; Shavelson and Webb 1991; Brennan 2001). Latent state–trait theory likewise separates enduring traits from occasion-specific states (Steyer et al. 1999). Variance-component and random-effects methods provide the corresponding estimation machinery (Searle et al. 1992; Robinson 1991). Our model uses the same conceptual separation but makes the prediction target a nonlinear functional of a correlated continuous-time state. Longitudinal and functional-data methods study covariance estimation, trajectory recovery, and measurement placement (Diggle et al. 2002; Ramsay and Silverman 2005; Ji and Müller 2017). Classical measurement-error theory distinguishes biological variation from noisy observation (Fuller 1987; Carroll et al. 2006). Sampling design and pseudoreplication also warn that correlated repeats do not equal independent support (Kish 1965; Cochran 1977; Hurlbert 1984; Pyper and Peterman 1998), while Bayesian design formalizes the value of choosing informative observations (Chaloner and Verdinelli 1995). We instead derive the maximum predictive information of a fixed protocol and show that its temporal value depends on the label functional. Weak supervision, learning from label proportions, and multiple-instance learning attach labels to bags rather than local instances (Dietterich et al. 1997; Quadrianto et al. 2009; Zhou 2018; Ilse et al. 2018). These frameworks do not by themselves determine what fraction of a long-horizon label is observable through a short temporal support. Our protocol ceiling is complementary: it bounds every learner using the same observations, independently of architecture or training-set size. Occupation functionals and Gaussian level crossings have a substantial probability literature (Kratz 2006; Adler and Taylor 2007); Hermite expansions and limit theory expose the role of higher-order correlations (Breuer and Major 1983). We use Plackett’s Gaussian derivative identity for threshold covariances (Plackett 1954). Discrete approximation of occupation functionals has sharp error theory (Altmeyer and Chorowski 2018; Altmeyer 2021), and Gaussian-process excursion-volume uncertainty and sequential design are well studied (Vazquez and Piera-Martinez 2006; Bect et al. 2012; Azzimonti et al. 2016; Bect et al. 2019). We do not claim novelty for excursion-volume posterior inference or adaptive point placement. Our focus is multiple independent objects, a stable trait plus correlated state, sparse noisy windows, and supervised labels whose temporal meaning changes with the aggregation functional. Temporal-Aggregate Learning For object i, let Zi(t)=αMi+1−αXi(t),t∈[0,T],Z_i(t)= α\,M_i+ 1-α\,X_i(t), t∈[0,T], (1) where Mi∼(0,1)M_i (0,1) is a stable trait, XiX_i is a zero-mean unit-variance stationary Gaussian process, and the two are independent. Write ρ(u) ρ(u) =CovXi(t),Xi(t+u), =Cov\X_i(t),X_i(t+u)\, (2) rα(u) r_α(u) =α+(1−α)ρ(u). =α+(1-α)ρ(u). (3) Thus Zi(t)Z_i(t) is marginally standard normal and α is its long-lag correlation. The unnormalized model μi+σsXi(t) _i+ _sX_i(t) is equivalent after rescaling, with α=σμ2/(σμ2+σs2)α= _μ^2/( _μ^2+ _s^2). For g∈L2(ϕ)g∈ L^2(φ), define the temporal-aggregate label Θi,g,T=1T∫0TgZi(t)dt. _i,g,T= 1T _0^Tg\Z_i(t)\\,dt. (4) We use g(z)=zg(z)=z for a mean label and gc(z)=z>cg_c(z)=1\z>c\ for an occupation-time label. A protocol π contains noisy linear window observations Yid=LdZi+εidY_id=L_dZ_i+ _id. A window of length wdw_d centered at tdt_d has LdZ=wd−1∫td−wd/2td+wd/2Z(u)duL_dZ=w_d^-1 _t_d-w_d/2^t_d+w_d/2Z(u)\,du. We distinguish the number of time windows D, their support lengths wdw_d, and the number of segments M used to estimate a fixed window. Increasing M can reduce measurement noise, but does not change the observed temporal set. Their primary roles are M M : precision, : precision, w w : local support, : local support, (D,t1:D) (D,t_1:D) : temporal coverage. : temporal coverage. Only w and (D,t1:D)(D,t_1:D) change temporal support; M refines measurement on support already observed. Under squared loss, the optimal protocol risk and explainability are Rπ∗=Var(Θg,T∣Yπ),ℐπ=1−Rπ∗Var(Θg,T).R_π^*=EVar( _g,T Y_π), _π=1- R_π^*Var( _g,T). (5) The quantity ℐπI_π is a protocol-level theoretical R2R^2, not the performance of a particular architecture. We also condition on MiM_i to isolate the state channel, preventing high cross-sectional prediction from being mistaken for successful tracking of within-person dynamics. Protocol Risk as a Common Tool Let (U,Vr)(U,V_r) be standard bivariate normal with correlation r, and define Cg(r)=Covg(U),g(Vr)C_g(r)=Cov\g(U),g(V_r)\. For linear Gaussian observations, let kπ(t)=CovZ(t),Yπk_π(t)=Cov\Z(t),Y_π\, Vπ=Var(Yπ)V_π=Var(Y_π), and mπ(t)=kπ(t)⊤Vπ−1Yπ,qπ(s,t)=kπ(s)⊤Vπ−1kπ(t).m_π(t)=k_π(t) V_π^-1Y_π, q_π(s,t)=k_π(s) V_π^-1k_π(t). (6) Here qπq_π is the covariance explained by the protocol; the posterior covariance is rα(s−t)−qπ(s,t)r_α(s-t)-q_π(s,t). Proposition 1 (Exact protocol-conditioned risk). For any g∈L2(ϕ)g∈ L^2(φ), Rπ∗=1T2∫0T∫0T[Cgrα(s−t)−Cgqπ(s,t)]dsdt.R_π^*= 1T^2 _0^T\! _0^T [C_g\r_α(s-t)\-C_g\q_π(s,t)\ ]\,ds\,dt. (7) The Bayes predictor is T−1∫0T[gZ(t)∣Yπ]dtT^-1 _0^TE[g\Z(t)\ Y_π]\,dt. Proof. Fubini’s theorem gives Var(Θg,T)=1T2∫0T∫0TCgrα(s−t)dsdt.Var( _g,T)= 1T^2 _0^T\! _0^TC_g\r_α(s-t)\\,ds\,dt. Now draw posterior-process replicas Z(1)Z^(1) and Z(2)Z^(2) independently conditional on YπY_π. Their unconditional marginals equal that of Z, while the Gaussian conditioning formula gives CovZ(1)(s),Z(2)(t) \Z^(1)(s),Z^(2)(t)\ =Cov(Z(s)∣Yπ), =Cov \E(Z(s) Y_π), (Z(t)∣Yπ) 33.99998ptE(Z(t) Y_π) \ =qπ(s,t). =q_π(s,t). (8) Conditional independence implies [gZ(1)(s)gZ(2)(t)] [g\Z^(1)(s)\g\Z^(2)(t)\] =[g(Z(s))∣Yπ =E [E\g(Z(s)) Y_π\ ⋅g(Z(t))∣Yπ], 27.0pt·E\g(Z(t)) Y_π\ ], (9) so the covariance of the two posterior means is Cgqπ(s,t)C_g\q_π(s,t)\. Integrating yields Var(Θg,T∣Yπ)=T−2∫0T∫0TCgqπ(s,t)dsdt.Var\E( _g,T Y_π)\=T^-2\! _0^T\!\! _0^TC_g\q_π(s,t)\\,ds\,dt. The law of total variance subtracts this quantity from Var(Θg,T)Var( _g,T), proving Eq. (7); conditional Fubini gives the stated predictor. ∎ For occupation labels, Cg(r)=Gc(r)=Pr(U>c,Vr>c)−Φ¯(c)2C_g(r)=G_c(r)= (U>c,V_r>c)- (c)^2; at c=0c=0, G0(r)=arcsin(r)/(2π)G_0(r)= (r)/(2π). Remark (Segments are useful, but they are not time). Let ℱMF_M be the information generated by an increasingly fine segmentation of a fixed observed set W. If ℱM↑σZ(t):t∈WF_M σ\Z(t):t∈ W\, Lévy’s upward theorem gives RM∗↓Var(Θg,T∣σZ(t):t∈W).R_M^* \! ( _g,T σ\Z(t):t∈ W\ ). More segments can reduce sensor noise and improve trait estimation, so the plateau may be lower than the risk of a coarse recording. They cannot reveal state innovations outside W; the title’s distinction therefore concerns temporal state coverage, not the usefulness of repeated measurements for precision. Protocol Ceilings and Learning Gaps For any measurable predictor f(Yπ)f(Y_π), the orthogonality of conditional expectation gives the exact decomposition Θg,T−f(Yπ)2=Rπ∗+f(Yπ)−fπ∗(Yπ)2.E\ _g,T-f(Y_π)\^2=R_π^*+E\f(Y_π)-f_π^*(Y_π)\^2. (10) Equivalently, its population coefficient of determination satisfies Rπ2(f)=ℐπ−f(Yπ)−fπ∗(Yπ)2Var(Θg,T).R_π^2(f)=I_π- E\f(Y_π)-f_π^*(Y_π)\^2Var( _g,T). (11) The first term is fixed by acquisition; only the second can be reduced by architecture, optimization, or more training objects. This separates two empirically similar forms of saturation: a large model gap under an informative protocol, and a small model gap near a low protocol ceiling. When ℐπI_π can be estimated from a calibrated temporal model, Rπ2(f)/ℐπR_π^2(f)/I_π is a descriptive ceiling-utilization ratio; it is meaningful only for the same target, loss, and protocol assumptions. Cross-sectional and monitoring benchmarks also answer different questions. Write Θg,T=hg(M)+Δg,T,(Δg,T∣M)=0. _g,T=h_g(M)+ _g,T, ( _g,T M)=0. Total explainability includes prediction of the stable hg(M)h_g(M) channel. A trait-conditioned state ceiling instead evaluates how much of Δg,T _g,T is recoverable from temporal windows. High total R2R^2 can therefore coexist with weak within-person tracking; reporting only the former can make protocol-limited state learning look like successful dynamics modeling. Trait–State Limits Center g and expand it in probabilists’ Hermite polynomials: g(z)−g(U) g(z)-Eg(U) =∑k≥1ak(g)k!Hk(z), = _k≥ 1 a_k(g)k!H_k(z), (12) Cg(r) C_g(r) =∑k≥1ak(g)2k!rk. = _k≥ 1 a_k(g)^2k!r^k. (13) Let τj=∫0∞ρ(u)jdu _j= _0^∞ρ(u)^j\,du. Theorem 1 (Trait–state decomposition). Assume 0≤ρ≤10≤ρ≤ 1 and the series below is finite. Then Var(Θg,T)=Vtrait(g)+Astate(g)T+o(T−1),Var( _g,T)=V_ trait(g)+ A_ state(g)T+o(T^-1), (14) where Vtrait(g) V_ trait(g) =Cg(α)=Var[g(Z(t))∣M], =C_g(α)=Var\! [E\g(Z(t)) M\ ], (15) Astate(g) A_ state(g) =2∫0∞[Cgα+(1−α)ρ(u)−Cg(α)]du =2 _0^∞ [C_g\α+(1-α)ρ(u)\-C_g(α) ]\,du (16) =2∑k≥1ak(g)2k!∑j=1k(kj)αk−j(1−α)jτj. =2 _k≥ 1 a_k(g)^2k! _j=1^k kjα^k-j(1-α)^j _j. (17) Proof. Stationarity and symmetry give Var(Θg,T)=2T2∫0T(T−u)Cgrα(u)du.Var( _g,T)= 2T^2 _0^T(T-u)C_g\r_α(u)\\,du. Writing Δg(u)=Cgrα(u)−Cg(α) _g(u)=C_g\r_α(u)\-C_g(α) separates this as Cg(α)+2T∫0∞u≤T(1−u/T)Δg(u)du.C_g(α)+ 2T _0^∞1\u≤ T\(1-u/T) _g(u)\,du. The assumed summability makes Δg _g integrable, so dominated convergence gives Astate(g)/T+o(T−1)A_ state(g)/T+o(T^-1). Substituting Eq. (13), expanding [α+(1−α)ρ(u)]k−αk[α+(1-α)ρ(u)]^k-α^k, and applying Tonelli’s theorem gives Eq. (17). Finally, [HkαM+1−αX∣M]=αk/2Hk(M),E\! [H_k\ αM+ 1-αX\ M ]=α^k/2H_k(M), and Hermite orthogonality yields Var[g(Z(t))∣M]=∑k≥1ak(g)2αk/k!=Cg(α)Var[E\g(Z(t)) M\]= _k≥ 1a_k(g)^2α^k/k!=C_g(α). ∎ The first term is a stable O(1)O(1) cross-sectional channel; the second is finite-horizon state variation. For one noisy point observation Y=Z(T/2)+εY=Z(T/2)+ , ε∼(0,ν2) (0,ν^2), limT→∞ℐY=Cgα2/(1+ν2)Cg(α)>0,α>0. _T→∞I_Y= C_g\α^2/(1+ν^2)\C_g(α)>0, α>0. (18) Indeed, qY(s,t)=rα(|s−T/2|)rα(|t−T/2|)/(1+ν2)q_Y(s,t)=r_α(|s-T/2|)r_α(|t-T/2|)/(1+ν^2) approaches α2/(1+ν2)α^2/(1+ν^2) away from a vanishing boundary fraction. Applying the exact risk identity and Cesàro convergence proves Eq. (18). Without a trait, a fixed snapshot explains a vanishing fraction of an increasingly long label. With a trait, cross-sectional prediction retains an O(1)O(1) channel, so apparent benchmark performance can remain substantial without learning temporal state dynamics. Corollary 1 (What D and M buy in the trait channel). For a mean label as T→∞T→∞, suppose D occasions are separated enough that their state terms are independent, and each occasion averages M segments with raw segment-noise variance σε2 _ ^2. Then the trait explainability is ℐtrait(D,M)=α+(1−α)/D+σε2/(DM).I_ trait(D,M)= α+(1-α)/D+ _ ^2/(DM). (19) Proof. The protocol average is Y¯=αMi+1−αX¯D+ε¯. Y= αM_i+ 1-α\, X_D+ . The independent terms have variances α, (1−α)/D(1-α)/D, and σε2/(DM) _ ^2/(DM). The Gaussian regression R2R^2 for predicting αMi αM_i from Y¯ Y is Cov(αMi,Y¯)2/[Var(αMi)Var(Y¯)]Cov( αM_i, Y)^2/[Var( αM_i)Var( Y)], giving Eq. (19). ∎ Same-time replication corresponds to D=1D=1: it removes measurement noise as M grows but plateaus at α. Temporally separated occasions also average transient state noise and can approach unit trait explainability. Thus more segments are useful for precision, while more time supplies an additional source of information. Importantly, Eq. (19) does not require the short-lag state kernel. After standardization, a conventional two-occasion test–retest design at a lag where the transient state correlation is negligible identifies α from cross-occasion covariance and σε2 _ ^2 from observed variance (or from within-occasion segments). Hence any suitable repeated-measurement dataset can already produce the trait-channel ceiling. Estimating ρ and τk≥2 _k≥ 2 is needed only for the state-channel quantities below. For mean labels, Vtrait=αV_ trait=α and Astate=2(1−α)τ1A_ state=2(1-α) _1. For occupation labels, every Hermite order contributes, making the state timescale task dependent. For an OU state kernel ρ(u)=e−u/τρ(u)=e^-u/τ, τj=τ/j _j=τ/j. Conditional on a trait value, an occupation label has standardized state threshold a and Aa=2τ∫01Ga(r)rdr,A0=τlog22.A_a=2τ _0^1 G_a(r)r\,dr, A_0= τ 22. (20) The closed form at the boundary independently matches the Hermite series. Task-Dependent Effective Time To isolate temporal information beyond the stable trait, condition on M=mM=m. Define hg(m)=X[gαm+1−αX(t)]h_g(m)=E_X\! [g\ αm+ 1-αX(t)\ ] and subtract this trait-conditional mean. For an occupation label, set a=(c−αm)/1−αa=(c- αm)/ 1-α. Let WwW_w be a standardized noisy state average over a window of length w, let rw(t)=CorrX(t),Wwr_w(t)=Corr\X(t),W_w\, and define Jk(w)=∫ℝrw(t)kdtJ_k(w)= _Rr_w(t)^k\,dt. Theorem 2 (Task-dependent state-effective span). Suppose the window remains interior as T→∞T→∞, rwk∈L1(ℝ)r_w^k∈ L^1(R) for every Hermite order with nonzero weight, and the weighted τk _k and Jk(w)2J_k(w)^2 series below are finite. Then ℐstate(w∣m)=ℓg(w;m)T+o(T−1).I_ state(w m)= _g(w;m)T+o(T^-1). (21) For a mean label, ℓmean(w) _ mean(w) =2τ1η(w)+ν2, = 2 _1η(w)+ν^2, (22) η(w) η(w) =2w2∫0w(w−u)ρ(u)du. = 2w^2 _0^w(w-u)ρ(u)\,du. (23) For an occupation label, ℓa(w)=∑k≥1Hk−1(a)2k!Jk(w)22∑k≥1Hk−1(a)2k!τk. _a(w)= _k≥ 1 H_k-1(a)^2k!J_k(w)^2 2 _k≥ 1 H_k-1(a)^2k! _k. (24) Proof. Conditional on M=mM=m, set gm(x)=g(αm+1−αx)g_m(x)=g( αm+ 1-αx) and expand gmX(t)−gm(U)=∑k≥1βk(m)k!HkX(t).g_m\X(t)\-Eg_m(U)= _k≥ 1 _k(m)k!H_k\X(t)\. Hermite orthogonality gives Var(Θg,T∣M=m) ( _g,T M=m) =Ag(m)T+o(T−1), = A_g(m)T+o(T^-1), (25) Ag(m) A_g(m) =2∑k≥1βk(m)2k!τk. =2 _k≥ 1 _k(m)^2k! _k. (26) Because WwW_w is standardized, Gaussian regression gives Hk(X(t))∣Ww=rw(t)kHk(Ww).E\H_k(X(t)) W_w\=r_w(t)^kH_k(W_w). Therefore Var(Θg,T∣Ww,M=m) \E( _g,T W_w,M=m)\ =1T2∑k≥1βk(m)2k! = 1T^2 _k≥ 1 _k(m)^2k! ×∫0Trw(t)kdt2. × \ _0^Tr_w(t)^k\,dt \^\!2. (27) For an interior window with integrable correlation-profile tails, the truncated integral converges to Jk(w)J_k(w); termwise convergence follows from the stated summability. Dividing explained variance by total state variance proves Eq. (21). For an occupation label, βk(m)=ϕ(a)Hk−1(a) _k(m)=φ(a)H_k-1(a), which gives Eq. (24). For a mean label only k=1k=1 remains. Since J1(w)=2τ1/η(w)+ν2J_1(w)=2 _1/ η(w)+ν^2, the ratio J1(w)2/(2τ1)J_1(w)^2/(2 _1) gives Eq. (23). ∎ For mean labels, only the first Hermite order remains. Occupation labels use all τk _k and all window profiles JkJ_k; the common ϕ(a)2φ(a)^2 factor cancels from Eq. (24). Thus physical duration w is not the statistical value of a window, and two kernels with equal τ1 _1 can still define different occupation-time information. For a noise-free OU window, ℓmean(w)=w2/[w−τ(1−e−w/τ)] _ mean(w)=w^2/[w-τ(1-e^-w/τ)], approaching 2τ2τ for w≪τw τ and w for w≫τw τ. Corollary 2 (Equal segment budget). Write ℓg(w;ν2,m) _g(w;ν^2,m) when the averaged window has noise variance ν2ν^2. Allocate N independent raw segments either to one fixed window or to N mutually separated windows. In the sparse long-horizon regime, ℐsame(N∣m) _ same(N m) =ℓg(w;σε2/N,m)T+o(T−1) = _g(w; _ ^2/N,m)T+o(T^-1) ⟶ℓg(w;0,m)T, _g(w;0,m)T, (28) ℐdispersed(N∣m) _ dispersed(N m) =Nℓg(w;σε2,m)T+o(N/T), = N _g(w; _ ^2,m)T+o(N/T), (29) until the sparse additivity approximation approaches saturation. Proof. For same-time segments, averaging N independent sensor errors changes only the local noise variance from σε2 _ ^2 to σε2/N _ ^2/N; Theorem 2 gives Eq. (28). For separated windows, cross-window covariance terms are negligible in the sparse regime, so their explained state variances add to first order. Summing N equal contributions and dividing by Ag(m)/TA_g(m)/T gives Eq. (29). ∎ Thus repeated segments can exhaust local measurement noise but have a finite state-information limit; separated windows purchase additional state support. For a mean label and point-like windows, the two expressions reduce to 2τ1T(1+σε2/N)and2Nτ1T(1+σε2), 2 _1T(1+ _ ^2/N) 2N _1T(1+ _ ^2), respectively. Theorem 3 (Boundary localization). Assume ρ(u)≥0ρ(u)≥ 0. For the trait-conditioned threshold a, Ga(r)=∫0rexp−a2/(1+s)2π1−s2ds.G_a(r)= _0^r \-a^2/(1+s)\2π 1-s^2\,ds. (30) Hence Aa=2∫0∞Gaρ(u)duA_a=2 _0^∞G_a\ρ(u)\\,du is even and strictly decreases with |a||a| whenever ρ is positive on a set of nonzero measure; it is maximized at a=0a=0. The factor ϕ(a)2φ(a)^2 cancels from Eq. (24), so ℓa(w) _a(w) varies only through reweighting of Hermite orders. Proof. Plackett’s identity gives ∂rPr(U>a,Vr>a)=exp−a2/(1+r)/[2π1−r2] _r (U>a,V_r>a)= \-a^2/(1+r)\/[2π 1-r^2]. At r=0r=0 the excess covariance is zero, so integration gives Eq. (30). For every r>0r>0 its integrand is even in a and strictly decreases with |a||a|. Integration over any nonnegative ρ preserves these properties and is strict when ρ>0ρ>0 on a set of positive measure. Finally, the occupation Hermite coefficient is ϕ(a)Hk−1(a)φ(a)H_k-1(a); the common squared factor ϕ(a)2φ(a)^2 cancels between the numerator and denominator of Eq. (24). ∎ Threshold-near individuals therefore have more state-driven label variance to explain. In the OU protocols studied below, ℓa(w) _a(w) decays far more slowly than AaA_a. For D separated sparse windows, the conditional residual state risk is approximately Rstate(D,w∣m)≈AaT1−Dℓa(w)T.R_ state(D,w m)≈ A_aT \1- D _a(w)T \. (31) Consequently, absolute residual error is largest near the threshold even when the fraction of state variance explained changes much less. The distinction is important: boundary-near individuals are not necessarily observed with a much less efficient window; rather, their labels contain substantially more state variation that the protocol must explain. Implications for ML Benchmarks Equations (10)–(11) give a ceiling-aware interpretation of benchmark progress. First, a reported score should be compared with the information available under the benchmark’s own D, w, M, and temporal placement, rather than with the unattainable value R2=1R^2=1. Second, total cross-sectional performance and state-tracking performance should be reported separately whenever the scientific claim concerns change. Subject-wise centering, repeated labels, or a calibrated latent trait can define the state target Δg,T _g,T; without such a decomposition, a model may rank individuals well while failing to monitor them. Third, acquisition and architecture should be treated as distinct experimental axes. Increasing training-set size estimates fπ∗f_π^* more accurately but does not alter ℐπI_π; changing the observation protocol alters the ceiling itself. A useful ablation therefore holds the raw measurement budget fixed while reallocating it between same-time replication and temporal coverage, as in Eqs. (28)–(29). The calibration burden is asymmetric. The trait-channel ceiling in Eq. (19) uses only α and σε2 _ ^2, available from ordinary test–retest data under the model, and can therefore be reported now for many existing benchmarks. Only state-channel claims require short-lag estimation of ρ and its higher-order integrals; for that purpose, a small densely sampled subset can be more informative than another large cross-sectional sample collected under the same snapshot protocol. Experiments and Benchmark Implications We test analytic predictions rather than compare architectures. OU paths use exact transitions, labels are computed on a fine grid, and predictors are exact posterior expectations under the discretized process. The main table uses 50 independent repetitions of 2,000 objects per scenario; the equal-budget experiment uses 30 repetitions of 1,500 objects. Monte Carlo means are reported with 95% half-widths. (a) Equal raw-segment budget: precision versus coverage. (b) Snapshot ceilings over label horizons and trait shares. Figure 1: Protocol ceilings, not model training curves. In (a), exact state-channel ceilings and Monte Carlo estimates show that same-time replication rapidly saturates while dispersed occasions continue to add temporal information. In (b), stable traits create nonzero long-horizon snapshot ceilings; dotted lines are the T→∞T→∞ limits. Equal segment budget: D versus M. Figure 1a isolates the state channel (α=0α=0) with an occupation label, T/τ=20T/τ=20, and unit noise per raw segment. For each N=DMN=DM, the same-time protocol uses D=1,M=ND=1,M=N at the midpoint, whereas the coverage protocol uses D=N,M=1D=N,M=1 at evenly spaced times. For c=0c=0, the exact ceiling is computed from Eq. (7) with Cg(r)=arcsin(r)/(2π)C_g(r)= (r)/(2π). Same-time replication reaches only 0.0970.097 at N=64N=64; dispersed occasions reach 0.8080.808 using the same 64 raw segments. Benchmark ceilings from the trait channel. Figure 1b verifies Eq. (18): when α=0α=0, one-snapshot explainability vanishes as the label horizon grows, whereas nonzero trait shares converge to positive plateaus. For a zero-threshold occupation label with T/τ=14T/τ=14 and point-noise variance 0.20.2, Eq. (7) gives ceilings ℐ=0.119I=0.119, 0.2560.256, and 0.3550.355 for α=0α=0, 0.200.20, and 0.350.35. A benchmark stalled near R2=0.25R^2=0.25 may therefore be close to its acquisition ceiling under a moderate trait channel. These values are model-based illustrations, not estimates for a particular dataset. Task dependence and boundary localization. Figure 2a plots ℓocc/ℓmean _ occ/ _ mean, making visible that OU and Matérn-3/23/2 kernels matched to the same τ1 _1 assign different relative value to the same window. Figure 2b separates magnitude from efficiency: AaA_a is sharply concentrated near a=0a=0, whereas ℓa(w) _a(w) decays much more slowly because the universal ϕ(a)2φ(a)^2 factor cancels. (a) Occupation-to-mean effective-span ratio. (b) State variance localizes more strongly than efficiency. Figure 2: The label defines temporal information. Both kernels in panel (a) have τ1=1 _1=1; panel (b) uses an OU window with w/τ1=0.5w/ _1=0.5. α c T/τT/τ Var(Θ)Var( ) ℐI Theory MC ± 95% half-width Theory MC ± 95% half-width 0.00 0 10 0.0314 0.0314 ± 0.0003 0.169 0.168 ± 0.0016 0.00 0 40 0.0085 0.0085 ± 0.0001 0.040 0.040 ± 0.0003 0.35 0 10 0.0800 0.0800 ± 0.0005 0.385 0.385 ± 0.0024 0.35 0 40 0.0631 0.0629 ± 0.0003 0.309 0.310 ± 0.0026 0.00 1 10 0.0143 0.0143 ± 0.0001 0.150 0.149 ± 0.0022 0.00 1 40 0.0038 0.0038 ± 0.0000 0.035 0.036 ± 0.0005 0.35 1 10 0.0363 0.0364 ± 0.0005 0.347 0.348 ± 0.0044 0.35 1 40 0.0275 0.0276 ± 0.0004 0.278 0.275 ± 0.0040 Table 1: Exact formulas versus Monte Carlo means. Each cell uses 50 repetitions of 2,000 objects; intervals are 95% Monte Carlo half-widths. Limitations and Conclusion The two channels have different calibration requirements. The trait-channel ceiling does not require a densely sampled calibration subset: under the normalized model, any suitable repeated-measurement or two-occasion test–retest dataset can estimate α and σε2 _ ^2 and directly evaluate Eq. (19). Purely cross-sectional data cannot identify that decomposition, but repeated measurements suffice without resolving the short-lag kernel. The state channel is more demanding. Mean-label state variation depends on τ1 _1, whereas occupation labels use τk≥2 _k≥ 2 and the full window profile. Sparse windows separated far beyond the correlation scale cannot identify this short-lag structure, so state-channel analysis requires a densely sampled subset, external short-lag longitudinal data, or a justified parametric kernel family. Figure 2a makes the issue visible: equal τ1 _1 does not imply equal occupation-time information. The occupation model is an abstraction of frequency-type labels, not a claim that every observed score is generated by one thresholded Gaussian state. The state-effective span is conditional on the trait; an average span defined as a ratio of expected explained and total state variance is weighted toward individuals with larger state variance and is not the span of a typical individual. The central conclusion is not that segments are useless. Segments improve denoising and can strengthen the trait channel. They are not additional time. Long-horizon benchmark performance combines an O(1)O(1) trait channel with a task-dependent state channel controlled by the aggregation functional, correlation structure, and temporal support. Consequently, a performance ceiling that appears architectural may instead be imposed by acquisition, and high cross-sectional accuracy need not imply that a model has learned temporal state dynamics. References R. J. Adler and J. E. Taylor (2007) Random fields and geometry. Springer, New York. External Links: Document Cited by: Related Work. R. Altmeyer and J. Chorowski (2018) Estimation error for occupation time functionals of stationary markov processes. Stochastic Processes and their Applications 128 (6), p. 1830–1848. External Links: Document Cited by: Related Work. R. Altmeyer (2021) Approximation of occupation time functionals. Bernoulli 27 (4), p. 2714–2739. External Links: Document Cited by: Related Work. D. Azzimonti, J. Bect, C. Chevalier, and D. Ginsbourger (2016) Quantifying uncertainties on excursion sets under a gaussian random field prior. SIAM/ASA Journal on Uncertainty Quantification 4 (1), p. 850–874. External Links: Document Cited by: Related Work. J. Bect, F. Bachoc, and D. Ginsbourger (2019) A supermartingale approach to gaussian process based sequential design of experiments. Bernoulli 25 (4A), p. 2883–2919. External Links: Document Cited by: Related Work. J. Bect, D. Ginsbourger, L. Li, V. Picheny, and E. Vazquez (2012) Sequential design of computer experiments for the estimation of a probability of failure. Statistics and Computing 22 (3), p. 773–793. External Links: Document Cited by: Related Work. R. L. Brennan (2001) Generalizability theory. Springer, New York. External Links: Document Cited by: Related Work. P. Breuer and P. Major (1983) Central limit theorems for non-linear functionals of gaussian fields. Journal of Multivariate Analysis 13 (3), p. 425–441. External Links: Document Cited by: Related Work. R. J. Carroll, D. Ruppert, L. A. Stefanski, and C. M. Crainiceanu (2006) Measurement error in nonlinear models: a modern perspective. 2 edition, Chapman and Hall/CRC, Boca Raton, FL. External Links: Document Cited by: Related Work. K. Chaloner and I. Verdinelli (1995) Bayesian experimental design: a review. Statistical Science 10 (3), p. 273–304. External Links: Document Cited by: Related Work. W. G. Cochran (1977) Sampling techniques. 3 edition, Wiley, New York. Cited by: Related Work. L. J. Cronbach, G. C. Gleser, H. Nanda, and N. Rajaratnam (1972) The dependability of behavioral measurements: theory of generalizability for scores and profiles. Wiley, New York. Cited by: Related Work. T. G. Dietterich, R. H. Lathrop, and T. Lozano-Pérez (1997) Solving the multiple instance problem with axis-parallel rectangles. Artificial Intelligence 89 (1–2), p. 31–71. External Links: Document Cited by: Related Work. P. J. Diggle, P. Heagerty, K. Liang, and S. L. Zeger (2002) Analysis of longitudinal data. 2 edition, Oxford University Press, Oxford. External Links: Document Cited by: Related Work. W. A. Fuller (1987) Measurement error models. Wiley, New York. External Links: Document Cited by: Related Work. S. H. Hurlbert (1984) Pseudoreplication and the design of ecological field experiments. Ecological Monographs 54 (2), p. 187–211. External Links: Document Cited by: Related Work. M. Ilse, J. Tomczak, and M. Welling (2018) Attention-based deep multiple instance learning. In Proceedings of the 35th International Conference on Machine Learning, p. 2127–2136. External Links: Link Cited by: Related Work. H. Ji and H. Müller (2017) Optimal designs for longitudinal and functional data. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 79 (3), p. 859–876. External Links: Document Cited by: Related Work. L. Kish (1965) Survey sampling. Wiley, New York. Cited by: Related Work. M. F. Kratz (2006) Level crossings and other level functionals of stationary gaussian processes. Probability Surveys 3, p. 230–288. External Links: Document Cited by: Related Work. R. L. Plackett (1954) A reduction formula for normal multivariate integrals. Biometrika 41 (3/4), p. 351–360. External Links: Document Cited by: Related Work. B. J. Pyper and R. M. Peterman (1998) Comparison of methods to account for autocorrelation in correlation analyses of fish data. Canadian Journal of Fisheries and Aquatic Sciences 55 (9), p. 2127–2140. External Links: Document Cited by: Related Work. N. Quadrianto, A. J. Smola, T. S. Caetano, and Q. V. Le (2009) Estimating labels from label proportions. Journal of Machine Learning Research 10, p. 2349–2374. External Links: Link Cited by: Related Work. J. O. Ramsay and B. W. Silverman (2005) Functional data analysis. 2 edition, Springer, New York. External Links: Document Cited by: Related Work. G. K. Robinson (1991) That BLUP is a good thing: the estimation of random effects. Statistical Science 6 (1), p. 15–32. External Links: Document Cited by: Related Work. S. R. Searle, G. Casella, and C. E. McCulloch (1992) Variance components. Wiley, New York. External Links: Document Cited by: Related Work. R. J. Shavelson and N. M. Webb (1991) Generalizability theory: a primer. Sage, Newbury Park, CA. Cited by: Related Work. R. Steyer, M. Schmitt, and M. Eid (1999) Latent state–trait theory and research in personality and individual differences. European Journal of Personality 13 (5), p. 389–408. External Links: Document Cited by: Related Work. E. Vazquez and M. Piera-Martinez (2006) Estimation of the volume of an excursion set of a gaussian process using intrinsic kriging. arXiv preprint math/0611273. External Links: Link Cited by: Related Work. Z. Zhou (2018) A brief introduction to weakly supervised learning. National Science Review 5 (1), p. 44–53. External Links: Document Cited by: Related Work. Technical Supplement The Label Defines the Timescale: Trait–State Limits of Temporal-Aggregate Learning Xizhe Zhang ORCID: 0000-0002-8684-4591 • zhangxizhe@gmail.com S1 Model, Notation, and Regularity Conditions For each independent object, the standardized latent process is Z(t)=αM+1−αX(t),t∈[0,T],Z(t)= αM+ 1-αX(t), t∈[0,T], (S1) where M∼(0,1)M (0,1), X is a zero-mean unit-variance stationary Gaussian process, and M⟂XM X. The state correlation is ρ(u)=CovX(t),X(t+u)ρ(u)=Cov\X(t),X(t+u)\, and the total-process correlation is rα(u)=α+(1−α)ρ(u).r_α(u)=α+(1-α)ρ(u). (S2) For g∈L2(ϕ)g∈ L^2(φ), where ϕφ is the standard-normal density, define Θg,T=1T∫0TgZ(t)dt. _g,T= 1T _0^Tg\Z(t)\\,dt. (S3) The observation protocol is a finite-dimensional linear Gaussian measurement Yπ=LπZ+ε,ε∼(0,Rπ),Y_π=L_πZ+ , (0,R_π), (S4) independent of Z. Window averages are a special case. We use the following sufficient conditions. They are stronger than necessary but make every interchange explicit. Assumption S1 (Summability for Main Theorem 1). The function g is square integrable under the standard-normal law. Whenever a long-horizon expansion is invoked, Δg(u)=Cgα+(1−α)ρ(u)−Cg(α) _g(u)=C_g\α+(1-α)ρ(u)\-C_g(α) (S5) is absolutely integrable on [0,∞)[0,∞). For an O(T−2)O(T^-2) remainder, we additionally assume ∫0∞u|Δg(u)|du<∞ _0^∞u| _g(u)|\,du<∞. Assumption S2 (Effective-span summability for Main Theorem 2). For the conditional state label under consideration, ∑k≥1bkτk<∞,∑k≥1bkJk(w)2<∞, _k≥ 1b_k _k<∞, _k≥ 1b_kJ_k(w)^2<∞, (S6) where bkb_k are the squared Hermite coefficients, τk=∫0∞ρ(u)kdu _k= _0^∞ρ(u)^k\,du, and Jk(w)=∫ℝrw(t)kdtJ_k(w)= _Rr_w(t)^k\,dt. The paper focuses on nonnegative correlations to state the boundary theorem cleanly. The exact risk identity itself permits negative correlations whenever CgC_g is evaluated on the corresponding interval. S2 Gaussian and Hermite Preliminaries Let HkH_k denote the probabilists’ Hermite polynomials, normalized by Hj(U)Hk(U)=k! 1j=k,U∼(0,1).E\H_j(U)H_k(U)\=k!\,1\j=k\, U (0,1). (S7) For centered g∈L2(ϕ)g∈ L^2(φ), g(z)−g(U)=∑k≥1ak(g)k!Hk(z),ak(g)=[(g(U)−g(U))Hk(U)].g(z)-Eg(U)= _k≥ 1 a_k(g)k!H_k(z), a_k(g)=E[(g(U)-Eg(U))H_k(U)]. (S8) If (U,Vr)(U,V_r) are standard bivariate normal with correlation r, Mehler’s identity gives Cg(r)=Covg(U),g(Vr)=∑k≥1ak(g)2k!rk.C_g(r)=Cov\g(U),g(V_r)\= _k≥ 1 a_k(g)^2k!r^k. (S9) For the threshold function ga(z)=z>ag_a(z)=1\z>a\, integration by parts yields, for k≥1k≥ 1, ak(a) a_k(a) =∫a∞Hk(z)ϕ(z)dz = _a^∞H_k(z)φ(z)\,dz (S10) =ϕ(a)Hk−1(a), =φ(a)H_k-1(a), (S11) because Hk(z)ϕ(z)=−Hk−1(z)ϕ(z)′H_k(z)φ(z)=-\H_k-1(z)φ(z)\ . Thus bk(a):=ak(a)2k!=ϕ(a)2Hk−1(a)2k!.b_k(a):= a_k(a)^2k!=φ(a)^2 H_k-1(a)^2k!. (S12) A second identity is used for the trait component. If M,XM,X are independent standard normals, then [HkαM+1−αX∣M]=αk/2Hk(M).E [H_k\ αM+ 1-αX\ M ]=α^k/2H_k(M). (S13) It follows either from the generating function of HkH_k or from the Gaussian Ornstein–Uhlenbeck semigroup. S3 Posterior Predictor Used in Validation The main paper proves the exact protocol-risk identity. Here we record the computational form used to verify it. Let kπ(t)=CovZ(t),Yπ,qπ(s,t)=kπ(s)⊤Var(Yπ)−1kπ(t).k_π(t)=Cov\Z(t),Y_π\, q_π(s,t)=k_π(s) Var(Y_π)^-1k_π(t). For gc(z)=z>cg_c(z)=1\z>c\, Gaussian conditioning gives Z(t)∣Yπ∼mπ(t),1−qπ(t,t),mπ(t)=kπ(t)⊤Var(Yπ)−1Yπ.Z(t) Y_π \m_π(t),1-q_π(t,t)\, m_π(t)=k_π(t) Var(Y_π)^-1Y_π. (S14) The corresponding exact risk is Rπ∗=1T2∫0T∫0T[Cgrα(s−t)−Cgqπ(s,t)]dsdt.R_π^*= 1T^2 _0^T\! _0^T [C_g\r_α(s-t)\-C_g\q_π(s,t)\ ]\,ds\,dt. (S15) Hence the Bayes predictor is Θ^c,T∗=1T∫0TΦ¯(c−mπ(t)1−qπ(t,t))dt. _c,T^\,*= 1T _0^T \! ( c-m_π(t) 1-q_π(t,t) )\,dt. (S16) At c=0c=0, Cg(r)=arcsin(r)/(2π)C_g(r)= (r)/(2π). S4 Fixed-Support Refinement The following lemma records the exact sense in which more segments are useful but do not become additional time. Lemma S1 (Fixed-support refinement supporting the main-paper discussion). Let ℱMF_M be an increasing sequence of sigma-fields generated by progressively finer observations on a fixed temporal set W, and suppose ℱM↑ℱWF_M _W. For every Θ∈L2 ∈ L^2, Var(Θ∣ℱM)↓Var(Θ∣ℱW).EVar( _M) ( _W). (S17) Proof. The martingale (Θ∣ℱM)E( _M) converges to (Θ∣ℱW)E( _W) in L2L^2 by the martingale convergence theorem. Since Var(Θ∣ℱM)=(Θ2)−[(Θ∣ℱM)2],EVar( _M)=E( ^2)-E[E( _M)^2], (S18) the result follows. Monotonicity is the projection property of conditional expectation. ∎ The limiting risk can be zero in exceptional analytically determined processes; the paper’s statement explicitly concerns processes for which unobserved temporal support retains innovations relevant to the target. Additional segments can reduce measurement noise, including noise that obscures the trait, but cannot change W. S5 Verification of Main Theorem 1: Trait–State Decomposition Theorem S1 (Trait–state decomposition; corresponds to Main Theorem 1). Assume 0≤ρ≤10≤ρ≤ 1 and Assumption S1. Then Var(Θg,T)=Cg(α)+Astate(g)T+o(T−1),Var( _g,T)=C_g(α)+ A_ state(g)T+o(T^-1), (S19) where Astate(g)=2∫0∞[Cgα+(1−α)ρ(u)−Cg(α)]du,A_ state(g)=2 _0^∞ [C_g\α+(1-α)ρ(u)\-C_g(α) ]\,du, (S20) equivalently given by Eq. (S25), and Cg(α)=Var[g(Z(t))∣M]C_g(α)=Var[E\g(Z(t)) M\]. Proof. By stationarity, Var(Θg,T)=2T2∫0T(T−u)Cgrα(u)du.Var( _g,T)= 2T^2 _0^T(T-u)C_g\r_α(u)\\,du. (S21) Because 2T−2∫0T(T−u)du=12T^-2 _0^T(T-u)\,du=1, Eq. (S21) can be written exactly as Var(Θg,T)=Cg(α)+2T2∫0T(T−u)Δg(u)du,Var( _g,T)=C_g(α)+ 2T^2 _0^T(T-u) _g(u)\,du, (S22) where Δg(u)=Cgα+(1−α)ρ(u)−Cg(α) _g(u)=C_g\α+(1-α)ρ(u)\-C_g(α). The main-paper asymptotic follows directly from the exact decomposition: TVar(Θg,T)−Cg(α) T\Var( _g,T)-C_g(α)\ =2∫0T(1−uT)Δg(u)du. =2 _0^T (1- uT ) _g(u)\,du. (S23) Dominated convergence gives Astate(g)=2∫0∞Δg(u)duA_ state(g)=2 _0^∞ _g(u)\,du. If the first absolute moment of Δg _g is finite, adding and subtracting the infinite integral bounds the remainder by an O(T−2)O(T^-2) term plus T−1∫T∞|Δg(u)|duT^-1 _T^∞| _g(u)|\,du. This finite-T bound is the appendix-level verification not needed for the leading statement in the main paper. Insert Eq. (S9) and expand binomially: Astate(g) A_ state(g) =2∑k≥1ak(g)2k!∫0∞[α+(1−α)ρ(u)k−αk]du =2 _k≥ 1 a_k(g)^2k! _0^∞ [\α+(1-α)ρ(u)\^k-α^k ]\,du (S24) =2∑k≥1ak(g)2k!∑j=1k(kj)αk−j(1−α)jτj. =2 _k≥ 1 a_k(g)^2k! _j=1^k kjα^k-j(1-α)^j _j. (S25) Tonelli’s theorem applies directly when 0≤ρ≤10≤ρ≤ 1. For completeness, applying Eq. (S13) term-by-term to Eq. (S8) verifies the trait identity Var[g(Z(t))∣M]=∑k≥1ak(g)2k!αk=Cg(α).Var [E\g(Z(t)) M\ ]= _k≥ 1 a_k(g)^2k!α^k=C_g(α). (S26) ∎ S5.1 One-Snapshot Explainability Let Y=Z(T/2)+εY=Z(T/2)+ with ε∼(0,ν2) (0,ν^2). The protocol-explained covariance is qY(s,t)=rα(|s−T/2|)rα(|t−T/2|)1+ν2.q_Y(s,t)= r_α(|s-T/2|)r_α(|t-T/2|)1+ν^2. (S27) If α>0α>0, Cesàro convergence and the exact-risk identity give limT→∞Var(Θg,T∣Y)=Cg(α21+ν2), _T→∞Var\E( _g,T Y)\=C_g ( α^21+ν^2 ), (S28) and therefore limT→∞ℐY=Cgα2/(1+ν2)Cg(α). _T→∞I_Y= C_g\α^2/(1+ν^2)\C_g(α). (S29) For the mean and occupation labels, if α=0α=0 and ρ is integrable, the explained variance is O(T−2)O(T^-2) while the label variance is O(T−1)O(T^-1), yielding ℐY=O(T−1)I_Y=O(T^-1). Corollary S1 (Trait-channel value of occasions and within-occasion segments; corresponds to Main Corollary 1). Consider the mean label as T→∞T→∞. Suppose D occasions are separated enough that their state terms are independent, and each occasion averages M independent raw segments with measurement-noise variance σε2 _ ^2 per segment. Then the protocol explainability for the limiting trait target αMi αM_i is ℐtrait(D,M)=α+(1−α)/D+σε2/(DM).I_ trait(D,M)= α+(1-α)/D+ _ ^2/(DM). (S30) Proof. The average across all observations can be written Y¯=αMi+1−αX¯D+ε¯, Y= αM_i+ 1-α\, X_D+ , (S31) with independent components and variances α, (1−α)/D(1-α)/D, and σε2/(DM) _ ^2/(DM). The Gaussian regression coefficient of determination for predicting αMi αM_i from Y¯ Y is Cov(αMi,Y¯)2/[Var(αMi)Var(Y¯)]Cov( αM_i, Y)^2/[Var( αM_i)Var( Y)], which simplifies to Eq. (S30). ∎ Same-time replication has D=1D=1: as M→∞M→∞ it removes measurement noise but leaves transient state variance. Increasing temporally separated occasions additionally averages the state component, so the two replication axes are statistically distinct even for the trait channel. After standardization, a two-occasion test–retest design at a lag with negligible transient correlation identifies α from cross-occasion covariance and σε2 _ ^2 from observed variance (or within-occasion segments). Thus this trait ceiling is available from ordinary repeated-measurement data; it does not require estimation of ρ or τk≥2 _k≥ 2. S6 Verification of Main Theorem 2: Effective Temporal Span Condition on M=mM=m. For the occupation label, the trait-conditioned threshold for the unit state process is a=c−αm1−α.a= c- αm 1-α. (S32) Let WwW_w be a standardized noisy average of X over a fixed window of length w centered at t0=T/2t_0=T/2, and write rw(u)=CorrX(t0+u),Ww,Jk(w)=∫ℝrw(u)kdu.r_w(u)=Corr\X(t_0+u),W_w\, J_k(w)= _Rr_w(u)^k\,du. (S33) Theorem S2 (Task-dependent state-effective span; corresponds to Main Theorem 2). Assume the window remains interior as T→∞T→∞, rwk∈L1(ℝ)r_w^k∈ L^1(R) for every Hermite order carrying nonzero weight, and the effective-span summability assumption holds. Then: 1. For the mean state label, ℐstate(w)=ℓmean(w)T+o(T−1),ℓmean(w)=2τ1η(w)+ν2,I_ state(w)= _ mean(w)T+o(T^-1), _ mean(w)= 2 _1η(w)+ν^2, (S34) where η(w)=2w2∫0w(w−u)ρ(u)du.η(w)= 2w^2 _0^w(w-u)ρ(u)\,du. (S35) 2. For the occupation state label Θa,T=T−1∫0TX(t)>adt _a,T=T^-1 _0^T1\X(t)>a\\,dt, ℐstate(w∣a)=ℓa(w)T+o(T−1),I_ state(w a)= _a(w)T+o(T^-1), (S36) with ℓa(w)=∑k≥1Hk−1(a)2k!Jk(w)22∑k≥1Hk−1(a)2k!τk. _a(w)= _k≥ 1 H_k-1(a)^2k!J_k(w)^2 2 _k≥ 1 H_k-1(a)^2k! _k. (S37) Proof. Mean label. For a unit state process, η(w)=Var1w∫−w/2w/2X(u)du=2w2∫0w(w−u)ρ(u)du.η(w)=Var \ 1w _-w/2^w/2X(u)\,du \= 2w^2 _0^w(w-u)ρ(u)\,du. (S38) If additive window noise has variance ν2ν^2, stationarity and Fubini’s theorem give J1(w)=∫ℝrw(u)du=2τ1η(w)+ν2.J_1(w)= _Rr_w(u)\,du= 2 _1 η(w)+ν^2. (S39) The long-horizon state-label variance is 2τ1/T+o(T−1)2 _1/T+o(T^-1). The variance explained by WwW_w equals 1T2∫0Trw(t−t0)dt2. 1T^2 \ _0^Tr_w(t-t_0)\,dt \^2. (S40) Because t0=T/2t_0=T/2 and rw∈L1(ℝ)r_w∈ L^1(R), ∫0Trw(t−t0)dt=∫−T/2T/2rw(u)du=J1(w)+o(1). _0^Tr_w(t-t_0)\,dt= _-T/2^T/2r_w(u)\,du=J_1(w)+o(1). (S41) Dividing the explained variance by the state-label variance and using Eq. (S39) proves Eq. (S34). Occupation label. Using Eq. (S11), Var(Θa,T)=AaT+o(T−1),Aa=2∑k≥1bk(a)τk.Var( _a,T)= A_aT+o(T^-1), A_a=2 _k≥ 1b_k(a) _k. (S42) For jointly normal X(t)X(t) and WwW_w, Hk(X(t))∣Ww=rw(t−t0)kHk(Ww).E\H_k(X(t)) W_w\=r_w(t-t_0)^kH_k(W_w). (S43) Orthogonality therefore gives Var(Θa,T∣Ww)=1T2∑k≥1bk(a)∫0Trw(t−t0)kdt2.Var\E( _a,T W_w)\= 1T^2 _k≥ 1b_k(a) \ _0^Tr_w(t-t_0)^k\,dt \^2. (S44) For each fixed k, absolute integrability and the interior placement imply ∫0Trw(t−t0)kdt=∫−T/2T/2rw(u)kdu=Jk(w)+o(1). _0^Tr_w(t-t_0)^k\,dt= _-T/2^T/2r_w(u)^k\,du=J_k(w)+o(1). (S45) The summability assumption permits passage of this limit through the Hermite series, yielding Var(Θa,T∣Ww)=Ba(w)T2+o(T−2),Ba(w)=∑k≥1bk(a)Jk(w)2.Var\E( _a,T W_w)\= B_a(w)T^2+o(T^-2), B_a(w)= _k≥ 1b_k(a)J_k(w)^2. (S46) Dividing by Eq. (S42) proves Eq. (S36). Substituting bk(a)=ϕ(a)2Hk−1(a)2/k!b_k(a)=φ(a)^2H_k-1(a)^2/k! yields Eq. (S37); the factor ϕ(a)2φ(a)^2 cancels exactly. ∎ For ρ(u)=e−u/τρ(u)=e^-u/τ and ν2=0ν^2=0, η(w)=2τw−τ(1−e−w/τ)w2,ℓmean(w)=w2w−τ(1−e−w/τ).η(w)= 2τ\w-τ(1-e^-w/τ)\w^2, _ mean(w)= w^2w-τ(1-e^-w/τ). (S47) Taylor expansion gives ℓmean(w)=2τ+(2/3)w+O(w2/τ) _ mean(w)=2τ+(2/3)w+O(w^2/τ) as w/τ↓0w/τ 0, and ℓmean(w)=w+τ+O(τ2/w) _ mean(w)=w+τ+O(τ^2/w) as w/τ→∞w/τ→∞. Corollary S2 (Equal raw-segment budget in the state channel; corresponds to Main Corollary 2). Write ℓg(w;ν2,m) _g(w;ν^2,m) for the trait-conditioned effective span when an averaged window has noise variance ν2ν^2. Allocate N independent raw measurements either to one fixed window, averaged to noise variance σε2/N _ ^2/N, or to N mutually separated windows, each with noise variance σε2 _ ^2. Under the sparse long-horizon conditions of the effective-span theorem above, ℐsame(N∣m) _ same(N m) =ℓg(w;σε2/N,m)T+o(T−1)⟶ℓg(w;0,m)T, = _g(w; _ ^2/N,m)T+o(T^-1) _g(w;0,m)T, (S48) ℐdispersed(N∣m) _ dispersed(N m) =Nℓg(w;σε2,m)T+o(N/T), = N _g(w; _ ^2,m)T+o(N/T), (S49) until the first-order additivity approximation approaches saturation. Proof. For the same-time allocation, averaging N conditionally independent measurements changes only the window-noise variance from σε2 _ ^2 to σε2/N _ ^2/N; Theorem S2 then gives Eq. (S48), and continuity of the posterior projection in the noise variance gives the limit. For mutually separated windows, the off-diagonal window covariances and the cross terms in the explained state variance are negligible in the sparse regime. Each window contributes Bg(w,m)/T2+o(T−2)B_g(w,m)/T^2+o(T^-2), while the state-label variance is Ag(m)/T+o(T−1)A_g(m)/T+o(T^-1). Summing the N contributions gives Eq. (S49). ∎ For a mean label and point-like windows these expressions reduce to 2τ1T(1+σε2/N)and2Nτ1T(1+σε2), 2 _1T(1+ _ ^2/N) 2N _1T(1+ _ ^2), (S50) respectively. Thus increasing M purchases precision on fixed support, whereas increasing D can purchase additional state support. S7 Verification of Main Theorem 3: Boundary Localization Define Ga(r)=Cov(U>a),(Vr>a).G_a(r)=Cov\1(U>a),1(V_r>a)\. (S51) Plackett’s identity gives ∂rPr(U>a,Vr>a)=exp−a2/(1+r)2π1−r2. ∂ r (U>a,V_r>a)= \-a^2/(1+r)\2π 1-r^2. (S52) Since Ga(0)=0G_a(0)=0, Ga(r)=∫0rexp−a2/(1+s)2π1−s2ds.G_a(r)= _0^r \-a^2/(1+s)\2π 1-s^2\,ds. (S53) Theorem S3 (Boundary localization; corresponds to Main Theorem 3). Assume ρ(u)≥0ρ(u)≥ 0. Then Aa=2∫0∞Gaρ(u)duA_a=2 _0^∞G_a\ρ(u)\\,du is even and nonincreasing in |a||a|. It is strictly decreasing in |a||a| whenever ρ is positive on a set of nonzero measure, and hence is maximized at a=0a=0. Proof. For each fixed s∈[0,1)s∈[0,1), the integrand in Eq. (S53) is even in a and strictly decreases with |a||a|. Integrating first in s and then in u gives the claimed properties of AaA_a; strictness holds whenever ρ is positive on a set of nonzero measure. ∎ Equation (S37) is also even because Hk(−a)2=Hk(a)2H_k(-a)^2=H_k(a)^2. Under locally uniform convergence it is differentiable and ℓa′(w)|a=0=0 _a (w)|_a=0=0. No general monotonicity is claimed for ℓa(w) _a(w): the remaining threshold dependence is through the relative reweighting of Hermite orders. The paper’s numerical comparison shows that it is substantially flatter than AaA_a for the investigated OU and Matérn windows. S8 Ornstein–Uhlenbeck Worked Example For ρ(u)=e−u/τρ(u)=e^-u/τ, τk=∫0∞e−ku/τdu=τk. _k= _0^∞e^-ku/τ\,du= τk. (S54) Using r=e−u/τr=e^-u/τ, Aa=2τ∫01Ga(r)rdr=2τϕ(a)2∑k≥1Hk−1(a)2k!k.A_a=2τ _0^1 G_a(r)r\,dr=2τφ(a)^2 _k≥ 1 H_k-1(a)^2k!\,k. (S55) For a=0a=0, G0(r)=arcsin(r)2π.G_0(r)= (r)2π. (S56) Therefore A0 A_0 =τπ∫01arcsin(r)rdr = τπ _0^1 (r)r\,dr (S57) =τlog22. = τ 22. (S58) The last integral follows by the substitution r=sinxr= x and the standard integral ∫0π/2xcotxdx=(π/2)log2 _0^π/2x x\,dx=(π/2) 2. S9 Sparse Multi-Window Consequence Suppose D equal windows are mutually separated so that all cross-window correlations are o(1)o(1) in the asymptotic regime, and suppose Dℓg(w)/T=o(1)D _g(w)/T=o(1). Then the explained state variances add to first order: ℐstate(πD)=Dℓg(w)T+o(D/T).I_ state( _D)= D _g(w)T+o(D/T). (S59) Under per-window cost c0+c1wc_0+c_1w and total per-object budget B, this gives the leading efficiency criterion maxwℓg(w)c0+c1w. _w _g(w)c_0+c_1w. (S60) This is only a sparse-regime consequence. Non-sparse placement must use the full qπq_π in Eq. (S15) and is closely related to existing excursion-set sequential-design problems. S10 Simulation Protocol and Additional Results The trait–state verification uses 50 repetitions and 2,000 independent objects per scenario. The equal-segment-budget experiment uses 30 repetitions and 1,500 objects. OU paths are generated on grids by the exact transition X(t+Δt)=e−Δt/τX(t)+1−e−2Δt/τζ,ζ∼(0,1).X(t+ t)=e^- t/τX(t)+ 1-e^-2 t/τ\,ζ, ζ (0,1). (S61) The final experiments were run on a Mac Studio with an Apple M2 Ultra CPU (24 cores) and 192 GB RAM under macOS 26.5.2, using Python 3.14.4, NumPy 2.4.4, SciPy 1.17.0, pandas 3.0.0, and Matplotlib 3.10.8; no GPU was used. For each object, the continuous occupation proportion is approximated on the fine grid. A noisy point observation at T/2T/2 is generated with variance 0.20.2. The predictor is the exact posterior mean of the discretized occupation proportion, not a fitted regression model. S10.1 Reported Quantities For every scenario, the analysis computes: • exact finite-T label variance; • exact one-snapshot explainability; • exact Bayes MSE R∗=Var(Θ)(1−ℐ)R^*=Var( )(1-I); • Monte Carlo label variance, explained variance, explainability, and MSE; • Monte Carlo means and 95% half-widths across repetitions; • equal-budget D–M curves comparing repeated same-time segments with dispersed occasions. The theory curves for effective spans are evaluated by numerical quadrature of τk _k and Jk(w)J_k(w) with stable normalized-Hermite recurrences. The OU boundary coefficient is also evaluated from its independent Plackett integral, providing a cross-check of the Hermite implementation. S10.2 Equal-Segment-Budget Experiment The state-only experiment sets α=0α=0, c=0c=0, T/τ=20T/τ=20, and raw segment-noise variance one. For each total budget N∈1,2,4,8,16,32,64N∈\1,2,4,8,16,32,64\, the same-time protocol uses one midpoint occasion with averaged noise variance 1/N1/N, while the coverage protocol uses N evenly spaced occasions with variance one each. Exact explainability is computed from Eq. (S15) using the arcsine transform; Monte Carlo uses the exact discrete OU posterior. At N=64N=64, exact explainability is 0.09700.0970 for same-time replication and 0.80830.8083 for dispersed occasions. S10.3 Numerical Cross-Checks The final values in the main-paper table are checked by two independent calculations: direct evaluation of the exact formulas and Monte Carlo estimation under the discretized posterior. The displayed half-widths summarize variation across the independent repetitions described above. S11 Interpretive Boundaries Trait channel. Purely cross-sectional observations do not identify α, but ordinary test–retest data do under the model. The trait ceiling needs only α and σε2 _ ^2 and therefore does not share the state channel’s dense short-lag calibration requirement. State channel. Occupation-time state variance and effective span depend on τ2,τ3,… _2, _3,…, not only τ1 _1. Two kernels matched at τ1 _1 can therefore agree on the leading long-horizon mean-label coefficient while disagreeing on occupation-label information. State-channel analysis requires a dense short-lag calibration subset, external longitudinal data, or a defensible parametric family. Average effective span. If one defines a population state-effective span as mBg(w,m)/mAg(m)E_mB_g(w,m)/E_mA_g(m), this is a ratio of expectations, not an expectation of individual ratios. It is weighted toward individuals with larger state-driven label variance, which for occupation labels are those near the threshold. It should not be interpreted as the span of a typical object.