Paper deep dive
A 1/R Law for Kurtosis Contrast in Balanced Mixtures
Yuda Bi, Wenjun Xiao, Linhao Bai, Vince D Calhoun
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 11:03:22 AM
Summary
The paper establishes a sharp redundancy law for Kurtosis-based Independent Component Analysis (ICA), proving that population excess kurtosis decays as O(1/R_eff) in wide, balanced mixtures. It derives a necessary condition for detectability linking mixture width R, sample size T, and source kurtosis, and proposes a 'purification' method to restore contrast by selecting sign-consistent source subsets.
Entities (6)
Relation Signals (4)
Independent Component Analysis → suffersfrom → Kurtosis Contrast Decay
confidence 95% · Kurtosis-based Independent Component Analysis (ICA) weakens in wide, balanced mixtures.
Effective Width → determinesdecayrate → Excess Kurtosis
confidence 93% · the population excess kurtosis obeys |κ(y)|=O(κ_max/R_eff)
Purification → restores → Kurtosis Contrast
confidence 90% · purification ... restores R-independent contrast Ω(1/m)
Sample Size → influences → Estimation Error
confidence 85% · surpassing the O(1/√T) estimation scale requires R≲κ_max√T
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Kurtosis-based Independent Component Analysis (ICA) weakens in wide, balanced mixtures. We prove a sharp redundancy law: for a standardized projection with effective width $R_{\mathrm{eff}}$ (participation ratio), the population excess kurtosis obeys $|\kappa(y)|=O(\kappa_{\max}/R_{\mathrm{eff}})$, yielding the order-tight $O(c_b\kappa_{\max}/R)$ under balance (typically $c_b=O(\log R)$). As an impossibility screen, under standard finite-moment conditions for sample kurtosis estimation, surpassing the $O(1/\sqrt{T})$ estimation scale requires $R\lesssim \kappa_{\max}\sqrt{T}$. We also show that \emph{purification} -- selecting $m\!\ll\!R$ sign-consistent sources -- restores $R$-independent contrast $\Omega(1/m)$, with a simple data-driven heuristic. Synthetic experiments validate the predicted decay, the $\sqrt{T}$ crossover, and contrast recovery.
Tags
Links
- Source: https://arxiv.org/abs/2602.22334v1
- Canonical: https://arxiv.org/abs/2602.22334v1
Trouble viewing inline? Open PDF directly →
Full Text
30,171 characters extracted from source content.
Expand or collapse full text
A 1/R Law for Kurtosis Contrast in Balanced Mixtures Yuda Bi, Wenjun Xiao, Linhao Bai, Vince Calhoun Y. Bi is with the Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA 30303 USA (e-mail: ybi@gsu.edu).W. Xiao is with the Department of Computer Science, The George Washington University, Washington, DC 20052 USA.L. Bai is with the School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA.V. D. Calhoun is with the Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, and Emory University, Atlanta, GA 30303 USA, and also with the School of Electrical and Computer Engineering, Georgia Institute of Technology, Atlanta, GA 30332 USA (e-mail: vcalhoun@gsu.edu). Abstract Kurtosis-based Independent Component Analysis (ICA) weakens in wide, balanced mixtures. We prove a sharp redundancy law: for a standardized projection with effective width ReffR_eff (participation ratio), the population excess kurtosis obeys |κ(y)|=O(κmax/Reff)|κ(y)|=O( _ /R_eff), yielding the order-tight O(cbκmax/R)O(c_b _ /R) under balance (typically cb=O(logR)c_b=O( R)). As an impossibility screen, under standard finite-moment conditions for sample kurtosis estimation, surpassing the O(1/T)O(1/ T) estimation scale requires R≲κmaxTR _ T. We also show that purification—selecting m≪Rm\! \!R sign-consistent sources—restores R-independent contrast Ω(1/m) (1/m), with a simple data-driven heuristic. Synthetic experiments validate the predicted decay, the T T crossover, and contrast recovery. Index Terms: Independent Component Analysis, Kurtosis, Redundancy, Source Separation I Introduction Independent Component Analysis (ICA) recovers statistically independent latent sources from linear mixtures and is identifiable whenever at most one source is Gaussian [11]. This underpins Infomax [4], FastICA [15], JADE [7], and related methods [14, 9], with broad use in neuroimaging [5, 6] and telecommunications. Excess kurtosis—the standardized fourth cumulant—is a central contrast function [8], and kurtosis-type nonlinearities remain standard in FastICA. While kurtosis-based estimators are well studied [10, 20, 17], existing analyses treat contrast strength as given rather than quantifying how much contrast survives as mixtures widen. As effective width R grows, the CLT drives standardized projections toward Gaussianity [11], flattening kurtosis contrast—a familiar effect at high model orders that lacks a population-level scaling law. Herrmann and Theis [12] studied how finite-sample kurtosis estimation degrades as sources approach Gaussianity, but targeted estimation rather than population contrast. Here we prove a population-level impossibility: under balanced projections, true excess kurtosis decays as 1/Reff1/R_eff, so increasing T alone cannot avert contrast collapse. Large-dimensional ICA studies [18, 3, 21] address recovery error and convergence, complementing the population decay characterized here. This regime arises naturally in neuroimaging: in group ICA [5, 6, 1], higher model orders [19] increase active sources; group-level projections activate blocks of R subject-level components, and balance holds when no single component dominates. Proposition 1 shows this is generic for well-conditioned blocks, with cbc_b growing at most logarithmically. Our law thus explains a common pattern: growing model order raises ReffR_eff, shrinking population kurtosis as 1/Reff1/R_eff and yielding noisy, unreproducible components. Similar effects appear in large-scale multimodal fusion [22, 23]. This decay is distinct from the usual O(1/T)O(1/\! T) estimation error: even with unlimited data, a broad balanced mixture has vanishing kurtosis. The bound characterizes contrast available to generic projections; ICA must locate the rare unbalanced directions that isolate sources, and this search becomes harder as the landscape flattens. We provide an explicit scaling law, a computable model-order ceiling, and a principled mechanism—purification—that restores non-vanishing contrast. This letter contributes: (i) A population-level impossibility law (Theorem 1): for balanced R-term mixtures, excess kurtosis scales as O(1/R)O(1/R), and this is order-tight. (i) A computable model-order diagnostic (Corollary 2): a necessary (not sufficient) viability condition is R≲κmaxTR _ T, linking mixture width to sample size. (i) A purification lower bound (Theorem 2): selecting m≪Rm R sources with a common kurtosis sign—equivalently, a targeted reduction of ReffR_eff—restores an R-independent contrast of order Ω(1/m) (1/m). I Model and Main Results I-A Setup and Notation We observe xt∈ℝpx_t ^p generated from independent sources st∈ℝks_t ^k via xt=Ast+ηt,A∈ℝp×k,ηt=0,x_t=As_t+ _t, A ^p× k,\ \ E _t=0, (1) where ηt _t is independent noise with finite fourth moments. Sources are standardized: stj=0Es_tj=0, stj2=1Es_tj^2=1 (and for Corollary 2 we assume finite eighth moments to control the variance of the sample kurtosis estimator). For any unit-variance v, define excess kurtosis κ(v):=[v4]−3κ(v):=E[v^4]-3. We present the square case p=kp=k with invertible A; the rectangular case follows via pseudoinverse. For a direction u∈ℝpu ^p, let :=j:aj⊤u≠0S:=\j:\,a_j u≠ 0\ be the active set and R:=||R:=|S| the mixture width. Define normalized projection coefficients wj:=aj⊤u∑ℓ∈(aℓ⊤u)2,j∈,w_j:= a_j u _ (a_ u)^2, j , so that the standardized noiseless projection can be written as y=∑j∈wjsjy= _j w_js_j with ∑j∈wj2=1 _j w_j^2=1. Unless stated otherwise, wjw_j are defined for the signal part x=Asx=As. Let κmax:=maxj∈|κ(sj)| _ := _j |κ(s_j)|. We call the projection balanced if maxj|wj|2≤cb/R _j|w_j|^2≤ c_b/R for some cb>0c_b>0. When convenient, we relabel =1,…,RS=\1,…,R\ without loss of generality. Proposition 1 (Block balance). Fix a block index set ⊆1,…,kS \1,…,k\ with ||=R|S|=R. If, for a nonzero direction u, the inner products cj:=|aj⊤u|c_j:=|a_j u| satisfy maxj∈cj2≤ρR∑ℓ∈cℓ2 _j c_j^2\;≤\; ρR _ c_ ^2 (2) for some ρ≥1ρ≥ 1, then maxj∈|wj|2≤ρ/R _j |w_j|^2≤ρ/R. In particular, if A⊤A=IRA_S A_S=I_R and u is uniform on the unit sphere in span(A)span(A_S), then maxj|wj|2≤C(logR)/R _j|w_j|^2≤ C( R)/R with probability at least 1−2R−c1-2R^-c for absolute constants C,c>0C,c>0. Proof. By definition, |wj|2=cj2/∑ℓ∈cℓ2|w_j|^2=c_j^2/ _ c_ ^2; taking the maximum and applying (2) gives maxj|wj|2≤ρ/R _j|w_j|^2≤ρ/R. For the probabilistic statement, under A⊤A=IRA_S A_S=I_R the vector A⊤uA_S u is uniformly distributed on the unit sphere; the claim follows from Gaussian maxima and chi-square concentration applied to the coordinate maxima of a random unit vector in ℝRR^R (see, e.g., [24, Ch. 3]). ∎ Thus, within a well-conditioned block, a generic direction spreads its energy across R columns, producing balance with cbc_b growing at most logarithmically in R. I-B Redundancy Law Theorem 1 (Sharp Redundancy Bound). Let sjj=1R\s_j\_j=1^R be independent sources with unit variance and bounded fourth moments. For any standardized projection y:=u⊤xVar(u⊤x)y:= u x Var(u x) in the noiseless case x=Asx=As, |κ(y)|≤κmax∑j=1R|wj|4.|κ(y)|≤ _ _j=1^R|w_j|^4. Under balanced mixing (maxj|wj|2≤cb/R _j|w_j|^2≤ c_b/R): |κ(y)|≤cbκmaxR.|κ(y)|≤ c_b\, _ R. (3) Order-tightness. For equal weights |wj|2=1/R|w_j|^2=1/R and aligned kurtoses κ(sj)=κ0κ(s_j)= _0, one has |κ(y)|=|κ0|/R|κ(y)|=| _0|/R exactly. Proof. By independence and unit variance, Var(u⊤As)=∑j=1R(aj⊤u)2Var(u As)= _j=1^R(a_j u)^2, so y=∑j=1Rwjsjy= _j=1^Rw_js_j with ∑j=1Rwj2=1 _j=1^Rw_j^2=1. Cross-cumulants vanish, giving κ(y)=∑j=1Rwj4κ(sj)κ(y)= _j=1^Rw_j^4κ(s_j) and |κ(y)|≤κmax∑j=1R|wj|4|κ(y)|≤ _ _j=1^R|w_j|^4. Under balance, ∑j|wj|4≤(maxj|wj|2)∑j|wj|2≤cb/R _j|w_j|^4≤( _j|w_j|^2) _j|w_j|^2≤ c_b/R. Sharpness: when |wj|2=1/R|w_j|^2=1/R and κ(sj)=κ0κ(s_j)= _0 for all j, we obtain |κ(y)|=|κ0|/R|κ(y)|=| _0|/R. ∎ Cancellation between positive and negative source kurtoses may further reduce |κ(y)||κ(y)| below this bound; the inequality concerns worst-case magnitude. Corollary 1 (Effective-width law). Define the effective mixture width Reff:=1/∑j∈wj4R_eff:=1/ _j w_j^4 (a participation-ratio measure of source dispersion). Then, for any projection, |κ(y)|≤κmax/Reff.|κ(y)|≤ _ /R_eff. (4) Balanced mixing implies Reff≥R/cbR_eff≥ R/c_b, recovering the O(1/R)O(1/R) rate of Theorem 1. Proof. Immediate from Theorem 1: κmax∑|wj|4=κmax/Reff _ Σ|w_j|^4= _ /R_eff. ∎ This formulation is fully general—it applies to any weight distribution without the balance assumption. For unbalanced weights, ReffR_eff can be much smaller than R, explaining the slower decay observed in Fig. 1(b). Remark (noise). If x=As+ηx=As+η with independent Gaussian noise, fourth-cumulant additivity gives κ(y)=SNR4κ(ys/σs)/(SNR2+1)2κ(y)=SNR^4\,κ(y_s/ _s)/(SNR^2+1)^2 where SNR2:=σs2/σn2SNR^2:= _s^2/ _n^2 is defined at the projected variance level for a fixed scalar projection; the signal kurtosis inherits the 1/R1/R decay under balance, attenuated further by the SNR factor. Figure 1: Empirical validation (mean ± SEM). (a) FastICA error err(W,A)err(W,A) increases with 1/Δκ1/ _κ (R2=0.78R^2=0.78). (b) Balanced Student-t mixtures (df=8=8, κ0=1.5 _0=1.5) show |κ^(y)|∝1/R| κ(y)| 1/R for R=2,…,50R=2,…,50 (R2=0.986R^2=0.986); unbalanced power-law weights decay more slowly. Inset at R=50R=50: std(κ^)∼σ0/Tstd( κ) _0/ T (σ0≈5.3 _0≈ 5.3); dotted line: |κ0|/R| _0|/R. (c) At R=50R=50 (df∈[6,30]∈[6,30]), purification (m=5m=5) increases contrast from ≈0.03≈ 0.03 to ≈0.43≈ 0.43 (oracle) / ≈0.44≈ 0.44 (sample-based), ∼14× 14×. Remark (consequence for FastICA). Under balance, Theorem 1 predicts attenuation of kurtosis contrast for standardized candidate projections, so the empirical separation Δκ:=minr≠ℓ|κ(yr)−κ(yℓ)| _κ:= _r≠ |κ(y_r)-κ(y_ )| between candidate components can become small. Since FastICA is known to be less well conditioned when kurtosis gaps are small [15, 14], this provides a plausible degradation mechanism (supported by Fig. 1a), rather than a standalone conditioning theorem. This does not imply that latent-source parameters κ(sj)κ(s_j) change; the shrinkage is mixing-induced attenuation of observable contrast, quantified by Theorem 1. We use σ0/T _0/ T as a conservative detectability scale (one-standard-deviation criterion); stronger guarantees would require algorithm-specific margins. Corollary 2 (Necessary model-order screening condition). Under the conditions of Theorem 1, assume the sample excess-kurtosis estimator κ κ from T i.i.d. observations satisfies std(κ^)≤σ0/Tstd( κ)≤ _0/ T for a moment-dependent constant σ0>0 _0>0 (e.g., under finite eighth moments; see [20]). A necessary condition for the kurtosis contrast of a balanced projection to exceed the estimation noise floor, |κ(y)|>σ0/T|κ(y)|> _0/ T, is that its active width R obeys R<cbκmaxTσ0.R\;<\; c_b\, _ \, T _0. (5) Equivalently, if an algorithm requires a minimum population contrast κ∗>0κ^*>0, then R≤cbκmax/κ∗R≤ c_b\, _ /κ^*. Proof. Theorem 1 gives |κ(y)|≤cbκmax/R|κ(y)|≤ c_b _ /R. Requiring cbκmax/R>σ0/Tc_b _ /R> _0/ T yields (5). Thus (5) is an impossibility-screening necessary condition (from an upper bound), not a sufficiency guarantee. ∎ The bound R<cbκmaxT/σ0R<c_b _ T/ _0 is a finite-sample ceiling, while R≤cbκmax/κ∗R≤ c_b _ /κ^* is a population ceiling; the operative constraint is the tighter of the two. The viable range expands only as T T (doubling tolerable R requires quadrupling T). Here R is projection-dependent and typically smaller than k; conservative screening sets R≈kR≈ k, while tighter screening uses the mean activation width. A practical proxy for ReffR_eff is the participation ratio 1/∑jwj41/\! _jw_j^4 of a component’s loading vector. Remark (computable screening). Given data X∈ℝp×TX ^p× T and candidate model order k: (1) whiten and project onto the leading k principal components; (2) compute sample kurtosis κ^i κ_i for each PC direction; (3) set κ^max=maxi|κ^i| κ_ = _i| κ_i| and estimate σ^0 σ_0 via bootstrap; (4) choose a conservative log-factor cbc_b consistent with Proposition 1 (e.g., cb=4logkc_b=4 k); and (5) compute kmax(screen)=cbκ^maxT/σ^0k_ ^(screen)=c_b κ_ T/ σ_0. If k≫kmax(screen)k k_ ^(screen), kurtosis contrast collapse is expected; reduce model order or apply purification. PC directions are a coarse proxy for ICA-active width and are intended to detect hopeless regimes, not to predict contrast precisely. I-C Purification via Sign-Consistent Subset Selection Theorem 1 implies contrast collapses when effective width is large; purification counteracts this by restricting to a sign-consistent subset of m sources and renormalizing, yielding effective width at most m and contrast lifted from O(1/Reff)O(1/R_eff) to Ω(1/m) (1/m), independent of R. By the pigeonhole principle, among R sources with non-zero kurtosis, at least R/2R/2 share the same kurtosis sign, so for any target size m≪Rm R a sign-consistent subset of size m can always be selected. Theorem 2 (Purification Lower Bound). Suppose sjj=1R\s_j\_j=1^R are independent with non-zero kurtosis. Let y=∑j=1Rwjsjy= _j=1^Rw_js_j with ∑j=1Rwj2=1 _j=1^Rw_j^2=1. Given a sign-consistent subset ℳ⊆1,…,RM \1,…,R\ with |ℳ|=m|M|=m and ∑j∈ℳwj2>0 _j w_j^2>0, define the purified mixture by restricting and re-normalizing the original weights: ypur=∑j∈ℳw~jsj,w~j:=wj∑ℓ∈ℳwℓ2,∑j∈ℳw~j2=1,y_pur= _j w_js_j, w_j:= w_j _ w_ ^2, _j w_j^2=1, which satisfies the R-independent lower bound |κ(ypur)|≥κmin,ℳm,|κ(y_pur)|≥ _ ,Mm, (6) where κmin,ℳ:=minj∈ℳ|κ(sj)| _ ,M:= _j |κ(s_j)|. Proof. By cumulant additivity, κ(ypur)=∑j∈ℳw~j4κ(sj)κ(y_pur)= _j w_j^4κ(s_j). Under sign-consistency, |κ(ypur)|≥κmin,ℳ∑j∈ℳw~j4|κ(y_pur)|≥ _ ,M _j w_j^4. By Cauchy–Schwarz, ∑j∈ℳw~j4≥(∑j∈ℳw~j2)2/m=1/m _j w_j^4≥( _j w_j^2)^2/m=1/m. ∎ Bound (6) is unconditional: it does not require |κ(y)|≤C/R|κ(y)|≤ C/R; Theorem 1 only motivates the regime where purification is most needed. Via Corollary 1, purification yields Reff(pur)≤mR_eff^(pur)≤ m, so (6) reads |κ(ypur)|≥κmin,ℳ/Reff(pur)|κ(y_pur)|≥ _ ,M/R_eff^(pur). Under additive Gaussian noise, the purified contrast inherits the same SNR attenuation factor as in the noise remark above, since Gaussian noise contributes zero fourth cumulant. Sign-consistency is essential: if selected sources have mixed kurtosis signs, cancellations can drive κ(ypur)κ(y_pur) toward zero. In many homogeneous families (e.g., sparse/speech vs. bounded/uniform-like sources), a dominant sign is plausible. A simple data-driven variant requires no oracle. Given candidate components s^j s_j from a preliminary separation (e.g., PCA + FastICA at moderate model order; even if components remain partially mixed due to contrast degradation, the additive nature of higher-order cumulants robustly ensures that dominant sources dictate the aggregate kurtosis signs), purification selects a subset: 1. Compute sample kurtoses κ^j κ_j of s^1,…,s^R s_1,…, s_R. 2. Set sign∗:=sign(∑jκ^j)sign^*:=sign\! ( _j κ_j ). 3. Among indices with sign(κ^j)=sign∗sign( κ_j)=sign^*, select the top-m by |κ^j|| κ_j|. 4. If |∑jκ^j|<τ | _j κ_j |<τ (a small positive empirical threshold, e.g., τ=0.1τ=0.1), evaluate both sign choices and keep the subset with larger |κ^(ypur)|| κ(y_pur)|. 5. Project the full data matrix onto the subspace spanned by the estimated mixing vectors of the selected candidates (e.g., using the corresponding columns of the pseudo-inverse of the preliminary unmixing matrix), and re-run ICA on this reduced-dimensional data. We evaluate this heuristic alongside the oracle in Fig. 1c, and find comparable contrast restoration. If the selected subset yields near-zero contrast, this indicates sign cancellation or weak non-Gaussianity; a practical check is to assess the stability of κ(ypur)κ(y_pur) under bootstrap resampling and then enlarge or re-select ℳM if needed. Figure 2: Paired kurtosis-gap comparison in COBRE group ICA (n=155n=155) for model orders k=53k=53 and k=100k=100, with FNC insets. Top: two-level FNC summaries (edge-level S and fingerprint-level F). Bottom: example connectivity fingerprint (|z||z|) and one representative subject’s static FNC matrix for component k∗=51k^*=51 in the k=53k=53 solution (with |z||z| clipped at the 98th percentile for visualization). FNC insets are qualitative context only and are not used in the redundancy-law inference. I Experiments Sources are independent Student’s t-variables (df>4df>4), standardized to zero mean and unit variance, with population excess kurtosis κt(df)=6/(df−4) _t(df)=6/(df-4). Results report mean ± SEM over independent trials. Sample excess kurtosis is κ^(v)=1T∑t((vt−v¯)/σ^v)4−3 κ(v)= 1T _t((v_t- v)/ σ_v)^4-3. Unmixing error for symmetric FastICA (g(y)=y3g(y)=y^3) is err(W,A)=minP,Λ‖WA−PΛ‖F/kerr(W,A)= _P, \|WA-P \|_F/ k, where P is a permutation matrix and Λ is a non-singular diagonal matrix. (Note: Strictly speaking, guaranteeing the asymptotic 1/T1/ T rate in Corollary 2 requires df>8df>8 for finite eighth moments. However, for practical finite sample sizes, we empirically observe a stable 1/T1/ T noise-floor crossing even at df=8df=8.) Conditioning (Fig. 1a). With a well-conditioned 5×55× 5 mixing matrix (T=5000T=5000, 30 trials), we vary the kurtosis gap Δκ _κ via source degrees-of-freedom. As 1/Δκ1/ _κ increases from 0.50.5 to 3333, mean unmixing error rises from ≈0.07≈ 0.07 to ≈0.14≈ 0.14 (R2=0.78R^2=0.78), consistent with a 1/Δκ1/ _κ-type amplification trend. Redundancy (Fig. 1b). Balanced mixtures y=1R∑j=1Rsjy= 1 R _j=1^Rs_j with i.i.d. df=8df=8 sources (κt=1.5 _t=1.5), T=20,000T=20,000, and R=2,…,50R=2,…,50 exhibit |κ^(y)|≈cfit/R| κ(y)|≈ c_fit/R with cfit≈1.60c_fit≈ 1.60 (R2=0.986R^2=0.986), confirming the O(1/R)O(1/R) law. An unbalanced power-law mixture decays more slowly, consistent with Reff≪R_eff R (Corollary 1). The inset shows std(κ^)∼σ0/Tstd( κ) _0/ T (σ0≈5.3 _0≈ 5.3), crossing the population contrast near T≈3×104T≈ 3× 10^4, matching (5). Purification (Fig. 1c). At R=50R=50 with heterogeneous sources (df∈[6,30]df∈[6,30], T=15,000T=15,000), the full balanced mixture yields weak contrast (|κ^|≈3×10−2| κ|≈ 3× 10^-2). Oracle purification selecting the top-m positive-kurtosis sources restores contrast to ≈0.43≈ 0.43 at m=5m=5 (∼14× 14× gain), while a sample-based sign estimator (τ=0.1τ=0.1) achieves comparable results. The monotone decay with m follows the 1/m1/m scaling of Theorem 2. COBRE real-data sanity check (supportive, time-course level). We test an observable implication of the redundancy law using resting-state fMRI group-ICA outputs from the COBRE cohort (n=155n=155). We run two group decompositions with model orders k∈53,100k∈\53,100\. For each subject, component time courses are z-scored across time, and we compute the sample excess kurtosis κ κ for each component. We summarize within-subject non-Gaussian contrast by the kurtosis-gap statistic G:=mean(top-5 |κ^|)−median(|κ^|)G:=mean(top-5 | κ|)-median(| κ|), and compare G across model orders using a paired Wilcoxon signed-rank test (one-sided, testing G53>G100G_53>G_100). Because time courses are standardized and inference is rank-based, this comparison is insensitive to scale differences between decompositions. Increasing model order from k=53k=53 to k=100k=100 produces a consistent reduction in G across subjects (Fig. 2; most paired points lie below the identity line), with a highly significant paired Wilcoxon result (p=1.74×10−27p=1.74× 10^-27). This provides a within-cohort sanity check of the predicted shrinkage of contrast with increasing model order, while we note that changing k also alters the signal subspace and effective degrees of freedom, so the comparison is supportive rather than a controlled causal test. In the FNC insets, edge-level summary S and fingerprint-level summary F also differ between model orders (Level-1 S: p=4.11×10−2p=4.11× 10^-2, d=0.12d=0.12; Level-2 F: p=7.58×10−23p=7.58× 10^-23, d=−1.14d=-1.14). The edge-level effect (d=0.12d=0.12) is small and reported as descriptive context; the primary real-data evidence is the reduction in G. IV Conclusion We established three results for kurtosis-based ICA as mixture width increases: (1) balanced mixing induces 1/R1/R decay of kurtosis contrast; (2) projection-level active width must scale at most as T T for kurtosis-contrast viability (a necessary screening condition); and (3) sign-consistent subset selection restores non-vanishing contrast. These results help explain high-model-order instability in group neuroimaging: for well-conditioned blocks, balance is typical (Proposition 1), so contrast collapse is structural rather than purely algorithmic. The theory is population-level and screening-oriented: it characterizes contrast availability and finite-sample viability, rather than providing a sufficiency guarantee for any specific ICA algorithm. In practice, it offers a simple rationale for model-order screening and for targeted contrast restoration via purification before re-estimation. A simple purification heuristic is near-oracle in our experiments. Our analysis is specific to kurtosis-based contrast and linear instantaneous ICA; other criteria (e.g., negentropy), structured/sparse mixing, and nonlinear or convolutive settings [13] require separate study. Future work includes adaptive purification within ICA and extensions to multi-domain settings such as IVA-L [2, 16]. References [1] A. Abou Elseoud, H. Littow, J. Remes, T. Starck, J. Nikkinen, J. Nissilä, M. Timonen, O. Tervonen, and V. Kiviniemi (2011) Group-ica model order highlights patterns of functional brain connectivity. Frontiers in systems neuroscience 5, p. 37. Cited by: §I. [2] M. Anderson, G. Fu, R. Phlypo, and T. Adalı (2014) Independent vector analysis: identification conditions and performance bounds. IEEE Transactions on Signal Processing 62 (17), p. 4399–4410. Cited by: §IV. [3] A. Auddy and M. Yuan (2025) Large-dimensional independent component analysis: statistical optimality and computational tractability. The Annals of Statistics 53 (2), p. 477–505. External Links: Document Cited by: §I. [4] A. J. Bell and T. J. Sejnowski (1995) An information-maximization approach to blind separation and blind deconvolution. Neural computation 7 (6), p. 1129–1159. Cited by: §I. [5] V. D. Calhoun, T. Adali, G. D. Pearlson, and J. J. Pekar (2001) A method for making group inferences from functional MRI data using independent component analysis. Human Brain Mapping 14 (3), p. 140–151. External Links: Document Cited by: §I, §I. [6] V. D. Calhoun and T. Adali (2012) Multisubject independent component analysis of fMRI: a decade of intrinsic networks, default mode, and neurodiagnostic discovery. IEEE Reviews in Biomedical Engineering 5, p. 60–73. External Links: Document Cited by: §I, §I. [7] J. Cardoso and A. Souloumiac (1993) Blind beamforming for non-Gaussian signals. IEE Proceedings F (Radar and Signal Processing) 140 (6), p. 362–370. External Links: Document Cited by: §I. [8] J. Cardoso (1989) Source separation using higher order moments. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Glasgow, UK, p. 2109–2112. Cited by: §I. [9] J. Cardoso (1998) Blind signal separation: statistical principles. Proceedings of the IEEE 86 (10), p. 2009–2025. External Links: Document Cited by: §I. [10] A. Chen and P. J. Bickel (2006) Efficient independent component analysis. The Annals of Statistics 34 (6), p. 2825–2855. Cited by: §I. [11] P. Comon (1994) Independent component analysis, a new concept?. Signal Processing 36 (3), p. 287–314. Cited by: §I, §I. [12] F. Herrmann and F. J. Theis (2007) Statistical analysis of sample-size effects in ICA. In International Conference on Intelligent Data Engineering and Automated Learning (IDEAL), p. 468–477. Cited by: §I. [13] A. Hyvärinen, I. Khemakhem, and H. Morioka (2023) Nonlinear independent component analysis for principled disentanglement in unsupervised deep learning. Patterns 4 (10). Cited by: §IV. [14] A. Hyvärinen and E. Oja (2000) Independent component analysis: algorithms and applications. Neural Networks 13 (4–5), p. 411–430. Cited by: §I, §I-B. [15] A. Hyvärinen (1999) Fast and robust fixed-point algorithms for independent component analysis. IEEE Transactions on Neural Networks 10 (3), p. 626–634. External Links: Document Cited by: §I, §I-B. [16] T. Kim, T. Eltoft, and T. Lee (2006) Independent vector analysis: an extension of ica to multivariate components. In International conference on independent component analysis and signal separation, p. 165–172. Cited by: §IV. [17] Z. Koldovský, P. Tichavský, and E. Oja (2006) Efficient variant of algorithm FastICA for independent component analysis attaining the Cramér–Rao lower bound. IEEE Transactions on Neural Networks 17 (5), p. 1265–1277. External Links: Document Cited by: §I. [18] C. J. Li and M. I. Jordan (2021) Stochastic approximation for online tensorial independent component analysis. arXiv preprint arXiv:2012.14415. Cited by: §I. [19] Y. Li, T. Adalı, and V. D. Calhoun (2007) Estimating the number of independent components for functional magnetic resonance imaging data. Human Brain Mapping 28 (11), p. 1251–1266. Cited by: §I. [20] J. Miettinen, K. Nordhausen, H. Oja, and S. Taskinen (2015) Fourth moments and independent component analysis. Statistical Science 30 (3), p. 372–390. Cited by: §I, Corollary 2. [21] F. Ricci, L. Bardone, and S. Goldt (2025) Feature learning from non-gaussian inputs: the case of independent component analysis in high dimensions. arXiv preprint arXiv:2503.23896. Cited by: §I. [22] R. F. Silva, S. M. Plis, T. Adalı, M. S. Pattichis, and V. D. Calhoun (2020) Multidataset independent subspace analysis with application to multimodal fusion. IEEE Transactions on Image Processing 30, p. 588–602. Cited by: §I. [23] R. F. Silva, E. Damaraju, X. Li, P. Kochunov, J. M. Ford, D. H. Mathalon, J. A. Turner, T. G. M. van Erp, T. Adali, and V. D. Calhoun (2024) A method for multimodal IVA fusion within a MISA unified model reveals markers of age, sex, cognition, and schizophrenia in large neuroimaging studies. Human Brain Mapping 45 (17), p. e70037. External Links: Document Cited by: §I. [24] R. Vershynin (2018) High-dimensional probability: an introduction with applications in data science. Vol. 47, Cambridge university press. Cited by: §I-A.