Paper deep dive
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/29/2026, 3:56:29 AM
Summary
This paper presents the first application of sparse autoencoder (SAE)-based mechanistic interpretability to a neutrino foundation model (PolarBERT) trained on IceCube data. The authors identify a validated atlas of physical concepts (e.g., detector quality, brightness) within the model's internal representations. They demonstrate that the primary direction reconstruction head largely ignores these interpretable features, whereas a newly trained uncertainty head causally relies on them. Utilizing this interpretable information, the uncertainty head significantly improves median angular resolution from 20.2° to 3.2° at 20% selection efficiency.
Entities (7)
Relation Signals (5)
PolarBERT → trainedon → IceCube
confidence 99% · Studying a neutrino foundation model pretrained on IceCube data
Uncertainty Head → improves → Angular Resolution
confidence 95% · At 20% selection efficiency, this interpretable estimator improves the median angular resolution from 20.2° to 3.2°.
Sparse Autoencoder (SAE) → usedfor → Mechanistic Interpretability
confidence 95% · We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics.
Uncertainty Head → dependson → Interpretable Latents
confidence 92% · Unlike the direction head, it depends causally on quality and brightness features from the atlas.
Direction Head → ignores → Interpretable Latents
confidence 90% · Causal interventions show that the direction head barely draws on this atlas.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.
Tags
Links
- Source: https://arxiv.org/abs/2608.26090v1
- Canonical: https://arxiv.org/abs/2608.26090v1
Trouble viewing inline? Open PDF directly →
Full Text
115,795 characters extracted from source content.
Expand or collapse full text
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders Raphaël Bonnet-Guerrini1,2, Johann Ioannou-Nikolaides3, Inar Timiryasov3, and Vincenzo Piuri1, 1 Computer Science Department, Università degli Studi di Milano, Milano, Italy 2 INFN Sezione di Milano, Milano, Italy 3 Niels Bohr Institute, University of Copenhagen, Copenhagen, Denmark Abstract We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model’s angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At %20\% selection efficiency, this interpretable estimator improves the median angular resolution from 20.2∘20.2 to 3.2∘3.2 . These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model’s internal representation and help design downstream tasks that exploit it. Contents 1 Introduction Modern neutrino telescopes reconstruct particle properties from sparse Cherenkov light recorded in large, irregular detector volumes [24, 2]. In IceCube, each event is a variable-length set of photomultiplier pulses on digital optical modules (DOMs) embedded in Antarctic ice [3, 1]. Direction reconstruction depends on the visible light pattern, but also on detector geometry, ice properties, noise and trigger conditions. As machine-learning-based reconstructions become more expressive, the learned representation that carries these effects becomes an object of physics interest in its own right. Foundation models are a natural fit in this domain. They use a shared representation from broad detector data and couple it to task-specific heads. This can reduce task-specific training cost and improve statistical efficiency, as various downstream tasks can read the same representation differently. This trend mirrors the broader success of scaling and pretraining in machine learning [22, 9], and is now emerging in particle physics through masked-particle modeling, pretraining, joint fine-tuning, and neutrino-specific foundation models [25, 20, 37, 35, 38]. Foundation models thus raise a concrete interpretability problem: which structures are encoded in the shared representation, and which of them are used by a given head? Standard post-hoc explanations alone do not answer this problem. Saliency maps, integrated gradients, Grad-CAM-like methods, and broader explainable-AI tools assign importance scores to inputs or intermediate quantities [33, 34, 31, 30]. While these explanations are useful, they are often local, approximate, and method-dependent, and do not expose the physical concepts latent within a representation. In particle physics, recent work has connected learned representations to jet-substructure observables using attribution methods, probes, path patching, Shapley values, and symbolic regression [36, 27]. To move from feature attribution methods to understanding the model’s internal representation, we turn to mechanistic interpretability [14, 32, 29]. The investigation of the internal variables and circuits that determine the model’s output raises two questions with regard to neutrino reconstruction. Can mechanistic interpretability recover concepts within the internal representation of neutrino foundation models? If yes, which features can be named, tested, and intervened on? We study these questions with sparse autoencoders (SAEs). An SAE learns an overcomplete dictionary of directions in activation space. A single neuron can contribute to several interpretable concepts at the same time, its representations are superposed [18, 11]. An SAE addresses this problem by reconstructing each activation from only a few learned directions, which aim to separate concepts that overlap in the original neuron basis [16, 19, 13]. The main criticism of SAEs is that a sparse feature is not automatically interpretable, and an interpretable feature is not automatically causal. Recent studies have shown that learned features can depend on the training run, underperform linear probes on labelled concepts, or fail to act as reliable steering directions even when they appear interpretable [21, 28, 26]. To address this gap, end-to-end sparse dictionary learning preserves downstream behavior rather than activation structure alone [10]. We therefore treat SAE interpretability as a hypothesis that must be rigorously validated. Candidate features must survive held-out tests, matched controls, dictionary variation, and interventions. We develop such a validation protocol and apply it on PolarBERT, a BERT-like foundation model trained on IceCube pulse sequences and fine-tuned for direction reconstruction [35, 17]. The public Monte Carlo setting provides true direction, detector-level auxiliary flag, and per-event angular errors [15], enabling rigorous hypothesis testing. Because sparse dictionaries are stochastic objects, we test our verdict across independent SAE training seeds, dictionary objectives, and model layers. To our knowledge, this is the first application of SAE-based mechanistic interpretability to a particle-physics foundation model. We show that PolarBERT’s frozen representations contain a rich atlas of physical concepts including detector quality, auxiliary activity, event brightness, and detector depth. Yet, the downstream direction head uses almost none of this atlas. Removing or overwriting identified features, both individually and in families, leaves the predicted direction essentially unchanged. Instead, one stable axis with no simple physics name dominates the aggregated Jacobian sensitivity, although local sensitivity varies substantially between events. This suggests that the shared representation contains substantial physical information that the direction objective leaves underused. Motivated by this underused information, we train an uncertainty head to predict the angular reconstruction error. Unlike the direction head, the uncertainty head depends causally on the bright-clean read-out and related clean-side features. At 20%20\% selection efficiency, it achieves a median angular resolution of 3.2∘3.2 , compared with 20.2∘20.2 for selection using the best single detector observable. The interventions show that this head causally uses physically interpretable information in the representation. The paper is organized as follows. Section 2 introduces PolarBERT, the sparse dictionaries, and the validation protocol. Section 3 maps the atlas of validated read-outs and its stability across dictionary draws. Section 4 shows that the causal use of this atlas depends on the downstream head. Section 4.4 uses the resulting uncertainty estimate for event selection. Section 5 summarizes the implications and future directions. 2 Setup This section defines the objects used throughout the paper. We first specify the frozen PolarBERT representation and the downstream heads that read it. We then define the sparse dictionaries trained on this representation. Finally, we describe the validation protocol used to separate readable features, validated physical read-outs, and causal features. 2.1 PolarBERT as a controlled testbed PolarBERT is an eight-block BERT-like transformer pretrained on IceCube pulse sequences and fine-tuned for neutrino direction reconstruction [35, 17]. During pretraining, the backbone learns through masked-DOM prediction and total-charge regression from the CLS state. During fine-tuning, the backbone and direction head are updated jointly. We freeze the fine-tuned direction backbone when training the uncertainty heads, so that both heads share the same representation. Across this study, we analyze the fine-tuned backbone. We only use the pretrained checkpoint alone in Section 4 to determine which physical information was already present before direction fine-tuning. A variable-length sequence of detected pulses represents each event. The arrival time, measured charge, digital optical module (DOM) identifier, and a Boolean auxiliary flag describe each pulse. In the public IceCube sample, this flag identifies pulses that were not fully digitized [15]. The full waveform only gets read out when the hard local coincidence (HLC) condition is met (for this at least one neighboring DOM on the same string also needs to record a signal within 1�s1μ s [12]). Pulses that do not meet this criterion are consequently more susceptible to uncorrelated photomultiplier noise, although the flag is not itself a truth-level noise label [4, 3]. We summarize the auxiliary content of each event as seen by the model using two observables defined on the tokenized input window. Each event contains at most L=127L=127 pulses before tokenization, with non-auxiliary pulses retained preferentially and the remaining slots filled by a uniform random sample of the auxiliary pulses. Let iT_i denote the resulting set of input pulses of event i (|i|=Npulse,i≤L|T_i|=N_pulse,i≤ L), let i⊆iA_i _i denote the subset carrying the auxiliary flag, and let qipq_ip be the measured charge of pulse p. We define faux,i=|i||i|,qaux,i=∑p∈iqip∑p∈iqip.f_aux,i= |A_i||T_i|, q_aux,i= _p _iq_ip _p _iq_ip. (1) Here, i indexes events and p indexes pulses within an event. The first quantity is the fraction of input pulses marked auxiliary, while the second is the fraction of the total charge in the input window carried by those pulses. For the 87%87\% of events that fit within the window, these quantities coincide with their full-event counterparts while for truncated events, they instead characterize the input actually presented to the model. We use the following event classes throughout: • Clean: faux<0.05f_aux<0.05 and qaux<0.05q_aux<0.05. • Auxiliary-dominated:faux>0.90f_aux>0.90 and qaux>0.80q_aux>0.80. • Within the clean class, we additionally define two brightness subsets: – Bright-clean:Qtot≥160Q_tot≥ 160. – Dim-clean: Qtot≤122.16Q_tot≤ 122.16. PolarBERT maps the pulse sequence to a contextualized pulse representation. In BERT-like models, the sequence begins with a learned CLS token, whose hidden state aggregates information from the pulse tokens. For event i, we denote the complete hidden-state sequence after transformer block ℓ by Hi(ℓ)H_i^( ), including both the CLS state and the pulse-token states. The final transformer-block CLS-token is hi=Hi,CLS(8)∈R256.h_i=H^(8)_i,CLS∈ R^256. (2) Every downstream head reads only this CLS representation. These internal activations of the frozen representation are the object investigated in this study. The frozen direction head maps the final CLS state to the sphere: gdir:R256→S2,u^i=gdir(hi),g_dir: R^256→ S^2, u_i=g_dir(h_i), (3) and is evaluated against the Monte Carlo truth direction uiu_i through �=iarccos(u^i⋅ui) _i= ( u_i· u_i ). Because gdirg_dir consumes only hih_i, interventions on the CLS state act on the complete input of the head, there is no side channel for directional information to bypass them. The final-block CLS representation defines the common vector space for sparse dictionaries, probes, principal components, and interventions. Later checks compare this final summary to information available at earlier layers. We report layer-wise scans in Appendix A. Crucially, our protocol trains the SAE purely on raw model activations, without requiring MC truth labels. This ensures that the methodology applies directly to real detector data, where truth information is absent or incomplete. Here, PolarBERT serves as a testbed: the main concept criteria are defined from detector-level observables, while Monte Carlo truth provides additional controls and evaluation quantities. The analysis uses a one-million-event sample forward-passed through PolarBERT’s fine-tuned backbone and partitioned into five disjoint sets. The sets, presented in Table 1, separate sparse-autoencoder training, reconstruction validation, concept discovery, independent concept validation, and causal testing. This separation is important because it ensures that we do not use the same Monte Carlo labels to find a candidate feature, tune its controls, and evaluate the final claim. Split name Role Fraction sae_train SAE weight training 75%75\% sae_reconstruction_val Reconstruction and fidelity validation 10%10\% concept_dev Candidate latent discovery 7%7\% concept_test Independent concept validation 6%6\% causal_test Independent causal interventions 2%2\% Table 1: Disjoint event splits used for SAE training and downstream validation. Fractions refer to the nominal one-million-event partition. Before training the SAE, we normalize the CLS activations using only the sae_train split. Let �μ and �σ denote the training-set mean and standard deviation. The normalized activation is h~i=hi−�. h_i= h_i-μσ. (4) We apply the same fixed transformation to all the other splits. We perform all SAE training and decoding in this normalized activation space. 2.2 Sparse dictionaries Intuitively, an SAE rewrites each CLS summary as a longer, low-frequency activation, so that the resulting few activations per event are more easily interpretable than the original dense vector. Architecture and sparsity. An SAE represents h~i h_i through an overcomplete sparse code zi∈Rmz_i∈ R^m with m≫dm d, where d=256d=256 is the PolarBERT activation dimension. The decoder reconstructs the normalized activation as h~^i=bdec+Wdeczi=bdec+∑j=1mzi,jfj, h_i=b_dec+W_decz_i=b_dec+ _j=1^mz_i,jf_j, (5) where bdec=D(0)b_dec=D(0) is the learned decoder bias, i.e. the baseline reconstruction when all latent activations are zero, and fj=(Wdec):,jf_j=(W_dec)_:,j is the j-th decoder direction. Each latent coordinate zi,jz_i,j is therefore a candidate feature activation, while the corresponding decoder vector fjf_j defines a direction in the normalized PolarBERT activation space. The encoder first maps the normalized CLS representation to a vector of latent scores, or pre-activations ai=Wench~i+benc.a_i=W_enc h_i+b_enc. Then, instead of using an L1L_1 penalty to encourage sparsity, we impose a hard sparse budget with BatchTopK [13]. For a batch of size B, BatchTopK keeps the largest BkB_k positive pre-activations across the batch and sets the rest to zero: zi,j=max(ai,j,0) 1[max(ai,j,0)≥�B],z_i,j= (a_i,j,0)\,1 [ (a_i,j,0)≥ _B ], (6) where �B _B is the threshold corresponding to the BkB_k-th largest positive activation in the batch. This enforces an average sparse budget of k active latents per event while allowing the number of active latents to vary from event to event. We use BatchTopK because its batch-level budget allows the number of active latents to adapt across events. A concept-independent sweep selects k=16k=16: the sparser k=8k=8 setting reduces fidelity and dictionary utilization, whereas k=32k=32 increases feature density for only a marginal reconstruction gain. Training objective. The main reconstruction loss is the mean squared error in normalized activation space, ℒrec=1B∑i=1B‖h~i−h~^i‖22.L_rec= 1B _i=1^B h_i- h_i _2^2. (7) To avoid dead latents, we add the AuxK revival loss used in scalable TopK sparse autoencoders [19]. AuxK prevents rarely used latents from becoming permanently inactive by asking them to reconstruct the part of the input that is not explained by the main sparse code. For each event, we define the detached reconstruction residual ri=stopgrad(h~i−h~^i).r_i=stopgrad\! ( h_i- h_i ). (8) Among latents that have not recently activated, AuxK retains the kauxk_aux largest positive pre-activations, giving an auxiliary sparse code ziauxz_i^aux. The corresponding auxiliary loss is ℒaux=1B∑i=1B‖ri−Wdecziaux‖22,L_aux= 1B _i=1^B \|r_i-W_decz_i^aux \|_2^2, (9) where the decoder bias is omitted because the auxiliary path reconstructs the residual around the main reconstruction. The full training objective is then ℒSAE=ℒrec+�auxℒaux.L_SAE=L_rec+ _auxL_aux. (10) AuxK is empirically important in this setting: without it, the final full-scale training run leaves 286286 dead latents, whereas nonzero auxiliary coefficients eliminate dead latents without materially changing reconstruction quality. Selecting the dictionary. To arrive at the dictionary studied in this paper, we select the optimal hyperparameter set based on reconstruction diagnostics alone without using any downstream concept label. We scan the SAE capacity, sparsity budget, and AuxK revival parameters on sae_reconstruction_val, using the activation reconstruction quality, recovered directional fidelity, latent utilization, feature density, decoder redundancy, and dead-latent count [19, 23]. We summarize the selected configuration based on hyperparameter studies in Table 2. Choice Values tested Selected Basis Expansion m/dm/d 2,4,8\2,4,8\ 44 intrinsic diagnostics BatchTopK budget k 8,16,32\8,16,32\ 1616 fidelity–sparsity tradeoff AuxK coefficient �aux _aux 0,2−8,2−6,2−5,2−4,2−3,2−2\0,2^-8,2^-6,2^-5,2^-4,2^-3,2^-2\ 2−52^-5 zero dead latents AuxK budget kauxk_aux 64,256,512\64,256,512\ 512512 validation diagnostics Dead-step threshold �dead _dead 28,29,210,212,213\2^8,2^9,2^10,2^12,2^13\ 282^8 validation diagnostics Table 2: Concept-independent SAE selection. We select the configuration using intrinsic reconstruction and sparsity diagnostics on sae_reconstruction_val set, before any concept validation or causal analysis. We use the selected expansion-44, k=16k=16 dictionary for all concepts and causal analyses. This label-independent selection is essential for interpreting later sparse features as unsupervised read-outs rather than as coordinates chosen for agreement with known physics labels. We measure the reconstruction quality by the raw coefficient of determination, Rraw2=1−∑i‖h~i−h~^i‖22∑i‖h~i−h~¯‖22,R^2_raw=1- _i h_i- h_i _2^2 _i h_i- h _2^2, (11) where h¯ h is the validation-set mean activation in normalized space. We also evaluate whether the SAE reconstruction preserves the information used by a frozen downstream head. For each validation event, we replace the original activation by its SAE reconstruction, forward the patched representation through the frozen downstream computation, and compare the induced angular deviation to a mean-ablation baseline. The recovered-fidelity score is fidrec=1−Ei[arccos(gdir(h^i)⋅gdir(hi))]Ei[arccos(gdir(h¯)⋅gdir(hi))]=1−Ei[�SAE,i]Ei[�mean,i],fid_rec=1- E_i [ (g_dir( h_i)· g_dir(h_i) ) ] E_i [ (g_dir( h)· g_dir(h_i) ) ]=1- E_i[ _SAE,i] E_i[ _mean,i], (12) where �SAE _SAE is the angular deviation induced by SAE reconstruction and �mean _mean is the angular deviation induced by replacing the activation with its validation-set mean. The canonical dictionary reaches Rraw2≃0.994R^2_raw 0.994 and fidrec≃0.989fid_rec 0.989 with no dead latents. In absolute terms, replacing hih_i with its SAE reconstruction changes the direction-head output by a mean angular deviation of 0.65∘0.65 , whereas mean ablation changes it by 56∘56 . These two numbers set the scale of every intervention below, and rule out a lossy reconstruction as the explanation of later nulls. To test robustness to the dictionary draw, we train four seeds with the same architecture, splits, and normalization procedure. Across seeds, Rraw2=0.9937±0.0001R^2_raw=0.9937± 0.0001 and fidrec=0.989±0.001fid_rec=0.989± 0.001, with zero dead latents in each draw. We also train per-layer dictionaries on the CLS state of each transformer block, to test whether some concepts reside upstream of the final summary. Finally, we train functionally oriented dictionaries whose objective preserves the frozen direction-head output. These provide a stress test of whether a dictionary optimized for downstream behavior recovers nameable physical mechanisms for direction reconstruction. The resulting dictionaries are sparse coordinate systems for PolarBERT activations, not explanations by themselves. The validation protocol below classifies latents as candidates, validated read-outs, or causal features. The reference dictionary. Unless stated otherwise, we report quantitative SAE results for the expansion-44, k=16k=16, seed-1 dictionary selected above using intrinsic metrics alone. This is a reporting convention, not a claim that this basis is privileged. Individual latent indices are specific to one training run. Section 3.3 therefore tests which read-outs, response patterns, and causal verdicts persist across independent seeds and dictionary architectures, while functionally trained dictionaries provide a separate stress test across training objectives. We use the reference dictionary to report concrete coordinates, but treat only the replicated structures and verdicts as conclusions about the model representation. 2.3 Validation protocol Our protocol validates sparse directions through three checks: read-out quality, nuisance selectivity, and causal relevance. Throughout zi,jz_i,j denotes the activation of latent j on event i, and G the tested object - either a single latent G=jG=\j\ or a family of related latents. We use three terms consistently throughout the paper: the association step proposes a candidate, a candidate that survives the validation and matched-control tests is a validated read-out, and a validated read-out that additionally passes downstream intervention tests is a causal feature (see Figure 1). concept_dev concept_test g causal_test Figure 1: Concept-validation protocol. A candidate SAE read-out is first identified by association, then tested for held-out selectivity, and finally intervened on to establish whether a particular downstream head uses it. Robustness checks are applied across the validation chain. Exact numerical criteria are given in Table 5. Association. For a concept indicated by the label y(c)y^(c), we first rank candidate latents by their association with y(c)y^(c) on concept_dev and then re-evaluate them on an independent validation split, concept_test. For binary concepts, we evaluate each latent using the AUROC of zjz_j as a classifier and the difference in mean activation between positive and negative events. We additionally report the firing rate, P(zj>0)P(z_j>0), and the enrichment of positive events among the highest activations Eq(j,c)=P(y(c)=1∣zj≥Qq(zj))P(y(c)=1),E_q(j,c)= P (y^(c)=1 z_j≥ Q_q(z_j) )P (y^(c)=1 ), (13) where Qq(zj)Q_q(z_j) is the q-quantile of the latent activation. For continuous observables, we formulate the same step through response profiles or monotonic trends rather than a single binary AUROC. Selectivity. A candidate should track the concept we claim it tracks, not a correlated nuisance variable. We therefore recompute the AUROC after matching positive and negative events jointly on five nuisance variables: pulse count, total charge, non-auxiliary pulse count, and, in this controlled study, true zenith and azimuth** * The only MC-exclusive values used in the entire protocol., and retain the candidate only if this stricter AUROC on the matched sample still clears the validation threshold. We also compare each candidate to latents with similar firing rates, excluding dense latents whose removal degrades reconstruction broadly, rather than for reasons specific to the tested concept. Causal relevance. We test causal relevance only after a feature passes the checks above, and always relative to a specific downstream head g. The simplest intervention is removal. We zero the corresponding sparse activations, decode the modified code, and pass the resulting patched activation through g, h~i(−G)=D(zi⊙j∉G). h^(-G)_i=D (z_i 1_j∉ G ). (14) Removal asks whether g needs the feature. We also test writing when a concept has a natural source and target population. In this case, we replace the activations in G by donor values before decoding, zi,j(write)=wi,j,j∈G,zi,j,j∉G,z^(write)_i,j= casesw_i,j,&j∈ G,\\ z_i,j,&j∉ G, cases (15) where wi,jw_i,j is the written activation. Writing asks whether the decoded feature can steer the head in the expected direction. We also study dose response, in which we rescale its activation, zj→szjz_j→ s\,z_j, and measure the resulting change in the downstream head. For each intervention we compare the effect against matched controls. In the case of removal, we use as controls latents or latent sets with a similar firing-rate distribution. Writing controls are latents with matching firing rates and the same donor activation. We fix every threshold before evaluation and report it in Appendix A. We call a feature causal for g only if its effect is large enough, exceeds the matched-control effects, and is specific to the target concept. The fixed discovery, validation, control-matching, and causal-verdict rules are summarized in Appendix A.2. Supervised baseline. Alongside the sparse latents, we use linear probes as the supervised reference: a probe is a linear model fit on the discovery split to predict y(c)y^(c) from the full activation h~i h_i [8]. Its held-out performance measures how much of the concept we can extract linearly from the full representation. This is the baseline against which we can judge localized sparse read-outs throughout the paper. Between these two extremes, k-sparse probes restricted to the k most informative latents (selected on the discovery split, refit and scored on held-out events) measure how many sparse coordinates a concept actually requires. 3 From candidates to a validated atlas This section reports what the SAE-dictionary reads out of the final-layer CLS representation (we use probes on all layers to estimate which carries most of the directional information in Appendix A.1). Of the 10241024 latents in the sparse dictionary, most do not correspond to any identifiable concepts we tested, and several concept candidates, such as event morphology and sub-detector geometry, do not correspond to any validated read-out. This is not surprising as association alone is a weak filter, and a look-elsewhere effect can produce convincing-looking candidates by chance across 10241024 unlabelled directions. Selectivity is the step that removes candidates associated with a particular concept: the few latents that pass every test organize into a few interpretable concepts. We first present these validated read-outs and the structure they reveal in the sample, compare them against two linear baselines, and show that none of it depends on the particular dictionary. All latent indices refer to the reference dictionary of Section 2.2 and their independence from this choice is established in Section 3.3. Figure 2: Illustration of five events ordered by increasing activation of zbcz_bc on causal_test (z=0.35z=0.35, 0.450.45, 0.580.58, 1.031.03, 2.672.67). Panels share camera, detector frame and charge scale. 3.1 The identified physical atlas Validated read-outs. The strongest final-layer read-out, zbcz_bc, is a bright-clean feature (illustrated in Figure 2): it activates on events whose standard, non-auxiliary pulses dominate light and whose total charge is high. In the reference dictionary this is latent 640640. On the independent validation split, it separates bright-clean events from strongly auxiliary, dim events with AUROC 0.9110.911, and remains informative under all matched controls. Feature Role Validation / response Min. control AUROC Response analog zbcz_bc bright-clean events 0.910.91 0.790.79 4/44/4 zbc′z_bc^\, clean events, secondary 0.690.69 0.840.84 4/44/4 zauxz_aux auxiliary activity 0.840.84 0.850.85 3/43/4 zbsz_bs brightness-associated support (dense) charge monot. +0.92+0.92 non-selective 4/44/4 zdepthz_depth detector depth (layer 2) 0.9860.986 0.9840.984 layer bank Table 3: Named sparse directions identified in the paper. The final column reports whether an analogous response appears across dictionary seeds; it does not report replication of the complete validation criteria. Section 3.3 gives the corresponding seed-by-seed verdicts. The complete clean- and auxiliary-side discovery rankings, including the candidates that were not retained, are reported in Appendix B.1. A weaker clean-side feature, zbc′z_bc^\, , corresponds to latent 10061006 and passes the same controls at lower discrimination. Latent 195195, zauxz_aux, captures the complementary event class which activates on auxiliary-dominated, low-charge events. Finally, zbsz_bs, latent 973973 in the reference dictionary, is a dense brightness-associated support feature. Its mean activation rises strongly with total charge, but it does not cleanly discriminate high- from low-charge events and fails the selectivity controls. We therefore do not treat it as a validated brightness read-out. evertheless, interventions on this coordinate produce a large change in the direction-head output. Section 4.1 shows that the direction reconstruction is strongly disrupted when this coordinate is moved away from its natural value. Figure 3: Response profiles of the main SAE latents along the auxiliary fraction and the total charge, on concept_dev set events. Bands show the min-max spread across dictionary seeds. In Figure 4, we show the most activated event for these latents. zbcz_bc fires on a compact bright deposit (391391 p.e., 1212 DOMs, 33 strings). zauxz_aux fires on sparse auxiliary-dominated activity across 3030 strings. The zbsz_bs panel is not a bright event (129129 p.e.; per-event rank correlation with total charge 0.050.05 on causal_test), consistent with it being a support direction rather than a concept read-out. Figure 4: Highest-activating event of each latent on causal_test. Markers are DOMs (charge summed per DOM), area grows with charge charge, color encodes hit time within the event, and filled versus hollow marks non-auxiliary versus auxiliary hits. Panels share axes and charge scale. Cleanliness and brightness are one axis. Auxiliary-dominated events are also systematically dimmer in this sample, so auxiliary activity and brightness form a strongly correlated axis. The bright-clean and dim-clean subsets defined in Section 2.1 allow us to test brightness while holding event quality fixed. We also test the analogous brightness contrast within auxiliary-dominated events, but find no validated sparse read-out. Within the clean class, zbcz_bc also distinguishes bright-clean from dim-clean events, with a validation AUROC of 0.8210.821. Both groups satisfy the same clean-event criteria and differ only in their total-charge selection, showing that zbcz_bc also carries a brightness dependence. We therefore describe zbcz_bc as a bright-clean read-out rather than as a pure event-quality feature. When applied to the other candidate latents, this within-clean-class test eliminates most of the weaker candidates. No latent that discriminates based on fauxf_aux survives the analogous within-class brightness test. Thus, some response-profile shapes are real physical structures, while others arise from correlations between observables. The response landscape. The protocol identifies and validates a few latents, we now ask whether they are a part of a larger family. Cleanliness and brightness are broad concepts, and we find that the validated read-outs are not isolated. Instead, they sit within a broader response landscape, which we characterize by dividing all events into ten bins of auxiliary fraction fauxf_aux. For each latent j, we compute its mean activation in each bin and take the Spearman correlation between these ten mean activations and the corresponding fauxf_aux bin centers. We denote this auxiliary-monotonicity score by amsjams_j. Thus, amsj≃−1ams_j -1 indicates a latent whose activation decreases toward auxiliary-dominated events, whereas amsj≃+1ams_j +1 indicates one whose activation increases. We define the latent sets used later for joint interventions as: • Clean pair: zbc,zbc′\z_bc,z_bc^\, \. • Clean family: all latents with ams≤−0.85ams≤-0.85, excluding the dense brightness-support direction zbsz_bs; this gives 194194 latents. • Clean core: the members of the clean family with firing rate P(zj>0)≥0.05P(z_j>0)≥ 0.05; this gives 55 latents. • Auxiliary family: all latents with ams≥+0.85ams≥+0.85; this gives 135135 latents. The exact reference-dictionary definitions and sizes of these intervention sets are summarized in Appendix B.2. A monotonic response, however, is not the same as a validated read-out. Many latents follow the global trend only weakly but do not pass every step of the validation protocol. The result is therefore not hundreds of clean or auxiliary features. It is a continuous quality-brightness axis spread across many sparse coordinates, with only a few coordinates sharp enough to validate individually. This matters for causality: if the direction head used the axis in a distributed way, removing one latent would not have a decisive impact. Section 4.2 therefore tests the causality of both the single validated read-outs and the full response-profile families. Discoveries beyond the target concepts. Because we train the dictionary without labels, it can also expose structure that we do not specify in advance. Some latents vary smoothly with zenith or azimuth direction (we report them in Appendix B.4.3). They are more active for events from one broad horizontal direction than from the opposite one, but no latents show the characteristic azimuthal dependence that arises from the detector’s hexagonal string spacing [5]. However, given our strict protocol, none of these candidates pass the selectivity test. We still report them to illustrate the use case of sparse dictionaries for discovery: unlike a supervised probe, which requires the concept to be named before training, an SAE can surface candidates that then require thorough validation. What the dictionary does not read, and why. While we find candidates for many concepts, the dictionary built from the last residual layer has an informative blind spot. DOM position information enters at the pulse-token level, so the network learns detector geometry early. At the final layer, however, no latent tied to event radius, depth, or sub-detector location survives the matched geometry contrasts. Using the same SAE recipe at layer 2 recovers a near-monosemantic depth feature, zdepthz_depth, with validation AUROC 0.9860.986. This does not mean, however, that geometry information is lost. At earlier layers, geometry is distributed across the sequence, and later transformer blocks can still integrate information from the pulse tokens into the final CLS state. The early CLS-level depth feature is therefore not necessarily a causal bottleneck for the direction prediction. Absence of a concept from one dictionary therefore does not imply absence from the model in general. A single-layer atlas should not be read as an inventory of everything the network represents. We similarly find no validated single-latent read-out for the elongated-versus-compact morphology contrast. We construct this contrast from charge-weighted pulse-pattern observables because the dataset provides no track/cascade labels. Because the Kaggle dataset lacks track/cascade labels, any such concept has to be identified using indirect features and would have to emerge indirectly from pulse-level structure. Linear probes nevertheless recover the continuous morphology proxies with held-out R2R^2 values between 0.460.46 and 0.650.65, showing that the representation contains morphology-related information without localizing it in one tested SAE coordinate. 3.2 Sparse read-outs versus linear baselines Section 3.1 reported what the SAE finds on its own. Before treating that atlas as a property of the representation, rather than an artifact of sparse dictionary learning, we check it against two linear baselines that make no sparsity projection: Principal Component Analysis (PCA), which asks what an unsupervised linear basis finds, and linear probes, which ask how much of a labelled concept the full representation carries. Linear probes are supervised models that combine all coordinates of the representation. We evaluate both in the same normalized final-layer CLS space as the SAE. PCA gives unsupervised linear directions vrv_r in the space of h~i h_i. We score event i on component r by si,rPCA=vr⊤(h~i−h¯PCA),s^PCA_i,r=v_r ( h_i- h_PCA ), (16) where h¯PCA h_PCA is the mean activation on the PCA training split. We fit the components on sae_train, select and sign-orient them on the concept_dev, and evaluate them on the independent validation split. A linear probe for concept c instead uses the labelled score si(c)=ac⊤h~i,s^(c)_i=a_c h_i, (17) with ac∈R256a_c∈ R^256. In the previous section we discussed that the SAE represents event quality and brightness through individually testable read-outs (zbcz_bc and zauxz_aux). We find that the tenth principal component (PC10) of the CLS summary represents event quality and brightness in a single entangled direction. While PCA and the SAE find and decompose this bright-clean axis differently, the fact that two unsupervised methods find the same axis is evidence that the axis is a property of the representation’s geometry, not an artifact of a specific method. This anticipates the result of Section 4: the quality-brightness axis organizes events physically, but the direction head assigns little causal weight to its validated sparse read-outs. Linear probes answer a slightly different question from a sparse dictionary. A probe is trained with the concept label and can combine all 256256 coordinates of the representation, so its performance measures how well the concept is linearly decodable from the full CLS state. A single SAE latent instead asks whether the same information has been localized into one unsupervised sparse coordinate. Accordingly, probes generally match or exceed the best single latent. For the angular reconstruction error, for example, a probe on the full CLS representation reaches AUROC 0.9250.925, while no individual SAE latent forms a validated error read-out. The error information is therefore present and linearly accessible in the representation, but distributed across multiple coordinates rather than localized in one sparse feature. This distributed behavior foreshadows Section 4.3: predicting angular reconstruction error requires a head that combines information distributed across the representation. Appendix B.4.1 uses k-sparse probes to test how strongly the morphology information concentrates in the SAE coordinates. No individual latent passes the fixed morphology criteria at layers 22, 33, or 88, while full linear probes recover the continuous morphology proxies. Thus, the information remains linearly accessible but does not localize in a single tested SAE coordinate. 3.3 Robustness of the atlas across dictionaries and inputs We further support our finding that the foundation model’s latent representation encodes a physically meaningful atlas by rerunning the same discovery, validation, and control pipeline across independent training seeds and across the different trained models of the hyperparameter studies. Across all seeds, we find latents that have a higher activation for cleaner events. An analog of zbcz_bc appears in every seed, and in three of four seeds it satisfies the same validation criteria used for the reference dictionary. The auxiliary read-out zauxz_aux passes the validation step only in the reference seed, although auxiliary-shaped latents appear in the other seeds as well. The clean and auxiliary families also appear in every seed with similar sizes. The zbsz_bs latent rising with charge is even more stable. In every dictionary draw we recover a dense, charge-monotonic but non-selective latent. In one of the seeds, we observe that the brightness-support behavior is carried jointly by two latents rather than concentrated in one. Detailed seed-by-seed mappings and validation results are reported in Appendix B.5. When re-running the clean-trigger isolation for m/d∈2,4,8m/d∈\2,4,8\ and k∈8,16,32k∈\8,16,32\, we find a control-surviving analog of zbcz_bc in every dictionary, with validation AUROC between 0.840.84 and 0.990.99. As a final check, we perturb the input events themselves and re-encode them. We inject synthetic auxiliary pulses, randomly remove pulses, or rescale all pulse charges before passing the event again through the frozen backbone and SAE. The verified latents respond in the expected directions: injected auxiliary pulses raise zauxz_aux and lower zbcz_bc activation. Pulse removal suppresses zbcz_bc and zbsz_bs activation, while charge rescaling moves zbsz_bs with the charge scale. These perturbations support the physical labels of the validated read-outs and, separately, the brightness association of zbsz_bs. We further test their causal role for the direction head in Section 4.2. 4 Causality is head-relative Section 3 established a validated, robust atlas of concepts with a clear interpretation. However, a concept can be reliably decodable and still play no role in what a downstream head computes. This section tests that gap directly, using the removal and writing interventions of Eqs. (14) and (15), always applied to the final-layer CLS state hih_i. We pass every intervention through the frozen direction head and then through an uncertainty head trained on hih_i. The two heads return opposite verdicts on the causality of the clean-side atlas. Finally, we compare the pretrained and fine-tuned backbones to test whether these physical read-outs are already present before direction fine-tuning. 4.1 The atlas is inert for the direction head Removal and writing on the direction head. As a scale reference, we replace the final CLS state with its validation-set mean. This intervention changes the direction-head output by 56.4∘56.4 . We use this value as the baseline. Compared to this, zeroing zbcz_bc shifts the angular error by only +0.06∘+0.06 , essentially the same as removing unrelated latents with matched firing rates (+0.05∘+0.05 ). This null also replicates across dictionary draws, with all analogs remaining within |�|≤0.1∘| |≤ 0.1 . We repeat the removal test at every scale of the atlas. In the reference dictionary, the clean pair, clean family, clean core, and auxiliary family all give null results. The clean pair gives a null result in all four seeds. The clean family and clean core give null results in three of four seeds. In the negative seed, both family ablations change the output by about 10∘10 because the family includes a reconstruction-critical support latent that does not validate as a clean concept. The auxiliary family gives a null result in all four seeds. Removing a latent is only one intervention, we also test the opposite, amplifying the activation by writing. This allows us to test whether a read-out can steer the head at all. We first write clean-like values of zbcz_bc and zbc′z_bc^\, into auxiliary events, and conversely write an auxiliary-like value of zauxz_aux into clean events. We also repeat these interventions while removing the latent signature of the original class. In all cases, the direction prediction is unchanged. Steering brightness within the clean class is likewise null (Appendix C.3). The direction head’s output remains unchanged in both directions: it neither needs the quality read-outs nor can it be steered by imposing them. zbsz_bs steers the direction head. One dense latent, however, behaves differently from the other validated read-outs. zbsz_bs, introduced in section 3.1, is functionally important. Setting it to zero increases the direction-head angular error by 15.0∘15.0 on clean-bright events and 10.9∘10.9 on clean-dim events. Moving zbsz_bs either below or above its natural value damages direction reconstruction systematically with a minimum angular error at its unperturbed value(Figure 5, left). The direction head therefore does not read larger zbsz_bs as “more brightness”. Rather, this charge-correlated coordinate supports a well-formed representation around its natural operating point. The zbsz_bs-like latents recovered in all four dictionary draws show the same V-shaped impact on the angular reconstruction error. Figure 5: Dose response of the brightness-support latent zbsz_bs (latent 973973) on causal_test: this single coordinate is rescaled as z→sz→ s\,z before decoding. In both panels the two curves are clean events (faux≤0.05f_aux≤ 0.05) split by brightness alone: bright (Qtot≥160Q_tot≥ 160 p.e.) and dim (Qtot≤122Q_tot≤ 122 p.e.). Left: the resulting change in the true angular error of the direction head. Right: the resulting shift M1M_1 in the log-error predicted by the uncertainty head (Section 4.3). Dictionary construction independence. Perhaps the activation dictionary is simply the wrong basis for this head. To test this, we use end-to-end sparse (e2e) dictionaries [10]. They are specifically design to preserve the output of a downstream head, pushing the recovery of causal features, which the reference dictionary misses. The learning objective of E2e dictionaries is ℒe2e=d(gdir(h^i),gdir(h~i))+�ℒrec,L_e2e=d (g_dir( h_i),\,g_dir( h_i) )+α\,L_rec, (18) with d an angular distance and �≥0α≥ 0 interpolating between a local dictionary and pure e2e dictionary. We test every causal latent against the full concept vocabulary: quality, auxiliary activity, charge, zenith, azimuth, depth, morphology, and error (Appendix C.4). The two pure end-to-end dictionaries contain no bright-clean analog, with a best AUROC of about 0.530.53. The two reconstruction-anchored dictionaries recover bright-clean analogs with AUROC 0.870.87-0.880.88, but these latents remain outside the causal sets. 4.2 The direction head’s causal axis The previous interventions show that the direction head depends strongly on the CLS representation, but only weakly on the physical read-outs identified by the SAE. We therefore ask which directions in the 256256-dimensional CLS space the frozen head is locally sensitive to. The local sensitivity of the frozen head to the summary is its Jacobian, Ji=∂u^i∂hi∈R3×256,J_i= ∂ u_i∂ h_i∈ R^3× 256, (19) computed exactly through the frozen head. We compute the Jacobian of the direction head for each event in the discovery split and combine these Jacobians to identify the representation directions to which the head is most sensitive. We detail the process in Appendix A.3. At the dataset level we find a direction that is stable across splits. We quantify how similar the directions across different splits are by their principal angle in the 256256-dimensional space. The leading axis obtained on the concept_dev differs by only 5.6∘5.6 from the axis that we compute independently on concept_test. This angle lies far from the average 90∘90 split of two random directions. It accounts for 98%98\% of the aggregate sensitivity on the latter. This dataset-level stability does not mean every event relies on the same direction locally. For a given event, we can ask what fraction of its own local sensitivity the fixed axis captures. The typical event captures none of it - the median is 0% because most individual Jacobians are close to degenerate to begin with. So, the head’s local response is nearly flat in most directions for most events. It is the minority of events that the shared axis explains well, which is what produces the higher sample mean of 14.9%. The stable axis is therefore a dataset-level summary of where sensitivity concentrates, not a bottleneck every event’s prediction passes through. We ask four questions about this stable axis, in the following order: does it match a named concept, is it simply the direction of largest generic variance, does the direction head depend on a physical quantity, and is it really causal rather than merely correlated with the output? It does not match any named concept and no verified latent exceeds |cos|=0.15| |=0.15, see Figure 6. However, it is also not simply a generic direction. Its strongest alignment with the two leading principal components is |cos|=0.62| |=0.62 and 0.520.52, meaning that it is related to the directions that dominate the representation’s overall variance, but clearly distinct from them. Each probe direction in Figure 6 is the linear-probe construction of Section 2.3, Eq. (17), which we use to test whether the Jacobian of the direction head is related to a labeled quantity. We fit each linear model to predict one labelled quantity (angular error, auxiliary (charge) fraction, charge, zenith) from the full CLS representation is only weakly aligned with the causal axis and at most |cos|=0.1| |=0.1 for the charge probe. The clearest evidence that the Jacobian and the direction head are causal and not merely correlated comes from the causal latents of the functionally trained dictionaries. The set of causal latents of the reconstruction-anchored functional dictionary aligns strongly with the direction-head causal subspace, with |cos|≃0.8| | 0.8. This direction is meaningful and its misalignment with the physical latents of the named concepts is a finding, not a coincidence. Finally, perturbing the CLS representation along the Jacobian-derived sensitivity directions changes the predicted direction, confirming that these directions have a causal effect. The induced change is event-dependent and does not correspond to a simple operation, such as a fixed rotation. Figure 6: Sensitivity structure of the direction and uncertainty heads. Left: normalized singular-value spectra of the aggregate head Jacobians on the concept_dev and concept_test splits. Right: absolute cosine similarity between each head’s leading sensitivity axis and selected SAE decoder directions, physics-probe directions, and the first two principal components. The dashed vertical line marks |cos|=0.14| |=0.14. The direction head, then, is not indifferent to its representation, but has a real, replicable causal axis, just not one that matches any concept we can name. This raises a sharper question: is that indifference to the atlas a property of the quality and brightness features themselves, or a property of this particular head? If inertness were a property of the features, it would persist for any reader. If it is instead head-dependent, a task that actually requires event-quality information should use them, even though the direction head does not. Direction reconstruction has no explicit need to assess whether an event is clean. However, error prediction does. 4.3 An uncertainty head with causal latents Training an error reader. We train an uncertainty head, a one-hidden-layer MLP UMLPU_MLP (256→128→1256→ 128→ 1, GELU), on the same frozen representation h~i h_i, leaving both the backbone and direction head unchanged. It predicts ui=log(1+�/ideg).u_i= (1+ _i/deg ). (20) We also trained a simple linear regressor on the same target, which reaches comparable performance, indicating that most of the error-relevant information is already linearly accessible in the representation. All causal interventions use UMLPU_MLP. UMLPU_MLP reaches a Spearman correlation of 0.600.60 with the true error on the independent validation split. As an additional ranking test, we use the predicted error to distinguish events in the highest and lowest quartiles of true angular error. The MLP reaches an AUROC of 0.9250.925, equal to a dedicated linear classifier trained on the full 256256-dimensional CLS representation and well above the 0.7870.787 obtained from a classifier using only pulse count, total charge, fauxf_aux, qauxq_aux, and zenith. The head recovers essentially all linearly available error information. And because that information is distributed rather than localized, the head must read broadly. It is the natural consumer of the atlas. Measuring intervention effects. To test whether the verified concepts we discovered in hih_i become causal for the uncertainty head, we repeat every intervention of Section 4.1. We measure the effects as the induced shift in the predicted log error, M1(G)=Ei[U(h~i(G))−U(h~i)],M_1(G)= E_i [U ( h^(G)_i )-U ( h_i ) ], (21) with U the uncertainty head. Because the target is a logarithm, M1M_1 has a direct multiplicative reading. Independent of an event’s baseline, an effect of M1M_1 corresponds to scaling (1+�/preddeg)(1+ _pred/deg) by eM1e^M_1. We judge effects, as before, by empirical exceedance of twenty matched control interventions and a fixed effect floor. The controls are latents matched by firing rate, with dense latents that contribute strongly to activation reconstruction excluded from the control pool when appropriate. We call a feature causal only when the intervention changes the predicted error by a meaningful amount, exceeds the effects of the matched controls, and moves the prediction in the expected direction. Appendix A.2 gives the exact thresholds. The clean latents become causal. As detailed in Table 4, the bright-clean latent becomes causal. Zeroing zbcz_bc gives M1=+0.31M_1=+0.31. This multiplies (1+�/preddeg) (1+ _pred/deg ) by approximately e0.31≃1.36e^0.31 1.36. The effect is far beyond its matched controls. We repeat the causal tests across all scales of the verified atlas. The clean pair gives M1=+0.32M_1=+0.32, the 194194-latent clean family gives M1=+0.78M_1=+0.78, and the dense clean core gives M1=+0.37M_1=+0.37. This clear causality holds across all dictionary seeds, with all effects well beyond their matched controls. The dense brightness-support coordinate zbsz_bs is causal for the uncertainty head as well. Unlike its V-shaped, non-monotone effect on the direction head, the uncertainty head’s predicted error relies monotonously on zbsz_bs in both event classes, see Figure 5, right (cross-seed details are in Appendix C.5). In contrast to zbcz_bc, we find that zauxz_aux remains inert for the uncertainty head. While events with a higher fraction of auxiliary pulses are more likely to be noisier, the uncertainty prediction does not causally depend on this latent, or on the broader auxiliary family. We see two plausible, non-exclusive explanations for this. First, clean and auxiliary activity are two ends of a similar underlying axis (see Section 3.1 and Figure 3). The bright-clean signature the head already relies on may carry enough of that shared structure to make a separate auxiliary-specific read-out redundant. Second, the head may still use auxiliary information, just spread across latents our family definition does not capture, which would mirror the error concept discussed in Section 3.2. The interventions run here cannot distinguish between these two explanations. Direction head Uncertainty head Intervention effect verdict seeds M1M_1 verdict seeds remove zbcz_bc +0.04∘+0.04 null 4/44/4 +0.31+0.31 causal 4/44/4 remove zauxz_aux +0.00∘+0.00 null 3/43/4 −0.059-0.059 null, sign only 3/43/4 remove clean pair +0.03∘+0.03 null 4/44/4 +0.32+0.32 causal 4/44/4 remove clean family −0.02∘-0.02 null 3/43/4 +0.78+0.78 causal 4/44/4 remove clean core +0.02∘+0.02 null 3/43/4 +0.37+0.37 causal 4/44/4 remove aux family −0.01∘-0.01 null 4/44/4 −0.085-0.085 null, sign only 4/44/4 scale zbsz_bs V-shaped support 4/44/4 monotone causal 3/43/4 Table 4: Intervention effects on the final layer for the direction head and the uncertainty head. The “seeds” columns report replication of the corresponding verdict across the four independently trained SAE dictionaries. Direction-head effects are measured as shifts of angular error in degrees. Uncertainty-head effects are measured as shifts of predicted log error, M1M_1 in Eq. (21). The causal basis becomes physically interpretable. The intervention results tie the uncertainty head directly to named sparse read-outs. Clean-side quality and brightness features that are inert for the direction head become causal for the uncertainty head, individually and in families, while the auxiliary family remains a controlled null. The Jacobian analysis gives a complementary view. In Figure 6, for the MLP uncertainty head, one aggregate sensitivity axis carries 78%78\% of the variance and is stable across the discovery and validation splits, with a principal angle of 0.4∘0.4 . Unlike the direction-head axis, it also captures a median of 77%77\% of each event’s own local sensitivity. This is a direction genuinely shared across events rather than an average over a mostly insensitive population. This means that the uncertainty head runs through one genuinely shared, low-dimensional mechanism across almost every event, while the direction head’s computation is fundamentally distributed and event-specific with no common bottleneck. This shared uncertainty-axis is not itself an latent or one of the principal components. As expected, it aligns most strongly with the angular-error probe, followed by event-quality probes. The SAE interventions therefore establish which named features the head uses, while the probe alignment identifies the broader physical quantity the uncertainty head relies on. Our validation protocol applied to the same final layer of the backbone gives opposite verdicts for two different heads. Represented but not used is not a property of a feature alone, but a relation between a representation and a specific downstream head. The physical read-outs predate direction fine-tuning. We finally compare the fine-tuned backbone with its pretrained-only checkpoint. Charge and detector geometry are already strongly recoverable before direction fine-tuning: total charge reaches R2=0.997R^2=0.997, compared with 0.9600.960 after fine-tuning, while depth and radius have comparable probe performance across the two checkpoints. Thus, physical structures are present and do not necessarily require fine-tuning. Instead, fine-tuning can partially erode unused directions. Such an analysis can reveal which useful structures are already present and help guide the choice of what to fine-tune, which layers to unfreeze, and which information should be preserved. Fine-tuning currently has no way to know that a concept it never explicitly needs might still be valuable to a reader trained later on the same representation, so nothing in the objective protects it. Since our protocol can track a concept’s matched-control quality across checkpoints, the same procedure could in principle run during fine-tuning itself to keep concepts intact for downstream heads that do not exist yet. We do not test this here, but the erosion we observe is direct evidence that fine-tuning is not neutral with respect to information represented, which is exactly the kind of blind spot mechanistic interpretability is positioned to catch. 4.4 Event selection and calibration The uncertainty head turns physical information underused by the direction head, into a per-event estimate of angular reconstruction error. We now ask whether this estimate can identify a sample with improved effective angular resolution. For each selection variable, we orient the score sis_i so that larger values indicate better-reconstructed events. At a target efficiency �ε, the figure of merit is the median angular error of the retained sample, R(�;s)=mediani∈S�(s)�,iR(ε;s)=median_\,i∈ S_ε(s) _i, (22) evaluated on the held-out causal_test split. We compare the uncertainty head with detector observables, the atlas feature zbcz_bc, and distributed supervised read-outs of the same frozen representation. Figure 7-left shows the resulting resolution-efficiency curves. Without selection, the median angular error is 51.3∘51.3 . Among the detector-level observables, total charge gives the strongest selection: retaining the highest-charge 20%20\% of events reduces the median to 20.2∘20.2 . The representation-derived uncertainty estimates are substantially sharper. At the same efficiency, the linear and MLP heads reach 3.16∘3.16 and 3.15∘3.15 , respectively. At 50%50\% efficiency they give 11.32∘11.32 and 11.24∘11.24 , compared with 37.72∘37.72 for total charge. Their performance is also close to a linear probe trained on the full CLS representation to distinguish high- from low-error events: the probe reaches 3.41∘3.41 at 20%20\% efficiency and marginally leads at 50%50\%, with 11.14∘11.14 . The comparison with zbcz_bc clarifies the distinct roles of sparse interpretation and distributed prediction. Although zbcz_bc is a validated physical read-out, it is not an effective standalone error-ranking score. It fires on only ∼14% 14\% of events, so quantile thresholds targeting larger efficiencies fall on its zero-activation plateau and return the unselected sample. In this analysis, the sparse latents provide localized variables for physical validation and causal intervention, whereas accurate event ranking requires combining information distributed across the full summary. Figure 7-right shows the calibration curve of the uncertainty head. The true median error rises with the predicted error, with an event-level Spearman correlation of �s=0.615 _s=0.615. In the low-error region relevant for the tightest selections, the binned medians remain close to the diagonal, meaning that the uncertainty head is well-calibrated. At intermediate predicted errors, the uncertainty head is overconfident, and the broad within-bin spread shows substantial event-to-event variation. The head is therefore well ordered and approximately calibrated in the most-certain regime, but its raw outputs should not be interpreted as globally calibrated event-by-event uncertainties without an additional calibration step. Figure 7: Left: median angular error of the retained sample as a function of selection efficiency for detector-level observables, distributed representation read-outs, the sparse feature zbcz_bc, and the uncertainty head. The horizontal line gives the unselected resolution, and bands show bootstrap intervals. Right: median true angular error in bins of the error predicted by UlinU_lin and UMLPU_MLP. The diagonal indicates perfect calibration, while the shaded regions show the within-bin spread of true angular errors. Together, these results show that the uncertainty head is both accurate and physically grounded. It ranks events nearly as well as a supervised probe built directly for the task, and Subsection 4.3 showed that this ranking causally depends on named, validated features rather than on opaque structure. 5 Outlook To our knowledge, we apply sparse-autoencoder-based mechanistic interpretability to a foundation model in particle physics for the first time. We ask what physical information is represented internally and which parts are used by downstream heads. The two questions have different answers. Across its layers, PolarBERT contains validated read-outs of event quality, auxiliary activity, brightness, and detector depth. The direction head shows marginal causal dependence on the validated physical read-outs across the dictionaries we trained, including functionally trained dictionaries. Its aggregate sensitivity is instead dominated by a stable causal axis with no simple physical interpretation. What the model demonstrably represents and what the direction task uses are almost disjoint. After this diagnosis, we investigate whether a more complex task could use this atlas. An uncertainty head, trained on the same event-level representation to predict the model’s own error reads the atlas that the direction head did not use. Intuitively, direction-reconstruction uncertainty should depend on how bright and clean an event is [7, 6]. Our results confirm this expectation. The SAE latents representing how clean and bright an event is are causally inert for the direction head and become causal for the uncertainty head across all four independently trained dictionaries, and the resulting estimator is a sharp selector, not just an interpretable one. At 20%20\% selection efficiency, the best single detector observable reaches the 20∘20 scale, while the uncertainty reader reaches 3.2∘3.2 . A black-box regressor trained on the same target might reach similar numbers, but its inputs would be opaque. Because we isolate, match-control and test by intervention the concept we are able to identify, we postulate that the uncertainty head relies on physics terms that can justify its selection rather than taken on faith. This strongly matters once working on real data points, where the correctness of a cut cannot be checked against Monte Carlo truth. We thus show that represented but not used is a relation between a shared representation and a particular downstream task. This is especially relevant for foundation models, where multiple objectives can reuse the same representations. Sparse dictionaries can reveal which physical information is available in that shared representation and which parts remain unused by a given head. This information can then guide the design of new downstream tasks or learning objectives that explicitly exploit the available structure. In this sense, mechanistic interpretability is not only a diagnostic tool, but also a way to inform how foundation-model representations can be used. We test our protocol on Monte Carlo with each step of the validation protocol extending to real detector data, where truth labels are incomplete and the atlas must stand on controls and replication alone. Natural next steps are to carry the uncertainty score into full likelihood analyses, repeat the interpretability study with models trained on more event-level labels such as morphology, and extend it to other particle-physics foundation models. Acknowledgements We thank Antonin Vacheret for suggesting the application of mechanistic interpretability to neutrino foundation models, and Jean-Loup Tastet for guidance on the use of Polar-BERT. Part of this work was supported by the European Union’s Horizon Europe research and innovation programme under the Marie Sklodowska-Curie grant agreement No 101168829, Challenging AI with Challenges from Physics: How to solve fundamental problems in Physics by AI and vice versa (AIPHY). We thank the CINECA Leonardo HPC cluster (via INFN grant PML4HEP) and the University of Copenhagen’s SCIENCE AI Centre for providing computing resources. References [1] I. C. M.G. Aartsen, M. Ackermann, J. H. Adams, J. A. Aguilar, M. Ahlers, M. Ahrens, D. Altmann, K. Andeen, T. Anderson, I. Ansseau, G. Anton, M. Archinger, C. A. Arguelles, R. Auer, J. Auffenberg, S. N. Axani, J. Baccus, X. Bai, S. Barnet, S. W. Barwick, V. Baum, R. Bay, K. Beattie, J. J. Beatty, J. K. B. Tjus, K. H. Becker, T. Bendfelt, S. Y. Benzvi, D. Berley, E. Bernardini, A. Bernhard, D. Z. Besson, G. Binder, D. Bindig, M. Bissok, E. Blaufuss, S. Blot, D. J. Boersma, C. Bohm, M. Borner, F. Bos, D. Bose, S. Boser, O. Botner, A. Bouchta, J. Braun, L. Brayeur, H. Bretz, S. Bron, A. Burgman, C. Burreson, T. Carver, M. Casier, E. Cheung, D. Chirkin, A. Christov, K. Clark, L. Classen, S. Coenders, G. H. Collin, J. M. Conrad, D. F. Cowen, R. Cross, C. T. Day, M. Day, J.P.A.M. de Andr’e, C. D. Clercq, E. del Pino Rosendo, H. Dembinski, S. D. Ridder, F. Descamps, P. Desiati, K. D. de Vries, G. de Wasseige, M. de With, T. DeYoung, J. C. Díaz-Vélez, V. di Lorenzo, H. Dujmovic, J. Dumm, M. Dunkman, B. Eberhardt, W. R. Edwards, T. Ehrhardt, B. Eichmann, P. Eller, S. Euler, P. A. Evenson, S. Fahey, A. R. Fazely, J. Feintzeig, J. Felde, K. Filimonov, C. Finley, S. Flis, C.-C. Fosig, A. Franckowiak, M. Frere, E. Friedman, T. Fuchs, T. K. Gaisser, J. Gallagher, L. M. Gerhardt, K. Ghorbani, W. Giang, L. E. Gladstone, T. Glauch, D. Glowacki, T. Glusenkamp, A. Goldschmidt, J. G. Gonzalez, D. Grant, Z. Griffith, L. Gustafsson, C. Haack, A. Hallgren, F. Halzen, E. V. Hansen, T. Hansmann, K. D. Hanson, J. Haugen, D. Hebecker, D. Heereman, K. Helbing, R. E. Hellauer, R. Heller, S. V. Hickford, J. Hignight, G. C. Hill, K. D. Hoffman, R. Hoffmann, K. Hoshina, F. Huang, M. E. Huber, P. O. Hulth, K. Hultqvist, S. In, M. Inaba, A. Ishihara, E. Jacobi, J. S. Jacobsen, G. S. Japaridze, M. Jeong, K. Jero, A. M. Jones, B. J. P. Jones, J. M. Joseph, W. Kang, A. Kappes, T. Karg, A. Karle, U. Katz, M. Kauer, A. Keivani, J. L. Kelley, J. Kemp, A. Kheirandish, J. H. Kim, M. S. Kim, T. Kintscher, J. Kiryluk, N. Kitamura, T. Kittler, S. R. Klein, S. Kleinfelder, M. von Kleist, G. Kohnen, R. Koirala, H. Kolanoski, R. Konietz, L. Kopke, C. Kopper, S. Kopper, D. J. Koskinen, M. Kowalski, M. Krasberg, K. M. Krings, M. Kroll, G. Kruckl, C. Kruger, J. Kunnen, S. Kunwar, N. Kurahashi, T. Kuwabara, M. L. M. Labare, K. Laihem, H. Landsman, J. Lanfranchi, M. J. Larson, F. Lauber, A. Laundrie, D. Lennarz, H. Leich, M. Lesiak-Bzdak, M. Leuermann, L. Lu, J. Ludwig, J. Lunemann, C. Mackenzie, J. Madsen, G. Maggi, K. B. M. Mahn, S. Mancina, M. Mandelartz, R. Maruyama, K. Mase, H. S. Matis, R. Maunu, F. McNally, C. P. Mcparland, P. Meade, K. J. Meagher, M. Medici, M. Meier, A. Meli, T. Menne, G. Merino, T. Meures, S. Miarecki, R. H. Minor, T. Montaruli, M. Moulai, T. Murray, R. Nahnhauer, U. Naumann, G. L. Neer, M. G. Newcomb, H. Niederhausen, S. C. Nowicki, D. R. Nygren, A. O. Pollmann, A. R. Olivas, A. O’Murchadha, T. Palczewski, H. Pandya, D. Pankova, S. J. Patton, P. Peiffer, O. Penek, J. A. Pepper, C. P. de los Heros, C. Pettersen, D. Pieloth, E. Pinat, P. B. Price, G. T. Przybylski, M. Quinnan, C. Raab, L. Radel, M. Rameez, K. Rawlins, R. Reimann, B. Relethford, M. Relich, E. Resconi, W. Rhode, M. Richman, B. Riedel, S. M. Robertson, M. Rongen, C. Roucelle, C. Rott, T. Ruhe, D. Ryckbosch, D. Rysewyk, L. Sabbatini, S. E. S. Herrera, A. Sandrock, J. Sandroos, P. Sandstrom, S. Sarkar, K. Satalecka, P. Schlunder, T. Schmidt, S. Schoenen, S. Schoneberg, A. Schukraft, L. Schumacher, D. Seckel, S. Seunarine, M. Solarz, D. Soldin, M. Song, G. M. Spiczak, C. Spiering, T. Stanev, A. Stasik, J. Stettner, A. Steuer, T. Stezelberger, R. G. Stokstad, A. Stossl, R. G. Strom, N. L. Strotjohann, K-H. Sulanke, G. W. Sullivan, M. S. Sutherland, H. Taavola, I. Taboada, J. Tatar, F. Tenholt, S. Ter-Antonyan, A. Terliuk, G. Tevsi’c, L. Thollander, S. Tilav, P. A. Toale, M. N. Tobin, S. Toscano, D. Tosi, M. Tselengidou, A. Turcati, E. Unger, M. Usner, J. Vandenbroucke, N. van Eijndhoven, S. Vanheule, M. van Rossem, J. van Santen, M. Vehring, M. Voge, E. Vogel, M. Vraeghe, D. Wahl, C. Walck, A. Wallace, M. Wallraff, N. Wandkowsky, C. N. Weaver, M. J. Weiss, C. H. Wendt, S. Westerhoff, D. Wharton, B. J. Whelan, S. Wickmann, K. Wiebe, C. H. Wiebusch, L. Wille, D. R. Williams, L. Wills, P. Wisniewski, M. Wolf, T. R. Wood, E. Woolsey, K. Woschnagg, D. L. Xu, X. W. Xu, Y. Xu, J. P. Yáñez, G. B. Yodh, S. Yoshida, and M. Zoll (2017) The icecube neutrino observatory: instrumentation and online systems. JINST 12 (03), p. P03012. Note: [Erratum: JINST 19, E05001 (2024)] External Links: 1612.05093, Document Cited by: §1. [2] M. G. Aartsen, R. U. Abbasi, M. Ackermann, J. Adams, J. A. Aguilar, M. Ahlers, D. Altmann, C. A. Arguelles, J. Auffenberg, X. Bai, M. Baker, S. W. Barwick, V. Baum, R. Bay, J. J. Beatty, J. K. B. Tjus, K. Becker, S. Y. Benzvi, P. Berghaus, D. Berley, E. Bernardini, A. Bernhard, D. Z. Besson, G. Binder, D. Bindig, M. Bissok, E. Blaufuss, J. Blumenthal, D. J. Boersma, C. Bohm, D. Bose, S. Böser, O. Botner, L. Brayeur, H. Bretz, A. M. Brown, R. Bruijn, J. Casey, M. Casier, D. Chirkin, A. Christov, B. J. Christy, K. Clark, L. Classen, F. Clevermann, S. Coenders, S. Cohen, D. F. Cowen, A. H. C. Silva, M. Danninger, J. D. Daughhetee, J. C. Davis, M. Day, C. D. Clercq, S. D. Ridder, P. Desiati, K. D. de Vries, M. D. With, T. DeYoung, J. C. Díaz-Vélez, M. Dunkman, R. Eagan, B. Eberhardt, B. Eichmann, J. Eisch, S. Euler, P. A. Evenson, O. O. Fadiran, A. R. Fazely, A. Fedynitch, J. Feintzeig, T. Feusels, K. Filimonov, C. Finley, T. Fischer-Wasels, S. Flis, A. Franckowiak, K. Frantzen, T. Fuchs, T. K. Gaisser, J. Gallagher, L. M. Gerhardt, L. E. Gladstone, T. Glüsenkamp, A. Goldschmidt, G. Golup, J. G. Gonzalez, J. A. Goodman, D. Góra, D. T. Grandmont, D. Grant, P. Gretskov, J. C. Groh, A. Groß, C. H. Ha, A. A. K. H. Ismail, P. Hallen, A. Hallgren, F. Halzen, K. D. Hanson, D. Hebecker, D. Heereman, D. Heinen, K. Helbing, R. E. Hellauer, S. V. Hickford, G. C. Hill, K. D. Hoffman, R. Hoffmann, A. Homeier, K. Hoshina, F. Huang, W. Huelsnitz, P. O. Hulth, K. Hultqvist, S. Hussain, A. Ishihara, S. Jackson, E. Jacobi, J. S. Jacobsen, K. Jagielski, G. S. Japaridze, K. Jero, O. Jlelati, B. Kaminsky, A. Kappes, T. Karg, A. Karle, M. Kauer, J. L. Kelley, J. Kiryluk, J. Kläs, S. R. Klein, J. Köhne, G. Kohnen, H. Kolanoski, L. Köpke, C. Kopper, S. Kopper, D. J. Koskinen, M. Kowalski, M. Krasberg, A. Kriesten, K. M. Krings, G. Kroll, J. Kunnen, N. Kurahashi, T. Kuwabara, M. L. M. Labare, H. Landsman, M. J. Larson, M. Lesiak-Bzdak, M. Leuermann, J. Leute, J. Lünemann, O. Macias, J. Madsen, G. P. Maggi, R. Maruyama, K. Mase, H. S. Matis, F. McNally, K. J. Meagher, Merck. Merck, T. Meures, S. Miarecki, E. Middell, N. Milke, J. Miller, L. Mohrmann, T. Montaruli, R. M. Morse, R. Nahnhauer, U. Naumann, H. Niederhausen, S. C. Nowicki, D. R. Nygren, A. Obertacke, S. Odrowski, A. Olivas, A. Omairat, A. O’Murchadha, L. Paul, J. Pepper, C. P. de los Heros, C. Pfendner, D. Pieloth, E. Pinat, J. Posselt, P. B. Price, G. T. Przybylski, M. Quinnan, L. Rädel, M. Rameez, K. Rawlins, P. C. Redl, R. Reimann, E. Resconi, W. Rhode, M. Ribordy, M. Richman, B. Riedel, S. M. Robertson, J. P. Rodrigues, C. Rott, T. Ruhe, B. Ruzybayev, D. Ryckbosch, S. M. Saba, H. G. Sander, M. Santander, S. Sarkar, K. Schatto, F. Scheriau, T. W. Schmidt, M. Schmitz, S. Schoenen, S. Schöneberg, A. Schönwald, A. Schukraft, L. Schulte, O. Schulz, D. Seckel, Y. Sestayo, S. Seunarine, R. Shanidze, C. Sheremata, M. W. E. Smith, D. Soldin, G. M. Spiczak, C. Spiering, M. Stamatikos, T. Stanev, N. A. Stanisha, A. Stasik, T. Stezelberger, R. G. Stokstad, A. Stößl, E. A. Strahler, R. Ström, N. L. Strotjohann, G. W. Sullivan, H. Taavola, I. Taboada, A. Tamburro, A. Tepe, S. Ter-Antonyan, G. Tesic, S. Tilav, P. A. Toale, M. N. Tobin, S. Toscano, M. Tselengidou, E. Unger, M. Usner, S. Vallecorsa, N. van Eijndhoven, A. V. Overloop, J. van Santen, M. Vehring, M. Voge, M. Vraeghe, C. Walck, T. Waldenmaier, M. Wallraff, C. N. Weaver, M. Wellons, C. H. Wendt, S. Westerhoff, B. J. Whelan, N. Whitehorn, K. Wiebe, C. H. Wiebusch, D. R. Williams, H. Wissing, M. Wolf, T. R. Wood, K. Woschnagg, D. L. Xu, X. W. Xu, J. P. Yáñez, G. B. Yodh, S. Yoshida, P. Zarzhitsky, J. Ziemann, S. Zierke, and M. Zoll (2014) Energy reconstruction methods in the icecube neutrino telescope. JINST 9, p. P03009. External Links: 1311.4767, Document Cited by: §1. [3] R. Abbasi, M. Ackermann, J. Adams, M. Ahlers, J. Ahrens, K. Andeen, J. Auffenberg, X. Bai, M. Baker, S.W. Barwick, R. Bay, J.L. Bazo Alba, K. Beattie, T. Becka, J.K. Becker, K.-H. Becker, P. Berghaus, D. Berley, E. Bernardini, D. Bertrand, D.Z. Besson, B. Bingham, E. Blaufuss, D.J. Boersma, C. Bohm, J. Bolmont, S. Böser, O. Botner, J. Braun, D. Breeder, T. Burgess, W. Carithers, T. Castermans, H. Chen, D. Chirkin, B. Christy, J. Clem, D.F. Cowen, M.V. D’Agostino, M. Danninger, A. Davour, C.T. Day, O. Depaepe, C. De Clercq, L. Demirörs, F. Descamps, P. Desiati, G. de Vries-Uiterweerd, T. DeYoung, J.C. Diaz-Velez, J. Dreyer, J.P. Dumm, M.R. Duvoort, W.R. Edwards, R. Ehrlich, J. Eisch, R.W. Ellsworth, O. Engdegård, S. Euler, P.A. Evenson, O. Fadiran, A.R. Fazely, T. Feusels, K. Filimonov, C. Finley, M.M. Foerster, B.D. Fox, A. Franckowiak, R. Franke, T.K. Gaisser, J. Gallagher, R. Ganugapati, L. Gerhardt, L. Gladstone, D. Glowacki, A. Goldschmidt, J.A. Goodman, R. Gozzini, D. Grant, T. Griesel, A. Groß, S. Grullon, R.M. Gunasingha, M. Gurtner, C. Ha, A. Hallgren, F. Halzen, K. Han, K. Hanson, R. Hardtke, Y. Hasegawa, J. Haugen, D. Hays, J. Heise, K. Helbing, M. Hellwig, P. Herquet, S. Hickford, G.C. Hill, J. Hodges, K.D. Hoffman, K. Hoshina, D. Hubert, W. Huelsnitz, B. Hughey, J.-P. Hülß, P.O. Hulth, K. Hultqvist, S. Hussain, R.L. Imlay, M. Inaba, A. Ishihara, J. Jacobsen, G.S. Japaridze, H. Johansson, A. Jones, J.M. Joseph, K.-H. Kampert, A. Kappes, T. Karg, A. Karle, H. Kawai, J.L. Kelley, J. Kiryluk, F. Kislat, S.R. Klein, S. Kleinfelder, S. Klepser, G. Kohnen, H. Kolanoski, L. Köpke, M. Kowalski, T. Kowarik, M. Krasberg, K. Kuehn, E. Kujawski, T. Kuwabara, M. Labare, K. Laihem, H. Landsman, R. Lauer, A. Laundrie, H. Leich, D. Leier, C. Lewis, A. Lucke, J. Ludvig, J. Lundberg, J. Lünemann, J. Madsen, R. Maruyama, K. Mase, H.S. Matis, C.P. McParland, K. Meagher, A. Meli, M. Merck, T. Messarius, P. Mészáros, R.H. Minor, H. Miyamoto, A. Mohr, A. Mokhtarani, T. Montaruli, R. Morse, S.M. Movit, K. Münich, A. Muratas, R. Nahnhauer, J.W. Nam, P. Nießen, D.R. Nygren, S. Odrowski, A. Olivas, M. Olivo, M. Ono, S. Panknin, S. Patton, C. Pérez de los Heros, J. Petrovic, A. Piegsa, D. Pieloth, A.C. Pohl, R. Porrata, N. Potthoff, J. Pretz, P.B. Price, G.T. Przybylski, K. Rawlins, S. Razzaque, P. Redl, E. Resconi, W. Rhode, M. Ribordy, A. Rizzo, W.J. Robbins, J.P. Rodrigues, P. Roth, F. Rothmaier, C. Rott, C. Roucelle, D. Rutledge, D. Ryckbosch, H.-G. Sander, S. Sarkar, K. Satalecka, P. Sandstrom, S. Schlenstedt, T. Schmidt, D. Schneider, O. Schulz, D. Seckel, B. Semburg, S.H. Seo, Y. Sestayo, S. Seunarine, A. Silvestri, A.J. Smith, C. Song, J.E. Sopher, G.M. Spiczak, C. Spiering, T. Stanev, T. Stezelberger, R.G. Stokstad, M.C. Stoufer, S. Stoyanov, E.A. Strahler, T. Straszheim, K.-H. Sulanke, G.W. Sullivan, Q. Swillens, I. Taboada, O. Tarasova, A. Tepe, S. Ter-Antonyan, S. Tilav, M. Tluczykont, P.A. Toale, D. Tosi, D. Turčan, N. van Eijndhoven, J. Vandenbroucke, A. Van Overloop, V. Viscomi, C. Vogt, B. Voigt, C.Q. Vu, D. Wahl, C. Walck, T. Waldenmaier, H. Waldmann, M. Walter, C. Wendt, S. Westerhof, N. Whitehorn, D. Wharton, C.H. Wiebusch, C. Wiedemann, G. Wikström, D.R. Williams, R. Wischnewski, H. Wissing, K. Woschnagg, X.W. Xu, G. Yodh, and S. Yoshida (2009) The icecube data acquisition system: signal capture, digitization, and timestamping. Nuclear Instruments and Methods in Physics Research Section A: Accelerators, Spectrometers, Detectors and Associated Equipment 601 (3), p. 294–316. External Links: ISSN 0168-9002, Link, Document Cited by: §1, §2.1. [4] R. Abbasi et al. (2009) The IceCube Data Acquisition System: Signal Capture, Digitization, and Timestamping. Nucl. Instrum. Meth. A 601, p. 294–316. External Links: 0810.4930, Document Cited by: §2.1. [5] R. Abbasi et al. (2012) Observation of an Anisotropy in the Galactic Cosmic Ray arrival direction at 400 TeV with IceCube. Astrophys. J. 746, p. 33. External Links: 1109.1017, Document Cited by: §3.1. [6] R. Abbasi et al. (2022) Low energy event reconstruction in IceCube DeepCore. Eur. Phys. J. C 82 (9), p. 807. External Links: 2203.02303, Document Cited by: §5. [7] R. Abbasi et al. (2026) Neural posterior estimation of the neutrino direction in IceCube using transformer-encoded normalizing flows on the sphere. External Links: 2604.19846 Cited by: §5. [8] G. Alain and Y. Bengio (2016) Understanding intermediate layers using linear classifier probes. arXiv e-prints, p. arXiv:1610.01644. External Links: Document, 1610.01644 Cited by: §2.3. [9] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Q. Davis, D. Demszky, C. Donahue, M. Doumbouya, E. Durmus, S. Ermon, J. Etchemendy, K. Ethayarajh, L. Fei-Fei, C. Finn, T. Gale, L. Gillespie, K. Goel, N. Goodman, S. Grossman, N. Guha, T. Hashimoto, P. Henderson, J. Hewitt, D. E. Ho, J. Hong, K. Hsu, J. Huang, T. Icard, S. Jain, D. Jurafsky, P. Kalluri, S. Karamcheti, G. Keeling, F. Khani, O. Khattab, P. W. Koh, M. Krass, R. Krishna, R. Kuditipudi, A. Kumar, F. Ladhak, M. Lee, T. Lee, J. Leskovec, I. Levent, X. L. Li, X. Li, T. Ma, A. Malik, C. D. Manning, S. Mirchandani, E. Mitchell, Z. Munyikwa, S. Nair, A. Narayan, D. Narayanan, B. Newman, A. Nie, J. C. Niebles, H. Nilforoshan, J. Nyarko, G. Ogut, L. Orr, I. Papadimitriou, J. S. Park, C. Piech, E. Portelance, C. Potts, A. Raghunathan, R. Reich, H. Ren, F. Rong, Y. Roohani, C. Ruiz, J. Ryan, C. Ré, D. Sadigh, S. Sagawa, K. Santhanam, A. Shih, K. Srinivasan, A. Tamkin, R. Taori, A. W. Thomas, F. Tramèr, R. E. Wang, W. Wang, B. Wu, J. Wu, Y. Wu, S. M. Xie, M. Yasunaga, J. You, M. Zaharia, M. Zhang, T. Zhang, X. Zhang, Y. Zhang, L. Zheng, K. Zhou, and P. Liang (2022) On the opportunities and risks of foundation models. External Links: 2108.07258, Link Cited by: §1. [10] D. Braun, J. K. Taylor, N. Goldowsky-Dill, and L. Sharkey (2024) Identifying functionally important features with end-to-end sparse dictionary learning. ArXiv abs/2405.12241. External Links: Link Cited by: §1, §4.1. [11] T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah (2023) Towards monosemanticity: decomposing language models with dictionary learning. Transformer Circuits Thread. Note: https://transformer-circuits.pub/2023/monosemantic-features/index.html Cited by: §1. [12] H. Bukhari, D. Chakraborty, P. Eller, T. Ito, M. V. Shugaev, and R. Ørsøe (2024) IceCube – Neutrinos in Deep Ice: The top 3 solutions from the public Kaggle competition. Eur. Phys. J. C 84 (6), p. 646. External Links: 2310.15674, Document Cited by: §2.1. [13] B. Bussmann, P. Leask, and N. Nanda (2024) BatchTopK sparse autoencoders. ArXiv abs/2412.06410. External Links: Link Cited by: §1, §2.2. [14] N. Cammarata, S. Carter, G. Goh, C. Olah, M. Petrov, L. Schubert, C. Voss, B. Egan, and S. K. Lim (2020) Thread: circuits. Distill. Note: https://distill.pub/2020/circuits External Links: Document Cited by: §1. [15] A. Chow, L. Heinrich, P. Eller, R. Ørsøe, and S. Dane (2023) IceCube - neutrinos in deep ice. Note: https://kaggle.com/competitions/icecube-neutrinos-in-deep-iceKaggle Cited by: §1, §2.1. [16] H. Cunningham, A. Ewart, L. R. Smith, R. Huben, and L. Sharkey (2023) Sparse autoencoders find highly interpretable features in language models. ArXiv abs/2309.08600. External Links: Link Cited by: §1. [17] J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) BERT: pre-training of deep bidirectional transformers for language understanding. External Links: 1810.04805, Link Cited by: §1, §2.1. [18] N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah (2022) Toy models of superposition. External Links: 2209.10652, Link Cited by: §1. [19] L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu (2024) Scaling and evaluating sparse autoencoders. ArXiv abs/2406.04093. External Links: Link Cited by: §1, §2.2, §2.2. [20] T. Golling, L. Heinrich, M. Kagan, S. Klein, M. Leigh, M. Osadchy, and J. A. Raine (2024) Masked particle modeling on sets: towards self-supervised high energy physics foundation models. Mach. Learn. Sci. Tech. 5 (3), p. 035074. External Links: 2401.13537, Document Cited by: §1. [21] S. Kantamneni, J. Engels, S. Rajamanoharan, M. Tegmark, and N. Nanda (2025) Are sparse autoencoders useful? a case study in sparse probing. ArXiv abs/2502.16681. External Links: Link Cited by: §1. [22] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020) Scaling laws for neural language models. External Links: 2001.08361, Link Cited by: §1. [23] A. Karvonen, C. Rager, J. Lin, C. Tigges, J. Bloom, D. Chanin, Y. Lau, E. Farrell, C. S. McDougall, K. Ayonrinde, M. Wearden, A. Conmy, S. Marks, and N. Nanda (2025) SAEBench: a comprehensive benchmark for sparse autoencoders in language model interpretability. ArXiv abs/2503.09532. External Links: Link Cited by: §2.2. [24] U. Katz and C. Spiering (2011) High-energy neutrino astrophysics: status and perspectives. Progress in Particle and Nuclear Physics 67, p. 651–704. External Links: Link Cited by: §1. [25] T. Kishimoto, M. Morinaga, M. Saito, and J. Tanaka (2023) Pre-training strategy using real particle collision data for event classification in collider physics. In 37th Conference on Neural Information Processing Systems, External Links: 2312.06909 Cited by: §1. [26] A. R. Kulkarni, T. Weng, V. S. Narayanaswamy, S. Liu, W. A. Sakla, and K. Thopalli (2025) Interpretable and steerable concept bottleneck sparse autoencoders. ArXiv abs/2512.10805. External Links: Link Cited by: §1. [27] P. Patel and S. Ganguly (2026) Explainable ai for jet tagging: a comparative study of gnnexplainer, gnnshap, and gradcam for jet tagging in the lund jet plane. ArXiv abs/2604.25885. External Links: Link Cited by: §1. [28] G. Paulo and N. Belrose (2025) Sparse autoencoders trained on the same data learn different features. ArXiv abs/2501.16615. External Links: Link Cited by: §1. [29] S. Rai and S. Ganguly (2026) Dissecting Jet-Tagger Through Mechanistic Interpretability. External Links: 2605.09881 Cited by: §1. [30] G. Schwalbe and B. Finzel (2023) A comprehensive taxonomy for explainable artificial intelligence: a systematic survey of surveys on methods and concepts. Data Mining and Knowledge Discovery 38 (5), p. 3043–3101. External Links: ISSN 1573-756X, Link, Document Cited by: §1. [31] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra (2019) Grad-cam: visual explanations from deep networks via gradient-based localization. International Journal of Computer Vision 128 (2), p. 336–359. External Links: ISSN 1573-1405, Link, Document Cited by: §1. [32] L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. Bloom, S. Biderman, A. Garriga-Alonso, A. Conmy, N. Nanda, J. Rumbelow, M. Wattenberg, N. Schoots, J. Miller, E. J. Michaud, S. Casper, M. Tegmark, W. Saunders, D. Bau, E. Todd, A. Geiger, M. Geva, J. Hoogland, D. Murfet, and T. McGrath (2025) Open problems in mechanistic interpretability. External Links: 2501.16496, Link Cited by: §1. [33] K. Simonyan, A. Vedaldi, and A. Zisserman (2014) Deep inside convolutional networks: visualising image classification models and saliency maps. External Links: 1312.6034, Link Cited by: §1. [34] M. Sundararajan, A. Taly, and Q. Yan (2017) Axiomatic attribution for deep networks. External Links: 1703.01365, Link Cited by: §1. [35] I. Timiryasov, J. Tastet, and O. Ruchayskiy (2024) PolarBERT: a foundation model for icecube. In NeurIPS 2024 Workshop: Machine Learning and the Physical Sciences, Cited by: §1, §1, §2.1. [36] S. Vent, R. Winterhalder, and T. Plehn (2026) The Physics Behind ML-based Quark-Gluon Taggers. SciPost Phys. 20, p. 084. External Links: 2507.21214, Document Cited by: §1. [37] M. Vigl, N. Hartman, and L. Heinrich (2024) Finetuning foundation models for joint analysis optimization in High Energy Physics. Mach. Learn. Sci. Tech. 5 (2), p. 025075. External Links: 2401.13536, Document Cited by: §1. [38] M. Vigl, N. Hartman, M. Kagan, and L. Heinrich (2026) Neural Scaling Laws for Boosted Jet Tagging. ICLR 2026 Workshop FM4Science Poster. External Links: 2602.15781 Cited by: §1. Appendix A Supplementary analysis protocol This section justifies the representational site, states the fixed verdict rules, and defines the head-aligned sensitivity basis. The following sections then report additional atlas results, robustness checks, and detailed interventions in the same order as the main text. A.1 Choice of representational site To justify analyzing the final-block CLS state, we fit ridge probes (�=1α=1) from the CLS residual state at every depth to the true neutrino direction. Index 00 denotes the post-embedding state, while indices 11–88 denote the transformer-block outputs. We train the probes on concept_dev (69,632 events) and evaluate them on the disjoint concept_test split (59,648 events). Figure 8 shows that linearly decodable directional information improves monotonically with depth and plateaus from block 6 onward. The final-block probe reaches a mean angular error of 61.0∘61.0 , close to the 59.0∘59.0 obtained by the frozen nonlinear direction head on the same events. This supports using the final-block CLS state as the common site for the dictionary, probes, and head interventions. Figure 8: Held-out mean angular error of a linear CLS probe at each layer. Lower values are better; the frozen direction head is shown as a reference. A.2 Fixed read-out and causal criteria We fix each verdict rule before its corresponding held-out evaluation and retain partial or failed replications. For a target intervention statistic T, its control-normalized score is z(T)=T−mean(Tctrl)std(Tctrl),z(T)= T-mean(T_ctrl)std(T_ctrl), (23) where TctrlT_ctrl is obtained from the matched-control interventions. For direction-head tests, the selectivity statistic is �+−�− _+- _-, with the control statistic computed in the same way. Table 5 collects the rules used throughout the analysis. Stage Fixed rule Primary binary candidate Discovery AUROC ≥0.75≥ 0.75 and survival of every evaluable matched-control scheme Secondary read-out Reported separately when it survives the nuisance controls but remains below the primary AUROC bar Independent confirmation Re-evaluate the fixed candidate on concept_test, without reselection or threshold tuning Family membership Auxiliary-monotonicity score ams≤−0.85ams≤-0.85 for the clean side or ams≥+0.85ams≥+0.85 for the auxiliary side Direction-head causality �+≥0.05∘ _+≥ 0.05 , z(�+)≥2z( _+)≥ 2, and z(�+−�−)≥2z( _+- _-)≥ 2 Uncertainty-head causality |M1|≥0.02|M_1|≥ 0.02, z(M1)≥2z(M_1)≥ 2, and the pre-specified sign Removal controls 2020 firing-rate-matched latents or sets; multi-latent controls also match firing-rate composition Writing controls The same donor values as the target write, with firing-rate-matched recipient coordinates Density exclusion Control-pool exclusion for firing rate >0.30>0.30 or top-decile mean nonzero activation Table 5: Fixed discovery, validation, and intervention rules. Each row gives the condition used at that stage. The density rule applies to matched-control pools, not automatically to the physical families. For the within-clean brightness analysis, the control pool additionally excludes the 4343 charge-tracking latents defined by that analysis. Latent 10061006 is retained as a secondary clean-side read-out because it survives the nuisance controls despite falling below the primary discovery AUROC. A.3 Head-aligned sensitivity basis For event i, let v∈R256v∈ R^256 be a unit direction in the standardized CLS activation space. A small perturbation �vε v changes a frozen head to first order as g(h~i+�v)−g(h~i)≃�Jiv,g( h_i+ε v)-g( h_i) ε J_iv, (24) where JiJ_i is the head Jacobian. Thus, ∥Jiv∥ J_iv measures local sensitivity along v. We compute the Jacobians by automatic differentiation, verify them against finite differences, and aggregate them over n=8000n=8000 discovery-split events as M=1n∑i=1nJi⊤Ji.M= 1n _i=1^nJ_i J_i. (25) All Jacobians are expressed in the standardized CLS frame used by the SAE, probes, and PCA. Because trM=1n∑i∥Ji∥F2,trM= 1n _i J_i _F^2, (26) events for which the head is more locally responsive contribute more strongly without an explicit event weight. The eigenvectors of M order activation-space directions by their average effect on the head output, and the normalized eigenvalues give their share of aggregate sensitivity. We retain a direction only when it replicates on the independent split. Appendix B Supplementary atlas results B.1 Isolation of the named read-outs Table 6 reports the leading clean-trigger coordinates on the discovery split. Latent 640640 is the only primary candidate. Latent 10061006 remains below the primary AUROC bar but is retained as a secondary read-out because it survives the nuisance controls and concentrates clean events in its activation tail. The sparse candidates 751751, 946946, and 369369 fail the controls. Latent 973973 has the largest mean-activation contrast but is dense and non-selective, motivating its separate treatment as a support coordinate rather than a validated concept read-out. The auxiliary contrast yields one primary candidate, latent 195195. Its discovery AUROC is 0.8320.832, and its held-out AUROC is 0.8400.840. The evaluable nuisance controls survive, with a minimum AUROC of 0.8540.854. schemes that additionally match the non-auxiliary pulse count produce no usable pairs because that count is nearly deterministic for the extreme clean and auxiliary classes. The remaining candidates have either weak discrimination or high enrichment at negligible firing rates and are not promoted. Latent AUROC Top-1%1\% enrich. Firing rate Min. ctrl AUROC Status Positive class: clean-trigger events 640 0.9130.913 15.40×15.40× 0.1390.139 0.8180.818 primary 10061006 0.6980.698 14.94×14.94× 0.0870.087 0.6990.699 secondary 751751 0.6310.631 6.13×6.13× 0.1340.134 0.5500.550 rejected 946946 0.6010.601 15.18×15.18× 0.0140.014 0.5100.510 rejected 418418 0.5910.591 3.89×3.89× 0.1260.126 0.5900.590 rejected 369369 0.5890.589 6.51×6.51× 0.0810.081 0.5430.543 rejected 973973 0.5610.561 2.07×2.07× 0.4060.406 0.5610.561 support only Positive class: auxiliary-dominated events 195 0.8320.832 7.07×7.07× 0.4930.493 0.8270.827 primary 261261 0.6610.661 19.01×19.01× 0.0240.024 0.6420.642 rejected 551551 0.6040.604 2.19×2.19× 0.1610.161 0.5710.571 rejected 7474 0.5820.582 3.17×3.17× 0.0850.085 0.5740.574 rejected 115115 0.5430.543 8.53×8.53× 0.0070.007 0.5390.539 rejected 289289 0.5160.516 3.90×3.90× 0.0030.003 0.5120.512 rejected Table 6: Discovery-split rankings for the clean-trigger and auxiliary-dominated contrasts. AUROC measures event-level discrimination, top-1%1\% enrichment is the prevalence of the positive class in the highest-activation tail relative to its overall prevalence, firing rate is P(zj>0)P(z_j>0), and “Min. ctrl AUROC” is the lowest AUROC across the matched-control schemes. Status records the role assigned after applying the fixed discovery criteria. B.2 Latent-family definitions Family Definition n clean_pair 640, 1006\640,\,1006\ 22 clean_family All ams≤−0.85ams≤-0.85, excluding the brightness-support coordinate zbsz_bs 194194 clean_core Members of clean_family with firing rate ≥0.05≥ 0.05 55 aux_family All ams≥+0.85ams≥+0.85 135135 Table 7: Reference-dictionary families. The final column gives the number of included latents. The family interventions test whether a single-latent null could arise because a concept is distributed across several coordinates with similar response profiles. Table 7 gives the reference-dictionary definitions used in these tests. B.3 PCA comparison Table 8 compares PC10, zbcz_bc, zauxz_aux, and zbsz_bs on the same validation events. For clean events the SAE’s sparse latent zbcz_bc is a sharper discriminant than PC10, while for auxiliary-dominated events, zauxz_aux and PC10 perform about the same. For for high-charge events, PC10 clearly outperforms zbsz_bs because zbsz_bs tracks charge only in its population-level average and is a poor per-event classifier, which is exactly why we do not treat it as a validated brightness read-out. Positive class Score Test AUROC Control AUROC clean PC1010 0.8460.846 0.7270.727 clean zbcz_bc 0.9110.911 0.7870.787 auxiliary PC1010 0.8470.847 0.8450.845 auxiliary zauxz_aux 0.8400.840 0.8540.854 high charge PC1010 0.8480.848† † High/low charge is a binary contrast between the top and bottom quintiles of total charge. 0.7500.750 high charge zbsz_bs (973973) 0.5320.532 ‡ ‡ zbsz_bs is strongly charge-monotonic in population means but has near-chance per-event high/low-charge discrimination. This is why we describe it as brightness-associated rather than as a validated brightness read-out. 0.5160.516 Table 8: PCA and SAE scores evaluated on the independent concept_test split. Control AUROC is the strictest matched-control value for each target. B.4 Distributed and exploratory physical information B.4.1 Morphology proxies To test whether linearly decodable morphology information is concentrated in a few SAE coordinates, we fit k-sparse probes to time extent, linefit speed, linearity, charge concentration, and vertical extent. A LARS path fitted on concept_dev selects the k coordinates, after which a ridge model is refitted on those columns and evaluated on concept_test. Dense ridge probes on the raw CLS state and on all SAE coordinates provide reference ceilings. Layer 8 Layer 3 Target CLS SAE k=1k=1 k90%k_90\% CLS SAE k=1k=1 k90%k_90\% linefit speed 0.480.48 0.420.42 0.160.16 128128 0.610.61 0.530.53 0.140.14 6464 linearity 0.520.52 0.530.53 0.120.12 128128 0.490.49 0.520.52 0.100.10 128128 time extent 0.780.78 0.780.78 0.360.36 3232 0.880.88 0.880.88 0.430.43 1616 top-5 charge frac. 0.600.60 0.510.51 0.050.05 256256 0.730.73 0.640.64 0.080.08 6464 z RMS 0.670.67 0.670.67 0.100.10 256256 0.740.74 0.710.71 0.190.19 6464 Table 9: Held-out morphology-probe results at layers 8 and 3. Within each layer, CLS and SAE are the R2R^2 values of dense ridge probes using the raw CLS representation and all SAE latents, respectively. The k=1k=1 column gives the best single-latent R2R^2, while k90%k_90\% is the smallest number of selected latents needed to reach 90%90\% of the dense-SAE R2R^2. Table 9 shows that time extent is moderately sparse, whereas the track-likeness and charge-concentration proxies require tens to hundreds of coordinates. The selected k values are upper bounds rather than exact circuit sizes because the selection path is greedy. Thus, the absence of a validated single morphology latent does not imply that morphology information is absent from the representation, however it is mostly distributed across the dictionary. B.4.2 Layer-dependent depth read-outs Layer Latent Depth AUROC Min. ctrl AUROC Direction-head �+ _+ 22 171171 0.9860.986 0.9840.984 +0.0064∘+0.0064 33 191191 0.9250.925 0.8990.899 −0.009∘-0.009 88 none retained 0.5980.598 — — Table 10: Depth isolation across layers. “Min. ctrl AUROC” is the lowest AUROC across the matched-control schemes, and �+ _+ is the direction-head angular-error shift after removing the selected latent. At layer 8, 0.5980.598 is the best raw AUROC, but no latent passed the isolation criteria. The final-layer dictionary contains no validated depth coordinate (Table 10), but this depends on the representational site. Repeating the same isolation procedure at layers 2 and 3 recovers strong depth read-outs, with the layer-2 latent reaching AUROC 0.9860.986 and remaining essentially unchanged under matched controls. Neither early-layer coordinate passes the direction-head causal criteria. Geometry is therefore sharply localized at an earlier CLS site without becoming a direction-head bottleneck. B.4.3 Exploratory geometry profiles We also search for latents whose activation changes with event direction, without first selecting candidates using a binary concept contrast. For azimuth, we divide the horizontal charge-centroid direction �c _c into 2424 bins and compute the mean activation of each latent in every bin. We then determine whether the resulting profile has one broad preferred direction (k=1k=1), two opposite preferred directions (k=2k=2), or the sixfold pattern expected from IceCube’s hexagonal string layout (k=6k=6). Of the 9191 latents active enough for this test, 6161 vary across azimuth with an anisotropy of at least 0.150.15. Among them, 5252 prefer one broad direction and 99 show a two-direction pattern. None shows the expected sixfold detector pattern. The observed azimuthal dependence is therefore more likely to reflect the non-uniform distribution of events in the sample than the hexagonal detector geometry. We perform a similar scan in zenith by measuring how each latent’s mean activation changes with cos�z _z. We find 207207 latents with a strong monotonic response, |�Spearman|≥0.85| _Spearman|≥ 0.85. However, the strongest responses come from dense latents that contribute broadly to reconstruction, and no individual latent survives the matched-control validation test. These direction-dependent latents are therefore reported as exploratory candidates rather than validated physical read-outs. B.5 Robustness across dictionaries We repeat the read-out analysis across four independently trained dictionaries. Table 11 groups functional analogues by role rather than by latent index. The clean-primary role passes every control scheme in three of four draws. The auxiliary role is less stable: only the reference seed passes all evaluable controls, and seed 3 yields no isolated auxiliary candidate. In contrast, the dense, charge-monotonic, non-selective support phenotype appears in every draw. Role Quantity Seed 1 Seed 2 Seed 3 Seed 4 Clean primary latent 640640 754754 617617 947947 test AUROC 0.9110.911 0.8740.874 0.9120.912 0.9960.996 min. ctrl AUROC 0.7870.787 0.7400.740 0.7810.781 0.9870.987 all ctrl schemes Yes No Yes Yes Clean secondary latent 10061006 662662 918918 808808 test AUROC 0.6940.694 0.7440.744 0.6910.691 0.6790.679 min. ctrl AUROC 0.8370.837 0.6970.697 0.8360.836 0.8310.831 all ctrl schemes Yes No Yes Yes Auxiliary latent 195195 539539 none 706706 test AUROC 0.8400.840 0.7480.748 — 0.7530.753 min. ctrl AUROC 0.8540.854 0.7150.715 — 0.7310.731 all ctrl schemes Yes No — No Brightness support carrier 973973 119119 661661 781781 charge monotonicity 0.9150.915 0.9150.915 0.8910.891 0.9150.915 clean–aux AUROC 0.5620.562 0.5690.569 0.5450.545 0.5680.568 Table 11: Functional read-out analogues across SAE seeds. Each block reports the coordinate and its held-out validation quantities. The larger response landscape is more stable than the individual coordinates. Every seed contains more than 100100 clean-monotonic and more than 100100 auxiliary-monotonic latents at |ams|≥0.85|ams|≥ 0.85. Across the tested expansion and sparsity grid, every dictionary also contains a control-surviving bright-clean analogue, with validation AUROC between 0.840.84 and 0.990.99. The robust object is therefore the physical role and the broader response axis, not a particular latent index. Appendix C Head-relative intervention details C.1 Family-level removals Table 12 reports family removals on causal_test. None of the clean- or auxiliary-side sets produces a selective direction-head effect. The control means are much larger than the medians for several families because a minority of high-density control sets removes substantial reconstruction mass. Family n Status �+ _+ �− _- Select. z+z_+ Ctrl med. Ctrl mean clean_pair 22 null +0.033∘+0.033 −0.007∘-0.007 +0.039∘+0.039 −0.33-0.33 +0.024∘+0.024 +0.844∘+0.844 clean_family 194194 null −0.020∘-0.020 −0.016∘-0.016 −0.005∘-0.005 −1.25-1.25 +7.651∘+7.651 +9.239∘+9.239 zauxz_aux 11 null +0.004∘+0.004 +0.033∘+0.033 −0.029∘-0.029 −0.22-0.22 −0.006∘-0.006 +0.247∘+0.247 aux_family 135135 null −0.008∘-0.008 +0.031∘+0.031 −0.039∘-0.039 −0.75-0.75 +0.025∘+0.025 +1.305∘+1.305 clean_core 55 null +0.018∘+0.018 −0.010∘-0.010 +0.028∘+0.028 −0.63-0.63 +0.033∘+0.033 +2.966∘+2.966 Table 12: Reference-dictionary removals on causal_test. “Select.” is �+−�− _+- _-; the last columns summarize matched controls. The clean-pair removal is the more reliable verdict for the secondary coordinate 10061006: removing the pair changes the positive-class error by only +0.033∘+0.033 and fails both the effect and control-normalized criteria. Removing the complete 194194-latent clean family is also null. The absence of an effect therefore does not result from testing only one sparse coordinate. C.2 Cross-seed single-latent replication We identify the latent that plays the same physical role in each dictionary and repeat the direction-head removal test. Table 13 shows that removing the primary clean or auxiliary read-out never passes the causal criteria in any evaluable seed. The direction-head null is therefore not specific to the reference dictionary. Role Seed 1 Seed 2 Seed 3 Seed 4 Clean primary +0.057∘+0.057 (null) +0.098∘+0.098 (null) +0.031∘+0.031 (null) −0.033∘-0.033 (null) Clean secondary +0.082∘+0.082 (pass) +0.113∘+0.113 (null) +0.042∘+0.042 (null) −0.019∘-0.019 (null) Auxiliary +0.156∘+0.156 (null) −0.117∘-0.117 (null) — −0.040∘-0.040 (null) Brightness support +13.51∘+13.51 (V) +14.81∘+14.81 (V) +11.61∘+11.61 (V) +14.59∘+14.59 (V) Table 13: Direction-head shifts across SAE seeds. Parentheses give the fixed-rule verdict or dose-response shape. The secondary clean read-out has one exception. In seed 1, removing latent 10061006 produces a small shift of +0.082∘+0.082 that formally passes the causal test because the matched-control effects have very little spread. However, the corresponding secondary read-outs are null in the other three seeds, and removing the clean pair 640,1006\640,1006\ is also null. We therefore report the seed-1 result as an isolated pass rather than a replicated causal effect. The brightness-support coordinates behave differently. Removing them changes the angular error by 11.6∘11.6 -14.8∘14.8 in every seed. Their full dose responses are V-shaped: moving the activation either below or above its natural value damages the prediction. These coordinates therefore support direction reconstruction, but the head does not interpret their activation as a monotonic brightness signal. C.3 Writing and swap interventions Removal tests whether the direction head uses a read-out. Writing tests whether changing that read-out can change the head’s prediction. We write clean-event activations into auxiliary events, write auxiliary-event activations into clean events, and swap brightness-related activations between bright and dim clean events. Table 14 reports the results. None of these interventions produces a larger effect than the matched controls. We therefore find no evidence that the direction head uses or responds to these validated read-outs. Experiment � recip. � donor z(�+)z( _+) z(sel.)z(sel.) Verdict write 640640 into aux −0.012∘-0.012 +0.025∘+0.025 −0.92-0.92 −0.71-0.71 null full swap (++zero 195195) −0.002∘-0.002 +0.034∘+0.034 −0.01-0.01 −0.51-0.51 null write 640,1006\640,1006\ into aux +0.003∘+0.003 +0.045∘+0.045 −0.21-0.21 +0.03+0.03 null full swap +0.013∘+0.013 +0.053∘+0.053 +0.38+0.38 −0.22-0.22 null write clean core (55 lat.) +0.006∘+0.006 +0.044∘+0.044 −0.06-0.06 +0.25+0.25 null full swap +0.016∘+0.016 +0.051∘+0.051 +0.28+0.28 +0.56+0.56 null write 195195 into clean +0.024∘+0.024 −0.003∘-0.003 −0.53-0.53 −1.04-1.04 null full swap (++zero 640,1006640,1006) +0.034∘+0.034 −0.003∘-0.003 +0.21+0.21 +0.69+0.69 null brightness swap: dim→ (640640) +0.016∘+0.016 +0.114∘+0.114 −0.44-0.44 −0.40-0.40 null brightness swap: bright→ (640640) +0.115∘+0.115 +0.013∘+0.013 +0.10+0.10 +0.70+0.70 null Table 14: Direction-head writing tests. Recipient and donor columns give the induced angular-error shifts. C.4 Functionally trained dictionaries For each dictionary, we first identify the latents whose removal changes the direction-head output beyond matched controls. We call these latents the causal set. We then test every latent in this set against the full concept vocabulary, including event quality, auxiliary activity, energy, direction, depth, morphology, and angular error. Separately, we search the complete dictionary for a latent representing bright clean events. Table 15 summarizes the four dictionaries. Dictionary Objective Bright-clean AUROC Causal set Best concept AUROC Validated pairs One-shot pure functional 0.530.53 1616 — 00 R1 pure functional 0.530.53 1111 0.590.59 00 R9 reconstruction anchored 0.870.87 66 0.690.69 00 R10 reconstruction anchored 0.8780.878 66 <0.75<0.75 00 Table 15: Results for the functionally trained dictionaries. ‘Bright-clean AUROC” gives the best bright-clean read-out, ‘Causal set” counts latents that affect the direction head, and ‘Best concept AUROC” tests those latents against the full concept vocabulary. ‘Validated pairs” counts the pairs that pass the complete validation procedure. The two pure functional dictionaries contain 1616 and 1111 causal latents, confirming that the training objective finds coordinates that affect the head. However, neither dictionary contains an identifiable bright-clean latent: the best AUROC is 0.530.53, close to chance. Adding the reconstruction term recovers a bright-clean latent in both anchored dictionaries, with AUROC 0.870.87 and 0.8780.878. In both cases, this latent lies outside the causal set and its removal does not affect the direction head. Conversely, none of the causal latents passes validation for any tested physical concept. The result therefore persists even when causal latents and an identifiable bright-clean latent coexist in the same dictionary. C.5 Uncertainty-head replication The reference-dictionary results are summarized in Table 4. Here we test whether the positive clean-side verdicts depend on the SAE draw. Removal of zbcz_bc, the clean pair, the clean family, and the clean core shifts the uncertainty head toward larger predicted error in every seed. The auxiliary-side interventions remain null, preserving the controlled contrast reported in the main text. Intervention Seed 1 Seed 2 Seed 3 Seed 4 remove zbcz_bc +0.310+0.310 +0.193+0.193 +0.295+0.295 +0.612+0.612 remove clean pair +0.321+0.321 +0.374+0.374 +0.307+0.307 +0.612+0.612 remove clean family +0.775+0.775 +0.837+0.837 +1.012+1.012 +0.805+0.805 remove clean core +0.373+0.373 +0.493+0.493 +0.681+0.681 +0.617+0.617 clean core after support exclusion +0.373+0.373 +0.493+0.493 +0.302+0.302 +0.617+0.617 Table 16: Cross-seed uncertainty-head removals. Entries are shifts M1M_1 in predicted log error. The first four rows reproduce the positive-sign reference-seed verdicts in every draw. The final row diagnoses the larger seed-3 core effect: after removing its support-carrier contamination, M1M_1 falls from +0.681+0.681 to +0.302+0.302 but remains significant (z=6.8z=6.8). Seed 3 also splits the support role between latents 661661 and 144144. The direction head is more sensitive to 661661, whereas the uncertainty head is more sensitive to 144144. Consequently, the pre-specified single-carrier uncertainty dose replicates in 3/43/4 seeds; a post-hoc joint dose on 661,144\661,144\ restores the monotone response in the remaining draw.