Paper deep dive
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
Dekun Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/17/2026, 3:54:28 AM
Summary
This paper presents a calibrated test of internal action maps in language models, specifically investigating whether hidden state signals support reusable, causally effective affine transformations (action maps) that satisfy algebraic closure properties like composition and commutativity. Using Qwen/Qwen3-4B and synthetic finite worlds (Z_2^11 and S_5), the study finds that while state signals are decodable and causally usable at certain layers (e.g., h28/h36), they fail to form globally closed affine maps. One-step reconstruction errors are significant, and composition gates are not passed, suggesting that state availability, causal use, and reusable closure are separable properties.
Entities (10)
Relation Signals (7)
Dekun Yang → affiliatedwith → Zhejiang University
confidence 99% · Dekun Yang 1,* 1 Zhejiang University
Qwen/Qwen3-4B → usedin → Alchemy
confidence 99% · The grounded experiments freeze the post-trained thinking/non-thinkingQwen/Qwen3-4Bcheckpoint for an Alchemy state-update task.
Qwen/Qwen3-4B → haslayer → h28
confidence 98% · In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519
S_5 → usedas → positive_control
confidence 97% · validate the geometric branch on a known affine S_5 carrier
h28 → fails → composition
confidence 96% · yet no refit passes composition.
h28 → exhibits → causal_use
confidence 95% · Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36.
Alchemy → tests → state_update
confidence 95% · The grounded experiments freeze the post-trained thinking/non-thinkingQwen/Qwen3-4Bcheckpoint for an Alchemy state-update task.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus .398 for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to .474 (.469 with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes.
Tags
Links
- Source: https://arxiv.org/abs/2608.13626v1
- Canonical: https://arxiv.org/abs/2608.13626v1
Trouble viewing inline? Open PDF directly →
Full Text
55,480 characters extracted from source content.
Expand or collapse full text
A Calibrated Test of Internal Action MapsarXiv preprint A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure Dekun Yang 1,* 1 Zhejiang University *Correspondence: pauliyangwork@gmail.com ORCID: Dekun Yang: 0009-0002-3496-3596 Preprint; evidence cutoff 12 August 2026 Abstract. A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine푆 5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trainedQwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus.398for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to.474(.469with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes. Keywords: mechanistic interpretability; causal intervention; compositional generalization; operator closure; representation geometry; permutation groups 1. Introduction Models that track a changing world must bind entities to their current attributes, update those attributes after actions, and carry the consequences forward. Language-model activations can encode entity states, and interventions on those activations can change later predictions [Li et al., 2021, Kim and Schuster, 2023]. Other work has identified computations that route or retrieve state information during fine-tuning and belief tracking [Prakash et al., 2024, 2026]. Together, these findings establish the presence and relevance of a state signal. They stop short of showing that each action implements a stable function that transfers to held-out sources and composes with other actions. Several weaker observations can look like evidence for such an operator. A probe may decode an off-manifold activation, or an affine map may beat a mean direction while still missing much of the natural target displacement. Correct and reversed action orders can also be ranked when neither endpoint is accurate. Static relation maps, task vectors, and function vectors make the operator hypothesis plausible [Hernandez et al., 2024, Hendel et al., 2023, Todd et al., 2024]. None of these observations alone establishes closure over changing world states. Our measurement framework is therefore a lattice rather than a sequential ladder. Once the availability of state information is established, causal local use and reusable closure/algebraic laws can be tested independently. Geometry without a detected output effect remains informative, as does causal use without held-source closure. Neither is the full claim. Under a behav- iorally eligible task, only evidence from both branches licenses the phrase “causally faithful reusable operator.” This partial order avoids implying that causal patching must precede every geometric diagnostic. We study one controlled language-model setting alongside two exact finite worlds. The grounded experiments freeze the post-trained thinking/non-thinkingQwen/Qwen3-4Bcheckpoint for an Alchemy state-update task. The other experiments train decoder-only Transformers in a 121-stateZ 2 11 system and the 120- state permutation group푆 5 . Alchemy retains language-model context dependence and permits paired interventions. The finite worlds provide complete transition tables, commuting labels, inverses, and training trajectories. A failed gate immediately raises a construct-validity question: would the test pass when the operator is known to exist? We answer it with an end-to-end positive control, a known contextual 푆 5 carrier evaluated under the same held-source logic. A parity-support split, smooth observation curvature, and held- domain coordinate conjugacy then stress that carrier. Exact recovery checks the implementation; the structured stresses probe behavior under misspecification. Their calibration is informative but limited, so we do not treat success on one favorable noise family as evidence of universal power. The grounded results separate reconstruction, causal use, and carrier identity. Within-test-domain cross-fitting improves h28 reconstruction, but randomized entity splits do not support maps dominated by entity identity. Action identity instead dominates the tested parameter geometry. One-step reconstruction is best at h4/h16, whereas independently regenerated paired interventions are effective at h28/h36. The lexical and state-probe controls do not settle whether h4 is a surface carrier. Finally, a frozen bridge associates most of the apparent h28 one-step improvement with final refitting rather than inverse-displacement weighting or a larger affine function class. Those refit maps still fail composition. Methodologically, the study provides a calibrated evidence lattice built around natural endpoints, held sources, distinct labels for negative and blocked outcomes, and independent audits of every formal follow-up. Empirically, the tested representations separate state availability, causal use, action-dominant parameter geometry, one-step specification sensitivity, relative algebraic 1 arXiv:2608.13626v1 [cs.AI] 13 Aug 2026 A Calibrated Test of Internal Action MapsarXiv preprint discrimination, and global affine closure. These findings apply to the tested carriers and hypotheses. They are not a general denial of internal operators. 2. Related Work 2.1 State tracking and causal state use Entity-tracking studies show that hidden activations can contain dynamic world information. Li et al. [2021] combined decod- ing with intervention, while Kim and Schuster [2023] studied behavioral tracking across entities and sequence conditions. In procedural text, Neural Process Networks update a designed entity-memory state with learned action operators [Bosselut et al., 2018]. Later work examined how fine-tuning reuses entity- tracking mechanisms and how lookback operations bind beliefs to earlier observations [Prakash et al., 2024, 2026]. Together, these studies address availability, binding, causal relevance, and architecturally specified state transformation. Our closure branch asks something else: whether an action-conditioned map recovered from an existing model’s residual stream predicts the natural post-action representation of sources excluded from fitting. 2.2 Linear transformations and composed functions Affine relation decoders can map subjects toward objects and support causal edits [Hernandez et al., 2024]. Task and function vectors can summarize computations induced by demonstrations and influence output behavior [Hendel et al., 2023, Todd et al., 2024]. Work on continuous latents and compositional primitives extends this idea to multi-step computation [Hao et al., 2025, Lippl et al., 2026]. Khandelwal and Pavlick [2026] study two-hop factual func- tions푔( 푓(푥)), identify residual-stream signatures of the in- termediate variable, and compare direct with compositional mechanisms. Our objects, representations, closure test, and causal claims differ. We condition maps on explicit world actions, score natural hidden endpoints, hold out entities or source states from fitting, and interpret relative laws only after absolute reconstruction. An intermediate representation can support function evaluation without forming a global affine action algebra. 2.3 Algebraic structure and world models Under suitable tasks and interfaces, structural studies show that algebraic computation can be learned or imposed. Li et al. [2025] identify associative and parity-associative mechanisms in Transformers trained on permutation words. An and Du [2026] connect representation-level homomorphism error with compositional generalization. Lee [2026] evaluates a structured recurrent interface using held-out transition pairs, drift, homo- morphism, and commutator diagnostics. These constructions motivate our positive carrier and show why failure in an ordinary residual stream is not an impossibility result. Behavior supplies an orthogonal standard. Transformers trained on Markov decision processes can encode transition dynamics [Chen et al., 2024], yet strong local predictions may coexist with an incoherent recovered world model [Vafa et al., 2024]. Probe controls raise a related concern: a measurement can reward its own fitting capacity instead of the intended construct [Hewitt and Liang, 2019]. A dated search log records databases, query families, screening bounds, and a closest-work feature matrix. Within that documented search through 12 August 2026, we found no prior study combining known-algebra calibration, held-source natural endpoints, matched interventions, law tests, behavioral eligibility, and independent artifact audits across an existing language model and exact finite systems. This is a search-bounded description of the integrated protocol, not a topic-level priority claim. 3. Evidence Framework 3.1 From a state signal to an action map Letℎ(푠,푐)be the representation of world state푠in context or history 푐. For action 푎, the primary family is 퐴 푎 (ℎ)= 푊 푎 ℎ+ 푏 푎 . Interpolating the observed pairs is not enough. A reusable 푊 푎 must transfer across entities, contexts, histories, or source states withheld from fitting. For Alchemy, row-relative endpoint error is∥퐴 푎 (ℎ)− ℎ ′ ∥ 2 /∥ℎ ′ − ℎ∥ 2 ; identity therefore has error one whenever the displacement is nonzero. In the finite worlds, normalized root-mean-square error (NRMSE) divides pooled squared error by a frozen target-variance reference. Both metrics compare the prediction with the natural post-action activation, rather than only checking a decoded label. Consider a held entity initially in stategreen, followed by fill_redandfill_blue. A state probe checks whether the final activation decodes as blue. An order test compares 퐴 blue 퐴 red ℎwith the reversed order. Absolute closure instead asks whether the composed point matches the held source’s natural final activation. The probe and order test can both succeed while that endpoint remains inaccurate. State availability is measured separately from causal local use. On held data, a fixed probe is evaluated against a deterministic random-label control. The preregistered residual replacement must change the target-versus-source logit margin more than a matched wrong-state or translation control. Because patching may move the activation off manifold, a successful intervention licenses a causally usable direction, not a naturally traversed transition. 3.2 Absolute closure and relative laws H1 compares one-step reconstruction with identity, mean trans- lation, and random-map controls. The grounded diagnostics also include a residual MLP, radial-basis-function kernel ridge, within-test-domain cross-fit, source-state-label-gated affine bank, and an optimistic in-sample capacity ceiling. Each reference answers a different question, and several are not deployable mechanisms. H2 scores퐴 푏 (퐴 푎 (ℎ))against the natural two-step endpoint and compares it with a directly fitted two-action map. Because the second setter overwrites the first, the setter domain also re- quires a last-action-only baseline. The frozen grounded decision combines endpoint error, direct-map discrepancy, order effect, and state-probe gain. An order contrast cannot pass H2 on its own. 2 A Calibrated Test of Internal Action MapsarXiv preprint H3 concerns commutativity. In Alchemy, its inferential units are matched same-entity and disjoint-entity pairs. In the finite worlds, the transition table labels every unordered action pair; AUROC then measures whether normalized commutators rank noncommuting pairs above commuting pairs. H4 applies the ground-truth inverse and scores the return to the source representation. At the checkpoint level, H5 tested whether algebraic violation covaried with behavioral incoherence. 3.3 Gates, branches, and evidence labels All thresholds, partitions, and inferential units were frozen before their corresponding formal results.NOT SUPPORTEDmeans that a testable joint gate failed.NOT TESTEDrecords an upstream- blocked experiment.UNTESTABLErecords a missing logical antecedent or usable observation.RIGHT-CENSOREDrecords an event not observed within a fixed budget. Diagnostic branch results cannot rewrite older frozen verdicts. 4. Experimental Settings 4.1 Controlled Alchemy and split support The grounded experiments used a fixed local snapshot of Qwen/Qwen3-4B, the post-trained thinking/non-thinking check- point rather thanQwen3-4B-Base. It has 36 Transformer layers and hidden width 2,560; weights were never updated. Prompts described containers in one of five states:empty,red,blue, green, oryellow. Actions emptied a container or filled it with one color. Representations were taken at the final prompt token. The original Phase 1 scan froze h28 because it was the earliest sampled location that passed both direct state decoding and a paired state-specific intervention. The split withholds entity identities and generated context/activation instances, not state categories or the two primary templates. For each of three dataset-seed regenerations from the same frozen checkpoint, H1 fits five action maps from 500 train, 100 validation, and 200 test pairs per action. H2 covers all 20 ordered distinct-action sequences and collapses reverse directions into ten semantic pairs for inference. H3 uses matched same-entity and disjoint-entity pairs. The planned natural-language inverse test was stopped when the behavior manipulation failed. 4.2 Known-algebra calibration and sampled-layer follow-up The positive control represents each of the 120 states of푆 5 by its5× 5permutation matrix plus a seven-dimensional repeat- specific context carrier that actions leave unchanged. A seeded orthogonal embedding maps the carrier into 64 observed di- mensions. Ten transpositions act linearly on the permutation component and identically on context. Three outer folds use 80 source states and repeats 0–7 for chart/map fitting; 40 states and repeats 8–11 remain excluded. The푆 5 permutation matrices span 1 2 + 4 2 = 17 dimensions (trivial plus standard representa- tion); adding seven context dimensions gives 24, and removing the constant direction absorbed by the affine bias yields the training-determined chart rank 23. A second embedding seed provides an independent reproduction. Phase 6 added target-side Gaussian nuisance at signal-RMS ratios0,.05,.10,.20,.40,.80, 20 nuisance seeds per ratio, with- out moving thresholds. It also froze h28 references and scanned h4, h16, h28, and h36 with full and rank-1/4/16/64 residual affine maps. Phase 7 then evaluated h4/h16 on all ordered action pairs with the unchanged composition construct; single-map choices were imported without sequence-data selection. 4.3 Structured diagnostics and attribution bridge Phase 8 was prospectively frozen against the Phase 6/7 artifacts. Structured calibration rebuilt the carrier under five embeddings. It tested an even-to-odd permutation support split, a common smooth quadratic/tanh observation warp, and a held-domain orthogonal conjugacy at strengths0,.05,.10,.20,.40,.80. The latter two preserve an exact latent action while progressively misspecifying a single observed-space affine map. Entity/context conditioning used h28 only. Stable-hash sam- pling retained 80 rows per entity-action cell. Seven cyclic outer splits fit four entities, selected on a fifth, and held out two; a five-fold within-entity reference used the same row budget. Residual affine parameter distance combined Frobenius matrix and intercept distance. Because entity, token, and context fea- tures covary, even a positive result would not identify a pure entity mechanism. Lexical controls fitted true and permuted-label state probes at h4/h16/h28/h36, including a conflict stratum in which the last mentioned state word disagreed with the queried current state. A token-only carrier averaged the last 16 static input embeddings. Neutral clauses inserted requested distances of 0, 16, or 64 tokens before the query, after which the same layer-local map protocol was applied. Metric-aligned fitting weighted training row푖by1/max(∥ℎ ′ 푖 − ℎ 푖 ∥ 2 2 , 10 −8 ), normalized within action and split, and reported train-defined displacement quintiles. Layer controls recorded residual norm, displacement norm, covariance participation ratio(tr퐶) 2 /tr(퐶 2 ), and one-step results after train-fitted scalar RMS normalization. Causal dataset-seed replication regenerated exactly 160 matched intervention pairs for each of three seeds and four sampled layers from the same frozen checkpoint. After the audited Phase 8 result, an explicitly outcome-aware Phase 8b bridge separated function class, final fitting rows, and weighting. It compared the original full-or-low-rank train fit, full unweighted train fit, full weighted train fit, and the corresponding full train-plus-validation refits. No test row entered selection or fitting. The unweighted and weighted refit maps were then passed unchanged to the frozen h28 H2 datasets, direct maps, probes, resampling procedures, and gates. Phase 8b is descriptive attribution, not a new confirmation or a revision of frozen H1/H2. 4.4 Exact learned transition systems The first learned world has 121 states(푥, 푦) ∈Z 2 11 and ten bijections: four translations, coordinate swap, joint negation, and four shears. Of 45 unordered action pairs, 21 commute. Models receive a start state and zero to six actions and predict only the endpoint. Three pre-norm decoder-only Transformer scales, 3.2M, 25.3M, and 85.2M parameters, each use three seeds and2 24 = 16.78million training examples. Evaluation adds lengths seven to twelve, inverse loops, route-equivalent histories, and all ordered distinct pairs. The second world uses the 120 states of푆 5 and ten self-inverse transpositions; 15 of 45 unordered pairs commute. Medium and 3 A Calibrated Test of Internal Action MapsarXiv preprint Table 1: Evidence lattice. Availability enables two independent branches. Only their conjunction under behavioral eligibility licenses the strongest operator claim. NodeMeasurementLicensed statementDoes not establish State availabilityHeld-data decoding versus label controls State information is available to the decoderCausal use or a natural update path Causal-use branchMatched intervention versus wrong-state/translation controls A state direction is locally usable by the output computation Held-source closure or an on-manifold tran- sition Closure/law branchNatural endpoints, then composition/inverse/commutativity The tested map transfers and satisfies the reported law Local output use or an untested carrier Behavioral eligibilityTask success and state differentiation Model-level interpretation is meaningfulA unique internal implementation Branch conjunctionCausal use plus reusable closure/laws under eligibility A causally faithful reusable operator in the tested carrier Universality across models, layers, tokens, or function classes Evidence lattice for internal action-map claims Availability enables two independent tests; neither branch substitutes for the other. Partial branch patterns remain informative. Directed edges denote evidential prerequisites, not temporal emergence or mediation. State availability Held-data state decoding Random-label and conflict controls Causal local use Matched residual intervention Wrong-state / translation controls Reusable closure and laws Held-source natural endpoints Composition, inverse, commutativity Behavioral eligibility Task success and state differentiation Guards model-level interpretation Causally faithful reusable operator Licensed only by the conjunction causal use + reusable closure/laws under an eligible behavioral task Figure 1: Evidence lattice for internal action-map claims. State availability enables independent tests of causal local use and reusable closure/laws. Behavioral eligibility guards model-level interpretation. Geometry without detected causal use and causal use without closure remain distinct partial outcomes; only their conjunction licenses a causally faithful reusable operator. Directed edges denote evidential prerequisites, not temporal emergence or mediation. The diagram contains no sampled observations. large trajectories extend to2 25 = 33.55million examples and 18 fixed checkpoints. E denotes task/state eligibility, D endpoint- order discrimination, J a shared rigid chart across zero-, one-, and two-action conditions, and K held-source affine closure. Only six medium/large trajectories enter the preregistered cross-world denominator; small models remain capacity controls. 4.5 Statistics, reproducibility, and auditing The grounded descriptive replicate is the regenerated dataset seed (푛= 3) from one frozen checkpoint. Prompts, actions, entities, layers, folds, bootstrap draws, and diagnostic cells are not independent model replicates. Pairwise permutation and bootstrap tests use ten semantic action pairs. Intervention intervals resample 160 matched pairs within each dataset seed; the cross-seed direction is reported separately. Positive-control recovery requires every held-source fold. Synthetic trajectory intervals resample complete training trajectories. The five preregistered grounded action-level comparisons use Benjamini–Hochberg false-discovery-rate control at푞= .05. Follow-up diagnostic families carry their prospectively frozen labels but do not alter Phase 1–7 decisions. Phase 8b is explicitly outcome-aware. Exact package versions are recorded by formal run. Every formal run verifies frozen input hashes, records config- uration, code commit, environment and resource telemetry, and writes an atomic artifact manifest beforeSUCCESS. Independent auditors do not import the fitting, scoring, bootstrap, or verdict helpers under test. Phase 8 independently rechecked structured masks and gates, entity splits and map distances, lexical probes and carriers, weighted fits, layer statistics, all 1,920 intervention rows, and frozen parity. Phase 8b independently refitted all 75 one-step maps, reproduced all three original H2 summaries, and rebuilt every composition gate. Both remote and locally copied artifacts passed their independent audits. 4 A Calibrated Test of Internal Action MapsarXiv preprint Table 2: Grounded split support. The protocol holds out entities and generated context/activation instances, not state categories or the two primary templates. SplitRows/seed Entities State categories Templates Role Train2,500A–Dall five0, 1 Map fitting Validation500Eall five0, 1 Rank/ridge selection Test1,000F–Gall five0, 1 Held-entity/context evaluation 5. Results 5.1 The lattice recovers a known algebra but structured calibration is limited The exact contextual carrier passed every Phase 6 positive- control gate across three formal and three reproduction folds. Held-state, one-step, two-step, and inverse decoding were 1.000. One-step, two-step, and inverse-cycle NRMSEs were approxi- mately3.31×10 −8 ,5.02×10 −8 , and6.55×10 −8 ; commuting-pair AUROC was 1.000. The independent embedding retained every gate direction, with all continuous comparisons inside the frozen 5% reproduction tolerance. Target-side nuisance produced a smooth specificity curve. At nuisance-to-signal ratios.05,.10,.20,.40,.80, mean one- step NRMSE rose to.0096,.0191,.0382,.0764,.1528; two-step error rose to.0136,.0272,.0544,.1090,.2188. This establishes exact-solution recovery and a graded response to isotropic target perturbation, not power against every structured failure. Phase 8 supplies the harder distinction. The even-to-odd support split and every zero-strength cell passed. Curved and held-domain families each had Spearman휌= 1.000between strength and median one/two-step error, but their scales dif- fered sharply. At strength.80, curved median one/two-step NRMSE was.0118/.0127; held-domain conjugacy reached .814/1.050while decoding remained 1.000. Across both maximum-strength families, 23/30 seed-fold cells crossed at least one closure gate, just below the frozen 24/30 criterion. The label isSTRUCTURED_CALIBRATION_LIMITED: the test detects a domain-dependent chart, whereas the curved family remains far below the decision region. Later failures cannot be attributed to an implementation that never passes, but one stress family cannot certify universal discriminative power. 5.2 The h28 transfer gap is action-dominant, not purely entity-specific Frozen h28 state-probe accuracy was82.87%± 1.81%, versus 20.80%± .20%for the random-label control. Across three datasets, the specific target-minus-wrong intervention effect was 2.334± .231logits. Original train-only affine maps had mean error.5189± .0286across 15 action-by-seed cells (cell-level sample SD), with 4/15 below the frozen.50gate. Translation averaged.9166. Within-test-domain five-fold cross-fit passed every cell at.3979± .0130, whereas radial-basis-function ridge averaged.5200and the source-state-label-gated affine bank .5689. The cross-fit is a descriptive within-domain reference, not a held-entity or deployable map. Direct tests did not support the stronger entity-specific in- terpretation. Across seven randomized outer splits, within- entity fitting averaged.424and cross-entity transfer.473. Per- seed within/cross ratios were.895,.900,.894, above the frozen .85cutoff, and only 6/21 seed-split directions were favor- able. Median same-action/across-entity parameter distances were.0281,.0271,.0269, smaller in every seed than different- action/within-entity distances.0292,.0287,.0297. The frozen label isACTION_DOMINANT_GEOMETRY, not entity conditioning. Action identity dominates this parameter-distance comparison, although entity, token, and contextual variation can still con- tribute to the transfer gap. Threshold sensitivity shows where the original result lies. No h28 cell passes at.40, 4/15 pass at.50, and all 15 first pass at.58. The frozen H1 verdict remainsNOT SUPPORTED. The pattern marks incomplete affine structure near a declared boundary, rather than absence of all action geometry. 5.3 Early geometry, lexical evidence, and causal depth re- main distinct Mean one-step error was.299at h4 and.347at h16; all 15 action-by-seed cells at each layer met the geometric rule. The h28 and h36 means were.519and.549. Scale alone does not explain this ordering. From h4 onward, mean residual norms were11.99,40.23,159.25, and147.73; displacement norms were.686,4.15,36.10, and25.82. After train-fitted scalar RMS normalization, mean errors remained approximately .295,.349,.519,.549. Effective rank, however, was only1.057 at h4 and1.085at h16, versus1.819and2.261at h28/h36; low-dimensional early geometry remains a live confound. The h4 carrier remains unresolved by the lexical con- trols. In the conflict stratum, true-state probe accuracy at h4 was.268,.262,.237, against permuted-label values of .192,.219,.184. This was weakly above control but far be- low the frozen.80criterion. Neutral material raised mean h4 action-map error from.299at distance 0 to.392at 16 and .434at 64; no seed passed all five actions at distance 64. The last-16-token static-embedding carrier also missed the all-action rule in every seed. The state-probe and distance gates fail, and the token-only carrier does not pass. The evidence there- fore supports neither a lexical explanation nor its exclusion: LEXICAL_ROLE_UNRESOLVED. The causal-depth pattern replicated across three dataset-seed regenerations from the same frozen checkpoint. Mean spe- cific effects were.0003± .0061logits at h4 and.0021± .0032 at h16; every within-seed pair-bootstrap interval contained zero. At h28, effects were2.424, 2.071, 2.506logits, and at h36 4.276, 4.399, 4.567. All six intervals excluded zero, and every sign-flip test gave푝= 10 −5 . Agreement across all 12 frozen di- rections yieldsCAUSAL_DEPTH_PATTERN_REPLICATED. These data establish a replicated sampled-depth dissociation within one checkpoint, not a temporal stage, mediation path, unique carrier, or model-level replication. 5 A Calibrated Test of Internal Action MapsarXiv preprint Phase 7 separately tested whether the early affine geometry composes. At h4, two-step endpoint error averaged.542and direct-map discrepancy.541; the h16 values were.726and .717. No layer-seed cell passed the joint H2 gate (0/6), and held-template endpoint errors exceeded 1.33. Under the lattice, these closure-branch diagnostics remain valid despite the absent causal effect; their label isEARLY_LAYER_GEOMETRY_ONLY, not a causally faithful early operator. 5.4One-step H1 depends on refitting scope, but composition remains unsupported Phase 8 initially appeared to show a metric-aligned rescue: weighted full maps refitted on train plus validation averaged .4686, with all 15 cells below.50. There were no exact no- op rows. Displacement stratification still showed denominator sensitivity: unweighted refit error fell from.558in the smallest- displacement quintile to.387in the largest; weighted values were .533 and .394. The outcome-aware descriptive Phase 8b bridge localizes the difference differently. Original selected train maps and full unweighted train maps were numerically identical at.5189 (4/15 below.50), excluding function-class selection as the explanation. Full weighted train maps reached only.5113 (5/15). Unweighted train-plus-validation refitting, by contrast, reached.4744(14/15), and weighted refitting reached.4686 (15/15). Refitting improved all 15 cells; median gains were .0395unweighted and.0399weighted. Weighting contributed median gains of only.0067on train and.0051after refitting, both below the frozen.01attribution criterion. The label is WEIGHTING_ATTRIBUTION_NOT_SUPPORTED. Within this de- scriptive bridge, the h28 one-step conclusion is associated mainly with final fitting scope, with a smaller increment from weighting. The older train-only H1 verdict remains frozen. This sensitivity does not extend to composition. Orig- inal h28 composition had mean endpoint error.8110 and direct-map discrepancy.7945. Unweighted refitting changed the pair to.7982/.8199, and weighted refitting to .7868/.8059. Endpoint error improved modestly, but the decisive direct-map gap stayed far above.50; every vari- ant failed H2 in all three datasets. The bridge verdict is METRIC_ALIGNED_COMPOSITION_NOT_SUPPORTED. 5.5Relative order signals remain weaker than endpoint closure At h28, reversed-order minus correct-order error was+.0895 across ten semantic pairs, with pair-bootstrap 95% CI [.0353,.1444], one-sided푝= .00114, and푑 푧 = .960. Every pair was positive in every dataset seed. Yet correct composition error was.811, a directly fitted two-action map reached.443, and composed-to-direct discrepancy was.794. Applying only the final setter was better than composing both maps (.685), as expected under last-write absorption. Probe gain was 16.1 per- centage points, below 20. H2 remainsNOT SUPPORTEDdespite the reproducible order contrast. The effect was concentrated in particular action families. Empty-versus-fill pairs had an order advantage of+.1978and a probe advantage of 37.1 points; fill-versus-fill pairs had+.0173 and 2.0 points. H3 showed the same boundary: the same-entity- minus-disjoint commutator difference was+.0768with 95% CI [.0336,.1240], but푑 푧 = .431, and only44.6%of matched pairs were positive. These decompositions motivate a destructive- clearing versus replacement hypothesis; they do not rescue all-pair H2 or H3. Grounded H4 remainsNOT TESTED. Even the best behavior scaffold confirmed the intended zero/one/two-action manipula- tion only 49.33% overall, with 28.5% in its weakest condition, below the 80% prerequisite. The result constrains that prompt manipulation, not every possible inverse representation. 5.6Shared charts and relative laws do not guarantee learned closure The learned푆 5 branch carries more inferential weight than Z 2 11 because all six medium/large trajectories passed task/state eligibility E, endpoint-order D, and shared-chart J. Stable E/D was observed between 6.29M and 12.58M examples, and J at 12.58M for medium and 25.17M for large. These left-truncated observation times do not define a causal emergence sequence. K was not observed in any trajectory within the fixed2 25 = 33.55 million-example budget, so the event isRIGHT-CENSORED. Final one-step NRMSE was1.084 ± .012for medium and 1.002± .029for large; two-step values were1.070± .010and 1.022± .027. Nested frozen-weight diagnostics found seen- source one-step NRMSE.260± .006versus held-source.786± .036 , held composition.835± .022, teacher-forced composition .540± .015, and a separately fitted depth-one map.497± .030. Source-state extrapolation and closed-loop accumulation both matter, but neither rank 119 nor residual multilayer perceptrons consistently rescue the result. We retain theZ 2 11 branch as a task-ineligible diagnostic. Medium/large ID accuracies exceed 99%, but length-OOD accuracy is only 6.5%–7.3%, and loop/route scores remain near chance. All nine final models nevertheless achieve commuting- pair AUROC 1.000 while the absolute H1, H2, and H4 metrics fail. Relative pair discrimination can coexist with poor endpoints, but this branch cannot adjudicate closure in an algorithmically qualified model. The preregistered H5 composite also failed as a measurement. Across 99 nonzero checkpoints, partial Spearman휌= .0911 with trajectory-cluster 95% CI[−.0352,.2486]did not support the predicted association. One trajectory illustrates why: at initialization, it decoded states at 3.34% while H1/H2 errors were spuriously low at.091/.120; after state differentiation, decoding reached 89.50% and the errors rose to.813/.920. Representation scale and an unmet task gate confounded the mixed composite. No favorable subset or alternative composite replaces it. 6. Discussion and Limitations Under the evidence lattice, the layer results are different partial outcomes rather than contradictions. h4/h16 are geometry- positive and causal-use-negative; h28 is causal-use-positive and frozen-closure-negative. Neither region licenses the conjunction. Selecting one layer because it decodes well or patches strongly would instead turn a local result into a model-wide mechanism. The positive carrier and structured stresses answer different calibration questions. Near-zero recovery validates the fit, chart, 6 A Calibrated Test of Internal Action MapsarXiv preprint One-step Two-step Inverse cycle 10 −9 10 −7 10 −5 10 −3 10 −1 NRMSE (log scale) a Known algebra is recovered 0.00.20.40.60.8 Target nuisance / signal RMS 0.00 0.05 0.10 0.15 0.20 0.25 0.30 NRMSE b Error grows smoothly with nuisance One-step Two-step Inverse cycle Cross- fit Frozen h28 RBFState oracle In- sample 0.0 0.2 0.4 0.6 Held-pair relative error c Domain shift, not a noise floor 0.250.300.350.400.450.500.55 One-step relative error 0 1 2 3 4 5 Specific intervention effect (logits) h4 h16 h28 h36 d Geometry and causal use separate reproducibly Figure 2: Calibration and sampled depth. a, Exact contextual푆 5 carrier results in three held-source folds; lines mark the.50and.60gates. b, Mean and empirical 95% nuisance-seed intervals over 20 seeds after averaging folds. c, Fifteen h28 action-by-dataset-seed cells for five references; cross-fit is fitted inside the test distribution, and the line is the unchanged H1 gate. d, One-step error averaged over five actions within each dataset seed versus matched intervention effects and within-seed 95% pair-bootstrap intervals. Each layer has three dataset-seed regenerations of 160 pairs from the same frozen checkpoint; colors denote layer. holdout, and gates. Target nuisance provides a graded specificity check, whereas held-domain conjugacy moves the gates into the grounded error range through structured chart mismatch. Because the curved family remains far below those gates, the stress suite does not establish universal statistical power. That limitation is part of the result. The h28 transfer gap does not warrant a pure entity-binding account. Within-domain cross-fit shows attainable local fit but has access to the evaluation domain. Randomized splits yield only modest within-entity gains, and the frozen map comparison is action-dominant. The supported description is a distribution- dependent map with action-structured parameters; entity, token, template, and contextual effects remain entangled. Specification matters for the one-step result. The observed h28 H1 estimate changes far more between train-only and train-plus- validation fits than between weighting choices, though neither procedure uses test data. The refit is a legitimate generalization estimate under a different protocol, not a post hoc revision of the original decision. Its failure to carry over to H2 sets the boundary: better one-step interpolation does not imply reusable composition. We distinguish confirmatory from descriptive analyses. The prospectively frozen set comprises the original H1–H5 gates, the Phase 6 positive carrier and target-nuisance tests, Phase 7 early-layer H2, and the Phase 8 structured, entity, lexical, metric, geometry, and causal-replication families. Phase 8b is outcome- aware and descriptive; semantic-family decompositions, frozen- weight mechanism grids, and sampled-layer geometry controls remain hypothesis-generating. Depending on the branch, the independent unit is the dataset seed, complete training trajectory, or held-source fold specified in Methods. Prompts, actions, checkpoints, and grid cells are not independent model replicates. Table 4 makes the limited carrier coverage explicit. The grounded evidence is restricted to one post-trained model family, an absorbing setter domain, and four sampled final-token layers. The conflict probe is weak, and entity and context factors are correlated. Causal replication uses three generated datasets from the same checkpoint, not independently pretrained checkpoints. In the learned world, K is right-censored. None of the results excludes nonlinear, attention-mediated, cross-token, cross-layer, KV-cache, or context-conditioned operators. Even with those limits, the measurement program suggests a reporting standard. Establish state availability, then test causal 7 A Calibrated Test of Internal Action MapsarXiv preprint 0.00.20.40.60.8 Structured perturbation strength 10 −8 10 −6 10 −4 10 −2 10 0 Median NRMSE (log scale) solid: domain shift dotted: curved 23/30 gates flip at strength .8 a Structured stress exposes limited calibration One-step Two-step Within / cross error Same-action / different-action map distance 0.85 0.90 0.95 1.00 Ratio (lower favors first term) equal magnitude frozen entity-transfer criterion b Pure entity-specificity is not supported PermutedTrue 0.150 0.175 0.200 0.225 0.250 0.275 0.300 Conflict-state accuracy c h4 probe 01664 Inserted neutral tokens 0.25 0.30 0.35 0.40 0.45 0.50 h4 relative error 0/3 pass all actions at 64 Distance stress Train U Train W Refit U Refit W 0.44 0.46 0.48 0.50 0.52 One-step relative error d H1 gap differs by fit scope FrozenRefit U Refit W 0.5 0.6 0.7 0.8 H2 direct-map gap Composition stays 0/3 Figure 3: Structured diagnostics localize the remaining alternatives. a, Median one- and two-step NRMSE under common curvature (dotted) and held-domain conjugacy (solid) across five embedding seeds and three folds. b, Three dataset-seed ratios for within/cross-entity error and same-action-across-entity/different-action-within-entity map distance. c, h4 conflict-state probe versus permuted control and h4 error after 0, 16, or 64 inserted neutral tokens. d, Seed-level h28 one-step means under train/refit and unweighted/weighted fits, alongside the frozen composition direct-map gap. U, unweighted; W, weighted. Red lines are frozen gates. Table 3: Principal frozen and follow-up verdicts. Diagnostic labels remain separate from earlier frozen decisions. SettingTestVerdictDecisive evidence Exact carrierZero/support recoveryPOSITIVE CONTROL RECOVERED All formal/reproduction/support folds pass near numerical preci- sion. Exact carrierStructured stressSTRUCTURED CALIBRATION LIMITED Monotone errors, but 23/30 rather than 24/30 gate flips. Qwen3-4B h28Frozen H1 train-onlyNOT SUPPORTEDMean .519; 4/15 below .50; cross-fit is descriptive. Qwen3-4B h28Entity/context diagnosticACTION-DOMINANT GEOMETRY Within/cross ratio about .90; action distance dominates. Qwen3-4B h4Lexical diagnosticLEXICAL ROLE UNRESOLVEDWeak conflict probe; distance and token-only controls do not adjudicate. Sampled layersPaired interventionsCAUSAL DEPTH PATTERN REPLICATED h4/h16 intervals include zero; h28/h36 positive in all datasets. Qwen3-4B h28H1 attributionWEIGHTING ATTRIBUTION NOT SUPPORTED Refit gain about .040; weighting gain only .005–.007. Qwen3-4B h28Refit-map H2 bridgeNOT SUPPORTEDAll variants 0/3; direct-map gaps .794–.820. Qwen3-4B h4/h16Layer-local H2EARLY LAYER GEOMETRY ONLY One-step geometry passes, but frozen H2 is 0/6. Learned 푆 5 K held-source closureRIGHT-CENSORED J passes 6/6; K absent through the fixed 33.55M-example budget. Z 2 11 trajectoriesH5 algebra-to-coherenceNOT SUPPORTED / NOT INTERPRETABLE CI crosses zero; task gate fails; early collapse invalidates low errors. use and held-source closure as separate branches. Report natural endpoints and continuous error, with calibration against a known operator under structured misspecification. Keep within-domain reference fits distinct from held-domain mechanisms. Com- 8 A Calibrated Test of Internal Action MapsarXiv preprint Table 4: Carrier coverage. Pass and not supported refer only to the listed test and frozen representation. Candidate carrierOne-stepCompositionCausal useCurrent status Final-token residual h4passnot supportednot detectedLow-dimensional geometry; lexical role unre- solved Final-token residual h16passnot supportednot detectedLayer-local affine geometry only Final-token residual h28fit-scope sensitivenot supportedreplicated positiveCausally usable signal without global affine closure Final-token residual h36not supportednot testedreplicated positiveCausal-use branch only Last-16 static token embeddingsnot supportednot testednot testedDoes not explain h4 by itself Other tokens/cross-layer spansnot testednot testednot testedLive alternative Attention/MLP pathsnot testednot testednot testedRequires path-specific natural targets KV cache/distributed statenot testednot testednot testedRequires a different carrier and metric position tests should include direct maps, last-action controls, commutativity, and inverse cycles. Behavioral eligibility, suffi- cient statistics, manifests, and independent audits complete the record. 7. Conclusion The held-source measurement recovers a known affine action al- gebra, while structured domain mismatch can move its gates into the empirical failure range. In Qwen3-4B, action-conditioned geometry, state decoding, and within-checkpoint dataset-seed causal-use results appear at different sampled depths. Direct entity tests do not support a purely entity-specific account; lexi- cal controls remain unresolved, and the observed h28 one-step result varies substantially with final fitting scope. The boundary is composition: neither early-layer nor h28 refit maps satisfy the unchanged composition construct. Learned finite worlds draw the same distinction between relative structure or shared charts and absolute held-source closure. The positive claim is deliberately bounded. Within the tested carriers, state availability, causal local use, action-structured geometry, and reusable affine closure are distinct empirical constructs. Calling an internal map a causally faithful reusable operator should require their conjunction in the same eligible carrier. Richer internal operators remain possible; this study specifies the evidence such a claim must reconstruct. Data Availability Source rows for Figures 2, 3, S1–S4 accompany this arXiv version underanc/source_data/, with result, manifest, audit, and source-table SHA256 values in the associated metadata JSON files. A compact reproducibility snapshot is provided as anc/reproducibility_bundle.zip. No human-participant or personal data were collected. Large model weights, raw activation tensors, intermediate checkpoints, and multi-gigabyte fitted artifacts are not redistributed. Code Availability The ancillary reproducibility bundle contains preregistrations, frozen configurations, experiment code, remote wrappers, in- dependent auditors, plotting scripts, tests, environment spec- ifications, audit reports, and source tables. The third-party state-probessubmodule is identified by its upstream URL and pinned commit but is not redistributed. Excluded large artifacts can be regenerated from the recorded model identifiers and frozen configurations; their hashes and passports remain in the included reports and metadata. Ethics Declaration This study used pretrained and from-scratch computational models, procedurally generated prompts, and finite synthetic transition systems. It involved no human participants, personal data, clinical data, or animal research. Author Contributions Dekun Yang: Conceptualization, Methodology, Software, Vali- dation, Formal analysis, Investigation, Data curation, Visualiza- tion, Writing – original draft, Writing – review and editing. Competing Interests The author declares no competing interests. Funding No external funding was received for this work. AI-Assistance Disclosure OpenAI Codex assisted with code generation, experiment orches- tration, validation scripts, figure production, evidence organiza- tion, and language drafting. Experimental claims and numerical values were checked against versioned machine-readable artifacts and independent audit outputs. The authors remain responsible for the scientific design, interpretation, and final text. 9 A Calibrated Test of Internal Action MapsarXiv preprint 8. Supplementary Results 8.1 Threshold sensitivity without verdict movement 0.40.50.60.7 H1 relative-error threshold 0.0 2.5 5.0 7.5 10.0 12.5 15.0 Action-seed cells passing (of 15) a H1 threshold sensitivity Frozen .50 gate All 15 pass at 0.58 0.40.60.8 Maximum direct-map discrepancy 0.10 0.15 0.20 0.25 0.30 Minimum probe advantage b H2 passes only after both gates move Frozen gate 0.700.750.800.850.90 Minimum h4 order d z 0.0 0.5 1.0 1.5 2.0 2.5 3.0 h4 seed cells passing (of 3) c Moving d z alone cannot rescue H2 Order sub-gate Joint H2 Figure S1: Frozen thresholds expose different degrees of boundary sensitivity. a, Number of 15 h28 action-by-seed cells below each H1 threshold. The red line is the frozen.50gate; all cells first pass at.58. b, H2 verdict over direct-map-discrepancy and probe-gain thresholds while other gates stay fixed; the cross is the frozen(.50,.20)point. c, number of Phase 7 h4 seed cells passing the order sub-gate or full H2 as the minimum 푑 푧 varies. Relaxing 푑 푧 alone never changes joint H2 because direct and probe gates still fail. No panel changes an original decision. 8.2 Failed H5 composite and collapse diagnostic 0.250.300.350.400.450.500.550.60 Algebra violation 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 1.00 World incoherence a Partial Spearman ρ = 0.091 Cluster 95% CI [-0.035, 0.249] Task/state gate failed Preregistered H5 stitching is unsupported Scale Small Medium Large 0.00.20.40.60.8 State decoding accuracy 0.0 0.2 0.4 0.6 0.8 H1 a ffi ne error b H1 gate Hollow: initialization Filled: final Apparent fit rises with state differentiation SmallMediumLarge H1 closure H2 composition H3† AUROC H4 inverse −0.6 −0.4 −0.2 0.0 0.2 Signed normalized gate margin c Positive = sub-gate passed; †diagnostic only Relative discrimination survives absolute failures 0.00.20.40.60.81.0 State decoding accuracy Figure S2: State differentiation exposes failure of the preregistered H5 composite. a, Algebra violation and world incoherence across 99 nonzero checkpoints; partial Spearman휌= .091with trajectory-cluster 95% CI[−.035,.249], and the task gate fails. b, One trajectory shows H1 error rising as collapsed state representations differentiate. c, Signed margins use(.5− H1)/.5,(.6− H2)/.6,(AUROC−.8)/.8, and(.5− H4)/.5. H3 remains diagnostic and cannot override absolute gates. 10 A Calibrated Test of Internal Action MapsarXiv preprint 8.3 Layer-local composition gates Endpoint error Direct-map gap Order dz Probe gain Joint H2 h4 · seed 20260808 h4 · seed 20260809 h4 · seed 20260810 h16 · seed 20260808 h16 · seed 20260809 h16 · seed 20260810 h28 · seed 20260808 h28 · seed 20260809 h28 · seed 20260810 0.540.540.79 9 p FAIL 0.540.540.78 8 p FAIL 0.540.540.77 6 p FAIL 0.720.710.724 ppFAIL 0.730.720.79-4 ppFAIL 0.720.710.855 ppFAIL 0.810.790.96 22 p FAIL 0.810.800.99 13 p FAIL 0.810.800.92 14 p FAIL Early-layer maps fail the frozen composition construct −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 Signed normalized gate margin Figure S3: Early-layer one-step geometry does not satisfy the frozen composition construct. Raw cell text reports endpoint error, direct-map gap, order푑 푧 , probe gain, and joint H2 status. Color is signed normalized distance from the frozen sub-gate. h28 rows are frozen compatibility comparators. No layer-seed cell passes joint H2. 8.4 Layer scale and effective dimension h4h16h28h36 50 100 150 Mean norm a Residual norm h4h16h28h36 0 10 20 30 Mean norm b Action displacement h4h16h28h36 1.0 1.5 2.0 Participation ratio c Effective rank h4h16h28h36 0.3 0.4 0.5 Relative error d RMS-normalized H1 Figure S4: Scalar scale control preserves the sampled-depth ordering, while early representations are low dimensional. a, Mean residual norm. b, Mean action displacement norm. c, Covariance participation ratio. d, One-step error after train-fitted scalar RMS normalization. Points are three activation datasets; black bars are seed means. 11 A Calibrated Test of Internal Action MapsarXiv preprint References Zhiyu An and Wan Du. Representational homomorphism predicts and improves compositional generalization in transformer language model, 2026. URL https://arxiv.org/abs/2601.18858. Antoine Bosselut, Omer Levy, Ari Holtzman, Corin Ennis, Dieter Fox, and Yejin Choi. Simulating action dynamics with neural process networks. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=SJUEXlDxf. Yuxi Chen, Suwei Ma, Tony Dear, and Xu Chen. Transformers learn transition dynamics when trained to predict markov decision processes. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 207–216. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.blackboxnlp-1.13. URL https://aclanthology .org/2024.blackboxnlp-1.13/. Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space. In Conference on Language Modeling, 2025. URL https://arxiv.org/abs/2412.06769. Roee Hendel, Mor Geva, and Amir Globerson. In-context learning creates task vectors. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 9318–9333. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.findings-emnlp.624. URL https://aclantholo gy.org/2023.findings-emnlp.624/. Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. Lin- earity of relation decoding in transformer language models. In Inter- national Conference on Learning Representations, 2024. URL https: //openreview.net/forum?id=w7LU2s14kE. John Hewitt and Percy Liang. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, pages 2733–2743. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1275. URL https://aclanthology.o rg/D19-1275/. Apoorv Khandelwal and Ellie Pavlick. How do language models compose functions?, 2026. URL https://arxiv.org/abs/2510.01685. Najoung Kim and Sebastian Schuster. Entity tracking in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3835–3855. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.acl-long.213. URL https://aclanthology.org/2023.acl-long.213/. Jeonghoon Lee. A held-out transition-pair falsifier for long-horizon non-abelian state tracking, 2026. URL https://arxiv.org/abs/2606.07254. Belinda Z. Li, Maxwell Nye, and Jacob Andreas. Implicit representations of meaning in neural language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1813–1827. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.acl-long.143. URL https://aclanthology.org/2 021.acl-long.143/. Belinda Z. Li, Zifan Carl Guo, and Jacob Andreas. (How) do language models track state? In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 34429–34452. PMLR, 2025. URL https://proceedings.mlr.press/v267/li25r .html. Samuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan, Pierce Zhang, Salma Ziadi, Oliver Eberle, and Ida Momennejad. Algorithmic primitives and compositional geometry of reasoning in language models, 2026. URL https://arxiv.org/abs/2510.15987. Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. Fine-tuning enhances existing mechanisms: A case study on entity tracking. In International Conference on Learning Representations, 2024. URL https://arxiv.org/abs/2402.14811. Nikhil Prakash, Natalie Shapira, Arnab Sen Sharma, Christoph Riedl, Yonatan Belinkov, Tamar Rott Shaham, David Bau, and Atticus Geiger. Language models use lookbacks to track beliefs, 2026. URL https://arxiv.org/abs/2505 .14685. Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. Function vectors in large language models. In International Conference on Learning Representations, 2024. URL https: //arxiv.org/abs/2310.15213. Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. Evaluating the world model implicit in a generative model, 2024. URL https://arxiv.org/abs/2406.03689. 12