Paper deep dive
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt, Mingyuan Wu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/20/2026, 2:51:38 PM
Summary
This paper introduces the first framework for transparent circuit tracing in Vision-Language Models (VLMs) to analyze multimodal reasoning. The authors utilize transcoders to decompose neural representations into interpretable monosemantic features, construct attribution graphs to map causal relationships, and perform attention-based analysis. They demonstrate that VLMs hierarchically integrate visual and semantic concepts, with distinct circuits handling mathematical reasoning and cross-modal associations. The framework is validated through feature steering and circuit patching, proving these circuits are causal and controllable.
Entities (8)
Relation Signals (6)
Gemma-3 4B IT â uses â SigLIP
confidence 95% · Gemma-3-4B-it processes images using a SigLIP [39] vision encoder
Transcoders â enables â circuit tracing
confidence 92% · We introduce the first framework for transparent circuit tracing in VLMs... By utilizing transcoders
Transcoders â decomposes â neural representations
confidence 90% · We insert and train per-layer-transcoders into VLMs to decompose multimodal, polysemantic representations
Attribution Graphs â reveals â causal relations
confidence 88% · We trace attribution graphs... to reveal causal relations between features
Circuit Patching â validates â causal circuits
confidence 85% · Validated through feature steering and circuit patching, our framework proves these circuits are causal
VLMs â performs â Multimodal Reasoning
confidence 80% · Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce... to systematically analyze multimodal reasoning.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and semantic concepts. We reveal that distinct visual feature circuits can handle mathematical reasoning and support cross-modal associations. Validated through feature steering and circuit patching, our framework proves these circuits are causal and controllable, laying the groundwork for more explainable and reliable VLMs.
Tags
Links
- Source: https://arxiv.org/abs/2602.20330v1
- Canonical: https://arxiv.org/abs/2602.20330v1
Trouble viewing inline? Open PDF directly â
Full Text
57,922 characters extracted from source content.
Expand or collapse full text
Circuit Tracing in VisionâLanguage Models: Understanding the Internal Mechanisms of Multimodal Thinking Jingcheng Yang1* Tianhu Xiong1* Shengyi Qian2â Klara Nahrstedt1 Mingyuan Wu1 1University of Illinois Urbana-Champaign 2Independent Researcher jy95, mw34, klara@illinois.edu Abstract Visionâlanguage models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and semantic concepts. We reveal that distinct visual feature circuits can handle mathematical reasoning and support cross-modal associations. Validated through feature steering and circuit patching, our framework proves these circuits are causal and controllable, laying the groundwork for more explainable and reliable VLMs. Our code and models are available at https://github.com/UIUC-MONET/vlm-circuit-tracing. â footnotetext: * Jingcheng and Tianhu contributed equally.â footnotetext: â This work was conducted independently and is not related to the authorâs employment at Meta FAIR. 1 Introduction Figure 1: Given an image and a prompt, how can we extract a circuit, as an internal computation graph, of open-source vision language models such as Gemma3-4B [34] from Google. We introduce the first framework for successful circuit tracing in VLMs, enabling analysis of the internal circuits and association of concepts in underlying multimodal reasoning. The rapid advancement of visionâlanguage models (VLMs) has fundamentally transformed how machines understand and reason about multimodal information. Contemporary VLMs such as CLIP [29], Flamingo [1], LLaVA [20], and GPT4-o [35] demonstrate remarkable capabilities in tasks ranging from visual question answering and image captioning to complex visual reasoning and embodied AI applications. These models seamlessly integrate visual perception with linguistic understanding, enabling machines to answer questions about images, generate detailed descriptions of visual scenes, and even perform multi-step reasoning that requires coordinating information across modalities. Despite these impressive empirical successes, a critical question remains largely unanswered: How do these models actually work internally? Understanding how VLMs work is essential for building trustworthy, controllable AI. Although they are used in high-stakes areas like medical imaging, autonomous driving, and content moderation, their decision-making remains opaque. This lack of interpretability makes it hard to diagnose errors, mitigate biases, and ensure alignment with human values. It also limits scientific insightâreverse-engineering their mechanisms can reveal how vision and language interact and guide the design of more capable, efficient architectures. Recent work in interpretability has begun revealing the internal algorithms used by language models through attention visualization [33], probing, and circuit discovery [3, 15]. Yet these methods focus almost entirely on text-only models. VLMs introduce deeper challenges: they must integrate two modalities with different statistics and semantics while discovering meaningful visualâlinguistic correspondences. How VLMs bind visual features to tokens, implement cross-modal reasoning, or coordinate visual and linguistic attention remains largely unknownâposing a more complex frontier than interpretability work in single-modality text models or early visual interpretability research [38, 26]. This work introduces the first framework for successful circuit tracing in VLMs, enabling systematic analysis of the internal computational mechanisms underlying multimodal reasoning. Our approach builds on recent advances in interpretability for language models. Specifically, we leverage transcoders [9] to decompose neural representations into interpretable features, and combine these with attribution-based circuit discovery methods [3, 15, 30] to map causal relationships between features. We are the first to extend these techniques to the multimodal setting. In doing so, we address the unique challenges posed by visionâlanguage integration and develop new methods to trace information flow from visual inputs through the modelâs reasoning process to final outputs. Extending circuit tracing to the multimodal domain is far more challenging than we thought, but brings richer insights and opportunities in VLMs than we anticipate. Make Circuit Tracing Happen in VLMs. We build three key components in VLMs: 1. Transcoders in VLM: We insert and train per-layer-transcoders into VLMs to decompose multimodal, polysemantic representations into interpretable, monosemantic features. (Sec. 3.1) 2. Attribution Graph: We trace attribution graphs that incorporate image-embedding residuals to reveal causal relations between features in the resulting multimodal computational graph. (Sec. 3.2) 3. Circuit Discovery of Visual and Text Tokens: We propose attention analysis to interpret unnamed multimodal features, allowing for fully interpretable multimodal circuit discovery. (Sec. 3.3) Representative Insights from Multimodal Circuits. These circuits are not mere post-hoc correlations but represent genuine causal mechanisms which help interpret VLM internal computation. In Section 5 we include: 1. Hierarchical Integration of Visual and Semantic Concepts: Features that simultaneously encode both semantic and visual concepts emerge only in higher layers of the network. 2. Highly Interpretable Case Studies: for Visual Math reasoning, Six Finger Hallucination, Association between Mars and Space Shuttle. 3. Distinct Visual Latent Space in Language Model: We find that visually similar features cluster and co-activate, indicating that the language-model component of the VLM preserves a distinctly visual representation space. Intervention with Multimodal Circuits. Intervention means changing the activation of features in the circuits, enabling more opportunities: 1. Steering: We modify the activation of certain features to observe the impact it has on the final output. (Sec. 3.5) 2. Circuit Patching: We transfer entire patches of one circuit to another circuit of similar structure and function, to see whether the transplanted circuit performs the same. (Sec. 4.5) In summary, our contributions are threefold: we establish the first circuit tracing framework for VLMs, uncover key insights into multimodal reasoning with circuits, and demonstrate through intervention experiments the potential for circuit-based model manipulation and control. 2 Related Work Mechanistic interpretability in LLMs aims to causally reverse-engineer how model components implement behavior, going beyond earlier correlational interpretability methods such as probing and attention analysis [36, 5, 16]. In language models, this literature includes circuit-level analyses of specific behaviors (e.g., induction heads, IOI, and factual recall/localization) [12, 27, 37, 22, 23], intervention-based localization methods such as activation patching and path patching [40, 14], and feature-level approaches motivated by superposition, where sparse autoencoders are used to recover more interpretable features than individual neurons [13, 11, 4]. The field has also become more scalable through partially automated circuit discovery (ACDC) and shared tooling platforms such as Neuronpedia [6, 18]. Sparse autoencoders and transcoders. Neural representations are often polysemantic, responding to multiple unrelated concepts [32]. Sparse autoencoders (SAEs) [10] mitigate this by learning sparse, overcomplete decompositions that yield more interpretable, often monosemantic features [8]. Transcoders [9] extend this idea by replacing MLP layers with SAEs trained to match the MLPâs outputs, preserving the modelâs computation while exposing feature-level structure. This enables end-to-end attribution and circuit tracing. While transcoders have been used to analyze language models [9], their application to VLMs remains unexplored. Visionâlanguage model interpretability. While interpretability methods have advanced significantly for language models, understanding visionâlanguage models remains challenging. Most existing work on VLM interpretability focuses on high-level analysis: attention visualization between image and text tokens [33], probing for learned visual concepts [26], and analyzing cross-modal alignment [29]. More recently, several works have attempted to localize visual knowledge in VLMs [25] or understand how visual information is processed through the network [17]. However, these methods typically analyze individual components in isolation or rely on correlation-based analysis rather than causal circuit discovery. Our work is the first to apply circuit tracing methodology to VLMs, enabling systematic identification of the computational mechanisms underlying multimodal reasoning. 3 Circuit Tracing in VLMs In this section, we present our framework for circuit tracing in visionâlanguage models. We first describe how we train transcoders to decompose MLP activations into sparse, interpretable features (Section 3.1). We then explain how attribution graphs are computed to capture causal relationships between features across layers (Section 3.2). Next, we detail our approach to analyzing and interpreting learned features through activation patterns and attention visualization (Section 3.3). We describe our circuit discovery algorithm that identifies minimal computational subgraphs responsible for specific capabilities (Section 3.4). Finally, we present intervention and steering methods to causally validate discovered circuits (Section 3.5). Figure 2: A summary of our method, we first train transcoders for Gemma-3-4b-it on our curated dataset, yielding a replacement model with monosemantic features. Then, we utilize feature activation analysis to obtain information that would provide interpretable information about the features. Then, we generate an attribution graph of a given prompt, and utilize human experts to derive the final circuit. 3.1 Transcoders Sparse Autoencoders (SAEs) [8] provide a sparse, interpretable basis for transformer activations, but they do not cleanly expose causal relations between features, limiting circuit discovery. Whereas SAEs are trained directly to reconstruct a transformerâs activations, transcoders [9] replace a transformerâs MLP sublayer with a sparse autoencoder, maintaining computational equivalence while enabling feature-level analysis. For each MLP layer in the visionâlanguage model, we train a transcoder consisting of an encoder and decoder. The encoder maps the MLP input xââdmodelx ^d_model to dfeatâ«dmodeld_feat d_model learned latent features via: zâ(x)=ReLUâ(Wencâx+benc),z(x)=ReLU(W_encx+b_enc), (1) where WencââdfeatĂdmodelW_enc ^d_featĂ d_model and bencââdfeatb_enc ^d_feat are learned parameters. The decoder reconstructs an approximation to the original MLP output: TCâ(x)=Wdecâzâ(x)+bdec,TC(x)=W_decz(x)+b_dec, (2) where WdecââdmodelĂdfeatW_dec ^d_modelĂ d_feat and bdecââdmodelb_dec ^d_model. Unlike the original transcoder implementation [9], which uses an â1 _1 penalty to encourage sparsity, we follow EleutherAI [10] and enforce sparsity directly via TopKâ(zâ(x),k)TopK(z(x),k), retaining only the k largest activations. This removes the need for a sparsity coefficient and yields more stable training with consistent sparse features. We evaluate reconstruction quality using the Fraction of Variance Unexplained (FVU): FVU=1nââi=1n(yiây^i)21nââi=1n(yiâyÂŻ)2=MSEVarâ(y).FVU= 1n _i=1^n(y_i- y_i)^2 1n _i=1^n(y_i- y)^2= MSEVar(y). (3) Where y is MLPâ(x)MLP(x) (the original MLP output) and y y is TCâ(x)TC(x). Training minimizes only the reconstruction error, while sparsity is controlled solely by the choice of k (rather than through a tunable penalty). Each transcoder feature is defined by a paired encoder column and decoder row, contributing additively to the output. The expanded latent space (dfeat=Nlatentsâ dmodelâ Nlayersd_feat=N_latents· d_model· N_layers) enables polysemantic MLP representations to be factorized into sparse, approximately monosemantic, and thus interpretable features. While transcoders have been extended to model cross-layer structure [3], we use the per-layer formulation as originally introduced by Dunefsky et al. [9]. After training a transcoder for each MLP block, we construct a replacement model by substituting each MLP with its corresponding transcoder, yielding a network expressed entirely in sparse, learned latent features. Because transcoders only approximate the original MLPs, we explicitly track the reconstruction residual eâ(x)=MLPâ(x)âTCâ(x),e(x)=MLP(x)-TC(x), (4) computed from cached MLP outputs and included as a separate error node in the circuit graph. This accounts for approximation error without altering the forward pass of the replacement model. Figure 3: Transcoder vs SAE; SAEs learn to reconstruct model activations, whereas transcoders imitate MLP sublayersâ input-output behavior. 3.2 Attribution Graphs To trace circuits, we compute an attribution graph as introduced by Lindsey et al. [19] and Ameisen et al. [3], adapted for per-layer transcoders by Hanna et al. [15], that linearly decomposes how features contribute to activations in upper layers and, ultimately, to the output logits on a fixed prompt. Because each transcoder replaces the MLP with a sparse linear readout from features, and all nonlinearities (ReLUs, attention patterns, and normalization factors) are frozen at their values on the given prompt, the model becomes locally linear around that input. Each node in the graph corresponds to a token embedding, an active transcoder feature at a particular (layer, position) pair, or an output logit. For a source node s and target node t, the attribution is defined as: Asât=asâwsât,A_sâ t=a_s\,w_sâ t, (5) where asa_s is the activation magnitude of source feature s and wsâtw_sâ t is the virtual weightâthe local derivative of tâs pre-activation with respect to sâs activation. For per-layer transcoders, this weight factors into the decoder of s, the frozen transformer Jacobian, and the encoder of t: wsât=fdec(s)â€âJ(s)â(t)âŒâfenc(t),w_sâ t=f_dec^(s)\, \,J _(s)â(t)\,f_enc^(t), (6) where fdec(s)ââdmodelf_dec^(s) ^d_model is the decoder vector for feature s, fenc(t)ââdmodelf_enc^(t) ^d_model is the encoder vector for feature t, and J(s)â(t)âŒââdmodelĂdmodelJ _(s)â(t) ^d_modelĂ d_model is the residual-stream Jacobian of the original transformer with stop-gradients on all nonlinearities (LayerNorm, attention softmax, and ReLU). Since the model is linearized, the pre-activation of each node is exactly the sum of its incoming attributions: ht=âsâpredâ(t)Asât,h_t= _s (t)A_sâ t, (7) yielding a complete additive explanation of how earlier representations influence later ones. We construct the full directed attribution graph G=(V,E)G=(V,E) for each prompt, where nodes V include all token embeddings, active features, and output logits, and edges E have weights AsâtA_sâ t. We prune edges with negligible attribution (|Asât|<Ï”|A_sâ t|<Δ) to produce a sparse, interpretable graph. 3.3 Feature Interpretation and Attention Analysis Feature activation analysis. To understand what each transcoder feature represents, we analyze its activation patterns across a diverse dataset of visionâlanguage inputs. For each feature fif_i, we collect the top-k activating examplesâimageâtext pairs (I,T)(I,T) that produce the highest activation ziz_iâand examine commonalities. We compute activation statistics including the featureâs activation frequency (fraction of inputs where zi>0z_i>0), mean activation magnitude, and position distribution (which tokens or image patches most frequently activate the feature). Vision encoder attention maps. While feature activations on text tokens are directly interpretable, activations on image tokens remain opaque. To visualize which image regions the SigLIP vision encoder attends to, we compute attention-rollout maps for each of the visual embeddings passed to the language model. In the SigLIP vision-encoder for Gemma 3, it processes a resized 896Ă896896Ă 896 image and reduces the representation to a fixed set of Nv=256N_v=256 visual tokens that are then input to the language model. We compute attention-rollout maps over the last K self-attention layers of the vision tower. Let the encoder output T tokens in total (including possibly a class token), of which NvN_v are the visual patch tokens entering the language model. For each of the last K layers â , we select the fraction q of attention heads with lowest mean entropy (i.e., the most focused heads), average their (TĂT)(TĂ T) attention matrices to obtain AÂŻ(â) A^( ), add an identity residual, and row-normalize to form a row-stochastic matrix A~(â) A^( ). The rollout is R=A~(LâK+1)âA~(LâK+2)ââŻâA~(L)ââTĂT.R\;=\; A^(L-K+1)\, A^(L-K+2)·s A^(L)\;â\;R^TĂ T. We then extract the NvĂNvN_vĂ N_v sub-matrix RvisR_vis corresponding to the visual tokens and interpret each row as an attention distribution over the 256 visual token positions. To provide spatial localization, we recover the original spatial grid of image patches (e.g., ghĂgwg_hĂ g_w) from the encoder, pool tokens within non-overlapping blocks of size bĂbĂ b, reshape into a ghĂgwg_hĂ g_w grid, upsample to the input image resolution, and normalize each map. The resulting grayscale heat-maps indicate which image regions are most attended to in the vision tower. We chose not to add additional refinements, denoising or post-processing so as to present the raw information for human experts to analyze. 3.4 Circuit Discovery A circuit is an abstract representation of the computational graph that explains a modelâs output logits for a given input. Instead of tracking every feature activation, through explaining the feature representations as described in Section 3.3, we can group features of similar functions into shared nodes, resulting in a simplified graph that explains the modelâs behavior. This mechanistic structure highlights the key computational pathways and interactions that drive the modelâs behavior, allowing us to steer and intervene in the modelâs activations towards predictable directions. While there are now methods for large-scale, partially automated circuit discovery [7], a human-annotated circuit remains the most accurate and interpretable representation of a modelâs mechanism. In this work we therefore use human experts to discover and annotate circuits. 3.5 Intervention and Steering To study how individual transcoder features influence model behavior, we directly modify feature activations during the forward pass and observe the resulting changes in the modelâs output. Feature representation. At layer â and position t, the transcoder produces feature activations zâ,t,iââ,z_ ,t,i , and each feature has an associated decoder vector dâ,iââdmodeld_ ,i ^d_model that determines how changes to that feature are written back into the residual stream. Interventions. An intervention specifies a target value vâ,t,iv_ ,t,i for a feature. During the forward pass we compute Îâzâ,t,i=vâ,t,iâzâ,t,iâ(x), z_ ,t,i=v_ ,t,i-z_ ,t,i(x), and apply the corresponding residual update hâ,tâhâ,t+Îâzâ,t,iâdâ,i.h_ ,t\;â\;h_ ,t+ z_ ,t,i\,d_ ,i. This simply adjusts the modelâs internal representation as if the feature had taken the desired value. Circuit patching refers to directly overwriting selected internal activations during the forward pass, or transplanting entire sub-circuits from another circuit, and observing how the modelâs output changes. In our case, we patch transcoder features by setting them to chosen target values (e.g., zero for ablation or a positive constant for amplification) at specific layers and positions. We observe whether transferring a patch from Circuit A onto B would recreate similar behaviors of A for B. All experiments are performed on the transcoder-replaced model, and we qualitatively inspect the differences in predictions to understand the behavioral role of the intervened features. 4 Experiments Model. We conduct our circuit tracing experiments on Gemma-3-4B-it [34], a state-of-the-art open-source visionâlanguage model. Gemma-3-4B-it processes images using a SigLIP [39] vision encoder, which produces patch tokens that are concatenated with text tokens and fed into a transformer decoder. This architecture enables joint visualâlinguistic processing and makes the model suitable for multimodal reasoning. Architecture. The model consists of a SigLIP encoder with patch size 1414, applied to 896Ă896896Ă896 input images. This produces a grid of 64Ă6464Ă 64 patches, yielding 40964096 initial patch tokens [34]. These tokens are then pooled into 256256 soft image tokens before being appended to the text sequence for decoding. Gemma-3 models employ bidirectional attention over image tokens, allowing the model to attend to the full image context throughout processing. The decoder is a 34-layer transformer with dmodel=2560d_model=2560, 3434 attention heads, and MLP dimension dff=10240d_f=10240. 4.1 Training Transcoders Name Content Size SmoLIM2-135M-10B [2] 2048 text tokens 144,000 ImageNet [31] 1 image; Caption 144,000 Cauldron [17] 1 image; QA Text 72,000 Table 1: Datasets used for Transcoder Training. Datasets. We train transcoders on a diverse corpus of visionâlanguage examples to ensure broad coverage of multimodal concepts and reasoning patterns. Our training dataset consists of a base set of pure text dataset for broad coverage of various domains [2], alongside ImageNet [31] for rich visual features and Cauldron [17] consisting of QA and visual reasoning tasks. For the cauldron split, we sampled from its 50 subsets evenly to ensure even coverage. Implementation details. For transcoder training, we implemented our framework with support for Gemma-3 on top of Sparsify [10]. Training configurations. For every MLP layers in Gemma-3-4B-it, we train a separate transcoder with dfeat=[Nlatentsâ dmodelâ 34]d_feat=[N_latents· d_model· 34] features. NlatentsN_latents is the number of features in the hidden layer of a single transcoder block, which we then multiply by 34, the number of decoder layers in Gemma-3-4B. During the forward pass of the transformers, we only activate the top k latents as described in Section 3.1, for our experiment we set k to 48. Figure 4: Top: Percentage of dead latents (Dead PCT) across layers for different values of NlatentsN_latents. Bottom: Fraction of variance unexplained (FVU) when training transcoders on a text-only dataset compared to our multimodal split. Transcoder training uses the AdamW [21] optimizer with a learning rate 2Ă10â4Ă214NlatentsĂdmodel2Ă 10^-4Ă 2^14N_latentsĂ d_model, with NlatentsN_latents = 64, dmodeld_model = 2560. We also apply a linear warmup over the first Nwarmup=1,000N_warmup=1,000 steps. Training uses a batch size of 12 and runs for 30,000 steps on 8 H100 GPUs for approximately 60 hours. Expansion Factor. We trained transcoders with different expansion factor Nlatentsâ32,64,128N_latentsâ\32,64,128\. As shown in the top panel of Figure 4, the choice of NlatentsN_latents substantially affects the proportion of dead latents. We define a latent as dead if it fails to activate above a minimal threshold for an extended portion of training. A high fraction of dead latents indicates that the model is not utilizing the full capacity of the latent space, suggesting limited diversity or overlap of concepts represented within the corresponding features. The curves also illustrate how this behavior varies across layers: early layers (e.g., Layer 3) exhibit substantially higher dead-latent ratios, whereas mid-network layers (e.g., Layer 15) show far denser activation patterns. FVU and multimodality. The bottom panel of Figure 4 compares the FVU obtained when training transcoders on our text-only split (SmoLIM2-135M-10B) versus our full multimodal (text+image) dataset. Although FVU is computed in-domain and does not directly reflect out-of-domain explainability, transcoders trained with multimodal supervision consistently achieve lower FVU across all layers. This suggests that visual features provide additional constraints that make the underlying representations more explainable. The gap is largest in the middle layers (e.g. Layer 15), consistent with our broader findings (Sec. 5) that this is where visual information is integrated into unified semantic representations; thus, explaining these circuits requires both modalities. At higher layers, the representation space becomes more uniform, and the FVU gap correspondingly decreases. 4.2 Computing Attribution Graphs Implementation details. We computed attribution graphs using the circuit tracer library by Hanna et al. [15], where we extended support for Gemma-3 VLMs. Images in Gemma-3 are not masked [34], achieving full attention and allowing the model to see every part of the image in a bidirectional manner. Our attribution graphs are adapted to reflect this. To properly monitor residual streams and activation of individual transformer blocks, we utilize the TransformerLens library by Nanda and Bloom [24], with extended support for VLMs [25]. Attribution. For attribution graph computation, we continuously add nodes and edges to our attribution graph until it reaches a cumulative influence thresholds 0.80.8 and 0.980.98 respectively. Circuit discovery uses the top-m features ranked by partial influence, where m controls the maximum number of feature nodes (m=7500m=7500), and we select up to 10 logit nodes whose cumulative probability mass is at least 0.950.95. All attribution experiments run in bfloat16 precision to reduce memory usage. Computing an attribution graph for a simple image-text QA task takes 20 minutes on a single H100 GPU. Figure 5: Feature activations on images with a mix of feature annotations resulting from full images (right) and feature analysis resulting from noisy attention maps (left). 4.3 Feature Analysis Datasets. Our dataset for feature analysis is a subset of our transcoder training dataset. Including imageâtext pairs sampled from ImageNet (18,000 images) and Cauldron (10,000 images) [31, 17] as well as a text-only split of 10,000 Ă 2048 tokens [2]. For circuit discovery and validation, we use a task-specific dataset consisting of a small, curated sample of 20-100 internet images that are semantically related to the image we want to trace. Our results demonstrate this to be a highly effective and efficient method at improving a graphâs interpretability. Implementation Details. For computing activations, we base our method upon SAE Dashboard [30] for LLM feature activation analysis. We insert hookpoints at every MLP sublayerâs input and output to record activation patterns using TransformersLens extended with VLM support [24, 25]. For activations upon image tokens, we store the index of the image tokens and the image in order to retrieve its attention map during circuit discovery. Ad Hoc Feature Analysis. Whereas circuit tracing tools [15, 18] for LLMs precompute all feature activations, we selectively compute only the features in a given attribution graph, which usually contains under 2000 features given our pruning parameters, significantly reducing the compute and storage costs of caching activation data. While our curated dataset allows for general interpretability of various features in the attribution graph, it may not be sufficient. In our attribution graph analyzing the image of a sea otter, further calculating the feature activations on a curated, small dataset of 30 sea otter images significantly increased the featureâs interpretability and allows us to be more certain of a featureâs function. Computing feature activation on a single attribution graph with approx. 1000 features takes 20 H100 GPU hours on our dataset, with a negligible increase when injected with a curated dataset. Attention Maps. We extract the custom SigLip encoder from Gemma-3-4B-it, and use that to compute attention maps for images using our method described in Section 3.3. Computing attention maps is fast, and we pre-compute attention maps for all the 28,000 images, which takes about 2 hours on a single H100 GPU and results in approx. 2TB of single channel masks that are recombined with the base image when queried. Figure 6: We conduct circuit tracing on the same prompt This is the planet with images of mars and earth. We then conduct activation patching by suppressing the mid-layer visual features of mars in the mars circuit, and set the activation of earth visual features discovered in the earth circuit in the mars circuit 4.4 Finding Circuits To obtain the most accurate circuits for validation, we use human experts (Sec. 3.4) to compute the subgraph. We manually identify features that exhibit a similar function (e.g. represent similar concepts, has similar effects on model behavior), and group them into distinct nodes. The attribution between nodes is the sum of the attribution exerted by the features in each node. The resulting graph of nodes is generally an explainable, simplified circuit. 4.5 Intervention Experiments We use a combination of steering and activation patching (Sec. 3.5) to validate our circuits. Detailed experiments are available in the appendix. One example of activation patching is shown in figure 6, where we suppressed features representing the visual concept of Mars, and instead activated features representing the visual concept of Earth identified in the earth circuit, and the subsequent feature activations and final output all changed to earth-related concepts. 5 Empirical Findings Our circuit tracing framework reveals several core principles underlying how VLMs such as Gemma3 integrate and manipulate multimodal information. The results validate our methodology and offer a clearer picture of the computational structure supporting vision-language reasoning. Hierarchical Formation of Multimodal Representations. We observe a progressive integration of visual and semantic information as network depth increases. Only in higher layers (emerging around Layer 20) do features jointly encode both visual and semantic concepts. Earlier layers remain largely modality-specific, supporting the progressive binding hypothesis in which cross-modal associations are gradually assembled across depth. Granularity and Monosemanticity Across Layers. Feature representations become increasingly abstract with depth. Early layers exhibit highly localized, fine-grained visual patternsâdown to digits or texturesâwhile later layers form object- and concept-level features, paralleling trends in vision models [38, 25] but now aligned with semantics. Figure 7: Tracing a purely visual attribution path from an image of Mars reveals internal visual associations (e.g., âspace shuttleâ) even without supporting cues. Visual Circuits in Simple Mathematical Reasoning. For image-based arithmetic (e.g., 1+21+2 rendered visually), the model appears to compute partially within visual space. Intermediate layers contain visual features corresponding to the resulting numeral (e.g., â3â), activating across contexts. We also identify visual encodings of digit ranges and modular arithmetic patterns, echoing textual findings from Lindsey et al. [19]. These results suggest that simple arithmetic over images can rely on visual circuits rather than purely semantic computation. Understanding Hallucination: The Six-Finger Problem. Our tracing suggests that hallucinationsâsuch as the six-finger caseâarise from an interaction between perceptual bias and internal circuit dynamics rather than from a single failure mode. The vision encoder appears to produce embeddings that heavily emphasize generic âhandâ semantics, while the modelâs internal circuits further amplify these features. As a result, visual features corresponding to the digit â6â are suppressed toward the level of unrelated numbers, whereas hand-related features strongly activate the âfiveâ circuit. Although the model possesses circuits capable of visual counting (Fig. 5), these can be overshadowed by more dominant semantic and perceptual signals, showing how both encoder-level weighting and feature competition jointly contribute to hallucination. Parallel Visual and Semantic Pathways with Late Convergence. Gemma3 maintains distinct visual and semantic representation streams deep into the network. We identify visually grounded associative featuresâsuch as âspace shuttleâ activations triggered by an image of Mars (Fig. 7) âreflecting internal visual associations independent of semantics. High-level layers also preserve visual similarity (e.g., consistent activations for sea otters, seals, and beavers) even when semantic categories diverge. These streams ultimately merge in the final layers, where they align into a unified multimodal representation supporting coherent reasoning. 6 Conclusion This work introduces the first circuit-tracing framework for visionâlanguage models, revealing the mechanisms underlying multimodal reasoning. Using transcoders to extract interpretable features and attribution graphs to trace causal structure, we isolate sparse circuits for tasks such as object recognition, counting, QA, and captioning. Intervention experiments confirm these circuits are causally meaningful, enabling targeted impairment and controllable steering. Beyond advancing scientific understanding, this framework provides practical tools for debugging, mitigating failures, and guiding more interpretable VLM designâsupporting the development of transparent, controllable, and aligned AI systems. 7 Limitations and Future Work. Despite these contributions, many important limitations remain. First, the vision-encoder attention maps used for interpreting visual features can be difficult to read: they sometimes fail to localize relevant regions or provide meaningful contextual cues, limiting their utility for feature annotation. In addition, for tasks requiring additional processing of visual inputs (e.g., math operations), it is difficult to distinguish between features that mediate the computation and features that represent its output. Second, our use of per-layer transcoders cannot capture cross-layer superposition [3], which could be a more significant drawback for VLMs due to the high feature density of image embeddings; we frequently observe many near-duplicate visual features firing in attribution graphs, suggesting the need for cross-layer or more adaptive transcoder designs. In addition, our work does not investigate in detail regarding specific choices of the transcoder configuration, under the assumption that they are largely the same. Research into specific SAE methods (e.g. JumpRelu, BatchTopK, etc) and varying configurations for optimal transcoder training for VLMs could prove to be useful. Thirdly, by largely mirroring the circuit-tracing method for LLMs [3], the computational cost for understanding and explaining multimodal features have risen significantly. While we accommodated for this limitation by computing feature activations ad-hoc, we believe methods for comprehensive or automated interpretation of features could greatly reduce the complexity of our current process. Current methods for automatic feature interpretation [28] are computationally too expensive, and we call for future works to improve this process. In addition, our current analysis is limited to one specific model. We have reason to suspect some complexities in circuit-tracing might be complicated by Gemma3âs SigLip and bidirectional-attention mechanism. Extending this work to accommodate a wider collection of VLMs could greatly solidify findings and conclusions in our current work. Finally, circuit discovery currently requires substantial human effort, making it difficult to introduce quantitative evaluation or apply our method directly to model fine-tuning or improvement. Automating feature labeling, simplifying attribution graphs, or developing cross-layer transcoders could help scale circuit tracing to larger models and enable more systematic evaluation. 8 Acknowledgements This research used the Delta advanced computing and data resource, which is supported by the National Science Foundation (award OAC-2005572) and the State of Illinois. Delta is a joint effort of the University of Illinois Urbana-Champaign and its National Center for Supercomputing Applications. This work was supported by the National Science Foundation under awards CNS-2106592 and CCF-2217144. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation. References [1] J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. BiĆkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan (2022) Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 23716â23736. External Links: Link Cited by: §1. [2] L. B. Allal, A. Lozhkov, E. Bakouch, G. M. BlĂĄzquez, G. Penedo, L. Tunstall, A. Marafioti, H. KydlĂÄek, A. P. LajarĂn, V. Srivastav, J. Lochner, C. Fahlgren, X. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf (2025) SmolLM2: when smol goes big â data-centric training of a small language model. External Links: 2502.02737, Link Cited by: §4.1, §4.3, Table 1. [3] E. Ameisen, J. Lindsey, A. Pearce, W. Gurnee, N. L. Turner, B. Chen, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. Marcus, M. Sklar, A. Templeton, T. Bricken, C. McDougall, H. Cunningham, T. Henighan, A. Jermyn, A. Jones, A. Persic, Z. Qi, T. Ben Thompson, S. Zimmerman, K. Rivoire, T. Conerly, C. Olah, and J. Batson (2025) Circuit tracing: revealing computational graphs in language models. Transformer Circuits Thread. External Links: Link Cited by: §1, §1, §3.1, §3.2, §7, §7. [4] T. Bricken et al. (2023) Towards monosemanticity: decomposing language models with dictionary learning. Note: Transformer Circuits External Links: Link Cited by: §2. [5] K. Clark, U. Khandelwal, O. Levy, and C. D. Manning (2019) What does bert look at? an analysis of bertâs attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP, External Links: Link Cited by: §2. [6] A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso (2023) Towards automated circuit discovery for mechanistic interpretability. arXiv preprint arXiv:2304.14997. External Links: Link Cited by: §2. [7] A. Conmy, A. N. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso (2023) Towards automated circuit discovery for mechanistic interpretability. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), External Links: 2304.14997 Cited by: §3.4. [8] H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey (2023) Sparse autoencoders find highly interpretable features in language models. External Links: 2309.08600, Link Cited by: §2, §3.1. [9] J. Dunefsky, P. Chlenski, and N. Nanda (2024) Transcoders find interpretable llm feature circuits. External Links: 2406.11944, Link Cited by: §1, §2, §3.1, §3.1, §3.1. [10] Sparsify: transformers with saes and transcoders Note: GitHub repository, accessed 13 November 2025 External Links: Link Cited by: §2, §3.1, §4.1. [11] N. Elhage, T. Hume, C. Olsson, et al. (2022) Toy models of superposition. arXiv preprint arXiv:2209.10652. External Links: Link Cited by: §2. [12] N. Elhage et al. (2021) A mathematical framework for transformer circuits. Note: Transformer Circuits External Links: Link Cited by: §2. [13] M. Geva, R. Schuster, J. Berant, and O. Levy (2021) Transformer feed-forward layers are key-value memories. In EMNLP, External Links: Link Cited by: §2. [14] N. Goldowsky-Dill, C. MacLeod, L. Sato, and A. Arora (2023) Localizing model behavior with path patching. arXiv preprint arXiv:2304.05969. External Links: Link Cited by: §2. [15] M. Hanna, M. Piotrowski, J. Lindsey, and E. Ameisen (2025) Circuit-tracer. Note: https://github.com/safety-research/circuit-tracerThe first two authors contributed equally and are listed alphabetically. Cited by: §1, §1, §3.2, §4.2, §4.3. [16] S. Jain and B. C. Wallace (2019) Attention is not explanation. In Proceedings of NAACL-HLT, External Links: Link Cited by: §2. [17] H. Laurençon, L. Tronchon, M. Cord, and V. Sanh (2024) What matters when building vision-language models?. External Links: 2405.02246 Cited by: §2, §4.1, §4.3, Table 1. [18] J. Lin (2023) Neuronpedia: interactive reference and tooling for analyzing neural networks. Note: Software available from neuronpedia.org External Links: Link Cited by: §2, §4.3. [19] J. Lindsey, W. Gurnee, E. Ameisen, B. Chen, A. Pearce, N. L. Turner, C. Citro, D. Abrahams, S. Carter, B. Hosmer, J. Marcus, M. Sklar, A. Templeton, T. Bricken, C. McDougall, H. Cunningham, T. Henighan, A. Jermyn, A. Jones, A. Persic, Z. Qi, T. B. Thompson, S. Zimmerman, K. Rivoire, T. Conerly, C. Olah, and J. Batson (2025) On the biology of a large language model. Transformer Circuits Thread. External Links: Link Cited by: §11.4, §3.2, §5. [20] H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 34892â34916. External Links: Link Cited by: §1. [21] I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. External Links: 1711.05101, Link Cited by: §4.1. [22] K. Meng, D. Bau, A. Andonian, and Y. Belinkov (2022) Locating and editing factual associations in gpt. arXiv preprint arXiv:2202.05262. External Links: Link Cited by: §2. [23] K. Meng, A. S. Sharma, A. Andonian, Y. Belinkov, and D. Bau (2022) Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229. External Links: Link Cited by: §2. [24] N. Nanda and J. Bloom (2022) TransformerLens. Note: https://github.com/TransformerLensOrg/TransformerLens Cited by: §4.2, §4.3. [25] Y. Nikankin, D. Arad, Y. Gandelsman, and Y. Belinkov (2025) Same task, different circuits: disentangling modality-specific mechanisms in vlms. External Links: 2506.09047, Link Cited by: §2, §4.2, §4.3, §5. [26] C. Olah, A. Mordvintsev, and L. Schubert (2017) Feature visualization. Distill. Note: https://distill.pub/2017/feature-visualization External Links: Document Cited by: §1, §2. [27] C. Olsson, N. Elhage, N. Nanda, et al. (2022) In-context learning and induction heads. Note: Transformer Circuits External Links: Link Cited by: §2. [28] G. Paulo, A. Mallen, C. Juang, and N. Belrose (2025) Automatically interpreting millions of features in large language models. External Links: 2410.13928, Link Cited by: §7. [29] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021-18â24 Jul) Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, M. Meila and T. Zhang (Eds.), Proceedings of Machine Learning Research, Vol. 139, p. 8748â8763. External Links: Link Cited by: §1, §2. [30] D. Research (2024) SAE Dashboard. Note: https://github.com/jbloomAus/sae-dashboard Cited by: §1, §4.3. [31] O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei (2015) ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV) 115 (3), p. 211â252. External Links: Document Cited by: §4.1, §4.3, Table 1. [32] A. Scherlis, K. Sachan, A. S. Jermyn, J. Benton, and B. Shlegeris (2025) Polysemanticity and capacity in neural networks. External Links: 2210.01892, Link Cited by: §2. [33] S. Seo, S. Yoo, H. Lee, Y. Jang, J. H. Park, and J. Kim (2025-04) A sentence-level visualization of attention in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations), N. Dziri, S. (. Ren, and S. Diao (Eds.), Albuquerque, New Mexico, p. 313â320. External Links: Link, Document, ISBN 979-8-89176-191-9 Cited by: §1, §2. [34] G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. RamĂ©, M. RiviĂšre, et al. (2025) Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Cited by: Figure 1, Figure 1, §4.2, §4, §4. [35] O. Team (2024) GPT-4o system card. External Links: 2410.21276, Link Cited by: §1. [36] I. Tenney, D. Das, and E. Pavlick (2019) BERT rediscovers the classical nlp pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, External Links: Link Cited by: §2. [37] K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt (2022) Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. arXiv preprint arXiv:2211.00593. External Links: Link Cited by: §2. [38] M. D. Zeiler and R. Fergus (2014) Visualizing and understanding convolutional networks. In Computer Vision â ECCV 2014, D. Fleet, T. Pajdla, B. Schiele, and T. Tuytelaars (Eds.), Cham, p. 818â833. External Links: ISBN 978-3-319-10590-1 Cited by: §1, §5. [39] X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer (2023) Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, p. 11975â11986. Cited by: §4. [40] F. Zhang and N. Nanda (2023) Towards best practices of activation patching in language models: metrics and methods. arXiv preprint arXiv:2309.16042. External Links: Link Cited by: §2. Supplementary Material 9 Transcoder Training 9.1 Dead Latents As shown in Figure LABEL:fig:all_dead, a feature is marked as dead if it fails to activate sufficiently over a given interval, and it is considered active once it activates sufficiently again. The percentage of dead latents provides insight into how features utilize the available expansion size. A high percentage does not necessarily imply high polysemanticity; it may instead reflect insufficient expansion capacity to support fully monosemantic features, causing the model to default to a denser cluster of max-k features that occupies less space than the nominal expansion factor. We hypothesize that the lower layers may either represent condensed, naturally superpositioned embeddings originating from the vision encoders, or may require a much larger expansion factor to be properly disentangled. In either case, we currently lack a mature method for accurately extracting low-level representationsâsuch as patterns, colors, or other subtle featuresâthat we believe are present in these layers, due to the difficulty of identifying consistent patterns when examining activation example images in aggregate. We plan to follow up on this in future works. Layer 11, being a global-attention layer, stands out from the local layers around it and creates a clear shift in the modelâs training curve. Because it suddenly incorporates full-context information, its behavior differs enough to change the curvature. Its transcoder is harder to interpret for the same reasonâit mixes information from across the entire sequence rather than nearby tokens. Even so, this complexity is contained within the layer, so it doesnât disrupt the rest of the circuit-tracing process. 10 Feature Discovery Feature discovery on multi-modal inputs is significantly more expensive than on text-only LLMs. For text activations, a single token with its surrounding context is often sufficient to explain a featureâs behavior. As a result, a paragraph of a few hundred tokens can provide hundreds of useful activation examples. In contrast, image inputs in VLMs produce far more tokens. For Gemma-3-4B, each image yields 256 tokens. Although a sufficiently dense image could, in principle, provide unique and informative signals across all 256 embeddings, most images do not contain enough complexity for all tokens to meaningfully activate distinct features. In typical or unfavorable cases, the full set of 256 image tokens may provide no more explanatory coverage than a single text token. Moreover, the vision encoder introduces additional computational overhead. These factors are the primary reason we have not yet analyzed all VLM features. We hope to optimize our pipeline and explore alternative methods for obtaining feature activations more efficiently. 11 Circuits 11.1 Hallucination - Fingers Figure 10: Circuit tracing analysis of prompt âThe number of fingers in the image is â with the result being 5, despite image containing six fingers. We analyzed the common hallucination in which the model predicts five fingers instead of six. Remarkably, for this model, the logit for the token 6 is no higher than for other unrelated digits (e.g., 7 or 1). Although we have previously identified features involved in visual counting tasks (e.g., features that activate on groups of three), we were unable to find any feature corresponding to the concept of six or six visual objects in the attribution graph. Instead, we found a direct circuit between the image embeddings, features associated with the number 5, and the final output for the â5â token. Notably, contrary to the typical assumption that the model would rely on a semantic âhandâ concept, the circuit appears to be driven by a visual feature representing the visual concept of a hand together with the number five. In other words, in this task, the visual concept of hand activates the concept of five, rather than the model performing a robust object-counting procedure. We also observed features related to the number 2 in upper layers, whichâpossibly due to the inputâbecome activated by the concept of a hand and by features representing the function âsay a number.â We hypothesize that this may arise because hands occur in pairs across animals, causing the modelâs number-related circuits to be influenced by hand priors rather than meaningful counting information from the encoder. Overall, this suggests that the error may stem from the encoder failing to provide a sufficiently strong representation of a six-fingered handâor from the model implicitly choosing to ignore such evidence. 11.2 Caption - Sea Otter Figure 11: Circuit tracing analysis of a seaotter, with interesting feature attributions. We conducted circuit tracing on an image of a seaotter, which required two-step circuit analysis due to seaotter being two tokens. We observe for the token âseaâ the logic was in the 90+% range, and we analyzed the circuit prompted seaotter. Interestingly, we also found that this image strongly activated a feature that represented sea lions - likely due to their visual similarities (the feature was also present in the first circuit without the sea token input), this shows that there may exist a purely visual latent space in the model. We also found another feature that included visual images of animals that contained the word âseaâ, such as sea urchins, sea slug, and sea cucumbers. There also exists a knowledge feature of a geographical range, which just happens to be the pacific coastline from California to Japan - where sea otters primarily live in. 11.3 Caption - Mars Figure 12: Analysis of the mars circuit, revealing a transition from purely visual features to a semantically grounded feature that controls the final output. We analyzed the mars prompt in Figure 12, which reveals a clean and interpretable circuit showing how raw visual representations from the encoder are gradually transformed into a unified multimodal feature representing the concept of Mars. This joint feature blends both the semantic textual notion of the word âmarsâ and the visual appearance of the planet, demonstrating how the model builds a shared representation that supports cross-modal grounding. The circuit further illustrates how early visual featuresâsuch as a generic âplanetâ detectorâfeed into progressively more specific features until they converge on a single joint concept that dominates the logit contribution for the final output token. Notably, the circuit also highlights a broader associative structure present in mid-lower layers. Features representing planets reliably activate features associated with rockets and space shuttles, even though these objects do not appear in the input image. This suggests that the model has learned a latent web of visual associations that mirrors human conceptual priors: planets evoke spacecraft, just as certain animals evoke particular habitats or behaviors. These associations arise purely from visual features rather than explicit textual grounding, indicating that VLMs develop an internal âassociation of ideasââwhere visually related objects co-activateâeven before higher-level semantic features take over. This phenomenon provides insight into how VLMs integrate visual context, prior knowledge, and semantic structure when generating grounded descriptions. 11.4 Reasoning - Simple Addition As shown in Figure 13, we analyzed a simple addition problem with an image of an incomplete equation. We found that at lower layers, the image embeddings activated features related to numerals, such as a feature representing a number between 0-5, showing that visual-semantic convergence may also occur at lower layers for low-level representations. We also note that the model contains a very diverse feature representation of numbers. For example, we identified a feature that primarily activates on images of charts, with the attention roughly focusing on the 3 range of the axis. This feature is potentially used for visual-numerical reasoning. We also found features activating on objects in groups of three, this shows that object count information is passed from the vision encoder, and that the VLM decoder contains features that can properly represent this. We note that this is entirely absent in the finger hallucination example which prompts us to incline towards the vision encoder not encoding such information properly. In the middle-upper layers, we discovered a group of semantic features that activates on mathematical operations, and we hypothesize that this is the subgraph that performs this operation in semantic space. Analyzing the purely semantic addition circuit is beyond the scope of our analysis, and we recommend referring to the circuit tracer paper by Lindsey et al. [19] for a deeper analysis. Figure 13: Circuit tracing analysis of a simple addition task. 12 Conclusion In this appendix, we presented additional analyses generated using our proposed VLM circuit-tracing methodology, illustrating its ability to reveal structure across visual, semantic, and multi-modal representations. Through examinations of transcoder behavior, feature activation patterns, and several representative circuits, we showed how VLMs blend visual features with language-level abstractionsâsometimes in unintended ways. These case studies highlight several recurring themes: the limitations of current vision encoders in providing robust object-level information; the presence of latent spaces that mix visual similarity, linguistic priors, and associative knowledge; and the emergence of compact, semantically meaningful features even within early or visually dominated layers. Our circuits serves as a useful tool to debug hallucinations, false knowledge, and internal mechanisms of VLMs for future improvements. Overall, these findings demonstrate both the promise and challenges of interpreting VLMs. While our method successfully uncovers many of the mechanisms driving model behavior, it also reveals gapsâsuch as missing low-level disentangled features or overreliance on semantic priorsâthat motivate further refinement of both architectures and interpretability tools. We hope that the insights provided here, along with the methodology introduced in the main paper, will help pave the way for a more systematic understanding of multimodal reasoning, its failure modes, and the underlying circuits that support these abilities in modern VLMs.