Paper deep dive
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
Xu Liu, Yongheng Zhang, Qiguang Chen, Yao Li, Sheng Wang, Libo Qin
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 2:33:58 AM
Summary
DaP-ICoT is an Interleaved-modal Chain-of-Thought reasoning framework that improves efficiency and performance in Multimodal Large Language Models (MLLMs) by introducing Dynamic Visual Thought Integration (DVTI) and Precise Visual Thought Guidance (PVTG). These components reduce redundant visual information and ensure semantic coherence, resulting in a 72.6% reduction in token consumption compared to standard ICoT methods.
Entities (5)
Relation Signals (4)
DaP-ICoT → incorporates → Dynamic Visual Thought Integration
confidence 100% · DaP-ICoT, which incorporates two key components: (1) Dynamic Visual Thought Integration
DaP-ICoT → incorporates → Precise Visual Thought Guidance
confidence 100% · DaP-ICoT... (2) Precise Visual Thought Guidance ensures visual semantically coherent
DaP-ICoT → reduces → Token Consumption
confidence 98% · DaP-ICOT significantly reduces the number of inserted images, leading to a 72.6% decrease in token consumption
Precise Visual Thought Guidance → uses → SAM2
confidence 95% · we apply the Segment Anything Model 2 (SAM2) to the original image I ori for object-level segmentation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thought Positioning, which statically inserts visual information at fixed steps, resulting in inefficient and inflexible reasoning; and (2) Broken Visual Thought Representation, which involves discontinuous and semantically incoherent visual tokens. To address these limitations, we introduce Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DaP-ICoT), which incorporates two key components: (1) Dynamic Visual Thought Integration adaptively introduces visual inputs based on reasoning needs, reducing redundancy and improving efficiency. (2) Precise Visual Thought Guidance ensures visual semantically coherent and contextually aligned representations. Experiments across multiple benchmarks and models demonstrate that DaP-ICoT achieves state-of-the-art performance. In addition, DaP-ICoT significantly reduces the number of inserted images, leading to a 72.6% decrease in token consumption, enabling more efficient ICoT reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2603.21754v1
- Canonical: https://arxiv.org/abs/2603.21754v1
Trouble viewing inline? Open PDF directly →
Full Text
43,437 characters extracted from source content.
Expand or collapse full text
Let’s Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts Xu Liu 1,2,3∗ , Yongheng Zhang 2 * , Qiguang Chen 2 , Yao Li 4 , Sheng Wang 4 , Libo Qin 1,2,3 1 Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen 2 School of Computer Science and Engineering, Central South University 3 Text Computing and Cognitive Intelligence Ministry of Education Engineering Research Center, Guizhou University 4 Shanghai Aviation Electric Co., Ltd, Aviation Industry Corporation of China, Shanghai Abstract Recently, Interleaved-modal Chain-of-Thought (ICoT) rea- soning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing at- tention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thought Positioning, which statically inserts visual information at fixed steps, resulting in inefficient and inflexible reasoning; and (2) Broken Visual Thought Representation, which involves discontinuous and semanti- cally incoherent visual tokens. To address these limitations, we introduce Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DAP-ICOT), which incorporates two key components: (1) Dynamic Visual Thought Integration adaptively introduces visual inputs based on reasoning needs, reducing redundancy and improving effi- ciency. (2) Precise Visual Thought Guidance ensures visual semantically coherent and contextually aligned representa- tions. Experiments across multiple benchmarks and models demonstrate that DAP-ICOT achieves state-of-the-art per- formance. In addition, DAP-ICOT significantly reduces the number of inserted images, leading to a 72.6% decrease in token consumption, enabling more efficient ICoT reasoning. Code — https://github.com/67L1/DaP-ICoT Introduction In recent years, Multimodal Large Language Models (MLLMs) (Achiam et al. 2023; Team et al. 2024; Qin et al. 2024), and Multimodal Chain-of-Thought (MCoT) (Zhang et al. 2023; Chen et al. 2024), have significantly advanced the reasoning capabilities of MLLMs across the complex real-world tasks (Wei et al. 2025; Chen et al. 2025a). How- ever, current MCoT mainly follows the traditional paradigm: cross-modal input, but reasoning output in the text modality, which limits the human’s ability to exploit modality com- plementarity, reducing its reasoning performance (Fei et al. 2024; Menon, Zemel, and Vondrick 2024; Wu et al. 2025). To address this limitation, researchers propose Interleaved- Modal Chain-of-Thought (ICoT) reasoning (Hu et al. 2024; Cheng et al. 2025b; Gao et al. 2025). ICoT allows the * Equal Contribution. Copyright © 2026, Association for the Advancement of Artificial Intelligence (w.aaai.org). All rights reserved. (b) Our DaP-ICoTReasoning 1. DynamicVisual Thought Integration! 2. PreciseVisual Thought Guidance! Step 1: First, ... a road sign... (continue Step 2) ... different countries. Step 2: Second, ... the national flag Precise Segmented Visual Thought Visual Thought 1 (Following Low Confidence) Step 1: First, analysis ... , is a road sign ... text written on the road sign. Step 2: Second, ... the national flag ... Step N: ... seems to be a convention center. (a) Current ICoTReasoning Options: (A) A convention center ... ... (D) A global office building. Question:Whatistheprobablefunctionofthebuilding withcountryflagsandatallbuildingnearby? Multi-Modal Query 1. StaticVisual Thought Positioning! 2. BrokenVisual Thought Representation! Visual Thought 1 (Following Each Step) Massive Image Token Cost Redundant Visual Thought Inefficient!! Limited Image Token Cost Feasible Visual Thought Efficient!! Step N: ... this is A global office building. Figure 1: (a) Current ICoT: While supporting multimodal inputs and outputs, it suffers from Static Visual Thought In- tegration, which requires the insertion of visual information after each step, and Broken Visual Thought Representation, in which the inserted visual tokens are lack coherence, re- sulting in inefficient reasoning. (b) Our DAP-ICOT: It pro- vides Dynamic Visual Thought Integration and Precise Vi- sual Thought Guidance, enabling efficient reasoning. model to receive multimodal input and simultaneously per- form multimodal reasoning output, enabling visual thoughts to effectively convey the image information during reason- ing (Meng et al. 2023; Shao et al. 2024; Zhang et al. 2024a; Li et al. 2025; Wang et al. 2025b; Chen et al. 2025b). Specifically, a growing body of research has focused on advancing ICoT reasoning to better convey visual thoughts. Cheng et al. (2025b) employ segmentation tools to extract key areas and generate images that form visual thoughts, thereby significantly enhancing the capabilities of MLLMs for more advanced, human-like multimodal reasoning and visual operations across diverse tasks. Hu et al. (2024) arXiv:2603.21754v1 [cs.CV] 23 Mar 2026 propose Sketchpad, a novel and effective sketching-based framework that enhances MLLM reasoning by enabling vi- sual thought expression through intuitive figure drawing on a dynamic visual sketchpad interface. Zhou et al. (2024) propose Image-of-Thought visual prompting, a structured method for step-by-step visual rationale extraction that fur- ther enhances multimodal reasoning by effectively convey- ing internal visual thoughts in MLLMs through designed vi- sual cues. Zhang et al. (2025a) propose Video-Text Inter- leaved CoT reasoning, a cognitively aligned temporal video reasoning paradigm that interleaves visual and textual in- formation to enhance video understanding and reasoning in MLLMs. Gao et al. (2025) further employ an attention- driven selection module to statically select some image to- kens that convey visual thoughts in each step. While achieving promising performance, as shown in Fig- ure 1 (a), current ICoT methods for visual thought con- veyance face two key challenges that impede their potential: (1) Static Visual Thought Positioning: Existing ICoT rea- soning approaches insert visual thoughts statically after each textual rationale generation step, resulting in rigid thinking patterns, redunant visual information, and a sig- nificant increase in computational overhead. (2) Broken Visual Thought Representation: Existing ICoT methods select discontinuous image tokens as visual thoughts, which are broken and lack coherence. Such broken visual thought impairs understanding and in- creases the risk of overlooking critical information. Motivated by these challenges, we introduce an Interleaved-modal Chain-of-Thought reasoning framework with Dynamic and Precise Visual Thoughts (DAP-ICOT). Specifically, to address the first challenge, as illustrated in Figure 1 (b), we propose Dynamic Visual Thought Integration, which dynamically and adaptively leverages visual information in response to evolving reasoning needs. In contrast to statically processing all available visual data, DAP-ICOT selectively invokes visual modalities based on contextual demands, ensuring that the integration of visual cues is both timely and relevant. This dynamic mechanism enhances multimodal reasoning by reducing redundant computation and focusing on salient visual cues. To address the second challenge, we introduce Precise Visual Thought Guidance, which emphasizes the integration of semanti- cally coherent and contextually relevant visual information. Instead of relying on broken cues, it employs precise visual representations that encapsulate complete semantics, ensuring tight alignment between visual input and the rea- soning trajectory. This precision-oriented design safeguards conceptual consistency and enhances interpretability and reasoning accuracy. Such two modules together minimize unnecessary visual information and better capture key visual thoughts, enabling more efficient ICoT reasoning. Experiments conducted on multiple widely used bench- marks and MLLMs demonstrate that DAP-ICOT consis- tently outperforms baselines. Further in-depth analysis re- veals that DAP-ICOT significantly reduces the number of image insertions, resulting in a 72.6% reduction in token consumption compared to the current ICoT approach. These results highlight its effective reasoning, balancing efficiency and effectiveness in multimodal understanding. Our contributions can be summarized as follows: (1) We highlight two drawbacks in the existing ICoT paradigm: Static Visual Thought Positioning that induces inefficient reasoning, and Broken Visual Thought Repre- sentation that undermines coherent visual thoughts. (2) We introduce the Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DAP-ICOT) to address drawbacks in ICoT, which in- corporates modules: Dynamic Visual Thought Integra- tion and Precise Visual Thought Guidance. (3) Extensive experiments demonstrate that DAP-ICOT achieves the state-of-the-art performance. In addition, DAP-ICOT effectively reduces the number of inserted images and achieves a 72.6% reduction in token con- sumption, enabling more efficient reasoning. DAP-ICOT Reasoning In this work, we introduce the Interleaved-modal Chain- of-Thought reasoning with Dynamic and Precise Visual Thoughts (DAP-ICOT) to address the inefficiencies in pre- vious ICoT approaches. Specifically, as shown in Figure 2, DAP-ICOT comprises Dynamic Visual Thought Integra- tion (§) and Precise Visual Thought Guidance (§). Dynamic Visual Thought Integration To fill the gap of the static visual thought positioning in ex- isting ICoT reasoning methods, we introduce Dynamic Vi- sual Thought Integration (DVTI), as shown in Figure 2 (a), a confidence-aware strategy that dynamically decides whether visual thought should be integrated during reason- ing, based on the confidence of MLLMs. 2.1.1 Confidence Estimation via Logit Margin Analysis Formally, given a generated textual rationale T t at reason- ing step t, we estimate the model’s confidence by analyzing token-level logit differentials during generation. This allows us to internally assess certainty without requiring external calibration. At each decoding position i, we define the local confidence δ i as the difference between the highest and the second-highest predicted logits: δ i = ℓ i,w (1) − ℓ i,w (2) , ∀i∈ 1,...,|T t |,(1) where ℓ i,w (1) and ℓ i,w (2) denote the top-1 and top-2 logits predicted at position i, respectively. This margin reflects the model’s decisiveness at each token position, larger values indicate stronger confidence in the token selection. To obtain a more robust and reliable overall confidence score for the entire rationale T t , we aggregate the local mar- gins by computing their mean value across positions: C t = 1 |T t | |T t | X i=1 δ i = 1 |T t | |T t | X i=1 ℓ i,w (1) − ℓ i,w (2) ,(2) where C t represents the average confidence across decoding steps, with each δ i quantifying the model’s certainty at po- sition i based on the logit gap between its top two predicted tokens, reflecting how decisively the selection of MLLMs. Visual Thought Insertion(§2.1.2) InterleavedReasoning(§2.2.2) Confidence Estimation(§2.1.1) Option: ... (D) A global office building. Question: What is the probable function of the building with country flags and a tall building nearby? Low Step 2: Second, ...the national flag (a) DynamicVisual Thought Integration(§2.1) (b) PreciseVisual Thought Guidance(§2.2) Token Logit Distribution Top-1 LogitTop-2 Logit-= Confidence 훴 Logit Margin Analysis Segment SAM2 Candidate Object Image Divide Need Visual Thought Scoreτ=0.8 High Don't need Visual Thought Scoreτ=0.4 ... flags of many countries. Object-Level Selection(§2.2.1) Scoreτ=0.8 High Don't need Visual Thought Step1:First,analysis...,thereisa roadsign. Answer: ... A global office building. Figure 2: An overview of Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DAP- ICOT), including Dynamic Visual Thought Integration (§), and Precise Visual Thought Guidance (§). 2.1.2 Visual Thought Insertion Guided by Confidence Based on the computed confidence score C t , we apply a thresholding mechanism to determine whether visual input is necessary for the next step. The determinant is as follows: I t+1 = I vision , if C t < τ ∅,otherwise (3) where I t+1 denotes the visual thought input used at the (t+1)-th reasoning step. If the confidence C t falls below the predefined threshold τ , the model is prompted to incorpo- rate visual context, retrieved via the Precise Visual Thought Guidance (§). Otherwise, no visual input is provided (∅), al- lowing for proceeding based on textual reasoning. DVTI is capable of reducing redundant information through dynamic visual information integration, thereby de- creasing token processing and improving efficiency. Precise Visual Thought Guidance To fill the gap of broken and semantically incoherent visual inputs in existing ICoT methods, we propose Precise Vi- sual Thought Guidance (PVTG), as shown in Figure 2 (b), which facilitates fine-grained object-level visual selection through cross-modal semantic relevance matching. 2.2.1 Object-Level Selection via Cross-Modal Relevance In the first, we apply the Segment Anything Model 2 (SAM2) (Ravi et al. 2024) to the original image I ori for object-level segmentation. This identifies multiple distinct object regions, each corresponding to semantically mean- ingful visual content. We then extract sub-images of ob- jects, forming a pool of candidate object-centric visual in- puts. Unlike patch-level image tokens, these object-centric sub-images preserve semantic information. Specifically: O = O 1 ,O 2 ,...,O N ,(4) where each O i denotes a complete and semantically coher- ent object sub-image extracted from the original image I ori . When Dynamic Visual Thought Integration (§) identifies the need for visual input at step t, we compute the semantic relevance between the textual rationale T t and each candi- date object image O i ∈O via cross-modal attention using a similarity function f attn (·,·): s i = f attn (T t ,O i ),(5) where s i denotes the attention-based similarity between ra- tionale T t and object image O i (Gao et al. 2025). We then select the most relevant object image with the highest score: ˆ O = argmax O i ∈O s i .(6) where ˆ O denotes the object image that exhibits the strongest semantic alignment with the current textual rationale. 2.2.2 Interleaved Reasoning with Aligned Visual Inputs Instead of treating ˆ O as a standalone input, we interleave it into the reasoning process by embedding it within the tex- tual rationaleR t , forming a multimodal reasoning sequence: R t→v =R t ⊕ ˆ O,(7) where ⊕ denotes the operation of embedding the selected object image ˆ O into the textual reasoning sequence R t , specifically by inserting the image token after the previously generated textual rationale. The model then continues the next reasoning step based on this interleaved multimodal input: R t+1 = argmax R P(R|Q, R t→v ,P t→v ),(8) where Q denotes the question, and P t→v represents the prompt constructed for interleaved reasoning. Through targeted selection and structured interleaving, PVTG provides precise visual information that preserves se- mantic coherence, reduces noise, and enhances the effective- ness and interpretability of the multimodal reasoning. ModelsMethods M 3 CoTScienceQAMME 0-Shot Acc.↑ 1-Shot Acc.↑ 0-Shot Acc.↑ 1-Shot Acc.↑ 0-Shot Score↑ 1-Shot Score↑ DIRECT22.523.843.143.4724.2942.9 MMCOTTMLR 202426.028.046.248.9435.8661.2 DDCOTNeurIPS 202329.830.347.448.4725.9953.8 SCAFFOLDACL 202531.031.248.650.7388.1634.5 CCOTCVPR 202425.126.342.844.3366.1487.9 ICOTCVPR 202526.132.144.545.3794.8928.9 Chameleon-7B (Team 2024) DAP-ICOT41.041.957.162.9832.31013.0 DIRECT23.225.721.923.5975.31079.8 MMCOTTMLR 202429.630.442.443.11140.21277.4 DDCOTNeurIPS 202325.426.529.331.0736.7991.8 SCAFFOLDACL 202526.628.337.739.81170.31264.3 CCOTCVPR 202434.235.533.835.71294.31349.5 ICOTCVPR 202534.635.041.746.71331.61421.6 LLaVA-V1.5-7B (Liu et al. 2023) DAP-ICOT36.337.650.451.11386.71526.9 DIRECT24.625.529.734.0995.41118.5 MMCOTTMLR 202432.132.956.158.31078.11224.3 DDCOTNeurIPS 202332.933.839.341.8800.91034.1 SCAFFOLDACL 202531.933.241.743.81231.21389.9 CCOTCVPR 202430.132.045.046.31249.31442.3 ICOTCVPR 202537.037.954.654.81405.41523.8 LLaVA-V1.5-13B (Liu et al. 2023) DAP-ICOT39.441.860.362.71556.31726.3 DIRECT14.424.564.365.2641.5741.3 MMCOTTMLR 202414.922.465.667.31102.81304.7 DDCOTNeurIPS 202337.939.365.868.4800.6967.0 SCAFFOLDACL 202540.343.666.769.41344.81536.2 CCOTCVPR 202420.237.764.266.5761.6867.5 ICOTCVPR 202535.837.360.467.0941.91453.9 Qwen2-VL-2B (Wang et al. 2024) DAP-ICOT47.351.068.473.61378.91862.4 DIRECT33.035.070.971.21599.31641.3 MMCOTTMLR 202444.447.570.873.81602.51874.3 DDCOTNeurIPS 202343.945.362.865.31752.41826.6 SCAFFOLDACL 202549.953.674.475.01668.21822.2 CCOTCVPR 202448.753.072.774.81866.31941.9 ICOTCVPR 202538.044.854.267.01587.31709.3 Qwen2-VL-7B (Wang et al. 2024) DAP-ICOT57.258.775.978.52012.22076.0 Table 1: The main experimental results. Bold indicates the best performance. For the M 3 CoT and ScienceQA, Acc. is used as the evaluation metric, while for the MME, the sum of the Perception and Cognition scores is used as the evaluation metric. Experiments and Analysis Experiments Setting Following Gao et al. (2025), in addition to the direct query approach, we also adopt the following methods as baselines: • MMCoT (Zhang et al. 2024b) separates rationale genera- tion and answer inference by incorporating both text and image modalities to improve reasoning performance. • DDCoT (Zheng et al. 2023) divides reasoning and recog- nition by combining LLM reasoning with visual recog- nition through negative-space prompting, enabling effec- tive and explainable multimodal CoT reasoning. • SCAFFOLD (Lei et al. 2025) prompts overlays a dot matrix on images as visual anchors and introduces coordinate-based textual references to effectively en- hance vision-language coordination in MLLMs. • CCoT (Mitra et al. 2024) first generates scene graphs with LMMs and then uses them in prompts to enhance compositional reasoning without annotations. • ICoT (Gao et al. 2025) generates paired visual and textual reasoning steps by inserting image regions via Attention-driven Selection to enhance reasoning. All methods are reproduced once using their official open- source implementations and evaluated under both 0-shot and 1-shot settings. To evaluate the effectiveness of DAP- ICOT, we conduct experiments on five MLLMs, including Chameleon-7B (Team 2024), LLaVA-V1.5-(7B, 13B) (Liu et al. 2023), and Qwen2-VL-(2B, 7B) (Wang et al. 2024). We use the default top-p and temperature settings provided by each MLLM. In DAP-ICOT, the confidence threshold τ is tuned on M 3 CoT validation set by searching within [0, 1] and selecting the value that yields the best performance. Main Results The experimental results are summarized in Table 1. Based on these results, we can observe the following: (1) DAP-ICOT achieves consistently superior perfor- mance. DAP-ICOT consistently achieves the highest reasoning accuracy across all settings. In particular, it significantly outperforms all baseline methods on the M 3 CoT benchmarks under both 0-shot and 1-shot set- tings. For example, on M 3 CoT task with the Chameleon- 7B, DAP-ICOT achieves a remarkable 0-shot accuracy of 41.0%, which is substantially higher than the second- best method, Scaffold, with a score of 31.0%. (2) DAP-ICOT demonstrates strong versatility across di- verse tasks. In addition to its promising performance on M 3 CoT, DAP-ICOT consistently outperforms all baselines across other challenging reasoning tasks and comprehensive multimodal benchmarks. Specifically, it achieves the highest average scores on ScienceQA and MME in all settings. This demonstrates its strong gen- eralization ability and adaptability to complex reasoning and holistic multimodal understanding tasks. (3) DAP-ICOT is generalizable across MLLMs of dif- ferent architectures and scales. DAP-ICOT consis- tently delivers superior performance across various MLLMs with diverse architectures and scales. From small MLLMs such as Qwen2-VL-2B to large MLLMs like Qwen2-VL-7B and LLaVA-V1.5-13B, DAP-ICOT maintains its leading performance. These results clearly suggest that DAP-ICOT is not only effective for specific models but also adaptable to a wide range of pretraining settings and reasoning capacities, demonstrating scala- bility and model-agnostic robustness. Analysis This section provides a more in-depth analysis of DAP- ICOT to demonstrate its effectiveness and efficiency. 1. Both the DVTI and PVTG modules are vital for ad- dressing key ICoT challenges. We perform thorough ab- lation studies on Qwen2-VL-7B to systematically assess the impact of Dynamic Visual Thought Integration (DVTI) and Precise Visual Thought Guidance (PVTG). The results, as shown in Figure 3, clearly confirm their effectiveness. • Removing the Dynamic Visual Thought Integration (DVTI) module results in a substantial performance degradation—specifically, a 14.4% drop on the M 3 CoT benchmark and a 20.8% drop on the ScienceQA bench- mark. These results underscore the critical role of DVTI in dynamically fusing multimodal information. • Removing the Precise Visual Thought Guidance (PVTG) module leads to a significant performance decline, with a 13.8% reduction on the M 3 CoT benchmark and 20.4% on the ScienceQA benchmark. This highlights the im- portance of PVTG in structuring the visual input space by leveraging fine-grained object-level segmentation. These findings demonstrate that both DVTI and PVTG are essential for efficient and precise ICoT reasoning. DaP-ICoT-w/o PVTG-w/o DVTI M3CoT 57.243.442.8 ScienceQA 75.955.555.1 40 48 56 64 72 Performance (%) Figure 3: Ablation Study on Qwen2-VL-7B: “w/o PVTG” indicates removal of Precise Visual Thought Guidance for Visual Cues, and “w/o DVTI” indicates removal of Dynamic Visual Thought Integration for Adaptive Reasoning 314 1,146 1,294 934 1,426 1,080 100300500700900110013001500 DaP-ICoT ICoT CCoT SCAF. DDCoT MCoT Ave r a g e To ke n C o n s u m p t i o n 72.6% Figure 4: A comparison of the total token consumption be- tween DAP-ICOT and baseline methods on the M 3 CoT benchmark using the Qwen2-VL-7B. DAP-ICOT achieves a 72.6% reduction in token consumption compared to ICoT. 2. DAP-ICOT significantly reduces the token consump- tion of MLLMs. To further verify the lightweight design of DAP-ICOT, we compare its total token consumption with several baseline methods on the M 3 CoT using the Qwen2- VL-7B. This experiment aims to assess whether DAP-ICOT can effectively reduce token usage while maintaining strong reasoning performance. As shown in Figure 4, DAP-ICOT achieves a significant reduction in token consumption, using an average of 314 tokens, which is a 72.6% decrease com- pared to ICoT that requires 1,146 tokens. Furthermore, most baselines consume considerably more tokens, with CCoT requiring 1,294 tokens and DDCoT reaching the highest at 1,426 tokens. Even relatively more efficient methods, such as MMCoT and SCAFFOLD, still require 1,080 and 934 to- kens, respectively. These results indicate that the Dynamic Visual Thought Integration module in DAP-ICOT effec- tively reduces the overhead associated with image embed- ding, thereby achieving significantly lower token consump- tion compared to existing baseline reasoning methods. 3.578 2.323 2.013 0.934 0.992 1.684 0 1 2 3 4 M3CoTScienceQAMME 240 156 135 16 21 41 0 50 100 150 200 250 300 Frequency Token Image Insertion Frequency Image Token Consumption ICoTDaP-ICoT Figure 5: A comparison of image insertion frequency and the number of inserted image tokens between DAP-ICOT and ICoT on the Qwen2-VL-7B model. M ! CoT Science QA MME The probability of an increase in confidence level.The valueof the increase. 80.5% 85.0% 76.7% 46.4% 41.1% 51.8% 0.15 0.17 0.13 0.02 -0.07 -0.08 ICoT DaP-ICoT Figure 6: The proportion of samples where the confidence level increases after image insertion, as well as the average confidence value improvement, for DAP-ICOT and ICoT. TheBlue Color denotes the baseline method ICoT, while the Green Color represents our DAP-ICOT. 3. DAP-ICOT reduces resource consumption from image insertions. To clarify the efficient feature of DAP-ICOT, we conduct a comparative analysis with ICoT on image us- age during reasoning, based on two key metrics: (1) The av- erage number of image insertions per sample and (2) The av- erage number of image tokens consumed after insertion. As shown in Figure 5, DAP-ICOT demonstrates a significant reduction in both image insertion frequency and token con- sumption. On average, DAP-ICOT inserts only 1.2 images per sample, whereas ICoT inserts an average of 2.6 images. Moreover, in terms of token usage, DAP-ICOT consumes merely 26 image tokens on average, which is substantially lower than ICoT. These results demonstrate the efficiency of DAP-ICOT, which achieves superior performance with minimal visual input and token consumption. This is due to its modules: Dynamic Visual Thought Integration, which adaptively reduces unnecessary image insertions, and Pre- cise Visual Thought Guidance, which effectively minimizes token usage by selecting compact object-level visual inputs. 0.10.20.30.40.50.60.70.80.91 M3CoT 56.857.256.756.157.156.053.052.953.552.3 52 53 54 55 56 57 58 Performance (%) Qwen2-VL-7B |0-Shot: 57.2 1-Shot: 58.7 Figure 7: Performance comparison across the Confidence Threshold τ range of (0, 1]: Optimal model performance achieved at τ = 0.2, with 57.2% Accuracy in the 0-shot setting and 58.7% in the 1-shot setting, effectively balanc- ing visual integration and textual reasoning. 4. DAP-ICOT effectively enhances confidence during the reasoning process. To further investigate why DAP- ICOT achieves superior performance, we conduct an in- depth analysis of the model’s reasoning confidence varia- tions after image insertion. Specifically, we compare DAP- ICOT and ICoT in terms of two key metrics: (1) The propor- tion of samples showing increased confidence after apply- ing Visual Thought, and (2) The average magnitude of the confidence improvement. As shown in Figure 6, DAP-ICOT consistently demonstrates a significantly higher probability of confidence improvement across all three benchmarks. On average, DAP-ICOT leads to an increase in confidence for 80.7% of the samples, while ICoT only improves confidence in 46.4% of cases. Furthermore, regarding the extent of con- fidence enhancement, DAP-ICOT consistently outperforms ICoT across all three benchmarks. These results clearly in- dicate that DAP-ICOT is more effective in leveraging vi- sual information to enhance confidence during reasoning, thereby contributing to its superior performance. 5. The search strategy for the confidence threshold τ . To gain a clearer understanding of the threshold selection process for the confidence threshold τ in the Dynamic Vi- sual Thought Integration (DVTI) module, we conduct a sys- tematic threshold search experiment using Qwen2-VL-7B on the M 3 CoT benchmark. Specifically, we vary the confidence threshold τ within the range of (0, 1] with an interval of 0.1 and evaluate the model’s performance under each setting. The experimental results are presented in Figure 7, illustrat- ing the relationship between the confidence threshold and overall performance. Reveal the impact of varying the con- fidence threshold τ on model performance. As τ increases, the performance exhibits a clear trend. The optimal perfor- mance of 57.2% is achieved at a threshold of 0.2. This trend suggests that a moderate threshold effectively balances the model’s reliance on visual information and textual reason- ing, enabling optimal reasoning performance. In contrast, overly conservative or overly aggressive visual interactions negatively affect the model’s ability to reason effectively. (a) CurrentICoTReasoning Apersonholdingtwoumbrellasovertheirhead... ...Theman'sarmsareraisedabovehisheadashe holdsontotwoumbrellas. So, Answer: (D) Waiting for someone SAM2 Segment O: (A) Selling umbrellas. (D) Waiting for someone. Q: What is the man in the walkway doing? Image Pool ... Static Visual Thought Step 1 BrokenVisual Thought Representation Static Visual Thought Step 2 BrokenVisual Thought Representation (b) OurDaP-ICoTReasoning Themaninthewalkwayisholdingtwoumbrellas. Thereisnoindicationthattheyare any... So, Answer: (A) Selling umbrellas DynamicVisual Thought Step 1 Precise Representation [DVTI] High Confidence LowConfidence [DVTI] [PVTG] Matched Image Figure 8: The case study. Figure (a) illustrates the reasoning process of the current ICoT, where after each reasoning step, it is necessary to insert image tokens that are most similar to the previous text. However, these tokens are broken image tokens, which ultimately leads to the incorrect answer (D). Figure (b) illustrates our DAP-ICOT. In contrast, DAP-ICOT only inserts images when the text confidence is low. Furthermore, the inserted images are precise and complete, having been segmented using the SAM2 model. After efficient interleaved visual-textual reasoning, the final correct answer (A) is obtained. 6. Qualitative Analysis. To better understand the perfor- mance of DAP-ICOT, we present a real-world example. As shown in Figure 8 (a), ICoT inserts incomplete image to- kens during reasoning, leading to an incorrect result (D). This example highlights the limitations of the current ICoT approach and its impact on reasoning accuracy. In contrast, Figure 8 (b) demonstrates DAP-ICOT, which inserts only complete, context-relevant images when text confidence is low. Compared to ICoT, DAP-ICOT inserts fewer images and employs a more selective approach. After ICoT reason- ing, the system correctly arrives at answer (A), which is ver- ified through the dynamic visual thought integration mecha- nism. This example illustrates the efficiency of the selective visual input mechanism in DAP-ICOT reasoning. Related Work In recent years, Multimodal Large Language Models (MLLMs) have witnessed rapid advancements (Liang et al. 2024; Qiu et al. 2025; Qiang et al. 2025), and the emer- gence of Multimodal Chain-of-Thought (MCoT) reasoning has further enhanced their performance (Wang et al. 2025b). Specifically, Zhang et al. (2024b) proposed Multimodal- CoT, a two-stage reasoning framework integrating text and image modalities. Chen et al. (2024) introduce M 3 CoT, a benchmark for multi-modal, multi-step, and multi-domain chain-of-thought reasoning, addressing key limitations of existing MCoT benchamrk. However, MCoT methods largely follow the conventional paradigm of taking cross- modal inputs while generating reasoning outputs only in the text modality. This limits the effective use of modality com- plementarity and diminishes reasoning performance (Wang et al. 2025a; Lin et al. 2025; Zhang et al. 2025b). To address this limitation, researchers have explored Interleaved-Modal Chain-of-Thought (ICoT) (Gao et al. 2025; Wu et al. 2025), which enhances the reasoning abil- ities of MLLMs through cross-modal Integrations (Hu et al. 2024; Cheng et al. 2025a). For example, Zhou et al. (2024) propose Image-of-Thought prompting to guide MLLMs in step-by-step visual extraction. Hu et al. (2024) introduce a sketching framework, enabling models to perform human- like drawing to reasoning. In addition, Cheng et al. (2025b) propose the CoMT for evaluating multimodal reasoning with visual and textual operations. Gao et al. (2025) intro- duce ICoT reasoning with an Attention-driven Selection for generating interleaved visual-textual reasoning. Compared to previous ICoT reasoning approaches, DAP- ICOT introduces both Dynamic Visual Thought Integration and Precise Visual Thought Guidance, enabling not only more efficient reasoning but also adaptive and context-aware visual clues for ICoT Reasoning. Conclusion In this work, we propose Interleaved-modal Chain-of- Thought reasoning with Dynamic and Precise Visual Thoughts (DAP-ICOT), achieving efficient reasoning. Specifically, DAP-ICOT adaptively integrates informative and context-relevant visual information and provides se- mantically coherent visual inputs. Extensive evaluations on multiple benchmarks and advanced MLLMs demonstrate that DAP-ICOT achieves superior performance. In addition, DAP-ICOT is capable of effectively reducing token con- sumption and the frequency of visual insertions, highlight- ing its strong potential in efficient multimodal reasoning. Acknowledgments This work was supported by the National Natural Sci- ence Foundation of China (NSFC) via grants 92570120 and 62306342. This work was supported by the Scien- tific Research Fund of Hunan Provincial Education Depart- ment (24B0001). This work was sponsored by the Excellent Young Scientists Fund in Hunan Province (2024J4070), the Science and Technology Innovation Program of Hu- nan Province under Grant 2024RC3024, and CCF-Zhipu Large Model Innovation Fund (NO.CCF-Zhipu202406). This study was also funded by the Open Project of the Text Computing and Cognitive Intelligence Ministry of Educa- tion Engineering Research Center (No. TCCI250101). Libo Qin is the corresponding author. References Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; Avila, R.; Babuschkin, I.; Balaji, S.; Balcom, V.; Baltescu, P.; Bao, H.; Bavarian, M.; Belgum, J.; Bello, I.; Berdine, J.; Bernadett-Shapiro, G.; Berner, C.; Bogdonoff, L.; Boiko, O.; Boyd, M.; Brakman, A.-L.; Brockman, G.; Brooks, T.; Brundage, M.; Button, K.; Cai, T.; Campbell, R.; Cann, A.; Carey, B.; Carlson, C.; Carmichael, R.; Chan, B.; Chang, C.; Chantzis, F.; Chen, D.; Chen, S.; Chen, R.; Chen, J.; Chen, M.; Chess, B.; Cho, C.; Chu, C.; Chung, H. W.; Cummings, D.; Currier, J.; Dai, Y.; Decareaux, C.; Degry, T.; Deutsch, N.; Deville, D.; Dhar, A.; Dohan, D.; Dowling, S.; Dunning, S.; Ecoffet, A.; Eleti, A.; Eloundou, T.; Farhi, D.; Fedus, L.; Felix, N.; Fishman, S. P.; Forte, J.; Fulford, I.; Gao, L.; Georges, E.; Gibson, C.; Goel, V.; Gogineni, T.; Goh, G.; Gontijo-Lopes, R.; Gordon, J.; Grafstein, M.; Gray, S.; Greene, R.; Gross, J.; Gu, S. S.; Guo, Y.; Hallacy, C.; Han, J.; Harris, J.; He, Y.; Heaton, M.; Heidecke, J.; Hesse, C.; Hickey, A.; Hickey, W.; Hoeschele, P.; Houghton, B.; Hsu, K.; Hu, S.; Hu, X.; Huizinga, J.; ; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Chen, Q.; Qin, L.; Liu, J.; Peng, D.; Guan, J.; Wang, P.; Hu, M.; Zhou, Y.; Gao, T.; and Che, W. 2025a. Towards rea- soning era: A survey of long chain-of-thought for reasoning large language models. arXiv preprint arXiv:2503.09567. Chen, Q.; Qin, L.; Zhang, J.; Chen, Z.; Xu, X.; and Che, W. 2024. M 3 CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), 8199–8221. Bangkok, Thailand: Association for Computational Linguistics. Chen, Q.; Yang, M.; Qin, L.; Liu, J.; Yan, Z.; Guan, J.; Peng, D.; Ji, Y.; Li, H.; Hu, M.; et al. 2025b. AI4Research: A Sur- vey of Artificial Intelligence for Scientific Research. arXiv preprint arXiv:2507.01903. Cheng, Z.; Chen, Q.; Xu, X.; Wang, J.; Wang, W.; Fei, H.; Wang, Y.; Wang, A. J.; Chen, Z.; Che, W.; et al. 2025a.Visual Thoughts: A Unified Perspective of Un- derstanding Multimodal Chain-of-Thought. arXiv preprint arXiv:2505.15510. Cheng, Z.; Chen, Q.; Zhang, J.; Fei, H.; Feng, X.; Che, W.; Li, M.; and Qin, L. 2025b. Comt: A novel benchmark for chain of multi-modal thought on large vision-language mod- els. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 23678–23686. Fei, H.; Wu, S.; Ji, W.; Zhang, H.; Zhang, M.; Lee, M.-L.; and Hsu, W. 2024. Video-of-thought: Step-by-step video reasoning from perception to cognition.arXiv preprint arXiv:2501.03230. Gao, J.; Li, Y.; Cao, Z.; and Li, W. 2025. Interleaved-modal chain-of-thought. In Proceedings of the Computer Vision and Pattern Recognition Conference, 19520–19529. Hu, Y.; Shi, W.; Fu, X.; Roth, D.; Ostendorf, M.; Zettle- moyer, L.; Smith, N. A.; and Krishna, R. 2024.Visual Sketchpad: Sketching as a Visual Chain of Thought for Mul- timodal Language Models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Lei, X.; Yang, Z.; Chen, X.; Li, P.; and Liu, Y. 2025. Scaf- folding Coordinates to Promote Vision-Language Coordina- tion in Large Multi-Modal Models. In Rambow, O.; Wan- ner, L.; Apidianaki, M.; Al-Khalifa, H.; Eugenio, B. D.; and Schockaert, S., eds., Proceedings of the 31st International Conference on Computational Linguistics, 2886–2903. Abu Dhabi, UAE: Association for Computational Linguistics. Li, C.; Wu, W.; Zhang, H.; Xia, Y.; Mao, S.; Dong, L.; Vuli ́ c, I.; and Wei, F. 2025.Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.arXiv preprint arXiv:2501.07542. Liang, Z.; Xu, Y.; Hong, Y.; Shang, P.; Wang, Q.; Fu, Q.; and Liu, K. 2024. A Survey of Multimodel Large Language Models. In Proceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineer- ing, 405–409. Lin, J.; Zeng, X.; Zhu, J.; Wang, S.; Shun, J.; Wu, J.; and Zhou, D. 2025. Plan and Budget: Effective and Efficient Test-Time Scaling on Large Language Model Reasoning. arXiv preprint arXiv:2505.16122. Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual in- struction tuning. Advances in neural information processing systems, 36: 34892–34916. Meng, F.; Yang, H.; Wang, Y.; and Zhang, M. 2023. Chain of images for intuitively reasoning. arXiv preprint arXiv:2311.09241. Menon, S.; Zemel, R.; and Vondrick, C. 2024. Whiteboard- of-Thought: Thinking Step-by-Step Across Modalities. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 20016–20031. Mitra, C.; Huang, B.; Darrell, T.; and Herzig, R. 2024. Com- positional Chain-of-Thought Prompting for Large Multi- modal Models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14420–14431. IEEE Computer Society. Qiang, C.; Wei, Z.; Han, X.; Wang, Z.; Li, S.; Lan, X.; Jiao, J.; and Han, Z. 2025. VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence. In Proceed- ings of the 33rd ACM International Conference on Multime- dia, 12698–12705. Qin, L.; Chen, Q.; Feng, X.; Wu, Y.; Zhang, Y.; Li, Y.; Li, M.; Che, W.; and Yu, P. S. 2024. Large language models meet nlp: A survey. arXiv preprint arXiv:2405.12819. Qiu, T.; Gao, J.; Li, J.; Leong, H.; Huang, X.; Wang, X.; Zhang, X.; Xu, K.; and Zhang, L. 2025. Intentvcnet: Bridg- ing spatio-temporal gaps for intention-oriented controllable video captioning. In Proceedings of the 33rd ACM Interna- tional Conference on Multimedia, 13822–13829. Ravi, N.; Gabeur, V.; Hu, Y.-T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; R ̈ adle, R.; Rolland, C.; Gustafson, L.; Mintun, E.; Pan, J.; Alwala, K. V.; Carion, N.; Wu, C.-Y.; Girshick, R.; Doll ́ ar, P.; and Feichtenhofer, C. 2024. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714. Shao, H.; Qian, S.; Xiao, H.; Song, G.; Zong, Z.; Wang, L.; Liu, Y.; and Li, H. 2024. Visual cot: Advancing multi-modal language models with a comprehensive dataset and bench- mark for chain-of-thought reasoning. Advances in Neural Information Processing Systems, 37: 8612–8642. Team, C. 2024.Chameleon: Mixed-modal early-fusion foundation models. arXiv preprint arXiv:2405.09818. Team, G.; Georgiev, P.; Lei, V. I.; Burnell, R.; Bai, L.; Gu- lati, A.; Tanzer, G.; Vincent, D.; Pan, Z.; Wang, S.; Mar- iooryad, S.; Ding, Y.; Geng, X.; Alcober, F.; Frostig, R.; Omernick, M.; Walker, L.; Paduraru, C.; Sorokin, C.; Tac- chetti, A.; Gaffney, C.; Daruki, S.; Sercinoglu, O.; Gleicher, Z.; Love, J.; Voigtlaender, P.; Jain, R.; Surita, G.; Mohamed, K.; Blevins, R.; Ahn, J.; Zhu, T.; Kawintiranon, K.; Firat, O.; Gu, Y.; Zhang, Y.; Rahtz, M.; Faruqui, M.; Clay, N.; Gilmer, J.; Co-Reyes, J.; Penchev, I.; Zhu, R.; Morioka, N.; Hui, K.; Haridasan, K.; Campos, V.; Mahdieh, M.; Guo, M.; Has- san, S.; Kilgour, K.; Vezer, A.; Cheng, H.-T.; de Liedekerke, R.; Goyal, S.; Barham, P.; Strouse, D.; Noury, S.; Adler, J.; Sundararajan, M.; Vikram, S.; Lepikhin, D.; Paganini, M.; Garcia, X.; Yang, F.; Valter, D.; Trebacz, M.; Vodrahalli, K.; Asawaroengchai, C.; Ring, R.; et al. 2024. Gemini 1.5: Un- locking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530. Wang, B.; Li, Y.; Zhou, Q.; Leong, H. Y.; Zhao, T.; Ye, L.; Deng, H.; Luo, D.; and Vasconcelos, N. 2025a. Do Vi- sion Language Models infer human intention without visual perspective-taking? Towards a scalable” One-Image-Probe- All” dataset. Wang, P.; Bai, S.; Tan, S.; Wang, S.; Fan, Z.; Bai, J.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Fan, Y.; Dang, K.; Du, M.; Ren, X.; Men, R.; Liu, D.; Zhou, C.; Zhou, J.; Lin, J.; et al. 2024. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191. Wang, Y.; Wu, S.; Zhang, Y.; Yan, S.; Liu, Z.; Luo, J.; and Fei, H. 2025b. Multimodal chain-of-thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605. Wei, Z.; Qiang, C.; Jiang, B.; Han, X.; Yu, X.; and Han, Z. 2025. ADˆ 2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions. arXiv preprint arXiv:2506.09557. Wu, X.; Liu, J.; Huang, D.; Li, X.; Wang, Y.; Chen, C.; Ma, L.; Cao, X.; and Xue, J. 2025. ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations. arXiv preprint arXiv:2505.14404. Zhang, Y.; Chen, Q.; Zhou, J.; Wang, P.; Si, J.; Wang, J.; Lu, W.; and Qin, L. 2024a. Wrong-of-thought: An integrated reasoning framework with multi-perspective verification and wrong information. In Findings of the Association for Com- putational Linguistics: EMNLP 2024, 6644–6653. Zhang, Y.; Liu, X.; Tao, R.; Chen, Q.; Fei, H.; Che, W.; and Qin, L. 2025a. ViTCoT: Video-Text Interleaved Chain-of- Thought for Boosting Video Understanding in Large Lan- guage Models. arXiv preprint arXiv:2507.09876. Zhang, Y.; Liu, X.; Zhou, R.; Chen, Q.; Fei, H.; Lu, W.; and Qin, L. 2025b. CCHall: A Novel Benchmark for Joint Cross- Lingual and Cross-Modal Hallucinations Detection in Large Language Models. arXiv preprint arXiv:2505.19108. Zhang, Z.; Zhang, A.; Li, M.; Karypis, G.; Smola, A.; et al. 2024b. Multimodal Chain-of-Thought Reasoning in Lan- guage Models.Transactions on Machine Learning Re- search. Zhang, Z.; Zhang, A.; Li, M.; Zhao, H.; Karypis, G.; and Smola, A. 2023. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923. Zheng, G.; Yang, B.; Tang, J.; Zhou, H.-Y.; and Yang, S. 2023. Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36: 5168–5191. Zhou, Q.; Zhou, R.; Hu, Z.; Lu, P.; Gao, S.; and Zhang, Y. 2024. Image-of-thought prompting for visual reason- ing refinement in multimodal large language models. arXiv preprint arXiv:2405.13872.