Paper deep dive
Retrieval Grounding Latent Reasoning for Dense Retrieval
Gang Zhou, Xiongxi Yu, Hu Tian, Yang Wei, Lu Pan, Ke Zeng, Shibiao Xu, Xiaolong Zheng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/17/2026, 5:04:17 AM
Summary
The paper introduces Retrieval Grounding Latent Reasoning (RGLT), a framework for dense retrieval that performs non-autoregressive reasoning in hidden space. RGLT addresses the limitation of existing latent reasoning methods by explicitly connecting intermediate latent transitions with retrieval improvements. It uses an instruction-conditioned latent reasoning trajectory constructed from silent tokens, combined with process-supervised explicit-to-implicit distillation and retrieval-grounded supervision. This approach ensures that each reasoning step contributes to incremental retrieval gains, outperforming strong baselines on reasoning-intensive benchmarks while maintaining efficient inference.
Entities (10)
Relation Signals (7)
RGLT → improves → Dense Retrieval
confidence 95% · We propose Retrieval Grounding Latent Reasoning (RGLT), a latent reasoning framework for dense retrieval...
RGLT → uses → Silent Tokens
confidence 95% · RGLT performs non-autoregressive reasoning in hidden space through an instruction-conditioned latent reasoning trajectory constructed from silent tokens.
RGLT → employs → Process-Supervised Explicit-to-Implicit Distillation
confidence 92% · It combines process-supervised explicit-to-implicit distillation with retrieval-grounded supervision...
RGLT → employs → Retrieval-Grounded Supervision
confidence 92% · It combines process-supervised explicit-to-implicit distillation with retrieval-grounded supervision...
Gang Zhou → affiliatedwith → Beijing University of Posts and Telecommunications
confidence 90% · Gang Zhou 1... 1 School of Artificial Intelligence, Beijing University of Posts and Telecommunications
RGLT → outperforms → strong baselines
confidence 90% · Experiments on reasoning-intensive retrieval benchmarks show that RGLT consistently outperforms strong baselines while preserving efficient embedding inference.
Gang Zhou → affiliatedwith → Meituan
confidence 85% · Gang Zhou 1,*,‡... 2 Meituan LongCat Interaction Team... This work was conducted during Gang Zhou’s internship at Meituan.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reasoning-intensive retrieval requires text representations to capture not only semantic similarity, but also the reasoning needed to determine relevance under a given retrieval instruction. Existing reasoning-enhanced embedding models improve retrieval by incorporating reasoning information into dense representations, yet their supervision is typically dominated by the final retrieval objective. As a result, latent reasoning trajectories may learn shortcut reasoning patterns that preserve retrieval performance without producing meaningful incremental retrieval gains. We propose Retrieval Grounding Latent Reasoning (RGLT), a latent reasoning framework for dense retrieval that explicitly connects intermediate latent transitions with retrieval improvements. RGLT performs non-autoregressive reasoning in hidden space through an instruction-conditioned latent reasoning trajectory constructed from silent tokens. It combines process-supervised explicit-to-implicit distillation with retrieval-grounded supervision, using stage-wise CoT reconstruction to shape intermediate latent states and retrieval-effect credit to optimize incremental retrieval gains across the latent reasoning trajectories. Experiments on reasoning-intensive retrieval benchmarks show that RGLT consistently outperforms strong baselines while preserving efficient embedding inference.
Tags
Links
- Source: https://arxiv.org/abs/2608.14107v1
- Canonical: https://arxiv.org/abs/2608.14107v1
Trouble viewing inline? Open PDF directly →
Full Text
50,422 characters extracted from source content.
Expand or collapse full text
Retrieval Grounding Latent Reasoning for Dense Retrieval Gang Zhou 1,*,‡ Xiongxi Yu 2,* Hu Tian 3,† Yang Wei 2,† Lu Pan 2 Ke Zeng 2 Shibiao Xu 1,§ Xiaolong Zheng 4,§ 1 School of Artificial Intelligence, Beijing University of Posts and Telecommunications 2 Meituan LongCat Interaction Team 3 School of Management Science and Engineering, Central University of Finance and Economics 4 Institute of Automation, Chinese Academy of Sciences zhougang2023@bupt.edu.cn yuxiongxi@meituan.com tianhu01@foxmail.com weiyang14@meituan.com panlu02@meituan.com zengke02@meituan.com shibiaoxu@bupt.edu.cn xiaolong.zheng@ia.ac.cn ABSTRACT Reasoning-intensive retrieval requires text representations to capture not only semantic similarity, but also the reasoning needed to determine relevance under a given retrieval instruction. Existing reasoning-enhanced embedding models improve retrieval by incorporating reasoning information into dense representations, yet their supervision is typically dominated by the final retrieval objective. As a result, latent reasoning trajectories may learn shortcut reasoning patterns that preserve retrieval per- formance without producing meaningful incremental retrieval gains. We propose Retrieval Grounding Latent Reasoning (RGLT), a latent reasoning framework for dense retrieval that explicitly connects intermediate latent transitions with retrieval improvements. RGLT performs non-autoregressive rea- soning in hidden space through an instruction-conditioned latent reasoning trajectory constructed from silent tokens. It combines process-supervised explicit-to-implicit distillation with retrieval- grounded supervision, using stage-wise CoT reconstruction to shape intermediate latent states and retrieval-effect credit to optimize incremental retrieval gains across the latent reasoning trajectories. Experiments on reasoning-intensive retrieval benchmarks show that RGLT consistently outperforms strong baselines while preserving efficient embedding inference. 1 Introduction Reasoning-intensive retrieval requires text representations to capture not only semantic similarity, but also the multi-stage reasoning needed to identify relevant evidence. For specialized domains like mathematics, science and programming, target documents may share little surface-level similarity with the query. Their relevance may emerge only after the implicit constraints and intermediate concepts are resolved. Recent reasoning-intensive benchmarks show that dense retrievers effective on conventional semantic retrieval may struggle on retrieval tasks requiring multi-stage reasoning, while explicit reasoning methods exhibit substantially greater potential (Su et al., 2025). These findings suggest that, in reasoning-intensive retrieval tasks, query representations should incorporate inferred evidence beyond literal semantics to identify relevant documents. One line of research addresses this problem by expanding the query with a generated explicit Chain-of-Thought (CoT), and then encoding the enriched text for retrieval. Despite their effectiveness, these methods introduce autoregressive decoding latency and make retrieval quality highly dependent on the verbosity and lexical form of the generated rationale. Subsequent studies on continuous reasoning have explored replacing textual reasoning steps with hidden states, showing that multi-stage reasoning can work without verbalizing every intermediate stage in natural language (Hao et al., 2025; Shen et al., 2025). More recently, embedding models have further internalized explicit CoT into latent tokens to learn reasoning-enhanced representations, allowing the final embeddings to capture reasoning information without long autoregressive generation (Jin et al., 2026; Cai et al., 2026). These works establish latent reasoning as a promising direction for combining reasoning ability with efficient dense retrieval. * Gang Zhou and Xiongxi Yu contributed equally to this work. † Hu Tian and Yang Wei are corresponding authors. ‡ This work was conducted during Gang Zhou’s internship at Meituan. § Shibiao Xu and Xiaolong Zheng, as Gang Zhou’s advisors, primarily provided guidance on topic selection for this work. arXiv:2608.14107v1 [cs.AI] 14 Aug 2026 RGLT (a) Existing Latent Reasoning Embedding State Alignment 푽 ퟏ ∗ 풁 ퟏ 풁 ퟐ 풁 ퟑ (b) RGLT Latent Reasoning Embedding 풁 ퟏ 풁 ퟐ 풁 ퟑ Latent CoT Token Latent CoT Token State similarity ≠ ranking gainEach transition is grounded in ranking gain Explicit CoT Stage 푑 + 푑 − 푽 ퟐ ∗ 푽 ퟑ ∗ 푽 ퟏ ∗ 푽 ퟐ ∗ 푽 ퟑ ∗ △ 2 ∗ △ 3 ∗ Explicit CoT Stage 푑 + 푑 − Ranking- effect Gain Figure 1: Comparison of latent reasoning supervision. Existing methods supervise latent states through state alignment, which does not guarantee retrieval gains during reasoning evolution. RGLT instead transfers stage-wise ranking effects ∆ ∗ k from explicit CoT to latent transitions, ensuring that each reasoning step contributes to retrieval improvement. However, constructing a latent reasoning trajectory does not necessarily translate into retrieval-relevant reasoning information. Existing approaches typically supervise latent reasoning through final retrieval objectives, distillation on final representations, or alignment between intermediate latent and explicit states (Figure 1a). However, these objectives mainly optimize final retrieval outcomes or representation similarity, without directly modeling whether intermediate state transitions produce incremental retrieval gains. As a result, latent reasoning trajectories may learn shortcut reasoning patterns that maintain retrieval quality without contributing meaningful retrieval improvements across reasoning stages. The core challenge is therefore to shape latent reasoning trajectories whose state transitions progressively improve retrieval discrimination between relevant documents and hard negatives, rather than merely producing plausible intermediate representations. To address this challenge, we propose Retrieval Grounding Latent Reasoning (RGLT), a representation framework that performs non-autoregressive reasoning directly in latent space. RGLT appends a fixed sequence of silent tokens to construct an ordered latent reasoning trajectory, while a retrieval-intent anchor injects retrieval-oriented guidance into the trajectory based on the query and retrieval instruction. During training, explicit CoT is divided into multiple reasoning stages and provides cumulative semantic supervision for the corresponding latent states. Direct contrastive supervision further maintains the retrieval discriminability of each stage. More importantly, RGLT assigns stage-level credit according to how much each explicit CoT stage improves the retrieval preference for relevant documents over hard negatives. These retrieval gains are then propagated to the corresponding latent transitions, enabling the latent reasoning trajectory to learn retrieval-effective multi-step reasoning without enforcing direct representation alignment between the explicit and latent trajectories (Figure 1b). The terminal latent state is used directly as the query embedding, requiring neither autoregressive CoT generation nor additional pooling or projection at inference. Our contributions are summarized as follows: • We identify an underexplored limitation of latent reasoning retrievers: intermediate transitions are not explicitly supervised for their incremental retrieval effects. •We introduce stage-level retrieval-effect credit assignment. Each latent transition must recover the ranking gain induced by its corresponding explicit CoT stage, ensuring that intermediate reasoning is grounded in actual retrieval improvements. •We propose RGLT, a non-autoregressive latent reasoner that surpasses autoregressive counterparts on three reasoning-intensive benchmarks, substantially reducing inference latency. 2 Related Work 2.1 Explicit Reasoning for Dense Retrieval Recent work shows that LLM-generated reasoning can substantially improve retrieval for such settings. BRIGHT (Su et al., 2025) demonstrates the benefit of explicit reasoning on reasoning-intensive benchmarks. Data-centric methods, 2 RGLT including ReasonIR (Shao et al., 2025), RaDeR (Das et al., 2025), and ReasonEmbed (Chen et al., 2026), construct reasoning-oriented training instances, while Search-R3 (Gui & Cheng, 2025) and GRACE (Sun et al., 2025) explicitly generate reasoning before forming the retrieval representation. Meanwhile, LLM-based retrievers improve dense retrieval through instruction tuning, synthetic supervision, and joint generative–representational learning (Wang et al., 2024; Muennighoff et al., 2025; Zhang et al., 2025a), while still encoding queries in a single forward pass. Query expansion methods instead externalize reasoning into text: HyDE (Gao et al., 2023) generates a hypothetical document, Query2doc (Wang et al., 2023) produces a pseudo-document, and DIVER (Sun et al., 2026) iteratively refines the query with reasoning and retrieved evidence. Despite their effectiveness, these methods introduce autoregressive latency and make retrieval quality dependent on the generated text. 2.2 Latent Reasoning for Dense Retrieval Recent work has explored replacing textual reasoning with latent reasoning performed directly in hidden space. Coconut (Hao et al., 2025) iteratively feeds hidden states back as continuous thoughts, while CODI (Shen et al., 2025) distills explicit CoT into continuous representations. In retrieval settings, GIRCSE (Tsai et al., 2025) generates soft embedding tokens, multi-query retrieval captures different relevance facets (Chen et al., 2025a), and AdaQR (Zhang et al., 2025b) approximates query reasoning through embedding-space transformations. However, these methods do not explicitly model whether intermediate latent transitions produce incremental retrieval gains. Among existing retrieval-oriented latent reasoning methods, LaSER (Jin et al., 2026), which aligns explicit and latent trajectories, is most closely related to our work. While this constrains intermediate states, it ignores whether transitions between them yield actual retrieval gains. Instead, our method transfers stage-specific retrieval effects from explicit CoT to latent transitions, optimizing intermediate steps directly for retrieval improvements rather than mere state alignment or final-outcome optimization. 3 Methodology In this section, we introduce the proposed RGLT in detail, whose overall architecture is illustrated in Figure 2 and whose training pipeline is illustrated in Figure 3. RGLT appends a fixed number of latent reasoning tokens to the end of the input sequence, forming an implicit CoT trajectory in a non-autoregressive manner. During training, both the reasoning signals encoded in explicit CoT and the retrieval effects induced by its individual reasoning stages jointly guide the evolution of the latent reasoning tokens. Stage-level retrieval supervision further requires each reasoning stage to yield tangible gains in document ranking. 3.1 Problem Formulation Given a user queryq, a retrieval instructionI, and a candidate document corpusD = d 1 ,...,d N , dense retrieval learns an encoderf θ (·)that maps queries and documents into a sharedm-dimensional embedding spaceR m . Retrieval is then performed by comparing the query and document embeddings using cosine similarity:s(q,d) = cos(v q ,v d ), wherev q = f θ ([q;I])andv d = f θ (d)denote the query and document embeddings, and[·;·]denotes concatenation under a fixed input template. Reasoning-enhanced embedding models extend standard dense retrieval by allowing the model to construct a CoT (chain-of-thought) trajectory before forming the final query embedding: [q;I]−→ r 1 −→·−→ r L −→ v ∗ q .(1) The intermediate reasoning unitsr 1 ,...,r L may take the form of explicit tokens generated autoregressively or latent states evolving in continuous space. Such multi-stage reasoning can improve retrieval by progressively refining the query representation. However, autoregressive reasoning increases decoding latency and makes the resulting retrieval representation sensitive to the lexical structure of the generated reasoning trajectory, reducing the efficiency advantages of embedding-based retrieval. In this work, we construct the latent reasoning trajectory directly within the encoder by appendingKfixed silent tokens, T 1 ,...,T K , to the instruction-conditioned query: [q;I;T 1 ;... ;T K ] f θ −→ v ∗ q .(2) The contextual states associated with these silent tokens form an implicit reasoning trajectory, where successive latent transitions progressively refine the retrieval representation. The entire latent reasoning process is completed within a single forward pass, avoiding long autoregressive decoding while maintaining efficient embedding inference. 3 RGLT 1 Input Construction [q][q][I][I][A][A] queryinstr.anchor a AnchorState X q = [q][q] [a][a][T₁][T₁][T₂][T₂] ⋯ [Tₖ][Tₖ] 2 Latent Reasoning in Encoder Transformer Encoder LayersTransformer Encoder Layers [q][q] [a][a] [T₁][T₁][T₂][T₂] ⋯ [Tₖ][Tₖ] W g • W d ෨ ℎ W u + γ ℎ KV Modulation on Silent Tokens only on T 3 Output [q][q][a][a][T₁][T₁][T₂][T₂] ⋯ [Tₖ][Tₖ] ⋯ 푣 푞 (Final QueryEmbedding) Final Silent Token Transformer Encoder LayersTransformer Encoder Layers ෨ ℎ Encoder ℎ 퐴 Figure 2: Architecture of instruction-conditioned latent reasoning. The query and instruction are first compressed into a unified anchor state. During encoding, the anchor explicitly modulates the key and value representations of appended silent tokens via a gated residual to guide reasoning. Finally, the hidden state of the terminal silent token is directly extracted as the final query embedding. 3.2 Instruction-Conditioned Latent Reasoning Trajectory Although the non-autoregressive latent reasoning process is defined over a fixed latent trajectory, the silent positions still need to be conditioned on the retrieval instruction. This conditioning is necessary because the same query may correspond to different retrieval objectives under different instructions. To keep the latent reasoning trajectory aligned with the intended retrieval objective, we compress the query and instruction into a unified contextual anchor. Specifically, we introduce a special anchor tokenA = <anchor>and first derive its hidden representation from the instruction- conditioned query context: a = Encoder([q;I;A]).(3) Subsequently, when formally constructing the latent reasoning sequence, we remove the instructionIfrom the sequence, directly inject the extracted vector a as the input embedding of the anchor, and append K fixed silent tokensT: X q = [q;a;T 1 ;... ;T K ].(4) With this structure, the retrieval instruction influences the latent reasoning trajectory primarily through the anchor representationa, which serves as the main conditioning signal for the subsequent latent transitions. To make the anchor representation consistently affect the evolution of the silent tokens, we modulate their attentionKandVrepresentations at each layer. Lethdenote the originalKorVvector of a silent token in the current layer. We then use the anchor hidden state h A from the same layer to produce an updated representation e h through a gated low-rank residual: e h = h + γW u [(W d h)⊙ σ(W g h A + b)].(5) 4 RGLT Query 풒 + Instruction 푰 Anchor + KSilent Tokens ExternalLLM Initial CoT Rationale Text ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ ~~~~ Rationale Segmentation ✂ Segmented Explicit CoTText Stages 3 Direct Retrieval Supervision ... − ⋯ Total loss =+ λ₁ + λ₂ 푟 푆 푟 2 푟 1 푧 푠 푧 2 푧 1 ... Text Reconstruction / CE Stage-level Retrieval Discrimination Contrastive Retrieval (pos vs. neg) Shared Retrieval Candidates Hard Negatives + Positive Encoder 2 Retrieval-Effect Transfer 푢 푆 ... 푢 1 푢 0 푧 0 푧 1 푧 푆 ... Align Transition Effects 푢 2 푧 2 ഥ Δ 0 푒푥푝 ഥ Δ 1 푒푥푝 ഥ Δ 푆 푒푥푝 ഥ Δ 2 푒푥푝 ഥ Δ 0 푙푎푡 ഥ Δ 1 푙푎푡 ഥ Δ 푆 푙푎푡 ഥ Δ 2 푙푎푡 Latent States / Tokens Explicit CoTText / States Latent trajectory (K silent tokens) ⋯ 풎=푲/푺 푺풔풕풂품풆풔 Stage-Terminal Extraction ⋯ 1 Text Reconstruction (CE) ⋯ ⋯ 퐻 ≤1 lat = 푞;푎;푇 1:푀 푟 푠 푟 2 푟 1 퐻 ≤2 lat = 푞;푎;푇 1:2푀 퐻 ≤푆 lat = 푞;푎;푇 1:푆푀 CE CE CE T 1 T 푚 T 푚+1 T 2푚 T 푠−1푚+1 T 푘 푧 푠 푧 2 푧 1 Latent Stage-Terminal States B B B B K B B B S S S S Figure 3: The training pipeline of RGLT. It illustrates the generation of the latent reasoning trajectory and the subsequent joint optimization using three complementary losses: text reconstruction loss (L rec ), retrieval-effect transfer loss (L eff ), and direct retrieval supervision loss (L ret ). Here,W g andbparameterize the gating function that produces a channel-wise modulation signal, whileW d andW u implement low-rank projections to reduce additional computation. The anchor representation therefore influences how information is propagated through the latent reasoning trajectory across layers. To preserve the reusability of document embeddings in dual-encoder retrieval, this modulation is applied only during query encoding and remains disabled for document representations. The silent tokens accumulate contextual information sequentially, allowing the final silent tokenT K to summarize the latent reasoning trajectory without additional pooling or projection layers. We therefore use its hidden state as the final query embedding: v q = f θ (X q ,T K ).(6) 3.3 Process-Supervised Explicit-to-Implicit Reasoning Distillation Despite the expressive capacity of silent tokens for latent reasoning, training only with the final retrieval objective provides little signal for intermediate reasoning stages. As a result, latent reasoning may converge to shortcut reasoning patterns that maintain retrieval performance through shallow matching cues rather than retrieval-effective state transitions. To mitigate this issue, we introduce Process-Supervised Explicit-to-Implicit Reasoning Distillation, which uses explicit CoT to shape the latent reasoning trajectory and imposes stage-wise reconstruction constraints on the intermediate latent states. For each training queryq, we use an external LLM to generate an explicit multi-stage CoT rationale, from whichS ordered reasoning stages(r 1 ,r 2 ,...,r S )are extracted via regular expression matching without hard truncation. We further construct cumulative reasoning prefixesr ≤k by concatenating the firstkstages. These cumulative prefixes serve as intermediate reconstruction targets for the latent reasoning trajectory. We divide theKsilent tokens intoSlatent reasoning stages with an equal stage sizeB = K/S. At the end of the k-th latent stage, the corresponding latent states are required to reconstruct the cumulative reasoning prefixr ≤k . We optimize this objective using a teacher-forced token-level cross-entropy loss: L rec = 1 S S X k=1 CE θ (r ≤k | [q;A(a);T 1:k×B ]).(7) 5 RGLT To avoid information leakage, only the silent tokens up to stagek× Bare exposed when computing the reconstruction loss forr ≤k . This process-supervised reconstruction objective encourages the latent reasoning trajectory to capture progressively richer reasoning information across stages. 3.4 Retrieval-Grounded Supervision for Latent Reasoning Semantic reconstruction encourages latent stages to capture progressively richer reasoning information, but does not directly constrain whether intermediate states and transitions improve retrieval discrimination. We therefore introduce retrieval-grounded supervision from two complementary perspectives: direct stage supervision, which preserves retrieval discrimination throughout the latent reasoning trajectory, and retrieval-effect credit, which models the incremental retrieval gains induced by adjacent reasoning transitions. Direct stage supervision. A latent reasoning trajectory should maintain retrieval discrimination throughout the reasoning process rather than only at the final stage. Relying solely on final-state supervision risks leaving early latent steps unconstrained, which may cause intermediate representations to drift outside the meaningful retrieval space. To encourage retrieval-effective intermediate states, we apply retrieval supervision to both the final representation and the intermediate latent stages. LetBbe the candidate pool,P q ⊆Bits positive subset, andτthe contrastive temperature. For a query state v, we define the multi-positive contrastive loss and the overall retrieval objective as: R(v;B) =− log P d∈P q exp(s(v,d)/τ ) P d∈B exp(s(v,d)/τ ) ,(8) L ret = R(v S ;B) + αR(v ∗ S ;B) + β S X k=1 w k R(v k ;B) (9) whereSis the number of reasoning stages,v S is the final embedding representation generated from the latent trajectory, andv ∗ S is the corresponding explicit representation (Hereafter, the superscript∗denotes states and associated quantities derived from the explicit Chain-of-Thought (CoT) trajectory). The non-negative stage weightsw k sum to one. The first two terms optimize the final latent and explicit retrieval representations, while the last term encourages intermediate latent stages to preserve retrieval discrimination throughout the reasoning trajectory. Retrieval-effect credit. Supervision on intermediate latent states alone does not capture whether a state transition produces meaningful retrieval gains. To explicitly model the retrieval effect of each reasoning transition, we measure how the retrieval preference over relevant documents and hard negatives changes between adjacent reasoning stages. Rather than relying on raw similarity differences, which can be trivially amplified without improving retrieval discrimination, we compute transition gains within a normalized local document distribution defined over a subsetS ⊆Bcontaining the positive document and several hard negatives: p v (d) = exp(s(v,d)/τ d ) P d ′ ∈S exp(s(v,d ′ )/τ d ) ,(10) whereτ d is the distillation temperature. This distribution compresses unbounded similarity scores into the(0, 1)interval and forces the probabilities over the local subsetSto sum to 1. This competitive normalization increases retrieval discrimination by forcing gains on relevant documents to come at the expense of hard negatives within the same candidate set. Based on this relative confidence space, the latent and explicit credits of stage k are strictly defined as: ∆ k (d) = logp v k (d)− sg logp v k−1 (d) , ∆ ∗ k (d) = logp v ∗ k (d)− logp v ∗ k−1 (d), (11) where the stop-gradient operatorsg[·]forcibly fixes the preceding latent state as a comparison baseline. This prevents the model from artificially increasing the transition gain by degrading the preceding state, forcing the incremental retrieval effect to be attributed to the current statev k . Meanwhile, the explicit increment∆ ∗ k (d)calculated from external CoT data officially serves as the “effect target” for learning here. Letπ k (d) = sg[p v ∗ k−1 (d)]be the distribution of the explicit preceding state. We zero-mean center both credits underπ k to eliminate global shifts independent of the candidate documents: ∆ k (d) = ∆ k (d)− E d ′ ∼π k [∆ k (d ′ )], ∆ ∗ k (d) = ∆ ∗ k (d)− E d ′ ∼π k [∆ ∗ k (d ′ )]. (12) The final effect loss is: L eff = S X k=1 E d∼π k ∆ k (d)− sg[∆ ∗ k (d)] 2 .(13) 6 RGLT Table 1: Overall performance on reasoning-intensive retrieval benchmarks. All values are percentages. ModelSize BRIGHTFollowIRBrowseComp-Plus nDCG@10R@10Robust04News21Core17Avg.Avg.R@5R@100R@1000 MAP@5nDCG@5MAP@5Scorep-MRR Standard dense retrievers BGE-M30.6B11.4014.391.5021.407.5010.10-3.904.1021.7047.20 E5-Mistral-7B-Instruct7B16.5020.122.7028.8013.7015.10-1.509.3037.1070.20 Qwen3-Embedding-8B8B14.0017.313.1025.5010.1012.907.207.7031.6061.30 Basic contrastive training Fair Baseline (Qwen3-0.6B)0.6B18.3022.142.1013.406.707.40-0.203.5021.3046.90 Fair Baseline (Qwen3-8B)8B25.7030.452.8018.9011.2011.001.7011.3037.4063.20 Fair Baseline (LLaMA3.1-8B)8B22.5026.862.5018.908.109.800.106.1025.4050.80 Explicit reasoning Rewrite-then-Retrieve (Qwen3-8B)8B28.1033.15– Search-R3 (Qwen2.5-1.5B)1.5B7.7010.243.2026.2010.1013.203.200.000.301.10 InBedder (LLaMA2-7B)7B5.567.502.608.501.604.20-0.101.408.1027.40 Latent reasoning GIRCSE (Qwen3-8B)8B26.0030.793.0022.608.5011.402.0013.0040.8068.10 LaSER (Qwen3-8B) ⋆ 8B29.9034.274.1021.8011.4012.502.4011.7038.4066.90 RGLT (Qwen3-8B)8B34.2039.9712.9023.604.6013.702.6213.4241.870.83 Rewrite-then-Retrieve uses an external LLM to generate a reasoning-enhanced query before retrieval. AoPS Bio Earth Econ Leet Pony Psych Robot Stack Sustain TheoQ TheoT 0 20 40 60 nDCG@10 LaSER RGLT Figure 4: Domain-level BRIGHT nDCG@10. RGLT outperforms LaSER on 9 of 12 domains. L eff transfers the stage-specific retrieval effects induced by explicit CoT reasoning to the corresponding latent transitions. Rather than constraining latent states to match explicit representations, the objective focuses on whether each latent transition produces comparable retrieval gains over relevant documents and hard negatives. This encourages the latent reasoning trajectory to learn retrieval-effective state transitions that progressively improve retrieval discrimination across reasoning stages. 3.5 Objective Function The complete optimization objective is defined as: L =L ret + λ 1 L rec + λ 2 L eff ,(14) whereλ 1 andλ 2 balance semantic reconstruction and retrieval-effect credit. These auxiliary weights are ramped up from zero early in training, allowing the encoder to establish a stable document space before enforcing latent trajectory constraints. 4 Experiments We organize the experiments around four research questions. RQ1: Effectiveness. Does RGLT improve reasoning- intensive retrieval over standard, explicit-reasoning, and latent-reasoning retrievers? RQ2: Components. Which parts of RGLT account for the improvement? RQ3: Representation evolution. Do the latent stages yield progressively stronger retrieval representations? RQ4: Efficiency and sensitivity. What is the query-side cost, and how sensitive is the method to the silent-token budget? 7 RGLT Table 2: Component ablations and architectural variants on BRIGHT. The lower block contains independently trained variants and is not a one-factor ablation. VariantnDCG@10R@10 RGLT34.2039.97 w/o stage retrieval-effect matching31.2636.15 w/o intent retrieval-effect matching31.6936.85 w/o ST16-vs-base advantage26.3932.56 replace effect matching with state alignment30.6435.69 Gated retrieval-credit fusion variant32.1437.94 Naive terminal-ST readout variant30.8737.01 Reasoning-query distillation variant32.0137.85 4.1 Experimental Setup Datasets and metrics.We train on the 81,659 instances from ReasonEmbed (Chen et al., 2026), where each query is paired with relevant documents, hard negatives, and an explicit reasoning path. We evaluate on BRIGHT (Su et al., 2025), FollowIR (Weller et al., 2025), and BrowseComp-Plus (Chen et al., 2025b). BRIGHT contains 1,384 queries and 1,145,164 documents from 12 technical and scientific domains; we report nDCG@10 and Recall@10. For FollowIR, we report the averaged task score and p-MRR. For BrowseComp-Plus, we report Recall@5, Recall@100, and Recall@1000. Baselines. We compare with four groups of methods. Standard retrievers include BGE-M3 (Chen et al., 2024), E5-Mistral-7B-Instruct (Wang et al., 2022), and Qwen3-Embedding-8B (Zhang et al., 2025a). The Fair Baseline fine-tunes the same Qwen3-8B backbone and training data using only contrastive learning. Explicit reasoning methods include Rewrite-then-Retrieve, Search-R3 (Gui & Cheng, 2025), and InBedder (Peng et al., 2024). Implicit reasoning methods include GIRCSE (Tsai et al., 2025) and LaSER (Jin et al., 2026). LaSER is closely related to our work because it connects explicit CoT reasoning with latent retrieval representations through intermediate trajectory supervision. Implementation details.We use Qwen3-8B. TheKsilent tokens form four cumulativeK/4stages, and the hidden state ofT K serves as the query embedding (document encoding omits silent tokens). We train for three epochs (LoRA rank 32) via Hugging Face Accelerate (torch.distributed) on an internal Linux cluster of 32 GPUs (4×8 A100, 64GB RAM/worker; Docker: Python 3.9, PyTorch 2.3.1, CUDA 12.4.0, FlashAttention 2.5.8, NCCL 2.20.5, OpenMPI 4, GCC 11). A per-device query batch of 4 (global 128) pairs 1 positive with 3 hard negatives per query, forming a 512-document pool via cross-device gathering. Online-generated explicit CoT states from the shared backbone act as stop-gradient targets. BRIGHT evaluation spans all 12 domains (1,384 queries) with fresh document embeddings and retrieval instructions enabled. With auxiliary weights optimally set toλ 1 = λ 2 = 0.10, all reported results are averaged over three independent runs. 4.2 Main Results RQ1: Effectiveness. Table 1 shows that RGLT achieves the best completed BRIGHT result, with 34.20 nDCG@10 and 39.97 Recall@10. It improves over LaSER by 14.4% and 16.6% respectively, and exceeds the Fair Baseline by 8.50 nDCG points. RGLT also outperforms GIRCSE and Rewrite-then-Retrieve by 8.20 and 6.10 nDCG points. These results show that retrieval-effect supervision is more effective than contrastive-only latent refinement and avoids explicit CoT generation at inference. Figure 4 shows that the gain is broad rather than domain-specific. RGLT improves 9 of 12 domains, with clear gains on AoPS, psychology, robotics, and both TheoremQA subsets. It also improves LeetCode, economics, and sustainable living. LaSER remains better on biology, earth science, and StackOverflow, but the gaps on the first two are small. Moreover, RGLT obtains higher Recall@10 on 11 of 12 domains. The macro improvement therefore comes from more balanced retrieval across domains, not from a single outlier. 4.3 Ablation and Architecture Analysis RQ2: Components. Table 2 first shows that a terminal silent state alone is insufficient. The naive terminal-ST variant reaches 30.87 nDCG@10, whereas RGLT reaches 34.20 by grounding the same stage terminals in explicit retrieval effects. The proposed stage-terminal design also outperforms the earlier gated-fusion variant, increasing nDCG@10 from 32.14 to 34.20 and Recall@10 from 37.94 to 39.97. Thus, the latent trajectory can directly serve as the retrieval representation without an external fusion module. All ablated variants underperform the full model. In particular, 8 RGLT Table 3: BRIGHT performance of cumulative latent stages. ST4–ST16 are the terminal states of four-token stages. Query representationnDCG@10R@10 Base query state26.3930.24 Stage 1 terminal (ST4)32.2534.96 Stage 2 terminal (ST8)32.4835.25 Stage 3 terminal (ST12)33.5637.14 Stage 4 terminal (ST16)34.2039.97 Table 4: Query-side inference efficiency on BRIGHT. Latency is measured over 80 queries with batch size 8 on one NVIDIA A100 80GB GPU. MethodLatencyRelative Text gen. Index cost (ms/query) Base retriever (Qwen3-8B)21.51.00×NoNone Rewrite-then-Retrieve4000186×YesNone LaSER (Qwen3-8B)32.51.51×NoNone RGLT (Qwen3-8B)24.01.12×NoNone replacing retrieval-effect matching with state alignment causes a clear drop, confirming that modeling stage-wise ranking changes is more effective than aligning intermediate states alone. 4.4 Retrieval-Grounded Representation Evolution RQ3: Representation evolution. Table 3 reports the retrieval performance of the base representation and the four stage-terminal states. nDCG@10 increases from 26.39 at the base state to 32.25, 32.48, 33.56, and 34.20 at ST4, ST8, ST12, and ST16, respectively. Recall@10 follows the same trend, increasing from 30.24 to 39.97. These results show that retrieval utility accumulates progressively along the latent trajectory rather than emerging only at the final stage. The best performance at ST16 also supports its direct use as the final query embedding. 4.5 Efficiency and Sensitivity Analysis RQ4: Efficiency and sensitivity. Table 4 details the query-side inference efficiency across different methods. As shown, explicitly generating textual CoT (Rewrite-then-Retrieve) introduces a prohibitive186×latency overhead. In contrast, RGLT operates without generating textual CoT, leaves the document encoding unchanged, and fully preserves single- vector nearest-neighbor search. While LaSER also avoids text generation, its recurrent construction of latent tokens incurs a1.51×latency penalty. RGLT overcomes this by exposing allKfixed silent positions simultaneously in a non-autoregressive forward pass, directly extracting the STKrepresentation. Because the extra computation is strictly limited to a fixed query-side sequence extension, RGLT achieves a latency of 24.0 ms/query under identical measurement setups, representing a minimal 12% overhead (1.12×) over the base retriever. Figure 5 evaluates retrieval performance across different latent budgets (K). Performance peaks at the default setting of K = 16(four tokens per stage) for both nDCG@10 and Recall@10. However, metrics drop atK = 24, indicating that performance does not scale indefinitely with the latent budget; an excessively largeKmay complicate optimization, making K = 16 the optimal choice. 8121624 30 32 34 nDCG@10 K 8121624 36 38 40 R@10 K Figure 5: Impact of the latent budget (K) on retrieval performance. 9 RGLT 5 Conclusion In this paper, we presented RGLT, a latent reasoning framework for reasoning-intensive dense retrieval. Unlike existing methods that supervise only the final representation or align intermediate states, RGLT explicitly grounds each latent transition in the retrieval effect of its corresponding CoT stage. This enables the model to learn an ordered evolution of retrieval representations, while the terminal latent state can be used directly as a single-vector query embedding without autoregressive CoT generation, additional pooling, or document-side modification. Experiments on reasoning-intensive retrieval benchmarks show that RGLT consistently outperforms standard, explicit-reasoning, and latent-reasoning retrievers. Analyses confirm that retrieval quality improves progressively across latent stages and that retrieval-effect matching is more effective than conventional state alignment. These results show that latent reasoning becomes more effective for retrieval when its intermediate computation is grounded in document-ranking improvements. 10 RGLT References Zhixin Cai, Jun Bai, Yang Liu, Jiaqi Li, Yichi Zhang, Taichuan Li, Zhuofan Chen, Zixia Jia, Zilong Zheng, and Wenge Rong. Xetrieval: Mechanistically explaining dense retrieval. arXiv preprint arXiv:2605.29507, 2026. Hung-Ting Chen, Xiang Liu, Shauli Ravfogel, and Eunsol Choi. Beyond single embeddings: Capturing diverse targets with multi-query retrieval. arXiv preprint arXiv:2511.02770, 2025a. Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. M3-Embedding: Multi-linguality, multi- functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216, 2024. Jianlyu Chen, Junwei Lan, Chaofan Li, Defu Lian, and Zheng Liu. ReasonEmbed: Enhanced text embeddings for reasoning-intensive document retrieval. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens (eds.), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1203–1221, San Diego, California, United States, July 2026. Association for Computational Linguistics. ISBN 979-8-89176-390-6. doi: 10.18653/v1/2026.acl-long.54. URLhttps://aclanthology.org/2026.acl- long.54/. Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Sahel Sharifymoghaddam, Andrew Liu, Joshua Green, Kshama Patel, Ruoxi Meng, et al. BrowseComp-Plus: A more fair and transparent evaluation benchmark of deep-research agent. In First Workshop on Multi-Turn Interactions in Large Language Models, 2025b. Debrup Das, Sam O’Nuallain, and Razieh Rahimi. RaDeR: Reasoning-aware dense retrieval models. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Confer- ence on Empirical Methods in Natural Language Processing, p. 19970–19997, Suzhou, China, November 2025. Association for Computational Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.1011. URL https://aclanthology.org/2025.emnlp-main.1011/. Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. Precise zero-shot dense retrieval without relevance labels. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1762–1777, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.99. URLhttps://aclanthology .org/2023.acl-long.99/. Yuntao Gui and James Cheng. Search-R3: Unifying reasoning and embedding generation in large language models. arXiv preprint arXiv:2510.07048, 2025. Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason E Weston, and Yuandong Tian. Training large language models to reason in a continuous latent space. In Second Conference on Language Modeling, 2025. Jiajie Jin, Yanzhao Zhang, Mingxin Li, Dingkun Long, Pengjun Xie, Yutao Zhu, and Zhicheng Dou. Internalizing explicit reasoning into latent space for dense retrieval. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’26, p. 689–700, New York, NY, USA, 2026. Association for Computing Machinery. ISBN 9798400725999. doi: 10.1145/3805712.3809575. URL https://doi.org/10.1145/3805712.3809575. Niklas Muennighoff, Hongjin SU, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. Generative representational instruction tuning. In Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (eds.), International Conference on Learning Representations, volume 2025, p. 45544–45613, 2025. URLhttps://proceedings. iclr.c/paper_files/paper/2025/file/70cf215430492f7d34830a24e744b3f1-Paper- Conference.pdf. Letian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa, Gaowen Liu, Zihan Wang, and Jingbo Shang. Answer is all you need: Instruction-following text embedding via answering the question. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 459–477, 2024. Rulin Shao, Rui Qiao, Varsha Kishore, Niklas Muennighoff, Xi Victoria Lin, Daniela Rus, Bryan Kian Hsiang Low, Sewon Min, Wen-tau Yih, Pang Wei Koh, et al. ReasonIR: Training retrievers for reasoning tasks. In Second Conference on Language Modeling, 2025. Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, and Yulan He. CODI: Compressing chain-of-thought into continuous space via self-distillation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 677–693, 2025. Hongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi, Niklas Muennighoff, Han-yu Wang, Liu Haisu, Quan Shi, Zachary Siegel, Michael Tang, et al. BRIGHT: A realistic and challenging benchmark for reasoning-intensive retrieval. In International Conference on Learning Representations, volume 2025, p. 48941–48991, 2025. 11 RGLT Duolin Sun, Meixiu Long, Dan Yang, Junjie Wang, Yecheng Luo, Yue Shen, Jian Wang, Hualei Zhou, Chunxiao Guo, Peng Wei, Jiahai Wang, and Jinjie Gu. DIVER: A multi-stage approach for reasoning-intensive information retrieval, 2026. URL https://arxiv.org/abs/2508.07995. Jiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong, Pengcheng Jiang, Bowen Jin, Peiran Li, Weijia Shi, and Jiawei Han. GRACE: Generative representation learning via contrastive policy optimization. arXiv preprint arXiv:2510.04506, 2025. Yu-Che Tsai, Kuan-Yu Chen, Yuan-Chi Li, Yuan-Hao Chen, Ching-Yu Tsai, and Shou-De Lin. Let LLMs speak embed- ding languages: Generative text embeddings via iterative contrastive refinement. arXiv preprint arXiv:2509.24291, 2025. Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533, 2022. Liang Wang, Nan Yang, and Furu Wei. Query2doc: Query expansion with large language models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 9414–9423, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.585. URL https://aclanthology.org/2023.emnlp-main.585/. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. Improving text embeddings with large language models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 11897–11916, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.642. URL https://aclanthology.org/2024.acl-long.642/. Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini. FollowIR: Evaluating and teaching information retrieval models to follow instructions. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 11926– 11942, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. ISBN 979-8-89176-189-6. doi: 10.18653/v1/2025.naacl-long.597. URL https://aclanthology.org/2025.naacl-long.597/. Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. Qwen3 embedding: Advancing text embedding and reranking through foundation models, 2025a. URL https://arxiv.org/abs/2506.05176. Yichi Zhang, Jun Bai, Zhixin Cai, Shuhan Qin, Zhuofan Chen, Jinghua Guan, and Wenge Rong. Your dense retriever is secretly an expeditious reasoner. arXiv preprint arXiv:2510.21727, 2025b. 12 RGLT Supplementary Material Overview This supplement provides the implementation details needed to reproduce RGLT without repeating the methodology in the main paper. We focus on three core aspects: how the CoT sequences are structured and segmented, how training is made tractable under the memory pressure imposed by ultra-long CoT branches and a cross-device document pool, chiefly through a deterministic gradient replay strategy, and how the evaluation protocols and efficiency optimizations (including KV-cache reuse) are implemented. Explicit CoT is used exclusively during training. At inference, a query is represented by the final silent state and each document remains a reusable single vector. A Structured Segmentation of CoT The CoT sequences in this dataset typically comprise four stages: problem identification, reasoning and relevant information, detailed solution, and verification and answer convergence. We introduce a robust parsing mechanism to systematically segment the continuous text into these four distinct stages. To overcome the lack of standardized paragraph boundaries in raw texts, our parser adopts a two-tier strategy. First, it prioritizes high-level semantic headings and structural step markers while carefully avoiding the misclassification of ordinary numbered calculations or code lines. Second, when encountering highly non-standard texts with insufficient reliable boundaries, the parser gracefully degrades by identifying natural line breaks or sentence boundaries near the four equal-length partitions of the text. This ensures reliable and structured supervision without disrupting the logical coherence of the chain of thought. Furthermore, to prevent this segmentation operation from disrupting the original logical context, we do not let the model learn these four segmented slices (r 1 ,r 2 ,r 3 ,r 4 ) in isolation. Instead, we formulate them as cumulative targets: r ≤k = r 1 ∥r 2 ∥·∥r k , k ∈1, 2, 3, 4,(15) where∥denotes concatenation with a blank line. Through this cumulative mechanism, the original four continuous slices are reorganized so that each subsequent reasoning stage naturally retains and observes the prior deduction history. This perfectly maintains the logical coherence of the chain of thought while enabling stage-wise supervision. B Training under Memory Constraints Table 5 details our training configuration, which relies on a frozen Qwen3-8B base augmented with LoRA (rank 32) and RGLT (rank 8) parameters. Training is based on a joint objective that combines four terms: final retrieval, per-stage retrieval credit, CoT reconstruction, and a retrieval-effect match between silent-stage transitions and the explicit CoT-prefix transition, while reusing the same frozen backbone, LoRA, and RGLT parameters throughout. Concretely, the forward passes that must coexist in memory on every step include the explicit CoT branches (up to 8,192 tokens) for the sampled cumulative prefixes, the document embeddings (one positive plus three hard negatives per query, shared across the batch), and the before/after pairs fed to the effect and reconstruction objectives, all multiplied by the cross-device document pool of 512. Retaining the computation graph for all of these simultaneously is infeasible even on4× 8A100 80GB GPUs. We therefore adopt deterministic gradient replay. The representations for these branches are first computed in micro-batches and detached as leaf tensors. Once the loss backward pass yields the representation gradients, the corresponding forwards are replayed to backpropagate those exact gradients into the weights. To guarantee mathematical equivalence between the initial forward pass and the replay, we enforce strict zero dropout across all model layers and LoRA modules. The much shorter 512-token latent query graph remains attached normally. C Evaluation Details and Inference Optimization C.1 Benchmark Protocol BRIGHT.We evaluate all 1,384 queries and complete corpora from the 12 domains. Query instructions are enabled, and the maximum length is 512. LeetCode is split into four document shards and merged before scoring. The full evaluation recomputes document vectors, and cache metadata contains the model and checkpoint fingerprint. We report the unweighted macro average of domain nDCG@10 and Recall@10. FollowIR. We use the official MTEB 1.38.32 Robust04, News21, and Core17 instruction retrieval tasks. Their standard metrics are MAP@5, nDCG@5, and MAP@5, respectively. The query and instruction are passed through the 13 RGLT Table 5: Training configuration. SettingValue BackboneQwen3-8B GPU 4× 8 NVIDIA A100 80GB Precision / attentionbfloat16 / FlashAttention 2 Per-GPU / global query batch4 / 128 Per-query candidates1 positive + 3 hard negatives Cross-device document pool512 Query/document maximum length512 tokens CoT maximum length8,192 tokens, no truncation Silent tokens / stages16 / 4 LoRA rank / scale / dropout32 / 64 / 0 RGLT rank8 LoRA / mechanism learning rate 10 −5 / 5× 10 −5 Optimizer / weight decayAdamW / 0.01 Warm-up / schedule100 steps / cosine decay Gradient clipping1.0 separate channels used by the contextual anchor; documents remain instruction independent. In addition to the standard task metrics, we use the official p-MRR implementation to measure whether changed instructions move newly relevant documents upward and newly irrelevant documents downward. We macro-average each metric over the three tasks. BrowseComp-Plus. We use the benchmark’s fixed corpus, queries, and relevance labels and report Recall@5, Recall@100, and Recall@1000. The document index is built once from the document-only path and is not modified at query time. C.2 KV-Cache Inference and Efficiency Measurement Query KV-Cache Reuse. Naively running the anchor and latent paths as two complete forward passes would encode the query twice, roughly doubling the latency. To solve this, our implementation strictly reuses the causal query-prefix key/value cache. LetP (q)be the complete tokenized prefix shared by both paths, ending immediately after the query. Inference is decomposed as follows: C q = Prefill(P (q)), a = Forward(I;<anchor>| C q ) <anchor> , z 4 = Forward A(a); <|embed_token|>;T 1:K | C q T K , (16) whereC q contains the per-layer keys and values of the shared prefix. The anchor branch and latent branch receive separate cache objects that map to the same read-only prefix tensors. Crucially, to prevent memory bloat, the instruction and temporary anchor cache are immediately discarded afterais obtained. The latent branch then spawns directly from the original C q , completely bypassing the instruction text. Preserving Positional Integrity. When sharing caches across variable-length batches, misaligned position IDs can silently corrupt the representations. To guarantee mathematical equivalence to full recomputation, we enforce strict positional tracking. Prefixes are left-padded, and suffix position IDs are dynamically computed per row starting exactly from each example’s non-padding prefix length. Furthermore, to maximize GPU utilization, allKsilent tokens are processed concurrently in a single causal block rather than being decoded autoregressively token-by-token. Computational Overhead Analysis. Because of the above designs, the query prefix is encoded only once. The additional computational burden relative to a standard base retriever is strictly bounded to two short suffix blocks: the instruction-plus-anchor suffix and the fixed (K + 2)-token latent block (injected anchor, one readout,Ksilent tokens). This architectural isolation is the precise reason why our measured latency overhead is only∼1.2×, rather than the2× cost of two full forwards. Strict Validation and Profiling. To ensure our optimization does not compromise retrieval accuracy, we enforce a strict cache-correctness test. Cached inference must yield identical top-krankings compared to full uncached recomputation, with the maximum absolute difference of ST(final) embeddings bounded near floating-point epsilon. For latency benchmarking, we isolate the components by usingtorch.cuda.synchronize()around the prefill, anchor suffix, and latent suffix regions separately, explicitly excluding I/O-bound operations like tokenization and index construction. 14