Paper deep dive
BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking
Hyunkyung Park, Arkaitz Zubiaga
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/18/2026, 1:34:09 AM
Summary
The paper introduces BiCon-Gate, a semantics-aware consistency gate designed to improve dialogue fact-checking by selectively applying de-colloquialisation rewrites. By using a staged pipeline (normalisation and coreference resolution) and a bidirectional NLI-based gate, the system mitigates semantic drift, improving both evidence retrieval and fact verification on the DialFact benchmark.
Entities (5)
Relation Signals (2)
Hyunkyung Park ā authored ā BiCon-Gate
confidence 100% Ā· BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking Hyunkyung Park and Arkaitz Zubiaga
BiCon-Gate ā improves ā DialFact
confidence 95% Ā· On the DialFact benchmark, our approach improves retrieval and verification
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automated fact-checking in dialogue involves multi-turn conversations where colloquial language is frequent yet understudied. To address this gap, we propose a conservative rewrite candidate for each response claim via staged de-colloquialisation, combining lightweight surface normalisation with scoped in-claim coreference resolution. We then introduce BiCon-Gate, a semantics-aware consistency gate that selects the rewrite candidate only when it is semantically supported by the dialogue context, otherwise falling back to the original claim. This gated selection stabilises downstream fact-checking and yields gains in both evidence retrieval and fact verification. On the DialFact benchmark, our approach improves retrieval and verification, with particularly strong gains on SUPPORTS, and outperforms competitive baselines, including a decoder-based one-shot LLM rewrite that attempts to perform all de-colloquialisation steps in a single pass.
Tags
Links
- Source: https://arxiv.org/abs/2604.14389v1
- Canonical: https://arxiv.org/abs/2604.14389v1
Trouble viewing inline? Open PDF directly ā
Full Text
66,422 characters extracted from source content.
Expand or collapse full text
BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking Hyunkyung Park Arkaitz Zubiaga Queen Mary University of London hyunkyung.park, a.zubiaga@qmul.ac.uk Abstract Automated fact-checking in dialogue involves multi-turn conversations where colloquial language is frequent yet understudied. To address this gap, we propose a conservative rewrite candidate for each response claim via staged de-colloquialisation, combining lightweight surface normalisation with scoped in-claim coreference resolution. We then introduce BiCon-Gate, a semantics-aware consistency gate that selects the rewrite candidate only when it is semantically supported by the dialogue context, otherwise falling back to the original claim. This gated selection stabilises downstream fact-checking and yields gains in both evidence retrieval and fact verification. On the DialFact benchmark, our approach improves retrieval and verificationāwith particularly strong gains on SUPPORTSāand outperforms competitive baselines, including a decoder-based one-shot LLM rewrite that attempts to perform all de-colloquialisation steps in a single pass. BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking Hyunkyung Park and Arkaitz Zubiaga Queen Mary University of London hyunkyung.park, a.zubiaga@qmul.ac.uk 1 Introduction Automated fact-checking determines whether a claim is supported, refuted, or there is not enough information (NEI) to verify it by grounding the claim in evidence from a reliable knowledge source such as Wikipedia Thorne and Vlachos (2018); Thorne et al. (2018). Most systems follow a retrieveāverify pipeline: evidence retrieval (Information Retrieval; IR) followed by fact verification (FV). Dialogue fact-checking extends this setting to multi-turn conversations, where a response claim depends on preceding context C Kim et al. (2021). In dialogue, colloquial phenomenaācontractions, missing punctuation, inconsistent casing, ellipsis, and pronominal referencesācan blur claim boundaries and entity mentions (Figure 1), making both retrieval and verification brittle Kim et al. (2021); Chamoun et al. (2023). Figure 1: Example multi-turn dialogue illustrating a context-dependent response claim about the Apollo 11 mission. In-claim pronouns in the response refer to antecedents introduced in earlier turns. A common remedy is to rewrite conversational inputs to be more self-contained (e.g., normalisation, de-contextualisation, or incomplete-utterance rewriting), which can improve recall on noisy queries Sundriyal et al. (2023); Li et al. (2023); Deng et al. (2024); Guo et al. (2024); Cao et al. (2024). This creates a rewrite-induced IRāFV trade-off: making a claim more explicit can increase retrieval recall, while even small semantic drift (e.g., resolving a pronoun to the wrong entity) can mislead verification and degrade end-to-end (E2E) performance Gupta et al. (2022); Chamoun et al. (2023). We focus on in-claim pronominal anaphora because it is frequent in dialogue claims and can often be addressed via minimal span substitution rather than broad paraphrasing; in DialFact, 42.1% of claims contain at least one in-claim pronoun (§3.2). We ask whether we can improve retrieval and verification with minimal, meaning-preserving edits to the claim surface. To this end, we first construct a conservative rewrite candidate via staged de-colloquialisation: de-contraction, punctuation restoration, true-casing, and scoped in-claim pronoun resolution. We then propose BiCon-Gate, a bidirectional natural language inference (NLI)-based consistency gate that routes each instance to the rewrite candidate only when it is semantically supported by the dialogue context; otherwise, it falls back to the original claim Zhang et al. (2023). Our contributions are three-fold: ⢠Controlling the IRāFV trade-off. We use a semantics-aware gate as a conservative control mechanism that selectively accepts rewrites to mitigate verification harm from semantic drift. ⢠Isolating where rewriting helps or hurts. We disentangle retrieval and verification effects via IR-only, FV-only (with gold evidence), and E2E evaluations. ⢠Demonstrating drift in one-shot LLM rewriting. We compare against a decoder-only, single-prompt rewrite baseline and show that aggressive one-shot rewrites can drift and hurt verification, whereas scoped edits paired with gating yield more reliable gains. To our knowledge, this is the first work that uses bidirectional NLI-style consistency signals as an instance-wise router for conservative dialogue-claim rewriting. To structure our evaluation, we test three hypotheses. These hypotheses are stated as mechanism-driven expectations based on how the claim surface is consumed by both stages of the pipeline: IR uses it as a query, while FV uses it as the verifierās hypothesis. In particular, R1R_1āR3R_3 make lightweight surface-form edits designed to preserve propositional content, whereas R4R_4 makes explicit reference substitutions that can change downstream behaviour, motivating H3ās test of whether semantic routing improves E2E robustness. We denote the cumulative pipeline outputs as R1R_1āR4R_4, and R5R_5 as the one-shot decoder rewrite (§4). (H1) Lightweight surface normalisation (R1R_1āR3R_3) is largely neutral for both retrieval and verification. (H2) Resolving in-claim pronouns (R4R_4) can improve fact verification, and BiCon-Gate increases robustness by falling back when a rewrite is not semantically supported. (H3) In end-to-end fact-checking, semantic routing mitigates the tension between retrieval gains and verification robustness. Figure 2: Overview of the staged de-colloquialisation pipeline (R1R_1-R5R_5), BiCon-Gate routing, and evaluation protocols (IR-only, FV-only, and E2E). BiCon-Gate combines bidirectional NLI entailment/contradiction signals (ent/cnt) with embedding cosine similarity (sim) and accepts a rewrite only when the gate score exceeds a task-specific threshold (Ļ); otherwise it falls back to R0R_0. 2 Related Work 2.1 Colloquial noise in retrieval and verification Informal dialogue phenomena (e.g., contractions, ellipsis, and pronominal references) can blur span boundaries and entity mentions, making retrieval and verification brittle. In dialogue fact-checking, making a claim more explicit can improve evidence retrieval, but incorrect rewrites (e.g., wrong antecedent substitutions) can introduce semantic drift that harms downstream verification. This yields an IRāFV trade-off: aggressive de-colloquialisation may raise recall, while conservative, semantics-preserving rewriting is needed to maintain verification accuracy. On the retrieval side, Kim et al. (2021) show that converting formal FEVER Thorne et al. (2018) claims into colloquial variants led to a large drop in document recall (90.00% ā 72.20%), highlighting how sensitive rankers are to informal phrasing. In dialogue, Gupta et al. (2022) similarly identify colloquialityāespecially pronouns and underspecified referencesāas an obstacle to fact-checking. A complementary line of work makes inputs more self-contained before retrieval. Sundriyal et al. (2023) show that normalisation and de-contextualisation improve retrieval on noisy social media text, and de-contextualising pronoun-heavy claims improves evidence finding in multi-document settings Deng et al. (2024). Related approaches rewrite incomplete utterances to mitigate ellipsis-driven failures in dialogue Li et al. (2023); Guo et al. (2024); Cao et al. (2024). However, in fact-checking the rewritten claim also serves as the verifierās hypothesis; thus, retrieval-oriented rewrites can hurt E2E performance if they drift semantically. This motivates instance-wise safeguards that preserve meaning while still improving retrieval robustness. Beyond rewriting, a related line of work improves fact-checking by explicitly aligning retrieval with verification objectives. Feedback-based Evidence Retriever (FER) feeds verifier signals back into the retriever, aligning document selection with end-task performance Zhang et al. (2023). Our work follows a similar principleāusing downstream-oriented signalsābut applies it to deciding when to trust a rewrite, rather than which documents to retrieve. 2.2 Pronominal Coreference in Conversation Recent mention-based coreference resolution systems such as Maverick Martinelli et al. (2024) achieve strong accuracy with efficient span-scoring, but they are trained to link textual mentions and are not designed for discourse-level references or informal dialogue. We focus specifically on in-claim pronominal anaphora because the claim text is used directly as (i) the retrieval query and (i) the verifierās hypothesis; unresolved pronouns therefore degrade both stages. Moreover, in-claim pronouns can often be resolved from the preceding dialogue without rewriting the entire conversation, making them a targeted and controllable intervention. LLMs demonstrate impressive referential reasoning, but converting that ability into stable gains on standard coreference benchmarksāespecially for pronominal mentions in long contextsāremains challenging Gan et al. (2024); Manikantan et al. (2025). Recent analyses suggest that remaining errors concentrate on pronominal mentions and overlapping mention structures in long contexts, making antecedent substitutions brittle even for strong LLMs. In dialogue fact-checking, Chamoun et al. (2023) find that naively substituting in-claim pronouns improves document recall (56.85% ā 67.00%) but slightly harms sentence-level evidence selection (54.19% ā 53.86%) and fact-verification accuracy (44.06% ā 42.71%). These findings suggest that imperfect coreference rewrites can propagate errors downstream. Their error analysis attributes this drop to incorrect coreference links that steer evidence selection toward similar sentences in the wrong document and can even change a claimās label with respect to the gold evidence, motivating our instance-wise semantic gate. Prior work typically studies individual rewrite phenomena in isolation and reports only end-to-end outcomes, which conflates retrieval and verification effects under a fixed backbone. It also rarely provides an explicit, conservative acceptance criterion to prevent semantic drift when applying coreference-based rewrites. We address these gaps by (i) measuring the component-wise impact of cumulative rewrites on the same instances using IR-only, FV-only, and E2E protocols, and (i) introducing an instance-wise semantic gate that adopts a rewrite only when it is consistent with the original claim. 2.3 NLI-based factual consistency and semantic gating A recurring theme in factual-consistency research is to compare a candidate text against a reference via NLI-style containment and semantic similarity. SummaC operationalises consistency by checking whether each summary sentence is entailed by, and not contradicted by, the source document Laban et al. (2022). AlignScore similarly defines alignment as being fully supported by the reference without contradiction Zha et al. (2023). MENLI further shows that NLI-based entailment and contradiction signals are more robust indicators of correctness than sentence-similarity scores alone Chen and Eger (2023). We adopt this principle in a different setting: rather than using NLI scores purely for evaluation, we use them operationally to decide whether a rewrite should be applied for a given instance. This design directly targets the IRāFV trade-off described above by accepting rewrites only when they are semantically consistent with the original claim. Motivated by NLI-based factual-consistency work (e.g., SummaC, AlignScore, MENLI), BiCon-Gate uses bidirectional entailment/contradiction signals together with semantic similarity to conservatively accept rewrites and otherwise fall back to the original claim. 3 Task & Data 3.1 Task Definition Given a multi-turn dialogue context C and a response claim R, our system retrieves evidence from a Wikipedia snapshot aligned with DialFact (see § 5.2 for details) and predicts a label in SUPPORTS, REFUTES, NEI. Because the claim surface is used both as the retrieval query and as the verifier hypothesis, rewriting can have opposite effects on IR and FV (e.g., higher recall but increased semantic drift). We denote the original claim R as R0R_0 and the variants produced by our de-colloquialisation pipeline as R1R_1āR5R_5. To isolate rewriting effects, we keep the retriever and verifier fixed and vary only the claim surface presented to them. We report results under three evaluation protocols: (i) IR-only, where the selected surface is used as the retrieval query; (i) FV-only, where the verifier receives gold evidence and the selected surface is used as the hypothesis; and (i) E2E, where both retrieval and verification consume the selected surface. Figure 2 summarises our pipeline and where rewriting interacts with retrieval and verification. Starting from the original claim R0R_0, we generate progressively de-colloquialised variants R1R_1āR4R_4 through lightweight surface normalisation (de-contraction, punctuation restoration, true-casing) followed by scoped in-claim coreference rewriting. We also construct an alternative one-shot LLM rewrite R5R_5. BiCon-Gate then routes between R0R_0 and a candidate rewrite (mainly R4R_4; R5R_5 in an ablation) using semantic consistency signals derived from bidirectional NLI and embedding similarity; implementation details are provided in § 4. 3.2 Data We use the official validation and test splits of DialFact Gupta et al. (2022), whose dialogues are categorised into factual or personal subsets (Table A1). DialFact features colloquial multi-turn contexts, context-dependent claims, and a large proportion of NEI cases. We focus on DialFact because it combines multi-turn dialogue with a subset of sentence-level gold evidence, enabling retrieval-focused evaluation while still supporting E2E fact verification. Sentence-level gold evidence annotations are provided only for the factual subset; accordingly, we compute IR metrics on that subset, while FV-only and E2E metrics are reported on the full test split. As shown in Table A1, 39.7% of validation claims and 44.3% of test claims contain at least one in-claim pronoun (42.1% overall). This prevalence motivates our focus on in-claim anaphora: unresolved pronominal references are common and directly affect both retrieval queries and verifier hypotheses. 4 Methodology Our methodology consists of a staged de-colloquialisation pipeline with a semantic consistency gate. We describe the rewriting steps (R1R_1-R5R_5) and then the consistency gate and evaluation protocols. 4.1 De-colloquialisation We adopt a cumulative pipeline that reduces colloquial noise and then resolves in-claim anaphora. Let R0R_0 be the original response; R1R_1āR5R_5 are derived as follows. R1R_1 De-contraction. We first restore missing apostrophes with conservative, regex-gated rules (e.g., im/ive/ill ā Iām/Iāve/Iāl), then expand contractions (e.g., itās ā it is) while protecting dotted acronyms (e.g., Ph.D., U.S.), honorifics, and URLs via placeholders. This stabilises token boundaries and negation cues. R2R_2 Turn-preserving punctuation restoration. We apply a multilingual punctuation model Vandeghinste and Guhr (2024) to each turn and insert only predicted commas and sentence-final marks, leaving existing tokens and punctuation intact. For non-question turns without sentence-final punctuation, we append a period. This improves sentence and NP boundaries for downstream resolution. R3R_3 True-casing. Using a BERT-based masked LM (bert-base-cased Devlin et al. (2019)), we true-case sentence onsets and proper names with a margin rule: for each alphabetic token, we compare the MLM log-probabilities of upper- vs. lower-initial variants and flip only when the margin exceeds a threshold. We keep spelling, whitespace, and punctuation intact, modifying only token-initial characters to aid coreference resolution without semantic change of the underlying text. R4R_4 Scoped coreference rewriting (with gate). Scope. We target only in-scope pronominal anaphora whose antecedents are present in C (e.g., he/she/they/you), excluding deictic (this/that) and expletive it (e.g., it is raining). Detect & propose. On R3R_3 we detect pronominal anchors (POS patterns) and let Maverick Martinelli et al. (2024) propose up to 10 candidate antecedent NPs from the true-cased context. Select & rewrite. We rank up to 10 candidates with an instruction-tuned LLM (Llama-3.1-8B-Instruct Grattafiori et al. (2024)) and substitute the selected antecedent span for the pronoun mention in R3R_3 to obtain a candidate R4R_4. Instruction-following LLMs have been shown to work effectively as controllable decision modules for claim matching in automated fact-checking, motivating our use of an instruction-tuned LLM as a lightweight selector over a small candidate referent set Pisarevskaya and Zubiaga (2025). We then apply BiCon-Gate (§4.2) as an instance-wise router. R5R_5 Decoder-based one-shot reformulation. A single Qwen2.5-14B-Instruct Qwen et al. (2025) prompt attempts all editing steps (R1R_1-R4R_4) in one shot given C+R0\C+R_0\, producing R5R_5. We use Qwen2.5-14B-Instruct as a representative strong instruction-tuned decoder to instantiate a competitive one-shot baseline. The rewrite is prompted as a constrained editing task (Appendix Table A5); the model is instructed to apply only surface normalisation and unambiguous pronoun substitution based on the provided context, without adding or changing meanings. For ablation experiments (§5.5) we optionally apply the same semantic gate as a router between R5R_5 and R0R_0: for each instance, if si(5)ā„Ļs_i^(5)ā„Ļ we use R5R_5, otherwise we fall back to R0R_0. Model identifiers for all third-party components used in our pipeline are listed in the Appendix Table A4. 4.2 Consistency Gate: BiCon-Gate For an instance i with context CiC_i and original response R0,iR_0,i, a rewrite Rk,iR_k,i (kā4,5kā\4,5\) is scored by ei e_i =minpent(Ci+R0,iāRk,i), = \p^ent (C_i+R_0,i R_k,i ), pent(Rk,iāCi+R0,i), p^ent (R_k,i C_i+R_0,i ) \, (1) ci c_i =maxpctr(Ci+R0,iāRk,i), = \p^ctr (C_i+R_0,i R_k,i ), pctr(Rk,iāCi+R0,i), p^ctr (R_k,i C_i+R_0,i ) \, (2) simi _i =cosā”(Ļā(R0,i),Ļā(Rk,i)). = (Ļ(R_0,i),\,Ļ(R_k,i) ). (3) Here pentā(ā )p^ent(Ā·) and pctrā(ā )p^ctr(Ā·) denote calibrated NLI probabilities for the entailment and contradiction classes, respectively, so that ei,ciā[0,1]e_i,c_iā[0,1]. ĻĻ is a sentence encoder and simisim_i is the cosine similarity between the original and rewritten claim. We take the minimum bidirectional entailment to require mutual semantic containment, and the maximum bidirectional contradiction to penalise any directional inconsistency. We use a three-way NLI model and score both directions by swapping the premise and hypothesis between Ci+R0,iC_i+R_0,i and Rk,iR_k,i; we take pentp^ent and pctrp^ctr from the entailment and contradiction softmax probabilities. We calibrate the NLI probabilities via temperature scaling. Concretely, we rescale NLI logits z as /Tz/T before softmax and learn a single scalar T on the DialFact validation split (used as a calibration set) by minimising NLL. We then apply the learned T=4.96T=4.96 to all bidirectional NLI scores at test time Xie et al. (2024); Guo et al. (2017). The gate score and decision are: si(k) s_i^(k) =αāei+βāsimi+γā(1āci), =α\,e_i+β\,sim_i+γ\,(1-c_i), (4) accepti(k) _i^(k) =ā[si(k)ā„Ļ], =I [s_i^(k)ā„Ļ ], (5) with non-negative weights α,β,γā„0α,β,γ\!ā„\!0 that sum to one (α+β+γ=1α+β+γ=1) and a threshold Ļā[0,1]Ļā[0,1]. For Rk,iR_k,i, if accepti(k)=1accept_i^(k)=1 we keep Rk,iR_k,i; otherwise we fall back to R0,iR_0,i. 4.3 Protocols and metrics We evaluate each claim surface RkR_k under three protocols that disentangle retrieval and verification. Unless stated otherwise, IR/E2E results focus on R0R_0āR4R_4, and R5R_5 is reported only as an FV-only ablation (§5.5); metrics are listed in Table A2. IR-only. The query is C+RkC+R_k and the retriever is fixed. We report document-level Recall@K and nDCG@K at the operating depths used in our retrieval stack (see §5.2). For gate tuning (§5.1) we additionally report micro Recall@K and 1āZHRā@āK1-ZHR@K. ZHRā@āKZHR@K denotes the zero-hit rate (the fraction of queries whose top-K retrieved set contains no gold evidence). FV-only. We isolate the effect of the claim surface RkR_k on fact verification by providing the verifier with gold evidence sentences as premises and using C+RkC+R_k as the hypothesis. We report fact-verification accuracy, macro-F1, and classwise F1 for SUPPORTS/REFUTES/NEI. End-to-End (E2E). IR and FV both consume RkR_k. Unlike FV-only, the verifier receives the top-1 passage retrieved by the IR component as evidence, so E2E reflects the combined effect of rewriting under retrieval noise. We report the same FV metrics as FV-only: accuracy, macro-F1, and classwise F1. 5 Experiments Our goal is to isolate the effect of claim de-colloquialisation and semantic routing, rather than to optimise the underlying retriever or verifier. Therefore, across all settings we keep the IR and FV backbones fixed and vary only the claim surface (R0R_0āR4R_4) and whether BiCon-Gate routes to a rewrite candidate; we additionally include a decoder-based one-shot rewrite (R5R_5) as an ablation. We evaluate on DialFact because it provides multi-turn dialogue contexts and evidence annotations that support IR-only, FV-only, and E2E protocols. 5.1 Gate parameters Following §2.3, we fix the gate weights to (α,β,γ)=(0.4,0.2,0.4)(α,β,γ)=(0.4,0.2,0.4) and tune task-specific thresholds on the validation split. We set α and γ symmetrically to weight entailment and non-contradiction equally, and down-weight cosine similarity (β) as a secondary signal because similarity alone can be high even under subtle semantic drift. To minimise hyperparameter tuning while keeping the gate interpretable, we keep these weights fixed across all experiments and tune only Ļ. We leave weight learning or re-tuning under larger distribution shifts to future work. Threshold for IR. Figure 3: IR gate sweep on the validation split. We report BM25 micro-Recall@180 and 1āZHRā@ā1801-ZHR@180 as a function of the R4R_4 gate threshold Ļ (where ZHRā@ā180ZHR@180 is the zero-hit rate). We use ĻIR=0.50 _IR=0.50 in the IR-only protocol. For IR, we tune the threshold using BM25 retrieval metrics on the validation split. For each Ļā[0.20,1.00]Ļā[0.20,1.00] in increments of 0.05, we build a gated claim surface exactly as in §2.3, run BM25 with depth K=180K=180, and compute micro-recall@180 and 1āZHR@1801-ZHR@180. Both curves exhibit a broad plateau for Ļā[0.4,0.6]Ļā[0.4,0.6] and start to degrade once Ļ>0.6Ļ>0.6. We set the IR gate threshold to ĻIR=0.50 _IR=0.50, which lies near the centre of this plateau and slightly maximises micro-recall@180 (Figure 3). Threshold for FV. Figure 4: FV-only gate sweep on the validation split (gold evidence). Macro-F1 and macro-Recall as a function of the R4R_4 gate threshold Ļ; we use ĻFV=0.70 _FV=0.70 in FV-only and E2E. For FV, we use gold evidence and sweep the gate threshold Ļā[0.20,1.00]Ļā[0.20,1.00] on the validation split. We use the gateSweep routine, which for each threshold Ļ constructs hypotheses with either R4R_4 or R0R_0 according to whether the BiCon-Gate score satisfies si(4)ā„Ļs_i^(4)ā„Ļ, and runs the verifier once on the resulting premiseāhypothesis pairs. As shown in Figure 4, macro-F1 and macro-Recall both peak at Ļā0.70Ļā 0.70 on valid split; performance improves as the gate starts accepting high-scoring R4R_4 rewrites, but once Ļ becomes too large we discard too many well-rewritten R4R_4 and the curves drop back towards the R0R_0 baseline. 5.2 IR-only BM25 (K=180) E5 dense (K=10) BGE-CE (K=1) Claim R@180 nDCG@180 R@10 nDCG@10 R@1 nDCG@1 R0R_0 73.89 48.37 70.47 61.98 53.51 57.01 R1R_1 73.22 47.96 69.95 61.56 53.14 56.59 R2R_2 73.22 47.96 69.94 61.60 53.18 56.64 R3R_3 73.22 47.96 69.96 61.67 53.21 56.66 R4R_4 76.54 50.94 73.18 64.45 55.53 59.15 +Gated 76.53 50.91 73.19 64.46 55.58 59.19 Table 1: IR-only retrieval results on the test (factual) split (macro Recall/nDCG, %) for claim variants R0R_0āR4R_4. R4R_4+Gated applies BiCon-Gate with ĻIR=0.50 _IR=0.50 and falls back to R0R_0 when the rewrite is rejected. We study how the claim surface (R0R_0āR4R_4) affects evidence retrieval. We index a 37 GB Wikipedia snapshot (2019-08-01) into 100-token passages with a stride of 50, matching the dump used to construct DialFact Gupta et al. (2022) and ensuring that all annotated gold evidence is, in principle, retrievable from our index. Since all passages have roughly the same length, we use a relatively weak length normalisation (b=0.4b=0.4) and set the term-frequency saturation parameter to k1=1.5k_1=1.5, following common BM25 settings in Pyserini/BEIR-style benchmarks Lin et al. (2021). Our retrieval pipeline has three steps. First, BM25 retrieves the top 300 passages per query. Based on an elbow analysis, we fix K=180K=180 as the working cut-off and pass these 180 candidates to the dense retriever. Second, we apply an E5-large bi-encoder Wang et al. (2024) to the K=180K=180 BM25 candidates, retrieve the top 20 semantically similar passages and keep the top 10. Finally, we apply a BGE-reranker-large cross-encoder Chen et al. (2024) to these 10 passages and select a single passage as the final evidence for IR-only evaluation. The original claim R0R_0 already retrieves most gold evidence (73.89% R@180 and 48.37% nDCG@180 with BM25), and the intermediate normalisation variants R1R_1āR3R_3 leave IR almost unchanged: all metrics remain within 0.7 points of R0R_0 at every stage. Relative to R0R_0, R4R_4 improves BM25 recall@180 from 73.89% to 76.54% and nDCG@180 from 48.37% to 50.94%; E5 dense retrieval recall@10 rises from 70.47% to 73.18% and nDCG@10 from 61.98% to 64.45%; and BGE cross-encoder recall@1 increases from 53.51% to 55.53% with a similar gain in nDCG@1. These gains indicate that making colloquial claims more self-contained by resolving in-scope pronouns helps all three retrieval stages focus on the correct entities. The gated variant uses the IR threshold ĻIR=0.50 _IR=0.50 chosen in §5.1. Figure 3 and Table 1 show that this threshold has almost no effect on retrieval: across BM25, E5, and BGE-CE, Gated stays within 0.05 points of R4R_4 on all recall and nDCG metrics. Overall, these IR-only results are consistent with H1: minimal surface normalisation (R1R_1āR3R_3) is retrieval-neutral across all three retrieval stages. In contrast, scoped pronoun rewriting (R4R_4) improves retrieval gains, and applying the IR-tuned gate (ĻIR=0.50 _IR=0.50) preserves these gains. 5.3 FV-only Claim Acc Macro-F1 F1(S) F1(R) F1(NEI) Ī 1 R0R_0 62.67 61.09 45.08 76.77 61.42 ā R1R_1 62.66 61.08 45.00 76.78 61.46 -0.01 R2R_2 62.71 61.16 45.33 76.69 61.47 +0.07 R3R_3 62.80 61.22 45.25 76.75 61.67 +0.13 R4R_4 63.15 61.85 47.59 76.20 61.77 +0.76 +Gated 63.77 62.93 50.59 76.12 62.10 +1.84 R5R_5 58.27 56.93 42.79 67.61 60.40 -4.16 +Gated 62.23 60.40 42.92 77.00 61.28 -0.69 Table 2: FV-only results on test split with gold evidence, comparing claim surfaces R0R_0-R5R_5 (Accuracy and macro/class-wise F1, %). R4R_4+Gated and R5R_5+Gated apply BiCon-Gate with ĻFV=0.70 _FV=0.70 (fallback to R0R_0), and Ī 1 is the macro-F1 change over R0R_0. We run FV-only with gold evidence and an NLI verifier He et al. (2023); Laurer et al. (2024); hypotheses use the last two dialogue turns by default, and R4R_4+Gated applies BiCon-Gate with the FV-tuned threshold ĻFV=0.70 _FV=0.70 (see §5.1). Table 2 summarises FV-only performance for R0R_0āR5R_5 and the gated variant of R4R_4. The original claim R0R_0 attains 61.09% macro-F1 and 62.67% accuracy. Cumulative claim de-colloquialisation (R1R_1āR3R_3) leaves FV almost unchanged when premises are gold evidence: macro-F1 stays within ±0.2± 0.2 points of R0R_0, and class-wise F1 for S/R/NEI shifts by at most 0.3 points. In contrast, scoped coreference rewriting R4R_4 yields a noticeable FV gain. R4R_4 improves macro-F1 from 61.09% to 61.85% (+0.76) and accuracy from 62.67% to 63.15%. Most of this gain comes from the SUPPORTS class: F1(S) rises from 45.08% to 47.59%, while F1(NEI) also slightly improves (61.42% ā 61.77%), and F1(R) decreases slightly (76.77% ā 76.20%). Making the claims more self-contained by replacing in-scope pronouns therefore helps the verifier distinguish supported facts without hurting NEI. This behaviour contrasts with observations by Chamoun et al. (2023), who report that generic coreference resolution and claim rewriting can improve document recall for conversational claims, but tend to degrade sentence-level evidence selection and claim verification on DialFact due to rewriting and resolution errors. In our setup, the scoped R4R_4 rewrites already improve FV over R0R_0. The gated variant attains the best FV-only performance with 62.93% macro-F1 (+1.84 over R0R_0) and 63.77% accuracy, while F1(S) increases further to 50.59%. For comparison, the decoder one-shot rewrite R5R_5 is markedly worse than R0R_0 (56.93% macro-F1), and even with BiCon-Gate it recovers only to 60.40% macro-F1, motivating the ablation discussion in §5.5. These trends are robust to context window size: Appendix Table A6 and Figures A1āA2 summarise Ī -F1 and class-wise Ī 1 over k. Importantly, we keep the FV threshold fixed at ĻFV=0.70 _FV=0.70 when varying the context window size k. The consistent gains of R4R_4+Gated across k in Appendix Table A6 therefore suggest that this operating point is reasonably stable under this moderate distribution shift in dialogue context length. Overall, these FV-only results are consistent with H1 and H2. With gold evidence provided as premises, light de-colloquialisation (R1R_1āR3R_3) remains largely verification-neutral, while scoped rewriting is the most reliable when filtered by BiCon-Gate. 5.4 End-to-End Claim Acc Macro-F1 F1(S) F1(R) F1(NEI) Ī 1 R0R_0 34.85 22.60 2.45 15.99 49.37 0.00 R1R_1 34.86 22.62 2.50 16.01 49.37 +0.02 R2R_2 34.96 22.71 2.55 16.13 49.45 +0.11 R3R_3 35.02 22.88 2.89 16.31 49.44 +0.28 R4R_4 34.10 20.73 3.03 9.57 49.60 ā1.87-1.87 +Gated 35.54 23.22 4.56 14.84 50.24 +0.62 Table 3: End-to-end results on the full test split using the top-1 retrieved passage as evidence (Accuracy and macro/class-wise F1, %). R4R_4+Gated uses BiCon-Gate with ĻFāV=0.70 _FV=0.70. For each claim surface RkR_k, we run the IR pipeline from §5.2 and take the BGE cross-encoderās top-1 passage as the premise. We then apply the NLI verifier from §5.3, constructing the hypothesis exactly as in FV-only (the last two turns in context followed by RkR_k). Since DialFact provides 1.32 gold evidence items per claim on average (test factual subset), top-1 is a conservative but practical choice that isolates claim-surface effects under retrieval noise; however, it may underestimate gains achievable with multi-evidence aggregation (Table 3). Table 3 shows that using R0R_0 yields 22.60 macro-F1 and 34.85% accuracy, substantially lower than in the FV-only setting with gold evidence, reflecting the difficulty of relying on a single retrieved passage. Light de-colloquialisation (R1R_1āR3R_3) has only a small impact on the full pipeline. Macro-F1 remains within 0.3 points of the R0R_0 (22.60 ā 22.88 for R3R_3), and accuracy increases only slightly from 34.85% to 35.02% for R3R_3. Class-wise F1 follows a similar pattern across R0-3R_0-3: NEI stays high and stable (ā49.3ā-ā49.5ā 49.3-49.5), while SUPPORTS and REFUTES remain much lower (F1(S) ā¤2.9⤠2.9, F1(R) ā16ā 16). This suggests that, once retrieval noise is introduced, the main bottleneck is deciding SUPPORTS vs. REFUTES under noisy top-1 evidence, rather than the exact surface form of the claim. Scoped coreference rewriting R4R_4 shows a different trade-off. Although R4R_4 improves document-level IR metrics (§5.2), it hurts E2E performance: macro-F1 drops from 22.60 to 20.73 and accuracy from 34.85 to 34.10. Most of this loss is due to REFUTES: F1(R) falls from 16.31 for R3R_3 to 9.57 for R4R_4, while F1(NEI) stays essentially unchanged (49.60%). Applying BiCon-Gate on top of R4R_4 largely recovers these losses: the gated variant (R4R_4+Gated) uses the FV-tuned threshold from §5.1 to decide, for each example, whether to use R4R_4 (if si(4)ā„ĻFVs_i^(4)ā„ _FV) or fall back to R0R_0 otherwise. We retrieve evidence using the same claim surface selected by the gate. This simple policy yields the best E2E performance: macro-F1 rises to 23.22 (+0.62 over R0R_0) and accuracy to 35.54% (+0.69). Class-wise, the gate boosts F1(S) to 4.56 and raises F1(R) back to 14.84. Taken together, these E2E results are consistent with H3: by selectively accepting rewrites and otherwise falling back to R0R_0, BiCon-Gate mitigates the rewrite-induced IRāFV trade-off and yields the best E2E performance. 5.5 Analysis This section provides three analyses of BiCon-Gate: (i) gate activation rates at the IR and FV/E2E operating points, clarifying how often the pipeline applies the scoped rewrite (R4R_4) rather than falling back to the original claim (R0R_0); (i) an ablation comparing the decoder-based one-shot rewrite (R5R_5) to our scoped coreference rewrite (R4R_4) to isolate the impact of aggressive reformulation on verification; and (i) qualitative error patterns that explain when rewriting helps and when it introduces semantic drift. Gate activation rates. We report the routing (activation) rates at the IR and FV/E2E operating points. At ĻIR=0.50 _IR=0.50, the gate routes 52.59% of factual test queries to R4R_4. At ĻFV=0.70 _FV=0.70, it routes 46.56% of test instances to R4R_4. In the R5R_5 ablation at the same FV threshold, the gate accepts only 0.92% of R5R_5 candidates (mean score 0.52), so R5R_5+Gated almost always falls back to R0R_0. BiCon-Gate acts as a semantic router, sending high-confidence scoped rewrites to R4R_4 and routing drift-prone paraphrases back to R0R_0. Decoder-based rewriting vs. scoped coreference Figure 5: Effect of gating on the decoder one-shot rewrite (R5R_5) in FV-only on the validation split. As Ļ increases (rejecting more R5R_5 instances), macro-F1 and macro-Recall approach the R0R_0 baseline, indicating that the gate mitigates noisy rewrites primarily via fallback to R0R_0. We compare the decoder one-shot rewrite R5R_5 (Table A5) against the scoped pronoun rewrite R4R_4, which performs a targeted antecedent substitution only when a pronoun is in-scope (i.e., its antecedent appears in the dialogue context C). Across context lengths (kā„2kā„ 2), R4R_4 consistently improves FV performance (best with R4R_4+Gated), whereas R5R_5 degrades FV (Appendix Table A6). Table 2 highlights the contrast clearly: while scoped rewriting R4R_4 improves FV (61.85% macro-F1) and its gated version further gains to 62.93%, the decoder rewrite R5R_5 collapses to 56.93% and its gated recovers only 60.40%. Consistent with the low acceptance rate for R5R_5 at ĻFV=0.70 _FV=0.70 (§5.5), this gap suggests that one-shot rewrites frequently introduce semantic drift, whereas targeted edits can be useful when selectively accepted. Class-wise analysis supports this interpretation. Appendix Figure A2 shows that R5R_5ās harm is driven primarily by a large drop in REFUTES across k, while R4R_4+Gated improves mostly via SUPPORTS gains. Further, Figure 5 shows that R5R_5 appears to improve mainly as the threshold increases and the system increasingly falls back toward R0R_0, whereas R4R_4 exhibits a clear interior optimum (Figure 4), indicating that a non-trivial subset of high-confidence scoped rewrites is genuinely beneficial. Qualitative Error Analysis. Appendix Table A3 provides representative R0R_0/R4R_4/R5R_5 triples; we summarise the recurring patterns here. R4R_4 is typically a minimal, semantics-preserving edit: it replaces an in-scope pronoun with its referent from the preceding text C (e.g., it ā Heartbreak Hotel, He ā Elvis), making the hypothesis more self-contained without broader paraphrasing. This aligns with the consistent SUPPORTS gains observed for R4R_4+Gated. In contrast, R5R_5 often performs surface-level normalisation (e.g., true-casing, punctuation, quoting) while leaving context-dependent pronouns unresolved; when it does paraphrase, it can blur cues that matter for contradiction, consistent with its REFUTES degradation (Appendix Figure A2). Gate scores reflect this difference: in these examples, R4R_4+Gated receives consistently higher scores than R5R_5+Gated (mean 0.79 vs. 0.54; Appendix Table A3), suggesting that BiCon-Gate favours verification-oriented de-contextualisation over cosmetic or potentially drifting rewrites. 6 Conclusions We study how colloquial, context-dependent dialogue claims degrade both retrieval and verification, and propose a conservative de-colloquialisation pipelineālightweight surface normalisation and scoped in-claim pronominal rewritingāpaired with BiCon-Gate, a bidirectional NLI-based router that accepts rewrites only when semantically supported and otherwise falls back to the original claim. On DialFact, surface-level normalisation is largely neutral, while scoped pronoun rewriting improves document retrieval across sparse, dense, and cross-encoder stages; yet in a strict top-1 E2E setting, ungated coreference rewrites can hurt verification despite better retrieval, highlighting a rewrite-induced IRāFV trade-off under retrieval noise. BiCon-Gate mitigates this trade-off by selectively accepting high-confidence rewrites, yielding the strongest FV-only gains (notably on SUPPORTS) and the best top-1 E2E performance among the claim variants, whereas one-shot decoder rewrites are less reliable and often harm verification. Overall, our results support H1āH3: minimal normalisation yields stable IR and FV with only marginal changes; scoped pronominal rewriting improves FV and BiCon-Gate further improves robustness by filtering rewrites that are not semantically supported; and semantic gating mitigates the IRāFV trade-off, making conservative rewriting a robust, controllable component in retrievalāverification pipelines. Limitations First, we evaluate only on the DialFact dataset using an English Wikipedia snapshot. The extent to which the observed gains transfer to other dialogue genres, longer contexts, or languages with different pronominal and morphological systems remains to be validated. Second, BiCon-Gate depends on multiple off-the-shelf components (a coreference resolver, a retriever, NLI model, and sentence encoders). Their biases, errors, and calibration properties can affect gate decisions; moreover, gate thresholds tuned on DialFact may require retuning under distribution shifts, and the resulting multi-model pipeline increases complexity, latency, and inference cost. Third, the scope of rewriting is intentionally narrowālimited to light normalisation and in-scope pronominal resolutionāleaving other colloquial phenomena (e.g., deixis, ellipsis, or filler words) unaddressed. In addition, incorrect antecedent substitutions can introduce errors that propagate to downstream retrieval and verification. Finally, our IR is passage-level and E2E setting uses a fixed retriever stack (BM25 ā E5 ā BGE-CE) and only the top-1 retrieved passage as evidence. Results may differ with multi-passage evidence aggregation, alternative retriever-verifier architectures, joint training, or recent evidence collections. Ethics Statement We use a publicly available dataset (DialFact) and an English Wikipedia snapshot, along with publicly released pretrained models, and we do not involve new data collection or human-subject studies. Our rewriting stepāespecially pronoun resolutionāmay introduce factual or attribution errors, which can lead to incorrect retrieval and verification outcomes. We therefore report aggregate benchmark results and caution against using outputs as definitive factual judgments without human oversight and transparent access to supporting evidence. References Z. Cao, P. Li, Y. Fan, and Q. Zhu (2024) Incomplete utterance rewriting with editing operation guidance and utterance augmentation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 7225ā7238. External Links: Link, Document Cited by: §1, §2.1. E. Chamoun, M. Saeidi, and A. Vlachos (2023) Automated fact-checking in dialogue: are specialized models needed?. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 16009ā16020. External Links: Link, Document Cited by: §1, §1, §2.2, §5.3. J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu (2024) BGE m3-embedding: multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. External Links: 2402.03216, Link Cited by: Table A4, §5.2. Y. Chen and S. Eger (2023) Menli: robust evaluation metrics from natural language inference. Transactions of the Association for Computational Linguistics 11, p. 804ā825. Cited by: §2.3. Z. Deng, M. Schlichtkrull, and A. Vlachos (2024) Document-level claim extraction and decontextualisation for fact-checking. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 11943ā11954. External Links: Link, Document Cited by: §1, §2.1. J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), p. 4171ā4186. Cited by: Table A4, §4.1. Y. Gan, M. Poesio, and J. Yu (2024) Assessing the capabilities of large language models in coreference: an evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M. Kan, V. Hoste, A. Lenci, S. Sakti, and N. Xue (Eds.), Torino, Italia, p. 1645ā1665. External Links: Link Cited by: §2.2. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. GuzmĆ”n, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Ćelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: Table A4, §4.1. C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger (2017) On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning-Volume 70, p. 1321ā1330. Cited by: §4.2. X. Guo, Q. Zhu, Q. Shi, X. Lin, L. Wang, D. DaqianLi, and Y. Chen (2024) Context-aware tracking and dynamic introduction for incomplete utterance rewriting in extended multi-turn dialogues. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 2138ā2148. External Links: Link, Document Cited by: §1, §2.1. P. Gupta, C. Wu, W. Liu, and C. Xiong (2022) DialFact: a benchmark for fact-checking in dialogue. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 3785ā3801. External Links: Link, Document Cited by: §1, §2.1, §3.2, §5.2. P. He, J. Gao, and W. Chen (2023) DeBERTaV3: improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. External Links: 2111.09543, Link Cited by: §5.3. B. Kim, H. Kim, S. Hong, and G. Kim (2021) How robust are fact checking systems on colloquial claims?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), Online, p. 1535ā1548. External Links: Link, Document Cited by: §1, §2.1. P. Laban, T. Schnabel, P. N. Bennett, and M. A. Hearst (2022) SummaC: re-visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics 10, p. 163ā177. Cited by: §2.3. M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers (2024) Building efficient universal classifiers with natural language inference. External Links: 2312.17543, Link Cited by: Table A4, §5.3. Z. Li, J. Li, H. Tang, K. Zhu, and R. Yang (2023) Incomplete utterance rewriting by a two-phase locate-and-fill regime. In Findings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 2731ā2745. External Links: Link, Document Cited by: §1, §2.1. J. Lin, X. Ma, S. Lin, J. Yang, R. Pradeep, and R. Nogueira (2021) Pyserini: an easy-to-use python toolkit to support replicable ir research with sparse and dense representations. External Links: 2102.10073, Link Cited by: §5.2. K. Manikantan, M. Tapaswi, V. Gandhi, and S. Toshniwal (2025) IdentifyMe: a challenging long-context mention resolution benchmark for LLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 768ā777. External Links: Link, Document, ISBN 979-8-89176-190-2 Cited by: §2.2. G. Martinelli, E. Barba, and R. Navigli (2024) Maverick: efficient and accurate coreference resolution defying recent trends. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 13380ā13394. External Links: Link, Document Cited by: Table A4, §2.2, §4.1. D. Pisarevskaya and A. Zubiaga (2025) Zero-shot and few-shot learning with instruction-following llms for claim matching in automated fact-checking. In Proceedings of the 31st International Conference on Computational Linguistics, p. 9721ā9736. Cited by: §4.1. Qwen, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025) Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: Table A4, §4.1. M. Sundriyal, T. Chakraborty, and P. Nakov (2023) From chaos to clarity: claim normalization to empower fact-checking. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 6594ā6609. External Links: Link, Document Cited by: §1, §2.1. J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal (2018) FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), M. Walker, H. Ji, and A. Stent (Eds.), New Orleans, Louisiana, p. 809ā819. External Links: Link, Document Cited by: §1, §2.1. J. Thorne and A. Vlachos (2018) Automated fact checking: task formulations, methods and future directions. In Proceedings of the 27th International Conference on Computational Linguistics, E. M. Bender, L. Derczynski, and P. Isabelle (Eds.), Santa Fe, New Mexico, USA, p. 3346ā3359. External Links: Link Cited by: §1. V. Vandeghinste and O. Guhr (2024) Fullstop: punctuation and segmentation prediction for dutch with transformers. Language Resources and Evaluation 58 (4), p. 1335ā1354. Cited by: Table A4, §4.1. L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei (2024) Multilingual e5 text embeddings: a technical report. External Links: 2402.05672, Link Cited by: Table A4, Table A4, §5.2. J. Xie, A. Chen, Y. Lee, E. Mitchell, and C. Finn (2024) Calibrating language models with adaptive temperature scaling. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 18128ā18138. Cited by: §4.2. Y. Zha, Y. Yang, R. Li, and Z. Hu (2023) AlignScore: evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 11328ā11348. Cited by: §2.3. H. Zhang, R. Zhang, J. Guo, M. de Rijke, Y. Fan, and X. Cheng (2023) From relevance to utility: evidence retrieval with feedback for fact verification. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 6373ā6384. External Links: Link, Document Cited by: §1, §2.1. Appendix A Appendix A.1 DialFact Dataset Statistics Split Type # S R NEI Ev.Item Turns(C) hasPronouns Valid factual 8,691 3,342 3,363 1,986 1.13 4.58 3,395 personal 1,745 0 0 1,745 1.01 4.36 747 total 10,436 3,342 3,363 3,731 1.11 4.54 4,142 Test factual 10,420 3,939 3,935 2,546 1.32 4.35 4,549 personal 1,389 0 0 1,389 1.10 3.80 679 total 11,809 3,939 3,935 3,935 1.29 4.28 5,228 Table A1: DialFact validation/test statistics and label counts. IR metrics are computed on the factual subset; FV/E2E use the full label distribution unless stated otherwise. "hasPronouns" counts samples containing at least one in-claim pronoun. A.2 Evaluation Metrics Protocol Metrics IR (gate tuning) micro-Recall@K, 1āZHRā@āK1-ZHR@K (BM25@180) IR (doc-level) macro-Recall@K, nDCG@K (BM25@180, E5@10, BGE-CE@1) FV-only Accuracy, macro-F1, classwise F1 (S/R/NEI) End-to-End Same FV metrics as FV-only, using IR-produced evidence Table A2: Evaluation metrics: IR (gate tuning) metrics are computed on the valid split to select ĻIR _IR, whereas IR (doc-level), FV-only, and E2E metrics are reported on the test split. ZHRZHR is the zero-hit rate; the proportion of queries for which no gold passage is retrieved in the top-K list. A.3 Pronoun Scope We rewrite only in-scope anaphoric pronouns whose antecedents appear in dialogue context C. We exclude deictic uses without textual antecedent and expletive it. ⢠Anaphora (coreferential, in-scope) A pronoun refers back to a preceding entity in C. Ex. "I have heard that Louis C.K. performed there in the past." / "I did not know that. He was pretty funny." (He ā Louis C.K.) ⢠Deictic (out-of-scope) Reference relies on extra-linguistic context, not on C. Ex. "This is amazing." (no textual antecedent) ⢠Expletive (out-of-scope) Non-referential "it" in weather/raising/extraposition. Ex. "It is raining." / "It seems that she left." / "It was John who called." / "It is important to exercise." A.4 Examples of Rewrites Example 1 (ID: 877___8--1) C Elvis is great, he was born in 1935. ⦠yeah he really brought rock and roll to the masses, thanks elvis. THANKS A LOT DUDE. R0R_0 Remember when heartbreak hotel came out? The public hated it at first! R4R_4 Remember when heartbreak hotel came out? The public hated heartbreak hotel at first! (gate(R4)=0.8178(R_4)=0.8178) R5R_5 Remember when āHeartbreak Hotelā came out? The public hated it at first! (gate(R5)=0.5445(R_5)=0.5445) Example 2 (ID: 445___6--2) C I love The Walking Dead, Iāve seen every episode since it premiered on October 31, 2010. ⦠Do you know what network I can find The Walking Dead on? R0R_0 It premiered on amc in the us on october 31, 2010, but you can probably find it on any basic cable channel like fox or hulu. R4R_4 The show premiered on amc in the us on October 31, 2010, but you can probably find The show on any basic cable channel like Fox or hulu. (gate(R4)=0.7619(R_4)=0.7619) R5R_5 It premiered on AMC in the US on October 31, 2010, but you can probably find it on any basic cable channel like Fox or Hulu. (gate(R5)=0.5368(R_5)=0.5368) Example 3 (ID: 617___2--0) C were you aware that the famous musician Elvisā middle name was Aaron? no, i wasnāt ! ⦠i donāt know much about Elvis. Where is he from? R0R_0 He was born in Tupelo Mississippi, but relocated to Memphis when he was 13. R4R_4 Elvis was born in Tupelo Mississippi, but relocated to Memphis when Elvis was 13. (gate(R4)=0.7891(R_4)=0.7891) R5R_5 He was born in Tupelo, Mississippi, but relocated to Memphis when he was thirteen. (gate(R5)=0.5490(R_5)=0.5490) Table A3: Representative DialFact instances used for qualitative analysis. Each block shows an excerpt of the dialogue context (C) and the resulting claim surface under the original claim (R0R_0), the scoped antecedent-substitution rewrite (R4R_4), and the decoder one-shot rewrite (R5R_5). Edited spans are marked with ā¦; BiCon-Gate scores for R4R_4 and R5R_5 are shown in parentheses. A.5 Model identifiers Component Model / identifier Punctuation restoration (R2R_2) oliverguhr/fullstop-punctuation-multilang-large Vandeghinste and Guhr (2024) True-casing (R3R_3) bert-base-cased Devlin et al. (2019) Coreference resolver (R4R_4) Maverick coreference resolver Martinelli et al. (2024) Antecedent selector (R4R_4) meta-llama/Llama-3.1-8B-Instruct Grattafiori et al. (2024) Decoder rewrite (R5R_5) Qwen/Qwen2.5-14B-Instruct Qwen et al. (2025) NLI verifier (FV-only) / BiCon-Gate (NLI scorer) MoritzLaurer/DeBERTa-v3-large-mnli-fever-anli-ling-wanli Laurer et al. (2024) BiCon-Gate embedding encoder intfloat/multilingual-e5-large Wang et al. (2024) Sparse retriever (IR) BM25 (Pyserini) Dense retriever (IR) E5-large Wang et al. (2024) Cross-encoder reranker (IR) BAAI/bge-reranker-large Chen et al. (2024) Table A4: Model identifiers for third-party components used in our de-colloquialisation pipeline, retrieval, and verification experiments. A.6 Prompts for Decoder-Based Rewriting Prompt for decoder-based rewrite R5R_5 System Follow the instructions exactly. Do not add or change facts. User You are an expert editor who rewrites informal, chatty utterances into well-formed declarative English without changing their meaning. You will receive: (i) Context, a list of previous dialogue turns; and (i) Response, the claim text to be normalised. Task: Rewrite Response into New_Response by applying only the following operations. (1) Add missing sentence-ending punctuation and fix spacing around punctuation. (2) Fix capitalisation at sentence starts and for proper nouns. (3) Insert missing apostrophes (e.g., dontā āt, cantā āt, imā ām). (4) Expand all contractions to full forms (e.g., isnātā not, arenātā not, wonātā not, wouldnātā not, Iāmā am, itāsā is, theyāreā are, donātā not, canātā ). (5) If a pronoun in Response (this/that/it/he/she/they/these/those) has a unique, clear antecedent in Context, replace it with that antecedent phrase; if ambiguous, leave it unchanged. (6) Do not add, remove, or correct any facts, numbers, names, or dates; preserve the claimās semantics exactly. (7) Output only the rewritten text as one or more sentences, with no explanations, lists, or markdown. The input is formatted as: Context (earliestā ): context_lines, Response: response_text, Output: (New_Response only; no explanations). Table A5: Prompt used with Qwen2.5-14B-Instruct to generate the decoder-based one-shot rewrite (R5R_5). The prompt constrains the model to apply only surface normalisation and unambiguous pronoun substitution based on the provided context, without changing claimās meaning. A.7 Additional FV-only Results: Context Window Sensitivity #Turns Claim FV-only (gold evidence) Acc Macro-F1 F1(S) F1(R) F1(NEI) Ī 1 0 R0R_0 65.31 63.56 46.47 80.37 63.84 ā R1R_1 65.20 63.43 46.17 80.27 63.84 -0.13 R2R_2 65.27 63.53 46.50 80.26 63.84 -0.03 R3R_3 65.36 63.58 46.34 80.35 64.04 +0.02 R4R_4 65.02 63.41 47.38 79.38 63.48 -0.15 +Gated (R4R_4) 65.90 64.83 51.15 79.06 64.28 +1.27 R5R_5 60.84 59.50 44.91 70.82 62.76 -4.06 +Gated (R5R_5) 65.31 63.56 46.47 80.38 63.84 +0.00 2 R0R_0 62.67 61.09 45.08 76.77 61.42 ā R1R_1 62.66 61.08 45.00 76.78 61.46 -0.01 R2R_2 62.71 61.16 45.33 76.69 61.47 +0.07 R3R_3 62.80 61.22 45.25 76.75 61.67 +0.13 R4R_4 63.15 61.85 47.59 76.20 61.77 +0.76 +Gated (R4R_4) 63.77 62.93 50.59 76.12 62.10 +1.84 R5R_5 58.27 56.93 42.79 67.61 60.40 -4.16 +Gated (R5R_5) 62.23 60.40 42.92 77.00 61.28 -0.69 4 R0R_0 62.90 61.73 49.67 74.51 61.12 ā R1R_1 62.94 61.76 49.51 74.64 61.13 +0.03 R2R_2 62.89 61.72 49.58 74.55 61.03 -0.01 R3R_3 62.86 61.66 49.46 74.38 60.66 -0.07 R4R_4 63.37 62.40 51.63 74.13 61.45 +0.67 +Gated (R4R_4) 63.91 63.29 54.07 73.98 61.82 +1.56 R5R_5 58.54 57.61 46.81 65.84 60.19 -4.12 +Gated (R5R_5) 62.83 61.58 48.66 74.86 61.22 -0.15 6 R0R_0 62.77 61.75 50.83 73.89 60.67 ā R1R_1 62.82 61.80 50.79 73.93 60.67 +0.05 R2R_2 62.81 61.78 50.85 73.85 60.64 +0.03 R3R_3 62.73 61.72 50.91 73.69 60.57 -0.03 R4R_4 63.33 62.52 53.04 73.57 60.95 +0.77 +Gated (R4R_4) 63.94 63.44 55.44 73.47 61.42 +1.69 R5R_5 58.21 57.46 47.73 65.07 59.58 -4.29 +Gated (R5R_5) 62.88 61.77 49.90 74.32 61.07 +0.02 8 R0R_0 62.69 61.71 50.98 73.87 60.56 ā R1R_1 62.78 61.81 50.98 73.82 60.62 +0.10 R2R_2 62.72 61.73 50.95 73.67 60.59 +0.02 R3R_3 62.63 61.66 51.00 73.53 60.45 -0.05 R4R_4 63.21 62.43 53.16 73.32 60.82 +0.72 +Gated (R4R_4) 63.64 63.18 55.40 73.23 60.91 +1.47 R5R_5 58.13 57.41 47.95 64.86 59.42 -4.30 +Gated (R5R_5) 62.61 61.57 50.08 73.80 60.83 -0.14 Table A6: FV-only context-window sensitivity (gold evidence). Results for kā0,2,4,6,8kā\0,2,4,6,8\ context turns, reporting Accuracy and macro/class-wise F1 (%) for each claim surface; Ī 1 denotes the macro-F1 change relative to R0R_0 at the same k. A.8 FV-only Macro-F1 Ī Heatmap Figure A1: FV-only heatmap of Ī -F1 across context turns kā0,2,4,6,8kā\0,2,4,6,8\, computed relative to R0R_0 at the same k (Table A6). Boxed rows highlight R4R_4 and R4R_4+Gated, and the main setting is k=2k=2. A.9 FV-only Class-wise Ī 1 Heatmaps Figure A2: Class-wise FV-only Ī 1 over context turns kā0,2,4,6,8kā\0,2,4,6,8\, computed from Table A6. Colors are clipped for readability; annotated values show the true deltas, and the dotted box marks the main setting (k=2k=2).