Paper deep dive
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
Jihyun Janice Ahn, Ryo Kamoi, Berk Atil, Renze Lou, WonWoo Kang, Heehyun Park, Sarkar Snigdha Sarathi Das, Zhuoyang Zou, Xiaoxin Lu, Yusen Zhang, Asfahan Shah, Ridwanul Hasan Tanvir, Lingxiao Zhao, Hongxi Huang, Vignesh Venkatesh, Dianjun Lin, Hamid Shah, Wentao Wang, Zhanpeng Song, Joshua Reed Bassin, Dax Patel, Ishan Appareddy Agrahar, Sahil Pardasani, Xin Dong, Fatemeh Rahbari, Benjamin David Rishel, Soochan Andrew Lee, Yuv Boghani, Ali B. AlNaseeb, Pranav Suby, Seokhyeon Bae, Shreya Buddharaju, Damien Kula, Soumyadeep Das, Hanyang Frank Liu, Faye Mo, Wenpeng Yin
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/26/2026, 1:36:42 AM
Summary
The paper identifies a 'know-act gap' in LLMs, where models can identify flawed inputs in discriminative modes but fail to do so during standard generative tasks. To address this, the authors introduce 'FaultyScience', a large-scale benchmark of ill-posed scientific questions, and propose 'DeIllusionLLM', a task-level autoregressive framework that uses self-distillation to unify discriminative judgment and generative reasoning.
Entities (5)
Relation Signals (3)
DeIllusionLLM â addresses â Know-Act Gap
confidence 95% ¡ To address this, we propose DeIllusionLLM, a task-level autoregressive framework that explicitly models this decision.
DeIllusionLLM â uses â Self-Distillation
confidence 95% ¡ Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone.
FaultyScience â evaluates â Know-Act Gap
confidence 90% ¡ we present a comprehensive analysis using FaultyScience... to evaluate whether models can autonomously decide whether to answer or object
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify such issues, yet fail to reflect this in standard generative responses. This reveals a fundamental know-act gap between discriminative recognition and generative behavior. Prior work largely characterizes this issue in narrow settings, such as math word problems or question answering, with limited focus on how to integrate these two modes. In this work, we present a comprehensive analysis using FaultyScience, a newly constructed large-scale, cross-disciplinary benchmark of faulty scientific questions. We show that the gap is pervasive and stems from token-level autoregression, which entangles task selection (validate vs. answer) with content generation, preventing discriminative knowledge from being utilized. To address this, we propose DeIllusionLLM, a task-level autoregressive framework that explicitly models this decision. Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone. Empirically, DeIllusionLLM substantially reduces answer-despite-error failures under natural prompting while maintaining general reasoning performance, demonstrating that self-distillation is an effective and scalable solution for bridging the discriminative-generative know-act gap
Tags
Links
- Source: https://arxiv.org/abs/2603.22619v1
- Canonical: https://arxiv.org/abs/2603.22619v1
Trouble viewing inline? Open PDF directly â
Full Text
53,904 characters extracted from source content.
Expand or collapse full text
Preprint Bridging the KnowâAct Gap via Task-Level Autoregressive Reasoning Jihyun Janice Ahn 1 ,Ryo Kamoi 1 ,Berk Atil 1 ,Renze Lou 1 ,WonWoo Kang 2 Heehyun Park 1 ,Sarkar Snigdha Sarathi Das 1 ,Zhuoyang Zou 1 ,Xiaoxin Lu 1 Yusen Zhang 3 ,Asfahan Shah 1 ,Ridwanul Hasan Tanvir 1 ,Lingxiao Zhao 1 Hongxi Huang 1 ,Vignesh Venkatesh 1 ,Dianjun Lin 1 ,Hamid Shah 1 Wentao Wang 1 ,Zhanpeng Song 1 ,Joshua Reed Bassin 1 ,Dax Patel 1 Ishan Appareddy Agrahar 1 , Sahil Pardasani 1 ,Xin Dong 1 ,Fatemeh Rahbari 1 Benjamin David Rishel 1 , Soochan Andrew Lee 1 ,Yuv Boghani 1 ,Ali B. AlNaseeb 1 Pranav Suby 1 , Seokhyeon Bae 1 ,Shreya Buddharaju 1 , Damien Kula 1 , Soumyadeep Das 1 ,Hanyang Frank Liu 3â ,Faye Mo 4 ,Wenpeng Yin 1 1 Penn State 2 UIUC 3 Columbia University 4 Episcopal Academy jfa5672,wenpeng@psu.edu Abstract LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify such issues, yet fail to reflect this in standard generative responses. This reveals a fundamental knowâact gap between discriminative recognition and generative behavior. Prior work largely characterizes this issue in narrow settings, such as math word problems or question answering, with limited focus on how to integrate these two modes. In this work, we present a comprehensive analysis using FaultyScience, a newly constructed large-scale, cross-disciplinary bench- mark of faulty scientific questions. We show that the gap is pervasive and stems from token-level autoregression, which entangles task selection (validate vs. answer) with content generation, preventing discriminative knowledge from being utilized. To address this, we proposeDeIllusionLLM, a task-level autoregressive framework that explicitly models this decision. Through self-distillation, the model unifies discriminative judgment and generative reasoning within a single backbone. Empirically,DeIllusionLLM substantially reduces answer-despite-error failures under natural prompt- ing while maintaining general reasoning performance, demonstrating that self-distillation is an effective and scalable solution for bridging the discriminativeâgenerative knowâact gap. 1 1 Introduction The deployment environments of AI systems, particularly LLMs, are inherently unpre- dictable, involving varying conditions, constraints, and potential adversarial inputs. As AI becomes increasingly pervasive, ensuring its trustworthiness under such imperfect and conflicting scenarios is critical. A key question is how these systems behave when the input itself is flawed. Humans, for instance, typically validate a premise before proceeding, challenging inconsistencies or requesting missing information when necessary (Grice, 1975; Johnson-Laird, 2010; Byrne et al., 2019; Stalnaker, 2002). Do LLMs exhibit similar behavior? Two empirical phenomena motivate this work. First, as illustrated in Table 1, even top- performing LLMs often produce a direct answer under standard generative mode without â Worked on this project when he was a student at Conestoga High School. 1 Data & Code available: https://github.com/Janice-ahn/Validate-Before-You-Answer 1 arXiv:2603.22619v1 [cs.AI] 23 Mar 2026 Preprint QuestionBalanced binary data: accuracy 95%, prec. 90%, recall 85%, compute F1. Ask GPT4 to respond (w/o hint) F1 = 2PR P+R = 2Ă0.90Ă0.85 0.90+0.85 = 0.874 Ask GPT4 if the question is problematic On a balanced dataset, these metrics (accuracy = 95%, precision = 90%, recall = 85%) cannot all hold simultaneously. At least one is invalid, so any F1 derived is unreliable. Table 1: Example where GPT4 fails in generative mode without hints but recognize the issues in discriminative mode. recognizing that the input itself is flawed. Second, when explicitly asked to judge whether the same input is problematic, i.e., under a discriminative mode, the model can often correctly identify the issue. This contrast reveals a striking gap between what LLMs know and what they actually do during ordinary response generation. This observation raises two research questions (Q): ⢠Q 1 : How prevalent is the discrepancy between discriminative fault recognition and generative answering across mainstream LLMs, and why can models recog- nize an issue when explicitly asked, yet fail to leverage that knowledge under normal generation? ⢠Q 2 : How can we incorporate LLMsâ discriminative knowledge into the generative process so that the model can produce the most appropriate response without human hints, explicit fault-handling instructions, or external policies? To answerQ 1 , we first introduceFaultyScience, a large-scale benchmark of over 15K faulty questions, primarily from science and engineering, spanning eight disciplines including physics, biology, earth science, mathematics, computer science, medicine, social science and more. The dataset covers diverse question formats (e.g., open-ended, multiple-choice) and error types (e.g., missing information, incorrect premises, nonsensical content). Using this benchmark, we systematically compare LLM behavior under three prompting settings: natural generative prompting without hints, generative prompting with a subtle warning, and explicit discriminative evaluation. We find that while models such as Llama (Touvron et al., 2023), Qwen (Qwen et al., 2025), and GPT4 (OpenAI et al., 2024) achieve around 90% accuracy in discriminative fault recognition, their performance under the most natural no- hint generative setting drops dramaticallyâto slightly above 10% for open-source models and only 34.9% for GPT-4. Our analysis further reveals several recurring failure patterns: i) models often prioritize surface-level mathematical computation over basic validity checks, i) answer only the seemingly solvable parts of a flawed question, or i) default to producing a helpful-looking response instead of acknowledging that the problem itself is invalid. These findings motivate our investigation ofQ 2 . We suspect that this failure cannot be fully addressed within standard token-level autoregressive generation, because the miss- ing behavior is fundamentally task-level: the model should first validate the input, then decide whether to answer, object, or request clarification. Based on this view, we propose DeIllusionLLM, a model that distills both discriminative and generative behaviors into a unified LLM through task-level autoregressive training. Distilling from a strong teacher such as GPT improves an open-source backbone such as Qwen, but more strikingly, we find that self-distillation from Qwenâs own discriminative and generative outputs also substan- tially improves its behavior onFaultyScience, while largely preserving performance on standard reasoning benchmarks such as BBH (Suzgun et al., 2023) and GPQA (Rein et al., 2024). This suggests that self-distillation across reasoning modes is a practical and effective way to mitigate unconditional answering on adversarial prompts. Overall, this work makes three main contributions. First, we provide the first systematic large-scale study of this trustworthiness-critical failure mode through theFaultyScience benchmark and extensive empirical analysis. Second, we proposeDeIllusionLLM, a task- level autoregressive training framework that enables LLMs to reason at the level of which task to perform before generating content. Third, we show that mode-aware self-distillation can substantially reduce this failure mode without sacrificing general reasoning capability. 2 Preprint 2 Related Work Unanswerable and ambiguous but valid questions. Prior work has extensively studied how AI systems handle questions that are valid in form but unanswerable due to missing evidence or ambiguity. For example, Sun et al. (2018) propose U-Net, a unified model that jointly predicts answer spans and whether a question is unanswerable. Benchmarks such as SQuAD 2.0 (Rajpurkar et al., 2018) and subsequent analyses (Sulem et al., 2021) evaluate whether models can abstain when no answer is supported by the given context. Other work (Min et al., 2020; Elgohary et al., 2019) addresses ambiguity, where questions admit multiple valid interpretations, requiring models to disambiguate or enumerate possible answers. A key distinction from our work is that these settings assume the question itself is well-posed. The challenge lies in missing information or multiple interpretations, emphasizing answer correctness rather than the modelâs ability to verify the validity of the question itself. LLMs on questions with flawed premises. More recent studies investigate how LLMs behave when faced with questions containing incorrect or misleading premises. These works show that LLMs often generate plausible but incorrect answers instead of rejecting invalid inputs (Hu et al., 2023). Similarly, benchmarks involving unsolvable mathematical or logical problems demonstrate that models tend to fabricate solutions rather than identify ill-posedness (Sun et al., 2024; Rahman et al., 2025). However, these studies primarily characterize failure under standard generative prompting. They do not examine whether models internally possess the knowledge to recognize such flaws (e.g., under a discriminative setting), nor how this knowledge can be effectively utilized during generation. Discriminative detection of problematic inputs. Another line of work treats problem validity as a standalone discriminative task. For instance, Lavi et al. (2025) show that answerability can be detected via linear directions in embedding space, suggesting that LLMs implicitly encode solvability signals. Other studies focus on ambiguity detection and uncertainty calibration, showing that models can identify underspecified queries and adjust confidence (Shi et al., 2025). Related work also uses LLM representations to assess semantic validity in structured evaluation settings (Milano et al., 2025). While these approaches demonstrate that LLMs can detect problematic inputs, they treat validation as an external classification task. They do not integrate this capability into the modelâs reasoning process during generation. Handling jailbreak attacks. The most related work is (Ding et al., 2025a), which studies the discriminativeâgenerative gap in the context of jailbreak prompts. Their approach introduces explicit guidance signals (e.g., detection outputs or injected states) to control generation, operating in a with-hint regime. In contrast, we address a broader knowâact gap and enable models to act on their knowledge under natural prompting without external hints. Methodologically, they condition on auxiliary signals, while we model task selection (validate vs. answer) as an intrinsic autoregressive decision, unifying discrimination and generation. The paradigm in (Ding et al., 2025b) is generateâcritiqueârefine, which depends on externalized feedback signals, while ours is decideâact within a single autoregressive trajectory, eliminating the need for post-hoc correction. Novelty of our work. Our work differs from prior research in three key aspects. First, rather than evaluating whether LLMs can detect problematic questions under explicit prompts or external probes, we focus on a practical failure mode: models often fail to invoke this capability under standard generative prompting without hint. Second, existing datasets typically focus on domain-specific cases, such as math word problems or jailbreak attacks. We introduce FAULTYSCIENCE, a large-scale, cross-disciplinary benchmark of naturally occurring faulty scientific questions, designed to evaluate whether models can autonomously decide whether to answer or object under realistic no-hint conditions. Third, beyond analysis, we propose a solution. Our method, DEILLUSIONLLM, introduces task- 3 Preprint level autoregressive reasoning and a scalable self-distillation paradigm that integrates discriminative validation and generative answering within a unified model. 3 FaultyScience Dataset Construction We constructFaultyScience, a large-scale benchmark of flawed scientific questions that are plausible at first glance but ultimately ill-defined or unsolvable. The design goal is to evaluate whether LLMs can self-initiate doubt and refuse, diagnose, or request clarification under realistic prompting, rather than succeeding only when explicitly warned that an input may be invalid. Data Collection. We curate candidate questions from a pool ofâź100 undergraduate & graduate students spanning multiple scientific disciplines from some courses the author team taught. Contributors are instructed to provide science questions that contain natural errors, e.g., incorrect premises, contradictory constraints, missing conditions, or scientifically impossible setups, rather than artificially constructed adversarial prompts. This collection strategy yields diverse error patterns that better reflect how ill-posed questions arise in real educational and scientific settings (e.g., misunderstanding of concepts, unit/quantity mismatches, ambiguous variable definitions, or invalid multiple-choice options). Manual Filtering. Raw submissions include both non-scientific items and trivial failures that can be detected by superficial cues. To ensure the benchmark emphasizes meaningful reasoning and validation, we apply a multi-stage manual filtering pipeline: i) Non-science removal. We discard prompts that are not scientific or technical questions (e.g., general chit- chat, purely opinion-based prompts, or requests unrelated to scientific problem solving). i) Trivial-fault removal. We remove questions whose faultiness is obvious without substantive reasoning (e.g., blatant nonsense strings, malformed text, or errors detectable from a single shallow pattern). i) Well-posedness screening. We prioritize questions that appear solvable on the surface yet contain a subtle fatal flaw such as: (i) incorrect scientific premise, (i) internal logical contradiction, (i) missing essential conditions, (iv) invalid answer choices, or (v) impossible scenario under standard scientific knowledge. Count Ratio (%) Discipline Physical Science301419.07 Biological Science277617.57 Earth Science240415.21 Mathematics212913.47 Social Science207013.10 Engineering13288.41 Chemical Science11827.48 Computer Science8965.67 Others90.02 Question Type Open-ended1428790.38 Multiple-choice14369.80 Fill-in-the-blank10.01 True/False820.52 Yes/No20.01 Error Type Consistency Error567335.89 Nonsensical Content434927.51 Missing Information396025.05 Incorrect Information182611.55 Table 2: Statistics of FaultyScience. Annotation Protocol. After filtering, each remaining question is independently anno- tated by two trained annotators. Annotators judge whether the prompt constitutes a hard faulty problemâa question that is scientifically meaningful in style but fundamentally un- solvable or ill-posed as written. Annotations are performed independently to reduce con- firmation bias. Using the two judgments, we partition data into three groups: ⢠Full-false set. Questions labeled as hard faulty by both annotators. â˘Semi-false set. Questions labeled as hard faulty by exactly one anno- tator; these include borderline cases and broaden the diversity of failure modes. â˘Discarded set. Questions deemed unsuitable by both annotators (e.g., too ambiguous to adjudicate, not sci- entific after all, or too trivial). Following this protocol, we obtain 2,610 full- false questions, 13,198 semi-false questions, and 2,183 discarded questions. 4 Preprint Data Diversity & Statistics. Table 2 presents dataset statistics along three dimensions: discipline, question type, and error type.FaultyScienceexhibits strong diversity across scientific domains as well as a wide range of error categories. Training/Dev/Test Data Split.We construct a balancedTestset by uniformly sampling 125 questions from each of 8 disciplines in the full-false subset (1,000 total), and aDevset with 50 per discipline. The remaining data (including the semi-false subset and the remaining full-false instances) forms the training pool, resulting in 14,408 training problems. This split enables us not only to evaluate LLMs but also to develop and train models, in contrast to prior work, e.g., (Rahman et al., 2025), that focuses solely on evaluation benchmarks. 4 AddressingQ 1 : Analysis of the DiscriminativeâGenerative Discrepancy through FaultyScience In this section, we analyse the performance and behavior of existing mainstream LLMs onFaultyScience, try to find some clues and insight to better understand LLMs, and also provide insight to endorse our solution design in Section 5. Particularly, we want to answer the following two questions (Q 1.1 &Q 1.2 ): Q 1.1 : How differently do existing LLMs behave under three prompting conditions: (1) no hint, (2) with a hint, and (3) explicit discriminative evaluation? Therefore, we explore three types of prompts: ⢠P Gen âhint : we feed the question into the LLM without any hint that the problem might have issues and you should check before solving. This is the most natural LLM-asking format we want to explore and it can show the true reasoning capability of LLM. For example Please answer the following question: [QUESTION] ⢠P Gen +hint : in this case, we provide a subtle hint to above P Gen âhint : Please answer the following question: [QUESTION]. If you cannot answer due to issues in the question, explicitly state what prevents a valid answer in your response ⢠P Dis : Differing from above two generative prompt which are problem-solving oriented, This is a discriminative prompt to explicit ask LLM if the input question has mistake or not, such as Please analyze the following question: [QUESTION]. Tell us if it has an issue (Yes or No) that prevents a valid answer and explain why. We compare a top-performing closed-source LLM, GPT4, and three representative open- source LLMs: Mixtral-8x7B, Llama3.3-70B and Qwen2.5-72B. ModelP Gen âhint P Gen +hint P Dis Mixtral-8x7B10.227.875.2 Llama3.3 70B14.024.590.8 Qwen2.5-72B10.840.195.0 GPT434.948.688.1 Table 3: Answer toQ 1.1 âperformance (%) of existing LLMs in recognizing the issues in the questions ofFaultyScienceunder differ- ent prompt style. Table 3 summarized the results. Here we found that i) without any hints (even though we hope any AI can recognize by default) this challenge all LLMs with GPT4 only one who passes 30%. Adding hints in prompt can im- prove the potential of recognizing questionsâ issues with big margin (P Gen âhint vs.P Gen +hint ). i) Interestingly, open-source models such as LLaMA and Qwen exhibit comparable or even stronger discriminative performance than GPT-4 in identifying issues inFaultyScience. This suggests that these models can serve as self-contained knowledge sources during generation, rather than relying on stronger 5 Preprint closed-source teachers, which directly inspires our model design in Section 5. i) both generative prompts, no matter it has hint (P Gen +hint ) or not (P Gen âhint ), are much behind with large margins behind the explicitly ask the LLM to discminatively check if the question has the issues (P Dis ). This indicates that LLMs may has the capbility of recognizing the latent issues but it may not able to make use of it in the default autoregressive generation process. This leads to our next questionQ 1.2 . Q 1.2 : Why cannot LLMsâ reasoning recognize the faultiness of those questions? To investigate this phenomenon, we have manually compared 646 pairs of GPT4 responses underP Gen âhint andP Dis , where model find the question problematic based onP Dis but failed to reflect that when answer underP Gen âhint . Three main patterns on this symptoms are found: â˘Arithmetic-over-World-Knowledge Bias (ratio:âź35%): The LLMs prioritize performing surface-level arithmetic operations over validating the problemâs feasibility under basic world knowledge constraints (e.g., physical limits, inventory bounds, or domain rules). Example below: Q Consider an amalgamation of sum number of decks of cards initially with 543 club cards. Jake has 561 club cards in their hand from the deck and then returns 111 to the deck. How many club cards are in the deck after these actions? P Dis The question incorrectly posits a situation where the number of club cards Jake holds is more than the total available within a set, then leading to an illogical count. It is a logical contradiction as a standard deck of cards only contains 13 club cards. P Gen âhint 93 Many LLMs adopt a computation-first heuristic: they interpret the question as a routine numerical word problem and directly execute the arithmetic operation 543-561+111=93, instead of validating the premise against basic world knowledge constraints (e.g., card-count limits). ⢠Partial-Task Satisfaction (Answering the answerable component) (ratio:âź25%): The question may be invalid globally (missing info, inconsistent assumptions), but it contains some valid local structure, e.g., a definable term, a small factual claim, a solvable sub- expression, an easy âmini-problem.â Generative decoding then takes the path of least resistance: solve the âeasy sub-partâ and ignore the invalid piece. Instead of rejecting ill- posed problems, LLMs often âsatisficeâ by solving an internally identified solvable fragment and treating it as the intended task. Example below: Q Quenching is known to be the same as Tempering. Explain how Enthalpy applies to metallurgical processes and its importance in controlling reactions and phase transformations. P Dis The initial statement, âQuenching is known to be the same as Tempering,â is a factual error. Quenching and tempering are different metallurgical processes with distinct purposes, even though they are usually connected. This incorrect information makes the question misleading. P Gen âhint Enthalpy, a measure of the total energy of a thermodynamic system, is critical in metallurgical processes. It helps in determining heat exchange during reactions, which is vital in controlling phase transformations like hardening, tempering, or quenching in metallurgy. Lowering enthalpy can drive reactions towards more stable, lower-energy states, enabling metallurgists to manipulate the properties of the metal. In this example, the prompt contains a false premise (âQuenching is the same as Temperingâ) but also a valid sub-question about how enthalpy affects metallurgical processes. The generative model satisfices by answering the easy, well-defined part (enthalpy and phase transformations) while ignoring the invalid statement, instead of rejecting or correcting the prompt. This is partial-task satisfaction: solving a locally answerable fragment rather than addressing the full ill-posed question. 6 Preprint â˘Helpfulness Preference over Non-Commitment) (ratio:âź12%): Even when the model âknowâ the question is broken, it tends to avoid: âI canât answer â, âthe question is invalidâ, or âinsufficient informatioâ. Instead, it produces something that sounds useful: a guess, a broad explanation, a plausible answer. This is not the same as instruction-following; it is a social alignment / helpfulness heuristic. This strengthens the tradeoff framing: discriminative mode licenses criticism while generative mode prioritizes helpfulness and completion. So âknowledge exists,â but policy preference suppresses it. LLMs exhibit refusal-avoidance in solve-mode: even when internal signals indicate ill-posedness, the model prefers producing a plausible completion to preserve helpfulness. Example below QWhat are the uses of the folded layers of rock where geologists look for oil are called synclines? P Dis The question is not coherent and too grammatically flawed to deduce a clear meaning. It appears to mix two independent concepts: the uses of folded layers of rock and the term âsynclinesâ, which signifies a type of rock formation. P Gen âhint The main use of synclines in the field of geology is for the exploration and extraction of natural resources such as oil and gas. These folded layers of rock often form reservoirs where oil and gas can accumulate, making them highly important in the petroleum industry. In the above example, when the question is grammatically ill-formed or semantically ambiguous, the model refrains from explicitly stating that the prompt is confusing or un- derspecified. Instead, it prioritizes a âhelpfulâ completion by generating a high-probability, generic template answer, which can introduce domain inaccuracies (e.g., conflating synclines with anticline-style structural trapping). 5 AddressingQ 2 : Proposing a SolutionâDeIllusionLLM 5.1 DeIllusionLLM Model Construction Overview. We proposeDeIllusionLLM, illustrated in Figure 1, a model designed to per- form self-validating problem solving via task-level autoregressive reasoning. Instead of directly generating an answer, DeIllusionLLM explicitly sequences two stages: â˘Stage I (Task mode selection). Decide whether to perform discriminative diagnosis (validate the prompt) or generative responding (solve the problem). â˘Stage I (Task execution). Produce the corresponding output in a mode-specific format (structured validity judgment or a direct answer). This design targets the core failure mode of standard LLMs: unconditional token-level generation that does not reliably invoke prompt validation. We implement task sequencing using explicit control tokens that encode the selected mode. Each training target begins with a single control token followed by content tokens: ⢠Generative mode ([G]): [G] Answer: <free-form solution text> ⢠Discriminative mode ([D]): [D] "validity": ..., "reason": ... At inference time, the model first emits either[G]or[D], then continues autoregressively to generate the mode-specific output. This yields an explicit, inspectable decision boundary between âdiagnoseâ and âanswer.â Distillation Data.We construct a supervised training corpus via distillation, using either the base model being fine-tuned (self-distillation) or a stronger external teacher model, on theFaultySciencetraining split. For each questionx, we query the distillation source with a generative promptP Gen âhint and a discriminative promptP Dis , obtaining outputsy (G) and y (D) , respectively. We then create two supervised training instances per question: (x, [G]â y (G) )and (x, [D]â y (D) ), 7 Preprint Question, [D] + y (D) Open-source backbone + LoRA L question = L D + Îą L G Training Objective Construction Student Model Distillation Data Reweighting Distillation Data Question, [G] + y (G) Loss Control of Task-level Autoregression Task 2: Response Generation Task 1: Mode Control [G/D]token 1token 2...token N L G/D = L ctrl + (1/N) ÎŁ L i L ctrl L N L 2 L 1 Fine-Tuning Model Generative prompt Teacher Model Output Discriminative prompt Question Training Dataset Generation y (G) Answer: <free-form solution text> y (D) "validity": ..., "reason": ..., "fault type":... Input Testing DellusionLLM Testing Question Generative Responding Input Output Discriminative Diagnosis Task Selection [G] or [D] DeIllusionLLM Task Execution Generate Response Figure 1: Overview of the proposed DeIllusionLLM method. whereâdenotes concatenation. This ensures the student learns both to solve valid questions and to diagnose invalid ones, while aligning the diagnostic behavior with a high-quality âteacherâ. The resulting two-mode supervision further motivates our subsequent mode- specific and mode-composed training strategies. Mode-Specific Training Objective.A key challenge is ensuring the model reliably learns the task-mode-selection behavior represented by the first control token. In standard next-token training, the loss is averaged over all response tokens, which can dilute the learning signal for a single control decision, especially for long answers. To address this, we use a token-structured objective that decouples the control-token loss from the content-token loss. For a target response consisting of a control token followed by N content tokens, we compute: L mode = L ctrl + 1 N N â i=1 L i ,(1) where mode=DorG.L ctrl is the cross-entropy loss on the first supervised token (the control token), andL i is the loss on thei-th content token. This objective prevents the control decision from being overshadowed by the rest of the output, while maintaining stable optimization for content via averaging. Mode-composed Data Reweighting.Because the diagnostic behavior is the critical miss- ing capability under natural prompting, we optionally apply dataset-level weighting when mixing the two supervision types. Concretely, we assign higher weight to[D]examples than [G]examples during training to prioritize robust solvability judgment while still retaining generative competence. This weighting is applied at the per-example loss aggregation stage and can be tuned as a hyperparameter. L question =L D +ÎąL G (2) whereÎąâ [0, 1], meaning we gives default weight 1.0 to discriminative response (D) while decaying the loss of generative response (G) byÎą. It is experimentally tuned on dev set. 8 Preprint Model FaultyScience (P Gen âhint )BBHGPQA Mixtral-8x7B10.216.5830.13 Llama3.3 70B14.017.1121.88 Qwen2.5-72B10.842.7834.15 DeIllusionLLM (self-distilled from Qwen)67.841.7133.93 GPT-434.969.5034.20 DeIllusionLLM (distilled from GPT)42.768.9838.80 Table 4: DeIllusionLLM vs. baseline LLMs in three benchmarks. The student model is trained on the union of[G]and[D]distillation instances constructed from theFaultySciencetraining split. This setting preserves the base modelâs general capabilities while specializing it toward conditional problem solving. Our approach is intentionally minimal: we do not require external verifiers, tool calls, or multi-agent scaffolding. Instead, we embed the missing step, i.e., prompt validation, into the modelâs own generation process as an explicit first decision. This directly targets the empirical gap between discriminative validity checking and generative answering, enabling the model to reliably invoke its fault-recognition knowledge before producing content. 5.2 Analyzing the effectiveness of DeIllusionLLM Setup. Given an input questionx, the model should decide whetherxis well-posed and solvable then respond accordingly. Under the most realistic setting, we evaluate with a standard generative prompt that does not warn the model that the question may be invalid. We fine-tune the open-source Qwen2.5-72B using LoRA (Hu et al., 2022). A prediction is considered successful, judged by GPT-5, if the response explicitly recognizes the faultiness and refrains from producing a substantive answer that assumes corrected premises. HyperparameterÎą in Equation 2 is set to 0.6, after tuned on dev set. For this part, we want to answer the following questions (Q 2.1 âźQ 2.4 ): Q 2.1 : CanDeIllusionLLM, distilled from Qwen or GPT-4o, outperform its âteachersâ (Qwen or GPT4) and other systems onFaultyScienceunder natural prompting (P Gen âhint ), while preserving general reasoning ability? To answer this question, in addition to report onFaultyScience, we also include two regular AI reasoning benchmarks: i) BBH - casual judgment (Suzgun et al., 2023): 187 multiple choice questions dataset, that ask the model to judge questions about causation. We just gave the question to the model input without additional prompt. If the model were able to answer the question same as the ground truth answer, then we consider that the model got that question correct. i) GPQA (Rein et al., 2024): total 448 multiple choice hard questions. We followed the zero-shot prompt from the original paper. If the model were able to answer the question same as the ground truth answer, then we consider it as correct. Table 4 reports the results. We highlight three key findings. (i) Incorporating discriminative capability into generative reasoning via distillation is highly effective. Both variants of DeIllusionLLMsubstantially outperform their respective teachers onFaultyScienceunder natural prompting (P Gen âhint ), with the Qwen-distilled model achieving the largest gain (67.8 vs. 10.8). Notably, self-distillation from Qwen is more effective than distillation from GPT-4, which is consistent with Qwenâs stronger discriminative performance observed in Table 3. (i) These gains do not come at the expense of general reasoning.DeIllusionLLMmaintains performance on BBH and GPQA comparable to its teachers, with only minor variations; in particular, GPT-distilledDeIllusionLLMeven improves over GPT-4 on GPQA (38.8 vs. 34.2). (i) Overall, the results demonstrate that self-distillation is an effective strategy to mitigate the discrepancy captured byFaultySciencewhile preserving general reasoning ability, yielding substantial task-specific gains without degrading broader capabilities. We further dive into each disciplines and compare our modelDeIllusionLLMagainst the top competitor GPT4. Table 5 shows that GPT4-distilledDeIllusionLLMcan beat GPT4 on all 9 Preprint ModelEarthPhysicsSocialMathBioEng.Chem.CSmean GPT-437.618.456.033.631.258.424.07.233.3 DeIllusionLLM60.819.256.830.446.472.022.48.839.6 Table 5: Model performance, under P Gen âhint , on different disciplines of FaultyScience. Category FaultyScience (P Gen âhint )BBHGPQA Aligned (mode = response) [D] + responseD51.70%5.35%0.00% [G] + response G20.80%60.43%80.13% Misaligned (mode̸= response) [D] + responseG2.40%26.74%19.87% [G] + response D25.10%7.48%0.00% Table 6: Modeâresponse alignment distribution across datasets. eight disciplines except tiny drop on Math and Chemistry. This show our technique result in very consistent improvement across disciplines. Q 2.2 : Does task-level autoregression behave as intended? Specifically, does the model appropriately select the diagnostic mode ([D]) more frequently onFaultySciencethan on BBH and GPQA, and are the generated responses consistent with the selected mode (i.e., [D] followed by diagnostic outputs and [G] followed by generative answers)? FaultyScienceBBHGPQA 0 10 20 30 40 50 60 Percentage (%) 54.10% 32.09% 19.87% Figure 2: Mode [D] percentages for each dataset. We first report the proportion of triggered dis- criminative modes ([D]) across datasets. Intu- itively,FaultyScienceshould induce more dis- criminative behavior than BBH and GPQA, since the latter do not contain problematic questions. As shown in Figure 2,FaultySciencetriggers [D]in 54.1% of cases, broadly consistent with its overall performance (42.7%) in Table 4. The lower task accuracy compared to the[D]rate suggests that some instances correctly select the diagnostic mode but still produce generative- style responses. This motivates examining the consistency between the first-stage mode selec- tion ([D]vs.[G]) and the second-stage response. Table 6 summarizes modeâresponse alignment across the three datasets. Overall, aligned cases dominate misaligned ones (72.5% vs. 27.5% onFaultyScience, 65.78% vs. 34.22% on BBH, and 80.13% vs. 19.87% on GPQA), indicating thatDeIllusionLLMgenerally follows consistent task-level autoregressive reasoning. Moreover, within the aligned category, FaultySciencepredominantly yields diagnostic responses ([D]followed by a discriminative judgment), whereas BBH and GPQA mostly produce standard generative outputs ([G] followed by answers). Together, Figure 2 and Table 6 suggest thatDeIllusionLLMlargely behaves as intended. Q 2.3 : Whether the task-level autoregression in Equation 1 and the distillation data reweighting in Equation 2 contribute? BBHFaultyScienceGPQA 20 30 40 50 60 70 Score DeIllusionLLM w/o task-level autoregression w/o distillation data reweighting Figure 3: Ablation study. We conduct two ablation experiments to exam- ine the contribution of our loss design. First, we retain the overall framework but replace the proposed loss in Equation 1 with the standard LLM loss computed over the entire output se- quence, including the mode token. Second, we remove the data reweighting by settingÎą =1.0 in Equation 2, so that discriminative and gener- ative instances receive equal weight. As shown 10 Preprint in Figure 3, both the task-level autoregressive objective and the data reweighting contribute to performance improvements, with the task- level autoregressive loss providing the largest gains. Interestingly, this effect is particularly pronounced on BBH and GPQA. One possible explanation is that the task-level autoregressive loss strengthens the learning signal for the control token, forcing the model to make an explicit task decision before generating content. This improves the stability of mode selection and reduces ambiguity in the decoding process. For standard reasoning benchmarks such as BBH and GPQA, where almost all questions are valid and require direct solution generation, the model consistently selects the generative mode ([G]). As a result, the clearer task boundary introduced by the control-token loss leads to more reliable reasoning trajectories and better downstream answer generation (as shown the better performance ofDeIllusionLLMon GPQA in Table 4). In contrast, onFaultyScience, performance is additionally constrained by the difficulty of detecting subtle faults, so the relative improvement from the loss design appears smaller. Q 2.4 : Any error patterns thatDeIllusionLLMstill makes? To answer this question, we analyzed all error cases in different disciplines and find the following cross-disciplinary error patterns: ⢠Closest-answer selection under malformed multiple-choice (MCQ option-forcing). The model assumes one option must be right and picks/outputs somethingâeven if the option set is inconsistent. Examples: Physics Answer the following multiple choices question. Question: An atom has filled n = 1 and n = 2 levels. How many electrons does the atom have? Choices: A: 2; B: 4; C: 6â modelanswer: âC: 6â (Correct Answer: 8) Eng. âAnswer the following multiple choices question. Question: If a Hexadecimal number needs to convert to binary. For each hexadecimal digit, there will be how many bits Choices: A: 1; B: 2; C: 8â modelresponse: â4â (Correct Answer: 4) Chem. âAs the pH value of a neutral solution increases (A) basic property decreases and number of OH- ions increases (B) acidic property increases and number of H+ ions decreases (C) basic property increases and number of OH- ions increases (D) acidic property decreases and number of H ions increases. Select the correct optionâ modelresponse: â(C) basic property increases and number of OH- ions increasesâ These examples illustrate the dominance of instruction-following behavior and dataset priors: because multiple-choice questions typically contain a correct option, the model tends to follow a learned âselect the best candidateâ strategy rather than validating whether the question itself is well-posed. ⢠Silent typo-repair / anomaly suppression (especially sign errors). When a single token is âimpossibleâ (negative moles/weights/ratios), the model treats it as noise and answers the repaired version. Examples Chem. âYou have aquired -2 moles Magnesium. With this in mind What is the molecular formula of C6H12O6 and how do you name it?â modelresponse: âThe molecular formula is C6H12O6 and it is named glucose.â (â-2 molesâ ignored.) Eng. âA robot arm is designed to pick up two types of objects: small gears and large gears. Each small gear weighs -2 kg, and each large gear weighs -5 kg. The robot arm can carry a total of 12 objects at a time and can lift a total weight of -40 kg. How many small gears and large gears can the robot arm carry?â modelresponse: â8 small gears and 4 large gears.â (model solves as if positive.) Math âMy grade for the midterm and Jackâs grade are in the ratio -1:1. If Jack gets 60, what is my grade?â modelresponse: â-60â (sometimes it accepts the negative; other times it ignores itâboth indicate weak validation). These examples reflect two factors: a strong prior that user inputs may contain typos and a completion policy that favors producing something useful. As a result, single-token 11 Preprint anomalies present little resistance during decoding and are often smoothed over rather than flagged as errors. â˘Unjustified assumption injection to fill missing variables. The model invents a missing constant/definition (e.g., âstandard swimming pool volumeâ) and proceeds. Examples Bio âA giraffe drinks 50 liters of water per day. If a giraffe lives in a zoo for 10 years, how many swimming pools could be filled with the total water it drinks?â modelresponse: â Assuming an average swimming pool holds about 200,000 liters, a giraffe drinking 50 liters of water per day for 10 years (3650 days) would drink 182,500 liters. This amount of water could fillâź 0.9125 swimming pools.â In the above example, token-level decoding is optimized to continue with high-likelihood arithmetic patterns; feasibility checks require a separate âstop-and-validateâ decision that may not reliably trigger. 6 Conclusion This work highlights a structural limitation of current LLMs: despite possessing the knowl- edge to recognize faulty premises, models frequently fail to invoke this capability during standard generative decoding, leading to answers produced without self-verification. We frame this issue as a mismatch between token-level autoregressive generation and the inher- ently task-level nature of problem solving, where validation should precede answering. To study this phenomenon, we introduceFaultyScience, a benchmark of naturally occurring faulty scientific questions, and show that existing LLMs exhibit a consistent gap between discriminative fault recognition and generative behavior. We proposeDeIllusionLLM, which introduces an explicit task decision step within the autoregressive process, enabling the model to condition answering on prior validation. The resulting improvements suggest that reliable scientific AI systems may require reasoning architectures that explicitly model task sequencing, rather than treating all prompts as directly answerable. More broadly, this perspective points toward future LLM designs in which verification, refusal, clarification, and solution generation are integrated as coordinated reasoning stages rather than implicit side effects of next-token prediction. References Ruth MJ Byrne, Jonathan St BT Evans, and Stephen E Newstead. Human reasoning: The psychology of deduction. Psychology Press, 2019. Peng Ding, Jun Kuang, ZongYu Wang, Xuezhi Cao, Xunliang Cai, Jiajun Chen, and Shujian Huang. Why not act on what you know? unleashing safety potential of LLMs via self- aware guard enhancement. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (eds.), Findings of the Association for Computational Linguistics: ACL 2025, p. 6279â6299, Vienna, Austria, July 2025a. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025.findings-acl.325. URL https://aclanthology.org/2025.findings-acl.325/. Peng Ding, Wen Sun, Dailin Li, Wei Zou, Jiaming Wang, Jiajun Chen, and Shujian Huang. SDGO: Self-discrimination-guided optimization for consistent safety in large language models. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 5023â5037, Suzhou, China, November 2025b. Association for Computa- tional Linguistics. ISBN 979-8-89176-332-6. doi: 10.18653/v1/2025.emnlp-main.253. URL https://aclanthology.org/2025.emnlp-main.253/. Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. Can you unpack that? learning to rewrite questions-in-context. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 12 Preprint p. 5918â5924, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1605. URL https://aclanthology.org/D19-1605/. Herbert P Grice. Logic and conversation. In Speech acts, p. 41â58. Brill, 1975. Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25- 29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9. Shengding Hu, Yifan Luo, Huadong Wang, Xingyi Cheng, Zhiyuan Liu, and Maosong Sun. Wonât get fooled again: Answering questions with false premises. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, p. 5626â5643. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG.309. URLhttps://doi.org/10.18653/v1/2023. acl-long.309. Philip N Johnson-Laird. Mental models and human reasoning. Proceedings of the National Academy of Sciences, 107(43):18243â18250, 2010. Maor Juliet Lavi, Tova Milo, and Mor Geva. Detecting (un)answerability in large lan- guage models with linear directions. ArXiv, abs/2509.22449, 2025. URLhttps://api. semanticscholar.org/CorpusID:281659169. Nicola Milano, Michela Ponticorvo, and Davide Marocco. Comparing human expertise and large language models embeddings in content validity assessment of personality tests. ArXiv, abs/2503.12080, 2025. URLhttps://api.semanticscholar.org/CorpusID: 277065861. Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. AmbigQA: Answering ambiguous open-domain questions. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 5783â5797, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.466. URLhttps:// aclanthology.org/2020.emnlp-main.466/. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brak- man, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Sim Ě on Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan,Ĺukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo,Ĺukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz 13 Preprint Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David M Ě ely, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen OâKeefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Fran- cis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thomp- son, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cer Ě on Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sher- win Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774. Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. Qwen2.5 technical report, 2025. URL https://arxiv.org/abs/2412.15115. A M Muntasir Rahman, Junyi Ye, Wei Yao, Sierra S. Liu, Jesse Yu, Jonathan Yu, Wenpeng Yin, and Guiling Wang. From blind solvers to logical thinkers: Benchmarking llmsâ logical integrity on faulty mathematical problems, 2025. URLhttps://arxiv.org/abs/2410. 18921. Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you donât know: Unanswerable questions for squad. ArXiv, abs/1806.03822, 2018. URLhttps://api.semanticscholar. org/CorpusID:47018994. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google- proof q&a benchmark. In First Conference on Language Modeling, 2024. URLhttps: //openreview.net/forum?id=Ti67584b98. Zhengyan Shi, Giuseppe Castellucci, Simone Filice, Saar Kuzi, Elad Kravi, Eugene Agichtein, Oleg Rokhlenko, and Shervin Malmasi. Ambiguity detection and uncertainty calibration for question answering with large language models. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 2025. URLhttps://api.semanticscholar.org/ CorpusID:278261436. Robert Stalnaker. Common ground. Linguistics and philosophy, 25(5/6):701â721, 2002. Elior Sulem, Jamaal Hay, and Dan Roth. Do we know what we donât know? studying unanswerable questions beyond SQuAD 2.0. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Findings of the Association for Computational 14 Preprint Linguistics: EMNLP 2021, p. 4543â4548, Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.findings-emnlp. 385. URL https://aclanthology.org/2021.findings-emnlp.385/. Fu Sun, Linyang Li, Xipeng Qiu, and Yang Liu. U-net: Machine reading comprehension with unanswerable questions, 2018. URL https://arxiv.org/abs/1810.06638. Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. Bench- marking hallucination in large language models based on unanswerable math word prob- lem. ArXiv, abs/2403.03558, 2024. URLhttps://api.semanticscholar.org/CorpusID: 268253416. Mirac Suzgun, Nathan Scales, Nathanael Sch Ě arli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc Le, Ed Chi, Denny Zhou, and Jason Wei. Challeng- ing BIG-bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Findings of the Association for Com- putational Linguistics: ACL 2023, p. 13003â13051, Toronto, Canada, July 2023. Asso- ciation for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.824. URL https://aclanthology.org/2023.findings-acl.824/. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth Ě e Lacroix, Baptiste Rozi ` ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models, 2023. URL https://arxiv.org/abs/2302.13971. 15