Paper deep dive
Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency
Zhihong Cao, Chen Huang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/27/2026, 3:56:15 AM
Summary
The paper introduces PASSING, a proactive conversational agent framework designed to clarify user expertise through targeted inquiries before generating responses. Unlike existing methods that struggle to infer expertise from queries alone, PASSING uses LLM-induced 'What-to-Ask' and 'How-to-Ask' strategies derived via self-play simulation to probe users. This approach significantly improves the accuracy of query-specific expertise estimation while minimizing conversational overhead.
Entities (10)
Relation Signals (8)
PASSING â usesstrategy â How-to-Ask
confidence 95% ¡ PASSING employs a strategy-guided probing mechanism, which utilizes LLM-induced strategies (i.e., What-to-Ask and How-to-Ask)
PASSING â usesstrategy â What-to-Ask
confidence 95% ¡ PASSING employs a strategy-guided probing mechanism, which utilizes LLM-induced strategies (i.e., What-to-Ask and How-to-Ask)
What-to-Ask â derivedfrom â LLM Self-Play
confidence 90% ¡ This is achieved by our What-to-ask and How-to-ask strategies, induced by LLM self-play.
How-to-Ask â derivedfrom â LLM Self-Play
confidence 90% ¡ This is achieved by our What-to-ask and How-to-ask strategies, induced by LLM self-play.
PASSING â evaluatedon â ARC-MCAS
confidence 90% ¡ we use questions from two multi-disciplinary datasets as input user queries: ... ARC-MCAS
PASSING â evaluatedon â MMLU-Pro
confidence 90% ¡ we use questions from two multi-disciplinary datasets as input user queries: MMLU-Pro
User Expertise â categorizedby â Dreyfus Model
confidence 88% ¡ we adopt the Dreyfus Model (Dreyfus and Dreyfus, 1980), which categorizes user expertise into five levels
PASSING â outperforms â ExpertPrompting
confidence 85% ¡ Our extensive experiments also show our superiority... ExpertPrompting... typically fails to accurately predict query-level expertise.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized assistants. A critical aspect of this evolution is the ability to tailor strategic interactions to a user's unique needs and expectations. Unlike existing studies that focus on proactively clarifying query ambiguities, we center on clarifying the user's expertise in order to tailor responses for better user comprehension. We find that existing agents struggle to determine user expertise from queries alone, a limitation that prevents them from dynamically adapting their responses. To address this gap, we introduce PASSING to empower the agent to proactively clarify a user's expertise through targeted inquiries. This is achieved by our What-to-ask and How-to-ask strategies, induced by LLM self-play. Our extensive experiments also show our superiority. We believe that PASSING represents a crucial step towards creating more human-centric conversational agents.
Tags
Links
- Source: https://arxiv.org/abs/2608.22266v2
- Canonical: https://arxiv.org/abs/2608.22266v2
Trouble viewing inline? Open PDF directly â
Full Text
93,422 characters extracted from source content.
Expand or collapse full text
Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency Zhihong Cao 1 Chen Huang 2,3 * 1 School of Computing and Data Science, The University of Hong Kong 2 Institute of Data Science, National University of Singapore 3 College of Computer Science, Sichuan University czh040302@163.com, huangc.scu@gmail.com Abstract In the context of information seeking, conversa- tional agents are undergoing an evolution from reactive tools to proactive, personalized assis- tants. A critical aspect of this evolution is the ability to tailor strategic interactions to a userâs unique needs and expectations. Unlike exist- ing studies that focus on proactively clarify- ing query ambiguities, we center on clarifying the userâs expertise in order to tailor responses for better user comprehension. We find that existing agents struggle to determine user ex- pertise from queries alone, a limitation that prevents them from dynamically adapting their responses. To address this gap, we introduce PASSING to empower the agent to proactively clarify a userâs expertise through targeted in- quiries. This is achieved by our What-to-Ask and How-to-Ask strategies, induced by LLM self-play. Our extensive experiments also show our superiority. We believe that PASSING rep- resents a crucial step towards creating more human-centric conversational agents. 1 Introduction Conversational agents have become a crucial tool for satisfying user information needs in multi-turn dialogues (Casheekar et al., 2024; Zamani et al., 2023). With the advent of large language models (LLMs), a new paradigm of Proactive Conversa- tional Agents has emerged that goes beyond simply reacting to user queries (Deng et al., 2025; Liao et al., 2023; Huang et al., 2026). These agents take the initiative to handle under-specified (Braslavski et al., 2017; Xu et al., 2019) or over-specified user requests (Wu et al., 2023; Min et al., 2019) during the information seeking process, rather than pas- sively following the userâs lead. Common proactive behaviors include asking questions to clarify query ambiguity (Zhang et al., 2024b; Chen et al., 2024b) or elicit user preference (Zhang et al., 2018, 2022; * Corresponding author. Figure 1: User expertise can vary depending on the specific query. LLM agent should tailor its response based on user expertise. Chen et al., 2023). For instance, in response to a broad query like "quantum computing", these agents might ask clarifying question: "Are you ask- ing about the history of quantum computing?". As such, the agent seeks to proactively comprehend user query to deliver a more accurate response. While current research prioritizes response ac- curacy, the crucial task of tailoring responses for user comprehension has been largely overlooked. Modern users expect responses that are not only factually correct but also personalized to their in- dividual domain expertise, adjusting the level of detail and complexity accordingly (Dreyfus and Dreyfus, 1980; Schank, 1999; Chen et al., 2024a; Yao et al., 2024; Jeck et al., 2025; Tawfik et al., 2021; Yaghoubzadeh and Kopp, 2017). As illus- trated in Figure 1, user expertise can vary signif- icantly depending on the specific query. LLM agents must estimate the userâs query-specific ex- pertise level prior to formulating a response. Oth- erwise, novices, for example, can be overwhelmed by technical details, making them susceptible to arXiv:2608.22266v2 [cs.AI] 26 Aug 2026 cognitive overload (Dreyfus and Dreyfus, 1980). To this end, our work centers on clarifying user expertise as a foundational step for generat- ing tailored responses 1 . We begin by experimen- tally investigating the limitations of current agents in query-specific expertise estimation (Section 3). Our findings reveal a significant expertise inconsis- tency in LLM-based agents: when presented with the same query, their predictions of user expertise are highly stochastic, with high inconsistency rate. This aligns with previous user study (Palta et al., 2025) and implies that a userâs expertise on the specified query is difficult to infer from the query text (Liu et al., 2024). For instance, a computer sci- ence PhD student may be a novice in physics but inquiry technical terms like âquantum computingâ without a full grasp of the concept. Therefore, we argue that proactive conversational agents must probe the userâs expertise on a given query via strategic interaction, a need previously identified (Chen et al., 2024a) but not yet fulfilled. In this paper, we propose PASSING, the first method forProActive uSer expertiSe probING through targeted inquiries. Recognizing that ex- pertise is often manifested in a userâs knowledge scope and reasoning processes (Feltovich and Hoff- man, 1997; Palta et al., 2025), PASSING employs a strategy-guided probing mechanism, which uti- lizes LLM-induced strategies (i.e., What-to-Ask and How-to-Ask) to direct PASSING in formulating informative inquiries. As illustrated in Figure 2, PASSING leverages these strategies to iteratively pose targeted questions at each turn, obtaining user information on the knowledge scope and reason- ing logic based on userâs responses. To balance the trade-off between maximizing information gain and minimizing conversational overhead, PASS- ING stops probing after a pre-defined number of turns. Subsequently, it estimates the userâs query- specific expertise and generates an appropriately tailored response. Before putting PASSING into use, PASSING implements an offline self-play sim- ulation to derive these strategies: dialogue histories between PASSING and simulated users are analyzed to extract effective probing strategies, which are then iteratively optimized throughout the self-play process. Notably, distinct from research on am- biguity resolution or preference elicitation, PASS- 1 Our research scope is confined to the task of query- specific user expertise estimation, a prerequisite for generating tailored responses. Response features required by experts and novices is provided in Appendix A. ING prioritizes the assessment of user expertise to facilitate better answer comprehension, addressing a vital gap in human-centric AI (Deng et al., 2024). To evaluate the effectiveness of PASSING, we conduct experiments with various baselines and LLM backbones using various datasets. Experi- mental results demonstrate that with an average of only 1.2 inquiry turns, PASSING achieves nearly a threefold improvement in the accuracy of user expertise estimation. Finally, we sum up our main contributions as follows. â˘We highlight the importance of user expertise estimation for proactive conversational agents, a step toward more human-centric interactions. â˘We propose PASSING, featuring accurate exper- tise estimation through targeted inquiries while minimizing conversational overhead. ⢠We experimentally demonstrate that existing agents fail to reliably infer user expertise from queries, and validate our effectiveness. 2 Related Work User Expertise Estimation. A body of theoretical (Dreyfus and Dreyfus, 1980; Schank, 1999) and user-centric (Chen et al., 2024a; Schubert et al., 2013; Palta et al., 2025) research highlights the sig- nificant differences between novices and experts in their perceptions and needs. In this case, to gener- ate such personalized response, existing research largely assumes that user expertise is a known in- formation (e.g., via pre-defined user profiles) (Jeck et al., 2025; Sun et al., 2025; Cheng et al., 2024). However, such information is insufficient for deter- mining query-specific expertise, because a userâs expertise can vary significantly from one query to the next. While recent methods like ExpertPrompt- ing (Xu et al., 2025) attempt to infer expertise via in-context learning, user studies indicate that LLM- based agents frequently miscalibrateâeither un- derestimating or overestimating the userâs knowl- edgeâresulting in reduced satisfaction (Palta et al., 2025). Therefore, proactive agents must be able to probe for this information before generating a response. To fill this gap, we propose PASSING, the first method designed to leverage an agentâs proactivity to probe and clarify a userâs expertise on the user query through targeted inquiries. Proactive Conversational Agents. Analogous to general proactive agents that explore an envi- ronment to acquire information that improves fu- ture decision-making (Lu et al., 2025; Guan et al., 2026), proactive conversational agents actively en- gage users to uncover task-relevant information and facilitate goal completion. These agents are distinguished by their ability to take initiative, steer- ing dialogues toward productive outcomes rather than passively following a userâs lead (Deng et al., 2024, 2025). To date, proactive efforts in informa- tion seeking have primarily focused on improving response accuracy. This is achieved through strate- gies such as clarifying query ambiguity (Zhang et al., 2024b; Chen et al., 2024b), eliciting user preferences (Zhang et al., 2018, 2022; Chen et al., 2023), or managing over-specified requests (Wu et al., 2023; Min et al., 2019). However, this in- tense focus on accuracy has left the crucial task of tailoring responses for user comprehension largely overlooked. User studies indicate a clear demand for agents that first assess a userâs knowledge level before providing an answer (Chen et al., 2024a). In response to this need, our work introduces a proac- tive conversational agent designed specifically to probe for a userâs expertise on a given query before generating a response. 2.1 Prediction Stability of Existing Methods Method MMLU-ProARC-MCAS KappaUnkKappaUnk Deterministic100.00.00100.00.00 Fully Random0.0016.670.0016.67 Direct Prompt41.2528.0030.6482.80 CoT28.4726.4070.3489.60 Self-consistency85.1232.0075.3586.80 Diverse-Aspect 39.9114.8047.2325.60 IDL*40.3123.7915.004.40 ExpertPrompting36.2525.6072.0094.00 Table 1: Prediction stability and unknown rate of exist- ing methods for query-specific user expertise estimation. 3 Preliminary Experiment This section investigates LLM-based agents can reliably determine user expertise on a given query. 3.1 Experiment Setup Experiment Overview. A single query can be posed by users with vastly different expertise levels. For instance, a query about quantum mechanics could originate from a physics novice, a physics expert, or even a computer science expert who is a novice in quantum mechanics. Considering the scarcity of open-source data labeled with query- specific user expertise, we propose an evaluation methodology focused onpredictionstabilityof an agentâs expertise judgments when it is presented with the same query multiple times. Baselines. As there are no established methods specifically for probing query-specific user ex- pertise, our baselines include two distinct cate- gories. 1) General LLM-based Agents:Direct Prompt,Chain-of-Thought(CoT) (Wei et al., 2023),Self-consistency(Wang et al., 2022), and our designedDiverse-Aspectthat first prompts the LLM to generate judgments from multiple specific aspects (e.g., terminology use) and then ensembles these aspect-specific judgments to produce the final prediction. 2) Expertise Estimation:IDL(Cheng et al., 2024) andExpertPrompting(Xu et al., 2025). Notably, IDL predicts general user expertise based solely on pre-collected user conversation history 2 and are inherently unable to provide an estimation specific to the given input query. To the best of our knowledge, ExpertPrompting is the only query- specific expertise estimation method. Finally, GPT- 4o serves as the LLM backbone. Dataset. To ground our evaluation in realistic and challenging scenarios, we use questions from two multi-disciplinary datasets as input user queries: MMLU-Pro (Wang et al., 2024) and ARC-MCAS 3 (Clark et al., 2018). For the output classes, we adopt the Dreyfus Model (Dreyfus and Dreyfus, 1980), which categorizes user expertise into five levels:Novice,AdvancedBeginner,Competent, Proficient, andExpert(detailed in Appendix A). We further augment these five expertise levels with an additional âUnknownâ category. This results in a six-way classification task where each baseline must classify the userâs expertise on the input query into one of these categories. A more comprehen- sive evaluation is presented in Section 5. Evaluation Metrics. Given the multi-class classifi- cation nature of our task, we measure prediction sta- bility usingFleissâKappa(Fleiss, 1971). This met- ric measures the inter-run agreement for individual predictions. We treat the set of predictions gen- erated under each random seed as an independent "rater" and calculate the Kappa score across all user queries to assess consistency between runs. Addi- tionally, we calculate the proportion of âUnknownâ predictions generated by the model, which serves 2 History is built from validation set of corresponding data. 3 Subset of the ARC-Challenge Strategy Induction via Self-play Strategy-guided Inquiry Generation User Iteration t Stage 1 User Simulator Group 1 Stage 2 User Simulator Group 2 Update âWhat-to-askâ strategies Update âHow-to-askâ strategies Iteration t+1 Stage 1 User Simulator Group 1 Stage 2 User Simulator Group 2 Update âWhat-to-askâ strategies Update âHow-to-askâ strategies Iteration T Stage 1 User Simulator Group 1 Stage 2 User Simulator Group 2 Update âWhat-to-askâ strategies Update âHow-to-askâ strategies Refine two strategy sets Refine two strategy sets Agent User answersthe inquiry 1. Use What-to-ask Strategies Decide WHAT information to ask about. Input Conversation history + Populated Schema Output Information to ask (what) 2. Use How-to-ask Strategies Decide HOW to ask that question. Input Information to ask (what)+ Conversation history Output Formulated question (how to ask) 3. Generate Inquiry Generate the inquiry (question) to ask the user. Input Formulated question (how)+ Conversation history Output Inquiry to the user 4. Fill Schema with User Response Extract and fill information into the schema. Input User response+ Conversation history Output Updated schema(filled slots) 5. Estimate User Expertise Level Estimate the user's level of expertise based on thecollected information. Input Filled schema+ Conversation history Output Estimated expertise level 6. Provide Answer to Initial Question Provide the final answer to the user's initial question. Input Filled schema+ Conversation history+ Estimated expertise level Output Final answer to theinitial question Agent provides the answer User asks a question User User Figure 2: Overview of PASSING. It utilizes LLM-induced strategies (i.e., What-to-Ask and How-to-Ask), obtained through self-play simulation. as a direct statistical measure of the modelâs self- assessed inability to determine user query-specific expertise. Refer to Appendix B.5 for details. Implementation Details. To evaluate the stability of each baselineâs judgment, we conducted 5 inde- pendent runs for every query. Each of these 5 runs was initialized with the same pre-defined random seeds. More details are in Appendix B. 3.2 Results LLM agents are unreliable for robustly esti- mating user expertise specific to a given query. As illustrated in Table 1, while existing methods demonstrate improved stability compared to fully random approaches when assessing user query- specific expertise, they still lack consistency com- pared to the deterministic method, often fluctuating in their judgments. The consistently low Kappa scores across methods reveal that even when mod- els do produce a prediction, their judgments are highly inconsistent across runs. Notably, while Self-consistency achieves a relatively high Kappa (85.12 on MMLU-Pro, 75.35 on ARC-MCAS), this stability is an artifact of its majority-voting mecha- nism, which is the aggregation of multiple samples naturally converges to a dominant prediction, rather than reflecting genuine expertise inference. Simi- larly, the high Kappa of certain methods on ARC- MCAS (e.g., CoT : 70.34, ExpertPrompting: 72.00) is largely driven by an extremely high Unknown rate (89.60% and 94.00% respectively), indicating that the model repeatedly abstains rather than mak- ing meaningful predictions. Importantly, Expert- Prompting, which is designed to estimate general user expertise, typically fails to accurately predict query-level expertise. This highlights that any user can copy-paste a jargon-heavy query from the web into an LLM, creating a façade of expertise despite having no actual proficiency in the domain. This motivates our proactive agent to probe the userâs query-level expertise. 4PASSING: Clarify User Expertise Overview. As illustrated in Figure 2, PASSING es- timates a userâs query-specific expertise through it- erative probing and slot-based state tracking. Given a user queryq, PASSING first constructs a query contextC 0 using the query and model-generated an- swer. It then engages the user in a multi-turn prob- ing process guided by predefined probing strategies S. Formally, letH t = (p i ,u i ) tâ1 i=1 denote the con- versation history before turnt, wherep i andu i represent the system inquiry and user response, re- spectively, withu 1 = q. LetL t denote the current slot-based expertise schema. At each turn, PASS- ING generates an inquiryp t based on the current state(C 0 ,L t ), receives the userâs responseu t , and updates both the conversation state and schema to incorporate newly acquired evidence about the userâs expertise. This process repeats until a termi- nation condition is met. Then PASSING produces a query-specific expertise assessment and generates a personalized response accordingly. Multi-turn Probing and Slot Filling. To capture query-specific expertise, we design an expertise schema based on the Dreyfus Model (Dreyfus and Dreyfus, 1980; Mangiante and Peno, 2021), where each slot corresponds to a key assessment aspect (Table 2). Following prior work on dialogue state tracking (Cao et al., 2025; Das et al., 2024), PASS- ING adopts a slot-filling paradigm that incremen- tally populates the schema using evidence collected during probing. After each user response, relevant slots are updated to reflect newly inferred expertise signals, and the updated schema is subsequently used to guide the next inquiry. To further improve inquiry generation, we incorporate LLM-induced probing strategies, described in Section 4.2. Query-specific Expertise Assessment. The prob- ing process terminates when either the maximum probing budgetTis reached or all schema slots have been populated. PASSING then estimates the userâs query-specific expertise using the populated schema together with the accumulated conversa- tion history, and generates a response tailored to both the userâs query and estimated expertise level. 4.1 Strategy Induction via Self-play The key challenge in expertise estimation lies in generating probing questions that effectively reveal user expertise while minimizing cognitive burden. This requires determining both what to ask and how to ask. To this end, PASSING learns these strategies through offline self-play with user simu- lators 4 and iteratively refines its probing strategies via trial-and-error. Specifically, strategy induction is performed in two stages using distinct simulators and optimization objectives. The first environment focuses on identifying missing or uncertain knowl- edge areas, inducing What-to-Ask strategies that guide the selection of probing targets. The second environment focuses on maximizing information gain while maintaining user engagement, inducing How-to-Ask strategies that govern the phrasing and presentation of inquiries. Diverse User Simulation. Each simulator is instan- tiated with a unique persona, including Big-Five Personality Traits 5 (Goldberg, 1992) and Decision- Making Styles 6 (Scott and Bruce, 1995), to encour- age behavioral diversity during interaction. To model varying expertise levels, each simulator is additionally assigned a query and a query-specific expertise label (e.g., novice). We further equip the 4 Refer to Appendix B.4.1 for implementation details 5 Openness, Conscientiousness, Extraversion, Agreeable- ness, and Neuroticism 6 Directive, Conceptual, Analytical, and Behavioral Algorithm 1 Strategy Induction via Self-play Require: N: Max simulation sessions,T: Max conversation turns 1:⡠W : What-to-Ask, H : How-to-Ask 2: for M âW, H do 3:Initialize strategy S M 0 4:Initialize user simulatorsU M 5:for k = 0 to N â 1 do 6:ExperienceE ââ 7:for all uâU M do 8:câ SELFPLAY(S M k , u, T) 9:E âE âŞc 10:end for 11:S M k+1 â EXTRACT&REFINE(S M k ,E) 12:end for 13: end for simulator with a query-specific knowledge depen- dency graph automatically generated by GPT-5, which contains the knowledge required to under- stand and answer the query. To control the simu- latorâs expertise, we randomly mask different por- tions of the knowledge dependency graph accord- ing to its assigned expertise level. Consequently, novice simulators possess only partial knowledge, whereas expert simulators retain more complete knowledge. This design provides a controllable en- vironment for learning What-to-Ask strategies. For How-to-Ask strategy induction, the simulator must also exhibit non-cooperative behaviors, as users may refuse to answer poorly phrased or overly de- manding questions. Thus, we follow Zhang et al. (2024a) and augment the simulators with refusal strategies and corresponding few-shot examples, enabling them to reject inquiries under appropriate circumstances. This forces PASSING to learn ques- tioning strategies that maximize information gain while maintaining user cooperation. Self-Play Pipeline. To iteratively optimize probing strategies, we conductNsessions of offline self- play simulations. We initialize the baseline strategy setS 0 using Gemini 2.5 Pro. In eachk-th session, PASSING employs the current strategies (S k ) to in- teract with a user simulator (GPT-4o Mini), aiming to estimate its expertise as described previously. A session terminates upon successful estimation or when the maximum turn limitTis reached. Fol- lowing each session, we analyze the dialogue logs as experience to induce new strategies, which are then refined to update the set for the next iteration. Refer to Algorithm 1 for details. â˘Strategy Extraction. Inspired by recent stud- ies in inductive learning, which empowers LLMs to synthesize broader patterns from granular ex- AspectsDescriptions Knowledge Scope (What they know) The extent, type, and organization of the knowledge of the person: knowing âfactsâ (explicit knowledge), knowing when to apply it (applied knowledge), or knowing âhowâ intuitively (tacit knowledge) Components (What they perceive) What types of knowledge or information the person relies on: context-free (rules, abstract principles), rule-based (context-aware rule selection), or situational (contextual, experiential patterns) Perspective (How they organize info) How the person filters and prioritizes information: overwhelmed easily, making a deliberate plan to filter noise, or instinctively seeing important information without trying Action (How they decide) How decisions are made: analytic (step-by-step reasoning) or intuitive (immediate, holistic response) Commitment (Emotional connection) The level of personal agency and emotional investment in the outcome: detached (disengaged, rule-following) or involved (personally invested, responsible for outcome) Table 2: Structured schema used for estimating user expertise based on the Dreyfus Model. amples (Cai et al., 2025; de Souza et al., 2025; Huang et al., 2024; Wang et al., 2025; Kim et al., 2025), the accumulated self-play experience is distilled into probing strategies that generalize across users and queries. For each self-play stage, we collect simulation dialogue logs together with the probing strategies employed during interac- tion. We then analyze these trajectories from two complementary perspectives: (1) the accu- racy of the final expertise assessment, and (2) the information gain obtained at each probing turn. Using the resulting dialogue trajectory and assess- ment outcome as context, an LLM is prompted to identify why a particular strategy contributed to successful or failed expertise estimation, or why a probing turn yielded high or low information gain. The extracted insights are subsequently distilled into strategies in a structured form: Probe [TAR- GET] to elicit information regarding [ASPECT], because [RATIONALE]. â˘Strategy Refinement. Since a single session comprises multiple turns, numerous raw strate- gies are extracted. To prevent these strategies from becoming overly specific to the query or its associated knowledge points, and to avoid re- dundancy, we employ an LLM-based semantic clustering approach, together with the strategies from previous session. Strategies within the same cluster are then abstracted to derive high-level probing tactics for use in subsequent simulations. Note that we use the same structured template for both strategy types to control for sentence- structure confounds while allowing their content to differ. Each template specifies the probing tar- get, the expertise-related information sought, and the probeâs utility. What-to-Ask strategies select diagnostic targets, whereas How-to-Ask strategies determine how those targets are elicited conversa- tionally. 4.2 Strategy-guided Inquiry Generation Given the populated schemaL t and conversation history H t , PASSING generates the next inquiry in a two-stage manner. First, a What-to-Ask strategy is selected based on the current schema and con- versation history to determine which aspect of the userâs expertise should be further probed. Condi- tioned on the selected strategy and conversation history, a How-to-Ask strategy is then chosen to determine how the inquiry should be phrased. The final inquiry is then generated based on the for- mulated question and conversation history. After receiving the userâs response, PASSING extracts and fills information into the schema based on the user response and conversation history, resulting in an updated schema with filled slots. This pro- cess continues until either all schema slots have been populated or the max conversation turn is reached. At that point, PASSING uses the filled schema and conversation history to estimate the userâs query-specific expertise level. The estimated level, together with the filled schema and conver- sation history, is then used to generate a person- alized response. Notably, the entire pipeline is implemented through the zero-shot capabilities of LLMs, including strategy selection and expertise estimation. While these components could alter- natively be replaced by dedicated trainable mod- ules, the lack of high-quality training data and our exploratory nature of query-specific expertise esti- mation motivate us to adopt a training-free design in this work. We leave more complex trainable formulations to future research. LLM Backbone Method MMLU-ProARC-MCAS Acc.âKappaâFAâComp.âUSâUnk.âAcc.âKappaâFAâComp.âUSâUnk.â Evaluation with LLM Simulators GPT-4o IDL (the best baseline)20.040.313.753.62-23.7920.015.004.593.42-4.40 Direct Prompt16.041.253.843.69-28.0020.030.644.593.67-82.80 PASSING77.164.903.853.834.010.087.471.284.593.874.180.0 PASSING w/o What-to-ask68.664.433.813.873.960.076.867.264.453.864.080.0 PASSING w/o How-to-ask77.164.243.933.863.970.073.771.314.583.844.100.0 DeepSeek V4-Pro IDL (the best baseline)16.045.974.333.08-2.020.055.704.493.62-0.0 Direct Prompt20.058.104.263.69-0.022.056.774.503.77-0.0 PASSING77.151.284.083.884.450.078.952.054.503.944.540.0 PASSING w/o What-to-ask77.150.843.993.874.240.068.456.544.463.924.440.0 PASSING w/o How-to-ask65.754.803.913.814.170.072.657.544.503.904.330.0 Evaluation with Human Participants GPT-4o IDL2040.313.913.413.3723.792015.003.733.473.554.40 PASSING68.752.954.524.624.560.076.776.654.354.744.580.0 DeepSeek V4-Pro IDL1645.973.363.263.312.02055.703.953.353.310.0 PASSING67.356.984.384.564.430.056.038.704.644.554.580.0 Table 3: Main results. Best results are in bold, and second-best results are underlined. 5 Experiments 5.1 Experimental Setup We adopt the setup from Section 3, with the follow- ing key additions to expand our evaluation: Baselines & Backbones. We further involve two LLM backbones (GPT-4o and DeepSeek-V4-Pro) for more comprehensive evaluation. We selected IDL and Direct Prompt for the main comparison. IDL is the strongest baseline overall, with a low rejection rate and stable Kappa scores, while Di- rect Prompt serves as the most fundamental base- line, representing an agent without explicit user- expertise modeling. Multi-turn Conversations. To facilitate the multi- turn conversations required for expertise probing, we employ a two evaluation strategy involving both user simulators and real human participants (cf. Appendix B.4). For simulator-based evaluation, we adopt the same simulation framework used in How-to-Ask strategy induction. To avoid evaluation bias, however, the evaluation simulators are entirely disjoint from those used during strategy induction, with no overlap in personas or refusal strategies. For human evaluation, we recruit participants to interact directly with PASSING under more realistic settings. Additionally, since expertise levels are defined for each userâquery pair rather than for the query alone 7 , for each query, participants are assigned a target expertise level and instructed to role-play accordingly. To ensure consistency with 7 For example, for the same query on quantum mechanics, a computer science PhD student may be a novice, whereas a physics graduate student may be proficient. the assigned query-specific expertise, we provide a knowledge base specifying the concepts expected to be known at that expertise level. Evaluation Metrics. Beyond prediction stability, we evaluate Prediction Accuracy (Acc.) for query- specific expertise estimation using the ground-truth expertise labels available in our multi-turn conver- sations. We additionally assess generated responses in terms of factual accuracy (FA), comprehensibil- ity (Comp.), which evaluates whether responses are appropriately adapted to the userâs expertise level; and user satisfaction (US), which measures the quality of the overall interaction and the ex- tent to which our probing strategy minimizes user burden. Furthermore, we conduct both human and LLM-based evaluations to assess FA, Comp. and US. Refer to Appendix B.6 for details. 5.2 Main Results PASSING consistently establishes superior per- formance across various LLM backbones and datasets. As shown in Table 3, PASSING consis- tently outperforms the strongest baseline across dif- ferent LLM backbones and evaluation benchmarks, achieving an average improvement of +288% in query-level expertise estimation accuracy. Benefit- ing from more accurate estimation of usersâ exper- tise levels, PASSING is able to generate responses that are better tailored to usersâ knowledge bound- aries, substantially improving answer comprehen- sibility (+4.1%). Importantly, despite introducing additional probing interactions with users, PASS- ING maintains factual accuracy comparable to exist- ing baselines, demonstrating that proactive probing does not compromise response correctness. These results suggest that proactively eliciting user exper- tise provides an effective mechanism for delivering personalized responses. PASSING demonstrates relatively reliable per- formance despite the inherent stochasticity of LLM-based probing. Compared with strong base- lines such as IDL and Direct Prompt, PASSING re- duces the unknown response rate to nearly zero while maintaining relatively high prediction stabil- ity (+42.5%, on average), indicating that the ob- served performance gains are not driven by random fluctuations. We also observe that PASSING does not always achieve the highest stability scores. We attribute this to the inherent stochasticity of LLMs in both selecting probing strategies across both the what-to-ask and how-to-ask stages, and estimating user expertise under the predefined schema. Al- though such variance does not diminish the overall advantage of PASSING over existing methods, it highlights a promising direction for future work: improving system stability through dedicated mod- ules for each stage, especially when supervised training data are available. Probing strategies in PASSING improve user modeling and help maintain user satisfac- tion. Beyond expertise estimation accuracy, PASS- ING achieves promising user satisfaction scores across different datasets and LLM backbones. Ab- lation results further show that removing either the What-to-Ask or How-to-Ask strategy leads to per- formance degradation, highlighting the importance of both selecting appropriate probing content and generating effective probing interactions. In par- ticular, removing the How-to-Ask strategy consis- tently decreases user satisfaction, suggesting that the manner of questioning substantially affects user experience. Interestingly, removing the What-to- Ask strategy also harms satisfaction. We find that, without explicit guidance on probing content, the generated questions become less controllable and occasionally mismatch user expertise levels: overly difficult questions tend to frustrate users, while sim- pler questions are generally better tolerated. Human evaluation further validates the effec- tiveness of PASSING. As shown in Table 3, PASS- ING consistently outperforms IDL across both two backbones under human participants and human evaluation. In particular, PASSING achieves sub- stantially higher expertise estimation accuracy and user satisfaction while reducing the unknown re- sponse rate to nearly zero. Meanwhile, the fac- S-01S-02S-03S-04S-05S-06S-07S-08S-09S-10S-11S-12S-13S-14S-15S-16S-17S-18S-19S-20S-21 S-01 S-02 S-03 S-04 S-05 S-06 S-07 S-08 S-09 S-10 S-11 S-12 S-13 S-14 S-15 S-16 S-17 S-18 S-19 S-20 S-21 Strategy similarity heatmap (text-embedding-3-small) 0.0 0.2 0.4 0.6 0.8 1.0 Cosine similarity Figure 3: Pairwise similarity among induced strategies (13 What-to-ask and 8 How-to-ask strategies in total) tual accuracy and completeness of generated re- sponses remain competitive, demonstrating that proactive probing can effectively improve personal- ized interactions without sacrificing response qual- ity. Following Bhagwatkar et al. (2025), Zhang et al. (2025), and Irawan et al. (2025), we further examine the consistency between human evaluators and LLM-based evaluators on human interaction data using Gwetâs AC2 score (Gwet, 2008). The humanâhuman AC2 reaches an average of 0.72, confirming consistent annotation quality across par- ticipants. The humanâLLM AC2 reaches an aver- age of 0.79, indicating strong agreement between human judgments and LLM-based evaluation re- sult. Detailed reliability computation and results are provided in Appendix B.8. 5.3 In-depth Analysis Induced strategies reveal a functional rather than topical organization. We further examine whether the induced What-to-Ask and How-to-Ask strategies occupy distinct regions in embedding space. Using the OpenAI text-embedding-3-small, we encode all induced strategies and compute pair- wise cosine similarities 8 (Figure 3). The two cate- gories are not semantically separated: the average intra-class similarities are 0.608 for What-to-Ask and 0.604 for How-to-Ask, while the inter-class similarity is 0.604, with a silhouette score of 0.006. This suggests that the two categories differ primar- ily in functional role rather than topical content: What-to-Ask strategies specify diagnostic targets, whereas How-to-Ask strategies operationalize these 8 More details are in Appendix B.7. 351013 50 60 70 80 90 Accuracy (%) What-to-ask Strategies 2468 50 60 70 80 90 Accuracy (%) How-to-ask Strategies 351013 3.5 4.0 4.5 US What-to-ask Strategies 2468 3.5 4.0 4.5 US How-to-ask Strategies GPT-4o / MMLU-Pro GPT-4o / ARC-MCAS DeepSeek V4-Pro / MMLU-Pro DeepSeek V4-Pro / ARC-MCAS Figure 4: Effect of the number of probing strategies on expertise estimation accuracy and user satisfaction. targets through conversational tactics. Maintaining sufficient strategies ensures the ef- fectiveness of proactive probing. We further in- vestigate the impact of the number of probing strate- gies on both expertise estimation accuracy and user satisfaction in PASSING. Specifically, given 13 What-to-Ask strategies and 8 How-to-Ask strate- gies, we randomly samplekstrategies from one category while keeping the other category com- plete, and evaluate the resulting performance. As shown in Figure 4, increasing the number of strate- gies generally improves performance, suggesting that richer probing diversity helps the model better capture user expertise and interaction preferences. 6 Conclusion We introduce PASSING, a proactive probing method that enables LLM agents to infer query-specific user expertise before generating responses. This is achieved through our LLM-induced two strategy sets. Generally speaking, our findings suggest that effective personalization should not solely depend on passively modeling users from limited observa- tions, but also on actively acquiring missing infor- mation through interaction. In this sense, proactive probing serves as a lightweight yet scalable mecha- nism in real-world interactions. We hope this work can inspire future research on interaction-centric personalization, adaptive information acquisition, and more human-aware LLM agents. Limitations As with prior studies on proactive agents by prompt- ing LLMs, the performance of PASSING may be influenced by prompt design. Prompt sensitiv- ity remains an inherent challenge for LLM-based systems, and exploring more robust prompting or training-based alternatives constitutes an important direction for future work. Another limitation of this work is that we do not systematically investi- gate the impact of self-play hyper-parameters, such as the number of self-play sessions or the maxi- mum conversation turns, on the quality of induced probing strategies. Although we analyze the ef- fect of the number of strategies in our experiments, the strategy induction process itself remains under- explored. Understanding how different self-play configurations influence strategy diversity, quality, and downstream personalization performance is an important direction for future work. Finally, al- though What-to-Ask and How-to-Ask strategies serve distinct roles, they inevitably exhibit seman- tic overlap. As an early study of query-specific expertise probing, we leave systematic disentangle- ment to future work. One promising approach is to condition second-stage extraction on first-stage strategies and explicitly discourage semantic redun- dancy. Ethical Considerations Similar to prior work (Zhang et al., 2024b; Chen et al., 2024b; Zhang et al., 2018, 2022; Chen et al., 2023), the primary experiments in this paper rely on LLM-based user simulators for large-scale evalua- tion. In addition, we conduct supplementary human evaluation experiments with voluntary participants. All participants were informed of the experimental setting and engaged in standard question-answering interactions without sensitive topics or psycholog- ically stressful tasks. Therefore, we believe the study poses minimal ethical risk. LLM Usage LLMs are used in this work as the backbone mod- els for conversational agents and evaluators in our experiments. In addition, LLMs are also used for language polishing during paper writing. References Rishika Bhagwatkar, Syrielle Montariol, Angelika Ro- manou, Beatriz Borges, Irina Rish, and Antoine Bosselut. 2025. Cave: Detecting and explaining com- monsense anomalies in visual environments. In Pro- ceedings of the 2025 Conference on Empirical Meth- ods in Natural Language Processing, pages 27110â 27151. Pavel Braslavski, Denis Savenkov, Eugene Agichtein, and Alina Dubatovka. 2017. What do you mean exactly? analyzing clarification questions in cqa. In Proceedings of the 2017 conference on conference human information interaction and retrieval, pages 345â348. Chengkun Cai, Xu Zhao, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, and Lei Li. 2025. The role of deductive and inductive reasoning in large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16780â16790. Jieming Cao, Chen Huang, Yanan Zhang, Ruibo Deng, Jincheng Zhang, and Wenqiang Lei. 2025. Breaking the stigma! unobtrusively probe symptoms in depres- sion disorder diagnosis dialogue. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 182â200, Albuquerque, New Mexico. Association for Computational Linguistics. Avyay Casheekar, Archit Lahiri, Kanishk Rath, Kaushik Sanjay Prabhakar, and Kathiravan Srini- vasan. 2024. A contemporary review on chatbots, ai-powered virtual conversational agents, chatgpt: Applications, open challenges and future research directions. Computer Science Review, 52:100632. John Chen, Xi Lu, Yuzhou Du, Michael Rejtig, Ruth Bagley, Mike Horn, and Uri Wilensky. 2024a. Learn- ing agent-based modeling with llm companions: Ex- periences of novices and experts using chatgpt & netlogo chat. In Proceedings of the 2024 CHI Con- ference on Human Factors in Computing Systems, pages 1â18. Yue Chen, Chen Huang, Yang Deng, Wenqiang Lei, Dingnan Jin, Jia Liu, and Tat-Seng Chua. 2024b. STYLE: Improving domain transferability of ask- ing clarification questions in large language model powered conversational agents. In Findings of the As- sociation for Computational Linguistics: ACL 2024, pages 10633â10649, Bangkok, Thailand. Association for Computational Linguistics. Yue Chen, Dingnan Jin, Chen Huang, Jia Liu, and Wen- qiang Lei. 2023. Travel: Tag-aware conversational faq retrieval via reinforcement learning. In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3861â3872. Chuanqi Cheng, Quan Tu, Wei Wu, Shuo Shang, Cunli Mao, Zhengtao Yu, and Rui Yan. 2024. âin-dialogues we learnâ: Towards personalized dialogue without pre-defined profiles through in-dialogue learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10408â10422, Miami, Florida, USA. Association for Computational Linguistics. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anas- tasios N. Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. Chat- bot arena: an open platform for evaluating llms by human preference. In Proceedings of the 41st Inter- national Conference on Machine Learning, ICMLâ24. JMLR.org. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question an- swering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457. Sarkar Snigdha Sarathi Das, Chirag Shah, Mengting Wan, Jennifer Neville, Longqi Yang, Reid Andersen, Georg Buscher, and Tara Safavi. 2024. S3-dst: Struc- tured open-domain dialogue segmentation and state tracking in the era of llms. In Findings of the As- sociation for Computational Linguistics: ACL 2024, pages 14996â15014. JoĂŁo Pedro Gandarela de Souza, Danilo Carvalho, and AndrĂŠ Freitas. 2025. Inductive learning of logical theories with llms: A expressivity-graded analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23752â23759. Yang Deng, Lizi Liao, Liang Chen, Hongru Wang, Wenqiang Lei, and Tat-Seng Chua. 2023. Prompt- ing and evaluating large language models for proac- tive dialogues: Clarification, target-guided, and non- collaboration. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10602â10621, Singapore. Association for Compu- tational Linguistics. Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, and Tat-Seng Chua. 2025. Proactive con- versational ai: A comprehensive survey of advance- ments and opportunities. ACM Transactions on In- formation Systems, 43(3):1â45. Yang Deng, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, and Tat-Seng Chua. 2024. Towards human- centered proactive conversational agents. In Proceed- ings of the 47th International ACM SIGIR Confer- ence on Research and Development in Information Retrieval, pages 807â818. Stuart E Dreyfus and Hubert L Dreyfus. 1980. A five- stage model of the mental activities involved in di- rected skill acquisition. Technical report. Paul J Feltovich and Robert R Hoffman. 1997. Expertise in context. AAAI Press Menlo Park, CA. Joseph L Fleiss. 1971. Measuring nominal scale agree- ment among many raters. Psychological bulletin, 76(5):378â382. Lewis R Goldberg. 1992. The development of mark- ers for the big-five factor structure. Psychological assessment, 4(1):26. Zhizhao Guan, Chen Huang, Ziming Liu, Hongru Liang, Wenqiang Lei, See-Kiong Ng, Tat-Seng Chua, and Anthony G Cohn. 2026. Clearing the fog: Towards installing and refining proactive exploration capabili- ties in llm agents. Preprint, arXiv:2608.14339. Kilem Li Gwet. 2008. Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psy- chology, 61(1):29â48. Chen Huang, Zitan Jiang, Zou Changyi, Wenqiang Lei, and See-Kiong Ng. 2026. Towards proactive informa- tion probing: Customer service chatbots harvesting value from conversation. In Findings of the Associa- tion for Computational Linguistics: ACL 2026, pages 28768â28792, San Diego, California, United States. Association for Computational Linguistics. Chen Huang, Yiping Jin, Ilija Ilievski, Wenqiang Lei, and Jiancheng Lv. 2024. Araida: Analogical reasoning-augmented interactive data annotation. In Proceedings of the 62nd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 10660â10675. Patrick Amadeus Irawan, Genta Indra Winata, Samuel Cahyawijaya, and Ayu Purwarianti. 2025. Towards efficient and robust VQA-NLE data generation with large vision-language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 4323â4340, Abu Dhabi, UAE. As- sociation for Computational Linguistics. Jakub Jeck, Florian Leiser, Anne HĂźsges, and Ali Sun- yaev. 2025. Tell-me: Toward personalized explana- tions of large language models. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA â25, New York, NY, USA. Association for Computing Machinery. Namyoung Kim, Kai Tzu-iunn Ong, Yeonjun Hwang, Minseok Kang, Iiseo Jihn, Gayoung Kim, Minju Kim, and Jinyoung Yeo. 2025. Principles: Synthetic strat- egy memory for proactive dialogue agents. In Find- ings of the Association for Computational Linguistics: EMNLP 2025, pages 21329â21368. Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive conversational agents in the post-chatgpt world. In Proceedings of the 46th international ACM SIGIR conference on research and development in information retrieval, pages 3452â3455. Qijiong Liu, Nuo Chen, Tetsuya Sakai, and Xiao-Ming Wu. 2024. Once: Boosting content-based recom- mendation with both open- and closed-source large language models. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM â24, page 452â461, New York, NY, USA. Association for Computing Machinery. Yaxi Lu, Shenzhi Yang, Cheng Qian, Guirong Chen, Qinyu Luo, Yesai Wu, Huadong Wang, Xin Cong, Zhong Zhang, Yankai Lin, and 1 others. 2025. Proac- tive agent: Shifting llm agents from reactive re- sponses to active assistance. In International Con- ference on Learning Representations, volume 2025, pages 47431â47457. Elaine M Silva Mangiante and Kathy Peno. 2021. Teaching and learning for adult skill acquisition: ap- plying the Dreyfus and Dreyfus model in different fields. IAP. Sewon Min, Victor Zhong, Luke Zettlemoyer, and Han- naneh Hajishirzi. 2019. Multi-hop reading compre- hension through question decomposition and rescor- ing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6097â6109, Florence, Italy. Association for Compu- tational Linguistics. Shramay Palta, Nirupama Chandrasekaran, Rachel Rudinger, and Scott Counts. 2025. Speaking the right language: The impact of expertise alignment in user- ai interactions. arXiv preprint arXiv:2502.18685. Roger C Schank. 1999. Dynamic memory revisited. Cambridge University Press. Christiane C Schubert, T Kent Denmark, Beth Crandall, Anna Grome, and James Pappas. 2013. Character- izing novice-expert differences in macrocognition: an exploratory study of cognitive work in the emer- gency department. Annals of emergency medicine, 61(1):96â109. Susanne G Scott and Reginald A Bruce. 1995. Decision- making style: The development and assessment of a new measure. Educational and psychological mea- surement, 55(5):818â831. Chenkai Sun, Ke Yang, Revanth Gangi Reddy, Yi Fung, Hou Pong Chan, Kevin Small, ChengXiang Zhai, and Heng Ji. 2025. Persona-db: Efficient large lan- guage model personalization for response prediction with collaborative data refinement. In Proceedings of the 31st International Conference on Computational Linguistics, pages 281â296. Andrew A Tawfik, Jessica D Gatewood, Jaclyn J Gish- Lieberman, and Charles W Keene. 2021. Explor- ing the differences between experts and novices on inquiry-based learning cases. Journal of Formative Design in Learning, 5(2):97â105. Borui Wang, Kathleen McKeown, and Rex Ying. 2025. Dystil: Dynamic strategy induction with large lan- guage models for reinforcement learning. arXiv preprint arXiv:2505.03209. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, and 1 others. 2024. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Ad- vances in Neural Information Processing Systems, 37:95266â95290. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elic- its reasoning in large language models. Preprint, arXiv:2201.11903. Zeqiu Wu, Ryu Parish, Hao Cheng, Sewon Min, Prithvi- raj Ammanabrolu, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. InSCIt: Information-seeking con- versations with mixed-initiative interactions. Trans- actions of the Association for Computational Linguis- tics, 11:453â468. Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. 2025.Expertprompting: Instructing large lan- guage models to be distinguished experts. Preprint, arXiv:2305.14688. Jingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, and Xu Sun. 2019. Asking clarification questions in knowledge- based question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natu- ral Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1618â1629, Hong Kong, China. Association for Computational Linguistics. Ramin Yaghoubzadeh and Stefan Kopp. 2017. Enabling robust and fluid spoken dialogue with cognitively impaired users. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 273â283, SaarbrĂźcken, Germany. Association for Computational Linguistics. Zonghai Yao, Nandyala Siddharth Kantu, Guanghao Wei, Hieu Tran, Zhangqi Duan, Sunjae Kwon, Zhichao Yang, and Hong Yu. 2024. README: Bridging medical jargon and lay understanding for patient education through data-centric NLP. In Find- ings of the Association for Computational Linguistics: EMNLP 2024, pages 12609â12629, Miami, Florida, USA. Association for Computational Linguistics. Hamed Zamani, Bhaskar Mitra, Everest Chen, Gord Lueck, Fernando Diaz, Paul N. Bennett, Nick Craswell, and Susan T. Dumais. 2020. Analyzing and learning from user interactions for search clarifi- cation. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR â20, page 1181â1190, New York, NY, USA. Association for Computing Machinery. Hamed Zamani, Johanne R Trippas, Jeff Dalton, Filip Radlinski, and 1 others. 2023. Conversational infor- mation seeking. Foundations and TrendsÂŽ in Infor- mation Retrieval, 17(3-4):244â456. Haiqi Zhang, Zhengyuan Zhu, Zeyu Zhang, and Chengkai Li. 2025. LLMTaxo: Leveraging large lan- guage models for constructing taxonomy of factual claims from social media. In Findings of the Asso- ciation for Computational Linguistics: ACL 2025, pages 19627â19641, Vienna, Austria. Association for Computational Linguistics. Tong Zhang, Chen Huang, Yang Deng, Hongru Liang, Jia Liu, Zujie Wen, Wenqiang Lei, and Tat-Seng Chua. 2024a. Strength lies in differences! improv- ing strategy planning for non-collaborative dialogues via diversified user simulation. In Proceedings of the 2024 Conference on Empirical Methods in Nat- ural Language Processing, pages 424â444, Miami, Florida, USA. Association for Computational Lin- guistics. Tong Zhang, Peixin Qin, Yang Deng, Chen Huang, Wen- qiang Lei, Junhong Liu, Dingnan Jin, Hongru Liang, and Tat-Seng Chua. 2024b. CLAMBER: A bench- mark of identifying and clarifying ambiguous infor- mation needs in large language models. In Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 10746â10766, Bangkok, Thailand. As- sociation for Computational Linguistics. Yiming Zhang, Lingfei Wu, Qi Shen, Yitong Pang, Zhi- hua Wei, Fangli Xu, Bo Long, and Jian Pei. 2022. Multiple choice questions based multi-interest policy learning for conversational recommendation. In Pro- ceedings of the ACM Web Conference 2022, W â22, page 2153â2162, New York, NY, USA. Associa- tion for Computing Machinery. Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W. Bruce Croft. 2018. Towards conversational search and recommendation: System ask, user respond. In Proceedings of the 27th ACM International Confer- ence on Information and Knowledge Management, CIKM â18, page 177â186, New York, NY, USA. As- sociation for Computing Machinery. A Background on User Expertise A.1 Dreyfus Model The Dreyfus Model (Dreyfus and Dreyfus, 1980) describes skill acquisition as a progression through five stages, each characterized by distinct differ- ences across five dimensions: knowledge scope, components, perspective, action, and commitment. Novices possess only explicit factual knowledge, rely entirely on context-free rules, are easily over- whelmed by complex information, reason purely step-by-step, and feel little personal investment in outcomes. Advanced Beginners have slightly broader knowledge but still treat all information as equally relevant and struggle to apply rules to real-world situations. Competent performers mark a turning point: they consciously plan and orga- nize information to filter out noise, and begin to feel personally responsible for their decisions. Pro- ficient users have accumulated rich experiential knowledge that allows them to instantly recognize what matters in a situation without deliberate effort, though they still verify their responses through an- alytic reasoning. Experts have fully internalized their domain knowledge, rather than recalling facts or following rules, they perceive situations holisti- cally and respond immediately and intuitively. A.2 Response Features Required by Experts and Novices The substantial differences across Dreyfus lev- els imply that responses must be tailored accord- ingly. Novices require short, jargon-free responses that explicitly explain basic concepts in simple language, as they can only process context-free facts and are easily overwhelmed by complex in- formation. Experts, by contrast, expect dense, conclusion-first, telegraphic responses that use do- main jargon without explanation and omit all foun- dational scaffolding. B Experiment Details We use Python 3.10.12, NumPy 1.26.4, and scikit- learn 1.3.2 for implementation and evaluation, in- cluding the computation of Fleissâ Kappa (Fleiss, 1971). For preliminary experiments, we randomly sample 50 user queries from the MMLU-Pro and ARC-MCAS datasets. For the main experiments, we use the full datasets. B.1 Details of Datasets MMLU-Pro (Wang et al., 2024) is a challeng- ing multi-task language understanding benchmark spanning 14 academic disciplines. We use 70 ques- tions from the test split for offline self-play strategy induction, and sample 50 questions from the vali- dation split for evaluation to avoid data leakage. ARC-MCAS (Clark et al., 2018) is a subset of the ARC-Challenge benchmark, consisting of science questions from the Massachusetts Comprehensive Assessment System (MCAS). We sample 95 ques- tions from the test split for evaluation. B.2 Implementation Details of Baselines Direct Prompt. The LLM is directly prompted to predict the userâs query-specific expertise level given the input query, without any additional rea- soning or ensemble steps. Chain-of-Thought (CoT). The LLM is prompted to reason step-by-step before predicting the userâs expertise level (Wei et al., 2023). Self-consistency. Multiple predictions are sam- pled independently under different random seeds, and the final prediction is determined by majority voting (Wang et al., 2022). Diverse-Aspect. The LLM first identifies multiple aspects of the query that are relevant to expertise assessment (e.g., terminology use, reasoning style). For each aspect, an independent expertise judg- ment is generated. These aspect-level assessments are then integrated into a final cohesive expertise prediction. IDL. The LLM is provided with pre-collected di- alogue history from the same user as reference. It extracts structured knowledge triples from the his- tory to form a persona constraint, and uses this constraint to assess the userâs expertise level on the new query (Cheng et al., 2024). ExpertPrompting. The LLM is prompted with few-shot examples to assess the userâs expertise level directly from the input query, and generates a tailored response accordingly (Xu et al., 2025). B.3 Implementation Details of PASSING As emphasized in prior conversational AI research, agent-initiated probing must account for both usersâ willingness to respond and the interaction burden it imposes. Although real users do engage with LLM agents and system-initiated questions in realistic information-seeking scenarios (Deng et al., 2023; Zamani et al., 2020; Chiang et al., 2024), demon- strating the practical relevance of this research set- ting, such engagement should not be assumed to be unconditional or cost-free. Our problem for- mulation should therefore explicitly model usersâ willingness to answer probing questions (i.e., How- to-Ask and What-to-Ask). In particular, our PASS- ING aims to provide responses tailored to the userâs query-specific expertise while maintaining user sat- isfaction. This is achieved by inducing two com- plementary strategy sets: What-to-Ask selects the most diagnostic expertise evidence, while How-to- Ask phrases the probe in a natural and low-burden way. To enhance the induction of What-to-ask strate- gies, we further involve a knowledge dependency graph into the PASSINGâs prompt. Given a user queryqand the model-generated answer 9 a, PASS- ING employs a structured approach to identify and represent the prerequisite knowledge required for comprehending the answer via prompts. During this process, key domain-specific concepts, includ- 9 The specific technique used for answer generation (e.g., retrieval augmentation or asking clarifying questions) is not the primary focus of our work. ing terminology entities and actions, are extracted as nodes, while their semantic relationships form the edges. Finally, a knowledge dependency graph Gis constructed. Note that each node is addition- ally assigned a min-level attribute indicating the minimum expertise level required to understand that concept; this attribute is used by the user sim- ulator to determine which concepts are known or unknown to a simulated user, but is masked from PASSING during probing to prevent information leakage. In particular, our masking mechanism per- forms content-based masking according to concept difficulty. Each node in the knowledge dependency graph has a min-level attribute indicating the min- imum Dreyfus level required to understand that concept. For a simulator assigned levell, we retain nodes whose min-level is no higher thanland mask the remaining nodes. Thus, a Novice retains only Novice-level concepts, an Advanced Beginner re- tains concepts from the first two levels, and so forth, while an Expert retains all concepts. Since graph size and level distribution vary by query, there is no fixed retention ratio for each level. The min-level attribute is visible only to the user simulator and hidden from PASSING during probing to prevent information leakage. B.4 Implementation Details of User Simulators and Human Participants B.4.1 User Simulators for Strategy Induction We detail the expertise Profile for different type of user as follows â˘The Novice operates with a Knowledge Scope limited to explicit facts and definitions, lacking any deep understanding of "how" things work in practice. They perceive Components strictly as context-free features, relying entirely on rigid rules and abstract principles regardless of the situ- ation. Lacking a functional Perspective, they are often unable to filter information effectively, forc- ing them to adhere strictly to instructions to avoid becoming overwhelmed. Their Action is purely analytic, characterized by slow, step-by-step rea- soning to execute instructions. Consequently, their Commitment remains detached; they view themselves as simply following the rules and feel little personal agency or responsibility for the outcome. ⢠The Advanced Beginner has expanded their Knowledge Scope slightly but still relies heav- ily on explicit facts. While they begin to rec- ognize some situational Components alongside context-free rules, they lack the experience to distinguish importance. As a result, their Per- spective is flawed; they treat all information as equally relevant, which leads them to be easily overwhelmed by the noise of a complex situation. Their Action remains analytic and deliberate as they struggle to apply guidelines to real-world scenarios. Like the Novice, their Commitment is largely detached, as they are focused on surviving the task rather than owning the result. â˘The Competent performer marks a major shift in Perspective. Faced with a complex situation, they do not get overwhelmed; instead, they con- sciously make a deliberate plan to organize in- formation and filter out noise. Their Knowledge Scope is now organized enough to handle this structuring. While they perceive both context- free and situational Components, they prioritize them based on a hierarchical goal. Their Action is still analyticâthey solve problems by reason- ing through their plan step-by-step. However, their Commitment shifts to being involved; be- cause they chose the plan, they feel a sense of personal responsibility and emotional investment in the outcome. â˘The Proficient user possesses a vast Knowledge Scope of tacit knowledge derived from experi- ence. Their perception of Components is now dominated by situational patterns rather than rules. Crucially, their Perspective has evolved from "making a plan" to "seeing the picture"; they instinctively distinguish important informa- tion from noise without conscious effort. Despite this intuitive perception, their Action remains an- alytic; they effectively diagnose the problem in- stantly but still rely on step-by-step reasoning to calculate the best specific response. Their Com- mitment is deeply involved in the perception of the problem, though they may remain analyti- cally detached when executing the solution. â˘The Expert operates with a deep, tacit Knowl- edge Scope where knowing "that" has completely transformed into knowing "how." They perceive Components entirely through situational patterns, reading the environment holistically. Their Per- spective is immediate and instinctive; important information jumps out at them, and they are never overwhelmed. The defining difference lies in their Action: it shifts from analytic to intuitive. They no longer rely on step-by-step reasoning to decide; the correct response arises immediately and fluidly. Their Commitment is fully involved, acting with a level of agency where the person and the task are effectively merged. What-to-Ask vs. How-to-Ask Simulator Design: â˘While both simulators instantiate users accord- ing to the five Dreyfus profiles described above, they differ in design. The What-to-Ask simula- tor injects only the expertise-level profile and knowledge boundaries (known/unknown con- cepts), and the simulatorâs sole task is to gen- erate a response consistent with the cognitive style of that level; responses are output as plain text with no explicit decision process, yielding five fixed behavioral distributions. The How- to-Ask simulator builds on this foundation by overlaying a personality trait drawn from the Big Five model (Goldberg, 1992) and a decision- making style drawn from four archetypes (Scott and Bruce, 1995), allowing users at the same expertise level to exhibit markedly different com- municative styles. It further introduces two addi- tional mechanisms: a level-specific refusal pat- tern that characterizes how users at each expertise level typically decline to answer, and an explicit answer/refuse decision step that requires the sim- ulator to weigh reasons for answering against rea- sons for refusing before producing a structured output containing both a decision and a response. The resulting training dialogues cover uncoop- erative behaviors including deflection, minimal engagement, and outright refusal, forcing the how-to-ask strategies to learn how to adapt their phrasing when users are not forthcoming. B.4.2 User Simulators for Experimental Evaluation For the main experiments, we employ GPT-4o Mini as the user simulator. Each simulator instance is initialized with a pre-defined expertise profile com- prising: (1) an expertise level drawn from the five Dreyfus levels; (2) a persona description includ- ing personality trait based on the Big Five model (Goldberg, 1992) and decision-making style (Scott and Bruce, 1995); (3) a set of known and unknown concepts derived from the knowledge dependency graph. To better simulate realistic user behavior, each expertise level is assigned a corresponding refusal pattern that allows the simulator to au- tonomously decide whether to answer or refuse a given inquiry. The simulator strictly respects its knowledge boundaries by never displaying knowl- edge of unknown concepts, enabling diverse and realistic simulation of users across different exper- tise levels. B.5 Implementation Details of Evaluation Metrics Prediction Stability (Kappa). We measure pre- diction stability using Fleissâ Kappa (Fleiss, 1971), defined as: Îş = Ě P â Ě P e 1â Ě P e , where Ě Pdenotes the mean observed agreement across all queries, and Ě P e denotes the expected agreement under chance. To quantify cross-run pre- diction consistency, we treat predictions produced under different random seeds as independent raters. The resulting Kappa score reflects how stable the method is across repeated runs. Factual Accuracy (FA). We employ GPT-5 as a third-party evaluator to assess the factual accuracy of the tailored responses. The evaluation covers four dimensions: core answer correctness, detail correctness, no factual error, and no misleading, each scored on a 1â5 scale. The final FA score is the average across all four dimensions. Comprehension (Comp.). Comprehension is eval- uated by the user simulator (GPT-4o Mini) from the perspective of the simulated user, initialized with a pre-defined expertise profile including known and unknown concepts. The simulator assesses four di- mensions: clarity, level match, completeness, and accessibility, each scored on a 1â5 scale. The final Comp. score is the average across all four dimen- sions. User Satisfaction (US). We employ GPT-5 as a third-party evaluator to assess user satisfaction based on the full conversation history. The eval- uation covers five dimensions: comfort, feeling understood, willingness to engage, conversation naturalness, and overall satisfaction, each scored on a 1â5 scale. The final US score is the average across all five dimensions. B.6 Implementation Details of Human Evaluation B.6.1 Role-playing Protocol We recruited ten graduate students specializing in computer science and artificial intelligence to par- ticipate in the human evaluation. All participants voluntarily took part in the study. Each participant role-played users at five exper- tise levels defined by the Dreyfus Model: Novice, Advanced Beginner, Competent, Proficient, and Expert. The expertise profiles used for role-playing follow the same Dreyfus-based descriptions as those used in our user simulators for strategy induc- tion, as detailed in Appendix B.4.1. Each partici- pant completed 10 conversations per level, yield- ing 50 conversations per participant. Experiments were conducted on two datasets, MMLU-Pro and ARC-MCAS, and two LLM backbones, GPT-4o and DeepSeek-V4-Pro. To ensure consistent role-playing behavior across participants, we provided participants with the expertise-level guide described in Ap- pendix B.4.1, together with representative example responses. This design allowed participants to fo- cus on controlling the reasoning style, uncertainty, and response granularity associated with each tar- get level. After each conversation, participants evaluated the systemâs tailored response on three dimensions using 5-point Likert scales. Factual Accuracy (FA, 4 items) measures factual consistency between the system response and a provided reference answer, independent of expertise-level adaptation. Compre- hensibility (Comp, 4 items) assesses whether the response is clear and understandable for the role- played user level. User Satisfaction (US, 5 items) measures whether the response is appropriate, help- ful, and well aligned with the participantâs assigned expertise level. B.7 Implementation Details of Experiment Semantic analysis of induced strategies. We provide additional details for the semantic analysis in Section 5.3. We encode all 21 induced strate- gies using an OpenAI-compatible embedding end- point based on text-embedding-3-small, obtaining a 1536-dimensional embedding for each strategy. The strategy set contains 13 What-to-Ask strategies and 8 How-to-Ask strategies. We compute pairwise cosine similarities and report intra-class similarity, inter-class similarity, the diversity gap, and the sil- houette score under the What-to-Ask/How-to-Ask labeling. The diversity gap is defined as: â = Sim What-to-Ask + Sim How-to-Ask 2 â Sim inter , MetricValue What-to-Ask intra-class similarity0.608 How-to-Ask intra-class similarity0.604 Inter-class similarity0.604 Diversity gap â0.003 Silhouette score0.006 Table 4: Semantic similarity statistics for induced strate- gies. whereSim What-to-Ask andSim How-to-Ask are the aver- age within-category cosine similarities excluding diagonal self-similarities, andSim inter is the aver- age cross-category cosine similarity. The closest cross-category pairs are S-03/S-17 (0.825), S-01/S-16 (0.807), and S-02/S-19 (0.759), corresponding to conceptual distinction, rationale probing, and concrete application, respectively. These nearest-neighbor pairs support the interpreta- tion that How-to-Ask strategies specify How-to-Ask questions that elicit evidence for diagnostic targets defined by What-to-Ask strategies. B.8 Human Evaluation Reliability Setup.To assess the reliability of the human eval- uation described in Section B.6.1, we compute inter-rater agreement from two perspectives. First, we measure humanâhuman agreement among the human participants to verify that the evaluation criteria were applied consistently. Second, we mea- sure humanâLLM agreement between the human ratings and the LLM-based evaluator to examine whether the automatic evaluator aligns with human judgment. Reliability is computed separately for the three post-conversation evaluation dimensions: factual accuracy (FA), comprehensibility (Comp.), and user satisfaction (US). Metric and computation. We use ordinal- weighted Gwetâs AC2 (Gwet, 2008) as the reli- ability metric. AC2 is a chance-corrected agree- ment coefficient that compares the observed agree- ment among raters with the agreement expected by chance: AC2 = A o â A e 1â A e , whereA o denotes the observed weighted agree- ment andA e denotes the chance-expected agree- ment estimated under Gwetâs occasional-guessing model. Since our evaluation uses 1â5 Likert-scale scores, we apply ordinal weights so that larger score differences incur larger disagreement penal- ties. For humanâhuman reliability, AC2 is computed over all independent human ratings for each evalu- ated session. For humanâLLM reliability, we com- pare the LLM-based evaluatorâs score with the cor- responding aggregated human score for each ses- sion. Since each evaluation dimension consists of multiple Likert-scale sub-items, we first aggregate the sub-items within each dimension and then com- pute AC2 on the resulting dimension-level scores. Results. As shown in Table 6, humanâhuman AC2 ranges from 0.52 to 0.87 (Îź = 0.72), indi- cating moderate-to-strong agreement among par- ticipants. HumanâLLM AC2 ranges from 0.74 to 0.85 (Îź = 0.79), which is comparable to humanâ human agreement and suggests that the LLM-based evaluator is well aligned with human judgment. Furthermore, Table 5 reports bootstrap 95% con- fidence intervals as error bars for the US scores in Table 3. We conduct participant-level paired analysis: for each participant, we first aggregate the US scores under IDL and PASSING then com- pute paired differences. We use an exact paired sign-flip test, with Holm correction across the four backboneâdataset settings. Across all four settings, PASSING achieves significantly higher US than IDL, with paired differences ranging from 1.028 to 1.280; all 95% CIs are strictly positive, and the Holm-corrected p-values are 0.007812. This confirms that PASSING Ě s user satisfaction is sig- nificantly higher than IDL under participant-level paired analysis. B.9 Additional Analysis For better understanding, we encode the five Drey- fus levels from 0 to 4 and compute mean absolute error between predicted and ground-truth levels. On MMLU-Pro and ARC-MCAS, the MAEs are 0.343 and 0.147 for GPT-4o, and 0.243 and 0.211 for DeepSeekV4-Pro, respectively. When pooling both datasets, the MAEs are 0.230 for GPT-4o and 0.224 for DeepSeekV4-Pro. Overall, 96.97% of predictions are within one level of the ground truth. We will report these ordinal metrics in the revised manuscript. We also include the corresponding 5Ă5 confu- sion matrices in Table 7, where rows denote ground- truth Dreyfus levels and columns denote predicted levels in the order Novice, Advanced Beginner, Competent, Proficient, Expert. C Case Study We provide a complete case study illustrating the full PASSING pipeline: initial queryâprobing questionsâuser responsesâschema fillingâ expertise estimationâ tailored response. Query: âWhy can large-aperture telescopes see fainter objects?â Turn 1 â˘What-to-Ask: S-03 (The strategy descrip- tions are provided at the end) ⢠How-to-Ask: S-14 ⢠Inquiry: âThis touches on a bit of optics â no worries if itâs unfamiliar! Do you have a sense of what light-gathering power ac- tually refers to, or how it differs from mag- nification?â ⢠User response: âI know a bigger aperture lets you see more, but I always thought that was mainly about zooming in, not really about brightness.â â˘Schema update: knowledge_scopeâex- plicit; perspectiveâ overwhelmed Turn 2 ⢠What-to-Ask: S-02 ⢠How-to-Ask: S-18 â˘Inquiry: âHereâs a comparison: if you wanted to collect more rainwater, what would you do? How might that relate to how a telescope collects light?â â˘User response: âUm, maybe a wider open- ing catches more rain? So maybe a bigger telescope catches more light too, I guess.â â˘Schema update: componentsâcon- text_free; actionâintuitive; commitment â detached Termination and Estimation All five schema slots are populated within 2 turns. PASSING estimates the user as Advanced Be- ginner. Final Tailored Answer Youâre on the right track! A telescopeâs light-gathering power depends on the area of its aperture, not just SettingPASSING US (95% CI)Difference (95% CI)Holm p GPT-4o / MMLU-Pro4.580 [4.540, 4.621]1.224 [1.073, 1.360]0.007812 GPT-4o / ARC-MCAS4.583 [4.573, 4.594]1.028 [0.999, 1.055]0.007812 DeepSeekV4-Pro / MMLU-Pro4.423 [4.408, 4.436]1.146 [1.056, 1.234]0.007812 DeepSeekV4-Pro / ARC-MCAS4.568 [4.558, 4.580]1.280 [1.240, 1.321]0.007812 Table 5: Error Bars and Statistical Significance in Human Evaluation. DatasetBackbone Gwetâs AC2â FAComp.US HumanâHuman ARC-MCAS DeepSeek-V4-Pro 0.7681 0.6996 0.6641 GPT-4o0.6949 0.8488 0.7614 MMLU-Pro DeepSeek-V4-Pro 0.8698 0.8121 0.7888 GPT-4o0.6213 0.5153 0.6477 Average0.7385 0.7190 0.7155 HumanâLLM ARC-MCAS DeepSeek-V4-Pro 0.7436 0.7563 0.8096 GPT-4o0.7351 0.7954 0.7887 MMLU-Pro DeepSeek-V4-Pro 0.8277 0.7479 0.7645 GPT-4o0.7766 0.8489 0.7419 Average0.7708 0.7871 0.7762 Table 6: Inter-rater agreement (Gwetâs AC2) for humanâ human and humanâLLM evaluation across datasets and backbones. FA: factual accuracy; Comp.: comprehensi- bility; US: user satisfaction. the width â and area scales with the square of the diameter. So doubling the aperture roughly quadruples the light collected, letting it pick up much fainter objects. Strategy IDs used in the case example: â˘S-02 (What-to-Ask): Probe its application in a different context (e.g., a practical sce- nario, a general rule, or a self-generated example) to elicit information regarding action, because this tests the ability to bridge declarative and procedural knowl- edge, distinguishing academic recall or guided analysis from a robust, applicable, and transferable understanding. â˘S-03 (What-to-Ask): Probe the key dis- tinction from a similar concept, the pre- cise mechanism behind their intuitive lan- guage, or the nuances omitted in their sim- plification to elicit information regarding knowledge_scope, because this moves be- yond simple recall to test the userâs under- standing of conceptual boundaries, surfac- ing foundational misunderstandings and revealing the true depth of their knowl- edge. â˘S-14 (How-to-Ask): Probe a single, sim- plified component of the difficult concept â first acknowledging the difficulty to normalize it, then asking a more focused question â to elicit information regard- ing knowledge_scope, because it reduces cognitive load and anxiety, preventing dis- engagement while establishing a baseline of knowledge from which to build. â˘S-18 (How-to-Ask): Probe their interpre- tation of or first step within a brief, con- crete situation provided as a shared refer- ence point to elicit information regarding perspective, because it lowers initial cog- nitive load and provides a common anchor, effectively grounding the conversation for all expertise levels and preventing novices from stalling. To understand why IDL struggles to estimate user expertise, we compare IDL with PASSING on two representative cases with the same gold level but different predictions. As shown in Table 8, IDL makes opposite errors: it underestimates a Com- petent user in computer science when the history resembles closed-form QA, while overestimating a Competent user in chemistry when the history contains advanced terms such as group-14 hydrides and EPR spectroscopy. This reveals a shared failure mechanism: IDL mistakes question-level signals for user-level understanding. These cases suggest that IDL fails because it relies on weak proxy signals rather than direct ev- idence of user reasoning ability. It observes his- torical question difficulty, domain consistency, and interaction format, but cannot observe how the user reasons through a concept. As a result, IDL infers Table 7: Confusion matrices across different LLM backbones and datasets. Rows denote ground-truth Dreyfus levels and columns denote predicted levels in the order Novice, Advanced Beginner, Competent, Proficient, Expert. GPT-4o / MMLU-ProGPT-4o / ARC-MCASDeepSeekV4-Pro / MMLU-Pro DeepSeekV4-Pro / ARC-MCAS N AB CP EN AB CPEN AB CPEN AB CPE N14000 0N181000N140 000N181000 AB01211 0AB217000AB68 000AB415000 C01 130 0C01 1800C02 930C03 1330 P000 14 0P000 190P01 0 130P001 162 E0157 1E0026 11E00 0410E000613 CaseDomainExpertiseIDL Pred. PASSING Pred. IDL Error PatternPASSING Evidence Case 1 Computer Science Competent Advanced Beginner CompetentUnderestimates the user because closed-form QA history appears shallow. Probes concept boundaries, mechanism understanding, and applied judgment. Case 2ChemistryCompetentProficientCompetent Overestimates the user because advanced terminology appears expert-like. Probes mechanism understanding, uncertainty, and guided reasoning. Table 8: Case-level comparison between IDL and PASSING. expertise from what topics the user has asked about, rather than what the user actually understands. In contrast, PASSING actively constructs diag- nostic evidence through follow-up questions. In the computer science case, PASSING probes concept boundaries, encryption mechanisms, and zero-day disclosure trade-offs. In the chemistry case, PASS- ING asks about MâH bond strength and further narrows the discussion to orbital overlap, reveal- ing whether the user understands the underlying mechanism or only recognizes the general trend. Overall, the case study shows that accurate user- level perception requires observing how users rea- son, not merely what topics they ask about. IDL performs proxy-based inference from historical questions, whereas PASSING performs reasoning- based diagnosis through adaptive questioning. Across cases, PASSING follows a consistent pattern: probing concept boundaries, mechanisms, uncer- tainty, and applied judgment. Thus, when historical signals are misleading, PASSING can still recover the correct user level. In short, IDL follows âhis- torical question difficultyâproxy-based expertise inferenceâ, whereas PASSING follows âadaptive follow-up dialogueâreasoning-based expertise diagnosisâ. D Induced Strategies D.1 What-to-Ask Strategies S-01Probe their step-by-step plan or the under- lying rationale for a specific choice to elicit information regarding perspective, because this externalizes the userâs mental model, re- vealing their procedural thinking, priorities, and whether their actions are based on rote memorization or goal-oriented understanding. S-02Probe its application in a different context (e.g., a practical scenario, a general rule, or a self-generated example) to elicit informa- tion regarding action, because this tests the ability to bridge declarative and procedural knowledge, distinguishing academic recall or guided analysis from a robust, applicable, and transferable understanding. S-03Probe the key distinction from a similar con- cept, the precise mechanism behind their intu- itive language, or the nuances omitted in their simplification to elicit information regarding knowledge_scope, because this moves beyond simple recall to test the userâs understanding of conceptual boundaries, surfacing founda- tional misunderstandings and revealing the true depth of their knowledge. S-04Probe a boundary case, exception, ambigu- ous scenario, or common pitfall that chal- lenges its standard application to elicit infor- mation regarding perspective, because this dif- ferentiates rigid, rule-based knowledge from a nuanced, adaptive understanding of a con- ceptâs limitations and failure modes in com- plex, real-world situations. S-05Probe the relative importance of factors, al- ternative approaches, trade-offs, or the limita- tions and critiques of their chosen approach to elicit information regarding perspective, be- cause this reveals whether the user has a holis- tic, prioritized mental model that considers the problem space broadly, including counter- arguments and limitations, rather than a linear or one-sided view. S-06Probe the precise, mechanistic interaction be- tween components or other interacting sys- temic factors to elicit information regard- ing components, because this distinguishes static, list-based knowledge from a dynamic, systems-level understanding by forcing the user to trace causal links and consider the sys- tem holistically. S-07Probe the informal signals, âgut feelings,â or intuitive patterns they use to guide analysis or form hypotheses to elicit information re- garding action, because this differentiates a competent userâs deliberate process from an expertâs intuitive pattern-matching, revealing the shift from analytical to holistic decision- making. S-08Probe the specific source of their hesitation or their method for reconciling the inconsistency to elicit information regarding commitment, because this forces the user to articulate the precise boundaries of their own knowledge, re- vealing how their knowledge is structured and whether they can resolve internal conflicts. S-09 Probe their personal experiences, profes- sional stance, or the most difficult âreal-worldâ aspect of the issue to elicit information re- garding commitment, because an expertâs deep experience is often linked to a strong sense of personal or professional investment, and eliciting this can help differentiate a detached analyst from an involved expert. S-10Probe the underlying first principles, policy reasons, or historical context for that rule to elicit information regarding knowledge_scope, because this tests for tacit understanding of a ruleâs origin and purpose (âthe whyâ) rather than just explicit knowledge of its content (âthe whatâ), a key differentiator between Competent practitioners and Proficient or Ex- pert thinkers. S-11Probe the limitations, critiques, or weak- nesses of that same concept or model to elicit information regarding perspective, because a Competent user can accurately apply a model, but an Expert understands its boundaries and when it breaks down. This probe forces a shift from a deliberate explanation of the modelâs parts to a holistic evaluation of the model it- self. S-12Probe a counter-intuitive scenario where that causal link is broken or reversed to elicit in- formation regarding action, because analytic reasoning follows predictable causal chains, which is characteristic of Competence. Ask- ing for exceptions or paradoxes tests for an intuitive grasp of the topic, which allows for non-linear thinking and is a hallmark of an Expert. S-13Probe a request for synthesis and prioritiza- tion across the discussed topics (e.g., âWhat is the single most overlooked factor?â) to elicit information regarding perspective, because such signals suggest a user may have a deeper, more integrated understanding than they have explicitly stated. Asking them to synthesize and prioritize forces them to move beyond explaining individual components and demon- strate a holistic view of the entire system, re- vealing Proficient or Expert-level thinking. D.2 How-to-Ask Strategies S-14Probe a single, simplified component of the difficult concept â first acknowledging the difficulty to normalize it, then asking a more focused question â to elicit information re- garding knowledge_scope, because it reduces cognitive load and anxiety, preventing dis- engagement while establishing a baseline of knowledge from which to build. S-15Probe the consequence or increasing com- plexity of their correct foundational answer â building on it affirmingly to incrementally raise the scope (e.g., âExactly. Building on that, what is the consequence of...?â) â to elicit information regarding perspective, be- cause it creates a logical and encouraging con- versational flow, confirming their base knowl- edge while smoothly gauging the depth of their understanding. S-16Probe the âwhyâ behind their âwhatâ â the underlying principle or reasoning they have not yet justified (e.g., âCan you walk me through your thinking?â or âWhatâs the un- derlying principle for that step?â) â to elicit information regarding perspective, because it elicits the userâs mental model, distinguishing deep, principle-based understanding from rote memorization. S-17Probe the precise distinction between two closely related concepts â whether the user has defined one correctly or only partially â (e.g., âWhatâs the primary difference between X and Y?â) to elicit information regarding knowledge_scope, because it tests the preci- sion and boundaries of their understanding, revealing whether their knowledge is superfi- cial or nuanced and interconnected. S-18 Probe their interpretation of or first step within a brief, concrete situation provided as a shared reference point to elicit information regarding perspective, because it lowers initial cognitive load and provides a common anchor, effectively grounding the conversation for all expertise levels and preventing novices from stalling. S-19Probe a concrete, self-generated example from their own experience that grounds the ab- stract concept they described (e.g., âCan you give me a specific example of when youâve seen that happen?â) to elicit information re- garding components, because it tests for ac- tive recall and grounds abstract knowledge in practical application, differentiating between theoretical knowledge and applied expertise. S-20Probe how their approach would shift when a trade-off, conflicting goal, or ambiguous con- straint is introduced into the scenario (e.g., âHow would you prioritize if you had limited resources?â or âWhat if constraint X were introduced?â) to elicit information regarding action, because it tests the robustness and flex- ibility of a userâs mental model, revealing their ability to adapt and make judgments under pressure â a hallmark of true expertise. S-21 Probe how a different stakeholder would view or evaluate the same concept or deci- sion (e.g., âHow would the finance team view this decision differently?â) to elicit informa- tion regarding perspective, because it probes for a systems-level understanding and empa- thy, which are hallmarks of higher proficiency, moving beyond self-contained knowledge. E Prompts The knowledge graph construction prompt Input: <Query>, <Answer> Output 1: <Knowledge Dependency Graph with min-level> (for user simulator) Output 2: <Knowledge Dependency Graph without min-level> (for probing) Table 9: The KG construction. The min-level attribute is masked from the probing agent to prevent information leakage. The What-to-Ask strategy selection prompt Your task is to decide what information to ask about, using the What-to-Ask strategy set. Inputs: <Conversation History> <Populated Schema> <What-to-Ask Strategy Set> Output: <Information to Ask> Table 10: The What-to-Ask prompt for deciding what information to ask about. The How-to-Ask strategy selection prompt Your task is to decide how to ask that question, using the How-to-Ask strategy set. Inputs: <Information to Ask> <Conversation History> <How-to-Ask Strategy Set> Output: <Formulated Question> Table 11: The How-to-Ask prompt for deciding how to ask that question. The inquiry generation prompt Your task is to generate the inquiry to ask the user. Inputs: <Formulated Question> <Conversation History> Output: <Inquiry> Table 12: The inquiry generation prompt. The user simulator prompt Your task is to roleplay as a real person and respond to the inquiry, strictly following the persona, expertise level, and knowledge bound- aries specified in the inputs below. Inputs: <Persona Description> <Dreyfus Level>, <Known Concepts>, <Un- known Concepts> <Refusal Strategies> <User Query>, <Conversation History>, <Inquiry> Output: <User Response> Table 13: The user simulator prompt. The slot filling prompt Your task is to extract and fill information into the schema based on the user response and conversation history. Inputs: <User Response> <Conversation History> Output: <Updated Schema with Filled Slots> Table 14: The slot filling prompt. The expertise estimation prompt Your task is to estimate the userâs expertise level based on the collected information. Inputs: <Filled Schema> <Conversation History> Output: <Estimated Expertise Level> Table 15: The expertise estimation prompt. The response generation prompt Your task is to provide the final answer to the userâs initial question. Inputs: <Filled Schema> <Conversation History> <Estimated Expertise Level> Output: <Final Answer> Table 16: The response generation prompt.