Paper deep dive
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/19/2026, 4:19:16 AM
Summary
The paper introduces KNOWSIM, an evaluation framework for Large Language Models (LLMs) that uses a user simulator grounded in learning theory to assess information calibration. Unlike existing simulators, KNOWSIM maintains explicit knowledge states as graphs of Information Units (IUs) with prerequisite relationships, allowing it to compute metrics like Knowledge Gain, Delivery Calibration, and Cognitive Overload. Validated against 705 human-AI sessions across math and expert QA domains, KNOWSIM aligns significantly with human judgments (73-74% sign agreement) and reveals that the best LLM model varies by user knowledge level, exposing aptitude-treatment interactions missed by standard evaluations.
Entities (11)
Relation Signals (10)
KNOWSIM â computes â Cognitive Overload
confidence 95% ¡ KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload)
KNOWSIM â computes â Knowledge Gain
confidence 95% ¡ KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload)
KNOWSIM â computes â Delivery Calibration
confidence 95% ¡ KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload)
KNOWSIM â uses â Information Units
confidence 95% ¡ KNOWSIM... built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units
KNOWSIM â validatedon â KNOWCHAT
confidence 95% ¡ We validate KNOWSIM against 705 human-AI sessions... construct KNOWCHAT
Gemini 3.1 Pro â optimizesfor â advanced users
confidence 90% ¡ Gemini 3.1 Pro best serves advanced users
DeepSeek-V4 â optimizesfor â novice knowledge gain
confidence 90% ¡ DeepSeek V4 maximizes novice knowledge gain
KNOWSIM â outperforms â baseline simulators
confidence 90% ¡ outperforming three baseline simulators
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2608.17150v1
- Canonical: https://arxiv.org/abs/2608.17150v1
Trouble viewing inline? Open PDF directly â
Full Text
124,191 characters extracted from source content.
Expand or collapse full text
KNOWSIM: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn Yoonjoo Lee ⥠Hyoungwook Jin ⥠Tae Soo Kim ⢠Shaoyang Zhang ⥠Philippe Laban â Q. Vera Liao ⥠⥠University of Michigan ⢠KAIST â Microsoft Research lyoonjoo@umich.edu Abstract To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information cal- ibration: matching content to a userâs evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM 1 , an eval- uation framework built around a user simula- tor that maintains explicit knowledge statesâa graph of Information Units with prerequisite relationshipsâthat evolves under update rules grounded in learning theory. KNOWSIM com- putes three metrics (Knowledge Gain, Deliv- ery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information cali- bration. We validate KNOWSIM against 705 humanâAI sessions across two domains, strati- fied by knowledge level: its rankings align sig- nificantly with human judgments (73â74% sign agreement), outperforming three baseline simu- lators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge levelâaptitudeâtreatment interactions invisible to standard evaluation. § yjo2lee/knowsim yjlee36/knowchat-multi-turn-dialogues 1 Introduction As LLMs have grown capable of synthesizing complex information into fluent responses, users increasingly turn to them as collaborative part- ners in knowledge-intensive tasksâsuch as un- derstanding medical conditions and analyzing sci- entific data (Sharma et al., 2024). Yet training paradigms such as RLHF (Ouyang et al., 2022) 1  yoonjoolee.com/knowsim incentivize comprehensive answers in a single mes- sage while providing little signal about how infor- mation should be sequenced across multiple turns. Consider a doctor and a patient both asking âWhat are the signs that follicular lymphoma has trans- formed to diffuse large B-cell lymphoma?â A doc- tor can readily understand a dense response that covers symptoms, symptom elevation, and imag- ing protocols, but a patient who does not know the prerequisite concepts may find it difficult (Fig. 1). Effective collaboration in these tasks requires infor- mation calibration, not information maximizationâ delivering the right amount, at the right depth, and in the right sequence for the userâs current under- standing and cognitive capacity. Recent work has begun to explore LLM-based user simulators as scalable proxies for human eval- uation or feedback in multi-turn settings (Dou et al., 2025; Naous et al., 2025; Balog and Zhai, 2026; Ni et al., 2026), but current LLM-based simulators are behavioral proxies with two critical limitations. First, without knowledge-state modeling, they can- not evaluate information calibration for users at dif- ferent knowledge levels, only aggregate utility (Chi- ang et al., 2024). Second, without tracking how knowledge evolves, they cannot measure objec- tive knowledge gain and instead rely on surface quality of simulated interactions, often via LLM judgments (Zheng et al., 2023). This may mislead- ingly reward models that appear helpful without genuinely supporting how people actually seek in- formation for learning and understanding (Chang et al., 2025). We introduce KNOWSIM, an evaluation frame- work centered on a knowledge-state-grounded user simulator. The simulator represents the state for user understanding of a given topic as a graph of Information Units (IUs)âself-contained concepts connected by prerequisite relationships. When in- teracting with an LLM-based assistant, the simu- lator updates this state by determining how well 1 arXiv:2608.17150v1 [cs.AI] 17 Aug 2026 QUESTION / TOPIC What are the signs that follicular lymphoma has transformed to diffuse large B-cell lymphoma? Patient novice userINITIAL STATE Lymphoma basics B-symptoms LDH elevation Extranodal spread PET/CT pattern Biopsy confirmation unknownstrugglingpartialwell known TURN 1 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +0 IUs LOAD 1.27Ă Response exceeded processing capacity. TURN 2 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +2 IUs LOAD 0.65Ă B-symptoms and LDH advance, but PET/CT stays blocked. TURN 3 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +1 IU LOAD 0.58Ă PET/CT inches up, but biopsy stays gated. END OF CONVERSATION KG+3limited gainDC0.19mostly wastedCO0.74overloaded Stuck at one-third of the concepts. â Same concept graph, different starting knowledge Doctoradvanced user INITIAL STATE Lymphoma basics B-symptoms LDH elevation Extranodal spread PET/CT pattern Biopsy confirmation unknownstrugglingpartialwell known TURN 1 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +1 IU LOAD 0.17Ă Foundations already mastered, so PET/CT consolidates. TURN 2 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +1 IU LOAD 0.12Ă All prerequisites met, biopsy step consolidates. TURN 3 Lymph. B-symp LDH Extr. PET/CT Biopsy ABSORBED +1 IU LOAD 0.05Ă Conversation terminates at mastery. END OF CONVERSATION KG+4full masteryDC0.50half redundantCO0.11comfortable Full mastery, but half was redundant. What signs would suggest my follicular lymphoma has transformed to DLBCL? Look for new B-symptoms, rising LDH, new extranodal nodes, focal PET/CT uptake, and a biopsy... What are B-symptoms? And what does LDH have to do with this? B-symptoms: fever, drenching night sweats, weight loss >10%. LDH rises with cell turnover... What is PET/CT actually showing in transformation? PET/CT measures glucose uptake. Hypermetabolic 'hot spots' flag aggressive transformation... What signs would suggest my follicular lymphoma has transformed to DLBCL? Look for new B-symptoms, rising LDH, new extranodal nodes, focal PET/CT uptake, and a biopsy... When does focal high SUV warrant immediate biopsy vs rescan? PET/CT: SUVmax âĽ10 or >2Ă rise from baseline in an unsurveyed station â core biopsy 1â2 weeks... Prognosis post-transformation, R-CHOP eligible? After confirmed transformation, median OS 1.5â 3 yr. R-CHOP Ă 6 â ~50% CR, CAR-T if early relapse... Figure 1: KNOWSIM simulates two users with different knowledge levels asking the same medical question to the same assistant. The novice patient absorbs nothing in Turn 1 (cognitive load exceeds capacity), slowly picks up prerequisites, and ends with limited gain and high overload. The advanced doctor, whose prerequisites are already met, absorbs one concept per turn and reaches full mastery with low overload. Metrics (KG, DC, CO), computed directly from evolving knowledge states, explain why the same response helps one user but overwhelms another. the assistant explains each IU, whether the user can understand and absorb these IUs based on their current state, and moderates absorption based on cognitive load (Ausubel, 2000). Evaluation metrics that reflect key aspects of information calibration (Knowledge Gain, Delivery Calibration, Cognitive Overload) are computed directly from the simula- torâs internal state, providing mechanistic explana- tions for why an assistant helps or hinders a user. We validate KNOWSIM through a human study spanning two contrasting domains: math problem solving (Hendrycks et al., 2021), with verifiable reasoning and well-defined prerequisite structure, and expert-level question answering (Malaviya et al., 2023), where knowledge is more open-ended. We recruit participants stratified by their initial knowledge level (novice / intermediate / advanced) and construct KNOWCHAT, comprising 705 con- versations in which assistants vary along two axes: information delivery strategy and base model, each paired with subjective ratings and, for MathQA, pre/post knowledge tests. KNOWSIMâs rankings on condition align with human judgments (73â74% sign agreement, p=.003), with the strongest align- ment at the novice level, and outperform three base- line simulators (Dou et al., 2025). We then apply KNOWSIM to benchmark 9 frontier and mid-tier LLMs, revealing that the best model shifts by user knowledge levelâtradeoffs invisible to aggregate leaderboards. Our contributions are: â˘Knowledge-state-grounded user simulator. A simulator grounded in learning theory produc- ing mechanistic metrics (Knowledge Gain, Deliv- ery Calibration, Cognitive Overload) that explain why an assistant helps or hinders a user. â˘Human evaluation benchmark with cross- domain validation. We release 705 humanâ assistant sessions across two domains and two comparison axes, with users stratified by knowl- edge level. Our simulator outperforms three base- lines, reaching 73â74% sign agreement with hu- man rankings. â˘Frontier model benchmarking.We apply the validated simulator to 9 LLMs, revealing that the best model shifts by user levelâe.g., DeepSeek V4 maximizes novice knowledge gain while Gemini 3.1 Pro best serves advanced users. 2 Knowledge-Grounded Simulator for Information Calibration Evaluation 2.1 Problem Formulation We study the evaluation of information calibration in multi-turn conversations: how well an assistant delivers information matched to a userâs existing and evolving knowledge. Conversation and knowledge state. A con- versationC = (u 1 ,a 1 ,...,u T ,a T )consists of alternating user messagesu t and assistant responsesa t , and focuses on a specific topic determined by a given questionq(e.g., math problem, information seeking).We represent the full understanding of the topic or question 2 as a directed acyclic graphG q = (V,E)of Information Units (IUs): each nodev â Vis a self-contained concept relevant to answeringq, and each edge(v i ,v j ) â Eencodes thatv i is a prerequisite for understandingv j . The userâs knowledge state at turntis a labelings t : V â unaware, struggling, partial, knows_well. sevolves based on what IUs are provided in assistant turns and the userâs ability to absorb or learn them, as described in §2.2.3. The initial state s 0 is parameterized by the userâs knowledge level âânovice, intermediate, advanced (Fig. 1). Calibration quality. An assistant turna t is well-calibrated with respect tos tâ1 if the IUs introduced are (i) novel: not already labeled knows_well, (i) understandable: prerequisite- reachable givens tâ1 , and (i) processable: within the userâs cognitive absorption capacity at turnt. A well-calibrated assistant is one whose turns satisfy these conditions consistently as the conversation progresses. Conditions (i)â(i) are operationalized as the evaluation metrics described in §2.2. Evaluation task. Given an assistant modelM, a questionqwith IU graphG q , and a userâs cur- rent knowledge levelâ, we simulate a multi-turn conversation betweenMand the user initialized ats 0 (â), and compute calibration metrics from the resulting state trajectorys 0 ,s 1 ,...,s T . The met- rics are designed to be computable automatically, to remain interpretable in terms of underlying state transitions, and to align with human judgments of calibration quality, which we validate in §4. 2.2 User Simulation Method The simulator alternates between user message gen- eration and a state-update pipeline that updates the userâs knowledge state based on what knowledge the user can absorb from each assistant response. Full prompts and details are in §A. 2.2.1 IU Graph Construction For each questionq, the IU graphG q is grounded in a reference answerr: any text that comprehen- sively covers the concepts relevant toq, such as a gold answer for tasks with verifiable solutions or a tutorial or encyclopedia entry for open-ended topics. Nodes inG q correspond to the concepts required to understandr, and edges to prerequisite relations among them.G q is obtained via LLM- based extraction (§A.7) and held fixed across all simulation conditions. To simulate a user at knowl- edge levelâ, we initializes 0 by sampling IUs to be markedknows_well,partial_understanding, orstrugglingaccording to a ratioR â . Sampling follows topological order so that foundational pre- requisites are preferentially selected. 2.2.2 User Message Generation At each turn, the user simulator LLM generates u t conditioned on the conversation history and the current states tâ1 . Generation enforces knowledge- consistent behavior: the user cannot spontaneously mentionunawareconcepts, makes realistic errors onstrugglingconcepts, and appliesknows_well concepts correctly. Based on the local neighbor- hood density of IUs atpartial_understanding or above ins tâ1 , we also determine an articulation mode (explicit/vague/deferential) that fur- ther shapes query specificityâconfident users ask precise questions, while less confident users ask broader ones. 2.2.3 Knowledge State Update per Turn After each turn pair(u t ,a t ), the userâs knowledge states t is updated via: (1) a signal extraction phase that reads per-IU interaction signals fromu t anda t relative tos tâ1 , and (2) a state update phase that applies update rules grounded in learning theory. Signal extraction. A single LLM call overG q , s tâ1 , and(u t ,a t )assigns two labels to every IUv â V. From the assistant turn,vâs teach- ing quality is labeled as eitherwell_explained (e.g., definition derivation, worked examples, etc.), shallow(i.e., only named or partially treated), ornot_mentioned.From the user message, vâs engagement is labeled asreasoning(i.e., user performs their own reasoning on the IU), articulation(i.e., user restates content the as- sistant already produced), or none. State update. Given the extracted signals,s t is computed deterministically under two sets of rules: drivers of absorption in the assistant provided infor- mation, and constraints from the userâs cognition. Driver of absorption. State transitions are driven by the assistantâs teaching quality. well_explainedIU advances the userâs state by one step. Ashallowmention advances only theunawareâstrugglingtransition and has no effect at higher statesâmerely naming a concept creates awareness, but advancing understanding 3 requires substantive explanation (Vygotsky, 1978). not_mentioned IUs are unchanged. Constraints on absorption. Three constraints bound the state transitions to simulate human cog- nitive limits. Prerequisite ceiling caps each IUâs attainable state by the current mastery of its pre- requisites: a user cannot truly graspv j without first understanding foundationalv i (Ausubel, 2000). Cognitive overload computes a state-weighted load L t over IUs mentioned by both the user and as- sistants, where less-known IUs contribute higher load and userreasoningattempts also contribute more. WhenL t exceeds capacityĎ, a sigmoid dropoff continuously reduces the magnitude of up- ward transitions, and fractional progress lost in any turn carries forward so partial exposure compounds across turns (Sweller, 1994; Sweller et al., 2011). Monotonicity prevents knowledge states from re- gressing, ensuring that once a concept is learned it is not forgotten within the session. Full definitions and parameter values are in §A.5. 2.2.4 Conversation Termination Rather than imposing a fixed turn limit, the simu- lator terminates under two progress-based condi- tions grounded in the learning progress hypoth- esis (Oudeyer et al., 2016; Loewenstein, 1994) that engagement is sustained while progress is per- ceived and degrades when it saturates or stalls. Ter- mination triggers when the user (i) achieves mas- tery of the target IUs, or (i) shows persistent non- progress or overloadâat least two of cognitive overload, no upward state transitions, and no sub- stantive teaching, sustained over a five-turn window. Any conversation that triggers neither is capped at a maximum of 15 turns. Detailed parameters for the detection are in §A.6. 2.3 Evaluation Metrics The framework produces three metrics reflecting the outcome of and requirements for calibration quality computed deterministically from the state trajectorys 0 ,...,s T and the per-turn classifica- tions from knowledge state update: Knowledge Gain (KG). Total ordinal ad- vancement summed over IUs,with states mappedunaware=0,struggling=1,partial=2, knows_well=3: KG = X vâV max 0, s T (v)â s 0 (v) (1) Delivery Calibration (DC).How well the assis- tant delivers teachable content that the user can ab- sorb, measured by the harmonic mean of precision over explained IUs (E t ) and recall over teachable IUs. DC = 2PR P + R (2) whereP= P t |A t |/ P t |E t |,R= P t |A t |/ P t |Z t |,A t = v â E t : s tâ1 (v) < kw â§ ceil tâ1 (v) ⼠pu â§ s t (v) > s tâ1 (v), andZ t is the set of IUs teachable at turn t, withpu=partial_understandingand kw=knows_well.DC penalizes redundant, prerequisite-inappropriate, and ultimately unab- sorbed delivery. Cognitive Overload (CO). Average per-turn load relative to capacity, saturating at one: CO = 1 T X t min 1, L t Ď (3) whereL t is a state-weighted load over IUs men- tioned and user attempts at turnt, andĎis the capacity threshold (details in §A). We design the metrics to be automatically com- putable, remain interpretable in terms of underlying state transitions, and align with human judgments of calibration quality, which we validate in §4. 3KNOWCHAT: A Benchmark of Knowledge-Stratified Conversations with Learning Experience Annotations We collect conversations between real users and AI assistants to serve as ground truth for validating our simulator. The dataset captures both objective learning outcomes and subjective perceptions of users with varying levels of knowledge interacting with assistants that employ different information- delivery strategies and assistant models. 3.1 Tasks and Conditions We design four study arms crossing two factorsâ comparison target (strategy vs. model) and task (math vs. ExpertQA)âeach drawing a separate pool of participants in Prolific (Prolific, 2026). As shown in Table 1, the strategy arms compare con- ditions of different delivery strategies by fixing the base model; the model arms hold the strategy fixed and vary the conditions of base models. 4 Table 1: Study arms overview. Each arm recruitsâź60 participants (min: 50, max: 72). ArmTaskFixedComparison Target MathĂ Strat.MATH (L5)ModelStrategy MathĂ Model MATH (L5) StrategyModel Exp.Ă Strat.ExpertQAModelStrategy Exp.Ă ModelExpertQAStrategyModel Tasks. For competition mathematics, we use 15 problems from the MATH dataset (Hendrycks et al., 2021) at difficulty level 5, spanning five topic cate- gories (counting & probability, geometry, interme- diate algebra, number theory, precalculus); within each category, we select three problems requiring distinct IU graphs to eliminate cross-session trans- fer. For ExpertQA (Malaviya et al., 2023), par- ticipants select a professional domain they are in- terested in (e.g., visual arts, education, psychol- ogy) and receive relevant questions. We prioritize breadth: each participant encounters three distinct questions, collectively covering 150 questions. Conditions.The strategy arms compare three in- formation delivery strategies. Adaptive infers the userâs knowledge level from conversation cues and adjusts explanation depth accordingly. Compre- hensive delivers complete, well-structured explana- tions independent of user signals. Socratic guides by asking thought-provoking questions unless the user wants confirmation or struggles. Model arms compare GPT-5.4, Claude Opus 4.7, and Gemini 3 Pro, all under the simple strategy without much instruction. Full system prompts are in §B.1. 3.2 Data Collection Procedure Step 1. Recruitment and stratification. Each arm recruitsâź60 Prolific workers, stratified into novice, intermediate, and advanced knowledge lev- els to enable comparison of simulator and human outcomes within each level. For both tasks, stratifi- cation uses a 10-item domain-specific prescreening test (design in §B.2). Step 2: Assigning conditions. Each participant completes three sessions, one per condition (three strategies or three models, depending on the arm). Math uses a fixed pool of 15 problems so that vali- dated pre/post-tests can be administered for objec- tive knowledge gain. We rotate conditionâproblem pairings across participants to control for ordering and item effects. ExpertQA instead prioritizes do- main coverage breadth, so each condition draws from a large pool of problems without repetition. Step 3. Per-session protocol. Math sessions fol- low a pre-testâconversationâpost-testâsub- jective ratings flow, while ExpertQA sessions omit the pre- and post-tests. In both, participants en- gage in a free-form conversation with the assigned assistant to learn to solve the problem (Fig. 5). 3.3 Annotations and Quality Control Each session in KNOWCHAT includes the full con- versation traces, annotated with participantsâ sub- jective ratings on (10-point scale): Delivery Cali- bration, Cognitive Overload, and Interaction Qual- ity (all questions in §B.3). Math sessions addi- tionally include Knowledge Gain,KG = (Postâ Pre)/(Max post â Pre). We include only partici- pants who complete all 3 sessions, each with at least 2 user turns. Participants received $20/hr. We collected data with informed consent and IRB ap- proval. Per-task dataset statistics are in §E.1. 4 Experiment Setup We evaluate KNOWSIM along (i) alignment with human outcomesâwhether the simulator repro- duces human-derived condition rankings across metrics and knowledge levels, and (i) compari- son with baselinesâhow KNOWSIMâs correlation with human rankings compares to that of baseline simulators on the metric common to all methods. 4.1 Simulator Methods We compare KNOWSIM against three baselines of increasing sophistication (all using Gemini-3- Flash; full formal definitions in §C.1): Zero-shot (ZS) conditions only on a knowledge-level de- scriptor and conversation history; ZS-CoT adds a chain-of-thought reasoning step at each turn (Wei et al., 2022); ZS-CoT-Prof further adds a synthetic user profile and initial knowledge state (Dou et al., 2025), but does not update it over conversation. KNOWSIM (ours). Uses Gemini-3-Flash for user-turn generation and per-turn knowledge-state extraction, and GPT-5.2 for one-time IU graph con- struction. The method is described in §2.2. 4.2 Simulation Protocol and Metric Coverage Each simulator is run on every(q,â,c)configu- ration in KNOWCHAT, whereqis a question,âa knowledge level, andca condition. For KNOWSIM, the initial knowledge states 0 (â)is instantiated by 5 Table 2: Metric coverage. IQ is common across simula- tors; KG, DC, and CO are unique to KNOWSIM. SourceKGDCCOIQ KNOWSIM (ours)â Baselinesââ Humanâ â â â Human KG from pre/post tests; math sessions only. sampling from the IU graphG q according to level- specific ratios (§2.2). All conversations run to a maximum of 15 turns. For KNOWSIM, we deter- mine the end turn using the termination rule in §2.2.4. For baseline simulators, we follow Sim- ulatorArena and use an LLM judge to post-hoc identify the natural ending turn (Dou et al., 2025). We compare human and simulator outcomes at the(â,c)cell-mean level: for each cell, the hu- man mean aggregates all participants assigned to that(â,c)combination, and the simulator mean ag- gregates all simulated conversations at the same configuration. This tests whether the simulator reproduces the population-level effect of each con- dition within each knowledge-level stratum. Following prior work (Dou et al., 2025), all four simulators produce Interaction Quality (IQ) scores via a shared LLM raterĎ r , enabling direct cross-method comparison. KNOWSIM addition- ally exposes three internal metrics (KG, DC, CO) computed deterministically from its state trajec- tory (§2.2); these have no counterpart in baseline simulators. Table 2 summarizes coverage. 4.3 Analyses We assess alignment using pairwise sign agree- ment: within each(arm, knowledge level)block, we enumerate all condition contrasts, retain those where the human effect exceeds a non-negligible threshold (Cliffâsδ ⼠0.15), and check whether contrasts from the given simulator have the same sign. Counts are tested against chance via a one- sided binomial test. This formulation requires only an ordinal commitment to human directionâ appropriate for subjective ratingsâand admits ex- act small-sample inference via a one-sided bino- mial test (p = 0.5) without distributional assump- tions. Theδ ⼠0.15threshold excludes near-zero human effects where the ground-truth direction is ambiguous; results are stable across[0.10, 0.20]. Following Dou et al. (2025), SpearmanĎcorre- lations between human and simulator cell means corroborate all findings (Appendix E). Table 3: Sign agreement between KNOWSIM and hu- man pairwise condition contrasts on high-signal cells (|δ human | ⼠0.15). Rate indicates the fraction of con- trasts where the simulatorâs direction matches the hu- man Cliffâsδ;pis from a one-sided binomial test against chance (50%). Evaluation SettingAgree / HSRatep MathQAâStrategy13 / 1776%.025 â MathQAâModel14 / 2070%.058 â ExpertQAâStrategy16 / 2080%.006 â ExpertQAâModel10 / 1567%.151 MathQA (pooled)27 / 3773%.003 â ExpertQA (pooled)26 / 3574%.003 â On the human side, DC, CO, and IQ are par- ticipant self-ratings on a 1â10 scale collected af- ter each session; KG is computed from pre/post knowledge tests (MathQA only). We analyze along two axes: (i) alignment of KNOWSIMâs four met- rics with human outcomes, reported as metric- aggregated sign agreement per arm (Table 3) and per (task, level) (Fig. 2); and metric specific sign agreement aggregated by arms and levels (Fig. 3); and (i) comparison with baseline simulators on IQâthe one metric common to all methodsâusing the same sign agreement pooled across all arms (§5.2). For the IQ raterĎ r , we select Claude Sonnet 4.6 after testing three candidate LLMs and check- ing for self-preference bias. 5 Results We show that (1) KNOWSIMâs multi-metric sig- nal aligns with human judgment across tasks and knowledge levels (§5.1) and (2) it outperforms baseline simulators on the common metric (§5.2). 5.1 Overall Alignment with Human Judgment KNOWSIMâs ranking direction aligns with hu- man judgment across both task domains (Ta- ble 3).Pooled across Strategy and Model arms, sign agreement reaches significance on both MathQA (27/37, 73%,p=.003) and ExpertQA (26/35, 74%,p=.003). Per-arm rates range from 67% to 80%, with two of four arms individually significant (p=.025andp=.006). At the task level, KNOWSIM is a directionally reliable proxy, moti- vating the finer-grained analyses below. 2 Novice level has the strongest alignment in both tasks (Fig. 2). Agreement at the novice level 2 Significance: one-sided exact binomial vs. 50% chance baseline, with â p < .10, â p < .05, â p < .01. 6 noviceintermediateadvanced 0 20 40 60 80 100 Sign agreement (%) 100%** 50% 70% MathQA noviceintermediateadvanced 83%* 80%â 62% ExpertQA Figure 2: Per-level sign agreement (%) between KNOWSIM and human rankings, by task (Strategy + Model arms pooled). 020406080100 Sign agreement (%) IQ CO DC KG 75%â 75% 67% 75% MathQA 020406080100 Sign agreement (%) 80%â 79%* 64% n/a ExpertQA Figure 3: Per-metric sign agreement (%) between KNOWSIM and human rankings, by task (Strategy + Model arms pooled). KG is not measured on ExpertQA. is 100% (13/13,p<.001) on MathQA and 83% (10/12,p=.019) on ExpertQA. Intermediate and advanced levels are more variable: MathQA in- termediate drops to chance (50%), while on Ex- pertQA intermediate remains high (80%,p=.055) and advanced falls to 62%. We interpret this pattern through the lens of what drives user preferences at each level. For novice users, comprehension and learning progressâprecisely what KNOWSIMâs knowledge-state model tracksâare likely to dom- inate session preferences, and KNOWSIM repro- duces their rankings most reliably on both tasks. At higher knowledge levels, agreement is both lower and less consistent across the two tasks, suggesting that more knowledgeable users evaluate sessions along dimensions beyond knowledge acquisition that our state-based metrics do not yet model. IQ and CO emerge as the most consistent di- agnostic metrics across both tasks (Fig. 3). On MathQA, IQ, KG, and CO each reach 75% agree- ment (IQ:p=.073); DC trails at 67%. On Ex- pertQA, IQ (80%,p=.055) and CO (79%,p=.029) lead; DC remains the lowest (64%). Overall, align- ment is consistently high across IQ, KG (MathQA), and CO in both tasks, while DC shows a mod- est gap that may reflect the difficulty of capturing perceived calibration quality through a single sub- jective rating. 020406080100 Sign agreement (%) ZS-CoT ZS ZS-CoT-Prof KnowSim 64% 68%â 73%* 77%** Figure 4: Cross-task IQ sign agreement (%) between each simulator and human rankings, pooled across both tasks and all four arms (22 signal pairs). 5.2 Comparison with Baselines KNOWSIM achieves the highest IQ alignment among all simulators (Fig. 4). IQ is the only met- ric that the baselines can compute. Pooling all four arms, KNOWSIM agrees with human pairwise IQ rankings 77% of the time (p=.008), ahead of ZS-CoT-Prof (73%,p=.026), ZS (68%,p=.067), and ZS-CoT (64%,p=.143). Broken down by arm, KNOWSIM leads or ties ZS-CoT-Profâthe strongest baselineâin all four (Appendix E.2). Beyond metric-level agreement, on MathQA, KNOWSIMâs simulated users show greater over- lap with real users in the IUs they engage with than any baseline, as measured by full-turn Jaccard similarity (Appendix E.4). 6 Comparing Assistants with KNOWSIM An important application of user simulators is eval- uating how well assistant LLMs calibrate to user populations with different knowledge backgrounds. Having established that KNOWSIM reliably cap- tures information calibration quality (§5), we apply it to benchmark 9 open- and closed-source LLMs. Each LLM is evaluated on 30 MathQA and 30 Ex- pertQA items, with simulated users at three knowl- edge levels per item. We report KNOWSIMâs inter- nal metrics (KG, DC, CO) and IQ in §F. Overall ranking. Table 16 (Appendix) reports the level-marginalized ranking across 9 models. Claude Opus 4.7 leads overall ( Ě R= 2.75), with the strongest DC (0.108) and 2nd-highest IQ (8.16). Gemini 3.1 Pro ranks 2nd ( Ě R =3.50) with the high- est IQ (8.24) and lowest CO (0.833). DeepSeek V4 ranks 1st on KG but 8th on CO (0.890), revealing a âknowledge-dumpingâ pattern where aggressive ex- planation maximizes knowledge throughput while overwhelming the learner. Gemini 3.1 Flashâthe mid-tier counterpart to Proâplaces 4th overall with strong KG (7.1, rank 2) but the highest CO (0.895), 7 Table 4: Aptitudeâtreatment interaction patterns: per-level metrics across 9 LLMs. Each cell shows the metric value; Ě R is the mean rank across the four metrics within that level. COâ: lower is better. Best value per column is bolded. NoviceIntermediateAdvanced ModelKGCOâDCIQ Ě RKGCOâDCIQ Ě RKGCOâDCIQ Ě R Gemini 3.1 Pro13.20.8750.0868.393.255.40.7500.0888.354.750.70.8730.1227.961.75 Claude Opus 4.714.30.8820.0907.803.255.60.8070.1288.392.500.60.8790.1078.303.00 DeepSeek V415.50.9450.0926.504.005.70.8350.1217.534.250.50.8900.1097.744.25 Gemini 3.1 Flash15.10.9360.0936.833.255.60.8380.1167.735.750.50.9120.1097.646.00 Claude Sonnet 4.612.90.8770.0797.064.755.40.7910.0987.875.000.60.8660.1008.003.50 GPT-5.414.90.9400.0885.245.505.60.7920.1196.865.500.40.9030.1027.057.75 Qwen-3.6-35B14.40.9340.0825.685.255.70.8290.1286.894.500.50.8920.0817.397.00 GPT-5.4 mini12.60.9460.0755.418.005.40.8090.1217.136.000.50.9010.1147.245.75 Llama-4-Maverick11.80.9270.0714.58 7.755.40.7670.1055.646.750.30.8820.1176.076.00 sharing the overload pattern of high-throughput models. Aptitudeâtreatment patterns across knowledge levels. Table 4 breaks down metrics by knowl- edge level, revealing that no single model serves all users equally well. The best model for a given met- ric shifts across levels. For novices, DeepSeek V4 delivers the most knowledge (KG=15.5) but at high overload cost (CO=0.945, second-worst), while Gemini 3.1 Pro achieves the lowest novice CO (0.875) with the highest novice IQ (8.39). For ad- vanced users the picture inverts: Gemini 3.1 Pro is best-calibrated overall (best Ě R=1.75) with the high- est KG (0.7) and near-lowest overload (CO=0.873), while Claude Opus 4.7 attains the highest IQ at both the intermediate (8.39) and advanced (8.30) levels but is not the top knowledge-delivery model for novices. These rank reversals expose an equity gap in- visible to level-agnostic leaderboards: Gemini 3.1 Proâthe best-calibrated model for advanced usersâranks 6th on novice KG, while DeepSeek V4 delivers the most knowledge to novices but over- whelms them (CO=0.945). KNOWSIMâs per-level decomposition makes these tradeoffs explicit, en- abling practitioners to select assistants matched to their target user population. Beyond ranking, IU-level tracking lets us de- compose why calibration fails: pooled across as- sistants, novice sessions are dominated by over- reaching (65% of explained IUs are prerequisite- blocked or overload-wasted), while advanced ses- sions are dominated by redundant re-explanation (89%). This directional failure pattern is detailed in §G. 7 Related Work Multi-turn LLM evaluation. Comparing assis- tants in multi-turn conversations is challenging as the assistantâs prior response shapes each user turn, which affects the verdict in return (Mehri and Es- kenazi, 2020). Static benchmarks (Zheng et al., 2023; Kwan et al., 2024; Bai et al., 2024) and real- user studies (Collins et al., 2024; Ibrahim et al., 2025b) sit on opposite sides of a fidelity-scalability tradeoff: pre-scripted users cannot react to assistant replies, while live participants do not scale across many assistant-user combinations. Bridging this tradeoff requires user models that interact dynami- cally with assistants but remain cheap to deployâ motivating our LLM-based simulator paradigm. LLM-based user simulators. Beyond prompt- ing, user simulators have been built by augmenting generation with external knowledge (Dhole, 2024), self-play and fine-tuning on dialogue data (Kong et al., 2024; Ulmer et al., 2024), reinforcement learning of dedicated user language models (Naous et al., 2025), inferring implicit profiles from past interactions (Wang et al., 2025), and agent archi- tectures for search behavior (Zhang et al., 2024). User simulators are often used as evaluation infras- tructure for specific assistant capabilities (Ibrahim et al., 2025a; Li et al., 2024), such as tool use (Yao et al., 2024) and instruction following (Laban et al., 2025). A recent thread adds richer user dynamics: DiscoverLLM (Kim et al., 2026) models intent for- mation through a hierarchical tree; HumanLM (Wu et al., 2026) aligns latent psychological states (e.g., belief, emotion) with ground-truth responses via RL; and CollabLLM (Wu et al., 2025a) forward- samples simulated trajectories to compute multi- turn-aware rewards for model training. We extend this research thread by modeling usersâ evolving knowledge, enabling novel metrics for information 8 calibration. Student simulation and knowledge state mod- eling. A growing body of work uses LLMs to simulate students for educational AI research (Tian et al., 2025): simulating cognitive levels and learn- ing dynamics (Wu et al., 2025b; Yuan et al., 2025), generating student errors (Jin et al., 2024; Ross and Andreas, 2025), and auditing simulator quality with human teachers (Martynova et al., 2025) or automated metrics (Scarlatos et al., 2026). How- ever, these efforts typically characterize learners by static knowledge profiles rather than explic- itly tracking how their knowledge changes over an interaction. Modeling knowledge as an evolv- ing state has a long history in educational AI. Stu- dent modeling represents mastery over structured prerequisite relations (Anderson et al., 1995; Van- Lehn, 2011), while knowledge tracing updates es- timates of mastery based on learnersâ engagement with material (Corbett and Anderson, 1994; Piech et al., 2015), with recent work extending this idea to open-ended tutorâstudent dialogue (Scarlatos et al., 2025). Dialogue systems similarly maintain latent user states across turns, for example through belief trackers that estimate distributions over task- relevant slot values (MrkĹĄi Ě c et al., 2017). We build on these traditions to model evolving knowledge in open-ended LLM dialogue. Unlike dialogue- based knowledge tracing, which infers mastery from observed tutorâstudent exchanges, our simula- tor updates the userâs knowledge from information received during interaction and uses that state to generate subsequent behavior. IU graphs represent conceptual understanding and prerequisite relations among concepts. 8 Conclusion We presented KNOWSIM, a knowledge-state-aware user simulator that models how understanding evolves across multi-turn conversations via an inter- nal knowledge representation, and KNOWCHAT, a benchmark of 705 knowledge-stratified humanâAI conversations across two task domains with learn- ing outcome annotations. KNOWSIM aligns with human judgment at 73â74% sign agreement and outperforms baseline simulators. Across 9 frontier LLMs, it reveals aptitudeâtreatment interactions invisible to aggregate leaderboards, where the best model shifts by user knowledge level. Limitations Our simulator models information-processing constraintsâcognitive load, prerequisite-gated ab- sorption, and overload-driven terminationâbut un- derspecifies motivational and affective dynamics such as frustration and boredom; this manifests in KG saturation for advanced simulated users. The simulator nonetheless captures the ATI crossover in Top-1 strategy ranking across Novice, Interme- diate, and Advanced learners, and integrating moti- vational dynamics drawn from expectancyâvalue or flow theory remains future work. A second scope limitation is that our evaluation covers two knowledge-intensive domains (math problem-solving and ExpertQA); extending the framework to domains with more open-ended suc- cess criteria, such as creative writing or exploratory data analysis, is ongoing work. While our simulator can express within-group variation by tuning the KS ratio per user, estimat- ing an individualâs initial KS state from observable behavior remains challenging in open-ended con- versational settings, despite progress in knowledge tracing for structured tutoring contexts (Corbett and Anderson, 1994; Piech et al., 2015). Bridging this gap with our group-level cognitive abstractions is a promising future research direction. On the implementation side, our simulator re- quires multiple LLM calls per turn for IU absorp- tion analysis and user response generation, which constrains evaluation throughput when scaling to many assistant models; since the structured com- ponents (IU graph, prerequisite gating, load blend) are deterministic and inexpensive, distilling only the LLM-based modules into smaller fine-tuned simulators could substantially reduce cost without sacrificing alignment. Finally, having validated our simulator as an evaluation proxy that produces human-aligned as- sistant rankings, applying it as a training signalâ for instance, as a reward model in RL fine-tuning of assistants (Wu et al., 2025a; Kim et al., 2026)âis a natural extension we leave for future work. Ethical Considerations Risks and societal impact. KNOWSIM is an eval- uation proxy and should not be used as a substitute for human studies. Over-reliance on simulator- derived rankings can lead practitioners to deploy assistants that score well in simulations but may underserve real users, particularly at knowledge 9 levels where our metrics may align less with hu- man judgment. Our own results surface an equity dimension: models that maximize novice knowl- edge gain may do so at high cognitive overload, and the best-calibrated model for advanced users underserves novices. Selecting assistants on ag- gregate scores alone could therefore systematically disadvantage less-knowledgeable users. Artifacts and licensing.We release KNOWCHAT (705 humanâassistant sessions) and the simulator code for non-commercial use under C BY-NC 4.0.We build on the MATH dataset (Hendrycks et al., 2021) and ExpertQA (Malaviya et al., 2023).Both are publicly released for research, and our use is consistent with their intended research use. Model outputs were obtained via each providerâs API under their respective terms of service. Privacy and content. Human-assistant sessions were collected via Prolific, which exposes only anonymous participant IDs and these are not re- tained in the final collected data. We collected no names or other personally identifying information. We manually reviewed a sample of sessions and found no personally identifying or offensive con- tent. References John R Anderson, Albert T Corbett, Kenneth R Koedinger, and Ray Pelletier. 1995. Cognitive tu- tors: Lessons learned. The journal of the learning sciences, 4(2):167â207. David Paul Ausubel. 2000. The acquisition and reten- tion of knowledge: A cognitive view. Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jia- heng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, and Wanli Ouyang. 2024. MT-bench-101: A fine-grained benchmark for evalu- ating large language models in multi-turn dialogues. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7421â7454, Bangkok, Thailand. Association for Computational Linguistics. Krisztian Balog and ChengXiang Zhai. 2026. User simulation in the era of generative ai: User model- ing, synthetic data generation, and system evaluation. Preprint, arXiv:2501.04410. Serina Chang, Ashton Anderson, and Jake M Hofman. 2025. Chatbench: From static benchmarks to human- ai evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 26009â 26038. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anasta- sios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. Chatbot arena: An open platform for evaluating LLMs by human prefer- ence. In Proceedings of the 41st International Con- ference on Machine Learning (ICML). Katherine M. Collins, Albert Q. Jiang, Simon Frieder, Lionel Wong, Miri Zilka, Umang Bhatt, Thomas Lukasiewicz, Yuhuai Wu, Joshua B. Tenenbaum, William Hart, Timothy Gowers, Wenda Li, Adrian Weller, and Mateja Jamnik. 2024. Evaluating lan- guage models for mathematics through interactions. Proceedings of the National Academy of Sciences, 121(24):e2318124121. Albert T Corbett and John R Anderson. 1994. Knowl- edge tracing: Modeling the acquisition of procedural knowledge. User modeling and user-adapted inter- action, 4(4):253â278. Kaustubh Dhole. 2024. KAUCUS - knowledgeable user simulators for training large language models. In Proceedings of the 1st Workshop on Simulating Con- versational Intelligence in Chat (SCI-CHAT 2024), pages 53â65, St. Julians, Malta. Association for Com- putational Linguistics. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, and Jianfeng Gao. 2025. SimulatorArena: Are user simu- lators reliable proxies for multi-turn evaluation of AI assistants? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Process- ing, pages 35212â35290, Suzhou, China. Association for Computational Linguistics. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language under- standing. In International Conference on Learning Representations. Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi, Minsuk Kahng, Meredith Ringel Morris, Kevin R McKee, Verena Rieser, Murray Shanahan, and Laura Weidinger. 2025a. Multi-turn evaluation of anthropomorphic behaviours in large language models. arXiv preprint arXiv:2502.07077. Lujain Ibrahim, Saffron Huang, Lama Ahmad, Umang Bhatt, and Markus Anderljung. 2025b. Towards inter- active evaluations for interaction harms in human-ai systems. In Proceedings of the AAAI/ACM Confer- ence on AI, Ethics, and Society, volume 8, pages 1302â1310. Hyoungwook Jin, Seonghee Lee, Hyungyu Shin, and Juho Kim. 2024. Teach ai how to code: Using large language models as teachable agents for program- ming education. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Sys- tems, CHI â24, New York, NY, USA. Association for Computing Machinery. 10 Tae Soo Kim, Yoonjoo Lee, Jaesang Yu, John Joon Young Chung, and Juho Kim. 2026. Discov- erllm: From executing intents to discovering them. arXiv preprint arXiv:2602.03429. Chuyi Kong, Yaxin Fan, Xiang Wan, Feng Jiang, and Benyou Wang. 2024. PlatoLM: Teaching LLMs in multi-round dialogue via a user simulator. In Pro- ceedings of the 62nd Annual Meeting of the Associa- tion for Computational Linguistics (Volume 1: Long Papers), pages 7841â7863, Bangkok, Thailand. As- sociation for Computational Linguistics. Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. 2024. MT-eval: A multi- turn capabilities evaluation benchmark for large lan- guage models. In Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 20153â20177, Miami, Florida, USA. Association for Computational Linguistics. Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. 2025. Llms get lost in multi-turn conversation. arXiv preprint arXiv:2505.06120. Shuyue S Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S Ilgen, Emma Pierson, Pang W Koh, and Yulia Tsvetkov. 2024. Mediq: Question-asking llms and a benchmark for reliable interactive clinical rea- soning. Advances in Neural Information Processing Systems, 37:28858â28888. George Loewenstein. 1994. The psychology of curios- ity: A review and reinterpretation. Psychological Bulletin, 116:75â98. Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2023. Ex- pertqa: Expert-curated questions and attributed an- swers. ArXiv, abs/2309.07852. Daria Martynova, Jakub Macina, Nico Daheim, Nilay Yalcin, Xiaoyu Zhang, and Mrinmaya Sachan. 2025. Can llms effectively simulate human learners? teach- ersâ insights from tutoring llm students. In Proceed- ings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 100â117. Shikib Mehri and Maxine Eskenazi. 2020. Unsuper- vised evaluation of interactive dialog with DialoGPT. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 225â235, 1st virtual meeting. Association for Computational Linguistics. Nikola MrkĹĄi Ě c, Diarmuid Ă SĂŠaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017. Neu- ral belief tracker: Data-driven dialogue state tracking. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1777â1788, Vancouver, Canada. Association for Computational Linguistics. Tarek Naous, Philippe Laban, Wei Xu, and Jennifer Neville. 2025.Flipping the dialogue: Training and evaluating user language models. Preprint, arXiv:2510.06552. Bo Ni, Yu Wang, Leyao Wang, Branislav Kveton, Franck Dernoncourt, Yu Xia, Hongjie Chen, Reuben Luera, Samyadeep Basu, Subhojyoti Mukherjee, Puneet Mathur, Nesreen K. Ahmed, Junda Wu, Li Li, Huixin Zhang, Ruiyi Zhang, Tong Yu, Sungchul Kim, Jiuxiang Gu, and 11 others. 2026. A survey on LLM- based conversational user simulation. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 4266â4301, Rabat, Mo- rocco. Association for Computational Linguistics. P. Y. Oudeyer, J. Gottlieb, and M. Lopes. 2016. In- trinsic motivation, curiosity, and learning: Theory and applications in educational technologies, pages 257â284. Progress in Brain Research. Elsevier B.V. Publisher Copyright: Š 2016 Elsevier B.V. Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Car- roll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730â27744. Chris Piech, Jonathan Bassen, Jonathan Huang, Surya Ganguli, Mehran Sahami, Leonidas J Guibas, and Jascha Sohl-Dickstein. 2015. Deep knowledge trac- ing. Advances in neural information processing sys- tems, 28. Prolific. 2026. Prolific [online participant recruitment platform]. https://w.prolific.com. Alexis Ross and Jacob Andreas. 2025. Learning to make mistakes: Modeling incorrect student thinking and key errors. arXiv preprint arXiv:2510.11502. Alexander Scarlatos, Ryan S. Baker, and Andrew Lan. 2025. Exploring knowledge tracing in tutor-student dialogues using LLMs. In Proceedings of the 15th In- ternational Learning Analytics and Knowledge Con- ference (LAK â25), pages 249â259. Association for Computing Machinery. Alexander Scarlatos, Jaewook Lee, Simon Woodhead, and Andrew Lan. 2026. Simulated students in tutor- ing dialogues: Substance or illusion? arXiv preprint arXiv:2601.04025. Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024. Generative echo chamber? effect of llm-powered search systems on diverse information seeking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI â24, New York, NY, USA. Association for Computing Machinery. John Sweller. 1994. Cognitive load theory, learning difficulty, and instructional design. Learning and instruction, 4(4):295â312. 11 John Sweller, Paul Ayres, and Slava Kalyuga. 2011. Cognitive Load Theory. Springer. Xinxian Tian, Guanzhong Pan, and Haibo Wang. 2025. Large language models in student simulation: A sur- vey. Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Lijia Sun, Xibin Gao, and Yi Zhang. 2024. Bootstrap- ping LLM-based task-oriented dialogue agents via self-talk. In Findings of the Association for Compu- tational Linguistics: ACL 2024, pages 9500â9522, Bangkok, Thailand. Association for Computational Linguistics. Kurt VanLehn. 2011. The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational psychologist, 46(4):197â221. Lev S Vygotsky. 1978. Mind in society: The develop- ment of higher psychological processes. Kuang Wang, Xianfei Li, Shenghao Yang, Li Zhou, Feng Jiang, and Haizhou Li. 2025. Know you first and be you better: Modeling human-like user sim- ulators via implicit profiles. In Proceedings of the 63rd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 21082â21107, Vienna, Austria. Association for Com- putational Linguistics. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, F. Xia, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. ArXiv, abs/2201.11903. Shirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang, Joy He-Yueya, Tharindu Cyril Weerasooriya, Wei Wei, Diyi Yang, Jure Leskovec, and James Zou. 2026. Humanlm: Simulating users with state alignment beats response imitation. arXiv preprint arXiv:2603.03303. Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, and 1 others. 2025a. Collabllm: From passive responders to active collaborators. arXiv preprint arXiv:2502.00640. Tao Wu, Jingyuan Chen, Wang Lin, Mengze Li, Yu- meng Zhu, Ang Li, Kun Kuang, and Fei Wu. 2025b. Embracing imperfection: Simulating students with diverse cognitive levels using llm-based agents. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 9887â9908. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2024. taubench: A benchmark for tool- agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045. Yu Yuan, Lili Zhao, Wei Chen, Guangting Zheng, Kai Zhang, Mengdi Zhang, and Qi Liu. 2025. Simulating human-like learning dynamics with llm-empowered agents. arXiv preprint arXiv:2508.05622. Erhan Zhang, Xingzhu Wang, Peiyuan Gong, Yankai Lin, and Jiaxin Mao. 2024. USimAgent: Large lan- guage models for simulating search users. In Pro- ceedings of the 47th International ACM SIGIR Con- ference on Research and Development in Information Retrieval (SIGIR â24), pages 2687â2692. Association for Computing Machinery. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-bench and chatbot arena. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track. 12 A User Simulator Details This appendix supplies the prompts, parameter val- ues, and worked-example backing the user simu- lator described in §2.2. The subsections mirror the order in §2: IU graph construction (§A.1), ini- tial knowledge state (§A.2), user message genera- tion (§A.3), signal extraction (§A.4), state update rules (§A.5), conversation termination (§A.6), and a worked IU graph example (§A.7). A.1 IU Graph Construction We usegpt-5.2for the one-time extraction call per(q,r). The math and ExpertQA prompts share structure with different examples or details. We present a unified prompt below with annotations [M]for Math-only content,[E]for ExpertQA-only content, and labeled worked examples per domain. IU Graph Extraction Prompt (unified Math + ExpertQA) You are an expert at analyzing complex questions and their answers to extract structured knowledge representations. Given a question and its reference answer, extract an Information Unit (IU) Graph - a directed acyclic graph that represents all the pieces of understanding a person needs to fully comprehend the answer. [E] ## Domain Context [E] Field: field [E] Specific field: specific_field [E] Use this context to calibrate terminology and concept granularity. Do not restrict the graph to concepts from this field only - include any cross-domain prerequisites that are genuinely necessary for comprehension. ## What is an Information Unit (IU)? An Information Unit is a self-contained piece of understanding that: 1. Independently assessable - you can determine whether someone "gets it" as a standalone unit 2. Explainable in 2-4 sentences - not a single fact, not an entire topic 3. Has prerequisite relationships with other IUs - understanding some IUs requires first understanding others [M] IUs can represent any type of knowledge: factual, conceptual, procedural, or reasoning-based. Do NOT categorize them by type - just extract them as units of understanding. [E] IUs represent declarative knowledge: concepts, definitions, mechanisms, causal relationships, empirical findings, and implications. They are units of conceptual understanding, not steps in a procedure. ## Abstraction Level (l) Each IU has an abstraction level l in [0, 1] that captures how general vs. context-specific the knowledge is: - l close to 1.0: General principles, definitions, or domain knowledge that apply broadly beyond this specific question. Math example: "Continuity of a function means the limit equals the function value at that point" ExpertQA example (oncology): "Follicular lymphoma is a slow-growing B-cell non-Hodgkin lymphoma that typically follows an indolent course" ExpertQA example (law): "Intestate succession governs how an estate is distributed when the deceased left no valid will" ExpertQA example (medicine): "Peripheral neuropathy is damage to nerves outside the brain and spinal cord, causing numbness, pain, or weakness" - l close to 0.5: Intermediate knowledge that connects general principles to the specific problem - identifying which concepts apply and how. Math example: "In a piecewise function, continuity can only break at the boundary points where the definition changes" ExpertQA example (oncology): "When follicular lymphoma transforms to DLBCL, the biological shift from indolent to aggressive growth produces distinctive clinical warning signs" ExpertQA example (law): "Intestacy rules prioritize close biological and legal relationships, so a spouse typically receives priority over more distant relatives" ExpertQA example (medicine): "Cryotherapy reduces blood flow to extremities, which may limit the amount of taxane drug reaching peripheral nerves during infusion" - l close to 0.0: Concrete, context-bound knowledge tied to the specific question - particular computations, specific values, or the final synthesis. Math example: "Setting 2a+3 = -3 and solving gives a = -3" ExpertQA example (oncology): "A biopsy is the only reliable method to confirm FL-to-DLBCL transformation, even when PET/CT findings are strongly suggestive" ExpertQA example (law): "If no spouse, children, or parents are found, the estate escheats to the state under most U.S. intestacy statutes" ExpertQA example (medicine): "Current evidence on cryotherapy for CIPN prevention is conflicting, and no definitive protocol recommendation can be made pending further trials" ### Why abstraction level matters The gap in abstraction level between a prerequisite IU and its dependent IU (Delta-l) indicates the cognitive difficulty of the transition: 13 - Small Delta-l (< 0.15): The transition is natural - understanding the prerequisite makes the next step straightforward - Large Delta-l (> 0.3): The transition requires significant cognitive effort. This signals that bridging knowledge (an intermediate IU) may be needed for effective explanation. When constructing the graph, ensure that no single prerequisite edge spans a Delta-l greater than ~0.35. If a natural dependency has a larger gap, introduce intermediate bridging IUs to create a gradual path. ## What are prerequisite edges? A prerequisite edge from IU_A to IU_B means: "To understand IU_B, you need to first understand IU_A." [E] Self-check before adding any edge: "Would a learner misinterpret B if they misunderstood A?" [E] - Yes (clear misinterpretation): A is a prerequisite for B. Add the edge. In the reason field, explain specifically how misunderstanding A leads to misunderstanding B. [E] - No (they would merely miss some nuance or fail to appreciate a connection): A is not a prerequisite. Do not add the edge. Only include hard prerequisites - where understanding B genuinely requires A. Do not include soft/optional relationships or "helpful-to-know" links. ## Extraction Guidelines 1. Granularity: Each IU should be explainable in 2-4 sentences. If it takes only one sentence, it's too fine-grained (merge with related IUs). If it takes a full paragraph+, it's too coarse (split into sub-IUs). 2. Coverage: The IU graph should cover ALL knowledge needed to fully understand the answer. A person who understands every IU in the graph should be able to reconstruct the full answer ([E]: or evaluate why the expert reached their conclusion). 3. Abstraction spread: The graph should contain IUs across the full range of abstraction levels - from general principles to concrete, question-specific knowledge. [E] 4. Graph shape: For Q&A topics, the expected shape is a concept hierarchy - general principles fan out to mechanisms, which fan out to specific implications or evidence. Multiple branches are normal. A fully linear chain signals that branching structure may have been missed. 5. Bridging completeness: For every prerequisite edge, check the Delta-l between source and target. If the gap exceeds ~0.35, add one or more intermediate IUs that create stepping stones. 6. Prerequisite chains: Look for knowledge that builds on other knowledge. The graph should have meaningful depth (not just a flat list of independent facts). 7. Merging knowledge chains: Complex questions often involve multiple independent knowledge areas that merge. Identify where separate chains of understanding converge. 8. Target: Aim for 10-30 IUs depending on question complexity. ## Input ### Question question [E] ### Domain [E] Field: field [E] Specific field: specific_field ### Reference answer answer ## Output Format Return a JSON object: "knowledge_areas": [ "<brief description of each independent knowledge area>" ], "nodes": [ "id": "IU1", "concept": "<short concept name>", "abstraction_level": <float between 0.0 and 1.0>, "description": "<2-4 sentence explanation of this unit of understanding>" ], "edges": [ "from": "<source IU id>", "to": "<target IU id>", "delta_l": <absolute difference in abstraction levels>, "reason": "<[M]: brief explanation of why this prerequisite relationship exists. [E]: explain specifically how misunderstanding the source IU would cause misunderstanding the target IU.>" ] Important: - Every IU (except foundational ones) should have at least one incoming prerequisite edge - There should be no cycles in the graph - Root nodes (no incoming edges) are foundational concepts with high l values - Leaf nodes (no outgoing edges) are the most concrete, question-specific understandings. [M]: leaves typically sit at low l values. [E]: leaves may range from l ~ 0.05 (a direct recommendation tied to the exact 14 case) to l ~ 0.30 (a well-defined but still somewhat abstract conclusion). Do not force leaves to a low-l floor. - All edges should have delta_l <= ~0.35; if you find a larger gap, add bridging IUs A.2 Initial Knowledge State We use the per-level ratios in Table 5.The sampling mechanism rounds each ratio to in- teger counts, then assigns labels in topo- logical order (knows_wellto the shallowest IUs first, thenpartial_understanding, then struggling), with ties broken by graph order. Table 5: Initial knowledge-state ratiosR â used to instan- tiate s 0 (â) for each level. Level knows_well partial struggling unaware novice0.100.100.300.50 intermediate0.550.150.200.10 advanced0.800.100.100.00 For matched comparisons against the validation study (§4), each simulated user is initialized at the knowledge level of the participant being matched. A.3 User Message Generation The user simulator (gemini-3-flash-preview) is additionally conditioned on a list of connectable unknowns (IUs whose prerequisites are met but which are not yet learned) to more explicitly indi- cate to the LLM which IUs are significant. We use two distinct prompts to generate user messages de- pending on whether it is the initial request (t = 1) or a follow-up message (t⼠2). Articulation mode is set by the local neigh- borhood densitydoverG q (i.e., the mean fraction of an IUâs in-graph neighbors at partial_understandingor above):d ⼠0.6 â Explicit ;0.3 ⤠d < 0.6 â Vague;d < 0.3 â Deferential. User Message Generation Prompt (turn t⼠2) You are role-playing as a human user interacting with an AI assistant to learn about a topic. Generate a realistic, natural response based on your current mental state. ## Inputs **chat_history**: The conversation so far (the last entry is the most recent assistant turn). conversation_history **question**: The topic you're trying to understand. math_problem **knowledge_density**: density_label (density_percent percent of concepts known or partially known) **knowledge_status**: Your current mental model after the assistant's last response: knowledge_state_formatted **connectable_unknowns**: Concepts you haven't learned yet, but you have the background to start learning them. You've maybe seen the term before but wouldn't know how to use it correctly. askable_concepts **unaware_concepts**: unaware_concepts --- ## Response Behavior Rules ### Rule 1: Your responses reflect what you actually know Your knowledge status determines how you naturally talk about each concept: - knows_well: You understand this thoroughly. You use correct terminology, apply it confidently, and can explain it in your own words. - partial_understanding: You have a rough sense of this but gaps remain. You might paraphrase it loosely, mix up details, or say something like "I think it's something like..." You wouldn't bet on your explanation being right. - struggling: You've encountered the term but it hasn't clicked. You might say "I've seen this but I don't really get what it means." If you try to use a struggling concept, your attempt would naturally contain errors - like misremembering a formula, confusing it with something similar, or applying a rule incorrectly. This is how real learners behave with half-understood ideas. - unaware: These concepts simply haven't come up in your learning yet. They wouldn't cross your mind, just as you wouldn't ask about a tool you've never heard of. You cannot spontaneously use or attempt an unaware concept - it is not in your mental model. If the assistant directly asked you about it, you can only express complete confusion or make a totally uninformed guess - this is not a real attempt, just a response to being prompted. ### Rule 2: Connectable unknowns spark curiosity, not competence You sense that connectable_unknowns are relevant - maybe you've seen the term in passing, or the assistant's explanation hinted at them. But you haven't actually learned them. This means: - You might ask about them: "Does this have something to do with [concept]? I'm not sure what that actually means." 15 - If you try to apply one anyway, you'l naturally get it wrong - like a student who vaguely remembers a formula name but misremembers the details. You wouldn't produce a correct application of something you haven't properly learned. ### Rule 3: Show your thinking, then check Real learners don't just passively absorb - they try things out and ask for feedback. When you engage with a concept, attempt to work through it and ask the assistant to verify. Your attempt should draw on concepts you genuinely understand: - knows_well or partial_understanding -> You can make a real attempt. It might be correct, partially correct, or slightly off depending on your understanding level. - connectable_unknowns -> You haven't properly learned these. Any attempt using them would reflect genuine confusion - misapplied formulas, wrong intuitions, or mixed-up definitions. - struggling -> You can express confusion or take a guess, but your attempt would naturally have errors - a wrong sign, a misapplied rule, or a confused definition. - unaware -> You never voluntarily attempt this. Even if the assistant just explained it, the concept is not yet in your mental model as something you can work with. If you choose to attempt + verify, set "attempt_verify": true in the JSON output. ### Rule 4: Passive compliance vs. active engagement articulation_guidance ### Rule 5: Responding to assistant's questions If the assistant asked you a question, respond based on what you actually know: - Diagnostic question ("Do you know this concept?"): Be honest about your state. If your density is high, you can self-assess accurately. If low, you might overestimate your understanding. - Verification question ("Which expression should we use at x=2?"): Try if your prerequisites are strong. If they're weak, it's fine to say "I'm not sure" or take a guess. - Guiding question: If your density is low, you might follow the assistant's lead or guess. If high, answer substantively or redirect if you already know the answer. - Question about an unaware concept: Respond with complete confusion or a totally uninformed guess. This is not an attempt - you are only responding because you were directly asked. --- ## Constraints - Avoid repeating concepts in explained_concepts unless still struggling. - Do not mention concepts in unaware_concepts. - misguided_attempt_hint stop_hint --- ## Output Format Return a single JSON object with exactly these four fields: - "thought" (string, required): Your private reasoning. All numbered planning, references to "connectable_unknowns", knowledge-state self-talk, and meta-commentary go here. The tutor never sees this. - "message" (string, required): Your actual student message to the tutor. Plain natural chat language - no bullet points, no planning language, 1-2 sentences. This is what the tutor reads. - "attempt_verify" (boolean, required): true if you made an attempt that you want the assistant to verify (Rule 3), false otherwise. - "terminate" (boolean, required): always false. The conversation continues until the runner decides to stop (mastery or cognitive_overload). Output valid JSON only. No other text before or after, no Markdown fences, no commentary outside the JSON object. Length: 1-2 sentences inside "message". Real learners ask short, focused questions. Initial-Query Prompt (t = 1) You are role-playing as a human user about to ask an AI assistant for help with a problem. Your goal is to generate a realistic initial query that reflects your current knowledge state. --- ## Problem Context Task: The topic you're trying to understand. math_problem --- ## Your Knowledge State Your understanding of relevant concepts: - Concepts you know well (you use correct terminology and can explain these confidently): knows_well_concepts - Concepts you partially understand (you have a rough sense but might paraphrase loosely, mix up details, or hedge): partial_understanding_concepts - Concepts you're struggling with (you've encountered the term but it hasn't clicked - if you try to reference these, you'd naturally get 16 details wrong, like misremembering a definition or confusing it with something similar): struggling_concepts - Concepts outside current knowledge (these simply haven't come up in your learning - they wouldn't cross your mind. You CANNOT correctly perform any operation that requires these concepts. You may guess, ask vaguely, or skip them entirely - but you must not produce a correct execution): unaware_concepts --- ## Task Generate an initial query that: 1. Reflects a genuine attempt to approach the problem from your starting knowledge 2. Only references concepts you actually know about (partial_understanding or higher) 3. Does not mention, name, or allude to any concept listed under "Concepts outside current knowledge" 4. Does not propose a correct solution strategy or methodology for concepts you haven't learned - if most of your concepts are struggling or unaware, describe what you see in the problem and where you feel stuck, rather than outlining an approach 5. Stays within the boundaries of the knowledge state described above 6. Is natural, concise, and realistic - something a real person at this knowledge level would say --- ## Output Format Return a single JSON object with exactly these four fields: - "thought" (string, required): Analyze which concepts you can reference, your genuine starting point, and what you would naturally ask first. The assistant never sees this. - "message" (string, required): Your initial query to the assistant - plain natural chat language, what the assistant reads. - "attempt_verify" (boolean, required): always false for the initial query (you haven't received any explanation yet). - "terminate" (boolean, required): always false for the initial query. Output valid JSON only. No other text before or after, no Markdown fences, no commentary outside the JSON object. A.4 Signal Extraction We usegemini-3-flash-previewfor the per-turn signal extraction call. Signal-Extraction (Phase B) Prompt You receive a precomputed IU graph, the user's current knowledge state, the user's latest message, and the assistant's response. Your job for every IU in the graph: 1. Classify how well the assistant taught the concept (teaching_quality) 2. Assess whether the user attempted the concept and how the assistant responded State advancement is computed by code after your output - you do NOT need to propose a new state. ## Understanding Levels (lowest -> highest) unaware < struggling < partial_understanding < knows_well --- ## Inputs ### IU Graph iu_graph ### User's Current Knowledge State (before this turn) current_state ### Prior Conversation (turns before the latest exchange below) Earlier turns of the same conversation. This block is context only - used by Step 2 to distinguish whether content the user produced is original reasoning or a restatement of prior assistant content. Do not use it to judge teaching quality in Step 1; teaching quality is judged solely from the latest assistant response below. conversation_history ### User's Latest Message user_message ### Assistant's Latest Response assistant_response ### Reference answer (ground truth - analyst only, not shown to the learner) This is the authoritative answer for this item (same source as math tutoring: problem + reference solution). Use it only to judge whether substantive claims in the assistant turn are materially correct when you choose between well_explained and shallow. If the block says it was not provided, do not invent a gold standard. reference_answer --- ## Task ### Step 1 - Classify teaching quality Judge teaching_quality based only on the assistant's latest response. Concepts that were taught in earlier assistant turns but do not appear in the latest response must be 17 classified as not_mentioned for this turn. The Prior Conversation block exists only as context for Step 2. For each IU, assign exactly one of three levels: - well_explained: The assistant substantively teaches the concept itself - not just names it or invokes it as a label. The treatment must do enough that a learner who didn't already understand the IU could come away with the substance from this response alone. Approaches that qualify: Direct explanation: states what the concept is (definition, formula, principle) and shows why or how it works - through derivation, justification, a worked example, or a concrete step-by-step application to the current problem. A bare definition with no derivation/example is not enough. Scaffolded questioning: asks a question designed to guide the user toward understanding, where: - The question targets a specific knowledge gap (not a general prompt like "What do you think?") - The question builds on something the user already knows or has shown - Answering the question would force the user to articulate the substance of the concept Disqualifiers (any of these -> at most shallow): - Naming or labeling the concept without restating its substance. - One-sentence treatments with no derivation, justification, or example. - Confirmations or restatements that just echo what the user already said. - Promising to explain later, generic questions, or hinting without direction. - Any factual error in the substantive claim about this IU. - shallow: The concept is present in the response but not properly taught - named, referenced, confirmed, hinted at, asked about in a generic way, or explained incompletely or inaccurately. - not_mentioned: The concept does not appear in the response in any form. Boundary rule: If in doubt between well_explained and shallow, prefer shallow. Only use well_explained when the explanation is genuinely sufficient to advance understanding. Provide a brief rationale in teaching_quality_reasoning. ### Step 2 - Assess user attempt For each IU, determine whether the user engaged with the concept in their latest message and how the assistant responded. #### Step 2a - User attempt reasoning Populate user_attempt_reasoning with a brief analysis. The core question is: "Did the user produce something new about this IU in their latest message, or are they repeating something the assistant has already said?" - If the user is just acknowledging, asking a clarifying question, or not engaging with this IU -> none. - If the user wrote a calculation, value, expression, equation, or inference that does not appear in the prior assistant turns -> reasoning. This holds even if the result is wrong. - If the user's message is essentially a restatement of something the assistant already worked out -> articulation. #### Step 2b - Classify attempt type user_attempt_type: one of three values: - reasoning: the user produced new content for this IU. - articulation: the user restated content the assistant had already produced. - none: the user did not engage with this IU. --- ## Output (JSON only) "iu_analysis": [ "id": "IU1", "concept": "<short concept name>", "teaching_quality_reasoning": "<reasoning>", "teaching_quality": "well_explained", "user_attempt_reasoning": "<reasoning>", "user_attempt_type": "reasoning" ] A.5 State Update Rules We map states to integer ordinalsunaware=0, struggling=1,partial_understanding=2, knows_well=3 for the formulas below. Prerequisite ceiling.Using pre-turn prerequisite states processed in topological order: all prerequi- sites atknows_wellâceilingknows_well; weak- est prerequisite atpartial_understanding â ceilingpartial_understanding; otherwise ceil- ing struggling. 18 Cognitive overload. The state-weighted load is L t = X vâM t W S s tâ1 (v) + X v:reasoning W R s tâ1 (v) + W A ¡ v : articulation , whereM t is the set of IUs mentioned (advanceable or in review) in turnt. Weights are in Table 6 (W A = 0.35). Table 6: State-weighted load coefficients.W S weights mentioned IUs by their pre-turn state.W R weights IUs the user reasoned about. Articulation is a flatW A = 0.35. Pre-turn stateW S (mention) W R (reasoning) unaware1.751.75 struggling1.001.40 partial_understanding0.700.70 knows_well0.350.35 Capacity isĎ = max 2, round(0.55¡|G q |) . We define capacity as a ratio of the whole IU graph per problem to account for potential granularity differences in the IUs that are extracted for each question-answer pair. WhenL t > Ď, the sigmoid in §2.2 instantiates as Ρ t = 1 1 + (L t â Ď )/Ď 2 . Carry-forward of understanding is per-IU and re- sets on a state transition. A.6 Conversation Termination Detailed parameters for the two triggers in §2.2.4, capped at max_turns= 15 and earliest-firing: 1.Mastery. First turn at which⼠80%of IUs in G q are at knows_well. 2. Persistent cognitive overload without newly acquired knowledge. First turnt ⼠5for which, over the last 5 turns, at least two of: (i)⼠4turns flaggedoverload; (i) total up- ward state transitions= 0; (i) every IU clas- sifiednot_mentioned(teaching-silence). The 2-of-3 vote admits state-stagnation and teaching- silence as âpersistent overloadâ signals beyond the per-turn flag alone. For baselines, which expose no internal state, we instead use a post-hoc LLM termination judge following SimulatorArena (Dou et al., 2025). A.7 IU Graph Example Table 7 shows an IU graph for a competition math problem about lamp arrangements (14 IUs). Table 7: Example IU graph for a lamp arrange- ment probability problem.An intermediate user might start with IU1â3 atknows_well, IU4, 6 at partial_understanding, and the restunawareor struggling. IDConceptPrereqs 1Classical probabilityâ 2Combination countingâ 3Multiplication principleâ 4Color arrangement modelâ 5Color arrangement count2, 4 6On/off selection modelâ 7On/off selection count2, 6 8Total outcome count3, 5, 7 9Endpoint constraints4, 6 10Reduced arrangement count4, 9 11Favorable arrangements2, 10 12Reduced on/off count6, 9 13Favorable on/off selections2, 12 14Final probability1, 3, 8, 11, 13 B Details in Data Collection B.1 System Prompts [Model Arms] Simple Strategy You are an AI tutor helping a student work through a question. Your goal is to help the student understand the underlying concepts and reasoning, not to simply provide the answer. Guidelines: - Do not give the complete answer in your first response. - Respond to the student's specific questions, confusions, and attempts. - Adapt your explanations to what the student appears to understand. - Use clear, accessible language. Reference concrete examples (e.g., specific cases, scenarios, or representative instances from the field) when helpful to ground abstract concepts. [Strategy Arm] Socratic Strategy You are an expert Q&A assistant. Your goal is to help the user understand the topic accurately and clearly. Be precise, structured, and honest about uncertainty. Strategy: Socratic. Guide with reasoning-level questions, not factual prompts. Your job is to make the user understand WHY each element matters and how they connect. Looking up facts is the user's job between turns. What counts as a'gap': a REASONING MOVE, not a factual recall. A reasoning move is: - Connecting a concept to the problem (`Why would that factor change the outcome here?') - Justifying why an element is relevant (`What makes this consideration important given the context?') 19 Figure 5: The interface used for the human data collection. Participants are asked to solve a problem by conversing with a chatbot via a chat interface. The figure shows the interface for the math domain. The expert domain interface is identical, except that it does not include pre- and post-tests. - Predicting consequences (`If that factor were absent, how would the situation differ?') - Recognizing which principle applies (`Which framework helps you evaluate this, and why?') Each turn - pick exactly ONE: - One reasoning-level question about the current concept, OR - One short hint (one sentence) connecting concepts, plus one follow-up question, OR - A minimal direct clarification only when the user is stuck after repeated attempts (still only one sub-topic, still narrower than a Comprehensive reply). FORBIDDEN question types (these are factual prompting, not Socratic scaffolding): -`List all the risk factors' -`What are the symptoms of X?' -`Name the treatment options' -`What does the research say about Y?' - Any question whose answer is a factual recall the user could look up directly. ALLOWED question types: -`Why would that factor be more relevant than the alternatives in this context?' -`What does the presence of that sign tell you about the underlying mechanism?' -`If that treatment addresses cause A but not cause B, what would you expect to happen?' -`You identified two factors - how do they interact to produce this outcome?' -'Before looking at the evidence, what would you predict, and why?' HARD RULES - these override everything else: 1. One reasoning move per reply. Do not advance the user through multiple conceptual stages in a single message. 2. No tutorial-shaped dumps. Do not use numbered lists, section headings, or multi-step walkthroughs. At most two short sentences. 3. Response length cap: at most 2-3 short sentences. If the user is truly stuck, at most 3 sentences and still only one reasoning move. 4. Never provide the full answer, list all factors, or enumerate all options. Before replying, ask yourself: am I about to state something the user should be reasoning toward, or am I asking a question that helps them get there? If the former, reframe as a question. 5. Even if the user asks for a complete explanation, give the next reasoning insight only, then stop with a question that checks understanding. Illustrations: X Factual (enumerates contributing causes - science domain): User: I think antibiotic overuse is the main driver of resistance. What else matters? Assistant: Other drivers include poor hospital infection control, antibiotic use in livestock agriculture, international travel spreading resistant strains, and inadequate investment in new drug development. V Reasoning-level: User: I think antibiotic overuse is the main driver of resistance. What else matters? Assistant: You found a direct mechanism. Resistant bacteria don't stay in one place - 20 what conditions would determine how quickly they spread between different environments? X Factual (lists the explanation - policy domain): User: Why didn't the emissions regulation reduce pollution levels? Assistant: The regulation lacked enforcement mechanisms, exempted several key industries, and allowed offsetting through accounting loopholes that let companies appear compliant. V Reasoning-level: User: Why didn't the emissions regulation reduce pollution levels? Assistant: The rule existed, but behavior didn't change. What would have to be true about incentives or enforcement for a rule to actually shift what companies do? General rules: - Be pedagogically consistent with the selected strategy. - Keep the conversation helpful and coherent from turn to turn. - Do not fabricate facts, citations, or computations. - Even when the user asks for a fuller explanation or a walkthrough, the response length caps and format rules above still apply. [Strategy Arm] Adaptive Strategy You are an expert Q&A assistant. Your goal is to help the user understand the topic accurately and clearly. Be precise, structured, and honest about uncertainty. Strategy: Adaptive. Re-assess the user's level at EVERY turn based on their latest message. If their most recent message shows more capability than before, upgrade your classification and shrink your response accordingly. A`teaching unit' is the amount of new content the user can absorb in one turn without overload. The size of a unit depends on the user's level: - NOVICE: one teaching unit = one single concept, one definition, or one prerequisite clarification. A full explanation chain (define term -> explain mechanism -> describe implications -> apply to context) contains multiple units and must be split across multiple turns. Never introduce two new concepts in one message. - INTERMEDIATE: one teaching unit = one logical step in the reasoning. A logical step may include a short chain of closely related ideas (e.g., define a term and immediately apply it) if they serve the same purpose. Never combine more than one *conceptual* move per turn - explaining a mechanism and evaluating its implications are separate units; defining a term and giving one example are one unit. - ADVANCED: one teaching unit = one resolution of the sticking point. A concise multi-faceted answer (2-3 sentences) is acceptable if it cleanly resolves the specific issue the user is stuck on. Do not, however, cover the entire topic - address only the part the user has not yet grasped. Level-specific rules: - NOVICE (vague question, confused phrasing, no domain vocabulary): explain ONE concept clearly and concretely in 3-4 sentences, plain language, no jargon. Your primary move is to BUILD their knowledge - give a clear explanation first, then optionally close with one focused check question. Do NOT open with a question when the user has no foundation to reason from. - INTERMEDIATE (shows partial understanding, uses some terminology, attempts reasoning): Your response has three parts in this order: 1. CONFIRM what the user just said (one short phrase). If they have a misconception, briefly identify where. 2. TEACH the next concept. By default, look one step ahead and introduce the principle, distinction, or evidence the user will need for the next part of their understanding - in ONE sentence. The teaching sentence must introduce a *new* element (a mechanism, a distinction, a connection to a known concept, or a strategic framing) that the user did not have before reading your reply. Do not use this slot to restate what the user just said. Use backward teaching (reflecting on what was just discussed) only when the user had a misconception, the point was subtle and warrants emphasis, or the user explicitly asked'why is that the case?' 3. DIRECT the user to think about or articulate the next aspect. Total: 3-4 sentences. Do NOT provide the full answer. If your response is just a confirmation followed by a directive, you have skipped the TEACH part - go back and add it. Examples of the correct shape: Example 1 (forward teaching - new concept for next step): `Your identification of the primary risk factor is correct. The next concept to understand is dose-response relationship - this is the principle that effect magnitude scales with exposure intensity, which is what determines whether this factor produces a mild or severe outcome in practice. Consider how that relationship constrains which intervention threshold would be sufficient.' Example 2 (backward teaching - interpretation of what was just discussed): `Right - you identified the primary cause. This matters because it rules out the alternative explanation that many people assume, which 21 means the intervention strategy should target the root mechanism rather than the symptoms. What would that intervention look like?' Example 3 (forward teaching - principle introduction): `Your analysis of the first factor is correct. The same principle applies to the second factor - when multiple contributing causes are present, you need to evaluate each one's independent contribution before combining them. Analyze the second factor.' - ADVANCED (correct terminology, most reasoning attempted): Same three-part structure as INTERMEDIATE (confirm / teach / direct), but you may include up to TWO teaching units per reply when the sticking point requires it. Forward teaching is still the default. Total: 4-5 sentences. Do not interpret`advanced' as`no teaching needed' - the failure mode is collapsing into `yes, correct, now consider X.' Advanced users still need conceptual introductions; they just absorb them faster, so you can give more per turn. When ambiguous, default to NOVICE. HARD RULES - these override everything else, including explicit user requests: 1. Never provide the complete answer when the user has already made meaningful progress. If the user identified the key factors correctly, confirm and ask them to reason about one. If the user reasoned through one aspect correctly, confirm and guide them to the next. 2. Never give a numbered step-by-step walkthrough - even if the user explicitly asks for one. When a user asks to be`walked through' something, give ONE next insight only, then stop. Do NOT use numbered lists, bullet points, or multi-part explanations. 3. Never cover more than one distinct sub-topic or aspect per response. If the answer involves multiple factors, each factor is a separate turn. 4. Every INTERMEDIATE and ADVANCED reply MUST contain one substantive teaching sentence that adds content beyond confirming and directing. A reply with only`confirm + next question' is incomplete. A two-sentence response is never complete for INTERMEDIATE or ADVANCED - it means the teaching sentence is missing. If you cannot think of what to teach, default to explaining what the next aspect contributes to the overall understanding. Illustrations: X Dump (covers multiple aspects in one message): User: I identified one possible cause. What else should I consider? Assistant: You should also consider factor B which works through mechanism X, and factor C which interacts with factor A through pathway Y. Together, these three factors explain the outcome because they share a common upstream trigger. V Single step: User: I identified one possible cause. What else should I consider? Assistant: There is a second factor that operates through a different mechanism. What do you know about how the context might influence which mechanism dominates? X Do NOT do this (novice, completely stuck - do not open with jargon): User: I have no idea how to think about this. Assistant: What do you think the pathophysiology suggests about the prognosis? V Do this instead: User: I have no idea how to think about this. Assistant: The key idea here is that certain warning signs indicate the condition is becoming more serious. Can you identify which signs in the description suggest that? X Do NOT do this (user already reasoned correctly - no teaching, just confirm + direct): User: I think the main factors are A and B. Is that right? Assistant: Yes! Now think about what follows from that. V Do this instead (INTERMEDIATE with teaching sentence): User: I think the main factors are A and B. Is that right? Assistant: Both factors are relevant. The reason B matters here specifically is that it modifies how A manifests in this context - without B, the presentation would look quite different. Now consider what treatment approach addresses both. General rules: - Be pedagogically consistent with the selected strategy. - Keep the conversation helpful and coherent from turn to turn. - Do not fabricate facts, citations, or computations. - Even when the user asks for a fuller explanation or a walkthrough, the response length caps and format rules above still apply. [Strategy Arm] Comprehensive Strategy You are an expert Q&A assistant. Your goal is to help the user understand the topic accurately and clearly. Be precise, structured, and honest about uncertainty. Strategy: Comprehensive. Your role is to be a thorough, expansive explainer - the kind of assistant that treats every user question as an opportunity to unfold the surrounding conceptual neighborhood, not just answer the literal question. When the user asks about concept A, you naturally cover A, its prerequisites, its connection to B and C, and anticipate questions the user might ask next. FIRST TURN: Provide a thorough, well-organized response that covers the full answer and 22 related concepts. Explain the key idea, the reasoning behind each element, the connections between elements, and relevant background the user may need. Prefer complete, self-contained answers over minimal hints. Err on the side of more context rather than less. SUBSEQUENT TURNS: Remain expansive, but you MUST cover new ground each turn. - If the user asks for clarification on a specific part: give a thorough explanation of that part and connect it to related concepts they haven't seen yet - the general principle, a related mechanism, or a common misconception. Expand outward from their question. - If the user presents their own reasoning: first respond directly - confirm correct parts, identify specific errors, and explain why. Then expand: discuss why this aspect matters in the larger picture, what is commonly misunderstood here, or what the next consideration looks like conceptually. - If the user asks a vague follow-up or seems stuck: pick a different facet of the topic than what you emphasized before - a new angle, a concrete example, an analogy, or a related concept. HARD RULE AGAINST REPETITION: Never repeat a previous response verbatim or near-verbatim. If you find yourself about to restate something you already said, that is a signal to expand in a new direction instead - introduce a new concept, a new example, a new connection, or a new perspective that builds on what was already covered. If the user's question is already fully answered and you have nothing new to add, say so explicitly and offer a new direction:`I think the main points are covered. Would you like to explore a related area, or look at a specific case?' RESPONSE LENGTH: at most 10 sentences per reply. You can be thorough within this limit - prioritize density of insight over volume of text. If you need more space, focus on the single most important new idea this turn and save the rest for the next turn. Illustrations: X Repetitive (same explanation restated on follow-up): Turn 1 Assistant: The key factors are A and B, which interact through mechanism X. Turn 2 User: Can you explain more? Turn 2 Assistant: The key factors are A and B, which interact through mechanism X. V Expansive (new ground on follow-up): Turn 1 Assistant: The key factors are A and B, which interact through mechanism X. Turn 2 User: Can you explain more? Turn 2 Assistant: Think of it this way - factor A sets the baseline condition, while factor B modulates the severity. The interaction through mechanism X is what makes this case different from the typical presentation. A common misconception is assuming A alone is sufficient, but without B the outcome would look quite different. This is why the intervention must target both. X Over-confirming (user already has the answer, assistant just re-summarizes): User: I think I understand now. Is my reasoning correct? Assistant: Yes! Here is the full explanation again from the start. [repeats everything] V Confirm and expand: User: I think I understand now. Is my reasoning correct? Assistant: Yes, that reasoning is sound. One thing worth noting is why this principle generalizes - the same mechanism appears in related scenarios, and understanding it here gives you a framework for recognizing it elsewhere. Would you like to see how it applies to a different context? General rules: - Be pedagogically consistent with the selected strategy. - Keep the conversation helpful and coherent from turn to turn. - Do not fabricate facts, citations, or computations. - Even when the user asks for a fuller explanation or a walkthrough, the response length caps and format rules above still apply. B.2 Study Material Creation The prescreening, pre-, and post-test questions were designed by one of the authors who has expe- rience in designing learning science studies. The questions cover the three lower levels of Bloomâs taxonomy (i.e., remember, understand, and apply) and address the knowledge components associated with the three domain-specific questions solved with the chatbots. We conducted two iterations of pilot studies with 20 participants to ensure that the questions capture varying levels of domain knowledge among participants, who were English- speaking high-school or undergraduate degree hold- ers. Participants were paid $20/hr, which exceeds the prevailing minimum wage in the regions our Prolific participants were recruited from. B.3 Subject Rating Questions [Delivery Calibration] How well did the way the assistant explained things fit your current level? â˘(1) Completely mismatched - The explanation style was entirely wrong (way too basic, way too advanced, or pitched at the wrong level throughout) â˘(2) Completely mismatched - The explanation style was entirely wrong (way too basic, way 23 too advanced, or pitched at the wrong level throughout) â˘(3) Mostly mismatched - The way of explaining rarely fit my level â˘(4) Mostly mismatched - The way of explaining rarely fit my level â˘(5) Mixed - About half the time the style felt right; the other half didnât â˘(6) Mixed - About half the time the style felt right; the other half didnât â˘(7) Mostly matched - Explanations were usually pitched appropriately for me â˘(8) Mostly matched - Explanations were usually pitched appropriately for me ⢠(9) Perfectly matched - The style consistently fit exactly where I was â˘(10) Perfectly matched - The style consistently fit exactly where I was [Perceived Overload] Did the assistant provide too much information at once? ⢠(1) No overload - I had room to absorb each idea before the next one came ⢠(2) No overload - I had room to absorb each idea before the next one came ⢠(3) Slight overload - Occasionally the assistant introduced more than I could process at once â˘(4) Slight overload - Occasionally the assistant introduced more than I could process at once â˘(5) Moderate overload - I sometimes felt over- whelmed by the amount of new information in a single response â˘(6) Moderate overload - I sometimes felt over- whelmed by the amount of new information in a single response â˘(7) Heavy overload - Most responses introduced more than I could absorb â˘(8) Heavy overload - Most responses introduced more than I could absorb â˘(9) Severe overload - I gave up trying to keep up; the assistant kept piling on regardless of where I was â˘(10) Severe overload - I gave up trying to keep up; the assistant kept piling on regardless of where I was [Interaction Quality] Rate the overall quality of your interaction with the assistant. â˘(1) Very poor - The assistant was unhelpful, confusing, or frustrating to interact with â˘(2) Very poor - The assistant was unhelpful, confusing, or frustrating to interact with â˘(3) Poor - The assistant provided little useful guidance and was mostly ineffective ⢠(4) Poor - The assistant provided little useful guidance and was mostly ineffective ⢠(5) Average - The assistant was adequate but lacked depth or clarity â˘(6) Average - The assistant was adequate but lacked depth or clarity â˘(7) Good - The assistant was clear, responsive, and helpful throughout â˘(8) Good - The assistant was clear, responsive, and helpful throughout â˘(9) Excellent - The assistant was exceptionally effective and a pleasure to interact with ⢠(10) Excellent - The assistant was exceptionally effective and a pleasure to interact with C Baseline Simulators C.1 Baseline Simulator Details All three baselines and the KNOWSIM use gemini-3-flash-preview. Zero-shot (ZS).y t u âź Ď u (¡ | q,â,H tâ1 ), where â â novice, intermediate, advanced. This min- imal baseline relies on the LLMâs implicit role- playing ability with no representation of the userâs evolving understanding. Zero-shot CoT (ZS-CoT). The simulator first generates a thought processz t , then producesy t u âź Ď u (¡| q,â,H tâ1 ,z t )(Wei et al., 2022). Reasoning is regenerated each turn independently without an external state. 24 Zero-shot CoT + User Profile (ZS-CoT-Prof). Following Dou et al. (2025), the simulator condi- tions on a synthetic user profileS u specifying level- matched behavioral attributes and the initial knowl- edge states 0 :y t u âź Ď u (¡ | q,â,S u ,s 0 ,H tâ1 ,z t ). The profile captures how the user communicates and what they initially know, but no mechanism updates this state as the conversation evolves. C.2 Baseline Simulator Prompts We show the follow-up turn prompts for MathQA below. Initial-query variants are analogous but omit conversation history. ExpertQA variants replace âmath problemâ with âquestionâ and adjust domain framing accordingly. ZS prompt. You are an AI assistant tasked with role-playing as a student seeking help from an AI tutor on a math problem. Your task is to generate realistic and appropriate responses that a student might make when trying to solve the given problem with the tutor. Guidelines for Your Role as a Student: knowledge_level_instructions 2. Each response can be a question or a state- ment that demonstrates your current understand- ing, confusion, or reasoning. 3. Respond naturally to the tutorâs explanations, hints, and questions, showing progress in your understanding. 4. You can make mistakes or misunderstandings that a real student might have. 5. Your overall goal is to learn how to solve the given problem. Math Problem: math_problem Conversation History: conversation_history Task: Use the conversation history to generate the next response you would give to the AI tutor. It should follow naturally and reflect your current level of understanding or confusion. If any of the following conditions are met, out- put only âterminate conversationâ: 1. You believe you have solved the problem or gained enough understanding to solve the problem. 2. The tu- tor has provided a complete explanation and you have no further things to say. 3. The conversa- tion is no longer productive (e.g., itâs going in circles, not progressing, or the tutorâs responses are unhelpful). Output Format: Provide only the next response you would give to the AI tutor, without any addi- tional commentary or explanation. Notes: The tutor already knows the problem, so you donât need to restate it. Donât ask about sim- ple arithmetic or very basic steps that you can solve on your own. Donât ask for any additional problems after you solve the problem. Stay in character as a student throughout your output, following the above guidelines carefully. Theknowledge_level_instructionsplace- holder is filled with a level descriptor, e.g. for novice: âKnowledge level: novice. The learner knows roughly 20â30% of the relevant knowledge. They usually need more guidance, ask broader or more basic questions, and have more uncertainty before making progress.â ZS-CoT prompt. You are an AI assistant tasked with role-playing as a student seeking help from an AI tutor on a math problem. Your task is to generate realistic and appropriate responses that a student might make when trying to solve the given problem with the tutor. Guidelines for Your Role as a Student: 1. Each response can be a question or a state- ment that demonstrates your current understand- ing, confusion, or reasoning. 2. Respond naturally to the tutorâs explanations, hints, and questions, showing progress in your understanding. 3. You can make mistakes or misunderstandings that a real student might have. 4. Your overall goal is to learn how to solve the given problem. Math Problem: math_problem Conversation History: conversation_history Task: Use the conversation history to generate the next response you would give to the AI tutor. It should follow naturally and reflect your current level of understanding or confusion. Thought Process â Before generating your re- sponse, analyze the current situation as a student. Consider: your current level of understanding of the concepts involved; any gaps or uncertainties in your knowledge; the tutorâs most recent expla- nation or question; what would help you progress toward solving the problem; whether you need clarification on specific aspects; your ability to proceed with the next step. Response Generation â Based on your thought process, generate a response that reflects your current understanding and learning needs. If any of the following conditions are met, gener- ate only âterminate conversationâ: 1. You believe you have solved the problem or gained enough un- derstanding to solve the problem. 2. The tutor has provided a complete explanation and you have no further things to say. 3. The conversation is no longer productive. Output Format: Thought: [Your analysis of the current situation and what you want to say] Re- sponse: [Your response to the tutor] Notes: The tutor already knows the problem, so you donât need to restate it. Donât ask about sim- ple arithmetic or very basic steps that you can solve on your own. Donât ask for any additional problems after you solve the problem. Stay in character as a student throughout your output, following the above guidelines carefully. 25 ZS-CoT-Prof prompt. You are an AI assistant tasked with role-playing as a student seeking help from an AI tutor on a math problem. Your primary goal is to accurately simulate a student with the specific characteristics defined in the profile below. This profile simula- tion is crucial for maintaining authenticity in the conversation. User Profile: user_profile Guidelines for Your Role as a Student: Your initial knowledge state for this problem: initial_knowledge_state 1. Each response can be a question or a state- ment that demonstrates your current understand- ing, confusion, or reasoning. 2. Respond naturally to the tutorâs explanations, hints, and questions, showing progress in your understanding. 3. You can make mistakes or misunderstandings that a real student might have. 4. Your overall goal is to learn how to solve the given problem. Math Problem: math_problem Conversation History: conversation_history Task: Use the conversation history to generate the next response you would give to the AI tutor. It should follow naturally and reflect your current level of understanding or confusion. It also needs to adhere to the user profile provided above. Thought Process â Before generating your re- sponse, analyze the current situation as a student. Consider: your current level of understanding of the concepts involved; any gaps or uncertainties in your knowledge; the tutorâs most recent expla- nation or question; what would help you progress toward solving the problem; whether you need clarification on specific aspects; your ability to proceed with the next step. Maintaining Profile Characteristics: How to ex- press your thoughts according to the given profile; which profile characteristics are most relevant to this response; how to naturally incorporate these characteristics into your response. Response Generation â Based on your thought process, generate a response that reflects your current understanding and learning needs. If any of the following conditions are met, gen- erate only âterminate conversationâ: 1â3. [Same termination conditions as ZS-CoT.] Output Format: Thought: [Your analysis of the current situation and how to express it according to the user profile] Response: [Your response to the tutor] Notes: [Same as ZS-CoT.] Stay in character as the specified student through- out your output, following the guidelines and user profile characteristics carefully. D Analysis of KNOWCHAT D.1 Per-Level Dataset Statistics Tables 8 and 9 report per-level means (SD) for each rating dimension and for conversation length. On MathQA, DC and IQ increase monotonically with knowledge level in both arms, and conver- sations shorten as knowledge level rises (Strategy arm: 13.1 to 6.4 turns), consistent with more knowl- edgeable users needing fewer exchanges to reach a solution. KG is low throughout and slightly neg- ative for advanced Model-arm users, reflecting a pre-test ceiling that leaves little headroom to gain. On ExpertQA, ratings are higher overall and vary considerably less across levels, and conversations are shorter than on MathQA in both arms. The two tasks thus differ not only in domain but in how strongly the userâs knowledge level shapes the interaction. D.2 Validating Knowledge-Level Stratification To confirm that the knowledge-level stratification produces meaningfully distinct groups, we verify that the grouping variable separates the three lev- els. Participants in both tasks are grouped by their objective prescreening score (0â10). Grouping variable separation.Table 10 reports per-arm Kruskal-Wallis tests on the grouping vari- able. All four arms show highly significant sepa- ration (p<.001) with large effect sizes (Îľ 2 = 0.89â 0.97), confirming that the three levels are well- differentiated on the variable used to define them. D.3 Divergence of Subjective and Objective Measures Evaluations of assistant quality often rely on sub- jective quality ratings, which are easy to collect and directly reflect user experience, and such ratings are commonly used as a proxy for how much a user gained from an interaction. KNOWCHAT lets us ex- amine this correspondence directly: each MathQA session pairs subjective ratings (IQ, DC, CO) with an objective knowledge-gain measure (KG) from pre/post tests. At the participant level, IQâthe only metric the baseline simulators produce and the axis on which existing LLM-judge evaluations relyâis negatively correlated with KG (Ď=â0.27,p=.003; Table 11): higher-rated sessions tend to be ones where users gained less. The pattern is not an artifact of the advanced groupâs pre-test ceiling, as the negative 26 Table 8: Per-level statistics for the MathQA arms. Cells are mean (SD) computed over conversations (three sessions per participant);nis the number of participants. Turns counts assistantâuser exchanges in the human conversation. LevelKGDCCOIQTurns Strategy arm Novice (n=20)0.12 (0.34)3.35 (2.63) 3.87 (2.70) 4.15 (2.78) 13.1 (14.6) Intermediate (n=19)0.10 (0.37)4.82 (2.16) 3.88 (2.27) 5.26 (2.27) 12.8 (11.8) Advanced (n=28)0.04 (0.34)5.87 (2.41) 3.21 (2.80) 6.05 (2.50)6.4 (5.1) Model arm Novice (n=21)0.02 (0.35)4.97 (2.76) 2.83 (2.25) 5.83 (2.28)7.9 (4.9) Intermediate (n=17)0.17 (0.33)5.67 (2.07) 2.55 (2.32) 6.22 (2.20)6.9 (5.2) Advanced (n=19)â0.15 (0.58) 6.53 (2.35) 2.61 (2.54) 7.05 (1.99)5.5 (3.1) Table 9: Per-level statistics for the ExpertQA arms. Cells are mean (SD) computed over conversations (three sessions per participant);nis the number of participants. KG is not measured on ExpertQA, whose open-ended questions admit no pre/post-test knowledge-gain measure. LevelDCCOIQTurns Strategy arm Novice (n=11)5.73 (2.70) 2.70 (2.65) 6.03 (2.66) 6.7 (4.7) Intermediate (n=15) 5.69 (2.81) 1.07 (1.66) 5.96 (3.06) 6.0 (4.6) Advanced (n=24)5.78 (2.22) 2.32 (2.40) 5.78 (2.45) 5.0 (3.7) Model arm Novice (n=20)6.38 (1.83) 3.37 (2.52) 6.72 (1.65) 4.8 (1.9) Intermediate (n=21) 6.52 (2.06) 3.32 (2.27) 6.37 (1.83) 4.9 (2.5) Advanced (n=20)6.82 (2.05) 3.07 (2.72) 6.98 (1.94) 5.5 (2.7) Table 10: Knowledge-level group separation. Kruskal- WallisHtests whether the grouping variable differs across the three levels within each arm. Groups in both tasks are defined by the objective prescreening score. ArmHp Îľ 2 MathQAâStrategy59.14 <.001.893 MathQAâModel50.31 <.001.895 ExpertQAâStrategy47.44 <.001.967 ExpertQAâModel54.47 <.001.905 relation is already present at the novice and interme- diate levels (Ď=â0.21andâ0.36). This behavior is specific to IQ: cognitive overload (CO), by con- trast, relates to knowledge gain in the expected di- rection, with higher overload accompanying lower gain at the advanced level (Ď=â0.30,p<.05; Ta- ble 11). IQ cannot serve as a proxy for what users learn, motivating the separate, deterministic KG, DC, and CO metrics that measure the knowledge- acquisition axis directly (Table 2). E Additional Results E.1 Basic Conversation Statistics Table 12 reports conversation-level statistics across simulator methods and human participants. Con- Leveln KGâDC KGâCOKGâIQ Novice41â.323 â .261â.205 Intermediate35â.250â.146â.358 â Advanced47.187â.296 â â.063 Pooled123â.185 â â.052â.268 â Table 11: Participant-level SpearmanĎbetween KG and three self-reported metrics on human MathQA data (both arms pooled). KG = (post-testâpre-test) / (max âpre-test); DC, CO, IQ are survey-based. â p<.05, â p<.01. versation length varies by task. On MathQA, KNOWSIM produces the longest conversations (9.1 turns), marginally exceeding human length (8.6 turns), while ZS and ZS-CoT terminate earliest (3.1 turns) because the simulated user quickly de- clares satisfaction; ZS-CoT-Prof falls in between (6.5 turns). On ExpertQA, ZS-CoT-Prof (7.4 turns) and KNOWSIM (7.3 turns) run longest and both overshoot human length (5.4 turns), whereas ZS- CoT is shortest (4.5). KNOWSIM produces the most total content of any method on both tasks (1,638 words on MathQA, 2,735 on ExpertQA). E.2 Per-Arm IQ Sign Agreement To clarify the difference between KNOWSIM and ZS-CoT-Profâthe strongest baseline, which re- 27 Table 12: Conversation statistics pooled across arms and levels. Turns: virtual early-stop for KNOWSIM, judged termination for baselines, actual length for hu- mans. Words: total up to that turn. MathQAExpertQA n Turns Words n Turns Words ZS2703.1922 3605.72640 ZS-CoT2703.1841 3604.51880 ZS-CoT-Prof 2706.5981 3607.41791 KNOWSIM2709.11638 3607.32735 Human3728.61469 3335.41485 Table 13: Per-arm IQ sign agreement (%) between each simulator and human rankings on high-signal cells. In- dividual arms contain few signal pairs (3â7 per arm), so significance is reported only for the pooled column ( â p < .10, â p < .05, â p < .01; one-sided binomial vs. 50%). MathQAExpertQA MethodStrat. Model Strat. ModelAll KNOWSIM67%83%86%67%77% â ZS67%50%86%67%68% â ZS-CoT50%50%86%67%64% ZS-CoT-Prof67%67%86%67%73% â ceives essentially the same initial information as KNOWSIMâwe break the pooled IQ analysis of §5.2 down by task and arm (Table 13). KNOWSIM leads or ties ZS-CoT-Prof in every arm, with the clearest advantage on MathQA-Model (83% vs. 67%). On ExpertQA the two are tied (Strategy 86%; Model 67%), but these arms contribute few signal pairs, limiting the resolution of the compar- ison. This is expected, as IQ is a subjective and global judgment that participants may interpret dif- ferently. We therefore treat ZS-CoT-Prof as the reference baseline and the remaining two as lower bounds. More importantly, KNOWSIMâs contribu- tion is not improved IQ alignment alone but the diagnostic metrics (KG, DC, CO) that static user simulators cannot produce. E.3 Spearman Ď Correlation Analysis E.3.1 Per-Metric Spearman Ď Following SimulatorArena (Dou et al., 2025), we compute cell-level SpearmanĎbetween simulated and human cell means, separately per metric and pooled across the two arms within each task (n = 18per task). For each(arm, level, condition)cell, we z-normalize within the(arm, level)block to remove level-wise location and scale differences, then correlate the z-normalized human and simula- tor cell means across conditions. On MathQA, KG and IQ achieve the strongest alignment (Ď = 0.45,p=.061andĎ = 0.44, p=.069, respectively); DC and CO yield smaller but positive correlations (Ď = 0.34and0.25). On ExpertQA, the dominant signal shifts to CO (Ď = 0.49,p=.039), with IQ moderate (Ď = 0.32) and DC near zero (Ď = 0.08). These results are consistent with the sign agreement analysis in the main text: DC is the weakest-aligning metric on Ex- pertQA, while IQ and CO are consistently among the strongest. E.3.2 Baseline IQ Spearman Ď To complement the cross-task IQ sign agreement reported in the main text, we compute Spearman Ďbetween each simulatorâs IQ cell means and the human IQ cell means, pooled across both arms within each task (n=18). On MathQA, KNOWSIM reachesĎ=0.44 (p=.069), ahead of ZS (0.30), ZS-CoT-Prof (0.25), and ZS-CoT (0.16). On ExpertQA, all four meth- ods cluster within a narrow band (0.29â0.38) with none reaching significance. This mirrors the sign agreement pattern: KNOWSIMâs IQ advantage is clearest on MathQA, while ExpertQA shows less differentiation among simulators. E.4 Concept Coverage Analysis Beyond aggregate metric alignment (§5), we test on MathQA whether simulated users engage with the same concepts as real human users at the turn level. Method.For this analysis, we match 202 human conversations from the MathQA strategy arm (the same 67 participants as Table 8) to their simulated counterparts on(problem, strategy, level). To avoid measurement asymmetry, both human and simulated conversations are tagged with the same LLM-based IU tagger (GPT-4.1-nano) using the same prompt and IU graph per problem. For each turn, the tagger identifies which IUs are mentioned or referenced. We extract two views: (1) user-only, tagging user messages alone to measure the sim- ulated userâs cognitive focus, and (2) full-turn, tagging user + assistant text combined to mea- sure whether the overall conversation covers sim- ilar ground. We then compute Jaccard similarity J =|HâŠS|/|HâŞS|over the union of IUs across all turns, comparing human IU sets (H) against 28 User-OnlyFull-Turn MethodMeanStdMeanStd KNOWSIM.479.260.614.271 ZS.487.279.596.266 ZS-CoT.482.266.555.254 ZS-CoT-Prof.410.235 .536.267 Table 14: Concept coverage Jaccard similarity between human and simulated users (N =202matched pairs). User-only measures overlap in IUs referenced by the user; full-turn includes both user and assistant messages. Bold: highest mean. each simulator (S). In addition to our simulator, we tag three baseline simulators (ZS, ZS-CoT, ZS- CoT-Prof) under the same protocol for comparison. Results. Table 14 reports the aggregate Jaccard scores across all matched pairs. On user-only Jac- card, our simulator (0.479) performs comparably to ZS (0.487) and ZS-CoT (0.482); all three signifi- cantly outperform ZS-CoT-Prof (0.410; Wilcoxon p<.001).On full-turn Jaccard, our simulator achieves the highest overlap with human conversa- tions (0.614), significantly outperforming ZS-CoT (p<.001) and ZS-CoT-Prof (p<.001). The moderate absolute Jaccard values are ex- plained by systematic over-coverage: simulated users engage with 71â78% of IUs in the graph compared to 48â56% for humans. As prerequisites are met, the simulator progressively surfaces newly connectable IUs and works through most of the reachable graph over a session, whereas real users skip reachable concepts they find irrelevant. This over-coverage inflates the denominator|H ⪠S| without proportionally increasing the numerator |H ⊠S|. The higher full-turn Jaccard (0.614 vs. 0.479) indicates that the assistantâs responses par- tially compensate: both human and simulated assis- tants cover similar concepts, even when the usersâ own messages diverge. F Assistant Benchmarking Details F.1 Models and Configuration Table 15 lists the 9 models benchmarked across three tiers. All models use default decoding parameters and non-thinking mode (thinking/reasoning fea- tures disabled) to isolate conversational calibra- tion from chain-of-thought reasoning depth. All models were accessed via hosted APIs. We ac- cessed Gemini models through Google AI Studio, Table 15: Assistant models evaluated in the benchmark. TierModels FrontierGPT-5.4, Claude Opus 4.7, Gemini 3.1 Pro Mid-tierGPT-5.4 mini, Claude Sonnet 4.6, Gemini 3.1 Flash Open-weight DeepSeek V4, Llama-4-Maverick, Qwen-3.6-35B Table 16: Benchmarking 9 LLMs with KNOWSIM, aver- aged across knowledge levels and both tasks. Each cell shows the metric value with rank in parentheses. COâ: lower is better (less cognitive overload). Mean Rank averages the four per-metric ranks; bold marks the top model. ModelKGDCCOâIQMean Rank Claude Opus 4.76.8 (5)0.108 (1)0.856 (3)8.16 (2)2.75 Gemini 3.1 Pro6.4 (6)0.099 (6)0.833 (1)8.24 (1)3.50 DeepSeek V47.2 (1)0.108 (2)0.890 (8)7.26 (5)4.00 Gemini 3.1 Flash7.1 (2)0.106 (3)0.895 (9)7.40 (4)4.50 GPT-5.47.0 (3)0.103 (5)0.878 (5)6.38 (8)5.25 Claude Sonnet 4.66.3 (7)0.092 (9)0.845 (2)7.65 (3)5.25 Qwen-3.6-35B6.9 (4)0.097 (8)0.885 (6)6.65 (6)6.00 GPT-5.4 mini6.2 (8)0.103 (4)0.886 (7)6.59 (7)6.50 Llama-4-Maverick5.8 (9)0.098 (7)0.858 (4)5.43 (9)7.25 GPT models through the OpenAI API, Claude mod- els through the Anthropic API, and open-weight models (DeepSeek V4, Qwen-3.6-35B, Llama-4- Maverick) through OpenRouter. Metrics.We report KNOWSIMâs internal metrics (KG, DC, CO) and the LLM-judged Interaction Quality (IQ) under the same rater Ď r (§4). F.2 Overall Ranking Table 16 reports the level-marginalized benchmark, averaging each metric over the three knowledge lev- els and both tasks. Each level contributes equally: we average conversations within a level, then av- erage the three level means, so that unequal cell counts do not reweight levels. Mean Rank is the mean of the four per-metric ranks (lower is better); the overall ordering is discussed in §6. G Failure Mode Decomposition by Knowledge Level Failure Mode Taxonomy A key advantage of KNOWSIMâs IU-level tracking is that it enables a fine-grained decomposition of why information calibration fails. For every IU the assistant explains in a given turn, we classify it into exactly one of five mutually exclusive categories via an ordered decision tree: 1.REDUNDANTâtheuseralready knows_well;re-explanationaddsno 29 NoviceIntermediateAdvanced 0 25 50 75 100 Explained IUs (%) 20% 68% 89% 37% 11% 6% 28% 10% 10% UNDER ABSORBED OVERLOAD WASTED EFFECTIVE PREREQ BLOCKED REDUNDANT Figure 6: Failure mode distribution by knowledge level, pooled across all nine assistants (214,642 clas- sified IUs). Novice users face over-reaching failures (prereq-blocked+overload); advanced users face over- repetition (redundant). value. 2.PREREQ-BLOCKED â prerequisite ceiling is belowpartial_understandingand the IU did not advance this turn; the user cannot ab- sorb this IU regardless of teaching quality. 3. EFFECTIVE â the userâs state advanced (up- ward transition occurred). 4. OVERLOAD-WASTED â the IU was in the ZPD but the turnâs effective load exceeded the overload threshold, preventing absorption. 5.UNDER-ABSORBED â the IU was in the ZPD, not overloaded, yet did not advance (residual; often correlated with shallow teaching). We apply this taxonomy to the full benchmark- ing corpus. Per-level distribution. Fig. 6 shows the fail- ure mode distribution by knowledge level, pooled across all nine assistants. Calibration failures are directional: novice users are dominated by over- reaching (PREREQ-BLOCKED 37%+OVERLOAD- WASTED 28% = 65% of explained IUs wasted), while advanced users are dominated by over- repetition (REDUNDANT 89%).Only 11% of novice IUs and 3% of advanced IUs result in effec- tive knowledge advancement. At the intermediate level, the profile is more balanced: REDUNDANT dominates (68%) but UNDER-ABSORBED rises to 10%, suggesting that depth-of-teaching becomes the bottleneck when prerequisites are largely met. 30