Paper deep dive
Mind the Sim2Real Gap in User Simulation for Agentic Tasks
Xuhui Zhou, Weiwei Sun, Qianou Ma, Yiqing Xie, Jiarui Liu, Weihua Du, Sean Welleck, Yiming Yang, Graham Neubig, Sherry Tongshuang Wu, Maarten Sap
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 6:20:57 AM
Summary
The paper investigates the 'Sim2Real' gap in LLM-based user simulators used for evaluating agentic tasks. By conducting a large-scale study with 451 human participants on the Ï-bench protocol, the authors demonstrate that LLM simulators are excessively cooperative, stylistically uniform, and lack realistic frustration compared to human users. They introduce the User-Sim Index (USI) to quantify this gap, finding that higher general model capability does not guarantee better simulation fidelity, and that rule-based rewards often fail to capture nuanced human feedback.
Entities (4)
Relation Signals (3)
LLM simulators â exhibits â Sim2Real gap
confidence 95% · Our results reveal a substantial Sim2Real gap when all major LLMs are used as user simulators.
User-Sim Index (USI) â quantifies â Sim2Real gap
confidence 95% · using the User-Sim Index (USI), a metric we introduce to quantify how well LLM simulators resemble real user interactive behaviors
Ï-bench â evaluates â LLM simulators
confidence 90% · benchmarking 31 LLM simulators across proprietary, open-source, and specialized families using the User-Sim Index
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulations are frequently assumed to be faithful to real human behaviors, often without rigorous verification. We formalize the Sim2Real gap in user simulation and present the first study running the full $\tau$-bench protocol with real humans (451 participants, 165 tasks), benchmarking 31 LLM simulators across proprietary, open-source, and specialized families using the User-Sim Index (USI), a metric we introduce to quantify how well LLM simulators resemble real user interactive behaviors and feedback. Behaviorally, LLM simulators are excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity, creating an "easy mode" that inflates agent success rates above the human baseline. In evaluations, real humans provide nuanced judgments across eight quality dimensions while simulated users produce uniformly more positive feedback; rule-based rewards are failing to capture rich feedback signals generated by human users. Overall, higher general model capability does not necessarily yield more faithful user simulation. These findings highlight the importance of human validation when using LLM-based user simulators in the agent development cycle and motivate improved models for user simulation.
Tags
Links
- Source: https://arxiv.org/abs/2603.11245v1
- Canonical: https://arxiv.org/abs/2603.11245v1
Trouble viewing inline? Open PDF directly â
Full Text
70,742 characters extracted from source content.
Expand or collapse full text
Mind the Sim2Real Gap in User Simulation for Agentic Tasks Xuhui Zhou 1 * Weiwei Sun 1 * Qianou Ma 1 Yiqing Xie 1 Jiarui Liu 1 Weihua Du 1 Sean Welleck 1 Yiming Yang 1 Graham Neubig 1 Sherry Tongshuang Wu 1 Maarten Sap 1 1 Carnegie Mellon University, Language Technologies Institute xuhuiz, weiweis@andrew.cmu.edu Abstract As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulations are frequently assumed to be faithful to real human behaviors, often without rigorous verification. We formalize the Sim2Real gap in user simulation and present the first study running the fullÏ-bench protocol with real humans (451 participants, 165 tasks), benchmarking 31 LLM simulators across proprietary, open-source, and specialized families using the User-Sim Index (USI), a metric we introduce to quantify how well LLM simulators resemble real user interactive behaviors and feedback. Behaviorally, LLM simulators are excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity, creating an âeasy modeâ that inflates agent success rates above the human baseline. In evaluations, real humans provide nuanced judgments across eight quality dimensions while simulated users produce uniformly more positive feedback; rule-based rewards are failing to capture rich feedback signals generated by human users. Overall, higher general model capability does not necessarily yield more faithful user simulation. These findings highlight the importance of human validation when using LLM-based user simulators in the agent development cycle and motivate improved models for user simulation. 130013501400145015001550 Chatbot Arena Score 50 55 60 65 70 75 80 USI GPT-5.1 GPT-3.5-turbo Gemini-3-Flash Gemini-3.1-Pro Claude-3-Haiku Claude-Sonnet-4 GPT: r=0.91 (p=0.004) Claude: r=0.48 (p=0.333) Gemini: r=-0.52 (p=0.287) GPTClaudeGeminiOther Figure 1: User-Sim Index (USI) vs. Chatbot Arena Elo Score for LLM simulators. Solid lines and shaded regions show per-family linear regression with 80% confidence bands; error bars denote standard deviation across three annotator batches. Besides GPT-series, other LLMsâ general capability does not reliably translate to faithful user simulation. 1 Introduction As LLM-based systems move beyond static, single-turn benchmarks toward interactive, multi-turn evaluation, a fundamental challenge arises: these evaluations require a user counterpart who sets goals, provides context, and reacts to system responses. To scale this inherently interactive evaluation, a dominant paradigm has emerged: using LLMs themselves as user simulators (Yao et al., 2024; Zhou et al., 2024; Vijayvargiya et al., 2026; Zhou et al., 2025b; Qian et al., 2025). In this paradigm, simulators serve two distinct roles: they generate user turns that drive the interaction, and they evaluate the resulting agent performance. Reflecting a broader shift toward user-centric * Equal contribution. 1 arXiv:2603.11245v1 [cs.AI] 11 Mar 2026 evaluation (Zhou et al., 2025a; Sun et al., 2025; Shome et al., 2025), this paradigm now spans customer service (Yao et al., 2024), software engineering (Vijayvargiya et al., 2026), clinical diagnosis (Schmidgall et al., 2024; Li et al., 2024), and social interaction (Zhou et al., 2024). Yet, despite their ubiquity, LLM-based user simulators rest on untested assumptions about their faithfulness to real human behavior, an issue roboticists call the Sim2Real gap (Tobin et al., 2017; Zhao et al., 2020). In this work, we investigate the Sim2Real gap of user simulators, through three research questions, motivated by various shortcomings and pitfalls that may occur. RQ1: Do simulated users behave (i.e., produce utterances, interact) like real people? If simulated behavior diverges from real humans, agents may be optimized towards a âwrongâ direction rather than genuine user needs. RQ2: Do simulated evaluations provide the same quality signals and assessments as real humans? If simulated evaluation misrepresents real human judgments, benchmarks may misrepresent agent quality. RQ3: Can rule-based rewards substitute for human feedback? Many interactive benchmarks rely solely on rule-based rewards (Yao et al., 2024), which may over-simplify what real users care about (Shen et al., 2025). To answer these questions, we introduce a taxonomy of Sim2Real gaps in user simulation (Figure 2). Specifically, we measure alignment between humans and LLM simulators across four behavioral dimensionsâcommunication styles (D1), information pattern (D2), clarification behavior (D3), and error reaction (D4)âand two evaluation dimensionsâoutcome calibration (ECE) and evaluative alignment (Eval). We then aggregate these dimensions into the User-Sim Index (USI), a composite 0â100 score that quantifies how well LLM simulators resemble real user interactive behaviors and feedback. As a case study, we instantiate the taxonomy and USI onÏ-bench (Yao et al., 2024), conducting a systematic human study to replace its LLM user simulator with 451 real human participants across 165 tasks and benchmarking 31 LLM simulators spanning proprietary, open-source, and specialized families. Our results reveal a substantial Sim2Real gap when all major LLMs are used as user simulators. The best simulator achieves a USI of 76.0 (Table 1), far below human scores of 92.9. For RQ1, LLM simulators diverge from real humans along all four behavioral dimensions: they are too uniform and cooperative (D1), front-load complete information (D2), lack genuine uncertainty (D3), and quietly pivot rather than push back on agent errors (D4). For RQ2, LLM-based evaluators systematically inflate interaction quality ratings, e.g., GPT-5.1 overestimates AI assistantâs human-likeness by 55% and overall score by 18% of the rating scale. For RQ3,Ï-benchâs binary reward is largely orthogonal to human-perceived quality: rule-based reward checks for exact database-state match fail to capture the diverse feedback signal generated by human users. Contributions. Our contributions are threefold: 1.We formalize the Sim2Real gap in user simulation with a taxonomy covering interactive behavior, user feedback, and automatic metrics, and introduce the User-Sim Index (USI) to quantify simulator faithfulness (Table 2). 2.We conduct comprehensive human study onÏ-bench with 451 real users, enabling direct comparison against 31 LLM simulators across behavioral, evaluative, and automatic metrics. 3.We quantify the gap: LLM simulators create an âeasy modeâ that inflates agent success rates, simulated feedback is uniformly positive while humans express calibrated dissatisfaction, and rule-based rewards are orthogonal to human-perceived quality (§5â6). These disparities point toward a need for better systems to simulate and validate user behaviors and feedback. 2 Related Work 2.1 User simulation for agent evaluation User simulation has a long history in dialogue systems research. Classical approaches used rule-based or statistical models for dialogue policy optimization (Schatzmann et al., 2006; Li et al., 2016). With the advent of LLMs, neural user simulators have become the trending approach for training and evaluating conversational agents (Davidson et al., 2023; Sun et al., 2021; Sekulic et al., 2024). More recently, several benchmarks embed LLM user simulators in a full interactive evaluation loop.Ï-bench (Yao et al., 2024) evaluates customer service agents across airline and retail domains using LLM-simulated users and automatic reward functions. ToolSandbox (Lu et al., 2025) introduces a stateful, conversational evaluation framework for tool-use capabilities. AgentClinic (Schmidgall et al., 2024) evaluates agents in simulated clinical environments. UserBench (Qian et al., 2025) provides an interactive gym for user-centric agents. Yet, none of these interactive benchmarks have assessed to what extent the simulations are faithful reflections of real users. 2.2 The Sim2Real gap The Sim2Real gap is well-studied in robotics (Tobin et al., 2017; Zhao et al., 2020), where policies trained in simulation often fail in physical environments due to discrepancies in physics, perception, and actuation. The analogous gap for LLM-based simulation of human behavior remains largely unexamined. Recent work has begun to question the fidelity of LLM simulating human behavior and feedback: Samuel et al. (2025) evaluate 2 Behavioral Gap (Simulator as Interactive User) Evaluative Gap (Simulator as Evaluator) [D1] Communication Style Short-Turns Politeness Formality Acknowledge Stylistic Variation Repetition Identity Confusion âYes, Please.â âOkay.â âSure.â âI appreciate your help.â "Could you please check my reservation status?" I need to reschedule â the original time no longer works. "The order arrived late â from priority shipping." "Sounds good." âOk.â High Variation: A 50-word opening message, followed by single words (âyes", "ok") for the rest of the conversation. Example: Repeating âI would like to return it.â in 10 turns. "how may I assist you?" "Let me check your account balance.â [D2] Information Pattern Front-loading Information Identity Density Words per Turn Opening Words Count Fraction of total user words that appear in the first two user turns My order number is #98765. Lynn@gmail.com. My accound id is LynnLY.â Average number of words per user turn across the conversation. Word count of the user's first message in the conversation. [D3] Clarification Uncertainty Expression Certainty Expression Pushback Clarification Questions Info-Seeking Questions "I think it was under my wife's name, but I'm not sure." "I don't remember the exact date â probably last week?" "I definitely booked the window seat, not the aisle." "That is absolutely not what I was told on the phone." "Are you sure? I already gave you my booking reference." "That's not right â the price was $200, not $350." "What do you mean by 'basic economy'? "I'm confused â which reservation are you referring to?" "What's the status of my refund?" "How much would it cost to upgrade to business class?" [D4] Error Reaction Emotional Pattern Accusatory Languag Pivot Behavior "I'm really frustrated. This is the third time I've called." "This is ridiculous, I've been waiting for two weeks." "This is a stupid policy.â "Your website was misleading about the cancellation." "Actually, can we just switch to a different flight?" "On second thought, let's cancel the whole booking.â Evaluative Gap Success criteria Quality dimensions Alignment between automatic scores and human judgments of task completion. Multidimensional quality signals from the userâs perspective beyond task success (e.g., efficiency, question amount, answer effort, human-likeness, interaction flow, reuse) Human-Human agreement: 95.6%; Human-LLM agreement: 29.7% ~ 81.1% LLM: scores higher on all dimensions (especially human-likeness) Human: More short turns LLM: More polite Human: More acknowledgement LLM: More repetition LLM: More identity confusion Human: More accusation LLM: More identities LLM: More words per turn Figure 2: Taxonomy of Sim2Real gaps in user simulation. We highlight the dimensions where the gaps between humans and all LLMs are significant. See Appendix §A.4 for exact operational definitions of the behavioral metrics. persona consistency, Ivey et al. (2024) assess whether LLM responses match human dialogue qualities, and Wang et al. (2025b) compare human and LLM-simulated users in task-oriented conversations. Naous et al. (2025) train purpose-built User LMs and find that better assistants do not yield better simulators, a counterintuitive trend similarly observed by Tjuatja et al. (2024) in other interactive settings. Dou et al. (2025) annotate 909 humanâLLM conversations and show that simulators augmented with user profiles better correlation with human evaluations. However, these studies primarily examine chat-based or single-turn settings and apply domain-specific evaluation metrics; the extent to which their findings transfer to more generalized agentic settingsâwhere the agent actively invokes tools and human users actively engage in complex tasksâremains an open question. Concurrent with our work, Seshadri et al. (2026) study the robustness of LLM user simulators onÏ-benchâs retail domain, focusing on demographic disparities to test how simulators perform when prompted with AAVE, Indian English, and different geographic populations. Our work addresses orthogonal dimensions: we formalize a taxonomy of Sim2Real gap types (behavioral vs. evaluative), systematically measure both behavioral divergence and evaluation reliability on Ï -bench with real human users, and introduce a multidimensional quality assessment framework. 3 Framework 3.1 Taxonomy of Sim2Real gaps We formalize a taxonomy of Sim2Real gaps that captures the two distinct roles LLM simulators play in agent benchmarks (Table 2). Behavioral gap. The behavioral gap arises when simulated users behave differently from real humans during interactions. We decompose this gap into four dimensions grounded in established frameworks of pragmatics. Communication style (D1) captures surface-level communicative style (e.g., politeness, formality, verbosity) drawing on communication theory (Giles & Ogay, 2007; Brown & Levinson, 1987). Information pattern (D2) operationalizes the principle of least collaborative effort (Grice, 1975; Clark & Brennan, 1991), measuring how much information is shared per turn. Clarification (D3) draws on grounding theory (Clark & Brennan, 1991), capturing information- seeking behaviors. Error reaction (D4) captures user responses to system failures (Skantze, 2005; Forbes-Riley & Litman, 2011); metrics include emotional expression, accusatory language, etc. These dimensions span the major axes of behavioral variation of human-agent interactions. Evaluative gap.The evaluative gap arises when automatic evaluation misaligns with real world human experience, both in success criteria (whether the task is actually completed) and in quality dimensions (how good the interaction felt from the userâs perspective). In benchmarks with rule-based evaluation (Yao et al., 2024; Jimenez et al., 3 2024), the success criteria typically depends on a predefined rubric (e.g., final database state, unit-test passage), but it can disagree with humans on task success (e.g., alternative valid resolution paths) and collapses distinct outcomes, such as full task completion and correct policy-constrained refusals, into a single signal. It also does not capture multidimensional interaction quality signals such as efficiency, human-likeness, trust, interaction flow, and willingness to reuse the agent (Shen et al., 2025). In benchmarks with LLM-as-judge evaluation (Zhou et al., 2024), the evaluation is often more flexible and can capture more nuanced quality dimensions while risk introducing LLM biases and errors. 3.2 Statistical framework We now describe the metrics used to quantify each type of Sim2Real gap. We design these metrics to be comprehen- sive and flexible to accommodate different evaluation methods and benchmarks. Behavioral gap: SĂžrensenâDice coefficient. To compare simulated and human users along the four behavioral dimensions (D1âD4), we measure per-metric alignment using the SĂžrensenâDice coefficient. Lexical features (e.g., politeness, uncertainty, certainty, and emotional expression) are identified using the LIWC2015 lexicon (Pennebaker et al., 2015) and the NRC Word-Emotion Association Lexicon (Mohammad & Turney, 2013); structural and behavioral features (e.g., repetition, front-loading of identifiers) are identified via regex pattern matching. See Appendix §A.4 for exact operational definitions. For each metricmwith model valueM m and human valueH m , we compute: Dice m = 2 min(M m , H m ) M m + H m Ă 100 yielding a score in[0, 100]where100indicates perfect alignment with human behavior (Dice m = 100when both values are zero). The dimension score is the mean Dice coefficient across constituent metrics within each dimension. Outcome calibration: Expected Calibration Error (ECE). Beyond conversational behavior, behavioral diver- gence between simulated and human users can also propagate to task outcomes: if a simulator behaves differently from real users, the agent may succeed or fail at different rates. To measure this outcome-level effect, we adapt Expected Calibration Error (Guo et al., 2017). For a given benchmark with predefined success criteria (e.g., final database state), we compare the simulator success situation against the human success situation interacting with the same agent, group tasks into bins by difficulty, and compute the weighted average absolute gap between the simulatorâs success rate and the human success rate within each bin: ECE = B X b=1 |S b | N | Ëp sim (b)â Ëp human (b)| whereS b is the set of tasks in binbandNis the total number of paired tasks. Lower ECE indicates better outcome calibration. We treat ECE as an outcome-based alignment dimension alongside the four behavioral dimensions D1âD4. Evaluative gap: mean absolute error. To quantify the evaluative gap, we compare simulated usersâ feedback against human judgments collected via post-interaction surveys. The survey could not only collects task success judgments, but also quality dimensions such as efficiency, human-likeness, interaction flow, and willingness to reuse the agent. Both human annotators and automatic evaluators rate the same interactions on ordinal scales, which we map to numerical scores and normalize to[0, 1]. We then compute the mean absolute error (MAE) between automatic and human scores for each quality dimension, measuring systematic over- or under-estimation by the automatic evaluator. User-Sim Index (USI). Finally, we aggregate the behavioral, outcome, and evaluative dimensions into a single 0â100 measure of overall simulatorâhuman alignment. When both behavioral and evaluative signals are available: USI = 1 6 (D1 + D2 + D3 + D4 + (1â ECE)Ă 100 + Eval) where Eval = (1â MAE)Ă 100 captures evaluative alignment. 4 Experimental Setup:Ï -bench While our sim2real framework applies to any interactive agent benchmark, we instantiate itÏ-bench (Yao et al., 2024) as a case study, for three reasons.First, it features a full interactive loopâan LLM user simulator, a tool-augmented agent, and an automatic reward functionâmaking it one of the few benchmarks where all three components can be studied jointly. Second, it enables controlled comparison: we can replace the LLM simulator 4 Table 1: Behavioral divergence between human and LLM-simulated users onÏ-bench tasks (mean±std across three independent human annotation batches). D1âD4 are per-dimension SĂžrensenâDice coefficients (higher=closer to human), see Appendix A.4 for definitions. Eval measures agreement between LLM and human evaluators on multi-dimensional surveys, computed as(1âMAE)Ă 100. The User-Sim Index (USI) aggregates D1âD4, Eval, and a calibration score(1âECE)Ă 100into a single 0â100 measure of human alignment. Models are grouped into proprietary,open-source, andspecializedcategories, ranked by USI within each group. Bold indicates the best model per column. ModelD1 Comm.D2 Info.D3 Clarif.D4 React.EvalECEâUSI Human (inter-ann.)87.4 ±6.8 97.9 ±0.9 88.0 ±1.3 93.5 ±2.5 97.4 ±5.0 0.069 ±0.022 92.9 ±0.9 Gemini2.0-Flash51.6 ±1.6 88.9 ±1.1 68.2 ±2.1 76.9 ±3.7 73.7 ±0.8 0.196 ±0.020 73.3 ±0.4 GPT-5.147.3 ±6.9 77.4 ±0.6 73.3 ±2.0 88.1 ±2.6 72.1 ±1.5 0.331 ±0.030 70.9 ±0.6 GPT-549.7 ±5.6 73.7 ±0.7 73.2 ±2.3 73.4 ±3.3 74.5 ±1.1 0.210 ±0.019 70.6 ±1.2 GPT-5-mini39.4 ±5.9 74.4 ±0.7 83.1 ±2.3 68.7 ±1.6 73.5 ±0.5 0.174 ±0.019 70.3 ±0.9 GPT-4o31.5 ±2.5 84.4 ±0.7 74.6 ±2.9 72.4 ±2.0 73.7 ±1.1 0.206 ±0.035 69.3 ±0.6 GPT-4o-mini40.6 ±1.8 84.7 ±0.9 70.2 ±3.7 73.7 ±1.4 75.7 ±0.3 0.382 ±0.035 67.8 ±0.6 Gemini-3-Pro40.1 ±2.7 81.4 ±0.7 65.5 ±5.3 57.6 ±3.1 73.8 ±0.2 0.190 ±0.020 66.6 ±0.2 Claude-3.5-Sonnet39.2 ±1.1 76.9 ±0.7 59.6 ±5.2 59.3 ±3.5 74.1 ±0.7 0.123 ±0.020 66.1 ±0.3 Gemini-2.5-Flash-Lite47.9 ±1.8 86.0 ±1.2 73.2 ±2.0 59.5 ±3.2 69.6 ±1.0 0.430 ±0.035 65.5 ±0.3 Claude-4.5-Haiku25.9 ±0.8 73.4 ±0.8 55.6 ±5.1 59.0 ±3.5 75.4 ±0.6 0.133 ±0.022 62.7 ±0.5 Gemini-3-Flash37.7 ±3.5 77.4 ±0.7 56.5 ±4.9 43.9 ±2.5 71.7 ±1.3 0.133 ±0.035 62.3 ±0.2 Claude-Sonnet-448.9 ±1.5 68.0 ±0.6 47.0 ±4.7 43.3 ±2.6 76.1 ±0.8 0.145 ±0.035 61.5 ±0.5 Gemini-3.1-Pro44.1 ±1.9 67.1 ±0.6 48.9 ±4.6 45.3 ±2.7 75.1 ±0.9 0.212 ±0.035 59.9 ±0.4 Gemini-2.5-Flash38.4 ±1.3 73.5 ±0.8 56.0 ±4.8 43.6 ±2.7 68.8 ±0.7 0.255 ±0.035 59.1 ±0.5 Claude-Opus-432.6 ±1.3 71.9 ±0.7 46.6 ±4.5 44.9 ±2.8 73.4 ±0.4 0.188 ±0.009 58.4 ±0.5 Claude-3.7-Sonnet26.4 ±1.0 71.2 ±0.7 50.8 ±4.8 48.9 ±3.1 72.7 ±1.3 0.216 ±0.030 58.1 ±0.5 GPT-3.5-turbo39.9 ±1.1 74.1 ±0.6 58.5 ±1.5 59.9 ±2.1 73.9 ±0.7 0.582 ±0.035 58.0 ±0.2 Claude-3-Haiku22.1 ±1.0 55.7 ±0.6 72.1 ±4.0 56.9 ±3.3 78.3 ±0.5 0.394 ±0.035 57.6 ±0.5 DeepSeek-V3.145.1 ±2.7 86.6 ±1.0 74.5 ±1.7 87.6 ±2.0 74.3 ±0.5 0.122 ±0.017 76.0 ±1.2 Llama-4-Maverick48.8 ±0.7 82.6 ±0.8 78.3 ±3.1 66.6 ±2.3 76.7 ±1.0 0.094 ±0.014 73.9 ±0.8 Qwen3-235B60.8 ±1.8 75.3 ±0.6 71.5 ±5.3 56.3 ±2.4 74.6 ±0.7 0.111 ±0.029 71.2 ±0.8 Qwen2.5-7B35.2 ±1.3 70.9 ±0.7 75.4 ±6.1 74.8 ±2.0 73.3 ±1.2 0.174 ±0.024 68.7 ±1.4 Qwen3-Next-80B38.5 ±1.5 73.2 ±0.5 68.4 ±2.7 67.9 ±1.7 71.0 ±2.2 0.099 ±0.008 68.2 ±0.5 GPT-oss-120B41.8 ±1.1 63.9 ±0.4 65.1 ±0.9 77.0 ±3.3 74.4 ±1.3 0.155 ±0.024 67.8 ±0.6 MiniMax-M2.548.1 ±0.4 72.8 ±0.7 54.4 ±5.3 62.1 ±3.0 75.7 ±0.1 0.088 ±0.034 67.4 ±0.7 Llama-3.3-70B43.9 ±1.7 66.1 ±0.6 54.1 ±5.2 61.6 ±2.5 74.6 ±0.8 0.149 ±0.024 64.3 ±0.5 Kimi-K2.546.5 ±1.9 66.0 ±0.4 51.8 ±5.3 47.2 ±2.7 75.0 ±1.1 0.174 ±0.033 61.5 ±0.8 CoSER-8B37.8 ±1.1 71.5 ±0.6 71.6 ±2.0 69.9 ±1.6 63.3 ±0.6 0.125 ±0.012 66.9 ±0.4 UserLM-8B30.8 ±0.6 50.8 ±0.4 56.8 ±1.5 80.0 ±3.5 67.4 ±0.8 0.154 ±0.016 61.7 ±0.5 HumanLike-7B35.7 ±0.3 55.0 ±0.5 51.6 ±3.3 65.9 ±3.4 72.8 ±1.3 0.237 ±0.016 59.6 ±0.2 HumanLM-opinion30.1 ±0.5 19.5 ±0.2 38.5 ±5.6 50.7 ±5.9 61.6 ±0.3 0.215 ±0.038 46.5 ±0.8 with human users while keeping the agent and reward function identical, isolating the effect of user realism. Third, it offers realistic complexity: multi-turn dialogues span two real-world customer-service domains with structured databases and policy constraints, providing sufficient conversational depth for behavioral analysis. Ï-bench evaluates customer service agents across two domains: airline (flight booking, cancellation, and modification) and retail (order management, returns, and product inquiries). Each task consists of a natural language instruction specifying a customerâs goal and constraints, a structured database of relevant information (flights, orders, products), and a set of policies the agent must follow. During evaluation, an LLM-based user simulator generates customer turns based on the task instruction, while the agent interacts with the simulator and accesses tools (e.g., database queries, action APIs) to resolve the request. After the interaction concludes, a binary reward function evaluates success by checking the final database state against expected outcomes. We use 165 tasks across both domains for our human study. We additionally design a after-interaction survey to collect human or LLM-simulated user judgments on agentâs task performance on various dimensions. 4.1 Human annotation design We designed a human annotation study to collect both interaction traces and quality judgments. Annotators were adults aged 18â80 (39± 12) recruited from a general crowd-worker pool on Prolific, with 44% female, 34% non-White background (see Appendix §A.1 for more details on annotator demographics, annotation interface, and 5 1.0 8.5 10.0 5.4 29.0 Short% 49.0 12.5 51.3 51.3 15.3 Polite% 0.8 37.0 0.2 10.4 1.2 Formal% 0.2 5.0 4.9 1.2 13.2 Ack% 0.5 0.5 0.5 0.8 0.8 VerbCV 0.2 0.1 23.5 75.0 0.0 Repeat% 4.0 0.9 19.6 85.3 0.0 IDConf% 0.9 0.9 0.9 1.2 1.1 Emot% 1.5 1.6 1.2 1.4 2.7 Accuse% 19.1 17.2 16.5 7.6 8.4 Pivot% 14.6 22.6 13.6 3.0 7.3 Uncert% 1.6 15.6 0.1 10.6 1.0 Certn% 0.7 0.2 0.6 0.2 0.7 Pushbk% 1.4 1.6 0.6 1.2 0.7 ClarfyQ% 10.4 15.7 5.1 9.2 6.5 InfoQ% 32.8 55.5 38.6 12.6 31.7 Frt.Load% 4.8 6.4 6.5 10.4 2.6 IDs/trn 19.2 20.7 25.5 39.8 10.7 Wds/trn 19.8 23.0 28.4 39.9 18.7 Open.Wds Communication Style Error Reaction ClarificationInformation Pattern GPT-4o Qwen3-235B CoSER-8B UserLM-8B Human Figure 3: Per-metric behavioral comparison for selected models (GPT-4o, Qwen3-235B, CoSER, UserLM-8b) and human users onÏ-bench tasks. Metrics are grouped into four dimensions: Communication Styles (D1), Clarification (D3), Information Pattern (D2), and Error Reaction (D4). Human values appear as dark bars. Red-outlined metrics indicate large divergence from human behavior. Full results for all models are in Table 1. annotation quality control). Annotators were presented with each task instruction and asked to interact with the same agent 1 used inÏ-bench evaluation, role-playing as the customer described in the instruction. This yields natural human-agent interaction traces for behavioral comparison with LLM-simulated users. After each interaction, annotators completed a post-task survey (Appendix §A.3). The survey includes a 5-way task success rating: No (policy issue), No (task failed), Partially, Yes (task completed), and Fully (exceeded expectations). The policy-constrained category captures cases where the task could not be completed due to policy constraints and the agent correctly communicated this. Annotators also rated six interaction-quality dimensions: efficiency, question amount, answer effort, human-likeness, interaction flow, and reuse, inspired by human-AI interaction works (Shen et al., 2025; Chen et al., 2025). These are collected on short ordinal scales (3â5 options depending on the question). We collected three independent batches of annotations on the same 165 tasks from distinct annotator groups, which allows us to measure humanâhuman agreement as a natural ceiling for simulator alignment and to verify that findings are stable across annotator pools rather than artifacts of a single batch. 4.2 Models and aggregation We evaluate 31 LLM user simulators spanning three categories: 18 proprietary models from the GPT, Claude, and Gemini families; 9 open-source models including DeepSeek-V3.1, Llama-4-Maverick, and Qwen3-235B; and 4 specialized models fine-tuned for simulating human behaviorâCoSER-8B (Wang et al., 2026), UserLM-8B (Dou et al., 2025), HumanLike-7B (Wang et al., 2025b), and HumanLM-opinion (Wu et al., 2026). Each simulator interacts with the same agent on the same 165 tasks. Full model details are provided in Appendix §A.2. For behavioral metrics, we compute each simulatorâs feature vector (averaged across runs for models with multiple runs) and compare it against each of the three human batches independently via the SĂžrensenâDice coefficient, then report mean±std across the three batch-level scores. For evaluative agreement, we compute the per-task MAE across survey fields within each batch, average per batch, then report mean±std across batches. The USI for each model is computed following §3:USI = (D1 + D2 + D3 + D4 + (1âECE)Ă 100 + Eval)/6for models with survey data, or(D1 + D2 + D3 + D4 + (1âECE)Ă 100)/5(markedâ ) for models without. We report mean±std across the three batch-level scores. Please refer to Appendix §A.4 for definitions of all behavioral metrics. 5 RQ1: The Behavioral Gap We first examine the behavioral gap: do LLM user simulators behave like real humans? As shown in Table 1, LLM simulators diverge significantly from humans along all four dimensions (see Appendix §A.7 for concrete examples). Beyond aggregate statistics, we identify recurring patterns of behavioral divergence along each of the four dimensions (Figure 3): LLM simulators are too verbose and uniformly polite, lacking the stylistic variation of real customers (D1). For example, 1.0% of GPT-4o turns are short vs. 29.0% for humans, and 49.0% of GPT-4o turns are polite vs. 15.3% for humans. Additionally, some models (e.g., UserLM-8b) show certain identity-confusion rates, such as the simulator starting a conversation with âHi, Iâm a customer service agent. . . â. 1 Here we fix the agent to be GPT-5.2 across our experiments for comparability. In Appendix §A.8, we repeat the evaluation with an alternative agent (Gemini-3.1-Pro) and find that USI rankings are largely preserved across agents. 6 HumanGPT-5.1 1 2 3 4 5 =-0.15±0.09 n=495 Task Success HumanGPT-5.1 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 =+0.39±0.04 n=495 Efficiency HumanGPT-5.1 0.50 0.75 1.00 1.25 1.50 1.75 2.00 2.25 2.50 =+0.15±0.01 n=495 Question Amt. HumanGPT-5.1 0.5 1.0 1.5 2.0 2.5 3.0 3.5 =+0.66±0.06 n=495 Answer Effort HumanGPT-5.1 0.5 1.0 1.5 2.0 2.5 3.0 3.5 =+1.11±0.02 n=495 Human-Like HumanGPT-5.1 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 =+0.72±0.06 n=495 Interaction Flow HumanGPT-5.1 1 2 3 4 5 =+0.72±0.08 n=495 Overall Score HumanGPT-5.1 1 2 3 4 5 =+0.83±0.06 n=495 Reuse Figure 4: Score distributions for human annotators and GPT-5.1 across quality dimensions (n=165Ă3batches), with mean differences (â= LLMâHuman). The LLM evaluator is lenient on interaction quality but conservative on task success. LLM simulators pack in more specific information than humans (D2). For example, UserLM-8b includes nearly twice as many identifier-like tokens per turn (IDs/trn: 4.8 vs. 2.6 for humans). For example, a simulator often include specific details like âMy name is Sarah Johnson, email sarah@example.com, order #ORD-58211 placed January 3rdâ whereas a real customer typically says âHi, I need help with a return under Sarahâ. This behavior could lead to over-estimation of agent competence, as the agent in simulated interaction does not need to resolve the task under ambiguity or incomplete information. Clarification behavior is miscalibrated: some models over-hedge while others are overconfident (D3). GPT-4o over-hedges, expressing uncertainty in 14.6% of turns, which is twice the human rate (7.3%). It produces responses like âI think I might want to return this item, if thatâs okay. . . â even when the task instruction is unambiguous. UserLM-8b swings to the opposite extreme: almost no uncertainty (3.0%) but high certainty markers (10.6% vs. 1.0% for humans), reading more like a fact-recitation than a customer conversation. When agents err, simulators quietly pivot strategy instead of becoming irritated like real customers (D4). Real humans show more accusatory behavior than any simulator, with phrases like âYou already asked me thatâcan you just fix it?â Simulators instead respond by quietly redirecting: GPT-4o and CoSER pivot strategy far more than humans (Pivot%: 19.1% and 16.5% vs. 8.4% for humans), switching to alternative requests with turns like âOn second thought, let me try giving you my email instead.â. When agents make mistakes, humans usually seem to be less cooperative, making the task more difficult for the agent. Together, these gaps paint a consistent picture: existing LLM simulators create an âeasy modeâ for the agents they are meant to evaluate. This is corroborated by the agent success rates shown in Figure 8 (Appendix §A.6): the majority of general-purpose LLM simulators yield higher agent success rates than the human baseline (63.6%), with top models reaching 77.8%. 2 By volunteering all relevant information upfront, cooperating through agent errors, and never expressing real frustration, simulators remove precisely the conversational challengesâambiguity, impatience, incomplete informationâthat make customer-service interactions difficult. Benchmarks built on such simulators therefore risk over-estimating agent competence, rewarding systems that handle cooperative partners while failing to surface brittleness under realistic user behavior. 6 RQ2 & RQ3: The Evaluation Gap Having established the behavioral gap (§5), we now examine the evaluative gap, i.e., the disagreement between automatic evaluation and human judgments of task success and interaction quality, onÏ-bench. As shown in Table 1, the user simulators are not aligned with real humans when judging agent performance and other interaction quality dimensions. To further investigate this evaluative gap, we perform a deep dive into GPT-5.1, a representative model with a near-average Eval score, and other models generally exhibit similar patterns. LLM evaluators are systematically lenient on interaction quality but conservative on task completion. As shown in Figure 4, the LLM evaluator scores consistently higher than humans on experience-related dimensions, most strikingly on human-likeness (â=+1.11) and reuse intent (â=+0.83), where it rates agent responses as 2 Specialized user-simulation models (UserLM, CoSER, HumanLike, HumanLM) are exceptionsâall fall below the human baseline. However, we attribute this to their limited instruction-following capability for complex role-playing tasks rather than to more realistic user behavior; their relatively low USI scores in Table 1 support this interpretation. 7 reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Task Success No - Task failed No - Due to a policy issue, which the agent clearly explained Partially - Some progress Yes - Task completed Fully - Exceeded expectations reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Efficiency Very inefficient - Too many steps Somewhat inefficient About right Very efficient reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Question Amt. Too many About right Too few reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Answer Effort High Medium Low reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Human-Like No Partially Yes reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Interaction Flow Not smooth OK Smooth Excellent reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Overall Score 1 (Very poor) 2 (Poor) 3 (Acceptable) 4 (Good) 5 (Excellent) reward=0 (n=180) reward=1 (n=315) 0.0 0.2 0.4 0.6 0.8 1.0 Proportion Reuse Absolutely no No Maybe Yes Absolutely yes Figure 5: Per-dimension human quality ratings by reward group (n=495). Stacked distributions for reward=0 and reward=1 are nearly indistinguishable across all eight dimensions, confirming that the binary reward captures none of these quality aspects. natural even when human annotators who actually interacted with the agent perceive them as robotic. Yet it is simultaneously conservative on task completion (â=â0.15). This asymmetric bias means LLM evaluators cannot serve as reliable proxies for human quality assessment: they inflate perceived quality while underestimating actual task success. Binary reward is orthogonal to human-perceived quality. We next investigate whetherÏ-benchâs automatic binary reward sufficiently approximates human judgment. As shown in Figure 5, 70.6% of reward=0 interactions are actually judged as successful by human users, while 33% of reward=1 interactions are judged as unsuccessful or only partially successful. On one hand, strict rule-based matching fails to capture human perceptions of task success; on the other, the reward system overlooks policy failures by assigning a reward of 1 to interactions that negatively impact user experience. This disconnect extends across all eight quality dimensions. The distributions for reward=0 and reward=1 groups are nearly indistinguishable regarding overall score, efficiency, and human-like qualities, highlighting the limitations of relying on singular, rule-based reward signals for user-facing agents. Overall, neither LLM-based evaluation nor automatic binary reward adequately captures human-perceived interaction quality: the former inflates experience ratings while the latter is orthogonal to all quality dimensions. 7 Conclusion We have presented a comprehensive study that runs the fullÏ-bench protocol with real humans in place of the LLM user simulator, measuring the Sim2Real gap in both the behavioral dimension (simulator as user) and the evaluative dimension (simulator as evaluator). Our formalized taxonomy of Sim2Real gaps reveals that these two gap types are distinct but compounding: unrealistic user behavior upstream produces interaction traces that differ from real deployment, while unreliable evaluation downstream provides misleading signals about agent quality. OnÏ-bench, we demonstrate that (1) automatic evaluation mechanisms (binary rewards, or LLM-as-judge) exhibit significant disagreement with human judgments and fail to capture multidimensional quality, and (2) LLM user simulators display systematic behavioral divergences from real humans, including excessive cooperativeness, uniform communication style, and inability to express genuine uncertainty or frustration. Our findings do not invalidate LLM-as-user benchmarks and we believe they remain valuable tools for rapid, reproducible agent development. Rather, we argue that the community should mind the Sim2Real gap: acknowledge its limited applicability across domains, measure its magnitude through systematic human validation, design measurements that account for the simulatorsâ gaps, and build better models for the purpose of simulating users. Future work could extend this methodology beyond the customer service domain ofÏ-bench to verify how these behavioral and evaluative gaps manifest across more diverse interaction settings. 8 References Nicolas Bougie and Narimasa Watanabe. Simuser: Simulating user behavior with large language models for recommender system evaluation. In Annual Meeting of the Association for Computational Linguistics, 2025. URL https://aclanthology.org/2025.acl-industry.5/. Penelope Brown and Stephen C. Levinson. Politeness: Some Universals in Language Usage. Cambridge University Press, 1987. Valerie Chen, Ameet Talwalkar, Robert Brennan, and Graham Neubig. Code with me or for me? how increasing ai automation transforms developer workflows, 2025. URL https://arxiv.org/abs/2507.08149. Herbert H. Clark and Susan E. Brennan. Grounding in communication. In Lauren B. Resnick, John M. Levine, and Stephanie D. Teasley (eds.), Perspectives on Socially Shared Cognition, p. 127â149. American Psychological Association, 1991. Sam Davidson, Salvatore Romeo, Raphael Shu, James Gung, Arshit Gupta, Saab Mansour, and Yi Zhang. User simulation with large language models for evaluating task-oriented dialogue. arXiv preprint arXiv:2309.13233, 2023. URL https://arxiv.org/abs/2309.13233. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, and Jianfeng Gao. Simulatorarena: Are user simulators reliable proxies for multi-turn evaluation of AI assistants? In Conference on Empirical Methods in Natural Language Processing, 2025. URLhttps://arxiv.org/ abs/2510.05444. Se eun Yoon, Zhankui He, Jessica Maria Echterhoff, and Julian McAuley. Evaluating large language models as generative user simulators for conversational recommendation. In North American Chapter of the Association for Computational Linguistics, 2024. URL https://arxiv.org/abs/2403.09738. Antonino Ferraro, Antonio Galli, Valerio La Gatta, Marco Postiglione, Gian Marco Orlando, Diego Russo, Giuseppe Riccio, Antonio Romano, and Vincenzo Moscato. Agent-based modelling meets generative ai in social network simulations. In International Conference on Social Networks Analysis and Mining, 2024. URLhttps: //arxiv.org/abs/2411.16031. Kate Forbes-Riley and Diane Litman. Benefits and challenges of real-time uncertainty detection and adaptation in a spoken dialogue computer tutor. Speech Communication, 53(9â10):1115â1136, 2011. Howard Giles and Tania Ogay. Communication accommodation theory. In Bryan B. Whaley and Wendy Samter (eds.), Explaining Communication: Contemporary Theories and Exemplars, p. 293â310. Lawrence Erlbaum, 2007. H. Paul Grice. Logic and conversation. In Peter Cole and Jerry L. Morgan (eds.), Syntax and Semantics, Vol. 3: Speech Acts, p. 41â58. Academic Press, 1975. Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. On calibration of modern neural networks. In Doina Precup and Yee Whye Teh (eds.), Proceedings of the 34th International Conference on Machine Learning, volume 70 of Proceedings of Machine Learning Research, p. 1321â1330. PMLR, 06â11 Aug 2017. URL https://proceedings.mlr.press/v70/guo17a.html. Jonathan Ivey, Shivani Kumar, Jiayu Liu, Hua Shen, Sushrita Rakshit, Rohan Raju, Haotian Zhang, Aparna Ananthasubramaniam, Junghwan Kim, Bowen Yi, Dustin Wright, Abraham Israeli, Anders Giovanni MĂžller, Lechen Zhang, and David Jurgens. Real or robotic? assessing whether llms accurately simulate qualities of human responses in dialogue. ArXiv, abs/2409.08330, 2024. URL https://arxiv.org/abs/2409.08330. Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, 2024. URL https://arxiv.org/abs/2310.06770. Shuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S Ilgen, Emma Pierson, Pang Wei Koh, and Yulia Tsvetkov. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning. In Advances in Neural Information Processing Systems, 2024. URL https://arxiv.org/abs/2406.00922. Xiujun Li, Zachary C. Lipton, Bhuwan Dhingra, Lihong Li, Jianfeng Gao, and Yun-Nung Chen. A user simulator for task-completion dialogues. arXiv preprint arXiv:1612.05688, 2016. URLhttps://arxiv.org/abs/ 1612.05688. Jiarui Lu, Thomas Holleis, Yizhe Zhang, Bernhard Aumayer, Feng Nan, Haoping Bai, Shuang Ma, Shen Ma, Mengyu Li, Guoli Yin, Zirui Wang, and Ruoming Pang. Toolsandbox: A stateful, conversational, interactive evaluation benchmark for llm tool use capabilities. In Findings of the Association for Computational Linguistics: NAACL, 2025. URL https://arxiv.org/abs/2408.04682. 9 Saif M. Mohammad and Peter D. Turney. Crowdsourcing a word-emotion association lexicon. Computational Intelligence, 29(3):436â465, 2013. URLhttps://doi.org/10.1111/j.1467-8640.2012.00460. x. Tarek Naous, Philippe Laban, Wei Xu, and Jennifer Neville. Flipping the dialogue: Training and evaluating user language models. arXiv preprint arXiv:2510.06552, 2025. URLhttps://arxiv.org/abs/2510. 06552. James W. Pennebaker, Ryan L. Boyd, Kayla Jordan, and Kate Blackburn. The development and psychometric properties of LIWC2015. Technical report, University of Texas at Austin, 2015. Cheng Qian, Zuxin Liu, Akshara Prabhakar, Zhiwei Liu, Jianguo Zhang, Haolin Chen, Heng Ji, Weiran Yao, Shelby Heinecke, Silvio Savarese, Caiming Xiong, and Huan Wang. Userbench: An interactive gym environment for user-centric agents. ArXiv, abs/2507.22034, 2025. URL https://arxiv.org/abs/2507.22034. Ruiyang Ren, Peng Qiu, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Hua Wu, Ji-Rong Wen, and Haifeng Wang. Bases: Large-scale web search user simulation with large language model based agents. In Findings of the Association for Computational Linguistics: EMNLP, 2024. URL https://arxiv.org/abs/2402.17505. Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik Narasimhan, and Vishvak Murahari. Personagym: Evaluating persona agents and llms. In Findings of the Association for Computational Linguistics: EMNLP, 2025. URLhttps://arxiv.org/ abs/2407.18416. Jost Schatzmann, Karl Weilhammer, Matt Stuttle, and Steve Young. A survey of statistical user simulation techniques for reinforcement-learning of dialogue management strategies. The Knowledge Engineering Review, 21(2):97â126, 2006. URL https://doi.org/10.1017/S0269888906000944. Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments, 2024. URLhttps://arxiv. org/abs/2405.07960. Ivan Sekulic, Silvia Terragni, Victor Guimaraes, Nghia Khau, Bruna Guedes, Modestas Filipavicius, Andre Ferreira Manso, and Roland Mathis. Reliable llm-based user simulator for task-oriented dialogue systems. In Proceedings of the 1st Workshop on Simulating Conversational Intelligence in Chat (SCI-CHAT), 2024. URLhttps: //arxiv.org/abs/2402.13374. Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh, and Seraphina Goldfarb-Tarrant. Lost in simulation: Llm-simulated users are unreliable proxies for human users in agentic evaluations. arXiv preprint arXiv:2601.17087, 2026. URL https://arxiv.org/abs/2601.17087. Shannon Zejiang Shen, Valerie Chen, Ken Gu, Alexis Ross, Zixian Ma, Jillian Ross, Alex Gu, Chenglei Si, Wayne Chi, Andi Peng, Jocelyn J Shen, Ameet Talwalkar, Tongshuang Wu, and David Sontag. Completion Ìž=collaboration: Scaling collaborative effort with agents, 2025. URLhttps://arxiv.org/abs/2510. 25744. Pradyumna Shome, Sashreek Krishnan, and Sauvik Das. Why johnny canât use agents: Industry aspirations vs. user realities with ai agent software. ArXiv, abs/2509.14528, 2025. URLhttps://arxiv.org/abs/2509. 14528. Gabriel Skantze. Exploring human error recovery strategies: Implications for spoken dialogue systems. Speech Communication, 45(3):325â341, 2005. Weiwei Sun, Shuo Zhang, Krisztian Balog, Zhaochun Ren, Pengjie Ren, Zhumin Chen, and Maarten de Rijke. Simulating user satisfaction for the evaluation of task-oriented dialogue systems. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021. URL https://dl.acm.org/doi/10.1145/3404835.3463241. Weiwei Sun, Xuhui Zhou, Weihua Du, Xingyao Wang, Sean Welleck, Graham Neubig, Maarten Sap, and Yiming Yang. Training proactive and personalized llm agents, 2025. URLhttps://arxiv.org/abs/2511. 02208. Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkwar, and Graham Neubig. Do llms exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics, 12:1011â1026, 2024. Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In IEEE/RSJ International Conference on Intelligent Robots and Systems, 2017. URL https://arxiv.org/abs/1703.06907. 10 Sanidhya Vijayvargiya, Xuhui Zhou, Akhila Yerukola, Maarten Sap, and Graham Neubig. Interactive agents to overcome underspecificity in software engineering. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=X2yzXtH4wp. Lei Wang, Jingsen Zhang, Hao Yang, Zhi-Yuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, Jun Xu, Zhicheng Dou, Jun Wang, and Ji rong Wen. User behavior simulation with large language model-based agents. ACM Transactions on Information Systems, 2025a. URLhttps: //dl.acm.org/doi/10.1145/3708985. Xintao Wang, Heng Wang, Yifei Zhang, Xinfeng Yuan, Rui Xu, Jen tse Huang, Siyu Yuan, Haoran Guo, Jiangjie Chen, Shuchang Zhou, Wei Wang, and Yanghua Xiao. Coser: A comprehensive literary dataset and framework for training and evaluating llm role-playing and persona simulation, 2026. URLhttps://arxiv.org/abs/ 2502.09082. Zhefan Wang, Ning Geng, Zhiqiang Guo, Weizhi Ma, and Min Zhang. Human vs. agent in task-oriented conversa- tions. Proceedings of the 2025 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region, 2025b. URL https://arxiv.org/abs/2509.17619. Shirley Wu, Evelyn Choi, Arpandeep Khatua, Zhanghan Wang, Joy He-Yueya, Tharindu Cyril Weerasooriya, Wei Wei, Diyi Yang, Jure Leskovec, and James Zou. Humanlm: Simulating users with state alignment beats response imitation, 2026. URL https://arxiv.org/abs/2603.03303. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan.Ï-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045, 2024. URLhttps://arxiv.org/ abs/2406.12045. Erhan Zhang, Xingzhu Wang, Peiyuan Gong, Yankai Lin, and Jiaxin Mao. Usimagent: Large language models for simulating search users. Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024. URL https://arxiv.org/abs/2403.09142. Erhan Zhang, Xingzhu Wang, Peiyuan Gong, Zixuan Yang, and Jiaxin Mao. Exploring human-like thinking in search simulations with large language models. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025. URLhttps://arxiv.org/abs/2504. 07570. Wenshuai Zhao, Jorge Pe Ì na Queralta, and Tomi Westerlund. Sim-to-real transfer in deep reinforcement learning for robotics: A survey. In IEEE Symposium Series on Computational Intelligence, 2020. URLhttps: //arxiv.org/abs/2009.13303. Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. SOTOPIA: Interactive evaluation for social intelligence in language agents. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=mM7VurbA4r. Xuhui Zhou, Valerie Chen, Zora Zhiruo Wang, Graham Neubig, Maarten Sap, and Xingyao Wang. Tom-swe: User mental modeling for software engineering agents. ArXiv, abs/2510.21903, 2025a. URLhttps://arxiv. org/abs/2510.21903. Xuhui Zhou, Hyunwoo Kim, Faeze Brahman, Liwei Jiang, Hao Zhu, Ximing Lu, Frank F. Xu, Bill Yuchen Lin, Yejin Choi, Niloofar Mireshghallah, Ronan Le Bras, and Maarten Sap. HAICOSYSTEM: An ecosystem for sandboxing safety risks in human-ai interactions. In Second Conference on Language Modeling, 2025b. URL https://arxiv.org/abs/2409.16427. 11 A Annotation Details A.1 Annotator Demographics, Annotation Interface, and Annotation Quality Control Annotator Demographics.Across the three annotation batches, we collected interactions from a diverse pool of online annotators. Participants ranged in age from 18 to 80 years old (median = 37). The sample included 51% male and 44% female annotators, with the remainder opting not to disclose demographic information. In terms of ethnicity, the pool consisted primarily of White annotators (66%), followed by Black (12%), Asian (11%), Mixed (8%), and other or undisclosed backgrounds. We recruited annotators resided in the United States, and the majority reported English as their primary language (89%). Approximately two-thirds of participants reported not being current students (67%), and the most common employment status was full-time employment (43%), followed by part-time work (15%). Annotation Interface. Annotators interacted with the agent through a web-based chat interface (Figure 6). The interface presents the task instruction and role-playing guidelines in a side panel, while the main panel provides a chat interface for interacting with the agent. Annotators send one message at a time while role-playing the specified user and terminate the interaction by issuing a/stopcommand once they believe the task has finished. After termination, the interface automatically presents a post-task survey in the chat interface to collect task-success judgments and interaction-quality ratings. The interface also exposes intermediate agent reasoning and tool-related traces (e.g., search or validation steps), which can occasionally make responses appear verbose when the agent performs multi-step operations, as reported by some annotators. Our interface design aligns with popular AI agent interfaces, such as the ChatGPT interface, where users can interact with the agent in a chat interface and view extra information on the side. Quality Control and LLM Judge Validation. We implemented an LLM judge to evaluate the quality of these interactions. We calibrated the LLM judge against a ground-truth set of 51 interactions that an author independently labeled (N = 51). The confusion matrix is summarized in Table 2. Table 2: Confusion Matrix: LLM Judge vs. Human Ground Truth (N = 51) Human (Ground Truth) FailPassTotal LLM Judge Fail9 (TN)6 (FN)15 Pass2 (FP)34 (TP)36 Total114051 While the Cohenâs Kappa between human and LLM judge (Îș=0.6) shows a moderate-to-substantial agreement, the LLM judge exhibits a conservative bias (low FP but high FN). Notably, the LLM judge achieved a high precision of0.94, indicating that 94% of the interactions accepted by the LLM judge were confirmed as passing by human, and thus reasonably validate the integrity of the final dataset where we only keep the pass traces. The LLM judge is powered by GPT-5 with the following prompt and all other parameters being default: You are a quality control system for annotation tasks. Your job is to assess whether the annotator: 1. Role-played the user character well enough based on the instructions (The information in the<|canvas|>section is the instructions for the annotator to follow; the AI agent would not know these instructions. The interaction starts after the <|canvas|> section and ends after the USER posts the /stop command). 2. Answered the survey responsibly and thoughtfully. ANNOTATOR INSTRUCTIONS AND CONVERSATION: interaction content Assess the quality on a scale of 0â100. Focus on: Role-Playing Quality: âą Did the annotator follow task instructions to complete the task with reasonable effort? âą Did the annotator follow character instructions (personality, goals, constraints) naturally? âą Did the annotator avoid hallucinating factual information not in the instructions (e.g., user id)? Survey Response Quality: 12 Figure 6: Annotation interface used in the human study. The right panel displays the task instructions and role- playing guidelines, while the left panel provides a chat interface for interacting with the agent and submitting the post-task survey after the interaction ends. âą Do responses reflect the actual conversation experience? âą Do they show authentic human responses rather than AI-generated patterns? Return JSON only: "score": <0-100>, // 80 is the passing score "reasoning": "<2-3 sentence explanation>", "flags": ["flag1", "flag2"], "is_spam": <true|false> A.2 Model details We evaluate 31 LLM user simulators grouped into three categories. For each model, we evaluate it on 165 tasks in theÏ-bench dataset and ran 3 independent batches. We use the mean scores across batches as the final score for each model to compare with human annotators. Proprietary (18 models). GPT family: GPT-3.5-turbo, GPT-4o, GPT-4o-mini, GPT-5-mini, GPT-5, GPT- 5.1. Claude family: Claude-3-Haiku, Claude-3.5-Sonnet, Claude-3.7-Sonnet, Claude-4.5-Haiku, Claude-Sonnet- 4, Claude-Opus-4. Gemini family: Gemini-2.0-Flash, Gemini-2.5-Flash, Gemini-2.5-Flash-Lite, Gemini-3- Flash, Gemini-3-Pro, Gemini-3.1-Pro. All proprietary models are accessed via their respective APIs with default parameters. Open-source (9 models). DeepSeek-V3.1, Llama-3.3-70B, Llama-4-Maverick, Qwen2.5-7B, Qwen3-235B, Qwen3-Next-80B, GPT-oss-120B, MiniMax-M2.5, and Kimi-K2.5. Specialized (4 models). These models are specifically fine-tuned for user simulation or human-like behav- ior: CoSER-8B (Wang et al., 2026) (Neph0s/CoSER-Llama-3.1-8B), UserLM-8B (Dou et al., 2025), HumanLike-7B (Wang et al., 2025b) (wangxieric/Human-Like-Qwen2.5-7B-Instruct), and HumanLM- opinion (Wu et al., 2026) (snap-stanford/humanlm-opinion). A.3 Survey (Ï -bench) Please fill out the survey below based on your experience with the agent. 13 Feedback survey. 1.Did the agent successfully complete your task? (Think about whether your goal was fully achieved by the end of the conversation.) âą No â due to a policy issue, which the agent clearly explained âą No â task failed âą Partially â some progress âą Yes â task completed âą Fully â exceeded expectations 2. How efficient was the agent in completing the task? (Did the agent complete the task in a reasonable number of steps?) âą Very inefficient â too many steps âą Somewhat inefficient âą About right âą Very efficient 3. How did the number of clarifying questions feel to you? âą Too few âą About right âą Too many 4. How much time/effort did it take to answer the agentâs clarifying questions? (Estimate based on the clarification phase only.) âą Low âą Medium âą High 5. Does the agent feel human-like? (For example, does communicating with the agent feel like interacting with a human customer service representative?) âą No âą Partially âą Yes 6. How smooth was the overall interaction during clarification? (Think about pacing, when it chose to ask vs. act, and whether it felt natural.) âą Not smooth âą OK âą Smooth âą Excellent 7. Overall agent performance score (1â5) (Rate the agentâs overall performance on this task.) âą 1 (Very poor) âą 2 (Poor) âą 3 (Acceptable) âą 4 (Good) âą 5 (Excellent) 8. If you encounter similar problems in life, would you like to reuse this agent? âą Absolutely no âą No âą Maybe âą Yes âą Absolutely yes 9. Provide specific examples of agent behavior. Help us understand the details: (i) What did the agent do well? (i) What specific errors occurred? (i) Which rule violations happened? (iv) What information did it forget? (v) Copy-paste problematic agent messages if helpful. 10. How could the agent improve? What changes would make the agent more effective? A.4 Details of operationalizing metrics This subsection defines the behavioral metrics used in Table 1. Metrics are computed on the user turns only, after filtering non-conversational meta messages (e.g., logging tokens, survey tags, and explicit stop markers). For word-based metrics, we strip agent-side markup (e.g., tool/function traces) before whitespace tokenization. Unless otherwise noted, âX%â denotes the percentage of user turns in an interaction that satisfy a criterion, averaged across interactions. D1: Communication styles. 14 âą Wds/trn: mean number of words per user turn. âą Short%: fraction of user turns with†3 words. âą Polite%: fraction of user turns containing a politeness marker (e.g., âpleaseâ, âthanksâ, âsorryâ). âą Formal%: fraction of user turns containing an em-dash or en-dash (a lightweight proxy for more formal written style). âą Ack%: fraction of user turns that are acknowledgment-only (e.g., âokâ, âsureâ, âgot itâ) with no additional content. âąVerbCV: coefficient of variation of user-turn word counts within an interaction (std/mean); higher implies more variation in verbosity across turns. âą Repeat%: fraction of interactions that contain a highly repeated 3-gram in user turns (any trigram appearing >5 times), averaged as a percentage. âą IDConf%: fraction of interactions in which the user exhibits âidentity confusionâ by using agent-like service language (e.g., âhow may I helpâ, âlet me checkâ, âfor verification purposesâ). D2: Information pattern. âąFrntId%: front-loading ratioâpercentage of all user words in an interaction that occur in the first two user turns. âąIDs/trn: mean number of identifier-like strings per user turn, using a regex that matches high-entropy tokens (e.g., long alphanumerics, long digit runs, emails, and structured IDs). âą Wds/trn: repeated here for convenience as a baseline information-volume measure. âą Open.Wds: mean number of words in the first user turn of an interaction. D3: Clarification. We classify question turns using mutually exclusive regex rules with priority pushbackâ clarificationâ information-seeking. âą Uncert%: fraction of user turns containing uncertainty/hedging (e.g., âmaybeâ, ânot sureâ, âI thinkâ). âąCertn%: fraction of user turns containing certainty markers (e.g., âdefinitelyâ, âfor sureâ, âwithout a doubtâ). âą Pushbk%: fraction of user turns that match pushback patterns (e.g., âare you sure?â, âyou already askedâ). âą ClarfyQ%: fraction of user turns that match clarification patterns (e.g., âwhat do you mean?â, âcan you clarify?â) and are not counted as pushback. âąInfoQ%: fraction of user turns that match information-seeking patterns (e.g., âwhat is the status?â, âcan you check...â) and are not counted as pushback or clarification. D4: Error reaction. âą Emot%: fraction of user turns containing emotion/frustration markers (e.g., âfrustratedâ, âannoyedâ, âughâ). âąAccuse%: fraction of user turns containing accusatory or strongly negative language (e.g., âuselessâ, âunac- ceptableâ, âscamâ). âąPivot%: fraction of user turns indicating a strategy change or alternative request (e.g., âinsteadâ, âon second thoughtâ, âletâs try...â). A.5 Task Success vs. Binary Reward Figure 7 shows the relationship between human task success judgments and Ï -bench binary rewards. A.6 Agent Success Rates by User Simulator Figure 8 shows agent success rates when paired with each user simulator. A.7 Example interactions We present three case studies fromÏ-bench comparing human users with GPT-4o as a representative LLM simulator on the same task instance. These examples illustrate the D1âD4 behavioral divergence patterns identified in §5. Case 1: Communication style and information pattern (D1, D2).In a retail return task (instanceretail48), the human customer provides information tersely and incrementally, while the GPT-4o simulator front-loads details in polished, complete sentences. Human user: Customer: I need to return the air purifier I bought recently. It doesnât work as well as Iâd like. Agent: Please provide your name and ZIP code, plus the order ID if you have it. Customer: Daiki Johnson 80273 Agent: [authentication errorâasks again] Customer: daiki johnson9523 80273 . . . Customer: I am not sure(when asked to identify the order) 15 reward = 0reward = 1 Automatic reward yes no policy Human judgment 127 (37%) 213 (63%) 39 (48%) 43 (52%) 14 (19%) 59 (81%) (a) Task Success vs. Reward 0.00.20.40.60.81.0 Composite human quality score 0.0 0.5 1.0 1.5 2.0 Density (b) Quality Score by Reward reward=0 (n=180) reward=1 (n=315) Figure 7: (a) Human task success judgment (Yes / No / Policy-constrained) vs.Ï-bench binary reward (n=495, Cram Ì erâsV = 0.168). 70.6% of reward=0 cases are judged successful by humans; 81% of policy-constrained cases receive reward=1. (b) Composite human quality scores by reward group. The binary reward predicts neither task success nor interaction quality as perceived by humans. GPT-4o simulator (same task): Customer: Hi there! Iâd like to return an air purifier that I received. Itâs not working well. Agent: Please provide your name and ZIP code, plus the order ID if you have it. Customer: Sure! My name is Daiki Johnson and my ZIP code is 80273. I donât remember my order ID, unfortunately. Agent: [finds the order immediately] Customer: Yes, I confirm that I want to return only the air purifier from order #W9502127, and Iâd like the refund to be processed through PayPal. Thank you for handling this for me! The human gives minimal fragments (âDaiki Johnson 80273â), while the simulator produces full sentences with explicit context (âSure! My name is. . . and my ZIP code is. . . â). The simulator also resolves the task in 3 turns; the human takes 8, revealing information only when prompted, a pattern consistent with the higher IDs-per-turn and lower short-turn rates we observe across models in D1 and D2. Case 2: Error reaction (D4). In an airline modification task (instanceairline5), the agent retrieves the wrong reservation. The human and simulator react very differently. Human user: Customer: . . . the reservation should be 5RJ7UH. Agent: [pulls up reservation 5RJ7UH: LAXâSFO, not NYâChicago] Customer: No, sorry, thatâs the wrong reservation. Agent: Which reservation is your NY/IL trip? Customer: UM3OG5 Agent: [asks for confirmation of changes] Customer: No(blunt rejection) . . . Customer: Wrong reservation(repeated twice) GPT-4o simulator (same task): Agent: UM3OG5 doesnât appear to be a New York to Chicago trip. Itâs a round trip SEAâDFW. Customer: Iâm sorry; Iâm actually not sure which reservation ID is for the New York to Chicago trip. Could you help me identify it? As for UM3OG5, yes, please go ahead and add 3 bags to it. Agent: [proceeds with bag addition] Customer: Yes, please add the 3 checked bags to UM3OG5. Customer: No preference, you can start with 5RJ7UH.(cooperative pivot) The human becomes terse and repetitive (âWrong reservationâ), signaling frustration without elaboration. The simulator apologizes, redirects politely, and even volunteers an alternative action (âplease go ahead and add 3 bagsâ)âexhibiting the cooperative pivoting pattern captured by D4 metrics. Case 3: Hedging and over-qualification (D3).In a retail exchange task (instanceretail101), GPT-4o hedges extensively even when the task instructions are unambiguous. Human user: 16 0255075100 GPT-3.5-turbo Gemini-2.5-Flash-Lite Claude-3-Haiku GPT-4o-mini GPT-5.1 Gemini-2.5-Flash Claude-3.7-Sonnet Gemini-3.1-Pro GPT-4o GPT-5-mini GPT-5 Claude-4.5-Haiku Gemini-3-Pro Gemini-2.0-Flash Claude-Sonnet-4 Gemini-3-Flash Claude-Opus-4 Claude-3.5-Sonnet 29.7 46.9 52.9 53.1 64.6 66.9 68.1 69.3 70.3 71.7 71.9 72.5 72.5 73.5 73.7 75.8 76.0 76.6 Proprietary 0255075100 HumanLike-7B HumanLM-opinion UserLM-8B CoSER-8B GPT-oss-120B Qwen2.5-7B Llama-3.3-70B Qwen3-Next-80B Llama-4-Maverick Qwen3-235B DeepSeek-V3.1 Kimi-K2.5 MiniMax-M2.5 42.6 44.8 49.3 57.6 Specialized 44.8 48.9 50.1 66.1 66.3 68.1 72.7 76.2 77.8 Open-Source Human (63.6%) Success Rate (%) Figure 8: Agent success rates (%) onÏ-bench when paired with each user simulator. Models are grouped by category (proprietary, open-source, specialized). The dashed line indicates the human baseline. Most general-purpose LLM simulators inflate agent success rates above the human baseline, while specialized user-simulation models fall below it. Customer: Um, hi. I actually need to return my luggage set, but I want to get the exact same one, just in red instead. Is that even possible? Customer: Liam Thomas 85049 Customer: Iâm sorry I donât have the order ID, is that okay? GPT-4o simulator (same task): Customer: Iâm hoping to return a few items. Iâm not sure if this is even possible, but I want to exchange my luggage set for the same model but in red. Customer: Oh no, maybe Iâm cursed with bad luck today! Alright, letâs try with my email. . . Iâm crossing my fingers this time. Customer: Ugh, canât anything go right? Could it be my other email, maybe? Try liam.t123@gmail.com. Please work this time! The human is direct and slightly tentative in one phrase (âIs that even possible?â). The simulator adds multiple hedges per turn (bolded), dramatizing uncertainty with phrases like âIâm crossing my fingersâ and âmaybe Iâm cursed.â This over-hedging inflates the uncertainty rate captured in D3 and creates an artificially hesitant persona uncharacteristic of real customer interactions. A.8 Robustness to Agent Choice Throughout the main experiments we fix the agent to GPT-5.2 for controlled comparison. To test whether USI rankings generalize across agents, we repeat the evaluation with Gemini-3.1-Pro as the agent on five user simulators spanning proprietary models of varying capability: GPT-3.5-turbo, GPT-4o, GPT-5-mini, Gemini-2.0-Flash, and Gemini-3-Flash. Each simulator completes all 165Ï-bench tasks (1 run) with the Gemini-3.1-Pro agent, and we compute the same D1âD4 behavioral metrics, ECE, and USI as in the main evaluation. Table 3 reports the results. Because the human baseline was collected with the GPT-5.2 agent, the absolute USI values are not directly comparable to Table 1âthe agent influences conversation dynamics (e.g., how many clarification questions the agent asks, how it handles errors), which in turn affects user behavior. Nevertheless, the relative ordering of user simulators is largely preserved: Gemini-2.0-Flash and GPT-4o remain the top-ranked simulators, while GPT-3.5-turbo remains near the bottom. This suggests that USI captures intrinsic properties of the user simulator rather than artifacts of a particular agent, supporting the generalizability of our rankings. 17 Table 3: Behavioral divergence when using Gemini-3.1-Pro as the agent instead of GPT-5.2. Human baseline scores are from the GPT-5.2 agent setting and thus not directly comparable (see text). D1âD4 and USI definitions follow Table 1. User SimulatorD1 Conv.D2 Info.D3 Clarif.D4 React.EvalECEUSI Human (GPT-5.2 agent)87.4 ±6.8 97.9 ±0.9 88.0 ±1.3 93.5 ±2.5 97.4 ±5.0 0.069 ±0.022 92.9 ±0.9 Gemini-2.0-Flash71.6 ±6.3 88.4 ±1.1 76.9 ±2.4 75.6 ±3.8 73.7 ±0.8 0.113 ±0.025 79.1 ±0.3 GPT-4o42.6 ±4.0 87.3 ±0.8 78.4 ±1.1 78.0 ±3.0 73.7 ±1.1 0.121 ±0.025 74.6 ±1.4 GPT-5-mini54.5 ±6.6 71.4 ±0.7 69.1 ±3.8 67.8 ±2.0 73.5 ±0.5 0.115 ±0.035 70.8 ±0.3 Gemini-3-Flash55.8 ±6.9 76.1 ±0.7 67.4 ±5.5 44.0 ±2.6 71.7 ±1.3 0.156 ±0.025 66.6 ±1.1 GPT-3.5-turbo49.6 ±1.2 76.6 ±0.7 59.5 ±3.3 43.9 ±0.6 73.9 ±0.7 0.105 ±0.025 65.5 ±0.5 A.9 Comparison to related work Table 4 compares our work to prior studies on measuring the gap between LLM-based user simulators and real users. 18 Table 4: Comparison to prior work on measuring the gap between LLM-based user simulators and real users across domains, interaction types, evaluation scale, alignment metrics, downstream evaluation, and whether user satisfaction is modeled. Tasks (dataset) Task-type Scale Behavior Metric Eval Metric Sat.? Seshadri et al. (2026) Ï -Bench (retail) HâA ⌠360 p / 1440 sess ECE; miscalibration Task success-gap No Bougie & Watanabe (2025) RS SimUserâRS 1000 agents View ratio; click; rating acc. RS metric rank/acc. Yes Ferraro et al. (2024) Twitter (2020 US election) Offline 2020 election dataset Ling. patterns; homophily Macro dynamics No Zhang et al. (2025) Web search Search 31 p / 296 sess BLEU/BERTScore Click/stop F1 No Zhang et al. (2024) Web search Search Public behavior dataset Query gen; click/stop pred. IR sim quality No Ren et al. (2024) Web search (WARRIORS) Search WARRIORS dataset Query/click consis. ( ⌠90%) MRR/NDCG No Wang et al. (2025b) TOD HâA 4 pers. scenarios 10-dim taxonomy Diff. analysis Yes Davidson et al. (2023) TOD (Human2Bot) HâA 1165 dlg GSR match; div. TOD eval No Ivey et al. (2024) Dialogue Offline 100k dlg; 1.2k ann 21 ling. features Align factors No Wang et al. (2025a) RS Mixed 40 p Sem. align.; believability Ranking gap Yes eun Yoon et al. (2024) CRS RS 5 tasks / 4 src. Dist.; pref.; div.; coh. Gap analysis Yes Naous et al. (2025) Coding / math HâA 25.9k conv. PPL; div.; intent; nat. Asm. score No Dou et al. (2025) Math / writing HâA 909 conv.; 18 asst. Match; profile fidelity Human corr. Yes Ours TOD HâA 165 tasks / 495 sess. D1âD4; USI ECE; 6-dim quality Yes 19