Paper deep dive
SAGE: A Service Agent Graph-guided Evaluation Benchmark
Ling Shi, Yuqin Dai, Ziyin Wang, Ning Gao, Wei Zhang, Chaozheng Wang, Yujie Wang, Wei He, Jinpeng Wang, Deiyi Xiong
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 1:52:03 AM
Summary
SAGE (Service Agent Graph-guided Evaluation) is a multi-agent benchmark designed to evaluate Large Language Models (LLMs) in customer service scenarios. It formalizes unstructured Standard Operating Procedures (SOPs) into Dynamic Dialogue Graphs to enable automated, dual-axis assessment of logical compliance and conversational quality. The framework utilizes a multi-agent setup with Judge Agents and a Rule Engine to generate deterministic ground truth, uncovering critical model limitations such as the 'Execution Gap' and 'Empathy Resilience' across 27 LLMs.
Entities (6)
Relation Signals (3)
SAGE → evaluates → LLMs
confidence 100% · SAGE... a universal multi-agent benchmark for automated, dual-axis assessment [of LLMs].
Judge Agent → analyzes → Service Agent
confidence 95% · Judge Agents and a Rule Engine analyze interactions between User and Service Agents.
Rule Engine → generates → Ground Truth
confidence 95% · The Rule Engine functions as a deterministic generator for procedural logic ground truth.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Existing benchmarks predominantly rely on static paradigms and single-dimensional metrics, failing to account for diverse user behaviors or the strict adherence to structured Standard Operating Procedures (SOPs) required in real-world deployments. To bridge this gap, we propose SAGE (Service Agent Graph-guided Evaluation), a universal multi-agent benchmark for automated, dual-axis assessment. SAGE formalizes unstructured SOPs into Dynamic Dialogue Graphs, enabling precise verification of logical compliance and comprehensive path coverage. We introduce an Adversarial Intent Taxonomy and a modular Extension Mechanism, enabling low-cost deployment across domains and facilitating automated dialogue data synthesis. Evaluation is conducted via a framework where Judge Agents and a Rule Engine analyze interactions between User and Service Agents to generate deterministic ground truth. Extensive experiments on 27 LLMs across 6 industrial scenarios reveal a significant ``Execution Gap'' where models accurately classify intents but fail to derive correct subsequent actions. We also observe ``Empathy Resilience'', a phenomenon where models maintain polite conversational facades despite underlying logical failures under high adversarial intensity. Code and resources are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.09285v1
- Canonical: https://arxiv.org/abs/2604.09285v1
Trouble viewing inline? Open PDF directly →
Full Text
88,620 characters extracted from source content.
Expand or collapse full text
SAGE: A Service Agent Graph-guided Evaluation Benchmark Ling Shi ∗ Tianjin University Tianjin, China Yuqin Dai ∗ Tsinghua University Beijing, China Ziyin Wang Tianjin University Tianjin, China Ning Gao Beihang University Beijing, China Wei Zhang Beijing University of Posts and Telecommunications Beijing, China Chaozheng Wang The Chinese University of Hong Kong Hong Kong, China Yujie Wang Independent Researcher Beijing, China Wei He Independent Researcher Beijing, China Jinpeng Wang Independent Researcher Beijing, China Deyi Xiong † Tianjin University Tianjin, China Abstract The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their perfor- mance remains challenging. Existing benchmarks predominantly rely on static paradigms and single-dimensional metrics, failing to account for diverse user behaviors or the strict adherence to structured Standard Operating Procedures (SOPs) required in real- world deployments. To bridge this gap, we propose SAGE (Service Agent Graph-guided Evaluation), a universal multi-agent bench- mark for automated, dual-axis assessment. SAGE formalizes un- structured SOPs into Dynamic Dialogue Graphs, enabling precise verification of logical compliance and comprehensive path cover- age. We introduce an Adversarial Intent Taxonomy and a modu- lar Extension Mechanism, enabling low-cost deployment across domains and facilitating automated dialogue data synthesis. Eval- uation is conducted via a framework where Judge Agents and a Rule Engine analyze interactions between User and Service Agents to generate deterministic ground truth. Extensive experiments on 27 LLMs across 6 industrial scenarios reveal a significant “Exe- cution Gap” where models accurately classify intents but fail to derive correct subsequent actions. We also observe “Empathy Re- silience”, a phenomenon where models maintain polite conversa- tional facades despite underlying logical failures under high ad- versarial intensity. Code and resources are available at https: //anonymous.4open.science/r/SAGE-Bench-4CD3/. ∗ Both authors contributed equally to this research. † Corresponding Author Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. KDD ’26, Jeju Island, Republic of Korea © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X CCS Concepts • Computing methodologies→Multi-agent planning; Dis- course, dialogue and pragmatics. Keywords Service Agent, Graph-guided Evaluation, Multi-agent Interaction ACM Reference Format: Ling Shi, Yuqin Dai, Ziyin Wang, Ning Gao, Wei Zhang, Chaozheng Wang, Yujie Wang, Wei He, Jinpeng Wang, and Deyi Xiong. 2026. SAGE: A Service Agent Graph-guided Evaluation Benchmark. In Proceedings of The 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD ’26). ACM, New York, NY, USA, 18 pages. https://doi.org/X.X 1 Introduction The rapid advancement of Large Language Models (LLMs) has significantly accelerated the automation process across various industrial sectors [2,42,46]. Among these, the domain of intelli- gent customer service has emerged as a pioneer in adopting these technologies to enhance operational efficiency and user experi- ence [7,28,43]. By leveraging the robust generative capabilities and context understanding of LLMs, enterprises are striving to tran- sition from traditional, rigid rule-based systems to highly adaptive intelligent service agents [44,47] capable of handling complex in- teractions [36,53], thereby addressing a broader range of customer needs while reducing labor costs. Consequently, benchmarking models on their strict adherence to Standard Operating Procedures (SOPs) for correct workflow transi- tions, a process requiring the generation of structured responses that align with predefined logic as shown in Figure 1, has become a priority for enterprises. However, existing service benchmarks face three fundamental limitations. First, insufficient evaluation dimensions: Current benchmarks typically assess either task com- pletion [20,25,31] or dialogue quality [41,59] in isolation. Real- world scenarios, however, demand both strict logical compliance with SOPs and appropriate communication skills. This metric insuffi- ciency causes evaluation bias and hinders precise error localization. arXiv:2604.09285v1 [cs.AI] 10 Apr 2026 KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. Classification Fields:...... Action: TransferHuman Path: [stage1, stage2, stage3,stage4’] Chat: ×...... ServiceAgent StructuredResponse Input Context ExtractionStandard Operating Procedures Replying to User 3 Stage1: Classification ... Stage2: View ConsumptionType Cancel Enquiry Stage3: View Penalty Stage5: View PackageStatus Stage4: View EmotionTag ChangeOrder Stage6: View ConsumptionProfile =0 !=0 Discontent Contracted Stage7: View ApplicationTendency Data/Voice Calm Goodbye Reject/Hesitate Change Chat TrasnferHuman SystemInformation Penalty PackageStatus ...... 1 2 4 OrderId Hello, I placed the wrong ...... Sorrytohearthat let me check... DialogueHistory ...... Iwanttochangemypackage...... Your order has not been shipped...... Figure 1: Service Agent SOP Example (Telecom Scenario). Second, static interaction paradigms: Relying on fixed datasets like scripts [5,34,58], traditional methods fail to test error recovery or cover diverse user behaviors, ranging from cooperative inquiries to adversarial conflicts. Consequently, these static evaluations are incomplete and struggle to reflect real-world performance. Finally, limited scalability: Benchmarks often depend on costly manual annotation [22,33] and are frequently over-fitted to the SOPs of a single domain, making it prohibitively expensive to adapt to diverse, multi-branch business scenarios. To address these challenges, we propose SAGE (Service Agent Graph-guided Evaluation). SAGE integrates three core contribu- tions: (1) Dynamic Multi-Turn Dialogue Graph Modeling, which formalizes SOPs into directed graphs (as shown in Figure 1) to en- able dynamic verification of logical compliance and ensure com- prehensive path coverage; (2) a Multi-Agent multi-dimensional Evaluation utilizes Judge Agents and a Rule Engine to analyze the interaction between User and Service Agents, generating determin- istic ground truth for a rigorous, multi-dimensional assessment of logical compliance and chat quality; and (3) a Scenario Extension Mechanism, which enabled the rapid deployment of six industrial scenarios in our study, demonstrating its practical feasibility. To validate SAGE, we evaluated 27 mainstream LLMs across 6 industrial scenarios. Our experiments reveal: (1) The gap be- tween open-source and closed-source models is narrowing, with DeepSeek-V3.2 surpassing several GPT-4 class models; (2) A signif- icant “Execution Gap” in complex scenarios, where high classifica- tion accuracy does not guarantee correct action execution, high- lighting the challenge of procedural reasoning; (3) Performance degradation in multi-turn dialogues due to context fatigue; and (4) “Empathy Resilience” under high adversarial intensity, where models maintain polite conversational facades despite underlying logical failures. In summary, our contributions are: • We propose SAGE, the first graph-guided multi-agent eval- uation benchmark that transforms unstructured SOPs into directed graphs to enable automated, dual-axis assessment of logical compliance and conversational quality. •We introduce dynamic graph modeling and an Adversarial Intent Taxonomy, bridging the gap between static testing and dynamic reality through diverse user behavior simula- tions. •We uncover critical phenomena such as the “Execution Gap” and “Empathy Resilience”, providing granular diagnostics for agentic capabilities, through extensive experiments on 27 LLMs across 6 distinct scenarios. • We design a modular Extension Mechanism for low-cost adaptation to new scenarios, which also facilitates the auto- mated synthesis of large-scale dialogue datasets for customer service. 2 Related Works 2.1 Evaluation for Large Language Models Large language model evaluation has evolved from basic instruction- following [60] to complex multi-constraint scenarios. Early bench- marks assessed format compliance [38,48], while recent work evalu- ates real-world complexity through Multi-IF [15], FollowBench [16], and InfoBench [32]. Guidebench [8] introduces domain-specific conditions, and Collie [52] systematically constructs constrained generation tasks. Domain-specific evaluation spans business pro- cess management [4,12–14,17,18,35], finance [49,50], and pro- cedural compliance. SOPBench [22] and SOP-Bench [27] evaluate tool-calling sequences but primarily focus on external manipula- tion rather than deep logical reasoning. Broader agent frameworks include AgentBench [25], WebArena [61], and WorkArena [9]. Com- plementing mathematical tasks, textual logical reasoning, and read- ing comprehension are assessed through LogiQA [24] for deductive reasoning, DROP [10] for discrete reasoning over paragraphs, and HotpotQA [51] for multi-hop reasoning, which are more closely aligned with the context understanding required in service sce- narios. Agent training advances include AGILE [30], ReAct [53], Reflexion [36], Agent-Pro [57], AgentTuning [54], Self-Refine [26], Self-Instruct [45], and human feedback training [29]. Unlike prior benchmarks limited by static datasets and single-dimensional met- rics, SAGE introduces a graph-guided multi-agent framework to dynamically evaluate both the logical compliance and conversa- tional quality of service agents. 2.2 Benchmarks for Service Dialogue Systems Service dialogue evaluation addresses customer-facing applications through DialogBench [28] for human-like conversation, ECom- Bench [43] for customer support resolution. Multi-turn complexity is captured by MG-ShopDial [3], Wizard of Shopping [21], and Par- rot [37]. Customer support dialogue is studied through evaluation frameworks [62], real-world conversation data [58], and recommen- dation as instruction following [56]. Existing benchmarks often rely rigidly on scenario-specific Standard Operating Procedures (SOPs), limiting their adaptability. SAGE addresses this with a modular, intent-based extension mechanism that enables rapid, code-free adaptation to new domains. 3 Methodology To address the complexity of evaluating customer service LLMs, we propose SAGE, a graph-guided multi-agent framework. As shown in Figure 2, the workflow integrates three key stages: (1) Dynamic Multi-Turn Dialogue Graph Modeling formalizes SOPs into di- rected graphs to enable dynamic verification of logical compliance against correct paths while ensuring comprehensive coverage of SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Graph-Guided Multi-Agent Evaluation Scenario Config SOP Design AdversarialUserIntents GraphModeling User Agent Service Agent Dual-axis Results ZeroWeakStrong RuleEngine PromptTemplate Intent Distribution Classification_Acc Action_AccPath_Acc LogicScore Judge Agent1 Judge Agent2 Judge Agent3 "CoreGoal": "Exchange", " ChatQualityScore": S1 "CoreGoal": "Return", " ChatQualityScore": S2 "CoreGoal": "Exchange", "ChatQualityScore": S3 Voting & Scoring ChatScore Metrics Right Label Dynamic Evaluator Multi - Turn Dialogue Simulator Classification Action Path ... ❌ TP ARER OEPS LD Rule Engine ... ... ... ...... Avg(S1,S2,S3) Intent User Config Persona Figure 2: Overview of SAGE evaluation framework. all potential scenarios; (2) Graph-Guided Multi-Agent Evalu- ation rigorously assesses these trajectories; and (3) an Scenario Extension Mechanism that leverages both user intents and SOPs to enable rapid adaptation to arbitrary scenarios through modu- lar configuration. We begin by detailing the graph formalization process. 3.1 Dynamic Multi-Turn Dialogue Graph Modeling Aiming at a systematic evaluation of LLMs applied in service agent, we transform SOPs into computable graphs to underpin our frame- work, thereby allowing for automated logical verification and com- prehensive scenario traversal. 3.1.1 Procedure Graph Formalization. Customer service Standard Operating Procedures (SOPs) typically exist as natural language documents containing numerous conditional branches as shown in Figure 1. Using such unstructured descriptions directly for auto- mated evaluation is prone to ambiguity. Therefore, we formalize SOPs as directed graphs퐺=(푉,퐸). The node set푉consists of three types: (1) Start/End Nodes; (2) Decision Nodes, which branch the flow based on specific conditions; and (3) Action Nodes, repre- senting concrete operations. The edges퐸define the transition logic, where transitions are triggered by the values of specific Classifica- tion FieldsF(e.g., Order Type) and System InformationS. This graph structure serves as the backbone of SAGE. It functions both as the navigation map for the Service Agent and the evaluation standard for the Rule Engine. Crucially, System InformationS is consistently propagated throughout the process: it is used to shape the User Agent’s persona, utilized by the Service Agent for decision-making, and finally employed by the Rule Engine to gen- erate process ground truth. This consistency ensures the fairness and determinism of the evaluation. 3.1.2 User-Agent Multi-turn Interaction. To overcome the limita- tions of traditional benchmarks, we generate dialogue trajecto- ries through dynamic interaction between user agents and service agents. User Agent generates user responses based on personas and dialogue context: 푎 푡 = UserAgent(ℎ <푡 ,푠 푡 ,S,I,P).(1) The user agent’s role is to provide dynamic adversarial testing for the service agent. The generated message푎 푡 comprehensively considers five factors: (1) dialogue historyℎ <푡 , which references previous exchanges; specifically, when history is empty, a dedi- cated initialization function is triggered to generate the opening utterance to populateℎ <푡 ; (2) agent state푠 푡 , utilized to align the user’s response with the agent’s current status, including handling opening protocols and detecting dialogue termination conditions; (3) system informationS, which grounds the user’s knowledge in reality (e.g., knowing their own payment status) to prevent hal- lucinations and ensure logical consistency; and crucially, (4) user KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. intentIand (5) user personaP. As further detailed in Section 3.3, Idetermines the user’s fundamental goal (e.g., refund vs. inquiry), whilePcharacterizes the user persona (encompassing traits such as emotional state, communication style, and cooperativeness). Both are dynamically instantiated to simulate diverse adversarial inten- sities. Service Agent is the evaluation target tasked with navigating the SOP graph to generate a structured response. The output at turn 푡 , denoted as 푏 푡 , comprises four distinct components: 푏 푡 = ServiceAgent(S,ℎ <푡 ,푎 푡 ,퐺)=푝 푡 , action 푡 ,F 푡 , chat 푡 .(2) Here,퐺represents the SOP graph described by text. The compo- nents of푏 푡 are defined as follows: (1)푝 푡 denotes the path, represent- ing the nodes the agent intends to traverse in the current turn; (2) action 푡 is the executed action, selected from the allowable action setAdefined by the current node; (3)F 푡 represents the classifica- tion fields, capturing the agent’s categorical judgment of specific classification fields (e.g., user emotion type); and (4)chat 푡 is the natural language chat response generated to interact with the user. The agent’s core objective is to accurately identify the classification fieldF 푡 , which dictates the subsequent transition path푝 푡 and the mandated actionaction 푡 within graph퐺, ultimately conditioning the generation of the response chat 푡 . Consequently, the User Agent acts as an environment, presenting a specific service scenario to the Service Agent. The Service Agent must then navigate multi-turn interactions to achieve specific goals while strictly adhering to the SOP. We evaluate these generated dialogue trajectories in the subsequent sections. 3.2 Graph-Guided Multi-Agent Evaluation Upon the completion of dialogue trajectory generation, SAGE exe- cutes a systematic evaluation through a graph-guided, multi-agent collaborative mechanism. To ensure that all potential business sce- narios and logical branches are rigorously tested, we enforce Path Coverage as the first component, verifying that the generated tra- jectories span the entire state space of the formalized SOP graph. Following this structural validation, SAGE performs a granular as- sessment of individual interactions. In this architecture, each node and its corresponding edge are evaluated through a hybrid mech- anism: a Judge Agent extracts categorical labels and linguistic nuances from the dialogue, while a Rule Engine cross-validates these outputs against the graph-defined logic. This dual-layered approach allows SAGE to simultaneously assess Logical Compli- ance and Chat Quality, ensuring that agents are both procedurally correct and contextually appropriate. 3.2.1 Path Coverage. To guarantee full logic coverage, we prede- fine all valid trajectoriesPbased on the SOP graph. This involves a two-stage process consisting of initial intent-balanced sampling followed by targeted supplementation for under-covered paths. For rare branches, we use an inverse configuration mechanism to deter- ministically synthesize the required user intents and system states. This ensures every logical path, including edge cases, is tested fre- quently, establishing a robust reference standard for the subsequent evaluation. With testing comprehensiveness secured, we next detail the methodology for evaluating each trajectory. 3.2.2 Judge Agent. The Judge Agent is designed to handle the semantic understanding of natural language interactions. For each turn푡, it analyzes the system informationS, dialogue historyℎ <푡 , user message푎 푡 , and the service agent’s natural language response chat 푡 (excluding internal reasoning steps). The judge outputs the classification ground truthF ∗ 푡 (e.g., user’s emotion type and user’s goal) and a conversational quality score 푠 chat : (F ∗ 푡 ,푠 chat )= JudgeAgent(S,ℎ <푡 ,푎 푡 , chat 푡 ).(3) To ensure robustness against individual model bias, we employ an ensemble of three Judge Agents. We apply majority voting to determine the consensus ground truthF ∗ 푡 of classification fields, while the quality score푠 chat is derived from the average rating. The classification ground truthF ∗ 푡 serves as the input for the Rule Engine, while 푠 chat directly quantifies the conversational quality. 3.2.3 Rule Engine. The Rule Engine functions as a deterministic generator for procedural logic ground truth. It receives the consen- sus classification ground truthF ∗ 푡 (derived from the Judge Agent), system informationS, and the SOP graph퐺as inputs. By perform- ing a deterministic search on the graph, it calculates the unique, theoretically correct execution path and action: (푝 ∗ 푡 , action ∗ 푡 )= RuleEngine(F ∗ 푡 ,S,퐺).(4) Here,푝 ∗ 푡 represents the reference path andaction ∗ 푡 denotes the refer- ence action. These outputs serve as the rigid standard for evaluating the service agent’s logical reasoning capabilities, specifically its path planning and action selection accuracy. 3.2.4 Dual-Axis Evaluation Metrics. With the ground truth estab- lished, we conduct a dual-axis evaluation comparing the service agent’s output against these standards. This composite design facil- itates the precise localization of model defects to specific granular dimensions. Logical Compliance Evaluation. We assess logical adherence across three dimensions, with weights푤 1 =0.4,푤 2 =0.4,푤 3 =0.2: (1) Classification Accuracy measures the alignment between the agent’s predicted fieldF 푡 and the ground truth label determined by the Judge Agent. Specifically, it quantifies the agent’s proficiency in correctly identifying the state-specific attributes that trigger graph transitions. A high Logic score indicates the agent correctly identifies intent and follows the SOP graph. Acc cls =|F| −1 ∑︁ 푓∈F I[F 푡 (푓)=F ∗ 푡 (푓)].(5) (2) Path Correctness measures the overlap between the agent’s planned path and the Rule Engine’s reference path: Sim path = |푝 푡 ∩ 푝 ∗ 푡 | |푝 ∗ 푡 | .(6) (3) Action Correctness verifies if the agent’s final executed action matches the Rule Engine’s reference action: Acc action = I[action 푡 = action ∗ 푡 ].(7) Chat Quality Evaluation. This metric assesses the linguistic and interactive performance of the service agent. To ensure a multi- dimensional evaluation, the Judge Agents evaluate each response across five key dimensions: Linguistic Quality, Anthropomorphism, SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea Content Utility, User Satisfaction, and Instruction Compliance. The final score for this metric is derived through a two-step aggregation: first, for each individual judge, a weighted sum is calculated based on these five dimensions to reflect their relative importance; second, the scores from the three independent Judge Agents are averaged to mitigate subjective bias and ensure the reliability of the evaluation. Score quality = 1 3 3 ∑︁ 푗=1 5 ∑︁ 푘=1 푠 푗,푘 ! .(8) Overall Assessment Score. The final overall score for a turn inte- grates both axes: Score overall = 0.8× Score logic + 0.2× Score quality .(9) For multi-turn dialogues, we adopt a turn-level evaluation strat- egy (assessing turns 1, 5, 10, 15, and the final turn) to measure performance across different conversation depths. This weighted integration prioritizes procedural rigor while accounting for user experience, facilitating the precise diagnosis of model defects across specific dimensions. Beyond robust evaluation, SAGE is engineered for scalability. We next detail how its modular design supports rapid adaptation to arbitrary scenarios. 3.3 Scenario Extension Mechanism Intent-based Adversarial Scenario Taxonomy. To guarantee a comprehensive simulation of realistic user behaviors, we estab- lish a taxonomy based on two critical dimensions: goal alignment (the extent to which user demands match agent capabilities) and emotional state (the intensity of user aggression or urgency): Zero- adversarial Intents reflect cooperative interactions where user goals align with agent services (e.g., standard payment inquiries). Characterized by clear requests and friendly attitudes, these sce- narios primarily test the agent’s basic procedural execution. Weak- adversarial Intents introduce procedural friction or rational criti- cism. Here, users present complex constraints (e.g., unpaid bills) or ambiguous needs, requiring agents to resolve contextual conflicts without facing direct hostility. Strong-adversarial Intents rep- resent high-stakes conflicts driven by emotional dissatisfaction or emergencies (e.g., safety hazards). Users employ aggressive strate- gies to demand immediate resolutions, rigorously testing the agent’s negotiation, de-escalation, and risk control capabilities. Extension Mechanism. We implement this taxonomy via prompt engineering, integrating user intentIand personaPinto the User Agent to generate diverse, high-fidelity scenarios (details in Ap- pendix B). Consequently, SAGE facilitates rapid scenario extension through a streamlined two-step process: formalizing the SOP graph and defining the user persona profile. This modular, “fill-in-the- blank” approach decouples scenario configuration from the core evaluation engine, significantly lowering the technical barrier for deployment. Beyond evaluation, this high-scalability architecture can be further extended to automated dialogue data synthesis, en- abling the large-scale generation of high-quality training corpora for customer service LLMs. 4 Experiments 4.1 LLM Configuration To ensure a comprehensive assessment of the current landscape, we select 27 representative Large Language Models (LLMs), catego- rized into closed-source and open-source families, covering a wide spectrum of parameter scales and architectures. Closed-Source Models. We evaluate state-of-the-art propri- etary systems accessed via official APIs. This includes the Claude series (Sonnet-4.5, Opus-4.5), known for strong reasoning capabili- ties; the GPT series (GPT-4.1) [1], serving as a standard baseline; and the Gemini series (2.5-Pro, 3-Pro/Flash) [39], representing multimodal-native architectures. We also include leading Chinese proprietary models such as Qwen-Max [2] and the Doubao series, which are widely optimized for Chinese application scenarios. Open-Source Models. We cover models ranging from light- weight (3B) to massive scale (1T) to analyze the impact of model size. Our selection features the Qwen2.5 and Qwen3 families [2], which provide a granular range of sizes (3B to 235B); the DeepSeek series (V3, V3.2, R1) [23], representing advanced Mixture-of-Experts (MoE) architectures; and the Llama-3 series [11]. Additionally, we evaluate high-performing models from other providers, including GLM-4.7 [55], Kimi-K2.5 [40], and MiniMax-M2.1 [6]. All open-source models are deployed locally using vLLM [19] to ensure consistent inference efficiency, while closed-source models are evaluated using their respective stable API endpoints. 4.2 Scenario Configuration To evaluate the generalization capability of service agents across diverse industrial domains, we constructed six distinct customer ser- vice scenarios (detailed in Appendix B.1.2). These scenarios range from standardized inquiries to complex, high-stakes disputes, cov- ering varying levels of SOP complexity. • Ecommerce Refund (ER): The most complex scenario, featur- ing a deep decision tree based on product status and credit levels. Agents must navigate multi-branch logic and negotiate terms with varied user temperaments. •Logistics Delivery (LD): Centered on supply chain exceptions (e.g., lost or delayed parcels), requiring proficiency in status track- ing and insurance claim processing. • Telecom Package (TP): A standardized scenario focusing on linear SOP execution for billing and plan upgrades, evaluating instruction-following and upselling protocols. •Property Service (PS): Emphasizes community coordination and emotional management (e.g., repair schedules or noise com- plaints) within offline service contexts. •Airline Refund (AR): A high-complexity scenario governed by rigid, time-sensitive policies. It tests the agent’s precision in cal- culating dynamic cancellation fees and de-escalating passenger anxiety. •Online Education (OE): Focuses on long-term contract disputes and rigorous risk control. Agents must identify potential mali- cious refunders and strictly adhere to intricate refund formulas. These scenarios collectively cover the spectrum from simple procedural execution to complex adversarial negotiation, ensuring a robust assessment of agentic capabilities. KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. Table 1: Main results across 6 scenarios (0-100 Scale). OA: Overall Assessment Score; Logic: Logical Compliance Score; Chat: Conversational Quality Score. The superscripts indicate the ranking within Closed-Source and Open-Source groups, respectively. Format Error: Percentage of outputs failing JSON parsing. Chat Length: Average character count of the response field. ModelAVG ScoreOA on 6 ScenariosFormatChat NameParamsOA Logic ChatERTPPSLDAROEErrorLength Closed-Source Large Language Models Claude-Sonnet-4.5-71.56 2 70.10 3 77.38 1 71.2175.0375.0767.3066.5274.240.75 %93.42 Claude-Opus-4.5-72.62 1 72.10 1 74.68 2 72.7373.1273.2571.2570.1275.220.94 %43.49 GPT-4.1- 66.79 6 68.31 4 60.70 8 70.3266.9065.4662.6963.6871.690.04 %18.59 Gemini-2.5-Pro-66.90 5 65.69 8 71.73 4 69.4666.4269.6064.9260.7170.280.00 %53.33 Gemini-3-Pro-Preview-71.40 3 71.27 2 71.89 3 70.9969.0973.6568.0564.4582.150.00 %48.60 Gemini-3-Flash-Preview- 68.28 4 68.16 5 68.75 5 67.5665.3671.6267.1062.4075.628.92 %50.16 Qwen-Max1T+66.54 7 67.57 6 62.44 7 66.8165.4764.3664.0363.8574.750.00 %27.55 Doubao-Seed-1.6-Flash230B 51.75 9 51.08 9 54.43 9 54.1851.5947.3446.9847.3763.040.49 %18.83 Doubao-Seed-1.8-66.14 8 65.78 7 67.62 6 70.2565.8969.0364.0162.1965.490.04 %26.79 Open-Source Large Language Models Qwen2.5-3B-Instruct3B56.16 17 55.16 17 60.15 12 68.9255.1654.5346.0159.6552.700.02 %26.76 Qwen2.5-7B-Instruct7B61.22 13 60.80 13 62.86 6 69.0562.5860.1947.2460.8867.350.00 %19.44 Qwen2.5-14B-Instruct14B60.90 14 60.67 14 61.80 8 67.6463.3358.1254.2462.9659.110.00 %27.23 Qwen2.5-32B-Instruct32B65.34 7 66.49 7 60.74 11 72.0768.2762.2662.9960.2566.200.00 %21.18 Qwen2.5-72B-Instruct72B65.27 8 65.51 11 64.30 4 64.9069.1665.7863.1659.8368.760.00 %44.67 Qwen3-4B4B57.86 15 58.56 15 55.05 16 64.7162.7256.5744.3656.4762.310.04 %24.12 Qwen3-8B8B 56.71 16 57.65 16 52.93 17 51.8862.8054.0347.7757.6766.080.00 %19.93 Qwen3-14B14B63.88 12 65.57 10 57.11 14 63.7968.0164.8156.1363.7866.740.43 %25.54 Qwen3-32B32B64.69 10 65.97 9 59.58 13 67.9465.4866.7158.7961.6667.561.48 %24.15 Qwen3-235B-A22B235B64.72 9 64.87 12 64.13 5 66.7364.9761.4764.7762.3768.021.34 %32.85 Deepseek-V3.2671B71.29 1 72.08 1 68.11 2 71.5170.6170.7170.8765.8478.170.00 %27.48 Deepseek-V3671B66.54 6 67.90 5 61.11 10 69.6271.6067.5760.0861.5868.770.00 %18.50 Deepseek-R1671B68.39 3 70.00 2 61.94 7 70.1670.3970.5865.7564.0369.400.97 %38.08 GLM-4.7355B67.25 5 68.68 4 61.52 9 68.1466.0269.6466.5060.8272.361.46 %17.58 Kimi-K2.51T68.71 2 69.16 3 66.93 3 69.6969.4171.8465.7464.4571.150.52 %34.17 MiniMax-M2.1229B68.33 4 66.59 6 75.29 1 68.3370.1264.0768.5466.3772.550.82 %61.27 Llama-3.1-8B-Instruct8B 37.62 18 37.36 18 38.64 18 49.5855.8337.5828.2638.1416.3044.33 %44.70 Llama-3.3-70B-Instruct70B64.02 11 66.06 8 55.85 15 68.6064.7463.8760.2762.9363.710.03 %13.79 4.3 Model Performance Across Scenarios In addition to the primary metrics (Logical Compliance Score, Chat Quality Score and Overall Assessment Score), we report two aux- iliary indicators to provide a more nuanced analysis of model be- havior: Format Error Rate and Average Chat Length. The former quantifies the frequency of JSON parsing failures to reflect the model’s instruction-following stability, while the latter measures response verbosity to identify potential issues with redundant gen- eration or lack of conciseness. We evaluated 27 mainstream Large Language Models (LLMs), covering both closed-source (e.g., GPT-4, Claude-3.5) and open-source (e.g., Qwen2.5, Llama-3) families. Ta- ble 1 presents the comprehensive performance across six diverse customer service scenarios. Superiority of Closed-Source Models and the Rising Open- Source Challengers. Closed-source models continue to define the performance frontier, with Claude-Opus-4.5 securing the highest Overall Assessment (OA) score (72.62 1 ), underpinned by its top- tier logical reasoning (72.10 1 ) and impressive chat quality score (74.68 2 ). However, the performance gap between proprietary and open-weight models is remarkably narrow. DeepSeek-V3.2 (671B) emerges as a formidable competitor, achieving an OA of 71.29 1 among open-source models—surpassing established closed-source giants like GPT-4.1 and Gemini-3-Pro-Preview. This indicates that state-of-the-art open-source architectures have reached a level of maturity capable of handling complex, graph-guided service logic previously reserved for proprietary systems. Decoupling Logical Compliance and Conversational Qual- ity. Our results reveal a nuanced trade-off between procedural rigor and linguistic flair. While Claude-Opus-4.5 leads in logic, its sibling Claude-Sonnet-4.5 dominates the Chat Quality category (77.38 1 ), suggesting a more empathetic persona. A notable outlier is MiniMax-M2.1, which, despite a moderate Logic score (66.59 6 ), SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea ARERLDOEPSTP 0.0 0.2 0.4 0.6 0.8 1.0 Score Classification Acc Path Acc Action Acc Figure 3: Logic performance gap analysis across six scenarios. 58 62 Score OA Score 56 60 Logic Score 63 64 Chat Score 35 37 AR Chat Length 67 68 Score 67 68 65 66 35 37 ER 60 61 Score 59 60 62 63 40 44 LD 68 72 Score 69 75 65 67 35 36 OE 57 63 Score 57 63 63 65 30 34 PS 151015 Turn 62 66 Score 151015 Turn 64 68 151015 Turn 58 60 151015 Turn 32.9 33.3 TP Figure 4: Performance evolution across dialogue turns in six scenarios. ZeroWeakStrong 0.0 0.2 0.4 0.6 0.8 1.0 Score Overall Score ZeroWeakStrong 0.0 0.2 0.4 0.6 0.8 1.0 Score Logic Ability ZeroWeakStrong Adversarial Intensity 0.0 0.2 0.4 0.6 0.8 1.0 Score Chat Quality ZeroWeakStrong Adversarial Intensity 0.00 0.02 0.04 0.06 0.08 0.10 0.12 Score JSON Error Rate Figure 5: Performance distribution across 3-level Adversarial intensities. achieves a stellar Chat score (75.29 1 ), the highest among open- source models. This suggests that certain models are specifically optimized for human-like interaction, which can occasionally com- pensate for minor procedural deviations in terms of overall user perception. In contrast, DeepSeek-R1 and GPT-4.1 lean heavily to- ward logic-first strategies, often at the expense of conversational warmth. The Dynamics of Response Verbosity and Stability. We observe that response length is a significant indicator of service quality, but only up to a threshold. Top-tier performers like Claude- Sonnet-4.5 (93.42 chars) and MiniMax-M2.1 (61.27 chars) tend to provide more elaborate guidance and empathetic de-escalation, which are crucial in high-complexity scenarios like ER and AR. Conversely, the failure of smaller models is often catastrophic rather than gradual; for instance, Llama-3.1-8B suffers from an ex- treme Format Error rate (44.33%), failing even the basic instruction- following required to output a valid JSON. Interestingly, Gemini-3- Flash-Preview maintains a competitive OA despite a higher-than- average error rate (8.92%), indicating that when it does follow the format, its reasoning is remarkably efficient. 4.4 Cross-Scenario Difficulty Analysis SAGE provides a standardized framework to quantify scenario com- plexity. As illustrated in Figure 3, we analyze three subdimension of logic performance across six domains. We define scenario difficulty by the “Execution Gap”—the disparity between the nested bars, which reveal two distinct failure modes in LLM reasoning. First, in high-complexity scenarios such as ER, AR, LD, and OE, we observe a significant gap between Classification_Acc (Light Gray) and Ac- tion_Acc (Dark Blue). This disparity highlights a Logic Deduction Barrier: even when models correctly identify the classification fields (F 푡 ), they frequently fail to derive the correct subsequent action. This suggests that for complex SOPs, correct semantic classification does not naturally guarantee successful procedural execution due to intricate conditional dependencies. Second, in more standardized scenarios like PS and TP, an inter- esting phenomenon occurs where Path_Acc (Light Blue) exceeds Classification_Acc (Light Gray). This stems from the metric’s defi- nition: Path Accuracy is calculated as the intersection length of the predicted and ground-truth paths divided by the total ground-truth length. Because this metric accounts for the successful traversal of early, simpler steps, models can achieve relatively high path scores by following partial correct fields, even if they stumble on specific complex classifications. This confirms that Path_Acc serves as a more lenient dimension, whereas the gap between classification and action acts as a more rigorous stress test for deep procedural reasoning. 4.5 Multi-turn Robustness Analysis To evaluate model stability over extended interactions, we analyze performance variations across different dialogue depths (Turn 1, 5, 10, 15). Figure 4 illustrates the evolution of OA, Logic, and Chat scores across six scenarios. A distinct performance trend “Inverted- U” Trajectory is observed across most scenarios (e.g., AR, ER, PS, TP): scores typically peak at Turn 5 and decline significantly by Turn 15. This phenomenon can be attributed to two phases. (1) Information Gain Phase (Turn 1→5): Performance generally improves from the first turn to the fifth (e.g., ER OA rises from ∼67 to∼68). In the initial turn (Cold Start), agents lack sufficient context. By Turn 5, through multi-turn interaction, agents gather critical user information (e.g., order IDs, specific complaints), en- abling more accurate intent classification and SOP navigation. (2) Context Fatigue Phase (Turn 10→15): As the dialogue extends beyond Turn 10, performance exhibits a marked decline (e.g., PS OA drops from∼63 to∼57). This degradation highlights the limita- tions of current LLMs in handling long-context dependencies. The accumulation of historical information introduces noise, leading to the “Lost in the Middle” phenomenon where models hallucinate or lose track of the current state within the SOP graph. 4.6 Impact of Adversarial Intensity To validate the effectiveness of our Adversarial Scenario Taxon- omy, we analyze model performance across three intensity levels: KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. 50556065707580 Chat Quality Score (0-100) 4 5 6 7 8 9 Sub-dimension Score (0-10) Sub-dimensions(CI=90%) Linguistic Quality (R²=0.917) Anthropomorphism (R²=0.955) Content Utility (R²=0.970) User Satisfaction (R²=0.963) Instruction Compliance (R²=0.905) Figure 6: Correlation analysis between the Chat Quality Score and its five sub-dimensions. The deviations from the regres- sion lines (e.g., Doubao-Seed-1.8’s high Linguistic vs. low Instruction scores) highlight specific model characteristics. Zero-, Weak-, and Strong-Adversarial. Figure 5 presents the distri- bution of four key metrics—Overall Score (OA), Logic Score, Chat Quality, and Format Error Rate—via box plots. Logic Degradation and Error Escalation. As adversarial in- tensity increases, we observe a consistent decline in Logical Com- pliance (Logic Score). This trend confirms that adversarial user behaviors (e.g., concealing information, emotional aggression) suc- cessfully introduce logical friction, making it harder for agents to adhere to SOPs. Concurrently, the Format Error Rate rises signif- icantly in Strong-Adversarial scenarios. This suggests that under high cognitive load or emotional pressure, models are more prone to instruction drift, failing to maintain the structured JSON output format required by the system. The Stability in Chat Quality. The distribution of Chat Quality scores remains stable across varying adversarial intensities. This “Empathy Resilience” reveals a decoupling between dialogue and logic: models maintain polite, de-escalating facades even when increased user aggression impairs their logical compliance. Such consistency underscores the need for dual-axis evaluation to distin- guish surface fluency from procedural robustness. Increased Discriminative Power. Crucially, the score distri- bution becomes significantly more dispersed (larger interquartile range and more outliers) as intensity increases. In Zero-Adversarial settings, most models perform comparably well. However, Strong- Adversarial scenarios widen the gap between top-tier and lower-tier models. This proves that high-intensity scenarios serve as a more effective filter for distinguishing robust agentic capabilities. 4.7 Correlation Analysis of Sub-Metrics To validate the internal consistency of our evaluation framework and diagnose fine-grained model capabilities, we analyze the corre- lation between the aggregated Chat Quality Score (X-axis, scaled to 0-1) and its five constituent Sub-dimension Scores (Y-axis, scaled 0-10). As illustrated in Figure 6, the regression analysis yields several key insights. High Metric Consistency. We observe an extremely strong linear correlation across all five dimensions—Linguistic Quality, Anthropomorphism, Content Utility, User Satisfaction, and Instruc- tion Compliance—with coefficient of determination (푅 2 ) values consistently exceeding 0.9. This statistical coherence confirms the validity of our metric design. In complex multi-agent evaluation, a high푅 2 indicates that each dimension provides a consistent and reliable contribution to the aggregated Chat Quality Score. If a specific dimension (e.g., User Satisfaction) exhibited a non-linear correlation—such as an S-shaped curve—or significant noise, it would suggest potential flaws in the scoring rubrics or latent hallu- cinations and inconsistencies in the LLM-as-judge. The absence of such anomalies in our results confirms that the SAGE evaluation framework effectively encapsulates the multi-faceted nature of con- versational ability, maintaining high interpretability and minimal bias across all predefined dimensions. Capability Profiling via Variance. While the general trend is linear, the vertical variance (residuals) from the regression lines serves as a fingerprint for specific model strengths and weaknesses. A point significantly deviating from the mean regression line indi- cates that a model’s capability in that specific dimension is dispro- portionate to its overall performance. Points located significantly above the green regression line (Linguistic Quality) represent mod- els with exceptional fluency and expression. Specifically, Doubao- Seed-1.8 and DeepSeek-R1 exhibit positive residuals in this di- mension, indicating that their linguistic generation capabilities are superior to the average level expected for their score range. Points falling below the purple regression line (Instruction Compliance) reveal deficits in following specific constraints. Notably, despite rea- sonable overall scores, Doubao-Seed-1.8 and Qwen3-8B appear significantly below the trend line for Instruction Compliance. This suggests a capability imbalance: these models are highly articulate (high Linguistic score) but prone to ignoring specific formatting or constraint instructions (low Instruction score). This granular analysis proves that SAGE can diagnose subtle trade-offs in model alignment, distinguishing between models that are merely chatty and those that are strictly compliant. 5 Conclusion In this paper, we addressed the limitations of existing customer service benchmarks—specifically their single-dimensional metrics, static interactions, and limited scalability—by proposing SAGE, a graph-guided multi-agent evaluation benchmark. By formalizing Standard Operating Procedures (SOPs) into dynamic graph struc- tures and incorporating adversarial user simulation with a dual-axis evaluation mechanism, SAGE achieves a comprehensive assess- ment of service agents in complex business logic and multi-turn interactions. Our extensive experiments not only validated SAGE’s effectiveness in distinguishing model capabilities but also revealed critical insights, such as the “Inverted-U” performance degrada- tion in long contexts and logical fragility under high adversarial intensity. Furthermore, SAGE’s modular design ensures rapid ex- tensibility to diverse vertical domains via simple configuration. We envision SAGE as a standardized metric tool to drive the evolution of intelligent customer service from simple chatbots to expert-level agents. Future work will explore more complex graph structures (e.g., nested subgraphs) and multimodal interaction scenarios. SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea References [1]Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [2]Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al.2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023). [3]Nolwenn Bernard and Krisztian Balog. 2023. MG-ShopDial: A multi-goal conver- sational dataset for E-commerce. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2775– 2785. [4]Alessandro Berti, Humam Kourani, and Wil MP van der Aalst. 2024. PM-LLM- Benchmark: Evaluating large language models on process mining tasks. In Inter- national Conference on Process Mining. Springer, 610–623. [5]Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018. Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling. arXiv preprint arXiv:1810.00278 (2018). [6]Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al.2025. MiniMax-M1: Scal- ing Test-Time Compute Efficiently with Lightning Attention. arXiv preprint arXiv:2506.13585 (2025). [7]Lei Cui, Shaohan Huang, Furu Wei, Chuanqi Tan, Chaoqun Duan, and Ming Zhou. 2017. Superagent: A customer service chatbot for e-commerce websites. In Proceedings of ACL 2017, system demonstrations. 97–102. [8] Lingxiao Diao, Xinyue Xu, Wanxuan Sun, Cheng Yang, and Zhuosheng Zhang. 2025. GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents. arXiv preprint arXiv:2505.11368 (2025). [9] Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, et al.2024. Workarena: How capable are web agents at solving common knowledge work tasks? arXiv preprint arXiv:2403.07718 (2024). [10] Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs. arXiv preprint arXiv:1903.00161 (2019). [11]Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv e-prints (2024), arXiv–2407. [12]Dirk Fahland, Fabiana Fournier, Lior Limonad, Inna Skarbovsky, and Ava JE Swevels. 2024. How well can large language models explain business processes? arXiv preprint arXiv:2401.12846 (2024). [13] Fabiana Fournier, Lior Limonad, and Inna Skarbovsky. 2024. Towards a Bench- mark for Causal Business Process Reasoning with LLMs. In International Confer- ence on Business Process Management. Springer, 233–246. [14]Michael Grohs, Luka Abb, Nourhan Elsayed, and Jana-Rebecca Rehse. 2023. Large language models can accomplish business process management tasks. In International conference on business process management. Springer, 453–465. [15]Yun He, Di Jin, Chaoqi Wang, Chloe Bi, Karishma Mandyam, Hejia Zhang, Chen Zhu, Ning Li, Tengyu Xu, Hongjiang Lv, et al.2024. Multi-if: Benchmark- ing llms on multi-turn and multilingual instructions following. arXiv preprint arXiv:2410.15553 (2024). [16]Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang. 2024. Followbench: A multi- level fine-grained constraints following benchmark for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 4667–4688. [17]Humam Kourani, Alessandro Berti, Jasmin Hennrich, Wolfgang Kratsch, Robin Weidlich, Chiao-Yun Li, Ahmad Arslan, Wil MP van der Aalst, and Daniel Schuster. 2025. Leveraging large language models for enhanced process model compre- hension. Decision Support Systems (2025), 114563. [18]Humam Kourani, Alessandro Berti, Daniel Schuster, and Wil MP van der Aalst. 2025. Evaluating large language models on business process modeling: frame- work, benchmark, and self-improvement analysis: H. Kourani et al. Software and Systems Modeling (2025), 1–36. [19]Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles. 611–626. [20]Minghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song, Hangyu Li, Haiyang Yu, Zhoujun Li, Fei Huang, and Yongbin Li. 2023. Api-bank: A comprehensive benchmark for tool-augmented llms. arXiv preprint arXiv:2304.08244 (2023). [21]Xiangci Li, Zhiyu Chen, Jason Ingyu Choi, Nikhita Vedula, Besnik Fetahu, Oleg Rokhlenko, and Shervin Malmasi. 2025. Wizard of shopping: Target-oriented e-commerce dialogue generation with decision tree branching. arXiv preprint arXiv:2502.00969 (2025). [22]Zekun Li, Shinda Huang, Jiangtian Wang, Nathan Zhang, Antonis Antoniades, Wenyue Hua, Kaijie Zhu, Sirui Zeng, Chi Wang, William Yang Wang, et al. 2025. Sopbench: Evaluating language agents at following standard operating procedures and constraints. arXiv preprint arXiv:2503.08669 (2025). [23]Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al.2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437 (2024). [24]Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang. 2020. Logiqa: A challenge dataset for machine reading comprehension with logical reasoning. arXiv preprint arXiv:2007.08124 (2020). [25]Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al.2023. Agentbench: Evaluating llms as agents. arXiv preprint arXiv:2308.03688 (2023). [26]Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 36 (2023), 46534–46594. [27] Subhrangshu Nandi, Arghya Datta, Nikhil Vichare, Indranil Bhattacharya, Huzefa Raja, Jing Xu, Shayan Ray, Giuseppe Carenini, Abhi Srivastava, Aaron Chan, et al. 2025. SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents. arXiv preprint arXiv:2506.08119 (2025). [28] Jiao Ou, Junda Lu, Che Liu, Yihong Tang, Fuzheng Zhang, Di Zhang, and Kun Gai. 2024. Dialogbench: Evaluating llms as human-like dialogue systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 6137–6170. [29]Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022), 27730–27744. [30]Feng Peiyuan, Yichen He, Guanhua Huang, Yuan Lin, Hanchong Zhang, Yuchen Zhang, and Hang Li. 2024. Agile: A novel reinforcement learning framework of llm agents. Advances in Neural Information Processing Systems 37 (2024), 5244–5284. [31] Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al.2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789 (2023). [32]Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu. 2024. Infobench: Evaluating instruction following ability in large language models. arXiv preprint arXiv:2401.03601 (2024). [33]Jun Quan, Shian Zhang, Qian Cao, Zizhong Li, and Deyi Xiong. 2020. RiSAWOZ: A large-scale multi-domain Wizard-of-Oz dataset with rich semantic annotations for task-oriented dialogue modeling. arXiv preprint arXiv:2010.08738 (2020). [34]Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020. Towards scalable multi-domain conversational agents: The schema- guided dialogue dataset. In Proceedings of the AAAI conference on artificial intelli- gence, Vol. 34. 8689–8696. [35] Adrian Rebmann, Fabian David Schmidt, Goran Glavaš, and Han van Der Aa. 2024. Evaluating the ability of llms to solve semantics-aware process mining tasks. In 2024 6th International Conference on Process Mining (ICPM). IEEE, 9–16. [36]Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (2023), 8634–8652. [37]Yuchong Sun, Che Liu, Kun Zhou, Jinwen Huang, Ruihua Song, Wayne Xin Zhao, Fuzheng Zhang, Di Zhang, and Kun Gai. 2024. Parrot: Enhancing multi-turn instruction following for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 9729–9750. [38]Xiangru Tang, Yiming Zong, Jason Phang, Yilun Zhao, Wangchunshu Zhou, Arman Cohan, and Mark Gerstein. 2024. STRUC-BENCH: Are Large Language Models Good at Generating Complex Structured Tabular Data?. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies (Volume 2: Short Papers). 12–34. [39]Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al.2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023). [40] Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, et al.2025. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534 (2025). [41] Vicuna Team. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90% chatgpt quality. Vicuna: An open-source chatbot impressing gpt-4 with 90 (2023). [42]Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, et al.2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023). KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. [43]Haoxin Wang, Xianhan Peng, Huang Cheng, Yizhe Huang, Ming Gong, Chenghan Yang, Yang Liu, and Jiang Lin. 2025. ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 276–284. [44]Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345. [45]Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. Self-instruct: Aligning language mod- els with self-generated instructions. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers). 13484–13508. [46]Walter F Wiggins and Ali S Tejani. 2022. On the opportunities and risks of foundation models for natural language processing in radiology. Radiology: Artificial Intelligence 4, 4 (2022), e220119. [47]Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al.2025. The rise and potential of large language model based agents: A survey. arXiv 2023. arXiv preprint arXiv:2309.07864 10 (2025). [48]Congying Xia, Chen Xing, Jiangshu Du, Xinyi Yang, Yihao Feng, Ran Xu, Wen- peng Yin, and Caiming Xiong. 2024. FOFO: A Benchmark to Evaluate LLMs’ Format-Following Capability. arXiv preprint arXiv:2402.18667 (2024). [49]Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al.2024. Finben: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems 37 (2024). [50] Qianqian Xie, Weiguang Han, Xiao Zhang, Yanzhao Lai, Min Peng, Alejandro Lopez-Lira, and Jimin Huang. 2023. Pixiu: A large language model, instruction data and evaluation benchmark for finance. arXiv preprint arXiv:2306.05443 (2023). [51] Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing. 2369–2380. [52]Shunyu Yao, Howard Chen, Austin W Hanjie, Runzhe Yang, and Karthik Narasimhan. 2023. Collie: Systematic construction of constrained text generation tasks. arXiv preprint arXiv:2307.08689 (2023). [53] Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. In The eleventh international conference on learning representations. [54]Aohan Zeng, Mingdao Liu, Rui Lu, Bowen Wang, Xiao Liu, Yuxiao Dong, and Jie Tang. 2024. Agenttuning: Enabling generalized agent abilities for llms. In Findings of the Association for Computational Linguistics: ACL 2024. 3053–3077. [55]Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al.2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414 (2022). [56]Junjie Zhang, Ruobing Xie, Yupeng Hou, Xin Zhao, Leyu Lin, and Ji-Rong Wen. 2025. Recommendation as instruction following: A large language model em- powered recommendation approach. ACM Transactions on Information Systems 43, 5 (2025), 1–37. [57] Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, and Weiming Lu. 2024. Agent-pro: Learning to evolve via policy-level reflection and optimization. arXiv preprint arXiv:2402.17574 (2024). [58] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhang- hao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al.2023. Lmsys- chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998 (2023). [59]Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al.2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems 36 (2023), 46595–46623. [60]Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911 (2023). [61]Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al.2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854 (2023). [62]Jie Zhu, Huaixia Dou, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang, and Fang Kong. 2025. Evaluating, Synthesizing, and Enhancing for Customer Support Conversation. arXiv preprint arXiv:2508.04423 (2025). A Experiment Results Supplementary This appendix provides supplementary data substantiating our main findings. We first validate our multi-agent ensemble via a Single- Judge Bias ablation study. Next, we present granular turn-level analysis to detail multi-turn robustness, followed by a breakdown of performance shifts under varying adversarial intensities. Finally, we decompose Logic and Chat scores into sub-metrics, quantifying the “Execution Gap” and capability imbalances. A.1 Impact of Single-Judge Bias To validate the robustness of our evaluation mechanism, we con- ducted an ablation study to investigate the bias inherent in using a single Large Language Model (LLM) as a judge. Specifically, we selected a subset of models (from the Qwen2.5 and Qwen3 families) to act as both "Service Agents" and "Judge Agents" in a round-robin evaluation setup. The results are visualized in three heatmaps: Overall Average Score (Figure 7), Chat Quality (Figure 8), and Logic Ability (Figure 9). In these plots, the Y-axis represents the Judge Model and the X-axis represents the Evaluated Agent. This experiment reveals two critical limitations of single-judge frameworks: 1. Egocentric Bias (Self-Preference). A distinct "diagonal dom- inance" is observable, particularly in Figure 8 (Chat Quality). Mod- els tend to assign higher scores to their own outputs (or outputs from the same model family) compared to external evaluators. For instance, the diagonal cells in Figure 8 often exhibit deeper colors than the off-diagonal cells in the same column, indicating that a model favors response styles similar to its own training distribution. 2. Systematic Scoring Bias. Significant horizontal variations exist across all three heatmaps, especially in Figure 9 (Logic Abil- ity). This indicates that different judges possess different strictness standards. Some judges (represented by rows with consistently lighter colors) act as "strict graders," systematically assigning lower scores across all agents, while others (darker rows) are more "le- nient." For example, Qwen3-8B as a judge might exhibit a different scoring distribution compared to Qwen2.5-32B. Justification for Multi-Agent Ensemble. These findings demon- strate that relying on a single judge introduces significant variance and bias, where the evaluation outcome is heavily contingent on the specific evaluator’s preferences rather than the intrinsic quality of the service agent. To mitigate this, SAGE employs an Ensemble of Three Judge Agents (using majority voting for classification and average scoring for quality). This collaborative approach effectively smooths out individual model biases, cancels out extreme scoring tendencies, and ensures a more objective and robust evaluation standard, as reflected in our main experimental results. A.2 Supplementary Analysis of Turn This chapter supplements the content of Section 4.5. Table 2 pro- vides a granular view of model performance evolution across di- alogue turns (T1, T5, T10, T15). Consistent with the “Inverted-U” trajectory discussed in the main text, most scenarios (e.g., AirlineRe- fund, PropertyService) exhibit a performance peak around Turn 5, followed by a decline at Turn 15. SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea qwen2.5-7B qwen2.5-14Bqwen2.5-32B qwen3-8B qwen3-14Bqwen3-32B Agent Model qwen2.5-7B qwen2.5-14B qwen2.5-32B qwen3-8B qwen3-14B qwen3-32B Judge Model 605553525149 616159585653 565453515352 506073545664 566270536564 596571576165 Overall 40 50 60 70 80 90 100 Score Figure 7: Overall Average Score (OA) Heatmap. Rows represent Judge models, columns represent Agent models. qwen2.5-7B qwen2.5-14Bqwen2.5-32B qwen3-8B qwen3-14Bqwen3-32B Agent Model qwen2.5-7B qwen2.5-14B qwen2.5-32B qwen3-8B qwen3-14B qwen3-32B Judge Model 757173576469 747671596670 676666536269 626468535767 727373617073 646664525768 Chat Quality 40 50 60 70 80 90 100 Score Figure 8: Chat Quality Heatmap. Darker diagonals indicate self-preference bias. qwen2.5-7B qwen2.5-14Bqwen2.5-32B qwen3-8B qwen3-14Bqwen3-32B Agent Model qwen2.5-7B qwen2.5-14B qwen2.5-32B qwen3-8B qwen3-14B qwen3-32B Judge Model 575148504843 585756575449 535149505148 475974555664 525970526462 586473586264 Logic Ability 40 50 60 70 80 90 100 Score Figure 9: Logic Ability Heatmap. Hori- zontal variances indicate differing judge strictness. •Context Accumulation (T1 to T5): The initial improve- ment in Logic Score (e.g., TelecomPackage: 68.0 to 68.5) in- dicates that models effectively gather user information in early turns to clarify intent. •Context Fatigue (T10 to T15): The subsequent drop (e.g., PropertyService Logic: 58.6 to 55.5) highlights the difficulty of maintaining logical consistency over long contexts. •Chat Stability: Interestingly, Chat Scores often remain sta- ble or decline less than Logic Scores (e.g., EcommerceRefund Chat: 66.5 to 66.2), suggesting that models maintain linguis- tic fluency even when their procedural reasoning falters. A.3 Supplementary Analysis of Adversarial Intensity This chapter supplements the content of Section 4.5. Table 3 details the impact of adversarial intensity on model performance. •Logic Degradation: As expected, Logic Scores generally decrease as intensity rises from Zero to Strong (e.g., Logis- ticsDelivery: 84.0 to 54.5), confirming that adversarial user behaviors successfully challenge the agent’s reasoning capa- bilities. •Format Error Escalation: The JSON Error rate consistently increases with intensity (Average: 1.2% to 3.7%), indicating that high-pressure scenarios induce “instruction drift,” caus- ing models to violate output formatting constraints. •Scenario Specifics: In some scenarios like AirlineRefund, performance actually improves from Zero to Weak, likely because the “Zero” setting involves trivial queries that some powerful models might over-complicate, whereas “Weak” scenarios provide clearer task structures. A.4 Supplementary Analysis of Correlation Analysis of Sub-Metrics This chapter supplements the content of Section 4.7. Table 4 breaks down the Logic and Chat scores into their constituent sub-dimensions. •The Execution Gap: A significant disparity exists between Classification Accuracy (Avg: 72.9) and Action Correctness (Avg: 41.7). This gap is most pronounced in EcommerceRe- fund (88.0 vs. 29.5), proving that while models are adept at understanding intent, they struggle to execute the correct sequence of actions in complex SOPs. •Chat Quality Imbalance: Within Chat Quality, Linguistic Quality (Avg: 69.7) and Instruction Compliance (Avg: 71.8) are high, whereas Anthropomorphism (Avg: 57.3) is notably lower. This suggests that current LLMs are polite and com- pliant but lack the “human touch” or empathy required for high-quality customer service. B Scenario Extending B.1 Scenario Configuration Template Our benchmark includes six customer service scenarios. Each sce- nario follows a unified configuration structure: B.1.1 General Structure. Each scenario configuration contains the following components: • Scenario Metadata: – scenario_id: Unique identifier for the scenario – scenario_name: Human-readable name – description: Brief description of the scenario •Classification Fields: Task-specific fields that the agent must classify based on dialogue context and system informa- tion. Each field has: – field_name: Name of the classification field – data_type: Data type (boolean, string, enum, etc.) – options: List of valid values – description: Detailed description of the field • System Variables: Backend information available to the agent (e.g., user credit level, package status, order status) •Actions: Possible actions the agent can take. Each action has: – action_name: Name of the action – description: What the action does •SOP (Standard Operating Procedure): A decision tree that defines: – Stage sequence (e.g., stage1, stage2, ...) KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. Table 2: Turn-by-Turn Performance Analysis by Scenario. (0-100 Scale). OA ScoreLogic ScoreChat ScoreChat Length ScenarioT1T5T10T15T1T5T10T15T1T5T10T15T1T5T10T15 AirlineRefund61.161.959.857.260.761.458.755.763.164.164.363.237363534 EcommerceRefund 66.967.968.166.567.568.568.566.664.665.566.566.237343836 LogisticsDelivery59.560.560.760.858.959.960.060.762.063.163.361.440394045 OnlineEducation 67.069.471.573.167.569.972.675.065.167.367.365.535353436 PropertyService64.164.759.856.963.964.658.655.564.665.464.462.830303235 TelecomPackage 65.966.464.162.168.068.565.062.957.858.360.358.933333333 Table 3: Adversarial Intensity Analysis: Average Performance Across All Models (0-100 Scale). Scenario OverallLogicChat QualityJSON Error (%) ZeroWeakStrongZeroWeakStrongZeroWeakStrongZeroWeakStrong AirlineRefund49.763.260.946.363.260.363.363.063.41.61.83.4 E-commerceRefund73.168.858.974.869.758.166.265.262.01.01.53.9 LogisticsDelivery 78.259.956.484.059.554.555.261.663.60.62.03.8 OnlineEducation75.774.462.678.176.762.066.165.164.92.32.94.5 PropertyService65.161.467.064.460.368.867.965.960.00.91.03.1 TelecomPackage66.269.554.867.772.655.260.157.153.31.01.63.1 Average68.066.260.169.267.059.863.163.061.21.21.83.7 Table 4: Sub-dimension Analysis: Average Performance Across All Models. Path: Path Correctness; Finals: Finals Correctness; Class.: Classification Accuracy; Ling.: Linguistic Quality; Anth.: Anthropomorphism; Cont.: Content Utility; Satis.: User Satisfaction; Instr.: Instruction Compliance (0-100 Scale). Scenario Logic AbilityChat Quality PathActionClass.Ling.Anth.Cont.Satis.Instr. AirlineRefund64.429.472.568.356.260.662.073.1 E-commerceRefund66.029.588.070.158.862.164.071.4 LogisticsDelivery64.332.366.968.955.558.660.071.6 OnlineEducation67.043.080.271.662.360.061.773.0 PropertyService67.257.164.071.358.360.862.175.0 TelecomPackage 74.658.765.967.952.852.153.966.3 Average67.341.772.969.757.359.060.671.8 –Branching conditions based on classification fields and system variables – Terminal actions at leaf nodes • Evaluation Metrics: –Path Accuracy: Whether the agent follows the correct SOP path –Action Accuracy: Whether the agent selects the correct final action –Classification Accuracy: Accuracy of classifying dialogue fields –Chat Quality: LLM-judged quality of agent responses (5 dimensions) – JSON Error Rate: Whether the agent outputs valid JSON B.1.2 Six Scenarios Overview. (1)Online Education: Student inquiry handling, complaint resolution, and refund negotiation (2)E-commerce Refund: Refund processing, order issue han- dling, and logistics anomalies (3)Telecom Package: Package subscription, modification, billing inquiry, and complaints (4)Property Service: Fee consultation, complaint handling, and maintenance services (5)Logistics Delivery: Package tracking, delivery issues, loss compensation, and returns (6)Airline Refund: Flight rebooking/refund, complaints, and flight information inquiry All six scenarios follow the same configuration structure but differ in their specific fields, actions, and SOP logic. SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea B.2 User Simulator Configuration Template To simulate realistic customer interactions, we configure user sim- ulators with: B.2.1 User Profile Components. • User Intent: The user’s goal in the conversation (e.g., "in- quiry", "complaint", "refund request") •Adversarial level of intent: Three levels to control conver- sation difficulty: –Zero Adversarial Intensity: Cooperative users who accept recommendations –Weak Adversarial Intensity: Users with mild concerns or questions –Strong Adversarial Intensity: Demanding or dissatisfied users requiring negotiation •Personality Traits: Defines user’s communication style (e.g., "friendly", "impatient", "detail-oriented") • Initial System State: Backend information about the user (e.g., order history, account status) B.2.2 Interaction Guidelines. User simulators follow these behav- ioral guidelines: Positive Behaviors (What users should do): • Use natural language consistent with their personality • Gradually disclose information over multiple turns • Stay focused on their intent • Respond appropriately to agent’s actions • Maintain consistent emotion and adversarial level Negative Behaviors (What users should NOT do): • Do not switch intents mid-conversation • Do not reveal all information at once • Do not end the conversation prematurely • Do not break character or mention being simulated B.3 Example: Telecom Package Scenario This subsection provides a complete example using the Telecom Package scenario to illustrate the configuration, prompts, and deci- sion paths. B.3.1 Scenario Configuration. Scenario: Telecom Package Sub- scription and Management Description: Handling customer inquiries about package sub- scription, package modification, billing inquiries, and complaint processing. Classification Fields: • ConsumptionType : User’s dialogue intention (Options: En- quiry, Change, Cancel) • ApplicationTendency: Whether user tends to subscribe to recommended package (Options: Agree, Reject, Hesitate) • ConsumptionProfile: Type of package user prefers (Op- tions: Data, Voice) • EmotionTag: User’s emotion in dialogue (Options: Calm, Dis- content) System Variables: • PackageStatus: User’s current package status (Contract- ed/NoContract) • Penalty: Penalty fee if user cancels contracted package (in- teger) Actions: • ChangeOrder: Process package change for user • GoodBye: Politely end the conversation • TransHuman: Transfer to human agent SOP Stages: (1) stage1 : Field Classification - Classify 4 fields based on dia- logue (2) stage2: User Consumption Intention Judgment - Branch by ConsumptionType (3) stage3 : User Consumption Profile Judgment - Branch by ConsumptionProfile (4) stage4 : User Package Status Judgment - Branch by Pack- ageStatus (5) stage5 : Contract Penalty Situation - Branch by Penalty amount (6) stage6: User Application Tendency Judgment - Branch by ApplicationTendency (7) stage7: User Emotion Judgment - Branch by EmotionTag B.3.2 Prompts. (1) Customer Service Agent Prompt Agent Prompt You are a professional intelligent customer service representative handling [telecommunications package subscriptions]. You must process user enquiries regarding package subscriptions according to the following Standard Operating Procedure (SOP) and system variables, outputting a complete response in JSON format. [System Variables Introduction] PackageStatus: User package status (Contracted/NoContract) Penalty: Penalty fee user needs to pay (int) [SOP Flow Introduction] 1. Field Classification (stage1): Classify the following 4 fields based on the given dialogue history, then jump to stage2. - ConsumptionType: User dialogue intent (Enquiry/Change/Cancel) - ApplicationTendency: Whether user tends to apply for recommended package (Agree/Reject/Hesitate) - ConsumptionProfile: Package type user prefers (Data/Voice) - EmotionTag: User emotion in dialogue (Calm/Discontent) 2. User Consumption Intent Judgment (stage2): Jump based on [ConsumptionType] field. - Jump Logic: Based on the value of [ConsumptionType], Enquiry→stage3; Change→stage4; Cancel→stage5. 3. User Consumption Profile Judgment (stage3): Jump based on [ConsumptionProfile] field. - Jump Logic: Based on the value of [ConsumptionProfile], jump to stage6. 4. User Package Status Judgment (stage4): Jump based on system variable [PackageStatus]. - Jump Logic: Based on the value of system variable [PackageStatus], Contracted→stage5; NoContract→ACTION=ChangeOrder→END. KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. 5. Contract Penalty Situation (stage5): Jump based on system variable [Penalty]. - Jump Logic: Based on the value of system variable [Penalty], Penalty=0→ACTION=ChangeOrder→END; Penalty≠0→stage7. 6. User Application Tendency Judgment (stage6): Jump based on [ApplicationTendency] field. - Jump Logic: Based on [ApplicationTendency] field, Agree→stage4; Reject/Hesitate→ACTION=GoodBye→END. 7. User Emotion Judgment (stage7): Jump based on [EmotionTag] field. - Jump Logic: Based on [EmotionTag] field, Calm→ACTION=ChangeOrder→END; Discontent→ACTION=TransHuman→END. [Action Descriptions] - ChangeOrder: Change package - GoodBye: Politely end the conversation - TransHuman: Transfer to human agent [Output Format Requirements] You must output in the following JSON format (do not include any other text): "classification_output": "ConsumptionType": "Enquiry"/"Change"/"Cancel", "ApplicationTendency": "Agree"/"Reject"/"Hesitate", "ConsumptionProfile": "Data"/"Voice", "EmotionTag": "Calm"/"Discontent" , "cot": "Briefly explain your classification reasoning and SOP flow jump logic", "now_path": ["stage1", "stage2", "stage3", ...], "finals": "Action": "ChangeOrder/GoodBye/TransHuman" , "chat": "Your friendly, professional response to the user based on Action" [Key Requirements - Must Follow] 1. Output pure JSON only, do not include any other content (such as explanations, notes, etc.). 2. JSON must contain the following fields: - classification_output (object) - cot (string) - now_path (array) - finals (object) - chat (string) 3. now_path must start from "stage1" and list the stages passed in order (e.g., ["stage1", "stage2", ...]). 4. The chat field must be limited to 40 words. It is a complete, concise user reply. The language must be consistent with the user’s language. The content must be enclosed in double quotes, and must not contain unescaped double quotes ("), backslashes (\), square brackets ([]), etc.; if you need to quote code or special content, please describe it in text instead of directly including code. 5. The complete JSON should be: ... (The outermost layer must have one and only one pair of curly braces). (2) User Simulator Prompt Template User Simulator Prompt You are role-playing as a telecom operator customer, preparing to consult with customer service or handle package-related services. [Your Identity] - User ID: user_id - Current Package: current_package - User Intent: user_intent - Adversarial level: adversarial_intensity_description - Your Personality: personality [Interaction Requirements] * What you should do: - Use natural language for communication. Preferably choose English, occasionally you can choose Chinese or other languages, but once you choose a language, you must remain consistent throughout the conversation - Respond accordingly based on the customer service representative’s reply and gradually advance the conversation - Strictly focus on your current intent (user_intent), do not deviate to other irrelevant topics - Avoid revealing all information at once, gradually disclose details to keep the conversation going for multiple rounds - Keep each reply concise and natural (5-15 words), simulating the rhythm of a real user’s conversation - When customer service asks for information, provide it gradually according to your personality and background, do not rush to end the conversation - When the problem is not fully resolved, continue to ask for details, confirm processes, or express concerns - If satisfied, express thanks and confirm follow-up steps; if not satisfied, continue to express your demands * What you should NOT do: - Do not cross intent boundaries: If your intent is "package inquiry", do not suddenly switch to "complaint" or "query other services" - Do not end the conversation too early: Do not easily say "OK thank you" and end before the problem is solved, ask for details in multiple rounds - Do not switch languages during the conversation: If you start with English, use English throughout; if you use Chinese, use Chinese throughout - Do not provide all information at once: Simulate real users’ gradual information disclosure - Do not deviate from role settings: Strictly act according to your personality and adversarial intensity - Do not mention that you are AI or simulating: Be fully immersed in the user role [Important Notes] - Maintain role consistency, strictly follow your intent boundaries - Adopt corresponding attitudes according to your adversarial intensity - Insist on your position when necessary - Let the conversation continue naturally, increase interaction rounds by asking, confirming, expressing emotions, etc. (3) Judge Prompt Template We use LLM-as-a-judge to evaluate chat quality. The judge evaluates five dimensions: Chat Quality Dimensions: SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea (a)Linguistic Quality (20%): Grammar, fluency, and professional language (b) Anthropomorphism & Emotion (25%): Natural human-like responses with appropriate emotional tone (c) Content Utility (25%): Accuracy and relevance of information provided (d) User Satisfaction (15%): Whether the response addresses user’s needs (e) Instruction Compliance (15%): Following SOP requirements and action descriptions Each dimension is scored on a 3-level scale (3/6/9 points), and the weighted sum yields a score from 0-100. Judge Prompt Template You are an expert evaluator assessing the quality of customer service agent responses in a telecom package scenario. [Your Task] Evaluate the agent’s response based on dialogue history, user message, and agent’s reply. Provide: 1. Classification of dialogue fields (ConsumptionType, ApplicationTendency, ConsumptionProfile, EmotionTag) 2. Chat quality evaluation across five dimensions [Chat Quality Dimensions] Each dimension is scored on a 3-level scale: 1. Linguistic Quality (Weight: 20%) - 9 points: Fluent, professional, grammatically perfect - 6 points: Generally clear with minor issues - 3 points: Poor grammar, unclear expression 2. Anthropomorphism & Emotion (Weight: 25%) - 9 points: Natural, empathetic, human-like interaction - 6 points: Somewhat robotic but appropriate tone - 3 points: Completely mechanical, inappropriate emotion 3. Content Utility (Weight: 25%) - 9 points: Accurate, comprehensive, directly addresses user needs - 6 points: Partially relevant, some information missing - 3 points: Irrelevant or incorrect information 4. User Satisfaction (Weight: 15%) - 9 points: Fully resolves user concerns, polite and helpful - 6 points: Partially addresses concerns - 3 points: Fails to address user needs, potentially frustrating 5. Instruction Compliance (Weight: 15%) - 9 points: Perfectly follows SOP and action requirements - 6 points: Minor deviations from instructions - 3 points: Significant violations of instructions [Output Format] "classification": "ConsumptionType": "Enquiry/Change/Cancel", "ApplicationTendency": "Agree/Reject/Hesitate", "ConsumptionProfile": "Data/Voice", "EmotionTag": "Calm/Discontent" , "chat_quality_dimensions": "linguistic_quality": 3 or 6 or 9, "anthropomorphism_emotion": 3 or 6 or 9, "content_utility": 3 or 6 or 9, "user_satisfaction": 3 or 6 or 9, "instruction_compliance": 3 or 6 or 9 , "classification_reasoning": "Explanation for classification decisions", "chat_quality_reasoning": "Explanation for each dimension’s score (e.g., 9 points means..., 6 points means..., 3 points means...)" [Requirements] 1. Only output valid JSON 2. Scores must be exactly 3, 6, or 9 for each dimension 3. Provide clear reasoning for both classification and quality evaluation B.3.3 Decision Paths (PathList). The PathList enumerates all valid SOP decision paths for the scenario. Each path specifies: • Classification Items: Values of classification fields • System Variables: Backend state (e.g., PackageStatus, Penalty) • Expected Path: Sequence of SOP stages • Final Output: Terminal action B.3.4 Example Paths. Below are three representative paths from the Telecom Package scenario: Path 1: Enquiry + Agree + Data + NoContract→ ChangeOrder "Classification_items": ["Enquiry", "Agree", "Data", "Calm"], "system_variables": "PackageStatus": "NoContract", "Penalty": 0, "expected_path": ["stage1", "stage2", "stage3", "stage6", "stage4"], "final_output": "Action": "ChangeOrder" Path 2: Change + Contracted + Penalty>0 + Discontent→Tran- sHuman "Classification_items": ["Change", "Agree", "Data", "Discontent"], "system_variables": "PackageStatus": "Contracted", "Penalty": 100, "expected_path": ["stage1", "stage2", "stage4", "stage5", "stage7"], "final_output": "Action": "TransHuman" Path 3: Enquiry + Reject + Voice→ GoodBye "Classification_items": ["Enquiry", "Reject", "Voice", "Calm"], "system_variables": "PackageStatus": "NoContract", "Penalty": 0, "expected_path": ["stage1", "stage2", "stage3", "stage6"], "final_output": "Action": "GoodBye" Note: The complete Telecom Package scenario has 36 valid paths in total. The other five scenarios have similar PathList structures with varying numbers of paths based on their complexity. KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. C SOP Graph of 6 Scenarios in Our Evaluation In this section, we visualize the Standard Operating Procedures (SOPs) formalized as directed graphs for the six industrial scenarios evaluated in SAGE. These graphs define the explicit logical con- straints, state transitions, and action spaces that serve as the ground truth for our Rule Engine evaluation. 1. Ecommerce Refund (ER): As illustrated in Figure 10d, this is the most complex scenario in our benchmark. The graph features a deep decision tree with intricate branching based onShippingStatus (Shipped/Unshipped/Signed) andResponsibilityattribution. Cru- cially, it incorporates dynamic checks on userCreditLevelto de- termine immediate refund eligibility. The agent must navigate this multi-branch logic to verify eligibility and negotiate terms, validat- ing its capability in handling high-complexity disputes. 2. Logistics Delivery (LD): Shown in Figure 10a, this scenario centers on supply chain exception handling. The SOP graph requires the agent to track package status and process insurance claims based onComplaintValidity. Key logic nodes includeRiskStatus(in- tercepting risky orders) andEmergencyLevel(prioritizing urgent registrations), testing the agent’s proficiency in managing urgency and strictly following insurance protocols. 3. Telecom Package (TP): Depicted in Figure 10c, this is a stan- dardized scenario with relatively linear logic. The flow focuses on billing inquiries and plan upgrades, guided byConsumptionType (Enquiry/Change/Cancel). The agent must checkPackageStatus andContractedstates before executing changes. The graph pri- marily evaluates the agent’s accuracy in instruction-following and executing upselling protocols within a structured framework. 4. Property Service (PS): As seen in Figure 10e, this scenario emphasizes community coordination. The graph divides flows into Payment, Complaint, and Repair, requiring checks onHouseStatus (Occupied/Vacant) andFeePaymentStatus. The logic is designed to test the agent’s ability to coordinate offline services (e.g., scheduling repairs) while managing resident emotions in noise complaint sub- branches. 5. Airline Refund (AR): Illustrated in Figure 10f, this is a high- complexity, high-pressure scenario governed by rigid policies. The graph enforces strict validation ofChangeReason(Personal vs. Air- line/Weather) andmemberLevel(VIP/Regular). The agent must precisely calculate dynamic cancellation fees based on these vari- ables. This structure rigorously tests the agent’s precision in policy adherence and its ability to de-escalate passenger anxiety under time constraints. 6. Online Education (OE): Shown in Figure 10b, this scenario focuses on rigorous Risk Control. The graph explicitly includes a isRiskUsercheck to identify potential malicious refunders based on historical records. The agent must strictly adhere to complex refund formulas and select specificPLANoptions (A-F) based on the user’s learning dependency. This design serves as a stress test for logical reasoning and compliance with anti-fraud protocols. SAGE: A Service Agent Graph-guided Evaluation BenchmarkKDD ’26, August 09–13, 2026, Jeju Island, Republic of Korea (a) Logistics Delivery (LD) SOP(b) Online Education (OE) SOP (c) Telecom Package (TP) SOP(d) E-commerce Refund (ER) SOP Figure 10: Standard Operating Procedures (SOPs) for Six Industrial Scenarios evaluated in SAGE. These directed graphs define the logical constraints and action spaces for each domain.(continued) KDD ’26, August 09–13, 2026, Jeju Island, Republic of KoreaLing Shi et al. (e) Property Service (PS) SOP (f) Airline Refund (AR) SOP Figure 10: Standard Operating Procedures (SOPs) for Six Industrial Scenarios.