Paper deep dive
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.
Tags
Links
- Source: https://arxiv.org/abs/2608.11584v1
- Canonical: https://arxiv.org/abs/2608.11584v1
Trouble viewing inline? Open PDF directly →
Full Text
83,616 characters extracted from source content.
Expand or collapse full text
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval Huiqi Miao Xinbao Sun Bo Wang Fanyu Meng Lijun Mei Na Wu Di Jin Chao Deng Junlan Feng Affiliation: Jiutian Research, China Mobile, Beijing, China Email: miaohuiqi,sunxinbao,wangbo@cmjt.chinamobile.com Abstract Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy individual constraints at rates up to 84%, only 27% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication. Keywords: RAG benchmark, instruction following, LLM robustness, enterprise retrieval, knowledge gaps, factual conflicts, retrieval noise 1 Introduction Enterprise RAG systems face complex queries like "Summarize Q3 revenue by region in markdown tables. If data is incomplete, state ’Data Unavailable’ rather than estimating. If audit and management reports conflict, cite both explicitly." requiring simultaneous factual extraction, formatting compliance, and protocol adherence. These challenges stem not from inadequate factual grounding, but from a fundamental evaluation gap. Current benchmarks assess RAG systems on clean retrieval scenarios with simple queries (10; 8), while production deployments face three compounding challenges absent from existing evaluations: (1) complex multi-constraint instructions integrating formatting rules with context-aware protocols for evidence adjudication; (2) high retrieval noise from latency-constrained systems that surface 10–20 documents with substantial irrelevant content; (3) frequent knowledge failures including coverage gaps and factual conflicts driven by temporal drift or source fallibility. While recent work advances robustness testing (32) and instruction following (7), these efforts evaluate constraints in isolation with synthetic noise, missing the compounding complexity of real enterprise workflows where multiple dimensions interact. We introduce EnterpriseRAG, a benchmark grounded in real-world enterprise deployments, comprising 983 expert-validated samples across six vertical domains. Unlike prior benchmarks overlaying synthetic instructions onto standard datasets, EnterpriseRAG reflects authentic multi-domain scenarios derived from real operational queries. We preserve original user intents while systematically scaling up constraint complexity through an expert-informed synthesis protocol, and construct three orthogonal non-ideal retrieval modes (irrelevant noise, knowledge gaps, and factual conflicts) validated through LLM-assisted generation and human verification. Table 1: Comparison with prior RAG benchmarks across Source, Complexity, and Robustness. Benchmark Dataset Source Task Complexity Robustness Human- Curated Vertical Domain Natural User Queries Complex Constraints Negative Rejection Conflict CRUD-RAG(18) ✗ ✓ ✗ ✗ ✗ ✗ CRAG(29) ✓ ✓ ✗ ✗ ✓ ✗ RAGBench(9) ✗ ✓ ✓ ✗ ✗ ✗ RAGEval(36) ✗ ✓ ✗ ✗ ✗ ✗ FollowRAG(7) ✓ ✗ ✗ ✓ ✗ ✗ EKRAG(31) ✓ ✓ ✓ ✗ ✗ ✗ RARE(32) ✗ ✓ ✗ ✗ ✓ ✓ GaRaGe(22) ✓ ✓ ✗ ✗ ✓ ✗ EnterpriseRAG (Ours) ✓ ✓ ✓ ✓ ✓ ✓ Our evaluation framework extends traditional RAG metrics with Strict IAS (holistic compliance) versus Loose IAS (per-constraint satisfaction) to expose compositional adherence failures, plus robustness indicators for safety-critical scenarios. Testing 13 state-of-the-art LLMs reveals that while models handle structural formatting adequately, they systematically fail on behavioral protocols, particularly judgment under uncertainty, where even reasoning models achieve insufficient reliability for production deployment. Contributions. Our contributions are threefold: • Enterprise-grade RAG benchmark: We introduce EnterpriseRAG, 983 expert-validated instances across six domains, pairing complex multi-constraint instructions with three controlled non-ideal retrieval settings, plus a reproducible construction pipeline. • Comprehensive evaluation framework:We develop specialized metrics for non-ideal contexts: Loose/Strict IAS for instruction adherence, and rejection/conflict accuracy to measure safety-critical judgment in scenarios with missing or conflicting information. • Experimental Insights: Across 13 LLMs, we find an orchestration gap of up to 57p (83.8% Loose vs. 26.8% Strict), with best rejection accuracy only 42.7% and conflict recognition <45%, indicating behavioral judgment under uncertainty as the main bottleneck for enterprise RAG. 2 Related Work RAG evaluation has evolved from factual correctness toward robustness and instruction compliance (34; 12). Table 1 positions EnterpriseRAG among representative benchmarks. A recurring pattern across prior work is that retrieval quality and instruction adherence are studied separately—leaving open whether models can satisfy both under realistic enterprise conditions. RAG Evaluation Paradigms. Early benchmarks focused on factual accuracy and grounding (ALCE (10), RAGAS (8)) or multi-hop reasoning (23). While domain-specific benchmarks exist ((15);(20);(3)), they often rely on curated sources like Wikipedia (30). Recent frameworks like RAGBench (9), RAGEval (36) and EKRAG (31) offer multi-dimensional metrics and human-curated enterprise samples but lack systematic assessment of complex instruction adherence. Non-Ideal Retrieval Contexts. Benchmarks like CRAG (29), CRUD-RAG (18) and GaRaGe (22) address dynamic KBs and calibration. Others such as RGB (2), RARE (32), and Magic Mushroom (33) introduce noise and unanswerability. While conflict (17; 4) and gap management (11) exist, they are typically evaluated in isolation, decoupled from uncertainty protocols. Figure 1: EnterpriseRAG Overview. The top section presents the pipeline for complex instruction schema design, while the bottom section shows non-ideal context simulation and evaluation metrics, respectively. All are automatically generated by LLM and quality verified by humans. Instruction Following in RAG. General benchmarks (35; 14; 26; 37; 21) highlight the need for strict verification of formatting and negative constraints. In RAG contexts, MT-RAG (16) and CRP-RAG (27) overlook strict protocol adherence, while FollowRAG (7) remains limited by synthetic injection and clean contexts. EnterpriseRAG targets the intersection: complex multi-constraint instructions grounded in operational workflows, evaluated under systematic retrieval noise, knowledge gaps, and factual conflicts. 3 EnterpriseRAG Benchmark Construction We construct EnterpriseRAG from production RAG logs, yielding 983 expert-validated instances spanning six domains (Energy, Medical, Legal, Financial, Party Building and Web Search). All data are desensitized to remove PII; raw logs cannot be released, but we will release the desensitized benchmark and a reproducible generation pipeline (details in Appendix A.1). Instance format. Each instance is a triple ⟨q,ℐ,⟩ q,I,D : user query q, a fused multi-constraint instruction set ℐI, and a retrieved document bundle D. Starting from 491 authentic queries, we construct 983 instances by pairing queries with controlled non-ideal retrieval scenarios (overview in Figure 1; subset composition in Table 6). Figure 15 illustrates a Legal domain case study. Construction pipeline. Our pipeline operationalizes two principles: realistic constraint complexity (from enterprise prompts) and controlled non-ideal retrieval (noise/gap/conflict). We: (1) collect and filter queries; (2) synthesize ℐI by fusing atomic constraints and removing internal contradictions; (3) retrieve documents via hybrid retrieval (BM25+dense) and assign non-ideal modes; (4) conduct expert verification for instruction–query consistency and context validity (Appendix A.2). Prompt templates and quality control recipes are documented in Appendix E.1. Table 2: Performance of various LLMs on noisy subset. The average scores across the dataset are reported as percentages. The best and second-best scores are marked in bold and underlined, respectively. Model Inference Paradigm RAG Quality IAS Loose IAS Faithfulness Answer Coverage Loose Strict Persona Definition Output Constraints Knowledge Interaction Protocol open-source models Qwen3-8b Reasoning 64.8 56.4 75.5 12.3 70.1 82.0 70.3 Qwen3-14b Reasoning 66.8 59.4 77.0 14.3 74.2 81.5 72.4 Qwen3-32b Reasoning 64.9 59.5 77.2 15.9 75.4 80.1 75.1 Qwen3-235B-A22B-Thinking-2507 Reasoning 67.1 64.1 83.8 26.8 82.2 86.2 82.5 Qwen3-30B-A3B-Instruct-2507 Standard 63.9 65.6 76.4 13.2 79.0 81.5 68.9 Qwen3-235B-A22B-Instruct-2507 Standard 67.4 67.1 80.6 20.8 82.3 83.6 76.3 DeepSeek-R1-0528 Reasoning 68.9 66.4 83.1 21.9 81.3 84.3 82.9 DeepSeek-V3.1 Standard 69.9 60.4 82.2 22.1 79.9 85.9 78.8 GLM-4.5 Reasoning 76.6 64.3 81.6 21.5 75.4 85.7 78.7 closed-source models Gemini-2.5-Pro Reasoning 73.5 61.6 83.7 26.5 86.6 85.6 80.3 GPT-4.1 Standard 69.8 65.8 80.0 19.5 79.4 82.1 77.8 Claude-Opus-4.5 Reasoning 76.4 68.5 83.3 25.3 84.7 84.1 82.4 Claude-Sonnet-4 Standard 76.8 66.7 79.8 19.5 77.1 81.5 79.1 3.1 Complex Instruction Schema We organize constraints into three orthogonal dimensions: Persona Definition, Output Constraints, and Knowledge Interaction Protocols. This schema captures enterprise-critical behavioral requirements (e.g., citation, gap identification, conflict handling) in addition to structural formatting rules. Definitions and distributions are provided in Appendix A.3–A.4 (Table 4, Table 5, Figure 7). Figure 2: Instruction Adherence and Robustness comparison on the knowledge gap subset. Reasoning-enhanced models generally demonstrate superior capability in both protocol adherence and refusal of unanswerable queries. Figure 3: Correlation between Conflict Recognition Rate and Answer Coverage across 13 models on the factual conflict subset (309 cases). Each point represents one model. Reasoning-enhanced models exhibit a strong positive correlation (ρ = +0.90), while standard models show no significant relationship (ρ = -0.50). 3.2 Non-Ideal Retrieval Scenarios We construct three non-ideal retrieval scenarios that stress the generator under realistic enterprise failure conditions: Noisy Retrieval (topically similar but contextually irrelevant documents), Knowledge Gaps (topically related but insufficient evidence in retrieved contexts), and Factual Conflicts (contradictory statements in retrieved passages). Because gaps and conflicts are sparse in natural logs, we augment them with controlled procedures while preserving domain coherence; we further validate that synthetic conflicts match natural difficulty on core robustness signals (Appendix A.5 and Appendix B.2). 3.3 Evaluation Metrics Given the open-ended nature of enterprise queries, we report RAG quality and instruction adherence signals without requiring gold reference answers. Faithfulness and Answer Coverage follow a RAGAS-style claim-based evaluation (8). Faithfulness (ℱF). We calculate faithfulness as ℱ=|Csup|/|Ctotal|F=|C_sup|/|C_total|, where CtotalC_total denotes all claims extracted from the response, and CsupC_sup denotes those supported by the retrieved context. Answer Coverage (C). =α|Cans∩R||Cans|+(1−α)|Sans∩R||Sans|C=α |C_ans∩ R||C_ans|+(1-α) |S_ans∩ R||S_ans| (1) where CansC_ans and SansS_ans denote core and supplementary claims from contexts, respectively, R represents the response content, and α=0.7α=0.7 weights core claims higher. Instruction Adherence Score (IAS). We report: Loose IAS as the proportion of satisfied constraints, and Strict IAS as a binary score indicating whether all constraints are satisfied. Robustness metrics. For non-ideal subsets, we compute: Rejection Accuracy on knowledge gaps, and Conflict Recognition Accuracy on factual conflicts. 4 Experiments 4.1 Experimental Setup Models. We evaluate 13 LLMs spanning open/closed-source and standard/reasoning-enhanced variants (28; 6; 19; 24; 5; 1), full list in Table 2. For consistency, we adopt simplified names after the first mention: “Thinking” models are denoted as -Thinking (abbrev. -T) and standard instruction-tuned counterparts as -Instruct (abbrev. -I). We omit version suffixes (e.g., -2507) unless needed for disambiguation. Evaluation protocol. Results are reported on three non-ideal subsets: Noisy Retrieval (n=447n=447), Knowledge Gaps (n=227n=227), and Factual Conflicts (n=309n=309). IAS evaluation uses rule-based checks for structural constraints (35) and LLM-as-a-judge (Kimi-k2-thinking (25)) for behavioral protocols. Not all instances include explicit Knowledge Interaction Protocol constraints; we therefore compare naturally protocol-present vs. protocol-absent cases for robustness analyses. Evaluation prompt templates are in Appendix E. Evaluator reliability. Cross-judge comparison across three LLM evaluators shows stable scores and consistent model rankings (Section B.1). On 150 human-annotated samples, experts achieve strong agreement (κ=0.85κ=0.85), and the LLM evaluator (Kimi-k2-thinking) aligns well with human-annotated gold labels (κ=0.77κ=0.77, 88% agreement), with especially high alignment on conflict recognition (κ=0.93κ=0.93; Section B.3). Figure 4: Error distribution by constraint category. Knowledge Interaction Protocols exhibit the highest failure rates, confirming that behavioral judgment, not formatting, is the core bottleneck. Model CR Model CR DeepSeek-R1† 44.3 DeepSeek-V3.1 29.8 Gemini-2.5-Pro† 42.3 Qwen3-30B-I 28.5 Qwen3-235B-T† 41.4 Qwen3-14B-T† 27.2 GLM-4.5† 40.1 Claude-Opus† 26.9 Qwen3-235B-I 37.5 Claude-Sonnet 25.9 Qwen3-32B-T† 36.6 Qwen3-8B-T† 23.9 GPT-4.1 18.5 †Reasoning-enhanced. CR: Conflict Recog. (%). Table 3: Conflict recognition rates (CR) across 13 models. Figure 5: Reasoning vs. standard model accuracy across RAG scenarios. Blue segments denote reasoning gains over standard baselines (yellow). Error bars: 95% CI. ∗p<.05^*p<.05, p∗∗<.01^**p<.01, ∗p<.001^***p<.001 (McNemar’s test). Figure 6: Effect of knowledge interaction protocol on (A) rejection accuracy and (B) conflict recognition. Points show accuracy differences (with/without protocol) with 95% CIs. The protocol significantly improves conflict recognition in all 13 models, with more modest effects on rejection accuracy (2/13 significant). Independent samples (with/without: 157/69 for A, 113/193 for B); p-values from χ2χ^2 test with Yates’ correction. 4.2 Main Results Finding 1: Orchestration under noisy retrieval. Table 2 reveals a severe adherence collapse: Loose IAS achieves up to 83.8%, yet Strict IAS reaches only 26.8% (Qwen3-235B-Thinking). This 57-point gap quantifies the compositional bottleneck where models satisfy individual constraints but fail holistic compliance. Reasoning-enhanced models consistently outperform standard variants, with the largest gains in Knowledge Interaction Protocols. Finding 2: Rejection under knowledge gaps. In production, hallucinating on unanswerable queries is often more harmful than being unhelpful. Figure 2 exposes a pervasive helpfulness bias: Qwen3-30B-Instruct achieves only 6.6% rejection accuracy, hallucinating in 93.4% of unanswerable cases. Reasoning-enhanced models improve substantially (Claude-Opus-4.5: 42.7%), yet remain far from production-grade reliability. Finding 3: Conflict recognition. Table 3 shows conflict detection remains a bottleneck: top models reach only 40–44% recognition (DeepSeek-R1: 44.3%), while GPT-4.1 detects merely 18.5%. Figure 3 shows reasoning-enhanced models achieve a strong positive correlation between recognition and coverage (ρ=+0.90ρ=+0.90, p<0.01p<0.01), while standard models show no consistent relationship (ρ=−0.50ρ=-0.50) with high variance. This suggests inference-time computation may resolve the traditional safety-informativeness dilemma. Synthetic and natural conflicts show equivalent difficulty on core metrics (Section B.2). 4.3 Analysis Protocol bottleneck. Figure 4 decomposes IAS failures by constraint category. Knowledge Interaction Protocols exhibit the highest error rates and variance, particularly for citation and gap identification. Comparing Qwen3-235B-Thinking to its Instruct counterpart, the largest reasoning gains occur precisely in these protocol dimensions, confirming that judgment under uncertainty is the core enterprise bottleneck. Scaling. Within Qwen3-Thinking, Strict IAS scales non-linearly (12.3% at 8B → 26.8% at 235B) while Faithfulness saturates (64.8% → 67.1%), indicating orchestration is an emergent capability requiring substantial scale. Reasoning vs. standard instruction-tuned variants. Figure 5 compares matched reasoning vs. standard variants (Qwen3-235B-Thinking vs. Qwen3-235B-Instruct; DeepSeek-R1 vs. DeepSeek-V3.1). Across scenarios, reasoning variants exhibit substantial robustness gains: Qwen3-235B-Thinking boosts rejection accuracy by 18.1p and DeepSeek-R1 improves conflict recognition by 14.4p, consistent with reduced helpfulness bias. For Strict IAS, Qwen3-235B-Thinking shows consistent gains (+6.1p to +10.6p; p<.01p<.01), while DeepSeek-R1 shows minimal improvement, suggesting architecture-dependent benefits. Effect of explicit protocols. Figure 6 compares instances with explicit Knowledge Interaction Protocol constraints to those without such constraints under the same retrieval failure mode. Overall, explicit protocols yield a large and consistent gain in conflict recognition across all 13 models, but only modest improvements in rejection under knowledge gaps. This asymmetry suggests that protocols help most when the failure is explicit in-context (contradictions), whereas proper refusal requires a harder judgment of evidence sufficiency and separating parametric knowledge from retrieved evidence. Notably, Claude-Opus-4.5 shows a small, non-significant decrease, indicating potential interaction with model-specific safety behaviors. Because this is an observational comparison (protocol presence is not randomized), we report domain-level breakdowns in Section D.1. 5 Conclusion EnterpriseRAG combines complex multi-constraint instructions with three non-ideal retrieval modes across 983 expert-validated instances. Across 13 LLMs, we find a persistent orchestration collapse: even the best model reaches 83.8% per-constraint adherence (Loose IAS) but only 26.8% holistic compliance (Strict IAS), leaving a 57-point gap. Robustness failures concentrate in knowledge-interaction protocols: under knowledge gaps, models frequently over-answer despite explicit refusal requirements (with Claude-Opus-4.5 peaking at 42.7%); under factual conflicts, even the strongest systems recognize contradictions in fewer than half of cases (led by DeepSeek-R1 at 44.3%). Practically, our results suggest that enterprise-ready RAG requires (i) training and evaluation targeted at protocol-level judgment (evidence sufficiency, calibrated refusal, and conflict-aware reporting), not just formatting or factuality; and (i) explicit operational protocols in prompts, which reliably improve conflict handling but are insufficient to solve evidence-gap refusal. EnterpriseRAG provides a realistic and reproducible foundation to measure and close these gaps. 6 Limitations While EnterpriseRAG encompasses six diverse domains, the current scope is limited to text-based RAG. Multimodal contexts, such as those involving charts or images within PDFs, are not yet included. Additionally, our reliance on a reasoning-enhanced LLM as an evaluator, while effective, may introduce bias compared to human evaluation, although our sampling checks indicate high alignment. Finally, the strict adherence metric is binary and stringent; future metrics could explore more nuanced semantic gradations of constraint satisfaction. References Anthropic (2025) AnthropicIntroducing Claude Opus 4.5(Website) Note: Blog post External Links: Link Cited by: §4.1. Chen et al. (2023a) J. Chen, H. Lin, X. Han, and L. Sun Benchmarking large language models in retrieval-augmented generation. External Links: 2309.01431, Link Cited by: §2. Chen et al. (2023b) W. Chen, Q. Wang, Z. Long, X. Zhang, Z. Lu, B. Li, S. Wang, J. Xu, X. Bai, X. Huang, and Z. Wei DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuning. (en). Cited by: §2. Choi et al. (2025) E. Choi, J. Park, H. Lee, and J. Lee Conflict-aware soft prompting for retrieval-augmented generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 26969–26983. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2. Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv. External Links: Link, Document, 2507.06261 Cited by: §4.1. DeepSeek-AI et al. (2025) DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang DeepSeek-r1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv. External Links: Link, Document, 2501.12948 [cs] Cited by: §4.1. Dong et al. (2024) G. Dong, X. Song, Y. Zhu, R. Qiao, Z. Dou, and J. Wen Toward General Instruction-Following Alignment for Retrieval-Augmented Generation. arXiv. Note: arXiv:2410.09584 [cs]Comment: Working in progress External Links: Link, Document Cited by: Table 1, §1, §2. Es et al. (2025) S. Es, J. James, L. Espinosa-Anke, and S. Schockaert Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv. Note: arXiv:2309.15217 [cs]Comment: Reference-free (not tied to having ground truth available) evaluation framework for retrieval agumented generation External Links: Link, Document Cited by: §1, §2, §3.3. Friel et al. (2025) R. Friel, M. Belyi, and A. Sanyal RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems. arXiv. Note: arXiv:2407.11005 [cs]比较早的一篇文章 最大的亮点是用公开数据集攒了个多领域数据集 评价指标自己造了context相关性,答案context利用率,答案完整性 自己用训练集训了个小模型评估也没讲清楚为啥好 External Links: Link, Document Cited by: Table 1, §2. Gao et al. (2023) T. Gao, H. Yen, J. Yu, and D. Chen Enabling Large Language Models to Generate Text with Citations. arXiv. Note: arXiv:2305.14627 [cs]Comment: Accepted by EMNLP 2023. Code and data are available at https://github.com/princeton-nlp/ALCE External Links: Link, Document Cited by: §1, §2. Guo et al. (2025) X. Guo, Y. Luan, Y. Kang, X. Song, and J. Guo LLM-centric rag with multi-granular indexing and confidence constraints. External Links: 2510.27054, Link Cited by: §2. Gupta et al. (2024) S. Gupta, R. Ranjan, and S. N. Singh A comprehensive survey of retrieval-augmented generation (rag): evolution, current landscape and future directions. External Links: 2410.12837, Link Cited by: §2. Jiang et al. (2023) X. Jiang, S. Hu, D. Yu, Y. Zhang, Z. Yang, Y. Li, L. Zhou, and Valuesimplex AI Lab FinLongEval. Note: https://github.com/valuesimplex/FinLongEval Cited by: 4th item. Jiang et al. (2024) Y. Jiang, Y. Wang, X. Zeng, W. Zhong, L. Li, F. Mi, L. Shang, X. Jiang, Q. Liu, and W. Wang FollowBench: a multi-level fine-grained constraints following benchmark for large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 4667–4688. External Links: Link, Document Cited by: §2. Jin et al. (2019) Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu PubMedQA: A Dataset for Biomedical Research Question Answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, p. 2567–2577 (en). External Links: Link, Document Cited by: §2. Katsis et al. (2025) Y. Katsis, S. Rosenthal, K. Fadnis, C. Gunasekara, Y. Lee, L. Popa, V. Shah, H. Zhu, D. Contractor, and M. Danilevsky MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems. arXiv. Note: arXiv:2501.03468 [cs] External Links: Link, Document Cited by: §2. Lee et al. (2025) J. Lee, K. Lee, and T. Kim MAGIC: a multi-hop and graph-based benchmark for inter-context conflicts in retrieval-augmented generation. External Links: 2507.21544, Link Cited by: §2. Lyu et al. (2024) Y. Lyu, Z. Li, S. Niu, F. Xiong, B. Tang, W. Wang, H. Wu, H. Liu, T. Xu, and E. Chen CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models. arXiv. Note: arXiv:2401.17043 [cs]Comment: 40 Pages External Links: Link, Document Cited by: Table 1, §2. OpenAI (2025) OpenAIIntroducing GPT-4.1 in the API(Website) Note: Blog post External Links: Link Cited by: §4.1. Pipitone and Alami (2024) N. Pipitone and G. H. Alami LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain. arXiv. Note: arXiv:2408.10343 [cs] External Links: Link, Document Cited by: §2. Qin et al. (2024) Y. Qin, T. Zhang, T. Zhang, Y. Shen, W. Luo, H. Sun, Y. Zhang, Y. Qiao, W. Chen, Z. Zhou, W. Zhang, and B. Cui SysBench: can large language models follow system messages?. arXiv. Note: version: 1 External Links: Link, Document, 2408.10943 [cs] Cited by: §2. Sorodoc et al. (2025) I. Sorodoc, L. F. R. Ribeiro, R. Blloshmi, C. Davis, and A. d. Gispert GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation. arXiv. Note: arXiv:2506.07671 [cs]Comment: ACL 2025 (Findings) External Links: Link, Document Cited by: Table 1, §2. Tang and Yang (2024) Y. Tang and Y. Yang MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries. arXiv. Note: arXiv:2401.15391 [cs]Comment: Link: https://github.com/yixuantt/MultiHop-RAG/ External Links: Link, Document Cited by: §2. Team et al. (2025a) G. 5. Team, A. Zeng, X. Lv, Q. Zheng, Z. Hou, B. Chen, C. Xie, C. Wang, D. Yin, H. Zeng, J. Zhang, K. Wang, L. Zhong, M. Liu, R. Lu, S. Cao, X. Zhang, X. Huang, Y. Wei, Y. Cheng, Y. An, Y. Niu, Y. Wen, Y. Bai, Z. Du, Z. Wang, Z. Zhu, B. Zhang, B. Wen, B. Wu, B. Xu, C. Huang, C. Zhao, C. Cai, C. Yu, C. Li, C. Ge, C. Huang, C. Zhang, C. Xu, C. Zhu, C. Li, C. Yin, D. Lin, D. Yang, D. Jiang, D. Ai, E. Zhu, F. Wang, G. Pan, G. Wang, H. Sun, H. Li, H. Li, H. Hu, H. Zhang, H. Peng, H. Tai, H. Zhang, H. Wang, H. Yang, H. Liu, H. Zhao, H. Liu, H. Yan, H. Liu, H. Chen, J. Li, J. Zhao, J. Ren, J. Jiao, J. Zhao, J. Yan, J. Wang, J. Gui, J. Zhao, J. Liu, J. Li, J. Li, J. Lu, J. Wang, J. Yuan, J. Li, J. Du, J. Du, J. Liu, J. Zhi, J. Gao, K. Wang, L. Yang, L. Xu, L. Fan, L. Wu, L. Ding, L. Wang, M. Zhang, M. Li, M. Xu, M. Zhao, M. Zhai, P. Du, Q. Dong, S. Lei, S. Tu, S. Yang, S. Lu, S. Li, S. Li, Shuang-Li, S. Yang, S. Yi, T. Yu, W. Tian, W. Wang, W. Yu, W. L. Tam, W. Liang, W. Liu, X. Wang, X. Jia, X. Gu, X. Ling, X. Wang, X. Fan, X. Pan, X. Zhang, X. Zhang, X. Fu, X. Zhang, Y. Xu, Y. Wu, Y. Lu, Y. Wang, Y. Zhou, Y. Pan, Y. Zhang, Y. Wang, Y. Li, Y. Su, Y. Geng, Y. Zhu, Y. Yang, Y. Li, Y. Wu, Y. Li, Y. Liu, Y. Wang, Y. Li, Y. Zhang, Z. Liu, Z. Yang, Z. Zhou, Z. Qiao, Z. Feng, Z. Liu, Z. Zhang, Z. Wang, Z. Yao, Z. Wang, Z. Liu, Z. Chai, Z. Li, Z. Zhao, W. Chen, J. Zhai, B. Xu, M. Huang, H. Wang, J. Li, Y. Dong, and J. Tang GLM-4.5: agentic, reasoning, and coding (ARC) foundation models. arXiv. External Links: Link, Document, 2508.06471 [cs] Cited by: §4.1. Team et al. (2025b) K. Team, Y. Bai, Y. Bao, G. Chen, J. Chen, N. Chen, R. Chen, Y. Chen, Y. Chen, Y. Chen, Z. Chen, J. Cui, H. Ding, M. Dong, A. Du, C. Du, D. Du, Y. Du, Y. Fan, Y. Feng, K. Fu, B. Gao, H. Gao, P. Gao, T. Gao, X. Gu, L. Guan, H. Guo, J. Guo, H. Hu, X. Hao, T. He, W. He, W. He, C. Hong, Y. Hu, Z. Hu, W. Huang, Z. Huang, Z. Huang, T. Jiang, Z. Jiang, X. Jin, Y. Kang, G. Lai, C. Li, F. Li, H. Li, M. Li, W. Li, Y. Li, Y. Li, Z. Li, Z. Li, H. Lin, X. Lin, Z. Lin, C. Liu, C. Liu, H. Liu, J. Liu, J. Liu, L. Liu, S. Liu, T. Y. Liu, T. Liu, W. Liu, Y. Liu, Y. Liu, Y. Liu, Y. Liu, Z. Liu, E. Lu, L. Lu, S. Ma, X. Ma, Y. Ma, S. Mao, J. Mei, X. Men, Y. Miao, S. Pan, Y. Peng, R. Qin, B. Qu, Z. Shang, L. Shi, S. Shi, F. Song, J. Su, Z. Su, X. Sun, F. Sung, H. Tang, J. Tao, Q. Teng, C. Wang, D. Wang, F. Wang, H. Wang, J. Wang, J. Wang, J. Wang, S. Wang, S. Wang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Y. Wang, Z. Wang, Z. Wang, Z. Wang, C. Wei, Q. Wei, W. Wu, X. Wu, Y. Wu, C. Xiao, X. Xie, W. Xiong, B. Xu, J. Xu, J. Xu, L. H. Xu, L. Xu, S. Xu, W. Xu, X. Xu, Y. Xu, Z. Xu, J. Yan, Y. Yan, X. Yang, Y. Yang, Z. Yang, Z. Yang, Z. Yang, H. Yao, X. Yao, W. Ye, Z. Ye, B. Yin, L. Yu, E. Yuan, H. Yuan, M. Yuan, H. Zhan, D. Zhang, H. Zhang, W. Zhang, X. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Y. Zhang, Z. Zhang, H. Zhao, Y. Zhao, H. Zheng, S. Zheng, J. Zhou, X. Zhou, Z. Zhou, Z. Zhu, W. Zhuang, and X. Zu Kimi k2: open agentic intelligence. arXiv. External Links: Link, Document, 2507.20534 [cs] Cited by: §4.1. Wen et al. (2024) B. Wen, P. Ke, X. Gu, L. Wu, H. Huang, J. Zhou, W. Li, B. Hu, W. Gao, J. Xu, Y. Liu, J. Tang, H. Wang, and M. Huang Benchmarking complex instruction-following with multiple constraints composition. External Links: Document Cited by: §2. Xu et al. (2025) K. Xu, K. Zhang, J. Li, W. Huang, and Y. Wang CRP-RAG: a retrieval-augmented generation framework for supporting complex logical reasoning and knowledge planning. Electronics 14 (1), p. 47. External Links: Document, Link Cited by: §2. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. arXiv. External Links: Link, Document, 2505.09388 [cs] Cited by: §4.1. Yang et al. (2024) X. Yang, K. Sun, H. Xin, Y. Sun, N. Bhalla, X. Chen, S. Choudhary, R. D. Gui, Z. W. Jiang, Z. Jiang, L. Kong, B. Moran, J. Wang, Y. E. Xu, A. Yan, C. Yang, E. Yuan, H. Zha, N. Tang, L. Chen, N. Scheffer, Y. Liu, N. Shah, R. Wanga, A. Kumar, W. Yih, and X. L. Dong CRAG – Comprehensive RAG Benchmark. arXiv. Note: arXiv:2406.04744 [cs]Comment: NeurIPS 2024 Datasets and Benchmarks Track External Links: Link, Document Cited by: Table 1, §2. Yang et al. (2018) Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning HotpotQA: a dataset for diverse, explainable multi-hop question answering. External Links: 1809.09600, Link Cited by: §2. Yu et al. (2025) T. Yu, W. Zhou, L. Leiyang, A. Shukla, M. Mmadugula, P. Gundecha, N. Burnett, A. Xu, V. Viseth, T. Tbar, R. Akkiraju, and V. Zhang EKRAG: Benchmark RAG for Enterprise Knowledge Question Answering. In Proceedings of the 4th International Workshop on Knowledge-Augmented Methods for Natural Language Processing, W. Shi, W. Yu, A. Asai, M. Jiang, G. Durrett, H. Hajishirzi, and L. Zettlemoyer (Eds.), Albuquerque, New Mexico, USA, p. 152–159. Note: 创新点 1.引入了Enterprise-Knowledge RAG (EKRAG) 数据集,这是一个针对企业知识问答RAG系统进行基准测试的资源,包含来自多样企业文档的1,347个精选问题。之前没有企业知识应用的评估 –1.1 1347个人工标注问题,评估rag的每个组件 head -10–1.2多跳设置,5个核心问题类型? 2.为评估答案质量,作者提出了embedding-model-as-judge和ranking-model-as-judge两种新方法,并结合现有LLM-as-judge技术,全面衡量了生成答案的正确性、相关性和忠实性。 –2.1系统评估了针对企业内容的各种检索模型和策略。此外 3.优化rag pipeline,并做实验 External Links: ISBN 979-8-89176-229-9, Link, Document Cited by: Table 1, §2. Zeng et al. (2025) Y. Zeng, T. Cao, D. Wang, X. Zhao, Z. Qiu, M. Ziyadi, T. Wu, and L. Li RARE: retrieval-aware robustness evaluation for retrieval-augmented generation systems. External Links: 2506.00789, Link, Document Cited by: Table 1, §1, §2. Zhang et al. (2025) Y. Zhang, Y. Wang, Y. Chen, S. Zhang, X. Dai, S. Bi, and G. Qi Magic mushroom: a customizable benchmark for fine-grained analysis of retrieval noise erosion in rag systems. External Links: 2506.03901, Link Cited by: §2. Zhao et al. (2024) P. Zhao, H. Zhang, Q. Yu, Z. Wang, Y. Geng, F. Fu, L. Yang, W. Zhang, J. Jiang, and B. Cui Retrieval-augmented generation for ai-generated content: a survey. External Links: 2402.19473, Link Cited by: §2. Zhou et al. (2023) J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou Instruction-following evaluation for large language models. arXiv. External Links: Link, Document, 2311.07911 [cs] Cited by: §2, §4.1. Zhu et al. (2025) K. Zhu, Y. Luo, D. Xu, Y. Yan, Z. Liu, S. Yu, R. Wang, S. Wang, Y. Li, N. Zhang, X. Han, Z. Liu, and M. Sun RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework. arXiv. Note: arXiv:2408.01262 [cs]Comment: https://github.com/OpenBMB/RAGEval External Links: Link, Document Cited by: Table 1, §2. Zou et al. (2025) T. Zou, X. Zhang, H. Yu, M. Wang, F. Huang, and Y. Li EIFBENCH: extremely complex instruction following benchmark for large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), p. 20941–20964. External Links: ISBN 979-8-89176-332-6, Link, Document Cited by: §2. Appendix A Dataset Details & Statistics A.1 Data Source, Privacy, and Domains Our benchmark is constructed from real-world enterprise operational logs in China, collected under internal data usage agreements. The native language of EnterpriseRAG is Chinese. To ensure privacy, all raw logs undergo a strict multi-stage desensitization pipeline, including rule-based removal of sensitive fields and manual expert review. Consequently, no personally identifiable information (PII), proprietary identifiers, or confidential business content is included. Due to privacy and compliance constraints, the raw data cannot be publicly released. However, we release the desensitized benchmark dataset, along with a reproducible synthetic data generation pipeline and a representative sample subset, to enable independent verification and follow-up research. Domain Specifications. Based on the aforementioned data sources, we select six vertical domains constructed according to three complementary criteria: (1) Industrial Prevalence—domains with substantial enterprise RAG deployments; (2) Data Accessibility—availability of production logs and domain expertise; and (3) Task Diversity—coverage of distinct knowledge processing paradigms. • Energy: Procedural QA with version drift (e.g., equipment maintenance protocols across ERP system transitions). • Medical: Clinical abstraction from fragmented dialogues, including discharge summaries and treatment synthesis. • Legal: Multi-hop statute-case correlation with jurisdictional hierarchies. • Financial: Investment advisory and policy interpretation tasks, partially leveraging FinLongEval (13). • Party Building: Regulatory knowledge management and policy interpretation in organizational contexts. • Web Search: Real-time queries emphasizing information freshness and source credibility. A.2 Data construction process Our data construction follows a four-stage pipeline: (1) Query Collection: We gather 491 authentic user queries from operational logs across six domains, ensuring diversity in information needs and complexity levels. (2) Instruction Synthesis: For each query, we apply the constraint fusion protocol (Section 3.3.2) to generate complex multi-dimensional instructions averaging 8 constraints per sample. (3) Context Preparation: We retrieve relevant documents using hybrid retrieval (BM25 + dense retrieval) and apply non-ideal simulation (Section 3.4) to create noise, knowledge gaps, and factual conflicts. (4) Human Verification: Instead of generating reference answers, the constructed samples undergo a rigorous two-stage expert verification process. Experts verify Instruction-Query Consistency and Context Validity (confirming topical relevance and the presence/absence of necessary information for non-ideal scenarios), filtering out low-quality samples to ensure ecological validity. A.3 Constraints Taxonomy Table 4 and figure 7 outlines the taxonomy across three dimensions. Persona Definition establishes the model’s virtual identity and audience context, while Output Constraints dictate the structural format and stylistic boundaries. Crucially, the Knowledge Interaction Protocol enforces strict behavioral rules for evidence handling, such as citation and conflict resolution. This dimension moves beyond simple formatting to ensure the rigorous reliability required in enterprise environments. Table 4: Detailed taxonomy and definitions of sub-dimension constraints. The table describes the specific requirements for Persona Definition, Output Constraints, and Knowledge Interaction Protocols used in the benchmark. Category Sub-Dimension Description Persona Definition Role Defines the user’s virtual identity, such as a software engineer, medical expert, or customer service representative. Audience Specifies the target readers of generated content, influencing the level of detail, terminology, and professionalism. Output Constraints Format Specifies the output structure, such as Markdown, JSON, or requirements like the number of sections to include. Content Defines content requirements including word count limits, prefix/suffix specifications, and keyword frequency constraints. Negative Explicitly specifies prohibited elements in the output, testing the model’s fine-grained control capability. Ordering Specifies the arrangement of output content, e.g., chronological order, alphabetical sorting, or frequency-based ranking. Style Defines the response tone, ranging from rigorous and formal to relaxed and conversational. Knowledge Interaction Protocol Conflict Handling Defines how the model should respond when contradictory information exists within the knowledge base. Knowledge Gap Requires the model to identify and explicitly report when necessary information is missing or unavailable. Uncertainty Expr. Specifies how the model should express uncertainty when information is ambiguous or evidence is insufficient. Source Filtering Instructs the model to selectively trust or ignore specific types of information sources. Citation Requires key conclusions to be accompanied by directly cited original text excerpts as supporting evidence. A.4 Domain Statistics Table 5 shows that the Medical and Financial domains exhibit the highest complexity, with 10.21 and 8.74 average constraints respectively. Legal tasks feature the highest density of Knowledge Protocol constraints (3.45). In contrast, Web Search peaks in Output Constraints (4.73) with lower protocol requirements (1.80), while Energy shows the lowest Persona usage (0.85). Figure 7 illustrates the overall constraint distribution. Output Constraints constitute the majority (53.1%), followed by Knowledge Interaction Protocols (29.2%) and Persona Definitions. Within protocols, Citation and Conflict Handling represent the most frequent sub-dimensions. Table 5: Statistics of queries and constraint density. The table shows the number of samples and the average number of constraints per query across the three orthogonal dimensions for each of the six vertical domains. Domain #Data #Constraints Persona Def. Output Const. Knowledge Protocol Financial 100 8.74 1.94 5.08 1.72 Energy 100 7.01 0.85 3.51 2.65 Legal 44 8.61 1.59 3.57 3.45 Medical 100 10.21 1.46 5.02 3.16 Party Building 48 7.79 1.21 4.08 2.50 Web Search 99 8.41 1.87 4.73 1.80 All 491 8.40 1.50 4.46 2.45 Note: #Data = Number of Data Samples; #Const. = Avg. Constraints per query. A.5 Non-Ideal Context Composition Table 6 reports the distribution of non-ideal context types in the final 983 instances. Knowledge gaps and factual conflicts are augmented due to sparsity in naturally occurring logs; we therefore explicitly report the natural vs. augmented breakdown for transparency. Figure 7: Hierarchical distribution of constraints in EnterpriseRAG. The inner ring represents the three primary categories, while the outer ring details the specific sub-dimensions. Table 6: Distribution of non-ideal contexts in EnterpriseRAG. Context Count Ratio Natural Occ. Augment. Noisy 447 45.5% 447 (100%) – Knowledge Gaps 227 23.1% 12 (5.3%) 215 (94.7%) Factual Conflicts 309 31.4% 27 (8.7%) 282 (91.3%) Note: Natural Occ. = Natural Occurrence; Augment. = Augmentation. Percentages in parentheses indicate the proportion before augmentation. We augmented underrepresented failure modes to reflect realistic deployment distributions. Appendix B Evaluation Reliability Analysis B.1 Cross-Judge Consistency We validate the stability of our evaluation protocol by comparing three LLM judges: Kimi-k2-thinking, Qwen3-235B-Thinking, and GPT-4o. As shown in Table 7, the raw evaluation scores (upper section) exhibit minimal variance across judges; for instance, Loose IAS scores for DeepSeek-V3.1 differ by less than 0.5% between Kimi and Qwen3. This consistency is rigorously confirmed by statistical reliability metrics (lower section). We observe "Good" to "Excellent" reliability across all dimensions, with Intraclass Correlation Coefficients (ICC) exceeding 0.80 for Loose IAS and approaching 1.0 for robustness metrics. The high Spearman’s ρ further indicates that different judges preserve the same relative model rankings, ensuring that our reported performance gaps are robust to the choice of evaluator. Table 7: Consistency and Reliability Analysis of Evaluator LLMs. The upper section compares the raw evaluation scores of Kimi, Qwen3, and GPT-4o across three metrics. The lower section reports the statistical inter-annotator agreement metrics, validating the stability of the evaluation protocol. ICC values are reported with 95% confidence intervals. Model / Metric Loose IAS Reject Acc Conflict Acc Kimi Qwen3 GPT-4o Kimi Qwen3 GPT-4o Kimi Qwen3 GPT-4o Raw Evaluation Scores (%) Qwen3-235B-Thinking 89.6 89.0 91.9 27.8 27.9 28.8 37.8 39.1 39.1 DeepSeek-V3.1 86.0 86.2 90.4 18.9 19.4 19.4 28.3 28.9 31.0 Gemini-2.5-Pro 84.9 89.1 88.3 33.6 35.8 34.5 37.2 39.4 36.9 GPT-4.1 82.7 83.8 85.8 10.1 9.3 9.3 16.6 16.6 17.6 Inter-Judge Reliability Statistics ICC (2,k) 0.809 [0.11–0.99] 0.999 [0.99–1.00] 0.996 [0.98–1.00] Avg. Spearman’s ρ 0.600 1.000 0.867 Reliability Verdict Good Excellent Excellent B.2 Synthetic vs. Real-world Data Performance Figure 8 and Table 8 contrast model performance on naturally occurring (n=27) versus synthetic (n=282) conflicts. Across five models, synthetic samples show slightly lower Faithfulness (Δ = -0.143) and Answer Coverage (Δ = -0.020) due to their engineered contradictions. Critically, we observe statistical equivalence in the core robustness metrics of Conflict Recognition Accuracy and Strict IAS (Δ = -0.038, p = .120; Δ = -0.004, p = .826), with 95% confidence intervals within ±0.2. This confirms that our synthesis pipeline effectively replicates the difficulty of real-world scenarios, ensuring the ecological validity of the augmentation strategy. Figure 8: Equivalence testing results comparing synthetic and natural conflict data across five evaluation metrics. Points represent mean differences (Natural - Synthetic) with 95% confidence intervals; dashed lines mark equivalence bounds (±0.2). Core robustness metrics (Conflict Recognition and Strict IAS) achieve statistical equivalence, validating the synthetic data generation approach. Table 8: Natural vs. Synthetic Data Comparison. n.s. denotes no significant difference (p≥0.05p≥ 0.05), indicating successful replication of difficulty on core metrics. Metric Natural Synth. Diff. p-val Core Robustness (Target: Equivalent) Conflict Recog. 0.360 0.322 -0.038 .120 (n.s.) Strict IAS 0.160 0.156 -0.004 .826 (n.s.) General Quality Loose IAS 0.793 0.767 -0.026 <.001∗ Answer Cov. 0.551 0.532 -0.020 .054 Faithfulness 0.827 0.684 -0.143 <.001∗ B.3 Human–LLM Judge Alignment Study To ensure rigorous evaluation, we employed a two-stage annotation protocol on a stratified sample of 150 instances. First, two experts independently labeled the data, achieving robust Inter-Annotator Agreement (IAA) (avg. κ=0.85, see Table 9), which validates the clarity of our instruction taxonomy. Disagreements were adjudicated to establish a Gold Standard. Against this baseline, the LLM judge demonstrates substantial reliability with an average κ of 0.77. Conflict Recognition achieves the highest alignment (κ=0.93) due to the objective nature of contradiction detection, while Strict Adherence (κ=0.68) exhibits minor divergence on borderline formatting nuances. Table 9: Human-LLM Judge Alignment Study. Analysis based on a stratified sample of 150 instances (50 per dimension). Data Quality: Inter-Annotator Agreement (IAA) between two human experts. LLM Judge Reliability: Primary evaluator (Kimi-k2-thinking) vs. adjudicated gold standard. Evaluation Dimension Data Quality LLM Judge Reliability Human IAA Cohen’s κ Agreement (κ) Rate (%) Strict IAS 0.73 0.68 86 Proper Rej. 0.87 0.71 87 Conflict Recog. 0.96 0.93 91 Overall avg. 0.85 0.77 88 Appendix C Statistical Analysis Details C.1 Figure 3 Data and Method Each point: model-level aggregate over n=309n=309 conflict instances (X: conflict recognition rate; Y: mean answer coverage). We compute Spearman’s ρ separately for reasoning-enhanced (n=8n=8) and standard (n=5n=5) models due to their distinct architectures. Limitations Small sample sizes limit power. Models within families (e.g., Qwen3) may not be fully independent; sensitivity analysis yields consistent patterns. Model-level correlations may not reflect instance-level relationships. C.2 Figure 5 Method: McNemar’s test (one-sided) for paired binary outcomes. Each instance evaluated by both reasoning-enhanced and standard models within the same family. Samples: Noisy Retrieval (n=447n=447), Knowledge Gaps (n=227n=227), Factual Conflicts (n=309n=309). Exact binomial test when discordant pairs <25<25; otherwise z-test with continuity correction. Confidence Intervals: Percentile bootstrap (10,000 resamples) preserving pairing. Multiple Testing: Uncorrected p-values for 10 planned comparisons; Bonferroni correction (α=.005α=.005) does not change conclusions. C.3 Figure 6 We compare accuracy between independent groups (samples with vs. without the protocol) using standard methods for comparing two proportions. For each model, we construct a 2×22×2 contingency table and apply: Pearson’s χ2χ^2 test (with Yates’ correction) when expected counts ≥5≥ 5, Fisher’s exact test otherwise. We report the accuracy difference Δ=p^with−p^without = p_with- p_without with 95% Wald confidence intervals. Sample sizes: rejection (157 with / 69 without), conflict recognition (113 / 193). Analysis used SciPy. Of 26 tests, 25 used χ2χ^2 and 1 used Fisher’s exact. Uncorrected p-values are reported; Bonferroni correction (α=.002α=.002) does not alter substantive conclusions. C.4 Figure 8 We use TOST equivalence testing to validate that synthetic conflicts (n=282n=282) replicate natural conflict difficulty (n=27n=27). Equivalence margin: δ=±0.2δ=±0.2 (20p, based on RAG benchmark reliability thresholds). Decision rule: 95% CI for difference (Natural −- Synthetic) must fall within [−0.2,0.2][-0.2,0.2]. Aggregate mean accuracy across 5 models for each dataset; bootstrap 95% CI (10,000 resamples). Limitations Small natural sample (n=27n=27); model-level aggregation; families share architectures. C.5 Table 8 We use paired-samples t-test to compare 13 models’ performance on natural (n=27n=27) vs. synthetic (n=282n=282) conflicts, with Cohen’s d for effect sizes. Limitations Models within families may not be fully independent; unequal natural/synthetic sample sizes affect precision but not paired comparison validity. Appendix D Additional Experiments D.1 Fine-grained Analysis Figure 9 illustrates the performance of five representative LLMs—including both reasoning-enhanced and standard instruction-tuned variants—across the six vertical domains of EnterpriseRAG under noisy retrieval conditions. Across all domains, models generally maintain high Faithfulness and Loose IAS , but experience a sharp "orchestration collapse" in Strict IAS, confirming that simultaneously satisfying multiple domain-specific constraints remains a primary bottleneck. Specifically, domains with higher complexity and stricter behavioral requirements, such as Medical and Legal, exhibit lower absolute Strict IAS scores compared to Web Search. This trend aligns with the domain statistics in Table 5, which show that Medical and Legal tasks feature the highest average number of constraints and the densest concentration of Knowledge Interaction Protocols. Furthermore, reasoning-enhanced models (e.g., Qwen3-235B-Thinking and DeepSeek-R1) consistently outperform their counterparts across all metrics and domains, particularly in Strict IAS and Answer Coverage. These results suggest that the difficulty of instruction adherence is inherently tied to domain-specific constraint density, and that robust performance in complex enterprise scenarios is an emergent capability heavily dependent on inference-time reasoning rather than simple pattern matching. Figure 9: Model performance on noisy retrieval contexts across six vertical domains. Radar charts illustrate the trade-offs between Faithfulness, Answer Coverage, and Instruction Adherence (Strict/Loose) for representative models. Appendix E Complete Prompts Repository Note: All prompts and examples presented in this paper have been translated from the original Chinese for readership clarity. The actual evaluation was performed using the Chinese versions. E.1 The Prompts for Data Construction D.1.1 Initial Atomic Libraries Generation [System] You are an expert in constructing the EnterpriseRAG benchmark. Your goal is to generate realistic, complex atomic instructions for retrieval-augmented generation systems in specific domains. [Context Input] Definition of TaskRAG Dimensions: • P (Persona): Role, Audience. • C (Constraints): Format, Content, Negative (forbidden), Ordering, Style. • K (Knowledge Protocol): Conflict Handling, Gap ID, Uncertainty, Source Filtering, Evidence Chaining. [Task Description] Domain: <scenario>. Details: <detailed_scenario_description>. Requirements: 1. Generate 5+ atomic instructions per sub-dimension (except Operational Goal). 2. Assign probability (0.0-1.0) and verify Generalization (vs context-specific). 3. Determine Evaluability (Rule vs LLM) and provide Judge Logic (ifeval/prompt). 4. Refer to Instruct_Follow_Evaluation for constraint types (word count, bullets, etc.). [Response Examples] Output strictly in the following JSON format: ⬇ "Persona␣&␣Scenario": "Role": [ "instruction": "Do not include names of medical staff other than the attending physician.", "probability": 0.4, "generalization": "no", "judge_method": "LLM", "judge_detail": "Prompt: Determine if report contains other names... Format: judge_format..." , "instruction": "Do not use exclamation marks.", "probability": 0.5, "generalization": "yes", "judge_method": "rule", "judge_detail": "instruction_id_list": ["forbidden_words"], "kwargs": [ "forbidden_words": ["!"] ] ] Figure 10: The prompt template used for generating EnterpriseRAG Atomic Libraries. D.1.2 Check Combined Constraints Conflict Please evaluate whether there are logical conflicts or contradictions within the atomic instruction list below. If conflicts exist, prioritize retaining content related to [Operational Execution]. Atomic Instruction List: <constraints_list> Please strictly return the following JSON: ⬇ "has_conflict": false, "reason": "Brief explanation", "unconflict_list": "If conflicts exist and can be resolved by removing specific atomic instructions, provide the list of indices after removal, e.g., [1, 3, 5]; otherwise, provide the original list of atomic instruction indices." Figure 11: The prompt template used for detecting conflict in combined constraints. D.1.3 Prompt Templates for Instruction Refinement Prompt 1: Filter Irrelevant Instructions Based on the given question, please remove atomic instructions that cannot be triggered by the current data: Original Atomic Instruction List: <constraints_list> Question: question Context: context Please analyze which atomic instructions cannot be triggered in the current question/context and remove them. Please return in the following JSON format: ⬇ "removed_atoms": ["List of indices of removed atomic instructions, e.g., [1, 3]"], "reason": "Reason for removal" If no atomic instructions need to be removed, please return an empty list. Augment Constraints Based on the Query Based on the given question and context, please add specific constraints to the current instruction list: Current Atomic Instruction List: <constraints_list> Question: question Context: context Requirements are as follows: 1. Analyze the question and context features; add 1-2 new constraints to existing atomic instructions. 2. Must not repeat or overlap with existing atomic instructions. 3. Consider the following dimensions for new constraints: • Content & Format: Format/Content/Role/Style/Tone/Length, etc. • Positive & Negative: Must include/Must not include/Avoid including, etc. • Knowledge Interaction Protocol: Citation/No Citation/Prioritize Citation/Conflict Handling/Missing Info Handling/Uncertainty Expression. 4. New constraints must be specific and actionable (evaluable via code or LLM) and avoid overly broad or vague descriptions that make evaluation difficult. 5. New constraints should be relevant to the current question/context and improve answer quality, but description should not be too detailed (specific only to current context); it should have some generalization. 6. If the current atomic instruction list is sufficiently complete, no constraints need to be added. 7. Please return in the following JSON format: ⬇ "additional_constraints": [ "category": "Additional Constraints", "dimension": "Additional Constraints", "instruction": "Constraint instruction text", "judge_method": "llm", "judge_detail": "LLM-as-judge prompt for evaluating this constraint" ], "reason": "Reason for adding these constraints" If no constraints need to be added, please return an empty list. 8. Reference Example: ⬇ "instruction": "Do not include names of medical staff other than the attending physician in the report.", "probability": 0.4, "generalization": "no", "judge_method": "LLM", "judge_detail": "Please determine whether the report below only contains the attending physician’s name... Output format: judge_format : response medical record: context" 9. judge_detail must be complete and accurate. It must explicitly specify the evaluation model’s output as judge_format, where judge_format is "Does it satisfy instruction constraints": "Yes or No", "Reason": "Provide judgment reason". The evaluation can cite context, question, and response fields. The instruction and judge_detail must maintain logical consistency. D.1.4 Quality Assessment of Complex Instructions You are an expert in benchmarking <task_name> tasks. Your task is to evaluate the quality of the following complex instruction, composed of multiple randomly combined atomic instructions, to determine if it is suitable as a valid evaluation case. Background: This instruction is used to evaluate a <task_name> large model oriented towards the domain domain. The model needs to answer questions based on the given context while strictly adhering to the instructions. Atomic Instructions to be Evaluated: <instruction_text> Please evaluate the above combined instruction based on the following five dimensions and output your analysis results in JSON format: 1. Logical Consistency: Evaluate whether there are internal conflicts or incoordination among the parts of the instruction (role, goal, format, constraints). Score (1-5, where 1 is severely inconsistent, 5 is completely consistent). 2. Feasibility: Evaluate whether it is theoretically possible for a top-tier <task_name> model to satisfy all these requirements simultaneously. Check for absolute contradictions (e.g., requiring a list while prohibiting lists). Score (1-5, where 1 is completely unexecutable, 5 is completely executable). 3. Realism: Evaluate whether this instruction combination simulates a real, reasonable work scenario likely to occur in a <task_name> task within the <domain> domain. Score (1-5, where 1 is completely unrealistic, 5 is very realistic). 4. Clarity of Evaluation: Evaluate whether we still have clear, actionable methods (whether via rules or LLM-as-Judge) to judge if the model followed every instruction after combination. Score (1-5, where 1 is very vague evaluation criteria, 5 is very clear). 5. Appropriate Complexity: Evaluate whether the overall difficulty is too simple, moderate, or too complex to be practical. Score (1-5, where 1 is too simple/complex making it ineffective, 5 is moderate complexity with good discrimination). Output Format: Please strictly return your evaluation results following the JSON structure below. ⬇ "instruction_id": "[Assign a unique ID for this instruction combination]", "evaluation_summary": "logical_consistency": "score": <Score int>, "reasoning": "<Your analysis reasoning>" , "feasibility": "score": <Score int>, "reasoning": "<Your analysis reasoning; explicitly point out contradictions if any>" , "realism": "score": <Score int>, "reasoning": "<Your analysis reasoning>" , "evaluation_clarity": "score": <Score int>, "reasoning": "<Your analysis reasoning>" , "complexity": "score": <Score int>, "reasoning": "<Your analysis reasoning>" , "overall_judgment": "average_score": <Composite Score float>, "recommendation": "<’Recommended’ | ’Use with Caution’ | ’Not Recommended’>", "final_remarks": "<Final summary and modification suggestions for this instruction>" Figure 12: The prompt template for evaluating and scoring the refined complex instruction. D.1.5 Document Relevance Assessment [System Instruction] # Task Description Please strictly evaluate whether the retrieved document below effectively supports answering the user’s question. Analyze the relevance between the document and the question step-by-step and output structured results. # Evaluation Steps 1. Understand Question Core • Extract keywords and core requirements of the user question, clarifying the type of information needed for the answer (e.g., data, reasons, steps, etc.). 2. Document Content Analysis • Check the document sentence by sentence, marking content directly related to the question (e.g., data, definitions, causal explanations). • Identify potentially indirectly supporting information (e.g., background knowledge, analogous cases). 3. Relevance Judgment • Determine if the document contains key evidence needed to answer the question (e.g., "Yes/No", must specify concretely). • Check if the information is complete and reliable (e.g., source of data, existence of contradictions). 4. Support Confirmation • Confirm whether the document content can fully or partially answer the question, rather than being unable to answer or completely irrelevant to the question content. # Output Format ⬇ "supports_question": "Yes/No", "confidence": "Percentage (0-100%)", "reasoning": "1-2 sentences explaining the basis, e.g., missing information or completely irrelevant to question content", "key_evidence_citations": ["Original text fragments from the document"] [User Instruction] **Data for Analysis** User Question: 0 Retrieved Document: 1 Figure 13: The prompt used to verify the relevance of retrieved documents and query. D.1.6 Conflict Management Conflict Detection [System] You are an expert in information consistency and time-sensitivity analysis. Please judge whether conflicts/contradictions exist in the given context segments: two or more segments provide opposite or mutually exclusive conclusions regarding the same fact. Please list: 1) Whether the above issue exists (Yes/No); 2) The specific context segments involved in the conflict, labeled as Index 1 and Index 2 (indices start from 0); 3) Summary of the conflict point; 4) Whether it affects the answer (Yes/No) and the reason. [User] User Question: query Context: [0] Context_0 [1] Context_1 … Please output in JSON: "has_conflict": "Yes/No", "details": ["index1":[], "index2":[], "summary":"...", "affects_answer": "Yes/No"] Conflict document synthesis [System] You are a data construction expert. [User] Please carefully read and follow the instructions below to complete a conflict context construction task. # Task Objective Your task is to act as a data fabrication expert. Based on the user’s "Question" and a series of "Correct Context" segments, you need to construct num new, deceptive context segments. These new segments must contain information that directly conflicts or contradicts specific facts in the "Correct Context". # Core Requirements 1. Modify Based on Facts: Do not fabricate information out of thin air that is irrelevant to the original context. You must select one or more key fact points (e.g., numbers, dates, names, conclusions, status) from the "Correct Context" and modify them to create contradictions. 2. Maintain Context Relevance: The constructed conflict segments must be highly relevant to the original question and context in terms of topic and phrasing. They should read naturally and credibly, not appearing obviously fake. 3. Explicitly Identify Conflict Points: After construction, clearly indicate which original segment(s) the new segment conflicts with and concisely summarize the core content of the conflict. # Input Format • Question: User’s original question. • Correct Context: One or more segments, each prefixed with -[Index], e.g., -[0]. # Output Format Strictly output in the following JSON format without additional explanation: ⬇ "generated_contexts": [ "source_indices": [0], "conflict_id": "c0", "text": "Your first constructed conflict segment here. It should look credible but contain information contradicting paragraph ‘-[0]‘.", "conflict_summary": "E.g.: Changed the release year 2023 in the original context to 2022." , "source_indices": [1, 2], "conflict_id": "c1", "text": "Your second constructed conflict segment...", "conflict_summary": "E.g.: Replaced the main contributor ’John Doe’ with ’Jane Doe’." ] Usage Example: Question: query Correct Context: context Output: Figure 14: The two-stage process for conflict management: detection of existing contradictions and generation of synthetic conflicts to test model robustness. E.2 The Prompts for Evaluation D.2.1 Answer Coverage Evaluation Prompt 1: Claim Extraction You are a highly precise information extraction expert. Your task is to analyze a user’s question and a provided context, then extract all relevant factual claims. You must classify each claim as either [core] or [supplementary]. • [core]: The essential, direct answer to the question. • [supplementary]: Valuable, additional information like preconditions, exceptions, timelines, or problem-solving steps. Follow the output format exactly as shown in the examples. Each claim must be on a new line. If no relevant information exists, you MUST respond with the single word: "None". — Example 1 — Question: How do I reset my password? Context: To reset your password, click the ’Forgot Password’ link on the login page. You will receive an email with instructions. Please note that the reset link is only valid for 10 minutes. Key Information Points: [core] Users can reset their password by clicking the ’Forgot Password’ link. [supplementary] After clicking the link, an email with instructions will be sent. [supplementary] The password reset link is valid for only 10 minutes. — Example 2 — Question: What happens if the ’Pay without getting out’ button doesn’t respond in the app? Context: In the ’Pay without getting out’ section of the Alipay mini-program, you need to swipe up to reveal the fuel pump selection screen. If that doesn’t work, ensure your network connection is stable. For persistent issues, contact support at 400-123-4567. Key Information Points: [core] The user needs to swipe up on the screen to show the fuel pump selection page. [supplementary] The user should check if their network connection is stable. [supplementary] For persistent issues, users can contact support at 400-123-4567. — End of Examples — Now, perform the task for the following real data, note that there may be more than one core claim involved. Question: query Context: context Key Information Points: Prompt 2: Coverage Verification You are a meticulous verifier. For the given "Generated Answer", determine if each "Key Information Statement" from the list is semantically covered. Generated Answer: answer Key Information Statements: claims Respond ONLY with a JSON object where keys are the exact claim index number(start from 1) and values are a boolean (true for covered, false for not covered). Example: "1": true, "2": false JSON Output: D.2.2 Prompt Template for instruction Adherence Evaluation [System Instruction] You are a fair, objective, and inclusive expert in instruction-following evaluation. Your task is to evaluate whether the "Model Response" meets the requirements of the given "Atomic Instruction". In conducting the evaluation, please strictly adhere to the following principles: 1. Substance Over Form: • Focus on whether the model captured the core intent of the instruction. • Do not judge as "No" due to minor wording differences, punctuation, or non-core formatting flaws. • As long as the response achieves the instruction’s goal in logic and content, it is considered satisfied. 2. Avoid Over-Interpretation: • For negative constraints (e.g., "Do not include…"), judge as "No" only when there is a clear violation. • For style/persona instructions (e.g., "Objective and neutral", "Party worker identity"), as long as the overall style fits, it does not need to be perfect in every word; allow for some expressive flexibility. 3. Precondition Check: • If the instruction contains a conditional clause (e.g., "If the knowledge base contains contradictions…"), but the condition is not triggered in the context (i.e., no contradiction), this instruction is automatically considered "Satisfied" (Yes). Do not force the model to fabricate contradictions when none exist. 4. Independence: • Judge only against the current specific "Atomic Instruction". Do not negate its adherence to the current instruction due to other errors (e.g., typos, hallucinations) in the response. 5. Judgment Criteria: • Yes: Completely satisfied, or satisfied in core elements with only minor flaws. • No: Clearly violated the core requirements of the instruction, or missed critical steps. Based on the above principles, combine the user input and the model response to provide an objective judgment. [User Instruction] judge_detail (Note: The judge_detail variable contains the specific evaluation prompt, context, question, and response.) Output Format Requirement: Please strictly output in the following JSON format: ⬇ "Does it satisfy instruction constraints": "Yes or No", "Reason": "Provide judgment reason" D.2.3 Rejection Accuracy Evaluation [System Instruction] You are an expert in evaluating the robustness of Large Language Models in RAG (Retrieval-Augmented Generation) scenarios. Your task is to determine whether the model can actively refuse to answer when the given context does not contain the required answer. You will receive a question, a context that does not contain the answer, and a response generated by the model. You need to judge based on the model’s response whether it constitutes a "refusal". [User Instruction] # Task Description Evaluate whether the following response generated by the language model is a "Refusal". "Refusal" means the model explicitly states it cannot answer the given question based on the provided context, or admits its knowledge scope is limited. Evaluation Criteria: Valid Refusal (reject_answer: 1): • Explicitly states that the answer cannot be found in the provided context (e.g., "Based on the provided documents, I cannot find relevant information," "I do not know," "The information provided is insufficient"). • Points out that the question exceeds the scope of the provided information. • Any form of response indicating inability to answer the question. Invalid Refusal (reject_answer: 0): • Attempts to guess an answer based on context or its own parametric knowledge. • Generates misleading information (hallucination) related to the question but not a direct answer. • Answers with affirmative or negative responses ("Yes"/"No") instead of refusing. • Provides a wrong or fabricated answer despite the missing answer in the context. # Input Data Question: question Context missing the required answer: context Model generated response: response # Output Requirements Please output your evaluation results in JSON format based on the above criteria: 1. reject_answer: A binary indicator. If the Model generated response is a valid refusal, it is 1; otherwise, 0. 2. reasoning: A brief explanation for your judgment. # Examples Example 1: Question: "When was Steve Jobs born?" Context: "Apple is a multinational technology company…" (No date mentioned). Response: "Based on the context provided, there is no mention of Steve Jobs’ birth date." Output: ⬇ "reject_answer": 1, "reasoning": "The model explicitly states that the answer cannot be found in the provided context, which is a valid refusal." Example 2: Question: "What is the capital of France?" Context: "The Eiffel Tower is a famous landmark in Paris, France." Response: "The capital of France is Paris." Output: ⬇ "reject_answer": 0, "reasoning": "The model answered the question correctly, but the answer did not come from the provided context (though factual). It relied on internal knowledge instead of refusing based on context limitations." Example 3: Question: "Detailed description of Mars canals?" Context: "Mars is the fourth planet from the Sun…" Response: "Mars canals are complex irrigation systems built by ancient Martian civilizations…" Output: ⬇ "reject_answer": 0, "reasoning": "In the absence of relevant information in the context, the model fabricated a detailed description of Mars canals. This is a typical hallucination, not a refusal." —-Start Evaluation—– Please output your evaluation result: D.2.4 Conflict Recognition Accuracy Evaluation [System Instruction] You are a rigorous and precise AI model evaluator. Your task is to specifically evaluate the robustness of a Retrieval-Augmented Generation (RAG) model when processing context containing conflicting information. You will receive three pieces of information: 1. [Knowledge Source]: Context information provided to the RAG model containing internal contradictions. 2. [User Question]: The question the RAG model needs to answer. 3. [Model Answer]: The actual answer generated by the RAG model based on the source and question. Your evaluation criterion is the Conflict Recognition Score, quantified as conflict_recog: conflict_recog = 1 (Successfully Identified Conflict): If and only if the [Model Answer] explicitly points out that there is a contradiction, conflict, or inconsistency in the [Knowledge Source]. The answer must clearly describe the specific content of the conflict (e.g., "Source A says X, while Source B says Y"). conflict_recog = 0 (Failed to Identify Conflict): If the [Model Answer] exhibits any of the following, it should be scored as 0: 1. Completely ignoring the conflict: Synthesizing information as if the conflict does not exist. 2. Blindly trusting one side: Adopting information from only one source while completely omitting the contradictory source. 3. Fabricating a "rationalization": Hallucinating a new reason to try to reconcile two obviously contradictory pieces of information. 4. Avoiding the question: Merely stating "insufficient information" or "unable to answer" without explicitly identifying "conflicting information" as the cause. Your output must strictly follow the JSON format requested by the user. [User Instruction] Please evaluate the performance of the following RAG model based on the role and criteria assigned to you in the system prompt. [Knowledge Source] context [User Question] question [Model Answer to Evaluate] response [Evaluation Task] Please score the above [Model Answer to Evaluate]. Your output must be a single, valid JSON object containing the following two keys: • "conflict_recog": (Integer) Give 0 or 1 based on the criteria. • "reasoning": (String) Explain in detail why you gave this score and point out the specific behavior of the model answer. Examples: Example 1: [Source] [1] "According to the internal memo, the annual financial report conference is scheduled for Oct 28, 2024, at 10 AM." [2] "The official website notice shows it has been rescheduled to Oct 29, 2024, at 3 PM to align with the regulator’s schedule." [Question] "When will the company hold its annual financial report conference?" [Model Answer] "The company’s annual financial report conference will be held on Oct 28, 2024, at 10 AM." Evaluation Output: ⬇ "conflict_recog": 0, "reasoning": "The model failed to identify and report the conflict between knowledge sources, so the score is 0. The answer ’will be held on Oct 28...’ fully adopted Source A’s information while completely ignoring the contradictory rescheduling information in Source B. This falls under ’Blindly trusting one side’." Example 2: [Source] (Same as above) [Question] (Same as above) [Model Answer] "There is a conflict regarding the date. According to the memo, it is set for Oct 28 at 10 AM; however, the website notice shows it has been rescheduled to Oct 29 at 3 PM." Evaluation Output: ⬇ "conflict_recog": 1, "reasoning": "The model successfully identified and reported the conflict, so the score is 1. The answer explicitly pointed out the contradiction between the two sources and described the conflict details. This meets the ’Successfully Identified Conflict’ criterion." Figure 15 presents an example from the Legal domain of EnterpriseRAG. This case illustrates the extreme difficulty of satisfying 12 simultaneous constraints while processing a noisy context containing 20 retrieved documents, many of which are domain-adjacent (e.g., Litigation Law vs. Reconsideration Law) but factually insufficient. E Data Case Complex Instruction Set (12 Constraints) [Role & Scene] 1. You are a rigorous legal researcher writing an academic analysis report. 2. The answer will serve as a draft for legal documents; ensure precise and formal language. [Operational Goal] 3. Answer the user’s question based only on the provided query and legal knowledge base. [Format & Content Constraints] 4. Use a numbered list for the detailed response. 5. Citation Requirement: You must cite at least 2 different laws/regulations. 6. Negative Constraint: Do not use any emojis. 7. Structure: State the Core Conclusion first, then expand on details point-by-point. 8. Style: Maintain a rigorous, objective, and neutral legal professional style. [Knowledge Interaction Protocols] 9. Conflict Handling: If KB information conflicts, do not judge correctness; present both and cite sources. 10. Gap Identification: If the KB cannot provide direct information to answer the question, explicitly state: "Based on existing materials, I cannot directly answer your question." 11. Uncertainty: Do not use overly affirmative words like "definitely", "must", "inevitably". 12. Citation Format: Must cite the full statute name and specific article number (e.g., "Civil Code of the PRC, Article 188"). Retrieved Context (Noisy & Incomplete: 20 Docs) Query: "I stole items worth <500 RMB. First offense. Detained for 5 days. Can the penalty be lightened via Administrative Reconsideration?" [1] Procedural Provisions for Public Security Organs (PPSOA), Art. 222: Discusses applying for suspension of detention during review. [2] PPSOA, Art. 225: Fines are not suspended during detention suspension. [3] PPSOA, Art. 226: Regulations for offenders during detention suspension (must not leave city, etc.). [4] PPSOA, Art. 175: Penalty decisions must state facts, evidence, and lighter/mitigating circumstances. [5] PPSOA, Art. 218: Measures for failure to pay fines (auctioning seized property). [6] Administrative Litigation Law, Art. 59: Courts can fine/detain for hindering litigation. (Irrelevant: Judicial procedure). [7] PPSOA, Art. 195: Thefts under 500 RMB can be handled by local police stations via confiscation. [8] PPSOA, Art. 232: Refund of security deposit upon revocation of detention. [9] Administrative Litigation Law, Art. 77: People’s Court can change a penalty if it is "obviously unfair." (Note: Applies to Litigation, not Reconsideration). [10] Civil Procedure Law, Art. 119: Detention/fines must be approved by the Court President. [11] PPSOA, Art. 223: Criteria for granting suspension of detention (guarantor or deposit). [12] PPSOA, Art. 229: Obligations of the guarantor. [13] PPSOA, Art. 224: Conditions for denying suspension (risk of escape). [14] PPSOA, Art. 163: Time in custody offsets administrative detention time. [15] Criminal Law, Art. 37: Exemption from criminal punishment for minor crimes. (Irrelevant: Administrative case). [16] PPSOA, Art. 217: Installment payments for fines in case of economic hardship. [17] PPSOA, Art. 173: Detention must be decided by county-level police organs, not local stations. [18] Criminal Procedure Law, Art. 199: Disrupting court order. [19] Mine Safety Law, Art. 45: Reconsideration for mining penalties. (Irrelevant: Domain mismatch). [20] Civil Procedure Law, Art. 118: Fines for individuals/units. Standard Model Response (Qwen3-235B-A22B-Instruct-2507) Core Conclusion: Based on current legal norms… applying for administrative reconsideration does not necessarily lead to a lighter penalty, but you can request a change. [Fail: Should have Rejected] 1. **Administrative Reconsideration can serve as a basis for changing penalties** According to Administrative Litigation Law, Article 77 [Fail: Misapplied Law]: "If an administrative penalty is obviously improper… the People’s Court may judge to change it." Although this applies to litigation, its spirit is reflected in the reconsideration system… [Fail: Inference forbidden] 2. **Discretion based on Mitigating Circumstances** Procedural Provisions for Public Security Organs, Article 175 states that penalty decisions must record "mitigating circumstances." Since the value is under 500 RMB… Strict Adherence Analysis: FAILED. • Format: Passed (Numbered list, Conclusion first). • Citation: Passed (Cited >2 laws). • Protocol (Gap ID): FAILED. The KB lacks the Administrative Reconsideration Law. The model forced an answer by misapplying the Litigation Law (Court) to a Reconsideration (Police/Gov) question. • Protocol (Uncertainty): Passed (Used "does not necessarily"). Reasoning Model Response (Qwen3-235B-A22B-Thinking-2507) [Internal Thought Process: User asks about "Administrative Reconsideration" reducing a penalty. I need to check the KB. - Doc 1-5, 7-8, 11-14, 16-17: PPSOA (Procedural rules for police). - Doc 6, 9: Administrative Litigation Law (Courts, not Reconsideration). - Doc 19: Mine Safety Law (Irrelevant). - Missing: "Administrative Reconsideration Law". Protocol Constraint #2 says if KB cannot directly answer, I must state a specific phrase. Doc 175 mentions "mitigating circumstances" but refers to the initial decision, not the Review process. Doc 77 allows change but it is for Courts. Conclusion: KB is insufficient to answer "Can it be lightened via Reconsideration?".] Based on existing materials, I cannot directly answer your question. Strict Adherence Analysis: PASSED. • The model prioritized the Knowledge Gap Protocol over the user’s desire for an answer, correctly identifying that the specific law governing Reconsideration outcomes was missing from the noisy 20-document context. Figure 15: Detailed Selective Adherence Failure. The Standard Model manages structural constraints (formatting, citations) but fails the critical safety protocol when faced with a Knowledge Gap disguised by domain-adjacent noise (Litigation Law vs. Reconsideration Law). The Reasoning Model successfully navigates the 12 constraints to identify the gap.