Paper deep dive
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
Arun Vignesh Malarkkan, Manan Roy Choudhury, Guangwei Zhang, Vivek Gupta, Qingyun Wang, Yanjie Fu, Denghui Zhang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:21:55 AM
Summary
FinRule-Bench is a new benchmark designed to evaluate the ability of Large Language Models (LLMs) to perform rule-based financial auditing. It uses real-world financial statements (Balance Sheets, Cash Flow, Income, and Equity) paired with human-curated accounting principles. The benchmark tests three progressive tasks: rule verification, rule identification, and joint rule diagnosis, utilizing a causal-counterfactual reasoning protocol to assess diagnostic completeness and identify systematic failure modes like mislocalization and incomplete coverage.
Entities (5)
Relation Signals (4)
FinRule-Bench → evaluates → Large Language Models
confidence 100% · FinRule-Bench provides a principled and reproducible testbed for studying rule-governed reasoning, diagnostic coverage, and failure modes of LLMs in high-stakes financial analysis.
FinRule-Bench → includestask → Rule Verification
confidence 100% · The benchmark defines three auditing tasks... (i) rule verification
FinRule-Bench → includestask → Rule Identification
confidence 100% · The benchmark defines three auditing tasks... (ii) rule identification
FinRule-Bench → includestask → Joint Rule Diagnosis
confidence 100% · The benchmark defines three auditing tasks... (iii) joint rule diagnosis
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains poorly explored. Existing benchmarks primarily evaluate question answering, numerical reasoning, or anomaly detection on synthetically corrupted data, making it unclear whether models can reliably verify or localize rule compliance on correct financial statements. We introduce FinRule-Bench, a benchmark for evaluating diagnostic completeness in rule-based financial reasoning over real-world financial tables. FinRule-Bench pairs ground-truth financial statements with explicit, human-curated accounting principles and spans four canonical statement types: Balance Sheets, Cash Flow Statements, Income Statements, and Statements of Equity. The benchmark defines three auditing tasks that require progressively stronger reasoning capabilities: (i) rule verification, which tests compliance with a single principle; (ii) rule identification, which requires selecting the violated principle from a provided rule set; and (iii) joint rule diagnosis, which requires detecting and localizing multiple simultaneous violations at the record level. We evaluate LLMs under zero-shot and few-shot prompting, and introduce a causal-counterfactual reasoning protocol that enforces consistency between decisions, explanations, and counterfactual judgments. Across tasks and statement types, we find that while models perform well on isolated rule verification, performance degrades sharply for rule discrimination and multi-violation diagnosis. FinRule-Bench provides a principled and reproducible testbed for studying rule-governed reasoning, diagnostic coverage, and failure modes of LLMs in high-stakes financial analysis.
Tags
Links
- Source: https://arxiv.org/abs/2603.11339v1
- Canonical: https://arxiv.org/abs/2603.11339v1
Trouble viewing inline? Open PDF directly →
Full Text
101,902 characters extracted from source content.
Expand or collapse full text
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles Arun Vignesh Malarkkan arun.malarkkan@asu.edu Arizona State University Tempe, Arizona, USA Manan Roy Choudhury mroycho1@asu.edu Arizona State University Tempe, Arizona, USA Guangwei Zhang changhu1633@gmail.com Pine AI Singapore, Singapore Vivek Gupta vgupt140@asu.edu Arizona State University Tempe, Arizona, USA Qingyun Wang qwang16@wm.edu The College of William & Mary Williamsburg, Virginia, USA Yanjie Fu yanjie.fu@asu.edu Arizona State University Tempe, Arizona, USA Denghui Zhang dzhang42@stevens.edu Stevens Institute of Technology New Jersey, USA Abstract Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements un- der explicit accounting principles remains poorly explored. Existing benchmarks primarily evaluate question answering, numerical rea- soning, or anomaly detection on synthetically corrupted data, making it unclear whether models can reliably verify or localize rule compli- ance on correct financial statements. We introduce FinRule-Bench, a benchmark for evaluating diagnostic completeness in rule-based financial reasoning over real-world financial tables. FinRule-Bench pairs ground-truth financial statements with explicit, human-curated accounting principles and spans four canonical statement types: Balance Sheets, Cash Flow Statements, Income Statements, and Statements of Equity. The benchmark defines three auditing tasks that require progressively stronger reasoning capabilities: (i) rule verification, which tests compliance with a single principle; (i) rule identification, which requires selecting the violated principle from a provided rule set; and (i) joint rule diagnosis, which requires de- tecting and localizing multiple simultaneous violations at the record level. We evaluate LLMs under zero-shot and few-shot prompting, and introduce a causal-counterfactual reasoning protocol that en- forces consistency between decisions, explanations, and counterfac- tual judgments. Across tasks and statement types, we find that while models perform well on isolated rule verification, performance de- grades sharply for rule discrimination and multi-violation diagnosis. Errors are dominated by incomplete coverage and mislocalization of violations, revealing systematic failures that are not captured by exist- ing financial or tabular benchmarks. FinRule-Bench provides a prin- cipled and reproducible testbed for studying rule-governed reason- ing, diagnostic coverage, and failure modes of LLMs in high-stakes financial analysis. We release the FinRule-Bench benchmark along with the complete evaluation codebase, including task-specific LLM prompting scripts for all three tasks across all financial statement types, at https://anonymous.4open.science/r/FinRule-Bench-6723/. 1 Introduction Large language models (LLMs) are increasingly deployed for fi- nancial analysis and decision support, including applications such as numerical question answering, report summarization, and anom- aly detection. However, a central task in financial auditing remains poorly understood from the perspective of LLM evaluation: deter- mining whether structured financial statements comply with explicit accounting principles and, if not, identifying and localizing the precise sources of non-compliance. This requirement goes beyond producing correct answers or detecting isolated inconsistencies; it demands diagnostic completeness, i.e., the ability to exhaustively en- force a formal rule system over structured, interdependent financial tables. Financial reporting underpins modern capital markets and serves as the primary mechanism through which firms communicate perfor- mance, risk, and regulatory compliance to investors and oversight bodies. The reliability of these reports is enforced through rigorous auditing processes overseen by regulatory authorities such as the PCAOB [34] and grounded in formal accounting frameworks includ- ing U.S. GAAP [12] and IFRS [20]. Auditing requires verifying that financial statements satisfy a large set of explicit and interdependent accounting principles, including arithmetic identities, hierarchical aggregation constraints, and conditional applicability rules. Effective auditing therefore requires structured reasoning that can be evaluated for completeness, requiring models to identify the violated principles and to attribute each violation localized to the specific records, rather than producing binary judgments or isolated answers. Existing benchmarks only partially capture these requirements. General table question answering datasets, such as TAT-QA [44], primarily evaluate single-table extraction and numerical reasoning over individual tables, without requiring validation against external rule systems. Financial question answering benchmark,s including FinQA [2] and ConvFinQA [3], emphasize retrieval and arithmetic computation but do not assess whether models can enforce account- ing principles or discriminate among competing rule violations. As a result, models can achieve high accuracy on existing benchmarks arXiv:2603.11339v1 [cs.AI] 11 Mar 2026 Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Balance Sheet (BS) Income Statement (IS) Statement of Cash Flow (CF) IV. SE 01-03 (Optional)Notes to the Financial Statement I. CF 01-03 I. SI 01-05 I. BS 01-05 e.g. Treasury stock should be reported as a deduction from total shareholders' equity. Data: Financial Statement Financial Principles Deterministic Validators Controlled Minimal Error Injection Validator: TRUE, FALSE Single-violationMulti-violation Task I. Rule VerificationTask I. Rule IdentificationTask I. Joint Rule Diagnosis Single Rule Verify specific rule compliance. + 1. zero-shot, 2. few-shot, 3. few-shot with causal and counterfactual reasoning Minimal Repair Main Failure Modes False Negative (Missed Violations) False Positive (Hallucinated Violations) Partial Coverage (Subset of Rules Found) Mislocalization (Right Rule, Wrong Record) Statement of changes in Equity (SE) eg: SI-03 Multi RulesMulti Rules Record-level Mapping: Highlights & Rule IDs Multi-violation Tables Identify the violated rule. Figure 1: FinRule-Bench Benchmark Overview. FinRule-Bench evaluates LLMs on financial auditing tasks that require both intra-table reasoning and joint reasoning across multiple financial statements under explicit accounting rules. while remaining unsuitable for auditing, where correctness depends on exhaustive rule enforcement and precise attribution rather than isolated answers. This gap motivates a distinct evaluation setting. Verifying ac- counting compliance over ground-truth financial statements differs fundamentally from question answering or anomaly detection. In realistic auditing scenarios, financial tables are often correct by con- struction, and failures arise from incorrect reasoning rather than from corrupted or noisy data. Consequently, evaluating LLMs in this setting requires benchmarks that isolate reasoning fidelity from data artifacts and directly test whether models can exhaustively identify violated rules and attribute each violation to the appropriate records in structured financial reports. To address this need, we introduce FinRule-Bench, a benchmark designed to evaluate rule-based financial reasoning over real-world financial statements. FinRule-Bench pairs structured financial ta- bles with explicit, human-curated accounting principles and defines three auditing tasks that require progressively stronger reasoning capabilities: rule verification, which tests compliance with a single principle; rule identification, which requires selecting the uniquely violated principle from a provided rule set; and joint rule diagnosis, which requires detecting and localizing multiple simultaneous vio- lations at the record level. Together, these tasks evaluate a model’s ability to compose accounting rules, distinguish violated principles, and localize violations to specific records under realistic auditing conditions. Beyond task formulation, FinRule-Bench introduces a causal- counterfactual reasoning protocol for structured financial analysis. This protocol requires models not only to output decisions, but also to justify them through explicit causal explanations and to assess whether minimal counterfactual modifications to the financial table would alter those decisions. Rather than aiming to improve accuracy uniformly, this protocol serves as a diagnostic instrument: by comparing decisions and counterfactual judgments, it reveals inconsistencies, incomplete coverage, and mislocalization errors that are not observable when evaluating final answers alone. Overall, FinRule-Bench provides a principled testbed for evaluat- ing rule-governed reasoning, diagnostic coverage, and failure modes of LLMs in high-stakes financial analysis. By focusing on rule com- pliance over correct financial statements, it complements existing tabular and financial benchmarks and enables precise analysis of where current models succeed and where they fundamentally fall short. We summarize our contributions as follows: (1)Benchmark. We introduce FinRule-Bench, a benchmark derived from real-world financial statements spanning Bal- ance Sheets, Cash Flow Statements, Income Statements, and Statements of Equity. The benchmark pairs ground-truth ta- bles with explicit accounting principles, enabling controlled evaluation of rule-based financial reasoning beyond question answering or anomaly detection. (2)Task Suite. We formalize three complementary auditing tasks that evaluate progressively stronger reasoning capabilities: rule verification, which tests compliance with a single ac- counting principle; rule identification, which requires select- ing the violated principle from a provided rule set; and joint rule diagnosis, which requires multi-violation detection and record-level localization. FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA (3)Reasoning and Evaluation Framework. We introduce a causal-counterfactual prompting protocol for structured finan- cial reasoning and introduce diagnostic evaluation analyses, including rule-complexity breakdowns, error-type categoriza- tion, and performance-cost scalability studies, that reveal systematic failures in diagnostic completeness not captured by aggregate accuracy alone. 2 Related Work LLMs for Table Understanding. Applying LLMs over tabular data poses challenges that differ fundamentally from textual reasoning, including modeling two-dimensional structure, performing precise numerical reasoning, and preserving semantic relationships across rows and columns [11]. Early work addressed these challenges through task-specific, encoder-based models such as TaPas [17], TaBERT [41], TURL [7], and TUTA [37], which explicitly modeled table structure and cell-level interactions. More recent approaches leverage instruction-tuned LLMs and prompting-based methods. Table-centric architectures like Table-GPT [26] and TableLlama [42] adapt LLMs for tabular reasoning, while agent-based frameworks including StructGPT [21] and DATER [40] incorporate external tools or structured representations. Across these approaches, seri- alization strategies and prompting techniques, including Chain-of- Thought reasoning [38], have been shown to substantially improve performance. Comprehensive surveys summarize tasks, datasets, and methodological trends in this area [11,32]. Despite these works, evaluation in tabular reasoning research remains dominated by ques- tion answering and cell-level prediction. These settings do not assess whether models can exhaustively identify and localize rule viola- tions in complex, real-world tables, which is a central requirement in auditing and compliance-oriented applications. LLMs for Financial Data Processing and Analysis. LLMs have been widely applied to financial data processing tasks [25,43], in- cluding algorithmic trading [19], investment analysis [13], market forecasting, sentiment analysis [27], fraud detection [23], and risk as- sessment [24]. A parallel line of work focuses on domain-adapted fi- nancial language models, such as FinGPT [31], trained on large-scale financial corpora, which have been adapted to improve performance. Recent work has also explored Retrieval-Augmented Generation (RAG) to improve factual grounding and contextual awareness in financial analysis [35]. Benchmark efforts such as FinDABench [30] and AuditBench [36] provide standardized evaluation settings for financial tasks involving retrieval, summarization, and prediction. However, these benchmarks largely evaluate information access or descriptive accuracy. They do not test whether models can apply explicit accounting principles to verify the internal consistency or diagnose violations in financial statements. Consequently, high per- formance on existing financial benchmarks does not imply suitability for auditing or compliance verification. Rule-Based and Logical Reasoning on Structured Data. Rule- based and logical reasoning over structured data has a long history in artificial intelligence [18], with recent work shifting toward neu- rosymbolic approaches [14]. Several methods augment LLMs with symbolic reasoning modules, constraint solvers, or logic-guided decoding to improve logical correctness and reasoning faithful- ness [6,33]. Related efforts study complex query answering [39], multi-hop inference [22], and logic-guided neural architectures [29]. Benchmark datasets such as LogicGame [15,28] provide insight into abstract reasoning ability but rely on synthetic rules with limited domain semantics. In contrast, FinRule-Bench evaluates rule-based reasoning grounded in real financial statements and formal account- ing principles, where violations correspond to realistic auditing er- rors and introduce challenges such as rule prioritization, attribution, and multi-violation coverage not captured by existing benchmarks. 3 FinRule-Bench: Dataset and Evaluation Framework Motivated by the limitations of existing tabular financial benchmarks, we introduce FinRule-Bench, a benchmark for evaluating rule-based reasoning over financial tables under explicit accounting principles. FinRule-Bench reflects the core demands of financial auditing, where correctness depends on verifying compliance between multiple in- terdependent rules and structured financial statements. Unlike prior benchmarks that emphasize question answering or anomaly detec- tion on synthetic corrupted data, FinRule-Bench evaluates reasoning over ground-truth financial data and focuses on rule-based reasoning fidelity as the primary objective. FinRule-Bench is built with real-world financial statements ex- tracted from corporate filings. This ensures that the evaluation is focused on realistic tabular data and not on synthesized datasets. Cor- rectness of data is validated by explicit, human-curated accounting principles and is implemented as deterministic validators, making the ground truth reliable and reproducible. To evaluate the LLMs on this benchmark, we curate a controlled, rule-aware error injec- tion to those clean financial statements at record-level. With the clean tabular financial statements as a ground-truth base dataset, the error-injected dataset instances are created via controlled, minimal edits that violate specific accounting principles. The task definitions correspond to different constrained perspectives over this perturbed dataset, and differ only in violation cardinality, required model out- puts, and complexity. In this section, we describe dataset construc- tion with statement-specific rules and validators, task formulation, evaluation protocols with defined prompting techniques, and dataset statistics. 3.1 Dataset Construction: Base Financial Statements FinRule-Bench is constructed from real-world corporate financial disclosures, with 2024 Form 10-K filings as the primary data source (Figure 7). These are standardized, regulator-approved financial fil- ings that adhere to formal accounting frameworks and principles. Financial statements are extracted from filings and converted into structured Markdown tables. We use the Qwen-2.5-VL-72B vision- language model [1] to assist in transcribing visual financial state- ments into Markdown tabular structures while preserving row hierar- chy, column semantics, and numerical values (Table 6). Markdown serves as the canonical representation because it is both human- readable and directly consumable by large language models. Each record in the markdown file corresponds to a single company finan- cial statement for a given fiscal year. The statement types include Balance Sheets, Cash Flow Statements, Income Statements, and Statements of Equity. When a statement of a specific company in Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. a single fiscal year spans multiple pages or images, all fragments are merged into a single coherent Markdown record. Each final- ized table is stored as an individual file indexed by statement type. All tables included in FinRule-Bench correspond to ground-truth financial statements. No numerical values are modified, perturbed, or synthetically corrupted during dataset construction. As a result, any detected violations arise exclusively from evaluating the tables against explicit accounting principles rather than from data noise or extraction artifacts. 3.2 Dataset Construction: Accounting Rules and Deterministic Validators Our benchmark curates explicit accounting principles specific to each statement type. We define four realistic accounting principle sets, each corresponding to the Balance Sheet, Cash Flow State- ments, Statements of Equity, and Income Statements, respectively (Table 22). Each rule set captures arithmetic consistency, hierarchical aggregation, cross-section dependencies, and conditional applica- bility relevant to the corresponding statement. The rule sets are non-overlapping and statement-specific rather than a single global rule pool. These principle/rule sets reflect the corresponding state- ments’ distinct semantics and structure. While some rules require reasoning across multiple rows or sections within a record, FinRule- Bench does not enforce reconciliation constraints across different statement types. This design isolates rule-based reasoning within individual statement types and avoids confounding intra-statement reasoning with cross-statement consistency. Each rule푟 ∈ 푅is associated with a deterministic validator func- tion푣 푟 (푇) ∈ True,False. Validators implement the accounting logic programmatically and are used exclusively to generate ground- truth labels and to evaluate model predictions. Validator outputs are treated as the sole source of truth for model evaluation throughout all experiments. 3.3 Dataset Construction: Controlled Error Injection For controlled evaluation of the models, we construct error-injected variants of the base dataset by deterministically violating the prede- fined accounting principles. Error injection is guided by the account- ing rule definitions per statement type, thereby ensuring realistic accounting errors rather than arbitrary perturbations which is done in some recent research papers [4,5]. The clean financial statements and the controlled, rule-aware error-injected statements together represent a unified dataset, designed to evaluate the joint reasoning capabilities of the models. The error-injection is carried out with respect to two task formulation- based regimes. For the single-violation regimes, exactly one ac- counting principle violation is injected per record. This constraint guarantees unambiguous rule-violation verification and identifica- tion, thereby isolating the intended reasoning requirement. For the multi-violation regime, multiple accounting principle violations are injected within the same financial statement. These violations may affect different rows, sections, or rule types and require multi-step joint reasoning that mirrors realistic auditing scenarios. Across both regimes, error injection preserves the original table structure, schema, and formatting. Only numerical values or derived fields directly rele- vant to the targeted principles are modified. As a result, performance reflects reasoning over accounting rules rather than sensitivity to synthetic formatting changes. 3.4 Task Formulation over FinRule-Bench FinRule-Bench defines three auditing tasks that evaluate progres- sively stronger reasoning capabilities over structured financial tables. Each task corresponds to a distinct auditing problem and evaluates whether a model can interpret accounting principles, apply them consistently, and localize violations when present. These tasks cor- respond to different formulations over the same unified dataset and differ in violation cardinality and required model outputs. Rule Verification.Given a financial markdown table and a single accounting rule, the model predicts whether the table complies with the specified rule. This task evaluates basic rule compliance with a binary decision. Rule Identification.Given a financial table and a set of accounting rules, the model identifies the violated rule at a record level. Each record in the statement-specific table in this setting contains ex- actly one violated rule, ensuring unambiguous model evaluation and requiring discrimination among competing accounting principles. Joint Rule Diagnosis.Given a financial table and a set of accounting rules, the model first determines whether any violations are present per record per statement-type and, if so, identifies all violated rules with record-level attribution. This task evaluates multi-step diagnos- tic joint reasoning and coverage under realistic auditing conditions where multiple violations may co-occur. All tasks operate within individual financial statements and do not assume explicit cross-statement consistency constraints. Formal task definitions are provided in Appendix A. 3.5 Prompting and Inference Protocols We compare three prompting protocols that differ in structure and model supervision. Zero-shot. Models recieve a task instruction, the serialized financial table푇, and the relevant rule or rule set. No exemplars are provided in this protocol. Few-shot. Models receive a small set of exemplars formatted in the same input/output style as the task, followed by the test instance. Exemplars are chosen to cover diverse rule types where possible and are fixed across the models evaluated to ensure fairness. Few-shot with causal and counterfactual reasoning. Beyond standard zero-shot and few-shot prompting, we introduce a causal- counterfactual prompting protocol to probe rule-based reasoning fidelity. We implement this prompting protocol using in-context learning, that vary systematically across three task formulations. Across all settings, causal and counterfactual reasoning structure is introduced exclusively through few-shot exemplars, rather than through architectural changes or additional supervision. In causal- counterfactual prompting, exemplars are constructed to make causal dependencies explicit rather than implicit. They illustrate both the condition that causes a rule violation and a minimal counterfactual modification that would remove the violation. Rule Verification Task In this setting, causal-counterfactual prompts are rule-specific and include compact exemplars that illustrate a single violation pattern together with a minimal counterfactual repair, such as correcting an arithmetic value or replacing a non-standard label. This evaluates whether exposure to causal and counterfactual FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA structure improves the detection of clear violations without requiring discrimination among multiple rules. Rule Identification Task Here, causal-counterfactual exemplars are designed to support rule discrimination by clarifying which table attributes are causally relevant to each rule and the corresponding counterfactual diagnosis of other violated rules and apparent irregu- larities. Joint Rule Diagnosis TaskIn this joint diagnosis reasoning task, the exemplars contain detailed causal reasoning and counterfactual test- ing. However, the model’s output at inference time is restricted to a structured list of rule codes; explanations are not required. This setting evaluates whether models understand causal structure within each record in the table with respect to their rule compliance, suffi- ciently to achieve broad coverage when multiple violations interact under joint compositional reasoning. Causal-counterfactual prompting does not enforce reasoning con- sistency and does not guarantee improved accuracy. Instead, it serves as a conditioning mechanism that exposes causal structure during in- context learning. Differences in performance across difficulty levels and prompting regimes therefore reflect the extent to which models internalize causal relevance and minimal interventions when rea- soning about accounting rule violations, particularly in settings that require rule discrimination and multi-rule coverage. We emphasize that causal-counterfactual reasoning is introduced solely through exemplar design and does not rely on chain-of-thought prompting or post-hoc explanation scoring. 3.6 Model Evaluation and Verification Protocol Model evaluation is carried out using the deterministic rule valida- tors corresponding to the defined accounting principles. The model outputs are parsed into structured representations and compared directly with the ground-truth. For rule verification and rule iden- tification tasks, evaluation is based on exact-match protocol. For joint-rule diagnosis, the multi-step correctness requires accurate vi- olation detection and accurate violated rule identification at record level. This protocol ensures clear objective fairness and reproducible evaluation. With respect to the evaluation protocols, FinRule-Bench uses task-aligned, deterministic evaluation metrics computed via exact matching against rule validators.Rule Verification.We report Accu- racy, Precision, Recall, and F1-score to capture overall correctness and the trade-off between false alarms and missed violations. Rule Identification.We report Top-1 Accuracy as the primary metric. To account for rule imbalance, we additionally report Balanced Accuracy, Macro-F1, and Weighted-F1. Joint Rule Diagnosis.For violation detection, we report binary Ac- curacy and F1-score. For correctness, we report Exact Match and Micro-F1 to quantify partial coverage in multi-violation cases. Efficiency Metrics. For all tasks, we report Total Tokens, Average Tokens per Instance, Cost per Instance, and Total Cost to analyze performance–efficiency trade-offs. 3.7 Dataset Statistics We summarize the composition of FinRule-Bench across statement types and task formulations to provide a quantitative overview of the dataset. FinRule-Bench spans four financial statement types: Balance Sheets (BS), Cash Flow Statements (CF), Statements of Table 1: Dataset statistics for FinRule-Bench. Each record cor- responds to a company’s financial statement. Single-violation records contain exactly one violated rule, while multi-violation records contain multiple simultaneous violations. Statement Type # Rules Ground-Truth Records Single-Violation Records Multi-Violation Records Balance Sheet (BS)52121060100 Cash Flow (CF)3318954100 Statement of Equity (SE)3287871100 Income Statement (SI)53001500100 Total1611174385400 26.1% 21.6% 8.8% 43.5% Balance Sheet (32.2 R × 4.4 S ÷ 4.0 C) Cash Flow (34.4 R × 3.4 S ÷ 4.0 C) Statement of Equity (10.1 R × 10.0 S ÷ 8.5 C) Income Statement (22.4 R × 7.9 S ÷ 3.0 C) Structural Complexity Distribution Complexity Score = Mean Rows (R) × Mean Sections (S) ÷ Mean Columns (C) Figure 2: Structural complexity distribution across financial statement types. Equity (SE), and Income Statements (SI). For each statement type, we define a statement-specific set of accounting rules and construct markdown tables with corresponding clean, single-violation, and multi-violation records, respectively. Each record in a table corre- sponds to a single company financial statement for a given fiscal period. Clean instances consist of fully compliant financial state- ments, while error-injected ones contain deterministically generated rule violations. For rule verification and rule identification, each instance contains at most one violation. For joint rule diagnosis, instances may contain multiple violations affecting different records or principles. Across all statement types, the dataset covers a diverse range of accounting principles, including arithmetic consistency rules, structural and classification rules, conditional applicability rules, and multi-section reconciliation rules. This diversity ensures that evaluation supports a comprehensive assessment of financial rule-based reasoning. Table 1 and figure 2 summarize the dataset statistics, rule sets, and structural complexity distribution. 3.8 Reproducibility and Artifacts All validators, error-injection prompts, exemplar prompts, and eval- uation scripts are deterministic and included in the release artifacts. Model outputs are parsed by deterministic parsers to structured forms used for exact matching. We provide the rule definitions, validator code, exemplar prompts, and injection scripts to ensure complete reproducibility of reported results. Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. 4 Experiments 4.1 Experimental Setup We describe the models, data, prompting protocols, inference con- figuration, and evaluation procedures used in our experiments. Models. We evaluate four representative large language models: GPT-4o [10], Gemini 2.5 Pro (thinking) [9], Gemini 2.0 Flash [16], and LLaMA 3.3 [8] that are both reasoning-focused and efficiency- oriented architectures. Data and splits. The dataset used for all the experiments is described in Section 3. Clean markdown tables are validated prior to instan- tiation, and error-injected tables are used according to the targeted tasks (Section 3.4). Test sets are disjoint from any exemplars used for few-shot prompting. All per-task instance counts and split statistics are reported in Table 1. Prompting Protocols.All models are evaluated under three prompt settings: zero-shot, few-shot, and few-shot with causal-counterfactual reasoning. Prompts include the task instruction, the serialized finan- cial table 푇 in markdown format, and the relevant rule or rule set. In-Context Exemplars.In the two few-shot strategies, we provide a fixed number of labeled exemplars preceding the test instance. Exemplars follow the same input-output format as the target task and are selected to cover diverse rule types. The same exemplar set is used across all models to ensure fair comparison. Decoding and Inference.All models are evaluated using determinis- tic decoding with temperature set to zero and a set output length. No external tools, retrieval mechanisms, or post-processing heuristics are used beyond deterministic output parsing. Evaluation Protocol.Each test instance is evaluated independently. For tasks involving stochastic components in generation, results are averaged over multiple runs with identical settings. Ground-truth labels are generated exclusively by deterministic rule validators. Cost and Token Accounting.We record prompt tokens, completion tokens, and total tokens for every model invocation. Monetary cost is computed using publicly documented pricing for each model at the time of evaluation. The recorded token statistics enable performance- efficiency trade-off analyses. 4.2 Experimental Results 4.2.1 Overall Performance Across Tasks. Table 2 reports model performance of GPT-4o, Gemini 2.5 Pro (thinking), Gemini 2.0 Flash, and LLaMA 3.3 across rule verification, rule identification, and joint rule diagnosis on Finrule-Bench. These tasks progressively increase reasoning demands from isolated rule checking to multi- rule diagnostic reasoning. On rule verification, all models achieve relatively high accuracy, indicating that basic compliance checking is largely within reach of modern LLMs. Performance differences among models are modest in this setting, suggesting that basic com- pliance checking against a single accounting principle is largely tractable for modern LLMs. Performance drops substantially on rule identification, where models must isolate the violated princi- ple among all accounting principles. Accuracy decreases across all models, and the performance gap between reasoning-oriented and efficiency-oriented models widens. Performance degrades further on joint rule diagnosis. While Step-1 violation detection remains rela- tively high, Step-2 diagnostic exact match is low across all models. Table 2: Overall Task performance on FinRule-Bench. ModelBaseline Prompt Rule Verification Rule Identification Diagnosis Step-1 Diagnosis Step-2 AccAccAcc Exact Match GPT-4o Zero-shot0.6400.5920.7130.234 Few-shot0.7360.6890.7860.297 Few-shot + CR0.7380.6610.8010.342 Gemini 2.5 Pro (Thinking) Zero-shot0.6140.5470.6950.216 Few-shot0.6290.6120.7280.259 Few-shot + CR0.6430.6010.7420.304 Gemini 2.0 Flash Zero-shot0.5640.5240.6310.152 Few-shot0.5660.5580.6670.189 Few-shot + CR0.5870.5690.6840.223 LLaMA 3.3 Zero-shot0.5530.4960.6190.147 Few-shot0.6920.5320.6540.181 Few-shot + CR0.6810.5410.6720.215 Even the strongest systems fail to localize all violated rules in most cases. However, this is where causal-counterfactual few-shot prompt- ing proves to be effective, giving the largest relative gains across different models. These results highlight the challenges in exhaus- tive, record-level reasoning under multiple interacting constraints. Also, we observe that the causal-counterfactual prompting yields clearer benefits for lightweight models like Gemini 2.0 Flash and LLaMA 3.3, where it consistently improves performance across veri- fication, detection, and diagnostic localization. In contrast, reasoning- oriented models, such as GPT-4o and Gemini 2.5 Pro (thinking), exhibit inconsistent gains, and in some cases degraded performance, particularly on rule identification. This suggests that explicit causal scaffolding can introduce redundancy or interfere with discrimi- native decision-making in models that already perform extensive internal reasoning. Overall, as task requirements progress from verifi- cation to discrimination and diagnosis, model performance degrades non-linearly, motivating the detailed diagnostic analyses presented in subsequent sections. 4.2.2 Rule Complexity and Error Analysis. In this analysis, we evaluate model performance with respect to stratified rule complex- ity, error type, and distribution across the formulated tasks. This study isolates the semantic properties of rules that drive failures across tasks, beyond what is visible from aggregate accuracy. Rule complexity analysis.Tables 3, 4, and 5 report performance strat- ified by rule category for rule verification, rule identification, and joint rule diagnosis tasks, respectively. Across all tasks, arithmetic rules are consistently the easiest, achieving high performance un- der zero-shot prompting and showing limited gains from additional reasoning structure. Structural rules exhibit intermediate difficulty and benefit most from few-shot and causal-counterfactual prompting strategies, reflecting improved schema and section-level interpreta- tion through exemplars. Conditional and multi-record rules are the most challenging across settings. Zero-shot performance is low, and although few-shot prompting improves results, accuracy remains substantially below that of simpler straight-forward rule types. We also observe that the Causal-counterfactual prompting yields its clearest benefits for multi-record constraints, particularly in the joint FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA Table 3: Rule Complexity Analysis on Rule Verification. Acc Category#Rules #Instances Zero-shot Few-shot Few-shot + CR Arithmetic24990.9160.6570.759 Structural54240.7350.7940.791 Conditional613570.5970.9780.991 Multi-record35100.9160.8220.790 Table 4: Rule Complexity Analysis for Rule Identification. Accuracy Rule Category #Rules #Instances Zero-shot Few-shot Few-shot + CR Arithmetic24990.9220.9100.896 Structural54240.5030.5590.540 Conditional613570.5900.7030.590 Multi-record36100.1590.1840.257 Table 5: Rule Complexity Analysis for Joint Rule Diagnosis. Metrics Rule Category#Rules#InstancesStep-1 (Accuracy)Step-2 (Micro-F1) Arithmetic2790.7970.545 Structural51360.8240.586 Conditional62850.9610.910 Multi-record31440.5760.591 Note: Results are micro-averaged over instances within each rule category and aggregated across all financial statement types on the Model GPT- 4o. Exact Match is not reported at the category level due to multi-rule aggregation. diagnosis setting, where reasoning must be composed across multi- ple records. However, these gains includes a trade off with increased computational cost. Error Type Analysis: Joint Rule Diagnosis.Here, we analyze the fail- ure modes on the joint rule diagnosis task. The results reveal a consis- tent gap between Step-1 violation detection and Step-2 diagnostic lo- calization across all rule categories. While models frequently detect the presence of violations, they fail to identify all violated rules or to localize them correctly, especially for conditional and multi-record constraints. Error-type analysis further shows that partial detection and incorrect localization errors are predominant across statement types. False positives are common for numerical constraints, while false negatives persist for context-dependent rules. These patterns indicate that joint diagnosis failures are driven primarily by incom- plete coverage and misattribution rather than by an inability to detect anomalies. Figure 3 further illustrates the distribution of error types across statement types and prompting protocols for joint rule diagno- sis. Across all settings, partial detection and mixed errors dominate, indicating that models frequently identify some violated rules but fail to produce complete or correctly localized diagnoses, confirming that joint diagnosis failures arise primarily from incomplete cover- age and misattribution rather than from a failure to detect anomalies. These results demonstrate that model failures in FinRule-Bench are structured and rule-dependent. Performance degrades systemat- ically as tasks require reasoning over contextual applicability and cross-section dependencies, confirming that the benchmark isolates fundamental limitations in diagnostic reasoning. BSCFSESI 0 20 40 60 80 100 Error Distribution (%) Statement Types Z = Zero-shot | FS = Few-shot | F+CR = Few-shot+CR ZFSF+CRZFSF+CRZFSF+CRZFSF+CR 18 41 7 20 22 9 19 15 41 4 64 24 10 35 15 9 11 8 26 15 29 30 15 21 9 19 18 31 34 18 13 45 10 22 20 15 7 27 15 25 19 60 23 13 19 10 30 3 45 29 20 42 9 12 17 8 Correct False Negative False Positive Multi-violation Mixed Errors Figure 3: Error distribution across prompting methods. 4.2.3 Performance-Cost Trade Offs. We analyze the trade-off between performance and token consumption of all the prompting strategies across tasks to assess the cost-effectiveness of structured reasoning in FinRule-Bench. Rule Verification.Figure 4 shows that moving from zero-shot to few- shot prompting yields consistent accuracy gains at a modest increase in token usage. However, few-shot with causal-counterfactual rea- soning substantially increases token consumption while providing limited statement-dependent accuracy improvements. For complex statements such as Balance Sheets and Cash Flow Statements, causal- counterfactual prompting can be cost-effective, whereas for simpler statements, it introduces token overhead without significant gains. Rule Identification.In this task, we observe that few-shot prompting improves identification accuracy relative to zero-shot across most statement types, but adding causal-counterfactual reasoning signif- icantly increases token usage while failing to improve, and often degrading, discrimination performance 4. This pattern holds across Balance Sheet, Statement of Equity, and Income Statement evalua- tions, indicating that verbose reasoning is poorly aligned with tasks that require identifying a single violated rule from a given rule set. Joint Rule Diagnosis.Figure 5 highlights a different trade-off. While causal-counterfactual prompting yields only modest gains for Step- 1 violation detection, it provides larger relative improvements for Step-2 exact-match localization, particularly for statements with multiple interacting constraints. However, these gains come at a substantial token cost, with few-shot+CR increasing token usage relative to standard few-shot prompting. These results show that structured causal reasoning is most beneficial when exhaustive multi- rule localization is required. A consistent statement-level pattern also emerges. Causal-counterfactual prompting yields its strongest and most stable gains on Income Statement (SI) evaluations, which exhibit the highest structural complexity among the statement types in FinRule-Bench. SI statements contain dense interdependencies between subtotals, conditional line items, and cross-section aggrega- tions, increasing the likelihood of interacting constraints 2. In this setting, explicit causal exemplars and minimal counterfactual repairs more effectively guide models toward exhaustive localization of violated rules. It clearly shows its effectiveness scales with structural complexity rather than with task difficulty alone. Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. 400500600700800 Tokens per Instance 40 50 60 70 80 90 Accuracy (%) BS CF SE SI (a) Performance vs. Token Consumption Trade-off BSCFSESI Statement Type 0 200 400 600 800 Tokens per Instance 761 760 478 397 785 781 537 458 789 797 545 563 (b) Token Consumption BSCFSESI Statement Type 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 Efficiency (Accuracy % / 100 Tokens) 9.8 5.3 14.3 18.2 10.3 7.9 13.1 17.6 11.2 8.1 11.5 14.2 (c) Cost-Efficiency Analysis Zero-shotFew-shotFew-shot+CR Figure 4: Performance-cost analysis for rule verification. 10002000300040005000 Tokens per Instance 75.0 77.5 80.0 82.5 85.0 87.5 90.0 92.5 Step-1 Accuracy (%) (Anomaly Detection) BS CF SE SI (a) Step-1: YES/NO Detection vs. Token Cost 10002000300040005000 Tokens per Instance 15 20 25 30 35 40 45 Step-2 Exact Match (%) (All Rules Correct) BS CF SE SI (b) Step-2: Multi-Violation Localization vs. Token Cost BSCFSESI Statement Type 0 1000 2000 3000 4000 5000 Tokens per Instance 1137 1133 939 1137 2751 1484 1327 1660 4862 1924 1572 3060 (c) Token Consumption Comparison BSCFSESI Statement Type 0 1 2 3 4 5 6 Composite Efficiency ((Step1+Step2)/2 per 100 Tokens) 4.2 5.8 5.3 5.1 1.7 4.3 4.1 3.3 1.0 3.5 3.5 1.9 (d) Cost-Efficiency: Two-Stage Performance Zero-shotFew-shotFew-shot+CR Figure 5: Performance-cost analysis for the joint rule diagnosis task - Step-1 anomaly detection and Step-2 localization. Overall, our performance efficiency analysis shows that structured prompting is not uniformly cost-effective. Additional reasoning over- head yields limited benefits for verification and discrimination tasks, but can improve on diagnostic localization tasks in joint rule diagno- sis. This underscores the need for task-aware prompting strategies that balance accuracy against computational efficiency. 5 Limitations FinRule-Bench evaluates rule-governed reasoning under the assump- tion that relevant accounting principles are provided and that finan- cial tables are complete and internally coherent. The benchmark does not address rule discovery, reasoning under noisy or missing data, cross-statement reconciliation, or domain knowledge retrieval 1000150020002500300035004000 Tokens per Instance 30 40 50 60 70 80 Accuracy (%) BS CF SE SI (a) Performance vs. Token Consumption Trade-off BSCFSESI Statement Type 0 1000 2000 3000 4000 Tokens per Instance 1241 1230 1015 1123 2165 1346 1096 1467 3752 2698 2314 2026 (b) Token Consumption BSCFSESI Statement Type 0 1 2 3 4 5 6 Efficiency (Accuracy % / 100 Tokens) 5.2 5.9 3.7 2.9 3.1 5.2 5.5 3.3 1.2 2.7 1.4 1.7 (c) Cost-Efficiency Analysis Zero-shotFew-shotFew-shot+CR Figure 6: Performance-cost analysis for rule identification. beyond the supplied rules. It is therefore intended as a focused diag- nosis of joint-reasoning fidelity and not as a general-purpose finan- cial analysis benchmark. We also do not study supervised fine-tuning or enforce chain-of-thought generation. Our goal is to evaluate the behavior of pretrained foundation models under fixed parameters and controlled task conditions, rather than to optimize performance through additional supervision or explanation constraints. While fine-tuning or explicit reasoning traces may improve accuracy, they also introduce confounding factors, optimization choices, and expla- nation verbosity that are outside the scope of our benchmark. 6 Conclusion We introduced FinRule-Bench, a benchmark for evaluating rule- based compliance reasoning over financial tables under explicit accounting principles. Unlike prior financial and tabular benchmarks that emphasize question answering or anomaly detection on syn- thetically perturbed data, FinRule-Bench evaluates enforcement of formal accounting rules on ground-truth financial statements. The benchmark defines three auditing tasks i.e., single-rule verification, violated-rule identification, and joint multi-error diagnosis captur- ing progressively stronger reasoning requirements. Across models and prompting strategies, results show that while large language models reliably handle isolated arithmetic checks, performance de- grades sharply for structural reasoning, conditional applicability, and composition across multiple interacting rules. Joint diagnosis in particular reveals a gap between detecting violations and achieving complete, record-level localization. These failures are systematic and rule-dependent, with performance declining as reasoning de- mands move from surface-level numerical consistency to exhaustive contextual and cross-record enforcement. By grounding evaluation in explicit accounting principles and realistic financial statements, FinRule-Bench provides a principled framework for assessing joint diagnostic reasoning in high-stakes tabular domains and motivates future work toward reliable end-to-end compliance verification. FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA Ethics Statement This work introduces FinRule-Bench, a research benchmark for evaluating the rule-based reasoning and diagnostic capabilities of large language models over structured financial statements. The benchmark is intended solely for research and evaluation purposes, with the goal of analyzing reasoning fidelity, coverage, and failure modes under explicit accounting principles. It is not designed for deployment in real-world auditing, regulatory enforcement, or finan- cial decision-making systems. All financial statements used in this benchmark are sourced from publicly available corporate disclosures, primarily Form 10-K filings released under regulatory requirements. The dataset operates exclusively on aggregated financial statement tables and accounting line items. No personally identifiable infor- mation, individual-level financial records, or confidential data are collected, inferred, or reconstructed. As such, the dataset does not introduce new privacy risks beyond those inherent in the original public filings. To enable controlled evaluation, FinRule-Bench includes deter- ministically generated, rule-aware error-injected variants of other- wise correct financial statements. These synthetic violations are introduced solely to test model reasoning and diagnostic complete- ness and are explicitly labeled as evaluation artifacts. Performance on this benchmark should not be interpreted as evidence that a model is suitable for real-world auditing or compliance verification without qualified human oversight. We acknowledge that the benchmark reflects the institutional and regulatory context of its source data, including accounting standards and reporting practices commonly used in U.S.-centric financial disclosures. Consequently, models eval- uated on FinRule-Bench may exhibit domain- or jurisdiction-specific behavior. These limitations are documented, and future extensions to additional accounting frameworks and jurisdictions are encouraged. Human involvement in dataset construction was limited to curating accounting principles, implementing and validating deterministic rule validators, and verifying task definitions. Large language mod- els were used only in a restricted auxiliary capacity (e.g., assisting with table transcription or manuscript editing) and were not used to generate ground-truth labels. All evaluation labels are produced by deterministic validators to ensure reproducibility and objectivity. Overall, this benchmark aims to support responsible data min- ing and AI research by enabling transparent analysis of reasoning failures in structured financial domains and by discouraging prema- ture or unqualified deployment of language models in high-stakes financial contexts. References [1] Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yi- heng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923 [cs.CV] https://arxiv.org/abs/2502.13923 [2] Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. FinQA: A Dataset of Numerical Reasoning over Financial Data. Proceedings of EMNLP 2021 (2021). [3]Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022.ConvFinQA: Exploring the Chain of Numeri- cal Reasoning in Conversational Finance Question Answering. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process- ing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 6279–6292. doi:10.18653/v1/2022.emnlp-main.421 [4]Manan Roy Choudhury, Adithya Chandramouli, Mannan Anand, and Vivek Gupta. 2026. Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities. arXiv:2511.00340 [cs.AI] https://arxiv.org/abs/2511. 00340 [5] Manan Roy Choudhury, Anirudh Iyengar Kaniyar Narayana Iyengar, Shikhhar Siingh, Sugeeth Puranam, and Vivek Gupta. 2025. TABARD: A Novel Benchmark for Tabular Anomaly Analysis, Reasoning and Detection. In Findings of the Associ- ation for Computational Linguistics: EMNLP 2025, Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng (Eds.). Association for Computational Linguistics, Suzhou, China, 21783–21817. doi:10.18653/v1/2025. findings-emnlp.1189 [6]Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022.Selection- Inference: Exploiting Large Language Models for Interpretable Logical Reasoning. arXiv:2205.09712 [cs.AI] https://arxiv.org/abs/2205.09712 [7]Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2022. TURL: Table Understanding through Representation Learning. SIGMOD Rec. 51, 1 (June 2022), 33–40. doi:10.1145/3542700.3542709 [8]Aaron Grattafiori et al. 2024.The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783 [9] Gheorghe Comanici et al. 2025. Gemini 2.5: Pushing the Frontier with Ad- vanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities. arXiv:2507.06261 [cs.CL] https://arxiv.org/abs/2507.06261 [10]OpenAI et al. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https: //arxiv.org/abs/2410.21276 [11]Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024. Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding – A Survey. arXiv:2402.17944 [cs.CL] https://arxiv.org/abs/ 2402.17944 [12] FASB. n.d.. Financial Accounting Standards Board. https://w.fasb.org. [13] George Fatouros, Kostas Metaxas, John Soldatos, and Manos Karathanassis. 2025. MarketSenseAI 2.0: Enhancing Stock Analysis through LLM Agents. arXiv:2502.00415 [q-fin.CP] https://arxiv.org/abs/2502.00415 [14] Artur d’Avila Garcez and Luís C. Lamb. 2023. Neurosymbolic AI: the 3rd wave. Artif. Intell. Rev. 56, 11 (March 2023), 12387–12406. doi:10.1007/s10462-023- 10448-w [15] Jiayi Gui, Yiming Liu, Jiale Cheng, Xiaotao Gu, Xiao Liu, Hongning Wang, Yuxiao Dong, Jie Tang, and Minlie Huang. 2024. LogicGame: Benchmarking Rule- Based Reasoning Abilities of Large Language Models. arXiv:2408.15778 [cs.AI] https://arxiv.org/abs/2408.15778 [16]Demis Hassabis and Koray Kavukcuoglu. 2024. Introducing Gemini 2.0: Our new AI model for the agentic era. https://blog.google/innovation-and-ai/models-and- research/google-deepmind/google-gemini-ai-update-december-2024/ [17] Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly Supervised Table Parsing via Pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 4320–4333. doi:10.18653/v1/2020.acl-main.398 [18]Aidan Hogan, Eva Blomqvist, Michael Cochez, Claudia D’amato, Gerard De Melo, Claudio Gutierrez, Sabrina Kirrane, José Emilio Labra Gayo, Roberto Nav- igli, Sebastian Neumaier, Axel-Cyrille Ngonga Ngomo, Axel Polleres, Sabbir M. Rashid, Anisa Rula, Lukas Schmelzeisen, Juan Sequeda, Steffen Staab, and An- toine Zimmermann. 2021. Knowledge Graphs. ACM Comput. Surv. 54, 4, Article 71 (July 2021), 37 pages. doi:10.1145/3447772 [19]Giorgos Iacovides, Thanos Konstantinidis, Mingxue Xu, and Danilo Mandic. 2024. FinLlama: LLM-Based Financial Sentiment Analysis for Algorithmic Trading. In Proceedings of the 5th ACM International Conference on AI in Finance (Brooklyn, NY, USA) (ICAIF ’24). Association for Computing Machinery, New York, NY, USA, 134–141. doi:10.1145/3677052.3698696 [20]IASB. n.d.. About the International Accounting Standards Board. https://w. ifrs.org. [21]Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Xin Zhao, and Ji-Rong Wen. 2023. StructGPT: A General Framework for Large Language Model to Reason over Structured Data. In Proceedings of the 2023 Conference on Empirical Meth- ods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 9237–9251. doi:10.18653/v1/2023.emnlp-main.574 [22]Jeonghoon Kim, Heesoo Jung, Hyeju Jang, and Hogun Park. 2024. Improving Multi-hop Logical Reasoning in Knowledge Graphs with Context-Aware Query Representation Learning. In Findings of the Association for Computational Lin- guistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 15978–15991. doi:10.18653/v1/2024.findings-acl.946 [23]Sukanth Korkanti. 2024. Enhancing Financial Fraud Detection Using LLMs and Advanced Data Analytics. In 2024 2nd International Conference on Self Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Sustainable Artificial Intelligence Systems (ICSSAS). 1328–1334. doi:10.1109/ ICSSAS64001.2024.10760895 [24]Andrei Lazarev and Dmitrii Sedov. 2024. Utilizing Modern Large Language Models (LLM) for Financial Trend Analysis and Digest Creation. In 2024 6th In- ternational Conference on Control Systems, Mathematical Modeling, Automation and Energy Efficiency (SUMMA). 317–321. doi:10.1109/SUMMA64428.2024. 10803746 [25]Jean Lee, Nicholas Stevens, and Soyeon Caren Han. 2025. Large Language Models in Finance (FinLLMs). Neural Computing and Applications (2025), 1–15. [26] Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024. Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks. Proc. ACM Manag. Data 2, 3, Article 176 (May 2024), 28 pages. doi:10.1145/3654979 [27]Chenghao Liu, Arunkumar Arulappan, Ranesh Naha, Aniket Mahanti, Joarder Kamruzzaman, and In-Ho Ra. 2024. Large Language Models and Sentiment Analysis in Financial Markets: A Review, Datasets, and Case Study. IEEE Access 12 (2024), 134041–134061. doi:10.1109/ACCESS.2024.3445413 [28]Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, and Yue Zhang. 2023.Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4. arXiv:2304.03439 [cs.CL] https://arxiv.org/abs/2304.03439 [29] Lihui Liu, Zihao Wang, and Hanghang Tong. 2025. Neural-Symbolic Reasoning over Knowledge Graphs: A Survey from a Query Perspective. SIGKDD Explor. Newsl. 27, 1 (July 2025), 124–136. doi:10.1145/3748239.3748249 [30]Shu Liu, Shangqing Zhao, Chenghao Jia, Xinlin Zhuang, Zhaoguang Long, Jie Zhou, Aimin Zhou, Man Lan, and Yang Chong. 2025. FinDABench: Benchmark- ing Financial Data Analysis Ability of Large Language Models. In Proceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 710–725. https://aclanthology.org/2025.coling-main.48/ [31] Xiao-Yang Liu, Guoxuan Wang, Hongyang Yang, and Daochen Zha. 2023. Fin- GPT: Democratizing Internet-scale Data for Financial Large Language Models. arXiv:2307.10485 [cs.CL] https://arxiv.org/abs/2307.10485 [32] Weizheng Lu, Jing Zhang, Ju Fan, Zihao Fu, Yueguo Chen, and Xiaoyong Du. 2025. Large language model for table processing: a survey. Frontiers of Computer Science 19, 2 (Jan. 2025). doi:10.1007/s11704-024-40763-6 [33]Qing Lyu, Shreya Havaldar, Adam Stein, Li Zhang, Delip Rao, Eric Wong, Mar- ianna Apidianaki, and Chris Callison-Burch. 2023. Faithful Chain-of-Thought Reasoning. arXiv:2301.13379 [cs.CL] https://arxiv.org/abs/2301.13379 [34] PCAOB. n.d.. Public Company Accounting Oversight Board. https://pcaobus.org. [35]Jingru Wang, Wen Ding, and Xiaotong Zhu. 2025. Financial Analysis: Intelligent Financial Data Analysis System Based on LLM-RAG. arXiv:2504.06279 [q- fin.ST] https://arxiv.org/abs/2504.06279 [36]Rushi Wang, Jiateng Liu, Weijie Zhao, Shenglan Li, and Denghui Zhang. 2025. AuditBench: A Benchmark for Large Language Models in Financial Statement Auditing. In International Workshop on AI for Transportation. Springer, 59–81. [37]Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021. TUTA: Tree-based Transformers for Generally Structured Table Pre-training. In KDD ’21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, Singapore, August 14-18, 2021, Feida Zhu, Beng Chin Ooi, and Chunyan Miao (Eds.). ACM, 1780–1790. doi:10.1145/3447548.3467434 [38]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’22). Curran Associates Inc., Red Hook, NY, USA, Article 1800, 14 pages. [39]Changyi Xiao and Yixin Cao. 2024. Complex Logical Query Answering by Calibrating Knowledge Graph Completion Models. In Findings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 13792–13803. doi:10.18653/v1/2024.findings-acl.819 [40]Yunhu Ye, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li. 2023. Large Language Models are Versatile Decomposers: Decomposing Evidence and Questions for Table-based Reasoning. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for Computing Machinery, New York, NY, USA, 174–184. doi:10.1145/3539618.3591708 [41] Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data. In Pro- ceedings of the 58th Annual Meeting of the Association for Computational Linguis- tics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa- tion for Computational Linguistics, Online, 8413–8426. doi:10.18653/v1/2020.acl- main.745 [42]Tianshu Zhang, Xiang Yue, Yifei Li, and Huan Sun. 2024. TableLlama: Towards Open Large Generalist Models for Tables. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin Duh, Helena Gomez, and Steven Bethard (Eds.). Association for Computational Linguistics, Mexico City, Mexico, 6024–6044. doi:10.18653/v1/2024.naacl-long.335 [43]Huaqin Zhao, Zheng Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Gengchen Mai, Ninghao Liu, and Tianming Liu. 2024. Revolutionizing Finance with LLMs: An Overview of Applications and Insights. ArXiv abs/2401.11641 (2024). https://api.semanticscholar.org/CorpusID: 267069342 [44]Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021. TAT-QA: A Question An- swering Benchmark on a Hybrid of Tabular and Textual Content in Finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, Online, 3277–3287. doi:10.18653/v1/2021.acl-long.254 FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA A Formal Task Definitions This appendix provides formal definitions of the tasks used in FinRule-Bench. These definitions are included for completeness and reproducibility. The main paper presents task descriptions in evaluation-oriented terms. A.1 Preliminaries Let푇denote a structured financial table consisting of푚records 푡 1 , . . . ,푡 푚 . Let푅= 푟 1 , . . . ,푟 푘 denote the set of accounting rules applicable to푇. Each rule푟 ∈ 푅is associated with a deterministic validator function 푣 푟 (푇) ∈ True,False, which evaluates whether푇satisfies rule푟. Validator outputs are treated as the ground truth for model evaluation. A.2 Rule Verification In the rule verification task, the model is given a financial table푇 and a single accounting rule푟 ∈ 푅. The objective is to determine whether 푇 satisfies the specified rule. Formally, the task is to produce a binary output 푦 ∈ True,False such that 푦= 푣 푟 (푇). This task corresponds to binary classification and evaluates a model’s ability to verify an explicit accounting principle on a struc- tured financial table. A.3 Rule Identification In the rule identification task, the model is given a financial table푇 and a set of accounting rules 푅. Exactly one rule 푟 ★ ∈ 푅 is violated. The objective is to output a rule identifier ˆ 푟 such that ˆ 푟= 푟 ★ , where the constraint | 푟 ∈ 푅 | 푣 푟 (푇)= False | = 1 holds for all records across the statement types in this task. Evaluation is done with exact-match criteria between the pre- dicted rule identifier and the violated rule. A.4 Joint Rule Diagnosis In the joint rule diagnosis task, the model is given a financial table푇 and a set of accounting rules푅. An arbitrary subset of rules may be violated. The task consists of two stages. Violation Detection. The model produces a binary decision 푑 ∈ Yes,No, where 푑= Yes ⇐⇒ ∃푟 ∈ 푅 such that 푣 푟 (푇)= False. Violation Identification. If푑= Yes, the model produces a map- ping 푀 :1, . . . ,푚→ 2 푅 , where푀(푖)denotes the set of accounting rules identified by the model as violated for record 푡 푖 . For a prediction to be correct, both of the following conditions must hold: (1) The violation detection decision 푑 matches the ground truth. (2)The predicted rule sets푀(푖)match the validator-defined vio- lated rules at the record level. This task evaluates diagnostic reasoning under multiple simulta- neous violations and requires both accurate detection and complete localization. B Accounting Rules and Validator Specifications Here, we provide the accounting rules used in FinRule-Bench and the deterministic validators employed to generate ground-truth labels. Validators provide an executable specification of each rule and error- representative algorithms are shown below. B.1 Validator Interface Each accounting rule푟is associated with a deterministic validator function: 푣 푟 : 푇 →True,False, where푇is a structured financial table represented as an ordered collection of records with numeric values and hierarchical metadata. A validator returnsTrueif the table satisfies the rule andFalse otherwise. Ground-truth labels for all tasks are computed by evaluating the relevant validators on the original or error-injected tables. B.2 Representative Rule Categories We group accounting rules into four categories based on the domi- nant reasoning operation required for validation. Arithmetic Rules. Arithmetic rules enforce numeric consistency relations. For example, a Balance Sheet must satisfy the accounting identity: Total Assets= Total Liabilities+ Total Equity. Algorithm 1: Arithmetic Validator (BS01) 1 def validate_bs01(table): 2 푎푠푒푡푠 ← 푡푎푏푙푒["Total Assets"] 3 푙푖푎푏푖푙푖푡푖푒푠 ← 푡푎푏푙푒["Total Liabilities"] 4 푒푞푢푖푡푦 ← 푡푎푏푙푒["Total Equity"] // Check accounting identity 5return 푎푠푒푡푠=푙푖푎푏푖푙푖푡푖푒푠+푒푞푢푖푡푦 Structural Rules. Structural rules constrain the presence, place- ment, or labeling of records. These rules do not require arithmetic computation but depend on correct classification within the table hierarchy. Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Algorithm 2: Structural Validator (SI03) 1 def validate_si03(table): // Non-controlled interests should not in Liabilities 2return "Non-controlled interests" ∈ 푡푎푏푙푒.푙푎푏푒푙푠() Conditional Rules. Conditional rules apply only when specific conditions are met. Validation requires checking both applicability and correctness. Algorithm 3: Conditional Validator (SI05) 1 def validate_si05(table): 2if 푡푎푏푙푒.푐표푛푡푎푖푛푠("Interest Expense") then // Must not be None if present 3return 푡푎푏푙푒.푣푎푙푢푒("Interest Expense")≠ None 4return True Multi-section Rules. Multi-section rules capture dependencies across multiple rows or time periods, such as rollforward or recon- ciliation constraints. Algorithm 4: Multi-section Validator (SE01) 1 def validate_se01(table): 2for 푦푒푎푟 ∈ 푡푎푏푙푒.푦푒푎푟푠() do 3푏푒푔푖푛 ← 푡푎푏푙푒.푣푎푙푢푒(푦푒푎푟, "Beginning Balance") 4푐ℎ푎푛푔푒 ← 푡푎푏푙푒.푣푎푙푢푒(푦푒푎푟, "Net Change") 5푒푛푑 ← 푡푎푏푙푒.푣푎푙푢푒(푦푒푎푟, "Ending Balance") // Rollforward check 6if푏푒푔푖푛+푐ℎ푎푛푔푒≠ 푒푛푑 then 7return False 8return True B.3 Error Injection via Validator Inversion Error-injected instances are generated by applying minimal edits that invert the outcome of a target validator while preserving table structure. For a rule 푟 , we construct a modified table 푇 ′ such that: 푣 푟 (푇)= True and 푣 푟 (푇 ′ )= False. For arithmetic rules, this typically involves perturbing a single numeric value. For structural or conditional rules, this involves delet- ing, renaming, or relocating a single record. For multi-record rules, a single component in the dependency chain is modified. All injections are deterministic and reversible. B.4 Validator Composition for Joint Diagnosis For the joint rule diagnosis task, ground truth is computed by evalu- ating all validators on the same table: 푉(푇)=푟 ∈ 푅 | 푣 푟 (푇)= False. A model prediction is considered correct if and only if it identifies exactly the set푉(푇)and associates each violated rule with the correct record(s). This strict criterion ensures that partial detection and mislocalization are treated as distinct failure modes. B.5 Reproducibility Guarantees Given the same input table, validators always produce identical outputs. This design guarantees that all reported results are fully reproducible and that any performance differences arise solely from model behavior rather than annotation noise or evaluator ambiguity. The accounting rules are provided in the table 22 C Error Injection and Dataset Instantiation This appendix describes the procedures used to create task-specific instances from the base corpus of clean financial statements. All procedures are deterministic, reversible, and implemented in the released artifact set. C.1 Notation and setup Let푇be a clean base table and푅= 푟 1 , . . . ,푟 푘 the rule set for푇 (statement-specific). Let푣 푟 (푇)be the deterministic validator for rule 푟, which for a clean table satisfies푣 푟 (푇)= Truefor all푟 ∈ 푅. An injection targeting rule푟constructs a perturbed table푇 ′ such that 푣 푟 (푇 ′ )= False while preserving structural properties of 푇 . In our dataset: • Balance Sheet (BS): 5 rules. • Cash Flow (CF): 3 rules. • Statement of Equity (SE): 3 rules. • Income Statement (SI): 5 rules. C.2 Single-violation injection For rule verification and rule identification tasks, we generate single- violation instances as follows. The process is deterministic and rule-aware: Algorithm 5: Single-violation injection 1 def inject_single(T, r): // Precondition: 푣 푟 (푇)== True 2Identify the validator checks and the minimal dependent field(s) for rule 푟 3Determine a minimal edit (numeric perturbation, label rename, or row relocation) that will cause 푣 푟 to return False 4Apply the edit to푇 to obtain푇 ′ 5Record the edit script 훿= diff(푇,푇 ′ ) for reversibility 6return (푇 ′ ,훿) C.3 Multi-violation construction (Hard task) For joint diagnosis task, we create multi-violation instances contain- ing multiple simultaneously violated rules. These are constructed deterministically from clean tables as follows: C.4 Dataset instantiation policy Task views are derived deterministically from the same base set of clean tables: FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA Algorithm 6: Multi-violation construction 1 def inject_multi(T, 푟 푎 ,푟 푏 ,...): // Precondition: 푣 푟 (푇)== True for all 푟 2Initialize푇 0 ←푇 3for each target rule 푟 푖 in a predefined sequence do 4Compute (푇 푖 ,훿 푖 ) ← inject_single(푇 푖−1 ,푟 푖 ) 5Verify 푣 푟 푗 (푇 푖 ) for all rules 푟 푗 ; if injection unintentionally flips non-target validators, apply secondary minimal edits to restore only desired flips (keeping determinism) 6Record combined edit scriptΔ=훿 1 ,훿 2 ,... 7return (푇 final ,Δ) •Clean instances: original extracted markdown tables with no edits. •Rule Verification / Rule-violation identification instances: for each rule, generate single-violation instances determin- istically; Also ensure uniqueness of the violated rule per instance. •Joint multi-step diagnosis instances: produce multi-violation records as described above; in our experiments, we produced 100 multi-violation records per statement type. This policy ensures that all task views are comparable, repro- ducible, and derived from the same underlying real-world disclo- sures. D Financial Statement Samples and Structural Characteristics This appendix provides representative examples of the financial statements used in FinRule-Bench. We include both original source documents and their extracted tabular representations to illustrate data provenance, extraction fidelity, and structural complexity. These examples are provided for transparency and are not exhaustive. D.1 Original Financial Statement Sources Figure 7 shows representative financial statements snapshots as reported in corporate filings. These images correspond to the original source documents from which FinRule-Bench tables are extracted. D.2 Extracted Markdown Table Representation Table 6 presents an example of a financial statement after extraction into the Markdown format used as model input. The extracted table preserves section hierarchy, record ordering, labels, and numerical values from the original filing snapshot 7. D.3 Structural Characteristics Financial statements in FinRule-Bench exhibit substantial variation in length, hierarchy, and internal dependencies. Tables range from compact statements with relatively flat structure to complex reports containing multiple sections, nested subsections, subtotal rows, and cross-record constraints. E Dataset Verification In this section, we report the inter-annotator agreement metrics used to assess annotation consistency and reliability. We also describe the annotation guidelines provided to the two annotators, which de- fine the verification criteria and procedures followed during manual review. E.1 Human Verification To assess the reliability of the dataset and the correctness of the controlled rule violations, we conducted a manual human verification on a randomly sampled 15% subset of the perturbed financial tables. This subset was partitioned into two disjoint samples, denoted as훼 1 and훼 2 , each covering 7.5% of the dataset, to independently evaluate annotation consistency. Results are shown in Table 8. Table 8: Inter-annotator agreement results for human verifica- tion SampleCohen’s Kappa (%)Jaccard Coefficient (%) 훼 1 93.397.5 훼 2 95.395.8 Two annotators with expertise in natural language processing and prior experience in the financial domain independently reviewed the sampled tables. Inter-annotator agreement was quantified using Cohen’s kappa to measure consistency in violated-rule identification and Jaccard’s coefficient to measure overlap in record-level violation localization. Agreement was computed separately for훼 1 and훼 2 to assess stability across independent samples. These agreement metrics serve as a validation of the dataset construction process and provide an additional check on the reliability of the deterministic rule validators used for large-scale annotation. E.2 Annotation Guidelines Annotators were provided with a detailed set of written guidelines prior to beginning the manual verification process to ensure consis- tency, clarity, and reproducibility across annotations. The guidelines were designed to standardize how accounting rule violations were interpreted, validated, and localized within structured financial ta- bles. Each annotator was instructed to independently review a perturbed financial table alongside its associated contextual information, in- cluding the stated accounting principle and the intended violation introduced through controlled error injection. Annotators were asked to evaluate whether the injected violation was (i) correctly instan- tiated according to the specified accounting rule, (i) contextually justified given the structure and semantics of the financial statement, and (i) accurately localized to the relevant rows or records in the table. Annotators were explicitly instructed not to infer additional violations beyond those introduced, and to focus solely on verifying the correctness and scope of the provided violation. The guidelines emphasized that annotations should be grounded in the explicit accounting principles supplied with each instance, rather than relying on implicit assumptions, external knowledge, or alternative accounting interpretations. In cases where a viola- tion appeared ambiguous, annotators were instructed to mark their Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Table 6: Example extracted Balance Sheet (MasterCard, December 31, 2023). The table illustrates hierarchical structure, subtotals, cross-section totals, and sign conventions used in FinRule-Bench. Individual line items are omitted for brevity. SectionSubsectionItemValue AssetsCurrent AssetsCash and cash equivalents8,588 . . .. . . Total current assets18,961 Non-Current AssetsGoodwill7,660 . . .. . . Total Assets42,448 LiabilitiesCurrent LiabilitiesAccounts payable834 . . .. . . Total current liabilities16,264 Non-Current LiabilitiesLong-term debt14,344 . . .. . . Total Liabilities35,451 Mezzanine EquityRedeemable Non-controlling Interests22 EquityStockholders EquityClass A treasury stock, at cost-60,429 . . .. . . Total Equity6,975 Total Liabilities and Equity42,448 Table 7: Structural complexity statistics. Statement Type # Records Avg Rows Min Rows Max Rows Avg Sections Frac. w/ Subsections Frac. w/ Subtotals Avg Cols (Min-Max) Balance Sheet (BS)21232.27494.41.001.004.0 (4-4) Cash Flow (CF)31834.47603.41.001.004.0 (4-4) Statement of Equity (SE)28710.112910.01.000.988.5 (5-16) Income Statement (SI)30022.42407.91.000.993.0 (3-3) judgment based on whether the violation could be reasonably sup- ported under the stated rule definition, without attempting to resolve ambiguity through subjective interpretation. Annotators were also provided with clarifying examples illus- trating common rule categories, including arithmetic consistency rules, structural classification rules, conditional applicability rules, and multi-record constraints. These examples were intended to align annotator judgments with the deterministic rule validators used in dataset construction, while preserving independent human assess- ment. Participation in the annotation process was entirely voluntary. Both annotators are NLP researchers with prior experience in financial- domain data and accounting-related tasks, and they chose to con- tribute due to their interest in the research problem and its potential impact on evaluating reasoning reliability in high-stakes financial settings. Annotators were not exposed to any sensitive or personally identifiable information and were not asked to perform tasks beyond verification of synthetic, research-oriented data artifacts. No adjudication or consensus-building was performed during annotation. Disagreements were retained and used exclusively for computing inter-annotator agreement metrics, which serve as an em- pirical measure of annotation consistency rather than as a mechanism for correcting labels. This design ensures that reported agreement re- flects genuine independent judgment rather than post-hoc alignment. FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA F Experimental Results This section presents additional experimental results that comple- ment the main findings in Section 4. Table 9: Performance on Rule Compliance Verification (GPT- 4o) Statement MethodAccPrecision Recall F1-Score Tokens (K) BS Zero-shot0.7490.8240.8900.855967.7 Few-shot0.8120.8820.8940.888998.0 Few-shot + CR 0.8840.9620.8970.9281003.1 CF Zero-shot0.4020.6960.3610.476982.5 Few-shot0.6200.8190.6340.7141009.2 Few-shot + CR 0.6430.8490.6380.7281030.8 SE Zero-shot0.6840.7340.9080.812548.3 Few-shot0.7060.8870.6960.780616.6 Few-shot + CR 0.6260.7530.7460.750616.1 SI Zero-shot0.7240.8340.8350.834713.9 Few-shot0.8060.8820.8850.884825.3 Few-shot + CR 0.7970.8810.8740.8781014.2 Overall Zero-shot0.6400.7720.7490.7443212.4 Few-shot0.7360.8680.7770.8163449.1 Few-shot + CR 0.7380.8610.7890.8213664.2 Table 10: Performance on Rule Compliance Verification (Gem- ini 2.5 Pro) Statement MethodAccPrecision Recall F1-Score Tokens (K) BS Zero-shot0.7390.8350.8570.8451272.0 Few-shot0.6760.8100.7980.8042617.9 Few-shot + CR 0.4730.7680.5260.6243316.7 CF Zero-shot0.6190.8030.6520.7202067.7 Few-shot0.6060.8000.7070.7511595.3 Few-shot + CR 0.6280.8860.5780.7002415.5 SE Zero-shot0.6340.8330.6410.7242114.6 Few-shot0.5850.7950.6020.6852411.0 Few-shot + CR 0.6510.8370.6630.7402820.8 SI Zero-shot0.4590.7570.5170.6152504.3 Few-shot0.6470.8060.7590.7822637.8 Few-shot + CR 0.6580.8330.7380.7833491.9 Overall Zero-shot0.6140.7920.6150.7267958.6 Few-shot0.6290.8030.7170.7569262.0 Few-shot + CR 0.6430.8320.6760.74612044.9 Table 11: Performance on Rule Compliance Verification (Gem- ini 2.0) Statement MethodAccPrecision Recall F1-Score Tokens (K) BS Zero-shot0.7590.8700.8360.8531109.4 Few-shot0.5190.9000.8620.8811142.3 Few-shot + CR 0.2970.8400.7390.7862257.7 CF Zero-shot0.8050.8030.6520.4171045.6 Few-shot0.8580.9190.5400.6801070.7 Few-shot + CR 0.9070.9530.6320.7602522.6 SE Zero-shot0.1050.7050.7140.710639.8 Few-shot0.1500.7170.7190.718761.8 Few-shot + CR 0.1900.8150.7520.782819.9 SI Zero-shot0.1300.7920.6610.720791.2 Few-shot0.2100.8430.8510.847906.0 Few-shot + CR 0.0870.8080.7710.7893769.9 Overall Zero-shot0.4500.7930.7160.6753586.1 Few-shot0.4340.8450.7430.7813880.8 Few-shot + CR 0.3700.8540.7240.7799370.1 Table 12: Performance on Rule Compliance Verification (LLaMA) Statement MethodAccPrecision Recall F1-Score Tokens (K) BS Zero-shot0.6060.8010.7090.7481006.1 Few-shot0.7490.8730.8180.8451036.6 Few-shot + CR 0.7590.9590.7430.8372016.7 CF Zero-shot0.4580.8080.3640.5021027.8 Few-shot0.5670.7650.6090.6781053.9 Few-shot + CR 0.5790.7620.6380.6952435.7 SE Zero-shot0.6280.8060.6630.728584.3 Few-shot0.7450.9150.7270.810653.2 Few-shot + CR 0.6880.8990.6570.760692.1 SI Zero-shot0.5200.7910.5770.667770.2 Few-shot0.7060.8390.8010.819883.3 Few-shot + CR 0.6980.8510.7740.8113376.8 Overall Zero-shot0.5530.8020.5780.6613388.4 Few-shot0.6920.8480.7390.7883627.0 Few-shot + CR 0.6810.8680.7030.7768521.3 Table 13: Performance on Single Violated Rule Identifica- tion (GPT-4o) Statement MethodAcc Bal. Acc M-Prec M-Rec M-F1 W-F1 BS Zero-shot0.6630.6630.7510.663 0.641 0.641 Few-shot0.7240.7240.7500.724 0.702 0.702 Few-shot + CR 0.7240.7240.7500.724 0.702 0.702 CF Zero-shot0.8410.8420.5100.505 0.504 0.843 Few-shot0.8680.8610.5160.516 0.516 0.868 Few-shot + CR 0.7820.7820.6450.563 0.587 0.816 SE Zero-shot0.4080.4090.2940.245 0.200 0.333 Few-shot0.4200.4210.3480.252 0.205 0.341 Few-shot + CR 0.4160.4160.4350.312 0.255 0.340 SI Zero-shot0.3840.3770.5710.377 0.338 0.345 Few-shot0.4140.4090.5350.409 0.368 0.375 Few-shot + CR 0.3890.3890.4460.324 0.300 0.360 Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Table 14: Performance on Single Violated Rule Identifica- tion ( Gemini 2.5 Pro) Statement MethodAcc Bal. Acc M-Prec M-Rec M-F1 W-F1 BS Zero-shot0.4720.6110.7300.611 0.616 0.691 Few-shot0.4280.6740.7080.674 0.652 0.671 Few-shot + CR 0.4910.7290.7590.729 0.666 0.668 CF Zero-shot0.9370.9381.0000.938 0.968 0.968 Few-shot0.9430.9430.9260.943 0.944 0.944 Few-shot + CR 0.6640.7430.5200.743 0.726 0.818 SE Zero-shot0.3400.5490.5900.549 0.552 0.566 Few-shot0.3260.5260.4580.526 0.476 0.488 Few-shot + CR 0.3460.6030.6420.603 0.608 0.612 SI Zero-shot0.3390.3390.1910.339 0.232 0.232 Few-shot0.3420.3450.1830.345 0.226 0.246 Few-shot + CR 0.3330.3720.2450.372 0.298 0.325 Table 15: Performance on Single Violated Rule Identifica- tion (Gemini 2.0) Statement MethodAcc Bal. Acc M-Prec M-Rec M-F1 W-F1 BS Zero-shot0.6580.6590.8190.659 0.620 0.621 Few-shot0.6780.6780.6940.678 0.618 0.618 Few-shot + CR 0.7040.7080.7820.708 0.665 0.664 CF Zero-shot0.8620.8620.8830.862 0.864 0.864 Few-shot0.8610.8610.8900.861 0.860 0.860 Few-shot + CR 0.8730.8730.9120.873 0.887 0.887 SE Zero-shot0.3600.3600.6670.360 0.232 0.232 Few-shot0.3450.3450.5180.345 0.200 0.200 Few-shot + CR 0.3620.3620.6080.362 0.240 0.240 SI Zero-shot0.3780.3780.2130.378 0.259 0.259 Few-shot0.3810.3820.2040.382 0.252 0.275 Few-shot + CR 0.4070.4090.2400.409 0.280 0.310 Table 16: Performance on Single Violated Rule Identifica- tion (LLaMA-3) Statement MethodAcc Bal. Acc M-Prec M-Rec M-F1 W-F1 BS Zero-shot0.5400.7020.7810.702 0.707 0.748 Few-shot0.6170.7630.8040.763 0.768 0.789 Few-shot + CR 0.4750.6440.7330.644 0.658 0.671 CF Zero-shot0.8090.8940.9230.894 0.902 0.904 Few-shot0.6870.8140.8420.814 0.819 0.821 Few-shot + CR 0.6990.8230.8610.823 0.830 0.833 SE Zero-shot0.3310.4970.5410.497 0.503 0.518 Few-shot0.3320.4990.5470.499 0.505 0.520 Few-shot + CR 0.3330.5000.5520.500 0.506 0.522 SI Zero-shot0.3960.5670.6010.567 0.571 0.584 Few-shot0.3910.5620.5960.562 0.566 0.579 Few-shot + CR 0.3770.5470.5830.547 0.552 0.565 Table 17: Performance Comparison of GPT-4o on Fi- nancial Statement Joint Multi-Step Diagnosis Task Method Step 1 ViolationViolation Classification Detection (Yes/No)(Multi-label) Acc Prec RecF1EM Prec 휇 Rec 휇 F1 휇 BS Zero-shot77.0 79.4 96.0 86.91 15.0 49.384.7 62.3 Few-shot76.0 80.1 93.5 86.3 18.0 52.178.3 62.6 Few-shot + CR 78.0 82.3 92.7 87.2 24.0 59.366.1 62.5 CF Zero-shot88.0 90.2 96.0 93.0 43.0 89.472.1 79.8 Few-shot89.0 98.2 88.3 92.9 40.0 91.264.5 75.5 Few-shot + CR 91.0 97.1 91.9 94.4 43.0 95.665.8 77.9 SE Zero-shot79.0 81.4 97.5 88.7 20.0 60.693.2 73.4 Few-shot81.0 83.2 96.0 89.1 29.0 66.585.2 74.7 Few-shot + CR 81.0 81.7 98.1 89.1 30.0 65.686.4 74.5 SI Zero-shot85.0 90.0 91.1 90.5 33.0 69.265.1 67.1 Few-shot81.0 87.4 88.1 87.7 26.0 66.562.1 64.2 Few-shot + CR 83.0 89.6 90.5 90.1 33.0 66.967.0 66.9 Table 18: Performance Comparison of Gemini 2.5 Pro on Financial Statement Joint Multi-Step Diagnosis Task Method Step 1 ViolationViolation Classification Detection (Yes/No)(Multi-label) Acc Prec RecF1 EM Prec 휇 Rec 휇 F1 휇 BS Zero-shot81.0 80.8 100.0 89.4 26.0 63.685.0 72.7 Few-shot83.0 82.5 100.0 90.4 35.0 66.584.4 74.4 Few-shot + CR 81.0 80.8 100.0 89.4 28.0 59.383.8 69.4 CF Zero-shot84.0 83.8 100.0 91.2 57.0 81.590.4 85.7 Few-shot87.0 89.8 95.2 92.4 49.0 89.974.7 81.6 Few-shot + CR 91.0 96.3 92.8 94.5 39.0 96.362.0 75.5 SE Zero-shot86.0 86.8 97.5 91.9 55.0 82.282.7 82.5 Few-shot90.0 90.8 97.5 94.0 62.0 85.285.2 85.2 Few-shot + CR 88.0 88.8 97.5 92.9 56.0 82.682.1 82.4 SI Zero-shot83.0 85.1 94.9 89.7 37.0 65.975.6 70.4 Few-shot81.0 83.9 93.6 88.5 36.0 64.276.9 70.0 Few-shot + CR 83.0 84.3 96.2 89.8 34.0 61.781.4 70.2 Note: All metrics reported as percentages. CR = Counterfactual Reasoning. EM = Exact Match (strict match of all violation codes). Prec 휇 /Rec 휇 /F1 휇 denote micro-averaged metrics across all violation types. FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA Table 19: Performance Comparison of Gemini 2.0 flash on Financial Statement Joint Multi-Step Diagnosis Task Method Step 1 ViolationViolation Classification Detection (Yes/No)(Multi-label) Acc Prec Rec F1 EM Prec 휇 Rec 휇 F1 휇 BS Zero-shot78.0 82.2 92.5 87.1 21.0 59.869.8 63.9 Few-shot81.0 82.1 97.5 89.1 29.0 65.181.9 72.5 Few-shot + CR 77.0 80.6 93.7 86.7 20.0 55.970.6 62.4 CF Zero-shot90.0 93.9 93.9 93.9 40.0 93.863.8 75.9 Few-shot91.0 95.1 93.9 94.5 47.0 94.168.0 79.0 Few-shot + CR 85.0 94.7 86.7 90.5 41.0 94.259.6 73.0 SE Zero-shot68.0 79.5 81.5 80.5 26.0 62.479.5 68.9 Few-shot81.0 82.2 97.5 82.2 34.0 65.993.2 77.2 Few-shot + CR 76.0 81.3 91.3 86.0 30.0 69.272.2 70.7 SI Zero-shot82.0 89.4 87.1 88.3 25.0 69.055.7 61.7 Few-shot83.0 85.8 93.5 89.5 23.0 62.167.3 64.6 Few-shot + CR 81.0 83.9 93.5 88.5 22.0 55.769.2 61.7 Note: All metrics reported as percentages. CR = Counterfac- tual Reasoning. EM = Exact Match (strict match of all viola- tion codes). Prec 휇 /Rec 휇 /F1 휇 denote micro-averaged metrics across all violation types. Table 20: Performance Comparison of LLaMA-3 on Fi- nancial Statement Joint Multi-Step Diagnosis Task Method Step 1 ViolationViolation Classification Detection (Yes/No)(Multi-label) Acc PrecRecF1 EM Prec 휇 Rec 휇 F1 휇 BS Zero-shot80.0 80.0 100.0 88.9 8.045.684.4 59.2 Few-shot84.0 84.897.5 90.7 16.0 45.791.9 61.0 Few-shot + CR 73.0 82.783.8 83.2 14.0 47.948.8 48.3 CF Zero-shot83.0 83.0 100.0 90.7 53.0 79.092.8 85.3 Few-shot81.0 83.595.0 88.9 11.0 45.789.4 60.5 Few-shot + CR 82.0 100.0 78.3 87.8 28.0 76.363.9 69.5 SE Zero-shot34.0 74.228.4 41.1 11.0 58.519.1 28.8 Few-shot55.0 92.948.1 63.4 20.0 62.435.8 45.5 Few-shot + CR 31.0 87.517.3 28.9 17.0 56.58.014.1 SI Zero-shot81.0 88.387.2 87.7 16.0 48.269.2 56.8 Few-shot79.0 87.186.2 87.7 15.0 49.066.0 56.3 Few-shot + CR 79.0 88.084.6 86.3 14.0 58.252.6 55.2 Note: All metrics reported as percentages. CR = Counterfactual Reasoning. EM = Exact Match (strict match of all violation codes). Prec 휇 /Rec 휇 /F1 휇 denote micro-averaged metrics across all violation types. Table 21: Error Type Dist. for Joint Rule Diagnosis under Few-shot + Counterfactual Prompting (GPT-4o). Statement Correct Error Breakdown Step-1 Error Partial Det. Partial Hall. Inc. Loc. BS18%22%10%41%9% CF41%9%35%4%11% SE29%19%10%34%8% SI20%18%22%31%9% Note: Results are computed at the instance level for the joint rule diagnosis task under few-shot + causal-counterfactual prompting. Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. (a) Balance Sheet(b) Statement of Cash Flow (c) Statement of changes in Equity(d) Income Statement Figure 7: Raw Financial Statement screenshot images FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA Table 22: Financial Statement Rules Example CategoryRule IDRule DescriptionFASB Reference FASB Guideline Text I. Balance Sheet Rules BS-01 Total Assets = Total Li- abilities + Total Owner- ship Interests (including parent company common and preferred equity, non- controlling interests, and any equity-classified mez- zanine items). FASB Concepts Statement No. 6 (CON 6) Equity or net assets is the residual interest in the assets of an entity that remains after deducting its liabilities. In a business enterprise, the equity is the ownership interest. BS-02 Assets and liabilities must be classified into cur- rent and non-current cate- gories. 210-10-45- 5/210-10- 05-4 A total of current liabilities shall be presented in clas- sified balance sheets./ The balance sheets of most entities show separate classifications of current as- sets and current liabilities (commonly referred to as classified balance sheets) permitting ready determi- nation of working capital. BS-03 The statement of cash flows must use descriptive terms like "cash" or "cash and cash equivalents". 230-10-45- 4 The statement shall use descriptive terms such as cash or cash and cash equivalents rather than am- biguous terms such as funds. BS-04 Anyappropriationof retained earnings must be presented separately within the shareholders’ equity section. 505-10-45- 3 Appropriation of retained earnings is permitted, pro- vided that it is shown within the shareholders’ equity section of the balance sheet and is clearly identified as an appropriation of retained earnings. BS-05Treasury stock should be reported as a deduction from total shareholders’ equity. 505-10-05- 1 The Equity Topic includes the following Subtopics: a. Overall b. Stock Dividends and Stock Splits c. Treasury Stock d. Subparagraph superseded by Ac- counting Standards Update No. 2018-07. e. Spinoffs and Reverse Spinoffs. I. Income Statement Rules SI-01The results of operat- ing activities must be reported separately from non-operating activities. 255-10-50- 3 An entity shall disclose all of the following informa- tion for each of the five most recent years: a. Net sales and other operating revenues b. Income from continuing operations on a current cost basis c. ...... SI-02 Depreciation expense for the period must be in- cluded in operating ex- penses. 360-10-35- 4 The cost of a productive facility is one of the costs of the services it renders during its useful economic life. Generally accepted accounting principles (GAAP) require that this cost be spread over the expected useful life of the facility in ...... SI-03 Earnings per share (EPS) for both basic and diluted shares must be presented on the face of the income statement. 260-10-45- 2 Entities with simple capital structures, that is, those with only common stock outstanding, shall present basic per-share amounts for income from continuing operations and for net income on the face of the income statement. ...... Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. CategoryRule IDRule DescriptionFASB Reference FASB Guideline Text SI-04On the income statement, a material gain from a transaction outside of pri- mary business operations must be presented as a dis- tinct line item, with the nature of that transaction clearly disclosed. 220-10- S99-2 Material amounts included under miscellaneous other income shall be separately stated in the state- ment of comprehensive income or in a note thereto, indicating clearly the nature of the transactions out of which the items arose. SI-05An event or transaction that is material and ei- ther unusual in nature or infrequent in occurrence must be presented as a separate component of in- come from continuing op- erations, and must not be reported on the face of the income statement net of in- come taxes. 220-20-45- 1 A material event or transaction that an entity con- siders to be of an unusual nature or of a type that indicates infrequency of occurrence or both shall be reported as a separate component of income from continuing operations........ I. Cash Flow Statement Rules CF-01Net Cash Flow from Op- erating Activities + Invest- ing Activities + Financing Activities = Net Change in Cash and Cash Equiv- alents (note: consider the effect of exchange rate changes). 230-10-45- 24 A statement of cash flows for a period shall report net cash provided or used by operating, investing, and financing activities and the net effect of those flows on the total of cash, cash equivalents, ..... CF-02Financial statements shall not report an amount of cash flow per share. 230-10-45- 3 Financial statements shall not report an amount of cash flow per share. Neither cash flow nor any com- ponent of it is an alternative to net income as an indicator of an entity’s performance, as reporting per-share amounts might imply...... CF-03 Clearly show changes in cash, cash equivalents, and restricted cash in the cash- flow statement, use precise terms (not “funds”). 230-10-45- 4 A statement of cash flows shall explain the change during the period in the total of cash, cash equiva- lents, and amounts generally described as restricted cash or restricted cash equivalents. The statement shall use...... IV. Statement of Equity Rules SE-01The Statement of Equity must follow certain rules, such as the following equa- tion regarding the compo- sition of Total Equity: To- tal Equity (year) = Com- mon Stock and Additional Paid-in Capital + Retained Earnings + Accumulated Other Comprehensive In- come + Non-controlling Interest + Treasury Stock. 505-10-50- 2 If both financial position and results of operations are presented, disclosure of changes in the separate accounts comprising shareholders’ equity (in addi- tion to retained earnings) and of the changes in the number of shares of equity securities during at least the most recent annual fiscal period and any subse- quent interim period presented is required to make the financial statements sufficiently informative. Dis- closure of such changes may take the form of sepa- rate statements or may be made in the basic financial statements or notes thereto. FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and PrinciplesConference’17, July 2017, Washington, DC, USA CategoryRule IDRule DescriptionFASB Reference FASB Guideline Text SE-02Net income (or loss) for the period must be shown as an addition (or deduc- tion) to retained earnings. 215-10-50- 1 If both financial position and results of operations are presented, disclosure of changes in the separate accounts comprising shareholders’ equity (in addi- tion to retained earnings) and of the changes in the number of shares of equity securities during at least the most recent annual fiscal period and any subse- quent interim period presented is required to make the financial statements sufficiently informative. Dis- closure of such changes may take the form of sepa- rate statements or may be made in the basic financial statements or notes thereto. SE-03Other comprehensive in- come (OCI) items must be presented as changes to ac- cumulated other compre- hensive income (AOCI). 220-10-45- 14 The total of other comprehensive income for a pe- riod shall be transferred to a component of equity that is presented separately from retained earnings and additional paid-in capital in a statement of finan- cial position at the end of an accounting period. A descriptive title such as accumulated other compre- hensive income shall be used for that component of equity. G Causal-Counterfactual Prompting Details This appendix provides the exact prompt templates and representa- tive screenshots used for causal-counterfactual prompting in FinRule- Bench. These materials are included for reproducibility and trans- parency. The main paper describes the prompting protocol at a high level in Section 3. G.1 Prompt Template Overview Causal-counterfactual prompting is implemented using in-context exemplars that explicitly specify (i) the causal source of a rule vi- olation and (i) a minimal counterfactual modification that would restore compliance. The templates below show the exact prompt structure used across tasks. Each exemplar contains a financial ta- ble, a target decision, a causal explanation identifying the violated principle, and a minimal counterfactual edit that would change the validator outcome. Exemplars are fixed across models and tasks to ensure comparability. G.2 Representative Prompt Representative causal-counterfactual prompt for the Rule Verification task (single violated rule identification). System Instruction: You are a strict financial auditor evaluating the following accounting rule. Rule Definition (BS-01): Total Assets MUST equal (Total Liabilities + Total Equity + Mezzanine Equity). Examples: • COMPLIANT→ T Assets : 42, 448= Liabilities : 35, 451+ Equity : 6, 975+ Mezzanine : 22 CAUSE: The accounting identity holds exactly. COUNTERFACTUAL: If Assets were changed to 42,449, the identity would be violated. • NON-COMPLIANT→ F Assets : 42, 500≠ Liabilities : 35, 451+ Equity : 6, 975+ Mezzanine : 22 CAUSE: Total Assets do not equal the sum of Liabilities, Equity, and Mezzanine. COUNTERFACTUAL: If Assets were changed to 42,448, the identity would hold. Evaluation Task: EVALUATE THE FOLLOWING BALANCE SHEET: BALANCE SHEET DATA: <INSERT MARKDOWN TABLE HERE> Respond with ’T’ if compliant, ’F’ if non-compliant. Answer: Conference’17, July 2017, Washington, DC, USAMalarkkan, A.V., et al. Representative causal-counterfactual prompt for the Rule Identification task (single violated rule identification). System Instruction: You are a strict financial auditor. Task: Determine whether the balance sheet violates ANY accounting rules. If violations exist, report ALL violated rules. Rule Priority: Check BS01 first, then BS02→ BS03→ BS04→ BS05. Rules: • BS01: Total Assets MUST equal (Total Liabilities + Total Equity). • BS02: ALL asset items MUST be classified as "Current_Assets" or "Non_Current_Assets". • BS03: Cash MUST use exact term "Cash and cash equivalents". • BS04: Retained earnings MUST be labeled "Retained earnings". • BS05: Treasury stock (if any) MUST be in Equity as a NEGATIVE value. Reference Examples: • NON-COMPLIANT→ YES: [BS03, BS04] Data Context: Assets | Current_Assets | Funds and funds equivalents | 5000 Equity | Stockholders_Equity | Accumulated profits | 10000 CAUSE: - BS03 violated due to non-GAAP cash terminology ("Funds..."). - BS04 violated because it says "Accumulated profits" instead of "Retained earnings" (non-standard label). COUNTERFACTUAL: - Renaming "Funds" to "Cash and cash equivalents" resolves BS03. - Renaming "Accumulated profits" to "Retained earnings" resolves BS04. • COMPLIANT→ NO Condition: BS01: 18000= 8000+ 10000 (Identity holds). BS02: Current_Assets/Non_Current_Assets present. BS05: No treasury stock. Output Format: YES: [rule1, rule2, ...] NO Representative causal-counterfactual prompt for the multi-rule diagnosis task. System Instruction: You are a strict financial auditor. TASK: Determine whether the balance sheet violates ANY accounting rules. If violations exist, report ALL violated rules. RULE PRIORITY: Check BS01 first, then BS02→ BS03→ BS04→ BS05. RULES: • BS01: Total Assets MUST equal (Total Liabilities + Total Equity). • BS02: ALL asset items MUST be classified as "Current_Assets" or "Non_Current_Assets". • BS03: Cash MUST use exact term "Cash and cash equivalents". • BS04: Retained earnings MUST be labeled "Retained earnings". • BS05: Treasury stock (if any) MUST be in Equity as a NEGATIVE value. REFERENCE EXAMPLES: • Scenario 1: Violations Found Input Data Snippet: Assets | Current_Assets | Funds and funds equivalents | 5000 Equity | Stockholders_Equity | Accumulated profits | 10000 OUTPUT: YES: [BS03, BS04] CAUSE: - BS03 violated due to non-GAAP cash terminology. - Says "Accumulated profits" instead of "Retained earnings" (non-standard label). COUNTERFACTUAL: - Renaming "Funds" to "Cash and cash equivalents" resolves BS03. - Renaming "Accumulated profits" to "Retained earnings" resolves BS04. • Scenario 2: Compliant Condition: BS01: 18000= 8000+ 10000 (Identity holds). BS02: Current_Assets/Non_Current_Assets present. BS05: No treasury stock. OUTPUT: NO OUTPUT FORMAT: YES: [rule1, rule2, ...] NO