Paper deep dive
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
Guy Stephane Waffo Dzuyo, Gaƫl Guibon, Christophe Cerisara, Luis Belmar-Letelier
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/25/2026, 1:31:53 AM
Summary
This paper introduces a robust framework for Financial Statement Fraud Detection (FSFD) using Large Language Models (LLMs) to integrate structured financial data and unstructured textual data (MD&A). It proposes a novel benchmark task, Company-Isolated FSFD (CI-FSFD), to evaluate generalization to unseen companies, addressing the limitations of random data splitting. The authors construct a comprehensive U.S. company dataset linking financial statements, summarized MD&A text, and fraud labels derived from SEC AAERs, demonstrating that their LLM-based approach achieves state-of-the-art performance on the CI-FSFD task.
Entities (14)
Relation Signals (13)
LLMs ā usedfor ā Financial Statement Fraud Detection
confidence 98% Ā· we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information
Company-Isolated FSFD ā proposedby ā Waffo Dzuyo et al.
confidence 95% Ā· To address these limitations and advance towards a more robust FSFD, we propose a novel framework... we introduce a novel task called Company-Isolated FSFD (CI-FSFD)
SEC ā publishes ā AAERs
confidence 95% Ā· fraud labels derived Accounting and Auditing Enforcement Releases (AAERs) from the U.S. Securities and Exchange Commissionās (SEC)
SEC ā publishes ā Form 10-Q
confidence 95% Ā· Quarterly financial data (Forms 10-Q) from 2009-2024 were sourced from the SEC website.
AAERs ā sourceof ā Fraud Labels
confidence 95% Ā· fraud labels derived Accounting and Auditing Enforcement Releases (AAERs) from the U.S. Securities and Exchange Commissionās (SEC)
MD&A ā usedin ā FSFD
confidence 95% Ā· leveraging the rich textual information within financial reports, alongside traditional structured financial data, can improve fraud detection performance
XGBoost ā comparedwith ā LLM-based Approach
confidence 90% Ā· We benchmark our approach against tree-based ensemble models: ... XGBoost
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.
Tags
Links
- Source: https://arxiv.org/abs/2607.19259v1
- Canonical: https://arxiv.org/abs/2607.19259v1
Trouble viewing inline? Open PDF directly ā
Full Text
156,879 characters extracted from source content.
Expand or collapse full text
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks Guy Stephane Waffo Dzuyo 1,2 , Ga Ģ el Guibon 2,3 , Christophe Cerisara 2 and Luis Belmar-Letelier 1 1 Forvis Mazars 2 LORIA, CNRS, Universit Ģ e de Lorraine 3 Universit Ģ e Sorbonne Paris Nord, CNRS, Laboratoire dāInformatique de Paris Nord, LIPN, F-93430 Villetaneuse, France guy.stephane.waffo, luis.belmar-letelier@forvismazars.com, gael.guibon@lipn.fr, christophe.cerisara@loria.fr Abstract Financial statement fraud detection (FSFD) is cru- cial for market integrity but faces challenges from increasingly sophisticated schemes and under- utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring prob- lem with the state of the art, we propose a robust FSFD framework leveraging Large Language Mod- els (LLMs) to integrate both structured financial data and unstructured textual information from fi- nancial reports. We provide a more realistic evalu- ation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a com- prehensive U.S. company dataset combining finan- cial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evalu- ation for reliable financial fraud detection. 1 Introduction The prevalence of financial statement fraud compro- mises the transparency and the integrity of financial markets and results in significant economic losses for stakeholders [ Rezaee, 2005 ] .Traditional deterministic and statistical models [ Altman, 1968; Beneish, 1999; Costa and Soares, 2022 ] are often insufficient against todayās complex and sophisticated fraudulent schemes, driving the need for more advanced detection approaches to identify them. Machine learning techniques, including tree-based models and deep learning, have shown promise in this domain by leveraging structured financial data [ Craja et al., 2020; Ali et al., 2022 ] . However, financial reports contain rich, unstructured textual information, such as Management Discussion and Analysis (MD&A) sections, which often contain qualitative signals and narratives that complement numerical data, and can be indicative of fraudulent intent or misrepresentation [ Kirkos et al., 2024 ] . Recent advance- ments in Large Language Models (LLMs) have demonstrated their capacity to process, understand, and reason over com- plex information across various domains [ Liu et al., 2025; Xu and Ding, 2025 ] . This capability presents a significant opportunity to leverage the textual components of financial reports more effectively [ Wang and Brorsson, 2025 ] for tasks like fraud detection. However, applying LLMs to financial data, especially for fraud detection through anomaly detec- tion, faces unique challenges [ Li et al., 2023 ] . A primary limitation in current Financial Statement Fraud Detection (FSFD) research is the evaluation methodology itself. Many studies simply rely on random data splitting, which can inflate performance metrics by allowing models to learn company-specific patterns or exploit temporal dependencies present in the training data, thus failing to generalize to un- seen companies or future periods [ Wang et al., 2023 ] , which is the main purpose of the FSFD task. This overestimation of predictive capabilities highlights the need for more robust and realistic evaluation frameworks that better reflect real- world deployment scenarios. To address these limitations and advance towards a more robust FSFD, we propose a novel framework grounded in two key hypotheses. Firstly, we hypothesize that (HYP1) generalization on companies in financial statement fraud detection is mandatory, which goes beyond the commonly used random train-test splitting. Indeed, realistic evaluation requires isolating models from specific company identities, which differs from what is currently standard practice.Secondly, we hypothesize that (HYP2) leveraging the rich textual information within financial reports, alongside traditional structured financial data, can improve fraud detection performance, particularly under these more challenging isolation conditions from HYP1. Driven by these hypotheses, we investigate three core research questions: (RQ1) How does evaluating FSFD performance under company isolation impact the modelās performance? (RQ2) To what extent does textual data from financial reports contribute to fraud detection, especially when structured data alone is less informative? (RQ3) Are LLMs capable of effectively detecting financial fraud over multimodal and multiformat financial data for robust fraud detection? arXiv:2607.19259v1 [cs.LG] 21 Jul 2026 In this paper, we contribute as follows: A Novel Task for Realistic FSFD Evaluation. We intro- duce a novel task called Company-Isolated FSFD (CI-FSFD) based on professional expertise. This novel task provides a more realistic assessment of model generalization capabili- ties compared to the current standard practice limited to tra- ditional random splitting. Multimodal Financial Dataset. We construct and make publicly available a comprehensive dataset of U.S. com- panies by integrating structured financial statement data with unstructured textual information from the Management Discussions and Analysis sections (MD&A). Crucially, we link these to fraud labels derived Accounting and Auditing Enforcement Releases (AAERs) from the U.S. Securities and Exchange Commissionās (SEC) through a two-stage, high-fidelity process. We detail our data extraction pipeline, including LLM-based summarization of MD&A and our manually-audited temporal linking of AAERs labels 1 , which ensures the accuracy of our ground-truth labels. We also de- tail our preparation pipeline and temporal linking of AAERs. LLM-based Framework. We propose and implement a novel LLM-based framework capable of effectively pro- cessing and combining structured numerical features and Summarized text data from MD&A (SMD&A), for binary fraud classification. Benchmark-leading Metrics. We demonstrate that our LLM-based approach achieves the top performance on the novel CI-FSFD task, validating our hypotheses and highlight- ing the important role of both textual information and robust evaluation in FSFD. By introducing these challenging evaluation and bench- marks we demonstrate the power of LLMs on multimodal and multiformat financial data. Our work lays a foundation for developing more reliable and generalizable financial fraud detection systems. We believe it will help the community to tackle financial statement analysis 2 . 2 Related Work Financial fraud involves the intentional misrepresentation of financial information to deceive stakeholders or gain an un- fair advantage [ Rezaee, 2005 ] . This illicit activity can man- ifest in various forms, including revenue misstatement, as- set misappropriation, and expense misstatement. The conse- quences of financial fraud are severe, leading to significant financial losses, legal repercussions, and reputational dam- age for companies and investors. While fraud can encom- pass issues such as disclosure violations, breaches of mar- ket regulations, bribery, and earning manipulation, the lat- ter is most likely to be revealed through financial statements. Therefore, in this work, we define fraud as earning manipu- lations and the false reporting of any accounting information 1 https://w.sec.gov/enforcement-litigation/ accounting-auditing-enforcement-releases 2 https://github.com/WaguyMz/Financial-Statements-Fraud -Detection intended to mislead investors, regulators, customers, or other parties [ Healy and Palepu, 2003; Mishkin, 2011 ] . Early Statistical Techniques. Early attempts to detect fi- nancial fraud relied on statistical techniques based on finan- cial ratios. They provide an estimate of the probability of fraud by analyzing various financial ratios and identifying patterns indicative of fraudulent behavior. The Altman Z- Score focuses on bankruptcy prediction [ Altman, 1968 ] , the Beneish M-Score on earnings manipulation [ Beneish, 1999 ] , and the Jones model focuses on accruals [ Costa and Soares, 2022 ] . In 2011, Dechow et al. [ 2011 ] proposed an efficient approach based on a logistic regression model and 7 features to predict earning mistatement. MachineLearningApproaches. Theriseofma- chine learning in the 2000s spurred research into more advanced techniques for FSFD, such as Deep Neu- ral Networks [ Krizhevsky et al., 2012 ] and Random Forests [ Breiman, 2001 ] . Later in the 2010s, the increased accessibility of powerful Natural Language Processing methods like LSTMs [ Hochreiter and Schmidhuber, 1997 ] enabled researchers to leverage textual information within financial reports. For example, Craja et al. [ 2020 ] used the Management Discussion and Analysis (MD&A) sections of Form-10K 3 , where executives explain financial performance, to enhance the performance of their classification model. In 2023, Wang et al. [ 2023 ] tackled financial statement fraud detection with a novel model, RCMA, emphasizing the importance of attentive mechanisms for distinguishing between modalities and coordinate financial ratios with textual data from financial reports. By addressing fusion ambiguity, their approach achieved strong fraud detection performance on CSMARD 4 . The Emergence of Large Language Models (LLMs). LLMs [ Radford et al., 2019; Touvron et al., 2023 ] offer new avenues for complex tasks like FSFD. Initial studies, such as Kirkos et al. [ 2024 ] using ChatGPT-4 on CEO letters and Kim et al. [ 2024 ] on general financial statement analysis, have highlighted LLMsā potential in understanding financial nar- ratives. However, these often rely on closed-source models, posing reproducibility challenges. Bhattacharya and Mick- ovic [ 2024 ] fine-tuned a BERT model using truncated MD&A sections, potentially missing key information. This under- scores the need for LLM-based FSFD approaches that can utilize the extensive textual data in financial reports, a gap our work addresses. Frameworks for Financial Statement Fraud Detection (FSFD). Prevailing FSFD evaluation using random data splitting often yields overoptimistic performance, as mod- els may learn company-specific or time-bound artifacts rather than generalizable fraud indicators, failing to reflect real- world deployment challenges. To address this, we introduce 3 Form 10-K is the comprehensive annual report that public com- panies file with the SEC. Form 10-Q is a quarterly report about the companyās financial performance during the quarter. 4 The China Stock Market & Accounting Research Database (CS- MARD) offers data on the China stock markets and the financial statements of Chinaās listed companies. more realistic evaluation via a novel task: Company-Isolated FSFD (CI-FSFD), evaluating generalization to unseen com- panies. To our knowledge, this is the first work to estab- lish dedicated benchmark for that specific setting, aiming for more reliable assessments of FSFD systems. Open Data and Datasets. A significant challenge in FSFD research is the scarcity of readily available, open datasets. While U.S. AAERs [ U.S. Securities and Exchange Com- mission, 2025 ] provide public fraud instances and Chinaās CSMARD [ CSMAR Database, 2025 ] offers extensive data for Chinese markets, integrating these primary sources with structured financial statements (often in XBRL format 5 ) and textual MD&A sections (both also available from the SEC) requires intensive, non-trivial preprocessing and accurate temporal linking. Curated datasets that perform this integra- tion, such as those available through commercial providers like the Compustat database [ S&P Global Market Intelli- gence, 2025 ] , often come at a significant cost, limiting ac- cessibility for widespread research. We address this gap by constructing and making publicly available a comprehensive FSFD dataset for U.S. companies, which combines financial statements, temporally linked AAERs, and processed MD&A text, and we detail our data collection and preprocessing pipeline. 3 Data Collection and Preprocessing Our FSFD dataset integrates financial data, textual informa- tion from Form 10-Qās MD&A sections, and fraud labels. This involved extracting, cleaning, and structuring these com- ponents from various sources, detailed below. 3.1 Financial Data Quarterly financial data (Forms 10-Q) from 2009-2024 were sourced from the SEC website. Using the US-GAAP tax- onomy, we mapped items to core accounts and, following Waffo Dzuyo et al. [ 2025 ] , processed XBRL data to extract raw metrics (e.g., Total Revenue) and impute missing val- ues. From these, we engineered 122 financial indicators (raw figures, change-based, ratios). For quality, reports with less than 25% of these features present were removed, resulting in 268,936 firm-quarter reports from 13,332 companies. Ap- pendix A details the extraction and list all features. 3.2 Text Data We collected 195,023 quarterly MD&A sections (Forms 10- Q, 2009-2024) via a paid SEC-API 6 . These raw HTML sec- tions were lengthy and variable (1k-150k tokens, avg. 14k), making direct LLM processing computationally challenging. To make this text tractable, we summarized each section us- ing the pretrained and open-source Qwen3 32B [ Yang et al., 2025 ] . This step filters out non-material boilerplate le- gal language. Because forensic accounting anomalies (e.g., transaction misstatements) are embedded within hard factual disclosures rather than subtle linguistic style shifts, utilizing 5 XBRL (eXtensible Business Reporting Language): https:// w.xbrl.org/the-standard/what 6 https://sec-api.io/ Qwen3 32B distills core factual triggers while minimizing context distraction for the classifier. This yielded concise summaries averaging 3,800 tokens, forming our Summarized MD&A dataset, referred as SMD&A. 3.3 Fraud Dataset Preprocessing Sourcing and aligning fraud labels is a critical step in con- structing a robust FSFD dataset. Our fraud labels are de- rived from 3,300 AAERs obtained via the SEC-API. A sig- nificant challenge arises because the machine-readable JSON summaries for these releases lack the specific fiscal years and quarters of the violations, preventing a direct link to our quar- terly financial data. To overcome this, we implemented a two- stage pipeline to guarantee the accuracy of our ground-truth labels. Stage 1: Automated Extraction. First, we scraped the full, detailed legal documents linked within each AAER summary. We then leveraged the long-context capabilities of the Qwen3 32B model as a powerful parsing assistant. Using a structured prompt, we tasked the LLM with identifying and extracting a preliminary set of key information from each document: the fraudulent company or companies involved, a description of the fraudulent scheme, a list of fine-grained fraud categories, based on the 11 earning misstatement types proposed by De- chow et al. [ 2011 ] , which we augmented with an additional Assets misstatement label (details in Appendix C). Stage 2: Manual Audit and Verification. Each of the 249 AAERs processed by the LLM was individually reviewed by one human expert in both Machine Learning and Audit- ing, who cross-referenced the extracted company, fiscal quar- ter(s), and fraud categories against the original legal source documents. This meticulous verification process confirmed that every extracted data point was correct, resulting in per- fect accuracy for our fraud labels. This process yielded a set of 1,451 firm-quarter reports identified as fraudulent between 2000 and 2022. These verified instances form our core binary fraud labels which will serve to further construct the dataset. Additionally, our fraud dataset preprocessing involves ex- tracting 12 fine-grained fraud labels. Although these labels could serve as valuable features for an advanced multi-label classification task, the scope of the current work is limited to binary classification to demonstrate the robustness of the novel CI-FSFD task. 3.4 Final Dataset Construction The final dataset construction involved 3 key steps: merging the datasets, handling class imbalance, and ensuring temporal and company consistency. Merging Datasets. We merged the financial data, text data, and fraud labels. The financial and text data were aligned using company identifiers and fiscal quarters, creating dis- tinct firm-quarter instances. The fraud labels, derived from AAERs, were linked to the financial and text data based on the extracted fiscal quarters. Critical Class Imbalance Handling. Financial fraud is an inherently rare event, leading to extreme class imbalance; in our raw dataset, fraud cases are only about 0.03% of firm- quarter observations. Training directly on such severe imbal- ance biases models towards the majority (non-fraud) class. Following rare-event ML paradigms, we target a stable 5% distribution. This preserves a realistic, severe class imbal- ance while ensuring gradient stability during training. Post- hoc threshold calibration via validation F1-maximization en- sures the model remains optimized for precision under these imbalanced constraints. This initial downsampling is done by preserving original industry and time distributions. Final Dataset Statistics. After merging and downsam- pling, the final dataset consists of 10,159 samples (511 fraud cases and 9,648 non-fraud cases). The distribution of samples across industries and time periods was maintained to ensure generalization and realistic evaluation. 4 Tasks Definition Classic FSFD. In the common setting of binary fraud detection, the dataset usually consists of sets of firm-quarters observations either labelled as fraud or not. The dataset is split randomly into train and test sets, with the goal of predicting whether a given firm-quarter observation is fraudulent or not. CI-FSFD: Company Isolated FSFD. In the classic FSFD, random splitting of the dataset can lead to overfitting, as the model may learn to recognize specific patterns of individual companies. In contrast, our CI-FSFD task requires the model to generalize across different compa- nies, ensuring that it can accurately identify frauds in firms it has never encountered before. This novel task is particu- larly relevant in real-world scenarios where models must be deployed to detect fraud in new companies. 5 Fraud Supervised Classification Input Data. To train and evaluate our models, we explored three feature sets derived from our processed data. First, we used Finan- cial Data Only (FIN), which comprises the 122 engineered financial indicators detailed in Appendix A. These indica- tors cover a range of metrics including raw figures, change- based values, financial ratios, and Beneish M-Score compo- nents. Second, we employed Text Data Only (SMD&A), consisting solely of the summarized quarterly MD&A sec- tions. Finally, we utilized Combined Financial and Text Data (FIN+SMD&A) to leverage information from both sources. For this combined input, we serialized the 122 struc- tured financial indicators into a key-value string (e.g., āTotal Revenue: 123456, Net Income: 7890, ...ā), which was then directly concatenated with the SMD&A text. This straight- forward fusion approach unified both modalities into a single text sequence for the model prompt (shown in Appendix F). Network Design. We propose a Large Language Model (LLM) based frame- work for financial fraud detection. Our primary model em- ploys a pretrained LLM. The input to the LLM is a care- fully constructed prompt that defines the binary fraud detec- tion task. This prompt includes the relevant financial (FIN) Figure 1: Financial Statement Fraud Classification. and/or textual (SMD&A) data for a given firm-quarter, and is structured to elicit a classification response. The LLM is not fine-tuned on the autoregressive language modeling objective (i.e., predicting every subsequent token in the in- put sequence), but rather on predicting the final target to- ken in the sequence, which represents the classification de- cision: either āYESā (indicating fraud) or āNOā (indicating non-fraud). Figure 1 shows an overview of our classification approach. The specifics of the fine-tuning methodology are detailed in section 6. Class Imbalance. Financial fraud is an inherently rare event, leading to highly imbalanced datasets. To mitigate the risk of models becoming biased towards the majority (non-fraud) class, we implement epoch-level undersampling during training. At each train- ing epoch, we dynamically undersample the non-fraud cases from the training set to match the number of fraud instances. This prevents the model from overfitting to the majority class while still utilizing the full diversity of the majority class sam- ples across different epochs. 6 Experiments and Results This section details the experimental setup, baseline models, evaluation metrics, and the results obtained for the different FSFD tasks (Classic FSFD and CI-FSFD). We also present re- sults for zero-shot performance of pretrained LLM on FSFD for further comparison. Data Splitting. For both the Classic FSFD and CI-FSFD tasks, we employ a 5-fold cross-validation strategy on the 10,159 firm-quarter observations. In Classic FSFD setting, folds are created by randomly splitting these observations. For CI-FSFD, com- pany isolation is enforced: all firm-quarter data from a spe- cific company belong exclusively to either the training or test set within a fold. This company-based splitting also main- tains the datasetās original industry sector distribution and approximates a 5% fraud ratio across folds. Detailed infor- mation on these data splitting methodologies is provided in Appendix D. 6.1 Baseline Models We compare our LLM-based approach against several estab- lished and contemporary baselines: Dechow Model. We include the logistic regression model proposed by Dechow et al. [ 2011 ] , hereafter referred to as LR-DECHOW. It is reference logistic regression model, built on 7 financial features and widely recognized bench- mark in prediction of earning misstatements. Its features are calculated from our 122 financial features. Appendix A.8 provides details on these features. Multi-Layer Perceptron (MLP). The MLP serves as a strong baseline for structured financial data. It is trained on the full set of 122 engineered financial indicators (FIN). Hy- perparameters, including the number of layers and neurons per layer, are optimized using a Bayesian optimization ap- proach via Hyperopt [ Bergstra et al., 2013 ] to maximize the average AUC over the 5 validation sets. Tree-Based Ensemble Models. We benchmark our ap- proach against tree-based ensemble models: Random For- est [ Breiman, 2001 ] , LightGBM [ Ke et al., 2017 ] , and XG- Boost [ Chen and Guestrin, 2016 ] . These methods are widely recognized for their robustness and efficacy in classification tasks, including financial fraud detection [ Ashtiani and Raa- hemi, 2022 ] . Random Forest aggregates multiple decision trees to improve stability, while LightGBM and XGBoost, advanced gradient boosting frameworks, are known to offer optimized performance and scalability. We train them all us- ing the 122 engineered financial indicators (FIN), with their respective hyperparameters tuned through Hyperopt as above. RCMA-adapted. We developed and benchmarked an adapted implementation of the Ratio-Chapter-Modality- Aware (RCMA) model [ Wang et al., 2023 ] . This adapta- tion was necessary due to two primary factors: the original modelās text subnetwork relies on legacy methods (Doc2Vec and LSTMs), and its source code and hyperparameters are not publicly available.Our principal modification was to replace those legacy methods by Jina Embedding V2- Small [ Nussbaum et al., 2025 ] , a modern, open-source Sen- tenceBERT model with a long-context architecture [ Reimers and Gurevych, 2019 ] , to align the model with current best practices. We fine-tuned this new component with LoRA [ Hu et al., 2021 ] . For all other hyperparameters, we performed a grid search optimization to find the best configuration. Crucially, the remainder of the RCMA architecture was replicated as faithfully as possible to Wang et al. [ 2023 ] ās de- scription, especially its core modality-aware attention mech- anisms. This ensures that our benchmark is a fair and up-to- date representation of the RCMA design. More details are provided in Appendix E. 6.2 Experimental Setup We fine-tune two foundation large language models: the general-purpose Llama-3.1 8B [ Touvron et al., 2023 ] and the domain-specific Fino1-8B [ Qian et al., 2025 ] . To ensure computational tractability, base models were loaded using 4- bit quantization [ Frantar et al., 2023; Zheng et al., 2024 ] . We utilized Low-Rank Adaptation (LoRA) [ Hu et al., 2021 ] for parameter-efficient fine-tuning, applying adapters to all lin- ear layers. Experiments were run on a single NVIDIA H100 GPU, requiring approximately 4 hours per fold. Hyperparam- eters are detailed in Appendix E. 6.3 Evaluation Methodology Our primary evaluation metric is the ROC AUC score, a widely recognized standard for tasks with significant class imbalance [ Fawcett, 2006 ] . We supplement this with stan- dard classification metrics: precision, recall, and F1-score. Our model selection and calibration process follows a two- stage approach on a validation set (10% of the training data). First, we select the model checkpoint that achieves the highest ROC AUC. Second, using this chosen model, we determine an optimal decision threshold by maximizing the F1-score on the same validation data. This final model is then used to generate predictions on the held-out test set. 6.4 Classic FSFD Results The results for the Classic FSFD task, detailed in Table 1, show high performance across most models. The Llama- 3.1 8B configuration achieved the best AUC of 0.96, and nearly all approaches surpassed an AUC of 0.89, with the LR- DECHOW model being the only exception. However, we argue that this high performance is more in- dicative of a methodological artifact than true generalization capability. The random splitting protocol results in signifi- cant data leakage, where the same companies appear in both training and evaluation sets. Our analysis confirms this is- sue: on average, each of the 321 fraudulent firms is present in 3.35 folds. This setup incites models to memorize company- specific patterns instead of learning robust fraud signals. The strong results reported here and in the literature [ Wang et al., 2023; Li et al., 2016 ] should be interpreted with caution, as they likely reflect this evaluation flaw. This observation re- sponds to our research question (RQ1). 6.5 Company-Isolated FSFD Results The company-isolated evaluation, presented in Table 2, pro- vides a more rigorous test of model generalization by pre- venting data leakage. The dramatic drop in performance for all models validates our hypothesis (HYP1) that the Classic FSFD task is prone to optimistic bias. Against this challenging backdrop, a clear pattern emerges. The Fino1-8B model, when leveraging only narrative ModelInputAUC± stdevF1± stdevPrecision± stdevRecall± stdev LR-DECHOWFIN0.68± 0.02720.15± 0.01350.09± 0.01210.53± 0.1530 MLPFIN0.89± 0.00980.40± 0.08270.30± 0.10800.70± 0.0800 LightGBMFIN0.95± 0.00980.74± 0.03550.84± 0.05410.66± 0.0451 XgBoostFIN0.96± 0.01080.76± 0.09300.84± 0.06040.69± 0.1240 Random ForestFIN0.92± 0.01190.54± 0.02450.52± 0.04430.57± 0.0366 RCMA-adaptedFIN+SMD&A0.89± 0.00810.39± 0.05980.28± 0.07050.71± 0.0572 Fino1 8BFIN0.90± 0.01950.46± 0.10320.38± 0.15640.70± 0.1162 Fino1 8BSMD&A0.95± 0.01780.66± 0.08180.56± 0.10490.83± 0.0584 Fino1-8BFIN+SMD&A0.94± 0.00940.60± 0.09200.50± 0.11690.80± 0.0573 Llama-3.1 8BFIN0.93± 0.01040.44± 0.08190.33± 0.09680.79± 0.0973 Llama-3.1 8BSMD&A0.96± 0.01840.76± 0.08420.71± 0.13540.84± 0.0430 Llama-3.1 8BFIN+SMD&A0.95± 0.01230.71± 0.08240.69± 0.16650.77± 0.0605 Table 1: Performance on the Classic FSFD task over 5 folds with standard deviation (stdev) ModelInputAUCF1PrecisionRecallāAUC (p-value) LR-DECHOWFIN0.67± 0.040.13± 0.020.07± 0.010.68± 0.08-0.074 (p=0.000) MLPFIN0.69± 0.060.14± 0.060.10± 0.040.39± 0.29-0.058 (p=0.000) LightGBMFIN0.68± 0.010.15± 0.030.11± 0.030.25± 0.08-0.085 (p=0.000) XgBoostFIN0.66± 0.040.13± 0.030.08± 0.020.38± 0.18-0.080 (p=0.000) Random ForestFIN0.70± 0.030.16± 0.030.10± 0.020.47± 0.18-0.042 (p=0.000) RCMA-adaptedFIN+SMD&A 0.65± 0.000.14± 0.010.08± 0.010.71± 0.08-0.134 (p=0.000) Fino1 8BFIN0.69± 0.040.14± 0.030.10± 0.030.49± 0.29-0.049 (p=0.000) Fino1 8BSMD&A0.74± 0.030.18± 0.040.16± 0.030.23± 0.08(Reference) Fino1 8BFIN+SMD&A 0.72± 0.010.17± 0.030.12± 0.050.46± 0.18-0.026 (p=0.002) Llama-3.1 8BFIN0.68± 0.040.12± 0.070.15± 0.080.24± 0.18-0.067 (p=0.000) Llama-3.1 8BSMD&A0.68± 0.040.14± 0.010.09± 0.010.43± 0.08-0.066 (p=0.000) Llama-3.1 8BFIN+SMD&A 0.68± 0.040.13± 0.040.08± 0.010.35± 0.20-0.067 (p=0.000) Table 2: Performance on the Company-Isolated FSFD (CI-FSFD) task over 5 folds. Metrics are reported as mean± standard deviation. The final column displays the results of a paired bootstrap test comparing each model against the top performer (Fino1 8B on SMD&A, in bold). This test reports the mean difference in AUC (ā AUC) and the corresponding empirical p-value, calculated from 5,000 bootstrap iterations. SMD&A data, significantly outperforms all other configura- tions, achieving a leading AUC of 0.74 and an F1-score of 0.18. This result also highlights the value of domain spe- cialization, as Fino1-8B consistently surpassed the general- purpose Llama-3.1 8B across all data modalities. Interest- ingly, this text-only model is more effective than the same LLM using financial data (AUC 0.69) or even the combina- tion of both data types (AUC 0.72). The superiority of this approach over the best-performing classical model, Random Forest (AUC 0.70), further highlights the unique advantage of LLMs in this context. 6.6 Discussion of Results Our LLM-based FSFD framework, particularly with sum- marized textual (SMD&A) data, yields significant insights, emphasizing the need for robust evaluation and thus vali- dating our first hypothesis (HYP1). The substantial perfor- mance drop observed in Company-Isolated (CI-FSFD) sce- narios vividly demonstrates how traditional random split- ting inflates real-world generalization estimates. Notably, SMD&A text proved highly valuable, consistently boost- ModelInputAUC± stdevF1± stdev Fino1 8BFIN0.48± 0.070.00 Fino1 8BSMD&A0.52± 0.050.00 Fino1 8BFIN+SMD&A0.51± 0.040.00 Llama-3.1 8BFIN0.49± 0.040.10± 0.01 Llama-3.1 8BSMD&A0.49± 0.040.10± 0.01 Llama-3.1 8BFIN+SMD&A0.52± 0.050.09± 0.00 Qwen3 32BFIN0.47± 0.040.02± 0.01 Qwen3 32BSMD&A0.48± 0.050.00± 0.01 Qwen3 32BFIN+SMD&A0.47± 0.040.02± 0.01 Table 3: Zero-shot FSFD performance (mean over 5 folds, with stan- dard deviation, stdev). All models perform extremely poorly, need- ing finetuning. Threshold for computing F1 was set to 0.5. ing model discrimination (higher AUC). Fino-1ās specialized financial fine-tuning enabled it to outperform the general- purpose Llama-3.1 8B model across all input types in the challenging Company-Isolated Financial Statement Fraud Detection (CI-FSFD) task. Interestingly, the combined in- Figure 2: Detection Performance per Misstatement type on the CI- FSFD task (Fino1 8B with SMD&A input).The average AUC per mistatement is also reported above the bars. put (FIN+SMD&A) underperformed compared to text alone, rejecting HYP2. We intentionally used simple serialization to establish a clean baseline; these results reveal a textual ānoise bottleneck,ā proving that naive concatenation distracts the LLM and highlighting the need for future non-linear cross- modal gating structures. Further, our statistical analysis, us- ing a bootstrap test with 1,000 iterations per fold, confirmed that this model significantly outperforms all others in terms of AUC score, with all observed empirical p-values being zero, except for one (Table 2). Finally, the consistently poor zero- shot LLM performance (AUC 0.50) confirms that LLMsā pre- trained alone are inadequate for fraud detection, highlighting the need for specialized FSFD fine-tuning. Fine-Grained Labels Analysis. To gain a deeper under- standing of our modelās performance on different types of financial misstatements, we conducted a post-training analysis using the 12 fine-grained fraud categories ex- tracted from the AAERs (details in Appendix C). Uti- lizing the predictions from our best-performing model (Fino1 8B with SMD&A input) on the CI-FSFD task, we computed the average AUC for each of these misstate- ment types.Figure 2 shows detection performance by category.āOther Expense/Shareholder Equity Accountā (AUC 0.64) and āRevenueā (AUC 0.71) are the most frequent and relatively well-detected misstatement types. In contrast, āAssets Valuationā recorded the lowest AUC (0.54), indicating it is particularly challenging to consider for fraud detection. 6.7 Explainability In order to explain the LLM classification, we employ At- tnLRP [ Achtibat et al., 2024 ] , a technique that calculates the relevancy of each input token to the LLM predictions using a gradient-perturbation of the input signal. For each sentence of the SMD&A document, we aggregate the relevance scores of all constituent tokens to derive a sentence-level relevance score. This approach allows us to identify and highlight key sentences that significantly influence the modelās decisions as shown in Figure 3. We acknowledge that those scores do not directly elicit explanations, but they can serve as clues, help- ing human experts to understand the modelās prediction. Figure 3: Attn-LRP sentence-level relevancy. Red highlights mean positive contribution to Fraud prediction and blue ones mean nega- tive contribution, with according intensity. 7 Limitations Our study, though it advances financial statement fraud detec- tion (FSFD), has some limitations. First, The low F1-score (0.18) reflects the severe difficulty of cross-company gener- alization without identity leakage. However, this baseline is valuable to help human experts narrow down audit spaces, rather than acting as an automated judge. Second, while CI-FSFD eliminates company identity leak- age, it does not strictly enforce chronological sequencing (e.g., historical-to-future splits). Merging company isolation with explicit rolling time windows is a crucial next trajectory for this benchmark to entirely prevent forward-looking bias. Third, our experiments are confined to the U.S. SEC dataset. To establish global generalizability, future work must extend this empirical evaluation to cross-country and multi- jurisdictional contexts using external databases such as CS- MARD [ CSMAR Database, 2025 ] . Finally, our approach assumes fraud signals reside primar- ily within factual disclosures, which may filter out subtle stylistic or linguistic anomalies. Future work should explore hybrid architectures that ingest original MD&A texts to cap- ture a broader spectrum of behavioral fraud indicators. 8 Conclusion Financial statement fraud detection is essential for market in- tegrity, yet it faces significant challenges due to the sophisti- cation of fraudulent schemes and subpar evaluation methods that often overestimate real-world performance. To address these issues, we introduced the novel Company-Isolated Fi- nancial Statement Fraud Detection (CI-FSFD) task that better evaluates modelsā ability to generalize to unseen companies compared to standard practice. We created and publicly re- leased a comprehensive dataset integrating structured finan- cial data, summarized Management Discussion and Analysis (SMD&A) texts, and fraud labels derived from SEC AAERs. Our experiments showed that fine-tuned LLMs, particularly the specialized Fino-1 8B model using SMD&A data, out- performed other models. These results highlight the critical value of both structured and textual data in fraud detection and underscore the importance of robust evaluation frame- works and the poor zero-shot performance of LLMs empha- sizes the necessity of task-specific fine-tuning. This work provides crucial benchmarks and resources, paving the way for more reliable fraud detection systems and evaluations. References [ Achtibat et al., 2024 ] Reduan Achtibat, Sayed Moham- mad Vakilzadeh Hatefi, Maximilian Dreyer, Aakriti Jain, Thomas Wiegand, Sebastian Lapuschkin, and Wojciech Samek. AttnLRP: Attention-aware layer-wise relevance propagation for transformers. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Pro- ceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learn- ing Research, pages 135ā168. PMLR, 21ā27 Jul 2024. [ Ali et al., 2022 ] Abdulalem Ali,Shukor Abd Razak, Siti Hajar Othman, Taiseer Abdalla Elfadil Eisa, Arafat Al-Dhaqm, Maged Nasser, Tusneem Elhassan, Hashim Elshafie, and Abdu Saif. Financial fraud detection based on machine learning: A systematic literature review. Applied Sciences, 12(19):9637, 2022. [ Altman, 1968 ] Edward I. Altman. Financial ratios, discrim- inant analysis and the prediction of corporate bankruptcy. The Journal of Finance, 23(4):589ā609, 1968. [ Ashtiani and Raahemi, 2022 ] Matin N. Ashtiani and Bijan Raahemi. Intelligent fraud detection in financial state- ments using machine learning and data mining: A sys- tematic literature review. IEEE Access, 10:72504ā72525, 2022. [ Beneish, 1999 ] Messod D. Beneish. The detection of earn- ings manipulation. Financial Analysts Journal, 55(5):24ā 36, 1999. [ Bergstra et al., 2013 ] James Bergstra, Daniel Yamins, and David D Cox. Making a science of model search: Hyper- parameter optimization in hundreds of dimensions for vi- sion architectures. In Proc. of the 30th International Con- ference on Machine Learning (ICML 2013), pages Iā115ā Iā23, June 2013. [ Bhattacharya and Mickovic, 2024 ] IndranilBhattacharya and Ana Mickovic.Accounting fraud detection using contextual language learning. International Journal of Accounting Information Systems, 53:100682, 2024. [ Breiman, 2001 ] Leo Breiman. Random forests. Machine Learning, 45(1):5ā32, 2001. [ Chen and Guestrin, 2016 ] Tianqi Chen and Carlos Guestrin. Xgboost:A scalable tree boosting system.CoRR, abs/1603.02754, 2016. [ Costa and Soares, 2022 ] Cristiano Machado Costa and Jos Ģ e Mauro Madeiros Vel Ė oso Soares. Standard jones and mod- ified jones: An earnings management tutorial. Revista de Administrac ̧ Ģ ao Contempor Ė anea, 26(2):e200305, 2022. [ Craja et al., 2020 ] Patricia Craja, Alisa Kim, and Stefan Lessmann. Deep learning for detecting financial statement fraud. Decision Support Systems, 139:113421, 2020. [ CSMAR Database, 2025 ] CSMAR Database. China stock market & accounting research (csmar) database. https:// w.csmar.com/en/, 2025. Accessed: 2025-05-16. [ Dechow et al., 2011 ] Patricia M. Dechow,Weili Ge, Chad R. Larson, and Richard G. Sloan.Predicting material accounting misstatements*: Predicting material accounting misstatements.Contemporary Accounting Research, 28(1):17ā82, 2011. [ Fawcett, 2006 ] Tom Fawcett. An introduction to roc analy- sis. Pattern Recognition Letters, 27(8):861ā874, 2006. [ Frantar et al., 2023 ] Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023. [ Healy and Palepu, 2003 ] Paul M. Healy and Krishna G. Palepu. The fall of enron. Journal of Economic Perspec- tives, 17(2):3ā26, June 2003. [ Hochreiter and Schmidhuber, 1997 ] Sepp Hochreiter and J Ģ urgen Schmidhuber. Long short-term memory. Neural computation, 9:1735ā80, 12 1997. [ Hu et al., 2021 ] Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language mod- els. CoRR, abs/2106.09685, 2021. [ Ke et al., 2017 ] Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. Lightgbm: A highly efficient gradient boost- ing decision tree. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Sys- tems, volume 30. Curran Associates, Inc., 2017. [ Kim et al., 2024 ] Alex G. Kim, Maximilian Muhn, and Va- leri V. Nikolaev. Financial statement analysis with large language models, 2024. [ Kirkos et al., 2024 ] Efstathios Kirkos, Georgia Boskou, Evrikleia Chatzipetrou, Eleftherios Tiakas, and Charalam- pos Spathis. Exploring the boundaries of financial state- ment fraud detection with large language models. SSRN Electronic Journal, 2024. [ Krizhevsky et al., 2012 ] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural in- formation processing systems, volume 25, 2012. [ Li et al., 2016 ] Bin Li, Julia Yu, Jie Zhang, and Bin Ke. De- tecting accounting frauds in publicly traded u.s. firms: A machine learning approach. In Geoffrey Holmes and Tie- Yan Liu, editors, Asian Conference on Machine Learning, volume 45 of Proceedings of Machine Learning Research, pages 173ā188, Hong Kong, 20ā22 Nov 2016. PMLR. [ Li et al., 2023 ] Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceedings of the Fourth ACM International Confer- ence on AI in Finance, ICAIF ā23, page 374ā382, New York, NY, USA, 2023. Association for Computing Ma- chinery. [ Liu et al., 2025 ] Shu Liu, Shangqing Zhao, Chenghao Jia, Xinlin Zhuang, Zhaoguang Long, Jie Zhou, Aimin Zhou, Man Lan, and Yang Chong. FinDABench: Benchmark- ing financial data analysis ability of large language mod- els. In Owen Rambow, Leo Wanner, Marianna Apidi- anaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert, editors, Proceedings of the 31st International Conference on Computational Linguistics, pages 710ā725, Abu Dhabi, UAE, January 2025. Association for Compu- tational Linguistics. [ Mishkin, 2011 ] Frederic S. Mishkin. Over the cliff: From the subprime to the global financial crisis. Journal of Eco- nomic Perspectives, 25(1):49ā70, March 2011. [ Nussbaum et al., 2025 ] Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. Nomic embed: Training a reproducible long context text embedder, 2025. [ Qian et al., 2025 ] Lingfei Qian, Weipeng Zhou, Yan Wang, Xueqing Peng, Han Yi, Jimin Huang, Qianqian Xie, and Jianyun Nie. Fino1: On the transferability of reasoning enhanced llms to finance, 2025. [ Radford et al., 2019 ] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learn- ers.https://cdn.openai.com/better-language-models/ language modelsareunsupervisedmultitasklearners. pdf, 2019. [ Reimers and Gurevych, 2019 ] Nils Reimers and Iryna Gurevych.Sentence-bert: Sentence embeddings using siamese bert-networks.In Conference on Empirical Methods in Natural Language Processing, 2019. [ Rezaee, 2005 ] Zabihollah Rezaee. Causes, consequences, and deterence of financial statement fraud. Critical Per- spectives on Accounting, 16(3):277ā298, 2005. [ S&P Global Market Intelligence, 2025 ] S&P Global Mar- ket Intelligence.Compustat via WRDS.https://wrds. wharton.upenn.edu/, 2025. Accessed: 2025-05-18. [ Touvron et al., 2023 ] Hugo Touvron, Thibaut Lavril, Gau- tier Izacard, Xavier Martinet, Marie-Anne Lachaux, Tim- oth Ģ e Lacroix, Baptiste Rozi ` ere, Naman Goyal, Eric Ham- bro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. Arxiv 2302.13971, 2023. [ U.S. Securities and Exchange Commission, 2025 ] U.S. Se- curities and Exchange Commission. Accounting and au- diting enforcement releases, 2025. Accessed: 2025-05-16. [ Waffo Dzuyo et al., 2025 ] Guy Stephane Waffo Dzuyo, Ga Ģ el Guibon, Christophe Cerisara, and Luis Belmar- Letelier. Linking industry sectors and financial statements: A hybrid approach for company classification. Proceed- ings of the AAAI Conference on Artificial Intelligence, 39(16):16444ā16452, Apr. 2025. [ Wang and Brorsson, 2025 ] Xinlin Wang and Mats Brors- son. Can large language model analyze financial state- ments well?In Chung-Chi Chen, Antonio Moreno- Sandoval, Jimin Huang, Qianqian Xie, Sophia Anani- adou, and Hsin-Hsi Chen, editors, Proceedings of the Joint Workshop of the 9th Financial Technology and Natural Language Processing (FinNLP), the 6th Financial Nar- rative Processing (FNP), and the 1st Workshop on Large Language Models for Finance and Legal (LLMFinLegal), pages 196ā206, Abu Dhabi, UAE, January 2025. Associa- tion for Computational Linguistics. [ Wang et al., 2023 ] Gang Wang, Jingling Ma, and Gang Chen. Attentive statement fraud detection: Distinguish- ing multimodal financial data with fine-grained attention. Decision Support Systems, 167:113913, 2023. [ Xu and Ding, 2025 ] Ruiyao Xu and Kaize Ding.Large language models for anomaly and out-of-distribution de- tection: A survey. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Com- putational Linguistics: NAACL 2025, pages 5992ā6012, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. [ Yang et al., 2025 ] An Yang, Anfeng Li, Baosong Yang, Be- ichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayi- heng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Hao- ran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jin- gren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report, 2025. [ Zheng et al., 2024 ] Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, Zhangchi Feng, and Yongqiang Ma.Llamafactory: Unified efficient fine- tuning of 100+ language models, 2024. A Appendix A : Details on Financial Features The financial data used in this study are derived from quarterly reports (Forms 10-Q and 10-K) sourced from the SEC, covering the period from 2009 to 2024. The process involved meticulous extraction, imputation, feature engineering, and quality control to construct a robust set of financial indicators for fraud detection. A.1 Data Preparation Overview Raw financial metrics were initially extracted by mapping reported items from company filings to the standardized US-GAAP (Generally Accepted Accounting Principles) taxonomy. This taxonomy provides a hierarchical structure for financial reporting elements. The quarterly financial datas are downloaded freely form the SEC -Website : https://w.sec.gov/data-research/ sec-markets-data/financial-statement-data-sets Taxonomy-based Data Imputation A significant challenge in processing financial statements is handling missing data. The hierarchical nature of the US-GAAP taxonomy was leveraged to impute missing values. For instance, if a parent account (e.g., Total Assets) is reported but some of its constituent child accounts are missing, their values can sometimes be inferred based on the reported parent value and other reported sibling accounts. This imputation helps in creating a more complete financial picture for each report. Figure 4 provides a simplified overview of the US-GAAP taxonomy structure. US GAAP Tree - Balance sheet tags Assets(1) Current Assets(2) Cash (3) Short-Term Investments(4) Non Current Assets(5) Liabilities and Stockholderās equity(6) Stockholderās equity(7)Liabilities(8) Current Liabilities(9)Non Current Liabilities(11) Figure 4: Simplified overview of the US-GAAP taxonomy tree structure. The full taxonomy is extensive and can be explored via the FASB website (https://xbrlview.fasb.org/yeti/resources/yeti-gwt/Yeti.jsp). Only a few top-level balance sheet tags are presented for illustration. Feature Engineering and Quality Control Following imputation, a comprehensive set of 122 financial indicators was engineered. These indicators are designed to capture a wide array of financial signals relevant to fraud detection. To ensure data quality, a cutoff filtering process was applied: reports with excessive missing information, specifically those where less than 25% of the 122 engineered features were present (i.e., non-zero and not NaN), were removed from the final dataset. The 122 engineered features are categorized into five groups as detailed below. Notation and Conventions: In the formulas presented, the subscript t denotes the current fiscal quarter, and tā 1 denotes the previous fiscal quarter. ⢠āX = X t ā X tā1 represents the change in feature X from the previous quarter to the current quarter. ⢠Avg(X) = (X t + X tā1 )/2 represents the average value of feature X over the current and previous quarters. ⢠Values for financial tags that are missing in a report are treated as 0. ⢠Safe Division (safe divide(num, den)): If the denominator den is 0 or NaN, or if the numerator num is NaN, the result is 0. Otherwise, it is num / den. ⢠Safe Summation (safesum(args...)): If any of the arguments args is NaN or 0, the result is 0. Otherwise, it is the sum of the arguments. This specific behavior is adopted for consistency in calculations. A.2 Basic Financial Numbers (43 features) These features are core financial metrics extracted directly from financial statements, standardized according to the US-GAAP taxonomy. The 43 basic financial numbers include: 1. AccountsPayableCurrentAndNoncurrent 2. AccountsReceivableNetCurrent 3. AccountsReceivableNetNoncurrent 4. AccumulatedOtherComprehensiveIncomeLossNetOfTax 5. AdditionalPaidInCapital 6. AmortizationOfIntangibleAssets 7. Assets 8. AssetsCurrent 9. CashCashEquivalentsAndShortTermInvestments 10. CommonStockHeldBySubsidiary 11. CommonStockValue 12. CostOfRevenue 13. DebtCurrent (Short-Term Debt) 14. DeferredTaxAssetsDeferredIncome 15. DeferredTaxLiabilitiesDeferredExpense 16. DeferredTaxLiabilitiesTaxDeferredIncome 17. DepreciationAndAmortization 18. Goodwill 19. GrossProfit 20. IncomeLossFromContinuingOperations 21. IntangibleAssetsNetIncludingGoodwill 22. InterestAndDebtExpense 23. InventoryNet 24. Liabilities (Total Liabilities) 25. LiabilitiesCurrent 26. LongTermDebtCurrent 27. LongTermDebtNoncurrent 28. MinorityInterest 29. NetCashProvidedByUsedInFinancingActivities 30. NetCashProvidedByUsedInInvestingActivities 31. NetCashProvidedByUsedInOperatingActivities 32. NetIncomeLoss 33. OperatingExpenses 34. OperatingIncomeLoss 35. PreferredStockValue 36. PropertyPlantAndEquipmentNet 37. ReceivableFromShareholdersOrAffiliatesForIssuanceOfCapitalStock 38. RetainedEarningsAccumulatedDeficit 39. Revenues 40. SellingGeneralAndAdministrativeExpense 41. TemporaryEquityCarryingAmountIncludingPortionAttributableToNoncontrollingInterests 42. TreasuryStockValue 43. UnearnedESOPShares A.3 Aggregated Measures (9 features) These features are composite values derived by summing related basic financial numbers to represent broader financial concepts. 1. agg ACCOUNTRECEIVABLES: AccountsReceivableNetCurrent t + AccountsReceivableNetNoncurrent t 2. agg LONGTERMDEBT : LongTermDebtCurrent t + LongTermDebtNoncurrent t 3. agg EQUITY : safe sum of: |CommonStockValue t |,|PreferredStockValue t |, AdditionalPaidInCapital t , RetainedEarningsAccumulatedDeficit t , AccumulatedOtherComprehensiveIncomeLossNetOfTax t , āTreasuryStockValue t ,āTemporaryEquityCarryingAmount... t ,āReceivableFromShareholders... t , āMinorityInterest t , UnearnedESOPShares t , CommonStockHeldBySubsidiary t 4. agg TOTALDEBT : DebtCurrent t + aggLONGTERMDEBT t 5. agg DEFTAXEXPENSE: DeferredTaxLiabilitiesDeferredExpense t ā DeferredTaxAssetsDeferredIncome t 6. aggACCRUALS: NetIncomeLoss t ā NetCashProvidedByUsedInOperatingActivities t 7. aggEBIT : Revenues t ā CostOfRevenue t ā OperatingExpenses t 8. aggEBITDA: aggEBIT t + DepreciationAndAmortization t 9. agg NETCASHFLOW: NetCashProvidedByUsedInOperatingActivities t +NetCashProvidedByUsedInFinancingActivities t + NetCashProvidedByUsedInInvestingActivities t A.4 Change-based Measures (16 features) This category includes 16 features designed to capture temporal changes in financial accounts and performance. These features are: 1. diff WCAccruals: āAssetsCurrentā āLiabilitiesCurrentā āCashCashEquivalentsAndShortTermInvestments 2. diffInventories: safedivide(āInventoryNet, Avg(Assets)) 3. diffReceivables: safe divide(āaggACCOUNTRECEIVABLES, Avg(Assets)) 4. diffCashSales: (safedivide(Revenues t , Avg(InventoryNet)) + safedivide(Revenues tā1 , Avg(InventoryNet))) /2 ā āaggACCOUNTRECEIVABLES 5. diffCashMargin: safedivide(safesum(CostOfRevenue t ,āāInventoryNet, āaggACCOUNTRECEIVABLES), diffCashSales t ) 6. diffDefTaxExpense: safedivide(āaggDEFTAXEXPENSE,Assets tā1 ) 7. diffEarnings: safedivide(āNetIncomeLoss, Avg(Assets)) 8. diffAverageAssets: Avg(Assets) 9. diffRevenues: āRevenues 10. diffCash: āCashCashEquivalentsAndShortTermInvestments 11. diffEBIT : āagg EBIT 12. diffEBITDA: āaggEBITDA 13. diff NetCashFlow: āaggNETCASHFLOW 14. diff Depreciation: āDepreciationAndAmortization 15. diffAssets: āAssets 16. diff Equity: āagg EQUITY A.5 Ratio-based Measures (45 features) These features are financial ratios calculated to assess profitability, liquidity, solvency, efficiency, and market valuation. Table 4 lists these ratios and their formulas. Table 4: List of 45 Ratio-based Measures No.Ratio Name (Feature ID)Formula 1ratioGrossProfitMarginsafedivide(GrossProfit t ,Revenues t ) 2ratioOperatingMarginsafedivide(OperatingIncomeLoss t ,Revenues t ) 3ratioNetProfitMarginsafedivide(NetIncomeLoss t ,Revenues t ) 4ratioEBITMarginsafedivide(aggEBIT t ,Revenues t ) 5ratio EBITDAMarginsafedivide(aggEBITDA t ,Revenues t ) 6ratio CashFlowMarginsafedivide(NetCashProvidedByUsedInOperatingActivities t , Revenues t ) 7ratioReturnOnAssetssafedivide(NetIncomeLoss t ,Assets t ) 8ratioReturnOnEquitysafedivide(NetIncomeLoss t , aggEQUITY t ) 9ratioCurrentRatiosafedivide(AssetsCurrent t ,LiabilitiesCurrent t ) 10ratioQuickRatiosafedivide(AssetsCurrent t āInventoryNet t ,LiabilitiesCurrent t ) 11ratioCashRatiosafedivide(CashCashEquivalentsAndShortTermInvestments t , LiabilitiesCurrent t ) 12ratioWorkingCapitalToTotalAssetssafedivide(AssetsCurrent t ā LiabilitiesCurrent t ,Assets t ) 13ratio DebtToAssetsRatiosafedivide(aggTOTALDEBT t ,Assets t ) 14ratioDebtToEquityRatiosafedivide(aggTOTALDEBT t , aggEQUITY t ) 15ratioInterestCoverageRatiosafedivide(OperatingIncomeLoss t ,InterestAndDebtExpense t ) 16ratioTotalLiabilitiesToAssetssafedivide(Liabilities t ,Assets t ) 17ratioAssetTurnoversafedivide(Revenues t , Avg(Assets)) 18ratioFixedAssetTurnoversafedivide(Revenues t , Avg(PropertyPlantAndEquipmentNet)) 19ratioReceivablesTurnoversafedivide(Revenues t , aggACCOUNTRECEIVABLES t ) 20ratio InventoryTurnoversafedivide(CostOfRevenue t , Avg(InventoryNet)) 21ratioSalesTurnoversafedivide(Revenues t , Avg(InventoryNet)) 22ratioEquityMultipliersafedivide(Assets t , aggEQUITY t ) 23ratioSGARatiosafedivide(SellingGeneralAndAdministrativeExpense t ,Revenues t ) 24ratioGoodwilltoAssetssafedivide(Goodwill t ,Assets t ) 25ratio CashFlowToDebtRatiosafedivide(NetCashProvidedByUsedInOperatingActivities t , aggTOTALDEBT t ) 26ratioCashFlowFinancingActivitiessafedivide(NetCashProvidedByUsedInFinancingActivities t , aggNETCASHFLOW t ) 27ratioCashFlowOperatingActivitiessafedivide(NetCashProvidedByUsedInOperatingActivities t , agg NETCASHFLOW t ) 28ratioEquityRatiosafedivide(aggEQUITY t ,Assets t ) 29ratioCashFlowToCurrentLiabilitiessafedivide(NetCashProvidedByUsedInOperatingActivities t , LiabilitiesCurrent t ) Table 4: List of 45 Ratio-based Measures (Continued) No.Ratio Name (Feature ID)Formula 30ratioCashFlowToRevenuesafedivide(NetCashProvidedByUsedInOperatingActivities t , Revenues t ) 31ratioCashFlowCoverageRatiosafedivide(NetCashProvidedByUsedInOperatingActivities t , agg TOTALDEBT t ) 32ratioNetWorkingCapital AssetsCurrent t ā LiabilitiesCurrent t 33ratioLongTermDebtToEquitysafedivide(aggLONGTERMDEBT t , aggEQUITY t ) 34ratioDegreeOfFinancialLeveragesafedivide(Revenues t ,NetIncomeLoss t ) 35ratio InvestedCapitalRatiosafedivide(PropertyPlantAndEquipmentNet t + InventoryNet t ,Assets t ) 36ratio CashToTotalAssetsafedivide(CashCashEquivalentsAndShortTermInvestments t , Assets t ) 37ratio DebtServiceCoveragesafedivide(OperatingIncomeLoss t , aggTOTALDEBT t ) 38ratioFinancialLeverageIndexsafedivide(OperatingIncomeLoss t ,Assets t ) 39ratioTimesInterestEarnedRatiosafedivide(NetIncomeLoss t + InterestAndDebtExpense t ,InterestAndDebtExpense t ) 40ratio CurrentAssetToRevenuessafedivide(AssetsCurrent t ,Revenues t ) 41ratio CurrentLiabilitiesToRevenuessafedivide(LiabilitiesCurrent t ,Revenues t ) 42ratioShortTermDebtToRevenuesafedivide(DebtCurrent t ,Revenues t ) 43ratioIntangibleAssetToRevenuesafedivide(IntangibleAssetsNetIncludingGoodwill t ,Revenues t ) 44ratioLongtermLeveragesafedivide(aggLONGTERMDEBT t ,Assets t ) 45ratioCFFsafedivide(safedivide (NetCashProvidedByUsedInFinancingActivities t , agg NETCASHFLOW t ), Avg(Assets)) A.6 Beneish M-Score Indicators (9 features) These features are components of the Beneish M-Score model, designed to detect earnings manipulation. The individual indicators are: 1. Beneish DSRI (Daysā Sales in Receivables Index): safedivide agg ACCOUNTRECEIVABLES t Revenues t , agg ACCOUNTRECEIVABLES tā1 Revenues tā1 2. BeneishGMI (Gross Margin Index): Let GM t =safedivide(Revenues t ā CostOfRevenue t ,Revenues t ).Then BeneishGMI= safedivide(GM tā1 , GM t ). 3. BeneishAQI (Asset Quality Index): Let N CA t = Assets t ā AssetsCurrent t ā PropertyPlantAndEquipmentNet t . Let AQ t = safedivide(N CA t ,Assets t ). Then BeneishAQI = safedivide(AQ t , AQ tā1 ). 4. Beneish SGI (Sales Growth Index): safedivide(Revenues t ,Revenues tā1 ) 5. BeneishDEPI (Depreciation Index): Let DepRate t = safedivide(DepreciationAndAmortization t ,DepreciationAndAmortization t + PropertyPlantAndEquipmentNet t ). Then BeneishDEPI = safedivide(DepRate tā1 , DepRate t ). 6. Beneish SGAI (SG&A Index): safedivide SellingGeneralAndAdministrativeExpense t Revenues t , SellingGeneralAndAdministrativeExpense tā1 Revenues tā1 7. BeneishACCRUALS (Total Accruals to Total Assets): safedivide(aggACCRUALS t ,Assets t ) 8. Beneish LVGI (Leverage Index): Let Lev t =safe divide(aggLONGTERMDEBT t + DebtCurrent t ,Assets t ).Then BeneishLVGI= safedivide(Lev t , Lev tā1 ). 9. BeneishPROBM (Beneish M-Score): This is the M-Score itself, calculated using the indicators above. The BeneishPROBM (M-Score) is calculated as: M-Score =ā4.84 + 0.920Ć DSRI + 0.528Ć GMI + 0.404Ć AQI + 0.892Ć SGI + 0.155Ć DEPI ā 0.172Ć SGAI + 4.679Ć ACCRUALS val ā 0.327Ć LVGIval Where DSRI, GMI, AQI, SGI, DEPI, SGAI are the values of the correspondingly named Beneish indicators (BeneishDSRI, Beneish GMI, etc.). ACCRUALSval is the value of the BeneishACCRUALS feature, and LVGIval is the value of the BeneishLVGI feature. A.7 Dataset Statistics Descriptive statistics for the number of engineered features per firm-quarter report and reports per company (CIK) are presented in Table 5. The āNumber of Extended Featuresā refers to the count of non-zero values among all potentially derived financial features for a given report prior to final selection for the model, indicating the richness of available data per report. āImportant Tagsā count refers to a predefined subset of raw US-GAAP tags deemed critical. Table 5: Descriptive Statistics for Feature Counts per Firm-Quarter Report FeatureCountMeanMedianMinMaxStd Dev Base numerical numbers (nimportanttags) *** 4320.6021.011.033.04.01 Aggregated Measures (n aggregates)91.351.00.07.01.36 Change-base Features (ndifffeatures) * 1611.2111.08.015.00.77 Ratio-based Measures (nratios)4519.7221.01.039.06.88 Beneish Features (n benishfeatures)92.622.00.09.01.57 All Features (nfeatures)12255.5056209312.35 * Refers to a broader set of calculated differential values tracked during preprocessing, from which the 8 change-based measures listed earlier are a specific subset. *** Count of non-zero values for a predefined subset of 43 important US-GAAP tags. A.8 Dechow Model Features The Dechow model [ Dechow et al., 2011 ] , as implemented in our study, incorporates the following features, based on the implementation from https://github.com/jdonadio/FSFraud: ⢠DECHOW RSSTACCRUALS (RSST Accrual): This measures discretionary accruals based on the Reverse-Salomon- Teoh (RSST) model. It is calculated as: RSST Accruals = āW C + āN CO + āF IN Average Total Assets where: ā āW C is the change in working capital, calculated as: āW C = [āCurrent Assetsā āCash and Short-term Investments]ā [āCurrent Liabilitiesā āShort-term Debt] ā āN CO is the change in net non-current operating assets, calculated as: āN CO = ā[Total AssetsāCurrent AssetsāInvestments]āā[Total LiabilitiesāCurrent LiabilitiesāLong-term Debt] ā āF IN is the change in financing activities, calculated as: āF IN = āShort-term Investmentsā ā[Long-term Debt + Short-term Debt + Preferred Stock] ⢠DECHOW CHREC (Change in Receivables): This is the change in accounts receivable scaled by total assets, calcu- lated as: Change in Receivables = āAccounts Receivable Average Total Assets ⢠DECHOWCHINV (Change in Inventory): This is the change in inventory scaled by total assets, calculated as: Change in Inventory = āInventory Average Total Assets ⢠DECHOWSOFTASSETS (Soft Assets): This is the ratio of intangible assets and goodwill to total assets, indicating the proportion of āsoftā or less tangible assets: Soft Assets = Intangible Assets + Goodwill Total Assets DECHOWCHCASHSALES (Change in Cash Sales): This is the change in cash sales, where cash sales are calculated as sales minus changes in accounts receivable: Change in Cash Sales = Salesā āAccounts Receivable ⢠DECHOWCHROA (Change in Return on Assets): This is the change in Return on Assets, indicating the trend in a companyās profitability relative to its assets: Change in ROA = Earnings t Average Total Assets t ā Earnings tā1 Average Total Assets tā1 ⢠DECHOW ISSUANCE (Issuance): This feature indicates whether the company issued new shares or new debt in the current period. It is a binary variable:1, if equity or debt issuance occurred, 0 otherwise B Appendix B. Details on MD&A Reports Dataset This appendix provides further details on the Management Discussion and Analysis (MD&A) sections used in our study. We elaborate on the characteristics of the raw MD&A data, the summarization process employed, the resultant Synthetically Summarized MD&A (SMD&A) dataset, and the associated costs for this data processing step. B.1 Raw MD&A Sections The raw MD&A sections were extracted using the API https://sec-api.io/. We subscribed for monthly plan which costs $55(https://sec-api.io/pricing and which is enough to download all the available quarterly MD&A sections. Specifically, we query the endpoint https://sec-api.io/docs/sec-filings-item-extraction-api to get the item part1item2 of Form-10Q which refers exactly to the desired MD&A sections. These sections are typically lengthy and contain a mix of textual narratives, financial figures, and sometimes tables, often embedded within HTML structures. Tokensā count Distribution of Raw MD&A The distribution of token counts for the raw MD&A sections is depicted in Figure 5. These sections exhibit considerable variability in length. Figure 5: Distribution of token counts in raw quarterly MD&A sections. The x-axis represents the number of tokens, and the y-axis represents the frequency. Table 6 presents the descriptive statistics for the token counts of the raw MD&A sections in our dataset. Table 6: Descriptive statistics for token counts of raw MD&A sections. StatisticValue Count10,159 Mean14,021 Standard Deviation12,580 Minimum18 25th Percentile7,049 50th Percentile11,370 75th Percentile17,180 Maximum270,346 Sample Raw MD&A Below is an excerpt from a sample raw MD&A section, illustrating its typical structure and content. Note the presence of HTML entities. Sample Raw MD&A Excerpt Item 2. Managementās Discussion and Analysis of Financial Condition and Results of Operations CAUTIONARY STATEMENT RELATING TO THE SAFE HARBOR PROVISIONS OF THE PRIVATE SECURITIES LITIGATION REFORM ACT OF 1995 This Quarterly Report contains forward-looking statements as that term is defined in the federal securities laws. The events described in forward-looking statements contained in this Quarterly Report may not occur. Generally, these statements relate to our business plans or strategies, projected or anticipated benefits or other consequences of our plans or strategies, financing plans, projected or anticipated benefits from acquisitions that we may make, or projections involving anticipated revenues, earnings or other aspects of our operating results or financial position, and the outcome of any contingencies. Any such forward-looking statements are based on current expectations, estimates and projections of management. We intend for these forward-looking statements to be covered by the safe-harbor provisions for forward-looking statements. Words such as ," ," ," ," ," ," ," ," ," and ," and their opposites and similar expressions are intended to identify forward-looking statements. We caution you that these statements are not guarantees of future performance or events and are subject to a number of uncertainties, risks and other influences, many of which are beyond our control that may influence the accuracy of the statements and the projections upon which the statements are based. Factors that could cause actual results to differ materially from those set forth or implied by any forward-looking statement include, but are not limited to, our ability to remain competitive with competitors, risks associated with the generic product industry, dependence on a limited number of suppliers, risks associated with healthcare reform and reductions in reimbursement rates, difficulty in predicting revenue stream and gross profit, industry and market changes, the effect of fluctuations in operating results on the trading price of our common stock, inventory levels, reliance on outside manufacturers, risks of incurring uninsured environmental and other industry specific liabilities, governmental approvals and regulations, risks associated with hazardous materials, potential violations of government regulations, product liability claims, reliance on Chinese suppliers, potential changes to Chinese laws and regulations, potential changes to laws governing our relationships in India , fluctuations in foreign currency exchange rates, tax assessments, changes in tax rules, global economic risks, risk of unsuccessful acquisitions, effect of acquisitions on earnings, indemnification liabilities, terrorist activities, reliance on key executives, litigation risks, volatility of the market price of our common stock, changes to estimates, judgments and assumptions used in preparing financial statements, failure to maintain effective internal controls, compliance with changing regulations , as well as other risks and uncertainties discussed in our reports filed with the Securities and Exchange Commission, including, but not limited to, our Annual Report on Form 10-K for the fiscal year ended June 30, 2011 and other filings. Copies of these filings are available at w.sec.gov. Any one or more of these uncertainties, risks and other influences could materially affect our results of operations and whether forward-looking statements made by us ultimately prove to be accurate. Our actual results, performance and achievements could differ materially from those expressed or implied in these forward-looking statements. We undertake no obligation to publicly update or revise any forward-looking statements, whether from new information, future events or otherwise. NOTE REGARDING DOLLAR AMOUNTS In this quarterly report, all dollar amounts are expressed in thousands, except for per-share amounts. The following Managementās Discussion and Analysis of Financial Condition and Results of Operations (MD&A) is intended to provide the readers of our financial statements with a narrative discussion about our business. The MD&A is provided as a supplement to and should be read in conjunction with our financial statements and the accompanying notes. Executive Summary We are reporting net sales of $212,024 for the six months ended December 31, 2011, which represents a 22.3% increase from the $173,343 reported in the comparable prior period. Gross profit for the six months ended December 31, 2011 was $39,163 and our gross margin was 18.5% as compared to gross profit of $26,410 and gross margin of 15.2% in the comparable prior period. Our selling, general and administrative costs (SG&A) for the six months ended December 31, 2011 increased $6,073 to $27,097 from the amount we reported in the prior period. Our net income increased to $7,621, or $0.29 per diluted share, compared to net income of $1,628, or $0.06 per diluted share in the prior period. Our financial position as of December 31, 2011 remains strong, as we had cash and cash equivalents and short-term investments of $28,700, working capital of $115,838 and shareholdersā equity of $161,571. Our business is separated into three principal segments: Health Sciences, Specialty Chemicals and Agricultural Protection Products. The Health Sciences segment is our largest segment in terms of both sales and gross profits. Products that fall within this segment include pharmaceutical intermediates, APIs, finished dosage form generic drugs and nutraceutical products. We typically partner with both customers and suppliers years in advance of a drug coming off patent to provide the generic equivalent. We believe we have a pipeline of new APIs poised to reach commercial levels over the coming years as the patents on existing drugs expire, both in the United States and in Europe. In addition, we continue to explore opportunities to provide a second-source option for existing generic drugs with approved abbreviated new drug applications (ANDAs). The opportunities that we are looking for are to supply the APIs for the more mature generic drugs where pricing has stabilized following the dramatic decreases in price that these drugs experienced after coming off patent. As is the case in the generic industry, the entrance into the market of other generic competition generally has a negative impact on the pricing of the affected products. By leveraging our worldwide sourcing, quality assurance and regulatory capabilities, we believe we can be an alternative economical, second-source provider of existing APIs to generic drug companies. On December 31, 2010, we acquired certain assets of Rising Pharmaceuticals, Inc. ( ") . We believe that the acquisition of Rising will establish another platform for our growth in our Health Sciences business by the expansion of our finished dosage form product offerings from both foreign and domestic facilities as well as complementing our core strength of sourcing active pharmaceutical ingredients. The addition of Rising provides Aceto with a presence as a developer and marketer of our own brand of generic pharmaceuticals, the Rising brand. According to an IMS Health press release on May 18, 2011, "global spending for medicines will reach nearly $1.1 trillion by 2015, reflecting a slowing compound annual rate of growth of 3 6 percent over the next five years. This compares with 6.2 percent annual growth over the past five years. Lower levels of spending growth for medicines in the U.S., the ongoing impact of patent expiries in developed markets, continuing strong demand in pharmerging markets and policy-driven changes in several countries are among the key factors that will influence future growth, according to IMS Institutes new study, The Global Use Of Medicines Outlook Through 2015". Aceto supplies the raw materials used in the production of nutritional and packaged dietary supplements, including vitamins, amino acids, iron compounds and biochemicals used in pharmaceutical and nutritional preparations. Acetoās identification of a change in the attitudes of Europeans towards nutritional products led to the decision to globalize this business and create an operating company to focus on it, Aceto Health Ingredients GmbH, headquartered in Germany. This globally structured business has become the model for all of our business segments, providing international reach and perspective for our customers. The Specialty Chemicals segment is a supplier to the many different industries that require outstanding performance from chemical raw materials and additives. Specialty Chemicals include a variety of chemicals which make plastics, surface coatings, textiles, fuels and lubricants perform to their designed capabilities. Dye and pigment intermediates are used in the color-producing industries such as textiles, inks, paper, and coatings. Many of our raw materials are also used in high-tech products like high-end electronic parts (circuit boards and computer chips) and binders for specialized rocket fuels. We continue to respond to the changing needs of our customers in the color producing industry by taking our resources and knowledge downstream as a supplier of select organic pigments. In addition, Aceto is a leader in the supply of diazos and couplers to the paper, film and electronics industries. According to a December 15, 2011 Federal Reserve Statistical Release, in the third quarter of calendar year 2011, the index for consumer durables, which impacts the Specialty Chemicals segment, grew at an annual rate of 11.4%. Item 3. Quantitative and Qualitative Disclosures About Market Risk Market risk is the risk of loss arising from adverse changes in market rates and prices, such as interest rates, foreign currency exchange rates and commodity prices. Our primary exposure to market risk is interest rate risk and foreign currency exchange rate risk. We do not use derivative financial instruments for trading or speculative purposes. We do not use any derivative contracts to hedge foreign currency or interest rate exposure. We seek to minimize foreign currency exchange rate risk through manage- ment of our current assets and liabilities which are denominated in foreign currencies. The principal foreign currencies to which we are exposed are the Euro, Indian Rupee and Chinese Yuan. Interest Rate Risk.Our interest expense is sensitive to changes in the general level of interest rates, as sub- stantially all of our borrowings are at variable rates.Our exposure to interest rate risk relates primarily to our Amended and Restated Credit Agreement, as amended (the āCredit Agreementā). As of December 31, 2011, we had 0outstandingunderourrevolvingcreditf acilityand5,000 outstanding under our term loan facility. Each of these facili- ties bears interest at a variable rate based on LIBOR or the Base Rate (as defined in the Credit Agreement). A 100 basis point increase in interest rates would increase our interest expense by $50 annually. Foreign Currency Exchange Rate Risk. We are exposed to foreign currency exchange rate risk related to our purchases and sales denominated in foreign currencies. Our foreign currency exchange rate risk is inherent in the sales and expenses of our foreign subsidiaries, which are denominated in their respective local currencies. The financial statements of our foreign subsidiaries are translated into U.S. dollars at exchange rates in effect at the balance sheet date for assets and liabilities and average exchange rates during the period for revenues and expenses. As a result, changes in exchange rates may affect the reported value of our foreign assets, liabilities, revenues and expenses, and could result in foreign currency translation gains or losses in our consolidated statements of operations. A hypothetical 10and Chinese Yuan would not have a material effect on our results of operations. B.2 MD&A Summarization Process To make the extensive textual data from MD&A sections more manageable for LLM processing while retaining core finan- cial insights, we employed a summarization strategy using Qwen3 -32B model. This model was chosen for its long-context capabilities and efficiency. System Prompt for Summarization The following system prompt was used to guide the Qwen3 -32B model in summarizing the raw MD&A sections. The āquar- ter infoā placeholder was dynamically filled with the specific quarter and year of the report (e.g., āQ4 2023ā). System Prompt for MD&A Summarization You are a highly skilled financial analyst with deep expertise in summarizing corporate disclosures.\\ You will be provided with the āManagementās Discussion and Analysisā (MD\&A) section of a financial report for quarter\_info.\\ Your task is to summarize it following the instructions below: Extract and present the ** distinct, and factual insights ** , along with subjective statements, management commentary, and qualitative explanations, organized into the following sections: --- ** 1. Strategic Priorities and Initiatives ** \\ Summarize key strategies, corporate objectives, growth plans, restructuring efforts, and major initiatives discussed by management.\\ Capture significant strategic shifts, operational transformations, ambitious targets, or business model changes.\\ Highlight subjective language, including optimistic tone, vague descriptions of progress, or assertions lacking clear supporting evidence. ** 2. Operational and Segment Performance ** \\ Summarize operational results and segment-level performance, including production metrics, KPIs, challenges, and improvements.\\ Pay special attention to:\\ Unexplained variances in performance.\\ Misalignment between narrative explanations and operational metrics.\\ Subjective, vague, or generic explanations (e.g., ," dynamics," excellence") without adequate quantification.\\ Unusual operational trends, sales fluctuations, production shifts, or inventory movements. ** 3. Financial Results and Key Trends ** \\ Capture ** all financial metrics ** , including revenue, profitability, margins, cost drivers, liquidity trends, capital structure, and debt along with financial ratios.\\ If the metrics are presented in tables, rewrite them in the section Figures and Tables" instead of here.\\ Also Include commentary on:\\ Revenue recognition patterns or timing shifts.\\ Significant margin changes or cost structure shifts.\\ Increases in accounts receivable, inventory, or other working capital components relative to sales without clear justification.\\ Use of non-recurring items, adjustments, or changes in estimates that materially impact results.\\ Use of non-recurring items, adjustments, or changes in estimates that materially impact results.\\ Subjective rationalizations for financial outcomes (e.g., references to demand" or efficiencies") that lack numeric validation. ** 4. Identified Risks and Uncertainties ** \\ Summarize disclosed risks, including operational, supply chain, regulatory, competitive, legal, and macroeconomic risks.\\ Capture both concrete risks and:\\ Subjective assessments of risk severity.\\ Ambiguous or hedged language (e.g., ," ," ").\\ Shifts in tone, emphasis, or presentation of risks compared to prior periods. ** 5. Forward-Looking Statements and Guidance ** \\ Capture managementās expectations, forecasts, assumptions, and outlook for future periods.\\ Highlight:\\ Changes in guidance or underlying assumptions.\\ Optimistic tone, hedging, or caveats (e.g., ," ," ").\\ Whether forward-looking statements are grounded in quantifiable drivers or rely mainly on qualitative assertions. ** 6. Significant Changes, Events, or Developments ** \\ Summarize material recent or upcoming events affecting the business, such as mergers, acquisitions, divestitures, leadership changes, legal proceedings, regulatory actions, or external shocks.\\ Note how management frames these events|whether impacts are clearly quantified or described with vague or qualitative language. ** 7. Important Figures and Tables ** \\ Extract key figures, tables, or financial data that are critical to understanding the MD\&A.\\ For each table:\\ Recreate the exact table content in clean markdown format preceded by the table title. ** 8. Management Explanations and Justifications ** \\ Capture how management explains or justifies operational and financial results, risks, or variances.\\ Pay attention to:\\ Vague, broad, or overly generic justifications.\\ Repetitive use of boilerplate terms (e.g., conditions," excellence") without specific detail.\\ Narratives that shift accountability to external factors or uncontrollable circumstances without precise quantification. ** 9. Accounting Estimates, Judgments, and Policy Changes ** \\ Summarize any disclosures related to:\\ Changes in accounting policies, methodologies, or estimates.\\ Adjustments to key assumptions (e.g., impairments, allowances, revenue recognition).\\ Areas where significant management judgment materially affects reported results.\\ Note whether explanations are clear, detailed, vague, hedged, or superficial. ** 10. Capital Allocation and Liquidity Management ** \\ Summarize commentary on:\\ Cash management strategies, liquidity preservation, and debt management.\\ Capital expenditures, share repurchases, dividend policies, and financing activities.\\ Highlight any:\\ Indications of liquidity stress.\\ Mismatches between optimistic narratives and defensive liquidity actions (e.g., drawing on credit lines despite claimed strong financial performance). ** 11. Legal, Regulatory, and Compliance Matters ** \\ Summarize discussions related to:\\ Ongoing or pending litigation.\\ Regulatory investigations or changes.\\ Compliance risks, including ESG-related disclosures that have material financial implications.\\ Note whether these issues are presented transparently, minimized, or framed with ambiguous language. --- ** Formatting Instructions: ** \\ Use section headers exactly as written above ie with the numbers and titles.\\ Present each point as a bullet () under the appropriate section.\\ Include both objective data and subjective commentary.\\ Explicitly note subjective explanations, optimistic framing, hedging, or vague descriptions wherever they appear.\\ Be precise and factual but the summary should be detailed\\ Avoid redundancy; each bullet must convey a distinct, meaningful insight. ** Critical Constraint: ** \\ Base the summary ** strictly on the content explicitly stated in the MD\&A. ** \\ Do not include any external knowledge, assumptions, interpretations, or analysis beyond the document provided.\\ Only output the summary | do not include any commentary, explanations, or meta-text. B.3 Summarized MD&A (SMD&A) Sections The summarization process resulted in the SMD&A dataset, consisting of condensed versions of the original MD&A narratives. Tokensā count Distribution of SMD&A Figure 6 illustrates the token count distribution for the SMD&A sections. As intended, these summaries are substantially shorter than the raw MD&A sections. Figure 6: Distribution of token counts in Summarized MD&A (SMD&A) sections. The x-axis represents the number of tokens, and the y-axis represents the frequency. Table 7 provides descriptive statistics for the token counts of the SMD&A sections. Table 7: Descriptive statistics for token counts of SMD&A sections. StatisticValue Count10,159 Mean3,807 Standard Deviation1,250 Minimum1 25th Percentile2,960 50th Percentile3,666 75th Percentile4,476 Maximum7,914 Sample SMD&A An excerpt from a sample SMD&A is provided below. This illustrates the more structured and condensed format achieved through the summarization process. Sample SSMD&A Excerpt ** 1. Strategic Priorities and Initiatives ** - The acquisition of St. Jude Medical, Inc. (St. Jude Medical) was completed on January 4, 2017, to expand Abbottās presence in the cardiovascular and neuromodulation markets. - Abbott is reshaping its business portfolio through divestitures, including the sale of its vision care business (AMO) to Johnson & Johnson for $4.325 billion in cash. - The company is pursuing the acquisition of Alere Inc. (Alere) to expand its global diagnostics presence. The purchase price was reduced from $56.00 to $51.00 per share in April 2017, with the acquisition expected to close by the end of Q3 2017, subject to regulatory approvals. - Abbott is implementing cost improvement initiatives across various functions and businesses, partially offsetting increased expenses from the St. Jude Medical acquisition. - The company is investing in research and development (R&D), with R&D expenses increasing significantly due to the integration of the St. Jude Medical business. - Abbott has a share repurchase program authorized by its board in 2014 for up to $3.0 billion, in addition to $512 million remaining from a prior program. - Abbott increased its quarterly dividend by approximately 2% in 2017 compared to 2016, indicating a focus on shareholder returns. ** 2. Operational and Segment Performance ** - Net sales for the Cardiovascular and Neuromodulation Products segment increased by 198.5% in the first six months of 2017 due to the St. Jude Medical acquisition. Excluding the acquisition and foreign exchange, sales in this segment decreased by 1.5% as lower coronary stent sales and a favorable 2016 royalty agreement resolution were partially offset by higher Structural Heart and endovascular sales. - Sales in the Established Pharmaceutical Products segment increased by 5.5% in the first six months of 2017. Excluding foreign exchange, sales in Key Emerging Markets increased 8.2% in the first half of 2017, driven by growth in Russia, China, and Latin America, partially offset by the impact of a new GST system in India. - Nutritional Products sales decreased slightly by 0.3% in the first six months of 2017, with International Pediatric Nutritionals declining by 8.0%. Challenging conditions in the Chinese infant formula market continued to impact international performance. - U.S. Pediatric Nutritionals increased by 7.7%, driven by momentum from recently launched infant formula products and growth in the PediaSure toddler brand. - International Adult Nutritionals increased by 3.3% compared to the first half of 2016, while U.S. Adult Nutritionals decreased by 4.5% due to competitive and market dynamics. - Diagnostic Products sales increased by 3.7% in the first six months of 2017, with a 5.1% increase excluding foreign exchange, driven by share gains in Core Laboratory and Point of Care markets in the U.S. and higher international sales. - The Other category decreased by 26.0% in the first six months of 2017, reflecting the sale of AMO partially offset by double-digit growth in Abbottās Diabetes Care business. - The decrease in Other Emerging Markets by 6.1% in the first six months of 2017 is attributed to the unfavorable impact of Venezuelan operations. Excluding Venezuela and foreign exchange, sales in Other Emerging Markets increased 4.1%. ** 3. Financial Results and Key Trends ** - For the three months ended June 30, 2017, total net sales were $6,637 million, a 24.4% increase from $5,333 million in the same period in 2016, with a 25.3% increase excluding foreign exchange. - For the six months ended June 30, 2017, total net sales were $12,972 million, a 27.0% increase from $10,218 million in the same period in 2016, with a 27.7% increase excluding foreign exchange. - U.S. sales increased by 42.5% in the second quarter of 2017 and by 47.0% in the first six months of 2017. - International sales increased by 16.3% in the second quarter of 2017 and by 17.9% in the first six months of 2017. - Gross profit margin decreased from 54.4% in the second quarter of 2016 to 46.3% in the second quarter of 2017, and from 53.8% to 45.0% for the first six months of 2017, primarily due to higher intangible amortization and inventory step-up amortization from the St. Jude Medical acquisition. - R&D expenses increased by $165 million in the second quarter of 2017 and by $333 million in the first six months of 2017, driven by the addition of St. Jude Medical. - Selling, general, and administrative (SG&A) expenses increased by 22.7% in the second quarter and 32.6% in the first six months of 2017, primarily due to the St. Jude Medical acquisition and integration costs, partially offset by cost improvement initiatives. - Interest expense (income), net increased by $100 million in the second quarter and $279 million in the first six months of 2017 compared to 2016, due to the $15.1 billion in debt issued in November 2016 to finance the St. Jude Medical acquisition. - Taxes on earnings from continuing operations in the first six months of 2017 included $430 million of tax expense related to the gain on the sale of the AMO business. - Earnings from discontinued operations, net of tax, were $46 million in the first six months of 2017, primarily reflecting net tax benefits from the resolution of tax positions related to AbbVieās operations prior to the 2013 separation. ** 4. Identified Risks and Uncertainties ** - Abbott operates in highly competitive and regulated markets, with ongoing debate over healthcare product availability, delivery, and payment methods, which could adversely affect its operations. - The company faces risks related to the integration of the St. Jude Medical acquisition, including potential challenges in combining operations and achieving expected synergies. - Foreign exchange fluctuations, particularly in Venezuela, pose a risk to financial reporting and cash flows. - Regulatory and legal risks are present, including the FDA warning letter related to the Sylmar, CA manufacturing facility acquired from St. Jude Medical. - The acquisition of Alere is subject to regulatory approvals and potential antitrust concerns, with the FTC and European Commission reviewing the transaction. - Alere is divesting certain businesses in connection with the regulatory review, including the Triage MeterPro and B-type Natriuretic Peptide assay businesses to Quidel Corporation and the subsidiary Epocal Inc. to Siemens Diagnostics Holding I B.V. ** 5. Forward-Looking Statements and Guidance ** - Abbott expects to complete the acquisition of Alere by the end of the third quarter of 2017, subject to customary closing conditions and regulatory approvals, which are now due by September 30, 2017. - The company expects to maintain an investment-grade debt rating. - Abbott expects to fund cash dividends, capital expenditures, and other investments with cash flow from operations, cash on hand, short-term investments, and borrowings. - Abbott expects to use the modified retrospective method to adopt the new revenue recognition standard (ASU 2014-09) and does not expect it to have a material impact on its consolidated financial statements. - The company expects to evaluate the impact of recently issued accounting standards, including ASU 2017-07, ASU 2016-16, and ASU 2016-02, on its consolidated financial statements. ** 6. Significant Changes, Events, or Developments ** - Abbott completed the acquisition of St. Jude Medical on January 4, 2017, for $23.6 billion, including $13.6 billion in cash and $10 billion in Abbott common shares. - Abbott sold its AMO segment to Johnson & Johnson for $4.325 billion in cash, completed on February 27, 2017, and recognized a pre-tax gain of $1.151 billion. - Abbott sold 50 million ordinary shares of Mylan N.V. in the first six months of 2017, generating approximately $1.9 billion in proceeds, reducing its ownership interest from 14% to 3.7%. - Abbott received a warning letter from the FDA in April 2017 regarding its Sylmar, CA manufacturing facility, which is part of the St. Jude Medical acquisition. - Abbott entered into a $2.8 billion term loan agreement in July 2017 to fund the acquisition of Alere. - Abbott commenced a tender offer to purchase its outstanding shares of Aleres Series B Convertible Perpetual Preferred Stock at $402 per share, subject to conditions. ** 7. Important Figures and Tables ** ** Net Sales to External Customers (in millions) ** ** Three Months Ended June 30: ** | Segment | 2017 | 2016 | Total Change | Impact of Foreign Exchange | Total Change Excl. Foreign Exchange | |---------|------|------|---------------|-----------------------------|------------------------------------| | Established Pharmaceutical Products | $1,021 | $981 | 4.1% | 0.6% | 3.5% | | Nutritional Products | $1,731 | $1,740 | (0.6)% | (1.1)% | 0.5% | | Diagnostic Products | $1,273 | $1,226 | 3.8% | (1.6)% | 5.4% | | Cardiovascular and Neuromodulation Products | $2,260 | $758 | 198.2% | (1.1)% | 199.3% | | Other | $352 | $628 | (44.0)% | (1.3)% | (42.7)% | | ** Net Sales ** | ** $6,637 ** | ** $5,333 ** | ** 24.4% ** | ** (0.9)% ** | ** 25.3% ** | | ** Total U.S. ** | ** $2,360 ** | ** $1,655 ** | ** 42.5% ** | ** 42.5% ** | | | ** Total International ** | ** $4,277 ** | ** $3,678 ** | ** 16.3% ** | ** (1.3)% ** | ** 17.6% ** | ** Net Sales to External Customers (in millions) ** ** Six Months Ended June 30: ** | Segment | 2017 | 2016 | Total Change | Impact of Foreign Exchange | Total Change Excl. Foreign Exchange | |---------|------|------|---------------|-----------------------------|------------------------------------| | Established Pharmaceutical Products | $1,971 | $1,868 | 5.5% | 1.0% | 4.5% | | Nutritional Products | $3,373 | $3,411 | (1.1)% | (0.8)% | (0.3)% | | Diagnostic Products | $2,431 | $2,344 | 3.7% | (1.4)% | 5.1% | | Cardiovascular and Neuromodulation Products | $4,363 | $1,467 | 197.4% | (1.1)% | 198.5% | | Other | $834 | $1,128 | (26.0)% | (1.3)% | (24.7)% | | ** Net Sales ** | ** $12,972 ** | ** $10,218 ** | ** 27.0% ** | ** (0.7)% ** | ** 27.7% ** | | ** Total U.S. ** | ** $4,684 ** | ** $3,186 ** | ** 47.0% ** | ** 47.0% ** | | | ** Total International ** | ** $8,288 ** | ** $7,032 ** | ** 17.9% ** | ** (1.0)% ** | ** 18.9% ** | ** Preliminary Allocation of Fair Value of St. Jude Medical Acquisition (in billions): ** | Item | 2017 | |------|------| | Acquired intangible assets, non-deductible | $15.0 | | Goodwill, non-deductible | $15.1 | | Acquired net tangible assets | $3.4 | | Deferred income taxes recorded at acquisition | ($4.6) | | Net debt | ($5.3) | | ** Total preliminary allocation of fair value ** | ** $23.6 ** | ** Assets and Liabilities Held for Disposition (in millions) as of December 31, 2016: ** | Item | 2016 | |------|------| | Trade receivables, net | $176 | | Total inventories | $82 | | Prepaid expenses and other current assets | $266 | | Current assets held for disposition | $524 | | Net property and equipment | $130 | | Intangible assets, net of amortization | $150 | | Goodwill | $1,966 | | Deferred income taxes and other assets | $503 | | Non-current assets held for disposition | $2,753 | | ** Total assets held for disposition ** | ** $3,266 ** | | Trade accounts payable | $145 | | Salaries, wages, commissions and other accrued liabilities | $108 | | Current liabilities held for disposition | $253 | | Post-employment obligations, deferred income taxes and other long-term liabilities | $34 | | ** Total liabilities held for disposition ** | ** $287 ** | ** 8. Management Explanations and Justifications ** - Management attributes the significant increase in Cardiovascular and Neuromodulation Products sales to the St. Jude Medical acquisition, but notes that sales excluding the acquisition and foreign exchange impact decreased by 1.5% due to lower coronary stent sales and the favorable 2016 royalty agreement resolution. - The decrease in International Pediatric Nutritionals is explained as being due to conditions in the Chinese infant formula market," a qualitative explanation without numeric validation. - The increase in U.S. Pediatric Nutritionals is attributed to momentum of several recently launched infant formula products" and of the PediaSure toddler brand," with no specific quantitative details provided. - The increase in Diagnostic Products sales is attributed to gains in the Core Laboratory and Point of Care markets in the U.S. and higher sales to various international markets," a broad explanation without specific market share or sales figures. - The decrease in the Other category is attributed to the of the AMO business," partially offset by -digit growth in Abbottās Diabetes Care business," with no further quantification of the Diabetes Care growth. - The increase in R&D expenses is explained as being due to the of the acquired St. Jude Medical business," a straightforward explanation without further detail. - The increase in SG&A expenses is attributed to the of the acquired St. Jude Medical business as well as the incremental expenses to integrate St. Jude Medical with Abbottās existing vascular business," partially offset by improvement initiatives," which is a general term without specifics. - The decrease in working capital is explained as being due to the of cash to fund the cash portion of the St. Jude Medical acquisition, repayments of debt, pension contributions and dividend payments," with a partial offset from the sale of Mylan shares and business dispositions. ** 9. Accounting Estimates, Judgments, and Policy Changes ** - Abbott received a warning letter from the FDA regarding its Sylmar, CA manufacturing facility, and has prepared a plan for corrective actions, which is progressing. - The preliminary allocation of fair value for the St. Jude Medical acquisition is based on estimates and may be subject to material changes as the valuation is finalized. - Abbott has recognized a $70 million credit to intangible amortization expense in the second quarter of 2017 due to measurement period adjustments to the value of intangibles. - Abbott is evaluating the impact of several new accounting standards, including ASU 2017-07, ASU 2016-16, ASU 2016-02, and ASU 2016-01, which will become effective in 2018 and 2019. - Abbott is currently evaluating the impact of ASU 2014-09 (revenue recognition) and expects to use the modified retrospective method for adoption, but has not yet quantified the potential impact. ** 10. Capital Allocation and Liquidity Management ** - Abbott reduced its cash and cash equivalents from $18.6 billion at December 31, 2016, to $9.7 billion at June 30, 2017, primarily due to the St. Jude Medical acquisition, debt repayments, pension contributions, and dividends. - Net cash from operating activities increased by $1.109 billion in the first six months of 2017 compared to 2016, due to the favorable impact of the St. Jude Medical acquisition and reduced pension contributions. - Abbott has $5.0 billion in unused lines of credit available, which expire in 2019. - Abbott entered into a $2.8 billion term loan agreement in July 2017 to fund the Alere acquisition. - Abbott is maintaining an investment-grade debt rating and has a long-term debt rating of B by Standard & Poorās and Baa3 by Moodyās. - Abbott is utilizing cash flow from operations, cash on hand, short-term investments, and borrowings to fund dividends, capital expenditures, and other business investments. ** 11. Legal, Regulatory, and Compliance Matters ** - Abbott received a warning letter from the FDA in April 2017 related to its Sylmar, CA manufacturing facility, which was acquired as part of the St. Jude Medical acquisition. - The FDA inspection findings have not yet resulted in a material impact on financial results, and Abbott is implementing a corrective action plan. - Abbott is subject to regulatory scrutiny in multiple jurisdictions, and tax authorities in various jurisdictions regularly review its income tax filings. - The company expects the recorded amount of gross unrecognized tax benefits to decrease by $200 million to $350 million, including cash adjustments, within the next twelve months as a result of concluding various domestic and international tax matters. - Abbottās U.S. federal income tax returns are settled through 2013, and St. Jude Medicalās federal income tax returns are settled through 2013 except for one item. C Appendix C. Details on AAER Dataset and Preprocessing This appendix details the acquisition, preprocessing, and filtering pipeline for the Accounting and Auditing Enforcement Re- leases (AAERs) used to generate fraud labels for our study. C.1 Raw AAER Data Acquisition and Initial Characteristics Raw AAER data was programmatically downloaded as JSON objects using the sec-api.io service, specifically querying their AAER Database API endpoint 7 . Approximately 3,300 AAERs were initially collected. Each JSON object contains structured information about an enforcement release. An example structure of a downloaded raw AAER JSON object is shown below: Sample Raw AAER JSON Object "id": "c9ac87509126bd0f1f62c89346cae52d", "dateTime": "2004-02-25T09:21:21-05:00", "aaerNo": "AAER-1964", "releaseNo": ["LR-18595"], "respondents": [ "name": "FOO", "type": "individual" ], "respondentsText": "FOO", "urls": [ "type": "primary", "url": "https://w.sec.gov/enforcement-litigation/litigation-releases/lr-18595" ], "summary": "The SEC filed a complaint against FOO...", "tags": ["disclosure fraud", "financial reporting fraud"], "entities": [ "name": "FOO", "type": "individual", "role": "defendant" , "name": "Just for Feet, Inc.", "type": "company", "role": "entity involved in the fraud", "cik": "918111", "ticker": "FEET" ], "complaints": [ "Ruttenberg was instrumental in the acquisition of fraudulent confirmations..." ], "parallelActionsTakenBy": ["United States Department of Justice", "..."], "hasAgreedToSettlement": false, "hasAgreedToPayPenalty": false, "penaltyAmounts": [], "requestedRelief": ["permanent injunction", "disgorgement", "..."], "violatedSections": ["Section 17(a) of the Securities Act of 1933", "..."], "otherAgenciesInvolved": ["name": "United States Department of Justice", ...] Each AAER JSON object includes a āprimaryurlā field, which typically links to a detailed document (often a PDF or HTML page) on the SEC website. Figure 7 shows an example of the first page of such a document. C.2 AAER Preprocessing Pipeline The downloaded AAERs underwent a multi-step preprocessing pipeline to extract relevant information and filter them for suitability in our fraud detection task. Parsing and Initial Data Extraction The pipeline began by iterating through all downloaded JSON files. For each AAER instance, key fields such as āaaerNoā (cleaned to a standard format, e.g., āAAER-Xā), ādateTimeā (date part extracted), ātagsā, āsummaryā, ācomplaintsā, the 7 https://sec-api.io/docs/aaer-database-api Figure 7: Example first page of a primary document linked from an AAER release. āprimaryurlā, and company-specific details (name, role, CIK) from the āentitiesā list were extracted. This information was structured into a tabular format for further processing. Extraction of Fiscal Quarter Violation Information A critical challenge is that raw AAER JSONs do not directly provide the specific fiscal years and quarters during which the violations occurred. This temporal information is essential for linking fraud events to quarterly financial reports. To address this, we implemented an automated extraction process: 1. Content Retrieval: For each AAER, the content of the document linked by its āprimary urlā was fetched. This involved using web automation tools (like Selenium with ChromeDriver) to download PDF documents or render HTML pages, followed by text extraction (using libraries like PyPDF2 for PDFs and BeautifulSoup for HTML). This process was paral- lelized to expedite the processing of numerous documents. 2. Fiscal Quarter Identification: The extracted text content from each AAER document was then processed by the Qwen3 32B model. This model was tasked with identifying the precise quarters (e.g., ā2018q1ā, ā2019q2ā) and companies associated with the financial violations described in the text. The process was guided by the system prompt detailed below, designed to ensure consistent and accurate extraction. API interactions were managed with rate limiting to ensure robust performance. 3. The promt was tailored to put emphasis on earning mistatements, as it is what our work considers as Fraud. System Prompt for Fiscal Quarter Extraction from AAER Content You are a specialized AI agent tasked with meticulously extracting detailed information about ** earnings misstatements ** (financial statement violations that directly impact the calculation of reported earnings, income, assets, or liabilities) from U.S. Securities and Exchange Commission (SEC) Accounting and Auditing Enforcement Releases (AAERs). Your goal is to deconstruct complex legal and financial text into structured, factual data. ### ** Objective ** Your primary objective is to identify every quarter in which a ** true earning misstatement ** occurred, specify the company responsible, detail the specific types of earning misstatements based on predefined categories, and describe the fraudulent scheme that led to these misstatements. You must adhere strictly to the formats and rules defined below. ### ** Key Definitions: āLIST_MISTATEMENT_TYPEā ** You must categorize all identified ** earnings misstatements ** using ** only ** the types from this predefined list. An "earnings misstatement" directly alters the reported financial performance or position (e.g., net income, assets, liabilities, equity balances). Violations related * solely * to disclosure failures that do not alter the numerical financial statements (e.g., failure to disclose related party relationships without affecting specific account balances, or control issues) should * not * be categorized here unless they clearly result in a numerical misstatement of an account listed below. * ** Revenue ** : Overstating or understating sales or income. This includes premature revenue recognition, fictitious sales, or improper income classification directly impacting the income statement. * ** Other Expense/Shareholder Equity Account ** : Manipulating expenses not directly related to cost of goods sold (e.g., operating, selling, general \& administrative expenses, R\&D), or directly misstating equity accounts (like retained earnings, common stock, additional paid-in capital) through improper accounting entries that affect net income or equity balances. Do not include in this category disclore fraud or governance issues that do not impact financial accounts. * ** Assets Valuation ** : Improperly recognition of assets or their values. This includes inflating asset values (e.g., property, plant \& equipment, intangible assets) or failing to recognize impairments, which directly impacts the balance sheet. * ** Capitalized Costs as Assets ** : Improperly recording expenses as long-term assets (e.g., property, plant \& equipment, intangible assets) to inflate current period income by reducing expenses. * ** Accounts Receivable ** : Overstating the money owed by customers. This includes recording fictitious sales, failing to write off uncollectible receivables, or otherwise inflating the asset balance. * ** Inventory ** : Overstating the value of goods for sale. This includes counting non-existent inventory, improper valuation methods, or misclassifying costs. * ** Cost of Goods Sold (COGS) ** : Understating the direct costs of production. Often linked to inventory manipulation (e.g., overstating inventory leads to understated COGS), directly impacting gross profit and net income. * ** Reserve Account ** : Manipulating funds set aside for future contingent liabilities (e.g., warranty, litigation reserves, bad debt reserves). This includes understating reserves to boost current income or overstating them to create "cookie jar" reserves for future manipulation, directly impacting expenses or liabilities. * ** Liabilities ** : Understating company obligations. This includes concealing debt, failing to record accrued expenses (e.g., unbilled services, payroll), or misclassifying liabilities to improve financial ratios or conceal obligations. * ** Marketable Securities ** : Misstating the value of short-term investments. This includes improper valuation (e.g., failing to mark to market when required) or failing to recognize impairment losses, directly impacting asset values and potentially income. * ** Allowance for Bad Debt ** : Understating the estimated uncollectible accounts receivable to inflate net receivables and income. This is a specific type of reserve manipulation. * ** Payables ** : Understating money owed to suppliers. This includes delaying invoice recording, concealing vendor liabilities, or manipulating cut-off dates, directly impacting liabilities and potentially expenses. --- ### ** Input Format ** You will receive a dictionary containing: * ā"aaerNo"ā: The unique identifier of the AAER. * ā"content"ā: The full text of the AAER. * ā"entities"ā: A list of dictionaries, each representing an entity (company or individual) involved in the AAER. --- ### ** Output Format ** You must generate a JSON list of dictionaries. Each dictionary represents a single fraudulent scheme by a specific company in a specific quarter. ājson [ "quarter": "YYYYqQ", "is_fiscal_quarter": true, "fraud_scheme_description": "A concise, factual description of the earning misstatement mechanics, focusing on how the numerical financial statements were altered. You must also justify the selection of the misstatements indicated in the āmisstatementsā field. The justification should be provided only when the āmisstatementsā field is not empty. The justification should be in the form of a list of sentences, each explaining why a specific misstatement type was selected for that quarter. For example: - Revenue because the company recorded fictitious sales transactions. - Accounts Receivable because the inflated sales led to an overstatement of amounts owed by customers.", "misstatements": ["Type1", "Type2", "Type3", ...], "misstating_company": "name": "Company Name", "role": "respondent", "cik": "0001234567" ] The extracted fiscal quarters were then associated with their respective AAERs in our structured dataset. Filtering and Refinement The dataset of AAERs, now augmented with extracted fiscal quarters, underwent several filtering steps to refine its suitability for our fraud detection task: ⢠Date Filtering: To align with the availability of our financial features (which start from 2009), only AAERs with identified violation years from 2009 onwards were considered for the primary dataset used in the experiments. This comprehensive filtering process significantly narrowed down the set of AAERs to those most pertinent for training and evaluating models for company-level financial statement fraud detection. C.3 Final Fraud Label Statistics After the complete preprocessing and filtering pipeline: ⢠From the initialā¼3,300 AAERs, the filtering steps (for relevance based on tags, direct company culpability, CIK presence, and violations occurring from 2009 onwards) resulted in 249 unique AAERs. ⢠These 249 AAERs were then merged with our financial features dataset. Due to factors such as non-overlapping CIKs between the AAER dataset and the companies present in our financial reports database, and further filtering based on the completeness threshold for financial features, the number of AAERs contributing to fraud labels in our final experimental dataset was reduced to 137. ⢠These 137 AAERs collectively identified 511 firm-quarters as fraudulent instances. These instances formed the positive class in our fraud detection experiments. This rigorous process ensures that the fraud labels used in our study are well-defined, temporally accurate, and directly linkable to the financial and textual data of the companies involved. D Appendix D: Final Dataset Construction and Data Splitting Methodology This appendix outlines the procedures for constructing the final dataset used in our experiments and details the different data splitting strategies employed to evaluate our models under various conditions. D.1 Final Dataset Construction The creation of our final experimental dataset involved several key steps: merging data from different sources, addressing class imbalance, and ensuring data consistency. Data Merging The foundational step was the integration of three primary data sources: ⢠Financial Data: As described in Section 3.1, this includes 122 engineered financial indicators derived from quarterly reports. ⢠Summarized MD&A (SMD&A): Textual data obtained from summarizing MD&A sections, as detailed in Appendix B. ⢠AAER-derived Fraud Labels: Binary fraud labels (fraud/non-fraud) for specific firm-quarters, processed as described in Appendix C. These datasets were merged based on common identifiers: Central Index Key (CIK), fiscal year, and fiscal quarter. Com- pany names were standardized (converted to lowercase) before merging to ensure consistency. A mapping between CIKs and company names was also created for reference. Any records that did not have corresponding entries across these essential dimensions or lacked MD&A data were excluded. Handling Class Imbalance: Stratified Downsampling of Non-Fraud Cases Financial fraud is a rare event, leading to highly imbalanced datasets. To create a tractable dataset for experimentation while preserving a significant level of imbalance, we downsampled the non-fraudulent cases to achieve a target fraud rate of ap- proximately 5.03% in the final dataset (as discussed in Section 3.4). This was performed carefully to maintain the underlying characteristics of the non-fraud data. The stratified downsampling algorithm is as follows: 1. Segregation: The merged dataset was divided into two subsets: fraudulent firm-quarter samples and non-fraudulent firm- quarter samples. 2. Target Calculation: ⢠Let N fraud be the total number of fraudulent samples. ⢠The target number of non-fraudulent samples (N non fraudtarget ) to achieve the desired overall fraud percentage (P fraud = 5.03%) was calculated as: N nonfraudtarget = N fraud Ć 100āP fraud P fraud . 3. Proportional Group Sampling (Initial Pass): ⢠The non-fraudulent subset was grouped by āyearā and āsicaggā. The āsicaggā field represents an aggregated Standard Industrial Classification code, corresponding to high-level industry sectors derived from the first two digits of the SIC code, based on classifications from official sources (e.g., https://siccode.com/). This grouping aims to preserve temporal and industry distributions. ⢠Aglobaldownsamplingfactor(F downsample )wascomputed:F downsample = N non fraudtarget /N nonfraudoriginal , where N nonfraudoriginal is the total count of non-fraudulent samples before downsampling. ⢠For each (āyearā, āsicaggā) group within the non-fraudulent data: ā The target number of samples for this group (N group target ) was calculated by multiplying the original size of this group by F downsample . ā N group target was adjusted to be at least 1 (if the group was non-empty and N grouptarget was positive) and no more than the actual number of samples available in that group. ā N grouptarget samples were randomly selected from this specific group. ⢠All samples selected from these groups were collected. 4. Refinement Pass (If Target Not Met): ⢠If the total number of non-fraudulent samples collected in the initial pass was less than N nonfraudtarget , a refinement step was performed. ⢠Groups that still contained unselected non-fraudulent samples were identified. ⢠Additional samples were iteratively drawn from these groups, prioritizing those with more remaining available sam- ples, until N non fraudtarget was reached or no more unselected samples were available. This ensures the target size is met more closely while still favoring the original distribution. 5. Final Dataset Assembly: The original set of N fraud fraudulent samples was combined with the N nonfraudtarget down- sampled non-fraudulent samples to form the final dataset used for all experiments. A fixed random seed was used during the sampling process to ensure reproducibility. This stratified downsampling ensures that while the dataset is made more balanced, the non-fraudulent samples still reflect the temporal and sectoral diversity of the original population. Final Dataset Composition After merging and downsampling, the final dataset used for our experiments consists of 10,159 firm-quarter observations. This includes: ⢠Fraudulent Samples: 511 (approximately 5.03%) ⢠Non-Fraudulent Samples: 9,648 (approximately 94.97%) ⢠Unique Companies: 5,658 It is important to note that a single company (identified by its CIK) can contribute multiple firm-quarter observations to the dataset, and these can include both fraudulent and non-fraudulent instances over different time periods. The dataset spans from 2009 to 2021. The overall distribution of samples per aggregated industry sector (based on āsicaggā) in this final dataset is presented in Table 8. Table 8: Overall industry sector distribution in the final experimental dataset (after downsampling). Numbers represent firm-quarter instances. Industry Sector (Aggregated SIC - āsicaggā)Total Samples Agriculture, Forestry, And Fishing39 Construction110 Finance, Insurance, And Real Estate2320 Manufacturing3789 Mining640 Public Administration5 Retail Trade426 Services1770 Transportation & Public Utilities786 Wholesale Trade274 D.2 Data Splitting Strategies for FSFD Tasks To evaluate model performance under different assumptions and levels of difficulty, we employed three distinct data splitting strategies, corresponding to the tasks defined in Section 4. For tasks involving cross-validation, 5 folds were used, and a fixed random seed was employed for fold generation to ensure reproducibility. Classic FSFD: Random K-Fold Cross-Validation This strategy represents the traditional approach to evaluating FSFD models. ⢠Methodology: The final dataset of 10,159 firm-quarter observations was split into 5 folds using random sampling. Strati- fication was applied based on the āis fraudā label to ensure that each fold maintained approximately the same 5.03% fraud ratio as the overall dataset. ⢠Characteristics: In this setup, observations from the same company can appear in both the training and testing sets of a given fold (though not the same observation). This allows the model to potentially learn company-specific patterns. ⢠Fold Statistics (Averages over 5 Folds): ā Test Set Size: ā¼2032 samples ā Fraud Samples in Test: ā¼102 (Fraud Ratio: ā¼5.03%) ā Industry distribution in test folds (e.g., āManufacturingā ā¼758 samples with ā¼52 fraud; āServicesā ā¼354 samples withā¼22 fraud) typically reflected the overall dataset distribution due to random sampling. ⢠Illustration: See Figure 8. Full Dataset (Firm-Quarters) Fold 1Fold 2 Fold 3 Fold 4 Fold 5 Test Train (F2,F3,F4,F5) Firm-quarters randomly assigned. Same com- pany may appear in train and test portions of a fold. Process repeated for each fold as test set. Figure 8: Conceptual diagram of Classic FSFD (Random K-Fold) splitting. Each fold serves as a test set once, with the remaining folds as training. Company-Isolated FSFD (CI-FSFD): Company-Based K-Fold Cross-Validation This strategy imposes a stricter evaluation by ensuring that companies seen during training are not present in the test set. ⢠Methodology: The dataset was split into 5 folds at the company (CIK) level. All firm-quarter observations belonging to a specific company (which may include a mix of fraud and non-fraud instances for that company) were assigned entirely to one fold. The assignment of companies to folds was performed aiming to balance the number of fraud reports and total reports per fold, and preserve the industry sector distribution within each foldās test set as much as possible. ⢠Characteristics: This split tests the modelās ability to generalize to entirely unseen companies, preventing it from relying on idiosyncratic patterns of companies present in the training data. ⢠Fold Statistics (Averages over 5 Folds): ā Test Set Size: ā¼2032 samples ā Unique Companies in Test: ā¼1132 ā Fraud Samples in Test: ā¼102 (Fraud Ratio: ā¼5.03%) ā Industry distributions were actively balanced. For example, across test folds, āManufacturingā had ā¼758 samples (withā¼52 fraud), and āServicesā hadā¼354 samples (withā¼22 fraud). ⢠Illustration: See Figure 9. Full Dataset (Grouped by Company) Companies ACompanies BCompanies CCompanies DCompanies E Fold 1 (Test) (e.g., Co. Group A) Train for Fold 1 (e.g., Co. Groups B, C, D, E) Groups of companies are assigned to folds. If a companyās data (all its firm-quarter instances) is in the test set of a fold, none of its data appears in the training set for that fold. Figure 9: Conceptual diagram of Company-Isolated FSFD (CI-FSFD) splitting. Companies (and all their associated firm-quarter instances) are assigned to folds. E Appendix E: Modelās Details and Hyperparameters This appendix provides detailed model configurations and hyperparameters for all models used in our experiments: Logistic Regression, MLP, Random Forest (LightGBM), XGBoost, RCMA-adapted, and LLM-based models. For all models, hyperpa- rameters were optimized using Hyperopt, targeting the maximization of the F1-score on a dedicated validation set (typically 10% of the training data for that specific fold/split). The ādecision thresholdā reported for classification models is the optimal threshold found on the validation set. It is important to note that the same set of optimized hyperparameters was applied across all folds for a given model and task (e.g., Classic FSFD or CI-FSFD). E.1 Logistic Regression (MLP-Classifier with no Hidden Layers) Our Logistic Regression baseline is implemented as a specialized case of the MLP Classifier with no hidden layers. Table 9: Optimized Hyperparameters for Logistic Regression. ParameterValue Features TypeDechow Learning Rate0.1 Batch Size64 Dropout Rate0.459 Epochs2 Patience20 OversampleTrue StandardizeTrue Decision Threshold0.5 E.2 Multi-Layer Perceptron (MLP) The MLP Classifier uses a feed-forward neural network architecture. Table 10: Optimized Hyperparameters for MLP. ParameterValue Classic FSFD Task Features TypeFinancial (122 features) Hidden Dims[512] Learning Rate0.1 Batch Size128 Dropout Rate0.413 CI-FSFD Task Features TypeFinancial (122 features) Hidden Dims[512, 512] Learning Rate0.001 Batch Size32 Dropout Rate0.145 Common Parameters Epochs2 Patience20 OversampleTrue StandardizeFalse Decision Threshold0.5 E.3 Random Forest (LightGBM Implementation) For the Random Forest baseline, we utilized the Random Forest mode of the LightGBM library [ Ke et al., 2017 ] . E.4 XGBoost The XGBoost classifier [ Chen and Guestrin, 2016 ] was optimized with Hyperopt. Table 11: Optimized Hyperparameters for Random Forest (LightGBM). ParameterValue Classic FSFD Task Features TypeFinancial (122 features) Learning Rate0.0799 Max Depth7 Num Estimators200 Num Leaves5 StandardizeTrue CI-FSFD Task Features TypeFinancial (122 features) Learning Rate0.0996 Max Depth78 Num Estimators200 Num Leaves50 StandardizeFalse Common Parameter Decision Threshold0.5 Table 12: Optimized Hyperparameters for XGBoost. ParameterValue Classic FSFD Task Features TypeFinancial (122 features) and Dechow Learning Rate0.05 Max Depth0 Num Estimators50 Num Leaves50 StandardizeTrue CI-FSFD Task Features TypeFinancial (122 features) Learning Rate0.090 Max Depth88 Num Estimators200 Num Leaves50 StandardizeTrue Common Parameter Decision Threshold0.5 E.5 RCMA-adapted Model The RCMA-adapted model architecture and training parameters were based on the work of Wang et al. [ 2023 ] , with modifica- tions for our SBERT-based text processing. E.6 LLM-based Models This section details the configuration for our Large Language Models, including Llama-3.1 8B and Fino1 8B, in various fine- tuning and zero-shot settings. All LLMs use LoRA for fine-tuning. Financial-only Models (Llama-3.1 8B and Fino1 8B) These models are fine-tuned exclusively on financial text. SMD&A-only Models (Llama-3.1 8B, Fino1 8B, Fino1 14B) These models are fine-tuned exclusively on Summary Management Discussion & Analysis (SMD&A) text. Table 13: Optimized Hyperparameters for RCMA-adapted Model. ParameterValue SBERT Model Nameājinaai/jina-embeddings-v2-small-enā SBERT Output Dim512 Trainable SBERT Layers1 Max SMD&A Length8192 tokens Financial Features122 Num Financial Groups7 Financial Embedding Dim512 Text Embedding Dim512 Dropout Rate0.05 Learning Rate1e-4 Batch Size8 Validation Batch Size8 Pos Weight Beta (for FocalLoss)0.75 Focal Gamma (for FocalLoss)2.0 Gradient Accumulation Steps4 Classic FSFD Task MLP Hidden Dims[128] Epochs10 Patience7 Consistency Loss Weight0.2 OversampleFalse CI-FSFD Task MLP Hidden Dims[512] Epochs20 Patience10 Consistency Loss Weight0 OversampleTrue Financial+SMD&A Models (Llama-3.1 8B and Fino1 8B) These models are fine-tuned on both Financial (122 features) and Summary Management Discussion & Analysis (SMD&A) sections. Zero-shot Model (Fino1 8B SMD&A) For the zero-shot evaluation, the Fino1 8B model was used without any fine-tuning (i.e., ānumlayerstofinetuneā is 0). Table 14: Hyperparameters for LLM-based Models (Financial-only). ParameterValue Llama-3.1 8B (Financial) Model URLāunsloth/Llama-3.1-8B-unsloth-bnb-4bitā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Max Context1500 tokens Batch Size4 Gradient Accumulation Steps2 Learning Rate1e-4 LoRA Target Modulesāq projā, āvprojā, āupprojā, ādownprojā, āgateprojā Fino1 8B (Financial) Model URLāTheFinAI/Fino1-8Bā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Max Context1500 tokens Batch Size4 Gradient Accumulation Steps2 Learning Rate1e-4 LoRA Target Modulesāq projā, āvprojā, āupprojā, ādownprojā, āgateprojā Common Parameters Epochs (Classic FSFD Task)10 Epochs (CI-FSFD Task)20 Max New Tokens1 Only CompletionTrue UndersampleTrue Run Eval on StartFalse Table 15: Hyperparameters for LLM-based Models (SMD&A-only). ParameterValue Llama-3.1 8B (SMD&A) Model URLāunsloth/Llama-3.1-8B-unsloth-bnb-4bitā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Max Context8500 tokens Batch Size8 Gradient Accumulation Steps1 Learning Rate1e-4 LoRA Target Modules (Classic FSFD Task)āq projā, āvprojā, āupprojā, ādownprojā, āgateprojā LoRA Target Modules (CI-FSFD Task)āqprojā, āvprojā, āupprojā, ādownprojā, āgateprojā, ālmheadā Fino1 8B (SMD&A) Model URLāTheFinAI/Fino1-8Bā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Max Context8500 tokens Batch Size8 Gradient Accumulation Steps1 Learning Rate1e-4 LoRA Target Modules (Classic FSFD Task)āq projā, āvprojā, āupprojā, ādownprojā, āgateprojā LoRA Target Modules (CI-FSFD Task)āqprojā, āvprojā, āupprojā, ādownprojā, āgateprojā, ālmheadā Fino1 14B (SMD&A) Model URLāTheFinAI/Fin-o1-14Bā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune40 Max Context8500 tokens Batch Size4 Gradient Accumulation Steps2 Learning Rate1e-4 LoRA Target Modulesāq projā, āvprojā, āupprojā, ādownprojā, āgateprojā Common Parameters Epochs (Classic FSFD Task)10 Epochs (CI-FSFD Task)20 Max New Tokens1 Only CompletionTrue UndersampleTrue Run Eval on StartFalse Use Full SummaryTrue Table 16: Hyperparameters for LLM-based Models (Financial+SMD&A). ParameterValue Llama-3.1 8B (Financial+SMD&A) Model URLāunsloth/Llama-3.1-8B-unsloth-bnb-4bitā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Learning Rate1e-4 LoRA Target Modulesāq projā, āvprojā, āupprojā, ādownprojā, āgateprojā Fino1 8B (Financial+SMD&A) Model URLāTheFinAI/Fino1-8Bā LoRA R8 LoRA Alpha8 LoRA Dropout0.05 Layers to Finetune32 Learning Rate1e-4 LoRA Target Modulesāq projā, āvprojā, āupprojā, ādownprojā, āgateprojā Common Parameters Max Context (Classic FSFD Task)9500 tokens Max Context (CI-FSFD Task)9500 tokens Batch Size (Classic FSFD Task)8 Batch Size (CI-FSFD Task)4 Gradient Accumulation Steps (Classic FSFD Task)1 Gradient Accumulation Steps (CI-FSFD Task)2 Epochs (Classic FSFD Task)10 Epochs (CI-FSFD Task)20 Max New Tokens1 Only CompletionTrue UndersampleTrue Run Eval on StartFalse Use Full SummaryTrue Table 17: Hyperparameters for Zero-shot Fino1 8B SMD&A Model. ParameterValue Model URLāTheFinAI/Fino1-8Bā Max Context9500 tokens Max New Tokens1 Batch Size1 Zero-ShotTrue F Appendix F: LLM System Prompts This appendix details the system prompts used for the Large Language Model (LLM) based classifiers, corresponding to the different input data configurations: Financials Only (FIN), Synthetically Summarized MD&A Only (SMD&A), and combined Financials + SMD&A. For each configuration, the model was provided with a specific user prompt outlining the task, the context (industry sector), and the relevant data. The LLM was then fine-tuned to generate a single token representing the classification: āYESā (indicating fraud) or āNOā (indicating non-fraud) immediately following the prompt. The prompts were designed to fit within the modelās context window, with specific token allowances made for the variable data portions (financial strings or MD&A content). The token counts provided below are approximate estimates for the fixed textual parts of each prompt, based on the Llama-3 tokenizer, and exclude the tokens from placeholder content like āindustry titleā, āfinancialsstrā, or āmdacontentā. F.1 Prompt for Financials Only (FIN) Input When using only financial data, the LLM was presented with the following prompt structure. LLM Prompt: Financials Only You are a financial forensic analyst. The company operates in the industry_title sector. Below are key financial indicators derived from its income statement, balance sheet, and cash flow statement: financials_str Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is engaging in Financial Manipulation Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"? F.2 Prompt for SMD&A Only (SMD&A) Input When using only the Synthetically Summarized MD&A text, the following prompt structure was employed.. LLM Prompt: SMD&A Only The company operates in the industry_title sector. Below is the summary of the Management Discussion and Analysis (MDA) section of the quarterly report: mda_content Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Manipulation Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"? F.3 Prompt for Combined Financials + SMD&A Input For the combined input scenario, the LLM received a prompt integrating both financial metrics and the SMD&A content. LLM Prompt: Financials + SMD&A The company operates in the industry_title sector. Here are financial variables derived from the income statement, balance sheet, and cash flow statement of the company. financials_str Also below is the structured summary of the Management Discussion and Analysis (MDA) section of the quarterly report: mda_content Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"? G Appendix G: Sample Prediction Details This appendix provides an illustrative example of a prediction made by our Fino-1 8B model using the combined Financial + SMD&A input. It shows the complete prompt provided to the model (truncated for brevity in this display, but the full content was used for prediction), the modelās generated answer, the ground truth label, and the associated prediction probabilities. The following example corresponds to a firm-quarter instance from the CI-FSFD task, where the model correctly identified a fraudulent case. Sample Model Input Prompt (Financials + SMD&A - Truncated for Display) The company operates in the RUBBER & PLASTICS FOOTWEAR sector. Here are financial variables derived from the income statement, balance sheet, and cash flow statement of the company. - Total Assets: $22,921,000,000 - Cash and Short-term Investments: $3,695,000,000 - Property, Plant, and Equipment: $4,688,000,000 - Degree Of Financial Leverage: 873% - Invested Capital Ratio: 20% - Cash To Total Asset: 16% - Debt Service Coverage: 36% - Financial LeverageIndex: 5.44% - Times InterestEarnedRatio: 9,275% - Current Asset To Revenues: 164% - Current Liabilities To Revenues: 76% - Short TermDebt To Revenue: 0.23% - Intangible Asset ToRevenue: 4.55% - LongtermLeverage: 15% - Asset Quality Index: 0.96 - Leverage Index: 0.99 ... Also below is the structured summary of the Summary Management Discussion and Analysis (SMD&A) section of the quarterly report: # 1. Strategic Priorities and Initiatives NIKEās goal is to deliver value to shareholders by building a profitable global portfolio of branded footwear, apparel, equipment, and accessories businesses. The companyās strategy is to achieve long-term revenue growth by creating innovative, must-have products, building deep personal consumer connections with its brands, and delivering compelling consumer experiences through digital platforms and at retail. In fiscal 2018, NIKE introduced the Consumer Direct Offense, a new company alignment designed to allow NIKE to better serve the consumer more personally, at scale. Through the Consumer Direct Offense, NIKE is focusing on the Triple Double strategy, with the objective of doubling the impact of innovation and increasing its speed to market and direct connections with consumers. --- # 2. Operational and Segment Performance For the third quarter of fiscal 2019, NIKE Brand delivered 8% revenue growth, with 12% growth on a currency-neutral basis, driven by higher revenues across all geographies, footwear and apparel, as well as growth in most key categories, led by Sportswear and the Jordan Brand. Converse revenues decreased 4% on a reported basis and 2% on a currency-neutral basis, primarily due to declines in the U.S. and Europe, partially offset by revenue growth in Asia. In North America, on a currency-neutral basis, revenues increased 7% for the third quarter and first nine months of fiscal 2019, driven by growth in several key categories for the quarter and nearly all key categories for the year-to-date period, led by Sportswear. NIKE Direct in North America increased 6% and 7% for the third quarter and first nine months, respectively, as higher digital commerce sales and the addition of new stores more than offset an 8% and 4% decline in comparable store sales, driven by NFS performance. In EMEA, on a currency-neutral basis, revenues grew 12%, driven by balanced growth across all territories and led by Sportswear and the Jordan Brand. NIKE Direct in EMEA increased 15% and 14% for the third quarter and first nine months, respectively, due to comparable store sales growth, higher digital commerce sales, and new store additions. In Greater China, on a currency-neutral basis, revenues increased ... --- # 3. Financial Results and Key Trends For the third quarter of fiscal 2019, revenues increased 7% to $9.6 billion, and net income was $1.1 billion with diluted earnings per share of $0.68, compared to a net loss of $921 million and diluted loss per share of $0.57 for the same period in fiscal 2018. Income before income taxes increased 11%, driven by revenue growth and gross margin expansion, partially offset by higher selling and administrative expense. Gross margin increased to 45.1% for the third quarter of fiscal 2019, compared to 43.8% in the same period in fiscal 2018. For the first nine months of fiscal 2019, revenues increased 9% to $28.9 billion, and net income was $3.0 billion, compared to $3.1 billion in the prior year. Gross margin for the nine months ended February 28, 2019, was 44.4%, compared to 43.5% in the prior year. Selling and administrative expense increased to $3.09 billion for the third quarter of fiscal 2019, representing 32.2% of revenues, compared to $2.767 billion and 30.8% of revenues for the third quarter of fiscal 2018. For the first nine months of fiscal 2019, selling and administrative expense increased to $9.296 billion, representing 32.1% of revenues, compared to $8.391 billion and 31.5% of revenues in the prior year. ... --- # 4. Identified Risks and Uncertainties The company is exposed to foreign currency market volatility, partly due to global trade uncertainty and geopolitical dynamics. Foreign currency exposures arise from transactions denominated in non-functional currencies and the translation of foreign currency-denominated results into U.S. Dollars. Argentina has been identified as a hyper-inflationary market, and the functional currency of the Argentina subsidiary was changed to U.S. Dollars in the second quarter of fiscal 2019. The translation of foreign currency-denominated profits and foreign exchange rate fluctuations had an unfavorable impact on income before income taxes for both the third quarter and first nine months of fiscal 2019. The functional currency change in Argentina did not have a material impact on the Companyās results of operations or financial condition, and management does not anticipate a material impact in future periods based on current rates. The Company may face challenges in accessing credit markets or increased interest costs due to future volatility. Foreign currency hedge gains and losses may impact operating performance, depending on actual market rates versus standard rates. ... --- # 5. Forward-Looking Statements and Guidance The company remains committed to its long-term financial goals, and continues to see opportunities to drive growth and profitability despite foreign currency volatility. NIKE Direct is expected to continue accelerating growth, driven by digital commerce and store expansion. Investments in data and analytics, digital commerce platforms, and a new enterprise resource planning tool are part of the end-to-end digital transformation strategy. The company plans to continue share repurchases under its new $15 billion four-year program, with funding expected from operating cash flows, excess cash, and debt proceeds. Management believes that existing cash, cash equivalents, short-term investments, and cash generated by operations, along with access to external funding, will be sufficient to meet capital needs in the foreseeable future. ... --- # 6. Significant Changes, Events, or Developments The Company completed the $12 billion share repurchase program authorized in November 2015 during the first nine months of fiscal 2019, repurchasing 192.1 million shares. A new four-year, $15 billion share repurchase program was authorized in June 2018, and 43.7 million shares were repurchased under this program during the first nine months of fiscal 2019. The functional currency of the Argentina subsidiary was changed to U.S. Dollars in the second quarter of fiscal 2019, due to hyper-inflationary conditions. ... --- # 7. Important Figures and Tables Revenues for the Three Months Ended February 28, 2019 and 2018 | Period | Revenues (in millions) | % Change | % Change |--------|------------------------|----------|----------- | 2019 | $9,611 | | | 11% | 2018 | $8,984 | 7% | | Revenues for the Nine Months Ended February 28, 2019 and 2018 | Period | Revenues (in millions) | % Change | % Change |--------|------------------------|----------|----------- | 2019 | $28,933 | | | 11% | 2018 | $26,608 | 8% | | Gross Profit and Gross Margin for the Three Months Ended February 28 | Period | Gross Profit (in millions) | % Change | Gross Margin | |--------|----------------------------|----------|--------------| | 2019 | $4,339 | | | 45.1% | | 2018 | $3,938 | 10% | 43.8% | ... --- # 8. Management Explanations and Justifications The increase in NIKE Brand revenues is attributed to growth across all geographies, footwear and apparel, and key categories like Sportswear and the Jordan Brand. The decline in Converse revenues is explained by revenue declines in the U.S. and Europe, partially offset by growth in Asia. ... --- # 9. Accounting Estimates, Judgments, and Policy Changes Revenue recognition is based on transfer of control to the customer, with variable consideration for sales returns, discounts, and miscellaneous claims estimated and recorded as a reduction to revenues. ... --- # 10. Capital Allocation and Liquidity Management Share repurchase activity increased significantly in fiscal 2019, with $3.386 billion spent on 43.7 million shares. ... --- # 11. Legal, Regulatory, and Compliance Matters The Company has no off-balance sheet arrangements that have or are reasonably likely to have a material effect on financial condition, results of operations, liquidity, or capital resources. ... Based on these informations and your knowledge of typical red flags in financial reporting, assess whether there is a high likelihood that this company is Financial Fraud. Do you think this company is engaging Fraud? Answer with "YES" or "NO"? Model Prediction and Ground Truth Modelās Generated Answer: NO Ground Truth Label: NO Prediction Probability for āYESā: 0.0073334336280823 (Note: The decision threshold optimized on the validation set for this fold was applied to this probability to yield the binary prediction.) Instance Identifiers: ⢠CIK: 320187 ⢠SIC (Aggregated): 3021 (RUBBER & PLASTICS FOOTWEAR) ⢠Quarter: 2019q3 This example illustrates how the model processes the combined textual and numerical information to arrive at a classification decision. The relatively low probability for a correct āYESā prediction in this specific case, despite being above the decision threshold for this fold, highlights the challenging nature of the CI-FSFD task. H Appendix H: Detailed Performance Analysis This appendix provides a more granular look at the performance of our LLM-based fraud detection framework (Fino-1 8B with SMD&A input) across different subgroups: individual companies, industry sectors, and performance on unseen companies in the classic setting. We present key metrics such as True Positives (TP - Detected Fraud), False Negatives (FN - Undetected Fraud), False Positives (FP - Non-Fraud Incorrectly Flagged as Fraud), Total Actual Fraud instances, Recall (TP / (TP + FN)), and Precision (TP / (TP + FP)). H.1 Fine-Grain Labels Performance Analysis Table 18 presents a detailed breakdown of the modelās performance on detecting different types of financial statement misstate- ments. The model exhibits varying recall across different misstatement categories, indicating that certain types of fraud are more challenging to detect than others. Table 18: Fraud Detection Performance by Misstatement Type (CI-FSFD Task). Model: Fino-1 8B with SMD&A. Misstatement TypeDetected Fraud (TP)Undetected Fraud (FN)Total Actual FraudRecallAUC misReserve Account74110.6360.819 misCapitalized Costs as Assets78150.4670.796 mis Cost of Goods Sold (COGS)1524390.3850.669 misAccounts Receivable2552770.3250.756 mis Allowance for Bad Debt1340.2500.955 misLiabilities1959780.2440.689 mis Revenue511602110.2420.707 misPayables1756730.2330.698 misOther Expense/Shareholder Equity Account552102650.2080.645 misAssets Valuation1085950.1050.536 misInventory236380.0530.565 H.2 Company-Level Performance Analysis (CI-FSFD) The Company-Isolated FSFD (CI-FSFD) task evaluates the modelās ability to generalize to companies not seen during training. Performance at the individual company level can vary significantly. Table 19 lists the top 30 companies where the model demonstrated the best performance (highest recall) in detecting fraud- ulent quarters under the CI-FSFD setting. Table 20 shows the 30 companies where the model struggled the most (lowest recall). The visual representations of True Positives and False Negatives for the top 10 from these company groups are in Figure 10 and Figure 11 respectively. Table 19: Top 30 Companies by Recall in Detecting Fraudulent Quarters (CI-FSFD Task), with False Positives and Precision. Model: Fino-1 8B with SMD&A. Industry SectorDetected Fraud (TP)Undetected Fraud (FN)False Positives (FP)Total Actual FraudRecallPrecision (Company Name) 3m company90191.0000.900 Akorn, inc.20021.0001.000 Hertz20021.0001.000 Ixia10011.0001.000 Jda software group, inc.20021.0001.000 Mcdermott international inc.10011.0001.000 Ocz technology group, inc.20021.0001.000 Roadrunner transportation systems, inc.60061.0001.000 Surgalign holdings, inc.20021.0001.000 Swisher hygiene inc.10011.0001.000 Taronis technologies, inc.10011.0001.000 Cognizant technology solutions corporation20221.0000.500 Kbr, inc.10211.0000.333 L3 technologies, inc.10211.0000.333 Tech data corp.10111.0000.500 Uti worldwide inc.20121.0000.667 Nci, inc.70171.0000.875 Newell brands inc.30131.0000.750 Quadrant 4 system corp.71080.8751.000 The kraft heinz co.41150.8000.800 Halliburton company31040.7501.000 Sciclone pharmaceuticals, inc.31040.7501.000 Corporate resource services, inc.21030.6671.000 Future fintech group inc.21030.6671.000 Mimedx group, inc.63090.6671.000 Synchronoss technologies, inc.953140.6430.750 Granite construction inc.54090.5561.000 Apex global brands inc.11020.5001.000 Evoqua water technologies corp.11020.5001.000 Gt advanced technologies inc.11020.5001.000 Table 20: Worst 30 Companies by Recall in Detecting Fraudulent Quarters (CI-FSFD Task). Model: Fino-1 8B with SMD&A. Industry SectorDetected Fraud (TP)Undetected Fraud (FN)False Positives (FP)Total Actual FraudRecallPrecision (Company Name) African gold acquisition corp.01010.0000.000 Amyris, inc.01010.0000.000 Andeavor llc01010.0000.000 Argo group international holdings, ltd.0150150.0000.000 Assisted living concepts, inc.02020.0000.000 Axesstel, inc.01010.0000.000 Barrett business services, inc.0100100.0000.000 Belden inc.01010.0000.000 Biomet, inc.04040.0000.000 Blue earth, inc.02020.0000.000 Brixmor property group inc.04040.0000.000 Cantaloupe, inc.05050.0000.000 Celadon group, inc.03030.0000.000 Celsius holdings, inc.02120.0000.000 China valves technology, inc.01010.0000.000 Chs inc.0150150.0000.000 Citigroup inc.01010.0000.000 Comscore, inc.05050.0000.000 Cpi aerostructures, inc.02020.0000.000 Dxc technology company02020.0000.000 Elanco animal health inc.04040.0000.000 Fmc technologies, inc.05050.0000.000 Fte networks, inc.02020.0000.000 General electric company07070.0000.000 General motors company07070.0000.000 Gtt communications, inc.02220.0000.000 Healthcare services group, inc.04040.0000.000 Home loan servicing solutions, ltd.05050.0000.000 Homestreet, inc.07070.0000.000 Iconix brand group, inc.03030.0000.000 Figure 10: Plot: Top 10 companies by number of correctly detected fraudulent quarters (True Positives vs False Negatives) in the CI-FSFD task. Model: Fino-1 8B with SMD&A. Figure 11: Plot: Top 10 companies by number of undetected fraudulent quarters (True Positives vs False Negatives) in the CI-FSFD task. Model: Fino-1 8B with SMD&A. H.3 Sectorial Performance Analysis (CI-FSFD) Performance also varies when aggregated by industry sector. Table 21 details the detection performance, including false posi- tives and precision, across major industry sectors for the CI-FSFD task. Table 21: Fraud Detection Performance by Industry Sector (CI-FSFD Task). Model: Fino-1 8B with SMD&A. Industry SectorDetected Fraud (TP)Undetected Fraud (FN)False Positives (FP)Total Actual FraudRecallPrecision Construction6420100.6000.231 Retail trade4625100.4000.138 Services34742061080.3150.142 Transportation & public utilities82147290.2760.145 Manufacturing611982772590.2360.180 Mining41510190.2110.286 Wholesale trade42042240.1670.087 Finance, Insurance, & Real Estate05257520.0000.000 Overall (CI-FSFD Task Total)1213906845110.2370.150 Note: The āOverallā row aggregates TP, FN, FP, and Actual Fraud across all test folds for the CI-FSFD task for the listed sectors and calculates overall Recall and Precision from these sums. Figure 12 visually represents the True Positives and False Negatives by sector. Figure 12: Plot: Detected Fraud (TP) vs. Undetected Fraud (FN) cases per industry sector in the CI-FSFD task. Model: Fino-1 8B with SMD&A. H.4 Performance on Unseen Companies in Classic FSFD Setting While the Classic FSFD setting involves random splitting, a small fraction of companies in the test set of each fold might still be entirely unseen during the training phase for that specific fold. Analyzing performance on these truly āunseenā companies within the classic random split provides insight into the modelās baseline generalization even when not explicitly forced by a company-isolated split. Table 22 summarizes key average metrics for the Fino-1 8B (SMD&A) model on these unseen company instances within the Classic FSFDās 5-fold cross-validation. The very low average number of unseen fraudulent instances (1.6 per fold) in the Classic FSFD setting makes it difficult to draw firm conclusions about the modelās ability to detect fraud in entirely new companies from these specific metrics (F1, Recall, Precision). Table 22: Average Performance Metrics on Unseen Companies within Classic FSFD Test Folds. Model: Fino-1 8B with SMD&A (Averages over 5 Folds). MetricMean Value± Std. Dev. Avg. Test Samples per Fold2604.2± 0.4 Avg. Unseen Samples in Test Fold757.6± 18.4 Avg. Test CIKs per Fold2163.2± 9.4 Avg. Unseen CIKs in Test Fold672.6± 14.1 Avg. Total Fraud Samples in Test Fold130.8± 15.7 Avg. Unseen Fraud Samples in Test Fold1.6± 1.2 Avg. Fraud Rate among Unseen Samples0.0021± 0.0015 F1 Score (on Unseen Samples)0.0± 0.0 Recall (on Unseen Fraud Samples)0.0± 0.0 Precision (on Unseen Samples)0.0± 0.0 AUC Score (on Unseen Samples)0.646± 0.280 Accuracy (on Unseen Samples)0.9924± 0.0022 The performance metrics such as F1 score, Recall, and Precision for unseen fraud cases are not statistically significant due to the very low average number of unseen fraud samples (1.6 per fold) in this random splitting setting. This underscores the importance of dedicated CI-FSFD for robustly evaluating generalization. I Appendix I: Explainability This appendix details the methodology employed for generating explanations of our LLMās predictions, specifically focusing on the textual components of the input. Understanding which parts of the financial text contribute most to a fraud prediction is crucial for interpretability and trustworthiness in high-stakes domains like financial anomaly detection. I.1 LRP-based Explanation Our explainability approach is built upon the Layer-wise Relevance Propagation (LRP) [ Achtibat et al., 2024 ] framework. LRP is a technique used to decompose the prediction of a deep neural network into contributions of its input features. It assigns a ārelevance scoreā to each input component (e.g., a token) indicating its importance to the final output. For a classification task, LRP propagates the prediction score backward through the network, layer by layer, until it reaches the input features. The core idea is to conserve the total relevance during propagation, ensuring that the sum of relevances at one layer equals the sum of relevances at the preceding layer. This property allows for a clear attribution of the final prediction to individual input elements. In our implementation, we leverage the LXT library, which provides an efficient and specialized LRP implementation for Transformer-based models. After the model makes a prediction (i.e., outputs logits for āFraudā or āNot Fraudā), we backprop- agate the relevance from the logit corresponding to the predicted class (or, more specifically, the āFraudā logit, regardless of the prediction, to understand drivers of potential fraud) back to the input embeddings. The LRP rules applied ensure that the relevance scores accurately reflect the contribution of each token in the input sequence to that specific logit. The process for generating LRP explanations for each test sample is as follows: ⢠The trained LLM is set to evaluation mode, and all its parameters are frozen ) for the input embeddings, which are set to requires grad=True to compute gradients for LRP. ⢠For each test sample, the input prompt (containing the textual and numerical financial data) is tokenized and fed into the model. ⢠The model performs a forward pass to obtain the logits for the FRAUD LABELID and NOTFRAUDLABELID tokens at the last position of the output sequence. ⢠The gradient of the FRAUD LABELID logit (representing the unnormalized score for the āFraudā class) with respect to the input embeddings is computed. ⢠The LRP relevance score for each input token is then calculated as the element-wise product of the input embeddings and their corresponding gradients, summed across the embedding dimensions. This yields a single relevance score for each token. ⢠These raw relevance scores are then normalized by their absolute maximum to scale them between -1 and 1, facilitating easier interpretation. I.2 Token-Level vs. Sentence-Level Relevance Given the nature of our input data, which consists of long financial reports (averaging around 3800 tokens per context), pro- viding token-level relevance scores directly to a human analyst can be overwhelming and impractical for actionable insights. A raw sequence of 3800 token relevance scores does not immediately highlight the key information at a glance. Therefore, focusing on individual token relevance is not the most pertinent approach for interpretability in this context. Instead, we emphasize sentence-level relevance. Financial analysts typically review reports section by section, and un- derstanding which sentences or clauses are most indicative of fraud is significantly more valuable than knowing the exact contribution of every single token. Sentence-level aggregation allows for a higher-level summary of the modelās reasoning, making the explanations more digestible and actionable. I.3 Sentence-Level Relevance Aggregation To derive sentence-level relevance from the token-level scores, we implemented a simple yet effective aggregation method: ⢠Summing of Relevance Scores: For every sentence, we sum the absolute relevance scores of all tokens belonging to that sentence. Using the absolute sum helps identify sentences that strongly contribute, either positively or negatively, to the fraud prediction. ⢠Normalization and Ranking: The summed relevance scores for sentences are then normalized and ranked. Sentences with higher absolute summed relevance are considered more impactful on the modelās prediction. This aggregation provides a concise summary of the most relevant sentences within a lengthy financial disclosure, allowing users to quickly pinpoint suspicious statements or critical pieces of information that drove the modelās classification decision. I.4 Highlighted Examples For visual interpretability, we generate PDF heatmaps that highlight the most relevant portions of the input. These heatmaps use color intensity to represent the magnitude of a tokenās (or aggregated wordās) relevance score, with different hues indicating positive or negative contributions to the fraud prediction. Figure 13: Example of a highlighted financial report section indicating token-level relevance for a fraud prediction. Red indicates higher positive relevance towards a āFraudā prediction, while blue indicates negative relevance.