Paper deep dive
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.
Tags
Links
- Source: https://arxiv.org/abs/2607.23813v1
- Canonical: https://arxiv.org/abs/2607.23813v1
Trouble viewing inline? Open PDF directly →
Full Text
22,766 characters extracted from source content.
Expand or collapse full text
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance Denglin Jiang, Haoran Zhou, Anshul Wadhawan, Brendan Fahy, Vinay Ramesh, David Weisberg, Dmitriy Derkachevskiy, Helen Sheehan, Srivas Prasad, Michele Franceschini Bloomberg, United States djiang108,hzhou245,awadhawan9,bfahy2,vramesh7,dweisberg6, dderkachevsk,hsheehan,sprasad60,mfrancesch10@bloomberg.net Abstract We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English- language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (i) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earn- ings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry la- bels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We re- port reproducible baselines for Whisper and Parakeet-TDT us- ing standardized scoring. Index Terms: automatic speech recognition, ASR benchmark, financial speech, earnings calls, speech dataset, industry-aware evaluation, long-form speech, domain adaptation 1. Introduction 1.1. Financial speech understanding is challenging Earnings calls are a critical channel for corporate communica- tion and pose unique challenges for automatic speech recogni- tion (ASR). They combine spontaneous dialogue with scripted remarks, heavy financial jargon, company and product names, frequent numeric expressions, rapid turn-taking, and overlap- ping speakers during Q&A—conditions that expose limitations of ASR systems trained primarily on read or general conversa- tional speech. These challenges are amplified by domain shift. Earnings calls often include operator boilerplate, variable recording qual- ity due to telephony compression or background noise, and fre- quent speaker transitions. Linguistically, the domain contains jargon and acronyms under-represented in general ASR training data, as well as accented English from international executives. Although self-supervised and weakly supervised pretraining has improved general ASR performance (e.g., wav2vec 2.0 [1] and Whisper [2]), domain mismatch remains a major source of error, motivating the need for recent, domain-specific speech resources. 1.2. Toward a comprehensive benchmark for financial ASR Despite growing interest in financial speech and NLP, few pub- licly available resources have been purpose-built for evaluating ASR on earnings calls. Existing datasets often emphasize train- ing data without standardized evaluation protocols or lack the metadata required for detailed error analysis. As a result, the field lacks a principled and reproducible benchmark for mea- suring progress on financial ASR. Beyond transcription accuracy, practical financial speech systems require evaluation of speaker attribution, diarization, role-aware performance (e.g., executives vs. analysts), and er- rors across industries and call structure.A comprehensive benchmark must therefore pair realistic long-form audio with rich metadata. Earnings25 is designed to provide aligned transcripts to- gether with speaker information, industry context, and struc- tured metadata for fine-grained evaluation. Unlike prior re- sources that emphasize scale or aggregate metrics, it is explic- itly constructed to reveal domain-specific and long-tail failure modes that are obscured by naturally distributed corpora. By combining long-form audio with industry-balanced evaluation, the benchmark enables systematic study of industry-aware ASR robustness and structured variation in financial speech. 2. Prior Work 2.1. Conversational and meeting speech benchmarks Earnings calls share several conversational properties with meeting and telephone speech, including rapid turn-taking, dis- fluencies, and speaker overlap.Standard benchmarks such as Switchboard (1992) [3], CallHome (1997) [4], and AMI (2007) [5] capture aspects of these phenomena, but generally lack the finance-specific terminology, dense numeracy, and do- main structure that significantly affect ASR performance on earnings calls. 2.2. Existing financial speech benchmarks Existing financial speech datasets fall short as comprehensive evaluation benchmarks in several key dimensions. SPGIS- peech, released in 2021 and expanded in 2025 [6, 7], provides over 5,000 hours of professionally transcribed earnings-call au- dio. However, it was designed primarily as a training corpus: the dataset includes very large train, development, and test splits (with approximately 2,000 hours in the test set), making it com- putationally expensive for controlled and reproducible bench- mark evaluation. The Earnings-21 (2021) and Earnings-22 (2022) bench- marks [8, 9] focus on open evaluation with aligned transcripts and accent- and country-level metadata. Earnings-22 provides approximately 160 hours of long-form global earnings calls, but lacks industry-balanced sampling: high-frequency sectors dom- inate evaluation metrics, obscuring performance on underrepre- sented domains. Moreover, neither dataset provides the struc- tured metadata—such as speaker roles, industry classification, or company identifiers—needed for fine-grained error analysis and speaker-aware evaluation. Overall, the financial speech domain lacks a comprehen- arXiv:2607.23813v1 [cs.CL] 26 Jul 2026 sive, balanced, and metadata-rich benchmark designed specif- ically for evaluating ASR systems under realistic earnings-call conditions. 3. Corpus Definition Earnings25 comprises two complementary test sets designed to support both long-form and segment-level evaluation, as summarized in Table 1. testset-full consists of 498 hours of complete earnings-call recordings from S&P 500 companies in 2025 Q4, preserving full conversational context with typical call durations of approximately one hour. testset-segmented is a curated 46-hour evaluation set sampled from over 2,000 U.S. earnings calls across 2025 (Q1–Q4), comprising 290 seg- ments—one per industry—to ensure balanced domain cover- age. Segments in testset-segmented are 5–10 minutes long. Both test sets include aligned transcripts and structured metadata and are constructed to preserve realistic conversa- tional flow, natural turn-taking, and cross-speaker dynamics. In addition to transcription, Earnings25 provides rich meta- data, including speaker segmentation, industry classification, and company identifiers, enabling fine-grained analysis beyond aggregate word error rate (WER) and supporting speaker- and structure-aware evaluation. 4. Corpus Generation We sample full earnings calls, apply CTC-based forced align- ment to obtain word-level timestamps, and aggregate shorter speech segments into 5–10 minute evaluation units. This sec- tion details the industry-stratified sampling, alignment, and seg- mentation procedures. 4.1. Industry-stratified sampling of the full earnings-call universe To construct a representative and industry-balanced evaluation set, we sample earnings calls from a pool of over 2,100 U.S. earnings calls spanning 2025 Q1–Q4. Sampling is performed using a two-stage, reproducible procedure designed to ensure broad industry coverage while preventing dominance by high- frequency sectors. In the first stage, we filter the corpus to retain only U.S.- domiciled companies within the target date range. In the sec- ond stage, we apply disproportionate stratified sampling based on industry classification. Calls are grouped by industry, and a fixed number of samples is drawn from each group, ensuring equal representation across industries regardless of their under- lying frequency in the corpus. When the number of industry groups exceeds the target evaluation size, a final random selec- tion is applied. All sampling steps are seeded for reproducibil- ity. 4.2. Alignment We perform forced alignment to obtain word-level timestamps, which enable the extraction of shorter segments from long earn- ings calls and thereby increase industry coverage. These time spans are also used to support speaker diarization and speaker- aware analysis. 4.2.1. CTC-based forced alignment Forced alignment is performed using Connectionist Tempo- ral Classification (CTC) models implemented in NVIDIA NeMo [10].CTC [11] defines a sequence-level objective that marginalizes over all monotonic alignments between input acoustic frames and output label sequences, producing frame- level posterior probabilities over output tokens and a blank sym- bol. These posteriors are used to align reference transcripts to audio and recover word-level timing information. Given a reference transcript tokenized into subword units, we compute frame-level CTC log-probabilities and apply Viterbi decoding to recover the most likely alignment path under the CTC lattice. Token-level timestamps are then ag- gregated into word boundaries. To mitigate alignment errors caused by trailing silence or low-confidence regions, we impose a maximum word-duration constraint of 0.5 seconds. The alignment pipeline consists of: (1) resampling au- dio to 16 kHz WAV format; (2) computing frame-level CTC log-probabilities; (3) Viterbi alignment between audio frames and reference tokens; and (4) word-boundary extraction from aligned tokens. 4.3. Segment extraction from full calls After obtaining word-level timestamps via forced alignment, we aggregate contiguous speech segments into non-overlapping blocks with durations constrained to a predefined range (e.g., 5– 10 minutes). Blocks are constructed greedily by accumulating consecutive segments until a minimum duration is reached and extending the block while the total duration remains below the maximum threshold. 4.3.1. Quality filtering We apply content-based filtering to remove blocks whose merged transcripts contain undesired patterns such as operator boilerplate or non-speech cues. From the remaining candidates, one block is sampled uniformly at random per call to avoid over- representation of individual calls. Prior to audio extraction, a fixed padding of 0.2 seconds is added to block boundaries to mitigate alignment jitter. This procedure produces a compact and diverse set of segments with clean transcripts, balanced across calls and suitable for segment-level ASR benchmarking. Speaker tags are used to further split the segments to ensure that segment boundaries do not bisect speaker turns. This proce- dure yields transcripts with reliable timing suitable for segment- level ASR evaluation, speaker-aware analysis, and downstream speech processing tasks. 5. Corpus Analysis 5.1. Geographic coverage and accent distribution testset-full includes earnings calls from all S&P 500 companies and spans a broad geographic footprint, covering English earn- ings calls from 12 countries. While U.S.-domiciled companies dominate the corpus (reflecting index composition), the dataset also includes calls from companies headquartered in Europe, Asia, and other regions. This geographic diversity introduces English accent variation and heterogeneous recording condi- tions that reflect real-world earnings-call audio. In contrast, testset-segmented restricts evaluation to U.S.- domiciled companies to reduce accent variability and provide a controlled evaluation setting. This design allows baseline ASR performance to be assessed under consistent linguistic condi- tions, while accent-robustness can be studied using the full-call test set. Table 1: Overview of the Earnings25 test sets. Propertytestset-fulltestset-segmented DescriptionFull earnings callsIndustry-stratified segments CoverageS&P 500 companiesU.S. companies Date range2025 Q4January–December 2025 Total audio498 hours46 hours Units ̃500 full calls290 segments (one per industry) Avg. duration58 minutes9.5 minutes Countries12 (US: 93.8%)1 (US only) Industries284 categories290 categories Sampling rate11–44.1 kHz (original)16 kHz Audio formatMP3WAV Table 2: Geographic distribution of testset-full. CountryNumber of calls United States482 Ireland8 United Kingdom6 India5 Switzerland4 Australia2 Bermuda2 Netherlands1 Mexico1 Mauritius1 China1 Canada1 5.2. Industry distribution Industry coverage is a core design consideration of Earnings25. testset-full spans 284 distinct industry categories, reflecting the natural sector distribution of the S&P 500. As summarized in Table 3, utilities and financial services are among the most frequently represented sectors, while the majority of industries appear only a small number of times, preserving the long-tail structure characteristic of financial data. To mitigate frequency bias in evaluation,testset- segmented is constructed via industry-stratified sampling from over 2,000 U.S. earnings calls spanning 2025 (Q1–Q4). All 290 segments are drawn exclusively from U.S.-domiciled companies to ensure consistent audio quality and reduce accent variability. The segmented test set spans 290 unique industry categories, with exactly one segment per industry, ensuring broad coverage of domain-specific vocabulary across diverse financial sub-domains—from 3D Printers and Adult Night- clubs to Wind Turbines and Wireline Telecom Equipment. This design enables fine-grained, industry-aware analysis of ASR performance while maintaining a controlled linguistic evaluation setting. 5.3. Speaking-style variation Earnings calls exhibit substantial variation in speaking style and interaction structure. Within a single call, speech alter- nates between scripted prepared remarks and spontaneous ana- lyst Q&A, often with rapid turn-taking and occasional overlap. testset-full preserves these dynamics across entire calls, while testset-segmented retains multi-speaker conversational struc- Table 3: Industry distribution of testset-full. Industry categoryNumber of calls Integrated Electric Utilities21 Financial Transaction Processors9 Biotech9 Banks9 Enterprise Software8 Crude Oil & Natural Gas E&P7 Apartment REIT7 Large Pharma6 Infrastructure Software6 Diversified Industrials6 Security & Cmdty Exchanges5 P&C Insurance Premiums5 Life Insurance5 Investment Management5 Other (270 categories)406 ture within each 5–10 minute segment. Speaker attribution metadata enables role-aware analysis of ASR performance across operators, executives, and ana- lysts. Operators typically deliver formulaic announcements, ex- ecutives present prepared remarks, and analysts contribute un- scripted questions. This variation allows evaluation of ASR ro- bustness across speaking styles, roles, and interaction patterns common in financial speech. 6. Transcription Experiments 6.1. Baseline Models and Evaluation Protocol We benchmark representative ASR systems spanning sequence- to-sequence and transducer architectures: OpenAI Whisper models [2] (base, medium, large-v2) and NVIDIA NeMo’s Parakeet-TDT-0.6B-v2 [12]. No external language model is used. Whisper inference: Decoding options are loaded from a fixed model configuration. We set a fixed random seed and use deter- ministic decoding settings. The language is set to English. No optional keyword prompting or boosting words are used during decoding. Parakeet-TDT inference: We use the pretrained checkpoint nvidia/parakeet-tdt-0.6b-v2 with greedy trans- ducer decoding and no external language model. Scoring and normalization: We report four consistent vari- Table 4: Baseline ASR results on the full earnings-call test set (498h) and the curated segmented test set (46h). Lower is better. testset-fulltestset-segmented ASR ModelWERWER-NWER-nc-npWER-N-nc-npWERWER-NWER-nc-npWER-N-nc-np Whisper-base0.178530.170680.115720.110430.193730.186040.12020.11606 Whisper-medium0.143990.136590.085940.081510.155360.149320.086760.08369 Whisper-large-v20.140390.134070.084220.080360.153330.147240.084590.08174 Parakeet-tdt-0.6b-v20.108370.103260.064130.061140.111090.104980.064730.06062 Table 5: Selected industry-level ASR performance of Parakeet-tdt-0.6b-v2 on testset-full (2–3 calls per subsector based on industry tags). This table is illustrative and not a comprehensive ranking. Lower is better. IndustryWERWER-NWER-nc-npWER-N-nc-np Biotech0.154400.147100.102650.09845 Pharma0.153430.148210.102550.10097 Crude Oil & Natural Gas E&P0.131310.125760.077590.07416 Health Care Software0.130330.123470.078490.07466 ants: WER (raw), WER-N (NeMo-normalized), WER-nc- np (lowercased, punctuation removed), and WER-N-nc-np (NeMo-normalized, lowercased, punctuation removed). For normalization-based variants, the same NeMo English text nor- malization [13] is applied to both reference and hypothesis. Reproducibility: All sampling and segmentation procedures use a fixed random seed (2025).Experimental scripts set PYTHONHASHSEED=2025 and framework RNG seeds, and log the full decoding configuration used for each run. 6.2. Results Table 4 reports baseline performance on testset-full (498 h) and the industry-balanced testset-segmented (46 h). Lower is better. Across models, performance on testset-segmented is slightly worse than on testset-full, despite shorter duration. This reflects the effect of industry stratification: while testset- full follows the natural frequency distribution of sectors, the segmented set enforces equal industry representation and in- creases exposure to long-tail terminology. To illustrate domain variability, Table 5 reports results for selected subsectors (2–3 calls each) using Parakeet-tdt-0.6b-v2. This is an illustrative subset and not a comprehensive indus- try ranking. Terminology-dense domains such as biotech and pharma show substantially higher WER (15.3–15.4%) than the aggregate testset-full WER (10.8%). Because the corpus-level metric is frequency-weighted across industries, high-volume and more repetitive sectors lower the overall average. In con- trast, the unweighted per-industry results emphasize more chal- lenging domains, demonstrating that aggregate WER can mask substantial domain-dependent variation. 7. Limitations and Conclusion Earnings25 provides value along three dimensions:(1) domain-specific evaluation, offering a challenging benchmark based on S&P 500 earnings calls with broad industry coverage and finance-specific terminology; (2) reproducible baselines, with standardized evaluations for contemporary ASR models, including Whisper and Parakeet-TDT; and (3) rich metadata, including industry, call-structure, and speaker annotations that enable stratified and speaker-aware analysis. Regarding limitations, Earnings25 focuses on English- language earnings calls and primarily reflects speech from U.S.- domiciled companies. While this design enables controlled and high-quality evaluation, it does not capture the full linguistic diversity of global earnings calls. Extending the benchmark to include multilingual earnings-call data across a broader range of countries and languages remains an important direction for future work. 8. Data Access and Licensing Earnings25 is released for research and benchmarking pur- poses.We redistribute the audio recordings, transcripts, metadata, annotations, and evaluation splits through Zenodo: https://doi.org/10.5281/zenodo.18762168. The transcripts, an- notations, metadata, evaluation splits, and alignments are re- leased under the Creative Commons Attribution 4.0 Interna- tional license (C BY 4.0). The redistributed audio recordings remain subject to any applicable terms of the original content providers, and users are responsible for ensuring compliance with those terms. 9. Generative AI Use Disclosure During preparation of this manuscript, the authors used a gener- ative AI assistant only for language editing and polishing (e.g., grammar and phrasing). The tool was not used to generate ex- perimental results, analyses, or conclusions, and no generative AI system is listed as an author. 10. References [1] A.Baevski,Y.Zhou,A.Mohamed,andM.Auli, “wav2vec 2.0:A framework for self-supervised learn- ingofspeechrepresentations,”inAdvancesinNeu- ral Information Processing Systems, vol. 33, 2020. [On- line]. Available: https://proceedings.neurips.c/paper/2020/hash/ 92d1e1eb1cd6f9fba3227870b6d7f07-Abstract.html [2] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large- scale weak supervision,” 2022. [Online]. Available:https: //arxiv.org/abs/2212.04356 [3] J. J. Godfrey, E. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in Proc. ICASSP 1992, 1992, p. 517–520. [4] “CALLHOME american english speech,” Linguistic Data Consortium (LDC), Catalog No. LDC97S42, 1997. [Online]. Available: https://catalog.ldc.upenn.edu/LDC97S42 [5] J. Carletta, “Unleashing the killer corpus: Experiences in creating the multimodal AMI meeting corpus,” Language Resources and Evaluation, vol. 41, no. 2, p. 181–190, 2007. [6] P. K. O’Neill, V. Lavrukhin, S. Majumdar, V. Noroozi, Y. Zhang, O. Kuchaiev, J. Balam, Y. Dovzhenko, K. Freyberg, M. D. Shul- man, B. Ginsburg, S. Watanabe, and G. Kucsko, “SPGISpeech: 5,000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition,” in Interspeech 2021, 2021, p. 1434–1438. [7] R. Grossman, T. Park, K. Dhawan, A. Titus, S. Zhi, Y. Shchadilova, W. Wang, J. Balam, and B. Ginsburg, “SPGIS- peech 2.0: Transcribed multi-speaker financial audio for speaker- tagged transcription,” in Interspeech 2025, 2025, p. 4048–4052. [8] M. Del Rio, N. Delworth, R. Westerman, M. Huang, N. Bhandari, J. Palakapilly, Q. McNamara, J. Dong, P. ̇ Zelasko, and M. Jett ́ e, “Earnings-21: A Practical Benchmark for ASR in the Wild,” in Interspeech 2021, 2021, p. 3465–3469. [9] M. D. Rio, P. Ha, Q. McNamara, C. Miller, and S. Chandra, “Earnings-22: A Practical Benchmark for Accents in the Wild,” 2022. [Online]. Available: https://arxiv.org/abs/2203.15591 [10] O. Kuchaiev, J. Li, B. Ginsburg et al., “NeMo: a toolkit for building AI applications using neural modules,” 2019. [Online]. Available: https://arxiv.org/abs/1909.09577 [11] A. Graves, S. Fern ́ andez, F. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,” in Proceedings of the 23rd International Conference on Machine Learning, 2006, p. 369–376. [12] H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” 2023. [Online]. Available: https://arxiv.org/abs/2304.06795 [13] Y. Zhang, E. Bakhturina, K. Gorman, and B. Ginsburg, “Nemo in- verse text normalization: From development to production,” arXiv preprint arXiv:2104.05055, 2021.