Paper deep dive
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
Yuefeng Zou, Yichen Lu, Jingxiao Yang, Bingtao Fu, Gaoyang Zhang, Xiongfei Bai, Tian Chen, Xiang Qi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/18/2026, 5:21:04 AM
Summary
The paper introduces LongDocBench, a benchmark designed to evaluate document-level structure recovery in long documents, specifically focusing on Table-of-Contents (TOC) Hierarchy Recovery and Contextual Relationship Recovery. The dataset comprises 85 real-world financial reports, textbooks, and academic papers totaling 2,582 pages, with human-verified annotations for 3,937 heading nodes and 3,258 contextual relationships. Experiments demonstrate that while existing parsers perform well on page-level tasks, they struggle significantly with document-level structure recovery. Furthermore, incorporating human-verified TOC hierarchies and contextual relationships improves long-document question-answering accuracy, with combined structures providing complementary benefits.
Entities (11)
Relation Signals (10)
LongDocBench → evaluatestask → TOC Hierarchy Recovery
confidence 95% · To benchmark these two tasks, we introduce LongDocBench... Table-of-Contents Hierarchy Recovery and Contextual Relationship Recovery.
LongDocBench → evaluatestask → Contextual Relationship Recovery
confidence 95% · To benchmark these two tasks, we introduce LongDocBench... Contextual Relationship Recovery.
TOC Hierarchy Recovery → usesmetric → TEDS
confidence 95% · We evaluate complete, ordered heading hierarchies using Tree Edit Distance Similarity (TEDS)...
GPT 5.6 Sol → achievesbestin → Contextual Relationship Recovery
confidence 90% · GPT-5.6-Sol achieves the highest Macro score of 0.63...
TextIn → achievesbestin → TOC Hierarchy Recovery
confidence 90% · TextIn performs best, achieving 0.49 Weighted TEDS and 0.55 Macro TEDS under the with-ignorable setting.
LongDocBench → containsdocuments → Financial Reports
confidence 90% · comprising 85 real-world financial reports, textbooks, and academic papers...
LongDocBench → containsdocuments → Textbooks
confidence 90% · comprising 85 real-world financial reports, textbooks, and academic papers...
LongDocBench → containsdocuments →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and figures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing benchmarks cannot directly evaluate two key document-level tasks: \emph{Table-of-Contents Hierarchy Recovery} and \emph{Contextual Relationship Recovery}. To benchmark these two tasks, we introduce \textsc{LongDocBench}, comprising 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships annotated across 2,680 table and figure objects. We further evaluate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual relationships improve reasoning, with their combination providing complementary benefits. Meanwhile, representative document parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release \textsc{LongDocBench} and its evaluation protocol and reproducible testbed for advancing document-level structure recovery in long documents.
Tags
Links
- Source: https://arxiv.org/abs/2608.15064v1
- Canonical: https://arxiv.org/abs/2608.15064v1
Trouble viewing inline? Open PDF directly →
Full Text
34,620 characters extracted from source content.
Expand or collapse full text
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents Yuefeng Zou 1∗ , Yichen Lu 1∗ , Jingxiao Yang 1,2∗ , Bingtao Fu 1 , Gaoyang Zhang 1 , Xiongfei Bai 1 , Tian Chen 1 , Xiang Qi 1 1 Ant Group 2 Zhejiang University Abstract Parsing visual documents into machine-readable representa- tions is fundamental to document intelligence. Existing bench- marks focus on page-level element recognition, reading order, formula recognition, and table structure. Long documents, however, also require document-level structure recovery. This includes reconstructing cross-page table-of-contents (TOC) hierarchies and identifying typed links from tables and fig- ures to their captions, notes, and sources, often in one-to-many form. Because these structures are covered only partially or subsumed within broader parsing protocols, existing bench- marks cannot directly evaluate two key document-level tasks: Table-of-Contents Hierarchy Recovery and Contextual Rela- tionship Recovery. To benchmark these two tasks, we intro- duce LongDocBench, comprising 85 real-world financial re- ports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes (mean node depth 3.55; maximum depth 9) and 3,258 contextual relationships anno- tated across 2,680 table and figure objects. We further eval- uate both the downstream utility and recoverability of these structures. Long-document question-answering experiments show that human-verified TOC hierarchies and contextual re- lationships improve reasoning, with their combination provid- ing complementary benefits. Meanwhile, representative doc- ument parsers remain limited on both recovery tasks despite strong page-level performance. To support further progress, we publicly release LongDocBench and its evaluation pro- tocol and reproducible testbed for advancing document-level structure recovery in long documents. 1 Introduction Document parsing aims to convert visual documents into machine-readable representations that support information extraction, retrieval-augmented generation, and domain- specific analysis (Xu et al. 2020; Lewis et al. 2020; Mathew, Karatzas, and Jawahar 2021). Advances in layout analysis, optical character recognition, and multimodal models have enabled increasingly reliable recognition of text, headings, formulas, tables, and figures on individual pages, as well as recovery of reading order and local layout structure (Huang et al. 2022; Wang et al. 2021; Blecher et al. 2024; Wang et al. 2024; Ouyang et al. 2025; Li et al. 2024; Zhou et al. 2026). Long documents, however, require capabilities be- yond page-level recognition. First, their content is often or- ∗ These authors contributed equally. Copyright © 2027, Association for the Advancement of Artificial Intelligence (w.aaai.org). All rights reserved. Figure 1: Representative challenges in document-level struc- ture recovery. (A) A parser preserves heading text but col- lapses the parent–child structure of the TOC hierarchy. (B) A source relationship is not recovered despite the presence of the figure and nearby contextual text. ganized across tens or hundreds of pages through deeply nested table-of-contents (TOC) hierarchies, as commonly seen in financial reports and textbooks (Bentabet et al. 2020; Ma et al. 2023; Zhang et al. 2024). An incorrect parent–child assignment changes not only a heading’s position in the hi- erarchy but also the sectional scope of the content governed by that heading. Second, tables and figures may be linked to captions, notes, and sources that are spatially distant, and a single object may be associated with multiple contextual el- ements (Smock, Pesala, and Abraham 2022; Xu et al. 2026). Therefore, Long-document parsing must consequently re- cover not only page-level elements, but also document-level TOC hierarchies and typed object-level contextual relation- ships from the document as a whole (Figure 1). Recent studies show that preserving multilevel document structure and recovering title hierarchies, cross-page con- ANT GROUP arXiv:2608.15064v1 [cs.AI] 15 Aug 2026 Figure 2: Overview of LongDocBench. (A) The benchmark contains complete long-document inputs from the finance, textbook, and paper domains. (B) Table-of-Contents Hierarchy Recovery reconstructs complete, ordered TOC hierarchies across pages. (C) Contextual Relationship Recovery identifies typed, potentially one-to-many links from tables and figures to their captions, notes, and sources. tinuity, and image–text associations improve downstream retrieval, generation, and question answering (Sarthi et al. 2024; Buchmann et al. 2024; Saad-Falcon et al. 2024; Lu et al. 2025; Xu et al. 2026). Existing benchmarks, how- ever, provide only partial diagnosis of such document or- ganization: general parsing benchmarks emphasize page- level recognition, reading order, and formula or table struc- ture, often through composite scores (Zhong, Tang, and Ji- meno Yepes 2019; Zhong, ShafieiBavani, and Jimeno Yepes 2020; Pfitzmann et al. 2022; Wang et al. 2021; Ouyang et al. 2025); multi-page benchmarks treat heading hierar- chy and cross-page merging as dimensions within broader protocols (Li et al. 2024; Zhou et al. 2026); and specialized post-processing evaluations cover title-hierarchy reconstruc- tion and image–text association without distinguishing cap- tion, note, and source relationships (Xu et al. 2026). Taken together, these limitations underscore the need for a unified, human-verified evaluation setting that directly assesses com- plete, ordered TOC hierarchies and provides type-specific evaluation of contextual relationships for tables and figures in real long documents. To enable more direct, fine-grained, and diagnostic eval- uation of document-level structure recovery, we carefully design LongDocBench around two complementary tasks (Figure 2): Table-of-Contents Hierarchy Recovery and Con- textual Relationship Recovery. LongDocBench contains 85 real-world financial reports, textbooks, and academic papers spanning 2,582 pages, with up to 105 pages per document. It provides human-verified annotations for 3,937 heading nodes, with a mean node depth of 3.55 and a maximum depth of 9, as well as 3,258 typed contextual relationships involving 2,680 table and figure objects. To examine whether these structures benefit long- document understanding, we conduct long-document question-answering experiments. The results show that both TOC hierarchies and contextual relationships improve question-answering accuracy, while their joint use yields further gains, indicating that they provide complementary organizational information. Although representative docu- ment parsers achieve strong and tightly clustered perfor- mance on conventional page-level parsing metrics, they re- main substantially weaker at document-level structure re- covery, reaching at best 0.55 Macro TEDS for TOC hier- archy recovery and a 0.63 Macro score for contextual rela- tionship recovery. These results show that strong page-level parsing performance does not yet translate into reliable re- covery of document-level TOC hierarchies and contextual relationships. The key contributions of this work are threefold: • We introduce and publicly release LongDocBench, a benchmark and unified evaluation protocol for two document-level structure-recovery tasks: Table-of- Contents Hierarchy Recovery and Contextual Relation- ship Recovery. It provides complete TOC hierarchies and type-specific contextual relationships for tables and fig- ures, supporting reproducible and fine-grained evalua- tion. • We evaluate the downstream utility of these structures Figure 3: Construction and annotation pipeline of LongDocBench. TextIn generates candidate TOC hierarchies, typed contextual relationships, and auxiliary cross-page continuations. These candidates undergo full-document manual correction, automated validation, and two-stage expert review, yielding human-verified annotations that are visually grounded and traceable to TextIn element IDs. through long-document question-answering experiments. Both TOC hierarchies and contextual relationships im- prove question-answering accuracy, while their joint use yields further gains, indicating that they provide comple- mentary organizational information. • We evaluate representative document parsers on both structure-recovery tasks and find that, despite strong page- level parsing performance, they remain limited in recov- ering complete TOC hierarchies and typed contextual re- lationships, indicating that both tasks remain challenging. 2 LongDocBench To evaluate document-level structure recovery in long- document parsing, we introduce LongDocBench, a bench- mark for two tasks: Table-of-Contents Hierarchy Recovery and Contextual Relationship Recovery. The former recon- structs headings distributed across pages into complete, or- dered TOC hierarchies, while the latter identifies typed links between tables or figures and their captions, notes, and sources. To ensure diversity and practical relevance, Long- DocBench comprises 85 real-world financial reports, text- books, and academic papers spanning 2,582 pages, covering complex hierarchical organization, heterogeneous layouts, and rich contextual relationships for tables and figures. 2.1 Data Collection We manually collect publicly accessible long-document PDFs from five source families: Chinese financial materials, international research reports, U.S. public company filings, textbooks, and academic papers. The collection includes pub- lic disclosures obtained through CNINFO and SEC EDGAR, together with scholarly articles identified through arXiv and DOI metadata (Shenzhen Securities Information Co., Ltd. 2026; U.S. Securities and Exchange Commission 2026; arXiv 2026; Crossref 2026). The first three are grouped into the finance domain, while textbooks and academic papers constitute the textbook and paper domains, respectively. Fi- nancial documents include annual and audit reports, prospec- tuses, debt-issuance announcements, regulatory filings, and institutional research reports, often featuring irregular lay- outs, deeply nested headings, and rich table/figure contextual relationships. Textbooks provide extended chapter structures across disciplines, whereas academic papers are generally shorter and contain regular sections with diverse formulas, tables, and figures. Each document retains its original page order to preserve document-level organization and cross- page dependencies. The collection contains 42 financial doc- uments, 18 textbooks, and 25 academic papers, totaling 85 documents and 2,582 pages, with up to 105 pages per doc- ument. It spans substantial variation in document length, layout, TOC depth, and table/figure contextual relationships. 2.2 Data Annotation We construct the annotations through three stages (Figure 3): Parser-Based Pre-annotation, Full-Document Manual Cor- rection, and Expert Quality Assurance and Annotation Re- lease. Manual correction is performed by trained annotators with backgrounds in document intelligence and visual doc- ument analysis, followed by automated validation and two- stage expert review by senior researchers with experience in document parsing, OCR, and structured document analysis. Parser-Based Pre-annotation. TextIn extracts page-level elements, including text, element types, page indices, IDs, and bounding boxes. Detected headings are organized into initial TOC hierarchies, while table and figure objects are assigned candidate caption, note, and source relationships. Figure 4: Dataset statistics of LongDocBench. (a) Document-length distributions; circles mark domain medi- ans. (b) Per-document TOC complexity. (c) Number of con- textual elements linked to each table or figure object. Potential cross-page continuations are also identified. All machine-generated candidates serve only as starting points for manual correction. Full-Document Manual Correction. Annotators inspect each complete document in a browser-based interface. They add or remove headings, correct heading text, order, depth, and coordinates, and mark ambiguous headings as ignorable. They also revise contextual links, classify them as captions, notes, or sources, and verify one-to-many, shared, and cross- page relationships. Expert Quality Assurance and Annotation Release. Au- tomated validation checks element IDs and coordinates, hi- erarchy consistency, relationship endpoints, duplicate links, type compatibility, and cross-page consistency. Senior re- Domain Docs Pages Avg. pages/doc. Max pages Finance 42 1,33131.69105 Textbook 18 88449.1176 Paper25 36714.6835 All85 2,58230.38105 Table 1: Document statistics by domain. Domain Heading nodes Mean node depth Max depth Finance2,5823.819 Textbook9103.316 Paper4452.495 All3,9373.559 Table 2: TOC hierarchy statistics by domain. searchers then conduct two-stage expert review, including a final document-level audit. Failed documents are revised and re-checked before release. The final annotations contain 3,937 human-verified heading nodes and 3,258 typed con- textual relationships annotated across 2,680 table and figure objects, together with auxiliary cross-page continuation la- bels. 2.3 Dataset Statistics Document Distribution. As shown in Table 1 and Fig- ure 4(a), LongDocBench contains 85 documents spanning 2,582 pages, with an average of 30.38 pages per document. The finance domain is the largest, comprising 42 documents and 1,331 pages, and contains the longest document at 105 pages. Although the textbook domain contains only 18 doc- uments, it has the greatest average length at 49.11 pages. Documents in the paper domain are comparatively compact, averaging 14.68 pages across 25 documents. The collec- tion therefore spans substantial variation in document length across domains. TOC Hierarchies. Table 2 summarizes the 3,937 anno- tated heading nodes. Depth is one-indexed, and mean depth is computed over all heading nodes within each domain. The finance domain contains 2,582 nodes with a mean depth of 3.81 and a maximum depth of 9, forming the largest and deepest TOC hierarchies in the benchmark. The textbook domain contains 910 nodes with a mean depth of 3.31 and a maximum depth of 6, while the paper domain contains 445 nodes with a mean depth of 2.49 and a maximum depth of 5. These distributions support evaluation across both shallow section structures and deeply nested hierarchies. Contextual Relationships. As reported in Table 3, the benchmark contains 1,169 tables and 1,511 figures, yield- ing 2,680 annotated objects. These objects are connected by 3,258 typed contextual relationships, including 1,703 cap- tion, 961 note, and 594 source relationships. Among them, 904 objects are linked to multiple contextual elements, di- rectly demonstrating the prevalence of one-to-many asso- ciations. Tables have more note relationships than figures CategoryTable Figure Total Objects1,169 1,511 2,680 Caption relationships 713 990 1,703 Note relationships587 374 961 Source relationships 206 388 594 All relationships1,506 1,752 3,258 Table 3: Object and contextual-relationship statistics. (587 versus 374), whereas figures have more source relation- ships (388 versus 206), reflecting different contextual pat- terns across object types. Together, these statistics provide diverse challenges for TOC hierarchy recovery and contex- tual relationship recovery. 3 Experiments 3.1 Evaluation Metrics TOC Hierarchy Recovery. We evaluate complete, or- dered heading hierarchies using Tree Edit Distance Simi- larity (TEDS), which jointly reflects errors in heading con- tent, order, and parent–child structure (Zhang and Shasha 1989; Zhong, ShafieiBavani, and Jimeno Yepes 2020). We report Macro TEDS, averaged equally across documents, and Weighted TEDS, weighted by document page count, under both with- and without-ignorable settings. Contextual Relationship Recovery. Predicted tables and figures are matched to same-type ground-truth objects on the same page based on bounding-box IoU. For each matched object, caption, note, and source relationships are evaluated using normalized text edit similarity (Levenshtein 1966). We report the score for each relationship type and their un- weighted mean as the final Macro score. Detailed normalization, matching, aggregation, and scor- ing procedures are provided in the appendix. 3.2 Experimental Setup We evaluate TOC hierarchy recovery on all 85 documents using TextIn, MinerU2.5-Pro, GLM-OCR, PaddleOCR-VL- 1.5, and PaddleOCR-VL-1.6, converting their native head- ing outputs into ordered trees through a unified construction procedure (Wang et al. 2024; Duan et al. 2026; Cui et al. 2026; Zhang et al. 2026). For contextual relationship re- covery, we fix TextIn parsing and benchmark GPT-5.6-Sol, MiniMax-M2.5, GLM-5.2, Kimi K2.6, and Qwen3.5 vari- ants on 2,680 benchmark-localized table and figure targets, using the parsed text and layout context to predict caption, note, and source relationships (MiniMax Team 2026; GLM- 5 Team 2026; Kimi Team 2025; Yang et al. 2025). Using the OmniDocBench v1.6 protocol, we evaluate the page-level parsing performance of LingDT-VL-OCR-4B, MinerU2.5- Pro, PaddleOCR-VL-1.5, PaddleOCR-VL-1.6, and GLM- OCR on all 2,582 pages of LongDocBench (Ouyang et al. 2025; Qian et al. 2026). To assess reasoning utility, we compare a structure-agnostic fixed-chunk BM25 baseline with structure-aware question answering (Lewis et al. 2020; ModelWeighted TEDS Macro TEDS w/ ign. w/o ign. w/ ign. w/o ign. TextIn0.49 0.47 0.55 0.53 MinerU2.5-Pro0.44 0.42 0.52 0.50 GLM-OCR0.45 0.43 0.52 0.48 PaddleOCR-VL-1.5 0.45 0.42 0.52 0.49 PaddleOCR-VL-1.6 0.44 0.42 0.51 0.48 Table 4: TOC hierarchy recovery on all 85 documents. We report page-count-weighted and document-macro TEDS un- der the with- and without-ignorable settings; higher is better. Robertson and Zaragoza 2009). The recovered-structure con- ditions use TOC hierarchies produced by TextIn and contex- tual relationships produced by Qwen3.5-397B with think- ing, while the verified conditions use human-verified Long- DocBench structures. All conditions use the same parsed content, questions, answer prompt, and GLM-5.2 answer model. Detailed experiment settings are provided in the ap- pendix. 3.3 TOC Hierarchy Recovery We first evaluate whether document parsers that perform well at page-level element recognition can reconstruct com- plete, ordered TOC hierarchies across pages. Table 4 reports the overall results, while Figure 5 compares performance across document domains. TextIn performs best, achieving 0.49 Weighted TEDS and 0.55 Macro TEDS under the with- ignorable setting. The remaining systems are closely clus- tered, with Weighted TEDS of 0.44–0.45 and Macro TEDS of 0.51–0.52. These scores indicate that recovering document- level heading organization remains difficult even for strong document parsers. Performance varies markedly across domains. Macro TEDS reaches 0.69–0.82 on papers but only 0.44–0.48 on financial documents and 0.37–0.46 on textbooks. This pat- tern is consistent with the dataset statistics: papers are gen- erally shorter and have shallower section structures, whereas financial documents and textbooks contain longer and more deeply nested hierarchies. Weighted TEDS is consistently lower than Macro TEDS, further indicating that recovery quality deteriorates as document length increases. Accounting for ignorable headings improves Macro TEDS by only 0.02–0.04 and Weighted TEDS by 0.02–0.03, with- out materially changing system rankings or domain-level trends. The remaining errors therefore arise primarily from failures to recover heading order, depth, and parent–child structure rather than from optional or ambiguous headings. 3.4 Contextual Relationship Recovery We next isolate contextual relationship recovery from object detection by providing all models with the same TextIn pars- ing outputs and benchmark-localized table and figure objects. As shown in Table 5, GPT-5.6-Sol achieves the highest Macro score of 0.63, followed by Qwen3.5-397B with thinking at 0.60. No model dominates all relation types: GPT-5.6-Sol performs best on captions and notes, whereas Qwen3.5-397B Figure 5: TOC hierarchy recovery by domain. Panels report (a) Weighted TEDS and (b) Macro TEDS. For each model and domain, the with-ignorable score is plotted upward and the without-ignorable score is mirrored downward; downward bars therefore encode nonnegative magnitudes rather than negative TEDS. Higher magnitudes indicate better recovery. Recovery modelCaption Note Source Macro GPT-5.6-Sol (xhigh)0.65 0.42 0.81 0.63 MiniMax-M2.5 (think)0.53 0.17 0.61 0.44 GLM-5.2 (think)0.64 0.30 0.74 0.56 Kimi K2.6 (w/o think)0.62 0.22 0.82 0.55 Qwen3.5-397B (w/o think) 0.61 0.30 0.85 0.59 Qwen3.5-397B (think)0.64 0.33 0.82 0.60 Qwen3.5-35B (w/o think) 0.57 0.20 0.78 0.52 Qwen3.5-9B (w/o think)0.59 0.16 0.80 0.52 Table 5: Contextual relationship recovery using common TextIn outputs and benchmark-localized objects. Columns score individual relationship types; Macro is their un- weighted mean. Higher is better. without thinking achieves the highest source score. The lim- ited best Macro score shows that reliable contextual relation- ship recovery remains difficult even when object locations and page-level parsing are fixed. The difficulty also differs substantially by relationship type. Source relationships obtain the highest scores, rang- ing from 0.61 to 0.85, followed by captions at 0.53–0.65. Notes are consistently the most challenging, with scores of only 0.16–0.42. Enabling thinking for Qwen3.5-397B raises its Macro score only marginally, from 0.59 to 0.60, and does not resolve its weakness on notes. These results identify note recovery as the principal bottleneck, reflecting the greater variability and weaker spatial regularity of note relationships. 3.5 Document Parsing Results We further evaluate whether strong conventional parsing per- formance is accompanied by reliable document-level struc- ture recovery. Table 6 reports page-level parsing results un- der the OmniDocBench v1.6 protocol (Ouyang et al. 2025; Zhong, ShafieiBavani, and Jimeno Yepes 2020). All five sys- tems achieve strong and closely clustered Overall scores of 93.54–95.96. MinerU2.5-Pro obtains the best Overall score and leads on text, formula, and table parsing, while LingDT- VL-OCR-4B achieves the lowest reading-order error. For- mula scores vary little across systems, whereas table parsing exhibits somewhat larger differences. These results contrast sharply with performance on the two document-level tasks. Despite Overall page-level scores above 93.5, the best systems reach only 0.55 Macro TEDS for TOC hierarchy recovery and 0.63 Macro for contextual rela- tionship recovery. Strong recognition of page elements and local structure therefore does not imply reliable reconstruc- tion of cross-page heading hierarchies or typed table/figure relationships. Document-level structure recovery remains a distinct capability that is not captured by conventional page- level parsing metrics. 3.6 Impact on Long-Document Reasoning We compare the structure-agnostic Fixed + BM25 baseline with conditions using recovered or human-verified TOC hier- archies and contextual relationships. All conditions share the same parsed document content and answer model. Table 7 reports the test-set results. Human-verified structures provide substantial gains over the Fixed + BM25 baseline of 32.98%. Verified TOC and Verified Relations achieve 45.95% and 47.62%, improving accuracy by 12.97 and 14.64 percentage points, respectively. Combining both structures yields the best result of 54.29%, a gain of 21.31 points over the baseline and additional gains of 8.34 and 6.67 points over Verified TOC and Verified Rela- tions alone. These results indicate that TOC hierarchies and ModelText↓ Formula (CDM)↑ Table (TEDS)↑ Table (TEDS-S)↑ Reading Order↓ Overall↑ LingDT-VL-OCR-4B 0.0295.3791.7593.900.1095.07 MinerU2.5-Pro0.0195.7893.3394.750.1195.96 PaddleOCR-VL-1.50.0495.3689.1591.450.1393.54 PaddleOCR-VL-1.60.0495.5089.6792.120.1393.67 GLM-OCR0.0395.7188.7691.250.1493.86 Table 6: Page-level document parsing on all 2,582 pages under the OmniDocBench v1.6 protocol. Down arrows denote error metrics (lower is better), whereas up arrows denote similarity or accuracy metrics (higher is better). ConditionAnswer Acc. Fixed + BM2532.98 TextIn TOC30.83 Verified TOC45.95 Qwen3.5 Relations31.90 Verified Relations47.62 TextIn TOC + Qwen3.5 Relations37.86 Verified TOC + Relations54.29 Table 7: Test-set Combined Answer Accuracy (%). Fixed + BM25 is the structure-agnostic chunk baseline. TextIn TOC and Qwen3.5 Relations use recovered structures, whereas Verified conditions use human-verified structures. All con- ditions share the same parsed content and answer model. The best result is shown in bold. contextual relationships contribute complementary organi- zational information to long-document reasoning. Automatically recovered structures realize only part of this potential. TextIn TOC and Qwen3.5 Relations obtain 30.83% and 31.90%, neither surpassing the fixed-chunk baseline, while their combination reaches 37.86%, exceeding the base- line by 4.88 points. The recovered conditions remain 15.12, 15.72, and 16.43 points below their verified TOC, relation, and combined counterparts, respectively. Thus, the verified results establish the utility of the two structures, whereas the recovered-to-verified gaps show that current recovery quality remains insufficient to realize their full downstream benefit. 4 Conclusion In this paper, we introduced LongDocBench, a human- verified benchmark for two document-level structure- recovery tasks: Table-of-Contents Hierarchy Recovery and Contextual Relationship Recovery. It contains 85 real-world documents spanning 2,582 pages, with 3,937 heading nodes and 3,258 typed contextual relationships annotated across 2,680 table and figure objects. We further demonstrate the downstream value of these structures: human-verified TOC hierarchies and contextual relationships jointly improve test- set question-answering accuracy from 32.98% with fixed- chunk BM25 retrieval to 54.29%. Finally, representative parsers achieve only 0.55 Macro TEDS and a 0.63 Macro score on the two recovery tasks despite strong page-level performance, while automatically recovered structures pro- vide much smaller reasoning gains. Although limited in scale and coverage, LongDocBench establishes document-level organization recovery as a useful, insufficiently addressed capability. References arXiv. 2026. About arXiv. https://info.arxiv.org/about/. Ac- cessed July 2026. Bentabet, N.-I.; Juge, R.; El Maarouf, I.; Mouilleron, V.; Valsamou-Stanislawski, D.; and El-Haj, M. 2020. The Finan- cial Document Structure Extraction Shared Task: FinTOC 2020. In Proceedings of the 1st Joint Workshop on Financial Narrative Processing and MultiLing Financial Summarisa- tion, 13–22. COLING. Blecher, L.; Cucurull, G.; Scialom, T.; and Stojnic, R. 2024. Nougat: Neural Optical Understanding for Academic Docu- ments. In International Conference on Learning Represen- tations. Buchmann, J.; Eichler, M.; Bodensohn, J.-M.; Kuznetsov, I.; and Gurevych, I. 2024. Document Structure in Long Docu- ment Transformers. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), 1056–1073. Associa- tion for Computational Linguistics. Crossref. 2026. Metadata Retrieval. https://w.crossref. org/documentation/retrieve-metadata/. Accessed July 2026. Cui, C.; Sun, T.; Liang, S.; Gao, T.; Zhang, Z.; Liu, J.; Wang, X.; Zhou, C.; Liu, H.; Lin, M.; Zhang, Y.; Zhang, Y.; Liu, Y.; Yu, D.; and Ma, Y. 2026. PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing. arXiv preprint arXiv:2601.21957. Duan, S.; Xue, Y.; Wang, W.; Su, Z.; Liu, H.; Yang, S.; Gan, G.; Wang, G.; Wang, Z.; Yan, S.; Jin, D.; Zhang, Y.; Wen, G.; Wang, Y.; Zhang, Y.; Zhang, X.; Hong, W.; Cen, Y.; Yin, D.; Chen, B.; Yu, W.; Gu, X.; and Tang, J. 2026. GLM-OCR Technical Report. arXiv preprint arXiv:2603.10910. GLM-5 Team. 2026. GLM-5: From Vibe Coding to Agentic Engineering. arXiv preprint arXiv:2602.15763. Huang, Y.; Lv, T.; Cui, L.; Lu, Y.; and Wei, F. 2022. Lay- outLMv3: Pre-Training for Document AI with Unified Text and Image Masking. In Proceedings of the 30th ACM Inter- national Conference on Multimedia, 4083–4091. ACM. Kimi Team. 2025. Kimi K2: Open Agentic Intelligence. arXiv preprint arXiv:2507.20534. Levenshtein, V. I. 1966. Binary Codes Capable of Correcting Deletions, Insertions and Reversals. Soviet Physics Doklady, 10(8): 707–710. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; Riedel, S.; and Kiela, D. 2020. Retrieval-Augmented Gen- eration for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, volume 33, 9459– 9474. Li, Z.; Abulaiti, A.; Lu, Y.; Chen, X.; Zheng, J.; Lin, H.; Han, X.; and Sun, L. 2024. READoc: A Unified Benchmark for Realistic Document Structured Extraction. arXiv preprint arXiv:2409.05137. Lu, W.; Chen, K.; Qiao, R.; and Sun, X. 2025. HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking. arXiv preprint arXiv:2509.11552. Ma, J.; Du, J.; Hu, P.; Zhang, Z.; Zhang, J.; Zhu, H.; and Liu, C. 2023. HRDoc: Dataset and Baseline Method toward Hierarchical Reconstruction of Document Structures. In Pro- ceedings of the AAAI Conference on Artificial Intelligence, volume 37, 1870–1877. Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021. DocVQA: A Dataset for VQA on Document Images. In Proceedings of the IEEE/CVF Winter Conference on Appli- cations of Computer Vision, 2200–2209. MiniMax Team. 2026. The MiniMax-M2 Series: Mini Ac- tivations Unleashing Max Real-World Intelligence. arXiv preprint arXiv:2605.26494. Ouyang, L.; Qu, Y.; Zhou, H.; Zhu, J.; Zhang, R.; Lin, Q.; Wang, B.; Zhao, Z.; Jiang, M.; Zhao, X.; Shi, J.; Wu, F.; Chu, P.; Liu, M.; Li, Z.; Xu, C.; Zhang, B.; Shi, B.; Tu, Z.; and He, C. 2025. OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24838–24848. Pfitzmann, B.; Auer, C.; Dolfi, M.; Nassar, A. S.; and Staar, P. W. J. 2022. DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis. arXiv preprint arXiv:2206.01062. Qian, S.; Bai, X.; Fu, B.; Lu, Y.; Zhang, G.; Yang, X.; and Zhang, P. 2026. LingDT-VL-OCR: Structure-Aware Document-Level Parsing with Fine-Grained Visual Refer- ence. arXiv preprint arXiv:2603.11044. Robertson, S.; and Zaragoza, H. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4): 333–389. Saad-Falcon, J.; Barrow, J.; Siu, A.; Nenkova, A.; Yoon, S.; Rossi, R. A.; and Dernoncourt, F. 2024. PDFTriage: Question Answering over Long, Structured Documents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, 153–169. Association for Computational Linguistics. Sarthi, P.; Abdullah, S.; Tuli, A.; Khanna, S.; Goldie, A.; and Manning, C. D. 2024. RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. In International Conference on Learning Representations. Shenzhen Securities Information Co., Ltd. 2026. CNINFO: Listed-Company Disclosure Platform. https://w.cninfo. com.cn/. Accessed July 2026. Smock, B.; Pesala, R.; and Abraham, R. 2022. PubTables- 1M: Towards Comprehensive Table Extraction from Un- structured Documents. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4634–4642. U.S. Securities and Exchange Commission. 2026. EDGAR: Search Filings. https://w.sec.gov/search-filings. Ac- cessed July 2026. Wang, B.; Xu, C.; Zhao, X.; Ouyang, L.; Wu, F.; Zhao, Z.; Xu, R.; Liu, K.; Qu, Y.; Shang, F.; Zhang, B.; Wei, L.; Sui, Z.; Li, W.; Shi, B.; Qiao, Y.; Lin, D.; and He, C. 2024. MinerU: An Open-Source Solution for Precise Document Content Extraction. arXiv preprint arXiv:2409.18839. Wang, Z.; Xu, Y.; Cui, L.; Shang, J.; and Wei, F. 2021. LayoutReader: Pre-Training of Text and Layout for Reading Order Detection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 4735– 4744. Association for Computational Linguistics. Xu, B.; Miao, Z.; Zhou, X.; Lin, Y.; Tang, Z.; Zhao, X.; Wu, F.; Tan, C.; Wang, B.; and He, C. 2026. MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing. arXiv preprint arXiv:2605.24973. Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020. LayoutLM: Pre-Training of Text and Layout for Docu- ment Image Understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1192–1200. ACM. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; et al. 2025. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388. Zhang, K.; and Shasha, D. 1989. Simple Fast Algorithms for the Editing Distance between Trees and Related Problems. SIAM Journal on Computing, 18(6): 1245–1262. Zhang, Y.; Zhang, Z.; Lai, W.; Zhang, C.; Gui, T.; Zhang, Q.; and Huang, X. 2024. PDF-to-Tree: Parsing PDF Text Blocks into a Tree. In Findings of the Association for Computational Linguistics: EMNLP 2024, 10704–10714. Association for Computational Linguistics. Zhang, Z.; Liu, H.; Liang, S.; Zhang, Y.; Xiang, Y.; Liu, J.; Sun, T.; Lin, M.; Zhang, Y.; Zhou, C.; Gao, T.; Cui, C.; Liu, Y.; Yu, D.; and Ma, Y. 2026. PaddleOCR-VL-1.6: Expand- ing the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training. arXiv preprint arXiv:2606.03264. Zhong, X.; ShafieiBavani, E.; and Jimeno Yepes, A. 2020. Image-Based Table Recognition: Data, Model, and Evalua- tion. In Computer Vision – ECCV 2020, 564–580. Springer. Zhong, X.; Tang, J.; and Jimeno Yepes, A. 2019. PubLayNet: Largest Dataset Ever for Document Layout Analysis. In 2019 International Conference on Document Analysis and Recognition (ICDAR), 1015–1022. IEEE. Zhou, B.; Xing, H.; Chen, Y.; Xu, J.; Zheng, Q.; Gao, F.; Yang, Z.; Bai, S.; Yan, M.; Ye, J.; and Xie, H. 2026. MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing. arXiv preprint arXiv:2605.22100.