Paper deep dive
Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining
Mahmoud Amiri, Jamile Mohammad Jafari, Sara Mostafapour, Thomas Bocklitz
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/15/2026, 1:40:00 AM
Summary
Lit2Vec is a reproducible workflow for constructing a legally screened, chemistry-specific full-text corpus from the Semantic Scholar Open Research Corpus (S2ORC). It utilizes multi-source license screening (Unpaywall, OpenAlex, Crossref), token-aware paragraph chunking, and dense embeddings (intfloat/e5-large-v2) to support downstream retrieval-augmented generation (RAG) and text-mining applications.
Entities (5)
Relation Signals (3)
Lit2Vec ā usesdataset ā S2ORC
confidence 100% Ā· We used the S2ORC dataset as the foundational source for constructing our corpus
Lit2Vec ā usesmodel ā intfloat/e5-large-v2
confidence 100% Ā· paragraph-level embeddings generated with the intfloat/e5-large-v2 model
Lit2Vec ā screenslicenseusing ā Unpaywall
confidence 95% Ā· Licensing was screened using metadata from Unpaywall, OpenAlex, and Crossref
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an internal study corpus of 582,683 chemistry-specific full-text research articles with structured full text, token-aware paragraph chunks, paragraph-level embeddings generated with the intfloat/e5-large-v2 model, and record-level metadata including abstracts and licensing information. To support downstream retrieval and text-mining use cases, an eligible subset of the corpus was additionally enriched with machine-generated brief summaries and multi-label subfield annotations spanning 18 chemistry domains. Licensing was screened using metadata from Unpaywall, OpenAlex, and Crossref, and the resulting corpus was technically validated for schema compliance, embedding reproducibility, text quality, and metadata completeness. The primary contribution of this work is a reproducible workflow for corpus construction and validation, together with its associated schema and reproducibility resources. The released materials include the code, reconstruction workflow, schema, metadata/provenance artifacts, and validation outputs needed to reproduce the corpus from pinned public upstream resources. Public redistribution of source-derived text and broad text-derived representations is outside the scope of the general release. Researchers can reproduce the workflow by using the released pipeline with publicly available upstream datasets and metadata services.
Tags
Links
- Source: https://arxiv.org/abs/2604.12498v1
- Canonical: https://arxiv.org/abs/2604.12498v1
Trouble viewing inline? Open PDF directly ā
Full Text
147,964 characters extracted from source content.
Expand or collapse full text
Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining Mahmoud Amiri a,b , Jamile Mohammad Jafari a,b , Sara Mostafapour a,b , Thomas Bocklitz (corresponding author) a,b a Leibniz Institute of Photonic Technology, Member of Leibniz Health Technologies, Member of the Leibniz Centre for Photonics in Infection Research (LPI), Albert-Einstein-Strasse 9, 07745 Jena, Germany. b Institute of Physical Chemistry (IPC) and Abbe Center of Photonics (ACP), Friedrich Schiller University Jena, Member of the Leibniz Centre for Photonics in Infection Research (LPI), Helmholtzweg 4, 07743 Jena, Germany We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an internal study corpus of 582,683 chemistry-specific full-text research arti- cles with structured full text, token-aware paragraph chunks, paragraph-level embeddings generated with the intfloat/e5-large-v2 model, and record-level metadata including abstracts and licensing information. To support downstream retrieval and text-mining use cases, an eligible subset of the corpus was addi- tionally enriched with machine-generated brief summaries and multi-label subfield annotations spanning 18 chemistry domains. Licensing was screened using metadata from Unpaywall, OpenAlex, and Crossref, and the resulting corpus was technically validated for schema compliance, embedding reproducibility, text quality, and metadata completeness. The primary contribution of this work is a reproducible work- flow for corpus construction and validation, together with its associated schema and reproducibility resources. The released materials include the code, reconstruction workflow, schema, metadata/prove- nance artifacts, and validation outputs needed to reproduce the corpus from pinned public upstream resources. Public redistribution of source-derived text and broad text-derived representations is outside the scope of the general release. Researchers can reproduce the workflow by using the released pipeline with publicly available upstream datasets and metadata services. Scientific Contribution: This work advances cheminformatics by introducing a reproducible workflow for constructing and validating a chemistry full-text corpus from public upstream resources using con- servative, metadata-based license screening. In contrast to prior chemistry text resources that are often abstract-only, insuļ¬iciently documented, or not ready for semantic retrieval, Lit2Vec integrates full-text reconstruction, multi-source license screening, token-aware paragraph chunking, dense embeddings, and structured workflow outputs in a single transparent framework. The resulting workflow, schema, meta- data/provenance artifacts, and validation resources provide a practical foundation for workflow-level reproducibility in chemistry literature retrieval and text mining under clearly documented access and licensing constraints. Keywords: Reproducible workflow, chemistry corpus, semantic literature mining, retrieval-augmented generation, scientific text resource, semantic retrieval 1 arXiv:2604.12498v1 [cs.DB] 14 Apr 2026 List of abbreviations Abbr.MeaningAbbr.Meaning ANNApproximate nearest neighborMAGMicrosoft Academic Graph BARTBidirectional and Auto-Regressive Transformers MLPMultilayer perceptron BERTScore Bidirectional Encoder Representations from Transformers Score MMR Maximal Marginal Relevance CCCreative CommonsNLPNatural language processing C0Creative Commons ZeroOAOpen access C-BYCreative Commons AttributionRAGRetrieval-augmented generation C-BY- NC Creative Commons Attribution- NonCommercial ReLURectified linear unit C-BY- NC-ND Creative Commons Attribution- NonCommercial-NoDerivatives ROUGE Recall-Oriented Understudy for Gisting Evaluation C-BY- NC-SA Creative Commons Attribution- NonCommercial-ShareAlike RTX Ray tracing texel eXtreme C-BY- ND Creative Commons Attribution- NoDerivatives S2AGSemantic Scholar Academic Graph C-BY-SA Creative Commons Attribution- ShareAlike S2ORCSemantic Scholar Open Research Corpus CUDACompute Unified Device Architecture SMILES Simplified Molecular Input Line Entry System DOIDigital Object IdentifierSPDXSoftware Package Data Exchange FAIRFindable, Accessible, Interoperable, and Reusable TL;DRToo long; didnāt read FAISSFacebook AI Similarity SearchJSONLJSON Lines F1F1 scoreL2Euclidean norm fp1616-bit floating pointLLMLarge language model GPT-4Generative Pre-trained Transformer 4 InChIInternational Chemical Identifier GPT-4o Generative Pre-trained Transformer 4 Omni IQRInterquartile range 1 Introduction Recent progress in computational chemistry increasingly depends on the ability to analyze large bodies of scientific text with large language models (LLMs) and modern natural language pro- cessing (NLP) methods. In this setting, chemistry is not limited by a lack of published knowl- edge, but by the diļ¬iculty of converting a vast and heterogeneous literature into resources that are machine-readable, semantically structured, and legally reusable for downstream analysis. This challenge is particularly acute because the chemistry literature is both large and opera- tionally fragmented. Although automated techniques such as information extraction, literature mining, and retrieval-augmented generation (RAG) offer a path beyond manual review, their scientific usefulness depends on reproducible full-text corpora with standardized structure, rich metadata, and reviewable licensing provenance. Existing open collections only partially satisfy these requirements: much of the literature remains inaccessible, and even available full text is 2 Figure 1: Overview of the Lit2Vec workflow. The primary pipeline constructs a legally screened chemistry corpus from S2ORC through domain filtering, text normalization, chunking, embedding, and license screening. Optional enrichment modules are applied to records with eligible abstracts to generate TL;DR summaries and subfield annotations. often not organized at the level of semantic segmentation and legal clarity needed for robust downstream workflows. Prior chemistry corpora have supported tasks such as named entity recognition, question an- swering, and pretraining, but they do not fully satisfy the requirements of retrieval-oriented and reproducible chemistry NLP workflows. Common limitations include incomplete full-text cov- erage, lack of paragraph-level segmentation and dense embeddings for retrieval, heterogeneous licensing constraints, and inconsistent metadata normalization across chemistry subdomains. In addition, broad scientific corpora are not chemistry-structured, while large chemistry datasets primarily optimized for pretraining rather than real-time retrieval or question answering. A detailed review appears in Supplementary Section7.1. To address this gap, we introduce Literature to Vector (Lit2Vec) as a reproducible work- flow for constructing a chemistry corpus for downstream retrieval and text-mining applications using conservative, metadata-based license screening. Using Semantic Scholar Open Research Corpus (S2ORC) as the upstream source, we assembled a study corpus of 582,683 full-text research articles through domain-based selection and rigorous screening checks. Licensing was screened by cross-referencing metadata from OpenAlex [ 1], Unpaywall [2], and Crossref [3], with inclusion in the study corpus requiring metadata agreement under the study rule between at least two sources. The resulting screened corpus underpins the analyses reported here, and the released code, schema, manifests, and reconstruction workflow enable third parties to repro- duce it from pinned public upstream sources. This screening provides a conservative, auditable basis for corpus construction by harmonizing license metadata from three independent services and excluding conflicting or insuļ¬iciently supported cases. It therefore improves transparency and reproducibility for workflow construction and auditing, while downstream redistribution 3 decisions still depend on the accuracy and version-specific relevance of upstream records. To our knowledge, Lit2Vec is the first chemistry-focused, reproducible workflow to recon- struct a legally screened full-text corpus from public upstream resources and organize it into a retrieval-ready framework with token-aware paragraph segmentation, dense embeddings, struc- tured metadata, validation artifacts, and chemistry-specific enrichment. Within this framework, articles are normalized into structured Markdown, segmented into approximately 28.8 million semantically coherent paragraph-level chunks, and embedded with the intfloat/e5-large-v2 [4] model to produce 1024-dimensional dense vector representations. For a workflow overview, see Figure1. For downstream artificial intelligence (AI) tasks such as RAG, document classification, scientific recommendation, and fact-grounded summarization, a large subset of records with eligible abstracts is further enriched with abstracts, abstract-level embeddings, LLM-generated Too Long; Didnāt Read (TL;DR) summaries, and multi-label topic annotations. A detailed comparison of Lit2Vec with existing chemistry-related and general- purpose corpora is provided in Table5. Lit2Vec supports a wide range of use cases across scientific research and infrastructure: ā¢RAG: High-recall paragraph-level retrieval enables grounding LLMs in trusted scientific evidence, supporting applications such as chemical safety analysis and reasoning systems (demonstrated in Section3.15.1). ā¢Semantic search and recommendation: Vector-based similarity enables discovery of related papers, techniques, or materials. We showcase a Facebook AI Similarity Search (FAISS)- powered recommender system in Section3.15.2. ā¢Model training and fine-tuning: The released workflow, together with locally reconstructed records, can support training custom language and embedding models. ā¢Knowledge graph construction: Locally reconstructed Lit2Vec embeddings and metadata can be integrated with external resources (e.g., PubChem [ 5], ChEMBL [6], and patent databases) to enhance chemical ontologies and build structured knowledge graphs. ā¢Benchmarking and evaluation: The corpus supports evaluation of LLM tasks including document retrieval, summarization, classification, question answering, hallucination de- tection, and citation fidelity. ā¢Trends analysis: With rich metadata and time-stamped documents, Lit2Vec enables trend analysis of materials, methods, or topics, for example the evolving use of gold versus silver in surface-enhanced Raman spectroscopy (SERS) (see Section 3.15.3). We therefore present Lit2Vec not simply as a static resource release, but as a reproducible foundation for chemistry-focused text infrastructure from public upstream sources. The re- leased code, schema, manifests, metadata/provenance artifacts, and validation resources enable transparent reporting, reproducible analysis, and practical downstream use, while broad public redistribution of source-derived text and text-derived representations remains outside the scope of the general release. 4 Table 2: Overview of the Lit2Vec pipeline, showing inputs, key processing steps, retained records, and retention rates relative to explicitly defined reference stages. Steps 9 and 10 (optional enrichment) were applied only to the subset of records with eligible abstracts. StepKey ProcessingRecords Re- tained % of Refer- ence Stage 1. Data AcquisitionDownload JSONL dump200,000,000 ā 2. Domain FilteringRestrict to fields_of_study = Chemistry18,621,0449.3% of S1 3. Abstract AlignmentJoin with abstract corpus via corpus_id8,660,97946.5% of S2 4. Full-Text AlignmentJoin with full-text corpus via corpus_id1,074,6385.8% of S2 5. Markdown Structuring Parse annotations; normalize section headers966,94689.9% of S4 6. ChunkingToken-aware text segmentation963,69799.7% of S5 7. Embedding Generation Encode with intfloat/e5-large-v2963,697100% of S6 8. License ScreeningMerge and verify license metadata across sources 582,68360.4% of S7 9. Topic ClassificationAssign to 18 chemistry subfields460,39779.0% of S8 10. TL;DR Summarization Generate two-sentence summary460,39779.0% of S8 2 Methods In this study, reproducibility refers to deterministic regeneration of the screened study cor- pus and reported pipeline outputs from the same pinned upstream release using the released code, schema, archived manifests, frozen license metadata, and documented procedures. Thus, reproducibility is defined at the workflow level through regeneration from a fixed upstream snap- shot, rather than through unrestricted redistribution of all source-derived content. A stepwise summary of the pipeline is given in Table2. 2.1 Data Collection We used the S2ORC dataset [7] as the foundational source for constructing our corpus, specifi- cally accessing the Semantic Scholar Academic Graph (S2AG) 2024-12-31 release via the oļ¬icial datasets API. This release includes structured metadata for approximately 200 million publica- tions, abstracts for 100 million, and full-text parsed content for nearly 10 million open-access papers. We selected S2ORC for its scale, disciplinary breadth, and rich structural annotations that support downstream natural language processing (NLP) tasks. To automate data acquisition, we developed a custom Python API client that interacts with the Semantic Scholar dataset API to retrieve metadata for the latest corpus release and generate download links for the papers, abstracts, and s2orc datasets. All files are distributed in compressed .json.gz format. We decompressed the papers dataset into JSONL (JSON Lines) format and ingested it into a MongoDB [ 8] database for eļ¬icient querying. Chemistry-related papers were identified directly from the fields_of_study metadata by selecting records labeled as Chemistry by Semantic Scholarās classification model, providing a clear and consistent basis for corpus construction. We aligned abstracts and full-text records to the filtered chemistry subset using the shared corpus_id field. Abstracts were decompressed, ingested via the same pipeline as the main metadata, and matched to chemistry-labeled papers. The s2orc full-text dataset was processed in parallel using the same method. Full-text entries were matched against the chemistry subset using corpus_id. Only valid matches were retained for downstream processing. Corpus yields and coverage outcomes are reported in Section 3.1. 5 2.2 Pre-processing Each JavaScript Object Notation (JSON) file in the full-text dataset contains raw text (con- tent.text) and character-level annotations (content.annotations) that identify titles, abstracts, section headers, and paragraphs using start/end positions. Annotations stored as strings are parsed into dictionaries. We extract relevant spans, apply light Markdown formatting, and sort them by document order. The converted output follows a standardized Markdown structure: the first title becomes a level-1 header (#), the abstract is added under a level-2 header (##), section headers are retained in order, and paragraphs are grouped with the nearest section. Content appearing before the first header is attached to it; content after the last header is added at the end. Post-processing includes two steps. First, we normalize headers and remove low-content sections to reduce boilerplate. Headers matching a curated list of common scientific section names (e.g., Introduction, Methods, Results) are kept as level-2 (##); others are demoted to level-3 (###). Sections with fewer than 10 words are discarded unless whitelisted. Second, we merge broken line wraps while preserving structural elements such as headers, empty lines, checklist markers, or simple equation labels. Paragraphs are separated by a single newline for downstream processing. Records with invalid annotations, offset errors, unreadable content, or insuļ¬icient structure (e.g., missing paragraphs or section headings) were excluded from downstream processing. The reproducible public pipeline mirrors this logic (see 6.2). 2.3 Text Chunking To accommodate the input length limitations of transformer-based language models, each pre- processed Markdown document was segmented into smaller, semantically coherent chunks using a recursive, token-aware chunking strategy. This choice was motivated in part by recent find- ings from the Chunk Twice, Embed Once study[9], which found that recursive, token-aware chunking strategy[10] consistently outperformed other chunking methods in chemistry-focused RAG systems. In addition, they found that retrieval-optimized models such as Intfloat E5 vari- ants significantly outperformed domain-specific models like SciBERT[11] and ChemBERTa[12] on chemistry-focused retrieval benchmarks. Based on these findings, we used intfloat/e5-large- v2 [4] for downstream embedding because of its strong performance on semantic search bench- marks. Chunking was implemented using HuggingFaceās AutoTokenizer, configured for e5-large-v2. Special tokens were excluded from token-length calculations. Our pipeline applies a hierarchical split: first on paragraph boundaries ( ), then on sentence-ending punctuation, and finally on whitespace if needed. Maximum chunk length was set to 200 tokens, with a 20-token overlap to preserve context. Chunks under 100 tokens were merged with adjacent segments. Documents with empty content or tokenization/encoding errors were skipped during chunk generation. 6 2.4 Embedding Generation We generated dense vector embeddings for each text chunk using the intfloat/e5-large-v2 model [4] via the HuggingFace sentence-transformers library. Chunks were tokenized with the modelās native tokenizer (max length: 512 tokens), prefixed with āpassage:ā following the modelās fine- tuning schema, and truncated as needed. Each resulting embedding is a 1024-dimensional, Euclidean-norm (L2)-normalized float32 vector suitable for downstream retrieval tasks. 2.5 License Screening To support conservative compliance review and downstream auditing, we implemented a multi- source license screening process for the chemistry corpus. For each full-text paper, the digital object identifier (DOI), when available, was extracted from the external_ids field. That DOI was used to query the Unpaywall, OpenAlex, and Crossref APIs in parallel to obtain license metadata for each paper. Unpaywall aggregates repository and publisher licenses, OpenAlex standardizes venue-level indicators, and Crossref provides publisher-reported terms. Retrieved licenses were normalized to Software Package Data Exchange (SPDX)-compatible categories (e.g., Creative Commons Attribution (C-BY), Creative Commons Zero (C0)). To improve both recall and reliability, we combined all three sources. A final license label for each paper was assigned only if at least two sources agreed on the license type, the third source was either missing or did not contradict the others, and there were no conflicting license reports. Only papers with metadata-based license signals that met our conservative study inclu- sion rule were retained in the internal screened corpus: c-by, Creative Commons Attribution- ShareAlike (C-BY-SA), c0, public domain, government works, as well as non-commercial variants such as Creative Commons Attribution-NonCommercial (C-BY-NC) and Creative Commons Attribution-NonCommercial-ShareAlike (C-BY-NC-SA). Papers with restrictive, conflicting, indeterminate licensing terms, or with insuļ¬icient agreement between sources were excluded. This rule provides a conservative and auditable basis for internal corpus construc- tion by retaining only records with consistent multi-source metadata support. It strengthens transparency and reproducibility for downstream auditing, while public redistribution decisions remain tied to the accuracy and version-specific relevance of upstream license records. Coverage and retained-corpus outcomes are reported in Section3.1. 2.6 Auxiliary Abstract Dataset for TL;DR and Subfield Enrichment To generate optional derivative annotations for a subset of the screened corpus, we assembled an auxiliary corpus of 19,992 chemistry abstracts from S2ORC with C-BY-compatible license metadata for model development. Each abstract was paired with two structured annotations: (i) a TL;DR-style summary (1ā2 sentences) that captures the core material or method, a key chemical finding, and at least one numeric result with standardized units; and (i) a multi-label subfield classification selected from a controlled vocabulary of 17 major chemistry subfields plus a fallback Others category (18 classes total). These annotations were automatically generated using a system-prompted GPT-4o model designed for structured information extraction. The supervision prompt required valid JSON 7 output only, constrained TL;DR summaries to 1ā2 sentences and no more than 50 words, re- quired at least one explicit numeric result with standard units when available, and restricted field assignments to a controlled chemistry vocabulary. For abstract-dependent enrichment, abstracts shorter than 100 characters were flagged as too short for text-quality purposes, and abstracts shorter than 1,000 characters were excluded from abstract embedding generation, sub- field prediction, and TL;DR generation. Full prompt text and schema details are provided in the Supplementary Materials (Section7.2). TL;DR summaries enable dense semantic embed- dings that improve retrieval precision and context eļ¬iciency during generation, while subfield labels support sharded indexing and prompt conditioning for domain-aware outputs. GPT-4o was used only for supervision; no proprietary APIs are required at inference time, supporting workflow-level reproducibility for the released downstream pipeline. To assess annotation quality, a random sample of 45 abstracts was manually reviewed by domain experts Sarah Mostafapour and Jamile Jafari. Their evaluation confirmed generally reasonable alignment between the abstracts, generated summaries, and subfield labels (see Sec- tion7.4). Release details for publicly available code and derivative resources are provided in Section6.1. 2.7 Abstractive Summarization Pipeline To transform complex and structurally varied chemistry abstracts into concise, structured TL;DR-style summaries, we fine-tuned a domain-adapted abstractive model using the sshleifer/distilbart- cnn-12-6 [13] checkpoint, a distilled version of Bidirectional and Auto-Regressive Transformers (BART) optimized for eļ¬iciency. Training was conducted on a C-BY enrichment dataset (Sec- tion2.6), using an 80/10/10 train/validation/test split. Abstracts were tokenized to a maximum input length of 1,024 tokens, and output summaries were capped at 128 tokens. Fine-tuning was performed with the Hugging Face Seq2SeqTrainer on a single NVIDIA Ray Tracing Texel eXtreme (RTX) 3090 graphics processing unit (GPU) using mixed precision, AdamW, a learning rate of 2e-5 and an effective batch size of 16 (per-device batch size of 4 with gradient accumu- lation over 4 steps) for 5 epochs. Evaluation, checkpoint saving, and logging were performed every 1,000, 1,000, and 500 steps, respectively. A fixed random seed (42) was applied across Python, NumPy, and PyTorch, with deterministic CUDA Deep Neural Network (CuDNN) set- tings enabled where applicable. Recall-Oriented Understudy for Gisting Evaluation metrics (ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum) and Bidirectional Encoder Represen- tations from Transformers Score (BERTScore) were computed using the Hugging Face evaluate framework with standard preprocessing. Full training and evaluation settings are documented in the Supplementary Materials (Section 7.5). Benchmark results are reported in Section 3.12, with complete metric definitions and evalu- ation details provided in Supplementary Materials Table 6. Access to trained model checkpoints and inference scripts is described in Section 6.2. Finally, using the fine-tuned model, we gener- ated TL;DR-style summaries for all papers with valid abstracts in our corpus. 8 2.8 Topic Classification Pipeline To support scalable organization and domain-specific retrieval of chemical knowledge, we de- veloped a reproducible multi-label topic classification pipeline. Each abstract is embedded with the intfloat/e5-large-v2 SentenceTransformer, yielding 1,024-dimensional unit-normalized vec- tors. These embeddings are fed to a two-layer multi-layer perceptron (MLP; 256 units per layer, rectified linear unit (ReLU), dropout = 0.3) with a sigmoid output over 18 subfields (including Others); abstracts may be assigned to multiple subfields to reflect interdisciplinary overlap. Training used batch normalization, weighted binary cross-entropy to address class imbalance, Adam with a learning rate of 1e-3, a batch size of 32, and up to 50 epochs with early stopping and reduce-on-plateau scheduling. Model selection used 5-fold cross-validation on the pooled training and validation data. A full description of the architecture, class-imbalance weighting, and training procedure is provided in Section7.7. Pointers to datasets, trained weights, and an interactive demo are provided in Sections6.1 and6.2. Benchmark results are reported in Section3.13. Finally, we applied the trained classifier to all papers with valid abstracts, assigning one or more subfield labels to each entry in the corpus. 2.9 Record Structure Each record is stored as a JSON object with required top-level fields for schema versioning, corpus identity, bibliographic metadata, abstract text, reconstructed full text, paragraph text units, paragraph embeddings, and license metadata from Unpaywall, Crossref, and OpenAlex. Optional enrichment fields include an abstract embedding, a TL;DR summary, and predicted subfield scores. Paragraph and embedding arrays are aligned by construction: each paragraph has a corresponding 1024-dimensional float32 embedding generated with intfloat/e5-large-v2, and both arrays have identical length. Full schema definitions, nested field dictionaries, and example records are provided in the Supplementary Materials (Section7.8). 2.10 Validation Workflow We performed a comprehensive validation of the screened study corpusās structure, metadata, semantics, alignment, and licensing across 582,683 records to ensure reliability for downstream research and reproducible reuse. Validation was conducted through three complementary pipelines: (i) schema and structural checks, (i) content and metadata quality assessment, and (i) alignment, reproducibility, and licensing verification. Core checks required the presence and correct typing of mandatory top-level fields, valid identifier and key patterns, fixed-length 1024-dimensional embedding vectors, paragraphāembedding alignment, and basic format con- straints for metadata, dates, and license fields. Each pipeline generated structured, timestamped reports with pass/warn/fail outcomes and detailed diagnostics, enabling full auditability. Com- plete validation rules and implementation details are provided in the Supplementary Materials (Sections 7.3to7.17). 9 3 Results 3.1 Corpus Yield and Coverage Starting from the S2AG 2024-12-31 papers release, our chemistry filter identified 18,621,044 doc- uments, corresponding to approximately 9.3% of the ingested corpus. Alignment by corpus_id yielded 8,660,979 abstracts (46.5% of the chemistry subset) and 1,074,638 full-text records (5.8% of the chemistry subset). As shown in Fig.2(a), 48.9% of chemistry-labeled records contained either an abstract or full text, 4.9% contained both, and 1.0% contained full text without an abstract. Figure 2: (a) Overlap between available full-text and abstract records for chemistry-labeled S2ORC papers. Most records contain only abstracts (43.0%), while 4.9% contain both and 1.0% contain only full text. (b) Coverage of license metadata sources (Unpaywall, OpenAlex, Crossref) for the chemistry full-text subset (N = 1,074,638). Overall, 89.9% of papers have license information from at least one source, with 49.5% covered by all three. Of the aligned full-text records, 966,946 were successfully converted to the normalized Mark- down representation used for downstream processing, and 963,697 produced valid chunked out- puts after token-aware segmentation. As shown in Fig.2(b), 966,384 papers (89.9%) had license metadata from at least one of Unpaywall, OpenAlex, or Crossref, and 531,784 (49.5%) were cov- ered by all three. Applying the multi-source agreement rule retained 582,683 full-text papers (60.4% of the embedding-generated set) in the final screened corpus used throughout this study. For a combined view of corpus coverage and license-source overlap, see Fig. 2. 3.2 Technical Validation Outcomes For a concise overview of the validation outcomes, see Table 3. 10 Table 3: Technical validation summary showing pass, warning, and non-complete rates (relative to the stated criteria) for each validation area. Validation AreaKey ChecksSummary of Results Schema & Structure JSON schema compliance, field typ- ing, embedding format 79% pass (460,397); 21% did not satisfy full en- richment/schema completeness criteria (122,286) ā mainly due to missing abstracts or abstract- dependent fields omitted for abstracts below the embedding eligibility threshold MetadataCompleteness of title, authors, venue, year, DOI, license 73% pass (423,519); 27% warn (159,164) ā most warnings from missing venue, or publication date. Subfield LabelsVocabulary match, valid confidence scores 79% pass (460,397); 21% did not satisfy full en- richment/schema completeness criteria (122,286) ā primarily attributable to missing or short up- stream abstracts under the abstract-embedding eligibility threshold. Text QualityLength limits, Unicode validity, abstractāfull-text alignment 68% pass (398,064); 32% warn (184,619) āā¼72% strong alignment; most records free of quality flags. ChunkingToken size limits, Unicode integrity, paragraphāembedding mapping 61% pass (354,064); 39% warn (228,619) ā warn- ings only for naturally short final chunks. EmbeddingsPresence, size, reproducibility100% pass (582,683) ā perfect cosine similarity agreement. IDs & IntegrityID consistency100% pass (582,683) ā no mismatches. LicensingCross-source open-access verifica- tion 100% pass (582,683) ā 88% C BY; no conflicts. Summaries ROUGE, BERTScore (evaluative) Moderate lexical overlap; strong semantic agree- ment. Figure 3: Schema validation results for 582,683 records in the chemistry full-text subset. Left: Pass/not-fully- complete distribution under schema/enrichment criteria, with non-complete cases broken down by the number of errors per record. Center: Most frequent non-completeness drivers are (i) short abstract embedding size (82,711 occurrences), (i) missing abstract-embedding key (39,575), and (i) missing predicted subfield annotation (39,575). Right: Normalized co-occurrence matrix showing two tight clusters: empty abstracts always co-occur with short abstract embedding size, while short abstracts (<1000characters) always co-occur with both missing abstract-embedding keys and missing predicted subfield annotations. These non-complete cases are primarily attributable to missing or short upstream abstracts under the abstract-length policy and are not indicative of embedding instability or licensing conflicts. 11 Figure 4: Metadata validation results for 582,683 records in the chemistry full-text subset. Left: Distribution of validation outcomes by category, showing that 72.7% of records pass all checks. The most common warnings are missing venue information (15.5%), missing open-access license metadata (10.4%), and missing publication date (1.0%). Less frequent issues include malformed author fields (0.3%), missing authors (0.2%), and short titles (<5 characters, rare). Right: Histogram of the number of warnings per record, with most affected records containing only a single warning. 3.3 Schema and Structural Validation Of the582,683records,460,397(79%) passed all checks (Fig.3-left), while122,286(21%) did not satisfy full enrichment/schema completeness criteria due to either one (67.6%) or two (32.4%) co- occurring structural errors (Fig.3-middle). These cases arise primarily from expected upstream abstract absence or abstract-length policy constraints rather than from embedding unreliability, licensing conflicts, or instability of the reconstruction workflow. Two main error clusters emerged (Fig.3-right): ā¢AāB: Empty abstracts paired with short or invalid abstract embeddings. ā¢CāDāE: Short abstracts co-occurring with missing abstract embeddings and missing predicted subfields. Review of source records confirmed that these issues originated upstream, where some doc- uments lacked abstracts entirely or contained malformed text segments due to PDF-to-text conversion errors in the S2ORC (see Fig. 2(a)). Accordingly, records with missing abstracts or abstracts below the enrichment threshold lacked abstract embeddings, predicted subfields, and TL;DR summaries by design. Future versions of the reconstruction workflow may address these structural gaps by sourcing valid abstracts from additional upstream repositories, thereby recovering high-quality records and reducing schema validation non-complete cases under these criteria. 12 3.4 Metadata Validation Bibliographic completeness and consistency were assessed for titles, authors, venues, publication dates, identifiers (e.g., DOI), fields of study, and open-access status. As shown in Fig.4-left, 423,519records (72.7%) passed without warnings. The most common warnings were miss- ing venue information (15.5%), missing license details (10.4%), and missing publication dates (1.0%). Warnings for malformed or missing author fields were rare (ā¤0.3%). The distribution of the number of warnings per record is shown in Fig.4-right. Most affected records contained only a single warning (21.6%of the corpus), while cases with three or more warnings were extremely rare (<0.1%). Overall, the screened study corpus demonstrates high metadata completeness, with deficiencies largely confined to a small subset of records affected by upstream source gaps. Notably, the license warnings primarily reflect gaps in the original S2ORC metadata rather than true absence of licensing evidence. We substantially improved record-level license prove- nance by enriching the corpus with metadata from OpenAlex, Crossref, and Unpaywall (see Section2.5), providing a stronger and more complete basis for auditing and downstream com- pliance review. Figure 5: Distribution of the number of predicted disci- plinary subfield labels per doc- ument for records that passed schema validation. Most doc- uments are assigned two sub- fields, followed by those with zero or one label, and a smaller portion with three labels. This reflects the multi-label nature of the classifier, which allows doc- uments to be categorized into multiple overlapping chemistry subfields. 3.5 Subfield Validation Predicted disciplinary subfields were evaluated against a controlled chemistry vocabulary, with confidence scores constrained to the range[0,1]. All122,286subfield cases that did not satisfy full enrichment/schema completeness criteria (21%) resulted from missing predictions caused by the abstract-dependent embedding policy 1 . Records passing without warnings totaled460,397 (79%). As shown in Fig.5, label distributions were stable across the corpus, with two subfields occurring most frequently. 1 These are design-expected omissions for short abstracts, not model failures on otherwise eligible records. 13 Figure 6: Character length distributions for abstracts (top) and full texts (bottom) in the chemistry full-text subset, shown on a log scale. For abstracts, the median length is 1,296 characters, with 83,877 flagged as too short (<100characters) and 82,711 having zero characters. For full texts, the median length is 29,077 characters, with no zero-length cases. Dashed vertical lines indicate quartiles and mean values. 14 Figure 7: Distribution of ROUGE-1 recall alignment scores between abstracts and full texts in the chemistry full- text subset, grouped into bins of width 0.05. Most documents have high alignment, with the largest peak in the 0.95ā1.0 range. Bars are color-coded by text validation flags, such as short abstracts, corrupted characters, low whitespace ratio, low American Standard Code for Information Interchange (ASCII) ratio, low sentence count, or missing/invalid abstract embeddings. The shaded area on the left indicates records with ROUGE-1 recall scores below 0.5. 3.6 Textual Content Quality We assessed abstract and full-text quality using multiple criteria, including character length, sen- tence count, Unicode integrity, ASCII/whitespace ratios, language identification, and abstractā full-text alignment (ROUGE-1 recall). As shown in Fig.6, abstracts were generally well-formed (median= 1,296characters), although83,877(38%) were under100characters, including82,711 empty abstractsāconsistent with upstream source gaps. Full texts were consistently substantive (median= 29,077characters;IQR = 19,100ā41,005), with outliers exceeding one million char- acters occurring rarely (<0.1%) and typically associated with unusually long supplementary sections. Alignment analysis (Fig. 7) showed that approximately355,000records (72%) achieved strong lexical alignment between abstracts and the introductory portions of their corresponding full texts (ROUGE-1ā„0.95). Fewer than10%of records had low alignment (<0.5), typically in combination with other quality flags. Most records (398,064) had no content-quality flags. These results confirm that the majority of records have strong textual integrity and that flagged cases can be programmatically identified and excluded from analyses sensitive to textual 15 quality. Figure 8: Paragraph-level chunk validation for the chemistry full-text subset. First row: Validation status showing that 61% of docu- ments pass without issues, while 39% contain short chunks. Second row: All affected documents contain exactly one short chunk. Third row: All short chunks occur in the final chunk of the document. Forth row: Distribution of mean chunk token length per document, with a median of 167 to- kens and most values falling between 162 (Q1) and 171 (Q3). 16 3.7 Chunk Level Validation For paragraph-level modeling, tokenized chunks and their embeddings were validated for token- length bounds, Unicode integrity, one-to-one paragraphāembedding mapping, dimensionality, finiteness, and near-unit vector norms. As shown in Fig.8-top-left,354,064documents (60.8%) passed without warnings; the remaining228,619(39.2%) received a single ātoo short chunkā warning. Fig.8-top-right shows that each affected document contained exactly one short chunk. Fig.8-bottom-left confirms these short chunks occur exclusively in the final paragraph position, typically containing80ā100tokens. Fig.8-bottom-right illustrates overall chunking consistency, with mean paragraph lengths tightly centered aroundā¼167tokens and a narrow interquartile range. Figure 9: Embedding reproducibility analysis for the chemistry full-text subset. Left: Distribution of mean cosine similarity between regenerated embeddings and their stored counterparts, with values tightly clustered around 1.0, indicating high reproducibility. Right: Distribution of maximum cosine drift (largest per-token deviation), showing that even the largest observed differences are on the order of10 ā7 , confirming numerical stability across regeneration. 3.8 Embedding Reproducibility Each paragraph embedding was validated for both structural integrity and semantic consistency. Checks confirmed correct dimensionality, finite values, and near-unit normalization. A strict one-to-one correspondence between paragraph IDs and embedding IDs was enforced, with any mismatches or anomalies flagged. As shown in Fig.9, all582,683records passed without errors (missing= 0.0%, extra= 0.0%, invalid= 0.0%). Cosine similarity between regenerated and stored embeddings was effectively perfect (mean = 1.000000; min= 0.9999999), with maximum numerical drift (±4.77Ć10 ā7 ) well within floating-point precision limits. Paragraphāembedding alignment was exact (Pearsonāsr= 1.0), confirming that embeddings are fully reproducible. 3.9 Identifier Integrity Internal identifiers (e.g., filename_id, corpus_id) and external identifiers (Crossref/Unpaywal- l/OpenAlex DOIs) were cross-validated and all records passed the test. 17 Figure 10: License validation results for 582,683 chemistry full-text records. Left: Sources from which license information was resolved, showing that most records (62.6%) were resolved by agreement across all three meta- data sources (Crossref, Unpaywall, OpenAlex), followed by Crossref+Unpaywall (24.6%), Unpaywall+OpenAlex (12.2%), and Crossref+OpenAlex (0.6%). Right: Distribution of resolved license types, with C-BY dominating (87.5%), followed by C-BY-NC (10.7%), C-BY-NC-SA (0.9%), C-BY-SA (0.5%), and public domain (0.4%). 3.10 Licensing Verification Licensing metadata from Crossref, Unpaywall, and OpenAlex were harmonized to a controlled vocabulary (e.g., C-BY, C-BY-NC). A record was considered a pass under the licensing validation step only when it satisfied metadata agreement under the study rule, meaning that at least two sources agreed on the open-access license without conflict. As shown in Fig.10-left, license resolution under this metadata-agreement rule was supported by agreement from all three sources for62.6%of records, from Crossref and Unpaywall for24.6%, from Unpaywall and OpenAlex for12.2%, and from Crossref and OpenAlex for0.6%. The resolved license distribution (Fig. 10-right) is dominated by C-BY (87.5%;509,638 records), followed by C-BY-NC (10.7%;62,473records). Other licensesāincluding C-BY- NC-SA, C-BY-SA, and Public Domaināaccount for less than2%combined. No license con- flicts were identified within the metadata sources under the study rule. 18 Figure 11: Relationship between abstract length and machine-generated summary evaluation scores for records with available summaries. Lines show binned mean scores with 95% confidence intervals for BERTScore_F1, ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum. The histogram (grey bars) indicates the number of abstracts per length bin, while the red gradient shows the proportion of abstracts exceeding the modelās input limit (ā¼1024 tokensā3500ā4500 characters, dashed vertical line). Scores generally peak at shorter lengths and decline as abstracts become longer, particularly beyond the model input limit. 3.11 Automated Summary Evaluation For records with available machine-generated summaries, we compared summaries to their corresponding abstracts using ROUGE-1/2/Lsum and BERTScore metrics (Fig.11). Median ROUGE values (ROUGE-1:32.0; ROUGE-2:18.1; ROUGE-Lsum:28.2) indicate moderate lexi- cal overlap, while BERTScore results demonstrate higher semantic capture (median recall:67.8; precision:36.3;F 1 :49.8). These results suggest that, although the summaries are not highly similar at the surface lexical level, they retain substantial semantic content from the abstracts. 3.12 TL;DR Summarization Benchmark Results Table6shows that lit2vec_tldr_finetuned outperforms both extractive and zero-shot abstrac- tive baselines across ROUGE and BERTScore metrics. Compared with bart_large_cnn, the fine-tuned model improves ROUGE-1 by +19.38, ROUGE-2 by +15.98, ROUGE-L by +18.38, and BERTScore F1 by +4.02, while decoding 36.6% faster than that larger baseline. Although extractive methods such as lead3 and lexrank remain faster in absolute runtime, they consis- tently underperform on summary quality. Qualitative inspection showed that the fine-tuned model typically preserves key experimental methods, midpoint numeric results, and brief application statements from the source abstracts. In high-performing examples, outputs closely match the GPT-4o-generated reference summaries in structure and specificity, whereas moderate examples retain the core findings with minor redundancy or phrasing issues (Section 7.6). These results support the use of the fine-tuned 19 DistilBART model as the main TL;DR generator in the Lit2Vec enrichment workflow. 3.13 Subfield Classification Benchmark Results On the held-out test set at decision threshold tau = 0.5, the classifier achieved micro-F1 = 0.81, weighted-F1 = 0.80, and macro-F1 = 0.75. Performance was strongest for frequent subfields such as Biochemistry (F1 = 0.92), Medicinal Chemistry (F1 = 0.82), and Materials Science (F1 = 0.80), whereas lower-resource domains such as Chemical Engineering (F1 = 0.53) remained more diļ¬icult. The full per-label benchmark is reported in Supplementary Table7. These results indicate that the embedding-based MLP captures major chemistry subfields reliably when adequate labeled support is available, while performance degrades for rarer classes. In practice, this makes the classifier suitable for broad corpus organization and retrieval support, with the main limitations concentrated in underrepresented labels. 3.14 Technical Validation Summary Across validation areas, the screened study corpus showed strong structural integrity, metadata completeness, textual quality, embedding reproducibility, and consistent metadata agreement under the study rule across 582,683 records. Across all checks, the corpus demonstrated high reliability:79%of records passed strict JSON Schema validation; metadata completeness ex- ceeded72%, with gaps largely limited to venue, license, and publication date fields; over68%of records showed strong abstractāfull-text alignment; chunking and embedding checks achieved near-perfect reproducibility; and licensing showed metadata agreement under the study rule, with100%of retained records passing the licensing-validation rule and88%of records screened as C-BY. Edge casesāprimarily short or missing abstracts and naturally short final chunksāare well characterized and can be programmatically filtered, supporting a robust and reproducible re- construction workflow. The core validated contribution of this work is therefore the reconstruc- tion workflow, together with the screened study corpus it produces under the stated access and licensing constraints. Enrichment layers should be understood as optional and inherently subset-dependent rather than universal outputs of the pipeline. In particular, abstract-related omissions are best interpreted as expected consequences of the validation and enrichment policy, especially where source metadata are sparse or incomplete, rather than as arbitrary processing failures. Remaining issues are thus tied mainly to upstream source quality, not to irreproducible downstream steps. 3.15 Example Applications We demonstrate Lit2Vec on three reference tasks: Trend Analysis, Semantic Paper Recom- mendation, and Retrieval-Augmented Generation. These examples illustrate representative downstream uses of the corpus and associated derivative resources. 20 3.15.1 RAG Pipeline To support scientific question answering grounded in published literature, we implemented a simple RAG system that integrates dense retrieval over paragraph-level embeddings with answer synthesis via an LLM. This system enables the generation of responses that are explicitly supported by source documents, with paragraph-level citations included inline. Valid embeddings are converted into NumPy arrays and indexed using FAISS with an In- dexFlatL2 structure. This enables eļ¬icient approximate nearest-neighbor search over paragraph- level vectors. Paragraphs and their corresponding metadata are stored in memory for lookup and retrieval. To perform retrieval, the system encodes the userās query into an embedding using the intfloat/e5-large-v2 model from the Sentence-Transformers library [ 4]. This query vector is nor- malized and searched against the FAISS index to retrieve the top-kmost semantically relevant paragraphs. Each retrieved result includes the paragraph text, a paragraph-level identifier, and associated metadata (title, authors, year, DOI). The retrieved paragraphs are assembled into a structured context block, where each para- graph is labeled with its unique identifier (e.g., 123456P2). This block, along with the userās question, is provided as a prompt to an OpenAI GPT-4 model (gpt-4o) using the chat comple- tion API. The system instructs the model to answer the question in a coherent paragraph format while citing supporting statements using paragraph IDs in parentheses. This approach ensures that each part of the generated answer can be traced back to its evidence source. In addition to the generated answer, the system outputs a reference list citing the original documents that contributed supporting paragraphs. Each reference includes the full title, au- thors, publication year, and DOI (if available), associated with the paragraph-level identifiers used in the inline citations. The full implementation is provided in a Jupyter notebook atexample_tasks/RAGon GitHub. 3.15.2 Semantic Paper Recommendation Pipeline We developed a lightweight semantic recommendation system designed to support literature discovery by identifying papers that are semantically related to a reference document or a free- text scientific query. The system is implemented in Python and leverages vector representations of scientific abstracts to perform similarity-based retrieval and ranking. During data loading, the system filters out any records with missing or malformed embed- dings. Specifically, it retains only those entries where the abstract_embedding field contains a non-empty, one-dimensional array of numeric values. This ensures that only valid documents are included in the subsequent indexing and retrieval steps. For semantic retrieval, we use the FAISS library [ 14] to build an eļ¬icient nearest-neighbor index. The abstract embeddings are L2-normalized and indexed using an IndexFlatIP structure, which enables fast cosine similarity search via inner products. The system supports two primary retrieval modes. In the first mode, a known paper (spec- ified by its corpus_id) is used as a query, and the most similar papers are returned based on 21 abstract-level semantic similarity. In the second mode, a free-text query is encoded into an em- bedding using the intfloat/e5-large-v2 model from the Sentence-Transformers library [4]. This query embedding is then used to retrieve relevant documents from the FAISS index. After retrieving a set of candidate papers, the system optionally applies a set of filters. These include filtering by publication year, enforcing a minimum confidence threshold for a predicted subfield label (e.g., ensuring that the score for āBiochemistryā exceeds 0.6), and selecting only open-access papers. To improve diversity among the top-ranked results, we implement maximal marginal rele- vance (MMR) 2 [15] reranking. MMR balances relevance to the query and dissimilarity among the results, thereby promoting coverage of multiple subtopics. This is particularly useful when the user query is broad or when diversity in retrieved papers is desired. For each retrieved document, the system returns its title, publication year, authors, venue, DOI (if available), and cosine similarity score relative to the query. If present, a Too Long; Didnāt Read (TL;DR) summary is also included. The full implementation is provided in a Jupyter notebook atrecomendation_systemon GitHub. 3.15.3 Trend Analysis One of the most immediate applications enabled by Lit2Vec is fine-grained temporal trend analysis across chemical subfields. The datasetās structured full text, standardized Markdown formatting, and paragraph-level segmentation allow targeted extraction of experimental and methodological mentions over time. In addition, its paragraph-level embeddings support both lexical and semantic filtering with minimal infrastructure. To demonstrate this, we extract and analyze trends in the use of surface-enhanced Raman spectroscopy (SERS) substrates and excitation wavelengths. Using regular expressions applied to paragraphs from relevant sections (e.g., Methods, Instrumentation, Results), we identify mentions of gold and silver nanoparticles co-occurring with SERS. We also track mentions of Raman excitation at 785 nm and 532 nm, restricting attention to sentences containing Raman- relevant context and instrumentation cues (e.g., ālaserā, āexcitationā). Documents are grouped by publication year using metadata fields (year, published, publi- cationDate), and binary flags are computed to indicate the presence of specific technologies per paper. Temporal smoothing is applied using a 3-year centered moving average. As shown in Figure 12, these methods capture meaningful domain signals. For instance, 532 nm excitation has become dominant in Raman experiments over the last decade, although 785 nm is resurging slightly in recent years. In SERS applications, gold and silver substrates are nearly balanced in recent years, with a slight shift toward gold after 2020. Because Lit2Vec includes paragraph-level embeddings, the same trend analysis workflow can be extended to embedding-based retrieval using tools like FAISS, or combined with domain- specific classifiers or keyword expansion models. These capabilities make Lit2Vec suitable for bibliometric studies, citation-aware recommendation systems, and retrospective analyses of re- search trends across multiple chemistry subdomains. 2 Maximal Marginal Relevance 22 A complete implementation of this example is available atexample_tasks/trend_analysis on GitHub. Figure 12: Temporal trends in SERS and Raman spectroscopy usage across 582,683 full-text chemistry papers in Lit2Vec. (Top) Document counts and smoothed shares for gold and silver nanoparticle use in SERS. (Bottom) Document counts and normalized shares for 785 nm and 532 nm excitation in Raman experi- ments. Trend analysis uses paragraph-level filtering and metadata-driven grouping, without fine-tuned models. 4 Discussion Lit2Vec should be interpreted not simply as a static corpus release, but as a reproducible work- flow for constructing a retrieval-ready chemistry text resource from public upstream sources using conservative, metadata-based license screening. This distinction matters because many existing chemistry text resources are limited to abstracts, insuļ¬iciently documented, or not prepared for semantic retrieval at the paragraph level. The present results show that, under the stated access and licensing constraints, the workflow can reconstruct a large screened cor- pus with stable structure, explicit provenance, and reproducible internal representations for downstream analysis. In that sense, the central validated contribution is the combination of the reconstruction pipeline and the screened study corpus it produces, rather than any single downstream artifact in isolation. The technical validation results support the reliability of this workflow for downstream use. Across 582,683 records, the corpus showed strong schema integrity, substantial metadata com- 23 pleteness, generally high textual quality, exact identifier consistency, reproducible embeddings, and consistent metadata agreement under the study rule for license resolution. Taken together, these checks suggest that the dataset is reliable enough for retrieval, recommendation, trend analysis, and related literature-mining tasks, because the main operational requirements for such uses are addressed jointly rather than one at a time. Equally important, the edge cases are well characterized: short or missing abstracts, naturally short final chunks, and sparse metadata appear in identifiable patterns and can be programmatically filtered when a downstream task requires stricter inclusion criteria. These cases therefore indicate bounded limitations of the available source material and enrichment policy, not instability of the reconstruction workflow itself. The enrichment results add a second layer of interpretation. Automated summary evalua- tion shows only moderate lexical overlap between generated TL;DRs and source abstracts, but substantially stronger semantic agreement, indicating that the summaries are more useful as compact semantic surrogates than as extractive replicas. This is consistent with the intended role of the summaries in retrieval and browsing workflows, where concise semantic compression is often more valuable than token-level similarity. The benchmark results further show that the fine-tuned lit2vec_tldr_finetuned model is the most practical main summarization compo- nent in the workflow, outperforming the tested baselines across ROUGE and BERTScore while remaining operationally eļ¬icient enough for large-scale enrichment. The same interpretation applies to the subfield classifier. Its held-out performance indicates that the embedding-based MLP is suitable for broad organization of the corpus and for retrieval support across major chemistry domains, especially where labeled support is adequate. At the same time, weaker performance on rarer classes means that these labels should be treated as useful navigational metadata rather than as uniformly definitive annotations for all subfields. More broadly, both the TL;DR summaries and predicted subfields should be understood as optional enrichment layers that improve usability for eligible subsets of records, not as universal outputs that define the validity of the corpus itself. The example applications help clarify why these validation and enrichment results matter in practice. The retrieval-augmented generation, semantic recommendation, and trend-analysis demonstrations show that paragraph-level chunking, dense embeddings, and license-screened metadata are suļ¬icient to support operational downstream workflows rather than only offline validation. In particular, the resource can support evidence-grounded retrieval, similarity-based literature discovery, and longitudinal analysis of technical trends with relatively modest addi- tional infrastructure. This practical usability is an important outcome of the study because it shows that the corpus is not only structurally well-formed, but also aligned with concrete chemistry AI and bibliometric use cases. Several limitations remain and should shape how Lit2Vec is reused. First, a substantial subset of records has missing abstracts or abstracts below the abstract-enrichment threshold, which limits abstract-level embeddings, summaries, and subfield annotations. However, many of these omissions are expected consequences of upstream source gaps and the enrichment pol- icy, especially where S2ORC records lack abstracts entirely or provide text too short for the abstract-dependent enrichment steps, and should not be interpreted as arbitrary downstream 24 processing failures. Second, exact regeneration remains upstream-dependent: the study corpus is not fully redistributed, and faithful reconstruction depends on continued access to the pinned source release, archived metadata snapshots, and compatible external services. Relatedly, redis- tribution constraints remain intrinsic to the problem setting, so the contribution is a transparent and reproducible workflow under legal screening rather than an unrestricted public mirror of all source-derived content. Additional uncertainty comes from noisy or imperfect source labels for interdisciplinary documents, possible bias in GPT-4o-supervised pseudo-labels and TL;DR targets, temporal drift in live metadata after collection, and the text-only scope of the re- source, which omits figures, chemical structures, and tabular data that are often scientifically important. Future work should therefore focus on extending coverage and tightening reproducibility without overstating what can be redistributed. The most immediate directions are to recover missing abstracts from additional upstream repositories where legally possible, improve clas- sification performance for rare labels, and expand the framework beyond text-only content to modalities such as figures, tables, or structured chemical representations. Reproducibility could also be strengthened further through more extensively pinned upstream snapshots, archived metadata resolutions, and continued release of metadata, provenance, validation, and other low-risk reproducibility artifacts, alongside any separately vetted permissive subsets that clearly support public redistribution. These steps would not change the core contribution of Lit2Vec, but would increase the completeness and practical reach of the workflow for chemistry-focused AI research. 5 Conclusions Lit2Vec provides a reproducible workflow for constructing and validating a chemistry corpus from S2ORC using conservative, metadata-based license screening, together with schema doc- umentation, reconstruction code, metadata/provenance artifacts, and validation resources that support workflow-level reproducibility. The resulting study corpus combines structured full text, token-aware paragraph chunks, dense embeddings, and optional enrichment layers in a form that is suitable for retrieval-oriented chemistry AI workflows under explicit access and licensing constraints. The release is intentionally scoped: the full reconstructed corpus is not publicly redis- tributed, and some enrichment outputs remain subset-dependent because they rely on abstract availability and upstream metadata quality. Within those limits, Lit2Vec establishes a transpar- ent foundation for workflow-level reproducibility and for building downstream systems grounded in license-screened scientific text. 5.1 Applications and limitations The full reconstructed corpus used in this study (582,683 records; 277 GB compressed,ā¼825 GB uncompressed) is not publicly redistributed at present because it is derived from upstream sources with heterogeneous licensing and redistribution conditions. The released materials sup- port workflow-level reproducibility rather than unrestricted corpus redistribution. Researchers 25 seeking to reuse the pipeline should use the released reconstruction workflow to rebuild the cor- pus from publicly available upstream datasets and metadata services, following the providersā published terms of use, licensing conditions, and normal download policies. Optional API cre- dentials may improve download speed, but are not required. Data Structure and Reconstruction: The internal corpus representation uses standalone JSON records with text, embeddings, and metadata. The released schema documents this structure so that users can reconstruct compatible records from upstream data and use com- patible local records or separately vetted permissive subsets with the same field conventions. Example loading code for reconstructed JSON files: from datasets import load_dataset dataset = load_dataset( "json", data_files="train": "reconstructed_records/*.jsonl", trust_remote_code=True,) # Peek at one example sample = dataset["train"][0] print(sample["metadata"]) # Raw JSON string with title, authors, etc. print(sample["predicted_subfield"]) # List of label, score # Extract title from metadata import json metadata = json.loads(sample["metadata"]) print("Title:", metadata.get("title")) # Display subfield predictions print("Predicted subfields:") for sub in sample["predicted_subfield"]: print(f" sub['label']: sub['score']:.3f") Loading the full reconstructed corpus in memory requires approximately 900 GB of RAM. Most users should use streaming or sharded loading after reconstructing it locally. All paragraph vectors are 1024-dimensional, float32, and L2-normalised. A FAISS-based retrieval baseline is provided in section 3.15.1. The average paragraph length (167 tokens) allows multiple hits to fit within an 8k-token LLM context window. The corpus supports tasks such as semantic retrieval, domain-specific LLM fine-tuning, chemistry knowledge graph construction, and subfield prediction. For modeling, predicted_subfield should be treated as multi-label. Quality Flags and Filtering: Each record includes validation flags such as short_abstract, missing_embeddings, and too_short_chunk. For high-recall applications like retrieval, flagged 26 records may be retained. For LLM training or evaluation, users may exclude records with hard fail statuses (see Section3.2). Compute Considerations: A flat inner-product index over the full set of 29M paragraph vectors requiresā¼120 GB RAM. Optimized variants (e.g., IVF-PQ or HNSW) can reduce this toā¤40 GB with minimal recall loss. Full-corpus embedding or fine-tuning benefits fromā„40 GB GPUs, though chunk-wise streaming is feasible on consumer-grade hardware. Licensing and Attribution: Licensing metadata is harmonised from OpenAlex, Unpaywall, and Crossref. Approximately 87.5% of records in the screened study corpus are C-BY and 10.7% are C-BY-NC. Each JSON includes explicit license fields (unpaywall_license, cross- ref_license, openalex_license). Users must verify source-specific terms before reconstructing, sharing, or redistributing any derivative outputs. Citation: Users should cite this manuscript and the corresponding code or derivative-resource release when using Lit2Vec materials. 6 Declarations 6.1 Availability of data and materials Third parties can deterministically reproduce the screened study corpus and the reported anal- yses by using the pinned S2AG 2024-12-31 release together with the archived reconstruction manifests, frozen license-resolution metadata, released code, and analysis scripts provided with this work. Minor differences should arise only when later upstream releases are substituted or live metadata services are queried instead of the archived snapshot. The full reconstructed Lit2Vec corpus analyzed in this study is not publicly redistributed because it is derived from upstream sources with heterogeneous licensing and redistribution conditions. Reproducibility is therefore supported through public release of the reconstruction and validation code, workflow documentation, machine-readable schema and example records, record-level reconstruction manifests, license-screening and license-resolution outputs, technical validation reports, figure- and table-level source data, and legally redistributable derivative resources generated in this work. All released materials are available athttps://huggingface.co/datasets/Bocklitz-Lab/Lit2Vec- dataset. The released code repository is available atBocklitz-Lab/Lit2Vec-code. Together, these resources provide the acquisition, preprocessing, chunking, embedding, filtering, and license- screening workflow required to reconstruct Lit2Vec-style records from the same pinned upstream release and associated metadata snapshot. Optional API credentials may improve throughput for some steps but are not required for reproduction. We do not redistribute the full reconstructed full-text corpus, restricted abstracts, paragraph- level corpus exports derived from non-redistributable source content, embedding vectors derived from non-redistributable corpus content, or other source-derived outputs whose redistribution is not permitted. The released schema documents the internal record structure used in this study, including both publicly releasable metadata/provenance fields and internal-or-local reconstruction fields such as text, chunk, embedding, and summary fields. Researchers can use this schema to 27 generate compatible local records after reconstructing the corpus from the same public upstream sources. The schema documents these structures for compatibility and reproducibility, but does not by itself imply general public redistribution of all field contents. Each reconstructed record retains machine-readable license metadata in the fields unpay- wall_license, crossref_license, and openalex_license. These fields are provided to support au- diting and downstream compliance checks, and users should verify record-level reuse conditions before redistribution or derivative use. A limited set of auxiliary resources used for model training and supervision is publicly available only where redistribution was separately scoped as appropriate for release: ā¢TL;DR Supervision Dataset (19,992 C-BY abstracts with GPT-4o-generated summaries) Bocklitz-Lab/lit2vec-tldr-bart ā¢Subfield Classification Dataset (multi-label annotations for 18 chemistry subfields) Bocklitz-Lab/lit2vec-subfield-classifier These released metadata, schema, and validation artifacts support inspection, auditing, and verification of the workflow outputs. Full regeneration depends on access to the same pinned upstream release and archived metadata snapshot distributed with this work. 6.2 Code availability All code used to generate the screened study corpus and reproduce the analyses is open-source and publicly available. A reproducible version of the workflow is released as a simplified, single- machine implementation. While the original pipeline was designed for high-throughput pro- cessing (e.g., parallel MongoDB filtering, distributed embedding inference), the released ver- sion faithfully replicates the core acquisition, preprocessing, chunking, embedding, and license- screening steps for local reconstruction from pinned upstream resources. Input/output formats, chunking logic, and embedding normalization remain identical. This code release should not be interpreted as a public redistribution of the general corpus or of broad text-derived stores. Main pipeline repository: ā¢Bocklitz-Lab/Lit2Vec-codeā contains the complete processing pipeline, including: āPaperProcessor.py: Parses full-text JSON into structured Markdown with normal- ized headers and section-aware grouping. āChunkProcessor.py: Performs recursive, token-aware chunking of text for transformer compatibility. āEmbeddingProcessor.py: Generates embeddings using intfloat/e5-large-v2. āLicenseValidator.py: Filters reusable content based on license metadata from Unpay- wall, OpenAlex, and Crossref. Optional enrichment repositories: 28 ā¢lit2vec-tldr-bartā scripts for TL;DR summarization, including model configs, training metrics, and evaluation logic. ā¢lit2vec-subfield-classifierā multi-label classification pipeline for scientific subfields, in- cluding training, class balancing, and inference. Example downstream applications: ā¢RAG Pipeline:example_tasks/RAGā dense paragraph-level retrieval using FAISS (In- dexFlatL2) and answer synthesis with gpt-4o, including inline paragraph ID citations and auto-generated references. ā¢Semantic Paper Recommendation:example_tasks/recommendation_systemā cosine similarity search over abstract embeddings using FAISS (IndexFlatIP); supports both paper-to-paper and free-text queries (encoded with intfloat/e5-large-v2), with optional filters and MMR reranking. ā¢Trend Analysis:example_tasks/trend_analysisā regex- and embedding-friendly pipelines for analyzing temporal trends in SERS substrates and Raman excitation wavelengths, us- ing centered 3-year smoothing. Trained model checkpoints and interactive demos: ā¢TL;DR Summarization: ālit2vec-tldr-bart(model) ālit2vec-tldr-bart-space(demo) ā¢Subfield Classifier: ālit2vec-subfield-classifier(model) ālit2vec-subfield-classifier-space(demo) 6.3 Competing interests The authors declare no competing interests. 6.4 Funding This work is supported by the BMFTR, funding program Photonics Research Germany (13N15466 (LPI-BT1-FSU), 13N15710 (LPI-BT3-FSU), 13N15715 (LPI-BT4-FSU)) and is integrated into the Leibniz Center for Photonics in Infection Research (LPI). The LPI initiated by Leibniz IPHT, Leibniz-HKI, Friedrich Schiller University Jena and Jena University Hospital is part of the BMFTR national roadmap for research infrastructures. 29 6.5 Authorsā contributions Mahmoud Amiri was responsible for the conceptualization and design of the pipeline, soft- ware implementation, data curation, validation, analysis, visualization, and preparation of the original manuscript draft. Prof. Dr. Thomas Bocklitz provided supervision, guidance on methodological design, and critical revision of the manuscript. Sara Mostafapourghasrodashti and Jamile Mohammad Jafari contributed to the validation of the study. Author roles follow the CRediT taxonomy; for more detail, see Table4. Table 4: Author contributions according to the CRediT taxonomy. RoleMahmoud A. Thomas B.Sara M.Jamile M.J. ConceptualizationX MethodologyXX SoftwareX ValidationXXX Formal AnalysisX InvestigationX Data CurationX VisualizationX Writing ā Original DraftX Writing ā Review & Editing X SupervisionX Project AdministrationX All authors approved the final version of the manuscript. 6.6 Acknowledgements We thank the Allen Institute for AI and the Semantic Scholar team for providing the well- documented S2ORC dataset. We also acknowledge the open-source contributions of Hugging Face, SentenceTransformers, and MongoDB, which made the development of a scalable pipeline possible. References [1]OurResearch. Openalex, 2025. URLhttps://openalex.org/. Accessed: 2025-06-04. [2]OurResearch. Unpaywall, 2025. URLhttps://unpaywall.org/. Accessed: 2025-06-04. [3]Crossref. Crossref, 2025. URLhttps://w.crossref.org/. Accessed: 2025-06-04. [4]Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. Text embeddings by weakly-supervised contrastive pre-training. arXiv preprint arXiv:2212.03533, 2022. [5]National Center for Biotechnology Information. Pubchem, 2025. URLhttps://pubchem. ncbi.nlm.nih.gov/ . Accessed: 2025-06-04. [6]European Bioinformatics Institute. Chembl, 2025. URLhttps://w.ebi.ac.uk/ chembl/. Accessed: 2025-06-04. 30 [7]Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Dan S Weld. S2orc: The semantic scholar open research corpus. arXiv preprint arXiv:1911.02782, 2019. [8]MongoDB, Inc. Mongodb, 2025. URLhttps://w.mongodb.com/. Accessed: 2025-06-04. [9]Mahmoud Amiri and Thomas Bocklitz. Chunk twice, embed once: A systematic study of segmentation and representation trade-offs in chemistry-aware retrieval-augmented gener- ation. arXiv preprint arXiv:2506.17277, 2025. [10]Harrison Chase and LangChain contributors. Langchain: Building applications with llms through composability.https://github.com/langchain-ai/langchain, 2022. Accessed: 2025-05-28. [11]Iz Beltagy, Kyle Lo, and Arman Cohan. Scibert: A pretrained language model for scientific text. EMNLP, 2019. [12]Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. Chemberta: large- scale self-supervised pretraining for molecular property prediction. arXiv preprint arXiv:2010.09885, 2020. [13]Hugging Face. sshleifer/distilbart-cnn-12-6, 2025. URLhttps://huggingface.co/ sshleifer/distilbart-cnn-12-6. Model checkpoint accessed: 2025-06-04. [14]Jeff Johnson, Matthijs Douze, and HervĆ© JĆ©gou. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7(3):535ā547, 2019. [15]Jaime Carbonell and Jade Goldstein. The use of mmr, diversity-based reranking for re- ordering documents and producing summaries. In Proceedings of the 21st annual inter- national ACM SIGIR conference on Research and development in information retrieval, pages 335ā336, 1998. [16]Adrian Mirza, Nawaf Alampara, MartiƱo RĆos-GarcĆa, Mohamed Abdelalim, Jack Butler, Bethany Connolly, Tunca Dogan, Marianna Nezhurina, Bünyamin Åen, Santosh Tiruna- gari, Mark Worrall, Adamo Young, Philippe Schwaller, Michael Pieler, and Kevin Maik Jablonka. Chempile: A 250 gb diverse and curated dataset for chemical foundation models, 2025. URLhttps://arxiv.org/abs/2505.12534. [17]Petr Knoth, Drahomira Herrmannova, Matteo Cancellieri, Lucas Anastasiou, Nancy Pon- tika, Samuel Pearce, Bikash Gyawali, and David Pride. Core: A global aggregation service for open access papers. Scientific Data, 10(1):366, 2023. [18]Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy, Jonathan Bragg, Alexan- dra Buraczynski, Isabel Cachola, Stefan Candra, Yoganand Chandrasekhar, Arman Cohan, et al. The semantic scholar open data platform. arXiv preprint arXiv:2301.10140, 2023. [19]Jiayuan He, Dat Quoc Nguyen, Saber A Akhondi, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Zubair Afzal, Zenan Zhai, Biaoyan Fang, Hiyori Yoshikawa, et al. Chemu 2020: Natural language processing methods are effective for information extraction from chemical patents. Frontiers in Research Metrics and Analytics, 6:654438, 2021. 31 [20]National Library of Medicine. Pubmed central open access subset, 2022.https://w. ncbi.nlm.nih.gov/pmc/tools/openftlist/. [21]Yu Gu, Ruiqi Tinn, Hao Cheng, et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH), 3(1):1ā23, 2021. [22]MartĆn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467, 2016. [23]Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, et al. Cord-19: The covid-19 open research dataset. arXiv preprint arXiv:2004.10706, 2020. [24]Martin Krallinger, Obdulia Rabal, AnĆ”lia LourenƧo, et al. The chemdner corpus of chem- icals and drugs and its annotation principles. Journal of Cheminformatics, 7(1):S2, 2015. [25]Jiao Li, Yueping Sun, Richard J Johnson, et al. Biocreative v cdr task corpus: a resource for chemical disease relation extraction. Database, 2016, 2016. [26]Ranit Islamaj, Sun Kim, Laritza Rodriguez, et al. Nlm-chem, a new resource for chemical entity recognition in pubmed full text literature. Scientific Data, 8(1):1ā12, 2021. [27]Jiayuan He, Dat Quoc Nguyen, Saber A Akhondi, Christian Druckenbrodt, Camilo Thorne, Ralph Hoessel, Zubair Afzal, Zenan Zhai, Biaoyan Fang, Hiyori Yoshikawa, et al. Overview of chemu 2020: named entity recognition and event extraction of chemical reactions from patents. In International Conference of the Cross-Language Evaluation Forum for European Languages, pages 237ā254. Springer, 2020. [28]Yujia Feng, Yida Shen, Tianze Xie, et al. Chemrxivquest: A benchmark for open-domain question answering in chemistry. arXiv preprint arXiv:2310.07699, 2023. [29]Kanishk Choudhary and David P Kelley. Chemnlp: An open-source toolkit for natural language processing in chemistry, 2023. https://github.com/OpenBioLink/ChemNLP. [30]Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. Camel: Communicative agents for āmindā exploration of large scale language model society, 2023. [31]National Library of Medicine. Pubmed, 2025. URLhttps://pubmed.ncbi.nlm.nih.gov/. Accessed: 2025-06-04. [32]Bocklitz Lab. Lit2vec dataset, 2026. URLhttps://huggingface.co/datasets/ Bocklitz-Lab/Lit2Vec-dataset . Accessed: 2026-03-30. [33]Bocklitz Lab. Lit2vec code, 2026. URLhttps://github.com/Bocklitz-Lab/ Lit2Vec-code . Accessed: 2026-03-30. 32 [34]Bocklitz Lab. lit2vec-tldr-bart dataset, 2026. URLhttps://huggingface.co/datasets/ Bocklitz-Lab/lit2vec-tldr-bart. Accessed: 2026-03-30. [35]Bocklitz Lab. lit2vec-subfield-classifier dataset, 2026. URLhttps://huggingface.co/ datasets/Bocklitz-Lab/lit2vec-subfield-classifier. Accessed: 2026-03-30. [36]Bocklitz Lab. lit2vec-tldr-bart, 2026. URLhttps://github.com/Bocklitz-Lab/ lit2vec-tldr-bart . Accessed: 2026-03-30. [37]Bocklitz Lab. lit2vec-subfield-classifier, 2026. URLhttps://github.com/Bocklitz-Lab/ lit2vec-subfield-classifier. Accessed: 2026-03-30. [38]Bocklitz Lab. Lit2vec example task: Rag, 2026. URLhttps://github.com/ Bocklitz-Lab/Lit2Vec-code/tree/main/example_tasks/RAG. Accessed: 2026-03-30. [39]Bocklitz Lab. Lit2vec example task: recommendation system, 2026. URL https://github.com/Bocklitz-Lab/Lit2Vec-code/tree/main/example_tasks/ recomendation_system. Accessed: 2026-03-30. [40]Bocklitz Lab. Lit2vec example task: trend analysis, 2026. URLhttps://github.com/ Bocklitz-Lab/Lit2Vec-code/tree/main/example_tasks/trend_analysis . Accessed: 2026-03-30. [41]Bocklitz Lab. Lit2vec annotation validation scripts, 2026. URLhttps://github.com/ Bocklitz-Lab/Lit2Vec-code/tree/main/annotation_validation. Accessed: 2026-03- 30. [42]Bocklitz Lab. lit2vec-tldr-bart model, 2026. URLhttps://huggingface.co/ Bocklitz-Lab/lit2vec-tldr-bart. Accessed: 2026-03-30. [43]Bocklitz Lab. lit2vec-tldr-bart-space demo, 2026. URLhttps://huggingface.co/ spaces/Bocklitz-Lab/lit2vec-tldr-bart-space. Accessed: 2026-03-30. [44]Bocklitz Lab. lit2vec-subfield-classifier model, 2026. URLhttps://huggingface.co/ Bocklitz-Lab/lit2vec-subfield-classifier. Accessed: 2026-03-30. [45]Bocklitz Lab. lit2vec-subfield-classifier-space demo, 2026. URLhttps://huggingface.co/ spaces/Bocklitz-Lab/lit2vec-subfield-classifier-space. Accessed: 2026-03-30. 7 Supplementary Materials 7.1 Related Work Early efforts in biomedical text mining focused on manually annotated, task-specific datasets. The CHEMDNER corpus [ 24] established a benchmark for chemical named entity recognition (NER), featuring approximately 84,000 chemical mentions across 10,000 PubMed [31] abstracts. The BioCreative V ChemicalāDisease Relation (CDR) corpus [ 25] extended this direction to relation extraction, annotating 1,500 abstracts with chemicalādisease pairs. NLM-Chem [26] 33 further advanced the field by providing high-quality chemical entity annotations in full-text ar- ticles. In parallel, domain-specific corpora emerged from patents and other non-journal sources. Notably, the ChEMU 2020 corpus [19] comprises approximately 1,500 annotated text segments drawn from around 170 chemical patent documents, with fine-grained entity annotationsāsuch as the roles of chemicals in reactionsāand detailed event extraction over procedural steps. More recently, initiatives have begun to unify multiple NLP tasks within chemistry. Chem- RxivQuest [28] automatically generates approximately 970 validated questionāanswer pairs from 155 ChemRxiv preprints across 17 subfields, enabling the development of chemistry-specific question answering (QA) systems. ChemNLP [29] provides an open-source library that aggre- gates and processes large-scale chemical literature, including around 1 million abstracts from arXiv and 100,000 compound records from PubChem, with pipelines for tasks such as named entity recognition, classification, clustering, summarization, and text generation. Recent efforts have also focused on creating large-scale corpora for training general-purpose chemical foundation models. ChemPile [16] is a 250 GB, multimodal dataset comprising over 75 billion tokens across educational materials, scientific articles, structured chemical representa- tions (e.g., Simplified Molecular Input Line Entry System (SMILES), Self-Referencing Embed- ded Strings (SELFIES), International Chemical Identifier (InChI), International Union of Pure and Applied Chemistry (IUPAC)), executable code, reasoning traces, and molecular images. Broad scientific text collections, such as the Semantic Scholar Open Research Corpus (S2ORC) [7, 18], the PubMed Central Open Access subset [20], and CORE [17], contain millions of articles across disciplines, including chemistry. These corpora have played a pivotal role in enabling large-scale text mining, citation analysis, document classification, and pretraining of general- purpose language models. Despite their individual contributions, existing corpora for chemical NLP face common lim- itations that hinder their integration into modern RAG and semantic search workflows. These include: (i) incomplete or inconsistent full-text coverage; (i) lack of paragraph-level, token- aware semantic segmentation suitable for transformer-based models; (i) absence of precom- puted dense embeddings for immediate use in retrieval systems; (iv) restrictive or heterogeneous licensing conditions that complicate downstream reuse and distribution; and (v) insuļ¬icient metadata normalization across chemistry subdomains. Furthermore, general-purpose scientific datasets lack chemical-specific structuring, while even large-scale efforts like ChemPile are op- timized for model pretraining rather than real-time retrieval or question answering. These gaps collectively limit reproducibility, scalability, and legal certainty for building open, domain-aware AI systems in chemistry. 7.2 System Prompt for Summary Generation and Field Classification The pipeline assembled a high-quality corpus of approximately 19,992 chemistry abstracts, each licensed under C-BY terms. Next, we used GPT-4o with a dedicated system prompt to generate high-quality summaries and field classifications. The system prompt instructed the model as follows: 34 CorpusDescriptionLimitations CHEMDNER [24]84K mentions in 10K PubMed abstracts; strong NER benchmark. Abstracts only; limited li- cense; no embeddings. CDR [25]1.5K abstracts with chemicalādisease rela- tions; expert-annotated. Small; abstracts only; no embeddings. NLM-Chem [26]Full-text NER in PMC; good chemical cov- erage. NER-only; no embeddings or structure. ChEMU 2020 [19]1.5K patent segments with NER/reaction events. Small; patent-focused; no embeddings. ChemRxivQuest [28]970 QA pairs from 155 preprints in 17 sub- fields. Small; QA-only; no full-text. ChemNLP [29]1M abstracts + 100K PubChem records for NLP. No full-text; mixed licens- ing; no embeddings. ChemPile [16]250 GB multimodal data: text, SMILES, code, images. No retrieval structure; no embeddings. S2ORC[7,18], PMC- OA[20], CORE[17] Large general corpora with partial full-text and metadata. Not chemistry-specific; li- cense variability. Lit2VecLit2Vec workflow / screened study corpus constructed from S2ORC; code, schema, and legally safe derivatives released. Full reconstructed corpus not publicly redistributed; text-only. Table 5: Comparison of chemistry and general scientific corpora by scope and features. You are an information-extraction model for chemistry research. TASK: Read chemistry abstracts, output valid JSON only (no explanations, Markdown, or extra text), precisely matching the schema below. General Rules: ā¢Always include numeric value units, using standard units only. ā¢Donāt invent or guess values or details not explicitly stated. ā¢Use only standard units: %, ⦠C, K, atm, bar, Pa, kPa, MPa, Torr, psi; h, min, s, ms, μs, ns; L, mL, μL; mol, mmol, μmol, mol L ā1 (M); g, mg, μg, kg; ppm, ppb, wt %, vol %, mol %; kJ mol ā1 , J mol ā1 , eV; W m ā2 , mW cm ā2 ; A cm ā2 , mA cm ā2 , μA cm ā2 ; V, mV; S cm ā1 , μS cm ā1 ; F g ā1 , mAh g ā1 , mAh cm ā2 ; nm, μm, Ć ; cm ā1 ; quantum yield (%), etc. Summary (1ā2 sentences,ā¤50 words): ā¢Clearly state material/method and main chemical problem. ā¢Clearly state key numeric results with midpoint values and brief significance or application. ā¢Use plain English, active voice, capital first letter, period at the end. ā¢Do not include citations, funding. ā¢The summary must always include at least one explicit numeric midpoint with its unit, unless no numeric data is available. Do not generalize or omit the number if itās in the abstract. Field Classification: ā¢āfield_classificationā must be an array of one or more strings. ā¢Each string must exactly match one of the approved main fields below. ā¢If the abstract clearly covers multiple domains, include multiple strings. ā¢Approved Main Fields (must match exactly): Catalysis, Organic Chemistry, Polymer Chemistry, Inorganic Chemistry, Materials Science, Analytical Chemistry, Physical Chemistry, Biochemistry, Environmental Chemistry, Energy Chemistry, Medicinal Chemistry, Chemical Engineering, Supramolecular Chemistry, Radiochemistry & Nuclear Chemistry, Forensic & Legal Chemistry, Food Chemistry, Chemical Education 35 Output Schema: "field_classification": [], "summary": "" OUTPUT Return only the pure JSON, matching this style. No explanations. No extra text. No formatting. Example Input: Abstract: We synthesized a TiO 2 /g-C 3 N 4 heterojunction that eļ¬iciently converts CO 2 to CO under visible light (Ī»>420nm, 100 mW cm ā2 ). At 25 ⦠C and 1 atm, the catalyst delivered a CO production rate of 115ā125 μmol g ā1 h ā1 with 85ā90 % CO selectivity over 5 h in the presence of 0.10 M triethanolamine (TEOA) sacrificial agent (pH 7). The apparent quantum yield was 15±1 %, and 93 % of the initial activity was retained after five cycles, indicating good stability. Photoluminescence and time-resolved spectroscopy confirmed suppressed charge recombination. This work highlights a robust, metal-free approach for solar-driven CO production from CO 2 . Output: "field_classification": ["Catalysis", "Materials Science"], "summary": "TiOļææ/g-CļææNļææ heterojunction photocatalyst reduces COļææ to CO under visible light, achieving 87.5 % CO selectivity and 120 μmol g￿¹ h￿¹ production rate, with 15 % quantum yield and 93 % activity retention after five cycles." 7.3 Subfield Classifier Resources and Label Mapping The subfield classifier was trained using chemistry-related categories (this set includes 17 chem- istry subfields and a fallback Others class, for a total of 18 labels), each mapped to a unique index as shown in the table below. These indices correspond to the modelās output classes during training and evaluation. SubfieldIndex SubfieldIndex SubfieldIndex Catalysis0Physical Chemistry6Supramolecular Chemistry12 Organic Chemistry 1Biochemistry7Radiochemistry & Nuclear Chemistry 13 Polymer Chemistry 2Environmental Chemistry 8Forensic & Legal Chemistry14 Inorganic Chemistry 3Energy Chemistry9Food Chemistry15 Materials Science4Medicinal Chemistry10Chemical Education16 Analytical Chemistry 5Chemical Engineering11Others17 7.4 Annotation Validation Two domain experts (Jamile and Sara) independently validated a stratified subset of 25 abstracts each, assessing (i) multi-label field classification and (i) TL;DR summary quality. Overall ratings were strongly positive. For field classification, average scores wereμ Jamile = 4.63±0.65 (median= 5) andμ Sara = 4.88±0.45(median= 5). For summaries, averages wereμ Jamile = 4.38±0.92andμ Sara = 5.00±0.00(all scores on a 1ā5 scale). 36 Experts also provided targeted corrective feedback, revising 28ā48% of items (Jamile) and 8ā12% (Sara), suggesting localized improvements rather than systematic issues. Cross-expert overlap comprised 5 items for field ratings and 4 for summary ratings (aligned by normalized titles). Because overlap was limited and one rater produced a degenerate distribution (all 5ās for summaries), inter-rater reliability measures such as CohenāsĪŗand correlation coeļ¬icients were undefined. As an alternative, we report per-expert distributions alongside agreement rates. For field classification, exact agreement was 60% and agreement within one point was 80%. For sum- maries, exact agreement was 75% and agreement within one point was 100%. These results indicate broad consistency across experts despite small overlap. The full validation results and analysis scripts are provided in our GitHub repository:an- notation_validation. 7.5 Abstractive Summarization: Training, Evaluation, and Benchmarks This section provides full implementation and evaluation details for the abstractive summa- rization pipeline described in Section2.7of the main text. Benchmark results are reported in Table6. We used the curated dataset described in Section2.6, comprising 19,992 chemistry research abstracts licensed under C-BY. We adopt the predefined train/validation/test splits provided with the dataset (80/10/10). Inputs are truncated to a maximum of 1,024 tokens and targets to 128 tokens. Tokenization uses the Hugging Face tokenizer (modern text_target API). We fine-tuned sshleifer/distilbart-cnn-12-6 using the Hugging Face Seq2SeqTrainer on a sin- gle NVIDIA GeForce Ray Tracing Texel eXtreme (RTX) 3090 graphics processing unit (GPU) with mixed precision (16-bit floating point (FP16) enabled when Compute Unified Device Archi- tecture (CUDA) is available). The optimizer used was AdamW with a learning rate of2Ć10 ā5 . Training was performed for 5 epochs with a per-device batch size of 4 and gradient accumulation over 4 steps, resulting in an effective batch size of 16. Evaluation, saving, and logging occurred at fixed intervals, with eval_steps= 1000, save_steps= 1000, and logging_steps= 500. A fixed random seed (42) was applied across Python, NumPy, and PyTorch to ensure reproducibility, with deterministic CuDNN settings enabled where applicable. During evaluation, the per-device batch size was set to 4. The āSec.ā column in Table6 reports the total wall-clock time required to summarize the test set using these settings on a single RTX 3090 GPU. ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Lsum scores were computed using the Hug- ging Face evaluate package with stemming enabled. Prior to scoring, predictions and references were stripped of whitespace and segmented into sentences using NLTK to match standard ROUGE evaluation practices. All ROUGE scores are reported as percentages, where higher values indicate better overlap with the reference summaries. BERTScore (precision, recall, and F1) was computed in a separate evaluation pass using the evaluate implementation with the RoBERTa-large model and IDF weighting. These scores are also reported in Table 6alongside ROUGE metrics. The results in Table 6show that our fine-tuned model (lit2vec_tldr_finetuned) consistently 37 ModelR-1 R-2 R-L ROUGE-Lsum BERT P BERT R BERT F1 Sec. lit2vec_tldr_finetuned56.38 31.00 44.2845.6490.81 92.14 91.46255.16 distilbart_cnn_12_6_zeroshot37.67 15.57 26.6830.5487.1488.0287.57318.94 bart_large_cnn37.00 15.02 25.9030.0187.1287.8087.44402.62 lexrank35.65 14.78 24.7328.5785.7488.7987.2215.05 lead333.74 13.79 23.5327.4485.4588.0686.7113.01 scibert_extractive33.74 13.79 23.5327.4485.4588.0686.7124.07 Table 6: Test-set performance across abstractive and extractive summarization models. ROUGE and BERTScore metrics are reported as percentages (higher is better). Sec. indicates total evaluation runtime on a single NVIDIA GeForce RTX 3090 under the default generation settings used by the training script. outperforms all baseline methods across both ROUGE and BERTScore metrics. Compared to the strongest zero-shot abstractive baseline (distilbart_cnn_12_6_zeroshot), our model improves ROUGE-1 by +18.71 and BERT F1 by +3.89, demonstrating the effectiveness of domain-specific fine-tuning. While extractive methods such as lead3 and lexrank offer signifi- cantly faster inference times (13ā15 seconds), they lag behind in all quality metrics, highlighting the limitations of purely extractive summarization for dense scientific text. Our model achieves a strong balance of accuracy and eļ¬iciency, decoding 36.6% faster than the larger bart_large_cnn baseline while yielding substantially better output quality. Overall, the trade-off between qual- ity and speed favors the fine-tuned DistilBART model for integration in production chemistry pipelines. 7.6 Qualitative Examples of Abstractive Summarization Performance To qualitatively assess the performance of our fine-tuned summarization model, we present three representative examples from the test set. Each example includes the original abstract, the reference summary generated using GPT-4o (used as the supervision target during fine-tuning), and the predicted summary produced by our fine-tuned DistilBART model. The examples illustrate cases of strong, moderate, and weak performance based on how well the predicted summary aligns with the key information, structure, and factual content of the reference. Example 1: High-Quality Prediction Abstract: Ultraviolet B (UVB; 290ā320 nm) irradiation-induced lipid peroxidation induces inflammatory responses that lead to skin wrinkle formation and epidermal thickening. Peroxisome proliferator- activated receptor (PPAR)α/γdual agonists have the potential to be used as anti-wrinkle agents because they inhibit inflammatory response and lipid peroxidation. In this study, we evaluated the function of 2-bromo-4-(5-chloro-benzo[d]thiazol-2-yl) phenol (MHY 966), a novel synthetic PPARα/γdual agonist, and investigated its anti-inflammatory and anti-lipid peroxidation effects. The action of MHY 966 as a PPARα/γdual agonist was also determined in vitro by reporter gene assay. Additionally, 8-week-old melanin-possessing hairless mice 2 (HRM2) were exposed to 150 mJ/cm 2 UVB every other day for 17 days and MHY 966 was simultaneously pre-treated every day for 17 days to investigate the molecular mechanisms involved. MHY 966 was found to stimulate the transcriptional activities of both PPARαandγ. In HRM2 mice, we found that the skins of mice exposed to UVB showed significantly increased pro-inflammatory 38 mediator levels (NF-ļææB, iNOS, and COX-2) and increased lipid peroxidation, whereas MHY 966 co-treatment down-regulated these effects of UVB by activating PPARαandγ. Thus, the present study shows that MHY 966 exhibits beneficial effects on inflammatory responses and lipid peroxidation by simultaneously activating PPARαandγ. The major finding of this study is that MHY 966 demonstrates potential as an agent against wrinkle formation associated with chronic UVB exposure. Reference Summary (GPT-4o): MHY 966, a novel synthetic PPARα/γdual agonist, reduces skin inflammation and lipid peroxidation induced by UVB exposure, showing potential as an anti-wrinkle agent. Predicted Summary (Fine-tuned model): MHY 966, a novel synthetic PPARα/γdual agonist, reduces UVB-induced lipid peroxidation and inflammatory responses in HRM2 mice exposed to 150 mJ/cm 2 UVB over 17 days, showing potential as an anti-wrinkle agent. Example 2: Moderate-Quality Prediction Abstract: Recent clinical trials of small interfering RNAs (siRNAs) highlight the need for robust delivery technologies that will facilitate the successful application of these therapeutics to humans. Ar- guably, cell targeting by conjugation to cell-specific ligands provides a viable solution to this problem. Synthetic RNA ligands (aptamers) represent an emerging class of pharmaceuticals with great potential for targeted therapeutic applications. For targeted delivery of siRNAs with aptamers, the aptamer-siRNA conjugate must be taken up by cells and reach the cytoplasm. To this end, we have developed cell-based selection approaches to isolate aptamers that inter- nalize upon binding to their cognate receptor on the cell surface. Here we describe methods to monitor for cellular uptake of aptamers. These include: (1) antibody amplification microscopy, (2) microplate-based fluorescence assay, (3) a quantitative and ultrasensitive internalization method (āQUSIMā) and (4) a way to monitor for cytoplasmic delivery using the ribosome in- activating protein-based (RNA-RIP) assay. Collectively, these methods provide a toolset that can expedite the development of aptamer ligands to target and deliver therapeutic siRNAs in vivo. Reference Summary (GPT-4o): The study addresses the targeted delivery of siRNAs using aptamers, with methods to monitor cellular uptake and cytoplasmic delivery of aptamer-siRNA conjugates, potentially enhancing therapeutic applications. Predicted Summary (Fine-tuned model): Cell-based selection approaches for aptamer ligands for targeted delivery of small interfer- ing RNAs are developed using cell-based methods like antibody amplification microscopy and microplate-based fluorescence assays to monitor cellular uptake and cytoplasmic delivery, aiding in the development of aptamer-siRNA conjugates for therapeutic applications. Example 3: Lower-Quality Prediction Abstract: This study evaluated the chemical composition, antioxidant, anti-inflammatory and anticancer 39 activities of a Euphorbia hirta L. extract. The antioxidant activities of whole E. hirta ethanol extract were determined by electron spin resonance spectrophotometric analysis of 1,1-diphenyl- 2-picryl-hydrazyl (DPPH), hydroxyl, and alkyl radical levels and by using an online high- performance liquid chromatography (HPLC)-2,2ā-azino-bis(3-ethylbenzothiazoline-6-sulfonic acid) assay. The E. hirta ethanol extract (0.5 mg/mL) exhibited DPPH-scavenging activity of 61.19% ±0.22%, while the positive control (0.5 mg/mL ascorbic acid) had 100%±0.22% activity. The concentration of the extract required to trap 50% of DPPH (IC50) was 0.205 mg/mL. Online HPLC analysis of the extract also showed strong antioxidant activity. The anti-inflammatory ac- tivity of the E. hirta extract was assessed in lipopolysaccharide-induced RAW 264.7 macrophages. The anti-inflammatory activity was highest in the presence of 200 μg/mL E. hirta extract, and nitric oxide production was decreased significantly (p <0.05). The extract also showed selective anticancer activity at a concentration of 100 μg/mL (p <0.05). These results indicated that E. hirta may warrant further investigation for the development of antioxidant, anti-inflammatory, and anticancer herbal medications. Reference Summary (GPT-4o): The study evaluated Euphorbia hirta extract for antioxidant activity, with a DPPH-scavenging rate of 61.19% at 0.5 mg/mL and an IC50 of 0.205 mg/mL, and observed significant anti- inflammatory and anticancer activities in cell models. Predicted Summary (Fine-tuned model): Euphorbia hirta L. extract exhibits strong antioxidant and anti-inflammatory activities, with 61.19% DPPH scavenging activity at 0.205 mg/mL and selective anticancer activity at 100 μg/mL, indicating potential for herbal medicine development. 7.7 Topic Classification: Training, Evaluation, and Benchmarks The classifier was trained on the enrichment dataset described in Section2.6, comprising 39,900 chemistry texts (abstracts and corresponding TL;DR summaries) and multi-label subfield an- notations. The taxonomy covers 18 labels, as described in Section7.3. We encode each text using the intfloat/e5-large-v2 SentenceTransformer [4], optimized for semantic similarity. Each document is mapped to a 1,024-dimensional, unit-normalized vector xāR 1024 (i.e.,ā„xā„ 2 = 1). No additional text normalization beyond the embedding modelās default preprocessing is applied. Embeddings serve as input to a feedforward multi-layer perceptron (MLP) implemented in TensorFlow [ 22]. The network has two hidden layers of 256 units each with ReLU activations, batch normalization for training stability, and dropout (rate= 0.3) after each hidden layer. The output layer uses elementwise sigmoid activations to produce per-label probabilities over C= 18subfields: Ė y=Ļ ( W 3 Ļ ( BN 2 (Ā·) ) +b 3 ) , Ļ(z) = max(0,z), Ė yā[0,1] C . Multi-label assignments are obtained by thresholding Ė yat a global decision thresholdĻ= 0.5. Training uses weighted binary cross-entropy to mitigate label imbalance. Lety ā ā0,1and Ėy ā ā[0,1]denote the ground-truth indicator and predicted probability for labelāā1, . . . , C. 40 WithNtraining samples andP ā positives for labelā, we set a per-label weight w ā = NāP ā P ā (ratio of negatives to positives). The loss for one instance is L= 1 C C ā ā=1 w ā [ āy ā log(Ėy ā )ā(1āy ā ) log(1āĖy ā ) ] . Optimization uses Adam (learning rate= 10 ā3 ), mini-batches of size 32, for up to 50 epochs. We employ early stopping (patience= 5) on validation loss and reduce the learning rate on plateau (factor= 0.5, patience= 3). Model parameters yielding the best validation loss are retained. We use 5-fold cross-validation on the pooled training and validation data (non-stratified due to the multilabel setting). The final test results are reported at a fixed decision threshold of Ļ= 0.5. Following standard practice, we report micro-averaged F1 (aggregating decisions across labels), macro-averaged F1 (unweighted mean across labels), weighted-average F1 (weighted by label support), and samples-average F1 (averaged over instances). 41 Figure 13: Analysis of predicted research subfields from document-level classification. Top row: (left) counts of all predicted subfields across documents and (right) counts of only the top-confidence subfield per document. Bottom row: distribution of prediction scores for each subfield when it appears as (left) the top-predicted category and (right) the second-predicted category (if applicable). Subfield names are abbreviated (legend) for clarity. Table7provides the per-label precision, recall, F1-score, and support on the held-out test set (thresholdĻ= 0.5). Consistent with the main text, frequent subfields (e.g., Biochemistry, Medicinal Chemistry, Materials Science) achieve higher F1-scores, while underrepresented do- mains (e.g., Chemical Engineering, Supramolecular Chemistry) remain challenging. Figure S 13summarizes the distribution of predicted subfields and confidence scores. Fig- ure S 14relates label support in the training set to test-set F1-scores, highlighting the impact of representation on performance. At inference, each abstract is embedded as above and passed through the MLP to obtain per- label probabilities. Unless otherwise specified, labels withĖy ā ā„Ļ(withĻ= 0.5) are assigned. For applications that prefer higher precision (e.g., curated knowledge graphs), users may raise Ļor apply label-specific thresholds calibrated on validation data. 42 Figure 14: Label support in the training set (bars, left axis) versus F1-score on the test set (red line, right axis) for each subfield. Frequent subfields such as Biochemistryand Medicinal Chemistry achieve high F1-scores, whereas underrepresented subfields including Chemical Engineering and Supramolecular Chemistry remain challenging. 43 Table 7: Per-label classification report on the test set (thresholdĻ= 0.5). Macro average is the unweighted mean across labels; micro average aggregates decisions across labels; weighted average weights by label support; samples average is the average across instances. LabelPrecision Recall F1-score Support Catalysis0.83 0.780.80197 Organic Chemistry0.85 0.600.70245 Polymer Chemistry0.80 0.650.72120 Inorganic Chemistry0.84 0.620.71203 Materials Science0.82 0.780.80917 Analytical Chemistry0.87 0.600.71633 Physical Chemistry0.78 0.530.63240 Biochemistry0.89 0.940.922106 Environmental Chemistry0.84 0.750.79508 Energy Chemistry0.75 0.830.79166 Medicinal Chemistry0.88 0.760.821343 Chemical Engineering0.75 0.410.53413 Supramolecular Chemistry0.68 0.680.6834 Radiochemistry & Nuclear Chemistry0.71 0.600.6520 Forensic & Legal Chemistry0.62 0.810.7016 Food Chemistry0.83 0.830.83282 Chemical Education0.85 0.850.8520 Others0.88 0.790.8319 Micro avg0.85 0.770.817482 Macro avg0.80 0.710.757482 Weighted avg0.85 0.770.807482 Samples avg0.87 0.790.807482 44 7.8 Data Records: Schema and Field Dictionaries KeyDescription _id (string)Internal unique identifier (optional). corpusid (integer)Mirror of corpus_id for traceability (optional). externalids (object)External IDs (any subset): DOI, PubMed, PubMedCentral, ArXiv, MAG, DBLP, ACL, CorpusId. url (string)Canonical landing page (e.g., Semantic Scholar). title (string)Full paper title. authors (array)List of authorId (string|int), name (string). venue (string|null)Journal or conference name. publicationvenueid (string|null) Venue identifier if available. year (integer)Publication year. referencecount, citation- count, influentialcita- tioncount (integer) Citation metrics. isopenaccess (boolean) Open access flag. s2fieldsofstudy (array) category (string), source (string). publicationtypes (string|array|null) Publication type(s). publicationdate (string|null) ISO 8601 date. journal (object|null)name (string), volume (string|null), pages (string|null). The internal corpus used in this study was represented as individual JSON records named corpus_id.json. We release the machine-readable JSON Schema and example record structure so that researchers can reconstruct compatible local records without requiring redistribution of the full corpus. For an overview of the top-level record fields, see Table 8. 45 FieldTypeDescription schema_versionstringSchema version (e.g., ā1.0ā). corpus_idintegerPrimary identifier (S2ORC CorpusId) for traceability and file naming. metadataobjectBibliographic metadata from S2ORC including title, authors, venue, year, URLs, and external IDs; see Table7.8. abstractstringOriginal abstract text from S2ORC. fulltextstringReconstructed full text (plain text). paragraphsarray[string]Paragraph-level text units used for retrieval (aligned with embed- dings by index). embeddingsarray[array[float]] Paragraph embeddings (model: intfloat/e5-large-v2; 1024-D, float32). abstract_embedding array[float]Embedding for the abstract (same model/dtype; length 1024). predicted_subfieldobjectMap from subfield label to confidence score in[0,1](17 chemistry subfields + Others; label mapping in section7.3). tldrstringTwo-sentence abstractive summary (DistilBART-CNN adapted to chemistry). unpaywall_licenseobject|nullLicense metadata from Unpaywall. crossref_licenseobject|nullLicense metadata from Crossref. openalex_licenseobject|nullLicense metadata from OpenAlex. Table 8: Top-level fields for each JSON record. Complete nested field dictionaries and a machine-readable schema are provided in the Supplementary Information. At the top level, a record is an object with the following required keys: schema_version, corpus_id, metadata, abstract, fulltext, paragraphs, embeddings, unpaywall_license, cross- ref_license, openalex_license. Optional keys: , abstract_embedding, tldr, predicted_subfield. The field paragraphs is an array of strings of lengthNā„1, and embeddings is a parallel array of lengthN, where each element is an array of 1024 float32 values generated with the intfloat/e5-large-v2 model; by construction, indexiin embeddings corresponds to indexiin paragraphs. The field abstract_embedding is a single array of 1024 float32 values (optional), while predicted_subfield is an object mapping strings to confidence scores in the range[0,1]. The license metadata fields unpaywall_license, crossref_license, and openalex_license are rep- resented as objects when available, or as null otherwise. Each license source (Unpaywall, Crossref, OpenAlex) is represented as an object (or null if unavailable). We preserve the upstream structure for provenance. Typical keys include: li- cense or best_oa_location.license (string), url (string), and where provided, embargo/host/type fields. Users should consult the upstream license URL to verify reuse terms. Paragraph text is produced by a token-aware recursive splitter with overlap to preserve coherence in downstream retrieval. For each paragraphp i , we store an embeddinge i āR 1024 computed with intfloat/e5-large-v2. Embeddings are serialized as arrays of float32. We guar- antee: 1.|paragraphs|=|embeddings|=N. 2.e i encodes paragraphs[i] (same index). 3.Model/dtype are fixed across all records (1024-D, float32). 46 Truncated example JSON record: "schema_version": "1.0", "corpus_id": 37254803, "metadata": "title": "...", "year": 2016, "externalids": "DOI": "10.x/x" , "url": "https://..." , "abstract": "Epigallocatechin gallate ...", "fulltext": "# Protective effect ...", "paragraphs": [37254803P0:"passage: ...", "..."], "embeddings": [37254803P0:[0.0123, -0.0456, ...], "..."], "abstract_embedding": [0.0345, -0.0789, 0.0123, "..."], "predicted_subfield": "Biochemistry": 0.997, "Medicinal Chemistry": 0.759, "tldr": "EGCG reduces lipid peroxidation ...", "unpaywall_license": "best_oa_location": "license": "c-by", "...", "crossref_license": "license": "http://creativecommons.org/licenses/by/4.0/", "...", "openalex_license": null A complete, untruncated example JSON record is provided in the code repository at exam- ples/record_full.json. It illustrates the full nested structures for metadata and license objects, the alignment between paragraphs and embeddings, and the numerical precision of stored vec- tors, without implying that the full corpus is redistributed. 7.9 Schema & Structural Validation We developed a schema-level validation framework to ensure all dataset records conform to a consistent, machine-readable standard. This is critical for maintaining data integrity, support- ing reliable downstream processing, and ensuring compatibility with external tools. The valida- tor enforces the Draft 2020-12 JSON Schema specification for Lit2Vec-style records, with fully deterministic and reproducible validation outcomes. In large datasets, small inconsistenciesā such as missing fields, incorrect types, or mis-sized arraysācan cause downstream errors, mis- interpretations, or ingestion failures. Schema validation addresses this by enforcing structural consistency, semantic correctness (e.g., plausible publication years), and interoperability with standard workflows and repositories. It also ensures auditability through reproducible, times- tamped error reporting. Tables9and10summarize the validation rules and provide output examples, respectively. Table 9: Schema Validation Rules and Diagnostic Flags for Lit2Vec-Style Records Validation RuleDescriptionFlag (on failure) Top-level Required Fields Must include: schema_version, corpus_id, metadata, ab- stract, paragraphs, embeddings, tldr, abstract_embedding, pre- dicted_subfield. missing_<field> ā Raised if any required field is missing (e.g., missing_abstract). Continued on next page 47 Validation RuleDescriptionFlag (on failure) Additional Proper- ties Extra fields not defined in the schema are disallowed at all levels (top-level and nested). additional_property_<field> ā Raised when unexpected field is present (e.g., ad- ditional_property_validation_info). Type Enforcement Fields must conform to expected types: strings, integers, booleans, arrays, or objects. type_mismatch_<field> ā Type does not match schema definition (e.g., string provided where object expected). Fixed-Length Arrays embeddingsandab- stract_embedding must be numeric arrays of exactly length 1024. too_short_<field> or too_long_<field> ā Array length is incorrect (e.g., too_short_abstract_embedding). Paragraph Key Pat- tern Keys in paragraphs and embed- dings must match regex +P +$. pattern_violation_<field> ā Raised when key format is invalid (e.g.,1Para2). Field Format Con- straints Fields like metadata.url must be valid URIs; publicationdate must be ISO 8601 date. invalid_format_<field> ā Raised for malformed dates or URLs (e.g., in- valid_format_metadata_url). Nullability RulesSome fields may be null (e.g., publi- cationvenueid), others must not be. type_mismatch_<field> ā Raised if null is used in a non-nullable field. Required Metadata Fields metadata must include: title, au- thors (with name), venue, year. missing_<field> ā Raised for any miss- ing required metadata subfield (e.g., miss- ing_metadata_title). Numeric Ranges in Metadata year must be between 1800ā2100; citation/reference counts must be ā„0. value_below_minimum_<field>or value_above_maximum_<field> ā Value out of bounds (e.g., value_below_minimum_metadata_year). Authors Field Struc- ture Each author must include name; no extra fields allowed. missing_name ā Name missing in au- thor object; additional_property_<field> ā Extra field present. Abstract and TLDR abstract and tldr must be strings. type_mismatch_<field> ā Raised if value is not a string (e.g., array given in- stead). Paragraph Content Each key in paragraphs must map to a string value. type_mismatch_paragraphs ā Raised when a paragraph value is not a string. Controlled Vocabu- lary predicted_subfield must contain only known keys, with float values in the range [0.0, 1.0]. type_mismatch_predicted_subfield, value_below_minimum_predicted_subfield, or value_above_maximum_predicted_subfield. Validation Outcome A record is marked āpassā only if it satisfies all schema constraints without triggering any flags. Oth- erwise, it is marked āfailā. Not a flag ā Outcome is tracked via sta- tus: āpassā or āfailā; includes summary and error count. 48 Table 10: Example schema validation outputs: failing (left) and passing (right). Example Failing RecordExample Passing Record "schema_validation": "status": "fail", "summary": "1 schema validation error(s) found.", "metrics": "validation_errors": [ "message": "[] is too short", "field": "abstract_embedding", "schema_path": "properties/abstract_embedding/minItems", "validator": "minItems" ], "flags": "too_short_abstract_embedding": true , "checked_at": "2025-07-30T23:41:30.136585" "schema_validation": "status": "pass", "summary": "Document conforms to the JSON schema.", "metrics": "validation_errors": [], "flags": , "checked_at": "2025-07-30T12:08:22.875185" 7.10 Metadata Validation To ensure structural integrity, completeness, and semantic correctness of the datasetās bibliographic records, we implemented a strict metadata validation pipeline. Each record is checked against required field definitions, expected types, semantic rules (e.g., date consistency, author structure), and dataset-specific constraints. The pipeline distinguishes between hard errors (e.g., type mismatches) and soft warnings (e.g., empty strings or missing licenses), enabling both robust schema enforcement and fine-grained data curation. Validation covers presence and typing of critical fields such as title, authors, year, venue, publicationdate, and s2fieldsofstudy, among others. Fields like authors are recursively validated to ensure proper structure (e.g., non- empty names and valid author IDs), while open-access metadata is cross-checked against expected license fields. Publication dates are parsed in a timezone-aware fashion and compared with the declared year. Discrepancies and future-dated entries are flagged. Tables11and12summarize the validation rules and provide output examples, respectively. Table 11: Metadata Validation Rules and Diagnostic Flags Validation RuleDescriptionFlag Missing Required FieldsRequired metadata fields are missing. missing:<field> Type MismatchField type does not match expected type (e.g., string vs int). type:<field>:<found_type>_to_<expected_type> Empty FieldsField is present but empty (e.g., empty string, list, or dict). empty:<field> Title Too ShortTitle has fewer than 5 characters (pos- sible placeholder). title_short Malformed Author EntriesAuthor missing name or ID, or not a valid dict. authors_malformed Publication Year MismatchParsed publicationdate year disagrees with year field. year_vs_date Continued on next page 49 Validation RuleDescriptionFlag Future Publication DatePublication date is after current Coor- dinated Universal Time (UTC) date. date_in_future Invalid Date Formatpublicationdate is unparseable.date_bad_format Year Out of Rangeyear is outside 1800ā(current_year + 1). year_out_of_range Missing Chemistry FOSNo s2fieldsofstudy entry with category āChemistryā. fos_no_chemistry Empty External IDsAll external ID fields (e.g., DOI) are empty or missing. externalids_empty Invalid Publication TypesOne or more publicationtypes items are not valid non-empty strings. pubtypes_bad_item Missing OA Licensesisopenaccess = true but no license fields are present. oa_missing_licenses Table 12: Example metadata validation outputs: warning (left) and passing (right). Warning RecordPassing Record "metadata_validation": "status": "warn", "summary": "Metadata field & content validation", "metrics": "error_count": 0, "warning_count": 1, "errors": [], "warnings": [ "empty:venue" ], "diagnostics": "missing_fields": [], "wrong_type_fields": [], "empty_fields": [ "venue" ], "malformed_authors": [], "title_short": false, "year_mismatch": false, "future_date": false, "date_issue": null, "no_chemistry_field": false, "oa_missing_licenses": [] , "checked_at": "2025-07-30T12:10:51.411572" "metadata_validation": "status": "pass", "summary": "Metadata field & content validation", "metrics": "error_count": 0, "warning_count": 0, "errors": [], "warnings": [], "diagnostics": "missing_fields": [], "wrong_type_fields": [], "empty_fields": [], "malformed_authors": [], "title_short": false, "year_mismatch": false, "future_date": false, "date_issue": null, "no_chemistry_field": false, "oa_missing_licenses": [] 7.11 License Validation To support record-level license provenance review and machine-readable compliance screening, we implemented an automated license validation pipeline that normalizes, compares, and resolves licensing information from 50 three metadata sources: Crossref, Unpaywall, and OpenAlex. Licensing inconsistencies, vague descriptors, or non-standard license formats are common in large bibliographic datasets, and can impact downstream reuse, redistribution, and compliance. Our validator applies a flexible parsing strategy that extracts the most explicit license statement from each source, normalizes it to a canonical identifier (e.g., c-by, c0, public-domain), and determines a resolved license by checking for agreement across sources. Non-informative values (e.g., unknown, other-oa) are excluded from agreement checks. When licenses from multiple sources agree, the record is treated as having stronger metadata support under the studyās screening rule. Conflicts between informative licenses are explicitly captured using conflict expressions (e.g., conflict:c- by_vs_closed). A record passes validation only when the resolved license is considered open-access, confirmed by at least two independent sources, and free from conflicts under the metadata-agreement rule. All other outcomesāincluding restrictive licenses (e.g., c-by-nd) or source disagreementāresult in a āfailā status. This validation step provides compliance evidence for the study rule; it does not by itself establish a legal guarantee for public redistribution of the exact reconstructed text. Each validation result includes the resolved license, contributing sources, raw normalized inputs, a conflict indicator, and a UTC timestamp. The validator supports parallel processing for large-scale evaluation and produces reproducible, auditable outputs. Tables 13and14summarize the validation rules and provide output examples, respectively. Table 13: License Validation Rules and Conflict Flags Validation RuleDescriptionFlag (on failure) Raw Field Extraction License fields are extracted from raw or nested metadata under Crossref, Unpaywall, and Ope- nAlex. missing_license_<source> ā Raised when no usable license is found in source. Normalization to Canon- ical ID Raw values are mapped to canon- ical forms using direct mappings or regex (e.g., c-by, c0, public- domain). unknown_license ā Raised if license is unrecognized or ambiguous. Source Agreement Logic If two or more sources provide the same informative license, that li- cense is accepted as resolved. license_conflict ā Raised when two or more informative sources disagree. Three-Way Conflict De- tection If all three sources provide differ- ent informative licenses, a multi- way conflict is recorded. license_conflict ā Set to true and conflict expression recorded (e.g., conflict:c-by_vs_closed_vs_c0). ResolvedLicense Recording The final resolved license is stored explicitly in resolved_license. Not a flag ā Used for downstream compliance filters. Open License Whitelist Licenses must be from an accepted set (e.g., c0, c-by, c-by-sa). fail ā If resolved license is missing, re- strictive, or not from allowed list. Conflict-Free Validation Final license must be agreed upon by at least two sources and be conflict-free. fail ā If only one informative source or a conflict is detected. 51 Table 14: Example license validation outputs: failing (left) and passing (right). Failing RecordPassing Record "license_validation": "status": "fail", "summary": "Checks licence consistency and open-access compliance.", "metrics": "resolved_license": "conflict:closed_vs_c-by-nc", "license_source": "unpaywall+openalex", "license_conflict": true, "input_licenses": "crossref": "missing", "unpaywall": "closed", "openalex": "c-by-nc" , "checked_at": "2025-07-30T23:41:30Z" "license_validation": "status": "pass", "summary": "Checks licence consistency and open-access compliance.", "metrics": "resolved_license": "c-by", "license_source": "crossref+unpaywall", "license_conflict": false, "input_licenses": "crossref": "c-by", "unpaywall": "c-by", "openalex": "missing" , "checked_at": "2025-07-30T12:08:22Z" 7.12 Text Validation To verify the structural integrity, linguistic quality, and semantic consistency of the textual components in the screened study corpus, we implemented an automated text validation pipeline. This validator operates on both the abstract and the fulltext fields of each record. It performs a comprehensive suite of checks spanning length requirements, Unicode integrity, formatting quality, language detection, semantic overlap, and embedding consistency. The pipeline first applies character and sentence-length thresholds to identify truncated or underspecified content. It then scans the text for corrupted characters (e.g., Unicode replacement characters), invisible for- matting symbols, and control characters, reporting examples when found. Additional formatting quality checks include whitespace density and ASCII letter share, both of which help identify structural anomalies such as excessive markup or malformed text. Language detection is applied to the beginning of the full text to verify alignment with the expected language (typically English), and semantic alignment between the abstract and body is measured using ROUGE-1 recall between the abstract and the first 2,000 characters of the full text. If an abstract_embedding vector is present, its L2 norm is compared against an expected range to ensure numerical consistency. Each record is assigned a validation status of āpassā, āwarnā, or āfailā depending on which checks are triggered. Failures are reserved for critical issues (e.g., missing embeddings or extremely short full text), while warnings denote potentially problematic but non-blocking anomalies. Diagnostic flags and metrics are included in the output for auditability and triage. Tables15and16summarize the validation rules and provide output examples, respectively. Table 15: Text Validation Rules and Diagnostic Flags Validation RuleDescriptionFlag (on failure) Abstract Character Length Abstract must be at least 100 charac- ters. abstract_too_short Full Text Character Length Full text must be at least 1000 char- acters. fulltext_too_short Continued on next page 52 Validation RuleDescriptionFlag (on failure) Abstract Sentence CountMust contain at least 2 sentence- ending punctuation marks. abstract_low_sentence_count Full Text Sentence CountMust contain at least 50 sentence- ending punctuation marks. fulltext_low_sentence_count Heading Markers in Full Text Detects Markdown or numbered headings (e.g., # Introduction). fulltext_missing_heading_markers Corrupted CharactersUnicodereplacementchar (U+FFFD) present in text. abstract_has_corrupted_chars, fulltext_has_corrupted_chars Whitespace DensityNon-whitespace density must exceed threshold (abstract: 0.75, full text: 0.83). abstract_low_whitespace_ratio, fulltext_low_whitespace_ratio ASCII Letter Share ASCII letter proportion must exceed threshold (abstract: 0.7, full text: 0.75). abstract_low_ascii_ratio, full- text_low_ascii_ratio Language DetectionMust match expected language (e.g., English) withā„0.9 confidence. language_mismatch_or_low_confidence Semantic Overlap (ROUGE- 1) ROUGE-1 recall between abstract and full text (first 2000 chars) must beā„0.5. low_rouge1_overlap Abstract Embedding Pres- ence Must be present and valid numeric vector. abstract_embedding_missing_or_invalid Abstract Embedding Norm Must be within tolerance of expected L2 norm (default: 1.0 ± 1e-3). abstract_embedding_norm_off Validation StatusStatus is āfailā if any critical flags are raised; otherwise āwarnā or āpassā. Not a flag ā See status field 53 Table 16: Example text validation output: warning (left). Warning Record(Placeholder for Passing Record) "text_validation": "status": "warn", "flags": "abstract_too_short": false, "fulltext_too_short": false, "abstract_has_corrupted_chars": false, "fulltext_has_corrupted_chars": false, "abstract_low_whitespace_ratio": false, "fulltext_low_whitespace_ratio": false, "abstract_low_ascii_ratio": false, "fulltext_low_ascii_ratio": false, "abstract_low_sentence_count": false, "fulltext_low_sentence_count": false, "fulltext_missing_heading_markers": false, "language_mismatch_or_low_confidence": false, "low_rouge1_overlap": true, "abstract_embedding_norm_off": false, "abstract_embedding_missing_or_invalid": false , "metrics": "abstract_chars": 1344, "fulltext_chars": 15561, "abstract_bad_chars": "valid": 1145, "corrupted": 0, "control": 0, "formatting": 0, "unassigned": 0, "whitespace": 199 , "fulltext_bad_chars": "valid": 13300, "corrupted": 0, "control": 0, "formatting": 0, "unassigned": 0, "whitespace": 2261 , "abstract_ws_ratio": 0.8519345238095238, "fulltext_ws_ratio": 0.8547008547008547, "abstract_ascii_ratio": 0.8050595238095238, "fulltext_ascii_ratio": 0.80573227941649, "abstract_sentence_count": 5, "fulltext_sentence_count": 104, "fulltext_heading_markers": 9, "language": "en", "language_score": 0.9999956877319424, "rouge1_recall": 0.45145631067961167, "abstract_embedding_norm": 1.0 , "examples": "abstract_problem_chars": [], "fulltext_problem_chars": [] , "checked_at": "2025-07-22T06:38:13.266722+00:00" 54 7.13 Chunk-Level Validation To ensure the internal consistency and semantic usability of paragraph-level text and embedding data, we im- plemented an automated chunk-level validation workflow. This process evaluates both the linguistic and vector components of each paragraph in a record to verify structure, dimensionality, and encoding correctness. These checks are critical in large-scale embedding datasets, where even minor irregularities can corrupt downstream model behavior. The validation pipeline analyzes each documentās paragraphs and their associated embedding vectors. It confirms the presence of both, measures token lengths using a consistent tokenizer (intfloat/e5-large-v2), and enforces dataset-wide length thresholds. Token lengths falling outside predefined minimum or maximum limits are flagged. Paragraph text is also inspected for Unicode compliance by counting valid, corrupted (ļææ), con- trol, formatting, and unassigned characters, revealing encoding issues that may not surface through standard inspection. Embeddings are validated in parallel. Each paragraph must be paired with a numeric array of fixed length (1024). Missing embeddings or incorrect-length vectors are flagged as critical errors. The overall chunk validation outcome is determined by the presence of such errors: records with any embedding failures are marked āfailā; records with token length violations but valid embeddings are marked āwarnā; records with no issues receive a āpassā. The validator scales across large corpora using multithreading or multiprocessing and provides detailed diagnostics per paragraph, aggregate metrics, and a UTC timestamp to support auditability and reproducibility. Tables 17and18summarize the validation rules and provide output examples, respectively. Table 17: Chunk-Level Validation Rules and Diagnostic Flags Validation RuleDescriptionFlag (on failure) Empty ParagraphsParagraph value is missing, empty, or not a string empty_chunks ā Count of non- usable paragraph entries Token Length Too ShortParagraph has fewer than 100 tokens chunks_too_short ā Count of under-length paragraphs Token Length Too LongParagraph has more than 300 tokens chunks_too_long ā Count of over-length paragraphs Corrupted CharactersParagraph contains replacement char- acters (ļææ) char_counts.corrupted ā Ag- gregated count per paragraph Control or Formatting Char- acters Paragraph includes invisible control (Cc), formatting (Cf), or unassigned (Cn) code points char_counts.control, format- ting, unassigned ā Diagnostic only Missing EmbeddingNo vector found for paragraph keymissing_embeddings ā Count of missing embeddings Invalid Embedding Shape Vector is not a list of length 1024invalid_embedding_vectors ā Count of malformed vectors Validation OutcomeEmbedding errors ā āfailā; Token length errors only ā āwarnā; No issues ā āpassā Not a flag ā outcome recorded in status, with human-readable summary 55 Table 18: Example chunk validation output: warning status with one chunk under the minimum token limit. Warning Record "chunk_validation": "status": "warn", "summary": "0 chunk(s) over max tokens; 1 under min tokens.", "metrics": "paragraph_count": 36, "token_length_distribution": "Q1": 147.75, "Q2": 165.5, "Q3": 183.25, "min": 56, "max": 239, "mean": 166.38888888888889 , "chunks_too_long": 0, "chunks_too_short": 1, "empty_chunks": 0, "missing_embeddings": 0, "invalid_embedding_vectors": 0, "max_token_limit": 300, "min_token_limit": 100 , "paragraphs": "9991603P0": "token_length": 201, "char_counts": "valid": 864, "corrupted": 0, "control": 0, "formatting": 0, "unassigned": 0 , "problem_characters": [] , "9991603P1": "token_length": 232, "char_counts": "valid": 1068, "corrupted": 0, "control": 0, "formatting": 0, "unassigned": 0 , "problem_characters": [] , . . . "9991603P35": "token_length": 56, "char_counts": "valid": 330, "corrupted": 0, "control": 0, "formatting": 0, "unassigned": 0 , "problem_characters": [] , "checked_at": "2025-07-22T06:37:27.593943+00:00" 56 7.14 Embedding Validation To ensure the structural integrity and semantic fidelity of paragraph-level vector representations in Lit2Vec- style records, we implemented a dedicated embedding validation framework. This framework checks alignment between paragraph texts and their corresponding embeddings, enforces vector shape and type constraints, and confirms semantic consistency through spot-checked re-encoding. Each embedding must be a finite, normalized 1024-dimensional float vector that accurately represents the source paragraph, as verified by cosine similarity to a reference model. The validation procedure begins by verifying that both paragraphs and embeddings fields are structured as dictionaries with matching keys. Paragraph identifiers must exist in both fields, with no missing or extra embeddings. Each vector must conform to strict structural constraints, including exact dimensionality, finiteness, and approximate unit norm. To test reproducibility, a random subset of paragraphāembedding pairs is re-encoded using a reference Sentence-Transformer model (e.g., intfloat/e5-large-v2), and cosine similarity is computed. Records pass this test only if all sampled pairs meet or exceed a defined similarity threshold. Validation is deterministic, reproducible, and supports both sequential and parallel processing. The final out- put includes status, summary, detailed metrics (including alignment and similarity scores), and a UTC timestamp for auditing. Tables 19and20summarize the validation rules and provide output examples, respectively. Table 19: Embedding Validation Rules and Diagnostic Flags for Lit2Vec-Style Records Validation RuleDescriptionFlag (on failure) Field Type Checkparagraphs and embeddings must be dictionaries. invalid_structure ā Raised when fields are not of type dict. ParagraphāEmbedding Alignment Each paragraph must have a corre- sponding embedding and vice versa. missing_embedding,ex- tra_embedding ā Raised for mis- matched IDs. Vector LengthEach embedding must have exactly 1024 float elements. invalid_shape_embedding ā Raised for incorrect dimensionality. Finite Float ValuesAll elements in embedding must be finite 32-bit floats. nonfinite_values_embeddingā Raised if NaNs or Infs are present. Unit Norm (Approxi- mate) Embedding must be approximately unit-normalized within a 0.05 toler- ance. unnormalized_embedding ā Raised if norm deviates from 1.0 beyond toler- ance. CosineSimilarity Threshold Randomly sampled paragraphā embedding pairs must match re-encoded vectors (cosine similar- ity ļææ threshold). cosine_mismatch ā Raised if any sampled vector falls below similarity threshold. Pass CriteriaAll structural and reproducibility checks must pass. No specific flag ā Outcome recorded as status: āpassā or āfailā. 57 Table 20: Example embedding validation outputs: failing (left) and passing (right). Failing RecordPassing Record "embedding_validation": "status": "fail", "summary": "Embedding validation complete", "metrics": "paragraphs": 5, "embeddings": 5, "alignment": "missing": [], "extra": [], "bad": ["p3"] , "determinism": "mean_cos": 0.821, "min_cos": 0.765, "max_delta": 0.235, "threshold": 0.01, "passed": false , "checked_at": "2025-07-30T22:05:49.882404" "embedding_validation": "status": "pass", "summary": "Embedding validation complete", "metrics": "paragraphs": 6, "embeddings": 6, "alignment": "missing": [], "extra": [], "bad": [] , "determinism": "mean_cos": 0.998, "min_cos": 0.996, "max_delta": 0.004, "threshold": 0.01, "passed": true , "checked_at": "2025-07-30T22:08:41.208336" 7.15 Predicted Subfield Validation To ensure that predicted disciplinary subfields are accurate, standardized, and suitable for downstream analysis, we implemented an automated subfield validation pipeline. This validation addresses common issues stemming from machine learningābased classification outputs or heterogeneous annotations, such as inconsistent label for- mats, incomplete predictions, and invalid probability scores. Each recordās predicted_subfield field is expected to be a dictionary mapping subfield names (drawn from a fixed controlled vocabulary) to floating-point confidence scores in the range [0.0, 1.0]. The controlled vocabulary includes core and specialized areas of chemistry, such as Catalysis, Supramolecular Chemistry, Forensic Chemistry, and others. The validator performs multiple checks: ā¢Presence and correct data type of the predicted_subfield field. ā¢Non-emptiness of the prediction map (optional but recommended). ā¢All labels must match entries in the controlled vocabulary. ā¢All scores must be floats within [0.0, 1.0]. Tables21and22summarize the validation rules and provide output examples, respectively. Table 21: Validation Rules and Diagnostic Flags for Predicted Subfield Assignments Validation RuleDescriptionFlag (on failure) Field PresenceThe predicted_subfield field must exist in the record. missing_predicted_subfield ā Field is entirely absent. Continued on next page 58 Validation RuleDescriptionFlag (on failure) Type ConstraintThe field must be a dictionary.not_a_dict ā Field is not a dict (e.g., list or string given). Empty Dictionary Check Empty predictions are optionally disal- lowed. empty_predictions ā No pre- dicted subfields found. Vocabulary MatchEach subfield label must match a known entry in the controlled vocabulary. contains_invalid_labels ā Un- known label(s) detected. Score ValidityEach score must be a float in the range [0.0, 1.0]. contains_invalid_scores ā Value is non-float or out-of- bounds. Validation OutcomeDetermines status: āpassā, āwarnā, or āfailā based on aggregated rule checks. Not a flag ā Derived from com- bination of failure conditions. Table 22: Example predicted subfield validation outputs: failing (left), warning (center), and passing (right). Failing RecordWarning RecordPassing Record "predicted_subfield_validation": "status": "fail", "flags": "missing_predicted_subfield": true, "not_a_dict": false, "contains_invalid_labels": false, "contains_invalid_scores": false, "empty_predictions": false , "metrics": "label_count": 0, "invalid_labels": [], "invalid_scores": [] , "checked_at": "2025-07-30T23:41:30.136585Z" "predicted_subfield_validation": "status": "warn", "flags": "missing_predicted_subfield": false, "not_a_dict": false, "contains_invalid_labels": true, "contains_invalid_scores": true, "empty_predictions": false , "metrics": "label_count": 2, "invalid_labels": [ "Quantum Wizardry" ], "invalid_scores": [ "Catalysis": 1.2 ] , "checked_at": "2025-07-30T23:41:31.004728Z" "predicted_subfield_validation": "status": "pass", "flags": "missing_predicted_subfield": false, "not_a_dict": false, "contains_invalid_labels": false, "contains_invalid_scores": false, "empty_predictions": false , "metrics": "label_count": 3, "invalid_labels": [], "invalid_scores": [] , "checked_at": "2025-07-30T23:41:32.288271Z" 7.16 Summary Quality Validation To assess the quality of automatically generated summaries in the internal record set used in this study, we implemented a reproducible, multi-metric evaluation framework that measures both lexical and semantic similar- ity between each recordās abstract and its associated tldr. This framework ensures that summaries are faithful, 59 relevant, and information-preserving with respect to the source text. The validation pipeline applies sentence-level tokenization to both the abstract and summary before com- puting two families of metrics: ā¢ROUGE (ROUGE-1, ROUGE-2, ROUGE-L, ROUGE-Lsum): Measures lexical overlap based on unigrams, bigrams, and longest common subsequences. ā¢BERTScore (Precision, Recall, F1): Measures semantic similarity using contextualized embeddings from pre-trained transformer models. These metrics are calculated using the Hugging Face evaluate library when available, with a fallback to native implementations for offline or high-throughput use. All scores are scaled to the range [0.0, 100.0]. Records with empty or invalid summaries or abstracts are assigned zero scores across all metrics. Tables23and24summarize the validation rules and provide output examples, respectively. Table 23: Summary Evaluation Rules and Diagnostic Flags Validation RuleDescriptionFlag (on failure) Non-empty Abstract and TLDR Both abstract and tldr must be non-empty strings. empty_abstract, empty_tldr Minimum ROUGE-Lsum Threshold ROUGE-Lsum must exceed a configurable minimum threshold (default: 10.0). low_rouge_lsum Minimum BERTScore-F1 Threshold BERTScore_F1 must exceed a config- urable minimum threshold (default: 30.0). low_bertscore_f1 Optional Composite Score Composite score may be defined as the av- erage of ROUGE-Lsum and BERTScore- F1. Records with low composite score may be flagged. low_summary_quality Evaluation FallbackIf online evaluation tools are unavailable, fall back to native libraries without com- promising score fidelity. Not a flag ā automatically handled by backend logic. Evaluation Output Fields Evaluation must produce all 7 scores: ROUGE-1, ROUGE-2, ROUGE-L, ROUGE-Lsum, BERTScore Precision, Recall, and F1. missing_score_<metric> 60 Table 24: Example summary validation outputs: failing (left) and passing (right). Failing RecordPassing Record "summary_validation": "scores": "ROUGE-1": 4.12, "ROUGE-2": 1.34, "ROUGE-L": 3.56, "ROUGE-Lsum": 5.89, "BERTScore_Precision": 22.74, "BERTScore_Recall": 24.65, "BERTScore_F1": 23.44 , "flags": "low_rouge_lsum": true, "low_bertscore_f1": true , "summary_score": 14.66, "evaluated_at": "2025-07-30T13:22:18.543672" "summary_validation": "scores": "ROUGE-1": 48.21, "ROUGE-2": 31.04, "ROUGE-L": 43.89, "ROUGE-Lsum": 46.75, "BERTScore_Precision": 88.93, "BERTScore_Recall": 90.14, "BERTScore_F1": 89.52 , "flags": , "summary_score": 68.13, "evaluated_at": "2025-07-30T13:25:14.781129" 7.17 Identifier and Consistency Validation To ensure that each record is internally consistent and correctly linked to external bibliographic resources, we implemented an automated identifier validation pipeline. This validator performs three core consistency checks: (1) alignment of internal corpus identifiers; (2) normalization and comparison of Digital Object Identifiers (DOIs) across multiple metadata sources; and (3) alignment between text content and vector embeddings. Corpus identifiers are verified by comparing the filename_id (from ingest), the top-level corpus_id, and the nested metadata.corpusid. All three must agree for the record to pass this check. DOI consistency is assessed by collecting DOI strings from three sourcesāmetadata.externalids.DOI, unpaywall.doi, and openalex.doiā normalizing them to canonical form, and checking for agreement. Finally, paragraphāembedding alignment ensures that each paragraph has a corresponding embedding, and vice versa, with up to five unmatched IDs reported per type. Tables25and26summarize the validation rules and provide output examples, respectively. Table 25: Consistency Validation Rules and Diagnostic Flags Validation RuleDescriptionFlag (on failure) ID Agreement CheckThe filename_id, corpus_id, and metadata.corpusid must all be present and equal. corpusid_mismatch ā Raised if values disagree. Missing Metadata IDmetadata.corpusid must be defined if other IDs are present. missing_metadata_id ā Raised if missing. DOI NormalizationDOIs from metadata, unpaywall, and openalex are normalized to a canonical form. Not a flag ā Used internally for com- parison. DOI Agreement Check DOIs from all available sources must match after normalization. doi_mismatch ā Raised if two or more DOIs differ. DOI Coverage CheckAt least one DOI must be available across all sources. doi_missing_sources ā Raised if all DOI sources are missing. Missing EmbeddingsAll paragraph IDs must have corre- sponding embedding IDs. missing_embeddings ā Raised if any paragraph is unrepresented. Continued on next page 61 Validation RuleDescriptionFlag (on failure) Orphan EmbeddingsEmbedding IDs must correspond to a known paragraph ID. orphan_embeddings ā Raised if extra embeddings are found. Status Assignmentāfailā if corpus ID mismatch exists; āwarnā for other issues; āpassā if all checks succeed. Not a flag ā Controlled by overall val- idation logic. Table 26: Example consistency validation outputs: failing (left), warning (center), and passing (right). Failing RecordWarning RecordPassing Record "consistency_validation": "status": "fail", "flags": "missing_metadata_id": false, "corpusid_mismatch": true, "doi_mismatch": false, "missing_embeddings": false, "orphan_embeddings": false, "doi_missing_sources": false , "metrics": "corpus_ids": "filename_id": 123456, "corpus_id": 123456, "metadata_id": 789012 , "doi_set": [], "para_count": 10, "embed_count": 10, "missing_para_ids": [], "extra_embed_ids": [], "doi_sources_present": "metadata_doi": false, "unpaywall_doi": false, "openalex_doi": false , "checked_at": "2025-07-30T19:08:15.015264" "consistency_validation": "status": "warn", "flags": "missing_metadata_id": false, "corpusid_mismatch": false, "doi_mismatch": true, "missing_embeddings": true, "orphan_embeddings": false, "doi_missing_sources": false , "metrics": "corpus_ids": "filename_id": 101010, "corpus_id": 101010, "metadata_id": 101010 , "doi_set": [ "10.1000/example", "10.1000/other" ], "para_count": 5, "embed_count": 3, "missing_para_ids": [ "1P3", "1P4" ], "extra_embed_ids": [], "doi_sources_present": "metadata_doi": true, "unpaywall_doi": true, "openalex_doi": false , "checked_at": "2025-07-30T19:35:42.554110" "consistency_validation": "status": "pass", "flags": "missing_metadata_id": false, "corpusid_mismatch": false, "doi_mismatch": false, "missing_embeddings": false, "orphan_embeddings": false, "doi_missing_sources": false , "metrics": "corpus_ids": "filename_id": 202020, "corpus_id": 202020, "metadata_id": 202020 , "doi_set": [ "10.1000/example" ], "para_count": 8, "embed_count": 8, "missing_para_ids": [], "extra_embed_ids": [], "doi_sources_present": "metadata_doi": true, "unpaywall_doi": true, "openalex_doi": true , "checked_at": "2025-07-30T20:05:08.123456" 62