Paper deep dive
MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
Weiyue Li, Ruizhi Qian, Yi Li, Yongce Li, Yunfan Long, Jiahui Cai, Yan Luo, Mengyu Wang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 3:38:23 AM
Summary
MedConclusion is a large-scale benchmark dataset containing 5.7 million PubMed structured abstracts designed to evaluate the ability of Large Language Models (LLMs) to perform evidence-to-conclusion reasoning. The dataset includes journal-level metadata such as biomedical categories and SJR scores, enabling subgroup analysis. The authors evaluate various LLMs using both reference-based metrics and LLM-as-a-judge, finding that conclusion generation is behaviorally distinct from summary writing and that evaluation scores are sensitive to judge identity.
Entities (5)
Relation Signals (3)
MedConclusion → contains → 5.7M PubMed structured abstracts
confidence 100% · We introduce MedConclusion, a large-scale dataset of 5.7M PubMed structured abstracts
MedConclusion → includes → SJR
confidence 100% · MedConclusion also includes journal-level metadata such as biomedical category and SJR
LLM-as-a-judge → evaluates → MedConclusion
confidence 95% · We evaluate diverse LLMs... and score outputs with both reference-based metrics and LLM-as-a-judge.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured biomedical evidence remain limited. We introduce $\textbf{MedConclusion}$, a large-scale dataset of $\textbf{5.7M}$ PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. As an initial study, we evaluate diverse LLMs under conclusion and summary prompting settings and score outputs with both reference-based metrics and LLM-as-a-judge. We find that conclusion writing is behaviorally distinct from summary writing, strong models remain closely clustered under current automatic metrics, and judge identity can substantially shift absolute scores. MedConclusion provides a reusable data resource for studying scientific evidence-to-conclusion reasoning. Our code and data are available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.06505v1
- Canonical: https://arxiv.org/abs/2604.06505v1
Trouble viewing inline? Open PDF directly →
Full Text
100,870 characters extracted from source content.
Expand or collapse full text
Preprint. Under review. MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts Weiyue Li ∗ Ruizhi Qian ∗, Yi Li ∗, Yongce LiYunfan LongJiahui Cai Yan Luo † Mengyu Wang †, Harvard AI and Robotics Lab, Harvard Medical School University of Southern California,Carnegie Mellon University,Stanford University Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University Abstract Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured biomedical evidence remain limited. We in- troduce MedConclusion, a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to- conclusion reasoning. MedConclusion also includes journal-level meta- data such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. As an initial study, we evaluate diverse LLMs under conclusion and summary prompting settings and score out- puts with both reference-based metrics and LLM-as-a-judge. We find that conclusion writing is behaviorally distinct from summary writing, strong models remain closely clustered under current automatic metrics, and judge identity can substantially shift absolute scores. MedConclu- sion provides a reusable data resource for studying scientific evidence- to-conclusion reasoning. Our code and data are available at:https: //github.com/Harvard-AI-and-Robotics-Lab/MedConclusion. 1 Introduction Large language models (LLMs) have shown strong reasoning capability across a wide range of demanding settings, including mathematical thinking, long-form creative writing, and scientific discovery assistance (Luo et al., 2025; Zheng et al., 2025; Li et al., 2026a; Zhang et al., 2025; Yu et al., 2025; Fein et al., 2026). As these capabilities improve, there is growing interest in using LLMs to support research workflows, not only to retrieve or summarize papers, but also to infer scientific conclusions from evidence. Structured abstracts provide a particularly convenient setting for studying this capability as we can frame this evidence-to-conclusion reasoning problem as: given Background, Methods, and Results, the model should infer a Conclusion without injecting unprovided context. However, finding large-scale data sources from diverse scientific domains is challenging, which motivates us to shift the focus to biomedicine, where structured abstracts are widely used. Existing work has explored this direction, but current resources remain limited in two ways. First, existing datasets are often narrow in scope, focusing on specific study types or spe- cialized report formats, such as randomized controlled trial abstracts or echocardiography notes, rather than broad biomedical literature (Shieh et al., 2019; Tang et al., 2022). Other work uses conclusion reconstruction mainly as a proxy for premise–conclusion alignment or as a training objective, rather than as a reusable data resource (Gao et al., 2024; Bastan et al., 2022). These resources also typically do not emphasize journal-level metadata such as biomedical category and SJR, which limits analysis of how difficulty varies across subfields ∗ Contributed equally as co-first authors. † Contributed equally as co-senior authors. 1 arXiv:2604.06505v1 [cs.CL] 7 Apr 2026 Preprint. Under review. DatasetStudy Design Name Data Size Broad Biomed. Struc. Abs. Gold References Journal Metadata Summary Contrast Judge Robust. Gao et al. (2024)17.4K✓✗ Shieh et al. (2019)195.7K✗✓✗ Tang et al. (2022)57.1K✗✓✗ Tang et al. (2023)200.2K✗✓✗ Bastan et al. (2022)633K✓✗✓✗ MedConclusion5.7M✓ Table 1: Comparison of MedConclusion with conclusion-centric prior work. Checks denote properties that are central and explicitly emphasized by each resource or study. For multi- part resources, total dataset size sums all released components used in the paper. or venue strata. Second, existing benchmarking designs do not fully isolate the reasoning problem of conclusion generation. Some adjacent biomedical resources focus on question answering, medical exam reasoning, treatment-effect inference, or claim verification rather than deriving the author-written conclusion itself (Jin et al., 2019; Tsatsaronis et al., 2015; Jin et al., 2021; Pal et al., 2022; Nye et al., 2018; Lehman et al., 2019; DeYoung et al., 2020; Wadden et al., 2020). Moreover, open-ended conclusion generation is difficult to evaluate reliably because reference-based metrics are incomplete, and LLM judges can vary substantially in calibration (Maynez et al., 2020; Zheng et al., 2023; Liu et al., 2023; Shi et al., 2025; Huang et al., 2025). To address these limitations, we present MedConclusion, a large-scale 5.7M dataset for biomedical conclusion generation from PubMed structured abstracts. Each instance pairs the non-conclusion sections of an abstract with its original author-written conclusion, yield- ing naturally occurring supervision for evidence-to-conclusion reasoning. In addition, MedConclusion includes journal-level metadata, including biomedical category labels and SJR records. This combination of large-scale author-written supervision, broad biomedi- cal coverage, and journal metadata enables analyses that are difficult to conduct in prior conclusion-generation settings. We further provide an initial empirical study using diverse LLMs, contrasting conclusion prompting against summary prompting, evaluating with a hybrid rule-based reference metrics and LLM judges, and examining robustness across judge backbones. In summary, our work makes three contributions. First, we curate MedConclusion, a 5.7M-example dataset of PubMed structured abstracts for biomedical conclusion generation. Second, we augment the dataset with journal-level metadata, enabling aggregate and subgroup analysis across biomedical domains and venue strata. Third, we provide a first empirical study of the dataset by evaluating diverse LLMs, contrasting conclusion versus summary prompting, and studying the sensitivity of automatic evaluation to judge identity. We hope MedConclusion serves as a reusable data resource for future study of scientific evidence-to-conclusion reasoning. 2 Related work Adjacent biomedical reasoning resources.A large body of biomedical NLP work studies reasoning over scientific and clinical text, but not specifically the task of inferring an abstract conclusion from preceding evidence. PubMedQA and BioASQ evaluate biomedical question answering (Jin et al., 2019; Tsatsaronis et al., 2015); MedQA and MedMCQA focus on exam-style medical reasoning (Jin et al., 2021; Pal et al., 2022); EBM-NLP and Evidence Inference study treatment-effect reasoning from structured evidence (Nye et al., 2018; Lehman et al., 2019; DeYoung et al., 2020); and SciFact evaluates scientific claim verification (Wadden et al., 2020). More recently, EvidenceBench studies sentence-level evidence extraction for biomedical hypotheses from full papers rather than conclusion 2 Preprint. Under review. generation from structured abstracts (Wang et al., 2025). Structured abstracts and discourse- aware scientific summarization have also been widely studied (Teufel & Moens, 2002; Cohan et al., 2018; Dernoncourt & Lee, 2017; Cachola et al., 2020; Yasunaga et al., 2019). MedConclusion differs from these resources by centering benchmarking the evidence-to- conclusion reasoning step itself. Conclusion-centric generation and reconstruction.The closest prior work to MedConclu- sion studies conclusion reconstruction directly. Gao et al. (2024) reconstruct conclusions of structured scientific abstracts to evaluate premise–conclusion alignment, treating conclusion generation mainly as an evaluation proxy. Shieh et al. (2019) study conclusion generation for randomized controlled trial abstracts, and Tang et al. (2022) focus on echocardiography notes. Tang et al. (2023) study factual consistency in conclusion-oriented clinical-study summarization, while Bastan et al. (2022) use large-scale PubMed conclusion generation as a training objective. In contrast, MedConclusion contributes a broader 5.7M-example PubMed resource with author-written targets and journal-level metadata, and uses it to study conclusion generation as a reasoning task rather than only as a proxy objective. Table 1 shows a detailed comparison. Evaluation of open-ended scientific reasoning.Evaluating conclusion generation is chal- lenging because open-ended outputs can differ in wording, scope, and detail while re- maining partially valid. Prior work in summarization has shown that lexical overlap and fluency metrics are incomplete proxies for factual correctness (Lin, 2004; Papineni et al., 2002; Reimers & Gurevych, 2019; Zhang et al., 2019; Maynez et al., 2020; Kry ́ sci ́ nski et al., 2020; Durmus et al., 2020; Scialom et al., 2021; Laban et al., 2022). LLM-as-a-judge pro- vides a scalable alternative (Zheng et al., 2023; Liu et al., 2023), but recent work documents sensitivity to judge identity and grading scale, verbosity bias, position bias, and broader generalization concerns (Dubois et al., 2024; Li et al., 2026b; Ye et al., 2024; Huang et al., 2025; Gu et al., 2024; Zhu et al., 2023). These findings motivate the evaluation protocol used in this paper, but they are secondary to our main goal of introducing MedConclusion as a large-scale dataset for biomedical conclusion generation. Figure 1: Overview of MedConclusion and the evaluation pipeline. Left: an example MedConclusion instance, including article metadata, subject categories, and a structured abstract, where the non-conclusion sections are used as model input and the author-written conclusion serves as the gold reference. Right: the non-conclusion abstract is paired with prompts and given to diverse LLM families to generate conclusions, which are then com- pared against the ground-truth conclusion using both rule-based reference metrics and multi-dimensional LLM-as-a-judge metrics. 3 Preprint. Under review. 3 Methodology 3.1 Data curation MedConclusion is constructed from PubMed articles with structured abstracts pub- lished between 2000 and 2025. We identify candidate papers using PubMed’s constraint hasstructuredabstract and collect the corresponding records for downstream processing. 3.1.1 Data collection pipeline Our data collection pipeline is implemented with Entrez Direct (EDirect) 1 and a custom XML parser. We first query PubMed for all UIDs satisfyinghasstructuredabstractwithin the target time span, and deduplicate the retrieved identifiers. We then batch-download the corresponding PubMed XML records via theepostandefetchcommands provided by EDirect, and parse each record into a JSONL representation containing article metadata, keywords, and structured abstract segments represented as (label,nlmcategory,text) tuples. Records with missing abstract labels are filtered out during parsing. After parsing, we perform record-level deduplication using PMID, DOI, and normalized title. We then apply a rule-based cleaning procedure that keeps only English-language records with non-empty core bibliographic fields, normalizes date fields and missing metadata, and retains only articles with at least three abstract segments and at least one conclusion section. Conclusion sections are identified by matching normalized labels against a curated set of conclusion variants (Appendix G). We further remove records with malformed labels and clean keywords by dropping empty, overlong, or non-ASCII entries. The resulting cleaned corpus contains 5,692,839 structured abstract records with at least one conclusion section. 3.1.2 Construction of MedConclusion and dataset statistics The structured abstract records in MedConclusion span a total of 3,772 unique journals. For each journal, we retrieved its subject category assignments and annual SJR scores from the SCImago Journal & Country Rank (SJR) database, 2 a publicly available bibliometric resource derived from Scopus. Across the full corpus, the 3,772 journals are distributed across 141 subject categories. For SJR scores, we collected annual values from each journal’s first indexed year in the SJR database through 2024. Dataset statistics and a formal definition of the SJR score are provided in Appendix F. 3.2 Task, evaluation, and experimental setup Given a structured abstract with its Conclusion removed, letxbe the concatenation of all remaining sections and let y ⋆ be the original conclusion. A model generates ˆ y = f θ (x). We evaluate four prompting modes. A (default) asks the model to write a formal academic conclusion, with no explicit length or style constraints.Basks the model to write a formal academic summary, with no explicit length or style constraints.Casks the model to write a formal academic conclusion, with explicit sentence- and word-count targets and instructions to match the abstract’s writing style.Dasks the model to write a formal academic summary, with the same sentence- and word-count targets and the same instruction to match the abstract’s writing style. Exact prompts are given in Appendix B. We score outputs with two classes of automatic metrics. First, we use multi-dimensional LLM-as-a-judge scoring. Given(y ⋆ , ˆ y), the judge outputs five scores in[0, 100]: semantic similarity, writing style similarity, non-contradiction, numeric consistency, and formality similarity. Second, we report lightweight diagnostics and reference-based metrics: word- count ratio, sentence-count ratio, embedding cosine similarity (Reimers & Gurevych, 2019), 1 Kans J. Entrez® Direct: E-utilities on the Unix Command Line. 2013 Apr 23 [Updated 2025 Mar 25]. In: Entrez® Programming Utilities Help [Internet]. Bethesda (MD): National Center for Biotechnology Information (US); 2010-. Available from: https://w.ncbi.nlm.nih.gov/books/NBK179288/ 2 https://w.scimagojr.com 4 Preprint. Under review. ROUGE-1/2/L (Lin, 2004), BLEU (Papineni et al., 2002), and perplexity on the original and generated conclusions under a fixed external language model (GPT-2). Details for reference-based metrics could be found in Appendix D. We evaluate a diverse set of LLMs spanning closed-source frontier models, open-source instruction-tuned models, multimodal models, reasoning-oriented models, and small mod- els. For each model and prompting mode, we generate one output per instance using the corresponding prompt template. All modes enforce a no-new-claims instruction, and we record length-ratio diagnostics to quantify format compliance. Due to cost constraints, we evaluate a randomly sampled 30K subset. Detailed model configurations are in Appendix H. To study evaluation sensitivity, we useGPT-5.4-mini(OpenAI, 2026) as the primary judge andGemini 3 Flash(Google DeepMind, 2025) as a secondary judge (prompts in Ap- pendix C). Our two main comparisons are conclusion prompting versus summary prompt- ing and judge-backbone robustness. The first tests whether models treat conclusion writing as a distinct discourse function (Teufel & Moens, 2002; Cohan et al., 2018); the second mea- sures how much absolute scores and model rankings depend on judge identity (Zheng et al., 2023; Dubois et al., 2024; Shi et al., 2025; Ye et al., 2024; Huang et al., 2025; Gu et al., 2024). Figure 1 shows our overall pipeline, and Appendix A shows an example data point. Model Semantic Sim. ↑ Writing Style Sim. ↑ Non-Contradiction Rate↑ Numeric Consistency↑ Formality Sim. ↑ General-purpose Models GPT-5.473.2271.2184.6188.2489.80 Gemini 3.1 Pro71.87 70.1382.0286.9289.49 Gemini 3 Flash71.3369.8781.7686.4589.17 Gemma-3-27B71.0369.1881.5584.1389.36 DeepSeek-V3.269.4768.2180.3186.2288.59 Llama-3.1-8B70.5366.6980.2479.8288.03 MiniMax-M2.171.2166.9581.8973.6588.83 Gemma-2-9B69.3167.4279.1275.0588.41 Qwen3-4B69.8066.3578.9671.7888.47 Qwen2.5-7B66.8765.7477.5077.3186.60 Llama-3.2-1B54.1750.6966.1482.6978.35 Reasoning Models Kimi-K269.7966.3680.9261.6288.62 DeepSeek-R168.9348.0679.6775.5875.91 Vision-Language Models GLM-4.6V70.8668.8380.5080.1988.87 Qwen2.5-VL-7B68.9664.7478.7371.8287.34 Table 2: LLM-as-Judge evaluation scores for conclusion generation. Models are grouped by primary capability. Bold andunderlinedenote the best and second-best scores, respectively. 4 Results 4.1 Overall performance under conclusion generation Table 2 shows thatGPT-5.4is the strongest model under the primary judge, leading all five judge dimensions. At the same time, models such asGemini 3.1 Pro,Gemini 3 Flash, DeepSeek-V3.2,Gemma-3-27B, andGLM-4.6Vlie within only a few points of the top model on most judge dimensions. This score compression suggests that current reference-comparison evaluation separates strong models only weakly, even though the task is not trivial. Table 3 tells a partly different story.DeepSeek-V3.2attains the best ROUGE-1/2/L and ties for the best BLEU, whileGPT-5.4remains the strongest model under judge-based semantic and non-contradiction scores. Likewise,Gemma-2-9Bhas the highest embedding similarity 5 Preprint. Under review. Model WC Ratio↓ SC Ratio↓ Embed. Sim. ↑ ROUGE- 1↑ ROUGE- 2↑ ROUGE- L↑BLEU↑ PPL Orig. ↓ PPL Gen. ↓ General-purpose Models GPT-5.42.191.710.770.340.100.210.0470.2640.39 Gemini 3.1 Pro2.281.750.730.330.100.210.0470.2634.54 Gemini 3 Flash2.121.690.740.340.100.210.0470.2633.87 Gemma-3-27B2.492.010.770.320.090.200.0470.2630.46 DeepSeek-V3.21.731.38 0.760.350.110.230.0570.2647.14 Llama-3.1-8B2.672.060.740.320.100.200.0470.2621.47 MiniMax-M2.13.112.390.760.300.090.190.0370.2630.69 Gemma-2-9B2.382.120.780.330.100.210.0470.2630.06 Qwen3-4B3.052.200.750.300.090.190.0370.2625.92 Qwen2.5-7B1.78 1.590.750.340.110.220.0570.2635.26 Llama-3.2-1B1.821.230.720.310.090.200.0470.2629.88 Reasoning Models Kimi-K22.902.790.750.300.090.180.0370.2660.76 DeepSeek-R19.4511.170.400.150.050.100.0170.2636.67 Vision-Language Models GLM-4.6V2.231.990.760.340.110.220.0570.2630.72 Qwen2.5-VL-7B2.892.490.750.310.100.200.0470.2623.52 Table 3: Rule-based evaluation scores for conclusion generation, same order as Table 2. despite clearly lower judge scores than the best closed models. Perplexity-based fluency diagnostics also decouple from task quality: several mid-sized open or multimodal models have substantially lower generated-text perplexity thanGPT-5.4, yet they do not approach its semantic or contradiction scores. These mismatches indicate that lexical overlap, embedding similarity, fluency, and judge agreement capture different aspects of biomedical conclusion generation, and indicate our hybrid evaluation approach provides a more comprehensive evaluation than traditional reference-based metrics. Mode Model Semantic Sim. ↑ Writing Style Sim. ↑ Non-Contradiction Rate↑ Numeric Consistency↑ Formality Sim. ↑ Prompt without Formatting Restriction A GPT-5.473.2271.2084.6188.2489.80 Gemini 3 Flash71.3369.8781.7686.4589.17 B GPT-5.472.1162.6083.9666.2488.47 Gemini 3 Flash71.0461.5582.4458.0988.11 Prompt with Length and Writing Style Restriction C GPT-5.470.9069.0782.1791.3687.54 Gemini 3 Flash68.5967.2478.4791.8286.51 D GPT-5.464.9960.1378.7674.0685.35 Gemini 3 Flash64.1662.2975.5177.6685.28 Table 4: LLM-as-Judge evaluation scores across prompt settings and generation modes. A andCare generating conclusions, whereasBandDare generating summaries. 4.2 Conclusion generation is not summary writing Table 4 shows a clear discourse-function effect. Relative to A , C slightly reduces semantic and style similarity for bothGPT-5.4andGemini 3 Flash, but it improves numeric consis- tency to about 91 for both models. This suggests that explicit style and length control helps models better match the reference conclusion’s level of numeric selectivity, even when it slightly constrains broader content choice. 6 Preprint. Under review. When the target is changed from a conclusion to a summary, performance shifts more substantially. UnderD, semantic similarity drops by about 7–8 points for both models relative toA, with even larger drops in writing style similarity and numeric consistency. B reveals an even more interesting pattern: semantic similarity rebounds to nearly the A level (within 1.11 points forGPT-5.4and 0.29 points forGemini 3 Flash), but writing style similarity remains more than 8 points lower and numeric consistency collapses by 22.00 and 28.36 points, respectively. This suggests that unconstrained summaries often preserve the broad meaning of the abstract while selecting different details, especially numbers, scope qualifiers, and level of detail, than the published conclusion. We therefore interpret these results as evidence that conclusion writing is behaviorally distinct from summary writing in this benchmark setting. Because the judge only compares outputs to the gold conclusion, the effect should be read as a difference in reference agreement and discourse targeting, not as proof that summary-mode outputs are unsupported by the input. Some of the additional detail produced in summary modes may still be compatible with the source abstract even when it lowers the agreement with the reference conclu- sion. We further examine whether this distinction holds across biomedical subfields in Appendix E.11. Model Semantic Sim. ↑ Writing Style Sim. ↑ Non-Contradiction Rate↑ Numeric Consistency↑ Formality Sim. ↑ Judge: GPT-5.4-mini GPT-5.473.2271.2084.6188.2489.80 Gemini 3.1 Pro71.8770.1382.0286.9289.49 Gemini 3 Flash71.3369.8781.7686.4589.17 Judge: Gemini 3 Flash GPT-5.484.3071.4997.5198.1892.50 Gemini 3.1 Pro82.6468.4196.5897.5390.70 Gemini 3 Flash82.5970.0496.6297.2891.46 Table 5: LLM-as-Judge evaluation scores: GPT-5.4-mini vs Gemini 3 Flash as judge. 4.3 Judge robustness: score scale shifts across judges Table 5 shows large absolute calibration shifts when the judge backbone is changed. Switching fromGPT-5.4-minitoGemini 3 Flashas judge raises semantic similarity, non- contradiction, and numeric consistency across the same three generation models. By contrast, writing style similarity changes only modestly. Thus, the absolute score scale is highly judge- dependent. However, ranking is relatively stable.GPT-5.4remains the top generator under both judges across all five dimensions, yet the middle ordering can flip occasionally. 5 Analysis We further analyze benchmark behavior across journal-level subgroups to understand where conclusion generation is relatively easier or harder. In particular, we study variation along two axes: journal prestige, measured by SJR score, and biomedical category. These analyses are descriptive rather than part of the benchmark definition itself, and are intended to reveal whether performance differences are associated with venue prestige or with the broader structure of biomedical subfields. All results in this section useGPT-5.4under settingAas the representative condition unless otherwise noted. 5.1 SJR score Figure 2 plots journal-level SJR scores against each evaluation metric forGPT-5.4under settingA. Most reference-based and judge-based metrics show small but statistically signifi- cant positive associations with SJR: ROUGE-1, ROUGE-2, Semantic Similarity, Writing Style 7 Preprint. Under review. 0.00.51.01.52.0 SJR 0.25 0.30 0.35 0.40 0.45 ROUGE-1 r = +0.067*** = +0.071 0.00.51.01.52.0 SJR 0.025 0.050 0.075 0.100 0.125 0.150 0.175 ROUGE-2 r = +0.065*** = +0.072 0.00.51.01.52.0 SJR 0.125 0.150 0.175 0.200 0.225 0.250 0.275 0.300 ROUGE-L r = +0.040* = +0.043 0.00.51.01.52.0 SJR 0.00 0.02 0.04 0.06 0.08 BLEU r = +0.063** = +0.090 0.00.51.01.52.0 SJR 10 20 30 40 50 60 70 Perplexity r = -0.035 = -0.022 0.00.51.01.52.0 SJR 55 60 65 70 75 80 85 90 Semantic Similarity r = +0.068*** = +0.048 0.00.51.01.52.0 SJR 60 65 70 75 80 Writing Style Sim. r = +0.098*** = +0.104 0.00.51.01.52.0 SJR 70 75 80 85 90 95 100 Non-Contradiction Rate r = -0.016 = -0.058 0.00.51.01.52.0 SJR 65 70 75 80 85 90 95 100 Numeric Consistency r = -0.052** = -0.093 0.00.51.01.52.0 SJR 87 88 89 90 91 92 93 Formality Similarity r = +0.086*** = +0.079 SJR vs Evaluation Metrics Figure 2: Scatter plots of journal-level SJR scores versus evaluation metrics forGPT-5.4 under the A setting (outliers removed). Each point represents one journal, aggregated by mean score. The top row shows reference-based metrics and the bottom row shows LLM-judge dimensions. Pearson (r) and Spearman (ρ) correlations are annotated in each panel (∗ p< 0.05,∗ p< 0.01,∗ p< 0.001). Similarity, and Formality Similarity are all significant atp<0.001. By contrast, Perplexity and Non-Contradiction Rate show no significant trend, while Numeric Consistency exhibits a small but significant negative correlation. These results suggest that journals with higher SJR scores tend to produce abstracts whose conclusions are slightly easier to match in terms of lexical overlap and writing style, yet this advantage does not extend to factual consistency dimensions. The overall effect sizes remain modest, indicating that venue prestige is a weak rather than dominant predictor of conclusion-generation difficulty. 5.2 Category Figure 3 compares the top-5 and bottom-5 biomedical categories under two ranking criteria: mean semantic similarity (panel a) and mean ROUGE-L (panel b), both forGPT-5.4under settingA, with all twelve metrics min-max normalized to [0, 1]. When categories are ranked by semantic similarity (panel a), the top-5 categories form large, uniformly filled polygons: high semantic similarity co-occurs with high writing style similarity, numeric consistency, non-contradiction rate, and lexical overlap metrics alike. In contrast, when categories are ranked by ROUGE-L (panel b), the top-5 polygons are visibly lopsided. ROUGE-1/2/L axes are high by construction, but judge-based dimensions such as writing style and numeric consistency do not follow. For example, Gerontology ranks among the top-5 by ROUGE-L yet falls well below the semantic-similarity top-5 on writing style and numeric consistency. This asymmetry reveals that lexical overlap with the reference conclusion is not a reliable proxy for overall conclusion quality. Categories whose generated conclusions share many surfacen-grams with the gold reference do not necessarily match its rhetorical style, numeric selectivity, or broader discourse structure. By contrast, high semantic similarity appears to act as a more holistic quality indicator that correlates with strong performance across both reference-based and judge-based dimensions. This observation reinforces the motivation for our hybrid evaluation protocol (Section 4.2): relying on ROUGE or BLEU alone would mask meaningful quality differences that only LLM-as-a-judge scoring can detect. The bottom-5 categories under both rankings share substantial overlap: Software, Computer Science Applications, and Applied Microbiology and Biotechnology appear in both lists, confirming that these interdisciplinary, non-clinical fields are consistently the hardest for conclusion generation regardless of the evaluation axis. Their radar profiles are highly irregular. For instance, Software has the lowest semantic similarity across all 112 categories (61.0) yet among the highest numeric consistency (96.4), illustrating that no single metric 8 Preprint. Under review. Semantic Sim. Writing Style Non-Contrad. Rate Numeric Cons. Formality Sim. Embed. CosSim ROUGE-1 ROUGE-2 ROUGE-L BLEU WC Ratio SC Ratio 0.2 0.4 0.6 0.8 1.0 Top 5 (Highest Semantic Sim.) Semantic Sim. Writing Style Non-Contrad. Rate Numeric Cons. Formality Sim. Embed. CosSim ROUGE-1 ROUGE-2 ROUGE-L BLEU WC Ratio SC Ratio 0.2 0.4 0.6 0.8 1.0 Bottom 5 (Lowest Semantic Sim.) Top 5 / Bottom 5 Experimental and Cognitive Psychology Endocrine and Autonomic Systems Advanced and Specialized Nursing Environmental Science Emergency Nursing Pollution Health, Toxicology and Mutagenesis Computer Science Applications Applied Microbiology and Biotechnology Software (a) Top 5 and bottom 5 biomedical categories ranked by mean Semantic Similarity. Semantic Sim. Writing Style Non-Contrad. Rate Numeric Cons. Formality Sim. Embed. CosSim ROUGE-1 ROUGE-2 ROUGE-L BLEU WC Ratio SC Ratio 0.2 0.4 0.6 0.8 1.0 Top 5 (Highest ROUGE-L) Semantic Sim. Writing Style Non-Contrad. Rate Numeric Cons. Formality Sim. Embed. CosSim ROUGE-1 ROUGE-2 ROUGE-L BLEU WC Ratio SC Ratio 0.2 0.4 0.6 0.8 1.0 Bottom 5 (Lowest ROUGE-L) Top 5 / Bottom 5 Advanced and Specialized Nursing Environmental Science Gerontology Histology Rheumatology Computer Science Applications Applied Microbiology and Biotechnology Linguistics and Language Software Plant Science (b) Top 5 and bottom 5 biomedical categories ranked by mean ROUGE-L. Figure 3: Radar chart comparison of the top 5 and bottom 5 biomedical categories ranked by both mean Semantic Similarity and mean ROUGE-L forGPT-5.4under setting A . Each axis is normalized to [0, 1]. suffices to characterize conclusion-generation difficulty. Appendix E provides representative generation examples from both high- and low-performing categories. 6 Conclusion We introduce MedConclusion, a large-scale benchmark of 5.7M structured PubMed abstracts for biomedical conclusion generation, paired with author-written conclusions and journal- level metadata. Our experiments show that conclusion generation is behaviorally distinct from summary writing, that strong LLMs remain closely clustered under current automatic metrics, and that absolute LLM-judge scores are sensitive to judge identity. Analyses across journal prestige and biomedical categories further show that task difficulty is heterogeneous and that lexical-overlap metrics alone do not adequately capture conclusion quality. 9 Preprint. Under review. References Mohaddeseh Bastan, Nishant Shankar, Mihai Surdeanu, and Niranjan Balasubramanian. SuMe: A dataset towards summarizing biomedical mechanisms. In Nicoletta Calzolari, Fr ́ ed ́ eric B ́ echet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, H ́ el ` ene Mazo, Jan Odijk, and Stelios Piperidis (eds.), Proceedings of the Thirteenth Language Resources and Evaluation Conference, p. 6922–6931, Marseille, France, June 2022. European Language Resources Association. URL https://aclanthology.org/2022.lrec-1.748/. Isabel Cachola, Kyle Lo, Arman Cohan, and Daniel S Weld. Tldr: Extreme summarization of scientific documents. In Findings of the Association for Computational Linguistics: EMNLP 2020, p. 4766–4777, 2020. Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. A discourse-aware attention model for abstractive summa- rization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), p. 615–621, 2018. DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025a. URL https://arxiv.org/abs/2501.12948. DeepSeek-AI. Deepseek-v3.2: Pushing the frontier of open large language models, 2025b. Franck Dernoncourt and Ji-Young Lee. Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 2: Short Papers), p. 308–313, 2017. Jay DeYoung, Eric Lehman, Benjamin Nye, Iain Marshall, and Byron C Wallace. Evidence inference 2.0: More data, better models. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, p. 123–132, 2020. Yann Dubois, Bal ́ azs Galambosi, Percy Liang, and Tatsunori B Hashimoto. Length-controlled alpacaeval: A simple way to debias automatic evaluators. arXiv preprint arXiv:2404.04475, 2024. Esin Durmus, He He, and Mona Diab. Feqa: A question answering evaluation framework for faithfulness assessment in abstractive summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics, p. 5055–5070, 2020. Daniel Fein, Sebastian Russo, Violet Xiang, Kabir Jolly, Rafael Rafailov, and Nick Haber. Lit- bench: A benchmark and dataset for reliable evaluation of creative writing. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7740–7755, 2026. Yingqiang Gao, Nianlong Gu, Jessica Lam, James Henderson, and Richard Hahnloser. Evaluating unsupervised argument aligners via generation of conclusions of structured scientific abstracts. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), p. 151–160, 2024. Gemma Team. Gemma. 2024. doi: 10.34740/KAGGLE/M/3301. URLhttps://w.kaggle. com/m/3301. Gemma Team. Gemma 3. 2025. URL https://goo.gle/Gemma3Report. Borja Gonz ́ alez-Pereira, Vicente P Guerrero-Bote, and F ́ elix Moya-Aneg ́ on. A new approach to the metric of journals’ scientific prestige: The sjr indicator. Journal of informetrics, 4(3): 379–391, 2010. Google DeepMind. Gemini 3 flash: Model card, December 2025. URLhttps://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-Flash-Model-Card.pdf. Pub- lished December 2025; accessed 2026-03-31. 10 Preprint. Under review. Google DeepMind. Gemini 3.1 pro: Model card, February 2026. URLhttps://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf . Pub- lished February 2026; accessed 2026-03-31. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Ro- driguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bob- bie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzm ́ an, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, OnurC ̧elebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Gird- har, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Ra- parthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Syd- ney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, V ́ ıtor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Chang- han Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana 11 Preprint. Under review. Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Gro- shev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bon- trager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satter- field, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783. Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge. The Innovation, 2024. Hui Huang, Xingyuan Bu, Hongli Zhou, Yingqi Qu, Jing Liu, Muyun Yang, Bing Xu, and Tiejun Zhao. An empirical study of llm-as-a-judge for llm evaluation: Fine-tuned judge model is not a general substitute for gpt-4. In Findings of the Association for Computational Linguistics: ACL 2025, p. 5880–5895, 2025. Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021. Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 conference 12 Preprint. Under review. on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), p. 2567–2577, 2019. Kimi Team, Yifan Bai, Yiping Bao, Y. Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, Jialei Cui, Hao Ding, Mengnan Dong, Angang Du, Chen- zhuang Du, Dikang Du, Yulun Du, Yu Fan, Yichen Feng, Kelin Fu, Bofei Gao, Chenxiao Gao, Hongcheng Gao, Peizhong Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Qizheng Gu, Xinran Gu, Longyu Guan, Haiqing Guo, Jianhang Guo, Xiaoru Hao, Tianhong He, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Yangyang Hu, Zhenxing Hu, Weixiao Huang, Zhiqi Huang, Zihao Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yongsheng Kang, Guokun Lai, Cheng Li, Fang Li, Haoyang Li, Ming Li, Wentao Li, Yang Li, Yanhao Li, Yiwei Li, Zhaowei Li, Zheming Li, Hongzhan Lin, Xiaohan Lin, Zongyu Lin, Chengyin Liu, Chenyu Liu, Hongzhang Liu, Jingyuan Liu, Junqi Liu, Liang Liu, Shaowei Liu, T. Y. Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yibo Liu, Yiping Liu, Yue Liu, Zhengying Liu, Enzhe Lu, Haoyu Lu, Lijun Lu, Yashuo Luo, Shengling Ma, Xinyu Ma, Yingwei Ma, Shaoguang Mao, Jie Mei, Xin Men, Yibo Miao, Siyuan Pan, Yebo Peng, Ruoyu Qin, Zeyu Qin, Bowen Qu, Zeyu Shang, Lidong Shi, Shengyuan Shi, Feifan Song, Jianlin Su, Zhengyuan Su, Lin Sui, Xinjie Sun, Flood Sung, Yunpeng Tai, Heyi Tang, Jiawen Tao, Qifeng Teng, Chaoran Tian, Chensi Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Jianzhou Wang, Jiaxing Wang, Jinhong Wang, Shengjie Wang, Shuyi Wang, Si Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yiqin Wang, Yuxin Wang, Yuzhi Wang, Zhaoji Wang, Zhengtao Wang, Zhengtao Wang, Zhexu Wang, Chu Wei, Qianqian Wei, Haoning Wu, Wenhao Wu, Xingzhe Wu, Yuxin Wu, Chenjun Xiao, Jin Xie, Xiao- tong Xie, Weimin Xiong, Boyu Xu, Jinjing Xu, L. H. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ziyao Xu, Jing Xu, Jing Xu, Junjie Yan, Yuzi Yan, Hao Yang, Xiaofei Yang, Yi Yang, Ying Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Xingcheng Yao, Wenjie Ye, Zhuorui Ye, Bohong Yin, Longhui Yu, Enming Yuan, Hongbang Yuan, Mengjie Yuan, Siyu Yuan, Haobing Zhan, Dehao Zhang, Hao Zhang, Wanlu Zhang, Xiaobin Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Haotian Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Shaojie Zheng, Longguang Zhong, Jianren Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Zhen Zhu, Weiyu Zhuang, and Xinxing Zu. Kimi k2: Open agentic intelligence, 2026. URL https://arxiv.org/abs/2507.20534. Wojciech Kry ́ sci ́ nski, Bryan McCann, Caiming Xiong, and Richard Socher. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), p. 9332–9346, 2020. Philippe Laban, Tobias Schnabel, Paul N Bennett, and Marti A Hearst. Summac: Re- visiting nli-based models for inconsistency detection in summarization. Transactions of the Association for Computational Linguistics, 10:163–177, 2022. Eric Lehman, Jay DeYoung, Regina Barzilay, and Byron C Wallace. Inferring which medical treatments work from reports of clinical trials. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 3705–3717, 2019. Weiyue Li, Mingxiao Song, Zhenda Shen, Dachuan Zhao, Yunfan Long, Yi Li, Yongce Li, Ruyi Yang, and Mengyu Wang. Llm review: Enhancing creative writing via blind peer review feedback. arXiv preprint arXiv:2601.08003, 2026a. Weiyue Li, Minda Zhao, Weixuan Dong, Jiahui Cai, Yuze Wei, Michael Pocress, Yi Li, Wanyan Yuan, Xiaoyue Wang, Ruoyu Hou, et al. Grading scale impact on llm-as-a-judge: Human-llm alignment is highest on 0-5 grading scale. arXiv preprint arXiv:2601.03444, 2026b. Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summariza- tion branches out, p. 74–81, 2004. 13 Preprint. Under review. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G- eval: Nlg evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing, p. 2511–2522, 2023. Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. Llm4sr: A survey on large language models for scientific research. arXiv preprint arXiv:2501.04306, 2025. Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics, p. 1906–1919, 2020. Meta.Llama 3.2:Revolutionizing edge ai and vision with open,cus- tomizablemodels,September2024.URLhttps://ai.meta.com/blog/ llama-3-2-connect-2024-vision-edge-mobile-devices/.Meta AI blog post, pub- lished 2024-09-25; accessed 2026-03-31. MiniMax. Minimax m2.1: Significantly enhanced multi-language programming, built for real-world complex tasks, December 2025. URLhttps://w.minimax.io/news/ minimax-m21. MiniMax news post, published 2025-12-23; accessed 2026-03-31. Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain Marshall, Ani Nenkova, and Byron C Wallace. A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 197–207, 2018. OpenAI.Introducing gpt-5.4, March 2026.URLhttps://openai.com/index/ introducing-gpt-5-4/. OpenAI product announcement, accessed 2026-03-31. Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Gerardo Flores, George H Chen, Tom Pollard, Joyce C Ho, and Tristan Naumann (eds.), Proceedings of the Conference on Health, Inference, and Learning, volume 174 of Proceedings of Machine Learning Research, p. 248–260. PMLR, 07–08 Apr 2022. URL https://proceedings.mlr.press/v174/pal22a.html. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, p. 311–318, 2002. Qwen Team. Qwen2.5: A party of foundation models, September 2024. URLhttps: //qwenlm.github.io/blog/qwen2.5/. Qwen Team. Qwen2.5-vl, January 2025a. URL https://qwen.ai/blog?id=qwen2.5-vl. Qwen Team. Qwen3 technical report, 2025b. URL https://arxiv.org/abs/2505.09388. Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP- IJCNLP), p. 3982–3992, 2019. Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Sta- iano, Alex Wang, and Patrick Gallinari. Questeval: Summarization asks for fact-based evaluation. In Proceedings of the 2021 conference on empirical methods in natural language processing, p. 6594–6604, 2021. Lin Shi, Chiyu Ma, Wenhua Liang, Xingjian Diao, Weicheng Ma, and Soroush Vosoughi. Judging the judges: A systematic study of position bias in llm-as-a-judge. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, p. 292–314, 2025. 14 Preprint. Under review. Alexander Te-Wei Shieh, Yung-Sung Chuang, Shang-Yu Su, and Yun-Nung Chen. Towards understanding of medical randomized controlled trials by conclusion generation. In Proceedings of the Tenth International Workshop on Health Text Mining and Information Analysis (LOUHI 2019), p. 108–117, 2019. Liyan Tang, Shravan Kooragayalu, Yanshan Wang, Ying Ding, Greg Durrett, Justin F Rousseau, and Yifan Peng. Echogen: generating conclusions from echocardiogram notes. In Proceedings of the 21st Workshop on Biomedical Language Processing, p. 359–368, 2022. Xiangru Tang, Arman Cohan, and Mark Gerstein. Aligning factual consistency for clinical studies summarization through reinforcement learning. In Tristan Naumann, Asma Ben Abacha, Steven Bethard, Kirk Roberts, and Anna Rumshisky (eds.), Proceedings of the 5th Clinical Natural Language Processing Workshop, p. 48–58, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.clinicalnlp-1.7. URL https://aclanthology.org/2023.clinicalnlp-1.7/. V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, Jing Chen, Jinhao Chen, Jinghao Lin, Jinjiang Wang, Junjie Chen, Leqi Lei, Letian Gong, Leyi Pan, Mingdao Liu, Mingde Xu, Mingzhi Zhang, Qinkai Zheng, Sheng Yang, Shi Zhong, Shiyu Huang, Shuyuan Zhao, Siyan Xue, Shangqin Tu, Shengbiao Meng, Tianshu Zhang, Tianwei Luo, Tianxiang Hao, Tianyu Tong, Wenkai Li, Wei Jia, Xiao Liu, Xiaohan Zhang, Xin Lyu, Xinyue Fan, Xuancheng Huang, Yanling Wang, Yadong Xue, Yanfeng Wang, Yanzi Wang, Yifan An, Yifan Du, Yiming Shi, Yiheng Huang, Yilin Niu, Yuan Wang, Yuanchang Yue, Yuchen Li, Yutao Zhang, Yuting Wang, Yu Wang, Yuxuan Zhang, Zhao Xue, Zhenyu Hou, Zhengxiao Du, Zihan Wang, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Minlie Huang, Yuxiao Dong, and Jie Tang. Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025. URL https://arxiv.org/abs/2507.01006. Simone Teufel and Marc Moens. Summarizing scientific articles: experiments with relevance and rhetorical status. Computational linguistics, 28(4):409–445, 2002. George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. An overview of the bioasq large-scale biomedical semantic indexing and question answering competition. BMC bioinformatics, 16(1):138, 2015. David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 7534–7550, 2020. Jianyou Wang, Weili Cao, Kaicheng Wang, Xiaoyue Wang, Ashish Dalvi, Gino Prasad, Qishan Liang, Hsuan-lin Her, Ming Wang, Qin Yang, et al. Evidencebench: A benchmark for extracting evidence from biomedical papers. arXiv preprint arXiv:2504.18736, 2025. Michihiro Yasunaga, Jungo Kasai, Rui Zhang, Alexander R. Fabbri, Irene Li, Dan Friedman, and Dragomir R. Radev. Scisummnet: a large annotated corpus and content-impact models for scientific paper summarization with citation networks. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’19/IAAI’19/EAAI’19. AAAI Press, 2019. ISBN 978-1-57735- 809-1. doi: 10.1609/aaai.v33i01.33017386. URLhttps://doi.org/10.1609/aaai.v33i01. 33017386. Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al. Justice or prejudice? quantifying biases in llm-as-a-judge. arXiv preprint arXiv:2410.02736, 2024. 15 Preprint. Under review. Zhouliang Yu, Ruotian Peng, Keyi Ding, Yizhe Li, Zhongyuan Peng, Minghao Liu, Yifan Zhang, Zheng Yuan, Huajian Xin, Wenhao Huang, et al. Formalmath: Benchmarking formal mathematical reasoning of large language models. arXiv preprint arXiv:2505.02735, 2025. Jie Zhang, Cezara Petrui, Kristina Nikoli ́ c, and Florian Tram ` er. Realmath: A continuous benchmark for evaluating language models on research-level mathematics. arXiv preprint arXiv:2505.12575, 2025. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt- bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023. Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 17744–17761, 2025. Lianghui Zhu, Xinggang Wang, and Xinlong Wang. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631, 2023. 16 Preprint. Under review. A Example data Example Datapoint from MedConclusion Metadata: PMID: 21401313 Title: Inositol phosphoglycan P-type in infants of preeclamptic mothers. Journal Title: The journal of maternal-fetal & neonatal medicine : the official journal of the European Association of Perinatal Medicine, the Federation of Asia and Oceania Perinatal Societies, the International Society of Perinatal Obstetricians Journal ISO: J Matern Fetal Neonatal Med Volume: 25Issue: 2 DOI: 10.3109/14767058.2011.557789 Language: eng Publication Date: 2012-02 Article Date: 2011-03-14 SJR 2012: 0.656 Subject Area and Category: Medicine - Obstetrics and Gynecology - Pediatrics, Perinatology and Child Health Structured Abstract: BACKGROUND: Inositol phosphoglycan P-type (P-IPG) has consistently found to be elevated during active preeclampsia, although the biosynthetic source has to be identified yet. This multicenter prospective cross-sectional case-control study evaluated the fetus/newborn as the source of P-IPG. METHODS: A urine specimen was collected longitudinally for three consecutive days after delivery from 90 newborns and their mothers, and ordered according to clinical diagnosis of preeclampsia, gestational hypertension, or healthy pregnancy. RESULTS: The urinary excretion of P-IPG on day 0 was higher in the mothers in all groups (p<0.05) with higher levels in preeclamptic women (p<0.01) in the mothers compared to their newborns in the preeclamptic group (p<0.01). The difference persisted at least two days post partum. CONCLUSION: Findings of this study confirm the specificity of the increase in urinary excretion of P-IPG in preeclamptic mothers at day of birth compared to healthy pregnancy and GH, but does not extend to their newborns. Table 6: Example datapoint used for prompt construction and evaluation. In the structured abstract, the non-conclusion sections are highlighted separately from theCONCLUSIONsection to indicate that the former are used as model input, while the conclusion serves as the ground-truth reference for evaluation. 17 Preprint. Under review. B Prompts for conclusion/summary generation B.1 Prompts for conclusion generation (AandC) Conclusion Generation (Avs.C) System: You are a senior scientist in generating conclusions for scientific papers. User: You will be provided with a structured abstract of a scientific paper. The abstract contains sections such as Background, Objective, Methods, Results, etc., but the corresponding Conclusion section is missing. Your task is to infer and write the most plausible CONCLUSION section that would appear in this abstract. Here are the requirements: - Output ONLY the text for the conclusion itself. Do NOT include any section headers or explanations. - Only use the information provided by the abstract to derive your conclusion. - Do NOT introduce new experiments, datasets, numerical values, or claims that are not supported by the abstract. Please generate the conclusion in the same writing style as the given abstract sections. The conclusion should be a concise paragraph withsen_numsentences totalling word_num words. Use formal academic writing style. Structured Abstract: <Abstract> abstract_text </Abstract> Table 7: Prompts for conclusion generation. The highlighted variants distinguish the constrained writing setting (C), which enforces sentence and word count targets and asks the model to match the abstract’s writing style, from the unconstrained setting (A), which instead only requires a formal academic style without length constraints. 18 Preprint. Under review. B.2 Prompts for summary generation (BandD) Summary Generation ( B vs. D ) System: You are a senior scientist in generating summaries for scientific papers. User: You will be provided with a structured abstract of a scientific paper. The abstract contains sections such as Background, Objective, Methods, Results, etc. Your task is to summarize the core information in the given abstract sections. Here are the requirements: - Output ONLY the text for the summary itself. Do NOT include any section headers or explanations. - Only use the information provided by the abstract to derive your summary. - Do NOT introduce new experiments, datasets, numerical values, or claims that are not supported by the abstract. Please generate the summary in the same writing style as the given abstract sections. The summary should be a concise paragraph withsen_numsentences totalling word_num words. Use formal academic writing style. Structured Abstract: <Abstract> abstract_text </Abstract> Table 8: Prompts for summary generation. The highlighted variants distinguish the con- strained writing setting (D), which requires the summary to follow the abstract’s writing style and satisfy sentence and word count targets, from the unconstrained setting (B), which only specifies a formal academic style and leaves length unrestricted. 19 Preprint. Under review. C Prompts for LLM judges LLM Judge Prompt System: You are an expert evaluator of scientific writing. User: Your task is to compare the Generated Conclusion against the Original (Reference) Conclusion and score multiple dimensions from 0 to 100 (decimals allowed). Use ONLY the two conclusions provided. Do NOT provide any explanations. Scoring dimensions: - semantic similarity: How similar the meaning and core claims are. - writing style similarity: How similar the tone, phrasing, structure, and rhetorical style are. - contradiction rate: Degree of contradiction between the Generated Conclusion and the Original Conclusion. - 100 = no contradiction - 0 = severe contradiction - numeric consistency: Consistency of all numerical information, quantities, direc- tions, and magnitudes. - 100 = fully consistent or no numeric content in either text - 0 = major numeric inconsistency - formality similarity: Similarity in academic/formal writing level and register. INPUT: - Original Conclusion (reference): original_conclusion - Generated Conclusion: generated_conclusion OUTPUT FORMAT (STRICT): "semantic similarity": <0-100>, "writing style similarity": <0-100>, "contradiction rate": <0-100>, "numeric consistency": <0-100>, "formality similarity": <0-100> Table 9: Prompt for LLM judge 20 Preprint. Under review. D Reference-based metrics In addition to LLM-as-a-judge scoring, we report a set of lightweight diagnostics and reference-based metrics that compare the generated conclusion ˆ yagainst the author-written reference conclusiony ⋆ . These metrics are inexpensive to compute, easy to reproduce, and provide complementary signals about lexical overlap, semantic proximity, length control, and fluency. We do not treat any single metric as a complete measure of conclusion quality; rather, we use them as a bundle of auxiliary indicators. Let|y| word denote the number of words in a texty, and let|y| sent denote its number of sentences. Word-count ratio.To measure length matching at the tokenized word level, we compute WCR( ˆ y, y ⋆ ) = | ˆ y| word |y ⋆ | word .(1) A value close to 1 indicates that the generated conclusion has similar length to the refer- ence. Values below 1 indicate shorter generations, while values above 1 indicate longer generations. Sentence-count ratio.To assess structural length control at the sentence level, we compute SCR( ˆ y, y ⋆ ) = | ˆ y| sent |y ⋆ | sent . (2) Embedding cosine similarity.To capture semantic similarity beyond surface lexical over- lap, we encode ˆ y andy ⋆ using an off-the-shelf ALL-MPNET-BASE-V2 3 sentence embedding model and compute the cosine similarity between the two vector representations (Reimers & Gurevych, 2019): CosSim( ˆ y, y ⋆ ) = φ( ˆ y) ⊤ φ(y ⋆ ) ∥φ( ˆ y)∥φ(y ⋆ )∥ ,(3) whereφ(·)denotes the embedding function. Higher values indicate greater semantic proximity. ROUGE. We report ROUGE-1, ROUGE-2, and ROUGE-L (Lin, 2004). ROUGE-1 and ROUGE-2 measure unigram and bigram overlap, respectively, while ROUGE-L measures longest-common-subsequence overlap. These metrics quantify the extent to which the generated conclusion reuses words and short phrases appearing in the reference. Because multiple valid conclusions may use different wording, ROUGE should be interpreted as a lexical-overlap signal rather than a direct measure of scientific correctness. BLEU.We also report BLEU (Papineni et al., 2002), which measuresn-gram precision of the generated text with a brevity penalty. As with ROUGE, BLEU is sensitive to phrasing and therefore is best viewed as an approximate indicator of closeness to the reference wording, not as a standalone measure of conclusion quality. Perplexity under an external language model. To estimate fluency and distributional typicality, we compute perplexity using a fixed external language model,GPT-2, on both the reference conclusiony ⋆ and the generated conclusion ˆ y. For a sequence of tokens y = (w 1 , . . . , w T ), perplexity under language model p is PPL(y) = exp − 1 T T ∑ t=1 log p(w t | w <t ) ! . (4) Lower perplexity indicates that the text is more probable under the external language model. Reporting perplexity for bothy ⋆ and ˆ yhelps contextualize whether model-generated conclusions are comparably fluent to author-written ones under the same scoring model. 3 https://huggingface.co/sentence-transformers/all-mpnet-base-v2 21 Preprint. Under review. Implementation notes.All reference-based metrics are computed between each generated conclusion ˆ yand its paired author-written conclusiony ⋆ . We aggregate scores over the evaluation set using the arithmetic mean unless otherwise noted. Since these metrics emphasize different aspects of generation quality, we recommend interpreting them jointly with the LLM-as-a-judge results in the main paper. 22 Preprint. Under review. E Additional category analysis E.1 Example 1 Example: Experimental and Cognitive Psychology (HIGH) Metadata: PMID: 32422422 Journal Title: Cortex Publication Year: 2020 SJR 2020: 1.786 Subject Area and Category: - Experimental and Cognitive Psychology Structured Abstract: BACKGROUND: Numerous studies have shown visuoperceptual/visuospatial deficits in dementia with Lewy bodies (DLB) and Alzheimer ’s disease (AD). Visual texture recognition is also impaired in patients with DLB and AD. Although patients with DLB often exhibit visual misidentifications of objects, there are few studies on the relationships between visual texture recognition and viewpoints for object recognition. OBJECTIVES: The aim of this study was to clarify how viewpoints, textures, and visual cognitive functions affect object recognition and result in visual misidentifi- cations in patients with DLB or AD. METHODS: A total of 37 patients with probable DLB and 58 with probable AD and 32 age-matched healthy controls underwent neuropsychological and visuopercep- tual assessments, and performed object identification tasks under four conditions (non-canonical view + blurry texture, non-canonical view + clear texture, canonical view + blurry texture, canonical view + clear texture). The relationship between object identification and other visuoperceptual functions was analyzed. RESULTS: Patients with DLB and AD exhibited significantly impaired object recog- nition under non-canonical viewing with blurry texture conditions, with the DLB patients exhibiting a significantly worse performance than the AD patients. Patients with DLB and AD exhibited visual misidentifications during object identification tasks under non-canonical viewing. In patients with DLB, the number of visual misidentifications was significantly correlated with the scores of visual texture recognition. CONCLUSION: The present study showed that significantly impaired object recog- nition in patients with DLB under the influences by both viewpoint and visual texture and in those with AD under the influence by viewpoint. Visual misiden- tification in object recognition could be associated with impaired visual texture recognition in DLB. GENERATED: Object recognition is particularly vulnerable in DLB and, to a lesser extent, in AD when objects are presented from non-canonical viewpoints with degraded texture information. The greater impairment in DLB and the association between visual misidentifications and visual texture recognition in this group suggest that deficits in processing texture cues, especially under atypical viewing conditions, contribute to visual misidentification. These findings indicate that altered viewpoint and texture processing are important mechanisms underlying visual recognition disturbances in DLB. Table 10: Example from Experimental and Cognitive Psychology. ROUGE-1: 0.480, ROUGE-2: 0.130, ROUGE-L: 0.272, BLEU: 0.068, Perplexity: 41.1, Semantic Sim.: 88, Writing Style Sim.: 84, Non-Contradiction Rate: 96, Numeric Consistency: 100, Formality Sim.: 92. 23 Preprint. Under review. E.2 Example 2 Example: Endocrine and Autonomic Systems (HIGH) Metadata: PMID: 37580720 Journal Title: Thyroid research Publication Year: 2023 SJR 2023: 0.492 Subject Area and Category: - Endocrine and Autonomic Systems Structured Abstract: PURPOSE: Lacrimal gland enlargement can be a feature of thyroid eye disease (TED). Unilateral or asymmetric lacrimal gland enlargement is poorly described and may impede diagnosis. We present the histological and clinical findings of four patients with asymmetric lacrimal gland enlargement. METHODS: A retrospective case note review was performed for patients over two tertiary orbital clinics (Royal Adelaide Hospital, South Australia and the Sussex Eye Hospital, Brighton, United Kingdom) presenting with an asymmetrical lacrimal gland enlargement with a background of TED that underwent biopsy to exclude alternate diagnoses. Baseline data was collected for each patient and histopathological images and reports were reviewed. RESULTS: All four patients were hyperthyroid at time of lacrimal gland biopsy. Biopsy demonstrated nonspecific, lymphoid aggregates, typically of B cell type, with no diagnostic findings to support lymphocyte clonality or IgG4-related disease. One biopsy specimen demonstrated evidence of some fibrosis. CONCLUSION: Asymmetrical lacrimal gland enlargement can occur as part of the TED spectrum but may require biopsy to exclude alternate pathology. Histology demonstrates a non-specific lymphocytic infiltrate. GENERATED: In patients with thyroid eye disease, asymmetric lacrimal gland enlargement may occur and can mimic alternative orbital pathology. In this series, biopsy showed only nonspecific benign lymphoid infiltration without evidence of lymphoma or IgG4-related disease, supporting asymmetric lacrimal gland in- volvement as a possible manifestation of TED. These findings may aid clinical recognition, although biopsy remains important when the presentation is atypical and alternative diagnoses must be excluded. Table 11: Example from Endocrine and Autonomic Systems. ROUGE-1: 0.358, ROUGE-2: 0.065, ROUGE-L: 0.232, BLEU: 0.020, Perplexity: 46.5, Semantic Sim.: 88, Writing Style Sim.: 82, Non-Contradiction Rate: 96, Numeric Consistency: 100, Formality Sim.: 94. 24 Preprint. Under review. E.3 Example 3 Example: Advanced and Specialized Nursing (HIGH) Metadata: PMID: 32951753 Journal Title: Complementary therapies in medicine Publication Year: 2020 SJR 2020: 0.58 Subject Area and Category: - Advanced and Specialized Nursing Structured Abstract: BACKGROUND AND OBJECTIVE: Walnut intake is considered a healthy dietary approach worldwide, particularly as a nutritional tool for the management of obe- sity and cardiometabolic disorders. Among these lines, leptin and adiponectin, as well as glycemic biomarkers, deserve further attention. We aimed to exam- ine the impact of walnut intake on circulation levels of leptin and adiponectin through a systematic review and meta-analysis of randomized clinical trials (RCTs); secondarily, assessing the glycemic profile as well. METHODS: The literature search was implemented in four following databases: Web of Science, Scopus, PubMed/Medline, and Google Scholar, thus, determining studies that measured the effects of walnut consumption on adiponectin, leptin, and glycemic biomarkers levels from 2004 up to December 2019. RESULTS: Fourteen trials were include in the meta-analysis, with an interven- tion period ranging from 5 weeks to 12 months.Walnut intake increased leptin (weighted mean difference (WMD): 2.502 ng/mL; 95 % CI: 2.147-2.856, p 0.001) and adiponectin (WMD: 0.440 ng/mL; 95 % CI: 0.323 to 0.557, p 0.001) levels. Pertaining to glycemic biomarkers, neither overall analyses nor sub-analyses corroborated with changes in fasting blood glucose (WMD: 0.500 mg/dL, 95 % CI: -0.596, 1.596, p = 0.371), insulin (WMD: -0.21 mg/dL, 95 % CI: -0.67, 0.24, p = 0.367), and glycated hemoglobin (WMD: 0.004 mg/dL, 95 % CI: -0.041, 0.049, p = 0.870) concentrations. CONCLUSION: Walnut intake may increase leptin and adiponectin levels but does not improve glycemic biomarkers. GENERATED: Walnut consumption appears to favorably modulate adipokine status by increasing circulating adiponectin and leptin levels, while showing no significant effect on fasting blood glucose, insulin, or glycated hemoglobin. These findings suggest that walnuts may contribute to obesity and cardiometabolic man- agement primarily through effects on adipokine regulation rather than glycemic control. Further well-designed randomized clinical trials are warranted to confirm these effects and clarify their clinical relevance. Table 12: Example from Advanced and Specialized Nursing. ROUGE-1: 0.198, ROUGE-2: 0.000, ROUGE-L: 0.123, BLEU: 0.004, Perplexity: 15.7, Semantic Sim.: 88, Writing Style Sim.: 82, Non-Contradiction Rate: 96, Numeric Consistency: 100, Formality Sim.: 90. 25 Preprint. Under review. E.4 Example 4 Example: Environmental Science (HIGH) Metadata: PMID: 31054526 Journal Title: Environmental research Publication Year: 2019 SJR 2019: 1.52 Subject Area and Category: Environmental Science Structured Abstract: BACKGROUND: Hypertension and air pollution are two important risk factors for cardiovascular morbidity and mortality. Although several studies suggest that air pollution has a significant impact on blood pressure, studies on long-term effects are sparse and still controversial. OBJECTIVE: To evaluate the effects of exposure of outdoor workers to different levels of traffic-generated PM2.5 on blood pressure. DESIGN: This is an observational panel study. PARTICIPANTS: 88 non-smoking workers exposed to different concentrations of air pollution were evaluated weekly along four successive weeks. MEASUREMENTS: In each week, personal monitoring of 24-h PM2.5 concentra- tion and 24-h ambulatory blood pressure were measured. The association between blood pressure variables and PM2.5, adjusted for age, body mass index, time in job, daily work hours, diabetes, hypertension and cholesterol was assessed by means of multiple linear regression models fitted by least squares. RESULTS: Exposure to PM2.5 (ranging from 8.5 to 89.7μg/m[3]) is significantly and consistently associated with an increase in average blood pressure. An eleva- tion of 10μg/m[3] in the concentration of PM2.5 is associated with increments of 3.9 m Hg (CI 95% = [1.5; 6.3]) in average systolic 24-h blood pressure for hypertensive and/or diabetic workers. CONCLUSION: Exposure to fine particles, predominantly from vehicular traffic, is associated with elevated blood pressure in hypertensive and/or diabetic workers. GENERATED: Personal exposure to traffic-generated PM2.5 was positively associ- ated with higher 24-hour blood pressure in this cohort of outdoor workers, with particularly pronounced effects among workers with hypertension and/or diabetes. These findings support the view that long-term exposure to fine particulate air pollution contributes to elevated blood pressure and may increase cardiovascular risk in susceptible populations. Table 13: Example from Environmental Science. ROUGE-1: 0.375, ROUGE-2: 0.231, ROUGE- L: 0.300, BLEU: 0.039, Perplexity: 15.8, Semantic Sim.: 88, Writing Style Sim.: 82, Non- Contradiction Rate: 98, Numeric Consistency: 100, Formality Sim.: 95. 26 Preprint. Under review. E.5 Example 5 Example: Emergency Nursing (HIGH) Metadata: PMID: 21458134 Journal Title: Resuscitation Publication Year: 2011 SJR 2011: 1.736 Subject Area and Category: - Emergency Nursing Structured Abstract: AIM: Body mass index (BMI) may influence the quality of cardiopulmonary re- suscitation and may influence prognosis after cardiac arrest. To review the direct effect of obesity on outcome after cardiac arrest, the following cohort study was conducted. METHODS: This study based on a cardiac arrest registry comprising all adult patients with cardiac arrest of non-traumatic origin and restoration of sponta- neous circulation (ROSC) admitted to the department of emergency medicine of a tertiary-care facility. Data were collected between January 1992 and December 2007 according to the Utstein criteria. We assessed the association between BMI accord- ing to the WHO classification (underweight, BMI18.5; normal weight, 18.5-24.9; overweight, 25.0-29.9; obese≥ 30), six-month survival and neurological recovery. RESULTS: Analysis was carried out on a total of 1915 adult patients (32% female). Patients had a median age of 59 years (interquartile range [IQR] 49-70) and a median BMI of 26.0 (IQR 23.9-29.1). Survival to six months was 50%. There was no significant difference in survival between the BMI groups (underweight 46%, normal weight 47%, overweight 52%, obese 51%). In a multivariate analysis neurological outcome was better in overweight patients as compared to subjects with normal BMI (odds ratio 1.35; 95% confidence interval 1.02-1.79). CONCLUSION: Body mass index may have no direct influence on six-month survival after cardiac arrest, but patients with moderately elevated BMI may have a better neurological prognosis. GENERATED: In this cohort of adult patients with non-traumatic cardiac arrest and ROSC, BMI was not significantly associated with six-month survival. How- ever, overweight patients showed a modestly better neurological outcome than patients with normal BMI after adjustment for confounders, suggesting that excess body weight does not adversely affect prognosis after cardiac arrest and may be associated with improved neurological recovery. Table 14: Example from Emergency Nursing. ROUGE-1: 0.341, ROUGE-2: 0.140, ROUGE- L: 0.182, BLEU: 0.027, Perplexity: 20.2, Semantic Sim.: 88, Writing Style Sim.: 82, Non- Contradiction Rate: 94, Numeric Consistency: 96, Formality Sim.: 90. 27 Preprint. Under review. E.6 Example 6 Example: Pollution (LOW) Metadata: PMID: 36497716 Journal Title: International journal of environmental research and public health Publication Year: 2022 SJR 2022: 0.828 Subject Area and Category: - Pollution Structured Abstract: BACKGROUND: Urinary incontinence (UI) and poor sleep negatively affect health- related quality of life (HRQoL). This study explored the UI-related factors and the relationships between UI, sleep quality, and HRQoL. METHODS: This cross-sectional study collected data from 237 women with type 2 diabetes. Multivariate logistic regression was conducted to identify the factors associated with UI. One-way analysis of variance was used to compare the mean sleep quality and HRQoL scores of women without UI and those who experienced UI of varying severities. Correlation coefficients were estimated, and multivariate linear regression was conducted to examine the relationships between UI severity, sleep quality, and HRQoL. RESULTS: Of the 237 women, 115 (48.52%) experienced UI and 139 (58.65%) were poor sleepers. The three factors associated with UI were advanced age, a higher body mass index, and a history of vaginal delivery. Significant associations between UI severity and sleep quality and between sleep quality and HRQoL were revealed. UI severity and night-time voiding frequency were both associated with sleep quality. CONCLUSION: One factor associated with UI (body mass index) is modifiable. UI severity is associated with sleep quality as the possible influence of night-time voiding frequency on sleep quality has been considered. GENERATED: Urinary incontinence is common among women with type 2 dia- betes and is associated with older age, higher body mass index, and a history of vaginal delivery. Greater UI severity, particularly with more frequent night-time voiding, is related to poorer sleep quality, and poor sleep is in turn associated with worse health-related quality of life. These findings highlight the importance of routine assessment and management of UI and sleep problems in this population to improve overall well-being. Table 15: Example from Pollution. ROUGE-1: 0.378, ROUGE-2: 0.165, ROUGE-L: 0.252, BLEU: 0.030, Perplexity: 18.8, Semantic Sim.: 62, Writing Style Sim.: 68, Non-Contradiction Rate: 78, Numeric Consistency: 100, Formality Sim.: 85. 28 Preprint. Under review. E.7 Example 7 Example: Health, Toxicology and Mutagenesis (LOW) Metadata: PMID: 28934092 Journal Title: Environmental health perspectives Publication Year: 2017 SJR 2017: 3.41 Subject Area and Category: - Health, Toxicology and Mutagenesis Structured Abstract: BACKGROUND: Some epidemiologic and laboratory studies suggest that insec- ticides are related to increased breast cancer risk, but the evidence is inconsistent. Women engaged in agricultural work or who reside in agricultural areas may experience appreciable exposures to a wide range of insecticides. OBJECTIVE: We examined associations between insecticide use and breast can- cer incidence among wives of pesticide applicators (farmers) in the prospective Agricultural Health Study. METHODS: Farmers and their wives provided information on insecticide use, de- mographics, and reproductive history at enrollment in 1993-1997 and in 5-y follow- up interviews. Cancer incidence was determined via cancer registries. Among 30,594 wives with no history of breast cancer before enrollment, we examined breast cancer risk in relation to the women’s and their husbands’ insecticide use using Cox proportional hazards regression to estimate adjusted hazard ratios (HRs) and 95% confidence intervals (CIs). RESULTS: During an average 14.7-y follow-up, 39% of the women reported ever using insecticides, and 1,081 were diagnosed with breast cancer ... Among women who did not apply pesticides, the husband’s use of fonofos was associated with elevated risk, although no exposure-response trend was observed. CONCLUSION: Use of several organophosphate insecticides was associated with elevated breast cancer risk. However, associations for the women’s and husbands’ use of these insecticides showed limited concordance. Ongoing cohort follow-up may help clarify the relationship, if any, between individual insecticide exposures and breast cancer risk. https://doi.org/10.1289/EHP1295. GENERATED: Overall insecticide use was not associated with breast cancer inci- dence among wives of pesticide applicators; however, use of specific insecticides, particularly the organophosphates chlorpyrifos and terbufos, was associated with modestly increased risk. These associations appeared to be driven primarily by premenopausal breast cancer, with little evidence of variation by estrogen receptor status. The observed elevation in risk with husbands’ fonofos use among non- applying wives, in the absence of an exposure-response trend, warrants cautious interpretation. These findings suggest that certain insecticides may contribute to breast cancer risk in this agricultural population and merit further investigation. Table 16: Example from Health, Toxicology and Mutagenesis. ROUGE-1: 0.381, ROUGE-2: 0.083, ROUGE-L: 0.245, BLEU: 0.027, Perplexity: 21.5, Semantic Sim.: 62, Writing Style Sim.: 78, Non-Contradiction Rate: 58, Numeric Consistency: 100, Formality Sim.: 92. 29 Preprint. Under review. E.8 Example 8 Example: Computer Science Applications (LOW) Metadata: PMID: 39014177 Journal Title: International journal of computer assisted radiology and surgery Publication Year: 2025 SJR 2025: 0.658 Subject Area and Category: - Computer Science Applications Structured Abstract: PURPOSE: Augmented reality guidance in laparoscopic liver resection requires the registration of a preoperative 3D model to the intraoperative 2D image. However, 3D-2D liver registration poses challenges owing to the liver ’s flexibility, particularly in the limited visibility conditions of laparoscopy. Although promising, the current registration methods are computationally expensive and often necessitate manual initialisation. METHODS: The first neural model predicting the registration (NM) is proposed, represented as 3D model deformation coefficients, from image landmarks. The strategy consists in training a patient-specific model based on synthetic data gener- ated automatically from the patient’s preoperative model. A liver shape modelling technique, which further reduces time complexity, is also proposed. RESULTS: The NM method was evaluated using the target registration error measure, showing an accuracy on par with existing methods, all based on numerical optimisation. Notably, NM runs much faster, offering the possibility of achieving real-time inference, a significant step ahead in this field. CONCLUSION: The proposed method represents the first neural method for 3D-2D liver registration. Preliminary experimental findings show comparable performance to existing methods, with superior computational efficiency. These results suggest a potential to deeply impact liver registration techniques. GENERATED: A patient-specific neural approach can accurately estimate 3D-2D liver registration from image landmarks while substantially reducing computa- tional cost relative to optimisation-based methods. By leveraging automatically generated synthetic data and an efficient liver shape modelling strategy, the pro- posed framework removes the need for costly numerical optimisation and supports the prospect of real-time augmented reality guidance in laparoscopic liver resection. Table 17: Example from Computer Science Applications. ROUGE-1: 0.303, ROUGE-2: 0.082, ROUGE-L: 0.182, BLEU: 0.025, Perplexity: 96.5, Semantic Sim.: 62, Writing Style Sim.: 54, Non-Contradiction Rate: 78, Numeric Consistency: 100, Formality Sim.: 88. 30 Preprint. Under review. E.9 Example 9 Example: Applied Microbiology and Biotechnology (LOW) Metadata: PMID: 26497155 Journal Title: Journal of applied microbiology Publication Year: 2016 SJR 2016: 0.84 Subject Area and Category: Biochemistry, Genetics and Molecular Biology Immunology and Microbiology Medicine - Applied Microbiology and Biotechnology - Biotechnology - Medicine Structured Abstract: AIMS: The effect of ohmic heating (OH) in a pilot plant system which had a zig- zag shaped (elbow-type) ohmic heater with five sequential voltage electrodes was investigated on Bacillus cereus spores in a commercial tsuyu sauce. METHODS AND RESULTS: The electrical field was fixed at 26·7 V cm(-1) with an alternating current frequency of 25 kHz. Raw tsuyu sauce (50 l) inoculated with B. cereus spores was submitted in a 4×3 factorial design to the OH system and heated at 95, 105, 115, and 125 ◦ C each for 30, 60, and 90 s. Survival of B. cereus spores and colour change in the commercial tsuyu sauce were both measured before and after treatment. As the treatment temperature and time increased, the number of surviving B. cereus spores decreased. The OH treatment in a bath-type process at 105 ◦ C for more than 30 s resulted in the total inactivation of the inoculated B. cereus spores (average 5·4 log reductions to undetectable levels after treatment). The OH protocol of heating at 105 ◦ C for 60 s which ensure complete eradication of the inoculated spores without compromising product quality was chosen and investigated for its suitability for commercial application on bulk quantities of samples (500 l). Reliable and reproducible reductions in B. cereus spore counts of 4·7-5·5 log CFU ml(-1) (mean±standard deviation = 5·1±0·3 CFU ml(-1)) were achieved by the selected protocol of the continuous OH treatment (105 ◦ C for 60 s). SIGNIFICANCE AND IMPACT OF THE STUDY: This procedure will enhance the microbiological quality of liquid foods while minimizing quality deterioration. CONCLUSION: This study suggests that OH treatment with five sequential elbow- type electrodes has great potential as an industrial sterilizing method for liquid food contaminated with B. cereus spores. GENERATED: Ohmic heating using the pilot-scale zig-zag electrode system effec- tively inactivated Bacillus cereus spores in commercial tsuyu sauce in a temperature- and time-dependent manner. Heating at 105 ◦ C for 60 s achieved complete spore eradication in the bath-type process and produced reliable, reproducible reductions during continuous treatment of bulk samples, while maintaining product colour quality. These findings indicate that the developed ohmic heating protocol is a prac- tical approach for commercial processing of liquid foods to improve microbiological safety with minimal quality deterioration. Table 18: Example from Applied Microbiology and Biotechnology. ROUGE-1: 0.177, ROUGE-2: 0.036, ROUGE-L: 0.088, BLEU: 0.006, Perplexity: 79.3, Semantic Sim.: 62, Writing Style Sim.: 58, Non-Contradiction Rate: 78, Numeric Consistency: 12, Formality Sim.: 84. 31 Preprint. Under review. E.10 Example 10 Example: Software (LOW) Metadata: PMID: 29157445 Journal Title: Computer methods and programs in biomedicine Publication Year: 2018 SJR 2018: 0.753 Subject Area and Category: - Software Structured Abstract: BACKGROUND AND OBJECTIVES: Diabetic retinopathy (DR) is one of the leading causes of preventable blindness in the world. Its earliest sign are red lesions, a general term that groups both microaneurysms (MAs) and hemorrhages (HEs). In daily clinical practice, these lesions are manually detected by physicians using fundus photographs. However, this task is tedious and time consuming, and requires an intensive effort due to the small size of the lesions and their lack of contrast. Computer-assisted diagnosis of DR based on red lesion detection is being actively explored due to its improvement effects both in clinicians consistency and accuracy. Moreover, it provides comprehensive feedback that is easy to assess by the physicians. Several methods for detecting red lesions have been proposed in the literature, most of them based on characterizing lesion candidates using hand crafted features, and classifying them into true or false positive detections. Deep learning based approaches, by contrast, are scarce in this domain due to the high expense of annotating the lesions manually. METHODS: In this paper we propose a novel method for red lesion detection based on combining both deep learned and domain knowledge. Features learned by a convolutional neural network (CNN) are augmented by incorporating hand crafted features. Such ensemble vector of descriptors is used afterwards to identify true lesion candidates using a Random Forest classifier. RESULTS: We empirically observed that combining both sources of information significantly improve results with respect to using each approach separately. Fur- thermore, our method reported the highest performance on a per-lesion basis on DIARETDB1 and e-ophtha, and for screening and need for referral on MESSIDOR compared to a second human expert. CONCLUSION: Results highlight the fact that integrating manually engineered approaches with deep learned features is relevant to improve results when the net- works are trained from lesion-level annotated data. An open source implementation of our system is publicly available at https://github.com/ignaciorlando/red-lesion- detection. GENERATED: The proposed red lesion detection framework, which integrates CNN-derived features with hand-crafted descriptors and classifies lesion candi- dates using a Random Forest, provides superior performance to either deep learned or hand-crafted features alone. Its strong results across DIARETDB1, e-ophtha, and MESSIDOR indicate that combining data-driven representation learning with domain knowledge is an effective strategy for accurate red lesion detection and for supporting diabetic retinopathy screening and referral decisions in clinical practice. Table 19: Example from Software. ROUGE-1: 0.264, ROUGE-2: 0.050, ROUGE-L: 0.198, BLEU: 0.008, Perplexity: 106.9, Semantic Sim.: 62, Writing Style Sim.: 58, Non-Contradiction Rate: 88, Numeric Consistency: 100, Formality Sim.: 84. 32 Preprint. Under review. E.11 The conclusion–summary distinction holds across categories Section 4.2 shows that summary-mode outputs recover most of the semantic similarity of conclusion-mode outputs while diverging sharply in writing style and numeric consistency. We now test whether this pattern is universal or driven by a subset of categories, by computing the per-category gap (ModeA−B) across all five judge dimensions forGPT-5.4. Across all 112 categories, the gap is always positive for writing style similarity (min+3.8, mean +8.3, max +13.7) and numeric consistency (min +11.3, mean +21.6, max +41.3). Table 20 shows the ten categories with the largest|∆Semantic Similarity|. Even within this group, the gap magnitudes vary substantially across dimensions: numeric consistency ranges from+18.6 to+33.0 and writing style from+4.5 to+13.7, indicating that the specific dimensions along which the two modes diverge are category-dependent. Biotechnology is a notable outlier: it is the only entry where summary-mode outputs are semantically closer to the reference (∆ =−4.0), yet numeric consistency still drops by +27.2 points. Table 21 isolates the ten categories where the semantic gap is smallest (|∆|<0.3). Even here, writing style still differs by+4.8 to+8.4 points and numeric consistency by+14.3 to+23.1 points. Semantic convergence does not imply behavioral equivalence: summaries differ from conclusions in phrasing, structure, and numeric detail even when their meaning is indistinguishable. This confirms that the conclusion–summary distinction is a structural property of the two discourse functions, not an artifact of category-level heterogeneity. Category ∆ Semantic Sim. ∆ Writing Style ∆ Non-Contrad. Rate ∆ Numeric Cons. ∆ Formality Sim. Immunology & Microbiology+6.0+13.0+4.8+28.0+2.3 Pharm., Toxicology & Pharmaceutics+5.5+13.5+4.9+33.0+2.4 Exp. & Cognitive Psychology+5.3+7.9+3.4+27.5+2.1 Arts & Humanities+5.1+12.3+6.5+32.4+2.8 Virology+4.8+13.0+7.2+23.2+3.1 Nursing+4.4+13.7+3.8+30.3+1.8 Social Psychology+4.1+11.4+2.9+18.6+2.1 Biotechnology−4.0+4.5−4.6+27.2+0.6 Education+3.9+12.3+3.6+27.6+1.7 Health Informatics+3.5+12.7+4.0+31.5+3.1 Table 20: Top 10 categories ranked by|∆Semantic Similarity|(Mode A − B ;GPT-5.4). Gap magnitudes vary substantially across dimensions within this group, indicating that the specific dimensions along which conclusion and summary modes diverge are category- dependent. Category ∆ Semantic Sim. ∆ Writing Style ∆ Non-Contrad. Rate ∆ Numeric Cons. ∆ Formality Sim. Cancer Research+0.2+7.6−0.5+16.7+0.7 Organic Chemistry−0.2+7.1−1.2+18.1+1.2 Endocrinology, Diabetes & Metab.+0.1+7.3−0.2+14.3+0.9 Cellular & Mol. Neuroscience −0.1+6.8+0.4+16.5+1.1 Hepatology−0.1+4.8+0.7+15.1+1.1 Biochemistry−0.1+5.4+0.6+16.4+0.2 Applied Psychology+0.0+5.9−1.2+23.1+2.5 Dentistry−0.0+8.4−0.5+19.6+1.1 Critical Care & ICM+0.0+5.6+0.6+18.6+0.4 Pharmaceutical Science−0.0+6.1−0.2+17.7+1.2 Table 21: Bottom 10 categories ranked by|∆Semantic Similarity|(ModeA−B;GPT-5.4). Despite near-zero semantic gaps, writing style (+4.8 to+8.4) and numeric consistency (+14.3 to +23.1) remain substantially positive. 33 Preprint. Under review. F MedConclusion dataset statistics SJR. The SJR score quantifies a journal’s scientific influence by accounting for both the volume of citations received and the prestige of the citing sources. It is computed via an iterative algorithm analogous to PageRank: a citation from a highly-ranked journal contributes more to a journal’s SJR score than one from a lower-ranked journal, and self- citations are down-weighted to mitigate inflation (Gonz ́ alez-Pereira et al., 2010). This design renders SJR size-independent and more robust to citation manipulation than raw impact factors. Because SJR scores are published annually, our dataset captures the longitudinal prestige trajectory of each journal from its earliest available record through 2024. F.1 Abstracts’ year distribution 200020052010201520202025 Publication year 0.0% 1.0% 2.0% 3.0% 4.0% 5.0% 6.0% 7.0% Share of abstracts Peak year: 2024 425,802 abstracts Publication-Year Density of Abstracts Bars show the annual share; the curve is a Gaussian-smoothed density over publication years. Figure 4: Publication-year density of abstracts in MedConclusion from 2000–2025. The median publication year is 2018. F.2 Journal category & SJR score distribution 02468101214 Proportion (%) Cancer Research Psychiatry and Mental Health Cardiology and Cardiovascular Medicine Radiology, Nuclear Medicine and Imaging Public Health, Environmental and Occupational Health Pharmacology Neurology Oncology Surgery Medicine 2.1% 2.2% 2.7% 2.8% 3.1% 3.1% 3.9% 4.0% 5.0% 13.8% (a) Top 10 Category Distribution 0.10.20.51251020 SJR Score 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 Density Median = 0.77 Mean = 0.98 (b) SJR Score Distribution Figure 5: Distribution statistics of MedConclusion. (a) Top 10 subject categories by propor- tion of abstracts in the dataset, with Medicine being the dominant category at 13.8%. (b) Distribution of SJR scores across the 3,772 journals (median=0.77, mean=0.98), showing a right-skewed distribution concentrated in the low-to-moderate prestige range. 34 Preprint. Under review. G Conclusion Label Variants We use the conclusion label variants shown in Table 22 to identify conclusion-type sections in structured abstracts. Conclusion Label Variants • CONCLUSION • CONCLUSIONS • CONCLUSION(S) • CONCLUSIONS AND RELEVANCE • CONCLUSION AND RELEVANCE • CONCLUSIONS AND IMPLICATIONS • CONCLUSION AND IMPLICATIONS • CONCLUSIONS AND IMPORTANCE • CONCLUSION AND IMPORTANCE • CONCLUSION AND SIGNIFICANCE • CONCLUSIONS AND SIGNIFICANCE • CONCLUSION AND INTERPRETATION • CONCLUSIONS AND INTERPRETATION • CONCLUSIONS AND CLINICAL RELEVANCE • CONCLUSION AND CLINICAL RELEVANCE • CONCLUSIONS AND CLINICAL IMPORTANCE • CONCLUSION AND CLINICAL IMPORTANCE • AUTHORS’ CONCLUSIONS • AUTHORS’ CONCLUSION • MAIN CONCLUSIONS • MAIN CONCLUSION Table 22: Conclusion label variants used to identify conclusion-type sections in structured abstracts. 35 Preprint. Under review. H Model configurations Short NameFull NameAccessCapabilityScale Max New TokensTemp. General-purpose Models GPT-5.4gpt-5.4 (OpenAI, 2026)ProprietaryGeneral-purposeLarge10240 Gemini 3.1 Progemini-3.1-pro (Google DeepMind, 2026)ProprietaryGeneral-purposeLarge10240 Gemini 3 Flashgemini-3-flash (Google DeepMind, 2025)ProprietaryGeneral-purpose Medium10240 DeepSeek-V3.2DeepSeek-V3.2 (DeepSeek-AI, 2025b)ProprietaryGeneral-purposeLarge10240 MiniMax-M2.1MiniMax-M2.1 (MiniMax, 2025)ProprietaryGeneral-purposeLarge10240 Gemma-3-27Bgemma-3-27b-it (Gemma Team, 2025)Open-weight General-purpose Medium10240 Llama-3.1-8BLlama-3.1-8B-Instruct (Grattafiori et al., 2024) Open-weight General-purposeSmall10240 Gemma-2-9Bgemma-2-9b-it (Gemma Team, 2024)Open-weight General-purposeSmall10240 Qwen2.5-7BQwen2.5-7B-Instruct (Qwen Team, 2024)Open-weight General-purposeSmall10240 Qwen3-4BQwen3-4B-Instruct-2507 (Qwen Team, 2025b) Open-weight General-purposeSmall10240 Llama-3.2-1BLlama-3.2-1B-Instruct (Meta, 2024)Open-weight General-purposeSmall10240 Reasoning Models Kimi-K2Kimi-K2-Thinking (Kimi Team et al., 2026)Open-weightReasoningLarge10240 DeepSeek-R1DeepSeek-R1 (DeepSeek-AI, 2025a)Open-weightReasoningLarge10240 Vision-Language Models GLM-4.6VGLM-4.6V (Team et al., 2025)ProprietaryVision-languageLarge10240 Qwen2.5-VL-7B Qwen2.5-VL-7B-Instruct (Qwen Team, 2025a) Open-weight Vision-languageSmall10240 Table 23: Model metadata and run-configuration summary for all evaluated models. For general-purpose models that expose a thinking mode, thinking=none was used. 36 Preprint. Under review. I The use of Large Language Models (LLMs) LLM is used only to aid writing quality (proofreading and polishing grammar). No ideas, claims, methods, results, or references are generated by LLMs. All content decisions and revisions are made by the authors. 37