Paper deep dive
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
Ting Yin, Danning Li, Chen Shu, Xiaoxia Yao, Boyu Fu, Yujing Chang, Tianyu Shi, Mengna Feng, Jie Chen, Jing Fu, Xiuli Xiao, Tianlin Li, Mumin Shao, Jiaxin Bi, Wenchuan Zhang, Xiaoyan Wu, Xiao Han, Zhang Zhang, Yuhao Yi, Hong Bu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/5/2026, 5:00:11 AM
Summary
The paper introduces CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7,901 paired whole-slide images and diagnostic reports. CorePath demonstrates superior performance in cancer detection, invasion assessment, and histological subtyping across multiple private and public cohorts compared to PRISM and other foundation models. Additionally, the authors developed CorePath-CRG, a risk-controlled report generation framework that integrates conformal prediction and Learn-Then-Test (LTT) methods to minimize hallucinations and enable selective report release, achieving zero non-breast hallucinations in released outputs.
Entities (9)
Relation Signals (8)
CorePath → finetunedfrom → PRISM
confidence 98% · CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM
CorePath-CRG → extends → CorePath
confidence 96% · CorePath-CRG, a Conformalized Report Generation extension of CorePath
CorePath → outperforms → PRISM
confidence 95% · CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping
CorePath-CRG → usesmethod → Learn-Then-Test
confidence 94% · combined conformal subtype-confidence gating with Learn-Then-Test risk control
CorePath-CRG → usesmethod → Conformal Prediction
confidence 94% · CorePath-CRG further combined conformal subtype-confidence gating
CorePath → evaluatedon → BCNB
confidence 92% · On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping
CorePath → evaluatedon → BRACS
confidence 92% · 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification
CorePath → trainedon → West China Hospital
confidence 85% · 7901 paired CNB whole-slide images and diagnostic reports from two centers... West China Hospital
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breast pathology benchmarks without task-specific retraining, CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping. It achieved weighted area under the receiver operating characteristic curves (AUCs) of 0.9526-0.9735 for five-class CNB histological subtyping across private centers. On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping, 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification. In report generation, CorePath reduced the overall non-breast hallucinations from 30.1% to 2.8%, demonstrating improved domain fidelity after breast-specific adaptation. CorePath-CRG further combined conformal subtype-confidence gating with Learn-Then-Test risk control to enable selective report release, subtype-level fallback, and deferral. CorePath-CRG achieved zero non-breast hallucinations among released outputs and showed the strongest overall performance in pathologist-validated LLM-based Evaluation Scores and quantitative report-generation metrics across most centers. These results demonstrate that domain-specialized foundation models with statistical risk control offer a promising approach for accurate breast CNB diagnosis and reliable report generation.
Tags
Links
- Source: https://arxiv.org/abs/2608.03079v1
- Canonical: https://arxiv.org/abs/2608.03079v1
Trouble viewing inline? Open PDF directly →
Full Text
129,103 characters extracted from source content.
Expand or collapse full text
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation Ting Yin1,2 Danning Li311footnotemark: 1 Chen Shu4 Xiaoxia Yao1 Boyu Fu5 Yujing Chang1 Tianyu Shi4 Mengna Feng1 Jie Chen2 Jing Fu6 Xiuli Xiao7,8 Tianlin Li7 Mumin Shao9 Jiaxin Bi9 Wenchuan Zhang1,2 Xiaoyan Wu1,2 Xiao Han10 Zhang Zhang1,222footnotemark: 2 Yuhao Yi1,2,422footnotemark: 2 Hong Bu1,2 1Department of Pathology, West China Hospital, Sichuan University, Chengdu, China; 2Institute of Clinical Pathology, West China Hospital, Sichuan University, Chengdu, China; 3The Hong Kong University of Science and Technology (Guangzhou); 4College of Computer Science, Sichuan University; 5Sichuan University-Pittsburgh Institute, Sichuan University; 6Department of Pathology, Sichuan Provincial People’s Hospital, Chengdu, China; 7Department of Pathology, The Affiliated Hospital of Southwest Medical University, Luzhou, China; 8Department of Pathology, The Fourth Affiliated Hospital of Southwest Medical University, Meishan, China; 9Department of Pathology, Shenzhen Traditional Chinese Medicine Hospital, The Fourth Clinical Medical College of Guangzhou University of Chinese Medicine, Shenzhen, China; 10College of Biomedical Engineering, Sichuan University, Chengdu, China yuhaoyi@scu.edu.cn, zhangzhang714@163.com, xiao_han@scu.edu.cn Equal contribution.Corresponding authors. Abstract Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breast pathology benchmarks without task-specific retraining, CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping. It achieved weighted area under the receiver operating characteristic curves (AUCs) of 0.9526-0.9735 for five-class CNB histological subtyping across private centers. On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping, 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification. In report generation, CorePath reduced the overall non-breast hallucinations from 30.1% to 2.8%, demonstrating improved domain fidelity after breast-specific adaptation. CorePath-CRG further combined conformal subtype-confidence gating with Learn-Then-Test risk control to enable selective report release, subtype-level fallback, and deferral. CorePath-CRG achieved zero non-breast hallucinations among released outputs and showed the strongest overall performance in pathologist-validated LLM-based Evaluation Scores and quantitative report-generation metrics across most centers. These results demonstrate that domain-specialized foundation models with statistical risk control offer a promising approach for accurate breast CNB diagnosis and reliable report generation. 1 Introduction Breast cancer is the most commonly diagnosed malignancy and a leading cause of cancer-related death among women worldwide Gradishar et al. (2024); Bray et al. (2024). Percutaneous core needle biopsy (CNB) is widely used as the initial diagnostic procedure for breast cancer, particularly for non-palpable lesions Yang et al. (2023). Despite its central role in clinical management, CNB remains a diagnostically challenging setting Collins (2021); Shaaban (2025). Unlike surgical excision specimens, CNB provides limited tissue and may incompletely sample heterogeneous lesions, thereby reducing diagnostically useful morphologic context Bilous (2010). Interpretation of CNB is further complicated by subtle morphologic overlap among diagnostically adjacent entities and by cross-institutional variation in tissue handling, staining, scanning, and reporting Quintana and Collins (2018); Wang et al. (2025b). Consequently, breast CNB interpretation is expertise-intensive and contributes substantial workload in routine pathology services Shaaban (2025). Recent advances in pathology foundation models and multimodal generative systems have created new opportunities for artificial intelligence (AI)-assisted diagnosis and report drafting from whole-slide images (WSIs) Huang et al. (2023); Lu et al. (2024); Shaikovski et al. (2024). By learning transferable visual-language representations from large-scale histopathology data, these models have shown encouraging performance in slide-level understanding and related downstream tasks. However, their direct application to routine breast CNB diagnosis remains challenging. Existing pathology foundation models are generally developed for broad pathology applications rather than for the specialized diagnostic setting of breast CNB. Moreover, although multimodal pathology systems can generate coherent free-text outputs, their clinical adoption in high-risk diagnostic workflows requires both reliable uncertainty estimation and the ability to defer cases that cannot be assessed with sufficient confidence Olsson et al. (2022). Fine-tuning refers to adapting a pretrained foundation model to a target clinical task using domain-specific supervision, while parameter-efficient approaches aim to retain broad pretrained knowledge by updating only a small subset of model parameters Sung et al. (2022a). Paired WSI-report data provide a natural source of supervision by linking tissue morphology with case-level diagnostic interpretation Ding et al. (2025). Motivated by these considerations, we fine-tuned PRISM using parameter-efficient fine-tuning on paired breast WSIs and reports to develop CorePath, a breast CNB pathology foundation model designed to improve the alignment between slide-level histomorphologic representations and downstream breast CNB diagnosis while preserving the general capabilities of the pretrained model Shaikovski et al. (2024). Importantly, improved alignment and stronger predictive performance alone are not sufficient for safe clinical integration. Many existing AI models generate deterministic point predictions without explicit uncertainty quantification, making it difficult for clinicians to assess the reliability of individual model outputs Olsson et al. (2022). This concern is particularly relevant for report generation, where clinically plausible language may still contain unsupported or hallucinated content Wang et al. (2025a). Conformal prediction provides a statistical framework for attaching calibrated uncertainty estimates to model outputs, thereby enabling explicit control of predictive risk and the detection of unreliable predictions Vovk et al. (2005); Angelopoulos et al. (2025). Previously, conformal prediction-based methods have found pathology applications in classification and segmentation tasks, but they have remained unexplored for open-ended pathology report generation Olsson et al. (2022); Zhang et al. (2026); Wieslander et al. (2021). Building on this principle, we develop CorePath-CRG, a Conformalized Report Generation extension of CorePath for uncertainty-aware classification and report generation. By issuing outputs only when confidence is sufficient and abstaining otherwise, CorePath-CRG is intended to support safer and more clinically deployable AI-assisted reporting for breast CNB, particularly across heterogeneous real-world settings. Here, we developed CorePath, a breast-specialized pathology foundation model adapted from a multimodal pathology foundation model through paired breast CNB WSI-report supervision and parameter-efficient Perceiver-only fine-tuning. We further developed CorePath-CRG, a risk-controlled report-generation framework designed to improve the reliability of pathology report generation. We evaluated CorePath across multicenter private cohorts and public breast pathology benchmarks to determine whether it can support breast disease classification at progressively finer diagnostic granularities, ranging from cancer detection and invasion-status categorization to histological subtyping. In addition, we assessed CorePath-CRG for hallucination control and clinical report quality. Collectively, these models provide a clinically oriented framework for breast CNB interpretation, aiming to improve diagnostic accuracy, cross-institutional robustness, and reporting reliability in high-risk clinical scenarios. 2 Methods 2.1 Breast Core Needle Biopsy Samples and Cohorts A total of 10946 Hematoxylin and Eosin (H&E)-stained breast CNB slides were retrospectively included. All slides were collected from West China Hospital (WCH), Sichuan Provincial People’s Hospital (SPH), the Affiliated Hospital of Southwest Medical University (SWH), West China Tianfu Hospital (WTH), Chengdu Shang Jin Nan Fu Hospital (SJH), and Shenzhen Traditional Chinese Medicine Hospital (SZH), forming training and test cohorts spanning from June 2019 to September 2024, enabling a comprehensive evaluation of cross-center generalization under distribution shifts Aubreville et al. (2023); Ehteshami Bejnordi et al. (2017). The WCH cohort was split into WCH-1 (n=5720 WSIs, June 2019 to April 2024) and WCH-2 (n=496 WSIs, May 2024 to August 2024). The SPH cohort was split into SPH-1 (n=2181 WSIs, January 2023 to April 2024) and SPH-2 (n=612 WSIs, May 2024 to September 2024). The WSI-report pairs of WCH-1 and SPH-1 were used for fine-tuning PRISM. Table 1 summarizes the characteristics of these datasets and details their training, calibration, and testing splits. The overview of the above datasets is shown in Table S1- S6. The inclusion criteria were as follows: (1) having undergone a breast biopsy with a definitive pathological diagnosis and (2) each slide should be clear and undamaged. Patients were excluded if the corresponding biopsy histopathological sections were unavailable. This study was approved by the ethics committee of participating hospitals (No. 997 in 2025) and abided by the Declaration of Helsinki before using tissue samples for scientific research purposes only. The requirement to obtain informed consent from the participants was waived by the ethics committee. Table 1: Summary of datasets used for model fine-tuning, conformal calibration, risk control, and zero-shot evaluation. Datasets WSIs Usage Tasks & Applications Model development SPH-1 & WCH-1 7901 Fine-tuning CorePath construction SPH-2 612 Calibration; Test CorePath-CRG construction; hierarchical diagnosis* Independent private validation cohorts SWH 1392 Test Report generation; hierarchical diagnosis* WCH-2 496 Test Report generation; hierarchical diagnosis* WTH 282 Test Report generation; hierarchical diagnosis* SJH 128 Test Report generation; hierarchical diagnosis* SZH 135 Test Cancer detection Independent public validation cohorts BCNB 1058 Test Invasive carcinoma subtype categorization BRACS 547 Test Lesion stratification; fine-grained classification *Hierarchical diagnosis includes cancer detection, invasion assessment, and histological subtyping. 2.2 Whole-Slide Image and Report Preprocessing All breast CNB WSIs underwent standardized tissue detection and patch extraction before feature encoding. Background regions and non-informative areas were excluded, and tissue-containing regions were tiled into patches matching the input resolution required by each model. For TITAN, CONCH patch-level features were extracted following the TITAN pipeline Ding et al. (2025). For PRISM and CorePath, WSI-level representations were obtained from patch features generated by the Virchow patch encoder Shaikovski et al. (2024). For report preprocessing, each WSI was paired with its original diagnostic pathology report. Because the reports were written in Chinese and exhibited substantial heterogeneity in wording and formatting across institutions, DeepSeek-R1-Distill-Qwen-32B was first used to extract structured diagnostic information, which was subsequently used to define diagnostic labels for downstream classification tasks DeepSeek-AI (2025). In parallel, the original Chinese diagnostic information was translated into English using OpenBioLLM-70B, an open-source biomedical large language model, to provide supervision for WSI-report alignment Ankit Pal (2024). When needed, the translated reports were manually reviewed to ensure the preservation of key diagnostic entities. Patient-identifying information was not retained in the processed text. Ultimately, our final datasets comprise paired WSI and pathology reports, with each WSI uniquely associated with one diagnostic report. 2.3 Development of CorePath We developed CorePath by fine-tuning the PRISM foundation model on 7901 breast CNB WSI-report pairs collected from WCH-1 and SPH-1 cohorts Shaikovski et al. (2024). These data covered routine diagnostic breast pathology cases and were used exclusively for supervised multimodal fine-tuning, with no overlap with any downstream evaluation or calibration datasets. This strict separation was designed to ensure that subsequent zero-shot classification and report-generation evaluations assessed genuine out-of-distribution generalization. In this study, we adopted a lightweight Perceiver-only fine-tuning strategy. All trainable parameters were restricted to the slide encoder, which was implemented as a Perceiver module for aggregating patch-level features into slide-level representations Jaegle et al. (2021). The text encoder and all cross-modal components outside the slide encoder were kept frozen throughout training Kirkpatrick et al. (2017). By restricting adaptation to the visual encoder, CorePath was encouraged to learn breast-specific histopathological representations while preserving the linguistic embedding space inherited from pretraining Sung et al. (2022b). Notably, keeping the language components frozen prevents the model from overfitting to fixed linguistic patterns in the training corpus and exploiting text distributional biases that are irrelevant to the visual content. Parameter-efficient fine-tuning was implemented by inserting adapters into targeted feed-forward sublayers within the Perceiver architecture Houlsby et al. (2019). These targeted modules include the feed-forward network of the final Transformer block, as well as those associated with the cross-attention and aggregation stages. Consequently, parameter updates are strictly confined to this compact subset of layers, while all remaining parameters of PRISM remain frozen. Comprehensive implementation details, including specific layer mappings, hyperparameter configurations, and learning rates, are provided in Supplementary Methods A.1. This strictly constrained fine-tuning protocol ensures that adaptation focuses on visual pattern abstraction at the whole-slide level, rather than the memorization of textual templates or reporting conventions. As a result, CorePath was suitable for downstream zero-shot evaluation, uncertainty-aware prediction, and risk-controlled inference across heterogeneous clinical centers Zhou et al. (2022). 2.4 Zero-Shot Diagnostic Classification Evaluation of CorePath To assess the diagnostic generalization and practical adaptability of CorePath, we evaluated the model in a zero-shot setting without task-specific retraining. Zero-shot classification assigns slides to predefined diagnostic categories using text prompts, rather than a task-specific classifier trained on labeled examples for each category set Radford et al. (2021). This strategy is well suited to breast CNB diagnosis, where hierarchical and workflow-dependent label spaces make exhaustive task-specific annotation impractical. For each task, we defined candidate diagnostic prompts and performed subtype prediction without updating model parameters. The prompts used in the zero-shot experiments are listed in Tables S7-S12. We evaluated PRISM, TITAN, and CorePath on private datasets from six independent centers, including SPH-2, SWH, WCH-2, WTH, SJH, and SZH, comprising 612, 1392, 496, 282, 128, and 135 cases, respectively. All private evaluation cohorts consisted of breast CNB specimens, consistent with the tissue context of the fine-tuning data, but none overlapped with the WCH-1 or SPH-1 cohorts used for CorePath construction. As summarized in Table 1, the private cohorts were used to evaluate a clinically progressive hierarchical diagnosis setting. For the internal cohorts, hierarchical diagnosis comprised three diagnostic levels: cancer detection, invasion assessment, and histological subtyping. Cancer detection distinguished cancer from noncancer. Invasion assessment separated noncancer, carcinoma in situ, and invasive breast carcinoma. Histological subtyping distinguished noncancer, carcinoma in situ, invasive ductal carcinoma (also referred to as invasive breast carcinoma of no special type, invasive breast carcinoma NOS), invasive lobular carcinoma, and other invasive breast carcinoma. These diagnostic levels were designed to reflect the stepwise workflow of breast CNB evaluation, ranging from initial cancer screening to assessment of invasion status and histological subtype determination. This design enabled us to examine whether CorePath improves not only coarse-level cancer detection but also more nuanced diagnostic distinctions that directly inform pathology reporting and downstream clinical management. The SZH cohort was available only for cancer detection and invasion assessment. Two public benchmarks were also included for external evaluation. The BCNB dataset included 1,058 CNB cases and was used for invasive carcinoma subtype categorization, with invasive ductal carcinoma, invasive lobular carcinoma, and other invasive breast carcinoma as candidate labels Xu et al. (2021). The BRACS dataset included 547 H&E-stained mastectomy or CNB specimens and was used for lesion stratification and fine-grained breast lesion classification Brancati et al. (2022). In the lesion stratification task, BRACS cases were grouped into benign, atypical, and malignant lesions. In the fine-grained classification task, the candidate labels comprised normal tissue, pathological benign lesion, usual ductal hyperplasia, flat epithelial atypia, atypical ductal hyperplasia, ductal carcinoma in situ, and invasive carcinoma. Because the label systems and specimen contexts differed between the private and public datasets, the public benchmarks were used primarily to assess representational transferability and external robustness rather than direct label-level comparability. 2.5 Development of CorePath-CRG for Reliable Report Generation CorePath-CRG was developed as a conformalized and risk-controlled extension of CorePath for reliable pathology report generation in breast CNB. CorePath-CRG included three operational stages: conformal calibration of diagnostic predictions, report scoring and Learn-Then-Test (LTT)-based threshold calibration, and three-tier selective release. Methodologically, these stages integrated five key elements in sequence: task-specific conformal calibration, Self-confidence Score-based candidate-report assessment, a separate ground-truth-aligned Judge Score for calibration-risk definition, LTT-based calibration of the release threshold, and agent-based finalization of diagnostic statements Vovk et al. (2005); Angelopoulos et al. (2025). The SPH-2 cohort was used as an independent calibration cohort for both diagnostic-confidence calibration and report-release calibration, whereas the SWH, WCH-2, WTH, and SJH cohorts were used only for subsequent report-quality evaluation (Table 1). For each WSI, CorePath first generated diagnostic probabilities through zero-shot inference, together with multiple candidate pathology reports. In the first stage, diagnostic confidence was calibrated from the probabilities produced by CorePath. For cancer detection, the probability of malignancy was compared with conformal lower and upper thresholds. For histological subtyping, the confidence margin between the two most probable subtypes was compared with a conformal margin threshold. These conformal rules retained reliable subtype-level predictions as auxiliary diagnostic evidence, whereas unstable classifier predictions were masked as “Unknown” and were not used as subtype-level evidence for report scoring or synthesis. In the second stage, candidate reports were assessed using the Self-Confidence Score, which served as an internal report-level reliability measure for CorePath-CRG. The Self-Confidence Score evaluated the clinical plausibility, breast-pathology relevance, diagnostic completeness, inter-report consistency, and linguistic coherence of candidate reports without access to ground-truth reference reports at inference. This score was combined with the conformal diagnostic confidence margin to obtain a fused release score SiS_i. During calibration, a separate ground-truth-aligned Judge Score was used to define binary risk labels for candidate outputs, and the LTT framework was applied to select a calibrated release threshold that controlled the risk of automatically released reports at a prespecified level. In the third stage, CorePath-CRG applied a three-tier selective release mechanism at inference. Cases with SiS_i above the calibrated release threshold were routed to trusted synthesis, in which a synthesizer agent integrated the retained candidate reports and the conformal-filtered diagnostic prediction into a concise diagnostic statement for automatic release. Cases below the release threshold but with a retained conformal classifier prediction were routed to a conservative subtype-only fallback pathway, in which unsupported narrative text was discarded and only the subtype-level diagnostic label was returned. Only when both the generated narrative failed the release criterion and the conformal classifier lacked a confident diagnostic label did the system output “Unknown” and defer the case for pathologist review. This three-stage design separated reference-based calibration from inference-time release decisions and agent-based output construction. During calibration, the Judge Score was used only to define calibration risk for LTT and select the release threshold. During inference, the Self-confidence Score and conformal diagnostic margin were combined into a fused release score, which was compared with the LTT-calibrated release threshold to decide whether CorePath-CRG should release the generated narrative, fall back to a subtype-level output, or defer the case. The Synthesizer Agent then constructed the final response according to this release decision and the available diagnostic evidence. Detailed implementations of conformal calibration, the Self-confidence Score, the Judge Score, LTT-based risk control, and agent-based synthesis and fallback handling are provided in Supplementary Methods A.2, A.3, A.4, A.5, and A.6, respectively, with the associated Supplementary Figures illustrating the Self-Confidence Score prompt (Figure S1), the Judge Score prompt (Figure S2), and the Synthesizer Agent prompt (Figure S3). 2.6 Report Quality Assessment We evaluated generated pathology reports across four independent medical centers, including SWH, WCH-2, WTH, and SJH. To evaluate report generation performance, we compared the reports generated by PRISM, CorePath, and CorePath-CRG. Question-answering foundation models were omitted from the comparison, as their responses are inherently contingent upon the specific questions posed. Report quality was assessed from three complementary perspectives: domain-level hallucination, reference-based textual similarity, and clinical semantic quality. For hallucination analysis, we focused on severe non-breast hallucinations, defined as generated reports containing disease entities or diagnostic descriptions unrelated to breast pathology, thereby reflecting a fundamental misidentification of the pathological domain. Reports generated by PRISM and CorePath were evaluated by a Qwen3 model111We use the official release checkpoint Qwen3-30B-A3B-Instruct-2507. to identify any non-breast hallucinations Team (2025). CorePath-CRG was designed for Conformalized Report Generation that incorporates constrained decoding and domain-specific guardrails to decrease hallucination rate. However, evaluation for CorePath-CRG is excluded from the primary comparison to maintain a fair evaluation under identical unconstrained generation conditions. To evaluate clinical utility and semantic fidelity beyond surface-level lexical overlap, we used a large language model (LLM)-as-a-judge Evaluation Score implemented through the DeepSeek API. For each case, the generated report was compared with the ground-truth reference report under an honesty-prioritized rubric in which correct information was favored over omitted or unknown information, and omitted or unknown information was favored over incorrect or hallucinated content. The Evaluation Score aggregated four clinically critical dimensions with predefined weights: Clinical Factuality & Honesty (40%), Clinical Completeness (30%), Logical Consistency (20%), and Professionalism & Fluency (10%). The resulting weighted composite score ranged from 0 to 1, with higher values indicating greater clinical accuracy, completeness, internal consistency, and professional readability. Additional implementation details for the Evaluation Score, including report extraction, scoring workflow, and the full LLM-as-a-judge rubric, are provided in Supplementary Methods A.7 and Figure S4. To validate the LLM-based Evaluation Score against pathologist assessment, we conducted an additional blinded, repeated pathologist review of a score-stratified subset of reports. The mean score across three independent review sessions was used as the pathologist reference score. Detailed procedures for report sampling, blinding, and scoring are provided in Supplementary Methods A.8. To quantify textual agreement with reference pathology reports, we also computed standard automatic text-similarity metrics, including BLEU, ROUGE, and METEOR (formal definitions provided in Supplementary Methods A.9). BLEU measures n-gram precision with a brevity penalty, capturing surface-level overlap with reference reports. ROUGE evaluates recall-oriented n-gram and subsequence matching, emphasizing content coverage. METEOR extends beyond exact matching by incorporating synonymy and stemming, offering a more semantically aware measure of lexical similarity. All metrics were computed using their standard definitions and default parameter settings. PRISM, CorePath, and CorePath-CRG were evaluated in SWH, WCH-2, WTH, and SJH. 2.7 Evaluation and Statistical Analysis Diagnostic classification performance was evaluated using accuracy, the weighted F1 score, and the weighted area under the receiver operating characteristic curve (AUC). For multiclass tasks, AUC was computed using a one-vs-rest approach and averaged across classes with class-frequency weighting. Classification metrics were calculated on the complete test sets, with 95% confidence intervals estimated by nonparametric bootstrapping using 1000 resamples. Report-generation quality was summarized using the hallucination, text-similarity, and Evaluation Score metrics described above. LLM-based Evaluation Scores were compared across the three matched model groups using a one-way repeated-measures analysis of variance with the Geisser-Greenhouse correction. When the overall difference was significant, Tukey’s multiple-comparisons test was used for pairwise comparisons. Statistical analyses were performed using GraphPad Prism 9.5 (GraphPad Software, San Diego, CA). A two-sided P<0.05P<0.05 was considered statistically significant. For pathologist validation of the LLM-based Evaluation Score, intra-rater reliability across the three independent review sessions was quantified using two-way mixed-effects, absolute-agreement intraclass correlation coefficients (ICC). ICC(A,1) represented the reliability of a single-session score, whereas ICC(A,3) represented the reliability of the mean of three session scores. Agreement between the LLM-based Evaluation Score and pathologist reference score was evaluated overall and separately for PRISM, CorePath, and CorePath-CRG. Spearman’s ρ assessed rank-order association, whereas a two-way mixed-effects, absolute-agreement, single-measure ICC, denoted ICC(A,1), quantified absolute agreement between the LLM-based Evaluation Score and the corresponding pathologist reference score. ICC estimates were reported with 95% confidence intervals. Mean absolute difference quantified the average magnitude of scoring error, and Bland-Altman analysis was used to estimate the mean signed bias (LLM-based Evaluation Score minus pathologist reference score) and the 95% limits of agreement. 3 Results 3.1 Overview of CorePath and CorePath-CRG Figure 1 summarizes the construction of CorePath and its risk-controlled CorePath-CRG extension. Paired breast CNB WSIs and pathology reports from two institutions are used to fine-tune PRISM and derive CorePath. The performance of CorePath was then assessed in independent multicenter cohorts and public datasets, with prespecified tasks spanning cancer detection, invasion assessment, histological subtyping, and fine-grained lesion classification. This design was intended to evaluate whether breast-specific adaptation improved diagnostic performance and generalizability across institutions, specimen sources, and hierarchical diagnostic tasks. We further developed CorePath-CRG as a three-tier selective report-generation framework. Conformal calibration retained high-confidence diagnostic predictions, while a fused score combining candidate-report Self-Confidence and diagnostic confidence was compared with an LTT-calibrated release threshold. Based on the release decision and the availability of a confident conformal prediction, CorePath-CRG automatically released a synthesized diagnostic statement, returned a subtype-only fallback, or output “Unknown” and deferred the case for pathologist review. Figure 1: Overview of CorePath and CorePath-CRG for breast CNB diagnosis and conformalized report generation. CorePath was developed by fine-tuning the PRISM pathology foundation model using paired CNB WSIs and pathology reports from two institutions. The resulting CorePath model was used for zero-shot diagnostic classification across predefined breast CNB tasks. For report generation, CorePath produced diagnostic probabilities and multiple candidate pathology reports for each case. CorePath-CRG first applied conformal thresholding to retain high-confidence subtype predictions as auxiliary diagnostic evidence or mask uncertain predictions as “Unknown”. Candidate reports were then assessed using a fused release score that combined the report-level Self-Confidence Score with the conformal diagnostic confidence margin, with the Judge Score used only for calibration-risk definition. An LTT-calibrated threshold enabled three-tier selective release: trusted synthesis and automatic release, subtype-only fallback, or “Unknown” output with deferral for pathologist review. CNB, core needle biopsy; WSI, whole-slide image; FFN, feed-forward network; LTT, learn-then-test. 3.2 CorePath Performance Across Breast Pathology Tasks CorePath demonstrated consistently strong zero-shot performance in breast tumor subtype prediction across independent external cohorts. Figure 2 and Tables S13-S16 summarize the results. Across the private multicenter CNB cohorts, CorePath consistently outperformed the base PRISM model throughout the hierarchical diagnostic workflow defined in Table 1, from cancer detection to invasion assessment and histological subtyping. For cancer detection, CorePath outperformed PRISM on all three metrics in every cohort and achieved the highest weighted AUC in all six cohorts, with values ranging from 0.9669 to 0.9989 (Figure 2a, Table S13). It also achieved the highest accuracy and weighted F1 score in SPH-2, SWH, WCH-2, and WTH. In SJH and SZH, TITAN achieved higher accuracy and weighted F1 scores, whereas CorePath achieved the highest weighted AUC. For invasion assessment, CorePath achieved the highest accuracy, weighted F1 score, and weighted AUC in all six private cohorts, with weighted AUCs ranging from 0.9643 to 0.9881 (Figure 2b, Table S14). Its advantage remained pronounced in histological subtyping, a more demanding task requiring discrimination among noncancer, carcinoma in situ, invasive breast carcinoma NOS, invasive lobular carcinoma, and other invasive breast carcinoma (Figure 2c, Table S15). CorePath achieved the highest values for all three metrics in evaluated centers, with weighted AUCs ranging from 0.9526 to 0.9735. These findings indicate that breast-specific multimodal fine-tuning improves generalization for clinically relevant subtype discrimination under cross-center distribution shifts. The public benchmarks provided complementary evidence of external generalization (Figure 2d, Table S16). In the BCNB cohort, an external CNB dataset for invasive carcinoma subtype classification, CorePath achieved the best overall performance, with the highest accuracy of 0.8989, weighted F1 score of 0.8919, and weighted AUC of 0.7780, supporting its generalizability beyond the participating institutions within the same specimen type. In the BRACS cohort, which contains H&E-stained tissue samples obtained by biopsy or surgical resection and thus represents a broader specimen context, CorePath achieved the highest values for all three metrics in both lesion stratification and fine-grained classification, with weighted AUCs of 0.8178 and 0.8252, respectively. Because the public benchmarks differ from the private CNB cohorts in specimen composition and label definitions, these results are interpreted as evidence of representational transferability and external robustness. Taken together, these observations show that fine-tuning PRISM on paired breast CNB WSI-report data yields consistent gains across cancer detection, invasion assessment, histological subtyping, invasive carcinoma subtype categorization, lesion stratification, and fine-grained classification, supporting the use of CorePath as a more generalizable diagnostic backbone for downstream breast CNB analysis. Figure 2: Zero-shot diagnostic classification performance across independent private and public breast pathology cohorts. (a) Cancer detection on independent private CNB cohorts. (b) Invasion assessment on independent private CNB cohorts. (c) Histological subtyping on independent private CNB cohorts. (d) Public benchmark evaluation on BCNB and BRACS, where the BCNB 3-class task corresponds to invasive carcinoma subtype categorization, and the BRACS 3-class and 7-class tasks correspond to lesion stratification and fine-grained classification, respectively. Bars show weighted AUCs for PRISM, TITAN, and CorePath; dashed horizontal lines indicate the average weighted AUC of each model within the corresponding panel. CNB, core needle biopsy; AUC, area under the receiver operating characteristic curve. 3.3 Safety and Clinical Quality of Generated Reports Beyond diagnostic classification, we evaluated the safety and clinical quality of CorePath and CorePath-CRG for pathology report generation. As shown in Figure 3a, CorePath substantially reduced non-breast hallucinations compared with PRISM, decreasing the overall hallucination rate from 30.1% to 2.8%. Across individual centers, hallucination rates ranged from 27.6% to 40.6% for PRISM and from 2.3% to 3.9% for CorePath. The high hallucination rates of PRISM underscore the inherent limitations of general-purpose models in specialized workflows, where they frequently fabricate extraneous disease entities. Conversely, the consistently low and cross-center stable performance of CorePath supports the effectiveness of our domain-specialized fine-tuning, which successfully anchors the generative distribution to the breast pathology context while maintaining robust generalization. Furthermore, CorePath-CRG eliminates the residual errors entirely via constrained decoding, achieving a strict 0% hallucination rate among released outputs. This design provides flexible deployment options, combining the greater generative flexibility of the base CorePath model with the more conservative, risk-controlled outputs of CorePath-CRG. The mean of three repeated pathologist-review sessions showed excellent intra-rater reliability (ICC(A,3) = 0.993; 95% CI, 0.99-1.00) (Table S17). LLM-based Evaluation Scores showed strong rank correlation and good absolute agreement with the pathologist reference score (Spearman’s ρ=0.911ρ=0.911; ICC(A,1) = 0.843; 95% CI, 0.74-0.90), supporting their use as a complementary report-quality measure (Table S18, Figure S5). The LLM-based Evaluation Scores in Figure 3b-e and Table S19 further showed progressively higher LLM-based Evaluation Scores from PRISM to CorePath and then to CorePath-CRG. Compared with PRISM, CorePath achieved higher mean and median Evaluation Scores across four independent centers. The improvement was statistically significant in SWH, WCH-2, and WTH, while the difference was not significant in SJH. CorePath-CRG further increased Evaluation Scores relative to CorePath across all four centers, with statistically significant gains in each cohort. These findings suggest that domain-specific fine-tuning improves the baseline quality of generated pathology reports, and that CorePath-CRG can enhance reference alignment and report reliability by preferentially retaining concise, high-confidence diagnostic outputs. In addition to hallucination evaluation and the LLM-based evaluation score, we also computed standard automatic report-generation metrics. As shown in Table 2, CorePath showed superior or comparable performance to PRISM across the four external centers. The clearest improvements were observed in WTH and WCH-2, where METEOR increased from 0.2124 to 0.2597 and from 0.1944 to 0.2313, respectively. These results suggest that domain-specific fine-tuning improved the agreement between generated reports and original reference reports, especially in relatively standardized diagnostic settings. The incorporation of CorePath-CRG further improved reference-based agreement in several centers, achieving the highest values for all eight metrics in SJH and for seven of eight metrics in both WCH-2 and WTH. For example, in WCH-2, ROUGE-1 increased from 0.2013 with CorePath to 0.3031 with CorePath-CRG, while BLEU-1 increased from 0.1702 to 0.2201. In SWH, CorePath-CRG achieved higher ROUGE-1, ROUGE-2, and ROUGE-L scores but lower BLEU and METEOR scores than CorePath, suggesting a potential trade-off between conservative risk control and lexical richness. This pattern may arise because CorePath-CRG preferentially retains concise high-confidence diagnostic outputs under uncertainty, thereby preserving essential diagnostic information while reducing longer free-text descriptions. To further distinguish the effect of risk-controlled rejection from the intrinsic quality of retained reports, we repeated the automatic text-metric evaluation after excluding samples assigned to “Unknown” by CorePath-CRG (Table S20). In this retained high-confidence subset, CorePath-CRG showed improved reference-based agreement compared with both PRISM and CorePath in most settings, achieving the best performance across all eight metrics in WCH-2, WTH, and SJH, and across seven of eight metrics in SWH. These findings suggest that the lower scores observed for some metrics in the full-set analysis were partly attributable to conservative “Unknown” outputs rather than poor generation quality among accepted reports. This interpretation is supported by the post-conformal diagnostic label distributions, which show the proportion of cases retaining confident cancer-status and subtype-level predictions after conformal filtering and provide context for the subtype-only fallback and full-deferral pathways (Table S21). The final-output and rejection statistics of CorePath-CRG further show that full rejection rates ranged from 19.15% to 38.58% and narrative rejection rates ranged from 78.43% to 83.59% across centers, confirming that many low-confidence cases were intentionally routed away from free-text report release (Table S22). Together, these findings indicate that domain-specific fine-tuning combined with risk-controlled selective generation can improve report reliability and reference alignment across heterogeneous medical centers. Figure 3: Hallucination reduction and LLM-based Evaluation Score for generated pathology reports. (a) Hallucination rates of unconstrained PRISM and CorePath outputs. CorePath-CRG is omitted because its 0% rate reflects risk-controlled selective outputs and is not directly comparable. (b-e) Violin plots of LLM-based Evaluation Score distributions for PRISM, CorePath, and CorePath-CRG in SWH, WCH-2, WTH, and SJH. Dashed horizontal lines denote the first quartile, median, and third quartile within each distribution. ns, not significant; **P<0.01P<0.01; ***P<0.001P<0.001; ****P<0.0001P<0.0001. Table 2: Quantitative evaluation of PRISM, CorePath, and CorePath-CRG using automatic text-similarity metrics across four datasets. B-1 to B-4 denote BLEU-1 to BLEU-4; R-1, R-2, and R-L denote ROUGE-1, ROUGE-2, and ROUGE-L. Best results are in bold. Dataset Model B-1 B-2 B-3 B-4 R-1 R-2 R-L METEOR SWH PRISM 0.1654 0.0545 0.0333 0.0253 0.1437 0.0225 0.1267 0.1705 CorePath 0.1810 0.0567 0.0355 0.0270 0.1672 0.0199 0.1330 0.1728 CorePath-CRG 0.0963 0.0448 0.0339 0.0254 0.1871 0.0762 0.1846 0.1323 WCH-2 PRISM 0.1510 0.0628 0.0405 0.0294 0.1659 0.0406 0.1507 0.1944 CorePath 0.1702 0.0717 0.0459 0.0321 0.2013 0.0507 0.1801 0.2313 CorePath-CRG 0.2201 0.0891 0.0709 0.0555 0.3031 0.0529 0.2992 0.1970 WTH PRISM 0.1648 0.0576 0.0351 0.0265 0.1449 0.0234 0.1357 0.2124 CorePath 0.1931 0.0665 0.0427 0.0306 0.2064 0.0341 0.1930 0.2597 CorePath-CRG 0.2870 0.0984 0.0736 0.0631 0.4029 0.0364 0.4019 0.2312 SJH PRISM 0.1483 0.0560 0.0381 0.0293 0.1241 0.0149 0.1139 0.1769 CorePath 0.1561 0.0489 0.0309 0.0233 0.1506 0.0169 0.1366 0.1931 CorePath-CRG 0.2454 0.0911 0.0704 0.0576 0.3134 0.0456 0.3134 0.2003 3.4 Visualization and Report Generation Examples Figure 4 presents three representative cases comparing the diagnostic outputs of PRISM, CorePath, and CorePath-CRG against the reference diagnoses. In Case 1, the reference diagnosis was invasive breast carcinoma, not otherwise specified. CorePath generated a breast-specific malignant diagnosis and partially overlapped with the reference, but its description included additional features that were not fully concordant. In contrast, PRISM produced an incorrect diagnosis with mismatched histological terminology. CorePath-CRG generated a concise diagnosis of invasive breast carcinoma NOS and achieved high agreement with the reference. This example suggests that breast-specific fine-tuning improves diagnostic relevance, while the risk-controlled framework can further promote clinically appropriate report release when the model output is sufficiently reliable. In Case 2, the reference diagnosis was invasive carcinoma involving the breast, with consideration of a primary breast origin. Both CorePath and PRISM produced incorrect diagnoses, as reflected by the discordant terms highlighted in red. Rather than releasing a potentially misleading diagnosis, CorePath-CRG returned “Unknown”, indicating abstention under uncertainty. This case illustrates the safety-oriented behavior of the proposed framework, in which uncertain or low-confidence cases are deferred instead of being automatically reported, prompting manual review by pathologists. In Case 3, the reference diagnosis was fibroadenoma. All three models generated reports consistent with the benign diagnosis. Importantly, CorePath-CRG retained and released the correct benign diagnosis rather than unnecessarily rejecting the case, suggesting that the risk-control mechanism does not simply increase rejection but aims to preserve reliable outputs. Additional examples are illustrated in Figure S6. Overall, these qualitative examples demonstrate the complementary benefits of breast-specific model adaptation and risk-controlled selective reporting. CorePath improves the domain specificity of generated reports compared with the general pathology foundation model, whereas CorePath-CRG further enhances clinical reliability by selectively releasing high-confidence reports and abstaining from uncertain or potentially wrong outputs. Figure 4: Representative report-generation examples across different models. For each case, the left panels show the whole-slide image and a magnified region of interest, and the green panel summarizes the reference diagnosis. Model-generated outputs are shown below each case. Red text indicates an incorrect or unsupported diagnosis. Scale bars: from top to bottom, 200μ , 250μ , and 200μ . 4 Discussion In this study, we developed CorePath, a breast-specialized multimodal pathology foundation model for CNB interpretation, and CorePath-CRG, a risk-controlled framework for selective AI-assisted report release. This work addresses the need to adapt broadly pretrained pathology foundation models to clinically specialized diagnostic settings while improving the reliability of generated diagnostic text in high-risk medical workflows. CorePath and CorePath-CRG serve complementary functions, in which CorePath enhances breast-specific diagnostic capability, and CorePath-CRG prevents uncertain or potentially unreliable content from being presented as definitive. This explicit separation between diagnostic inference and output authorization provides the conceptual basis for interpreting our findings and guiding future clinical deployment. CorePath demonstrated strong diagnostic generalization, suggesting that domain-specific adaptation can further align large pathology foundation models with specialized clinical settings while preserving the flexibility of language-guided zero-shot inference Ramanathan et al. (2025). Breast CNB represents a specialized diagnostic setting in which limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can make subtype discrimination challenging Schnitt (2019). Although general pathology foundation models have demonstrated strong transferability across histopathology tasks, their representations may not be optimally aligned with the diagnostic hierarchy of breast CNB. By fine-tuning PRISM on paired CNB WSI-report data, CorePath strengthened the association between slide-level morphology and breast-specific diagnostic language, with the most pronounced gains observed in histological subtyping. Importantly, this specialization did not require training separate supervised classifiers for different label systems. Instead, CorePath can be queried using clinically defined diagnostic categories, making it well suited to hierarchical CNB workflows that progress from cancer detection to invasion assessment and histological subtyping. The same principle may also support future prompt-defined label spaces for borderline or clinically challenging entities, such as atypical ductal hyperplasia versus ductal carcinoma in situ, or carcinoma in situ with suspected microinvasion. Prior work on medical report generation has shown that readable narratives may still fail to capture the clinical accuracy required for clinician-facing workflows, and that generated reports can contain unsupported or hallucinated clinical statements Ramesh et al. (2022); Asgari et al. (2025). Such unsupported diagnostic details may mislead clinicians even when the generated report appears stylistically plausible. In our study, CorePath substantially reduced non-breast hallucinations relative to the base PRISM model, suggesting that domain-specific fine-tuning more closely aligns model generation with breast pathology. However, the residual hallucinations observed with CorePath also show that fine-tuning alone is insufficient for safe report generation. CorePath-CRG adopts a conservative operating mode, releasing concise reports only when the calibrated evidence is sufficient, falling back to subtype-level output when narrative content is unreliable, and abstaining when both report generation and subtype prediction are uncertain. This tiered behavior is consistent with uncertainty-aware medical AI, in which selective prediction and abstention are used to surface uncertainty and defer low-confidence cases rather than forcing an answer Kompa et al. (2021). At the same time, the risk-control guarantees provided by CorePath-CRG require careful translation into practice. Conformal prediction provides distribution-free, finite-sample uncertainty guarantees under standard calibration assumptions, whereas LTT calibrates predictive algorithms to achieve explicit finite-sample risk-control guarantees without model refitting Angelopoulos et al. (2023, 2025). In pathology workflows, however, calibration must remain relevant to the deployment setting, because external scanners, laboratories, and assessment practices can introduce systematic differences that affect reliability Olsson et al. (2022). Real-world deployment will therefore require local calibration, continuous monitoring, and periodic recalibration as case mix, scanners, staining protocols, and reporting practices change. The formal guarantee applies to the auto-released report tier, whereas fallback and deferral pathways serve as additional safeguards. Their effects on diagnostic accuracy, turnaround time, workload, and user trust should be evaluated prospectively. Several limitations should be acknowledged. First, this was a retrospective study, and prospective validation in real clinical workflows is required before clinical deployment. Second, although the evaluation included multiple independent private centers and two public benchmarks, the fine-tuning data were drawn from only two centers. Additional external cohorts are needed to assess robustness across broader variations in tissue processing, staining, scanning platforms, and reporting conventions. Third, the statistical guarantees of the conformal and LTT procedures rely on calibration and test examples being exchangeable, and their validity may be affected by substantial distribution shifts, changes in clinical workflow, or modifications to the underlying model and prompting strategy after calibration. Fourth, our hallucination analysis focused primarily on non-breast hallucinations, which capture severe domain-level errors but do not encompass all possible factual inaccuracies, such as unsupported tumor grade, biomarker status, laterality, or subtle morphologic overinterpretation. However, this information could be reflected indirectly in the LLM-as-a-judge clinical quality evaluation. Fifth, while the LLM-based Evaluation Scores demonstrate strong rank correlation and absolute agreement with pathologist reference scores, Bland-Altman analysis reveals a conservative bias, overestimating clinical risk. Future evaluations could integrate a domain-specific knowledge base to better resolve semantic equivalence and entailment in diagnostic narratives. Taken together, CorePath demonstrates that paired WSI-report supervision can specialize a multimodal pathology foundation model for breast CNB diagnosis while preserving zero-shot flexibility across clinically relevant diagnostic tasks. CorePath-CRG further shows that breast-specific adaptation can be integrated with explicit risk control to enable selective report release, subtype-level fallback, and deferral of uncertain cases. These findings suggest that domain-specialized and statistically risk-controlled pathology AI may offer a practical path toward more reliable clinical decision support. Before clinical implementation, prospective studies involving pathologist review of AI-assisted diagnoses and reports are required to assess diagnostic accuracy, reporting efficiency, calibration stability, patient safety, pathologist workload, and user trust in real-world clinical workflows. References [1] A. Angelopoulos, S. Bates, J. Malik, and M. I. Jordan (2020) Uncertainty sets for image classifiers using conformal prediction. arXiv preprint arXiv:2009.14193. Cited by: §A.2. [2] A. N. Angelopoulos, S. Bates, E. J. Candès, M. I. Jordan, and L. Lei (2025) Learn then test: calibrating predictive algorithms to achieve risk control. The Annals of Applied Statistics 19 (2), p. 1641–1662. Cited by: §A.5.4, §A.5.4, §1, §2.5, §4. [3] A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster (2022) Conformal risk control. arXiv preprint arXiv:2208.02814. Cited by: §A.2. [4] A. N. Angelopoulos, S. Bates, et al. (2023) Conformal prediction: a gentle introduction. Foundations and trends® in machine learning 16 (4), p. 494–591. Cited by: §A.2, §4. [5] A. N. Angelopoulos and S. Bates (2021) A gentle introduction to conformal prediction and distribution-free uncertainty quantification. arXiv preprint arXiv:2107.07511. Cited by: §A.2. [6] M. S. Ankit Pal (2024) OpenBioLLMs: advancing open-source large language models for healthcare and life sciences. Hugging Face. Note: https://huggingface.co/aaditya/OpenBioLLM-Llama3-70B Cited by: §2.2. [7] E. Asgari, N. Montaña-Brown, M. Dubois, S. Khalil, J. Balloch, J. A. Yeung, and D. Pimenta (2025) A framework to assess clinical safety and hallucination rates of llms for medical text summarisation. NPJ digital medicine 8 (1), p. 274. Cited by: §4. [8] M. Aubreville, N. Stathonikos, C. A. Bertram, R. Klopfleisch, N. Ter Hoeve, F. Ciompi, F. Wilm, C. Marzahl, T. A. Donovan, A. Maier, et al. (2023) Mitosis domain generalization in histopathology images-the MIDOG challenge. Medical Image Analysis 84, p. 102699. Cited by: §2.1. [9] M. Bilous (2010) Breast core needle biopsy: issues and controversies. Modern Pathology 23, p. S36–S45. Cited by: §1. [10] N. Brancati, A. M. Anniciello, P. Pati, D. Riccio, G. Scognamiglio, G. Jaume, G. De Pietro, M. Di Bonito, A. Foncubierta, G. Botti, et al. (2022) Bracs: a dataset for breast carcinoma subtyping in h&e histology images. Database 2022, p. baac093. Cited by: §2.4. [11] F. Bray, M. Laversanne, H. Sung, J. Ferlay, R. L. Siegel, I. Soerjomataram, and A. Jemal (2024) Global cancer statistics 2022: globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: A Cancer Journal for Clinicians 74, p. 229 – 263. External Links: Link Cited by: §1. [12] C. J. Clopper and E. S. Pearson (1934) The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26 (4), p. 404–413. Cited by: Remark A.2. [13] L. C. Collins (2021) Precision pathology as applied to breast core needle biopsy evaluation: implications for management. Modern Pathology 34, p. 48–61. Cited by: §1. [14] DeepSeek-AI (2025) DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning. External Links: 2501.12948, Link Cited by: §2.2. [15] T. Ding, S. J. Wagner, A. H. Song, R. J. Chen, M. Y. Lu, A. Zhang, A. J. Vaidya, G. Jaume, M. Shaban, A. Kim, et al. (2025) A multimodal whole-slide foundation model for pathology. Nature medicine, p. 1–13. Cited by: §1, §2.2. [16] B. Ehteshami Bejnordi, M. Veta, P. Johannes van Diest, B. Van Ginneken, N. Karssemeijer, G. Litjens, J. A. Van Der Laak, C. consortium, M. Hermsen, Q. F. Manson, et al. (2017) Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer. Jama 318 (22), p. 2199–2210. Cited by: §2.1. [17] W. J. Gradishar, M. S. Moran, J. Abraham, V. Abramson, R. Aft, D. Agnese, K. H. Allison, B. Anderson, J. Bailey, H. J. Burstein, et al. (2024) Breast cancer, version 3.2024, NCCN clinical practice guidelines in oncology. Journal of the National Comprehensive Cancer Network 22 (5), p. 331–357. Cited by: §1. [18] N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly (2019) Parameter-efficient transfer learning for NLP. In International conference on machine learning, p. 2790–2799. Cited by: §2.3. [19] Z. Huang, F. Bianchi, M. Yuksekgonul, T. J. Montine, and J. Zou (2023) A visual–language foundation model for pathology image analysis using medical twitter. Nature medicine 29 (9), p. 2307–2316. Cited by: §1. [20] A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira (2021) Perceiver: general perception with iterative attention. In International conference on machine learning, p. 4651–4664. Cited by: §2.3. [21] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al. (2017) Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences 114 (13), p. 3521–3526. Cited by: §2.3. [22] B. Kompa, J. Snoek, and A. L. Beam (2021) Second opinion needed: communicating uncertainty in medical machine learning. NPJ Digital Medicine 4 (1), p. 4. Cited by: §4. [23] E. L. Lehmann and J. P. Romano (2005) Testing statistical hypotheses. Springer. Cited by: §A.5.4. [24] M. Y. Lu, B. Chen, D. F. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, L. P. Le, G. Gerber, et al. (2024) A visual-language foundation model for computational pathology. Nature Medicine 30, p. 863–874. Cited by: §1. [25] H. Olsson, K. Kartasalo, N. Mulliqi, M. Capuccini, P. Ruusuvuori, H. Samaratunga, B. Delahunt, C. Lindskog, E. A. Janssen, A. Blilie, et al. (2022) Estimating diagnostic uncertainty in artificial intelligence assisted pathology using conformal prediction. Nature communications 13 (1), p. 7761. Cited by: §A.2, §1, §1, §4. [26] L. M. Quintana and L. C. Collins (2018) Assessing intraductal proliferations in breast core needle biopsies. Diagnostic Histopathology 24 (2), p. 49–57. Cited by: §1. [27] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §2.4. [28] V. Ramanathan, T. Xu, P. Pati, F. Ahmed, M. Goubran, and A. L. Martel (2025) Modaltune: fine-tuning slide-level foundation models with multi-modal information for multi-task learning in digital pathology. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 23912–23923. Cited by: §4. [29] V. Ramesh, N. A. Chi, and P. Rajpurkar (2022) Improving radiology report generation systems by removing hallucinated references to non-existent priors. In Machine Learning for Health, p. 456–473. Cited by: §4. [30] S. J. Schnitt (2019) Problematic issues in breast core needle biopsies. Modern Pathology 32, p. 71–76. Cited by: §4. [31] A. M. Shaaban (2025) Diagnostic pitfalls in needle core biopsy of the breast: an update. Diagnostic Histopathology. Cited by: §1. [32] G. Shaikovski, A. Casson, K. Severson, E. Zimmermann, Y. K. Wang, J. D. Kunz, J. A. Retamero, G. Oakley, D. Klimstra, C. Kanan, et al. (2024) Prism: a multi-modal generative foundation model for slide-level histopathology. arXiv preprint arXiv:2405.10254. Cited by: §1, §1, §2.2, §2.3. [33] Y. Sung, J. Cho, and M. Bansal (2022-06) VL-adapter: parameter-efficient transfer learning for vision-and-language tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 5227–5237. Cited by: §1. [34] Y. Sung, J. Cho, and M. Bansal (2022) Vl-adapter: parameter-efficient transfer learning for vision-and-language tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5227–5237. Cited by: §2.3. [35] Q. Team (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §2.6. [36] R. J. Tibshirani, R. Foygel Barber, E. Candes, and A. Ramdas (2019) Conformal prediction under covariate shift. Advances in neural information processing systems 32. Cited by: Remark A.1. [37] V. Vovk, A. Gammerman, and G. Shafer (2005) Algorithmic learning in a random world. Springer. Cited by: §1, §2.5. [38] C. Wang, W. Zhou, S. Ghosh, K. Batmanghelich, and W. Li (2025) Semantic consistency-based uncertainty quantification for factuality in radiology report generation. In Findings of the Association for Computational Linguistics: NAACL 2025, p. 1739–1754. Cited by: §1. [39] J. Wang, J. Yu, H. Yang, Y. Zhu, L. Jiang, X. Li, J. Zhang, and Y. Xu (2025) Self-supervised stain normalization empowers privacy-preserving and model generalization in digital pathology. npj Digital Medicine. Cited by: §1. [40] H. Wieslander, P. J. Harrison, G. Skogberg, S. Jackson, M. Fridén, J. Karlsson, O. Spjuth, and C. Wählby (2021) Deep learning with conformal prediction for hierarchical analysis of large-scale whole-slide tissue images. IEEE Journal of Biomedical and Health Informatics 25 (2), p. 371–380. Cited by: §1. [41] F. Xu, C. Zhu, W. Tang, Y. Wang, Y. Zhang, J. Li, H. Jiang, Z. Shi, J. Liu, and M. Jin (2021) Predicting axillary lymph node metastasis in early breast cancer using deep learning on primary tumor biopsy slides. Frontiers in oncology 11, p. 759007. Cited by: §2.4. [42] Y. Yang, Z. Liu, J. Huang, X. Sun, J. Ao, B. Zheng, W. Chen, Z. Shao, H. Hu, Y. Yang, et al. (2023) Histological diagnosis of unprocessed breast core-needle biopsy via stimulated raman scattering microscopy and multi-instance learning. Theranostics 13 (4), p. 1342. Cited by: §1. [43] X. Zhang, T. Wang, C. Yan, F. Najdawi, K. Zhou, Y. Ma, Y. Cheung, M. C. F. Yeung, and B. A. Malin (2026) Implementing trust in non-small cell lung cancer diagnosis with a conformalized uncertainty-aware ai framework. Nature Biomedical Engineering. External Links: Document, Link Cited by: §1. [44] K. Zhou, J. Yang, C. C. Loy, and Z. Liu (2022) Learning to prompt for vision-language models. International Journal of Computer Vision 130 (9), p. 2337–2348. Cited by: §2.3. Supplementary A Supplementary Methods A.1 Implementation Details Hardware and Training Environment All fine-tuning experiments were conducted in a distributed manner using Distributed Data Parallel (DDP) across 8 NVIDIA RTX 3090 GPUs. Optimization and Hyperparameters The model was optimized for a total of 6 epochs. We utilized a per-GPU batch size of 5, coupled with 16 gradient accumulation steps, resulting in an effective global batch size of 640 (5×16×85× 16× 8). The learning rate was initialized at 2×10−42× 10^-4 with a weight decay of 5×10−65× 10^-6. The overall training objective comprised multiple loss terms: the weights for both the contrastive loss and the distillation loss were set to 1.0, while the generative loss component was explicitly disabled (weight = 0.0) for this specific fine-tuning phase. Input Processing Configuration For whole-Slide Image (WSI) processing, the maximum number of tiles was capped at 2,560. To introduce stochasticity and mitigate overfitting during training, random tile sampling was enabled. For the text modality, the maximum sequence length was restricted to 768 tokens. Parameter-Efficient Fine-Tuning via Adapters To efficiently adapt the pre-trained PRISM architecture while keeping the computational footprint manageable, we injected lightweight adapter modules exclusively into the Slide Encoder (Perceiver) component. The adapter modules were configured with a bottleneck dimension of 8, a ReLU activation function, a dropout rate of 0.0, and a residual scaling factor of 1.0. Rather than updating the full parameter space, we strategically targeted the feed-forward networks (FFNs) of specific self-attention and cross-attention layers within the Perceiver blocks. The exact targeted modules include the FFNs in the 6th transformer layer of the first block (layers.0.tf.5.f), as well as the FFNs of the cross-attention layers in the first and second blocks (layers.0.xattn.f and layers.1.xattn.f). A.2 Conformal Calibration of Diagnostic Predictions Let X denote the slide-level representation extracted by CorePath, and let Y denote the corresponding subtype label. Given a pretrained classifier that outputs predictive probabilities, conformal prediction aims to construct a prediction rule whose error rate is controlled at a predefined level α, such that the probability of making an incorrect but confident prediction is upper bounded by α [25]. For cancer detection task, CorePath produces a likelihood probability p=ℙ(Y=malignant∣X)p=P(Y=malignant X) via prompt-based zero-shot inference. Instead of enforcing a single decision threshold, we introduce an ambiguity-aware conformal scheme that explicitly allows abstention [4]. Using a held-out calibration set, we collect predictive probabilities from correctly classified samples only, separating malignant cases with p>0.5p>0.5 and benign cases with p<0.5p<0.5. Two conformal thresholds are then estimated via empirical quantiles: a lower threshold tlowert_lower derived from the upper tail of benign predictions, and an upper threshold tuppert_upper derived from the lower tail of malignant predictions. For a new sample, predictions with p>tupperp>t_upper are deemed reliably malignant, predictions with p<tlowerp<t_lower are deemed reliably benign, and predictions falling in between are labeled as unstable. This construction ensures that, with probability at least 1−α1-α, a prediction assigned to either class is correct, while uncertain cases are explicitly deferred [3]. For histological subtyping, we adopt a margin-based conformal strategy. Given class probabilities pkk=1K\p_k\_k=1^K produced by CorePath through multi-class prompt inference, we define the nonconformity score as the classification margin between the top two predicted classes, namely Δ=p(1)−p(2) =p_(1)-p_(2) [5], where p(1)p_(1) and p(2)p_(2) denote the largest and second-largest probabilities, respectively [1]. Calibration is performed by collecting margins from correctly classified calibration samples and estimating a threshold tmargint_margin as the α-quantile of this distribution. At inference time, a prediction is considered reliable only if its margin exceeds tmargint_margin; otherwise, the model abstains. This margin-based conformal rule provides a principled mechanism for identifying ambiguous subtype predictions while maintaining finite-sample coverage guarantees. A.3 Self-Confidence Score for Candidate-Report Reliability Assessment Overview and Workflow. The Self-Confidence Score is a continuous metric (ranging from 0.00 to 1.00) designed to evaluate the holistic clinical and linguistic quality of AI-generated pathology reports. For each case, the evaluation pipeline takes as input a set of 5 candidate reports (generated via sampling or ensemble methods) alongside the reference classifier’s predictions. To ensure the evaluation is grounded in reliable references, we first apply a Conformal Prediction mechanism to determine the reliability of the reference classifier labels. Subsequently, a Large Language Model (LLM) auditor is prompted to assess the reports and output a structured analysis and a final continuous score. Uncertainty-Aware Auxiliary Context via Conformal Prediction. While the reference report serves as the primary standard for evaluation, the upstream classifier’s predictions are provided to the LLM auditor as auxiliary context to aid the auditing process. To prevent the auditor from unfairly penalizing AI-generated reports that contradict uncertain classifier predictions, we utilize Conformal Prediction to dynamically filter out low-confidence auxiliary labels. • Binary Malignancy Status: Given the classifier’s confidence score sbis_bi, the auxiliary prediction is deemed reliable only if it falls outside the ambiguity region (i.e., sbi>τuppers_bi> _upper or sbi<τlowers_bi< _lower). Otherwise, the auxiliary status is masked as Unknown. • Multi-Class Subtype Status: Given the classifier’s probability distribution over subtypes, we compute the prediction margin Δ=pmax−psecond_max =p_max-p_second\_max. If Δ>τmargin > _margin, the auxiliary subtype is considered reliable; otherwise, it is masked as Unknown. By marking uncertain classifier predictions as Unknown, the scoring prompt explicitly instructs the LLM auditor to bypass the classifier-alignment constraint for these specific cases, ensuring that the evaluation relies solely on the reliable reference and clinical logic when upstream predictions are noisy. LLM Auditor and Prompt Design. We employ the DeepSeek API (deepseek-chat) with a temperature of 0.0 to guarantee deterministic and reproducible scoring. The LLM is instructed to act as a breast pathology report quality auditor, evaluating the candidate reports across five critical dimensions: (1) Breast relevance, (2) Diagnostic completeness, (3) Inter-report consistency, (4) Information alignment (conditionally applied based on Conformal Prediction), and (5) Linguistic coherence. The exact prompts are detailed in Figure S1. Prompt Template for Self-Confidence Score (Self-Confidence Scoring Agent) ⬇ ### SYSTEM PROMPT ### You are a breast pathology report quality auditor. Given 5 AI-generated pathology reports and a reference classifier output, score the overall report quality on a continuous scale from 0.00 to 1.00. Focus on these aspects when forming your score: 1. Breast relevance -- do the reports clearly discuss breast tissue or breast disease? 2. Diagnostic completeness -- do they contain a meaningful clinical conclusion? 3. Consistency -- do the reports agree with each other on malignancy and subtype? 4. Information alignment -- do the generated reports align with the reference classifier? 5. Linguistic coherence -- are the reports written as coherent medical prose? Scoring guidance (use the full continuous range, avoid rounding to 0.1 steps): - Reports that are relevant, complete, consistent, and reference classifier output-aligned -> near 1.00 - Partial relevance or mild inconsistency -> 0.50-0.79 - Mostly irrelevant, incoherent, or contradicting reference classifier output -> below 0.40 - Completely off-topic or nonsensical -> near 0.00 Output format (use these exact tags, nothing else outside them): <analysis> Breast relevance: [brief note] Diagnostic completeness: [brief note] Consistency: [brief note] Information alignment: [brief note or "reference classifier unknown -- skipped"] Coherence: [brief note] </analysis> Self-Confidence_Score: X.X ### USER PROMPT ### Reference Classifier: Subtype=mu_status | Malignancy=bi_status (low-confidence -- treat as unknown) [Conditionally appended] Reports: R1: report_1 R2: report_2 R3: report_3 R4: report_4 R5: report_5 Figure S1: System and User prompts used for calculating the Self-Confidence Score via the LLM auditor. Response Parsing. The LLM’s raw output is parsed using regular expressions to extract the textual analysis enclosed within the <analysis> tags and the numerical score following the Self-Confidence_Score: prefix. Fallback patterns are implemented to capture standard floating-point formats (e.g., 0.x or 1.00) to ensure robust extraction even if the LLM slightly deviates from the strict formatting constraints. A.4 Judge Score for Ground Truth-Aligned Clinical Quality Assessment Overview and Workflow. The Judge Score is designed to evaluate the factual accuracy and diagnostic alignment of the AI-generated reports against expert-annotated references. While the Self-Confidence Score assesses holistic quality and internal consistency using reference classifier predictions, the Judge Score directly compares the generated text with human reference translations to measure clinical correctness. Ground Truth Alignment and Report Filtering. For each case, the pipeline retrieves multiple equivalent reference translations (typically 5 versions) of the ground truth (GT) diagnosis from an external clinical database, establishing a high-confidence diagnostic consensus. To ensure the LLM auditor focuses only on meaningful outputs and is not distracted by degenerate generations, we apply a heuristic filtering mechanism to the AI-generated candidate reports. A report is retained for evaluation only if it exceeds a minimum character length, contains appropriate sentence-level punctuation, and includes core pathological terminology from a predefined lexicon (e.g., malignant, benign, carcinoma, invasive, fibroadenoma). LLM Auditor and Prompt Design. Similar to the Self-Confidence Score, we employ the DeepSeek API with a temperature of 0.0 to ensure deterministic evaluation. The LLM is instructed to act as a Ground Truth Judge, focusing on three primary dimensions: (1) GT Alignment (correct identification of core malignancy and histological subtype), (2) Inter-Report Consistency, and (3) Linguistic Coherence. The exact prompts are detailed in Figure S2. Prompt Template for Judge Score ⬇ ### SYSTEM PROMPT ### You are a breast pathology report quality auditor (Ground Truth Judge). Given the Ground Truth (GT) reference translations (multiple equivalent versions) and a set of AI-generated pathology reports, score the factual accuracy and alignment of the AI reports against the GT on a continuous scale from 0.00 to 1.00. Note: The GT section contains 5 equivalent translation versions of the same diagnosis. Focus on the core diagnostic consensus (e.g., Malignancy, Histological Subtype) present in these translations. Focus on these aspects when forming your score: 1. GT Alignment -- do the AI reports correctly identify the core malignancy and histological subtype stated in the GT translations? 2. Inter-Report Consistency -- do the AI reports agree with each other on the core diagnosis? 3. Linguistic Coherence -- are the AI reports written as coherent, professional medical prose? Scoring guidance (use the full continuous range, avoid rounding to 0.1 steps): - Perfect GT alignment, high consistency, and coherent -> near 1.00 - Minor GT mismatches (e.g., subtype ambiguity) or mild inconsistency -> 0.60-0.89 - Major GT contradictions (e.g., benign vs malignant) or high variance -> 0.20-0.59 - Completely wrong diagnosis, nonsensical, or no valid reports -> near 0.00 Output format (use these exact tags, nothing else outside them): <analysis> GT Core Entities: [Malignancy, Subtype extracted from translations] AI Consensus: [Malignancy, Subtype] Alignment: [brief note] Consistency: [brief note] Coherence: [brief note] </analysis> JUDGE_Score: X.X ### USER PROMPT ### GT REFERENCE TRANSLATIONS [Confidence: gt_confidence]: gt_text AI REPORTS (N=n_valid_reports): R1: report_1 R2: report_2 ... Figure S2: System and User prompts used for calculating the Judge Score via the LLM auditor. Response Parsing. The parsing logic mirrors that of the Self-Confidence Score, utilizing regular expressions to extract the structured analysis and the continuous JUDGE_Score. Cases where the GT is unavailable in the external database are explicitly flagged with a low-confidence indicator in the user prompt, instructing the LLM to adjust its alignment expectations accordingly and preventing undue penalization of the AI models. A.5 Learn-Then-Test Calibration for Risk-Controlled Report Release This section provides the complete theoretical foundation for the Learn-Then-Test (LTT) risk control framework applied to automated pathology report release. We establish finite-sample, distribution-free guarantees for the clinical risk rate of auto-released reports, present the full algorithmic procedure, and describe the clinical implementation. A.5.1 Notation and Problem Setup Data Domain. Let X denote the space of whole slide images (WSIs), Y the space of ground-truth pathology reports, and Y the space of AI-generated candidate reports. Define the joint data domain =×^Z=X×Y× Y. Let cal=zii=1n=(xi,yi,y^i)i=1nD_cal=\z_i\_i=1^n=\(x_i,\,y_i,\, y_i)\_i=1^n denote the calibration set of n samples, where xi∈x_i is the input WSI, yi∈y_i is the ground-truth report, and y^i∈ y_i∈ Y is the corresponding set of AI-generated candidates. Exchangeability Assumption. Assumption A.1 (Exchangeability). The calibration samples z1,…,znz_1,…,z_n are exchangeable, i.e., for any permutation σ of 1,…,n\1,…,n\, (z1,…,zn)=(zσ(1),…,zσ(n)).(z_1,…,z_n)\; d=\;(z_σ(1),…,z_σ(n)). Samples drawn i.i.d. from any fixed distribution ℙP over Z satisfy Assumption A.1. Exchangeability is strictly weaker than i.i.d.: it permits marginal heterogeneity so long as the joint distribution is invariant under permutation. It is the minimal condition required for the distribution-free statistical guarantees derived below. Throughout this work we assume i.i.d. sampling, which implies exchangeability. We defer a discussion of potential violations in multi-site pathology settings to Remark A.1. Risk Indicator. Let J:→ℝJ:Z be a reference-aligned judge scoring function that evaluates the clinical acceptability of a generated report. We define the binary risk indicator R:→0,1R:Z→\0,1\ as: R(z)=(J(z)<τjudge),R(z)=I\! (J(z)< _judge ), (1) where τjudge∈ℝ _judge is a pre-specified clinical quality threshold and R(z)=1R(z)=1 indicates a high-risk (clinically unacceptable) report. The true risk of a selection policy ⊆A is: Rtrue()=z∼ℙ[R(z)∣z∈].R_true(A)=E_z \! [R(z) z ]. A.5.2 Score Construction and Continuousization LLM-based judge scores J(z)J(z) are typically integer-valued, yielding a discrete empirical distribution over calD_cal that precludes a strict total ordering of samples and leads to ambiguous threshold selection. We address this by constructing a continuous fused score via integration with complementary signals from the upstream conformal classifier. Upstream Classifier Scores. Let p^i(c) p_i^(c) denote the predicted class probability for breast cancer subtype c∈c from the upstream conformal classifier applied to sample ziz_i. We define two scalar statistics derived from the classifier’s softmax output: • Self-Confidence Score: Pi=maxc∈p^i(c)P_i= _c \, p_i^(c), the maximum predicted class probability. • Predictive Margin: Mi=p^i(c1)−p^i(c2)M_i= p_i^(c_1)- p_i^(c_2), where c1=argmaxcp^i(c)c_1= *arg\,max_c\, p_i^(c) and c2=argmaxc≠c1p^i(c)c_2= *arg\,max_c≠ c_1 p_i^(c) are the top-1 and top-2 predicted subtypes, respectively. Both PiP_i and MiM_i are continuous and derived from the classifier’s output independently of the judge score JiJ_i, providing orthogonal evidence for risk stratification. Fused Score. For each sample ziz_i, define the fused score: Si=Pi+λMi+ϵi,S_i=P_i+λ M_i+ _i, (2) where λ>0λ>0 is a weighting hyperparameter and ϵi∼i.i.d.(0,η) _i i.i.d. U(0,η) with η>0η>0 is a negligible noise term added solely for tie-breaking. Lemma A.1 (Strict Total Ordering). With probability 1, all fused scores Sii=1n\S_i\_i=1^n are distinct, inducing a strict total ordering on calD_cal. Proof. Since each ϵi∼(0,η) _i (0,η) is drawn from a continuous distribution, ℙ(ϵi=ϵj)=0P( _i= _j)=0 for any i≠ji≠ j. Therefore, for any pair i≠ji≠ j, ℙ(Si=Sj)=ℙ(ϵi=ϵj−(Pi+λMi−Pj−λMj))=0P(S_i=S_j)=P( _i= _j-(P_i+λ M_i-P_j-λ M_j))=0, since this event requires ϵi _i to take a specific deterministic value. Applying the union bound over all (n2) n2 pairs, ℙ(∃i≠j:Si=Sj)≤(n2)⋅0=0.P\! (∃\,i≠ j:S_i=S_j )\;≤\; n2· 0=0. Hence all scores are almost surely distinct. ∎ Lemma A.1 ensures that the acceptance sets defined in Section A.5.3 are well-defined and the LTT threshold selection is unambiguous. In practice we set η=10−6η=10^-6, which is negligible relative to the scale of PiP_i and MiM_i and does not materially alter the ranking. A.5.3 Learn-Then-Test Risk Control Framework Acceptance Sets. Let S(1)≥S(2)≥⋯≥S(n)S_(1)≥ S_(2)≥·s≥ S_(n) denote the sorted fused scores in descending order, and let R(j)R_(j) denote the risk indicator associated with the j-th ranked sample. For any k∈1,…,nk∈\1,…,n\, define the acceptance set: k=z∈cal:S(z)≥S(k),A_k= \z _cal:S(z)≥ S_(k) \, (3) comprising the k samples with the highest fused scores. The corresponding empirical risk is: R^k=1k∑j=1kR(j). R_k= 1k _j=1^kR_(j). (4) Risk Control Objective. Given a target risk level α∈(0,1)α∈(0,1) and a confidence parameter δ∈(0,1)δ∈(0,1), we seek the largest acceptance set k∗A_k^*, or equivalently the highest auto-release coverage k∗/nk^*/n, such that the true risk is controlled with high probability: ℙ(Rtrue(k∗)≤α)≥1−δ.P\! (R_true(A_k^*)≤α )≥ 1-δ. (5) Maximizing k∗k^* subject to (5) achieves a Pareto-optimal risk-coverage operating point. A.5.4 Finite-Sample Statistical Guarantees Clopper-Pearson Exact Confidence Bound For a fixed acceptance set kA_k of size k, let Vk=∑j=1kR(j)V_k= _j=1^kR_(j) denote the observed risk count. Under Assumption A.1, the risk indicators R(j)j=1k\R_(j)\_j=1^k are exchangeable Bernoulli random variables. Let p=Rtrue(k)p=R_true(A_k) denote the true risk probability. We note that VkV_k stochastically dominates Binomial(k,p)Binomial(k,p) under exchangeability in the sense that ℙ(Vk≥v)≥ℙ(Bin(k,p)≥v)P(V_k≥ v) (Bin(k,p)≥ v) for all v [23], ensuring that the Clopper-Pearson bound derived below remains valid (conservative) in this setting. Theorem A.2 (Clopper-Pearson Exact Bound). Let VkV_k denote the number of high-risk samples among the k calibration samples in kA_k. The (1−δ)(1-δ)-confidence upper bound on the true risk p=Rtrue(k)p=R_true(A_k) is: U(k,Vk,δ)=Beta−1(1−δ;Vk+1,k−Vk),U(k,\,V_k,\,δ)=Beta^-1\! (1-δ;\;V_k+1,\;k-V_k ), (6) where Beta−1(⋅;a,b)Beta^-1(·;\,a,b) is the quantile function of the Beta distribution with shape parameters a and b. This bound satisfies: ℙ(p≤U(k,Vk,δ))≥1−δ,P\! (p≤ U(k,V_k,δ) )≥ 1-δ, (7) with equality when p is the parameter of an exact binomial model. Proof. We exploit the classical duality between the binomial CDF and the regularized incomplete Beta function. For X∼Binomial(k,p)X (k,p) and an observed value v∈0,…,kv∈\0,…,k\: ℙ(X≤v)=I1−p(k−v,v+1),P(X≤ v)=I_1-p(k-v,\;v+1), (8) where Ix(a,b)=B(x;a,b)/B(a,b)I_x(a,b)=B(x;\,a,b)/B(a,b) is the regularized incomplete Beta function and B(a,b)B(a,b) is the Beta function. Define U(k,v,δ)U(k,v,δ) as the value p∗p^* solving ℙp∗(X≤v)=δP_p^*(X≤ v)=δ, i.e. I1−p∗(k−v,v+1)=δ.I_1-p^*(k-v,\;v+1)=δ. Using the symmetry identity I1−p(k−v,v+1)=1−Ip(v+1,k−v)I_1-p(k-v,\,v+1)=1-I_p(v+1,\,k-v) yields 1−Ip∗(v+1,k−v)=δ⟹Ip∗(v+1,k−v)=1−δ,1-I_p^*(v+1,\;k-v)=δ\; \;I_p^*(v+1,\;k-v)=1-δ, so p∗=I−1(1−δ;v+1,k−v)=Beta−1(1−δ;v+1,k−v),p^*=I^-1(1-δ;\;v+1,\;k-v)=Beta^-1(1-δ;\;v+1,\;k-v), confirming (6). To verify (7), note that for any p≤p∗p≤ p^*, the binomial distribution is stochastically dominated, so ℙp(X≤v)≥ℙp∗(X≤v)=δP_p(X≤ v) _p^*(X≤ v)=δ, which means ℙp(X>v)≤1−δP_p(X>v)≤ 1-δ. Equivalently, ℙp(p∗≥p)≥1−δP_p(p^*≥ p)≥ 1-δ, i.e. ℙ(p≤U(k,Vk,δ))≥1−δP(p≤ U(k,V_k,δ))≥ 1-δ. This bound is non-asymptotic: it does not rely on normal approximations and is valid for all k≥1k≥ 1 and v∈0,…,kv∈\0,…,k\. ∎ Main Theorem: Distribution-Free Risk Control Theorem A.3 (Distribution-Free Risk Control). Under Assumption A.1, let k∗=maxk∈1,…,n:U(k,∑j=1kR(j),δ)≤α,k^*= \! \k∈\1,…,n\:U\! (k,\; _j=1^kR_(j),\;δ )≤α \, (9) with the convention k∗=0k^*=0 if no such k exists, in which case no report is auto-released. Then: ℙ(Rtrue(k∗)≤α)≥1−δ.P\! (R_true(A_k^*)≤α )≥ 1-δ. (10) This guarantee is distribution-free: it holds for any unknown distribution ℙP over Z without parametric assumptions on the data-generating process, the score function S(⋅)S(·), or the risk indicator R(⋅)R(·). Proof Sketch. The formal proof of this result is established in [2], Theorem 1, to which we refer the reader for complete details. We outline the key argument here. Step 1: P-value construction. For each k∈1,…,nk∈\1,…,n\, define a p-value for the null hypothesis Hk:Rtrue(k)>αH_k:R_true(A_k)>α as: pk=ℙp=α(Bin(k,α)≤Vk),p_k=P_p=α\! (Bin(k,α)≤ V_k ), (11) the probability that a Binomial(k,α)Binomial(k,α) random variable does not exceed the observed count VkV_k. By the Beta-binomial duality established in the proof of Theorem A.2, the condition U(k,Vk,δ)≤αU(k,V_k,δ)≤α is equivalent to pk≤δp_k≤δ. Under Assumption A.1, each pkp_k is a valid p-value: ℙ(pk≤t∣Hkis true)≤tP(p_k≤ t H_k~is true)≤ t for all t∈(0,1)t∈(0,1). Step 2: Reformulation of the selection rule. The index k∗k^* in (9) is equivalently k∗=maxk∈1,…,n:pk≤δ,k^*= \! \k∈\1,…,n\:p_k≤δ \, i.e., the largest candidate index whose associated null hypothesis is rejected. Step 3: Familywise error rate control. A naive union bound over all n tests would inflate the error probability by a factor of n; however, since k∗k^* is selected as the maximum rejecting index over a finite ordered family of nested acceptance sets 1⊆⋯⊆nA_1 ·s _n, the procedure possesses a monotone structure that avoids this inflation. Specifically, [2] Theorem 1 establishes that for a finite family of valid p-values pkk=1n\p_k\_k=1^n and the maximum-rejecting selection rule, the probability of false acceptance satisfies: ℙ(Rtrue(k∗)>α)≤δ.P\! (R_true(A_k^*)>α )≤δ. The distribution-free property follows because pkp_k depends only on the binomial model, which holds under Assumption A.1 for any ℙP. ∎ Corollary A.4 (Inference-Time Guarantee). Define the release threshold λ^thresh=S(k∗) λ_thresh=S_(k^*), the fused score of the k∗k^*-th ranked calibration sample. At inference time, any new report with fused score S(z)≥λ^threshS(z)≥ λ_thresh is auto-released. Under Assumption A.1 applied jointly to calibration and test samples (i.e., test samples are drawn exchangeably from the same distribution ℙP), the set of auto-released reports satisfies: ℙ([R(z)∣S(z)≥λ^thresh]≤α)≥1−δ.P\! (E\! [R(z) S(z)≥ λ_thresh ]≤α )≥ 1-δ. (12) Proof. Corollary A.4 follows directly from Theorem A.3 and the definition λ^thresh=S(k∗) λ_thresh=S_(k^*). The extension from calibration to test samples holds because, under i.i.d. sampling from ℙP, calibration and test samples are jointly exchangeable, and the release event S(z)≥λ^thresh\S(z)≥ λ_thresh\ is measurable with respect to the calibration-determined threshold. ∎ A.5.5 Algorithmic Procedure Algorithm 1 summarizes the complete LTT calibration procedure. Algorithm 1 LTT Calibration for Risk-Controlled Report Release 1:Calibration set cal=(xi,yi,y^i)i=1nD_cal=\(x_i,y_i, y_i)\_i=1^n, target risk α∈(0,1)α∈(0,1), confidence level δ∈(0,1)δ∈(0,1), margin weight λ>0λ>0, judge threshold τjudge _judge 2:Release threshold λ^thresh λ_thresh 3:Phase 1: Score and risk computation 4:for i=1i=1 to n do 5: Compute Pi=maxcp^i(c)P_i= _c\, p_i^(c) and Mi=p^i(c1)−p^i(c2)M_i= p_i^(c_1)- p_i^(c_2) from the upstream classifier 6: Sample ϵi∼(0,10−6) _i (0,10^-6) 7: Set Si←Pi+λMi+ϵiS_i← P_i+λ M_i+ _i 8: Set Ri←(Ji<τjudge)R_i (J_i< _judge) via the reference-aligned judge 9:end for 10:Phase 2: Sorting 11:Sort samples by SiS_i descending to obtain S(1)≥⋯≥S(n)S_(1)≥·s≥ S_(n) with associated risks R(1),…,R(n)R_(1),…,R_(n) 12:Phase 3: Threshold selection 13:Initialize V←0V← 0, k∗←0k^*← 0 14:for k=1k=1 to n do 15: V←V+R(k)V← V+R_(k) 16: Compute Uk←Beta−1(1−δ;V+1,k−V) U_k ^-1\! (1-δ;\;V+1,\;k-V ) 17: if Uk≤αU_k≤α then 18: k∗←k^*← k 19: end if 20:end for 21:Phase 4: Output 22:if k∗=0k^*=0 then 23: return λ^thresh←+∞ λ_thresh←+∞ ⊳ No auto-release permitted 24:else 25: return λ^thresh←S(k∗) λ_thresh← S_(k^*) 26:end if Computational Complexity. Phase 2 (sorting) requires O(nlogn)O(n n) time. Phase 3 makes n sequential evaluations of the Beta quantile function, each of which is O(1)O(1) via standard library implementations, yielding O(n)O(n) for the linear scan. Total complexity: O(nlogn)O(n n). No iterative root-finding or optimization is required. A.5.6 Properties and Remarks Remark A.1 (Exchangeability in Multi-Site Settings). Assumption A.1 holds when slides are collected and processed independently under consistent protocols, with no systematic batch effects or temporal ordering biases. In multi-site settings, slides acquired under different scanners, staining protocols, or patient populations may introduce covariate shift that violates exchangeability. In such cases, weighted conformal methods that reweight calibration samples by the likelihood ratio between the calibration and test marginals [36] provide a principled extension. In this work, the calibration set is drawn from a single institutional cohort under consistent processing conditions, providing reasonable support for Assumption A.1. Remark A.2 (Tightness of the Clopper-Pearson Bound). The Clopper-Pearson (CP) bound is the tightest exact (non-asymptotic) upper confidence bound for binomial proportions: no other exact bound achieves uniformly smaller coverage error [12]. In comparison, commonly used asymptotic alternatives such as the Wald interval, Hoeffding’s inequality, or the Bernstein inequality, yield intervals that can be substantially looser when k is moderate, or when Vk/kV_k/k is near 0 or 1. This tightness is especially important in our setting, where the calibration set is of moderate size (n≈500n≈ 500), the target risk α is small (e.g. 10%10\%), and the observed risk rate R^k R_k is near zero in the high-score regime. Using a loose bound under these conditions would unnecessarily inflate the confidence interval and reduce auto-release coverage. Remark A.3 (Monotonicity and Computational Efficiency). The Clopper-Pearson upper bound U(k,v,δ)U(k,v,δ) satisfies two monotonicity properties that are central to the algorithm: (i) Monotone in v: U(k,v,δ)U(k,v,δ) is non-decreasing in v for fixed k and δ. Adding observed risks raises the upper bound, correctly reflecting increased uncertainty. (i) Monotone in k: U(k,v,δ)U(k,v,δ) is non-increasing in k at a fixed empirical risk rate v/kv/k. Larger acceptance sets yield tighter bounds due to increased sample evidence. These properties also confirm that the selection rule k∗=maxk:U(k,Vk,δ)≤αk^*= \k:U(k,V_k,δ)≤α\ is well-defined: the set k:U(k,Vk,δ)≤α\k:U(k,V_k,δ)≤α\ is non-empty whenever the empirical risk is sufficiently low, and the linear scan in Algorithm 1 correctly identifies its maximum element. Remark A.4 (Coverage-Risk Trade-off). The parameters α and δ jointly control the risk-coverage trade-off. Decreasing α enforces stricter clinical risk control at the cost of lower auto-release coverage (k∗/nk^*/n). Decreasing δ demands higher statistical confidence, making the CP bound more conservative and further reducing coverage. The LTT framework maximizes coverage subject to both constraints simultaneously, yielding a Pareto-optimal operating point on the risk-coverage curve; this is the sense in which k∗k^* is optimal. In clinical practice, α should be determined by domain experts based on acceptable diagnostic error rates. We use α=0.10α=0.10 and δ=0.10δ=0.10 as default values, corresponding to a 90%90\%-confidence guarantee that no more than 10%10\% of auto-released reports are clinically unacceptable. Summary of Guarantees. The LTT risk control framework provides finite-sample, distribution-free guarantees for the selected acceptance set. Specifically, the true risk satisfies Rtrue(k∗)≤αR_true(A_k^*)≤α with probability at least 1−δ1-δ (Theorem A.3), without requiring any parametric assumptions on ℙP. The guarantee is non-asymptotic and valid for any calibration sample size n≥1n≥ 1. It is based on the exact Clopper–Pearson upper confidence bound for the binomial proportion, which is the tightest exact bound of this form (Remark A.2). Moreover, the score function S(⋅)S(·) may be any measurable function and does not require calibration. Among all candidate acceptance sets satisfying the risk constraint, LTT selects the one with the largest empirical coverage, namely k∗/nk^*/n (Remark A.4). A.6 Agent-Based Report Finalization and Fallback Workflow Overview and Tiered Workflow. While the Learn-Then-Test (LTT) framework provides a statistically rigorous threshold λ^thresh λ_thresh for risk control, the actual deployment requires an intelligent mechanism to process the AI outputs based on this risk assessment. We design a three-tier Agent Synthesis and Feedback Mechanism that dynamically routes the generated content. Instead of blindly releasing all LLM-generated text, the system employs specialized LLM agents to either synthesize a comprehensive diagnosis, gracefully degrade to a safer output, or intercept the case for human review, ensuring that the feedback provided to the clinician is always commensurate with the system’s statistical confidence. Scoring Agent and Fused Risk Metric. Before synthesis, a Scoring Agent acts as an internal quality auditor. It evaluates the ensemble of candidate reports (e.g., 5 generated variations) alongside the upstream classifier’s predictions. The agent outputs a continuous Self-Confidence Score (Pi∈[0.00,1.00]P_i∈[0.00,1.00]) assessing breast relevance, diagnostic completeness, consistency, information alignment, and linguistic coherence (Supplementary Methods A.3). This score is then fused with the conformal predictive margin (MiM_i) to compute the final risk metric Si=Pi+λMi+ϵiS_i=P_i+λ M_i+ _i. This continuousization resolves the discrete clustering of LLM scores and enables precise thresholding. Synthesizer Agent and Tiered Feedback Strategy. Based on the fused score SiS_i, the system executes one of three feedback pathways: 1. Trusted Synthesis (Si≥λ^threshS_i≥ λ_thresh): The reports are deemed clinically safe. A Synthesizer Agent integrates the candidate reports and the classifier’s subtype into a single, professional diagnostic statement for auto-release. 2. Confident Fallback (Si<λ^threshS_i< λ_thresh, but MiM_i is sufficient): The narrative reports are deemed potentially hallucinated or inconsistent, but the upstream classifier remains confident. The system discards the LLM-generated text and falls back to outputting only the classifier’s subtype label, prioritizing safety over narrative detail. 3. Human-in-the-Loop Rejection (Both uncertain): The system outputs an “Unknown” status, strictly intercepting the case and flagging it for mandatory expert pathological review. LLM Agents and Prompt Design. Both the Self-Confidence Scoring and Synthesizer agents are powered by the DeepSeek API (temperature = 0.0) to ensure deterministic and rigorous clinical reasoning. System prompts for the Self-Confidence Scoring Agent used to derive the continuous Self-Confidence Score are detailed in Figure S1. The exact system prompts governing synthesizer agents are detailed in Figure S3. Prompt Templates for Synthesis Agent ⬇ ### SYNTHESIZER SYSTEM PROMPT ### You are an expert breast pathologist. Your task is to synthesize a set of AI-generated pathology reports and a reference classifier subtype into a single, concise, and professional final diagnosis string. Rules: - The output must be a single diagnostic statement. - Do NOT include any explanations, greetings, or extra text. Output ONLY the diagnosis string. - Follow the style and brevity of these exact examples: 1. Invasive carcinoma, favor invasive lobular carcinoma. 2. Fibroadenoma with focal calcification. 3. The breast tissue shows extensive acute and chronic inflammatory cell infiltration. Figure S3: Prompt for the Synthesizer Agent used to generate the final auto-released clinical diagnosis under the trusted pathway. Response Parsing and Execution Guarantee. The system employs strict regex-based parsing to extract the Self-Confidence_Score from the <analysis> block, ensuring that the continuousization pipeline is robust against LLM formatting variations. By decoupling the evaluation (Self-Confidence Scoring Agent) from the final output generation (Synthesizer Agent), and gating the latter behind a statistically validated conformal threshold, this mechanism guarantees that the AI copilot never presents a highly detailed but potentially hallucinated narrative to the clinician when the underlying statistical confidence is low. A.7 Evaluation Score Methodology and Prompt Design Overview and Workflow. While the Self-Confidence score and Judge score evaluate ensemble consistency and translation alignment, the Evaluation Score is designed to assess the clinical utility and safety of a single generated pathology report as it would be presented to a clinician in a real-world deployment. For each case, the pipeline retrieves a single reference report and compares it against one AI-generated report. To accommodate different generation paradigms, we employ a dynamic report extraction strategy: for PRISM and CorePath that generate multiple candidate reports (delimited by |||), one report is randomly sampled for evaluation; for the CorePath-CRG model that produces a single deterministic output, the entire text is evaluated directly. Honesty Priority and Evaluation Dimensions. A core principle of this evaluation is the Honesty Priority, which strictly enforces the clinical safety hierarchy: Correct Information > Unknown/Omitted > Incorrect/Hallucinated Information. Fabricating medical facts or making critical diagnostic errors (e.g., confusing benign and malignant lesions) is heavily penalized, whereas omitting uncertain information is strongly preferred over guessing. The LLM auditor evaluates the report across four weighted dimensions: 1. Clinical Factuality & Honesty (40%): Accuracy of medical entities and strict adherence to the honesty priority. 2. Clinical Completeness (30%): Coverage of essential diagnostic information present in the ground truth, including malignancy status, histological subtype, grade when available, and relevant morphologic findings. 3. Logical Consistency (20%): Internal coherence, ensuring the microscopic description aligns with the final diagnosis. 4. Professionalism & Fluency (10%): Use of standard breast pathology terminology and coherent medical prose. LLM Auditor and Prompt Design. We utilize the DeepSeek API with a temperature of 0.0 and an increased maximum token limit (1024 tokens) to allow for more detailed analytical reasoning. The exact prompts, including the scoring guidance and dimension weights, are detailed in Figure S4. Prompt Template for Evaluation Score ⬇ ### SYSTEM PROMPT### You are an expert pathology report quality auditor. Given the Ground Truth (GT) reference report and an AI-generated pathology report, evaluate the AI report on a continuous scale from 0.00 to 1.00. [Core Principle: Honesty & Factuality] Priority: Correct Information > Unknown/Omitted > Incorrect/Hallucinated Information. Fabricating medical facts, hallucinating non-existent lesions, or making critical diagnostic errors (e.g., benign vs. malignant) must be heavily penalized. Omitting uncertain information is strictly preferred over guessing incorrectly. [Evaluation Dimensions & Weights] 1. Clinical Factuality & Honesty (40%): Accuracy of stated diagnostic information. Strictly apply the honesty priority. 2. Diagnostic Coverage (30%): Coverage of essential diagnostic information present in the ground truth, including malignancy status, histological subtype, grade when available, and relevant morphologic findings. Do not penalize omission of information that is not present in the ground truth or not assessable on core needle biopsy. 3. Logical Consistency (20%): Internal coherence (e.g., microscopic description aligns with final diagnosis). 4. Professionalism & Fluency (10%): Standard breast pathology terminology and coherent medical prose. [Scoring Guidance] - Near 1.00: Accurate, honest, complete, and professional. - 0.60-0.89: Minor omissions or slight terminology flaws, but NO critical factual errors. - 0.20-0.59: Minor factual contradictions. - Near 0.00: Severe hallucinations, critical diagnostic errors, or nonsensical. Output format (use these exact tags, nothing else outside them): <analysis> Factuality & Honesty: [brief note] Completeness: [brief note] Consistency: [brief note] Professionalism: [brief note] </analysis> EVAL_Score: X.X ### USER PROMPT ### GT REFERENCE REPORT: gt_text AI GENERATED REPORT: ai_report Figure S4: System and User prompts used for calculating the Evaluation Score, emphasizing the Honesty Priority and weighted clinical dimensions. Response Parsing and Clinical Significance. The response parsing follows the established regex-based extraction protocol to isolate the <analysis> block and the final EVAL_Score. By strictly enforcing the honesty priority through the prompt design, this metric provides a highly reliable proxy for clinical safety that can be compared with expert assessment. Crucially, it ensures that models like the Risk-Control (CorePath-CRG) variant are appropriately rewarded for their factual reliability, rather than being penalized by traditional lexical overlap metrics (e.g., BLEU/ROUGE) for missing uncertain information. A.8 Pathologist Validation of the LLM-Based Evaluation Score To assess the alignment between the Evaluation Score assigned by the DeepSeek API and pathologist judgment, we conducted a blinded and repeated pathologist review using a subset of generated reports selected through score stratification. Twelve cases were selected from each of the four external centers (SWH, WCH-2, WTH, and SJH), yielding a total of 48 cases. Within each center, cases were sampled from four predefined intervals of the CorePath-CRG Evaluation Score, with three cases selected from each interval. For each selected case, the reports generated by PRISM, CorePath, and CorePath-CRG were all included in the evaluation. Consequently, the same 48 cases were evaluated across all three models, resulting in 144 reports for evaluation. All reports were anonymized and presented in a fully randomized order. Information regarding the generating model, originating center, and associations among reports from the same case was concealed from the reviewer. A breast pathologist independently evaluated all reports in three sessions conducted at intervals of one week. The report order was randomly shuffled, and new anonymous identifiers were assigned before each session. Ratings from previous sessions were not available during subsequent assessments. The pathologist applied the same rubric that prioritized diagnostic honesty and used the same scoring range from 0 to 1 as the DeepSeek API evaluator. The mean score across the three independent review sessions was used as the pathologist reference score. A.9 Definitions of Automated Text-Similarity Metrics We provide the formal definitions of the automatic metrics used in our quantitative evaluation. BLEU. The BLEU score is computed as: BLEU=BP⋅exp(∑n=1Nwnlogpn),BLEU=BP· ( _n=1^Nw_n p_n ), (13) where pnp_n denotes the modified n-gram precision for order n, wnw_n is the weight for each n-gram order (typically wn=1/Nw_n=1/N), and N is the maximum n-gram order (commonly N=4N=4). The brevity penalty BP is defined as: BP=1if c>r,exp(1−r/c)if c≤r,BP= cases1&if c>r,\\ (1-r/c)&if c≤ r, cases (14) where c is the length of the candidate generation and r is the effective reference length. ROUGE. ROUGE-N measures the recall of n-gram overlap between the generated text and reference: ROUGE-N=∑S∈ℛ∑gramn∈SCountmatch(gramn)∑S∈ℛ∑gramn∈SCount(gramn),ROUGE-N= _S _gram_n∈ SCount_match(gram_n) _S _gram_n∈ SCount(gram_n), (15) where ℛR denotes the set of reference reports, Count(gramn)Count(gram_n) is the total occurrence of an n-gram in the references, and Countmatch(gramn)Count_match(gram_n) is the maximum number of matching n-grams between the candidate and references. ROUGE-L, based on the longest common subsequence (LCS), is also commonly reported to capture sentence-level structural similarity. METEOR. METEOR computes a harmonic mean of unigram precision and recall, adjusted by a fragmentation penalty: METEOR=Fmean⋅(1−Pen),METEOR=F_mean·(1-Pen), (16) where Fmean=P⋅RαP+(1−α)R,F_mean= P· Rα P+(1-α)R, (17) with P and R denoting unigram precision and recall, respectively, and α is a weighting parameter (typically α=0.9α=0.9). The fragmentation penalty Pen is defined as: Pen=γ⋅(chm)θ,Pen=γ· ( chm )^θ, (18) where ch is the number of chunks (contiguous matched sequences), m is the total number of matched unigrams, and γ,θγ,θ are penalty parameters. METEOR additionally incorporates synonym matching and stemming to capture semantic equivalence beyond exact string matches. A.10 Code and Data Availability The private institutional WSIs and diagnostic reports used for model adaptation and validation are not publicly available due to patient privacy, institutional data governance, and ethical restrictions. Public benchmark datasets analyzed in this study are publicly available from their original repositories, including BCNB (https://bcnb.grand-challenge.org/) and BRACS (https://w.bracs.icar.cnr.it/download/). The pretrained weights of the baseline models are publicly available from Hugging Face, including PRISM (https://huggingface.co/paige-ai/Prism) and TITAN (https://huggingface.co/MahmoodLab/TITAN). The preprocessing, training, evaluation, statistical analysis, and risk-control code, together with the trained CorePath model weights, will be made publicly available (https://github.com/danninglee/CorePath) upon acceptance of the manuscript to support reproducibility. Qualified academic researchers with approved access to PRISM may reproduce and use the adaptation pipeline under the applicable license terms. B Supplementary Results B.1 Supplementary Figures (a) (b) (c) (d) Figure S5: Bland-Altman analysis of agreement between LLM-based Evaluation Scores and pathologist reference scores. Panels show the results for (a) all reports (N=144N=144), (b) reports generated by PRISM (N=48N=48), (c) reports generated by CorePath (N=48N=48), and (d) reports generated by CorePath-CRG (N=48N=48). For each report, the pathologist reference score was defined as the mean of three repeated assessments. The horizontal axis represents the mean of the LLM-based Evaluation Score and the pathologist reference score, whereas the vertical axis represents their difference, calculated as the LLM-based Evaluation Score minus the pathologist reference score. The solid horizontal line indicates the mean difference (bias), and the upper and lower dashed horizontal lines indicate the upper and lower 95% limits of agreement, respectively, calculated as the bias ±1.96± 1.96 times the standard deviation of the paired differences. The bias and 95% limits of agreement are annotated in each panel. LLM, large language model; LoA, limits of agreement; N, the number of reports. The results show that the LLM-based Evaluation Scores align with pathologist reference scores but are generally more conservative. Figure S6: Supplementary report-generation examples across different models. For each case, the left panels show the whole-slide image and a magnified region of interest, and the green panel summarizes the reference diagnosis. Model-generated outputs are shown below each case. Red text indicates an incorrect or unsupported diagnosis. Scale bars: from top to bottom, 250μ , 200μ , and 200μ . B.2 Supplementary Tables Table S1: Information on different tasks about the SPH-2 breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 367/245 Tumor category classification Noncancer/in situ/invasive 245/35/332 Five-class histological subtype Noncancer/ILC/IDC/other IBC/in situ 245/20/282/30/35 Total 612 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma; ILC, invasive lobular carcinoma; IDC, invasive ductal carcinoma; other IBC, other invasive breast carcinoma. Table S2: Information on different tasks about the SWH breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 585/807 Tumor category classification Noncancer/in situ/invasive 807/44/541 Five-class histological subtype Noncancer/ILC/IDC/other IBC/in situ 807/8/508/25/44 Total 1392 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma; ILC, invasive lobular carcinoma; IDC, invasive ductal carcinoma; other IBC, other invasive breast carcinoma. Table S3: Information on different tasks about the WCH-2 breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 336/160 Tumor category classification Noncancer/in situ/invasive 160/48/288 Five-class histological subtype Noncancer/ILC/IDC/other IBC/in situ 160/11/270/7/48 Total 496 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma; ILC, invasive lobular carcinoma; IDC, invasive ductal carcinoma; other IBC, other invasive breast carcinoma. Table S4: Information on different tasks about the WTH breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 207/75 Tumor category classification Noncancer/in situ/invasive 75/17/190 Five-class histological subtype Noncancer/ILC/IDC/other IBC/in situ 75/4/185/1/17 Total 282 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma; ILC, invasive lobular carcinoma; IDC, invasive ductal carcinoma; other IBC, other invasive breast carcinoma. Table S5: Information on different tasks about the SJH breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 76/52 Tumor category classification Noncancer/in situ/invasive 52/6/70 Five-class histological subtype Noncancer/ILC/IDC/other IBC/in situ 52/2/67/1/6 Total 128 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma; ILC, invasive lobular carcinoma; IDC, invasive ductal carcinoma; other IBC, other invasive breast carcinoma. Table S6: Information on different tasks about the SZH breast biopsy dataset. Task Categories Slides per class Cancer detection Cancer/noncancer 75/60 Tumor category classification Noncancer/in situ/invasive 60/3/72 Total 135 Note. In situ, carcinoma in situ; invasive, invasive breast carcinoma. Table S7: Prompt class names used for cancer detection. Task Class Class names Cancer detection Noncancer noncancer benign Cancer cancer carcinoma Table S8: Prompt class names used for invasion assessment. Task Class Class names Invasion assessment Noncancer noncancer benign Invasive breast carcinoma invasive lobular carcinoma breast invasive lobular carcinoma invasive lobular carcinoma of the breast invasive carcinoma of the breast, lobular pattern breast ILC invasive ductal carcinoma breast invasive ductal carcinoma invasive ductal carcinoma of the breast invasive carcinoma of the breast, ductal pattern breast IDC tubular carcinoma cribriform carcinoma mucinous carcinoma mucinous cystadenocarcinoma invasive micropapillary carcinoma carcinoma with apocrine differentiation metaplastic carcinoma rare and salivary gland-type tumours neuroendocrine neoplasms neuroendocrine tumour neuroendocrine carcinoma In situ carcinoma lobular carcinoma in situ ductal carcinoma in situ Table S9: Prompt class names used for histological subtyping. Task Class Class names Histological subtyping Noncancer noncancer benign In situ carcinoma lobular carcinoma in situ ductal carcinoma in situ Invasive breast carcinoma NOS invasive ductal carcinoma breast invasive ductal carcinoma invasive ductal carcinoma of the breast invasive carcinoma of the breast, ductal pattern breast IDC Invasive lobular carcinoma invasive lobular carcinoma breast invasive lobular carcinoma invasive lobular carcinoma of the breast invasive carcinoma of the breast, lobular pattern breast ILC Other invasive breast carcinoma tubular carcinoma cribriform carcinoma mucinous carcinoma mucinous cystadenocarcinoma invasive micropapillary carcinoma carcinoma with apocrine differentiation metaplastic carcinoma rare and salivary gland-type tumours neuroendocrine neoplasms neuroendocrine tumour neuroendocrine carcinoma Table S10: Prompt class names used for invasive carcinoma subtype categorization of BCNB. Task Class Class names Invasive carcinoma subtype categorization Invasive breast carcinoma NOS invasive ductal carcinoma breast invasive ductal carcinoma invasive ductal carcinoma of the breast invasive carcinoma of the breast, ductal pattern breast IDC Invasive lobular carcinoma invasive lobular carcinoma breast invasive lobular carcinoma invasive lobular carcinoma of the breast invasive carcinoma of the breast, lobular pattern breast ILC Other invasive breast carcinoma tubular carcinoma cribriform carcinoma mucinous carcinoma mucinous cystadenocarcinoma invasive micropapillary carcinoma carcinoma with apocrine differentiation metaplastic carcinoma rare and salivary gland-type tumours neuroendocrine neoplasms neuroendocrine tumour neuroendocrine carcinoma Table S11: Prompt class names used for lesion stratification of BRACS. Task Class Class names Lesion stratification Benign tumors benign breast tumor benign breast lesion benign breast neoplasm non-malignant breast tumor Atypical tumors atypical breast lesion breast atypia atypical proliferative lesion of breast premalignant breast lesion Malignant tumors malignant breast tumor breast cancer breast carcinoma breast malignancy malignant breast neoplasm Table S12: Prompt class names used for fine-grained classification of BRACS. Task Class Class names Fine-grained classification Normal normal breast tissue normal mammary tissue healthy breast tissue non-lesional breast tissue Pathological benign benign breast lesion benign breast pathology benign breast disease non-malignant breast lesion Usual ductal hyperplasia usual ductal hyperplasia benign ductal hyperplasia usual epithelial hyperplasia proliferative lesion without atypia UDH Flat epithelial atypia flat epithelial atypia flat epithelial lesion with atypia columnar cell lesion with atypia atypical columnar cell change FEA Atypical ductal hyperplasia atypical ductal hyperplasia ductal hyperplasia with atypia atypical intraductal proliferation atypical hyperplasia of the breast atypical ductal proliferation ADH Ductal carcinoma in situ ductal carcinoma in situ in situ ductal carcinoma intraductal carcinoma non-invasive ductal carcinoma ductal carcinoma (in situ) DCIS Invasive carcinoma invasive breast carcinoma invasive breast carcinoma invasive breast cancer infiltrating breast carcinoma invasive malignant breast tumor carcinoma of the breast, invasive Table S13: Zero-shot cancer detection performance on independent private datasets. Values are original test-set point estimates with 95% bootstrap confidence intervals in parentheses. Bold indicates the best result per dataset per metric. Dataset Model Accuracy Weighted F1 Weighted AUC SPH-2 PRISM 0.8758 (0.8480-0.9020) 0.8717 (0.8426-0.8989) 0.9572 (0.9384-0.9737) TITAN 0.9183 (0.8954-0.9395) 0.9173 (0.8940-0.9390) 0.9669 (0.9495-0.9810) CorePath 0.9330 (0.9134-0.9526) 0.9322 (0.9119-0.9521) 0.9783 (0.9653-0.9891) SWH PRISM 0.8556 (0.8369-0.8736) 0.8561 (0.8374-0.8741) 0.9756 (0.9682-0.9827) TITAN 0.9016 (0.8865-0.9167) 0.9022 (0.8870-0.9170) 0.9886 (0.9839-0.9929) CorePath 0.9555 (0.9447-0.9655) 0.9557 (0.9450-0.9657) 0.9948 (0.9909-0.9977) WCH-2 PRISM 0.9133 (0.8851-0.9375) 0.9112 (0.8817-0.9362) 0.9701 (0.9538-0.9830) TITAN 0.9254 (0.9012-0.9476) 0.9247 (0.8993-0.9471) 0.9734 (0.9594-0.9838) CorePath 0.9536 (0.9335-0.9718) 0.9532 (0.9327-0.9716) 0.9912 (0.9825-0.9971) WTH PRISM 0.9291 (0.9007-0.9574) 0.9262 (0.8942-0.9564) 0.9665 (0.9408-0.9874) TITAN 0.9326 (0.9007-0.9610) 0.9301 (0.8956-0.9597) 0.9858 (0.9714-0.9954) CorePath 0.9539 (0.9291-0.9752) 0.9527 (0.9260-0.9749) 0.9907 (0.9821-0.9968) SJH PRISM 0.8594 (0.7969-0.9141) 0.8570 (0.7905-0.9130) 0.9370 (0.8840-0.9762) TITAN 0.9375 (0.8982-0.9766) 0.9371 (0.8956-0.9765) 0.9603 (0.9167-0.9965) CorePath 0.9297 (0.8826-0.9688) 0.9293 (0.8789-0.9687) 0.9669 (0.9271-0.9931) SZH PRISM 0.8593 (0.7926-0.9111) 0.8539 (0.7850-0.9091) 0.9942 (0.9843-0.9998) TITAN 0.9630 (0.9259-0.9926) 0.9629 (0.9254-0.9926) 0.9891 (0.9668-1.0000) CorePath 0.9556 (0.9111-0.9852) 0.9552 (0.9115-0.9852) 0.9989 (0.9956-1.0000) Table S14: Zero-shot invasion assessment performance on independent private datasets. Values are original test-set point estimates with 95% bootstrap confidence intervals in parentheses. Bold indicates the best result per dataset per metric. Dataset Model Accuracy Weighted F1 Weighted AUC SPH-2 PRISM 0.7010 (0.6634-0.7369) 0.6798 (0.6354-0.7192) 0.8705 (0.8448-0.8947) TITAN 0.5654 (0.5261-0.6046) 0.5004 (0.4568-0.5447) 0.9635 (0.9481-0.9753) CorePath 0.8268 (0.7941-0.8578) 0.8234 (0.7892-0.8552) 0.9643 (0.9487-0.9775) SWH PRISM 0.5999 (0.5754-0.6279) 0.6061 (0.5799-0.6365) 0.8425 (0.8219-0.8601) TITAN 0.3886 (0.3635-0.4145) 0.3249 (0.2989-0.3526) 0.9745 (0.9668-0.9815) CorePath 0.8355 (0.8168-0.8542) 0.8397 (0.8210-0.8582) 0.9860 (0.9800-0.9908) WCH-2 PRISM 0.7460 (0.7056-0.7863) 0.7371 (0.6948-0.7787) 0.8683 (0.8377-0.8954) TITAN 0.6089 (0.5645-0.6512) 0.5371 (0.4897-0.5838) 0.9405 (0.9203-0.9594) CorePath 0.8710 (0.8407-0.8992) 0.8660 (0.8330-0.8953) 0.9793 (0.9680-0.9878) WTH PRISM 0.7730 (0.7234-0.8192) 0.7576 (0.7014-0.8103) 0.8758 (0.8354-0.9116) TITAN 0.7021 (0.6454-0.7589) 0.6443 (0.5818-0.7082) 0.9666 (0.9415-0.9856) CorePath 0.8972 (0.8617-0.9291) 0.8924 (0.8521-0.9271) 0.9840 (0.9703-0.9944) SJH PRISM 0.6797 (0.5938-0.7578) 0.6671 (0.5748-0.7530) 0.7718 (0.6892-0.8482) TITAN 0.5703 (0.4844-0.6563) 0.4894 (0.3957-0.5859) 0.9681 (0.9339-0.9928) CorePath 0.8125 (0.7500-0.8750) 0.8067 (0.7347-0.8760) 0.9804 (0.9561-0.9949) SZH PRISM 0.5704 (0.4815-0.6519) 0.5419 (0.4434-0.6337) 0.6866 (0.5992-0.7658) TITAN 0.1259 (0.0741-0.1852) 0.1658 (0.0976-0.2431) 0.6138 (0.5318-0.6923) CorePath 0.8074 (0.7333-0.8741) 0.7908 (0.7078-0.8604) 0.9881 (0.9740-0.9969) Table S15: Zero-shot histological subtyping performance on independent private datasets. Values are original test-set point estimates with 95% bootstrap confidence intervals in parentheses. Bold indicates the best result per dataset per metric. Dataset Model Accuracy Weighted F1 Weighted AUC SPH-2 PRISM 0.5343 (0.4918-0.5719) 0.5379 (0.4925-0.5797) 0.8412 (0.8160-0.8644) TITAN 0.4706 (0.4314-0.5114) 0.3694 (0.3267-0.4144) 0.9312 (0.9152-0.9444) CorePath 0.7337 (0.6993-0.7680) 0.7285 (0.6904-0.7644) 0.9557 (0.9412-0.9695) SWH PRISM 0.3384 (0.3147-0.3628) 0.4124 (0.3845-0.4405) 0.7299 (0.7104-0.7477) TITAN 0.3534 (0.3290-0.3779) 0.2676 (0.2418-0.2937) 0.9573 (0.9480-0.9660) CorePath 0.7636 (0.7407-0.7852) 0.7834 (0.7599-0.8030) 0.9735 (0.9658-0.9801) WCH-2 PRISM 0.4960 (0.4516-0.5383) 0.5379 (0.4894-0.5829) 0.7338 (0.6953-0.7663) TITAN 0.5202 (0.4738-0.5625) 0.4315 (0.3799-0.4776) 0.9029 (0.8767-0.9257) CorePath 0.7621 (0.7218-0.7984) 0.7405 (0.6959-0.7819) 0.9526 (0.9361-0.9675) WTH PRISM 0.4858 (0.4291-0.5426) 0.5834 (0.5217-0.6416) 0.7487 (0.6982-0.7906) TITAN 0.6525 (0.5957-0.7163) 0.5751 (0.5081-0.6482) 0.9364 (0.9066-0.9610) CorePath 0.8511 (0.8085-0.8901) 0.8531 (0.8046-0.8949) 0.9704 (0.9511-0.9848) SJH PRISM 0.4766 (0.3906-0.5703) 0.5763 (0.4845-0.6629) 0.7257 (0.6592-0.7892) TITAN 0.5078 (0.4219-0.6016) 0.3993 (0.3086-0.5020) 0.9377 (0.8979-0.9696) CorePath 0.7656 (0.6875-0.8359) 0.7545 (0.6720-0.8307) 0.9613 (0.9154-0.9905) Table S16: Zero-shot classification performance on public benchmark datasets. Values are original test-set point estimates with 95% bootstrap confidence intervals in parentheses. Bold indicates the best result per dataset per metric. Dataset Task Model Accuracy Weighted F1 Weighted AUC BCNB Invasive carcinoma subtyping PRISM 0.2760 (0.2476-0.3034) 0.3562 (0.3227-0.3881) 0.5748 (0.5214-0.6304) TITAN 0.7968 (0.7741-0.8214) 0.8280 (0.8064-0.8497) 0.7666 (0.7141-0.8182) CorePath 0.8989 (0.8800-0.9159) 0.8919 (0.8698-0.9114) 0.7780 (0.7279-0.8257) BRACS Lesion stratification PRISM 0.5868 (0.5448-0.6271) 0.6313 (0.5918-0.6684) 0.8016 (0.7700-0.8295) TITAN 0.2687 (0.2321-0.3053) 0.2164 (0.1788-0.2507) 0.7937 (0.7586-0.8237) CorePath 0.6362 (0.5941-0.6746) 0.6772 (0.6400-0.7111) 0.8178 (0.7856-0.8458) BRACS Fine-grained classification PRISM 0.4168 (0.3766-0.4570) 0.4268 (0.3848-0.4699) 0.8106 (0.7844-0.8347) TITAN 0.3821 (0.3400-0.4241) 0.3491 (0.3061-0.3919) 0.8121 (0.7908-0.8314) CorePath 0.4442 (0.4022-0.4863) 0.4602 (0.4192-0.5016) 0.8252 (0.8011-0.8485) Table S17: Intra-rater reliability of repeated pathologist report-quality assessments. Reliability measure N reports ICC 95% CI Single assessment, ICC(A,1) 144 0.980 0.97–0.99 Mean of three assessments, ICC(A,3) 144 0.993 0.99–1.00 Note. ICCs were estimated using a two-way mixed-effects, absolute-agreement model. ICC(A,1) denotes the reliability of a single pathologist assessment, whereas ICC(A,3) denotes the reliability of the mean of three repeated assessments. N, number of reports; CI, confidence interval. Table S18: Agreement between LLM-based Evaluation Scores and pathologist reference scores. Group N Spearman ρ ICC(A,1), 95% CI MAD Bias 95% LoA Overall 144 0.911 0.843 (0.74–0.90) 0.112 −0.083-0.083 −0.459-0.459 to 0.293 PRISM 48 0.918 0.828 (0.70–0.90) 0.109 −0.074-0.074 −0.467-0.467 to 0.318 CorePath 48 0.850 0.778 (0.58–0.88) 0.137 −0.105-0.105 −0.523-0.523 to 0.312 CorePath-CRG 48 0.907 0.887 (0.78–0.94) 0.090 −0.070-0.070 −0.385-0.385 to 0.245 Note. The pathologist reference score for each report was defined as the mean of three repeated assessments. ICC(A,1) denotes the two-way mixed-effects, absolute-agreement, single-measure ICC between LLM-based Evaluation Score and the corresponding pathologist reference score. Spearman ρ measures rank association; MAD, mean absolute difference; bias, mean signed difference (LLM-based Evaluation Score minus pathologist reference score); LoA, limits of agreement; N, number of reports; CI, confidence interval. Table S19: LLM-based Evaluation Score across different datasets (weighted composite score, 0-1). Mean and median values are reported to reflect both average performance and distributional stability. Best results are in bold. Dataset Model Mean Median SWH PRISM 0.335 0.150 CorePath 0.373 0.300 CorePath-CRG 0.532 0.450 WCH-2 PRISM 0.250 0.100 CorePath 0.319 0.250 CorePath-CRG 0.589 0.650 WTH PRISM 0.224 0.075 CorePath 0.328 0.200 CorePath-CRG 0.681 0.950 SJH PRISM 0.229 0.025 CorePath 0.277 0.125 CorePath-CRG 0.614 0.850 Table S20: Quantitative evaluation of PRISM, CorePath, and CorePath-CRG using automatic text-similarity metrics on the retained samples (excluding “Unknown” rejections) across four datasets. N denotes the number of valid samples after exclusion. Best results are in bold. Dataset Model B-1 B-2 B-3 B-4 R-1 R-2 R-L METEOR SWH (N=855) PRISM 0.1590 0.0469 0.0292 0.0225 0.1368 0.0105 0.1159 0.1497 CorePath 0.1893 0.0521 0.0311 0.0236 0.1873 0.0105 0.1394 0.1657 CorePath-CRG 0.1562 0.0727 0.0551 0.0411 0.3035 0.1235 0.2995 0.2146 WCH-2 (N=379) PRISM 0.1526 0.0607 0.0384 0.0272 0.1653 0.0338 0.1510 0.1988 CorePath 0.1716 0.0692 0.0435 0.0297 0.2040 0.0464 0.1829 0.2412 CorePath-CRG 0.2880 0.1166 0.0928 0.0726 0.3967 0.0693 0.3916 0.2579 WTH (N=228) PRISM 0.1644 0.0540 0.0325 0.0247 0.1399 0.0158 0.1334 0.2177 CorePath 0.1976 0.0658 0.0413 0.0291 0.2153 0.0324 0.2032 0.2778 CorePath-CRG 0.3550 0.1217 0.0911 0.0780 0.4984 0.0451 0.4971 0.2860 SJH (N=94) PRISM 0.1451 0.0533 0.0348 0.0256 0.1158 0.0084 0.1085 0.1792 CorePath 0.1581 0.0455 0.0283 0.0217 0.1598 0.0132 0.1447 0.2117 CorePath-CRG 0.3341 0.1240 0.0958 0.0784 0.4267 0.0621 0.4267 0.2727 Table S21: Post-conformal-filtering diagnostic label distribution across four centers. Values are percentages. Label SWH WCH-2 WTH SJH Binary Cancer Status Cancer 43.46% 66.53% 76.60% 57.03% Noncancer 55.17% 33.27% 23.40% 40.62% Unknown 1.36% 0.20% — 2.34% Histological Subtype Noncancer 18.39% 13.91% 8.51% 14.84% In situ carcinoma 0.36% 2.02% 1.77% — Invasive breast carcinoma NOS 42.17% 60.08% 70.21% 58.59% Invasive lobular carcinoma 0.07% 0.20% 0.35% — Other invasive breast carcinoma 0.43% 0.20% — — Unknown 38.58% 23.59% 19.15% 26.56% Note. The table summarizes the distribution of diagnostic labels after conformal filtering. “Unknown” denotes predictions that did not satisfy the conformal confidence criterion and were therefore not retained as confident diagnostic evidence. NOS, not otherwise specified. Table S22: CorePath-CRG final output statistics across four centers. Values are percentages. Metric SWH WCH-2 WTH SJH Full Rejection Rate 38.58% 23.59% 19.15% 26.56% Narrative Rejection Rate 81.03% 78.43% 82.27% 83.59% Full Rejection ∣ Narrative Rejected 47.61% 30.08% 23.28% 31.78% Note. Full Rejection Rate denotes the proportion of cases in which both the narrative report and subtype prediction were rejected, requiring pathologist review. Narrative Rejection Rate denotes the proportion of cases in which the CorePath-generated narrative report failed the release criterion. Full Rejection ∣ Narrative Rejected denotes the conditional proportion of full rejection among cases with rejected narrative reports, indicating how often the subtype prediction was also rejected when the narrative failed.