Paper deep dive
Multimodal MRI Report Findings Supervised Brain Lesion Segmentation with Substructures
Yubin Ge, Yongsong Huang, Xiaofeng Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 1:33:30 PM
Summary
The paper introduces MS-RSuper, a report-supervised learning framework for brain lesion segmentation in multimodal MRI. It addresses limitations of classical report-supervised methods by explicitly parsing global quantitative and modality-wise qualitative findings. The framework employs existence/absence losses for modality-specific substructures (e.g., T1c enhancement for Enhancing Tumor), one-sided lower-bound losses for partial quantitative cues (e.g., largest lesion size), and anatomical priors to distinguish between meningioma (extra-axial) and metastases (intra-axial). Evaluated on 1238 BraTS-MET/MEN scans, it outperforms baselines.
Entities (10)
Relation Signals (9)
MS-RSuper → uses → Report-supervised learning
confidence 95% · We introduce a unified, one-sided, uncertainty-aware formulation (MS-RSuper) ... Report-supervised (RSuper) learning seeks to alleviate the need for dense tumor voxel labels
MS-RSuper → evaluatedon → BraTS-MEN
confidence 92% · We validated its effectiveness on the combined BraTS-MET and BraTS-MEN
MS-RSuper → evaluatedon → BraTS-MET
confidence 92% · We validated its effectiveness on the combined BraTS-MET and BraTS-MEN
FLAIR → constrains → Edema
confidence 90% · FLAIR findings (edema, non-specific hyperintensity) constrain the Edema (ED) probability map.
T1c → constrains → Enhancing Tumor
confidence 90% · T1c findings (enhancement) constrain the Enhancing Tumor (ET) probability map.
BraTS-MEN → contains → Meningioma
confidence 90% · BraTS-MEN (Meningioma): A collection of 1000 subjects... with meningioma.
BraTS-MET → contains → Metastases
confidence 90% · BraTS-MET (Metastases): A collection of 238 subjects... with brain metastases.
Llama 3.1 70B → usedby → MS-RSuper
confidence 85% · The LLM parser was implemented using Llama 3.1 70B
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Report-supervised (RSuper) learning seeks to alleviate the need for dense tumor voxel labels with constraints derived from radiology reports (e.g., volumes, counts, sizes, locations). In MRI studies of brain tumors, however, we often involve multi-parametric scans and substructures. Here, fine-grained modality/parameter-wise reports are usually provided along with global findings and are correlated with different substructures. Moreover, the reports often describe only the largest lesion and provide qualitative or uncertain cues (``mild,'' ``possible''). Classical RSuper losses (e.g., sum volume consistency) can over-constrain or hallucinate unreported findings under such incompleteness, and are unable to utilize these hierarchical findings or exploit the priors of varied lesion types in a merged dataset. We explicitly parse the global quantitative and modality-wise qualitative findings and introduce a unified, one-sided, uncertainty-aware formulation (MS-RSuper) that: (i) aligns modality-specific qualitative cues (e.g., T1c enhancement, FLAIR edema) with their corresponding substructures using existence and absence losses; (ii) enforces one-sided lower-bounds for partial quantitative cues (e.g., largest lesion size, minimal multiplicity); and (iii) adds extra- vs. intra-axial anatomical priors to respect cohort differences. Certainty tokens scale penalties; missing cues are down-weighted. On 1238 report-labeled BraTS-MET/MEN scans, our MS-RSuper largely outperforms both a sparsely-supervised baseline and a naive RSuper method.
Tags
Links
- Source: https://arxiv.org/abs/2602.20994v1
- Canonical: https://arxiv.org/abs/2602.20994v1
Trouble viewing inline? Open PDF directly →
Full Text
19,267 characters extracted from source content.
Expand or collapse full text
Multimodal MRI Report Findings Supervised Brain Lesion Segmentation with Substructures Abstract Report-supervised (RSuper) learning seeks to alleviate the need for dense tumor voxel labels with constraints derived from radiology reports (e.g., volumes, counts, sizes, locations). In MRI studies of brain tumors, however, we often involve multi-parametric scans and substructures. Here, fine-grained modality/parameter-wise reports are usually provided along with global findings and are correlated with different substructures. Moreover, the reports often describe only the largest lesion and provide qualitative or uncertain cues (“mild,” “possible”). Classical RSuper losses (e.g., sum volume consistency) can over-constrain or hallucinate unreported findings under such incompleteness, and are unable to utilize these hierarchical findings or exploit the priors of varied lesion types in a merged dataset. We explicitly parse the global quantitative and modality-wise qualitative findings and introduce a unified, one-sided, uncertainty-aware formulation (MS-RSuper) that: (i) aligns modality-specific qualitative cues (e.g., T1c enhancement, FLAIR edema) with their corresponding substructures using existence and absence losses; (i) enforces one-sided lower-bounds for partial quantitative cues (e.g., largest lesion size, minimal multiplicity); and (i) adds extra- vs. intra-axial anatomical priors to respect cohort differences. Certainty tokens scale penalties; missing cues are down-weighted. On 1238 report-labeled BraTS-MET/MEN scans, our MS-RSuper largely outperforms both a sparsely-supervised baseline and a naive RSuper method. †footnotetext: † contribute equally. ∗corresponding author: xiaofeng.liu@yale.edu Index Terms— Report supervision, multimodal MRI, meningioma, brain metastases, segmentation. 1 Introduction Accurate delineation of lesion structures is a fundamental step in clinical diagnosis, intervention, and treatment planning [liu2023incremental]. However, voxel-wise annotation for 3D multimodal MRI is costly and subjective, especially when substructures of tumor core (TC), enhancing tumor (ET), and edema (ED) must be delineated across sequences with different contrast mechanisms [liu2023incremental]. How to utilize existing, routinely summarized radiology reports to aid (or assist) segmentation model training is of great importance for more practical utilization of big medical data. To exploit the valuable and relatively large-scale text information, early attempts usually form a multi-task learning with an auxiliary task in addition to the segmentation with either the extracted tumor present/absent label for classification [zhang20213d], or a contrastive language-image pre-training objective [blankemeier2024merlin]. While the benefits of auxiliary tasks to segmentation are indirect and occasionally minor, the recently developed Report-supervised learning (RSuper) [bassi2025learning] promises to directly leverage the detailed volumes, counts, and sizes extracted from abundant abdominal CT radiology reports and enforces the corresponding loss functions alongside the conventional segmentation loss for only a small portion of segmentation labeled samples (e.g., 50 scans). This approach therefore largely reduces the burden of manual labeling. However, in MRI studies of brain tumors, we often involve multi-parametric scans of T1, T1c, T2, and FLAIR as well as substructures of TC, ET, and ED. In which the fine-grained modality/ parameter-wise finding reports are usually provided along with the global finding reports, and correlated to different substructures. How to systematically exploit both the global descriptors (falx-/skull-base adjacency vs. deep parenchyma, approximate size, multiplicity, edema/midline shift) and modality-wise descriptors (T1c enhancement pattern; FLAIR hyperintensity) is largely underexplored. Moreover, the reports often describe only the largest lesion, provide qualitative or uncertain cues (“mild,” “possible”). Specifically, in brain lesions, reports are often partially specified: many cases provide only the largest lesion size dmaxd_ , omit axes for diameters, and/or use certainty qualifiers (mild, possible, equivocal); and the counts are sometimes qualitative (“multiple”). Simply adopt the classical RSuper volume loss, which penalizes the difference of predicted and labeled sum volume of all tumors, may (a) learn to suppress small lesions that the report does not enumerate, or (b) unduly shrink tumors to match a partial volume hint. Finally, when combining image-report records from multiple diseases, cohort-specific priors are lost. For example, BraTS-MET (metastases) typically includes multiple intra-axial parenchymal lesions, often with ring enhancement. In contrast, BraTS-MEN (meningioma) is typically extra-axial and dural-based (e.g., falx, skull base) with solid enhancement. A report-derived supervision should transfer across cohorts without imposing contradictory biases. A naive RSuper loss cannot leverage these strong anatomical priors. To address these limitations, we propose a novel multimodality with substructure RSuper framework for brain lesion segmentation. Our main contributions are: ∙ -Substructure Alignment: We introduce a loss that links modality-specific report findings (e.g., T1c enhancement, FLAIR edema) directly to their corresponding segmentation substructures (ET and ED, respectively). ∙ -Sided Partial-Report Loss: We propose a ”lower-bound” size loss and ”minimal-multiplicity” count loss to handle incomplete reports that only describe the largest lesion or use qualitative counts, avoiding penalty for valid, unreported lesions. ∙ -Specific Priors: We integrate an anatomical prior loss that penalizes intra-axial predictions for MEN and extra-axial predictions for MET, guided by cohort-level cues from the reports. We validated its effectiveness on the combined BraTS-MET and BraTS-MEN with segmentation and report111https://huggingface.co/datasets/JiayuLei/RadGenome-Brain_MRI/tree/main. 2 Methodology Our framework trains a 3D segmentation network using a partially segmented dataset, i.e., a large set of image-report pairs (DRD_R) and a small set of fully-masked data (DMD_M). The model is trained with a composite loss. For data in DMD_M, we use a standard supervised segmentation loss, LsegL_seg (e.g., a combination of Dice and Cross-Entropy loss). For data in DRD_R, we introduce a novel report-supervised loss, LreportL_report, designed to handle the hierarchical, qualitative, and partial nature of multimodal MRI reports. 2.1 Hierarchical Report Parsing and Mapping We first employ a Large Language Model (LLM) with domain-specific prompts to parse each free-text radiology report. Critically, we categorize the extracted cues into two distinct types based on their nature and scope: (A) Quantitative Global Cues: These are specific measurements, typically found in the ”global findings” section, that apply to the entire lesion or provide a total count. These are often partial (e.g., “largest lesion measuring 45x39x47 m,” ”multiple punctate… lesions”). (B) Qualitative Modality-Specific Cues: These are usually descriptive, non-numeric findings tied to a specific MRI sequence, which inherently map to tumor substructures. Though sometimes the tumor size is provided, it is the same as the global finding, which does not provide incremental information. For example, T1c often includes “obvious enhancement,” “ring enhancement,” “no enhancement,” while FLAIR includes “surrounding extensive edema,” “mild hyperintense signal.” Based on this parsing, we propose to establish a Modality-Substructure Alignment Principle. This is not a loss function itself, but a crucial mapping rule that directs how constraints are applied: ∙ T1c findings (enhancement) constrain the Enhancing Tumor (PETP_ET) probability map. ∙ FLAIR findings (edema, non-specific hyperintensity) constrain the Edema (PEDP_ED) probability map. ∙ T1 or T2 findings (e.g., ”hypointense core”) constrain the Tumor Core (PTCP_TC) map. ∙ Global cues (e.g., total size, count) constrain the Whole Tumor (PWTP_WT) map, where PWT=PET+PED+PTCP_WT=P_ET+P_ED+P_TC. Uncertainty cues (e.g., ”possible,” ”mild”) are parsed into a scaling weight λ∈[0,1]λ∈[0,1] for the corresponding loss term. Fig. 1: Overview of our proposed report-supervised framework. An LLM parses hierarchical findings from reports. Our losses align modality-specific findings, handle partial cues, and enforce anatomical location priors. 2.2 Unified Report Constraint Loss (ℒreportL_report) Our primary report loss, ℒreportL_report, combines constraints from both qualitative and quantitative cues, applying them to the aligned substructure maps identified in §2.1 2.1. 2.2.1 Substructure Qualitative Existence and Absence Loss As identified, most modality-specific cues are qualitative (e.g., ”edema is present”) and lack quantitative volumes. We cannot use a symmetric L1/L2 volume loss [bassi2025learning]. Instead, we formulate a loss based on the existence or absence of a finding. For a given substructure class k (e.g., k∈ET, ED, TCk∈\ET, ED, TC\), let VkV_k be the predicted volume of Pk()≥0.5P_k(x)≥ 0.5 for that substructure. If the report confirms the presence of substructure k (e.g., ”surrounding edema” →k=ED→ k=ED) with confidence λk,pos _k,pos, we apply an ”Existence Loss.” This loss penalizes the model only if it fails to predict any presence (volume <1<1) of that substructure ℒexist(k)=max(0,1−Vk),L_exist^(k)= (0,1-V_k), which encourages the model to segment at least 1 voxel for class k, without hallucinating a specific target volume. Conversely, if the report explicitly confirms the absence of a substructure k (e.g., ”no enhancement” →k=ET→ k=ET), we apply an loss that penalizes any prediction for that class ℒexist(k)=Vk.L_exist^(k)=V_k. Therefore, we have ℒexist(k)=max(0,1−Vk)if presence confirmed,Vkif absence confirmed,0otherwise.L_exist^(k)= cases (0,1-V_k)&if presence confirmed,\\ V_k&if absence confirmed,\\ 0&otherwise. cases (1) 2.2.2 Global One-Sided Partial Cue Loss (Size and Count) We handle the quantitative but partial nature of global cues: ∙ Size Loss: Reports often provide only the 3D-dim or diameter(s) of the largest lesion, dmaxd_ . Let CpredC_pred be the set of predicted connected components for the whole tumor (PWT≥0.5P_WT≥ 0.5). Let dcd_c be the volume of a component c∈Cpredc∈ C_pred. The loss is: ℒsize=|dmax−maxc∈Cpreddc|.L_size=|d_ - _c∈ C_predd_c|. This loss used mean absolute error, which is more robust to small inaccurate measure of dmaxd_ . ∙ Count Loss: Reports often use qualitative counts like ”multiple” or ”a few.” We parse this to a minimal integer NqualN_qual (e.g., ”multiple” →Nqual=2→ N_qual=2). We apply a one-sided count loss: ℒcount=max(0,Nqual−|Cpred|).L_count= (0,N_qual-|C_pred|). It penalizes the model if it predicts fewer than NqualN_qual lesions. Therefore, we haveℒglobal=wsizeℒsize+wcountℒcountL_global=w_sizeL_size+w_countL_count. 2.3 Cohort-Specific Anatomical Prior Loss (ℒpriorL_prior) Finally, we leverage global location cues (e.g., “falx,” “parenchymal”) to identify the cohort (MEN or MET) and apply a strong anatomical prior. We use pre-defined binary masks for the dura/extra-axial space (MduralM_dural) and the brain parenchyma/ intra-axial space (MparenchM_parench). ∙ If the report suggests Meningioma (MEN), which is extra-axial, we penalize any intra-axial predictions: ℒprior=∑(PWT()⋅Mparench()).L_prior= _x(P_WT(x)· M_parench(x)). -3.0pt ∙ If the report suggests Metastases (MET), which are intra-axial, we penalize extra-axial predictions: ℒprior=∑(PWT()⋅Mdural()).L_prior= _x(P_WT(x)· M_dural(x)). -3.0pt This loss effectively guides the model to search in the correct anatomical compartment, resolving ambiguity and reducing false positives. 2.4 Total Loss Function The model is first pre-trained on DMD_M using LsegL_seg. It is then fine-tuned on the combined dataset DM∪DRD_M∪ D_R. For a mixed batch B=BM∪BRB=B_M∪ B_R, the total loss is: ℒtotal=1|BM|∑i∈BMℒseg(i)+wr|BR|∑j∈BRℒreport(j),L_total= 1|B_M| _i∈ B_ML_seg^(i)+ w_r|B_R| _j∈ B_RL_report^(j), where wrw_r is a balancing weight for the report-based supervision, and ℒreportL_report is the sum of our proposed constraint losses: ℒreport=∑k3ℒexist(k)+wsizeℒsize+wcountℒcount+wpriorℒprior, _report= _k^3L_exist^(k)+w_sizeL_size+w_countL_count+w_priorL_prior, where wsizew_size, wcountw_count, and wpriorw_prior are weights for each component of the report-supervised loss. 3 Experiments and Results We used two large-scale, multi-modal MRI segmentation datasets. Their associated radiology reports are manually generated in RadGenome-Brain_MRI dataset [lei2024autorg]1. Each subject has a global finding and four modality-wise findings. ∙ BraTS-MEN (Meningioma): A collection of 1000 subjects (4000 3D mpMRI scans) with meningioma. Reports frequently describe extra-axial, dural-based lesions (e.g., ”falx cerebri,” ”skull base,” ”cerebellopontine angle”) with ”marked,” ”uniform” enhancement on T1c. ∙ BraTS-MET (Metastases): A collection of 238 subjects (952 3D mpMRI scans) with brain metastases. Reports describe ”multiple,” ”parenchymal” (intra-axial) lesions, often with ”ring enhancement” and ”extensive surrounding edema” on FLAIR. We held out 50 MEN and 50 MET for testing, and remaining subjects for training (all with reports, while 50 MEN and 50 MET has segmentation masks). The LLM parser was implemented using Llama 3.1 70B as in [bassi2025learning] with prompts engineered to extract the hierarchical attributes, uncertainty weights, and cohort priors. We used the 3D nnU-Net framework as our base segmentation architecture due to its strong performance. The model was pre-trained on BraTS2018 for Glioblastoma (different tumor from MEN or MET) with the supervised CE loss for labeled substructures of TC, ET, and ED [liu2023incremental]. No report is available. Loss weights were set empirically as wr=0.2w_r=0.2, wsize=1.0w_size=1.0, wcount=0.5w_count=0.5, and wprior=0.2w_prior=0.2. We compare three methods: (1) Masks-Only (Baseline): fine-tuned only on the 100 labeled scans (DMD_M). (2) R-Super [bassi2025learning] finetuned on DM∪DRD_M∪ D_R, using the summed volume and count applied to the ”Whole Tumor” (WT = ET+ED+TC) prediction. (3) Ours MS-RSuper: multimodal with substructure supervised by ℒreportL_report. As shown in Table 1, our method largely outperforms both baselines. The Masks-Only model suffers from poor generalization, as expected from only 50 labels in each disease. The RSuper [bassi2025learning] baseline dose not provides improvement, since its symmetric “summed volume” loss is confused by the partial reports (e.g., only the largest lesion), leading to suboptimal performance. Our MS-RSuper achieves the highest Dice scores across all substructures and both cohorts. The gains are promising in the MET dataset, where our ℒcountL_count (handling “multiple” lesions) and our qualitative losses (ℒexistL_exist) (which align ’edema’ to PEDP_ED and ‘enhancement’ to PETP_ET) are critical. For the MEN dataset, ℒpriorL_prior (enforcing extra-axial location) was key to reducing false positives in the brain parenchyma. Table 1: Dice Score on the held-out test sets. Method Test Set WT (DSC) TC (DSC) ET (DSC) Masks-Only MEN 0.481 0.323 0.370 R-Super [bassi2025learning] MEN 0.452 0.301 0.353 MS-RSuper(Ours) MEN 0.554 0.428 0.489 Masks-Only MET 0.420 0.385 0.321 RSuper [bassi2025learning] MET 0.443 0.391 0.333 MS-RSuper(Ours) MET 0.529 0.494 0.452 An ablation study (Table 2) on the MET dataset confirms that each of our proposed loss components contributes to the final performance. The quantitative, partial-cue losses (ℒexistL_exist) provide the first major boost by handling size and count. Adding the qualitative modality-aligned losses (ℒglobalL_global) further improves performance by correctly using the T1c and FLAIR cues. Finally, the cohort prior (ℒpriorL_prior) provides an additional gain by penalizing anatomically implausible predictions. 4 Conclusion We introduced a novel report-supervised learning framework tailored for the complexities of multi-parametric brain MRI and substructure segmentation. Unlike prior work on CT, our method addresses three key challenges: (1) it aligns qualitative modality-specific findings with their corresponding segmentation substructures using novel existence and absence losses, (2) it uses one-sided, uncertainty-aware losses to robustly handle partial quantitative reports (e.g., “largest lesion only,” “multiple”), and (3) it integrates cohort-level anatomical priors (intra- vs. extra-axial) derived from report keywords. We evaluate on a large dataset of 1238 meningioma and metastases scans, our approach largely outperformed both a sparsely-supervised baseline and a naive application of existing RSuper methods. This work demonstrates that by designing losses that faithfully reflect the hierarchical and often-incomplete nature of radiology reports, we can effectively leverage large-scale text data to improve multi-class segmentation in multimodal imaging. Table 2: Ablation study on the BraTS-MET Test set (WT Dice). Each component of our proposed loss (ℒexistL_exist, ℒglobalL_global, ℒpriorL_prior) provides a cumulative benefit. Method WT (DSC) Masks-Only (Baseline) 0.420 + ℒexistL_exist (Partial size/count) 0.475 + ℒexist+ℒglobalL_exist+L_global (Adds qualitative) 0.513 + ℒexist+ℒglobal+ℒpriorL_exist+L_global+L_prior (Full MS-RSuper) 0.529 5 COMPLIANCE WITH ETHICAL STANDARDS This retrospective study used open-access human subject data; no additional ethical approval was required. 6 ACKNOWLEDGMENTS Supported in part by NIH R21EB034911 and NVIDIA Academic Grant Program. References