Paper deep dive
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation
Kaito Baba, Risa Kishikawa, Satoshi Kodera
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 3:21:56 AM
Summary
The paper introduces MARL-Rad, a multi-modal multi-agent reinforcement learning framework for automated radiology report generation. It decomposes chest X-ray interpretation into region-specific agents (left, right, central) and a global integrating agent, jointly optimizing them using clinically verifiable rewards (RadGraph, CheXbert, GREEN). Experiments on MIMIC-CXR and IU X-ray datasets demonstrate state-of-the-art performance in clinical efficacy metrics compared to single-agent and training-free baselines.
Entities (12)
Relation Signals (14)
MARL-Rad â evaluatedon â IU X-ray
confidence 98% ¡ Experiments on the MIMIC-CXR and IU X-ray datasets
MARL-Rad â evaluatedon â MIMIC-CXR
confidence 98% ¡ Experiments on the MIMIC-CXR and IU X-ray datasets
MARL-Rad â consistsof â Right-Region Agent
confidence 95% ¡ Our system consists of three region-specific agents... a right-region agent
MARL-Rad â consistsof â Global Integrating Agent
confidence 95% ¡ a global integrating agent that synthesizes their outputs
MARL-Rad â consistsof â Central-Region Agent
confidence 95% ¡ Our system consists of three region-specific agents... a central-region agent
MARL-Rad â consistsof â Left-Region Agent
confidence 95% ¡ Our system consists of three region-specific agents... a left-region agent
MARL-Rad â uses â MedGemma
confidence 95% ¡ We use MedGemma-4B as the base model for all agents
MARL-Rad â outperforms â MedGemma (agent)
confidence 92% ¡ This variant performs worse than the vanilla MedGemma model... MARL-Rad... achieves higher scores
â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework for radiology report generation that trains the entire agentic system on policy within its deployed radiology workflow. MARL-Rad addresses the limitation of post-hoc agentization, where fixed LLMs are organized into hand-designed agentic workflows without being optimized for their assigned roles. Our framework decomposes chest X-ray interpretation into region-specific agents and a global integrating agent, and jointly optimizes them using clinically verifiable rewards. Experiments on the MIMIC-CXR and IU X-ray datasets show that MARL-Rad consistently improves clinical efficacy metrics such as RadGraph, CheXbert, and GREEN scores, achieving state-of-the-art clinical efficacy performance. Further analyses show that MARL-Rad improves laterality consistency and produces more accurate and detailed reports. A blinded clinician evaluation further suggests that MARL-Rad produces reports clinically comparable to ground-truth reports.
Tags
Links
- Source: https://arxiv.org/abs/2603.16876v2
- Canonical: https://arxiv.org/abs/2603.16876v2
Trouble viewing inline? Open PDF directly â
Full Text
92,606 characters extracted from source content.
Expand or collapse full text
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Kaito Baba &Risa Kishikawa &Satoshi Kodera & Department of Cardiovascular Medicine The University of Tokyo Hospital, Tokyo, Japan baba-kaito662@g.ecc.u-tokyo.ac.jp Abstract We propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework for radiology report generation that trains the entire agentic system on policy within its deployed radiology workflow. MARL-Rad addresses the limitation of post-hoc agentization, where fixed LLMs are organized into hand-designed agentic workflows without being optimized for their assigned roles. Our framework decomposes chest X-ray interpretation into region-specific agents and a global integrating agent, and jointly optimizes them using clinically verifiable rewards. Experiments on the MIMIC-CXR and IU X-ray datasets show that MARL-Rad consistently improves clinical efficacy metrics such as RadGraph, CheXbert, and GREEN scores, achieving state-of-the-art clinical efficacy performance. Further analyses show that MARL-Rad improves laterality consistency and produces more accurate and detailed reports. A blinded clinician evaluation further suggests that MARL-Rad produces reports clinically comparable to ground-truth reports. 1 Introduction Recent advances in large language models (LLMs) and large visionâlanguage models (LVLMs) have demonstrated remarkable reasoning and generation capabilities across a wide range of domains, including dialogue systems, mathematical reasoning, code synthesis, scientific discovery, and medical diagnosis [23, 86, 31, 123]. To further translate these capabilities into complex real-world tasks, agentic systems have emerged as an increasingly important paradigm for leveraging LLMs beyond single-turn generation. By decomposing tasks into manageable subtasks, agentic systems can support structured reasoning and workflow-oriented problem solving [119, 33, 37, 57, 137, 4, 92, 52, 44, 59]. Medicine is one of the most important application domains of LLMs and LVLMs. Among medical modalities, chest X-ray (CXR) is one of the most widely used diagnostic tools in clinical practice, enabling physicians to assess thoracic structures and identify conditions such as pneumonia, pneumothorax, and pleural effusion [81, 11, 12, 2]. Radiologists carefully inspect CXR images and manually compose diagnostic reports based on their interpretations. However, radiology reporting is time-consuming and labor-intensive, and the growing demand for diagnostic imaging can delay diagnosis and degrade report quality [21, 10, 95, 26]. To alleviate this burden, automated radiology report generation (RRG) using LLMs has attracted growing attention as a promising approach [70, 90, 140, 122, 55, 1, 115, 24, 61, 22, 112, 100, 5, 121], with recent studies increasingly exploring agentic-system designs to improve performance [133, 78, 28, 126, 50]. Despite these recent developments, most existing agentic systems do not train the underlying LLMs for their deployed workflows. They typically keep pretrained or domain-adapted LLMs fixed and rely on prompts and hand-designed workflows to elicit role-specific behavior [137, 92, 52, 44, 59]. As a result, role specialization and coordination are imposed only through prompt context: the workflow expects agents to act as specialized and coordinated components, while the underlying LLMs remain fixed policies that have not been optimized for their positions in the workflow. This mismatch is particularly consequential in real-world tasks that require specialized expertise, such as medicine. In such settings, domain-specific fine-tuning alone is insufficient; the full agent system must be trained on policy within the deployed workflow. Indeed, in SectionË4.3, we show that even MedGemma [94], a medical adaptation of Gemma 3 [106], performs poorly when its parameters are kept fixed and it is merely organized into a multi-agent workflow for RRG. To address this limitation, we propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework that trains the entire agentic RRG system end-to-end on policy within its deployed workflow. MARL-Rad consists of region-specific agents responsible for localized observations and a global integrating agent that synthesizes their outputs. These agents are jointly trained with clinically verifiable rewards, optimizing the agentic system within its deployed workflow. This workflow mirrors the real practice of radiologists, who meticulously examine each anatomical region before composing a comprehensive diagnostic report. Experiments on the MIMIC-CXR and IU X-ray datasets demonstrate that MARL-Rad consistently improves clinical efficacy (CE) metrics such as RadGraph F1 [43], CheXbert F1 [101], and GREEN scores [87], achieving state-of-the-art performance. Moreover, deeper analyses show that MARL-Rad improves laterality consistency and produces more detailed and clinically accurate descriptions. Finally, a small blinded clinician evaluation suggests that MARL-Rad produces clinically comparable reports to ground-truth reports. Our key contributions are summarized as follows: ⢠Workflow-aligned training beyond fixed agentification: We show that simply organizing a fixed LLM into an agentic workflow is insufficient and can even degrade performance. MARL-Rad addresses this limitation by training the entire agentic system on policy within the same radiology-inspired workflow in which it is deployed. ⢠State-of-the-art performance on CE metrics: Experiments on the MIMIC-CXR and IU X-ray datasets demonstrate that MARL-Rad achieves state-of-the-art performance on various CE metrics, including RadGraph F1, CheXbert F1, and GREEN scores. ⢠Enhanced laterality consistency and accurate, detail-informed reports: MARL-Rad improves laterality consistency and generates more detailed and clinically accurate reports compared to single-agent RL baselines. 2 Related work Here we review the prior work most relevant to ours; further discussion is provided in AppendixËA. LLMs for radiology report generation (RRG). Early RRG systems mostly followed encoderâdecoder paradigms with Transformers, such as R2Gen [16], which established strong baselines on the MIMIC-CXR [48] and IU X-ray [25] datasets. Recent LLM-centric approaches align a medical visual encoder with a frozen or fine-tuned LLM to produce more fluent, clinically grounded reports, such as R2GenGPT [114], XrayGPT [107], MAIRA-1 [42], MAIRA-2 [7], and CheXagent [17], along with other approaches [74, 91, 124, 134, 128, 38, 56, 45, 20, 117, 79, 83, 131, 26, 30, 140, 122, 70, 90, 55, 1, 115, 24, 100, 32, 84, 3, 141, 49, 36, 116, 120, 73, 111, 136, 108, 39, 34, 135, 121, 51, 125, 8, 35, 46, 127, 40]. Agentic systems for RRG. Multi-agent frameworks are beginning to appear in RRG, where they align with clinical reasoning stages or combine retrieval-augmented generation (RAG). RadAgents [133] proposes a radiologist-like, multi-agent workflow for chest X-ray interpretation. CXRAgent [78] introduces a director-orchestrated, multi-stage agent for chest X-ray interpretation that validates tool outputs with an evidence-driven validator and coordinates diagnostic planning and team-based reasoning. Yi et al. [126] decomposes CXR reporting into retrieval, draft, refinement, vision, and synthesis agents aligned with stepwise clinical reasoning. Elboardy et al. [28] proposes a model-agnostic, ten-agent framework that unifies radiology report generation and evaluation, coordinated by an orchestrator and including an LLM-as-a-judge. However, these agentic RRG systems are mostly training-free, relying on pretrained models without end-to-end optimization of entire systems. Reinforcement learning for RRG. Recent studies have incorporated reinforcement learning (RL) to enhance clinical accuracy in RRG. LM-RRG [140] integrates clinical-quality RL by directly optimizing RadCliQ [129] as a reward signal. BoxMed-RL [47] couples chain-of-thought supervision with spatially verifiable RL that ties textual findings to bounding-box evidence. DeepMedix-R1 [65] employs a three-stage pipeline: instruction-fine-tuning, synthetic reasoning sample exposure, and online RL. OraPO [18] proposes an oracle-educated group relative policy optimization (GRPO) for RRG that leverages fact-level rewards. Med-R1 [54] applies GRPO-based RL to medical VLMs across eight imaging modalities and five VQA task types. However, these RL approaches primarily focus on optimizing a single model rather than optimizing entire multi-agent systems. 3 Method 3.1 Preliminaries In this section, we introduce the notations necessary for this paper, and as background, briefly describe Group Sequence Policy Optimization (GSPO) [139], which is a sequence-level variant of GRPO [96] that improved performance. Let D denote the data distribution over queryâanswer pairs (q,a)(q,a). Like GRPO, GSPO samples a group of G responses (xii=1GâźĎθold(â âŁq))(\x_i\_i=1^G _ _old(¡ q)) for each query q, where Ďθold _ _old denotes the policy used to generate outputs before updating the current policy. Each sampled response xix_i is then evaluated by a verifiable reward râ(xi,a)r(x_i,a), and group-relative advantage is computed as follows: A^iârâ(xi,a)âmeanâĄ(râ(xi,a)i=1G)stdâĄ(râ(xi,a)i=1G). A_i r(x_i,a)-mean(\r(x_i,a)\_i=1^G)std(\r(x_i,a)\_i=1^G). (1) GSPO then performs policy optimization by maximizing the following objective: GSPOâ(θ)â(q,a)âź,xii=1GâźĎθold(â âŁq)â[1Gââi=1GminâĄ(siâ(θ)âA^i,clipâĄ(siâ(θ),1âÎľlow,1+Îľhigh)âA^i)], \!\!\!J_GSPO(θ)\! \!E_(q,a) ,\x_i\_i=1^G\! _ _old\!(¡ q)\! [\! 1G\! _i=1^G \! (\!s_i(θ) A_i,clip(s_i(θ),1- _low,1+ _high) A_i\! )\!\! ]\!,\!\! (2) where siâ(θ) s_i(θ) â(Ďθâ(xiâŁq)Ďθoldâ(xiâŁq))1/|xi|=(ât=1|xi|Ďθâ(xi,tâŁq,xi,<t)ât=1|xi|Ďθoldâ(xi,tâŁq,xi,<t))1/|xi| ( _θ(x_i q) _ _old(x_i q) )^\!1/|x_i|= ( _t=1^|x_i| _θ(x_i,t q,x_i,<t) _t=1^|x_i| _ _old(x_i,t q,x_i,<t) )^\!1/|x_i| (3) is the importance ratio based on sequence likelihoods, and |xi||x_i| denotes the number of tokens of response xix_i. Here, xi,tx_i,t and xi,<tâ(xi,1,âŚ,xi,tâ1)x_i,<t (x_i,1,âŚ,x_i,t-1) denote the t-th tokens and the preceding tokens of the response xix_i, respectively. 3.2 Multi-agent GSPO In this section, we introduce a simple yet effective multi-agent formulation of GSPO. Although the formulation is a simple extension of GSPO, it enables the entire agentic system to be optimized end-to-end in an on-policy manner, encouraging coordination among agents within the deployed workflow. The adaptation to RRG is discussed in SectionË3.3. Let there be K agents with policies Ďθ(k)k=1K\ _θ^(k)\_k=1^K, where the indices k=1,âŚ,Kk=1,âŚ,K follow the order in which the agents are activated. For each (q,a)âź(q,a)\! \!D, we sample a group of G joint rollouts: xii=1GâźĎθoldagent(â âŁq)ââk=1KĎθold(k)(â âŁci(k)), \x_i\_i=1^G _ _old^agent(¡ q) _k=1^K _ _old^(k)(\,¡\, c^(k)_i), (4) where xiâ(xi(1),âŚ,xi(K)) x_i (x^(1)_i,âŚ,x^(K)_i). Here, ci(k) c^(k)_i denotes the context observed by agent k when generating xi(k)x^(k)_i. This context may include the query q or the outputs of preceding agents. Each joint rollout is evaluated by a system-level verifiable reward riârâ(xi(K),a), r_i r(x_i^(K),a), and the resulting group-relative advantage computed by Eq.Ë1 is shared across all agents. Then, we define the multi-agent GSPO (MA-GSPO) objective as follows: MA-GSPOâ(θ1,âŚ,θK) \!\!J_MA-GSPO( _1,âŚ, _K) =(q,a)âź,xiâźĎθoldagent(â âŁq)â[1Kââk=1K1Gââi=1GminâĄ(si(k)â(θk)âA^i,clipâĄ(si(k)â(θk),1âÎľlow,1+Îľhigh)âA^i)], \!\!\!=\!E_(q,a) ,\x_i\ _ _old^agent(¡ q)\! [\! 1K\! _k=1^K\! 1G\! _i=1^G \! (\!s^(k)_i\!( _k) A_i,clip\! (s^(k)_i\!( _k),1- _low,1+ _high ) A_i\! )\! ]\!,\!\! (5) where the importance ratio for agent k is: si(k)â(θk)=(Ďθkâ(xi(k)âŁci(k))Ďθkoldâ(xi(k)âŁci(k)))1/|xi(k)|.s^(k)_i( _k)= ( _ _k\! (x^(k)_i c^(k)_i ) _θ^old_k\! (x^(k)_i c^(k)_i ) )^\!1/|x^(k)_i|. (6) This reduces to standard GSPO when K=1K=1. Although the advantage is shared, each agent is updated through its own sequence-level importance ratio, so the objective optimizes the joint agentic system while preserving agent-specific policy updates. Chest X-ray Left-Region AgentCentral-Region AgentRight-Region AgentGlobal Integrating Agent### FindingsThe lungs are clear âŚ### ImpressionNo acute cardiopulmonaryâŚFinal Generated Report ### FindingsThe lungs are clear âŚ### ImpressionNo acute cardiopulmonaryâŚGround Truth Report Clinically Verifiable Reward Examines left lung,left hilar structures, âŚExamines cardiac silhouette,mediastinum, trachea, âŚExamines right lung,right hilar structures, âŚDiagnosis ofleft regionDiagnosis ofcentral regionDiagnosis ofright region Figure 1: Overview of the proposed multi-agent RL framework. Region-specific agents and the global integrating agent collaboratively generate the radiology report, and the entire agent system is jointly optimized through RL based on clinically verifiable rewards. 3.3 Multi-agent RL on RRG The overall architecture of MARL-Rad is illustrated in Fig.Ë1. Our system consists of three region-specific agents and one global integrating agent that collaborate to generate the final diagnostic report. Each region-specific agent focuses on a distinct anatomical region of the chest X-ray: a left-region agent examines structures such as the left lung, left hilar structures, left costophrenic angle, and left clavicle; a right-region agent examines structures such as the right lung, right hilar structures, right costophrenic angle, and right clavicle; and a central-region agent examines structures such as the cardiac silhouette, mediastinum, aortic arch, trachea, and spine. This design follows the actual diagnostic workflow of radiologists, who meticulously examine each anatomical region of a chest X-ray before composing the complete report. The global integrating agent then receives the outputs from the three region-specific agents, incorporates their diagnoses, and produces a final comprehensive report reflecting the overall condition of the chest X-ray. For the reinforcement learning, we employ clinically verifiable rewards using the CheXbert [101] accuracy, which measures the correctness of predicted clinical findings based on automatically labeled disease categories, and the RadGraph F1 [43], which evaluates factual and relational consistency. Additionally, we incorporate the ROUGE-L [64] to encourage lexical alignment with reference reports. The final reward is defined as the unweighted sum of these three components, and the entire agent system is jointly optimized as we described in SectionË3.2. 4 Experiments 4.1 Experimental setup In this section, we describe the overview of the experimental setup. Further details are provided in AppendixËB. Datasets We train our agents using the MIMIC-CXR [48] dataset. Following prior studies [16, 55, 5, 41, 85, 72], we exclude samples without the corresponding reports. We adopt the official split of MIMIC-CXR and use the training set for training our agents. Evaluation is performed on the official test split of MIMIC-CXR and the IU X-ray [25] dataset. For IU X-ray, we follow the standard 70%/20%/10% split used in prior work [16, 15, 67, 36, 120, 15, 39, 97, 113, 41] and use only the test split for cross-dataset evaluation. Table 1: Comparison with previous state-of-the-art methods on the MIMIC-CXR dataset [48]. The best score for each metric within each section is highlighted in bold. Our approach achieves the best performance across all clinical efficacy (CE) metrics, which are considered to align more closely with cliniciansâ assessments [105, 69, 9, 129]. Method NLG Metricsâ\ CE Metricsâ\ BLEU-1 BLEU-4 METEOR ROUGE-L RadGraph F1 CheXbert F1 GREEN MIMIC-CXR Findings R2Gen [16] 0.353 0.103 0.142 0.277 - 0.276 - R2GenCMN [15] 0.353 0.106 0.142 0.278 - 0.278 - PPKED [68] 0.360 0.106 0.149 0.284 - - - CMCL [67] 0.344 0.097 0.133 0.281 - - - SA [122] - - - - 0.228 - - RGRG [104] 0.373 0.126 0.168 0.264 - 0.447 - METransformer [113] 0.386 0.124 0.152 0.291 - 0.311 - KiUT [41] 0.393 0.113 0.160 0.285 - 0.321 - CoFE [58] - 0.125 0.176 0.304 - 0.405 - MAN [97] 0.396 0.115 0.151 0.274 - 0.389 - Med-LMM [72] - 0.128 0.161 0.289 - 0.395 - SEI [71] - 0.135 0.158 0.299 0.249 0.460 - FMVP [75] 0.389 0.108 0.150 0.284 - 0.336 - HERGen [109] 0.395 0.122 0.156 0.285 - 0.317 - CMN [15] 0.353 0.106 0.142 0.278 - 0.278 - CXRMate [85] - 0.079 - 0.262 0.272 0.357 - I3+C2FD [66] 0.402 0.128 0.175 0.291 - 0.473 - MLRG [70] 0.411 0.158 0.176 0.320 0.291 0.505 0.353 DART [90] 0.437 0.137 0.175 0.310 - 0.533 - LM-RRG [140] - 0.122 0.165 0.296 - 0.484 - DeepMedix-R1 [65] 0.340 0.105 0.289 0.329 0.238 0.239 - SRRG [36] 0.416 0.128 0.294 0.162 - 0.486 - MPO [120] 0.416 0.139 0.162 0.309 - 0.353 - DAMPER [39] 0.402 0.193 0.289 0.301 - 0.507 - OISA [121] 0.428 0.129 - - 0.244 0.486 0.322 MedGemma [94] 0.285 0.014 0.238 0.203 0.195 0.494 0.361 MARL-Rad (Ours) 0.533 0.056 0.290 0.275 0.294 0.544 0.396 MIMIC-CXR Findings + Impression MedRegA [110] 0.405 0.126 0.319 0.276 - - - CXRMate [85] - 0.074 0.158 0.255 - 0.378 - MLRG [70] 0.402 0.152 0.172 0.327 0.289 0.509 - Flamingo-CXR [105] - 0.101 - 0.297 0.205 0.519 - MedGemma [94] 0.333 0.052 0.338 0.197 0.209 0.523 0.388 MARL-Rad (Ours) 0.598 0.142 0.397 0.292 0.316 0.556 0.398 Evaluation metrics Following previous studies [16, 15, 122, 104, 113, 41, 58, 97, 72, 71, 75, 109, 15, 85, 66, 70, 90, 140, 55, 1, 65, 133, 36, 120, 39, 121, 110], we evaluate the generated reports using both natural language generation (NLG) metrics and CE metrics. For NLG metrics, we report BLEU-1, BLEU-4 [88], METEOR [6], and ROUGE-L [64]. For CE metrics, we report RadGraph F1 [43], CheXbert F1 [101], and GREEN scores [87]. CheXbert is a 14-label classification models that assess the presence of common thoracic findings (e.g., pneumonia, atelectasis, edema) from generated reports. RadGraph measures the correctness of clinical entity and relation extraction, reflecting the structural and factual consistency of the report. The GREEN score is a LLM-based metric that evaluates clinical efficacy by grading report-level clinical correctness beyond surface-level lexical overlap. Consistent with prior work, evaluations are performed on the Findings section alone [36, 70, 15, 55, 85] and on both the Findings and Impression sections combined [70, 105, 55]. Used models We use MedGemma-4B [94] as the base model for all agents, including the three region-specific agents and the global integrating agent. All agents are trained jointly end-to-end from the same base checkpoint, but they do not share parameters, yielding distinct parameter sets. Table 2: Comparison with previous state-of-the-art methods on the IU X-ray dataset [25]. The best score for each metric within each section is highlighted in bold. Our approach achieves the best performance across all clinical efficacy (CE) metrics, which are considered to align more closely with cliniciansâ assessments [105, 69, 9, 129]. Method NLG Metricsâ\ CE Metricsâ\ BLEU-1 BLEU-4 METEOR ROUGE-L RadGraph F1 CheXbert F1 GREEN IU X-ray Findings R2Gen [16] 0.470 0.165 0.187 0.371 - - - R2GenCMN [15] 0.475 0.170 0.191 0.376 - - - PPKED [68] 0.483 0.168 0.190 0.376 - - - METransformer [113] 0.483 0.172 0.192 0.380 - - - KiUT [41] 0.525 0.185 0.242 0.409 - - - CoFE [58] - 0.175 0.202 0.438 - - - CMCL [67] 0.473 0.162 0.186 0.378 - - - MAN [97] 0.501 0.170 0.213 0.386 - - - CMN [15] 0.475 0.170 0.191 0.375 - - - Med-LMM [72] - 0.168 0.381 - - - - FMVP [75] 0.485 0.169 0.201 0.398 - - - LM-RRG [140] - 0.208 0.216 0.387 - - - CXRMate [85] - 0.046 - 0.282 0.291 0.277 - I3+C2FD [66] 0.499 0.184 0.208 0.390 - - - SRRG [36] 0.533 0.218 0.219 0.418 - - - MPO [120] 0.548 0.209 0.224 0.415 - - - DAMPER [39] 0.520 0.225 0.284 0.397 - - - Multi-Agent [126] - 0.047 0.362 0.247 - - - OISA [121] 0.431 0.131 - - 0.282 0.219 0.481 MedGemma [94] 0.179 0.012 0.287 0.215 0.270 0.427 0.632 MARL-Rad (Ours) 0.468 0.046 0.347 0.292 0.337 0.501 0.644 IU X-ray Findings + Impression CXRMate [85] - 0.046 - 0.282 0.291 0.277 - MedGemma [94] 0.244 0.058 0.432 0.212 0.265 0.469 0.653 MARL-Rad (Ours) 0.591 0.182 0.531 0.348 0.365 0.516 0.659 4.2 Main results: Comparison with previous state-of-the-art methods The results are shown in TablesË1 and 2. Our method outperforms the baselines on all CE metrics. As demonstrated in Tanno et al. [105], where evaluations were conducted by 27 board-certified radiologists from two regions (the United States and India), traditional NLG metrics based solely on linguistic similarity do not align with cliniciansâ assessments, whereas CE metrics are more consistent with clinical judgments. Multiple studies have similarly reported that conventional NLG metrics fail to reflect cliniciansâ evaluations, underscoring the need to prioritize CE metrics [69, 9, 129]. Considering this, it is reasonable that our method does not align with some NLG metrics, while the strong performance on CE metrics suggests improved clinical relevance. The consistent improvement on both the MIMIC-CXR and IU X-ray datasets demonstrates the effectiveness across different evaluation settings. Notably, despite being trained exclusively on MIMIC-CXR training set, our agent attains high performance on IU X-ray, suggesting promising robustness and cross-dataset generalization. We also report the VRAM usage and inference throughput of MARL-Rad in AppendixËC, showing that despite the additional computation introduced by agentization, it maintains a practical inference speed for real-world radiology reporting workflows. 4.3 Ablation study We conduct ablation studies to examine the effects of workflow-aligned RL, regional decomposition, and reward design. The results are presented in TableË3. 4.3.1 Training-free agentic workflow is insufficient First, we evaluate a training-free agentic workflow in which MedGemma is kept fixed and executed with the same agentization as MARL-Rad. The results are presented as âMedGemma (agent)â in TableË3. This variant performs worse than the vanilla MedGemma model, despite following the same multi-agent workflow as our full method. This suggests that simply organizing a fixed LLM into an agentic workflow is insufficient and can even degrade performance, likely because the underlying model has not been optimized for its assigned roles or for coordinated reasoning within the deployed workflow. This finding highlights the importance of jointly optimizing the entire agent system, rather than relying on naive agentification without learning. Furthermore, our method also achieves higher scores than GSPO training without agentization (presented as âMedGemma + RLâ), indicating that our regional multi-agent workflow itself is effective when the agent system is properly optimized for the deployed workflow. We conduct a more in-depth analysis of this advantage in SectionË4.4. Table 3: Ablation study results. âMedGemma (vanilla)â represents the unmodified model without reinforcement learning or agentization. âMedGemma (agent)â employs the same agentic workflow as MARL-Rad but without any RL optimization. âMedGemma + RLâ applies GSPO to MedGemma in a single-agent setting. âCounterfactual Rewardâ uses agent-specific rewards based on counterfactual removal of each regional agentâs output from the global agentâs context. The best and worst scores for each metric within each dataset are highlighted in bold and underline, respectively. Method NLG Metricsâ\ CE Metricsâ\ BLEU-1 BLEU-4 METEOR ROUGE-L RadGraph F1 CheXbert F1 GREEN MIMIC-CXR Findings + Impression MedGemma (vanilla) 0.333 0.052 0.338 0.197 0.209 0.523 0.388 MedGemma (agent) 0.252 0.044 0.267 0.144 0.106 0.309 0.344 MedGemma + RL 0.465 0.116 0.373 0.283 0.299 0.502 0.377 Counterfactual Reward 0.575 0.136 0.375 0.272 0.292 0.518 0.369 MARL-Rad (Ours) 0.598 0.142 0.397 0.292 0.316 0.556 0.398 IU X-ray Findings + Impression MedGemma (vanilla) 0.244 0.058 0.432 0.212 0.265 0.469 0.653 MedGemma (agent) 0.210 0.053 0.380 0.163 0.141 0.348 0.649 MedGemma + RL 0.598 0.181 0.529 0.339 0.358 0.495 0.629 Counterfactual Reward 0.561 0.166 0.513 0.323 0.349 0.478 0.614 MARL-Rad (Ours) 0.591 0.182 0.531 0.348 0.365 0.516 0.659 4.3.2 Shared versus counterfactual rewards A natural concern with using a shared system-level reward is that it may suffer from a credit assignment problem: since all agents receive the same final reward, it may be unclear which agent contributed to the improvement or degradation of the final report. To examine this issue, we consider a naive credit-assignment variant based on counterfactual rewards. Let rir_i denote the reward of the original final report for the i-th joint rollout. For a regional agent k, let ri(âk) r_i^( k) denote the reward obtained after regenerating the final report with the global integrating agent while removing agent kâs intermediate output from its input context. We then assign agent k the difference reward ri(k)âriâri(âk) r_i^(k) r_i-r_i^( k). That is, each regional agent is rewarded according to how much the final report reward decreases when its regional output is removed from the integrating agentâs context. For the global integrating agent, we keep the original system-level reward rir_i. We report this variant as âCounterfactual Rewardâ in TableË3. As shown in TableË3, this counterfactual reward variant does not improve over our shared-reward formulation and instead yields slightly worse clinical metrics. This suggests that explicitly decomposing the final reward in this naive manner is not necessarily beneficial for RRG. One reason is that the counterfactual report is generated from an ablated context that differs from the deployed workflow. Thus, the resulting reward difference reflects not only the contribution of the removed regional agent, but also the global integrating agentâs ability to compensate under an out-of-distribution context. In contrast, the success of our simple shared-reward formulation shows that explicit counterfactual reward decomposition is not necessary for this setting. For structured short-horizon agentic RRG workflows, optimizing the complete system with a shared clinically grounded reward is sufficient and preferable to noisy counterfactual reward decomposition. 4.4 Analysis of laterality consistency using RadGraph Table 4: Laterality-specific RadGraph F1 scores. âLeft Agentâ and âRight Agentâ denote the pre-integration outputs of the corresponding regional agents. Higher scores indicate better laterality consistency. The best score among the final-report methods, i.e., MedGemma w/ RL and MARL-Rad (Ours), is highlighted in bold. Method Left Right Left lung Right lung MIMIC-CXR Findings + Impression MedGemma w/ RL 0.144 0.142 0.219 0.102 MARL-Rad (Ours) 0.184 0.277 0.406 0.146 Left Agent 0.214 0.011 0.369 0.066 Right Agent 0.001 0.257 0.000 0.156 IU X-ray Findings + Impression MedGemma w/ RL 0.063 0.044 0.035 0.057 MARL-Rad (Ours) 0.082 0.166 0.032 0.168 Left Agent 0.130 0.005 0.155 0.082 Right Agent 0.000 0.169 0.000 0.175 To assess the effect of decomposing the task into region-specific agents, we conduct a detailed analysis of the RadGraph outputs to examine region-level consistency. RadGraph extracts entities classified as either âAnatomyâ or âObservationâ along with the relations between them, enabling the identification of clinically meaningful regionâfinding pairs, such as âleftâlungâpneumonia.â Among region-level attributes, laterality (left vs. right) is the one most directly handled by our left/right agents, and it carries critical clinical importance: errors in leftâright identification are strictly unacceptable in real clinical practice. Therefore, we extract representative laterality-related entities from the RadGraph output such as âleft,â âright,â âleft lung,â and âright lung,â along with the entities connected to them via relations. Based on these subsets, we define a laterality-specific RadGraph F1 score by applying the standard RadGraph evaluation procedure only to these laterality-focused subgraphs and use it for our analysis. The detailed procedure for computing the laterality-specific RadGraph F1 score is described in AppendixËD. Input Chest X-ray Image Output of Left-Region Agent ### Findings The left lung is relatively clear. The left hilar structures are unremarkable. The left costophrenic angle is sharp. The left clavicle is intact. The left hemithorax is unremarkable. The left apical zone is unremarkable. ### Impression There is no evidence of acute infiltrate in the left lung. The left hilar structures are unremarkable. The left costophrenic angle is sharp. The left clavicle is intact. The left hemithorax is unremarkable. The left apical zone is unremarkable. Output of Central-Region Agent ### Findings The heart size is normal. The mediastinum is unremarkable. The aortic arch is visualized. The trachea and carina are midline. The SVC is visualized. The atria and cavoatrial junction are visualized. The abdomen is not well visualized. The spine shows degenerative changes. No acute findings in the lungs. ### Impression No acute findings. Output of Right-Region Agent ### Findings The right lung is clear. The right hilar structures are unremarkable. The right costophrenic angle is sharp. The right hemidiaphragm is visible. The right clavicle is visible. The right apical zone is visible. ### Impression The right lung is clear. There is no evidence of infiltrate, pneumothorax, or pleural effusion. Output of Global-Integrating Agent ### Findings The lungs are clear without focal consolidation, pleural effusion, or pneumothorax. The heart size is normal. The mediastinal contours are unremarkable. There is evidence of prior median sternotomy with surgical clips noted in the anterior mediastinum. There is mild emphysema, particularly in the lung apices. There is also evidence of right pleural thickening. ### Impression No acute cardiopulmonary process. Ground Truth Report ### Findings The cardiomediastinal and hilar contours are stable. The aorta is again noted to be tortuous. The patient is status post CABG with median sternotomy wires in place. The second most superior median sternotomy wires again noted to be fractured. There is no pleural effusion or pneumothorax. The lungs are well-expanded with stable scarring at the right costophrenic angle. There is no new focal consolidation concerning for pneumonia. There is no overt pulmonary edema. The upper abdomen is unremarkable aside from surgical clips. ### Impression No acute cardiopulmonary process. Figure 2: Example output from MARL-Rad. Region-specific agents focus on their assigned regions and generate regional diagnoses. The global integrating agent synthesizes these drafts into a concise and coherent report while adding relevant global findings. The resulting final report is detailed and aligns well with the ground truth, correctly concluding with âNo acute cardiopulmonary process.â The results are shown in TableË4. We observe that MARL-Rad captures laterality more accurately than the single-agent RL baseline. This improvement may be attributed to a limitation identified in recent studies, which report that standard visionâlanguage models tend to behave as global image parsers and struggle to reason about spatial relations at the region level [19, 13]. By introducing region-specific agentsâincluding dedicated left/right agentsâour approach decomposes the interpretation of image regions and enables each agent to focus on its corresponding region, thereby enhancing the modelâs ability to capture laterality. To further verify that the improvement is not merely driven by the global integrating agent, we also report the laterality-specific RadGraph F1 scores of the left and right agents in TableË4. These scores evaluate each regional agentâs output before global integration. As shown in the table, each left/right agent achieves a score comparable to that of the final integrated report on its assigned side, while its score on the opposite side remains much lower. This indicates that MARL-Rad does not simply rely on the global agent to correct or infer regional information; rather, the regional agents themselves learn to produce clinically meaningful localized diagnoses. 4.5 Case study Fig.Ë2 shows an example case from our multi-agent RL system. In this example, the left, right, and central region agents each focus on their respective anatomical responsibilities and consistently produce findings and impression texts that are restricted to their assigned regions. This behavior indicates that the intended task decomposition is functioning as designed. Furthermore, the global integrating agent leverages these region-specific drafts and produces a concise and coherent final report, rather than simply concatenating the regional outputs. For instance, given the left and right agentsâ statements âThe left lung is relatively clearâ and âThe right lung is clear,â the integrating agent appropriately compresses these into âThe lungs are clear.â Similarly, central findings such as âThe heart size is normalâ are preserved and incorporated into the final report. In addition, the integrating agent also adds global findings, such as post-surgical changes (âThere is evidence of prior median sternotomy with surgical clips noted in the anterior mediastinumâ), which are not tied to any specific region. As a result, the final report accurately captures both the absence of acute pulmonary abnormalities and the patientâs status post cardiac surgery, aligning with the ground-truth conclusion of âNo acute cardiopulmonary process.â ### Findings The lungs are clear. There is no focal consolidation, pleural effusion, or pneumothorax. The heart size is normal. The mediastinal and hilar contours are unremarkable. ### Impression No acute cardiopulmonary process. Figure 3: An example case generated by the single-agent RL model (MedGemma w/ RL) using the same input image as in Fig.Ë2. For comparison, Fig.Ë3 presents an example generated by the single-agent RL model on the same image, where MedGemma is optimized with RL without agentization. Although the content of the report and its final conclusion are broadly consistent with the ground truth, the report lacks detailed analysis and remains somewhat vague, omitting several clinically relevant findings present in the ground-truth report. This contrast underscores the advantage of agentization, which explicitly examines each anatomical region and yields more accurate, detailed reports. Additionally, as illustrated in this example, agentization improves interpretability: each regional agent produces an explicit, region-focused diagnosis, making it clear how the model assessed each anatomical region. This transparency is particularly valuable for clinicians or end-users who may not be familiar with the internal behavior of LVLMs. In contrast, a single-agent model provides only a monolithic summary, making it difficult to verify whether the underlying regional findings were adequately considered. 5 Clinician evaluation Figure 4: Physician preferences comparing MARL-Rad and the ground-truth reports. Bars show 5-point Likert ratings for five criteria; the rightmost panel shows the overall 2-choice vote. During evaluation, report sources were anonymized and randomly shuffled for each case. Automated metrics may not fully capture clinical quality, and improvements on such metrics may not always align with clinician judgment. To partially address this concern, we conduct a small blinded clinician comparison between MARL-Rad outputs and ground-truth reports. To minimize bias, the two reports for each case were anonymized and randomly ordered. Physicians evaluated each pair together with the corresponding chest X-ray image, without knowing the source of either report. For each case, physicians rated five aspects: completeness, correctness, conciseness, readability, and clinical utility. Detailed definitions of these evaluation criteria are provided in AppendixËE. Each aspect was evaluated on a 5-point comparative Likert scale with the options âReport A is better,â âReport A is slightly better,â âAbout the same,â âReport B is slightly better,â and âReport B is better.â In addition, physicians provided an overall preference vote between reports A and B. We randomly sampled 10 images from the MIMIC-CXR test set, and six physicians completed the evaluation. The results are shown in Fig.Ë4. Across all five criteria, physician ratings show that MARL-Rad outputs are clinically comparable to the ground-truth reports. The overall preference vote is nearly balanced, with a slight preference for MARL-Rad. These results complement the automated evaluation and suggest that the gains of MARL-Rad do not merely reflect overfitting to automatic metrics, but correspond to clinically comparable report quality. 6 Conclusion In this work, we introduced MARL-Rad, a novel multi-modal multi-agent reinforcement learning framework for radiology report generation. MARL-Rad coordinates region-specific agents with a global integrating agent and jointly optimizes the entire agent system end-to-end to produce clinically consistent reports. By optimizing agents on policy within their deployed workflow, MARL-Rad overcomes the limitations of training-free agentification of fixed LLMs. Our method achieves state-of-the-art performance on CE metrics on both MIMIC-CXR and IU X-ray datasets, enhances laterality consistency, and yields more accurate, detailed reports. Although this work focuses on CXR-based RRG, the framework may be extended to other workflow-structured applications. Future work includes applying this framework to other modalities in medical AI, such as electrocardiogram (ECG) and echocardiography, as well as other real-world tasks. References [1] S. B. Ahsan, M. Ikhalas, M. M. Khan, S. Ullah, and M. Z. Zaheer (2025) ARDGen: augmentation regularization for domain-generalized medical report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, p. 6526â6535. External Links: Document Cited by: §1, §2, §4.1. [2] D. J. Alapat, M. V. Menon, and S. Ashok (2022) A review on detection of pneumonia in chest X-ray images using neural networks. Journal of Biomedical Physics and Engineering 12 (6), p. 551â558. External Links: Document Cited by: §1. [3] S. Albastaki, A. Sohail, I. I. Ganapathi, B. Alawode, A. Khan, S. Javed, N. Werghi, M. Bennamoun, and A. Mahmood (2025) Multi-resolution pathology-language pre-training model with text-guided visual representation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 25907â25919. Cited by: §2. [4] K. Baba, C. Liu, S. Kurita, and A. Sannai (2025) Prover Agent: an agent-based framework for formal mathematical proofs. arXiv preprint arXiv:2506.19923. Cited by: §A.2, §1. [5] K. Baba, R. Yagi, J. Takahashi, R. Kishikawa, and S. Kodera (2024) JRadiEvo: a japanese radiology report generation model enhanced by evolutionary optimization of model merging. arXiv preprint arXiv:2411.09933. Cited by: §1, §4.1. [6] S. Banerjee and A. Lavie (2005) METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, p. 65â72. Cited by: §4.1. [7] S. Bannur, K. Bouzid, D. C. Castro, A. Schwaighofer, A. Thieme, S. Bond-Taylor, M. Ilse, F. PĂŠrez-GarcĂa, V. Salvatelli, H. Sharma, F. Meissen, M. Ranjit, S. Srivastav, J. Gong, N. C. F. Codella, F. Falck, O. Oktay, M. P. Lungren, M. T. Wetscherek, J. Alvarez-Valle, and S. L. Hyland (2024) MAIRA-2: grounded radiology report generation. arXiv preprint arXiv:2406.04449. Cited by: §2. [8] X. Bo, F. Yang, F. Xu, and X. Zhang (2025) Cross-counter-repeat attention for enhanced understanding of visual semantics in radiology report generation. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 4242â4250. External Links: Document, ISBN 9798400720352 Cited by: §2. [9] W. Boag, T. H. Hsu, M. Mcdermott, G. Berner, E. Alesentzer, and P. Szolovits (2020) Baselines for chest X-ray report generation. In Proceedings of the Machine Learning for Health NeurIPS Workshop, Proceedings of Machine Learning Research, Vol. 116, p. 126â140. Cited by: §4.2, Table 1, Table 2. [10] G. W. L. Boland, A. S. Guimaraes, and P. R. Mueller (2008) Radiology report turnaround: expectations and solutions. European Radiology 18 (7), p. 1326â1328. External Links: Document, ISBN 1432-1084 Cited by: §1. [11] J. Broder (2011) Imaging the chest: the chest radiograph. In Diagnostic Imaging for the Emergency Physician, p. 185â296. External Links: Document, ISBN 978-1-4160-6113-7 Cited by: §1. [12] S. Candemir and S. Antani (2019) A review on lung boundary detection in chest X-rays. International Journal of Computer Assisted Radiology and Surgery 14 (4), p. 563â576. External Links: Document, ISBN 1861-6429 Cited by: §1. [13] B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia (2024-06) SpatialVLM: endowing vision-language models with spatial reasoning capabilities. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14455â14465. Cited by: §4.4. [14] M. Chen, L. Sun, T. Li, H. Sum, Y. Zhou, C. Zhu, H. Wang, J. Z. Pan, W. Zhang, H. Chen, F. Yang, Z. Zhou, and W. Chen (2025) ReSearch: learning to reason with search for LLMs via reinforcement learning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, Cited by: §A.1. [15] Z. Chen, Y. Shen, Y. Song, and X. Wan (2021) Cross-modal memory networks for radiology report generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 5904â5914. External Links: Document Cited by: §4.1, §4.1, Table 1, Table 1, Table 2, Table 2. [16] Z. Chen, Y. Song, T. Chang, and X. Wan (2020) Generating radiology reports via memory-driven transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 1439â1449. External Links: Document Cited by: §2, §4.1, §4.1, Table 1, Table 2. [17] Z. Chen, M. Varma, J. Delbrouck, M. Paschali, L. Blankemeier, D. V. Veen, J. M. J. Valanarasu, A. Youssef, J. P. Cohen, E. P. Reis, E. Tsai, A. Johnston, C. Olsen, T. M. Abraham, S. Gatidis, A. S. Chaudhari, and C. Langlotz (2024) CheXagent: towards a foundation model for chest X-ray interpretation. In AAAI 2024 Spring Symposium on Clinical Foundation Models, Cited by: §2. [18] Z. Chen, H. Yu, Y. Xu, Y. Luo, L. Duong, and Y. Li (2025) OraPO: oracle-educated reinforcement learning for data-efficient and factual radiology report generation. arXiv preprint arXiv:2509.18600. Cited by: §2. [19] A. Cheng, H. Yin, Y. Fu, Q. Guo, R. Yang, J. Kautz, X. Wang, and S. Liu (2024) SpatialRGPT: grounded spatial reasoning in vision-language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: §4.4. [20] N. C. F. Codella, Y. Jin, S. Jain, Y. Gu, H. H. Lee, A. B. Abacha, A. Santamaria-Pang, W. Guyman, N. Sangani, S. Zhang, H. Poon, S. Hyland, S. Bannur, J. Alvarez-Valle, X. Li, J. Garrett, A. McMillan, G. Rajguru, M. Maddi, N. Vijayrania, R. Bhimai, N. Mecklenburg, R. Jain, D. Holstein, N. Gaur, V. Aski, J. Hwang, T. Lin, I. Tarapov, M. Lungren, and M. Wei (2024) MedImageInsight: an open-source embedding model for general domain medical imaging. arXiv preprint arXiv:2410.06542. Cited by: §2. [21] I. A. Cowan, S. L. S. MacDonald, and R. A. Floyd (2013) Measuring and managing radiologist workload: measuring radiologist reporting times using data from a radiology information system. Journal of Medical Imaging and Radiation Oncology 57 (5), p. 558â566. External Links: Document, ISSN 1754-9485 Cited by: §1. [22] D. C. de Castro, A. Bustos, S. Bannur, S. L. Hyland, K. Bouzid, M. T. Wetscherek, M. D. SĂĄnchez-Valverde, L. Jaques-PĂŠrez, L. PĂŠrez-RodrĂguez, K. Takeda, J. M. Salinas-Serrano, J. Alvarez-Valle, J. Galant-Herrero, and A. Pertusa (2025) PadChest-GR: a bilingual chest X-ray dataset for grounded radiology report generation. NEJM AI 2 (7), p. AIdbp2401120. External Links: Document Cited by: §1. [23] DeepSeek-AI (2025) DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948. Cited by: §1. [24] J. Delbrouck, J. Xu, J. Moll, A. Thomas, Z. Chen, S. Ostmeier, A. Azhar, K. Z. Li, A. Johnston, C. Bluethgen, E. P. Reis, M. S. Muneer, M. Varma, and C. Langlotz (2025) Automated structured radiology report generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 26813â26829. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §1, §2. [25] D. Demner-Fushman, M. D. Kohli, M. B. Rosenman, S. E. Shooshan, L. Rodriguez, S. Antani, G. R. Thoma, and C. J. McDonald (2015-07) Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association 23 (2), p. 304â310. External Links: Document, ISSN 1067-5027 Cited by: Appendix B, §2, §4.1, Table 2. [26] F. Dong, S. Nie, M. Chen, F. Xu, and Q. Li (2025) Keyword-based ai assistance in the generation of radiology reports: a pilot study. npj Digital Medicine 8 (1), p. 490. External Links: Document, ISBN 2398â6352 Cited by: §1, §2. [27] G. Dong, Y. Chen, X. Li, J. Jin, H. Qian, Y. Zhu, H. Mao, G. Zhou, Z. Dou, and J. Wen (2025) Tool-Star: empowering LLM-brained multi-tool reasoner via reinforcement learning. arXiv preprint arXiv:2505.16410. Cited by: §A.1. [28] A. T. Elboardy, G. Khoriba, and E. A. Rashed (2025) Medical AI consensus: a multi-agent framework for radiology report generation and evaluation. arXiv preprint arXiv:2509.17353. Cited by: §1, §2. [29] J. Feng, S. Huang, X. Qu, G. Zhang, Y. Qin, B. Zhong, C. Jiang, J. Chi, and W. Zhong (2025) ReTool: reinforcement learning for strategic tool use in LLMs. External Links: 2504.11536 Cited by: §A.1. [30] A. Fink, A. Rau, M. Reisert, F. Bamberg, and M. F. Russe (2025) Retrieval-augmented generation with large language models in radiology: from theory to practice. Radiology: Artificial Intelligence 7 (4), p. e240790. External Links: Document Cited by: §2. [31] Google (2025) Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §1. [32] A. Heiman, X. Zhang, E. Chen, S. E. Kim, and P. Rajpurkar (2025) FactCheXcker: mitigating measurement hallucinations in chest X-ray report generation models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 30787â30796. Cited by: §2. [33] S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber (2024) MetaGPT: meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, Cited by: §A.2, §1. [34] W. Hou, Y. Cheng, K. Xu, H. Li, Y. Hu, W. Li, and J. Liu (2025) RADAR: enhancing radiology report generation with supplementary knowledge injection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 26366â26381. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §2. [35] X. Hou, X. Li, M. Lu, S. Wang, and Y. Zhang (2025-08) RRG-Mamba: efficient radiology report generation with state space model. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, p. 7410â7418. External Links: Document Cited by: §2. [36] X. Hou, Y. Li, and S. Wang (2025) Knowledge-driven query network with adaptive cross-view attention for structured radiology report generation. In IEEE/CVF International Conference on Computer Vision Workshops, p. 1234â1243. Cited by: §2, §4.1, §4.1, Table 1, Table 2. [37] M. Hu, Y. Zhou, W. Fan, Y. Nie, Z. Ye, B. Xia, T. Sun, Z. Jin, Y. Li, Z. Zhang, Y. Wang, Q. Ye, B. Ghanem, P. Luo, and G. Li (2025) OWL: optimized workforce learning for general multi-agent assistance in real-world task automation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, Cited by: §A.2, §1. [38] S. Huang, L. Shen, M. P. Lungren, and S. Yeung (2021) GLoRIA: a multimodal global-local representation learning framework for label-efficient medical image recognition. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), p. 3922â3931. External Links: Document Cited by: §2. [39] X. Huang, W. Chen, J. Liu, Q. Lu, X. Luo, and L. Shen (2025) DAMPER: a dual-stage medical report generation framework with coarse-grained mesh alignment and fine-grained hypergraph matching. AAAI Conference on Artificial Intelligence 39 (4), p. 3769â3778. External Links: Document Cited by: §2, §4.1, §4.1, Table 1, Table 2. [40] X. Huang, Y. Han, Y. L, R. Li, P. Wu, and K. Zhang (2025) CmEAA: cross-modal enhancement and alignment adapter for radiology report generation. In Proceedings of the 31st International Conference on Computational Linguistics, p. 8546â8556. Cited by: §2. [41] Z. Huang, X. Zhang, and S. Zhang (2023) KiUT: knowledge-injected u-transformer for radiology report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19809â19818. External Links: Document Cited by: §4.1, §4.1, Table 1, Table 2. [42] S. L. Hyland, S. Bannur, K. Bouzid, D. C. Castro, M. Ranjit, A. Schwaighofer, F. PĂŠrez-GarcĂa, V. Salvatelli, S. Srivastav, A. Thieme, N. Codella, M. P. Lungren, M. T. Wetscherek, O. Oktay, and J. Alvarez-Valle (2024) MAIRA-1: a specialised large multimodal model for radiology report generation. arXiv preprint arXiv:2311.13668. Cited by: §2. [43] S. Jain, A. Agrawal, A. Saporta, S. Truong, D. N. Duong, T. Bui, P. Chambon, Y. Zhang, M. Lungren, A. Ng, C. Langlotz, P. Rajpurkar, and P. Rajpurkar (2021) RadGraph: extracting clinical entities and relations from radiology reports. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, Vol. 1. Cited by: §1, §3.3, §4.1. [44] F. Jiang, C. Pan, L. Dong, K. Wang, O. A. Dobre, and M. Debbah (2025) From large AI models to agentic AI: a tutorial on future intelligent communications. arXiv preprint arXiv:2505.22311. Cited by: §A.2, §1, §1. [45] H. Jiang, X. Hao, Y. Huang, C. Ma, J. Zhang, Y. Pan, and R. Zhang (2025) Advancing medical radiograph representation learning: a hybrid pre-training paradigm with multilevel semantic granularity. In European Conference on Computer Vision Workshops, p. 16â33. External Links: ISBN 978-3-031-91721-9 Cited by: §2. [46] Y. Jiang, J. Chen, D. Yang, M. Li, S. Wang, T. Wu, K. Li, and L. Zhang (2025) CoMT: chain-of-medical-thought reduces hallucination in medical report generation. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1â5. External Links: Document Cited by: §2. [47] P. Jing, K. Lee, Z. Zhang, H. Zhou, Z. Yuan, Z. Gao, L. Zhu, G. Papanastasiou, Y. Fang, and G. Yang (2025) Reason like a radiologist: chain-of-thought and reinforcement learning for verifiable report generation. arXiv preprint arXiv:2504.18453. Cited by: §2. [48] A. E. W. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C. Deng, R. G. Mark, and S. Horng (2019) MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 6 (1), p. 317. External Links: Document, ISBN 2052-4463 Cited by: Appendix B, §2, §4.1, Table 1. [49] H. Kalisch, F. HĂśrst, J. Kleesiek, K. Herrmann, and C. Seibold (2025) CT-GRAPH: hierarchical graph attention network for anatomy-guided CT report generation. arXiv preprint arXiv:2508.05375. Cited by: §2. [50] Y. Kim, C. Park, H. Jeong, Y. S. Chan, X. Xu, D. McDuff, H. Lee, M. Ghassemi, C. Breazeal, and H. W. Park (2024) MDAgents: an adaptive collaboration of LLMs for medical decision-making. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: §1. [51] Y. Kim, J. Wu, S. H. Kim, P. Vasudev, J. Shen, and H. Wu (2025) Look & mark: leveraging radiologist eye fixations and bounding boxes in multimodal large language models for chest X-ray report generation. In Findings of the Association for Computational Linguistics: ACL 2025, p. 17680â17694. External Links: Document, ISBN 979-8-89176-256-5 Cited by: §2. [52] A. Koubaa (2025) From pre-trained language models to agentic AI: evolution and architectures for autonomous intelligence. Preprints. External Links: Document Cited by: §A.2, §1, §1. [53] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica (2023) Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, p. 611â626. External Links: ISBN 9798400702297, Document Cited by: Appendix B. [54] Y. Lai, J. Zhong, M. Li, S. Zhao, Y. Li, K. Psounis, and X. Yang (2025) Med-R1: reinforcement learning for generalizable medical reasoning in vision-language models. arXiv preprint arXiv:2503.13939. Cited by: §2. [55] K. Lee, S. Yoon, and H. Lim (2025) CLARIFID: improving radiology report generation by reinforcing clinically accurate impressions and enforcing detailed findings. arXiv preprint arXiv:2507.17234. Cited by: §1, §2, §4.1, §4.1. [56] S. Lee, J. Youn, H. Kim, M. Kim, and S. H. Yoon (2025) CXR-LLaVA: a multimodal large language model for interpreting chest X-ray images. European Radiology 35 (7), p. 4374â4386. External Links: Document, ISBN 1432-1084 Cited by: §2. [57] G. Li, H. A. A. K. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for âmindâ exploration of large language model society. In Thirty-seventh Conference on Neural Information Processing Systems, Cited by: §A.2, §1. [58] M. Li, H. Lin, L. Qiu, X. Liang, L. Chen, A. Elsaddik, and X. Chang (2024) Contrastive learning with counterfactual explanations for radiology report generation. In European Conference on Computer Vision, p. 162â180. External Links: ISBN 978-3-031-72774-0, Document Cited by: §4.1, Table 1, Table 2. [59] X. Li, S. Wang, S. Zeng, Y. Wu, and Y. Yang (2024) A survey on llm-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth 1 (1), p. 9. External Links: Document, ISBN 3005-060X Cited by: §A.2, §1, §1. [60] X. Li, H. Zou, and P. Liu (2025) ToRL: scaling tool-integrated RL. arXiv preprint arXiv:2503.23383. Cited by: §A.1. [61] Y. Li, Y. Liu, Z. Wang, X. Liang, L. Liu, L. Wang, and L. Zhou (2025) S-RRG-Bench: structured radiology report generation with fine-grained evaluation framework. Meta-Radiology, p. 100171. External Links: Document, ISSN 2950-1628 Cited by: §1. [62] Z. Li, H. Zhang, S. Han, S. Liu, J. Xie, Y. Zhang, Y. Choi, J. Zou, and P. Lu (2025) In-the-flow agentic system optimization for effective planning and tool use. arXiv preprint arXiv:2510.05592. Cited by: §A.2. [63] T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, S. Shi, and Z. Tu (2024) Encouraging divergent thinking in large language models through multi-agent debate. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 17889â17904. External Links: Document Cited by: §A.2. [64] C. Lin (2004) ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, p. 74â81. Cited by: §3.3, §4.1. [65] Q. Lin, Y. Zhu, B. Pu, L. Huang, H. Luo, J. Ma, Z. Peng, T. Zhao, F. Xu, J. Zhang, K. He, Z. Ou, S. Mishra, and M. Feng (2025) A foundation model for chest X-ray interpretation with grounded reasoning via online reinforcement learning. arXiv preprint arXiv:2509.03906. Cited by: §2, §4.1, Table 1. [66] C. Liu, Y. Tian, W. Chen, Y. Song, and Y. Zhang (2024) Bootstrapping large language models for radiology report generation. AAAI Conference on Artificial Intelligence 38 (17), p. 18635â18643. External Links: Document Cited by: §4.1, Table 1, Table 2. [67] F. Liu, S. Ge, and X. Wu (2021) Competence-based multimodal curriculum learning for medical report generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 3001â3012. External Links: Document Cited by: §4.1, Table 1, Table 2. [68] F. Liu, X. Wu, S. Ge, W. Fan, and Y. Zou (2021) Exploring and distilling posterior and prior knowledge for radiology report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13748â13757. External Links: Document Cited by: Table 1, Table 2. [69] G. Liu, T. H. Hsu, M. McDermott, W. Boag, W. Weng, P. Szolovits, and M. Ghassemi (2019) Clinically accurate chest X-ray report generation. In Proceedings of the 4th Machine Learning for Healthcare Conference, Proceedings of Machine Learning Research, Vol. 106, p. 249â269. Cited by: §4.2, Table 1, Table 2. [70] K. Liu, Z. Ma, X. Kang, Y. Li, K. Xie, Z. Jiao, and Q. Miao (2025) Enhanced contrastive learning with multi-view longitudinal data for chest X-ray report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10348â10359. External Links: Document Cited by: §1, §2, §4.1, Table 1, Table 1. [71] K. Liu, Z. Ma, X. Kang, Z. Zhong, Z. Jiao, G. Baird, H. Bai, and Q. Miao (2024) Structural entities extraction and patient indications incorporation for chest X-ray report generation. In proceedings of Medical Image Computing and Computer Assisted Intervention, Vol. LNCS 15003. Cited by: §4.1, Table 1. [72] R. Liu, M. Li, S. Zhao, L. Chen, X. Chang, and L. Yao (2024) In-context learning for zero-shot medical report generation. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 8721â8730. External Links: ISBN 9798400706868, Document Cited by: §4.1, §4.1, Table 1, Table 2. [73] T. Liu, J. Wang, Y. Hu, M. Li, J. Yi, X. Chang, J. Gao, and B. Yin (2025) HC-LLM: historical-constrained large language models for radiology report generation. In AAAI Conference on Artificial Intelligence, p. 5595â5603. Cited by: §2. [74] X. Liu, H. Liu, G. Yang, Z. Jiang, S. Cui, Z. Zhang, H. Wang, L. Tao, Y. Sun, Z. Song, T. Hong, J. Yang, T. Gao, J. Zhang, X. Li, J. Zhang, Y. Sang, Z. Yang, K. Xue, S. Wu, P. Zhang, J. Yang, C. Song, and G. Wang (2025) A generalist medical language model for disease diagnosis assistance. Nature Medicine 31 (3), p. 932â942. External Links: Document, ISBN 1546-170X Cited by: §2. [75] Z. Liu, Z. Zhu, S. Zheng, Y. Zhao, K. He, and Y. Zhao (2024) From observation to concept: a flexible multi-view paradigm for medical report generation. IEEE Transactions on Multimedia 26 (), p. 5987â5995. External Links: Document Cited by: §4.1, Table 1, Table 2. [76] Z. Liu, C. Chen, W. Li, P. Qi, T. Pang, C. Du, W. S. Lee, and M. Lin (2025) Understanding R1-Zero-like training: a critical perspective. In Second Conference on Language Modeling, Cited by: §A.1. [77] Z. Liu, J. Liu, Y. He, W. Wang, J. Liu, L. Pan, X. Hu, S. Xiong, J. Huang, J. Hu, S. Huang, J. Obando-Ceron, S. Yang, J. Wang, W. Su, and B. Zheng (2025) Part I: tricks or traps? a deep dive into RL for LLM reasoning. arXiv preprint arXiv:2508.08221. Cited by: §A.1. [78] J. Lou, Y. Yang, Z. Yu, Z. Fu, W. Han, Q. Huang, and J. Yu (2025) CXRAgent: director-orchestrated multi-stage reasoning for chest X-ray interpretation. arXiv preprint arXiv:2510.21324. Cited by: §1, §2. [79] C. Ma, H. Jiang, W. Chen, Y. Li, Z. Wu, X. Yu, Z. Liu, L. Guo, D. Zhu, T. Zhang, D. Shen, T. Liu, and X. Li (2024) Eye-gaze guided multi-modal alignment for medical representation learning. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: §2. [80] X. Mai, H. Xu, Z. Li, X. W, W. Wang, J. Hu, Y. Zhang, and W. Zhang (2025) Agent RL scaling law: agent rl with spontaneous code execution for mathematical problem solving. arXiv preprint arXiv:2505.07773. Cited by: §A.1. [81] J. P. Metlay, G. W. Waterer, A. C. Long, A. Anzueto, J. Brozek, K. Crothers, L. A. Cooley, N. C. Dean, M. J. Fine, S. A. Flanders, M. R. Griffin, M. L. Metersky, D. M. Musher, M. I. Restrepo, and C. G. Whitney (2019) Diagnosis and treatment of adults with community-acquired pneumonia. an official clinical practice guideline of the american thoracic society and infectious diseases society of america. American Journal of Respiratory and Critical Care Medicine 200 (7), p. e45âe67. External Links: Document Cited by: §1. [82] S. R. Motwani, C. Smith, R. J. Das, R. Rafailov, P. Torr, I. Laptev, F. Pizzati, R. Clark, and C. S. de Witt (2025) MALT: improving reasoning with multi-agent LLM training. In Second Conference on Language Modeling, Cited by: §A.2. [83] T. Moutakanni, P. Bojanowski, G. Chassagnon, C. Hudelot, A. Joulin, Y. LeCun, M. Muckley, M. Oquab, M. Revel, and M. Vakalopoulou (2024) Advancing human-centric ai for robust X-ray analysis through holistic self-supervised learning. arXiv preprint arXiv:2405.01469. Cited by: §2. [84] V. Nath, W. Li, D. Yang, A. Myronenko, M. Zheng, Y. Lu, Z. Liu, H. Yin, Y. M. Law, Y. Tang, P. Guo, C. Zhao, Z. Xu, Y. He, S. Harmon, B. Simon, G. Heinrich, S. Aylward, M. Edgar, M. Zephyr, P. Molchanov, B. Turkbey, H. Roth, and D. Xu (2025) VILA-M3: enhancing vision-language models with medical expert knowledge. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14788â14798. Cited by: §2. [85] A. Nicolson, J. Dowling, D. Anderson, and B. Koopman (2024) Longitudinal data and a semantic similarity reward for chest X-ray report generation. Informatics in Medicine Unlocked 50, p. 101585. External Links: ISSN 2352-9148, Document Cited by: §4.1, §4.1, Table 1, Table 1, Table 2, Table 2. [86] OpenAI (2024) GPT-4 Technical Report. arXiv preprint arXiv:2303.08774. Cited by: §1. [87] S. Ostmeier, J. Xu, Z. Chen, M. Varma, L. Blankemeier, C. Bluethgen, A. E. M. Md, M. Moseley, C. Langlotz, A. S. Chaudhari, and J. Delbrouck (2024) GREEN: generative radiology report evaluation and error notation. In Findings of the Association for Computational Linguistics: EMNLP 2024, p. 374â390. External Links: Document Cited by: §1, §4.1. [88] K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, p. 311â318. External Links: Document Cited by: §4.1. [89] C. Park, S. Han, X. Guo, A. E. Ozdaglar, K. Zhang, and J. Kim (2025) MAPoRL: multi-agent post-co-training for collaborative large language models with reinforcement learning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 30215â30248. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §A.2. [90] S. Park, K. Heo, D. Shin, Y. Son, J. Oh, and T. Kam (2025) DART: disease-aware image-text alignment and self-correcting re-alignment for trustworthy radiology report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 15580â15589. External Links: Document Cited by: §1, §2, §4.1, Table 1. [91] C. Pellegrini, E. Ăzsoy, B. Busam, N. Navab, and M. Keicher (2025) RaDialog: a large vision-language model for radiology report generation and conversational assistance. arXiv preprint arXiv:2311.18681. Cited by: §2. [92] A. Plaat, M. van Duijn, N. van Stein, M. Preuss, P. van der Putten, and K. J. Batenburg (2025) Agentic large language models, a survey. arXiv preprint arXiv:2503.23037. Cited by: §A.2, §1, §1. [93] C. Qian, E. C. Acikgoz, Q. He, H. Wang, X. Chen, D. Hakkani-TĂźr, G. Tur, and H. Ji (2025) ToolRL: reward is all tool learning needs. arXiv preprint arXiv:2504.13958. Cited by: §A.1. [94] A. Sellergren, S. Kazemzadeh, T. Jaroensri, A. Kiraly, M. Traverse, T. Kohlberger, S. Xu, F. Jamil, C. Hughes, C. Lau, J. Chen, F. Mahvar, L. Yatziv, T. Chen, B. Sterling, S. A. Baby, S. M. Baby, J. Lai, S. Schmidgall, L. Yang, K. Chen, P. Bjornsson, S. Reddy, R. Brush, K. Philbrick, M. Asiedu, I. Mezerreg, H. Hu, H. Yang, R. Tiwari, S. Jansen, P. Singh, Y. Liu, S. Azizi, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. RamĂŠ, M. Riviere, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Buchatskaya, J. Alayrac, D. Lepikhin, V. Feinberg, S. Borgeaud, A. Andreev, C. Hardin, R. Dadashi, L. Hussenot, A. Joulin, O. Bachem, Y. Matias, K. Chou, A. Hassidim, K. Goel, C. Farabet, J. Barral, T. Warkentin, J. Shlens, D. Fleet, V. Cotruta, O. Sanseviero, G. Martins, P. Kirk, A. Rao, S. Shetty, D. F. Steiner, C. Kirmizibayrak, R. Pilgrim, D. Golden, and L. Yang (2025) MedGemma technical report. arXiv preprint arXiv:2507.05201. Cited by: §1, §4.1, Table 1, Table 1, Table 2, Table 2. [95] V. Shah, Y. R. Chillakuru, A. Rybkin, Y. Seo, T. Vu, and J. H. Sohn (2022) Algorithmic prediction of delayed radiology turn-around-time during non-business hours. Academic Radiology 29 (5), p. e82âe90. External Links: Document, ISSN 1076-6332 Cited by: §1. [96] Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo (2024) DeepSeekMath: pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300. Cited by: §A.1, §3.1. [97] H. Shen, M. Pei, J. Liu, and Z. Tian (2024) Automatic radiology reports generation via memory alignment network. AAAI Conference on Artificial Intelligence 38 (5), p. 4776â4783. External Links: Document Cited by: §4.1, §4.1, Table 1, Table 2. [98] G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu (2025) HybridFlow: a flexible and efficient rlhf framework. In Proceedings of the Twentieth European Conference on Computer Systems, p. 1279â1297. External Links: ISBN 9798400711961, Document Cited by: Appendix B. [99] J. Singh, R. Magazine, Y. Pandya, and A. Nambi (2025) Agentic reasoning and tool integration for LLMs via reinforcement learning. arXiv preprint arXiv:2505.01441. Cited by: §A.1. [100] I. SĂŽrbu, I. SĂŽrbu, J. Bogojeska, and T. Rebedea (2025) GIT-CXR: end-to-end transformer for chest X-ray report generation. Information 16 (7). External Links: Document, ISSN 2078-2489 Cited by: §1, §2. [101] A. Smit, S. Jain, P. Rajpurkar, A. Pareek, A. Ng, and M. Lungren (2020) Combining automatic labelers and expert annotations for accurate radiology report labeling using BERT. In Empirical Methods in Natural Language Processing, p. 1500â1519. External Links: Document Cited by: §1, §3.3, §4.1. [102] H. Song, J. Jiang, Y. Min, J. Chen, Z. Chen, W. X. Zhao, L. Fang, and J. Wen (2025) R1-Searcher: incentivizing the search capability in LLMs via reinforcement learning. arXiv preprint arXiv:2503.05592. Cited by: §A.1. [103] Y. Su, D. Yu, L. Song, J. Li, H. Mi, Z. Tu, M. Zhang, and D. Yu (2025) Crossing the reward bridge: expanding rl with verifiable rewards across diverse domains. arXiv preprint arXiv:2503.23829. Cited by: §A.1. [104] T. Tanida, P. MĂźller, G. Kaissis, and D. Rueckert (2023) Interactive and explainable region-guided radiology report generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 7433â7442. External Links: Document Cited by: §4.1, Table 1. [105] R. Tanno, D. G. T. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, C. Lau, T. Tu, S. Azizi, K. Singhal, M. Schaekermann, R. May, R. Lee, S. Man, S. Mahdavi, Z. Ahmed, Y. Matias, J. Barral, S. M. A. Eslami, D. Belgrave, Y. Liu, S. R. Kalidindi, S. Shetty, V. Natarajan, P. Kohli, P. Huang, A. Karthikesalingam, and I. Ktena (2025) Collaboration between clinicians and visionâlanguage models in radiology report generation. Nature Medicine 31 (2), p. 599â608. External Links: Document, ISBN 1546-170X Cited by: §4.1, §4.2, Table 1, Table 1, Table 2. [106] G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. RamĂŠ, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. GyĂśrgy, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. Klimczak-PluciĹska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. PĂľder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot (2025) Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Cited by: §1. [107] O. C. Thawakar, A. M. Shaker, S. S. Mullappilly, H. Cholakkal, R. M. Anwer, S. Khan, J. Laaksonen, and F. Khan (2024) XrayGPT: chest radiographs summarization using large medical vision-language models. In Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, p. 440â448. External Links: Document Cited by: §2. [108] A. Wang, Z. Zhang, D. Wang, F. Wang, H. Hu, J. Guo, Y. Zhou, C. Pang, and S. Wen (2025) Overcoming heterogeneous data in federated medical vision-language pre-training: a triple-embedding model selector approach. AAAI Conference on Artificial Intelligence 39 (7), p. 7500â7508. External Links: Document Cited by: §2. [109] F. Wang, S. Du, and L. Yu (2025) HERGen: elevating radiology report generation with longitudinal data. In European Conference on Computer Vision, p. 183â200. External Links: ISBN 978-3-031-73001-6 Cited by: §4.1, Table 1. [110] L. Wang, H. Wang, H. Yang, J. Mao, Z. Yang, J. Shen, and X. Li (2025) Interpretable bilingual multimodal large language model for diverse biomedical tasks. In The Thirteenth International Conference on Learning Representations, Cited by: §4.1, Table 1. [111] S. Wang, C. Chen, X. Le, Q. Xu, L. Xu, Y. Zhang, and J. Yang (2025) CAD-GPT: synthesising cad construction sequence with spatial reasoning-enhanced multimodal llms. AAAI Conference on Artificial Intelligence 39 (8), p. 7880â7888. External Links: Document Cited by: §2. [112] X. Wang, F. Wang, Y. Li, Q. Ma, S. Wang, B. Jiang, and J. Tang (2025) CXPMRG-Bench: pre-training and benchmarking for X-ray medical report generation on chexpert plus dataset. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 5123â5133. External Links: Document Cited by: §1. [113] Z. Wang, L. Liu, L. Wang, and L. Zhou (2023) METransformer: radiology report generation by transformer with multiple learnable expert tokens. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 11558â11567. External Links: Document Cited by: §4.1, §4.1, Table 1, Table 2. [114] Z. Wang, L. Liu, L. Wang, and L. Zhou (2023) R2GenGPT: radiology report generation with frozen LLMs. Meta-Radiology 1 (3), p. 100033. External Links: Document, ISSN 2950-1628 Cited by: §2. [115] Z. Wang, K. Lee, Q. Deng, T. Y. So, W. H. Chiu, B. Zhou, and E. S. Hui (2024) Expert insight-enhanced follow-up chest X-ray summary generation. In Artificial Intelligence in Medicine, p. 181â193. External Links: Document, ISBN 978-3-031-66534-9 Cited by: §1, §2. [116] Z. Wang, S. Yan, K. Yin, X. Zhang, and W. K. Cheung (2025) CURV: coherent uncertainty-aware reasoning in vision-language models for X-ray report generation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, Cited by: §2. [117] Z. Wang, Z. Wang, B. Srinivasan, V. N. Ioannidis, H. Rangwala, and R. Anubhai (2024) BioBridge: bridging biomedical foundation models via knowledge graphs. In International Conference on Learning Representations, Cited by: §2. [118] X. Wen, Z. Liu, S. Zheng, S. Ye, Z. Wu, Y. Wang, Z. Xu, X. Liang, J. Li, Z. Miao, J. Bian, and M. Yang (2025) Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs. arXiv preprint arXiv:2506.14245. Cited by: §A.1. [119] Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang (2024) AutoGen: enabling next-gen LLM applications via multi-agent conversations. In First Conference on Language Modeling, Cited by: §A.2, §1. [120] T. Xiao, L. Shi, P. Liu, Z. Wang, and C. Bai (2025) Radiology report generation via multi-objective preference optimization. In AAAI Conference on Artificial Intelligence, p. 8664â8672. Cited by: §2, §4.1, §4.1, Table 1, Table 2. [121] T. Xiao, L. Shi, Y. Zhang, H. Yang, Z. Wang, and C. Bai (2025) Online iterative self-alignment for radiology report generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 27799â27814. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §1, §2, §4.1, Table 1, Table 2. [122] B. Yan, R. Liu, D. Kuo, S. Adithan, E. Reis, S. Kwak, V. Venugopal, C. OâConnell, A. Saenz, P. Rajpurkar, and M. Moor (2023) Style-aware radiology report generation with RadGraph and few-shot prompting. In Findings of the Association for Computational Linguistics: EMNLP 2023, p. 14676â14688. External Links: Document Cited by: §1, §2, §4.1, Table 1. [123] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025) Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §1. [124] J. Yang, B. Su, X. Zhao, and J. Wen (2024) Unlocking the power of spatial and temporal information in medical multimodal pre-training. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, p. 56382â56396. Cited by: §2. [125] L. Yang, Z. Ni, Y. Wen, Y. Liu, L. He, and H. T. Shen (2025) Self-supervised anatomical consistency learning for vision-grounded medical report generation. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 2958â2967. External Links: Document, ISBN 9798400720352 Cited by: §2. [126] Z. Yi, T. Xiao, and M. V. Albert (2025) A multimodal multi-agent framework for radiology report generation. arXiv preprint arXiv:2505.09787. Cited by: §1, §2, Table 2. [127] H. Yin, S. Zhou, P. Wang, Z. Wu, and Y. Hao (2025) KIA: knowledge-guided implicit vision-language alignment for chest X-ray report generation. In Proceedings of the 31st International Conference on Computational Linguistics, p. 4096â4108. Cited by: §2. [128] K. You, J. Gu, J. Ham, B. Park, J. Kim, E. K. Hong, W. Baek, and B. Roh (2023) CXR-CLIP: toward large scale chest X-ray language-image pre-training. In Medical Image Computing and Computer Assisted Intervention â MICCAI 2023, p. 101â111. External Links: ISBN 978-3-031-43895-0 Cited by: §2. [129] F. Yu, M. Endo, R. Krishnan, I. Pan, A. Tsai, E. P. Reis, E. K. U. N. Fonseca, H. M. H. Lee, Z. S. H. Abad, A. Y. Ng, C. P. Langlotz, V. K. Venugopal, and P. Rajpurkar (2023) Evaluating progress in automatic chest X-ray radiology report generation. Patterns 4 (9), p. 100802. External Links: Document, ISSN 2666-3899 Cited by: §2, §4.2, Table 1, Table 2. [130] Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, W. Dai, T. Fan, G. Liu, L. Liu, X. Liu, H. Lin, Z. Lin, B. Ma, G. Sheng, Y. Tong, C. Zhang, M. Zhang, W. Zhang, H. Zhu, J. Zhu, J. Chen, J. Chen, C. Wang, H. Yu, Y. Song, X. Wei, H. Zhou, J. Liu, W. Ma, Y. Zhang, L. Yan, M. Qiao, Y. Wu, and M. Wang (2025) DAPO: an open-source LLM reinforcement learning system at scale. arXiv preprint arXiv:2503.14476. Cited by: §A.1. [131] J. M. Zambrano Chaves, S. Huang, Y. Xu, H. Xu, N. Usuyama, S. Zhang, F. Wang, Y. Xie, M. Khademi, Z. Yang, H. Awadalla, J. Gong, H. Hu, J. Yang, C. Li, J. Gao, Y. Gu, C. Wong, M. Wei, T. Naumann, M. Chen, M. P. Lungren, A. Chaudhari, S. Yeung-Levy, C. P. Langlotz, S. Wang, and H. Poon (2025) A clinically accessible small multimodal radiology model and evaluation metric for chest X-ray findings. Nature Communications 16 (1), p. 3108. External Links: Document, ISBN 2041-1723 Cited by: §2. [132] J. Zhang, J. Xi, Z. Song, J. Lu, Y. Ke, T. Sun, Y. Yang, J. Zhang, S. Zhang, and Z. Xie (2025) L0: reinforcement learning to become general agents. arXiv preprint arXiv:2506.23667. Cited by: §A.1. [133] K. Zhang, C. D. Barrett, J. Kim, L. Sun, T. Taghavi, and K. Kenthapadi (2025) RadAgents: multimodal agentic reasoning for chest X-ray interpretation with radiologist-like workflows. arXiv preprint arXiv:2509.20490. Cited by: §1, §2, §4.1. [134] K. Zhang, R. Zhou, E. Adhikarla, Z. Yan, Y. Liu, J. Yu, Z. Liu, X. Chen, B. D. Davison, H. Ren, J. Huang, C. Chen, Y. Zhou, S. Fu, W. Liu, T. Liu, X. Li, Y. Chen, L. He, J. Zou, Q. Li, H. Liu, and L. Sun (2024) A generalist visionâlanguage foundation model for diverse biomedical tasks. Nature Medicine 30 (11), p. 3129â3141. External Links: Document, ISBN 1546-170X Cited by: §2. [135] X. Zhang, Z. Meng, J. Lever, and E. S. L. Ho (2025) Libra: leveraging temporal images for biomedical radiology analysis. In Findings of the Association for Computational Linguistics: ACL 2025, p. 17275â17303. External Links: Document, ISBN 979-8-89176-256-5 Cited by: §2. [136] X. Zhang, Y. Shi, J. Ji, C. Zheng, and L. Qu (2025) MEPNet: medical entity-balanced prompting network for brain ct report generation. AAAI Conference on Artificial Intelligence 39 (24), p. 25940â25948. External Links: Document Cited by: §2. [137] B. Zhao, L. G. Foo, P. Hu, C. Theobalt, H. Rahmani, and J. Liu (2025) LLM-based agentic reasoning frameworks: a survey from methods to scenarios. arXiv preprint arXiv:2508.17692. Cited by: §A.2, §1, §1. [138] Y. Zhao, A. Gu, R. Varma, L. Luo, C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li (2023) PyTorch FSDP: experiences on scaling fully sharded data parallel. Proceedings of the VLDB Endowment 16 (12), p. 3848â3860. External Links: ISSN 2150-8097, Document Cited by: Appendix B. [139] C. Zheng, S. Liu, M. Li, X. Chen, B. Yu, C. Gao, K. Dang, Y. Liu, R. Men, A. Yang, J. Zhou, and J. Lin (2025) Group sequence policy optimization. arXiv preprint arXiv:2507.18071. Cited by: §A.1, Appendix B, §3.1. [140] Z. Zhou, M. Shi, M. Wei, O. Alabi, Z. Yue, and T. Vercauteren (2024) Large model driven radiology report generation with clinical quality reinforcement learning. arXiv preprint arXiv:2403.06728. Cited by: §1, §2, §2, §4.1, Table 1, Table 2. [141] Y. Zou and Z. Yin (2025) MVCM: enhancing multi-view and cross-modality alignment for medical visual question answering and medical image-text retrieval. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, p. 180â190. External Links: Document Cited by: §2. Appendix A Extended related work A.1 Reinforcement learning for LLMs Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as an alternative to traditional Reinforcement Learning from Human Feedback (RLHF), providing objective, outcome-based feedback through deterministic verification functions rather than human preference models. By rewarding correctness or rule-based validity, RLVR has improved reasoning capabilities in structured domains, such as mathematics and code generation [118, 103, 96, 130, 139, 76, 77]. More recently, agentic reinforcement learning has gained attention, where models interact with external tools such as Python interpreters or web search engines to improve factual accuracy [99, 80, 27, 132, 29, 93, 60, 102, 14]. However, most of these methods still focus on training a single, monolithic model, rather than jointly optimizing multiple agents. Consequently, genuine cooperative optimization among agents in real-world workflows remains largely unexplored. A.2 Agentic systems with LLMs Agentic systems leverage LLMs to perform goal-oriented reasoning, planning, and collaboration through structured interactions among multiple agents. Recent frameworks such as AutoGen [119], MetaGPT [33], CAMEL [57], have shown that pretrained LLMs can engage in cooperative behaviors via carefully designed prompts and workflows [37, 4, 137, 92, 52, 44, 59, 63]. However, most of these systems are training-free, relying on pre-trained models without end-to-end optimization, resulting in suboptimal coordination and limited adaptability when applied to complex, real-world workflows. Recent works have started to explore reinforcement learning within agentic systems. For example, MALT [82] performs post-training using trajectories collected from agent executions, but its optimization remains off-policy and detached from real-world workflow interactions. MAPoRL [89] applies multi-agent reinforcement learning to improve collaboration among language models, yet the optimization of realistic, role-structured workflows in applied domains remains largely unexplored. To address this gap, AgentFlow [62] introduces on-policy reinforcement learning within an actual multi-agent workflow, but the optimization is limited to a single key planner agent rather than the entire agentic system. Appendix B Detailed experimental setup SectionË4.1 provides an overview of the experimental setup. In this section, we describe additional details that were not included in SectionË4.1. Datasets MIMIC-CXR [48] is a large-scale publicly available chest X-ray dataset containing more than 370,000 images paired with free-text radiology reports. The dataset covers a wide range of thoracic conditions and includes both frontal and lateral views. IU X-ray (Indiana University Chest X-Ray dataset) [25] consists of chest X-ray images paired with corresponding radiology reports, including both Findings and Impression sections. Each study contains frontal and lateral views with detailed narrative annotations. Implementation details The reinforcement learning implementation is built on top of verl [98], using vLLM [53] for rollouts and Fully Sharded Data Parallel (FSDP) [138] for parameter updates. We set the batch size to 16, and 16 rollouts are generated for each training sample. The maximum rollout length is set to 2048 tokens. We train each model for 100 reinforcement learning update steps. Following the original GSPO paper [139], we set the clipping thresholds in objectives to Îľhigh=0.0004 _high=0.0004 and Îľlow=0.0003 _low=0.0003. The rollout temperature is fixed at 1.0. We use a learning rate of 1e-6, a warmup ratio of 0.05, and a weight decay of 0.1. We use AdamW as the optimizer. The batch size for parameter updates is set to 4. We use google/medgemma-4b-it111https://huggingface.co/google/medgemma-4b-it as the base checkpoint. We use verl v0.5.0, vLLM v0.10.1, Transformers v4.53.3, and PyTorch v2.7.1. All experiments are conducted on 4 Ă NVIDIA H100 GPUs, each with 80 GB of memory. Each training run took approximately 18 hours. For inference, we follow the default vLLM generation settings except that the maximum number of generated tokens is set to 2048. Appendix C Computational cost and deployment considerations To assess the practical deployment cost of MARL-Rad, we measured inference latency and GPU memory usage on H100 GPUs using 600 reports from the MIMIC-CXR test set. When the three region-specific agents were executed in parallel, MARL-Rad required 772 seconds in total, corresponding to 1.28 seconds per sample. When all agents were executed sequentially, the total inference time was 1586 seconds, corresponding to 2.64 seconds per sample. This latency is identical to that of the training-free agentic baseline used in our ablation study, i.e., âMedGemma (agent),â because both methods share the same multi-agent architecture at inference time. For reference, the non-agent baseline, âMedGemma (vanilla),â required 398 seconds in total, corresponding to 0.66 seconds per sample. The GPU memory usage was approximately 9 GB per agent. Although MARL-Rad introduces additional computational overhead due to agentization, the region-specific agents can be executed in parallel, substantially reducing latency compared with sequential execution. The resulting inference speed remains within a practical range for real-world radiology reporting workflows. Appendix D Details of the laterality-specific RadGraph F1 score computation To evaluate region-level consistency focused on laterality, we compute a laterality-specific RadGraph F1 score derived from the standard RadGraph evaluation. We first construct the RadGraph for both the prediction and the ground truth using the standard RadGraph extraction procedure. From each constructed graph, we then extract the entities corresponding to laterality-related anatomical regions, which include the tokens âleftâ, ârightâ, âleft lungâ, and âright lung.â From these entities, we further extract the subgraph consisting of all nodes and relations connected to them, resulting in a laterality-specific subset of each RadGraph. We then calculate the laterality-specific RadGraph F1 score by applying the standard RadGraph matching procedure but restrict the comparison to this laterality-specific subgraph only. Appendix E Clinician evaluation criteria Physicians evaluated each report pair according to the following five criteria. ⢠Completeness: Whether the report includes all clinically relevant findings visible in the chest X-ray, without omitting important abnormalities or normal findings that should be documented. ⢠Correctness: Whether the described findings are clinically accurate and consistent with the chest X-ray, without introducing incorrect diagnoses, false findings, or contradictions. ⢠Conciseness: Whether the report is appropriately concise, avoiding unnecessary repetition, irrelevant descriptions, or overly verbose statements while preserving clinically important information. ⢠Readability: Whether the report is clearly written, well organized, and easy for clinicians to understand, with coherent phrasing and appropriate radiological terminology. ⢠Clinical Utility: Whether the report would be useful in clinical practice for supporting diagnosis, communication, and downstream decision-making. Physicians were instructed to compare two anonymized reports for each chest X-ray case and select which report was better for each criterion. The report order was randomized, and the physicians were blinded to whether each report was generated by MARL-Rad or taken from the ground truth. The evaluation was conducted with approval from the affiliated hospital according to the applicable institutional requirements. Potential risks to participants were minimal, as the task involved expert assessment of anonymized reports without identifiable patient information. Physicians performed the evaluation during regular working hours under the affiliated hospitalâs approval, and no additional honorarium was provided specifically for this task. Appendix F Limitations and broader impact F.1 Limitations Although MARL-Rad achieves strong performance on standard RRG benchmarks, this work has several limitations. First, our experiments focus on chest X-ray report generation using MIMIC-CXR and IU X-ray, and the effectiveness of the proposed framework on other imaging modalities, institutions, and clinical settings remains to be validated. Second, our reinforcement learning rewards rely on automatic clinical metrics such as CheXbert and RadGraph. While these metrics are clinically motivated and widely used, they may not fully capture all aspects of report quality and may contain annotation or evaluation noise. We partially address this concern through blinded clinician evaluation, but the evaluation is limited in scale and should be expanded in future work. Third, MARL-Rad requires multiple model invocations at inference time, increasing computational cost and GPU memory usage compared with a single-model baseline, although parallel execution mitigates the added latency as discussed in AppendixËC. Finally, although our framework represents a step toward clinically useful agentic RRG systems by improving clinical efficacy metrics and laterality consistency, real-world deployment would require rigorous prospective validation, safety monitoring, and integration into radiologist-supervised workflows. F.2 Broader impact MARL-Rad aims to improve radiology report generation by training agentic systems within the workflows in which they are deployed. If properly validated, such systems could help reduce the burden of report writing, improve the consistency of generated reports, and support radiologists in clinical documentation. At the same time, medical report generation is a high-stakes application. Incorrect or hallucinated findings could negatively affect clinical decision-making if used without appropriate oversight. Future deployment should require careful clinical validation, bias assessment across patient populations and institutions, privacy-preserving data handling, and mechanisms for human review and correction.