Paper deep dive
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/20/2026, 5:22:14 AM
Summary
The paper introduces MedUAG, a unified medical multimodal model for understanding and generation (UAG). It addresses gaps in the medical domain by presenting MedUAGCorpus (a large-scale dataset with 6M+ instances across 14 modalities) and MedUAGBench (a benchmark with 12 generation tasks). MedUAG is an end-to-end model trained on these resources, demonstrating strong performance in both understanding (VQA, MRG) and generation (synthesis, translation, reconstruction, prediction) tasks compared to existing baselines.
Entities (14)
Relation Signals (12)
MedUAGCorpus â containsinstances â 6 million
confidence 95% · comprising over 6 million instances across 14 imaging modalities
MedUAGBench â coverstasks â 12
confidence 95% · expands medical generation evaluation to 12 diverse tasks
MedUAG â evaluatedon â MedUAGBench
confidence 95% · we introduce MedUAGBench... Extensive experiments demonstrate that MedUAG achieves strong performance
MedUAG â supportstask â Medical Report Generation
confidence 95% · On the understanding side, it covers VQA and MRG.
MedUAG â supportstask â Visual Question Answering
confidence 95% · On the understanding side, it covers VQA and MRG.
MedUAG â uses â MedUAGCorpus
confidence 95% · leveraging these resources, we develop MedUAG... MedUAGCorpus... containing over 6M instances
MedUAG â outperforms â UniMedVL
confidence 90% · Table II and III show MedUAG achieving higher scores in various metrics compared to UniMedVL
MedUAG â outperforms â HealthGPT
confidence 90% · Table II and III show MedUAG achieving higher scores in various metrics compared to HealthGPT
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.
Tags
Links
- Source: https://arxiv.org/abs/2608.18937v1
- Canonical: https://arxiv.org/abs/2608.18937v1
Trouble viewing inline? Open PDF directly â
Full Text
46,786 characters extracted from source content.
Expand or collapse full text
MedUAG: Unified Understanding and Generation for Medical Multimodal Models Zijie Meng 1,â , Yuncheng Zhang 1,â , Hualiang Wang 2,â , Yitian Tang 1 , Xiaotang Gai 1 , Chen Shen 1 , Songtao Jiang 1 , Shaosheng Cao 3 , Jian Wu 1 , Xian Wu 4,â , Zuozhu Liu 1,â 1 Zhejiang University, 2 Hong Kong University of Science and Technology, 3 Tsinghua University, 4 Tencent Jarvis Lab AbstractâRecentMultimodalLargeLanguageModels (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems. Index TermsâMedical Multimodal Learning; Unified Under- standing and Generation; Benchmark and Dataset I. INTRODUCTION Recent Multimodal Large Language Models (MLLMs) have evolved from isolated understanding or generation systems into unified understanding and generation (UAG) frameworks across heterogeneous modalities [1], as shown in Figure 1. Driven by the scaling and integration of data across diverse UAG tasks, these unified training paradigms have demon- strated the potential to learn universal representations within a single system, substantially reducing the reliance on task- specific fine-tuning for downstream applications [2]â[6]. However, the few pioneering works in medical UAG [7], [8] are hindered by two significant limitations. (1) The absence of comprehensive UAG training corpora and evaluation benchmarks. While training and evaluating MLLMs neces- sitate large-scale data across diverse tasks, current studies often rely on a limited scope of medical imaging modalities and clinical applications. They resort to a narrow selection of common or readily accessible tasks, such as basic text- conditioned image synthesis, CT-MRI translation or MRI super resolution. Consequently, some challenging generation tasks of profound clinical significance remain overlooked. (2) The lack of a broadly validated unified medical model. Beyond data and benchmarks, current medical UAG studies â Equal contribution. â Corresponding authors. TargetImages Generation Model Source Images Prompt Answers Understanding Model Medical Images Questions Answers TargetImages Unified Model (Ours) Medical Images Questions Source Images Prompt Fig. 1. Comparison of different medical multimodal modeling paradigms. Our framework unifies understanding and generation within a single model. have yet to establish a unified model that is trained on large- scale and diverse unified medical corpora and systematically validated across comprehensive understanding and generation benchmarks. Although recent approaches [7], [8] have shown promising initial results, their empirical validation remains limited in task breadth, modality coverage, and evaluation consistency. As a result, it remains unclear whether a sin- gle end-to-end model can robustly handle diverse medical multimodal demands, ranging from semantic understanding and clinical reasoning to structurally faithful and clinically meaningful image generation. This limitation not only leaves the generality and robustness of unified medical modeling insufficiently demonstrated, but also makes it difficult to assess its potential for practical application. To address these gaps, we introduce MedUAGBench, a comprehensive benchmark for unified medical generation. MedUAGBench extend the 5 tasks provided by UniMedVL [8] to 12 tasks in terms of the medical image generation, includ- ing various modalities under standardized prompts, metrics, and evaluation settings, enabling systematic and reproducible assessment of medical generative capability. In parallel, we construct MedUAGCorpus, the largest unified medical under- standing and generation dataset to date, containing over 6M instances across 14 imaging modalities and diverse task types. Together, these resources provide the large-scale and diverse supervision foundation needed to support the comprehensive training and systematic validation of unified medical models, as summarized in Table I and Figure 2. Building upon these resources, we develop MedUAG, an end-to-end model for unified medical understanding and gen- arXiv:2608.18937v1 [cs.CL] 19 Aug 2026 TABLE I COMPARISON WITH EXISTING UNIFIED MEDICAL UNDERSTANDING AND GENERATION WORKS. Modalities in Generation UnderstandingGeneration Tasks (#) Data Scale VQAMRGSynthesisTranslationReconstructionPrediction HealthGPT [7]11â122â1.5M UniMedVL [8]10â21115.6M MedUAG (ours)14â33426.4M Generation Synthe sis Reconstruction Predic - tion Generate an ultrasound image of the liver showing the following features: The liver is ... Generate an abdominal CT slice driven solely by the mask labels: red=liver, green=kidney... Imagine the patient â s condition changed as described. Edit the source X-ray to show: There is ... From this 3T T1-weighted sagittal scan of the brain, create its FLAIR counterpart. Create the corresponding CT image for this MRI of the thorax. This input is a mask of the retinal blood vessels.Produce a fundus image... Create a high-quality, full-dose equivalent CT image from this low- quality, axial low-dose chest scan source. Transform this noisy coronal PET MIP, representing a 1/50 fraction of the standard dose, into a standard- dose projection. Generate a clean, high- quality 3T MRI image of the brain based on this T2-weighted axial low-field (64mT) scan. Given this 4x under- sampled sagittal MRI scan of a knee, remove the artifacts and restore the image to full-sampled quality. Generate a 2D dose map for the provided CT slice. The model should predict the radiation dose in Gray (Gy) for PTV56 and brainstem. Convert this H&E stained image to IHC stained image. Text - based Mask - based Counterfactual Synthesis MRI MRI MRI CT Retinal Vessel Fundus Image Low-Dose CT Denoising Low-Dose PET Denoising Low-Field MRI Enhancem ent Under- sampled MRI Artifact Removal Radiotherapy Dose Map Prediction H&E IHC Translation Understanding Visual Question Answering (VQA) Medical Report Generation (MRG) Question: is the opacity located in the middle of the image inside the patient or superficial to the patient's skin? Answer: superficial to the patient's skin Please generate a medical report for this image in the format of Findings and Impression. Make it as detailed as possible. The trachea is midline. The cardiomediastinal silhouette is normal ... No evidence of pneumothorax ... Vague density in the medial right lung apex most representing overlying shadows of bony structures, which is stable. MedUAG: Unified Understanding and Generation forMedical Multimodal Models Fig. 2. Overview of MedUAG. Our unified framework supports both medical understanding and generation. On the understanding side, it covers VQA and MRG. On the generation side, it supports four task categories: synthesis, translation, reconstruction, and prediction across diverse medical modalities. Representative demonstrations of each task type are provided. eration. Extensive experiments show that MedUAG achieves strong performance across a broad range of tasks, establish- ing a competitive baseline for future research. Overall, our work provides a unified foundation spanning benchmark, data corpus, and model, which takes a step toward more general and practically useful next-generation medical multi-modal systems. Our main contributions are as follows: âą We introduce MedUAGBench, a comprehensive bench- mark for unified medical generation, covering 12 tasks and various modalities under standardized evaluation protocols. âą We construct MedUAGCorpus, a large-scale unified un- derstanding generation corpus with over 6M instances across 14 medical modalities and diverse task types. âą We develop MedUAG, an end-to-end unified medical model, and demonstrate strong performance across a wide range of understanding and generation tasks. I. METHODS A. Definition of Tasks As illustrated in Figure 2, we study unified medical mul- timodal modeling through two primary task domains: under- standing and generation. 1) Understanding Tasks: Understanding tasks evaluate the modelâs ability to interpret medical images and extract clini- cally relevant semantics. We consider two primary tasks: (i) Visual Question Answering (VQA), which requires answering natural language questions grounded in specific medical im- ages and evaluates localized diagnostic reasoning and response precision; and (i) Medical Report Generation (MRG), which requires generating a complete clinical report from medical images. Unlike VQA, MRG requires the model to identify clinically significant findings and summarize them in a struc- tured and coherent form. MixtureofTransformerExperts E gen E und Text Tokenizer LM Head D gen Next Token PredictionVelocity Prediction â Your task is radiotherapy dose prediction. Given the 2D CT slice...The goal is to produce a plan...", âWhat is the specific type of cancer present in the image?â âAdenocarcinoma of the left lower lobe.â (a) Constructionpipeline of dataset.(b) Model architecture UnifiedSample Standardization Padding & Resizing MetadataPacking SourceData Curation Dataset Collection ManualReview & Organization kaggle 512Ă512; only for Gen Task-SpecificConstruction Data Alignment Source EmulationTarget Generation Label Restructuring Image Registration Noise Injection Undersampling (Und; Task 1,2,3,6,11,12) (Task 4,5) (Task 7,8) (Task 10) Slicing & Sampling (Task 4,5,7-10) Mask Colorization (Task 2) Dose Map Rendering (Task 11) Fig. 3. The construction process of datasets and model architecture of MedUAG. The dataset construction pipeline includes source data curation, task-specific construction, and unified sample standardization. The model adopts dual image encoders and mixture of transformer experts to support both medical image understanding and generation within a single framework. 2) Generation Tasks: Generation tasks evaluate the modelâs ability to produce medically meaningful images under di- verse conditions. We categorize them into four paradigms. (i) Synthesis creates new medical images or edits existing ones under semantic guidance, including text-based, mask- based, and counterfactual settings. This requires the model to capture anatomical structure and pathological semantics, enabling plausible image generation or clinically specified image modification. Such capability is useful for data aug- mentation, medical education, and visualization of potential disease progression. (i) Translation transforms images across domains while preserving the underlying anatomical content. It includes intra-modality translation, inter-modality transla- tion, and structure-to-image generation, evaluating whether the model can preserve shared structures while adapting domain- specific appearance. Clinically, it can synthesize comple- mentary views or modalities from existing scans, enriching diagnostic information without additional acquisition. (i) Reconstruction recovers high-quality images from degraded inputs. In our dataset, this category includes low-dose CT de- noising, low-dose PET denoising, low-field MRI enhancement, and undersampled MRI restoration. Successful reconstruction requires preserving diagnostically relevant details while reduc- ing acquisition-induced distortions, supporting more efficient and lower-risk imaging protocols. (iv) Prediction generates clinically informative outputs beyond direct restoration or cross-domain translation. This category includes radiotherapy dose distribution prediction from CT images with annotated target volumes and organs at risk, as well as H&E-to-IHC pathological stain conversion. Compared with other generation tasks, prediction emphasizes outputs with more direct clinical relevance and application potential. B. Construction of Datasets As illustrated in Figure 3, we construct a systematic and extensible pipeline to build the unified corpus and benchmark. 1) Source Data Curation: We first collect public datasets according to the task taxonomy, and then manually review and normalize them into a unified data pool for task-aware processing. For understanding tasks, we mainly use medical multimodal data from Hulu-Med [9]. For generation tasks, we collect public datasets covering synthesis, translation, reconstruction, and prediction [10]â[29]. After collection, all datasets are manually inspected to remove unusable samples and harmonize file structures and annotation formats, resulting in normalized raw data for subsequent construction. 2) Task-Specific Construction: Following curation, task- specific construction further exploits the information available in each dataset and derives samples that match different downstream scenarios. This stage consists of three compo- nents: (i) Data alignment includes label restructuring and image registration, which adapt existing annotations to task- specific requirements and, when necessary, align images of the same anatomical structures across different modalities. (i) Source emulation simulates degraded acquisition condi- tions, such as noise injection and undersampling. (i) Tar- get generation performs task-dependent operations, including slicing and sampling volumetric data into 2D images, mask colorization for mask-based synthesis, and dose map rendering for radiotherapy dose prediction. Through these procedures, heterogeneous raw data are transformed into unified input- output pairs for different tasks. 3) Unified Sample Standardization: Finally, all processed samples are standardized into a unified format and split into the training corpus and benchmark at the patient level. Specifically, all images, slices, and derived samples from the same patient are assigned exclusively to either the training corpus or the benchmark, preventing data leakage between training and evaluation. Furthermore, to satisfy the fixed input size required by some generative models, images in generation tasks are resized and padded to 512Ă 512 while preserving the original anatomical aspect ratio as much as possible. Each (b) MedUAGBench 2.- 3(.# -$- -%- 3(.# -$- )/(. ,!./& 3(.# -$- $).# ,*3 )- * , $.$)( .) )&),$4.$)( ( ,-'*& ,.$!. ')0& )1$ & (#( ' (. )1)- ()$-$(" )1)- ()$-$(" -- &.)/(/- ,(-&.$)( ( ,(-&.$)( +/ ( ,(-&.$)( #! ! !"! ! (a) MedUAGCorpus Fig. 4. Statistical overview of dataset. (a) MedUAGCorpus: distribution of samples across anatomical systems and imaging modalities, together with the composition of the two training stages, including MedUAG-Corpus Align and MedUAG-Corpus SFT. âT2Iâ, âI2Tâ, âReconâ, âUndâ and âGenâ indicate text- to-image, image-to-text, reconstruction, understanding and generation, respectively. (b) MedUAGBench: distribution of benchmark samples across anatomical systems and imaging modalities, along with the hierarchical task composition covering reconstruction, translation, synthesis, and prediction. sample is then packaged with its image path, metadata, unique identifier, and task-specific text prompt. This step ensures format consistency across samples and supports unified multi- task training and evaluation. 4) Dataset Statistics: As shown in Figure 4, our dataset exhibits broad diversity in both the corpus and benchmark. MedUAGCorpus contains over 6M training instances, includ- ing a 1.76M domain-alignment set and a 4.61M instruction- tuning set. It comprises more than 5.3M images and spans 14 anatomical systems and 14 imaging modalities. In contrast, MedUAGBench is more compact and evaluation-oriented. Since it focuses primarily on medical generation, its sys- tem and modality coverage is smaller, yet it still spans 10 anatomical systems and 11 imaging modalities. Moreover, the benchmark remains task-diverse, covering four core generation categories, thereby enabling comprehensive and fine-grained evaluation of medical image generation. C. Implementation of MedUAG a) Model Architecure: As shown in Figure 3, MedUAG follows the design of Bagel [4] to balance the different gran- ularity requirements of understanding and generation within a unified architecture. Specifically, it uses a ViT encoder [30] to extract high-level semantic visual tokens for understanding, and a VAE encoder-decoder [31] to model low-level latent representations for image generation. Built on a decoder-only transformer backbone [32], the model instantiates two task- specific branches for understanding and generation, decoupling their optimization while preserving architectural consistency. For understanding, the model takes text tokens and ViT visual tokens as input and performs autoregressive next-token prediction, following standard MLLMs [32], [33]. The objec- tive is the cross-entropy loss: L und =âE (x,v,y)âŒD ïŁź ïŁ° |y| X t=1 logp Ξ (y t | x, v,y <t ) ïŁč ïŁ» ,(1) where x denotes the input text tokens, v the ViT visual tokens, and y the target output sequence. For generation, the model operates in the VAE latent space and predicts the target velocity under the flow-matching ob- jective [34], [35]: L gen =E (z t ,u t ,c)âŒD h â„f Ξ (z t ,t, c)â u t â„ 2 2 i ,(2) where z t is the noised latent feature at timestep t, c is the conditioning information, and u t is the target velocity. The final text and image outputs are decoded by the LM head and VAE decoder, respectively. The overall objective is: L = λ und L und + λ gen L gen ,(3) where λ und and λ gen balance these two losses. b) Domain Alignment: MedUAG is trained in two stages. The first stage adapts the pretrained backbone to the medical domain before unified instruction tuning. We construct the alignment corpus from three representative tasks: medical image reconstruction, text-to-image generation, and image captioning. These tasks provide complementary supervision: reconstruction offers dense pixel-level guidance for the genera- tion branch, text-to-image generation strengthens medical text- guided synthesis, and captioning improves image-language semantic alignment for the understanding branch. This stage establishes domain-aware representations and provides a stable initialization for subsequent instruction tuning. c) Instruction Tuning: The second stage trains MedUAG with unified multimodal instructions covering both generation and understanding. The generation data include synthesis, translation, reconstruction, and prediction, while the under- standing data include VQA and MRG. Compared with domain alignment, this stage emphasizes task diversity, compositional conditioning, and instruction responsiveness. Through joint instruction tuning, the understanding branch learns language- guided reasoning and reporting, while the generation branch learns diverse conditional image generation objectives, yield- ing a general-purpose medical multimodal model. TABLE I AVERAGE RESULTS ACROSS THE FOUR GENERATION CATEGORIES. FOR SYNTHESIS, WE REPORT TASK-LEVEL MACRO AVERAGES OF FID, GFID, AND BIOCS. FOR TRANSLATION, RECONSTRUCTION, AND PREDICTION, WE REPORT TASK-LEVEL MACRO AVERAGES OF LPIPS, MSE, PSNR, AND SSIM. Medical Synthesis Avg.Translation Avg.Reconstruction Avg.Prediction Avg. FIDâgFIDâBioCSâLPIPSâMSEâPSNRâSSIMâLPIPSâMSEâPSNRâSSIMâLPIPSâMSEâPSNRâSSIMâ BLIP3-o [5]â320.35316.740.3490.6250.06412.5360.3480.5130.05613.1770.3730.7130.09110.5960.268 UniWorld-V1 [36]â 334.51332.540.3210.6090.05313.1410.3770.6020.06911.9310.2870.6740.08411.1400.150 Bagel [4]â 261.39248.720.3610.6130.1869.1470.3070.5230.12312.3010.3290.6020.1658.4960.249 HealthGPT [7]â211.63202.480.3810.4660.05713.6590.4160.3700.02417.0000.5330.6790.07612.0030.262 UniMedVL [8]â192.21179.320.4030.4070.05413.6160.4860.3650.06313.4390.4550.5310.12911.2370.354 MedUAG (ours)â181.72167.720.3960.3640.06316.9920.5750.1130.00725.3570.8090.3200.03917.8970.529 TABLE I COMPARISON AMONG VARIOUS MODELS ON UNDERSTANDING BENCHMARKS, WHERE OM.VQA INDICATES OMNIMEDVQA. ModelUnifiedVQA-RADSLAKEPathVQAOM.VQAAvg. Proprietary Baselines GPT-4.1â65.072.255.575.567.1 Claude Sonnet 4â67.670.654.265.564.5 Gemini-2.5-Flashâ68.575.855.471.067.7 General-purpose Baslines Qwen2.5-VL-32Bâ71.871.241.968.263.3 InternVL3-38Bâ65.472.751.079.867.2 Llama3.2-11Bâ58.865.832.943.850.3 Qwen2.5VL-7Bâ 63.266.844.163.659.4 InternVL3-8Bâ65.472.848.679.166.5 Janus-Pro-7Bâ49.755.235.459.650.0 Bagelâ60.158.939.171.157.3 Medical Baselines LLaVA-Med-7Bâ46.651.935.234.842.1 RadFMâ50.634.614.323.530.8 HuatuoGPT-V-7Bâ67.668.144.874.363.7 HealthGPT-M3â55.956.439.768.555.1 HealthGPT-L14â58.364.544.474.460.4 UniMedVLâ61.975.453.585.869.2 MedUAG (ours)â75.678.054.577.071.3 I. EXPERIMENTS A. Implementation Details We train MedUAG in two stages initializing from the pretrained BAGEL-7B-MoT. In the alignment stage, the model is trained for 10k steps with a learning rate of 5 Ă 10 â5 on 1.76M samples, including 1.14M reconstruction samples (64.6%), 191K text-to-image samples (10.8%), and 432K image-to-text samples (24.5%). In the subsequent SFT stage, training is resumed from the aligned checkpoint and continued for 20k steps with a learning rate of 2 Ă 10 â5 on 4.61M instruction-tuning samples, consisting of 3.33M generation samples (72.3%) and 1.28M single-image understanding sam- ples (27.7%). Training is conducted on 32 NVIDIA H800 80GB GPUs. In both stages, we set the maximum number of tokens per sample to 16,384, and adopt a constant learning- rate schedule with 2,000 warmup steps, the AdamW optimizer (ÎČ 1 = 0.9, ÎČ 2 = 0.95, Δ = 10 â15 ), EMA decay of 0.9999, and frozen VAE weights. The loss weights λ und and λ gen are both set to 1.0. B. Benchmarks, Baselines and Metrics a) Benchmarks: We evaluate MedUAG on both medical generation and understanding tasks to comprehensively assess its unified capability. For generation, we use MedUAGBench, X-ray Dose Map MRI OCT CT Histopathology PET Microscopy Ultrasound Dermoscopy Fig. 5. Visualization of the MedUAG generation capability. a unified medical generation benchmark comprising 5,000 test instances over 12 tasks, which cover diverse medical imaging scenarios and require the model to handle heterogeneous generation objectives. For understanding, we evaluate on four widely used medical benchmarks, including VQA-RAD [37], SLAKE [38], PathVQA [39], and OmniMedVQA [40]. To- gether, these benchmarks provide a broad assessment of medical visual understanding ability, covering complementary reasoning demands across radiology, pathology, and general clinical question answering. b) Baselines: We compare MedUAG with a diverse set of baselines covering both unified multimodal models and spe- cialized medical MLLMs. For generation, we include represen- tative unified multimodal models from the general domain, in- cluding BLIP3-o [5], UniWorld-V1 [36], and Bagel [4], as well as medical-domain unified models, including HealthGPT [7] and UniMedVL [8]. For understanding, we consider three groups of baselines: proprietary systems (GPT-4.1 [41], Claude Sonnet 4 [42], and Gemini-2.5-Flash [43]), general-purpose open-source models (e.g., Qwen2.5-VL [32], InternVL3 [44], Llama3.2 [45], Janus-Pro [3], and Bagel [4]), and special- ized medical models (e.g., LLaVA-Med [46], RadFM [47], HuatuoGPT-V [48], HealthGPT [7], and UniMedVL [8]). c) Metrics: For MedUAGBench, we follow the standard evaluation protocol for each generation category. Specifically, for synthesis, we report FID [49], generation FID (gFID), and BiomedCLIP Score (BioCS) [50] to evaluate image realism, generative fidelity, and medical semantic consistency, respec- tively. For translation, reconstruction, and prediction, we report LPIPS [51], MSE, PSNR, and SSIM [52] to measure per- TABLE IV ABLATION STUDY OF DOMAIN ALIGNMENT AND INITIALIZATION. Align.Init. SynthesisTranslationReconstructionPrediction (gFIDâ)(LPIPSâ)(PSNRâ)(MSEâ) âScratch332.400.47822.4490.053 âBagel128.280.39624.4760.040 âScratch309.500.47523.0780.057 âBagel 167.720.36425.3570.039 TABLE V PSNR GAINS FROM DOMAIN ALIGNMENT ON RECONSTRUCTION TASKS. Init. Low-Dose CT Low-Dose PET Low-Field MRI Undersampled MRI Avg. From Scratch+1.13+1.00+0.33+0.05+0.63 From Bagel+1.14+0.05+1.89+0.44+0.88 ceptual similarity, pixel-wise error, reconstruction quality, and structural consistency. For medical understanding benchmarks, we use accuracy as the evaluation metric, where open-ended questions are assessed by Qwen3-VL-30B-A3B-Instruct [53]. C. Main Results 1) Comparison on MedUAGBench: Table I reports the quantitative results on MedUAGBench. MedUAG consis- tently outperforms both general-domain unified models and prior medical unified models across most generation set- tings, demonstrating the benefit of large-scale medical unified training. Its advantage is especially clear on reconstruction and prediction, where preserving anatomical structure and producing clinically grounded outputs are essential. These results suggest that MedUAG does not merely improve image realism, but also learns generation capabilities aligned with medical structure and task semantics. The qualitative examples in Figure 5 further show that MedUAG produces realistic and structurally faithful outputs across diverse imaging modalities. 2) Comparison on Understanding Benchmarks: Table I reports results on medical understanding benchmarks. Med- UAG achieves the best average accuracy of 71.3 against 16 baselines, leading on VQA-RAD and SLAKE with scores of 75.6 and 78.0, respectively, while remaining competitive on PathVQA and OmniMedVQA. Notably, although MedUAG is jointly trained for both generation and understanding, with understanding data comprising less than 30% of the corpus in both stages, it still demonstrates strong visual reasoning across radiology, pathology, and general medical QA. Its gap to UniMedVL on OmniMedVQA suggests that data composition, the understanding-generation mixing ratio, and task balancing remain important directions for future unified medical multi- modal models. D. Ablation Studies 1) Effect of Domain Alignment: To assess domain align- ment, we compare models with and without this stage under both from-scratch and Bagel initializations, as shown in Ta- ble IV. Domain alignment mainly benefits structure-preserving UnderstandingPredictionReconstructionTranslation Synthesis 0 20 40 60 80 Accuracy / PSNR 64.9 17.2 24.1 13.6 66.8 18.1 25.3 16.9 71.2 18.3 25.2 16.7 71.3 17.9 25.4 17.0 5%10%50%Full 0.0 0.1 0.2 0.3 0.4 0.5 BiomedCLIP Score 0.362 0.348 0.385 0.396 Fig. 6. Comparison of different SFT data ratios. tasks: with Bagel initialization, it lowers translation LPIPS by 0.032 and improves reconstruction PSNR by 0.881, with consistent gains across reconstruction benchmarks in Table V. However, synthesis performance decreases after alignment, revealing a trade-off between anatomical fidelity and genera- tive diversity. This suggests that alignment-induced structural priors should be retained, while synthesis-oriented objectives should be progressively strengthened in later training to better balance structural fidelity and generative realism. 2) Effect of Base Model: Table IV further illustrates the efficacy of different initialization strategies. Compared to training from scratch, Bagel initialization consistently yields substantial gains across nearly all tasks, most notably in syn- thesis and reconstruction. Without domain alignment, Bagel drastically reduces gFID from 332.40 to 128.28 and boosts reconstruction PSNR by 2.027. This superiority persists even after domain alignment is introduced, where Bagel still pro- vides the optimal foundation for translation, reconstruction, and prediction. These results underscore that general-domain unified pretraining establishes a robust representational prior, which significantly lowers the barrier for downstream medical adaptation and ensures a better optimization starting point. 3) Effect of Data Scale: We further analyze the effect of SFT data scale as shown in Figure 6. Increasing instruction- tuning data improves performance across all tasks, but the gains gradually saturate at larger scales. Understanding and synthesis show more sustained improvements than reconstruc- tion and prediction, likely because they occupy smaller por- tions of the SFT set and therefore benefit more from increased diversity and coverage as the total data grows. These results suggest that, beyond sufficient scale, further gains depend more on data diversity, task balance, and quality control than on simply adding more samples. IV. APPLICATION EXPLORATION Beyond unified understanding and generation, MedUAG can also synthesize multimodal data for downstream training in low-resource settings, where paired clinical data are often scarce, costly, or privacy-restricted [54]. To evaluate this potential, we conduct an exploratory experiment on MRG task from CheXpert [55] using Qwen3-VL-4B as the backbone. We compare three training settings: 1K real image-report pairs, 1K TABLE VI EFFECT OF SYNTHETIC DATA ON MEDICAL REPORT GENERATION. #Real#Synthetic#TotalMETEORBERTScoreLLM JudgeAverage 1K01K0.0980.3940.0370.176 1K1K2K0.1910.4950.1000.262 2K02K0.1960.4910.1250.271 real pairs augmented with 1K synthetic images generated by MedUAG from the corresponding reports, and a 2K-real upper bound. As shown in Table VI, synthetic augmentation im- proves the 1K-real baseline by 0.086 on average and narrows the gap to the 2K-real upper bound to only 0.009. These results suggest that MedUAG-generated images can provide effective supervision when real paired data are limited, highlighting the potential of unified medical multimodal models for scalable and cost-effective data augmentation. V. RELATED WORK Recent medical vision-language models adapt general- purpose MLLMs [32], [45], [53] to clinical scenarios [46], [47], [56]â[58]. These models have shown strong capabilities in medical VQA, report generation, and clinical reasoning by leveraging medical image-text alignment, instruction tuning, and large-scale domain-specific data. However, most Med- MLLMs are still centered on visual comprehension and text generation. They cannot directly synthesize, translate, recon- struct, or manipulate medical images at the pixel level, limiting their applicability to generative clinical tasks. UAG has recently become an active direction in general- domain multimodal learning [3]â[6], [36]. Existing methods explore different design choices, including coupling LLMs with generative modules, discretizing images into visual to- kens, decoupling visual encoders, and adopting expert-based architectures to support both perception and generation. De- spite their strong open-domain generalization, these models are primarily trained on natural images. As a result, they often struggle with clinical semantics, anatomical fidelity, and fine- grained pathology in medical imaging. Recent works have begun to explore unified multimodal models for medicine. HealthGPT [7] unifies comprehension and generation with heterogeneous knowledge adaptation, and UniMed-VL [8] models further mark important progress toward unified medical AI. Nevertheless, existing efforts re- main constrained by limited data scale, narrow generation task coverage, and insufficiently standardized evaluation. In contrast, our work provides a more comprehensive foundation by constructing a large-scale corpus, a systematic generation benchmark, and an end-to-end unified medical model. VI. CONCLUSION We present MedUAG, a unified medical multimodal model, together with MedUAGBench and MedUAGCorpus for sys- tematic training and evaluation of medical understanding and generation. MedUAG achieves strong performance across both task families, with clear gains in reconstruction and prediction, demonstrating the promise of unified modeling for general- purpose medical AI. We also identify key challenges, including limitations in some understanding settings and a trade-off between structural fidelity and synthesis diversity. Future work will further improve data composition, task balancing, and training strategies toward more robust and clinically useful medical multimodal systems. REFERENCES [1] S. Zhao, X. Zhang, J. Guo, J. Hu, L. Duan, M. Fu, Y. X. Chng, G.- H. Wang, Q.-G. Chen, Z. Xu et al., âUnified multimodal understanding and generation models: Advances, challenges, and opportunities,â arXiv preprint arXiv:2505.02567, 2025. [2] Y. Ge, S. Zhao, J. Zhu, Y. Ge, K. Yi, L. Song, C. Li, X. Ding, and Y. Shan, âSeed-x: Multimodal models with unified multi- granularity comprehension and generation,â 2025. [Online]. Available: https://arxiv.org/abs/2404.14396 [3] X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan, âJanus-pro: Unified multimodal understanding and generation with data and model scaling,â arXiv preprint arXiv:2501.17811, 2025. [4] C. Deng, D. Zhu, K. Li, C. Gou, F. Li, Z. Wang, S. Zhong, W. Yu, X. Nie, Z. Song, G. Shi, and H. Fan, âEmerging properties in unified multimodal pretraining,â 2025. [Online]. Available: https://arxiv.org/abs/2505.14683 [5] J. Chen, Z. Xu, X. Pan, Y. Hu, C. Qin, T. Goldstein, L. Huang, T. Zhou, S. Xie, S. Savarese et al., âBlip3-o: A family of fully open unified multimodal models-architecture, training and dataset,â arXiv preprint arXiv:2505.09568, 2025. [6] X. Wang, Y. Cui, J. Wang, F. Zhang, Y. Wang, X. Zhang, Z. Luo, Q. Sun, Z. Li, Y. Wang et al., âMultimodal learning with next-token prediction for large multimodal models,â Nature, p. 1â7, 2026. [7] T. Lin, W. Zhang, S. Li, Y. Yuan, B. Yu, H. Li, W. He, H. Jiang, M. Li, X. Song et al., âHealthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation,â arXiv preprint arXiv:2502.09838, 2025. [8] J. Ning, W. Li, C. Tang, J. Lin, C. Ma, C. Zhang, J. Liu, Y. Chen, S. Gao, L. Liu et al., âUnimedvl: Unifying medical multimodal understanding and generation through observation-knowledge-analysis,â arXiv preprint arXiv:2510.15710, 2025. [9] S. Jiang, Y. Wang, S. Song, T. Hu, C. Zhou, B. Pu, Y. Zhang, Z. Yang, Y. Feng, J. T. Zhou, J. Hao, Z. Chen, R. Wu, T. Tang, J. Lv, H. Xu, H. Wang, J. Xiao, B. Feng, F. Zhu, K. Li, W. Xie, J. Sun, J. Wu, and Z. Liu, âHulu-med: A transparent generalist model towards holistic medical vision-language understanding,â 2025. [Online]. Available: https://arxiv.org/abs/2510.08668 [10] P. Chambon, J.-B. Delbrouck, T. Sounack, S.-C. Huang, Z. Chen, M. Varma, S. Q. Truong, C. T. Chuong, and C. P. Langlotz, âChexpert plus: Augmenting a large chest x-ray dataset with text radiology reports, patient demographics and additional image formats,â 2024. [Online]. Available: https://arxiv.org/abs/2405.19538 [11] J. Yang, R. Shi, D. Wei, Z. Liu, L. Zhao, B. Ke, H. Pfister, and B. Ni, âMedmnist v2 - a large-scale lightweight benchmark for 2d and 3d biomedical image classification,â Scientific Data, vol. 10, no. 1, Jan. 2023. [Online]. Available: http://dx.doi.org/10.1038/ s41597-022-01721-8 [12] Y. Xie, C. Zhou, L. Gao, J. Wu, X. Li, H.-Y. Zhou, S. Liu, L. Xing, J. Zou, C. Xie, and Y. Zhou, âMedtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine,â 2025. [Online]. Available: https://arxiv.org/abs/2408.02900 [13] J. R Ì uckert, L. Bloch, R. Br Ì ungel, A. Idrissi-Yaghir, H. Sch Ì afer, C. S. Schmidt, S. Koitka, O. Pelka, A. B. Abacha, A. G. Seco de Herrera, H. M Ì uller, P. A. Horn, F. Nensa, and C. M. Friedrich, âRocov2: Radiology objects in context version 2, an updated multimodal image dataset,â Scientific Data, vol. 11, no. 1, Jun. 2024. [Online]. Available: http://dx.doi.org/10.1038/s41597-024-03496-6 [14] J. Ma, Y. Zhang, S. Gu, C. Zhu, C. Ge, Y. Zhang, X. An, C. Wang, Q. Wang, X. Liu, S. Cao, Q. Zhang, S. Liu, Y. Wang, Y. Li, J. He, and X. Yang, âAbdomenct-1k: Is abdominal organ segmentation a solved problem?â IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, p. 6695â6714, 2022. [15] C. Ma, Y. Ji, J. Ye, L. Zhang, Y. Chen, T. Li, M. Li, J. He, and H. Shan, âTowards interpretable counterfactual generation via multimodal autoregression,â 2025. [Online]. Available: https: //arxiv.org/abs/2503.23149 [16] ULF-EnC Organizers, âUlf-enc challenge: Ultra-low-field mri image en- hancement challenge,â https://w.synapse.org/Synapse:syn65485242/ wiki/, 2025, accessed 2026-04-02. [17] IXI Consortium, âIxi dataset,â https://brain-development.org/ixi-dataset/, 2007, public MRI dataset with T1, T2 and PD-weighted images; accessed 2026-04-02. [18] A. Thummerer, E. van der Bijl, A. J. Galapon, F. Kamp, M. Savenije, C. Muijs, S. Aluwini, R. J. H. M. Steenbakkers, S. Beuel, M. P. Intven, J. A. Langendijk, S. Both, S. Corradini, V. Rogowski, M. Terpstra, N. Wahl, C. Kurz, G. Landry, and M. Maspero, âSynthrad2025 grand challenge dataset: Generating synthetic cts for radiotherapy from head to abdomen,â Medical Physics, vol. 52, no. 7, Jul. 2025. [Online]. Available: http://dx.doi.org/10.1002/mp.17981 [19] K. Jin, X. Huang, J. Zhou, Y. Li, Y. Yan, Y. Sun, Q. Zhang, Y. Wang, and J. Ye, âFives: A fundus image dataset for artificial intelligence based vessel segmentation,â Scientific Data, vol. 9, no. 1, p. 475, 2022. [Online]. Available: https://doi.org/10.1038/s41597-022-01564-3 [20] J. Staal, M. Abramoff, M. Niemeijer, M. Viergever, and B. van Gin- neken, âRidge-based vessel segmentation in color images of the retina,â IEEE Transactions on Medical Imaging, vol. 23, no. 4, p. 501â509, 2004. [21] T. R. Moen, B. Chen, D. R. Holmes, 3rd, X. Duan, Z. Yu, L. Yu, S. Leng, J. G. Fletcher, and C. H. McCollough, âLow-dose CT image and projection dataset,â Medical Physics, vol. 48, no. 2, p. 902â911, 2021, pMID: 33202055; PMCID: PMC7985836. [22] S. G. Armato, I, G. McLennan, L. Bidaut, M. F. McNitt- Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Henschke, E. A. Hoffman et al., âData From LIDC- IDRI,â The Cancer Imaging Archive, 2015. [Online]. Available: https://doi.org/10.7937/K9/TCIA.2015.LO9QL9SX [23] S. Xue, H. Wang, Y. Chen, F. Liu, H. Zhu, M. Viscione, R. Guo, A. Rominger, B. Li, and K. Shi, â UDPET: Ultra-low Dose PET Imaging Challenge Dataset ,â in proceedings of Medical Image Computing and Computer Assisted Intervention â MICCAI 2025, vol. LNCS 15972. Springer Nature Switzerland, September 2025. [24] J. Zbontar, F. Knoll, A. Sriram, T. Murrell, Z. Huang, M. J. Muck- ley, A. Defazio, R. Stern, P. Johnson, M. Bruno et al., âfastmri: An open dataset and benchmarks for accelerated mri,â arXiv preprint arXiv:1811.08839, 2018. [25] S. Liu, C. Zhu, F. Xu, X. Jia, Z. Shi, and M. Jin, âBci: Breast cancer immunohistochemical image generation through pyramid pix2pix,â 2022. [Online]. Available: https://arxiv.org/abs/2204.11425 [26] F. Li, Z. Hu, W. Chen, and A. Kak, âAdaptive supervised patchnce loss for learning h&e-to-ihc stain translation with inconsistent groundtruth image pairs,â 2023. [Online]. Available: https://arxiv.org/abs/2303.06193 [27] A. Babier, B. Zhang, R. Mahmood, K. L. Moore, T. G. Purdie, A. L. McNiven, and T. C. Y. Chan, âOpenkbp: The open-access knowledge-based planning grand challenge and dataset,â Medical Physics, vol. 48, no. 9, p. 5549â5561, Jun. 2021. [Online]. Available: http://dx.doi.org/10.1002/mp.14845 [28] R. Gao, M. Diallo, H. Liu, A. Magliari, J. Sackett, W. Verbakel, S. Meyers, R. Mcbeth, M. Zarepisheh, S. Arberet, M. Kraus, F. C. Ghesu, and A. Kamen, âAutomating rt planning at scale: High quality data for ai training,â 2025. [Online]. Available: https://arxiv.org/abs/2501.11803 [29] AAPM Grand Challenge Organizers, âGdp-hmm challenge: Generaliz- able dose prediction for heterogenous multi-cohort and multi-site radio- therapy planning,â https://w.aapm.org/GrandChallenge/GDP-HMM/, 2025, associated with the HMM-RT dataset; accessed 2026-04-02. [30] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., âAn image is worth 16x16 words: Transformers for image recognition at scale,â arXiv preprint arXiv:2010.11929, 2020. [31] D. P. Kingma and M. Welling, âAuto-encoding variational bayes,â arXiv preprint arXiv:1312.6114, 2013. [32] S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang et al., âQwen2.5-vl technical report,â eprint arXiv: 2502.13923, 2025. [33] H. Liu, C. Li, Q. Wu, and Y. J. Lee, âVisual instruction tuning,â Advances in neural information processing systems, vol. 36, p. 34 892â 34 916, 2023. [34] X. Liu, C. Gong, and Q. Liu, âFlow straight and fast: Learning to generate and transfer data with rectified flow,â arXiv preprint arXiv:2209.03003, 2022. [35] Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, âFlow matching for generative modeling,â arXiv preprint arXiv:2210.02747, 2022. [36] B. Lin, Z. Li, X. Cheng, Y. Niu, Y. Ye, X. He, S. Yuan, W. Yu, S. Wang, Y. Ge, Y. Pang, and L. Yuan, âUniworld-v1: High-resolution semantic encoders for unified visual understanding and generation,â 2025. [Online]. Available: https://arxiv.org/abs/2506.03147 [37] J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, âA dataset of clinically generated visual questions and answers about radiology images,â Scientific data, vol. 5, no. 1, p. 180251, 2018. [38] B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu, âSlake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,â in 2021 IEEE 18th international symposium on biomedical imaging (ISBI). IEEE, 2021, p. 1650â1654. [39] X. He, Y. Zhang, L. Mou, E. Xing, and P. Xie, âPathvqa: 30000+ questions for medical visual question answering,â arXiv preprint arXiv:2003.10286, 2020. [40] Y. Hu, T. Li, Q. Lu, W. Shao, J. He, Y. Qiao, and P. Luo, âOmnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm,â in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, p. 22 170â22 183. [41] OpenAI. (2025, Apr.) Introducing gpt-4.1 in the api. [Online]. Available: https://openai.com/index/gpt-4-1 [42] Anthropic. (2025, May) Introducing claude 4. [Online]. Available: https://w.anthropic.com/news/claude-4 [43] G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen et al., âGem- ini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,â arXiv preprint arXiv:2507.06261, 2025. [44] J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao et al., âInternvl3: Exploring advanced training and test-time recipes for open-source multimodal models,â arXiv preprint arXiv:2504.10479, 2025. [45] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., âThe llama 3 herd of models,â arXiv preprint arXiv:2407.21783, 2024. [46] C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, âLlava-med: Training a large language-and-vision assistant for biomedicine in one day,â Advances in Neural Information Processing Systems, vol. 36, p. 28 541â28 564, 2023. [47] C. Wu, X. Zhang, Y. Zhang, H. Hui, Y. Wang, and W. Xie, âTowards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data,â Nature Communications, vol. 16, no. 1, p. 7866, 2025. [48] J. Chen, C. Gui, R. Ouyang, A. Gao, S. Chen, G. H. Chen, X. Wang, Z. Cai, K. Ji, X. Wan et al., âTowards injecting medical visual knowledge into multimodal llms at scale,â in Proceedings of the 2024 conference on empirical methods in natural language processing, 2024, p. 7346â 7370. [49] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, âGans trained by a two time-scale update rule converge to a local nash equilibrium,â Advances in neural information processing systems, vol. 30, 2017. [50] S. Zhang, Y. Xu, N. Usuyama, H. Xu, J. Bagga, R. Tinn, S. Preston, R. Rao, M. Wei, N. Valluri et al., âBiomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs,â arXiv preprint arXiv:2303.00915, 2023. [51] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, âThe unreasonable effectiveness of deep features as a perceptual metric,â in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, p. 586â595. [52] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, âImage quality assessment: from error visibility to structural similarity,â IEEE transactions on image processing, vol. 13, no. 4, p. 600â612, 2004. [53] S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge et al., âQwen3-vl technical report,â arXiv preprint arXiv:2511.21631, 2025. [54] J. Wang, K. Wang, Y. Yu, Y. Lu, W. Xiao, Z. Sun, F. Liu, Z. Zou, Y. Gao, L. Yang et al., âSelf-improving generative foundation model for synthetic medical image generation and clinical applications,â Nature Medicine, vol. 31, no. 2, p. 609â617, 2025. [55] J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Mark- lund, B. Haghgoo, R. Ball, K. Shpanskaya et al., âChexpert: A large chest radiograph dataset with uncertainty labels and expert comparison,â in Proceedings of the AAAI conference on artificial intelligence, vol. 33, no. 01, 2019, p. 590â597. [56] J. Chen, C. Gui, R. Ouyang, A. Gao, S. Chen, G. H. Chen, X. Wang, Z. Cai, K. Ji, X. Wan et al., âTowards injecting medical visual knowledge into multimodal llms at scale,â in Proceedings of the 2024 conference on empirical methods in natural language processing, 2024, p. 7346â 7370. [57] S. Jiang, Y. Wang, S. Song, T. Hu, C. Zhou, B. Pu, Y. Zhang, Z. Yang, Y. Feng, J. T. Zhou et al., âHulu-med: A transparent generalist model towards holistic medical vision-language understanding,â arXiv preprint arXiv:2510.08668, 2025. [58] K. Zhang, R. Zhou, E. Adhikarla, Z. Yan, Y. Liu, J. Yu, Z. Liu, X. Chen, B. D. Davison, H. Ren et al., âA generalist visionâlanguage foundation model for diverse biomedical tasks,â Nature medicine, vol. 30, no. 11, p. 3129â3141, 2024.