Paper deep dive
Dense Temporal Contrast Synthesis via Conditioned Latent Transport
Smriti Joshi, Apostolia Tsirikoglou, Daniel M. Lang, Richard Osuala, Noah Márquez Varaa, Alejandro Guzman, Grzegorz Skorupko, Sebastian Ibarra Arregui, Lidia Garrucho, Akane Ohashi, Dimitra Ntoula, Eugen Divjak, Oğuz Lafcı, Jan C. Peeken, Julia A. Schnabel, Fredrik Strand, Oliver Diaz, Karim Lekadir
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/3/2026, 2:58:12 AM
Summary
The paper proposes a novel conditioned latent transport framework for synthesizing dynamic contrast-enhanced MRI (DCE-MRI) from pre-contrast images. This method predicts contrast enhancement in a single forward pass by anchoring latent trajectories to pre-contrast anatomy and applying continuous time conditioning. It outperforms state-of-the-art models in spatial, perceptual, and temporal metrics, significantly improves downstream tumor segmentation performance, and demonstrates clinical viability in a reader study with radiologists.
Entities (9)
Relation Signals (6)
Conditioned Latent Transport Framework → synthesizes → DCE-MRI
confidence 98% · predicts contrast enhancement in a single forward pass... synthesizes patient-specific contrast evolution
DCE-MRI → usedfor → Breast Cancer
confidence 97% · Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management
Conditioned Latent Transport Framework → outperforms → state-of-the-art models
confidence 95% · The proposed approach outperforms baseline and the state-of-the-art models across spatial, perceptual, temporal, and distributional metrics.
Conditioned Latent Transport Framework → improves → Tumor Segmentation
confidence 92% · our synthetic contrast enhancement significantly improved downstream tumor segmentation performance
Conditioned Latent Transport Framework → reducesrelianceon → Gadolinium-based contrast agents (GBCAs)
confidence 90% · suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows
Synthesized Images → supports → clinical management decisions
confidence 85% · synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindicated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow iterative sampling, underutilize structural priors, and lack clinical validation. We propose a novel conditioned latent transport framework that predicts contrast enhancement in a single forward pass. By anchoring the latent trajectory to the pre-contrast anatomy and applying continuous time conditioning, the model synthesizes patient-specific contrast evolution at any acquisition time. The proposed approach outperforms baseline and the state-of-the-art models across spatial, perceptual, temporal, and distributional metrics. Evaluated on an independent external cohort, the method demonstrates robustness to domain shifts induced by scanner noise as well as differing acquisition protocol. Furthermore, our synthetic contrast enhancement significantly improved downstream tumor segmentation performance, yielding a 22.4% relative increase in Dice coefficient (0.60 vs. 0.49 baseline pre-contrast, p < 0.01), reducing boundary segmentation error by over 39%, while outperforming all other generative model baselines. Finally, a reader study involving four breast radiologists evaluated the image quality, kinetic fidelity, and diagnostic viability of our synthesized sequences across 40 randomly selected cases. The results demonstrated that in 70% of cases, synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI, suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows.
Tags
Links
- Source: https://arxiv.org/abs/2607.29394v1
- Canonical: https://arxiv.org/abs/2607.29394v1
Trouble viewing inline? Open PDF directly →
Full Text
119,778 characters extracted from source content.
Expand or collapse full text
Dense Temporal Contrast Synthesis via Conditioned Latent Transport Smriti Joshi a,∗ , Apostolia Tsirikoglou c, † , Daniel M. Lang d,e, † , Richard Osuala a,g, † , Noah Márquez Vara a,h , Alejandro Guzman a , Grzegorz Skorupko a , Sebastian Ibarra Arregui a , Lidia Garrucho a , Akane Ohashi i,j , Dimitra Ntoula k , Eugen Divjak l,m , Oğuz Lafcı n , Jan C. Peeken g , Julia A. Schnabel d,e,f , Fredrik Strand c , Oliver Diaz a and Karim Lekadir a,b a Departament de Matemàtiques i Informàtica, Universitat de Barcelona, Barcelona, Spain b Institució Catalana de Recerca i Estudis Avançats (ICREA), Barcelona, Spain c Department of Oncology-Pathology, Karolinska Institutet, SE-171 77, Stockholm, Sweden d Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Munich, Germany e School of Computation, Information and Technology, Technical University of Munich, Munich, Germany f School of Biomedical Engineering and Imaging Sciences, King’s College London, London, United Kingdom g Department of Radiation Oncology, TUM University Hospital Rechts der Isar, TUM School of Medicine and Health, Technical University of Munich, Munich, Germany h Department of Computer Science and Engineering, Chalmers University of Technology, Gothenburg, Sweden i Department of Translational Medicine, Diagnostic Radiology, & CIRCE – the Center for Interdisciplinary Research on Cancer and Equity in Women, Lund University, Lund, Sweden j Department of Imaging and Physiology, Skåne University Hospital, Malmö, Sweden k Department of Radiology, Karolinska University Hospital, 17177, Stockholm, Sweden l University of Zagreb, School of Medicine, Zagreb, Croatia m University Hospital Dubrava, Zagreb, Croatia n Department of Biomedical Imaging and Image-Guided Therapy, Medical University of Vienna, Vienna, Austria A R T I C L E I N F O Keywords: Contrast Enhancement Breast Cancer Magnetic Resonance Imaging Clinical Reader Study Tumor Segmentation Generative Models A B S T R A C T Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based contrast agents (GBCAs) restricts use in contraindi- cated populations, prolongs scan protocols, and presents environmental toxicity concerns. Contrast synthesis offers a non-invasive alternative; however, existing approaches struggle to balance spatial realism with temporal continuity, suffer from slow iterative sampling, underutilize structural priors, and lack clinical validation. We propose a novel conditioned latent transport framework that predicts contrast enhancement in a single forward pass. By anchoring the latent trajectory to the pre-contrast anatomy and applying continuous time conditioning, the model synthesizes patient-specific contrast evolution at any given acquisition time. The proposed approach outperforms baseline and the state- of-the-art models across spatial, perceptual, temporal, and distributional metrics. Evaluated on an independent external cohort, the method demonstrates robustness to domain shifts induced by scanner noise as well as differing acquisition protocol. Furthermore, our synthetic contrast enhancement significantly improved downstream tumor segmentation performance, yielding a 22.4% relative increase in Dice coefficient (0.60 vs. 0.49 baseline pre-contrast, 푝 < 0.01), reducing boundary segmentation error by over 39%, while outperforming all other generative model baselines. Finally, a reader study involving four breast radiologists evaluated the image quality, kinetic fidelity, and diagnostic viability of our synthesized sequences across 40 randomly selected cases. The results demonstrated that in 70% of cases, synthesized images provided sufficient clinical information to support the same management decisions as real DCE-MRI, suggesting a path toward safer and faster contrast-free or contrast-reduced imaging workflows. 1. Introduction Magnetic resonance imaging (MRI) is extensively used in clinical practice due to its high diagnostic sensitivity and soft-tissue contrast. It plays an integral role in breast cancer staging (Selvi et al., 2018), treatment monitoring, and surgical planning (Parvaiz et al., 2016; Kuhl et al., 2017). Unlike other imaging modalities like mammography and ultrasound, which are known to underestimate tumor † These authors contributed equally to this work. ∗ Corresponding author smriti.joshi@ub.edu (S. Joshi) ORCID(s): 0000-0001-8480-023X (S. Joshi) extent (Dixon et al., 2016; Uematsu et al., 2008), MRI en- ables the accurate delineation of tumor margins, potentially reducing surgical re-excision rates (Gonzalez et al., 2014). In addition to treatment management, breast MRI is a criti- cal screening modality for high-risk populations, including women carrying genetic mutations, those with a history of chest irradiation before age 30, or those presenting dense breast tissue (Saslow et al., 2007; Monticciolo et al., 2018). The clinical value of MRI in screening is also significant; in one study, single-screening MRI has been shown to depict 18.1 additional cancers per 1,000 women with a history of breast-conserving therapy (BCT) (Gweon et al., 2014). Moreover, the evaluation of time-intensity curves detailing S. Joshi et al.: Preprint submitted to ElsevierPage 1 of 25 arXiv:2607.29394v1 [cs.CV] 31 Jul 2026 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 1: Overview of the study design with validation frame- work. Our proposed (left) method leverages a pre-contrast anchor and continuous acquisition time (휏) to synthesize tem- porally consistent post-contrast DCE-MRIs (center). Clinical viability is evaluated via a three part validation strategy (right) comprising: (1) multi-site evaluation for domain robustness, (2) downstream tumor segmentation for biological fidelity, and (3) a clinical reader study to assess diagnostic impact. contrast enhancement serves as a versatile biomarker, pro- viding, among others, a reliable, non-invasive estimation of tumor malignancy (Kuhl et al., 1999; Partridge et al., 2014). The high sensitivity of MRI is largely owed to contrast agents (Pesapane et al., 2025), most commonly gadolinium- based contrast agents (GBCAs). However, routine reliance on GBCAs introduces universal procedural burdens, along- side well-recognized toxicity risks (Rogosnitzky and Branch, 2016; Fraum et al., 2017; Alahari and Benstetter, 2017). From an operational perspective, contrast administration re- quires intravenous administration, which invariably length- ens scan preparation time and carries a risk of access failure requiring specialized staff assistance. Furthermore, safety protocols mandate the immediate physical availabil- ity of a supervising physician, imposing geographical and scheduling restrictions on where and when breast MRI can be performed. In terms of patient safety, GBCAs pose significant risks and contraindications for vulnerable popula- tions, including pregnant women (Peterson et al., 2023) and patients with renal failure (Grobner and Prischl, 2007). They have also been associated with retention and accumulation in the brain (Kanda et al., 2015) and bone tissue (Darrah et al., 2009), particularly following exposure to linear agents. Although newer macrocyclic agents appear to reduce tissue retention, they may not completely eliminate it (Murata et al., 2016), and the clinical significance of this persistence remains uncertain. The environmental consequences of GBCAs are also considerable. Because these agents are predominantly ex- creted renally without undergoing metabolization, they resist degradation in wastewater treatment plants and are contin- uously emitted into aquatic ecosystems (Bau and Dulski, 1996; Laczovics et al., 2023). Ecologically, the 퐺푑 3+ ion is capable of mimicking essential cations, including cal- cium, zinc, magnesium, and iron (Cotruvo, 2019), thereby threatening to disrupt critical cellular and biochemical path- ways (Krasznai et al., 2003). The potential for these agents to impact the food chain (Arciszewska et al., 2022) through plants (Lindner et al., 2013), terrestrial as well as aquatic life (Lingott et al., 2016), alongside their unknown long- term ecological effects, constitutes a pressing and unresolved concern. To mitigate these risks and reduce overall scan times, research has focused on synthesizing dynamic contrast- enhanced (DCE) MRI from non-contrast, early-phase, or low-dose acquisitions. Explored methods include generative adversarial networks (GANs) (Dar et al., 2019; Müller- Franzes et al., 2023; Osuala et al., 2025; Fonnegra et al., 2025), probabilistic diffusion models (Osuala et al., 2024; Kishore Kumar et al., 2024; Ibarra et al., 2025; Fan et al., 2025; Kong et al., 2026), and deterministic approaches (Lang et al., 2025; Chung et al., 2025a; Chen et al., 2026). However, three primary limitations persist across these paradigms. First, enforcing temporal consistency across DCE sequences often compromises spatial realism. Realistic high-frequency textures are typically generated via stochastic mechanisms, such as noise injection or adversarial sampling; however, this inherent randomness introduces non-physiologically- grounded variations across sequential frames, disrupting smooth temporal continuity. Second, iterative generative models require extensive inference time, limiting computa- tional efficiency. Third, existing methods underutilize pre- contrast anatomical priors; requiring a network to synthesize baseline structures from scratch distracts it from accurately modeling true contrast dynamics. Beyond these techni- cal challenges, the true clinical utility of existing models remains difficult to assess due to a widespread lack of downstream task validation and expert reader evaluation. To address these gaps collectively, we propose a frame- work for dense temporal contrast synthesis, and rigorously evaluate it to position the method in terms of generalizability and clinical utility. A respective overview of this work is presented in Figure 1. Our principal contributions are sum- marized as follows: • We introduce a novel conditioned latent transport network that synthesizes patient-specific DCE-MRI at any continuous acquisition time (휏) from its corre- sponding pre-contrast image. • We demonstrate that anchoring the generative trajec- tory to the pre-contrast anatomy enforces macroscopic structural fidelity, while tailored spectral and struc- tural loss functions preserve high-frequency details and perceptual realism. S. Joshi et al.: Preprint submitted to ElsevierPage 2 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 2: Overview of the proposed Latent Generative Model Architecture. Pre-contrast (푥 푝푟푒 ) and post-contrast (푥 푝표푠푡 ) images are compressed into a high-fidelity latent space via a frozen encoder () of a custom pretrained autoencoder. Then, the intermediate state (푧 푡 ) is constructed by interpolating between the non-enhanced structural anchor (푧 푝푟푒 ) and the target state (푧 푝표푠푡 ) with added stochasticity (휎). Conditioned on 푧 푝푟푒 , a sinusoidal pharmacokinetic time embedding (휏), and interpolation timestep t, the central U-Net acts as a target-predictor, directly regressing the constant latent subtraction map (̂푧 푒푛ℎ ≈ 푧 푝표푠푡 − 푧 푝푟푒 ). Adding the 푧 푝푟푒 to the predicted enhancement, ̂푧 푝표푠푡 is obtained to compute the latent supervision loss ( 푀푆퐸 ). Then, ̂푧 푝표푠푡 is passed through a frozen decoder () to upsample the image to pixel space. Structural and perceptual fidelity are strictly enforced via perceptual loss ( 퐿푃퐼푃푆 ), while a log-amplitude Fourier loss ( 퐹퐿 ) explicitly penalizes frequency deviations to preserve fine micro-vascular details. • We evaluate our method on an independent, external validation cohort, showing improvement over the pre- contrast image and competing state-of-the-art meth- ods, while also systematically analyzing the domain shift and its effect on performance. • We validate the anatomical fidelity of synthesized im- ages through downstream tumor segmentation, show- ing statistically significant improvements over com- peting methods across two complementary evaluation paradigms. • We conduct a comprehensive clinical reader study with four expert radiologists to assess the synthesized images, demonstrating their diagnostic viability and potential to reduce reliance on contrast administration in clinical practice. The remainder of this manuscript is organized as follows: Section 2 reviews the current state-of-the-art in contrast synthesis. Section 3 outlines our proposed methodology. Section 4 discusses the implementation details. Section 5 discusses the quantitative and qualitative results. Finally, Section 6 summarizes our core findings and highlights po- tential directions for future research. 2. Related Work Early methodologies in contrast phase synthesis lever- aged GAN based architectures such as pix2pix (Isola et al., 2017), for image-to-image generation of post-contrast phases from pre-contrast sequences (Osuala et al., 2025; Müller- Franzes et al., 2023). GANs have also been applied to cross- phase synthesis, demonstrating the feasibility of predicting late post-contrast phases from early-phase acquisitions (Fon- negra et al., 2025; Müller-Franzes et al., 2024). While these models successfully capture high-frequency structural details, they often struggle with training instability and the hallucination of non-existent anatomical features. Diffusion models (Ho et al., 2020; Rombach et al., 2022) provide an alternative to GANs because of their stable training dynamics and high-fidelity sample generation. In the medical imaging domain, diffusion models have been widely adopted across a spectrum of downstream tasks, including image synthesis (Konz et al., 2024; Pinaya et al., 2022), artifact denoising (Gao et al., 2023), and semantic segmentation (Yan et al., 2024a). Specifically, within con- trast synthesis, ContrastControlNet (CCNet) (Osuala et al., 2024) trains the latent diffusion model for pre to post contrast synthesis in breast imaging, while parallel frameworks have been optimized for prostate imaging (M et al., 2024). To further constrain the generative process and ensure anatom- ical fidelity, several approaches have incorporated multi- modal conditioning inputs, integrating raw imaging data with clinical metadata and explicit outline guidance (Ibarra et al., 2025; Konz et al., 2024; Fan et al., 2025; Osuala et al., 2024). Probabilistic diffusion models are valuable for mod- eling data diversity; however, the preference for consis- tent, repeatable outputs in clinical workflows motivates the shift toward deterministic synthesis. In this realm, works on breast DCE-MRI contrast synthesis rely on hierarchical networks (Zhang et al., 2023), temporal neural cellular au- tomata (Lang et al., 2025), and iterative networks (Chung et al., 2025b). These works promise a strong direction by demonstrating improved anatomical and temporal fidelity, which are frequently sacrificed in favor of perceptual realism within traditional stochastic models. Another line of research are cold diffusion models, which replace stochastic Gaussian perturbations with deterministic degradation operators (e.g. S. Joshi et al.: Preprint submitted to ElsevierPage 3 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport blurring) and learn the corresponding inverse restoration process (Bansal et al., 2023). Although not yet explored for contrast synthesis, this formulation has demonstrated promising results across a variety of medical imaging appli- cations, including segmentation (Yan et al., 2024b; Zaman et al., 2024; Qi et al., 2025), anomaly detection (Naval Ma- rimont et al., 2024), denoising (Zhang et al., 2025; Tang et al., 2025), and image reconstruction (Shen et al., 2024; Habijan et al., 2025). Importantly, the deterministic degrada- tion trajectory provides a more structured restoration process compared to stochastic sampling, which can facilitate effi- cient inference. A parallel direction in generative modeling is represented by flow-based approaches, including flow matching and rectified flow formulations, which learn a time-dependent velocity field to transport samples between source and target distributions (Lipman et al., 2022; Liu et al., 2022). These methods have been explored in latent generative frameworks (Dao et al., 2023), Computed To- mography (CT) image synthesis (Wang et al., 2025), multi- modal MRI translation (Tur et al., 2026), and more recently for contrast synthesis (Chen et al., 2026), where the task is formulated as a missing modality prediction problem. Despite their impressive generative capabilities, both stochastic diffusion models and flow-based generative meth- ods typically rely on multi-step denoising or numerical inte- gration during inference. While effective for unconstrained image synthesis, repeated integration is not inherently re- quired for contrast enhancement, where the source anatomy is already observed and only the physiological contrast up- take must be modeled, an observation also explored in conditional flow formulations (Tur et al., 2026). Cold Dif- fusion replaces iterative denoising targets with direct end- point prediction from deterministically corrupted intermedi- ate states (Bansal et al., 2023). Concurrently, Rectified Flow frames generation as a transport problem learned from inter- polated intermediate states, providing supervision along the transformation path rather than only at the endpoints (Lip- man et al., 2022). Motivated by these observations, we formulate contrast synthesis as a conditional latent transport problem that combines endpoint prediction with structured interpolation, enabling efficient single-step inference while retaining the richer supervision provided by intermediate- state training. 3. Methodology Our method is presented visually in Figure 2. We detail the mathematical formulation, architectural pipeline, and inference strategy of the proposed method below. 3.1. Problem Formulation Let ⊂ℝ 퐶×퐻 denote the pixel space of 2D MR image slices, where 퐻 and 푊 represent the spatial height and width, respectively. For a given patient, standard clinical protocols acquire a non-enhanced pre-contrast baseline slice 푥 푝푟푒 ∈ and a target contrast-enhanced slice 푥 푝표푠푡 (휏 푖 ) ∈ captured at a discrete, protocol-specific physical timepoint 휏 푖 . However, while real scanner acquisitions are fundamen- tally discrete and temporally sparse, physiological contrast enhancement is a continuously evolving process. Therefore, our objective is to learn a continuous generative transport mapping 휃 ∶ (푥 푝푟푒 ,휏)↦ ̂푥 푝표푠푡 (휏) that accepts any arbitrary timepoint 휏 ∈ℝ + . This formulation allows us to accu- rately synthesize the non-linear pharmacokinetic hemody- namics of localized tumor regions across the entire temporal domain, while strictly preserving the underlying patient- specific anatomical structure and global tissue topology. 3.2. Latent Space Compression and Custom Autoencoder To mitigate the severe computational bottlenecks associ- ated with high-resolution medical imaging, we map the pixel space into a compressed, low-dimensional latent manifold ⊂ℝ 푐× 퐻 푓 × 푊 푓 using a Variational Autoencoder (VAE), where 푓 denotes the spatial downsampling factor. Let and denote the encoder and decoder, such that 푧 = (푥) and 푥 ≈ (푧). Standard stable diffusion VAEs (Rombach et al., 2022) typically employ an 푓 = 8 spatial downsampling factor, achieved through three consecutive convolutional blocks with a stride of 2. While computationally efficient, we em- pirically observed that this aggressive compression discards the fine-grained micro-vasculature and parenchymal textures essential for clinical diagnostics (Appendix A.3). To address this spatial bottleneck, we train a custom AutoencoderKL 1 architecture from scratch, strictly constrained to a milder 푓 = 4 spatial downsampling factor. We optimize this encoder-decoder pair on our pre-contrast and post-contrast sequences using the standard VAE objective: 푉 퐴퐸 = 푟푒푐표푛 (푥,((푥))) + 훽 퐾퐿 (푞 휙 (푧|푥)||푝(푧)) , where 푟푒푐표푛 is a composite reconstruction loss combining Mean Squared Error (MSE) for signal fidelity and Learned Perceptual Image Patch Similarity (LPIPS) (Zhang et al., 2018) for structural sharpness, and 퐾퐿 is the Kullback- Leibler divergence used to regularize the latent manifold (Kingma and Welling, 2013) by forcing the learned posterior 푞 휙 (푧|푥) to approximate a standard normal prior푝(푧). We observe that this4× configuration balances computational efficiency with the retention of micro-scale structures required for accurate contrast synthesis. Once trained, and are frozen. The resulting latent representations (푧 푝푟푒 and 푧 푝표푠푡 ) provide the mathematical domain for the contrast synthesis network. 3.3. Forward Process and Temporal Conditioning The core objective is to transport the non-enhanced base- line latent푧 푝푟푒 to the target enhanced state푧 푝표푠푡 (휏). We frame this generation as a conditioned latent process governed by a linear interpolant. The network, parameterized as an acquisition-time conditioned U-Net 휃 , is trained to map a corrupted intermediate state 푧 푡 to the fully enhanced target 1 https://huggingface.co/docs/diffusers/en/api/models/autoencoderkl S. Joshi et al.: Preprint submitted to ElsevierPage 4 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport state. The forward process corrupts the target state directly toward the patient’s non-enhanced baseline 푧 푝푟푒 : 푧 푡 = (1 − 푤 푡 )푧 푝표푠푡 + 푤 푡 푧 푝푟푒 + (푤 푡 ⋅ 휎 푛표푖푠푒 )휖 , where 푤 푡 ∈ [0,1] dictates the temporal degradation schedule, 푡 ∈ [0,1000] is the timestep, and 휖 ∼ (0,퐼) introduces stochasticity to prevent deterministic collapse. The degradation weight 푤 푡 governs the structural denoising process and is strictly decoupled from the physical acqui- sition time 휏. The stochastic noise variance scales linearly with 푤 푡 , maximizing at the baseline (푤 푡 = 1) and decaying to zero at the target (푤 푡 = 0). We treat the physical acquisition time 휏 as a continuous conditional prior. It is encoded via Sinusoidal Positional Embeddings (Vaswani et al., 2017) and passed through a Multi-Layer Perceptron (MLP) to match the generative network’s feature dimensionality. This condition, alongside the timestep 푡, is injected into the U-Net residual blocks via Adaptive Group Normalization (AdaGN) (Dhariwal and Nichol, 2021). Conditioning explicitly on 휏 enables the model to learn spatially-varying intensity mappings, assign- ing distinct physiological kinetics to different biological structures. The spatial input is constructed via channel-wise concatenation, yielding 푧 푖푛 = [푧 푡 ∥ 푧 푝푟푒 ]. The generative mapping is thus formalized as ̂푧 푝표푠푡 = 휃 (푧 푖푛 ,푡,휏). 3.4. Residual Parameterization and Training Objective Because the non-enhanced anatomical topology remains strictly static during the localized acquisition window, the pharmacokinetic transformation can be expressed in residual form as Δ푧 = 푧 푝표푠푡 − 푧 푝푟푒 . To improve representational efficiency, we parameterize the latent U-Net 휃 to predict this residual component directly, yieldingΔ̂푧 = 휃 (푧 푖푛 ,푡,휏). This residual formulation decoupling the representation of the transformation (via 푧 푡 ) from the prediction objective (via Δ푧) prevents the network from allocating capacity to redun- dant anatomical reconstruction to simplify optimization. The final enhanced latent is reconstructed as ̂푧 푝표푠푡 = 푧 푝푟푒 + Δ̂푧. We train the model using a composite objective function to ensure signal fidelity, perceptual realism, and the retention of fine frequency details: 푡표푡푎푙 = 휆 푀푆퐸 푀푆퐸 + 휆 퐿푃퐼푃푆 퐿푃퐼푃푆 + 휆 퐹퐿 퐹퐿 Signal Fidelity ( 푀푆퐸 ): We apply a standard MSE loss between the predicted and ground-truth latents, defined as ||̂푧 푝표푠푡 −푧 푝표푠푡 || 2 2 , to guarantee macroscopic pharmacokinetic accuracy and correct global intensity scaling. Perceptual Realism ( 퐿푃퐼푃푆 ): To overcome the deter- ministic blurring inherent to purely pixel-wise objectives, we evaluate LPIPS (Zhang et al., 2018) loss between the decoded images 푥 푝표푠푡 and ̂푥 푝표푠푡 . By comparing deep feature activations, this term enforces high-level structural coher- ence and prevents regression-to-the-mean artifacts. Spectral Fidelity ( 퐹퐿 ): Standard spatial losses often fail to resolve chaotic, high-frequency micro-vasculature. We incorporate a Focal Frequency Loss (FFL) (Jiang et al., 2021) to explicitly penalize discrepancies in the Fourier domain. Let denote the 2D orthogonal Fast Fourier Trans- form. We map the predicted and ground-truth latents through the frozen decoder and penalize the logarithmic difference of their amplitude spectra: 퐹퐿 = ‖ ‖ ‖ log(1 +|((̂푧 푝표푠푡 ))|) − log(1 +|((푧 푝표푠푡 ))|) ‖ ‖ ‖ 1 This logarithmic scaling ensures that the network is mathe- matically incentivized to recover peripheral high-frequency clinical textures rather than being overpowered by the zero- frequency (DC) component. The values of 휆 푀푆퐸 , 휆 퐿푃퐼푃푆 , and 휆 퐹퐿 were empir- ically determined on the validation set as 1, 5, and 50, respectively. 3.5. One-Step Inference Strategy During inference, given a pre-contrast anchor 푧 푝푟푒 and a requested target temporal phase 휏, the model evaluates the state at the maximum degradation timestep 푡 푚푎푥 (effectively resulting in pre-contrast sequence with added noise 휎). To preserve the generative capacity and mimic realistic scanner textures while ensuring smooth temporal consistency across continuous samples of 휏, we inject a fixed, patient-level latent noise map 휖 푓푖푥푒푑 . The noisy input state is constructed as: 푧 푖푛푝푢푡 = 푧 푝푟푒 + (푤 푡 푚푎푥 ⋅ 휎 푛표푖푠푒 )휖 푓푖푥푒푑 , where푤 푡 푚푎푥 represents the terminal degradation weight and 휎 푛표푖푠푒 controls the scale of the injected generative variance. The U-Net receives this noisy input state 푧 푖푛푝푢푡 , concate- nated with the clean anatomical anchor 푧 푝푟푒 , the timestep 푡 푚푎푥 , and the continuous temporal condition 휏. In a single predictive pass, the network outputs the predicted residual Δ̂푧, reconstructing the fully enhanced target endpoint as ̂푧 푝표푠푡 = 푧 푝푟푒 + Δ̂푧. Finally, the synthesized latent repre- sentation is mapped back to the high-resolution pixel space via the frozen VAE decoder, yielding the synthetic image ̂푥 푝표푠푡 = (̂푧 푝표푠푡 ). 4. Implementation Details This section describes the datasets, evaluation metrics, and baseline methods. Additional information, including data preprocessing, and network hyperparameters, can be found in Appendix A. 4.1. Datasets We use two publicly available datasets as well as one private dataset for training and validating our method. MAMA-MIA: As detailed in (Garrucho et al., 2025), the MAMA-MIA dataset is a multi-center cohort of 1,506 patients featuring pre-treatment T1-weighted DCE breast MRIs. Compiled from the ISPY-1 (Newitt et al., 2016), ISPY-2 (Li et al., 2022), NACT (Newitt and Hylton, 2016), and Duke-Breast Cancer MRI (Saha et al., 2021) collections, the dataset demonstrates high technical diversity, including S. Joshi et al.: Preprint submitted to ElsevierPage 5 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport (a) Intensity Histogram(b) Peak Enhancement Time(c) Spectral Difference Figure 3: Characterization of the domain shift between the internal and external validation cohorts. (a) Log-scaled marginal intensity histograms of the post-contrast phase exhibit general macroscopic alignment, indicating relatively comparable global contrast distributions. (b) However, the probability density function of peak enhancement time (휏) reveals a severe temporal domain shift; the external cohort demonstrates significantly faster pharmacokinetic wash-in dynamics compared to the broader, delayed uptake of the internal dataset. (c) The 2D log power spectral difference between the cohorts highlights high-frequency spatial discrepancies, characteristic of differing scanner hardware, resolution, and acquisition protocols. various scanner vendors (GE, Siemens, Philips), magnetic field strengths (1.5T and 3T) and multiple acquisition planes (axial, sagittal). More importantly, MAMA-MIA provides expert tumor segmentations as well as harmonized acqui- sition times extracted in a uniform format from all the collections. Following the provided protocol, we maintain the designated split, utilizing 300 cases for the internal validation set. Duke-Breast-Cancer-MRI (Saha et al., 2021): This dataset was collected from 2000-2014 and contains MRI scans of 922 patients acquired from two vendors, namely GE and Siemens, with 7 scanners of 1.5T and 3T magnetic field strengths from a single center in the United States. The acquisition plane is axial. MAMA-MIA includes 259 cases from this study. We incorporate the additional 663 cases from this collection in our training set. Karolinska Instituet: This private dataset was collected from 2012-2020 and contains MRI scans of 192 patients, acquired from a single vendor GE with 1.5T and 3T magnetic field strengths from a single center in Sweden. The acquisi- tion plane is axial. We use this dataset to perform external validation and evaluate our model under domain shift. 4.2. Metrics We evaluate the synthetic post-contrast MRIs using eight complementary metrics: MSE quantifies pixel-wise recon- struction fidelity. Peak Signal-to-Noise Ratio (PSNR) mea- sures the overall reconstruction quality relative to the ref- erence image. Structural Similarity Index Measure (SSIM) (Wang et al., 2004) evaluates the preservation of local struc- tural information, accounting for luminance, contrast, and texture consistency. LPIPS (Zhang et al., 2018) measures perceptual similarity using deep feature representations. To assess distribution-level realism, we report DINOv2-based Fréchet Distance (FID-DINOv2) (Heusel et al., 2017; Oquab et al., 2023). Specifically, we replace Inception features with self-supervised DINOv2 embeddings to better capture semantic and structural similarity in medical images. We also report the Fréchet Radiomics Distance (FRD) (Konz et al., 2026), computed from radiomic features extracted within the expert-annotated tumor masks. Unlike image- level metrics, FRD evaluates whether the generated images preserve clinically relevant tumor characteristics, including intensity statistics, texture, and shape-related descriptors, providing an assessment of radiomic fidelity. Furthermore, the quantitative evaluation of temporal dynamics in contrast synthesis remains largely absent from current literature. To address this gap and validate the tem- poral realism of the generated pharmacokinetic curves, we evaluate all methods using two targeted metrics: Time-to-Peak Error (PTE): In clinical BI-RADS as- sessment, the Time-to-Peak (TTP) is a critical diagnostic biomarker. To quantify temporal shift, we define the Peak Tim- ing Error (PTE) as the mean absolute difference in sec- onds between the predicted and ground-truth time-to-peak across all temporal curves in the validation set. Let 푇 = 휏 1 ,휏 2 ,...,휏 퐾 be the set of acquisition times, and 푐(휏) be the mean intensity of the tumor region at time 휏. The error is defined as: TTPE = 1 푁 푁 ∑ 푖=1 | | | | | argmax 휏∈푇 푐 (푖) 푝푟푒푑 (휏) − argmax 휏∈푇 푐 (푖) 푔푡 (휏) | | | | | A value of zero indicates that the model correctly identified the peak enhancement phase, while non-zero values corre- spond to a misalignment of one or more DCE phases (typi- cally multiples of ∼80–100 s, depending on the acquisition protocol). Dynamic Time Warping (DTW): To measure shape preservation independent of localized temporal distortions, we employ normalized Dynamic Time Warping (DTW)(Chung et al., 2025a) on the mean intensity value inside the ground truth mask. For a predicted sequence 푃 and ground truth sequence 퐺, the optimal alignment path 휋 is computed to minimize the cumulative 퐿 1 cost, normalized by the S. Joshi et al.: Preprint submitted to ElsevierPage 6 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 1 Quantitative evaluation of pharmacokinetic synthesis across pixel-level accuracy, structural fidelity, perceptual realism, and temporal alignment metrics. Baseline refers to the results on the pre-contrast image. MSE, DTW and DTW-ROI are reported in 10 −2 scale. pix2pix model does not consider acquisition time 휏 during training. Therefore, temporal metrics are not reported for this method. Best results are highlighted in bold, second-best are underlined . Statistical significance was evaluated using the Wilcoxon signed-rank test for all metrics, with the exception of PTE, which was assessed using a binomial sign test. Unless otherwise indicated, all reported values are statistically significant at 푝 < 0.001. Markers denote lower significance levels: * indicates 푝 < 0.01, ** indicates 푝 < 0.05, and † indicates a statistically insignificant difference. MethodMSE↓PSNR↑SSIM↑LPIPS↓FID↓FRD↓PTE↓DTW↓DTW-ROI↓ Internal Validation Baseline1.93 (1.39)18.24 (3.25)0.70 (0.11)0.19 (0.05)166.685.12265.30 (114.43)2.91 (1.39)18.55 (5.46) U-Net1.01 (0.67)20.84 (2.81)0.71 (0.11)0.19 (0.05)184.835.3441.24 (83.58) † 0.84 (0.77) ∗ 6.18 (4.91) pix2pix1.05 (0.60)20.36 (2.21) 0.73 (0.08) 0.18 (0.04)274.564.50– CCNet1.60 (0.87)18.60 (2.43)0.58 (0.10)0.26 (0.05)205.927.40101.46 (108.67)1.21 (0.83)8.84 (0.05) TeNCA0.99 (0.65)20.90 (2.69) 0.73 (0.11)0.21 (0.05)181.665.5774.47 (110.09) ∗ 1.03 (0.86)5.97 (4.94) Ours0.84 (0.55) 21.61 (2.77) 0.71 (0.11) 0.18 (0.05) 154.564.98 44.57 (81.59)0.68 (0.56)3.76 (3.33) External Validation Baseline2.52 (1.62)17.20 (3.73) 0.69 (0.09) ∗ 0.20 (0.04) 220.345.03237.24 (104.89)3.31 (1.53)19.44 (4.54) U-Net1.26 (0.71) † 19.75 (2.71)0.70 (0.09)0.20 (0.04)257.744.9343.59 (82.23) † 1.11 (0.84) † 7.32 (4.67) † pix2pix1.22 (0.55)19.55 (1.90) 0.73 (0.06) 0.19 (0.03) † 345.824.57– CCNet1.95 (1.01)17.68 (2.29)0.57 (0.09)0.26 (0.04)254.075.44106.12 (95.11)1.27 (0.65)9.60 (4.60) TeNCA1.11 (0.56) † 20.12 (2.40)0.72 (0.08)0.21 (0.04)244.235.2240.29 (77.88) † 1.00 (0.62) † 8.25 (4.65) Ours1.08 (0.57) 20.32 (2.57) 0.69 (0.09) 0.19 (0.04) 226.264.7837.86 (70.13)0.98 (0.68)6.28 (3.95) sequence lengths: DTW(푃,퐺) = 1 |푃| +|퐺| min 휋 ∑ (푖,푗)∈휋 |푃 푖 − 퐺 푗 | A lower DTW score indicates that the generative model suc- cessfully learned the continuous pharmacokinetic manifold, synthesizing biologically plausible enhancement trajectories even if subtly shifted in absolute time. 4.3. Baseline Methods To rigorously evaluate our proposed framework, we compare it against a diverse set of baseline and state-of- the-art techniques spanning multiple generative paradigms. First, we establish a standard U-Net (Ronneberger et al., 2015; Schreiter et al., 2024) which models sequential post- contrast sequences through different output channels. Sec- ond, we evaluate the adversarial pix2pix (Osuala et al., 2025) architecture. Because this model does not incorporate continuous acquisition time conditioning, we exclude it from temporal metric evaluations and assess it solely on spatial image fidelity. Third, we compare against CCNet (Osuala et al., 2024), a latent diffusion model conditioned on both pre-contrast anatomy and acquisition time. Because CCNet synthesizes contrast by sampling from random noise, it exhibits high stochasticity and frame-to-frame variance. Finally, we benchmark against Temporal Neural Cellular Automata (TeNCA) (Lang et al., 2025), a deterministic approach based on neural cellular automata that generates smooth, dense temporal trajectories, originally designed to overcome the stochastic limitations of diffusion models like CCNet. U-Net and the proposed method rely on residual prediction, while the remainder of the methods directly predict post-contrast phases. 5. Results This section presents results on in-domain dataset, ex- amines out-of-domain generalization, and ablates different components of the proposed method. The quantitative re- sults are summarized in Table 1. 5.1. In-Domain Performance At the pixel level, our method strictly dominates tra- ditional reconstruction metrics, achieving the lowest MSE (0.84 x 10 −2 ) and highest PSNR (21.61). While determin- istic continuous-time models (TeNCA) and 2D adversarial networks (pix2pix) marginally outperform our framework on SSIM (0.73 vs. 0.71), this stems from inherent algorith- mic biases: TeNCA favors smooth, regression-to-the-mean approximations, whereas pix2pix relies on hyper-realistic, albeit hallucinatory, textural synthesis (see Figure 4). This adversarial optimization also rewards pix2pix with lowest FRD of 4.50 on the test set, followed by our method at 4.98. Our formulation prioritizes true high-frequency structural coherence. This is quantitatively validated by our superior performance on LPIPS (0.18) and global feature distribution via FID-Dinov2 (154.56). Ultimately, the most significant advantage of our framework lies in its preservation of phar- macokinetic dynamics; our method drastically outperforms all baselines in temporal alignment, achieving a DTW of 0.68 (x 10 −2 ) and a highly localized DTW-ROI of 3.76 (x 10 −2 ). Our framework establishes a new state-of-the-art across the internal MAMA-MIA validation cohort, uniquely balancing spatial fidelity, perceptual realism, and temporal dynamics where existing baselines force a compromise. Dense temporal contrast synthesis results can be viewed in Supplementary Videos 1 - 3. S. Joshi et al.: Preprint submitted to ElsevierPage 7 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 4: Qualitative comparison of synthetic contrast generation. This figure presents highly challenging cases where tumor margins are unclear in the pre-contrast image, complicating accurate contrast injection for all models. The images displayed are selected and correspondingly synthesized at the earliest acquisition time available in the ground truth. Additional comparisons are provided in Appendix Figure A.1. Columns from left to right: non-enhanced pre-contrast input, baseline methods (U-Net, pix2pix, CCNet, TeNCA), our method, and Ground Truth (GT). U-Net and TeNCA exhibit over-smoothing. pix2pix and CCNet demonstrate unpredictable behavior, capturing the structure in some instances but altering the underlying anatomy in others. Although our method also struggles to replicate exact tumor boundaries compared to the GT, it improves spatial fidelity and tumor localization. Note that the figure presents magnified views centered directly on the tumor for better visibility, Images are normalized to [0, 255] for display. 5.2. Out-of-Domain Performance To fully contextualize the performance on the exter- nal KI cohort, we analyze the fundamental domain shifts between the two institutions in Figure 3. First, the global post-contrast intensity distributions (Figure 3a) exhibit near- perfect alignment. This confirms the absence of macroscopic intensity covariate shifts; the physiological boundaries of contrast uptake remain strictly consistent across domains. This macroscopic alignment directly explains our model’s robust preservation of structural realism (LPIPS: 0.19) and overall pixel accuracy. However, density estimation of the acquisition priors (Figure 3b) exposes a severe temporal concept shift. The external KI protocol captures peak enhancement signifi- cantly earlier (휏 ≈ 280s) than the internal MAMA-MIA training distribution (휏 ≈ 450s). While our temporal align- ment metrics exhibit a predictable degradation compared to our internal validation baseline (DTW: 0.68 vs. 0.98), our method still achieves the best absolute temporal performance (Table 1). Under this severe domain shift, the differences in global temporal alignment (PTE and DTW) between our approach, U-Net, and TeNCA do not reach statistical significance (푝 > 0.05). However, our framework maintains a statistically significant advantage over CCNet globally. Furthermore, within the localized tumor regions (DTW- ROI), our approach significantly outperforms both TeNCA and CCNet (푝 < 0.001), demonstrating robust preservation of tumor-specific kinetic fidelity despite the accelerated, out- of-distribution injection protocol. Finally, while macroscopic statistics align, the 2D Spec- tral Difference Map (Figure 3c) reveals a highly directional high-frequency covariate shift. The distinct spectral lobes indicate that the external scanner possesses a fundamentally S. Joshi et al.: Preprint submitted to ElsevierPage 8 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 2 Ablation study. Evaluating the impact of predicting subtraction, stochastic regularization and different loss components. Sub. refers to using subtraction between post-contrast and pre-contrast images as the prediction target, instead of the post-contrast image. 휖 refers to the added noise. 퐿푃퐼푃푆 and 퐹퐿 refer to the perceptual and spectral losses, respectively. Best results are bolded. MSE, DTW and DTW-ROI are reported in 10 −2 scale. Sub.휖 퐿푃퐼푃푆 퐹퐿 MSE↓PSNR↑SSIM↑LPIPS↓FID-Dinov2↓FRD↓PTE↓DTW↓DTW-ROI↓ ✓0.93 (0.53) 20.98 (2.49) 0.71 (0.11) 0.18 (0.05)162.835.1745.35 (77.37) 0.94 (0.64) 3.48 (3.47) ✓1.06 (0.63) 20.53 (2.65) 0.68 (0.12) 0.22 (0.06)357.527.8148.94 (80.26) 1.70 (1.07)3.54 (3.32) ✓0.89 (0.54) 21.31 (2.66) 0.71 (0.11) 0.18 (0.05)152.145.1047.87 (78.60) 0.80 (0.61)4.02 (3.61) ✓0.86 (0.54) 21.47 (2.74) 0.71 (0.11) 0.18 (0.05)145.065.0545.65 (83.22) 0.77 (0.62)3.65 (3.23) ✓ 0.84 (0.55) 21.61 (2.77) 0.71 (0.11) 0.18 (0.05)154.564.98 44.57 (81.59) 0.68 (0.56) 3.76 (3.33) Figure 5: Qualitative evaluation of added ablation components. These images are normalized to range [0, 255] for display. different microscopic noise floor and k-space reconstruc- tion profile. Therefore, the degradation observed in deep- feature metrics like FID (226.26) on the external cohort is likely an artifact of localized hardware noise and disparate institutional statistics, rather than a failure of the network’s physiological synthesis. 5.3. Ablation Study 5.3.1. The Impact of Loss Components Our ablation study (Table 2) and corresponding qualita- tive visualizations (Fig. 5) demonstrate the critical, additive benefits of integrating structural, perceptual, and frequency- based constraints into the synthesis pipeline. Training with 푀푆퐸 maps the macro-level contrast uptake, but visually results in a highly over-smoothed, blurry lesion devoid of internal heterogeneity. Introducing perceptual loss mitigates this blurring, recovering structural dimension and sharpen- ing macroscopic tumor boundaries. The subsequent addi- tion of stochastic regularization (+ noise) injects essential micro-textural realism into the tissue, visually harmoniz- ing the generated distribution and quantitatively improving deep-feature semantic fidelity (optimizing FID-Dinov2 to 145.06). Finally, the integration of Fourier-domain guidance proves essential for recovering high-frequency morphological de- tails. As observed visually, the Fourier constraint improves contrast in fine vascular structures surrounding the tumor area, closely aligning the generated output with the ground truth. Quantitatively, this comprehensive formulation yields the best overall pixel fidelity (MSE: 0.0084, PSNR: 21.61), radiomic texture preservation (FRD: 4.97), and temporal pharmacokinetic accuracy (DTW: 0.0068). 5.3.2. Effect of pre-contrast concatenation Table 3 evaluates the effect of the explicit morphological anchor (푧 푝푟푒 ) on generative fidelity. The unanchored model Table 3 Ablation Study. Effect of Pre-conditioning. Best results are highlighted in bold. MSE, DTW and DTW-ROI are reported in 10 −2 scale. Metricw/o pre conditioningw pre conditioning MSE↓1.04 (0.66)0.84 (0.55) PSNR↑20.67 (2.73)21.61 (2.77) SSIM↑0.71 (0.11)0.71 (0.11) LPIPS↓0.18 (0.05)0.18 (0.05) PTE↓58.94 (92.79)44.57 (81.59) FID-Dinov2↓141.87154.56 FRD↓5.674.98 DTW↓1.08 (1.01)0.68 (0.56) DTW-ROI↓4.45 (3.73)3.76 (3.33) achieves a lower FID-Dinov2 score (141.91 vs. 155.17) and a similar SSIM. By synthesizing both baseline anatomy and contrast through a single pathway, the unanchored network produces smooth deep-feature representations. However, ex- plicitly conditioning the vector field on 푧 푝푟푒 improves the clinical and temporal metrics. The anchored network focuses on contrast synthesis, achieving better pixel-level calibration (MSE: 0.0084), radiomic texture preservation (FRD: 4.97), and reduced trajectory error (DTW: 0.0068 vs. 0.0098). Furthermore, while both CCNet and our approach use stochastic noise to model output distributions, the pre- contrast anchor and the amount of injected noise dictate its specific function in image generation. CCNet initializes from pure noise conditioned on the pre-contrast image and acqui- sition time. As shown in Figure 6, while this conditioning preserves the macroscopic structure of the breast, different random noise result in varying tumor extents and anatomical hallucinations, across four predictions at a given acquisition time. In contrast, our method uses the pre-contrast image S. Joshi et al.: Preprint submitted to ElsevierPage 9 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 6: Qualitative comparison of multiple stochastic pre- dictions. Columns refer to pre-contrast image, synthesized images from CCNet (Osuala et al., 2024) and our method respectively, and GT (real post-contrast image). Four rows correspond to four different random noise initializations. In C- Net, the injected noise leads to structurally different anatom- ical predictions, including the shape of the tumor region. In the proposed method, anatomy is strictly preserved and noise accounts for subtle contrast variations and inherent MRI noise. Note that the figure presents magnified views centered directly on the tumor for better visibility, These images are normalized to range [0, 255] for display. as an anchor to retain the macroscopic integrity of both the anatomy and the pathology. Consequently, the injected noise models epistemic uncertainty, manifesting as realistic variations in contrast and standard MRI noise. The relevance of this property for uncertainty estimation is discussed in Section 5.5. 5.3.3. How to add noise for temporal continuity? Standard clinical practice typically acquires 4-5 post- contrast phases during MRI exam. Transitioning from this sparse sampling to modeling a dense temporal trajectory requires balancing spatial realism with temporal continuity. While injecting latent noise mimics uncertainty in contrast and scanner texture, independent random sampling across frames causes non-physiological flickering. We compared two noise injection strategies during inference, Independent random noise per phase restores spatial texture but disrupts temporal continuity. Patient-level noise holds a single noise map constant across all phases for a given patient, ensuring the temporal condition solely drives image changes. While we compare CCNet with independent noise to replicate the original setting of the paper, we also demonstrate that patient-level noise can improve the method to yield smooth trajectories. Specifically, this can be done because CCNet is based on denoising diffusion implicit models, which do not Table 4 Quantitative downstream segmentation performance on the MAMA-MIA internal validation set. Results evaluate the performance on synthesized images against real pre-contrast (Baseline) and real post-contrast (Upper Bound) targets. Statistical significance was computed via a paired Wilcoxon signed-rank test (푝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). MethodDice Coefficient↑HD95↓ Segmentation model trained on post-contrast Baseline0.17 (0.30) 160.69 (104.48) U-Net0.44 (0.36) 82.20 (104.68)* pix2pix0.22 (0.33) 147.79 (108.24) CCNet0.45 (0.35) 76.81 (102.32) TeNCA0.30 (0.35) 122.31 (110.69) Ours0.51 (0.37)68.78 (99.62) Upper Bound0.68 (0.33)38.59 (77.73) Segmentation model trained on pre-contrast Baseline0.49 (0.37)71.48 (99.90) U-Net0.56 (0.35)53.05 (89.01) pix2pix0.44 (0.35)70.10 (96.55) CCNet0.44 (0.35) 78.51 (102.96) TeNCA0.51 (0.36)62.59 (93.86) Ours0.60 (0.33)43.38 (80.12) Upper Bound0.63 (0.34)47.46 (84.74) require noise injection at each denoising step. In Supplemen- tary Videos 4 - 7, we show the comparison for both CCNet and the proposed method with different noise strategies. We observe that patient-level noise results in smoother trajec- tories for both methods. In contrast to our method, which maintains reasonable consistency across different random noise seeds, CCNet exhibits a high degree of variance and becomes unreliable due to its inherent dependence on noise, reinforcing the observations made in the previous section. 5.4. Downstream Tumor Segmentation To evaluate the preservation of biological structures, we compare downstream tumor segmentation performance on real and generated scans (Table 4). Specifically, two 2D nnUNet (Isensee et al., 2021) mod- els are trained on the MAMA-MIA training set, on (a) first post-contrast images only, and (b) pre-contrast images only. We evaluate downstream segmentation using the Dice coef- ficient and the 95th percentile Hausdorff Distance (HD95). These metrics are reported exclusively for the first post- contrast phase, as the ground truth masks were delineated on this specific timepoint. Training details, extended evalu- ations across all subsequent temporal sequences, as well as comprehensive results on the external validation cohort, are detailed in the Appendix B. Can a segmentation model trained on real post contrast images find the tumor on virtually contrast injected images? S. Joshi et al.: Preprint submitted to ElsevierPage 10 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 7: Qualitative downstream segmentation results. Each row depicts a distinct patient case across real and synthesized images. Note that the figure presents magnified views centered directly on the tumor for better visibility, These images are re-normalized to range [0, 255] for display. The accompanying Kernel Density Estimation (KDE) maps (far right) illustrate the intensity distribution within each ground truth masks. The top row features a large, highly lobulated malignant lesion exhibiting limited signal on the pre-contrast scan. The pix2pix model artificially spikes pixel intensity (evidenced by the sharp green peak in the KDE plot) in and around the tumor, resulting in poor boundary discrimination. CCNet hallucinates high level of texture and overestimates the tumor region. While U-Net underestimates the tumor extent and TeNCA demonstrates moderate localization, our model achieves the closest alignment with the ground truth (GT) in both morphological boundary and intensity distribution. The middle row presents a diffuse, lower-contrast malignant lesion, visible merely as a darker hypointense region in the baseline scan but still, correctly identified by the pre-contrast trained model. Here, CCNet, and U-Net underestimates the region, pix2pix fails to provide meaningful enhancement, and TeNCA suffers from a spatial hallucination (false positive). Conversely, our model successfully recovers the broad enhancement profile, matching both the GT boundaries and the true KDE distribution. The bottom row depicts a distinct malignant focal mass. Although pix2pix achieves a KDE profile closer to the GT, it suffers from severe image degradation and, similar to TeNCA, massively overestimates the tumor boundary. While U-Net, CCNet and our proposed method accurately localize the mass, our method demonstrates superior statistical consistency with the true post-contrast KDE profile. On the pre-contrast baseline, the post-contrast trained segmentation network effectively fails to localize the le- sion (Dice: 0.17), establishing its critical reliance on en- hancement. While the competing methods including pix2pix and TeNCA previously demonstrated high SSIM, their cor- responding downstream segmentation performances only show a small improvement over the non-enhanced base- line, In contrast, our proposed method explicitly preserves biological boundaries, achieving the highest segmentation performance (Dice: 0.51, HD95: 68.78) and recovering the vast majority of the true clinical signal (Upper Bound Dice: 0.67). If the goal is to avoid contrast, why not simply rely en- tirely on pre-contrast scans? In other words, why synthesize at all? Segmentation networks trained directly on post-contrast data rely highly on the magnitude of pixel intensity (Joshi et al., 2024), as evident by low performance on pre-contrast data. Consequently, they also have higher sensitivity to vari- ations in contrast dynamics, for instance, degree of contrast uptake. Conversely, as shown in Table 4, a network explicitly trained on pre-contrast data learns robust morphological and textural priors and already localizes tumor better in absence of contrast signal (Dice: 0.49). When we route these images through the generative frameworks, synthetic contrast sur- faces the latent physiological structures, and all determinis- tic methods improve over the baseline. Our method performs the best with Dice of 0.60 and an HD95 of 43.38 (Upper bound: 0.63), thereby maximizing the algorithmic utility of raw scans while bypassing the physiological risks of actual gadolinium. Visual inspection (Figure 7) shows three representative examples where our model produces intensity distributions (pink) that closely mirror the true post-contrast profiles (green), and yields automated segmentation boundaries that conform to the ground truth. Additionally, our proposed framework exhibited the lowest absolute failure rate across all evaluated methods and training paradigms, yielding com- plete missed detections in only 29 of 300 cases. This perfor- mance was second only to the true post-contrast ground truth S. Joshi et al.: Preprint submitted to ElsevierPage 11 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure 8: Clinical evaluation across varying case complexities. The figure contrasts a representative low-complexity case (Top) with a high-complexity case (Bottom) from the reader study. For both cases, the image panels display the pre-contrast baseline, the predicted synthetic time-series, the corresponding pixel-wise uncertainty maps (휎), and the ground truth (GT) time-series. The kinetic plots illustrate the enhancement dynamics, comparing the predicted mean tumor intensity against the ground truth over time. The consensus tumor characterization across synthetic and real images, and corresponding diagnostic impact are summarized for each case in text format. (16/300). A detailed analysis of catastrophic segmentation failures (Dice = 0) is provided in the Appendix Section B.4 and qualitative examples are added to Appendix Figure B.2. 5.5. Clinical Reader Study To systematically evaluate the diagnostic equivalence, realism, and clinical utility of the synthetic contrast-enhanced MRI images, we conducted a comprehensive multi-part clinical reader study with four radiologists, with 5, 8, 13, and 15 years of experience respectively, evaluating randomly selected forty cases across three stratification variables: tumor size, tumor shape and enhancement pattern. The study design, reader profile, statistical tests and additional results are added in Appendix C. Readers can directly access platform here: https://smriti-joshi.github.io/reader-stu dy-contrast-synthesis/. Which method do the readers prefer? In a direct comparative evaluation against established temporal contrast synthesis baselines (UNet and TeNCA), readers demonstrated an strong preference for the proposed method. Specifically, out of 160 total evaluations (40 cases x 4 readers), our method was selected as the superior image in 83.8% of cases (134 votes). In contrast, the compared architectures tied for a second, with UNet and TeNCA each capturing 8.1% (13 votes) of the total preference. Statistical analysis confirmed that this dominance was highly signif- icant (푝 < 0.001) and not due to chance. This preference for our method was consistent across all four individual readers, while the two alternative methods were statistically indistinguishable from one another (푝 = 1.0). How do real and synthetic image characteristics com- pare? The reader study comparing synthetic and ground truth images demonstrated a robust overall mean agreement rate of 69.8% (median 66.7%). As seen in Table 5, coarse categor- ical judgments showed substantial GT-to-synthetic agree- ment, whereas margin and internal-enhancement assess- ments were only moderate. Readers noted a divergence in ordinal assessments such as image quality (56.2% agree- ment) and kinetic plausibility (48.8%), where real images maintained a clear scoring advantage. Nonetheless, for both criterion, synthetic and real images are scored between 2 and 3, suggesting that the readers found the kinetic curves plau- sible and the image quality acceptable on average (further demonstrated in Appendix Figure B.3). How does synthetic data impact patient care? In the side-by-side review of real and synthetic images, readers judged that relying on the synthetic image would produce a major change in clinical management in 30.0% of assessments. However, 34.4% of evaluations resulted in minor diagnostic deviations that did not alter the clinical S. Joshi et al.: Preprint submitted to ElsevierPage 12 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table 5 Lesion characterization: GT vs. synthetic. Top: categorical tasks (Cohen’s 휅 with bootstrap 95% CI). Bottom: ordi- nal tasks (GT/synthetic means, mean difference Δ, signed rank-biserial 푟, Wilcoxon 푝). All 푝-values survive Benjamini– Hochberg correction. Categorical taskN% agreeCohen’s 휅95% CI Lesion Type16086.20.72[0.62, 0.83] Shape11989.90.73[0.57, 0.86] Margins15377.10.45[0.30, 0.59] Enhancement16068.80.43[0.31, 0.55] Ordinal taskGTSynthΔ푟푝 Kinetic Plausibility 2.712.27-0.43 -0.62 <0.001 Image Quality2.812.37-0.44 -0.77 <0.001 Table 6 Predictors of clinical impact. Proportional-odds ordinal logistic regression; standardized predictors. OR > 1 indicates higher odds of greater clinical impact. PredictorOR95% CI푝 Perceived synthetic appearance 3.60 2.26–5.74 <0.001 Diagnostic complexity2.14 1.46–3.12 <0.001 Image quality0.86 0.56–1.33 0.507 Reader experience0.54 0.38–0.77 <0.001 Tumor size (large)0.82 0.59–1.14 0.235 Shape (irregular)1.24 0.89–1.72 0.197 Peak enhancement (late)0.84 0.60–1.18 0.318 management plan, and 35.6% resulted in no change at all. In other words, in 70.0% of the cases evaluated, replac- ing real images with synthetic ones would not negatively alter patient care. Further analysis using a proportional- odds ordinal logistic regression model identified perceived synthetic appearance (OR 3.60, 푝 < 0.001) and diagnostic complexity (OR 2.14, 푝 < 0.001) as significant independent predictors of greater clinical impact (i.e., a higher likelihood of altered patient management). Conversely, higher reader work experience served as a significant protective factor against these deviations (OR 0.54, 95% CI 0.38–0.77, 푝 < 0.001). Image quality was not independently predictive of clinical impact once perceived appearance and complexity were accounted for in the model (OR 0.86, 푝 = 0.51; Table 6). How does uncertainty map affect reader confidence? When presented with synthetic image and corresponding uncertainty maps, readers indicated that they would avoid or defer physical contrast injection in 64% of the evaluated cases. To help guide these critical management decisions, the integration of uncertainty served as a valuable, though imperfect, adjunct. This map was computed by adding ran- dom noise for 10 different iterations to produce slightly different outputs. As seen in Figure 8, the areas where the predictions agreed were displayed with high model con- fidence (black) while disagreement indicated uncertainty (yellow). In the reader study, map provided actionable guid- ance in 49% of evaluations overall, meaningfully shifting radiologist confidence by either validating trustworthy syn- thetic images (increasing confidence in 32% of cases) or flagging unreliable ones (decreasing confidence in 17%). These shifts directly influenced downstream actions: when the map decreased confidence, it acted as an effective safety mechanism, prompting readers to request true contrast injec- tion in 88% of those instances. Conversely, when it increased confidence in unchanged assessments, 82% of readers pro- ceeded without requesting real contrast. Furthermore, statistical analysis revealed that the map’s directional influence is significantly modulated by case com- plexity (Kendall’s 휏 = 0.218, 푝 = 0.003). While the map predominantly increased confidence in low-complexity cases (39% increase vs. 11% decrease), it functionally in- verted in high-complexity scenarios, where it more fre- quently decreased confidence (41% decrease vs. 14% in- crease) to flag unreliable generations. Consequently, its over- all informative rate peaked during these highly complex evaluations (55%). This utility also demonstrated a trending association with lesion type (푝 = 0.059), proving especially beneficial for challenging non-mass enhancement (NME) lesions, which achieved a 67% informative rate. However, these metrics must be interpreted with distinct clinical cau- tion. Inter-reader agreement regarding the map’s effect was notably low (Fleiss’ 휅 = 0.02–0.10), indicating that reliance on the map remains highly subjective and reader-dependent. Furthermore, qualitative feedback highlighted specific lim- itations with the map’s reliability, noting critical instances where it incorrectly displayed high confidence despite the underlying synthetic image being clinically inaccurate. This susceptibility to overconfident errors underscores that while the uncertainty map is a helpful decision-support tool in the aggregate, it is not a definitive fail-safe, and must continue to be scrutinized carefully to prevent diagnostic missteps. 6. Discussion and Conclusion This work presents a conditioned latent transport frame- work for single-step DCE-MRI contrast synthesis. Quanti- tative evaluations demonstrate that the approach effectively balances spatial fidelity with precise pharmacokinetic tem- poral alignment, outperforming baseline models on internal data. Independent external validation evaluates the model’s adaptability to varying scanner noise profiles and tempo- ral acquisition shifts. While the framework demonstrates promising resilience, the performance degradation under differing clinical protocols indicates that cross-institutional generalization remains an ongoing challenge Furthermore, the ablation studies evaluate the design choices required for temporal contrast synthesis. Specifi- cally, anchoring to a pre-contrast image enforces physio- logical adherence, allowing the model to focus on temporal synthesis, individual loss components improve visual and temporal fidelity, while a fixed noise strategy maintains continuity in densely sampled predictions. In the downstream tumor segmentation task, our pro- posed framework outperforms competing state-of-the-art methods across two distinct evaluation strategies. Success on a network natively trained on post-contrast images confirms S. Joshi et al.: Preprint submitted to ElsevierPage 13 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport that the synthesized contrast uptake accurately mirrors the ground truth. Superior performance on a pre-contrast trained network demonstrates that the synthesized data strictly preserves the underlying morphological and textural char- acteristics. These quantitative successes translate directly to clinical viability. In a blinded reader study conducted by expert breast radiologists, our synthesized images were preferred over competing baselines in approximately 84% of cases. Highlighting the framework’s potential to safely reduce contrast burden, the study further confirmed that substituting real post-contrast MRI with our synthetic coun- terparts would result in no major negative deviation to patient management in 70% of cases. Qualitative feedback from the reader study indicates clinical optimism regarding the potential of synthetic contrast- enhanced MRI to reduce patient contrast burden. At the current maturity level of the state-of-the-art in contrast synthesis, conventional DCE-MRI remains indispensable for precise preoperative surgical planning. However, radiol- ogists acknowledged that upon rigorous clinical validation, the technology shows promise as a diagnostic adjunct. Specifically, in their opinion, integrating synthetic contrast with high-resolution diffusion-weighted imaging (DWI) or mammography could elevate baseline screening confidence without necessitating immediate contrast administration. This aspect is not explicitely tested in this study and can make for an interesting future direction. In addition, further model refinements must address specific morphological and kinetic blind spots. Targeted improvements should prior- itize enhancing margin sharpness in dense breast tissue, accurately reproducing internal tumor heterogeneity (e.g., central necrosis), and reliably capturing the true spatial extent of small lesions and non-mass enhancement (NME). Furthermore, it is important to increase robustness in failure detection through uncertainty estimation and other comple- mentary methods. We add one failure case from the reader study in Appendix Figure C.6 which obtained identical diag- nostic assessment, but the tumor is completely mislocalized. Strictly in terms of methodology, our work has several limitations. First, our current pipeline relies on qualitative assessments of pre- and post-contrast image registration, meaning that patient motion between acquisitions was not strictly quantified or corrected. Second, the training dataset exhibits a long-tailed temporal distribution, with sparse rep- resentation of early (< 100 s) and late (> 500 s) acquisition phases (Appendix Figure D.7). Consequently, the dataset contains fewer examples of tumors captured during the rapid wash-in period or exhibiting late-phase enhancement. We explored sampling techniques to address this imbalance, but they failed to improve synthesis quality in these low- density temporal regions without simultaneously degrading performance on the broader test set (Appendix Table D.7). Addressing this domain gap to ensure kinetic fidelity in low- data regimes remains a critical area for future investigation. Third, in our current formulation, the model’s integration timestep 푡 and the physical acquisition time 휏 are fully decoupled. Given the highly non-linear nature of contrast uptake in malignant tumors, characterized by a rapid wash- in phase and a gradual wash-out phase, future work could explicitly couple these variables by utilizing a physiological pharmacokinetic (PK) model as the interpolation function. While a strict PK-driven trajectory could enforce rigorous physical priors in highly controlled, single-center protocols, it presents a unique challenge for generalized models. Be- cause the MAMA-MIA dataset aggregates data from over 25 institutions with widely heterogeneous injection proto- cols and temporal resolutions, enforcing a single, idealized PK curve could act as an overly restrictive inductive bias. Future architectures must balance these physiological priors with the flexibility required to model multi-center clinical variance. Finally, owing to limited data availability, we did not explicitly evaluate synthetic contrast uptake in benign lesions. Because benign and malignant masses typically exhibit distinct pharmacokinetic enhancement profiles, ex- tending this framework to reliably model and differentiate between these varying pathologies remains an essential ob- jective for future clinical translation. In conclusion, the aim of this work is to contextualize the current progress of contrast synthesis breast DCE-MRI. We hope that our evaluation framework serves an anchor for future research, ensuring that progress in this domain remains transparent and clinically meaningful. We also hope that the community builds upon the gaps identified in the study to advance temporal contrast generation. CRediT authorship contribution statement Smriti Joshi: Conceptualization, Data curation, Method- ology, Investigation, Project administration, Formal Anal- ysis, Software, Writing - original draft, Writing - review & editing. Apostolia Tsirikoglou: Data curation, Formal analysis, Writing - original draft, Writing - review & editing. Daniel M. Lang: Formal analysis, Writing - review & edit- ing. Richard Osuala: Conceptualization, Writing - review & editing. Noah Márquez Vara: Software. Alejandro Guz- man: Data curation. Grzegorz Skorupko: Software. Se- bastian Ibarra Arregui: Writing - review & editing. Lidia Garrucho: Writing - review & editing. Akane Ohashi: Data curation, Validation. Dimitra Ntoula: Data curation, Validation. Eugen Divjak: Validation, Writing - review & editing. Oğuz Lafcı: Validation, Writing - review & editing. Jan C. Peeken: Writing - review & editing. Julia A. Schnabel: Writing - review & editing. Fredrik Strand: Resources, Writing - review & editing. Oliver Diaz: Super- vision, Writing - review & editing. Karim Lekadir: Super- vision, Resources, Funding acquisition, Writing - review & editing. Acknowledgements This project has received funding from the European Union’s Horizon 2020 research and innovation programme under grant agreement No 101057699 (RadioVal). Addi- tionally, this work was partially supported by the project FUTURE-ES (PID2021-126724OB-I00) and project AIMED S. Joshi et al.: Preprint submitted to ElsevierPage 14 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport (PID2023-146786OB-I00) from the Ministry of Science and Innovation of Spain. D.M.L. and J.A.S. received fund- ing from HELMHOLTZ IMAGING, a platform of the Helmholtz Information and Data Science Incubator. G.S. re- ceived funding from the European Union’s Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart). The authors acknowledge the role of Gemini 3.1 Pro for enhancing clarity in the text, as well as Claude Sonnet 4.6 and DeepSeek V4 Flash for assistance in coding. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article. 7. Ethical Approval Statement For the publicly available datasets used in this study, formal ethics committee approval and written informed con- sent were not required. For external validation data, ethical approval was obtained from the Swedish Ethical Review Authority (Etikprövningsmyndigheten) (Approval Number: 2020-00488, with subsequent amendment 2022-06777-02). The requirement for written informed consent was waived due to the retrospective nature of the study. 8. Declaration of competing interest Given their role as Medical Image Analysis Associate Editor, Julia Schnabel had no involvement in the peer-review of this article and has no access to information regarding its peer-review. Full responsibility for the editorial process for this article was delegated to another journal editor. The authors declare that they have no additional competing fi- nancial interests or personal relationships that could have appeared to influence the work reported in this paper. 9. Data Availability The MAMA-MIA dataset (Garrucho et al., 2025) and DUKE-Breast-Cancer-MRI dataset (Saha et al., 2021) are publicly available on Synapse 2 and TCIA 3 , respectively. The external validation KI dataset is private and the authors do not have permission to release it. 10. Code Availability The source code will be made publicly available upon acceptance. Appendix A Proposed Model The implementation details and hyperparameter con- figuration of the proposed architecture is detailed in this section. Figure A.1 shows additional qualitative comparison between different methods. 2 https://w.synapse.org/Synapse:syn60868042/wiki/628716 3 https://w.cancerimagingarchive.net/collection/duke-breast- cancer-mri/ A.1 Preprocessing This work follows the same preprocessing pipeline es- tablished in TeNCA (Lang et al., 2025), with resampling to a uniform voxel spacing of 1 m and intensity values are linearly rescaled between zero and one based on the 0.02 and 99.98 percentiles of the respective pre-contrast image. Additionally, images are cropped to patches of size 168×168. A.2 Network and Hyperparameters Variational autoencoder: The model is trained with a composite loss function with MSE weighted at 1.0, percep- tual (LPIPS) loss weighted at 5.0, complemented by a small KL regularization term (1e-6), a Sobel gradient-based edge- sharpness loss (0.5). Training runs for 200 epochs with a batch size of 16, a learning rate of 1e-4, and gradient clipping at 1.0. This yields a 4-channel latent space derived from a VAE with a 4× spatial compression and a latent scaling factor of 1.0259. Latent UNet: The core latent space architecture utilizes DiffusionModeUNet (from MONAI Generative (Pinaya et al., 2023)), with channel configuration of [128,256,512,512], incorporating two residual blocks per resolution stage and spatial self-attention at the three highest-resolution levels. The noise level fixed at 휎 = 0.1. We employ classifier- free guidance, applying a conditioning dropout rate of 0.1 during training and a guidance scale of 1.0 during inference. Training is conducted over 200 epochs with a batch size of 8 using the AdamW optimizer (learning rate: 1×10 −4 , weight decay: 3×10 −4 ). To ensure training stability and efficiency, we apply gradient clipping at 1.0 and utilize mixed-precision training. The composite objective function comprises an MSE loss in the latent space (휆 = 1.0), a perceptual loss (휆 = 5.0), and a focal frequency loss (휆 = 50). Finally, an exponential moving average (EMA) of the network weights is maintained with a decay rate of 0.999. A.3 The Impact of VAE reconstruction Table A.1 evaluates the effect of the autoencoder (VAE) architecture on reconstruction fidelity (Phase 1) and tem- poral contrast synthesis (Phase 2). Using a pretrained au- toencoder (Stability VAE, 8x downsampling) 4 creates a spa- tial information bottleneck. Despite medical-domain fine- tuning, the 8x compression degrades high-frequency anatom- ical structures, resulting in a Phase 1 SSIM of 0.79. A custom VAE, trained from scratch with a 4x downsampling factor, preserves these details and improves structural reconstruc- tion (SSIM: 0.92, PSNR: 30.80). This reconstruction performance directly impacts the Phase 2 generative temporal synthesis. The 4x latent space retains finer details, enabling the network to map the phar- macokinetic trajectory while maintaining spatial bound- aries. This yields improved pixel fidelity (MSE:0.84×10 −2 ) and spatiotemporal alignment across the full image and isolated lesions (DTW: 0.68, DTW-ROI: 3.76). The 8x latent 4 https://huggingface.co/stabilityai/sd-vae-ft-mse S. Joshi et al.: Preprint submitted to ElsevierPage 15 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure A.1: Additional Qualitative Results for Synthetic Contrast Generation. S. Joshi et al.: Preprint submitted to ElsevierPage 16 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table A.1 Ablation Study. Comparing autoencoder reconstruction fidelity (Phase 1) and its downstream impact on temporal synthesis (Phase 2). Best results are highlighted in bold. MSE is reported in 10 −2 scale. Metric Stability VAE (Zero-Shot) Stability VAE (Fine-tuned) VAE 4x (Scratch) Phase 1: Pure VAE Reconstruction MSE↓0.33 (0.21)0.31 (0.20)0.11 (0.08) PSNR↑25.82 (3.25)26.13 (3.27) 30.80 (3.14) SSIM↑0.78 (0.07)0.79 (0.07)0.92 (0.04) LPIPS↓0.12 (0.02)0.11 (0.02)0.10 (0.03) FID-Dinov2↓79.1378.18120.99 Phase 2: Synthesis MSE↓–0.97 (0.60)0.84 (0.55) PSNR↑–20.93 (2.67) 21.61 (2.77) SSIM↑–0.70 (0.11)0.71 (0.11) LPIPS↓–0.18 (0.05)0.18 (0.05) FID-Dinov2↓–131.19154.56 FRD↓–5.724.98 DTW↓–0.74 (0.65)0.68 (0.56) DTW-ROI↓–4.07(3.72)3.76 (3.33) space yields better perceptual deep-feature scores (FID- Dinov2: 131.19 vs. 154.56), likely due to its closer alignment with the natural-image distributions expected by the pre- trained metric. However, the 4x architecture performs better on temporal and localized metrics, including FRD (4.98 vs. 5.72). Appendix B Segmentation B.1 Training Details For downstream segmentation, we trained a 2D network using the standard nnUNet framework (Isensee et al., 2021), utilizing both the pre-contrast and first post-contrast images of the MAMA-MIA training set. The automated config- uration heuristic generated a deep 2D U-Net architecture tailored to the median image dimensions and spacing of the dataset. Preprocessing included resampling all modali- ties to an in-plane resolution of 0.703 × 0.703 m (using third-order spline interpolation for image data and nearest- neighbor for segmentation masks) alongside global Z-score normalization. Model optimization was executed on 256 × 256 pixel patches using a batch size of 49 and a batch- aggregated Dice loss function. B.2 Performance on all post contrast phases Table B.2 extends the downstream segmentation evalu- ation presented in the main manuscript by reporting perfor- mance across all synthesized post-contrast sequences, rather than exclusively the first acquisition phase. This extended analysis provides strong evidence that our generative frame- work successfully synthesizes an identifiable and persistent contrast enhancement trajectory as temporal acquisition pro- gresses. However, we note an expected slight degradation in the absolute metrics, including the real post-contrast Upper Bound, which decreases from a Dice of 0.68 to 0.62 on the post-contrast trained network. This is due to a structural limitation in the evaluation paradigm: the reference ground truth (GT) segmentation masks were clinically delineated based strictly on the first post-contrast sequence. As the contrast agent naturally diffuses into surrounding tissues during later acquisition phases, the physiological enhance- ment boundaries shift relative to this static, early-phase GT mask. Consequently, the multi-phase metrics reported here accurately reflect sustained temporal contrast identifiability, but serve only as a proxy for absolute boundary precision in late-phase dynamics. B.3 External Validation Table B.3 presents the downstream tumor segmen- tation performance evaluated on the independent, multi- institutional external validation cohort. As anticipated, the absolute quantitative metrics are lower compared to internal validation, including the real post-contrast upper bound (Dice: 0.60), which reflects the inherent domain gap when applying a fixed segmentation network to external data. Despite this overall reduction in absolute performance, the relative hierarchical trends established in the internal vali- dation remain consistent, where our proposed method main- tained robust biological feature amplification, significantly outperforming all synthetic baselines (Dice: 0.43 and 0.46 with post-contrast and pre-contrast trained segmentation networks respectively). B.4 Failure Cases An analysis of complete segmentation failures (Dice = 0) further highlights the robustness of our method. Natively trained post-contrast networks proved sensitive to any de- viation from post-contrast distribution, yielding 238 unique patient failures across all methods, compared to 150 for the robust pre-contrast network. As seen in Table B.4, our pro- posed framework exhibited the lowest catastrophic failure rate (46 cases), and even surpassing the real post-contrast ground truth (54 cases) when tested with pre-contrast seg- mentation network. Overall, we have the lowest failure rate across synthesis methods where 29/300 case were missed by both segmentation networks. We add representative failure cases of the proposed method in Figure B.2. Appendix C Clinical Reader Study To systematically evaluate the diagnostic equivalence, realism, and clinical utility of the synthetic contrast-enhanced MRI images, we conducted a comprehensive multi-part reader study. Readers can directly access platform here: https://smriti-joshi.github.io/reader-study-contras t-synthesis/. C.1 Study Design C.1.1 Selection of Cases To ensure a balanced evaluation, a subset of 40 cases was selected from the internal test set using a stratified random sampling approach. Stratification was based on three objective metrics extracted directly from the ground- truth segmentations: tumor volume, boundary sphericity, S. Joshi et al.: Preprint submitted to ElsevierPage 17 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure B.2: Failure Cases in Downstream Tumor Segmentation Task. and peak enhancement latency. To establish the phenotypic subgroups, each metric was dichotomized using a robust median split approach, where the median value across the entire test-set cohort served as the definitive cutoff threshold. Specifically, tumor volume was assessed using the maximal cross-sectional area in pixels, categorizing cases as Small (less than or equal to the median) or Large (greater than the median). Boundary sphericity was quantified using a standard circularity metric, mathematically defined as 4휋×Area Perimeter 2 , with cases classified as Irregular (less than or equal to the median) or Regular (greater than the median). Finally, peak enhancement latency was measured as the S. Joshi et al.: Preprint submitted to ElsevierPage 18 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table B.2 Downstream Segmentation Performance on Internal Validation Data (extended with evaluation on all sequences). Baseline and Upper Bound refer to inference on pre-contrast images and real post-contrast images, respectively. Statistical significance was computed via a paired Wilcoxon signed-rank test (푝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). MethodFirstAll Post-ContrastPost-Contrast Dice↑HD95↓Dice↑HD95↓ Segmentation model trained on post-contrast Baseline0.17 (0.30) 160.69 (104.48) 0.17 (0.30) 160.69 (104.48) Upper Bound 0.68 (0.33)38.59 (77.73)0.62 (0.32)51.91 (79.03) U-Net0.44 (0.36) 82.20 (104.68)* 0.44 (0.35)83.32 (98.23)* pix2pix0.22 (0.33) 147.79 (108.24) 0.22 (0.33) 147.79 (108.24) CCNet0.45 (0.35)76.81 (102.32)0.40 (0.30)84.67 (87.48) TeNCA0.30 (0.36) 122.31 (110.69) 0.25 (0.32)145.04 (94.63) Ours0.52 (0.36) 63.07 (95.85) 0.48 (0.36) 75.48 (97.23) Segmentation model trained on pre-contrast Baseline0.49 (0.37)71.48 (99.90)0.49 (0.37)71.48 (99.90) Upper Bound 0.63 (0.34)47.46 (84.74)0.63 (0.31)41.33 (71.74) U-Net0.56 (0.35)53.05 (89.01)0.57 (0.33)49.98 (83.36) pix2pix0.44 (0.35)70.10 (96.55)0.44 (0.35)70.09 (96.55) CCNet0.44 (0.35)78.51 (102.96)0.46 (0.31)67.66 (83.13) TeNCA0.51 (0.36)62.59 (93.86)0.53 (0.34)59.52 (86.86) Ours0.60 (0.33) 43.38 (80.12) 0.60 (0.33) 44.26 (79.37) Table B.3 Downstream Segmentation Performance on External Validation Data. Baseline and Upper Bound refer to inference on pre-contrast images and real post-contrast images, respectively. Statistical significance was computed via a paired Wilcoxon signed-rank test (푝 < 0.01). Results that did not reach statistical significance are marked with an asterisk (∗). MethodFirstAll Post-ContrastPost-Contrast Dice↑HD95↓Dice↑HD95↓ Segmentation model trained on post-contrast Baseline0.07 (0.21)201.92 (81.43)0.07 (0.21)201.92 (81.43) Upper Bound 0.60 (0.35)53.55 (90.56)0.54 (0.33)67.61 (89.76) U-Net0.38 (0.33)* 84.19 (105.11)* 0.38 (0.31)* 83.30 (97.52)* pix2pix0.09 (0.23)191.52 (89.97)0.09 (0.23)191.52 (89.96) CCNet0.36 (0.32)76.45 (101.29)0.33 (0.28)89.29 (81.12) TeNCA0.19 (0.30) 152.64 (107.67) 0.19 (0.29)153.31 (97.06) Ours0.43 (0.33) 59.28 (90.86) 0.40 (0.32) 76.32 (91.76) Segmentation model trained on pre-contrast Baseline0.36 (0.35) 107.26 (110.60) 0.36 (0.35) 107.26 (110.60) Upper Bound 0.51 (0.37)76.18 (101.95)0.52 (0.33)65.45 (86.69) U-Net0.39 (0.34)88.55 (105.97)0.41 (0.33)77.12 (94.96) pix2pix0.33 (0.31)95.46 (107.05)0.33 (0.31)95.46 (107.05) CCNet0.34 (0.33)93.74 (107.18)0.37 (0.31)82.77 (86.70) TeNCA0.41 (0.33)80.45 (101.59)0.37 (0.30)93.90 (87.15) Ours0.46 (0.33) 65.86 (95.69) 0.46 (0.32) 62.07 (90.57) (a) Paired Kinetic(b) Paired Quality Figure B.3: Overall composite caption describing the kinetic analysis, quality metrics, and case agreement. S. Joshi et al.: Preprint submitted to ElsevierPage 19 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table B.4 Analysis of catastrophic segmentation failures (Dice = 0). The table reports the absolute number of unique patient cases where the segmentation networks completely failed to localize the lesion, evaluated across both pre-contrast and post- contrast training paradigms. "Overlap" indicates the number of cases that failed simultaneously under both paradigms. MethodPre-ContrastPost-ContrastOverlap Network FailuresNetwork Failures(Both) Pre-contrast8620379 Post-Contrast54 4016 U-Net609645 pix2pix8618478 CCNet928751 TeNCA6315353 Ours4682 29 Lesion type Shape Margins Enhancement pattern Kinetic plausibility Image quality 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Fleiss' Inter-Reader Agreement: Fleiss' with 95% Bootstrap CI (GT vs Synthetic) =0.2 (slight) =0.4 (fair) =0.6 (moderate) GT Synthetic Figure B.4: Inter-reader agreement on GT and synthetic images (Fleiss’ 휅 with bootstrap 95% CI) across tasks. acquisition time in seconds required for the mean signal intensity inside the tumor mask to reach its maximum across all post-contrast phases, categorizing cases as Early (less than or equal to the median) or Late (greater than the me- dian). This categorization yielded eight distinct phenotypic subgroups. To prevent evaluation bias toward any single, frequently occurring tumor presentation, exactly five cases were randomly sampled from each of the eight subgroups. The detailed distribution is provided in Table C.5. C.1.2 Lesion Characterization and Image Quality In the first section, readers were presented with isolated images (either real or synthetic, presented in a blinded fash- ion) and asked to characterize the breast tumor and assess the overall image quality. The characterization was based on the standard lexicon, requiring readers to classify the lesion type as a mass, non-mass enhancement (NME), focus/foci, or in- determinate. For lesions identified as masses, readers further specified the shape (round, oval, or irregular) and margins (circumscribed, irregular, spiculated, or indeterminate). The enhancement pattern was categorized as homogeneous, het- erogeneous, rim enhancement, dark internal septations, or not assessable. Beyond standard characterization, readers evaluated the kinetic plausibility of the enhancement using a 5-point scale ranging from 1 (implausible/non-physiologic pattern) to 5 (fully physiologic). Finally, readers rated the overall quality of the MRI image, independent of the tumor, on a 5-point Likert scale ranging from 1 (non-diagnostic due to severe artifacts) to 5 (excellent high-quality image with minimal artifacts). C.1.3 Clinical Decision-Making and Diagnostic Equivalence In the second section, readers performed a side-by- side comparative evaluation of the synthetic image against the corresponding real acquisition to determine diagnostic equivalence. Readers assessed the potential impact of the synthetic image on clinical decision-making, categorizing it as causing no change, a minor change (no impact on clinical management), or a major change (affecting clinical manage- ment or diagnosis). If a major change was indicated, readers specified the primary discrepancy, such as incorrect tumor localization, inaccurate extent/size estimation, implausible enhancement kinetics, image quality issues, or missed/false detections. Additionally, readers rated the inherent diagnostic com- plexity of the case (low, moderate, or high) and evaluated the perceived synthetic appearance of the generated image on a 4-point scale from "None" (fully natural) to "Strong" (clearly synthetic). During this phase, readers were also provided with an AI-generated anomaly map indicating epistemic uncertainty. Readers reported how this maps influenced their diagnostic confidence (increased, decreased, or no change) and selected their next logical clinical step if presented with this data in a real workflow (e.g., proceed normally, review with caution, or request standard contrast injection). C.1.4 Method Preference and Comparative Benchmarking In the final phase, the proposed generation method was benchmarked against alternative state-of-the-art methods for temporal contrast synthesis. Readers were presented with the actual MRI acquisition alongside images generated by com- peting models, UNet (Ronneberger et al., 2015) and TeNCA (Lang et al., 2025). They were tasked with selecting their preferred synthetic image based on a holistic assessment of tumor characterization accuracy, enhancement quality, and overall image fidelity. An optional free-text field was pro- vided for readers to leave additional qualitative comments regarding their selection. C.1.5 Broader Implications Upon completion of the image-specific evaluations, readers participated in a final post-study survey designed to capture their broader perspectives on the current state and future utility of AI-based contrast synthesis. To assess the perceived maturity of the technology, readers were asked to evaluate how close the field is to solving the problem of contrast enhancement synthesis for clinical use, selecting from four levels of progress: very far (fundamental limita- tions remain), moderate progress (significant gaps remain), getting close (most cases are convincing), or nearly solved (clinically equivalent in most scenarios). The remainder of the survey consisted of open-ended, free-text questions aimed at gathering qualitative insights to S. Joshi et al.: Preprint submitted to ElsevierPage 20 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table C.5 Stratified case selection from the internal test set based on tumor volume, boundary sphericity, and peak enhancement latency. Tumor VolumeBoundary SphericityPeak EnhancementSelected CasesTotal Pool LargeIrregularEarly533 LargeIrregularLate558 LargeRegularEarly534 LargeRegularLate524 SmallIrregularEarly531 SmallIrregularLate528 SmallRegularEarly552 SmallRegularLate539 Table C.6 Paired vs. inter-reader agreement. GT-to-synthetic Cohen’s 휅 exceeds GT inter-reader Fleiss’ 휅 for every categorical task. TaskGT↔Synth 휅Inter-reader 휅Δ Lesion Type0.720.43 +0.30 Shape0.730.31 +0.41 Margins0.450.31 +0.14 Enhancement0.430.12 +0.31 guide future technical development and clinical implemen- tation. Readers were prompted to identify where current AI- based synthesis methods are most lacking and to specify which technical improvements they would prioritize (e.g., spatial resolution, temporal consistency, kinetic accuracy, or artifact reduction). Furthermore, the survey explored the readers’ clinical vision for the technology, specifically ask- ing for their thoughts on integrating virtual contrast MRI with standard mammography for breast cancer screening workflows, as well as its potential utility when combined with other unenhanced MRI sequences (such as diffusion- weighted imaging) for comprehensive diagnostic assess- ments. A final open-text field was provided to capture any remaining general feedback or observations regarding the study. C.2 Participants Four board-certified breast-imaging radiologists from three academic centers participated in the reader study. Sub- specialty experience ranged from 5 to 15 years (median 10.5, values 5, 8, 13, 15). Each reader evaluated all 40 cases across all four sections, yielding 2,927 analyzable responses (overall item-level missingness 12.9%, predominantly the conditional and optional items). C.3 Statistical analysis Analyses were paired at the case×reader level wherever the design permitted. Categorical agreement between GT and synthetic characterization was quantified with Cohen’s 휅 (unweighted for nominal tasks, linear-weighted for ordinal tasks) and percent agreement; 95% confidence intervals were obtained by case-level bootstrap resampling (1,000 itera- tions). Ordinal ratings (kinetic plausibility, image quality) were compared with the Wilcoxon signed-rank test, with a signed rank-biserial correlation 푟 (computed as (푇 + − 푇 − )∕(푇 + + Lesion Type Shape Margins Enhancement Pattern 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Kappa Paired GTSynthetic agreement exceeds inter-reader agreement GT inter-reader (Fleiss ) GTSynthetic (Cohen ) Figure C.5: Paired agreement exceeds inter-reader agree- ment. For every categorical task, GT-to-synthetic Cohen’s 휅 (blue) exceeds GT inter-reader Fleiss’ 휅 (gray). 푇 − ), negative values indicating lower synthetic ratings) as the effect size, and Kendall’s 휏 for association. Inter-reader agreement on GT and on synthetic images was summarized with Fleiss’ 휅 and bootstrap CIs. The clinical-impact outcome (no change / minor / ma- jor) was modeled with a proportional-odds ordinal logistic regression on standardized predictors (perceived synthetic appearance, diagnostic complexity, image quality, reader experience, and the three stratification axes); odds ratios with 95% CIs are reported. Method preference was tested globally with Cochran’s 푄 (the appropriate 퐾=3 extension for related binary outcomes) and post-hoc with exact McNe- mar tests; preference against chance used the binomial test. All families of tests were corrected for multiple comparisons within each analysis phase using the Benjamini–Hochberg false-discovery-rate procedure at 푞 = 0.05; significance (훼) was set at 0.05. C.4 Additional results C.4.1 Synthetic images add less variability than radiologists themselves For every categorical task, paired GT-to-synthetic Co- hen’s 휅 exceeded the GT inter-reader Fleiss’ 휅 (Figure C.5). In other words, the discrepancy a synthetic image introduces relative to its own GT is smaller than the disagreement that already exists between radiologists reading the same real images. S. Joshi et al.: Preprint submitted to ElsevierPage 21 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Figure C.6: Case demonstrating a major impact on clinical management. Although the synthetic lesion is characterized identically to the ground truth, its incorrect spatial localization fundamentally alters the diagnostic outcome. C.4.2 Real images show have better image quality and kinetic plausibility Extending the analysis in the main manuscript on this theme, Figure B.3 show absolute number of cases binned in each category. Most real images are classified as having good kinetic curves, while for majority of synthetic curves are deemed plausible. Similarly, while the distribution of cases is similar in real and synthetic cases, the mean score is lower due to higher number of non-diagnostic and poor quality cases. C.4.3 Inter-reader agreement Inter-reader agreement was moderate for lesion type (Fleiss’ 휅 = 0.43 GT, 0.36 synthetic) and decreased for finer tasks, reaching near-zero for the two ordinal scales on both image types, indicating that kinetic-plausibility and quality ratings are inherently reader-subjective regardless of image origin. C.4.4 Failure mode Blinded uncertainty propagated to clinical concern: when a reader had answered “cannot determine” in the blinded phase, their later side-by-side impact rating was markedly higher indicating worse impact (mean 1.48 vs. 0.85; 65% vs. 24% “major”; Kendall’s 휏=0.25, 푝 < 0.001). Four cases drew unanimous (4/4) major-impact verdicts. Three were explained by visible quality degradation, but one (case 31) was rated indistinguishable from GT (quality Δ ≈ 0) yet drew unanimous major-impact verdicts because of spatial errors, specifically incorrect tumor localization and extent, corroborated by three independent readers (Case 31; Fig. C.6). This dissociation between perceptual realism and spatial reliability is the study’s cautionary finding: a synthetic image can look entirely natural while placing or sizing disease incorrectly. Appendix D MISCELLANEOUS Training Data Distribution: Figure D.7 are referenced in the main manuscript Section IV to highlight lower sample sizes early and late in acquisition period. Figure D.7: Temporal distribution of training samples and ground truth contrast enhancement. The overlaid histogram (grey bars, right axis) illustrates the number of available training samples across the acquisition time 휏, highlighting the long-tailed distribution with sparse representation in early (< 100 s) and late (> 500 s) phases. The solid blue line and shaded region (left axis) denote the population mean and ±1 standard deviation of the ground truth enhancement trajectory, respectively. Sampling: Table D.7 shows an ablation with and with- out using sampling strategy to increase representation in early (< 100푠) and late (> 500푠) acquisition period. This is referenced in main manuscript Section IV. We employ a stochastic temporal latent augmentation strategy during training. For a given fraction of the batch, the method gener- ates synthetic training targets by sampling new acquisition times (휏 푛푒푤 ) located within sparsely populated temporal region. Once 휏 푛푒푤 is sampled, the algorithm identifies the two closest available real acquisition phases that bracket this time point. The corresponding images are encoded into the latent space, and a linear interpolation is performed between these two bracketing latents based on the relative temporal position of 휏 푛푒푤 . The original post-contrast training target is then replaced with this newly interpolated latent and its corresponding timestamp. This data-driven augmentation provides the model with a smooth, continuous supervisory signal across the entire temporal range, enforcing physically plausible kinetic transitions in low-density regions. References Alahari, A., Benstetter, M., 2017. Ema’s final opinion confirms restrictions on use of linear gadolinium agents in body scans. Medical Writing 26, 52. Arciszewska, Ż., Gama, S., Leśniewska, B., Malejko, J., Nalewajko- Sieliwoniuk, E., Zambrzycka-Szelewa, E., Godlewska-Żyłkiewicz, B., 2022. The translocation pathways of rare earth elements from the environment to the food chain and their impact on human health. Process safety and environmental protection 168, 205–223. Bansal, A., Borgnia, E., Chu, H.M., Li, J., Kazemi, H., Huang, F., Gold- blum, M., Geiping, J., Goldstein, T., 2023. Cold diffusion: Inverting ar- bitrary image transforms without noise. Advances in Neural Information Processing Systems 36, 41259–41282. Bau, M., Dulski, P., 1996. Anthropogenic origin of positive gadolinium anomalies in river waters. Earth and Planetary Science Letters 143, S. Joshi et al.: Preprint submitted to ElsevierPage 22 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Table D.7 Ablation Study. For a fraction of training samples, we replace the real post contrast objective with a synthetic sample where tau falls in a sparsely-populated region of the acquisition-time distribution. The post-contrast latent is linearly interpolated between the two nearest available phases (including pre at tau=0). This aims to give the model smooth, learnable supervisory signal across the entire 휏 range, especially in regions where real training data is scarce. Using this sampling strategy, the model still produces contrast and shows improvement over baseline. However, it is not able to learn physiologically accurate temporal kinetics, with marked increase in peak timing error. Best results are highlighted in bold. MSE, DTW and DTW-ROI are reported in 10 −2 scale. MethodMSE↓PSNR↑SSIM↑LPIPS↓FID - Dinov2↓FRD↓PTE↓DTW↓DTW-ROI↓ MAMA-MIA (Internal Validation) With Sampling0.88 (1.55) 18.19 (4.29) 0.71 (0.11) 0.18 (0.05)161.655.2960.79 (83.63) 0.74 (0.60)3.76 (3.33) Without Sampling 0.84 (0.55) 21.61 (2.77) 0.71 (0.11) 0.18 (0.05)154.564.98 44.57 (81.59) 0.68 (0.56) 3.76 (3.33) 245–255. Chen, Y., Yin, F., Chen, H., Wu, J., Li, C., 2026. Contrast-x: A multi- modal contrast image synthesis benchmark and universal modality flow matching. URL: https://arxiv.org/abs/2601.15884, arXiv:2601.15884. Chung, W., Kang, J., Park, G.E., Kim, S.H., Nam, Y., 2025a. Synthesizing delayed-phase contrast-enhanced breast mr images from early-phase images using an iterative deep network, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 583–592. Chung, W., Kang, J., Park, G.E., Kim, S.H., Nam, Y., 2025b. Synthesizing delayed-phase contrast-enhanced breast mr images from early-phase images using an iterative deep network, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 583–592. Cotruvo, J., 2019. The chemistry of lanthanides in biology: recent discov- eries, emerging principles, and technological applications. acs cent sci 5: 1496–1506. Dao, Q., Phung, H., Nguyen, B., Tran, A., 2023. Flow matching in latent space. arXiv preprint arXiv:2307.08698 . Dar, S.U., Yurt, M., Karacan, L., Erdem, A., Erdem, E., Cukur, T., 2019. Image synthesis in multi-contrast mri with conditional generative adver- sarial networks. IEEE transactions on medical imaging 38, 2375–2388. Darrah, T.H., Prutsman-Pfeiffer, J.J., Poreda, R.J., Ellen Campbell, M., Hauschka, P.V., Hannigan, R.E., 2009. Incorporation of excess gadolin- ium into human bone from medical contrast agents. Metallomics 1, 479– 488. Dhariwal, P., Nichol, A., 2021. Diffusion models beat gans on image synthesis, in: Advances in Neural Information Processing Systems, p. 8780–8794. Dixon, J., Newlands, C., Dodds, C., Thomas, J., Williams, L., Kunkler, I., Bing, A., Macaskill, E.J., 2016. Association between underestimation of tumour size by imaging and incomplete excision in breast-conserving surgery for breast cancer. Journal of British Surgery 103, 830–838. Fan, X., Xu, L., Zheng, B., Zeng, X., Li, W., Huang, Z., 2025. Pre-to post- contrast medical image synthesis with outline-guide accelerate diffusion model. Neural Networks , 107851. Fonnegra, R.D., Hernández, M.L., Caicedo, J.C., Díaz, G.M., 2025. Synthe- sizing late-stage contrast enhancement in breast mri: A comprehensive pipeline leveraging temporal contrast enhancement dynamics. Comput- ers in Biology and Medicine 196, 110660. Fraum, T.J., Ludwig, D.R., Bashir, M.R., Fowler, K.J., 2017. Gadolinium- based contrast agents: a comprehensive risk assessment. Journal of Magnetic Resonance Imaging 46, 338–353. Gao, Q., Li, Z., Zhang, J., Zhang, Y., Shan, H., 2023. Corediff: Contextual error-modulated generalized diffusion model for low-dose ct denoising and generalization. IEEE Transactions on Medical Imaging 43, 745– 759. Garrucho, L., et al., 2025. A large-scale multicenter breast cancer dce-mri benchmark dataset with expert segmentations. Scientific Data 12, 453. Gonzalez, V., Sandelin, K., Karlsson, A., Åberg, W., Löfgren, L., Iliescu, G., Eriksson, S., Arver, B., 2014. Preoperative mri of the breast (pomb) influences primary treatment in breast cancer: a prospective, randomized, multicenter study. World journal of surgery 38, 1685–1693. Grobner, T., Prischl, F., 2007. Gadolinium and nephrogenic systemic fibrosis. Kidney international 72, 260–264. Gweon, H.M., Cho, N., Han, W., Yi, A., Moon, H.G., Noh, D.Y., Moon, W.K., 2014. Breast mr imaging screening in women with a history of breast conservation therapy. Radiology 272, 366–373. Habijan, M., Krpić, Z., Perić, J., Galić, I., 2025. Diffusion models for mri reconstruction: A systematic review of standard, hybrid, latent and cold diffusion approaches. Electronics 15, 76. Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S., 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30. Ho, J., Jain, A., Abbeel, P., 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33, 6840–6851. Ibarra, S., del Riego, J., Catanese, A., Cuba, J., Cardona, J., Leon, N., Infante, J., Lekadir, K., Diaz, O., Osuala, R., 2025. Comparing con- ditional diffusion models for synthesizing contrast-enhanced breast mri from pre-contrast images, in: Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, Springer. p. 226–236. Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H., 2021. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 203–211. Isola, P., Zhu, J.Y., Zhou, T., Efros, A.A., 2017. Image-to-image translation with conditional adversarial networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1125–1134. Jiang, L., Dai, B., Wu, W., Loy, C.C., 2021. Focal frequency loss for image reconstruction and synthesis, in: Proceedings of the IEEE/CVF international conference on computer vision, p. 13919–13929. Joshi, S., Osuala, R., Garrucho, L., Tsirikoglou, A., del Riego, J., Gwoździewicz, K., Kushibar, K., Diaz, O., Lekadir, K., 2024. Leverag- ing epistemic uncertainty to improve tumour segmentation in breast mri: an exploratory analysis, in: Medical Imaging 2024: Image Processing, SPIE. p. 292–300. Kanda, T., Fukusato, T., Matsuda, M., Toyoda, K., Oba, H., Kotoku, J., Haruyama, T., Kitajima, K., Furui, S., 2015. Gadolinium-based contrast agent accumulates in the brain even in subjects without severe renal dysfunction: evaluation of autopsy brain specimens with inductively coupled plasma mass spectroscopy. Radiology 276, 228–232. Kingma, D.P., Welling, M., 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 . Kishore Kumar, M., Ramanarayanan, S., Sadhana, S., Sarkar, A., Gayathri, M.N., Ram, K., Sivaprakasam, M., 2024. Dce-diff: Diffusion model for synthesis of early and late dynamic contrast-enhanced mr images from non-contrast multimodal inputs. in 2024 ieee, in: CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), p. 5174–5183. Kong, J., He, Y., Xia, C., Ge, R., Li, S., 2026. Mri contrast enhancement kinetics world model, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 1288–1299. Konz, N., Chen, Y., Dong, H., Mazurowski, M.A., 2024. Anatomically- controllable medical image generation with segmentation-guided diffu- sion models, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 88–98. S. Joshi et al.: Preprint submitted to ElsevierPage 23 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Konz, N., Osuala, R., Verma, P., Chen, Y., Gu, H., Dong, H., Chen, Y., Marshall, A., Garrucho, L., Kushibar, K., et al., 2026. Fréchet ra- diomic distance (frd): A versatile metric for comparing medical imaging datasets. Medical Image Analysis , 103943. Krasznai, Z., Morisawa, M., Krasznai, Z.T., Morisawa, S., Inaba, K., Bazsáné, Z.K., Rubovszky, B., Bodnár, B., Borsos, A., Márián, T., 2003. Gadolinium, a mechano-sensitive channel blocker, inhibits osmosis- initiated motility of sea-and freshwater fish sperm, but does not affect human or ascidian sperm motility. Cell motility and the cytoskeleton 55, 232–243. Kuhl, C.K., Mielcareck, P., Klaschik, S., Leutner, C., Wardelmann, E., Gieseke, J., Schild, H.H., 1999. Dynamic breast mr imaging: are signal intensity time course data useful for differential diagnosis of enhancing lesions? Radiology 211, 101–110. Kuhl, C.K., Strobel, K., Bieling, H., Wardelmann, E., Kuhn, W., Maass, N., Schrading, S., 2017. Impact of preoperative breast mr imaging and mr-guided surgery on diagnosis and surgical outcome of women with invasive breast cancer with and without dcis component. Radiology 284, 645–655. Laczovics, A., Csige, I., Szabó, S., Tóth, A., Kálmán, F.K., Tóth, I., Fülöp, Z., Berényi, E., Braun, M., 2023. Relationship between gadolinium- based mri contrast agent consumption and anthropogenic gadolinium in the influent of a wastewater treatment plant. Science of the Total Environment 877, 162844. Lang, D.M., Osuala, R., Spieker, V., Lekadir, K., Braren, R., Schnabel, J.A., 2025. Temporal neural cellular automata: Application to modeling of contrast enhancement in breast mri, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 604–614. Li, W., et al., 2022. I-SPY 2 breast dynamic contrast enhanced MRI trial (version 1) [data set]. The Cancer Imaging Archive URL: https: //doi.org/10.7937/TCIA.D8Z0-9T85, doi:10.7937/TCIA.D8Z0-9T85. Lindner, U., Lingott, J., Richter, S., Jakubowski, N., Panne, U., 2013. Speciation of gadolinium in surface water samples and plants by hy- drophilic interaction chromatography hyphenated with inductively cou- pled plasma mass spectrometry. Analytical and bioanalytical chemistry 405, 1865–1873. Lingott, J., Lindner, U., Telgmann, L., Esteban-Fernández, D., Jakubowski, N., Panne, U., 2016. Gadolinium-uptake by aquatic and terrestrial organisms-distribution determined by laser ablation inductively coupled plasma mass spectrometry. Environmental Science: Processes & Im- pacts 18, 200–207. Lipman, Y., Chen, R.T., Ben-Hamu, H., Nickel, M., Le, M., 2022. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747 . Liu, X., Gong, C., Liu, Q., 2022. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003 . M, K.K., Ramanarayanan, S., S, S., Sarkar, A., Gayathri, M.N., Ram, K., Sivaprakasam, M., 2024. Dce-diff: Diffusion model for synthesis of early and late dynamic contrast-enhanced mr images from non-contrast multimodal inputs, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, p. 5174–5183. Monticciolo, D.L., Newell, M.S., Moy, L., Niell, B., Monsees, B., Sickles, E.A., 2018. Breast cancer screening in women at higher-than-average risk: recommendations from the acr. Journal of the American College of Radiology 15, 408–414. Müller-Franzes, G., Huck, L., Bode, M., Nebelung, S., Kuhl, C., Truhn, D., Lemainque, T., 2024. Diffusion probabilistic versus generative adversarial models to reduce contrast agent dose in breast mri. European radiology experimental 8, 53. Müller-Franzes, G., Huck, L., Tayebi Arasteh, S., Khader, F., Han, T., Schulz, V., Dethlefsen, E., Kather, J.N., Nebelung, S., Nolte, T., et al., 2023. Using machine learning to reduce the need for contrast agents in breast mri through synthetic images. Radiology 307, e222211. Murata, N., Gonzalez-Cuyar, L.F., Murata, K., Fligner, C., Dills, R., Hippe, D., Maravilla, K.R., 2016. Macrocyclic and other non–group 1 gadolin- ium contrast agents deposit low levels of gadolinium in brain and bone tissue: preliminary results from 9 patients with normal renal function. Investigative radiology 51, 447–453. Naval Marimont, S., Siomos, V., Baugh, M., Tzelepis, C., Kainz, B., Tarroni, G., 2024. Ensembled cold-diffusion restorations for unsuper- vised anomaly detection, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 243–253. Newitt, D., Hylton, N., 2016. Single site breast DCE-MRI data and segmen- tations from patients undergoing neoadjuvant chemotherapy (version 3) [data set]. The Cancer Imaging Archive URL: https://doi.org/10.793 7/K9/TCIA.2016.QHsyhJKy, doi:10.7937/K9/TCIA.2016.QHsyhJKy. Newitt, D., et al., 2016. Multicenter breast DCE-MRI data and segmen- tations from patients in the I-SPY 1/ACRIN 6657 trials. The Cancer Imaging Archive URL: https://doi.org/10.7937/K9/TCIA.2016.HdHpgJL K, doi:10.7937/K9/TCIA.2016.HdHpgJLK. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al., 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 . Osuala, R., Joshi, S., Tsirikoglou, A., Garrucho, L., Pinaya, W.H., Lang, D.M., Schnabel, J.A., Diaz, O., Lekadir, K., 2025. Simulating dynamic tumor contrast enhancement in breast mri using conditional generative adversarial networks. Journal of Medical Imaging 12, S22014–S22014. Osuala, R., Lang, D.M., Verma, P., Joshi, S., Tsirikoglou, A., Skorupko, G., Kushibar, K., Garrucho, L., Pinaya, W.H., Diaz, O., et al., 2024. Towards learning contrast kinetics with multi-condition latent diffusion models, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 713–723. Partridge, S.C., Stone, K.M., Strigel, R.M., DeMartini, W.B., Peacock, S., Lehman, C.D., 2014. Breast dce-mri: influence of postcontrast timing on automated lesion kinetics assessments and discrimination of benign and malignant lesions. Academic radiology 21, 1195–1203. Parvaiz, M.A., Yang, P., Razia, E., Mascarenhas, M., Deacon, C., Matey, P., Isgar, B., Sircar, T., 2016. Breast mri in invasive lobular carcinoma: a useful investigation in surgical planning? The breast journal 22, 143– 150. Pesapane, F., Sorce, A., Battaglia, O., Mallardi, C., Nicosia, L., Mari- ano, L., Rotili, A., Dominelli, V., Penco, S., Priolo, F., et al., 2025. Contrast agents in breast mri: State of the art and future perspectives. Biomedicines 13, 829. Peterson, M.S., Gegios, A.R., Elezaby, M.A., Salkowski, L.R., Woods, R.W., Narayan, A.K., Strigel, R.M., Roy, M., Fowler, A.M., 2023. Breast imaging and intervention during pregnancy and lactation. Radiographics 43, e230014. Pinaya, W.H., Graham, M.S., Kerfoot, E., Tudosiu, P.D., Dafflon, J., Fer- nandez, V., Sanchez, P., Wolleb, J., Da Costa, P.F., Patel, A., et al., 2023. Generative ai for medical imaging: extending the monai framework. arXiv preprint arXiv:2307.15208 . Pinaya, W.H., Tudosiu, P.D., Dafflon, J., Da Costa, P.F., Fernandez, V., Nachev, P., Ourselin, S., Cardoso, M.J., 2022. Brain imaging generation with latent diffusion models, in: MICCAI workshop on deep generative models, Springer. p. 117–126. Qi, M., Shi, R., Cai, Y., He, L., Wang, W., Ma, L., 2025. A boundary- aware cold-diffusion model for electron microscopy segmentation, in: International Conference on Medical Image Computing and Computer- Assisted Intervention, Springer. p. 13–23. Rogosnitzky, M., Branch, S., 2016. Gadolinium-based contrast agent toxicity: a review of known and proposed mechanisms. Biometals 29, 365–376. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B., 2022. High- resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recogni- tion, p. 10684–10695. Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation. URL: https://arxiv.org/abs/1505 .04597, arXiv:1505.04597. Saha, A., et al., 2021. Dynamic contrast-enhanced magnetic resonance images of breast cancer patients with tumor locations [data set]. The Cancer Imaging Archive URL: https://doi.org/10.7937/TCIA.e3sv-re9 3, doi:10.7937/TCIA.e3sv-re93. S. Joshi et al.: Preprint submitted to ElsevierPage 24 of 25 Dense Temporal Contrast Synthesis via Conditioned Latent Transport Saslow, D., Boetes, C., Burke, W., Harms, S., Leach, M.O., Lehman, C.D., Morris, E., Pisano, E., Schnall, M., Sener, S., et al., 2007. American cancer society guidelines for breast screening with mri as an adjunct to mammography. CA: a cancer journal for clinicians 57, 75–89. Schreiter, H., Eberle, J., Kapsner, L.A., Hadler, D., Ohlmeyer, S., Erber, R., Emons, J., Laun, F.B., Uder, M., Wenkel, E., et al., 2024. Virtual dynamic contrast enhanced breast mri using 2d u-net architectures, in: Deep Breast Workshop on AI and Imaging for Diagnostic and Treatment Challenges in Breast Care, Springer. p. 85–95. Selvi, V., Nori, J., Meattini, I., Francolini, G., Morelli, N., Di Benedetto, D., Bicchierai, G., Di Naro, F., Gill, M.K., Orzalesi, L., et al., 2018. Role of magnetic resonance imaging in the preoperative staging and work-up of patients affected by invasive lobular carcinoma or invasive ductolobular carcinoma. BioMed research international 2018, 1569060. Shen, G., Li, M., Farris, C.W., Anderson, S., Zhang, X., 2024. Learning to reconstruct accelerated mri through k-space cold diffusion without noise. Scientific Reports 14, 21877. Tang, R., Zhang, X., Guo, P., Qiang, X., 2025. Residual pre-training assisted cold diffusion for denoising low-dose computed tomography images, in: Eighth International Conference on Artificial Intelligence and Pattern Recognition (AIPR 2025), SPIE. p. 1127–1135. Tur, Y., Stojkovic, M., Bagci, U., 2026. Wfm: 3d wavelet flow matching for ultrafast multi-modal mri synthesis, in: Medical Imaging with Deep Learning. Uematsu, T., Yuen, S., Kasami, M., Uchida, Y., 2008. Comparison of magnetic resonance imaging, multidetector row computed tomography, ultrasonography, and mammography for tumor extension of breast can- cer. Breast cancer research and treatment 112, 461–474. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I., 2017. Attention is all you need. Advances in neural information processing systems 30. Wang, J., Reynaud, H., Erick, F.X., Kainz, B., 2025. Ctflow: Video- inspired latent flow matching for 3d ct synthesis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 6750– 6758. Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 600–612. Yan, P., Li, M., Zhang, J., Li, G., Jiang, Y., Luo, H., 2024a. Cold segdiffusion: A novel diffusion model for medical image segmentation. Knowledge-Based Systems 301, 112350. Yan, P., Li, M., Zhang, J., Li, G., Jiang, Y., Luo, H., 2024b. Cold segdiffusion: A novel diffusion model for medical image segmentation. Knowledge-Based Systems 301, 112350. Zaman, F.A., Jacob, M., Chang, A., Liu, K., Sonka, M., Wu, X., 2024. Surf-cdm: Score-based surface cold-diffusion model for medical image segmentation, in: 2024 IEEE International Symposium on Biomedical Imaging (ISBI), IEEE. p. 1–5. Zhang, J., Liu, G., Chen, J., Cheng, Y., 2025. Multi-scale adaptive residual cold diffusion model for low-dose ct denoising. Expert Systems with Applications 294, 128817. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O., 2018. The unreasonable effectiveness of deep features as a perceptual metric, in: Proceedings of the IEEE conference on computer vision and pattern recognition, p. 586–595. Zhang, T., Han, L., D’Angelo, A., Wang, X., Gao, Y., Lu, C., Teuwen, J., Beets-Tan, R., Tan, T., Mann, R., 2023. Synthesis of contrast-enhanced breast mri using t1-and multi-b-value dwi-based hierarchical fusion net- work with attention mechanism, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 79–88. S. Joshi et al.: Preprint submitted to ElsevierPage 25 of 25