Paper deep dive
Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis
Xuan Yang, Xiaohan Yuan, Hao Li, Lingyu Chen, Yanan Liu, Qingya Li, Lei Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/21/2026, 3:31:23 AM
Summary
The paper proposes a phenotype-guided latent flow framework for synthesizing full-cycle biventricular motion from a single end-diastolic (ED) mesh. The method integrates motion-informed functional parcellation with a conditional rectified-flow model to capture spatially heterogeneous and phenotype-dependent cardiac dynamics. Experiments on ACDC, M&Ms, and M&Ms-2 datasets demonstrate superior geometric accuracy and functional fidelity compared to existing methods.
Entities (10)
Relation Signals (9)
Phenotype-Guided Latent Flow → evaluatedon → M&Ms
confidence 99% · Experiments on ACDC, M&Ms, and M&Ms-2 demonstrate consistent improvements
Phenotype-Guided Latent Flow → evaluatedon → M&Ms-2
confidence 99% · Experiments on ACDC, M&Ms, and M&Ms-2 demonstrate consistent improvements
Phenotype-Guided Latent Flow → evaluatedon → ACDC
confidence 99% · Experiments on ACDC, M&Ms, and M&Ms-2 demonstrate consistent improvements
Phenotype-Guided Latent Flow → takesinput → End-Diastolic Mesh
confidence 97% · investigate full-cycle biventricular motion synthesis from a single ED mesh
Phenotype-Guided Latent Flow → uses → Rectified Flow
confidence 95% · A phenotype-conditioned rectified-flow model subsequently maps the ED anatomy to full-cycle motion latents
Phenotype-Guided Latent Flow → uses → Functional Parcellation
confidence 92% · integrates motion-informed functional parcellation with conditional latent flow
Phenotype-Guided Latent Flow → achievesmetric → ASSD
confidence 90% · our method achieves biventricular ASSD... of 1.49±0.34 mm
Phenotype-Guided Latent Flow → achievesmetric → HD95
confidence 90% · our method achieves biventricular... HD95... of 3.77±1.06 mm
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventricular meshes are not routinely available, whereas end-diastolic (ED) anatomy can often be obtained reliably. We therefore investigate full-cycle biventricular motion synthesis from a single ED mesh. This task is challenging because cardiac deformation is spatially heterogeneous and phenotype dependent, while conventional global generative models often obscure localized motion patterns. In this study, we propose a region-specific and phenotype-adaptive framework that integrates motion-informed functional parcellation with conditional latent flow. A functional partition learned from reconstructed motion organizes the ventricular surface into regions with coherent dynamics and enables topology-aware regional feature exchange. A phenotype-conditioned rectified-flow model subsequently maps the ED anatomy to full-cycle motion latents through fine-grained conditioning and prototype-routed motion adapters. An optional control branch further incorporates available motion descriptors for controllable synthesis. Experiments on ACDC, M\&Ms, and M\&Ms-2 demonstrate consistent improvements in geometric accuracy and functional fidelity. Under ED-only synthesis, our method achieves biventricular ASSD, HD95, and vRMSE of \(1.49\pm0.34\)~mm, \(3.77\pm1.06\)~mm, and \(3.31\pm1.03\)~mm, respectively, outperforming all competing methods. Complementary functional and robustness evaluations further demonstrate that the synthesized sequences preserve physiologically plausible ventricular dynamics and generalize across cohorts and disease phenotypes. The code will be released publicly upon acceptance of the manuscript for publication.
Tags
Links
- Source: https://arxiv.org/abs/2608.19738v1
- Canonical: https://arxiv.org/abs/2608.19738v1
Trouble viewing inline? Open PDF directly →
Full Text
73,967 characters extracted from source content.
Expand or collapse full text
Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion SynthesisJournal: Medical Image Analysis Xuan Yang Address: Department of Biomedical Engineering, National University of Singapore, Singapore Xiaohan Yuan Address: Department of Biomedical Engineering, National University of Singapore, Singapore Address: School of Automation, Southeast University, Nanjing, China Hao Li Address: Department of Biomedical Engineering, National University of Singapore, Singapore Lingyu Chen Address: School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing, China Yanan Liu Address: Department of Biomedical Engineering, National University of Singapore, Singapore Qingya Li Address: Department of Biomedical Engineering, National University of Singapore, Singapore Lei Li* URL: lei.li@nus.edu.sg Address: Department of Biomedical Engineering, National University of Singapore, Singapore Abstract Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventricular meshes are not routinely available, whereas end-diastolic (ED) anatomy can often be obtained reliably. We therefore investigate full-cycle biventricular motion synthesis from a single ED mesh. This task is challenging because cardiac deformation is spatially heterogeneous and phenotype dependent, while conventional global generative models often obscure localized motion patterns. In this study, we propose a region-specific and phenotype-adaptive framework that integrates motion-informed functional parcellation with conditional latent flow. A functional partition learned from reconstructed motion organizes the ventricular surface into regions with coherent dynamics and enables topology-aware regional feature exchange. A phenotype-conditioned rectified-flow model subsequently maps the ED anatomy to full-cycle motion latents through fine-grained conditioning and prototype-routed motion adapters. An optional control branch further incorporates available motion descriptors for controllable synthesis. Experiments on ACDC, M&Ms, and M&Ms-2 demonstrate consistent improvements in geometric accuracy and functional fidelity. Under ED-only synthesis, our method achieves biventricular ASSD, HD95, and vRMSE of 1.49±0.341.49± 0.34 m, 3.77±1.063.77± 1.06 m, and 3.31±1.033.31± 1.03 m, respectively, outperforming all competing methods. Complementary functional and robustness evaluations further demonstrate that the synthesized sequences preserve physiologically plausible ventricular dynamics and generalize across cohorts and disease phenotypes. The code will be released publicly upon acceptance of the manuscript for publication. Keywords: Myocardial Infarction , Cine MRI , 3D Infarct Reconstruction , Contrast Free , Electrophysiological Simulation , Cardiac Digital Twins 1 Introduction Cardiovascular disease is a major cause of death worldwide (41). Cardiac abnormalities are reflected not only in changes in anatomy, but also in altered functional and motion patterns throughout the cardiac cycle (1; 36). Cardiac magnetic resonance imaging (MRI) can characterize myocardial shape and motion during contraction and relaxation with high spatial and temporal resolution (46). However, analyzing such spatiotemporal information only in the voxel or label space makes it difficult to directly represent continuous anatomical surfaces, topological structures, and geometrical changes that are aligned across time. In contrast, time-resolved 3D mesh sequences provide a more natural representation of cardiac surface geometry and temporal correspondence, and offer a useful basis for patient-specific visualization and physics-based computational modeling (31; 19; 29). Despite these advantages, complete and high-quality 4D (3D+t) cardiac mesh sequences remain difficult to obtain in practice (53; 3), and in real clinical settings only a few key states can usually be acquired reliably. Previous studies (24; 21; 28; 10) have shown that even under very limited observations, such as a single frame, sparse slices, or a reference phase, it is still possible to recover patient-specific cardiac dynamics. Among these states, end-diastole (ED) is commonly used as a reference because it provides a relatively stable and complete anatomical configuration and serves as an anchor for subsequent dynamic modeling (54; 31; 45). These observations motivate using the ED anatomy as an accessible patient-specific condition for synthesizing full-cycle cardiac dynamics. Generative modeling has become increasingly valuable in cardiac analysis, not only for completing missing dynamics, but also for expanding the available sample space, constructing virtual cohorts, and supporting in silico clinical trials (32; 14; 44; 13). Early efforts have largely focused on static cardiac anatomy modeling (21; 6; 14; 5), including patient-specific mesh reconstruction from cardiac images and conditional generation of anatomical shapes or virtual populations. These methods provide important tools for characterizing inter-subject morphological variability. However, static anatomy alone cannot describe functional changes over the cardiac cycle, nor can it capture the coupled evolution of shape and motion that underlies cardiac function. Recent studies have increasingly shifted toward dynamic cardiac modeling (38; 44; 13; 37), enabling tasks such as motion sequence completion, dynamic shape generation, and joint shape-motion distribution modeling. Despite these advances, many existing dynamic generative models mainly emphasize population-level plausibility and temporal consistency, while explicit modeling of regional motion variation remains limited. This is particularly important because cardiac motion is spatially heterogeneous, with regional wall motion differences reflecting local myocardial function and pathology (1; 48). Without explicitly accounting for such regional characteristics, dynamic generation models may struggle to represent localized abnormalities and fine-grained functional variation. A second challenge is modeling phenotype-dependent motion variability. Many existing dynamic generation methods are trained primarily on healthy subjects or large normative cohorts (37; 13), which may limit their ability to represent pathological motion patterns. Although several studies have introduced conditional control for cardiac sequence generation (38; 44; 13), the conditioning signals are often restricted to general clinical covariates or diagnostic labels, and thus may not fully capture fine-grained, patient-specific functional differences. Clinically meaningful cardiac motion generation therefore requires more informative conditioning mechanisms, particularly phenotype-aware representations that can encode individual disease characteristics and guide the generation of region-specific dynamic patterns. To address the above challenges, we propose a region-specific and phenotype-adaptive latent generative framework for full-cycle biventricular motion synthesis on unified-topology meshes. The framework uses a patient-specific ED mesh as the anatomical condition and synthesizes full-cycle cardiac dynamics under phenotype guidance, aiming to capture anatomy-dependent and disease-associated motion variation. Specifically, our method first learns a motion-informed functional partition from real cardiac sequences, providing explicit regional priors for modeling local dynamics. The priors are then used to organize a motion variational autoencoder (VAE) for reconstructing full-cycle cardiac sequences on the learned regional structure. Finally, a rectified-flow generator is trained in the learned VAE latent space to enhance phenotype-conditioned motion synthesis from ED anatomy, with optional motion descriptors used when available for controllable generation. The main contributions of this work are as follows: i. We establish a phenotype-adaptive latent motion generation framework for synthesizing full-cycle biventricular motion from a single ED mesh. i. We introduce a motion-informed regional prior and integrate it into a region-structured motion VAE, enabling compact representation of full-cycle cardiac dynamics while preserving regional motion heterogeneity. i. We develop a rectified-flow latent generator with prototype-routed motion adapters, improving phenotype-conditioned motion generation while supporting optional functional control. iv. Extensive experiments on multiple public cardiac MRI datasets demonstrate the effectiveness of the proposed method in full-cycle motion synthesis, functional consistency, and cross-disease generalization. This paper significantly extends a preliminary conference version of the work presented in 50. Compared with the previous framework, this study redesigns the overall pipeline into a region-specific latent generative framework for full-cycle motion synthesis, strengthening both the structural prior and the latent generation mechanism. Specifically, we replace the fixed binary regional adjacency with a hop-based region-topology prior, improving the modeling of gradual motion propagation across neighboring regions. We also replace the previous VAE-prior-based motion generation with a rectified-flow latent generator in the learned latent space, where fine-grained ED anatomical and phenotype conditioning support more detailed morphology-aware motion synthesis. In addition, an optional functional-control branch enables controllable modulation of the generated dynamics when motion descriptors are available. We further conduct more systematic experiments, including comparisons with competing methods and ablation studies, to evaluate the contribution of these modeling components. 2 Related Work 2.1 Cardiac Shape and Motion Generative Modeling Cardiac statistical shape analysis has been used to characterize anatomical variation and its relationship to disease, supporting diagnosis, risk assessment, and treatment planning (1; 8; 2). With the development of deep learning, research has gradually shifted toward more expressive generative models. For instance, 14 proposed a conditional VAE to synthesize virtual anatomical populations of the left ventricle under covariate constraints. 4; 5 used point clouds and VAE to model three-dimensional cardiac shape variation, and applied it to myocardial infarction-related shape analysis and virtual heart synthesis. 20 further extended conditional anatomical generation to more complex topologies, enabling controllable synthesis of congenital heart disease anatomy through type-shape disentanglement. Overall, static cardiac mesh generation has evolved from patient-specific anatomical reconstruction to controllable anatomical distribution learning. However, these methods remain limited to static anatomy and cannot directly describe functional changes throughout the cardiac cycle. To solve this, 38 proposed a conditional spatiotemporal generative model to jointly model 4D cardiac anatomy and its relationship with non-imaging clinical factors. 37 directly learned the shape-motion distribution of 3D+t biventricular meshes and explored cardiac morphology and motion patterns in large-scale population data. 13 further generated dynamic virtual heart populations through spatiotemporal disentanglement. Nevertheless, existing dynamic models mainly focus on globally coherent motion patterns and are often developed on healthy populations or under global condition signals, leaving disease-related dynamic variation largely underexplored. 2.2 Patient-Specific Dynamic Modeling from Sparse Observations Patient-specific cardiac dynamic modeling has mainly relied on registration, motion tracking, and spatiotemporal deformation estimation from cine imaging (39; 52; 27; 40; 26; 49). With the development of deep learning, these approaches have gradually shifted from traditional optimization-based registration to data-driven dynamic recovery. One group of methods (40; 54; 27), learns temporal deformation representations in the image domain. For example, 40 achieved left ventricular myocardial motion tracking through a biomechanically constrained latent deformation manifold. DragNet (54) further recovered temporally consistent full cardiac dynamics from a single-frame cine MRI. However, the outputs of these methods are still essentially image-domain spatiotemporal deformations or generated image sequences, rather than mesh sequences with unified topological constraints. To obtain more explicit geometric representations, recent studies moved from image-domain displacement fields to mesh-based modeling. DeepMesh (31) represented patient-specific cardiac anatomy by deforming a template mesh to the ED geometry and subsequently estimated three-dimensional mesh motion. 22 further recovered patient-specific 4D cardiac meshes from 2D echocardiographic videos under a weakly supervised setting. Compared with general image-domain motion fields, mesh sequences can more naturally preserve surface geometry, fixed topology, and vertex correspondence across time, making them more suitable for subsequent regional prior modeling and local dynamic analysis. Overall, existing patient-specific dynamic modeling studies have shown that recovering full cardiac dynamics from key states, single-frame images, or limited observations is a reasonable and important problem setting. However, their main goal is still tracking, reconstruction, or globally plausible motion recovery, rather than unified generative modeling of disease-related dynamic processes. 2.3 Latent and Flow-Based Motion Generation In complex temporal motion modeling, latent-variable generative models are among the most representative technical routes (18; 11; 42). Their core idea is to compress high-dimensional temporal motion into a low-dimensional latent space and then model complex motion distributions within that space. Action2Motion (15) used a temporal VAE to model human motion under action conditions. ACTOR (35) further employed a Transformer to learn action-aware latent motion representations. Furthermore, AnimateAnyMesh (47) extended conditional generation in compressed latent spaces to three-dimensional mesh objects with arbitrary topology, enabling unified animation generation for multiple types of mesh objects. This suggests that latent-space modeling is not only suitable for general temporal signals, but can also support motion generation for geometrically more complex and topologically more flexible 3D objects. Diffusion models (16; 12; 43) achieve high-fidelity generation through iterative denoising and have shown strong performance in modeling complex distributions, but they usually rely on multi-step sampling and large training scales. To improve training and sampling efficiency, recent studies have turned to continuous flow-based modeling. Flow Matching (23) characterizes transport between distributions by learning continuous-time velocity fields along fixed probability paths. Rectified Flow (25) further favors learning more direct transport trajectories. Motion Flow Matching (17) has applied this idea to human motion generation, completion, and interpolation. However, these methods have mostly been developed for general human motion or generic 3D geometric object scenarios, and they still lack specialized designs for jointly encoding anatomical states, disease-related conditions, and regional priors. Therefore, combining latent-space modeling with flow-based generation, and adapting it to disease-related and regionally heterogeneous cardiac dynamics, remains an important direction for further study. 3 Methods 3.1 Problem Formulation Given only the ED biventricular mesh of a subject, our goal is to recover the full dynamic biventricular mesh sequence. Let xt∈ℝN×3x_t ^N× 3 denote the biventricular mesh at cardiac frame t, where N is the number of vertices and t=0,…,T−1t=0,…,T-1. The full cardiac mesh sequence is defined as =xtt=0T−1∈ℝT×N×3x=\x_t\_t=0^T-1 ^T× N× 3, where x0x_0 corresponds to the ED mesh. All biventricular meshes are represented using a unified template with shared topology and vertex-wise correspondence across subjects and cardiac frames. This allows cardiac motion to be formulated as an ED-referenced displacement trajectory, reducing subject-specific spatial variation and enabling the model to focus on learning cardiac deformation rather than absolute vertex coordinates. Specifically, the displacement at frame t is defined as ut=xt−x0u_t=x_t-x_0. The complete displacement trajectory is therefore given by =utt=0T−1∈ℝT×N×3u=\u_t\_t=0^T-1 ^T× N× 3. Under this formulation, cardiac sequence recovery is defined as learning a mapping from the patient-specific ED mesh x0x_0 to the full displacement trajectory u, as shown in Fig. 1. The dynamic mesh sequence is then recovered by adding the generated displacement field to the ED mesh: x^t=x0+u^t,t=0,…,T−1. x_t=x_0+ u_t, t=0,…,T-1. (1) To predict the displacement trajectory u, we learn a motion generator conditioned on the patient-specific ED mesh x0x_0 and its ED-derived phenotype descriptors cEDglobalc_ED^global and cEDfinec_ED^fine. Following the cardiac MRI-derived phenotype definitions in 2, cEDglobalc_ED^global summarizes global chamber and myocardial morphology, whereas cEDfinec_ED^fine further captures regional ED myocardial wall-thickness information (see Sec. 4.1 for details). When ventricular functional measurements are available, the motion descriptor cmotionc^motion can be additionally incorporated as optional inputs. Figure 1: Illustration of the task and its formulation. ED: end-diastolic. Figure 2: Illustration of the proposed regional motion-informed and phenotype-guided biventricular motion synthesis framework. (a) Motion-driven functional regionalization. (b) Region-topology aware motion variational autoencoder (VAE). (c) Phenotype-guided latent flow for anatomy-to-motion generation. DyMeshVAE: dynamic mesh VAE; PE: positional encoding. 3.2 Motion-Informed Cardiac Functional Parcellation To model spatially heterogeneous cardiac deformation, we learn a data-driven functional parcellation directly from training motion sequences, avoiding predefined anatomical divisions that may not reflect regional motion similarity. Defined on the unified biventricular template, the resulting partition provides a fixed, subject-independent regional prior. Specifically, given full cardiac mesh sequences x, we first train a dynamic mesh reconstruction network, DyMeshVAE (47), to extract vertex-wise motion features f(i)f(i), as shown in Fig. 2 (a). To obtain a subject-independent descriptor for each template vertex, we average the vertex features across the training set and thus obtain f¯(i) f(i). We then apply K-means to the template-level descriptors and partition the template vertices into NRN_R motion regions: r(i)=argminρ∈1,…,NR‖f¯(i)−μρ‖22,r(i)∈1,…,NR,r(i)= _ρ∈\1,…,N_R\ \| f(i)- _ρ \|_2^2, r(i)∈\1,…,N_R\, (2) where i is the index of mesh vertex and μρ _ρ is the centroid of the ρ-th motion region. The resulting regional prior comprises the fixed vertex-to-region assignments r(i)i=1N\r(i)\_i=1^N and the region adjacency matrix AdjAdj. Two regions are considered adjacent if at least one mesh edge connects vertices assigned to the two regions. Cardiac motion involves coordinated deformation across multiple interconnected regions, which may not be fully captured by one-hop connectivity alone. We therefore introduce multi-hop regional attention bias (MRAB) to extend the binary adjacency mask used in 50 by encoding the multi-hop structure of the learned region graph in the attention logits. Let Dreg(ρ,ρ′)D_reg(ρ,ρ ) denote the shortest-path distance between regions ρ and ρ′ρ . Rather than treating all admissible region pairs equally, MRAB converts this distance into an additive attention bias: Breg(ρ,ρ′)=0,Dreg(ρ,ρ′)=0,(Dreg(ρ,ρ′)−1)logη,1≤Dreg(ρ,ρ′)≤Lhop,−∞,Dreg(ρ,ρ′)>Lhop,B_reg(ρ,ρ )= cases0,&D_reg(ρ,ρ )=0,\\ (D_reg(ρ,ρ )-1 ) η,&1≤ D_reg(ρ,ρ )≤ L_hop,\\ -∞,&D_reg(ρ,ρ )>L_hop, cases (3) where 0<η<10<η<1 controls the attenuation across successive hops and LhopL_hop specifies the maximum interaction range. Accordingly, self- and one-hop interactions remain unbiased, whereas every additional hop introduces a progressively stronger negative bias. Region pairs outside the prescribed neighborhood are excluded. MRAB thus provides a distance-aware topological prior for controlled information exchange beyond immediately adjacent regions. The vertex-to-region assignments and BregB_reg are subsequently used for regional feature aggregation and topology-aware attention, respectively, in the motion VAE described in Sec. 3.3. 3.3 Region-Structured Motion VAE Figure 3: (a) Region-specific injection module. (b) Phenotype-conditioned rectified-flow block with optional functional motion control. FPS: farthest point sampling; MLP: multilayer perceptron. To learn a compact latent representation for full-cycle cardiac motion, we design a region-structured motion VAE, as illustrated in Fig. 2 (b). The VAE adopts a two-stream design in which the ED stream encodes the static anatomical state x0x_0, while the motion stream encodes the ED-relative trajectory u. This separation allows the motion stream to model temporal deformation relative to a stable anatomical reference, reducing interference between static anatomical coordinates and temporal displacement patterns. Inspired by DyMeshVAE (47), we apply distinct positional encoding (PE) to x0x_0 and u, yielding x0PEx_0^PE and PEu^PE, respectively. Region-balanced farthest point sampling (FPS) then selects NaN_a anchor vertices, and the same sampling indices are used to gather features from both streams, forming paired ED and motion anchor tokens a0,au∈ℝNa×da_0,a_u ^N_a× d, where d denotes the token feature dimension. Each token pair inherits the region label of its sampled vertex from the motion-informed partition defined in Sec. 3.2. To preserve the regional motion heterogeneity captured by this partition, each encoder block incorporates a region-specific injection module, as shown in Fig. 3 (a). The module mean-pools the anchor tokens within each region to obtain ED and motion region tokens R¯0,R¯∈ℝNR×d R_0, R_u ^N_R× d. At the regional level, the two streams retain distinct roles. The ED region tokens determine how information is exchanged between regions, whereas the motion region tokens provide the dynamic information propagated through these interactions. Accordingly, BregB_reg is added to the attention logits computed from R¯0 R_0, and the resulting attention map is applied to both streams: R^0 R_0 =Softmax(R¯0R¯0⊤d+Breg)R¯0+R¯0, =Softmax ( R_0 R_0 d+B_reg ) R_0+ R_0, (4) R R_u =Softmax(R¯0R¯0⊤d+Breg)R¯+R¯. =Softmax ( R_0 R_0 d+B_reg ) R_u+ R_u. The updated region tokens are then mapped back to their corresponding anchors and injected into both streams through FiLM-based residual modulation (34). Stacking KmK_m such encoder blocks progressively incorporates topology-constrained regional context into the anchor features while preserving fine-grained spatial representations. After the final encoder block, the ED anchor tokens are projected into the deterministic anatomical latent x¯0∈ℝNa×dz x_0 ^N_a× d_z, while the motion anchor tokens are projected into μq _q and σq _q, which parameterize the motion posterior. qϕ(z∣x0,)=(μq,diag(σq2)),q_φ(z x_0,u)=N ( _q,diag( _q^2) ), (5) where z∈ℝNa×dzz ^N_a× d_z denotes a motion latent sampled from the posterior and dzd_z is the per-anchor latent feature dimension. The decoder takes the deterministic anatomical latent x¯0 x_0 and motion latent z as inputs, projects them back to the ED and motion anchor representations, and refines both streams through SyncAttention. The dense ED features x0PEx_0^PE then query the decoded ED anchor representations through cross-attention and aggregate the corresponding motion anchor representations to reconstruct ^∈ℝT×N×3 u ^T× N× 3. We define the reconstruction term as the mean squared error between the reconstructed and reference dynamic mesh sequences: ℒrec=1T1N∑t=0T−1∑i=1N‖x^t(i)−xt(i)‖22.L_rec= 1T 1N _t=0^T-1 _i=1^N \| x_t(i)-x_t(i) \|_2^2. (6) The motion VAE is optimized with dynamic mesh reconstruction and KL regularization: ℒVAE=ℒrec+λKLKL[qϕ(z∣x0,)∥(0,I)],L_VAE=L_rec+ _KLD_KL [q_φ(z x_0,u)\|N(0,I) ], (7) where λKL _KL is a balancing parameter. After training, the motion VAE is frozen, providing the latent space and decoder used by the phenotype-guided motion generator described in Sec. 3.4. 3.4 Phenotype-Guided Latent Motion Generation Latent Rectified-Flow Training As described in Sec. 3.3, the frozen motion VAE represents each training sequence using a deterministic anatomical latent x¯0 x_0 and a posterior motion latent z, from which its decoder reconstructs the displacement trajectory u. Its Gaussian prior, however, is independent of subject-specific ED anatomy and phenotype. We therefore train a conditional rectified-flow (RF) model whose velocity field is conditioned on x¯0 x_0, cEDglobalc_ED^global, and cEDfinec_ED^fine, thereby transporting Gaussian noise into subject-specific motion latents, as illustrated in Fig. 2 (c). Following the standard RF formulation (25; 23), for each motion latent z sampled from the frozen VAE posterior, we draw ϵ∼(0,I)ε (0,I) and τ∼(0,1)τ (0,1), and construct zτ=(1−τ)ϵ+τz_τ=(1-τ)ε+τ z with target velocity z−ϵz-ε. The velocity network comprises KrK_r RF blocks, as shown in Fig. 3 (b). The representations of zτz_τ and x¯0 x_0 are summed to initialize the hidden representation. Within each block, cEDfinec_ED^fine and τ are incorporated through AdaLN (33; 51), whereas x¯0 x_0 and cEDglobalc_ED^global jointly determine the prototype-specific adapter route described below. The RF model is optimized by ℒRF=z,ϵ,τ[‖vθ(zτ,x¯0,cEDglobal,cEDfine,τ)−(z−ϵ)‖22].L_RF=E_z,ε,τ [ \|v_θ (z_τ, x_0,c_ED^global,c_ED^fine,τ )-(z-ε) \|_2^2 ]. (8) Prototype-Routed Motion Adapters A shared RF backbone models motion patterns common across subjects. To adapt this shared velocity field to heterogeneous ED anatomy and phenotypes without training separate generators, we introduce NeN_e lightweight motion adapters, each associated with one routing prototype. For each subject, we form the routing feature grouteg_route by concatenating the mean-pooled anatomical latent MeanPool(x¯0)MeanPool( x_0) with the global phenotype descriptor cEDglobalc_ED^global. Applying K-means to the training routing features yields NeN_e routing prototypes ejj=1Ne\e_j\_j=1^N_e. Each subject is assigned to its nearest routing prototype: j∗=argminj∈1,…,Ne‖groute−ej‖22,j^*= _j∈\1,…,N_e\ \|g_route-e_j \|_2^2, (9) where eje_j denotes the j-th routing prototype. The selected index j∗j^* activates the corresponding adapter in each RF block, allowing the shared velocity field to specialize according to subject-specific ED anatomy and global phenotype. Optional Motion Control ED anatomy does not completely determine functional properties such as contraction strength and motion amplitude. When the motion descriptor cmotionc^motion is available, it is encoded by a multilayer perceptron (MLP) and injected into each RF block through an additional Motion Adapter. As shown in Fig. 3 (b), the resulting residual is scaled by the control scale s and added to the hidden representation: hℓ+1←hℓ+1+sℳℓ(MLP(cmotion)),h_ +1← h_ +1+s\,M_ (MLP(c^motion) ), (10) where hℓ+1h_ +1 denotes the output hidden representation of the ℓ -th RF block and ℳℓM_ denotes its Motion Adapter. The coefficient s controls the strength of functional modulation. Inference Given an ED mesh x0x_0, the frozen ED encoder produces the anatomical latent x¯0 x_0. The routing feature grouteg_route is then constructed from x¯0 x_0 and cEDglobalc_ED^global, and Eq. (9) selects the adapter route j∗j^*, which remains fixed throughout RF integration. Starting from zτ=0=ϵz_τ=0=ε, where ϵ∼(0,I)ε (0,I), the terminal motion latent is obtained by z^=zτ=1=ϵ+∫01vθ(j∗)(zτ,x¯0,cEDfine,τ)τ, z=z_τ=1=ε+ _0^1v_θ^(j^*) (z_τ, x_0,c_ED^fine,τ )\,dτ, (11) where vθ(j∗)v_θ^(j^*) denotes the RF velocity field with the j∗j^*-th adapter route activated. When cmotionc^motion is available, the motion adapter additionally modulates the velocity field throughout integration; otherwise, this branch is disabled. Finally, the terminal latent z^=zτ=1 z=z_τ=1, together with x¯0 x_0, is passed to the frozen motion decoder to reconstruct the full-cycle ED-relative displacement trajectory u. Table 1: Composition of the study cohort across datasets and diagnostic groups. NOR: normal; DCM: dilated cardiomyopathy; HCM: hypertrophic cardiomyopathy; RVA: right-ventricular abnormality. Dataset NOR DCM HCM RVA Total ACDC 30 30 30 30 120 M&Ms 89 97 85 15 286 M&Ms-2 75 60 60 65 260 Total 194 187 175 110 666 4 Experiments and Results Table 2: Quantitative comparison under the ED-conditioned full-cycle biventricular motion synthesis. ASSD: average symmetric surface distance; HD95: 95th percentile Hausdorff distance; vRMSE: vertex-wise root mean square error; BiV: biventricular; LV: left ventricular; RV: right ventricular. Values are reported as the mean ± standard deviation across test subjects. For each subject, each metric was first averaged across 20 independently generated motion sequences. Method ASSD (m) ↓ HD95 (m) ↓ vRMSE (m) ↓ LV RV BiV LV RV BiV LV RV BiV CVAE (42) 2.16±0.662.16± 0.66 2.08±0.632.08± 0.63 2.15±0.542.15± 0.54 4.34±1.304.34± 1.30 4.78±1.774.78± 1.77 4.71±1.544.71± 1.54 3.86±1.343.86± 1.34 4.25±1.544.25± 1.54 4.11±1.254.11± 1.25 ACTOR (35) 2.10±0.642.10± 0.64 2.00±0.802.00± 0.80 2.11±0.602.11± 0.60 4.28±1.414.28± 1.41 4.85±2.334.85± 2.33 4.93±1.984.93± 1.98 3.71±1.393.71± 1.39 4.14±1.834.14± 1.83 4.01±1.524.01± 1.52 Action2Motion (15) 2.28±0.712.28± 0.71 2.18±0.782.18± 0.78 2.22±0.642.22± 0.64 4.81±1.654.81± 1.65 5.07±2.185.07± 2.18 5.16±1.985.16± 1.98 4.20±1.534.20± 1.53 4.51±1.854.51± 1.85 4.21±1.524.21± 1.52 CHeart (38) 2.21±0.652.21± 0.65 2.04±0.722.04± 0.72 2.09±0.532.09± 0.53 4.48±1.434.48± 1.43 4.73±1.984.73± 1.98 4.97±1.524.97± 1.52 3.80±1.373.80± 1.37 4.10±1.704.10± 1.70 3.98±1.223.98± 1.22 MeshHeart (37) 2.09±0.652.09± 0.65 2.01±0.632.01± 0.63 2.08±0.532.08± 0.53 4.23±1.304.23± 1.30 4.64±1.774.64± 1.77 4.68±1.544.68± 1.54 3.78±1.343.78± 1.34 4.16±1.544.16± 1.54 4.13±1.254.13± 1.25 4DCardioSynth (13) 2.07±0.612.07± 0.61 2.04±0.652.04± 0.65 2.12±0.552.12± 0.55 4.22±1.294.22± 1.29 4.73±1.724.73± 1.72 4.76±1.704.76± 1.70 3.63±1.253.63± 1.25 4.11±1.504.11± 1.50 3.94±1.333.94± 1.33 RePCM (50) 1.79±0.601.79± 0.60 2.00±0.762.00± 0.76 1.69±0.361.69± 0.36 3.94±1.483.94± 1.48 4.54±2.044.54± 2.04 4.27±1.834.27± 1.83 3.43±1.423.43± 1.42 4.08±1.744.08± 1.74 3.61±1.453.61± 1.45 Ours 1.63±0.481.63± 0.48 1.85±0.501.85± 0.50 1.49±0.341.49± 0.34 3.56±1.093.56± 1.09 4.24±1.234.24± 1.23 3.77±1.063.77± 1.06 3.13±1.093.13± 1.09 3.74±1.123.74± 1.12 3.31±1.033.31± 1.03 Figure 4: Qualitative comparison of ED-conditioned biventricular motion synthesis. Surface-error maps are shown for two representative cardiac phases from NOR, DCM, HCM, and RVA subjects. Each column corresponds to one competing method, and the color scale denotes the point-wise surface error relative to the reference mesh. The LV and RV volume curves on the right indicate the selected phases using vertical dashed lines. 4.1 Dataset and Pre-processing We used three public cine cardiac MRI datasets: ACDC (7), M&Ms (9), and M&Ms-2 (30). These datasets cover different acquisition protocols and cardiac conditions. Diagnostic labels were harmonized as normal (NOR), dilated cardiomyopathy (DCM), hypertrophic cardiomyopathy (HCM), and right ventricular abnormality (RVA). The DCM group combined ACDC/ M&Ms DCM with M&Ms-2 dilated-left-ventricle cases. The RVA group combined ACDC/ M&Ms right-ventricular-abnormality cases with M&Ms-2 arrhythmogenic-cardiomyopathy and dilated-right-ventricle cases. These groupings reflect shared LV- or RV-dominant phenotypes and were used only for stratification and subgroup evaluation. Five additional diagnoses were reserved exclusively for diagnostic out-of-distribution (OOD) evaluation. M&Ms hypertensive heart disease (HHD) and ischemic heart disease (IHD) were withheld because of limited sample sizes. M&Ms-2 tetralogy of Fallot (FALL), interatrial communication (CIA), and tricuspid regurgitation (TRI) were not merged into RVA because their defining abnormalities involve anatomical structures or flow phenomena not represented by the modeled biventricular surfaces. For each subject, the segmentation masks across the cardiac cycle were converted into a unified-topology biventricular surface-mesh sequence. We adopted an SSM-based fitting procedure based on the biventricular atlas of 1. The template was fitted to the multi-phase contours and refined through global alignment, non-rigid deformation, and temporal smoothing. This process produced anatomically consistent meshes with fixed topology and vertex-wise correspondence across subjects and cardiac frames. All sequences were temporally resampled to T=25T=25 frames. For model training, all meshes were aligned to the template coordinate system through center-of-mass matching and rigid registration, followed by normalization using a global scaling factor. The same spatial transformation was applied to all frames of each subject to preserve the relative motion trajectory. During evaluation, the generated meshes were transformed back to their original physical scale before geometric and functional measurements were computed. Following the MRI-derived phenotype definitions in 2, we extracted ED-derived phenotype descriptors from the unified biventricular meshes. The global descriptor cEDglobal∈ℝ6c_ED^global ^6 comprised the logarithmically transformed LV ED volume, RV ED volume, and LV mass; global myocardial wall thickness; the logarithmic LV-to-RV ED volume ratio; and the logarithmic LV-mass-to-LV-volume ratio. The fine-grained descriptor cEDfine∈ℝ22c_ED^fine ^22 additionally included LV myocardial wall-thickness measurements from the 16 AHA regions. When functional information was available, the optional motion descriptor cmotion∈ℝ6c^motion ^6 comprised the logarithmically transformed LV and RV end-systolic volumes, the logarithmically transformed LV and RV stroke volumes, and the LV and RV ejection fractions. The final in-distribution cohort contained 666 subjects, as summarized in Table 1. The data were divided at the patient level into training, validation, and test sets using a ratio of 7:1:27:1:2. 4.2 Gold Standard and Evaluation The SSM-fitted full-cycle mesh sequences were used as the reference standard. All methods were evaluated under the same ED-conditioned synthesis setting defined in Sec. 3.1. Before evaluation, the generated meshes were transformed back to the original physical coordinate system. Geometric agreement between the generated and reference sequences was evaluated using the average symmetric surface distance (ASSD), the 95th-percentile Hausdorff distance (HD95), and the vertex-wise root mean square error (vRMSE). Because all meshes shared fixed vertex correspondence, vRMSE was used to quantify point-wise trajectory errors over the complete cardiac cycle. For each test subject, every method generated 20 stochastic motion samples. Each metric was computed separately for the 20 samples and then averaged within the subject. Functional agreement was assessed using the LV and RV volume trajectories derived from the synthesized sequences. Agreement in LVEF and RVEF was quantified using Pearson’s correlation coefficient r and mean absolute error (MAE). Population-level agreement in LVESV, RVESV, LVEF, and RVEF was further evaluated using the KL divergence and Wasserstein distance. 4.3 Implementation All models were implemented in PyTorch and trained on a single NVIDIA RTX A5500 GPU. The motion-informed functional partition contained NR=16N_R=16 regions, and the motion VAE employed Na=512N_a=512 region-balanced anchor tokens. The token feature dimension and motion latent dimension were set to d=256d=256 and dz=32d_z=32, respectively. Both the encoder and decoder contained Km=8K_m=8 attention blocks, with four attention heads in each block. The maximum hop range and attenuation factor of MRAB were set to Lhop=4L_hop=4 and η=0.5η=0.5, respectively. The motion VAE was trained for 500 epochs using Adam with a learning rate of 1×10−41× 10^-4, a batch size of 8, and a KL-divergence weight of λKL=5×10−4 _KL=5× 10^-4. The rectified-flow network contained KrK_r = 4 blocks with a hidden dimension of 128. The fine-grained descriptor cEDfine∈ℝ22c_ED^fine ^22 conditioned the RF blocks through AdaLN. The global descriptor cEDglobal∈ℝ6c_ED^global ^6 was combined with the anatomical latent solely for selecting among Ne=4N_e=4 prototype-routed adapters. The RF model was trained for 200 epochs using AdamW with a learning rate of 1×10−51× 10^-5, a batch size of 8, and a weight decay of 5×10−35× 10^-3. During inference, the rectified-flow ODE was solved using NEuler=4N_Euler=4 uniform Euler steps over τ∈[0,1]τ∈[0,1]. In the optional functional-control experiments, the six-dimensional motion descriptor cmotionc^motion was used with a control strength of s=1s=1. This branch was disabled in the primary ED-conditioned setting. Figure 5: Functional fidelity of the synthesized biventricular sequences. (a) Diagnostic group-stratified relative volume-change trajectories for the LV (top) and RV (bottom). Solid and dashed curves represent the reference and synthesized sequences, respectively, and shaded bands indicate inter-subject variability. (b) Agreement between reference and synthesized left- and right-ventricular ejection fractions (LVEF and RVEF), with the identity line shown for reference. (c) Normalized KL divergence and Wasserstein distance (WD) between the reference and synthesized left- and right-ventricular end-systolic volumes (LVESV and RVESV), LVEF, and RVEF distributions. Lower values indicate better distributional agreement. MAE: mean absolute error. 4.4 Comparison Study Table 2 reports the quantitative comparison under the ED-conditioned full-cycle motion synthesis setting. We compared the proposed method with representative sequence generation and cardiac motion synthesis baselines, including a conditional VAE (42), ACTOR (35), Action2Motion (15), CHeart (38), MeshHeart (37), 4DCardioSynth (13), and RePCM (50). For fairness, all baselines used the same pre-processed mesh sequences, ED input, temporal resolution, and patient-level split. One can see that the proposed method achieved the best overall performance across all metrics of different chambers and their combination, indicating more accurate full-cycle motion synthesis than both generic motion-generation baselines and cardiac-specific competing methods. Compared with all external baselines, our method consistently achieved lower biventricular (BiV), left-ventricular (LV), and right-ventricular (RV) surface and vertex-wise errors. Compared with the preliminary RePCM framework, our method significantly reduced BiV ASSD from 1.69 to 1.49 m (p<0.01p<0.01), BiV HD95 from 4.27 to 3.77 m (p<0.01p<0.01), and BiV vRMSE from 3.61 to 3.31 m (p<0.05p<0.05), as determined by subject-level paired Wilcoxon signed-rank tests. The RV results are also noteworthy because RePCM showed comparatively limited gains over the external baselines for this chamber. RV motion synthesis is particularly challenging because of the chamber’s thin myocardial wall, complex geometry, and substantial inter-subject variability. Compared with RePCM, our method reduced RV ASSD from 2.00 to 1.85 m and RV vRMSE from 4.08 to 3.74 m, indicating more accurate synthesis of challenging RV dynamics. Figure 4 provides representative surface-error maps for all four diagnostic groups. Competing methods frequently produced spatially extended errors around the basal ventricular surfaces and the RV free wall, particularly near phases of maximal contraction. In contrast, the proposed method yielded smaller and more spatially localized errors across the selected systolic and relaxation phases. The improvement was consistent across NOR, DCM, HCM, and RVA cases, supporting the quantitative findings in Table 2. Figure 5 presents whether the synthesized sequences preserve clinically relevant ventricular dynamics. The diagnostic group-stratified relative LV and RV volume trajectories closely followed the reference contraction-relaxation patterns, including the reduced LV contraction observed in the DCM group and the stronger relative LV volume reduction in HCM. Across individual subjects, the synthesized LVEF achieved a correlation of r=0.90r=0.90 with an MAE of 8.1%, while RVEF achieved r=0.76r=0.76 with an MAE of 9.6%. The proposed method also produced the smallest normalized KL divergence and Wasserstein distance across LVESV, RVESV, LVEF, and RVEF among the compared methods. Thus, the geometric improvements translated into more faithful ventricular volume trajectories and functional-index distributions. 4.5 Ablation Study 4.5.1 Effect of Motion-Informed Parcellation Table 3 presents the contribution of the motion-informed regional prior and the effect of partition granularity. Compared with the AHA-like partition, the motion-informed parcellation with NR=8N_R=8 and NR=16N_R=16 reduced all BiV errors, confirming the benefit of deriving functional regions directly from cardiac dynamics. The best overall performance was obtained with NR=16N_R=16, whereas increasing the number of regions to NR=32N_R=32 removed these gains. This indicates that performance does not improve monotonically with finer regionalization, as excessive fragmentation may weaken stable regional aggregation. Figure 6: Regional analysis of motion-informed parcellation. (a) Surface-group vRMSE for the AHA-like partition and motion-informed parcellation with different numbers of regions. (b) Region-wise vRMSE reduction obtained by the NR=16N_R=16 motion-informed partition relative to the AHA-like partition, reported for each diagnostic group and the full cohort. Positive values indicate lower errors with motion-informed parcellation. Figure 6 provides a more detailed analysis across ventricular surfaces and learned motion regions. As shown in Fig. 6 (a), the NR=16N_R=16 partition achieved the lowest vRMSE across all four ventricular surface groups, while the RV free wall remained the most challenging surface under all partition strategies. Fig. 6 (b) further reports the region-wise vRMSE reduction of the NR=16N_R=16 partition relative to the AHA-like partition, where positive values indicate lower errors with motion-informed parcellation. At the cohort level, improvements were observed across all 16 learned regions. Similar improvements were found across most diagnostic group-region combinations, with only a few isolated exceptions. These results show that the benefit of motion-informed parcellation was distributed across the ventricular surface rather than being driven by a small subset of regions. Table 3: Effect of regional partition strategies on biventricular motion synthesis. Partition ASSD (m) ↓ HD95 (m) ↓ vRMSE (m) ↓ AHA-like 1.59±0.441.59± 0.44 4.11±1.744.11± 1.74 3.62±1.433.62± 1.43 NR=8N_R=8 1.50±0.371.50± 0.37 3.88±1.393.88± 1.39 3.40±1.193.40± 1.19 NR=16N_R=16 1.49±0.341.49± 0.34 3.77±1.063.77± 1.06 3.31±1.033.31± 1.03 NR=32N_R=32 1.59±0.381.59± 0.38 4.19±1.364.19± 1.36 3.62±1.213.62± 1.21 4.5.2 Effect of Multi-Hop Regional Attention Bias Table 4 summarizes the effect of the MRAB used in the motion VAE. Removing the topology bias increased BiV vRMSE from 3.31±1.033.31± 1.03 m to 3.51±1.083.51± 1.08 m, confirming the contribution of topology-guided regional communication. Introducing MRAB improved performance across most hop ranges, indicating that graded topology-aware communication among the learned regions facilitates coherent motion propagation. The Lhop=4L_hop=4 configuration achieved comparable mean accuracy to Lhop=2L_hop=2, while exhibiting consistently lower inter-subject variability across all three metrics. Table 4: Effect of MRAB on biventricular motion synthesis performance. MRAB: multi-hop regional attention bias. Topology setting ASSD (m) ↓ HD95 (m) ↓ vRMSE (m) ↓ w/o MRAB 1.56±0.341.56± 0.34 3.92±1.273.92± 1.27 3.51±1.083.51± 1.08 Lhop=1L_hop=1 1.55±0.421.55± 0.42 3.81±1.713.81± 1.71 3.35±1.383.35± 1.38 Lhop=2L_hop=2 1.49±0.381.49± 0.38 3.75±1.553.75± 1.55 3.31±1.243.31± 1.24 Lhop=3L_hop=3 1.49±0.431.49± 0.43 3.86±1.803.86± 1.80 3.37±1.443.37± 1.44 Lhop=4L_hop=4 1.49±0.341.49± 0.34 3.77±1.063.77± 1.06 3.31±1.033.31± 1.03 Lhop=5L_hop=5 1.55±0.401.55± 0.40 3.94±1.453.94± 1.45 3.41±1.313.41± 1.31 4.5.3 Effect of Phenotype Conditioning and Prototype Routing Table 5: Ablation of phenotype conditioning, prototype routing, and optional functional control in the rectified-flow generator. Superscripts G and F denote the global descriptor cEDglobalc_ED^global and fine-grained descriptor cEDfinec_ED^fine, respectively. The final row uses the additional motion descriptor cmotionc^motion and is not part of the primary ED-only setting. cEDc_ED Adapter cmotionc^motion ASSD (m)↓ HD95 (m)↓ vRMSE (m)↓ Condition Route Control × × × 1.64±0.421.64± 0.42 4.17±1.784.17± 1.78 3.60±1.433.60± 1.43 ✓G ^G × × 1.56±0.371.56± 0.37 4.09±1.534.09± 1.53 3.47±1.253.47± 1.25 ✓F ^F × × 1.51±0.361.51± 0.36 3.92±1.123.92± 1.12 3.36±1.163.36± 1.16 × ✓G ^G × 1.52±0.361.52± 0.36 3.99±1.473.99± 1.47 3.40±1.163.40± 1.16 × ✓F ^F × 1.54±0.371.54± 0.37 3.97±1.443.97± 1.44 3.43±1.173.43± 1.17 ✓G ^G ✓G ^G × 1.51±0.361.51± 0.36 3.83±1.523.83± 1.52 3.41±1.193.41± 1.19 ✓G ^G ✓F ^F × 1.56±0.351.56± 0.35 4.00±1.234.00± 1.23 3.54±1.183.54± 1.18 ✓F ^F ✓G ^G × 1.49±0.341.49± 0.34 3.77±1.063.77± 1.06 3.31±1.033.31± 1.03 ✓F ^F ✓F ^F × 1.54±0.371.54± 0.37 3.97±1.443.97± 1.44 3.39±1.163.39± 1.16 ✓F ^F ✓G ^G ✓ 1.46±0.391.46± 0.39 3.74±1.493.74± 1.49 3.21±1.233.21± 1.23 Figure 7: Diagnostic-group distributions across prototype-routed motion adapters. Row-normalized adapter assignments are shown for routing using the anatomical latent together with (a) the global ED descriptor cEDglobalc_ED^global or (b) the fine-grained ED descriptor cEDfinec_ED^fine. Each cell reports the percentage of subjects in a diagnostic group assigned to the corresponding adapter. Table 5 evaluates the two descriptor-based conditioning mechanisms used in Sec. 3.4. Relative to the baseline without explicit phenotype conditioning or routing, fine-grained ED conditioning consistently improved all three metrics, while global-descriptor-based routing also yielded clear gains across the three measures. Combining cEDfinec_ED^fine for latent conditioning with cEDglobalc_ED^global for adapter routing achieved the best ED-only result, with an ASSD of 1.491.49 m, HD95 of 3.773.77 m, and vRMSE of 3.313.31 m. The two mechanisms therefore provide complementary benefits: fine-grained descriptors guide subject-specific latent generation, whereas global descriptors provide a more discriminative signal for routing among the motion adapters. Figure 7 explains the different roles of the two descriptors. Routing based on cEDglobalc_ED^global assigned DCM and HCM subjects predominantly to distinct adapters, while NOR and RVA subjects were distributed mainly across the remaining routes. In contrast, routing with cEDfinec_ED^fine produced greater overlap across diagnostic groups and distributed HCM subjects across several adapters. Fine-grained descriptors are therefore useful as continuous generation conditions, but may contain route-irrelevant local variation that weakens the separation of global motion modes. Accordingly, the final model uses cEDfinec_ED^fine for latent conditioning and cEDglobalc_ED^global for prototype routing. Table 6: Effect of the number of routing prototypes/ adapters NeN_e on biventricular motion synthesis performance. NeN_e ASSD (m) ↓ HD95 (m) ↓ vRMSE (m) ↓ 4 1.49±0.341.49± 0.34 3.77±1.063.77± 1.06 3.31±1.033.31± 1.03 6 1.47±0.381.47± 0.38 3.82±1.583.82± 1.58 3.36±1.303.36± 1.30 8 1.50±0.401.50± 0.40 3.94±1.723.94± 1.72 3.41±1.383.41± 1.38 Figure 8: Geometric errors across in-distribution (ID) and out-of-distribution (OOD) diagnostic groups. (a) Four ID groups shared by ACDC, M&Ms, and M&Ms-2. (b) Five OOD groups excluded from training: hypertensive heart disease (HHD) and ischemic heart disease (IHD) from M&Ms, and tetralogy of Fallot (FALL), interatrial communication (CIA), and tricuspid regurgitation (TRI) from M&Ms-2. Figure 9: Illustration of full-cycle synthesized mesh sequences for three representative OOD cases and their corresponding LV and RV volume trajectories. Figure 10: Controllable modulation of synthesized cardiac motion. (a) Mean LV and RV displacement trajectories for the reference, the ED-only synthesis, and predictions obtained using the optional functional-control branch. With cmotionc^motion uses the subject-specific motion descriptor of the illustrated NOR case, whereas the DCM-mean and HCM-mean curves use the corresponding group-average descriptors. (b) Response of a representative NOR case to control scale ranging from 0 to 2, including the relationship between the control coefficient and peak normalized mean displacement and the corresponding spatial maps of the cmotionc^motion-induced deformation at s=0s=0, s=1s=1, and s=2s=2. Table 6 evaluates the sensitivity to the number of routing prototypes/ adapters NeN_e. Performance was relatively stable across different values of NeN_e. Although Ne=6N_e=6 achieved a slightly lower ASSD, Ne=4N_e=4 produced the best HD95 and vRMSE with lower variance. We used Ne=4N_e=4 as the default setting, providing a compact routing module while maintaining stable motion synthesis accuracy. 4.6 In- and Out-of-Distribution Diagnostic-Group Analysis Figure 8 (a) reports the geometric errors across four in-distribution (ID) diagnostic groups shared by ACDC, M&Ms, and M&Ms-2. Performance remained broadly stable across the four groups, with no marked degradation observed in any single group. DCM and HCM exhibited somewhat higher median ASSD than NOR and RVA, reflecting the greater variability associated with pathological ventricular remodeling and contraction. Nevertheless, their interquartile ranges remained substantially overlapped. Fig. 8 (b) further evaluates five OOD diagnostic groups from M&Ms and M&Ms-2 that were completely excluded from training: HHD, IHD, FALL, CIA, and TRI. Although some OOD groups exhibited greater error variability, particularly CIA, the overall error distributions remained within a comparable range across the two datasets. These results suggest that the learned motion representation generalizes beyond the diagnostic groups observed during training rather than relying on a single dominant disease pattern. Representative synthesized sequences are shown in Fig. 9 for OOD cases with CIA, FALL, and HHD. The generated biventricular surfaces evolve smoothly throughout the cardiac cycle and exhibit distinct LV and RV volume trajectories across the three cases. These examples illustrate that the model generates temporally coherent motion with distinct case-specific ventricular dynamics across the three OOD cases. 4.7 Optional Functional Motion Control The primary model infers motion solely from ED anatomy and ED-derived phenotype descriptors. When additional motion descriptors are available, the optional cmotionc^motion branch enables explicit modulation of the generated motion. As shown in the final row of Table 5, incorporating this control reduced BiV ASSD from 1.491.49 to 1.461.46 m and vRMSE from 3.313.31 to 3.213.21 m. Fig. 10 (a) shows that the controlled LV and RV mean-displacement trajectories more closely followed the reference than the uncontrolled synthesis. Conditioning on the mean motion descriptors derived from the DCM and HCM groups produced distinct displacement amplitudes, demonstrating descriptor-dependent modulation of cardiac motion. Fig. 10 (b) further evaluates a representative NOR case under control scale ranging from 0 to 2. The peak normalized mean displacement increased monotonically and approximately linearly with the control scale, while the corresponding spatial maps showed smooth and spatially varying deformation. These results indicate that the optional branch provides interpretable control over motion magnitude without changing the primary ED-only inference setting. 5 Discussion and Conclusion This study presents a region-specific and phenotype-adaptive framework for synthesizing full-cycle biventricular motion from a single ED mesh. Across three public cine MRI datasets and four in-distribution diagnostic groups, the proposed ED-conditioned model consistently outperformed generic motion generators and cardiac-specific synthesis methods in surface- and correspondence-based metrics. The improvement over RePCM indicates that replacing global latent generation with structured motion representation and phenotype-conditioned RF generation provides a more effective formulation for this task. The gains were particularly evident for the RV, whose complex geometry and greater inter-subject variability make its motion more difficult to recover. Importantly, the geometric improvements were accompanied by better functional fidelity: the synthesized ventricular volume trajectories preserved diagnostic-group-specific ventricular motion patterns, and the derived LVEF and RVEF showed meaningful agreement with their reference values. The lower distributional discrepancies in ventricular volumes and ejection fractions further suggest that the model captures clinically relevant population-level variation rather than merely minimizing point-wise geometric errors. The OOD evaluation on M&Ms and M&Ms-2 further indicates that the learned representation generalizes beyond the diagnostic groups observed during training, although the increased variability in some OOD groups highlights the remaining challenge of uncommon pathological anatomies. The ablation studies provide complementary evidence for the principal design choices. Motion-informed functional parcellation consistently outperformed the AHA-like partition, demonstrating that regions learned directly from deformation patterns are better suited to generative motion modeling than predefined anatomical divisions. The strongest performance was obtained with 16 regions, whereas excessive partitioning weakened regional aggregation, indicating that the regional prior should balance local specificity with motion coherence. MRAB further improved synthesis by permitting graded communication between neighboring functional regions while suppressing unrelated interactions. In the latent generator, fine-grained ED phenotype descriptors were more effective for conditioning motion generation, whereas global ED phenotype descriptors produced clearer and more stable prototype routing. This distinction suggests that detailed regional morphology is useful for specifying subject-level dynamics, while global anatomy is better suited to identifying broader motion modes. The optional functional-control branch further modulated motion magnitude in a continuous manner, but its results should be interpreted separately from the primary ED-only setting because it requires additional motion descriptors unavailable from the ED anatomy alone. Several limitations remain. First, the reference sequences were obtained through SSM fitting of segmentation masks and therefore inherit errors introduced by image segmentation, template fitting, and temporal smoothing; evaluation against independently tracked or manually verified 4D meshes would provide a stronger assessment of motion accuracy. Second, the current study focuses on unified-topology biventricular surfaces and four relatively broad diagnostic groups. Its applicability to atrial motion, myocardial-layer deformation, congenital abnormalities, and more heterogeneous or subtle disease subtypes remains to be investigated. Third, although ED anatomy provides an accessible patient-specific condition, it cannot uniquely determine functional properties such as contraction strength, electromechanical delay, or regional dysfunction, as also reflected by the remaining errors in RVEF estimation. Future work should therefore incorporate complementary imaging, clinical, or physiological conditions when available, evaluate generalization on independent clinical cohorts and alternative mesh-construction pipelines, and extend the framework toward controllable whole-heart motion generation and virtual population synthesis for downstream computational modeling and in silico studies. References Bai et al. (2015) W. Bai, W. Shi, A. de Marvao, T. J. Dawes, D. P. O’Regan, S. A. Cook, and D. Rueckert A bi-ventricular cardiac atlas built from 1000+ high resolution mr images of healthy subjects and an analysis of shape and motion. Medical image analysis 26 (1), p. 133–145. Cited by: §1, §1, §2.1, §4.1. Bai et al. (2020) W. Bai, H. Suzuki, J. Huang, C. Francis, S. Wang, G. Tarroni, F. Guitton, N. Aung, K. Fung, S. E. Petersen, et al. A population-based phenome-wide association study of cardiac and aortic structure and function. Nature medicine 26 (10), p. 1654–1662. Cited by: §2.1, §3.1, §4.1. Banerjee et al. (2022) A. Banerjee, E. Zacur, R. P. Choudhury, and V. Grau Automated 3d whole-heart mesh reconstruction from 2d cine mr slices using statistical shape model. In 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), p. 1702–1706. Cited by: §1. Beetz et al. (2021) M. Beetz, A. Banerjee, and V. Grau Generating subpopulation-specific biventricular anatomy models using conditional point cloud variational autoencoders. In International Workshop on Statistical Atlases and Computational Models of the Heart, p. 75–83. Cited by: §2.1. Beetz et al. (2025) M. Beetz, A. Banerjee, L. Li, J. Camps, B. Rodriguez, and V. Grau 3D cardiac shape analysis with variational point cloud autoencoders for myocardial infarction prediction and virtual heart synthesis. Computerized Medical Imaging and Graphics 124, p. 102587. Cited by: §1, §2.1. Beetz et al. (2022) M. Beetz, J. Corral Acero, A. Banerjee, I. Eitel, E. Zacur, T. Lange, T. Stiermaier, R. Evertz, S. J. Backhaus, H. Thiele, et al. Interpretable cardiac anatomy modeling using variational mesh autoencoders. Frontiers in cardiovascular medicine 9, p. 983868. Cited by: §1. Bernard et al. (2018) O. Bernard, A. Lalande, C. Zotti, F. Cervenansky, X. Yang, P. Heng, I. Cetin, K. Lekadir, O. Camara, M. A. G. Ballester, et al. Deep learning techniques for automatic mri cardiac multi-structures segmentation and diagnosis: is the problem solved?. IEEE transactions on medical imaging 37 (11), p. 2514–2525. Cited by: §4.1. Biglino et al. (2017) G. Biglino, C. Capelli, J. Bruse, G. M. Bosi, A. M. Taylor, and S. Schievano Computational modelling for congenital heart disease: how far are we from clinical translation?. Heart 103 (2), p. 98–103. Cited by: §2.1. Campello et al. (2021) V. M. Campello, P. Gkontra, C. Izquierdo, C. Martin-Isla, A. Sojoudi, P. M. Full, K. Maier-Hein, Y. Zhang, Z. He, J. Ma, et al. Multi-centre, multi-vendor and multi-disease cardiac segmentation: the m&ms challenge. IEEE Transactions on Medical Imaging 40 (12), p. 3543–3554. Cited by: §4.1. Chen et al. (2026) Y. Chen, J. Yang, D. S. Mercadier, H. Le, J. Schwitter, and P. Fua End-to-end 4d heart mesh recovery across full-stack and sparse cardiac mri. External Links: 2509.12090 Cited by: §1. Chung et al. (2015) J. Chung, K. Kastner, L. Dinh, K. Goel, A. C. Courville, and Y. Bengio A recurrent latent variable model for sequential data. Advances in neural information processing systems 28. Cited by: §2.3. Dhariwal and Nichol (2021) P. Dhariwal and A. Nichol Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34, p. 8780–8794. Cited by: §2.3. Dou et al. (2025) H. Dou, J. Huang, A. Zakeri, Z. Zhou, T. Mu, J. Duan, and A. F. Frangi 4D cardiosynth: synthesising dynamic virtual heart populations through spatiotemporal disentanglement. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 3–12. Cited by: §1, §1, §2.1, §4.4, Table 2. Dou et al. (2023) H. Dou, N. Ravikumar, and A. F. Frangi A conditional flow variational autoencoder for controllable synthesis of virtual populations of anatomy. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 143–152. Cited by: §1, §2.1. Guo et al. (2020) C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, and L. Cheng Action2motion: conditioned generation of 3d human motions. In Proceedings of the 28th ACM international conference on multimedia, p. 2021–2029. Cited by: §2.3, §4.4, Table 2. Ho et al. (2020) J. Ho, A. Jain, and P. Abbeel Denoising diffusion probabilistic models. Advances in neural information processing systems 33, p. 6840–6851. Cited by: §2.3. Hu et al. (2023) V. T. Hu, W. Yin, P. Ma, Y. Chen, B. Fernando, Y. M. Asano, E. Gavves, P. Mettes, B. Ommer, and C. G. M. Snoek Motion flow matching for human motion synthesis and editing. External Links: 2312.08895 Cited by: §2.3. Kingma and Welling (2013) D. P. Kingma and M. Welling Auto-encoding variational bayes. External Links: 1312.6114 Cited by: §2.3. Kong and Shadden (2022) F. Kong and S. C. Shadden Learning whole heart mesh generation from patient images for computational simulations. IEEE Transactions on Medical Imaging 42 (2), p. 533–545. Cited by: §1. Kong et al. (2024) F. Kong, S. Stocker, P. S. Choi, M. Ma, D. B. Ennis, and A. L. Marsden SDF4CHD: generative modeling of cardiac anatomies with congenital heart defects. Medical image analysis 97, p. 103293. Cited by: §2.1. Kong et al. (2021) F. Kong, N. Wilson, and S. Shadden A deep-learning approach for direct whole-heart mesh reconstruction. Medical image analysis 74, p. 102222. Cited by: §1, §1. Laumer et al. (2023) F. Laumer, M. Amrani, L. Manduchi, A. Beuret, L. Rubi, A. Dubatovka, C. M. Matter, and J. M. Buhmann Weakly supervised inference of personalized heart meshes based on echocardiography videos. Medical image analysis 83, p. 102653. Cited by: §2.2. Lipman et al. (2023) Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le Flow matching for generative modeling. External Links: 2210.02747 Cited by: §2.3, §3.4. Liu et al. (2026) X. Liu, X. Yuan, M. Y. Chan, C. Sia, and L. Li CineMesh4D: personalized 4d whole heart reconstruction from sparse cine mri. External Links: 2605.13994 Cited by: §1. Liu et al. (2022) X. Liu, C. Gong, and Q. Liu Flow straight and fast: learning to generate and transfer data with rectified flow. External Links: 2209.03003 Cited by: §2.3, §3.4. López et al. (2023) P. A. López, H. Mella, S. Uribe, D. E. Hurtado, and F. S. Costabal WarpPINN: cine-mr image registration with physics-informed neural networks. Medical Image Analysis 89, p. 102925. Cited by: §2.2. Lu et al. (2023) J. Lu, R. Jin, M. Wang, E. Song, and G. Ma A bidirectional registration neural network for cardiac motion tracking using cine mri images. Computers in Biology and Medicine 160, p. 107001. Cited by: §2.2. Luo et al. (2026) Y. Luo, D. Sesia, F. Wang, Y. Wu, W. Ding, K. Hasan, J. Huang, F. Shi, A. Shah, A. Kaura, et al. Explicit differentiable slicing and global deformation for cardiac mesh reconstruction. Medical image analysis, p. 103999. Cited by: §1. Marsden and Feinstein (2015) A. L. Marsden and J. A. Feinstein Computational modeling and engineering in pediatric and congenital heart disease. Current opinion in pediatrics 27 (5), p. 587–596. Cited by: §1. Martín-Isla et al. (2023) C. Martín-Isla, V. M. Campello, C. Izquierdo, K. Kushibar, C. Sendra-Balcells, P. Gkontra, A. Sojoudi, M. J. Fulton, T. W. Arega, K. Punithakumar, et al. Deep learning segmentation of the right ventricle in cardiac mri: the m&ms challenge. IEEE Journal of Biomedical and Health Informatics 27 (7), p. 3302–3313. Cited by: §4.1. Meng et al. (2023) Q. Meng, W. Bai, D. P. O’Regan, and D. Rueckert DeepMesh: mesh-based cardiac motion tracking using deep learning. IEEE transactions on medical imaging 43 (4), p. 1489–1500. Cited by: §1, §1, §2.2. Niederer et al. (2020) S. A. Niederer, Y. Aboelkassem, C. D. Cantwell, C. Corrado, S. Coveney, E. M. Cherry, T. Delhaas, F. H. Fenton, A. V. Panfilov, P. Pathmanathan, et al. Creation and application of virtual patient cohorts of heart models. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 378 (2173). Cited by: §1. Peebles and Xie (2023) W. Peebles and S. Xie Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4195–4205. Cited by: §3.4. Perez et al. (2017) E. Perez, F. Strub, H. de Vries, V. Dumoulin, and A. Courville FiLM: visual reasoning with a general conditioning layer. External Links: 1709.07871 Cited by: §3.3. Petrovich et al. (2021) M. Petrovich, M. J. Black, and G. Varol Action-conditioned 3d human motion synthesis with transformer vae. In Proceedings of the IEEE/CVF international conference on computer vision, p. 10985–10995. Cited by: §2.3, §4.4, Table 2. Puyol-Anton et al. (2017) E. Puyol-Anton, M. Sinclair, B. Gerber, M. S. Amzulescu, H. Langet, M. De Craene, P. Aljabar, P. Piro, and A. P. King A multimodal spatiotemporal cardiac motion atlas from mr and ultrasound data. Medical image analysis 40, p. 96–110. Cited by: §1. Qiao et al. (2025) M. Qiao, K. A. McGurk, S. Wang, P. M. Matthews, D. P. O’Regan, and W. Bai A personalized time-resolved 3d mesh generative model for unveiling normal heart dynamics. Nature Machine Intelligence 7 (5), p. 800–811. Cited by: §1, §1, §2.1, §4.4, Table 2. Qiao et al. (2023) M. Qiao, S. Wang, H. Qiu, A. De Marvao, D. P. O’Regan, D. Rueckert, and W. Bai Cheart: a conditional spatio-temporal generative model for cardiac anatomy. IEEE transactions on medical imaging 43 (3), p. 1259–1269. Cited by: §1, §1, §2.1, §4.4, Table 2. Qiao et al. (2020) M. Qiao, Y. Wang, Y. Guo, L. Huang, L. Xia, and Q. Tao Temporally coherent cardiac motion tracking from cine mri: traditional registration method and modern cnn method. Medical Physics 47 (9), p. 4189–4198. Cited by: §2.2. Qin et al. (2023) C. Qin, S. Wang, C. Chen, W. Bai, and D. Rueckert Generative myocardial motion tracking via latent space exploration with biomechanics-informed prior. Medical Image Analysis 83, p. 102682. Cited by: §2.2. Roth et al. (2020) G. A. Roth, G. A. Mensah, C. O. Johnson, G. Addolorato, E. Ammirati, L. M. Baddour, N. C. Barengo, A. Z. Beaton, E. J. Benjamin, C. P. Benziger, et al. Global burden of cardiovascular diseases and risk factors, 1990–2019: update from the gbd 2019 study. Journal of the American college of cardiology 76 (25), p. 2982–3021. Cited by: §1. Sohn et al. (2015) K. Sohn, H. Lee, and X. Yan Learning structured output representation using deep conditional generative models. Advances in neural information processing systems 28. Cited by: §2.3, §4.4, Table 2. Song et al. (2022) J. Song, C. Meng, and S. Ermon Denoising diffusion implicit models. External Links: 2010.02502 Cited by: §2.3. Sørensen et al. (2024) K. Sørensen, P. Diez, J. Margeta, Y. El Youssef, M. Pham, J. J. Pedersen, T. Kühl, O. De Backer, K. Kofoed, O. Camara, et al. Spatio-temporal neural distance fields for conditional generative modeling of the heart. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 422–432. Cited by: §1, §1. Upendra et al. (2021) R. R. Upendra, S. K. Hasan, R. Simon, B. J. Wentz, S. M. Shontz, M. S. Sacks, and C. A. Linte Motion extraction of the right ventricle from 4d cardiac cine mri using a deep learning-based deformable registration framework. In 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), p. 3795–3799. Cited by: §1. Wan et al. (2015) M. Wan, W. Huang, J. Zhang, X. Zhao, R. S. Tan, X. Wan, and L. Zhong Variational reconstruction of left cardiac structure from cmr images. PloS one 10 (12), p. e0145570. Cited by: §1. Wu et al. (2025) Z. Wu, C. Yu, F. Wang, and X. Bai Animateanymesh: a feed-forward 4d foundation model for text-driven universal mesh animation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 13557–13568. Cited by: §2.3, §3.2, §3.3. Xue et al. (2018) W. Xue, G. Brahm, S. Leung, O. Shmuilovich, and S. Li Cardiac motion scoring with segment-and subject-level non-local modeling. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 437–445. Cited by: §1. Yang et al. (2024) J. Yang, Y. Lin, B. Pu, and X. Li Bidirectional recurrence for cardiac motion tracking with gaussian process latent coding. Advances in Neural Information Processing Systems 37, p. 34800–34823. Cited by: §2.2. Yang et al. (2026) X. Yang, X. Yuan, H. Li, L. Chen, Y. Liu, and L. Li RePCM: region-specific and phenotype-adaptive bi-ventricular cardiac motion synthesis. External Links: 2605.21237 Cited by: §1, §3.2, §4.4, Table 2. Yang et al. (2025) Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al. Cogvideox: text-to-video diffusion models with an expert transformer. In International Conference on Learning Representations, Vol. 2025, p. 83048–83077. Cited by: §3.4. Ye et al. (2021) M. Ye, M. Kanski, D. Yang, Q. Chang, Z. Yan, Q. Huang, L. Axel, and D. Metaxas Deeptag: an unsupervised deep learning method for motion tracking on cardiac tagging magnetic resonance images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 7261–7271. Cited by: §2.2. Yuan et al. (2023) X. Yuan, C. Liu, and Y. Wang 4D myocardium reconstruction with decoupled motion and shape model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 21252–21262. Cited by: §1. Zakeri et al. (2023) A. Zakeri, A. Hokmabadi, N. Bi, I. Wijesinghe, M. G. Nix, S. E. Petersen, A. F. Frangi, Z. A. Taylor, and A. Gooya DragNet: learning-based deformable registration for realistic cardiac mr sequence generation from a single frame. Medical Image Analysis 83, p. 102678. Cited by: §1, §2.2.