Paper deep dive
A Unified Model for Highly Accurate ECG-Free Dynamic Coronary Roadmapping Using Spatio-Temporal Transformers
Saahil Islam, Sebastian Piat, Venkatesh N. Murthy, Serkan Cimen, Puneet Sharma, Andreas Maier, Florin C. Ghesu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/14/2026, 3:28:58 AM
Summary
This paper presents a unified deep learning framework for ECG-free Dynamic Coronary Roadmapping (DRM) during Percutaneous Coronary Intervention (PCI). The method utilizes a large-scale spatio-temporal transformer encoder pretrained on 16 million X-ray frames to simultaneously perform cardiac phase matching and catheter tip tracking. By integrating auxiliary tasks like ECG R-peak detection and employing a majority-voting post-processing strategy, the model achieves state-of-the-art accuracy and robustness while minimizing the need for extensive manual annotations, ultimately reducing radiation exposure and contrast agent risks in clinical practice.
Entities (10)
Relation Signals (8)
Spatio-Temporal Transformer Encoder β pretrainedon β 16 Million X-ray Frames
confidence 97% Β· Our method employs a large-scale spatio-temporal encoder pretrained on 16 million X-ray frames to learn cardiac motion dynamics.
Spatio-Temporal Transformer Encoder β performstask β Cardiac Phase Matching
confidence 96% Β· Our method employs a large-scale spatio-temporal encoder pretrained on 16 million X-ray frames to learn cardiac motion dynamics.
Siemens Healthineers β affiliatedwith β Saahil Islam
confidence 95% Β· Digital Technology and Innovation, Siemens Healthineers, Princeton, USA
Dynamic Coronary Roadmapping β reducesriskof β Percutaneous Coronary Intervention
confidence 94% Β· DRM reduces these risks by overlaying a precomputed angiographic vessel map onto live fluoroscopy and continuously updating it throughout the procedure.
Unified Model β usesauxiliarytask β ECG R-peak Detection
confidence 93% Β· We further introduce auxiliary tasks based on ECG R-peak detection and catheter tip tracking, improving optimization while eliminating the need for extensive catheter mask annotations.
Majority-Voting Post-processing β improves β Cardiac Phase Matching
confidence 91% Β· Finally, a majority-voting postprocessing strategy aggregates temporal predictions, improving robustness and providing a confidence score that correlates with phase-matching error.
Triplet Loss β optimizes β Cardiac Phase Matching
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Percutaneous Coronary Intervention (PCI) is a minimally invasive procedure used to restore coronary blood flow obstructed by atherosclerotic plaque. During PCI, repeated injections of iodine-based contrast agents are required to visualize the coronary arteries and guide interventional devices. However, frequent contrast injections increase radiation exposure and the risk of contrast-induced nephropathy, with acute kidney injury reported in up to 30% of patients with renal impairment. Dynamic Coronary Roadmapping (DRM) reduces these risks by overlaying a precomputed angiographic vessel map onto live fluoroscopy and continuously updating it throughout the procedure. Accurate DRM relies on precise cardiac phase matching between angiography and fluoroscopy, together with reliable catheter tip tracking for motion compensation. These tasks remain challenging in ECG-free settings and when only limited manual annotations are available. We present a unified DRM framework that simultaneously performs cardiac phase matching and catheter tip tracking for accurate real-time guidance. Our method employs a large-scale spatio-temporal encoder pretrained on 16 million X-ray frames to learn cardiac motion dynamics. To the best of our knowledge, this is the first application of large-scale spatio-temporal pretraining for motion compensation in DRM. We further introduce auxiliary tasks based on ECG R-peak detection and catheter tip tracking, improving optimization while eliminating the need for extensive catheter mask annotations. Finally, a majority-voting postprocessing strategy aggregates temporal predictions, improving robustness and providing a confidence score that correlates with phase-matching error. Comprehensive evaluation on clinical X-ray datasets demonstrates state-of-the-art performance, achieving low temporal misalignment and robust phase-matching accuracy suitable for real-time DRM.
Tags
Links
- Source: https://arxiv.org/abs/2607.09805v1
- Canonical: https://arxiv.org/abs/2607.09805v1
Trouble viewing inline? Open PDF directly β
Full Text
59,136 characters extracted from source content.
Expand or collapse full text
A Unified Model for Highly Accurate ECG-Free Dynamic Coronary Roadmapping Using Spatio-Temporal Transformers β Saahil Islam a,c,β,1 , Sebastian Piat b , Venkatesh N. Murthy b , Serkan Cimen b , Puneet Sharma b , Andreas Maier a and Florin C. Ghesu c a Pattern Recognition Lab, Friedrich Alexander University, Erlangen, Germany b Digital Technology and Innovation, Siemens Healthineers, Princeton, USA c Digital Technology and Innovation, Siemens Healthineers, Erlangen, Germany A R T I C L E I N F O Keywords: Dynamic Coronary Roadmapping ECG-free Motion Compensation Self-supervised spatio-temporal encoder Cardiac Phase Matching Unified Model Image-Guided Intervention A B S T R A C T Percutaneous Coronary Intervention (PCI) is a minimally invasive procedure designed to restore coronary blood flow obstructed by atherosclerotic plaque. During PCI, repeated injections of iodine- based contrast agents are required to visualize the coronary arteries and guide interventional devices. However, frequent injections increase radiation exposure and the risk of contrast-induced nephropathy, with acute kidney injury reported in up to 30% of patients with renal impairment. Dynamic Coronary Roadmapping (DRM) has emerged as an effective strategy to mitigate these risks by overlaying a precomputed angiographic vessel map onto live fluoroscopy and continuously updating it during the procedure. Accurate DRM depends on precise cardiac phase matching between angiography and fluoroscopy, as well as reliable catheter tip tracking for motion compensation. These tasks remain challenging in electrocardiogram-free settings or when manual annotations are limited. In this work, we introduce a unified DRM framework that simultaneously performs cardiac phase matching and catheter tip tracking, enabling accurate and real-time guidance. Our approach employs a large-scale spatio-temporal encoder pretrained on 16 million frames to model cardiac motion dynamics. To our knowledge, it is the first time this kind of pretraining has been used for motion compensation in DRM. Furthermore, we propose auxiliary tasks based on ECG R-peak detection and catheter tip tracking, which stabilize optimization and eliminate the need for extensive catheter mask annotations. Finally, a majority-voting postprocessing strategy aggregates temporal predictions to enhance robustness and provides a confidence score that correlates with the phase-matching error. Comprehensive evaluations on clinical X-ray datasets show that the proposed method achieves state-of-the-art performance, producing low temporal misalignment and consistent phase-matching accuracy suitable for real-time DRM applications. 1. Introduction Percutaneous Coronary Intervention (PCI) is a mini- mally invasive cardiology procedure performed to restore blood flow in coronary arteries obstructed by atherosclerotic plaque deposits, a condition commonly referred to as steno- sis (Khan and Ludman (2022)). Under X-ray angiographic guidance, a catheter is navigated into the ostium of the coro- nary artery. Through this guiding catheter, a balloon catheter equipped with a stent is advanced over a guidewire to the stenotic site. Inflation of the balloon expands the stent, which is subsequently deployed to maintain vessel patency and prevent restenosis. To visualize the coronary vasculature and guide the catheter and guidewire to the stenotic region, an iodine-based contrast agent is injected during the procedure. However, opacification of coronary arteries only lasts for a short period of time and the contrast agent is injected multiple times for navigation of devices. Moreover, cardiol- ogists have to mentally reconstruct the position of vessels and stenosis based on previous angiograms. This practice poses risks such as radiation exposure and contrast-induced β β Saahil Islam saahil.islam@fau.de (S. Islam) (S. Islam) ORCID(s): 1 nephropathy. In patients with pre-existing renal impairment, the incidence of acute kidney injury can be as high as 30% (Piayda et al. (2018); Tehrani et al. (2013)). 1.1. Dynamic Coronary Roadmapping Dynamic Coronary Roadmapping (DRM) has emerged as a promising technique to mitigate these risks by reducing both the radiation dose and the use of contrast agents (Elion (1989); Zhu et al. (2010); Manhart et al. (2011); Kim et al. (2018); Piayda et al. (2018); Ma et al. (2020); Liu et al. (2024)). DRM leverages a detailed coronary map precom- puted from angiography, which is superimposed on live fluoroscopy. The map is continuously updated in real time with each acquired fluoroscopic frame during PCI, providing cardiologists with immediate visual feedback throughout the intervention. This facilitates accurate guidewire naviga- tion to the appropriate coronary branch and ensures precise placement of the stent at the stenotic site, while simulta- neously reducing patient exposure to radiation and contrast agents. Developing a DRM system necessitates the precise su- perimposition of a precomputed coronary artery map onto live fluoroscopic images. This process is particularly chal- lenging because fluoroscopy provides only limited vessel in- formation, complicating the compensation of motion caused by both cardiac activity and respiration. During the cardiac : Preprint submitted to ElsevierPage 1 of 13 arXiv:2607.09805v1 [eess.IV] 9 Jul 2026 Figure 1: Example of dynamic coronary roadmap. The image on the left shows a fluoroscopy frame and the image on the right shows the overlayed vessel roadmap on the fluoroscopy frame. An accurate overlayed roadmap typically has the guidewire under the correct vessel branch. cycle, the coronary vessel tree undergoes dynamic deforma- tion due to myocardial contraction and relaxation, requiring the selection of an anatomically appropriate vessel config- uration for accurate mapping. In comparison, respiratory motion can often be approximated as translational; however, it still introduces additional displacement on top of cardiac- induced motion, thereby further complicating the alignment process. The earliest DRM system, proposed by Elion (1989), generated roadmaps through digital subtraction of contrast- enhanced and mask sequences spanning a full cardiac cycle. These roadmaps were later synchronized with live fluo- roscopy by aligning R-waves in the ECG. Although this method accounted for cardiac motion, it did not address res- piratory motion. Later works, such as Zhu et al. (2010) and Manhart et al. (2011), introduced image-based respiratory compensation methods. These approaches assumed an affine respiratory motion model in ECG-gated fluoroscopic frames and estimated displacement from soft-tissue motion while handling static structures separately. Their effectiveness, however, depends on sufficient visibility of relevant tissue in the field of view and is limited to cardiac-gated frames. Kim et al. (2018) proposed creating binary vessel masks from at least one cardiac cycle of angiographic images to serve as roadmaps. Cardiac motion was compensated by temporally aligning angiographic roadmaps with fluo- roscopy via cross-correlation of ECG signals, while respira- tory motion was corrected by aligning the guidewire center- line in fluoroscopy with vessel contours from angiography. Although this method demonstrated feasibility in phantom studies, it lacked quantitative evaluation of spatiotemporal accuracy and relied heavily on robust vessel and guidewire extraction, which remains challenging in X-ray images. In contrast to direct roadmapping, model-based ap- proaches predict motion in fluoroscopic frames using sur- rogate signals such as ECG-derived cardiac motion or respi- ratory motion from diaphragm tracking. Some studies model both cardiac and respiratory motion (Shechter et al. (2005); Timinger et al. (2005); Faranesh et al. (2013); Fischer et al. (2017)), while others focus solely on respiratory motion using cardiac-gated images (Schneider et al. (2010); King et al. (2009); Peressutti et al. (2013)). A major limitation of these methods is their patient-specific nature, requiring model retraining for each subject. Moreover, when surrogate values at inference fall outside the training range (e.g., due to abnormal motion), extrapolation is needed, which may compromise motion compensation accuracy. Recent approaches employ deep learning for respiratory motion compensation, often by tracking the tip of the guid- ing catheter within a Bayesian filtering framework (Ma et al. (2020)). However, this method is restricted to fluoroscopy and relies on manual detection in angiographic sequences, where catheter tracking is particularly challenging due to oc- clusions from contrast injection. To address this, Demoustier et al. (2023) proposed an optical flow-based solution, achiev- ing superior tracking results compared to existing tracking methods in natural imaging (Yan et al. (2021); Cui et al. (2022); Li et al. (2018)). Islam et al. (2024b,a) leveraged pretrained transformers for improved robustness. Although these techniques demonstrate promising results for catheter tracking in both angiography and fluoroscopy, they have not been evaluated in the context of DRM. Furthermore, they require a manual initialization to achieve high robustness limiting full automation. Despite advances in respiratory motion correction, car- diac motion compensation in these methods still depends on ECG signals, necessitating ECG acquisition during both angiography and live fluoroscopy. Liu et al. (2024) show that cardiac phase matching can be achieved using image- based features obtained from deep-learning based feature extractors for angiography and fluoroscopy alignment. How- ever, the network struggles to learn and converge unless a large dataset is used containing catheter mask annotations. This is due to the methodological limitation of encoder typi- cally focusing on first image feature extraction and temporal fusion on highly downsampled image features. Hence, the motion cue is introduced into the model from the motion of the catheter masks. Obtaining such annotations for both angiography and fluoroscopy is highly labor-intensive and costly. Moreover, their framework relies on separate models for catheter tip tracking and phase matching, which limits its efficiency for real-time DRM. 1.2. Spatio-temporal features and feature matching Different strategies have been developed to learn spatio- temporal representations from both labeled and unlabeled data, demonstrating strong effectiveness in tasks where tem- poral dynamics play a important role. Most existing methods rely on self-supervised learning from large-scale unlabeled datasets and have been predominantly applied to action recognition in natural images (Tong et al. (2022); Wang et al. (2023); Bardes et al. (2024); Assran et al. (2025)). How- ever, while action recognition models capture long-range temporal dependencies, they often overlook subtle inter- frame variations. In contrast, approaches such as SAM2 : Preprint submitted to ElsevierPage 2 of 13 (Ravi et al. (2024)) achieve fine-grained frame-level tracking and segmentation but require extensive labeled data for training. On the other hand, approaches such as dino (Caron et al. (2021); Oquab et al. (2023)) are pretrained only on images to learn spatial features and performs well on video segmentation tasks, it lacks consistency over the frames and fails during object occlusions. FIMAE (Islam et al. (2024b)), trained in a self-supervised manner on unlabeled angiography and fluoroscopy datasets, addresses these lim- itations by learning representations sensitive to fine inter- frame correspondences. Furthermore, HiFT (Islam et al. (2024a)) enhances the learned feature space by incorporating auxiliary supervision from vessel segmentation tasks on annotated data. In parallel, contrastive and metric learning approaches aim to structure the embedding space by bringing similar (positive) examples closer while pushing dissimilar (neg- ative) examples apart (Schroff et al. (2015); Khosla et al. (2020); Chen et al. (2020); He et al. (2020); Caron et al. (2020)). The classical contrastive loss (Hadsell et al. (2006)) enforces this separation by minimizing the distance be- tween positive pairs and ensuring negatives remain beyond a defined margin. While initially introduced for supervised learning (Khosla et al. (2020); Ghojogh et al. (2020)), con- trastive learning has since become central to self-supervised representation learning. Triplet loss extends this concept by operating on triplets (ν,ν,ν)βan anchor ν, a positive ν (similar or same class), and a negative ν (dissimilar or different class)βand constraining the embedding space such that the anchor-positive distance is at least a margin smaller than the anchor-negative distance (Schroff et al. (2015); Zhao et al. (2019)). This formulation was recently employed by (Liu et al. (2024)) for cardiac phase matching in DRM, where frames of the same cardiac phase are pulled closer in the embedding space, while those of different phases are pushed apart. 1.3. Contributions Despite recent progress, existing approaches for DRM face several major limitations. Many methods rely on ECG signals for cardiac phase alignment, which are not always available or reliable in clinical practice. Image-based ap- proaches, on the other hand, often require large amounts of dense annotations, such as catheter masks, to provide sufficient motion cues for learning. Furthermore, prior works typically address cardiac phase matching and catheter tip tracking using separate models, leading to increased com- putational complexity and limiting real-time applicability. In addition, differences in image quality between angiography and fluoroscopy, as well as motion inconsistencies caused by respiration, patient movement, or device manipulation, further complicate robust alignment. These challenges high- light the need for a unified and annotation-efficient frame- work that can generalize across diverse imaging conditions. To address these challenges, we propose a unified frame- work for DRM that simultaneously addresses cardiac phase matching and catheter tip tracking for accurate and real- time guidance. Our main contributions are summarized as follows: β’ We use a pretrained spatio-temporal encoder trained on 16 million frames for cardiac phase matching. To the best of our knowledge, this is the first work to leverage spatio-temporal pretrained features for car- diac motion compensation, achieving superior phase- matching performance compared to existing methods. β’ We propose a novel pipeline that employs ECG R- peak detection and catheter tip tracking as auxiliary tasks, enabling stable training and convergence of phase matching without requiring additional catheter mask annotations. β’ We design a unified model that jointly performs phase matching and catheter tip tracking, resulting in a fast inference suitable for real-time DRM. β’ We develop a majority-voting post-processing strat- egy that incorporates predictions from previous flu- oroscopy frames, enhancing robustness and produc- ing a reliable confidence score. Through quantitative analysis, we demonstrate how this score correlates with the phase-matching error, making it valuable in interventional settings. β’ We provide both visual and quantitative evaluations on clinical angiography and fluoroscopy sequences, demonstrating low errors for DRM and accurate catheter tip tracking. 2. Methods In this section, we explain our proposed methodology for training and inference. 2.1. Training The overview of our training pipeline is shown in Fig. 2. Given a live fluoroscopy frame νΌ νΉ , our objective is to iden- tify the corresponding frame in the angiography sequence, ν ν΄ = [νΌ ν΄ 1 ,νΌ ν΄ 2 ,...,νΌ ν΄ ν ] that matches its cardiac phase. Since a single fluoroscopy frame lacks sufficient information to capture cardiac motion, we instead consider a fluoroscopy sequence ν νΉ = [νΌ νΉ 1 ,νΌ νΉ 2 ,...,νΌ νΉ ν ], assuming the last frame is the one we are concerned with matching and the other (past) frames are responsible for the motion context. We employ a pretrained spatio-temporal encoder ν ν , trained in a self-supervised manner on large-scale angiogra- phy and fluoroscopy datasets (Islam et al. (2024b)), to extract features from both ν ν΄ and ν νΉ . Shared weights are used for the two modalities: ν ν΄ = ν ν (ν ν΄ ), ν νΉ = ν ν (ν νΉ ). Note that since the pretrained network was trained with 10 frames, we limit νβ€ 10. To reduce memory con- sumption during training, we alternately stop gradient flow : Preprint submitted to ElsevierPage 3 of 13 Figure 2: Overview of our method: We train our model in 3 stages; (a) First a spatio-temporal transformer encoder is pretrained on a large unlabeled dataset of angiography and fluoroscopy, (b) The pretrained encoder is used to train a model to track the catheter tip and the trained model is then used to generate pseudo catheter tip labels for a labeled ECG dataset. (c) Finally, a unified DRM model is trained with the pretrained backbone jointly on the ECG labels and pseudo catheter tip labels. for either angiography or fluoroscopy feature extraction. The spatio-temporal features ν ν΄ ,ν νΉ are then processed by a lightweight CNN ν ν (4 convolutional layers with max pooling followed by global average pooling), producing 768- dimensional frame-level embeddings: Μν ν΄ = ν ν (ν ν΄ ), Μν νΉ = ν ν (ν νΉ ). Finally, embeddings from all frames are concatenated along the temporal dimension and combined with learnable positional encodings ν to distinguish between angiography and fluoroscopy frames and to provide temporal context for each frame. These positional encodings are distinct from those used in the pretrained encoder and are learned specif- ically for the joint sequence. In particular, separate sets of positional embeddings are maintained for angiography and fluoroscopy, with a maximum capacity of 10 embeddings per modality. For a given input, only the required number of embeddings is selected (e.g., 5 out of 10 for a 5-frame sequence), allowing the model to encode both the temporal order and the modality-specific identity of each frame. We apply multi-headed self-attention layers on joint frame em- beddings, Μν: Μν = Concat([Μν ν΄ , Μν νΉ ]) + ν. This step is followed by multi-headed self-attention (MHA) layers on the joint embeddings for the model to learn the motion relationship between the angiography and fluoroscopy frame embeddings as well as for the model to learn a unique embedding for each cardiac phase in order to match between the sequences. For each attention head β, queries, keys, and values are computed as ν β = Μνν ν β , νΎ β = Μνν νΎ β , ν β = Μνν ν β , where ν ν β ,ν νΎ β ,ν ν β ββ νΓν β are learnable projection matrices,ν is the embedding dimension of the input features, and ν β is the embedding dimension of a single attention head. The attention output for head β is defined as Attention(ν β ,νΎ β ,ν β ) = Softmax ( ν β νΎ β€ β β ν β ) ν β . The outputs of all ν» heads are concatenated and linearly projected to obtain the final representation, ν = MHA(Μν). Although cross-attention is often used in matching tasks, our early experiments showed that self-attention performed better than cross-attention. As noted in Liu et al. (2024), the loss landscape for cardiac phase matching is unsmooth, since frame-to-frame : Preprint submitted to ElsevierPage 4 of 13 motion is minimal and a full cardiac cycle may not be visible when the encoder is limited to10 frames. To aid convergence and avoid local minima, we introduce an auxiliary task in which the network predicts the R-peaks of the ECG signal. For this, ν is passed through a linear projection followed by a sigmoid activation, yielding a scalar output Μν¦ β [0,1] indicating whether a frame corresponds to an R-peak (1) or not (0). Since R-peaks are sparse (often none or at most a few in a sequence), we use an adaptive weighted binary cross- entropy loss that balances positive and negative classes dynamically: ξΈ ν ν = β 1 ν ν β ν=1 ν€ ν [ ν¦ ν logν(ν§ ν )+(1βν¦ ν )log(1βν(ν§ ν )) ] , where ν§ ν are the logits, ν¦ ν β 0,1 are the ground-truth labels, ν(β ) is the sigmoid function, and ν€ ν is the adaptive weight. Let ν 0 and ν 1 denote the number of negative and posi- tive samples in the batch, respectively, with ν = ν 0 + ν 1 . The adaptive weights are defined as ν€ ν = β§ βͺ β¨ βͺ β© ν 2ν 0 , ν¦ ν = 0, ν 2ν 1 , ν¦ ν = 1, ensuring equal contribution from positive and negative sam- ples regardless of class imbalance. For the cardiac phase matching, we employ a triplet loss on the attended embeddings ν. Given an anchor embedding ν νΉ ν from a fluoroscopy frame, a positive embedding ν ν΄ ν from an angiography frame of the same cardiac phase, and a negative embeddingν ν΄ ν from a different phase, the triplet loss is defined as: ξΈ ν ν = max ( 0, ν(ν νΉ ν ,ν ν΄ ν ) β ν(ν νΉ ν ,ν ν΄ ν ) + νΌ ) , where ν(β ,β ) is cosine similarity and νΌ is the margin with a value of 0.8. Finally, we introduce catheter tip tracking as an auxiliary task, making the network unified for all the requirements of dynamic coronary roadmapping. This task provides addi- tional cues for phase matching by enforcing the network to learn catheter motion. Since only a few frames per sequence are annotated (3β5 on average), we employ Historical Fea- ture Guided Tracker (HiFT) Islam et al. (2024a) to generate pseudo-labels for all frames. For both angiography and fluoroscopy sequences, the spatio-temporal features ν ν΄ and ν νΉ are passed through a CNN-based upsampling head ν Ξ¦ . The outputs are concate- nated and projected to obtain the predicted heatmap Μ ν». We supervise Μ ν» with pseudo-label heatmapsν» using Dice loss: ξΈ νΆ ν‘ = 1 β 2 β ν Μ ν» ν ν» ν β ν Μ ν» 2 ν + β ν ν» 2 ν . Figure 3: Proposed inference strategy with majority voting post-processing The final loss for training is given by: ξΈ = ν ν ν ξΈ ν ν + ν ν ν ξΈ ν ν + ν νΆ ν‘ ξΈ νΆ ν‘ Where ν ν ν , ν ν ν and ν νΆ ν‘ are loss weights chosen as 0.5, 1.0 and 0.5 respectively. 2.2. Inference The inference procedure is divided into an offline phase and an online phase. 2.2.1. Offline phase In the offline stage, we first apply a contrast detection model together with R-peak detection to identify one car- diac cycle (peak-to-peak) of highly contrasted frames. To construct the angiography sequence, an offset of two frames is applied. Since the network is limited to processing at most 10 frames, we use a temporal window of 10 frames with a stride of 4 as the angiography sequence input. For each input sequence, all angiography frame embeddings Μν ν΄ are precomputed and stored. A Res-UNet trained for vessel segmentation is used to extract vessel roadmaps, which are also stored for later use. 2.2.2. Online phase During the online phase, real-time inference is per- formed on incoming fluoroscopy frames. We employ a sliding temporal window of 10 frames with a stride of 1, where the last frame in the sequence is treated as the current live frame. For each fluoroscopy sequence, frame embed- dings Μν νΉ are computed and concatenated with every stored angiography embedding Μν ν΄ . The embeddings are added with their respective positional encodings ν and passed MHA layers, consistent with the training setup. Cosine similarity is then computed between fluoroscopy and angiography embeddings across all stored sets. To determine the final matching index, we employ a majority-voting post-processing strategy. Specifically, for each fluoroscopy frame embedding, the angiography frame index with maximum cosine similarity is retrieved from every stored set. The matched indices are interpolated to estimate the most likely corresponding index for the last fluoroscopy frame in the sequence. The final index for the : Preprint submitted to ElsevierPage 5 of 13 Table 1 Comparison of different strategies and models for automatic image-based phase matching for dynamic coronary roadmapping. Angio-Fluoro refers to our internal large unlabeled dataset of angiography and fluoroscopy dataset and LVD-142M (Oquab et al. (2023)) refers to their curated dataset consisting of 142 million natural images, used to train Dinov2. StrategyPretrainedFeature Extraction Distance error (Frames) Distance error (Time in ms) Distance error (Percentage) mean Β± std median maxmean Β± stdmean Β± std AIT (Liu et al. (2024)) ImageNetCNN-Transformer 1.20 Β± 0.731.05.0 98.51 Β± 68.357.67 Β± 4.25 Ours ImageNetCNN-Transformer 1.02 Β± 0.83 0.77 3.69 84.85 Β± 76.786.65 Β± 5.49 LVD-142M Dinov2-Transformer 1.08 Β± 1.01 0.91 3.40 89.38 Β± 85.148.48 Β± 7.0 Angio-Fluoro Dinov2-Transformer 0.94 Β± 0.68 0.81 3.45 77.8 Β± 63.816.19 Β± 4.83 NoneSpatio-Temporal2.78 Β± 1.252.77.0 207.56 Β± 95.67 16.70 Β± 7.8 Angio-Fluoro Spatio-Temporal 0.64 Β± 0.66 0.55 3.15 56.31 Β± 62.43 4.38 Β± 4.51 Figure 4: Percentile distribution of phase-matching errors and comparison of different backbones with our proposed strategy. STPM refers to spatio-temporal phase matching (ours). current live frame is then obtained via majority voting over all interpolated indices. An illustration of this majority- voting post-processing is shown in Fig. 3. 3. Experiments 3.1. Dataset and training details Our phase-matching dataset comprises 1,434 training pairs of angiography and fluoroscopy sequences, totaling 116,256 frames. The validation set includes 204 sequences with 16,223 frames, while the test set contains 79 pairs with 6,676 frames, of which 1,404 are fluoroscopy frames. Each sequence is associated with synchronized ECG data, from which R-peaks are identified. To represent cardiac phase continuity, we assign a normalized phase value of 0 to R- peak frames, with intermediate frames linearly interpolated between 0 and 1. When the number of frames between con- secutive R-peaks differs across angiography and fluoroscopy sequences, the nearest phase value between the two is used to establish correspondence. Additionally, we use a catheter tracking dataset compris- ing 2,314 training sequences with a total of 198,993 frames, among which 44,957 frames include catheter tip annotations. A subset of this dataset contains catheter mask annotations (β 10% of the total number of frames). Since most sequences have sparse, non-consecutive annotations, we employ the pretrained HiFT model Islam et al. (2024a) to generate pseudo labels for unannotated frames. The pseudo catheter tip coordinates are converted to gaussian heatmap with a standard deviation of β 2.5m serving as ground truth for catheter tip detection. This enables multi-task training for both phase matching and catheter tip tracking across all frames. The dataset containing paired ecg data is a subset of this larger dataset. For pretraining the spatio-temporal encoder, we use an internal unlabeled coronary X-ray dataset similar to FIMAE- SC Islam et al. (2024a). This dataset comprises 241,362 se- quences collected from 21,589 patients, totaling 16,342,992 frames across both angiography and fluoroscopy modalities. Supplementary cues for FIMAE-SC pretraining are derived from a ResUNet trained on 3,300 angiography sequences (with 91 for testing), where coronary arteries were annotated with centerline points and approximate vessel radii for five highly contrasted frames to generate target vesselness maps. All frames are resized to 512Γ512 and augmented with random affine transformations, including translation in the range (β0.15,0.15), rotation between (β10 β¦ ,10 β¦ ), scaling between (0.8,1.2), and random horizontal and vertical flips. The model is trained for 500 epochs using a cosine annealing scheduler with a linear warmup. The learning rate is ini- tialized at 4eβ6, linearly increased for the first 30 epochs to 8eβ6, and then gradually decayed following a cosine schedule to 1eβ7. The final model is chosen based on its phase-matching performance on the validation set. This model is then fine- tuned using catheter-tip labels generated by HiFT, during which only the CNN decoder for tip detection is updated while the remaining components are kept frozen. In prelimi- nary experiments, training with the weighted loss introduced a substantial number of false positives, some of which over- lapped with the true positives in the catheter-tip heatmaps, making it harder to remove with any kind of post-processing. These spurious responses were largely eliminated after the dedicated decoder fine-tuning stage. : Preprint submitted to ElsevierPage 6 of 13 3.2. Results for phase matching We conduct a comprehensive evaluation of our method against state-of-the-art approaches, analyzing the effective- ness of the proposed strategy and further examining whether the modelβs confidence scores exhibit a correlation with phase-matching errors in DRM. Consistent with the evalua- tion protocol of Liu et al. (2024), we report distance errors in terms of both frames and frame percentage. To enhance in- terpretability and provide a more tangible understanding of temporal misalignment, we additionally express the distance errors in milliseconds. The complete results are summarized in Table 1. Overall, our approach attains state-of-the-art perfor- mance on the test dataset. The results demonstrate that the proposed strategy consistently surpasses AIT when employ- ing the same backbone architecture (CNN-Transformer). In particular, our method achieves an average distance error of 56.31 millisecondsβcorresponding to 4.38% of the cardiac cycleβwith a low standard deviation of 0.66, indicating highly stable and reliable phase-matching perfor- mance across diverse sequences. It is worth noting that only 10% of the training dataset includes catheter mask anno- tations, making a direct comparison with Liu et al. (2024) less straightforward. Despite this, the findings underscore that integrating auxiliary tasks within the model provides sufficient supervisory signals, substantially mitigating the reliance on extensive manual catheter mask annotations. Moreover, the results highlight the significance of jointly modeling spatial and temporal dynamics for effective phase matching. A unified spatio-temporal feature extractor con- sistently outperforms designs that decouple spatial encoding from temporal fusion. Finally, domain-specific pretraining on angiography and fluoroscopy data yields substantially superior performance compared to pretraining on unrelated datasets. This trend is evident when comparing DINOv2 pretrained on LVD-142M Oquab et al. (2023) with the same model pretrained on our internal large-scale dataset of angiography and fluoroscopy. Notably, DINOv2 pretrained on LVD-142M performs worse than the CNN-Transformer baseline, which can be attributed to the domain mismatch and the limited size of the downstream dataset. These find- ings further reinforce the importance of domain-adapted representation learning for achieving robust performance in DRM. Figure 4 illustrates a percentile-wise comparison of phase-matching errors, measured in frame distance, across different models. The proposed Spatio-Temporal Phase Match- ing (STPM) approach consistently exhibits superior per- formance across nearly the entire distribution of test sam- ples, achieving lower distance errors than both the CNN and DINOv2 baselines. These results indicate that STPM generalizes more robustly to diverse cardiac motion patterns and varying imaging conditions. A notable observation is that STPM maintains a distance error below one frame for approximately 82% of all test cases, underscoring its reliability and fine-grained temporal alignment capability. In contrast, DINOv2 and CNN reach Figure 5: Distribution of error in each bins of confidence score obtained from the voting frequency of our model. Note that the distance error depicted here is in the frame level and minimum confidence score obtained in the test set was 0.16. this error threshold considerably earlier, reflecting a sharper degradation in more challenging frames. Moreover, for about 22% of the samples, the error is exactly zero, signifying perfect frame-level phase alignment between angiography and fluoroscopy sequences. Beyond the 80th percentile, all models exhibit a sharper increase in error, which can be attributed to difficult imaging conditions such as extreme view angles that obscure cardiac motion or cases where the vessel tree in the angiography frame is partially cropped. Nevertheless, even within these challenging regions, STPM maintains the lowest error margin, demonstrating its ability to capture subtle inter-frame dynamics that other models fail to represent effectively. 3.2.1. Confidence score and distance error The majority-voting post-processing strategy employed during inference additionally enables the estimation of a confidence score for each prediction. This confidence score is defined as the ratio between the number of candidate predictions that agree on the same frame while being the majority and the total number of candidates considered for that prediction. In the context of image-guided therapy, the availability of an interpretable confidence measure is highly beneficial, as it allows surgeons or technicians to assess the reliability of the systemβs output and appropriately weigh it against their own expertise. For such a confidence score to be meaningful, it is desirable that it exhibits an inverse relationship with the prediction error. The relationship between the confidence score and the distance error is illustrated in Fig. 5 using a violin plot. Here, distance errors are evaluated at the frame level rather than being averaged over an entire sequence. The plot depicts the error distributions for confidence score bins of width 0.1. For confidence scores below 0.2, the error distribution is approximately uniform over the range of 1 to 8 frames, indicating that predictions with very low confidence are : Preprint submitted to ElsevierPage 7 of 13 Figure 6: Distance error as a function of coverage (% of data remaining for confidence score thresholds of 0.1). (Top) Mean distance error. (Bottom) Maximum distance error. Note that the distance error depicted here is in the frame level. equally likely to be either close to the correct cardiac phase or to exhibit large errors. As the confidence score increases, a consistent reduction is observed in the mean, median, 5β 95 percentile range, and maximum error. Notably, for confi- dence scores between 0.4 and 0.8, both the mean and median errors fall below one frame; however, isolated outliers with errors as large as five frames are still present. Beyond the overall reduction in error statistics, Fig. 5 also reveals how the variability of the predictions evolves with the confidence score. For low-confidence predictions, the error distributions are broad and exhibit substantial spread, indicating highly inconsistent behavior across samples. As the confidence score increases, these distributions become progressively narrower and increasingly concentrated around zero error. This suggests that the confidence score reflects not only the expected accuracy of a prediction but also its robustness, with higher confidence predictions exhibiting markedly re- duced variability. An interesting transition can be observed for confidence scores above approximately 0.8. In this regime, the error distributions collapse to near-zero values, with the median error reaching zero frames and the upper percentiles remain- ing tightly bounded. This indicates that the model is able to identify a subset of predictions for which the inferred cardiac phase is consistently accurate. Such behavior is particularly Figure 7: Effect of adding different modules in our spatio- temporal phase matching model relevant in practical settings, where it may be preferable to rely only on predictions that the model itself deems reliable. To further characterize predictive uncertainty, we ana- lyze the relationship between distance error and data cov- erage as a function of the confidence score, as shown in Fig. 6. Coverage is defined as the percentage of test samples retained after discarding predictions with confidence scores below a specified threshold. Starting from a threshold of 0.1, the threshold is incremented in steps of 0.1. A trend similar to the error distribution statistics over the confidence score is reflected in the coverage-based analysis. As predictions with lower confidence scores are progressively discarded, the mean distance error decreases in a near-linear manner, followed by a sharper drop when only the most confident predictions are retained. In contrast, the maximum error remains relatively high over a wide range of coverage values, indicating the presence of occasional large errors even when the average performance is reasonable. These outliers are largely removed only when the coverage falls below approx- imately 40%, suggesting that confidence-based filtering is particularly effective at mitigating worst-case errors rather than merely improving average accuracy. Overall, these observations indicate that the proposed confidence score is well aligned with the underlying predic- tion error. It provides a meaningful measure of uncertainty that can be used to balance accuracy and coverage, and offers a practical mechanism for identifying reliable predictions in live dynamic coronary roadmapping where erroneous outputs may have significant consequences. 3.2.2. Ablation for phase matching Figure 7 presents the performance of our model under different architectural configurations, illustrating the contri- bution of each module within the pipeline. We observe that incorporating a simple R-peak detection auxiliary task en- hances the modelβs performance by 62.7%. As noted in Liu et al. (2024), the loss landscape associated with triplet loss in cardiac phase matching is inherently unsmooth. While Liu et al. (2024) address this issue by introducing auxiliary catheter mask inputs, our findings demonstrate that even : Preprint submitted to ElsevierPage 8 of 13 Table 2 Comparison of the performance of different models for catheter tip tracking. Note that the difference in performance between HiFT and the unified model under automatic initialization is not statistically significant (ν > 0.1). Init typeModeldistance error (m) Mean Β± std Median Max Manual ConTrack1.63 Β± 1.701.08 13.32 SimST1.44 Β± 1.351.02 10.23 HiFT1.21 Β± 0.68 1.044.04 Auto AIT1.91 Β± 1.751.32 15.12 ConTrack2.87 Β± 2.362.29 17.26 SimST2.24 Β± 2.191.61 18.66 HiFT1.45 Β± 1.30 1.059.29 Unified Model (Ours) 1.58 Β± 1.421.11 14.60 a lightweight auxiliary task, such as ECG R-peak detec- tion, is sufficient to stabilize training and facilitate conver- genceβthereby eliminating the need for extensive manual catheter mask annotations. Moreover, integrating the proposed majority-voting post- processing leads to a further and substantial improvement in performance. This gain is expected, as the aggregation of predictions across multiple candidates yields more robust and accurate outcomes. Finally, we observe that incorpo- rating catheter tip tracking within the same model not only unifies the DRM framework but also yields additional gains in phase-matching accuracy, underscoring the mutual benefit of multi-task learning. 3.3. Results for catheter tip tracking from unified model Table 2 presents a quantitative comparison of catheter tip tracking performance across different models and ini- tialization strategies. Most existing approaches rely on man- ual initialization of the catheter landmarks, which typically leads to improved accuracy but limits their applicability in a fully automated DRM workflow. As expected, methods evaluated under manual initialization consistently achieve lower mean and median errors compared to those operating with automatic initialization. Among the manually initialized models, HiFT achieves the best overall performance, with substantially reduced maximum error compared to other approaches, indicating improved robustness in challenging cases. Under automatic initialization, performance degradation is observed across all methods, highlighting the difficulty of the task in the absence of manual guidance. In this setting, HiFT again attains the lowest errors, while the unified model achieves comparable performance with only a modest increase in mean error. It is worth noting that for ConTrack, SimST, and HiFT, a separate model trained with the same encoder as the tracker is used to detect the catheter tip in the first frame, serving as the initialization. In contrast, AIT and the proposed unified model perform both initialization and tracking within a single model, reducing the overall training Figure 8: Comparison of model speed and phase matching error for different backbones in different setups time and eliminating the need for an additional initialization network. The unified model exhibits slightly lower accuracy than the automatically initialized HiFT model, which is expected given that it is trained using pseudo-labels generated by HiFT due to the lack of frame-level manual annotations across the full training set. Notably, the difference in per- formance between the two models is not statistically signif- icant (ν > 0.1), suggesting that the observed gap is small. Moreover, the unified model maintains competitive median error while reducing reliance on model-specific initialization assumptions, making it better aligned with the requirements of a scalable and fully automated DRM pipeline. 3.4. Phase matching error and inference speed The advantage that arises from a unified model is greater speed than having separate models for different tasks. Com- parison of distance error and speed of different models under a two stage setup and our unified setup is shown in Fig. 8. The two stage setup performs catheter tip tracking and cardiac phase matching sequentially using separate models, whereas the unified setup performs both tasks jointly within a single model. As shown, unifying the pipeline consistently im- proves computational efficiency across all backbones, with the most pronounced gains observed for STPM. While the two-stage STPM already achieves the lowest distance error among the evaluated methods, its unified counterpart fur- ther reduces inference latency without sacrificing accuracy, demonstrating the effectiveness of integrating spatial and temporal reasoning within a single model. It is also worth noting that although CNN-based models are typically faster than transformer-based approaches, their advantage is substantially diminished in the conventional two-stage dynamic coronary roadmapping pipeline, where overall throughput is reduced to approximately 20 fps due to sequential processing. In contrast, the single-stage unified formulation removes this bottleneck and allows the CNN- based model to operate at up to 84 fps, better reflecting the inherent efficiency of convolutional architectures. : Preprint submitted to ElsevierPage 9 of 13 Figure 9: Qualitative result of our model on a few consecutive frames of an example angiography-fluoroscopy pair. The zoomed version shows the overlap of the guidewire with the predicted roadmap. The structural similarity of the shape of the guidewire and the roadmap signifies accurate cardiac phase matching. The numbers on top of frame depicts the frame number. Unified STPM (ours) occupies a favorable region of the speedβaccuracy trade-off, achieving the lowest error over- all while operating at substantially higher throughput than its two-stage formulation and other unified baselines. This result highlights the complementary nature of our contri- butions: STPM provides a robust spatio-temporal matching mechanism, while the unified design removes redundant computation and enables efficient end-to-end inference. To- gether, these design choices yield a model that is both accu- rate and practical for real-time distance estimation, aligning with the objectives outlined throughout this work. : Preprint submitted to ElsevierPage 10 of 13 Figure 10: Example of DRM results on frames randomly selected from some challenging cases. (a) to (d) refers to 4 different angiography-fluoroscopy sequence pairs. 3.5. Qualitative results Figure 9 presents qualitative results of the proposed method on consecutive frames from an angiography flu- oroscopy pair. The overlaid roadmap closely follows the guidewire trajectory across frames, indicating accurate car- diac phase matching. In particular, the structural alignment between the guidewire in fluoroscopy and the vessel tree from angiography remains consistent despite cardiac mo- tion. The temporal consistency across frames suggests that the model effectively captures the underlying cardiac dy- namics rather than relying on frame-wise appearance cues. The zoomed-in regions further highlight that fine vessel structures are well aligned with the guidewire, demonstrat- ing that the model preserves spatial details while main- taining temporal coherence. This behavior is essential for reliable dynamic coronary roadmapping, where even small misalignments can affect clinical usability. Figure 10 shows qualitative results on more challenging cases. In Fig. 10(a) and Fig. 10(b), the guidewire or stent is fully occluded by the overlaid vessel branches. This is expected in accurate dynamic coronary roadmapping, where the projected vessel structure should coincide with the un- derlying device. The consistent overlap between the device and the vessel tree indicates correct cardiac phase matching and precise spatial alignment between angiography and flu- oroscopy. In Fig. 10(c) and Fig. 10(d), the fluoroscopy frames exhibit lower spatial resolution compared to the angiography reference, resulting in noticeable discrepancies in spatial appearance between the two modalities. In addition, fluo- roscopy provides limited structural information, primarily from background anatomy and interventional devices such as the catheter or guidewire. This reduced and indirect set of cues makes phase matching more challenging. Despite this, the model achieves good alignment by leveraging spatio- temporal cues and capturing the underlying cardiac motion dynamics, leading to accurate phase matching even under such appearance variations. Fig. 10(d) also illustrates a case with a slight angulation change between angiography and fluoroscopy, which can be observed from the image borders. This geometric discrepancy introduces additional misalign- ment in the overlay. Overall, these examples highlight the robustness of the proposed method under occlusions, reso- lution differences, and mild geometric inconsistencies, while also illustrating remaining limitations in such scenarios. 4. Conclusion In this work, we presented a unified and fully automated framework for ECG-free dynamic coronary roadmapping that jointly addresses cardiac phase matching and catheter tip tracking. By leveraging a large-scale spatio-temporal encoder pretrained on millions of unlabeled angiography and fluoroscopy frames, the proposed method effectively cap- tures fine-grained cardiac motion dynamics that are essential for accurate temporal alignment. Unlike prior approaches that rely heavily on extensive frame-level dense annotations, our framework incorporates lightweight auxiliary tasks, in- cluding ECG R-peak detection and catheter tip tracking, to stabilize training and provide meaningful motion cues without requiring additional manual supervision. While ex- isting methods typically achieve strong performance through extensive manual annotations, our approach instead relies on large-scale unlabeled data to learn a foundation model that : Preprint submitted to ElsevierPage 11 of 13 not only reduces annotation requirements for dynamic coro- nary roadmapping, but also has the potential to generalize to a broader range of angiography-based tasks. Comprehensive evaluations on clinical X-ray datasets demonstrate that the proposed approach achieves state-of- the-art performance in cardiac phase matching, yielding low temporal misalignment with high consistency across diverse imaging conditions. The results further highlight the importance of domain-specific spatio-temporal pretraining, as well as the benefit of jointly modeling spatial and temporal information within a unified architecture. In addition, the majority-voting post-processing strategy improves robust- ness during inference and naturally provides a confidence score that correlates well with phase-matching error, en- abling uncertainty-aware deployment in clinical settings. We further showed that the unified model achieves com- petitive catheter tip tracking performance under automatic initialization, while eliminating the need for separate task- specific models. Although a modest performance gap is observed compared to methods trained with manual ini- tialization, this difference is not statistically significant and is outweighed by the advantages of reduced system com- plexity, faster inference, and improved suitability for fully automated DRM workflows. Despite these promising results, several limitations re- main and point toward directions for future work. In partic- ular, while the proposed phase-matching approach achieves low average error, a non-negligible number of outliers per- sist in challenging cases. Achieving a fully reliable system may therefore require additional robustness, which could be obtained through increased annotated training set or by introducing further auxiliary tasks that provide complemen- tary motion cues, such as guidewire detection. Moreover, although catheter tip tracking errors are low, closing the remaining gap to methods such as HiFT with manual initial- ization within a unified model remains an open challenge. Addressing this may require more advanced optimization strategies tailored for multi-task learning, enabling better task balancing and more effective feature sharing across objectives. Overall, the proposed framework represents a step to- ward practical, real-time dynamic coronary roadmapping with minimal reliance on manual annotations or external signals. By demonstrating that large-scale spatio-temporal pretraining combined with carefully designed auxiliary tasks can effectively reduce the dependence on extensive labeled data, this work opens avenues for extending the learned representations to other interventional imaging tasks where motion understanding and robustness are critical. Disclaimer The concepts and information presented in this paper are based on research results that are not commercially available. Future commercial availability cannot be guaranteed. References Assran, M., Bardes, A., Fan, D., Garrido, Q., Howes, R., Muckley, M., Rizvi, A., Roberts, C., Sinha, K., Zholus, A., et al., 2025. V-jepa 2: Self- supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985 . Bardes, A., Garrido, Q., Ponce, J., Chen, X., Rabbat, M., LeCun, Y., Assran, M., Ballas, N., 2024. Revisiting feature prediction for learning visual representations from video. arXiv preprint arXiv:2404.08471 . Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., Joulin, A., 2020. Unsupervised learning of visual features by contrasting cluster assignments. Advances in neural information processing systems 33, 9912β9924. Caron, M., Touvron, H., Misra, I., JΓ©gou, H., Mairal, J., Bojanowski, P., Joulin, A., 2021. Emerging properties in self-supervised vision transformers, in: Proceedings of the IEEE/CVF international conference on computer vision, p. 9650β9660. Chen, T., Kornblith, S., Norouzi, M., Hinton, G., 2020. A simple frame- work for contrastive learning of visual representations, in: International conference on machine learning, PmLR. p. 1597β1607. Cui, Y., Jiang, C., Wang, L., Wu, G., 2022. Mixformer: End-to-end tracking with iterative mixed attention, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 13608β 13618. Demoustier, M., Zhang, Y., Narasimha Murthy, V., Ghesu, F.C., Comaniciu, D., 2023. Contrack: contextual transformer for device tracking in x- ray, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 679β688. Elion, J.L., 1989. Dynamic coronary roadmapping. US Patent 4,878,115. Faranesh, A.Z., Kellman, P., Ratnayaka, K., Lederman, R.J., 2013. Inte- gration of cardiac and respiratory motion into mri roadmaps fused with x-ray. Medical physics 40, 032302. Fischer, P., Faranesh, A., Pohl, T., Maier, A., Rogers, T., Ratnayaka, K., Lederman, R., Hornegger, J., 2017. An mr-based model for cardio- respiratory motion compensation of overlays in x-ray fluoroscopy. IEEE transactions on medical imaging 37, 47β60. Ghojogh, B., Sikaroudi, M., Shafiei, S., Tizhoosh, H.R., Karray, F., Crow- ley, M., 2020. Fisher discriminant triplet and contrastive losses for training siamese networks, in: 2020 international joint conference on neural networks (IJCNN), IEEE. p. 1β7. Hadsell, R., Chopra, S., LeCun, Y., 2006. Dimensionality reduction by learning an invariant mapping, in: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPRβ06), p. 1735β1742. doi:10.1109/CVPR.2006.100. He, K., Fan, H., Wu, Y., Xie, S., Girshick, R., 2020. Momentum contrast for unsupervised visual representation learning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9729β9738. Islam, S., Murthy, V.N., Neumann, D., Cimen, S., Sharma, P., Maier, A., Comaniciu, D., Ghesu, F.C., 2024a. A novel tracking framework for devices in x-ray leveraging supplementary cue-driven self-supervised features, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. p. 25β34. Islam, S., Murthy, V.N., Neumann, D., Das, B.K., Sharma, P., Maier, A., Comaniciu, D., Ghesu, F.C., 2024b. Self-supervised learning for interventional image analytics: toward robust device trackers. Journal of Medical Imaging 11, 035001β035001. Khan, S.Q., Ludman, P.F., 2022. Percutaneous coronary intervention. Medicine 50, 437β444. Khosla, P., Teterwak, P., Wang, C., Sarna, A., Tian, Y., Isola, P., Maschinot, A., Liu, C., Krishnan, D., 2020. Supervised contrastive learning. Advances in neural information processing systems 33, 18661β18673. Kim, D., Park, S., Jeong, M.H., Ryu, J., 2018. Registration of angiographic image on real-time fluoroscopic image for image-guided percutaneous coronary intervention. International journal of computer assisted radi- ology and surgery 13, 203β213. King, A.P., Boubertakh, R., Rhode, K.S., Ma, Y., Chinchapatnam, P., Gao, G., Tangcharoen, T., Ginks, M., Cooklin, M., Gill, J.S., et al., 2009. A : Preprint submitted to ElsevierPage 12 of 13 subject-specific technique for respiratory motion correction in image- guided cardiac catheterisation procedures. Medical Image Analysis 13, 419β431. Li, B., Yan, J., Wu, W., Zhu, Z., Hu, X., 2018. High performance visual tracking with siamese region proposal network, in: Proceedings of the IEEE conference on computer vision and pattern recognition, p. 8971β 8980. Liu, Y., Zhao, L., Chen, E.Z., Chen, X., Chen, T., Sun, S., 2024. Auxiliary input in training: Incorporating catheter features into deep learning models for ecg-free dynamic coronary roadmapping, in: International Conference on Medical Image Computing and Computer-Assisted In- tervention, Springer. p. 67β77. Ma, H., Smal, I., Daemen, J., van Walsum, T., 2020. Dynamic coronary roadmapping via catheter tip tracking in x-ray fluoroscopy with deep learning based bayesian filtering. Medical image analysis 61, 101634. Manhart, M., Zhu, Y., Vitanovski, D., 2011. Self-assessing image-based respiratory motion compensation for fluoroscopic coronary roadmap- ping, in: 2011 IEEE International Symposium on Biomedical Imaging: From Nano to Macro, IEEE. p. 1065β1069. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al., 2023. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 . Peressutti, D., Penney, G.P., Housden, R.J., Kolbitsch, C., Gomez, A., Rijkhorst, E.J., Barratt, D.C., Rhode, K.S., King, A.P., 2013. A novel bayesian respiratory motion model to estimate and resolve uncertainty in image-guided cardiac interventions. Medical image analysis 17, 488β 502. Piayda, K., Kleinebrecht, L., Afzal, S., Bullens, R., Ter Horst, I., Polzin, A., Veulemans, V., Dannenberg, L., Wimmer, A.C., Jung, C., et al., 2018. Dynamic coronary roadmapping during percutaneous coronary intervention: a feasibility study. European journal of medical research 23, 36. Ravi, N., Gabeur, V., Hu, Y.T., Hu, R., Ryali, C., Ma, T., Khedr, H., RΓ€dle, R., Rolland, C., Gustafson, L., et al., 2024. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:2408.00714 . Schneider, M., Sundar, H., Liao, R., Hornegger, J., Xu, C., 2010. Model- based respiratory motion compensation for image-guided cardiac inter- ventions, in: 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE. p. 2948β2954. Schroff, F., Kalenichenko, D., Philbin, J., 2015. Facenet: A unified em- bedding for face recognition and clustering, in: Proceedings of the IEEE conference on computer vision and pattern recognition, p. 815β823. Shechter, G., Shechter, B., Resar, J.R., Beyar, R., 2005. Prospective motion correction of x-ray images for coronary interventions. IEEE transactions on medical imaging 24, 441β450. Tehrani, S., Laing, C., Yellon, D.M., Hausenloy, D.J., 2013. Contrast- induced acute kidney injury following pci. European journal of clinical investigation 43, 483β490. Timinger, H., Krueger, S., Dietmayer, K., Borgert, J., 2005. Motion compensated coronary interventional navigation by means of diaphragm tracking and elastic motion models. Physics in Medicine & Biology 50, 491. Tong, Z., Song, Y., Wang, J., Wang, L., 2022. Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35, 10078β10093. Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., Qiao, Y., 2023. Videomae v2: Scaling video masked autoencoders with dual masking, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 14549β14560. Yan, B., Peng, H., Fu, J., Wang, D., Lu, H., 2021. Learning spatio-temporal transformer for visual tracking, in: Proceedings of the IEEE/CVF inter- national conference on computer vision, p. 10448β10457. Zhao, X., Qi, H., Luo, R., Davis, L., 2019. A weakly supervised adaptive triplet loss for deep metric learning, in: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, p. 0β0. Zhu, Y., Tsin, Y., Sundar, H., Sauer, F., 2010. Image-based respiratory motion compensation for fluoroscopic coronary roadmapping, in: In- ternational Conference on Medical Image Computing and Computer- Assisted Intervention, Springer. p. 287β294. : Preprint submitted to ElsevierPage 13 of 13