Paper deep dive
Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction
Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen, Renyang Gu, Xinyu Liu, Yongsheng Pan, Yong Xia
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/19/2026, 4:33:10 AM
Summary
The paper introduces LiftXR, a geometry-guided framework for bi-planar X-ray-to-CT reconstruction. It addresses the ill-posed nature of reconstructing 3D volumes from sparse 2D X-rays by explicitly modeling anatomical layout as a spatial constraint. The framework interleaves layout generation (from X-rays) and layout perception (from reconstructed CT) to refine anatomical boundaries and calibrate intensities, achieving state-of-the-art performance on CT-RATE and LIDC-IDRI datasets.
Entities (13)
Relation Signals (12)
LiftXR → evaluatedon → LIDC-IDRI
confidence 95% · Extensive experiments on two public datasets... LIDC-IDRI
LiftXR → evaluatedon → CT-RATE
confidence 95% · Extensive experiments on two public datasets... CT-RATE
LiftXR → usescomponent → Layout Lifter
confidence 95% · Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays
LiftXR → usescomponent → Intensity Renderer
confidence 95% · providing spatial guidance for an intensity renderer to reconstruct a CT volume
LiftXR → usescomponent → Anatomical Parser
confidence 95% · An anatomical parser then performs volumetric perception on the reconstruction
LiftXR → usescomponent → Intensity Calibrator
confidence 95% · an intensity calibrator uses the parsed layout to correct region-specific intensity inconsistencies
Intensity Calibrator → calibrates → CT Volume
confidence 90% · an intensity calibrator uses the parsed layout to correct region-specific intensity inconsistencies
Layout Lifter → generates → Anatomical Layout
confidence 90% · a layout lifter first generates a 3D anatomical layout from bi-planar X-rays
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.
Tags
Links
- Source: https://arxiv.org/abs/2608.17255v1
- Canonical: https://arxiv.org/abs/2608.17255v1
Trouble viewing inline? Open PDF directly →
Full Text
49,305 characters extracted from source content.
Expand or collapse full text
Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction Yifei Wu Yicheng Wu Qiang Ma Qi Chen Renyang Gu Xinyu Liu Yongsheng Pan Yong Xia Abstract X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released. Figure 1: X-ray projection and CT reconstruction. The projected value p(u,v)p(u,v) records the accumulated attenuation μ()μ(x) along its corresponding ray ru,vr_u,v. We characterize this process by an anatomical layout capturing the spatial organization of organs and tissues, and an intensity field specifying their voxel-wise CT values. Recovering both aspects from sparse X-ray views is highly ill-posed since geometric structure and intensity information are strongly entangled. Then, an oracle experiment on CT-RATE shows that reconstruction conditioned on ground-truth anatomical layouts achieves 37.04 dB PSNR and 95.07% SSIM, demonstrating the value of anatomical layout as an effective spatial constraint. Introduction Computed tomography (CT) reconstructs volumetric attenuation fields from multi-angle X-ray projections, typically requiring hundreds of views for accurate 3D imaging (21). However, dense-view acquisition increases radiation exposure, limiting its applicability to longitudinal monitoring and large-scale screening (3). Sparse-view CT reconstruction reduces this burden by using only a few dozen projections (26; 28; 35; 13). bi-planar X-ray-to-CT reconstruction offers an even more acquisition-efficient alternative, relying only on frontal and lateral radiographs that are routinely acquired in clinical practice and widely available in public datasets (47; 38). Nevertheless, recovering a 3D CT volume from only two projections remains severely ill-posed. The ill-posedness originates from the severe loss of depth information during projection acquisition (4; 41). As illustrated in Fig. 1, each X-ray pixel records attenuation accumulated along a ray path, collapsing depth-dependent anatomical information into a two-dimensional measurement. We describe the underlying CT volume by an anatomical layout exhibiting the spatial extent and arrangement of organs and tissues, and a volumetric intensity field specifying their voxel-wise CT values. Because distinct 3D layouts and intensity distributions can produce similar 2D projections, the two aspects remain highly entangled under limited observations. Existing X-ray-to-CT methods (46; 22; 33) have introduced geometric consistency, learnable priors, or anatomy-related regularization to facilitate CT reconstruction. However, anatomical organization is typically encoded implicitly within image features or inferred only through the reconstructed intensities, rather than being modeled as an explicit region-level 3D layout. As a result, reconstructed volumes may perform favorably on voxel-wise metrics while still exhibiting inaccurate regional extents, blurred boundaries, or inconsistent spatial relationships (24; 36; 8). To further investigate the role of anatomical layout, we conduct an oracle study on CT-RATE (14), in which the reconstruction model is conditioned on ground-truth anatomical masks. The resulting reconstruction achieves a PSNR of 37.04 dB and an SSIM of 95.07%. This result demonstrates that the anatomical layout provides strong spatial constraints for CT reconstruction. By localizing major body regions in 3D space, the layout reduces the geometric uncertainty that would otherwise need to be resolved jointly with voxel-wise intensity estimation. Motivated by these observations, we propose LiftXR, an end-to-end, geometry-guided framework that explicitly incorporates anatomical layout recovery into bi-planar X-ray-to-CT reconstruction. Specifically, a layout lifter first generates an initial 3D anatomical layout from the bi-planar X-rays. This layout provides explicit spatial guidance for an intensity renderer to reconstruct a CT volume. Although imperfect, the reconstructed volume contains spatially resolved boundary and intensity cues, and a subsequent anatomical parser predicts region masks from the reconstructed CT to obtain a refined layout. Finally, an intensity calibrator uses the parsed layout to correct region-specific intensity inconsistencies. Importantly, the two anatomical layouts play fundamentally different roles. The lifted layout is inferred directly from bi-planar projections and serves as a generative geometric prior for volumetric reconstruction. In contrast, the parsed layout is obtained by perceiving anatomical regions from the initially reconstructed CT volume, constituting an explicit volumetric anatomical parsing task. This transition from layout generation to layout perception enables the reconstructed volume to provide feedback for geometry refinement, which subsequently guides region-specific intensity calibration. Our contributions are threefold: • We formulate the bi-planar X-ray-to-CT reconstruction task by explicitly modeling anatomical layout as a spatial constraint for volumetric generation. By specifying where anatomical regions are located in 3D space, the layout constrains and guides the estimation of what CT intensities they exhibit. • We propose LiftXR, a geometry-guided framework that couples volumetric generation with anatomical perception. LiftXR integrates projection-conditioned layout generation, conditional CT rendering, reconstruction-conditioned anatomical parsing, and region-specific intensity calibration, forming a unified generation-perception feedback process. • Extensive experiments on two public chest CT datasets demonstrate that LiftXR achieves state-of-the-art reconstruction performance. Its reconstructed CT volumes further yield superior downstream segmentation results, showcasing improved anatomical fidelity. Related Work X-ray-to-CT Reconstruction X-ray-to-CT reconstruction aims to recover a 3D CT volume from one or two 2D X-ray projections, which is highly challenging due to the severe ambiguity introduced by the projection process. X2CT (46) first demonstrated the feasibility of reconstructing volumetric CT from bi-planar X-rays, and subsequent studies introduced geometric, generative, and structural priors to reduce this ambiguity. PerX2CT (22) incorporates perspective projection modeling to improve spatial localization, DiffuX2CT (30) introduces a conditional diffusion prior to improve volumetric consistency, and DSDF (33) separates global structure modeling from local texture synthesis to enhance reconstruction quality. However, these methods mainly optimize voxel-level appearance or represent anatomical structure implicitly, leaving explicit constraints on organ boundaries, spatial extents, and inter-organ relationships underexplored (24). LiftXR addresses this limitation by explicitly predicting anatomical structures, where predicted layouts guide CT estimation and reconstructed volumes provide cues for further layout parsing. This interleaved layout-intensity interaction constrains the solution space and promotes anatomically consistent reconstruction with clear anatomical constraints derived from bi-planar X-rays. Sparse-view CT Reconstruction Sparse-view CT reconstruction aims to reduce radiation exposure by recovering CT volumes from limited X-ray projections. Analytical methods such as FBP and FDK require sufficient angular sampling and often introduce streak artifacts and structural distortions under severe undersampling (10). Traditional approaches (45; 32) alleviate this ill-posedness with handcrafted regularization priors, but they typically require expensive optimization and careful parameter tuning. Deep learning methods improve reconstruction by learning effective data-driven priors from large-scale datasets (5; 25). Image-domain approaches, such as FBPConvNet (19), refine degraded analytical reconstructions, whereas projection-domain approaches recover missing measurements before analytical reconstruction through DNN-based sinogram synthesis (23) or learned interpolation (9). Nevertheless, these approaches still depend on multiple projections collected along a scanning trajectory. In contrast, X-ray-to-CT reconstruction relies on only one or two radiographs, leading to a more severely under-constrained inverse problem that requires stronger anatomical priors. Figure 2: Overview of LiftXR, an interleaved, geometry-guided framework coupling volumetric generation with anatomical perception. A Layout Lifter first performs projection-conditioned anatomical generation, which guides an Intensity Renderer to reconstruct an initial CT volume. A Layout Parser then performs reconstruction-conditioned anatomical perception using the spatially resolved boundary and intensity cues. The parsed layout is subsequently fed back to an Intensity Calibrator to correct region-specific intensity distributions. Medical Image Generation Generative models, particularly CNN-GANs and diffusion models, have been widely applied to medical image reconstruction (37; 27; 34). CNN-GANs learn supervised input–target mappings through convolutional generators, while adversarial discriminators promote realistic textures and plausible structures (11; 16). They have been applied to accelerated MRI, low-dose CT, and cross-modality translation (44; 29; 1; 42; 43), offering efficient inference and structural control. Diffusion models generate high-quality images through iterative denoising (15) and have been extended to MRI, sparse-view CT, and low-dose CT reconstruction (40; 7). Despite their strong generative capability, diffusion models require multi-step sampling, resulting in considerable computational overhead for 3D reconstruction and structural uncertainty (18; 6). Therefore, we adopt a CNN-GAN framework, which provides efficient deterministic inference and better integrates with anatomy-conditioned CT reconstruction. Method LiftXR reconstructs a CT volume from paired PA and lateral X-rays by interleaving anatomical layout recovery with volumetric intensity estimation. As illustrated in Fig. 2, the two projections are first up-sampled and fused into an explicit 3D X-ray volume. A Layout Lifter predicts a coarse anatomical layout, which constrains an Intensity Renderer to generate an initial CT volume. Although imperfect, this reconstruction provides spatially resolved boundary and intensity cues for a Layout Parser to refine the anatomical layout, further regularizing regional occupancy and spatial relationships. The refined layout is then combined with the volumetric representation to guide an Intensity Calibrator in correcting region-specific intensity inconsistencies. Through this reciprocal interaction between layout and intensity, LiftXR progressively resolves geometric ambiguity and improves anatomical fidelity. Bi-planar X-ray Initialization Given two PA and LAT projections, we first construct an explicit X-ray volume to bridge the gap between two-dimensional observations and the target three-dimensional space. Each projection is replicated along its unobserved spatial dimension and aligned within a shared volumetric coordinate system for subsequent aggregation. VPA=Replicate(XPA,D),VLAT=Replicate(XLAT,W),V_PA=Replicate(X_PA,D),V_LAT=Replicate(X_LAT,W), (1) where VPA,VLAT∈ℝH×W×DV_PA,V_LAT ^H× W× D denote the lifted representations of the two views. We then aggregate these representations to obtain the X-ray volume by voxel-wise averaging Vxray=Average(VPA,VLAT),V_xray=Average (V_PA,V_LAT ), (2) which preserves the ray-integrated observations from both projections in a unified representation. This strategy can naturally extend to additional projection angles by lifting each view into the same volumetric coordinate system and applying the same aggregation operation. However, a direct replication operation cannot recover the missing depth information or resolve the inherent projection ambiguity. Layout Lifting and Intensity Rendering Therefore, starting from the X-ray volume VxrayV_xray, LiftXR performs geometry-guided volumetric lifting to establish an initial volumetric state. Instead of directly predicting intensity from sparse projections, a Layout Lifter θll _l first estimates a 3D anatomical layout that provides explicit geometric constraints for subsequent CT reconstruction. The layout is predicted solely based on VxrayV_xray as Llifted=θll(Vxray),L_lifted= _l(V_xray), (3) where LliftedL_lifted represents the lifted layout. By encoding organ occupancy, spatial extent, and relative arrangement, LliftedL_lifted provides anatomical configurations compatible with the bi-planar observations. Conditioned on the projection-derived representation and predicted layout, the Intensity Renderer θir _ir reconstructs the coarse CT volume as Vrendered=θir(Vxray⊕Llifted)V_rendered= _ir (V_xray L_lifted ) (4) where ⊕ is the channel-wise concatenation. This layout-conditioned lifting process transforms the X-ray volume into a structured volumetric state that provides a geometry-aware initialization for subsequent refinement. Layout Parsing and Intensity Calibration The lifted layout and CT volume remain imperfect due to limited information from bi-planar projections. LiftXR therefore introduces an interleaved refinement stage that parses the anatomical layout and calibrates the CT intensity field. Lparsed=θlp(Vxray⊕Vrendered),L_parsed= _lp (V_xray V_rendered ), (5) where VxrayV_xray preserves projection evidence and VrenderedV_rendered provides spatially resolved boundary and attenuation cues for the layout refinement by a parser θlp _lp. Conditioned on the refined layout LparsedL_parsed, the Intensity Calibrator θic _ic further updates the rendered CT volume as Vcalibrated=θic(Vxray⊕Vrendered⊕Lparsed).V_calibrated= _ic (V_xray V_rendered L_parsed ). (6) Overall, although both stages predict anatomical layouts, they solve different tasks. The Layout Lifter performs projection-conditioned layout generation to initialize the missing 3D geometry, whereas the Layout Parser performs reconstruction-conditioned anatomical perception to recover spatially resolved organ boundaries and regional identities. Interleaved Layout-Intensity Supervision There are two aspects to supervise the anatomical layout prediction and CT reconstruction. We denote predicted layouts as =Llifted,LparsedA=\L_lifted,L_parsed\ and reconstructed volumes as =Vrendered,VcalibratedV=\V_rendered,V_calibrated\. Layout supervision constrains 3D anatomical geometry, while CT supervision preserves intensity and volumetric appearance. Anatomical Layout Supervision: Accurate anatomical layouts provide explicit constraints on organ occupancy, boundary placement, and spatial organization. We extract seven parenchymal and major thoracoabdominal regions, including tissue, bone, heart, liver, left lung, right lung, and esophagus, for spatial localization. The target anatomical layout is extracted as LGT=Eseg(VGT),L_GT=E_seg (V_GT ), (7) where Eseg(⋅)E_seg(·) denotes the anatomical mask extractor. Both the lifted layout and parsed layout are supervised by ℒlayout=∑a∈[CE(a,LGT)+Dice(a,LGT)],L_layout= _a [CE (a,L_GT )+Dice (a,L_GT ) ], (8) where a is selected from A. This supervision guides both initial layout lifting and subsequent layout parsing toward anatomically plausible three-dimensional structures. CT Reconstruction Supervision: While anatomical layouts provide explicit geometric constraints, they cannot fully determine region-specific intensity distributions. Therefore, both rendered and calibrated CT volumes are supervised using the ground-truth CT volume. Following (46), for each reconstructed volume v∈v , the multi-scale feature matching loss is defined as ℒMSL(v)=‖MS(v)−sg(MS(VGT))‖1,L_MSL^(v)= \|MS(v)-sg (MS(V_GT) ) \|_1, (9) where MS(⋅)MS(·) denotes the multi-scale feature transformation, and sg(⋅)sg(·) denotes the stop-gradient operation. The CT reconstruction objective is defined as ℒrec=∑v∈[ℒ1(v,VGT)+ℒMSL(v)+ℒadv(v)],L_rec= _v [L_1 (v,V_GT )+L_MSL^(v)+L_adv^(v) ], (10) where ℒ1L_1 is the mean absolute error (MAE) loss and ℒadv(v)L_adv^(v) denotes the adversarial loss that encourages reconstructed CT volumes to match the distribution of real CT volumes, following the typical setting (46; 33). This objective preserves voxel-level intensities, multi-scale structural consistency, and realistic volumetric appearance throughout the lifting and calibration stages. Overall Training Objective: The anatomical layout and CT reconstruction objectives jointly optimize the shape and intensity branches of LiftXR. The overall training objective is defined as ℒtotal=ℒrec+λ×ℒlayout,L_total=L_rec+λ×L_layout, (11) where λ is a loss weight to balance the two aspects. Method CT-RATE LIDC-IDRI Reconstruction Segmentation Reconstruction Segmentation PSNR (dB)↑ SSIM (%)↑ Dice (%)↑ HD95 (voxel)↓ PSNR (dB)↑ SSIM (%)↑ Dice (%)↑ HD95 (voxel)↓ X2CT (46) 25.96±1.0025.96_± 1.00 71.88±0.5171.88_± 0.51 63.39±0.7663.39_± 0.76 16.05±2.1216.05_± 2.12 25.06±1.1325.06_± 1.13 69.47±0.5769.47_± 0.57 56.80±0.7356.80_± 0.73 19.81±1.4519.81_± 1.45 PerX2CT (22) 25.84±0.8425.84_± 0.84 71.29±0.3671.29_± 0.36 55.44±0.5955.44_± 0.59 19.19±1.6419.19_± 1.64 25.34±0.8725.34_± 0.87 70.24¯±0.60 70.24_± 0.60 58.96±0.7858.96_± 0.78 18.23±1.0318.23_± 1.03 DSDF (33) 25.82±0.9925.82_± 0.99 73.29¯±0.39 73.29_± 0.39 66.09¯±0.79 66.09_± 0.79 13.92¯±1.81 13.92_± 1.81 25.73¯±0.94 25.73_± 0.94 69.43±0.5869.43_± 0.58 58.38±0.8958.38_± 0.89 16.87¯±1.01 16.87_± 1.01 GAAL (31) 22.82±1.2122.82_± 1.21 65.62±0.4065.62_± 0.40 28.21±0.6628.21_± 0.66 94.19±3.5094.19_± 3.50 21.02±1.2321.02_± 1.23 51.53±0.5251.53_± 0.52 30.09±0.9830.09_± 0.98 85.56±4.9885.56_± 4.98 DX2CT (17) 25.76±0.4925.76_± 0.49 70.42±0.1570.42_± 0.15 64.11±0.4964.11_± 0.49 14.98±1.2314.98_± 1.23 25.06±1.3425.06_± 1.34 69.27±0.4469.27_± 0.44 59.02±1.0959.02_± 1.09 17.45±1.4917.45_± 1.49 DiffuX2CT (30) 26.03¯±0.60 26.03_± 0.60 72.82±0.1972.82_± 0.19 65.58±0.5165.58_± 0.51 14.82±1.1214.82_± 1.12 25.12±1.3525.12_± 1.35 70.16±0.4070.16_± 0.40 59.22¯±0.95 59.22_±0.95 17.07±1.4017.07_± 1.40 LiftXR (Ours) 26.45±1.00∗∗26.45_± 1.00^** 76.44±0.40∗∗76.44_± 0.40^** 68.64±0.80∗∗68.64_± 0.80^** 12.51±1.60∗∗12.51_± 1.60^** 26.03±1.08∗26.03_± 1.08^* 71.72±0.60∗∗71.72_± 0.60^** 62.16±0.87∗∗62.16_± 0.87^** 15.91±1.14∗∗15.91_± 1.14^** Table 1: Quantitative comparisons on the CT-RATE and LIDC-IDRI datasets. Reconstruction and segmentation results are reported as mean ± standard deviation. The best and second-best results are highlighted in bold and underlined, respectively. ∗ and ∗ indicate that LiftXR significantly outperforms the second-best method at p<0.05p<0.05 and p<0.01p<0.01, respectively. Experiment and Result Implementation Details Dataset and Pre-processing: We evaluate LiftXR on two benchmarking chest CT datasets. On the large-scale CT-RATE dataset (14), we randomly select 2,000 subjects for training and 300 for testing. On the LIDC-IDRI dataset, which contains 1,018 CT volumes (2), we use 818 volumes for training and 200 for testing. Following the official setting, anatomical labels are generated on both datasets by using a frozen TotalSegmentator model (39). Following prior X-ray-to-CT studies (46; 12), we generate PA and Lateral radiographs from each CT volume using digitally reconstructed radiography (DRR) techniques, yielding synthetic pairs. We use the DiffDRR (12) tool11 1 https://github.com/eigenvivek/DiffDRR, with the setting using a source-to-detector distance of 3,000 m and a detector pixel spacing of 1.0 m. All CT volumes are resized to 128×128×128128× 128× 128. The intensity range is clamped to [−1024,1024][-1024,1024] HU and then normalized to [0,1][0,1]. Experimental Settings: We train LiftXR using the Adam optimizer with a learning rate of 2e−42e^-4 and a batch size of 4. Based on the sensitivity analysis, λ is set to 1. Eseg(⋅)E_seg(·) is set to TotalSegmentator to obtain the anchor organs. 100,000 iterations are used for training, taking approximately 60 hours on a single NVIDIA GeForce RTX 4090 GPU. Each model inside our LiftXR is a 3D U-Net architecture, and the total computational complexity is 6480 GFLOPs with 172.62M parameters. To ensure fair comparisons, all experiments use a fixed random seed 43 and follow identical settings. For external downstream segmentation, we mainly focus on thoracoabdominal regions, and 20 specific categories are selected, which are elaborated in the supplementary material. Main Results Quantitative Analysis: Tab. 1 compares LiftXR with representative X-ray-to-CT reconstruction methods on CT-RATE and LIDC-IDRI. LiftXR achieves the best performance across all reconstruction and segmentation metrics on both datasets, demonstrating consistent effectiveness in recovering CT intensities and anatomically meaningful three-dimensional structures. The improvements in PSNR are relatively moderate compared with those in anatomical metrics, as voxel-wise errors are dominated by large homogeneous regions and are less sensitive to localized corrections around organ boundaries and small structures. In contrast, the consistent SSIM improvements indicate that LiftXR better preserves local structural patterns and tissue contrast during reconstruction. These results suggest that explicit anatomical layout modeling reduces the depth ambiguity inherent in bi-planar projections and constrains the reconstruction toward anatomically plausible volumetric configurations. These gains indicate that the predicted layout explicitly constrains organ extent, shape, and spatial relationships, preventing anatomically implausible solutions that may still achieve comparable voxel-wise similarity. Meanwhile, the reconstructed CT volume provides boundary and tissue-specific intensity cues to refine uncertain layouts. The refined layout subsequently guides CT calibration to reduce tissue mixing, geometric distortions, and inconsistent intensity distributions. Through this interleaved refinement process, LiftXR progressively improves both organ occupancy and boundary accuracy. Figure 3: Qualitative comparisons on the CT-RATE dataset. The top two rows show axial and coronal slices, while the bottom two rows present the corresponding segmentation layouts in 2D and 3D. LiftXR more faithfully preserves anatomical structures and recovers fine-grained details, such as maintaining rib continuity and reconstructing a smooth liver contour. Qualitative Analysis: Fig. 3 compares axial and coronal CT slices, sagittal segmentation overlays, and 3D anatomical renderings on the CT-RATE dataset. Existing methods generally recover the overall thoracic structure but exhibit noticeable errors in both slice-level details and volumetric anatomy. X2CT, PerX2CT, and DSDF preserve the general lung regions but produce blurred mediastinal boundaries, missing fine structures, and inaccurate organ shapes. GAAL, a multi-view reconstruction method, exhibits severe over-smoothing and substantial loss of anatomical details, likely due to the difficulty of inferring fine structures from limited projection constraints. DX2CT and DiffuX2CT, diffusion-based approaches, recover cleaner global structures with improved volumetric coherence, but still introduce local boundary errors and inconsistencies in inter-organ relationships. In comparison, LiftXR produces results closer to the ground truth, with clearer lung contours and more accurate vertebral and rib structures. The final row removes surrounding tissues to visualize internal anatomy, where LiftXR maintains more continuous ribs, better spine alignment, more symmetric lungs, and more plausible organ arrangements. These localized geometric improvements explain the moderate PSNR gains and the larger improvements in anatomical metrics, which are more sensitive to organ extent and boundary accuracy. The qualitative observations agree with the quantitative results, demonstrating that interleaved refinement of anatomical layout and CT intensity enables more accurate volumetric reconstruction. Layout Interleaved Training Reconstruction Segmentation PSNR (dB)↑ SSIM (%)↑ Dice (%)↑ HD95 (voxel)↓ - - 25.68 72.89 63.54 15.45 ✓ 26.07 75.34 67.03 13.31 ✓ - 25.95 74.10 65.45 13.45 ✓ ✓ 26.45 76.44 68.64 12.51 Table 2: Ablation studies on the CT-RATE dataset. Best and second-best results are highlighted. Both anatomical layout and interleaved training are evaluated for both reconstruction and downstream segmentation. Ablation Study: Tab. 2 reports the ablation results on CT-RATE. Both the layout and interleaved training contribute to reconstruction and segmentation performance, while their combination provides complementary benefits. Anchor organs establish explicit organ-level references, constraining the possible anatomical configurations during the initial lifting process and reducing large deviations in organ extent and spatial arrangement. The interleaved training further improves the reconstruction by enabling interleaved interaction between the anatomical layout and CT volume, where the estimated geometry and intensity information mutually correct residual structural and attenuation errors. While anchor organs mainly provide global anatomical guidance, the interleaved training focuses on improving local details, including organ boundaries, tissue appearance, and inter-organ relationships. Their effects demonstrate that reliable anatomical prediction together with interleaved layout-intensity refinement is effective in X-ray-to-CT reconstruction. Figure 4: Intermediate visualization of the layout-intensity refinement process in LiftXR. The first row shows the initial lifted anatomical layout and rendered CT volume, the second row presents the refined layout and calibrated CT after interleaved optimization, and the third row provides the ground-truth reference on the CT-RATE dataset. Analysis of Intermediate Processes: As shown in Fig. 4, the visualization further illustrates the effectiveness of the interleaved layout-intensity refinement process in LiftXR. Starting from the initial lifted layout and rendered CT volume, the reconstructed CT provides spatially resolved boundary and intensity cues that help refine uncertain anatomical regions. Meanwhile, the refined layout introduces more accurate structural constraints for subsequent CT intensity calibration. Through this reciprocal interaction, both anatomical layouts and CT volumes are progressively improved, leading to more accurate organ localization, clearer anatomical boundaries, and more consistent intensity distributions compared with the initial reconstruction. Layout Selection Loss Weight λ PSNR (dB)↑ SSIM (%)↑ HU-based Regions 1 26.07 75.53 Anchor Organs 26.45 76.44 All Organs 26.37 76.31 Anchor Organs 0.1 26.28 75.70 1 26.45 76.44 10 26.09 75.46 Table 3: Sensitivity analysis of anatomical layout selection and the loss weight λ on the CT-RATE dataset. Discussions Effect of Layout Selection: Tab. 3 evaluates different anatomical supervision strategies on CT-RATE. HU-based regions (33) provide coarse tissue constraints from intensity ranges, while organ-level supervision introduces explicit anatomical boundaries and spatial relationships. Both organ anchors and full anatomical labels outperform HU-based regions, demonstrating the importance of organ-level guidance for resolving projection ambiguity. However, incorporating all possible anatomical targets from TotalSegmentator does not further improve performance since hollow and fine organs may introduce additional segmentation uncertainty, leading to sub-optimal performance. Effect of λ: Tab. 3 further investigates the balance between anatomical layout and CT supervision by varying the loss weight λ. The best performance is achieved at λ=1λ=1, suggesting that accurate bi-planar X-ray-to-CT reconstruction requires a joint optimization of anatomical geometry and intensity. When λ is reduced, insufficient anatomical supervision weakens the geometric constraints provided by the predicted layouts, limiting the recovery of organ structures from ambiguous projections. In contrast, an excessively large λ places excessive emphasis on anatomical consistency, which may degrade intensity reconstruction. Therefore, λ=1λ=1 provides an effective balance between anatomical guidance and CT intensity recovery and is adopted in all experiments. Generalization to Real X-ray Images: Fig. 5 shows the results using real frontal and lateral X-rays on the MIMIC dataset (20). Following (46; 22), we first train a CycleGAN (48) to transform the style between real X-rays and our synthetic projections, and then reconstruct CT volumes using X2CT, DSDF, and LiftXR. In the enlarged regions, X2CT produces noticeable boundary artifacts, suggesting inaccurate anatomical recovery under real X-ray inputs. DSDF preserves more coherent global structures but still fails to recover clear textures. In contrast, LiftXR produces sharper organ boundaries and more plausible anatomical structures, benefiting from the layout guidance and interleaved layout-intensity refinement. It demonstrates that LiftXR maintains cross-domain generalization potential on real radiographs. Figure 5: Generalization analysis on real frontal and lateral X-rays on the MIMIC dataset. LiftXR synthesizes more plausible organ textures while reducing boundary distortions. Method [0∘0 , 88∘88 ] [0∘0 , 92∘92 ] PSNR (dB)↑ Δ SSIM (%)↑ Δ PSNR (dB)↑ Δ SSIM (%)↑ Δ X2CT 25.49 0.47↓ 70.79 1.09↓ 25.61 0.35↓ 70.92 0.96↓ DSDF 25.20 0.62↓ 72.19 1.10↓ 25.18 0.64↓ 72.12 1.17↓ LiftXR 26.30 0.15↓ 76.14 0.30↓ 26.27 0.18↓ 75.94 0.50↓ Table 4: Robustness evaluation under different angle deviations on the CT-RATE dataset. The values report the reconstruction performance and the corresponding degradation (Δ ) compared with the original acquisition setting. Robustness to Projection Angle Variations: To further validate the robustness of LiftXR under realistic acquisition variations, we investigate the impact of projection angle deviations from the standard orthogonal bi-planar setting. In clinical scenarios, frontal and lateral X-rays may not be perfectly aligned due to patient positioning and acquisition conditions, introducing geometric inconsistencies between the observed projections and the ideal reconstruction configuration. As shown in Tab. 4, we perturb the projection angles to [0∘,88∘][0 ,88 ] and [0∘,92∘][0 ,92 ] while using the same reconstruction pipeline. LiftXR consistently achieves the best reconstruction performance under both perturbed settings, while exhibiting the smallest performance degradation compared with existing methods. This robustness to small projection perturbations is attributed to the proposed interleaved layout-intensity refinement, where anatomical layouts provide explicit geometric constraints and reconstructed CT volumes further refine the estimated layouts. Such reciprocal interaction enables LiftXR to partially compensate for geometric deviations and maintain reliable volumetric reconstruction under practical and non-ideal projection scenarios. Limitation and Future Work: While LiftXR demonstrates promising performance in bi-planar X-ray-to-CT reconstruction, several limitations remain. First, reconstructed CT volumes are currently assessed through image quality and anatomical segmentation metrics, while their effectiveness for clinical applications requires further exploration. Second, the interleaved refinement process introduces additional computational costs, which may limit the deployment to specific applications such as mobile scenarios. Furthermore, bi-planar X-ray-to-CT reconstruction remains inherently constrained by limited information, making the recovery of highly variable or small pathological anatomical structures challenging. Future work will investigate more efficient architectures and conduct broader clinical evaluations to improve the generalizability and applicability of LiftXR. Conclusion We present LiftXR, an interleaved geometry-guided framework for bi-planar X-ray-to-CT reconstruction. LiftXR transitions from projection-conditioned anatomical layout generation to reconstruction-conditioned anatomical perception, allowing the reconstructed volume to refine its geometry and the perceived anatomy to further calibrate region-specific intensities. This generation-perception coupling reduces geometric ambiguity and enables more anatomically consistent volumetric reconstruction. Experiments on two public chest CT datasets demonstrate state-of-the-art reconstruction performance and impressive downstream segmentation results, indicating superior anatomical fidelity. References Armanious et al. (2020) K. Armanious, C. Jiang, M. Fischer, T. Küstner, T. Hepp, K. Nikolaou, S. Gatidis, and B. Yang MedGAN: medical image translation using gans. Computerized Medical Imaging and Graphics 79, p. 101684. Cited by: Medical Image Generation. Armato I et al. (2011) S. G. Armato I, G. McLennan, L. Bidaut, M. F. McNitt-Gray, C. R. Meyer, A. P. Reeves, B. Zhao, D. R. Aberle, C. I. Henschke, E. A. Hoffman, et al. The lung image database consortium (lidc) and image database resource initiative (idri): a completed reference database of lung nodules on ct scans. Medical Physics 38 (2), p. 915–931. Cited by: Dataset and Pre-processing:. Cai et al. (2024) Y. Cai, J. Wang, A. Yuille, Z. Zhou, and A. Wang Structure-aware sparse-view x-ray 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11174–11183. Cited by: Introduction. Chen et al. (2025a) H. Chen, Z. Hao, L. Guo, and L. Xiao Mitigating data consistency induced discrepancy in cascaded diffusion models for sparse-view ct reconstruction. IEEE Transactions on Medical Imaging 44 (7), p. 3012–3024. Cited by: Introduction. Chen et al. (2018) H. Chen, Y. Zhang, Y. Chen, J. Zhang, W. Zhang, H. Sun, Y. Lv, P. Liao, J. Zhou, and G. Wang LEARN: learned experts’ assessment-based reconstruction network for sparse-data ct. IEEE Transactions on Medical Imaging 37 (6), p. 1333–1347. Cited by: Sparse-view CT Reconstruction. Chen et al. (2025b) J. Chen, Y. Lin, Y. Qin, H. Wang, and X. Li Cross-view generalized diffusion model for sparse-view ct reconstruction. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 140–150. Cited by: Medical Image Generation. Chung and Ye (2022) H. Chung and J. C. Ye Score-based diffusion models for accelerated mri. Medical Image Analysis 80, p. 102479. Cited by: Medical Image Generation. Deo et al. (2025) Y. Deo, Y. Jia, T. Lassila, W. A. Smith, T. Lawton, S. Kang, A. F. Frangi, and I. Habli Metrics that matter: evaluating image quality metrics for medical image generation. arXiv preprint arXiv:2505.07175. Cited by: Introduction. Dong et al. (2019) X. Dong, S. Vekhande, and G. Cao Sinogram interpolation for sparse-view micro-ct with deep learning neural network. In Medical Imaging 2019: Physics of Medical Imaging, Vol. 10948, p. 692–698. Cited by: Sparse-view CT Reconstruction. Feldkamp et al. (1984) L. A. Feldkamp, L. C. Davis, and J. W. Kress Practical cone-beam algorithm. Journal of the Optical Society of America A 1 (6), p. 612–619. Cited by: Sparse-view CT Reconstruction. Goodfellow et al. (2014) I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio Generative adversarial nets. Advances in Neural Information Processing Systems 27. Cited by: Medical Image Generation. Gopalakrishnan and Golland (2022) V. Gopalakrishnan and P. Golland Fast auto-differentiable digitally reconstructed radiographs for solving inverse problems in intraoperative imaging. In Workshop on Clinical Image-Based Procedures, p. 1–11. Cited by: Dataset and Pre-processing:. Guo et al. (2024) T. Guo, Y. Liu, P. Zhang, Y. Liu, and Z. Gui MAIR-net: a sparse-view ct reconstruction network based on a combination of mixed attention and iterative optimization learning. Journal of Instrumentation 19 (08), p. P08029. Cited by: Introduction. Hamamci et al. (2026) I. E. Hamamci, S. Er, C. Wang, F. Almas, A. G. Simsek, S. N. Esirgun, I. Dogan, O. F. Durugol, B. Hou, S. Shit, et al. Generalist foundation models from a multimodal dataset for 3d computed tomography. Nature Biomedical Engineering, p. 1–19. Cited by: Introduction, Dataset and Pre-processing:. Ho et al. (2020) J. Ho, A. Jain, and P. Abbeel Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems 33, p. 6840–6851. Cited by: Medical Image Generation. Isola et al. (2017) P. Isola, J. Zhu, T. Zhou, and A. A. Efros Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 1125–1134. Cited by: Medical Image Generation. Jeong et al. (2025) Y. S. Jeong, H. B. Yoo, and I. Y. Chun DX2CT: diffusion model for 3d ct reconstruction from bi or mono-planar 2d x-ray (s). In International Conference on Acoustics, Speech and Signal Processing, p. 1–5. Cited by: Table 1. Jiang et al. (2025) H. Jiang, M. Imran, T. Zhang, Y. Zhou, M. Liang, K. Gong, and W. Shao Fast-ddpm: fast denoising diffusion probabilistic models for medical image-to-image generation. IEEE Journal of Biomedical and Health Informatics 29 (10), p. 7326–7335. Cited by: Medical Image Generation. Jin et al. (2017) K. H. Jin, M. T. McCann, E. Froustey, and M. Unser Deep convolutional neural network for inverse problems in imaging. IEEE Transactions on Image Processing 26 (9), p. 4509–4522. Cited by: Sparse-view CT Reconstruction. Johnson et al. (2019) A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C. Deng, R. G. Mark, and S. Horng MIMIC-cxr, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 6 (1), p. 317. Cited by: Generalization to Real X-ray Images:. Jung (2021) H. Jung Basic physical principles and clinical applications of computed tomography. Progress in Medical Physics 32 (1), p. 1–17. Cited by: Introduction. Kyung et al. (2023) D. Kyung, K. Jo, J. Choo, J. Lee, and E. Choi Perspective projection-based 3d ct reconstruction from biplanar x-rays. In IEEE International Conference on Acoustics, Speech and Signal Processing, p. 1–5. Cited by: Introduction, X-ray-to-CT Reconstruction, Table 1, Generalization to Real X-ray Images:. Lee et al. (2018) H. Lee, J. Lee, H. Kim, B. Cho, and S. Cho Deep-neural-network-based sinogram synthesis for sparse-view ct image reconstruction. IEEE Transactions on Radiation and Plasma Medical Sciences 3 (2), p. 109–119. Cited by: Sparse-view CT Reconstruction. Lin et al. (2025) T. Lin, X. Li, C. Zhuang, Q. Chen, Y. Cai, K. Ding, A. L. Yuille, and Z. Zhou Are pixel-wise metrics reliable for sparse-view computed tomography reconstruction?. arXiv preprint arXiv:2506.02093. Cited by: Introduction, X-ray-to-CT Reconstruction. Lin et al. (2019) W. Lin, H. Liao, C. Peng, X. Sun, J. Zhang, J. Luo, R. Chellappa, and S. K. Zhou DuDoNet: dual domain network for ct metal artifact reduction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10512–10521. Cited by: Sparse-view CT Reconstruction. Lin et al. (2023) Y. Lin, Z. Luo, W. Zhao, and X. Li Learning deep intensity field for extremely sparse-view cbct reconstruction. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 13–23. Cited by: Introduction. Lin et al. (2026) Y. Lin, H. Sun, Y. Li, R. Aslam, L. F. Tse, T. Cheng, C. S. Chui, W. F. Yau, V. R. Le Meur, M. Amangeldy, et al. Real-time reconstruction of 3d bone models via very-low-dose protocols. npj Digital Medicine 9 (1), p. 353. Cited by: Medical Image Generation. Lin et al. (2024) Y. Lin, H. Wang, J. Chen, and X. Li Learning 3d gaussians for extremely sparse-view cone-beam ct reconstruction. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 425–435. Cited by: Introduction. Liu and Ding (2023) W. Liu and H. Ding Solving low-dose ct reconstruction via gan with local coherence. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 524–534. Cited by: Medical Image Generation. Liu et al. (2024a) X. Liu, Z. Qiao, R. Liu, H. Li, J. Zhang, X. Zhen, Z. Qian, and B. Zhang DiffuX2CT: diffusion learning to reconstruct ct images from biplanar x-rays. In European Conference on Computer Vision, p. 458–476. Cited by: X-ray-to-CT Reconstruction, Table 1. Liu et al. (2024b) Z. Liu, Y. Fang, C. Li, H. Wu, Y. Liu, D. Shen, and Z. Cui Geometry-aware attenuation learning for sparse-view cbct reconstruction. IEEE Transactions on Medical Imaging 44 (2), p. 1083–1097. Cited by: Table 1. Niu et al. (2014) S. Niu, Y. Gao, Z. Bian, J. Huang, W. Chen, G. Yu, Z. Liang, and J. Ma Sparse-view x-ray ct reconstruction via total generalized variation regularization. Physics in Medicine and Biology 59 (12), p. 2997–3017. Cited by: Sparse-view CT Reconstruction. Pan et al. (2025) Y. Pan, Y. Ye, Y. Zhang, Y. Xia, and D. Shen Draw sketch, draw flesh: whole-body computed tomography from any x-ray views. International Journal of Computer Vision 133 (5), p. 2505–2526. Cited by: Introduction, X-ray-to-CT Reconstruction, CT Reconstruction Supervision:, Table 1, Effect of Layout Selection:. Pan et al. (2026) Z. Pan, H. Zhou, Q. Ren, X. Hou, J. Dai, X. Liu, Y. Che, X. Zhao, Y. Xie, Z. Li, et al. X2Shape: ct-free 3d multi-organ reconstruction with biplanar x-rays. Medical Image Analysis, p. 104085. Cited by: Medical Image Generation. Shi et al. (2026) J. Shi, D. M. Pelt, and K. J. Batenburg DM4CT: benchmarking diffusion models for computed tomography reconstruction. arXiv preprint arXiv:2602.18589. Cited by: Introduction. Shi et al. (2025) W. Shi, Y. Hu, Y. Sun, G. Chang, Y. Yang, Y. Song, H. Qian, Z. Wei, L. Zhao, M. Li, et al. Exploration of optimal thresholds for predicting the invasive nature of stage t1 lung adenocarcinoma using artificial intelligence-based 3d solid component volume segmentation. Quantitative Imaging in Medicine and Surgery 15 (1), p. 249–258. Cited by: Introduction. Song et al. (2026) T. Song, Y. Wu, M. Hu, X. Luo, L. Wei, G. Wang, Y. Guo, F. Xu, and S. Zhang Learning modality-aware representations: adaptive group-wise interaction network for multimodal mri synthesis. IEEE Transactions on Medical Imaging 45 (5), p. 2306–2316. Cited by: Medical Image Generation. Wang et al. (2017) X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers ChestX-ray8: hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 2097–2106. Cited by: Introduction. Wasserthal et al. (2023) J. Wasserthal, H. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang, et al. TotalSegmentator: robust segmentation of 104 anatomic structures in ct images. Radiology: Artificial Intelligence 5 (5), p. e230024. Cited by: Dataset and Pre-processing:. Webber and Reader (2024) G. Webber and A. J. Reader Diffusion models for medical image reconstruction. BJR| Artificial Intelligence 1 (1), p. ubae013. Cited by: Medical Image Generation. Wu et al. (2025) J. Wu, J. Lin, X. Jiang, W. Zheng, L. Zhong, Y. Pang, H. Meng, and Z. Li Dual-domain deep prior guided sparse-view ct reconstruction with multi-scale fusion attention. Scientific Reports 15 (1), p. 16894. Cited by: Introduction. Wu et al. (2024) Y. Wu, X. Luo, Z. Xu, X. Guo, L. Ju, Z. Ge, W. Liao, and J. Cai Diversified and personalized multi-rater medical image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11470–11479. Cited by: Medical Image Generation. Wu et al. (2026) Y. Wu, T. Song, Z. Wu, J. Ye, Z. Ge, W. Bai, Z. Chen, and J. Cai Virtual full-stack scanning of brain mri via imputing any quantised code. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 21026–21035. Cited by: Medical Image Generation. Yang et al. (2017) G. Yang, S. Yu, H. Dong, G. Slabaugh, P. L. Dragotti, X. Ye, F. Liu, S. Arridge, J. Keegan, Y. Guo, et al. DAGAN: deep de-aliasing generative adversarial networks for fast compressed sensing mri reconstruction. IEEE Transactions on Medical Imaging 37 (6), p. 1310–1321. Cited by: Medical Image Generation. Yang et al. (2025) L. Yang, J. Huang, G. Yang, and D. Zhang CT-sdm: a sampling diffusion model for sparse-view ct reconstruction across various sampling rates. IEEE Transactions on Medical Imaging 44 (6), p. 2581–2593. Cited by: Sparse-view CT Reconstruction. Ying et al. (2019) X. Ying, H. Guo, K. Ma, J. Wu, Z. Weng, and Y. Zheng X2CT-gan: reconstructing ct from biplanar x-rays with generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10619–10628. Cited by: Introduction, X-ray-to-CT Reconstruction, CT Reconstruction Supervision:, CT Reconstruction Supervision:, Table 1, Dataset and Pre-processing:, Generalization to Real X-ray Images:. Yu et al. (2025) T. Yu, X. Tian, J. Yang, D. He, J. Yu, X. Wang, and Y. Zhang SPIDER: structure-preferential implicit deep network for biplanar x-ray reconstruction. arXiv preprint arXiv:2507.04684. Cited by: Introduction. Zhu et al. (2017) J. Zhu, T. Park, P. Isola, and A. A. Efros Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision, p. 2223–2232. Cited by: Generalization to Real X-ray Images:.