Source-linked AI summary
Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction
Yifei Wu, Yicheng Wu, Qiang Ma, Qi Chen, Renyang Gu, Xinyu Liu, Yongsheng Pan, Yong Xia
TL;DR
Reconstructing 3D CT from two X-ray projections is severely ill-posed because depth, anatomy, and intensity are entangled. LiftXR interleaves anatomical layout recovery with intensity estimation, achieving state-of-the-art reconstruction and downstream segmentation on two public chest CT datasets.
Problem
Recovering 3D CT volumes from bi-planar X-rays remains severely ill-posed because projections collapse depth and entangle anatomical layouts with intensity distributions.
Method
LiftXR interleaves projection-conditioned layout generation, conditional CT rendering, reconstruction-conditioned anatomical parsing, and region-specific intensity calibration.
Results
LiftXR achieves state-of-the-art reconstruction performance and superior downstream segmentation on two public chest CT datasets.
Takeaways & Limitations
The generation-perception coupling enables more anatomically consistent volumetric reconstruction within the bi-planar X-ray-to-CT setting.
Takeaways & Limitations
Clinical effectiveness remains unexplored, while added computational costs and limited information challenge deployment and recovery of small or highly variable pathological structures.
Abstract
from arXiv · showhide
X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.
Introduction
LiftXR addresses the severe ill-posedness of bi-planar X-ray-to-CT reconstruction by explicitly recovering 3D anatomical layout before estimating CT intensities. Its geometry-guided framework couples layout generation, intensity rendering, and anatomical perception, producing anatomically improved reconstructions.
- Motivation: CT typically requires hundreds of X-ray views, while sparse-view reconstruction uses only a few dozen to reduce radiation exposure.Dense-view acquisition limits longitudinal monitoring and large-scale screening because of increased radiation exposure.
- Challenge: Reconstructing a 3D CT volume from only two projections is severely ill-posed because projection collapses depth-dependent anatomical information into two-dimensional measurements.Each X-ray pixel records attenuation accumulated along its corresponding ray path, entangling spatial layout and intensity information.
- Challenge: Existing methods model anatomical organization implicitly or through reconstructed intensities, risking inaccurate regional extents, blurred boundaries, and inconsistent spatial relationships.Such volumes may perform favorably on voxel-wise metrics while remaining anatomically inaccurate.
- Motivation: 37.04 dB PSNR and 95.07% SSIM are achieved when conditioning reconstruction on ground-truth anatomical masks, demonstrating that layout provides strong spatial constraints.The oracle study on CT-RATE shows that localizing major body regions in 3D reduces geometric uncertainty for CT reconstruction.
- Method: LiftXR first generates a 3D anatomical layout from bi-planar X-rays, then uses it to guide intensity rendering and anatomical perception for refined region-specific calibration.The reconstructed volume supplies spatially resolved boundary and intensity cues for the anatomical parser.
- Results: LiftXR’s reconstructed CT volumes yield superior downstream segmentation results, indicating improved anatomical fidelity.The framework explicitly models layout as a spatial constraint and couples volumetric generation with anatomical perception.
Related Work
X-ray-to-CT reconstruction is highly ambiguous because 2D projections collapse volumetric information, motivating geometric, generative, and structural priors. Related work also includes sparse-view reconstruction methods and generative models for medical image reconstruction.
- X-ray-to-CT Reconstruction: X-ray-to-CT reconstruction recovers 3D CT volumes from one or two 2D projections, with X2CT demonstrating bi-planar feasibility and later work introducing geometric, generative, and structural priors.PerX2CT incorporates perspective projection modeling to improve spatial reconstruction.
- Sparse-view CT Reconstruction: Sparse-view CT methods address limited projections for reduced radiation exposure, while FBP and FDK can produce streak artifacts and structural distortions under severe undersampling.Handcrafted regularization priors alleviate ill-posedness but typically require expensive optimization.
- Generative Reconstruction Models: CNN-GANs and diffusion models are widely used for medical image reconstruction, supporting efficient inference and structural control across modalities including MRI, CT, and cross-modality translation.CNN-GANs learn supervised input–target mappings, whereas diffusion models generate images through iterative denoising.
Method
LiftXR interleaves geometry-guided anatomical layout recovery with volumetric intensity estimation for bi-planar X-ray-to-CT reconstruction. It lifts paired projections into a shared 3D representation, alternates layout generation and parsing with intensity rendering and calibration, and jointly supervises anatomical geometry and CT appearance.
- Geometry-Guided Lifting: A Layout Lifter predicts a coarse 3D anatomical layout from the X-ray volume, and an Intensity Renderer uses it to reconstruct an initial CT volume.The layout encodes organ occupancy, spatial extent, and relative arrangement as explicit geometric constraints.
- X-ray Volume Construction: LiftXR replicates and aligns PA and lateral projections in a shared volumetric coordinate system, then averages them voxel-wise into an explicit X-ray volume.This unified representation preserves ray-integrated evidence from both views but cannot recover missing depth or resolve projection ambiguity by itself.
- Interleaved Refinement: A Layout Parser refines spatially resolved organ boundaries and regional identities from the rendered CT, while an Intensity Calibrator updates region-specific intensity inconsistencies.The parser is reconstruction-conditioned, unlike the projection-conditioned Layout Lifter, enabling reciprocal layout–intensity refinement.
- Training Objectives: LiftXR jointly supervises lifted and parsed layouts with anatomical masks and rendered and calibrated CT volumes with ground-truth CT targets.Layout supervision constrains 3D geometry, while CT supervision preserves intensity and volumetric appearance.
- Training Objectives: The CT objective combines mean absolute error, multi-scale feature matching, and adversarial losses to preserve voxel intensities, structural consistency, and realistic volumetric appearance.The anatomical layout and CT reconstruction objectives are balanced by a loss weight λ.
Experiment and Result
LiftXR consistently outperforms existing methods on reconstruction and segmentation across CT-RATE and LIDC-IDRI, while qualitative and ablation analyses support its interleaved layout–intensity refinement. Additional experiments demonstrate anatomical improvements on real X-rays and under projection-angle variation, with computational and clinical limitations remaining.
- Quantitative Analysis: LiftXR achieves the best reconstruction and segmentation performance across all reported metrics on both CT-RATE and LIDC-IDRI.The comparison reports consistent effectiveness in recovering CT intensities and anatomically meaningful three-dimensional structures.
- Ablation Study: Ablations show that anatomical layout and interleaved training each improve reconstruction and segmentation, while their combination provides complementary benefits.Anchor organs constrain global anatomical configurations, whereas interleaved training corrects residual local structural and attenuation errors.
- Qualitative Analysis: Qualitative comparisons show that LiftXR more faithfully preserves anatomical structures and fine-grained details, including rib continuity and a smooth liver contour.Existing methods exhibit blurred mediastinal boundaries, missing fine structures, and inaccurate organ shapes.
- Ablation Study: Organ anchors and full anatomical labels outperform HU-based regions, whereas including all TotalSegmentator targets can introduce uncertainty and reduce performance.Hollow and fine organs may add segmentation uncertainty, making the complete target set sub-optimal.
- Effect of λ: The best performance occurs at λ = 1, indicating that reconstruction benefits from jointly optimizing anatomical geometry and CT intensity.Lower λ weakens anatomical supervision and the geometric constraints needed to recover organ structures from ambiguous projections.
- Generalization and Limitations: On real MIMIC X-rays, LiftXR is compared with X2CT and DSDF, while projection-angle experiments assess robustness to deviations from orthogonal bi-planar acquisition.The study also notes additional computational costs and the need for further clinical validation.
Conclusion
LiftXR couples projection-conditioned anatomical layout generation with reconstruction-conditioned perception to reduce geometric ambiguity and calibrate region-specific intensities. Experiments on two public chest CT datasets show state-of-the-art reconstruction and strong downstream segmentation, indicating superior anatomical fidelity.
- Conclusion: LiftXR is an interleaved, geometry-guided framework that couples layout generation with anatomical perception for bi-planar X-ray-to-CT reconstruction.The framework transitions from projection-conditioned layout generation to reconstruction-conditioned perception, allowing reconstructed geometry and perceived anatomy to guide subsequent refinement.
- Conclusion: This generation-perception coupling reduces geometric ambiguity and enables region-specific intensity calibration, producing more anatomically consistent volumetric reconstructions.The reconstructed volume refines geometry, while perceived anatomy calibrates region-specific intensities.
- Conclusion: Experiments on two public chest CT datasets demonstrate state-of-the-art reconstruction performance and impressive downstream segmentation results, indicating superior anatomical fidelity.The conclusion attributes the downstream segmentation performance to improved anatomical fidelity in the reconstructed CT volumes.