Source-linked AI summary
THA-Flow Generative Model: Prosthesis Geometry Prediction from Preoperative CT
Yiping Wang, Jie Li, Jingyu Shen, Liao Wang
TL;DR
THA planning often treats a one-to-many clinical problem as selection of a single prosthesis configuration. THA-Flow addresses this by generating three-dimensional prosthesis geometry from preoperative CT with bone conditioning and optional prosthesis controls, producing anatomically matched variation across major stem models while retaining a defined scope and known limitations.
Problem
THA planning is inherently one-to-many because the same osseous anatomy may admit several clinically reasonable prosthesis configurations and placements.
Method
THA-Flow uses conditional rectified-flow generation in a prosthesis latent space, conditioned spatially on preoperative bone and optionally on stem model and size.
Results
THA-Flow generated complete acetabular and femoral geometries across seven major stem models representing 93.4% of the cohort, with repeated samples preserving principal interfaces while allowing limited local variation.
Takeaways & Limitations
The model represents patient-matched conditional prosthesis geometry rather than a fixed template and supports product-conditioned variation while preserving anatomical fit.
Takeaways & Limitations
The single-centre cohort may limit generalisability, and strong bone conditioning can produce anatomically adapted deformations rather than faithful catalogue geometry.
Abstract
from arXiv · showhide
Preoperative planning for total hip arthroplasty (THA) is commonly framed as selecting a single prosthesis configuration and placement for a patient's osseous anatomy. In practice, however, the same anatomy may admit several clinically reasonable solutions, making planning inherently a one-to-many problem that is better represented by a conditional probability distribution. We present THA-Flow, a conditional flow-matching model that generates three-dimensional prosthesis geometry directly from preoperative CT. Separate AutoencoderKL models compress preoperative bone anatomy and prosthesis geometry, while a three-dimensional UNet learns a rectified flow from Gaussian noise to the prosthesis latent space under spatial bone conditioning and optional structured prosthesis parameters. The retrospective cohort comprised 1,355 hips from 1,149 patients undergoing primary THA. Following rigid registration of postoperative CT to preoperative CT, the actual postoperative prostheses were transformed independently according to the pelvic and femoral registrations and represented as a dual-channel truncated signed distance field. The prosthesis autoencoder achieved a peak signal-to-noise ratio of 47.11 dB and a structural similarity index of 0.9964 on the validation set. Complete acetabular and femoral geometries were generated across seven major stem models representing 93.4% of the cohort. Repeated bone-conditioned sampling preserved component position, alignment, and the principal bone-prosthesis interfaces while allowing limited local geometric variation. To our knowledge, THA-Flow represents the first application of generative AI to three-dimensional surgical planning for THA.
1 Introduction
THA planning is typically reduced to selecting one prosthesis configuration from measurements, rules, and predefined inventories, although the same anatomy may support multiple plausible plans. THA-Flow instead frames planning as conditional generation of three-dimensional prosthesis geometry while preserving the patient’s original bone anatomy.
- Existing planning workflows: Current THA workflows use measurements, heuristic rules, and commercial prosthesis libraries to select shape, size, and placement.AI-based systems improve segmentation and landmark detection but retain downstream rule-based selection from predefined inventories.
- Motivation for generation: Single-configuration planning does not represent the distribution of plausible prosthesis plans for the same anatomy.Generative modeling permits repeated sampling from a conditional distribution.
- Prior generative approach: Prior generative THA work synthesized two-dimensional postoperative radiographs rather than complete three-dimensional prosthesis geometry.Generating the whole postoperative image also redraws bone, soft tissue, and the prosthesis interface.
- Proposed formulation: THA-Flow conditions three-dimensional prosthesis generation on preoperative bone and restricts the target to an implicit prosthesis representation.This allows the output to be superimposed on unaltered patient anatomy and converted into a surface for planning.
- Technical formulation: The method addresses spatially consistent supervision through independent pelvic and femoral registration, while latent-space generation reduces the burden of high-resolution three-dimensional volumes.Generating only prosthesis geometry avoids modeling postoperative soft tissue, metal artefacts, and acquisition-specific features.
2 Results
THA-Flow generates three-dimensional prosthesis candidates from preoperative CT, using bone anatomy as a spatial anchor and optional stem information as semantic constraints. Across stem designs and repeated samples, generated geometries preserved anatomical fit, component positioning, and alignment while allowing controlled variation.
- 2.1 Conditional Inference Architecture: The model encodes preoperative bone anatomy spatially and optional femoral stem model or size as semantic constraints before decoding prosthesis geometry in preoperative CT space.Bone-only, bone-plus-model, and bone-plus-model-and-size modes support progressively more specific candidate generation.
- 2.1 Conditional Inference Architecture: The neck discontinuity reflects independent pelvic and femoral registration rather than a generation failure, while joint generation preserves the cup–head relationship and supports leg-length correction estimation.The dual-channel representation separates the stem from the acetabular cup and femoral head while retaining a shared coordinate system.
- 2.2 Patient-Matched Generation Across Stem Designs: 1,265 hips, or 93.4% of the cohort, were covered by the seven most prevalent stem models, and both bone-only and model-conditioned modes generated complete acetabular and femoral geometries.Model conditioning constrained stem length, proximal width, and lateral profile without disrupting anatomical fit.
- 2.3 Bone–Prosthesis Fit and Fill: Generated stems selectively filled the effective canal, respected the calcar boundary, and retained regular distal geometry, while generated cups followed the acetabular wall curvature.The evaluation emphasizes region-specific interface behavior rather than coarse component centre, axis, or canal occupancy alone.
- 2.4 Conditional Distribution Consistency: Repeated bone-only sampling kept component position, alignment, and principal interfaces stable while producing limited local contour variation around the clinical solution.Non-overlapping regions formed narrow bands near the cup rim, proximal stem boundary, and stem tip; distribution-level consistency is more informative than reproducing one postoperative solution.
3 Discussion
THA-Flow reframes THA planning as conditional generation of three-dimensional prosthesis distributions rather than prediction of a single plan. Its outputs adapt to anatomy, preserve principal interfaces, and support product-conditioned variation, while remaining limited by generalisability and clinical-actionability constraints.
- THA-Flow generated complete acetabular and femoral geometry across cases representing seven major stem models.
- Separate pelvic and femoral registrations return acetabular components with the pelvis and stems with the femur in preoperative space.
- Strong spatial conditioning and truncated distance-field supervision produced selective canal fit and fill rather than indiscriminate canal occupation.
- Repeated samples preserved component position, alignment, and principal interfaces while allowing limited boundary variation, supporting conditional distribution consistency.
- Specifying a stem model modified stem length, proximal width, and lateral profile while preserving anatomical fit.
- The single-centre cohort may limit generalisability, and outputs require model or size identification plus reduction and assembly to complete a surgical plan.
4 Methods
The study used a retrospective, matched preoperative–postoperative CT cohort and a two-stage latent-generation pipeline. Separate autoencoders were pretrained and then frozen before conditional rectified-flow matching.
- Training used latent-space pretraining followed by conditional rectified-flow matching.
- Separate AutoencoderKL models learned bone and prosthesis representations, and both were frozen after reconstruction validation.
- The flow model concatenated bone latents with prosthesis path states and encoded stem model, size, and component parameters as semantic tokens.
- The retrospective study matched temporally proximate preoperative and postoperative CT examinations from the same primary THA episode without additional intervention.
- 1,355 hips from 1,149 patients formed the final cohort, split into 1,188 training, 84 validation, and 83 held-out test hips.
- Nine stem models were used for model-stratified splitting, covering 1,280 hips, while 75 hips lacked reliable or sufficiently frequent records.
4.4 Prosthesis Metadata and CT Verification
Prosthesis metadata were extracted from billing records and mapped through manufacturer catalogues, then key dimensions were verified on postoperative CT. Registration used separate pelvic and femoral transformations to recover prostheses in preoperative space.
- Billing catalogue numbers indexed stem model, size, cup, head, and liner metadata through manufacturer catalogues.
- The cohort included nine stem models and 70 model-specific sizes, with cup diameters of 40–62 mm and head diameters of 22–40 mm.
- Three-dimensional CT verification corrected a billed 54-mm cup diameter to a CT-supported value of 56 mm.
- Catalogue-derived head and liner offsets were retained, while image-derived liner offset was used only for cup-shell localisation.
- Metal artefacts required sampling bone surfaces remote from contaminated or surgically altered regions during pelvic and femoral registration.
- Separate transformation matrices accommodated independent pelvic and femoral motion, with automatic overlays followed by manual review.
4.6 Watertight Reconstruction and Dual-Channel TSDF
The prosthesis was reconstructed as a watertight dual-channel TSDF and compressed with a dedicated latent autoencoder. Validation metrics indicate high-fidelity prosthesis reconstruction, while the TSDF preserves continuous surface-distance information for generation and recovery.
- Metal voxels were converted into a closed, watertight surface defining the TSDF zero level set, with distances truncated at ±5 mm.
- Unlike binary masks, TSDFs provide continuous distance gradients on both sides of the surface while limiting remote-background loss contributions.
- The dual-channel representation separated acetabular and femoral prosthesis geometry in preoperative space.
- The three-dimensional solid prosthesis could be recovered directly from the TSDF zero level set.
- All images were resampled to 1-mm isotropic voxels, with dimensions padded to multiples of 32 and density-normalised preoperative CT inputs.
- Separate KL-regularised autoencoders compressed bone and prosthesis TSDF volumes using four latent channels and four-fold spatial downsampling.
- The prosthesis branch used Eikonal-gradient and solid-interior constraints during reconstruction training.
4.8 Bone-Conditioned Rectified Flow Matching
THA-Flow learns a single-stage rectified flow that maps Gaussian-noise prosthesis latents toward bone-conditioned prosthesis latents. The model combines spatial bone conditioning with optional structured prosthesis information.
- The model regresses the velocity from Gaussian noise toward the prosthesis latent along a linear generative path.The prosthesis latent y and bone latent x are combined through channel-wise concatenation, with τ denoting normalized generative time.
- The three-dimensional UNet receives 12 channels and predicts eight across channel levels of 96, 192, and 384.It uses two residual blocks per level and self-attention only at the deepest level.
- Bone latents enter at the first convolution, preserving spatial correspondence between anatomy and the prosthesis path state.Structured parameters are injected through deepest-level cross-attention using six 256-dimensional tokens.
- Optional structured conditioning includes stem model, model-specific size, cup outer diameter, head diameter, head offset, and liner offset.
4.9 Conditioning, Optimisation, and Inference
The generator uses classifier-free condition dropout and long-run optimization with dynamic gradient accumulation and EMA weights. Evaluation combines visual, interface, distributional, and latent-space checks while distinguishing reconstruction metrics from clinical planning quality.
- Conditioning: Six conditioning modes were sampled with equal probability, spanning unconditional generation, prosthesis-only conditioning, bone-only conditioning, and progressively richer bone–prosthesis combinations.No classifier-free guidance amplification was applied during training.
- Optimisation: Dynamic gradient accumulation adjusted to voxel count mitigated variance from physical batch size one and variable volume dimensions.EMA reduced sampling perturbations caused by individual updates during prolonged optimization.
- Optimisation: 1,000 epochs reduced mean training velocity-field MSE from 1.1524 to 0.0187, with final inference using EMA weights from epoch 999.Median validation MSE for bone-only conditioning over the final 100 epochs was 0.02099.
- Inference and evaluation: Evaluation compared bone-only and stem-model-conditioned generations across seven stem designs using projections, consecutive sections, and shared axial interface levels.Reviews examined alignment, fit and fill, calcar boundaries, cup–wall relationships, and conditional variation across initial noise.
- Inference and evaluation: L1, PSNR, SSIM, Eikonal error, and velocity-field MSE verified latent reconstruction and training convergence rather than stand-alone clinical planning quality.
Data Availability
The study data and processed data are unavailable for public release because of privacy, ethics, and institutional governance requirements.
- Patient privacy, ethics requirements, and institutional data-governance restrictions prevent public release of the raw and processed data.
Code Availability
The study provides public access to its source code through a GitHub repository.
- Source code is available at https://github.com/doidio/nonahip.
Competing Interests
Y.W. and J.S. report employment at Changzhou Jinse Medical Information Technology Co., Ltd., while J.L. and L.W. declare no competing interests.
- Y.W. and J.S. are employees of Changzhou Jinse Medical Information Technology Co., Ltd.; J.L. and L.W. declare no competing interests.