Source-linked AI summary
Reliability of PET/CT shape and heterogeneity features in functional and morphological components of Non-Small Cell Lung Cancer tumors: a repeatability analysis in a prospective multi-center cohort
Marie-Charlotte Desseroit, Florent Tixier, Wolfgang Weber, Barry A Siegel, Catherine Cheze Le Rest, Dimitris Visvikis, Mathieu Hatt
TL;DR
PET/CT shape and heterogeneity features may be sensitive to acquisition and analysis conditions, making their repeatability important for longitudinal use. This study evaluates test-retest reliability across PET and low-dose CT and finds that repeatability varies substantially by feature and quantization method.
Problem
PET/CT shape and heterogeneity features can be sensitive to image noise, segmentation, and reconstruction settings, while their repeatability matters for longitudinal analyses.
Method
The study analyzed prospectively acquired multicenter PET/CT test-retest datasets to assess feature repeatability across PET and low-dose CT using two image-quantization strategies.
Results
Repeatability varied greatly among features, and quantization affected modalities differently, with quantizationW producing worse PET but better low-dose CT repeatability.
Takeaways & Limitations
PET and low-dose CT features should be selected with careful attention to metric-specific repeatability and modality-specific quantization choices.
Takeaways & Limitations
The authors identify limitations in their analysis of low-dose CT and PET images.
Abstract
from arXiv · showhide
Purpose: The main purpose of this study was to assess the reliability of shape and heterogeneity features in both Positron Emission Tomography (PET) and low-dose Computed Tomography (CT) components of PET/CT. A secondary objective was to investigate the impact of image quantization.Material and methods: A Health Insurance Portability and Accountability Act -compliant secondary analysis of deidentified prospectively acquired PET/CT test-retest datasets of 74 patients from multi-center Merck and ACRIN trials was performed. Metabolically active volumes were automatically delineated on PET with Fuzzy Locally Adaptive Bayesian algorithm. 3DSlicerTM was used to semi-automatically delineate the anatomical volumes on low-dose CT components. Two quantization methods were considered: a quantization into a set number of bins (quantizationB) and an alternative quantization with bins of fixed width (quantizationW). Four shape descriptors, ten first-order metrics and 26 textural features were computed. Bland-Altman analysis was used to quantify repeatability. Features were subsequently categorized as very reliable, reliable, moderately reliable and poorly reliable with respect to the corresponding volume variability. Results: Repeatability was highly variable amongst features. Numerous metrics were identified as poorly or moderately reliable. Others were (very) reliable in both modalities, and in all categories (shape, 1st-, 2nd- and 3rd-order metrics). Image quantization played a major role in the features repeatability. Features were more reliable in PET with quantizationB, whereas quantizationW showed better results in CT.Conclusion: The test-retest repeatability of shape and heterogeneity features in PET and low-dose CT varied greatly amongst metrics. The level of repeatability also depended strongly on the quantization step, with different optimal choices for each modality. The repeatability of PET and low-dose CT features should be carefully taken into account when selecting metrics to build multiparametric models.
INTRODUCTION
Radiomics enables quantitative characterization of NSCLC tumors from PET and CT, but feature sensitivity to imaging conditions and uncertain repeatability complicate longitudinal use. This study therefore evaluated repeatability across both PET and low-dose CT in a large prospective multi-center cohort and examined image quantization.
- Radiomics extracts intensity, shape, and heterogeneity features from medical images and may offer greater value than standard metrics while enabling combined PET and low-dose CT analysis.
- Many radiomic features are sensitive to image noise, segmentation, and reconstruction settings, creating challenges for therapy-response monitoring and early prediction.
- Prior repeatability studies used small single-center cohorts and did not report repeatability for low-dose CT features from PET/CT, limiting assessment of combined-component models.
- A secondary goal was to evaluate the impact of the image quantization step on textural-feature analysis and repeatability.
- The primary goal was to evaluate repeatability of shape and heterogeneity metrics from both PET and low-dose CT components in a large prospective multi-center cohort.
MATERIALS AND METHODS · Patient cohort and imaging · PET and CT analysis
This secondary analysis used prospectively acquired, deidentified PET/CT test-retest data from 74 patients with stage IIIB-IV NSCLC across two multicenter trials. PET and low-dose CT tumor volumes were independently delineated, then shape, first-order, and texture features were computed using fixed-bin and fixed-width intensity quantization approaches.
- Patient cohort and imaging: The cohort comprised 74 patients enrolled across the Merck MK-0646-008 and ACRIN 6678 multicenter trials.The trials included 40 patients at 17 sites and 34 patients at 14 sites, respectively.
- Patient cohort and imaging: The present analysis used deidentified PET/CT images and extended prior SUV-only analyses by computing texture features and shape parameters on both PET and CT.The analysis was approved by ACRIN and conducted in compliance with HIPAA.
- PET and CT analysis: PET and low-dose CT images were processed independently, with PET metabolically active volumes segmented using the Fuzzy Locally Adaptive Bayesian algorithm.The PET segmentation covered the primary tumor and up to three additional lesions.
- PET and CT analysis: Low-dose CT anatomical volumes of primary tumors were delineated using a validated semi-automatic approach in 3D SlicerTM.Additional lesions were analyzed when they could be reliably delineated.
- PET and CT analysis: The feature set included 3D shape descriptors, first-order intensity and histogram metrics, and second- and third-order texture metrics.Examples included sphericity, irregularity, maximum and mean Hounsfield units or SUV, skewness, kurtosis, energy, entropyHIST, CHAUC, GLCM, NGTDM, and grey-level zone size matrix features.
- PET and CT analysis: Quantization was applied before constructing texture matrices, whereas first-order metrics did not require this resampling step.The texture matrices used all 13 orientations simultaneously.
Statistical analysis
Repeatability was assessed primarily with Bland–Altman analysis, using differences between paired measurements and repeatability limits based on 1.96×SD. Metrics were also correlated with Spearman coefficients and classified according to variability relative to VOI repeatability.
- Bland–Altman analysis assessed each metric’s repeatability by reporting the mean and standard deviation of differences between the two measurements.
- Repeatability limits were calculated as ±1.96×SD, after log-transformation when distributions were non-normal.
- Correlations between metrics were evaluated using Spearman rank coefficients (r_s), while intra-class correlation coefficients were provided in supplementary material.
- Metrics were categorized as very reliable, reliable, moderately reliable, or poorly reliable according to thresholds of 0.5×, 1.5×, and 2× VOIrepSD.
RESULTS
The analysis included 73 datasets, with PET measurements from 73 primary tumors and 32 additional lesions. Low-dose CT analysis excluded two patients, leaving 71 primary tumors and five additional lesions.
- 73 datasets were analyzed because one dataset was unavailable.
- PET analysis included 73 primary tumors and 32 additional nodal or distant metastatic lesions.Mean MAV was 47.8 cm3, with a median of 24.9 cm3 and SD of 55.4 cm3.
- Two patients were excluded from low-dose CT analysis because repeatable volume delineation could not be ensured.
- Low-dose CT analysis included 71 primary tumors and five additional lesions.Mean AV was 52.4 cm3, with a median of 37.5 cm3 and SD of 53.0 cm3.
PET and low-dose CT volumes · PET features · Shape descriptors and 1st-order metrics
PET and low-dose CT volume determinations showed repeatability around zero, while PET shape features were generally highly repeatable and intensity-based features varied substantially. The most repeatable PET first-order metrics were CHAUC and entropyHIST, whereas energy and skewness were least repeatable.
- PET and low-dose CT volumes: MAV repeatability was -1.4±11.1%, with upper and lower limits of +20.3% and -23.2%; smaller volumes were less repeatable.The volume association was rs=-0.41, p<0.0001.
- PET and low-dose CT volumes: AV repeatability was -0.4±10.5%, with upper and lower limits of +20.3% and -21.0%, and weaker volume dependence.The volume association was rs=-0.32, p=0.006.
- Shape descriptors and 1st-order metrics: PET shape features were very repeatable overall: irregularity and sphericity had only 4.8% SD.3D surface and major axis were reliable but more variable, at 9.0% and 8.4%, respectively.
- PET features: CHAUC and entropyHIST were the most repeatable PET intensity-based first-order features, each showing -0.2 ± 3.6%.These metrics were more repeatable than the other reported first-order features.
- PET features: SUVmean and SUVmax were moderately reliable, with repeatability limits of -30.4% to 36.3% and -34.3% to 41.3%, respectively.The reported limits indicate greater variability for SUVmax than SUVmean.
2nd-order metrics · 3nd-order metrics
Second-order feature repeatability varied substantially: several GLCM features were most repeatable with quantizationB, while quantizationW reduced reliability and increased outliers. For third-order metrics, quantizationB yielded reliable or very reliable grey-level zone-size features, whereas quantizationW classified all such features as poorly reliable.
- 2nd-order metrics: With quantizationB, entropyGLCM (-0.1 ± 2.6%), sum entropy (-0.2 ± 2.1%) and difference entropy (-0.2 ± 3.0%) were the most repeatable GLCM features.
- 2nd-order metrics: Most other GLCM features were reliable, while five were moderately reliable and three were unreliable.
- 2nd-order metrics: Correlation appeared very poorly repeatable because Bland-Altman analysis was sensitive to a few outliers near zero; excluding them yielded reproducibility limits below ±20%.After excluding the outliers, correlation could be re-categorized as moderately reliable.
- 2nd-order metrics: The five NGTDM features were less repeatable than the best GLCM features but remained reliable, with SD ~14-17% except contrastNGTDM (27.6%).
- 2nd-order metrics: QuantizationW changed the feature hierarchy and produced much lower reliability, with notably more outliers and higher variability overall.
- 3nd-order metrics: For third-order metrics, all grey-level zone size matrix features were poorly reliable with quantizationW, whereas quantizationB classified two as very reliable and three as reliable.With quantizationB, small zone size emphasis and zone size percentage had SD <4%, while large zone size emphasis, gray-level non-uniformity and zone size non-uniformity had SD ~11-14%.
- 3nd-order metrics: Among the least repeatable third-order features were those focusing on small zones and/or low grey values, including LZLGE, SZLGE and LGLZE.
Low-dose CT features … 3nd-order metrics
Low-dose CT feature repeatability varied substantially across shape, first-order, second-order, and third-order metrics. Quantization strongly influenced reliability, with quantizationw generally outperforming quantizationB for second- and third-order features.
- Shape descriptors and 1st-order metrics: Morphological irregularity, sphericity, and 3D surface were the most repeatable shape descriptors, with SDs of 3.3%, 10.0%, and 11.6%.Major axis was less reliable at 3.8 ± 18.4%.
- Shape descriptors and 1st-order metrics: Maximum, mean intensity, kurtosis, and skewness showed poor reliability, with variability of 4.7 ± 38.6%, -4.2 ± 43.6%, 4.8 ± 37.4%, and 11.1 ± 202.2%.The reported metrics were maximum, mean intensity, kurtosis, and skewness.
- Shape descriptors and 1st-order metrics: EntropyHIST and CHAUC were very reliable, with repeatability values of -0.1 ± 2.5% and 0.7 ± 9.1%.These metrics contrasted with the poorly reliable histogram measures.
- 2nd-order metrics: Quantizationw improved second-order repeatability compared with quantizationB; entropyGLCM, sum entropy, and difference entropy were most repeatable under both methods.Their quantizationB versus quantizationw values were -1.9 ± 12.0% vs. -0.4 ± 5.2%, -1.4 ± 10.0% vs. 0.1 ± 0.4%, and -2.3 ± 13.1% vs. -0.3 ± 1.9%, respectively.
- 2nd-order metrics: NGTDM also showed higher repeatability with quantizationw, while Complexity was the only parameter categorized as reliable under both methods.Complexity measured 0.5 ± 14.3% with quantizationB and -0.5 ± 12.3% with quantizationw.
- 3nd-order metrics: Eight third-order parameters were moderately reliable or better with quantizationw, compared with only two using quantizationB.The quantization method therefore had an important impact on third-order feature reliability.
- 3nd-order metrics: Small zone size emphasis and zone size emphasis were the most repeatable third-order features.Their quantizationB versus quantizationw values were -0.6 ± 4.8% vs. -0.5 ± 2.6% and -2.8 ± 17.4% vs. -0.9 ± 11.9%, respectively.
Impact of quantization method
Quantization affected repeatability in opposite ways in PET and low-dose CT because features tracked different combinations of volume and maximum intensity. In PET, quantizationB improved repeatability through stronger association with MAV, whereas in CT it worsened repeatability because maximum intensity was less repeatable than volume.
- Impact of quantization method: PET features from quantizationW correlated with SUVmax rather than MAV, whereas quantizationB features correlated with MAV rather than SUVmax.
- Impact of quantization method: QuantizationB produced higher PET repeatability because MAV was much more repeatable than SUVmax.
- Impact of quantization method: In low-dose CT, quantizationB features correlated with both volume and maximum intensity, while quantizationW features were less or not correlated with either.
- Impact of quantization method: QuantizationB worsened low-dose CT repeatability because maximum intensity was much less repeatable than volume.
DISCUSSION
Shape descriptors were generally reliable in PET and low-dose CT, whereas repeatability varied substantially across first-order and higher-order heterogeneity features. Quantization had modality-specific effects, underscoring the need to account for repeatability and segmentation limitations when selecting features for PET/CT models.
- Feature reliability: Shape descriptors were reliable, sometimes highly repeatable, in both PET and low-dose CT, consistent with repeatable segmentation.The study was the first, to the authors’ knowledge, to report repeatability of these features in low-dose CT.
- Limitations: Repeatability may be lower clinically because the evaluation included segmentation variability and used only one expert, while less accurate segmentation could particularly affect volume-correlated features.The images were analyzed separately with independent test and retest segmentations.
- Feature reliability: Repeatability varied greatly among first-order and textural features; some were unreliable in both modalities, while others were reliable across all categories.Examples to avoid included 1st-order skewness, 2nd-order Angular Second Moment, contrastGLCM, contrastNGTDM, and 3rd-order metrics quantifying low grey values or small zones.
- Quantization: Quantization had modality-specific effects: quantizationW worsened repeatability in PET but improved it in low-dose CT.B=64 was described as a good compromise, although repeatability of some metrics depended on B.
- Model development: PET/CT clinical models should carefully account for repeatability, especially when tracking features across therapy or seeking generalization to external cohorts.Feature selection should also consider discriminative power, robustness, and redundancy.
- Limitations: Without respiratory gating, motion may introduce quantitative bias between test and retest images and between PET and low-dose CT, making reported repeatability less broadly applicable.The authors note that motion correction or less motion-prone body regions could yield different repeatability.
CONCLUSION
Test-retest repeatability of PET/CT shape and heterogeneity features varied greatly among metrics and depended on the quantization step, with different optimal choices for PET and low-dose CT. Repeatability should therefore be carefully considered when selecting metrics for multiparametric models.
- Repeatability of shape and heterogeneity features varied greatly among PET/CT metrics.
- Quantization strongly affected repeatability, with different optimal choices for PET and low-dose CT.The difference reflected distinct relationships between metrics and volume or intensity.
- Repeatability should be carefully considered when selecting metrics to combine in multiparametric models.
FIGURE CAPTIONS
The figures depict repeatability rankings for shape and intensity-based metrics across FDG PET and low-dose CT, including comparisons of quantization approaches. They also illustrate relationships between textural features and tumor volume or maximum intensity.
- Volume repeatability: Figure 1 presents Bland–Altman analysis and correlations between volume and repeatability for MAV and AV determination.The caption identifies both volume measures as part of the repeatability analysis.
- 1st-order and shape features: Figure 2 ranks 1st-order metrics and 3D shape descriptors by repeatability in FDG PET and low-dose CT, from highest to lowest.Features are categorized as very reliable, reliable, moderately reliable, or poorly reliable.
- Higher-order metrics: Figures 3 and 4 rank 2nd- and 3rd-order metrics by repeatability for FDG PET and low-dose CT under quantizationB or quantizationW.The figures distinguish very reliable, reliable, moderately reliable, and poorly reliable features.
- Correlative relationships: Figure 5 illustrates correlations between GLCM dissimilarity and either volume or maximum intensity in PET and low-dose CT, depending on quantization approach.PET and low-dose CT are shown separately, with volume relationships in the first row and maximum-intensity relationships in the second.