Source-linked AI summary
Assessing robustness of radiomic features by image perturbation
Alex Zwanenburg, Stefan Leger, Linda Agolli, Karoline Pilz, Esther G. C. Troost, Christian Richter, Steffen Löck
TL;DR
Radiomic feature robustness can vary with phenotype, imaging conditions, and segmentation, while test-retest imaging may be difficult to obtain. The study compared image-perturbation methods with test-retest robustness and found that chained intensity and segmentation perturbations identified robust features with few false positives.
Problem
Test-retest imaging may be unavailable, and feature robustness information cannot necessarily be transferred across phenotypes or imaging modalities.
Method
The study compared image-perturbation methods with test-retest imaging, measuring robustness using ICC and classifying features with ICC ≥0.90 as robust.
Results
Test-retest robustness was higher in NSCLC than HNSCC, while chained perturbations combining intensity changes with segmentation updates produced few false positives.
Takeaways & Limitations
Chained perturbations that alter image intensity and segmentation may be used instead of test-retest imaging to determine radiomic feature robustness.
Takeaways & Limitations
Test-retest ICC estimates were limited by usually having only two test-retest images, which may reduce precision.
Abstract
from arXiv · showhide
Image features need to be robust against differences in positioning, acquisition and segmentation to ensure reproducibility. Radiomic models that only include robust features can be used to analyse new images, whereas models with non-robust features may fail to predict the outcome of interest accurately. Test-retest imaging is recommended to assess robustness, but may not be available for the phenotype of interest. We therefore investigated 18 methods to determine feature robustness based on image perturbations. Test-retest and perturbation robustness were compared for 4032 features that were computed from the gross tumour volume in two cohorts with computed tomography imaging: I) 31 non-small-cell lung cancer (NSCLC) patients; II): 19 head-and-neck squamous cell carcinoma (HNSCC) patients. Robustness was measured using the intraclass correlation coefficient (1,1) (ICC). Features with ICC$\geq0.90$ were considered robust. The NSCLC cohort contained more robust features for test-retest imaging than the HNSCC cohort ($73.5\%$ vs. $34.0\%$). A perturbation chain consisting of noise addition, affine translation, volume growth/shrinkage and supervoxel-based contour randomisation identified the fewest false positive robust features (NSCLC: $3.3\%$; HNSCC: $10.0\%$). Thus, this perturbation chain may be used to assess feature robustness.
1 Introduction
Radiomic feature robustness matters for reproducible models, but test-retest evidence is difficult to obtain and may not transfer across phenotypes, modalities, or computational settings. The study therefore evaluates whether perturbing single images can identify features that are non-robust under test-retest imaging.
- Radiomics computes quantitative features within regions of interest to support model-based treatment decisions.
- Patient positioning, image acquisition, and segmentation affect radiomic features to varying degrees, motivating robustness assessment for model generalisability.
- Test-retest imaging identifies non-robust features by comparing the same region of interest imaged twice within minutes to days.
- Feature-robustness results may not transfer across phenotypes, modalities, voxel sizes, discretisation settings, or computational configurations.
- The study tests whether perturbations of single images can identify most features that are non-robust under test-retest imaging while minimising false positive robust features.
2 Results
The study compared test-retest and perturbation robustness for 4032 features across NSCLC and HNSCC CT cohorts using chained image and mask perturbations. Robust-feature fractions and feature-wise false positives varied by cohort and perturbation chain.
- 2.1 Comparison between NSCLC and HNSCC cohorts: The analysis included 31 NSCLC patients and 19 HNSCC patients with CT images and 4032 features computed after gross tumour volume delineation.
- 2.1 Comparison between NSCLC and HNSCC cohorts: Robustness was quantified with ICC(1,1), with features considered robust when ICC ≥0.90.
- 2.1 Comparison between NSCLC and HNSCC cohorts: 73.5% of NSCLC features versus 34.0% of HNSCC features were robust under test-retest imaging.
- 2.2 Robustness under image perturbations: Noise alone identified the highest robust-feature fractions, whereas chains including mask adaptation or contour randomisation identified fewer robust features.Noise yielded 96.6% robust features in NSCLC and 99.3% in HNSCC; the lowest NSCLC fraction was 43.6% for TVC, and the lowest HNSCC fraction was 30.7% for NTVC.
- 2.3 Feature-wise comparison of perturbation and test-retest robustness: No perturbation recovered every test-retest non-robust feature, and average false-positive fractions were 7.8% in NSCLC versus 30.3% in HNSCC.
- 2.3 Feature-wise comparison of perturbation and test-retest robustness: The lowest false-positive fractions were 2.6% for RC in NSCLC and 10.0% for NTVC in HNSCC.
3 Discussion
Chained perturbations that alter both image intensity and segmentation can assess radiomic feature robustness, while test-retest imaging remains an imperfect reference. The study also identifies methodological and scope limitations and outlines possible uses of repeated measurements in modelling.
- Findings: NTVC produced few false-positive robust features in both cohorts, while TVC, RNVC and RVC showed similar performance.The authors conclude that any of these chains may be used to assess feature robustness.
- Findings: Combining intensity-altering perturbations with volume adaptation and contour randomisation reduced false positives relative to single perturbations or simple rotation–translation combinations.Single perturbation methods, including noise addition, rotation and translation, performed poorly.
- Limitations: Test-retest ICC estimates were uncertain because only two images were usually available, with average 95% confidence-interval widths of 0.12 for NSCLC and 0.35 for HNSCC.Repeated perturbations can provide more precise ICC estimates; the NTCV chain had average widths of 0.11 in NSCLC and 0.18 and 0.17 in HNSCC.
- Limitations: Test-retest images may remain too similar when the same equipment, protocols and annotator are used, so some features classified as non-robust by perturbation may be false negatives relative to the reference.This similarity can limit how well test-retest imaging captures relevant variability.
- Limitations: The study was limited to CT, and delineation uncertainty was not directly assessed against a multiple-delineation dataset.The proposed methodology should be evaluated for modalities such as PET and MRI, and contour perturbations should be compared with multiple delineations.
- Implications: Repeated measurements could enter radiomic modelling through robust-feature selection, feature averaging, or inclusion of all perturbed values during model development.The authors note that these alternatives should be compared in future work.
4 Methods
The methods generated radiomic features from perturbed CT images and compared their ICC-based robustness across two test-retest cohorts. Image perturbations, IBSI-compliant processing, multiple feature scales and discretisation settings were systematically combined.
- Cohorts: 31 NSCLC and 19 HNSCC patients provided test-retest CT cohorts for the robustness analysis.The NSCLC cohort was publicly available, whereas the HNSCC cohort was collected in-house.
- Image processing: CT images and GTV masks were processed using IBSI-based steps including rotation, Gaussian noise addition, sub-voxel translation and interpolation.The image and mask were transformed together where applicable before feature computation.
- Perturbations: Five perturbation types were implemented: rotation, noise addition, translation, volume adaptation and contour randomisation.Rotation used an affine transformation in the axial plane, while volume adaptation changed the ROI mask and contour randomisation altered the contour.
- Perturbation chains: Eighteen chained perturbation combinations generated new distorted images, with settings, repetitions and translation directions permuted across instances.Each perturbed image was used to calculate 4032 features.
- Feature computation: 4032 features were computed from 182 IBSI features across 1–4 mm isotropic voxel spacings and multiple discretisation settings.Fixed bin size and fixed bin width discretisation algorithms were both used.
- Robustness analysis: Robustness was quantified with ICC(1,1), treating features with ICC ≥0.90 as robust.Test-retest ICCs compared the two original images, while perturbation ICCs compared images within each perturbation chain and were averaged over test and retest images.
Supplementary note 1: image acquisition parameters
CT test-retest data were collected for NSCLC and HNSCC cohorts using repeat scans acquired under different timing and acquisition conditions.
- The NSCLC cohort included a second CT scan acquired 15 minutes after the first, after patient repositioning.
- The HNSCC cohort included a second CT scan acquired within 4 days to determine PET attenuation corrections.
- Table S1 reports acquisition parameters and cohort characteristics for both image datasets.
Supplementary note 2: pre-interpolation low-pass filtering
Pre-interpolation Gaussian low-pass filtering was designed to reduce aliasing before resampling, while balancing suppression of high-frequency detail against image smoothing.
- Rationale: Down-sampling can cause aliasing, so images were low-pass filtered before interpolation to suppress high-frequency content.
- Filter design: The Gaussian filter width σ is defined per image axis because the original CT coordinate grid is typically non-uniform.
- Filter design: The smoothing parameter β ranges from 0 to 1 and specifies the Gaussian response at the Nyquist frequency.
- Evaluation: The study evaluated β values from 0.50 to 0.97, plus no low-pass filtering, using test-retest ICCs and confidence intervals.
- Visual effects: Down-sampling without interpolation caused visible artefacts, whereas wide Gaussian smoothing produced images lacking detail in both example cohorts.
- Results: In HNSCC, β = 0.93 yielded 34.0% robust features, while β = 0.90 yielded 43.0%, representing a compromise between aliasing and loss of detail.
Supplementary note 3: image features
The study computed a broad set of standardized radiomic features from tumour regions using multiple feature families, discretisation schemes and interpolation spacings.
- Feature computation: Features were extracted according to Image Biomarker Standardisation Initiative definitions.
- Feature families: The feature set included morphology, intensity-volume histogram, co-occurrence, run-length, size-zone, distance-zone, neighbourhood-difference and dependence features.
- Texture computation: Three-dimensional GLCM and GLRLM features were calculated across 13 directions and subsequently averaged.
- Discretisation: Texture-related feature families requiring discretisation used fixed bin numbers of 8, 16, 32 or 64, or fixed bin widths of 6, 12, 18 or 24 HU.
- Feature computation: A total of 4032 features were computed using four interpolation spacings and two discretisation methods with four settings each.
Supplementary note 4: Image perturbation algorithms
The perturbation algorithms model positioning, intensity noise, segmentation-volume variation and contour uncertainty through reproducible image and mask transformations.
- Implementation: The implementation used Python 3.6.1 with NumPy, SciPy, scikit-image and PyWavelets.
- Rotation: Rotation applies an affine transformation to image and mask in the axial plane, with bilinear intensity sampling and integer HU rounding.
- Noise addition: Noise perturbation estimates image noise variance and adds Gaussian noise with matching standard deviation to voxel intensities.
- Translation: Translation shifts the interpolation grid along the x, y and z axes by sub-voxel fractions using tri-linear interpolation.
- Volume adaptation: Volume adaptation mimics delineation variability by growing or shrinking the ROI through iterative dilation or erosion followed by rim voxel adjustment.
- Contour randomisation: Contour randomisation crops the image around the ROI, rescales intensities, segments supervoxels with SLIC and uses their overlap with the expert mask.
Supplementary note 5: image perturbation settings
The study constructed 18 perturbation chains from rotation, noise addition, translation, volume adaptation, and contour randomisation. Chains were limited to roughly 30 perturbed images to control the number of permutations.
- Perturbation design: 18 perturbation chains combined five basic operations: rotation, noise addition, translation, volume adaptation, and contour randomisation.Each chain generated distorted images for feature calculation.
- Permutation constraints: Chains combining rotation, translation, and volume adaptation were excluded because five angles, two translation fractions, and five volume factors would produce 200 permutations.The chains were designed to produce roughly 30 permutations instead.
- Perturbation design: m denotes the total number of perturbed images generated for a perturbation setting.The listed perturbation settings included 30 repetitions for noise addition and contour randomisation.
- Permutation constraints: Noise addition was not included in every possible chain because its effect was marginal when combined with other perturbations.Several listed chains used one repetition of noise addition or contour randomisation, while others used repeated contour randomisation.
Supplementary note 6: Robustness differences between perturbed test and retest images
Perturbation ICC differences between test and retest images were evaluated for bias using pooled statistical testing. The differences were not significantly displaced from zero, and their distributions were visualized with box plots.
- Bias assessment: Perturbation ICCs were averaged across test and retest images to facilitate comparison with test-retest ICCs.The analysis first calculated feature-wise differences between perturbation ICCs for the two images.
- Bias assessment: n = 1 was used for complete pooling because correlated features violate the independence assumption required for n = 4032.The test assessed the pooled difference distribution against mean 0.
- Results: None of the perturbations produced ICC differences significantly different from 0.The analysis used |z| ≥1.96 as the threshold corresponding to p ≤0.05.
- Results: Figure S6 shows box plots of ICC differences, with medians, interquartile ranges, and whiskers extending to 1.5 times the IQR.The figure summarizes the distributions of differences between test and retest images for the perturbation chains.