Source-linked AI summary
Radiomics strategies for risk assessment of tumour failure in head-and-neck cancer
Martin Vallières, Emily Kay-Rivest, Léo Jean Perrin, Xavier Liem, Christophe Furstoss, Hugo J. W. L. Aerts, Nader Khaouam, Phuc Felix Nguyen-Tan, Chang-Shu Wang, Khalil Sultanem, Jan Seuntjens, Issam El Naqa
TL;DR
The study asks whether radiomic information from head-and-neck cancer imaging can improve risk assessment of locoregional recurrences and distant metastases. Using machine-learning models that combine PET/CT radiomics with clinical information, it shows that radiomics provide important prognostic information for these outcomes.
Problem
The study addresses whether imaging-derived tumour data can support risk assessment of locoregional recurrences and distant metastases in head-and-neck cancer.
Method
The authors analyzed PET/CT radiomics and clinical information from 300 patients across four institutions to develop machine-learning prediction models for early outcome-risk assessment.
Results
Radiomics provided important prognostic information for assessing the risks of locoregional recurrences and distant metastases.
Takeaways & Limitations
Radiomics may support earlier risk assessment of locoregional recurrences and distant metastases in head-and-neck cancer.
Takeaways & Limitations
The models require validation in larger patient cohorts and robust clinical trials before benefits for patient survival can be established.
Abstract
from arXiv · showhide
Quantitative extraction of high-dimensional mineable data from medical images is a process known as radiomics. Radiomics is foreseen as an essential prognostic tool for cancer risk assessment and the quantification of intratumoural heterogeneity. In this work, 1615 radiomic features (quantifying tumour image intensity, shape, texture) extracted from pre-treatment FDG-PET and CT images of 300 patients from four different cohorts were analyzed for the risk assessment of locoregional recurrences (LR) and distant metastases (DM) in head-and-neck cancer. Prediction models combining radiomic and clinical variables were constructed via random forests and imbalance-adjustment strategies using two of the four cohorts. Independent validation of the prediction and prognostic performance of the models was carried out on the other two cohorts (LR: AUC = 0.69 and CI = 0.67; DM: AUC = 0.86 and CI = 0.88). Furthermore, the results obtained via Kaplan-Meier analysis demonstrated the potential of radiomics for assessing the risk of specific tumour outcomes using multiple stratification groups. This could have important clinical impact, notably by allowing for a better personalization of chemo-radiation treatments for head-and-neck cancer patients from different risk groups.
Introduction
Radiomics uses quantitative analysis of routine medical images to characterize tumour heterogeneity and support cancer risk assessment. This study evaluates machine-learning models combining radiomic and clinical information to predict locoregional recurrence and distant metastasis in head-and-neck cancer before treatment.
- Radiomics rationale: Intratumoural heterogeneity is associated with treatment resistance, progression, metastasis, and recurrence, motivating imaging-based risk assessment.Tumour heterogeneity reflects spatial and temporal variation across genes, proteins, microenvironments, tissues, and anatomical landmarks.
- Radiomics rationale: Radiomics is the quantitative extraction and analysis of high-dimensional mineable data from medical images to support clinical decision-making.FDG-PET and CT images provide minimally invasive data for decoding tumour phenotype.
- Study objective: The study constructs prediction models for locoregional recurrence and distant metastasis before chemo-radiation in head-and-neck cancer.The authors hypothesize that radiomic features are important prognostic factors for specific head-and-neck cancer outcomes.
- Study design: 1615 radiomic features were extracted from 300 patients across four institutions and combined with clinical attributes using random forests and imbalance-adjustment strategies.Radiomic features quantified intensity, shape, and texture; clinical attributes included age, head-and-neck cancer type, and tumour stage.
- Study design: Two cohorts trained the models, while two independent cohorts evaluated binary outcome prediction and time-to-event prognostic performance.The study also compares locoregional recurrence and distant metastasis models with overall-survival models and evaluates radiomic, clinical, and volumetric variables.
- Clinical significance: Integrating radiomic features into clinical prediction models may support accurate risk stratification and eventual personalization of radiation doses and chemotherapy regimens.The proposed clinical impact is adapting treatment according to locoregional recurrence and distant metastasis risk before treatment.
Results
Radiomic and clinical variables supported outcome-specific prediction and prognostic risk assessment in head-and-neck cancer, with strongest performance for distant metastases and significant multi-group stratification for selected outcomes.
- Univariate associations: 0 %, 63 % and 12 % of PET radiomic features, versus 0 %, 61 % and 34 % of CT features, were significantly associated with LR, DM and OS, respectively.The associations were identified after multiple testing corrections.
- Univariate associations: The strongest radiomic associations were LZHGEGLSZM for LR (rs = −0.15, p = 0.007), ZSNGLSZM for DM (rs = −0.29, p = 2 × 10−7), and GLVGLRLM for OS (rs = 0.24, p = 4 × 10−5).All three highest associations were obtained from CT scans.
- Prediction performance: For DM, the radiomic-plus-clinical model had sensitivity 0.86, specificity 0.76, accuracy 0.77, CI 0.88, and Kaplan-Meier p-value 3×10−6, with radiomic variables more important than clinical variables.Clinical variables alone did not perform well, whereas volume was significant but inferior to radiomic variables.
- Final models: The final random-forest classifiers used radiomic and clinical features for LR and DM, but clinical variables only for OS.The LR classifier included PET-GLN GLSZM, CT-CorrelationGLCM, CT-LGZE GLSZM, age, H &N type, T-Stage, and N-Stage; the DM classifier included CT-LRHGE GLRLM, CT-ZSV GLSZM, CT-ZSN GLSZM, age, H &N type, and N-Stage.
Discussion
PET/CT radiomics combined with clinical information enabled risk assessment of locoregional recurrences and distant metastases in head-and-neck cancer, with models capturing outcome-specific tumour phenotypes. Radiomics outperformed tumour volume for recurrence and metastasis prediction but not overall survival, while larger cohorts and clinical trials remain necessary before implementation.
- Contribution: PET/CT radiomics and clinical information produced prediction models for early assessment of locoregional recurrence and distant metastasis risk.The models were developed using advanced machine learning.
- Outcome-specific modelling: Radiomic models used distinct, predominantly textural features for locoregional recurrences and distant metastases, suggesting outcome-specific tumour phenotypes.The locoregional recurrence model combined one PET and two CT metrics, whereas the distant metastasis model used three CT metrics.
- Comparative performance: Radiomic models performed considerably better than tumour volume alone for predicting locoregional recurrences and distant metastases.Clinical variables alone performed well for locoregional recurrences but poorly for distant metastases, while combining clinical and radiomic variables positively affected prediction and prognostic assessment.
- Overall survival: Tumour volume matched or outperformed radiomic models for overall survival, whereas clinical variables achieved the best global overall-survival performance.Radiomic features showed low correlations with tumour volume in the final models, but they still did not outperform tumour volume for overall survival.
- Clinical translation: The models significantly separated patients into two locoregional-recurrence risk groups and three distant-metastasis risk groups, but clinical implementation requires larger cohorts and robust trials.Cohort heterogeneity and varying image-acquisition parameters may affect model power and generalizability.
Methods · Data sets availability
The study analyzed imaging and clinical data from 300 head-and-neck cancer patients across four institutions, with cohorts divided into training and testing sets. Patients underwent pre-treatment FDG-PET/CT, and anonymized imaging, clinical, contour, and ROI data were made publicly available.
- Data sets availability: 300 H&N cancer patients from four institutions received radiation alone or chemo-radiation with curative intent.Radiation alone comprised n = 48 (16 %) and chemo-radiation n = 252 (84 %).
- Data sets availability: The H&N1 and H&N2 cohorts formed the training set, comprising 92 and 102 HNSCC patients, respectively.H&N1 was treated at HGJ de Montréal and H&N2 at CHUS, Québec, Canada.
- Data sets availability: Contours from separate treatment-planning CT scans were propagated to the FDG-PET/CT reference frame using intensity-based free-form deformable registration in MIM.This procedure applied to the 207 patients whose contours were drawn on a different CT scan.
- Data sets availability: Pre-treatment FDG-PET/CT images, clinical data, RTstruct contours, ROI-reading MATLAB routines, and associated patient data were made available through TCIA after anonymisation.Institutional review boards approved the study, and McGill University Health Center approved online publication of the anonymized clinical and imaging data.
Sample size and division of cohorts
The four patient cohorts were divided into a combined training set of 194 patients and a combined testing set of 106 patients. Resampling and model construction used the training set, while fully independent validation used the testing set.
- Eligibility and exclusions: Patients with recurrent or metastatic-at-presentation H&N cancer, palliative treatment, or inadequate recurrence/metastasis follow-up were excluded.Patients without locoregional recurrence or distant metastases and with follow-up shorter than 24 months were also excluded.
- Cohort division: 194 patients formed the combined training set from H&N1 and H&N2, while 106 patients formed the combined testing set from H&N3 and H&N4.The resulting training-to-testing sample-size ratio was approximately 2:1.
- Training procedures: Bootstrap resampling and stratified random subsampling estimated performance metrics and constructed final prediction models using training-set patients.Partition sampling maintained approximately similar proportions of locoregional recurrences and distant metastases across the training and testing sets.
- Independent validation: Fully independent validation results were computed using patients from the testing set.Combining different cohorts in training allowed the models to account for some institutional variability.
Extraction of radiomic features
The study extracted 1615 radiomic features from tumour regions in pre-treatment FDG-PET and CT images, covering intensity, shape, and texture characteristics. Texture features quantified intratumoural heterogeneity using multiple matrix types and extraction-parameter combinations.
- Extraction of radiomic features: 1615 radiomic features were extracted from PET and CT tumour regions defined by the “GTVprimary + GTVlymph nodes” contours.PET images were converted to standard uptake value maps, while CT images remained in raw Hounsfield Unit format.
- Extraction of radiomic features: The feature set comprised first-order intensity statistics, morphological shape features, and texture features computed across 40 extraction-parameter combinations.The groups included 10 intensity features, 5 shape features, and 40 texture features.
- Extraction of radiomic features: Texture features measured intratumoural heterogeneity by describing the spatial distribution of intensities within the tumour region.Features were derived from the GLCM, GLRLM, GLSZM, and NGTDM matrices using 3D analysis and 26-voxel connectivity.
- Extraction of radiomic features: Texture extraction varied isotropic voxel size, quantization algorithm, and the number of gray levels across all possible parameter combinations.The tested voxel sizes were 1–4 mm, the quantization algorithms were equal-probability and uniform, and gray levels were 8, 16, 32, or 64.
Construction of radiomic models
Prediction models for three radiomic feature sets and three head-and-neck cancer outcomes were constructed in H&N1/H&N2 training cohorts (n = 194), then tested in H&N3/H&N4 cohorts (n = 106). Feature reduction and multivariable logistic-regression selection used bootstrap AUC optimization to identify parsimonious models.
- Training-set construction: Models were constructed for PET, CT, and combined PET–CT feature sets across three outcomes using H&N1 and H&N2 cohorts (n = 194).The three initial feature sets were I: PET features, II: CT features, and III: PET and CT features.
- Feature reduction: Stepwise forward selection reduced each initial feature set to 25 features balancing predictive power, measured by Spearman’s rank correlation, and non-redundancy, measured by maximal information coefficient.The reduction used the Gain equation.
- Model selection: Feature combinations were selected by maximizing the 0.632+ bootstrap AUC through 25 experiments per model order, using 100 logistic-regression models per starting feature and repeating selection through model order 10.Each model used bootstrap resampling with 100 samples, and combinations were retained for model orders 1 to 10.
- Model selection: For each feature set and outcome, the model order with the best parsimonious properties was selected after estimating performance with 0.632+ bootstrap AUC.The selected model orders were identified from prediction estimates shown in Supplementary Fig. S2.
- Final models and testing: Final coefficients for 3 feature sets × 3 outcomes were averaged across 100 bootstrap samples, and the resulting logistic-regression models were tested in H&N3 and H&N4 cohorts (n = 106).The final models were directly tested in the defined testing set.
Combination of radiomic and clinical variables
Prediction models were built by combining radiomic and clinical variables using random forests trained on 194 patients from the H&N1 and H&N2 cohorts. Clinical predictors included age, H&N type, and outcome-specific tumour-stage variables before incorporation into final radiomic-clinical classifiers.
- Combination of radiomic and clinical variables: 194 patients from the H&N1 and H&N2 cohorts formed the training set for combined radiomic and clinical prediction models.
- Combination of radiomic and clinical variables: Random forest classifiers evaluated Age, H &N type, and Tumour stage for LR, DM, and OS outcomes.H &N type comprised oropharynx, hypopharynx, nasopharynx, or larynx.
- Combination of radiomic and clinical variables: Four tumour-stage groupings were assessed: T-Stage, N-Stage, T-Stage and N-Stage, and TNM-Stage.
- Combination of radiomic and clinical variables: Stratified random sub-sampling and imbalance adjustments were used during feature selection and random forest training.The adjustments accounted for disproportion between event occurrence and non-occurrence.
- Combination of radiomic and clinical variables: The selected staging variables were T-Stage and N-Stage for LR, N-Stage for DM, and T-Stage and N-Stage for OS.
- Combination of radiomic and clinical variables: Clinical variables identified for each outcome were incorporated with variables from three radiomic feature sets into final random forest classifiers.
Imbalance-adjustment strategy
An adapted imbalance-adjustment strategy was used because recurrence, metastasis, and death events were underrepresented relative to non-events. It trained ensembles of balanced classifiers rather than a single unbalanced classifier.
- Imbalance-adjustment strategy: Imbalance adjustment was needed because locoregional recurrences, distant metastases, and death events comprised smaller proportions than non-events in the training and testing sets.
- Imbalance-adjustment strategy: Each bootstrap sample was partitioned into P = [N−/N+] balanced subsets, with all minority-class instances copied into every partition.
- Imbalance-adjustment strategy: Majority-class instances were randomly sampled without replacement so each partition contained either ⌊N−/P⌋ or ⌈N−/P⌉ instances.
- Imbalance-adjustment strategy: For N− = 168 and N+ = 32, five partitions contained 33 or 34 majority-class instances and all 32 minority-class instances.
- Imbalance-adjustment strategy: Logistic regression averaged coefficients across partition-specific classifiers, whereas random forests appended one decision tree from each partition.
Random forest training
Random forests were trained on 194 patients from the H&N1 and H&N2 cohorts using 100 bootstrap samples and imbalance adjustment. Final forests contained 582, 661, and 518 trees for LR, DM, and OS, respectively, with outcome-specific oversampling weights.
- Random forest training: 100 bootstrap samples trained each random forest on the H&N1 and H&N2 training cohorts (n = 194).Each bootstrap sample used the described imbalance-adjustment strategy.
- Random forest training: Each bootstrap sample generated multiple decision trees, one per partition, so final tree counts depended on event proportions for each outcome.The number of trees per forest therefore varied with the actual bootstrap-sample event proportions.
- Random forest training: 582, 661 and 518 decision trees were used for LR, DM and OS, respectively.These were the three final random forest models developed in the work.
- Random forest training: Oversampling weights of 0.5 to 2 in increments of 0.1 were tested to further correct data imbalance during random forest training.Stratified random sub-sampling estimated the optimal weight using maximal average AUC across 10 subtraining and sub-testing splits with a 2:1 size ratio and equal event proportions.
- Random forest training: Final oversampling weights were 1.4, 1.6 and 1.7 for LR, DM and OS, respectively.These weights were used with the previously described imbalance-adjustment strategy to train the forests’ decision trees.
Calculation of prediction/prognostic performance metrics
Models were trained on H&N1/H&N2 and independently tested on H&N3/H&N4. Prediction used ROC-based classification metrics, whereas prognosis used concordance and Kaplan–Meier log-rank analyses.
- Models were trained on H&N1 and H&N2 cohorts (n = 194) and independently tested on H&N3 and H&N4 cohorts (n = 106).
- Prediction performance was assessed using AUC, sensitivity, specificity, and accuracy for binary locoregional recurrence, distant metastasis, and death endpoints.
- Prognostic performance was assessed using the concordance-index and Kaplan–Meier log-rank p-values comparing risk groups.
- Logistic-regression and random-forest outputs used probability thresholds of 0.5 for calculating sensitivity, specificity, and accuracy.
- Cox-regression outputs directly calculated the CI, while the training-set median separated testing patients into two Kaplan–Meier risk groups.
Code and models availability
The study’s software code is freely shared under the GNU General Public License on GitHub. An organized experiment script and standalone MATLAB GUI are available to run experiments and test final models on new patients.
- Code availability: All software code used to produce the study’s results is freely shared under the GNU General Public License on GitHub.The repository is available at https://github.com/mvallieres/radiomics.
- Models and tools availability: A single organized script is available to run all experiments performed in this work.
- Models and tools availability: A standalone MATLAB application with a graphical user interface can test the final models on new patients.