Source-linked AI summary
Predicting Risk of Developing Diabetic Retinopathy using Deep Learning
Ashish Bora, Siva Balasubramanian, Boris Babenko, Sunny Virmani, Subhashini Venugopalan, Akinori Mitani, Guilherme de Oliveira Marinho, Jorge Cuadros, Paisan Ruamviboonsuk, Greg S Corrado, Lily Peng, Dale R Webster, Avinash V Varadarajan, Naama Hammel, Yun Liu, Pinal Bavishi
TL;DR
Diabetic retinopathy screening faces a growing burden, motivating risk stratification to personalize screening frequency. This study developed and validated fundus-photograph deep learning systems that retained prognostic value across internal and external datasets, including after adjustment for available risk factors.
Problem
Growing diabetic populations make screening burdensome, while existing fundus-photograph grading systems do not further stratify patients without apparent retinopathy by future risk.
Method
The study trained deep learning systems using three-field or single-field color fundus photographs and multi-task learning, then validated them on internal and external datasets.
Results
The DLS demonstrated good performance alone and after adjustment for available risk factors across internal and external validation datasets.
Takeaways & Limitations
The findings support using fundus photographs to identify prognostic information for diabetic retinopathy development beyond available risk factors.
Takeaways & Limitations
Several known risk factors were unavailable for both datasets, limiting comparisons with or adjustment for those factors.
Abstract
from arXiv · showhide
Diabetic retinopathy (DR) screening is instrumental in preventing blindness, but faces a scaling challenge as the number of diabetic patients rises. Risk stratification for the development of DR may help optimize screening intervals to reduce costs while improving vision-related outcomes. We created and validated two versions of a deep learning system (DLS) to predict the development of mild-or-worse ("Mild+") DR in diabetic patients undergoing DR screening. The two versions used either three-fields or a single field of color fundus photographs (CFPs) as input. The training set was derived from 575,431 eyes, of which 28,899 had known 2-year outcome, and the remaining were used to augment the training process via multi-task learning. Validation was performed on both an internal validation set (set A; 7,976 eyes; 3,678 with known outcome) and an external validation set (set B; 4,762 eyes; 2,345 with known outcome). For predicting 2-year development of DR, the 3-field DLS had an area under the receiver operating characteristic curve (AUC) of 0.79 (95%CI, 0.78-0.81) on validation set A. On validation set B (which contained only a single field), the 1-field DLS's AUC was 0.70 (95%CI, 0.67-0.74). The DLS was prognostic even after adjusting for available risk factors (p<0.001). When added to the risk factors, the 3-field DLS improved the AUC from 0.72 (95%CI, 0.68-0.76) to 0.81 (95%CI, 0.77-0.84) in validation set A, and the 1-field DLS improved the AUC from 0.62 (95%CI, 0.58-0.66) to 0.71 (95%CI, 0.68-0.75) in validation set B. The DLSs in this study identified prognostic information for DR development from CFPs. This information is independent of and more informative than the available risk factors.
Funding
Google LLC funded the study and participated in its approval for publication.
- Funding: Google LLC funded the study and had a role in approving it for publication.The funder was based in Mountain View, California.
Research in context
A deep learning system using color fundus photographs predicted two-year development of mild-or-worse diabetic retinopathy in patients without DR and improved risk stratification across two independent datasets. The tool may help clinicians personalize screening intervals by monitoring high-risk patients more closely and screening lower-risk patients less frequently.
- Study contribution: The DLS used color fundus photographs to predict development of mild-or-worse (“Mild+”) DR within two years in diabetic patients without DR.It was developed from a large retrospective longitudinal dataset collected through a DR screening program.
- Study contribution: Significant improvement in risk stratification was observed across two independent datasets from different countries and ethnicities.
- Clinical implications: Clinicians could use the tool to optimize screening intervals by following high-risk patients more closely and lower-risk patients less frequently.This approach aims to improve visual outcomes for high-risk patients while reducing the burden of screening for lower-risk patients.
Introduction
Diabetic retinopathy is a major, preventable cause of blindness, but rising diabetes prevalence makes regular screening and follow-up increasingly difficult. This study hypothesized that deep learning applied to fundus photographs could stratify two-year DR-development risk, including among eyes without visible DR.
- Clinical need: 415 million diabetic patients in 2015 are expected to increase to 642 million by 2040, intensifying the burden of DR screening and follow-up.Regular screening is important because early treatment can slow, halt, or reverse DR progression.
- Clinical need: Most diabetic patients develop DR during the first two decades of diabetes, and untreated disease can progress to vision-threatening DR.Major organizations recommend screening, with no or mild DR generally screened every 12 to 24 months.
- Risk stratification: Fundus photographs show diabetes-related retinal microvascular changes, while DR development and progression are also influenced by systemic, demographic, and genetic risk factors.Known modifiable factors include hyperglycemia, hypertension, dyslipidemia, obesity, smoking, anemia, pregnancy, health literacy, healthcare access, and treatment adherence.
- Study rationale: Because microvascular changes are not known to be detectable before DR develops, current photograph-based grading systems do not directly provide preclinical risk stratification.The study therefore focused on combining fundus photographs with known risk factors to improve risk stratification.
- Study objective: The study created a deep learning system to predict whether eyes would develop Mild+ DR within two years and evaluated risk stratification using photographs alone or with available risk factors.The authors hypothesized that fundus photographs could stratify risk in patients showing no signs of DR and assessed generalization across outcomes and time points.
Methods
The study used de-identified CFPs from U.S. and Thai datasets to predict 2-year development of Mild+ DR in eyes initially without DR. An Inception-v3 DLS was trained with multi-task learning, evaluated using prespecified statistical methods, and examined with saliency techniques.
- Datasets and imaging: 430,917 retinal examination visits from 362,283 patients formed the U.S. EyePACS dataset, using three 45° CFP fields per eye.Images came from multiple camera devices used at screening sites.
- Datasets and imaging: 6,791 patients from Thailand provided an external dataset using standard 45° primary-field CFP imaging, with 4,762 eyes retained for validation set B.Thai patients were randomly identified from the national diabetic patients registry across 13 health regions.
- Outcome definition: The endpoint was whether an eye without DR at baseline developed Mild+ DR within 2 years, with a 28-day buffer accounting for visit scheduling and between-visit progression.A Mild+ grade during a visit implied progression between the prior and current visits.
- Reference grading: DR lesions were graded by certified graders using a modified ETDRS protocol, mapping lesion findings to five DR levels and diabetic macular edema.EyePACS eyes were graded using all three fields; Thailand images were also graded under the study protocol.
- Model development: An Inception-v3 DLS was trained on ⅞ of the development dataset and tuned on ⅛, using multi-task learning to leverage lesion and DR-grade labels when 2-year outcomes were unavailable.Only 25,211 of 503,527 training-set eyes had known 2-year Mild+ DR outcomes.
- Analysis and interpretation: Saliency methods and survival-analysis procedures were used to interpret DLS predictions and quantify outcomes, with AUC confidence intervals calculated by the DeLong method.Interpretability methods included Integrated Gradients, Guided Backprop, and Blur Integrated Gradients.
Results
The DLS showed predictive discrimination for DR development across validation datasets, with generally high negative predictive values and sharply higher predicted risk in the highest-risk group. It was well calibrated in validation set A, while validation set B required scaling to resolve overestimation.
- Discrimination: 0.79 AUC was achieved by the 3-field DLS for predicting DR development in validation set A.The 95% CI was 0.77-0.81.
- Discrimination: 0.70 AUC was achieved by the 1-field DLS in validation set B, compared with 0.78 AUC in validation set A.The 1-field DLS’s 95% CIs were 0.67-0.74 in set B and 0.76-0.80 in set A.
- Calibration: Both DLS versions were well calibrated in validation set A, whereas the 1-field DLS overestimated DR development in validation set B until simple scaling using 5% of that set.Validation set B had a lower DR incidence than validation set A, 15% versus 19%.
- Risk stratification: NPVs generally exceeded 80% and surpassed 95% for the lowest-risk group, while the highest-risk decile had predicted DR risks of 40% to 60%.These findings occurred in the context of incidence rates below 20%.
- Comparison with risk factors: 0.68 AUC was achieved by HbA1c alone, while all risk factors combined by multivariable logistic regression achieved 0.72 AUC in validation set A.The remaining individual risk factors had AUCs ranging from 0.59 to 0.64.
- Temporal association: At 4-8 years before DR developed, median predicted risk remained approximately 10%, then increased to approximately 20% 3 years before, 25% 2 years before, and 40% 1 year before development.Kaplan-Meier analysis also examined associations between DLS-defined risk groups and DR development over time.
Discussion
The deep learning system stratified patients by their risk of developing diabetic retinopathy from color fundus photographs, retained prognostic value beyond available risk factors, and may support personalized screening and targeted interventions. Its generalizability was supported across internal and external datasets, although calibration and risk-factor data remained limitations.
- Study contribution: The DLS predicted DR development within 2 years and showed good performance both alone and after adjustment across internal and external validation datasets.The datasets were predominantly Hispanic and Thai, respectively; Kaplan–Meier analyses supported prognostication across different time points.
- Study contribution: The model’s predictions were also associated with progression beyond mild DR, including moderate DR and vision-threatening DR.
- Limitations: The model was well calibrated on validation set A but overestimated DR incidence on validation set B, while missing risk factors and grading variability limited interpretation.Possible contributors included lower incidence in set B, lower high-HbA1c prevalence, grading-protocol differences, unavailable blood pressure data, uncertain HbA1c timing, and variability in subtle findings.
- Study contribution: The DLS retained high prognostic value after adjustment for several risk factors, supporting independent information from color fundus photographs for patients without DR.The study addressed risk stratification using color fundus photographs and risk factors available in most screening settings.
- Clinical utility: Personalized risk assessment could optimize screening intervals by following high-risk patients more frequently and low-risk patients less frequently.Potential applications also include targeted lifestyle counseling, stricter pharmacologic blood-sugar control, clinical-trial selection, and alerts for missed screening visits.
Tables · A B C
The tables summarize baseline characteristics, longitudinal DR outcome cohorts, and predictive-performance comparisons involving three-field and one-field DLS configurations. They also define the reported outcomes, abbreviations, and validation datasets.
- Tables: Table 1 reports baseline characteristics for the study cohorts, including age, sex, ethnicity, and HbA1c.
- Tables: 7,976 eyes in one validation cohort and 4,762 eyes in another had at least two sets of gradable images without DR at the first visit.
- Tables: 3,678 and 2,345 eyes, respectively, had known outcomes for mild-or-worse (“Mild+”) DR within 2 years.
- Tables: The tables separately define moderate-or-worse (“Moderate+”) DR within 2 years and vision-threatening DR (VTDR) within 2 years as outcomes.
- Tables: The tables use abbreviations including DR, IQR, HbA1c, AUC, CFP, and DLS; VTDR includes diabetic macular edema.
- Tables: Table 2 compares DLS predictive performance with and without known risk factors in the EyePACS validation dataset and the Thailand dataset.
- A B C: The listed model configurations include a 3-field DLS for validation set A and 1-field DLSs for validation sets A and B.
D E F
Figures 1 and 2 evaluate DLS discrimination, calibration, and predictive values for diabetic retinopathy development across validation sets and model inputs. ROC analyses compare DLS performance with risk factors and their combination, while PPV and NPV are shown with 95% confidence intervals.
- Discrimination and calibration: Figure 1 presents ROC curves for 3-field and 1-field DLS models, risk factors, and their combinations in validation set A.ROCs were plotted only for patients with known risk-factor values to enable comparison.
- Predictive values: Figure 2 shows positive and negative predictive values for DLS prediction of diabetic retinopathy development across validation sets A and B.The analyses include 3-field and 1-field models, with shaded regions indicating 95% confidence intervals.
- Evaluation settings: The reported analyses distinguish 3-field DLS in validation set A from 1-field DLS evaluations in validation sets A and B.These configurations are listed as separate evaluation settings.
A B C · Supplement
The supplementary figures illustrate DLS-based DR risk stratification, image-level prediction explanations, and the contributions of different fundus fields and image regions.
- A B C: Kaplan–Meier plots show DR incidence stratified into high-, moderate-, and low-risk groups by the DLS.Risk groups were defined using the upper and lower quartiles in the tuning dataset.
- A B C: The plots include 3-field and 1-field DLS results for validation set A and 1-field results for validation set B.Validation sets A and B had different follow-up durations.
- A B C: Integrated Gradients saliency heatmaps visualize how baseline CFP regions contributed to predictions of developing or not developing DR.Red indicates contribution toward developing DR, yellow toward not developing DR, and blue indicates little contribution.
- A B C: Example cases pair baseline CFPs and saliency heatmaps with CFPs from follow-up visits.The figure presents the baseline image, explanation map, and follow-up image in separate columns.
- Supplement: The supplementary analyses compare DLS performance when using one or two images from three-field examinations with performance using all three images.Mean performance and standard deviation across three training runs are shown.
- Supplement: The figures also examine the importance of different fundus fields and different parts of the image.Figure 5 includes sample images of each of the three fields and sample images of the primary field.
Supplementary Methods · Visualization of differences between patients with vs those without followup · Handling outlier values for risk factors
The supplementary methods compared baseline variables and DLS predictions between patients with and without follow-up visits, while processing extreme risk-factor values before logistic regression.
- Visualization of differences between patients with vs those without followup: Baseline variables were visualized to assess differences between patients with and without follow-up visits.The visualization included several baseline variables.
- Visualization of differences between patients with vs those without followup: DLS predictions were also visualized when comparing patients with versus without follow-up visits.These comparisons were reported in Supplementary Table S7.
- Visualization of differences between patients with vs those without followup: The comparison was designed to identify differences between patients who had follow-up visits and those who did not.The analysis used visualization rather than a stated inferential test.
- Handling outlier values for risk factors: Age values less than 1 or higher than 122 were removed before logistic regression modeling.Age values from 90 to 122 were instead set to 90.
- Handling outlier values for risk factors: HbA1c values below 1% or above 18% were removed to limit the influence of potential outliers.The stated motivation included possible data-entry errors.
- Handling outlier values for risk factors: “Years with diabetes” values were clipped to the range from 1 to 20 before modeling.This processing was intended to prevent risk-factor outliers from disproportionately influencing logistic regression models.
Comparison of different modeling approaches … Supplementary Figures
Supplementary analyses compared modeling strategies, assessed training-data sufficiency, and validated the DLS against an automated reference standard. Supplementary tables and figures documented grading definitions, hyperparameters, risk-factor analyses, and validation-set characteristics.
- Comparison of different modeling approaches: AUC of 0.58 was achieved by a network predicting DR on the five-point scale.The study also explored logistic regression applied to the continuous predicted distribution over that scale.
- Comparison of different modeling approaches: Performance had not plateaued with the available training data and could potentially improve with more data.This analysis was presented in Supplementary Figure S6.
- Using a reference standard based on automated grading: AUC of 0.954 was achieved on the tune set for classifying no DR versus Mild+ DR using an automated-grading model trained on three CFP fields.The model was trained and tuned using the development sets before being applied to validation images.
- Supplementary Tables: Supplementary Tables S1 and S2 define DR and DME levels for the EyePACS and Thailand datasets, respectively, using lesion-based grading mappings.EyePACS grading followed the ETDRS protocol, whereas Thailand grades were provided directly by retina specialists.
- Supplementary Tables: Supplementary Table S3 reports the model hyperparameters, while Supplementary Table S4 compares different modeling approaches.These tables document methodological choices underlying the supplementary experiments.
- Supplementary Tables: Supplementary Tables S5 and S6 report DLS predictive performance alongside risk factors and univariable and multivariable Cox analyses for validation datasets A and B.Table S5 also describes the available race/ethnicity categories, and Table S6 gives sample sizes for analyses with complete risk-factor data.
- Supplementary Figures: Supplementary Figure S1 shows visit-time distributions for validation sets A and B, and Supplementary Figure S2 shows HbA1c distributions in both datasets.These figures characterize follow-up timing and glycemic-marker distributions across the validation cohorts.
A B C
Supplementary analyses characterize DLS discrimination, predicted risk over time before DR onset, saliency-method examples, training-data effects, and baseline differences by follow-up status.
- Discrimination: Supplementary Figure S3 presents ROC curves for 3-field and 1-field DLS models across validation sets A and B.The figure examines discrimination for predicting DR incidence in all patients, rather than only those with corresponding risk factors.
- Risk over time: Supplementary Figure S4 shows predicted DR-development probabilities in eyes that eventually developed DR as a function of time before onset.Box plots display medians, quartiles, 1.5-times-interquartile-range whiskers, outliers, and confidence intervals across medians.
- Saliency analysis: Supplementary Figure S5 provides example cases using three saliency techniques instead of one.These examples correspond to those presented in Figure 4.
- Training data: Supplementary Figure S6 examines how training-data amount affects 1-field DLS performance using tune-set AUCs with 95% confidence intervals.Confidence intervals were computed with the DeLong method from one trained model per point.
- Follow-up status: Supplementary Figure S7 compares baseline variables and DLS predictions between patients with and without follow-up.