Source-linked AI summary
Interpretable Fundus Image Classification via Ring-Based Retinal Vasculature Features
Xiaoyan Li, Shixin Xu, Arvind Gupta, Huaxiong Huang
TL;DR
The paper tackles limited interpretability in fundus-image classification by introducing an optic-disc-centered ring representation of retinal vascular measurements. It combines geometry, color, oxygenation-related appearance, and vessel–background entropy across rings, achieving strong results across public datasets and 91.1% HRF accuracy matching RETFound. The study also examines dataset shift, segmentation quality, and acquisition-related cues affecting pretrained image models, while identifying boundaries for broader validation and physiological interpretation.
Problem
Deep fundus-image classifiers can perform strongly while relying on latent representations that are difficult to interpret in clinically meaningful terms.
Method
The framework aggregates vascular geometry, color, oxygenation-related appearance, and vessel–background entropy within optic-disc-centered rings from vessel masks and optical-density measurements.
Results
91.1% HRF accuracy using automatically generated vessel masks matched RETFound, while ablation and cross-dataset analyses showed complementary feature information and useful transfer under dataset shift.
Takeaways & Limitations
Ring-wise vascular descriptors provide an interpretable and complementary representation for disease classification, retinal phenotype characterization, and inspection of measurable vascular evidence.
Takeaways & Limitations
Public datasets lacked harmonized acquisition, demographic, and clinical metadata, and SO2-proxy descriptors were relative RGB appearance features rather than calibrated oxygen-saturation measurements.
Abstract
from arXiv · showhide
Retinal fundus photography is widely used for screening and monitoring ocular diseases, but many modern classification pipelines rely on deep latent representations and provide limited interpretability. This study develops an interpretable fundus image classification framework based on a ring-structured representation of the retinal vasculature centered on the optic disc. The method quantifies vessel geometry, color appearance, oxygenation-related vascular appearance, and vessel--background entropy within concentric retinal regions. These physiologically motivated descriptors are derived from vessel masks, image intensities, and optical-density measurements and aggregated across rings to capture spatial variation in vascular properties. Using only quantitative vascular descriptors, the proposed method achieved strong classification performance across three public fundus datasets. On HRF, it achieved 91.1\% accuracy using automatically generated vessel masks, matching RETFound, a vision transformer pretrained on large-scale retinal fundus image data, under the same evaluation setting. Additional analyses suggest that pretrained image models are sensitive to acquisition-related spatial cues, including fundus scale and retinal position within the field of view, as well as broader non-vessel image characteristics. This framework may support interpretable disease classification, quantitative retinal phenotyping, and retinal biomarker discovery without requiring large task-specific training datasets.
1 Introduction
The paper addresses the interpretability limits of deep fundus-image classifiers by representing retinal vasculature through optic-disc-centered annular measurements. It combines physiologically motivated vascular features with spatial analysis of pretrained models and compares their transparency and classification utility.
- Vascular alterations motivate clinically meaningful fundus descriptors for diabetic retinopathy and glaucoma classification.The paper links diabetic retinopathy to microvascular remodeling and glaucoma to peripapillary vascular attenuation and optic-disc-centered changes.
- The proposed representation quantifies geometry, color, oxygenation-related appearance, and vessel–background entropy across concentric optic-disc-centered rings.Features are computed from vessel masks and optical-density measurements, then concatenated into ring-wise vectors for a shallow classifier.
- Prior work often examined vascular and oxygenation-related information separately, whereas this framework integrates them in one ring-wise representation.The unified representation is evaluated across multiple retinal disease categories and against pretrained image-based models including RETFound.
- Controlled experiments examine whether pretrained models respond to field-of-view size, retinal position, and other acquisition-related spatial characteristics.These cues may contribute to classification performance without directly corresponding to retinal abnormalities used in clinical assessment.
- 91.1% accuracy on HRF matched RETFound under the same evaluation setting using automatically generated vessel masks.The result is presented as evidence that compact quantitative vascular descriptors can provide strong classification performance while remaining interpretable.
2 Results
The ring-based vascular representation performed competitively across datasets while providing feature- and region-level explanations. Additional analyses show that pretrained image models are sensitive to acquisition-related spatial variation and broad non-vessel appearance, whereas vascular performance depends partly on mask quality.
- Benchmarking: On HRF, predicted vessel masks yielded 0.911 accuracy, matching RETFound; expert masks increased the proposed method’s accuracy to 1.000.The predicted-mask result remained higher than the other ImageNet-pretrained baselines.
- Benchmarking: On SUSTech-SYSU, the proposed representation achieved 0.896 accuracy, while ConvNeXt-B and ViT-B/16 slightly outperformed RETFound.All methods performed strongly for binary diabetic-retinopathy-versus-healthy classification.
- Benchmarking: On FIVES, the proposed method achieved 0.766 accuracy with expert masks and 0.724 with predicted masks, while RETFound reached 0.806.The expert-mask result was comparable to ConvNeXt-B at 0.769 and ViT-B/16 at 0.753.
- Acquisition-related spatial characteristics: FOV standardization reduced pretrained-model accuracy, including RETFound from 0.927 to 0.806 and ConvNeXt-B from 0.829 to 0.769.Disease groups also differed in fundus-radius and FOV-center distributions, while optic-disc-center distributions largely overlapped.
- Non-vessel appearance: Broad background perturbations reduced RETFound accuracy to 0.725, 0.763, and 0.760 after color, contrast, and brightness standardization, respectively.Joint standardization reduced accuracy to 0.682, whereas lesion-aware inpainting changed accuracy only from 0.806 to 0.803.
- Interpretability: The proposed explanations assign signed predicted-versus-runner-up evidence to optic-disc-centered rings and vessel pixels, linking decisions to measurable vascular phenotypes.Ring-wise scores aggregate feature evidence, with larger positive scores indicating stronger support for the predicted class.
- Segmentation and feature importance: Appearance-related feature groups were generally more important under segmentation-based settings, while high-quality HRF ground-truth masks produced a mixture of geometric and appearance features.These rankings suggest that image quality and vessel delineation affect which vascular cues remain stable.
- Ablation study: Removing SO2-proxy features reduced HRF accuracy from 0.911 to 0.778, while removing vessel–background entropy reduced it to 0.800.Removing geometry or red–green color features reduced accuracy to 0.844; ablation measures incremental utility, unlike standalone feature-family analysis.
3 Discussion
Across three public datasets, ring-wise vascular descriptors provided useful and interpretable classification information, while results highlighted transferability, segmentation, acquisition, and validation boundaries.
- The framework linked predictions to measurable vascular characteristics across HRF, FIVES, and SUSTech-SYSU.
- Removing SO2-proxy features reduced HRF accuracy from 0.911 to 0.778, supporting complementary feature families.The SO2-proxy family had weaker standalone performance than geometry but contributed substantial information within the complete representation.
- Cross-dataset evaluation provided preliminary evidence that anatomically normalized ring-wise descriptors may retain useful information under dataset shift.Adding a within-image-normalized SO2-proxy descriptor improved accuracy from 0.733 to 0.778 after raw color features were excluded.
- Field-of-view and background perturbations suggested that whole-image models may use acquisition-related information beyond retinal vascular pathology.Explicit vascular measurements remain inspectable, but robustness across acquisition settings requires further validation.
- The evaluation is bounded by dependence on reliable segmentation and localization, uncalibrated SO2-proxy measurements, public datasets, and non-real-time feature extraction.Broader validation across institutions, devices, populations, clinical settings, and better-characterized spectral systems remains necessary.
4.1 Datasets
The study evaluates the method on three public fundus datasets differing in image characteristics, disease composition, annotation availability, and classification task.
- Three publicly available datasets were used to assess robustness across differing dataset settings.
- HRF contains 45 images evenly divided among healthy, diabetic retinopathy, and glaucoma classes, with expert vessel and FOV masks.
- FIVES contributes 800 images with manual vessel annotations and quality labels, from which 381 high-quality healthy, glaucoma, and diabetic retinopathy images were selected.
- SUSTech-SYSU provides 1,016 high-resolution images for binary diabetic retinopathy-versus-healthy classification.
4.2 Framework Overview
The framework organizes vascular measurements in optic-disc-centered rings, combines complementary feature families, and classifies the standardized representation with logistic regression.
- Binary vessel masks supply geometry, red–green color, optical-density-derived SO2-proxy, and vessel–background entropy features.Entropy incorporates information from nearby perivascular pixels.
- Concentric annular rings centered on the optic disc preserve vascular variation across retinal eccentricity.The optic disc serves as an anatomically interpretable reference because major vessels emerge from the optic nerve head and branch outward.
- Ring-wise descriptors are concatenated and standardized before classification.
- Logistic regression enables direct examination of feature-level and annular-region contributions to individual predictions.Fig. 4 summarizes this processing pipeline.
4.3 Ring-Based Features
The ring-based representation summarizes retinal vasculature within optic-disc-centered annuli using complementary geometry, caliber, tortuosity, curvature, and connectivity descriptors. Ring-wise robust statistics preserve spatial variation and accommodate segmentation and labeling limitations.
- Ring-wise aggregation: Median and IQR summarize within-ring measurements, preserving typical values and spatial heterogeneity across peripapillary and peripheral retina.IQR is relatively robust to skew, segmentation noise, spurious vessel fragments, and focal atypical regions.
- Density and branching: VAF, VLD, and BPD quantify vessel area, vessel length, and branch-point density within each annular region.BPD is derived from the skeletonized vascular network.
- Density and branching: BPD represents projected-network junction density when combined vessel masks are used, because artery–vein crossings are not removed.It therefore is not necessarily a count of anatomically verified bifurcations.
- Caliber and thickness: Local vessel diameter is estimated as twice the Euclidean distance transform at skeleton pixels, with within-ring variability summarized by the IQR.The estimate assumes skeleton pixels lie near vessel centerlines.
- Caliber and thickness: Thin-vessel length and thick-to-thin diameter ratio provide label-free measures of small-vessel content and within-ring caliber heterogeneity.These groups should not be interpreted as direct substitutes for anatomically defined artery–vein ratios.
- Shape and connectivity: Arc-to-chord tortuosity, local curvature, and graph connectivity characterize vessel shape and network organization across rings.Connectivity descriptors complement BPD by capturing neighboring skeleton connections and possible continuity or fragmentation.
4.3.2 OD-derived SO2-Proxy Features
The SO2-proxy features use RGB optical-density modeling to derive oxygenation-related vascular appearance descriptors within optic-disc-centered rings. The approach combines approximate spectral weighting, local vessel-free references, and regularized nonnegative inversion, while treating the output as an appearance proxy rather than calibrated saturation.
- Motivation: SO2-proxy descriptors are included because prior studies reported altered retinal oxygenation patterns in diabetic retinopathy and glaucoma.Reported patterns include increased venous oxygen saturation in diabetic retinopathy and reduced arteriovenous oxygen difference in glaucoma.
- Optical-density model: The RGB formulation combines near-isosbestic green, oxygen-sensitive red, and additional short-wavelength blue information under a Beer–Lambert approximation.The model represents oxy- and deoxyhaemoglobin contributions through effective channel-averaged extinction coefficients.
- Reference estimation: A local vessel-free reference reduces sensitivity to pigmentation, background reflectance, multiplicative brightness changes, and smooth shading.The reference is estimated using vessel-mask exclusion, inpainting, and smoothing.
- Spectral approximation: RGB channel measurements use approximate band supports and Gaussian-shaped within-band weighting because camera spectral responses are typically unavailable.The experiments use nominal blue, green, and red center wavelengths and report similar trends under moderate passband perturbations.
- Proxy inversion: Ridge-regularized least squares estimates nonnegative haemoglobin coefficients, from which the local OD-derived SO2-proxy is computed and summarized by ring-wise IQR.When artery–vein masks are available, the proxy can also be evaluated by vessel type.
- Interpretation and caveat: The RGB-derived values are oxygenation-related appearance descriptors, not calibrated physiological oxygen-saturation measurements.Approximate passbands and omission of scattering and camera-specific processing limit physiological calibration.
4.3.3 Red–Green Color Features
The framework extracts ring-wise vascular appearance, geometry, oxygenation-related, and vessel–background entropy descriptors from fundus images. These features preserve distinct image-processing pathways for color, oxygenation-related appearance, geometry, and entropy measurements.
- Red–Green Color Features: Red–green color descriptors quantify within-ring vascular color variability and typical chromatic contrast from original RGB intensities.The red-to-green ratio is summarized by its IQR, while the normalized red–green difference is summarized by its median.
- Vessel–Background Entropy: Photometric normalization corrects illumination variation and vignetting while preserving local retinal contrast for vessel–background entropy measurements.The pathway uses channel-wise illumination correction, Lab luminance enhancement, and mild green-channel unsharp blending.
- Feature Processing: Geometry-based features use vessel masks and skeletons, whereas color and SO2-proxy features use the original RGB image.This separation preserves the vessel-to-background intensity relationships required by appearance measurements and leaves geometry independent of image normalization.
- Vessel–Background Entropy: Vessel-centerline and near-vessel background entropy are computed separately within each annular region.The near-vessel band spans non-vessel pixels 2–6 pixels from the nearest vessel, reducing boundary contamination while retaining localized perivascular context.
- Descriptive Results: HRF entropy separation was 0.133 in healthy images, 0.051 in glaucoma, and 0.001 in DR.The reported values are based on differences between vessel-centerline and near-vessel background entropy medians.
- Limitations: Entropy patterns are descriptive because local entropy is nonspecific and may be affected by image quality and acquisition characteristics.The smaller separation in glaucoma is interpreted less directly than the reduced separation in DR.
4.4 Implementation and Evaluation Protocol
The study evaluates ring-wise vascular descriptors with standardized cross-validation and elastic-net logistic regression, comparing them with pretrained retinal and ImageNet image models. It also measures the influence of vessel-mask quality and inference-stage runtime.
- Feature and Mask Processing: Automatically generated vessel masks from RIP-AV were used for the proposed method, with ground-truth masks additionally evaluated on HRF and FIVES.The additional evaluation quantifies the influence of vessel-segmentation quality.
- Feature and Mask Processing: Ring-wise descriptors were concatenated into one standardized feature vector per image before classification.Standardization parameters were fitted within each training fold and applied to held-out data.
- Classification: Elastic-net-regularized logistic regression classified the descriptors using consistent solver, penalty, iteration, and class-weight settings across datasets.Only regularization strength varied by dataset: C = 1.0 for HRF and SUSTech-SYSU, and C = 0.1 for FIVES.
- Baselines: The ring-based representation was compared with RETFound and three ImageNet-pretrained backbones: ResNet-50, ViT-B/16, and ConvNeXt-B.RETFound used a publicly available color-fundus-photography checkpoint pretrained on approximately 1.6 million unlabeled retinal images.
- Baselines: Lightweight fine-tuning trained the final two transformer blocks or convolutional stages together with each model’s classification head.The trainable components differed according to backbone architecture.
- Evaluation: Leave-one-out cross-validation was used for HRF and FIVES, while SUSTech-SYSU used stratified five-fold cross-validation.Accuracy, macro-F1, ROC-AUC, and per-class measures were computed from aggregated out-of-fold predictions.
- Runtime: Runtime profiling used GPU acceleration for vessel segmentation and pretrained encoders, with CPU processing for feature extraction, standardization, and logistic-regression prediction.Detailed stage-wise runtimes and hardware specifications were reported in Supplementary Section S17.
4.5 Use of Generative Artificial Intelligence
The authors used ChatGPT to improve manuscript language, clarity, and organization. They state that authors retained responsibility for the scientific and experimental work.
- AI Assistance: ChatGPT assisted with improving manuscript language, clarity, and organization.The authors report that all AI-assisted text was critically reviewed, verified, and revised.
- Author Responsibility: The authors retained responsibility for the study’s conception, design, experiments, analysis, interpretation, conclusions, and final manuscript.
Data Availability
The study analyzes three publicly available fundus image datasets and generated no new image datasets. The source datasets are HRF, FIVES, and SUSTech-SYSU.
- Public Datasets: The analyzed datasets were the publicly available HRF, FIVES, and SUSTech-SYSU fundus image datasets.
- Data Generation: No new image datasets were generated during the study.