Source-linked AI summary
Retinal OCTA Phenotyping with LLM Reporting for Alzheimer's Disease
Progga Paromita Dutta, Jeba Maliha, Md Rafiul Kabir
TL;DR
Population-scale Alzheimer’s disease identification remains difficult, while many OCTA approaches require diagnostic labels and offer limited measurement-level interpretation. This paper develops an explainable OCTA pipeline for annotation-aware segmentation, label-free vascular phenotyping, and measurement-grounded language-model reporting, finding internally consistent vascular structure in nine held-out subjects without supporting clinical interpretation.
Problem
Existing Alzheimer’s assessment methods can be costly or burdensome for population-scale screening, while OCTA approaches often require diagnostic labels and provide limited measurement-level interpretation.
Method
The pipeline combines annotation-aware vessel segmentation, layer-appropriate biomarker extraction, label-free clustering, and measurement-grounded reporting across three language models.
Results
The segmentation models achieved ROC-AUC values of 0.916–0.970 and Dice scores of 0.695–0.781, while nine held-out subjects showed an internally consistent lower-density, lower-fractal-dimension phenotype.
Takeaways & Limitations
The framework provides a transparent, non-diagnostic connection between retinal vascular measurements, exploratory phenotyping, and evidence-linked interpretation for Alzheimer’s research.
Takeaways & Limitations
The absence of diagnostic labels and the small evaluation cohort preclude clinical interpretation.
Abstract
from arXiv · showhide
Early identification of Alzheimer's disease (AD) remains challenging because established assessment methods can be costly, resource-intensive, or unsuitable for population-scale screening. Optical coherence tomography angiography (OCTA) provides non-invasive visualization of retinal microvasculature, but existing approaches often require diagnostic labels and provide limited measurement-level interpretation. We present an explainable OCTA pipeline that integrates annotation-aware vessel segmentation, layer-specific vascular biomarker extraction, label-free phenotyping, and measurement-grounded LLM reporting. Using 117 ROSE-1 images from 39 subjects, we apply annotation-matched segmentation models to superficial vascular complex (SVC), deep vascular complex (DVC), and combined SVC+DVC representations. The models achieve ROC-AUC values of 0.916-0.970 and Dice scores of 0.695-0.781. Six density and fractal-dimension biomarkers form subject-level profiles for exploratory clustering. Analysis of nine held-out subjects identifies an internally consistent lower-density, lower-fractal-dimension phenotype, although the absence of diagnostic labels prevents clinical interpretation. Reports generated using GPT, Gemini, and Llama are evaluated for measurement grounding, citation faithfulness, and diagnostic caution. Overall, the framework provides a transparent, non-diagnostic connection between retinal vascular measurements, exploratory phenotyping, and evidence-linked interpretation for Alzheimer's research.
I. INTRODUCTION
The paper addresses limitations of OCTA-based Alzheimer’s research by linking annotation-aware vascular measurement with label-free phenotyping and evidence-grounded reporting. Its scope is transparent research analysis rather than clinical diagnosis.
- OCTA offers non-invasive, capillary-resolution visualization of retinal microvasculature, while established Alzheimer’s assessments can be costly or burdensome for population-scale screening.
- Existing OCTA approaches often require diagnostic labels and provide prediction scores without identifying the vascular measurements underlying them.
- The proposed pipeline adapts segmentation, biomarker definitions, and evaluation to heterogeneous vessel annotations and retinal vascular layers.
- Six density and fractal-dimension biomarkers support exploratory clustering without diagnostic labels and report generation using cohort-relative z-scores and curated evidence.
- The reporting pipeline evaluates GPT, Gemini, and Llama for grounding accuracy, citation faithfulness, and diagnostic caution.
II. RELATED WORK
Prior OCTA studies report vascular abnormalities associated with Alzheimer’s disease, but findings vary across cohorts and methods. The paper addresses this variability by integrating annotation-aware segmentation, layer-appropriate biomarkers, label-free phenotyping, and measurement-grounded reporting.
- Prior studies report reductions in vessel density and vascular fractal dimension, while findings for affected plexuses and foveal avascular zone area remain inconsistent.
- Reported OCTA variability is associated with differences in populations, imaging devices, scan protocols, retinal-layer definitions, and processing methods.
- Feature-based diagnostic methods are relatively interpretable but remain sensitive to segmentation and feature-selection choices and may overfit small cohorts.
- Segmentation benchmarks and retrieval-augmented generation rarely provide a complete connection from annotation-aware vascular outputs to downstream phenotyping and structured interpretation.
- The proposed framework links annotation-aware segmentation, layer-appropriate biomarkers, label-free phenotyping, and measurement-grounded reporting without assigning an Alzheimer’s diagnosis.
A. Dataset and Annotations
The framework processes heterogeneous ROSE-1 OCTA annotations through representation-specific segmentation, biomarker extraction, clustering, and stability analysis. It uses a six-dimensional subject profile while excluding a quality-compromised FAZ biomarker.
- Dataset and annotations: ROSE-1 contains 117 en-face OCTA images from 39 subjects covering SVC, DVC, and combined SVC+DVC representations with different annotation formats.
- Segmentation framework: The pixel-annotated SVC and SVC+DVC models use separate thick-vessel and centerline decoder branches, whereas centerline-annotated DVC uses a reduced single-branch model.
- Biomarker extraction: Vessel density is calculated for SVC and SVC+DVC, normalized centerline density for DVC, and fractal dimension for all three representations.
- Biomarker extraction: FAZ area was excluded from the final biomarker set because qualitative quality control revealed frequent background leakage.
- Phenotyping: Standardized biomarkers are clustered with k-means without diagnostic labels, using silhouette, leave-one-subject-out, permutation, and cross-layer agreement analyses.
- Phenotyping: Independent feature permutation is a restricted null test because it does not preserve cross-feature covariance.
C. Evidence-Grounded LLM Report Generation
The reporting pipeline converts quantitative OCTA biomarkers into structured subject-level narratives. It combines cohort-relative measurements, unsupervised assignments, and a locked evidence bank to constrain generated claims.
- The pipeline inputs patient-level biomarker values, cohort statistics, z-scores, and unsupervised cluster assignments to generate standardized Findings, Interpretation, and Important sections.
Biomarker Input and Phenotype Encoding:
Subject profiles combine six layer-specific vascular biomarkers with cohort-relative standardization and unsupervised cluster assignments.
- Subject profiles include measured values, cohort statistics, layer-specific z-scores, and unsupervised cluster assignments.
- Six biomarkers cover vessel density and fractal dimension for SVC and SVC+DVC, plus normalized centerline density and fractal dimension for DVC.
- Each biomarker is represented by its measured value and cohort-relative z-score.
LLM Prompt for Report Generation:
The reporting prompt structures LLM outputs around complete biomarker findings and research-oriented interpretation rather than clinical decision-making.
- The system prompt requires Findings, Interpretation, and Important sections.
- Findings must report all six biomarkers with actual values and z-score deviations from the cohort mean.
- Table I reports held-out ROSE-1 segmentation performance, with Dice values requiring representation-specific interpretation.
- Generated output is intended for research interpretation rather than clinical decision-making.
Literature-Grounded Citation Control:
Report citations are controlled through a locked evidence bank and evaluated alongside segmentation and diagnostic-caution criteria.
- A locked reference bank maps approved sources to specific biomarker patterns and permitted claims.
- Extracted citations are checked against the subject profile’s evidence mapping so claims remain traceable to approved sources.
- The evaluation uses a subject-disjoint split with 30 subjects for segmentation development and nine held-out subjects.
- Report-generation quality is evaluated for grounding accuracy, citation faithfulness, and diagnostic caution.
B. Result Analysis of Retinal Image
Held-out segmentation and clustering analyses identify an internally consistent lower-density, lower-fractal-dimension phenotype across retinal vascular representations.
- Segmentation Performance: ROC-AUC ranged from 0.916 to 0.970 across SVC, DVC, and SVC+DVC representations.SVC achieved strict pixel-level Dice of 0.781, combined SVC+DVC achieved 0.712, and DVC achieved tolerance-aware centerline Dice of 0.695.
- Exploratory Phenotyping: The same three subjects formed the lower-density, lower-fractal-dimension group across SVC, DVC, and SVC+DVC.
- Exploratory Phenotyping: Two clusters contained three and six subjects, with one group showing lower density and lower fractal dimension across the representations.
- Exploratory Phenotyping: The primary two-cluster solution had silhouette 0.626, mean leave-one-subject-out ARI 1.000, empirical permutation p-value 0.0005, and mean cross-layer ARI 1.000.These findings indicate internal consistency but require cautious interpretation and larger independent validation because the cohort contains nine subjects.
C. Result Analysis of LLM-Generated Reports
The evaluation measures whether generated OCTA reports accurately reflect biomarkers and faithfully support claims with citations. It assesses grounding, citation faithfulness, and consistency with subject-level measurements.
- Grounding accuracy checks whether reports correctly present six biomarkers, their cohort-relative directions, and overall z-score patterns.Mixed profiles receive particular attention.
- Citation faithfulness verifies claim support, biomarker attribution, and the absence of invented or mismatched references against a locked reference bank.A deterministic 1–5 score rewards exact claim support.
- Fig. 5 summarizes the LLM-based evaluation metric used for report assessment.
Clinical Caution:
The report-generation pipeline evaluates whether outputs preserve non-diagnostic language, while error analysis identifies FAZ leakage and cohort-specific limitations. These constraints limit clinical interpretation of the exploratory findings.
- Clinical Caution: A 1–5 clinical-caution scale evaluates whether reports preserve appropriate non-diagnostic language around Alzheimer’s-related retinal vascular phenotypes.
- Clinical Caution: Across GPT, Llama, and Gemini, grounding scores ranged from 3.11 to 3.56, citation-faithfulness scores from 3.00 to 3.22, and caution scores from 3.33 to 3.56.Grounding was generally higher than citation faithfulness.
- Error Analysis and Discussion: FAZ extraction was excluded after leakage through inter-vessel gaps produced implausibly large regions.The failure is illustrated in Fig. 6.
- Error Analysis and Discussion: The nine-subject analysis yields cohort-specific z-scores and cluster assignments that show internal consistency rather than clinical validity.The same subjects were used for standardization, phenotyping, and reporting.
V. CONCLUSION
The paper concludes with an explainable OCTA framework linking segmentation, vascular biomarkers, exploratory phenotyping, and evidence-linked reporting. Its findings remain non-diagnostic and motivate validation on labeled external cohorts.
- The framework integrates annotation-aware vessel segmentation, layer-appropriate biomarker extraction, label-free phenotyping, and measurement-grounded report generation.
- ROC-AUC values of 0.916–0.970 and Dice scores of 0.695–0.781 were achieved across SVC, DVC, and SVC+DVC representations.
- Nine held-out subjects showed an internally consistent lower-density, lower-fractal-dimension phenotype, but absent diagnostic labels and the small cohort preclude clinical interpretation.
- Reports from three language models showed stronger measurement grounding than citation faithfulness.
- Future work will use labeled external cohorts, controlled segmentation baselines, improved FAZ extraction, and stronger citation verification.