Source-linked AI summary
LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology
Marie-Lisa Eich, Kai Standvoss, Timo Milbich, Alexander Möllers, Miriam Hägele, Philipp Anders, Lars Tharun, Hanna Kontradiuk, Sebastian Kons, Nader Aldoj, Recepcan Adigüzel, Adam Narai, Lukas Hönig, Jonathan Striebel, Binru Yang, Mihnea P. Dragomir, Marvin Sextro, Philipp Keyl, Philipp Jurmeister, Rosemarie Krupar, Evelyn Ramberger, James Wells, Julika Ribbat-Idel, Andreas Kunft, Hussam Shuaib, Christian Grohé, Reinhard Büttner, David Horst, Klaus-Robert Müller, Lukas Ruff, Maximilian Alber, Frederick Klauschen, Simon Schallenberg
TL;DR
Lung cancer pathology requires integrating complex multimodal evidence, but manual assessment is laborious and variable and existing AI tools often lack comprehensive prospective validation. LUCAID combines an interactive diagnostic agent with nine validated modules spanning the routine workflow. It achieved 93.0% concordance with an expert-panel reference standard in prospective clinically actionable assessments.
Problem
Manual lung cancer pathology must integrate complex histological, immunohistochemical, and molecular features, while existing AI tools cover limited tasks and often lack expert-level performance and prospective validation.
Method
LUCAID uses an agent to orchestrate nine validated modules covering quality control, tumor analysis, TME profiling, cellularity, IHC phenotyping, biomarker scoring, and structured reporting.
Results
93.0% concordance with an expert-panel adjudicated reference standard was achieved prospectively across clinically actionable decisions, compared with 68.3–81.1% for five pathologists.
Takeaways & Limitations
LUCAID provides quantitative, reproducible analysis across the lung cancer diagnostic workflow and supports integrated case-level pathology reporting.
Takeaways & Limitations
The system is focused on lung cancer, its effect on routine clinical decision-making remains unestablished, and report evaluation covered only ten cases as proof of principle.
Abstract
from arXiv · showhide
Lung cancer tissue diagnostics is complex, as therapy decisions in precision oncology rely on the integration of histomorphological, immunohistochemical, and molecular features. Yet pathological assessment remains largely visual and semi-quantitative and shows interobserver variability, while existing artificial intelligence (AI) tools cover only selected tasks, rarely reach generalizable expert-level performance, and lack prospective clinical validation. To address these challenges, we developed and clinically validated LUCAID, an agentic AI system for precision lung cancer pathology. An integrative agent couples diagnostic reasoning with nine modules that cover the full routine workflow, from quality control, tumor detection and segmentation, histological subtyping, tumor microenvironment profiling, tumor cellularity quantification, and predictive biomarker scoring (PD-L1, MET, TROP-2) to automated structured report generation. LUCAID enables users to interactively query the module outputs and generate reports that contextualize the results. Against large-scale expert ground-truth annotations, the analysis modules achieved F1 scores of 0.82-0.95. In prospective clinical validation, LUCAID reached 93.0% concordance with an expert-panel adjudicated reference standard across clinically actionable decisions, compared with 68.3-81.1% for five experienced thoracic pathologists.
1 Introduction
Lung cancer pathology increasingly requires integrating complex histological, immunohistochemical, and molecular evidence, while manual quantification remains laborious and variable. LUCAID addresses this gap with an agentic system that coordinates validated modules across the diagnostic workflow and supports prospective clinical evaluation.
- Precision lung cancer treatment decisions increasingly depend on integrating histomorphological, immunohistochemical, and molecular features.
- Manual assessment of emerging biomarkers, multiple scoring systems, cell populations, compartments, and spatial relationships is laborious and prone to interobserver variability.
- Existing AI tools often address selected tasks but may not achieve pathologist-level performance or provide complete, clinically validated workflow integration.
- LUCAID coordinates validated modules for quality control, tumor analysis, TME profiling, cellularity, IHC phenotyping, and biomarker scoring through an interactive agent.
- 93.0% concordance with an expert-panel reference standard exceeded the 68.3–81.1% concordance of five individual pathologists in prospective clinically actionable assessments.
2 Results
LUCAID implements a modular, agent-orchestrated workflow for lung cancer pathology, with validation spanning expert annotations, biological concordance, report grounding, and prospective clinical assessment. Its modules support quality control, tissue and cell analysis, diagnostic subtyping, biomarker scoring, cellularity assessment, and structured reporting.
- LUCAID deploys nine modules covering WSI quality control, tissue segmentation, tumor subtyping, TME profiling, cellularity, IHC phenotyping, biomarker scoring, and structured report generation.
- The system was validated using 1,620 multicentric lung cancer cases, H&E and nine IHC markers, dedicated hold-out sets, molecular references, report assessment, and prospective validation.
- Quality control: 0.91 and 0.88 overall F1 scores were achieved for H&E and IHC quality control, respectively.
- Tissue segmentation: 0.93 average F1 was achieved across seven tissue categories, while carcinoma detection reached F1 scores of 0.98 across carcinoma subtypes and 0.99 in metastatic samples.
- Tumor subtyping: 93.3% to 100% accuracy was achieved for lineage-marker scoring used to distinguish LUAD from LUSC, including TTF-1, CK7, p40, and CK5/6.
- Tumor cellularity: Cellularity heatmaps visualize regional tumor-content heterogeneity and may support more standardized tissue-area selection for molecular profiling.
- IHC phenotyping: IHC cell phenotyping achieved F1 scores of 0.86 for TROP-2, 0.86 for MET, and 0.90 for PD-L1 across all cells.
NSCLC NEC
LUCAID combines quantitative pathology modules with structured, evidence-grounded reporting for lung cancer diagnosis and biomarker assessment. Prospective evaluation showed high concordance with expert-panel decisions across clinically actionable tasks, while report outputs were largely grounded and clinically correct.
- Structured reporting: LUCAID integrates validated module outputs into predefined report sections with case-level assessment and literature-supported clinical interpretation.Quantitative values are transferred from validated modules, while interpretations may be supported by PubMed-retrieved references.
- Report evaluation: 1,365 atomic statements from ten reports included 134 module-referencing statements and 1,231 clinical interpretations.All module-referencing statements were grounded, and 97.0% of interpretive statements were judged correct.
- Report evaluation: 82.4% of 159 report citations fully or partially supported their associated statements.Citation support was assessed separately from the groundedness of quantitative module references and correctness of interpretations.
- Report evaluation: No clinically relevant omissions occurred across 70 section-level assessments, while three contradictions occurred in two reports.Among 111 errors, 86.5% had no potential for harm and 13.5% had mild-to-moderate potential; none were severe.
- Clinical validation: LUCAID had the lowest mean absolute error for tumor cellularity assessment at 3.46 percentage points, compared with 8.9–21.23 percentage points for pathologists.Correlations with the reference were strong across the evaluated tasks, including ρ = 0.975 for tumor cellularity.
- Clinical validation: 93.0% overall clinical action concordance was achieved with the expert-panel reference standard, compared with 68.3–81.1% for individual pathologists.LUCAID achieved 98.6% concordance for molecular-testing eligibility, 89.9% for PD-L1 treatment stratification, and 92.5% for MET classification.
3 Discussion
LUCAID combines validated quantitative modules with an agent that orchestrates complete lung cancer pathology workflows and generates structured reports. Prospective validation and report review support reproducible clinical assessment, while important scope and deployment questions remain open.
- Clinical validation: 93.0% concordance with the expert-panel reference standard was achieved across clinically actionable decisions, exceeding the five pathologists’ 68.3–81.1%.The prospective multicenter validation assessed real-world lung cancer cases across decisions guiding systemic therapy selection.
- Workflow coverage: Nine quantitative modules cover quality control, tumor detection, subtyping, TME profiling, cellularity, IHC phenotyping, biomarker scoring and structured reporting.The agent orchestrates analyses from H&E-based tissue profiling through immunohistochemical assessment and case-level reporting.
- Reproducibility: Pathologists disagreed on the clinical action category in 122 of 338 assessments, underscoring variability in conventional semi-quantitative evaluation.The authors frame standardization and reproducibility as a principal potential benefit of AI in complex diagnostic decisions.
- Generalizable measurements: Validated biological primitives generalize across histological types and support cohort-scale measurements, including regression grading within 1.6–2.9 percentage points of joint pathologist assessment.The same measurements also support research-oriented immune-phenotype stratification based on lymphocyte density and PD-L1 scoring.
- Grounded reporting: Complete workflow coverage enables reports that integrate case parameters and ground interpretations in validated module measurements rather than unrestricted image-language reasoning.Threshold-based biomarker decisions require whole-slide counting and density estimation that vision-language models do not reliably perform.
- Limitations: Report evaluation covered ten cases, real-world clinical influence and routine pathologist interaction remain unestablished, and the system currently focuses on lung cancer.The authors describe report review as proof of principle rather than full-scale clinical validation.
4 Methods
The study assembled heterogeneous discovery, testing and prospective cohorts across institutions, scanners and staining contexts to support evaluation of LUCAID’s generalization. Prospective validation used consecutive multicenter lung cancer cases.
- Study design: Cohorts were assembled to maximize morphological and technical heterogeneity across tissue types, staining modalities and imaging domains.This design was intended to support model generalization.
- Image datasets: The H&E segmentation and cell-phenotyping test cohort included 158 primary and 47 metastatic lung tumors scanned across two platforms.The H&E quality-control cohort included 132 WSIs, while the IHC quality-control cohort included 106 WSIs.
- Cohorts: 1,001 surgically resected NSCLC tumors formed the bicentric discovery cohort, comprising 581 LUAD, 402 LUSC and 18 ASC cases.Cases were collected between 2006 and 2019 at Charité and University Hospital Cologne.
- Prospective validation: 70 consecutive lung cancer cases were prospectively analyzed between November 2024 and March 2025 in the nNGM program.The cases originated from Charité and five referring pathology sites in Berlin and Brandenburg.
4.2 Immunohistochemical Tumor Classification
Immunohistochemical classification used a marker panel and cell-level phenotyping followed by intensity-based aggregation into biomarker scores. Tumor cellularity was quantified with complementary cell-proportion and nuclear-area measures.
- IHC classification: 274 marker-specific WSIs comprised the IHC lung cancer classification panel spanning TTF-1, p40, CK7, CK5/6, SYP and CgA.The dataset included both positive and negative staining patterns for each marker.
- Tumor cellularity: Tumor cellularity was estimated as carcinoma-cell proportion among detected cells and carcinoma nuclear area relative to total nuclear area.AI estimates were generated within pathologist-annotated tumor-rich regions before comparison with routine assessments.
- Validation: AI-derived and pathologist cellularity estimates were compared with KRAS variant allele fraction using Spearman correlation and mean absolute error.The comparison evaluated agreement with a molecular surrogate for tumor-derived DNA fraction.
- Scoring: MET and TROP-2 H-scores combine percentages of weakly, moderately and strongly positive tumor cells weighted by 1, 2 and 3.PD-L1 expression uses tumor proportion score, defined as the percentage of PD-L1-positive tumor cells.
4.5 Image Analysis Pipeline
LUCAID’s image-analysis pipeline builds on the Atlas foundation model and converts cell-level classifications into slide-level biomarker scores. An orchestrating agent retrieves validated outputs and generates reports contextualized by clinical thresholds or reference distributions.
- Model pipeline: Atlas-based modules support H&E quality control, tissue segmentation and cell phenotyping, while LUCAID adds IHC phenotyping and PD-L1, MET and TROP-2 expression scoring.The modules follow the Atlas H&E-TME development framework and extend it to lung cancer precision pathology.
- Expression scoring: Carcinoma cells are classified by staining intensity, and slide-level H-scores or tumor proportion scores aggregate predictions across all classified tumor cells.MET and TROP-2 use negative, weak, moderate and strong intensity categories.
- Agent orchestration: Each diagnostic module is independently callable through the Model Context Protocol, forming the LUCAID toolbox orchestrated by Claude Opus 4.8.The architecture exposes modules as tools for LLM-based agent models.
- Report generation: Reports use pre-computed validated module readouts, so quantitative values are not produced by the language model itself.Retrieved outputs are assembled into predefined report sections corresponding to diagnostic modules.
- Metric contextualization: Diagnostic and therapeutic markers are interpreted against clinical cut-offs, whereas tissue-composition and microenvironment metrics use reference distributions from an independent lung cancer cohort.This distinction provides contextualization for metrics lacking a single accepted threshold.
4.8 Molecular Analysis
Molecular profiling was conducted within the routine diagnostic workflow, from pathologist-guided tumor-region selection through targeted sequencing and variant filtering. The clinical reference standard was established by joint expert-panel review after a washout period, with LUCAID heatmaps available only for qualitative spatial inspection.
- Workflow: Pathologists identified and annotated tumor-rich regions by light microscopy before DNA extraction.Five to twenty serial 5-µm FFPE sections were prepared for molecular profiling.
- Workflow: DNA was extracted semi-automatically, quantified with a Qubit HS assay, and used for targeted library preparation.Libraries used 80 ng genomic DNA input with the AmpliSeq for Illumina Cancer Hotspot nNGM Panel v3.
- Sequencing: The targeted panel covered hotspot regions in 53 cancer-associated genes.The listed genes included actionable and clinically relevant loci such as EGFR, KRAS, MET, RET, and ROS1.
- Variant calling: Sequencing reads were aligned to hg19, variants were called with SEQUENCE Pilot, and variants below a 5% minimum allele frequency were filtered out.Pathogenic or likely pathogenic variants were retained for downstream analysis.
- Reference standard: All 70 cases were jointly re-evaluated by five thoracic pathologists to establish task-specific clinical reference classifications.Adjudication followed a washout period of at least three months after the initial assessment.
- Reference standard: During adjudication, LUCAID heatmaps supported qualitative spatial inspection, while numeric module readouts and categorical classifications were withheld.This design limited panel review to tissue and cell-distribution overlays rather than LUCAID’s quantitative outputs.
4.10 Clinical Action Concordance
Clinical action concordance compared categorical LUCAID module outputs and pathologist assessments with an expert-panel reference standard. The evaluated decisions included molecular-testing adequacy, PD-L1 expression, MET expression, and membranous or cytoplasmic TROP-2 expression.
- Concordance definition: CAC measured the proportion of cases in which each rater’s categorical assessment matched the reference-standard classification.LUCAID’s module readouts were evaluated separately from the agent’s free-text report output.
- Clinical categories: Molecular-testing adequacy was classified by tumor cellularity as < 10% versus ≥10%.These thresholds represented the clinical decision categories used for concordance analysis.
- Clinical categories: PD-L1 expression was categorized by TPS as < 1%, 1–49%, or ≥50%.The categories were compared independently against the reference standard.
- Clinical categories: MET and membranous or cytoplasmic TROP-2 expression were categorized using H-score ranges for negative or weak, moderate, and high expression.The ranges were 0–100, 100–200, and 200–300, respectively.
- Additional analyses: Spatial tumor-microenvironment features were derived from seven AI-phenotyped cell classes and related to overall survival in stage-adjusted Cox models.Analyses reported hazard ratios, 95% confidence intervals, and log-rank p values with false-discovery-rate control.
4.13 Statistical Analysis
Statistical analyses evaluated categorical module performance against pathologist annotations and continuous predictions against reference standards. The study also specified exclusion, bootstrap, significance, and ethical-reporting procedures.
- Performance metrics: F1 scores quantified quality-control, segmentation, cell-phenotyping, and IHC expression-scoring performance against pathologist annotations.Scores were reported per class and as macro averages at the pixel or cell level, as appropriate.
- Continuous analyses: Continuous predictions were assessed with Pearson and Spearman correlations and mean absolute error with 95% confidence intervals.Confidence intervals were derived from 1,000 bootstrap iterations using two-tailed tests with α = 0.05.
- Exclusions: Slides containing fewer than 100 tumor cells were excluded from percentage-based analyses.Tumor-cellularity predictions were compared with expert-panel consensus or KRAS VAF, depending on the analysis.
- Ethics and data: Ethical approval was granted by Charité – Universitätsmedizin Berlin, and patients provided written informed consent for scientific use of archived tissue and clinical data.Race, ethnicity, and socioeconomic status were unavailable because they were not collected in routine diagnostic documentation.
Competing interests
Several authors reported affiliations with Aignostics, while the other authors declared no conflict of interest.
- Disclosures: F.K., M.A., and K.R.M. are co-founders of Aignostics.Additional reported relationships included advisory, board, technical-advisor, employment, and scientific-advisory roles at Aignostics.
Supplementary Tables And Supplementary Figures
The supplementary materials provide representative visual examples, cohort characteristics, feature definitions, and discordant-case comparisons supporting LUCAID’s pathology analyses.
- Supplementary Figures: Representative figures show LUCAID overlays for immunohistochemical marker scoring, H&E tumor detection, cell phenotyping, and prospective validation analyses.The materials include positive and negative marker examples, tissue-segmentation overlays across metastatic lung cancer subtypes, molecular validation cases, and rater-correlation analyses.
- Supplementary Figures: LUCAID estimated tumor cellularity at 7.0% and 8.8%, whereas all pathologists estimated 15.0–40.0% in the discordant cases.The cases illustrate disagreement between LUCAID and pathologists for tumor cell content estimation.
- Supplementary Tables: The supplementary tables cover prospective and discovery cohort characteristics, AI-quantified spatial features, feature definitions, and clinical analysis variables.Listed fields include formulas, units, hazard ratios, confidence intervals, p-values, false-discovery-rate-adjusted q-values, split rules, cutoffs, sample sizes, and events.
- Supplementary Tables: Spatial measurements include tissue composition, tumor cell density, interface immunity, lymphocyte distance, necrosis, dispersion, tissue NLR, and UICC stage.The listed measures combine established composition or density features with novel spatial distances and neighborhood-related quantities.
B Report Evaluation
Report evaluation separates statement-level fidelity and clinical interpretation from whole-report consistency and clinical-harm assessment.
- Claim-based Evaluation: Claim-based evaluation decomposes reports into individual statements and checks fidelity to referenced module outputs and correctness of clinical inferences.Groundedness preserves both the direction and magnitude of the underlying measurement.
- Report-level Evaluation: Report-level evaluation checks internal consistency and whether clinically relevant findings appear where a reader would expect them.This assessment considers the report as a whole rather than only its atomic statements.
- Clinical Harm Evaluation: Errors are additionally rated by their potential clinical harm so that review priorities reflect possible effects on patient care.The harm assessment considers both the extent and likelihood of errors.
- Claim-based Evaluation: Grounded statements preserve direction and magnitude, partially grounded statements preserve direction but misstate degree, and ungrounded statements lack support or contradict the data.An output of 35% described as “extensive” is partially grounded, while 8% described as “moderate–high” is ungrounded.
C Report Prompt
The report prompt specifies pathology-focused content, marker interpretation rules, actionable molecular findings, formatting constraints, and deterministic traffic-light assignments.
- Diagnostic Marker Section: The diagnostic marker section reports lineage markers as positive or negative for subtype classification without therapeutic or molecular-testing recommendations.The prompt specifies marker-profile rules for ADC, LCNEC, and squamous cell carcinoma patterns.
- Molecular Panel Section: The molecular panel section reports only detected mutations from the nNGM sequencing panel and identifies actionable alterations with relevant therapies or thresholds.The prompt lists KRAS G12C, EGFR alterations, BRAF V600E, ALK/ROS1/RET fusions, MET exon 14 skipping, ERBB2, and NTRK fusions.
- Formatting and Style: The prompt requires concise American English pathology-report sentences, factual reporting, preserved marker names, critical thresholds, and no fabricated data.It also prohibits dashes and semicolons and limits each row to one short clinical sentence.
- Traffic-Light Rules: Traffic-light rules assign green or amber to therapeutic markers, reserve red for poor prognosis or insufficient cellularity, and return null for tissue-composition and lineage rows.Therapeutic markers above threshold receive green, below threshold receive amber, while tissue-composition and TME rows are colored deterministically from a reference cohort.