Source-linked AI summary

An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

Weike Zhao, Chaoyi Wu, Yanjie Fan, Xiaoman Zhang, Pengcheng Qiu, Yuze Sun, Xiao Zhou, Yanfeng Wang, Xin Sun, Ya Zhang, Yongguo Yu, Kun Sun, Weidi Xie

arXiv:2506.20430v3cs.CLcs.AIcs.CVcs.MA

TL;DR

Rare-disease diagnosis is difficult because cases are heterogeneous, data are scarce, and knowledge evolves rapidly, while clinical workflows require interpretable evidence. DeepRare addresses this gap with a multi-agent LLM system that integrates heterogeneous patient data and specialized knowledge tools to produce ranked, evidence-linked diagnoses. It consistently outperformed existing methods across diverse evaluations, while physician validation supported the traceability of its reasoning chains.

  • Problem

    Rare-disease diagnosis requires multidisciplinary reasoning despite heterogeneous symptoms, scarce cases, rapidly evolving knowledge, and clinical demands for transparency.

  • Method

    DeepRare uses a three-tier agentic LLM architecture that orchestrates specialized tools and medical sources to analyze free text, HPO terms, genomic data, and external cases.

  • Results

    DeepRare consistently outperformed existing methods across diverse datasets, specialties, and input modalities, with evidence-based reasoning chains validated through expert clinical assessment.

  • Takeaways & Limitations

    DeepRare provides rare-disease decision support that combines diagnostic performance with transparent, verifiable reasoning for clinical information gathering and decision-making.

  • Takeaways & Limitations

    The system primarily targets undiagnosed patients already aware of rare diseases, while non-specialist screening remains a future direction.

Abstract

from arXiv · show

Rare diseases affect over 300 million individuals worldwide, yet timely and accurate diagnosis remains an urgent challenge. Patients often endure a prolonged diagnostic odyssey exceeding five years, marked by repeated referrals, misdiagnoses, and unnecessary interventions, leading to delayed treatment and substantial emotional and economic burdens. Here we present DeepRare, a multi-agent system for rare disease differential diagnosis decision support powered by large language models, integrating over 40 specialized tools and up-to-date knowledge sources. DeepRare processes heterogeneous clinical inputs, including free-text descriptions, structured Human Phenotype Ontology terms, and genetic testing results, to generate ranked diagnostic hypotheses with transparent reasoning linked to verifiable medical evidence. Evaluated across nine datasets from literature, case reports and clinical centres across Asia, North America and Europe spanning 14 medical specialties, DeepRare demonstrates exceptional performance on 3,134 diseases. In human-phenotype-ontology-based tasks, it achieves an average Recall@1 of 57.18%, outperforming the next-best method by 23.79%; in multi-modal tests, it reaches 69.1% compared with Exomiser's 55.9% on 168 cases. Expert review achieved 95.4% agreement on its reasoning chains, confirming their validity and traceability. Our work not only advances rare disease diagnosis but also demonstrates how the latest powerful large-language-model-driven agentic systems can reshape current clinical workflows.

1 Main

Rare diseases impose a substantial global burden but remain difficult to diagnose because of heterogeneous presentations, scarce cases, evolving knowledge, and limited interpretability. DeepRare addresses these challenges with an agentic LLM system that integrates heterogeneous clinical inputs, specialized tools, and traceable evidence.

  • Over 300 million people worldwide are affected by more than 7,000 rare diseases, with approximately 80% having genetic origins.
  • Patients experience a diagnostic odyssey averaging over five years, involving repeated referrals, misdiagnoses, and unnecessary interventions.
  • Rare-disease AI must handle multidisciplinary symptoms, scarce training cases, rapidly changing knowledge, and demands for transparency and traceability.
  • Agentic LLM systems orchestrate specialized tools and external knowledge sources while supporting few-shot and zero-shot scenarios where annotated data are scarce.
  • DeepRare processes free-text descriptions, HPO terms, and genomic results to produce ranked diagnoses with reasoning linked to verifiable medical evidence.
  • Across 8 datasets covering 3,134 diseases and 14 specialties, DeepRare consistently achieves superior diagnostic accuracy.

2 Results

DeepRare combines a multi-tier agentic architecture with external evidence retrieval and self-reflection, and it consistently outperforms comparison methods across datasets, specialties, modalities, and long-tail diseases. Its reasoning chains show high physician agreement, although errors remain concentrated in phenotype weighting and phenotypic mimicry.

  • Framework: DeepRare uses a central LLM host, specialized agent servers, and heterogeneous web-scale medical sources to orchestrate diagnostic reasoning.
  • HPO-wise analysis: 57.18% top-1 diagnosis recall was achieved in HPO-wise evaluation, exceeding Claude-3.7-Sonnet-thinking at 33.39%.
  • Dataset comparisons: 78% and 85% Recall@1 and Recall@3 were achieved on RareBench, surpassing PubCaseFinder by 30% and 20%.
  • Disease-level analysis: 31.8% of long-tail diseases reached Recall@1 > 0.8, compared with 23.5% for DeepSeek-V3 and 26.6% for DeepSeek-R1.
  • Expert comparison: Recall@5 reached 78.5%, exceeding clinicians’ average accuracy of 65.6%.
  • Traceable reasoning validation: 95.4% average reference accuracy was achieved at the case level, while incorrect references involved hallucinated URLs or sources unrelated to the true disease.
  • Ablation study: Agentic orchestration improved GPT-4o’s average Recall@1 from 26.11% to 54.67% and DeepSeek-V3’s from 26.99% to 56.94%.

3 Discussion

The discussion presents DeepRare as a multimodal, evidence-grounded decision-support system with consistent benchmark performance and interpretable reasoning. The authors position it as potentially useful for clinical information gathering and for physicians with limited rare-disease expertise.

  • DeepRare accepts chief complaints, genetic data, and detailed phenotypes while generating transparent reasoning chains for clinical decision support.
  • Across datasets, disease categories, medical centers, specialties, and input modalities, DeepRare consistently outperformed existing methods.
  • Evidence-based reasoning chains with verifiable references were reported to reduce literature-review and case-research time and minimize misdiagnosis-related costs.
  • Consistent performance across specialties suggests potential value for non-specialist physicians and settings with limited access to specialized care.

4 Limitations

DeepRare's current scope leaves several components for future development, including additional knowledge sources, refined retrieval, screening, and validated patient interaction.

  • The agentic architecture does not yet fully incorporate potentially valuable data sources.Its MCP-like plugin interface is intended to support future integration of additional rare disease knowledge systems and bioinformatics tools.
  • Phenotypic information is currently processed in aggregate rather than through refined, adaptive retrieval.Future retrieval mechanisms could optimize knowledge curation and potentially improve diagnostic precision.
  • DeepRare primarily targets patients aware of rare diseases but lacking a precise diagnosis, not initial screening in non-specialist settings.The authors identify screening as a future direction and report a preliminary attempt in Supplementary Material 13.7.
  • Patient-interaction modules have not yet been experimentally evaluated because suitable validation datasets are unavailable.The authors plan further investigation as such datasets become available.
  • The authors characterize these issues as areas for ongoing development rather than fundamental limitations.Future work also aims to extend the framework to rare disease treatment and prognosis prediction.

6 Author Contributions

The authors describe their respective contributions to the study's conception, computational design, clinical design, medical aspects, and data acquisition.

  • All listed authors meet the ICMJE four criteria for authorship, and W.Z., C.W., and Y.F. contributed equally.Y.Z., Y.Y., K.S., and W.X. are identified as corresponding authors.
  • W.Z. and C.W. led computational algorithm design, while Y.F. led clinical design and medical aspects.The authors collectively contributed to the conception and design of the study and data acquisition.

7 Competing Interests

The authors declare no competing interests or other interests perceived to influence the paper's reported results or discussion.

  • The authors declare no competing interests as defined by Nature Portfolio.They also declare no other interests that might be perceived to influence the reported results or discussion.

8 Methods

DeepRare is a modular multi-agent system that combines phenotype and genotype processing, external evidence retrieval, diagnostic synthesis, self-reflection, and traceable rationale generation. Its evaluation uses heterogeneous datasets and case-retrieval components to support rare disease differential diagnosis.

  • System architecture: DeepRare comprises a central host with memory, specialized local agent servers, and heterogeneous diagnostic data sources.The host coordinates operations, local servers interface with tailored tools, and external sources provide evidence.
  • System architecture: Patient input consists of phenotype and genotype information, with phenotype represented by free-text descriptions, HPO terms, or both.Either phenotype or genotype may be absent, and the formal input is I = {P, G} with P = (T, H).
  • Diagnostic synthesis: Final rationales are free-text explanations linked to medical sources, with reference verification removing invalid URLs to mitigate hallucination.The system produces ranked diagnoses and evidence-grounded explanations traceable to literature, guidelines, and similar cases.
  • Workflow: The system operates through information collection followed by self-reflective diagnosis, orchestrated by the central host.These stages coordinate specialized evidence gathering and subsequent diagnostic decision-making.
  • Phenotype analysis: Phenotype processing standardizes free-text into HPO entities, retrieves supporting documents and similar cases, and applies bioinformatics tools for diagnostic suggestions.Search agents consult external sources and the memory bank to avoid duplicating previously recorded items.
  • Diagnostic synthesis: The central host retains original free-text reports alongside processed inputs to preserve temporal and contextual details potentially lost during HPO extraction.It generates an initialized disease list and subsequently synthesizes collected information for diagnosis.
  • Genotype analysis: Genotype processing includes VCF annotation, variant ranking, and synthetic analysis when genotype data are provided.Variant prioritization considers functional impact, allele frequency, conservation, and predicted pathogenicity; the host then interprets variants and gene–phenotype relationships.
  • Evaluation: Case retrieval uses embedding-based top-50 candidate selection followed by MedCPT-Cross-Encoder re-ranking, and this two-stage approach outperforms alternatives.The method is designed to balance computational efficiency with the clinical relevance of retrieved cases.

9 Data Availability

The study uses six data sources, combining public datasets with hospital datasets deposited in public repositories. Genetic and clinical records from the hospital datasets require controlled-access applications for research use.

  • The study utilized six data sources: RareBench, MyGene2, DDD, MIMIC-IV, Xinhua Hospital, and Hunan Hospital.
  • The first four datasets are publicly available through designated dataset repositories.
  • The Xinhua Hospital and Hunan Hospital datasets are deposited in public repositories at China’s National Genomics Data Center.
  • Variant data and clinical metadata are archived under separate Genome Variation Map and OMIX accession numbers.
  • Access to genetic data and clinical records requires a research proposal and institutional IRB approval through the Xinhua Hospital Data Access Committee.

11 Figure Legends

The figure legends present traceable reasoning outputs and the system architecture’s core elements, including diagnostic candidates, evidence sources, and supporting knowledge resources.

  • The output includes traceable reasoning for a diagnostic candidate such as Ehlers-Danlos Syndrome.
  • The system’s decision-making layer combines a memory bank with LLM-powered agent orchestration.
  • Specialized components include a disease normalizer and case searcher.
  • The supporting knowledge environment includes case banks, literature, guidelines, knowledge bases, and genetic resources.
  • The example output describes Ehlers-Danlos Syndrome using clinical references and subtype-related information.

Human

Figures 4 and 5 assess DeepRare’s traceable reasoning and examine how central-host models and agentic components contribute to diagnostic performance.

  • Figure 4 evaluates human expert validation of DeepRare’s traceable reasoning chains and failure modes.
  • Figure 5 compares five LLMs as central hosts across eight rare disease datasets.
  • Figure 5 compares baseline LLM performance with the corresponding powered agentic DeepRare systems.
  • Figure 5 analyzes contributions from similar case retrieval, web knowledge integration, and self-reflection relative to baseline GPT-4o.

12 Extended Data

The extended data describe DeepRare’s architecture, cohort curation, benchmark characteristics, and web-application workflow. Together, these materials show how inputs are processed, datasets allocated, and diagnostic reports produced.

  • Extended Data Figure 1: DeepRare accepts free-text information, structured HPO IDs, or combinations of these inputs.
  • Extended Data Figure 1: Its architecture comprises a central host with memory, specialized agent servers, and external medical data sources.
  • Extended Data Figure 1: The main workflow has information collection and self-reflection diagnosis stages.
  • Extended Data Figure 2: MIMIC-IV-Note was reduced from 331,794 to 9,185 cases, while Xinhua Hospital was reduced from 352,425 to 5,820 after exclusions and filtering.
  • Extended Data Table 1: The benchmark table covers case distributions, HPO-based phenotypic complexity, disease spectrum, provenance, and genetic annotation status.
  • Extended Data Figure 3: The web application progresses through clinical data entry, systematic inquiry, HPO mapping, diagnostic analysis, and report downloading.

13 Supplementary

The supplementary analyses examine DeepRare’s workflow, component contributions, phenotype extraction, and screening extension. Across these evaluations, the system combines complementary agentic modules with evidence-linked reasoning and supports rare-disease triage in mixed cohorts.

  • Main Workflow: DeepRare’s workflow accepts free-text, HPO items, and genetic variants, producing a final diagnosis list with rationale explanations.The workflow includes information collection, phenotype retrieval, genotype analysis, self-reflective diagnosis, and final rationale generation.
  • Ablation Study: +2.28, +8.38, and +4.57 in R@1 are reported for knowledge tools and reflective reasoning in sparse-case or unseen-disease settings.The unseen-disease subset contains diagnoses absent from the case bank, while retrieval can still identify diseases with similar manifestations.
  • Ablation Study: +60.00 and +34.56 in R@1 are reported for case retrieval on RareBench MME and RAMEDIS, respectively, where similar typical cases are available.Tool-calling with reflection also achieves a +40.00 gain on MME over the base LLM.
  • Ablation Study: The full system outperforms individual configurations, often by more than 10 points on R@1, indicating complementary contributions from its components.Ablations compare LLM-only, case search, and tool-calling with reflection across RareBench and an unseen-disease subset.
  • Phenotype Extractor Comparison: Phenotype extraction errors include non-canonical HPO phrasing, associative drift, and loss of qualifiers during summarization.The reported examples motivate canonical-term normalization, ontology-constrained decoding, and qualifier-preserving summaries.
  • Screening Module: 93.94% accuracy and a 0.8526 balanced F1-score are achieved by the rare-versus-common disease triage extension on 4,825 common or healthy and 1,227 rare-disease cases.The extension is implemented as a callable Rare Disease Discriminator agent and is reported to reduce false positives earlier in the care pathway.

13.9 Case Study

Across case studies, DeepRare ranked diagnoses by combining phenotype–genotype concordance with diagnostic reasoning, while explicitly downgrading candidates supported mainly by weak phenotype overlap or uncertain variants.

  • Optic and ocular phenotypes: OPA1-related optic atrophy ranked first because visual impairment, fundus abnormalities, and a likely pathogenic stop-gain variant showed strong concordance.Bietti dystrophy remained second, while several high Exomiser-ranked candidates were lowered for limited ocular relevance.
  • Cross-case interpretation: The case studies illustrate that DeepRare can prioritize phenotype-concordant diagnoses over candidates elevated primarily by computational ranking or uncertain variants.Examples include deprioritizing KANSL1, PDGFRB, and RYR1-associated candidates when hallmark clinical features were absent.
  • Endocrine phenotype: FGFR1-associated hypogonadotropic hypogonadism ranked first for micropenis, testicular atrophy, and short stature consistent with impaired GnRH secretion.Olfactory assessment was proposed to distinguish Kallmann syndrome from normosmic congenital hypogonadotropic hypogonadism.
  • Renal and hematologic phenotypes: SLC4A1-related distal renal tubular acidosis ranked highest despite a benign or likely benign ClinVar label, based on a 0.993 Exomiser score and phenotype concordance.The recessive dRTA4 differential required hemolysis assessment, biallelic confirmation, and second-hit evaluation.
  • Neurodevelopmental phenotype: Kleefstra syndrome 2 ranked first for developmental delay, microcephaly, and delayed speech, supported by a KMT2C stop-gain variant, Exomiser score 0.997, and phenotype score 0.969.CAMK2B-related intellectual developmental disorder remained a high-ranked alternative pending segregation or functional evidence.

13.10 Exomiser Config  

The Exomiser configuration specifies reference data, output settings, population and pathogenicity filters, inheritance models, and prioritization modules.

  • GRCh37 is set as the genome assembly, with TSV and HTML outputs written using a configurable prefix.
  • Population-frequency resources include Thousand Genomes, TOPMed, UK10K, ESP, and multiple gnomAD ancestry-specific datasets.
  • Pathogenicity annotations use PolyPhen, MutationTaster, and SIFT, with analysis restricted to passing variants.
  • The template includes autosomal-dominant, autosomal-recessive, and X-recessive inheritance models with distinct homozygous and compound-heterozygous weights.
  • Variant processing includes failed-variant, variant-effect, frequency, pathogenicity, and inheritance filters before OMIM and hiPhive prioritization.
Loading 2506.20430v3…