Source-linked AI summary

Towards Conversational Diagnostic AI

Tao Tu, Anil Palepu, Mike Schaekermann, Khaled Saab, Jan Freyberg, Ryutaro Tanno, Amy Wang, Brenna Li, Mohamed Amin, Nenad Tomasev, Shekoofeh Azizi, Karan Singhal, Yong Cheng, Le Hou, Albert Webson, Kavita Kulkarni, S Sara Mahdavi, Christopher Semturs, Juraj Gottweis, Joelle Barral, Katherine Chou, Greg S Corrado, Yossi Matias, Alan Karthikesalingam, Vivek Natarajan

arXiv:2401.05654v1cs.AIcs.CLcs.LG

TL;DR

Clinical LLMs have not yet been rigorously evaluated for history-taking and diagnostic dialogue, despite the importance of acquiring information under uncertainty and maintaining patient rapport. This paper introduces AMIE, an LLM system optimized for clinical dialogue, and evaluates it against primary care physicians in simulated text-based consultations, finding more accurate differential diagnoses and comparable information elicitation.

  • Problem

    Prior health-focused LLM research had not rigorously examined clinical history-taking and diagnostic dialogue or compared these capabilities with expert clinicians.

  • Method

    AMIE is an LLM-based system optimized for clinical history-taking and diagnostic dialogue, using chain-of-reasoning and evaluated against PCPs in a randomized, double-blind crossover study with human simulated patients.

  • Results

    AMIE produced more accurate and complete differential diagnoses than board-certified PCPs and was as adept as PCPs at eliciting pertinent information during simulated consultations.

  • Takeaways & Limitations

    AMIE's simulated-consultation performance represents a milestone toward conversational diagnostic AI assessed across multiple clinically relevant axes.

  • Takeaways & Limitations

    The synchronous text-chat experiments did not emulate the expected quality of diagnostic dialogue in real clinical practice and excluded several settings and specialties.

Abstract

from arXiv · show

At the heart of medicine lies the physician-patient dialogue, where skillful history-taking paves the way for accurate diagnosis, effective management, and enduring trust. Artificial Intelligence (AI) systems capable of diagnostic dialogue could increase accessibility, consistency, and quality of care. However, approximating clinicians' expertise is an outstanding grand challenge. Here, we introduce AMIE (Articulate Medical Intelligence Explorer), a Large Language Model (LLM) based AI system optimized for diagnostic dialogue. AMIE uses a novel self-play based simulated environment with automated feedback mechanisms for scaling learning across diverse disease conditions, specialties, and contexts. We designed a framework for evaluating clinically-meaningful axes of performance including history-taking, diagnostic accuracy, management reasoning, communication skills, and empathy. We compared AMIE's performance to that of primary care physicians (PCPs) in a randomized, double-blind crossover study of text-based consultations with validated patient actors in the style of an Objective Structured Clinical Examination (OSCE). The study included 149 case scenarios from clinical providers in Canada, the UK, and India, 20 PCPs for comparison with AMIE, and evaluations by specialist physicians and patient actors. AMIE demonstrated greater diagnostic accuracy and superior performance on 28 of 32 axes according to specialist physicians and 24 of 26 axes according to patient actors. Our research has several limitations and should be interpreted with appropriate caution. Clinicians were limited to unfamiliar synchronous text-chat which permits large-scale LLM-patient interactions but is not representative of usual clinical practice. While further research is required before AMIE could be translated to real-world settings, the results represent a milestone towards conversational diagnostic AI.

1 Introduction

AMIE addresses the unmet need for AI systems that can conduct clinically meaningful diagnostic dialogue, beyond single-turn medical question answering. The work introduces AMIE and evaluates it against primary care physicians across diagnostic, communication, and patient-centred dimensions.

  • Prior medical LLM research had not rigorously examined clinical history-taking and diagnostic dialogue against expert clinicians.
  • Diagnostic dialogue combines history-taking, diagnosis, management reasoning, rapport, respect, and communication efficacy.
  • AMIE is an LLM-based AI system optimized for clinical history-taking and diagnostic dialogue.
  • AMIE uses simulated self-play with automated feedback and inference-time chain-of-reasoning to scale learning across medical contexts and improve dialogue quality.
  • The evaluation framework assessed history-taking, diagnostic reasoning, communication skills, and empathy from clinician-centred and patient-centred perspectives.
  • 149 case scenarios enabled randomized comparison of AMIE with primary care physicians in remote text-based OSCE consultations with validated patient actors.

2 AMIE: An LLM based AI System for Diagnostic Dialogue

AMIE combines real-world medical data with simulated self-play to develop diagnostic dialogue and reasoning capabilities. Its training environment uses iterative feedback, broad condition coverage, and inference-time chain-of-reasoning to refine conversational responses.

  • AMIE was instruction fine-tuned on medical question-answering, reasoning, summarization, and real-world dialogue datasets.
  • The real-world dialogue dataset contained 98,919 transcripts from over 1,000 clinicians, spanning 51 specialties and 168 medical conditions or visit reasons.
  • Real-world transcripts were limited by incomplete medical coverage and noisy language, motivating a simulated learning environment.
  • An inner loop used critic feedback to refine simulated conversations, while an outer loop added refined dialogues to later fine-tuning iterations.
  • Each self-play iteration generated 11,686 dialogues from 5,230 medical conditions, expanding training across diverse scenarios.
  • Simulated dialogues addressed limited high-quality labelled conversation data and improved generalization and adaptability across medical contexts.
  • AMIE was built on PaLM 2 and used a three-step chain-of-reasoning process before generating each dialogue-turn response.

3 Evaluation

The evaluation used a blinded remote OSCE with standardized text-chat consultations, comparing AMIE with PCPs across 149 scenarios and assessing diagnostic dialogue through clinician, patient-actor, and automated measures.

  • Remote OSCE Study Design: 20 board-certified PCPs and 20 trained patient actors participated in the remote OSCE comparison across scenarios from Canada, India, and the UK.The study produced 149 distinct simulated patients for randomized consultations.
  • Remote OSCE Study Design: Each simulated patient completed two synchronous text-chat consultations, one with a PCP and one with AMIE, with randomized ordering and patient actors blinded to the agent identity.PCPs and patient actors were primed and completed pilot consultations before the study.
  • Evaluation Criteria: Patient actors rated consultations using GMCPQ, PACES, and PCCBP-derived measures, while OSCE agents supplied ranked differential diagnoses and management recommendations.The questionnaires captured both patient-centred communication and clinical decision outputs.
  • Evaluation Criteria: Specialist physicians evaluated diagnostic accuracy, differential comprehensiveness, escalation, investigations, treatment, management, follow-up, communication, and confabulations.Specialists had access to the scenario pack, ground-truth differential, and additional accepted differentials.
  • Automated Evaluation: Model-based auto-evaluation assessed dialogue quality and diagnostic accuracy as economical alternatives to specialist assessment.The dialogue evaluator aligned with human ratings and was comparable to inter-specialist agreement on four PACES axes.

4 Results

AMIE outperformed PCPs in specialist-rated diagnostic accuracy and in most patient-actor and specialist-rated conversation and reasoning measures, with gains also supported by automated evaluation.

  • Diagnostic Accuracy: AMIE achieved significantly higher top-k diagnostic accuracy than PCPs across all k values for both ground-truth and accepted-differential matching.The comparison covered 149 scenarios, and bootstrap testing found all top-k differences significant after FDR correction.
  • Diagnostic Accuracy: AMIE matched or surpassed PCP diagnostic performance across all six specialties, with the largest improvements in respiratory and cardiovascular specialties.Specialty-level results are reported in Figure A.8.
  • Automated Evaluation: Auto-evaluation trends aligned with specialist assessments, and simulated dialogues after self-play were preferred more often than baseline dialogues without self-critique.The auto-evaluation methods showed similar diagnostic trends despite marginal differences in computed accuracy values.
  • Diagnostic Accuracy: AMIE’s superior differential-diagnosis performance remained when it processed PCP conversations, suggesting the gain was linked more to interpretation than information acquisition.Both AMIE conditions outperformed PCP differentials, while AMIE’s performance was markedly similar across the two information sources.
  • Conversation Quality: Patient actors rated AMIE significantly better than PCPs on 24 of 26 conversation-quality axes.The only non-significant differences were respecting patient privacy and acknowledging mistakes.
  • Conversation Quality: Specialist physicians rated AMIE significantly better than PCPs on 28 of 32 evaluation axes, including consultations, diagnoses, and management plans.The specialist-rating differences were statistically significant after FDR correction.

5 Related Work

Prior medical conversational-AI work largely emphasized symptom checking, transcription, or limited dialogue metrics, leaving clinical history-taking and diagnostic dialogue insufficiently evaluated against clinician expertise.

  • Medical Communication: Patient-centred communication frameworks identify relationship-building, information gathering, information provision, decision-making, emotional response, and behaviour support as core functions.These functions frame communication quality in clinical encounters.
  • Conversational AI: Conversational AI has advanced through transformers, large language models, alignment, self-improvement, and scalable oversight, but rigorous clinical task evaluation remains limited.The related work describes renewed interest in goal-oriented dialogue alongside gaps in clinical conversational assessment.
  • Medical Conversational AI: Most medical-AI consultation studies focused on symptom checkers, transcription, or generating plausible dialogue from clinical notes rather than full natural dialogue.Clinical dialogue datasets have also been used without comprehensive evaluation.
  • Evaluation Frameworks: Existing human-evaluation frameworks used broad criteria such as fluency, relevance, informativeness, expertise, or human likeness rather than detailed clinical communication and history-taking standards.The cited frameworks were less comprehensive and specific than criteria taught and practiced by medical professionals.

6 Discussion

AMIE outperformed PCPs in simulated text-based diagnostic consultations across multiple clinically meaningful axes, but the study’s design and evaluation framework limit how directly these findings generalize to routine care.

  • Overall findings: AMIE outperformed PCPs on simulated diagnostic conversations across multiple clinically meaningful axes in a randomized, double-blind crossover study.The study used human simulated patients in an OSCE-style evaluation, but was not designed to represent usual clinical conventions or practice.
  • Diagnostic performance: AMIE’s differential diagnoses were more accurate and complete than those of board-certified PCPs when specialist physicians evaluated them.Unlike fixed-input diagnostic tasks, the study required AMIE to acquire relevant information through conversation under uncertainty.
  • Variation across settings: Both AMIE and PCPs performed worse in obstetric/gynecology and internal medicine scenarios, while Canada-site diagnostic accuracy exceeded India-site accuracy without statistically significant differences.In 40 scenarios enacted in both sites, AMIE and PCP performance was equivalent; the study was not powered to compare specialties.
  • Conversational performance: AMIE was rated higher than PCPs for empathy and communication skills by both patient actors and specialist raters.These dimensions constituted a majority of the evaluated axes, although the text-only format may have disadvantaged clinicians by removing voice and non-verbal communication.
  • Study boundaries: The text-chat interface enabled a scalable and familiar comparison for LLM interaction but did not emulate diagnostic dialogue quality in routine clinical practice.Physicians may be more familiar with telephone or video consultations, or asynchronous text for specific needs, than synchronous diagnostic text chat.
  • Evaluation limitations: The evaluation framework was more granular than prior AI-dialogue studies, yet its axes were non-exhaustive and often subjective, with limited rater breadth and generalizability.The authors call for broader replication, more diverse raters, and participatory development involving patients and clinical and health-equity experts.

7 Conclusion

The study presents AMIE as a promising conversational diagnostic AI system, while emphasizing that substantial research and development remain before real-world use.

  • AMIE was assessed across multiple clinically relevant axes for conversational diagnostic medical AI.These included history-taking, diagnostic dialogue, communication, empathy, and related clinical considerations.
  • Translating simulated history-taking and diagnostic dialogue into real-world tools requires further work on safety, reliability, fairness, efficacy, and privacy.
  • The authors describe AMIE’s simulated-consultation performance as a milestone toward conversational diagnostic AI.
  • The paper suggests that conversational medical AI could support next-generation learning health systems if these requirements are met.

Data Availability

The paper provides limited data-availability information, identifies some datasets and scenario materials as accessible, and withholds AMIE’s model code and weights for safety reasons.

  • Some development datasets are open-source, including MedQA, and UK OSCE scenario packs are available online.
  • AMIE’s model code and weights are not open-sourced because unmonitored medical use poses safety concerns.
  • The authors plan to work with research partners, regulators, and providers to validate and explore safe onward uses of AMIE.
  • The study was funded by Alphabet Inc. or a subsidiary, and all authors are Alphabet employees who may own company stock.

Appendix

The appendix supplies additional analyses, evaluation materials, and examples concerning AMIE’s diagnostic-dialogue performance and automated assessment.

  • The appendix includes OSCE evaluation rubrics, simulated dialogues after self-critique, AMIE interfaces, and example consultations with OSCE agents.
  • It reports diagnostic-differential top-k accuracy by matching degree, specialty, location, and dialogue-turn count.
  • The appendix examines auto-evaluation of differential-diagnosis accuracy and qualitative criteria.
  • Additional analyses compare auto-evaluation rank ordering with specialist judgments and assess simulated dialogues generated through self-play.

A.1 OSCE Evaluation Rubrics

The OSCE evaluation rubrics assess clinical history-taking, reasoning, communication, patient-centered behavior, and empathy using structured questions and rating scales.

  • History-taking: The rubrics assess history-taking domains including the presenting complaint, systems review, past medical history, family history, and medication history.
  • Clinical communication: Communication ratings cover the accuracy, clarity, structure, comprehensiveness, and professionalism of clinical explanations.
  • Patient-centered communication: The instruments include criteria for seeking and addressing patient concerns and confirming patients’ knowledge and understanding.
  • Empathy: Empathy is assessed on a 5-point scale ranging from not at all empathic to extremely empathic by specialists and patient actors.
  • Patient-centered communication: Patient-centered communication includes fostering relationships, gathering information, providing information, making decisions, and enabling disease- and treatment-related behavior.
  • Clinical reasoning: Clinical reasoning is evaluated through the appropriateness and comprehensiveness of the differential diagnosis and the quality of the management plan.

A.2 Example of Simulated Dialogue After Self-critique

AMIE modifies its behavior during simulated dialogue through in-context self-play critique, illustrating iterative feedback in the simulated environment.

  • AMIE revises its behavior based on in-context feedback during inner-loop self-play.The example illustrates one preliminary round of iterative feedback; average dialogue-quality improvements are shown separately across four PACES criteria.

A.3 AMIE User Interfaces

The study used separate interfaces for online text-based consultation, patient actor ratings, and specialist physician evaluation.

  • The study included an interface for conducting online text-based consultations.
  • Patient actors used a dedicated interface to rate consultations.
  • Specialist physicians used a dedicated interface for evaluating consultations.

A.4 Example Consultation with OSCE Agents

The appendix presents paired example consultations in which the same scenario pack and patient actor interacted with AMIE and a PCP.

  • The examples compare AMIE and a PCP consulting with the same patient actor for the same scenario pack.

A.4.1 Example AMIE Consultation

This section documents diagnostic-accuracy analyses across matching thresholds, specialties, automated evaluation, dialogue sources, consultation length, locations, and self-play refinement. AMIE matched or surpassed PCP performance across all reported specialties, while automated evaluation aligned with specialist assessments and refined dialogues scored higher on average.

  • Specialty and matching analyses: Specialist-rated DDx accuracy was analyzed across increasingly strict matching levels, with significant AMIE–PCP differences at Relevant, Extremely Relevant, and Exact Match thresholds.
  • Specialty and matching analyses: AMIE matched or surpassed PCP diagnostic performance across all specialties evaluated by specialists.
  • Auto-evaluation: Automated evaluation showed overall performance trends aligned with specialist assessments despite marginal differences in accuracy values.
  • Auto-evaluation: Automated top-k DDx accuracy comparisons were evaluated against both the ground-truth diagnosis and the accepted differential, with significance emerging for k > 2 and k > 4, respectively.
  • Consultation inputs and dialogue length: AMIE’s diagnostic quality remained consistent whether it processed its own consultation or the corresponding PCP consultation.
  • Consultation inputs and dialogue length: Patient-word and turn counts were consistent between groups, but AMIE produced substantially more verbose responses that may have influenced specialist qualitative ratings.
  • Consultation inputs and dialogue length: AMIE’s average diagnostic accuracy plateaued within 10 turns for both AMIE and PCP conversations, with diminishing returns from additional information gathering.
  • Location and self-play analyses: AMIE showed higher average diagnostic performance in Canada than India, while matched-location scenarios produced consistent AMIE and PCP performance across locations.
Loading 2401.05654v1…