Source-linked AI summary
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
Luming Yang, Haoxian Liu, Siqing Li, Rong Jia, Yue Xiao, Guanhua Chen, Li Lu
TL;DR
Clinical interview training needs scalable alternatives to resource-intensive standardized patients. This study developed and evaluated a scaffolding-oriented multi-agent AI standardized patient system with patient, tutor, and evaluator agents in a randomized study of medical students. The system improved final examination performance and communication-related behaviors without improving diagnostic accuracy, while the authors caution against using it for autonomous competency or credentialing decisions.
Problem
Clinical education needs scalable ways to prepare medical students for safe, coherent patient interviews under uncertainty beyond resource-intensive standardized patient training.
Method
The study evaluated a scaffolding-oriented multi-agent AI standardized patient system using patient simulation, Socratic tutoring, and turn-level clinical-progress evaluation in randomized medical-student training.
Results
The multi-agent system improved final examination performance, especially communication and consultation conduct, while diagnostic accuracy did not differ between groups.
Takeaways & Limitations
Specialized LLM agents enhanced the process quality of simulated clinical interviews without artificially inflating examination outcomes.
Takeaways & Limitations
The system should not autonomously determine clinical competence, make progression or credentialing decisions, or replace human-led remediation.
Abstract
from arXiv · showhide
Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.
1 Introduction
The study addresses whether LLM agents organized around explicit scaffolding functions can improve simulated clinical interviewing beyond structured non-LLM materials. MeduAI-SP combines specialized agents with an annotated dataset to support scalable, pedagogically grounded training.
- Traditional standardized patient training is resource-intensive and difficult to scale, limiting repeated practice, timely feedback, and exposure to diverse cases.
- The central research question is whether explicitly scaffolded LLM agents improve simulated clinical interviewing compared with structured non-LLM learning materials from the same cases.
- Realistic patient interaction alone is insufficient; effective training must also support information elicitation, empathy, diagnostic comparison, and reflective feedback.
- MeduAI-SP assigns distinct instructional functions to an AI standardized patient, a Socratic teaching agent, and a turn-level evaluator without directly revealing diagnostic answers.
2 Results
The multi-agent AI standardized patient produced higher final examination performance than the structured control condition, with the clearest advantages in communication and selected history-taking behaviors. Diagnostic accuracy did not differ, while annotations indicated that scaffolding needs increased as consultations progressed.
- Checklist behaviors: The multi-agent condition also improved weighted checklist scores, while history-taking completeness showed a positive but nonsignificant association.Weighted checklist score: β = 11.7; 95% CI, 0.2 to 23.1; P = 0.046. History-taking completeness: β = 0.34; 95% CI, −0.05 to 0.73; P = 0.091.
- Final examination performance: The multi-agent group scored higher on the final examination than the control group, with means of 71.8% versus 55.6% and a significant between-group difference.The difference had P = 5.51 × 10−5 and Hedges’ g = −0.81; the negative sign reflected comparison coding.
- OSCE domains: Communication showed the strongest domain-level advantage, with mean OSCE scores of 3.53 versus 2.64 and a significant group difference.Regression estimated a 0.90-point increase in communication score for multi-agent assignment.
- Diagnostic outcomes: Diagnostic accuracy was similar between groups: 86% of control cases and 84% of multi-agent cases had correct initial diagnoses.The difference was not significant (P = 1.000), and worksheet completion also did not differ significantly.
- Checklist behaviors: Empathy showed the largest item-level difference, occurring 31 percentage points more often in the multi-agent group than in the control group.The difference remained significant after Holm correction (P = 8.30 × 10−4).
- Scaffolding needs: Approximately 24.1% of student utterances required instructional scaffolding, increasing from 16.1% early in consultations to 34.8% late.Intervention reasons included conversational impasse, knowledge gaps, communication breakdown, and premature closure.
3 Discussion
Compared with structured progressive-disclosure materials, multi-agent AI-SP training improved final examination performance, especially communication and selected consultation behaviors, without improving diagnostic accuracy. The findings support process-level evaluation and bounded, complementary use of AI-SP systems while leaving human-feedback equivalence and broader generalization unresolved.
- Primary finding: Diagnostic accuracy did not differ between groups, so the observed benefit was concentrated in how learners conducted the consultation.The discussion interprets this pattern as improved observable consultation behavior rather than unequal access to diagnostic knowledge.
- Communication and consultation behavior: Communication showed the strongest improvement, including higher OSCE communication scores and the largest checklist difference for empathic expression.The result concerns observable empathic communication during the simulated encounter, not a stable personal trait or superiority to human feedback.
- Communication and consultation behavior: The multi-agent group was more likely to use patient-centered communication, elicit relevant information, and complete selected consultation behaviors.Item-level gains included observable empathic expression, medication allergy history, and summarizing or confirming information.
- Interpretation: The findings suggest that early educational gains may appear in question asking, information confirmation, avoidance of premature closure, and patient-centered communication before diagnostic endpoints change.These behaviors are presented as enabling clinical reasoning rather than substitutes for diagnostic accuracy.
- Interpretation: The annotated consultation corpus provides process-level evidence for where instructional support is needed, including learners’ relative underuse of examination, testing, management, and empathic communication.The dataset can inform future tutor-agent design by identifying when and why learners need support.
4 Ethics Approval and Consent
The study received ethics approval, obtained informed consent from all participants, and protected participant privacy through de-identification.
- The study was reviewed and approved by the Ethics Committee of Guangzhou Medical University (Approval No. 202605032).
- Students confirmed informed consent through an online form before participation.
- Collected data were handled under strict privacy procedures and de-identified before analysis.
5 Author Contributions
The authors distributed contributions across conceptualization, methodology, development, data work, analysis, writing, clinical expertise, supervision, and technical support.
- Contributors covered conceptualization, methodology, model development, experimental implementation, data analysis, and manuscript preparation.
- The team conducted data curation, collection, annotation, formal analysis, visualization, writing, editing, and educational or medical consultation.
- Other contributions included LLM development support, supervision, project administration, and data collection.
8 Methods
The study used a randomized two-arm comparison of multi-agent scaffolded learning and structured progressive disclosure for simulated clinical interviewing. Students completed shared learning and examination procedures, while outcomes combined OSCE-aligned ratings, checklist behaviors, and a consultation corpus.
- Study design: The two-arm randomized controlled study compared multi-agent learning with structured control learning across two learning encounters and one examination encounter.Participants were third-year clinical medicine undergraduates randomized 1:1.
- Cases and platform: The experiment used three acute abdominal cases, with appendicitis and pancreatitis as learning cases and perforated peptic ulcer as the examination case.
- Multi-agent intervention: The platform included a patient agent, Socratic teaching agent, turn-based evaluator, and final evaluator for scaffolded learning.Tutor prompts addressed history taking, diagnostic reasoning, information summarization, and rapport-building communication without revealing diagnoses or answers.
- Control condition: The structured control used fixed-stage case information and omitted free AI or human dialogue, real-time Socratic prompts, tutor feedback, and formative feedback.
- Examination and outcomes: Both groups took an identical patient-only examination without student-facing teaching or monitoring agents, while a background evaluator applied the OSCE-aligned framework.
- Examination and outcomes: The outcome framework combined overall scores, four 1–5 global domains, checklist behavior coverage, diagnostic accuracy, and a corpus of 207 consultation sessions with 4,815 messages.The corpus was drawn from multi-agent learning and examination sessions because the control condition produced no free-text dialogue.
9 Data Availability
The study provides a de-identified, multi-expert annotated dataset of clinical interview interactions and aligned evaluation outputs. The dataset is intended for future research on pedagogically grounded AI standardized-patient systems.
- The dataset will be publicly available upon manuscript acceptance with a permanent DOI, while reviewers may request access during peer review.
- The dataset includes de-identified interview transcripts, structured checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes.
S.1 Annotated Dialogue Examples
The annotated examples show tutor scaffolding targeting premature closure, redundant questioning, inaccurate summarization, and requests for unavailable test results. Interventions use hints or corrections to promote systematic, accurate, and safer interviewing.
- Premature closure: Tutor interventions addressed premature closure by redirecting students toward foundational history-taking or confirmatory examination and testing.Examples covered leading questions before establishing the presentation and moving to low-yield questions after recognizing high-risk features.
- Repetition / failure to use known information: Hints encouraged students to avoid redundant questioning and build on information already obtained during the encounter.The example intervention explicitly identified repeated questions about pain radiation and recommended moving to other associated symptoms or focused examination.
- Critical error: Corrections targeted distorted summaries by requiring students to restate the patient’s history accurately without adding unconfirmed symptoms.The example contrasts the patient’s reported timing and pain description with unsupported details introduced in the student’s summary.
- Critical error: Tutor corrections also prevented students from requesting results before tests had been performed.The prompt clarified that no blood test had occurred and redirected the student to consider ordering a complete blood count first.
S.2 Case materials
Three acute abdominal scenarios were developed using a shared case structure, with appendicitis and pancreatitis serving as learning cases and perforated peptic ulcer serving as the examination case.
- S.2 Case materials: Three acute abdominal disease scenarios—appendicitis, pancreatitis, and perforated peptic ulcer—were constructed for the experiment.Appendicitis and pancreatitis were used during learning, while perforated peptic ulcer was reserved for examination.
S.3 Supplementary assessment and intervention details
The supplementary materials document the study workflow, platform and analysis environment, assessment reliability and scoring anchors, tutor prompts, and alignment between cases and OSCE constructs.
- Supplementary materials: The supplementary tables cover participant workflow, platform architecture, examination outcomes, questionnaire reliability, scoring rubrics, tutor prompts, and case-to-construct alignment.The listed materials include Tables S1–S7 and supplementary case figures for appendicitis, pancreatitis, and perforated peptic ulcer.
- Study scope: The primary analysis used complete cases from volunteer learners, and no biological samples or identifiable patient clinical data were collected.This defines the participant and data-collection scope of the supplementary assessment materials.
- Statistical analysis: Continuous outcomes were analyzed with Mann–Whitney U tests, binary outcomes with chi-square tests, and negative Hedges’ g values indicate higher MA-group scores.The negative effect-size convention is an export convention rather than a change in the underlying score direction.
- Assessment: OSCE-aligned domains used shared 1–5 anchors, with empathy contributing to both checklist items and the global communication-and-empathy domain.The anchors range from clearly insufficient to excellent, and case-specific items map to systematic history taking, clinical thinking, communication and empathy, diagnostic accuracy, and weighted checklist coverage.
- Intervention details: Tutor prompts were Socratic questions or brief hints designed to guide interviewing without directly disclosing the final diagnosis or case answer.The platform automatically recorded dialogue records, timestamps, worksheets, agent outputs, evaluation results, questionnaire responses, and prompt-version snapshots.