Source-linked AI summary
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Cheng Liang, Pengcheng Qiu, Ya Zhang, Yanfeng Wang, Chaoyi Wu, Weidi Xie
TL;DR
Static, single-turn medical benchmarks do not adequately assess the sequential information gathering, intervention, monitoring, and adaptation required in clinical encounters. The paper introduces MedSP1000, an SP-derived interactive benchmark, and finds that GPT-5.5 completes only 60.4% of expert-defined rubric items while medical-specialized models perform worse.
Problem
Static and limited interactive benchmarks do not adequately assess the sequential clinical behaviours required across evolving patient encounters.
Method
MedSP1000 provides a standardized-patient-based interactive evaluation framework for clinical agents.
Results
GPT-5.5 completed 60.4% of rubric items, while the strongest medical-specialized model reached 40.0%.
Takeaways & Limitations
Current LLMs should remain under strict human supervision and be deployed only as assistive tools.
Takeaways & Limitations
MedSP1000 evaluates idealized teaching simulations with LLM-driven patient and environment agents rather than real clinical encounters.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically delivers care across an encounter: gathering information, planning treatment, and adapting longitudinal management across successive patient states. Medical education has long addressed an analogous challenge through standardized patients (SPs): trained actors who consistently portray clinical cases, enabling realistic practice and objective, scripted assessment. Here we introduce MedSP1000, an SP-derived interactive benchmark for clinical-agent evaluation, including 1,638 SP cases with 24,602 trajectory-level peer-reviewed rubrics. MedSP1000 converts peer-reviewed SP teaching cases into executable scenarios with defined SP case scripts, clinical environment contexts, and human-validated structured rubric. In each simulation evaluation run, a clinical agent interacts in closed loop with a patient agent and an environment controller, and its behaviour is scored throughout the encounter against expert criteria specified in the original materials. Applying MedSP1000 to a range of general-purpose and medically specialized LLMs, we find that performance on static benchmarks does not reliably translate to such educational scenarios. The best-performing model, GPT-5.5, completes only 60.4% of expert-defined rubric items, whereas the strongest medically specialized model reaches 40.0%; increasing test-time compute produces no measurable gain. These results suggest that current LLMs, including agentic systems tuned for medicine, are not yet reliable enough to be safely integrated into actual clinical practice. More broadly, MedSP1000 shows how process-level, SP-style evaluation can reveal clinically relevant failure modes that single-turn benchmarks miss.
1 Introduction
Existing LLM evaluations emphasize static question answering and incompletely assess the longitudinal, interactive behaviours required for clinical care. MedSP1000 addresses this gap by adapting standardized-patient cases into closed-loop clinical-agent evaluations scored with expert-derived rubrics.
- Evaluation gap: Static medical benchmarks largely reduce clinical practice to isolated prompts and final answers, limiting assessment of interactive clinical-agent capabilities.Clinical care requires gathering incomplete information, selecting examinations and tests, interpreting evolving results, and initiating management over time.
- Evaluation gap: Recent interactive evaluations often focus on diagnostic history-taking or online consultation, leaving broader behaviours such as interventional decision-making insufficiently assessed.
- MedSP1000: MedSP1000 provides a standardized-patient-based interactive evaluation framework for clinical AI agents.Standardized patients enable realistic encounters, while predefined faculty rubrics support objective, quantitative assessment beyond diagnostic accuracy.
- MedSP1000: 1,638 SP cases and 24,602 trajectory-level rubrics form MedSP1000, spanning 17 specialties and six ACGME core competencies.The cases were curated from peer-reviewed medical education materials in MedEdPORTAL.
- Initial results: 60.4% of expert-defined rubric items were completed by GPT-5.5, while the best medical-specialized model reached 40.0%, trailing by 20.4 percentage points.These results indicate a substantial gap between static medical competence and interactive clinical performance.
2 Results · 2.1 Introduction of MedSP1000
MedSP1000 introduces an executable standardized-patient benchmark that reorganizes educational cases into structured multi-agent simulations scored against expert rubrics. Its construction and runtime execution received strong clinician validation, while performance is summarized using rubric-completion rates across cases, specialties, and competencies.
- 2 Results: The Results section first introduces MedSP1000 and its evaluation protocol, then reports model performance across simulated clinical cases and synthesizes capability patterns.
- 2.1.1 Standardized Patient Construction: MedSP1000 comprises 1,638 executable cases curated from MedEdPORTAL standardized-patient resources containing objectives, setups, role assignments, clinical progression, and grading points.
- 2.1.1 Standardized Patient Construction: The benchmark converts dispersed source materials into role-specific packets for the clinician, patient, environment controller, and evaluator agents.The packets encode scenario context, role information, dynamic events, patient responses, environment behavior, and scoring references for multi-agent simulation.
- 2.1.1 Standardized Patient Construction: Cases span broad clinical specialties, with each case assigned a HealthBench Professional taxonomy specialty and each rubric item mapped to one of six ACGME core competencies.
- 2.1.2 Human Validation of MedSP1000: 4.66, 4.85, 4.80, and 4.81 were the mean scores across four data-construction quality dimensions, while annotators differed by only 0.41 points on average.Clinicians independently reviewed 100 automatically constructed cases using a 5-point scale.
- 2.1.2 Human Validation of MedSP1000: 97.7% of runtime-fidelity ratings fell in the 4-5 range, with exact agreement for 83.3% of paired ratings and differences of no more than 1 point for 96.2%.Twelve clinicians reviewed 100 completed runs, with each run independently scored by two annotators.
- 2.1.3 Evaluation Metric: The primary metric is case-level rubric completion rate, calculated as the fraction of rubric items completed in a run and averaged across cases.Specialty-level rates average per-case rates, whereas competency-level rates aggregate rubric items within each competency.
2.2 Evaluation Results across Core Competencies
MedSP1000 evaluation shows frontier general-purpose models leading overall rubric completion, while performance varies substantially across ACGME competencies. Patient care and professionalism are strongest across models, whereas practice-based learning and improvement is consistently weakest.
- Overall performance: GPT-5.5 leads overall rubric completion at 60.4%, with the four highest-scoring models clustering within a six-point micro-metric band.All four highest-scoring models are general-purpose frontier models.
- Overall performance: Competency-macro scores are uniformly 3–6 points below micro scores because micro weighting is dominated by patient care and interpersonal and communication skills.Macro weights the six competencies equally, including lower-volume, lower-scoring dimensions such as PBLI.
- Competency profile: Across all seven models, patient care and professionalism rank strongest, medical knowledge, systems-based practice, and interpersonal and communication skills occupy the middle range, and PBLI ranks lowest.This ordering is consistent across the six ACGME core competencies.
- Competency profile: PBLI completion does not exceed 30% for any model and falls below 20% for both medical-domain models.PBLI is the lowest-performing competency across every evaluated model.
2.3 Effect of Context Length on Performance · 2.4 Evaluation Results across Specialties
Across most models, longer input contexts improve rubric completion rather than causing degradation, except when simulations approach a model’s context limit. Specialty performance varies substantially, with general-purpose models generally leading and the strongest results concentrated in several higher-volume acute and protocol-driven specialties.
- 2.3 Effect of Context Length on Performance: The context-length analysis relates rubric completion to input length across all models using individual-run points and ten equally sized bins.Each bin contains 10% of runs and is plotted at its mean input length and mean rubric completion.
- 2.3 Effect of Context Length on Performance: Completion rate increases monotonically with input length for most models with ample context windows, with no performance degradation at the upper end.The longest benchmark cases reach only about 40,000 tokens, below current models’ maximum context windows.
- 2.3 Effect of Context Length on Performance: Longer contexts reflect more thorough clinical exploration and evidence gathering, while modern mainstream context windows adequately handle most interactive medical SP scenarios.Current limits are 1M for four general-purpose models, 256K for Qwen-3.5, and 128K for MedGemma.
- 2.3 Effect of Context Length on Performance: Baichuan-M3’s completion rate declines near the upper context range because its 40,960-token maximum closely matches the longest simulated contexts.This pattern is consistent with the longest scenarios approaching or exceeding its capacity.
- 2.4 Evaluation Results across Specialties: Specialty analysis covers 15 specialties with at least five cases, excluding two of the 17 specialties because their case counts are too small for stable estimates.Figure 5 reports each model’s per-specialty macro rubric completion rate with 95% confidence intervals.
- 2.4 Evaluation Results across Specialties: GPT-5.5 ranks first in 10 of 15 specialties, while Gemini-3.1-Pro leads Critical Care at 71.4%, ahead of DeepSeek-V4-Pro at 70.5% and GPT-5.5 at 68.0%.The two medical-tuned models occupy the lowest two positions in 12 of 15 specialties.
- 2.4 Evaluation Results across Specialties: Infectious disease, Dermatology, and ENT & Ophthalmology are exceptions, but each has n ≤17 and wide confidence intervals make within-specialty ordering unreliable.Baichuan-M3 rises into the range of general-purpose models in these three smallest specialties.
- 2.4 Evaluation Results across Specialties: Among specialties with n ≥36, mean completion is highest in Emergency medicine at 63.4% and lowest in Primary care at 47.8%, while full-distribution extremes are 26.8% in ENT & Ophthalmology and 67.9% in Hematology and oncology.Other higher-volume means include Internal medicine at 63.1%, Surgery at 62.4%, Critical Care at 61.0%, Geriatrics at 51.2%, and Obstetrics and gynecology at 51.6%; the extremes have fewer than 20 cases.
2.5 Comparative Analysis of General-Purpose and Medical-Specialized Models · 2.6 Analysis of the Effectiveness of Test-Time Scaling Strategies
General-purpose models substantially outperform medical-specialized models on MedSP1000, while test-time scaling strategies yield no statistically significant overall improvement for GPT-5.5. Best-of-5 sampling and multidisciplinary consultation produce only marginal macro gains, with a consistent improvement limited to interpersonal and communication skills.
- 2.5 Comparative Analysis of General-Purpose and Medical-Specialized Models: 40.0% and 39.5%: Baichuan-M3 and MedGemma’s micro completion rates trail Qwen-3.5 at 51.5% and GPT-5.5 at 60.4%.The medical-specialized models score at least 11 points below the weakest general-purpose model and more than 20 points below GPT-5.5.
- 2.5 Comparative Analysis of General-Purpose and Medical-Specialized Models: Medical-specialized models occupy the bottom two positions in 12 of 15 specialties and trail general-purpose models across all six ACGME competencies.Per-case distributions likewise place general-purpose models, especially GPT-5.5, closer to high completion and the ceiling.
- 2.5 Comparative Analysis of General-Purpose and Medical-Specialized Models: The results suggest that static, single-turn medical benchmarks may encourage domain-specialized models to overfit formats that do not test sequential clinical decision-making.The passage contrasts these models’ limitations with general-purpose models’ broader medical knowledge and long-horizon agentic capabilities.
- 2.6 Analysis of the Effectiveness of Test-Time Scaling Strategies: The test-time scaling experiment evaluates GPT-5.5 on MedSP1000-Verified, the human-validated subset, because the strategies impose substantial computational costs.The unscaled results used direct answer production at each turn.
- 2.6 Analysis of the Effectiveness of Test-Time Scaling Strategies: Best-of-5 sampling independently runs each case five times and uses self-consistency majority voting, whereas MDT coordinates five specialist clinicians with an integrating role.These are the two representative test-time scaling strategies examined.
- 2.6 Analysis of the Effectiveness of Test-Time Scaling Strategies: Neither additional sampling nor multidisciplinary collaboration produces a discernible overall improvement, indicating that more complex prompting does not overcome base-model capability constraints.The reported differences lack statistical significance.
- 2.6 Analysis of the Effectiveness of Test-Time Scaling Strategies: 0.57→0.61: both scaling strategies consistently improve only interpersonal and communication skills, while other ACGME dimensions remain largely flat or lower.The passage attributes some lower-dimensional performance to overconfidence and premature commitment under multi-agent prompting.
2.7 Case Study
Three representative standardized-patient trajectories show that GPT-5.5 can satisfy most rubric items while still exhibiting clinically important omissions in assessment, counseling, and emergency management. The cases illustrate that strong information gathering and partial task completion do not ensure complete or safe care.
- Representative encounters: Three representative SP encounters provide full turn-by-turn trajectories with rubric items marked completed or not.The encounters are presented as concrete demonstrations of the reported patterns.
- Acute stroke: In an acute ischaemic stroke case, GPT-5.5 completes the initial assessment within the mandated window and reaches a correct thrombolysis decision, scoring near ceiling across several domains.The agent establishes last-known-well, obtains finger-stick glucose, orders labs and imaging, and performs a focused NIHSS.
- Prenatal nutrition counseling: In prenatal nutrition counseling, GPT-5.5 gathers detailed quantitative dietary information but scores 3/5 on patient care and 2/7 on interpersonal and communication skills.It never states the recommended fish-intake guideline or addresses how preparation affects intake.
- PICU emergency management: In a PICU case involving a 2-year-old with altered mental status, a 3-to-2 subspecialist majority ends the encounter at the seventh turn before core resuscitation occurs.Three subspecialists judge the child stabilized, overriding emergency-medicine and critical-care agents who note that core resuscitation has not yet occurred.
3 Discussion
MedSP1000 introduces a process-level, interactive benchmark built from standardized-patient cases to evaluate LLMs as sequential clinical decision-makers. Results show that medical specialization and additional test-time compute do not reliably improve interactive performance, while the benchmark remains an idealized simulation rather than a direct proxy for live clinical readiness.
- Evaluation framework: The benchmark uses a closed-loop, multi-agent framework to assess information gathering, state updating, and time-sensitive clinical actions across multiple time steps.The model interacts with a patient agent and environment controller under a standardized state-transition protocol.
- Benchmark contribution: MedSP1000 is the first sequential clinical decision-making evaluation dataset built from educational standardized-patient cases.It repurposes medical-training materials into dynamic clinical scenarios with quantifiable scoring.
- Main findings: Medical-domain models rank below general-purpose LLMs, with Baichuan-M3 at 40.0% and MedGemma at 39.5% micro completion rates.Baichuan-M3 falls 11.5 points below Qwen-3.5 at 51.5% and 20.4 points below GPT-5.5 at 60.4%.
- Main findings: Additional test-time compute does not materially improve GPT-5.5’s interactive performance: the single-pass baseline scores 67.1%, Best-of-5 scores 67.8%, and five-specialist consultation scores 68.0%.The three approaches differ by at most 0.9 points and are statistically indistinguishable.
- Limitations: MedSP1000 provides reproducible, expert-defined evaluation but uses idealized teaching simulations and LLM-driven patient and environment agents, limiting direct inference about live clinical deployment readiness.The authors therefore identify a gap between benchmark performance and readiness for real clinical encounters.
4 Methods
MedSP1000 transforms MedEdPORTAL standardized-patient teaching materials into executable, role-specific clinical simulations with structured scoring rubrics. Its framework uses controlled information access and interacting agents to evaluate clinical models across dynamically progressing encounters.
- Role-specific packets: Each retained case is partitioned into four role-specific packets: scenario initialization, patient script, environment-controller materials, and scoring rubric.The packets define the information accessible to the assessed model, patient agent, environment controller, and evaluator.
- Role-specific packets: Strict visibility rules prevent information leakage by restricting the assessed model to encounter-appropriate information and withholding privileged or future disease information.The assessed model can progressively access content dynamically revealed during interaction but cannot non-causally access future progression.
- Automated processing: Codex performs metadata extraction, role-packet construction, and leakage and consistency auditing before rubric parsing standardizes educators’ free-text scoring points into discrete items.The auditing phase checks evaluator-item leakage, patient-packet sufficiency, and encounter-splitting consistency.
- Data curation: 1,017 source articles across 17 specialties pass text-completeness filtering, including 931 retained directly and 86 retained after manual review.Articles are unified into Markdown using format-specific conversion and OCR-based extraction when needed.
- Multi-agent evaluation: The evaluation framework combines a patient agent for multi-turn standardized-patient dialogue, an environment controller for resolving actions and scenario progression, and an evaluator that judges trajectories against rubrics.The model under evaluation interacts with the environment through a structured interface.
Data availability
MedSP1000 is derived from an open-access, peer-reviewed repository whose source materials are Creative Commons licensed and were processed under their original terms for non-commercial research.
- Data availability: MedSP1000 was constructed from 1,638 source standardized-patient and simulation articles in MedEdPORTAL, retaining original author attribution and following their CC BY, CC BY-NC, or CC0 licences.The materials were accessed and processed solely under their original licence terms for non-commercial research.
5 Supplementary
The supplementary materials characterize attachment diversity in the MedEdPORTAL source corpus at both coarse modality and fine-grained file-type levels.
- Attachment distributions: Supplementary Figure 1 groups corpus attachments into text, image, audio, video, and programmatic-resource categories.This provides a coarse-grained modality-level distribution of attachments.
- Attachment distributions: Supplementary Figure 2 reports attachment distributions by original file extension and source file type.This provides a fine-grained view of file-level diversity in the source corpus.
5.1 Full Results on the Human-Annotated Subset
On the human-annotated subset of 100 cases, results preserve the full-benchmark model ordering and competency profile while yielding higher absolute completion rates. GPT-5.5 leads the best medical model by 15.4 points, and PBLI remains every model’s weakest competency.
- Model performance: GPT-5.5 remains strongest on both micro and macro metrics, with general-purpose models ahead of the two medical-specialized models.The subset results closely track the full benchmark, preserving the relative ordering of systems.
- Model performance: 64.4% micro for GPT-5.5 leads MedGemma at 49.0% micro by 15.4 points.The general–medical performance gap persists at a similar magnitude on the human-annotated subset.
- Competency profile: PBLI is the weakest dimension for every model, not exceeding 16.7% in any system.Patient care and professionalism remain among the strongest competencies.
- Subset comparison: Absolute completion rates are higher across all models than on the full benchmark.The subset was filtered for construction quality and spans a narrower, more cleanly specified set of scenarios.
- Subset comparison: Consistent model ordering and competency structure indicate that full-benchmark conclusions are not driven by construction noise.The human-annotated subset was constructed and verified by twelve clinicians and is also used for the test-time scaling experiment.
5.2 Case Study
The case studies show that GPT-5.5 can reproduce broad clinical workflows while still missing fine-grained protocol and counseling requirements. MedAgents’ multi-agent deliberation instead terminates prematurely, leaving basic resuscitation and communication tasks unreached.
- Successful encounter: GPT-5.5 completes a largely successful nine-turn stroke encounter, compressing initial assessment and thrombolysis decisions into the mandated window.It establishes presence, gathers onset and last-known-well, checks glucose, screens risks, orders tests, performs NIHSS, and decides on thrombolysis at T7.
- Successful encounter: GPT-5.5 scores near competency ceilings in the stroke case: PC 13/14, MK 4/4, and SBP 2/2.The two missed items involve a 20 mg rather than guideline-mandated initial 10 mg labetalol dose and absent explicit risk–benefit–alternative consent for alteplase.
- Failing encounter: GPT-5.5 scores only 3/5 on Patient Care and 2/7 on ICS in a failing prenatal nutrition encounter despite thorough quantitative dietary data collection.It omits recommended fish-intake guidance, preparation-related contaminant counseling, cardiovascular-benefit evidence, and answers to the patient’s two actionable questions.
- Multi-agent encounter: MedAgents terminates at T7 after three of five subspecialists vote that the child is already stabilized, despite dissent from Emergency Medicine and Critical Care.The majority decision occurs during an initial ED presentation and prevents completion of broader resuscitation and monitoring.
- Multi-agent encounter: The MedAgents aggregation leaves basic Patient Care and ICS items unreached, including a 20 cc/kg saline bolus, glucose, blood gas, lactate, naloxone, nurse listening, and team introduction.The case indicates that multi-agent deliberation and majority-vote aggregation can be counterproductive rather than consistently beneficial.
5.3 Material Processing Prompts
Phase 1 evaluates whether source material contains a text-simulatable clinical or communication encounter, using a scenario-centered gate rather than role-by-role checks. A valid case requires a clear task, portrayable patient-side counterpart, and sufficient scene information to launch and advance the interaction.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: Phase 1 judges the whole case rather than pre-disqualifying it by checking four downstream roles individually.The gate asks whether the material already provides the minimum skeleton for a text-enactable patient encounter or communication scenario.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: A clear interactive encounter task is the first required scenario element.Examples include triage, history taking, breaking bad news, informed consent, discharge instructions, family communication, crisis management, and teaching mini-encounters.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: A portrayable patient, family member, or other counterpart is the second required element.A full character script is preferred, but sufficiently clear role positioning, background, presenting situation, emotions, or response boundaries can also support stable portrayal.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: Sufficient scene information to open and advance the interaction is the third required element.Relevant information may include location, relationships, current events, starting state, observable data, workflow notes, SP responses, examination reports, or branching cues.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: When all three elements are present, the case is normally marked simulatable=true.Formal rubrics, pre-split role materials, and fully quantified patient data are not required if role boundaries and opening information are adequate.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: A case is marked simulatable=false only when no encounter can be extracted, no interactive counterpart can be portrayed, or the scene cannot be launched.Pure lectures, abstract teaching content, course schedules, and discussion outlines do not qualify as runnable encounters.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: Each scenario represents one independently runnable clinical or communication encounter, and separate entries require explicit distinguishing anchors in the source.Variants may differ by counterpart, setting, age or identity, task, or other explicitly stated information.
- 5.3.1 Phase 1: Material Triage and Scenario Extraction: Communication-training material normally yields scenarios=1 when it has a clear role, task, and opening setup, while simulatable=false normally yields scenarios=[].Missing formal assessment tools or isolated details alone should not force an empty scenario array, and unspecified fields must remain null rather than being guessed.
5.4 Rubric Extraction
Rubric extraction freezes scenario-specific scoring items before model evaluation, using only evaluator materials and assigning each item to exactly one ACGME core competency. Items must remain observable, independently judgeable, semantically faithful, and verbatim where extracted.
- Frozen rubric source: The rubric is frozen before evaluation and reused unchanged across examinee models, independent of model behavior or transcripts.A Codex agent extracts the rubric from evaluator materials for the evaluator agent to consume during scoring.
- Frozen rubric source: Scoring items are derived exclusively from readable files in the scenario’s evaluator/ directory, not other role directories or pipeline products.The extraction process reads evaluator/ recursively while excluding examinee/, sp_actor/, environment_controller/, and unrelated generated or temporary files.
- Item eligibility: Each scoring item must be a decidable, independently judgeable statement about examinee behavior or judgment with matching semantics in the evaluator materials.Narrative facts, structural labels, and overarching learning objectives are excluded when concrete scoring-level descriptions are available.
- Competency assignment: Each item is assigned exactly once to the ACGME dimensions PC, MK, SBP, ICS, PBLI, or PROF according to its overall semantic competency.Ambiguous items go to the nearest dimension, with no catch-all category or duplicate categorization.
- Output format: The output is strict JSON containing only the case metadata, rubric version, and six competency arrays of unique verbatim scoring-item strings.Items receive no completion flags or nested metadata, and empty dimensions remain [].
5.5 Simulation Agent Prompts
The environment controller must preserve a strict boundary between objective in-world state and examinee-driven clinical decisions. It must also maintain reference-grounded healthcare-team presence and role-consistent speech while handling progression and unavailable information from the on-site perspective.
- 5.5.1 Environment Controller Agent: In-world text must sound like records, monitors, or on-site observers, using declarative or completed wording without modal, imperative, or future-intention language.Reference-material concepts and meta-language about missing information are prohibited in these fields.
- 5.5.1 Environment Controller Agent: In-world fields record only the patient’s current objective state and completed events, never recommendations about what the examinee should do next.Injecting next-action guidance would contaminate downstream evaluation and reduce scoring discrimination.
- 5.5.1 Environment Controller Agent: Scenario progression distinguishes objective changes driven by the patient, disease, or time from events that require an examinee-triggered decision.Only the former should be proactively written into in-world fields; the latter must remain for the examinee to reason about or request.
- 5.5.1 Environment Controller Agent: Unavailable information is handled according to cause: clinical waiting or no change appears naturally in in-world fields, whereas out-of-scope actions are reported only as unsupported with a rationale.The latter must not be reflected in feedback, events, or patient_status.
- 5.5.1 Environment Controller Agent: actors_present must list the full current set of reference-grounded healthcare-team roles, including roles carried over from prior turns.Patient-side roles and the examinee are excluded, and roles cannot be invented from clinical plausibility.
- 5.5.1 Environment Controller Agent: Healthcare-team roles may be added only when required by a plot node, explicitly called by the examinee, or activated by a reference-specified trigger.Opening-turn roles explicitly described as present are included immediately rather than waiting for them to speak.
- 5.5.1 Environment Controller Agent: Feedback may contain objective statements or role-prefixed speech, but speech is limited to present healthcare-team members whose names already appear in actors_present.Patient, family, companion, and guardian dialogue belongs to the SP output, not the environment controller.