Source-linked AI summary
SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training
Zihan Wang, Anita Marie Slominska, Rennie Bimman, Elizabeth Di Flumeri, Amanda Mayappo-Neeposh, Conall Francoeur, Tamara Ellen Carver, Xiao-Wen Chang, Doina Precup, Esin Darici Haritaoglu, Ismail Haritaoglu, Akshatha Arodi, Naomi Goloff
TL;DR
Pediatric SIC demands scalable training that captures emotional, relational, multi-party, and feedback-dependent behavior, but existing simulators mainly optimize generic dialogue quality. The paper introduces clinician-anchored turn- and dialogue-level benchmarks plus SIC-Agents, a curriculum-guided self-improving parent simulator with an inspectable skill document. Its central finding is that curriculum guidance is necessary: unguided self-improvement can improve surface dialogue quality while degrading clinically critical contingent behaviors.
Problem
Existing pediatric SIC training is difficult to scale, while dialogue simulators largely optimize generic quality rather than curriculum-contingent emotional and relational behavior.
Method
The paper develops PitfallBench and DialogueBench and a curriculum-conditioned self-improving parent simulator that revises a clinician-editable skill document using generated practice and judge feedback.
Results
Curriculum guidance is necessary because unguided self-improvement can improve surface dialogue quality while degrading clinically critical contingent behaviors.
Takeaways & Limitations
Pediatric SIC provides a testbed for emotionally grounded, curriculum-conditioned dialogue modeling in high-stakes, data-scarce settings.
Takeaways & Limitations
Whether the benchmarks and behaviors transfer across languages, cultures, and healthcare systems remains untested, and the simulator is built exclusively on Claude models.
Abstract
from arXiv · showhide
Pediatric serious illness communication (SIC) is critically important, yet scalable communication training for clinicians remains limited. Compared with other dialogue simulation settings, pediatric SIC poses additional challenges, including multi-party interactions, response to parental distress and strong dependence on feedback dynamics. Existing LLM-based simulators optimize generic dialogue quality rather than curriculum-contingent behavior required for effective SIC training. In collaboration with educators and pediatric clinicians, we introduce the first benchmark suite and simulation framework tailored to pediatric SIC training. Our benchmarks, PitfallBench and DialogueBench, evaluate simulators both at the turn-level and across full dialogues. We further propose SIC-Agents, a self-improving framework that generates a clinician-editable skill document to guide simulator behavior. Our experiments show that SIC-Agents outperforms static expert prompting. To support future research, we release our benchmarks for parent simulation in pediatric SIC at https://github.com/Beikewzh/sic-benchmarks
1 Introduction
Pediatric serious illness communication requires emotionally responsive, adaptive interaction beyond medical information exchange, yet clinicians report feeling underprepared and existing training is difficult to scale. The paper introduces clinician-anchored benchmarks and SIC-Agents, a curriculum-guided, self-improving simulator with an editable skill document.
- Motivation: Clinicians report feeling underprepared despite the critical importance of pediatric serious illness communication skills.
- Motivation: Existing training relies largely on resource-intensive standardized-patient simulations, limiting scalable repeated practice.LLM simulators offer an on-demand alternative, but prior systems generally emphasize generic dialogue quality rather than curriculum-contingent behavior.
- Motivation: Pediatric SIC requires clinicians to address emotional impact, meaning, trust, and engagement alongside medical information.The paper describes communication as feedback-dependent, with clinician language shaping parental emotion and subsequent behavior.
- Contributions: PITFALLBENCH and DIALOGUEBENCH evaluate pediatric SIC simulators at turn level and across full dialogues.They compare supportive and suboptimal clinician strategies in matched scenarios and assess emotional responsiveness, relational adaptation, and conversational grounding.
- Contributions: SIC-Agents self-improves a clinician-editable skill document while using curriculum guidance to reduce medically and behaviorally inappropriate parent responses.The framework is designed for auditing, correction, and curriculum control by clinicians and instructors.
- Contributions: The framework integrates the simulator into a rapid-cycle deliberate-practice web interface for scalable evaluation and clinical education.
2 Background and Related Work
The paper frames pediatric SIC simulation as requiring both localized adaptation to clinician communication and coherent grounding across extended, multi-turn encounters. It develops clinician-informed behavioral benchmarks and embeds pediatric SIC curriculum feedback in an inspectable self-improvement process.
- Learning Opportunities: PitfallBench isolates trigger turns where learners commit clinically meaningful communication mistakes, enabling targeted formative feedback.The taxonomy contains ten common learner pitfalls across sharing and supporting medical information and exploring goals and values.
- Conversational Flow: Dialogue-level evaluation measures whether an AI parent simulator maintains grounded, topically consistent, human-like interaction across an entire encounter.This complements trigger-turn evaluation because locally appropriate responses can still produce a disjointed dialogue.
- Background and Related Work: Pediatric SIC training requires repeated feedback-driven interaction, but one-off standardized-patient workshops cannot provide scalable practice.
- Background and Related Work: Prior medical dialogue simulators mainly target generic realism, while pediatric SIC requires contingent adaptation to emotional and relational dynamics.
- Background and Related Work: The benchmarks evaluate both trigger-turn contingent adaptation and grounding across extended dialogue in pediatric SIC.The design responds to privacy constraints that limit large public pediatric SIC corpora.
- Background and Related Work: SIC-Agents embeds pediatric SIC curriculum feedback in a clinician-editable form rather than relying only on standard self-improvement components.
3 Datasets
The paper introduces clinician-anchored resources for evaluating pediatric SIC parent simulators at trigger turns and across full dialogues. PitfallBench tests contingent reactions to communication quality, while DialogueBench tests grounding and topic consistency during scripted encounters.
- Datasets: PitfallBench contains 1,000 turn-level items, while DialogueBench contains 152 gold dialogues for full-encounter evaluation.The resources are accompanied by an 8,391-dialogue silver corpus.
- PitfallBench: PitfallBench pairs a curriculum-defined pitfall move with its skilled alternative under the same case and dialogue history.The simulator generates the next parent utterance, which is classified as challenge_signal, reward_signal, or neither.
- PitfallBench: PitfallBench covers ten learner pitfalls organized across sharing and supporting medical information and exploring goals, values, and alignment.Each pitfall is linked to a curriculum rule specifying the dialogue state, clinician moves, and expected parent reaction.
- PitfallBench: PitfallBench validates gold labels through an explicit acceptance criterion iteratively calibrated between pediatric SIC clinicians and an independent LLM reviewer.Iteration ends when reviewer–clinician agreement reaches 95% on a stratified validation sample of 50 items.
- DialogueBench: DialogueBench evaluates parent responses across full multi-turn encounters with clinician turns held fixed and scores Conversational Grounding and Topic Consistency.The two-axis 1–3 rubric is judged by two independent LLM judges calibrated against annotations from 11 pediatric clinicians.
- Datasets: SIC-SimCorpus provides a synthetic substrate varying eight generator families, six pediatric disease categories, three trainee proficiency levels, and 12–28-turn dialogues.The release separates the full 8,391-dialogue silver corpus from the 152-dialogue human-reviewed gold benchmark.
4 SIC-Agents
SIC-Agents is a curriculum-conditioned self-improving parent simulator that adapts an editable natural-language skill document rather than latent model parameters. It combines judge-based interaction feedback with curriculum-level coaching to refine behavior for pediatric SIC practice.
- Framework objectives: SIC-Agents generates training evidence without private clinical dialogue data while preserving interpretability through an editable skill document.The document represents adaptive parent behavior in natural language rather than latent model parameters.
- Agents: A fixed clinician agent follows the pediatric SIC curriculum to provide a stable learner-side interlocutor and surface relevant clinician moves.This supports coherent practice turns and self-improvement runs.
- Agents: The parent agent conditions on the case, dialogue history, and a mutable skill document encoding expression patterns, exemplars, transitions, and behavioral constraints.Constraints include length, role fidelity, emotional calibration, and off-topic generation.
- Self-improvement cycle: Each turn samples K=3 candidate responses, filters invalid generations, scores candidates on Conversational Grounding and Topic Consistency, and retains selected and rejected responses as contrastive evidence.The highest-scoring candidate is inserted into the dialogue.
- Self-improvement cycle: The Skill Editor uses batch evidence to make localized skill-document changes, accumulating interaction-derived rules and exemplars across simulation cycles.Adaptation changes the interpretable document rather than the underlying LLM parameters.
- Feedback: Self-improvement combines judge-derived interaction feedback with Coach-derived curriculum feedback about pitfalls, skills, and appropriate parent reactions.The judge assesses grounding, topic maintenance, response length, and role fidelity, while the Coach supplies pedagogical context.
- Training infrastructure: The resulting skill document powers a training platform where learners practice multi-turn interactions with curriculum-grounded feedback and faculty can route challenging cases for refinement.The framework is positioned for confidential, data-scarce, feedback-intensive domains.
5 Experiments
The experiments compare static prompts with self-improvement variants to test whether curriculum guidance preserves clinically contingent behavior while retaining dialogue quality. The full system achieves the strongest primary PitfallBench outcome and highest composite, whereas generic feedback alone permits clinical drift.
- Evaluation: PitfallBench accuracy tests curriculum-contingent reactions at matched trigger turns, while DialogueBench measures grounding and topic consistency across full dialogues.PitfallBench accuracy is the primary outcome; a composite averages it with rescaled DialogueBench Grounding.
- Experimental settings: The study compares five deployable conditions, including static prompts and two five-cycle self-improvement systems sharing the same minimal seed.The no-curriculum ablation retains candidate sampling, judges, and the Skill Editor but removes the Coach signal.
- Main results: 0.78 ± 0.03 PitfallBench and 2.96 ± 0.02 Grounding are achieved by the full self-improvement condition across three independent seeds.The full condition is strongest on the primary PitfallBench outcome and yields the highest composite score.
- Main results: The expert-prompt baseline reaches PitfallBench 0.85 versus 0.75 for the seed but has the weakest grounding score, while the full-taxonomy prompt remains below the full system on all three metrics.Grounding is 2.32 for the full-taxonomy prompt versus 2.96 for the full system.
- Curriculum ablation: Generic judge feedback without curriculum guidance keeps Grounding high but lowers PitfallBench accuracy below the initial seed, especially for skilled-reward responses.The curriculum ablation uses the same seed, candidate count, judges, and five edit cycles; only the Coach’s taxonomy and expected reactions are removed.
- Mechanism: The full system reduces drift by converting judge-identified failures and curriculum interpretations into editable rules, constraints, and exemplars in the skill document.The Coach identifies missing SIC skills such as permission or values elicitation and specifies corresponding parent responses.
- Caveats: The trajectory analysis uses a smaller evaluation subset, and pitfall-resistance remains easier for the full system to learn than skilled-reward.Interpretation therefore emphasizes the across-cycle pattern rather than cycle-to-cycle changes.
- Robustness: The condition ordering full system > expert prompt > no-curriculum holds under GPT-5, DeepSeek R1, and Qwen3 judges.The authors report small score differences across the three non-Claude judges.
6 Conclusion
The paper introduces clinician-anchored pediatric SIC benchmarks and SIC-Agents, a self-improving parent-simulation framework. Its central conclusion is that generic dialogue optimization alone can improve surface quality while degrading clinically critical contingent behavior, whereas curriculum-guided evidence distillation supports inspectable skill updates.
- Conclusion: PITFALLBENCH and DIALOGUEBENCH evaluate pediatric SIC parent simulation at turn level and across full dialogues, respectively.The resources are clinician-anchored and designed for pediatric SIC communication behaviors.
- Conclusion: SIC-Agents generates practice evidence, interprets failures with expert curriculum labels, and writes editable results into a clinician-inspectable skill document.This design targets data-scarce settings where supervised fine-tuning is infeasible and hand-crafted prompts are brittle.
- Conclusion: Without curriculum guidance, self-improvement improves surface dialogue quality while degrading clinically critical contingent behaviors.The conclusion frames curriculum-aware feedback as necessary for preserving clinically grounded simulation.
Limitations
The study’s validation and generalizability are limited by criterion-level expert review, untested transfer across settings, and reliance on a single model family.
- Expert validation used stratified samples and criterion-level review rather than item-by-item annotation of all benchmark examples.
- Whether the benchmark behaviors transfer across languages, cultures, and healthcare systems remains untested.
- The simulator and Skill Editor use only Claude models, leaving possible model-family-specific stylistic and distributional biases.
- Replicating the self-improvement loop across model families is identified as future work.
Ethical Considerations
The work is intended for clinician training and research rather than medical decision-making or direct patient interaction, using synthetic, de-identified cases and clinician ratings. The authors emphasize continued safety, bias, generalization, and scope limits requiring human oversight and supervised training.
- SIC-Agents is intended for clinician training and research, not medical decision-making or direct patient interaction.
- The study uses synthetic, de-identified cases and clinician ratings, with no patient or trainee involvement.
- LLMs may still produce unsafe, biased, or emotionally inappropriate responses and may not generalize equally across populations.
- Simulated parents cannot fully capture the diversity of real families, so the simulator complements rather than replaces supervised clinical training.
C Benchmark Diversity
PitfallBench achieves diversity through 100 pediatric cases spanning twelve care settings, while DialogueBench uses an 8,391-dialogue factorized corpus varying generators, trainee proficiency, and dialogue length. The benchmarks also use structured turn-level evaluation and clinician-linked validation.
- PitfallBench: PitfallBench contains 1,000 items from 100 pediatric cases across twelve care settings, ages 0–18, and varied family configurations.
- PitfallBench: UMAP shows PitfallBench items spread broadly across embedding space rather than collapsing onto a few templates.
- DialogueBench: 8,391 DialogueBench dialogues cross eight LLM generators, three trainee proficiency levels, and 12–28-turn conversations averaging 20 turns.
- DialogueBench: DialogueBench responses spread broadly in embedding space, reflecting diversity in the released dialogue corpus.
- Evaluation: Each parent turn is scored for Grounding and Topic Consistency on 1–3 scales, with hard failures for unsupported facts, excessive length, or ignored questions.
- Evaluation: The rubric judge was calibrated against 119 expert ratings from 11 pediatric clinicians on 28 source dialogues.
F SIC-Agents Implementation Details
SIC-Agents uses clinician-readable skill documents, filtered candidate generation, judge-based selection, and iterative Skill Editor updates. The implementation supports auditable curriculum-guided adaptation and deliberate-practice deployment, while example edits improve response behavior but can cause document bloat.
- Knowledge layer: The parent agent uses skill documents encoding six expression patterns, exemplars, transitions, and constraints on length, role fidelity, emotion, and off-topic generation.
- Per-turn simulation: At each parent turn, the system samples K=3 candidates, hard-filters violations, scores survivors on Grounding and Topic Consistency, and emits the highest-scoring response.
- Skill Editor: The Skill Editor reads selected and rejected candidates with judge-based evidence, then proposes one localized edit to the weakest response pattern.
- Skill Editor: Edits are auditable natural-language diffs using ADD_EXAMPLE, ADD_PATTERN, TIGHTEN_CONSTRAINT, or REVISE_DESCRIPTION.
- Coach signal: The Coach supplies the Skill Editor with the curriculum pitfall map, core communication skill, and expected parent reaction during full self-improvement.
- Example edits: Example edits tightened length, role-fidelity, and grounding constraints and added exemplars; the document grew from 164 bytes to roughly 7 KB over five cycles.
- Limitation: The editor appends more readily than it consolidates, leaving near-duplicate blocks and deferring full MERGE/DEDUP functionality.
H Experiment Configuration
The experiments compare static prompts with self-improvement conditions under controlled candidate sampling, judging, and curriculum-signal settings. Results are reported for PitfallBench and DialogueBench using their respective skill, resistance, grounding, topic, and composite metrics.
- Models and loop: The parent agent is Claude Sonnet 4.6, judges are Claude Haiku 4.5, and the Skill Editor is Claude Opus.
- Models and loop: Each condition runs five cycles, practicing on ten dialogues per cycle and sampling K=3 candidate utterances per parent turn.
- Baselines: The expert-prompt baseline uses a 270-word expert-written prompt with SIC principles but no taxonomy, candidate sampling, judge feedback, or Skill Editor updates.
- Baselines: The full-taxonomy prompt adds the P1–P10 taxonomy, core skill labels, and expected reactions without running the self-improvement loop.
- Metrics: PitfallBench reports skill, pitfall-resistance, and mean scores on 0–1, while DialogueBench reports Grounding and Topic on 1–3 and a 0–1 composite.
- Ablations: The curriculum-signal ablation compares identical self-improvement pipelines while removing the Coach signal from the no-curriculum condition.
- Judging: PitfallBench uses a gated classifier judge, whereas DialogueBench uses a judge calibrated against clinician ratings.
I Detailed Results
The detailed results examine condition-, pitfall-, and judge-level behavior, while also delimiting the benchmark’s curriculum coverage and reporting deployment and release details. The analyses show headroom for oracle access, broad full-system advantages, judge-robust ordering, and a supervised training interface with public benchmark resources.
- Condition-level results: The oracle condition reveals the curriculum signal at inference, effectively sees the expected reaction, and is therefore non-deployable.Its purpose is to quantify remaining headroom on the hardest skilled-response cases.
- Per-pitfall results: The full system leads on almost every per-pitfall cell, while no-curriculum collapses on skilled-reward for pediatric-specific pitfalls P8 and P10.Table 6 reports cycle-5 accuracy separately for pitfall-resistance and skilled-reward.
- Judge robustness: The ordering full system > expert prompt > no-curriculum is preserved under GPT-5, DeepSeek R1, and Qwen3 judges.The non-Claude judges produce only small absolute score differences, while the parent agent and Skill Editor remain Claude-only.
- Curriculum coverage: PitfallBench primarily covers information sharing, family-perspective understanding, and agreement, but excludes opening, relationship-building, and closure.It should therefore be read as covering the information-sharing and goal-alignment core of a pediatric SIC encounter.
- Training and release: The skill document also powers a supervised browser-based practice interface, and the benchmarks, corpus, metadata, prompts, and evaluation drivers are publicly released.The trainee-facing interface is available institutionally to educators on request, while the released resources are licensed CC BY-NC-SA 4.0.