Source-linked AI summary

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Ruoyu Wu, Shenfu Xie, Yinqian Sun, Haibo Tong, Feifei Zhao

arXiv:2608.23397v1cs.AI

TL;DR

Interactive clinical agents need to gather decisive evidence and convert it into safe, grounded actions under partial observability. MediSkill-Evo addresses this by evolving governed process knowledge in four typed banks and using a process-constrained harness to select actions. On fixed suites, complete-system evaluations report improved diagnosis, treatment-intent, and evidence coverage alongside fewer automatically scored critical failures, while remaining descriptive rather than causal or clinically validating.

  • Problem

    Under partial observability, a correct diagnosis alone does not establish that an agent gathered valid evidence or respected care-process constraints.

  • Method

    MediSkill-Evo evolves four typed banks under provenance, support, replay, and safety checks, then ranks actions with a process-constrained harness and Clinical Process Critic.

  • Results

    Complete-system comparisons report higher treatment-intent and evidence coverage and fewer automatically scored safety-related failures across fixed suites and backbone endpoints.

  • Takeaways & Limitations

    The results provide descriptive end-to-end evidence that governed process knowledge can support evidence-grounded clinical interaction within the evaluated benchmark settings.

  • Takeaways & Limitations

    A failure-boundary case shows that governance checks cannot determine whether an image interpretation is clinically true, motivating independent image adjudication.

Abstract

from arXiv · show

Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final diagnosis alone does not show that an agent respected evidence and care-process constraints. We introduce MediSkill-Evo, a clinical agent that evolves governed process knowledge without backbone fine-tuning. It separates experience into four typed banks for clinical skills, process rules, symbolic schemas, and measurement procedures. Provenance, support, replay, and controller-defined safety checks govern publication to a frozen test-time snapshot. A Process-Constrained Preference Harness binds evidence to its source, rejects controller-invalid candidates, and ranks actions with a safety-prioritized Clinical Process Critic. We evaluate complete agent systems across two backbone endpoints and six controlled stress dimensions under the same Doctor-turn limit. On 300 held-out Qwen encounters, MediSkill-Evo improves diagnosis accuracy from 61.33 percent to 69.00 percent and treatment-intent coverage from 33.62 percent to 66.44 percent, while reducing automatically scored critical failures from 31.00 percent to 16.33 percent relative to AgentClinic. On 180 hard-isolation conditions derived from 30 cases, target recovery reaches 93.61 percent under patient-behavior pressure, 100.00 percent for temporal evidence, and 92.22 percent for triage red flags. An exploratory 100-case MedSAM comparison evaluates request-gated tool-interface feasibility. These results provide descriptive end-to-end evidence for the complete system on fixed evaluation suites, not causal evidence for an individual bank or clinical validation of the automatic judge.

1 Introduction

MediSkill-Evo addresses the challenge of improving clinical agents under partial observability while preserving evidence boundaries and care-process obligations. It combines governed four-bank self-evolution with process-constrained action selection and reports complete-system gains on fixed clinical interaction benchmarks.

  • Motivation: Clinical agents must gather decisive history and examinations, update assessments, and recommend treatment without treating unavailable tests as negative findings.A correct diagnosis can still conceal an unsupported or unsafe interaction trajectory.
  • Scope: The system turns evolving external knowledge into controller-valid, evidence-grounded actions without claiming independent clinical legality or safety.Its harness binds results to valid requests, preserves unavailable evidence as unknown, and applies benchmark-defined safety checks without hidden-diagnosis access.
  • Approach: MediSkill-Evo evolves clinical experience through four typed banks for strategies, workflow rules, evidence semantics, and visual measurement procedures.Provenance, replay, support, and safety checks govern publication into the next frozen snapshot.
  • Approach: The Process-Constrained Preference Harness retrieves state-relevant knowledge, rejects controller-invalid actions, and selects candidates with a safety-prioritized Clinical Process Critic.Hard-invalid candidates cannot return through finite penalties; persistent safety-threshold failure leads to safe termination and human escalation.

3 Experiments

The experiments evaluate MediSkill-Evo as a complete system across fixed end-to-end, stress, multimodal, and component-comparison suites under controlled interaction settings. Results show stronger process completion on several dimensions, but also reveal bounded performance and prevent single-component causal conclusions.

  • Experimental setup: The evaluation uses two hosted backbone endpoints with fixed cases, environments, evaluation code, and Doctor-turn ceilings; multimodal conditions additionally share cases and interaction settings but not internal compute.The study reports fixed-suite point estimates and explicitly makes no computational-efficiency claim.
  • Cross-backbone FullChain results: 69.00% diagnosis accuracy and 66.44% treatment-intent coverage are achieved by MediSkill-Evo on 300 held-out Qwen3.6-Flash encounters, versus 61.33% and 33.62% for AgentClinic.Evidence recall rises by 36.12 points, required-history recall by 86.74 points, and critical failures fall from 31.00% to 16.33%.
  • Cross-backbone FullChain results: 47.30% treatment-intent coverage, 40.11% evidence recall, and 73.65% history recall are reported on DeepSeek-V4-Flash, with lower safety violations, critical failures, and unnecessary examinations than AgentClinic.The authors characterize this as consistency within the evaluated endpoint families, not model-family-wide generalization.
  • Process Correctness and Safety under Controlled Stress: 83.79% stress-process score, 88.30% required-action recall, 82.40% treatment-intent coverage, and 80.03% core score are reported in auxiliary stress results, while diagnosis accuracy is 93.89%.These results support stronger process completion in selected dimensions rather than uniform dominance in diagnosis or safety.
  • Modular Visual Measurement on NEJM Cases: The exploratory MedSAM-enabled condition is 3.24 points higher in core-case score than the original-image condition on 100 NEJM cases.The authors describe this as tool-interface feasibility evidence, not a MedSAM improvement or segmentation-only causal estimate.
  • Component Analysis: Component diagnostics show no profile dominates every endpoint, and one frozen run per profile does not establish that any bank is necessary or causally beneficial.Observed shifts support a tradeoff interpretation among banks rather than a single-component explanation of the end-to-end result.

4 Conclusion

MediSkill-Evo frames clinical-agent evolution as governed process knowledge, combining typed update paths with stress testing of evidence-grounded interaction. The reported results remain bounded system-level evidence rather than clinical validation or causal attribution to individual components.

  • 4 Conclusion: Four typed banks provide distinct update and validation paths for reusable strategies, workflow rules, evidence semantics, and visual procedures.The Process-Constrained Preference Harness evaluates these elements jointly during interaction.
  • 4 Conclusion: Stress testing shows stronger target recovery under patient-behavior, temporal, and triage pressure, alongside weaker diagnostic-discriminator and unavailable-test acquisition.
  • 4 Conclusion: The study does not establish clinical safety, population-level generalization, judge construct validity, or causal credit for individual components.
  • 4 Conclusion: The evaluation uses deidentified MIMIC-IV-derived records and published NEJM image cases under their original access, licensing, and redistribution terms.Restricted records and images are not released through the benchmark artifact.

S1.1 Clinical Skill Bank

The supplementary design separates clinical skills, workflow constraints, evidence semantics, and visual procedures into typed banks with explicit provenance, safety, and lifecycle controls. Post-episode proposals are restricted to reusable, controller-valid updates rather than case answers.

  • S1.1 Clinical Skill Bank: The Clinical Skill Bank stores reusable case strategies with applicability conditions, evidence-acquisition or management sequences, and misuse warnings.Symbolic preconditions and semantic gating remove skills that conflict with visible evidence or diagnostic boundaries.
  • S1.1 Clinical Skill Bank: Process Rules encode cross-disease workflow constraints, including triggers, required or prohibited actions, release conditions, and priorities.Rules constrain registered state without creating clinical facts or overriding evidence semantics.
  • S1.1 Clinical Skill Bank: Symbolic Schemas define legitimate observation sources, permitted state transitions, request–result relations, and provenance-bearing fact consumers.Missing, pending, and unavailable states remain distinct and cannot default to normal or negative.
  • S1.1 Clinical Skill Bank: Measurement procedures specify modality, targets, prerequisites, steps, report fields, quality checks, and failure modes while separating procedures from case findings.
  • S1.1 Clinical Skill Bank: Reflection runs only after training encounters and may propose at most one Process Rule and one Symbolic Schema mutation using existing active identifiers.Proposals must be reusable, avoid hidden targets, and pass anti-leakage and runtime-judgability checks.
  • S1.1 Clinical Skill Bank: Clinical Skill mutations may add, merge, patch, deprecate, or discard artifacts, with failed cases producing correction strategies rather than memorized answers.Medication, procedure, escalation, and monitoring policies include relevant safety checks.

S2.3 Measurement Agent prompts

The Measurement Agent uses staged visual localization and review to produce evidence reports for the Doctor without issuing final diagnoses. Its prompts separate image observations from non-image evidence and govern reusable measurement-bank updates after completed cases.

  • S2.3 Measurement Agent prompts: A two-stage image prompt first identifies modality and defensible regions, then combines original pixels, optional MedSAM outputs, context, and retrieved measurement skills.
  • S2.3 Measurement Agent prompts: The visual locator returns one structured entry per panel with modality, segmentation applicability, visible findings, confidence, and at most two ROI boxes.ROI coordinates use a 0..1000 panel frame and are limited to defensible abnormal or decision-salient regions.
  • S2.3 Measurement Agent prompts: The reviewer independently verifies localization against original pixels, treats MedSAM as a localization aid, and separates non-image evidence from image observations.Unreliable masks are reported as unavailable rather than implying segmentation findings.
  • S2.3 Measurement Agent prompts: The visual report contains panel findings, mask-derived observations, cross-panel synthesis, limitations, and a segmentation assessment for the Doctor.
  • S2.3 Measurement Agent prompts: Post-training reflection evaluates whether visual measurement helped, what evidence was missed or overstated, and whether a reusable modality–task checklist should be maintained.Measurement-bank proposals prefer discard when no generalizable visual lesson exists and require verification of mask-derived measurements against original pixels.

S2.4 Online candidate generation and preference criticism

Online candidate generation produces distinct next actions from visible state, while a Clinical Process Critic ranks them using evidence alignment, safety, triage, and process constraints. Final-turn candidates must be complete diagnosis-ready plans.

  • S2.4 Online candidate generation and preference criticism: The Doctor receives visible dialogue, observations, the Process Rule ledger, retrieved Clinical Skills, and Symbolic Schema state at each non-deterministic turn.
  • S2.4 Online candidate generation and preference criticism: The generator creates three distinct next actions, including a focused question, at most one atomic test request, and another question or diagnosis-ready action.
  • S2.4 Online candidate generation and preference criticism: A requested test must separate named leading diagnoses, change a decision, and include a stop rule; unavailable tests must not be repeated or treated as negative evidence.
  • S2.4 Online candidate generation and preference criticism: The Clinical Process Critic scores candidates using visible state and general safety, prioritizing targeted evidence acquisition, diagnostic specificity, prerequisites, escalation, and efficient testing.
  • S2.4 Online candidate generation and preference criticism: On the final step, every candidate must be DIAGNOSIS_READY and contain complete diagnosis, evidence, treatment, safety, and follow-up fields.The critic also evaluates diagnosis support, treatment completeness, monitoring, triage, and follow-up as a coherent plan.

S2.5 Final risk audit, rewrite, and certification

The evaluation pipeline independently audits clinical risk, rewrites the final response, and certifies release using visible evidence, process obligations, and strict output checks.

  • Final risk audit, rewrite, and certification: The Final Rewriter produces one complete JSON response containing diagnosis, differential diagnoses, evidence, tests, treatment, safety checks, and follow-up or escalation.It cannot request more evidence and must incorporate every evidence-supported required correction.
  • Final risk audit, rewrite, and certification: The Release Certifier independently reconstructs the highest-risk problem and releases the answer only when no material evidence-integrity, treatment, medication, or disposition defect remains.The certifier assumes the rewritten diagnosis may be wrong.
  • Final risk audit, rewrite, and certification: The diagnosis-blind safety frame identifies visible red flags, high-harm pathways, missing prerequisites, and required monitoring before seeing a proposed diagnosis or treatment.The frame forbids guessing hidden labels, inventing findings, or treating missing evidence as negative.
  • Final risk audit, rewrite, and certification: The final risk auditor reconciles medications and procedures with allergies, contraindications, physiology, interactions, prerequisites, red flags, escalation needs, and time-critical care.Unknown prerequisites require active acquisition, a safe alternative, or deferral.
  • Final risk audit, rewrite, and certification: Deterministic validators reject unavailable-result citations, controller-invalid sources, missing final fields, invalid action contracts, schema errors, and unsafe memory mutations.The evaluator combines deterministic rules with a semantic judge only where exact matching cannot represent clinical equivalence or observable process quality.
  • Final risk audit, rewrite, and certification: The evaluator scores treatment-intent coverage rather than raw drug-string equality and penalizes missing critical intents most strongly.It also caps treatment accuracy at 0.4 for contraindicated or materially unsafe treatment.

S3.2 Controlled Clinical Stress Evaluator Prompt

Stress V2 evaluates whether agents recover and use controller-released evidence safely across diagnosis, evidence, behavior, treatment, temporal, and triage dimensions.

  • S3.2 Controlled Clinical Stress Evaluator Prompt: Stress V2 separates deterministic target release and unavailability from semantic assessment of visible clinical use and safety.Recovery credit cannot be created by the semantic judge, and hidden source values do not enter runtime prompts merely because they are later evaluated.
  • S3.2 Controlled Clinical Stress Evaluator Prompt: The evaluator returns used target IDs, integrated timeline IDs, Boolean process and safety fields, transcript evidence indexes, and a rationale constrained to released targets and existing turns.The controller log is authoritative about release and unavailability.
  • S3.2 Controlled Clinical Stress Evaluator Prompt: Diagnosis difficulty measures delayed-discriminator release and use, while evidence completeness measures unavailable-history or unavailable-test requests and safe handling of unavailable responses.Hallucinating an unavailable value is penalized.
  • S3.2 Controlled Clinical Stress Evaluator Prompt: Patient-behavior and treatment dimensions measure focused-question recovery and delayed medication, allergy, pregnancy, or renal prerequisite recovery, respectively.The patient-behavior dimension also records unrecovered-target failure, while treatment evaluates prerequisite-sensitive prescription handling.
  • S3.2 Controlled Clinical Stress Evaluator Prompt: Temporal and triage dimensions measure delayed timeline-fact recovery and integration, delayed-red-flag recovery, escalation adequacy, and unsafe reassurance.The stress evaluator distinguishes safe conditional treatment, alternatives, or deferral from unsafe action.
  • S3.3 Case-Level and Aggregate Formulas: Case-level formulas define diagnosis, treatment, history, test, safety, critical-failure, efficiency, and stress-process quantities from eligible sets and validated outputs.The auxiliary required-action value combines dimension-specific recovery and use obligations.
  • S3.3 Case-Level and Aggregate Formulas: Unavailable components are removed and remaining weights renormalized; diagnosis and treatment receive half the auxiliary stress composite’s nominal weight, so it is reported only as an auxiliary summary.Empty required sets receive recall one, empty request sets unnecessary-test rate zero, and missing required treatment receives zero.
  • S3.3 Case-Level and Aggregate Formulas: Aggregate percentages are macro-averages over declared eligible sets, with fixed-denominator recovery and adverse-event metrics using 30 cases per dimension.HqR and KTR use 10 construction-assigned cases, while conditional metrics follow their eligibility rule.

S3.4 Registered Comparator and Artifact Ledger

The comparator ledger standardizes training, testing, endpoints, and interaction budgets while documenting limits on reproducibility and clinical interpretation.

  • S3.4 Registered Comparator and Artifact Ledger: All learned Qwen comparators use the same 700 training encounters, publish method-native memory before testing, and remain frozen for the same 300-case test.Only the Doctor varies; Patient, Measurement, moderator/evaluator, case order, and the six-turn ceiling are shared.
  • S3.4 Registered Comparator and Artifact Ledger: Agent-KB, ExPeL, MemP, Reflexion, and SkillWeaver use adapted native memory prompts without the MediSkill-Evo harness, audit, rewrite, or certification path.The ledger distinguishes these comparator adaptations from the proposed system.
  • S3.4 Registered Comparator and Artifact Ledger: The registered calls use Qwen3.6-Flash and DeepSeek-V4-Flash aliases through specified endpoints at temperature zero, without portable seeds or immutable provider revisions.The shared budget is a Doctor-turn ceiling rather than a match on semantic items, words, patient burden, internal calls, tokens, latency, or cost.
  • S3.4 Registered Comparator and Artifact Ledger: Method blindness and evidence indexing improve internal comparability but cannot eliminate controller–evaluator rubric alignment, automatically generated target validation, or same-backbone moderator calibration to independent clinical judgment.Absolute clinical interpretation remains unsupported.

S3.5 Existing-Trace Paired Transitions

The paired-trace section documents frozen FullChain case transitions and one observable clinical encounter, including request handling, evidence use, treatment planning, and scoring.

  • S3.5 Existing-Trace Paired Transitions: Table S3 compares already frozen AgentClinic and MediSkill-Evo outputs across 300 shared FullChain case indices without rollout or re-judging.Metric direction determines which transition is labeled better, and released artifacts include hashes and per-case rows.
  • S3.5 Existing-Trace Paired Transitions: The recorded case concerns a 24-year-old woman with acute right upper-quadrant pain, with focused history, examination, laboratory, diagnostic, and treatment objectives.The protocol includes registered requestable tests.
  • S3.5 Existing-Trace Paired Transitions: The retrieved memory context supplies biliary triage and acute-abdomen baseline protocols for the encounter.These memories are shown as part of the observable interaction context.
  • S3.5 Existing-Trace Paired Transitions: The interaction elicits fever, chills, nausea, vomiting, jaundice, medication use, allergies, substance use, exposures, family history, and pregnancy status; one requested test is unavailable.The transcript records the unavailable result explicitly rather than treating it as negative.
  • S3.5 Existing-Trace Paired Transitions: The scored prediction identifies acute cholecystitis and cites persistent fatty-meal-triggered pain, nausea, vomiting, leukocytosis with neutrophil predominance, and absent fever, chills, or jaundice.The recorded outcome marks the diagnosis-ready prediction correct against the gold diagnosis.
  • S3.5 Existing-Trace Paired Transitions: The treatment plan requests baseline laboratory testing, withholds medications pending renal, hepatic, bleeding, allergy, and pregnancy verification, and seeks alternative imaging because ultrasound is unavailable.It also specifies urgent admission or emergency transfer, monitoring, serial examinations, and escalation triggers.

S5 NEJM Boundary Case: Original Record and Paired Interaction

The NEJM boundary case uses a frozen interactive record in which the agent must gather focused evidence and request available tests before diagnosing. Evaluation reserves reference information and unrevealed results from the interaction.

  • The frozen record distinguishes interaction-visible information from evaluation-reserved fields, and correctness refers only to the automatic final-diagnosis score.
  • The case concerns a 25-year-old woman with blurred vision, headaches, transient visual obscurations, and severe obesity.
  • The doctor objective requires focused history, review of the supplied examination, selective requests for available tests or imaging, and one most likely diagnosis.
  • Unlisted results are unavailable rather than normal, preserving the record’s evidence boundary.

S5.2 Paired test protocol and outcomes

The paired protocol compares identical interactive records under local no-MedSAM and remote with-MedSAM conditions. It controls the model, banks, evaluator, request gate, and inference budget, while acknowledging that the comparison is not tool-only.

  • The protocol uses the same source index, interactive test record, frozen Doctor banks, model, inference budget, request gate, and evaluator in both runs.
  • The registered tool condition determines whether the Measurement Agent may invoke MedSAM.
  • Because the Measurement Bank evolves on the corresponding training condition, the paired comparison is not a tool-only intervention.
  • The local condition disables MedSAM and produces zero nonempty masks with an incorrect final diagnosis score.
  • The remote condition enables MedSAM and produces three nonempty masks with a correct final diagnosis score.

S5.3 Local condition: Measurement learning without MedSAM

The local condition develops Measurement behavior without MedSAM and conducts a request-gated multimodal encounter using qualitative morphology rather than automated segmentation metrics. Its final diagnosis is scored incorrect against the recorded gold diagnosis.

  • The interaction requires the agent to request NEJM_Medical_Image before returning DIAGNOSIS READY, while available names do not reveal results.
  • The patient interaction elicits headache features, transient visual obscurations, medication history, substance use, exposures, family history, and pregnancy information.
  • The Measurement Agent reports flattened posterior globes and an empty sella, with no MedSAM segmentation or automated quantitative metrics.
  • The generated plan identifies suspected secondary intracranial hypertension and recommends specialist evaluation, monitoring, and deferred or verified treatment steps.
  • The recorded gold diagnosis is idiopathic intracranial hypertension, and the no-MedSAM prediction is scored incorrect.

S5.4 Remote condition: Measurement learning with MedSAM

The remote condition enables MedSAM and produces segmented multimodal observations during the same style of request-gated encounter. The system reaches the recorded diagnosis, but the trace remains a qualitative failure-boundary illustration rather than clinical or causal evidence.

  • The encounter requests Head_MRV and reports transverse sinus stenoses without obstruction or thrombosis, alongside flattened posterior globes and an empty sella.
  • MedSAM-generated observations localize bilateral orbital foci and a midline inferior brain abnormality, while preserving raw-image findings.
  • The remote plan recommends urgent neurosurgical and ophthalmologic evaluation, monitored admission, and withholding acetazolamide until pregnancy and baseline laboratory checks are verified.
  • The MedSAM prediction is scored correct for idiopathic intracranial hypertension, but that label does not endorse the visual evidence or management plan.
  • The authors characterize the pair as a qualitative failure-boundary illustration, not positive clinical evidence or a causal estimate that MedSAM improved the case.
Loading 2608.23397v1…