Source-linked AI summary

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou, Aleksey Ropan, Pavel Satalkin

arXiv:2609.09070v1cs.CLcs.AIcs.HC

TL;DR

The paper asks whether clinical AI evaluation should assess diagnosis and management after adaptive information gathering rather than fixed-information tasks. It compares Doctorina, physicians, and standalone frontier language models in 150 synthetic Polish-language primary-care consultations, finding that Doctorina outperformed physicians across primary diagnosis, diagnostic workup, and initial treatment, with advantages reproduced in a second execution.

  • Problem

    Comparative clinical AI performance depends on the task, available information, comparator expertise, interaction design, and outcome definition, motivating evaluation after adaptive information gathering.

  • Method

    The study compared Doctorina with eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care adaptive consultations using shared cases and endpoint definitions.

  • Results

    Doctorina outperformed physicians across diagnosis and management after adaptive consultation, and its positive physician-comparison differences were reproduced in a second eligible execution.

  • Takeaways & Limitations

    Primary-diagnosis selection, diagnostic workup, and initial treatment should be assessed as distinct dimensions in adaptive clinical consultation evaluations.

  • Takeaways & Limitations

    The purposive, non-randomized physician sample and reader-by-case structure constrain generalization to broader clinician populations.

Abstract

from arXiv · show

Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.

Introduction

Adaptive clinical consultation requires respondent-directed questioning, sequential diagnostic reasoning, investigation selection, and initial management planning. This study evaluates the integrated Doctorina system against physicians and standalone frontier language models in standardized Polish-language consultations.

  • Introduction: Adaptive consultation makes performance contingent on each respondent’s information-seeking decisions.Respondents determine what to ask and when to complete the encounter.
  • Introduction: Doctorina combines task-specific instructions, agent coordination, and consultation-state management to support adaptive, multi-turn consultations.The integrated product produces a most-likely diagnosis, differential diagnoses, diagnostic workup, and initial treatment.
  • Introduction: Direct comparison of an integrated patient-facing clinical AI product with physicians and standalone frontier language models addresses performance after adaptive consultation.The comparison evaluates diagnostic and management outputs from each system conducting adaptive consultations.
  • Introduction: The study used 150 synthetic Polish-language primary-care cases administered as adaptive simulated consultations.Doctorina was compared with eight physicians and four standalone models, with Doctorina–physician Top-1 concordance designated as the primary comparison.
  • Introduction: The six-group analyses comprised 150 equally weighted diagnostic cases and 149 complete management case pairs.Doctorina and physician outputs were reviewed together, while standalone-model outputs were reviewed in a second aligned block.

Diagnostic concordance

Doctorina showed higher diagnostic concordance than physicians and closely ranked with the strongest standalone models, while management estimates placed Opus, Doctorina, and Kimi near one another. The figure summarizes case-standardized diagnostic and management outcomes across six groups.

  • Diagnostic concordance: 82.0% versus 57.0% Top-1 concordance favored Doctorina over physicians by 25.0 percentage points.The 95% paired-case bootstrap CI was 17.7 to 32.7 percentage points.
  • Diagnostic concordance: 121 of 150 cases had the same Top-1 status for Doctorina and Kimi, while Doctorina alone was concordant in 17 cases and Kimi alone in 12.Doctorina and Kimi therefore had 80.7% case-level agreement in Top-1 status.
  • Diagnostic concordance: 97.3% versus 85.0% primary-or-reference-differential concordance favored Doctorina over physicians by 12.3 percentage points.The 95% CI was 7.0 to 18.0 percentage points.
  • Diagnostic-workup and initial-treatment outcomes: 89.4 versus 66.9 diagnostic-workup scores favored Doctorina over physicians across 149 complete case pairs.Opus had the highest diagnostic-workup point estimate at 90.3, followed by Doctorina at 89.4 and Kimi at 86.4.
  • Diagnostic-workup and initial-treatment outcomes: 83.7 versus 61.2 initial-treatment scores favored Doctorina over physicians.Opus, Doctorina, and Kimi had closely spaced management point estimates, with intervals spanning modest differences in either direction.
  • Diagnostic-workup and initial-treatment outcomes: 5.3% of Doctorina cases versus 20.7% of physician cases received poor or very poor workup verdicts.Treatment proportions receiving poor or very poor verdicts were 7.3% for Doctorina and 26.3% for physicians.

Consistency across observed physician assignments

Doctorina’s advantages over physicians remained positive across assignment blocks, complexity strata, outcomes, and a second execution, while management comparisons with the strongest standalone models varied by execution. The findings also distinguish diagnostic prioritization, workup, and treatment as related but separate dimensions.

  • Assignment consistency: Positive Doctorina-minus-physician differences appeared in every physician packet block for Top-1 concordance, diagnostic workup, and treatment.Top-1 differences ranged from 10.0 to 46.7 percentage points; workup differences ranged from 13.3 to 30.8 normalized-score points and treatment differences from 14.7 to 38.3 points.
  • Complexity consistency: 83.5%, 78.9%, and 80.0% were Doctorina’s Top-1 concordance estimates across complexity levels 1, 2, and 3, versus 63.9%, 47.4%, and 36.7% for physicians.Interaction intervals were compatible with similar Doctorina–physician margins across recorded complexity levels.
  • Second execution: 23.7 percentage points, 10.3 points, 18.3 normalized-score points, and 16.9 points were the reproduced Doctorina-minus-physician differences for the four outcomes.The respective 95% confidence intervals remained positive across Top-1 concordance, the broader diagnostic endpoint, workup, and treatment.
  • Standalone-model comparisons: Management comparisons varied by execution, although Doctorina’s diagnostic point estimates relative to Opus and Kimi remained positive in both executions.The second execution produced negative workup and treatment differences versus both Opus and Kimi, while diagnostic comparisons remained positive.
  • Outcome interpretation: The weak association between diagnostic concordance and management scores, together with discordant diagnosis–management combinations, supports separate assessment of diagnosis, workup, and treatment.Case-level errors were partly shared and partly complementary across AI systems, suggesting disagreement may identify cases for additional review.
  • Execution agreement: Agreement across executions was high for diagnostic status but less exact for management, and run-level management differences were influenced by technical completion.Valid-pair intervals included zero, while comparisons with Opus and Kimi also varied by execution.

Study design and analytical scope

The study compared complete AI implementations and physicians in adaptive, Polish-language simulated consultations using a shared synthetic case corpus and standardized outcomes. The design explicitly treated respondent-directed information gathering as part of the evaluated task.

  • Study design: Six respondent types were compared in Polish-language simulated primary-care consultations: physicians, Doctorina, Gemini 3.1 Pro Preview, Claude Opus 5, GPT-5.6-sol, and Kimi K3.Eight physicians completed allocated subsets, while two Doctorina executions and one execution of each standalone model attempted all 150 cases.
  • Study design: All groups used the same case corpus, constructed references, clinical task, and outcome definitions.Case-standardized effects evaluated each complete AI implementation or the observed physician cohort under the shared task.
  • Adaptive consultation: Final diagnostic and management outputs were contingent on the questions selected by each respondent during adaptive consultation.The evaluation therefore included respondent-directed information seeking before scored outputs.
  • Reporting: Relevant items from STARD 2015, STARD-AI, and TRIPOD-LLM informed reporting.The study design and information boundaries were summarized in Fig. 2.

Setting and physician participants

The physician comparison used eight purposively recruited Polish clinicians in a controlled, approximately six-hour in-person session. Participants included specialists and residents who used a common interface and response format.

  • Setting: The physician comparison was conducted during one approximately six-hour in-person session in Warsaw, Poland.Recruitment used purposive direct professional outreach, with eligibility requiring Polish licensure and current family, general, or internal medicine practice or training.
  • Participants: Eight physicians completed the study: five specialists and three second-year residents.All participants practised the interface and response format with one non-analytical test case before assessment.
  • Procedure: All eight physicians worked contemporaneously in the same controlled space under common instructions.An independent observer unaffiliated with Doctorina was present during the assessed session.

Clinical-vignette corpus and physician allocation

The study used a clinician-reviewed synthetic primary-care corpus built from common eligible diagnoses and assigned physicians paired 30-case packets with reordered case membership.

  • Corpus construction: 150 synthetic cases were constructed from eligible diagnoses informed by 2024 National Health Fund patient counts.The most common eligible conditions were selected and scaled to define the diagnostic distribution.
  • Corpus construction: Every case field underwent review by three clinicians, with revisions followed by renewed review before acceptance.The final accepted fields constituted the constructed case references.
  • Corpus composition: 82 cases described female patients and 68 described male patients.The corpus contained 127 distinct disease labels and 128 distinct ICD-10 values.
  • Physician allocation: Ten 30-case packets formed five paired case-membership sets, with each pair containing the same cases in different orders.Each physician received one packet, and eight completed assignments yielded 240 retained physician diagnostic consultations.

Consultation environment and physician procedure

Respondents completed adaptive, case-bounded consultations from the same controlled information environment and submitted diagnoses, differentials, workup, and treatment outputs. Doctorina was evaluated through two provenance-eligible 150-case executions alongside standalone language-model comparators.

  • Consultation environment: Adaptive text dialogues began with each presenting complaint and allowed respondents to determine the sequence and content of questioning.The simulator generated brief, truthful replies from each case’s permitted-facts bank within a fixed evidence environment.
  • Physician procedure: Physicians worked independently through an external text interface using dialogue-only evidence and submitted a diagnosis, differentials, diagnostic workup, and initial treatment.They determined when to end each consultation under standardized closed-resource conditions.
  • Doctorina arm: Doctorina was evaluated as a complete clinical AI product with coordinated doctor, assistant, recommendation, translation, attachment-processing, reference, and follow-up functions.The execution protocol used isolated case sessions and required the same four clinical domains as the physician procedure.
  • Doctorina arm: Runs 479 and 481 were the provenance-eligible Doctorina executions, with run 481 primary and run 479 used for separate analysis.Both runs attempted the same 150 cases; run 481 supplied the primary estimates.
  • Standalone comparators: Gemini 3.1 Pro Preview, Claude Opus 5, GPT-5.6-sol, and Kimi K3 each underwent corpus-wide standalone execution under the same case-specific consultation design.Gemini, Opus, and Kimi produced final outputs for all 150 cases, whereas GPT produced 148.

Data capture and quality control

Data capture linked source records, respondent outputs, provenance, completion states, and analytical outcomes, with separate validation of physician and standalone-model datasets.

  • Data integration: A case–respondent ledger linked source values and formulas to identifiers, provenance, completion states, inclusion flags, and analytical outcomes.Analyses integrated case-distribution, packet, rating, execution-log, and source-case data.
  • Physician data: 240 assigned physician diagnostic observations were retained after reconstructing assignments, while the management cohort contained 239 records across all 150 cases.One unmatched physician response remained outside the analytical cohort.
  • Model data: Standalone-model datasets were validated for unique case ordinals, complete start and finish markers, opening-complaint agreement, and completion-status consistency.

Physician review and outcomes

Physician and standalone-model outputs were evaluated in aligned review blocks against the same constructed case references, with escalation procedures for uncertain judgments.

  • Review process: Physician and Doctorina outputs underwent manual evaluation by resident-physician reviewers, with uncertain judgments escalated to a senior physician.One primary evaluator reviewed each output.
  • Review process: Standalone-model outputs were evaluated by four physician reviewers using the same case references, endpoint definitions, and rating direction.For each case, the same reviewer assessed all four standalone-model outputs.

Diagnostic concordance

Diagnostic outcomes compared submitted primary diagnoses with a designated primary diagnosis or reference differential, while workup and treatment were rated independently on a five-category scale and normalized for comparison.

  • Diagnostic endpoints: Top-1 concordance compared the submitted most-likely diagnosis with the designated primary diagnosis, accepting clinically equivalent terminology and Polish–English equivalents.For Doctorina, the patient-visible terminal recommendation supplied the primary diagnosis; standalone LLMs used their first listed diagnosis.
  • Diagnostic endpoints: Primary-or-reference-differential concordance counted a match when the submitted primary diagnosis matched either the designated primary diagnosis or a condition in the reference differential.
  • Management endpoints: Diagnostic workup and initial treatment were rated independently of diagnostic correctness on a five-category scale from Excellent to Very Poor, with lower verdicts indicating better performance.Adequate was anchored to essential investigations or basic therapy addressing the patient’s main needs.
  • Record handling: Technical errors and incomplete attempts were assigned outcome value 0 in all-attempt analyses, while clinically discordant diagnoses and numeric management ratings remained distinct.

Statistics and reproducibility

The study prespecified a case-standardized analytical hierarchy, paired bootstrap uncertainty, execution-level sensitivity analysis and exploratory robustness assessments, with reproducibility supported by accompanying data and code.

  • Analysis plan: The Doctorina–physician Top-1 comparison was primary; primary-or-reference-differential concordance was secondary, while management and standalone-model comparisons were exploratory.The standalone-model analysis plan was finalized before completion of the standalone-model physician-review block.
  • Case-standardized design: The primary comparison used all 150 cases and eight completed physician assignments, with physicians represented by within-case means and effects reported as paired Doctorina-minus-comparator differences.Management comparisons comprised 149 complete case pairs.
  • Uncertainty: 10,000 paired whole-case bootstrap resamples produced two-sided 95% percentile confidence intervals while preserving the observed reader-by-case and review-block structure.The intervals quantified case-level variation conditional on the observed physician sample, system executions and retained final ratings.
  • Weighting: Case weighting gave every rated case one unit of probability mass, splitting weight equally between two physician ratings and assigning full weight to single-rating cases.
  • Sensitivity analysis: Run 479 replaced primary run 481 in the execution-level sensitivity analysis, with all 150 cases contributing to each sensitivity outcome.Figure 3 summarizes the observation structure, endpoints and case-standardized estimand.
  • Robustness: Exploratory analyses assessed poor-management rates, Doctorina-centred win–tie–loss summaries and assignment-stratified consistency.
  • Run agreement: Observed agreement between runs 479 and 481 used raw agreement and Cohen’s κ for binary outcomes, with exact and within-one-category agreement for technically valid management records.
  • Reproducibility: Supplementary data provide metadata, allocation, rating ledgers, case-standardized analyses, sensitivity results, arm estimates, contrasts and robustness materials; supplementary code reproduces primary estimates and validates files.

Competing interests

Several authors had professional or technical affiliations with Doctorina, while other listed authors declared no competing interests.

  • Author disclosures: H.P. was Doctorina’s co-founder and Chief Medical Officer, while other authors held Doctorina clinical, technical-development or affiliated-medical roles.P.G., J.M. and V.H. declared no competing interests.
Loading 2609.09070v1…