Source-linked AI summary

From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

Xinyu Li, Zijian Li, Mengyu Xia, Luzhen Tang, Naping Chen, Changmin Lin, Danijela Gasevic, Dragan Gasevic, Yizhou Fan

arXiv:2608.28619v1cs.CLcs.AIcs.HC

TL;DR

Medical history-taking logs are rich but difficult to use for formative review because transcripts overwhelm routine analysis while final scores conceal the reasoning process. This study coded 1,030 GenAI virtual-patient dialogues and applied prevalence, co-occurrence, and transition analyses anchored to weekly teacher ratings. High-rated consultations showed more coordinated information gathering, communication, checking, organisation, synthesis, and reasoning-oriented follow-up, supporting process-focused feedback.

  • Problem

    GenAI virtual-patient transcripts are too detailed for routine review, while final scores obscure how learners follow cues, check uncertainty, and organise questioning.

  • Method

    The study analysed coded GenAI virtual-patient dialogues using behavioural prevalence, Epistemic Network Analysis, and Transition Network Analysis against weekly teacher-rated consultation scores.

  • Results

    High-rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, synthesis, and reasoning-oriented follow-up.

  • Takeaways & Limitations

    Layered analysis of GenAI virtual-patient logs can reveal teacher-interpretable process patterns associated with higher-rated history taking and support process-focused feedback.

  • Takeaways & Limitations

    The learner-only coding, single course and GenAI environment, five chest-pain cases, and GPT-3.5 setting constrain claims about full interactions and stable process signatures.

Abstract

from arXiv · show

Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue. However, these logs are educationally difficult to use directly. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning. This study examined whether coded GenAI VP dialogues can provide teacher-interpretable process evidence of clinical reasoning. We analysed 1{,}030 GenAI VP dialogues from 210 second-year medical learners across five weeks chest-pain cases. Each consultation was teacher-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high- or low-rated using the weekly median score. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis. High-rated consultations involved more history taking activity, but differences were not simply about volume. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis. Summarising and organising moves more often led to verification or mechanism-oriented follow-up. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.

1 Introduction

History taking requires learners to reason through dialogue while gathering, clarifying, organising, and integrating patient information. GenAI virtual patients make repeated practice and full dialogue capture feasible, motivating layered analysis of coded consultations against teacher-rated performance.

  • History taking combines clinical conversation with ongoing reasoning about which information matters and which questions should follow.
  • Final scores and full transcripts provide incomplete formative evidence about how learners followed cues, checked uncertainty, and organised questioning.
  • GenAI virtual patients support natural-language questioning, varied consultation paths, scalable practice, and preservation of turn-by-turn dialogue data.
  • The study compared high- and low-rated consultations within each weekly case rather than classifying learners as generally strong or weak.
  • Behavioural prevalence, local co-occurrence, and sequential transition analyses examined presence, coordination, and ordering of coded history taking behaviours.

2 Literature Review

History taking is a dialogue-based form of clinical reasoning whose observable behaviours can provide process evidence, although dialogue records do not directly measure cognition. The study addresses how coded behaviours relate to consultation quality while preserving their local and sequential context.

  • 2.1 History taking as observable metacognitive regulation of clinical reasoning: History taking involves eliciting details, clarifying ambiguity, following cues, and reorganising information during the consultation.
  • 2.1 History taking as observable metacognitive regulation of clinical reasoning: Dialogue records preserve behavioural traces of planning, monitoring, uncertainty regulation, integration, and movement toward reasoning-oriented questioning, but not metacognitive states directly.
  • 2.2 Virtual patients and dialogue data: Earlier virtual patients improved repeatability and recording but often constrained history taking through predefined options or scripted branches.
  • 2.2 Virtual patients and dialogue data: GenAI virtual patients enable natural-language questions and multiple learner-generated consultation paths, shifting analysis toward complete dialogue corpora.
  • 2.3 Learning analytics from dialogue records to process evidence: A coded learner turn gains educational meaning from neighbouring utterances, so counts alone cannot show how behaviours are positioned around other moves.
  • 2.3 Learning analytics from dialogue records to process evidence: The study anchored process analysis to teacher-rated rubric scores and posed questions about behavioural prevalence, local co-occurrence, and sequential transitions.

3.1 Participants

The study involved 210 second-year medical students enrolled in a symptomatology and history taking course. Weekly analyses used 205–207 valid dialogue logs, with 197 students contributing complete five-week records.

  • 210 second-year medical students from an anonymised medical college in China participated during April and May 2024.
  • Weekly analytic samples comprised W1 n = 206, W2 n = 207, W3 n = 205, W4 n = 206, and W5 n = 206.
  • 197 students had complete dialogue logs across all five weeks.

3.2 Study design and learning environment

Participants completed an orientation followed by five history taking tasks with a GenAI virtual patient in an integrated learning platform. Analyses used dialogue logs and weekly teacher-rated scores based on the full consultation dialogue.

  • The study received ethics approval, obtained informed consent, and allowed participants to withdraw with their data removed.
  • Tasks were completed on an anonymised Moodle-integrated platform hosting instructional resources and a GPT-3.5 case-specific GenAI VP chatbot.
  • After orientation and training, participants completed five consecutive tasks involving instructions, GenAI VP history taking, optional notes or resources, and a diagnostic conclusion.
  • Weekly history taking scores assessed the full consultation dialogue rather than the submitted diagnosis alone.
  • The scoring rubric combined the Kalamazoo Essential Elements checklist, national licensing examination criteria, and course teaching requirements.
  • Twenty participants' five exercises were independently scored by three blinded raters to examine scoring consistency.

3.3 Dialogue dataset preparation

The dataset converted complete GenAI VP consultation dialogues into deidentified, quality-checked learner-behaviour records using a validated coding scheme. The codes represented observable dialogue behaviours rather than inferred metacognitive states.

  • Dataset preparation: Deidentified user IDs replaced identifiable information, and task-level checks retained records with complete consultation dialogue logs.Inclusion was determined week by week because analyses were conducted separately for each week.
  • Behavioural coding: Learner utterances were represented using a previously developed and validated 12-code behavioural scheme.The prior study established inter-rater reliability through iterative calibration and predefined thresholds.
  • Behavioural coding: The coding scheme operationalised task-relevant behaviours including hypothesis-driven questioning, follow-up to patient cues, organisation, and integrative synthesis.These codes treated history taking as a dialogue-based clinical reasoning process rather than a checklist of questions.
  • Analytic representation: Analyses used observed behavioural-code frequencies, local co-occurrences, and transitions rather than recoding utterances as metacognitive states.Dialogue logs captured what learners said, not their consciously planned, monitored, or regulated states.

3.4 Analysis plan

The analysis compared high- and low-rated consultations within each weekly case using behavioural prevalence, local co-occurrence, and sequential transition analyses. Additional modelling examined whether behavioural patterns remained associated with continuous weekly history-taking scores.

  • Grouping and scope: Analyses were conducted separately for weeks W1 to W5, treating each chest-pain task as a distinct case context rather than estimating longitudinal growth.The weekly cases differed in clinical content and context.
  • Grouping and scope: High-rated consultations were defined as weekly teacher-rated total scores at or above the median, while low-rated consultations fell below it.Group membership could vary across weeks, so contrasts represented consultation-level performance within a weekly case.
  • Behavioural prevalence (RQ1): Individual behaviour counts were compared between groups with two-sided Mann-Whitney U tests, false-discovery-rate control, and rank-biserial correlations.The reported effect size was r_rb = 2U/(n_high n_low)−1, with positive values indicating higher counts in the high-rated group.
  • Behavioural prevalence (RQ1): Rate-profile PERMANOVA used R^2 as a profile-level effect size and served as a length-control check.This complemented analyses of consultation length, individual behaviours, and overall behavioural composition.
  • Sensitivity analysis: Continuous-score sensitivity analyses fitted separate mixed-effects regressions using code rates to predict weekly HT-Total scores while accounting for total turns, week, and learner.The model was fitted separately for each of the 11 retained codes, with estimates and fit indices reported in Supplementary Table S13.
  • Local co-occurrence (RQ2): Epistemic Network Analysis modelled local co-occurrence among 11 retained utterance states within moving windows, separately by week and learner dialogue.OS was filtered out before ENA.
  • Sequential transitions (RQ3): Transition Network Analysis estimated first-order conditional transition probabilities between 11 retained behavioural states separately for each week and performance group.Transition differences were defined as Δp = P(high-rated)−P(low-rated).

4 Results

Across five weekly within-case contrasts, high-rated consultations showed more activity and distinct behavioural organisation. Layered analyses indicated that high-rated dialogues differed in local coordination and in how organising, checking, summarising, and cue-following behaviours connected across turns.

  • Behavioural prevalence: High-rated consultations contained more coded learner turns than low-rated consultations in every week.The descriptive gap was largest in W1 and smallest in W2.
  • Behavioural prevalence: SI and RQ differed between performance groups in all five weeks, while LO, CC, and SS differed in four weeks.Across significant raw-count contrasts, mean rank-biserial correlations ranged from 0.164 for RR to 0.424 for LO.
  • Behavioural prevalence: Rate-based differences were concentrated in LO, SI, CK, and RQ, narrowing the raw-count interpretation toward organisation, summarising/integrating, and checking.Significant rate contrasts occurred for LO in W3 and W5, SI in W1 and W3, CK in W1, and RQ in W3.
  • Behavioural composition: Raw-count behavioural composition differed significantly in all five weeks, whereas rate-profile composition differed only in W1 and W4.Raw-count PERMANOVA effect sizes ranged from R2 = .0225 in W2 to R2 = .1306 in W1; rate-profile effects were small where nonsignificant.
  • Local co-occurrence: High-rated consultations showed consistent separation in local coordination across all five weeks, with the RQ-CC pairing stronger in W1, W2, and W4.The first ENA dimension separated groups in every week, and sensitivity checks reproduced the separation across three-, five-, and six-turn windows.
  • Sequential transitions: Weekly consultations shared a common transition routine, but performance groups diverged in transitions involving CC and SI.High-rated dialogues more often moved from organising or cue-following to CC, and from SI to mechanism-oriented or checking moves; low-rated dialogues more often returned to routine questioning or repeated questions.

5 Discussion

The discussion interprets high-rated history taking as coordinated process rather than mere dialogue volume, while outlining instructional uses and important limits on generalisation and causal interpretation.

  • Performance-related process evidence: High-rated consultations had more coded learner turns, but length-adjusted and continuous-score analyses showed that performance differences were not independent of consultation length.Raw behavioural volume was the most stable group difference; LO, SI, and CC were the clearest positive score-related behaviours.
  • Performance-related process evidence: Symptom-specific questioning appeared more often in high-rated consultations by raw count but was negatively associated with total history-taking score after controlling for week and coded-turn count.Its educational meaning therefore depended on whether symptom details supported later organisation, checking, or synthesis.
  • Performance-related process evidence: High-rated consultations more often connected routine and symptom-specific questioning with communication, organisation, checking, and summarising.These behaviours gained meaning from their local coordination rather than from isolated frequency.
  • Performance-related process evidence: Summaries followed by checking or hypothesis-oriented questioning suggested verification or mechanism-focused follow-up, whereas summaries followed by repetition suggested limited redirection.Communication moves likewise depended on their sequence position, including whether they followed organisation or returned to routine questioning.
  • Instructional uses: Teacher-facing reports could flag questioning stretches that lack organisation, checking, or synthesis, while sequence-aware systems could prompt clarification or mechanism-oriented follow-up.Aggregated patterns could also inform curriculum review when learners repeatedly gather symptoms without integrating them into consultation structure.
  • Limitations and future work: These patterns are not ready-made scoring rules because the study identified performance-linked dialogue patterns without testing whether reports, prompts, or pattern-based review improve later performance.Future work should assess teacher interpretability, learner actionability, and effects on GenAI VP, SP, or OSCE performance.
  • Limitations and future work: The performance anchor was a course-based history-taking score, and future studies should test alignment with OSCEs, expert ratings, diagnostic reasoning assessments, and later clinical interviews.The analyses also contrasted weekly performance rather than estimating individual longitudinal growth.
  • Limitations and future work: Because only learner utterances were coded in one course, one GenAI VP environment, and five GPT-3.5 chest-pain cases, the full interaction and broader stability of signatures remain constrained.Replication with other models, symptoms, institutions, and VP designs is needed.

6 Conclusion

The study shows that coded GenAI VP dialogues can transform open-ended history-taking logs into interpretable process evidence linked to teacher-rated performance.

  • Conclusion: High-rated consultations coordinated questioning, symptom exploration, communication, clarification, organisation, and synthesis into an account used for verification or reasoning-oriented follow-up.The findings connect teacher-rated performance with concrete dialogue patterns that may inform future feedback and clinical reasoning practice.

Declarations

The manuscript anonymises identifying details for double-blind peer review and states that they will be supplied in the final manuscript.

  • Declarations: Author names, affiliations, funding sources, ethics approval numbers, and author contributions are anonymised during double-blind peer review.These details will be provided on the title page and in the final manuscript.

Funding

The declarations report no competing interests, ethics approval and informed consent, and data availability on reasonable request.

  • Declarations: The authors declare that they have no competing interests.
  • Ethics: The study received institutional ethics approval, and all participants provided written informed consent after receiving study information.
  • Data availability: Study datasets are available from the corresponding author on reasonable request.
  • Editorial policies: The manuscript lists Springer and Nature Portfolio editorial-policy links.
Loading 2608.28619v1…