Source-linked AI summary

AI as Teammate: Rethinking Task Distribution in Medical Training

Fendi Tsim, Alina Gutoreva, Anthony Weiss, Nicole Dubosh

arXiv:2608.28373v1cs.HCcs.AI

TL;DR

The paper addresses concerns that AI in medical training may erode clinical competence by reframing the problem as misclassification of task-appropriate AI interaction. It proposes SCAN, a metacognition-based task-allocation framework, and concludes that task-level classification and engagement distinguish upskilling from multiple forms of skill failure, while emphasizing that these claims require empirical testing.

  • Problem

    Medical training has incorporated AI faster than the conceptual vocabulary needed to evaluate its effects, while 97.1% of graduate medical trainees report no formal AI instruction.

  • Method

    The paper develops SCAN, a human-centric framework for systematic task allocation grounded in task classification and metacognitive regulation.

  • Results

    SCAN theoretically distinguishes how task type, learner stage, and engagement mode shape upskilling and the triad of skill failure in clinical reasoning development.

  • Takeaways & Limitations

    Reframing AI problems from misuse to misclassification gives educators a basis for curriculum design, supervision, and assessment focused on task classification and engagement.

  • Takeaways & Limitations

    SCAN remains a claim to be tested through evidence rather than an established empirical account.

Abstract

from arXiv · show

Integrating Artificial Intelligence (AI), particularly generative AI, into medical training has prompted concerns about learner over-reliance, misuse, and erosion of foundational clinical competencies. We propose a conceptual reframing at the decision level: the problem is not misuse but misclassification - a mechanistic failure of real-time metacognitive evaluation in selecting a subzone-inappropriate AI interaction mode. Drawing on "SCAN" (Substitute, Complement, Aid, Non-Negotiable), a human-centric decision-making framework for generative AI task allocation grounded in Vygotsky's Zone of Proximal Development and metacognition, we advance the emerging social-constructivist conversation around AI in medical education by offering a testable account of AI's role in clinical reasoning development. This framework yields testable predictions for how misclassification can be detected, mitigated, and, more importantly, prevented in the clinical learning environment. Regarding clinical reasoning development, we show how trajectories of skill acquisition (upskilling) and failure (the triad of skill failure: de-skilling, never-skilling, and mis-skilling) operate at the individual task level in ways that fixed-phase, cohort-wide treatments fail to capture. We further identify passive engagement within correctly classified AI-scaffolded tasks as a particularly insidious, detection-resistant pathway to mis-skilling - one requiring subzone re-identification from AI assistance to expert assistance, with human experts serving as epistemic auditors. The paper operationalizes SCAN for clinical curriculum design, supervision, and assessment, and opens an empirical research agenda grounded in cognitive science. This paradigm shift from misuse to misclassification is not semantic: it offers educators a clear perspective on what to look for, what to assess, and what to intervene on.

1. Introduction

The paper reframes AI over-reliance and misuse in medical training as task-level misclassification: selecting an AI interaction mode inappropriate to a learner’s developmental subzone. SCAN provides a metacognitive, task-level framework for distinguishing harmful from beneficial AI engagement and guiding curriculum, supervision, and assessment.

  • AI adoption in medical education is expanding faster than the conceptual vocabulary needed to evaluate its effects.
  • 97.1% of graduate medical trainees report no formal AI instruction, despite widespread reported AI use among trainees.The passages also report that 85.5% use AI tools for clinical decision support, academic writing, and research.
  • Existing institutional responses mainly operate at the cohort or policy level, while learners make moment-by-moment AI-use decisions for individual clinical tasks.
  • Studies link misleading or fluent AI outputs with reduced diagnostic accuracy, impaired confidence calibration, and potential erosion of foundational clinical reasoning.Misleading explanations can leave confidence disconnected from correctness, while explanation style shifts cognitive load and confidence calibration.
  • SCAN extends the Zone of Proximal Development by organizing tasks into Substitute, Complement, Aid, and Non-Negotiable subzones.The framework is grounded in social constructivism, Vygotsky’s ZPD, and metacognition.
  • Misclassification is defined as a real-time metacognitive failure that assigns a subzone-inappropriate interaction mode, disrupting the cognitive engagement needed for clinical reasoning development.The paper presents this as a falsifiable construct with operational pathways for curriculum design, supervision, and assessment.
  • The paper argues that clinical reasoning development and failure should be analyzed at the learner-task level rather than through fixed, cohort-wide treatments.It distinguishes developmental trajectories of upskilling and skill failure and proposes testable predictions about harmful and beneficial AI engagements.

2. Theoretical Foundation

SCAN frames generative-AI use in medical training as a task-allocation problem governed by the learner’s developmental position and metacognitive evaluation. Its four subzones specify when AI substitutes, complements, or aids learning, and when human expertise remains non-negotiable.

  • Framework: SCAN is a human-centric framework for systematic task allocation between learners, human experts, and AI.It organizes tasks by AI’s appropriate cognitive responsibility and the developmental consequences of that responsibility.
  • Task Paradigm: The Task Paradigm divides AI-supported work into Substitute, Complement, Aid, and Non-Negotiable subzones.The framework scans task identification, processing, and evaluation throughout task completion.
  • Task Paradigm: Substitute tasks are automated when learners lack task-specific knowledge and cannot independently verify the output.Complement tasks require both learner expertise and AI capability, whereas Aid tasks preserve active learner judgment.
  • Task Paradigm: Non-Negotiable tasks require human experts when relational, ethical, or embodied judgment is irreplaceable or automation’s developmental cost is unacceptable.Human experts function as more knowledgeable others in this subzone.
  • Metacognition: Metacognition is SCAN’s primary operational mechanism for monitoring and adjusting a learner’s competence boundary in real time.It includes metacognitive knowledge, monitoring, and control, which together support zone identification and task reassignment.
  • Metacognition: As AI reliance increases, accuracy in detecting AI errors decreases while user confidence does not, creating a performance–metacognition dissociation.Without adequate metacognition, learners may over-rely on AI or avoid it, impairing skill acquisition and clinical performance.

3. Clinical Reasoning Development

SCAN describes clinical reasoning development as task-level movement through distinct developmental trajectories rather than a uniform, cohort-wide process. Correct engagement supports upskilling, while misclassification and passive engagement produce distinguishable forms of skill failure requiring different interventions.

  • Developmental Trajectories: The framework identifies Internalization and Symbiosis trajectories as complementary routes through which clinical expertise develops with AI.Internalization produces situated expertise, while Symbiosis develops the metacognitive capacity to direct and evaluate AI.
  • Developmental Trajectories: Complete clinical formation requires both trajectories rather than treating either as a substitute for the other.The proposed endpoints are tacit clinical mastery and metacognitive responsibility for AI as an extended self.
  • Misclassification: Treating an Aid task as a Substitute task prevents progression by eliminating the cognitive struggle needed for learning.The learner delegates judgment they do not yet possess and cannot reliably supervise.
  • Misclassification: Treating an Aid task as a Complement task produces pseudo-competence by creating a false impression that the learner can direct and evaluate AI confidently.Plausible AI errors may pass through verification that feels rigorous but is not.
  • Developmental Trajectories: SCAN distinguishes upskilling from the triad of skill failure: de-skilling, never-skilling, and mis-skilling.These outcomes arise from how task classification and engagement shape clinical reasoning development.
  • Curriculum Implications: SCAN’s task-level account captures within-cohort heterogeneity because the same learner may require AI-free training for one task while still needing assistance for another.The framework therefore links developmental harm to task type, learner stage, and engagement mode.
  • Mis-skilling: Passive engagement within a correctly selected Aid subzone can still produce mis-skilling because the learner fails to exercise epistemic activity.The proposed response is subzone re-identification from AI assistance to human expert assistance, with experts auditing the learner’s reasoning.

4. Applications

SCAN translates AI use in clinical training from a binary permission question into an assessable decision about task-specific subzone classification and engagement. It applies this distinction to curriculum design, supervision, and clinical reasoning development, while identifying misclassification and passive offloading as risks to skill acquisition.

  • Curriculum design: SCAN makes subzone classification a teachable and assessable component of clinical curricula.The framework treats faculty development and visible, discussable subzone boundaries as necessary implementation supports.
  • Curriculum design: A task’s appropriate AI role depends on the learner’s expertise, so novice differential diagnosis may be Non-negotiable while resident reasoning may use AI as a scaffold.The same clinical task can occupy different subzones as learners progress from novice to expert.
  • Clinical supervision: Supervisors use SCAN to evaluate whether delegation is appropriately classified and whether engagement with AI output is active.Zone verification asks learners to justify assignments; output interrogation examines reconstruction of reasoning, failure detection, and clinical integration.
  • Clinical supervision: SCAN shifts supervision from output evaluation to process evaluation through zone verification and output interrogation.These actions are intended to surface miscalibration before it consolidates into habit.
  • Clinical supervision: Early learners’ overconfidence and limited metacognitive monitoring make faculty calibration structurally necessary because SCAN assignment accuracy has not yet been empirically tested.The paper links this concern to possible assignment of Aid or Complement when the learner’s knowledge state warrants Substitute.
  • Clinical reasoning development: Diagnostic tasks were associated with passive engagement and offloading, whereas management tasks prompted Aid-subzone engagement, illustrating divergent AI-supported learning outcomes.The reported comparison connects the performance gain and added time in AI-supported work to different subzone classifications.

5. Research Agenda

The paper proposes an empirical agenda to test whether metacognitive accuracy in SCAN classification predicts clinical reasoning, whether misclassification pathways are mechanistically distinct, and how passive engagement and mis-skilling can be detected or remediated. It connects these tests to process-sensitive supervision and metacognitive remediation.

  • Detection and remediation: Longitudinal process measures would examine whether passive engagement during correctly classified Aid interactions leaves detectable reasoning signatures before mis-skilling appears in assessment outcomes.Suggested measures include think-aloud protocols, verify-and-trust tasks, pre/post-consultation reasoning comparisons, and supervisory ratings.
  • Detection and remediation: The proposed remediation study would test whether mis-skilled reasoning is remediable and which behavioral intervention features are necessary.The paper frames mis-skilling as a consolidated reasoning schema and emphasizes metacognitive monitoring, control, scaffolded practice, and corrective feedback.
  • Prediction and measurement: A prospective cohort study would test whether baseline and changing zone-classification accuracy predict subsequent clinical reasoning development.Clinical reasoning would be assessed with the Clinical Reasoning Examination and structured OSCE performance.
  • Prediction and measurement: The proposed study predicts that changes in zone-classification accuracy will precede changes in clinical reasoning performance.This specifies a temporal prediction rather than only a cross-sectional association.
  • Mechanism-specific intervention: The paper proposes testing whether two never-skilling pathways have distinguishable predictive signatures.The pathways are characterized as no subzone migration from Substitute and misclassification.
  • Mechanism-specific intervention: A randomized 2x2 trial would test whether pathway-matched training reduces misclassification only for the targeted pathway.A double dissociation would support mechanistic distinctness, whereas uniform improvement would falsify that claim.

6. Discussion

The paper reframes AI-related developmental harm in medical training as misclassification rather than misuse, offering task-level concepts and testable predictions for clinical education. It extends this account through distinct skill-failure pathways, practical redesign proposals, and explicit implementation limitations.

  • Theoretical Contributions: Misclassification specifies how developmental harm varies across task types, learner stages, and engagement modes, while identifying corresponding interventions.This reframing shifts the explanatory locus from behaviour to cognition and from compliance to competency.
  • Theoretical Contributions: SCAN distinguishes upskilling from three theoretically distinct skill-failure pathways: de-skilling, never-skilling, and mis-skilling.The paper argues that these pathways were previously conflated under AI-related skill erosion, with consequences for intervention design.
  • Theoretical Contributions: Passive Aid engagement is identified as an insidious, detection-resistant mis-skilling pathway requiring task re-identification to Non-Negotiable and human experts as epistemic auditors.This extends the framework beyond developmental analysis to a proposed mechanism for prevention and empirical inquiry.
  • Practical Contributions: The framework offers educators a shared, testable vocabulary for task-level evaluation and redesign across curriculum design, supervision, and assessment.Proposed applications include explicit instruction in metacognitive classification, supervision that preserves reasoning observability, and assessment of process-level failures before output deficits.
  • Broader Implications: Task-level specificity addresses AI-related skill failure that fixed-phase, cohort-wide curricular structures fail to capture.The discussion also connects the framework to broader professional education where generative AI supplies substitutes for cognitive processes required for development.
  • Limitations: SCAN requires metacognitive capacity and domain-specific knowledge that novice learners may lack, while implementation requires supervisor training and demanding process-oriented debriefing.The framework remains empirically untested, and its claims should be read as theoretically grounded hypotheses pending validation; never-skilling also requires direct clinical evidence.

7. Conclusion

The paper reframes generative AI integration in medical training as a cognitive and behavioral decision problem centered on the moment learners choose how to engage with AI. It argues that educators should teach and assess this competency while treating the framework as a testable proposal.

  • The paper addresses how to integrate generative AI into clinical training without undermining clinical reasoning.
  • SCAN is presented as a cognitive and behavioral solution that should earn its claims through evidence and be tested rather than asserted.
  • SCAN focuses on the moment before a learner opens an AI tool and decides what engagement suits the task and developmental stage.
  • The paper identifies this decision point as where clinical reasoning may be formed or deformed, shifting attention toward learner understanding rather than only monitoring and restriction.
  • The proposed shift treats AI-related cognitive judgment as a foundational competency for clinical practice in the AI era, alongside history-taking and physical examination skills.
Loading 2608.28373v1…