Source-linked AI summary

From Few-Shot Segmentation to Clinician-in-the-Loop Medical Image Analysis

Yazhou Zhu

arXiv:2609.10001v1cs.CV

TL;DR

Few-shot medical image segmentation fixes task-defining evidence before inference, leaving rare, shifted, ambiguous, or degraded cases insufficiently represented and clinically risky. This Perspective reframes it as a bounded sequential clinician-model decision problem, where support and interaction budgets are distinct and feedback is requested, used, or deferred according to expected value and safety controls. Its supported conclusion is that clinical reliability requires governed attention allocation and evaluation beyond segmentation quality alone.

  • Problem

    FSMIS assumes a small fixed support set contains enough evidence for each query, an assumption challenged by acquisition shift, atypical pathology, ambiguous boundaries, and poor image quality.

  • Method

    The Perspective formulates FSMIS as a sequential accept-query-defer system with separate support and interaction budgets, response-conditioned querying, bounded adaptation, and safety gates.

  • Results

    The Perspective synthesizes an integration framework linking few-shot, cross-domain, interactive, selective, and adaptive segmentation under explicit clinical-risk and expert-attention evaluation.

  • Takeaways & Limitations

    Clinician attention should be allocated only when expected feedback benefit exceeds its burden and can reduce clinically relevant risk under governed adaptation.

  • Takeaways & Limitations

    The framework is normative and does not itself establish prospective clinical effectiveness; segmentation quality is not identical to clinical benefit.

Abstract

from arXiv · show

Few-shot medical image segmentation (FSMIS) seeks to delineate unseen structures from a small support set, but its standard formulation fixes task-defining evidence before inference. This assumption is fragile when query cases exhibit acquisition shift, atypical pathology, ambiguous boundaries, or poor image quality. Prototype learning, cross-domain matching, interactive segmentation, uncertainty estimation, test-time adaptation, and promptable foundation models address parts of this problem, yet have not been jointly evaluated under a common model of expert attention and clinical risk. This Perspective reframes FSMIS as a sequential clinician-model decision problem with a static support budget $K$ and a distinct interaction budget $B$. At each step, a system accepts the current segmentation, requests feedback, or defers to full expert review. Queries vary in location and modality and are selected by response-conditioned net expected value of information; clinician-provided feedback informs bounded adaptation only after prespecified provenance, consistency, and safety gates. The framework separates distributional atypicality from predicted clinical failure and treats clinician responses as informative but fallible observations. We synthesize the transition from few-shot and cross-domain segmentation to interactive and selective adaptation, delineate the integration gap, and define four research directions with falsifiable hypotheses. Evaluation spans external-domain calibration, quality-effort trade-offs, reader studies, and prospective workflow assessment. The central claim is not that interaction alone resolves domain shift, but that scarce expert attention should be allocated only when it is expected to reduce clinically relevant risk.

1 Introduction

FSMIS is motivated by expensive expert annotation but becomes fragile when fixed support evidence cannot resolve rare, shifted, ambiguous, or degraded query cases. The Perspective therefore frames analysis as a bounded sequential clinician-model system that spends attention when expected clinical risk reduction justifies its cost.

  • 1 Introduction: Fixed support evidence can fail on rare disease, atypical anatomy, postoperative change, severe artifact, unfamiliar acquisition, or ambiguous boundaries.
  • 1 Introduction: Clinical attention is costly and fallible because queries interrupt work, modalities impose unequal burdens, experts disagree, and mistaken corrections can amplify risk.
  • 1 Introduction: The proposed system links few-shot segmentation, cross-domain generalization, interaction, active acquisition, selective prediction, and adaptation through accept-query-defer decisions.
  • 1 Introduction: Clinical difficulty reflects interacting data, coverage, and attention scarcity, so improving average Dice under one-shot protocols addresses only part of the problem.
  • 1 Introduction: The framework separates static support budget K from interaction budget B, preventing few-shot learning from being conflated with ordinary interactive segmentation.
  • 1 Introduction: Reliable deployment should allocate limited clinician attention only when expected reduction in clinical risk justifies its cost, rather than relying on confidence or model scale alone.

3 From Fixed Support to Cross-Domain Robustness

Research progresses from representing sparse support evidence toward robust matching across domains, but these branches remain complementary rather than strictly sequential. Stronger representations and domain-robust matching still cannot recover missing case-specific anatomy or establish whether clinician feedback is worth requesting.

  • 3.1 The canonical few-shot assumption: The canonical few-shot pipeline compresses a heterogeneous target structure into a small, potentially biased support description.
  • 3.2 Representing heterogeneous support evidence: Regional decomposition, relation-aware refinement, and episode-conditioned filtering treat support elements as unequally relevant and case-dependent.
  • 3.3 Cross-domain FSMIS: Cross-domain FSMIS adds transfer to unfamiliar target domains, where source-trained encoders and similarity functions may provide poorly aligned representations.
  • 3.4 Integration gap: The field's branches improve sparse-evidence representation, shifted-domain matching, or promptable interfaces, but they are complementary rather than a strict chronology.
  • 3.4 Integration gap: Existing methods do not establish whether missing case-specific evidence is worth requesting from a clinician.

4 Adjacent Paradigms and the Integration Gap

Existing methods provide interaction, adaptation, selection, and deferral capabilities, but the field lacks a jointly evaluated policy for deciding when and how clinical information should be acquired and controlled.

  • 4.1 Interactive segmentation is necessary but not sufficient: Interactive segmentation can correct current masks, but evidence from target-mask studies supports annotation efficiency rather than independent patient-level boundary judgment.ScribblePrompt improved annotation efficiency relative to SAM ViT-B in a study of 16 academic-hospital imaging researchers shown target masks.
  • 4.1 Interactive segmentation is necessary but not sufficient: Related systems explore automatic click prediction, reinforcement-guided interaction, sequential memory, simulated test-time optimization, and cross-image correction-driven adaptation.These methods extend interaction beyond purely reactive correction.
  • 4.1 Interactive segmentation is necessary but not sufficient: Human correction, adaptation, context accumulation, selection, deferral, and preference alignment are established separately rather than as a novel unified contribution.The evidence spans multiple systems with differing publication and evaluation levels.
  • 4.2 Active learning, in-context learning, and deferral: Active learning formalizes acquisition through information gain or epistemic uncertainty, while medical systems add representativeness, consistency, diversity, or simplified interactive labels.Interactive segmentation usually targets the current mask, whereas active learning targets information useful for future model improvement.
  • 4.2 Active learning, in-context learning, and deferral: The central gap is a controlled policy for sequential model-initiated clinical querying under few-shot, cross-domain, and case-level OOD conditions, jointly evaluating risk, effort, adaptation, and deferral.Queries should specify location, modality, and clinical question, update an auditable episode state, and stop when further interaction no longer dominates acceptance or deferral.

5 Decision-Theoretic Formulation

The formulation models few-shot medical image analysis as a bounded accept-query-defer process in which clinician feedback can update the episode, but only under cost, risk, retention, and safety constraints.

  • 5.1 Episode state, actions, and bounded update: Each episode starts with a query image, K support image-mask pairs, an initial model, and a separate interaction budget B.The state also tracks interaction history, the current representation, remaining budget, and local-to-case-level risk evidence.
  • 5.1 Episode state, actions, and bounded update: The controller chooses accept, query, or defer, with queries specifying a location, modality, and clinical question and responses potentially updating prompts, prototypes, memory, features, or restricted parameters.Updates are applied through a bounded transition operator using the clinician response.
  • 5.1 Episode state, actions, and bounded update: Acceptance incurs clinically weighted release loss, whereas deferral incurs full-review burden and residual expert or joint-result loss, so neither terminal action is free or perfect.The cost ledger fixes hard resource use and soft burden before evaluation to avoid double counting.
  • 5.2 Response-conditioned value of clinical feedback: A query is selected by response-conditioned net expected value of information, incorporating expected post-response stopping cost, query burden, and a retention penalty for degrading protected capabilities.Uncertainty alone is insufficient; a plausible response must change the mask, lower residual risk, or support deferral.
  • 5.2 Response-conditioned value of clinical feedback: Greedy control queries only when the best action has positive value, while non-myopic control accounts for how one response changes the value of later questions.Clicks, boundary corrections, reference cases, and text differ in bandwidth and semantic content, with the latter channels requiring empirical testing.
  • 5.3 Accept, query, or defer: Selective interaction should query when a compact intervention can resolve elevated risk, accept only under externally validated thresholds, and defer outside the validated adaptation envelope or when ambiguity is irreducible.Deferral also applies when support evidence is inadequate or another interaction has low expected value.
  • 5.3 Accept, query, or defer: Clinician feedback is informative but fallible, so provenance, modality, timing, confidence, disagreement, and plausible masks must be retained rather than treating the clinician as an infallible oracle.Figure 2 assigns interaction and stopping to a separate controller and requires independent reassessment of each adaptation update against clinical risk and protected capabilities.

6 Design Principles and Research Questions

The framework organizes clinician-in-the-loop FSMIS around calibrated risk, semantically distinct feedback, bounded adaptation, and governed personalization. It treats expert responses as valuable but fallible evidence that must be controlled by safety and provenance requirements.

  • 6.1 Calibrated recognition: Three linked uncertainty levels should localize errors, assess structure quality, and govern case-level acceptance or deferral.Accuracy and confidence must remain separate because overconfidence and distribution shift can degrade calibration.
  • 6.1 Calibrated recognition: OOD scores should be combined with predicted mask quality and clinical consequence because atypicality alone does not predict segmentation failure.Controlled factorial shifts and incremental decision value are proposed for testing operational proxies.
  • 6.2 Cost-sensitive multimodal feedback: Points, boxes, scribbles, masks, references, and text should be evaluated as distinct supervision channels whose value and human cost remain unresolved.The key open problem is response-conditioned selection among channels coupled to accept, query, and defer decisions.
  • 6.3 Bounded rapid adaptation: Adaptation should begin with reversible episode-local updates protected by source anchors, trust regions, snapshots, frozen monitors, and rollback gates.Persistent memory requires separate governance, provenance, privacy controls, and protected-task testing before affecting later patients.
  • 6.4 Selective and personalized teaming: Clinician response time and correction style cannot directly measure competence because case difficulty, interface familiarity, fatigue, and prior allocation confound behavior.Personalization therefore requires consent, interpretable features, minimum evidence, workload audits, reversibility, and unconditional override.

7 Foundation Models as an Enabling Substrate

Foundation models are best positioned as reusable representation and interaction substrates rather than complete safety solutions. Their interfaces can support prompting, but calibration, stopping, patient-level risk, and adaptation governance remain unresolved.

  • 7 Foundation Models as an Enabling Substrate: Foundation models can provide reusable features and interaction interfaces, but they do not resolve calibration, stopping, or adaptation safety.The framework deliberately leaves the backbone unspecified.
  • 7 Foundation Models as an Enabling Substrate: Promptable systems support points, boxes, and masks without task-specific retraining, yet medical performance varies substantially with task and prompt type.Medical adaptations such as MedSAM primarily use box prompting, leaving broader interaction questions open.
  • 7 Foundation Models as an Enabling Substrate: MAUP bridges multi-center representations and uncertainty-aware prompt selection to frozen SAM inference without parameter updates.Its uncertainty ranking does not establish patient-level risk calibration or justified abstention.
  • 7 Foundation Models as an Enabling Substrate: A modular architecture should separate foundation encoding, episode memory, task prediction, risk estimation, action control, and constrained adaptation.This separation allows individual safety claims to be tested independently.

8 Experimental and Clinical Validation

Validation must assess the complete clinician-model trajectory across shifts, risk, effort, adaptation, and workflow rather than final-mask quality alone. The proposed ladder combines calibrated external evaluation, realistic interaction studies, and prospective controlled assessment.

  • 8.1 External-domain evaluation: Evaluation should span controlled and natural acquisition shifts, unseen structures, rare lesions, artifacts, weak boundaries, and poor image quality across multi-institutional MRI and CT.Stratification should include site, scanner, sequence, target size, pathology, and difficulty.
  • 8.1 External-domain evaluation: Patient- or structure-level probability of complete-mask failure should be calibrated against a blinded acceptability criterion using reliability curves, Brier scores, and log scores.Voxel-level expected calibration error is secondary because it can conceal case-level miscalibration.
  • 8.2 Evaluation of the coupled system: Interaction outcomes should be measured against elapsed clinician time, corrected area or slices, actions, latency, time to acceptable mask, and quality-effort trade-offs.False queries require randomized query/no-query assignment or validated offline policy evaluation because safe no-query outcomes are counterfactual.
  • 8.2 Evaluation of the coupled system: Adaptation must be evaluated after every response for gain, calibration, monotonicity, protected-task retention, within-volume transfer, error propagation, and rollback frequency.Increasing confidence while increasing clinically weighted error is defined as a safety failure.
  • 8.3 From robot users to clinicians: Simulated users enable scalable ablations but cannot establish clinical interaction efficiency because they omit search, navigation, cognitive switching, fatigue, and misleading-output recovery.User protocols and interface design can alter conclusions.
  • 8.3 From robot users to clinicians: Validation should progress from multiple robot policies and prompt perturbations through disjoint manual traces, randomized reader studies, silent prospective evaluation, and limited workflow studies.Stopping and rollback criteria should accompany the final workflow stage.

9 Falsifiable Hypotheses

The paper proposes five falsifiable hypotheses covering dynamic querying, multilevel risk, task-representation updates, bounded adaptation, and governed personalization. Each hypothesis specifies a comparison and concrete failure signals involving effort, calibration, transfer, safety, or equity.

  • 9 Falsifiable Hypotheses: Value-guided dynamic queries should achieve acceptable masks more often than random, fixed-click, or passive-correction policies at matched clinician time.The hypothesis is falsified if the gain disappears after elapsed-time costs or with real users.
  • 9 Falsifiable Hypotheses: Combining local error localization, structure quality, and case atypicality should improve accept/query/defer decisions over voxel entropy alone.Failure is defined by absent held-out-institution selective-risk improvement or failed patient-level calibration after frozen thresholds.
  • 9 Falsifiable Hypotheses: Updating support or prompt representations should improve unedited slices within the same volume more than locally overwriting corrected pixels at equal risk.The hypothesis is falsified by negligible transfer, propagated error, or excessive protected-task loss.
  • 9 Falsifiable Hypotheses: Trust regions, anchors, frozen gates, and rollback should retain most target-domain gain while reducing catastrophic accumulation and forgetting.The hypothesis fails if constraints suppress useful adaptation without improving worst-case or longitudinal safety.
  • 9 Falsifiable Hypotheses: With consent and sufficient observations, governed personalization should reduce time and unnecessary queries without increasing error or over-reliance.It is falsified by robot-only benefits, inequitable task allocation, or behavior that cannot be explained, audited, and overridden.

10 Levels of Evidence Required for Translation

Translation requires staged evidence, beginning with benchmarked decision behavior, then testing interaction and adaptation, and finally assessing prospective clinician-model workflows.

  • Level I: benchmark and decision evidence: A unified benchmark protocol should establish calibration, accept/query/defer baselines, clinically weighted edge cases, and interaction-cost models before parameter adaptation.An auditable controller can be evaluated before any adaptation is permitted.
  • Level II: interaction and adaptation evidence: Comparative interaction studies must test multiple feedback modalities and bounded updates for risk reduction, transfer beyond edited regions, rollback, and clinical acceptability.Clinical collaboration is needed because interaction cost and acceptability cannot be inferred from benchmark masks alone.
  • Level III: prospective clinician-model evidence: Prospective studies should freeze thresholds and report model, clinician, and joint-system performance while measuring shift, failure prevalence, trigger rates, and workload.Silent studies should precede limited workflow studies, with longitudinal cross-case learning separately consented and governed.

11 Limitations, Governance, and Scope

The framework remains constrained by unreliable uncertainty under shift, fallible feedback, human-factor risks, governance requirements, response-model misspecification, and the gap between segmentation quality and clinical benefit.

  • Uncertainty and shift: Uncertainty can become unreliable under severe shift, while detected shift can block adaptation or trigger review but cannot prove error.Independent signals and conditional calibration are proposed mitigations, not guarantees.
  • Feedback reliability: Clinician feedback is fallible because corrections may reflect ambiguity, fatigue, or preference, requiring provenance, repeated evidence, confidence, and multi-reader models before persistence.The framework treats responses as informative observations that still require safeguards before entering memory.
  • Human factors: Human factors require measuring override behavior, workload, and recovery from imperfect suggestions because plausible masks and repeated low-value queries can create safety risks.Erroneous AI advice degrading physician decisions is identified as a primary safety endpoint.
  • Governance and privacy: Persistent site- or clinician-specific memory requires access control, audit trails, deletion mechanisms, and separate consent and governance for longitudinal cross-case learning.Per-case updates should remain local and ephemeral unless persistence has governance and measurable transfer support.
  • Response modeling: A misspecified response distribution can produce confidently low-value queries, so sensitivity analysis and prospective re-estimation are required.The response model must represent disagreement, no response, latency, and interface effects.
  • Clinical scope: Segmentation quality does not establish clinical benefit, so studies must prespecify the use case, tolerance, failure cost, and downstream endpoint.Retrospective gains, aggregate metrics, or foundation-model scale cannot establish prospective clinical safety; the framework is normative rather than effectiveness evidence.

12 Conclusion

Clinician-in-the-loop FSMIS asks how to obtain the smallest, safest, and most informative dialogue for each clinical case. The framework therefore evaluates the trajectory of a joint system, combining querying, calibrated decisions, bounded adaptation, and governed memory rather than relying on confidence or model scale alone.

  • 12 Conclusion: Clinician-in-the-loop FSMIS shifts the question from learning from few labels to finding the smallest, safest, and most informative dialogue for the current case.This extends conventional and cross-domain FSMIS with an explicit evidence-acquisition problem.
  • 12 Conclusion: The framework changes the unit of analysis from a final mask to a joint clinician-model trajectory using active querying, calibrated acceptance and deferral, bounded adaptation, and governed memory.Foundation models provide a reusable interface, not a safety argument.
Loading 2609.10001v1…