Source-linked AI summary

Human-Centered Tools for Coping with Imperfect Algorithms during Medical Decision-Making

Carrie J. Cai, Emily Reif, Narayan Hegde, Jason Hipp, Been Kim, Daniel Smilkov, Martin Wattenberg, Fernanda Viegas, Greg S. Corrado, Martin C. Stumpe, Michael Terry

arXiv:1902.02960v1cs.HCcs.CY

TL;DR

Medical image retrieval algorithms may not capture the similarity features pathologists need in every diagnostic case, limiting relevance and trust. The paper develops interactive refinement tools for guiding deep-learning image search, and two evaluations found higher information utility and trust without reduced diagnostic accuracy. Users also repurposed the tools to understand the algorithm and test diagnostic hypotheses.

  • Problem

    Similarity algorithms may return images that are visually similar but clinically irrelevant to pathologists’ changing diagnostic needs, undermining trust in opaque systems.

  • Method

    The paper identifies pathologists’ needs and develops SMILY with region, example, and concept-based interactive refinement tools for on-the-fly search guidance.

  • Results

    Refinement tools increased diagnostic information utility and user trust, were preferred over a traditional interface, and produced no significant effect on accuracy (p=0.55) or confidence (p=0.7).

  • Takeaways & Limitations

    Refinement gave experts agency to hypothesis-test, understand ML behavior, and apply domain knowledge while leveraging automated image retrieval.

  • Takeaways & Limitations

    Iterative refinement may encourage confirmation bias when users search mainly for evidence consistent with existing beliefs.

Abstract

from arXiv · show

Machine learning (ML) is increasingly being used in image retrieval systems for medical decision making. One application of ML is to retrieve visually similar medical images from past patients (e.g. tissue from biopsies) to reference when making a medical decision with a new patient. However, no algorithm can perfectly capture an expert's ideal notion of similarity for every case: an image that is algorithmically determined to be similar may not be medically relevant to a doctor's specific diagnostic needs. In this paper, we identified the needs of pathologists when searching for similar images retrieved using a deep learning algorithm, and developed tools that empower users to cope with the search algorithm on-the-fly, communicating what types of similarity are most important at different moments in time. In two evaluations with pathologists, we found that these refinement tools increased the diagnostic utility of images found and increased user trust in the algorithm. The tools were preferred over a traditional interface, without a loss in diagnostic accuracy. We also observed that users adopted new strategies when using refinement tools, re-purposing them to test and understand the underlying algorithm and to disambiguate ML errors from their own errors. Taken together, these findings inform future human-ML collaborative systems for expert decision-making.

1 INTRODUCTION

Medical image retrieval can support pathologists, but imperfect similarity models may return clinically irrelevant results and erode trust. This paper develops interactive refinement tools that let users guide similarity according to changing diagnostic needs.

  • Pathologists use visually similar images from previously diagnosed patients to support differential diagnosis and clinical decision-making.
  • Because relevant visual features vary across cases and moments, algorithmically similar images may not match a pathologist’s immediate diagnostic needs.
  • Clinically irrelevant results can reduce physician trust, while opaque ML systems can worsen the interactive experience.
  • The paper proposes interactive refinement techniques that let pathologists emphasize important image characteristics during search.
  • SMILY includes refine-by-region, refine-by-example, and refine-by-concept tools for guiding image retrieval.
  • Across two evaluations, refinement increased information utility and trust, supported algorithm understanding and diagnosis probing, and preserved diagnostic accuracy.

Content-based Image Retrieval

CBIR systems retrieve images by content, but extracted visual content can diverge from users’ semantic interpretations. This work exposes deep neural network representations and concept directions to support interactive refinement.

  • Content-based image retrieval organizes image databases by content and can retrieve medically similar images for clinical decision support.
  • The semantic gap arises when image content extracted by a CBIR system does not correspond to the user’s semantic interpretation.
  • Prior approaches reduce this gap through engineered visual features or human relevance feedback, while this paper emphasizes the surrounding user experience.
  • Deep neural networks produce embeddings whose relative positions can encode high-level concepts and support similar-image search.
  • Concept Activation Vectors are directions learned with linear classifiers, and this work exposes clinical CAVs through refine-by-concept sliders.

3 USER NEEDS

Pathologists generate and compare diagnostic hypotheses while seeking similar images that reveal distinctions across diagnoses. Their needs include controlling which visual features matter, including features that are localized, pervasive, or relevant only at particular moments.

  • Needs During Similar Image Search: In iterative design sessions with three pathologists, researchers used paper prototypes, interviews, think-alouds, and functional prototypes to identify search needs.
  • Needs During Clinical Decision-Making: Pathologists generate hypotheses, compare evidence, and determine the most likely diagnosis during differential diagnosis.
  • Needs During Clinical Decision-Making: Diagnostic decisions can be difficult because cancers in different grades share visual characteristics and diagnosis determines subsequent treatment.
  • Needs During Clinical Decision-Making: Pathologists seek similar images across diagnostically distinct categories to maintain a safety net against missed diagnoses.
  • Needs During Similar Image Search: Users wanted to de-emphasize irrelevant features and emphasize clinically relevant ones in search results.
  • Needs During Similar Image Search: Feature importance can vary over time, requiring users to shift attention among localized features, overall architectures, and visual patterns.

4 USER INTERFACE AND SYSTEM DESIGN

SMILY retrieves nearest neighbors in a deep neural network embedding space and provides region, example, and concept-based refinement tools. Concept sliders use learned clinical directions to shift queries, while formative findings exposed limitations involving opposing concepts and confounding variables.

  • System Architecture: SMILY uses a pre-trained deep neural network to embed query images and retrieve nearest neighbors from previously diagnosed pathology cases.
  • System Architecture: The system presents approximately 15 nearest neighbors per page to balance result variety with avoiding user overload.
  • Refinement Tools: Refinement mechanisms guide retrieval toward meaningful directions in the DNN’s high-dimensional embedding space without altering the deep neural network itself.
  • Refinement Tools: Refine-by-region replaces the query with a user-selected crop, while refine-by-example averages selected result embeddings to retrieve more like them.
  • Refinement Tools: Refine-by-concept sliders let users request more or less of a clinical concept, including concepts absent from the original query, to test diagnostic hypotheses.
  • Refinement Tools: Concept sliders use linear classifiers to learn CAV directions from labeled positive and negative examples, with 20 labels producing cosine similarity of ~0.9 to all-label CAVs.
  • Refinement Tools: Users sometimes expected a negative CAV direction to represent the opposing concept, motivating relative CAVs trained with opposing concepts as negatives.

5 TOOL EVALUATION STUDY

The Tool Evaluation Study tested whether refinement mechanisms changed retrieved images in intended clinical directions. Across refine-by-example, refine-by-concept, and refine-by-region evaluations, refinement generally increased the presence of desired concepts or features.

  • Study design: The study first validated whether refinement mechanisms updated search results as intended, then evaluated their effects on user experience and search practices.Pathologists rated paired results with and without refinement, with conditions randomized and concealed.
  • Refine-by-region Evaluation: Refine-by-region targeted small glandular structures that could be overshadowed by physically prominent but irrelevant image features.The evaluation used ten images in which the key glandular component occupied approximately 25% of the image.
  • Refine-by-example Evaluation: Refine-by-example evaluated four representative concepts spanning structural components, artifacts, and morphological patterns using 40 query images.A pathologist selected ten images per concept, another applied refinements, and two blinded pathologists rated concept presence.
  • Refine-by-example Evaluation: Refined results contained more of the desired concept than unrefined results (µ = 4.6 vs. 2.8, p < 0.0005, F = 108.7).Refinement produced greater concept presence in 82% of cases, tied in 11%, and produced less in 7%.
  • Refine-by-concept Evaluation: CAV-based refinement increased desired concept presence from µ = 2.6 to µ = 5.4 (p < 0.0005, F = 584.6), with greater presence in 99% of cases.It also produced fewer unintended consequences than refine-by-example.
  • Refine-by-concept Evaluation: An independent pathologist identified most CAV concepts correctly, although Fused Glands also elicited biologically correlated and other confounding concepts.The correlated concepts were judged less surprising than concepts not biologically related to Fused Glands.

6 USER STUDY

The User Study examined whether SMILY improved diagnostic information utility, workload, and trust relative to a conventional n-best list interface. It also investigated how pathologists used refinement tools and the trade-offs among tool types.

  • Study motivation: The study followed the Tool Evaluation Study by examining how refinement mechanisms affected end-user experience and actual search practices.The authors considered whether added interaction complexity and workload would outweigh the tools' utility.
  • Research questions: The first research question addressed SMILY's effects on diagnostic utility, workload, and trust compared with a conventional n-best list interface.The second addressed pathologists' refinement practices and trade-offs among refinement tools.
  • Research questions: The study also asked how pathologists used refinement tools during search and decision-making, including trade-offs between different refinement mechanisms.This broadened evaluation beyond whether the tools technically changed retrieved results.

Measures

The user study measured diagnostic utility, mental support, workload, trust, and intended clinical use using ratings and behavioral records from twelve pathologists performing prostate-image tasks.

  • Outcome measures: All study outcome items were rated on a 7-point scale across utility, workload, and attitudes toward the system.The measures included diagnostic utility, mental support for decision-making, effort, frustration, and trust.
  • Outcome measures: Trust was assessed through Mayer's capability and benevolence dimensions, which were selected because of their use in prior studies and HCI work.Participants answered Likert-scale questions about both dimensions.
  • Interface labeling: The two systems were labeled Version N and Version B during the study and later described as SMILY and the conventional interface.Counterbalanced labels were used to avoid biasing participants.
  • Participants and procedure: Twelve pathologists participated, with 1–20 years of post-residency pathology experience (µ = 9.8).Each completed a tutorial, analyzed six prostate images, and then completed a questionnaire and interview.
  • Participants and procedure: The six cases comprised two borderline, two asymmetric, and two diverse images selected from previously contested cases.The cases were chosen to span difficult diagnostic scenarios and were tested in randomized conditions and categories.

7 USER STUDY RESULTS

Compared with the conventional interface, SMILY provided more diagnostically useful information with less effort and increased trust, mental support, and intended clinical use. Users preferred SMILY, while diagnostic accuracy and confidence showed no significant interface effect.

  • User study outcomes: Diagnostic utility was higher with SMILY (µ = 4.7) than with the conventional interface (µ = 3.7), with a significant interface effect (p = 0.025, t = 2.3).The analysis used interface and image type as fixed effects and participant as a random effect.
  • User study outcomes: Effort was lower with SMILY (µ = 2.8) than with the conventional interface (µ = 3.3; p = 0.034, t = -2.2), with no difference in frustration.Image type also affected effort, with diverse images requiring more effort than borderline images.
  • User study outcomes: Trust ratings favored SMILY for capability (µ = 6 vs. 4.7, p = 0.01, t = 3.08) and benevolence (µ = 5.8 vs. 2.6, p < 0.001, t = 6.04).These were post-study questionnaire ratings.
  • Search practices: Refinement tools were interleaved and repurposed for coarse-to-fine searching, backtracking, testing hypotheses, and understanding algorithm behavior.Participants often used region refinement before example- and concept-based adjustments.
  • Preference and practice: All 12 participants preferred SMILY over the conventional interface, with 7 preferring it totally, 4 much more, and 1 slightly more.Participants also viewed SMILY as more practically useful in day-to-day work.
  • Accuracy and confidence: SMILY produced no significant interface effect on diagnostic accuracy (p = 0.55) or confidence (p = 0.7).Detecting subtle differences would require a full clinical trial, which was outside the study's scope.

Refine-by-Example

Pathologists used refinement tools in complementary ways to steer image retrieval toward diagnostically relevant similarity. Refine-by-concept was faster and supported systematic exploration, while refine-by-example could become slow and ambiguous when result changes were subtle.

  • Refine-by-example let pathologists mark useful results and retrieve more visually similar examples, supporting perceptual, pattern-based search.
  • Users struggled to judge refine-by-example updates when changes were subtle or selected examples contained confounding features.Users provided a mean of 2 examples per update.
  • Refine-by-concept let users adjust the presence of clinical concepts, helping them reason systematically about diagnostic features.
  • Refine-by-concept sliders helped users examine continuous diagnostic distinctions, including the boundary between cancer grades 3 and 4.
  • In the conventional interface, users scanned multiple result pages, whereas with SMILY they usually refined from the first page.Users traversed a mean of 2.6 pages conventionally versus 0.4 with SMILY.

9 DECISION MAKING AND COPING WITH BLACK-BOX ML

Pathologists used iterative refinement not only to navigate diagnostic hypotheses but also to reduce cognitive complexity. The tools helped them track confidence, generate ideas, and manage intermediate reasoning states.

  • Tracking the Likelihood of a Decision Hypothesis: Increasing numbers of visually similar images in a hypothesized category reassured users, while divergent updates signaled that they might be on the wrong path.
  • Generating New Ideas: Iterative refinement helped pathologists reflect on how they reached a conclusion and notice unexplored diagnostic features.Concept sliders could surface possibilities that had not otherwise come to mind.
  • Generating New Ideas: Refinement tools could help users think too deeply, potentially sending them down tangential diagnostic paths.
  • Reducing Problem Complexity: Reference images containing diverse feature mixtures were cognitively difficult because users had to analyze features separately while considering their interactions.
  • Reducing Problem Complexity: Refinement could focus users on one component at a time, potentially offloading intermediate state from working memory to the interface.

Refinement Strategies for Coping with ML

Pathologists used refinement to cope with mismatches between their reasoning and the opaque algorithm. They adjusted concepts, tested hypotheses, and investigated whether surprising results reflected ML or human error.

  • Attempting to Narrow the Semantic Gap: Users narrowed the semantic gap by emphasizing relevant medical concepts and reducing irrelevant ones when SMILY’s similarity differed from their own.Unexpected or irrelevant results degraded trust, while refinement adapted retrieval to users’ interests.
  • Developing a Mental Model of the ML: Unexpected results prompted users to form theories about which image features the algorithm weighted and to test those theories through refinement.
  • Disambiguating ML Errors from Self Errors: Refinement helped pathologists distinguish algorithmic noise from features they had themselves overlooked.They removed suspected confounding variables and repeated the search to test whether ML caused the error.
  • Fear of Over-influencing the Algorithm: Users sometimes worried that extreme slider settings could over-influence the algorithm and return examples beyond the intended concept.
  • Developing a Mental Model of the ML: Users wanted control over visual characteristics while preserving diagnostic categories during refinement.

10 DISCUSSION

The discussion frames refinement as a lightweight way to customize imperfect medical image retrieval while preserving expert agency. It also identifies transparency, disambiguation, and confirmation bias as important design considerations.

  • Refinement tools let users customize retrieval for changing case-specific needs without retraining the original model.
  • Users repurposed refinement to track diagnoses and build mental models of ML, contributing to greater trust.
  • SMILY’s refinement tools helped experts begin disambiguating algorithmic errors from their own oversights during uncertain decisions.
  • Refinement may introduce confirmation bias if users search only for evidence consistent with existing beliefs or mistake apparent improvement for real improvement.The authors report no evidence of deteriorated decision-making with SMILY and suggest exposing unexplored refinement paths.
  • Interactive refinement could make users active interpreters of opaque algorithms by enabling hypothesis-testing of their intuitions.
  • The paper concludes that refinement can help ML augment rather than replace expert intelligence in critical decision-making.
Loading 1902.02960v1…