Source-linked AI summary

"Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction

Sunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong, Andrés Monroy-Hernández

arXiv:2210.03735v2cs.HCcs.AIcs.CVcs.CY

TL;DR

End-users’ explainability needs and behaviors remain understudied in real-world AI contexts, limiting understanding of how XAI can support human-AI interaction. The authors conducted a mixed-methods study with 20 Merlin users and found that participants preferred practically useful, part-based explanations that support collaboration beyond merely understanding AI outputs.

  • Problem

    End-users’ explainability needs and behaviors around XAI explanations remain insufficiently studied in real-world contexts.

  • Method

    The authors conducted a mixed-methods study with 20 Merlin users spanning diverse AI and birding backgrounds, using interviews, a survey, and mock-up XAI explanations.

  • Results

    Participants preferred practically useful information and part-based explanations, intending to use XAI for trust calibration, skill improvement, better inputs, and developer feedback.

  • Takeaways & Limitations

    XAI design should account for end-users’ collaboration-oriented needs and uses beyond understanding AI outputs.

  • Takeaways & Limitations

    The findings may not generalize beyond the Merlin app, and some background subgroups had relatively few participants.

Abstract

from arXiv · show

Despite the proliferation of explainable AI (XAI) methods, little is understood about end-users' explainability needs and behaviors around XAI explanations. To address this gap and contribute to understanding how explainability can support human-AI interaction, we conducted a mixed-methods study with 20 end-users of a real-world AI application, the Merlin bird identification app, and inquired about their XAI needs, uses, and perceptions. We found that participants desire practically useful information that can improve their collaboration with the AI, more so than technical system details. Relatedly, participants intended to use XAI explanations for various purposes beyond understanding the AI's outputs: calibrating trust, improving their task skills, changing their behavior to supply better inputs to the AI, and giving constructive feedback to developers. Finally, among existing XAI approaches, participants preferred part-based explanations that resemble human reasoning and explanations. We discuss the implications of our findings and provide recommendations for future XAI design.

1 INTRODUCTION

This study connects XAI research with end-users of a real-world AI application to examine their explainability needs, intended uses, and perceptions. It finds that participants valued actionable collaboration support and preferred part-based explanations resembling human reasoning.

  • The study addresses limited understanding of end-users’ explainability needs and behaviors around XAI explanations.Existing XAI methods have often been developed without fully embracing the spectrum of end-user needs.
  • 20 Merlin users spanning low-to-high AI and domain backgrounds participated in a mixed-methods study of XAI needs, uses, and perceptions.The study used interviews, a survey, an interactive feedback session, and mock-ups of four XAI approaches.
  • Participants unanimously wanted practically useful information that could improve collaboration with the AI, while curiosity about technical details varied by AI background and bird interest.Participants with higher AI backgrounds or stronger bird interest generally expressed greater curiosity about system details.
  • Participants intended to use explanations to calibrate trust, improve task skills, provide better inputs, and give constructive feedback to developers.These intended uses extend beyond understanding the AI’s outputs.
  • Participants preferred part-based concept and prototype explanations because they resembled human reasoning and explanations.These approaches were considered most useful for the purposes participants identified.

2 RELATED WORK

Related work has increasingly shifted from algorithm-centered XAI toward human-centered study of people’s needs in context. This paper extends that direction through an everyday, real-world application and qualitative investigation of how users may collaborate with AI.

  • The paper frames its contribution within human-centered XAI and human-AI collaboration, focusing on users’ needs, goals, contexts, and shared work with AI.The collaboration framing emerged from participants’ descriptions of a two-way exchange with Merlin.
  • XAI research has often focused on developers and algorithmic transparency rather than end-users’ needs in deployment contexts.The paper identifies this as a critical gap because end-users may require different forms of explainability.
  • Prior application-specific studies offered rich insights but often examined hypothetical or prototype systems, leaving real-world end-user needs less studied.The paper addresses this open question through an in-context study of a deployed application.
  • Unlike prior work centered on high-stakes medical applications, this study examines an ordinary application used by people with diverse AI and domain knowledge.It also explicitly studies differences associated with participants’ AI and domain backgrounds.
  • The study differs from lab experiments that prescribe how participants should use explanations by using a qualitative descriptive design to investigate open-ended collaboration.Earlier experiments commonly used simple tasks and recruited participants from Amazon Mechanical Turk.

3 STUDY APPLICATION: MERLIN BIRD IDENTIFICATION APP

Merlin provides a grounded setting for studying explainability with diverse end-users in ordinary bird-identification activities. Its computer-vision focus also aligns with substantial prior XAI research on bird classification.

  • Merlin is a mobile app with over a million downloads that uses computer vision to identify birds in user-input photos and audio recordings.It is used outdoors by people with diverse birding and AI knowledge.
  • Users employ Merlin for bird identification in everyday scenarios, making it a real-world context for studying end-user interaction with AI.The app was selected because it fit the study’s requirements for ordinary use and diverse knowledge backgrounds.
  • The app provides a grounded context where users’ experience can be augmented with mock-up XAI explanations.The study used this setting because many computer-vision XAI methods have been developed and evaluated on bird image classification.

4 METHODS

The study recruited a diverse sample of Merlin users and examined their explainability needs, uses, and perceptions through interviews, surveys, and interactive feedback on mocked-up XAI approaches.

  • 4.1 Participant recruitment and selection: 20 Merlin users were selectively recruited to maximize diversity in bird-domain and AI backgrounds.Participants used Merlin’s Photo ID and/or Sound ID features, with varied usage frequencies.
  • 4.2 Study instrument: The one-hour interviews combined contextual questions, an XAI-needs survey, and interactive evaluations of four explanation approaches.The survey covered ten categories of questions, while the interactive session used real identifications paired with mock-up explanations.
  • 4.2 Study instrument: The study evaluated heatmap-, example-, concept-, and prototype-based explanations using three real Merlin Photo ID outputs.The outputs included one correct identification and two misidentifications, while the explanations were designed mock-ups rather than actual Merlin explanations.
  • 4.3 Conducting and analyzing interviews: Two authors developed an initial descriptive codebook from five transcripts, then all authors refined it into shared conceptual themes.The analysis interpreted expressed needs such as learning the AI’s capabilities, identifying errors, and supplying better inputs as themes of improved human-AI collaboration.

5 RESULTS

Participants were broadly interested in learning about Merlin’s AI, but their willingness to pursue technical details varied by background and domain interest. Across participants, practically useful information was valued for improving collaboration, verification, and trust in the AI’s outputs.

  • 5.1 XAI needs: Participants unanimously wanted practically useful information to improve collaboration with Merlin, while active pursuit of system details varied by background and interest.High-AI-background or highly bird-interested participants were more willing to expend effort seeking technical information.
  • 5.1.2 Practically useful information: Participants wanted information about Merlin’s capabilities and limitations so they could judge reliability and supply better inputs.Understanding when identifications were more or less reliable was linked to improving the quality of submitted photos or recordings.
  • 5.1.2 Practically useful information: Participants frequently requested confidence displays, including percentage-based scores, to better determine when to trust the AI’s output.Some participants also wanted qualified or general outputs when the exact species could not be identified.
  • 5.1.2 Practically useful information: Participants wanted more detailed outputs that would make identifications easier to verify, especially for recordings containing multiple bird sounds.Requested details included the relevant time period and the type of sound detected, such as juvenile, flock, or alarm calls.
  • 5.1.2 Practically useful information: These practical explainability needs emerged before participants saw the XAI mock-ups, indicating they were not prompted solely by exposure to explanations.Participants expressed these needs while discussing their actual, real-world use of Merlin.

5.2 XAI uses: Participants intended to use explanations for calibrating trust, improving their own task skills, collaborating more effectively with AI, and giving constructive feedback to developers

Participants viewed XAI explanations as tools for collaborating with Merlin, not merely understanding its outputs. They wanted explanations to calibrate trust, build bird-identification skills, improve inputs, and provide developers with actionable feedback.

  • Calibrating trust: Participants intended to use explanations to determine when to trust Merlin’s identification results.They described explanations as information that could increase or decrease confidence in an output.
  • Learning task skills: More participants intended to use explanations to improve their bird-identification skills independently.They viewed Merlin as a teacher that could reveal features to look for when birding without the AI.
  • Improving collaboration: Participants wanted explanations to show how their behavior could help Merlin perform better by supplying improved inputs.They sought guidance for changing photos or recordings so the AI could identify birds more accurately.
  • Actionability: Participants criticized explanations that were interesting but did not help them change their behavior or help Merlin become more correct.Actionable feedback was a central criterion for explanations supporting collaboration.
  • Feedback to developers: Participants with high AI backgrounds intended to use explanations to give developers more detailed feedback for improving Merlin.Suggested mechanisms included correcting labeled regions, similar-bird examples, and other training information.

5.3 XAI perceptions: Participants preferred part-based explanations that resemble human reasoning and explanations

Participants generally preferred explanation forms that exposed specific, human-digestible parts of bird identification, especially prototypes and concepts. Preferences were qualified by concerns about coarse, generic, cluttered, numerical, or potentially unfaithful explanations.

  • Heatmap-based explanations: Heatmaps received mixed reviews because some participants found them intuitive, while others found them unintuitive, coarse, or uninformative.Critics also said heatmaps often showed what regions mattered without explaining why or how users should act.
  • Example-based explanations: Example-based explanations were widely considered understandable but often too general to reveal the features driving identifications.Their utility varied, including for trust calibration, but many participants found them insufficiently informative for broader intended uses.
  • Concept-based explanations: Concept-based explanations were praised because they break outputs into parts that human birders use when identifying birds.Participants could agree or disagree with the AI at the level of individual concepts and appreciated confidence scores.
  • Caveats: Some participants found concept-based explanations overwhelming, while AI-experienced participants questioned whether displayed explanations faithfully represented Merlin’s reasoning.These concerns show that usability and interpretation varied across users and explanation forms.
  • Prototype-based explanations: Prototype-based explanations were the most frequently selected favorite, with 14 participants preferring them.Participants valued their visual, part-based resemblance to how birders think and teach identification.
  • Design improvements: Participants recommended combining prototype explanations with concept, heatmap, or example-based approaches.They also suggested clearer descriptions, less clutter, and prototypes curated around salient features with experts and end-users.

6 DISCUSSION

The discussion reframes XAI as a medium for human-AI collaboration: explanations should help users improve decisions, provide better inputs, and give feedback. The authors recommend designing with end-users, supporting specific and multimodal explanations, and evaluating faithfulness and use-case effects.

  • 6.1 XAI as collaboration: Participants wanted explanations to help them improve Merlin’s accuracy and achieve better outcomes together, beyond improving decisions based on existing outputs.The paper distinguishes this collaborative goal from the XAI literature’s focus on usability.
  • 6.1 XAI as collaboration: The authors argue that XAI should provide actionable feedback to and from end-users, drawing an analogy to VizWiz prompts that improve user-provided photos.The proposed direction treats explanations as part of a richer interaction between people and AI.
  • 6.2 Designing Merlin’s XAI: The proposed Merlin design combines prototype- and concept-based explanations, letting users inspect matched image regions and their labels.Users would begin with boxed prototype regions and tap each box for a prototype plus a short concept description.
  • 6.3 Implications for future XAI research: End-users should participate in XAI design because they exposed mismatches between generic concepts and birders’ field-mark language.The paper identifies this as a creator-consumer gap and suggests developing concept banks with end-users.
  • 6.3 Implications for future XAI research: Participants wanted explanations that address why features mattered, although explaining causal relationships in computer-vision models remains an open problem.Heatmaps were criticized for identifying important regions without explaining their significance.
  • 6.3 Implications for future XAI research: The authors recommend multiple explanation forms and modalities because participants proposed combining approaches and using evidence such as photos, sound, and location.They also recommend evaluating explanations for both method goals and use-case goals to address faithfulness and negative effects.

7 LIMITATIONS AND FUTURE WORK

The study’s findings may not generalize beyond the Merlin context, and some participant subgroups were small. The study also excluded other stakeholder groups whose XAI needs may differ.

  • Findings may not generalize to other contexts because interview questions and study materials primarily concerned Merlin.The authors describe this as an intentional trade-off favoring depth in one specific context.
  • Some background subgroups had relatively few participants, motivating future research with larger subgroup samples.
  • The study did not include Merlin developers or deployers, whose XAI needs might differ from those of end-users.The authors plan comparative research across stakeholder groups.

8 CONCLUSION

This study examines end-users’ explainability needs and behaviors in a real-world setting through empirical research with 20 Merlin users. Participants wanted explanations that support collaboration with the AI, preferred part-based explanations resembling human reasoning, and revealed a creator-consumer gap in XAI.

  • A qualitative, descriptive, empirical study with 20 Merlin users examined real-world XAI needs and usage.
  • Participants wanted to use explanations to improve their collaboration with the AI system.
  • Participants preferred part-based explanations that resemble human reasoning and explanations among four representative XAI approaches.
  • Participant feedback revealed a creator-consumer gap in XAI and highlighted the need to involve end-users in XAI design.
  • The authors provide recommendations for future XAI research and design based on their findings.
Loading 2210.03735v2…