Source-linked AI summary

Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems

Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, Elena L. Glassman

arXiv:2001.08298v1cs.AIcs.HC

TL;DR

Explainable AI systems are sociotechnical systems, but evaluations often use proxy tasks or subjective measures instead of human+AI performance on real decisions. Across two online experiments and an in-person think-aloud study, the paper found that proxy-task results did not correspond to actual-task results and that trust and preference did not predict performance. The authors therefore caution against relying on these measures alone, while noting that realistic-task studies are costly and can slow research.

  • Problem

    Explainable AI is rarely evaluated by measuring human+AI teams on actual decision-making tasks, leaving the predictive value of proxy tasks and subjective measures uncertain.

  • Method

    The authors conducted two online experiments and one in-person think-aloud study comparing proxy and actual nutrition decisions with inductive and deductive explanations.

  • Results

    Proxy-task evaluations did not predict actual-task results, and participants performed significantly better with inductive explanations when the AI made erroneous recommendations despite preferring and trusting deductive explanations more.

  • Takeaways & Limitations

    Trust and preference should complement rather than replace performance measures, and evaluations should assess complete human+AI systems on real tasks.

  • Takeaways & Limitations

    Realistic human-subject experiments are expensive in time and resources and may negatively affect the pace of innovation.

Abstract

from arXiv · show

Explainable artificially intelligent (XAI) systems form part of sociotechnical systems, e.g., human+AI teams tasked with making decisions. Yet, current XAI systems are rarely evaluated by measuring the performance of human+AI teams on actual decision-making tasks. We conducted two online experiments and one in-person think-aloud study to evaluate two currently common techniques for evaluating XAI systems: (1) using proxy, artificial tasks such as how well humans predict the AI's decision from the given explanations, and (2) using subjective measures of trust and preference as predictors of actual performance. The results of our experiments demonstrate that evaluations with proxy tasks did not predict the results of the evaluations with the actual decision-making tasks. Further, the subjective measures on evaluations with actual decision-making tasks did not predict the objective performance on those same tasks. Our results suggest that by employing misleading evaluation methods, our field may be inadvertently slowing its progress toward developing human+AI teams that can reliably perform better than humans or AIs alone.

1 INTRODUCTION

Human+AI teams are expected to outperform people or AI alone, yet they often perform worse. The paper argues that proxy tasks and subjective measures may mislead evaluations of explainable AI systems intended for real decision-making.

  • Human+AI teams often perform worse than AIs alone despite their expected complementary strengths.
  • The paper identifies proxy tasks and subjective measures as critical evaluation mistakes when the goal concerns human+AI performance on complex decisions.
  • Proxy tasks may alter behavior by forcing attention to explanations, unlike realistic decisions where people may prefer less demanding cognition.
  • Two online experiments and an in-person think-aloud study compared proxy and actual nutrition decisions using inductive and deductive explanations.
  • Proxy-task subjective results did not generalize to actual decisions, and subjective results on actual tasks did not predict objective performance.
  • The authors conclude that explainable AI evaluations should demonstrate how complete human+AI sociotechnical systems perform on real tasks.

2 RELATED WORK

Decision-support systems combine AI assistance with human judgment, making evaluation of the complete sociotechnical system important. Related work distinguishes actual-task evaluations from proxy tasks and questions whether subjective measures predict performance.

  • 2.1 Decision-making and Decision Support Systems: People use heuristics to reduce cognitive effort, but these shortcuts can sometimes produce systematic biases and poor decisions.
  • 2.1 Decision-making and Decision Support Systems: AI-enabled decision-support systems increasingly assist decisions, but overall accuracy depends on both the system and the human final decision-maker.
  • 2.1 Decision-making and Decision Support Systems: Cognitive forcing strategies promote self-awareness and self-monitoring and have improved decision performance in some human- and AI-assisted settings.
  • 2.2 Evaluating AI-Powered Decision Support Systems: Explainable AI evaluation taxonomies distinguish application-grounded, human-grounded, and functionally grounded approaches.
  • 2.2 Evaluating AI-Powered Decision Support Systems: Actual-task studies assess human and system performance together, whereas proxy studies ask users to simulate model decisions or decision boundaries.
  • 2.2 Evaluating AI-Powered Decision Support Systems: Trust, satisfaction, confidence, and related subjective measures are informative but do not necessarily predict users’ performance with AI systems.

3 EXPERIMENTS

The experiments test whether proxy-task results predict realistic decision-making results and whether self-reported trust and preference predict ultimate human+AI performance.

  • H1 proposes that proxy-task results may not predict outcomes when users focus on realistic decision-making.
  • H2 proposes that self-reported trust and preference for explanation designs may not predict ultimate human+AI performance.

3.1 Proxy Task

The proxy task used nutrition images and simulated AI explanations to ask participants to predict AI decisions. Participants compared inductive and deductive explanations while reporting performance and subjective evaluations.

  • 3.1 Proxy Task: Participants predicted whether a simulated AI would classify the fat content of each of 24 pictured food plates.
  • 3.1 Proxy Task: Inductive explanations used examples requiring participants to recognize relevant ingredients and infer a conclusion.
  • 3.1 Proxy Task: Each simulated AI was 75% accurate, misclassifying images or misrecognizing ingredients in 25% of cases.
  • 3.1 Proxy Task: The procedure was an online, two-block within-subjects study with inductive explanations in one block and deductive explanations in the other.
  • 3.1 Proxy Task: Figure 1 contrasts an inductive explanation with appropriate examples and a deductive explanation containing misrecognized ingredients.
  • 3.1 Proxy Task: Participants completed questionnaires assessing trust, preference, mental demand, understanding, and direct comparisons between the two simulated AIs.
  • 3.1 Proxy Task: 200 adults in the United States were recruited through Amazon Mechanical Turk, with 183 retained for final analyses.
  • 3.1 Proxy Task: Proxy-task performance was measured as the percentage of correct predictions of the AI’s decisions, alongside perceived appropriateness of its examples or ingredients.

3.2 Actual Decision-making Task

The actual decision-making task asked participants to make nutrition decisions with or without simulated AI recommendations and explanations. It compared no-AI, no-explanation, and explanation conditions, including inductive and deductive explanations.

  • Task design: Participants judged whether each food plate exceeded a specified fat-content threshold, assisted by simulated AI recommendations and explanations.The task used 24 food images and asked participants to make their own decisions rather than predict the AI’s decision.
  • Conditions: The study compared no-AI, no-explanation, and recommendation-plus-explanation conditions.The no-AI condition provided neither recommendations nor explanations; the no-explanation condition provided recommendations without explanations.
  • Conditions: Within the explanation condition, participants saw either inductive or deductive explanations.The explanation-type factor was applied only when simulated AI recommendations were accompanied by explanations.
  • Measures: The procedure measured performance, understanding, trust, and mental demand through repeated questions and questionnaires.Performance was based on correct answers, while subjective measures used 5-point Likert scales for understanding, trust, and mental demand.
  • Procedure: The experiment recruited 113 U.S. adults through Amazon Mechanical Turk, retaining 102 for final analyses.The task lasted 10 minutes on average, and each worker was paid 5 USD per task.

4 RESULTS

Actual-task performance improved with AI recommendations and improved further when explanations were provided, while explanation type did not change overall accuracy. However, inductive explanations helped participants detect erroneous recommendations, whereas subjective evaluations favored different explanation types across tasks.

  • Subjective evaluations: Trust and preference favored inductive explanations in the proxy task but deductive explanations in the actual decision-making task.In the proxy task, participants trusted inductive explanations more, while 63% preferred deductive explanations in the actual task.
  • Actual decision-making performance: M = 0.64 versus M = 0.64: explanation type produced no significant difference in overall performance.Performance also did not significantly differ between explanation types on questions with incorrect explanations when considered as an overall comparison.
  • Subjective evaluations: Participants rated explanations as more helpful and understood the AI better when explanations were present, with deductive explanations rated more helpful.Helpfulness was M = 3.78 with explanations versus M = 3.26 without, while perceived understanding was M = 3.84 versus M = 3.67.
  • Actual decision-making performance: M = 0.74 versus M = 0.68: explanations significantly improved overall accuracy compared with recommendations without explanations.AI recommendations overall also outperformed no-AI assistance, with M = 0.72 versus M = 0.46.
  • Actual decision-making performance: M = 0.63 versus M = 0.48: participants were more accurate with inductive than deductive explanations when AI recommendations were incorrect.When recommendations were correct, performance was similar with inductive and deductive explanations: M = .78 versus M = .81.
  • Replication: The two experiments were replicated with almost identical setups and produced the same main significance results.The replication covered the principal findings reported in this results section.

5 QUALITATIVE STUDY

The qualitative study examined how participants reasoned with inductive and deductive explanations during the actual task. Think-aloud participants often preferred inductive explanations, but the method may have shifted behavior toward analytical engagement and proxy-like evaluations.

  • Preference of one explanation type over another: Eight of 11 participants preferred inductive explanations, often viewing similar images as evidence that the AI had supporting data.Participants preferring deductive explanations instead regarded recognized ingredients as reliable evidence.
  • Preference of one explanation type over another: Participants used inductive explanations to confirm an initial judgment, whereas deductive explanations prompted more evaluation before deciding.The observed difference concerned how participants integrated explanations and recommendations into their decision process.
  • Cognitive Demand: Ten of 11 participants found inductive explanations easier to understand, and participants reported spending more time thinking with deductive explanations.The qualitative pattern aligned with the proxy-task results for perceived cognitive demand and ease of understanding.
  • Errors and Over-reliance: Nine of 11 participants claimed to trust inductive explanations more, even though some participants agreed with erroneous inductive recommendations.Some participants missed differences between the query image and comparison images or judged them similar enough to support the recommendation.
  • Impact of the Think-Aloud method on participant behavior: The think-aloud procedure may have increased analytical thinking about recommendations and explanations relative to the prior actual-task experiment.The authors characterize think-aloud as a possible cognitive forcing intervention that can affect performance on cognitively demanding tasks.

6 DISCUSSION

The experiments support both hypotheses: proxy-task evaluations may not predict realistic-task outcomes, and subjective trust or preference may not predict performance. The discussion attributes these mismatches partly to differences in cognitive effort and task interaction, while noting practical limits on realistic evaluation.

  • H1 and H2 state that proxy-task results and subjective trust or preference measures may not predict ultimate human+AI performance.
  • Participants preferred and trusted inductive explanations in the proxy task, but preferred and trusted deductive explanations in the actual decision-making task.
  • In actual tasks with incorrect AI recommendations, participants provided correct answers significantly more often with inductive than deductive explanations.
  • These contradictory experiment results support H1, potentially because proxy tasks require analytical engagement while actual tasks let users choose whether and how deeply to engage.
  • Subjective measures should complement rather than replace performance measures because participants preferred and trusted deductive explanations but recognized AI errors better with inductive explanations.
  • Realistic human-subject experiments are expensive and resource-intensive, motivating research into lower-burden techniques that predict deployment outcomes accurately.
  • Cognitive effort may explain evaluation differences: users may avoid demanding explanations, while think-aloud instructions may induce additional effort and alter evaluation behavior.

7 CONCLUSION

The study concludes that proxy tasks, subjective measures, and think-aloud studies can produce evaluation results that diverge from realistic human+AI decision-making performance. It therefore emphasizes holistic sociotechnical evaluation as explainable AI enters critical decision-making domains.

  • The study used online experiments and an in-person study to show that assumptions about evaluation can produce misleading results.
  • Proxy-task evaluations may shift users’ focus toward the AI and produce explanation preferences that reverse in actual decision-making tasks.
  • Trust and preference may not correspond to performance: users trusted and preferred deductive explanations but recognized AI errors better with inductive explanations.
  • Think-aloud studies may not represent realistic decisions because their actual-task results aligned more closely with proxy-task results.
  • The authors urge cautious evaluation design and holistic assessment of explainable AI interfaces as sociotechnical systems in critical domains.
Loading 2001.08298v1…