Source-linked AI summary
Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies
Vivian Lai, Chacha Chen, Q. Vera Liao, Alison Smith-Renner, Chenhao Tan
TL;DR
Human-AI decision making needs empirical understanding because fully automated decisions can be undesirable in high-stakes settings, while manual decisions can be inaccurate and time consuming. This survey synthesizes empirical human-subject studies across decision tasks, AI assistance, and evaluation metrics, identifying field-wide gaps and highlighting the need for common frameworks for generalizable knowledge.
Problem
Empirical human-AI decision-making studies are distributed across communities with diverse questions and methods, leaving a lack of coherent overview and shared frameworks for generalizable knowledge.
Method
The paper surveys empirical human-subject studies of human-AI decision making and analyzes study-design choices across decision tasks, AI assistance elements, and evaluation metrics.
Results
The survey finds that most existing studies focus on high-stake domains such as justice systems and medicine, while experts are seldom the subjects when tasks require domain expertise.
Takeaways & Limitations
The survey highlights the need for common frameworks that help researchers make rigorous design choices and build generalizable scientific knowledge.
Takeaways & Limitations
A task category driven by AI benchmark datasets has unclear stakes or societal significance.
Abstract
from arXiv · showhide
As AI systems demonstrate increasingly strong predictive performance, their adoption has grown in numerous domains. However, in high-stakes domains such as criminal justice and healthcare, full automation is often not desirable due to safety, ethical, and legal concerns, yet fully manual approaches can be inaccurate and time consuming. As a result, there is growing interest in the research community to augment human decision making with AI assistance. Besides developing AI technologies for this purpose, the emerging field of human-AI decision making must embrace empirical approaches to form a foundational understanding of how humans interact and work with AI to make decisions. To invite and help structure research efforts towards a science of understanding and improving human-AI decision making, we survey recent literature of empirical human-subject studies on this topic. We summarize the study design choices made in over 100 papers in three important aspects: (1) decision tasks, (2) AI models and AI assistance elements, and (3) evaluation metrics. For each aspect, we summarize current trends, discuss gaps in current practices of the field, and make a list of recommendations for future research. Our survey highlights the need to develop common frameworks to account for the design and research spaces of human-AI decision making, so that researchers can make rigorous choices in study design, and the research community can build on each other's work and produce generalizable scientific knowledge. We also hope this survey will serve as a bridge for HCI and AI communities to work together to mutually shape the empirical science and computational technologies for human-AI decision making.
1 INTRODUCTION
Human-AI decision-making research is growing, but empirical studies remain fragmented across domains, tasks, assistance elements, and evaluation practices. This survey organizes the field around three study-design aspects and calls for common frameworks to support coherent, generalizable knowledge.
- Empirical human-AI decision-making studies are increasingly important for evaluating AI assistance and understanding how people interact with AI when making decisions.
- The literature lacks a coherent overview and shared practices because studies span diverse communities, research questions, methodologies, domains, and decision tasks.This fragmentation hinders concerted research effort and the emergence of scientific knowledge.
- Researchers face unresolved questions about how task scope, AI assistance elements, and evaluation metrics affect findings and their generalizability.Individual studies often examine only a small set of assistance elements, while evaluation methodologies remain at an early stage.
- The survey examines studies intended to evaluate, understand, or improve human performance and experience on decision tasks rather than improve AI models.
- It analyzes decision tasks, AI models and assistance elements, and evaluation metrics, summarizing trends, identifying gaps, and recommending future research.
- The paper calls for common frameworks covering the design and research spaces of human-AI decision making and for mutual shaping between empirical science and computational technologies.
2 METHODOLOGY
The survey defines its scope around empirical human-subject studies of human-AI decision making and identifies papers through targeted literature searches and iterative author coding. Inclusion centers on evaluative decision tasks involving human decision makers.
- The survey includes empirical human-subject studies that evaluate, understand, or improve human performance or experience on human-AI decision-making tasks.
- It excludes purely formative studies, non-decision tasks, and papers focused on AI automation or improving the model rather than studying human decision makers.
- An initial search yielded over 130 papers, which the authors narrowed to over 80 in-scope papers before coding their study-design choices.
- The coding covered decision tasks, AI models and assistance elements, and evaluation metrics, with authors merging related codes and resolving ambiguities through discussion and consensus.
- The resulting summary tables provide a quick overview of the literature space.
3 DECISION TASKS
The survey organizes human-AI decision tasks by application domains and examines how task characteristics affect interpretation and generalizability. It identifies recurring gaps in task selection and reporting, then recommends frameworks and clearer justification for future studies.
- 3.1 Tasks Grouped by Domains: Researchers group human-AI decision tasks by application domains, including law and civic activities, medicine and healthcare, finance and business, education, leisure, professional work, artificial tasks, and generic tasks.Table 1 summarizes these domain groupings.
- 3.2 Task Characteristics: The survey characterizes tasks along four dimensions: risk, required expertise, subjectivity, and AI for emulation versus discovery.These dimensions are intended to help interpret results and assess how task choices may affect generalizability.
- 3.3 Summary & Takeaways: The survey recommends characterizing tasks with shared frameworks, justifying task choices, reporting task characteristics and participant expertise, and expanding dataset availability.It also urges researchers to examine the boundary between AI emulation and discovery tasks because results may not generalize across that boundary.
- 3.2 Task Characteristics: Human performance in some deception-detection tasks was close to random guessing, and human decisions are also prone to biases.These findings complicate the use of manual annotation for obtaining ground truth in such tasks.
- 3.2 Task Characteristics: Most surveyed studies focus on high-stakes domains, while artificial and generic tasks are also used; experts are seldom study subjects even when tasks require domain expertise.Most existing studies report objective rather than subjective decision tasks.
4 AI MODELS AND AI ASSISTANCE ELEMENTS
The survey examines the models, data, and assistance elements used in empirical human-AI decision-making studies, identifying recurring design patterns and gaps. It highlights a field dominated by shallow models and explanations, with comparatively limited attention to broader system elements, realistic settings, and task-driven frameworks.
- AI models and data: Wizard of Oz studies let researchers control model accuracy, error types, and explanation styles without investing in technical development.However, unrealistic simulations can impair the validity and generalizability of results.
- AI assistance elements: The survey organizes AI assistance into predictions, information about predictions, information about models or training data, and other system elements affecting agency or experience.These elements can help people assess prediction reliability and make informed decisions beyond simply following or ignoring a prediction.
- AI models and data: More empirical studies use shallow models than deep learning models, despite deep learning’s popularity and typically greater predictive power.Shallow models may be favored because they are easier to train, debug, and explain; with few features, they can achieve competitive accuracy.
- AI assistance elements: A large portion of prior work studies explanations, including both local explanations for individual predictions and global explanations for models.The survey identifies explanation-focused research as a central pattern in the literature.
- AI assistance elements: A small portion of studies explores system elements beyond the model, including workflow and user control that affect user agency and action spaces.These elements broaden the design space from model outputs and explanations to how people can act within the system.
- Recommendations: The survey recommends human-centered frameworks for the design space, studies beyond decision trials, and task-driven research complemented by feature-driven studies.It specifically calls for field and longitudinal studies that account for real user needs, prior experience, and individual differences.
5 EVALUATION OF HUMAN-AI DECISION MAKING
The survey organizes evaluation of human-AI decision making around task outcomes and AI-related perceptions or interactions, distinguishing objective from subjective measurements. It finds diverse constructs and substantial variation in measurement practices, motivating more common and validity-conscious metrics.
- Gaps and recommendations: The survey finds a wide range of evaluated constructs and significant variation in construct, content, and formulation, including multiple ways to measure trust.Many studies use home-grown subjective items, and papers often omit survey scales or questionnaires, limiting interpretation and replication.
- Evaluation framework: Evaluation metrics are grouped by whether they assess the decision task or the AI, then classified as objective or subjective measurements.The task dimension includes efficacy and efficiency, while AI-related evaluation covers constructs such as understanding, trust, satisfaction, and fairness.
- Task evaluation: Accuracy is the most commonly used objective metric for decision-task performance, with other measures selected for imbalanced, regression, gamified, or label-free tasks.Examples include F1, precision, recall, AUC-ROC, error rates, continuous accuracy counterparts, win rate, cumulative award, and inter-annotator agreement.
- Task evaluation: Subjective measures capture perceived accuracy, performance improvement, confidence, satisfaction, and related perceptions that objective task metrics do not directly measure.One reviewed approach combines subjective and objective measures to assess the soundness of users’ mental models.
- Task evaluation: Efficiency complements efficacy by measuring how quickly participants make decisions, most commonly through task time.The survey also notes self-reported task efficiency as common, while subjective efficiency metrics are absent from the reviewed papers.
- Gaps and recommendations: The authors recommend aligning metrics with research questions and targeted constructs, developing common metrics, and continually examining whether measures reflect stakeholder and societal values.They emphasize measurement validity, shared understanding of methods, and reflection on the long-term effects of prioritizing particular measures.
6 SUMMARY: TOWARDS A SCIENCE OF HUMAN-AI DECISION MAKING
The survey synthesizes more than 100 empirical studies to identify barriers to cumulative knowledge in human-AI decision making. It calls for shared frameworks, stronger reproducibility practices, and closer collaboration between HCI and AI to guide both empirical research and technical development.
- Summary and motivation: The survey summarizes design choices in more than 100 empirical human-AI decision-making papers to support shared knowledge and future research.Its synthesis focuses on decision tasks, AI and assistance elements, and evaluation metrics.
- Building on each other’s work: Advancing empirical science requires replication, cross-study meta-analysis, rigorous methodology, metrics development, theory building, and publication practices that enable reuse and reproduction.The survey recommends articulating rationales behind design choices and making study materials and knowledge more accessible.
- Developing common frameworks: Common frameworks should characterize decision tasks, AI assistance elements, and evaluation metrics so researchers can identify gaps and compare studies.Frameworks can make latent factors explicit, distinguish study setups and measurement coverage, and support more robust knowledge and theories.
- Developing common frameworks: The field should examine AI assistance across the entire decision process, including holistic system experience, contextual factors, and individual factors.The authors specifically urge moving beyond support for discrete decision trials and considering what should be measured for different stakeholders.
- Bridging AI and HCI communities: HCI and AI should mutually shape human-AI decision making by connecting empirical understanding of human needs with the development of effective and human-compatible AI.AI assistance frameworks can inform needed techniques, while evaluation frameworks can guide technical optimization.
- Bridging AI and HCI communities: The survey positions principled frameworks and empirical studies as ways for HCI insights to guide AI research and for AI work to incorporate human needs and behaviors.The authors note that cross-disciplinary collaboration requires cultural change and translative research to bridge perspectives.