Source-linked AI summary
Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
Michael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, Hanna Wallach
TL;DR
Fairness tools and practices may not align with the organizational contexts in which practitioners use them. This paper studies that gap for disaggregated evaluations through interviews and workshops with 33 practitioners from 10 teams at three technology companies, finding challenges in selecting metrics, relevant groups, and suitable datasets, alongside organizational pressures that shape fairness work.
Problem
Prior research identifies gaps between fairness practices’ intended designs and their use in particular organizational contexts.
Method
The authors conducted semi-structured interviews and structured workshops with 33 practitioners from 10 teams at three technology companies.
Results
Practitioners faced challenges choosing performance metrics, identifying relevant stakeholders and demographic groups, and collecting datasets for disaggregated evaluations.
Takeaways & Limitations
Fairness work is shaped by limited stakeholder or domain-expert engagement, customer-prioritizing business imperatives, and the drive to deploy AI systems at scale.
Takeaways & Limitations
The findings may not generalize to all teams and companies because most workshop participants developed language technologies and participation was limited to three companies.
Abstract
from arXiv · showhide
Various tools and practices have been developed to support practitioners in identifying, assessing, and mitigating fairness-related harms caused by AI systems. However, prior research has highlighted gaps between the intended design of these tools and practices and their use within particular contexts, including gaps caused by the role that organizational factors play in shaping fairness work. In this paper, we investigate these gaps for one such practice: disaggregated evaluations of AI systems, intended to uncover performance disparities between demographic groups. By conducting semi-structured interviews and structured workshops with thirty-three AI practitioners from ten teams at three technology companies, we identify practitioners' processes, challenges, and needs for support when designing disaggregated evaluations. We find that practitioners face challenges when choosing performance metrics, identifying the most relevant direct stakeholders and demographic groups on which to focus, and collecting datasets with which to conduct disaggregated evaluations. More generally, we identify impacts on fairness work stemming from a lack of engagement with direct stakeholders or domain experts, business imperatives that prioritize customers over marginalized groups, and the drive to deploy AI systems at scale.
1 INTRODUCTION
The paper examines how practitioners design disaggregated evaluations to uncover demographic performance disparities, focusing on their processes, challenges, and organizational support needs. It argues that fairness practices must be understood within the organizational contexts shaping their use.
- Motivation: Disaggregated evaluations assess AI performance separately across demographic groups to uncover disparities that aggregate assessment can obscure.Such disparities may otherwise be discovered only after deployment and after people experience fairness-related harms.
- Research questions: The study asks how practitioners design disaggregated evaluations, what organizational support they need, and how organizational contexts affect their work.
- Approach: The authors investigate these questions through semi-structured interviews and structured workshops with 33 practitioners from 10 teams at three technology companies.Participants held roles including program or product management, data science, and user experience design.
- Findings: Practitioners encounter challenges choosing performance metrics, identifying relevant stakeholders and demographic groups, and collecting suitable evaluation datasets.
- Implications: The paper highlights tensions in practitioners’ desires for organizational guidance and resources when designing and conducting disaggregated evaluations.
- Implications: The discussion connects fairness-work challenges to business imperatives that prioritize customers over marginalized groups and to AI deployment at scale.
2 RELATED WORK
Prior work frames disaggregated evaluations as a way to identify demographic performance disparities, but shows that organizational contexts and incentives shape how fairness work is conducted. This paper builds on that literature to examine tensions between evaluation practices and pressures to deploy AI systems quickly and broadly.
- Fairness assessment: Disaggregated evaluations can be conducted by external third parties, but relying solely on them may create drawbacks despite potential credibility benefits.
- Fairness assessment: Internal-auditing frameworks aim to align AI evaluations with organizational principles, yet practitioners still face incentives to ship products quickly.Those incentives may conflict with the slow and careful work needed to design and conduct disaggregated evaluations.
- Anticipating harms: Anticipatory ethics and related practices emphasize understanding use cases and deployment contexts when identifying potential harms from emerging technologies.
- Organizational contexts: Organizational contexts shape AI practitioners’ work and can contribute to gaps between fairness practices’ intended designs and their use.
- Organizational contexts: Prior fairness research links organizational cultures, incentives, and business imperatives to practitioners’ ability to conduct fairness work.
3.1 Study design
The study used a two-phase protocol combining interviews with program or product managers and structured workshops involving multiple team members. Participants came from varied AI product teams and professional roles across three technology companies.
- Study protocol: The protocol comprised 30–60-minute semi-structured interviews followed by two 90-minute structured workshop sessions.The second phase involved multiple members of each recruited manager’s team.
- Study protocol: Interviews examined teams’ development practices and existing fairness work, while informing preparation for the workshop sessions.
- Participants: Thirty-three practitioners from 10 teams at three technology companies participated, although only seven teams completed both study phases.
- Participants: Participants included program or product managers, data scientists, applied scientists, software developers, user experience designers, and technical managers.
- Participants: Teams developed products and services including text suggestion, chatbots, text summarization, and fraud detection systems.
3.2 Designing disaggregated evaluations
The workshop protocol adapted Barocas et al.’s framework into a sequence of design questions for disaggregated evaluations. Teams considered performance, contexts, stakeholders, demographic groups, and the data required to evaluate them.
- Contexts and stakeholders: Teams identified relevant use cases and deployment contexts before deciding which stakeholders might experience poor performance.
- Evaluation design: The protocol organized evaluation design around what performance means, which contexts and groups matter, what data is needed, and how system behavior should be assessed.
- Performance: Teams first defined good performance for their AI systems and selected performance metrics to operationalize that definition.They also considered what existing metrics might miss, including aspects more likely to exhibit disparities.
- Contexts and stakeholders: For each relevant direct stakeholder, teams considered which demographic groups might be most at risk of poor performance.Direct stakeholders include system users, operators, and people directly affected by someone else’s use of the system.
- Data requirements: A suitable evaluation dataset must contain enough labeled examples from each group of interest and capture realistic within-group variation.It must also include relevant information about other factors that could affect performance, such as lighting conditions.
3.3 Data analysis
The authors analyzed interview and workshop transcripts through inductive thematic analysis, iteratively refining codes and themes across the research team.
- Three authors independently open-coded transcripts, beginning by coding the same transcript and using Atlas.ti.The coding process was inductive and software-assisted.
- All five authors discussed emerging themes, while three authors repeatedly reorganized and hierarchically structured the codes.The team merged redundant codes during iterative refinement.
- Two authors wrote reflective memos to develop analytic questions and connect themes across study phases.The memos incorporated interview and workshop experience and workshop artifacts.
3.4 Positionality
The authors situate their analysis within their experiences, positionality, and disciplinary backgrounds as U.S.-based industry-oriented researchers working on AI fairness.
- The researchers identify as U.S.-based researchers primarily working in industry with experience collaborating with AI practitioners on fairness projects.They explicitly discuss these experiences as shaping their research perspectives and approaches.
- Their disciplinary backgrounds include AI and HCI, which informed their research on fairness in AI systems.
4 FINDINGS
Practitioners face challenges in designing disaggregated evaluations, especially when selecting metrics, identifying relevant stakeholders and demographic groups, and collecting evaluation datasets.
- Practitioners reported challenges choosing performance metrics, identifying relevant stakeholders and demographic groups, and collecting datasets for disaggregated evaluations.
4.1 Challenges when designing disaggregated evaluations
Designing disaggregated evaluations involves contested metric choices, incomplete stakeholder engagement, organizational pressure for rapid scale, and difficulties obtaining demographic data.
- 4.1.1 Challenges when choosing performance metrics: Teams sometimes reused aggregate-performance metrics, but most lacked standard metrics for their AI systems.For some systems, established metrics such as word error rate made metric selection relatively straightforward.
- 4.1.1 Challenges when choosing performance metrics: Some teams considered usage, precision, or recall inadequate for assessing fairness-related harms and developed new quality-oriented metrics.One team sought metrics capturing output quality and catastrophic failures rather than usage alone.
- 4.1.1 Challenges when choosing performance metrics: Fraud-detection stakeholders faced a deeper conflict among affected transaction subjects, platform companies, and government auditors.The team eventually brought two stakeholders together, but the underlying tension remained.
- 4.1.1 Challenges when choosing performance metrics: Stakeholder discussions about aggregate performance did not resolve fairness-metric choices because they excluded direct stakeholders.The omitted direct stakeholders were people whose transactions might be mistakenly labeled fraudulent.
- 4.1.1 Challenges when choosing performance metrics: Metric decisions were shaped by business imperatives that prioritized customers over marginalized groups and reflected disagreements about system goals.A fraud-detection team described pressure to use the metric preferred by platform customers.
- 4.1.2 Challenges when identifying direct stakeholders and demographic groups: Participants wanted to identify at-risk groups using geographic understandings of marginalization and engagement with direct stakeholders and domain experts.
- 4.1.2 Challenges when identifying direct stakeholders and demographic groups: Development practices usually engaged customers and users, but not other direct stakeholders or users as sources of knowledge about marginalization.User engagement often focused on feedback rather than informing fairness-related decisions.
- 4.1.2 Challenges when identifying direct stakeholders and demographic groups: Using team members’ experiences to identify relevant groups was problematic because many AI teams were demographically homogeneous.Participants recognized a possible disconnect between harms customers experience and harms teams identify or measure.
4.2 Priorities that compound existing inequities
Teams prioritized direct stakeholders and demographic groups using harm severity, data availability, mitigation ease, customer or market needs, and perceived reputational impact. These choices could favor already powerful or visible groups and compound existing inequities.
- Perceived severity of fairness-related harms: Teams considered harm severity, affected population size, accumulated harm, and tradeoffs between benefits and harms when prioritizing stakeholders and demographic groups.Participants often could not directly measure harm severity and instead considered perceived severity or estimated breadth.
- Perceived ease of data collection or mitigation: Teams also prioritized groups they already worked with or for which data were available, making existing relationships and data access practical selection criteria.
- Perceived ease of data collection or mitigation: Three of seven teams prioritized disparities that seemed easier to mitigate, even when the affected group was not especially marginalized or severely harmed.British English speakers were cited as relatively easy to support, despite not typically being marginalized by AI systems or society.
- Perceived PR or brand impacts and needs of customers or markets: Teams prioritized groups according to potential PR, brand, customer, and market impacts, which could shift attention toward visible users, high-value customers, or organizational image.Participants described social-media visibility, newspaper headlines, brand association, and customer needs as prioritization considerations.
- Reflexive and reactive approaches: These prioritization approaches are backward-looking because they focus on disparities uncovered after deployment, rather than reflecting on values and impacts during development.
- Needs of customers or markets: Customer-focused prioritization could favor already over-represented and privileged groups, while deployment decisions based on market opportunity or feedback could compound geographic inequities.Participants linked customer obsession and business outcomes to reduced attention to marginalized groups and direct stakeholders.
4.3 Needs for organizational support
Practitioners wanted organization-wide guidance, reusable data strategies, and resources for disaggregated evaluations, while recognizing that teams must adapt support to varied systems and contexts. Business priorities and limited budgets constrained what fairness work teams could conduct.
- Guidance: Participants requested organizational guidance on selecting relevant direct stakeholders, demographic groups, fairness harms, and datasets for disaggregated evaluations.
- Guidance: Participants favored global lists or guidance that teams could adapt, supplement, and prioritize for specific models, products, services, and deployment contexts.They wanted broad coverage requirements without assuming every demographic factor would be equally relevant across systems.
- Strategies for collecting datasets: Participants proposed centralized data warehouses or clearinghouses and organizational strategies for collecting demographic data while balancing data collection with privacy.
- Tensions in organizational support: Organization-wide guidance and dataset strategies may be difficult to establish because AI systems, use cases, and deployment contexts vary substantially.Centralized approaches may also homogenize understandings of demographic groups across geographic contexts.
- Resources: Teams needed money, time, and personnel for evaluations, stakeholder and domain-expert engagement, data collection, and labeling, but budgets constrained what they could achieve.
- Resources: Business imperatives determined whether fairness resources became available, with executive buy-in helping fairness priorities reach immediate managers and teams.
- Resources: Practitioners faced a cycle in which leadership wanted evidence of harms before funding evaluations, although evaluations were needed to produce that evidence.
5 DISCUSSION
The discussion shows that organizational priorities and deployment scale shape how practitioners design disaggregated evaluations, often limiting engagement with marginalized stakeholders and local conditions. It also identifies study boundaries and argues that fairness practices may need to change as AI systems expand across contexts.
- 5.1 Implications of business imperatives that shape disaggregated evaluations: Business imperatives shape practitioners’ choices about performance metrics, direct stakeholders, demographic groups, and datasets when designing disaggregated evaluations.
- 5.1 Implications of business imperatives that shape disaggregated evaluations: Stakeholder identification often prioritizes high-value customers and higher-tier markets, reflecting instrumental rationales tied to organizational performance or profitability.
- 5.1 Implications of business imperatives that shape disaggregated evaluations: Without engagement with direct stakeholders or domain experts, practitioners may rely on personal experiences, potentially deprioritizing lower-tier markets and groups.
- 5.2 Implications of deploying AI systems at scale: Participants on every team reported pressure to expand AI deployments geographically, with these pressures affecting nearly every evaluation-design decision.
- 5.2 Implications of deploying AI systems at scale: The authors argue that human-centered requirements should receive equal consideration with technical feasibility and that development practices should respond to local conditions as systems expand.
- 5.3 Limitations: The study is limited by anonymized organizational details, a U.S.-centered researcher positionality, predominantly language-technology workshop participants, and non-longitudinal data that may not generalize across teams or companies.