Source-linked AI summary
Improving fairness in machine learning systems: What do industry practitioners need?
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, Miro Dudík, Hanna Wallach
TL;DR
ML systems can amplify social inequities, yet fair ML research has often been designed without systematic evidence about industry practitioners’ needs. Through 35 interviews and a survey of 267 practitioners, this paper identifies alignments and disconnects between practice and research, highlighting directions for more relevant ML and HCI support.
Problem
Fair ML research has rarely been guided by systematic evidence about the challenges and support needs of industry practitioners developing fairer systems.
Method
The study combines semi-structured interviews with 35 practitioners and an anonymous survey of 267 ML practitioners across commercial product contexts.
Results
Industry teams report needs that have been neglected in fair ML research, including emphasis on data collection and difficulty applying existing auditing and de-biasing methods.
Takeaways & Limitations
The findings identify ML and HCI research opportunities aimed at reducing technical and organizational barriers to improving fairness in industry products.
Takeaways & Limitations
Snowball recruitment and targeting teams with prior fairness-related media coverage may have overrepresented practitioners unusually motivated to address fairness.
Abstract
from arXiv · showhide
The potential for machine learning (ML) systems to amplify social inequities and unfairness is receiving increasing popular and academic attention. A surge of recent work has focused on the development of algorithmic tools to assess and mitigate such unfairness. If these tools are to have a positive impact on industry practice, however, it is crucial that their design be informed by an understanding of real-world needs. Through 35 semi-structured interviews and an anonymous survey of 267 ML practitioners, we conduct the first systematic investigation of commercial product teams' challenges and needs for support in developing fairer ML systems. We identify areas of alignment and disconnect between the challenges faced by industry practitioners and solutions proposed in the fair ML research literature. Based on these findings, we highlight directions for future ML and HCI research that will better address industry practitioners' needs.
1 INTRODUCTION
ML systems increasingly affect consequential life outcomes, while fair ML research has emphasized statistical definitions and algorithmic bias mitigation. This paper argues that industry tools should be designed around practitioners’ actual challenges and investigates those needs systematically.
- ML systems increasingly influence healthcare, education, media exposure, employment, incarceration, and policing, creating concern that they may amplify social inequities.
- Fair ML research has concentrated on statistical fairness definitions, bias assessment, mitigation algorithms, and integrated toolkits for wider usability.
- Fairness tools may have limited industry impact when their design follows available algorithms more than practitioners’ real-world needs.
- 35 practitioners across 25 product teams and 10 companies were interviewed, followed by an anonymous survey of 267 practitioners to assess theme prevalence and generality.
- Industry practitioners’ needs both align with and diverge from fair ML research, including greater emphasis on data collection and challenges applying individual-level auditing methods.
- The study identifies opportunities for ML and HCI research to better address industry needs and improve fairness in practice.
3 METHODS
The study combines interviews and a broader survey to examine commercial ML practitioners’ fairness practices, challenges, and support needs. Practitioners are defined broadly across roles on ML product teams, and the study received ethical review.
- 35 practitioners from 25 ML product teams at 10 major companies participated in semi-structured interviews, followed by an anonymous survey of 267 practitioners.
- Practitioners included anyone working in any role on a team developing products or services involving ML.
- The study was ethically reviewed and IRB-approved, with interview protocols and survey questions provided in supplementary materials.
Interview Study
The interview study used formative and in-depth rounds to examine how commercial ML teams understand fairness, work across development stages, and need support. Interviews were analyzed through contextual-design interpretation and affinity diagramming.
- The interview study began with six formative interviews and then expanded to more in-depth interviews based on emerging themes.
- The first round interviewed product managers about products, customers, team structures, and whether fairness entered team workflows.
- Fairness was left open-ended, with clarification defined as ML systems performing differently across groups in potentially undesirable ways.
- The main round included 29 practitioners across 19 ML product teams from 10 technology companies and sought multiple roles where possible.
- Recruitment was difficult because contacts feared reputational harm, researcher distrust, and accidental disclosure of trade secrets.
- Interviews examined fairness at each pipeline stage, from data collection and dataset design through product development, assessment, and mitigation.
- An oracle exercise encouraged participants to describe support needs without restricting them to currently available solutions.
- Approximately 25 hours of interview audio were analyzed using interpretation sessions and iterative bottom-up affinity diagramming.
Survey
The survey quantitatively supplemented the interviews by testing the prevalence and generality of their themes among ML product practitioners. It used snowball recruitment, branching questions, and responses from 267 completed participants.
- The anonymous online survey recruited ML product practitioners through contacts at over 40 companies, social media, and online ML and AI communities.
- Figure 1 summarizes the top 10 self-reported technology areas and team roles, with respondents able to select multiple options.
- Survey structure mirrored the main interviews and measured demographic backgrounds, current practices, challenges, and support needs.
- 287 people started the survey, but analyses included only the 267 respondents who completed at least one section beyond demographics.
4 RESULTS AND DISCUSSION
The results identify common industry challenges and neglected support needs around fairness in ML, while emphasizing that technical tools must accompany organizational processes, policies, and education.
- Teams reported common fairness challenges and needs across diverse application domains, including data collection, blind spots, proactive auditing, and auditing complex ML systems.The themes were derived from interviews and supplemented with survey responses.
- The study highlights research and design opportunities addressing industry needs that have received little attention in the fair ML literature.
- Many findings represent barriers faced even by practitioners who are motivated to improve fairness but feel unsupported by team or company leadership.Interviewees were recruited through snowball sampling and from teams whose products had received media coverage about ML unfairness.
- Fairness efforts must combine technically focused research with improvements to organizational processes, policies, and education because unfairness is fundamentally socio-technical.
Fairness-aware Data Collection
Practitioners often see data collection and curation as the most important points for improving fairness, contrasting with fair ML research that commonly treats training data as fixed. They need guidance for collecting representative data and designing test sets that expose bias.
- Most interviewees reported lacking processes for collecting and curating balanced or representative datasets, with teams often ingesting available data without a fairness review.
- Practitioners requested active guidance during data collection to target underrepresented groups and allocate costly labeling or scoring effort effectively.The automated essay-scoring example concerned obtaining enough high-scoring essays from African American students.
- 60% of respondents whose teams had some control over data collection rated active guidance at least “Very” useful.This percentage is among the 25% who indicated such control.
- 66% of respondents rated tools scaffolding test-set design “Very” or “Extremely” useful for checking subgroup coverage and bias across slices.This percentage is among the 70% asked the question.
- Future research should explicitly design data and test-set tools to support fairness and equity in downstream ML models, not only overall predictive accuracy.
Challenges Due to Blind Spots
Industry teams struggle to anticipate relevant subpopulations and discover fairness problems before complaints or negative media coverage. Addressing these blind spots requires sharing cultural and domain knowledge beyond any single team.
- Most interviewees needed support identifying application-specific subpopulations so teams could collect sufficient data or balance existing datasets.Participants described the relevant subpopulations as highly context- and application-dependent.
- Teams often discovered serious fairness issues only after customer complaints or negative media coverage, rather than through proactive monitoring.
- Teams reported that no individual member has expertise in all forms of bias, particularly across different cultures.
- 67% of surveyed respondents rated tools for pooling knowledge about potential fairness issues at least “Very” useful.This percentage is among the 71% asked the question.
- Blind spots can hamper efforts to obtain corrective training data when teams lack knowledge of affected people or contexts.An image-captioning team struggled to identify unfamiliar celebrities who were underrepresented in training data.
- Future research should support sharing and reuse of cultural and domain knowledge across teams or companies, including shared test sets, case studies, and external experts.
Needs for More Proactive Auditing Processes
Practitioners described fairness auditing as reactive, difficult to scale, and poorly integrated into routine workflows. They need domain-specific, proactive methods that work despite limited demographic data and can distinguish isolated incidents from systemic problems.
- Fairness auditing is often reactive, centered on customer complaints rather than proactive detection.
- Teams struggle to audit fairness without individual-level demographics, sometimes considering proxy-based inference but worrying that proxies introduce additional bias.
- Limited user-study samples and context dependence make comprehensive, scalable automated auditing difficult.
- 62% of surveyed respondents asked this question considered tools for finding other instances of an issue at least Very useful, supporting diagnosis of systemic problems.
- Practitioners need domain-specific metrics, processes, and tools to navigate application-dependent fairness challenges.
Needs for More Holistic Auditing Methods
Practitioners argued that fairness should be assessed at the level of the complete user-facing system, not only through de-contextualized metrics for individual models. Complex interactions and user expectations require empirical and simulation-based approaches.
- Fairness may be a system-level property because individual model metrics do not always map cleanly to system utility for users.
- Interactions among components and other design elements can make system behavior difficult to predict without empirical studies involving actual users.
- User studies revealed that many users viewed manipulating image-search results as unethical, despite an alternative goal of representing more image types.
- Simulation tools could help developers prototype conversational systems and identify risky conversation patterns in context-dependent interactions.
- Future auditing and prototyping methods should address sensitive contexts and real-world impacts beyond de-contextualized quantitative metrics.
Addressing Detected Issues
After detecting fairness issues, teams need help selecting effective interventions, estimating data requirements, anticipating trade-offs, and avoiding harmful side effects. Practitioners also emphasized broader system-design changes and ethical guidance, not only model-level debiasing.
- Teams need to anticipate trade-offs between fairness definitions and other system desiderata, including user satisfaction.
- Teams struggle to isolate causes and compare the cheapest, most effective strategies for particular fairness issues.
- Different teams make different intervention choices because strategy costs vary by context.
- 71% of respondents considered tools for understanding UX side effects of fairness strategies at least Very useful.
- 66% of surveyed respondents considered support for estimating minimum additional data per subpopulation at least Very useful.
- Fifty-five percent of respondents rated domain-specific ethical resources at least Very useful, reflecting concerns about the morality of fairness interventions.
- Forty percent of survey respondents reported using fail-soft strategies that change broader system responses to model outputs.
Biases in the Humans in the Loop
Practitioners recognized that human judgments embedded throughout the ML pipeline can introduce bias. They therefore need tools and processes to understand and mitigate bias in labeling, scoring, and other human activities.
- Bias can enter through humans embedded across the ML development pipeline, including annotators and user-study participants.
- 69% of respondents asked about it considered tools to reduce human bias in labeling or scoring at least Very useful.
- Among respondents asked about current practice, 69% reported that their teams at least Sometimes mitigate labeling or scoring biases.
- Future research should help teams understand and mitigate human biases throughout the ML development pipeline.
5 CONCLUSIONS AND FUTURE DIRECTIONS
Industry practitioners face technical and organizational barriers to improving fairness, while fair ML research has rarely been guided by their challenges. Future work should address neglected practical needs, including fair data practices, context-specific support, coarse-grained auditing, and fairness-focused debugging.
- Industry practitioners often face technical and organizational barriers even when motivated to improve fairness in their products.
- Fair ML research should support collecting and curating high-quality datasets with fairness in downstream models, not focus only on algorithmic de-biasing.
- Because fairness depends on context and application, practitioners need domain-specific educational resources, metrics, processes, and tools.
- Future auditing methods should accommodate teams that lack individual-level demographic data and instead possess only coarse-grained information.
- Fairness-focused debugging tools should help determine whether individual complaints are isolated incidents or symptoms of broader systemic problems.