Source-linked AI summary

Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices

Manish Raghavan, Solon Barocas, Jon Kleinberg, Karen Levy

arXiv:1906.09208v3cs.CYcs.AIcs.LG

TL;DR

Algorithms are increasingly used in hiring to address bias, yet little is known about how these assessments are built, validated, and examined in practice. The paper reviews vendors’ public claims and practices from technical and legal perspectives, finding heterogeneous approaches shaped by data, prediction targets, validation constraints, and antidiscrimination law.

  • Problem

    Little is known about how algorithmic employment assessments are developed, validated, and examined for bias, especially when models and data are not publicly accessible.

  • Method

    The authors systematically identify assessment vendors and analyze their public disclosures about development, validation, bias mitigation, data collection, prediction targets, and legal context.

  • Results

    Vendor practices are heterogeneous and reveal choices and trade-offs involving data collection, representative validation, prediction targets, and discrimination law.

  • Takeaways & Limitations

    Public statements can yield valuable information about proprietary assessment systems and support questions about their design and deployment contexts.

Abstract

from arXiv · show

There has been rapidly growing interest in the use of algorithms in hiring, especially as a means to address or mitigate bias. Yet, to date, little is known about how these methods are used in practice. How are algorithmic assessments built, validated, and examined for bias? In this work, we document and analyze the claims and practices of companies offering algorithms for employment assessment. In particular, we identify vendors of algorithmic pre-employment assessments (i.e., algorithms to screen candidates), document what they have disclosed about their development and validation procedures, and evaluate their practices, focusing particularly on efforts to detect and mitigate bias. Our analysis considers both technical and legal perspectives. Technically, we consider the various choices vendors make regarding data collection and prediction targets, and explore the risks and trade-offs that these choices pose. We also discuss how algorithmic de-biasing techniques interface with, and create challenges for, antidiscrimination law.

1 Introduction

Algorithmic hiring is promoted as a way to mitigate bias, but formal fairness language can lend vague de-biasing claims undue credibility. Because models and data are usually inaccessible, this study examines vendors’ public disclosures to characterize their practices, choices, and unresolved concerns.

  • Formal fairness definitions and metrics can paradoxically give vague claims about “de-biasing” and “fairness” undue credibility.
  • Limited public access to hiring models and sensitive employee data makes direct empirical characterization of industry practices difficult.Models and data are generally kept private, and auditing outputs or training data can create privacy risks.
  • The study identifies 18 pre-employment-assessment vendors and documents their disclosed development, validation, and bias-removal practices.It uses these disclosures to assess industry attempts to address bias and identify critical issues left unaddressed.
  • The analysis examines vendor choices about target variables, training data, representative validation populations, and the interaction between bias prevention and discrimination law.These choices create trade-offs, while observed practices remain heterogeneous and lack clear guidance.
  • The paper analyzes vendors’ website statements for consistent comparison, while noting that other public sources could provide additional details.
  • The authors aim to provide a realistic picture of algorithmic pre-employment assessment and recommendations for effective and appropriate use, rather than an exposé.

2 Background

Pre-employment assessments occupy the screening stage of consequential hiring pipelines and are evaluated through psychological validity concepts and legal discrimination standards. Their history includes persistent equity concerns, while modern algorithmic tools introduce new assessment methods that research and regulation are still addressing.

  • Hiring pipelines comprise sourcing, screening, interviewing, and selection; this paper focuses on algorithmic pre-employment assessments used for screening.
  • Academic research has struggled to keep pace with rapidly evolving assessment technology, allowing vendors to extend practices without rigorous independent research.
  • Cognitive assessments have historically disadvantaged minority populations despite evidence that minority candidates can have similar real-world job performance.
  • The APA describes predictive bias as systematic over- or under-prediction for a group and rejects equal outcomes as a requirement for fairness.
  • Validity is the legal gold standard for assessment: test outcomes should say something meaningful about a candidate’s potential as an employee.Evidence may be criterion, content, or construct validity.
  • Employment assessments can face disparate-treatment or disparate-impact challenges under Title VII and the Uniform Guidelines.Disparate treatment concerns explicit differential treatment, while disparate impact concerns more nuanced differences in selection outcomes.
  • The 4/5 rule flags potential disparate impact when one protected group’s selection rate is below four-fifths of the highest group’s rate.Employers may defend significant disparities by showing that selection procedures are valid and business-necessary.

3 Empirical Findings

The study reviewed publicly disclosed practices of 18 vendors offering algorithmic pre-employment assessments, focusing on assessment design, validation, and bias claims. Vendors commonly emphasized outcome equality and the 4/5 rule, while validation procedures and de-biasing details were unevenly disclosed.

  • Vendor landscape: 18 vendors provided algorithmically driven pre-employment assessments, with most offering small, often U.S.-based organizations.Among 16 vendors with available funding information, funding ranged from around $1 million to $93 million; 14 had 50 or fewer employees and 9 were based in the United States.
  • Assessment types: Questions were the most common assessment type, followed by video interview analysis and gameplay.Eleven vendors offered questions, while 6 offered video interview analysis and 6 offered gameplay; many vendors offered multiple assessment types.
  • Target variables and training data: 15 vendors offered custom or customizable assessments, and 8 built assessments using clients’ past and current employee data.Clients generally determined the outcomes to predict, including performance reviews, sales numbers, and retention time.
  • Validation: Vendor websites generally did not clearly disclose whether models were validated, which methodologies or data were used, or how validation was tailored to clients.Good & Co. provided psychometric validation and demographic score audits but did not provide comparable documentation for its algorithmic “culture fit” recommendations.
  • Accounting for bias: 15 vendors referred abstractly to bias, but only 7 explicitly discussed compliance or adverse impact, including 3 that mentioned the 4/5 rule.HireVue and pymetrics described removing or downweighting features correlated with protected attributes when adverse impact was detected; other vendors made less detailed claims.
  • Accounting for bias: All examined vendors with concrete bias claims focused on equality of outcomes and compliance with the 4/5 rule.Some vendors claimed naturally unbiased assessments with similar score distributions across demographic groups, while others iteratively modified models or data after detecting adverse impact.

4 Analysis of Technical Concerns

The paper examines how data choices, prediction targets, and assessment formats create technical challenges for algorithmic hiring. It contrasts custom and pre-built assessments, emphasizing trade-offs between broad validation and client-specific fit.

  • Data Choices: Data sources and target variables are interdependent choices that can introduce bias into the machine-learning pipeline.Vendors may rely on client employee data and outcomes such as performance reviews, sales numbers, or retention.
  • Data Choices: Client data often limits models to existing employees, providing no direct evidence about how rejected applicants would have performed.The available data may also be incomplete or inaccessible, constraining the practitioner’s options.
  • Data Choices: Predicting future employee success from current employees can reproduce existing hiring patterns, while job evaluations may themselves be biased against minorities.The target variable therefore embeds assumptions about what counts as success and whose performance is represented.
  • Data Choices: Team-fit models can be difficult to audit because each narrowly customized model raises separate questions about what bias means and how it should be assessed.Some vendors explicitly seek candidates who fit existing employees or teams.
  • Assessment Formats: Pre-built assessments can use larger and more diverse populations, but generic constructs and differing job contexts may limit their fit for particular clients.Traits such as “grit” or “openness” lack a single objective benchmark, and representative validation populations vary across locations, companies, and roles.
  • Assessment Formats: Pooling data across companies or locations reduces small-sample variance but can bias conclusions away from a client’s specific needs.The paper frames this as a domain-adaptation and bias-variance trade-off with no clear universal best practice.
  • Assessment Formats: Game- and video-based assessments can discover predictive relationships automatically, making theoretical justification and detection of indirect protected characteristics more difficult.Rich inputs may encode legally proscribed characteristics, and facial-analysis systems have shown gender- and racial disparities in error rates.

5 Algorithmic De-Biasing

The paper analyzes algorithmic de-biasing through the legal framework of disparate impact, focusing on vendors’ reliance on outcome parity and the 4/5 rule. It shows that technical efforts to reduce disparities can create tensions with validation requirements, disparate-treatment law, and broader assessments of harm.

  • Legal Context: Vendors making concrete de-biasing claims generally frame them around equal outcomes and compliance with the 4/5 rule.The analysis examines how this focus interacts with the stages of a disparate-impact lawsuit.
  • Legal Context: Under disparate-impact doctrine, employers may defend a disparity by demonstrating validity and business necessity, but may remain liable if an avoidable alternative exists.The rule does not prohibit every disparity; it targets unjustified or avoidable disparate impact.
  • Legal Context: Reliance on the 4/5 rule can reduce the practical need for employers to establish business necessity through a rigorous validation process.Clients may still request validation to support the goal of selecting qualified candidates.
  • De-biasing Methods: Algorithmic de-biasing may become an alternative business practice when it reduces adverse impact without significantly harming predictive ability and at trivial cost.Failure to adopt such a technique could expose an employer to liability under the paper’s analysis.
  • De-biasing Methods: Within-group scoring equalizes selection rates by reporting percentiles within each group and preserves within-group rank order, unlike feature-removal approaches.The paper identifies this as a theoretically optimal way to equalize selection rates, while noting its legal history.
  • De-biasing Methods: Within-group reporting could satisfy the 4/5 rule but was prohibited because it was treated as disparate treatment under U.S. civil-rights law.The paper uses this example to illustrate tension between disparate-impact and disparate-treatment doctrines.
  • Limitations of Outcome Measures: Outcome disparities depend on the population and job context used for evaluation, making representative samples crucial for validation.Selection-rate comparisons can differ across applicant populations, locations, and roles.
  • Limitations of Outcome Measures: Satisfying the 4/5 rule does not replace analysis of how bias and harm arise throughout the assessment-development pipeline.The paper calls for examining inputs, outputs, feature justification, ranking effects, and other system-level concerns.

6 Discussion and Recommendations

The paper argues that evaluating algorithmic hiring requires examining vendors’ practices alongside the legal, historical, and social context of deployment. It recommends greater transparency, broader bias assessment, and attention to the limits and legal implications of de-biasing.

  • Discussion: Public statements can provide valuable evidence about proprietary hiring systems when traditional external audits are infeasible.The analysis uses vendors’ practices to raise relevant questions about models deployed in real-world settings.
  • Discussion: Technical systems should be analyzed together with the legal, historical, and social influences surrounding their use and deployment.The authors argue that these surrounding contexts are necessary for understanding vendors’ design decisions.
  • Recommendations: Additional transparency is necessary to craft effective policy and enable meaningful oversight of algorithmic hiring systems.Vendors are generally not forthcoming about their practices, with some exceptions.
  • Recommendations: Disparate impact is not the only indicator of bias; vendors should also monitor metrics such as differential validity.The recommendation broadens bias evaluation beyond selection-rate disparities.
  • Recommendations: Outcome-based bias measures require representative applicant datasets, do not critically examine individual predictors, and depend on protected attributes that may be unavailable.These limitations constrain the power of measures including disparate impact and differential validity.
  • Legal implications: Machine learning may discover ethically problematic correlations even when an assessment is statistically valid under existing legal standards.The paper therefore questions whether validity standards in the Uniform Guidelines adequately address machine-learning assessments.
  • Legal implications: Algorithmic de-biasing automates the search for less discriminatory alternatives, creating implications for alternative business practices and antidiscrimination law.The paper recommends vendor exploration of these techniques and clearer guidance from the EEOC.

A Administrative Information on Vendors

The supplied table is identified only as providing administrative information about vendors.

  • Table 3 is labeled “Administrative information,” but the supplied passage contains no administrative findings to summarize.
Loading 1906.09208v3…