Source-linked AI summary
Algorithmic Fairness
Dana Pessach, Erez Shmueli
TL;DR
As AI increasingly makes consequential decisions, the paper addresses how such systems can reproduce bias and why fairness cannot be assumed from automation alone. It surveys causes, definitions, measures, improvement mechanisms, datasets, and emerging subfields, emphasizing scenario-dependent choices and trade-offs. The survey provides a comprehensive, up-to-date overview intended to support researchers and practitioners working on algorithmic fairness.
Problem
AI algorithms can preserve historical and other data-related biases in consequential decisions, making it important to assess and improve their fairness.
Method
The paper surveys causes of unfairness, fairness definitions and measures, enhancement mechanisms and their trade-offs, datasets, and emerging research subfields.
Results
The survey provides a comprehensive and up-to-date overview of algorithmic fairness, including guidance on mechanisms and measures for different settings.
Takeaways & Limitations
The paper supplies researchers and practitioners with knowledge and tools for entering, studying, and applying results from the algorithmic fairness field.
Takeaways & Limitations
Fairness mechanisms and metrics may need to be selected for each scenario because conclusions can differ with missing versus complete information.
Abstract
from arXiv · showhide
An increasing number of decisions regarding the daily lives of human beings are being controlled by artificial intelligence (AI) algorithms in spheres ranging from healthcare, transportation, and education to college admissions, recruitment, provision of loans and many more realms. Since they now touch on many aspects of our lives, it is crucial to develop AI algorithms that are not only accurate but also objective and fair. Recent studies have shown that algorithmic decision-making may be inherently prone to unfairness, even when there is no intention for it. This paper presents an overview of the main concepts of identifying, measuring and improving algorithmic fairness when using AI algorithms. The paper begins by discussing the causes of algorithmic bias and unfairness and the common definitions and measures for fairness. Fairness-enhancing mechanisms are then reviewed and divided into pre-process, in-process and post-process mechanisms. A comprehensive comparison of the mechanisms is then conducted, towards a better understanding of which mechanisms should be used in different scenarios. The paper then describes the most commonly used fairness-related datasets in this field. Finally, the paper ends by reviewing several emerging research sub-fields of algorithmic fairness.
1 INTRODUCTION
AI algorithms increasingly control consequential decisions, but automated decision-making can preserve historical and structural biases rather than guarantee fairness. This survey reviews how algorithmic fairness is defined, measured, improved, and applied across scenarios, while highlighting trade-offs and open challenges.
- Motivation: AI algorithms are increasingly used in decisions across healthcare, transportation, education, hiring, loans, and criminal justice.
- Causes of unfairness: Training data can encode historical biases, so algorithms may reproduce unfair patterns even without intentional discrimination.
- Motivation: Examples include criminal-risk predictions that falsely labeled African-Americans at twice the rate of white people and advertising that favored men for higher-paying executive jobs.
- Fairness trade-offs: Fairness improvement involves an inherent trade-off because pursuing greater fairness may compromise accuracy.
- Scope of the survey: The survey covers fairness definitions and measures, fairness-enhancing mechanisms, datasets, mechanism trade-offs, and emerging research areas.
- Fairness measures: Fairness measures include disparate treatment, disparate impact, group parity measures, and individual fairness based on similarity between people.
- Fairness trade-offs: Demographic parity and disparate impact can judge a fully accurate classifier unfair when groups have different base rates, and may treat similar people differently across groups.
Fairness measures trade-offs
The surveyed fairness measures can be mutually incompatible, requiring application-specific choices rather than simultaneous optimization of every fairness goal. Table 1 summarizes the measures and their definitions.
- Incompatibility: Multiple fairness notions cannot generally be satisfied simultaneously.The incompatibility result excludes only trivial cases in some settings.
- Practical choice: When calibration and equalized odds are incompatible, practitioners should choose one goal according to the application’s requirements.The selected fairness measure should also be evaluated in its legal, social, and ethical context.
- Overview: Table 1 presents the measures and definitions used in this fairness discussion.
Fairness-accuracy trade-off
The literature describes an inherent trade-off between fairness and accuracy. This trade-off has received both theoretical analysis and empirical support.
- Increasing fairness may compromise predictive accuracy.
- The fairness–accuracy trade-off has been studied theoretically and supported empirically across multiple papers.
- Table 1 summarizes the fairness measures discussed in the section, with references provided for further reading.
4 FAIRNESS-ENHANCING MECHANISMS
Fairness-enhancing mechanisms are categorized as pre-process, in-process, or post-process interventions, each with distinct benefits, limitations, and application requirements. Comparisons indicate that no single mechanism consistently dominates; suitability depends on the dataset, fairness definition, and available information.
- Mechanism categories: Fairness mechanisms are typically categorized into pre-process, in-process, and post-process approaches.The paper reviews each category and compares when each type should be used.
- Pre-process mechanisms: Pre-process methods alter labels, weights, or feature representations before training so a subsequent classifier becomes fairer.They can be applied with any classification algorithm but may reduce explainability and create uncertainty about final accuracy.
- In-process mechanisms: In-process methods incorporate fairness during training through regularization terms or constraints tied to measures such as equalized odds or disparate impact.They can explicitly impose the accuracy–fairness trade-off but are tightly coupled to the learning algorithm.
- Post-process mechanisms: Post-process methods modify classifier outputs using decision flips, group-specific thresholds, or separate classifiers for different groups.They can be used with any classifier, but may require sensitive attributes at decision time and can treat similarly featured individuals differently across groups.
- Choosing a mechanism: Method selection depends on ground-truth availability, sensitive-attribute availability at test time, and the fairness definition desired for the application.The paper also identifies a need for more research on robust mechanisms and metrics or scenario-specific choices.
- Comparing mechanisms: Comparative studies found no conclusively dominant method: performance varies across datasets, fairness measures, train–test splits, and data conditions.In-process methods sometimes outperform pre-process methods, while strong under-representation of unprivileged groups can favor pre-process methods.
5 FAIRNESS-RELATED DATASETS
The survey reviews commonly used fairness-related datasets spanning criminal justice, income, credit, and promotion prediction. These datasets differ in population, task, size, features, and sensitive attributes.
- Criminal justice: The ProPublica dataset contains 6,167 individuals from the COMPAS risk assessment system and predicts two-year recidivism.Features include previous felonies, charge degree, age, race, and gender; race or gender may serve as the sensitive attribute.
- Income prediction: The Adult dataset uses 1994 US census data to predict whether annual income exceeds $50,000.Age, gender, and race have been used as sensitive attributes, and preprocessing can reduce the dataset from 48,842 to 45,222 individuals.
- Credit risk: The German dataset contains 1,000 individuals with 20 attributes and predicts good or bad credit risk from financial and demographic features.Gender and age have been used as sensitive attributes.
Ricci promotion dataset
The Ricci dataset contains promotion-exam results for 118 individuals and supports prediction of promotion decisions. Its sensitive attribute is race.
- Dataset description: The Ricci dataset contains exam results for 118 individuals whose promotion decisions are predicted from exam features and current position.The dataset originated from a case brought to the United States Supreme Court.
- Sensitive attribute: Race is the sensitive attribute used in the Ricci promotion dataset.
The Dutch Census dataset
The Dutch Census dataset contains 189,725 individuals and 13 attributes for predicting whether a person holds a highly prestigious occupation. The sensitive feature is gender.
- Dataset description: The Dutch Census dataset includes 189,725 individuals and 13 attributes from the IPUMS repository.Some studies use only 60,420 individuals who are not underaged.
- Task and sensitive feature: The prediction task is whether an individual holds a highly prestigious occupation, using demographic, household, geographic, educational, and economic features.Gender is the sensitive feature utilized in the cited studies.
The Communities and Crimes dataset
The Communities and Crimes dataset is a U.S. community-level benchmark for predicting violent-crime rates, with a race-related sensitive attribute added for fairness research.
- The dataset contains 1,994 U.S. community instances and 128 attributes.
- Its prediction target is violent crimes per 100,000 individuals.
- Features include demographic characteristics such as age, marital status, number of children, and race.
- The dataset is publicly available through the UCI repository.
- A sensitive attribute indicates whether the African-American population exceeds 0.06% of the community.
6 EMERGING RESEARCH ON ALGORITHMIC FAIRNESS
Emerging fairness research extends beyond batch classification to sequential, adversarial, embedding, multimodal, recommender, and causal settings, each introducing distinct fairness definitions and challenges.
- 6.1 Fair Sequential Learning: Sequential-learning fairness must account for feedback loops, evaluating decisions at each step because short-term actions can affect long-term outcomes.
- 6.1 Fair Sequential Learning: Sequential learning requires balancing exploitation of existing knowledge against exploration that gathers data from different populations.
- 6.1 Fair Sequential Learning: A key open challenge is that time-dependent fairness definitions depend on the selected period and discount factors, while exploration itself may be unethical.
- 6.2 Fair Adversarial Learning: Fair adversarial learning uses minimax objectives that improve outcome prediction while reducing the adversary’s ability to infer sensitive attributes.
- 6.3 Fair Word Embedding: Fair word-embedding methods may hide rather than remove gender bias, because an SVM can recover much of the gender information.
- 6.4 Fair Visual Description: Fair visual-description systems must address fairness in both NLP and computer-vision models, alongside inaccurate or stereotyped annotations and unbalanced datasets.
- 6.5 Fair Recommender Systems: Fair recommender systems consider multiple stakeholders, but groups discriminated by more than one sensitive attribute require new fairness definitions and computational improvements.
- 6.6 Fair Causal Learning: Causal fairness models may address challenges in fair prediction, but obtaining the correct causal model is difficult and removing correlated features can reduce accuracy.
7 DISCUSSION AND CONCLUSION
The paper synthesizes algorithmic-fairness concepts, mechanisms, datasets, and emerging sub-fields, then identifies dataset bias, evaluation choices, fairness–accuracy balance, and transparency as open challenges.
- The survey covers causes of unfairness, fairness definitions and measures, enhancement mechanisms, benchmark datasets, and emerging research sub-fields.
- Representative datasets are difficult to achieve because labels may reflect unfair processes and populations or labels may be under-represented.
- The field lacks clear standards for evaluating new mechanisms, including which measures, datasets, and comparison mechanisms to use.
- Interpretability and transparency are important for understanding and trusting algorithmic decisions, and may be legally required in some domains.
- The paper concludes that fairness efforts should address both algorithms and data, potentially integrating humans and algorithms in the decision pipeline.
A ADDITIONAL MEASURES FOR ALGORITHMIC FAIRNESS
The appendix reviews additional algorithmic-fairness measures, including parity, calibration, error-rate, predictive-value, probability-balance, unawareness, dependence, and mean-difference criteria. These measures differ in what they compare across groups and may be incompatible or require assumptions about ground truth and legitimate features.
- Overall accuracy equality requires similar accuracy across groups.
- Predictive parity requires similar positive predictive values across groups, but can be incompatible with equalized odds and equal opportunity when prevalence differs.
- Equal calibration compares groups at each predicted probability value, but is incompatible with equalized odds and insufficient to ensure accuracy or equitable decisions.
- Conditional statistical parity compares selection rates while controlling for legitimate features, whose definition is difficult because features may not be independent of sensitive attributes.
- Predictive equality requires similar false-positive rates across groups, while conditional use accuracy equality compares positive and negative predictive values but does not guarantee equalized odds.
- Treatment equality compares false-negative and false-positive ratios, whereas balance measures compare mean predicted probabilities separately for positive and negative classes.
- Fairness through unawareness excludes sensitive attributes, but proxies, sample bias, and selection bias can still produce discrimination, while excluding information can sometimes harm decisions.
- Mutual information and mean difference assess dependence or prediction-mean gaps without using actual outcomes; lower values indicate better fairness, and mean difference supports binary, categorical, or numerical predictions.