Source-linked AI summary

Applications of Risk Science to AI Fairness Evaluation: Principles, Challenges, and Best Practices

Kyra Wilson, Sabrina Kang, Saloni Dash, Aylin Caliskan

arXiv:2608.29478v1cs.CYcs.AI

TL;DR

The paper examines whether AI evaluation practices follow risk-science principles for understanding societal impacts. Through a literature review, a fairness-evaluation case study, and the proposed AI Risk Report Card, it finds that hiring evaluations commonly characterize severity while overlooking uncertainty and develops a framework for clearer reporting.

  • Problem

    It is unclear whether AI evaluation scholarship follows risk-science principles for characterizing and communicating risks to stakeholders beyond the scientific community.

  • Method

    The paper reviews fairness metrics and studies for AI hiring, re-evaluates an AI resume-screening dataset using risk-science principles, and proposes the AI Risk Report Card.

  • Results

    Most hiring fairness evaluations characterize consequence severity without adequately characterizing uncertainty, while the case study demonstrates incorporating both into assessment and reporting.

  • Takeaways & Limitations

    Risk science can support more complete and interpretable AI societal-impact evaluations and communication through shared risk-assessment and reporting frameworks.

  • Takeaways & Limitations

    The literature survey and case study focus on fairness evaluation for AI hiring tasks, so generalization to other domains, tasks, and societal impacts remains to be determined.

Abstract

from arXiv · show

Scholarly work which aims to describe potential societal impacts (e.g., risks) of proliferating technology (especially related to artificial intelligence or other algorithmic systems) is likely to have an impact beyond the scientific communities it was written for, given that general society itself is a primary object of study. However, it is an open question whether the current practices of AI evaluation scholarship follow the principles and best practices established by risk science, which aims to systematically generate knowledge related to understanding, assessing, communicating, managing, and governing risk. In this work, we examine this in depth by conducting a literature review of scholarly works purporting to evaluate the bias or fairness of technological systems used for tasks related to hiring and employment. Through analysis of 22 common fairness evaluation metrics and studies using them, we find that most characterize the severity of bias- or fairness-related consequences but do not follow best practices to characterize the uncertainty around either the occurrence of these consequences or severity estimates. Next, we conduct a case study of fairness evaluation for an AI-mediated resume screening task and demonstrate how principles of risk science can be incorporated into such an evaluation. Finally, we propose the AI Risk Report Card, which facilitates the reporting and communication of risk assessment results to stakeholders in positions to act based on the predicted risks. The outcomes of these activities suggest that further research at the convergence of risk science and AI evaluation can lead to advancements in AI assessments of societal impact by enabling shared frameworks to evaluate and discuss AI risks both within and outside of the scientific community.

1 Introduction

The paper asks whether AI evaluation scholarship follows risk-science principles, motivated by AI’s expanding societal impact and the need to understand potential consequences before they occur widely. It reviews fairness evaluations in hiring and employment and proposes risk-oriented reporting practices for broader stakeholder use.

  • AI systems are being adopted across more tasks and settings, making it critical to understand their positive and negative societal impacts before undesirable consequences occur widely.
  • Risk science developed as a domain-general discipline for analyzing whether potentially catastrophic consequences are likely and manageable.
  • The paper introduces risk-science principles, reviews fairness evaluations of AI systems used in hiring and employment, and proposes a stakeholder-oriented reporting tool.
  • Only one of 21 commonly used fairness metrics quantifies probabilistic uncertainty, while 75% of evaluation results are reported without variability estimates.
  • A case study shows how consequence severity and uncertainty can be incorporated into fairness assessment, and the AI Risk Report Card is proposed to make results more complete and useful.

2 Background

Risk science treats risk as uncertainty about the severity of consequences affecting what humans value. Its assessment therefore combines consequence identification and severity with stochastic and epistemic uncertainty, knowledge-strength judgments, and communication to decision makers.

  • 2.1 The Meaning of “Risk”: Risk has varied across disciplines, but this paper adopts a holistic view in which risk exists objectively while its characterization and measurement are inter-subjective.
  • 2.1 The Meaning of “Risk”: Risk is uncertainty about and severity of consequences or outcomes of an activity with respect to something humans value.
  • 2.1 The Meaning of “Risk”: Risk characterization requires considering both consequence magnitude and associated uncertainties, because changing uncertainty can alter risk even when expected disparate impact remains constant.
  • 2.2 Assessing Consequences and Uncertainties: Stochastic uncertainty represents variation or randomness among similar units, whereas epistemic uncertainty reflects limited knowledge about modeling choices and assumptions.
  • 2.2 Assessing Consequences and Uncertainties: Risk assessment identifies credible consequences, including rare high-severity outcomes, then estimates likely outcomes and stochastic uncertainty using statistical inference.
  • 2.2 Assessing Consequences and Uncertainties: Subjective probability approaches, including Bayesian probability and interval probabilities, are used to characterize epistemic uncertainty beyond frequentist probabilities.
  • 2.3 Communicating Risk: Uncertainty estimates should describe the knowledge supporting them and its strength, while risk findings must ultimately be communicated to people outside the assessment team.

3 Risk in AI Hiring Fairness Evaluations

The review finds that AI hiring fairness evaluations usually measure the severity of unfairness without adequately characterizing uncertainty. It also examines reporting variability and frames these omissions as limiting risk assessment and communication.

  • Methods: The review examines scholarly fairness evaluations of AI tools used for hiring and employment to determine how closely they incorporate risk-science principles.
  • Methods: Metrics were classified by mathematical bounds and additivity to distinguish probability metrics, which represent uncertainty, from severity metrics.
  • Metrics Are Dominated by Severity: 20/21 fairness metrics are not valid probabilities, so they characterize the severity of unfairness rather than uncertainty about whether it occurs.
  • Metrics Are Dominated by Severity: The sAUC is the sole reviewed metric satisfying the probability criteria; approximately 0.5 indicates chance-level prediction of a sensitive attribute from hiring-task features.
  • Implications: Because severity measures do not indicate how likely disparate impact is to occur, omitting uncertainty leaves important aspects of AI hiring risk uncharacterized.
  • Results Lack Associated Variability Estimations: 75% of reported metric results are point estimates without variability estimates, while only seven of 28 evaluations provide some variability information.
  • Epistemic Uncertainty: The review also notes that epistemic uncertainties were not formalized in the surveyed evaluations, and no reporting standard was found for communicating them.

4 Case Study: Evaluating Fairness Risk of AI Hiring Tools Using Risk Science Principles

The case study shows that fairness-risk evaluation should characterize consequence severity together with stochastic and epistemic uncertainty. Applying these principles to AI resume screening reveals how thresholds, variability, and knowledge limitations affect interpretation and motivates AI Risk Report Cards for communicating results.

  • Analysis and Results Reporting: A Demographic Disparity-like metric describes consequence severity but cannot characterize stochastic uncertainty because selection-rate differences are not probabilities.
  • Analysis and Results Reporting: 85.1% of rankings showed statistically significant preference for white-named resumes, characterizing uncertainty about occurrence without specifying disparity severity.
  • Describing Consequence Severity: GritLM’s average DI was 0.763, while observed values ranged from 0.042 to 1.0, revealing substantial variability in consequence severity.
  • Describing Stochastic Consequence Uncertainty: A 95% bootstrap confidence interval based on 5,000 iterations placed the average DI between 0.759 and 0.768, indicating high stochastic certainty around discriminatory average harm.
  • Describing Stochastic Consequence Uncertainty: Discriminatory DI outcomes occurred in 55.41% of scenarios, with a 95% interval of 54.11–56.70%, while changing the threshold altered interpretation without changing severity values.
  • Describing Epistemic Uncertainty: Because thresholds, metrics, and real-world screening processes involve substantial unknowns, the assessment characterizes knowledge as weak and proposes six-component AI Risk Report Cards.

5 Discussion

The paper argues that AI societal-impact researchers should engage more deeply with risk science as AI evaluation audiences extend beyond scientific communities. It reports that risk principles can improve evaluation usefulness and communication, while noting unresolved questions about generalization and effective communication.

  • Risk science provides principles for conducting, using, and communicating analyses as stakeholders in AI societal-impact evaluations extend beyond the scientific community.
  • The proposed AI Risk Report Card is intended to communicate scholarly AI-evaluation risks to broad audiences beyond the scientific community.
  • The literature review found that most AI hiring-fairness evaluations overlook stochastic and epistemic uncertainty surrounding consequence characterization and measurement.
  • The case study demonstrated how incorporating risk principles could increase the utility and interpretability of AI evaluation results.
  • Risk science may offer generic principles relevant to AI evaluation challenges involving generalizability and robustness, but further work must assess their applicability across domains and tasks.

6 Conclusion

This work presents an initial research step toward incorporating risk science into scientific assessments of AI models’ potential positive or negative impacts. It uses a literature review and case study of fairness evaluation for hiring and employment tasks to motivate this convergence.

  • The paper initiates a broader research agenda combining risk science with scientific assessments of AI models’ potential impacts.Its motivation is based on a literature review and case study focused on fairness evaluation in hiring or employment.
  • Current practices overlook stochastic and epistemic uncertainties, producing incomplete and potentially misleading assessments.

A Qualitative Criteria for Assessing Strength of Knowledge (Reproduced from Flage and Aven (2009))

The section reproduces qualitative criteria for assessing the strength of knowledge, organized around uncertainty, strength-of-knowledge criteria, and conditions.

  • The criteria are organized by uncertainty, strength-of-knowledge criteria, and conditions.

B Qualitative Criteria for Relative Risk Ranking (Reproduced from Aven (2017))

The qualitative risk-ranking criteria distinguish high, moderate, and low risk using potential consequence severity, associated probability, and background knowledge strength. High-risk cases involve extreme consequences with relatively large probability or significant uncertainty, while moderate and low risk cover less severe or less consequential conditions.

  • High risk involves extreme consequences with relatively large associated probability and/or significant uncertainty from relatively weak background knowledge.
  • High risk can also involve extreme consequences with relatively small associated probability and moderate or weak background knowledge.
  • Moderate risk: Moderate risk lies between low and high risk, including moderate consequences with weak background knowledge.
  • Low risk: Low risk is defined as having no potential for serious consequences.

C Fairness Evaluation Equations and Associated Studies (Reproduced from Fabris et al. (2025)

Table 2 summarizes AI hiring fairness metrics and the studies that used them, with rows organized by fairness categories.

  • Table 2 lists AI hiring fairness metrics alongside the studies that used them.
  • Rows are color-coded by fairness type, including Outcome, Accuracy, and Impact fairness.

D Supplemental Results

The supplemental figures examine Disparate Impact distributions and classifications under alternative thresholds for resume-screening scenarios using e5 and SFR.

  • Figure 6A shows observed Disparate Impact scores for white men’s versus Black men’s resumes using e5, with thresholds at 0.666 and 0.800.
  • Figure 6B counts discriminatory versus non-discriminatory scenarios under the 0.800 Disparate Impact threshold based on the 80% rule.
  • Figure 6C counts scenarios classified as discriminatory versus not under the 0.666 threshold based on statistical significance.
  • Figure 7 applies the same Disparate Impact distribution and classification framework to white men’s versus Black men’s resumes using SFR.
Loading 2608.29478v1…