Source-linked AI summary
The Flaws of Policies Requiring Human Oversight of Government Algorithms
Ben Green
TL;DR
Government policies increasingly rely on human oversight to prevent harms from algorithmic decision-making, despite limited empirical grounding for its effectiveness. This paper surveys 41 such policies, finds that oversight often fails and legitimizes flawed algorithms, and proposes institutional oversight requiring agency justification and democratic review before adoption.
Problem
Government algorithms can produce errors, biases, inequitable outcomes, and rigid decisions, while the effectiveness of required human oversight remains insufficiently grounded in empirical evidence.
Method
The paper surveys 41 government-algorithm policies and evaluates their human-oversight requirements against research on human-algorithm interactions.
Results
The vast majority of evidence suggests that people cannot adequately perform envisioned oversight functions, while the assumption of effective oversight legitimizes flawed and unaccountable algorithms.
Takeaways & Limitations
Institutional oversight should require agencies to justify algorithm adoption and proposed human oversight, followed by democratic public review before adoption.
Takeaways & Limitations
Even effective human oversight cannot prevent harms from algorithms that violate human rights, expand surveillance, or entrench inequity.
Abstract
from arXiv · showhide
As algorithms become an influential component of government decision-making around the world, policymakers have debated how governments can attain the benefits of algorithms while preventing the harms of algorithms. One mechanism that has become a centerpiece of global efforts to regulate government algorithms is to require human oversight of algorithmic decisions. Despite the widespread turn to human oversight, these policies rest on an uninterrogated assumption: that people are able to effectively oversee algorithmic decision-making. In this article, I survey 41 policies that prescribe human oversight of government algorithms and find that they suffer from two significant flaws. First, evidence suggests that people are unable to perform the desired oversight functions. Second, as a result of the first flaw, human oversight policies legitimize government uses of faulty and controversial algorithms without addressing the fundamental issues with these tools. Thus, rather than protect against the potential harms of algorithmic decision-making in government, human oversight policies provide a false sense of security in adopting algorithms and enable vendors and agencies to shirk accountability for algorithmic harms. In light of these flaws, I propose a shift from human oversight to institutional oversight as the central mechanism for regulating government algorithms. This institutional approach operates in two stages. First, agencies must justify that it is appropriate to incorporate an algorithm into decision-making and that any proposed forms of human oversight are supported by empirical evidence. Second, these justifications must receive democratic public review and approval before the agency can adopt the algorithm.
1. Introduction
Governments increasingly use algorithms for consequential decisions, while human oversight has become a prominent regulatory safeguard despite limited empirical support. The article surveys these policies, identifies two flaws, and proposes institutional oversight instead.
- Algorithms promise more accurate, fair, and consistent decisions, but practical uses have produced errors, biases, and insufficient attention to individual circumstances.These competing benefits and harms motivate regulatory efforts to govern algorithmic decision-making.
- Human oversight policies permit government algorithms only when humans retain some oversight or control over the final decision.Such policies distinguish permissible algorithm-assisted decisions from solely automated judgments.
- Policymakers rarely ground human oversight requirements in empirical evidence that people actually advance human-rights and dignity goals.Some guidance instead acknowledges risks such as people over-relying on algorithmic advice.
- A survey of 41 policy documents finds that human oversight is not empirically supported and that people generally cannot reliably perform the desired oversight functions.The policies thereby create a false sense of security and shift accountability for algorithmic harms toward lower-level operators.
- The proposed institutional oversight model requires agencies to justify algorithm adoption and proposed oversight, followed by public review and approval before use.Algorithms that violate rights, poorly fit the process, or lack trustworthy evidence should not be adopted.
- Institutional oversight is intended to prevent human oversight from becoming a superficial response to injustices associated with algorithmic decision-making.
2. Discretion, algorithms, and decision-making in government
Government algorithm use reflects a tension between consistency and discretion. Algorithms promise accuracy and standardization, while human discretion is valued for adapting decisions to complex individual circumstances.
- Street-level bureaucrats make consequential decisions in courts, police departments, schools, welfare agencies, and social services.These settings are central sites of controversial government algorithm use.
- Rules promote consistency and predictability, whereas standards permit flexibility and sensitivity to specific cases.Street-level bureaucracies must navigate both forms of decision-making.
- Discretion enables officials to respond to ambiguous, uncertain, and unpredictable situations that rigid formal logic cannot accommodate.Restricting discretion can prevent adaptation to complex or novel circumstances.
- Algorithms attract policymakers and publics because they promise accuracy, objectivity, and consistency.
- Evidence suggests government algorithms can be inaccurate and unfair in criminal justice, education, policing, and welfare.
- Algorithmic rules can conflict with the principle that government decisions should respond to individual circumstances.Unlike human decision-makers, algorithms cannot reflexively adapt to novel situations.
3. Survey of human oversight policies
The surveyed policies use several forms of human oversight to regulate government algorithms. The survey covers 41 documents and identifies restrictions on solely automated decisions, human discretion, and post hoc intervention.
- The survey covers 41 policy documents, including legislation, government guidance, and materials concerning controversial risk assessments.The documents were inductively coded for their principles and mechanisms of human involvement.
- Policies take three approaches to human oversight, ordered from least to most stringent requirements on human involvement.
- Restricting “solely” automated decisions: Twenty of the 41 policies prohibit or restrict decisions made through solely automated means, and all 20 are proposed or passed legislation.
- Restricting “solely” automated decisions: The GDPR exemplifies restrictions on solely automated decisions that produce legal or similarly significant effects.
- Restricting “solely” automated decisions: Twelve of 19 policies explicitly grant subjects of solely automated decisions a right to post hoc human intervention.This allows a person to request human inspection and possible alteration after the automated decision.
- Emphasizing human discretion: Fourteen of the 41 policies require human discretion, including legislation, policy guidance, and risk-assessment materials.
- Emphasizing human discretion: Risk-assessment policies emphasize that supervisors or judges retain discretion and may override algorithmic recommendations.Examples include the AFST, COMPAS, and the Public Safety Assessment.
4. Two flaws with human oversight policies
The article identifies two flaws in human oversight policies: people generally cannot provide the envisioned protections, and the assumption that they can legitimizes flawed algorithms while reducing accountability.
- The vast majority of evidence suggests that people cannot reliably oversee algorithmic errors, biases, and inflexibility.
- Assuming effective human oversight creates a false sense of security and reduces accountability for vendors and policymakers facing algorithmic harms.
4.1. Flaw 1: Human oversight policies are not supported by empirical evidence
Human oversight policies have little empirical grounding, and evidence across domains suggests people cannot reliably detect, correct, or appropriately override algorithmic errors and biases.
- Empirical foundation: Human oversight policies have little basis in empirical evidence, while research suggests people cannot reliably oversee algorithms.The paper evaluates three forms of oversight and finds each unlikely to provide the desired protection against algorithmic errors, biases, and inflexibility.
- Solely automated decisions: Policies restricting solely automated decisions often have no effect where algorithms already operate with human involvement and final decision-makers.High-stakes public-sector settings commonly include human decision-makers, including criminal justice and child welfare.
- Solely automated decisions: Nominal human involvement can bypass regulation by encouraging superficial rubber-stamping of automated decisions.The narrow scope of solely automated decision restrictions makes them easy to avoid through minimal human participation.
- Solely automated decisions: Post hoc human intervention places the burden of seeking review on people after harm has occurred, and remedies may be slow or difficult to obtain.Many affected individuals lack the means or knowledge to request review, while requested remedies can be onerous and delayed.
- Human discretion: People struggle to judge algorithmic output quality and decide when to override recommendations, often discounting accurate advice or relying on inaccurate advice.Research also finds racial bias in how people incorporate algorithmic advice.
- Human discretion: Police and judges can use discretion in harmful ways: police overestimated facial-recognition accuracy, while judges’ overrides increased detention and racial disparities.The cited evidence shows human discretion may add inconsistency and bias rather than correct algorithmic problems.
- Meaningful oversight: Meaningful oversight lacks a defined standard, and its proposed components are either ineffective or difficult to achieve.The paper identifies definitional and functional problems with allowing overrides, requiring understanding, and avoiding reliance on algorithms.
- Meaningful oversight: Explanations and transparency can hinder oversight by increasing trust in incorrect recommendations and reducing error detection.Studies cited in the paper find that explanations may strengthen reliance even when explanations are unrelated to actual algorithmic functioning.
4.2. Flaw 2: Human oversight policies legitimize flawed and unaccountable algorithms in government
Because human oversight does not reliably prevent algorithmic harms, oversight policies can reduce scrutiny, legitimize flawed government algorithms, and shift accountability away from leaders and vendors.
- Policy effects: Human oversight policies reduce scrutiny without reliably reducing algorithmic harms, creating a false sense of security and shifting accountability to frontline operators.The paper characterizes this as a loophole allowing agencies to adopt flawed algorithms while avoiding responsibility for resulting harms.
- Policy effects: Human oversight can provide cover for fundamental concerns about error-prone, biased, and inflexible algorithms, thereby justifying inappropriate government adoption.The proposed safeguard fails to mitigate the underlying concerns that motivated scrutiny of algorithmic decision-making.
- Human overrides: Human overrides cannot remedy the concerns motivating them and instead create the appearance of quality control around flawed and controversial algorithms.Overrides may legitimize tools without addressing their underlying problems.
- Human overrides: People tend to override algorithms detrimentally rather than beneficially, weakening overrides as a safeguard against inaccurate or unfair decisions.The paper presents this pattern as evidence that human discretion does not reliably correct algorithmic shortcomings.
- Human overrides: If policymakers distrust an algorithm enough to require many overrides, the appropriate remedy is to improve or reject the algorithm rather than rely on human discretion.The paper argues that structural reform is necessary when an algorithm remains too flawed to use safely.
- COMPAS example: Judicial discretion in the COMPAS example alleviated concerns about risk assessments without addressing their quality or due-process problems.The paper argues that scrutiny should instead focus on algorithmic quality and whether the tool should affect sentencing.
- Accountability: Oversight policies shift responsibility from agency leaders and vendors to frontline operators who have limited control over system design and political objectives.Operators can become scapegoats for harms produced by system-level decisions.
- Accountability: Governments and vendors can praise algorithms when outcomes are favorable while invoking human oversight to deflect scrutiny when harms occur.This dual rhetoric enables goodwill for algorithmic benefits alongside escape from accountability for algorithmic harms.
5. From human oversight to institutional oversight
The paper argues that human oversight cannot reliably protect against harms from government algorithms and may legitimize unjust systems while weakening accountability. It proposes institutional oversight requiring agencies to justify algorithm use and oversight empirically, followed by democratic review and approval.
- Limits of human oversight: Human oversight policies fail because people cannot reliably perform the scrutiny and balancing tasks required of them.Automation bias persists despite training, and human oversight is difficult when people must quality-control systems adopted for superior predictive performance.
- Limits of human oversight: Human oversight cannot address harms from algorithms that violate rights, expand surveillance, or entrench inequity, even when predictions become more accurate.More effective oversight may further legitimize agencies’ adoption of such systems and diminish accountability for agencies and vendors.
- Institutional oversight: The paper proposes institutional oversight that increases rigor and democratic participation in decisions about whether and how governments use algorithms.The approach places a greater burden on agencies before implementation rather than treating human oversight as sufficient permission.
- Institutional oversight: Institutional analysis should assess algorithm trustworthiness relative to decision stakes, distinguishing validated advice from whether individual decision-makers trust it.The proposed framework combines algorithm trustworthiness with the decision’s need for human discretion when assigning decision-making roles.
- Institutional oversight: After agency justification, a public or democratically accountable body should review and approve the proposed algorithm adoption.This second stage is intended to enable democratic accountability over decisions to adopt algorithms.
- Institutional oversight: Agencies should empirically justify both an algorithm’s appropriateness and the effectiveness of any proposed human oversight before implementation.The default should be skepticism toward human oversight, with agencies required to provide affirmative evidence that it improves outcomes.
6. Conclusion
The study finds that human oversight policies fail to ensure effective oversight and can legitimize flawed algorithms while weakening accountability. It proposes institutional oversight based on agency justification and democratic review before adoption.
- Evidence suggests people cannot adequately perform the oversight functions envisioned by human oversight policies.The study identifies this as the first major flaw in the global policy trend.
- Human oversight policies can legitimize flawed and unaccountable government algorithms instead of preventing their associated risks.The study links this problem to the incorrect assumption that human oversight is effective.
- Institutional oversight would require agencies to justify algorithm adoption and support proposed human oversight with empirical evidence.These justifications would precede an agency’s ability to adopt an algorithm.
- Agencies would also need to make their written justifications public and obtain approval through a democratic review process.The proposed process places public review before algorithm adoption.
- Effective regulation must account for algorithms’ social contexts and empirical implementation evidence rather than relying on intuitively appealing but ineffective rules.The paper frames this as a sociotechnical and evidence-based approach to more democratic and equitable algorithmic governance.
Appendix: Summary of human oversight policies
The appendix summarizes 41 policy documents and classifies them by document type and approach to human oversight. The approaches distinguish restrictions on solely automated decisions, emphasis on human discretion, and requirements for meaningful human input.
- 41 policy documents were reviewed in the study.The table covers legislation, government or government-appointed policy guidance, and selected manuals, policies, and court cases.
- Document Classification distinguishes proposed or passed legislation, policy guidance, and manuals, policies, or court cases from specified US criminal justice settings.The categories are coded 1, 2, and 3, respectively.
- Approach to Human Oversight distinguishes restricting solely automated decisions, emphasizing human discretion, and requiring meaningful human input.The table codes these approaches as 1, 2, and 3 and groups documents accordingly.
Author Information
Ben Green is a Postdoctoral Scholar in the Michigan Society of Fellows and an Assistant Professor in the Gerald R. Ford School.
- Ben Green is a Postdoctoral Scholar in the Michigan Society of Fellows and an Assistant Professor in the Gerald R. Ford School.