Source-linked AI summary

The Fallacy of AI Functionality

Inioluwa Deborah Raji, I. Elizabeth Kumar, Aaron Horowitz, Andrew D. Selbst

arXiv:2206.09511v2cs.LG

TL;DR

The paper argues that AI policy and research often overlook whether deployed systems function or provide benefits, despite documented failures. Using case studies, it develops a taxonomy of functionality issues and identifies policy and organizational responses that can help protect affected communities.

  • Problem

    AI research and policy often prioritize ethical or value-aligned deployment without first establishing whether systems function or provide benefits.

  • Method

    The paper analyzes case studies to construct a taxonomy of AI functionality issues and examines overlooked policy and organizational responses.

  • Results

    The case studies document functionality failures including impossible tasks, engineering failures, deployment failures, and communication failures.

  • Takeaways & Limitations

    Functionality is a meaningful AI policy challenge and a necessary first step toward protecting affected communities from algorithmic harm.

  • Takeaways & Limitations

    Self-regulatory approaches are constrained by clear conflicts of interest, while some reported model failures cannot be established with certainty.

Abstract

from arXiv · show

Deployed AI systems often do not work. They can be constructed haphazardly, deployed indiscriminately, and promoted deceptively. However, despite this reality, scholars, the press, and policymakers pay too little attention to functionality. This leads to technical and policy solutions focused on "ethical" or value-aligned deployments, often skipping over the prior question of whether a given system functions, or provides any benefits at all. To describe the harms of various types of functionality failures, we analyze a set of case studies to create a taxonomy of known AI functionality issues. We then point to policy and organizational responses that are often overlooked and become more readily available once functionality is drawn into focus. We argue that functionality is a meaningful AI policy challenge, operating as a necessary first step towards protecting affected communities from algorithmic harm.

1 INTRODUCTION

Deployed AI systems frequently fail in consequential ways, yet public and policy discussions often assume they function. The paper makes functionality a primary AI policy concern and develops a case-based taxonomy of failures and related interventions.

  • Deployed AI systems misclassify, misallocate, misread, and misdiagnose, producing harms in employment, housing, education, policing, healthcare, and resource distribution.Examples include false fraud flags, wrongful arrests, lost benefits, unsafe-content flags, and medical errors.
  • 100% error rate was reported for New York MTA’s facial-recognition pilot, yet the program proceeded.
  • These functionality failures can disproportionately affect minoritized groups, darker-skinned women, Black patients, and lower-income patients.
  • Scholars and policymakers often assume deployed systems work, causing accountability proposals to overlook functionality issues and functional safety’s role in harm.
  • Functionality assessments can be empirically measured despite contested stakeholder expectations, grounding claims about harm.
  • The paper treats functionality as a primary concern, catalogs failure forms and harms through AAAIRC case studies, and discusses overlooked accountability tools.

2 RELATED WORK

Prior work recognizes engineering, evaluation, marketing, and deployment problems in AI, but functionality has not been systematically treated as its own range of problems.

  • Existing literature identifies a broad AI functionality problem, yet little work systematically discusses functionality-specific problems.
  • AI research faces scientific-validity and evaluation problems, including reproducibility failures and advances that weaken under closer scrutiny or broader application.
  • ML systems are difficult to engineer because practitioners must adapt software practices, increase testing effort, and define requirements under distinctive challenges.
  • The AI label can support inflated functionality claims, while some supposedly intelligent public-sector systems use manually crafted heuristics and constrained, context-specific data.
  • Commercial and political narratives have been criticized as hype, including the “snake oil” metaphor and the “criti-hype” pattern in technology criticism.
  • Recommendation systems and Cambridge Analytica’s product were described as having less behavioral or predictive ability than their public claims suggested.

3 THE FUNCTIONALITY ASSUMPTION

AI discourse and policy frequently presuppose that systems work, emphasizing trust, fairness, explainability, democratization, or speculative advanced capabilities before basic functionality is established. The paper argues that functionality should be examined first because many practical harms arise from systems that do not work as expected.

  • The functionality assumption appears when AI discourse treats systems as working and focuses on trust, fairness, explainability, democratization, or advanced misuse.
  • Functionality is rarely explicit in AI principles and policy proposals, and when considered it often receives less emphasis than other concerns.
  • Many trust guidelines emphasize public confidence and expectations, placing responsibility on people to trust rather than institutions to make systems reliably operational.
  • Policy discussions of democratization expanded access to AI tools and data while some COVID-19 systems’ functionality and utility remained untested.
  • Fears about runaway systems, alignment, and misuse presume that AI can execute declared objectives or produce highly capable outputs, while practical limitations may be neglected.
  • The EU draft regulation addresses functionality to some degree, but its prominent concerns were characterized as bordering on the fantastical.
  • Fairness audits may focus on the 4/5ths selection-rate rule while rarely discussing product validation, presuming the prediction task is valid.

4 THE MANY DIMENSIONS OF AI DYSFUNCTION

The taxonomy presents AI dysfunction as a broad set of failure points spanning impossible tasks, implementation, deployment, interaction, adversarial conditions, communication, and inadequate safeguards. Case studies show that these failures can produce unsafe outputs, wrongful decisions, wasted labor, and harm to affected communities.

  • Taxonomy and scope: The taxonomy consolidates disparate AI product failures and provides language for research, policy, and intervention discussions.The authors identify many failure points and use case studies to connect examples with distinct causes or elements of failure.
  • Implementation failures: Synthetic oncology data produced unsafe and incorrect treatment recommendations, while fragmented leukemia records impeded reliable extraction of therapy timelines.These cases illustrate failures arising from unsuitable data and difficulty handling time-dependent clinical information.
  • Impossible tasks and opaque decisions: 97% of UK tests were flagged as suspicious, including 58% classified invalid and 39% questionable, before visa cancellations began.Error-rate estimates ranged from 1% to 30%, while only 3,600 of 12,500 appeals succeeded.
  • Evaluation and deployment failures: Poor calibration across hospitals and discrimination against under-represented demographics suggested insufficient pre-deployment testing of the Epic sepsis model.The cited review also found high risk of bias, unavailable datasets or code in over 90% of studies, and suboptimal reporting standards.
  • Safety and operational failures: Missing safeguards created operational hazards, including accidental deactivation of a carbon-monoxide detector and fatal crashes involving Uber and Tesla vehicles.The NTSB cited absent warnings, remote monitoring, and inadequate safety culture; these safeguards were considered more serious hazards than some engineering failures.
  • Robustness and interaction failures: Changing formats, negation, adversarial actions, and unsuitable facial-recognition probe images can induce failures after deployment.LVMPD used non-suitable probe images in almost half of its searches, increasing the chances of false identification.

5 DEALING WITH DYSFUNCTION: OPPORTUNITIES FOR INTERVENTION ON FUNCTIONAL SAFETY

The paper frames AI functionality failures as a recurring safety challenge and identifies existing governance infrastructure in related industries as a source of possible interventions.

  • AI faces a challenge resembling earlier industries that experienced fraudulent or dysfunctional products and later developed governance responses.Examples include food safety, medicine, finance, aviation, and automobiles.
  • Healthcare and transportation offer relevant governance models through rigorous evaluation processes and thorough accident investigations.The paper highlights healthcare evaluation and transportation investigations as existing infrastructure.
  • The paper outlines legal and organizational interventions for addressing functionality issues in the contexts where AI is developed and deployed.
  • Functionality failures can cause harm when systems are deployed without working well.The paper connects this concern to functional safety in engineering design.

5.1 Legal/Policy Interventions

The paper argues that existing consumer-protection, products-liability, and administrative tools can address dysfunctional AI, although products-liability doctrine faces uncertainty for software.

  • Consumer Protection: Consumer-protection law provides the primary legal category for addressing products that fail to work correctly.The discussion is U.S.-based, while analogous tools exist in most jurisdictions.
  • Consumer Protection: Misleading functionality claims can support deceptive-practices claims, including claims about capabilities a product cannot achieve.The paper notes that implied design claims may also be misleading.
  • Consumer Protection: Foreseeable and harmful post-deployment failures may fall within unfair-practices authority even when external actors contribute to them.
  • Regulatory Authorities: Federal, state, and subject-specific agencies provide additional routes including standards, certifications, investigations, bans, recalls, and enforcement.Relevant bodies include the FTC, consumer-product, transportation, and consumer-finance regulators, as well as state authorities.
  • Products Liability: Products liability could apply to basic AI functionality failures, especially when the error was foreseeable and a working alternative should have been built.However, defective software has not produced a products-liability verdict, and courts have not clearly established software as a product for these purposes.
  • Products Liability: Failures involving construct validity or arbitrary decisions can make legal cases easier by undermining claims that AI decisions rest on a sound basis.

5.2 Organizational interventions

Organizational interventions include audits, documentation, certification, standards, procurement requirements, safety culture, and continuous monitoring, but self-regulation has important limits.

  • Limits of Self-Regulation: Self-regulatory approaches are constrained by conflicts of interest, although they provide an immediate path for addressing functionality issues.
  • Certification & Standards: Aerospace processes such as FMEDA, FHA, and FDALS could help identify functional-safety issues before AI deployment.
  • Internal Audits & Documentation: Internal audits and documentation can improve quality control, deployment evaluation, and engineering reflection on functionality.Independent parties may provide a fresh perspective, while documentation can support internal quality-assurance standards.
  • Internal Audits & Documentation: Audit results are rarely disclosed externally and are not mandatory or incentivized, so assessments mainly serve internal use.
  • Certification & Standards: Certification and standards audits can raise deployment criteria and mature the product-development pipeline.Possible formats range from documentation exercises to challenge datasets used as benchmarks.
  • Industry Coordination: Industry-wide standards and collective decision-making can raise expectations and awareness around functionality challenges.The paper points to standards organizations, industry groups, and model-documentation efforts as templates.
  • Other Interventions: Procurement requirements, internal safety culture, logging, evaluation, and continuous monitoring offer additional organizational levers.These measures can set functionality expectations and support internal inspection of deployed systems.

6 CONCLUSION : THE ROAD AHEAD

The conclusion argues that AI functionality must be assessed and communicated before mass deployment because poorly vetted products create overlooked harms and accountability gaps.

  • Assuming that AI products work can obscure harms and legal or organizational remedies for functionality failures.
  • Limited attention to functionality leaves these issues inadequately emphasized and poorly addressed by available accountability tools.
  • Faulty AI products already on the market make functionality failures an urgent policy problem, while claimed benefits often go unchallenged.
  • The paper calls for developers to understand, explore, and communicate the limits of their products more honestly.
  • Adequate functionality assessment and communication should be minimum requirements for mass deployment, and nonfunctional products should not affect people’s lives.
Loading 2206.09511v2…