Source-linked AI summary

The Normalization of Deviance in AI Development

Emilio Barkett, Alexander Kimpton, Daniel Graham, Yusuf Kundgol

arXiv:2609.05749v1cs.AIcs.ETeess.SY

TL;DR

AI risk research has largely emphasized dangerous system capabilities, leaving less attention to whether AI-building organizations themselves drift toward failure. Using historical case studies and mapping their recurring mechanisms onto AI development, the paper argues that existing safety infrastructure may not prevent catastrophe when organizations can satisfy formal processes while those mechanisms operate through them.

  • Problem

    Research has focused predominantly on capability risk, with less attention to whether the institutions building AI are structurally predisposed to drift toward failure.

  • Method

    The paper compares the Challenger, Three Mile Island, and Boeing 737 MAX disasters to identify common structural mechanisms and maps them onto contemporary AI development.

  • Results

    The paper finds that AI development exhibits the historical mechanisms of production pressure, false assurance, structural secrecy, and weakened independent oversight, while safety processes can still be completed.

  • Takeaways & Limitations

    The pre-disaster period of AI development remains underway, making these organizational dynamics legible while they may still be interrupted.

  • Takeaways & Limitations

    The historical analogy is consistent with normalization of deviance in AI development but is not proof, because the full evidentiary record becomes visible only after failure.

Abstract

from arXiv · show

Work on the risks of artificial intelligence has focused predominantly on capability risk: the danger that systems become too powerful, too autonomous, or too misaligned with human values. Far less attention has been paid to the organizational level---to whether the institutions building these systems are themselves predisposed to drift toward failure. This paper argues that they are. Regardless of how capable AI systems become, the organizations building them face the same structural dynamics that preceded past major technological disasters. Drawing on case studies of the Space Shuttle Challenger, the Three Mile Island accident, and the Boeing 737 MAX crashes, this paper identifies the common structural mechanisms preceding each failure and maps them onto contemporary AI development. The findings suggest that existing safety infrastructure may provide less protection than it appears, as organizations can complete safety processes in full compliance and still produce catastrophic outcomes. The pre-disaster period of AI development is still underway; the purpose of this paper is to make these dynamics legible while they can still be interrupted.

1 Introduction

AI risk research has focused mainly on dangerous system capabilities, while this paper examines whether the organizations building AI are structurally predisposed to drift toward failure. It argues that safety processes can be completed while organizations still move toward catastrophic outcomes.

  • AI risk research has focused predominantly on systems becoming too powerful, autonomous, or misaligned to control.
  • The paper shifts attention to whether AI-building organizations are structurally predisposed to drift toward failure through organizational dynamics rather than negligence or malice.
  • Existing safety infrastructure may provide less protection than it appears because organizations can complete safety processes and still drift toward failure.

2 Background

The paper frames normalization of deviance as organizational drift in which repeated non-disaster makes safety-norm violations appear normal. Four mechanisms recur across organizations, especially where emerging technologies lack mature safety standards.

  • Normalization of deviance describes incremental organizational drift in which repeated non-disaster redefines a safety-norm violation as normal.
  • The framework does not require villains: competent, experienced decision-makers can sincerely believe a system is safe when organizational structures supply misleading information and interpretive frameworks.
  • Production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight are four recurring mechanisms of normalized deviance.
  • These mechanisms apply especially forcefully to emerging technologies because safety standards develop alongside deployment without a stock of operational experience.
  • A review of 33 studies across five sectors found evidence of these dynamics in every sector examined.

3 Case Studies

The case studies show organizations drifting toward disaster as warning signals are reinterpreted, workarounds become routine, and oversight loses independence. Challenger, Three Mile Island, and Boeing 737 MAX illustrate distinct mechanisms within a common pattern.

  • The cases were selected for extensive documentation across independent industries and decades, establishing how normalization dynamics operate rather than how often they produce failure.
  • 3.1 The Challenger Launch Decision (1986): Challenger broke apart 73 seconds after launch, killing all seven astronauts after engineers’ concerns about O-ring failure were overruled.
  • 3.1 The Challenger Launch Decision (1986): At Challenger, repeated flights without catastrophe redefined O-ring erosion as acceptable while production pressure shifted the burden of proof onto those urging caution.
  • 3.2 The Accident at Three Mile Island (1979): At Three Mile Island, years of minor anomalies were handled through improvised workarounds that gradually calcified into standard practice.
  • 3.2 The Accident at Three Mile Island (1979): Three Mile Island’s complex, opaque system produced contradictory indicators and a widening gap between formal procedures and actual practice, preventing safety information from reaching decision-makers.
  • 3.3 The Boeing 737 MAX Crashes (2018–2019): The 737 MAX crashes followed MCAS repeatedly pushing the aircraft nose down in response to faulty sensor data, while the system compensated for a physical design flaw.
  • 3.3 The Boeing 737 MAX Crashes (2018–2019): Boeing chose modification over a new aircraft to reduce time, cost, and simulator-training requirements, after which MCAS gained expanded authority and dependence on a single sensor.
  • 3.3 The Boeing 737 MAX Crashes (2018–2019): FAA oversight eroded as Boeing employees performed much of the certification work and FAA managers overruled engineers’ concerns.

4 Normalization of Deviance in AI Development

AI development exhibits the same four mechanisms identified in the historical cases, intensified by immature institutions, opaque systems, and unusually strong competitive pressure. Safety information can be documented yet fail to influence deployment decisions, while developers retain substantial control over standards and oversight.

  • AI development combines production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight while frontier laboratories remain relatively young institutions.
  • Production pressure: Commercial and state competition makes safety work appear to hinder speed, shifting the burden of proof toward those arguing for caution.
  • False assurance from prior success: Benchmark performance and incident-free deployments can be treated as safety evidence, progressively raising the implicit safety baseline for later systems.
  • Structural secrecy: Division among safety, product, and leadership teams can prevent safety concerns and red-team findings from reaching or influencing deployment decisions.
  • Erosion of independent oversight: Responsible scaling policies and model evaluations are defined, conducted, and enforced by the same organizations whose deployment decisions they are meant to constrain.

5 Recommendations

The paper recommends structural interventions because normalization of deviance operates through ordinary rule-following, making additional compliance requirements insufficient on their own. Its recommendations address organizational independence, epistemic culture, external governance, and technical opacity.

  • Overall recommendation: Adding rules, compliance requirements, or accountability mechanisms alone is unlikely to prevent normalization because the process operates through rule-following.The recommendations therefore target the structural conditions producing the four mechanisms.
  • Organizational: Safety functions require structural independence from production incentives that can pressure organizations to minimize safety friction.The case studies identify dependence on production management or commercial interests as a recurring weakness.
  • Epistemic and cultural: Organizations should explicitly recognize normalization of deviance and treat prior successful deployments as precedent rather than proof of safety.Each more capable system should receive affirmative safety evidence independent of earlier deployment history.
  • Regulatory and governance: Strong internal safety culture requires external, independent oversight because industry self-regulation can legitimize overrides under commercial pressure.The paper proposes third-party evaluation infrastructure independent of developer cooperation and funding.
  • Technical: Opacity can drive normalization when operators cannot form accurate mental models, so interpretability is treated as a safety prerequisite.The paper connects this problem to both Three Mile Island operators and AI evaluators assessing deployment safety.

6 Limitations

The paper’s central analogy between AI development organizations and organizations behind past technological disasters is conceptually useful but empirically and structurally limited. Its evidence cannot yet distinguish normalization from organizational learning, and AI failure modes differ from those of physical engineering systems.

  • Evidentiary limitation: The paper’s central structural analogy is not proof that normalization of deviance is occurring within AI organizations.The historical cases establish organizational dynamics, but the AI claim remains preliminary and requires empirical evidence.
  • Evidentiary limitation: The framework’s full evidentiary record becomes available only after the failure it predicts, limiting pre-disaster verification.Current indicators, including policy revisions and researcher departures, are consistent with both normalization and organizational learning.
  • Analogy boundaries: The analogy is imperfect because AI failure modes are diffuse, less measurable, and harder to distinguish from normal system behavior than physical-engineering failures.The paper notes that the original framework concerned relatively legible failures such as inspectable O-rings.
  • Field-specific boundary: AI safety organizations are more self-aware about these risks than NASA, the NRC, or Boeing were at comparable stages.The field has recognized safety disciplines, substantial safety teams, and published frameworks for organizational risk.
  • Field-specific boundary: Awareness of normalization risks may be insufficient when the organizational structures producing those risks remain intact, and this distinction is difficult to verify empirically.The paper does not claim that AI organizations are indifferent to safety.
  • Scope boundary: The paper partly conflates corporate and state-level AI development despite their different organizational dynamics and incentive structures.A fuller treatment would separate the two and develop their governance implications independently.

7 Future Work

The paper identifies empirical research needed to test whether normalization of deviance is occurring in AI development. Proposed directions examine governance documents, personnel testimony, and state-level AI organizations.

  • Research agenda: The paper’s primary contribution is conceptual rather than empirical, so future work must test whether the historical process is occurring within AI organizations.The authors describe the application to frontier AI as preliminary and identify empirical research as necessary.
  • Corporate laboratories: The authors propose longitudinally coding frontier laboratories’ governance artifacts and relating policy changes to competitive events and safety incidents.Normalization would predict increasingly permissive procedural language and more entrenched competitive-adjustment provisions, unlike learning.
  • Corporate laboratories: Systematic analysis of personnel testimony could aggregate first-hand accounts describing standards eroding after deployments in which nothing visibly went wrong.The proposed evidence includes public accounts from departing safety and governance researchers.
  • State-level analysis: The framework should be extended to nation-states competing over AI capability, whose strategies, controls, procurement, and testing standards can be analyzed comparatively.The comparison would ask whether safety commitments erode differently under state principals than under market discipline.

8 Conclusion

The paper argues that AI development exhibits the organizational conditions associated with normalized deviance, while opaque systems and immature regulation make warning signals and correction more difficult. Because AI safety failures may be unrecoverable or illegible, safety institutions should be established before competitive pressures intensify.

  • AI development shows production pressure, false assurance from prior success, structural secrecy, and erosion of independent oversight.These mechanisms are presented as observable organizational conditions rather than failures requiring negligence or malice.
  • Opaque AI systems and immature regulation make warning signals harder to interpret while competitive pressure encourages treating safety work as friction.Practices that would violate mature safety norms can become normalized when developers effectively write the norms themselves.
  • Unlike earlier disasters, a sufficiently large AI safety failure may be neither recoverable nor legible enough to support institutional learning and reform.Prior investigators could reconstruct those disasters and design reforms to prevent recurrence; AI failures may not offer the same opportunity.
  • AI development is still in a pre-disaster period when norms and institutions are being designed, making early safety infrastructure more robust than infrastructure built after failure.The paper frames this period as analogous to the years before the major accidents examined in the paper.
Loading 2609.05749v1…