Source-linked AI summary

Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies

Miles Brundage, Noemi Dreksler, Aidan Homewood, Sean McGregor, Patricia Paskov, Conrad Stosz, Girish Sastry, A. Feder Cooper, George Balston, Steven Adler, Stephen Casper, Markus Anderljung, Grace Werner, Soren Mindermann, Vasilios Mavroudis, Ben Bucknall, Charlotte Stix, Jonas Freund, Lorenzo Pacchiardi, Jose Hernandez-Orallo, Matteo Pistillo, Michael Chen, Chris Painter, Dean W. Ball, Cullen O'Keefe, Gabriel Weil, Ben Harack, Graeme Finley, Ryan Hassan, Scott Emmons, Charles Foster, Anka Reuel, Bri Treece, Yoshua Bengio, Daniel Reti, Rishi Bommasani, Cristian Trout, Ali Shahin Shamsabadi, Rajiv Dattani, Adrian Weller, Robert Trager, Jaime Sevilla, Lauren Wagner, Lisa Soder, Ketan Ramakrishnan, Henry Papadatos, Malcolm Murray, Ryan Tovcimak

arXiv:2601.11699v4cs.CY

TL;DR

Frontier AI companies lack mechanisms for confident independent verification of their safety and security claims. This paper proposes organization-level audits with calibrated assurance levels and concludes that effective auditing requires comprehensive scope, secure access, continuous verification, and rigorous processes.

  • Problem

    Frontier AI companies lack mechanisms for confident third-party verification of safety and security claims, despite increasingly capable and widely deployed systems.

  • Method

    The paper proposes organization-level frontier AI audits and calibrated AI Assurance Levels that require progressively greater access, resources, and sophistication.

  • Results

    The paper concludes that effective frontier AI auditing requires comprehensive scope, secure access to non-public information, continuous verification, auditor independence, and rigorous traceable processes.

  • Takeaways & Limitations

    Calibrated assurance levels provide an adaptable vocabulary for communicating how confidently a frontier AI audit’s conclusions can be trusted.

  • Takeaways & Limitations

    Selective participation may disadvantage responsible developers while leaving the public exposed to systemic risks from weaker industry participants.

Abstract

from arXiv · show

We outline a vision for frontier AI auditing, which we define as rigorous third-party verification of frontier AI developers' safety and security claims, and evaluation of their systems and practices against relevant standards, based on deep, secure access to non-public information. Frontier AI audits should not be limited to a company's publicly deployed products, but should instead consider the full range of organization-level safety and security risks, including internal deployment of AI systems, information security practices, and safety decision-making processes. We describe four AI Assurance Levels (AALs), the higher levels of which provide greater confidence in audit findings. We recommend AAL-1 as a baseline for frontier AI generally, and AAL-2 as a near-term goal for the most advanced subset of frontier AI developers. Achieving the vision we outline will require (1) ensuring high quality standards for frontier AI auditing, so it does not devolve into a checkbox exercise or lag behind changes in the industry; (2) growing the ecosystem of audit providers at a rapid pace without compromising quality; (3) accelerating adoption of frontier AI auditing by clarifying and strengthening incentives; and (4) achieving technical readiness for high AI Assurance Levels so they can be applied when needed.

Key paper takeaways

AI systems receive less rigorous third-party scrutiny than many other widely relied-upon systems, creating an increasingly untenable gap as AI capabilities and deployment expand. Transparency alone cannot establish well-calibrated trust in frontier AI because key safety and security details are confidential, require expert interpretation, and should not be assessed solely by developers.

  • AI systems face less rigorous third-party scrutiny than consumer products, corporate financial statements, and food supply chains.
  • As AI becomes more capable and widely deployed, this scrutiny gap increasingly inhibits confident deployment in high-stakes contexts.
  • Transparency alone cannot support well-calibrated trust in frontier AI because safety- and security-relevant details may be confidential, require expert interpretation, and make self-assessment suspect.Third-party skepticism is warranted given the track record of companies “checking their own homework” in other industries.

Frontier AI auditing motivations

AI’s rapid development and deployment have outpaced institutions for ensuring safety and reliability, creating an especially urgent gap for frontier systems. Frontier AI auditing is proposed as rigorous, independent verification of developers’ safety and security claims against relevant standards, using deep access to non-public information.

  • Motivation: AI systems increasingly inform or autonomously make consequential decisions affecting billions, while institutional safeguards have not kept pace with development and deployment.AI is becoming critical societal infrastructure, but institutions ensuring that systems work safely and as advertised have lagged behind.
  • Motivation: Frontier systems face risks including harmful failures and weaponization, making the institutional gap particularly important for the most capable AI.The passage defines frontier systems as general-purpose models no more than a year behind the state of the art.
  • Motivation: Stakeholders need reliable verification of technical safeguards, but complexity, rapid change, proprietary information, and the limits of public transparency make this difficult.Many important details require confidentiality and expert judgment to interpret.
  • Proposed response: Frontier AI auditing would provide rigorous third-party verification of safety and security claims and evaluate systems and practices against relevant standards using deep access to non-public information.The proposal aims to give stakeholders, including skeptics and uncertain parties, justified confidence that frontier AI is being developed safely and securely.
  • Proposed response: A private-sector ecosystem of for-profit and nonprofit auditors could broaden confidence in frontier AI while avoiding both companies grading their own homework and exclusive reliance on governments.The proposed ecosystem is intended to address concerns about government technical expertise, capacity, and agility.

Summary of the proposal

The proposal sets out eight interlinked design principles for ambitious frontier AI auditing as capabilities and risks advance. It emphasizes comprehensive, organization-level assessment with calibrated assurance, secure access, continuous monitoring, and trustworthy communication.

  • Scope of risks: Audits should cover intentional misuse, unintended system behavior, information security, and governance risks, verifying company claims and evaluating practices against policies, regulations, and best practices.The ecosystem should ensure comprehensive coverage across all three components of assessing safety and security claims, even when individual audits focus on particular domains.
  • Organizational perspective: Auditors should assess companies’ safety and security practices as a whole, because organizational risk emerges from interactions among models, systems, and broader practices.This organization-level perspective is intended to avoid abstraction errors from evaluating individual components in isolation.
  • Levels of assurance: AI Assurance Levels should calibrate and communicate confidence in audit conclusions, with deeper, more secure access to non-public information proportional to the audit’s scope and assurance level.Proposed access includes model internals, training processes, compute allocation, governance records, and staff interviews.
  • Continuous monitoring: Audits should be living assessments whose findings state validity conditions and are updated through periodic deep assessments, event-triggered reviews, and continuous monitoring of fast-changing surfaces.AI systems can change through model and software adjustments and shifts in user behavior, making previously accurate conclusions misleading within days or weeks.
  • Trustworthy audit practice: Trustworthy audits require independent, expert auditors; rigorous, traceable, adaptive methods; and clear reports communicating scope, assurance level, conclusions, reasoning, and recommendations.Safeguards include financial-relationship disclosure, standardized engagement terms, cooling-off periods, auditor autonomy, and automated, transparent, reproducible procedures where feasible.

Challenges and next steps

Frontier AI auditing faces four urgent challenges: maintaining quality, expanding the provider ecosystem, strengthening adoption incentives, and achieving technical readiness for high assurance levels. Recommended next steps include independent verification, accreditation, safe harbors, clearer insurance and procurement requirements, and immediate investment in pilots and research.

  • Challenges: The most urgent challenges are maintaining high-quality audits, rapidly expanding providers without compromising quality, strengthening adoption incentives, and achieving technical readiness for high AI Assurance Levels.These challenges are substantial but considered achievable through carefully controlled practices modeled on mature industries.
  • Ensuring high quality standards: AI companies, philanthropists, investors, and insurers should publicly fund analysis of audit quantity and quality, while policymakers should establish a PCAOB-style nonprofit “auditor of auditorsˮ.The proposed oversight body would receive final government approval of standards, hold auditors accountable, and innovate at private-sector speed.
  • Growing the ecosystem: The evaluation ecosystem should create a Frontier AI Auditor Accreditation Program with tiered certifications, specialty endorsements, and meaningful accountability mechanisms.This recommendation is intended to expand the auditing ecosystem while preserving quality and auditor accountability.
  • Accelerating adoption: Policymakers and developers should create targeted safe harbors, clarify AI-risk insurance coverage, and incorporate frontier AI auditing requirements into procurement, especially for health and defense systems.Safe harbors should protect good-faith research and auditing without creating a liability gap and should depend on compliance with established best practices.
  • Achieving technical readiness for high AALs: Immediate investment in auditing pilots, technical research, and policy research is needed to keep auditing aligned with rapid AI progress and deployment.The paper argues that urgency is essential for frontier AI auditing to mature and scale alongside AI development.

1 Introduction

Frontier AI systems’ development and safeguards remain largely opaque to the external stakeholders most affected by their impacts, while public-information assessments provide insufficient assurance. The paper proposes frontier AI auditing—rigorous, independent verification based on secure access to non-public information—and standardized AI Assurance Levels to calibrate confidence in audit conclusions.

  • Motivation: Frontier AI systems are rapidly transforming society, but their development, evaluation, and safeguards remain largely opaque to users, insurers, investors, and policymakers.These external stakeholders bear both benefits and risks while having limited ability to scrutinize the systems.
  • Problem: Most third-party assessments rely on public information and products, which rarely reveal evaluation conditions, generalizability, or whether negative findings were softened or withheld.The paper argues that these limitations become more consequential as frontier AI systems grow more capable and more widely deployed.
  • Problem: Meaningful assurance requires independent access to non-public technical and organizational information, yet developers typically control access through opaque bespoke contracts and can influence assessment scope, timing, and publication.Standards for developer collaboration with third-party AI assessment organizations remain nascent.
  • Proposed approach: The paper defines frontier AI auditing as rigorous third-party verification of developers’ safety and security claims throughout development and deployment, plus evaluation against relevant standards using deep, secure access to non-public information.The approach addresses both systems and organizational practices rather than only publicly deployed products.
  • Proposed approach: The paper proposes standardized AI Assurance Levels (AALs) as an adaptable vocabulary that clarifies the assumptions required to trust a particular audit’s conclusions.The framework is intended to accommodate different contexts and levels of assurance rather than impose a one-size-fits-all audit model.
  • Paper roadmap: The paper’s later sections develop frontier AI auditing’s scope and assurance levels, then address implementation challenges and propose directions for achieving the vision.The discussion covers organization-level focus, access, continuous monitoring, independence, rigor, communication, and funding-related topics.

How to read this paper

The paper offers tailored reading paths: newcomers should begin with Section 2 and Appendix A, while readers familiar with AI assessment limitations can proceed from Section 4.1 to Section 5.

  • Reading paths: Readers unfamiliar with the paper’s terminology should read Section 2 and consult Appendix A as needed.
  • Reading paths: Most readers should read Section 3 to understand the problems frontier AI auditing would address.
  • Reading paths: Readers familiar with current AI assessment limitations can skip to Section 4.1, then proceed directly to Section 5.

2 Key Terminology and Scope

Frontier AI auditing is defined as rigorous third-party evaluation and verification of developers’ systems, practices, and safety and security claims, based on deep, secure access to non-public information. The scope centers on frontier AI developers and systems across development and deployment, using assurance levels to match audit effort and confidence to the risks of increasingly capable systems.

  • Definition and scope: Frontier AI auditing combines rigorous third-party evaluation against relevant standards with verification of safety and security claims, using deep, secure access to non-public information.The framework distinguishes audits from assessments by emphasizing systematic, evidence-based examination by a qualified party.
  • Definition and scope: Frontier AI includes general-purpose models and systems within one year of the state-of-the-art on a broad suite of general capability benchmarks.The most capable systems are less well understood because they are new, making rigorous auditing particularly important.
  • Definition and scope: The scope covers companies that train frontier models or significantly extend their capabilities, and systems throughout development and deployment.Frontier systems evolve through new developers, product offerings, deployment contexts, and increasingly autonomous agentic designs.
  • Access and assurance: Third-party assessments should use non-public information because public or unprivileged information may miss major risks and vulnerabilities below the surface.This is especially important when dangerous systems are deployed internally or access is restricted to selected customers.
  • Access and assurance: The framework offers a menu of assurance levels, with higher levels intended for systems whose improving capabilities warrant proportionally greater analytical effort and confidence.The authors expect some frontier systems will ultimately merit the highest AI Assurance Levels.

3 Motivations: Why Frontier AI Auditing is Needed

Frontier AI auditing is needed because internal self-assessment is limited by knowledge gaps and conflicting incentives, while independent assessments can improve safety, security, ecosystem learning, and adoption confidence. A primarily private-sector regime, complemented by government standards and oversight, distributes expertise and accountability while avoiding dependence on either sector alone.

  • Why auditing is needed: Internal evaluations are insufficient because developers may not fully understand external risks and face incentives misaligned with adequate safety precautions.External auditing challenges internal narratives and addresses the resulting safety and security risks.
  • Why auditing is needed: External auditors add skepticism and broader expertise, and third-party assessments have surfaced safety and security issues that developers later remedied.Independent auditing helps guard against groupthink and improve development and deployment decisions.
  • Why auditing is needed: Auditing enables ecosystem-level learning by making practices more comparable, revealing systemic risks, and allowing auditors to share patterns, best practices, and mitigations across companies.Wide participation is important to capture these benefits and discourage companies from cutting corners for short-term advantage.
  • Why auditing is needed: Credible third-party audits lower due-diligence costs, support adoption and investment, document responsible decisions amid unsettled standards, and create competitive differentiation for audited developers.Shared independent assessments are especially useful to enterprises and governments that lack frontier-level technical expertise.
  • Why a private-sector regime: A private-sector-led auditing regime is preferable to relying primarily on governments, but governments must set standards, oversee auditors, and enforce accountability.Private auditing distributes oversight across institutions with different incentives, expertise, and failure modes while retaining public involvement.

4 Lessons from Related Domains and Current AI Assessment

Lessons from established assurance regimes emphasize layered testing, lifecycle risk management, adversarial assessment, auditor independence, and clear assurance limits. Current frontier AI assessment is growing but remains uneven in access, rigor, transparency, continuity, scope, standardization, scale, and independence.

  • Lessons from Related Domains: The goal of frontier AI auditing is to drive major safety and security progress without requiring catastrophe to trigger action.Historically, rigorous oversight in several industries followed serious incidents, but frontier AI auditing aims to avoid that pattern.
  • Lessons from Related Domains: Food safety and consumer-product testing show that effective assurance requires defense in depth across lifecycle stages and failure modes, supported by trusted independent certification.Independent organizations such as Underwriters Laboratories demonstrate that firms may opt into certification when consumers value third-party assurance.
  • Lessons from Related Domains: Safety-critical engineering and penetration testing support proactive lifecycle risk management and active adversarial searches for unexpected, chained weaknesses rather than static checklists.Aviation and nuclear practices use hazard analysis and safety cases, while penetration testing combines adversarial analysis with collaborative remediation.
  • Lessons from Related Domains: Financial auditing demonstrates the value of independent review, standardized metrics, and sensitive-information access, while warning against conflicts of interest, procedural box-ticking, and unclear assurance expectations.These cautions motivate auditor independence, explicit assurance levels, and attention to systemic issues.
  • Current AI Assessment: Current frontier AI assessments vary substantially, with most relying on black-box access and one-off technical evaluations while reporting, threat modeling, deeper access, organizational scope, standards, and independence remain limited.Assessments are often confidential and bespoke, systems change without updated third-party review, and evaluators depend on companies for access and sometimes funding.
  • Current AI Assessment: The field nevertheless has a growing ecosystem, proposed frameworks, emerging best practices, evaluator coordination, government-institute pilots, and early reviews of company risk assessments.The paper identifies METR’s review of Anthropic’s sabotage risk report as among the first AAL-1 audits.

5 A Vision for Frontier AI Auditing

The paper proposes frontier AI auditing as mature, third-party assessment of both frontier systems and the companies building them, extending beyond current assurance practice. Its vision emphasizes comprehensive risk coverage, organization-level analysis, calibrated assurance levels, secure deep access, continuous monitoring, and independent expertise.

  • Design principles: The vision organizes frontier AI auditing around eight interlinked principles, including comprehensive risk coverage, organization-level analysis, calibrated assurance, secure access, continuous monitoring, and independent expertise.The proposed approach aims beyond the status quo because current assurance needs are not fully met and future systems may create new demands.
  • Scope of risks: Audits should cover risks directly linked to company action or inaction, including unintended system behavior, information security failures, and emergent social phenomena.Examples include harmful irreversible actions, model-weight or data exfiltration, sabotage, addiction, AI-induced psychosis, and facilitation of self-harm.
  • Organizational perspective: Auditors should emphasize the company as a whole because individual systems or components can illustrate risk but never capture the full story of organizational impact.Audit conclusions about specific systems or artifacts should be framed explicitly within the larger company context, because abstraction errors can mislead stakeholders about overall risk posture.
  • Levels of assurance: AI Assurance Levels calibrate and communicate confidence in conclusions, with higher levels assessing company-wide risk more confidently and progressively ruling out abstraction and other audit errors.Higher assurance requires deeper access across the lenses used to form a composite company risk profile, supported by standardized processes and safety-case-style evidence.
  • Assurance-level capabilities: AAL-1 can reveal negligence, policy-practice gaps, cherry-picking, and basic security failures, but cannot reliably detect unsampled problems, concealment, shadow systems, or sophisticated fraud.AAL-1 engagements provide meaningful evidence beyond self-assessment, although they may not qualify as audits under some standards.
  • Assurance-level capabilities: AAL-3 is at least very difficult today and AAL-4 infeasible today, while AAL-4 is intended to address deliberate deception through mechanisms such as hardware attestation and cryptographic logs.Research and pilots for both levels are therefore identified as priorities.

6 Challenges and Next Steps

Frontier AI auditing must balance rigorous assurance with adaptability, while developing accountable oversight, auditor capacity, legal protections, and incentives that can sustain adoption. The paper recommends combining independent quality control, accreditation and training, targeted safe harbors, insurance and procurement requirements, and public ecosystem monitoring.

  • Audit quality standards: Auditing standards must provide meaningful assurance while adapting to rapid industry change, avoiding rigid box-ticking, excessive flexibility, Goodhart effects, and temporal mismatch.The paper recommends durable outcome-oriented goals paired with evolving practices.
  • Adoption incentives: Public reporting on audit quantity and quality, including at least quarterly ecosystem-health reports, can create interim soft pressure before formal incentives and requirements mature.The paper recommends publicly available assessments to inform investment, insurance, procurement, and regulatory decisions while avoiding pressure that encourages corner-cutting.
  • Independent oversight: A PCAOB-style nonprofit “auditor of auditors” should set and inspect standards, investigate failures, enforce accountability, and retain legitimacy through government approval and globally credible governance.Possible enforcement includes revoking accreditation; proposed methods include re-performing audits with equivalent access and developing hard-to-game quality indicators.
  • Auditor capacity: A tiered Frontier AI Auditor Accreditation Program, supplemented by training and multi-organizational experience, should address talent bottlenecks and progressively raise competency standards.Academic researchers are identified as a promising supplemental talent pool, provided implementation concerns such as revolving-door risks are addressed.
  • Legal protections: Targeted safe harbors should protect good-faith safety research and auditing while remaining conditional on established best practices and avoiding liability gaps.Developers can provide interim testing permissions, disclosure channels, and non-retaliation commitments before legislation is enacted.
  • Adoption incentives: Because markets alone may not ensure adoption, policymakers should use insurance and procurement requirements, especially for frontier AI systems deployed in high-stakes domains.Recommendations include resolving insurers’ coverage exclusions and requiring explicit coverage of AI-related risks in government procurement.

7 Conclusion … E.2 Consumer product safety

The paper defines frontier AI auditing as rigorous, independent verification of frontier developers’ safety and security claims using deep, secure access to non-public information, while recognizing major implementation challenges and important scope limitations. It argues that auditing can support accountability, risk pricing, international stability, and credible compliance evidence, but must assess organizations, systems, practices, and relevant standards comprehensively.

  • 7 Conclusion: Frontier AI auditing addresses the absence of reliable mechanisms for confirming companies’ safety and security claims or compliance with relevant standards.The proposed alternative is rigorous third-party verification based on deep, secure access to non-public information.
  • 7 Conclusion: Effective audits should cover misuse, unintended behavior, information security, emergent social phenomena, organizational practices, non-public evidence, calibrated assurance, and continuous assessment.The organizational perspective treats culture, governance, and security as relevant alongside specific AI systems.
  • 7 Conclusion: Substantial challenges remain in maintaining audit quality, expanding the provider ecosystem, accelerating adoption, and achieving technical readiness for higher assurance levels.The paper also notes limitations concerning open-weight models, fine-tuning providers, downstream deployers, cross-country scaling, and reliance on complementary institutions.
  • 7 Conclusion: The paper recommends AAL-1 for frontier AI generally and AAL-2 for the leading subset, while leaving later assurance decisions to further research and pilots.Higher assurance levels require greater confidence in audit conclusions, but the paper presents only interim near-term recommendations.
  • Enabling risk price discovery through insurance: Third-party audits provide insurers with verified, standardized information needed to differentiate AI risk profiles and support meaningful coverage and pricing.Insurance pricing can translate uncertainty into a continuously updating signal for policymakers, the public, developers, regulators, insurers, and civil society.
  • Maintaining international stability: Independent auditors can verify adherence to shared safety commitments, reduce competitive risk-taking, and provide a foundation for future international AI safety treaties.Audits serve as trusted intermediaries when governments are unlikely to grant one another deep access for meaningful verification.
  • Ensuring accountability for risk creation: Auditing supports public and governmental accountability by providing external validation of safety claims and mechanisms for verifying compliance with standards and voluntary commitments.Third-party auditors can distribute authority, provide specialized expertise, and scale more readily than government agencies.
  • C Access types (non-exhaustive): The framework relies on varied access to models, systems, governance, operations, communications, records, interviews, public outputs, regulatory disclosures, research, and external analyses.Audit reports should communicate scope, assurance level, conclusions, reasoning, validity conditions, and remediation recommendations; higher assurance may require secure technical pathways such as hardware guarantees or zero-knowledge proofs.

E.3 Safety-critical systems engineering · E.4 Aviation safety

Safety-critical systems engineering shows that frontier AI auditing should combine technical analysis with organizational practice audits, continuous lifecycle risk management, and evidence-based safety arguments. Aviation safety further demonstrates the need for genuinely independent, technically capable auditors because self-certification and weak oversight can produce severe consequences.

  • E.3 Safety-critical systems engineering: Technical system analysis should be paired with audits of company practices because safety emerges from complex sociotechnical interactions and control relationships.Safety-critical systems engineering treats system safety as an emergent property rather than solely a technical-system attribute.
  • E.3 Safety-critical systems engineering: Continuous lifecycle risk management with formal acceptance and verification is preferable to one-off certification for systems that change over time.This approach supports proactive identification and management of hazards as systems evolve.
  • E.3 Safety-critical systems engineering: Near-misses and incidents can provide early warning of failures that later cause significant harm, making lessons from them important to organizational safety.Safety-critical engineering emphasizes that catastrophic accidents are often anticipatable and avoidable through process attention and learning from near-misses.
  • E.3 Safety-critical systems engineering: Auditors should review hazard analyses, safety cases, and other documented evidence in addition to raw evaluation results and system-property measurements.These artifacts support explicit, evidence-based arguments about the implications of system properties.
  • E.4 Aviation safety: 346 fatalities across two Boeing 737 MAX crashes exposed critical weaknesses in aviation certification and the dangers of excessive manufacturer self-certification.Boeing self-certified 96% of the parts for the 737 MAX, illustrating the scale of delegated certification.
  • E.4 Aviation safety: Aviation’s strong safety regimes still leave residual risk, so frontier AI auditing should set ambitious expectations while remaining realistic about what audits can achieve.Government agencies may also struggle to retain specialized expertise, supporting a private-sector auditing ecosystem with public oversight.
  • E.4 Aviation safety: Delegating certification to the entities being certified creates conflicts of interest, so frontier AI auditors must be genuinely independent third parties.Commercial and internal pressures can discourage employees from raising safety concerns, turning self-assessment into self-dealing.
  • E.4 Aviation safety: Auditors and regulators cannot effectively oversee systems they do not understand, and insufficient technical expertise risks reducing oversight to rubber stamping.Public and third-party auditors need resources to maintain independent technical expertise.

E.5 Penetration testing

Frontier AI audits should use active, adversarial penetration testing to uncover realistic attack paths and high-impact vulnerabilities that checklist reviews may miss. These engagements should complement in-house security work, be supported by safe harbors and iterative remediation, and connect private findings to broader ecosystem improvement.

  • Core role: Penetration testing actively probes networks, applications, or infrastructure like a real attacker, creatively chaining weaknesses to pursue concrete, high-impact goals.It is controlled and permissioned, but goes beyond checking documented requirements.
  • Limitations: A failed simulation of a motivated nation-state attacker does not establish that real state-level attackers would fail, limiting conclusions about defensive strength.Finding and exploiting a vulnerability does show defenses are unlikely to withstand attacks requiring similar expertise and effort.
  • Core role: Active adversarial testing should be core to audits for security- and misuse-related risks rather than relying only on checklist-style reviews.Realistic attacks can expose inadequate defense-in-depth and critical vulnerabilities missed by qualified in-house teams.
  • Engagement design: Legal safe harbors for good-faith researchers are essential to constructive engagement with companies on high-risk product aspects.Adversarial analysis can coexist with collaboration when auditors and companies iteratively fix issues rather than treating audits as one-off pass/fail exercises.
  • Engagement design: Penetration-style engagements should complement, not substitute for, in-house security work, focusing audits on realistic attack paths and prioritized remediation guidance.Audits can reveal vulnerabilities missed internally while building on existing security work.
  • Ongoing scrutiny: Bug bounty-style programs can add continuous, incentive-aligned scrutiny from broad researcher pools, with payment tied to impact and expectations for rapid remediation.Private engagement can provide space to mitigate risks, but findings should eventually be surfaced publicly to drive ecosystem-wide improvements.

E.6 Financial auditing · F Contemporary third-party frontier AI assessment · F.1 Reporting

Financial auditing offers frontier AI auditing a model for independent review of sensitive information, professional standards, and public accountability, while exposing risks from compromised independence and weak fraud detection. Contemporary third-party assessment provides a foundation, but inconsistent reporting and broader gaps must be addressed to realize rigorous frontier AI auditing.

  • E.6 Financial auditing: Financial audits parallel frontier AI audits by independently testing public claims and internal safeguards through access to highly sensitive, non-public information.Financial auditing demonstrates that independent review can assess both external credibility and internal controls.
  • E.6 Financial auditing: Financial auditing shows that professional standards, standardized risk-and-control comparisons, and combined private and public demand signals can support high-stakes decisions.Its mature ecosystem spans firms of many sizes, private standard-setting, public regulation, and extensive case law.
  • E.6 Financial auditing: Frontier AI auditing can borrow financial auditing’s conflict-management, evidence-evaluation, fraud-detection, professional-judgment, and public-duty concepts.These tools provide institutional precedents without requiring frontier AI auditing to reinvent every conceptual foundation.
  • E.6 Financial auditing: Enron and Wirecard show that concentrated client dependence, consulting conflicts, and retention pressures can compromise auditor independence and objectivity.These scandals demonstrate why conflict disclosure and safeguards against auditor capture are essential.
  • E.6 Financial auditing: Frontier AI auditing should pair standardized evidence requirements with professional judgment, adapt faster than traditional rules, automate coverage, and prioritize duties to users and the public.Because audits have weak fraud-detection records, deception risks require explicit scoping or unusually deep access, high-effort testing, and strong disincentives.
  • F Contemporary third-party frontier AI assessment: Third-party frontier AI assessment has grown significantly and provides an important foundation, but substantial gaps remain between current practices and the proposed auditing vision.The assessment examines reporting, access, rigor, standardization, continuous monitoring, scope, scale, independence, and ecosystem maturity.
  • F.1 Reporting: Public reporting remains inconsistent across and within frontier AI developers, with substantial variation in templates, substance, and style and some analyses finding declining quality.System cards and related publications are currently the most common channels for communicating audit results.

F.2 Access … F.9 Ecosystem Maturity

Frontier AI auditing is constrained by limited access, uneven rigor, weak standardization, delayed monitoring, narrow scope, incomplete adoption, dependence on developer cooperation, and an immature evaluator ecosystem. Addressing these gaps requires broader secure access, stronger and more reproducible methods, continuous change detection, organization-level assessment, expanded coverage, and greater independence and capacity.

  • F.2 Access: Third-party access usually remains developer-controlled and black-box, while high-assurance assessments may require richer interfaces, non-public documentation, and gray-box or white-box access.Evaluators rarely receive training data, chain-of-thought, model internals, or detailed training and deployment documentation; practical obstacles include model changes, API bugs, rate limits, and time constraints.
  • F.3 Rigor: Third-party assessment methodology varies substantially, with benchmark, red-teaming, construct-validity, coverage, time, resource, standardization, and reproducibility limitations weakening confidence in findings.Assessment windows rarely exceed a few weeks and often compress to days, while divergent threat models, scoring rubrics, protocols, proprietary prompts, and undocumented implementation details hinder comparison and replication.
  • F.4 Standardization: Frontier AI assessment standards are evolving, but confidential bilateral contracts and absent standardized contractual and reporting requirements often omit details needed for independent verification.Recent frameworks, industry best practices, and the AI Evaluator Forum are promising but lag behind evaluation regimes in more established industries.
  • F.5 Continuous monitoring: Developers often change systems without giving third parties early access for updated risk assessments, increasing the need for automated measurement of changes and reduced reliance on self-reporting.The paper identifies model-fingerprint checking by stampr-AI as a nascent effort toward continuous public and auditor inspection.
  • F.6 Scope: Current assessments emphasize technical, consumer-facing systems and selected capability, propensity, or mitigation risks, while generally lacking visibility into organizational processes, governance, culture, and internally deployed AI.Platform-level assessments can require different operating conditions from system-level assessments, including safe-harbor protections for security researchers.
  • F.7 Scale: Participation in third-party assessment is inconsistent across developers and regions, and universal high-assurance coverage would exceed immediate evaluator capacity despite likely feasibility at lower assurance levels.Third-party assessment of frontier Chinese AI systems remains particularly limited, while engagement also varies among fast followers and open-weight developers.
  • F.8 Independence: Voluntary participation and developers’ contractual leverage can pressure assessors to avoid findings or disclosures that might jeopardize continued access, although external-evaluation requirements can improve consistency.The EU General-Purpose AI Code of Practice links independent external evaluations to a presumption of conformity with the EU AI Act for signatories.
  • F.9 Ecosystem Maturity: The specialized frontier AI auditing ecosystem remains nascent, with dedicated safety capacity concentrated in a handful of organizations employing likely only a few hundred full-time staff, alongside emerging coordination.Private organizations increasingly coordinate through the AI Evaluator Forum, while governments coordinate through the Network for Advanced AI Measurement, Evaluation, and Science.

G Risks of and alternatives to “auditee pays” models

Auditors are currently funded partly by philanthropy and frontier AI companies, but company-paid competition risks biased “rubber stamping.” Alternative models may better align incentives, though evidence remains limited and no model is recommended yet.

  • Alternative funding models: Alternative funding models include insurer payments, regulator-administered pools, downstream enterprise-user payments, industry-wide levies, and hybrid approaches.These mechanisms could diversify funding beyond philanthropy and frontier AI company payments.
  • Risks of auditee pays: Company-awarded contracts create a significant risk of “rubber stamping” because auditors may develop conscious or unconscious bias toward favorable client outcomes.The paper compares this risk to financial auditing before the 2008 crisis.
  • Risks of auditee pays: Insurer- or regulator-funded models merit particular attention because they better align auditor incentives with accurate risk assessment.The paper does not claim these models are proven superior; it identifies them as especially worthy of further consideration.
  • Open questions: Empirical evidence on alternative models is limited, the ideal frontier AI funding model remains unclear, and the paper recommends research rather than a specific model.The authors note few exceptions to the limited evidence base, including reference.

H Frontier AI definitions and different thresholds for triggering audits

Frontier AI is defined by general-purpose capabilities within roughly 12 months of the state of the art, but temporal and capability-based thresholds are imperfect proxies. Audit triggers should therefore remain adaptable and balance coverage, burden, quality, and risks from systems outside the general-capability definition.

  • Definition: Frontier AI means general-purpose models and systems whose broad benchmark performance is no more than one year behind the state of the art.This definition is similar to the Frontier Model Forum’s definition, which uses performance relative to models widely deployed for at least 12 months.
  • Threshold rationale: The 12-month threshold focuses attention on actual capabilities, but temporal thresholds remain imperfect proxies that may be easier to administer and communicate.The paper acknowledges problems with relying on a single threshold while noting that proxy-based rules can be more predictable for policy implementation.
  • Adaptability: Thresholds tied to audit requirements must be changeable over time because a metric that works today may become less suitable as AI capabilities evolve.This concern applies to both legislators codifying requirements and other actors designing audit-triggering rules.
  • Scope limitations: General-capability thresholds can exclude systems with dangerous domain-specific capabilities, so the auditing approach may need multiple sufficient triggers or adaptations.Systems trained on certain data may pose area-specific risks despite poor performance on broad benchmarks.
  • Threshold trade-offs: Threshold design must balance false positives and false negatives while stimulating rigorous audits without encouraging auditors to churn out low-quality work.Adaptive triage can reduce errors, but imprecision is inevitable and market analysis and quality standards can only soften the trade-offs.
Loading 2601.11699v4…