Source-linked AI summary

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims

Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, Tegan Maharaj, Pang Wei Koh, Sara Hooker, Jade Leung, Andrew Trask, Emma Bluemke, Jonathan Lebensold, Cullen O'Keefe, Mark Koren, Théo Ryffel, JB Rubinovitz, Tamay Besiroglu, Federica Carugati, Jack Clark, Peter Eckersley, Sarah de Haas, Maritza Johnson, Ben Laurie, Alex Ingerman, Igor Krawczuk, Amanda Askell, Rosario Cammarota, Andrew Lohn, David Krueger, Charlotte Stix, Peter Henderson, Logan Graham, Carina Prunkl, Bianca Martin, Elizabeth Seger, Noa Zilberman, Seán Ó hÉigeartaigh, Frens Kroeger, Girish Sastry, Rebecca Kagan, Adrian Weller, Brian Tse, Elizabeth Barnes, Allan Dafoe, Paul Scharre, Ariel Herbert-Voss, Martijn Rasser, Shagun Sodhani, Carrick Flynn, Thomas Krendl Gilbert, Lisa Dyer, Saif Khan, Yoshua Bengio, Markus Anderljung

arXiv:2004.07213v2cs.CY

TL;DR

AI developers need ways to substantiate responsible-development claims and enable outside scrutiny. This report analyzes mechanisms for making claims about safety, security, fairness, and privacy more verifiable, offering an incremental toolbox for developers and stakeholders.

  • Problem

    Existing regulations, norms, and non-binding ethics principles provide limited means to assess whether AI developers’ actions align with responsible-development claims.

  • Method

    The report analyzes institutional, software, and hardware mechanisms for improving the verifiability and assessment of claims about AI development.

  • Results

    The report identifies ten mechanisms and associated recommendations addressing specific gaps in assessing claims about AI developers’ safety, security, fairness, and privacy practices.

  • Takeaways & Limitations

    Verifiable claims can provide a toolbox for developers, civil society, and regulators seeking to support more trustworthy AI development.

  • Takeaways & Limitations

    The report focuses on verifiable claims about safety, security, fairness, and privacy, which are only one aspect of trustworthy AI development.

Abstract

from arXiv · show

With the recent wave of progress in artificial intelligence (AI) has come a growing awareness of the large-scale impacts of AI systems, and recognition that existing regulations and norms in industry and academia are insufficient to ensure responsible AI development. In order for AI developers to earn trust from system users, customers, civil society, governments, and other stakeholders that they are building AI responsibly, they will need to make verifiable claims to which they can be held accountable. Those outside of a given organization also need effective means of scrutinizing such claims. This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems. We analyze ten mechanisms for this purpose--spanning institutions, software, and hardware--and make recommendations aimed at implementing, exploring, or improving those mechanisms.

Executive Summary

Existing regulations, norms, and non-binding ethics principles do not sufficiently ensure responsible AI development or enable outsiders to assess and hold developers accountable. The report proposes a toolbox of institutional, software, and hardware mechanisms to make claims about AI systems and development processes more verifiable, especially regarding safety, security, fairness, and privacy.

  • Problem: Existing regulations, norms, and non-binding ethics principles leave outsiders unable to reliably assess whether developers’ actions match their stated commitments or hold them accountable.The report identifies a gap between adopting principles and translating them into observable, enforceable actions.
  • Goals: The proposed mechanisms focus on evidence for AI systems’ safety, security, fairness, and privacy protection, supporting more trustworthy development and oversight.Greater verifiability can help regulators, the public, and other developers assess responsible behavior and reduce incentives to cut corners.
  • Approach: The report calls for a robust toolbox of mechanisms that helps developers substantiate claims and enables users, policymakers, and civil society to make more specific demands.These mechanisms address gaps that currently prevent effective assessment of developers’ claims.
  • Institutional Mechanisms: Institutional mechanisms emphasize third-party auditing to shape incentives and increase visibility into developers’ behavior and responsible-AI practices.The report treats institutions as foundational because people remain ultimately responsible for AI development.
  • Software Mechanisms: Software mechanisms address accountability, system understanding, and privacy through audit trails, interpretability, and privacy-preserving machine learning.These mechanisms target critical development and deployment information, scrutiny of system characteristics, and stronger privacy commitments.
  • Hardware Mechanisms: Hardware mechanisms use secure hardware, compute measurement, and related capabilities to substantiate privacy and security claims, improve transparency, and shape access to verification resources.The report organizes these mechanisms alongside institutional and software mechanisms as intertwined components of AI systems and development processes.

List of Recommendations

The recommendations call for stronger scrutiny of AI through third-party auditing, red teaming, incentives for broad-based review, and improved incident disclosure. They also promote standards, interpretability, privacy, hardware security, precise compute measurement, and expanded academic resources for verification.

  • Scrutiny and accountability: Stakeholders should create a task force on funding and conducting third-party AI audits, while developers run red teaming and share related practices and tools.These measures target independent scrutiny and systematic exploration of risks.
  • Scrutiny and accountability: Developers should pilot bias and safety bounties and share AI incident information through collaborative channels to strengthen broad-based scrutiny and learning.The recommendations emphasize both incentives for external review and collaborative incident disclosure.
  • Standards and technical foundations: Standards bodies should develop audit trail requirements for safety-critical AI, while organizations and funders support interpretability research focused on risk assessment and auditing.Together, these recommendations strengthen documentation and the technical basis for evaluating AI systems.
  • Technical foundations: Developers should develop, share, and use privacy-preserving machine-learning tool suites measured against common standards, while industry and academia advance secure hardware practices.Secure-hardware work includes AI accelerator features and secure enclaves on commodity hardware.
  • Measurement and research capacity: AI labs should estimate project computing power in great detail and report adoption potential, while governments substantially increase computing resources for academic verification research.The recommendations combine high-precision compute measurement with greater academic access to computing resources.

1 Introduction

The report argues that responsible AI development requires verifiable claims supported by evidence, because principles alone have not secured public trust. It surveys mechanisms across institutions, software, and hardware for demonstrating and scrutinizing claims about safety, security, fairness, and privacy.

  • Growing concern about AI’s societal impacts reflects development and deployment that can conflict with developers’ stated values.
  • Ethics principles increasingly converge on safety, security, fairness, and privacy, but principles alone are insufficient to ensure beneficial AI outcomes.
  • Verifiable claims are falsifiable statements supported by evidence and arguments, enabling stakeholders to assess which AI properties and processes can be credibly demonstrated.
  • The report organizes mechanisms for claim verification into institutional, software, and hardware pillars within AI’s sociotechnical development processes.
  • The report focuses on organizations’ verifiable claims about AI systems and external scrutiny, especially claims concerning safety, security, fairness, and privacy, while noting that verifiability is not equivalent to trustworthy development.

2 Institutional Mechanisms and Recommendations

Institutional mechanisms shape incentives, clarify goals and values, increase transparency, and foster information exchange to support accountable and verifiable AI development. The section recommends stronger external auditing and broader scrutiny through red teaming, bias and safety bounties, and incident sharing.

  • Institutional mechanisms: Institutional mechanisms clarify organizational goals and values, increase transparency, create incentives for responsible behavior, and foster information exchange.These functions help stakeholders evaluate claims and scrutinize AI development practices.
  • Third-party auditing: Third-party auditors can independently verify developers’ safety, security, privacy, and fairness claims while receiving privileged, secured access to sensitive information.The section recommends that a stakeholder coalition research options for conducting and funding third-party AI auditing.
  • Red teaming: Organizations developing AI should run red-teaming exercises to uncover risks and share best practices and tools, while addressing scope and proprietary-information constraints.Red teams adopt an attacker’s mindset to find flaws and vulnerabilities, including risks developers may fail to anticipate.
  • Bias and safety bounties: AI developers should pilot bias and safety bounties to extend bug-bounty practices and strengthen incentives for broad-based scrutiny of AI systems.These bounties can complement efforts to document datasets and models’ performance limitations and other properties.
  • Incident sharing: AI developers should share more information about incidents, including through collaborative channels, because incident sharing reveals overlooked risks and improves external scrutiny.Repeated sharing can also provide evidence that particular organizations find and acknowledge incidents, although additional mechanisms are needed.

3 Software Mechanisms and Recommendations

Software mechanisms can make AI claims more verifiable through audit trails, interpretability, and privacy-preserving machine learning (PPML). The section recommends developing standards, research, and open tools to address gaps in auditability, model understanding, and privacy-preserving implementation.

  • Audit trails: AI systems lack traceable records of problem definition, design, development, and operation, undermining accountability for claims about their properties and impacts.Audit trails could improve claim verification, but remain immature for AI and are expected to become increasingly important in safety-critical applications.
  • Audit trails: Standards bodies should define audit-trail requirements and guidance for safety-critical AI, balancing efficiency, completeness, tamperproofing, and other design considerations.Potential trails include code changes, training-run logs, and model outputs, while existing application-specific standards have not yet been established for AI.
  • Interpretability: Interpretability research should support risk assessment and auditing by developing shared criteria, studying model provenance, and constraining models to be interpretable by default.Interpretability tools depend on the target user and downstream task, but may become prerequisites for auditing sensitive-domain systems.
  • Privacy-preserving machine learning: PPML can safeguard sensitive data and models, but lacks evaluation standards and requires trade-offs among privacy, model quality, developer productivity, and computational overhead.Its implementation currently lies outside the typical AI developer’s skill set, and techniques must therefore be applied judiciously.
  • Privacy-preserving machine learning: AI developers should create, share, and use open PPML tools benchmarked against common performance standards to reduce implementation barriers and improve scrutiny.Repositories of real-world cases, implementation guides, and interoperability work could further support adoption and standardization.

4 Hardware Mechanisms and Recommendations

Hardware mechanisms can support verifiable claims about AI security, privacy, and development practices, but specialized hardware, performance costs, and incomplete verification limit current assurance. The section recommends stronger accelerator security, detailed compute accounting, and increased academic access to computing resources.

  • Hardware security: Secure enclaves remain less mature for machine learning, impose performance overhead, and are unavailable on specialized accelerators without significant development costs.Existing demonstrations have primarily used commodity CPUs and GPUs or outsourced portions of computation to less secure hardware.
  • Hardware security: Secure hardware can provide strong privacy and security assurances, while secure enclaves translate guarantees into verifiable hardware designs.Secure enclaves have emerged as a way to support claims that software alone cannot achieve.
  • Hardware security: ML-specific security features with remote attestation could guarantee that a model never leaves a particular chip, supporting more complex privacy and security policies.The section recommends collaboration between industry and academia to develop security features for AI accelerators or establish best practices for secure hardware use.
  • Compute reporting: More precise and standardized compute reporting could improve cross-organization comparisons and help third-party auditors identify discrepancies between available and reported computing power.The section recommends that one or more AI labs conduct a comprehensive compute-accounting pilot, while acknowledging costs, unclear success metrics, and limits to uniform reporting.
  • Access to computing resources: Government funding bodies should substantially increase academic access to computing resources so researchers can better verify claims made by industry.The recommendation connects researchers’ computing resources with their ability to scrutinize industry claims.

5 Conclusion

The report presents verifiable claims and mechanisms for assessing them as a way to earn trust in AI development, while emphasizing that these mechanisms enable incremental progress rather than decisively solving verification. Their effectiveness also depends on adoption, regulation, and collaborative action by AI developers and other stakeholders.

  • Contribution: Verifiable claims offer a way for AI developers to earn societal and mutual trust through mechanisms for making and assessing claims about AI development.The report frames these mechanisms as a richer toolbox for supporting trustworthy AI development.
  • Contribution: Adopting mechanisms for verifiable claims is presented as a second step after ethical principles establish standards against which behavior can be judged.The authors hope this framing will inspire meaningful dialogue within the AI community about approaching verifiability collaboratively.
  • Limitations: The proposed mechanisms support incremental improvements but do not decisively solve verification or ensure that AI developers behave responsibly.The report identifies three reasons for these limitations.
  • Limitations: Narrow system properties are easier to verify than broad claims, creating tension between claim verifiability and the generality of socially important claims.For example, performance on a particular safety metric is easier to verify than safety writ large, and circumscribed claims are easier to verify than broad societal-impact claims.
  • Limitations: Available verification mechanisms may not be demanded in practice, and power asymmetries may prevent corrective action even when claims are shown false.Consumers may prioritize convenience over stated values, while marginalized communities may lack the political power to resist harmful technologies; regulation is therefore required.
  • Path forward: Collaborative efforts to improve claim verifiability are necessary because wider deployment of AI in high-stakes tasks brings a growing range of risks.Without concerted action by developers and other stakeholders, society’s concerns about AI development are likely to grow.

Appendices · I Workshop and Report Writing Process

The report evolved from an interdisciplinary effort on trust in AI development toward a focused examination of verifiable claims. Its workshop and semi-modular, multi-stakeholder writing process incorporated expert and reviewer feedback but had representation and review limitations.

  • I Workshop and Report Writing Process: The project began with an interdisciplinary expert workshop in San Francisco in April 2019, involving academia, industry labs, and civil society organizations.The project initially addressed productive work related to trust in AI development.
  • I Workshop and Report Writing Process: During writing, the report shifted from trust in AI development broadly to verifiable claims in particular.
  • I Workshop and Report Writing Process: The workshop drew experts from dimensions of trust identified in a pre-workshop white paper, including secure enclaves, third-party auditing, and privacy-preserving machine learning.
  • I Workshop and Report Writing Process: Fewer than one third of workshop participants were women, and the workshop lacked greater gender diversity.
  • I Workshop and Report Writing Process: The corresponding authors led a multi-stakeholder writing, editing, and feedback process, later adding authors with expertise complementary to the original workshop participants.
  • I Workshop and Report Writing Process: The authors acknowledge that not all trust dimensions or verifiable-claim perspectives were represented and that not every author fully supported all content.Footnotes were included where appropriate to clarify process and attribution.
  • I Workshop and Report Writing Process: Experts drafted familiar mechanism or research-area subsections, which were substantially revised in response to author and reviewer feedback and changes in the report’s framing.
  • I Workshop and Report Writing Process: External reviewers assessed specific portions for clarity and accuracy, but the report as a whole was not formally peer reviewed.

II Key Terms and Concepts

The report frames AI systems as software running on hardware under human direction within institutions, making all three dimensions relevant to verifying developer claims. It treats trustworthy AI development as grounded in evidence that substantiates claims about behavior and calibrates trust.

  • AI system: An AI system is a software process running on physical hardware under human direction within an institutional context.The software, hardware, and institutions involved may each affect the verifiability of claims made by an AI developer.
  • AI development: AI development encompasses researching, designing, testing, deploying, and monitoring AI systems.The report uses this broad catch-all because the same mechanisms can apply across multiple development phases.
  • Responsible AI development: Responsible AI development seeks acceptably low risks of harm and ideally increases systems’ likelihood of benefiting society.It includes safety and security testing, pre-release social-impact evaluation, and willingness to abandon or delay projects that fail a high safety bar.
  • Transparency: Transparency means making information about an AI developer’s operations or systems available to actors inside and outside the organization.Implementing transparency requires institutional mechanisms and legal structures, while open publication may be limited by privacy, safety, and competition.
  • Trust and trustworthiness: The report treats verifiable claims as a building block of trustworthy AI development, using evidence about behavior to substantiate claims and calibrate trust.Scrutinizing an AI developer’s claims and commitments helps assess whether trust is appropriate in a given context.

III The Nature and Importance of Verifiable Claims

Verifiable claims are precise, falsifiable statements for which evidence and arguments can inform their likelihood of being true, although verifying AI-development claims is difficult because of complexity, distributed actors, rapid development, and vagueness. Their verifiability enables scrutiny, higher standards for responsible development, and reduced risk of harmful competitive pressures or uninformed decisions.

  • What makes claims verifiable: Verifiable claims are sufficiently precise to be falsifiable, with attainable certainty varying across contexts.The report uses “verifiable” in a broader sense than formal verification, which seeks mathematical proof under specified assumptions.
  • Challenges to verification: AI-development claims are difficult to verify because AI and its infrastructure are complex and heterogeneous, ecosystems are dispersed, development is fast, and claims are often vague.
  • Why verifiability matters: Verifiable claims let affected parties, civil society, policymakers, and users scrutinize developers’ claims, potentially reducing harm and improving societal outcomes.
  • Why verifiability matters: Without verifiable claims, developers may enter a race to the bottom that trades safety, security, privacy, or fairness for competitive advantage.In commercial and non-commercial contexts, verifiable claims may instead foster cooperation rather than race-like behavior.
  • Assurance cases and CAE: Assurance cases provide documented evidence and convincing arguments for top-level claims, while CAE structures claims, arguments, and evidence for safety, security, reliability, and dependability.CAE has been used in aviation, nuclear, and defense and is increasingly applied to AI safety analysis.

IV AI, Verification, and Arms Control

AI arms control depends on credible commitments and verification, but the technology’s dual-use nature, unclear boundaries, and difficult-to-monitor development make restraint challenging. Hardware may offer more governable verification points, while AI researchers can help design and scrutinize governance mechanisms.

  • Arms control foundations: Arms control relies on state cooperation, reciprocity, and verification regimes because treaties coordinate behavior but do not directly enforce compliance.States themselves must respond to violations through sanctions, reciprocal weapons development, military action, or other measures.
  • Arms control foundations: Verification difficulties can cause states to assume competitors are cheating, incentivizing reciprocal development of prohibited technologies.This dynamic is especially problematic when states believe secret development could provide a military advantage.
  • AI-specific challenges: AI-specific restraint is difficult because AI is widely available, dual-use or “omni-use,” hard to divide into acceptable and unacceptable applications, and challenging to verify.Broad bans are unlikely to succeed, whereas prohibitions on specific military applications could work if states agree to compatible limits.
  • AI community contributions: AI researchers can contribute technical expertise by evaluating governance proposals, limiting proliferation, and strengthening human accountability in military AI.The paper specifically suggests collaboration with arms control experts to scrutinize defensively oriented systems targeting lethal autonomous weapons but not humans.
  • Hardware-based verification: Hardware is uniquely governable in principle because chips and robots depend on physical materials and supply chains that are countable, trackable, and inspectable.By contrast, AI insights, data, code, and models can be reproduced and distributed at negligible marginal cost, making their spread difficult to control.

V Cooperation and Antitrust Laws

Collaborations among competing AI labs can support verifiable claims and trust but may raise US and international antitrust concerns. Appropriate governance should preserve procompetitive benefits while preventing collaborations that harm consumer welfare.

  • Legal context: Collaborations between competing AI labs can raise antitrust issues, warranting attention to international legal implications because AI development and markets are global.The section primarily addresses US antitrust law.
  • Legal context: US antitrust law seeks to prevent unreasonable restraints on trade, with courts generally guided by a consumer welfare test.Recent academic and popular proposals challenge this test, but consumer welfare remains the guiding principle for antitrust courts.
  • Collaboration tradeoffs: Collaborations between competitors are not always anticompetitive because they can enable new products, share useful know-how, and realize economies of scale and scope.These benefits must be balanced against possible harms from collaboration, including reduced competition.
  • Governance implications: With appropriate antitrust governance, joint activities among competitive AI labs can enhance consumer welfare and intra-industry trust without using verifiable claims to cover harmful practices.Practices that harm consumer welfare could erode trust between society and AI labs collectively.

VI Supplemental Mechanism Analysis · A Formal Verification

Formal verification for ML-based AI systems remains immature, owing to difficulties in specifying and modeling ML behavior and scaling verification to real-world models. The section identifies needs for AI-specific specifications and techniques, collaboration between ML and verification researchers, and verification of supporting software, while noting concrete implementation faults and gaps for Python.

  • A Formal Verification: Formal verification techniques for ML-based AI systems are still in their infancy.The section frames this immaturity as a central challenge for assuring AI systems.
  • A Formal Verification: ML verification requires formal claims and proofs despite unclear or environment-dependent model behavior, while traditional formal properties must be reconceived.Model outputs may differ in the field from behavior observed under testing.
  • A Formal Verification: Verification is hindered by difficulties modeling some ML systems mathematically and by real-world models exceeding existing techniques’ capacity.These limitations concern both formalizing system building blocks and handling model size.
  • A Formal Verification: Formal verification can reinforce non-ML software used to construct ML models, including through proof assistants that establish implementations are free of errors.This provides an avenue even while deep-neural-network verification remains difficult.
  • A Formal Verification: A demonstrated analysis found memory leaks leading to unpredictable behavior, crashes, and corrupted data.The cited faults included files opened without being closed and temporarily allocated data not freed.
  • A Formal Verification: Additional implementation risks included unchecked free-call results that could load incorrect weights and divide-by-zero faults that could crash online training.The section also identifies floating-point divide-by-zero issues in network cost calculation.
  • A Formal Verification: Formal verification for Python is inherently limited, leaving a major gap as safety-critical AI systems such as autonomous vehicles use the language.Python’s dynamic typing creates different errors from C and C++, and runtime type annotations are not enforced.
  • A Formal Verification: The section calls for AI-specific specifications and mathematical frameworks, novel verification techniques, and collaboration producing deep-learning systems more amenable to verification.These needs address assurance challenges created by the rapid introduction of ML into safety-critical environments.

B Verifiable Data Policies in Distributed Computing Systems · C Interpretability · What has interpretability research focused on?

Project Oak addresses the absence of enforceable data-use policies by combining secure infrastructure, verification, and transparency, while interpretability research spans explanations of predictions, global behavior, models, visualizations, and neural-network components.

  • B Verifiable Data Policies in Distributed Computing Systems: Project Oak provides open-source infrastructure for verifiably secure storage, processing, and exchange of any type of data.It aims to address the lack of mechanisms enforcing sharing and anonymity restrictions when data is shared with another party.
  • B Verifiable Data Policies in Distributed Computing Systems: Oak attaches enforceable policies to data and uses encrypted enclaves, remote attestation, and verifiable processing records to constrain access and use.Remote attestation ensures only appropriate code directly accesses secured data within the limits of a configurable policy.
  • B Verifiable Data Policies in Distributed Computing Systems: Oak combines enclaves, formal verification, remote attestation, and binary transparency so data moves only when receiving code is verified to obey its policy.Open-source development allows independent researchers, regulators, and consumer advocates to examine whether the implementation matches expected behavior.
  • B Verifiable Data Policies in Distributed Computing Systems: Oak’s usability and auditability depend on four actors: end users, application developers, policy authors, and verifiers.Verifiers add credibility by checking whether an Oak app upholds its policy.
  • B Verifiable Data Policies in Distributed Computing Systems: Oak’s success requires user-centered design for understanding policies, capturing preferences, delegating trust decisions, and handling changes to apps, policies, and assessments.Developers also need deliberate support for building, deploying, and debugging applications to avoid critical mistakes.
  • B Verifiable Data Policies in Distributed Computing Systems: Oak does not provide privacy by default; its utility depends on specifying and maintaining privacy-preserving policies and translating policy goals into enforceable rules.An Oak node without a correct and useful policy is useless.
  • What has interpretability research focused on?: Interpretability research has focused on explaining individual predictions, explaining global model behavior, building interpretable models, enabling interactive visualization, and analyzing neural-network sub-components.Prediction explanations may identify influential input regions or training examples, while global explanations may approximate complex models with simpler representations such as shallow decision trees.

Current directions in interpretability research … Hardware Mechanisms and Recommendations

The paper highlights interpretability research that combines interpretable-by-design models, interactive exploration, and practitioner tools while noting gaps in standardized benchmarks and novice-oriented software. Its recommendations call for institutional, software, and hardware mechanisms to strengthen scrutiny, auditing, privacy, security, and researchers’ ability to verify industry claims.

  • Current directions in interpretability research: Interpretability research pursues models constrained to be interpretable by design alongside tools for humans to explore, modify, and understand ML systems.Interactive tools include dashboards for predictions and errors, explanations of model representations, and direct interaction with AI agents.
  • Current directions in interpretability research: Most open-source interpretability code emphasizes new methods rather than standardized benchmarks and comparisons, limiting support for novice users and systematic evaluation.The paper calls for packages that help novices use interpretability techniques and provide standardized benchmarks for comparing methods.
  • Institutional Mechanisms and Recommendations: Institutional recommendations include third-party-audit task forces, organizational red teaming, bias and safety bounties, and collaborative sharing of AI-incident information.These measures are intended to broaden scrutiny, improve risk exploration, and strengthen incentives and processes for examining AI systems.
  • List of Recommendations for Reference: The recommendations prioritize third-party auditing and broad-based scrutiny as reference mechanisms for improving accountability in AI development.The proposed mechanisms include task forces, red teaming, bounties, and collaborative incident reporting.
  • Software Mechanisms and Recommendations: Software recommendations call for audit-trail requirements in safety-critical AI applications and research into interpretability focused on risk assessment and auditing.Standards bodies, academia, industry, AI developers, and funding organizations are identified as relevant participants.
  • Software Mechanisms and Recommendations: AI developers should develop, share, and use privacy-preserving machine-learning tool suites that measure performance against common standards.The recommendation emphasizes both practical tools and standardized performance measurement.
  • Hardware Mechanisms and Recommendations: Hardware recommendations include developing security features for AI accelerators, establishing secure-hardware practices, and estimating project computing power in high precision.Secure enclaves on commodity hardware are included, and AI labs are asked to report on wider adoption of high-precision compute measurement.
  • Hardware Mechanisms and Recommendations: Government funding bodies should substantially increase computing resources for academic researchers so they can better verify claims made by industry.The recommendation concerns computing-power resources for researchers in academia.
Loading 2004.07213v2…