Source-linked AI summary
Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem
Sasha Costanza-Chock, Emma Harvey, Inioluwa Deborah Raji, Martha Czernuszenko, Joy Buolamwini
TL;DR
AI audits are proliferating while audit processes remain unstandardized, poorly understood, and unsupported by widely used standards or regulatory guidance. This paper scans the emerging audit ecosystem through a catalog, survey, and interviews, finding rapid growth and overwhelming support for mandating audits alongside perceived regulatory gaps.
Problem
AI audit practices remain poorly understood and lack widely used standards or regulatory guidance, while audit services proliferate without standardized processes.
Method
The paper provides an overview of the AI audit ecosystem using a catalog of relevant participants, an anonymous survey, and interviews with industry leaders.
Results
95% of non-auditors and auditors alike agree that AI audits should be mandated, while the ecosystem is growing rapidly and practitioners believe current regulation is lacking.
Takeaways & Limitations
The findings support developing clearer standards and regulation for AI audits as the ecosystem expands.
Takeaways & Limitations
The authors identify limits to the accuracy and generalizability of their findings and call for future research on algorithmic audits in the Global South.
Abstract
from arXiv · showhide
AI audits are an increasingly popular mechanism for algorithmic accountability; however, they remain poorly defined. Without a clear understanding of audit practices, let alone widely used standards or regulatory guidance, claims that an AI product or system has been audited, whether by first-, second-, or third-party auditors, are difficult to verify and may exacerbate, rather than mitigate, bias and harm. To address this knowledge gap, we provide the first comprehensive field scan of the AI audit ecosystem. We share a catalog of individuals (N=438) and organizations (N=189) who engage in algorithmic audits or whose work is directly relevant to algorithmic audits; conduct an anonymous survey of the group (N=152); and interview industry leaders (N=10). We identify emerging best practices as well as methods and tools that are becoming commonplace, and enumerate common barriers to leveraging algorithmic audits as effective accountability mechanisms. We outline policy recommendations to improve the quality and impact of these audits, and highlight proposals with wide support from algorithmic auditors as well as areas of debate. Our recommendations have implications for lawmakers, regulators, internal company policymakers, and standards-setting bodies, as well as for auditors. They are: 1) require the owners and operators of AI systems to engage in independent algorithmic audits against clearly defined standards; 2) notify individuals when they are subject to algorithmic decision-making systems; 3) mandate disclosure of key components of audit findings for peer review; 4) consider real-world harm in the audit process, including through standardized harm incident reporting and response mechanisms; 5) directly involve the stakeholders most likely to be harmed by AI systems in the algorithmic audit process; and 6) formalize evaluation and, potentially, accreditation of algorithmic auditors.
1 INTRODUCTION
AI audits are increasingly used for algorithmic accountability, but audit practices, standards, and regulatory guidance remain underdeveloped. This paper scans the ecosystem to identify practices, barriers, and policy recommendations for more effective audits.
- Audit services have proliferated while audit processes remain unstandardized and poorly understood.
- AI audits evaluate automated decision systems against criteria and may assess bias, effectiveness, transparency, security, compliance, consent, labor practices, and energy use.
- Without shared practices, standards, or regulatory guidance, claims that systems were audited are difficult to verify and may exacerbate harm.
- The study catalogs 438 individuals and 189 organizations, surveys 152 respondents, and interviews 10 industry leaders.
- The authors identify emerging best practices, commonly used methods and tools, barriers to accountability, and six policy recommendations.
- Recommendations include independent audits against defined standards, notification, peer-review disclosure, harm reporting, stakeholder participation, and auditor evaluation or accreditation.
2 BACKGROUND
Algorithmic auditing has expanded across companies, contractors, journalists, civil society, regulators, and researchers, while methods and accountability expectations remain contested. The background highlights recurring concerns about disclosure, stakeholder inclusion, real-world harm, and operationalizing principles.
- Algorithmic systems have produced documented harms involving discrimination, wrongful benefit denials, kidney-transplant and mortgage decisions, and facial-recognition arrests.
- Auditing is presented as one potential way to improve algorithmic accountability and expose evidence that deployments fall short of performance claims.
- The ecosystem includes first-party internal teams, second-party contractors, and third-party independent auditors, with each arrangement offering different access and independence conditions.
- First-party audit results are typically undisclosed, while second-party examples have raised concerns about company funding, employee co-authorship, misrepresentation, and confidentiality.
- Third-party audits have helped create public awareness of algorithmic harms through work by journalists, civil society groups, regulators, and researchers.
- Few widely adopted standards exist, and debates continue over audit content, suitable tools, stakeholder involvement, use context, socioeconomic impacts, and alternative accountability approaches.
- The study focuses on practitioner methods, tools, emerging best practices, real-world harm, stakeholder engagement, and similarities and differences among auditor groups.
3 METHODS
The study combines a field scan, interviews, and a survey to examine the emerging audit ecosystem, while acknowledging limits in representativeness, geographic coverage, and the boundaries of the audit framing.
- 3 METHODS: The field scan identified 438 individuals from 189 organizations involved in auditing or related work, including auditors, advocates, regulators, and researchers.
- 3 METHODS: The researchers conducted semi-structured interviews with 10 field leaders and then surveyed contacts at all 189 identified organizations.
- 3 METHODS: 152 people responded to the survey, including 56 who reported personally participating in an AI audit.
- 3.1 Limitations: The findings may not generalize to the entire auditor population because only 56 respondents reported direct audit experience and newer audits were not fully represented.
- 3.1 Limitations: Survey coverage was concentrated in the US, UK, and EU, with minimal or no responses from several other regions.
- 3.1 Limitations: The professional-auditor focus may have excluded practitioners who use other terms or approaches, including community-based, participatory, or evocative audits.
- 3.1 Limitations: The study did not collect auditors’ demographic information and identifies systematic demographic research as future work.
4 KEY FINDINGS
The field scan finds that AI audits are predominantly quantitative, customized, and weakly documented, while auditors broadly support regulation but disagree about its scope and institutional design.
- Quantitative over qualitative: 77% of auditors assess algorithmic accuracy, fairness, and statistical soundness, while 51% assess systems for reporting real-world harm.Training-data quality is also assessed by 77% of auditors.
- Quantitative over qualitative: The four most common audit methods are quantitative checks of training-data suitability, representativeness, input-data bias, and individual-level accuracy.Each method is used by over 70% of respondents.
- Customized tools and thin standards: Only 7% use a standardized audit framework and toolset, while 38% use none of the specified pre-built audit tools.Practitioners describe their frameworks and tools as custom-built and tailored to particular use cases.
- Protected classes and intersectionality: Most auditors assess legally protected classes, but evidence for intersectional fairness assessments is limited and legal concerns may constrain de-biasing efforts.Age, race/ethnicity, and sex are the three most frequently assessed classes; only 65% self-report conducting intersectional assessments, and few provided examples of methods or outcomes.
- Disclosure and peer review: Only seven of 43 auditors provide documentation of their audit process, and just four link to audit results, despite 82% supporting public availability in principle.Respondents identify best-in-class auditors as those that publish methodologies and results.
- Regulation and standards: 95% support mandated audits, but respondents divide over whether mandates should cover only high-stakes systems or all AI systems.53% support high-stakes-only mandates, while 42% support mandates for all AI systems; 73% favor decentralized domain-specific government regulation.
5 DISCUSSION AND RECOMMENDATIONS
The AI audit ecosystem is growing rapidly but lacks consistent standards, regulation, disclosure, and enforcement. The authors identify strong support for mandated audits and disclosure while highlighting gaps between auditors’ stated priorities and current practice, motivating six policy recommendations.
- Discussion: N=438 individuals and N=189 organizations indicate a rapidly growing but still nascent AI audit ecosystem, increasing the need for standards and regulatory oversight.The authors link this growth to discrepancies between desired best practices and reality.
- Discussion: 95% support mandating AI audits, while 82% support disclosing audit results in part or in full.This consensus coexists with concerns about overly broad requirements producing cursory checks and overly specific requirements producing narrow assessments.
- Discussion: Auditors report limited buy-in, enforcement power, data access, and disclosure, constraining whether audits lead to changes in AI systems or engineering practices.Over half of surveyed auditors report neither commitment from auditees to address problems nor power to require changes.
- Discussion: Auditors often value intersectional analysis, real-world harm assessment, and stakeholder involvement, but these practices are rarely implemented.65% express interest in intersectional analysis, 65% consider real-world harm important, and 41% consider stakeholder inclusion important in theory.
- Recommendations: The authors recommend independent audits against defined standards, notification, key-results disclosure, harm reporting, stakeholder involvement, and auditor evaluation or accreditation.The first four recommendations have broad support, while stakeholder involvement and auditor accreditation are more controversial among practitioners.
- Recommendations: Notification can enable individuals to seek information, contest decisions, and support community oversight, but policy alone does not guarantee implementation.The authors therefore recommend evaluating notification, opt-out, and functional appeal mechanisms, noting that GDPR notification duties are not always respected.
6 CONCLUSION
The paper maps a rapidly growing AI audit ecosystem and finds broad agreement that current regulation is inadequate. It identifies consensus around mandatory audits, notification, and disclosure, while highlighting practical gaps and policy recommendations.
- The AI audit ecosystem is growing rapidly, while practitioners overwhelmingly view current regulation as lacking.
- Auditors broadly support mandatory audits, individual notification, and disclosure of key audit findings, though implementation details remain debated.
- Practitioners report a mismatch between what auditors consider important and what they can accomplish in practice.
- The recommendations call for independent audits against clear standards, applicable to AI owners and operators.
- They also recommend notification, peer-reviewable disclosure, real-world harm reporting, stakeholder participation, and possible auditor accreditation.
A.1 Field Scan
The field scan assembled a broad registry of people and organizations involved in AI auditing or closely related work. It included auditors, advocates, researchers, regulators, and other relevant practitioners.
- 438 individuals from 189 organizations were identified as involved to some degree in AI auditing.
- The registry included first-, second-, and third-party auditors, civil-society and advocacy members, academic researchers, and regulators.
- Ten field leaders were selected for semi-structured interviews.
- The survey was distributed to contacts at all 189 identified organizations.
A.2 Interviews
The interview study used ten semi-structured interviews with AI-auditing leaders to examine organizational roles, industry practices, barriers, regulation, and approaches to real-world harm.
- The researchers conducted ten semi-structured interviews with individuals identified as leaders in AI auditing.
- Interviews lasted 45–60 minutes, were conducted and recorded via Zoom, and followed informed consent procedures.
- Interview questions covered participants’ organizations, industry perspectives, and navigation of potential or actual harm.
- The interviews addressed audit methods, tools, emerging standards, best practices, barriers, and desired regulatory oversight.
- Questions on harm examined stakeholder involvement, harm investigation across the AI lifecycle, and reporting for deployed systems.
- Interviewers used examples involving incident reporting, marginalized-community participation, and harms during development or after deployment.
A.3 Survey
The survey gathered 152 responses from a geographically diverse group identified through direct invitations and public outreach. Respondents included both auditors and non-auditors, with most respondents based in the United States.
- The survey was sent to identified individuals through email, LinkedIn, and Twitter, then released publicly through social media and a newsletter.
- Respondents represented six continents and twenty-five countries, with 59% from the United States.
- 37% of respondents had personally worked on an AI-system audit, while 63% identified as non-auditors.
- Among auditors, 61% were from the United States, representing 34 of 56 individuals.
B AI AUDIT FIELD SCAN
The field scan provides linked supplementary materials documenting the ecosystem of organizations and individuals involved in algorithmic audits and the interview process.
- A linked spreadsheet catalogs organizations and individuals involved to some extent in algorithmic audits.
- A linked interview guide contains an introduction, directions, and a research interview consent agreement.
D SURVEY INSTRUMENT
The survey instrument is accompanied by a linked PDF containing the survey questions and the complete list of AI audit toolkits mentioned in Section 4.1.2.
- A linked PDF provides the survey questions and includes the full list of AI audit toolkits mentioned in Section 4.1.2.