Source-linked AI summary
Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, Daniel E. Ho
TL;DR
Algorithmic accountability policies have emphasized audits while neglecting the institutional conditions needed for effective third-party oversight. The paper synthesizes audit systems and evidence from other domains to identify design choices for AI auditing, concluding that audits alone are insufficient for meaningful accountability.
Problem
Algorithmic accountability approaches have paid insufficient attention to how third parties can participate effectively in oversight.
Method
The paper examines current AI audit challenges, surveys audit systems across other domains, and synthesizes evidence around five institutional design dimensions.
Results
The survey finds that audits vary substantially in scope, independence, access, standards, and post-audit actions, so third-party audit effectiveness depends on institutional design.
Takeaways & Limitations
AI governance should deliberately design audit ecosystems that enable third parties to confront, verify, and scrutinize performance claims.
Takeaways & Limitations
The evidence review is a selective and incomplete first step across diverse audit fields, and the interaction among design dimensions remains for future work.
Abstract
from arXiv · showhide
Much attention has focused on algorithmic audits and impact assessments to hold developers and users of algorithmic systems accountable. But existing algorithmic accountability policy approaches have neglected the lessons from non-algorithmic domains: notably, the importance of interventions that allow for the effective participation of third parties. Our paper synthesizes lessons from other fields on how to craft effective systems of external oversight for algorithmic deployments. First, we discuss the challenges of third party oversight in the current AI landscape. Second, we survey audit systems across domains - e.g., financial, environmental, and health regulation - and show that the institutional design of such audits are far from monolithic. Finally, we survey the evidence base around these design components and spell out the implications for algorithmic auditing. We conclude that the turn toward audits alone is unlikely to achieve actual algorithmic accountability, and sustained focus on institutional design will be required for meaningful third party involvement.
1 INTRODUCTION
Third parties have exposed consequential algorithmic harms, yet policy proposals often center audits controlled by the companies being assessed. Lessons from ESG and algorithmic auditing suggest that audit design must address independence, scope, access, standards, and transparency.
- Third-party oversight: Third parties have uncovered algorithmic bias and harm across domains including content curation, hiring, criminal justice, and public health.These actors include regulators, academics, civil society, journalists, and specialized law firms.
- Policy gap: Many algorithmic accountability proposals emphasize internal audits commissioned, executed, paid for, and controlled by targeted companies.The paper identifies this pattern in the GDPR, the proposed Algorithmic Accountability Act, and U.S. state and municipal proposals.
- Lessons from ESG: ESG illustrates how broad and ambiguous audit objectives can produce costly certifications that risk becoming “cheap talk” rather than evidence of social responsibility.The paper also notes conflicts involving auditor selection, payment, employment, access, diffuse standards, and nondisclosure agreements.
- Implications for AI audits: Algorithmic audits face analogous risks because auditor conflicts can compromise quality, while unclear scope and standards can make assessments unfocused and expensive.The paper highlights third-party auditor scope, independence, access, standards, and transparency as neglected design dimensions.
- Implications for AI audits: Effective algorithmic accountability requires an intentionally designed audit ecosystem that enables third-party participation rather than relying on audits alone.The proposed design priorities include target identification, independent selection and compensation, data and system access, clear standards, and post-audit transparency.
2 THE CURRENT U.S. AI AUDIT POLICY LANDSCAPE
The U.S. AI audit landscape largely frames audit activity as internal compliance and gives limited attention to affected communities, conflicts of interest, auditor diversity, and access. These constraints have contributed to audit failures and weakened external oversight.
- Audit concepts: AI audits evaluate systems against articulated expectations, standards, or claims, but may not capture broader downstream impacts of deployment.Audits can be qualitative or quantitative and focus on comparing product behavior with clearly specified expectations.
- Audit concepts: First-party audits are conducted by companies, second-party audits by contractual counterparties, and third-party audits by entities specifically engaged under predetermined standards.External oversight differs from internal oversight because external parties lack direct employment or contractual ties to the target.
- Internal-audit focus: Current policy proposals commonly assume that internal stakeholders will conduct or coordinate algorithmic impact assessments and compliance audits.This pattern appears in guidance for internal audits, the Algorithmic Accountability Act, and mainstream interpretations of algorithmic impact assessments.
- Internal-audit focus: Internal accountability can miss affected communities’ concerns because organizations often prioritize compliance, customer, or user interests over impacted non-users.Examples include policing tools prioritizing police needs and moderation systems that may further marginalize users experiencing hate online.
- Internal-audit focus: Conflicts of interest can suppress internal critiques and manipulate consultant reporting when companies control publication or communication of audit results.The paper cites Google, Facebook, and HireVue’s ORCAA audit as examples.
- External-audit capacity: Policy proposals also narrow external oversight to selected academics, regulators, or enforcement agencies, while limited audit capacity has produced missed demographic disparities and assessment challenges.The paper contrasts these limits with broader third-party participation and cites the Gender Shades intervention as exposing shortcomings in facial-recognition testing.
3 THE INSTITUTIONAL DESIGN OF AUDIT SYSTEMS
Audit systems across finance, environment, telecommunications, transportation, and healthcare use varied institutional arrangements rather than a single model. Their effectiveness depends on how scope, independence, access, professional standards, and post-audit actions are designed.
- Design dimensions: The paper’s typology covers target identification and scope, auditor independence, auditor privileges, professionalization and conduct standards, and post-audit actions.These dimensions address who is audited, who can audit, what access and protections auditors receive, how they operate, and what follows the audit.
- Survey scope: The paper examines audit systems across financial services, environmental protection, telecommunications, transportation, and healthcare, including systems already intersecting with AI.Examples include NTSB analysis of self-driving crashes and FDA approval of AI-enabled medical devices.
- Cross-domain patterns: Across surveyed systems, specific standards commonly scope audits, while third-party arrangements vary between public agencies and private parties.Public inspections may rely on complaints or risk-based heuristics because regulatory capacity is limited.
- Cross-domain patterns: All surveyed audit systems provide some access to otherwise confidential information, contrasting with the access challenges faced by third-party AI oversight.Access may include entering facilities, inspecting records, or reviewing financial statements.
- Cross-domain patterns: A third-party audit can function like a first-party audit when the auditee controls auditor selection, compensation, scope, access, and post-audit actions.The paper identifies this arrangement as questionably effective for achieving downstream accountability.
4 THE EVIDENCE BASE FOR EFFECTIVE AUDITS
The paper evaluates which institutional design choices may improve audit effectiveness by organizing evidence from social science and other industries around five design categories. It then develops implications for AI auditing.
- Evidence synthesis: The evidence review examines audit design considerations, relevant findings from social science and other industries, and implications for the AI context.The discussion is organized along the five design categories introduced earlier.
4.1 Target Identification & Audit Scope
Effective audit ecosystems must identify relevant targets and define precise scopes, while creating reporting channels that surface harms despite uneven disclosure and reporting behavior.
- Audit Selection: Risk-based or complaint-based selection focuses scarce audit resources on parties or issues most likely to involve harm.Incident reporting systems provide real-time signals for prioritizing secondary inspections.
- Audit Selection: Complaint data can predict later performance and genuine safety concerns, but voluntary reporting may underrepresent low-income and minority communities.Reported biases make active solicitation and adjustment for differences in reporting propensity important.
- Audit Precision: Audits should be precisely scoped because vague or overly broad mandates make results difficult to interpret or translate into enforcement.Existing policy definitions provide limited guidance for translating broad system descriptions into specific audit scopes.
- Audit Selection: A national incident reporting system could broaden third-party complaints, but its design must address reporting disparities and solicit minority perspectives.Companies may also use notices to encourage harm reporting, although excessive notices can produce consent fatigue and distract oversight.
- Audit Selection: Transparent disclosure of algorithmic products is necessary to identify audit targets and incidents of algorithmic harm.Current harm discovery often relies on ad hoc notice and reporting through public forums.
4.2 Independence
Audit independence depends heavily on who selects and pays auditors and on how long auditors remain familiar with auditees. Evidence supports reducing direct dependence on targets, while rotation involves an expertise tradeoff.
- Selection and Compensation: Cross-selling non-audit services is associated with more client-favorable or lower-quality audits in several domains.The evidence is not uniform: one study found no effect on independent expert reports for takeover targets.
- Selection and Compensation: Auditors should ideally not be selected or paid directly by auditees because target control can undermine independence and audit quality.A randomized trial found environmental audits were more accurate when paid from a government-distributed common pool.
- Auditor Tenure: Longer auditor–client familiarity can increase leniency, although tenure also builds firm-specific expertise needed to understand complex data ecosystems.Evidence from supply-chain audits and food-safety inspections illustrates the tension between independence and expertise.
- Auditor Tenure: Evidence on mandatory audit rotation is mixed, with some financial-audit evidence suggesting fresh-look benefits alongside limited overall benefits.Financial restatements increased during the first two years after rotation in one study.
- AI Auditing Implications: AI policy proposals use divergent definitions of independent audits, including arrangements where the target directly chooses and pays the auditor.Voluntary Pymetrics and HireVue audits illustrate how target involvement can compromise independence or control scope and communication.
4.3 Auditor Access
Third-party audits require meaningful access to data, systems, premises, and records, but current arrangements often leave access under target control or expose outsiders to legal risk.
- Access Requirements: Robust third-party review requires access to sufficient data, systems, premises, records, and documentation.Across regulatory domains, privileged information access is a recurring feature of external audits.
- Access Requirements: Announced inspections can permit targets to game the process, making unannounced access relevant to measuring actual performance.Research on NHS hospitals examined whether cleanliness differed between announced and unannounced inspections.
- Access Constraints: Management pre-screening of information and interviews can cause audits to omit major violations and contain serious deficiencies.The scope and conditions of access therefore affect audit quality, not merely audit convenience.
- Access Constraints: Third parties may face legal risk, corporate retaliation, or obstruction when seeking access to algorithmic systems.Some organizations obtained access to analyze Facebook’s ad-delivery discrimination only through a legal settlement.
- Access Constraints: Target-controlled access can produce highly limited audits that review documentation without independently evaluating data or models.The cited HireVue audit examined documentation for one assessment and excluded independent evaluation of its data and models.
- AI Auditing Implications: Proprietary-information concerns do not require withholding access; controlled mechanisms such as custom APIs can protect models while enabling review.The paper identifies lack of access as the most significant vulnerability in the current AI audit ecosystem.
4.4 Professionalization
Professionalization can improve audit quality through training, standards, accreditation, and public accountability, but overly rigid standards and licensing raise important tradeoffs.
- Training: Training and peer review improve audit accuracy and consistency, while experienced and client-knowledgeable teams conduct more effective audits.Trained supply-chain monitors found significantly more violations, and food-safety inspection consistency improved with training and peer review.
- Standardization: Clear, rigorous standards can reduce audit errors, limit auditor shopping, and improve quality by prompting lower-quality auditors to exit.However, overly detailed standards may encourage checklist compliance and reduce auditor agency in outlier situations.
- Accreditation and Licensing: Accreditation is widespread, but evidence on professional licensing and quality is mixed.Studies of CPA and dental licensing found no meaningful improvement in several quality or outcome measures.
- Transparency: Public or vetted access to audit outcomes can improve auditor conduct, accountability, and communication quality.Some U.S. financial auditors must make communications and records accessible to shareholders.
- Implications for AI Auditing: The paper supports training, standardization, and accreditation for third-party AI auditors, while preferring professional accreditation over broad occupational licensing.The rationale includes licensing’s potential to reduce service supply and government agencies’ limited near-term capacity.
- Implications for AI Auditing: Certification should extend beyond academic researchers to public-interest groups, law firms, and journalists with different audit strategies and community-relevant concerns.The paper links this recommendation to the capacity gap and the limits of focusing audit access on academic researchers.
- Implications for AI Auditing: Professional standards can define limited legal-immunity conditions, while serious misuse of accessed data can justify credential revocation.Examples include selling target information, conducting unrelated activities with accessed data, or violating security and privacy measures.
4.5 Post-Audit Actions
Post-audit disclosure and repeated assessment can strengthen accountability, verification, and standard-setting, but transparency requirements also impose costs and can conflict with proprietary interests.
- Disclosure and accountability: Public audit results can prevent companies from hiding undesirable outcomes and incentivize better behavior.Examples include searchable vehicle-safety results and market consequences after disclosed safety and efficacy overstatements.
- Disclosure and accountability: Public release enables third-party verification and is associated with more accurate audits, especially when auditors face conflicts of interest.Confidential audits can shield practices from verification by other researchers.
- Transparency trade-offs: Mandatory disclosure can be costly, inconsistent, and resource-intensive, while broad AI audits may be especially expensive for newer startups.The cited evidence includes average Sarbanes-Oxley compliance costs of $2.2 million and deficiencies in restaurant grading across ten jurisdictions.
- Transparency trade-offs: Restricted transparency can balance proprietary protections with external accountability by registering reports and releasing summaries publicly or upon vetted request.The proposed registry is intended to avoid both pre-publication review bias and secrecy under nondisclosure agreements.
- Repeated assessment and standard-setting: Repeated assessments can reveal continued discriminatory behavior after earlier assurances, while audit findings can inform concrete general standards.The paper connects recurring audits of Facebook’s job and housing ads to later standard-setting, including IEEE P7013 and the Gender Shades audit.
5 LIMITATIONS
The paper presents its cross-domain audit synthesis as an initial, selective account and clarifies that it does not reject internal audits or sharply separate inspections from audits.
- Scope and evidence: The evidence base is necessarily selective and incomplete because the paper draws on many different regulatory areas.The authors describe the work as a first step toward learning from cognate audit schemes.
- Scope and evidence: The paper does not deprecate internal audits, but argues that internal audits alone are insufficient for accountability.The authors note that internal audits are valued in other industries, including finance.
- Scope and evidence: The analysis commingles public and private inspections with audits because regulatory schemes and scholarship often treat them as functionally similar.The comparison centers on shared issues such as product defects.
- Unexamined interactions: The paper acknowledges that interactions among institutional design dimensions may matter, such as access depending on auditor independence.It leaves examination of these interactions to future work.
- Unexamined interactions: The paper does not specify whether proposed reforms should be implemented through legislation, regulation, enforcement, or self-regulation.The authors identify implementation choice as an important next step.
6 CONCLUSION
The paper concludes that calling for audits is insufficient for algorithmic accountability. It advocates deliberate institutional design that enables third parties to scrutinize corporate performance claims and address complaints of harm.
- Conclusion: Audits alone cannot address algorithmic accountability challenges without interventions supporting effective third-party participation.The conclusion characterizes reliance on audits without such design as a danger.
- Conclusion: AI policy can draw on audit schemes from other industries to develop an ecosystem in which third parties can survive and thrive.The proposed ecosystem supports confronting, verifying, and scrutinizing corporate performance claims.
- Conclusion: The proposed ecosystem is intended to support external scrutiny and complaints of harm from impacted populations.The conclusion frames this scrutiny as directed toward protecting people vulnerable to harm.
A APPENDIX
The appendix provides additional context and definitions for the policy survey, including tables summarizing audit ecosystem features and selected U.S. audit programs.
- Appendix: The appendix provides additional context and definitions for the policy survey.
- Appendix: Table 1 summarizes audit systems by audit ecosystem features and groups them as first-, second-, or third-party audits.Rows represent distinct audit systems, while columns represent five main design features.
- Appendix: Table 2 presents an overview of selected U.S. audit programs.