Source-linked AI summary
From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good
Nitesh V. Chawla, Paulo Benanti
TL;DR
AI can reveal and alter institutional failures, so the paper links system evaluation to the rupture into which deployment occurs. It proposes RISE AI and bounded claims while preserving moral and political questions beyond measurement.
Problem
Responsible AI needs to evaluate both AI systems and the institutional ruptures they enter, because AI can reveal and intervene in existing failures of responsiveness, belonging, care, and accountability.
Method
The paper develops a rupture test and RISE AI architecture that connect institutional baselines, governance commitments, system evaluation, claim-evidence relations, context, power, and legitimate constraints.
Results
The paper concludes that evidence-bounded deployment limits claims to what has been evaluated, while measurement-bounded governance records constraints that favorable evidence cannot override.
Takeaways & Limitations
Responsible AI requires engineering and institutional repair to proceed together with continued moral and political judgment about dignity, power, and legitimacy.
Takeaways & Limitations
Evaluation can establish properties of a deployed system while leaving ownership, agenda-setting, sustaining labor, and institutional asymmetries unaddressed.
Abstract
from arXiv · showhide
Artificial Intelligence does more than create a governance problem. It can also reveal where institutions have already failed to provide responsiveness, belonging, care, and accountability. Once deployed, AI becomes an intervention in those conditions. It can repair, compound, substitute for, or conceal the failures it encounters. Responsible AI must therefore evaluate both the system and the institutional rupture into which it is introduced. The move from principles to protocols is already underway. The EU AI Act, NIST AI RMF, ISO/IEC 42001, and assurance practices translate commitments into roles, requirements, records, oversight, and assessment. The harder questions are what these protocols actually establish, whose power they leave untouched, and where measurement must stop. Pope Leo XIV's Magnifica Humanitas provides a broader moral frame centered on dignity, technological power, and the common good. Drawing on that frame, we develop a rupture test that links institutional baselines to system evaluation. We distinguish evidence-bounded deployment, which limits claims to what has actually been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. Within those limits, RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment. Responsible AI requires better engineering, institutional repair, and continued moral and political judgment.
1 Introduction
AI did not create longstanding institutional reductions of intelligence, education, relationships, and human worth; it scales them and makes them more visible. Because AI can intervene in the weaknesses it reveals, responsible evaluation must address both systems and institutional conditions.
- AI did not create reductions of intelligence to performance, education to measurement, relationships to transactions, or human worth to productivity; institutions had already begun making these reductions.
- People turn to AI for recognition, companionship, guidance, and affirmation partly because communities, educational systems, workplaces, and civic institutions often fail to provide them.
- AI reveals institutional failures while also intervening in them, potentially repairing, compounding, substituting for, or concealing those failures.
- Ethical AI already has a shared vocabulary translated unevenly into standards, curricula, management systems, assurance practices, and legal requirements.
- The paper proposes treating AI as both intervention and revelation while making evidence, authority, and unresolved questions explicit through RISE AI.
2 Human-Centered AI and the Social Order
A moral frame centered on dignity, power, and the common good extends responsible AI beyond technical artifacts to the institutional and political order surrounding them. The paper translates commitments into protocols while preserving questions that engineering evidence cannot settle.
- Magnifica Humanitas asks responsible AI to examine the economic, institutional, and political structures through which technology is conceived, financed, controlled, and used.
- Dignity governs institutional legitimacy and treatment of persons, while AI can foster participation and justice or intensify inequality, control, and exclusion.
- The encyclical connects dignity and the common good to impact assessment, vulnerable-population inclusion, digital literacy, worker protections, understandable decisions, contestation, and oversight.
- From Commitments to Protocols: The protocol chain moves from principle through design objective, system requirement, implementation mechanism, and evaluation protocol to a bounded evidence claim.
- From Commitments to Protocols: A rupture test identifies the prior institutional or relational failure, establishes a human or non-AI baseline, and checks whether AI repairs, compounds, substitutes for, or conceals it.
- Meaningful control requires disclosure, understandable reasons, alternatives, override and appeal channels, authoritative human review, and evaluation of their practical use.
- Responsible AI concerns the sociotechnical system rather than only the model, because workflow integration, overreliance, absent recourse, and weak institutions can make deployment unsafe or disempowering.
3 Design Patterns and Legal Protocols
The paper connects governance objectives to implementable design patterns and context-specific evaluation, while limiting deployment claims to collected evidence. It argues that legal and technical compliance cannot by itself establish empowerment, institutional repair, or legitimacy.
- Design Patterns and Legal Protocols: Answerability, contestability, agency-preserving interfaces, context-aware evaluation, and evidence-bounded deployment connect governance objectives to engineering mechanisms and evidence.
- Design Patterns and Legal Protocols: Across the design patterns, deployment claims should not exceed collected evidence, with gaps, conflicts, and expiry conditions triggering qualification, evaluation, remediation, or withdrawal.
- Design Patterns and Legal Protocols: Technical requirements can be satisfied while a system still violates a categorical constraint or operates within an illegitimate institutional order.
- EU AI Act Requirements and Mechanisms: The EU AI Act translates commitments into risk management, data governance, documentation, logging, human oversight, impact assessment, and conformity-assessment requirements for high-risk systems.
- EU AI Act Requirements and Mechanisms: Oversight requirements do not establish that reviewers have the time, authority, understanding, and institutional protection needed for meaningful control.
- EU AI Act Requirements and Mechanisms: Compliance alone does not establish empowerment, institutional repair, or a just distribution of technological power.
4 Limits of Design and Measurement
Design-centered evaluation cannot establish political legitimacy or resolve dignity, power, and institutional-order questions. RISE therefore pairs evidence-bounded claims with explicit constraints that favorable measurements cannot override.
- Political Economy and System Boundaries: A system may preserve user control and contestation while leaving ownership, financing, labor, infrastructure, and institutional asymmetries unexamined.System validity within a deployment context does not establish that the larger political-economic order is just.
- Dignity and Construct Validity: Measurement cannot establish whether a well-performing system respects dignity or treats people merely as means.Dignity is an inherent human status, not a variable, indicator, or outcome that can be offset by other measures.
- Dignity and Construct Validity: RISE evaluates evidence for variable sociotechnical conditions rather than measuring dignity, empowerment, responsibility, inclusion, or safety as latent attributes.Validity attaches to the inference from evidence to a bounded claim, not to the normative value itself.
- Measurement-Bounded Governance: Measurement-bounded governance records uses, actions, populations, and relationships that evidence cannot legitimize or substitute away.These commitments function as explicit, reviewable constraints rather than constructs to be measured.
- Measurement-Bounded Governance: Responsible AI requires better system design alongside moral and political judgment because benchmarks cannot establish legitimacy or settle questions of power and accountability.Technical and ethical agendas must proceed together.
5 RISE AI Evidence Architecture
RISE AI is a claim-evidence architecture for bounded system claims, linked to institutional context and power. It uses explicit baselines, indicators, provenance, thresholds, and expiry conditions to test whether evidence supports deployment claims.
- RISE AI Evidence Architecture: RISE organizes bounded claims around Responsibility, Inclusivity, Safety, and Empowerment rather than a universal governance score.Empowerment is assessed relative to an explicit baseline through capability, choice, contestation, and human or institutional support.
- RISE AI Evidence Architecture: A versioned claim-evidence graph links each claim to system versions, populations, contexts, owners, indicators, thresholds, provenance, limitations, and expiry triggers.Evidence may be direct, proxy-based, conflicting, missing, or stale.
- Operational Protocol and Institutional Baseline: RISE adds a context-and-power record covering ownership, financing, labor, resources, agenda-setting authority, and risk-benefit distribution.These fields expose the institutional order surrounding a supported claim and can qualify or preclude deployment.
- Operational Protocol and Institutional Baseline: The protocol scopes institutional rupture and baseline, formulates claims with affected parties, and predeclares indicators, methods, and thresholds before evaluation.Thresholds are contextual rather than universal.
- RISE AI Evidence Architecture: The rupture test compares an AI-mediated system with human and non-AI alternatives, testing both intended repair and plausible displacement.Improved accuracy or throughput may not support empowerment if access to human assistance or contestation declines.
- RISE AI Evidence Architecture: Across education and clinical decision support, performance metrics alone do not establish empowerment or sociotechnical safety.The required evidence includes learning transfer and authorship in education, and responsibility, override, subgroup outcomes, incidents, and correction in clinical settings.
6 Research Roadmap
The proposed architecture remains to be tested through participatory, reproducibility, consequential-validity, change-sensitivity, field, and decision studies. The paper also limits its claims because RISE is a proof of concept whose judgments remain shaped by institutional power and measurement risks.
- Research Roadmap: RISE should be evaluated for participatory content validity, inter-evaluator reproducibility, consequential validity, and sensitivity to change.These tests ask whether claims cover what matters, evaluators classify evidence consistently, profiles reveal failures, and conclusions respond to system changes.
- Research Roadmap: Field studies should involve regulators, providers, deployers, workers, and affected populations rather than treating compliance artifacts as self-interpreting.Comparative studies can test whether existing records and oversight measures support the inferences stakeholders need.
- Research Roadmap: Decision studies should compare choices made with and without RISE profiles to determine whether the framework changes consequential decisions.Documentation improvements alone are insufficient if they do not affect decisions.
- Limitations: The worked trace and cross-domain probes are proof-of-concept specifications, not empirical validation of RISE or an implementation study of the AI Act.Claim formulation and threshold setting still require normative judgment and are shaped by institutional power.
- Limitations: Evidence may underrepresent affected people, while logging and provenance can create privacy and surveillance risks.Political-economic analysis can also become superficial, and measurement cannot resolve the underlying power relations.
7 Conclusion
Responsible AI must evaluate both deployed systems and the institutional ruptures they enter. Protocols and evidence can improve accountability, but they cannot replace judgments about dignity, power, legitimacy, and institutional repair.
- AI exposes institutional failures and can repair, compound, substitute for, or conceal them, making evaluation of the surrounding rupture necessary.
- Existing protocols translate values into roles, requirements, records, assessments, and decision gates, but their evidence and omissions still require scrutiny.
- Magnifica Humanitas extends responsible AI beyond system design to ownership, infrastructure, financing, labor, political authority, and technological power.
- RISE AI links claims to evidence and context-and-power records while distinguishing evaluated claims from constraints that favorable evidence cannot override.
- Responsible AI requires engineering, institutional repair, and moral-political judgment to proceed together rather than treating improved performance as sufficient.