Source-linked AI summary
Assessing Company Contributions to Societal Resilience: Extending the Societal Capacity Assessment Framework to Agentic AI
Catherine Simons, Alexander K. Saeri, Peter Slattery, Neil Thompson
TL;DR
Downstream companies shape the sociotechnical conditions and societal resilience associated with agentic AI, but existing assessment approaches do not adequately capture that role. The paper extends SCAF into a 16-indicator company-level framework and applies it to Microsoft, surfacing stronger coverage where resilience aligns with product strategy and weaker coverage where governance commitments conflict with commercial interests.
Problem
Existing assessment approaches focused on internal risk management and model capabilities do not capture how companies that develop and deploy agentic AI shape societal resilience.
Method
The paper adapts SCAF from country-level societal conditions to company contributions, operationalizes 16 agentic-AI resilience indicators, and assesses Microsoft’s public-facing documents.
Results
Microsoft documents show stronger indicator coverage where resilience-building aligns with product strategy and weaker coverage for commitments that conflict with commercial interests.
Takeaways & Limitations
Company-level SCAF can identify risk-response gaps in corporate AI governance and connect missing mitigations to societal resilience.
Takeaways & Limitations
Because the corpus contains only public documents, the assessment measures Microsoft’s disclosed posture and should treat absent evidence as not disclosed rather than absent in fact.
Abstract
from arXiv · showhide
Companies that deploy AI agents and make them available to others are creating the sociotechnical circumstances under which this technology integrates into existing social and economic structures. AI-deploying companies are institutional actors that actively shape society's capacity to withstand and govern the consequences of agentic AI. In view of these societal impacts, companies can build societal resilience by designing and promoting safer implementations of AI agents. To operationalize this goal, this paper adapts the indicator-based Societal Capacity Assessment Framework (SCAF) to measure how a company's deployment decisions contribute to societal resilience, inverting its original measurement of societal resilience as a backdrop for deployment decisions (Gandhi et al., 2025). Our procedure has two steps: a conceptual step in which we design a suite of indicators that define what SCAF's vulnerability, coping, and adaptive capacities mean when assessing a company's agentic AI deployment decisions; and a measurement step in which we apply this framework in a structured assessment of public-facing Microsoft documents.
1 Introduction
The paper argues that agentic AI deployment decisions shape societal impacts and asks how company contributions to societal resilience can be assessed. It answers by adapting SCAF into resilience indicators and applying them to Microsoft’s public documents.
- Motivation: Agentic AI risks depend not only on technical capabilities but also on deployment choices and the sociotechnical conditions surrounding users, operators, institutions, and communities.The paper frames safer diffusion as a matter of choices made when systems are first introduced.
- Concepts: The paper defines societal resilience as the capacity to absorb immediate shocks, adapt to changing risks, and transform access to resources amid present and future shocks.AI agents are systems that iteratively plan, reason, and use tools for real-world tasks; agentic AI denotes the broader products and platforms built around them.
- Motivation: Companies that deploy AI agents shape which use cases spread, how much autonomy is default, and which safeguards protect against downstream societal effects.These decisions affect third-party risk exposure, cultural norms, and institutional functioning.
- Research gap: SCAF offers an indicator-based approach for vulnerability, coping, and adaptive capacity, but its country-oriented prototype lacked a comparable company-level assessment.The paper extends SCAF because companies’ deployment actions can shape societal impacts.
- Contribution: The study prototypes resilience indicators grounded in high-priority agentic-AI risk-management actions and applies them to Microsoft’s public documents as an illustrative case study.This operationalizes the research question of how to assess company contributions to societal resilience.
2 Background and Motivation
The background motivates examining downstream deployers as contributors to societal resilience, while warning that resilience must remain a concrete analytical concept rather than a slogan. The paper therefore adapts SCAF’s indicator-based approach from societal conditions to company contributions.
- Conceptual caveat: The paper cautions that resilience can become an empty buzzword or shift responsibility onto downstream actors expected merely to adjust to technological change.This concern motivates keeping the concept tied to concrete governance practices and indicators.
- SCAF adaptation: SCAF provides a template for measuring vulnerability, coping capacity, and adaptive capacity, and this paper reverses its original use as a backdrop for organizational risk management.The adaptation measures company contributions to societal resilience instead of societal conditions alone.
- Governance gap: Existing AI-governance assessments largely examine frontier AI companies, leaving downstream deployers that determine how agents reach society comparatively unexamined.Frontier-company assessments are partly favored because concentrated, compute-intensive supply chains create potential intervention points.
- Societal resilience: Societal resilience is presented as a response to wicked AI-governance problems because it supports absorbing, adapting to, and recovering from unspecified surprises while preserving actor agency.This approach does not depend on enumerating every failure or pretesting every solution.
- Governance approach: The framework supports decentralized AI governance by assuming that many companies can build societal resilience independently of centralized governance focused on frontier AI companies.The paper also notes that organizations struggle to scale agents because of governance and organizational constraints rather than technological limits.
3 Methods
The method translates SCAF’s resilience capacities into company actions and agentic-AI indicators, using expert-prioritized mitigations and public-document evidence. The resulting framework contains 16 indicators and is applied through a structured Microsoft document assessment.
- Methods: The study conducts a two-step process: designing a company-level SCAF for agentic-AI resilience contributions, then applying it in an illustrative case study.The design remains faithful to the original SCAF where possible while shifting from societal snapshots to dynamic company contributions.
- Resilience capacities: Company-level analysis translates country-oriented resilience descriptions into general company actions contributing to vulnerability, coping, and adaptive capacities.Agentic AI was selected because its risks fall under large companies’ purview and have consequential societal effects.
- Agentic AI threat model: The threat model treats agentic AI risk as relational and spanning systems with varying degrees of access and autonomy, including risks from automation, security challenges, and broader societal impacts.The paper identifies labor-bottleneck removal, multi-agent unpredictability, tool exploitation, and prompt injection as relevant concerns, while describing the threat list as illustrative rather than exhaustive.
- Indicator design: The indicator suite uses multiple risk mitigations as empirical proxies for resilience capacities and selects high-priority actions from the 2026 CLTC Agentic AI Risk Management Profile.The selection rule addresses the effectively unbounded mitigation landscape and the report’s large-company, agentic-AI scope.
- Indicator design: Researchers searched governance documents from 201 companies, manually reviewed six companies with the most matches, and used that evidence to design publicly assessable indicators.This broader review informed which practices could in principle be measured beyond the eventual Microsoft case.
- Evidence collection: The process produced a company-level framework with 16 indicators, followed by a Microsoft review using 306 keywords and 142 included excerpts.The most frequently matched indicators were Autonomy Scope (23), Behavioral Logging (19), Error Remediation (15), and Risk Thresholds (25).
4 Results & Discussion
The company-level SCAF prototype identifies Microsoft’s disclosed resilience profile and connects missing mitigations to corporate governance gaps. Results suggest adaptive-capacity dominance, stronger coverage where resilience aligns with product strategy, and weaker coverage where commitments conflict with commercial interests.
- Framework application: The company-level SCAF extended the framework from countries to companies and demonstrated structured measurement of organizational activities using agentic AI as the risk domain.The assessment treats Microsoft as an illustrative case rather than a sector-wide comparison.
- Microsoft profile: Microsoft’s SCAF profile is adaptive-capacity-dominant, consistent with organizational risk management’s inward-facing incentives.Socially beneficial incentives are less obvious for vulnerability and transformative capacities.
- Microsoft profile: Microsoft’s documents describe escalating autonomy toward “digital workers,” broad diffusion ambitions, and sensitive-risk notification rather than prohibition in many critical scenarios.Recommended-use guidance shifts responsibility for safe deployment downstream to customers.
- Microsoft profile: Indicator coverage was strongest where resilience-building aligned with product strategy, including identity binding, capability evaluation, defensive security products, and research.Research addressed reversibility, multi-agent failure, and misalignment, but governance commitment lagged behind this engagement.
- Governance gaps: Coverage was weaker for mitigations that could conflict with commercial interests, including redlines, emergency shutdown obligations, remediation, and accountability-inviting incident disclosure.The framework connects these missing mitigations to a structured gap analysis for improving societal outcomes.
- Limitations: The assessment cannot generalize from one company, verify Microsoft’s implementation, or treat absent public disclosure as proof that a mitigation is absent in fact.Its indicator design also inherits assumptions about CLTC proxies, transformative capacity, and the specificity–sensitivity trade-off in coding.
- Implications and future work: SCAF can support deployers’ risk management and help policymakers align company-level practices with societal goals through incentives or future compliance requirements.Future work includes inter-rater reliability, comparative profiles, incident-based construct validation, and policy interventions targeting undercovered indicators.
5 Conclusion
Companies that develop and deploy agentic AI shape the resilience of the societies into which their products diffuse, but existing approaches focused on internal risk management and model capabilities miss this relationship. The paper extends SCAF to companies and uses Microsoft as an illustrative case to identify corporate risk-response gaps linked to societal resilience.
- Companies that develop and deploy agentic AI shape the resilience of the societies into which their products diffuse.
- The company-level SCAF framework provides a structured method for assessing whether deployers build societal resilience as agentic systems take consequential actions.
A Appendix A
The appendix presents the prototype indicators for large companies and applies them to Microsoft as an illustrative case. Microsoft was selected because its extensive public governance materials enable a more informative demonstration of the methodology.
- Table 3 presents prototype indicators for large companies focused on agentic AI risks, while Table 4 applies them to Microsoft.
- Microsoft was selected for its extensive public-facing AI governance materials and its prominent promotion of agent adoption by “Frontier Firms.”Its product reach and pro-automation stance make it an illustrative case for the assessment methodology.
A.3 Microsoft Case: Findings
The Microsoft case measures disclosed posture and stated commitments rather than an exhaustive review or verified implementation. The assessment therefore provides an illustrative account of public evidence, not proof that the described measures were implemented.
- The Microsoft SCAF profile is an illustrative assessment of disclosed posture and stated commitments, not an exhaustive review or measure of implementation.
A.3.1 Vulnerability
Microsoft’s documents indicate broad ambitions for diffusing agentic AI, while safeguards leave important deployment boundaries and default human oversight unresolved.
- A.3.1 Vulnerability: Microsoft documents broad ambitions for widespread diffusion of agentic AI, with limited concrete information about safety-critical deployments.High-stakes examples, including medicine, appear only as isolated case studies.
- A.3.1 Vulnerability: Digital-worker framing implies that human-in-the-loop operation is a temporary stage rather than a lasting deployment condition.The documents describe human-in-the-loop operation as a starting point to move beyond.
- A.3.1 Vulnerability: Sensitive-use governance generally requires notification and additional oversight rather than prohibiting critical-risk deployments.Recommended-use guidance also shifts responsibility downstream to customers instead of making products safe by default.
A.3.2 Coping Capacity
Microsoft’s documents show strong behavioral logging but limited evidence of rapid intervention, emergency shutdown, mandatory human intervention, or remediation for harmed parties.
- A.3.2 Coping Capacity: Strong evidence of behavioral logging supports review of agent activity and outputs.The documents describe logging, but not rapid intervention.
- A.3.2 Coping Capacity: The documents describe no hardware-enabled kill switches or emergency-shutdown commitments for agent behavior.They do describe AI-enabled security products, but those do not constitute documented shutdown commitments.
- A.3.2 Coping Capacity: Tiered oversight exists, but the documents do not clearly specify which risk categories require human intervention.Their dominant framing automates routine action and reserves humans for ambiguous cases.
- A.3.2 Coping Capacity: Remediation language is not a commitment to compensate or redress parties harmed by AI agents.The term is generally used in its cybersecurity sense rather than for restoring affected parties.
A.3.3 Adaptive Capacity
Microsoft’s documents strongly cover identity binding and capability evaluation, but provide thinner incident disclosure and retain internal review and mitigation rather than firm red lines for critical capabilities.
- A.3.3 Adaptive Capacity: Identity binding and capability evaluation are very well covered in Microsoft’s documents.These controls address agent identity and the assessment of capabilities before or during deployment.
- A.3.3 Adaptive Capacity: Incident disclosure is thin and inconsistent, and the definition of “incident-level events” remains unclear.The headline claim about no 2024 AI-malfunction events therefore raises a definitional question.
- A.3.3 Adaptive Capacity: Critical-risk capabilities, including full automation of AI research and development, receive further review and mitigations rather than deployment prohibition.This leaves no stated red line for capabilities whose harms could be unrecoverable.
- A.3.3 Adaptive Capacity: Independent review is valued but remains internal through AIRT, while announced CAISI and AISI agreements represent clearer movement toward external scrutiny.Customer governance is well covered, though its status as transformation rather than adaptation remains unresolved.
A.3.4 Conclusions
The Microsoft case illustrates how the adapted SCAF organizes company-level evidence about agentic-AI governance and identifies commitments that are present, hedged, or missing.
- A.3.4 Conclusions: The adapted SCAF connects social-ecological resilience literature with downstream corporate AI governance to assess companies.The methodology is presented as a potential comparative tool for structuring corporate commitments.
- A.3.4 Conclusions: Evidence density per indicator could serve as a proxy for where governance aligns with product lines or imposes costs on the company.This supports using the framework to compare documented commitments and omissions.
- A.3.4 Conclusions: Table 4 applies the prototype company-level indicators to Microsoft as an illustrative case study.The prototype is limited to large companies and agentic-AI risks.
- A.3.4 Conclusions: Microsoft documents include human-in-the-loop approval, risk-based authority, and confidence thresholds for agent actions.These examples show documented governance mechanisms across products and enterprise guidance.
- A.3.4 Conclusions: Microsoft reports no 2024 incident-level events caused by AI malfunction or benign use and documents vulnerability disclosure and bug-bounty participation.The reported incidents involved malicious attempts to bypass security measures or misuse AI products.
- A.3.4 Conclusions: High- and critical-risk models require internal notification, while sensitive uses require notification to the Office of Responsible AI and additional governance steps.The framework tracks capabilities such as advanced autonomy and CBRN risks through defined thresholds.
- A.3.4 Conclusions: Microsoft’s evaluation practices include agentic-AI capability tests, external experts, automated adversarial simulation, and independent or external red teams.The documents also describe agreements with CAISI and AISI to advance AI-evaluation science.
- A.3.4 Conclusions: The corpus documents containment, runtime checks, identity controls, audit logging, and recoverability guidance for deployed agents.Examples include Agent ID, webhook-based blocking, audit histories, and a unified control plane.
B.1 Agentic AI Keywords
This section organizes agentic-AI resilience indicators around deployment characteristics and risk-management actions, with associated keywords for document analysis.
- Indicators mapped to risk-management actions: The framework maps agentic-AI resilience indicators to actions from the 2026 Agentic AI Risk Management Profile.The mapped actions include autonomy-scope definitions, acceptable-use policies, rapid intervention, behavioral logging, and tiered oversight.
- Deployment characteristics: Deployment reach and autonomy scope characterize where agents operate and the authority they receive.Deployment-reach keywords include critical infrastructure and enterprise deployment, while autonomy keywords include delegated decision-making, privileged actions, tool use, and long-horizon operation.
- Deployment characteristics: Use boundaries and anthropomorphic framing identify policy constraints and how companies describe agents’ roles.Relevant terms include acceptable-use policies, prohibited uses, operational constraints, AI workers, and digital employees.
- Risk-management actions: Coping and control indicators cover rapid intervention, behavioral logging, tiered oversight, error remediation, identity binding, and incident disclosure.These actions include disabling agents, monitoring behavior, escalating high-risk actions for review, repairing erroneous interactions, authenticating communications, and reporting incidents.