Source-linked AI summary
Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems
Paul Vautravers, Oliver Chalkley, Gabriel Downer, Kate S, Damian Ruck
TL;DR
AI adoption in complex sociotechnical systems exposes a gap between model-level evaluation and system-level risk assessment. The paper addresses this gap by integrating structured hazard analysis, component testing, and probabilistic system modelling, and its RTGS worked example shows how local AI behaviour can be traced to quantified system outcomes.
Problem
Policymakers and regulators lack a clear way to anticipate system-level harms from bottom-up AI adoption when they have limited visibility and control over individual components.
Method
The framework combines STPA-based hazard analysis, empirical component testing, and probabilistic system modelling to connect hypothesised AI harms with system consequences.
Results
The framework provides a traceable pathway from hypothesised harms to quantified system outcomes, illustrated by adversarially induced shifts in asset allocation propagated through a stylised financial network model.
Takeaways & Limitations
The approach supports anticipatory, evidence-gated oversight by helping practitioners assess plausible harm pathways before widespread deployment.
Takeaways & Limitations
The framework depends on analytical judgement and simplifying assumptions when mapping component-level qualities to system-level impacts, especially for harms emerging at scale or through interactions.
Abstract
from arXiv · showhide
Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI evaluation remains model-centric, offering little insight into how observed behaviours might translate into system-level risk. We propose a framework that links structured hazard analysis, component-level testing, and probabilistic system modelling to bridge this gap. By providing a traceable pathway from model behaviour to system-level outcomes, the framework enables practitioners to answer the "so what?" of AI failures, quantify their systemic impact, and move toward evidence-based and anticipatory governance of AI in complex systems. Applied to the UK's Real Time Gross Settlement (RTGS) system as an illustrative worked example, we derive AI-driven loss scenarios using Systems Theoretic Process Analysis (STPA) and examine adversarial manipulation of LLM-based trading as one such loss scenario. Component-level experiments show that simple adversarial inputs induce measurable behavioural shifts where AI recommendations are followed. Under the component-to-system mapping used here for a financial contagion model, these shifts alter system resilience, increasing bank failures and lowering the threshold at which shocks lead to cascading disruption, particularly under widespread or monopolistic AI adoption.
Introduction
AI adoption in complex sociotechnical systems creates system-level risks that model-centric evaluation does not adequately connect to harms. The paper proposes a proportionate framework linking hazard analysis, component testing, and system modelling, illustrated through the UK RTGS system.
- Complex sociotechnical systems combine technical, human, organisational, and legacy elements whose interactions can produce nonlinear outcomes.
- Rapid AI change, incentives for widespread adoption, and limited institutional expertise leave policymakers reasoning about system-level harms under uncertainty.
- Existing AI evaluations often focus on models in isolation, leaving adopters to bridge separate AI-specific and system-safety assessment practices.
- The proposed framework uses structured hazard analysis, component testing, and simulation-based modelling to identify harm pathways, measure concerning behaviours, and quantify system consequences.
- In the RTGS worked example, STPA identified unsafe behaviours and AI-enabled loss scenarios, while prompt-injection testing and contagion modelling connected component qualities to system-level harm.
- The RTGS application is an illustrative, high-level worked example rather than a high-fidelity analysis.
Background
The paper situates its framework within safety engineering for complex systems, where interactions between functioning components can generate harms. It combines proportionate modular assessment with STPA to capture emergent pathways that component-failure methods may miss.
- The UK RTGS settles over £700 billion on an average working day and operates within a complex, adaptive financial system.
- Effective system-level AI risk evaluation should draw on established safety-critical methods rather than treating AI assessment as an isolated discipline.
- The framework is modular, allowing practitioners to stop at stages A, B, or C according to findings, capabilities, and risk tolerance.
- FMEA analyses harms arising from component failures but can miss unsafe outcomes produced by interactions between correctly functioning components.
- STPA treats safety as an emergent property and is suited to identifying interaction-driven pathways to harm in complex sociotechnical systems.
Methodology
The framework traces AI-enabled harm pathways through structured hazard analysis, tests component behaviours, and projects measured shifts into system-level simulation. In the RTGS worked example, STPA guided selection of an adversarial trading scenario for financial-contagion modelling.
- Structured hazard analysis: STPA identified stakeholders, losses, hazards, constraints, control structures, unsafe control actions, and causal loss scenarios.The process supports a qualitative taxonomy of AI-enabled risk pathways for early prioritisation before component testing.
- Structured hazard analysis: The analysis applied STPA to the existing RTGS control structure before mapping candidate AI use-case classes onto control actions, feedback, and decision-making.This avoids repeating the full STPA process separately for every possible deployed AI system.
- Scenario selection: LS-18.2.2-A tested highly delegated GenAI trading operations as a technically feasible and organisationally plausible upper-bound adoption scenario.The scenario was selected for stress testing, not as a forecast of typical deployment.
- Component testing: Component testing used a lightweight LLM decision tool that consumed controlled market-news articles and allocated portfolios across equities, mortgage-backed securities, corporate bonds, and government bonds.Each episode produced asset weights between 1 and 0 that summed to 1 across the four asset types.
- Component testing: Adversarial attacks targeted already distressed assets by manipulating LLM preferences rather than bypassing safety guardrails.The study focused on low-cost, accessible attacks intended to amplify fireselling under stress.
Results
STPA produced a broad set of RTGS hazards and AI-driven loss scenarios, while component experiments showed measurable allocation shifts from simple adversarial attacks. Mapping these behaviours into the contagion model associated stronger AI adoption and adversarial manipulation with higher bank failure rates and reduced resilience.
- Hazard analysis: Four high-level losses were linked to six hazard states, with over 100 AI-agnostic unsafe control actions and eight AI-driven loss scenarios identified.The scenarios were developed within the RTGS system boundary and its stakeholder outcomes.
- Scenario selection: One representative scenario examined adversarial manipulation of LLM-based portfolio advice and its potential to amplify bank failures during stress.The qualitative scenario connected causal factors producing unsafe control actions to component-level testing.
- Component testing: Component tests across three foundation models measured asset-allocation shifts under simple adversarial attacks, averaged across multiple attack instances.The experiments also exposed model-specific non-adversarial preferences, including aversion to mortgage-backed securities in Neutral settings.
- Component testing: Simple adversarial attacks caused small absolute but large relative allocation shifts, making distressed assets appear more distressed in some single-shot trading responses.Attack effectiveness varied across individual instances.
- Mitigation scope: The unmitigated analysis projected component failures into system-level risk, while secure design, input filtering, human review, and prompt hardening were identified as mitigation directions.The prompt-hardened setting was reported separately in the appendix.
- System-level outcomes: Increasing firesale intensity increased maximum bank failure rates, with adversarial settings and GPT-5-mini monopoly adoption producing the greatest rates in the illustrated mapping.The relationship became self-limiting through plateauing bank-failure behaviour and persisted across mapping-parameter choices.
- System-level outcomes: Higher firesale intensity lowered the mortgage-shock threshold required to reach a given bank-failure rate, indicating reduced resilience under the modelled conditions.This provides a quantitative bridge between environmental shocks and the loss of wider banking-sector functioning.
Discussion
The framework connects hazard analysis, component testing, and quantitative modelling to explain how AI behaviours can contribute to system-level harms. In the RTGS illustration, it also supports anticipatory, prioritised, and accountable oversight, while relying on explicit judgements and assumptions.
- Model-level failures matter only in relation to the system where outputs are used and the losses they may contribute to.
- Adversarial asset-allocation shifts increased fire-sale intensity, raised stressed bank failure rates, and lowered the external-shock threshold for widespread failure under the mapping assumptions.
- Anticipatory policymaking: The framework supports anticipatory policymaking by starting from system-level losses and assessing whether hypothesised harm pathways are material under particular assumptions.
- Prioritisation at adoption: Focusing on AI adoption and use cases makes testing and mitigation more actionable around input–output behaviours that directly drive stakeholder harm.
- Monitoring and accountability: Linking losses, scenarios, behaviours, and system outcomes supports targeted reporting and monitoring of unsafe or higher-risk operating conditions.
- Limitations and future work: The framework inherits variation from qualitative assumptions and normative judgements, while component-to-system mappings require simplifying analytical choices in complex sociotechnical settings.
- Limitations and future work: The RTGS analysis is conceptual and public-data-based, with stakeholder engagement needed to validate assumptions and test applicability in other CNI sectors.
Conclusions
The paper bridges AI component evaluation and system-level risk by integrating hazard analysis, empirical testing, and probabilistic modelling. Its RTGS worked example shows how adversarial asset-allocation shifts can be traced to quantified financial-network consequences, supporting anticipatory and actionable oversight.
- The framework integrates structured hazard analysis, component testing, and probabilistic system modelling into a traceable account of systemic AI risk.
- In the RTGS example, STPA, empirical testing, and simulation-based modelling connect small adversarial asset-allocation shifts to quantified financial-network outcomes.
- The staged framework supports anticipatory governance, prioritised oversight at AI adoption, and evidence-gated risk assessment.
- Qualitative judgements and interdisciplinary expertise remain necessary, but the framework makes these dependencies explicit and auditable.
A STPA Analysis Details
The STPA analysis defines RTGS losses, hazards, control structures, unsafe control actions, and AI-driven loss scenarios. It also narrows the analysis to selected financial-system use cases and excludes broader or finer-grained impacts.
- The appendix documents STPA artefacts for AI integration in or adjacent to the UK RTGS system, with full tables supplied separately.
- The losses cover high-value payment disruption, settlement incorrectness, degraded bank functioning, and loss of confidence in Bank of England policy and infrastructure.
- Information disclosure is treated as a hazard leading to loss of confidence, rather than as a high-level loss itself.
- Loss of socioeconomic welfare was excluded because tracing and modelling it would require substantial real-economy or national-governance detail.
- The hierarchical control structure captures system relations through control and frames analysis around preventing unacceptable outcomes and losses.
- Control actions and unsafe control actions were used to construct eight AI-driven loss scenarios, selected by overlap with financial AI use cases and systemic risks.
- The analysis did not cover all AI integrations, excluding indirect economy-wide effects and retail-level fraud, AML, and counter-terrorist-financing use cases.
A.7 Information Gathering and Effective Hazard Analysis
The analysis uses publicly available information and STPA-derived loss scenarios to structure AI-related hazard analysis, while acknowledging that RTGS safeguards and operational details cannot be fully assessed publicly.
- Hazard analysis: The RTGS analysis identifies broad unsafe control actions and uses them to construct AI-driven loss scenarios for subsequent testing.The scenarios are intended to structure analysis rather than determine which existing safeguards already mitigate particular actions.
- Scope and information: The analysis is preliminary because detailed RTGS controls and a complete HCS or STPA analysis cannot be openly disclosed due to infrastructure sensitivity.Future work is intended to engage RTGS operators, supervisors, and participating institutions to relate unsafe actions to existing controls and adoption pathways.
- Loss scenarios: Loss Scenario 18.2.2-A examines an LLM-based trading tool whose raw media input contains a strong prompt injection.The scenario concerns an investment manager generating near-real-time trade strategies from aggregated news and media.
- Loss scenarios: Prompt injection, overreliance, and excessive agency are identified as factors that can allow a flawed AI-generated strategy to influence trading.The described secure-design concern is that raw form text can permit hijacked control flow, particularly where unintended tool use is available.
- Consequences: A flawed trade could impair an organisation’s operations and, under a targeted stress-oriented attack, contribute to systemic liquidity shortages.This is the stated consequence of the trading loss scenario, not a general claim about all AI-assisted trading.
- Loss scenarios: A separate monitoring scenario describes an ML classifier failing to detect RTGS stress, preventing migration to an established contingency and degrading uptime and payment integrity.The scenario attributes the failure to an algorithm design or implementation flaw combined with staff over-reliance and limited technical understanding.
B.1 Component Testing Experimental Detail
The component-testing design evaluates how LLM-enabled trading decisions respond to article ordering, simple adversarial inputs, and prompt-hardening across three foundation models, with mitigation results limited to GPT-5-mini.
- Experimental design: Each episode represents one LLM-enabled trading decision and contains exactly one article per asset class.With four assets, each episode contains four articles sampled from a predefined asset-by-sentiment pool.
- Experimental design: Article ordering defines distinct episodes because presentation order may affect the recorded decision process.Permutations of the same articles are therefore treated as separate experimental inputs.
- Experimental design: Experimental runs are generated deterministically from a random seed with model origin fixed per experiment unless otherwise specified.Run objects contain the seed, model origin, generated episodes, and condition-specific generation procedures.
- Adversarial testing: The attacks apply simple preference-manipulation inputs to an illustrative LLM-based trading function.The attack inputs are inspired by adversarial search-engine-optimisation attacks described in prior literature.
- Mitigation: Prompt-hardening inserts a safety directive into the system and article-processing prompts to prevent direct instructions from steering asset decisions.The directive specifically addresses asset dropping, escalatory or overly polite language, and direct reallocation instructions.
- Results presentation: Unmitigated figures report attack-vector results across three foundation models, whereas mitigated results were collected only for OpenAI GPT-5-mini.The mitigated figures cover average responses and the five named attack types, including escalation, model address, alignment association, appeal to authority, and system instruction attacks.
C Complex Model Specification
The contagion model represents banks, portfolios, interbank exposures, correlated or uncorrelated asset shocks, counterparty losses, and integrated fire-sale effects in a discrete-round simulation.
- Model structure: The model extends Eisenberg–Noe clearing with portfolio holdings, correlated shocks, fire-sale contagion, and priority of claims.It combines network-contagion simulation with a stylised interbank structure.
- Balance sheets: Each bank has external assets, interbank assets, interbank liabilities, and external liabilities, with solvency defined by positive equity.External assets include tradable securities, while interbank assets and liabilities represent claims and obligations between banks.
- Network: The interbank network is a directed graph with core and periphery banks and tier-dependent edge probabilities and exposure distributions.The default parameterisation uses 50 banks, including 10 core and 40 periphery banks.
- Shock application: Asset shocks may be deterministic, correlated stochastic, or uncorrelated stochastic, and banks with non-positive equity after initial shocks are marked failed.Correlated shocks use a common market factor plus asset-specific components, with fresh realisations drawn for each Monte Carlo run.
- Contagion: After initial shocks, contagion propagates through counterparty credit losses in discrete rounds until no new failures occur.Interbank creditors receive residual recoveries after external creditors under the default priority-of-claims setting.
- Fire sales: Fire sales are integrated into each contagion round, producing mark-to-market losses for surviving banks and potentially creating further failures.Price impacts are normalised by initial market size using a non-compounding linear formulation.
E.3 Experiment Configurations
The experiments are summarised in a configuration table, with the detailed experimental conditions presented as a consolidated set of configurations.
- Experiment configurations: Table 5 summarises the experiments performed.The supplied passage identifies the table as the summary location but does not enumerate its configurations.
- Experiment configurations: Table 5 is presented as a summary of experimental configurations.The passage provides the table title without further methodological or quantitative detail.
F Complex Model Sensitivity Analysis
The analysis maps component-level AI behaviour to fire-sale intensity and examines how increasing intensity changes systemic stability. Parameter sweeps show steeper, less predictable, more amplified, and more compressed failure responses, while adversarial amplification depends on the mapping and adoption scenario.
- Fire-sale sensitivity: Increasing λ raises the maximum single-step increase in mean failure rate from 12–14 pp to 16–17 pp.The resulting steeper response reduces the policy-intervention window as mortgage shocks increase.
- Fire-sale sensitivity: Failure-outcome standard deviation increases from 26–35% to 42–44% as λ rises, making identical macroeconomic conditions produce more variable outcomes.The passage describes this high variance as approaching a bimodal limit and undermining stress-test forecast reliability.
- Fire-sale sensitivity: Mean fire-sale losses rise from approximately 5% of total losses at baseline to 15–18% at high λ.This measures amplification of the indirect fire-sale channel relative to direct interbank contagion.
- Adversarial mapping: Adversarial fire-sale values produce appreciably higher bank failure rates across a large range of mapping parameters, with effects varying by adoption scenario.Amplification is averaged across adoption scenarios excluding 0% AI adoption; monopolistic adoption can show larger differences, while excessively large mapped fire-sale values can leave scenarios outside the model’s sensitivity range.
H.3 Additional System-Level Results
Additional simulations examine bank-failure responses and shock thresholds across three component-to-system mappings. The examples also illustrate how fire-sale intensity could be related to operating zones, while explicitly limiting that use to illustration.
- K = 1.0: For K = 1.0 and fH = 0.05, simulations show bank-failure rates across mortgage shocks and minimum shocks for specified failure rates across AI-driven fire-sale values.These results correspond to Figures 28 and 29.
- K = 1.5: For K = 1.5 and fH = 0.033, simulations provide the same bank-failure and shock-threshold analyses across AI-driven fire-sale values.These results correspond to Figures 30 and 31.
- K = 2.0: For K = 2.0 and fH = 0.025, simulations again examine bank-failure rates and minimum mortgage shocks across AI-driven fire-sale values.The accompanying operating-zone plot is presented purely for illustration, using unchanged thresholds as desirable, the worst measured value as dangerous, and intermediate values as cautionary.
I Code and STPA Artefact Availability
The project makes its simulation code and STPA-related artefacts available through a repository, alongside figures illustrating system-level results. The paper also bounds interpretation through explicit methodological, contextual, and dual-use caveats.
- Artefact availability: The repository contains STPA tables, component-level testing code, financial-contagion simulation code, and instructions for running simulations.The artefacts could not reasonably be included in the extended paper and are organised in repository subdirectories.
- Scope and review: The RTGS worked example uses a level of abstraction intended to avoid operationally sensitive disclosure and was reviewed by domain and safety experts.The analysis is described as a stylised financial setting focused on systemic consequences rather than advancing attack capability.
- Limitations: The framework’s conclusions depend on input quality, analytical expertise, and model realism, so partial analyses should inform judgement rather than provide standalone assurance.The authors warn that treating the framework as a comprehensive safety assessment or compliance checklist can create false assurance and leave residual risks unmodelled.