Source-linked AI summary

Stress-testing university AI governance: A prospective method for locating policy breakpoints

Biranchi Poudyal

arXiv:2608.28925v1cs.CY

TL;DR

Universities often publish AI principles without specifying accountable decision pathways for unfamiliar forms of AI agency. This paper develops IAGST, a prospective documentary stress-testing method that combines frozen governance corpora, controlled capability escalation, six response-chain dimensions, non-compensatory rules, and breakpoint diagnosis. Across 75 university-scenario cases, governed pathways declined from augmentation to delegation and disappeared at autonomous substitution, while the method remained a diagnostic of documented preparedness rather than implementation.

  • Problem

    Universities’ AI policies may state defensible principles while leaving authority, decision criteria, procedures, safeguards, review, and remedy unspecified for unfamiliar AI cases.

  • Method

    IAGST tests a frozen public governance corpus through controlled capability escalation using six response-chain dimensions, non-compensatory decision rules, and university-scenario breakpoint diagnosis.

  • Results

    Across 75 cases, six were resolved, 14 were resolved through structured discretion, and 55 were indeterminate; governed pathways fell from 16 of 25 augmentation cases to four delegation cases and none at autonomous substitution.

  • Takeaways & Limitations

    IAGST shows how universities can test whether authority, procedures, safeguards, and review remain connected as AI capabilities evolve, beyond inventories of policies and principles.

  • Takeaways & Limitations

    IAGST measures public documentary preparedness at one point in time, not implementation, informal coordination, culture, resources, or live-case behaviour, and its Western Australian case set is not nationally representative.

Abstract

from arXiv · show

Universities are producing AI principles and use policies faster than they are building decision pathways for unfamiliar forms of AI agency. This study develops Institutional AI Governance Stress Testing (IAGST), a prospective documentary method for locating where publicly documented governance ceases to yield an accountable response. IAGST adapts established policy stress-testing and wind-tunneling logic. Its originality lies in combining controlled capability escalation, a frozen documentary corpus, a six-dimensional governance response chain, non-compensatory decision rules, and case-level breakpoint diagnosis. The method was demonstrated using 133 substantive public documents from five Western Australian universities and 15 quality-screened scenarios, resulting in 75 university-scenario encounters. Six cases were resolved, 14 were resolved through structured discretion, and 55 were indeterminate. Governed pathways fell from 16 of 25 augmentation cases to four delegation cases and none at autonomous substitution. The dominant weakness was not the complete absence of responsible roles: all 50 authority-gap cases named a role at only a generic level but lacked sufficient decision criteria or process. The findings show how universities can move beyond policy inventories and principal statements by testing whether authority, procedures, safeguards, and reviews remain connected as AI capabilities evolve. IAGST is a reproducible diagnostic for policy learning, not a ranking or measure of implementation.

Introduction

Universities have expanded AI policies, but documented principles often do not specify who decides unfamiliar cases, which criteria apply, what process affected people can use, or where review and remedy sit. IAGST adapts policy stress-testing to locate these governance breakpoints across escalating forms of AI agency.

  • Problem: Policy presence is a weak proxy for managerial preparedness because documents may omit decision authority, criteria, procedures, review, or remedy.These omissions become more consequential as AI moves from bounded assistance toward delegation, representation, coordination, and cross-system action.
  • Methodological gap: Conventional policy stress-testing and anticipatory governance provide relevant foundations but are difficult to reproduce, compare, or audit when based mainly on deliberative workshops.IAGST retains the prospective challenge while making the documentary evidence trail and classification rules inspectable.
  • IAGST: IAGST tests a frozen public governance corpus through controlled escalation from augmentation to delegation and autonomous substitution.The unit of analysis is a university-scenario encounter, and the object under test jointly distributes authority, procedure, rights, safeguards, and revision responsibilities.
  • Contribution: The method combines a frozen multi-document corpus, evidence-grounded scenarios, a six-part governance response chain, non-compensatory rules, and university-scenario breakpoint diagnosis.Its originality is methodological and configurational rather than the invention of scenario planning, stress testing, document analysis, or responsible-AI principles.
  • Scope: IAGST tests documented response capacity rather than implementation, organisational culture, staff competence, policy compliance, or actual AI-system safety.The empirical demonstration uses five Western Australian universities and produces no institutional ranking or composite readiness score.

Conceptual framework

The framework treats preparedness as a relational, reviewable connection among governance provisions rather than as isolated policy coverage. It uses six diagnostic links and controlled capability escalation to test whether accountable institutional action remains possible as AI agency expands.

  • Documentary preparedness: Public documents are consequential but incomplete representations of governance because they do not disclose informal coordination, resources, implementation quality, or live-case behaviour.They still establish categories, formal responsibilities, legitimate procedures, and grounds for challenge, making the public corpus an accessible interface of institutional authority.
  • Documentary preparedness: Preparedness is relational: individually adequate privacy, assessment, appeals, and security provisions may fail collectively to assign ownership of a cross-system AI agent or a route to contest its action.The analysis therefore examines relationships among provisions rather than document-by-document coverage.
  • Governance response chain: IAGST models six links: scope recognition, authority allocation, normative coherence, procedural actionability, safeguard coverage, and adaptive capacity.Together, they describe documentary conditions for an institutionally usable and reviewable answer, not a causal model or interchangeable indicators.
  • Governance response chain: Structured discretion requires a named authority, stated criteria, a reachable procedure, and a review or appeal route.Generic instructions to consult a lecturer, committee, or service unit do not satisfy this standard when the decision basis and next steps are unspecified.
  • Decision logic: IAGST uses non-compensatory thresholds because strong normative coherence cannot make a case actionable when authority or review is absent.The resulting classification is deliberately conservative and preserves institutional accountability as a necessary condition.
  • Capability escalation: Scenarios escalate within recognisable activities from augmentation to delegation and autonomous substitution rather than forecasting what will occur.The sequence varies human control, identity, persistence, or system reach to test capability-governance fit.
  • Capability escalation: A breakpoint is the first level with a strict documentary failure, while its failure signature records the reasons for breakdown.This distinguishes IAGST from policy inventories, maturity models, and deliberative wind-tunnelling.
  • Capability escalation: Existing university AI governance is better suited to human-authored content questions than to systems that represent people, coordinate groups, generate research evidence, or operate across institutional systems.This inherited focus motivates testing governance under expanded AI agency.

Method

The study freezes public governance documents and tests them against 15 quality-screened scenarios across five Western Australian universities. Cases are scored across six dimensions and classified with fixed, conservative rules whose auditability depends on explicit evidence and visible researcher judgment.

  • Design and corpus: The multiple-case design covers four public and one private Western Australian university, with five universities and 15 scenarios producing 75 university-scenario cases.The bounded set demonstrates the method within one state context rather than estimating national prevalence or ranking providers.
  • Design and corpus: Institutional websites were systematically searched for current public materials covering AI, integrity, assessment, privacy, conduct, appeals, teaching, delegation, procurement, risk, and governance.The search targeted the wider governance corpus rather than a single AI policy.
  • Design and corpus: The register contained 133 substantive documents and two absence or access-limitation records.Substantive totals were 38 for Curtin, 17 for Edith Cowan, 28 for Murdoch, 30 for Notre Dame Australia, and 20 for the University of Western Australia.
  • Design and corpus: Evidence precedence ran from binding instruments through institution-wide and unit guidance to other official webpages, with lower-authority sources unable to override binding provisions.Each entry retained metadata, official URLs, dates where stated, access information, and a frozen local copy.
  • Scenarios: Fifteen scenarios formed five capability families, each containing augmentation, delegation, and autonomous-substitution cases.The sequence held the activity recognisable while varying AI control, authority, identity, persistence, or system reach.
  • Scenarios: A 27-source capability-evidence matrix supported scenario cards that fixed the actor, capability, action, oversight, identity, data, domain, risks, boundaries, and diagnostic question.Cards were internally screened for evidence, plausibility, relevance, distinctiveness, governance focus, escalation continuity, disciplined futurity, and related criteria.
  • Scoring and classification: Each governance dimension was scored independently from 0 to 3 using explicit anchors, with a score of 3 requiring scenario-specific operational evidence and a score of 1 generally indicating generic or analogical applicability.Scores were not summed into a readiness index.
  • Scoring and classification: A case was indeterminate when scope was absent, authority absent or merely generic, procedure unavailable, or a directly critical safeguard absent.Structured discretion additionally required a delegating mode, authority and procedure scores of at least 2, and an established review or appeal route.

Results

Across 75 university-scenario cases, governed pathways declined sharply as AI agency escalated, while failure analysis showed that nominally assigned authority often lacked actionable criteria, procedures, or review power. Results also varied by capability family, revealing distinct institutional reform problems rather than one uniform deficit.

  • Overall outcomes: 6 of 75 cases (8.0%) were resolved, 14 (18.7%) through structured discretion, and 55 (73.3%) were indeterminate.No case met the strict definition of documentary conflict.
  • Escalation pattern: Governed pathways fell from 16 of 25 augmentation cases to four delegation cases and none at autonomous substitution.At delegation, none of the four resolved cases used structured discretion; all 25 autonomous substitution cases were indeterminate.
  • Six-dimensional profiles: Mean scores declined from augmentation to autonomous substitution across scope recognition, authority allocation, normative coherence, procedural actionability, safeguard coverage, and adaptive capacity.Scope recognition fell from 2.56 to 0.40, while safeguard coverage fell from 2.52 to 0.24.
  • Capability families: Assessment autonomy was the only capability family generally governed through delegation, with four of five Level 2 cases resolved.Representation and identity failed earliest: all 15 cases were indeterminate despite applicable privacy, attendance, and appeals provisions.
  • Capability families: All five course-level memory scenarios were resolved through structured discretion, whereas every cross-course profile and cross-system agent was indeterminate.Existing governance could allocate local responsibility but did not connect ownership across institutional domains.
  • Failure signatures: 50 cases triggered authority gaps, compared with 19 safeguard omissions, 15 scope failures, and 13 procedural dead ends.Every authority-gap case scored one for authority allocation, indicating that generic role assignment lacked documented decision criteria, procedure, or review authority.
  • Failure signatures: Five indeterminate near misses lacked a complete evidence threshold despite having no strict failure trigger, while procedural actionability had the lowest overall mean at 1.20.The 73.3% indeterminate result therefore comprised different configurations, including incomplete pathways requiring targeted completion.

Discussion

The discussion reframes preparedness as whether governance connections remain intact as AI capabilities change, rather than whether policies merely exist. IAGST identifies inspectable breakpoints and translates distinct failure signatures into targeted governance maintenance.

  • From policy presence to capability-governance fit: Almost three-quarters of cases were indeterminate despite strong normative coherence and no direct conflict.Ethical agreement did not reliably connect ownership, procedure, safeguards, and review.
  • From policy presence to capability-governance fit: AI delegates can disrupt identity, authority, representation, organisational boundaries, and accountability relationships assumed by assessment policies.The governance challenge changes when AI attends classes, represents students, evaluates group members, or combines data across courses.
  • Methodological contribution: IAGST combines a frozen governance corpus, controlled capability escalation, a six-part response chain, deterministic rules, and case-level breakpoint diagnosis.These adaptations test the whole public governance system rather than a favoured AI policy.
  • Methodological contribution: IAGST identifies the first plausible capability condition where public documents fail to yield a complete response without ranking universities or predicting adoption.Its separable corpus, scenarios, anchors, and algorithms support replication and falsifiability.
  • Implications for university policy and management: The 50 authority-gap cases named generic roles but lacked the criteria and procedures needed for accountable action.A governable allocation couples an authorised role, decision criteria, a reachable procedure, and a route for review.
  • Implications for university policy and management: Fourteen augmentation cases were governable through structured discretion, showing that flexibility can coexist with consistency and due process.Unstructured referrals instead risk transferring uncertainty to students or frontline staff.
  • Implications for university policy and management: Stress-testing should become a recurring maintenance cycle that records breakpoints, assigns remediation owners, and tests whether reforms move them.Repeated testing can reveal new procedural dead ends or ownership problems when capabilities and vendors change.
  • Implications for university policy and management: Governance performed best within assessment and research-integrity roles and worsened when AI crossed identity, participation, memory, data, or administrative boundaries.Future governance should address capability and accountability relationships alongside traditional functional silos.

Limitations

The study’s claims are bounded to public documentary preparedness in a specific case set and time, not implementation or live organisational behaviour. Interpretation, scenario design, thresholds, and capability coverage also require further validation.

  • Scope and evidence: IAGST measures public documentary preparedness at one point in time, not implementation, resources, informal coordination, culture, or live-case behaviour.Public evidence may also underestimate institutions whose operative procedures are internal.
  • Scope and evidence: The Western Australian case set demonstrates the method but does not represent Australian universities generally.The study’s geographic and institutional scope limits external generalisation.
  • Interpretation and validation: One researcher constructed, coded, and audited the cases, so interpretive independence and intercoder reliability are not established.This constrains confidence in coding consistency across researchers.
  • Interpretation and validation: Scenario quality depends on capability evidence and boundaries fixed at the freeze date, while the dimensions and conservative thresholds require external validation and sensitivity testing.Workable improvisation may exist where documents remain indeterminate.
  • Scope and evidence: The five capability families do not exhaust plausible AI futures, and autonomous-substitution scenarios are stress conditions rather than adoption forecasts.The scenarios should not be read as predictions of institutional uptake.

Conclusion

IAGST tests what happens when AI capabilities outgrow the use cases assumed by university documents. In the Western Australian demonstration, governed pathways weakened from augmentation to delegation and disappeared at autonomous substitution, with thin authority—not total ownerlessness—as the dominant weakness.

  • Conclusion: IAGST adapts policy stress-testing to frozen documents, controlled capability escalations, a governance response chain, non-compensatory rules, and breakpoint diagnosis.It provides a reproducible way to examine whether documented governance yields an accountable response.
  • Conclusion: Governed pathways were common during augmentation, scarce during delegation, and absent during autonomous substitution.The conclusion presents this progression as the central empirical pattern of the demonstration.
  • Conclusion: The dominant weakness was thin authority: roles could be named without the criteria and procedures needed for accountable action.This distinguishes the practical reform problem from policy conflict or complete ownerlessness.
  • Conclusion: Used iteratively, IAGST can help universities connect foresight to governance redesign before novel AI practices harden into unmanaged institutional dependencies.The method turns broad claims of unpreparedness into a practical reform agenda.
Loading 2608.28925v1…