Source-linked AI summary

Generative AI for trustworthy systems - Towards a health check model

Jan Bosch, Rick Kazman, Henry Muccini, Helena Holmström Olsson

arXiv:2609.10595v1cs.SE

TL;DR

The paper addresses the limits of unidimensional maturity models for representing how organizations establish trust in GenAI-assisted software engineering. It develops an inductively grounded health-check model from practitioner interviews, finding four trust paradigms and eight dimensions that support configurational comparison rather than a single progression path. The paper frames alignment between AI authority and supporting dimensions as a testable hypothesis while limiting the model to diagnosis rather than prediction.

  • Problem

    Existing trustworthy-AI, maturity-model, and governance literature does not provide an empirically grounded multidimensional structure for characterizing how organizations operationalize trustworthy autonomy across diverse industrial contexts.

  • Method

    The paper uses inductive analysis of interviews with practitioners and synthesizes empirically supported themes into a five-level health-check model of dimensions and trust paradigms.

  • Results

    The model identifies four trust paradigms and eight dimensions spanning system and organizational properties, supporting cross-organizational comparison without imposing a single progression path.

  • Takeaways & Limitations

    Trustworthiness is characterized through context-dependent configurations, with the same nominal dimensional position potentially healthy in one domain and pathological in another.

  • Takeaways & Limitations

    The model is diagnostic rather than predictive, and its alignment hypothesis remains a testable claim for future empirical work.

Abstract

from arXiv · show

The adoption of generative AI in software-intensive systems is proceeding rapidly, but current analytical instruments - principally unidimensional maturity models - compress important configurational variation into a single progressive axis. Drawing on an inductive interview study of eighteen senior practitioners across telecommunications, automotive, defence, aviation, banking, energy, government, and enterprise software services, this paper presents the Trustworthy Autonomy Health Check Model: a structured, multidimensional instrument for characterizing how organizations establish trust in GenAI-assisted software engineering. The model organizes eight empirically grounded dimensions into a system layer (Scope of Agent Authority, Assurance Mechanisms, Data Trustworthiness, Architectural Containment, Traceability & Comprehensibility) and an organizational layer (Governance, Human Oversight Posture, Workforce Capability Sustainability), each on a five-level ordinal scale. A cross-cutting overlay of four trust paradigms - operational, engineering, statistical, and containment-based - captures how trust is established, complementing the dimensions capturing what must be trustworthy. The model is diagnostic rather than prescriptive: it supports cross-organizational comparison, surfaces configurational trade-offs, and locates an organization in a shared space without imposing a single progression path. Higher levels are not inherently better; the goal is alignment across dimensions appropriate to the organization's domain and chosen trust paradigm. This is formalized as the alignment hypothesis: effective trustworthiness is constrained by the fit between the scope of agent authority and the dimensions enabling it, with misalignment producing pathological risk or unnecessary friction. We discuss how practitioners can apply the model and outline directions for empirical validation.

1 Introduction

Generative AI adoption creates diverse trustworthiness challenges that cannot be represented adequately on a single maturity axis. The paper responds with an inductively derived, multidimensional health-check model and an alignment hypothesis for comparing organizational configurations.

  • 1 Introduction: Organizations adopting GenAI face different combinations of authority, architecture, governance, assurance, and other trustworthiness choices rather than a single adoption challenge.The study frames trustworthy autonomy as a configurational problem across multiple dimensions.
  • 1 Introduction: A maturity model’s single progression axis conceals meaningful differences among organizations operating in distinct multidimensional configurations.The paper argues that unidimensional framing can conceal more than it reveals.
  • 1 Introduction: The study asks how organizations operationalize trustworthy autonomy and which dimensions and trust paradigms characterize variation across domains.The two research questions form an inductive arc from practitioners’ descriptions to a structured framework.
  • 1 Introduction: The Trustworthy Autonomy Health Check Model organizes eight empirically grounded dimensions into a configuration space for comparison without imposing a single progression path.The model covers five system properties and three organizational properties.
  • 1 Introduction: Four trust paradigms describe how trust is established, while the dimensions describe what aspects of systems and organizations require trustworthiness.The paradigms are operational, engineering, statistical, and containment-based.
  • 1 Introduction: The alignment hypothesis proposes that effective trustworthiness depends on fit between granted AI authority and the dimensions supporting it, conditional on the trust paradigm.The authors state that misalignment can produce pathological risk or unnecessary friction and offer the hypothesis for future quantitative testing.

2 Background and Related Work

Existing trustworthy-AI, maturity-model, and AI-engineering governance literature provides properties, ladders, and issue-specific treatments, but lacks an empirically grounded multidimensional account of operational practice. The paper positions its health-check model as filling that gap.

  • 2 Background and Related Work: Trustworthy-AI literature defines properties and assurance vocabularies but does not centrally address how organizations configure partially achievable properties as a portfolio.The gap concerns operational trade-offs under heterogeneous GenAI adoption conditions.
  • 2 Background and Related Work: Maturity models typically represent organizations through progressive levels on a single axis, despite evidence that organizations and organizational units occupy mixed positions.The paper connects this limitation to the need for multidimensional characterization.
  • 2 Background and Related Work: Recent autonomy and AI maturity models broaden covered practice areas, but still do not provide the paper’s empirically grounded configuration structure across diverse contexts.The reviewed models address autonomy progression or areas such as governance, skills, data, and responsible AI.
  • 2 Background and Related Work: Governance and assurance research for AI-assisted software engineering addresses concrete concerns such as supply-chain security and prompt injection but remains fragmented across individual issues.The literature lacks a structure tying these concerns together as a portfolio of trustworthiness practices.
  • 2 Background and Related Work: Across the three streams, no reviewed framework provides an empirically grounded, multidimensional structure for characterizing trustworthy autonomy across industrial contexts.The paper presents its health-check model as the contribution filling this gap.

3 Research Method

The paper uses a purposive, inductive interview study conducted by three geographically distributed research teams. Cross-team coding, saturation judgments, and model synthesis produced the dimensions and paradigms while preserving methodological heterogeneity as a consideration for interpretation.

  • 3 Research Method: Three research teams conducted an inductive expert interview study focused on GenAI’s role in constructing and evolving trustworthy software-intensive systems.The teams were based in Sweden, Italy, and the United States and converged on a shared topic.
  • 3 Research Method: Participants were purposively selected for variation in application domain, system criticality, and organizational role, with senior practitioners directly involved in GenAI adoption or embedded systems.The sample prioritized analytical richness rather than statistical representativeness.
  • 3 Research Method: The study treats parallel team protocols as distinct perspectives, while acknowledging implications for construct validity and interpretation.The teams used a shared topic but differed in angles, and consensus disagreements were resolved through discussion or conservative readings.
  • 3 Research Method: Analysis proceeded through open, axial, and selective coding, with shared schemas, cross-team review, and candidate dimensions retained only when supported across multiple cases.The final axial-coding phase produced the eight-dimension structure.
  • 3 Research Method: Researchers used LLM-assisted summarization as a reviewed supporting tool, while identifying risks including paraphrasing drift, sub-claim aggregation, and loss of verbatim grounding.Reported quotations were extracted directly from transcripts rather than AI-generated summaries.
  • 3 Research Method: The four-paradigm framing and eight-dimension structure emerged inductively as teams compared interviews and revised, discarded, or retained candidate themes.Containment-based trust was added after the initial three-paradigm framing as further rounds were incorporated.
  • 3 Research Method: Saturation was judged when later interviews reinforced or refined the dimensional structure rather than adding structural novelty, although substantively different contexts might alter it.The paper specifically notes possible change in underrepresented industries or geographies.

4 Findings

The findings section presents inductive empirical patterns from eighteen interviews, organized thematically rather than by the model’s dimensions. Direct transcript quotations are used to ground the analysis while preserving respondent anonymity.

  • 4 Findings: Findings are organized around inductively emerging themes in five steps, with the health-check model presented later as their synthesis.This organization separates empirical findings from the model construction.
  • 4 Findings: Interview quotations are drawn directly from transcripts and anonymized by respondent reference rather than attributed by name or organization.The anonymity practice preserves analytical context while withholding identifiable affiliations.

4.1 What ”Trustworthy” Means in Practice

Practitioners converge on trustworthiness as predictable, repeatable performance that fulfills intended behavior without unintended actions. This operational standard is difficult for probabilistic GenAI systems, motivating distinct strategies for establishing trust.

  • Trustworthy systems do what they are supposed to do, do so consistently, and do nothing they are not supposed to do.
  • Predictable outputs, stable responses to repeated inputs, and the absence of unintended behavior are central practical requirements.
  • Trustworthiness also encompasses robustness, repeatability, reliability, human oversight, security, privacy, and grounded results across application domains.
  • Respondents describe trust as built incrementally through reliable and consistent system behavior rather than as an inherent property.
  • GenAI’s probabilistic behavior creates an “8 out of 10 correct” problem because users cannot reliably identify incorrect outputs or validate them.
  • The four trust paradigms are strategies for reconciling practitioners’ convergent operational definition with GenAI’s probabilistic reality.

4.2 Four Trust Paradigms

The interviews identify four non-progressive trust paradigms—operational, engineering, statistical, and containment-based—that establish trust differently across system configurations. Their limitations show why trust practices must match feedback latency, assurance scope, probabilistic behavior, and runtime change.

  • Four trust paradigms emerged: operational, engineering, statistical, and containment-based; they are distinct epistemologies rather than maturity levels.Organizations may combine several paradigms across subsystems or lifecycle stages.
  • Operational trust: Operational trust relies on measured outcomes and feedback loops, extending established approaches to complex software rather than treating GenAI as categorically different.In banking, human approval remains an explicit trust anchor alongside outcome measurement.
  • Operational trust: Operational trust weakens when outcome feedback is delayed, weak, or absent, especially where failures are silent or consequences are catastrophic.The paradigm works best when feedback loops close quickly and poorly when detection is delayed.
  • Engineering trust: Engineering assurance must remain continuous and future-facing because deployment can outpace verification and autonomy can drift beyond the validated scope.Practitioners described forward-looking verification across several releases and runtime drift detection.
  • Engineering trust: Engineering trust decomposes autonomy into validated components and uses independent guardrails, quality gates, and staged assurance to control risk.One practitioner described separating the AI that generates outputs from the AI that checks them; another broke an autonomous-vehicle vision into 250 constituent parts.
  • Statistical trust: Statistical trust accepts variable outputs when factual data remains correct and defines success through flexible distributional criteria rather than binary first-time correctness.Agents may iterate rapidly over candidate solutions until they reach an acceptable margin of error.
  • Statistical trust: Current statistical-trust mitigations can be costly and incomplete: adversarial manipulation may trigger a full reset that discards prior learning.This was presented as an unsatisfactory description of the current state.

4.3 The dimensions along which organizations vary

Organizations vary across eight dimensions of trustworthy autonomy, spanning system properties and organizational capabilities. These dimensions are independent of a single maturity progression and include authority, assurance, data, containment, traceability, governance, oversight, and workforce sustainability.

  • Model structure: The model organizes trustworthy autonomy around five system-layer and three organizational-layer dimensions.The system layer covers authority, assurance, data trustworthiness, architectural containment, and traceability/comprehensibility; the organizational layer covers governance, human oversight, and workforce capability sustainability.
  • Model structure: Each dimension uses five ordered analytical levels that denote scope or capability, not inherently better organizational states.The levels characterize positions along distinct axes rather than stages in a maturity progression.
  • Scope of Agent Authority: Agent authority varies across lifecycle activities and architectural layers, with organizations often granting AI broad coding authority while retaining human control over architecture.Examples include autonomous reverse engineering and modernization alongside human control of high-stakes architecture, and different authority levels for system, component, and module work.
  • Assurance Mechanisms: Assurance is established through evaluation, formal safety frameworks, and emergent process-based validation, while current methodologies can lag rapid deployment.Enterprise organizations emphasize constant validation, ground-truth datasets, feedback loops, and cross-model checks; safety-critical contexts use frameworks such as ISO PAS 8800.
  • Data Trustworthiness: Data trustworthiness is constrained by sovereignty, export controls, privacy, ownership, licensing, and limited access to organizational or customer data.These constraints can determine model and infrastructure choices, prevent raw PII exposure, and limit deployment where data rights or ownership remain unresolved.
  • Architectural Containment: Architectural containment separates AI from deterministic cores and isolates generator and guardian components so that AI outputs are independently checked or grounded.The corpus includes independently developed generator-guardian systems, cross-model validation, and retrieval-augmented grounding.

4.3.5 Traceability and Comprehensibility (system layer)

Traceability and comprehensibility address whether decisions can be reconstructed and whether people retain system-level understanding as AI-generated components accumulate. The organizational layer complements this with governance, human oversight, and sustainable workforce capability.

  • Traceability and Comprehensibility: Traceability concerns reconstructing decisions, while comprehensibility concerns retaining understanding of the whole system over time.The paper combines them because both depend on the organizational competence to know what the system does and explain it.
  • Traceability and Comprehensibility: Shadow IT, absent provenance, and limited explainability make traceability a significant issue in regulated and public-sector settings.Respondents identified traceability as a current problem and a future improvement priority, while standards work aims to formalize it.
  • Traceability and Comprehensibility: Loss of architectural visibility and hidden bugs arise when AI-generated code outpaces human capacity to debug, validate, and maintain it.The paper characterizes “vibe coding” as code that appears to work but cannot be adequately validated or maintained.
  • Governance: Governance differs sharply by sector: regulated organizations adapt to external rules, whereas enterprise services develop internal governance as adoption evolves.Banking, defence, and automotive contexts use explicit mandates or standards, while other organizations rely on emerging centers of excellence and internal methodology.
  • Human Oversight Posture: Human-in-the-loop is the current default, with a widely anticipated shift toward human-on-the-loop oversight as confidence develops.The transition is described as incremental, while legal requirements can prevent human-on-the-loop deployment for lethal autonomous decisions.
  • Workforce Capability Sustainability: Workforce capability sustainability concerns whether organizations preserve expertise and develop practitioners able to oversee, validate, and evolve AI-driven systems.Respondents describe K-shaped productivity, skill degradation, and new roles such as AI trainers, intelligence engineers, and intelligence architects.

4.4 Domain-specific patterns

The corpus shows three dominant, overlapping domain patterns: safety-critical and regulated organizations, enterprise software services, and exploratory or transformation contexts. Their trust paradigms, authority levels, governance, and binding constraints differ.

  • Safety-critical and regulated: Safety-critical and regulated organizations favor engineering or containment paradigms, constrained authority, strong assurance requirements, and mandated or policy-based oversight.Their data practices are also constrained by sovereignty, export controls, or sector-specific requirements.
  • Enterprise software services: Enterprise software services combine high development-time agent authority with low product-runtime authority and internally evolving governance.Validation burden and workforce capability sustainability are the principal concerns, with the K-shaped problem more prominent than in safety-critical contexts.
  • Exploratory and transformation contexts: Exploratory and transformation contexts have less developed trust paradigms, with authority negotiated organizationally and data and governance questions dominating technical ones.At early adoption stages, organizations are still deciding what to do with the technology rather than only how to assure it.
  • Cross-domain qualification: These patterns are dominant tendencies rather than exhaustive or mutually exclusive organizational categories.The paper notes that many organizations exhibit mixtures of the three patterns.

4.5 Recurring tensions

Four recurring tensions structure the corpus: speed versus assurance, development use versus product use, continuity versus discontinuity, and productivity versus comprehensibility. These tensions expose why trustworthy autonomy cannot be reduced to a single progression.

  • Speed versus assurance: GenAI iteration often outpaces assurance construction, creating a speed-versus-assurance tension across safety-critical and enterprise settings.Practitioners describe fast generation alongside slow validation and assurance methods designed for more discrete release cycles.
  • AI for development versus AI in products: Trustworthiness concerns differ between using AI for development and embedding AI in products, even though many dimensions apply to both.Respondents explicitly distinguish the two uses while noting that the same dimensions may carry different levels of concern.
  • Continuity versus discontinuity: Organizations differ in whether they view GenAI as an acceleration of existing opacity or as a discontinuous technology requiring new methods.This continuity-versus-discontinuity reading shapes the trust paradigm an organization adopts.
  • Productivity versus comprehensibility: GenAI can generate code faster than people can comprehend, validate, and maintain it, linking productivity gains to architectural opacity and workforce risks.The corpus connects this tension with K-shaped productivity, loss of architectural visibility, and “vibe coding,” especially where authority is high and comprehensibility is weak.
  • Cross-cutting implication: The four tensions are presented as observable patterns that any account of trustworthy autonomy in GenAI-assisted software engineering must confront.

5 The Trustworthy Autonomy Health Check Model

The Trustworthy Autonomy Health Check Model represents trustworthy autonomy as a multidimensional configuration rather than a maturity ladder. It combines eight dimensions, four cross-cutting trust paradigms, and an alignment hypothesis linking agent authority to enabling capabilities.

  • 5 The Trustworthy Autonomy Health Check Model: The model is a structured configuration space, not a maturity ladder or single aggregated score.Dimensions retain separate analytical weight, and healthy configurations depend on alignment rather than uniformly reaching the highest level.
  • 5.2.1 System-Layer Dimensions: The five system-layer dimensions cover agent authority, assurance, data, containment, and traceability, with empirical variation across organizations.Low traceability is especially problematic under high AI authority, while level-5 containment is exemplified by zero-trust and generator-guardian architectures.
  • 5.2.2 Organization-Layer Dimensions: The organizational layer covers governance, human oversight, and workforce capability sustainability, with human-in-the-loop oversight the current default and workforce capability often eroding or unmanaged.No respondent described operational level 4 or 5 workforce capability sustainability, while level 4 oversight remains exceptional.
  • 5.3 The trust paradigms as cross-cutting: The same dimensional position can have different meanings under different trust paradigms, so dimension-only and paradigm-only readings each obscure important variation.Level-3 assurance may suffice operationally when outcomes are measurable but remain insufficient under engineering or statistical trust and irrelevant under containment-based trust.
  • 5.4 The alignment hypothesis: The alignment hypothesis states that trustworthy autonomy depends on matching agent authority with assurance, data, containment, and oversight capabilities under the dominant paradigm.Misalignment produces pathological risk when authority exceeds enabling dimensions, or unnecessary friction when enabling capability exceeds authority.
  • 5.4 The alignment hypothesis: The model is intended as a practical diagnostic that asks organizations whether authority, enabling dimensions, and trust paradigm are aligned.Its purpose is to support configuration diagnosis rather than prescribe a universal progression.

6 Threats to Validity

The paper identifies validity threats arising from heterogeneous interview protocols, inductive interpretation, AI-assisted coding, respondent selection, and limited transferability. It reports mitigations including cross-corpus testing, audit trails, convergence across coding teams, and transparent scope claims.

  • 6 Threats to Validity: The most significant construct-validity concern is that three research teams applied the common interview protocol differently.The protocol differences could influence which constructs respondents discussed and how they were expressed.
  • 6 Threats to Validity: The authors mitigate construct concerns by retaining dimensions supported across the corpus and corroborating trust paradigms across interview sets.They also report constructs in practitioners’ own terms where possible.
  • 6 Threats to Validity: Inductive interpretation cannot establish that the eight dimensions are the uniquely correct decomposition or that the four paradigms are exhaustive.Another research team could reasonably derive a different decomposition from the same corpus.
  • 6 Threats to Validity: The synthesis may over-weight articulate or analytically distinctive respondents, potentially producing different emphases than a reading centered on typical cases.The authors identify this as a structural feature of their cross-case synthesis rather than a defect.
  • 6 Threats to Validity: Transferability is limited by geographic, organizational, professional, and participation biases in the sample.Small startups, several world regions, compliance and legal perspectives, and opaque organizations are underrepresented.
  • 6 Threats to Validity: Three coding teams produced converging interpretations, but the study reports no formal inter-coder reliability coefficient and LLM-assisted summarization may be difficult to reproduce exactly.Structured codings provide an audit trail from transcripts to synthesis.

7 Conclusion and Future Work

The paper concludes that trustworthy autonomy is a context-dependent configuration rather than a single progression path. It proposes the health check as a multidimensional diagnostic and identifies quantitative, longitudinal, recursive, and domain-specific extensions.

  • 7 Conclusion and Future Work: Organizations converge on trustworthiness as predictability, repeatability, and absence of unintended behavior but differ in how they establish it.The paper identifies operational, engineering, statistical, and containment-based trust paradigms.
  • 7 Conclusion and Future Work: The model identifies eight empirically derived dimensions spanning system properties and organizational properties.The system layer contains five dimensions, while the organizational layer contains three.
  • 7 Conclusion and Future Work: The alignment hypothesis links agent authority to enabling dimensions and predicts pathological risk or unnecessary friction when they are misaligned.The hypothesis is presented as a proposition for future quantitative testing.
  • 7 Conclusion and Future Work: The configurations show that identical nominal dimensional positions can be healthy in one domain and pathological in another.This supports interpreting the model through domain and trust-paradigm context.
  • 7 Conclusion and Future Work: Future work includes quantitative tests of misalignment, longitudinal studies of configuration change, and recursive health checks across architectural layers.Quantitative studies could relate misalignment indicators to incident rates, defect density, productivity, and other outcomes.
  • 7 Conclusion and Future Work: The paper’s central contribution is making trustworthy-autonomy configurations visible and assessable at multidimensional analytical resolution.It argues that trustworthy autonomy is a configuration to align rather than a destination on one road.

Declarations

The declarations report funding, ethics and consent procedures, authorship, data-access restrictions, conflicts of interest, and the nonapplicability of clinical-trial registration.

  • Declarations: The research was supported by Software Center and the U.S. National Science Foundation, whose funders had no role in the study or publication decision.The NSF grant number is 2232721.
  • Declarations: The study involved consenting adult professionals in semi-structured interviews and did not involve patients, vulnerable populations, or clinical or behavioral intervention.Participants consented to recording and automated transcription and could withdraw without consequence.
  • Declarations: All authors contributed to study conception, design, execution, data collection, coding, synthesis, and model formalization.The author-contribution statement describes collaborative cross-corpus analysis and drafting responsibilities.
  • Declarations: Interview transcripts and per-interview codings are not publicly available because of confidentiality commitments to participants and organizations.The restriction covers sensitive internal practices, projects, and organizational decisions shared under limited-access conditions.
  • Declarations: The authors disclose industry consulting and research partnerships overlapping with some interview domains and describe safeguards against influence.No respondent was interviewed about a project involving an interviewing author’s ongoing consulting relationship.
  • Declarations: Clinical-trial registration is not applicable.
Loading 2609.10595v1…