Source-linked AI summary
LAAF: A Layered Accountability Architecture Framework for LLM Applications
Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi, Muhammad Waqas, George Loukas, Alireza Jolfaei, Lucas Cordeiro, Pierre Dantas
TL;DR
LLM applications can produce fluent but ungrounded outputs in high-stakes settings, raising unresolved questions about answerability and traceable responsibility. The review synthesises accountability concepts and mechanisms, classifies them across four layers, and maps them to regulatory and sectoral frameworks. It identifies persistent gaps and consolidates the evidence into LAAF, while presenting the architecture as unvalidated.
Problem
LLM hallucinations are structurally persistent, while existing surveys do not integrate who answers for harm with the mechanisms, obligations, and actors involved.
Method
The review synthesises accountability mechanisms across four families, organises the corpus through four layers and three cross-cutting properties, and maps the result to regulatory frameworks.
Results
The review identifies four persistent gaps and consolidates the surveyed evidence into LAAF, an integrated four-layer accountability architecture with cybersecurity aligned to the OWASP LLM Top 10 (2025).
Takeaways & Limitations
Accountability for LLM applications is presented as specifiable and decomposable across the lifecycle through layered mechanisms, roles, traceability, and monitoring.
Takeaways & Limitations
LAAF is most directly applicable to deployments under the analysed frameworks, and the review’s author-scored comparison and search boundaries remain threats to validity.
Abstract
from arXiv · showhide
Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated as authoritative even when ungrounded or incorrect. When such an output contributes to harm, who is answerable, and through what mechanisms can responsibility be traced, explained, and acted upon? Following PRISMA guidance, five databases were searched from January 2022 to March 2026 against four review questions; of 4,512 records identified, 122 primary studies were included, together with 12 regulatory and standards documents analysed as primary sources. The review consolidates a sociotechnical account of accountability as an actor-forum relation resolved into five dimensions, and synthesises mechanisms across four families: technical controls, human oversight, organisational governance, and documentation and traceability, each with a maturity assessment. The corpus is read through a four-layer classification device spanning provenance, application logic, human oversight, and governance and redress, cross-cut by traceability, role clarity, and continuous monitoring. Both are mapped onto the EU AI Act, whose high-risk obligations have applied since 2 August 2026, the NIST AI RMF with its Generative AI Profile, ISO/IEC 42001, and sectoral guidance in healthcare, consumer finance, education, and the public sector. Four persistent gaps emerge: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation, alongside five structural tensions that no surveyed instrument resolves. The review closes by consolidating the classification device into an integrated accountability architecture, LAAF, with cybersecurity aligned to the OWASP LLM Top 10 (2025); it is a synthesis of the surveyed evidence rather than a validated artefact.
1 Introduction
LLM accountability remains difficult because fluent outputs can be ungrounded, while responsibility, evidence, and follow-up are not reliably assigned across sociotechnical deployments. This review integrates conceptual, technical, organisational, regulatory, sectoral, and cybersecurity perspectives into a layered account.
- The accountability imperative: LLMs can produce fluent, factually wrong statements even without the knowledge required, making hallucinations a high-stakes accountability concern.Formal results indicate that hallucinations persist under calibration constraints and for computable LLMs, although mitigation is not thereby futile.
- The accountability imperative: Existing responsible-AI concepts do not by themselves specify who answers for harm or what happens next.The review identifies this unresolved actor-and-consequence question as the gap it addresses.
- Review focus: The review investigates four recurring gaps: under-specified human oversight, absent shared accountability metrics, disciplinary disconnection, and limited empirical evaluation.The corpus also distinguishes accountability from transparency, explainability, and liability rather than treating them as interchangeable.
- Review scope and contributions: The study systematically synthesises accountability mechanisms across technical controls, human oversight, organisational governance, and documentation and traceability.Each family receives a maturity assessment and regulatory anchor.
- Review scope and contributions: It maps the literature onto the EU AI Act, NIST AI RMF with its Generative AI Profile, ISO/IEC 42001, and sectoral guidance across four sectors.The sectoral mapping covers healthcare, consumer finance, education, and the public sector and employment.
- Review scope and contributions: LAAF consolidates the surveyed evidence into an integrated accountability architecture with three cross-cutting properties and cybersecurity aligned to the OWASP LLM Top 10 (2025).The framework is presented as a synthesis of the corpus rather than a validated artefact.
2 Background
LLMs are dynamic probabilistic systems whose lifecycle-spanning harms motivate accountability mechanisms that can detect, assign, and review responsibility. The background distinguishes five lifecycle phases, five hallucination types, four detection families, and four accountability layers.
- LLM properties: LLMs generate outputs from statistical regularities rather than validated knowledge, while their opaque and changing behaviour complicates accountability.Providers may update weights and safety layers, while deployers alter retrieval corpora and prompt logic.
- Lifecycle harms: Harms originate across five interdependent lifecycle phases and accumulate rather than remaining confined to a single point.This motivates treating accountability as a lifecycle property.
- Hallucinations: Five recurring hallucination types are factual contradiction, contextual irrelevance, logical inconsistency, temporal disorientation, and ethical violation.The classification is descriptive because the underlying studies use incompatible definitions, prompts, and domains.
- Detection formalisms: Hallucination detection branches into four method families based on external-reference availability and white-box access to model states.The families are black-box reference-free, white-box internal-state, retrieval alignment, and retrieval combined with an internal check.
- Detection formalisms: Retrieval alignment flags outputs when claim support falls below a deployment-specific threshold, while dispersion can operate without a reference corpus but misses confident consistent errors.A clinical example flags g(x, R) = 0.75 against τ = 0.85 and routes the output to review.
- Accountability architecture: The survey reads accountability through four layers—provenance, application logic, human oversight, and governance and redress—cross-cut by traceability, role clarity, and continuous monitoring.The device organises heterogeneous literature rather than constituting a proposed design at this stage.
3 Methodology
The review used a prespecified PRISMA-informed protocol, five database searches, staged screening, structured extraction, and narrative thematic synthesis. The process yielded 122 included studies and 12 regulatory documents, with explicit quality scoring and recorded limitations.
- Protocol: The review protocol specified its questions, search strategy, eligibility criteria, extraction template, and synthesis approach before searching began.It followed PRISMA reporting guidance and evidence-based systematic-review practice.
- Limitations: The protocol was not registered with PROSPERO or OSF, and residual scoring subjectivity and English-only inclusion were recorded as limitations.The scoring rubric used two independent passes reconciled by re-reading the source text, while structural tables were single-rater scored.
- Search strategy: Searches covered five peer-reviewed databases from 1 January 2022 to 31 March 2026 using Boolean blocks for LLMs, accountability, and application domains.The databases were ACM Digital Library, IEEE Xplore, ScienceDirect, SpringerLink, and Wiley Online Library.
- Eligibility: Eligibility required peer review, indexing in the selected databases, publication between January 2022 and March 2026, accountability relevance, and English-language publication.Authoritative regulatory and standards documents were added as primary sources outside database screening.
- Screening: 4,512 database records were identified, 3,825 remained after deduplication, and 122 studies were included after sequential screening, alongside 12 regulatory documents.Citation chasing added 9 eligible studies from 38 candidates; regulatory documents were analysed separately and not counted among the 122.
- Synthesis: Synthesis combined narrative synthesis with Braun and Clarke’s six-phase thematic analysis for heterogeneous evidence unsuitable for quantitative meta-analysis.Themes were derived deductively, refined inductively, and applied to the full corpus in a second pass.
- Quality appraisal: Each record received equally weighted scores for empirical rigour, reproducibility, and conceptual contribution, with totals of 7–9 high, 4–6 medium, and 3 conceptual-only.Scores weighted the synthesis narratively rather than serving as an inclusion filter.
4 Review Process Reporting and Dataset Characteristics
The review screened and characterised a corpus of 122 included studies, supplemented by 12 regulatory and standards documents, using a staged PRISMA process and an explicit scoring rubric. The evidence base was methodologically diverse, recent, and dominated by medium-quality or conceptual work, while available benchmarks measured outputs rather than accountability mechanisms.
- Review Process: 4,512 records were identified, and 122 studies were included, representing 2.7% of the initial records.Records were deduplicated and screened through a distinct introduction-and-conclusion stage before full-text assessment.
- Review Process: 12 regulatory documents were added as primary sources but were not counted among the 122 included studies.They were analysed separately and used as mapping targets.
- Dataset Characteristics: The 122 studies comprised conceptual or survey 39 (32%), empirical case studies 28 (23%), technical mechanism 33 (27%), and regulatory or governance 22 (18%).Publication was concentrated in 2025 and early 2026, with 45 (37%) studies in 2025 and 16 (13%) from January–March 2026.
- Dataset Characteristics: 28 studies (23%) scored high, 71 (58%) medium, and 23 (19%) conceptual-only under the review rubric.The predominance of medium-quality, simulated, and conceptual work motivated the empirical-evaluation gap.
- Evaluation Resources: Existing public benchmarks cover hallucination, factuality, bias, fairness, and security but do not measure oversight, governance closure, or enforcement responsiveness.The measurement gap concerns accountability dimensions rather than output quality.
5 Conceptual Foundations of Accountability in the Literature
The review defines accountability in LLM applications as a sociotechnical actor-forum relation requiring standards, explanation, process, and consequences. It resolves this relation into five dimensions and highlights distributed responsibility, weak oversight specification, missing metrics, disciplinary separation, and limited empirical evaluation.
- Conceptual Definition: Accountability requires an actor to explain and justify conduct to a forum against standards, through a defined process with possible consequences.The relation is distinct from transparency, explainability, or responsibility alone.
- Conceptual Definition: The review defines accountability as identifying lifecycle actors, requiring justification of design and deployment choices, preserving output traceability, and enabling oversight, correction, and redress.This definition treats accountability as broader than model interpretability.
- Structural Complications: Accountability is complicated by attributability gaps, multi-layered dependencies, and a dual-phase interaction that obscures the locus of human responsibility.These conditions make responsibility distributed across models, interfaces, retrieval systems, moderation layers, and human decision phases.
- Five Dimensions: Five dimensions are disclosure and justification, oversight, organisational and governance, moral and professional responsibility, and enforcement and redress.Together they connect technical choices, human review, institutional roles, responsibility, and responses to failure.
- Persistent Gaps: Human oversight is under-specified, shared accountability metrics are absent, technical and governance literatures are disconnected, and empirical evaluation is limited.Oversight and disclosure dominate the literature, while enforcement and redress are treated least often.
6 Accountability Mechanisms: A Systematic Synthesis
The synthesis organises accountability mechanisms by four intervention families and four application layers, treating them as complementary rather than interchangeable. Maturity is strongest for grounding, retrieval, logging, and provenance documentation, while oversight and governance remain less empirically established and several mechanisms document intent without proving outcomes.
- Architecture: Technical controls, human oversight, organisational governance, and documentation and traceability are distinct mechanism families, and no family or layer suffices alone.Documentation has different accountability owners depending on whether it is a model card or an incident record.
- Layer 1: Provenance: Model and system cards, datasheets and data provenance, evaluation and red-team reports, and watermarking are the dominant provenance mechanisms.Cards and datasheets are high maturity, while watermarking is emerging without a standardised production scheme.
- Layer 2: Application Logic: Grounding and retrieval is the most widely adopted application-logic control, with RAG constraining generation to retrieved documents and producing significant hallucination reductions.Guardrails filter inputs and outputs, while uncertainty quantification, abstention, logging, audit trails, and output-level provenance provide additional controls.
- Layer 2: Application Logic: RAG and logging are high maturity, guardrails and uncertainty quantification medium, and output-level provenance emerging without a standardised schema.These mechanisms operate on output surfaces rather than establishing correctness or exposing underlying reasoning.
- Layer 3: Human Oversight: Substantive oversight involves active interrogation of evidence, uncertainty, and reasoning through review, monitoring, escalation pathways, and override authority.Its maturity is medium because independent evaluation is mostly simulated and effectiveness depends on reviewer capacity, expertise, and institutional willingness to override.
- Layer 4: Governance and Redress: Role allocation, governance committees, red teaming, incident response, post-deployment monitoring, and formal AI management systems dominate organisational governance.Maturity is medium, and structures may exist without altering deployment decisions.
- Deployment Conditions: Mechanism selection depends on whether an authoritative reference corpus exists and whether a human sees the output before action.Where grounding is unavailable or outputs act directly, accountability weight shifts toward human oversight and other layers.
7 Regulatory and Standards Convergence
The review maps accountability across binding regulation, operational risk guidance, certifiable management systems, and sector-specific guidance. These instruments converge on lifecycle accountability but differ in legal force, operational specificity, certification, and domain calibration, supporting a common layered architecture with sector adaptation.
- Scope: The mapped landscape covers the EU and United States, NIST AI RMF and GenAI Profile, ISO/IEC 42001, and four sectoral regimes.The mapping excludes the United Kingdom’s regulator-led approach and China’s generative-AI measures.
- EU AI Act: The EU AI Act assigns provider Chapter V obligations and deployer high-risk obligations when an LLM is used in a high-risk context.This regulatory division mirrors the review’s sociotechnical allocation of accountability.
- EU AI Act: The EU AI Act’s four core obligations are human oversight, documentation and logging, post-market monitoring and incident reporting, and cybersecurity.These obligations connect Articles 14, 11–12, 72–73, and 15 to the architecture’s layers.
- EU AI Act: High-risk obligations under Article 6(2) have applied since 2 August 2026, with non-compliance penalties reaching €15 million or 3% of worldwide turnover.Incorrect, incomplete, or misleading information to authorities has a separate tier of up to EUR 7.5 million or 1% of turnover.
- NIST: NIST AI RMF organises practice around Govern, Map, Measure, and Manage, while its 2024 GenAI Profile identifies twelve generative-AI risks.The RMF is voluntary but gained practical force through procurement, regulators, and state legislation.
- ISO/IEC 42001: ISO/IEC 42001, published in December 2023, is the first international standard specifying requirements for an artificial-intelligence management system.Its controls cover policy, roles, impact assessment, lifecycle, data, transparency, and third-party relationships, but adoption evidence remains thin.
- Instrument Convergence: The EU AI Act, NIST AI RMF, ISO/IEC 42001, and sectoral guidance are complementary: binding force, operational specificity, certifiable conformity, and domain calibration.No single instrument supplies all four properties.
- Lifecycle and Sectoral Calibration: All four instrument types require accountability before, during, and after deployment, while sectoral guidance calibrates mechanisms to healthcare, finance, education, and public-sector settings.Their shared lifecycle requirements support one architecture serving multiple regimes, although the cost reduction has not been measured empirically.
8 Consolidated Findings Against the Review Questions
The review consolidates accountability as an answerability relation and finds that mechanisms, regulatory requirements, and evaluation conditions remain uneven across the LLM application lifecycle. Four persistent gaps and five structural tensions define the requirements that LAAF is designed to navigate.
- Conceptual findings: Accountability is an answerability relation between an actor and a forum, judged against standards through a defined process with consequences.The review resolves this account into five dimensions within a sociotechnical framing.
- Regulatory convergence: The literature converges on common lifecycle requirements across regulatory and standards instruments, but the cost of serving several jurisdictions remains unmeasured.The comparison covers the EU AI Act, NIST AI RMF, ISO/IEC 42001, and sectoral guidance.
- Persistent gaps: The corpus leaves operational detail unresolved for output provenance, substantive oversight under Article 14, and contestability of adverse decisions.These shortcomings limit how directly accountability requirements can be implemented.
- Cross-cutting patterns: Mechanism maturity declines with distance from the model, while regulatory specificity increases, concentrating risk where obligations are most specific and evidence is thinnest.Detection, grounding, retrieval, and logging are more mature than oversight and governance mechanisms.
- Persistent gaps: Four gaps concern under-specified human oversight, absent shared metrics, disciplinary disconnection, and limited empirical evaluation.Few studies test whether oversight intercepts harmful outputs or incident response produces durable change.
- Structural tensions: Five tensions involve transparency and security, automation and oversight, standardisation and contextual fit, traceability and privacy, and deployment speed and governance.LAAF treats these tensions as design constraints to navigate rather than problems fully resolved by existing instruments.
9 Synthesis: Toward an Integrated Accountability Architecture
LAAF consolidates the review’s evidence into a four-layer architecture that assigns responsibilities, information flows, and intervention points across the lifecycle. Three cross-cutting properties make the layers auditable, while regulatory mapping and cybersecurity integration connect the architecture to implementation contexts.
- Architecture overview: LAAF comprises Foundation Model and Provenance, Application Logic and Guardrails, Human Oversight and Review, and Governance, Audit, and Redress.The layers are ordered by distance from the model and cross-cut by traceability, role clarity, and continuous monitoring.
- Architecture overview: Accountability operates bidirectionally between adjacent layers: artefacts provide evidence upward, while policy binds lower-layer review and guardrail regimes.Authority flows downward only within the deploying organisation’s control, not across the provider boundary.
- Cross-cutting properties: Traceability reconstructs consequential decisions, role clarity assigns overrides and sign-offs to named individuals, and continuous monitoring triggers defined responses to deviations.Monitoring includes thresholds for abstention, override, retrieval failure, and drift, with escalation to review and incident records.
- Regulatory mapping: The layers map to EU AI Act, NIST AI RMF, and ISO/IEC 42001 requirements, with layer-specific indicators and cybersecurity anchors.Layer 4 carries the broadest regulatory mapping, while Layer 1 aligns with Chapter V, NIST Map, and design-and-development controls.
- Cybersecurity integration: Cybersecurity is treated as integral to accountability because manipulated outputs cannot be fully attributed to an accountable deployment.Layer 2 carries the largest OWASP share because retrieved content is an injection surface; Layer 3 requires reviewers trained to detect adversarial outputs.
- Sector specialisation: LAAF preserves a constant four-layer, three-property structure while calibrating its content across healthcare, finance, and public-sector examples.The architecture is presented as sector-agnostic in structure but sector-calibrated in instantiation.
- Comparative positioning: Compared with fourteen existing instruments, LAAF is described as the only one covering all seven criteria, although the comparison is author-scored.Coverage concentrates on oversight and regulatory mappability; cybersecurity and sector calibration are uncommon among existing instruments.
10 Open Questions and Research Priorities
The research agenda targets measurement, technical artefacts, governance practice, and collaboration needed to evaluate and operationalise accountability in real deployments. Priorities include benchmarks, provenance and uncertainty standards, effective oversight criteria, sector calibration, and cross-community work.
- Measurement: Future work needs accountability benchmarks for oversight substantiveness, governance closure, enforcement responsiveness, and the five accountability dimensions.Empirical case studies of accountability practice in real deployments are also called for.
- Measurement: LAAF evaluation requires criteria for completeness, end-to-end traceability, oversight substantiveness, and governance closure, none of which currently has an instrument.A retrospective analysis of documented deployment failures is proposed as a cheaper first step.
- Technical artefacts: Technical priorities include auditable RAG records linking claims to retrieved passages, human-meaningful uncertainty, and standardised output-level provenance metadata.These artefacts are intended to support faithfulness assessment, reviewer decisions, and Layer 2 interoperability.
- Governance practice: Governance research should operationalise Article 14 oversight, assign outputs to review regimes, match review volume to reviewer capacity, and calibrate Layers 3 and 4 by sector.Healthcare, consumer finance, and public administration are identified for sector-specific duty-of-care calibration.
- Collaboration: Cross-community priorities include shared vocabulary, joint computer-science and human-factors work, academic–regulator engagement, adversarial benchmarks, and integrated AI-and-information-security management.The proposed management integration combines ISO/IEC 42001 and ISO/IEC 27001.
11 Limitations and Threats to Validity
The review’s validity is constrained by qualitative synthesis, author-scored comparisons, temporal and language limits, non-registration, and separate treatment of preprints. LAAF itself is scoped to non-agentic deployments mapped to the analysed instruments and assumes sufficient governance capacity.
- Threats to validity: The review faces construct, internal, external, and conclusion threats, including imprecise accountability boundaries, researcher and publication bias, and qualitative synthesis.The protocol’s definitional stance was not coded as a countable field.
- Internal validity: The framework comparison is author-scored; publishing the rubric mitigates but does not eliminate this residual internal threat.The same limitation is recorded for the comparative table’s scores.
- External validity: External scope is limited by the January 2022–March 2026 temporal window, English-only searching, sector-linked query terms, non-registration, and non-systematic preprint screening.LAAF is most directly applicable to deployments governed by the frameworks analysed in Section 7.
- Evidence limits: Quantitative comparison of accountability mechanisms across studies is not yet possible because the synthesis is qualitative.This limitation is itself reported as a review finding.
- Framework scope: LAAF does not cover agentic deployments, assumes governance capacity that small deployers may lack, and is mapped only to the instruments analysed.It would be disconfirmed by substantive accountability failures that remain undiagnosed despite all layers and properties being present.
12 Conclusion
The paper answers its motivating accountability question by treating LLM harms as distributed across sociotechnical lifecycles. It synthesises mechanisms, maps the literature onto regulation, and consolidates identified gaps and tensions into LAAF.
- LLM harms are distributed across the lifecycle because LLMs are sociotechnical systems.
- The paper synthesises accountability mechanisms across four families and maps the literature onto the binding regulatory landscape.
- Four persistent gaps and five structural tensions are consolidated into the LAAF accountability architecture.
Data Availability
The submission reports its review and scoring procedures and provides a replication package covering the included studies, extraction data, quality scores, and comparison rubrics. Generative AI was limited to language editing and was not used in the review's substantive processes.
- The search strategy, inclusion criteria, and quality-assessment rubric are reported in full, with scoring procedures documented separately.The scoring rubric for Tables 2 and 15 is given in Section 3.7.
- A replication package contains records for 122 included studies, the completed extraction sheet, per-study quality scores, and scored comparison rubrics.
- Generative AI was used only for language editing and not for study selection, data extraction, quality scoring, comparison scoring, or findings generation.