Source-linked AI summary

An Evaluation Framework for National AI Regulation

Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory, Sreeparvathy Sajeev, Utkarsh Tomar, Avyay M Casheekar

arXiv:2608.15417v1cs.CYcs.AI

TL;DR

Comparing national AI policies is difficult because portfolios combine instruments with different legal status, scope, and operating detail. This paper evaluates documented policy design and implementation readiness across eight jurisdictions, finding that portfolios combine statutes, agencies, guidance, and development measures in different ways.

  • Problem

    Existing comparisons do not provide a traceable assessment of national AI policy portfolios across instruments with different legal status, scope, and operating detail.

  • Method

    The framework evaluates versioned portfolios across technical risk, institutions, lifecycle coverage, affected people, public benefit, and economic policy using evidence-linked criteria.

  • Results

    The eight jurisdictions combine statutes, executive action, sectoral regulation, guidance, coordination, development policy, and data or cybersecurity law in different ways.

  • Takeaways & Limitations

    The portfolio, rather than one prominent law or strategy, is the appropriate unit for comparing national AI governance.

  • Takeaways & Limitations

    The framework evaluates documented design and implementation readiness but does not establish enforcement, compliance, or improvements in safety or individual rights.

Abstract

from arXiv · show

Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these national approaches is difficult. A binding rule and a detailed voluntary framework can address the same problem but create different duties. The resources needed to carry them out also differ by jurisdiction. This paper develops an evaluation framework for the documented design and implementation readiness of national AI policy. The comparison covers China, India, Japan, Singapore, South Korea, the United Kingdom and the United States. The European Union is included as a supranational comparator. The framework evaluates a versioned portfolio of official instruments rather than one prominent law or strategy. Its criteria ask whether the portfolio governs serious AI risks and whether responsible institutions can implement its commitments. They examine coverage across the AI lifecycle and the protections available to people affected by AI systems. Public benefit and responsible innovation remain a separate part of the assessment. Each sub-criterion is scored through ordered anchors and tied to the provision that supports the judgment. The protocol also records the source search, missing evidence, included instruments and cutoff date. The result is a traceable comparison of policy content that keeps category differences visible. It evaluates what a portfolio provides on paper. It does not estimate enforcement success or policy outcomes.

1 Introduction

Governments are adopting AI for administrative and public-service benefits amid rapid investment and falling costs, while confronting discrimination, privacy, accountability and implementation challenges. Because policy instruments create different duties and capacities, the paper develops a traceable framework for comparing national AI policy portfolios.

  • AI already supports translation, image recognition, fraud detection, administrative processing, resource allocation, public services and decisions affecting citizens.
  • Global corporate AI investment reached $252.3 billion in 2024, while fixed-performance querying costs fell more than 280-fold between late 2022 and late 2024.US private investment reached $109.1 billion, compared with $9.3 billion in China and $4.5 billion in the United Kingdom.
  • Public-sector AI can reproduce discrimination, expose sensitive records and create accountability questions when outputs influence consequential decisions.Institutions must determine who reviews outputs and who can act when systems produce errors or require changes.
  • Implementation requires procurement and monitoring staff, reliable data, technical infrastructure, continuing funding, workforce planning and coordination among regulators.
  • Statutes, strategies and voluntary guidance are not interchangeable because they differ in enforceability, detail, resource allocation and institutional capacity.The framework therefore evaluates versioned national policy portfolios by what they provide on paper, how clearly they provide it and the evidence supporting each judgment.

2 Background and Related Work

The paper situates its comparison across eight jurisdictions with different technical capacities, economic conditions and governance structures. It treats these conditions as implementation context while evaluating documented policy portfolios rather than inferring policy quality or outcomes from national capability.

  • Scope and context: These contextual conditions indicate which policy commitments may be feasible and where external dependence can hinder implementation, but they do not determine regulatory quality.The paper reports economic capacity and related ecosystem measures as context rather than converting them into higher policy scores.
  • Scope and context: The comparison covers China, India, Japan, Singapore, South Korea, the United Kingdom, the United States and the European Union as a supranational comparator.These jurisdictions were selected purposively to reflect different levels of technical capacity and forms of governance.
  • Scope and context: National AI ecosystems are assessed through compute infrastructure, economic and political conditions, domestic demand, research institutions and talent.Scientific supercomputers, private training clusters and commercial cloud regions provide different forms of access and should not be treated as equivalent.
  • Policy portfolios: Because the portfolios differ in legal force and institutional structure, the framework evaluates a portfolio of instruments instead of searching for one national AI law.The European Union’s risk-classification model and other jurisdictions’ differing development and protective emphases illustrate why instrument-level comparison is needed.
  • Emerging coverage: Agentic systems expose a newer coverage problem because deployed agents can take multi-step actions through tools while public safety disclosure remains limited.Singapore issued a dedicated government framework for agentic AI in January 2026, while other portfolios may rely on general risk duties, sectoral law, oversight or cybersecurity provisions.
  • Analytical boundaries: The framework combines principles, technical risks, individual rights, distributional effects, enforcement and resources while reporting these dimensions separately.It evaluates what declared portfolios provide and how specifically they provide it, without treating readiness rankings or policy language as evidence of outcomes.

3 Evaluation Framework · 3.1 Framework Development and Design

The framework translates serious AI risks and implementation concerns into traceable questions for evaluating versioned national policy portfolios. Ordered scoring, source and cutoff records, and category reporting support comparison and diagnosis while leaving reliability, validity, and enforcement outcomes unresolved.

  • 3.1 Framework Development and Design: The framework uses purposive design synthesis to construct an evaluation instrument rather than estimate concept prevalence in a fixed literature.Technical AI safety research identifies misuse, control, testing, intervention, and longer-term resilience concerns.
  • 3.1 Framework Development and Design: Policy concerns are translated into separately scoreable questions covering testing, access control, incident reporting, enforcement, accountability, and responsibility allocation.The questions make broad concerns answerable from public policy evidence and keep distinct institutional functions separate.
  • 3.1 Framework Development and Design: Candidate concepts are divided or merged to preserve distinctions, avoid double counting, and retain only features evaluators can identify and justify from declared sources.Legal force and substantive detail remain distinct because they can differ in practice.
  • 3.1 Framework Development and Design: The design addresses advanced-system risks, misuse, testing, intervention, infrastructure resilience, implementation, scope, individual rights, distributional effects, development policy, and responsible innovation.Safety is treated as a sustained concern without becoming the whole of national policy.
  • 3.1 Framework Development and Design: Ordered anchors distinguish absent, general, partial, clear, and fully specified provisions, while category and criterion results provide more diagnostic information than an overall mean.The overall number remains useful because it requires evaluators to state which evidence changes the judgment.
  • 3.1 Framework Development and Design: The unit of analysis is a versioned national policy portfolio containing adopted or in-force official instruments, implementation guidance, institutional mandates, and budget commitments within the declared scope.Every score is tied to sources and a portfolio cutoff, with later changes creating a new version while preserving the older result’s interpretability.
  • 3.1 Framework Development and Design: The protocol supports auditable comparison, category-level reporting, and rescoring when portfolios change, but it does not establish rubric reliability or validity across settings.The instrument still requires testing by multiple evaluators and assessment of its sensitivity to weighting.

3.2 Framework Structure

The framework organizes national AI policy evaluation into five categories and 25 criteria, with sub-criteria serving as the scoring units and indicators directing evidence searches. It assesses safety, implementation capacity, lifecycle and risk coverage, protections for affected people, and public benefit and responsible innovation without scoring outcomes themselves.

  • Framework hierarchy: Five categories contain 25 criteria, while sub-criteria provide the scoring units and indicators guide evidence searches without receiving separate scores.This hierarchy distinguishes broad policy objectives, governance dimensions, evaluative questions, and locatable portfolio provisions.
  • Safety and risk governance: Safety-risk criteria examine controls for high-consequence systems, including standards, testing, red teaming, auditing, interfacing, misuse, concentration, and infrastructure resilience.The rubric considers whether testing connects to accepted methods or responsible institutions and whether reviewers have sufficient independence and access.
  • Implementation capacity: Implementation-capacity criteria assess mandates, authority, staffing, expertise, funding, independence, coordination, clarity, measurability, feasibility, and incentive alignment.Measurability requires objectives, baselines, indicators, reporting periods, and assigned bodies, while feasibility includes administrative, technical, workforce, and regulatory burdens.
  • Lifecycle and risk coverage: Coverage criteria assess responsibility across the AI lifecycle and risks including technical failure, misuse, discrimination, privacy, security, labor, environmental cost, participation, review, and redress.Lifecycle stages run from research and design through data, training, evaluation, deployment, monitoring, incident response, modification, and decommissioning.
  • Protections for affected people: Protection criteria separately evaluate fairness, non-discrimination, mitigation, monitoring, transparency, explainability, accountability, oversight, liability, complaints, appeals, and meaningful redress.The framework distinguishes public system descriptions from information needed to understand a particular decision and separates notice from explanation.
  • Public benefit and responsible innovation: Public-benefit and innovation criteria assess intended beneficiaries, public need, access, evaluation, research, infrastructure, startups, technology transfer, competition, workforce, equity, and trust measures.The category scores documented commitments and implementation provisions, not later economic growth, public-confidence changes, or distributional outcomes.

3.3 Scoring Methodology and Aggregation.

The methodology freezes the policy portfolio and search protocol, then assigns reviewable anchor-based scores to sub-criteria using cited provisions and rationales. It aggregates scores with transparent defaults while distinguishing documented absence from missing evidence and preserving coverage limitations.

  • Scoring procedure: Evaluators assign one score per sub-criterion, recording the cited provision, rationale, and source-search notes rather than scoring indicators separately.The evidence trail makes judgment reviewable without eliminating evaluator judgment.
  • Scoring procedure: The protocol freezes the portfolio and search process, inventories official instruments, and checks surrounding text, amendments, replacements, and binding actors before selecting an anchor.Official consolidated text and primary sources are preferred, with search terms, searched documents, language procedures, and dates recorded.
  • Anchor selection: The evaluator selects the highest whole-anchor description supported by the evidence; anchor numbers enable aggregation but do not imply equal intervals.Indicators are evidence prompts rather than independent questions, and “comprehensive” refers to sub-criterion coverage and specification, not guaranteed implementation success.
  • Missing evidence: Not scorable items are excluded from parent means with renormalized weights, while the reported coverage rate distinguishes insufficient evidence from a supported finding of absence.A score of 1 is reserved for declared searches that support absence; an entirely unscorable parent produces no parent score.

4 Discussion and Future Work

The framework enables traceable comparison of national AI policy portfolios while preserving differences in legal status, institutional capacity, rights protection, and public benefit. Its document-based scores improve auditability but remain limited by interpretation, uneven public evidence, and the absence of outcome validation.

  • Contribution: The framework converts heterogeneous national policy instruments into a traceable portfolio comparison without treating different legal statuses or scopes as equivalent.It connects broad governance comparison with provision-level evidence while limiting its claim to policy content.
  • Contribution: Separate categories assess technical risk governance, institutional authority and resources, affected people’s protection and agency, and public benefit.This structure prevents development programs from silently compensating for weaknesses in rights protection or other categories.
  • Legal force in context: Legal status and force must be recorded separately from policy quality: binding EU and Korean instruments score 5, while Singapore’s nonbinding guidance scores 2 on sub-criterion 8.4.Guidance may describe operating practices more clearly, while statutes can create authority without specifying obligations or individual rights.
  • Using and interpreting the framework: Assessments define the portfolio and cutoff, document sources and translations, score each sub-criterion with cited evidence, and use independent evaluators before resolving disagreements.Equal weighting is a starting point, but narrower questions may justify different weights.
  • Limits of document review: The framework improves auditability but can underrate informal practices, effective unpublished coordination, soft law, and implementation evidence that remains private or unavailable.Interpretive disagreement, cross-jurisdictional legal meaning, changing instrument versions, and the time required for detailed application further constrain document review.
  • Future work: Future studies can test whether policy-design scores predict enforcement, compliance, institutional operation, complaints, adoption, and broader outcomes using longitudinal causal evidence.Potential evidence includes enforcement actions, compliance records, incident data, institutional budgets, public complaints, and adoption data.

5 Conclusion

The paper evaluates national AI governance as a versioned portfolio of laws, institutions, funding and guidance rather than through a single prominent instrument. Its rubric assesses documented risk coverage, protections, public benefit and implementation readiness while not establishing enforcement success or policy outcomes.

  • Conclusion: National AI governance combines binding rules, guidance, institutions and resources differently across jurisdictions, so prominent laws or strategies alone are insufficient for comparison.Binding provisions may depend on later guidance and institutional authority, while development strategies and voluntary frameworks can shape practice without creating enforceable rights.
  • Conclusion: The European Union and Republic of Korea center comprehensive AI statutes, whereas the United States, United Kingdom, Singapore and Japan distribute governance across agencies, regulators, guidance and coordination.India combines development policy and data protection with an emerging AI governance structure.
  • Conclusion: 25 criteria evaluate whether serious risks are governed, responses can be implemented, policy scope is adequate, and affected people and AI development receive separate consideration.Each sub-criterion requires a cited provision and ordered judgment, with category and criterion results showing how the final score was reached.
  • Conclusion: A strong score indicates documented design and implementation readiness, but does not establish enforcement, organizational compliance, or improved safety and individual rights.Innovation and public welfare require separate outcome evidence.

A Compute Infrastructure Review

The review inventories computing infrastructure across thirteen jurisdictions to inform purposive selection, while treating the inventory as contextual rather than part of the policy score. It records public scientific capacity, commercial cloud provision, and private AI infrastructure without equating these categories.

  • Scope and interpretation: The thirteen-jurisdiction inventory informed purposive selection but did not enter the policy score.The review explicitly limits the inventory to contextual use.
  • Singapore: Singapore combines public and commercial capacity, including a government commitment of S$270 million to expand the National Supercomputing Centre and related training.Commercial providers operate private cloud and accelerator capacity, while AWS, Microsoft Azure, and Google Cloud operate Singapore cloud regions.
  • Japan: Japan’s ABCI 3.0 became publicly available in January 2025 with 6,128 NVIDIA H200 GPUs and a stated peak of 6.22 EFLOPS at FP16.Domestic and global providers also operate Japanese cloud regions, but the comparison does not equate commercial regions with scientific supercomputers.
  • Comparative infrastructure: The inventory documents mixed infrastructure models, including India’s public PARAM and National Supercomputing Mission capacity alongside commercial cloud regions.This illustrates the review’s inclusion of both public high-performance computing and commercial provision.

B Cross-National Policy Overview

The overview reports documented policy instruments and their stated implementation status through August 15, 2026, without inferring causal effects. It illustrates how national AI portfolios combine binding rules, voluntary guidance, executive actions, research programs, and sectoral measures with different implementation limits.

  • Scope and evidentiary limits: The tables report instrument status through August 15, 2026 and describe cited sources without inferring causal effects.This establishes the overview’s evidentiary and temporal cutoff.
  • United States: Executive Order 14110 was revoked on January 20, 2025 and no longer forms the operative federal executive-order framework.The order had directed action on safety, security, privacy, civil rights, research, workforce issues, reporting, evaluations, standards, and guidance.
  • United States: The United States portfolio at the cutoff comprises Executive Order 14179, the America’s AI Action Plan, federal programs, guidance, and standards work.The portfolio reorients federal policy toward US AI leadership, infrastructure, innovation, and international diplomacy while directing revision of the NIST AI RMF and renaming the AI Safety Institute as CAISI.
  • United States: The AI Bill of Rights Blueprint and NIST AI RMF provide nonbinding or voluntary governance guidance rather than independent approval or penalty mechanisms.The blueprint addresses safe systems, discrimination, privacy, notice, explanation, and human alternatives; the NIST AI RMF uses Govern, Map, Measure, and Manage functions across the lifecycle.
  • Singapore: Singapore combines binding, technology-neutral data-protection duties with nonbinding AI governance and sectoral guidance.The PDPA establishes applicable duties concerning consent, purpose, notification, access, correction, protection, retention, and transfer, while healthcare and financial guidance address validation, monitoring, fairness, accountability, explainability, and patient-centered design.

C Existing Policy Evaluation Framework Review

The appendix presents a selected methodological comparison of frameworks that directly informed the design synthesis, rather than a systematic literature review.

  • Scope of the review: The review covers frameworks contributing directly to the design synthesis and is explicitly limited to a selected methodological comparison.It does not claim to be a systematic literature review.

D Complete National AI Policy Evaluation Framework

The framework evaluates national AI policy portfolios through equally weighted categories, criteria, and sub-criteria, using ordered indicators to assess governance, implementation readiness, scope, rights, and innovation. It covers institutional capacity, compliance and enforcement, measurability, lifecycle coverage, risk range, stakeholder inclusion, adaptability, international harmonization, and user protections.

  • Weighting: Each category receives 20%, with criteria equally weighted within categories and sub-criteria equally weighted within criteria.These weights implement the default rule in Section 3.3.
  • Effectiveness and Feasibility: Institutional capacity is assessed through oversight mandates, operational independence, legal authority, funding, staffing, and technical, legal, and ethical expertise.The indicators distinguish absent, vague, limited, clear, and substantively guaranteed institutional arrangements.
  • Effectiveness and Feasibility: The framework evaluates whether policies specify compliance monitoring, enforcement mechanisms and penalties, legal force, data collection and reporting, and performance metrics.Ordered anchors range from no provisions to comprehensive mechanisms, requirements, or metrics aligned with policy objectives.
  • Comprehensiveness and Scope: Scope criteria examine coverage across the AI lifecycle, the range of technical, ethical, societal, and governance risks, stakeholder inclusion, adaptability, and international harmonization.The lifecycle extends from design and data collection through development, deployment, monitoring, and decommissioning.
  • User Rights, Protection and Agency: User-rights criteria cover fairness and non-discrimination, transparency and explainability, and accountability, while a separate category addresses socioeconomic impact and innovation.The framework also includes international cooperation through information sharing, joint research, and coordinated policy development.
Loading 2608.15417v1…