Source-linked AI summary
Law of Large Numbers: Accuracy as Statistical Measure for AI Compliance and Competition
Rabanus Derr, Alina Wernick, Robert C. Williamson
TL;DR
The paper asks whether machine learning and legal communities mean the same thing by accuracy, a term central to both machine learning and EU AI Act compliance. It analyzes accuracy as a boundary object across the two communities and finds that their differing expectations cannot be reduced to one universal notion. The paper uses these frictions to identify translation gaps and practical needs at the technology-law boundary.
Problem
Accuracy is central to machine learning and EU AI Act compliance, but the Act does not define it and neither community supplies one universal notion.
Method
The paper uses accuracy as a case study and boundary object to compare legal and technical sense-making, expectations, uses, and aims.
Results
Legal L-accuracy frames accuracy as normative performance compliance tied to intended purpose, while technical E-accuracy frames it as statistical, formal, empirical model performance with limited validity.
Takeaways & Limitations
The resulting frictions reveal translation gaps and support recommendations for both legal and technical communities rather than a single definition of accuracy.
Takeaways & Limitations
Accuracy measurements have limited validity when data, metrics, constructs, or interventions mismatch the deployment context, and aggregate measures poorly justify individual accuracy.
Abstract
from arXiv · showhide
The machine learning community progresses (in part) by improving the "accuracy" of its systems. The EU AI Act explicitly refers to "accuracy" as part of its compliance measures for high-risk AI systems. Are we talking about the same thing? This work presents "accuracy" as a case-study for differing requirements of social worlds, the technological machine learning community and the legal community. While competition on accuracy contributes to technological development, machine learning scholars simultaneously recognize accuracy's shortcomings regarding the usefulness and effectiveness of machine learning systems. The legal counterpart embraces the vagueness of "accuracy," leaving interpretative flexibility for technological and societal changes. At the same time, accuracy is a core element of compliance within the EU AI Act. We elaborate on five main tensions, (a) nature of accuracy, (b) notion of performance, (c) scope of validity, (d) ends, and (e) statisticalness, to show that the two communities project disparate, and sometimes contradictory, expectations on accuracy. Both legal and technical communities lack precise understanding of "accuracy" beyond the contextual boundaries of their community. The resulting frictions, \eg, based on the empirical or normative understanding of accuracy, are symptoms of an unresolved (and unresolvable) debate on what accuracy is. We constructively use the frictions to recommend baselines and interventional studies in standardization, and demand for tools to extend the validity of accuracy measurements.
1 How to Read This Paper?
The paper presents accuracy through separate legal and engineering perspectives, showing that the term carries different expectations across machine learning and law. It uses these differences to examine translation gaps and frictions at their intersection.
- The paper is designed for separate legal and engineering reading paths, with complementary perspectives on accuracy.The main sections use parallel L and E threads, while supplementary material provides their detailed arguments.
- Accuracy is central to both machine learning development and EU AI Act compliance, but remains undefined in the Act.In machine learning, accuracy evaluates model outputs; in law, it is tied to performance and intended purpose.
- The paper uses accuracy as a case study of differing expectations, aims, and uses where machine learning meets society.It focuses on what accuracy does in technical and legal domains rather than seeking one isolated definition.
- The authors argue that neither technical nor legal communities can provide a single universal notion of accuracy.They instead emphasize constructive friction, translation gaps, misunderstandings, and recommendations for both communities.
- Machine learning treats accuracy as a technical quality criterion, while law treats it as a quality certification connected to intended use and societal effects.The legal framing includes transferability, robustness, harms, and effects on individuals.
2 Accuracy as a Boundary Object
The paper treats accuracy as a boundary object whose meaning is shaped by distinct legal and technical practices. Its interpretative flexibility enables interaction without consensus but also produces friction and translation problems.
- Accuracy functions as a boundary object shared by lawyers and engineers, each of whom interprets it according to distinct informational and epistemic requirements.Boundary objects permit different groups to work together without reaching consensus.
- The concept of accuracy has developed through changing expectations across machine learning, ethics guidance, and EU legislation.The paper traces a path from classical classification accuracy to a broader, undefined legislative requirement.
- Friction between L-accuracy and E-accuracy drives mutual interaction between their sense-making loops.Figure 2 marks this interaction as a transition between the legal and engineering conceptualizations.
- Undefined technical terms in law create interpretive uncertainty because they acquire multiple technical meanings and become subject to legal interpretation.Providers and legal actors may therefore participate in giving accuracy its practical meaning.
- Legal and technical communities reproduce different forms of sense-making: formal technical language in machine learning and normative, teleological consistency in law.Each discipline can treat its own understanding as the right or true one.
3 L What is L-Accuracy? AI Lawyer’s View on Accuracy
L-accuracy is a legal compliance concept tied to the intended purpose of high-risk AI systems. It is normative because accuracy is used to mitigate risks to health, safety, and fundamental rights.
- The EU AI Act requires high-risk AI systems to achieve an appropriate level of accuracy, although it does not explicitly define the term.Accuracy is also referenced in documentation and other compliance requirements.
- L-accuracy evaluates performance with respect to an AI system’s intended purpose.Examples include predicting school drop-out risk and allocating support based on that risk.
- The Act’s lifecycle-oriented risk mitigation framing makes L-accuracy inherently normative.The legal approach emphasizes transparent measurement and the validity of the arguments supporting it.
- Accuracy in the EU AI Act can protect subgroups and individuals by addressing risks to fundamental rights.The Act explicitly mentions accuracy in relation to subgroups and individuals.
What is E-Accuracy? ML Engineer’s View on Accuracy
E-accuracy is a formal, empirical, and statistical measure of machine learning output performance. Although it supports model comparison and optimization, its validity is limited when data, metrics, constructs, or interventions do not match the deployment context.
- Machine learning systems are statistical technologies that generate predictions, recommendations, or other artifacts from aggregates of data.Their operation is therefore oriented toward aggregate rather than individual-level claims.
- E-accuracy formally compares desired and generated outputs using a metric over an aggregate of instances.Classical classification accuracy is the match-rate between predicted and actual classes.
- Accuracy serves as an abstract performance objective and defines model performance within a limited scope.Its limits become more visible near deployment when validity mismatches arise.
- Benchmarking measures optimization progress by ranking models through standardized accuracy comparisons.These comparisons primarily support relative competition between models.
- Individual accuracy is poorly justified because statistical systems operate on aggregates rather than single unobserved instances.The paper describes individual accuracy as a hardly justified and potentially oxymoronic concept.
4 From Clashes to Recommendations
The paper argues that translating machine-learning accuracy into EU AI Act compliance exposes persistent tensions over meaning, performance, validity, purpose, and statistical measurement. It recommends decision-accuracy and interventional studies for interventions, baseline comparisons for validity, and context-sensitive safeguards against metric gaming.
- Nature of Accuracy: E-accuracy is formal, empirical, and computable, whereas L-accuracy is teleological, legally elusive, and responsive to changing social aims.The paper argues that these concepts are not fully translatable because one is momentary and measurable while the other must remain adaptable.
- Scope of Validity: Metric, dataset, and target-construct mismatches limit validity, so reports should justify metric semantics, deployment representativeness, lifecycle scope, and the target’s relation to intended purpose.These mismatches are presented as translation difficulties between technical accuracy and legally relevant performance.
- Notion of Performance: Accuracy should address intended purpose: interventional systems require accurate treatment assignment with maximal causal effect, not merely accurate outcome prediction.The proposed interpretation requires specifying the causal effect, using randomized evaluation data, and modeling effects.
- Scope of Validity: Comparative accuracy measurements transfer more reliably across test scenarios than direct scores, motivating baseline and comparison-system reporting in EU AI Act compliance.Baselines also reveal whether an AI system improves on existing methods, including simple statistical approaches or human performance.
- Scope of Validity: Comparison systems must be chosen and justified contextually because humans, annotators, and simple methods may differ in accuracy or relevance across purposes.The paper rejects a naive universal comparison system.
- Statisticalness and Ends: Statistical AI systems rely heavily on statistical accuracy because they lack diverse non-statistical certifications, while compliance metrics can be strategically made irrelevant and easily satisfied.The paper calls for remedies against “accuracy-hacking” and warns that statistical performance may diverge from broader societal consensus.
5 Lessons Learned – A Satellite Perspective
The paper treats accuracy as a boundary concept whose competing technical and legal interpretations cannot fully converge, yet their friction can refine shared understanding.
- The authors shift from asking what accuracy is to asking what accuracy does in different social and technical contexts.
- Machine learning formalization and legal attention to competing social interests generate non-overlapping interpretations of technical terms.
- Statistical machine learning and individual-focused legal perspectives produce persistent disagreement about how accuracy should be understood.
- Machine learning’s drive toward new methods and solutionism conflicts with law’s ideal of consistency when technical terms cross social boundaries.
- These clashes are productive because they refine concepts while warning against treating abstract nouns as having one universally stable meaning.
- Future work should examine how power structures shape relevant terms and interpretations, including accuracy in AI Act standard setting.
L What is L-Accuracy? AI Lawyer’s View on Accuracy
The EU AI Act makes accuracy a central requirement for high-risk AI systems while leaving its meaning undefined within a broader risk-based regulatory framework.
- The Act regulates AI through product-safety principles intended to protect health, safety, and fundamental rights while supporting innovation.
- Its risk-based structure prohibits unacceptable-risk systems and subjects high-risk applications to detailed compliance requirements.
- The accuracy obligation applies specifically to high-risk AI systems, as assumed throughout the discussion.
- The Act distinguishes AI systems from machine learning models because systems may include hardware and input or output processing beyond the model.
L.2 L-accuracy is a legal compliance measure
L-accuracy is embedded in high-risk AI compliance as a lifecycle performance requirement supported by documentation, measurement, and related robustness obligations.
- Accuracy, robustness, and cybersecurity are among the key compliance obligations for providers of high-risk AI systems.
- Providers must demonstrate compliance before market placement or service, with breaches punishable by substantial financial penalties.
- Article 15 requires high-risk systems to achieve appropriate accuracy and perform consistently throughout their lifecycle.
- Accuracy obligations are presented with robustness, which concerns resilience to errors, faults, inconsistencies, and operating-environment interactions.
- Accuracy-related requirements are distributed across the Act, including documentation and information duties for providers and deployers.
- The Act uses accuracy and performance in overlapping ways, including reporting accuracy degrees for specific persons or groups.
L.3 L-accuracy certifies the fitness-to-purpose of machine learning systems
The paper interprets L-accuracy as fitness to an AI system’s intended purpose, linking appropriate accuracy to context, objectives, and potential intervention.
- Accuracy is an ex ante compliance criterion considered during system design and before market placement or service.
- Intended purpose is the provider-specified use, including the context and conditions described in instructions, materials, and technical documentation.
- A felicity-oriented interpretation asks whether outputs meet their intended purpose, with accuracy measuring how well that purpose is achieved.
- The paper distinguishes explicit objectives, encoded by developers, from implicit objectives inferred from system behavior, training data, or environmental interaction.
- The appropriate accuracy level depends on the system’s intended purpose and use context rather than being interpretively independent of them.
- Because AI outputs can influence environments and generate risks, intended purposes may include direct or indirect interventions.
L.4 L-accuracy is inherently normative and used for risk mitigation
Under the EU AI Act, accuracy is not merely a technical metric but a normatively interpreted criterion tied to safety, transparency, fairness, and intended purpose. Compliance therefore requires providers to establish, justify, communicate, and monitor accuracy across the system lifecycle.
- The AI Act interprets accuracy through its broader objectives and normative concepts rather than treating it as a purely technical criterion.
- Accuracy contributes to mitigating risks to health, safety, and fundamental rights by supporting consistent system performance throughout the lifecycle.
- Accuracy extends beyond algorithms to input-data quality, design logic, and the operational conditions in which the system is developed.
- Accuracy compliance spans three lifecycle functions: establishing performance through testing, communicating expectations, and enabling performance monitoring.
- The Act requires transparency about accuracy levels, relevant metrics, validation and testing procedures, and the data used.
- Providers must justify why their chosen accuracy metrics and evaluation procedures are legitimate and well-founded, not merely disclose them.
L.5 L-accuracy exists to protect subgroups and individuals
The EU AI Act frames accuracy in relation to intended users, including specific individuals and groups, reflecting its rights-oriented structure. Yet ex ante accuracy information for individuals or subgroups can be impossible or technically infeasible to provide.
- The Act requires information about performance and accuracy limitations for specific persons or groups on whom the system is intended to be used.
- Accuracy is linked to the legal system’s focus on protecting individual rights, while AI systems can produce scalable effects on groups beyond individual awareness.
- The rationale for referring to specific persons is unclear because providers must document accuracy before systems reach the market or enter service.
- Ex ante accuracy information for individuals or subgroups may be impossible or technically infeasible, although documentation duties lack an explicit feasibility qualifier.
- The Act treats accuracy as a quality of performance that supports risk mitigation and requires consistent performance over the system’s lifetime.
E.2 E-accuracy is a formal and empirical measure
In machine learning, E-accuracy is a formal, empirical comparison between generated and desired outputs over data instances. Its metric and test data encode context-dependent utilities, while the data underlying the measurement are often undervalued.
- Classical accuracy compares binary predictions with desired labels, while broader forms compare arbitrary generated outputs with desired counterparts using a utility function.
- For semantically rich outputs such as text, the desired output may concern broad properties and remain ambiguous between equivalent expressions.
- Accuracy definitions are output-focused and statistical because they aggregate comparisons across sets of instances rather than inspect internal model workings.
- Different metrics express different utilities, so the appropriate accuracy measure depends on the scenario and intended use of the model’s outputs.
- E-accuracy requires model outputs, desired outputs, a formal comparison function, and a dataset of instances.
- Machine learning research often treats test data as given truth, although benchmark datasets may rely on convenience samples, artificial generation, or imperfect annotation.
- Data-production histories and data-selection choices are frequently neglected, despite their methodological importance for accuracy measurement.
E.3 E-accuracy defines the performance of machine learning models
Machine learning commonly uses accuracy to define and compare model performance, especially against state-of-the-art benchmarks. However, accuracy has limited validity beyond its measurement setting and may fail to represent deployment performance or intervention success.
- Accuracy is frequently used in experiments to demonstrate a method’s goodness by comparing it with the state of the art.
- Accuracy is increasingly used to support decisions about which methods to develop, deploy, or use as tasks become more concrete.
- Accuracy does not necessarily translate into successful application because its measurement may lack validity beyond the particular test scenario.
- Metric mismatch occurs when the selected accuracy measure fails to reflect the use-case utility or the unequal costs of errors.
- Dataset mismatch occurs when test distributions differ from deployment conditions, making measured accuracy claims non-transferable.
- Target-construct mismatch arises when the predicted target differs from the abstract construct the system is intended to assess.
- When system outputs alter the data or outcomes they predict, interventional success cannot generally be measured by accuracy alone.
E.4 E-accuracy is inherently statistical and used for ranking models
E-accuracy is a statistical, comparative measure used to rank machine-learning models, but absolute scores have limited validity and interpretability beyond their measurement context.
- E.4 E-accuracy is inherently statistical and used for ranking models: Accuracy can evaluate outputs from non-statistical systems, including physical simulators and human clinicians, although statistical systems are favored by statistical metrics.The metric itself does not prohibit non-statistical systems from being measured.
- E.4 E-accuracy is inherently statistical and used for ranking models: Benchmarking tracks machine-learning progress through repeated comparisons of models on shared training and test datasets.ImageNet exemplifies how competition over categorical classification accuracy helped trigger deep-learning research.
- E.4 E-accuracy is inherently statistical and used for ranking models: E-accuracy measures relative performance and is used to select better models in development or deployment.Its exact value is unnecessary when the valid ranking of models is sufficient.
- E.4 E-accuracy is inherently statistical and used for ranking models: Relative model rankings are more stable than absolute accuracy scores across experimental setups.This reduced validity claim supports competition even when test-set scores are not indicative beyond the original setup.
- E.4 E-accuracy is inherently statistical and used for ranking models: Accuracy scores may fail to generalize to unseen test data and can be influenced by small random fluctuations.Test-set reuse and dataset mismatch further undermine internal validity.
- E.4 E-accuracy is inherently statistical and used for ranking models: Absolute E-accuracy scores cannot define performance reliably without baseline comparisons.Their usefulness is therefore primarily comparative rather than absolute.
E.5 E-accuracy is poorly suited to subgroups and individuals
E-accuracy can be measured for individual instances, but aggregate statistical accuracy does not justify particular assignments and permits multiple equally accurate models.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Individual accuracy is measurable on a singleton instance, but justifying accuracy on an unseen individual instance is problematic.The difficulty concerns the argument supporting the assignment, not the formal measurement itself.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Several accurate models may assign different predictions to the same individuals because of modeling choices and training-data choices.This phenomenon is also called predictive multiplicity, predictive inconsistency, or individual arbitrariness.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Individual arbitrariness is not inherently problematic because it can help avoid algorithmic monoculture or accommodate fairness conditions.Its central harm is instead that choosing among equally accurate models can lack a principled rationale.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Aggregate accuracy naturally permits model multiplicity because models differing on one individual can remain indistinguishable by aggregate quality measures.Removing this contradiction would require removing the statistical aspect of accuracy.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Statistical statements about individuals depend on selecting a reference class from which aggregate frequencies or accuracy are derived.The reference-class choice is therefore part of the justification problem for individualized claims.
- E.5 E-accuracy is poorly suited to subgroups and individuals: Individual assignments lack justification when they are rationalized only through aggregate accuracy.The resulting circularity makes the individual accuracy claim empty.