Source-linked AI summary
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Max Lamparth, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
TL;DR
AI evaluation reporting lacks shared conventions, limiting comparison and interpretation across sources. Evaluation Cards unifies reporting infrastructure and reveals widespread gaps in reproducibility, documentation completeness, and multi-source reporting.
Problem
AI evaluation reporting lacks shared conventions and compatible formats, limiting readers’ ability to interpret and compare results across sources.
Method
Evaluation Cards develops a literature- and practitioner-grounded reporting framework and a traceable hierarchy for evaluation evidence.
Results
98.2% of model-benchmark pairs are reported by only one party, while median benchmark documentation completeness is 10.7%.
Takeaways & Limitations
Evaluation Cards provides an interpretable reporting layer for surfacing assumptions, provenance, reproducibility, and comparison context in evaluation evidence.
Takeaways & Limitations
Reporting completeness does not assess comprehensiveness, and the corpus excludes results not ingested into EVALUATION CARDS.
Abstract
from arXiv · showhide
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.
1 Introduction
The paper presents EVALUATION CARDS as a live reporting layer that unifies evaluation infrastructure and makes evaluation evidence more interpretable across sources. It derives a reporting framework from literature and stakeholder needs, structures evidence hierarchically, and audits reporting gaps at scale.
- Motivation: Existing AI evaluation reporting uses incompatible formats, omits interpretive fields, and lacks a standard basis for cross-source comparison.These gaps affect leaderboards, model cards, benchmark papers, and company blogs, while inconsistent self-reporting makes evaluation omissions difficult to define.
- Motivation: Evaluation fragmentation prevents readers from turning reported scores into actionable claims for deployment, regulation, or scientific assessment.Interpretation depends on what is measured, the assumptions used, and how results are presented to the reader.
- Contributions: EVALUATION CARDS unifies existing evaluation infrastructure and surfaces interpretive signals through a live interactive reporting layer.The paper frames the system as an operational layer rather than an isolated reporting artifact.
- Contributions: A reporting framework derived from a structured review of 52 papers and interviews with 12 stakeholders specifies what should accompany evaluations for reproduction, context, and comparison.The stakeholders represented technical, developer, and policy roles.
- Empirical audit: 0.0% versus 16.6% field population on paired first- and third-party reports shows the widest developer self-reporting gap.The same audit reports 10.7% median per-benchmark documentation completeness, 98.2% single-party reporting of model–benchmark pairs, and 51.9% cross-party divergence above the 5% threshold in multi-organization metric groups.
2 Related Work
Prior work spans technical reporting documentation, evaluation infrastructure and data schemes, and systematic reviews of evaluation practice. However, no prior artifact combines these components with reader-specific rendering and continuous monitoring of public evaluation reporting.
- Prior work covers technical reporting documentation, evaluation infrastructure and data schemes, and systematic reviews of evaluation practice.These areas address parts of the evaluation lifecycle and evaluation practice.
- No prior artifact joins a reporting framework, evaluation run data, and benchmark metadata while differentiating rendering by reader type.
- No prior artifact provides a continuous monitoring instrument for the state of public evaluation reporting.
3 The EVALUATION CARDS Framework
The EVALUATION CARDS framework is a permissive, hierarchical reporting format that organizes evaluation evidence across benchmark metadata, run data, and model metadata. It was synthesized from a systematic review and refined through stakeholder interviews, with platform operationalization focused on artifact-side reporting.
- Framework structure: EVALUATION CARDS are a permissive reporting standard: partial population can still produce a useful card, rather than requiring every field.The framework is explicitly designed as an alternative to a prescriptive checklist.
- Framework structure: EVALUATION CARDS organize evaluation evidence in a hierarchical format spanning benchmark metadata, evaluation run data, and model metadata.The format is introduced in Section 3.1, with its hierarchical structure described in Section 3.2.
- Framework development: 52 papers informed the framework, after a systematic review identified requirements for comprehensive evaluation reporting.The review covered AI evaluation practice published between 2020 and 2025.
- Operational scope: The platform operationalizes artifact-side reporting while linking process-side decisions to complementary mechanisms when no published trace exists.This scope avoids requiring fields that evaluators cannot substantiate from published evidence.
- Framework development: 12 stakeholders across 9 organizations and 3 geographic regions helped refine and evaluate the EVALUATION CARDS platform through semi-structured interviews.Participants included technical evaluators, AI engineers, and policy actors.
1. Design
The design specifies the evaluation’s goals, constructs, validity, task development, and ethical context, while defining protocols for preregistration, scoring, validation, splits, and pre-reporting. It also addresses pilots, baselines, contamination, gaming, and participant awareness.
- The design defines evaluation goals, tested constructs, context, preregistration, validity, task types, item development, and human-subjects ethics.
- It establishes protocols for pre-runs, scoring, validation, data splits, holdouts, pilots, and baselines.
- The design also considers contamination, gaming, participant awareness, and pre-reporting procedures.
2. Before execution · 3. Execution
The execution section centers on reproducibility-oriented run logging, documenting mitigations and adaptations, and analyzing differences between runs.
- 3. Execution: Run logging is identified as a core component of execution reporting.The passage specifically connects run logging with reproducibility capture.
- 3. Execution: Execution reporting includes explicit reproducibility capture.
- 3. Execution: Mitigations are recorded as part of the execution process.
- 3. Execution: Adaptations are included alongside execution mitigations.
- 3. Execution: Analysis is treated as a distinct execution-reporting concern.
- 3. Execution: Differences between runs are included in execution analysis.
4. Lifecycle
The lifecycle schema organizes evaluation reporting around stages where reporting choices are made, while the rollout hierarchy represents each score through its internal benchmark structure. This path-based representation supports scoped integrity signals, drill-down from aggregates, and comparability across multi-source results.
- Lifecycle stages: Each top-level category corresponds to a stage of the evaluation lifecycle at which reporting choices are made.The categories cover data availability and access, later use and maintenance, reporting and publication, process reporting, transparency, and replication and reproducibility.
- Rollout hierarchy: The five-level rollout hierarchy replaces flat model–benchmark–score records with paths through Family, Composite, Benchmark, and lower-level evaluation structure.It reflects internal structure such as MATH subject splits, SWE-bench language and setup variants, and composite benchmark aggregates.
- Path-based interpretation: Integrity signals attach to specific model–metric paths rather than benchmark names, allowing warnings to distinguish subtasks, metrics, and variants.Scores resolve to paths through the hierarchy, preventing conflation when the same model or benchmark label differs in subtask or metric.
- Path-based interpretation: The hierarchy enables drill-down from aggregate family-level claims to the specific metric supporting each claim and exposes uneven evidentiary support.Readers can see where a claim is well-evidenced and where it depends on a single reported number.
- Path-based interpretation: Distinct paths make multi-source results comparable within EVALUATION CARDS by preventing conflation across differing subtasks or metrics.This addresses a known gap in existing sources that defer deduplication to the analysis layer.
4 EVALUATION CARDS: Pipeline & Interpretive Layer
EVALUATION CARDS builds unified records from benchmark, evaluation-run, and model metadata, standardizing identifiers before computing four interpretive signals. It renders those signals through summary and research modes that surface omissions, methodological detail, provenance, risk, and score comparability for different audiences.
- 4.1 Pipeline: EVALUATION CARDS combines Auto-BenchmarkCards, EEE, and community model catalogs into a unified schema spanning benchmark, execution, reporting, and model metadata.Auto-BenchmarkCards supplies benchmark metadata, EEE supplies aggregate and instance-level run data, and catalogs enrich models with fields such as release date, parameter count, and weight accessibility.
- 4.1 Pipeline: A standardization layer maps inconsistent model and benchmark names across sources to stable identifiers.Examples include model variants such as gpt-4 and gpt-4-0613 and benchmark references expressed through paper names, leaderboard slugs, or version-qualified identifiers.
- 4.2 Interpretive Layer: Four interpretive signals answer whether readers have enough information to contextualize and trust an evaluation result for a decision.The signals address reproducibility, reporting completeness, provenance and risk, and score comparability; missing fields can trigger several signals and are explicitly shown to readers.
- 4.2 Interpretive Layer: Summary and research modes operate on identical records but differ in which fields they surface, compress, or reframe for distinct audiences.Research mode foregrounds methodology, configuration, missing reproducibility fields, and setup differences, while summary mode uses plain-language accountability and narrative caveats.
- 4.2 Interpretive Layer: Participant feedback was generally positive, with interviewees describing EVALUATION CARDS as better than other review methods and as saving substantial review time.Further systematic usability evaluation is planned as post-deployment work.
5 Empirical findings from the EVALUATION CARDS corpus
The EVALUATION CARDS corpus provides a structured view of AI evaluation reporting across thousands of models, benchmarks, and results. It reveals pervasive gaps in result reproducibility and benchmark documentation, while multi-source reporting is uncommon and often divergent.
- Corpus construction and coverage: The corpus contains 5,816 models, 635 single-benchmarks, and 101,955 reported results from 30 organizations and two source types.The benchmarks comprise 62 families and 10 composites; 211 carry matched Auto-BenchmarkCards records.
- Finding 1: result-level reproducibility is the dominant reporting gap: 96.5% of 50,461 model-benchmark-metric triples lack at least one field in the minimal reproducibility sub-schema.max_tokens is absent from 95.6% of triples and temperature from 93.9%; eval_plan and eval_limits are missing from 100% of agentic-benchmark triples.
- Finding 2: benchmark-level documentation is thin: Median benchmark-level completeness is 10.7% across 635 benchmarks, with documentation and provenance fields often unpopulated.eee.metric_config.score_type and eee.score have 100.0% population, while evalcards.preregistration_url and evalcards.lifecycle_status have 0.0%.
- Finding 3: multi-source reporting is rare, and frequently divergent when it occurs: 98.2% of 49,865 model-benchmark pairs are reported by only one party, while 7.2% of multi-party pairs exceed the 5% score-divergence threshold.Among 181 multi-organization metric groups, 94 (51.9%) exceed the threshold.
- Implications for evaluation reporting: These patterns leave readers without the re-execution inputs and contextual documentation needed to interpret evaluation scores for downstream decisions.The reproducibility gap is widest in developer self-reporting, while independent reporting is least common in agentic and general categories.
6 Community adaptability and adoption … A.3 Model-level walkthrough: GPT-5
EVALUATION CARDS is designed as an open, extensible infrastructure layer that composes existing evaluation efforts for diverse audiences. A GPT-5 walkthrough shows how its signals expose reproducibility gaps, documentation differences, provenance patterns, and score divergence.
- 6 Community adaptability and adoption: EVALUATION CARDS is participatory, openly governed, and extensible, with openly released code that supports self-hosted deployments.The initiative is intended to evolve with the AI evaluations community rather than remain a one-off research artifact.
- 6 Community adaptability and adoption: The system addresses fragmented evaluation reporting by composing benchmark metadata, run data, and reporting practice into one interpretive layer.It surfaces missing information and renders results in forms different audiences can act on.
- A EVALUATION CARDS in Practice: The practical examples cover rollout hierarchies, model and benchmark walkthroughs, and corpus-level aggregations rendered from the interface on June 4, 2026.The examples include sections A.1, A.3, A.4, and A.5.
- A.1 Hierarchy: Figure 4 depicts the hierarchy for the composite Artificial Analysis, which comprises 15 benchmarks.The figure is presented as an example of the five-level rollout hierarchy.
- A.2 Corpus: 5,816 models, 635 single-benchmarks, 62 families, 10 composites, and 101,955 reported results comprise the corpus as of June 4, 2026.The corpus also includes 211 Auto-BenchmarkCards contributed by 30 organizations, with results ingested through converters, leaderboard scrapes, and community contribution.
- A.3 Model-level walkthrough: GPT-5: 202 of 213 GPT-5 results (95%) lack at least one temperature or max_tokens field, while 64 benchmarks have matching Auto-BenchmarkCards scoring 93% completeness and the remainder score 11%.Only 11 results have the reproducibility fields completed; the most populated benchmarks include livecodebench-pro and global-mmlu-lite.
- A.3 Model-level walkthrough: GPT-5: 27 GPT-5 results are first-party (13%) and 186 are third-party or independent (87%), while MATH-500 scores range from 84.7% to 98.9% across 3 organizations.The interface also identifies first-party-only benchmarks such as FrontierMath and HumanEval and exposes divergent scores that individual leaderboards may omit.
A.4 Evaluation-level walkthrough: MMLU-Pro · A.5 Corpus-level analysis
The MMLU-Pro walkthrough shows substantial missing reproducibility metadata and score divergence across reporting sources, while corpus-level analysis finds these documentation and provenance gaps are systematic. Across the corpus, reproducibility fields are usually incomplete, schema population varies sharply, and cross-party reports frequently diverge.
- A.4 Evaluation-level walkthrough: MMLU-Pro: MMLU-Pro covers 401 documented models from 8 organizations and 5,079 reported results, aggregating across sources rather than presenting a single-model view.The walkthrough is designed to expose benchmark-level completeness and evaluation-level comparability across reporting organizations.
- A.4 Evaluation-level walkthrough: MMLU-Pro: 98% of MMLU-Pro results, or 4,975 of 5,079, have at least one missing field in the minimal reproducibility sub-schema.Only 104 results report the minimum reproducibility sub-schema.
- A.4 Evaluation-level walkthrough: MMLU-Pro: 95.5% of MMLU-Pro models have third-party-only results, 1.5% first-party-only results, and 3.0% both; 41.4% have reports from at least two organizations.The source distribution spans 383 third-party-only, 6 first-party-only, and 12 mixed-source models, with 166 models receiving multi-source reporting.
- A.4 Evaluation-level walkthrough: MMLU-Pro: 11 MMLU-Pro cases show same-model score divergence above the comparability threshold under different setups, while 6 multi-source entries exceed the threshold across organizations.One example is Llama 3.2, reported at 20.9% by Hugging Face and 61.8% by Arcadia Impact.
- A.5 Corpus-level analysis: 48,698 of 50,461 corpus (model, benchmark, metric-path) triples, or 96.5%, have at least one missing minimal-reproducibility field.Missingness is concentrated in max_tokens at 95.6% and temperature at 93.9%; agentic reports additionally require eval_plan and eval_limits.
- A.5 Corpus-level analysis: Across 180 paired (model, benchmark) cases, first-party rows populate 0.0% of base reproducibility fields on average, compared with 16.6% for third-party rows.This comparison indicates lower per-field documentation in first-party rows for the same paired cases.
- A.5 Corpus-level analysis: Per-field schema population ranges from 100.0% for eee.metric_config.score_type and eee.score to 0.0% for evalcards.preregistration_url and evalcards.lifecycle_status.Median per-benchmark completeness is 10.7% across 635 benchmarks with warehouse completeness rows; raw score fields are more populated than documentation and comparison-context fields.
- A.5 Corpus-level analysis: 98.2% of 49,865 corpus (model, benchmark) pairs are reported by only one party; among multi-party pairs, 7.2% exceed the comparability threshold.At the scored multi-organization metric-group level, 94 of 181 groups, or 51.9%, have cross-party divergence flags.
B Limitations & Future Work … 1. When you look at evaluation results, what are you trying to decide?
EVALUATION CARDS inherits limitations from its sources, canonicalization, scope, LLM focus, and interpretation boundaries, while future work targets contamination reporting. The accompanying interviews and guide examine stakeholder roles, information needs, evaluation-review challenges, and reactions to the tool.
- B Limitations & Future Work: EVALUATION CARDS inherits source limitations because accurate but incomplete fields can pass validation, while EEE is not a complete census of public reporting.Reporting completeness therefore cannot guarantee comprehensive coverage of central information or all public evaluation results.
- B Limitations & Future Work: Completeness scores measure artifact-side documentation adequacy, not the thoroughness or rigor of the underlying evaluation.Process-side decisions such as pre-run protocol commitments and pilot execution are outside the system’s scope, creating a possible safetywashing risk.
- B Limitations & Future Work: EVALUATION CARDS does not judge whether reported scores satisfy safety, regulatory, or deployment-fitness thresholds, which depend on jurisdiction and use case.Its research and policy reader modes were informed by preliminary semi-structured interviews with ten practitioners, so additional reader types may emerge.
- B Limitations & Future Work: EVALUATION CARDS currently supports only LLM evaluation reporting, with broader AI-system and modality coverage remaining on the development roadmap.The authors identify expansion beyond LLMs as a priority rather than a current capability.
- B Limitations & Future Work: Future work will structure contamination-control information, including exposure-assessment methodology, detection mechanisms, and access controls for sensitive benchmark items.The current schema captures contamination only in free-text limitations fields, which do not contribute to reporting completeness.
- C Interview Methodology: 12 interviews with technical and policy stakeholders informed the study, using fluent interview languages and one or two researchers per interview.Participants were told their responses would support EVALUATION CARDS and appear in an academic paper, potentially with direct quotation.
- C.1 Interview Participants: Some interviewees represented both policy and technical perspectives, although P1, P3, and P4 were tagged primarily as policy stakeholders.The interviews nevertheless contained insights into both stakeholder groups’ needs when interpreting evaluation results.
- C.2 Interview Guide: The interview guide asked about stakeholder roles, decision information, missing metadata, downstream users, evaluation trustworthiness, comparison difficulties, and desired tool features.It also included tool testing, asking participants to rate fit from 1–10 and identify the most and least useful features and any ambiguities.
C.3 Limitations … D.3 Worked Examples
The paper notes geographic and recruitment limits in its interviews, then describes a normalization pipeline that reconciles evaluation and benchmark stores into comparable records and frontend views, with worked examples tracing this process end to end.
- C.3 Limitations: Ten interviews primarily represented North America, potentially omitting technical evaluator and policy perspectives from the Global South because recruitment used interviewers’ networks.The authors also acknowledge that critical viewpoints may have been missed.
- D.1 Inputs: The normalization layer reconciles independently maintained stores and naming conventions into a corpus where each result has stable model, benchmark, and metric identities.Where possible, descriptive metadata is also attached to explain what was measured.
- D.2 Pipeline: The pipeline decomposes nested records into atomic (model, benchmark leaf, metric) triples, canonicalizes identities, and joins rows to generate pre-materialized comparison views.These include per-model, per-benchmark, per-developer summaries and a model-by-benchmark matrix, enabling filter-and-project frontend queries instead of runtime cross-store joins.
- D.2.1 Data Cleaning: Deterministic cleaning lowercases identifiers, standardizes separators, separates families from versions, detects splits, and parses metrics while preserving parameters such as k in pass@k.For example, MMLU-Pro, mmlu pro, and mmlu/pro produce the same key.
- D.2.2 Entity Matcher: The entity matcher links models, benchmarks, metrics, harnesses, and organizations through canonical aliases and confidence-ordered exact, normalized, and narrow fuzzy matches.The registry is populated with seed entities and public metadata sources, while unresolved strings remain in downstream processing rather than being dropped.
- D.2.2 Entity Matcher: 98.3% model, 77.4% benchmark, and 86.7% metric resolver accuracy was measured by manually labeling 200 randomly sampled entities per type from the EEE corpus.The authors characterize these as in-domain results optimized for labeling coverage, and retain unresolved entities for downstream processing.
- D.3 Worked Examples: Two worked examples trace real upstream evaluation records through the benchmark-metadata join into the frontend representation.They illustrate the path from source records to what the platform serves.
D.3.1 Example 1: a single-benchmark, single-metric record … H.3 Agreement metrics
EvalCards normalizes heterogeneous evaluation records into canonical, composable artifacts while preserving source identity, benchmark hierarchy, metric specificity, and snapshot variants. Its interpretive methodology then computes reproducibility, completeness, provenance, comparability, aggregation, and agreement signals, supported by standardized infrastructure, reader profiles, and open governance.
- D.3.1 Example 1: a single-benchmark, single-metric record: A single-benchmark record is normalized into a canonical model–benchmark–metric row while preserving raw identifiers and enabling model, benchmark, and developer views.The example maps moonshotai/Kimi-K2-Thinking, GPQA Diamond, and Accuracy to canonical identifiers while retaining upstream strings and joining benchmark metadata.
- D.3.2 Example 2: a composite benchmark with sub-tasks and a roll-up: Six atomic results are extracted from the HELM Capabilities record, with five canonical leaf benchmarks and a Mean score aggregate excluded from corpus-wide leaderboards and averages.The decomposition preserves suite grouping, while the aggregate is tagged is_summary_score: true.
- E Compute Resources: Standard CPU infrastructure runs a daily hosted pipeline that processes upstream data into versioned Parquet files with JSON sidecars, with a full corpus rebuild completing in under 20 minutes as of May 7, 2026.The runner uses ubuntu-24.04 LTS, 4 vCPUs, 16 GiB RAM, and a 75-minute timeout.
- F User Personas; Persona 1: Technical Evaluator; Persona 2: Policy Actor; Persona 3: Model Developer: Reader modes target distinct needs: technical evaluators assess reproducibility and comparability, policy actors assess accountability and risks, and model developers use completeness and standardization as a self-audit checklist.The profiles emphasize missing generation configurations, reporting parties, benchmark risk categories, plain-language metric limitations, and consistency between aggregates and underlying scores.
- H.1.3 Provenance; H.1.4 Comparability: Provenance distinguishes first-party, third-party, and collaborative reporting, detects multi-party coverage, and propagates benchmark risk categories; comparability separately evaluates variant and cross-party divergence using θ = 0.05.The overall comparability signal is the maximum of variant and cross-party divergence, with differing setup fields rendered alongside flags.
- H.2 Aggregation Across Views; H.3 Agreement metrics: Signals aggregate across model, benchmark, and corpus views, while agreement analysis reports raw agreement, Cohen’s κ, Krippendorff’s α, exact set match, and mean Jaccard similarity after excluding blank ratings.Multi-label tags are expanded into binary item-by-tag decisions for pooled κ and α computation.
H.4 Benchmark Categorization … J.3.1 Selected Studies and Characteristics
The paper develops a taxonomy- and infrastructure-based approach to make AI evaluation reporting more interpretable, combining benchmark categorization with a systematic synthesis of evaluation guidance. Its review identifies diverse, overlapping recommendations and documents the scope, methods, synthesis process, and limitations of the selected literature.
- H.4 Benchmark Categorization: 518 benchmarks came from EEE’s flat evaluation-record export, while 37 additional benchmarks came from its family hierarchy.Fifteen near-duplicate pairs were collapsed through surface normalization, producing the deduplicated set of 635 benchmarks.
- H.4 Benchmark Categorization: 635 unique benchmarks were assigned one or more tags from an 18-category flat taxonomy derived from Ni et al.The categorization is a first-pass annotation intended to support corpus-level filtering in EVALUATION CARDS.
- H.4 Benchmark Categorization: The taxonomy annotations were generated with claude-haiku-4-5 from benchmark names, then manually inspected, corrected, and recategorized.The first batch of 50 was reviewed before the full run; afterward, three direct corrections and 14 recategorizations were applied, while four ambiguous entries were retained.
- I Related Work Extended; J.1 Background and Previous Work: Prior reporting artifacts and evaluation infrastructures address separate pipeline components, motivating a unified framework for consistent AI evaluation practice.The review contrasts Model Cards, Datasheets, Data Cards, BenchmarkCards, Audit Cards, Eval Factsheets, HELM, AILuminate, leaderboards, and Auto-BenchmarkCards.
- J Systematic Literature Review; J.1 Background and Previous Work; J.2 Methods; J.2.3 Item Extraction and Characterization; Extracted Item-Level Characterization: The systematic review synthesized recommendations without resolving disagreements, because evaluation guidance uses overlapping categories, differing vocabularies, and context-dependent goals.The extraction was deliberately broad, while later coding characterized items by workflow stage, artifact, objective, type, class, and supercategory.
- J.2.1 Search, Screening, and Inclusion; J.3.1 Selected Studies and Characteristics: The review used a preregistered step-down search across Google Scholar, arXiv, Elicit, expert suggestions, and citation chasing, yielding 748 sources before screening.Searches targeted AI/ML evaluation and audit guidance, standards, frameworks, and best-practices documents.
- J.2.2 Study-Level Characterization; J.2.4 Best Fit Framework; J.2.5 Synthesis & Groupings: The synthesis mapped evaluation frameworks into a shared structure and grouped extracted items primarily by workflow stage and objective, using a best-fit iterative approach.The framework was intended to support later Delphi elicitation and to integrate recommended practices across the reviewed literature.
- J.2.6 Risk of Bias within studies, in synthesis and in reporting biases; J.3 Results; J.3.1 Selected Studies and Characteristics: 52 selected studies were characterized by domain, authorship, audience, and publication type, with developers as the dominant intended audience.The studies covered artificial intelligence (50%), machine learning (23%), natural language processing (17%), and applied science (10%); 31/52 (60%) had five authors or fewer, and 38/52 (73%) targeted developers.
J.3.2 Extracted Item Characteristics … L Fields Ingested by EVALUATION CARDS
The review extracted and synthesized 730 recommendations from 52 studies into a five-part AI evaluation framework spanning design, execution, lifecycle, and reporting. Its findings emphasize transparency, accountability, context-dependent validity, and the need to improve operational evaluation practices.
- J.3.2 Extracted Item Characteristics: 44% of items addressed Reporting & Transparency, while 73% covered multiple workflow stages rather than a single stage.Reporting & Transparency appeared in 320/730 items; 197/730 addressed one stage and 533/730 addressed multiple stages.
- J.3.2 Extracted Item Characteristics: 78% of items pursued Transparency & Accountability and 73% discussed the evaluation Process, with outputs, systems, data/prompts, models, and environments receiving less attention.The most common objective was Transparency & Accountability (568/730), and the most frequent artifact was Process (534/730).
- J.3.3 Best Fit Framework: The review identified 13 framework papers, selected seven Evaluation Stages / Needs frameworks, and derived an initial framework of 11 candidate categories.Frameworks also covered construct/task types, ethics concerns, uncertainty types, long-form task processes, impacts, and audit stages/needs.
- J.3.4 Synthesis & Groupings: Expert feedback consolidated the 11 candidate categories into five higher-level groups: Evaluation Design, Before Evaluation Execution, Evaluation Execution, Evaluation Lifecycle, and Evaluation Reporting & Publication.The final framework refined category structure and item labels based on thematic overlap and expert feedback.
- Before Execution / Execution / Lifecycle / Reporting & Publication: The framework specifies preregistration, scoring and validation, splits and holdouts, pilots and baselines, contamination controls, run logging, adaptations, data access, maintenance, reporting, usage transparency, and reproducibility.These recommendations span the stages before execution, execution, lifecycle, and reporting and publication.
- J.4 Discussion: The discussion characterizes evaluation weaknesses primarily as disclosure failures, while emphasizing that results are context-dependent and provisional and that construct validity remains a central challenge.The review also identifies tensions between standardization and fit-for-purpose, and between feasibility and complexity.
- J.4 Discussion / J.5 Conclusion / J.8 Codebook: The review is a time-bounded organizational contribution rather than a finalized standard, and its consolidated framework combines Design, Before Execution, During Execution, Lifecycle, and Reporting & Publication.Limitations include a mid-2025 search cut-off, possible mainstream-literature bias, and subjective interpretive judgment mitigated through expert feedback and two reviewers at each step.
M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals · N Code and Demo
The systematic literature review is mapped to EVALUATION CARDS fields, ingested sources, lifecycle categories, and interpretive outputs, with traceability shown across the schema. The implementation is publicly available through an online code repository and live demo.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: The literature review items map to EVALUATION CARDS fields ingested from different sources and to interpretive signals.Figures 19 and 20 show these mappings.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: The mapping distinguishes sources including AutoBenchmarkCard, EEE aggregate eval, and EEE instance-level eval.These source keys appear in the schema mapping.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: Schema requirements are categorized as required, conditional, or optional.The requirement key defines these three statuses.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: The schema organizes attributes across five lifecycle categories: design goals and context, before execution, execution, lifecycle, and reporting and publication.These categories are listed alongside attribute groups and sources.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: Figure 19 maps input groups to individual items and the sources ingested by EVALUATION CARDS.The diagram presents the mapping as a Sankey diagram.
- M Mapping from Systematic Literature Review to EVALUATION CARDS Schema and Interpretive Signals: Figure 20 provides traceability from the literature-derived framework to EVALUATION CARDS fields and interpretive outputs.It links the review framework to both schema fields and downstream signals.
- N Code and Demo: The EVALUATION CARDS code is available in a Hugging Face Space repository.The repository URL is https://huggingface.co/spaces/evaleval/general-eval-card/tree/main.
- N Code and Demo: A live EVALUATION CARDS demo is available at https://evalcards.evalevalai.com.The demo is provided alongside the public code repository.