Source-linked AI summary
AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games
Lance Ying, Ryan Truong, Prafull Sharma, Kaiya Ivy Zhao, Nathan Cloos, Kelsey R. Allen, Thomas L. Griffiths, Katherine M. Collins, José Hernández-Orallo, Phillip Isola, Samuel J. Gershman, Joshua B. Tenenbaum
TL;DR
Existing benchmarks often test narrow, static task spaces rather than the generality and adaptability of human intelligence. The paper proposes AI GameStore, which uses LLMs and humans-in-the-loop to create standardized variants of human games for open-ended evaluation. Across 100 games, leading VLMs scored less than 10% of median human players on average and struggled with memory, world-model learning, and long-horizon reasoning, while the platform remains limited in scope and implementation.
Problem
Conventional benchmarks assess isolated capabilities in narrow task spaces, leaving the generality and adaptability of human intelligence insufficiently evaluated.
Method
AI GameStore uses LLMs and humans-in-the-loop to source, standardize, adapt, and generate variants of games from digital platforms for open-ended evaluation.
Results
Less than 10% of median human players’ scores was achieved on average by leading VLMs across 100 games, with major difficulties in episodic memory, world-model learning, and long-horizon reasoning.
Takeaways & Limitations
The AI GameStore functions both as a quantitative benchmark and as a diagnostic tool for identifying missing capabilities in machine cognition.
Takeaways & Limitations
The current implementation is limited in the capacities it can assess and the insights it can yield, while platform heterogeneity remains a major standardization challenge.
Abstract
from arXiv · showhide
Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technological advance. Conventional AI benchmarks typically assess only narrow capabilities in a limited range of human activity. Most are also static, quickly saturating as developers explicitly or implicitly optimize for them. We propose that a more promising way to evaluate human-like general intelligence in AI systems is through a particularly strong form of general game playing: studying how and how well they play and learn to play \textbf{all conceivable human games}, in comparison to human players with the same level of experience, time, or other resources. We define a "human game" to be a game designed by humans for humans, and argue for the evaluative suitability of this space of all such games people can imagine and enjoy -- the "Multiverse of Human Games". Taking a first step towards this vision, we introduce the AI GameStore, a scalable and open-ended platform that uses LLMs with humans-in-the-loop to synthesize new representative human games, by automatically sourcing and adapting standardized and containerized variants of game environments from popular human digital gaming platforms. As a proof of concept, we generated 100 such games based on the top charts of Apple App Store and Steam, and evaluated seven frontier vision-language models (VLMs) on short episodes of play. The best models achieved less than 10\% of the human average score on the majority of the games, and especially struggled with games that challenge world-model learning, memory and planning. We conclude with a set of next steps for building out the AI GameStore as a practical way to measure and drive progress toward human-like general intelligence in machines.
1 Introduction
The paper argues that evaluating AI on the broad space of human-designed games can better test human-like general intelligence than narrow, static benchmarks. It introduces AI GameStore as a scalable, open-ended platform for generating and evaluating such games, and finds substantial capability gaps between frontier VLMs and humans.
- The evaluation challenge: Traditional benchmarks measure isolated domains and therefore capture only fragments of the broad range of human capabilities and activities.Examples include chess, language understanding, domain-specific questions, mathematics, and coding.
- The Multiverse of Human Games: The Multiverse of Human Games evaluates how well AI systems play and learn to play conceivable human games against representative human adults with comparable resources.The proposed space is restricted to games that humans could create and that some broad segment of people could enjoy.
- The Multiverse of Human Games: Games provide broad coverage because they abstract human activities and exercise skills including strategic planning, resource management, social interaction, deception, pattern recognition, and physical navigation.This breadth motivates their use as a testbed for diverse cognitive capabilities.
- Practical challenges: The approach remains difficult to realize because actual human games form only a finite subset of conceivable games and commercial titles pose interface, intellectual-property, and data-contamination challenges.These constraints motivate synthetic adaptation and open-ended game generation.
- AI GameStore: AI GameStore uses LLMs to adapt games from digital marketplaces into standardized, containerized environments, with humans refining games and creating novel variants.The pipeline is intended to support a never-ending benchmark rather than a fixed suite.
- Initial evaluation: 100 games were evaluated with frontier VLMs and human participants, and the models achieved less than 30% of the human baseline on average while taking 15–20× more computation time.The models primarily struggled with robust world-model acquisition and long-term memory.
2 Measuring General Intelligence with the Multiverse of Human Games
The paper proposes evaluating human-like general intelligence through representative samples of the open-ended Multiverse of Human Games, while using digital games as a practical approximation. Games are motivated as abstractions of real-world activities, but direct evaluation over commercial games faces major scale, access, and technical constraints.
- 2.1 Why is this a good measurement of general intelligence?: The Multiverse of Human Games comprises all human-designed games and is proposed as a comprehensive testbed for machine general intelligence.The paper frames this space as infinite and open-ended, with games designed to be enjoyable and learnable by people.
- 2.1 Why is this a good measurement of general intelligence?: Games abstract real-world activities and provide practice for adapting to complex problems, spanning capabilities such as planning, reasoning, and social interaction.The paper links play to learning and describes games as cultural artifacts that transmit abstractions of real-world challenges.
- 2.2 Practical challenges of evaluating general intelligence with human games: Because the full space of human games is unbounded, the paper proposes starting with finite collections of digital-platform games that retain broad cognitive diversity.The digital-game space is presented as a practical subset of the broader Multiverse.
- 2.2 Practical challenges of evaluating general intelligence with human games: Directly benchmarking commercial digital games is constrained by copyright, heterogeneous platforms, limited human-data access, latency, and contamination risk.These barriers make a universal, monolithic benchmark difficult to implement and can undermine evaluation validity.
- 2.2 Practical challenges of evaluating general intelligence with human games: AI GAMESTORE addresses these constraints by sourcing and rebuilding synthetic, standardized game environments through an automated pipeline with human refinement.The platform is intended as a tractable, scalable, open-ended approximation to the broader Multiverse vision.
3 The AI GAMESTORE
AI GAMESTORE constructs a living benchmark by sourcing popular digital games, generating standardized adaptations, refining them with simulated play and humans, profiling cognitive demands, and evaluating models against people. The pipeline also supports novel variants and continual expansion to reduce benchmark saturation.
- 3 The AI GAMESTORE: AI GAMESTORE is a unified sandbox for standardized adaptations of popular human-designed digital games, enabling comparison between AI systems and human performance.It uses synthetic games rather than directly evaluating the original platform titles.
- 3 The AI GAMESTORE: Approximately 30 minutes is required on average to generate and refine a new game with human-in-the-loop, supporting continual benchmark expansion.Crowdsourced participants enable the suite to grow as new games emerge on digital marketplaces.
- 3 The AI GAMESTORE: The four-stage pipeline sources and filters games, generates and refines playable versions, annotates cognitive demands, and evaluates models and humans through a standardized interface.The stages combine LLM-based generation and scoring with automated testing, human feedback, cognitive profiling, and performance analysis.
- 3 The AI GAMESTORE: Human-generated variants add novel mechanics to source games and provide multiple valid tasks while mitigating benchmark saturation.Variants are created through the same refinement interface used to improve base games.
- 3 The AI GAMESTORE: The initial benchmark contains 100 games, with 10 released publicly and 90 retained as a private test set to limit overfitting and saturation.The public/private split supports demonstration and open experimentation while preserving an unreleased evaluation portion.
4 Model and Human Experiments
The experiments compare seven frontier vision-language models with human participants on short interactions across the AI GAMESTORE games. Humans play through a customized interface, while models use a lightweight harness that supplies game state, screenshots, actions, and a persistent scratchpad.
- 4 Model and Human Experiments: The benchmarking study evaluates state-of-the-art AI models and human participants on 100 curated AI GAMESTORE games.The study is designed to compare model performance with human play across the corpus.
- 4 Model and Human Experiments: 106 human participants each played 10 randomly assigned games, receiving two minutes per game through a customized interface.The human study was conducted under an MIT IRB-approved protocol.
- 4 Model and Human Experiments: Seven frontier VLMs were evaluated, with each model run three times per game and performance and runtime averaged across runs.The models included GPT-5.2, GPT-5-MINI, GEMINI-2.5-PRO, GEMINI-2.5-FLASH, CLAUDE-OPUS-4.5, QWEN-3-VL-32B, and LLAMA-4-MAVERICK.
- 4 Model and Human Experiments: Because current models cannot interact with games like humans in real time, the study uses a lightweight harness for model-game interaction.The harness is motivated partly by model API latency exceeding normal human response times.
- 4 Model and Human Experiments: Each model prompt includes game instructions, a scratchpad, recent actions and screenshots, and the currently available actions.The scratchpad functions as model memory by recording gameplay history and other information across API calls.
5 Results
Across 100 games, frontier VLMs remain far below human performance, are substantially slower, and struggle most when tasks require multiple or sophisticated cognitive capabilities. Models often make initial progress but then stagnate or fail to progress.
- Less than 10% of the human baseline was achieved by GPT-5.2, Gemini-2.5-Pro, and Claude-Opus-4.5 in geometric mean score.The differences between the top six models were not statistically significant.
- 12–18 times longer runtime was required on average for models than for humans playing the games.Humans played for 120 seconds, while models commonly required more than 1,200 seconds to complete the API calls.
- 30–40% of games produced less than 1% of the median human score for all models.On roughly two-thirds of games, models made some progress, but most of those scores were only 10–30% of median human performance.
- Memory, planning, and world-model learning were the cognitive demands on which models struggled especially.Performance also decreased significantly as planning and social-reasoning demands became more sophisticated.
- Games requiring more distinct cognitive capabilities caused significantly lower relative model performance.Success often requires integrating capabilities rather than merely possessing sufficient competence in each one.
- Models typically made early progress before stagnating or advancing much more slowly than humans, while all models failed to progress on a significant fraction of games.Humans made steady progress across the 10 public games shown in Figure 9.
6 Discussion and Future Directions
The AI GAMESTORE is a proof of concept for scalable, open-ended evaluation using human games, but its current suite remains limited in difficulty, diversity, and diagnostic precision. Future work targets richer social interaction, longer-horizon gameplay, more challenging level generation, and improved capability profiling.
- The current platform is a modest first step toward operationalizing the full Multiverse of Human Games vision.
- Current games emphasize rapid learning of relatively simple, novel games, with human participants playing mainly for fun rather than material reward.
- Many games use naive NPCs, limiting assessment of complex social reasoning such as recursive mentalizing.
- The suite primarily contains easy, short-duration, or casual games, so it does not capture important long-horizon and multi-timescale human activities.
- LLMs can generate playable casual games but often produce levels that are too trivial or impossible, motivating more scalable testing and iteration.
- Interactive-game scores often confound multiple skills, motivating ability-profiling methods that estimate latent capabilities across overlapping skill sets.
A.1 AI evaluation
The paper positions AI GAMESTORE as a response to static, narrow benchmarks by grounding continual evaluation in the broad distribution of human-created games. It extends general game playing from performance across environments to learning and playing human-designed games under human-like constraints.
- Traditional benchmarks target specific cognitive domains and remain fundamentally static despite enabling standardized model comparisons.
- Benchmark optimization has raised concerns about saturation, contamination, and narrow capability measurement.
- AI GAMESTORE proposes a continually generated meta-benchmark covering a broad spectrum of human activities and skills.
- The Multiverse of Human Games shifts general game playing toward learning and playing the full space of human-designed games under human-like constraints.
A.3 LLM-guided game generation
AI GAMESTORE combines LLM-based environment generation with human refinement while anchoring games in popular human game concepts. This approach aims to preserve playability, diversity, and evaluative relevance as the evaluation suite scales.
- LLMs have been used to generate interactive environments, translate descriptions into playable worlds, design levels, and produce full games through iterative pipelines.
- Fully automated generated-task evaluation can suffer from weak playability, representativeness, and evaluation validity.
- AI GAMESTORE anchors generation in existing popular game concepts and combines LLM generation with human-in-the-loop refinement.
- The pipeline produces variants intended to preserve playability, diversity, and evaluative relevance while remaining aligned with human-created games.
B Sourcing Games
The first AI GAMESTORE corpus was sourced from major Apple App Store and Steam game rankings, then curated into a diverse 100-game suite. Standardized interaction and scoring constraints make model–human comparisons more tractable.
- Sourcing Games: The source pool comprised 7,500 Apple App Store games from five categories across 15 countries, plus 500 top Indie games from Steam.
- Sourcing Games: The final 100-game suite spans diverse genres, with action as the largest cohort and many puzzle and board games.
- Implementation constraints: Games are written in JavaScript with p5.js, with optional three.js and matter.js support for graphics and physics.
- Implementation constraints: Keyboard-only interaction restricts the action space to key presses, while future versions may support cursor movement.
- Implementation constraints: Games can be paused and resumed to accommodate the high latency of current AI-model API calls.
- Evaluation constraints: Scoring, progressive difficulty, and sufficient lives or resets support measurable comparison while allowing players to learn during sessions.
D Game Refinement
AI GameStore refines source games through human play and LLM feedback, then generates variants whose mechanics can target different cognitive demands.
- Game Refinement: Human players interact with generated games, describe issues or requested features, and ask a target LLM to apply fixes until the game is satisfactory.The refinement interface supports real-time play, automated testing, natural-language feedback, and iterative fixes.
- Game Refinement: The authors performed all iteration and refinement themselves, while proposing online recruitment as a scalable future alternative.This limits the current refinement process to the coauthors’ involvement.
- Game Variants: Multiple variants can be generated from few source games by augmenting mechanics, enabling targeted evaluation of cognitive capability demands.Human players propose novel mechanics, which can be implemented through the refinement interface.
- Cognitive-Demand Rubrics: The annotation rubrics define capability demands independently for each game without assuming independence or correlation among capabilities.The rubrics are intended to provide interpretable, human-readable construct definitions and flexible capability annotations.
- Cognitive-Demand Rubrics: The rubric spans spatial-temporal coordination, visual object identification, memory, and discovery of hidden mechanics.Examples range from static or turn-based tasks to games requiring long-horizon memory, visual inference, and active hypothesis testing.
G Comparing Public and All Games
The 10 publicly released games provide a high-fidelity proxy for the full 100-game benchmark, with comparable human-centric ratings and model performance.
- Public Subset Representativeness: Comparable mean ratings for Funness, Challengingness, and model performance support the 10-game public subset as representative of all 100 games.The comparison evaluates representativeness against the full dataset.
H Model Experiment Details
The evaluation harness queries models at fixed intervals, providing game state, prior actions, screenshots, available keys, and scratchpad context before executing predicted action sequences.
- Action Interface: Each API call requests five action objects, each controlling a 0.2-second interval through available keyboard inputs.The five objects together specify the model’s actions for the next game second.
- Action Interface: The harness supports simultaneous key presses, instant or held actions, and level restarts through RETRY.HOLD actions persist for the full 0.2-second interval, whereas regular presses apply once.
- Model Prompt: The prompt includes game instructions, screenshots, previous actions, reasoning, and a scratchpad that carries the model’s gameplay understanding across calls.The scratchpad is used to maintain state and plans between API queries.
I Additional Experimental Results
Additional analyses report model scores across all 100 games and test whether slow reaction time explains performance on games with low spatio-temporal demands.
- Aggregate Performance: Model median normalized scores and geometric mean scores are reported across all 100 games.The results are summarized in Table 10.
- Reaction-Time Analysis: Low-demand games with Spatial Temporal Coordination scores of 2 or below produced little aggregate performance difference.These games typically include puzzles and turn-based strategy games.
- Reaction-Time Analysis: Top-model performance, including GPT-5.2 and GEMINI-2.5-PRO, did not change significantly on the low-coordination subset.The analysis indicates that slow reaction time alone does not explain model failure on these games.