Source-linked AI summary
A Definition of AGI
Hendrycks, Dan, Song, Dawn, Szegedy, Christian, Lee, Honglak, Gal, Yarin, Brynjolfsson, Erik, Li, Sharon, Zou, Andy, Levine, Lionel, Han, Bo, Fu, Jie, Liu, Ziwei, Shin, Jinwoo, Lee, Kimin, Mazeika, Mantas, Phan, Long, Ingebretsen, George, Khoja, Adam, Xie, Cihang, Salaudeen, Olawale, Hein, Matthias, Zhao, Kevin, Pan, Alexander, Duvenaud, David, Li, Bo, Omohundro, Steve, Alfour, Gabriel, Tegmark, Max, McGrew, Kevin, Marcus, Gary, Tallinn, Jaan, Schmidt, Eric, Bengio, Yoshua
TL;DR
The paper addresses the lack of a concrete AGI definition by defining AGI relative to the cognitive versatility and proficiency of a well-educated adult. It operationalizes this definition with CHC-based cognitive batteries and ten core domains, finding a jagged profile in which current systems are strong in knowledge-intensive areas but weak in foundational machinery. The resulting scores quantify rapid progress while showing that substantial gaps remain before human-level general intelligence.
Problem
The lack of a concrete AGI definition obscures the gap between specialized AI and human-level cognition.
Method
The paper adapts CHC-based human cognitive batteries across ten core domains to produce a standardized AGI Score.
Results
Current AI systems show a highly jagged cognitive profile, with strength in knowledge-intensive domains and critical deficits in foundational cognitive machinery.
Takeaways & Limitations
The framework quantifies rapid progress while showing that substantial gaps remain before human-level general intelligence.
Takeaways & Limitations
The framework focuses on core cognitive capabilities of well-educated individuals, excluding physical abilities and direct prediction of economic automation or diffusion.
Abstract
from arXiv · showhide
The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 57%) concretely quantify both rapid progress and the substantial gap remaining before AGI.
1 Introduction
The paper defines AGI as matching the cognitive versatility and proficiency of a well-educated adult, then operationalizes that definition through CHC-based cognitive assessment. Applying the framework reveals strong specialized performance alongside substantial deficits in foundational cognition.
- 1 Introduction: AGI is defined as an AI matching or exceeding the cognitive versatility and proficiency of a well-educated adult.
- 1 Introduction: The framework adapts human cognitive batteries from CHC theory to evaluate AI abilities and produce a standardized AGI Score from 0% to 100%.A score of 100% signifies AGI.
- 1 Introduction: Contemporary AI systems solve roughly half of the assessed foundational abilities despite impressive performance on complex benchmarks.The paper characterizes current systems as narrower than well-educated humans overall but much stronger on some specific tasks.
- 1 Introduction: The framework organizes general intelligence into ten equally weighted cognitive components spanning knowledge, language, mathematics, reasoning, memory, and other abilities.The components are derived from CHC broad abilities and each receives a 10% weight.
- 1 Introduction: The framework is intended as a holistic, multimodal diagnostic tool for identifying AI strengths and weaknesses rather than as an automatic evaluation or fixed dataset.Its scope emphasizes cognitive abilities of well-educated individuals, excluding physical abilities and direct economic-automation prediction.
General Knowledge (K)
The General Knowledge domain measures knowledge familiar to well-educated adults or important enough that most adults have encountered it. GPT-4 has substantial general knowledge, while GPT-5 partially addresses its remaining gaps.
- General Knowledge (K): General Knowledge measures information familiar to most well-educated people or important enough that most adults have encountered it.
- General Knowledge (K): GPT-4 has substantial general knowledge, and GPT-5 partially fills its remaining gaps.
4 Reading and Writing Ability (RW)
Reading and Writing Ability covers the declarative knowledge and procedural skills needed to consume and produce written language. The supplied performance descriptions report GPT-4 limitations and GPT-5 improvements across reading, proofreading, and mathematical ability.
- 4 Reading and Writing Ability (RW): Reading and Writing Ability covers the declarative knowledge and procedural skills used to consume and produce written language.
- 4 Reading and Writing Ability (RW): GPT-4 struggles with token-level understanding, long documents, substring analysis, and careful proofreading, while GPT-5 addresses these issues.The stated causes include GPT-4’s small context window and imprecise working memory.
- 4 Reading and Writing Ability (RW): GPT-4 has limited mathematical capabilities, whereas GPT-5 has exceptional mathematical capabilities.
- 4 Reading and Writing Ability (RW): GPT-4 has negligible on-the-spot reasoning capabilities, while GPT-5 has only some remaining gaps.
On-the-Spot Reasoning (R)
The framework evaluates cognitive abilities including on-the-spot reasoning, working memory, and long-term memory. Its application shows that current AI systems have uneven capabilities, with particular weaknesses in long-term memory storage and retrieval reliability.
- On-the-Spot Reasoning (R): On-the-spot reasoning concerns flexible attention control for solving novel problems without relying exclusively on learned habits, schemas, or scripts.
- Working Memory (WM): Working-memory evaluation includes concrete assessment procedures and examines performance on working-memory tasks.
- Long-Term Memory: The memory assessment framework distinguishes acquisition and storage from retrieval, including fluency and hallucination avoidance.
- Long-Term Memory Storage (MS): Both GPT-4 and GPT-5 lack appreciable long-term memory storage capabilities.
- Long-Term Memory Retrieval (MR): GPT-4 and GPT-5 rapidly retrieve many parameterized concepts but frequently hallucinate.
Visual Processing (V)
Visual Processing (V) measures the ability to analyze and generate natural and unnatural images and videos. GPT-4 lacks visual perception and generation, while GPT-5 shows appreciable but highly incomplete visual capabilities.
- Visual Processing (V): Visual Processing (V) evaluates the ability to analyze and generate natural and unnatural images and videos.
- Visual Processing (V): The framework directs readers to Appendix H for further details on concretely assessing visual processing capabilities.
- Visual Processing (V): GPT-4 had no ability to perceive or generate images, while GPT-5 had appreciable but highly incomplete visual processing capabilities.
MS MR
The framework distinguishes AGI from economically valuable or strategically capable AI and treats intelligence as constrained by weaknesses across cognitive abilities.
- The Engine Analogy: The framework treats intelligence as an engine whose overall capability is limited by defective components, even when other abilities are highly optimized.
- Contamination: The framework warns that test contamination requires robustness checks under rephrased or related but distinct questions.
- The operationalization uses task specifications and illustrative datasets as necessary but insufficient evidence, allowing evaluations to evolve beyond fixed benchmarks.
- Evaluators may manually grade batteries, but varying precision means different graders can produce different AGI estimates.
- Limitations: Its scope excludes some faculties, uses English- and culturally specific examples, selectively samples knowledge, and makes results dependent on discretionary scoring weights.
- AGI is defined as matching or exceeding the cognitive versatility and proficiency of a well-educated adult, rather than achieving economic value or broad task replacement.
A General Knowledge (K)
General Knowledge measures familiarity with widely exposed or important information across science, social science, history, culture, and commonsense.
- General Knowledge contributes up to 10% of the AGI score through five 2% areas: commonsense, science, social science, history, and culture.
- Commonsense: Commonsense covers shared background knowledge about how the world works, including intuitive physics, procedures, temporal expectations, and morality.
- Science: Science proficiency spans physics, chemistry, and biology, with 1% awarded for exactly one subject and 2% for at least two.
- Social Science: Social science covers psychology, microeconomics, macroeconomics, geography, and comparative government, using the same 1%-or-2% scoring rule.
- History: History assesses European, US, world, and art history, with 1% for one proficient subject and 2% for two or more.
- Culture: Culture concerns literacy in widely recognized art, music, literature, media, and public figures, divided equally between current affairs and popular culture.
A.5.2 Popular Culture
The framework treats popular culture as cultural literacy and assesses reading, writing, English usage, and mathematical knowledge through structured ability components.
- A.5.2 Popular Culture: Popular culture evaluates knowledge of widely recognized art, music, literature, media, and public figures across text, audio, and visual modalities.
- Reading and Writing: Reading and writing ability is decomposed into letter-word ability, reading comprehension, writing ability, and English usage knowledge, totaling up to 10%.
- Reading Comprehension: Reading comprehension spans sentence, paragraph, and document levels, including whether questions are underdetermined by context.
- Writing Ability: Writing ability is assessed at sentence, paragraph, and essay levels, with a GRE Analytical Writing score of at least 4 out of 6 sufficient for its full 3%.
- Mathematical Ability: Mathematical ability covers arithmetic, algebra, geometry, probability, and calculus, awarding 1% for rudimentary and 2% for proficient performance in each area.
D On-the-Spot Reasoning (R)
On-the-spot reasoning measures novel problem solving through deduction, induction, theory of mind, planning, and adaptation, while working memory spans multiple modalities.
- D On-the-Spot Reasoning (R): On-the-spot reasoning is treated as a measure of abstract reasoning, adaptability, and algorithmic complexity handling, not a strong proxy for overall AI intelligence.
- D On-the-Spot Reasoning (R): The ability is decomposed into deduction, induction, theory of mind, planning, and adaptation, with weights of 2%, 4%, 2%, 1%, and 1%.
- Induction: Induction uses verbal and visual Raven’s Progressive Matrices, mapping below-average performance to no increase, below the 90th percentile to 1%, and at least the 90th percentile to 2%.
- Working Memory: Working memory maintains, manipulates, and updates information across textual, auditory, visual, and cross-modal forms, totaling up to 10%.
- Working Memory: Visual working memory receives the largest modality weight because textual and auditory working memory are partly tested in reading-writing and auditory abilities.
- Textual Working Memory: Textual working memory tests recall and transformation of short sequences, including list operations such as append, remove, sort, reverse, and set intersection.
E.2.2 Transformation Sequence
The framework tests whether AI systems can remember and transform information across modalities, procedures, and extended experiences. It separates short-term working-memory tasks from long-term storage and association.
- Working memory: The framework evaluates visual working memory through recall, transformation sequences, spatial navigation memory, and long-video question answering.Each component contributes 1% to the AGI score.
- Transformation sequence: Transformation-sequence tests require an AI to apply ordered operations such as adding, deleting, rotating, denoising, deblurring, or colorizing visual inputs.The task is linked to the CHC ability Visualization (Vz).
- Long-term storage: Long-term storage is tested separately from working memory by presenting information in one session and testing it in a new session without external tools.The framework distinguishes associative, meaningful, and verbatim memory.
- Associative memory: Associative-memory tests examine whether one stimulus later retrieves an originally unrelated stimulus across modalities, personalization contexts, and procedures.These tests cover cross-modal association, personalization adherence, and procedural association.
- Contextual memory: The framework also tests whether AI systems retain context-specific preferences, factual overrides, and multi-step procedures over time.Examples include remembering user-specific stylistic rules, updated facts, and data-cleaning instructions.
F.3.2 Set Recall
Set recall measures whether an AI can retrieve elements from previously presented word or image collections after a delay. The evaluation emphasizes both correct recall and resistance to intrusions or hallucinations.
- Set recall: Set recall requires naming elements from previously presented word or image collections after 48 hours.Word-set evaluation uses 10–20 words and measures the proportion recalled correctly.
- Evaluation criterion: The target performance for set recall is at least 90% precision and recall, while also limiting intrusions from items absent from the original set.The criterion applies to recalling elements of the presented set.
- Related abilities: The broader memory framework distinguishes set recall from spatial arrangement, long-term retrieval fluency, and hallucination avoidance.These neighboring abilities assess visual structure, access to stored knowledge, and resistance to confabulation.
- Related abilities: Visual processing is assessed through perception, generation, reasoning, and spatial scanning rather than through set recall alone.Perception includes image recognition, captioning, anomaly detection, clip captioning, and video anomaly detection.
H.2 Visual Generation
Visual generation evaluates whether AI systems can synthesize simple and complicated images and short natural videos. The paper treats these outputs as measurable proxies for conceptual and imaginative synthesis rather than direct human analogues.
- Visual generation: Visual generation covers simple natural images, complicated images, and simple natural videos, with 1% awarded for one task, 2% for two, and 3% for all three.Examples include generating a dog in a park, an anatomically unusual horse, and a person typing.
- Rationale: The framework includes image synthesis because translating abstract concepts into novel visual information is treated as a critical component of modern general intelligence.The resulting score is a measurable proxy rather than a direct human-equivalent ability.
- Visual reasoning: Visual reasoning spans gestalt reasoning, mental rotation and folding, embodied reasoning, figure question answering, and related spatial skills.The framework assigns 2% when the system is proficient across these visual-reasoning tasks.
- Spatial scanning: Spatial scanning tests visual search across complex fields, including mazes, object finding, map analysis, and target matching.Proficiency across these tasks contributes 1% to the AGI score.
I.3 Voice
The framework evaluates voice quality, musical abilities, and processing speed as distinct components of AI cognition. Speed tests compare AI latency or throughput with the average performance of a well-educated adult, counting artificial delays against the system.
- Voice: Voice evaluation covers natural speech and natural conversation, contributing 2% and 1% respectively.The tests assess non-robotic utterances and conversational fluidity without long delays or excessive interruptions.
- Musical judgment: Musical judgment tests whether systems can discriminate simple musical patterns without requiring musical jargon.Examples compare pitch, dissonance, and anomalous musical passages.
- Processing speed: Processing speed includes perceptual search and comparison, reading and writing, arithmetic, reaction time, and pointer fluency.The framework further distinguishes simple from choice reaction time.
- Processing speed: Processing-speed scores compare AI latency or throughput with a well-educated adult baseline, and artificial delays count toward the time limit.Each tested area receives 1% when the AI meets or exceeds the human baseline.