Source-linked AI summary
Evaluating Cognitive Age Alignment in Interactive AI Agents
Yifan Shen, Jiawen Zhang, Jian Xu, Junho Kim, Ismini Lourentzou, Xu Cao, Meihuan Huang
TL;DR
Current AI evaluations offer limited ways to determine whether MLLM agents match human developmental stages rather than merely maximizing correctness. ChildAgentEval provides a WISC-inspired interactive benchmark and shows that standard age prompting fails, while skill guidance improves alignment unevenly across cognitive domains.
Problem
Current evaluation paradigms provide limited tools for determining whether MLLM agents align with distinct stages of human cognitive development.
Method
ChildAgentEval is a WISC-inspired interactive benchmark that evaluates MLLM agents across cognitive factors using age-specific items, difficulty levels, and developmental constraints.
Results
Standard age prompting fails to produce reliable age trajectories, while skill guidance improves differentiation in stronger models but alignment remains uneven across cognitive domains.
Takeaways & Limitations
Authentic developmental alignment requires cognitive constraints on perception, memory, and reasoning rather than stylistic role-play alone.
Takeaways & Limitations
The constraint design does not generalize uniformly across model families and requires sufficient baseline controllability and instruction-following capacity.
Abstract
from arXiv · showhide
While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life to advanced scientific research, a profound gap remains between artificial and human intelligence. Despite the integration of powerful tools and advanced MLLMs, state-of-the-art AI agents frequently fail at foundational, seemingly simple tasks that a child can resolve with ease. Inspired by the Wechsler Intelligence Scale for Children (WISC), we introduce ChildAgentEval, the first psychometrically grounded interactive benchmark for evaluating cognitive age alignment in MLLM-based agents. ChildAgentEval systematically compares the reasoning performance of various MLLM-based interactive agents against age-specific human developmental stages, exposing where current agentic AI systems can and cannot simulate age-specific cognitive behavior.
1 Introduction
The paper defines cognitive age alignment as calibrating interactive-agent behavior to human developmental structures rather than maximizing raw capability. It introduces ChildAgentEval, a WISC-inspired benchmark, and finds that standard prompting yields unstable age trajectories while skill guidance improves differentiation unevenly across cognitive domains.
- Most agent benchmarks prioritize task correctness and advanced completion, rarely evaluating whether reasoning is developmentally appropriate for a specific user.
- Cognitive age alignment requires structured developmental constraints, including simpler language and restricted working memory for younger targets and stronger reasoning for older targets.
- ChildAgentEval uses a WISC-IV-informed framework to assess verbal comprehension, perceptual and fluid reasoning, working memory, age-normed scores, subtest behavior, developmental trends, and language complexity.
- Standard age prompting does not reliably induce developmental alignment, whereas skill guidance improves differentiation and produces more monotonic trajectories in stronger proprietary models.
- The proposed skill-guided distillation strategy converts developmental markers into executable constraints and exposes limitations in working memory and visuospatial reasoning.
2 Related Works
Prior work benchmarks language and vision-language models with psychological assessments and uses generative agents to simulate human cognition. However, existing literature does not systematically distill age-specific skills from child data or test whether agents reason like particular age groups.
- Psychological and cognitive benchmarking: Psychological benchmarking extends beyond traditional evaluation through IQ, EQ, and PQ frameworks and assessments such as the Wisconsin Card Sorting Test.These approaches evaluate LLMs and VLMs from human-centered or classical psychological perspectives.
- Psychological and cognitive benchmarking: Recent studies compare generative models with population-normed intelligence distributions and human psychometric tests.This work positions psychometric comparison against human normative distributions as an evaluation direction for foundation models.
- Human cognition simulation: LLMs and generative agents are increasingly used as computational tools to simulate human behavior, psychological processes, roles, and Theory of Mind.Existing systems include behavioral prediction, psychological simulation, character-based cognitive modeling, executive-agent design, and real-world Theory of Mind studies.
- Research gap: Existing literature lacks systematic methods to distill age-specific skills from real child data and psychometrically measure whether agents reason like specific age groups.This gap distinguishes child cognitive-age evaluation from prior adult-focused or general-purpose simulation work.
- Child cognitive simulation: Child-focused work evaluates model safety, language patterns, and alignment with young users’ preferences across developmental stages.Related studies simulate child agents and conduct psychology experiments probing the computational capabilities of models such as LaMDA and GPT.
3 ChildAgentEval
ChildAgentEval adapts WISC-inspired cognitive constructs into a psychometrically grounded web benchmark for interactive AI agents. Its ten subtests assess four cognitive domains across ages 6–16 using standardized administration, scoring, and error reporting.
- Design Principles and Grounding: Interactive web tasks preserve cognitive constructs by requiring dynamic browser actions, working-memory maintenance, and timed or sequential interactions rather than static items.The platform uses text inputs and separates sequences across pages to prevent context-window leakage.
- Design Principles and Grounding: Ten interactive subtests map to the CHC model, covering crystallized intelligence, fluid and visual-spatial reasoning, working memory, and processing speed.Crystallized intelligence includes Similarities, Vocabulary, and Comprehension.
- The Interactive Web Environment: A Finite State Machine independently administers each subtest and applies Reversal and Discontinuation rules to control item progression.Reversal returns agents to foundational items after failure at a higher age starting point, while Discontinuation ends a subtest after consecutive zero scores.
- Evaluation Protocol: Raw subtest scores are converted using age-based normative tables, aggregated into domain Index Scores, and synthesized into FSIQ with categorized error tags.The benchmark reports detailed performance metrics alongside the final outputs.
4 Age-Specific Cognitive Skill Distillation
ChildAgentEval derives age-specific cognitive constraints from real child and adolescent interaction data rather than stereotypes or simple role prompts. It distills these features into cognitive profiles and executable agent filters controlling vocabulary, memory, reasoning, visual reliance, and social perspective.
- Age-Specific Cognitive Skill Distillation: Age-specific settings are extracted from real interactions of children and adolescents and translated into executable constraints for language-model agents.The approach uses a parameterized cognitive distillation architecture instead of subjective stereotype-based construction or simple role prompts.
- Data Collection and Age Slicing Normalization: A multi-source corpus covering ages 6 to 17 captures developmental features through age-appropriate spoken, multimodal, classroom, interview, and narrative data.Lower-age data captures vocabulary boundaries, attention spans, and self-repair markers, while higher-age data includes classroom discussions, psychological interviews, and narrative writing.
- Cognitive Profile Vector Representation: Six cognitive-profile dimensions encode age-specific limits on vocabulary abstraction, reasoning depth, temporary information retention, visual reliance, processing speed, and social perspective.The dimensions include Gc, Gf, Gv, WM, PSI, and an auxiliary Social variable for perspective switching and egocentric bias.
- Two Stage Skill Distillation Pipeline: The two-stage pipeline extracts linguistic and psychological statistics, then uses a teacher language model to produce standardized cognitive skill cards.The cards specify age-group vocabulary boundaries, multi-step reasoning limits, preferred resolution strategies, and expected logical error patterns.
- Cognitive Filter Module and Agent Integration: Five cognitive filter modules implement distilled skills across prompting, memory, and reasoning planning, including vocabulary, working-memory, reasoning-budget, visual-reliance, and social-perspective controls.These filters restrict academic concepts and retained information for lower ages, direct observation matching versus hypothesis verification, reproduce visual illusions, and constrain social explanations by developmental perspective.
5 Experiments
Experiments compare baseline prompting with skill-guided age-specific configurations across four anchor ages using normalized performance, developmental trajectory, age-normed factor deviations, and linguistic-age metrics. Skill guidance produces age-ordered trajectories in stronger proprietary models, but alignment depends on model capability and remains uneven across cognitive factors.
- Experimental Design: The evaluation covers ages 7, 10, 13, and 16, contrasting age-labeled standard prompting with distilled age-specific skill configurations.Metrics include normalized total and subtest scores, trajectory monotonicity and trend statistics, age-normed z-scores for FSIQ and four cognitive factors, and linguistic age fidelity.
- Developmental Trajectories: Baseline agents fail to show stable age-ordered progression across highly capable proprietary models; baseline GPT-5.4 scores 0.53 at age 7, 0.46 at age 13, and 0.52 at age 16.The trajectory decreases before recovering rather than increasing consistently with target age.
- Developmental Trajectories: Skill-guided prompting induces monotonic total-score increases from age 7 to age 16 across evaluated proprietary models.For GPT-5.4-based agents, scores rise from 0.41 at the 6–8 age band to 0.50 at the 15–17 age band.
- Capability Threshold: Skill-guided alignment depends on sufficient baseline controllability and instruction-following capacity, with Qwen3.5-27B and Gemma-4-31B remaining comparatively weak.The constraint design does not generalize uniformly across model families.
- Factor-Level Profiles: Age alignment is factor-specific: Gc and language-related dimensions scale with age, whereas Gf/Gv, WMI, and PSI are less sensitive to developmental constraints.Gemini-3.1-Pro shows positively shifted Gc and WM indices, while Gf/Gv falls behind normative trajectories and PSI remains below norm.
- Quantitative Differentiation: GPT-5.4’s Spearman correlation improves from -0.40 at baseline to 1.00 with skill guidance, while its age-16 versus age-7 score gap increases to 0.09.The absolute score range remains modest despite the consistent monotonic shift.
6 Discussion
Cognitive age alignment depends on selectively reconfiguring behavior, not uniformly reducing capability. Standard prompting prioritizes correctness over age-consistent behavior, while skill guidance improves calibration but cannot overcome uneven alignment and architectural bottlenecks.
- Behavioral Reconfiguration: Cognitive age alignment requires selective behavioral reconfiguration rather than uniform capability reduction.The discussion frames age alignment as behavioral adaptation rather than a general decrease in capability.
- Prompting Limitations: Standard prompting fails to produce age-ordered trends because agents prioritize task correctness over behavioral consistency.This prioritization disrupts consistent simulation of age-specific behavior.
- Calibration: Skill guidance improves calibration by constraining reasoning, memory, and vocabulary, but alignment remains uneven.The gains from guidance do not yield uniformly age-aligned behavior across these dimensions.
- Architectural Bottlenecks: MLLMs readily adapt linguistic style, but architectural bottlenecks hinder authentic reproduction of human-like memory decay and perception.The passage contrasts flexible language adaptation with limitations in reproducing other human cognitive characteristics.
7 Conclusion … A.1 Execution modes: vision-only vs. DOM-assisted.
ChildAgentEval evaluates developmental alignment in MLLM agents using a WISC-grounded framework and skill distillation, showing that nominal age instructions alone are insufficient while cognitive filters produce age-ordered behavior. The appendix documents implementation details, execution modes, calibration, scoring, reproducibility, statistical framing, factor analyses, data processing, impacts, and limitations.
- 7 Conclusion: ChildAgentEval combines a WISC-grounded interactive framework with data-driven skill distillation to evaluate and implement developmental alignment in MLLM agents.The framework is presented as the paper’s central contribution.
- 7 Conclusion: Nominal age instructions are insufficient because general-purpose agents default to their maximum capabilities, whereas targeted cognitive filters yield monotonic scores and age-ordered linguistic patterns in high-performing models.The conclusion links these findings to authentic alignment.
- Appendix Contents: The appendix contents cover implementation details, execution modes, calibration, scoring, computational reproducibility, statistical framing, WISC normative scoring, factor analyses, data processing, broader impacts, and limitations.These topics are listed as appendix sections rather than described substantively in the supplied passages.
- Appendix Contents: The appendix further lists factor-level trajectories, linguistic profiles and cross-model cognition, data collection and processing details, ChildAgentEval’s role, broader impacts, and limitations.These entries extend the appendix’s documented scope.
- A.1 Execution modes: vision-only vs. DOM-assisted.: The browser environment supports vision-only interaction from rendered screenshots and DOM-assisted interaction using screenshots plus a sanitized accessibility tree.The accessibility tree describes visible interactive elements, roles, and bounding boxes.
- A.1 Execution modes: vision-only vs. DOM-assisted.: DOM assistance excludes hidden states, answer keys, and backend data attributes, and both modes require Playwright interactions rather than direct textual responses to the evaluator.The filtering is intended to support spatial action grounding without cognitive shortcuts.
A.2 Calibration Data and Separation from Evaluation … B Factor-Level Analysis and Multidimensional Profiles
The benchmark calibrates developmental constraints independently of evaluation scores, uses deterministic scoring with blinded human verification for open-ended items, and records reproducible age-profile summaries under fixed protocols. Its analyses emphasize descriptive developmental trajectories and standardized age-stratified human norms, while omitting formal uncertainty estimates and significance tests.
- A.2 Calibration Data and Separation from Evaluation: Skill configurations derive from age-stratified developmental corpora and summaries rather than directly fitting benchmark evaluation scores.The constraints target vocabulary abstraction, memory capacity, reasoning depth, and related behavioral markers.
- A.2 Calibration Data and Separation from Evaluation: Benchmark items, answer keys, and item-level scores are excluded as tuning objectives to prevent data leakage.The calibration stage specifies developmentally motivated constraints instead of numerically matching benchmark tables.
- A.2 Calibration Data and Separation from Evaluation: Calibration is framed as developmental constraint design, with future validation planned through larger held-out corpora and filter-transferability ablations.This explicitly distinguishes the current pipeline from score matching on the evaluation set.
- A.3 Scoring and human verification: Objective subtests are scored deterministically from browser state using exact-match, ground-truth selection, or correct operations within a time limit.Examples include Digit Span, Letter–Number Sequencing, Picture Concepts, and Matrix Reasoning.
- A.3 Scoring and human verification: Open-ended verbal responses receive blinded human verification, with GPT-5.4 limited to pre-annotation and ambiguity flagging rather than final scoring.Responses are anonymized, and raters are blind to model identity, setting, target age, and trajectory-level hypothesis.
- A.3 Scoring and human verification: Two raters independently apply a standard 0/1/2 rubric, while disagreements undergo third-reviewer or child-psychology-trained adjudication.Final benchmark tables use the human-verified scoring process.
- A.4 Computational Execution and Reproducibility: Greedy decoding at temperature 0.0 and independent Playwright browser contexts support deterministic, parallel administration, which takes approximately one hour per complete administration.Workers control one model, setting, and age configuration at a time.
- A.5 Statistical Framing of the Main Results: Trajectory metrics such as Spearman rank correlation, age-gap scores, and regression slopes summarize deterministic age-ordered behavior without confidence intervals, significance tests, or repeated-run variance estimates.The main analyses are descriptive and future work may introduce stochastic sampling and bootstrap confidence intervals.
B.1 Fine-grained Factor-Level Trajectories
Factor-level trajectories show that baseline GPT-5.4 performance is generally flat, non-monotonic, or extreme, while skill-guided constraints produce clear age scaling only for Gc. WM remains ceiling-saturated and PSI fluctuates without a developmental trend, whereas Gf/Gv is not responsive to current prompt-based constraints.
- Gc trajectory: Only Gc shows clear monotonic age scaling under skill-guided constraints, rising from 0.64 at age 7 to 0.84 at age 16.The pattern indicates that vocabulary boundaries and abstractness filters function as intended.
- Overall factor trajectories: Baseline trajectories remain largely flat, non-monotonic, or pegged at performance extremes across all four psychometric factors.Figure 6 compares normalized Gc, WM, Gf/Gv, and PSI scores across four anchor ages.
- Gf/Gv trajectory: Gf/Gv is not responsive to the current prompt-based developmental constraints, consistent with evidence of a broader visual cognition gap in multimodal LLMs.The passage contrasts this insensitivity with Gc’s clear upward scaling under skill-guided constraints.
- WM and PSI trajectories: WM remains saturated at the performance ceiling across the simulated age span, while PSI fluctuates at a lower tier without a clear developmental trend.These patterns are described as distinct insensitivities to the simulated age settings.
B.2 Linguistic Profiles and Cross-Model Cognitive.
The section defines Lang. as a supplementary language-complexity dimension for open-ended verbal subtests and uses it alongside cognitive profiles to examine age-calibrated agent behavior. Across target ages, language complexity changes with cognitive capacity, including a GPT-5.4-based agent increase from 0.09 at age 7 to 0.38 at age 16.
- Language-complexity dimension: Lang. quantifies response length, lexical diversity, abstraction, categorization, and explanatory structure across Similarities (T2), Vocabulary (T6), and Comprehension (T9).It combines seven component metrics, including MLU, MATTR, abstractness, category rate, causal rate, definition rate, and average number…
- Language-complexity dimension: Lang. uses max-normalized component metrics averaged across evaluated model conditions.MLU measures response length, while MATTR measures lexical diversity with a 50-unit sliding window.
- Language-complexity dimension: The remaining language features identify abstract lexical items, superordinate categorization, causal connectives, and definitional structures in responses.These components are implemented as response-level binary or frequency features.
- Cross-model cognitive profiles: Figure 7 compares normalized subtest and factor-level profiles for age 7 versus age 16, integrating language complexity with four cognitive factors.The visualizations show how cognitive performance and linguistic behavior vary across models and target ages.
- Cross-model cognitive profiles: 0.09 at age 7 to 0.38 at age 16 marks the GPT-5.4-based agent’s language composite expansion.The analysis reports that language complexity scales synchronously with cognitive capacity, while factor balance changes structurally across ages.
- Cross-model cognitive profiles: Subtest improvements in language-related and verbally mediated tasks drive outward expansion in higher-order dimensions such as Gc and Lang.The section presents synchronous growth between cognitive capacity and language complexity as a critical scientific finding.
C Data Collection and Processing Details · D Role of ChildAgentEval · E Broader Impacts
ChildAgentEval builds age-aligned cognitive evaluation from multi-source corpora spanning ages 6–17 and operationalizes it as a model-agnostic, web-based infrastructure. It also targets developmentally appropriate child-facing AI while cautioning against treating benchmark scores as clinical measurements.
- C Data Collection and Processing Details: Ages 6–11 are represented mainly through spoken and multimodal interaction corpora capturing everyday vocabulary, attention spans, and self-repair markers.The sources include CHILDES, OCSC, and Frog Story.
- C Data Collection and Processing Details: Ages 12–17 are represented through classroom discussions, psychological interviews, and narrative writing focused on adolescent language and reasoning features.These data support analysis of abstract vocabulary, long-range logical organization, and adolescent egocentric bias.
- C Data Collection and Processing Details: The corpora are divided into four metadata-based groups: 6–8, 9–11, 12–14, and 15–17 years old.Lexical diversity is measured with MATTR, while structural and fluency metrics include Mean Length of Utterance, grammatical depth, and mid-course corrections.
- D Role of ChildAgentEval: ChildAgentEval is an evaluation infrastructure, not a static benchmark dataset, providing standardized web-based cognitive-task administration with consistent interaction and scoring protocols.The platform is designed to evaluate agents under controlled procedures.
- D Role of ChildAgentEval: The system is model-agnostic, allowing external researchers to integrate their own large language models as interactive agents for evaluation.Users conduct evaluations through age-specific cognitive skill distillation procedures.
- E Broader Impacts: ChildAgentEval may support safer, more developmentally appropriate child-facing AI in education, tutoring, and assistive interaction.It evaluates alignment in language, reasoning, memory use, and explanation style with target developmental stages, beyond raw task accuracy.
- E Broader Impacts: A stated risk is over-interpreting benchmark scores as clinical measurements.The broader-impacts discussion also flags potential misuse of age-simulation ability.
F Limitations
ChildAgentEval operationalizes cognitive factors through constrained task performance rather than biological reconstructions. Its controlled browser setting supports standardized evaluation but leaves backend effects and other interaction formats for future work.
- Operational interpretation: Processing-speed scores reflect overall pipeline efficiency, including model inference, API latency, and tool-orchestration overhead, rather than pure cognitive execution speed.Identical task protocols cannot completely isolate cognitive execution speed from backend delays.
- Operational interpretation: Working-memory measures do not reproduce the model’s intrinsic architecture, human memory decay, or biological attentional bottlenecks.Both processing speed and working memory are constrained operational metrics rather than direct structural implementations of human cognitive limits.
- Scope: The benchmark’s controlled browser environment enables standardized administration and detailed logging but excludes voice-based and long-horizon tutoring scenarios.These interaction formats are left for future work.
- Scope: The benchmark evaluates representative age bands to obtain stable developmental trajectories, while finer-grained extensions remain future work.The passage identifies representative age bands as the current evaluation scope.