Source-linked AI summary
What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models
Paolo Ciancarini, Remo Pareschi
TL;DR
Chess research on strategic reasoning is rich but fragmented across agent types, concepts, and evidence standards. This systematic mapping makes those traditions comparable and finds strong coverage of assessment, evaluation, and action selection, alongside underdeveloped planning, explanation, metacognition, and human–AI complementarity.
Problem
Chess studies address strategic reasoning across heterogeneous agents and standards, making it difficult to determine which components have been investigated and evaluated.
Method
The paper systematically maps chess research across human, classical-engine, neural, reinforcement-learning, LLM, and hybrid approaches using shared strategic-reasoning and evaluation frameworks.
Results
Research is strongest on assessment, evaluation, and action selection, while explicit planning, explanation, metacognition, and human–AI complementarity remain comparatively underdeveloped.
Takeaways & Limitations
The framework makes distinct research traditions comparable while showing that hybrid architecture and human–AI synergy require finer-grained evaluation distinctions.
Takeaways & Limitations
The mapping remains subject to coder subjectivity because independent duplicate coding was not performed across the complete corpus.
Abstract
from arXiv · showhide
Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has renewed the relevance of chess as a controlled environment for investigating strategic reasoning and comparing human and artificial decision-making. We present a systematic mapping study of recent research spanning human players, classical chess engines, neural and reinforcement-learning systems, LLMs, and hybrid approaches. The final map comprises 84 core study families, classified according to agent type, strategic-reasoning stages, and evaluation dimensions. The map reveals a literature strongly concentrated on situation assessment, evaluation, and action selection, while explicit planning, explanation, metacognition, and human--AI collaboration remain less explored. LLM research places particular emphasis on state representation and generalization, whereas grounded explanation appears more frequently in hybrid approaches combining language models with engines, expert knowledge, or other external structures. Two distinctions emerge that the map aggregates rather than resolves: hybrid systems differ in where and when heterogeneous capabilities combine, and evaluations that show improved human performance do not thereby establish human--AI synergy. We propose both as extensions of the mapping framework. We argue that chess provides a useful bridge between cognitive and computational perspectives on strategic reasoning, and identify explicit planning, grounded and faithful explanation, metacognitive calibration, and human--AI complementarity as directions for future research.
1 Introduction
Chess research now spans human cognition, classical and neural engines, LLMs, and hybrid systems, but these traditions operationalize strategic reasoning differently. This study maps their shared components and evaluation practices to clarify what has been investigated and where gaps remain.
- Emerging LLM research: LLMs extend chess research beyond explicit game-tree search toward symbolic state processing, move prediction, question answering, and natural-language explanation.Their evaluation must distinguish coherent board and rule representations from apparent knowledge derived from training patterns.
- Motivation: The literature is fragmented across agent types, concepts, and evidential standards, making cross-tradition coverage and underexplored combinations difficult to determine.Human studies emphasize expertise and metacognition, whereas computer-chess research emphasizes search, evaluation, and playing strength.
- Approach: The study uses a systematic mapping across agent type, strategic-reasoning stage, and evaluation dimension rather than synthesizing a common effect size.This creates a common comparison space without assuming that agents use identical reasoning processes or notions of success.
- Contributions: The paper maps research on human cognition, classical computer chess, neural and reinforcement-learning systems, LLMs, and hybrid approaches while distinguishing reasoning components from evaluation methods.It identifies underexplored intersections involving explicit planning, grounded and faithful explanation, metacognitive calibration, and human–AI complementarity.
2 Conceptions of Strategic Reasoning in Chess
Chess conceptions of strategic reasoning developed from positional principles and patterns toward explicit procedures for generating, comparing, and selecting actions. Together, these traditions treat strategic competence as both reusable knowledge and control over how and when reasoning is deployed.
- Historical foundations: Steinitz grounded positional analysis in general ideas and analogies, extending strategy beyond immediate tactical opportunities.Positional features could be related to general maxims guiding later action.
- Historical foundations: Nimzowitsch organized concepts such as blockade, outposts, overprotection, pawn-chain strategy, and prophylaxis into a vocabulary for conscious interpretation and plan construction.The contribution was the deliberate organization of strategically relevant relationships, not merely a list of principles.
- Procedural reasoning: Kotov distinguished candidate generation, calculation, comparison, and commitment through the procedural metaphor of a tree of analysis.This complements positional knowledge by specifying how alternatives are investigated.
- Human and computational reasoning: Shannon’s Type A and Type B strategies frame strategic choice as allocating limited reasoning resources between uniform and selective search.Chase and Simon’s chunking and pattern recognition explain how experts rapidly identify salient configurations that shape further analysis.
- Synthesis: Strategic reasoning combines reusable principles, patterns, and heuristics with processes that determine their relevance and support decision commitment.Competence also includes knowing when to deploy knowledge and when further analysis is warranted.
3 Research Questions
The study asks how recent chess research distributes attention across agent types, reasoning stages, and evaluation practices. Combining these dimensions reveals which intersections are populated and which remain comparatively underexplored.
- Research questions: Four research questions cover the agent landscape, operationalized reasoning processes, evaluation practices, and sparse intersections in the literature.The fourth question shifts analysis from individual studies to the structure of the research field.
- RQ1 — Agent landscape: RQ1 asks which types of agents are studied in recent chess research on strategic reasoning.The framework distinguishes who or what is treated as the reasoning agent.
- RQ2 — Strategic-reasoning processes: RQ2 asks which strategic-reasoning stages are operationalized and evaluated across agent types.This examines which components of the reasoning process are made explicit.
- RQ3 — Evaluation practices: RQ3 asks which evaluation dimensions establish strategic capabilities and how those practices vary across agent types.The question concerns what evidence is accepted as demonstrating strategic competence.
- RQ4 — Research gaps: RQ4 identifies combinations of agent types, reasoning stages, and evaluation dimensions that remain comparatively underexplored.The two principal maps are Agent Type × Strategic-Reasoning Stage and Agent Type × Evaluation Dimension.
4 Method
The study systematically maps recent chess research using study families and multi-label classifications of agents, reasoning stages, and evaluation dimensions. Its 84-family core corpus is designed to expose populated and sparse research intersections while preserving the field’s heterogeneity.
- Study design: The systematic mapping study classifies heterogeneous research rather than estimating a common effect size across incompatible designs and criteria.The map uses agent type, strategic-reasoning stage, and evaluation dimension, with its protocol and dataset supplied in a replication package.
- Record consolidation: Records are consolidated into study families using DOI or database identifiers, normalized-title comparison, and manual inspection before screening.Preprints and published versions of the same contribution are merged, while substantively distinct studies from one project remain separate.
- Eligibility and classification: Eligibility requires chess to be a substantive empirical or computational domain for directly investigating constructs represented in the framework.Relevant conceptual or background work and studies where chess is secondary are retained as Context studies rather than mapped Core studies.
- Corpus construction: The final mapped corpus contains 84 Core study families and 19 Context study families.Context studies support interpretation but are excluded from the quantitative mapping matrices.
- Coding framework: The framework uses independent multi-label dimensions covering agent types, nine reasoning stages, and ten evaluation dimensions.Stages include state representation, assessment, candidate generation, search, planning, evaluation, action selection, explanation, and metacognition; evaluations include legality, move quality, strength, understanding, explanation, faithfulness, human-likeness, calibration, generalization, and synergy.
- Validity: Coding is conservative but remains limited by interpretive judgment and the absence of complete-corpus duplicate coding or an inter-rater agreement statistic.The framework is a mapping instrument, not a claim that strategic reasoning proceeds through nine discrete sequential stages.
5 Results
The 84-family map is concentrated on action selection, assessment, and evaluation, but profiles differ substantially across agent types and evaluation practices. Planning, explanation and faithfulness, metacognition, and human–AI synergy remain comparatively underexplored, while LLM and hybrid studies show distinctive emphases.
- 63 of 84 study families (75.0%) address decision/action selection, compared with 43 (51.2%) for situation assessment and 40 (47.6%) for evaluation/tradeoffs.
- Decision/action selection is the most consistently studied reasoning stage across Human, Classical Engine, Neural/RL, LLM, and Hybrid research.
- Classical Engine studies emphasize search, evaluation, and decision, while Neural/RL studies combine decision-oriented research with learned representations and assessment.
- LLM studies particularly emphasize state representation, with R1 appearing in 20 of 28 studies, alongside situation assessment and decision-making.
- Human studies emphasize decisions, assessment, and evaluation, with reflection/metacognition comparatively prominent but direct strategic-planning operationalization rare.
- Hybrid studies contain more explanation-focused work and commonly use engines, expert models, symbolic rules, knowledge graphs, or taxonomies to ground explanations.
- Classical Engine evaluations center on move quality and playing strength, human evaluations are more heterogeneous, Neural/RL studies combine strength with human-likeness, and LLM studies emphasize legality, state consistency, and generalization.
- The map identifies gaps in planning, explanation quality and faithfulness, and human–AI synergy; no LLM study directly evaluates human–AI synergy, while hybrid studies include seven such evaluations.
6 Discussion
The map separates decision quality from evidence about strategic reasoning and shows that hybrid approaches require finer structural and temporal distinctions. It also identifies state representation, grounded explanation, metacognition, and human–AI evaluation as important unresolved areas.
- 6.1 Decision Quality Is Not Equivalent to Strategic Reasoning: The map shows that strong decision quality does not by itself establish a strategic reasoning process.Behavioural measures can establish whether a move is strong, but process claims require evidence about representation, alternatives, look-ahead, planning, or revision.
- 6.1 Decision Quality Is Not Equivalent to Strategic Reasoning: Internal analyses of computation complement behavioural measures by examining evidence such as future-move representations and causally relevant activations.Ruoss et al. demonstrate strong play without explicit inference-time search, while Jenner et al. identify internal information linked to future moves and recover a move two turns ahead.
- 6.2 State Representation and Generalization: LLM chess evaluation must verify board-state and rule consistency because strategic reasoning cannot operate reliably on an incorrectly represented position.This requirement helps explain the prominence of state tracking and motivates tests of transfer beyond familiar configurations through systematic changes to rules, initial states, information, or board structure.
- 6.3 Hybridization Operates at Different Levels: Diachronic hybridization transfers machine-derived knowledge into later human decisions without requiring a persistent joint decision architecture.AlphaZero-derived concepts were taught to elite grandmasters, whose improved post-learning performance provided evidence of knowledge transfer; this is typically evaluated as augmentation rather than synergy.
- 6.3 Hybridization Operates at Different Levels: Hybridization varies by structural locus and temporal coupling, so the HYB category does not represent one architectural or cognitive paradigm.The framework distinguishes where capabilities combine, whether they operate synchronically or diachronically, and whether agents retain independent agency.
6.4 Grounded Explanation Is Emerging as a Hybrid Capability
Strategic explanation should be distinguished from fluent language generation by checking correctness, usefulness, and faithfulness to supporting evidence. The literature increasingly uses hybrid architectures to connect language models with independently verifiable chess structure.
- 6.4 Grounded Explanation Is Emerging as a Hybrid Capability: Explanation research distinguishes producing an explanation from making it strategically correct, useful, and faithful to the decision process.This distinction matters especially for LLMs, whose linguistic fluency can appear authoritative despite incorrect or weakly grounded chess analysis.
- 6.4 Grounded Explanation Is Emerging as a Hybrid Capability: Hybrid explanation systems use engines, expert models, symbolic components, or structured knowledge to supply independently checkable grounding.These components can provide evaluations, variations, concepts, rules, constraints, and relations that an LLM converts into natural-language explanations.
- 6.5 Metacognition Remains Underexplored: Human metacognitive constructs concern whether further reasoning is worthwhile, evaluations are reliable, and prior conclusions should be reconsidered.Comparable questions remain less developed for artificial agents, although confidence, computation, feedback, and contradictory evidence offer natural chess-based tests.
6.6 From Human Augmentation to Dynamic Complementarity
The paper distinguishes human augmentation from genuine human–AI synergy and frames collaboration as a dynamic relation that changes with competence, task allocation, interaction, and authority.
- Human augmentation and synergy: Human performance can improve with AI assistance without the combined human–AI system outperforming the stronger component alone.Positive augmentation requires Δaug > 0, whereas positive synergy requires the combination to exceed both individual components.
- Dynamic complementarity: Chess collaboration shifted from humans contributing strategic judgment alongside calculation toward human orchestration of increasingly autonomous artificial capabilities.Evidence from conventional, centaur, and engine chess indicates both substitution and complementation as engines strengthened.
- Dynamic complementarity: The dynamic complementarity boundary marks where integration yields positive synergy versus where the stronger component performs at least as well alone.Its position changes with relative competence, subtask allocation, interaction design, and decision authority; contributions can become redundant or detrimental at one level while complementarity emerges elsewhere.
- Human augmentation and synergy: Evaluations claiming synergy should compare human, artificial-agent, and team performance while reporting component competence and decision authority.When no meaningful AI-alone baseline exists, improvement over unaided human performance supports augmentation but not synergy.
6.7 Chess as a Bridge Domain
Chess bridges cognitive and computational approaches by allowing the same strategic decision to be studied through expertise, search, learned representations, language, confidence, and collaboration. Its controlled structure supports richer evaluations beyond playing strength.
- Bridge domain: Chess enables direct comparison of human expertise and metacognition with engine search, learned policies, language-model explanations, and collaborative decision support.A single position can be examined across these perspectives with externally verifiable decision quality.
- Bridge domain: Standard chess is limited to one strategic-environment class, while variants can test uncertainty, information acquisition, altered rules, and transfer.Such variants may separate capabilities tied to familiar chess structures from more general strategic reasoning.
- Evaluation priorities: Chess research should evaluate representation, planning, explanation, monitoring, and collaboration rather than relying only on move quality or game outcomes.Planning needs evidence of look-ahead, persistence, or revision; explanations should be assessed for correctness and faithfulness; confidence should be tested for calibration.
- Evaluation priorities: Four priorities are explicit planning, grounded and faithful explanation, metacognitive calibration, and human–AI complementarity.These priorities target sparse regions of the systematic map rather than constituting an exhaustive agenda.
- Evaluation priorities: Complementarity studies should compare team performance with relevant constituent baselines and examine competence, task allocation, interaction design, and decision authority.These comparisons are needed to investigate how complementarity changes across configurations.
7 Threats to Validity
The mapping is constrained by incomplete literature retrieval, interpretive coding, post-hoc extensions, temporal coverage, and the limited meaning of sparse map regions.
- Search and selection: Interdisciplinary terminology and publication venues make complete retrieval difficult, so relevant studies may remain missing despite complementary database searches.The search combined five Scopus families, Web of Science, and a curated Zotero collection, but may still omit other databases and grey literature.
- Coding and construct validity: The multi-label R1–R9 and E1–E10 frameworks require interpretive coding, although verification removed or refined overly broad assignments.Strong move quality alone was not coded as evidence of search or strategic planning.
- Scope of extensions: The proposed distinctions on hybridization and synergy were not applied as coding dimensions, leaving their prevalence, discriminant value, and reproducibility untested.They remain extensions for future empirical evaluation rather than validated corpus-wide categories.
- Temporal validity: The map covers research beginning in 2023, excluding foundational earlier work from quantitative analysis while publication speeds differ across fields.Rapid LLM preprints and conferences are contrasted with slower cognitive and behavioral journal cycles.
- Interpretation of gaps: Sparse map regions identify underrepresented research combinations but do not establish scientific importance, capability deficits, or field immaturity.Dense cells likewise indicate research attention rather than quality or resolution of a problem.
8 Conclusion
The map finds a rich but uneven chess literature, with LLMs expanding representation and explanation research while key reasoning and collaboration areas remain underdeveloped. It proposes framework extensions and four priorities for using chess to study strategic reasoning.
- Main findings: Research is strongest on assessment, evaluation, and action selection, while explicit planning, explanation, metacognition, and human–AI complementarity remain comparatively underdeveloped.
- Main findings: LLMs expand research toward learned state representations, language-mediated reasoning, robustness, and explanation, but explanation-oriented work is concentrated in hybrid systems.
- Framework extensions: The framework should distinguish hybrid systems by where heterogeneous capabilities combine and whether coupling occurs during decisions or across successive learning and play phases.
- Framework extensions: The framework should also distinguish improved human performance from genuine human–AI synergy because augmentation alone cannot establish synergy.
- Future directions: Future research should prioritize explicit planning, grounded and faithful explanation, metacognitive calibration, and human–AI complementarity.These priorities reposition chess as a controlled environment for studying how agents represent, plan, explain, monitor, and combine decisions.
Statements and Declarations
The authors disclose project funding, no relevant competing interests, open research materials, and reviewed use of generative AI for editing and repository preparation.
- Remo Pareschi received partial support from the Ermete project, funded by Italy’s Ministry of Enterprises and Made in Italy in 2023.
- The authors report no relevant financial or non-financial competing interests.
- The study protocol, search strategies, codebook, canonical study-family dataset, and analysis script are openly available in a permanently archived Zenodo replication package.
- ChatGPT and Claude supported language editing, textual refinement, manuscript organization, and replication-repository preparation, with all outputs reviewed and revised by the authors.