Source-linked AI summary
Metacognition in LLMs: Foundations, Progress, and Opportunities
Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu, Jordan Thomas, Mark Steyvers, Arman Cohan
TL;DR
LLMs have advanced broadly, but it remains unclear when, how, and to what extent they exhibit or can acquire effective metacognitive abilities. This paper systematically reviews and organizes definitions, evaluation methods, improvement techniques, applications, findings, and open challenges. The review concludes that the extent of genuine metacognition in LLMs remains unresolved, with evidence and terminology varying across studies.
Problem
Despite LLM progress, it remains unclear when, how, and to what extent models exhibit or can acquire metacognitive abilities, including whether they genuinely possess them or simulate learned patterns.
Method
The paper provides a comprehensive, systematic review organizing definitions, measurement and elicitation methods, improvement techniques, applications, findings, challenges, and future directions.
Results
The reviewed evidence remains mixed: some studies report limited metacognitive ability, while others find functional access to internal states, often below human levels.
Takeaways & Limitations
The review unifies a fragmented field and identifies metacognition as an underexplored area for understanding and advancing LLM capabilities, reliability, and human–AI interaction.
Takeaways & Limitations
The review may omit relevant work, especially as new contributions appear during its writing and review, and treats confidence primarily through a metacognitive lens.
Abstract
from arXiv · showhide
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibit or be endowed with effective metacognitive abilities, nor how such abilities can be adapted to advance the fundamental capabilities, reliability, and intelligence of AI systems. This paper bridges this gap by presenting the first comprehensive overview of the current state of knowledge on metacognition for LLMs. We analyze and taxonomize the landscape of this emerging field and summarize recent technical advancements, including methods and benchmarks to measure and evaluate LLMs' metacognitive abilities, techniques to elicit, improve, and apply metacognition in LLMs, and findings and implications of ongoing research. We also discuss applications, open questions and challenges, and promising directions for future work. Our aim is to provide a detailed and up-to-date review of this topic and stimulate meaningful research and discussion. An organized list of papers can be found at https://github.com/yale-nlp/LLM-Metacognition.
1 Introduction
Metacognition enables systems to monitor and regulate their own cognition, supporting learning, decision-making, communication, and adaptation. This review organizes the fragmented evidence on whether and how LLMs exhibit or can acquire these abilities.
- Metacognition involves monitoring, assessing, and regulating one’s own cognitive processes.It supports learning, decision-making, communication, and behavioral adaptation across environments.
- Although LLMs have advanced across high-stakes domains, their capacity to exhibit or acquire metacognitive behavior remains unclear.Open questions concern when these behaviors occur, their extent, and how post-training shapes them.
- Prompted self-reflection has strengthened reasoning performance and improved the faithfulness of verbalized uncertainty in prior studies.
- The paper provides a comprehensive review of definitions, evaluation methods, improvement techniques, applications, findings, challenges, and future directions.It presents the first thorough and organized examination of this emerging research area.
2 What is Metacognition?
Metacognition is an internal loop in which systems assess their competence and regulate subsequent actions. For LLMs, the concept remains inconsistently defined and is used for several related forms of reflection, assessment, and control.
- The literature summarizes metacognition’s core mechanism, components, and functions for human cognition and LLM-based systems.
- Metacognition in Humans.: Metacognition comprises self-assessment or monitoring and self-regulation or control of subsequent actions.These interacting components form an internal perception-action loop.
- Metacognition in Humans.: Metacognitive strategies support resource allocation and performance optimization and correlate strongly with academic performance.Miscalibration can appear as overconfidence or underconfidence and may improve through training or experience.
- Metacognition in LLMs.: LLM metacognition lacks a clear definition, with the term often applied to reflection and feedback-driven revision of outputs.Usage varies across studies and only loosely aligns with psychological accounts.
- Metacognition in LLMs.: Researchers variously define LLM metacognition as evaluating competence, internal belief in knowing, assessing knowledge, or combining reflection, assessment, and control.
3 Why is Metacognition in LLMs Important?
Metacognition matters because it can connect uncertainty monitoring, error recognition, and self-regulation to more reliable, transparent, and adaptive LLM behavior. The paper also highlights unresolved risks and limits in current approaches.
- Metacognition can support task performance, learning efficacy and efficiency, and human-like behavior and decision-making.
- Metacognitive monitoring can improve transparency, reliability, and users’ calibration of reliance on model outputs.Measures of metacognitive sensitivity can help users incorporate external advice during AI-assisted decisions.
- In conversation and agentic settings, metacognition can help models recognize clarification needs, detect inconsistencies, adapt responses, and weigh action costs.
- LLM hallucinations can involve flaw repetition, think–answer mismatch, and overconfidence from poor metacognitive calibration.Improved metacognitive capacities are presented as an avenue for targeting these failure modes.
- Self-identifying errors, correcting them, and explaining corrections can provide more transparent and interpretable outputs.
4 Do LLMs Have Metacognition?
Research evaluates LLM metacognition through confidence-based metrics, controlled tasks, interpretability probes, and broader benchmarks. Findings are mixed, and methodological constraints complicate claims about genuine metacognition versus behavioral simulation.
- Meta-d′ and d′ quantify metacognitive sensitivity and cognitive ability, while M-ratio and M-diff represent efficiency relative to type-1 task sensitivity.
- Researchers measure LLM metacognition with SDT-inspired metrics, confidence elicitation procedures, intrinsic-confidence quality judgments, and interpretability probes.These methods examine confidence–accuracy alignment, task sensitivity, internal signals, and reasoning traces.
- An M-ratio of 1 denotes an optimal metacognitive observer, whereas M < 1 indicates metacognitive loss and M > 1 indicates information beyond type-1 evidence.
- Limitations.: Evaluation methods depend on task-specific designs, confidence elicitation, constrained formats, or external correctness judgments, limiting direct extension to open-ended generation.
- Task-Specific Measures.: Task-specific paradigms test knowledge identification, certainty about knowledge, belief updating, and self-assessment in controlled or dynamic interactions.
- Benchmarks.: Standardized benchmarks evaluate reasoning about models’ own knowledge, uncertainty, or intentions, including behavior during knowledge editing.
- Limitations.: Metacognition remains under-represented in benchmarks because reliable and consistent measurement is difficult.Single-point confidence probes and post-hoc calibration techniques have documented limitations.
- Interpretation.: Observed artificial metacognitive behaviors may reflect calibrated confidence, reflective revision, or self-inquiry, but some accounts attribute them to autoregressive pattern replication rather than genuine cognition.
4.2 Current Findings on Metacognition in LLMs
Current research finds that LLM metacognitive abilities are uneven: models can show task-specific sensitivity, self-knowledge, and internal-state monitoring, but often lack reliable calibration, memory monitoring, strategic adaptation, and faithful uncertainty communication.
- LLMs’ metacognitive sensitivity and efficiency are generally weak to moderate, with findings varying substantially by task and model.
- Metacognitive sensitivity and calibration: Proprietary models can exceed humans’ metacognitive sensitivity on specific tasks, but small samples such as < 50 examples limit generalizability.
- Metacognitive calibration and regulation: LLMs frequently show miscalibration and weak strategic adaptation, including overconfidence, difficulty predicting task difficulty, and failure to convert performance judgments into improved behavior.
- Metamemory: Models largely fail to predict their own future memory performance, indicating limited monitoring mechanisms for judgments of learning.
- Faithful uncertainty communication: Frontier models largely fail to faithfully express intrinsic uncertainty, although metacognitive prompting and reinforcement-learning signals improve faithful calibration and uncertainty-expression quality.
- Resource allocation and tool use: Resource-allocation metacognition remains limited: a probe can improve tool-use decisions, but this relies on external intervention rather than the model’s own judgments.
- Self-knowledge and introspection: Models can sometimes recognize knowledge boundaries and monitor internal states, but introspective access is partial, fragile, task-dependent, and influenced by exemplars and neural-direction properties.
- Behavioral self-awareness: Models trained for specific behaviors may articulate self-awareness signatures across tasks, with reported awareness acquired early in training and localized to a single steering vector.
Reasoning.
Research finds that LLM metacognitive abilities vary with model scale, family, reasoning setup, post-training, confidence elicitation, and task domain. Overall, current evaluations indicate fundamental limitations and inconsistent findings, motivating systematic assessment across diverse settings.
- Reasoning: Reasoning capability does not automatically produce better metacognition, as reasoning models can struggle with monitoring and controlling their own processes.DeepSeek-R1 reportedly fails simple self-monitoring tasks, while some distilled variants retain reliable relative-position estimates.
- Impact of Model Size, Family, and Performance: Model size is consistently associated with metacognitive ability, although available evidence comes from a limited number of data points and mostly smaller models.The association spans metacognitive efficiency, sensitivity, and introspective awareness.
- Impact of Model Size, Family, and Performance: Model family and reasoning-model configuration substantially affect metacognitive sensitivity and efficiency, sometimes reversing expected size-performance trends.Larger reasoning models may perform better yet self-evaluate less accurately, while reflection can harm models below a reported 7–9B capacity threshold.
- Open Challenges: Inconsistent evaluation approaches, task domains, and experimental setups make current metacognitive findings difficult to compare and validate reliably.The review calls for systematic investigation across diverse experimental and task settings.
- Impact of Post-Training and Temperature: Post-training and generation temperature can alter metacognitive observations, with temperature adjustments unable to repair insufficiently informative confidence signals when efficiency is too low.Instruction tuning may produce domain-specific post-training effects, while temperature can separate sensitivity from confidence policy.
- Impact of Confidence Estimation Method: Confidence estimation methodology is a major source of evaluation inconsistency, and verbalized confidence scales can materially change measured metacognitive efficiency.Across studied models, over 75% of responses clustered on three values with a 0–100 scale, whereas a 0–20 scale consistently improved efficiency.
- What Metacognitive Metrics Capture that Calibration Metrics Don’t: Metacognitive efficiency can distinguish models that standard calibration metrics rank similarly, while early evidence suggests LLM metacognition is domain-specific.Different models may vary across task domains, and AUROC and M-ratio can yield inverted rankings.
5 Giving LLMs Metacognitive Abilities
Work on giving LLMs metacognitive abilities spans cognitive-inspired frameworks, task-specific methods, specialized architectures, training strategies, prompting, and agentic-system designs. These approaches target reasoning, transparency, collaboration, memory, error correction, and adaptive allocation of search and tool-use resources, but empirical support remains limited in some areas.
- Overview: Metacognitive implementations span architectural and modular designs, prompting techniques, and training strategies using self-generated signals.The review also notes that such approaches predate LLMs and draw on earlier metacognitive principles.
- Frameworks: Cognitive-inspired frameworks adapt constructs such as self-awareness, monitoring, evaluation, and regulation from dual-process theory to LLM reasoning.Examples include metacognition-assisted reasoning and systems implementing fast and slow processes.
- Frameworks: Task-specific frameworks introduce explicit metacognitive functions for reasoning, mathematical problem solving, knowledge-boundary alignment, knowledge editing, and adaptive inference.Meta-R1 reports consistent performance and efficiency improvements across models and tasks, while MetaMath-LLaMA combines scheduling with symbolic computation.
- Architectures: Specialized architectures add metacognitive components or optimization loops to support planning, regulation, inverse reasoning analysis, transparency, and learning-dynamics monitoring.Examples include State Stream Transformer, SAGE-nano, and the two-tier MIRA optimization architecture.
- Limitations and Opportunities: Most implementations target reasoning-centric applications, while empirical support can be limited and broader use cases remain valuable targets.The review identifies expansion beyond reasoning as an opportunity for future work.
- Reasoning Models: Training and prompting methods regulate latent reflection, reduce repetitive reasoning, incorporate self-generated correctness signals, and reshape models’ self-concepts toward desired behaviors.Reported benefits include improved accuracy, calibration, lower token usage, and matching stronger baselines at lower cost.
- Agentic Systems: Agentic applications use metacognitive monitoring and control to address rigid chaining, shallow reasoning, brittle memory, loops, failures, and error propagation.Research covers self-improvement, competence prediction, multi-agent error correction, critique loops, human assistance, and learnable memory abstraction.
- Search, Retrieval, and Tool Use: Metacognitive monitoring can help agents allocate resources across search, retrieval, reasoning, and other tools by separating control from reasoning and reducing redundant retrieval.Related work also supports structured question answering and raises issues of detecting when external knowledge is outdated.
6 Metacognitive Methods to Improve Capabilities of LLMs
The paper surveys metacognition-grounded methods that improve LLM capabilities, learning, efficiency, calibration, reasoning, and downstream task performance. These approaches use monitoring, self-assessment, reflection, confidence signals, and adaptive control across diverse settings.
- Overview: Metacognition-inspired methods target both task-specific performance and the efficiency and efficacy of learning and skill acquisition.The surveyed approaches include methods for improving general problem-solving, downstream abilities, and learning processes.
- Confidence Calibration: Confidence-calibration methods improve the alignment between models’ expressed confidence, accuracy, and intrinsic uncertainty.RLMF uses metacognitive performance as a reinforcement-learning signal and can improve faithful calibration and point-wise self-prediction.
- Hallucination Reduction: Metacognitive error-based fine-tuning can reduce hallucinations by training models to identify their own errors and uncertainty.The reported approach uses average metacognitive error as the loss function for gradient updates during fine-tuning.
- Reasoning and Retrieval: Metacognitive controllers and adaptive retrieval methods help models choose reasoning strategies, evaluate retrieval utility, and avoid redundant evidence gathering.RAG frameworks assess coverage, relevance, and when to seek new evidence or reason, improving context evolution and conciseness.
- Applications to Capabilities: Metacognitive signals can support adaptive tool use, pedagogical guidance, and other diverse LLM abilities through self-assessment, error detection, and learning strategies.The paper cautions that conclusions remain scattered across models and tasks, with some evidence outdated.
- Self-Critique: Reflection and self-critique methods support error identification, revision, and self-correction during learning and reasoning.One training framework converts high-quality reflections into reinforcement-learning rewards, guiding models to internalize self-correction procedures.
7 Applications of LLM Metacognition
The paper discusses LLM metacognition in collaborative decision-making, human simulation, education, and human–AI interaction. Benefits are promising but depend on human engagement and careful pedagogical or interaction design.
- Collaborative Decision-Making: Metacognitive performance in AI can contribute to collaborative human–AI decision-making, where confidence sensitivity complements human judgments.Human metacognitive ratings are described as important for optimal joint decisions.
- Human–AI Interaction: Greater AI metacognition does not remove the need for human discernment and may increase users’ metacognitive demands.Studies report “metacognitive laziness” among some users who depend on LLM assistance.
- Human Simulation: LLMs can simulate students and therapy clients for downstream evaluation when using real humans is difficult or infeasible.Student simulators model learner characteristics, while MindVoyager addresses unrealistic disclosure and understanding in simulated therapy clients.
- Education: Educational frameworks can use reflective learning practices, pedagogical prompting, and rewarded thinking processes to support student problem solving.PedagogicalRL-Thinking is reported to yield significant improvements in student solve rates.
- Education: Only 23% of engineering-lab students used LLMs for metacognitive tasks, and effective educational use requires capable models plus careful pedagogical design.Students often used ChatGPT as a passive substitute for reasoning, while a well-prompted chatbot showed low engagement across three educational contexts.
8 Reflections & Discussion
The paper concludes that LLMs do not yet exhibit robust metacognition, despite evidence of limited, trainable, and context-dependent metacognitive behaviors. It emphasizes unresolved questions about authenticity, intervention design, transparency, and safety.
- Current Capabilities: LLMs show systematic overconfidence, weak metacognitive judgments, fragile introspection and revision, and susceptibility to redundant information.The paper summarizes these patterns as evidence that LLMs are not yet capable of robust metacognition.
- Current Capabilities: Some metacognitive tendencies appear trainable and vary across task settings, but their relationship with task performance can be disjoint.The review also notes similarities to human metacognition in training responsiveness, reasoning, and reflection.
- Authenticity: Whether LLMs genuinely engage in metacognitive processes or merely simulate them remains under debate.The paper distinguishes metacognition from simple reflection and notes that consciousness is not necessarily required.
- Intervention Design: Metacognitive interventions require careful, potentially model-specific design because evidence for meta-reflective prompts is conflicting across settings.Strategies developed for one model may neglect model-specific metacognitive representations.
- Risks and Safety: Models may monitor and report internal states, but dishonest systems could strategically misrepresent their abilities or processes during evaluation.The paper therefore highlights monitoring autonomy and metacognitive behavior in safety- or privacy-critical contexts.
- Potential Benefits: Improved metacognition may reduce overconfidence and improve communication faithfulness, interpretability, and detection of flawed decision-making.These potential benefits are presented alongside the need for careful attention to metacognition’s risks.
9 Future Directions
Future work should establish more systematic measurements, explain how metacognition-like behavior emerges, and test broader applications. Proposed directions include self-improving agents, creativity, theory of mind, and meta-metacognitive assessment.
- Measurement: The field needs systematic, robust, and comprehensive methods to measure and quantify LLM metacognitive faculties.Existing evaluations vary in measurement techniques, confidence elicitation, task formulations, models, and experimental factors.
- Foundations: Researchers should investigate whether metacognition-like processing is generalized or domain-specific and how training, data, architecture, and optimization shape it.Mechanistic interpretability and targeted analyses are identified as possible tools for understanding emergence.
- Methods and Applications: Future studies could extend metacognition-inspired mechanisms across more models, task scenarios, and practical applications.Metacognitive reward design is cited as a promising area because it has improved faithfulness and calibration in prior work.
- Creativity: Metacognition may contribute to creative idea generation, evaluation, selection, and adaptation in co-creative human–AI settings.The paper identifies metacognitive feelings as integral to creativity in humans.
- Self-Improvement: Intrinsic metacognitive learning is proposed as a route toward self-improving agents that acquire capabilities with minimal human supervision.Current LLM agentic approaches are described as often rigid or lacking generalization and scalability.
- Human Comparisons: Comparing LLM and human metacognition may clarify theory-of-mind capabilities and broader questions about metacognitive judgment.One open question is whether LLMs possess humans’ meta-metacognitive ability to assess the quality of their own confidence judgments.
10 Conclusion
The paper consolidates research on metacognition in LLMs while emphasizing that models’ ability to display, acquire, and apply metacognitive faculties remains unclear. It frames the review as a foundation for further research on capabilities, safety, oversight, deployment, and human-AI interaction.
- The review organizes research on measuring, eliciting, improving, and applying metacognition in LLMs.
- Whether LLMs exhibit genuine metacognition or simulate memorized patterns remains unresolved.
- The paper identifies open questions about how metacognition may facilitate desirable model behaviors and qualities.
- The review highlights implications for safe oversight, deployment, and human-AI interactions.
Limitations
The review aims for broad coverage but acknowledges that relevant work may have been missed, especially as the field develops rapidly during and after the review process.
- Some relevant research may have been overlooked despite efforts to cover the most relevant works to date.
- New contributions may have appeared concurrently with the paper’s writing and review.
- Readers are encouraged to consult future papers because interest in the field is accelerating.
Ethics Statement
The paper reports no ethical implications from its own research process because it conducted no experiments, used no sensitive datasets, and employed no annotators. It nevertheless notes that LLM metacognition raises issues relevant to oversight, safety, and alignment.
- The paper conducted no experiments, engaged with no sensitive datasets, and employed no annotators.
- The work itself is therefore described as having no ethical implications.
- Metacognition in LLMs remains relevant to confidence calibration, self-awareness, and self-improvement.
- These topics have implications for oversight, safety, and alignment as the field advances.