Source-linked AI summary

LLM Agents for Education: Advances and Applications

Zhendong Chu, Shen Wang, Jian Xie, Tinghui Zhu, Yibo Yan, Jinheng Ye, Aoxiao Zhong, Xuming Hu, Jing Liang, Philip S. Yu, Qingsong Wen

arXiv:2503.11733v2cs.CYcs.AIcs.CLcs.HC

TL;DR

Educational LLM agents address limitations of traditional educational data mining by supporting complex pedagogical tasks and personalized learning. This survey synthesizes their technical foundations, task-centered applications, resources, challenges, and domain-specific extensions. It concludes that reliable and ethical deployment requires addressing ethical issues, hallucination and overreliance, ecosystem integration, and resource constraints.

  • Problem

    Traditional educational data mining faces shallow contextual understanding, limited interaction, and difficulty generating adaptive, personalized learning materials.

  • Method

    The survey reviews LLM agents’ technical foundations, educational tasks, datasets, benchmarks, evaluation methodologies, and deployment challenges.

  • Results

    The survey proposes a task-centric taxonomy spanning Teaching Assistance and Student Support and synthesizes applications, capabilities, resources, and deployment challenges.

  • Takeaways & Limitations

    Reliable and ethical educational deployment requires addressing ethical issues, hallucination and overreliance, and integration with existing educational ecosystems.

  • Takeaways & Limitations

    The review may omit some recent advancements because its search covered publications through May 2025.

Abstract

from arXiv · show

Large Language Model (LLM) agents are transforming education by automating complex pedagogical tasks and enhancing both teaching and learning processes. In this survey, we present a systematic review of recent advances in applying LLM agents to address key challenges in educational settings, such as feedback comment generation, curriculum design, etc. We analyze the technologies enabling these agents, including representative datasets, benchmarks, and algorithmic frameworks. Additionally, we highlight key challenges in deploying LLM agents in educational settings, including ethical issues, hallucination and overreliance, and integration with existing educational ecosystems. Beyond the core technical focus, we include in Appendix A a comprehensive overview of domain-specific educational agents, covering areas such as science learning, language learning, and professional development.

1 Introduction

The survey reviews LLM agents for education through a task-centric framework, emphasizing their capabilities, applications, deployment challenges, and research resources.

  • Motivation: Traditional educational data mining still faces shallow contextual understanding, limited interaction, and difficulty generating adaptive, personalized materials.
  • Agent capabilities: LLM agents support education through memory, tool use, planning, personalization, explainability, and multi-agent communication.
  • Survey organization: The survey organizes educational agents into Teaching Assistance and Student Support roles across classroom simulation, feedback generation, curriculum design, adaptive learning, knowledge tracing, and error correction.
  • Challenges: It examines ethical issues, hallucination and overreliance, and integration into existing educational ecosystems as deployment challenges.
  • Resources: The survey compiles datasets and benchmarks to support research on LLM-driven educational solutions.
  • Scope: Appendix A reviews domain-specific agents for science learning, language learning, and professional development.

2 LLM Agents for Education

LLM agents for education combine capabilities such as memory, personalization, explainability, and collaboration with task-specific applications and research taxonomies.

  • Core capabilities: Educational LLM agents combine memory, tool use, planning, personalization, explainability, and multi-agent communication for flexible, learner-tailored support.
  • Applications: Representative research covers feedback comment generation, classroom simulation, professional learning, curriculum design, and related educational applications.
  • Task taxonomy: Table 1 maps primary agent capabilities to educational tasks while omitting secondary capabilities used in some studies.

3 Agent for Teaching Assistance

Teaching-assistance agents support educators through classroom simulation, feedback generation, and curriculum design. These tasks combine capabilities such as memory, planning, tool use, personalization, explainability, and multi-agent communication, while remaining constrained on complex feedback tasks.

  • Overview: Teaching-assistance agents support educators across classroom simulation, feedback comment generation, and curriculum design.Their objectives include improving teaching quality, enriching student learning experiences, and reducing educators’ workload.
  • Classroom Simulation: Classroom simulations model student-teacher dialogues, collaborative activities, and problem-solving so educators can test strategies and assess student reactions.Agents use memory, planning, and multi-agent communication to model fine-grained behavior across diverse personas and learning patterns.
  • Classroom Simulation: Iterative reflection and simulation enable strategy testing for diverse student profiles and can refine teaching plans.These systems reduce educators’ task loads while broadening exploration of student profiles.
  • Feedback Comment Generation: Feedback agents generate automated comments, with multi-agent systems using one agent for drafting and another for evaluation and refinement.The second agent is used to prevent overpraise and excessive inferences.
  • Feedback Comment Generation: Agent-generated feedback still faces challenges on complex programming tasks and professional academic-paper review, where inaccurate feedback may occur.Suggested directions include external tools, stronger memory, and improved personalization.
  • Curriculum Design: Curriculum design relies on memory, tool use, planning, personalization, and explainability to align learning paths with students’ knowledge levels.LLM-based agents support dynamic sequencing and content adaptation through retrieval-based or generation-based methods.
  • Curriculum Design: Future curriculum systems may integrate adaptive learning systems, hybrid retrieval and generation, and multimodal resources to respond to performance and diversify content.The proposed directions include interactive media, video, and immersive simulations.

4 Agent for Student Support

Student-support agents provide real-time, personalized assistance through adaptive learning, knowledge tracing, and error correction. They model learner states and coordinate specialized agents to tailor instruction, while multimodal and complex-task support remain active areas of development.

  • Overview: Student-support agents aim to provide real-time personalized assistance without direct teacher involvement through adaptive and interactive feedback.Core tasks include adaptive learning, knowledge tracing, and error correction and detection.
  • Adaptive Learning: Adaptive learning agents maintain structured student profiles that inform content selection, pacing, and feedback strategies.Profiles can be updated dynamically from student performance and external-tool inputs.
  • Adaptive Learning: Learner representations can encode cognitive, affective, psychological, preference, and personality dimensions for instructional adaptation.Affective modeling supports changes in feedback tone, encouragement, and pacing to maintain engagement.
  • Adaptive Learning: Multi-agent adaptive-learning systems assign specialized roles such as gap identification, learner profiling, simulation, path scheduling, and content creation.These roles support goal-oriented, personalized instruction.
  • Knowledge Tracing: Knowledge-tracing systems use administrators, judgers, and critics to delegate assessment, discuss cognitive states, and evaluate whether criteria are met.This creates a structured yet flexible knowledge-tracing process.
  • Knowledge Tracing: Other approaches trace knowledge through simulated learner personas or dialogue-driven probing of students’ conceptual boundaries.These methods refine knowledge estimation across varied learning profiles and conversational exchanges.
  • Error Correction and Detection: Error-detection agents provide context-aware feedback across writing, programming, and mathematics by adapting responses to learner proficiency.State representations and adaptive inference mechanisms track error patterns and misconceptions dynamically, including from handwritten or digital drafts.

5 Challenges and Future Directions

The survey identifies privacy, bias and fairness, hallucination and overreliance, ecosystem integration, and multimodal deployment as major barriers to reliable and ethical educational LLM agents. It outlines research directions addressing data protection, structured integration, multimodal alignment, and efficient real-time operation.

  • Privacy, Bias and Fairness: Educational LLM agents face privacy risks from sensitive personal data and bias that can reinforce stereotypes and disparities.The survey calls for stronger data protection and bias-mitigation strategies to support equitable learning experiences.
  • Future Directions: Proposed directions include machine unlearning, standardized integration frameworks, robust multimodal fusion, context-aware personalization, and latency-aware optimization.These directions target privacy preservation, practical integration, cultural alignment, and efficient multimodal interaction.
  • Hallucination and Overreliance: Hallucinated information can mislead learners, while overreliance may hinder skill acquisition and reduce engagement with learning materials.Confidently presented false information can contribute to misconceptions.
  • Integration with Existing Educational Ecosystems: Integration remains constrained by limited structured frameworks, difficult validation across real-world settings, and unequal access to AI infrastructure.Project-based learning also lacks guidance frameworks, while scalable deployment must accommodate institutions with different resource levels.
  • Multimodal LLM Agents for Education: Multimodal educational agents struggle with heterogeneous modality integration, cultural and contextual alignment, and low-latency reasoning.Current systems may overfit to text, misinterpret classroom communication, and respond inefficiently across modalities.

6 Conclusion

The survey reviews the technical foundations and educational potential of LLM agents, proposes a task-centric taxonomy, and examines deployment challenges. It also compiles datasets, benchmarks, and evaluation methodologies while emphasizing technical rigor and thoughtful system design for future development.

  • The survey reviews LLM agents’ technical foundations and potential applications in personalized learning, intelligent tutoring, and pedagogical automation.
  • Its task-centric taxonomy organizes agents into Teaching Assistance and Student Support around capabilities including memory augmentation, tool use, planning, and personalization.
  • The survey identifies ethical issues, hallucination and overreliance, and ecosystem integration as challenges for reliable and ethical deployment.
  • It compiles critical datasets, benchmarks, and evaluation methodologies to support continued research on educational LLM agents.

Limitations

The survey may omit some of the newest advances because of its publication timing, despite efforts to include foundational and representative work. Its scope includes LLM agent-based methods while excluding studies focused solely on LLM-based approaches, and it covers domain-specific applications.

  • Some of the most recent advances may not have been captured at the time of writing.The authors nevertheless report efforts to include foundational and representative works.
  • The review includes studies focusing on LLM agent-based methods.
  • Studies focused solely on LLM-based approaches were excluded.
  • The survey also examines domain-specific agents in science learning, language learning, and professional development.

A.1 Agent for Science Learning

Science-learning agents use LLMs to support personalized, interactive acquisition and application of scientific knowledge. The survey covers applications across mathematics, physics, chemistry, biology, and broader scientific discovery, including tool use, multi-agent reasoning, simulation, and hypothesis generation.

  • Science-learning agents provide personalized, interactive support for acquiring and applying scientific knowledge.Their reported roles include tailored feedback, conceptual understanding, and active engagement with complex scientific ideas.
  • Mathematics: Mathematics agents address complex problems through tool-integrated reasoning and multi-agent collaboration.TORA integrates tools for mathematical problem-solving, while MACM uses Thinker, Judge, and Executor agents for complex logical deduction.
  • Physics: Physics agents support conceptual learning and interactive simulation of physical phenomena.The survey describes NEWTON, Physics Reasoner’s knowledge-augmented pipeline, and a calculus-based physics course case study.
  • Scientific Discovery: Scientific generative agents can generate and revise hypotheses through bilevel optimization and exploit-and-explore strategies.

A.1.3 Chemistry

The survey presents chemistry-focused LLM agents that support interactive explanation, autonomous chemical synthesis, and more rigorous experimentation through specialized reasoning, grounding, and coordination modules.

  • Chemistry agents explain molecular structures, chemical reactions, and experimental processes interactively for educational use.
  • ChemCrow performs autonomous planning and execution of chemical syntheses, including an insect repellent and three organocatalysts.
  • ChemAgent improves on ChemCrow by emphasizing reasoning and grounding as core cognitive abilities for chemistry problem-solving.
  • Curie introduces intraagent rigor, interagent rigor, and experiment knowledge modules to improve reliability, systematic control, and interoperability.
  • Biology-oriented multi-agent systems extend related scientific learning and discovery capabilities through protein analysis, design, mutation analysis, inverse folding, and visualization.

A.3 Agent for Professional Development

LLM agents for professional development provide scalable, adaptive, and context-aware learning experiences tailored to domain-specific needs, including medical, computer science, and law education.

  • Professional-development agents target medical, computer science, and law education with domain-specific learning experiences.
  • Healthcare agents support personalized, interactive, and scalable systems, while LLMs can help create curricula, adaptive learning plans, and dynamic assessments for medical education.

A.3.2 Computer Science Education

The survey describes LLM agents for computer science education as tools for personalized coding guidance, debugging, and conceptual learning, alongside legal-education agents for interpretation, simulation, and case analysis.

  • Computer Science Education: Computer-science agents provide personalized guidance on coding, debugging, and understanding computer-science principles.
  • Computer Science Education: CodeAgent supports repository-level code generation by incorporating external tools such as WebSearch and DocSearch.
  • Law Education: Legal-education agents support judicial interpretation, moot-court simulation, and case analysis through pretrained legal knowledge, interaction, and reasoning.
  • Law Education: LawLuo uses multi-agent retrieval-augmented generation to simulate multi-turn legal consultations and improve personalization and ambiguity handling.
  • Law Education: AgentCourt, AgentsCourt, and AgentsBench extend training through courtroom interaction, judicial decision-making, multi-agent legal reasoning, and case analysis.

B Datasets & Benchmarks

The survey compiles datasets and benchmarks for evaluating educational LLM agents across pedagogical tasks and specialized domains, organizing resources by goals, users, subjects, levels, languages, modalities, and sizes.

  • Table 2 summarizes publicly available datasets and benchmarks for evaluating educational LLM agents across multiple domains.
  • The resource taxonomy categorizes datasets by primary goal, target users, subject domain, education level, language, modality, and dataset size.
  • ASSIST09 and Junyi support knowledge tracing in K-12 mathematics, EduAgent supports adaptive learning, and Virtual Teacher and MathCCS assess error correction and detection.
  • ScienceAgentBench and TheoremExplainBench evaluate scientific reasoning and theorem explanation, while ML-Bench and MLAgentBench focus on machine-learning education.
  • LawBench, LegalBench, and AgentCourt evaluate legal knowledge application, case analysis, and court simulations; MedBench and OmniMedVQA test clinical reasoning and medical knowledge retrieval.
Loading 2503.11733v2…