Source-linked AI summary

Practical and Ethical Challenges of Large Language Models in Education: A Systematic Scoping Review

Lixiang Yan, Lele Sha, Linxuan Zhao, Yuheng Li, Roberto Martinez-Maldonado, Guanliang Chen, Xinyu Li, Yueqiao Jin, Dragan Gašević

arXiv:2303.13379v2cs.CLcs.AIcs.CY

TL;DR

Educational LLM innovations may automate laborious text-based tasks, but their practicality and ethicality remain concerns for authentic educational adoption. This paper conducts a systematic scoping review of 118 peer-reviewed papers and identifies 53 educational use cases alongside practical and ethical challenges. It concludes with recommendations focused on model updating, open sourcing, and human-centred development.

  • Problem

    The practical and ethical challenges of LLM-based innovations across educational tasks remain insufficiently understood.

  • Method

    The study systematically reviewed 118 peer-reviewed papers and assessed innovations across practicality and ethicality dimensions.

  • Results

    The review identified 53 LLM application scenarios for automating educational tasks and several practical and ethical challenges.

  • Takeaways & Limitations

    Future studies should update existing innovations, support open-source models and systems, and adopt a human-centred development approach.

  • Takeaways & Limitations

    The review did not assess the quality of included studies, limiting interpretation of findings, particularly performance metrics.

Abstract

from arXiv · show

Educational technology innovations leveraging large language models (LLMs) have shown the potential to automate the laborious process of generating and analysing textual content. While various innovations have been developed to automate a range of educational tasks (e.g., question generation, feedback provision, and essay grading), there are concerns regarding the practicality and ethicality of these innovations. Such concerns may hinder future research and the adoption of LLMs-based innovations in authentic educational contexts. To address this, we conducted a systematic scoping review of 118 peer-reviewed papers published since 2017 to pinpoint the current state of research on using LLMs to automate and support educational tasks. The findings revealed 53 use cases for LLMs in automating education tasks, categorised into nine main categories: profiling/labelling, detection, grading, teaching support, prediction, knowledge representation, feedback, content generation, and recommendation. Additionally, we also identified several practical and ethical challenges, including low technological readiness, lack of replicability and transparency, and insufficient privacy and beneficence considerations. The findings were summarised into three recommendations for future studies, including updating existing innovations with state-of-the-art models (e.g., GPT-3/4), embracing the initiative of open-sourcing models/systems, and adopting a human-centred approach throughout the developmental process. As the intersection of AI and education is continuously evolving, the findings of this study can serve as an essential reference point for researchers, allowing them to leverage the strengths, learn from the limitations, and uncover potential research opportunities enabled by ChatGPT and other generative AI models.

Practitioner notes

LLMs can automate laborious text generation and analysis in education, but practical deployment requires stronger reporting and human-centred development. The section highlights 53 potential educational tasks and recommendations for more practical and ethical innovations.

  • 53 educational tasks could potentially benefit from LLM-based innovations.
  • A structured assessment examined the practicality and ethicality of existing innovations across seven important aspects.
  • Updating existing innovations with state-of-the-art models may reduce manual effort when adapting models to different educational tasks.
  • Improving empirical reporting standards is needed for educational technologies using large language models.
  • A human-centred approach throughout development could help address practical and ethical challenges in education.

1 | INTRODUCTION

LLMs are increasingly used to automate educational text generation and analysis, yet their practical and ethical challenges remain insufficiently understood across educational tasks. This review addresses that gap through a systematic scoping review of 118 peer-reviewed publications.

  • LLMs can complete natural language tasks with little or no additional training, potentially lowering barriers to educational innovation.
  • Educational LLM innovations target tasks including question generation, feedback provision, essay scoring, and administrative recommendations.
  • No prior work had systematically reviewed the practical and ethical challenges of LLM-based innovations across different educational tasks.
  • The review included 118 peer-reviewed publications from four databases using the PRISMA protocol and inductive thematic analysis.
  • Practicality was assessed through technological readiness, model performance, and replicability, while ethicality covered transparency, privacy, equality, and beneficence.
  • The paper contributes a list of 53 educational tasks, a structured seven-aspect assessment, and three recommendations for future research.

2 | BACKGROUND

The background frames practicality as readiness, performance, replicability, and alignment with educational contexts, while ethicality concerns transparency, privacy, equality, and beneficence. Earlier reviews generally addressed individual applications, leaving broader LLM-related challenges insufficiently synthesized.

  • Practicality: The practicality index emphasizes adoption feasibility, cost-benefit ratio, and alignment with existing practices and beliefs.
  • Practicality: Practicality concerns technological readiness, model performance, and methodological replicability in authentic educational contexts.
  • Ethicality: Ethicality is defined as systematizing appropriate and inappropriate functionalities and outcomes of LLM-based innovations.
  • Ethicality: Transparency makes information, decisions, decision-making processes, and assumptions available to stakeholders.
  • Ethicality: The review considers privacy, equality, and beneficence as fundamental ethical issues for educational LLM innovations.
  • Related work: Previous systematic reviews mainly focused on specific applications, while broader practical and ethical issues across LLM-based education remained underexplored.

3 | METHODS

The review used a systematic scoping approach to map LLM applications in education and assess their practicality and ethicality. It followed PRISMA-guided searching, screening, extraction, and thematic analysis.

  • Review design: The review adopted systematic scoping methods to survey an emerging field and identify its concepts, methods, evidence, and challenges.The authors note that study quality was often not assessed because the aim was to provide a broader picture of the field.
  • Search strategy: The search covered Scopus, ACM Digital Library, IEEE Xplore, Web of Science, Google Scholar, and ERIC for English publications from 2017 through 2022.Search terms included large language model, pre-trained language model, GPT, BERT, education, student, and teacher.
  • Screening and selection: Two researchers independently screened eligible studies using predetermined inclusion and exclusion criteria focused on empirical educational applications of LLMs.Studies were excluded when they lacked educational automation, used unspecified LLMs, were non-empirical, non-peer-reviewed, short-format, or non-English.
  • Screening and selection: 854 publications were initially identified, 663 remained after duplicate removal, 197 underwent full-text review, and 118 were selected for data extraction.Inter-rater reliability was substantial during title-and-abstract screening (Cohen’s kappa = 0.75) and full-text review (Cohen’s kappa = 0.73).
  • Data extraction and analysis: The analysis extracted educational tasks, stakeholders, models, and machine-learning tasks, then evaluated technology readiness, performance, replicability, transparency, privacy, and equality.Two researchers independently coded samples, resolved disagreements through discussion or consultation, and cross-checked the remaining studies.

4.1 | The Current State — RQ1

The review found that LLM research already spans many educational automation tasks and stakeholders, but the literature remains concentrated on older models and classification-oriented applications.

  • Educational task coverage: 53 educational use cases were identified across nine categories: profiling/labelling, detection, grading, teaching support, prediction, knowledge representation, feedback, content generation, and recommendation.Examples include student confusion detection, essay grading, conversational agents, knowledge graphs, feedback, MCQs, and course recommendations.
  • Stakeholders: 85 studies targeted teacher-related activities, 54 addressed student activities, 20 supported institutional practices, and 14 empowered researchers.Teacher examples included question grading and generation; student examples included feedback and resource recommendation.
  • Models and tasks: BERT and its variants appeared in 109 studies, while GPT-2 and GPT-3 appeared in five and three studies, respectively.BERT-based models often required manual fine-tuning, reported in 90 studies.
  • Models and tasks: 74 studies investigated classification, compared with 24 generation studies and 23 prediction studies.The review also identified limited use of OpenAI Codex and T5, each appearing in two studies.
  • Models and tasks: Most innovations relied on older models such as BERT and GPT-2, whereas state-of-the-art models such as GPT-3 had not yet been widely applied.The authors suggest commercial, closed-source models may increase development and operating costs.

4.2 | Practical Challenges — RQ2

Practical readiness was limited despite promising task-level performance: most innovations remained early-stage, were insufficiently validated in authentic settings, and were difficult to replicate.

  • Technology readiness: None of the reviewed innovations reached TRL-6, meaning successful operations had not been verified.Some systems reached TRL-5 through integrated components, including an intelligent virtual standard patient and a university-admission chatbot.
  • Technology readiness: The review found no evidence that existing innovations improved teaching, learning, or administrative processes in authentic educational contexts.The authors characterize the systems as capable of automating certain tasks but still early in development and testing.
  • Performance: Prediction and grading results were often strong, including QWK values of 0.80 for off-topic and gibberish essays and 0.94 for paraphrased answers.Automatic short-answer grading correlated with human ratings at Pearson’s correlation values between 0.75 and 0.82.
  • Performance: Generation systems achieved an F1 score of 0.92 for single-word-answer MCQs, while fine-tuned Codex solved 81% of advanced mathematics problems.Other systems generated plausible programming exercises and solutions around three out of four times.
  • Replicability: 107 studies lacked sufficient methodological detail for replication, and only 11 released both code and data publicly.Among the 107 non-replicable studies, 87 required fine-tuning, limiting further evaluation across datasets.

4.3 | Ethical Challenges — RQ3

The review identified ethical weaknesses involving transparency, privacy, equality, beneficence, and fairness. Human stakeholders were rarely involved throughout system development, and ethical reporting was generally limited.

  • Transparency and human involvement: Stakeholders were usually involved only in post-hoc evaluation rather than throughout development, limiting knowledge of system principles and weaknesses.Evaluations included distinguishing AI-generated from human-generated content and assessing student satisfaction.
  • Privacy: Privacy issues were rarely investigated, and studies fine-tuning LLMs on student text did not explicitly explain consent or data-protection strategies.Student language may contain private or sensitive information, including identifying details shared through digital platforms.
  • Equality: Most reviewed studies used LLMs only for English content, although applications also covered Chinese and 11 other languages.The review identified 19 Chinese-language studies and 10 studies covering Vietnamese, Spanish, Italian, and German.
  • Beneficence and fairness: Seven studies discussed possible beneficence violations, and only nine reported descriptive information about sample groups such as gender and ethnicity.The review notes that demographic balancing may improve both model fairness and accuracy.
  • Beneficence and fairness: Accurate models may still produce bias or discrimination, while inaccurate models may harm students’ learning experiences.Suggested mitigations included deferring model decisions and requiring teacher revision before determining correctness.

5 | DISCUSSION

The review identified broad educational applications for LLMs while finding substantial practical and ethical barriers to their use in authentic contexts. It recommends newer models, open systems, transparent reporting, and human-centred development to address these barriers.

  • Use cases: 53 LLM application scenarios were identified across nine categories, including profiling, detection, grading, teaching support, prediction, knowledge representation, feedback, content generation, and recommendation.
  • Model use: 92% of reviewed studies used BERT-based models, indicating that many innovations were developed with older models rather than state-of-the-art LLMs.State-of-the-art models may reduce fine-tuning effort and potentially achieve similar performance through zero-shot approaches.
  • Research focus: 63% of studies focused on classification, while 72% of innovations primarily supported teachers rather than students.
  • Practical challenges: Most innovations had low technological readiness because they had not been fully integrated and validated in authentic educational contexts.The review highlights a need for in-the-wild studies.
  • Ethical challenges: 92% of innovations were transparent only to AI researchers and practitioners, with nine studies reaching transparency for educational technology experts and enthusiasts.The review attributes low transparency primarily to limited human-in-the-loop components.
  • Recommendations: The authors recommend updating innovations with state-of-the-art models, open-sourcing models and systems, and adopting human-centred development with stakeholder involvement.They also call for systematic reporting of model use, prompts, datasets, and implementation details to improve replicability.
  • Limitations: The review is limited because it did not assess included-study quality, excluded non-English and non-peer-reviewed work, and omitted some practical and ethical dimensions.Its transparency index also did not assess transparency to students, and academic integrity was outside scope.

6 | CONCLUSION

The study systematically reviewed educational LLM research and identified practical and ethical challenges that must be addressed for these innovations to become beneficial and impactful. It proposes three recommendations intended to support practical and ethical applications across a wide range of educational tasks.

  • The review identified practical and ethical challenges that need to be addressed for LLM-based innovations to become beneficial and impactful.
  • The authors recommend updating existing innovations with state-of-the-art models, open-sourcing models and systems, and using a human-centred approach throughout development.
  • These recommendations could support innovations that are implemented in authentic contexts to automate a wide range of educational tasks.
Loading 2303.13379v2…