Source-linked AI summary

"Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education

Emma Harvey, Allison Koenecke, Rene F. Kizilcec

arXiv:2502.14592v1cs.CY

TL;DR

LLM-based edtech is expanding, but its downstream effects remain understudied and existing risk taxonomies are not tailored to education’s distinctive population and goals. Through semi-structured interviews with six edtech providers and 23 educators, the paper develops an education-specific account of harms and finds that providers emphasize technical harms while educators emphasize broader interactional impacts.

  • Problem

    The downstream impacts of LLM-based edtech remain understudied, while existing harm taxonomies are not tailored to education.

  • Method

    The study uses semi-structured interviews and thematic analysis with six edtech providers and 23 educators.

  • Results

    Providers primarily focus on harms measurable from LLM outputs, whereas educators are more concerned about harms requiring observation of interactions among students, educators, school systems, and edtech.

  • Takeaways & Limitations

    The paper develops an education-specific overview of LLM harms and recommends centering educators in edtech design and development.

  • Takeaways & Limitations

    The provider findings should not be broadly generalized because the sample includes only six edtech employees.

Abstract

from arXiv · show

Education technologies (edtech) are increasingly incorporating new features built on large language models (LLMs), with the goals of enriching the processes of teaching and learning and ultimately improving learning outcomes. However, the potential downstream impacts of LLM-based edtech remain understudied. Prior attempts to map the risks of LLMs have not been tailored to education specifically, even though it is a unique domain in many respects: from its population (students are often children, who can be especially impacted by technology) to its goals (providing the correct answer may be less important for learners than understanding how to arrive at an answer) to its implications for higher-order skills that generalize across contexts (e.g., critical thinking and collaboration). We conducted semi-structured interviews with six edtech providers representing leaders in the K-12 space, as well as a diverse group of 23 educators with varying levels of experience with LLM-based edtech. Through a thematic analysis, we explored how each group is anticipating, observing, and accounting for potential harms from LLMs in education. We find that, while edtech providers focus primarily on mitigating technical harms, i.e., those that can be measured based solely on LLM outputs themselves, educators are more concerned about harms that result from the broader impacts of LLMs, i.e., those that require observation of interactions between students, educators, school systems, and edtech to measure. Overall, we (1) develop an education-specific overview of potential harms from LLMs, (2) highlight gaps between conceptions of harm by edtech providers and those by educators, and (3) make recommendations to facilitate the centering of educators in the design and development of edtech tools.

1 Introduction

LLM-based edtech is increasingly used in classrooms, but its downstream effects remain understudied and prior findings on engagement and learning outcomes conflict. Because education involves children, learning goals beyond prediction, and direct student–educator interactions, research on LLM harms must be tailored to education.

  • The downstream impacts of LLM use in education remain understudied, with conflicting findings on student engagement and learning outcomes.
  • Education is not simply a prediction problem because learning, especially critical thinking and social skills, is neither automatable nor directly observable.
  • K-12 education involves children, creating additional privacy and ethical considerations for LLM-based edtech.
  • Domain-agnostic risk research therefore requires adaptation to education to have practical value.
  • LLM-based edtech harms must be studied in the context of interactions among students, educators, school systems, and technology.

Contributions.

The study combines interviews with six edtech providers and 23 educators to develop an education-specific account of LLM harms. It finds a mismatch between providers’ focus on technical harms and educators’ concern about broader impacts that are harder to measure and mitigate.

  • Interviews with six edtech providers and 23 educators examined how each group anticipates, observes, and accounts for LLM harms in education.
  • The study identifies technical, human–LLM interaction, and broader-impact harms from LLM-based edtech.
  • Technical harms include toxic or biased content, privacy violations, and hallucinations, while interaction harms include academic dishonesty.
  • Broader impacts include inhibited learning and social development, increased educator workload, reduced educator autonomy, and exacerbated systemic inequalities.
  • Providers primarily focus on technical harms, whereas educators are more concerned about broader harms that providers currently cannot measure or mitigate.
  • The paper recommends educator-centered design and identifies responsibilities for school leaders, regulators, and researchers when design alone cannot mitigate harms.

2 Background and Related Work

AI-powered edtech is expanding from established educational technologies, with LLMs enabling language-generation capabilities alongside significant risks. Existing LLM harm taxonomies provide a starting point but are domain agnostic and require adaptation to education.

  • AI advances have increased interest in AI-powered edtech intended to improve teaching processes and learning outcomes.
  • LLMs are trained on massive text datasets to predict the next token and can generate fluent text through sequential token prediction.
  • Prior LLM risk taxonomies synthesize broad potential harms but are not tailored to educational contexts.
  • The paper investigates how edtech providers and educators anticipate, measure, and mitigate LLM harms in education.
  • Existing work on AI in education addresses ethical concerns, providing a foundation for examining harms from LLM-based edtech.

AI in Edtech.

AI in edtech can personalize learning and reduce teacher workload, but existing evidence and responsible-AI frameworks also identify risks involving privacy, discrimination, inaccurate instruction, autonomy, and equity. The paper responds by examining harms with both providers and educators and by emphasizing trust, co-design, and educator-centered development.

  • Observed risks: Prior research reports discrimination, privacy and autonomy harms, and inaccurate instruction from AI-based edtech.
  • Potential benefits and frameworks: AI tools may personalize instruction and support teachers through feedback, administrative assistance, lesson planning, and grading.
  • Potential benefits and frameworks: AI-in-education frameworks emphasize human involvement, evidence-based pedagogy, privacy, explainability, nondiscrimination, learner autonomy, and equity.
  • Gaps in existing guidance: Existing frameworks do not fully address LLM-specific risks such as hallucinations and academic dishonesty.
  • Gaps in existing guidance: Some prior work has called for pausing LLM-based edtech adoption until risks are better understood and responsible-AI frameworks are established.
  • Study motivation: The paper examines how providers and educators understand potential harms to support trust-building and educator-centered co-design.

3 Method

The study used 29 semi-structured interviews to examine how participants anticipate, measure, and mitigate harms from LLM-based edtech. Recruitment sought diversity across educator roles and subjects, and interviews were conducted under standard consent and ethics procedures.

  • Study design: The researchers conducted 29 semi-structured interviews between November 2023 and February 2024.Interviews lasted 30–60 minutes and were conducted and recorded via Zoom.
  • Recruitment: Educator recruitment used professional networks, listservs, and snowball sampling to broaden diversity across roles and academic subjects.Snowball sampling was added after STEM teachers were over-represented in the initial recruitment.
  • Research procedures: Educators received $50 compensation, participants provided informed consent, and the study protocol was deemed exempt by the institutional review board.

Participants: Edtech Providers.

The provider sample comprised six distinct edtech organizations representing international leaders and varied organizational types. Their products served substantial student populations and incorporated several LLM-based functionalities, but the sample was not representative of the broader market.

  • Sample: Six individuals employed by distinct edtech providers participated in the study.The providers represented international leaders, with each product reporting at least 10,000 active student users.
  • Product functionality: Their tools supported live chat or feedback, content pre-generation, and report creation for educators or moderators.The sample therefore covered multiple forms of LLM-based functionality in edtech products.
  • Provider diversity: The providers included for-profit, nonprofit, general-purpose, and STEM-focused organizations.Participants held leadership or development roles and had research backgrounds.
  • Scope: The six-employee provider sample was too small to represent the large and diverse edtech market.The authors avoid broad generalizations, while noting that participants came from widely used and well-regarded providers.

Participants: Educators.

The educator sample included 23 professionals spanning classroom, support, counseling, and administrative roles. Participants varied in their prior LLM experience, enabling perspectives across different levels of familiarity with LLM-based edtech.

  • Sample: The study interviewed 23 educators, including teachers, instructional support staff, guidance counselors, and school administrators.These roles broadened the sample beyond classroom teachers alone.
  • Participant documentation: The educators interviewed were identified individually in the study and described in a participant-background table with fuller profiles in an appendix.
  • LLM experience: Educators reported no, limited, regular, or significant prior use of LLMs.Significant use included advanced features such as creating custom chatbots.

Appendix B.

The study combined diverse educator interviews with provider interviews and inductive-deductive thematic coding to examine anticipated and experienced harms. Findings distinguish technical harms from interactional and learning-related concerns, while emphasizing limits on generalization.

  • Sample limitations: The educator sample was diverse but not representative, with prior LLM experience and positive opinions over-represented.White people, men, and STEM teachers were also over-represented despite efforts to recruit differing perspectives.
  • Interview procedure: Interviews began with open-ended questions, then probed domain-agnostic harms and participants’ measurement or mitigation efforts.The researchers used this sequence to identify salient concerns before testing correspondence with existing frameworks.
  • Analysis: Five consecutive interviews without a newly raised harm defined saturation for the study.All provider harms were proactively raised by the first provider participant, while almost all educator harms appeared within the first four educator interviews.
  • Analysis: The coding process combined pre-identified harm categories with inductively identified harms raised by participants.Codes covered whether harms were prompted, concern and experience levels, and mitigation strategies.
  • Provider findings: Providers primarily addressed technical harms measured from LLM outputs, whereas most did not explore harms arising from human–LLM interactions.The technical harms included toxic or biased content, privacy violations, and hallucinations.
  • Learning-related findings: Most providers reported trying to prevent LLMs from inhibiting student learning, but they struggled to account for outputs that diverged from optimal pedagogy.Providers described LLM outputs as potentially reproducing an average internet instructional style rather than research-supported presentation choices.

5 How Educators Account for LLM Harms

Educators reported awareness of technical and human–LLM harms, but were especially concerned about broader effects on learning, social development, workload, autonomy, and inequality. Their confidence in mitigating harms often depended on remaining able to mediate students’ interactions with LLMs.

  • Most educators reported awareness of and efforts to mitigate technical harms, including toxic or biased content, privacy violations, and hallucinations.
  • Educators generally felt able to mitigate technical and human–LLM interaction harms, but were more concerned when LLMs disrupted the teacher–student relationship.
  • Educators were less certain about broader impacts of LLMs on education than about technical harms.
  • Educators proactively raised concerns that LLMs could inhibit learning and social development, increase workload while reducing autonomy, and exacerbate systemic inequalities.
  • Some broader harms were open questions for providers, while others may require action by school leaders, regulators, or researchers.

Toxic Content, Stereotyping, and Bias.

Educators described mixed concern about toxic or biased content, privacy, and hallucinations, often expressing confidence in classroom mediation. Concern increased when students used LLMs without educator oversight or when LLMs displaced educators’ intermediary role.

  • Toxic Content, Stereotyping, and Bias.: Educators reported mixed concern about toxic or biased content, with multiple educators having personal experience of biased outputs and stereotypes.
  • Toxic Content, Stereotyping, and Bias.: Some educators limited LLM use after biased outputs, while others addressed bias through classroom discussion and teachable moments.
  • Toxic Content, Stereotyping, and Bias.: Educators’ concerns intensified when they could not mediate student–LLM interactions and therefore could not correct harmful outputs.
  • Toxic Content, Stereotyping, and Bias.: Educators reported mixed concern about privacy, but many believed they could reduce risks by teaching students not to disclose personal data.
  • Toxic Content, Stereotyping, and Bias.: A prominent privacy concern was that students might disclose abuse or other sensitive details to chatbots instead of educators.
  • Toxic Content, Stereotyping, and Bias.: Educators had direct experience with hallucinations and concrete mitigation strategies, but were most concerned when students used LLMs without oversight.

Academic Dishonesty.

Educators identified academic dishonesty as a salient human–LLM interaction harm, but generally felt able to address it through teaching practices and relationships with students. They also raised broader concerns about learning, social development, workload, autonomy, and inequality.

  • Academic Dishonesty.: Most educators proactively raised academic dishonesty as a relevant potential harm and reported mixed levels of concern about it.
  • Academic Dishonesty.: Educators generally felt confident mitigating academic dishonesty through teaching practices, student relationships, and judgments about meaningful assignments.
  • Academic Dishonesty.: Educators often relied on intuition rather than external AI detectors because they viewed detectors as fallible and wanted flexibility in addressing suspected misuse.
  • Inhibiting Student Learning.: Educators worried that LLM reliance could inhibit critical thinking, weaken students’ voices, reduce human interaction, and erode trust between students and teachers.
  • Educator Workload and Autonomy.: Educators reported spending substantial professional and personal time vetting LLM tools and learning how to mitigate harms from tools developed without their input.
  • Educator Workload and Autonomy.: Educators said procurement and unclear guidance could reduce classroom autonomy, while unequal costs could worsen systemic inequality between districts.

6 Discussion

The discussion organizes education-specific LLM harms into technical, interaction, and broader-impact categories, then contrasts providers’ and educators’ priorities. It proposes educator-centered design, procurement, regulation, and co-design practices to address these gaps.

  • Education-specific harms: The paper identifies technical harms, interaction harms, and broader-impact harms affecting learning, social development, educator workload and autonomy, and systemic equality.Technical harms include toxic or biased content, privacy violations, and hallucinations; interaction harms include academic dishonesty.
  • Gaps in harm conceptions: Edtech providers primarily focus on technical harms, whereas educators are more concerned about broader harms arising through interactions with students, educators, and school systems.These broader harms are often not currently measurable or mitigable by providers.
  • Educator-centered development: The authors recommend making educators’ concerns more salient to providers and clarifying providers’ mitigation strategies to support trust and co-design.The intended outcome is educator-centered design and development of LLM-based edtech tools.
  • Educator mediation: Providers should build tools that facilitate educator mediation, because educators report being able to mitigate several harms through oversight of student interactions with tools.The proposed approach can increase educator autonomy while supporting mitigation of toxic or biased content, privacy violations, hallucinations, academic dishonesty, and reduced helpfulness.
  • Procurement and governance: Regulators should provide centralized, clear, independent reviews and searchable repositories of vetted LLM-based edtech to reduce educators’ workload and information gaps.The discussion points to existing organizations such as the What Works Clearinghouse as a possible foundation for this work.
  • Procurement and governance: The authors recommend educator-centered procurement, including educator input, non-penalization for opting out, and risk-benefit analysis of whether tools improve processes without marginalizing educators.Regulators and school leaders are identified as responsible for these procurement practices.

Limitations.

The study’s findings are narrowly focused on education in WEIRD, English-speaking countries and are not intended to generalize across all edtech providers. The interview design also involved nonrepresentative sampling and potential participant-identification risks.

  • Scope: The results are narrowly focused on education in WEIRD, English-speaking countries.The authors note that cultural biases, poorer performance in low-resource languages, and uneven labor and environmental costs were not surfaced by participants.
  • Sampling: The interview sample was not representative, despite including 23 educators and reaching saturation across both populations.The authors sought participants with diverse backgrounds rather than representative samples, and acknowledge that some perspectives were over-represented.
  • Sampling: The study does not attempt to generalize standard practices across the universe of edtech providers.Only six providers were interviewed, although they represented leaders whose practices may reflect emerging best practices.
  • Study design: The study’s Zoom interview method placed demands on educators’ already limited time.The researchers provided $50 compensation per educator, corresponding to an hourly rate of $50 to $100 depending on interview length.
  • Ethics: Participant anonymity created a risk that interviewees could be identified.The researchers addressed this by anonymizing quotes, limiting demographic granularity, securing data, and deleting original recordings after transcription.
Loading 2502.14592v1…