Source-linked AI summary

When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk

Liliana Santos-Deonizio, James Malamut, Ramón Martínez, Dorottya Demszky

arXiv:2608.23780v1cs.CLcs.AIcs.HC

TL;DR

Text-only validation of LLM-based student-talk measures can miss classroom context and classifications that students contest. This study uses ethnographically oriented methods to re-contextualize interaction, finding that these methods recover dimensions unavailable in transcripts and that students challenge model classifications and coding.

  • Problem

    Text-only validation approaches are insufficient because transcripts miss classroom context and students may contest the resulting classifications.

  • Method

    The study re-contextualized classroom interaction using ethnographically oriented methods alongside text-based measures of student talk.

  • Results

    Ethnographically oriented methods recovered interactional dimensions that text-only transcripts cannot represent, while students contested model classifications and the coding scheme.

  • Takeaways & Limitations

    Interpreting student participation requires attention to dimensions beyond text-only transcripts and to students’ challenges to classifications and coding.

  • Takeaways & Limitations

    The claims are bounded by the study’s intentionally broader framing of student learning and the small number of conversations it could re-contextualize.

Abstract

from arXiv · show

LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal contributions, which de-contextualize student language. Common practices for validating these measures include comparing outputs against expert annotations by adults, using held out evaluation sets and F1 scores. We argue that these approaches are insufficient to ensure that such measures are meaningful and equitable for teaching and learning, particularly for racially and linguistically marginalized youth. In order to center the youth whose talk is being analyzed, re-contextualizing these classroom conversations and engaging youth in the research process is necessary. Sharing epistemic authority with youth, ultimately, centers their point of view and adds crucial nuance to the analysis of their talk that adult experts, researchers, and LLMs cannot provide. In a case study of multilingual youth in one 8th-grade math classroom, we address the epistemic exclusion of youth by employing multiple ethnographically-oriented methods to re-contextualize student conversations and center youth as epistemic authorities in conversation with researchers and LLMs. We conducted participant observations, interviews, focus groups, and member checks with four focal students. Findings reveal that there were misalignments between students' interpretations of their own math talk experiences and the LLM-based measures of their talk. Students contested both the LLM classifications and the coding scheme used to measure their talk, highlighting the need for youth to be involved in the epistemic process of producing knowledge about their experiences.

1 Introduction

The paper argues that validating LLM-based measures of student talk requires re-contextualizing transcripts and involving youth as epistemic authorities on their own talk. A case study shows that youth interpretations can reveal limitations in both LLM classifications and the coding schemes used to define math talk.

  • Motivating example: Ximena contested an LLM’s “Off-task” classification, explaining that she was discussing math even though negatively.Her response illustrates how students may interpret their talk differently from text-only classifications.
  • Epistemic authority: Validation is framed as knowledge production that raises questions about who holds epistemic authority over interpretations of student language.The paper positions students as experts on their own experiences, whose involvement can deepen analysis beyond what adults and LLMs achieve alone.
  • Epistemic shift: The proposed epistemic shift re-contextualizes transcripts as situated, social, cultural, and embodied experiences and includes youth as authorities rather than merely annotation subjects.This approach treats language as more than a text-based representation and centers youth participation in validation.
  • Case study: In a case study of multilingual youth in one 8th-grade math classroom, participant observations, interviews, focus groups, and member checks re-contextualized four students’ talk.These methods surfaced relationships, routines, physical space, gesture, prosody, and language ideologies obscured by text-only representations.
  • Findings: Students’ perspectives revealed misalignments with LLM classifications that required modifying the coding scheme for math talk, not merely prompt-engineering.The resulting validation model makes youth participants in deciding what the data means, rather than only treating them as data sources.

2 Related Work

Prior work uses computational methods to analyze classroom discourse at scale, typically validating models against expert annotations and standardized performance metrics. Related scholarship shows that transcript-based measures can miss contextualized, multilingual participation, motivating participatory validation that treats youth as epistemic authorities in the research process.

  • Computational discourse analysis: Computational methods analyze student and teacher discourse, including talk moves, collaboration, and classroom participation, using LIWC, BERT, and other classifiers.These studies relate discourse features to learning outcomes or collaborative behaviors.
  • Computational discourse analysis: Models are typically validated with expert-annotated transcripts, held-out test sets, F1 scores, AUROC, or agreement with human experts.These practices assess how well models reproduce expert labels.
  • Limits of existing validation: Transcript-based validation treats transcripts as sufficient representations of classroom interaction and rarely incorporates students’ perspectives on their measured discourse.Thus, it provides limited evidence about whether outputs align with students’ understandings or broader contexts.
  • Contextualized multilingual participation: Ethnographic and multilingual mathematics education research shows that discourse meaning and participation depend on social relationships, identity, language ideologies, race, and embodied context.Multilingual students’ broad linguistic repertoires may not be fully legible through transcript-based measures alone.
  • Participatory and epistemic approaches: Participatory AI and HCI evaluation surfaces harms, biases, and mismatches overlooked by standardized metrics, while this work applies epistemic exclusion frameworks to measuring and representing youth learning.The study engages youth as experts of their experiences who interrogate LLM classifications of their talk.

3 LLM-Based Measures of Student Talk

The study developed preliminary GPT-5.1 measures of multilingual students’ math language practices using expert-annotated classroom transcripts and seven frequent talk codes. Although revisions improved F1 scores, validation exposed context-dependent ambiguities and persistent weaknesses in conceptual, infrequent features.

  • Data and codebook: The project collected multilingual middle-school math classroom transcripts, transcribed by humans or automated speech recognition with human cleanup before annotation.Students worked on math tasks with peers, and at least one student per group reported Spanish as a home language.
  • Data and codebook: LLMs coded collaborative and academic talk moves using a multi-theoretic scheme; analysis focused on Off-task, Understanding, Recording, Question, Claim, Next-step, and Disagree.The codebook combined collaborative problem solving and mathematics learning frameworks, with definitions developed by human experts.
  • Validation and prediction: Five transcripts totaling 41,196 utterances formed the human-expert validation set, with sentence-level utterances tagged as English, Spanish, or both.Claude 3.7, GPT 5.0, and GPT 5.1 were evaluated, and GPT-5.1 performed best based on F1 scores.
  • Validation and prediction: After revising prompts and resolving human-annotation inconsistencies, F1 scores increased for every feature.Claim increased by +0.104, Question by +0.23, and the remaining analytic-sample codes by +0.021 - +0.061.
  • Limitations and motivation: Validation required context beyond text, including intonation, intent, and references, and GPT-5.1 initially struggled to reach a 0.7 F1 score on most features.Scores improved after adjusting definitions but remained low for some features, notably Next-Step and Disagree, motivating re-contextualization and additional validation approaches.

4 Ethnographically-Oriented Methods

The study re-contextualized student talk and shared validation with youth through ethnographically oriented methods. In one 8th-grade math classroom, four focal students reviewed their conversations and LLM annotations, enabling analysis of contextual factors and misalignments in classifications.

  • Methods: The validation approach combined re-contextualizing student talk with bringing youth into the validation process through participant observations, interviews, focus groups, and member checks.These methods were used as part of the study’s epistemic shift in validating LLM-based measures of student talk.
  • Focal classroom and participants: The case study examined one 8th-grade math classroom serving multilingual youth from African-American and Latinx families, focusing on four students who frequently worked in the same pairs.The focal students were Ximena, Destiny, Diego, and Drake, whose collaborative dynamics could therefore be discussed.
  • Data collection: Across the school year, the first author conducted seven participant observations, four interviews, two focus groups, and four member checks with the focal students.Two recent classroom transcripts supplied the excerpts used across interviews, focus groups, and member checks.
  • Student interviews: Interviews used language mapping to document students’ linguistic repertoires and participant retrospection to elicit what students were doing, thinking, and feeling during math conversations.Language mapping addressed students’ language use across daily spaces, audiences, and interlocutors, within and beyond math class.
  • Participant retrospection and member checks: During member checks, students reviewed LLM annotations of their own transcript excerpts, indicated agreement or disagreement, and explained what the model had missed.Focus groups revisited the same excerpts to discuss group dynamics, collaboration, and math-class talk with peers.
  • Analysis: Analysis grouped observation notes into environmental, social, multimodal, and relational factors absent from transcripts, then categorized student disagreements as challenges to LLM classifications or the coding scheme.These groupings surfaced misalignments between students’ views of their talk and the LLMs’ and coding scheme’s representations.

5 Results

The results show that transcript-only LLM measures omit contextual, multimodal, relational, and intentional dimensions of student participation. Ethnographically oriented methods and student member checks revealed misalignments between coding schemes, LLM classifications, and students’ interpretations of their talk.

  • Contextualizing Student Talk: Transcript-only LLM measures operate on an incomplete representation of classroom activity because they omit physical, sociocultural, multimodal, relational, curricular, and language-ecology dimensions.Ethnographically oriented methods recover much of this context and support more holistic representations of student math talk and participation.
  • Contextualizing Student Talk: Students’ physical environment and independent work patterns demonstrated that participation was not reducible to verbal contributions during partner talk or whole-class sharing.A majority of youth kept their heads down and continued working independently during interaction opportunities.
  • Contextualizing Student Talk: Nonverbal actions, tone, prosody, silence, gesture, and posture conveyed mathematical thinking and meaning that transcribed speech and text-based LLM measures did not capture.Coding schemes centered on spoken academic talk therefore risk reflecting differences in students’ preferred modes of engagement as deficits.
  • Contextualizing Student Talk: Interpersonal relationships with peers and teachers shaped students’ comfort, willingness to talk, and participation in ways absent from recordings and LLM coding schemes.These relationships helped explain participation variation that transcript-based measures would otherwise obscure.
  • Student Member Checks: Student member checks exposed deeper misalignments between LLM classifications and students’ intended functions, including frustration, counting, and guessing being coded as off-task or Claim.Students’ interpretations corrected coding assumptions and surfaced missing categories such as brainstorming or guessing.

6 Discussion

The discussion identifies two epistemic shifts: LLM-based measures structurally omit contextual dimensions of classroom interaction, and validation must share interpretive authority with youth. Participatory, multi-method validation can make discourse measures more meaningful, equitable, and representative for marginalized students.

  • Epistemic shifts in data: LLM-based measures reduce embodied classroom interactions to text and discrete labels, structurally omitting relational, spatial, sociocultural, and multimodal dimensions.These omissions result from applying text-based classification to complex social interaction and cannot be corrected solely through improved prompting or larger datasets.
  • Epistemic shifts in data: Taking conversations out of context and excluding student input de-centers youth whose talk is being measured.The discussion frames this as a design-level consequence of how LLM-based systems represent student discourse.
  • Epistemic shifts in validation: Validation determines whose interpretation of classroom data counts, so involving youth shifts epistemic authority from researchers and models to students.Student perspectives help assess whether talk measures are meaningful and accurate to their original intentions while revealing what youth see as part of math learning.
  • Findings and implications: Students’ disagreements with LLM classifications and coding schemes reflected alternative interpretations of their talk, not simply classification errors.Ximena’s challenge to an off-task label illustrated that math conversations can include social and affective dimensions, while students’ added nuance exposed coding misalignments.
  • Recommendations: Participatory validation should be central to responsible AI in education, using multiple methods and community input to create more equitable, representative measures.The discussion recommends grounding measures in youth’s lived experiences to reduce misattribution, missed contributions, and harms disproportionately affecting marginalized youth.

7 Limitations and Ethical Considerations

The paper’s claims are limited by its small, single-classroom sample and validation-only youth participation, while LLM-based measures raise concerns about bias, harm, environmental impacts, and broader community effects.

  • Sample and methodological scope: Methodological triangulation re-contextualized a small number of conversations but constrained how many conversations could be analyzed.The approach treats student learning as more than spoken contributions, creating a tradeoff between contextual depth and analytic scale.
  • Sample and methodological scope: Building relationships over months limited the study to one classroom and four students, preventing generalization from the sample.A larger or more varied sample could reveal additional participation patterns and disagreements.
  • Participatory engagement: Youth participated in validating LLM outputs and the coding scheme but did not help design the coding scheme or prompts.A fuller co-design process could define math talk before classification and surface mismatches missed by member checks.
  • Ethical considerations: LLM-based measures raise concerns about bias and harms to users, environmental impacts, and effects on communities near data centers.These concerns are especially salient for larger models that cannot be run locally.
  • Ethical considerations: Future research must account for its effects beyond the focal classroom, including on communities most affected by LLM-dependent infrastructure.The authors emphasize that research impact matters even though the validation set is small.

8 Future work

Future work should broaden the sample, deepen youth co-design, and develop smaller locally runnable models. It should also examine how discourse measures operate across settings and downstream uses, and whether participatory validation applies to other educational AI tools.

  • Future studies should use broader samples, fuller co-design with youth, and smaller, locally runnable models.
  • Research should extend beyond mathematics classrooms to after-school, extended-day, and other subject-area settings.This would clarify how youth language practices and descriptive categories shift across contexts.
  • Future work should examine how LLM-based discourse outputs shape participation, instruction, and student-teacher relationships.The relevant downstream audiences include teachers, students, and administrators.
  • Participatory validation should be applied to other educational AI tools, including automated feedback, assessment, and behavior-detection systems.This can test whether observed misalignments are specific to discourse measurement or reflect broader gaps between AI-generated representations and students’ experiences.

9 Conclusion

Standard validation practices do not fully capture the meaning of student participation or ensure that LLM-based discourse measures are meaningful and equitable. Ethnographic re-contextualization and shared epistemic authority with youth make validation a participatory knowledge-production practice.

  • 9 Conclusion: Standard validation practices, including expert annotation, F1 scores, and held-out test sets, are insufficient to capture student participation fully or ensure meaningful, equitable measures.These practices are identified as insufficient for validating LLM-based measures of classroom discourse.
  • 9 Conclusion: Ethnographically oriented methods recovered dimensions of classroom interaction that text-only transcripts could not represent.The finding comes from a case study of multilingual youth in one 8th-grade math classroom.
  • 9 Conclusion: Students contested both LLM classifications and the coding schemes used to produce them.This contestation demonstrates that student interpretations can diverge from the measures applied to their talk.
  • 9 Conclusion: Treating youth as epistemic authorities reframes validation as knowledge production in which measured participants help determine what measurements mean.Students are described as experts regarding their own experiences.
  • 9 Conclusion: Participatory and community-engaged approaches that center students’ voices are essential for making educational LLM tools accurate and equitable.Epistemic authority should be shared with students when evaluating and designing tools to interpret their talk.

A Research Methods · A.1 Student Interviews

This section includes an overview table for the participant retrospection and member check protocol, alongside manuscript-submission placeholder text. No further methodological details are provided in the supplied passages.

  • A.1 Student Interviews: The section includes a placeholder stating “Manuscript submitted to ACM.”This placeholder appears in two supplied paragraphs.
  • A.1 Student Interviews: Table 6 is titled “Overview of Participant Retrospection and Member Check Protocol.”The supplied passage provides the table title but not its contents.
Loading 2608.23780v1…