Source-linked AI summary

Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery

Qingchuan Lyu, Yingxin Li, Albert Yang

arXiv:2608.16963v1cs.LGcs.CYstat.AP

TL;DR

Learning analytics often assumes that interpretable clusters of intelligent tutoring-system behavior reveal learner types that predict learning. This paper tests that assumption with temporally separated EdNet-KT3 behavior clusters and outcomes, finding that the styles predict later engagement but not later unassisted accuracy. Thus, behavioral profiles describe study practice and engagement rather than mastery.

  • Problem

    It remains unclear whether interpretable, well-separated clusters of intelligent tutoring-system logs predict learning outcomes rather than merely describe study behavior.

  • Method

    The study derives and bootstrap-validates hierarchical study-strategy clusters, fits them on early practice, and evaluates later engagement and unassisted accuracy under a respond-count split.

  • Results

    Early behavioral clusters forecast later engagement, especially continued activity across late sessions, but not later unassisted accuracy; supervised mastery structure is nearly independent of behavior styles.

  • Takeaways & Limitations

    Behavioral clusters can support engagement-oriented analytics, but cluster membership should not be treated as a standalone proxy for mastery.

  • Takeaways & Limitations

    The findings concern active learners on one TOEIC-oriented platform and may not generalize automatically to less active users, other subjects, or other tutoring systems.

Abstract

from arXiv · show

Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning. We test that assumption on EdNet-KT3. Clustering study-strategy features (resource use, revision, video, problem practice) for 5{,}000 active learners yields a silhouette-selected parent cut ($k=5$) with 4 contrast poles (reading-focused, video-heavy, revision-heavy, and problem-first) plus a large near-mean residual ($\sim$64.9\%). Reclustering that residual adds four finer styles, giving a bootstrap-stable hierarchy of 8 named strategies. We split each learner's timeline by respond count so clusters use only the early half and outcomes only the late half. Early clusters predict later engagement (continuing to practice and finishing late sessions, especially persistence, $η^{2}\approx 0.106$; completion $η^{2}\approx 0.021$) but not later unassisted accuracy (correctness on late first-attempts without help; $p_{\mathrm{adj}}\approx 0.093$). Volume rises with some styles, yet volume-only clustering barely matches strategy labels (ARI$=0.064$). A knowledge-tracing model (SAKT) on the seven TOEIC exam sections predicts next correctness only modestly better than a baseline that knows only how hard each section usually is (AUC lift $+0.051$; CI $[+0.045,+0.058]$), and that mastery signal is nearly independent of behavior styles (ARI$=0.007$). Behavioral clustering here describes study styles and engagement, not knowledge gains.

1. Introduction

This study tests whether unsupervised study-strategy profiles from ITS logs forecast later learning or only later engagement. Using EdNet-KT3, it separates early cluster construction from late outcome assessment to prevent behavioral leakage and reports a stable hierarchy of study styles.

  • Motivation: ITS clustering is often treated as revealing learner types that predict learning and can support adaptation, despite limited stress-testing beyond interpretability and separation.ITS logs include video watching, reading explanations, answering questions, and revising questions.
  • Research question: The research question asks whether unsupervised ITS study-strategy profiles forecast later learning or only later engagement, defined as continuing to practice.The analysis uses the public EdNet-KT3 corpus.
  • Evaluation design: Early-half clustering and late-half outcome scoring prevent late behavior from contributing to both cluster features and prediction outcomes.Styles are first discovered and named on full logs, then prediction clusters are rebuilt from the early half of each learner’s question attempts.
  • Study-style hierarchy: k=5 yields four sharp contrast poles plus four softer styles within the large near-average majority.The poles are reading focused, video heavy, revision heavy, and problem attempt first; softer styles include mild in-session study before problems, video-leaning majority, read passages then attempt, and mild revisers.
  • Study-style hierarchy: Mean ARI ≈0.984 for the parent cut and ≈0.938 for the hierarchy indicate bootstrap stability.The reported hierarchy contains eight named study styles in total.

2. Literature Review · 2.1 Intelligent tutoring and adaptive instruction · 2.2 Learning styles and log-based profiling

The literature positions this study at the intersection of ITS, adaptive personalization, learning-style and strategy profiling, SRL, and the distinction between engagement and mastery. It narrows the question to whether unsupervised study-strategy clusters from interaction logs predict later mastery or mainly later engagement.

  • 2.1 Intelligent tutoring and adaptive instruction: Cognitive tutors showed that model-tracing ITS can improve classroom outcomes through immediate, concise feedback grounded in cognitive domain models.Later work extended tutor principles to help-seeking and self-explanation while identifying challenges in student modeling, metacognition, and authoring.
  • 2.1 Intelligent tutoring and adaptive instruction: Systematic reviews emphasize evaluating ITS techniques—including rule-based methods, Bayesian networks, and data mining—on learner outcomes.A 2025 review of 28 K–12 studies (N=4,597) found generally positive ITS effects, often reduced versus non-intelligent tutoring systems, making personalization and adaptivity critical.
  • 2.1 Intelligent tutoring and adaptive instruction: The study inherits ITS’s goal of actionable learner models but asks whether unsupervised study-strategy clusters forecast later mastery or mainly later engagement.This frames the paper’s contribution as a narrower test of behavioral clusters derived from interaction logs.
  • 2.2 Learning styles and log-based profiling: Adaptive-learning research commonly identifies learning styles to personalize instruction, using questionnaires, hybrid neural recognizers, and clustering of web-usage traces.The Felder–Silverman taxonomy remains influential in this literature.
  • 2.2 Learning styles and log-based profiling: This approach instead clusters simple study-strategy summaries, including video watching, reading, attempting, revising, pre-answer study, and post-error returns.Clusters are evaluated against both later engagement and later mastery rather than mapped onto a fixed questionnaire taxonomy.
  • 2.2 Learning styles and log-based profiling: Silhouette and similar scores help select the number of clusters but do not establish educational meaning.The study therefore treats cluster-count selection as insufficient evidence that labels represent meaningful learning constructs.
  • 2.2 Learning styles and log-based profiling: Labels must remain stable under resampling, and prediction claims must use held-out late outcomes.These requirements distinguish reproducible behavioral profiles from clusters that merely fit the data used to construct them.

2.3 Self-regulated learning and metacognition in ITS … 2.6 Personalization in the era of large language models

The section frames learner behavior through self-regulation and metacognition, while distinguishing engagement from achievement and testing knowledge-tracing predictions against section-difficulty baselines. It then situates ITS personalization within the opportunities and risks introduced by large language models.

  • 2.3 Self-regulated learning and metacognition in ITS: Self-regulated learning models describe tutoring behavior as recursive task definition, planning, studying tactics, adaptation, and metacognitive monitoring.The COPES model treats monitoring as a gateway to regulation.
  • 2.3 Self-regulated learning and metacognition in ITS: Prompted self-explanation acts as a lightweight metacognitive scaffold that can improve learning gains and transfer in cognitive tutors.
  • 2.3 Self-regulated learning and metacognition in ITS: AI–SRL research emphasizes adaptive personalization, prediction, profiling, tutoring systems, and assessment more than motivational constructs.A 2025 systematic mapping covered 84 AI–SRL studies, while log traces more readily capture strategy and engagement than motivation.
  • 2.4 Engagement versus achievement: Engagement indicators such as participation, persistence, and time-on-task can diverge from learning achievement.Learners may be highly active without commensurate gains, and interaction efficiency or process-quality improvements need not produce immediate learning gains.
  • 2.5 Knowledge tracing and outcome prediction: Knowledge tracing predicts next-response correctness from interaction sequences using latent knowledge updates, with BKT as a foundational baseline and SAKT as a standard neural baseline.
  • 2.5 Knowledge tracing and outcome prediction: The study uses SAKT on EdNet-KT3 to test whether behavioral structure predicts later in-app mastery against a section-difficulty prior.The prior predicts correctness from how often the cohort answers each exam section correctly, using the same guess for every learner.
  • 2.6 Personalization in the era of large language models: Large language models create personalization opportunities for dialogic tutoring and scalable feedback alongside risks involving pedagogical quality, equity, and assessment integrity.Recent adaptive and predictive tutoring systems continue combining mastery estimation with personalization.

2.7 Hierarchy, stability, and predictive validity · 2.8 Formal background · 2.9 Summary of the gap

These sections frame the study as a test of whether behavioral clusters are stable and externally valid, while separating ability-like mastery signals from behavioral study styles. The remaining gap is a unified design combining hierarchy, stability checks, temporal outcome tests, and explicit treatment of engagement–achievement divergence.

  • 2.7 Hierarchy, stability, and predictive validity: High silhouette scores or visually clean UMAPs do not establish external validity, motivating bootstrap stability checks and temporally held-out outcome tests.The methodological literature also notes that bootstrap ARI floors and feature-definition ablations are less commonly reported alongside these tests.
  • 2.7 Hierarchy, stability, and predictive validity: Hierarchical decomposition of a large residual majority is under-reported compared with flat clustering, supporting a hierarchy beyond a single chosen k.Education data-mining practice often reports a chosen k and named profiles without the full stability and ablation evidence.
  • 2.8 Formal background: Item-response theory models correctness through latent ability θ and item difficulty b, distinguishing mastery-related variation from behavioral study patterns.The Rasch form is identified as the one-parameter logistic model of this ability–difficulty relationship.
  • 2.8 Formal background: Correctness-trained models, including neural knowledge-tracing latents, are expected to recover an ability-like direction rather than a behavioral typology.Clustering such latents can separate mastery without implying behavioral study styles.
  • 2.8 Formal background: First-order sequence models represent within-session actions with transition probabilities P(a′|a), whereas flat features retain only action frequencies and time shares.Exploratory order probes therefore assess structure discarded by aggregate behavioral features.
  • 2.8 Formal background: ARI≈1 denotes reproducible structure beyond chance, and the design gates clustering on bootstrap ARI before testing outcomes with ANOVA, Holm correction, and bootstrap confidence intervals.The confidence intervals are applied to mastery gaps.
  • 2.9 Summary of the gap: The literature supports ITS learning under well-designed feedback and adaptive modeling, but engagement and achievement can diverge.Learning-style profiles are widely used for personalization, while KT and LLM tutors can improve prediction or process without validating behavior-only mastery claims.
  • 2.9 Summary of the gap: The underexplored gap is a unified empirical treatment linking interpretable behavioral profiles with stability, temporal validation, and mastery-related analyses.This gap is positioned against existing work on ITS, learning styles, self-regulated learning, engagement, achievement, knowledge tracing, and LLM tutors.

3. Methods

The study defines a locked EdNet-KT3 cohort and preprocessing pipeline, discovers a hierarchical eight-pattern typology, and evaluates it with stability checks and leakage-controlled temporal splits. Additional analyses test whether styles reduce to practice volume and compare behavioral clustering with SAKT predictions across seven TOEIC sections.

  • Cohort and preprocessing: 45,621 users met the active-learner gate of at least 50 question respond events; profile discovery and audits used 5,000 seeded learners, with 4,986 retained after feature filtering.Active status was based on respond events rather than total actions or calendar days.
  • Cohort and preprocessing: 30-minute inactivity gaps segmented streams, sessions with fewer than five actions were dropped, and each log was capped at the earliest 5,000 actions.Actions were mapped to video watching, reading, problem attempts, revisions, idle returns, and related study events.
  • Hierarchical profile discovery: k=5 parent clusters were selected using silhouette and Davies–Bouldin ranks, while reclustering the ∼64.9% near-average residual at k=4 produced 8 named learning patterns.The typology retained four contrast poles as parent labels and added four majority styles.
  • Robustness checks: 80% subsample reclustering assessed label agreement with adjusted Rand index, while HDBSCAN was excluded after labeling everyone as noise at the parent stage.Feature-definition ablations also examined alternative reading categories, session lengths, feature subsets, and log-counts.
  • Temporal evaluation design: The leakage-controlled evaluation split question attempts by count, rebuilt features from the early half, refit clustering at locked k=5, and scored outcomes on the late half.Equalizing experience share accommodates different start times and practice volumes, although late accuracy estimates can be noisier for learners with few late items.
  • Comparative analyses: Volume-only clustering used average responds, actions, and sessions at k=5, while SAKT predicted next-response correctness on first-attempt sequences labeled by 7 TOEIC sections.The SAKT configuration used d=64, 4 heads, maximum length 100, dropout 0.2, Adam 10^-3, and validation-AUC early stopping.

4. Results

The results establish a bootstrap-stable eight-style hierarchy whose early labels predict later engagement but not later unassisted accuracy. Volume and correctness-supervised knowledge signals remain largely distinct from the behavioral typology.

  • Behavioral typology: Parent k=5 identifies four contrast styles plus a large near-average majority, which reclustering divides into four softer styles.The parent analysis covers 4,986 learners; the typical majority contains 3,236 users.
  • Behavioral typology: Bootstrap-stable hierarchical encodings support reporting the eight-style typology despite modest within-majority silhouette.Within-majority silhouette is approximately 0.199, but stability—not density—is the reporting criterion.
  • Early/late outcomes: η2≈0.106 for late-session persistence and η2≈0.021 for late completion, whereas terminal unassisted accuracy shows η2≈0.003 and p_adj≈0.093.Clusters use the early timeline and outcomes use the late timeline, limiting leakage between predictors and outcomes.
  • Knowledge probe: AUC lift +0.051 over the part-difficulty prior, with CI [+0.045,+0.058], is modest, and the correctness-supervised signal barely agrees with behavior styles.SAKT uses seven TOEIC section-level knowledge components; its embeddings largely re-encode ability rather than a dense behavioral mastery typology.

5. Discussion

The stable, interpretable hierarchy of behavior-based study styles predicts later engagement rather than later unassisted accuracy. These findings caution against treating behavioral profiles as mastery measures and support pairing style dashboards with independent correctness or assessment evidence.

  • Core interpretation: Stable study-style clusters forecast later engagement, not later unassisted accuracy, cautioning against inferring learning from interpretable behavioral typologies.Feature ablations indicate the hierarchy is not merely an artifact of maximizing silhouette.
  • Core interpretation: −12.7 pp vs. majority; bootstrap CI [−20.9,−4.2] describes the small video-heavy pole, while difficulty-adjusted OLS does not overturn the overall mastery null.Parent styles barely differ in mean empirical item difficulty, and external learning gains require independent outcome labels.
  • Implications for analytics practice: Policies keyed only to labels such as “revision-heavy” or “video-heavy” partly track practice volume and later activity, not who knows more.Practitioners should pair style dashboards with independent mastery evidence, including item correctness or external assessments.
  • Theoretical reading: Correctness-trained models can recover ability-like structure, but correctness-based mastery and behavior-only profiles need not align.The SAKT probe matches the theoretical expectation that correctness models recover ability, while study-strategy features omit correctness.

6. Limitations

The study’s conclusions are limited by in-app outcome proxies, a narrow active-learner sample, and design choices affecting reliability and label interpretation. The descriptive typology does not by itself establish actionable interventions or generalize beyond this platform and cohort.

  • “Mastery” is later in-app terminal unassisted first-attempt accuracy, not a post-test, grade, certification, or retention measure, and remains confounded by guessing and item reach.
  • Results describe active learners with ≥50 responds on one TOEIC-oriented platform and may not generalize to less active users, K–12, other subjects, or other ITS products.
  • Absolute SAKT AUC (0.605) is reported for a ∼5k active cohort, so the study emphasizes lift over the part-difficulty prior (+0.051; CI [+0.045,+0.058]) rather than state-of-the-art comparison.The KT embeddings use an early-stopped validation checkpoint and serve as a descriptive supervised probe, not a pure hold-out embedding study.
  • The respond-rank 50/50 split leaves unequal late volume, ranging from about 10 to 1,000 responds, making terminal-accuracy reliability heteroskedastic and estimates noisier for sparse late or unassisted items.
  • Full-timeline profile names and stability ARIs should not be transferred to strict early labels, which use an independent fixed k=5 fit rather than fresh model selection.
  • The reviser versus straight-through-answerer contrast is descriptive only because removing idle-return events eliminates bootstrap stability, so it is excluded from the reported 8-style typology.
  • A stable engagement typology does not prescribe interventions; educator-facing supports require practitioner co-design, institutionally aligned outcomes, multi-platform replication, and external criteria.

7. Conclusion

The study finds that behavioral strategy styles describe how learners practice and forecast later engagement, but not later unassisted accuracy or mastery. Accordingly, style labels may support engagement interventions but should not serve as high-stakes knowledge judgments.

  • Conclusion: 8 stable study-strategy styles capture contrasting poles and finer variations within the typical majority.The hierarchy combines sharp contrasts at the first clustering cut with softer styles inside the large near-mean group.
  • Conclusion: Early behavior-only profiles forecast later engagement, especially whether learners remain active across late sessions, rather than later unassisted accuracy.Practice volume accompanies some parent styles but does not define them.
  • Toward helping learners thrive: Behavior-only profiles are actionable for engagement support but insufficient alone as mastery signals.Future work should test nudges for learners showing early signs of dropping practice, rather than targeting a named style such as revision-heavy.
  • Ethics and stakeholder use: In live ITS use, style labels should describe practice behavior and cue engagement support, not measure knowledge or inform high-stakes decisions.The analyses used aggregated behavioral features without attempting to re-identify learners in the public EdNet research corpus.
  • Reproducibility: Configuration files store paths, thresholds, and seeds, while each pipeline stage records run metadata.Analysis code and notebooks for the locked results are available from the author upon reasonable request.

Funding

The research received no specific grant funding from public, commercial, or not-for-profit sectors.

  • Funding: No specific grant supported the research from public, commercial, or not-for-profit funding agencies.

Appendix

Appendix exemplars show that parent study styles differ sharply in temporal density and action mix. Typical-majority and revision-heavy activity arrives in short bursts, while other styles span gaps, repeated returns, or longer calendar periods.

  • Exemplar parent-style timelines: Parent-style exemplar timelines reveal sharply different temporal density and action mix across learners.The exemplars show the first 200 actions for each parent style.
  • Exemplar parent-style timelines: Typical-majority and revision-heavy exemplars concentrate activity in a short early burst.
  • Exemplar parent-style timelines: Reading-focused activity includes a second session after a long idle gap, whereas video-heavy activity mixes video with repeated returns.
  • Exemplar parent-style timelines: Problem-first activity stretches sparse problem-solving across a long calendar span.
Loading 2608.16963v1…