Source-linked AI summary
Student Success Prediction in MOOCs
Josh Gardner, Christopher Brooks
TL;DR
Predictive modeling research in MOOCs has expanded without a comprehensive critical synthesis of its features, methods, outcomes, and theoretical grounding. This paper surveys 87 studies across those dimensions and finds methodological gaps involving evaluation, experimental populations, and realism of prediction settings. It recommends stronger evaluation, broader populations, and realistic contexts to support accurate, actionable, and theory-building models.
Problem
MOOC predictive-modeling research lacks comprehensive synthesis and methodological evaluation despite its relevance to student success prediction and intervention.
Method
The paper surveys and categorizes MOOC predictive-modeling studies by features, outcomes, theoretical models, data, feature engineering, algorithms, evaluation, and prediction architecture.
Results
Clickstream features were the strongest predictors of completion, while adding them improved a linguistic-only model by about 10%.
Takeaways & Limitations
Future MOOC research should pursue robust evaluation, broader experimental populations, and realistic contexts for accurate, actionable, and theory-building models.
Takeaways & Limitations
Many surveyed experiments use post hoc data or same-course evaluation unavailable to real-time models in active courses.
Abstract
from arXiv · showhide
Predictive models of student success in Massive Open Online Courses (MOOCs) are a critical component of effective content personalization and adaptive interventions. In this article we review the state of the art in predictive models of student success in MOOCs and present a categorization of MOOC research according to the predictors (features), prediction (outcomes), and underlying theoretical model. We critically survey work across each category, providing data on the raw data source, feature engineering, statistical model, evaluation method, prediction architecture, and other aspects of these experiments. Such a review is particularly useful given the rapid expansion of predictive modeling research in MOOCs since the emergence of major MOOC platforms in 2012. This survey reveals several key methodological gaps, which include extensive filtering of experimental subpopulations, ineffective student model evaluation, and the use of experimental data which would be unavailable for real-world student success prediction and intervention, which is the ultimate goal of such models. Finally, we highlight opportunities for future research, which include temporal modeling, research bridging predictive and explanatory student models, work which contributes to learning theory, and evaluating long-term learner success in MOOCs.
1 Introduction
MOOC predictive-modeling research expanded rapidly after 2012, but lacked synthesis of its findings and methodology. This review surveys the field to identify consensus, gaps, and directions for models that can support learners in active courses.
- The review surveys n = 87 studies to identify consensus, research gaps, unanswered questions, and needed action in MOOC dropout research.
- It evaluates not only prior findings but also feature extraction, modeling techniques, and methodology across a multidisciplinary research community.
- The review targets predictive models that can support real-time, personalized interventions for learners in active courses.
- The study organizes prior work around the task, available data, predictive-model construction and evaluation, and a matrix of n = 87 studies.
- MOOCs differ from traditional educational settings through massive, open, online, low- or no-stakes, asynchronous, and heterogeneous participation conditions.
- Low completion rates among intended completers make effective prediction and support an important concern for MOOC research.
2 Student Success Prediction in MOOCs
The review defines student success broadly because MOOCs support diverse learner goals and lack consensus on a single success metric. It therefore surveys predictive models covering completion, engagement, learning, and future achievement.
- The review frames student success prediction as a task requiring clear definition before comparing prior work and drawing conclusions.
- The surveyed outcomes include completion, certification, overall course grades, and exam grades.
- Student success includes metrics for course completion, engagement, learning, or future achievement related to a MOOC’s content or goals.
- Using several metrics captures diverse learner goals, reflects measurement disagreement, and permits tests of model robustness across outcomes.
- The literature review catalogs common success metrics and their frequencies, while noting alternative long-term measures as a separate line of work.
2.2 Why Model Student Success in MOOCs?
MOOC success models are intended to support personalized interventions, adaptive learning experiences, and understanding of learner behavior. The review distinguishes accuracy, actionability, and theory-building while noting limited real-time use.
- Predictive models are developed to provide personalized support, adapt content and pathways, and improve understanding of learner behavior and outcomes.
- Targeted interventions matter because MOOC populations are large while instructional support resources are scarce.
- Dropout or stopout prediction is the most common success metric, alongside final exam grade, course grade, and pass/fail outcomes.
- Accurate predictions are necessary for interventions, but most prior architectures are difficult to implement in actively running courses.
- The review presents predictive models along three dimensions—accuracy, actionability, and theory-building—with no strict trade-off among them.
- Very little prior research has used student-success predictions for true real-time intervention, adaptive content, or learner pathways.
2.3 Data for Student Success Prediction in MOOCs
MOOC platforms provide rich, high-granularity data from clickstreams, forums, assignments, course materials, and demographics. These sources support varied feature engineering but also create substantial processing, coverage, and bias challenges.
- The review catalogs raw data sources, formats, schemas, behaviors, and metrics used across predictive-modeling studies.
- MOOC platforms provide rich, high-granularity data at a scale generally unavailable in traditional educational contexts.
- Clickstream logs record server interactions in JSON, but their raw form requires labor-intensive feature extraction before modeling.
- A single course can generate tens of gigabytes of clickstream data, making parsing and aggregation computationally expensive.
- Forum data supports engagement, natural-language measures of mastery or affect, and social-network features, and is second only to clickstreams in surveyed use.
- Assignment data is less common because few registrants complete assignments and assignment practices vary substantially across courses.
- Demographic data is often collected through optional surveys subject to response bias, limiting its availability and use in predictive studies.
2.4 Relation to Other MOOC Research
MOOC predictive-modeling research sits within a broader literature addressing participation, interventions, learner discourse, demographics, and academic honesty. Within this landscape, predictive modeling supports either understanding factors associated with outcomes or informing learner-support systems.
- MOOC research spans learner discourse, completion interventions, demographics, participation, course activity, plagiarism, and academic honesty.
- Predictive modeling is used either to understand factors associated with dropout or to support systems intended to improve student experiences or outcomes.
3 Predictive Models of Student Success in MOOCs: A Feature, Outcome, and Model-Based Taxonomy
The review surveys predictive models of student success in MOOCs using a structured approach to organize prior work. It begins by introducing the review methodology and relevant categorizations.
- The review organizes prior predictive-modeling research on MOOC student success through an overview of its methodology and relevant categorizations.
3.1 Categorization Scheme
The review classifies MOOC predictive-modeling studies by input features, prediction outcomes, and theoretical models, while emphasizing feature extraction and documenting unevenly studied feature–outcome combinations. It also cautions that differences in populations, evaluation methods, and outcomes limit direct performance comparisons.
- Categorization Scheme: Studies are grouped by input features, prediction outcomes, and theoretical models, with overlapping categories discussed when applicable.
- Categorization Scheme: The categorization is presented as novel for MOOC predictive-modeling research and exposes research gaps through observed feature–outcome pairings.
- Categorization Scheme: Only two surveyed works used performance-based features to predict course completion, indicating a specific underexplored feature–outcome pairing.
- Categorization Scheme: Feature extraction is treated as a critical dimension because prior work identifies preprocessing and variable construction as especially important to predictive performance.
- Categorization Scheme: Activity-based feature sets dominate the survey, reflecting abundant granular platform data and the prevalence of dropout or persistence outcomes.
- Categorization Scheme: Comparing predictive performance across studies is unreliable because experimental populations, evaluation methods, metrics, and outcomes vary substantially.
3.2 Survey Methodology and Criteria for Inclusion
The survey uses broad inclusion criteria to unify predictive-modeling research across disciplines and MOOC-like settings. It includes relevant exploratory or descriptive studies, while acknowledging departures from the core definition and interpretive trade-offs between fit and interpretability.
- Survey Methodology and Criteria for Inclusion: The review includes peer-reviewed studies applying predictive modeling to broadly defined student success in MOOCs or sufficiently similar contexts.
- Survey Methodology and Criteria for Inclusion: Studies were generally included when they provided enough methodological detail about data, feature engineering, models, and experimental results.
- Survey Methodology and Criteria for Inclusion: The survey searched prominent peer-reviewed venues in learning analytics and educational data mining, among related fields.
- Survey Methodology and Criteria for Inclusion: Some included studies evaluated for-credit courses or unclear contexts, typically because of novelty, importance, completeness, or a preference for inclusion.
- Survey Methodology and Criteria for Inclusion: The review includes some non-predictive studies because predictive and data-understanding models both use statistical relationships between attributes and outcomes.
- Survey Methodology and Criteria for Inclusion: Black-box models may fit data better but are harder to interpret, whereas interpretable data models may sacrifice fit quality.
3.3 Activity-Based Models
Activity-based models dominate MOOC student-success prediction because platforms provide abundant granular behavioral data, especially for engagement outcomes. The surveyed work emphasizes temporal behavior, higher-order feature engineering, and deployment-aware evaluation, while highlighting data and methodological constraints.
- Activity-Based Models: Activity-based features and outcomes are overwhelmingly the most common surveyed approach because MOOC platforms provide abundant, granular behavioral data.Clickstream data offer interaction-level records unavailable for most other feature categories.
- Models Utilizing Early Course Activity: Early engagement predicts persistence, with first-week quiz or peer-assessment completion significantly associated with continued course participation.The finding remains significant when controlling for other behaviors, indicating that learners reveal useful behavioral information early.
- Temporal and Sequential Activity Models: Students who never check course progress show weekly dropout rates of 20-40%, whereas those checking at least four times remain below 5%.The progress-page feature serves as an observable proxy for learners’ intention to complete.
- Temporal and Sequential Activity Models: Temporal models can outperform SVM, IOHMM variants, and logistic regression, but LSTM results require replication across larger course samples.Comparisons are difficult when studies use different populations, outcome definitions, or implementations.
- Higher-Order Activity-Based Features: Feature engineering, rather than predictive algorithms, is identified as the primary driver of improvements in MOOC predictive modeling.The review recommends higher-order and distinctive features that capture information relevant to student success.
- Novel Feature Extraction and Prediction Architectures: Retrospective a posteriori models often provide optimistic estimates and may lose performance when transferred to new course offerings.The review calls for evaluation on successive offerings when models are intended for real-time use.
3.4 Discussion Forum and Text-Based Models
Discussion forum and text-based models use linguistic, social, and behavioral information from learner-generated text to predict dropout, completion, and performance. Their usefulness varies with feature choice and learner activity, while sparse participation limits generalizability.
- Forum-Based Predictive Work: Forum models combine cohort, post, and social-network features to predict dropout among forum participants.In one study, first-week membership, longer-than-average posts, and higher authority were associated with lower dropout probability.
- Text Features: Pre-course unigram features improved dropout prediction over demographics alone for students intending to complete the course, whereas LIWC features were not significant.The result concerns a selected subgroup defined by intended completion.
- Performance Prediction: Discourse features explained 5% of final-grade variance overall but 23% among the most active students.A model combining discourse and participant features explained 93% of variance, suggesting discourse usefulness was concentrated among highly active learners.
- Feature Comparisons: Clickstream activity features were stronger completion predictors than NLP features, while adding clickstream data improved linguistic-only performance by about 10%.The comparison used students who both posted in a forum and completed an assignment.
- Sentiment Analysis: Sentiment measures showed associations with assignment performance and attrition, including topic-dependent sentiment–attrition relationships.Assignment-related sentiment had a modest negative correlation with performance, while forum sentiment increased modestly over the course.
- Limitations: Forum and text data are often sparse because optional activities reach only a fraction of learners, sometimes as little as 5%.This limits the population to which text-based predictive findings can directly generalize.
3.5 Social Models
Social models represent relationships among learners and connect network structure with MOOC performance and dropout. Findings support social-network features as useful predictors, but measurement and generalization remain difficult.
- Social Networks: Students who earned certificates or distinctions were more likely to interact through social networks than non-completers.The finding came from two programming MOOCs offered in English and Spanish.
- Social Centrality: Discourse features explained about 10% of performance variance through predicted social centrality, compared with 92% for discourse-plus-participant features.The comparison indicates that participant information substantially expanded explanatory coverage.
- Subcommunities: Membership in some probabilistic subcommunities significantly predicted dropout, with two to four significant subcommunities across three MOOCs.The models considered up to 20 subcommunities per course.
- Interaction Types: Student–student, student–teacher, and student–resource interactions were significantly related to performance in online courses but not in courses with an in-person component.This evidence comes from non-MOOC online courses and is presented as relevant comparative evidence.
- Limitations: Further MOOC research is needed because social networks are difficult to measure in existing data and external network data are rarely used.The limitation is especially pronounced in relatively small, single-course samples.
3.6 Cognitive Models
Cognitive models infer learner states, attitudes, emotions, or higher-order thinking from questionnaires, forum behavior, activity traces, and sensors. The surveyed work finds predictive signals but relies heavily on novel or self-reported data collection.
- Cognitive Data: Cognitive data remain relatively uncommon in MOOCs because collecting them is harder than collecting activity or forum data.Studies have used biometric tracking and contemporaneous questionnaires to access cognitive states.
- Higher-Order Thinking: Active and constructive forum behaviors were associated with learning outcomes in hand-coded analyses of student discourse.The study used a cognitive-science learning-activity classification scheme to evaluate multiple outcomes.
- Motivation and Engagement: Forum-derived cognitive engagement and human-coded motivation features significantly predicted dropout.These models inferred motivation and confusion from learner-generated forum content.
- Activity-Based Inference: An information processing index inferred cognitive states from expert-defined interaction sequences and manually assigned action-group weights.The authors caution that manually defined features can introduce experimenter bias.
- Emotions: Anxiety, confusion, frustration, and hope were each significantly correlated with dropout, although another questionnaire study found no emotion-proportion differences between completers and non-completers.Sentiment analysis also suggests that emotional information in learner text can support dropout prediction.
- Novel Data Collection: Mobile-phone heart-rate tracking has been used as a proof of concept for inferring mind wandering, interest, and confusion.The work points toward multimodal learning analytics and sensor-based student models.
- Future Directions: Future cognitive-model research should move beyond questionnaires and self-reports as the sole source of learner data.Increasingly affordable sensors and ubiquitous mobile devices may support broader data collection.
3.7 Learning-Based Models
Learning-based models use performance, learning behavior, and learning theory to predict achievement and related outcomes. Results highlight the importance of prior knowledge and interactive resources, while showing that learning and dropout may diverge.
- Scope of Research: Learning-based features and outcomes have been used surprisingly little in predictive MOOC research despite the central role of learning.Existing work often adapts methods such as Item Response Theory and Bayesian Knowledge Tracing.
- Learning Prediction: Personalized linear regression outperformed KT-IDEM for predicting homework scores across two MOOCs.The comparison targeted quiz and homework grade prediction.
- Assessment Measures: The Cloze Test was positively associated with exam performance and overall course grade but not with online open-book quizzes or projects.The contrast was attributed to differences in student control and time constraints across task settings.
- Prior Knowledge: Prior knowledge and problem-solving abilities accounted for 83% of performance variance in a discrete optimization MOOC.The predictors were measured using two performance tasks in a course with relatively high prior-knowledge requirements.
- Interactive Resources: Interactive resource use had an estimated learning impact more than six times that of watching or reading, but it did not significantly predict dropout.Quiz scores and quiz participation, rather than interactive-resource use, significantly predicted dropout.
- Time-on-Task: Time spent on homework and labs predicted higher assignment achievement, while discussion-board and book time was less predictive or not statistically significant.Time on ungraded in-video quiz problems was more predictive than time on lecture videos.
- Additional Approaches: Predictive models have also used peer evaluation, intermediate correctness predictions, incremental training, and detailed IDE interaction data.These approaches extend learning-based prediction beyond conventional grades and platform activity.
3.8 Demographics-Based Models
Demographic features have mixed associations with MOOC success, but evidence suggests they generally add little predictive value beyond learner activity. Findings vary across outcomes, learner groups, course domains, and demographic variables.
- Demographics-based models use static learner attributes to predict success during a course.
- Prior coursework, parental engineering background, offline teacher contact, age, prior education, and prior MOOC experience were associated with selected achievement, completion, or dropout outcomes.
- Gender associations with activity and certification differed between science and non-science XuetangX courses.
- Comparing demographic findings across studies is difficult because analyses use different controls and outcomes.
- In related online education research, demographics were not associated with dropout, whose reported reasons varied across personal, job-related, program-related, and technology-related factors.
- Demographics often have limited predictive value compared with activity data, even early in a course when activity information is sparse.Demographic features provided no discernable improvement over activity-only models and degraded performance later in the course.
4 Synthesis: Trends in Predictive Models of Student Success in MOOCs
The survey finds concentration in MOOC data sources, features, outcomes, algorithms, and evaluation practices. Activity-based modeling dominates, while uneven coverage and metric choices limit comparison and leave important research areas underexplored.
- A small number of platforms and raw data sources support most surveyed MOOC predictive-modeling research.Coursera and edX dominate platform coverage, while non-English platforms such as XuetangX are less represented.
- Clickstreams are the dominant raw data source, but their complex semistructured formats require substantial human and computational effort to parse.
- Most clickstream modeling uses simple counts despite the temporal complexity of learner interaction logs.Additional data sources are increasingly combined with MOOC data, but privacy protections make integration difficult.
- Feature engineering is central to predictive performance, including when models use relatively simple algorithms.The survey highlights rigorous comparisons among activity, NLP, demographic, forum, and assignment feature sets as a continuing need.
- Activity-based feature extraction is the dominant approach because activity data are prevalent, fine-grained, and rich in behavioral patterns.
- Dropout prediction was more than twice as common as any other outcome, with 39 works predicting dropout or stopout.
- Tree-based and generalized linear models were most common, but nearly half of surveyed work used an algorithm not used elsewhere.Only random forests appeared in more than 10 surveyed works, limiting comparisons among specific tree algorithms.
- Evaluation metrics lacked strong agreement, and relying only on accuracy was discouraged because additional metrics require little extra computation.
5 Methodological and Research Gaps
The survey identifies methodological conditions that make MOOC prediction findings difficult to compare, generalize, and use in active courses. Major concerns involve filtered populations, weak evaluation and replication practices, and prediction settings that rely on unavailable future information.
- Differences in populations, metrics, and experimental practices make findings difficult to compare and undermine reliable identification of a MOOC modeling state of the art.
- Many experiments exclude over 80% of MOOC participants, limiting relevance to broad learner populations.Highly filtered samples may represent unusually fluent, motivated, or survey-completing learners rather than the overall course population.
- More than half of surveyed works evaluated only one MOOC, making generalization to other courses difficult.
- Uncorrected multiple comparisons and weak inferential practices create risks of Type I errors and reduce confidence in reported findings.The survey found little appropriate significance testing and no acknowledgement of multiple-comparison concerns in the reviewed work.
- Average cross-validation performance was used in 31 studies, but unadjusted comparisons can be inappropriate for predictive-model evaluation.
- The scarcity of replication prevents researchers from determining how robust many MOOC prediction findings are.The authors recommend direct replication on new, larger data alongside more rigorous inference and evaluation.
- Many surveyed experiments are not realistic for deployment because prediction uses labels or information unavailable at intervention time.Same-course post hoc prediction can be optimistically biased, while future-course prediction generally performs worse.
- Real-time prediction tasks require techniques aligned with information available during an active course, unlike explanatory analyses focused only on data understanding.
6 Opportunities for Future Research
The survey proposes future work that better models temporal behavior, connects prediction with explanation and learning theory, and evaluates success beyond course participation. These directions aim to improve the informativeness, interpretability, and scope of MOOC student models.
- Future studies should use large, unfiltered, multi-MOOC populations and rigorous evaluation tests to compare features and algorithms.
- The survey identifies four gaps: temporal modeling, bridging predictive and explanatory models, theory-building, and long-term student-success modeling.
- MOOC prediction is inherently temporal because learner behavior, available data, and course activity evolve throughout a course.
- Weekly feature sets capture broad time periods but do not explicitly model complex temporal patterns.
- More extensive sequence, survival, and time-series modeling could uncover informative patterns and improve predictive performance.
- Combining accurate predictive models with interpretability methods could support both theory-building and actionable student interventions.
- MOOC predictive research has contributed relatively little to learning theory and should be grounded more often in established theories of learning or engagement.
- Long-term evaluation should connect MOOC performance with career or academic outcomes, although privacy protections and optional questionnaires create barriers.
7 Conclusion
The survey finds that MOOC predictive-modeling research has expanded substantially but adopted robust, generalizable, and actionable methodologies only incompletely. It recommends stronger evaluation, broader populations, realistic contexts, and replication tools to support useful student-success models.
- The review identifies incomplete adoption of methodologies needed for predictive models to be accurate, interpretable, and generalizable.
- The authors recommend robust model evaluation, broader experimental populations, and realistic experimental contexts.These priorities are presented as necessary for developing accurate, actionable, and theory-building models.
- Restrictive populations can limit applicability to large learner segments, while unrealistic contexts may require data unavailable when predictions are made.
- Replication studies using more rigorous statistical evaluation and tools for benchmarking published models are identified as valuable future contributions.
- The survey reports a literature matrix spanning predictive-modeling studies and records missing or inapplicable information explicitly.The matrix uses ellipses for unreported information, asterisks for nonapplicable fields, and MANY for at least 10 unique values.