Source-linked AI summary

Causal Modelling of Support Interventions for Student Competency Assessment

Francesca Mangili, Alessandro Antonucci, Rafael Cabañas

arXiv:2608.24632v1cs.AI

TL;DR

Educational assessment needs models that represent competencies and interventions beyond associative belief updating. This paper develops a structural causal modelling protocol with expert-elicited equations and illustrates interventional and counterfactual analyses on algorithmic-skill assessment data.

  • Problem

    Standard psychometric and associative models do not directly support the interventional and counterfactual reasoning needed to analyse interventions such as hints.

  • Method

    The paper specifies PSCMs with expert-elicited structural equations for skills, help-seeking, answers, and noise, then uses assessment data to derive compatible FSCMs.

  • Results

    The illustrative assessment supports population skill and help analyses, student profiling, and counterfactual estimates of question performance under alternative assistance.

  • Takeaways & Limitations

    Structural causal modelling offers a framework for separating help-seeking behaviour from proficiency and analysing intervention effects in educational assessment.

  • Takeaways & Limitations

    The learner model is illustrative rather than validated: it was tested on a small dataset against a directly learned BN only for predictive accuracy.

Abstract

from arXiv · show

Accurate assessment of student competencies is essential for enabling educators to identify individual needs, design targeted interventions, and evaluate the effectiveness of educational strategies. Empirical assessment procedures are typically grounded in psychometric models, such as item response theory, which relate student competence levels to performance on assessment tasks. In this paper, we advocate adopting a structural causal modelling approach to educational assessment, moving beyond probabilistic belief updating toward a framework that explicitly supports interventional and counterfactual reasoning. We propose a corresponding protocol for its construction and analyse the practical relevance of forms of reasoning that remain inaccessible to standard associative models, including the explicit modelling of interventions such as hints and the related counterfactual scenario analysis. Although our protocol requires the structural equations to be elicited from experts, the necessary information is purely logical and does not rely on probabilistic, less tenable assumptions. We illustrate the approach using data from an assessment that employs complex tasks designed to measure compulsory school student algorithmic skills.

1. Introduction

Educational assessment models support competency monitoring and personalised interventions, but standard approaches have limitations in representing mixed competencies and causal reasoning. The paper therefore proposes moving learner modelling from Bayesian networks to structural causal models.

  • Accurate competency assessment supports population-level skill monitoring and personalised interventions such as tutoring.
  • Item response theory models single competences, while multidimensional extensions are not designed to represent mixtures of competencies within individual items.
  • Probabilistic graphical models provide flexible and interpretable representations of relationships between multiple competencies and observable behaviours.
  • The paper evolves learner-modelling protocols from Bayesian networks to structural causal models to support interventional and counterfactual reasoning.The proposed transition is motivated by the bipartite skill-question structure and the correspondence between noisy gates and structural equations.
  • The paper presents a causal elicitation protocol, discusses counterfactual inferences, and illustrates the approach on algorithmic-skill assessment tasks.

2. Background

The background introduces structural causal models as extensions of Bayesian-network reasoning, distinguishing latent exogenous variables from observed endogenous variables. Unlike observational queries, SCMs support interventions and counterfactual queries, although latent-variable estimation is required in practice.

  • In an SCM, structural equations map exogenous inputs to endogenous variables, whereas conditional probability tables represent probabilistic relationships.
  • The model treats student answers and hint usage as observed endogenous variables generated by latent skill, luck, and help-seeking variables.
  • A fully specified SCM combines structural equations with marginal probability distributions over exogenous variables and induces a Bayesian-network factorisation.
  • An intervention replaces a variable’s structural equation with a constant, enabling post-interventional queries and counterfactual scenarios that differ from observed values.
  • Because exogenous variables are unobserved, datasets provide observed-variable distributions, while causal EM can derive compatible fully specified SCMs for counterfactual inference.

3. Educational Tests by PSCMs

The protocol constructs educational-test PSCMs by defining observable answers and help requests as functions of latent skills, propensity, luck, and deviations. Domain experts specify the structural relationships that give these latent states operational meaning.

  • The protocol begins with model variables and structural-equation elicitation for educational tests.
  • Answer equations exclude previous answers so the model remains suitable for adaptive tests whose question sequence may change.
  • Skills are latent exogenous variables that may be Boolean, ordinal, or continuous, and a shared skill can confound multiple questions.
  • Domain experts identify the skills relevant to each question, while propensity parents hint nodes and luck or deviation variables represent unexplained variability.Luck and deviations account for slips, guessing, contextual effects, and other variability; their state spaces should allow observable outcomes to remain reachable.
  • In the two-question example, Q2 is modelled as Q2 = (S1 ∧S2) ∨LQ2, allowing a correct answer through either both skills or luck.

4. Educational Assessments by Causal Inference

The causal assessment framework combines PSCM specifications with data-derived compatible FSCMs to support group-level and individualised profiling, intervention analysis, and counterfactual reasoning. It exposes both informative findings and substantial uncertainty in the illustrative assessment.

  • Group Inferences: Causal EM derives a compatible collection of FSCMs from a PSCM and empirical observed-variable distribution, after which queries are answered over that fixed collection.
  • Group Inferences: Group inferences estimate population skills, evaluate question quality, and analyse the influence of help variables on assessment outcomes.
  • Group Inferences: Counterfactual necessity and sufficiency probabilities quantify whether help was necessary for observed success or could enable success under lower help.
  • Group Inferences: Observational relationships such as P(Q|HQ) cannot isolate help’s causal effect because proficiency can confound help-seeking and performance.
  • Group Inferences: Twin networks compute generalised necessity and sufficiency by duplicating real and hypothetical scenarios while sharing exogenous parents except for intervened variables.
  • Personalised Inferences: Individualised inferences estimate student profiles and counterfactual performance under alternative help levels from each student’s observed answers and assistance.
  • Personalised Inferences: Counterfactual help comparisons are meaningful only when hypothetical help and answer levels change consistently with the expected effect of help.

5. Use Case

The use case applies the proposed causal model to algorithmic-skill assessment with complex cross-array tasks, encoding how skills, help, luck, and question outcomes interact.

  • The CAT battery assesses compulsory-school pupils’ algorithmic skills using complex cross-array tasks and data from 109 students.
  • Question outcomes combine algorithmic skill, help alignment with student autonomy, luck-related shifts, and clipping to the permitted complexity range.
  • Hint variables normally reflect each student’s propensity for help, with question-specific deviations increasing or decreasing hint usage within its admissible range.

6. Results

The results use predictive checks, group and individual posterior inferences, and counterfactual analyses to examine the CAT assessment and inferred learner profiles. They reveal informative intervention and skill relationships alongside substantial uncertainty in some competency estimates.

  • The expert-elicited FSCM is evaluated against a data-learned BN using five-fold cross-validation, held-out log-likelihood, and posterior predictions for answers and hints.
  • Group Inferences: P(Saut = feedback) ∈ [0.92, 1] while P(Saut = scheme) ∈ [0, 0], indicating highly likely low autonomy and an essentially impossible intermediate level.
  • Group Inferences: The first six questions have very-good-luck probability at most 0.15, whereas Q7–Q9 can exceed 0.30, suggesting poorer calibration in the latter tasks.
  • Group Inferences: The probability of necessity for Salg = 1D in Q6 = 1D ranges from 0.11 to 0.95, making that inference nearly vacuous.
  • Group Inferences: The lower bound for sufficiency of H6 = feedback for Q6 = 0D is 0.94, while Salg = 1D is sufficient for Q6 = 2D with probability at most 0.06.
  • Group Inferences: As question complexity rises, necessity of help and skill tends to increase while sufficiency tends to decrease.
  • Individual Inferences: Completing every task at maximum level cannot establish the highest skill with high confidence, partly because luck strongly influences algorithmic-skill inference.
  • Individual Inferences: For the illustrated student, removing scheme help from Q5 still yields probability one for a 2D answer, whereas adding feedback to Q12 leaves uncertainty between 1D and 2D.

7. Limitations and Conclusions

The paper presents SCMs for educational assessment as a methodological exploration, illustrating causal, interventional, and counterfactual reasoning while identifying substantial validation and deployment limitations.

  • Conclusions: SCMs represent causal relationships and support interventional and counterfactual reasoning in learner modelling.The approach is illustrated through assessment objectives including instrument evaluation, intervention analysis, and learner profiling.
  • Conclusions: Counterfactual reasoning can help disentangle help-seeking behaviour from underlying proficiency when students influence task conditions.The paper identifies adaptive assessment as a possible application in which support decisions could use interventional queries.
  • Limitations: The learner model is illustrative because its structure and parameterisation were tested only on a small dataset and evaluated only for predictive accuracy.The study was compared against a BN learned directly from data, without full validation of the causal model.
  • Limitations: The paper does not claim evidence about impacts on decision-making or the correctness of its counterfactual estimates.The authors instead present and exemplify a structural causal modelling approach to support further methodological development.
  • Limitations: Further work should address broader validation, sensitivity analysis, adaptive-testing deployment, and scalability to larger skill sets.The authors identify larger datasets or simulations and real-time deployment as necessary continuation points.

Appendix A. Inferential Complexity

The protocol’s computational inferences use BN inference within models whose graph topology yields question-specific cliques, with cutset conditioning simplifying their connections.

  • Inferential Complexity: Both causal EM and subsequent inference over the returned FSCMs rely on BN inference with a topology analogous to the paper’s learner model.The relevant graph structure is analysed around each question and its parent skills.
  • Inferential Complexity: For each question Q, moralisation around its parents yields cliques (Q, H_Q, S_Q, L_Q) and (H_Q, S_Q, W_Q, R).S_Q denotes the skills relevant to answering Q.
  • Inferential Complexity: The resulting graph is not a clique tree because hint-related cliques share R and different questions connect when their relevant-skill sets overlap.These two sources create additional edges across the graph.
  • Inferential Complexity: Cutset conditioning on R removes the edges caused by R appearing in every hint-associated clique.Variables L_Q and W_Q can also be handled separately because each appears in only one clique.

Appendix B. Details on the Use-Case Assessment Protocol

The CAT is an unplugged assessment of algorithmic skills using 12 cross-array tasks in which pupils verbally instruct a tutor to reproduce target arrays with optional assistance.

  • Assessment Protocol: The CAT assesses the algorithmic-skills component of computational thinking in pupils aged 3 to 16.It uses a battery of unplugged tasks.
  • Assessment Protocol: Students verbally instruct a tutor to reproduce 12 target arrays, each containing 20 coloured circles, on a blank cross array.A barrier prevents students from seeing how the tutor is colouring.
  • Assessment Protocol: Students may request a blank cross array to support verbal instructions by pointing to circles to colour.This is one of the two forms of assistance available during the assessment.
  • Assessment Protocol: Students may remove the barrier to obtain visual feedback on the result of their instructions.This constitutes the second available form of assistance.
  • Assessment Protocol: Figure 5 depicts the blank cross array alongside the 12 CAT schemes used as target arrays.The figure caption identifies the left panel as the blank cross array.
Loading 2608.24632v1…