Source-linked AI summary

Correctness Is Not Homogeneous Evidence: A Correctness-conditioned Evidence-aware Knowledge Tracing Model

Fuzheng Zhao

arXiv:2608.22267v1cs.HC

TL;DR

The study examines whether behavioral condition scores are associated with future same-skill performance after correctness is held fixed. It evaluates CE-KT’s recurrent gating design and finds stronger predictive performance overall, while proxy-specific and calibration benefits remain limited.

  • Problem

    The study asks whether behavioral condition scores are associated with future same-skill performance within groups separated by current response correctness.

  • Method

    CE-KT uses behavioral process information and a recurrent full-state gate with separate correct- and incorrect-response pathways, evaluated through comparisons, negative controls, ablations, and sensitivity analyses.

  • Results

    CE-KT outperformed ordinary behavioral-input models, output-only modulation, NoGate, and SAKT performance context, while calibration worsened for the post-error improvement candidate group.

  • Takeaways & Limitations

    The full recurrent gate configuration is a strong predictive structure over mechanism-matched baselines, but the evidence does not uniquely attribute improvements to proxy scores, correctness-specific routing, or one modulation operation.

  • Takeaways & Limitations

    CE-KT did not preserve the higher subsequent-performance candidates within incorrect interactions, and improvements on the correct side were small.

Abstract

from arXiv · show

Knowledge tracing models usually use response correctness as a central observation for estimating students' latent knowledge states. However, the same correct or incorrect response may arise from different behavioral contexts, such as rapid guessing, hint use, or repeated attempts. Treating correctness as uniformly informative may therefore introduce ambiguity into recurrent state updates. This study proposes Correctness-conditioned Evidence-aware Knowledge Tracing (CE-KT), which uses observable response-process features to condition how correctness is written into recurrent states. CE-KT derives weakly supervised behavioral proxy scores from response time, hint use, attempt count, and behavioral history. These scores are used as behavioral signals, not as direct measures of mastery, response quality, or cognitive state. CE-KT then uses current correctness to select a correct-response or incorrect-response gate. The selected gate modulates both the LSTM hidden state and cell state, and the modulated states are fed back into later recurrent updates. Experiments on ASSISTments data show that behavioral condition scores are associated with future same-skill performance within fixed correctness groups, especially for incorrect interactions. CE-KT generally outperforms several behavior-fusion alternatives on the main predictive metrics, although its calibration advantage is not consistent. Ablation analyses provide partial support for correctness-specific recurrent modulation and recurrent feedback. These findings suggest that behavioral information can help condition the interpretation of response correctness in knowledge tracing, but the proposed proxy scores should not be treated as direct evidence of true mastery or causal learning effects.

5.1 RQ1: Behavioral Heterogeneity Within Fixed Correctness

Behavioral condition scores remained associated with future same-skill performance even after current response correctness was fixed, but the associations differed sharply between correct and incorrect interactions.

  • -0.0111 lower future same-skill performance for high- versus low-score correct interactions, with p = .038.The 95% confidence interval was [-0.0207, -0.0005], and the magnitude was about 1.11 percentage points.
  • 0.0907 higher future same-skill performance for high- versus low-score incorrect interactions, with p = .006.The 95% confidence interval was [0.0151, 0.1995].
  • Behavioral heterogeneity was supported within fixed correctness groups, while the positive and negative sides differed in direction, magnitude, and interpretability.

5.2 RQ2: Stratified Prediction Bias Across Behavioral Proxy Scores

Standard KT architectures showed systematic residual differences across behavioral-proxy strata, while CE-KT reduced but did not eliminate these stratified biases.

  • The mixed Q1-Q4 strata examine residual variation across behavioral-proxy regions rather than forming a unified response-quality scale.The positive- and negative-side scores were trained from different weak-supervision tasks and are not directly comparable.
  • 0.0760 was the largest maximum cross-stratum residual gap among the standard architectures, observed for SAKT.DKT, DKVMN, and AKT-S had maximum gaps of 0.0649, 0.0683, and 0.0509, respectively.
  • Across standard KT architectures, Q1 showed overprediction, whereas Q3 or Q4 showed underprediction to varying degrees.
  • 0.0498 was CE-KT’s maximum cross-stratum residual gap, the smallest among the five models.Its cross-stratum Wald test remained significant at p = .000823, so residual differences were reduced but not eliminated.

5.3 RQ3: Recurrent Correctness-Specific State Modulation Versus Other Behavior-Fusion Strategies

RQ3 evaluates whether conditioning recurrent state updates on correctness and behavioral context improves prediction over alternative behavior-fusion strategies. CE-KT generally improves ranking and probabilistic error, while calibration gains and unique attribution to proxy scores or routing remain unsupported.

  • Main comparisons: CE-KT outperformed DKT, DKT+FullBehavior, CE-KT-OutputOnly, and CE-KT-NoGate on AUC, LogLoss, and Brier score.Mean ECE was also lower, but its bootstrap confidence intervals included zero.
  • Main comparisons: Relative to DKT, CE-KT improved AUC by 0.0294 and reduced LogLoss and Brier score by 0.0168 and 0.0070.Relative to DKT+FullBehavior, AUC improved by 0.0206, while LogLoss and Brier score decreased by 0.0128 and 0.0053.
  • Main comparisons: Relative to CE-KT-OutputOnly, CE-KT improved AUC by 0.0208, LogLoss by 0.0116, and Brier score by 0.0046.Relative to CE-KT-NoGate, AUC improved by 0.0295, LogLoss decreased by 0.0169, and Brier score decreased by 0.0070.
  • Negative controls: CE-KT-RandomProxy was not reliably different from CE-KT, so the results do not establish a stable additional predictive increment from the three proxy scores when full behavioral features are available.The study therefore supports the full recurrent gate configuration over mechanism-matched baselines without uniquely attributing gains to the proxy scores, routing, or one modulation operation.
  • Structural ablations: CE-KT improved AUC over CE-KT-SharedGate, CE-KT-PosOnly, and CE-KT-NegOnly by 0.0133, 0.0120, and 0.0129, respectively.LogLoss and Brier score were also reliably lower, but the ablations had fewer parameters than the complete model.
  • Structural ablations: CE-KT improved AUC by 0.0110 over CE-KT-GammaOnly and by 0.0042 over CE-KT-BetaOnly.The scale-plus-shift combination mainly improved ranking and probabilistic error, while calibration gains were not stable across all ablation comparisons.
  • Gate-input sources: CE-KT-ProxyOnly achieved an AUC of 0.7077 versus 0.6936 for DKT+FullBehavior, with an AUC difference of 0.0140 and a 95% confidence interval of [0.0114, 0.0166].LogLoss and Brier score also decreased, but stable calibration improvement was not established.
  • Gate-input sources: CE-KT-FullContext improved AUC over CE-KT-ProxyOnly by 0.0035 and reduced LogLoss and Brier score by 0.0027 and 0.0012.The original behavioral features therefore provided additional ranking and probabilistic-error gains beyond the proxy scores.

5.4 RQ4: Mechanism Diagnosis Within Correct and Incorrect Interactions

RQ4 diagnoses whether CE-KT differentiates interactions sharing the same correctness label according to behavioral proxy scores. It shows the clearest correction for low-score incorrect interactions, but only partial support for the broader mechanism.

  • Diagnostic setup: RQ4 compares observed target correctness rates, model predictions, residuals, and bias reduction across four operational behavioral groups.The groups are proxy-based labels and do not identify externally validated cognitive types or causal effects.
  • Correct interactions: For correct interactions, observed target rates were 0.7785 and 0.7916 for low- and high-score groups, a difference of about 1.31 percentage points.CE-KT increased predictions for both groups and reduced absolute bias by 0.34 and 0.17 percentage points, respectively.
  • Incorrect interactions: For incorrect interactions, the low-score error group had an observed target rate of 0.4946, while DKT predicted 0.6106.DKT overpredicted by 11.59 percentage points in this operational group.
  • Incorrect interactions: CE-KT reduced the low-score error prediction to 0.5533, lowering overprediction to 5.87 percentage points and absolute bias by 5.73 percentage points.The incorrect-response gate therefore suppressed DKT’s overoptimistic prediction for this proxy-defined group.
  • Incorrect interactions: The post-error improvement candidate group had an observed target rate of 0.6278, but CE-KT decreased prediction from DKT’s 0.6171 to 0.5903.Absolute bias consequently increased from 1.07 to 3.75 percentage points, so the higher subsequent-performance signal was not captured in the expected direction.
  • Overall diagnosis: Across the four groups, CE-KT had the smallest absolute residual for high-score correct and low-score error groups, whereas DKT was best for the post-error improvement candidate group.RQ4 therefore provides partial support and boundary evidence for behavior-conditioned gate modulation.

5.5 Robustness and Sensitivity Analyses

Sensitivity analyses found that the main model conclusions were stable across rapid-guessing and post-error thresholds, while gate amplitude produced a performance–calibration trade-off. Proxy rankings were generally consistent across several construction windows and decay weights.

  • Threshold sensitivity: Changing rapid-guessing percentiles from 5% to 1% or 10% did not produce stable changes in AUC, LogLoss, Brier score, or ECE.Bootstrap confidence intervals relative to the default included zero.
  • Threshold sensitivity: Changing the post-error improvement threshold from 0.20 to 0.10 or 0.30 left model metrics stable and preserved positive negative-side future-performance differences.The differences remained between 0.0872 and 0.0881, with confidence intervals above zero.
  • Gate-amplitude sensitivity: An amplitude of 0.25 reduced AUC to 0.7080 and worsened LogLoss, Brier score, and ECE, indicating insufficient recurrent modulation.The result suggests that overly weak gating was ineffective.
  • Proxy-construction sensitivity: Proxy rankings correlated 0.9319–1.0000 with the default local-gain definition across tested windows and decay weights, although candidate coverage changed.Post-error improvement candidates represented 21.41%–28.44% of incorrect interactions.

6. Discussion and Conclusion

The discussion supports behavioral heterogeneity within fixed correctness groups and generally favors recurrent state modulation over alternative behavior-fusion strategies. However, proxy validity, calibration, scope, and causal interpretation remain limited, so CE-KT is presented as an explicit conditioning pathway rather than a measure of true cognitive state.

  • Discussion and conclusion: CE-KT generally outperformed ordinary behavioral-input models, output-only modulation, NoGate, and SAKT performance context, while reducing bias in most conditional-score regions.The mixed Q1–Q4 analysis found improved overall probabilistic metrics but no improvement in Q2.
  • Discussion and conclusion: Incorrect interactions showed expected long-range performance differences associated with behavioral condition scores, whereas correct-side differences were small, negative, and not evidence of stable mastery.RQ1 therefore provides clearer support for behavior-related heterogeneity on the incorrect side.
  • Mechanism and proxy findings: Recurrent state modulation generally outperformed ordinary behavioral input and output-only modulation, but learned proxies lacked a stable marginal increment over CE-KT-RandomProxy.CE-KT-ProxyOnly nevertheless remained significantly better than DKT+FullBehavior after full behavioral inputs were removed.
  • Mechanism and proxy findings: CE-KT reduced overprediction for the low-score error group but lowered predictions for post-error improvement candidates and failed to preserve their higher observed target rate.This indicates asymmetric learning of negative modulation across incorrect-interaction groups.
  • Limitations and future work: The study used high-activity students and frequent ASSISTments skills, heuristic proxies, incomplete action-order reconstruction, and observational data that cannot support causal inference.Additional limitations included absent student-level cross-fitting, incomplete covariate adjustment, and unstable CE-KT-RandomProxy differences.
Loading 2608.22267v1…