Source-linked AI summary

Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

Aaron Ceross

arXiv:2609.09788v1cs.LG

TL;DR

The paper asks how to compare retraining policies when subgroup disparities accumulate across the entire sequence of deployed models. It evaluates scheduled, loss-triggered, and subgroup-gap-triggered updating against freezing using paired simulated and replay analyses. All three policies had lower mean cumulative disparity in the evaluated simulations, but trajectory conclusions and policy rankings depended on measurement population, actions, drift conditions, and the non-confirmatory study design.

  • Problem

    Retraining-policy evaluation needs to account for subgroup error rates across all models and deployment windows rather than only post-update performance.

  • Method

    The study compares complete scheduled, loss-triggered, subgroup-gap-triggered, and frozen policies on shared observations and delayed labels using paired cumulative absolute TPR and FPR gaps.

  • Results

    All three updating policies had lower mean cumulative disparity than freezing in two simulated drift regimes, with reductions of 0.04 to 0.88 percentage points in average gap per window.

  • Takeaways & Limitations

    Policy choice requires group-specific rates, action distributions, and an explicit evaluation population because lower mean disparity alone does not establish stable trajectory-level benefit or improved group outcomes.

  • Takeaways & Limitations

    The evidence is limited to specified drift regimes, one learner and horizon, exploratory replay settings, and shared observations with complete delayed labels; all analyses are non-confirmatory.

Abstract

from arXiv · show

Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affect these comparisons. In a follow-up sample of 400 new trajectories per condition across two simulated drift regimes, all three policies had lower mean cumulative disparity, equivalent to reductions of 0.04 to 0.88 percentage points in the average gap per window. Evaluating the unchanged models against the known generating distributions preserved all mean directions, but finite-window and population comparisons agreed on whether updating increased, reduced or left cumulative disparity unchanged in 69 to 92 percent of trajectories. Under subgroup-specific drift, smaller true-positive-rate gaps accompanied lower sensitivity in both groups. In an exploratory American Community Survey replay, person weighting reversed all three race false-positive-rate mean comparisons without changing predictions or actions; all three weighted intervals included zero. Policy comparisons require group-specific rates, action distributions, and an explicit evaluation population alongside mean disparity. These analyses are non-confirmatory. Shared replay requires policy-independent observations and complete labels after the specified delay.

1 Introduction

The paper evaluates retraining policies by comparing subgroup error-rate disparities across the full sequence of deployed models and intervening windows. It argues that policy comparisons must jointly represent subgroup rates, actions, observations, and the evaluation population.

  • Updating can change cumulative disparity because the models used between refits contribute to the outcome, not only the model fitted immediately after retraining.
  • Scheduled, loss-triggered, and subgroup-gap-triggered policies are compared with freezing using paired cumulative absolute TPR and FPR gaps over deployment windows.
  • The simulation estimates lower mean cumulative disparity for scheduled and monitored retraining in both evaluated drift regimes, while population evaluation preserves mean directions but changes many trajectory signs.
  • Under combined drift, loss-triggered retraining uses fewer refits than scheduled retraining but produces greater TPR and FPR disparity.
  • Person weighting in the exploratory Census replay reverses all three race-FPR mean comparisons without changing predictions or actions.
  • Interpretation requires the policy action rule, subgroup outcome measure, and explicit evaluation population, while shared replay assumes policy-independent observations and complete delayed labels.

4 Experimental design

The experiments use controlled simulated drift with fixed learner and monitoring choices, paired follow-up trajectories, and retrospective statistical checks. The design also documents calibration, alternative analyses, and non-confirmatory limitations.

  • The simulation uses ten evaluation windows, 5,000 observations per window, a binary group with prevalence 0.2, and a separate no-drift initial-training window.
  • A fixed L2 logistic-regression learner with threshold 0.5 receives features without group membership or group-feature interactions.
  • Loss and subgroup-gap CUSUM monitors trigger refits after strict threshold crossings, with upward accumulation for loss and two-sided accumulation for signed gaps.
  • The follow-up generates 400 trajectories per drift condition, with separate samples for random-reference and hindsight analyses and additional trajectories for action matching.
  • The reanalysis uses a centred studentised bootstrap with B = 10,000 and Holm adjustment within each eight-comparison run family.
  • The analyses are non-confirmatory because the follow-up hypothesis followed inspection of initial results, the execution failed a fixed-source requirement, and measurement checks were retrospective.

5 Retraining effects on subgroup disparity

Across follow-up simulations, scheduled, loss-triggered, and gap-triggered retraining reduced mean cumulative subgroup disparity versus freezing, but individual trajectory results and predictive-rate trade-offs varied by drift regime and evaluation window.

  • All twelve follow-up policy–condition–rate mean differences were negative, with mean reductions of approximately 0.04–0.88 percentage points per evaluation window.The centred studentised reanalysis supported all eight monitored-policy reductions at Holm level 0.05, with adjusted p-value estimates from 0.00080 to 0.01340.
  • Temporal and between-trajectory variation: Although ten-window means were negative, 20–45% of finite-window trajectory contrasts were positive, showing that endpoint averages can conceal earlier increases.Five-window mean contrasts ranged from −0.009 to +0.087 percentage points, whereas ten-window means were all negative.
  • Population integration preserved all twelve negative mean directions, but finite-window and population trajectory signs agreed in only 69.25–92.25% of cases.Observed-minus-population mean differences were at most 0.032 percentage points in absolute magnitude.
  • Group-specific predictive performance: Under subgroup-specific drift, smaller TPR gaps accompanied lower sensitivity in both groups, so reduced disparity did not by itself establish improved welfare.For the gap trigger, TPR changes were about −0.87 and −0.82 percentage points across the two groups, while FPR changes were about −0.69 and −0.81 percentage points.
  • Comparisons with scheduled retraining: Under combined drift, the loss trigger used about 1.8 fewer refits than cadence but produced 0.237 percentage points greater TPR disparity and 0.357 percentage points greater FPR disparity.The gap trigger used 0.075 more refits and had 0.087 and 0.185 percentage points lower TPR and FPR disparities, respectively.

6 Action patterns and schedule comparisons

Action incidence, timing, and count distributions shape cumulative disparity comparisons, while finite-window hindsight gains may not survive population evaluation.

  • 6.1 Action incidence and conditional disparity: The loss-triggered policy’s adverse share is 14.3% across states but 62.5% among acting states, because unconditional summaries partly reflect inaction.It has positive differences in five of 35 states and acts in eight.
  • 6.2 Alarm direction and reference persistence: Under combined drift, the gap policy’s FPR alarms usually accompany smaller absolute gaps: 98.6% improve versus the initial reference and 64.4% versus the preceding observation.The two-sided signed-change rule can retrigger after improvement when the stream remains displaced from its original reference.
  • 6.3 Action-count matching: Similar mean refit counts can conceal different action patterns: any-refit probabilities differ by up to 15.5 percentage points and total variation reaches 0.207.The disparity contrast therefore combines timing, realised counts, and the probability of acting.
  • 6.3 Action-count matching: The selected loss-threshold approximation fails its total-variation and any-refit-difference criteria, so the planned disparity comparison was withheld.Its total variation is 0.290 against a 0.20 bound, and its any-refit difference is 0.145 against a 0.10 bound.
  • 6.4 Refit timing and count in the hindsight benchmark: Most mean hindsight advantage comes from rescheduling at the same refit count, while allowing fewer refits adds 0.00077 to 0.00906 cumulative units.A strictly positive fewer-refit contribution occurs in 6.25–38.0% of trajectories.
  • 6.4 Refit timing and count in the hindsight benchmark: Population re-evaluation reduces every mean policy–oracle disparity difference, with negative differences in 5.0–20.5% of trajectories.Finite-window schedule selection can exploit measurement variation and need not improve population disparity for each trajectory.

7 Evaluation-population sensitivity in Census data

The Census replay evaluates fixed predictions and actions under alternative record and person-weighted populations. Person weighting reverses all three race-FPR mean comparisons, but the weighted intervals include zero and the analyses do not establish a population effect.

  • 7 Evaluation-population sensitivity in Census data: The Census replay uses repeated population samples rather than follow-up of the same individuals, with equal state weighting and an absent 2020 wave.Eligibility and window construction impose additional sample restrictions.
  • 7 Evaluation-population sensitivity in Census data: Person weighting reverses all three race-FPR mean comparisons without changing fitted models, predictions, monitoring history, or refit actions.Cadence changes from −0.05585 to +0.03741, loss-trigger from −0.01040 to +0.00516, and gap-trigger from −0.06091 to +0.04688 cumulative units.
  • 7 Evaluation-population sensitivity in Census data: All three person-weighted race-FPR intervals include zero, so the reversals are mean point-estimate changes rather than supported nonzero effects.The sex-cadence TPR mean also changes sign.
  • 7 Evaluation-population sensitivity in Census data: For race-FPR, person weighting moves the state contrast upward in 26 of 29 cadence states and 24 of 29 gap-trigger states.All three mean reversals survive removing any single state.
  • 7 Evaluation-population sensitivity in Census data: The weighting comparison concerns which evaluation population a disparity comparison describes, not a population effect, because state resampling does not account for interstate dependence.The analyses use descriptive intervals and do not establish a population effect.

8 Discussion

The discussion shows that maintenance-policy comparisons depend on action behavior, disparity measurement, evaluation populations, and the limits of the simulated and exploratory settings. Lower mean cumulative disparity did not by itself establish stable trajectory-level benefit or improved outcomes for each group.

  • Implications for maintenance decisions: Under combined drift, loss-triggered retraining used about 1.8 fewer refits than scheduled retraining but had 0.237 and 0.357 percentage-point greater average TPR and FPR gaps.The trade-off concerns both disparity and the unmeasured operational costs of retraining.
  • Implications for maintenance decisions: Gap-monitor FPR alarms can occur while absolute disparity improves because the monitor responds to departures from a fixed signed reference.Repeated alarms may therefore reflect the two-sided change definition and persistent reference rather than worsening absolute disparity.
  • Implications for maintenance decisions: Reporting refit probability and conditional outcomes is necessary when many trajectories retain the initial model and contribute zero disparity contrasts.Matching mean refit counts does not ensure comparable count distributions or action probabilities.
  • Measurement and evaluation: Population evaluation preserves mean directions but changes many trajectory signs, making average comparisons more stable than identifying individual trajectories with increased disparity.The comparison evaluates the same prediction sequence at a more precise measurement level rather than redefining the policy.
  • Limitations and future work: The controlled evidence covers two specified drift environments, one learner, a ten-window horizon, and non-confirmatory analyses with shared observations and complete delayed labels.Policy-dependent observations or selective labels would require an identification design representing feedback.
  • Interpretation of disparity: Under subgroup-specific drift, smaller TPR gaps accompany lower sensitivity in both groups, so disparity reduction does not establish improved predictive outcomes.The result links subgroup disparity to component rates rather than treating the gap alone as sufficient.

B Population integration and numerical validation

Population integration reconstructs each archived model and evaluates its rates against the known generating distribution, while numerical checks verify the calculation. This separates evaluation noise from policy behavior without replacing the original monitoring streams.

  • Population integration: All 800 trajectories passed reconstruction checks, with recomputed empirical TPR and FPR gaps agreeing within 10^-9 and zero maximum observed discrepancy.Models are reconstructed at archived refit boundaries and reused across schedules.
  • Rate calculation: Numerical integration computes subgroup rates from the joint normal outcome-logit and classifier-score distributions, including conditional outcome probabilities and threshold events.The calculation separately supplies positive-prediction and joint positive-outcome probabilities for FPR and TPR components.
  • Numerical validation: The population calculation was cross-checked with four scrambled Sobol estimates using 2^18 Gaussian feature draws per group, with maximum mean-rate difference 0.000407.The probe check is a separate numerical validation, not a policy recalibration.
  • Population integration: Population integration conditions on the original training samples, models, alarms, and refit schedules, changing only evaluation rather than policy decisions.Replacing noisy monitoring streams with population gaps would define a different policy requiring separate evaluation.
  • Interpretation: Finite-window and population comparisons represent the same prediction sequence at different measurement levels, while observed trajectory differences remain subject to finite-window measurement error.Population integration is available here because the generating distribution is known; observed-data replays retain uncertainty in estimated subgroup rates.

C Directional inference and statistical validation

The statistical-validation section revises the historical bootstrap analysis using centred and studentised paired differences while retaining the original directions and comparison families. It also records the limits of the historical evidence and unassessed finite-sample guarantees.

  • Current inference: The follow-up studentised-bootstrap tests retain the original directional comparisons and eight-test family structure while applying the revised procedure to archived paired differences.Each test uses 10,000 resamples with inclusive ties and explicit handling of zero-variance resamples.
  • Statistical validation: The bootstrap plus-one convention and Holm adjustment provide Monte Carlo estimates and multiplicity control, but zero tail events do not establish an exact hypothesis-test bound.The corresponding two-sided 95% Monte Carlo interval has an upper endpoint of approximately 0.000369.
  • Initial and follow-up inference: The initial-sample analysis tested higher expected disparity under monitored retraining, and none of its eight adjusted tests rejected.The initial point estimates and nominal intervals describe Run A, while current centred studentised tests are reported separately.
  • Validation limits: Historical Monte Carlo checks evaluate the original percentile-tail procedure rather than validating the current studentised bootstrap or its behavior under arbitrary distributions.No BCa coverage study was performed, and the original power target must not be reinterpreted as current validation.

D Initial-sample estimates and policy outcomes

The initial-sample estimates describe cumulative disparity contrasts, refit timing, action counts, and predictive-performance changes across the policy comparisons. Cadence is descriptive, while monitored-policy summaries separately report alarms and refits.

  • Initial-sample inference: Run A tested whether monitored retraining increased expected disparity, and none of its eight adjusted tests rejected that direction.The follow-up sample is analyzed separately in Run B.
  • Cumulative disparity estimates: Table 10 reports direct trajectory means, relative disparity reduction, and the trajectory share with positive cumulative disparity contrast.These estimates summarize the initial-sample policy comparisons against the retained model.
  • Cumulative disparity estimates: Figure 5 applies the follow-up cumulative-disparity calculation to Run A, with cadence refits at t = 3, 6, and 9.The figure also distinguishes subgroup-specific and combined drift archives.
  • Refit timing and counts: Figure 6 shows the proportion of trajectories not yet refitted through each boundary, with cadence first acting at window 3 and endpoints including paths that never act.Table 11 complements this with refit-count distributions over all 400 trajectories.
  • Refit timing and counts: Tables 12 and 13 summarize first-alarm timing and no-alarm counts for monitored policies across 400 trajectories in Runs A and B.These alarm summaries do not apply to frozen or scheduled retraining.
  • Predictive performance: Under B4, all policies lower both groups’ mean TPR while also lowering cumulative absolute TPR disparity.The predictive-performance entries are descriptive and outside the directional testing families.

E Action incidence, no-drift controls, and cadence comparisons

Action records separate inaction from zero disparity contrasts, while no-drift controls and cadence comparisons clarify how policy actions relate to cumulative disparity outcomes.

  • Action incidence: Action records identify inaction directly because refitting can also produce a zero disparity contrast.Conditional means and adverse shares describe policy-selected trajectory sets, which differ across policies and therefore lack causal comparability.
  • Cadence comparisons: The cadence comparison includes paired changes in log loss, accuracy, balanced accuracy, and both groups’ TPR and FPR.
  • Action incidence: Conditional cross-policy differences have no causal interpretation because policies select different paths on which to act.
  • No-drift controls: No-drift controls retain each condition’s calibrated baseline and environmental draws while removing the drift process.Thresholds and initial/training-window sizes remain unchanged, and the controls are non-confirmatory.
  • Cadence comparisons: Evaluation-horizon prefixes reuse the same observations and policy actions, so cumulative and per-window contrasts answer different scaling questions.The retained prefixes include the first 5, 8, or 10 windows, with each average-gap contrast divided by its own horizon.

G.1 Alarm direction and recurrence under a fixed reference

Alarm analyses distinguish fixed-reference accumulation from changes relative to the preceding observation, then assess recurrence, refit effects, and random-count comparators under calibrated uncertainty.

  • Alarm direction: CUSUM evidence accumulates relative to a fixed reference, whereas a new observation can improve relative to the immediately preceding value.
  • Alarm direction: Refit-change diagnostics compare old and new models on the same next evaluation sample after replay and cannot influence trigger decisions.The diagnostic measures prediction change conditional on the realised refit, not an isolated causal trigger effect.
  • Alarm recurrence: Each refit is classified using streams newly alarming at that boundary; worsening means increased absolute disparity since the prior observation, not since the CUSUM reference.Multiple alarms can lead to one action, and mixed changes include opposing absolute-gap directions among active streams.
  • Random comparator: The random policy’s binomial count distribution is compared with empirical monitored-policy counts before interpreting paired disparity differences.Matching expected counts does not constrain the probability of any action or the realised count distribution.
  • Random comparator: Outcome bootstrap intervals condition on the calibrated probability and omit uncertainty from repeating calibration and evaluation together.

G.3 Threshold selection and matching validation

Threshold matching evaluates action-count similarity on separate selection and validation samples, but the selected approximation fails key validation criteria and its planned disparity comparison is withheld.

  • Threshold selection: Threshold candidates are ranked by action-count agreement with the fixed gap-trigger policy, then assessed on a separate validation sample.The validation sample evaluates all five matching criteria after selection.
  • Threshold selection: The selected approximation’s validation-stage metrics are disjoint from selection-stage candidate descriptions.Candidate numbering follows increasing threshold rather than selection rank.
  • Matching validation: The selected threshold passes mean-count, paired-interval, and common-support validation criteria but fails count-distribution and any-refit criteria.Because the selected approximation failed required criteria, the planned matched disparity comparison was not run.
  • Oracle comparison: Oracle schedules are reconstructed by exhaustive subset enumeration at every exact count, including zero, with policy-specific realised-count caps.
  • Reconstruction: Reconstruction reproduces archived aggregate disparity contrasts and action counts across all attribute–state combinations and both evaluation weightings.Historical execution lacked invocation metadata and individual predictions, limiting verification against original outputs.
  • Census weighting: Person weighting changes the evaluation of fixed Census predictions without changing fitted models, monitoring, refit actions, or state weighting.Weighted race-FPR leave-one-state-out means retain sign reversals, while all exchangeable-state BCa intervals include zero.

I.1 Monitoring calibration and action targets

Monitoring calibration keeps the loss-trigger policy below the three-refit target across tested settings, while audit and provenance records delimit what can be verified historically.

  • Monitoring calibration: Across every tested reference value, the loss-triggered policy remains below the three-refit calibration target.Tables report selected parameters and maximum attained counts across the tested thresholds.
  • Monitoring calibration: The Census monitoring target is three mean refits, with reference values and thresholds expressed in calibration-stream scale units.
  • Monitoring calibration: The maximum calibration-state mean count is reported for each reference value, including additional diagnostic settings that do not alter policy selection.
  • Sensitivity audit: The cap audit excludes Utah at 10,000 records per state-window, while the complete 25,000-row archive reproduces the main state-level contrasts.Historical audit metadata cannot attribute observed differences to cap size.
  • Reproducibility: The reproduction package binds supplied source and outputs with file hashes and retains calibrated inputs, trajectory outcomes, actions, and diagnostic records.
  • Reproducibility: Historical reconstruction establishes present comparisons but does not validate an old invocation because repeat-output directories and machine-readable manifests were not retained.
Loading 2609.09788v1…