Source-linked AI summary
Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry
Kihun Rhee
TL;DR
Recent clinical offline RL studies report improvements over physician decisions, but the validity of those estimates is uncertain. This paper systematically evaluates five RL families and 14 reward designs in a large stroke registry, finding that apparent improvement attenuates after reward deconfounding and does not support a clinically meaningful aggregate benefit.
Problem
The paper addresses whether reported offline RL improvements in clinical treatment decisions remain credible when observational proxy rewards encode baseline severity and prognosis.
Method
The study evaluates five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients, using FQE, factorial reward deconfounding, and non-RL triangulation.
Results
After GBM reward residualization, the FQE estimate attenuated to +0.0033 (p = 0.132), while full deconfounding yielded +0.0025 (p = 0.291); complementary analyses converged away from clinically meaningful aggregate improvement.
Takeaways & Limitations
Reward-channel diagnostics and the six-step evaluation checklist are presented as safeguards against premature positive policy-improvement claims in clinical offline RL.
Takeaways & Limitations
Remaining heterogeneity is hypothesis-generating for prospective trial design rather than evidence of treatment benefit, and hospital-level disagreement does not persist after full deconfounding.
Abstract
from arXiv · showhide
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). Standard Fitted Q-Evaluation (FQE) yields an apparent policy-improvement estimate of +0.0069; adding an Early Neurological Deterioration penalty increases it to +0.0101. We identify reward-embedded confounding, in which a proxy terminal reward encodes baseline severity and prognosis as well as treatment efficacy. A 2 x 2 factorial analysis finds that terminal reward confounding accounts for 218.6% of the observed signal change, so its removal overshoots the null. After DML-inspired GBM reward residualization, the FQE estimate attenuates to +0.0033 (p = 0.132), and full deconfounding yields +0.0025 (p = 0.291). FQE-based diagnostics, T-learner analyses, and direct recurrence analyses converge away from a clinically meaningful aggregate improvement. A 1-year mRS factorial analysis replicates the attenuation. We provide an empirically motivated six-step evaluation checklist. NIHSS-stratified heterogeneity is hypothesis-generating for prospective trial design; hospital-level disagreement does not persist after full reward deconfounding.
1. Introduction
The paper evaluates offline RL for post-stroke antithrombotic treatment as a diagnostic exercise and finds that apparent policy improvement is vulnerable to reward-embedded confounding. It argues that standard external-confounding checks do not address bias carried through the reward channel.
- +0.0069 (p = 0.048) was the initial standard FQE policy-improvement estimate in 44,894 post-2018 acute ischemic stroke patients.
- +0.0101 (p = 0.0002, E-value = 18.37) followed addition of an Early Neurological Deterioration penalty, potentially supporting a premature positive claim.
- Reward-embedded confounding occurs because proxy terminal rewards encode baseline severity and prognosis alongside treatment-efficacy information.
- The E-value is necessary but insufficient because it targets external confounding, whereas this bias operates through the reward channel.
- Offline RL evaluation is used to expose pitfalls that can produce premature positive policy-improvement claims, not to deploy a clinical policy.
- The paper proposes a six-step checklist combining FQE confirmation, support, confounding checks, non-RL triangulation, and subgroup stability.
2. Methods
The methods construct a two-step offline RL evaluation of six antithrombotic strategies using registry data, multiple algorithms, reward variants, and layered diagnostics. Central analyses compare raw and prognosis-residualized rewards with FQE and complementary methods.
- Data and cohort: The primary RL cohort contains 44,894 post-2018 patients from a prospective registry spanning 20 tertiary hospitals across South Korea.
- MDP formulation: The two-step MDP represents acute and discharge phases with 421 clinical features, six antithrombotic strategies, and discount factor γ = 0.99.
- Algorithms: Five offline RL families were trained: CQL, BCQ, IQL, Decision Transformer, and Behavior Cloning.
- Reward design: The study evaluates 14 reward variants, with central variants including standard, END-augmented, DML-residualized, fully deconfounded, and 1-year-outcome rewards.
- Reward deconfounding: GBM residualization subtracts baseline prognosis from the terminal reward to test whether apparent policy advantage survives removal of the severity–outcome pathway.
- Evaluation: FQE estimates V (imp) = V (πRL) − V (πphysician) using 5-fold cross-validation across 3 random seeds, alongside E-value analysis and four triangulation strategies.
3. Results
Standard offline RL evaluation produced positive policy-improvement signals, but reward deconfounding progressively attenuated them to a result consistent with no aggregate effect. Independent causal and structural diagnostics supported this interpretation, while NIHSS-stratified heterogeneity remained exploratory.
- 3.1. Act 1: Standard Offline RL Evaluation Yields Positive Signal: +0.0101 policy improvement under END-augmented FQE exceeded the +0.0069 estimate under the standard reward.The END estimate had p = 0.0002 and E-value = 18.37; rstd had p = 0.048 and E-value = 4.70.
- 3.3. MDP Structure Diagnosis: Direct Q-evaluation estimates were 5–131× larger than FQE across approximately 450 runs, while IS-based OPE was infeasible with ESS = 0.21%.FQE was retained as the least problematic primary evaluator, not an assumption-free method.
- 3.2. Act 2: Progressive Deconfounding Attenuates the Signal to Null: +0.0033 FQE policy improvement after GBM reward residualization was a 52% reduction from rstd and was not statistically significant.The 95% CI was [−0.001, +0.008] with p = 0.132; full deconfounding further reduced the estimate to +0.0025, p = 0.291.
- 3.2. Act 2: Progressive Deconfounding Attenuates the Signal to Null: 218.6% of the net rEND →rfull signal change was attributed to terminal reward confounding in the factorial decomposition.Terminal residualization alone overshot the null, while residualized intermediate reward partially offset it; the terminal deconfounding main effect was −0.0101.
- 3.2. Act 2: Progressive Deconfounding Attenuates the Signal to Null: The 1-year mRS analysis replicated attenuation, reducing r1yr’s +0.0133 estimate to −0.0004 after GBM residualization.The deconfounding main effect was −0.0069, p = 0.005, and the residualized estimate had p = 0.843.
- 3.2. Act 2: Progressive Deconfounding Attenuates the Signal to Null: T-learner and direct recurrence analyses converged away from a clinically meaningful aggregate improvement.The T-learner showed 75% attenuation in the full cohort and 90% in the overlap zone, while IPW/AIPW found no significant recurrence reduction.
- 3.4. Clinical Heterogeneity: Hypothesis Generation for Prospective Evaluation: NIHSS-stratified heterogeneity remained hypothesis-generating, whereas hospital-level disagreement disappeared after full reward deconfounding.The T-learner found a 4.6× severe/mild CATE ratio, directionally consistent with a 3.4× direct-Q gradient; hospital correlation changed from ρ = −0.495 to ρ = −0.026.
4. A 6-Step Evaluation Checklist for Clinical Offline RL
The paper proposes a six-step empirical checklist to determine when positive offline RL policy-value estimates require additional evidence before being interpreted as clinical benefit. Retrospective application found major reporting gaps, and the checklist itself remains externally unvalidated.
- 4. A 6-Step Evaluation Checklist for Clinical Offline RL: The six-step checklist operationalizes seven evaluation pitfalls with prespecifiable diagnostic checks for positive policy-improvement estimates.It is intended to identify when additional evidence is needed before interpreting V (imp) as clinical benefit.
- 4. A 6-Step Evaluation Checklist for Clinical Offline RL: External validation of the checklist remains future work because it was derived from diagnostic experience rather than preregistered or externally validated.This limits the current evidentiary status of the checklist as an evaluation instrument.
- 4. A 6-Step Evaluation Checklist for Clinical Offline RL: Retrospective application to three widely cited clinical RL papers found that Step 1 could not be assessed as satisfied because none reported FQE or equivalent model-based OPE.Steps 4–6 were also not reported, leaving positive findings unexamined for reward-channel confounding.
5. Discussion
The analysis identifies conditions limiting offline RL’s added value in this registry and proposes diagnostics for interpreting policy-improvement estimates. Deconfounded effects are statistically null and clinically small, while generalizability and sequential-data limitations constrain the conclusions.
- When Can Offline RL Add Value in Clinical Settings?: Reward-channel diagnostics, meaningful sequential dynamics, and balanced action support are identified as conditions under which offline RL may add value beyond simpler causal inference.This study had 98.3% state invariance and IS-ESS of 0.21%, limiting sequential optimization and importance-sampling evaluation.
- When Can Offline RL Add Value in Clinical Settings?: The six-step checklist organizes checks for OPE reliability, confounding robustness, causal triangulation, and generalizability.External validation of the checklist remains future work.
- Limitations: The near-null result is specific to antithrombotic selection in Korean tertiary stroke registries, where more than 80% of AF patients receive anticoagulation.Different populations, hospitals, treatment dimensions, or richer data modalities may have different treatment heterogeneity and scope for improvement.
- Power Analysis and Effect Size Bounds: V (imp)adj ≈ +0.0045, corresponding to ≤2.7% of MCID, after accounting for estimated over-adjustment from GBM residualization.Semi-synthetic calibration estimated 24–26% absorption, while a conservative propensity-based bound estimated 27%.
- Limitations: External validity remains unvalidated across other ethnic populations, community hospitals, and healthcare systems with different practice patterns.The two-step MDP also has 98.3% state invariance, and FQE realizability violations cannot be ruled out.
- When Can Offline RL Add Value in Clinical Settings?: NIHSS-stratified heterogeneity is presented as a trial-design signal rather than a treatment recommendation, and hospital-level disagreement disappears after full reward deconfounding.The T-learner’s 4.6× severe/mild CATE ratio motivates stratified prospective evaluation.
6. Conclusion
Standard offline RL evaluation produced a positive apparent improvement, but reward deconfounding substantially attenuated the estimate. Across complementary analyses, the evidence does not support a clinically meaningful aggregate improvement.
- Conclusion: +0.0101 (p = 0.0002, E-value = 18.37) under standard evaluation was reduced to +0.0025 (p = 0.291) after full deconfounding.Terminal residualization accounted for 218.6% of the net rEND-to-rfull change.
- Conclusion: Reward-embedded confounding is a reward-channel pathway that external-confounding E-values do not evaluate.The pathway arises when observational outcomes used as terminal rewards carry baseline severity and prognosis information.
- Conclusion: No algorithm family produced positive evaluated evidence in the common-reward screen, while T-learner and recurrence AIPW analyses provided triangulation.In primary CQL, residualization attenuated the estimate to null.
- Conclusion: Positive offline RL results based on observational terminal rewards should remain hypothesis-generating until reward-channel diagnostics are applied.This conclusion concerns interpretation of observational-reward evaluations rather than deployment of a clinical policy.
Data and Code Availability
The paper describes a restricted, partially crossed evaluation dataset and reports substantial instability in importance-sampling OPE and selected causal analyses. The supplied availability material indicates that individual-level registry data cannot be publicly deposited.
- Data Availability: Individual-level CRCS-K data cannot be publicly deposited because of consent, ethics, privacy, regulatory, and data-sharing restrictions.Qualified researchers may request a de-identified minimum dataset subject to committee and institutional approvals.
- Cohort Construction: The primary RL cohort contains 44,894 post-2018 episodes derived from 129,033 registry admissions across 20 hospitals.The cohort was restricted temporally after the 2018 Korean Stroke Society guideline update.
- OPE Diagnostics: ESS = 93.6 patients, or 0.21% of 44,894, under the undeconfounded rstd reward, versus a recommended IS threshold of at least 10%.The maximum cumulative importance weight reached 9.8 × 10^14.
- Evidence Map: Positive signals clustered in non-deconfounded reward designs and attenuated after GBM residualization across the evidence map.Figure 4 encodes V (imp) radially relative to the physician-baseline null and uses color for result category.
- Experimental Design: The experiment grid was partially crossed: five algorithm families were screened under a common reward, followed by reward deconfounding for prespecified primary CQL.Untested algorithm–reward combinations were not imputed.
- Causal Analyses: The DAPT experiment had TWFE ATT = +0.134 (p = 0.170), but parallel trends were violated, making the results exploratory.The NOAC estimate was unstable across inference methods and modest confounding could explain both estimates.
Appendix H. Direct Recurrence Analysis
Direct recurrence analyses found no significant reduction in ischemic stroke recurrence across treatment comparisons, while highlighting confounding and positivity constraints that limit causal interpretation.
- No comparison achieves significant ischemic stroke recurrence reduction.The analysis used IPW/AIPW treatment-effect estimation and cause-specific Cox regression across three binary comparisons and four endpoints.
- HR = 0.938 (p = 0.001) for mortality under DAPT, the most robust observed signal.This mortality finding is consistent with prior CHANCE/POINT trial evidence.
- ATE = +0.010 (p = 0.019) for any-cause recurrence in the AF anticoagulation contrast should not be interpreted as evidence that anticoagulation increases recurrence risk.Higher stroke severity among anticoagulated patients and a small control group limit reliable counterfactual estimation.
- Positivity violation limits causal interpretation for C1 statin, with 90.2% treated and an ESS ratio of 3.2% at trim = 0.10.
- The 3.4× NIHSS severity gradient in mRS-based RL is absent for recurrence across all severity strata.This indicates that the heterogeneity pattern is a property of the mRS composite outcome.
Appendix I. MDP Design Sensitivity: 2-Step Anchor Experiment (E8F)
The E8F anchor experiment tested whether correcting the MDP from three steps to two materially changed the study’s null finding. It did not, while improving FQE TD error.
- V (imp) = +0.0020 versus +0.0024 (t = 0.08, p = 0.94) across the original and E8F configurations.The comparison used the same full registry, CQL α=4.0, R4 reward, and 15-fold cross-validation, differing only in MDP formulation.
- The E8F-versus-E8B difference of +0.005 is attributed to the post-2018 temporal restriction, not the MDP formulation.This is consistent with the guideline-shift analysis in Appendix B.
- FQE TD error decreases 3.7× from the three-step to the two-step formulation on the same comparison.
Appendix J. E-Value Sensitivity Analysis
E-value sensitivity analysis supports a high apparent robustness signal, but the paper argues that E-values can mislead when the terminal reward itself carries prognostic severity information.
- E-values were reported under three standard-error assumptions because the Z-to-risk-ratio mapping is non-standard.The mapping has no closed-form conversion for continuous policy-value differences and likely overestimates risk ratios because mRS is ordinal.
- Reward-embedded confounding is distinct from external unmeasured confounding because baseline severity and prognosis can be encoded directly in the observational terminal reward.Therefore, E-value sensitivity analysis does not target the reward-definition problem addressed here.
- rEND’s E-value remains high under all three assumptions.The authors interpret this as evidence that high E-values are misleading when reward-embedded confounding is present.
- The study’s implementation uses CQL α=4.0 as the primary configuration, with other hyperparameters treated as ablations.The GBM hyperparameters apply to both the rDML outcome model and the T-learner.
- The GBM outcome model uses 743 treatment-free covariates, whereas the MDP state space uses 421 state features after additional filtering.Both pipelines exclude treatment variables to prevent leakage; the GBM prioritizes outcome prediction and the MDP emphasizes temporal variation and interpretability.
Appendix M. Sensitivity to Treatment Effect Absorption
Sensitivity analyses examined whether reward residualization could absorb genuine treatment-benefit variation and whether subgroup patterns were stable. The results support substantial attenuation, while demographic heterogeneity remains uncertain under limited treatment overlap.
- The propensity-based bound predicts an attenuation ratio of 0.73, implying at most 27% of treatment effect is absorbed.Only 22% of patients fall in the strict overlap region, while 78% have e(X) > 0.9.
- The semi-synthetic calibration produces attenuation ratios of 0.74–0.77 across injected treatment effects.The τ=0 baseline bias is −0.026 and biases residualized estimates in the negative direction.
- The observed rDML estimate is +0.0033, 1.6× lower than the conservative bound implied by rstd’s +0.0069.The authors infer that confounding removal accounts for at least 38% of the reduction; the adjusted estimate remains ≤2.7% MCID with p > 0.10, and rfull is +0.0025 (p = 0.291).
- Statin prescription rates are 87–92% across demographic cells, leaving only 22% within the strict overlap region.This positivity limitation applies across demographic subgroups.
- ATE ≈−0.068 in the 65–74 group contrasts with ATE ≈+0.010–+0.023 in younger and older groups.The non-monotone pattern likely reflects T-learner extrapolation because the untreated arm provides insufficient support for stable counterfactual prediction.
- The significant sex and age interaction terms may reflect large-sample sensitivity rather than clinically meaningful heterogeneity.The interaction test rejects homogeneity, but the subgroup estimates are raw, not rDML-residualized, and remain confounded.
- The analysis finds no evidence of demographic disadvantage because directional CATE uncertainty is uniform across sex-by-age cells.The same high and near-uniform statin uptake mechanism limits causal inference in both subgroup and full-cohort analyses.
- The 1-year mRS cohort shrinks from N = 44,895 to N = 35,744, primarily because 2024–2025 lacked sufficient follow-up time.The retained 2018–2023 cohort has 97.2% 1-year mRS coverage.
Appendix P. Extended Limitations
The appendix qualifies the study’s conclusions through assumptions about FQE, incomplete safety-rule data, and the limited temporal scope of the endpoint.
- Limitations: FQE assumes Q-function realizability, supported indirectly by low TD error, convergence across additional methods, and a full-dataset 2-step validation.The anchor experiment estimates V (imp) = +0.0024 (p = 0.235, 95 % CI [−0.001, +0.006]).
- Limitations: +0.0024 V (imp) in the full 116,215-episode 2-step dataset was statistically identical to earlier estimates, indicating temporal restriction did not create the null.The 2-step formulation narrowed the confidence interval from ±0.020 to ±0.006.
- Safety constraints: Action masking enforced guideline constraints at inference time, but two additional safety rules were inoperative because required columns were absent from CRCS-K.The inoperative rules concerned statin/liver disease and drug allergy.
- Safety constraints: The masking framework required proxies for some contraindications, including an intracranial-hemorrhage flag that does not distinguish hemorrhage subtypes.Severe renal failure was proxied by creatinine above 3.0 mg/dL when eGFR was unavailable.
Appendix R. T-Learner Convergence Analysis
The convergence analysis compares T-learner and sequential RL diagnostics, while also exposing limited treatment overlap and weak sequential dynamics. Across these checks, deconfounding removes most apparent signal, leaving only a small, extrapolation-sensitive T-learner residual.
- T-Learner Convergence Analysis: 75–90 % of the T-learner ATE attenuates after deconfounding, but a small residual remains and should be interpreted as an upper bound.The T-learner residual is clinically small, below 6 % of the MCID, and depends on extrapolation for most patients.
- T-Learner Convergence Analysis: Only 22 % of T-learner patients fall in the overlap zone, while 78 % require extrapolation beyond the observed treatment distribution.The statin prescription rate is 90.2 % and propensity AUC is 0.84 in the T-learner cohort.
- MDP Structure Diagnosis: 98.3 % of state features have Pearson r > 0.99 across timesteps, and 73.4 % of patients receive identical CQL-preferred actions at both timesteps.The near-zero transition structure makes sequential optimization resemble discounted 1-step optimization.
- MDP Structure Diagnosis: 3.7× reduction in FQE TD error occurs when moving from the 3-step to the 2-step MDP, from 8.0 × 10−3 to 2.2 × 10−3.Both errors are measured on the same 116,215-episode dataset.
- OPE Diagnostics: 0.21 % IS ESS in Phase 8B is 48× below the 10 % viability threshold, with cumulative weights reaching 9.8 × 10^14.Action-distribution mismatch, especially around anticoagulant recommendations, makes IS-based OPE effectively inapplicable.
- Policy Disagreement Profiling: 57 % of RL–physician disagreements involve statin co-prescription, while agreement rises from 62.1 % at t=0 to 71.3 % at t=1.AF history is the strongest patient-level driver, with Cohen’s d = 0.443.
- Confounding Analysis: The causal structure separates baseline severity’s effects on treatment and outcome from treatment’s effect on outcome, which external E-values do not evaluate.The standard reward encodes the full severity-to-outcome pathway, motivating reward deconfounding.
- Statistical Details: Stricter variance and multiple-testing checks leave rEND significant while rstd, rDML, and rfull remain non-significant.The Nadeau–Bengio correction increases standard error by approximately 12 %.