Source-linked AI summary
Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning
Lars van der Laan, Nathan Kallus
TL;DR
Existing occupancy-ratio estimators can retain Bellman-balance violations without a direct supervised validation loss. The paper introduces isotonic Bellman calibration, a one-dimensional monotone FORE-based post-processing method that preserves ranking while correcting fitted ratios. It establishes conditional fixed-point and calibration–refinement results, finite-sample guarantees, and downstream guarantees for occupancy functionals including policy-value estimation.
Problem
Existing marginalized importance-weighting estimators can retain occupancy-balance violations, while their fixed-point objectives generally provide no direct supervised validation loss for tuning or model selection.
Method
Isotonic Bellman calibration applies FORE over nondecreasing transformations of an arbitrary initial occupancy-ratio estimate, preserving its ordering while adjusting scale and shape.
Results
The method achieves finite-sample calibration guarantees and a KL regret bound showing that, up to statistical error, calibration preserves and may improve occupancy-ratio estimation accuracy.
Takeaways & Limitations
When initial fitted values remain informative and behavior data adequately cover the target occupancy, calibration can improve Bellman balance and downstream off-policy value estimation.
Takeaways & Limitations
The method is most useful with adequate target-occupancy coverage and meaningful initial ranking; under limited coverage or small calibration samples, simpler corrections may offer a better bias–variance tradeoff.
Abstract
from arXiv · showhide
Marginalized importance weighting evaluates a target policy by reweighting offline state-action samples with its discounted occupancy ratio, characterized by an adjoint Bellman equation. Existing minimax, primal-dual, and fitted fixed-point estimators can leave residual occupancy-balance violations because of function-class approximation, regularization, or incomplete optimization. These violations are difficult to diagnose and reduce because the objectives generally lack a direct supervised validation loss for hyperparameter tuning, model selection, and early stopping. We introduce isotonic Bellman calibration, a one-dimensional, model-agnostic post-processing method that reduces these violations while preserving the ranking information in any initial occupancy-ratio estimate. The method corrects the estimate's scale and shape by applying fitted occupancy-ratio evaluation (FORE) over a one-dimensional class of nondecreasing transformations. We characterize Bellman calibration as a conditional fixed-point property equivalent to occupancy-balance against every test function of the calibrated ratio. More generally, we derive a calibration-refinement bound showing that any fitted ratio with small calibration error performs nearly as well as the best post-processing based on its fitted values. For isotonic Bellman calibration, we establish finite-sample calibration guarantees and a KL oracle inequality relative to the best monotone transformation of the initial estimate. Consequently, isotonic Bellman calibration achieves small calibration error and KL risk within statistical error of the best monotone correction, with guarantees for downstream target-occupancy functionals, including policy-value estimation.
1. Introduction
Marginalized importance weighting uses discounted occupancy ratios to evaluate target policies, but existing estimators can retain Bellman imbalance without a direct validation loss. The paper introduces isotonic Bellman calibration as model-agnostic post-processing that preserves ranking information while correcting fitted ratios.
- Motivation: Marginalized importance weighting corrects marginal state–action occupancies without multiplying importance ratios across time.
- Motivation: Existing minimax, primal–dual, and fitted fixed-point estimators can leave population Bellman imbalance because of approximation, regularization, optimization error, or early termination.These methods also lack observed labels or a canonical held-out loss for tuning and model selection.
- Core idea: Bellman calibration requires the conditional mean adjoint Bellman update to agree with the fitted occupancy ratio and preserves the ordering encoded by the initial estimate.This condition is equivalent to balancing adjoint Bellman moments against every function of the fitted ratio.
- Method: Isotonic Bellman calibration applies FORE over a one-dimensional class of nondecreasing transformations of any initial occupancy-ratio estimate.Each iteration solves a convex isotonic optimization problem whose optimality conditions imply an exact empirical Bellman-balance identity.
- Guarantees: The paper establishes a calibration–refinement bound, finite-sample calibration guarantees, and a KL regret bound showing calibration preserves and may improve ratio-estimation accuracy up to statistical error.
2. Discounted occupancy ratios and Bellman balance
The paper defines discounted occupancy ratios relative to the offline state–action distribution and characterizes them through an adjoint Bellman fixed point. Balance-based estimators can still miss residual imbalance when critics are insufficiently rich or fitting is imperfect.
- Discounted MDPs: A discounted MDP comprises state and action spaces, a transition kernel, and discount factor γ ∈ [0, 1).
- Occupancy ratios: The discounted occupancy ratio wπ is the Radon–Nikodym derivative dµπ/dν of the target discounted occupancy measure relative to the offline sampling distribution.
- Bellman balance: The target occupancy measure satisfies an adjoint Bellman equation, yielding occupancy-balance moments evaluated with one-step transitions and the target initial distribution.
- Existing estimators: Minimax occupancy balancing estimates ratios by minimizing empirical violations over a critic class, whose richness determines which residual directions are detectable.
- Residual imbalance: Insufficient critic richness, finite-sample error, regularization, early stopping, and imperfect optimization can leave residual imbalance despite small empirical objectives.
3. Bellman Calibration for Marginalized Importance Weights
Bellman calibration makes fitted occupancy ratios self-consistent after conditioning on their own fitted values. A calibration–refinement bound links calibration error to the best transformation-based recovery of the true ratio and controls downstream occupancy functionals.
- Calibration: An initial ratio estimate can preserve predictive information while exhibiting systematic Bellman imbalance from scale or shape distortion.
- Calibration: Perfect calibration requires adjoint Bellman moment identities against every test function of the calibrated ratio.Equivalently, the conditional average adjoint Bellman update equals the fitted ratio within fitted-value strata.
- Conditional fixed point: Perfect calibration is equivalent to the calibrated weight satisfying an adjoint Bellman fixed-point equation after conditioning on its own fitted values.
- Accuracy: The calibration–refinement bound decomposes ratio error into Bellman calibration error and the smallest error achievable by transforming the fitted values.Thus, calibration supports accuracy when the fitted values can be transformed accurately into the true occupancy ratio.
- Downstream functionals: The same bound controls every bounded linear functional of the target discounted occupancy distribution, including rewards, costs, feature moments, visitation probabilities, and policy values.
4. Fitted isotonic Bellman calibration
Fitted isotonic Bellman calibration post-processes an initial occupancy-ratio estimate using nondecreasing transformations. It preserves ranking while adapting scale and shape through one-dimensional isotonic risk minimization.
- Method: The procedure improves Bellman calibration while preserving the ordering induced by an initial occupancy-ratio estimate.
- Method: FORE is applied over nondecreasing transformations, reducing each update to a one-dimensional isotonic risk-minimization problem.
- Transformation class: The transformation class includes the identity, rescaling, clipping, and flexible nonlinear corrections.
- Implementation: Formal guarantees use an external calibration sample independent of training data, while the empirical implementation uses cross-fitted out-of-fold predictions.
- Optimization: Each isotonic update is convex and one-dimensional, and can be solved after sorting in linear time using a generalized pooled adjacent violators algorithm.
5. Finite-sample guarantees for isotonic Bellman calibration
Isotonic Bellman calibration enforces empirical occupancy balance through one-dimensional isotonic updates, yielding finite-sample calibration guarantees. Under stated assumptions, its calibration and KL accuracy approach the best admissible monotone correction up to statistical, approximation, and iteration errors.
- Bellman calibration error: Each isotonic update satisfies an exact empirical occupancy-balance condition, which approximates perfect calibration when successive iterates are close.The resulting iteration difference is directly observable and can provide a stopping criterion.
- Bellman calibration error: The finite-sample calibration theorem converts exact empirical balance into population calibration under Conditions A1–A3.The guarantee is conditional on the training data and holds with high probability over the calibration sample.
- Bellman calibration error: Once iteration changes fall below statistical error, L2 calibration error is order (n∧m)^−1/3 up to logarithmic and high-probability factors.The squared calibration error is order (n∧m)^−2/3, matching the classical isotonic calibration rate.
- KL regret and accuracy preservation: The KL oracle inequality compares calibrated ratios with the best normalized monotone transformation while accounting for iteration, approximation, and calibration-sample statistical errors.Logarithmically many iterations make the geometric iteration error negligible, and the resulting guarantee is within statistical error when other errors are controlled.
- KL regret and accuracy preservation: KL analysis uses a strictly positive constrained variant because logarithmic loss is singular at zero, although the unconstrained procedure is recommended in practice.The constrained transformation lies in [εb, TN], while inactive constraints make constrained and unconstrained updates coincide.
6. Experimental evaluation
The evaluation tests cross-fitted isotonic calibration of four occupancy-ratio estimators on D4RL and InfiniteCartPole. Calibration generally reduces projected calibration and policy-value errors, but benefits vary by ESS stratum and base estimator.
- Evaluation design: The protocol evaluates calibrated neural FORE, DualDICE, MWL, and NeuralDICE using ten-fold grouped cross-fitting and independent evaluation data.The study includes 960 paired D4RL comparisons and 160 paired CartPole comparisons.
- Calibration and policy-value accuracy: Calibration reduces both projected calibration error and mean policy-value error across the two benchmarks.Mean policy-value error falls from 2.99 to 0.710 in D4RL and from 0.176 to 0.110 in CartPole.
- Task and ESS heterogeneity: Calibration lowers value error in eleven of twelve D4RL dataset–policy pairs, with improvements varying substantially across ESS strata.Mean policy-value error decreases by 80.8% in high ESS and 74.7% in moderate ESS, but increases by 16.8% in low ESS.
- Base estimators: Calibration reduces projected calibration error for every base estimator, while policy-value error improves for DualDICE, MWL, and NeuralDICE but worsens for Neural FORE.Neural FORE value error increases from 0.465 to 1.03 in D4RL and from 0.123 to 0.133 in CartPole.
7. Conclusion
Isotonic Bellman calibration post-processes occupancy-ratio estimates with a monotone transformation that preserves ranking while correcting residual Bellman imbalance. The conclusion also identifies coverage and ranking quality as practical conditions, and notes extensions and implementation details for the isotonic fit.
- Conclusion: Isotonic Bellman calibration adjusts fitted weights through a data-driven monotone transformation while preserving their ordering.The approach is low-dimensional post-processing for residual occupancy-balance violations.
- Conclusion: When initial fitted values remain informative about the true ratio, calibration can improve downstream off-policy value estimation.
- Conclusion: The method is most useful when behavior data adequately cover the target occupancy and the initial estimate preserves meaningful ranking information.With limited coverage or small calibration samples, log-linear corrections may offer a better bias–variance tradeoff.
- Conclusion: Extensions include coverage-stopped FORE, richer multicalibration conditions, and finite-horizon or nonstationary settings with time-indexed ratios.The finite-horizon and nonstationary extension requires calibration conditions that may vary across decision times.
- Generalized PAVA: The finite-dimensional fitting step represents calibration as a nondecreasing step function over ordered fitted-value design points.Generalized PAVA pools adjacent blocks that violate monotonicity and updates each pooled block using C_B/A_B.
A.2. Reference Python implementation
The reference implementation solves isotonic FORE on scalar fitted values using generalized PAVA, normalization, and iterative stopping. Experimental calibration uses grouped cross-fitting, pooled maps, median aggregation, and independent evaluation data.
- PAVA implementation: The implementation solves a convex isotonic problem, requiring positive a and nonnegative b inputs before applying generalized PAVA.The objective is min_{u increasing} sum_j a_j u_j - b_j log u_j.
- Normalization: The fitted weights are normalized by dividing by the weighted sum of the isotonic outputs.
- Iteration: Isotonic FORE iterates until the root-mean-square weight change falls below the stopping tolerance or the iteration limit is reached.
- Calibration and aggregation: Grouped cross-fitting fits base estimators on nine folds, pools out-of-fold scores into one map, and aggregates calibrated predictions by the median.The comparison uses the same base estimator fitted on the full training sample without calibration.
- Evaluation: Projected Bellman calibration error is evaluated on independent data partitioned into subsamples for bin construction and Bellman-moment estimation.D4RL additionally uses disjoint training and evaluation episodes.
B.2. Base estimators and benchmarks
The benchmark compares fixed-hyperparameter occupancy-ratio estimators on D4RL and InfiniteCartPole using policy-value error and projected Bellman calibration error. Evaluation uses paired comparisons, independent data, and reference values from target-policy rollouts.
- Estimator settings: The study fixes one hyperparameter set per base estimator across tasks, replications, folds, and full-sample fits.DualDICE, MWL, and NeuralDICE use the cited public implementations.
- D4RL: D4RL includes twelve matched dataset–policy pairs across HalfCheetah, Hopper, and Walker2d, with sample sizes 10,000 and 50,000.The design uses γ = 0.99, ten replications, and 960 comparisons across four base estimators.
- InfiniteCartPole: InfiniteCartPole uses four full-support behavior settings, 400 training trajectories, 50,000 transitions, γ = 0.995, and ten replications.The benchmark produces 160 comparisons across four base estimators and evaluates calibration on 400 additional trajectories.
- Metrics: Policy-value error is the absolute difference between the estimated value and rollout-based reference value, using the training behavior sample.
- Metrics: Projected Bellman calibration error estimates the squared L2(ν) norm of residual projections onto fitted-ratio quantile-bin indicators.Three disjoint evaluation subsamples support bin construction, Gram-matrix estimation, and moment estimation.
- Comparison design: Primary comparisons pair calibrated and full-sample estimators within each experimental setting, replication, sample size, discount, and base estimator.Percentile 95% confidence intervals use 10,000 cluster-bootstrap resamples.
B.4. Additional results
Additional results report calibration and policy-value errors by base estimator and matched D4RL dataset–policy pair. Calibration reduces projected calibration error for every base estimator and lowers mean policy-value error in eleven of twelve D4RL pairs.
- Base estimators: Table 3 disaggregates projected calibration and policy-value error by base estimator, comparing uncalibrated and ten-fold cross-calibrated outputs.Lower error is preferred.
- Task-level heterogeneity: Calibration lowers mean policy-value error in eleven of the twelve matched D4RL dataset–policy pairs.Table 4 averages over sample sizes, replications, and base estimators.
C. Fixed-image KL calibration–refinement identity
This section decomposes adjoint Bellman KL discrepancy into calibration and refinement components. Calibration measures the removable error from post-processing fitted values, while refinement measures the best remaining discrepancy attainable through such transformations.
- The adjoint Bellman KL discrepancy measures departure from the occupancy-ratio fixed-point equation and controls L1(ν) ratio error.
- The exact fixed-image identity decomposes KL discrepancy into refinement error plus calibration error.
- With the adjoint Bellman image fixed, refinement error is the smallest KL discrepancy achievable by transforming only fitted values.
- Calibration error is zero exactly when the fitted ratio is perfectly calibrated.
- The calibration–refinement decomposition identifies calibration error as the maximal KL improvement available through fitted-value post-processing.
- Squared calibration error controls the KL calibration component when the fitted ratio and calibrated image are bounded below.The bound is CalKL(ω) ≤ (1/(2m)) Cal_2^2(ω).
D.5. Proof of the isotonic FORE KL-regret theorem
The proof establishes the isotonic FORE KL-regret theorem by analyzing a bounded monotone transformation class of the initial fitted ratio. Its complexity and approximation properties support an oracle bound with geometric, approximation, and statistical terms.
- The isotonic class has localized complexity yielding a critical radius of order N−1/3 up to class-dependent constants.
- The bounded isotonic calibration class consists of nondecreasing transformations of the initial fitted ratio constrained between log εb and log TN.
- The empirical class is parameterized by ordered values at the distinct fitted-value design points, forming a compact convex set.
- The restricted recursion has approximation error exactly εiso,N.
D.6. Proof of finite-sample Bellman calibration
This proof combines exact empirical Bellman balance with localized empirical-process control for bounded-variation calibration classes. The resulting argument transfers empirical balance into finite-sample calibration guarantees under smoothing, tail, and boundedness conditions.
- The Bellman-balance functional equals a conditional calibration residual evaluated over functions of the fitted ratio.
- Exact empirical Bellman balance holds at every iteration for every bounded measurable transformation of the fitted values.
- Bounded isotonic iterates and calibration residuals are placed in bounded-variation composition classes controlled by the initial fitted ratio.
- Entropy bounds and local maximal inequalities provide simultaneous empirical-process control at the N−1/3 scale, where N = n ∧m.
- The proof uses truncation to compare the calibrated residual with a bounded-variation surrogate and then controls the resulting empirical and population discrepancies.
- For sufficiently large N, logarithmic truncation keeps the relevant population masses bounded, placing the iterates in the required class.