Source-linked AI summary

Deciding When to Decide: Testing Operational Suboptimality Under Distributional Shift

Minxing Zheng, Holly Wiberg, Shixiang Zhu

arXiv:2608.29465v1stat.MLcs.LGstat.ME

TL;DR

The paper asks when operating-condition changes make a retained decision materially suboptimal, since distribution-shift tests can detect changes irrelevant to the decision. RADAR infers latent preferences through inverse optimization and tests the deployed decision’s target-distribution optimality gap, distinguishing harmful from harmless shifts across experiments. Its validity depends on the forward objective model and sufficiently informative incumbent decision; preference ambiguity can make adequacy indeterminate.

  • Problem

    Re-optimization is costly, while standard distribution-shift tests do not determine whether a deployed decision has become materially suboptimal.

  • Method

    RADAR combines inverse optimization with regret-based inference to test the deployed decision’s optimality gap under the target distribution.

  • Results

    RADAR distinguishes decision-relevant shifts from decision-irrelevant changes across synthetic and real-world problems, with sequential monitoring detecting harmful suboptimality while baseline change-point tests alarm continuously.

  • Takeaways & Limitations

    Targeting the incumbent decision’s optimality gap provides a basis for deciding whether distributional change warrants re-optimization.

  • Takeaways & Limitations

    RADAR’s guarantees are conditional on the specified forward model, and weakly rationalized decisions can leave preference ambiguity that makes the audit indeterminate.

Abstract

from arXiv · show

Deployed decisions are often optimized once and retained because updates impose operational, regulatory, or switching costs. As operating conditions change, when should such decisions be re-optimized? We study this question for stochastic optimization when the objective's functional form is known but the decision maker's trade-offs are encoded by an unknown preference parameter. Standard distribution-shift tests are poorly aligned with this goal: they can flag detectable yet decision-irrelevant changes without determining whether the incumbent decision has become materially suboptimal. We propose \texttt{RADAR} (Regret-based Assessment of Decision Adequacy and Risk), a decision-focused framework that uses inverse optimization to infer latent preferences and tests the deployed decision's optimality gap under the current distribution. By targeting regret, \texttt{RADAR} ignores decision-irrelevant shifts while detecting changes that warrant re-optimization. We develop two-sample and sequential changepoint procedures and establish asymptotic guarantees for Type-I error and power. Across synthetic optimization problems, a semi-synthetic capacity allocation task, and police-zone planning, \texttt{RADAR} more reliably distinguishes harmful from harmless shifts than decision-agnostic alternatives.

1 Introduction

The paper asks when distributional change makes an implemented decision worth re-optimizing, arguing that decision adequacy depends on the deployed decision’s optimality gap rather than distributional change alone. RADAR addresses this by inferring latent preferences and testing regret in two-sample and sequential settings.

  • Re-optimizing deployed decisions can be costly, risky, and slow because revised plans may require validation, regulatory review, coordination, or workflow changes.
  • Detectable distribution shifts may be decision-irrelevant, while small shifts can create large suboptimality, so shift magnitude alone is misaligned with re-optimization.
  • Assessing adequacy requires inferring the latent preference structure that rationalizes the deployed decision and evaluating its suitability in the changed environment.
  • RADAR combines inverse optimization with regret-based inference to test whether a deployed decision remains approximately optimal under distribution shift.
  • The framework targets the incumbent decision’s optimality gap, distinguishing decision-relevant shifts from changes that do not warrant re-optimization.

2 Problem Setup

The setup compares a fixed decision optimized under a baseline distribution with its adequacy under a potentially different target distribution. The objective form is known, but the preference parameter governing trade-offs is unknown.

  • The baseline domain contains historical contexts and a deployed decision assumed optimal for a stochastic optimization problem.
  • The target domain contains new contexts from a distribution that may differ from the baseline, while the deployed decision remains fixed and feasible.
  • The stochastic objective has known functional form but an unknown preference parameter vector encoding latent trade-offs such as cost, reliability, fairness, and risk.
  • The target-domain risk-optimality gap measures the deployed decision’s excess expected objective value relative to the optimal decision under the target distribution and true preferences.
  • The hypothesis test asks whether the target-domain gap exceeds tolerance τ, with rejection indicating excess risk beyond tolerance and suggesting re-optimization.

3 Proposed Method

RADAR assesses whether a deployed decision remains approximately optimal after distributional shift by inferring latent preferences, constructing a target-domain benchmark, and testing the resulting optimality gap. It supports static and sequential monitoring while accounting for inverse-optimization ambiguity and estimation error.

  • Decision-adequacy objective: RADAR tests whether the deployed decision remains approximately optimal under the target distribution rather than testing for generic context-distribution change.This targets decision-relevant shifts while avoiding shifts that do not induce a nontrivial deployment optimality gap.
  • Step 3: Testing Procedure: Sample-split evaluation estimates the deployment gap using held-out target-domain loss differences and supports a one-sided Wald test.For nonsmooth risk functionals or small samples, the framework allows a bootstrap test extension.
  • Step 1: Preference Parameter Estimation: Inverse optimization estimates the latent preference parameter that rationalizes the observed baseline decision.The estimator need only be risk-consistent; the inferential guarantees do not depend on a particular inverse-optimization algorithm beyond that property.
  • Step 2: Constructing a Deployment Benchmark: RADAR constructs a data-driven target-domain optimizer because the target distribution and latent preferences are unknown.The target samples are randomly split so one subsample constructs the benchmark and the other independently evaluates the deployment gap.
  • Sequential Decision-Adequacy Monitoring Extension: The sequential extension raises an alarm when the estimated current optimality gap exceeds the tolerance.It estimates preferences during a baseline burn-in period, then evaluates recent-observation windows using benchmark and evaluation splits.
  • Theoretical Guarantees: The estimated gap consistently approaches the oracle gap up to inverse-optimization ambiguity, with separate preference, benchmark, and evaluation-error contributions.When ambiguity is nonzero, the unadjusted test has guarantees on separated nulls and alternatives; a sensitivity-adjusted threshold provides a conservative test when an ambiguity bound is available.

4 Experiments

Experiments evaluate RADAR across synthetic optimization, Fashion-MNIST capacity allocation, and Atlanta police districting, comparing decision-focused adequacy tests with distributional and risk-based alternatives. Across these settings, RADAR better controls false alarms under harmless shifts while detecting shifts that create incumbent suboptimality.

  • Experimental settings: RADAR is evaluated on synthetic QP and LP problems, a semi-synthetic Fashion-MNIST capacity-allocation task, and a real Atlanta police districting study.Synthetic and capacity-allocation experiments use two-sample tests with oracle-labeled decision-irrelevant and decision-relevant shifts; the police study uses temporal monitoring.
  • Synthetic optimization: Under synthetic decision-irrelevant shifts, RADAR stays near the α = 0.05 nominal level, while baselines often reject and inflate Type-I error.The synthetic experiments include expectation-risk QP and CVaR-risk LP settings with different notions of decision relevance.
  • Synthetic optimization: Under synthetic decision-relevant shifts, RADAR’s power increases with shift magnitude δ as the oracle optimality gap grows.The comparison includes generic context two-sample tests and risk-value baselines.
  • Capacity allocation: Fashion-MNIST H0 shifts alter fine image classes while preserving group shares and the forward objective, whereas balanced H1 shifts change shares from (0.60, 0.25, 0.15) to (1/3, 1/3, 1/3).An orthogonal H1 shift changes the optimal allocation with little impact on realized cost, separating decision adequacy from deployed risk.
  • Capacity allocation: In Fashion-MNIST results, context tests increasingly reject the harmless H0 image shift, while Risk-Value and X-Mean miss the orthogonal H1 shift; RADAR targets the optimality gap.Table 1 reports empirical rejection rates with 90% Wilson confidence intervals over 50 repetitions at α = 0.05.
  • Police districting: In Atlanta, RADAR first crosses its sequential detection threshold in June 2016 after nearly five years, while the baseline rejects at every monitoring time, including the prealarm period.The June 2016 alarm precedes measured workload deterioration by roughly a year and district reconfiguration by two and a half years; the later plan reduced workload imbalance by 43% and high-priority response time by 5.8%.
  • Robustness: A set-identified preference analysis finds decision-irrelevant shifts can yield a determinate zero gap, while larger decision-relevant shifts warrant re-optimization and smaller ones may remain indeterminate.The conclusion is read from the target-gap range over the entire inverse-feasible preference set relative to tolerance τ.

5 Discussion and Future Work

RADAR is designed to assess whether distributional change has made a deployed decision suboptimal, while its validity depends on the forward model, preference identification, and evaluation data. The paper uses one-sided Wald and bootstrap procedures, with sample splitting separating learned components from gap evaluation.

  • Scope and limitations: RADAR’s estimated optimality gap need not reflect realized operational loss when the specified objective family is misspecified.The framework’s interpretation is conditional on the forward model.
  • Scope and limitations: Weakly rationalized incumbent decisions widen the identified gap range, and audits are indeterminate when that range straddles τ.The limitation concerns ambiguity in the latent preference parameter.
  • Scope and limitations: A non-rejection is inconclusive because it is consistent with adequacy, preference ambiguity, or limited evaluation data.Thus, failure to reject does not by itself establish that the incumbent remains adequate.
  • Testing procedures: Sample splitting separates preference and challenger construction from final gap evaluation, supporting valid inference for the deployment optimality gap.The evaluation data are not reused for the learned components.
  • Testing procedures: The Wald test rejects when its one-sided statistic exceeds z1−α, using the standard normal approximation for the evaluation-sample estimator.The variance estimate is computed from the evaluation subsample.
  • Testing procedures: Bootstrap testing is offered for general or nonsmooth risk functionals, small evaluation samples, and settings where closed-form variance estimation is inconvenient.It recomputes estimated gaps over bootstrap samples drawn from the evaluation subsample.

B Decision-Relevant Change-Point Detection

Decision-relevant sequential monitoring tests whether the current window makes the deployed decision inadequate, rather than merely detecting a distributional change. RADAR uses trailing windows, sample-split gap estimation, and a first-crossing alarm, thereby ignoring safe shifts and responding to sustained inadequacy.

  • Monitoring objective: RADAR tests whether the post-change distribution makes z0 inadequate, whereas standard CPD tests whether the context distribution changes.The decision-safe set Sτ(z0) contains distributions under which z0 is τ-optimal.
  • Monitoring objective: The sequential hypotheses distinguish no change, a safe change with P1 ∈Sτ(z0), and an unsafe change with P1 ∉Sτ(z0).The baseline is assumed to be decision-safe.
  • Decision-relevant monitoring: Standard CPD flags both changes, but RADAR ignores the variance-only shift at γ1 = 1000 and keeps rejecting after the mean shift at γ2 = 2000 makes z0 inadequate.CPD rejection decays after each change leaves its comparison windows, whereas RADAR remains alarmed in the unsafe regime.
  • Window construction: Each monitoring time uses a trailing window of w recent observations, with no reuse of burn-in data.The window is Wt = {t − w + 1, …, t}.
  • Window construction: RADAR randomly splits each window into benchmark and evaluation samples, constructs a challenger, and estimates the incumbent’s current optimality gap on held-out data.The evaluation split also supplies the standard deviation used in the test statistic.
  • Alarm rule: The alarm time is the first monitoring time at which Tt exceeds the Bonferroni critical value q = z1−α/N.The procedure returns infinity if no monitored statistic crosses the threshold.
  • Alarm rule: Because the statistic uses a trailing window, an alarm can lag the onset of inadequacy by at most w observations.The procedure reports when re-examination is warranted rather than when the environment changed in the past.

Theoretical guarantees

The theoretical results establish validity under safe windows and power when the optimality gap exceeds the tolerance by a resolvable margin. Convexity protects against false alarms from windows that mix safe regimes, while detection delay is governed by windowing and the monitoring grid.

  • Setup: The fixed monitoring grid contains N equally spaced times, with asymptotics taken in the sample sizes while N, stride s, and α remain fixed.This matches a deployed monitor with a fixed review schedule and growing data behind each review.
  • Window mixtures: A window straddling a change has distribution Pt = λtP0 + (1 − λt)P1, so its law is a mixture of the pre- and post-change regimes.The fraction λt denotes the portion of the window before the change.
  • Window mixtures: The decision-safe set is convex because the incumbent’s optimality gap is convex in the operating distribution.Therefore, mixtures of safe distributions remain safe.
  • Type-I error: A safe change cannot trigger a transition alarm: every window remains safe, so the no-false-alarm guarantee applies across the monitoring horizon.The result uses convexity together with the local validity assumption and Bonferroni thresholding.
  • Power: The companion alarm-time result requires no changepoint structure; it only requires inadequacy at a monitored time by a margin resolvable by the evaluation split.This supports detection under sustained or gradual degradation, not only abrupt changes.
  • Power: If some monitoring time has ε = Δt⋆ − τ > 0 and the studentized margin diverges beyond q, RADAR alarms by t⋆ with probability tending to one.The relevant signal is the evaluation-sample gap excess divided by its estimated standard deviation.
  • Detection delay: After an abrupt change at γ, the first fully post-change window is monitored no later than γ + w − 1 + (s − 1), under the theorem’s margin condition.The bound combines the trailing-window width with the spacing of the monitoring grid.

C Inverse Optimization Ambiguity

Inverse optimization may identify a set of preferences rather than a single parameter, but the resulting adequacy conclusion can remain stable. RADAR quantifies this identification sensitivity through target-gap ranges over compatible preferences.

  • Preference identification: In constrained or piecewise-linear problems, a vertex decision is rationalized by every preference parameter in its normal cone, producing set rather than point identification.The inverse-feasible set can therefore be non-singleton even when the deployed decision is observed.
  • Preference identification: A non-singleton inverse-feasible set does not invalidate adequacy testing when compatible preferences induce nearly identical target-domain optimality gaps.The downstream adequacy question can remain identified despite preference non-identification.
  • Decision-irrelevant shifts: Under expectation risk, shifts preserving the mean leave the target gap unchanged across compatible preferences, yielding δinv = 0 even without point identification.Positive rescaling of the cost vector also preserves the gap for positively homogeneous objectives.
  • Sensitivity analysis: RADAR reports lower and upper target-gap estimates over approximate rationalizers; conclusions are adequate below τ, warrant re-optimization above τ, and indeterminate when the range straddles τ.If the empirical range crosses τ, the update decision is identification-sensitive.
  • Sensitivity analysis: In the linear-program experiment, the inverse-feasible normal cone covered 40.6% of the preference simplex, while decision-relevant shifts warranted re-optimization once the lower gap exceeded τ.The experiment used K = 6 candidate decisions, d = 3 preference dimensions, and τ = 0.005.

D Theoretical Proofs

The theoretical analysis establishes testing guarantees under general risk-based objectives, using compactness, regularity, integrability, and inverse-stability assumptions. The framework covers expectation risk and OCE/CVaR-type risks with independent samples across estimation stages.

  • Proof strategy: The proofs establish consistency and convergence rates for each error component, the test statistic, and the resulting hypothesis-testing guarantees.The analysis decomposes preference estimation, benchmark construction, and evaluation effects.
  • Risk framework: The theory uses a general risk-based objective and optimality gap, with expectation risk recovered as a special case.The covered class includes expectation risk and OCE/CVaR-type risks.
  • Assumptions: The statistical validity conditions require independent samples, conditional variance control, a Lindeberg condition, and consistent variance estimation.These conditions support the Wald-style test analysis.
  • Assumptions: Compact parameter and decision spaces, objective regularity, envelope conditions, and Lipschitz continuity control existence and uniform convergence.These conditions support empirical-process arguments for the induced function classes.
  • Assumptions: The inverse-stability condition requires the baseline gap to certify local proximity to the inverse-feasible set at a polynomial rate.Its degree κ determines how parameter estimation error translates into distance from compatible preferences.

Proposition, Lemmas and Theorems

The lemmas prove uniform convergence and consistency of inverse preference estimation, target optimization, and evaluation. Under inverse stability, the deployment-gap estimation rate is governed by κ, while the Wald test controls size and achieves power under separated alternatives.

  • Uniform convergence: The OCE function class is P-Glivenko–Cantelli under compactness, continuity, and an integrable envelope.This supplies uniform convergence of empirical OCE risks over the indexed decision, preference, and auxiliary-parameter class.
  • Consistency: The empirical inverse estimator is set-consistent: dist(ˆθ, Θinv) converges to zero in probability.Any limit point of the estimator lies in the inverse-feasible set.
  • Consistency: The benchmark and evaluation components are consistent under target-domain square-integrability of the envelope and Lipschitz modulus.The decomposition treats benchmark construction and evaluation as separate error terms.
  • Rates: Under κ-degree inverse stability and εn = Op(n^-1/2), preference distance and the associated deployment-gap error scale as Op(n^-1/(2κ)) when δinv = 0.The rate follows from uniform empirical-risk convergence and Lipschitz continuity of the deployment gap.
  • Testing guarantees: Under δinv = 0, the one-sided Wald test has rejection probability at most α + o(1) under the null, while power tends to one when ∆⋆ ≥ τ + δinv + η.A sensitivity-adjusted threshold τ + ¯δinv offsets the largest admissible identification error.

Implementation Configuration

The experiments use differentiable optimization layers for forward and inverse optimization, with multi-start gradient training to estimate latent preferences. Synthetic and Fashion-MNIST runs take under one minute per repeated run, while Atlanta evaluations take several seconds per sliding window.

  • Optimization pipeline: Forward optimization uses cvxpylayers with ECOS, while inverse optimization estimates latent preferences through multi-start gradient training with a suboptimality loss.The default configuration uses 300 epochs, learning rate 5×10^-3, and 10 random initializations.
  • Data configuration: Preference estimation, challenger construction, and evaluation use 3000, 2000, and 2000 contexts, respectively.All experiments were implemented in PyTorch with differentiable optimization layers.
  • Compute: Synthetic and Fashion-MNIST experiments require less than one minute per repeated run, whereas Atlanta districting requires several seconds per sliding-window evaluation.Runs used an Apple M4 machine and a Windows PC with an NVIDIA RTX 5070 Ti GPU.

Details of Synthetic Experiments

The synthetic and application experiments compare RADAR with tests of context or realized-risk changes across optimization settings and structured distribution shifts. Evaluation focuses on empirical rejection rates, with additional monitoring applied to police districting data.

  • Baselines: RADAR is compared with tests targeting context means, full context distributions, and realized risk of the deployed decision.The alternatives include X-Mean, X-Distr, and Risk-Value.
  • Evaluation: RADAR tests whether the incumbent decision is more than tolerance τ suboptimal relative to the target-domain optimum.Empirical rejection rates represent Type-I error when the null holds and power when it is violated.
  • Optimization settings: The synthetic study spans LP with CVaR risk and QP with expectation risk, representing tail-sensitive and average-performance regimes.The LP and QP settings use simplex-constrained context, preference, and decision variables.
  • Distributional shifts: Context shifts are generated by perturbing Dirichlet population parameters from α to α′, with magnitude δ indexing the perturbation size.The structured shifts are summarized in Table 3.
  • Capacity allocation: The capacity-allocation experiment uses Fashion-MNIST images grouped into clothing, footwear, and accessory categories with latent shortage priorities.Expected loss can be expressed through induced group probabilities while the stochastic problem remains defined over image contexts.
  • Police-zone planning: The police-zone monitoring study uses quarterly windows from December 2011 through March 2021 and evaluates districting gaps at relative tolerance τ = 0.2 and level α = 0.01.Each 52-week window is split into benchmark-construction and evaluation data, and the decision-agnostic comparison is a two-sample Hotelling test.
Loading 2608.29465v1…