Source-linked AI summary

Quantifying Distributional Model Risk via Optimal Transport

Jose Blanchet, Karthyek R. A. Murthy

arXiv:1604.01446v2math.PRmath.ST

TL;DR

The paper addresses how to quantify the effect of model misspecification on general expected values, especially for stochastic processes and path-dependent functionals. It uses optimal-transport ambiguity sets around a baseline model to derive robust expectation bounds, dual formulations, and applications including ruin probabilities and first-passage events. It also discusses non-parametric calibration of the tolerance region and conditions governing existence of worst-case transport plans.

  • Problem

    Model misspecification complicates performance and risk evaluation, creating a need for robust estimates of expected values across plausible probability models.

  • Method

    The paper optimizes expectations over transport-cost neighborhoods around a baseline measure, deriving strong duality, a one-dimensional dual reformulation, and coupling-based characterizations.

  • Results

    The framework applies to ruin probabilities, multidimensional first-passage probabilities, and decision-making under model ambiguity, including stochastic-process settings.

  • Takeaways & Limitations

    Optimal-transport ambiguity sets provide computable bounds for expectations while accommodating tractable surrogate models and path-dependent stochastic-process applications.

  • Takeaways & Limitations

    A primal optimal transport plan need not exist in general, although additional convexity and concavity conditions can ensure existence and uniqueness.

Abstract

from arXiv · show

This paper deals with the problem of quantifying the impact of model misspecification when computing general expected values of interest. The methodology that we propose is applicable in great generality, in particular, we provide examples involving path dependent expectations of stochastic processes. Our approach consists in computing bounds for the expectation of interest regardless of the probability measure used, as long as the measure lies within a prescribed tolerance measured in terms of a flexible class of distances from a suitable baseline model. These distances, based on optimal transportation between probability measures, include Wasserstein's distances as particular cases. The proposed methodology is well-suited for risk analysis, as we demonstrate with a number of applications. We also discuss how to estimate the tolerance region non-parametrically using Skorokhod-type embeddings in some of these applications.

1. Introduction.

The paper develops a framework for assessing model misspecification in expected performance or risk measures by bounding expectations over transport-based neighborhoods around a baseline model. It establishes general optimal-transport results and applies them to stochastic processes, path-dependent expectations, and risk analysis.

  • Model misspecification can arise from limited data or specific parametric choices, motivating robust estimates of performance and risk.
  • The framework bounds an expectation regardless of the probability measure, provided it remains within a prescribed tolerance of a suitable baseline model.The tolerance neighborhood interpolates between no ambiguity at δ = 0 and greater model uncertainty as δ increases.
  • Transport costs provide flexible discrepancy measures that include Wasserstein distances and avoid the absolute-continuity requirement associated with relative entropy.This flexibility is especially relevant for stochastic-process models, where relative entropy can be infinite even when two measures generate similar samples with high probability.
  • The paper derives a dual formulation with strong duality and reduces the infinite-dimensional dual problem to a tractable one-dimensional reformulation.Under sufficient conditions, the paper also characterizes an optimizer through a coupling involving the baseline measure and the transport cost.
  • Applications cover robust ruin probabilities, multidimensional first-passage probabilities, and decision-making under model ambiguity.The approach can use tractable surrogates such as Brownian motion, diffusions, and reflected Brownian motion.
  • A Skorokhod-type embedding is proposed for non-parametric calibration of the tolerance parameter δ in tractable surrogate settings.The paper presents this calibration as a way to choose the feasible region around a baseline model.

2. Our Main Result.

The paper formulates model uncertainty through optimal-transport neighborhoods around a baseline measure and establishes strong duality for worst-case expectations. The resulting one-dimensional characterization supports optimizer structure and tractable worst-case probability calculations.

  • Optimal transport formulation: The optimal transport cost is the minimum expected cost over all couplings between two probability measures, with Wasserstein distance as a special case.When the cost is symmetric and satisfies the triangle inequality, the transport cost defines a metric; unlike KL divergence, Wasserstein neighborhoods need not preserve the baseline support.
  • Optimal transport formulation: The ambiguity set consists of probability measures within tolerance δ of a baseline measure under an optimal-transport cost.The tolerance controls the level of model uncertainty, while the worst-case objective bounds the expectation regardless of which measure in the neighborhood is used.
  • Strong duality: Under the stated assumptions, the primal worst-case expectation equals its dual value, and a dual optimizer exists in the form (λ, φλ).The theorem also characterizes when feasible primal and dual solutions attain the same finite value.
  • Strong duality: The dual problem reduces the infinite-dimensional optimization to a univariate infimum involving only the baseline measure µ.This makes the formulation easier to work with when the baseline model is characterized or sampled directly.
  • Optimizer structure: A worst-case transport plan moves each x to a maximizer of f(y) − λ∗c(x,y), while sufficiently large ambiguity budgets can move all mass to global maximizers of f.If the local maximizer is unique µ-almost everywhere, the primal optimal transport plan is unique; when λ∗ = 0, the objective reaches supz∈S f(z).
  • Worst-case probabilities: For worst-case probabilities, the optimal value can be expressed as the baseline probability of an inflated neighborhood of the target set.Under the theorem’s continuity condition, this yields I = µ(x : c(x,A) ≤ 1/λ∗); if the condition fails, an optimal coupling may require randomization between extreme cases.

3. Computing ruin probabilities: A first example.

The example computes worst-case ruin probabilities around a Brownian baseline using transport-based ambiguity, then interprets the resulting capital requirements for insurance risk.

  • Model and objective: The Cramer-Lundberg model uses initial reserve, safety loading, claim-arrival rate, and claim-size distribution to characterize insurance ruin.The target is the probability that the insurer becomes bankrupt before a specified duration T.
  • Model and objective: Because exact finite-horizon ruin probabilities can be difficult to compute, a Brownian diffusion approximation based on the first two claim-size moments provides a tractable substitute.The paper notes that the approximation’s accuracy is difficult to verify, motivating robust estimates around the Brownian baseline.
  • Robust computation: Worst-case ruin probabilities optimize over measures within a transport-cost neighborhood of the Brownian motion measure.The path space is the Skorokhod space D([0,T], R) with the J1 topology, and the ruin event is represented by a closed set.
  • Robust computation: Model ambiguity reduces the initial reserve to a modified level ũ = u − (λ*)^-1/p, so the robust probability becomes a Brownian level-crossing probability at that lower level.This gives the worst-case estimate an elementary interpretation as crossing a reduced reserve threshold.
  • Numerical and regulatory implications: For T = 100, ν = 1, p = 2, and η = 0.1, the Brownian approximation requires u ≥ 60, whereas the robust estimate requires u ≈ 200 to keep ruin probability below 0.01.The robust requirement is roughly three times the Brownian model’s capital requirement.
  • Numerical and regulatory implications: The robust capital requirement can serve as a regulatory minimum and shows that increasing premium income alone cannot arbitrarily reduce ruin probability.The robust criterion requires sufficiently large initial capital as well as an appropriate safety loading.

4. Proof of the duality theorem.

The proof establishes strong duality and optimizer existence for transport-based distributionally robust expectations, first on compact Polish spaces and then through extensions to noncompact settings.

  • Compact spaces: On compact Polish spaces with continuous cost, the primal and dual problems have equal values and a primal optimizer exists.This is stated in Proposition 5 under the relevant assumptions on f and c.
  • Compact spaces: For lower semicontinuous costs, the same equality and optimizer existence follow by approximating the cost with an increasing sequence of continuous costs.The proof uses tightness, weak convergence, monotone convergence, and weak duality to pass to the limit.
  • Noncompact spaces: Extending the result beyond compact spaces requires controlling supports through increasing compact subsets and addressing universal measurability of the dual envelope.The support product is σ-compact, while projections of Borel sets are analytic and therefore universally measurable.
  • Dual formulation: The dual feasible functions are represented through pairs (λ, φ), with λ ≥ 0 and φ satisfying φ(x) + λc(x,y) ≥ f(y).The construction uses bounded continuous functions and convex-analytic functionals before identifying their conjugates.
  • Dual formulation: Fenchel duality applies because the relevant feasible sets have suitable relative-interior and epigraph conditions, yielding equality between the primal and dual values.The supremum is attained by a feasible measure π* in the compact-space argument.
  • Noncompact spaces: Proposition 7 supplies the key intermediate result for extending strong duality to noncompact sets under integrability, marginal, and feasibility conditions.The proposition assumes a probability measure π with finite objective value and the prescribed first marginal.

5. On the existence of a primal optimal transport plan.

The paper distinguishes strong duality from attainment of a primal optimal transport plan: primal optimizers may fail to exist, but additional compactness and upper-semicontinuity conditions ensure existence. Under further growth assumptions, these conditions yield existence, while a concavity-convexity condition can also imply uniqueness.

  • A primal optimizer need not exist because the feasible set Φµ,δ is not necessarily compact, even when strong duality holds.Example 2 demonstrates nonattainment with a metric transport cost on a Polish space.
  • Example 2 has dual value I = 1 but no primal optimizer, because f(x) < 1 everywhere while the supremum is approached but not attained.
  • P-Compactness and P-USC provide sufficient conditions for extracting a weakly convergent feasible subsequence and passing optimality to its limit.P-Compactness supplies tightness, while P-USC gives the upper-semicontinuity step needed for optimality.
  • Under Assumptions (A1)–(A4), local compactness, and λ* > 0, a primal optimizer exists and satisfies I(π*) = I = J = J(λ*, φλ*).
  • If c(x, y) is convex in y and f is concave, the primal optimizer exists and is unique when the local maximizer is unique for every x.

6. A few more examples.

The examples apply the robust transport framework to first-passage and ruin probabilities, including multivariate reserve processes and capital allocation. Worst-case first passage can be reduced to baseline first passage into an inflated target set, while robust reinsurance changes the optimal allocation and increases worst-case loss.

  • 6.1. Applications to computing general first passage probabilities.: Worst-case first-passage probability equals the baseline probability of reaching an appropriately inflated target set.This extends the one-dimensional level-crossing result to general target sets.
  • 6.1. Applications to computing general first passage probabilities.: In the multivariate insurance example, ruin occurs when a two-dimensional reserve process reaches a rule-dependent ruin set B within time T.The parameter β controls the extent of capital transfer permitted between the two business lines.
  • 6.1. Applications to computing general first passage probabilities.: The worst-case first-passage problem optimizes over probability measures within transport distance δ of the baseline measure, using the closure of the hitting event when jumps prevent closedness.
  • 6.1. Applications to computing general first passage probabilities.: The capital requirement is evaluated across δ values so that the worst-case probability of ruin remains below 0.01.
  • 6.1. Applications to computing general first passage probabilities.: Choosing an anisotropic norm can inflate the ruin set differently across directions, but the effects and practical appropriateness of such cost choices remain an important application question.
  • 6.2. Distributionally robust optimal decisions.: For θ = 0.3, the Brownian approximation gives b = 0.66 and L = 17.63, whereas the robust solution gives b = 0.42 and worst-case L′ = 28.86.

Appendix A. Brownian embeddings for the estimation of δ in Example 1

The appendix describes a data-driven coupling of a compound Poisson risk process with Brownian motion to estimate the transport tolerance δ. The construction uses observed claim-size realizations and simulated coupled paths.

  • A coupling between the risk process and Brownian motion provides an upper-bound approach for estimating the optimal transport distance.
  • The embedding procedure assumes Brownian motion, the first claim-size moment, and independent realizations of claim sizes.
  • The construction recursively defines stopping times and auxiliary processes, then applies a time change to produce the coupled process Z(t).
  • The resulting process is closely coupled with Brownian motion, and independent replications are used to simulate coupled risk processes and choose δ.

Appendix B. Some technical proofs

Appendix B states the technical results used in Examples 1, 3, and 4, then proves supporting lemmas for Section 4 and Theorem 1. It closes by clarifying notation for the primal problem I.

  • The section first states and proves technical results used in Examples 1, 3, and 4.
  • Lemmas 15 and 16 provide technical results used in Section 4 to complete Theorem 1’s proof.
  • Lemma 17 and Corollary 3 clarify notation adopted for the primal problem I in Section 2.2.

B.1. Technical results used in Examples 1 and 4. (

This section establishes technical properties of path spaces and the J1 metric used in the paper’s examples. It proves closedness and distance formulas for threshold-crossing sets, then derives a dual optimization identity.

  • The results use S = D([0,T],R) equipped with the J1-topology and metric dJ1.
  • For threshold sets Au, the section proves closedness and characterizes c(x,Au) through the J1 distance.
  • Time changes can be restricted to the identity map when evaluating the relevant inner infimum.
  • The optimization yields φλ(x) = f(x) + a2^2/4λ after matching lower and upper bounds.
  • The section also sets up the J1-topology on D([0,T],R2) for later two-dimensional technical results.

B.2. Proofs of results used in Section 6.1.

This section proves technical results for two-dimensional path constraints under the J1-topology. It establishes closure properties and explicit distances to the constraint sets used in Section 6.1.

  • The two-dimensional path space S = D([0,T],R2) is equipped with the J1-topology induced by the relevant metric.
  • For nonnegative a1 and a2 with a1a2 > 0, the set of paths satisfying a^T x(t) ≤ 0 at some time is closed after taking its closure characterization.
  • The closure is obtained by approximating paths in the nonpositive-infimum set with uniformly shifted paths entering the time-indexed constraint set.
  • For A = A1 ∪ A2, the distance to A is the minimum of the distances to A1 and A2.
  • The distance to A1 is determined by the squared minimum of βx1(t) + x2(t), scaled by 1 + β2; A2 has the analogous expression.

B.3. Technical results used in Section 4 to complete the proof of Theorem 1.

This section supplies technical measure-theoretic and approximation results used to complete Theorem 1. It handles signed measures, increasing feasible sets, and convergent penalty parameters.

  • For compact Polish S × S and upper semicontinuous f, a signed measure with negative mass on a measurable set leads to an unbounded negative infimum over the specified continuous-function class.
  • Lemma 16 studies increasing subsets Sn and parameters λn converging to λ*, using g(y,λ) = f(y) − λc(x,y).
  • The proof relies on g being non-increasing in λ and suprema over Sn being non-decreasing as the sets expand.
  • The approximation argument treats both finite and infinite limiting suprema over the union of the increasing sets.
  • The resulting limit identifies the supremum over the union of Sn at λ* as the desired right-hand side.

B.4. A technical result to add clarity to the notation (49).

Under the stated assumptions, an admissible coupling can be modified without changing its first marginal so that it is concentrated on pairs satisfying f(y) ≥ f(x), while retaining the distance tolerance.

  • Lemma 17: Lemma 17 constructs an admissible coupling π′ concentrated on pairs where f(y) ≥ f(x).The construction applies to every π ∈ Φµ,δ under Assumptions (A1) and (A2).
  • Technical conditions: The proof also establishes that the relevant integral involving f(y) is well-defined under the stated conditions.This completes the claims associated with the technical result and its corollary setup.
  • Construction: The construction sets X′ := X and replaces Y by Y when f(X) ≤ f(Y), otherwise by X.This transformation preserves the marginal distribution of X′ as µ.
  • Cost preservation: Because c(x, x) = 0 and c is non-negative, the modified coupling does not increase the transportation cost.The proof uses these cost properties to retain membership in the tolerance set Φµ,δ.
  • Application: An optimal coupling attaining dc(µ, ν) exists by lower semicontinuity of c, enabling the lemma to be applied whenever dc(µ, ν) ≤ δ.The resulting second marginal ν′ satisfies dc(µ, ν′) ≤ δ.
Loading 1604.01446v2…