Source-linked AI summary
Distributionally Robust Mean-Variance Portfolio Selection with Wasserstein Distances
Jose Blanchet, Lin Chen, Xun Yu Zhou
TL;DR
The paper addresses the sensitivity of empirical mean–variance portfolios to distributional uncertainty by using a Wasserstein ambiguity set around the empirical measure. It converts the robust problem into regularized empirical variance minimization and develops data-driven inference for the ambiguity size and robust target return. The resulting robust strategies are reported to improve out-of-sample performance while retaining comparable computational tractability.
Problem
Mean–variance portfolios are sensitive to uncertain return distributions and empirical estimates that can deviate substantially from the true mean and covariance.
Method
The paper models distributional ambiguity with a Wasserstein-based empirical ambiguity set, derives an equivalent regularized empirical optimization, and extends robust Wasserstein profile inference for data-driven parameter selection.
Results
Robust strategies enhance out-of-sample performance with essentially the same computational tractability as standard mean–variance selection.
Takeaways & Limitations
The framework provides a data-driven distributionally robust Markowitz model whose uncertainty size and worst-case return target are informed by return data rather than chosen exogenously.
Takeaways & Limitations
The paper uses an lq norm for the Wasserstein distance and identifies alternative transportation costs and dynamic DRO Markowitz models as directions for generalization.
Abstract
from arXiv · showhide
We revisit Markowitz's mean-variance portfolio selection model by considering a distributionally robust version, where the region of distributional uncertainty is around the empirical measure and the discrepancy between probability measures is dictated by the so-called Wasserstein distance. We reduce this problem into an empirical variance minimization problem with an additional regularization term. Moreover, we extend recent inference methodology in order to select the size of the distributional uncertainty as well as the associated robust target return rate in a data-driven way.
1 Introduction
The paper formulates mean–variance portfolio selection under distributional ambiguity around the empirical distribution, using Wasserstein distance to model discrepancies. It derives a regularized empirical optimization and proposes data-driven selection of ambiguity and robust-return parameters.
- Motivation: Markowitz portfolio selection chooses stock weights to maximize risk-adjusted expected return, but its solutions are sensitive to estimated means and covariances.The true return distribution is unknown, and empirical estimates—especially the mean—can substantially deviate from the truth.
- Robust formulation: The distributionally robust formulation introduces an adversarial probability measure within a Wasserstein-based ambiguity set around the empirical measure.The ambiguity set controls distributional uncertainty, while the feasible region imposes a worst-case mean-return target.
- Parameter choice: A larger ambiguity radius δ gives the adversary more power, but choosing δ too large relative to sample size can make portfolio selection unnecessarily conservative.The robust target return ᾱ must also account for δ; setting ᾱ equal to the original target ρ tends to produce overly aggressive portfolios.
- Contributions: The robust problem is equivalent to empirical variance minimization with an added regularization term, providing a theoretical justification for regularization in mean–variance selection.The equivalent problems have the same optimal solutions and value, and the resulting optimization remains convex and tractable.
- Contributions: The paper reports that robust strategies can enhance out-of-sample performance with essentially the same computational tractability as standard mean–variance selection.The paper also situates the work within prior Wasserstein-based links to regularization and machine-learning algorithms.
- Contributions: The paper extends robust Wasserstein profile inference to select the ambiguity size δ and worst-case target return ᾱ from historical data under suitable mixing conditions.The procedure combines optimization principles with statistical theory rather than choosing these parameters exogenously.
2 Formulation and Main Results
The paper defines Wasserstein-based distributional ambiguity around the empirical measure and reformulates the resulting robust portfolio problem into tractable convex optimization problems. Under suitable transportation costs, the formulation becomes an empirical-measure problem with cost-dependent regularization.
- 2.1 Basic notation and assumptions: The Wasserstein discrepancy is defined as the minimum transportation cost over couplings between two probability measures.The cost function is assumed lower semicontinuous, non-negative, and zero when both arguments coincide.
- 2.1 Basic notation and assumptions: The ambiguity set is centered on the empirical probability measure Pn constructed from n observed return realizations.The paper uses a Wasserstein distance of order 2 and represents Pn through the observed returns Ri.
- 2.2 Computational tractability: A dual argument reduces the infinite-dimensional variational problem to a two-dimensional optimization in λ1 and λ2.With additional structure in the cost function, including quadratic lq costs, the reduction can be simplified further.
- 2.2 Computational tractability: The robust formulation fixes the portfolio's expected return at α and uses this constraint to make the innermost maximization problem linear in P.This reformulation enables the subsequent optimization over the distributional ambiguity set.
- 2.2 Computational tractability: For a general lower-semicontinuous non-negative cost, Proposition 2 gives the optimal value of the inner distributional problem.The proposition supplies the key value-function characterization used in the computational reformulation.
- 2.2 Computational tractability: The primal robust problem and its dual have the same optimal solutions and optimal value.The resulting formulation is therefore an equivalent optimization representation rather than merely a relaxation.
- 2.2 Computational tractability: The resulting problems are convex optimization problems because the portfolio variance mapping and feasible region are convex.The formulation also provides a justification for regularization techniques commonly used in practical mean–variance portfolio selection.
3 Choice of Model Parameters
The paper selects the uncertainty size δ and robust target return ᾱ from statistical principles, balancing confidence in the classical optimum against the strength of robustification. Under its assumptions, the recommended uncertainty size has order O(n^-1), while ᾱ is calibrated asymptotically using a normal approximation.
- Assumptions and calibration: The parameter-selection analysis assumes stationary, ergodic returns with finite fourth moments, an identifiability condition, and a unique classical optimum.The uniqueness assumption simplifies calibration; if multiple classical optima exist, the analysis would require a different confidence-region criterion.
- Choice of δ: δ should be chosen from the data rather than arbitrarily, because excessive ambiguity weakens data relevance while insufficient ambiguity makes robustification negligible.The paper defines the uncertainty region large enough for the classical optimal portfolio to be plausible with a specified confidence level.
- Choice of δ: Any δn of order o(n^-1) is too small, so an appropriate asymptotic order is O(n^-1).This order follows from the convergence rate of the difference between the classical and empirical optimal values.
- Choice of δ: Λδ(Pn) is treated as a confidence region containing portfolios optimal for probability measures within Wasserstein distance δ of the empirical measure.The selected δ is the smallest uncertainty size making the true optimal portfolio belong to this region at the desired confidence level.
- Choice of δ: The RWP function provides a computational route for determining the confidence-calibrated uncertainty size, although the direct statistic is cumbersome because it requires a mean-and-variance minimization.An alternative statistic using only empirical mean and variance yields an upper bound while preserving the target convergence rate O(n^-1).
- Choice of ᾱ: After δ is selected, ᾱ is chosen just large enough that the classical optimum remains feasible with the user-specified confidence level.The calibration uses a consistent empirical optimizer and an approximately normal statistic, whose 1−ϵ quantile determines υ0 and hence ᾱ.
4 Concluding Remarks
The paper presents a data-driven distributionally robust theory for Markowitz mean–variance portfolio selection, reducing the robust model to an empirical problem with regularization. It also develops a scheme that infers the uncertainty-region size from return data, while identifying alternative transportation costs and dynamic extensions as generalization directions.
- 4 Concluding Remarks: The robust Markowitz model is equivalent to empirical variance minimization with an additional regularization term.This connects distributional robustness with regularization-based variance minimization.
- 4 Concluding Remarks: The uncertainty-region size is informed by return data rather than specified exogenously.The paper develops a scheme for data-driven selection of this size.
- 4 Concluding Remarks: The chosen lq norm could be replaced by other transportation costs, including costs related to adaptive regularization or industry clusters.The paper identifies these alternatives as possible generalizations.
- 4 Concluding Remarks: A dynamic discrete-time or continuous-time version of the distributionally robust Markowitz model is another significant extension direction.
A Proof of Proposition 1
The proof of Proposition 1 begins by formulating the relevant optimization problem and then deriving its dual using Slater’s condition and a cited Wasserstein-distance result.
- A Proof of Proposition 1: The proof first states the optimization problem used for Proposition 1.
- A Proof of Proposition 1: Slater’s condition and Proposition 4 of Blanchet, Kang and Murthy (2016) are used to obtain the dual problem.
B Proof of Proposition 2
The proof transforms the dual formulation by introducing a deterministic slack variable, rewriting constraints, and applying an interiority condition to establish equality with a dual problem.
- B Proof of Proposition 2: A slack random variable S ≡ v is introduced, with v deterministic, to recast the problem.
- B Proof of Proposition 2: The reformulation imposes an expected transportation-cost constraint, fixes the return marginal at Pn, and sets S equal to v.
- B Proof of Proposition 2: The transformed formulation uses indicator constraints for the observed returns Ri and the deterministic slack value v.
- B Proof of Proposition 2: The proof defines augmented functions and measures, then uses the fact that q̃ lies in the interior of Q f̃ when φ ≠ 0.
- B Proof of Proposition 2: Proposition 6 of Blanchet, Kang and Murthy (2016) is invoked to equate the optimal value of the primal problem with that of its dual.
- B Proof of Proposition 2: The dual constraints require the transformed expression involving v, c(u,r), and s to dominate (φTu)^2 over the feasible domain.
- B Proof of Proposition 2: Substituting r = Ri yields the corresponding inequality for each empirical return observation.
- B Proof of Proposition 2: The proof successively rewrites the dual problem, including transformations and replacements involving λ1, λ2, and auxiliary coefficients.
C Proof of Proposition 3
The proof of Proposition 3 separates several cases, identifies one non-trivial case, optimizes with respect to λ2, and derives feasibility conditions for the primal problem.
- C Proof of Proposition 3: The proof handles the first three cases separately before addressing a non-trivial fourth case.
- C Proof of Proposition 3: Taking the partial derivative with respect to λ2 and setting it to zero provides a necessary optimization condition.
- C Proof of Proposition 3: The resulting condition is substituted back into the preceding expression to obtain the next formulation.
- C Proof of Proposition 3: When p < 0, the problem’s optimal value is −∞, indicating that the primal problem is infeasible.
- C Proof of Proposition 3: The proof concludes by rewriting problem (7) under the derived conditions.
D Proof of Theorem 2
The proof characterizes the limiting random quantities by combining asymptotic assumptions, a scaled parameterization, and derivative-based matrix conditions. It then uses these relations to establish the variance conclusion under A3).
- Asymptotic setup: The proof invokes a prior proposition for fixed μ and Σ before introducing the asymptotic scaling.The scaling sets Δ=¯Δ/n^1/2, ¯λ=λn^1/2, and ¯Λ=Λn^1/2.
- Asymptotic setup: Under A2), ¯Δ and ¯λ can be restricted to compact sets with high probability, making ¯Δ^T¯Λ¯Δ/n^1/2 asymptotically negligible.This compactness argument is attributed to the proof technique of Blanchet, Kang and Murthy (2016).
- First-order conditions: Differentiating with respect to each row ¯Λi· produces an equation that is solved through successive equivalence conditions.The displayed derivation proceeds from the row-wise derivative to a condition that holds if and only if another relation holds.
- Conclusion: The proof concludes by applying A3) to the resulting variance expression.The supplied passage introduces the final variance step but does not include the complete displayed expression.
- Limit characterization: The proof identifies the limiting random quantity Z through the normalized sample sum and covariance fluctuation Hn:=Σn−Σ∗.The normalized sum converges to Z0∼N(0,Υg1) under A1).