Source-linked AI summary
Orthogonal Statistical Learning
Dylan J. Foster, Vasilis Syrgkanis
TL;DR
The paper addresses statistical learning when the target risk depends on an unknown nuisance parameter estimated from data. It proposes a two-stage sample-splitting framework using Neyman orthogonality, showing second-order nuisance effects and oracle rates under complexity conditions. The results broaden excess-risk guarantees to weaker assumptions and complex nonparametric target classes.
Problem
The paper studies how to obtain low excess risk for a target parameter when its population risk depends on an unknown nuisance parameter.
Method
A two-stage sample-splitting meta-algorithm combines arbitrary target and nuisance estimators, using Neyman orthogonality to reduce nuisance sensitivity.
Results
Neyman orthogonality makes nuisance estimation error second order and, under metric-entropy conditions, permits oracle rates for complex target and nuisance classes.
Takeaways & Limitations
Excess-risk analysis supports causal and other nuisance-dependent learning with weaker assumptions, including misspecification, nonidentifiability, and rich nonparametric targets.
Takeaways & Limitations
The framework focuses on plug-in nuisance estimation and relies on Neyman orthogonality, whereas specialized estimators or additional structure can yield more refined guarantees.
Abstract
from arXiv · showhide
We provide non-asymptotic excess risk guarantees for statistical learning in a setting where the population risk with respect to which we evaluate the target parameter depends on an unknown nuisance parameter that must be estimated from data. We analyze a two-stage sample splitting meta-algorithm that takes as input arbitrary estimation algorithms for the target parameter and nuisance parameter. We show that if the population risk satisfies a condition called Neyman orthogonality, the impact of the nuisance estimation error on the excess risk bound achieved by the meta-algorithm is of second order. Our theorem is agnostic to the particular algorithms used for the target and nuisance and only makes an assumption on their individual performance. This enables the use of a plethora of existing results from machine learning to give new guarantees for learning with a nuisance component. Moreover, by focusing on excess risk rather than parameter estimation, we can provide rates under weaker assumptions than in previous works and accommodate settings in which the target parameter belongs to a complex nonparametric class. We provide conditions on the metric entropy of the nuisance and target classes such that oracle rates of the same order as if we knew the nuisance parameter are achieved.
1 Introduction
The paper studies statistical learning when the evaluated population risk depends on an unknown nuisance parameter, with applications including causal prediction, treatment effects, and policy optimization. It develops sample-splitting and orthogonality-based guarantees showing that nuisance error can have only a higher-order effect while supporting oracle rates for complex classes.
- Motivation: Causal applications require estimating nuisance quantities to de-bias predictive models and learn targets with causal interpretations from observational data.Examples include propensity scores for policy rewards and nuisance functions used to construct treatment-effect proxy labels.
- Problem formulation: The framework evaluates learners by excess risk at the unknown true nuisance parameter, whose interpretation can include mean squared error or policy regret.The target risk depends on both the target parameter and nuisance parameter, so the goal is low excess risk relative to the true nuisance value.
- Method: A two-stage sample-splitting meta-algorithm reduces nuisance-component learning to standard statistical learning while treating target and nuisance estimators as black boxes.The analysis assumes only that the target-learning algorithm performs well when a nuisance value is fixed, rather than specifying a particular model family.
- Main results: Neyman orthogonality makes nuisance estimation error enter overall excess risk at second order, enabling oracle-rate guarantees relative to knowing the nuisance parameter.The guarantee supports more complex nuisance models and, under complexity conditions, target rates matching the known-nuisance benchmark up to lower-order terms.
- Fast rates: Under strong convexity, nuisance error contributes order R_G^4 and n^-1/4 nuisance RMSE suffices for parametric targets.Strong convexity is imposed on prediction rather than parameter space, and the treatment-effect loss is an example.
- Slow rates and extensions: Without strong convexity, slow-rate settings such as policy optimization have nuisance impact of order R_G^2, again allowing n^-1/4 nuisance RMSE for parametric targets.Metric-entropy conditions extend oracle-rate results beyond parametric targets, while specialized analyses cover ERM and variance-penalized ERM.
- Scope of guarantees: Focusing on excess risk yields guarantees under weaker assumptions, including misspecification, nonidentifiability, and sparse high-dimensional prediction without restricted eigenvalue assumptions.The framework is applied to heterogeneous treatment effects, offline policy optimization, domain adaptation, and learning with missing data.
2 Framework: Statistical Learning with a Nuisance Component
The framework studies learning a target predictor when its population loss depends on an unknown nuisance predictor. It uses sample splitting to estimate the nuisance and target while evaluating performance through excess risk at the true nuisance value.
- Framework setup: The learner observes i.i.d. samples and selects target functions from Θ and nuisance functions from G, defined on subsets of an abstract observation space.The target and nuisance values lie in finite-dimensional vector spaces, while Θ and G have associated pre-norms.
- Framework setup: Performance is measured by a population loss LD(θ, g), which maps a target predictor and nuisance predictor to a real-valued loss.Typically, LD is the expectation of a pointwise loss, although the general framework does not require that structure explicitly.
- Learning objective: The nuisance parameter g0 is unknown, so the goal is to produce a target predictor with low excess risk evaluated at g0.The target predictor is judged against the best achievable population risk under the true nuisance value.
- Learning objective: The method always uses a sample-splitting meta-algorithm that takes a nuisance predictor bg as an input to target learning.The framework separates nuisance prediction from the subsequent target-prediction stage.
- Assumptions: The framework does not generally assume that a target minimizer θ0 exists, although convexity and a first-order condition can make θ0 serve as a population-risk minimizer.When θ0 belongs to Θ and the population risk is convex, the minimizer can be taken as θ0 without loss of generality.
3 Orthogonal Statistical Learning
Orthogonal statistical learning uses algorithm-independent sample splitting to learn a target under an unknown nuisance parameter, making nuisance error higher-order under suitable loss conditions. The framework provides excess-risk guarantees under weaker prediction-space assumptions and supports complex target classes, including causal-learning examples.
- The framework uses rate guarantees for arbitrary first- and second-stage algorithms, requiring assumptions on the population risk rather than particular learning procedures.The meta-algorithm reduces nuisance-component learning to standard statistical learning and treats target and nuisance estimators as black boxes.
- Under fast-rate conditions, nuisance error of order ε contributes at order ε4, whereas slow-rate results contribute at order ε2.Fast rates require stronger assumptions on the loss, and which regime is preferable depends on target- and nuisance-class complexity.
- Neyman orthogonality makes nuisance estimation affect target excess risk at higher order rather than first order.A de-biasing correction can reduce nuisance-error effects on the loss gradient, and orthogonality is defined through a vanishing functional cross-derivative.
- 3.2 Beyond Strong Convexity: Slow Rates: Theorem 2 gives a slower excess-risk guarantee with a squared nuisance-rate term and requires weaker assumptions than the fast-rate theorem.For generic Lipschitz losses, orthogonality changes the sufficient nuisance condition for oracle rates from ∥bg −g0∥G = o(n−1) to ∥bg −g0∥G = o(n−1/2).
- 3.1 Fast Rates Under Strong Convexity: Theorem 1 gives an excess-risk bound combining the second-stage rate with a fourth-power nuisance-rate term under strong convexity in prediction space.The relevant curvature assumption is weaker than parameter-space strong convexity because the analysis targets prediction error rather than parameter recovery.
- Examples and extensions: The framework supplies oracle-rate conditions for metric-entropy-controlled nuisance and target classes, including nonparametric classes, and applies to residualized and doubly robust losses.In a treatment-effect example, an O(n−1/4) nuisance L2-error rate suffices to achieve the optimal no-nuisance rate; doubly robust formulations yield products of nuisance estimation rates.
4 Instantiating the Main Results: Plug-In Empirical Risk Minimization
This section instantiates the general orthogonal-learning framework with plug-in ERM and derives oracle excess-risk guarantees using empirical-process complexity measures. It treats both fast and slow rates, including nuisance robustness and applications to complex target classes.
- Algorithm: Plug-in ERM estimates the nuisance component on one split, then minimizes the empirical loss obtained by plugging that estimate into the second-stage risk.The analysis uses localized empirical-process tools to control the second-stage rate.
- General guarantee: The nuisance estimate affects the oracle excess-risk bound only through lower-order terms, so classical ERM bounds carry over up to constants.The result is obtained by bounding the second-stage rate and applying the main theorems.
- Fast rates: In the fast-rate regime, the critical radius of the target class Θ continues to govern the rate under strongly convex losses.The framework extends local Rademacher-complexity analysis to estimated nuisance components and covers sparse linear models, neural networks, and kernels.
- Slow rates: In the slow-rate regime, a moment-penalized plug-in ERM has a leading term equal to the critical radius multiplied by the optimal target loss variance.This improves over prior variance-penalized bounds whose leading term typically depends on single-scale metric entropy.
- Nuisance robustness: Orthogonality makes nuisance dependence scale as ∥bg −g0∥4 when r = 0, allowing complex nuisance classes without spoiling the target-class rate.The proof first controls the second-stage rate using empirical-process tools, then invokes orthogonality for the final guarantee.
- Complexity conditions: Metric-entropy conditions on target and nuisance classes yield oracle rates and extend the framework beyond parametric target classes.The plug-in ERM results are expressed through standard complexity measures for Θ and can support complex nonparametric classes.
5 Instantiating the Main Results: Sufficient Conditions for Oracle Rates
This section develops oracle-rate guarantees for orthogonal statistical learning using metric-entropy conditions, generic estimation algorithms, and sample splitting. It extends minimax-rate results to settings with nuisance parameters, including complex nonparametric target classes.
- Complexity conditions: Metric entropy quantifies class complexity: p = 0 covers parametric classes, whereas p > 0 includes nonparametric classes such as Lipschitz, smooth, and kernel spaces.The theorem indexes nuisance and target complexity through p1 and p2, defined from output-coordinate projections.
- Square-loss setting: For square losses, nuisance parameters are modeled through regression functions, with orthogonality characterized using gradients of the loss with respect to target and nuisance arguments.The nuisance is defined through conditional expectations, while the sufficient orthogonality condition uses the gradients ∇ζ and ∇γ.
- Algorithms: Skeleton Aggregation extends the framework beyond plug-in ERM by constructing data-dependent covers and aggregating them across sample splits.The method supplies first-stage rates and an extension for relating target rates evaluated at estimated versus true nuisance parameters.
- Complexity conditions: p1 < 2p2 + 2 is a sufficient relationship between nuisance and target complexity for the well-specified oracle-rate result.The condition is imposed alongside the theorem’s regularity and entropy assumptions.
- Oracle rates: Theorem 5 matches the minimax rate without nuisance parameters, using Skeleton Aggregation for stage one and either plug-in ERM or Skeleton Aggregation for stage two.Plug-in ERM suffices when p2 ≤ 2; Skeleton Aggregation is used for stage two when p2 > 2.
- Oracle rates: When the target class is parametric, it suffices to require p1 < 2, recovering the usual semiparametric-inference setup.The resulting oracle rate is of order Θ(n^(-2/(2+p2))) as stated for the well-specified case.
6 Discussion
The discussion frames orthogonality as useful for excess-risk guarantees with nuisance parameters, including misspecified models and large nonparametric target classes. It also identifies specialized structure and loss assumptions as boundaries or opportunities for refinement.
- Discussion: Orthogonality supports prediction-error and excess-risk guarantees with nuisance parameters under possible model misspecification and large nonparametric target classes.The paper presents this as a systematic study of prediction error and excess risk in this setting.
- Discussion: The appendix reports experiments on orthogonal risk-minimization methods and supplies supplemental theory, applications, and proofs.The appendices include experimental results, theorem variants, orthogonal-loss constructions, applications, and proofs of the main results.
A Experiments
The experiments evaluate orthogonal-loss methods for CATE estimation and policy learning on synthetic data spanning varied nuisance, treatment-effect, sample-size, dimension, and noise settings. Orthogonal methods generally match or outperform non-orthogonal losses, approach oracle performance in some CATE settings, and produce policies comparable to the oracle.
- Experimental setup: Experiments use synthetic data with confounders, binary treatment, and outcomes generated from propensity, base-response, treatment-effect, and Gaussian-noise components.They vary sample size n ∈ {500, 1000, 3000}, confounder dimension d ∈ {6, 12}, and noise scale σ ∈ {.5, 1, 2}, running 100 experiments per setting.
- Experimental setup: Six data-generating setups capture challenges including complex propensity functions, complex base responses, constant or discontinuous treatment effects, and smooth treatment effects.Setups A–F vary these properties to stress CATE estimation and policy learning.
- CATE estimation: Orthogonal-loss methods significantly outperform or match the non-orthogonal loss for CATE estimation in most domains.Their advantage disappears with small samples and high noise, potentially because nuisance-function errors become large.
- CATE estimation: For relatively small samples, orthogonal CATE methods perform comparably to the oracle using the true nuisance models.Sample splitting provides a substantial performance boost, while cross-fitting generally adds only a small improvement.
- Policy learning: In policy learning, orthogonal losses are comparable to non-orthogonal losses in most settings and significantly outperform them in setups E and F.Learned policies perform comparably to the oracle policy.
B Additional Algorithms
The paper extends its two-stage learning framework with cross-fitting variants and more general second-stage procedures. These variants support plug-in empirical risk minimization and broader algorithms while retaining second-order nuisance dependence and oracle-rate guarantees.
- Cross-fitting variants: Meta-Algorithms 2 and 3 replace pure sample splitting with K-fold cross-fitting, corresponding to DML1- and DML2-style procedures.Each fold estimates nuisance functions on the complementary data and either estimates a target per fold or aggregates across folds.
- Plug-in ERM: Meta-Algorithm 1 uses plug-in ERM with centered second-moment penalization after splitting the sample into three subsets.One subset estimates the nuisance, another fits the target with a variance penalty, and the third estimates the minimum empirical risk.
- General second-stage algorithms: The framework permits arbitrary target-estimation algorithms that achieve an average plug-in excess-risk guarantee.This separates the meta-algorithm’s analysis from the particular second-stage learner.
- General second-stage algorithms: For plug-in ERM, localized Rademacher analysis yields an average excess-risk bound of order δ_n/K · ∥bθ − θ⋆∥_L2(ℓ2,D) + δ_n^2/K.Combined with the main theorems, this gives oracle excess-risk bounds with second-order nuisance dependence.
C Orthogonal Statistical Learning: User-Friendly Tools
The main theorems provide excess-risk guarantees for a sample-splitting meta-algorithm under Neyman orthogonality and generic nuisance and target estimators. Supporting lemmas relate oracle risk, plug-in risk, and estimation errors under self-bounding and derivative conditions.
- Main guarantees: Theorem 1 and Theorem 2 give excess-risk bounds for the sample-splitting meta-algorithm with generic nuisance and target estimators under Neyman orthogonality.The guarantees rely on abstract assumptions rather than a specific estimation algorithm.
- Supporting consequences: Lemma 1 derives a sample-splitting estimate bound from target-estimator self-bounding functions ε_n and α_n, including nuisance-rate terms.The result applies to plug-in empirical risk minimization and related estimators satisfying the stated property.
- Supporting consequences: Lemma 2 bounds the first derivative of the oracle risk at the target optimum by target and nuisance estimation-rate terms.The bound has the form Rate_D(Θ, S2, δ/2; bθ, bg) + C·Rate_D(G, S1, δ/2).
- Supporting consequences: Lemma 3 uses orthogonality and second-order Taylor expansions to upper-bound plug-in excess risk through target estimation error and nuisance estimation error.This intermediate result is used to analyze second-stage procedures such as Skeleton Aggregation.
D Construction of Orthogonal Losses
The paper constructs orthogonal losses by adding debiasing corrections to initial losses whose risks are not orthogonal. The construction introduces auxiliary nuisance functions, supports Riesz-representer formulations, and yields concrete treatment-effect and strategic-competition examples.
- General construction: A central problem is how to modify a non-orthogonal loss so that the main orthogonal-learning theorems apply.The paper addresses this by constructing a new loss from the initial loss and additional nuisance components.
- General construction: The generic construction augments the nuisance parameter with a function a0 satisfying a derivative-representation condition, then adds ⟨a(w), u − g(w)⟩·θ(x) to the loss.This correction is designed to make the resulting population risk orthogonal.
- Verification: Under the stated moment and derivative conditions, the constructed population risk satisfies the paper’s first-order and orthogonality assumptions.The proof uses conditional moment restrictions and the law of total expectation.
- Automatic debiasing: The construction’s initial target estimator need not exploit orthogonality because its error enters the final bound through estimation of a0 and therefore has higher-order impact.The paper notes that details of this argument are beyond its scope.
- General construction: The added debiasing correction reduces the impact of nuisance-estimation errors on the loss gradient.The construction can be understood through Riesz representers for linear operators acting on nuisance functions.
- Concrete examples: For strategic-competition utility estimation, a0 can be a known function of θ0 and g0, so preliminary estimates can replace a separate regression for a0.The resulting construction applies to utility functions in models of incomplete information.
- Concrete examples: In treatment-effect estimation, the generic correction produces an orthogonal loss equivalent to the squared residual-on-residual form used in the introduction.The example uses a Riesz representer a0(w) and shows equivalence between two loss formulations.
E.1 Fast Rates
The section develops sufficient conditions for fast oracle rates across a broad class of single-index losses. These conditions connect target and nuisance norms, orthogonality, curvature, and algorithm-specific rates.
- Loss structure: The population loss is expressed as an expectation of a point-wise loss applied to nuisance and target predictions.The single-index form is ℓ(θ(x), g(w); z) = Φ(⟨Λ(g(w), v), θ(x)⟩, g(w), z).
- Assumptions: Fast-rate conditions trade off the nuisance distance exponent p against regularity requirements governed by its dual exponent q.Larger p strengthens the nuisance metric but weakens the corresponding regularity condition.
- Assumptions: Assumption 9 supplies curvature, smoothness, and related conditions that imply the assumptions of the main oracle-rate theorem.Lemma 5 establishes the implication, and combining it with Theorem 1 yields an oracle excess-risk bound.
- Oracle rates: For square loss, the nuisance estimation requirement is especially favorable because Tsi = τsi = 1.When relevant problem-dependent parameters remain fixed and nuisance estimation is sufficiently fast, their asymptotic impact is negligible.
- Oracle rates: Under universal orthogonality, the nuisance contribution enters the slow-rate bound quadratically through the nuisance estimation rate.Corollary 2 gives RateD(Θ, S2, δ/2; bθ, bg) + βsi · (RateD(G, S1, δ/2))2.
- Misspecification: The misspecified square-loss analysis permits more nuisance metric-entropy complexity when the target class is large, while parametric targets require p1 < 2 for an n−1/4 first-stage rate.For p2 = 5, the condition changes from p1 < 12 in the well-specified case to p1 < 18 under misspecification.
F.2 Oracle Rates for Generic Lipschitz Losses
This section extends oracle-rate guarantees to generic Lipschitz losses, including settings without strong convexity. It identifies entropy conditions under which nuisance estimation does not prevent minimax rates and illustrates the framework across policy-learning models.
- Oracle-rate target: The generic-loss analysis targets the minimax rate attainable when nuisance parameters are absent.The main theorem applies to bounded Lipschitz losses with Lipschitz gradients in the nuisance argument.
- Oracle rates: When the nuisance metric-entropy parameter is not too large relative to the target parameter, suitable first- and second-stage algorithms achieve the oracle rate.The proof combines Theorem 2 with bounds for plug-in ERM and skeleton aggregation.
- Assumptions: Unlike the square-loss results, the generic Lipschitz theorem does not require strong convexity or moment comparison for the nuisance class.It does require the additional universal orthogonality condition.
- Oracle rates: For parametric target classes, taking p1 < 2 suffices for the standard first-stage n−1/4 rate.This matches the corresponding condition in the well-specified and misspecified square-loss settings.
- Applications: The framework covers multiple finite treatments and continuous-treatment policy learning through unbiased or randomized-policy loss formulations.The multiple-treatment population risk satisfies universal orthogonality, while continuous treatments can be represented using randomly perturbed policies.
- Applications: In counterfactual risk minimization, an unknown propensity is a nuisance parameter, but the standard inverse-propensity loss is not orthogonal to it.The framework therefore motivates orthogonalizing the population risk when propensity scores are estimated.
G.2 Domain Adaptation and Sample Bias Correction
The framework is applied to covariate-shift domain adaptation and missing-data regression, where density or missingness mechanisms act as nuisance parameters. Orthogonality preserves prediction guarantees while allowing nuisance estimation errors to contribute at higher order.
- Domain adaptation: In covariate shift, the target risk is represented as a source-distribution weighted loss using an unknown density ratio.The goal is prediction under a target covariate distribution when both source and target densities are unknown.
- Domain adaptation: Treating the density ratio as a nuisance makes the covariate-shift loss orthogonal.The mixed derivative vanishes because the conditional target-loss gradient has mean zero given covariates.
- Domain adaptation: For square loss under a lower bound g(x) ≥ η > 0, the general fast-rate conditions apply to domain adaptation.The resulting dependence on η−1 is asymptotically negligible when nuisance estimation is sufficiently fast.
- Domain adaptation: If the nuisance rate is o(n−1/8), the dominant excess-risk term attains the stated target-class rate.Variance-penalized ERM can keep η−1 out of leading terms for VC classes under the stated variance and capacity conditions.
- Missing data: In missing-data regression, the learner models both the missingness propensity and an additional outcome-related nuisance when W differs from X.The extra nuisance h is unnecessary when W = X, where h0 = 0.
- Missing data: With the true nuisance plugged in, excess risk precisely corresponds to prediction accuracy, and the model satisfies the required orthogonality conditions.Proposition 4 verifies orthogonality for the target, propensity, and outcome-related nuisance directions.
- High-dimensional prediction: The high-dimensional examples obtain prediction guarantees without restricted eigenvalue assumptions, including optimal rates for hard-sparsity-constrained ERM.This contrasts prediction analysis with parameter recovery, for which restricted-eigenvalue-type conditions are required.
H.2.1 Proof of Theorem 8
The proof establishes a variance-sensitive excess-risk bound by controlling local Rademacher complexity for a plug-in loss class. It combines concentration, covering-number, and critical-radius arguments to obtain the final guarantee.
- Setup: The proof studies the class F = {z ↦ ℓ(θ(x), bg(w); z)} and uses boundedness and Lipschitz assumptions to control its complexity.The analysis focuses on the plug-in nuisance estimate bg and assumes ∥f∥∞ ≤ 1 in the normalized setting.
- Critical radius: The critical radius is bounded by combining the local complexity inequality with the capacity-function comparison and nuisance estimation error.The argument separates the cases δ > ε0 and uses rescaling for general Lipschitz and boundedness constants.
- Initial bound: Theorem 4 supplies the initial variance-penalized ERM excess-risk inequality that the proof refines through local complexity bounds.The subsequent argument relates the local Rademacher complexity at bg to a capacity function defined at g0.
- Covering argument: The proof controls the plug-in class by comparing its empirical L2 covering number with a Hamming-error cover of the target class.The construction replicates sample points so Hamming error bounds the corresponding empirical L2 error.
- Complexity control: Haussler’s VC covering bound and symmetrization yield a local Rademacher complexity control for the centered class F − f⋆.The proof then combines concentration and Markov inequalities with union bounds to obtain high-probability control.
- Orthogonality argument: Taylor expansions and universal orthogonality cancel the first-order nuisance terms, leaving a higher-order contribution in the excess-risk analysis.The proof relates mixed derivatives at the estimated and true target predictors before applying smoothness bounds.
J.2 Proofs for Examples
The appendix verifies the assumptions behind the paper’s examples by establishing orthogonality, regularity, and complexity conditions, then invokes the main theorems to obtain guarantees.
- Example proofs: Orthogonality is established by conditioning residual products, yielding zero directional derivatives at the target parameter.The proof reduces the relevant conditional expectation to E[ε1 · ε2 | X = x] = 0 and concludes DθLD(θ0, g0)[θ − θ0] = 0.
- Example 1: For the RKHS example, approximation error scales as O(n^−(1−2α)/(p+(1−2α))) when c ∝ n^(α/(p+(1−2α))).This approximation is then combined with an oracle excess-risk analysis against the constrained optimizer θ⋆.
- Example 1: Choosing c = n^(α/(p+(1−2α))) gives excess risk O(n^−(1−2α)/(p+(1−2α))).The bound combines the approximation result with the fixed-c oracle excess-risk bound.
- Doubly robust example: The doubly robust loss satisfies orthogonality because E[∇γφ(f0, e0; z) | X] = 0 by double robustness.The proof also verifies regularity under bounded outcomes, bounded predictors, and overlap η ≤ e(X) ≤ 1 − η.
- Technical lemmas: The technical appendix develops constrained M-estimation tools using Lipschitz losses, bounded vector-valued classes, peeling, concentration, and vector-valued contraction.These lemmas control empirical-process deviations and establish exponential probability bounds for constrained estimators.
K.1 Proofs of Lemmas for Constrained M-Estimators
These proofs establish concentration and complexity controls for constrained empirical risk minimization, then apply them to the two-stage plug-in ERM analysis.
- Lemma proofs: The constrained M-estimator analysis bounds empirical-process deviations for Lipschitz losses over bounded vector-valued function classes.The argument uses local Rademacher complexity, vector-valued contraction, and concentration inequalities.
- Lemma proofs: A peeling argument partitions functions by L2 distance from f⋆ and applies localized bounds on each radius shell.The shell-wise probabilities are combined with a union bound over at most log(2/δn) shells.
- Lemma proofs: For linear losses, the proof uses a star-hull reduction that avoids requiring a lower bound on δn^2.The reduction rescales a violating function to the boundary radius δn while preserving the relevant linear comparison.
- Two-stage analysis: The two-stage plug-in ERM uses sample splitting, with nuisance estimation on S1 and target estimation on S2.The nuisance-dependent excess-risk terms enter through the generic theorem’s εn and αn quantities.
L.2 Proof of Theorem 4
The section develops moment-penalized plug-in ERM and complexity tools for deriving rates under parametric and nonparametric metric-entropy assumptions.
- Theorem 4: Moment-penalized plug-in ERM centers losses using a preliminary estimate of the optimal risk value.The construction applies a second-moment bound to the centered loss class.
- Theorem 4: A vanishing preliminary-risk error affects the final regret only at second order.The proof writes the final regret bound in terms of εn and uses AM-GM to control the resulting terms.
- Theorem 4: The preliminary risk estimate is controlled by two-sided uniform convergence over the loss class using vanilla Rademacher complexity.Lipschitz dependence on the nuisance estimate relates plug-in and true population risks.
- Complexity conditions: Parametric and nonparametric cases are distinguished for both nuisance and target classes through metric-entropy conditions.The nuisance and target classes receive separate dimension parameters d1, d2 or entropy exponents p1, p2.
M.3 Overview of Proofs
The proof overview reduces both learning stages to standard regression procedures, then uses orthogonality to preserve target-stage rates when nuisance estimates are plugged in.
- Overview: The paper’s strategy is to use out-of-the-box learning algorithms for nuisance and target estimation, while analyzing their interaction through orthogonality.Optimal algorithm choice depends on the complexity of the nuisance and target classes.
- First stage: First-stage rates are provided for Global ERM and Skeleton Aggregation under parametric and nonparametric metric-entropy assumptions.The informal proposition includes a nonparametric rate involving K1 n^−2/(2+p1), with logarithmic factors suppressed.
- First stage: Skeleton Aggregation is optimal for all nuisance entropy exponents p1, whereas Global ERM is optimal only when p1 ≤ 2.The stated minimax rate is Ω(n^−2/(2+p1)).
- Second stage: The second stage maps the target problem to square-loss regression using auxiliary covariates and outcomes generated from the nuisance estimate.Standard algorithms can then be applied to the resulting auxiliary dataset.
- Second stage: In the well-specified setting, orthogonality controls plug-in misspecification so the target rate is close to the rate obtained with the true nuisance parameter.In the misspecified setting, the analysis instead uses a worst-case nuisance-dependent target rate.
- Overview: The final excess-risk guarantee combines the target-stage rate with a squared nuisance-stage rate.The bound is LD(bθ, g0) − LD(θ⋆, g0) ≤ C · RateD(Θ, S2, δ; bθ, bg) + C′(RateD(G, S1, δ))^2.