Source-linked AI summary

Transfer Learning for High-dimensional Linear Regression: Prediction, Estimation, and Minimax Optimality

Sai Li, T. Tony Cai, Hongzhe Li

arXiv:2006.10593v1stat.MEstat.ML

TL;DR

The paper addresses limited methods and theory for high-dimensional linear regression transfer learning. It develops procedures based on auxiliary-target similarity and shows optimality, improved gene-expression prediction, and robustness to non-informative auxiliary samples.

  • Problem

    High-dimensional linear-regression transfer learning lacks sufficient methods and fundamental theoretical results, including rate-optimal estimation and prediction when informative auxiliary samples are known.

  • Method

    The paper characterizes informative auxiliary studies by sparse contrasts with the target model and develops Oracle Trans-Lasso and data-driven Trans-Lasso procedures for estimation and prediction.

  • Results

    The proposed Oracle Trans-Lasso is minimax optimal over a range of parameter spaces, while Trans-Lasso is robust to non-informative auxiliary samples and achieves a 17% average prediction-error gain over Lasso for JAM2 expression.

  • Takeaways & Limitations

    Informative auxiliary samples can improve target estimation, prediction, and gene-expression prediction across tissues, while Trans-Lasso limits harm from adversarial auxiliary samples.

  • Takeaways & Limitations

    When auxiliary sample size is sufficiently large, adding more auxiliary samples may no longer improve the convergence rate, making protection against adversarial samples more important.

Abstract

from arXiv · show

This paper considers the estimation and prediction of a high-dimensional linear regression in the setting of transfer learning, using samples from the target model as well as auxiliary samples from different but possibly related regression models. When the set of "informative" auxiliary samples is known, an estimator and a predictor are proposed and their optimality is established. The optimal rates of convergence for prediction and estimation are faster than the corresponding rates without using the auxiliary samples. This implies that knowledge from the informative auxiliary samples can be transferred to improve the learning performance of the target problem. In the case that the set of informative auxiliary samples is unknown, we propose a data-driven procedure for transfer learning, called Trans-Lasso, and reveal its robustness to non-informative auxiliary samples and its efficiency in knowledge transfer. The proposed procedures are demonstrated in numerical studies and are applied to a dataset concerning the associations among gene expressions. It is shown that Trans-Lasso leads to improved performance in gene expression prediction in a target tissue by incorporating the data from multiple different tissues as auxiliary samples.

1. Introduction

The paper studies high-dimensional linear regression transfer learning, where auxiliary studies can improve target estimation and prediction when they are sufficiently similar. It develops procedures for known and unknown informative sets, with robustness to non-informative auxiliary samples.

  • Motivation: High-dimensional regression often has more covariates than observations, motivating transfer learning with additional auxiliary studies.The target coefficient is sparse, while auxiliary studies provide related regression data.
  • Informative auxiliary samples: Informative auxiliary studies are defined by contrast vectors between auxiliary and target coefficients having sufficiently low sparsity.For q = 0, this means at most h nonzero contrast coordinates.
  • Research gap: The literature lacks sufficient methods and fundamental theory for high-dimensional linear-regression transfer learning, especially when the informative set is unknown.Known-set methods also lacked rate-optimal estimation and prediction procedures.
  • Contributions: Oracle Trans-Lasso achieves minimax-optimal estimation and prediction when the informative set is known under mild conditions.Its convergence rate is faster when the informative set is nonempty and contrast sparsity is sufficiently smaller than target sparsity.
  • Contributions: Trans-Lasso adapts to an unknown informative set by aggregating candidate estimators and remains not much worse than Lasso despite adversarial auxiliary samples.Under proper conditions, it transfers knowledge from informative subsets and can perform as if the informative set were known.
  • Applications and validation: The proposed algorithms are theoretically and numerically justified under heterogeneous design distributions.The paper also develops applications involving biological and gene-expression data.

2. Estimation with Known Informative Auxiliary Samples

With the informative auxiliary set known, Oracle Trans-Lasso combines auxiliary and primary data, then corrects the auxiliary-induced bias using primary data. Under stated design and noise conditions, it is minimax optimal and can improve rates over target-only Lasso.

  • Algorithm: Oracle Trans-Lasso first estimates a coefficient using primary and informative auxiliary samples, then corrects its bias with primary data.The initial estimator is biased because auxiliary coefficients generally differ from the target coefficient.
  • Conditions: The analysis assumes Gaussian, identically distributed designs for the target and informative auxiliary samples, with bounded covariance eigenvalues and sub-Gaussian noise.Non-informative auxiliary samples require no assumptions because Oracle Trans-Lasso does not use them.
  • Convergence rates: When the informative set is empty, Oracle Trans-Lasso has the target-only Lasso rate O_P(s log p/n0).With informative samples, the rate is sharper when their contrasts are sufficiently sparse and the combined sample size is substantially larger.
  • Theoretical analysis: The proof exploits a decomposition of the auxiliary-limit coefficient into sparse components and establishes restricted-eigenvalue control for both algorithmic steps.The tuning parameters depend on response moments and noise levels, with cross-validation available in practice.
  • Minimax optimality: Oracle Trans-Lasso is minimax rate optimal over the considered parameter space under the theorem’s conditions.The lower bound includes the ideal rate s log p/(nA+n0) and the unhelpful-auxiliary rate h(log p/n0)^1/2.

3. Unknown Set of Informative Auxiliary Samples

When the informative auxiliary set is unknown, Trans-Lasso constructs candidate estimators and aggregates them, aiming to retain useful transfer while remaining robust to uninformative samples. Its candidate sets exploit estimated sparsity differences between informative and non-informative studies.

  • 3.1. The Trans-Lasso Algorithm: Trans-Lasso uses sample splitting, candidate estimators, and aggregation to handle an unknown informative auxiliary set.Candidate estimators are built independently from the aggregation sample so Q-aggregation can compare them under its theoretical guarantees.
  • 3.1. The Trans-Lasso Algorithm: The candidate dictionary includes the original Lasso using only primary samples, providing a baseline that protects against harmful auxiliary studies.Aggregation guarantees performance not much worse than this baseline under the stated conditions.
  • 3.3. Theoretical Properties of Trans-Lasso: The conditions used to establish transfer gains may require assumptions that are not guaranteed in practice, whereas robustness to non-informative samples does not require those conditions.The paper also notes that increasing auxiliary sample size eventually stops improving the convergence rate once the relevant first term no longer dominates.
  • 3.2. Constructing the Candidate Sets for Aggregation: The naive construction has 2^K candidates, so the proposed procedure seeks far fewer candidates to reduce aggregation cost while preserving the desired selection property.The paper notes that exhaustive enumeration can be computationally burdensome and that aggregation cost can be of order K/n0.
  • 3.2. Constructing the Candidate Sets for Aggregation: Candidate sets are formed by ranking auxiliary studies according to estimated sparsity indices and progressively including those with the smallest values.This construction targets studies whose contrast vectors are sparser, as expected for informative auxiliaries.
  • 3.2. Constructing the Candidate Sets for Aggregation: When all or none of the auxiliary samples are informative, the candidate collection contains the true informative set and Trans-Lasso is not much worse than Oracle Trans-Lasso.Mixed cases with 0 < |A| < K are identified as more challenging.
  • 3.3. Theoretical Properties of Trans-Lasso: The aggregation result makes Trans-Lasso robust to other candidates: its performance depends on the best candidate rather than on poorly performing candidates.The estimation guarantee requires a sufficiently small constructed dictionary, such as K ≤ c n0, while a slower bound is available for arbitrarily large dictionaries.
  • 3.3. Theoretical Properties of Trans-Lasso: Under separation conditions, estimated sparsity indices distinguish selected informative studies from non-informative ones, enabling adaptation to the unknown set.When the relevant conditions hold, Trans-Lasso errors are comparable to those of the known-set Oracle Trans-Lasso.

4. Extensions to Heterogeneous Designs and ℓq-sparse Contrasts

The paper extends Trans-Lasso to heterogeneous designs and to contrasts characterized by ℓq-sparsity, including exact sparsity. It establishes convergence and minimax results while identifying design similarity and auxiliary-sample size as key conditions for transfer gains.

  • 4.1. Heterogeneous Designs: Heterogeneous-design extensions allow Trans-Lasso to handle auxiliary and target samples with different covariance structures.The extension introduces a relaxed condition and characterizes covariance differences through CΣ.
  • 4.1. Heterogeneous Designs: When nA0 ≫ n0 and CΣh(log p/n0)^(1/2) ≪ s, informative auxiliary samples yield a sharper estimation rate than using primary data alone.The gain is favored by small covariance-discrepancy parameter CΣ and sparse contrasts.
  • 4.1. Heterogeneous Designs: The heterogeneous-design Oracle Trans-Lasso has an upper-bound guarantee under the stated theorem conditions, with numerical experiments used to study this setting.The guarantee is stated in Corollary 1 and follows assumptions including Gaussian rows and bounded covariance eigenvalues.
  • 4.2. ℓq-sparse Contrasts: The framework extends from ℓ1-sparse contrasts to ℓq-norm constraints for q ∈[0, 1), with exact sparse contrasts treated through a dedicated Oracle Trans-Lasso algorithm.The q = 0 procedure is presented for a known informative set, while q ∈(0, 1) results are deferred to the Appendix.
  • 4.2. ℓq-sparse Contrasts: For q = 0, the proposed estimator has a convergence-rate result and a matching minimax lower-bound argument under the relevant conditions.Theorem 6 establishes minimax optimality of the q = 0 estimator, while unknown informative sets can be handled by replacing the oracle procedure in Trans-Lasso.
  • 4.2. ℓq-sparse Contrasts: The q = 0 setting is less practical because ℓ0-sparse contrasts require coefficient equality in most coordinates and can be altered by standardization.These considerations motivate the paper’s focus on ℓ1-sparse contrasts for applications; the q = 0 procedure also has higher computational cost when the informative set is large and relies heavily on homogeneous designs.

5. Simulation Studies

Simulation studies compare Lasso, Naive Trans-Lasso, Oracle Trans-Lasso, and Trans-Lasso across covariance structures and auxiliary-sample configurations. Trans-Lasso methods generally improve as more informative samples are added, while unknown-set adaptation can fail in harder cases.

  • Experimental setup: The study compares Lasso, Oracle Trans-Lasso, Trans-Lasso, and a naive method that treats every auxiliary sample as informative.The experiments use p = 500, n0 = 150, n1,...,nK = 100, and K = 20.
  • Identity covariance designs: Estimation errors for all three Trans-Lasso-based methods decrease as the number of informative auxiliary samples increases, whereas Lasso performance is unchanged.Errors increase with the contrast sparsity parameter h.
  • Identity covariance designs: Naive Trans-Lasso performs worse than Lasso when the informative set is relatively small, showing that it cannot uniformly adapt to an unknown informative set.Oracle Trans-Lasso and Trans-Lasso are comparable in most settings.
  • Homogeneous designs: Under homogeneous covariance designs, Trans-Lasso and Oracle Trans-Lasso have reliable performance when informative and non-informative studies have different covariance structures.The proposed sparsity index consistently separates informative auxiliary samples in both configurations.
  • Interpretation: The definition of the informative set need not be the subset that yields the smallest estimation errors, which can make Trans-Lasso slightly outperform its oracle comparator.Cross-fitting also means the compared procedures use different empirical samples.
  • Heterogeneous designs: Trans-Lasso remains comparable to Oracle Trans-Lasso in most heterogeneous-design settings, but a larger gap appears when h = 12 in configuration (i).In that case, the estimated sparsity index separates only a subset of informative studies from the others.

6. Application to Genotype-Tissue Expression Data

The application evaluates transfer learning for gene-expression regression across GTEx tissues, using target-tissue samples and other tissues as auxiliary data. Trans-Lasso generally improves prediction over Lasso and is more reliable than naive transfer when auxiliary tissues differ from the target.

  • Data and targets: The analysis uses GTEx gene-expression data from 49 tissues and focuses on JAM2 and related genes across multiple target tissues.For JAM2, 47 tissues with more than 120 measurements are used, with 1,079 final covariates.
  • Evaluation: Prediction is evaluated by repeated 5-fold cross-validation comparing Lasso, Naive Trans-Lasso, and Trans-Lasso.Trans-Lasso constructs candidate estimators and aggregates them using separate portions of the training data.
  • JAM2 prediction: Trans-Lasso achieves the smallest JAM2 prediction errors across tissues, with an average gain of 17% relative to Lasso.Its improvement is especially pronounced in Amygdala and Nucleus accumbens basal ganglia.
  • JAM2 prediction: In Pituitary, both transfer methods provide only mild improvement, consistent with a relatively distinct target regression model.Trans-Lasso nevertheless remains the best-performing method in the reported tissue comparisons.
  • Additional genes: Across 25 additional Chromosome 21 genes, Trans-Lasso has the best overall performance among target tissues.Improvements are significant in Cerebellar Hemisphere, Cortex, and Frontal Cortex.
  • Additional genes: Naive Trans-Lasso is comparable to Lasso in most additional-gene cases, indicating weak overall similarity between auxiliary and target tissues.This contrasts with the more selective gains obtained by Trans-Lasso.

7. Discussion

The discussion summarizes minimax-optimal transfer learning for known informative sets and adaptive aggregation for unknown sets. It also identifies statistical inference and alternative similarity measures as directions for further work.

  • Main findings: The paper characterizes auxiliary-study similarity through sparsity of the contrast vector and develops estimation and prediction procedures for high-dimensional linear regression.Theoretical, numerical, and real-data results support the proposed transfer-learning framework.
  • Main findings: When the informative set is known, Oracle Trans-Lasso is minimax optimal over a range of parameter spaces and can achieve faster convergence rates.The faster rates occur when the informative set is non-empty and h is sufficiently smaller than s.
  • Adaptation: When the informative set is unknown, adaptation is achieved by aggregating a collection of candidate estimators.The discussion reports numerical and real-data support for this result.
  • Future directions: Future work includes confidence intervals, hypothesis testing, and minimax-optimal inference under transfer learning.The discussion also suggests studying alternative measures of similarity between target and auxiliary models.

A. Proofs in Section 2

The appendix proves the Section 2 results through oracle inequalities, restricted-eigenvalue arguments, concentration conditions, and lower-bound constructions. The proof fragments also cover cases with uninformative auxiliary samples and varying sparsity regimes.

  • Upper bounds: The upper-bound proofs derive oracle inequalities under restricted eigenvalue conditions and sample-size assumptions.The arguments combine lemmas controlling estimator errors with high-probability events.
  • Upper bounds: Lemma 3 controls the auxiliary-estimator error through bounds on both its ℓ2 and ℓ1 norms.The stated orders depend on sparsity, tuning parameters, and a covariance-related quantity.
  • Probability control: The proofs establish the required high-probability events using restricted eigenvalue conditions and sub-Gaussian errors.These events support the desired estimation and prediction bounds.
  • Lower bounds: Lower bounds are constructed separately across sparsity regimes, including cases with sparse coefficients and small contrast radius h.The proof considers classes with prescribed support sizes and coefficient magnitudes.
  • Lower bounds: The lower-bound argument explicitly includes settings where all auxiliary samples contain no information about the target parameter.It also treats known informative sets through bounds involving n_Aq + n0.

B.1. Proof of Lemma 1 and Remark 1

The proof decomposes estimation and prediction errors using dictionary and covariance constructions, then applies Gaussian-design arguments and singular-value decomposition. These steps establish the stated error bounds and support screening results for auxiliary studies.

  • Error decomposition: The proof represents candidate regression vectors through a dictionary whose columns include estimated target-related coefficients.The dictionary is formed with columns such as ˆβ( bG_l).
  • Error decomposition: The target covariance term is written using the empirical covariance matrix of the primary design restricted to an index set.The proof defines bΣ(0,c) from the restricted primary design matrix.
  • Spectral argument: A singular value decomposition bB = UΛV^⊺ is used to analyze the dictionary-dependent term.Λ contains the singular values of bB.
  • Error decomposition: The estimation error is bounded by separating dictionary-estimation error from approximation error relative to the target coefficient.The displayed inequality bounds the squared error by two terms involving bB(ˆθ−θ∗) and bBˆθ∗−β.
  • Gaussian-design argument: Gaussian assumptions on the primary design allow conditioning arguments based on independence between U and X(0).The proof explicitly uses the Gaussian property and independence of U from the primary design.
  • Screening argument: The screening proof controls auxiliary-study statistics uniformly using sub-exponential bounds, trimming arguments, and a union bound over non-informative studies.The argument invokes Gaussian tails, sub-exponential behavior, and logarithmic control of |A^c|.

C.1. Proof of Theorem 5

The proof of Theorem 5 establishes bounds for transfer learning with informative auxiliary samples by combining effective sample-size control with restricted-eigenvalue and concentration arguments. It treats estimation and prediction errors under explicit sample-size and sparsity conditions.

  • Rate decomposition: The proof separates the baseline rate s log p/n0 from additional transfer-learning terms when informative auxiliary samples are available.The baseline rate is the convergence rate when A0 is empty, while nonempty A0 permits use of auxiliary samples.
  • High-probability control: Under Theorem 5 conditions, Lemma 4 supplies a high-probability bound uniformly for every informative auxiliary study.The probability is at least 1−exp(−c1 log p)−exp(−c2nA0).
  • Design conditions: The proof combines Gaussian design results with sub-Gaussian noise and sub-exponential empirical-covariance properties.These ingredients are used to verify the relevant high-probability event.
  • Rate decomposition: The informative auxiliary and primary samples contribute through the combined effective sample size nA0 + n0.The proof uses bounds involving log p/(nA0 + n0).
  • High-probability control: The proof requires the estimation error to be controlled in a regime where s log p/(nA0 + n0) = o(1).This condition is stated before the estimation-error argument.
  • Design conditions: Restricted-eigenvalue control is guaranteed under a sample-size condition linking contrast sparsity h, dimension p, and primary and auxiliary sample sizes.The stated condition is h log p/n0 = o((log p/(n0 + nA0))1/4).

C.2. Minimax optimal rates for q ∈(0, 1)

This section establishes minimax lower and achievable upper bounds for estimation with ℓq-sparse contrasts when q ∈(0, 1). The proof uses a transfer-learning algorithm and derives rates involving both target sparsity and auxiliary-study information.

  • Minimax lower bound: The section first proves a minimax lower bound for estimation over the parameter class Θq(s,h).The lower-bound theorem assumes Conditions 1 and 2 and a small-error regime.
  • Minimax lower bound: The lower-bound result applies for fixed q ∈(0, 1) under a condition involving hq(log p/n0)1/2−q/4 and s log p/(nA0 + n0).Both terms are required to be o(1).
  • Achievability: The upper-bound argument begins by specifying an algorithm that takes primary data and informative auxiliary samples as input.The algorithm input is listed as (X(0), y(0)) together with informative auxiliary samples.
  • Achievability: The achievable rate contains a term log p/(n0 + nAq), reflecting the combined primary and informative-auxiliary sample sizes.The displayed rate is stated with a sufficiently large constant c1.
  • Achievability: Theorem B establishes achievability of the upper bound under Conditions 1 and 2 and its associated regularity assumptions.The theorem is explicitly labeled as achievability for q ∈(0, 1).
  • Proof mechanism: The proof controls contrast-estimation errors using ℓq and ℓ2 bounds before solving a resulting quadratic constraint.The argument invokes a bound on ∥ˆv(k)∥q and then selects the positive root.

D. More results on simulation

The simulation section evaluates screening and transfer-learning performance across settings, while the accompanying tables summarize auxiliary-study and sample-size configurations. Results indicate reliable performance and decreasing estimation error as more auxiliary studies are used.

  • Screening evaluation: A screening proportion close to 1 is favorable because it indicates separation of informative auxiliary studies from the others.The empirical probability is defined through the first |A| smallest bR(k) values.
  • Screening evaluation: The simulations exclude the trivial cases A = ∅ and A = {1, . . . , K}.Only nontrivial informative-set configurations are considered in Table 1.
  • Screening evaluation: Table 1 reports the proportion of simulations in which informative-study scores are no larger than non-informative-study scores.The reported event is maxk∈A bR(k) ≤ maxk∈Ac bR(k).
  • Gene-expression application: Table 2 lists analyzed genes together with the number of auxiliary studies and average primary and auxiliary sample sizes.Its columns are Aux. studies, Avg. pri. ss, and Avg. aux. ss.

E. More results on data application

The data application reports gene analyses across multiple target tissues, including detailed comparisons of prediction error and overall performance for several transfer-learning methods.

  • 137 genes were analyzed across multiple target tissues.
  • Figure 6 compares average prediction error for Lasso, Naive-Trans-Lasso, and Trans-Lasso on gene JAM2 across multiple tissues.Errors were evaluated using 5-fold cross-validation.
  • Figure 7 summarizes overall prediction performance for Lasso, Naive-Trans-Lasso, and Trans-Lasso across 25 genes on Chromosome 21 and in Module.
Loading 2006.10593v1…