Source-linked AI summary

A knockoff filter for high-dimensional selective inference

Rina Foygel Barber, Emmanuel J. Candes

arXiv:1602.03574v3stat.MEmath.ST

TL;DR

High-dimensional studies may have more features than observations, with correlated features and non-identifiability complicating reliable association discovery. The paper combines feature screening with knockoff-based inference and proves finite-sample directional FDR control without assumptions on the design matrix or regression coefficients.

  • Problem

    High-dimensional association studies can involve more features than observations, correlated features, and non-identifiability, complicating reliable variable selection.

  • Method

    The paper screens features, applies a knockoff filter to the reduced model, and develops recycling to reuse screening data for greater power.

  • Results

    Finite-sample directional FDR is rigorously controlled without assumptions on the design matrix or regression coefficients.

  • Takeaways & Limitations

    The screen + knockoff procedure provides a form of reproducibility for selected variables and their effect directions through directional FDR control.

  • Takeaways & Limitations

    Reusing the full data after screening does not control FDR because screening changes the distribution of the response.

Abstract

from arXiv · show

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split into two groups, where the first group is used to screen for a set of potentially relevant variables, whereas the second is used for inference over this reduced set of variables; we also develop strategies for leveraging information from the first part of the data at the inference step for greater power. In our work, the inferential step is carried out by applying the recently introduced knockoff filter, which creates a knockoff copy-a fake variable serving as a control-for each screened variable. We prove that this procedure controls the directional false discovery rate (FDR) in the reduced model controlling for all screened variables; this says that our high-dimensional knockoff procedure 'discovers' important variables as well as the directions (signs) of their effects, in such a way that the expected proportion of wrongly chosen signs is below the user-specified level (thereby controlling a notion of Type S error averaged over the selected set). This result is non-asymptotic, and holds for any distribution of the original features and any values of the unknown regression coefficients, so that inference is not calibrated under hypothesized values of the effect sizes. We demonstrate the performance of our general and flexible approach through numerical studies, showing more power than existing alternatives. Finally, we apply our method to a genome-wide association study to find locations on the genome that are possibly associated with a continuous phenotype.

1 Introduction

The paper addresses association discovery when features outnumber observations, correlations undermine sparsity, and measured variables may be proxies rather than causal factors. It proposes screening followed by knockoff-based directional FDR control, with finite-sample validity that does not require design or coefficient assumptions.

  • Motivation: GWAS and similar studies often measure hundreds of thousands of variables with fewer observations than features.
  • Motivation: When p > n, fixed-design linear regression is not identifiable, while correlated features and approximate non-sparsity weaken individual null testing.
  • Proposed framework: The proposed two-stage procedure screens potentially relevant features using one data portion, then applies knockoffs to the reduced model using another.
  • Scope: The framework targets reliable discovery of large effects and permits proxies for unmeasured causal mutations, rather than requiring a sparse causal model.
  • Directional FDR: Directional FDR counts both selected null effects and nonzero effects assigned the wrong sign, making it relevant when many coefficients are near zero.
  • Relation to prior work: Unlike asymptotic approaches requiring sparsity, beta-min, or design assumptions, the paper emphasizes finite-sample inference after selection and directional FDR control.

2 Knockoffs

The knockoff filter augments each feature with a matched control whose exchangeability enables estimation of false discoveries. Antisymmetric statistics then identify features preferred over their knockoffs, and a threshold controls the target FDR without assumptions on the design or unknown coefficients.

  • Knockoff construction: For each feature X_j, the method constructs a knockoff copy eX_j that preserves the original correlation structure and cross-correlations.
  • Knockoff construction: Exchangeability means swapping null features with their knockoffs leaves the relevant distribution unchanged, making positive and negative evidence equally likely for nulls.
  • Selection statistics: Running OMP or Lasso on the augmented design compares original and knockoff coefficients, using selected knockoffs to estimate false positives.
  • Selection statistics: The statistic W_j must satisfy sufficiency and antisymmetry, with positive values favoring X_j and negative values favoring its knockoff.
  • Thresholding: The filter chooses a positive threshold from the target FDR level q, using the knockoff count as an overestimate of false discoveries.
  • Guarantees: Knockoff+ controls FDR, while knockoff controls a slight FDR modification, without assumptions on the design, regression coefficients, or noise level.

3 Controlling Errors of Type S

The paper extends knockoff inference to directional errors, controlling mistakes in selected effect signs without calibrating to null coefficients. Its proof uses a martingale inequality for Bernoulli variables with varying success probabilities.

  • Directional errors: The directional criterion treats selecting a nonzero effect with the wrong sign as an error, making sign reliability central to regression discoveries.The paper motivates this because researchers want to report effect directions confidently.
  • Directional FDR guarantees: Knockoff+ controls directional FDR at level q, while knockoff controls the corresponding modified directional FDR.The guarantees apply to the selected set defined by the knockoff procedure.
  • Directional errors: Unlike the original unsigned-FDR result, the new theorem does not rely on calibrating claims under βj = 0.Nonzero coefficients near zero can produce asymmetric positive and negative probabilities for Wj.
  • Contribution: The knockoff and knockoff+ guarantees improve on earlier results that controlled only unsigned FDR or modified unsigned FDR.The earlier guarantees concerned false discoveries defined by zero coefficients rather than incorrect effect directions.
  • Proof strategy: The proof reduces directional FDR control to a martingale inequality for Bernoulli variables whose success probabilities may vary.This generalizes the simpler exchangeable-sign argument, where null signs are i.i.d. Bernoulli(0.5).

4 Knockoffs in High Dimensions

Because knockoff construction is impossible when p > n, the paper screens features using one data split and applies knockoffs to the remainder. Data recycling and signed statistics recover power while preserving directional FDR control under stated screening and model conditions.

  • Feature screening and data splitting: When p > n, direct knockoff construction is impossible, motivating feature screening followed by knockoff inference on a reduced set.The reduced model contains fewer screened features than the observations available for inference.
  • Feature screening and data splitting: Data splitting uses disjoint observations for screening and selection, with directional FDR control relying on sure screening of relevant features.Including extra null features in the screened set does not create the same issue as omitting relevant variables.
  • Increasing power with data recycling: Reusing screened data through data recycling preserves FDR control while raising power substantially toward the sensitivity of full-data analysis.The recycled procedure concatenates original features from the first split with knockoffs on the second split, then runs the filter on all observations.
  • Increasing power with data recycling: Data recycling increases power because non-null Wj statistics appear earlier along the knockoff path, improving the filter’s ability to select them.The method treats the first split as fixed after screening, allowing its information to be reused without losing the guarantee.
  • Increasing power with signed statistics: Signed statistics use screening-stage coefficient signs to focus inference on effects consistent with the estimated direction, generally increasing power.This extracts sign information from the first data portion before selection on the second portion.
  • Sure screening and directional FDR control: Under Gaussian responses, the procedure controls modified directional FDR at level q, while knockoff+ controls directional FDR at level q.If the sure-screening event has probability near one, the corresponding unconditional guarantees incur only the probability of its complement.
  • Directional FDR control in the reduced model: Missing relevant variables during screening can bias null-versus-knockoff comparisons, undermining directional FDR control in the reduced model.Correlated null features may then be more likely to be selected than their knockoffs, invalidating the null-counting argument.
  • Directional FDR control in the reduced model: The Gaussian design assumption is strong, although the authors expect FDR control to extend to a broader range of models.The method does not require knowledge or estimation of the Gaussian design’s mean and covariance parameters.

5 Simulations

The simulations compare BH, selective inference for the square-root Lasso, and knockoff filtering with data splitting or recycling across AR and GWAS designs. All methods generally control FDR at the target level, while data recycling provides higher restricted power than data splitting.

  • Simulation settings: The AR design uses n = 2000 observations and p = 2500 correlated features, while the GWAS design uses SNP features filtered to pairwise correlations at most 0.5.AR correlations vary as ρ ∈ {0, 0.25, 0.5, 0.75}.
  • Methods: The simulations evaluate BH, exact selective inference for the square-root Lasso, and knockoff filtering with data splitting or data recycling.Methods are applied after screening for a low-dimensional submodel.
  • Implementation: Screening uses n0 = 750 observations and kmax = 450 features for AR, versus n0 = 1800 and kmax = 720 for GWAS.The remaining observations are used for inference after the initial Lasso screening step.
  • Knockoff implementation: The knockoff filter forms Wj = |β̂j| − |β̃j| using the square-root Lasso before applying the knockoff threshold.Data splitting uses the remaining n1 observations, whereas data recycling also reuses the first data portion as fixed information.
  • Results: All methods generally control FDR at the desired level, except least squares + BH shows overly high FDR for large ρ in the AR setting.The authors attribute this exception to differences between reduced-model and full-model FDR under highly correlated designs.
  • Results: Data recycling shows somewhat higher restricted power than the other methods and a clear gain over data splitting, while non-restricted power remains low because weak signals are difficult to distinguish from zero.Directional FDR is slightly higher than FDR across settings because weak signals can be selected with the wrong sign.

6 Real Data Experiments

The GWAS experiment applies a multi-stage screen-and-knockoff pipeline to NFBC1966 data, using repeated random rotations and held-out inference to identify LDL- and HDL-associated genomic regions.

  • Data and preprocessing: The NFBC1966 data contain 5,402 SNP arrays and phenotype measurements from subjects born in Northern Finland in 1966.The final processed design contains 328,934 SNP features.
  • Data and preprocessing: Population stratification was corrected by regressing responses and SNP features on the top five design-matrix principal components.Phenotype-specific processing included log transformations and exclusions based on medication, fasting, pregnancy, and missingness criteria.
  • Screening and inference pipeline: Random rotations enabled splitting the data into screening and inference subsets without matrix-conditioning problems, with n0 = 1900 used for screening.The procedure was repeated with a new rotation matrix for 10 total repetitions.
  • Results: Across 10 LDL trials, 44 SNPs were selected and grouped into 29 regions using a 10^6-base-pair proximity rule.Figure 1 reports regional selection frequencies and overlaps with regions identified by prior analyses.
  • Results: Across 10 HDL trials, 13 SNPs were discovered and grouped into 10 regions, with an estimated FDR of 4.29% against the Willer et al. meta-analysis.Figure 2 uses the same interpretation as Figure 1.

7 Summary

The paper presents screen + knockoff as a controlled variable-selection approach for high-dimensional regression, with recycling improving power while directional FDR remains controlled in finite samples.

  • Summary: Screen + knockoff controls directional FDR, combining Type I and Type S error control for reported variables and their effect directions.The guarantee holds without assumptions on the design matrix or regression-coefficient values and is finite-sample.
  • Summary: Recycling reuses screening data during inference rather than discarding it, improving the procedure’s power.

A.1 A general form of directional FDR control

The appendix states a general directional-FDR theorem for knockoff statistics under Gaussian responses, pairwise exchangeability, and sufficiency and antisymmetry conditions.

  • General theorem: The paper derives its specific results as special cases of this general statement and proves the theorem in Appendix B.
  • General theorem: Theorem 4 assumes y is Gaussian with mean µ and covariance Θ, while fixed X and eX satisfy pairwise exchangeability.The theorem also requires a statistic W satisfying the stated sufficiency and antisymmetry properties.
  • General theorem: Under these conditions, knockoff+ controls a directional FDR, while the knockoff method controls a modified directional FDR.
  • Proof applications: For the fixed-design model, the proof takes µ = Xβ and uses sign((Xj − eXj)ᵀµ) = sign(βj) whenever Xj ≠ eXj.The trivial case Xj = eXj yields Wj = 0, so the feature cannot be selected.
  • Proof applications: The proof identifies a sign error with an incorrect sign of βj on the screening event, using the Gram-matrix condition and sj > 0 when Xj ≠ eXj.
  • Proof applications: The reduced-model proof conditions on screened features and y(0), while the random-design proof conditions additionally on X(0) and screened X(1) features.The random-design argument replaces β with expected partial regression coefficients and increases the variance level because of missed-signal randomness.

B Additional proofs

Appendix B presents supporting lemmas before proving the general directional-FDR control theorem.

  • Proof organization: The appendix first proves Lemma 1 and additional supporting lemmas, then proves Theorem 4, the general error-control result.

B.1 Key lemmas

The key lemmas establish the conditional sign structure needed for the high-dimensional knockoff argument. They show that knockoff statistics have conditionally independent signs with a bias toward negative signs, enabling false-discovery control.

  • The proof relies on Lemma 1 for non-i.i.d. Bernoulli sequences and a lemma concerning knockoff statistics W.
  • Conditioning on V, the absolute statistics |W| and auxiliary sign vector S are fixed, while the signs of W are mutually independent.The construction uses sufficiency and antisymmetry of the knockoff statistics.
  • The covariance structure follows from pairwise exchangeability, making the relevant contrast terms mutually independent and independent of the sum terms.
  • For every coordinate, P{sign(Wj) = −1 | V} ≥ 1/2, so negative signs are at least as likely as positive signs conditionally on V.This follows from the normal distribution of (Xj − eXj)⊤y and the sign relation defining Wj.
  • The resulting sign property generalizes the original knockoff setting, where null signs were i.i.d. unbiased, to conditionally independent signs with at most a 50% chance of being positive.
  • Corollary 1 converts these conditional sign probabilities into a bound for the knockoff or knockoff+ filter over a random subset M.

B.2 Proof of Theorem 4

The proof of Theorem 4 rewrites directional sign errors in terms of knockoff statistics and applies the conditional sign bound. The same strategy handles both knockoff and knockoff+ thresholds.

  • Directional discoveries are characterized by Wj ≥ T, allowing the modified directional FDR to be written using the adaptive threshold T.
  • The sign-error set is expressed through the auxiliary sign vector and sign(Wj), linking directional errors to the conditional sign analysis.
  • Conditionally on V, the relevant W signs are independent and have negative-sign probability at least 1/2, so Corollary 1 supplies the required bound.
  • The knockoff+ proof proceeds analogously, using its slightly more conservative threshold definition.

B.3 Proof of Lemma 1

Lemma 1 is proved by constructing Bernoulli variables through a random subset and then controlling a reverse-time stopping-time ratio. A supermartingale argument completes the bound.

  • A reverse-time stopping-time argument reduces the needed bound to exchangeability of selected Bernoulli variables under an appropriate filtration.
  • The ratio Mj = (1 + j)/(1 + B1 + ··· + Bj) is shown to be a supermartingale, using conditional exchangeability and the Bernoulli structure.
  • Optional stopping then yields the desired expectation bound, completing the lemma after applying the exchangeability corollary.
  • The proof constructs Bernoulli variables using a random set A and independent Bernoulli variables Qi, preserving mutual independence with probabilities ρi.

C Detailed results for GWAS experiment

The detailed GWAS results are organized by phenotype, with separate tables for LDL and HDL. The tables report SNP-level and region-level selection information alongside genomic identifiers.

  • LDL phenotype: Table 2 reports the GWAS experiment results for the LDL phenotype.
  • HDL phenotype: Table 3 reports the GWAS experiment results for the HDL phenotype.
  • LDL phenotype: The LDL results table includes SNP name, chromosome, base-pair position, and selection frequencies by SNP and by nearby-SNP region.
  • The tables use physical reference positions drawn from Human Genome Build 37/HG19.
Loading 1602.03574v3…