Source-linked AI summary
Parsimonious Tensor Response Regression
Lexin Li, Xin Zhang
TL;DR
The paper addresses regression for scientific data combining high dimensionality with complex structure and underlying correlations. It proposes parsimonious tensor response regression estimated through the envelope method, with asymptotically efficient estimation and demonstrated effectiveness in simulations and real-data analyses.
Problem
Scientific data can combine high dimensionality with complex structure, while conventional approaches may ignore underlying correlations among observations.
Method
The paper proposes a parsimonious tensor response regression model and estimates it using the envelope method to focus on material information.
Results
The resulting estimator is asymptotically efficient, and simulations and real-data analyses demonstrate its effectiveness.
Takeaways & Limitations
The envelope-based estimator reduces the information used for estimation while retaining effectiveness across simulations and real-data analyses.
Takeaways & Limitations
The selected envelope dimension reflects a bias–variance trade-off, and the resulting estimator can be more variable despite being unbiased.
Abstract
from arXiv · showhide
Aiming at abundant scientific and engineering data with not only high dimensionality but also complex structure, we study the regression problem with a multidimensional array (tensor) response and a vector predictor. Applications include, among others, comparing tensor images across groups after adjusting for additional covariates, which is of central interest in neuroimaging analysis. We propose parsimonious tensor response regression adopting a generalized sparsity principle. It models all voxels of the tensor response jointly, while accounting for the inherent structural information among the voxels. It effectively reduces the number of free parameters, leading to feasible computation and improved interpretation. We achieve model estimation through a nascent technique called the envelope method, which identifies the immaterial information and focuses the estimation based upon the material information in the tensor response. We demonstrate that the resulting estimator is asymptotically efficient, and it enjoys a competitive finite sample performance. We also illustrate the new method on two real neuroimaging studies.
1 Introduction
The paper addresses tensor response regression for high-dimensional, structurally complex scientific data, especially comparing tensor images across groups while adjusting for covariates. It proposes a parsimonious model that jointly represents all voxels and uses envelope-based estimation to reduce parameters and improve efficiency.
- Problem: Tensor response regression models multidimensional-array responses with vector predictors, supporting image comparisons across groups after covariate adjustment.
- Limitations of existing approaches: Existing neuroimaging approaches often regress one voxel at a time, ignoring correlations among voxels.
- Methodological contribution: The proposed approach jointly models all image voxels while incorporating their intrinsic spatial correlations.
- Scientific interpretation: The model treats the image tensor as response and covariates such as disease status and age as predictors, targeting changes in brain regions across groups.
- Generalized sparsity: A generalized sparsity principle assumes that some tensor-response information is irrelevant to regression, reducing free parameters and improving interpretation.
- Estimation and evaluation: The envelope method identifies and jointly estimates material information, and simulations and real-data analyses show improved performance over alternative solutions.
2 Preparations
The preparations introduce tensor notation and the envelope framework for multivariate response regression. The generalized sparsity principle isolates predictor-relevant response information, reducing model complexity and yielding asymptotically more efficient estimation.
- Tensor notation: A tensor is a multidimensional array; its order is also called its dimension, way, or mode, and fibers generalize matrix rows and columns.
- Tensor operations: Tensor operations include vectorization, mode-k matricization, mode-k matrix products, and mode-k vector products.
- Envelope model: The envelope model assumes that some response components remain stochastically constant as predictors vary and are irrelevant to regression.
- Generalized sparsity: The generalized sparsity principle permits irrelevant linear combinations of response variables, unlike usual sparsity principles that focus on individual variables.
- Material and immaterial information: Under the envelope assumption, regression focuses on the material subspace while immaterial variation contributes extraneous estimation variability.
- Efficiency and complexity: p(r −u) parameters are eliminated relative to the unconstrained model, and the envelope estimator is asymptotically more efficient than ordinary least squares.
3 Models
The paper formulates tensor response regression with a vector predictor and separable covariance, then uses generalized sparsity to isolate response information relevant to regression. The resulting tensor envelope representation reduces dimensionality by focusing on material subspaces.
- Tensor response linear model: The tensor coefficient B captures the relationship between an mth-order tensor response Y and vector predictor X.The model assumes tensor errors are independent of X, mean zero, and have separable Kronecker covariance.
- Parsimony: Because the envelope dimensions uk are usually smaller than the response dimensions rk, the representation substantially reduces free parameters and improves estimation feasibility.The covariance is also constrained by a separable Kronecker structure to reduce its parameter count.
- Generalized sparsity principle: The generalized sparsity principle separates each mode into material and immaterial response components, with the latter independent of X and non-informative for the former.The decomposition writes Y as material part P(Y) plus immaterial part Q(Y).
- Parsimonious representation: The parsimonious representation projects the tensor response onto mode-specific subspaces and models the resulting core tensor as a surrogate response.The core tensor has dimensions u1 × ... × um, typically smaller than the original response dimensions.
- Tensor envelope: The tensor envelope TΣ(B) is the minimum and unique tensor product of reducing subspaces containing the coefficient information.It can equivalently be constructed from mode-specific envelopes and combined using a Kronecker product.
4 Estimation
The paper estimates the tensor envelope model with an iterative likelihood-based procedure and a non-iterative one-step alternative. The one-step estimator separates envelope-basis estimation, accelerates computation, retains consistency, and is recommended in practice.
- Estimation framework: The primary estimation target is B, parameterized through envelope bases Γk, reduced coefficient Θ, and covariance components Ωk and Ω0k.The objective function estimates B and the separable covariance jointly under the tensor envelope parameterization.
- Iterative estimator: Under normal errors, the iterative procedure minimizes the negative log-likelihood; without normality, it provides a moment-based estimator.The iterative algorithm alternates parameter updates until the objective function converges.
- Iterative estimator: The iterative algorithm initializes B with elementwise OLS and updates covariance components using residuals and tensor-mode calculations.Its covariance updates follow the approach of Dutilleul and Manceur and Dutilleul.
- Computational considerations: The iterative estimator can be computationally intensive because envelope-basis estimation uses non-convex Grassmann optimization with multiple local minima.The procedure may therefore require multiple starting values and internal iterations.
- One-step estimator: The one-step estimator replaces envelope-basis optimization with a modified objective and runs the algorithm only once.The modified objective makes estimation of the mode-specific envelope bases separable.
- One-step estimator: The one-step estimator is computationally much faster, has competitive finite-sample performance, remains consistent, and is recommended for practice.Its basis estimates can use a sequential moment algorithm that is described as faster and more stable than Grassmann optimization.
5 Asymptotics
The asymptotic analysis establishes root-n consistency and asymptotic normality for the envelope estimators under weak moment conditions. It also shows that envelope estimation is asymptotically more efficient than OLS, while practical covariance substitution yields conservative p-values.
- Consistency: Both the iterative and one-step estimators are root-n consistent for the true tensor coefficient, even when errors are not normally distributed.The result is established under finite fourth moments and is stated in terms of projection matrices for the envelope subspaces.
- Asymptotic normality: Under normal errors, the iterative envelope estimator is maximum likelihood and has asymptotic covariance no greater than OLS.The corresponding asymptotic normality result compares UENV with UOLS.
- Efficiency: The envelope estimator is asymptotically more efficient than OLS for estimating the tensor coefficient.The theoretical efficiency gain increases when immaterial variation dominates material variation.
- Known basis: With a known envelope basis, the envelope estimator is root-n consistent, asymptotically normal, and has covariance no greater than OLS.The efficiency difference grows as immaterial covariance variation becomes more dominant.
- Unknown basis: When the envelope basis is unknown, both OLS and envelope parameter estimators remain root-n consistent and asymptotically normal.The envelope and covariance estimators are asymptotically correlated, so the coefficient covariance lacks a simple explicit form.
- Voxelwise inference: For voxelwise inference, substituting UOLS for the envelope covariance is computationally simpler but produces more conservative p-values.The resulting p-values can be thresholded directly or adjusted using false discovery rate correction.
6 Simulations
Simulations compare the envelope estimator with OLS and tensor predictor regression, examine envelope dimension and immaterial variation, and extend evaluation to three-way responses.
- Simulation design: The simulations compare the new estimator with one-at-a-time OLS and Zhou et al.’s tensor predictor regression.They also assess model misspecification, envelope dimension, immaterial variation, and three-way tensor responses.
- Comparison with OLS: With weak signals at SNR 0.01 or 0.1, OLS failed to identify meaningful signal, whereas the envelope estimator recovered it with n = 20.The comparison varied signal shape and strength.
- Immaterial information: Greater dominance of immaterial information improved envelope-estimator performance, while OLS continued to miss meaningful signal.The result supports recognizing immaterial information during estimation.
7 Real Data Analysis
The real-data analyses apply the method to EEG and MRI tensor responses, where envelope estimates produce clearer or more informative signal patterns than OLS.
- EEG analysis: The EEG study used 77 alcoholic individuals and 44 controls, with responses formed from 64 electrodes and 256 time points.Averaging four consecutive time points yielded a 64 × 64 matrix response.
- EEG analysis: For EEG data, envelope estimation identified relevant channel and time ranges distinguishing alcoholic from control groups.The highlighted ranges were channels about 15–30 and 45–60, at times 30–120 and about 200–240.
- EEG analysis: The EEG OLS estimator was more variable, with less clearly defined signal regions than the envelope estimator.Both coefficient estimates and thresholded p-value maps were reported.
- ADHD analysis: In the ADHD analysis, OLS revealed essentially no useful information, whereas envelope estimation showed distinctive activity in the cuneus and fusiform gyrus.These regions were reported as consistent with prior literature.
8 Discussion
The paper proposes parsimonious tensor response regression that focuses estimation on relevant tensor-response information, reducing free parameters while achieving asymptotic efficiency. Simulations and real-data analyses support the estimator’s effectiveness, while practical use requires tuning the envelope dimension and balancing coefficient- and p-value-map interpretations.
- Contribution: The proposed parsimonious model uses generalized sparsity and envelope-based estimation for tensor responses with vector predictors.The approach focuses on material information in the tensor response rather than relying solely on penalty functions.
- Contribution: Envelope-based estimation identifies material tensor-response information and effectively reduces the number of free parameters.The resulting estimator is described as asymptotically efficient.
- Evidence and practice: Simulations and real-data analyses demonstrate the effectiveness of the new estimator.The paper also recommends combining coefficient maps and p-value maps to help identify relevant signal regions.
- Evidence and practice: Coefficient maps may recover weak signals but can include many small signal regions, whereas p-value maps are usually cleaner and may be conservative.The discussion specifically notes conservativeness when the true signal is weak.
- Limitations and tuning: The working envelope dimension is the main tuning parameter and creates a bias–variance trade-off.Selecting a dimension smaller than the truth yields a biased estimator, while selecting a larger dimension yields an unbiased but potentially more variable estimator.
- Limitations and extensions: Envelope-based estimation can be coupled with penalty functions for further regularization, but this extension remains under investigation.The method complements penalty-function-based solutions rather than replacing their possible use.