Source-linked AI summary
Tucker Tensor Regression and Neuroimaging Analysis
Xiaoshan Li, Hua Zhou, Lexin Li
TL;DR
Neuroimaging regression needs models that preserve multidimensional image structure while avoiding the prohibitive dimensionality of vectorized arrays. The paper proposes generalized linear tensor regression using Tucker-decomposed coefficient arrays, develops its estimation and theory, and compares it with CP regression. The model provides sound recovery of complex signals and consistently estimates the best Tucker approximation in Kullback-Leibler distance, while offering flexible dimension-specific ranks.
Problem
Vectorizing neuroimaging arrays creates ultrahigh-dimensional regression problems and destroys spatial structure; PCA-based reduction can lose regression-relevant information.
Method
The paper develops generalized linear tensor regression whose coefficient arrays follow a Tucker decomposition, with associated estimation, asymptotic theory, and regularization.
Results
The model provides sound recovery of complex signals and consistently estimates the best Tucker-structure approximation to the full array model in Kullback-Leibler distance.
Takeaways & Limitations
Tucker regression offers a flexible neuroimaging framework that permits distinct orders across tensor dimensions and supports principled imaging-data downsizing.
Takeaways & Limitations
Real imaging signals rarely have exact low rank, so Tucker and CP models approximate rather than exactly represent the underlying signal.
Abstract
from arXiv · showhide
Large-scale neuroimaging studies have been collecting brain images of study individuals, which take the form of two-dimensional, three-dimensional, or higher dimensional arrays, also known as tensors. Addressing scientific questions arising from such data demands new regression models that take multidimensional arrays as covariates. Simply turning an image array into a long vector causes extremely high dimensionality that compromises classical regression methods, and, more seriously, destroys the inherent spatial structure of array data that possesses wealth of information. In this article, we propose a family of generalized linear tensor regression models based upon the Tucker decomposition of regression coefficient arrays. Effectively exploiting the low rank structure of tensor covariates brings the ultrahigh dimensionality to a manageable level that leads to efficient estimation. We demonstrate, both numerically that the new model could provide a sound recovery of even high rank signals, and asymptotically that the model is consistently estimating the best Tucker structure approximation to the full array model in the sense of Kullback-Liebler distance. The new model is also compared to a recently proposed tensor regression model that relies upon an alternative CANDECOMP/PARAFAC (CP) decomposition.
1 Introduction
Neuroimaging regression must handle massive multidimensional covariates without discarding their spatial structure. The paper develops Tucker tensor regression as a flexible, supervised low-rank alternative to vectorized, feature-extracted, PCA-based, and CP-based approaches.
- Motivation: Feature extraction and unsupervised dimension reduction transform imaging regression into classical vector regression, while PCA may lose regression-relevant information.The paper notes no consensus on the best brain-image summary even for a single modality.
- Comparison and implications: For a 128-by-128-by-128 MRI with five covariates, parameters decrease from 2,097,157 to 389 under rank 1 and 1,157 under rank 3.The related CP model was reported to recover even high-rank signals, motivating comparison with Tucker regression.
- Tucker tensor regression: The proposed models assume a Tucker decomposition for the coefficient array, using a core tensor and factor matrices with potentially distinct orders across dimensions.This extends generalized linear tensor regression to multidimensional array-valued imaging covariates.
- Tucker tensor regression: Tucker regression performs supervised low-rank dimension reduction, supports continuous, binary, and count responses, and is designed for efficient estimation at imaging scale.Regularization is also studied, including sparsity-oriented models and incorporation of prior scientific knowledge.
2 Model
The model section develops generalized linear regression for tensor-valued covariates by imposing a Tucker structure on the coefficient array. It explains the supervised estimation framework, its dimensionality reduction, and its parameter-size and empirical-fit relationship to CP regression.
- Tensor preliminaries: A tensor is a multidimensional array whose fibers generalize matrix rows and columns, with operators such as vec, matricization, and mode-d multiplication supporting the model formulation.Mode-d multiplication transforms fibers by a matrix, and its mode-d matricization equals matrix multiplication of the corresponding matricization.
- Tucker regression model: The generalized linear tensor regression model imposes a Tucker decomposition on coefficient array B and estimates its core tensor and factor matrices jointly from responses and covariates.This is presented as a supervised version of Tucker decomposition, with the response guiding estimation of the factor matrices.
- Tucker regression model: The Tucker model includes CP regression as a special case with equal mode ranks and a super-diagonal core, while allowing distinct ranks and higher overall rank.CP uses R1 = · · · = RD = R and has rank at most R; Tucker can have rank as high as R^D.
- Dimensionality reduction: Tucker duality converts the inner product with a Tucker coefficient tensor into an inner product between the core tensor and transformed data, effectively reducing the predictor dimensions.When factor matrices are fixed and Rd ≪ pd, fitting the model becomes a lower-dimensional tensor regression; the transformed data provide the corresponding reduced predictors.
- Regularization: The model can use fixed or over-complete basis matrices, with regularization on the core tensor through hard or soft thresholding and basis choices such as wavelets or Fourier functions.The section describes low-rank decompositions as hard thresholding and penalized regression as soft thresholding for regularizing G.
- Model size: For D = 2, Tucker and CP have the same model size; for D = 3, Tucker costs R(R −1)(R −2) more parameters when R > 2, while unequal Tucker ranks can make it more parsimonious.For D = 4 and R = 3, Tucker takes 54 more parameters than CP under equal ranks, but the equal-rank assumption does not hold generally.
- Empirical comparison: In the simulated 3D imaging example, Tucker regression generally achieves better model fit than CP regression at the same number of free parameters, as indicated by lower deviance.The comparison uses n = 1,000 subjects and evaluates CP ranks R = 1, . . . , 5 against several Tucker orders.
3 Estimation
The Tucker regression is estimated by alternating updates of factor matrices and the core tensor, reducing each block update to a low-dimensional GLM. Under stated regularity conditions, the block-relaxation algorithm converges to stationary points and is locally attracted to strict local maxima.
- Alternating estimation: The proposed estimator uses maximum likelihood for Tucker tensor regression and alternately updates factor matrices and the core tensor.The objective is linear in each block separately, enabling blockwise GLM updates.
- Alternating estimation: Updating a factor matrix turns the problem into a low-dimensional GLM with that factor matrix as parameter.The remaining factor matrices and core tensor are held fixed during the update.
- Alternating estimation: Updating the core tensor likewise produces a GLM with vecG as the parameter and a transformed tensor predictor.This block has Q parameters, making the update computationally manageable.
- Algorithmic design: The algorithm alternates block updates without enforcing factor-matrix column orthogonality because the goal is tensor-signal approximation rather than modewise principal components.The procedure is summarized as Algorithm 1 and uses random initial factor matrices and core tensor.
- Convergence: Under continuity, coercivity, strict block concavity, and isolated stationary points, the iterates converge globally to a stationary point and locally to a strict local maximum.The local result applies when the initialization lies in the attraction region of the strict local maximum.
4 Statistical Theory
The statistical theory addresses identifiability, asymptotic behavior, and inference for Tucker tensor regression under constrained parameterizations. It establishes consistency and asymptotic normality under regularity conditions, while showing that misspecified low-rank models target the best Tucker approximation in Kullback–Leibler distance.
- Asymptotic framework: Large-sample theory is relevant when the sample size exceeds the Tucker model’s effective parameter count, while regularization addresses smaller or moderate samples.The asymptotic analysis targets the regime in which n is considerably larger than pT.
- Asymptotic framework: The model consistently estimates the best Tucker-structure approximation to the full array model in Kullback–Leibler distance.This result applies even when the true coefficient array is not exactly low rank.
- Score and information: The score, Hessian, and Fisher information are derived for statistical estimation and inference, with gradients expressed through the factor matrices and core tensor.The Hessian includes derivatives of the nonlinear systematic part, while canonical-link special cases simplify it.
- Identifiability: Tucker parameters are identifiable only up to nonsingular transformations, so constrained parameterizations are required for asymptotic analysis.The paper fixes selected factor-matrix entries or imposes alternative restrictions to remove transformation indeterminacy.
- Identifiability: Local identifiability holds when the information matrix has constant rank near a parameter point and the stated rank condition is satisfied.The relevant condition is expressed through the factor matrices and their Jacobian blocks.
- Asymptotic results: The MLE is consistent for normal, binary, and Poisson tensor regression, and asymptotically normal under an interior-point nonsingular-information condition.The asymptotic normal distribution has covariance matrix I^-1(B0) after sqrt(n) scaling.
5 Regularized Estimation
Regularized Tucker regression is motivated by ultrahigh dimensionality, limited sample sizes, and the need to incorporate structural knowledge. The paper regularizes the core tensor to combine sparsity and shrinkage while retaining computationally tractable updates.
- Motivation: Regularization addresses situations where the effective parameter count can exceed the number of observations and may stabilize estimation.The paper also motivates regularization as a way to improve risk properties and incorporate prior scientific knowledge.
- Scientific structure: Regularization can encode prior brain-structure knowledge, including symmetry constraints for MRI parameters along the coronal plane.The choice of regularization depends on the practical purpose of the scientific study.
- Regularization choices: Two regularization strategies are considered: penalizing the core tensor alone or penalizing both the core tensor and factor matrices.The presented illustration focuses on regularization of the core tensor.
- Core regularization: The proposed regularized formulation penalizes only core-tensor elements, producing sparsity in Tucker outer products together with shrinkage.The penalty family includes lasso, ridge, elastic net, SCAD, and MC+ choices.
- Computation and tuning: Regularized estimation changes the core update to a penalized GLM while leaving the other alternating-update steps unchanged.The implementation uses a sparse-regression toolbox, and tuning can use cross-validation or BIC.
6 Numerical Study
Numerical experiments evaluate signal recovery, sample-size behavior, comparisons with CP regression, and an ADHD MRI application. Tucker regression recovers varied signals, improves with sample size, uses fewer parameters than CP while achieving more accurate estimation, and performs best among the reported ADHD models when regularized.
- Study design: The simulations examine signal-shape recovery, consistency as sample size increases, comparison with CP regression, and a real MRI application.The experiments include both synthetic tensor predictors and ADHD imaging data.
- Signal recovery: The Tucker model provides sound recovery of high-rank and natural-shaped two-dimensional signals, including disk and butterfly patterns.Figure 2 uses 64 × 64 image covariates and a sample size of 1000; TR(r) denotes an r-by-r core tensor.
- Tucker versus CP: The Tucker model uses fewer free parameters than CP and achieves more accurate estimation in the reported 100-replication comparison.The comparison evaluates root mean squared error under different tensor-rank setups.
- ADHD application: In the ADHD MRI analysis, the regularized Tucker model performs best on the testing data among the reported models.The reported testing misclassification rate for regularized Tucker is 0.32 to 0.36.
7 Discussion
The paper develops Tucker tensor regression as a flexible framework for imaging covariates, with fast estimation, regularization, and asymptotic guarantees. It compares Tucker with CP models and notes that low-rank approximations are useful when real imaging signals are not exactly low rank.
- Tucker regression contains CP tensor regression as a special case while providing a more flexible framework for imaging covariates.
- The Tucker framework includes a fast estimation algorithm, a general regularization procedure, and associated asymptotic properties.
- The paper provides detailed analytical and numerical comparisons between Tucker and CP tensor models.
- Because real imaging signals rarely have exact low rank, low-rank Tucker and CP models can provide reasonable approximations under limited sample sizes.
- Numerical studies show that the Tucker model can achieve sound recovery even for complex signals.
- The framework is general enough to support extensions such as multimodal imaging, imaging classification, and longitudinal imaging analysis, which are left for future research.