Source-linked AI summary

Separable covariance arrays via the Tucker product, with applications to multivariate relational data

Peter D. Hoff

arXiv:1008.2169v1stat.ME

TL;DR

The paper addresses how to model correlations across multiple index sets in multidimensional data beyond the matrix-normal setting. It constructs array normal distributions with separable covariance through Tucker-product transformations, develops maximum-likelihood and Bayesian inference, and illustrates the model on multivariate trade data. The discussion reports that the array model provides a better fit than a matrix normal model that covers only two of the four data modes, while noting that MLE existence and uniqueness remain open questions.

  • Problem

    Multidimensional data may contain dependencies along several index sets, but matrix-normal models provide covariance matrices for only two sets.

  • Method

    The paper constructs array normal distributions by applying multilinear Tucker-product transformations to arrays of independent standard normal variables and develops maximum-likelihood and Bayesian estimation.

  • Results

    The array model provides a better fit than a matrix normal model that includes covariance parameters for only two of the four data-array modes.

  • Takeaways & Limitations

    Separable array-normal models provide a framework for describing covariance within the index sets of multidimensional data.

  • Takeaways & Limitations

    Conditions for existence and uniqueness of the maximum-likelihood estimator remain an open area of research.

Abstract

from arXiv · show

Modern datasets are often in the form of matrices or arrays,potentially having correlations along each set of data indices. For example, data involving repeated measurements of several variables over time may exhibit temporal correlation as well as correlation among the variables. A possible model for matrix-valued data is the class of matrix normal distributions, which is parametrized by two covariance matrices, one for each index set of the data. In this article we describe an extension of the matrix normal model to accommodate multidimensional data arrays, or tensors. We generate a class of array normal distributions by applying a group of multilinear transformations to an array of independent standard normal random variables. The covariance structures of the resulting class take the form of outer products of dimension-specific covariance matrices. We derive some properties of these covariance structures and the corresponding array normal distributions, discuss maximum likelihood and Bayesian estimation of covariance parameters and illustrate the model in an analysis of multivariate longitudinal network data.

1 Introduction

The article extends separable covariance modeling from matrices to arrays, using multilinear transformations to construct array normal models and support covariance estimation. It motivates the approach with relational and longitudinal data whose dependencies may occur along multiple index sets.

  • The paper constructs and estimates a class of covariance models and Gaussian distributions for multidimensional arrays.
  • Array data can represent relational measurements collected across objects, conditions, time points, or other index sets.
  • The motivating data consist of commodity exports between countries across commodity classes and years, represented as a four-way array.
  • Researchers may seek similarities among network nodes and correlations among adjacent time points or other array indices.
  • For matrices, separable covariance uses dimension-specific covariance matrices, offering a stable and parsimonious alternative when unrestricted covariance estimation is unstable or unavailable.
  • The paper generalizes this framework to arbitrary-dimensional arrays through the Tucker product and develops maximum-likelihood and Bayesian estimation procedures.
  • The model is illustrated with trade-volume data for pairs of 30 countries, 6 commodity types, and 10 years.

2 Separable covariance via array-matrix multiplication

The Tucker product applies separate matrix transformations along each array mode, generating separable covariance structures from arrays with uncorrelated entries. This provides the array-level analogue of matrix separability and shows that the class is closed under single-mode transformations.

  • Array notation and basic operations: A K-array is a multidimensional object indexed by K sets, with each index set corresponding to one mode.
  • Array notation and basic operations: The k-mode product multiplies a selected mode by a matrix while preserving the other modes, and successive mode products form the Tucker product.
  • Separable covariance via the Tucker product: For matrices, separable covariance is represented as Cov[vec(Y)] = Σ2 ⊗Σ1, with separate covariance matrices for rows and columns.
  • Separable covariance via the Tucker product: Applying invertible mode-specific transformations to an array with uncorrelated mean-zero variance-one entries produces covariance factors for each mode.
  • Separable covariance via the Tucker product: A transformation along mode k replaces only that mode’s covariance factor Σk with GΣkG^T while leaving the other factors unchanged.
  • Separable covariance via the Tucker product: Separable covariance arrays arise through repeated single-mode multiplications from an array whose covariance is Im1 ◦· · · ◦ImK, and the class is closed under these transformations.

3 Covariance estimation with array normal distributions

The array normal model represents tensor-valued observations with separable covariance, formed from multilinear transformations of independent standard normal arrays. The section develops conditional distributions, maximum-likelihood estimation, and Bayesian estimation for the dimension-specific covariance matrices.

  • Array normal model: Array normal distributions use an outer-product covariance structure, Σ1 ◦ · · · ◦ ΣK, across the array’s modes.The model is generated by transforming an array of independent standard normal entries with mode-specific matrices.
  • Array normal model: Conditioning on selected array slices preserves the array normal class, changing only the relevant mode covariance to its conditional covariance.For subsets along the first mode, Σ1,b|a = Σ1[b,b] − Σ1[b,a](Σ1[a,a])^-1Σ1[a,b].
  • Maximum likelihood: The maximum-likelihood procedure updates each covariance matrix conditionally on the others through an iterative block-coordinate descent algorithm.Each iteration increases the likelihood, and the updates use standardized residual arrays and mode-specific effective sample sizes.
  • Maximum likelihood: The covariance-factor scales are not separately identifiable, so replacing Σ1 and Σ2 by cΣ1 and Σ2/c leaves the distribution unchanged.Consequently, maximum-likelihood scale estimates depend on their initial values.
  • Bayesian estimation: Bayesian estimation uses semiconjugate priors, yielding an array normal conditional distribution for M and inverse-Wishart conditionals for covariance parameters.The conditional mean is [κ0M0 + nȲ]/[κ0 + n], with covariance Σ1 ◦ · · · ◦ ΣK/[κ0 + n].
  • Bayesian estimation: A total-variance parameter γ reparameterizes the prior to address difficult interpretation of the nonidentifiable covariance-factor scales.The prior expected total variation can be set through γ or assigned a gamma prior, and Gibbs sampling approximates posterior quantities.

4 Example: International trade

The trade example models yearly changes in international commodity flows as a four-mode array with exporter, importer, commodity, and time dependence. Posterior predictive checks reject independent exporter and importer modes, while estimated correlations reveal geographic structure and no substantial commodity-mode lack of fit.

  • Data and model: The four-mode array models residual covariance across exporters, importers, commodities, and time points.The analysis compares a model with identity exporter and importer covariance against a fuller model allowing those modes to be correlated.
  • Posterior inference: Posterior inference uses Gibbs sampling because temporal dependence makes direct integration of the joint posterior difficult.Each model used 205,000 iterations, discarded the first 5,000, and retained 5,000 parameter values.
  • Posterior predictive checks: The independent exporter-importer model exhibits substantial lack of fit for summary statistics t1 and t2.These checks indicate that assuming i.i.d. structure along the first two array modes does not fit the data.
  • Posterior predictive checks: Neither model exhibits substantial lack of fit for covariance among commodities, as measured by t3.Thus the posterior predictive comparison distinguishes the first two modes from the commodity mode.
  • Estimated correlations: Estimated exporter and importer correlations are geographically structured, while the correlation-matrix eigenvalues suggest possible lower-dimensional modeling.Countries with similar eigenvector values are typically geographically near each other.

5 Discussion

The discussion presents separable array-normal covariance models as a flexible framework for modeling array data, with extensions for factor-analytic structure and non-normal outcomes. A longitudinal trade-data example shows that the full array-normal model fits better than a matrix-normal model, while maximum-likelihood existence remains unresolved.

  • Model and applications: The proposed array-normal class represents separable covariance structures through array-matrix multiplication, supporting covariance modeling across array index sets.The construction uses multilinear transformations and motivates the Tucker-product formulation.
  • Model and applications: In longitudinal trade data, the full array-normal model provides a better fit than a matrix-normal model covering only two of four array modes.The comparison concerns a model with covariance structure across all array modes versus one restricted to two modes.
  • Limitations and open problems: Conditions for existence and uniqueness of the maximum-likelihood estimator remain unresolved, even for the matrix-normal case.A sufficient-looking iterative-algorithm condition can hold while the likelihood is unbounded and no MLE exists; stronger conditions are known but may not be necessary, and their generalization to arrays is unclear.
  • Model extensions: A factor-analytic extension can impose simplifying structure on component matrices, including rank-reduced covariance estimates for high-dimensional modes.The proposed array factor model extends vector factor analysis and can combine structured and unstructured covariance components across modes.
  • Model extensions: The framework can fit factor-analytic covariance structure to some modes while leaving others unstructured, yielding a separable array-normal covariance model.This mixed structure is intended for arrays whose modes differ substantially in dimension or covariance complexity.
  • Model extensions: Extensions to non-normal continuous or discrete data could embed the array-normal model in generalized linear or ordered-probit models.A latent array-normal probit construction is described for three-way binary arrays.

Appendix

The appendix derives covariance identities and distributional results for array-normal models. It also develops density and conditional-distribution calculations underlying estimation procedures.

  • Covariance derivations: The covariance of vectorized transformed arrays is derived from multilinear transformations of independent standard-normal elements.The derivations use mode-wise transformations and the Tucker-product representation.
  • Density: The array-normal density is obtained by re-expressing the multivariate normal density of vec(Y − M).The resulting covariance is expressed as ΣK ⊗· · · ⊗Σ1.
  • Estimation: For covariance estimation, the number of columns of Y(k) serves as the sample size for Σk.This quantity is defined directly from the mode-k matricization.
  • Conditional distributions: The appendix derives conditional distributions and variance expressions using partitioned-matrix results and matrix-normal quadratic forms.The calculations introduce full conditional distributions for covariance components and identify a row variance as Ψ−1.
  • Conditional distributions: For an array-normal distribution, a matricized mode has matrix-normal structure with row covariance Σ1 and column covariance ΣK ⊗· · · ⊗Σ2.This representation supports conditional-distribution calculations by reducing them to the matrix-normal case.
Loading 1008.2169v1…