Source-linked AI summary

Multilinear tensor regression for longitudinal relational data

Peter D. Hoff

arXiv:1412.0048v2stat.ME

TL;DR

Relational data exhibit dependence between dyads, but unrestricted regression can require too many parameters for stable estimation. The paper develops multilinear tensor regression with separable regression and covariance structures; in applications, multiplicative models outperform additive alternatives and identify geographically patterned predictive dependencies, while extensions cover network effects and longer-term lags.

  • Problem

    Longitudinal relational analysis must quantify dependence between dyadic time series, while unrestricted VAR models can be unstable because they require very many parameters.

  • Method

    The paper uses multilinear tensor regression with Kronecker-structured regression parameters and separable covariance, including a multiplicative model for relational autoregression.

  • Results

    13.2% versus 5.8% R2: the multiplicative model outperformed the additive model, while separate dyad-specific fits had −2.4% average predictive R2 from severe overfitting.

  • Takeaways & Limitations

    The model quantifies which countries predict others’ future actions, identifies geographic and cross-relation dependencies, and represents effects such as reciprocity, transitivity, and longer-term lag dependence.

  • Takeaways & Limitations

    The framework is based on least-squares and normal-error modeling, motivating extensions for binary, ordinal, sparse, or otherwise non-Gaussian data.

Abstract

from arXiv · show

A fundamental aspect of relational data, such as from a social network, is the possibility of dependence among the relations. In particular, the relations between members of one pair of nodes may have an effect on the relations between members of another pair. This article develops a type of regression model to estimate such effects in the context of longitudinal and multivariate relational data, or other data that can be represented in the form of a tensor. The model is based on a general multilinear tensor regression model, a special case of which is a tensor autoregression model in which the tensor of relations at one time point are parsimoniously regressed on relations from previous time points. This is done via a separable, or Kronecker-structured, regression parameter along with a separable covariance model. In the context of an analysis of longitudinal multivariate relational data, it is shown how the multilinear tensor regression model can represent patterns that often appear in relational and network data, such as reciprocity and transitivity.

1. Introduction.

The paper develops a parsimonious regression approach for dependence between dyadic time series, using shared multiplicative parameters to model longitudinal relational data. In country-interaction data, this model outperforms additive and separate dyad-specific alternatives while supporting extensions to tensor-valued data.

  • Motivation: Longitudinal relational data comprise time series of directed relationships among node pairs, creating dependence across dyads even without shared nodes.The paper studies country-pair actions represented as matrices over time.
  • Motivation: Unrestricted VAR estimation is unstable or unavailable because the regression matrix has m^4 entries, motivating lower-dimensional parameterizations.The problem becomes especially severe when time series are not extremely long.
  • Model: The multiplicative specification represents effects that are strongest when both source- and target-side relationships are influential, unlike an additive model that requires only one side to be influential.This structure can capture relational patterns such as dependencies associated with paired alliances.
  • Empirical comparison: 13.2% versus 5.8%: the multiplicative model’s R2 more than doubled the additive model’s R2 despite equal parameter counts.Both models explained only a small fraction of total variation.
  • Empirical comparison: Separate dyad-specific rank-one fits achieved 26.5% in-sample R2 but −2.4% average predictive R2, indicating severe overfitting.These fits use roughly 2m^3 parameters rather than the multiplicative model’s roughly 2m^2.
  • Extensions: The framework extends to multilinear tensor regression with Tucker-product coefficients and Kronecker-structured covariance, accommodating multivariate relational measurements.The paper also develops least-squares and Bayesian estimation approaches.

2. The bilinear regression model.

The bilinear regression model represents matrix outcomes as AXB^T, reducing a potentially enormous regression parameter to a Kronecker product while retaining interpretable row- and column-specific effects. The paper develops least-squares estimation and establishes consistency under correct specification and positive-definite covariate covariance.

  • Model formulation: The model regresses a matrix Y on X through the mean AXB^T, equivalently using the Kronecker-structured coefficient B ⊗ A.A and B are unknown matrices, and the parameters are identifiable only up to reciprocal scaling.
  • Estimation: Alternating least squares estimates A and B by repeatedly minimizing the residual mean squared error with respect to one matrix while holding the other fixed.The procedure is a block coordinate descent algorithm that converges to a local minimum under conditions on the data.
  • Large-sample behavior: Under correct specification and positive-definite Σxx, the pseudotrue parameters equal the true parameters and the least-squares estimator is asymptotically consistent.Related propositions also connect zero conditional relationships to zero pseudotrue elements of A under stated covariance conditions.

3. Extension to correlated multiway data.

The bilinear model extends to multilinear tensor regression through Tucker products, while a separable covariance model accommodates dependence across tensor modes. The resulting framework supports least-squares, generalized least-squares, and Bayesian estimation for correlated multiway data, including longitudinal relational arrays.

  • Multilinear tensor regression: The bilinear model is a special case of multilinear tensor regression for mapping an explanatory tensor X to an outcome tensor Y.The mapping applies mode-specific matrices B1,...,BK through the Tucker product.
  • Bayesian inference: A semiconjugate Bayesian model with a Gibbs sampler provides posterior simulation for parameter values and richer characterization of parameter uncertainty.The Bayesian approach is presented for the joint multilinear mean and covariance model.
  • Multilinear tensor regression: The Tucker product applies linear transformations along each tensor mode and can be computed through matricization and a sequence of matrix multiplications.Mode-k matricization reshapes an array into a matrix that exposes the corresponding mode transformation.
  • Multilinear tensor regression: For each tensor element, b1,i1,j1 represents the multiplicative effect of slice j1 of X on slice i1 of Y.This generalizes the mode-specific interpretation of the bilinear model to multiple array modes.
  • Data representation: The model handles replicated observations by stacking arrays and fixing the parameter matrix for the added replication mode to the identity.This represents n replications as a (K + 1)-way model with BK+1 = In.
  • Application: In the international-relations application, residual correlation patterns among country groups contradict the patternless residual structure expected under i.i.d. errors.The analysis uses four relational measurements between 25 countries across 543 observations and examines eigenvectors of mode-specific residual correlation matrices.
  • Correlated errors: The array-normal error model imposes Kronecker-structured covariance, retaining tensor-mode correlations while avoiding unrestricted covariance estimation when sample sizes are limited.Conditional generalized least-squares estimates are obtained by coordinate descent after incorporating covariance along the other modes.

4. Analysis of longitudinal multirelational IR data.

The analysis compares multilinear models for four-country-action time series and extends the best-fitting framework to reciprocity, transitivity, and longer-term dependence. Bayesian estimation then characterizes cross-dyad, action-type, reciprocal, transitive, and lag-scale effects.

  • Data and candidate models: The data contain weekly counts of four action types between 25 countries from 2004 through mid-2014.
  • Data and candidate models: The joint multilinear model shares coefficient structure across action types while allowing one action type to predict another.This can improve estimation when coefficient matrices are similar across event types.
  • Reciprocity and transitivity: Reciprocity is modeled by adding reversed dyad lag predictors, so past actions from i2 to i1 can predict future actions from i1 to i2.The expanded predictor tensor has eight action-related slices.
  • Reciprocity and transitivity: Transitivity is represented with predictors based on relations connecting each endpoint to common targets, using a direction-agnostic third-order measure.These predictors occupy four additional slices, expanding the tensor to 12 predictor variables along its third mode.
  • Longer-term dependence: The longer-term model uses separability to combine weekly and monthly lag effects without doubling the dimension of B3.B4 encodes the relative effect of one-week versus one-month lagged data.
  • Model comparison: The relational multilinear model outperformed the competing models in predictive performance and beat the joint multilinear model on all ten test data sets.These results suggest that the relational model was not overfitting relative to simpler alternatives.
  • Parameter estimation and interpretation: Posterior estimates show that own-dyad, same-action history is generally the strongest predictor, followed by other actions in the same dyad or relations sharing an actor or target.Diagonal B1 and B2 effects were positive, with diagonal posterior means about 35 times larger than off-diagonal elements on average.
  • Parameter estimation and interpretation: The posterior ratio of one-week to one-month lag effects had mean 1.98 and a 95% interval of (1.94, 2.03).Both lag coefficients were positive in every Gibbs-sampler iteration.

5. Discussion.

The article develops a general multilinear tensor regression model and applies it to longitudinal relational data, quantifying dependencies among country actions while supporting network and temporal effects.

  • The multilinear tensor regression model handles correlated outcome tensors with multiplicative rather than additive regression coefficients.
  • The model can represent reciprocity, transitivity, and temporal effects from lags beyond a first-order autoregressive model.
  • The international-relations application identifies countries whose actions predict a given country's future actions and quantifies those dependencies.
  • The strongest predictive dependencies are generally between geographically close countries, with exceptions involving
  • The model may require extensions for binary or ordinal data, and sparse binary tensors may not provide enough information for stable estimates.

APPENDIX A: PROOFS

The appendix derives properties of pseudotrue parameters under misspecification, including conditions under which elements of the multiplicative coefficient matrix vanish or reflect row-specific associations.

  • When the relevant expectation is invertible, the pseudotrue parameter for A is characterized through the model's conditional least-squares criterion.
  • If B is zero, the corresponding pseudotrue elements of A can be set to zero without changing the asymptotic criterion function.
  • When x_j is mean zero and independent of other rows of X, the relevant element of A is determined by a row-specific expression; if y_i and x_j are uncorrelated, it is zero.
  • With Kronecker-structured predictor covariance, the appendix expresses the pseudotrue parameter using blockwise matrix operations and the Hadamard product.
  • Under the proposition's zero-effect assumption, the elements of A describing effects from row j of X to row i of Y are zero.

APPENDIX B: DETAILS OF THE MCMC ALGORITHM

The MCMC implementation uses multiple Gibbs samplers with convergence checks and post-processing, while traceplots indicate good mixing for the sampler initialized at least-squares estimates.

  • Four Gibbs samplers approximated the posterior: three used random starts and one started at the least-squares estimates.
  • Each sampler ran for 5500 iterations, including 500 iterations allocated for convergence to the stationary distribution.
  • The least-squares-initialized sampler appeared to converge essentially immediately, whereas randomly initialized samplers took about 50 to 250 iterations.
  • Normalization preserved the relative magnitudes of the coefficient matrices while leaving the magnitude of their Kronecker product unchanged.
  • Traceplots of B3's 48 entries indicated very good Gibbs-sampler mixing after convergence.
Loading 1412.0048v2…