Source-linked AI summary

Generalized Network Psychometrics: Combining Network and Latent Variable Models

Sacha Epskamp, Mijke Rhemtulla, Denny Borsboom

arXiv:1605.09288v4math.STstat.ME

TL;DR

Psychometric network models need to accommodate both latent variables and violations of local independence. The paper embeds Gaussian graphical models in SEM, introduces latent and residual network formulations, and finds that exploratory search methods identify relevant structures adequately in simulations. The resulting framework is implemented in lvnet for confirmatory testing, exploratory search, and empirical modeling.

  • Problem

    Standard SEM relies strongly on local independence, while network models assume observed covariances are not caused by latent variables.

  • Method

    The paper develops Latent Network Modeling and Residual Network Modeling within SEM and implements stepwise and penalized maximum likelihood search in lvnet.

  • Results

    Simulation studies found high specificity and sensitivity that increased with sample size for stepwise and penalized maximum likelihood estimation of latent or residual network structures.

  • Takeaways & Limitations

    The framework supports fitting and comparing SEM, network, LNM, and RNM models while allowing exploratory searches for latent or residual network structures.

  • Takeaways & Limitations

    The empirical example is highly explorative, and the RNM structure should be replicated in independent samples before substantial interpretation.

Abstract

from arXiv · show

We introduce the network model as a formal psychometric model, conceptualizing the covariance between psychometric indicators as resulting from pairwise interactions between observable variables in a network structure. This contrasts with standard psychometric models, in which the covariance between test items arises from the influence of one or more common latent variables. Here, we present two generalizations of the network model that encompass latent variable structures, establishing network modeling as parts of the more general framework of Structural Equation Modeling (SEM). In the first generalization, we model the covariance structure of latent variables as a network. We term this framework Latent Network Modeling (LNM) and show that, with LNM, a unique structure of conditional independence relationships between latent variables can be obtained in an explorative manner. In the second generalization, the residual variance-covariance structure of indicators is modeled as a network. We term this generalization Residual Network Modeling (RNM) and show that, within this framework, identifiable models can be obtained in which local independence is structurally violated. These generalizations allow for a general modeling framework that can be used to fit, and compare, SEM models, network models, and the RNM and LNM generalizations. This methodology has been implemented in the free-to-use software package lvnet, which contains confirmatory model testing as well as two exploratory search algorithms: stepwise search algorithms for low-dimensional datasets and penalized maximum likelihood estimation for larger datasets. We show in simulation studies that these search algorithms performs adequately in identifying the structure of the relevant residual or latent networks. We further demonstrate the utility of these generalizations in an empirical example on a personality inventory dataset.

Introduction

The paper formalizes network psychometrics within SEM and introduces latent and residual network generalizations. These frameworks support exploratory estimation alongside confirmatory model comparison.

  • Network psychometrics: The Gaussian Graphical Model formalizes network modeling for multivariate normal psychometric data within a framework familiar to psychometricians.It relates network modeling to SEM by modeling the inverse covariance matrix and can be estimated in SEM.
  • Generalizations: LNM models networks among latent variables, enabling exploratory estimation of conditional independence relationships.Model search addresses equivalent-model problems that complicate exploratory SEM estimation.
  • Generalizations: RNM models a network among SEM residuals, allowing identifiable models with structured violations of local independence.It also accounts for covariance between items that may partly arise from latent factors.
  • Exploratory estimation: Stepwise search adds or removes edges when fit improves, whereas penalized maximum likelihood estimates a sparse network.The paper evaluates both approaches in four simulation studies and implements them in the free-to-use lvnet package.

Modeling Multivariate Gaussian Data

The multivariate Gaussian framework represents centered item responses through a covariance matrix and estimates a model-implied covariance matrix from sample covariances. Model fitting seeks a closely matching covariance structure with positive degrees of freedom.

  • Data and covariance structure: The response vector y is assumed centered and distributed according to a multivariate Gaussian density.The framework considers P observed items and N independent samples.
  • Data and covariance structure: The sample covariance matrix is defined as S = 1/(N−1)YᵀY from the response matrix Y.Rows of Y contain realized response vectors for individual subjects.
  • Estimation: Maximum likelihood estimation uses S to minimize −2 times the log-likelihood and obtain the model-implied covariance matrix ˆΣ.Perfect fit occurs when ˆΣ = S.
  • Model fit: A saturated model uses K = P(P + 1)/2 parameters and reproduces S exactly, whereas the target is a parsimonious model with K < P(P + 1)/2.The reduced-parameter model should still make ˆΣ resemble S closely.

Structural Equation Modeling

CFA and SEM model observed responses through latent variables, factor loadings, structural relations, and residual covariances. Their identifiability typically depends on restricting residual covariances through local independence.

  • CFA and SEM: CFA represents observed variables as linear effects of centered latent variables plus residuals, with factor loadings collected in Λ.The implied covariance structure depends on the latent variance-covariance matrix Ψ and residual matrix Θ.
  • CFA and SEM: SEM extends latent covariance modeling by adding structural linear relations through regression matrix B and latent residuals ζ.The latent residual covariance matrix becomes Ψ = Var(ζ).
  • CFA and SEM: Path analysis is obtained by setting Λ = I and Θ = O, allowing direct causal effects between observed variables to be modeled.This is a special case of the broader SEM formulation.
  • Local independence: Local independence typically sets Θ diagonal, meaning indicators are independent after conditioning on the latent variables.A fully estimated Θ is saturated, so systematic nonzero residual covariances threaten identification.
  • Local independence: Local independence may be violated by direct causal effects, semantic overlap, or reciprocal interactions among indicators.Such violations have also been questioned for symptoms conditional on a latent mental disorder.

Network Modeling

Network models represent variables as nodes connected by pairwise interactions or conditional associations. For Gaussian data, the GGM uses partial correlations and the precision matrix, but assumes observed covariances are not caused by latent variables.

  • Pairwise network models: Pairwise Markov Random Fields encode conditional associations: a missing edge means independence after conditioning on all other nodes.The example permits all three variables to correlate while encoding conditional independence between Y1 and Y3 given Y2.
  • Pairwise network models: Network interactions are not equivalent to marginal correlations: correlated variables may be unconnected, while connected variables may be uncorrelated.Common causes and common effects can produce these different patterns.
  • Gaussian Graphical Models: For multivariate Gaussian data, the GGM uses zero partial correlations to identify conditional independence and treats partial correlations as edge weights.The GGM is therefore a model of the inverse covariance rather than the covariance itself.
  • Gaussian Graphical Models: The precision matrix ˆK is the inverse of ˆΣ, and its off-diagonal elements are proportional to partial correlations after conditioning on all other variables.Standardizing the precision matrix yields the proposed network parameterization.
  • Model assessment: GGM structures can be tested against saturated and independence models using χ2 fit statistics and indices such as RMSEA and CFI.This provides confirmatory model assessment for psychometric networks.
  • Assumptions: The GGM assumes that observed-variable covariances are not caused by latent or unobserved variables.This assumption limits its use when a latent factor generates the observed covariance structure.

Generalizing Factor Analysis and Network Modeling

The paper proposes two SEM generalizations that represent latent or residual covariance structures as networks, extending network modeling to models with latent variables.

  • Latent Network Modeling: LNM models the latent-variable variance–covariance matrix as a Gaussian graphical model representing undirected conditional-independence relationships.It replaces directed latent effects with an undirected network.
  • Residual Network Modeling: RNM models the residual variance–covariance matrix as a Gaussian graphical model of pairwise residual interactions.This permits confirmatory factor structures even when local independence is systematically violated and residuals are correlated.
  • Modeling Frameworks: The four frameworks combine latent variables, manifest indicators, residuals, directed parameters, and undirected pairwise interactions in distinct model structures.Figure 2 contrasts SEM, LNM, RNM, and network-model configurations.
  • Modeling Frameworks: The CFA decomposition is used for LNM because its main application is exploratory estimation of relationships between latent variables.The paper uses CFA rather than SEM as the primary framework for this generalization.

Latent Network Modeling

Latent Network Modeling represents conditional independence among latent variables with an undirected Gaussian graphical model, avoiding directional and acyclic assumptions and supporting unique exploratory structures.

  • Latent Network Modeling: LNM models a CFA model’s latent variance–covariance matrix as a Gaussian graphical model.The latent network is specified through the latent covariance structure.
  • Latent Network Modeling: LNM represents conditional independence relationships between latent variables without assuming directionality or acyclicness.These assumptions are implicit in directed acyclic graph formulations of SEM.
  • Model Equivalence: LNM can model conditional-independence patterns that have no equivalent representation in restricted DAG or GGM structures.Figure 3 illustrates both equivalent representations and patterns unavailable under particular edge constraints.
  • Exploratory Estimation: Unlike exploratory DAG estimation, LNM assigns one undirected model to each set of conditional-independence relationships.This avoids equivalent undirected models with the same nodes.
  • Measurement Error: LNM allows network construction from concepts measured by multiple indicators while controlling for measurement error.This extends network analysis beyond networks based on single indicators.

Residual Network Modeling

Residual Network Modeling represents SEM residual structure as a Gaussian graphical model, preserving the factor structure while modeling structured residual interactions.

  • Residual Network Modeling: RNM models the residual structure of SEM as a Gaussian graphical model.The residual covariance matrix is parameterized through a residual network.
  • Residual Network Modeling: RNM permits identifiable factor models despite systematic violations of local independence and correlated residuals.Residual associations are represented as pairwise interactions rather than ordinary correlations.
  • Residual Network Modeling: Residual interactions represent pairwise linear effects that may reflect causal influence or partial indicator overlap remaining after latent effects are controlled.The framework distinguishes these interactions from residual correlations.
  • Confirmatory Factor Modeling: RNM improves confirmatory factor-model fit by adding residual interactions rather than allowing additional cross-loadings.The specified factor structure remains exactly intact.

Exploratory Network Estimation

The lvnet framework supports confirmatory and exploratory estimation of latent and residual networks, with model-search procedures designed for different dataset dimensions.

  • Confirmatory Estimation: Both LNM and RNM support confirmatory testing of network structures using estimation procedures similar to SEM.The modeled network is ΩΨ for latent structures or ΩΘ for residual structures.
  • Software: lvnet returns fit indices, parameter estimates, and model-comparison tests for specified models.Reported fit indices include RMSEA, CFI, and χ2.
  • Exploratory Estimation: The package provides step-wise model search and penalized maximum-likelihood estimation for unknown latent or residual network structures.These are the two exploratory search algorithms included for both network frameworks.
  • Evaluation Metrics: Sensitivity is the ratio of detected true edges to all true edges, whereas specificity is the ratio of detected missing edges to all absent edges.The supplied definitions quantify true-positive and true-negative network recovery.
  • Evaluation Metrics: High specificity is preferred to limit false positives, while sensitivity should increase with sample size.This preference supports sparse and interpretable network estimates.

Simulating Gaussian Graphical models

The simulations construct positive-definite inverse-covariance matrices by generating an unweighted network, assigning signed weights, and setting diagonal elements from row-wise absolute-weight sums.

  • Networks were first generated as unweighted structures before edge weights were assigned randomly.Weights were drawn uniformly between 0.5 and 1 and made negative with 50% probability.
  • Diagonal elements of K were set to 1.5 times each row’s absolute-weight sum, or to 1 when that sum was zero.
  • This construction was used to obtain a positive definite inverse-covariance matrix K.

Stepwise Model Search

Stepwise search estimates latent or residual network structures by iteratively adding or removing edges according to fit tests or information criteria. Simulations evaluated these procedures across sample sizes and showed improving edge recovery as sample size increased, with criterion-specific sensitivity and specificity trade-offs.

  • Search target: Stepwise search targets network structures in either ΩΨ for LNM or ΩΘ for RNM.The search can use χ2 difference testing or optimize AIC, BIC, or EBIC.
  • χ2 difference testing: The χ2 algorithm adds the edge that most improves fit or removes the edge that least worsens fit until no further change is supported.Adding requires significant improvement, while removing requires no significant worsening at α = 0.05.
  • Information-criterion optimization: Information-criterion search changes the edge producing the greatest AIC, BIC, or EBIC improvement until no changed edge improves the criterion.
  • Simulation study 1: 20 000 total simulated datasets evaluated stepwise search across five sample sizes and four criteria in the LNM study.The design used sample sizes from 50 to 1 000 and replicated each condition 1 000 times.
  • Simulation study 1: Sensitivity improved with sample size in LNM, AIC performed best and EBIC worst, and all criteria performed well from sample sizes of 500 onward.The study assessed sensitivity for detecting true edges and specificity for avoiding zero edges.
  • Simulation study 2: In RNM, sensitivity increased with sample size and AIC performed best, while specificity was very high across criteria and EBIC performed best.All four criteria performed well overall; EBIC favored caution and AIC favored discovery, with true-edge discovery generally good above 250 observations.

LASSO Regularization

LASSO regularization provides a faster exploratory alternative for larger latent and residual networks by shrinking parameters and selecting sparse structures across tuning parameters.

  • Stepwise search becomes very slow when networks exceed about 10 nodes, motivating LASSO for higher-dimensional structures.This is especially relevant for RNM, where indicator counts can exceed 10 even in small models.
  • LASSO penalizes the sum of absolute parameter values and yields sparse models with many relationships estimated as zero.The tuning parameter ν controls the penalty level.
  • The exploratory LASSO algorithm evaluates regularized models, counts parameters, selects the best AIC, BIC, or EBIC model, and refits it without LASSO.Parameters with absolute estimates below ϵ are fixed to zero during refitting.
  • lvnetLasso tests a logarithmically spaced sequence of 20 tuning parameters by default, spanning 0.01 to 1.
  • Simulation studies: LASSO performance was evaluated in latent-network simulations with 24 observed variables and residual-network simulations with 20 observed variables across 5 000 datasets each.Both studies varied sample sizes from 100 to 2 500 and selected models using AIC, BIC, or EBIC.

Empirical Example: Personality Inventory

The empirical example applies lvnet and LASSO to a 2 800-observation, 25-item Big Five inventory dataset, comparing CFA with residual and latent network extensions. The network extensions improve fit and yield interpretable latent and residual structures, but the exploratory result requires independent replication before substantial interpretation.

  • Dataset and estimation: The analysis estimates a CFA model, then applies LASSO to the RNM model using 100 tuning parameters and EBIC.The lvnet package is used for the confirmatory and exploratory analyses.
  • Dataset and estimation: The example uses 2 800 observations of 25 Big Five items, with five items designed to measure each of five personality traits.
  • Model comparison: The RNM model fits substantially better than CFA, while the RNM+LNM model removes five latent-network edges after accounting for residual interactions.
  • Network interpretation: In the final latent network, Agreeableness is conditionally independent of three traits given Extraversion, which is directly linked to all other traits.The residual network contains many meaningful connections but only 30% of all possible edges and has 176 degrees of freedom.
  • Caveat: The example’s procedures are highly exploratory, so the structure should be replicated in independent samples before substantial interpretation.

Conclusion

The paper presents LNM and RNM as generalizations that combine latent-variable and network modeling, enabling exploratory latent-network estimation and residual-network modeling without local independence. Simulations support the search algorithms’ specificity and sample-size-dependent sensitivity, while the authors note important scope and data limitations.

  • LNM networks latent variables, enabling exploratory searches for conditional independence relationships without prior theory, while RNM networks indicator residuals without assuming local independence.Together, the frameworks also account for measurement error and latent common causes within network modeling.
  • Step-wise search and penalized maximum likelihood estimation achieved high specificity and increasingly high sensitivity as sample size increased in simulations.Higher sample sizes led the algorithms to detect more true edges.
  • AIC produced the best sensitivity and EBIC the best specificity across the four simulation studies.The authors caution that information-criterion choice depends on whether discovery or caution is prioritized.
  • The framework is implemented in the freely available lvnet software, which can estimate the discussed model combinations.The current paper focuses on distinct latent- and residual-network benefits, although additional combinations are possible.
  • The presented modeling and simulation results are limited to multivariate normal data, particular model setups, and the sample-size ranges examined.More advanced models are possible, but some are not yet implemented in lvnet.
  • The optimized expressions are truly applicable only to complete data because they rely on summary statistics; FIML is not implemented in lvnet.

Introduction

The paper formalizes network modeling within SEM and introduces LNM and RNM to represent network structure among latent variables or indicator residuals. These frameworks support confirmatory testing and exploratory estimation while relaxing limitations involving equivalent models, latent variables, and local independence.

  • Network modeling in SEM: The Gaussian Graphical Model formalizes psychometric network modeling by modeling the inverse covariance matrix rather than the covariance matrix.It can be incorporated into SEM, enabling confirmatory testing and fit comparisons with saturated and independence models.
  • Latent Network Modeling: Latent Network Modeling represents relationships among latent variables as a network, allowing exploratory estimation of conditional independence without directional or acyclic assumptions.LNM avoids the equivalent undirected models that complicate exploratory SEM estimation.
  • Residual Network Modeling: Residual Network Modeling places a network structure on SEM residuals, allowing local independence to be violated while accounting for covariance attributable to latent factors.Residual interactions are constrained through the inverse residual covariance matrix, while latent factor interpretations can remain intact.
  • Exploratory estimation: The framework provides step-wise and penalized maximum-likelihood search algorithms for estimating unknown latent or residual network structures.Step-wise search adds or removes edges according to fit or information criteria, while penalized estimation supports larger datasets.
  • Simulation evidence: Simulation studies found high specificity and increasing sensitivity with sample size, with AIC favoring sensitivity and EBIC favoring specificity.For the step-wise procedures, sample sizes above 500 produced high sensitivity and specificity across criteria; the reported results were comparable to established network-estimation techniques.
  • Empirical example: In the personality-inventory example, RNM substantially improved fit over CFA, and the final RNM+LNM model removed five latent-network edges after accounting for residual interactions.Extraversion was the most central latent-network node, while the residual network contained 30% of all possible edges; the exploratory structure requires replication in independent samples.
Loading 1605.09288v4…