Source-linked AI summary
Copula Gaussian graphical models and their application to modeling functional disability data
Adrian Dobra, Alex Lenkoski
TL;DR
The paper addresses graphical-model determination for observational data containing mixed variable types, including the challenge of uncertain graph structure in high-dimensional, small-sample settings. It develops Bayesian CGGMs using a semiparametric Gaussian copula and extended rank likelihood, then applies them to functional-disability contingency data. The approach averages over graphs and accommodates latent-variable conditional independences, while its observed-variable interpretation is limited for discrete marginals and excludes non-binary, non-ordinal discrete variables.
Problem
Graphical-model methods need to handle mixtures of binary, ordinal, and continuous variables, while high-dimensional small-sample settings make selecting one graph unreliable.
Method
CGGMs embed Bayesian graphical-model selection in a semiparametric Gaussian copula and use extended rank likelihood with stochastic graph search and model averaging.
Results
The CGGMs identify conditional associations in the 16-dimensional functional-disability data and achieve squared error 407.04 for expected cell counts, versus 284.79 for the all two-way interaction model.
Takeaways & Limitations
Bayesian model averaging lets CGGMs account for many plausible graphs rather than selecting one particular graph, producing sensible results for sparse contingency tables.
Takeaways & Limitations
The framework excludes discrete variables that are neither binary nor ordinal, and latent-variable graphs need not match conditional-independence graphs among observed variables with discrete marginals.
Abstract
from arXiv · showhide
We propose a comprehensive Bayesian approach for graphical model determination in observational studies that can accommodate binary, ordinal or continuous variables simultaneously. Our new models are called copula Gaussian graphical models (CGGMs) and embed graphical model selection inside a semiparametric Gaussian copula. The domain of applicability of our methods is very broad and encompasses many studies from social science and economics. We illustrate the use of the copula Gaussian graphical models in the analysis of a 16-dimensional functional disability contingency table.
1. Introduction.
The paper develops CGGMs to identify conditional independence relationships for observational data containing binary, ordinal, and continuous variables. It motivates the approach with functional disability data and a 16-dimensional contingency table.
- Graphical models determine conditional independence relationships, a key component of analyzing observational studies.
- The motivating NLTCS table contains 16 binary disability measures covering six ADLs and ten IADLs among adults aged 65 and above.
- The proposed methodology determines complex conditional associations among the 16 daily living activities, extending prior analyses that left this issue unexplored.
- The method targets observational studies with mixtures of binary, ordinal, and continuous variables rather than only contingency tables.
- CGGMs impose conditional-independence constraints on the inverse correlation matrix within a Gaussian copula, with multivariate normality applying to latent variables.
2. Gaussian graphical models.
Gaussian graphical models represent conditional independence through graph structure and zeros in a constrained precision matrix. Because the graph space is enormous, inference requires stochastic graph search and Bayesian prior and posterior distributions.
- An undirected graph represents variables as vertices and conditional independence relationships through absent edges.
- The graph space contains 2^p(p−1)/2 undirected graphs, so stochastic search traverses neighboring graphs differing by one added or deleted edge.
- A GGM assumes X follows a multivariate normal distribution with precision matrix K, whose off-diagonal zeros encode missing graph edges.
- The precision matrix K is restricted to the cone P_G of symmetric positive definite matrices with zeros outside the graph’s edges.
- A G-Wishart prior on K is conjugate to the likelihood, yielding a G-Wishart posterior with updated parameters δ+n and D+U.
3. Incorporating binary and ordinal categorical variables.
The framework incorporates ordinal variables through continuous latent variables and thresholds, while Hoff’s rank-based formulation avoids explicitly using those thresholds. Observed samples constrain the latent samples according to the observed ordering.
- An ordinal variable with values 1 through d_v is represented by a continuous latent variable Z_v underlying the observed categories.
- The observed ordinal values are linked to the latent variable through ordered thresholds from −∞ to ∞.
- Fixing the first threshold identifies the threshold model, but the proposed rank-likelihood approach does not explicitly involve the thresholds.
- Given observed data, latent samples are constrained to preserve the ordering implied by the observed ordinal values.
4. Copula Gaussian graphical models.
CGGMs model mixed observed variables through latent Gaussian graphical models while leaving marginal distributions unspecified. Bayesian inference combines an extended rank likelihood with graph and precision-matrix sampling, including stochastic graph exploration and model averaging.
- Model specification: CGGMs represent binary, ordinal, count, and continuous observed variables through latent variables whose conditional independences are encoded by zeros in a precision matrix.The latent variables follow a multivariate normal model, while the observed-variable marginals are treated as nuisance parameters.
- Model specification: The observed-data likelihood is decomposed so that the extended rank likelihood p(D|K) is the only component relevant for inference on K and does not depend on the marginal distributions.This avoids formal parametric assumptions for the univariate distributions.
- Bayesian inference: Bayesian graph determination imposes graph-specific zero constraints on K, with a G-Wishart prior for K conditional on G and a uniform prior over candidate graphs.The joint posterior concerns both the precision matrix and the graph.
- Scope and limitation: CGGM Markov properties translate to the observed variables when all marginals are continuous, whereas discrete marginals can create additional or mismatched observed-space dependencies.Thus latent-variable edges need not correspond exactly to observed-variable conditional independences when discrete variables are present.
- Bayesian inference: MCMC sequentially resamples latent data, the free elements of the precision matrix, and the graph using edge additions or deletions.Precision-matrix updates use Metropolis–Hastings, while graph moves change parameter-space dimensionality and use reversible-jump methodology.
- Bayesian inference: When several graphs have comparable posterior probabilities, averaging over many graphs is preferred to relying only on the highest-posterior graph.This concern is especially relevant in high-dimensional data sets with few observed samples.
5. Examples.
The examples apply CGGMs to sparse Rochdale survey data and highly imbalanced NLTCS disability data, identifying dependence structures through Bayesian graph averaging. Across both analyses, CGGMs capture observed cell patterns and distinguish latent connectivity from observed-variable association.
- 5.1. The Rochdale data.: 407.04 squared error for CGGMs compares with 284.79 for the all two-way interaction model and 905.78 for Whittaker’s log-linear model across 256 cells.CGGMs performed as well as the all two-way interaction model for the largest observed count, 57.
- 5.1. The Rochdale data.: Posterior inclusion probabilities for Rochdale’s four strongest interactions were 1 for (b,d), 0.96 for (b,h), 0.98 for (e,f), and 1 for (a,g).These correspond to the strongest pairwise interactions identified by Whittaker.
- 5.2. The NLTCS functional disability data.: NLTCS connectivity differed across measures: IADL4 and IADL10 stood out in latent space, whereas IADL1, IADL2, and IADL3 had the highest cumulative observed Cramér’s V associations.IADL1 was highly connected in both spaces and was identified as key to assessing disability level.
- 5.2. The NLTCS functional disability data.: CGGMs indicate that disability measures should not be weighted equally, because reported disability counts can misrepresent the seriousness of particular disabilities.The authors caution that simply counting disabilities may be misleading when evaluating overall disability.
6. Discussion.
CGGMs extend Gaussian graphical models to mixed observed data by modeling conditional independence among latent variables, while separating dependence from univariate margins. Bayesian model averaging supports inference across graphs, particularly for sparse contingency tables, but the framework requires ordered observed values and omits parametric marginal distributions.
- CGGMs extend Gaussian graphical models to data where the observed variables are not plausibly multivariate normal.
- The models impose conditional independence constraints on latent variables in a Gaussian copula, with one latent variable corresponding to each observed variable.
- The framework does not include a parametric representation of univariate marginal distributions, although the authors identify combining it with marginal-model methods as a possible extension.
- The framework applies to observational studies containing binary, ordinal, or continuous variables, provided each observed variable has ordered possible values.
- Bayesian model averaging over latent-space graphs is useful for sparse contingency tables because closely competing graphs need not be reduced to one selected model.
Supplement: C++ implementation of copula Gaussian graphical models
The authors provide source code for the CGGM methodology, using cluster computing to run multiple Markov chains in parallel and replicate the Rochdale and NLTCS analyses.
- Source code is provided for the CGGM methodology and includes sample input files for the Rochdale and NLTCS analyses.
- The program uses cluster computing to run several Markov chains in parallel.
- The supplied code supports replication of the paper’s two reported data analyses.