Source-linked AI summary

Multilevel functional principal component analysis

Chong-Zhi Di, Ciprian M. Crainiceanu, Brian S. Caffo, Naresh M. Punjabi

arXiv:0906.1457v1stat.AP

TL;DR

Large, repeated EEG datasets such as the SHHS pose challenges from dimensionality, measurement error, and within- and between-subject variability, while existing FPCA methods do not address multilevel observations. The paper develops MFPCA to decompose hierarchical functional variation and estimate subject- and visit-level components. In SHHS, it identifies dominant slow-wave sleep variation and associations between the first principal component and hypertension, while providing a framework for subject-specific prediction and downstream analysis.

  • Problem

    Existing FPCA methods are not designed for multilevel data with subject-level curves observed at several visits, limiting analysis of hierarchical sleep-EEG functions.

  • Method

    MFPCA decomposes multilevel functional variability into intra- and inter-subject components, using mixed-model methods to estimate principal component scores.

  • Results

    The first principal component was associated with lower odds of hypertension, and this effect persisted after including the respiratory disturbance index.

  • Takeaways & Limitations

    MFPCA provides a parsimonious hierarchical representation, shrinkage score estimators, subject-specific predictions with uncertainty, and functional correlation measures.

  • Takeaways & Limitations

    The number of eigenfunctions lacks a theoretically satisfactory selection solution, so practical choices include cross-validation or AIC.

Abstract

from arXiv · show

The Sleep Heart Health Study (SHHS) is a comprehensive landmark study of sleep and its impacts on health outcomes. A primary metric of the SHHS is the in-home polysomnogram, which includes two electroencephalographic (EEG) channels for each subject, at two visits. The volume and importance of this data presents enormous challenges for analysis. To address these challenges, we introduce multilevel functional principal component analysis (MFPCA), a novel statistical methodology designed to extract core intra- and inter-subject geometric components of multilevel functional data. Though motivated by the SHHS, the proposed methodology is generally applicable, with potential relevance to many modern scientific studies of hierarchical or longitudinal functional outcomes. Notably, using MFPCA, we identify and quantify associations between EEG activity during sleep and adverse cardiovascular outcomes.

1. Introduction.

The SHHS provides extensive repeated sleep-EEG data for studying associations between sleep and health, but its scale, hierarchy, and variability create substantial analytical challenges. The paper introduces MFPCA to extract and represent multilevel functional signals for subsequent outcome analyses.

  • The SHHS recruited 6,441 participants and collected in-home polysomnograms at baseline and approximately five years later.
  • More than 3,000 participants had repeat measurements, while the study generated over 1.5 terabytes of unprocessed EEG data.
  • Data processing and statistical challenges: EEG signals were transformed into adjacent 30-second normalized δ-power intervals to reduce dimensionality and target frequency bands relevant to sleep research.
  • Data processing and statistical challenges: The resulting functional data exhibit within- and between-subject heterogeneity, within-subject clustering, measurement error, and repeated visits.
  • Data processing and statistical challenges: MFPCA decomposes variability according to the data hierarchy and extracts signals that can replace high-dimensional functions in later analyses.
  • Methods for the analysis of functional data: Existing FPCA methods were not designed for multilevel data with subject-level curves observed across several visits.

2. Models and framework.

MFPCA extends FPCA to clustered functional data by decomposing subject- and visit-level variation into separate principal components. It estimates these components through covariance decomposition and eigenanalysis while accommodating unbalanced observation designs and measurement error.

  • Multilevel FPCA: MFPCA combines FPCA with multilevel mixed models to decompose subject-level and visit-level functional variation.The framework represents level 1 and level 2 functions with separate Karhunen–Loève expansions and principal component scores.
  • Multilevel functional model: The model treats overall and visit-specific means as fixed functions, with subject- and visit-specific deviations modeled as uncorrelated stochastic processes.The methodology does not require every subject to be observed at every visit.
  • Model assumptions: Level 1 and level 2 eigenfunctions are each orthonormal bases but need not be mutually orthogonal.This permits distinct multilevel components while allowing their bases to overlap.
  • Estimation: The estimation algorithm first estimates mean and covariance functions, then applies eigenanalysis separately to between-subject and within-subject covariance functions before estimating scores.Within-subject covariance is formed as the difference between total and between-subject covariance.
  • Measurement error: The framework accommodates noisy functional observations by modeling observed data as latent multilevel functions plus white-noise error.The error process is assumed to have variance σ^2.

3. Principal component scores.

Multilevel functional principal component scores are estimated through full and projection models because level-specific eigenfunction bases are not mutually orthogonal. The projection model reduces computational burden, while model choice depends on data density and the relationship between level-specific eigenfunctions.

  • Estimating scores is harder in multilevel functional data because the two functional bases are not mutually orthogonal.This motivates the proposed approach for estimating scores with and without measurement error.
  • The full score model PC-F is a linear mixed model whose random effects are the subject- and visit-level principal component scores.Gaussian score distributions are assumed for convenience, and measurement error is represented by εij(t) when present.
  • MCMC estimates principal component scores while providing posterior means and full posterior distributions; BLUP estimation is also possible.The fixed functional effects, eigenvalues, and eigenfunctions are treated as estimated quantities when fitting the score model.
  • In the SHHS application, model (3.1) involves more than 2.5 million observations, motivating a computational solution based on projection.With I subjects, J visits, T grid points, and N1 and N2 retained dimensions, the model has IJT observations and I(N1 + JN2) random effects.
  • The projection model PC-P summarizes each function with low-dimensional vectors and is less computationally intensive than PC-F.For dense functional data the methods yield similar results, whereas PC-F typically performs better for sparse functional data.
  • When level-specific eigenfunctions are identical, projection identifies only the sum of subject- and visit-level scores, so estimation quality depends on their inner-product matrix C.The projection model incorporates measurement error and residual variation from discarded dimensions in its integrated residuals.

4. Simulation studies.

Extensive simulations assess MFPCA under orthogonal and nonorthogonal multilevel eigenfunction structures and increasing noise. The results show that smooth MFPCA reduces eigenvalue bias and that the methodology recovers signals and eigenfunction shapes, while score estimation worsens with noise and nonorthogonality.

  • Simulation design: The simulations used 200 subjects, two visits per subject, four components at each level, and 1000 replicates across noise scenarios σ = 0, 1, and 2.The study focused primarily on the more realistic case of mutually nonorthogonal bases, while orthogonal-basis simulations assessed correlation effects on score estimation.
  • Signal recovery: With 200 samples, MFPCA recovered true subject-level signals even when noise was large (σ = 2).Increasing noise made subject-level patterns less recognizable, but the methodology still recovered the underlying signals.
  • Eigenvalues and eigenfunctions: Unsmooth MFPCA recovered eigenvalues without bias at σ = 0, but bias increased for smaller components as noise rose, especially at σ = 2.The third and fourth components showed the most pronounced bias under moderate and large noise.
  • Eigenvalues and eigenfunctions: Smooth MFPCA practically removed the eigenvalue bias observed with the unsmooth algorithm under noisy conditions.Figure 4 compares estimated eigenvalues with the true eigenvalues across the noise scenarios.
  • Eigenvalues and eigenfunctions: The methods separated level 1 and level 2 variation and captured individual eigenfunction shapes, including when noise was large.The no-noise evaluation used 20 randomly selected simulations, while the broader noisy-data assessment found that both smooth and unsmooth methods retained overall shape.
  • Principal component scores: Score RMSEs increased with noise and were generally larger for nonorthogonal bases; PC-F performed slightly better than PC-P.In Case 2, an inner product of 0.96 between two eigenfunctions corresponded to a 16 degree angle and coincided with relatively large RMSE for those components.

5. The analysis of sleep data from the SHHS.

MFPCA was applied to SHHS sleep EEG δ-power data to characterize between-subject and within-subject functional variation and relate subject-level components to hypertension. Subject-level variation was concentrated in a few components, while visit-level variation was distributed more broadly; the first subject-level component was associated with lower odds of hypertension.

  • SHHS data and preprocessing: 3,201 subjects with complete data were analyzed over the first 4 hours of sleep, comparing raw and smoothed MFPCA approaches; only the raw-data smooth MFPCA results were presented because results were similar.The selected sample had sleep duration exceeding 4 hours at both visits.
  • Subject-level variation: 80.6% of subject-level variation was explained by the first eigenvalue, while the second and third explained 7.7% and 3.7%, together exceeding 91%.The first component represented an overall shift in sleep EEG δ-power, while the second represented an early-night shift.
  • Visit-level variation: 90% of visit-level variability required the first 14 principal components, with 50% explained by the first 4, reflecting substantial within-subject heterogeneity.The first visit-level component represented a shift in average δ-power that was more pronounced after the first hour; later eigenfunctions were typically periodic.
  • Subject-level clustering: 21.3% of sleep EEG δ-power variability was attributable to subject-level variability, equivalent to an average same-subject correlation of 0.213.A bootstrap confidence interval for ρ was [0.210,0.236], covering the estimated 0.213.
  • PC-score distributions: Females tended to have higher average percent sleep EEG δ-power than males, while δ-power tended to decrease with age and respiratory disturbance index; BMI showed no clear association.Former and current smokers appeared to have less percent sleep EEG δ-power than nonsmokers, with a smaller percentage among current smokers.
  • Hypertension associations: The first subject-level PC score was strongly and negatively associated with hypertension across all six adjustment models.With full confounder adjustment, the reported odds ratio was e−0.86 = 0.423 per unit increase in the first PC score; the standardized odds ratio was e−0.205 = 0.815 in the sex-and-smoking model.

6. Discussion.

MFPCA addresses the SHHS data’s large dimensionality, measurement error, and within- and between-subject variability by decomposing hierarchical functional variation. The discussion also identifies methodological benefits, the necessity of multilevel analysis in this application, and limitations requiring further work.

  • 6. Discussion: MFPCA provides a robust, computationally feasible, parsimonious decomposition of subject-level and visit/subject-level functional characteristics.These characteristics would be difficult to detect by inspecting plots across thousands of subjects.
  • 6. Discussion: The SHHS application indicates that multilevel analysis is necessary to distinguish subject-specific from subject/visit-specific variability.Averaging replicate functions and applying single-level FPCA identified 16 or 17 directions, whereas MFPCA identified 3 directions for long-term subject averages.
  • 6. Discussion: Multilevel analysis supplies hierarchical decomposition, shrinkage score estimators, subject-specific predictions with uncertainty, and functional-correlation measures.The shrinkage estimators typically improve mean square error relative to raw estimates and can yield unbiased functional-regression effect estimators.
  • 6. Discussion: The methodology approximates finite-dimensional functional spaces with finite, small-dimensional subspaces, making the approach inherently parametric after conditioning on those subspaces.Allowing the number of principal components to increase appropriately with sample size can make the method nonparametric and target the true process.
  • 6. Discussion: Further work is needed on positive-definite efficient covariance estimation, functional-space dimension selection, and nonlinear second-level analyses using estimated scores.Ignoring score variability in nonlinear second-level analyses may produce biased results.

SUPPLEMENTARY MATERIAL

The supplementary material documents additional methodological and application details for MFPCA. It covers component-number selection, Bayesian score estimation, simulation and SHHS results, and residual variance-covariance details.

  • SUPPLEMENTARY MATERIAL: The supplement assesses criteria for selecting the number of principal components and provides Bayesian MCMC details for estimating principal component scores.It also reports additional simulation and SHHS application results.
  • SUPPLEMENTARY MATERIAL: The supplement provides technical details on the variance and covariance of residuals from the projection model.
  • SUPPLEMENTARY MATERIAL: Additional simulation and SHHS application results are included in the supplementary material.
Loading 0906.1457v1…