Source-linked AI summary
Demixed principal component analysis of population activity in higher cortical areas reveals independent representation of task parameters
Dmitry Kobak, Wieland Brendel, Christos Constantinidis, Claudia E. Feierstein, Adam Kepecs, Zachary F. Mainen, Ranulfo Romo, Xue-Lian Qi, Naoshige Uchida, Christian K. Machens
TL;DR
Complex, mixed selectivity in higher-cortical population activity makes task-related information difficult to interpret. The paper introduces dPCA, an unbiased dimensionality-reduction method that preserves data while producing task-interpretable components. Applied to rat and monkey cortical recordings, dPCA successfully demixes activity, summarizes major population features, and reveals condition-independent components.
Problem
Diverse, mixed neural tuning in higher cortical areas obscures how population activity represents stimuli, decisions, rewards, and other task parameters.
Method
dPCA reduces population activity to latent components that are interpretable with respect to task parameters while preserving the original data as much as possible.
Results
Across rat and monkey recordings, dPCA successfully demixes population activity, summarizes previously described features in single figures, and reveals condition-independent activity.
Takeaways & Limitations
Task-related information can be linearly demixed at the population level even when it is mixed in individual neurons.
Takeaways & Limitations
Task-informed parametric approaches can miss important data structures when neural activities do not conform to assumed dependencies, such as through nonlinearities.
Abstract
from arXiv · showhide
Neurons in higher cortical areas, such as the prefrontal cortex, are known to be tuned to a variety of sensory and motor variables. The resulting diversity of neural tuning often obscures the represented information. Here we introduce a novel dimensionality reduction technique, demixed principal component analysis (dPCA), which automatically discovers and highlights the essential features in complex population activities. We reanalyze population data from the prefrontal areas of rats and monkeys performing a variety of working memory and decision-making tasks. In each case, dPCA summarizes the relevant features of the population response in a single figure. The population activity is decomposed into a few demixed components that capture most of the variance in the data and that highlight dynamic tuning of the population to various task parameters, such as stimuli, decisions, rewards, etc. Moreover, dPCA reveals strong, condition-independent components of the population activity that remain unnoticed with conventional approaches.
Introduction
Higher-cortical population recordings contain rich, mixed neural responses that complicate interpretation. dPCA addresses this challenge by reducing dimensionality while preserving information and separating activity according to task parameters.
- Motivation: Population recordings become difficult to interpret as the number of neurons and the diversity of task-related responses increase.This challenge is especially pronounced in higher-order areas such as prefrontal cortex.
- Motivation: Averaging or pre-selecting neurons can identify processed information but discards the richness of individual-cell activity.Population averages also obscure how information is represented at the neuronal level.
- Motivation: Mixed selectivity arises because neurons in higher cortical areas often encode several task parameters simultaneously.These parameters can include stimuli, rewards, and actions.
- Prior approaches: Existing dimensionality-reduction methods address spike-train or dynamical structure, while task-informed approaches can depend on assumptions about firing-rate relationships.Parametric methods may miss important structures when neural activities are nonlinear.
- Contribution: dPCA develops an unbiased dimensionality-reduction technique that seeks interpretable latent components while preserving the original data as much as possible.The method removes unnecessary orthogonality constraints from earlier methodological work.
- Contribution: Across rat and monkey datasets, dPCA successfully demixes population activity and summarizes important features in a single figure for each task.It also reveals condition-independent activity and supports visual comparison across tasks and brain areas.
Results
Across monkey PFC and rat OFC tasks, dPCA separates condition-independent, stimulus, decision, and interaction components despite mixed selectivity, while preserving the population response. It recovers task-specific dynamics and shows that representations shift across neural state space over time.
- Somatosensory working memory: 832 prefrontal neurons from two monkeys performing a somatosensory working memory task were summarized across the whole trial into four component categories.The categories were condition-independent, stimulus, decision, and stimulus-decision interaction components.
- Somatosensory working memory: Stimulus and decision information was fully demixed at the population level despite being mixed within individual neurons.The demixing was linear, making the components theoretically retrievable through synaptic weighting and dendritic integration.
- Somatosensory working memory: ∼90% of signal variance was captured by condition-independent components representing temporal modulation throughout the trial.These components were not usually analyzed or shown, and their activity extended beyond conventional descriptions of delay-period ramping.
- Somatosensory working memory: Three stimulus components shared monotonic tuning but occupied distinct S1, delay, and S2 periods, indicating that the stimulus representation rotated through firing-rate space.The components were #10 during S1, #6 during the delay, and #11 during S2.
- Technical properties: dPCA explained almost the same cumulative signal variance as standard PCA, showing that demixing imposed little variance loss.The resulting components therefore retained an accurate representation of population activity.
- Visuospatial working memory: In visuospatial working memory, dPCA again separated condition-independent, stimulus, decision, and interaction components, with stimulus and decision representations active in different periods.The overall population structure up to the second stimulus was almost identical to the somatosensory task, although component power differed.
- Visuospatial working memory: ∼75% cross-validated single-trial accuracy was achieved for match versus non-match classification after 2 s using fixed linear decoders.This approximately matched previously reported state-of-the-art classification performance.
- Olfactory tasks: In rat OFC, interaction components separated rewarded from unrewarded conditions, while condition-independent components accounted for over 60% of total variance.Decision and stimulus information were localized to distinct time periods, and decision components could also reflect movement direction or position.
Discussion
dPCA addresses the complexity of higher-order cortical responses by extracting demixed latent components from single-cell activity. Across the presented cases, it makes previously reported major population-activity features directly visible in compact summary figures.
- Discussion: dPCA extracts latent components whose individual representations avoid mixed selectivity, simplifying exploration and interpretation of higher-order cortical data.The method is presented as specifically targeted to the complexity of prefrontal and orbitofrontal responses.
- Discussion: In all presented cases, dPCA summary figures directly reveal the major aspects of population activity previously reported.
- Discussion: The discussion compares dPCA with alternative approaches and revisits what the method reveals about neural activity in higher-order areas.
I. Percentage of significantly responding neurons
Counting significantly responding neurons summarizes parameter tuning as percentages but omits important temporal, tuning-shape, population-distribution, and multi-parameter information. dPCA addresses these limitations by producing time-dependent, jointly analyzed components expressed across the population.
- I. Percentage of significantly responding neurons: The conventional analysis counts neurons significantly responsive to a parameter during a selected time period and reports the resulting percentage.The approach was used in the original publications reanalyzed here.
- I. Percentage of significantly responding neurons: dPCA preserves the time course of neural tuning through time-dependent components rather than restricting analysis to one time bin or averaged period.
- I. Percentage of significantly responding neurons: For parameters with more than two values, dPCA components can reveal tuning-curve shape, whereas percentages of significantly tuned neurons cannot.Vertical slices through stimulus-dependent components yield tuning curves, including rainbow-like stimulus components.
- I. Percentage of significantly responding neurons: Because every demixed component is expressed across the population with varying strength, dPCA avoids implying two discrete tuned and untuned subpopulations.
- I. Percentage of significantly responding neurons: dPCA analyzes multiple parameters together, avoiding confounding from unequal trial numbers across conditions.Separate tests can be confounded in unbalanced experimental designs.
- I. Percentage of significantly responding neurons: Sliding-window multi-way ANOVA addresses the time-course and confounding limitations, but retains the limitations concerning tuning-curve shape and arbitrary significance cutoffs.
II. Population averages
Averaging PSTHs across neurons selected for significant tuning can highlight some population dynamics, but it discards much of the response heterogeneity and complexity.
- II. Population averages: Averaging PSTHs over neurons selected as significantly tuned can highlight some dynamics of the population response while ignoring the full complexity of the data.The heterogeneous activity is largely averaged out, so the resulting representation does not faithfully capture the population response.
- II. Population averages: The population-average approach selects a subset of neurons rather than analyzing the full neural population.
- II. Population averages: The approach can therefore fail to faithfully represent the heterogeneous, time-varying population activity.
III. Demixing approach based on multiple regression
Multiple-regression demixing constructs axes from linear firing-rate regressions but has important scope and representation limitations. Unlike dPCA, it cannot currently handle missing data or continuous task parameters.
- III. Demixing approach based on multiple regression: Multiple regression demixes neural activity by regressing firing rates on several parameters and using the regression-coefficient vectors as demixing axes.
- III. Demixing approach based on multiple regression: The regression approach ignores condition-independent components, assumes linear neural tuning, and finds only one demixed component per parameter.
- III. Demixing approach based on multiple regression: Multiple-regression demixing cannot demix when axes for different parameters are far from orthogonal.The authors illustrate these disadvantages by applying the method to their datasets.
- III. Demixing approach based on multiple regression: The regression approach can handle missing data or continuous task parameters, whereas dPCA currently cannot.Future dPCA extensions could combine parametric dependencies with the advantages of the current method.
IV. Decoding approach
Linear classifiers decode task parameters from population firing rates by measuring time-dependent, cross-validated classification accuracy.
- Linear classifiers predict task parameters from population firing rates, with cross-validated classification accuracy quantifying neural tuning.Separate classifiers can be applied in each time bin to obtain time-dependent accuracy.
V. Linear discriminant analysis
LDA incorporates class labels into dimensionality reduction by maximizing class separation, whereas dPCA also handles multiple parameters and reconstructs the original data. The broader analysis finds strong condition-independent activity, time-varying parameter representations, near-orthogonal encoding axes, and practical methodological limits.
- V. Linear discriminant analysis: LDA maximizes between-class while minimizing within-class variance, unlike PCA, which maximizes total variance without class labels.
- Insights obtained from applying dPCA to the four datasets: 70–90% of total task-locked variance is captured by condition-independent components, partly reflecting overall firing-rate increases during task periods.
- Insights obtained from applying dPCA to the four datasets: Parameter tuning shifts between components over time, with separate stimulus components appearing during stimulus, delay, and later stimulus periods.
- Insights obtained from applying dPCA to the four datasets: Only 22 of 420 encoding-axis pairs were non-orthogonal, indicating that condition-dependent task parameters were mostly represented independently.
- Limitations: dPCA is limited to discrete parameters with complete condition combinations, typically needs about 100 neurons, and here uses trial-averaged PSTHs.
Materials and Methods
The study reanalyzed neural recordings from monkey and rat task datasets using condition-averaged PSTHs and dPCA. The analysis separated stimulus, decision, interaction, and condition-independent activity while evaluating demixing quality.
- Materials and Methods: Recordings came from multiple sessions in monkey prefrontal and rat orbitofrontal datasets, so most neurons were not recorded simultaneously.
- Materials and Methods: Each trial was labeled by stimulus and decision, with reward omitted because deterministic protocols allowed it to be inferred from those parameters.
- Materials and Methods: Spike trains were Gaussian-filtered with σ = 50 ms and averaged across trials within each condition to produce smoothed PSTHs.
- Materials and Methods: The centered data were decomposed into stimulus, decision, interaction, and condition-independent marginalizations for dPCA analysis.
- Materials and Methods: Time-period separation produced distinct stimulus-component time courses without noticeable loss of explained variance and highlighted rotating parameter representations.
- Materials and Methods: 0.97±0.02 was the average demixing index for the first 15 dPCA components versus 0.76±0.16 for PCA in the somatosensory dataset.The difference was significant at p = 0.00016 using a Mann-Whitney-Wilcoxon ranksum test.
Supplementary Information
The supplementary material includes the paper’s author and contributor information.
- Supplementary Information: The listed contributors include D Kobak, W Brendel, C Constantinidis, C Feierstein, A Kepecs, Z Mainen, R Romo, X-L Qi, N Uchida, and C Machens.
S1 Supplementary Figures
The supplementary analyses illustrate neural-response heterogeneity, compare standard PCA with dPCA, and test dPCA across classification, datasets, clustering, alignment, and regularization choices.
- Heterogeneity: Forty randomly chosen monkey neurons show a large number of neurons with mixed selectivity.The neurons had average firing rates between 20 Hz and 40 Hz.
- PCA versus dPCA: Standard PCA components exhibit mixed selectivity, whereas dPCA components have higher demixing ratios.For the first 15 components, the average demixing index was 0.76 ± 0.16 for PCA versus 0.97±0.02 for dPCA, with p = 0.00016.
- Classification: Classification accuracy was evaluated across datasets using the first three stimulus, decision, and interaction dPCs against shuffled chance distributions.Black lines show linear-classifier accuracies, while shaded gray regions show chance distributions from 100 shuffling iterations.
- Dataset analyses: Supplementary analyses applied dPCA to monkey visuospatial memory data before training and rat olfactory discrimination data without re-stretching.The pretraining monkey analysis used 673 neurons, and the non-re-stretched rat analysis produced qualitatively the same results as the re-stretched analysis.
- Clustering and comparison: Supplementary figures examined neuron clustering from dPCA encoding weights and compared the demixing approach of Mante et al. across datasets.Clustering used density peaks or Gaussian mixture models applied to encoder-weight representations.
- Procedures and robustness: The supplementary material documents trial re-stretching, marginalization, time-period separation, and regularization analyses for dPCA.Without regularization, the olfactory-task components overfit and cross-validation error was large.
S2 Supplementary notes
The supplementary notes define marginalized averages as a complete, variance-partitioning segregation of parameter-dependent data and relate the construction to categorical linear models and ANOVA.
- S2.1.1 Intuition: A function of time and stimulus is decomposed into a mean, time component, stimulus component, and interaction component.The interaction captures dependence on both parameters after removing the mean and single-parameter contributions.
- S2.1.1 Intuition: Marginalized averages are constructed by averaging over parameters that are not being represented in each component.The mean averages over all parameters, while single-parameter terms estimate the remainder with the other parameter undetermined.
- S2.1.1 Intuition: The segregation is complete, so it does not lose information about the original function.Each marginalized average also averages to zero over any parameter on which it depends.
- S2.1.2 Analysis of Variance: Marginalized components are covariance-independent, allowing the variance of the original function to equal the sum of marginalized variances.For time, stimulus, and interaction terms, Var(x(t, s)) = Var(z(t)) + Var(z(s)) + Var(z(t, s)).
- S2.1.2 Analysis of Variance: The marginalization procedure is exactly equivalent to a linear model with categorical predictors used in ANOVA.With one value per parameter combination, ANOVA has little power for significance testing, but the underlying linear model remains equivalent to marginalization.
- S2.1.3 Generalization: The framework generalizes from scalar functions and two parameters to vector-valued data and arbitrary subsets of parameters.A multidimensional data matrix uses the first index for observables and subsequent indices for task parameters.
S2.2 dPCA optimization
dPCA optimizes low-dimensional reconstructions of marginalized neural data while controlling rank and regularization, using cross-validation to select stable solutions.
- Data representation: dPCA represents multidimensional neural data with observables in the first index and task parameters in subsequent indices.For example, Xnts denotes the mean response of neuron n at time t and stimulus s.
- Optimization objective: The dPCA objective projects full data into low-dimensional dPCs and reconstructs a specified marginalized dataset.The objective favors variance in the selected marginalization and penalizes variance from other marginalizations.
- Marginalization choices: Decision and decision-time effects can be pooled into one marginalization when decision variance is intrinsically time-dependent.The paper gives Xd + Xtd → Xtd as an example of this pooling.
- Variance preservation: Across the analyzed neural datasets, demixing alone did not strongly decrease explained variance, so subsequent analyses used the dPCA objective alone.The optimization is paired with rank constraints that select components carrying the most variance.
- Regularization: Without regularization, decoder directions can overfit low-variance kernel directions and amplify noise on validation data.The problem is especially relevant when neuron count is similar to the number of independent data points.
- Regularization: Cross-validation selects the regularization parameter by holding out one random trial per neuron and choosing the minimum reconstruction error.The procedure used 10 train-test splittings, and two error formulas yielded the same optimal λ.
S2.4 Relation to Linear Discriminant Analysis
The section compares dPCA with LDA as approaches to demixing neural population activity. LDA can produce similar components, but dPCA additionally represents the original data and quantifies explained variance.
- LDA formulation: LDA finds discriminant axes that maximize between-class variance while minimizing within-class variance.
- LDA formulation: LDA is one-way, typically uses few classes with many data points, and is therefore difficult to apply to these neural datasets.In the somatosensory working memory example, the data form 3006 classes with only 2 points per class.
- Comparison with dPCA: Regularized LDA can produce components similar to dPCA, but estimating within-class covariance in the 832-dimensional space causes severe over-fitting without regularization.The regularized estimator replaces ΣW with (1 −λ)ΣW + λI.
- Comparison with dPCA: For a single decoder and encoder, the dPCA objective is similar to LDA because it maximizes between-class and minimizes within-class variance.
- Comparison with dPCA: dPCA is considered more principled because it uses both encoding and decoding axes, represents the original dataset, and provides explained-variance measures, unlike LDA.
- Sampling artifacts: Pooling neurons across sessions with small response-onset shifts can generate additional PCA components that resemble temporal derivatives of the source component.This provides an explanation for derivative-like decision components observed in the working memory task.
S2.6 Mathematical proofs
The mathematical proofs establish that marginalization is complete, linearly parameter-independent, and unique. These properties support the covariance decomposition used by dPCA.
- Marginalization properties: Marginalized averages form a complete segregation of the original population activity.
- Reformulation and uniqueness: The closed-form marginalized averages satisfy completeness and linear parameter independence, completing the proof of the construction.
- Marginalization properties: Each marginalized average has zero mean when averaged over any parameter it contains, establishing linear parameter independence.
- Marginalization properties: Different marginalized averages are independent under the defined marginalization scheme.
- Reformulation and uniqueness: The marginalized averages can be reformulated using parameter-specific averages without depending on the order of evaluation.
- Reformulation and uniqueness: Any segregation satisfying completeness and linear parameter independence is unique.