Source-linked AI summary
Pairwise maximum entropy models for studying large biological systems: when they can and when they can't work
Yasser Roudi, Sheila Nirenberg, Peter Latham
TL;DR
The paper addresses whether pairwise models inferred from small biological subsystems can reliably describe larger systems. It analyzes this extrapolation problem for pairwise models and finds that, generally, sufficiently small systems have almost no predictive value for larger ones, while predictive behavior depends on a crossover in system size.
Problem
The paper asks whether pairwise models supported by studies of small subsystems generalize to realistic biological systems.
Method
The paper studies the extrapolation problem for pairwise models using analytical results and simulations, comparing pairwise approximations with true distributions.
Results
In the perturbative regime, pairwise models have almost no predictive value for what will happen in a large system.
Takeaways & Limitations
Pairwise-model predictions for whole biological systems require determining whether the system lies below or above the crossover point.
Takeaways & Limitations
The analysis includes a catch-22 in which small time bins avoid ignored correlations but are precisely the regime where those correlations are ignored.
Abstract
from arXiv · showhide
One of the most critical problems we face in the study of biological systems is building accurate statistical descriptions of them. This problem has been particularly challenging because biological systems typically contain large numbers of interacting elements, which precludes the use of standard brute force approaches. Recently, though, several groups have reported that there may be an alternate strategy. The reports show that reliable statistical models can be built without knowledge of all the interactions in a system; instead, pairwise interactions can suffice. These findings, however, are based on the analysis of small subsystems. Here we ask whether the observations will generalize to systems of realistic size, that is, whether pairwise models will provide reliable descriptions of true biological systems. Our results show that, in most cases, they will not. The reason is that there is a crossover in the predictive power of pairwise models: If the size of the subsystem is below the crossover point, then the results have no predictive power for large systems. If the size is above the crossover point, the results do have predictive power. This work thus provides a general framework for determining the extent to which pairwise models can be used to predict the behavior of whole biological systems. Applied to neural data, the size of most systems studied so far is below the crossover point.
1 Introduction
Accurate statistical descriptions of biological systems are difficult because they involve many interacting elements, enormous configuration spaces, and limited data. Pairwise models offer a possible simplification, but evidence from small subsystems raises whether they predict whole systems; the paper finds that, generally, they do not.
- Biological systems contain many interacting elements and exponentially many possible configurations, making statistical description difficult.
- Neuroscience illustrates the data constraint: current technology records simultaneously from about 100 neurons out of approximately 100 billion in the human brain.
- Pairwise models estimate the full distribution from probabilities over pairs, reducing the number of quantities from exponential to quadratic in system size.
- Evidence for pairwise-model efficacy has come from small subsystems where the true distribution could be measured and directly compared with the model.
- The central question is whether pairwise models that predict a subset's distribution also predict the distribution of the whole system.
- For a biologically relevant class of systems, the answer is generally no, and the paper explains why analytically and with simulations.
2 Results
The paper tests whether pairwise models fitted to small neural subsystems can predict larger populations, finding a crossover between uninformative and potentially predictive regimes. In the perturbative regime, goodness-of-fit worsens predictably with population size, while behavior beyond it must be assessed empirically.
- The extrapolation problem: Pairwise models match the true distribution's means and pairwise correlations, but the paper asks whether that agreement persists as population size N grows.The comparison uses p_pair and p_true, with Δ_N measuring their distance: zero indicates equality and values near one indicate a large difference.
- The extrapolation problem: For small N, Δ_N increases linearly with N, a generic perturbative-regime behavior independent of the true distribution.The perturbative scaling is derived when νδt ≪1 and is described as independent of details of p_true.
- The extrapolation problem: Linear growth of Δ_N provides no information about the large-N distribution, whereas saturation beyond the perturbative regime can make extrapolation meaningful.The paper therefore requires data beyond the perturbative regime before using small-subsystem results to infer large-system behavior.
- The extrapolation problem: A system is perturbative when its neuron count is small compared with 100, but small time bins create a conflict between independent bins and ignored temporal correlations.Small bins support the independent-bin approximation in one sense, yet temporal correlations cannot validly be ignored in that same regime.
- The extrapolation problem: The crossover size N_c separates an uninformative regime, N ≪ N_c, from a regime, N ≫ N_c, where full-system inferences may be possible.The paper interprets N_c as the point where pairwise-model behavior changes between perturbative and potentially predictive regimes.
- The dangers of extrapolation: Naive entropy and divergence extrapolations can become physically inconsistent, including decreasing true entropy, negative entropy, or Δ_N > 1.The characteristic quantities N_ind and N_pair may occur beyond the range where the extrapolations remain meaningful, underscoring the weak connection between perturbative and large-N behavior.
3 Is there anything wrong with using small time bins?
Using small time bins can make a pairwise model appear increasingly accurate under an independent-bin approximation while worsening its description of the true temporally correlated spike-train distribution. When bins are smaller than the spike-train correlation time, temporal correlations introduce more error than ignoring correlations across neurons.
- The apparent improvement from shrinking time bins is misleading because the independent-time-bin approximation worsens as adjacent bins become highly correlated.The approximation is good when bins exceed the spike-train correlation time but poor when bins are smaller.
- In this small-bin regime, ignoring temporal correlations causes more error than ignoring correlations across neurons.Thus, a model that captures pairwise neuronal interactions but assumes independent time bins can be substantially misrepresentative.
- The full time-series distribution is represented over M bins of size δt, with M = T/δt for spike trains of length T.The derivation combines the full distribution with the independent-bin and pairwise-model assumptions to quantify the discrepancy.
- The analysis compares the error from assuming independence across time with the error from assuming independence across neurons.This comparison quantifies when temporal correlations matter more than cross-neuron correlations.
- For bins smaller than the spike-train correlation time, the pairwise model does a very bad job describing full, temporally correlated spike trains.The analysis identifies this regime as the case in which the temporal-independence assumption fails most strongly.
4 Discussion
The discussion argues that pairwise models can fit small, temporally independent subsystems while failing to predict large populations, unless the system lies beyond the perturbative regime. Neural applications face an additional tradeoff: small time bins improve pairwise fits but ignore temporal correlations, whereas larger bins weaken that guarantee.
- The framework asks whether pairwise models applied to small subsystems are reliable models of the full biological system.
- Small subsystems can match temporally independent distributions without showing that the same pairwise model will work for larger time bins or populations.
- Pairwise models in the perturbative regime have almost no predictive value for large populations, making extrapolation from small to large systems dangerous.
- Nearest-neighbor models may extrapolate because their interaction range is fixed: the number of nearest neighbors does not grow with population size.
- 4.1 Time bins and population size: Temporal independence is especially problematic for natural stimuli because their long correlation times are inherited by spike trains.
- 4.2 Other biological problems that have been approached with pairwise models, e.g, protein folding: Protein studies differ from many neural studies because the proteins usually considered are outside the perturbative regime, leaving evidence for pairwise models promising.
- The authors developed tests for identifying the perturbative regime and propose using them to analyze experiments and design studies.
- 4.1 Time bins and population size: For neural data, small time bins favor pairwise fits but ignore temporal correlations, while larger bins reduce adjacent-bin correlations without guaranteeing that pairwise models work.
5.1 The behavior of the true entropy in the large N limit
This section analyzes how the entropy of a neural population changes as neurons are added. It relates the entropy increment to single-neuron entropy minus mutual information and concludes that entropy is linear in population size when mutual information saturates.
- The entropy increment satisfies S_N+1 − S_N = S_1 − I(1; N), linking population entropy growth to mutual information.
- When activity in N neurons does not constrain neuron N + 1, single-neuron entropy exceeds mutual information and entropy increases with N.
- Population entropy is linear in N if mutual information between one neuron and the remaining population saturates as N grows.
- Synaptic failures make saturation of mutual information a reasonable assumption for neural networks.
5.2 Perturbative Expansion
The perturbative analysis expands the KL divergence between the true distribution and simpler approximations in powers of Nδ. It uses moment matching and the Sarmanov–Lancaster expansion to identify the leading corrections and their scaling.
- The main quantitative result is an expression for the KL divergence between the true distribution and its approximations.
- The analysis expands the independent and pairwise distributions perturbatively in powers of Nδ.
- The Sarmanov–Lancaster expansion supplies relationships between moments and model parameters that simplify the KL-divergence calculation.
- Moment matching is convenient because distribution parameters correspond closely to moments, allowing matched first and second moments to define the pairwise model.
- Terms involving unmatched higher-order structure survive after lower-order terms vanish, producing the leading correction to the pairwise approximation.
- The perturbative expressions retain corrections of O(Nδ), so the derived local fields are controlled only to the stated order.
5.3 Generating synthetic data
The synthetic-data procedure generates a base distribution by sampling higher-order and pairwise parameters, then choosing local fields to obtain target firing rates. The construction uses a 15-neuron base population.
- Synthetic data depend on three parameter sets, including local fields and higher-order interaction parameters.
- The base distribution contains N* = 15 neurons, from which local fields are selected using Eq. (15a).
- In the perturbative regime, the generated distribution has firing rates approximately equal to the target rates r*i.
- The higher-order interaction parameters are drawn from Gaussian distributions with means 0.05 and 0.02 and standard deviations 0.8 and 0.5, respectively.
5.4 Bin size and the correlation coefficients
Bin size affects normalized correlation coefficients according to the correlogram’s shape, which determines how ΔN scales with bin size. A finite correlation time guarantees a sufficiently small-bin regime with linear ΔN dependence.
- If δt exceeds the correlation time, g_ind and g_pair depend on δt, so ΔN is no longer strictly linear in bin size.
- Because the correlation time is finite, there is always a bin-size range below which the linear relationship ΔN ∼ δt is guaranteed.
- For smooth correlograms, normalized second- and third-order coefficients approach constants as δt becomes small.
- For sharp central peaks wider than the bin, second-order coefficients scale as 1/δt and third-order coefficients as 1/δt^2.
- The correlogram’s shape strongly determines how normalized correlation coefficients, and therefore ΔN, depend on bin size.Smooth correlograms approach constant coefficients as δt becomes small, whereas sharp central peaks produce inverse-bin scaling.
- In the sharp-peak regime with large bins relative to correlation time, g_ind scales as Nδt and g_pair as N^2δt, making ΔN approximately independent of δt.
5.5 Assessing goodness of fit for independence across time assumption
The goodness-of-fit analysis compares temporally correlated data with an independence-across-time assumption. As the number of time bins increases, the assumed temporal independence can make the derived expression a lower bound on γ.
- The analysis starts from the definition of γ and rewrites its numerator using the true distribution over temporally correlated responses.
- The derivation decomposes entropy terms using a distribution that preserves within-neuron temporal correlations while making neurons independent.
- The pairwise maximum entropy model permits replacing the true distribution with the maximum entropy distribution because their first and second moments match.
- As M increases, S_true/M becomes increasingly different from S_maxent because the former assumes independence across time bins.
- Equation (22) should therefore be treated as a lower bound on γ.