Source-linked AI summary
Inverse statistical problems: from the inverse Ising problem to data science
H. Chau Nguyen, Riccardo Zecchina, Johannes Berg
TL;DR
The paper reviews how inverse statistical physics infers model parameters from observed data, focusing on the inverse Ising problem and its applications. It synthesizes equilibrium and non-equilibrium approaches, emphasizing pseudolikelihood among approximate reconstruction methods. The review connects these methods to applications including neural connectivity, protein structure, and gene-regulatory inference.
Problem
Inverse statistical physics must recover microscopic model parameters from observations, a challenge arising prominently in high-throughput biological data and related many-body systems.
Method
The review synthesizes inverse Ising applications and reconstruction methods across equilibrium and non-equilibrium settings, including likelihood, variational, tree-based, regression, and pseudolikelihood approaches.
Results
Pseudolikelihood provides a reconstruction close to the optimal equilibrium approach and is asymptotically close in sample number, while regression links successful equilibrium and non-equilibrium methods.
Takeaways & Limitations
Inverse Ising methods provide a common framework for inferring interactions or model parameters from observed statistics across biological and physical applications.
Abstract
from arXiv · showhide
Inverse problems in statistical physics are motivated by the challenges of `big data' in different fields, in particular high-throughput experiments in biology. In inverse problems, the usual procedure of statistical physics needs to be reversed: Instead of calculating observables on the basis of model parameters, we seek to infer parameters of a model based on observations. In this review, we focus on the inverse Ising problem and closely related problems, namely how to infer the coupling strengths between spins given observed spin correlations, magnetisations, or other data. We review applications of the inverse Ising problem, including the reconstruction of neural connections, protein structure determination, and the inference of gene regulatory networks. For the inverse Ising problem in equilibrium, a number of controlled and uncontrolled approximate solutions have been developed in the statistical mechanics community. A particularly strong method, pseudolikelihood, stems from statistics. We also review the inverse Ising problem in the non-equilibrium case, where the model parameters must be reconstructed based on non-equilibrium statistics.
I. INTRODUCTION AND APPLICATIONS
The inverse Ising problem reverses the usual statistical-physics workflow by inferring microscopic interactions and fields from observed system behavior. This review connects the problem to biological and physical applications, while distinguishing equilibrium and non-equilibrium settings.
- Motivation: Inverse statistical problems infer microscopic model parameters from observations rather than deriving observables from known laws.The inverse Ising problem specifically learns spin interactions and fields from magnetisations, correlations, configurations, or other observables.
- Model: Binary variables with pairwise couplings and local fields provide the basic Ising representation of many systems.The spins can represent degrees of freedom such as neural states, while the couplings need not form a regular lattice or be uniformly positive.
- Model: The inverse task is to choose couplings and fields so model moments match observed magnetisations and correlations.When first and second moments are insufficient, maximum-entropy models can incorporate higher-order interactions such as three-spin terms.
- Computational difficulty: Inferring Ising parameters is algorithmically hard because changing one coupling can affect many correlations and inference regimes may be easy, hard, solvable, or unsolvable.These threshold phenomena motivate a statistical mechanics of inverse problems whose phase space consists of quantities normally treated as fixed parameters.
- Applications: Applications include reconstructing neural and gene-regulatory networks, determining protein structures, designing many-body systems, and modelling biological responses.The review emphasizes high-throughput biological data, including neural states, DNA or protein sequences, and RNA concentrations.
- Equilibrium and non-equilibrium settings: Equilibrium inverse Ising models use symmetric couplings and Boltzmann or maximum-entropy descriptions, whereas non-equilibrium systems require reconstruction from non-equilibrium statistics.The review surveys approximate solutions, including methods developed in statistical mechanics and pseudolikelihood from statistics.
C. Protein structure determination
Protein structure determination can use evolutionary sequence variation to infer residue couplings and predict three-dimensional contacts. Pairwise models reduce correlation-driven false contacts while also reproducing some higher-order sequence statistics.
- Motivation: Protein structure determination is challenging because proteins are numerous and their three-dimensional structures determine physical, chemical, and interaction properties.Experimental structure determination requires crystallization and X-ray diffraction, while sequence-based computational modeling can require extensive resources.
- Evolutionary information: Evolutionary correlations across related protein sequences provide information about contacts between residues distant along the linear sequence.Interacting amino acids can co-vary across species while preserving the folded structure.
- Contact prediction: Direct-information predictions show excellent agreement with experimentally determined protein contacts, whereas raw correlations contain many false contact predictions.In the SigmaE example, the fraction of false predictions is considerably reduced for Potts-model couplings; contact maps compare experimental contacts, mutual information, and direct information.
- Inverse modeling: Inverse Potts models infer pairwise amino-acid couplings from observed frequencies and correlations, then use strong couplings to predict proximal sequence positions.Each position has 21 possible states, including the 20 amino acids and a gap; couplings are summarized using matrix norms or direct information.
- Model scope: Pairwise-interaction models can predict non-trivial three-point amino-acid correlations, although adding higher-order interactions may risk overfitting.The approach has also been applied to protein families, transcription-factor binding sequences, and IgM antibody sequences.
- Related applications: Pairwise sequence models have been used to infer fitness landscapes from HIV protein sequence frequencies, while related maximum-entropy methods extend structure analysis beyond evolutionary data.Applications include chromosome conformation capture data and structure or function modeling across protein families.
F. Interactions between species and between individuals
Inverse Ising methods model interactions among organisms, individuals, or market variables from observed collective data. Their interpretation depends on coupling symmetry, which distinguishes equilibrium from non-equilibrium settings.
- Interactions between species: Ecological interaction networks can be inferred from co-culturing and perturbation experiments involving species growth and removal.The underlying relationships include predation and metabolic effects among microorganisms.
- Interactions between individuals: Pairwise models describe collective behavior in bird flocks by coupling the directional velocities of individuals.Because birds change neighbors during flight, couplings may depend on inter-individual distance or entire trajectories.
- Financial interactions: Stock-market analyses represent bullish or non-bullish conditions for company shares as binary spins with pairwise interactions.The model is applied to correlations among price changes across companies.
- Coupling symmetry: Symmetric couplings correspond to maximum-entropy models and detailed balance, whereas asymmetric couplings lead to non-equilibrium steady states.This symmetry distinction organizes applications and determines whether the steady state is described by a Boltzmann distribution.
- Inverse Ising framework: The inverse Ising problem infers pairwise couplings and local fields from observed spin configurations, with equilibrium statistics determined by the Boltzmann distribution.Temperature is absorbed into the couplings and fields by setting k_BT = 1, while only βJ_ij and βh_i are inferable.
- Maximum likelihood: Maximum likelihood uses observed magnetizations and pair correlations as sufficient statistics for estimating Ising-model parameters.With finite data, estimates may be biased; in the large-sample limit, maximum likelihood is consistent and has no consistent estimator with smaller mean-squared error.
1. Exact maximization of the likelihood
Exact likelihood maximization recovers parameters by matching model and sample magnetisations and correlations, but evaluating the required averages over 2^N configurations is generally infeasible.
- Maximum-likelihood estimates occur when Boltzmann expectations of pair correlations and magnetisations match their sample averages.
- The log-likelihood is concave, so convex optimization can in principle find its maximum; Boltzmann machine learning uses gradient descent for this purpose.
- Thermal averages and the partition function require sums over 2^N configurations, making exact evaluation generally infeasible except for the smallest systems.
- Pseudolikelihood provides an efficient alternative that can use correlations beyond the first two moments of the data.
- The likelihood maximum is unique under finite parameters because the Hessian is positive-definite when the relevant observables fluctuate.
3. Maximum entropy modelling
Maximum-entropy modelling derives the Boltzmann Ising distribution as the least-structured distribution matching observed first and second moments, while highlighting when pairwise modelling may be inadequate.
- Maximum entropy under constraints on magnetisations and pair correlations yields the Boltzmann distribution with Ising couplings and fields.
- The Ising parameters are Lagrange multipliers chosen to reproduce the data’s first and second moments.
- Pairwise couplings reflect the choice to match only the first two moments, while higher-order interactions may become viable as sample sizes increase.
- Pairwise equilibrium models can be assessed by comparing their predicted three-point correlations with those observed in the data.
- Graphical model selection seeks to recover nonzero-coupling edges and coupling values from independently sampled data.
- At least c max{d^2, α^-2} ln N samples are required when maximum graph connectivity d grows with system size.
5. Thermodynamics of the inverse Ising problem
Thermodynamic potentials recast inverse Ising inference: correlations and magnetisations determine couplings and fields through Legendre transforms and derivatives, with variational approximations needed in practice.
- The inverse problem uses a Legendre transform of the Helmholtz free energy with respect to couplings and fields when correlations and magnetisations are given.
- The transformed potential is the entropy function up to sign and is linked directly to maximum-likelihood inference.
- Connected correlations link susceptibility responses to field changes and provide the relation used to infer couplings.
- Differentiating the relevant thermodynamic potentials at sample correlations and magnetisations yields the couplings and magnetic fields.
- For most systems, free energies and other thermodynamic potentials cannot be evaluated, motivating approximation schemes based on variational principles.
- The variational free-energy principle minimizes a functional over trial distributions, with constrained families producing upper bounds to the Helmholtz free energy.
7. Mean-field theory
Mean-field theory approximates the inverse Ising problem with factorised spin distributions, producing correlation-based coupling estimates that are computationally efficient but can introduce nonphysical diagonal terms.
- The mean-field ansatz factorises the Boltzmann distribution across sites, making spins statistically independent in the trial distribution.
- Mean-field reconstruction relates observed connected correlations to couplings through the inverse correlation matrix.
- Matrix inversion requires O(N^3) operations, compared with the exponentially large computation needed for partition functions or likelihood maximization.
- Mean-field magnetisation parameters minimize the KL divergence between the factorised ansatz and the Boltzmann distribution.
- The linear-response formulation becomes over-determined when diagonal couplings are constrained to zero, so common procedures ignore those constraints and later discard nonphysical diagonal terms.
- The TAP correction adds an Onsager term representing fluctuation effects transmitted through neighbouring spins.
9. Couplings without a loop: mapping to the minimum
For tree-structured couplings, the inverse Ising problem reduces to selecting a minimum spanning tree using pairwise mutual informations, after which the couplings can be recovered efficiently. Bethe–Peierls provides an exact tree solution but becomes uncontrolled when loops are present.
- Computational setting: Exponential correlation-computation costs arise from loops, whereas trees permit efficient inverse Ising reconstruction.The first efficient tree method maps reconstruction onto a minimum spanning tree problem.
- Tree reconstruction: The tree model is optimized by minimizing KL divergence over possible tree topologies using empirical pairwise mutual informations.For a fixed tree, the optimal marginals equal the empirical marginals.
- Tree reconstruction: The optimal topology is a minimum spanning tree whose edge weights are negative pairwise mutual informations.The resulting tree connects all vertices while minimizing the total edge-weight sum.
- Computational setting: Minimum spanning trees can be found by greedy procedures in O(|V|^2 ln |V|) steps.This avoids exhaustive exploration of all possible tree topologies.
- Parameter recovery: After identifying the tree topology, the coupling values can be computed with the independent-pair approximation.On a tree, the factorized probability measure makes this step straightforward.
- Scope and limitations: Bethe–Peierls reconstruction is exact for trees but an uncontrolled approximation on graphs with loops, where its distribution may not be normalized.The ansatz is expected to work better for locally treelike graphs with long loops than for graphs with many short loops.
11. Belief propagation and susceptibility propagation
Belief propagation computes Ising marginals exactly on trees and approximately on locally treelike graphs, supporting inverse reconstruction through parameter updates. Related cluster expansions extend pairwise approximations toward short-loop corrections, but their computational cost grows rapidly with cluster size.
- Belief propagation: Belief propagation computes marginal distributions exactly on trees and can approximate models whose coupling graphs are locally treelike.Its approximation is most effective when existing loops are long.
- Belief propagation: Cavity recursions compute tree partition functions recursively from leaves, then normalize them to obtain one- and two-point marginals.The two-point marginal combines cavity partition functions with the direct coupling between spins.
- Inverse reconstruction: Inverse Ising implementations use belief propagation either to estimate magnetisations and correlations for Boltzmann learning or through susceptibility propagation.Susceptibility propagation introduces messages and update rules for spin susceptibilities.
- Inverse reconstruction: A coupling-updating susceptibility-propagation variant matches Bethe–Peierls reconstruction but may fail to converge even when analytical approximate solutions remain valid.The convergence issue distinguishes the iterative implementation from the corresponding analytical reconstruction.
- Pair and cluster approximations: The independent-pair approximation is exact when the known coupling topology is a tree and the entropy expansion includes only interacting pairs.With unknown topology, summing over all spin pairs yields a rather bad entropy approximation.
- Pair and cluster approximations: Adaptive-cluster expansion adds clusters whose entropy contributions exceed a threshold, continuing until no new cluster qualifies.Higher-order residual entropy evaluations require 2^k steps and quickly become intractable.
- Scope and limitations: The approaches based on mean-field or Bethe–Peierls ansätze are uncontrolled approximations, motivating controlled small-coupling or small-correlation expansions.Short loops are a central reason pairwise Bethe–Peierls approximations can break down.
13. The Plefka expansion
The Plefka expansion approximates the Gibbs free energy by systematically expanding in coupling strength, then derives inverse Ising reconstructions from its derivatives. Its loop terms explain why Bethe–Peierls is exact on trees and require higher-order corrections when loops contribute.
- Expansion strategy: Mean-field theory emerges from the zeroth- and first-order terms of Plefka’s expansion.The resulting free-energy approximation can be used to derive inverse Ising solutions.
- Expansion strategy: Plefka’s method expands the Gibbs free energy in powers of a perturbation parameter distinguishing orders in coupling strength.The parameter is set to one after the expansion is calculated.
- Inverse reconstruction: Mean-field and TAP coupling reconstructions follow from derivatives of the approximate Gibbs free energy with respect to magnetisations.The corresponding magnetic-field reconstructions follow from first derivatives with respect to magnetisations.
- Loop contributions: Terms such as JijJjkJki quantify coupling strengths around closed loops and vanish when the coupling graph is a tree.For graphs with loops, these terms contribute to the Gibbs free energy.
- Loop contributions: Plefka terms resum to the Bethe–Peierls result, establishing Bethe–Peierls exactness on trees within this expansion.Higher-order terms may matter for specific coupling distributions even when lower-order terms are extensive.
14. The Sessak–Monasson small-correlation expansion
The Sessak–Monasson approach expands the entropy around an uncorrelated system in connected correlations, producing successive approximations for fields and couplings. Its resummations combine mean-field and independent-pair contributions, while specialized loop resummations can improve robustness to sampling noise.
- Expansion strategy: The Sessak–Monasson expansion expresses entropy as a perturbative series in connected correlations around the non-interacting case.When connected correlations vanish, the reconstructed couplings are zero and h_i^(0) = artanh(m_i).
- Inverse reconstruction: Successive evaluations of the expansion yield reconstructed fields and couplings after empirical magnetisations and connected correlations are substituted.The expansion begins from the uncorrelated solution and proceeds through higher orders.
- Resummation: The independent-pair approximation appears as a resummed sub-series of the Sessak–Monasson expansion.Its combined expression adds mean-field and independent-pair sub-series while subtracting their overlap.
- Resummation: The mean-field contribution is itself a sub-series of the Sessak–Monasson expansion.The overlap term represents the mean-field reconstruction of two spins treated independently.
- Resummation: Resumming loop diagrams has been used to obtain a reconstruction described as particularly robust against sampling noise.Other sub-series can also be resummed for special cases.
B. Logistic regression and pseudolikelihood
Pseudolikelihood reformulates inverse Ising inference through conditional statistics, turning reconstruction into tractable regression problems. Across tested equilibrium settings, it is accurate and more robust than approximate methods, especially at strong coupling.
- Logistic regression and pseudolikelihood: Pseudolikelihood has polynomial computational complexity in N and M, unlike exact likelihood maximization, whose complexity is exponential in system size.It is exact in the infinite-sample limit.
- Logistic regression and pseudolikelihood: The reconstruction decomposes into N separate problems, replacing Boltzmann averages over 2^N−1 states with averages over observed configurations.This uses information from all spin configurations, not only sample magnetisations and correlations.
- Logistic regression and pseudolikelihood: Conditioning on all other spins makes each field and coupling row the intercept and coefficients of a multivariate logistic regression.The conditional statistics of one spin are fully captured by its field and couplings to the remaining spins.
- Reconstruction performance: At β = 1.3, pseudolikelihood accurately reconstructs couplings and fields, with inferred correlations and magnetisations showing little bias alongside exact likelihood.Mean-field overestimates large couplings, while TAP and Bethe–Peierls reduce that error.
- Reconstruction performance: Pseudolikelihood outperforms other non-exact methods across β and remains valid in the glassy phase, while increased samples can compensate for its strong-coupling error.The other approximate methods break down at high β in ways that additional samples cannot compensate.
- Reconstruction performance: On random graphs and square lattices, pseudolikelihood and Bethe–Peierls perform well, whereas adaptive cluster expansion is highly accurate at weak coupling but breaks down at strong coupling.Adaptive cluster expansion can avoid sampling-noise error at weak coupling by setting some couplings exactly to zero.
D. Reconstruction and non-ergodicity
Non-ergodicity complicates inverse Ising reconstruction because samples may represent one thermodynamic state rather than the full Gibbs measure. Restricting inference to the sampled ergodic component can recover the appropriate likelihood, but generic extensions remain open.
- Reconstruction and non-ergodicity: Strong couplings can produce non-ergodic sampling, making parameter reconstruction fundamentally difficult and limiting current understanding.The system may remain confined to one free-energy valley for long times.
- Reconstruction and non-ergodicity: Spin-glass ergodicity breaking creates multiple thermodynamic states and causes mean-field, TAP, and Bethe–Peierls methods to fail in the glassy phase.Their self-consistent equations develop multiple magnetisation solutions corresponding to different frozen orientations.
- Reconstruction and non-ergodicity: Mean-field equations describe individual thermodynamic states rather than the mixture represented by the full Boltzmann measure.At strong coupling, identifying observed quantities with one self-consistent state becomes the central difficulty.
- Reconstruction and non-ergodicity: Belief-propagation fixed points can approximate marginal distributions within individual non-ergodic components, especially when loops are sufficiently long.For non-tree couplings, this approximation can become exact in the thermodynamic limit under the stated condition.
- Reconstruction and non-ergodicity: Restricting the partition function and empirical averages to the sampled component allows parameters for all thermodynamic states to be inferred from single-state data.This replaces the full likelihood with a likelihood computed over the relevant configuration subset.
- Reconstruction and non-ergodicity: Inverse problems for non-ergodic systems remain challenging, including whether the single-state approach extends from the Hopfield model to generic couplings.The review identifies this extension as an open question.
E. Parameter reconstruction and criticality
The review connects inferred-model criticality with empirical magnetisations and susceptibilities, then contrasts equilibrium inference with non-equilibrium reconstruction under asymmetric dynamics. Exact non-equilibrium identities exist, but their evaluation generally requires unknown steady-state statistics.
- Parameter reconstruction and criticality: In the retinal example, 160-neuron recordings are converted into binary spins over 20 ms intervals, with correlations and magnetisations estimated across 297 stimulation courses.Pseudolikelihood reconstruction is applied to the first 30 neurons.
- Parameter reconstruction and criticality: At β = 1, the reconstructed model reproduces magnetisations and correlations close to those in the original retinal data, while changing β changes the magnetisations.Fixed fields are not rescaled to preserve the original magnetisations away from β = 1.
- Parameter reconstruction and criticality: Empirical magnetisations and correlations can place inferred Ising models near a ferromagnetic critical point, even when that transition is not present in the original system.The inferred transition arises when correlations and magnetisations are sufficient and magnetisations are neither +1 nor −1.
- Parameter reconstruction and criticality: High susceptibility links inferred-model criticality to the Fisher information matrix used to calculate covariances of maximum-likelihood estimates.The review presents this as a connection between criticality and information geometry.
- Non-equilibrium reconstruction: Asymmetric couplings violate detailed balance and produce non-equilibrium steady states, so equilibrium reconstruction results do not apply.This motivates reconstruction from time-series data or non-equilibrium steady-state samples.
- Non-equilibrium reconstruction: Exact relations for non-equilibrium magnetisations and correlations can support inference, but they are difficult to evaluate because the relevant steady-state averages are generally unknown.The relations involve statistics of spins in the non-equilibrium steady state.
2. Parallel Glauber dynamics
Parallel Glauber dynamics turns time-series reconstruction into an explicitly evaluable likelihood, connecting non-equilibrium inference to regression. Mean-field and TAP approximations extend through expansions around a factorising ansatz.
- Dynamics: Parallel updates define a Markov chain in which all spins stochastically update between successive time points.For symmetric couplings, the steady state can still be represented using Peretto’s pseudo-Hamiltonian.
- Likelihood reconstruction: Time-series likelihoods can be evaluated over arbitrary intervals, including before the nonequilibrium steady state is reached.The same approach also reconstructs symmetric couplings when time-series data are available.
- Likelihood reconstruction: The non-equilibrium likelihood is computationally tractable because its normalisation is contained in the cosh term.Its derivatives can be evaluated in MN^2 computational steps and optimized by gradient ascent.
- Regression connection: Parallel Glauber dynamics yields logistic-regression equations for predicting a spin at the next time from the current configuration.Time-series likelihood derivatives can therefore be written as averages over regression data.
- Approximations: Mean-field and TAP equations apply to both sequential and parallel dynamics because their magnetisation and correlation relationships are identical.The same expansions also extend to times before the nonequilibrium steady state is reached.
3. The Gaussian approximation
The Gaussian approximation describes effective local fields with Gaussian statistics and links their parameters to spin observables. For asymmetric Ising models, it provides a reconstruction framework that improves on mean-field estimates but requires iterative inference of latent-field parameters and couplings.
- In some regimes and in the thermodynamic limit, effective local fields in asymmetric Ising models follow Gaussian distributions characterized by means, standard deviations, and covariances.
- The method transforms between spin time series and effective local fields, then evaluates correlation functions within Gaussian theory.
- Gaussian theory captures local-field fluctuations more accurately than mean-field theory, which neglects them in the relevant term.
- The effective-field variance-related quantity cannot be determined from data alone, so an iterative scheme jointly infers effective-field parameters, couplings, and magnetic fields.
- For asymmetric couplings, simulations compare mean-field, Gaussian, and exact-likelihood reconstructions; mean-field significantly underestimates couplings.
C. Outlook: Reconstruction from the steady state
Reconstructing non-equilibrium models from steady-state samples is underdetermined by pairwise correlations alone because asymmetric couplings produce symmetric correlation observables. The review therefore discusses higher-order correlations, perturbations, and broader inference methods, alongside their assumptions and practical scope.
- C. Outlook: Reconstruction from the steady state: Pairwise correlations are insufficient for non-equilibrium reconstruction because correlations are symmetric while asymmetric couplings contain twice as many free parameters as observables.
- C. Outlook: Reconstruction from the steady state: Three-spin correlations can infer couplings, but connected three-spin correlations are small when effective local fields are approximately multivariate Gaussian.
- C. Outlook: Reconstruction from the steady state: Perturbation-based inference compares pair correlations before and after known coupling, field, or spin-fixing perturbations and has been applied to gene regulatory networks.
- C. Outlook: Reconstruction from the steady state: Genetic variation and expanding single-cell analysis provide additional perturbations that may extend network inference beyond gene regulatory systems.
- C. Outlook: Reconstruction from the steady state: For equilibrium data, pseudolikelihood is asymptotically close to optimal, while time-series likelihoods are comparatively easy to evaluate in the non-equilibrium regime.
- C. Outlook: Reconstruction from the steady state: Pseudolikelihood and time-series likelihood can both be formulated as regression, creating a common framework across equilibrium and non-equilibrium inference.
- C. Outlook: Reconstruction from the steady state: Interaction screening uses convex optimization to infer each coupling-matrix column and field by balancing exponential sample sums, with regularization enabling sparsity.
- C. Outlook: Reconstruction from the steady state: The interaction-screening objective has a unique global minimum at the true parameters and nearly reaches information-theoretic bounds for sparse Ising reconstruction.