Source-linked AI summary
The free energy principle for action and perception: A mathematical review
Christopher L. Buckley, Chang Sub Kim, Simon McGregor, Anil K. Seth
TL;DR
The FEP is considered in the context of the search for a unified brain theory, while questions about priors remain open. This paper presents its essential mathematical aspects and implementation, reports formal simplifications, and highlights sensitivity to precisions alongside unresolved Gaussian-representation questions.
Problem
The brain sciences have long searched for a unified brain theory, while the origins of priors remain a key open question.
Method
The paper presents the essential mathematical aspects of the FEP and its implementation, including free energy formed from the Kullback-Leibler divergence.
Results
The implementation can be simplified while remaining formally equivalent, and behaviour is extremely sensitive to precisions.
Takeaways & Limitations
Clarifying the mathematical structure and assumptions of the FEP helps clarify its scientific contributions and significance for the brain sciences.
Takeaways & Limitations
It remains an open question whether representing the world in terms of Gaussian distributions is appropriate.
Abstract
from arXiv · showhide
The 'free energy principle' (FEP) has been suggested to provide a unified theory of the brain, integrating data and theory relating to action, perception, and learning. The theory and implementation of the FEP combines insights from Helmholtzian 'perception as inference', machine learning theory, and statistical thermodynamics. Here, we provide a detailed mathematical evaluation of a suggested biologically plausible implementation of the FEP that has been widely used to develop the theory. Our objectives are (i) to describe within a single article the mathematical structure of this implementation of the FEP; (ii) provide a simple but complete agent-based model utilising the FEP; (iii) disclose the assumption structure of this implementation of the FEP to help elucidate its significance for the brain sciences.
1. Introduction
The FEP is presented as a candidate unified theory relating action, perception, and learning, while this paper provides a mathematical appraisal, agent-based model, and analysis of its assumptions.
- Theoretical background: The framework combines Helmholtzian perceptual inference, Bayesian and machine-learning ideas, and thermodynamic free energy.Its historical development includes Helmholtz machines, expectation-maximization, autoencoders, and population-code learning.
- Theoretical background: The FEP describes free-energy minimization through action that changes sensory input or perception that updates internal models.This gives action and perception roles within one framework for studying their interactions with learning.
- Motivation: The FEP aims to provide a unified account of cognition spanning action, perception, and learning.It is also claimed to relate concepts including memory, attention, value, reinforcement, and salience.
- Theoretical background: Free energy is an information-theoretic proxy for sensory surprise that an organism can evaluate using sensory input and an internal environmental model.The principle motivates minimizing atypical events associated with maintaining viable biological organization.
- Scope: The FEP’s broad explanatory claims motivate close mathematical examination, but the paper does not attempt to resolve those claims.The framework has been proposed to unify several brain theories and extend across multiple biological timescales.
- Paper objectives: The paper first supplies a complete technical account, then a simple complete agent-based model, and finally an analysis of the framework’s assumption structure.The authors aim to clarify the FEP’s scientific contributions and disclose non-obvious assumptions.
2. An overview of the FEP
The FEP frames organisms as maintaining probabilistic models of their environments and reducing atypical sensory exchanges through perception and action. Its implementation uses informational free energy to approximate Bayesian inference and bound sensory surprisal.
- Surprisal: Surprisal quantifies sensory atypicality as the negative logarithm of the probability of observed data.It is large for unlikely observations and zero when observations have probability 1.
- Probabilistic modelling: The FEP proposes that organisms maintain probabilistic models of environmental states and update them from sensory signals.These models are represented by recognition and generative densities encoded through physical brain variables.
- Informational free energy: Informational free energy is a non-negative divergence-based quantity that can be evaluated using the recognition and generative densities.It is distinct from thermodynamic free energy and depends on interpreting brain variables as encoding probability densities.
- Action: Minimizing informational free energy also provides an upper bound on sensory surprisal rather than minimizing surprisal directly.Organisms are proposed to reduce surprisal indirectly by acting on the environment and changing sensory input.
- Perception: Minimizing informational free energy makes the recognition density approximate the posterior distribution of environmental variables given sensory data.This provides an approximate Bayesian-inference account when exact inference is difficult.
- Action: In the implementation considered here, action is treated as control that reduces deviations between actual and desired environmental trajectories.Other proposed roles for action, such as disambiguating competing models, are not considered.
3. Informational free energy
The paper formulates informational free energy as an evaluable objective for approximate Bayesian inference when posterior calculations are intractable. Minimizing it improves posterior approximation and supplies an upper bound on sensory surprisal, with action providing an indirect route to reducing surprisal.
- Generative model: The agent infers environmental states from sensory input using a generative density that combines prior beliefs with a sensory likelihood.The generative density is factorized into the prior over environmental states and the probability of sensory input given those states.
- Approximate inference: Exact posterior calculation can be intractable because continuous integrals may lack analytic solutions and discrete sums can grow exponentially with state number.Variational Bayes addresses this by introducing an auxiliary recognition density and an optimization problem.
- Informational free energy: Informational free energy can be evaluated directly from the recognition density and generative density even though the true posterior is unavailable.This follows from rewriting the divergence between the recognition density and posterior in terms of quantities the agent can specify.
- Posterior approximation: Minimizing informational free energy with respect to the recognition density minimizes its Kullback–Leibler divergence from the true posterior.The resulting recognition density therefore approximates environmental states conditional on sensory data.
- Surprisal bound: Informational free energy provides an upper bound on sensory surprisal, becoming equal to surprisal only under a specific condition.The paper emphasizes that this procedure estimates and bounds surprisal but does not minimize it directly.
- Action: Action can minimize informational free energy indirectly by changing the environment and thereby changing sensory input.Within this treatment, action is linked to control of deviations from desired environmental trajectories.
4. The R-density: How the brain encodes environmental states
The R-density encodes the brain’s recognition distribution over environmental states through physical variables such as neuronal activity. The paper develops a Gaussian Laplace approximation that simplifies informational free energy into a function of means, variances, and sensory inputs, while relying on consequential assumptions.
- Neural encoding: The implementation requires the brain to encode the R-density through neuronal quantities that parameterize its sufficient statistics.The R-density is treated as a family of probability densities over environmental states, selected by brain states.
- Approximations: The paper considers factorized and Gaussian approximations because the general R-density optimization is intractable.Factorization yields iterative mean-field updates, whereas the Gaussian approach is developed in detail.
- Laplace approximation: Under the Laplace approximation, Gaussian means and variances become parameters optimized numerically to minimize informational free energy.The derivation is presented first for a univariate Gaussian and then extended to the multivariate case.
- Assumptions: The analytic model relies on a sharply peaked Gaussian recognition density and a smooth energy function, assumptions with non-trivial implications for interpreting brain function.The paper returns to these implications in its later discussion.
- Variance optimization: The derivation removes the variance dependence from informational free energy by optimizing the Gaussian variance.The resulting expression depends on Gaussian means and sensory inputs rather than variances.
- Multivariate formulation: The multivariate Laplace-encoded energy provides the general approximation to informational free energy used in the remainder of the study.It is formulated using vectors of brain states and sensory data corresponding to environmental variables.
- Interpretation: The Laplace encoding represents the most likely environmental causes of sensory data while also encoding uncertainty through variances or inverse variances.The interpretation is therefore not limited to a single point estimate of environmental states.
5. The G-density: Encoding the brains beliefs about environmental causes
The paper constructs generative densities that encode beliefs about environmental causes of sensory data, then derives Laplace-encoded free-energy expressions for static and dynamic models. In the simplified multivariate case, the resulting energy is a precision-weighted sum of prediction errors plus logarithmic variance terms.
- Generative densities: The G-density encodes beliefs about environmental causes of sensory signals and supports specification of the IFE.The construction proceeds from a generative model to brain-state expectations and then to an IFE expression.
- Simplest generative model: The simplest model assumes sensory data arise from a nonlinear mapping of environmental-state beliefs with additive Gaussian noise.The environmental-state belief is centered on an a priori mean and includes Gaussian fluctuations.
- Belief representations: The model distinguishes current environmental beliefs in the R-density from future-state expectations and confidence encoded in the G-density.µ and ζ describe uncertain beliefs about the current environment, whereas ¯µ and σw describe expected future states and confidence in them.
- Laplace-encoded energy: The Laplace-encoded energy combines sensory and state residual errors, each weighted by the inverse variance representing relative confidence.The residuals are εz = ϕ − g(µ; θ) for sensory prediction error and εw = µ − ¯µ for deviation from expected state.
- Multivariate extension: The multivariate formulation includes correlated noise sources and covariance matrices, but the paper simplifies the general case by assuming statistical independence.Under independence, prior and likelihood densities factorise into uncorrelated Gaussian forms.
- Dynamic generative models: For dynamic environments, generalized coordinates and Langevin-type dynamics extend the formulation to arbitrary dynamical orders.The dynamic model replaces the static state equation with a drift function plus random fluctuation and yields a multivariate IFE approximation across arbitrary orders.
4806. IFE minimisation: How organisms infer environmental states
The paper describes recognition dynamics as gradient descent on Laplace-encoded energy, allowing brain states to implement approximate inference. The generalized-coordinate formulation requires a distinction between ordinary motion and the trajectory encoded by the generalized state.
- IFE minimisation: Recognition dynamics update brain states by gradient descent on IFE, recursively moving them toward lower Laplace-encoded energy.The update uses a learning rate κ and modifies states between sequential time steps.
- IFE minimisation: The stationary solution occurs when the energy gradient vanishes, so brain-state dynamics settle at a minimum of Laplace-encoded energy.For ordinary states, this corresponds to the relevant time derivative becoming zero.
- Generalized coordinates: Generalized coordinates require extending the update rule to higher dynamical orders and converting sequential updates into differential equations.The formulation uses a temporal-derivative operator to represent generalized-state dynamics.
- Generalized coordinates: A complication is that ordinary temporal derivatives are already contained in generalized coordinates, preventing the gradient-descent procedure from reaching a stationary solution at every order.The paper addresses this by distinguishing motion in generalized state-space from the trajectory encoded in the brain.
- Generalized coordinates: With the additional distinction between generalized velocity and the mode of the path, the FEP restates gradient descent so the energy gradient vanishes when 9˜µα = D˜µα.In static situations, D˜µα is set to zero by construction.
- Neuronal implementation: Continuously integrating the resulting first-order ordinary differential equations in sensory-data streams would continuously minimise IFE and implement approximate inference.The paper notes that additional assumptions about implementation are needed for a strongly biologically plausible process theory.
5107. Active inference
Active inference extends perceptual inference by using action to alter sensory data and reduce IFE. In a simple temperature-control agent, action reconciles conflicting sensory evidence and desired states by moving the agent through its environment.
- Active inference: The framework treats action as minimising IFE by changing sensory input rather than directly changing the formulation of IFE.Action relies on an inverse model linking actions to changes in sensory data.
- Agent-based model: The agent-based model contains a mobile agent on a one-dimensional plane that moves to achieve a desired local temperature.Temperature depends on position relative to a simple source, and the agent senses local temperature and its temporal derivative.
- Agent-based model: The agent’s internal generative model has a stable equilibrium at Tdesire, even though the actual environment has different dynamics.The agent therefore acts to make the environment conform to its internally modelled dynamics.
- Perceptual inference: With greater confidence in sensory input than in its internal model, the agent successfully infers local temperature and its derivatives.Under this condition, the gradient-descent scheme is equivalent to least mean square estimation on sensory data.
- Perceptual inference: When sensory and internal-model variances are equally weighted, inferred temperature falls between the desired and sensed temperatures.The two sources of information cannot both be satisfied when perception conflicts with the agent’s desire.
- Active inference: The example demonstrates that IFE minimisation can underpin both perception and action, reconciling tension between desires and perception through action.The model provides a simple complete agent-environment implementation of this relationship.
5958. Hierarchical Inference and Learning
The hierarchical FEP extends generative models across cortical layers and dynamical orders, enabling empirical priors, learning, inference, and action within one full construct. Its state units encode conditional expectations and prediction errors, yielding posterior expectations of environmental causes.
- Hierarchical Inference and Learning: The paper identifies developing agent-based models using the full construct as a next step for extending the framework.The earlier agent model used a generative model of a simple environment, whereas the full construct incorporates the richer hierarchical and dynamical structure.
- Hierarchical Inference and Learning: Hierarchical generative models use higher levels to provide empirical priors or constraints on lower levels, avoiding an explicit fixed prior.The hierarchy is proposed as a route to empirical Bayes and learning arbitrary environmental dynamics.
- Hierarchical Inference and Learning: The model places sensory data at the lowest cortical level and lets higher-level dynamics be governed by random fluctuations across statistically independent hierarchical levels.Each level is connected through a generative function, with fluctuations specifying inter-level variability and observation noise.
- Hierarchical Inference and Learning: A large top-level noise variance leaves the level below effectively unconstrained, producing an empirical-Bayes-like inference without a prior at that level.The corresponding top-level term in the Laplace-encoded energy is approximately zero.
- Hierarchical Inference and Learning: The full construct combines multi-layer cortical hierarchies with multi-scale generalized-coordinate dynamics within each layer.Inter-layer links use causal states, while hidden states mediate intra-layer dynamics and the resulting prediction errors enter the energy formulation.
- Hierarchical Inference and Learning: At the end of inference, the optimal brain state represents the posterior expectation of the environmental cause of observed sensory data.The state units encode conditional expectations and associated prediction errors during this process.
9. Discussion
The discussion presents the FEP implementation as mathematically tractable and compatible with a hierarchical cortical architecture, while emphasizing unresolved assumptions about Gaussian representations, optimization, generalized motions, and active behavior. The paper therefore frames the implementation as a structured process theory whose practical and biological adequacy remains open.
- 9. Discussion: The paper presents the FEP implementation as an ambitious mathematical framework while stressing that its assumptions and approximations require further evaluation.The authors specifically leave the interplay between brain states and precisions in complex active behavior unresolved.
- 9. Discussion: The Laplace implementation represents expectation values rather than full distributions over environmental states, with uncertainty encoded through precisions on brain-state expectations.This shifts uncertainty representation toward uncertainty in the model linking hidden causes and sensory signals.
- 9. Discussion: Gaussian assumptions simplify FEP implementation and make it formally equivalent to predictive coding.They also support a proposed neuronal interpretation involving message passing through cortical hierarchies.
- 9. Discussion: The implementation maps inferred causes to firing rates, generative models to synaptic connectivity, and free-energy minimization to neuronal relaxation dynamics.Hierarchical generative models correspond to cortical network organization, with top-down predictions and bottom-up prediction errors.
- 9. Discussion: Whether Gaussian representations suffice for complex real-world sensorimotor interactions remains an open question.Alternative schemes using multimodal distributions or Bayesian sampling may offer more versatile implementations, but their FEP and neuronal compatibility remains unresolved.
- 9. Discussion: Gradient-descent minimization simplifies inference and learning, but its convergence behavior, local minima, and timing remain insufficiently understood.Learning-rate choices are also important for timely inference and control, without a clear consensus on how to incorporate them into process theories.
- 9. Discussion: Generalized-motion inference assumes differentiable sensory noise and linear interactions among derivatives, while the practical value of higher-order derivatives remains unclear.Signals beyond the second derivative may be small and noisy, limiting their usefulness in some applications.
- 9. Discussion: Active inference requires an inverse model linking actions to sensations, whose specification is non-trivial in the general case.The framework proposes that motor actions fulfill proprioceptive predictions, but the inverse-model requirement remains a substantive assumption.
Appendix A. Variational Bayes: Ensemble learning
The appendix derives the ensemble-learned R-density by variationally optimizing the IFE under factorization and normalization constraints. The resulting density is directed toward the posterior, while the factorized environmental sub-states are treated through averaged interactions and distinctive time-scales.
- Variational derivation: Variational optimization of the IFE with respect to one R-density, holding the others fixed and enforcing normalization, yields an optimal density.A Lagrange multiplier implements the normalization constraint.
- Variational derivation: The optimal R-density has a canonical-ensemble-like functional form, with partially averaged energy determining its exponential weighting.The paper relates this form to the equilibrium canonical ensemble in statistical physics.
- Ensemble learning: Under the factorization approximation, the R-density is expressed using a total partition function for environmental states and summed partially averaged energies.The partially averaged energies incorporate the average effects of interactions among environmental partitions.
- Ensemble learning: The self-consistent solution iteratively updates the R-density using the partially averaged energy until estimation and evaluation converge.The procedure begins with an ansatz for the optimal R-density and repeatedly updates it.
- Inference consequence: Equation (A.12) directs the R-density toward the posterior, and minimizing the IFE makes the R-density approximate the true posterior.The minimum IFE is also described as identical to surprisal under the stated assumptions.
- Assumptions: The factorization approximation assumes environmental sub-states vary on distinctive, ordered time-scales, allowing their complicated interactions to be averaged out.Each sub-state has an associated time-scale satisfying τ1 < τ2 < ... < τN.
Appendix B. Dynamic Bayesian Thermostat
The Dynamic Bayesian Thermostat implements sensory input, prediction errors, variational free energy, and gradient-based recognition dynamics in a simple agent model. Its glossary connects these computations to hierarchical brain states, generative mappings, environmental causes, and Bayesian quantities.
- Implementation: The thermostat initializes sensors, brain-state variables, error terms, variational energy, generative-model parameters, and action-related variables before iterating the simulation.The implementation explicitly separates sensory input, error terms, variational energy, inference parameters, and generative-process parameters.
- Variational energy: The instantaneous variational free energy sums precision-weighted squared sensory and dynamic prediction errors with a logarithmic precision term.The same structure is used for the initial energy and the iterated energy IFE(i).
- Recognition dynamics: Recognition dynamics update generalized brain states by gradient descent on prediction errors, with separate updates for mu_0, mu_1, and mu_2.The update equations weight error terms by their corresponding precisions and advance states using dt and the inference rate k.
- Action: Active inference updates the action variable from the sensory prediction error after time exceeds 25, coupling action to the inferred sensory state.The action update uses the sensory mapping derivative Tx(i), the error epsilon_z1, its precision, and an action rate k_a.
- Bayesian interpretation: The glossary defines free energy as an upper bound on surprisal that permits approximation of the posterior, while generalized brain states collect successive time derivatives.The model also represents sensory data, environmental states, priors, likelihoods, and posteriors within a hierarchical generative framework.