Source-linked AI summary
Empowerment -- an Introduction
Christoph Salge, Cornelius Glackin, Daniel Polani
TL;DR
The chapter introduces empowerment to formalize an agent’s preparedness and perceptible control as a task-independent utility function. It presents the channel-capacity formulation, surveys applications and approximations, and reports varied control behavior while identifying horizon, sensing, and locality limitations.
Problem
Agents need a task-independent proxy for preparedness and control because directly evaluating long-term fitness or every possible task is impractical.
Method
Empowerment is formulated as channel capacity from sequences of agent actions to later sensor states, with applications across sensorimotor settings and a continuous-domain approximation.
Results
Empowerment produces varied observed behaviors across applications, including pendulum swing-up, oscillation, lower-position resting, maze centrality, and sensor evolution.
Takeaways & Limitations
The same generic empowerment principle can organize behavior across different agents, environments, and goal-oriented tasks.
Takeaways & Limitations
The relationship between local empowerment and global properties is not generally guaranteed, especially when relevant structure lies beyond the n-step horizon or environments contain heterogeneous regions.
Abstract
from arXiv · showhide
This book chapter is an introduction to and an overview of the information-theoretic, task independent utility function "Empowerment", which is defined as the channel capacity between an agent's actions and an agent's sensors. It quantifies how much influence and control an agent has over the world it can perceive. This book chapter discusses the general idea behind empowerment as an intrinsic motivation and showcases several previous applications of empowerment to demonstrate how empowerment can be applied to different sensor-motor configuration, and how the same formalism can lead to different observed behaviors. Furthermore, we also present a fast approximation for empowerment in the continuous domain.
4.1 Introduction
Empowerment formalizes preparedness as an information-theoretic measure of an agent’s perceptible control over its environment. The chapter surveys its task-independent formulation, applications, parameter choices, and continuous-domain challenges.
- 4.1 Introduction: Empowerment quantifies an agent’s degrees of freedom as a proxy for preparedness and prospective fitness.It avoids evaluating the full fitness landscape directly.
- 4.1 Introduction: The measure combines an agent’s control over its environment with its ability to sense that control.This supports both behavioral decision-making and analysis of agent-environment interaction.
- 4.1 Introduction: Empowerment is intended as a local, universal, and task-independent utility function expressed in information-theoretic terms.Local computation requires knowledge of the agent’s local dynamics rather than the whole system.
- 4.1 Introduction: The formalism defines empowerment as channel capacity between earlier actuator states and later sensor states, measuring reliable perceptible influence.For n-step empowerment, the input is a sequence of the next n actions and the output is the resulting later sensor state.
- 4.1 Introduction: The chapter reviews prior empowerment work, identifies application choices that affect computation, and discusses continuous settings requiring fast approximations.Its stated goal is to provide an overview and accessible entry point to the field.
4.2 Related Work
Empowerment is situated among information-theoretic, embodied-cognition, and intrinsic-motivation approaches. Unlike curiosity-oriented measures, it emphasizes preferred states based on local control rather than exploration or learning progress.
- 4.2 Related Work: Empowerment draws on information theory applied to embodied biological systems, including sensor redundancy and information bottlenecks.This places the measure within a tradition linking perception, cognition, and information processing.
- 4.2 Related Work: Its conceptual basis is the immediate relationship between a situated, embodied agent and its surroundings.This connects empowerment to the Umwelt principle, perception-action loops, enactivism, and embodied robotics.
- 4.2 Related Work: Intrinsic-motivation research includes pain-like adaptations, exploration, artificial curiosity, flow, learning progress, and predictive information.These approaches explain behavior without relying solely on explicit external rewards.
- 4.2 Related Work: Some learning-oriented mechanisms tend toward stable, easily predictable environments, requiring additional provisions to preserve behavioral richness.The cited limitation concerns predictability without sufficient potential for instability.
- 4.2 Related Work: Curiosity-based mechanisms often target learning or prediction, whereas empowerment identifies preferred states once local environmental dynamics are known.High empowerment can remain satisfactory even when the environment is not well understood.
- 4.2 Related Work: The chapter also considers whether empowerment-like behavior could arise from physical principles such as maximum entropy production.This is presented as a possible route beyond biological motivation hypotheses.
4.3 Empowerment Hypotheses
The chapter presents empowerment as a task-independent motivation for artificial and biological behavior, while treating its biological hypotheses as open and requiring scenario-specific testing. It also identifies useful behaviors and clear task boundaries.
- 4.3 Empowerment Hypotheses: The chapter frames empowerment hypotheses as motivations rather than conclusive arguments, with generic forms that may require task-specific operationalization.Specific scenarios can make the hypotheses experimentally testable.
- 4.3 Empowerment Hypotheses: One motivation is enabling artificial agents to decide flexibly without a task designed into every situation; another is finding proxies for prospective fitness.These motivations connect empowerment to general AI and organismal adaptation.
- 4.3 Empowerment Hypotheses: The behavioral hypothesis proposes that evolved organisms behave as if maximizing empowerment in the absence of specific goals.A weaker version requires only behavior that approximates empowerment-driven action, not explicit computation.
- 4.3 Empowerment Hypotheses: The framework is expected to be local and applicable across different organisms and sensorimotor configurations, but these properties remain part of the hypothesis.The text presents locality and universality as conditions supporting plausibility.
- 4.3 Empowerment Hypotheses: Another hypothesis states that natural evolution increases the empowerment of resulting organisms through more efficient sensor and actuator use.Empowerment is defined as actuator-to-sensor channel capacity, so imperceptible actuator effects or irrelevant sensing are inefficient.
- 4.3 Empowerment Hypotheses: Empowerment can generate beneficial behavior across tasks, including pole balancing and maze centrality, but it will not seek externally desired states that are not highly empowered.Such goals remain the standard domain of traditional AI algorithms.
4.4 Formalism
Empowerment recasts an agent’s sensorimotor influence as an information-theoretic channel capacity, with causal interventions linking actions to later sensor states. The formalism extends from single-step and deterministic settings to memory, n-step action sequences, context-dependent analysis, and nondeterministic continuous or discrete calculations.
- 4.4 Formalism: Entropy, conditional entropy, mutual information, and channel capacity provide the information-theoretic quantities used to formalize empowerment in bits.Channel capacity is the maximum reliably transmissible information rate, and empowerment specializes it to the actuation-perception channel.
- 4.4.1 The Causal Interpretation of Empowerment: Empowerment uses intervention-based causal information flow rather than the observed joint distribution, maximizing mutual information over interventional action distributions.The interventional conditional p(y|ˆx) defines the causal channel, whose capacity can be computed with methods such as Blahut-Arimoto.
- 4.4.1 The Causal Interpretation of Empowerment: Computing empowerment from observed actuation-sensing data requires a causal pair without common causes or reverse influence from later sensors to earlier actions.This condition allows the observed conditional distribution to represent the relevant causal channel.
- 4.4.2 Empowerment in the Perception Action Loop: Empowerment measures both environmental influence and the agent’s perception of that influence as a continuous information-theoretic analogue of controllability and observability.Unlike dimensional controllability and observability, it can take any non-negative real value and apply beyond linear spaces or manifolds.
- 4.4.2 Empowerment in the Perception Action Loop: In the perception-action loop, empowerment is the channel capacity between actuator states and later sensor states, with actions, sensors, memory, and environment represented in a time-unrolled causal Bayesian network.The basic one-step form relates A_t to S_t+1; multiple actuators can be represented as one actuator variable.
- 4.4.3 n-step empowerment: n-step empowerment evaluates action sequences over the next n steps against the resulting later sensor state, here using open-loop sequences selected before the final observation.This generalization can reveal interaction characteristics that are not distinguished by a single action and subsequent sensor readout.
- 4.4.4 Context-dependent Empowerment: Context-dependent empowerment evaluates how the current world state changes action effects and identifies a possibly non-unique minimal context Kopt that maximizes contextual empowerment.Such contexts can support intrinsic internal representations based on the agent’s sensory-motor capacity and interaction with the world.
- 4.4.6 Discrete Deterministic Empowerment: In discrete deterministic worlds, empowerment reduces to the logarithm of the number of distinguishable sensor states reachable with available actions, while nondeterministic continuous cases require difficult integrals or costly approximations.For finite discrete sensor states, the relevant integral is replaced by a sum in the Blahut-Arimoto calculation.
4.5 Discrete Examples
Discrete examples show that empowerment depends on local reachability, sensor and actuator configurations, and the chosen time horizon. These factors can produce informative gradients, connect local control to global structure under conditions, or flatten the landscape when sensing and action capacity become mismatched.
- 4.5.2 Average Distance vs. Empowerment: Local n-step empowerment can correlate negatively with a location’s average distance to all other maze locations.This suggests that local reachability may encode a global property, although the relationship is not generally established.
- 4.5.3 Sensor and Actuator Selection: For a pushable but unobservable box, the empowerment map is flat because the box affects neither outcomes nor sensor inputs.This provides no empowerment gradient for action selection.
- 4.5.3 Sensor and Actuator Selection: Perceiving a pushable box raises empowerment near it by distinguishing outcomes that otherwise share the same final agent location.The increased value reflects more distinguishable sensor outcomes, not more reachable world states.
- 4.5.3 Sensor and Actuator Selection: A non-pushable box lowers nearby empowerment by blocking movement, while observability does not change the resulting map.The obstacle reduces the number of states reachable within the 5-step sequence.
- 4.5.4 Horizon Extension: Longer horizons reveal more structure, but action sequences grow exponentially with horizon length, making noisy cases quickly infeasible.Extending the horizon can also miss changes beyond the agent’s current cone of view, and overly large action or sensor spaces can flatten empowerment.
- 4.5.6 Sensor and Actuator Evolution: Empowerment-driven sensor evolution produces different layouts at different starting positions, including a central blob that eventually collapses into a one-dimensional heading sensor.Some locations admit several nearly equivalent solutions, whereas others are more constrained.
4.6 Continuous Empowerment
Continuous empowerment extends the channel-capacity framework to continuous actuator and sensor variables, where naive discretization becomes computationally difficult and general analytic solutions are unavailable.
- 4.6 Continuous Empowerment: Continuous motion and motor-control problems create large state and action spaces when represented through increasingly fine discretizations.The resulting number of options can make direct empowerment computation computationally prohibitive.
- 4.6 Continuous Empowerment: Channel capacity remains well defined for continuous input and output spaces, but introduces conceptual differences from the discrete case.The chapter treats discrete time while allowing actuator and sensor variables to be compound vectors across multiple time steps.
- 4.6 Continuous Empowerment: Noise-free continuous actions and sensors can transmit arbitrarily precise real-valued information, making channel capacity theoretically infinite.A deterministic continuous channel can reproduce real-valued inputs exactly through its outputs.
- 4.6 Continuous Empowerment: Because no general analytic solution exists for continuous channel capacity, the chapter reviews approximation methods, emphasizing a fast quasi-linear Gaussian approximation.It also briefly discusses naive binning and Monte Carlo integration.
- 4.6 Continuous Empowerment: Differential entropy may be infinite or negative, but differences of differential entropy terms still yield mutual information with the usual interpretation.The continuous mutual information is written as I(X; Y) := h(X) − h(X|Y).
4.6.2 Infinite Channel Capacity
Continuous channels can have infinite capacity under deterministic, noiseless copying, so empowerment requires approximation or additional modeling assumptions for practical computation.
- 4.6.2 Infinite Channel Capacity: For some continuous conditional distributions, channel capacity becomes infinite because differential entropy can reach negative infinity.A Dirac δ distribution is concentrated at a single point and represents the noiseless deterministic case.
- 4.6.2 Infinite Channel Capacity: A noiseless channel that copies s = a has H(S) = H(A), which is the largest possible value and therefore the channel capacity.The exact-copy example demonstrates why continuous mutual information can diverge.
- 4.6.2 Infinite Channel Capacity: Continuous mutual information can become infinite while remaining consistent with the ability of continuous variables to store unlimited information.The chapter notes that this does not substantially change the interpretation of mutual information.
- 4.6.2 Infinite Channel Capacity: Practical empowerment computation therefore approximates the world model, beginning with discretization and considering binning, Monte Carlo integration, and other estimators.The chapter presents these approaches as ways to compute empowerment when exact continuous capacity is unavailable.
- 4.6.2 Infinite Channel Capacity: Adaptive binning requires consistent bins across measurements, while sparse samples can create spurious nonzero mutual information.Both binning approaches require samples from the relevant continuous conditional distribution.
4.6.5 Evaluation of Binning
Binning approximates empowerment by discretizing action outcomes, but arbitrary boundaries and granularity can distort mutual-information estimates. The Monte Carlo alternative instead samples outcomes from a transition model and iteratively estimates empowerment and its maximizing action distribution.
- Binning limitations: Binning can create empowerment differences that are artefacts of arbitrary state boundaries rather than true sensor resolution.With proper rounding, an agent moving by 0.2 can appear more empowered at 1.5 than 1.0 despite a uniform underlying continuous structure.
- Binning limitations: Choosing too few bins loses structural correlations, whereas too many produce sparse samples that can significantly overestimate mutual information.Some bins may contain only one sample or none.
- Monte Carlo Integration: Monte Carlo Integration estimates empowerment by sampling outcomes from representative available action sequences.It can handle generic conditional distributions p(s|a).
- Monte Carlo Integration: The method requires a transition model that evaluates how probable each sampled outcome is under every action, providing a distinguishability-related distance measure.A Gaussian transition model specifies each action-conditioned outcome distribution through its mean and covariance.
- Monte Carlo Integration: The iterative algorithm initializes a uniform action distribution, updates action weights and an empowerment estimate, and outputs the estimated empowerment with the maximizing distribution.Iteration stops when successive estimates differ by less than ϵ or the iteration limit is reached.
4.6.7 Evaluation of Monte Carlo Integration
Monte Carlo Integration avoids binning artefacts while retaining general outcome distributions, but it introduces a noise-model assumption and can be computationally impractical for online robotic use.
- Advantages and limitations: Monte Carlo Integration removes artefacts caused by arbitrary bin boundaries while remaining applicable to generic distributions p(s|a).The Jung et al. solution assumes Gaussian noise.
- Advantages and limitations: Its computational cost grows with the number of representative action sequences needed for accurate approximations.Applications in Jung et al. (2011) were computed off-line and were considered infeasible for robotic applications.
4.6.8 Quasi-Linear Gaussian Approximation
The quasi-linear Gaussian approximation accelerates continuous empowerment by replacing a nonlinear noisy actuation-sensing mapping with parallel Gaussian channels when local linearity and Gaussian-noise assumptions are adequate.
- Model assumptions: The approximation models perception as a deterministic actuation mapping plus Gaussian noise, with noise potentially dependent on the current state.The noise is part of the channel model and limits channel capacity.
- Model assumptions: The actuation-to-perception mapping must be sufficiently well approximated by an affine function over the allowed action distributions.The chapter assumes this informally without making the notion precise or deriving error bounds.
- Approximation: Under linear-transformation and additive-Gaussian-noise assumptions, continuous channel capacity reduces to parallel Gaussian channels solvable with established algorithms.This reduction provides the quasi-linear Gaussian approximation for empowerment.
- Approximation: A quadratic action-power constraint prevents unbounded action amplitudes from making all outcomes distinguishable and empowerment infinite.The constraint is treated as a physically plausible limitation, while its choice is motivated by tractability in this approximation.
4.6.9 MIMO channel capacity
The MIMO formulation decomposes a linear Gaussian empowerment channel into independent parallel channels using singular-value decomposition. Capacity is then optimized by allocating available power across channels according to their noise levels.
- Channel model: With independent Gaussian noise across sensor dimensions, the system can be interpreted as multiple sensor channels with dimension-specific noise.Equal unit variance yields a linear MIMO channel with additive isotropic Gaussian noise.
- Scope: The quadratic power constraint is a tractability choice in the quasi-linear Gaussian case rather than a general consequence of empowerment’s informational definition.Other constraints may be more appropriate for particular real systems.
- MIMO reduction: Singular-value decomposition transforms the MIMO channel into independent dimensions, reducing total capacity computation to parallel Gaussian-channel capacities.The transformed variables use unitary rotations of sensor, action, and noise vectors.
- Power allocation: Gaussian input distributions achieve capacity, and water-filling allocates power first to the least noisy channels before distributing it more broadly.Depending on total available power, some channels may receive no power.
- Noise requirement: Noise variance must be positive because vanishing noise makes channel capacity infinite; practical applications generally include actuator, system, or sensor noise.Noise creates overlap between outcome states, allowing meaningful empowerment values.
4.6.10 Coloured Noise
Coloured Gaussian sensor noise can be transformed into an equivalent isotropic-noise channel without changing mutual information or channel capacity.
- 4.6.10 Coloured Noise: Coloured noise is modeled as zero-mean Gaussian noise with covariance matrix K_s across sensor dimensions.The covariance captures dependencies between noise components.
- 4.6.10 Coloured Noise: Rotation, translation, and scaling preserve mutual information, enabling reduction to an i.i.d. isotropic-noise channel.The transformation uses the singular value decomposition of K_s.
- 4.6.10 Coloured Noise: All noise singular values must be strictly positive; otherwise a zero-noise channel would yield infinite capacity.The empowerment maximizer could otherwise place all power in the noiseless component.
- 4.6.10 Coloured Noise: After redefining the transformation matrix, the problem becomes S = TA + Z with isotropic Gaussian noise and unchanged channel capacity.This permits use of the channel-capacity solution described for isotropic noise.
4.6.11 Evaluation of QLG Empowerment
The quasi-linear Gaussian approximation makes continuous empowerment fast to compute, but its Gaussian, local-linearity, and power-constraint assumptions limit applicability.
- 4.6.11 Evaluation of QLG Empowerment: QLG empowerment’s computational bottleneck is a singular value decomposition with dimensions matching the sensors and actuators.This makes the approximation quick to compute.
- 4.6.11 Evaluation of QLG Empowerment: QLG assumes Gaussian noise and a locally linear action–sensor relationship, so it cannot represent locally nonlinear relationships.Abruptly emerging degrees of freedom are softened by the Gaussian actuation model.
- 4.6.11 Evaluation of QLG Empowerment: The approximation introduces the power constraint P as an additional free parameter.Its effects are examined in a later pendulum example.
4.7 Continuous Examples
Continuous empowerment control selects actions by expected successor empowerment, producing behaviors that depend strongly on dynamics, horizon, power, and model availability.
- 4.7. Continuous Examples: Greedy control selects actions according to the immediate expected empowerment of their successor states, but empowerment is not a true cumulative value function.The expectation weights each successor state by p(s|a).
- 4.7. Continuous Examples: Continuous actions require sampling candidate actions, such as maximal positive, zero, and maximal negative acceleration, before selecting among them.The sampled actions are compared through their resulting successor empowerment.
- 4.7. Continuous Examples: Varying time-step length and power constraint can make the pendulum swing upright, oscillate cyclically, or rest in the lower position.A longer time step provides more look-ahead but worsens the local-linear approximation.
- 4.7. Continuous Examples: Pendulum empowerment can decrease along a greedy trajectory because acceleration is controllable while position changes are mediated by current velocity.Dynamics can force transitions through lower-empowerment regions toward later higher-empowerment states.
- 4.7. Continuous Examples: The success of local greedy optimization remains an open question when sharp empowerment changes lie beyond the local action-sequence horizon.The pendulum’s upright basin appears sufficiently broad for local optimization to find it, but the required dynamical properties are uncharacterized.
- 4.7. Continuous Examples: Increasing power generally increases empowerment, but changing power can invert the ranking of states in the empowerment landscape.At low power the largest singular value dominates; at higher power, the combination of singular values matters.
- 4.7. Continuous Examples: Empowerment’s dynamics are sensitive to the power constraint, with higher-power pendulum examples remaining in the lower rest position.The cited interpretation contrasts fine-tuned behavior for weak actuators with simpler behavior for strong actuators.
- 4.7. Continuous Examples: Computing empowerment requires a local actuator-to-sensor dynamics model, whose acquisition or adaptation lies outside the formalism.The model may be supplied in advance or learned using another technique.
4.8 Conclusion
Empowerment applies one task-independent principle across diverse agent–world behaviors, while its usefulness is bounded by local modeling, computational cost, and the absence of explicit goals.
- 4.8 Conclusion: The same empowerment principle has organized swarms, controlled pendulums, selected maze locations, approached manipulable objects, and shaped sensor evolution.These examples illustrate the formalism’s cross-domain universality.
- 4.8 Conclusion: Task independence can produce observer-approved default behaviors, but integrating explicit non-default goals remains fully open.Empowerment is therefore not necessarily suitable when a desired goal differs from its default behavior.
- 4.8 Conclusion: The local character of empowerment reduces computation and model-acquisition costs, but its proxy value depends on local dynamics reflecting global dynamics.Only a small part of the state space needs exploration under the local assumption.
- 4.8 Conclusion: Computational feasibility remains a recurring problem, motivating methods for continuous empowerment and deeper horizons.The limitation affects both behavioral hypotheses and real-time AI or robotics applications.
- 4.8 Conclusion: Because empowerment often produces intuitive default behaviors, its relationship to default adaptations in nature remains relevant to investigate.The chapter presents empowerment as a candidate bridge between biological adaptation and artificial devices.