Source-linked AI summary
Plasticity as the Mirror of Empowerment
David Abel, Michael Bowling, André Barreto, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh
TL;DR
The paper asks how agents can be shaped by observations, complementing the established question of how agents influence future observations. It develops generalized directed information to define plasticity, showing that plasticity mirrors empowerment and that the two capacities face a formal tension. The authors argue that this relationship matters for understanding and designing agents.
Problem
The paper addresses the lack of a general agent-centric measure for how observations influence an agent, alongside the established measure of empowerment.
Method
The paper defines plasticity using generalized directed information, which extends directed information to arbitrary sequence intervals while preserving key properties.
Results
Plasticity is the mirror of empowerment because both use the same measure with reversed influence direction, and their achievable levels are constrained by a formal tension.
Takeaways & Limitations
Agent analysis and design should consider plasticity and empowerment together because emphasizing one can affect the other.
Takeaways & Limitations
The proposed desiderata do not completely axiomatize plasticity, so subtle aspects of the definition may require future adjustment.
Abstract
from arXiv · showhide
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has served as a vital framing concept across artificial intelligence and cognitive science. This former capacity, however, is equally foundational: In what ways, and to what extent, can an agent be influenced by what it observes? In this paper, we ground this concept in a universal agent-centric measure that we refer to as plasticity, and reveal a fundamental connection to empowerment. Following a set of desiderata on a suitable definition, we define plasticity using a new information-theoretic quantity we call the generalized directed information. We show that this new quantity strictly generalizes the directed information introduced by Massey (1990) while preserving all of its desirable properties. Under this definition, we find that plasticity is well thought of as the mirror of empowerment: The two concepts are defined using the same measure, with only the direction of influence reversed. Our main result establishes a tension between the plasticity and empowerment of an agent, suggesting that agent design needs to be mindful of both characteristics. We explore the implications of these findings, and suggest that plasticity, empowerment, and their relationship are essential to understanding agency
1 Introduction
The paper frames plasticity—the capacity to be shaped by observations—as a foundational counterpart to empowerment, the capacity to influence future observations. It formalizes plasticity with generalized directed information and shows that plasticity mirrors empowerment while creating a design tension between them.
- 1 Introduction: Plasticity measures how an agent can be shaped by what it observes, complementing empowerment’s measure of how the agent can influence its observable future.The paper presents both capacities as fundamental aspects of agency across designs, substrates, and goals.
- 1 Introduction: Generalized directed information (GDI) strictly generalizes Massey’s directed information and provides the formal basis for defining plasticity under minimal assumptions.The paper develops GDI to support arbitrary sequence intervals while retaining directed information’s relevant structure.
- 1 Introduction: Plasticity is the mirror of empowerment: both use the same measure, with the direction of influence reversed between agent and environment.The paper attributes this equivalence to the symmetry of exchanging agent and environment symbols and functions.
- 1 Introduction: The paper establishes a tension between plasticity and empowerment, motivating agent designs that consider both characteristics rather than optimizing only one.The introduction identifies this connection and tension as having implications for agent design and analysis.
- 1 Introduction: The paper also derives necessary and sufficient conditions for non-zero plasticity and several further properties, providing a toolkit for studying agents under minimal assumptions.These results are presented as consequences of the paper’s formal development of plasticity.
2 Preliminaries: Directed Information, Agents, and Empowerment
The preliminaries model agent–environment interaction as bidirectional information exchange over finite action and observation interfaces. They use directed information to characterize influence and empowerment, then motivate GDI as the needed generalization.
- 2 Preliminaries: Directed Information, Agents, and Empowerment: Directed information captures influence between sequences and is especially suited to agent–environment systems with feedback.It is non-negative, upper-bounded by mutual information, and decomposes mutual information into opposing directions of influence.
- 2 Preliminaries: Directed Information, Agents, and Empowerment: An interface consists of finite action and observation sets, while agents and environments are stochastic functions mapping histories to actions and observations, respectively.Swapping the action and observation sets converts an agent into an environment and vice versa.
- 2 Preliminaries: Directed Information, Agents, and Empowerment: Empowerment quantifies an agent’s ability to influence observable outcomes, extending from open-loop mutual information to closed-loop directed information.The closed-loop formulation maximizes directed information over a richer class of agent functions.
3 Generalized Directed Information
The paper introduces generalized directed information to measure influence between arbitrary sequence intervals, including partially overlapping or disjoint windows. GDI recovers directed information in the standard case while preserving decomposition, temporal consistency, and conservation properties.
- 3 Generalized Directed Information: GDI extends directed information to sequences of arbitrary lengths and start times, addressing the standard measure’s equal-length, time-zero restriction.It is designed for systems with bidirectional communication and flexible observation windows.
- 3 Generalized Directed Information: When both intervals begin at time one and have the same endpoint, GDI recovers directed information, so it strictly generalizes the original quantity.This is stated as Proposition 3.2.
- 3 Generalized Directed Information: GDI is temporally consistent: if the source interval begins after the target interval ends, the directed influence is zero.Thus, non-zero GDI requires at least some source variables to precede target variables.
- 3 Generalized Directed Information: GDI can be decomposed or subdivided across interval boundaries while preserving the total quantity.The paper shows that each resulting term remains a valid GDI term.
- 3 Generalized Directed Information: The paper proves a GDI conservation law analogous to directed information’s conservation law, with conditioning used to remove potential confounders from earlier variables.The conservation result is identified as critical for analyzing the plasticity–empowerment tension.
- 3 Generalized Directed Information: Additional GDI properties, including an extension of the data-processing inequality, are deferred to the appendix because of space constraints.This is a scope boundary on the presentation rather than a claim that the properties are absent.
4 Plasticity and Empowerment as Generalized Directed Information
The paper defines plasticity with generalized directed information (GDI), establishing when observations influence actions and relating that capacity to empowerment. It shows that plasticity mirrors empowerment, while both capacities face a tight trade-off over shared intervals.
- 4.1 Plasticity: Plasticity measures how much an observation sequence influences an agent’s action sequence, with nonzero plasticity exactly when the observations provide information about those actions.Open-loop agents and agents whose actions depend only on history length or past actions have zero plasticity.
- 4.1 Plasticity: Generalized directed information extends directed information to arbitrary, noninitially aligned sequences and supports a flexible definition of plasticity.The GDI preserves directed information as a special case and retains key structural properties.
- 4.2 Revisiting Empowerment through GDI: Plasticity is the mirror of empowerment: both are defined by GDI between action and observation, with the direction of influence reversed.Because agents and environments are symmetric at the interface, an agent’s empowerment equals the environment’s plasticity, and vice versa.
- 4.3 Main Result: Plasticity and Empowerment Tension: Some environments force every interacting agent to have zero plasticity or empowerment, because they can ignore actions or provide no information.The paper illustrates the mirror relationship with two interacting agents: one agent’s plasticity equals the other’s empowerment.
- 4.3 Main Result: Plasticity and Empowerment Tension: The main theorem gives a tight upper bound on the sum of plasticity and empowerment, with m = min{(b − a + 1) log |O|, (d − c + 1) log |A|}.Under mild interface and interval assumptions, the bound is achievable by pairs that maximize one quantity while forcing the other to zero.
- 4.3 Main Result: Plasticity and Empowerment Tension: An agent cannot simultaneously maximize plasticity and empowerment over the same pair of intervals.The achievable level of either quantity is strictly determined by the other, and extreme plasticity can coincide with zero empowerment, or vice versa.
- 4.4 A Simple Experiment: In Q-learning experiments, lower ϵ produces higher plasticity, while varying the initial Q value tests how optimism and pessimism affect plasticity and empowerment.The experiments estimate these quantities over intervals [1 : 3] and [2 : 5], and compare their sum with the theorem’s upper bound.
5 Discussion
The discussion frames generalized directed information as a foundation for defining plasticity and reveals plasticity–empowerment trade-offs, scope boundaries, and future research directions.
- Contributions: The generalized directed information supports a general definition of plasticity, characterizes when plasticity is non-zero, and establishes a trade-off with empowerment.The paper presents these results as a toolkit for studying agency under minimal assumptions.
- Plasticity via Action, Policy, or Agent State?: The paper's behavioral plasticity measure captures how observations influence actions, while the same formalism can instead target policies or internal agent state.The action-focused version is chosen for simplicity, but policy- and state-focused variants preserve much of the formal structure.
- Optionality, Information, and Causality: Every deterministic environment forces zero behavioral plasticity because the environment offers no alternative unfoldings to which the agent can react.Memory limits can make an environment appear stochastic from the agent's perspective, motivating internal-state variants of plasticity.
- The Plasticity-Empowerment Tension: In the corridor example, rooms interpolate between maximal plasticity and maximal empowerment as control over the light shifts between the environment and the mouse.With a non-stationary latent process, the controllable end closes off learning about its evolution, whereas the environment-driven end maximizes that learning potential.
- Future Directions: The framework omits goals and treats incorporating goal-directedness as a major direction for future work.The authors also suggest balancing plasticity and empowerment, since over-optimizing one can affect the other.
A Notation
Appendix A introduces notation and points readers to a consolidated notation table.
- A Notation: Table 1 summarizes the notation used throughout the paper.The appendix states that the table collects all notation and core definitions.
B Proofs of Presented Results
Appendix B organizes proofs of the paper's presented results into results on generalized directed information and results on plasticity and empowerment.
- B Proofs of Presented Results: The proofs are divided between Section 3 results on generalized directed information and Section 4 results on plasticity and empowerment.Appendix C separately presents additional generalized-directed-information results, including a data-processing inequality.
B.1 Proofs of Results from Section 3 on GDI
The appendix proves the principal generalized-directed-information results, including strict generalization, interval decomposition, temporal consistency, and conservation.
- GDI Results: Proposition 3.2 establishes that generalized directed information strictly generalizes directed information when both intervals span the full sequence.The proof is identified in the appendix, while the proposition states the generalization result.
- GDI Results: Proposition 3.4 decomposes generalized directed information over either output or input intervals, providing the summation identities used in later proofs.The appendix proves the identities by splitting the defining sums across interval boundaries and recombining terms.
- Conservation Law: The conservation-law proof decomposes both directional GDI terms into boundary, overlapping, and conditional mutual-information components before recombining them.The proof uses an interval decomposition lemma, induction on interval length, and the mutual-information chain rule.
- Core Properties: Theorem 3.5 states the conservation law of generalized directed information, while Proposition 3.3 establishes temporal consistency through zero terms for non-overlapping intervals.The temporal-consistency proof follows from the summation bounds and the condition that later input intervals cannot influence earlier output intervals.
B.2 Proofs of Results from Section 4 on Plasticity and Empowerment
The proofs establish when plasticity is positive, characterize agents with zero plasticity, and derive its mirror relationship with empowerment and their shared tight upper bound.
- Lemma 4.2: Plasticity is non-zero exactly when an agent’s actions are influenced by observations at some time-step.The proof establishes necessity and sufficiency through conditional mutual information between observations and actions.
- Theorem 4.3: Theorem 4.3 shows plasticity is nonnegative, can be positive for a deterministic agent, and is monotone with respect to enlarging the environment set.It also identifies several agent classes with zero plasticity for every environment set.
- Theorem 4.3: Closed-loop agents, constant-policy agents, history-length agents, and agents depending only on past actions have zero plasticity relative to every environment set.For these agents, conditioning on observations does not change the action distribution.
- Proposition 4.6: The agent’s empowerment equals the environment’s plasticity, while the agent’s plasticity equals the environment’s empowerment.The proof obtains both equalities by applying the same information measure with the direction of influence reversed.
- Theorem 4.8: m = min{(b −a + 1) log |O|, (d −c + 1) log |A|} is a tight upper bound on both empowerment and plasticity, and attaining it forces the other quantity to zero.The tightness construction uses an environment emitting uniform observations and an agent that mirrors observations as actions under stated interface and interval conditions.
C Other Results
The appendix establishes a data-processing inequality for generalized directed information, derives upper-bound and entropy decompositions, and recovers Kramer’s identity as a special case.
- Theorem C.1: The generalized directed information satisfies a data-processing inequality when each Z_i is conditionally independent of X_1:i given Y_i.The proof follows the standard mutual-information argument with a base case and induction over both intervals.
- Proposition C.2: The generalized directed information is nonnegative and upper bounded through its conservation law and standard conditional-mutual-information and entropy bounds.These properties support the bound used elsewhere for empowerment and plasticity.
- Proposition C.3 and Corollary C.4: The generalized directed information decomposes into conditional entropy terms, providing a generalized Kramer decomposition for arbitrary indices.When a = c = 1 and b = d = n, the appendix recovers Kramer’s identity as a special case.