Source-linked AI summary

Driven by Compression Progress: A Simple Principle Explains Essential Aspects of Subjective Beauty, Novelty, Surprise, Interestingness, Attention, Curiosity, Creativity, Art, Science, Music, Jokes

Juergen Schmidhuber

arXiv:0812.4360v2cs.AIcs.NE

TL;DR

The paper asks how computationally limited, self-improving observers can find data that becomes interesting through improved prediction or compression. It proposes curiosity as intrinsic reward for compression progress and argues that this principle qualitatively explains attention, novelty, creativity, art, science, and related phenomena. The framework also motivates artificial-agent and psychological tests while depending on assumptions about predictability and the limits of probabilistic compression.

  • Problem

    The paper addresses how subjective interest, curiosity, attention, and creativity can be explained for computationally limited observers seeking regularities whose compressibility is not yet known.

  • Method

    The paper proposes a framework combining a continually improving predictor or compressor, a computable compression-progress reward, and reinforcement learning that selects actions to maximize expected reward.

  • Results

    The paper argues that compression progress qualitatively explains fundamental aspects of attention, novelty, surprise, interestingness, curiosity, creativity, subjective beauty, jokes, science, and art.

  • Takeaways & Limitations

    The framework provides a common account of passive observation and creative action, and motivates implementations in artificial systems plus controlled psychological and neurophysiological tests.

  • Takeaways & Limitations

    The approach assumes that better explanations of the past improve future prediction and task search, while probabilistic compression may miss broader algorithmic regularities such as those in π.

Abstract

from arXiv · show

I argue that data becomes temporarily interesting by itself to some self-improving, but computationally limited, subjective observer once he learns to predict or compress the data in a better way, thus making it subjectively simpler and more beautiful. Curiosity is the desire to create or discover more non-random, non-arbitrary, regular data that is novel and surprising not in the traditional sense of Boltzmann and Shannon but in the sense that it allows for compression progress because its regularity was not yet known. This drive maximizes interestingness, the first derivative of subjective beauty or compressibility, that is, the steepness of the learning curve. It motivates exploring infants, pure mathematicians, composers, artists, dancers, comedians, yourself, and (since 1990) artificial systems.

1 Store & Compress & Reward Compression Progress

The framework treats curiosity as compression progress: agents store experience, improve predictive compression, reward saved bits, and choose actions that seek further learnable regularities. It motivates active exploration even when external rewards are rare, while leaving the specific compressor and reinforcement learner open.

  • Algorithmic Framework: The framework encourages active exploration for previously unknown regularities, even when external reward is rare or absent.Its objective is to improve understanding and predictability through intrinsic rewards for discoveries in action-dependent data.
  • Algorithmic Framework: Agents should store the complete raw history of actions, sensory observations, and reward signals as the basis for knowledge.The paper argues that full storage is feasible for human-scale sensory streams and increasingly affordable technical systems.
  • Algorithmic Framework: Adaptive compression improves subjective explanations by discovering regularities that reduce the description length of historical data.The proposed compressor may learn to predict or postdict historical data from other historical data.
  • Algorithmic Framework: Intrinsic curiosity reward is proportional to compression progress, measured by the number of bits saved when encoding historical data.The agent monitors improvements in its adaptive compressor and rewards learning progress.
  • Algorithmic Framework: Reinforcement learning selects actions expected to maximize intrinsic curiosity reward by directing attention toward aspects of the world that enable new discoveries.The controller observes the adaptive compressor and seeks action sequences associated with future reward.
  • Algorithmic Framework: The framework specifies objectives rather than a unique implementation, leaving the adaptive compressor and reinforcement-learning algorithm open.The paper presents possible concrete instances separately rather than fixing one compressor-controller combination.

2 Consequences of the Compression Progress Drive

The compression-progress principle links subjective beauty, novelty, curiosity, art, science, and consciousness to an observer’s ability to discover or exploit regularities. Interestingness is temporary: it rises while newly discovered regularities continue improving subjective compressibility.

  • Compression and consciousness: Efficient compression creates internal representations for recurring patterns, including a code representing the agent itself; activating that representation can constitute self-awareness or consciousness.Consciousness is presented as a by-product of ongoing world modeling and problem solving through data compression.
  • Subjective beauty: Subjective beauty depends on the shortest description available to the observer’s current encoding method, making regular, symmetric, or prototype-like patterns especially beautiful.A new face close to an internal prototype requires fewer encoded deviations than a dissimilar face.
  • Subjective beauty: Observers may prefer faces resembling their own because frequently seen faces influence the subjective prototype and improve coding efficiency.
  • Interestingness and novelty: Interestingness is the first derivative of subjective beauty: data remains rewarding while learning makes its regularities increasingly compressible.Once a regularity is fully assimilated, the associated data becomes less interesting even if it remains beautiful.
  • Interestingness and novelty: True novelty lies between predictability and randomness: dark, unchanging input is boring because it is already compressible, while white noise is boring because it offers no compressible regularity.The framework therefore rejects surprise defined solely by Boltzmann or Shannon information.
  • Art and science: Discoveries are unusually large compression breakthroughs, such as gravity’s short description of many observations; art similarly connects patterns so their combined description becomes shorter.Artists and scientists both seek non-random, non-arbitrary data with previously unknown regularities, although science formalizes the regularity while art may leave it subconscious.
  • Intrinsic and external reward: The framework separates intrinsic compression-progress rewards from external rewards such as praise or money, while defining beauty and interestingness solely through intrinsic learning progress.This scope excludes pleasure based on external rewards, including social approval or associations with pleasurable memories.

3 Previous Concrete Implementations of Systems Driven by (Approximations of) Compression Progress

Previous implementations approximated compression-progress motivation through predictive, probabilistic, and program-based systems. The paper emphasizes improving prediction or compression rather than rewarding persistent error, while noting limits of naive statistical methods.

  • Early predictive systems: Early artificial-curiosity systems used recurrent neural networks to predict sensory inputs and rewards, with curiosity rewards proportional to prediction errors.This implicitly assumed that high prediction error would lead to future predictor improvement.
  • Learning progress versus surprise: Follow-up work argued that probabilistic environments require rewarding predictor improvements rather than errors, because noise and randomness can sustain high errors without improving compressibility.
  • Probabilistic variants: A 1995 information-theoretic variant measured curiosity through the Kullback-Leibler distance between subjective prior and posterior probability distributions after new observations.
  • Probabilistic variants: In 2005, Bayesian surprise was reported to explain certain human visual-attention patterns better than certain earlier approaches.
  • Limits of probabilistic compression: Naive probabilistic compression cannot discover general algorithmic regularities such as the short generating program for π, requiring broader program-search techniques.Finite π digit sequences can have random-looking statistics even though the entire expansion is algorithmically computable.
  • Program-based systems: Later systems increased controller and predictor power through co-evolving, symmetric, opposing modules based on self-modifying probabilistic programs in a universal programming language.A related approach framed the method as system identification through co-evolution of computable models and tests.
  • Empirical consequence: The cited implementations demonstrated experimentally that intrinsic or curiosity rewards can speed the collection of external reward.

4 Visual Illustrations of Subjective Beauty and its First Derivative Interestingness

Subjective beauty is illustrated as the result of discovering simple, learnable regularities in visual data. The reward comes from compression progress—the steepness of the learning curve—rather than from static simplicity alone.

  • The framework treats unsupervised attention and artistic creativity as by-products of compression-progress drives in human observers.
  • A female face is considered beautiful by some observers because its essential features follow a simple geometrical pattern requiring few bits to specify.
  • A butterfly and flower vase can be generated by a simple algorithm based on fractal circle patterns, and understanding that algorithm increases appreciation.
  • Beauty increases as observers discover how a visual pattern can be described more compactly, moving from a longer to a shorter description.
  • The associated reward reflects the first derivative of subjective beauty: how rapidly the observer’s compression or understanding improves.

5 Conclusion & Outlook

The paper proposes compression progress as a simple principle for explaining diverse forms of curiosity and creativity. It outlines artificial-agent, psychological, and neurophysiological tests, while noting that successful validation could motivate robot implementations.

  • Compression and compression progress are presented as a unifying principle for attention, novelty, surprise, interestingness, curiosity, creativity, beauty, jokes, science, and art.
  • The formal framework combines an improving predictor or compressor, a computable progress measure for intrinsic rewards, and reinforcement learning to select rewarding actions.
  • Controlled psychological experiments are proposed to test whether intrinsic reward peaks when predictions improve most rapidly and vanishes when improvement stops.
  • Neurophysiological studies could localize intrinsic rewards and relate them to improvements in neural predictors.
  • Success in testing the principle would provide additional motivation to implement it on robots.

A Appendix

The appendix frames curiosity as an intrinsic reward for discovering compressible structure in an agent’s history. It formalizes the agent-environment setting and separates curiosity from externally supplied reward.

  • Discoveries correspond to large compression improvements found by an application-dependent compressor-improvement algorithm.
  • The agent operates in discrete time, receiving real-valued environmental inputs and executing real-valued actions that may affect future inputs.
  • The history h(t) contains the current input, action, and reward, while expected utility is evaluated under a possibly unknown environmental distribution.
  • The reward is split into external and intrinsic components, with this paper focusing especially on intrinsic or curiosity reward.
  • The paper initially assumes curiosity is useful, sets external reward to zero, and generates curiosity reward when the predictor or history compressor improves.
  • Reinforcement learning translates the compression goal into action sequences that expose previously unknown but learnable regularities.

A.1 Predictors vs Compressors

Prediction and compression are closely related because accurate predictions can encode history compactly. However, MDL-based compressor-oriented prediction need not converge as quickly as Solomonoff induction, although both converge in the limit under general conditions.

  • A predictor that correctly predicts many historical inputs can encode the history compactly using only prediction errors and their time steps.
  • MDL-based compressor-oriented prediction does not necessarily converge to correct predictions as quickly as Solomonoff’s universal inductive inference.
  • Both MDL-based prediction and Solomonoff induction converge in the limit under general conditions.

A.2 Which Predictor or History Compressor?

The framework allows predictors ranging from simple linear mappings to more complex adaptive recurrent neural networks using longer histories.

  • Predictors can use a linear mapping to predict x(t + 1) from x(t) and y(t + 1).
  • More complex adaptive RNN predictors can use nonlinear mappings and possibly the entire history h(≤t).

A.3 Compressor Performance Measures

Compressor performance is evaluated on the observed history, with shorter compressor programs representing greater compressibility and predictive regularity.

  • At time t, C(p, h(≤t)) denotes a compressor program’s performance on the history observed so far.
  • The compressor description length l(p) is measured in bits, so shorter programs indicate greater algorithmic regularity, compressibility, predictability, and lawfulness.
  • The ultimate limit is K*(h(≤t)), defined as the length of the shortest program that computes an output beginning with the observed history.

A.4 Compressor Performance Measures Taking Time Into Account

Time-aware compressor measures account for the computation time required to process the history, trading storage efficiency against runtime.

  • An alternative performance measure incorporates the time τ(p, h(≤t)) spent by compressor p computing h(≤t).
  • From an asymptotic optimality perspective, this approach trades off storage and computation time.
  • Compression by one bit is treated as equivalent to reducing runtime by a factor of 1

A.5 Measures of Compressor Progress / Learning Progress

Curiosity depends on compressor improvement rather than absolute compression performance: reward compares old and new compressors on the same history, often measuring the gain as a difference.

  • Curiosity reward responds to the compressor’s performance improvement between times t and t + 1, not its performance alone.
  • The intrinsic reward is computed by applying f to the old and new compressors’ performance on history h(≤t + 1).
  • Using f(a, b) = a − b yields a discrete-time version of maximizing the first derivative of subjective data compressibility.
  • Both compressors must be tested on the same history so the reward reflects progress rather than differences in evaluated data.

A.6 Asynchronous Framework for Creating Curiosity Reward

The framework separates an acting controller from an asynchronously improving compressor, turning compression progress into intrinsic rewards that guide action selection. It also contrasts computable approximations with theoretically optimal but generally uncomputable predictors and policies.

  • Controller: The controller selects actions, observes new inputs, incorporates any available intrinsic reward, and updates its policy through reinforcement learning.The controller operates from the history and may also use the latest compressed data.
  • Compressor: An asynchronous compressor repeatedly stores the current history, evaluates its old compressor, improves it, and evaluates the resulting new compressor.Compressor improvement can use adaptive predictors such as neural networks and may take many time steps.
  • Caveat: The asynchronous design can create long delays between actions and curiosity rewards, making temporal credit assignment difficult for the controller.The text notes that specialized reinforcement-learning algorithms may still address this burden.
  • Curiosity objective: Curiosity rewards are generated from improvements in compression, directing reinforcement learning toward previously unknown but learnable regularities.Pure curiosity is defined relative to the computational limitations of the chosen compressor class.
  • Optimal action selection: Solomonoff mixture prediction provides a theoretically optimal basis for action selection, but its infinite distribution class makes the resulting scheme incomputable.AIXI extends this prediction scheme to select action sequences maximizing predicted future reward, while computable variants trade away practicality or incur slowdowns.
  • Self-improvement: Gödel machines can rewrite their own code after proving a rewrite useful, yielding globally optimal self-improvements when such utility is provable.They can also remove computational slowdowns hidden by asymptotic notation when the speed-up is provably useful.
Loading 0812.4360v2…