Source-linked AI summary
Embodied Intelligence via Learning and Evolution
Agrim Gupta, Silvio Savarese, Surya Ganguli, Li Fei-Fei
TL;DR
The principles linking environmental complexity, evolved morphology, and learnable intelligent control remain difficult to study at scale. DERL jointly evolves morphologies and learns controllers in increasingly complex environments, revealing that environmental complexity produces bodies that support faster, broader learning and that evolution rapidly selects faster-learning, stable, energy-efficient morphologies.
Problem
Relations between environmental complexity, evolved morphology, and the learnability of intelligent control remain elusive because large-scale evolution-and-learning experiments are difficult.
Method
DERL asynchronously evolves articulated morphologies while reinforcement learning trains controllers from low-level egocentric sensory observations across flat, variable, and manipulation-variable terrain.
Results
More complex environments produce morphologies that learn novel tasks better and faster, while evolution rapidly selects morphologies with faster learning, greater stability, and energy efficiency.
Takeaways & Limitations
The experiments suggest that morphological intelligence and a morphological Baldwin effect can emerge through bodies that simplify control by leveraging passive physical interactions.
Abstract
from arXiv · showhide
The intertwined processes of learning and evolution in complex environmental niches have resulted in a remarkable diversity of morphological forms. Moreover, many aspects of animal intelligence are deeply embodied in these evolved morphologies. However, the principles governing relations between environmental complexity, evolved morphology, and the learnability of intelligent control, remain elusive, partially due to the substantial challenge of performing large-scale in silico experiments on evolution and learning. We introduce Deep Evolutionary Reinforcement Learning (DERL): a novel computational framework which can evolve diverse agent morphologies to learn challenging locomotion and manipulation tasks in complex environments using only low level egocentric sensory information. Leveraging DERL we demonstrate several relations between environmental complexity, morphological intelligence and the learnability of control. First, environmental complexity fosters the evolution of morphological intelligence as quantified by the ability of a morphology to facilitate the learning of novel tasks. Second, evolution rapidly selects morphologies that learn faster, thereby enabling behaviors learned late in the lifetime of early ancestors to be expressed early in the lifetime of their descendants. In agents that learn and evolve in complex environments, this result constitutes the first demonstration of a long-conjectured morphological Baldwin effect. Third, our experiments suggest a mechanistic basis for both the Baldwin effect and the emergence of morphological intelligence through the evolution of morphologies that are more physically stable and energy efficient, and can therefore facilitate learning and control.
DERL: Computational framework for creating embodied agents
DERL decouples learning and evolution asynchronously to scale searches over embodied morphologies and controllers. Agents learn from low-level egocentric observations, enabling large parallel experiments across thousands of morphologies.
- DERL: Computational framework for creating embodied agents: DERL uses asynchronous tournament-based evolution to decouple lifetime learning from evolutionary selection.Each run begins with P = 576 agents with unique topologies, avoiding simultaneous replacement of the entire population.
- DERL: Computational framework for creating embodied agents: Agents receive low-level egocentric proprioceptive and exteroceptive observations and learn stochastic neural policies with PPO.
- DERL: Computational framework for creating embodied agents: The evolutionary dynamics preserve diverse high-fitness lineages, including descendants from founders with lower initial fitness.Lineage abundance can increase through accumulated beneficial mutations rather than high initial fitness alone.
- DERL: Computational framework for creating embodied agents: DERL scales experiments across 1152 CPUs, training 4000 morphologies with 5 million learning iterations per morphology.Training 288 morphologies in parallel allows the process to finish in less than 16 hours.
UNIMAL: A UNIversal aniMAL morphological design space
UNIMAL provides an expressive articulated-body design space for evolving diverse animal-like morphologies. Bilateral-symmetry-preserving mutations retain a physical simplification while spanning approximately 10^18 morphologies.
- UNIMAL: A UNIversal aniMAL morphological design space: UNIMAL represents agents as kinematic trees of articulated 3D rigid parts connected by motor-actuated hinge joints.Spheres form roots and cylinders form limbs; mutations grow or delete limbs and modify limb and joint properties.
- UNIMAL: A UNIversal aniMAL morphological design space: Agents evolved in variable and manipulation-variable terrain tend to be longer forward and shorter in height than flat-terrain agents.Flat-terrain agents are less space filling, while evolved agents have similar masses and degrees of freedom.
- UNIMAL: A UNIversal aniMAL morphological design space: Paired mutations preserve bilateral symmetry, placing every agent’s center of mass on the sagittal plane.This reduces the control required for left-right balancing.
- UNIMAL: A UNIversal aniMAL morphological design space: The design space contains approximately 10^18 unique morphologies with fewer than 10 limbs.
Successful evolution of diverse morphologies in complex environments
DERL evolves successful and diverse morphologies across flat, variable, and manipulation-variable environments. Environmental demands correspond to distinct body plans and movement or manipulation strategies.
- Successful evolution of diverse morphologies in complex environments: The environments increase in complexity from flat terrain to variable terrain and non-prehensile manipulation in variable terrain.Variable terrain samples new obstacle sequences, while manipulation-variable terrain adds box-to-target control through contact dynamics.
- Successful evolution of diverse morphologies in complex environments: DERL finds successful morphological solutions in all three environments while preserving diverse evolutionary lineages.Asynchronous tournaments allow lower-fitness founders to contribute abundant high-fitness descendants.
- Successful evolution of diverse morphologies in complex environments: Environmental conditions strongly shape morphology: variable-terrain agents are longer and shorter, while flat-terrain agents are less space filling.
- Successful evolution of diverse morphologies in complex environments: Across seven test tasks, manipulation-variable agents outperform flat-terrain agents after training selected morphologies from scratch.The test suite spans stability, agility, and manipulation, with median reward and cost-of-work measurements.
- Successful evolution of diverse morphologies in complex environments: Variable-terrain agents develop paired limbs for stability and maneuverability, whereas manipulation-variable agents evolve forward-reaching pincer- or claw-like arms.These arms guide boxes toward goal positions.
Environmental complexity engenders morphological intelligence
The study evaluates morphological intelligence by testing how well evolved bodies support learning novel tasks. More complex evolutionary environments produce morphologies that learn better and faster, alongside stability and energy-efficiency patterns.
- Environmental complexity engenders morphological intelligence: Morphological intelligence is measured by how much a body facilitates learning a suite of novel tasks.The approach parallels evaluating latent neural representations through downstream transfer tasks.
- Environmental complexity engenders morphological intelligence: Within 10 generations, average learning time to a criterion fitness is cut in half across all three environments.The study presents this rapid reduction as evidence for a morphological Baldwin effect.
- Environmental complexity engenders morphological intelligence: The eight test tasks span agility, stability, and manipulation, with controllers learned from scratch for each morphology.Learning from scratch isolates performance differences attributable to morphology.
- Environmental complexity engenders morphological intelligence: Across seven tasks, MVT morphologies outperform FT morphologies, while VT morphologies outperform FT in 5 out of 6 agility and stability tasks.VT and MVT morphologies show more pronounced advantages when learning iterations are reduced from 5 million to 1 million.
Demonstration of a stronger form of the conjectured morphological Baldwin effect
Evolution rapidly selected morphologies that learned faster, demonstrating a stronger morphological Baldwin effect without direct selection for learning speed.
- The result extends the long-conjectured Baldwin effect, which concerns learned behaviors becoming available earlier through evolutionary change.
- Evolution rapidly reduced the learning time needed to reach a criterion fitness across the top 100 agents in all three environments.Average learning time was cut in half within 10 generations.
- The eighth-generation agent achieved the first-generation agent’s final fitness in one-fifth the time and ended learning with twice its performance.
- Fitness was determined solely by performance at the end of learning, with no explicit selection pressure for faster learning.
- The study reports the first evidence for a morphological Baldwin effect in either in vivo or in silico morphological evolution.
A mechanistic underpinning for morphological intelligence and the strong Baldwin effect
The experiments link morphological intelligence and faster learning to evolved physical properties, especially energy efficiency and passive stability, in increasingly complex environments.
- The authors hypothesize that intelligent morphologies efficiently exploit passive physical dynamics in body–environment interactions.
- Evolution selected energy-efficient morphologies without direct selection for energy efficiency, while total body mass increased across all three environments.
- More energy-efficient morphologies performed better and learned faster at any fixed generation.
- Evolution also selected more passively stable morphologies over time, with a higher stable fraction in variable-terrain environments than flat terrain.
- Energy efficiency and stability improved over evolutionary time in a manner tightly correlated with learning speed.
- More complex-environment agents showed better performance on novel tasks, especially with fewer learning iterations, and were more energy efficient than flat-terrain agents.
Conclusion
DERL enabled large-scale studies showing how evolution, learning, and environmental complexity can produce morphologies that simplify control through passive physical interactions.
- DERL simulations yielded insights into interactions among learning, evolution, environmental complexity, and intelligent morphologies.
- Evolution transferred fitness from learned phenotypic ability to genotypically encoded morphology within a few generations through a Baldwin effect.
- Evolved morphologies supported better and faster learning in novel tasks, likely through increased passive stability and energy efficiency.
- Figures 7 and 8 show selected and example agent morphologies evolved across three evolutionary runs and different environments.
Appendix A: Evolutionary setup
DERL asynchronously evolves and trains articulated morphologies across environments of increasing complexity, using tournament selection, mutation, and lifetime reinforcement learning.
- Distributed Asynchronous Evolution: Each run begins with P = 576 agents having unique randomly generated morphologies, whose controllers are learned in parallel for 5 million interactions.End-of-lifetime reward over approximately the last 100,000 iterations defines morphological fitness.
- Distributed Asynchronous Evolution: Evolution asynchronously selects T = 4 agents for tournaments, mutates the highest-fitness morphology, and evaluates the child through lifetime learning.The child begins with a randomly initialized controller, so only morphology is inherited.
- UNIversal aniMAL (UNIMAL) design space: UNIMAL represents agents as articulated 3D rigid parts in a kinematic tree, with a sphere head and cylinders forming motor-actuated limbs.
- UNIversal aniMAL (UNIMAL) design space: Mutation operations grow or delete limbs and modify limb physical properties, density, joint degrees of freedom, gear ratios, and joint angles.
- UNIversal aniMAL (UNIMAL) design space: The UNIMAL hyperparameter choices estimate a morphology space of 10^18 possible morphologies.
- Environments: DERL evaluates flat terrain, variable terrain, and non-prehensile manipulation in variable terrain, with progressively more sub-tasks.
Appendix B: Learning algorithm
The learning algorithm uses egocentric sensory observations, task-specific rewards, and PPO to learn stochastic neural-network policies that maximize discounted cumulative reward.
- Reinforcement learning: At each time step, reinforcement learning maps an observation to an action and reward, optimizing expected cumulative reward R = Σ γ^t r_t over horizon H.The discount factor satisfies γ ∈ [0, 1).
- Observations: Agents receive low-level egocentric proprioceptive and exteroceptive observations, including morphology-dependent sensors, terrain, goals, objects, and obstacles.Terrain is represented as a non-uniform 2D heightmap sampled from behind the agent to ahead and laterally.
- Rewards: Locomotion rewards combine forward velocity with a weak penalty on large actuator inputs, while tournament selection compares only forward progress.The specified weights are w_x = 1 and w_c = 0.001.
- Rewards: Manipulation rewards combine progress toward the object, progress of the object toward its goal, and a weak actuator penalty.The distance terms use changes in geodesic distance, with w_ao = w_og = 100 and w_c = 0.001; a sparse reward is also provided near the goal.
- Policy architecture: Each stochastic policy uses deep policy and critic networks, with observation encodings concatenated before further processing.Each observation type is encoded by a two-layer MLP with hidden dimensions [64, 64], followed by a 64-dimensional representation.
- Optimization: PPO constrains the likelihood ratio between new and old policies, enabling multiple minibatch updates while reducing training instability and improving sample efficiency.Generalized Advantage Estimation supplies the advantage estimates.
Appendix C: Evaluation task suite
The evaluation suite measures how morphology supports learning across eight tasks spanning agility, stability, and manipulation, with controllers trained from scratch for each task.
- Evaluation design: Eight test tasks assess agility, stability, and manipulation, and controllers are learned from scratch so performance differences reflect morphology differences.The suite includes patrol, point navigation, obstacle, exploration, escape, incline, push box incline, and manipulate ball.
- Agility: Patrol requires repeated fast movement and direction changes between goals 10m apart, while point navigation requires reaching randomly placed goals from the arena center.Patrol flips the goal after the agent reaches within 0.5m and supplies a sparse reward of 10.
- Agility: Obstacle navigation requires traversing 50 randomly initialized obstacles with heights and bases varying from 0.5m to 3m.Obstacle information is provided through a terrain heightmap.
- Agility: Exploration maximizes the number of distinct 1 × 1m2 grid squares visited in a 100 × 100m2 arena.Its reward is sparse relative to dense locomotion rewards.
- Stability: Escape tests balance on random hilly terrain by maximizing geodesic distance from the initial location.The agent begins in a bowl-shaped terrain surrounded by small hills.
- Manipulation: Manipulate ball requires moving a radius-0.2m ball between random source and target locations while maintaining balance under complex contact dynamics.The agent starts at the center of a flat 30 × 30m2 arena.
Appendix D: Evaluation
Evaluation controls random-seed variation, measures energy efficiency through cost of work, assesses passive stability without control, and defines thresholds for beneficial mutations.
- Reporting methodology: Evolutionary runs use repeated random seeds, selecting the top three agents across three runs to identify seed-robust morphologies.Within each run, lifetime learning across morphologies shares the same seed; a run typically yields 15–20 surviving lineages.
- Energy efficiency: Cost of transportation measures energy per unit mass per unit distance, while cost of work adapts the normalization to energy per unit mass per unit reward.COT uses total energy E, mass M, distance D, and gravitational acceleration g.
- Energy efficiency: Cost of work is computed as energy divided by mass and reward, with lower COW indicating higher energy efficiency.Energy is measured as the absolute sum of all joint torques and applied to evolutionary and test environments.
- Stability: Passive stability is measured without control by checking whether the head remains above 50% of its original height after 400 time steps.This operationalizes standing stability using contact geometry and the support polygon.
- Beneficial mutations: A mutation is beneficial when child fitness exceeds parent fitness by at least 300 for FT or 100 for VT and MVT.The threshold excludes small RL-related changes judged potentially statistically meaningless under single-seed fitness evaluation.