Source-linked AI summary

Efficient collective swimming by harnessing vortices through deep reinforcement learning

Siddhartha Verma, Guido Novati, Petros Koumoutsakos

arXiv:1802.02674v1physics.flu-dyncs.AIphysics.comp-ph

TL;DR

The paper asks how fish can extract energetic benefits from complex vortex wakes during collective swimming, a question constrained by limited quantitative mechanistic evidence. It combines deep reinforcement learning with high-fidelity flow simulations to discover autonomous follower strategies, finding improved swimming efficiency through wake interception and synchronized body motion, with implications for robotic swarms.

  • Problem

    The energetic role and physical mechanisms of hydrodynamic interactions in fish schooling remain limited by insufficient quantitative information.

  • Method

    The study combines deep reinforcement learning with direct numerical or high-fidelity flow simulations of autonomous leader–follower swimmers.

  • Results

    IS η shows 32% higher average swimming-efficiency and 36% lower Cost of Transport than SS η, while three-dimensional followers increase group efficiency by 7.4% versus isolated swimmers.

  • Takeaways & Limitations

    The results support that fish may harvest energy from peer-generated vortices and that deep reinforcement learning can produce navigation algorithms for complex flow-fields.

  • Takeaways & Limitations

    The elastically rigid swimmer is restricted from storing energy supplied by the flow, yielding a conservative estimate of potential Cost-of-Transport savings.

Abstract

from arXiv · show

Fish in schooling formations navigate complex flow-fields replete with mechanical energy in the vortex wakes of their companions. Their schooling behaviour has been associated with evolutionary advantages including collective energy savings. How fish harvest energy from their complex fluid environment and the underlying physical mechanisms governing energy-extraction during collective swimming, is still unknown. Here we show that fish can improve their sustained propulsive efficiency by actively following, and judiciously intercepting, vortices in the wake of other swimmers. This swimming strategy leads to collective energy-savings and is revealed through the first ever combination of deep reinforcement learning with high-fidelity flow simulations. We find that a `smart-swimmer' can adapt its position and body deformation to synchronise with the momentum of the oncoming vortices, improving its average swimming-efficiency at no cost to the leader. The results show that fish may harvest energy deposited in vortices produced by their peers, and support the conjecture that swimming in formation is energetically advantageous. Moreover, this study demonstrates that deep reinforcement learning can produce navigation algorithms for complex flow-fields, with promising implications for energy savings in autonomous robotic swarms.

2. Biological Sciences/Biophysics and Computational Biology

Fish schooling occurs in complex, vortex-rich flow-fields, but the energetic role of hydrodynamic interactions remains quantitatively unresolved. The study uses deep reinforcement learning to discover efficient coordinated swimming strategies for autonomous swimmers.

  • Fish navigate flow-fields containing mechanical energy distributed across multiple scales by vortices from obstacles and other swimming organisms.
  • Hydrodynamic interactions may benefit schooling fish, but quantitative evidence explaining the physical mechanisms behind those energetic benefits remains limited.
  • The study examines two self-propelled swimmers in a leader–follower arrangement to identify mechanisms enabling efficient, sustained coordinated swimming.
  • Two interacting-swimmer cases use deep reinforcement learning for autonomous decisions, while solitary-swimmer cases control for the absence of a leader’s wake.

Reinforcement learning for autonomous swimmers

Deep reinforcement learning is combined with direct numerical simulation to train autonomous followers behind a steadily swimming leader. The learned policies produce distinct positional behaviors, including efficient wake-following without an explicit positional reward.

  • Deep reinforcement learning is combined with direct numerical simulations for two autonomous swimmers in a tandem configuration.The leader represents an adult zebrafish of length L, swims at velocity U, and the simulations use Re = UL/ν ≈5000.
  • IS d maintains its position behind the leader effectively, consistent with a reward penalizing lateral displacement.Its reward is Rd = 1 −|∆y|/L, and the learned follower maintains ∆y ≈0.
  • IS η also settles near the center of the leader’s wake despite receiving reward only for swimming-efficiency, not relative position.
  • Both IS d and IS η maintain a downstream separation of approximately ∆x ≈2.2L.

Intercepting vortices for efficient swimming

The efficient follower adapts its position and deformation to intercept wake vortices and synchronize with their induced flow. These interactions reduce energetic cost and increase thrust, while controlled three-dimensional followers also improve group efficiency.

  • Energetic outcomes: 11% higher average speed, 32% higher average swimming-efficiency, and 36% lower Cost of Transport distinguish IS η from SS η.Over t = 20 to t = 30, the benefit also includes 29% lower deformation effort and 53% higher average thrust-power.
  • Simulation and policy context: Figure 1 depicts leader–follower simulations, diverging vortex-ring wakes, two-dimensional versus three-dimensional vorticity, and visual observations used by the smart follower.
  • Energetic outcomes: IS η and IS d are considerably more energetically efficient than IS d and SS d comparisons indicate hydrodynamic benefits of coordinated swimming.
  • Vortex-interception mechanism: IS η synchronizes head motion with lateral wake velocity and intercepts incoming vortices in a slightly skewed manner.The interception splits each vortex into stronger and weaker fragments, linking wake structure to the follower’s efficient motion.
  • Vortex-interception mechanism: Wake vortices can create suction forces that either assist or oppose body undulations, depending on alignment with deformation velocity.For example, W1L produces negative PDef and favourable PThrust on part of the swimmer’s lower surface.

Energy-saving mechanisms in coordinated swimming

Energy savings arise from synchronizing body deformation with wake vortices, while coordinated three-dimensional followers also benefit through vortex-ring interactions. Deep reinforcement learning discovers these flow-navigation strategies, including responses to erratic leaders.

  • Two-dimensional mechanism: Energy gains occur where wake-vortices interact with body deformation, especially near the midsection rather than through reduced relative velocity alone.The predominant energetics gain occurs in high-relative-velocity regions, including near the high-speed spot generated by vortex L2.
  • Learning and simulation: Without additional training, the smart follower can extract energetic benefit from an erratic leader by deliberately interacting with its unsteady wake.Most reported results use a steadily swimming leader, so this extension broadens the demonstrated flow setting.
  • Three-dimensional extension: 11% higher average swimming-efficiency and 5% lower Cost of Transport are achieved by each controlled three-dimensional follower.The group shows a 7.4% efficiency increase compared with three isolated, noninteracting swimmers.
  • Learning and simulation: Deep reinforcement learning combined with direct numerical simulations produces autonomous navigation policies for self-propelled swimmers in complex flows.The simulations use two- and three-dimensional Navier–Stokes formulations and simplified zebrafish-based swimmer geometries.
  • Three-dimensional extension: In three dimensions, intercepted wake-vortex rings generate lifted-vortex rings that initially reduce efficiency before benefiting the follower downstream.The efficiency modulation follows the lifted ring as it travels along the body.

Supporting Information - Methods

The methods combine incompressible Navier–Stokes simulations with reinforcement learning to control self-propelled swimmers in flow. The simulations compute flow-induced forces and energetics while the learning agent observes swimmer state and selects body-curvature actions.

  • Flow simulations: The simulations are based on incompressible Navier–Stokes equations and use penalty coupling to represent swimmers on the computational grid.The characteristic function identifies each swimmer, while λχ(us −u) couples swimmer and fluid velocities.
  • Flow simulations: Two-dimensional simulations use a vorticity formulation and adaptive wavelet grid, while three-dimensional simulations use pressure projection and a parallelized uniform grid.The three-dimensional domain uses a 2048×1024×256 grid, and pressure-related equations are solved with distributed numerical methods.
  • Energetics: Flow-induced pressure and viscous forces are integrated over swimmer surfaces to obtain thrust, drag, deformation power, swimming efficiency, and Cost of Transport.Surface integrands can also provide spatial distributions of thrust-, drag-, and deformation-power.
  • Energetics: Negative deformation-power values are neglected when computing efficiency and Cost of Transport, producing a conservative estimate of potential savings.The restriction reflects that the elastically rigid swimmer cannot store energy supplied by the flow and prevents overstating benefits.
  • Reinforcement learning: The reinforcement-learning agent observes six state variables and selects discrete curvature-wave control amplitudes to manoeuvre the swimmer.The observed variables include relative position, orientation, recent actions, and tail-beat stage; available amplitudes are 0, ±0.25, and ±0.5.
  • Reinforcement learning: The agent learns a policy maximizing discounted future rewards through a neural-network approximation of the action-value function, with recurrent processing addressing partial observability.Training uses an asynchronous recurrent DQN procedure with target-network updates and multiple independent simulations.

Supporting Information - Supplementary Text, Figures, and Movies

The supporting analyses show that wake interactions alter swimming energetics through synchronized flow-following, but performance remains sensitive to trajectory deviations and swimmer objectives. Supplementary figures and movies compare power distributions, body deformation, flow-induced forces, and two- and three-dimensional configurations.

  • Energetics comparisons: 32% higher swimming-efficiency and 36% lower energy per unit distance distinguish interacting swimmer IS η from solitary swimmer SS η.IS η also requires 29% less body-undulation power and generates 52% higher thrust-power than SS η.
  • Sensitivity and recovery: Small trajectory deviations can desynchronize IS η from incoming wake vortices and noticeably reduce efficiency, although the swimmer autonomously recovers its optimal behavior.The comparison indicates that similar body shape and deformation velocity can nevertheless produce different efficiencies when the surrounding flow differs.
  • Power distribution: The largest reduction in deformation-power occurs along the body midsection, while IS η uniquely develops negative deformation-power near the head.This pattern supports a mechanism distinct from drag reduction through lower relative velocity or thrust increase through channeling.
  • Wake synchronization: IS η synchronizes head movement with lateral wake velocity and intercepts vortices skewedly, splitting them into stronger and weaker fragments.The associated flow organization is linked to the efficient swimming state and is examined through vorticity, velocity, and force fields.
Loading 1802.02674v1…