Source-linked AI summary
Emergent aggregation from collective foraging
Gorka Muñoz-Gil, Andrea López-Incera, Vide Ramsten, Giovanni Volpe, Thomas Müller, Hans J. Briegel
TL;DR
The paper asks whether collective aggregation can arise without direct rewards for grouping, when agents optimize individual foraging while seeing only conspecifics. It trains independent reinforcement-learning foragers in environments with invisible replenishable targets and analyzes their learned strategies. Increasing visual range produces a crossover to scale-agnostic collective search that coincides with spatial aggregation, although the authors stress that this is not a thermodynamic phase transition.
Problem
The paper examines whether collective order can emerge from an individual foraging objective when the objective contains no direct social term.
Method
Independent reinforcement-learning agents optimize target collection while sharing no information, seeing only conspecifics, and never observing targets directly.
Results
As visual range increases, agents switch from environment-tuned individual search to scale-agnostic collective search, accompanied by spatial aggregation.
Takeaways & Limitations
The findings identify indirect, resource-driven reward as a route by which collective phenomena can emerge without rewarding grouping, alignment, or approach.
Takeaways & Limitations
The observed crossover is not a thermodynamic phase transition, according to the finite-size analysis of the aggregation order parameter.
Abstract
from arXiv · showhide
Collective behaviour in living systems is usually modelled as the outcome of a \emph{direct} social drive: agents are rewarded, or hard-wired, to align with or approach their neighbours. Here we show that aggregation can instead emerge from an \emph{indirect} objective. We let reinforcement learning foragers, initially performing a random walk, optimize their dynamics from a purely individual reward for finding replenishable targets, while perceiving only their conspecifics and never the targets themselves. As the visual range grows, the agents undergo a sharp crossover from an environment-tuned individual search to a scale-agnostic collective one, and this crossover coincides with the onset of spatial aggregation. Thus a collective phase arises as a by-product of optimal foraging, without any direct reward for grouping. A minimal analytical first-passage model reproduces the transition as a crossover between the two search strategies. Our results identify indirect, resource-driven reward as a generic route to emergent collective phenomena.
INTRODUCTION
The paper asks whether collective order can emerge from an individual foraging objective rather than a direct social drive. It uses reinforcement learning to test whether agents rewarded only for finding unseen replenishable targets develop collective behaviour through conspecific cues.
- Collective behaviour is typically modelled through direct social interactions that align agents with or move them toward neighbours.
- The central question is whether collective order can emerge when the objective contains no reference to grouping or social alignment.
- Finding optimal environment-specific search strategies is difficult, and interactions make each agent’s optimum depend on the strategies of others.
- The study trains independent reinforcement learners that share no information, interact through vision, and receive reward solely for collecting replenishable targets they cannot see.
- As visual range increases, agents switch from environment-tuned individual search to collective search, accompanied by a transition from disorder to spatial aggregation.
METHODS
Agents forage in a periodic two-dimensional environment, receive individual collection rewards, and observe conspecifics through a coarse-grained visual channel. Reinforcement learning adapts their movement policies while targets remain invisible.
- Agents move in a periodic square box, take unit-length steps, and collect randomly placed targets that are depleted for a recovery time τ.
- Agents see conspecifics but never targets through a cone of half-width π/8 and range rv, represented by three visual states.
- Each agent selects continue or turn actions using a policy conditioned on movement counter c and visual state v, trained by model-free reinforcement learning with reward R = 1 per target collected.
- Mean reward per step η remains at blind single-agent efficiency below a parameter-dependent visual range, then jumps to a higher plateau.
RESULTS
As visual range increases, independently learned foragers switch sharply from environment-tuned individual search to a scale-agnostic collective strategy. This transition coincides with spatial aggregation and is organized by target spacing, depletion rules, and search-time trade-offs.
- Efficiency transition: The learned efficiency η remains at the blind single-agent value below a visual-range threshold, then jumps to a higher plateau.Larger τR shifts the jump to smaller rv and lowers the plateau because longer tags are more frequent but noisier cues.
- Two dynamical strategies: At large rv in the cooperative case, agents reverse their no-agent and non-rewarded policies, turning frequently instead of exploiting environmental length scales.The reversal begins at rv = 6, matching the efficiency jump; competitive agents invert under other environmental parameters.
- Two dynamical strategies: The collective strategy turns until a rewarded conspecific appears, then moves ballistically toward it, whereas the alternative exploits environment-tuned scales.Agents converge to one strategy or the other in both cooperative and competitive settings, with the competitive crossover occurring at larger rv.
- Emergence of aggregation: The Clark–Evans index R maps the strategy transition onto spatial organization: disordered states have R ≈1, while aggregation has R < 1.Before clustering, increasing rv can produce over-dispersion because frequent non-rewarded-agent sightings induce high turning probabilities.
- Emergence of aggregation: The transition is sharp, with agents collapsing onto a common R on either side and exhibiting only individual or collective macroscopic states.A finite-size analysis identifies a well-defined onset r*_v, but the transition width does not close as agent number grows, so it is not a thermodynamic phase transition.
- Role of depletion and target density: Higher target density pushes the crossover to larger rv because blind search remains efficient, while the visual-range threshold is nearly insensitive to τR.The threshold is set mainly by whether following a rewarded agent beats exploiting environmental scales, linked here to mean target spacing dt = 5.
- Physical modelling of the crossover: The minimal first-passage model qualitatively reproduces the aggregation boundary by comparing single and collective search strategies across target densities.Because simulations do not strictly satisfy the model’s dilute assumption, the model identifies the selection mechanism and scale rather than quantitatively predicting the boundary.
DISCUSSION
Independent reinforcement learners can develop collective foraging and spatial aggregation while optimizing only individual search efficiency. The collective phase emerges from indirect information about neighbours’ foraging success rather than any direct grouping objective.
- Independent learners spontaneously develop collective foraging and spatial aggregation while optimizing only their own search efficiency.
- No agent is rewarded for grouping, aligning, or approaching others; collective behaviour instead emerges from an indirect cue: neighbours’ foraging success.
- The transition separates an environment-tuned individual search from a scale-agnostic collective search as visual range increases.
- The findings contrast with models where collective motion is imposed through hard-wired alignment or rewards for matching neighbours.
Appendix A: Training
The training procedure uses independent Projective Simulation reinforcement-learning agents whose policies are updated from target-collection rewards while coupling occurs through visual observations.
- Each agent uses Projective Simulation, an off-policy reinforcement-learning algorithm storing its policy as weighted clip-transition networks.
- Agents observe state s = [c, v], sample actions from π(a|s), and reinforce transition sequences after collecting a target.
- The agents hold independent networks and exchange no weights, so their coupling is mediated by the shared environment through visual state v.
- All reported policies were trained for 10^4 episodes of 5000 steps, with target fields resampled at each episode start.
Appendix B: Theoretical model for strategy transition
The theoretical model explains the strategy transition by comparing simplified single-agent and collective search strategies through their mean first-passage times. It assumes dilute targets and approximates collective search as random motion followed by ballistic motion.
- The single-agent strategy is represented by an optimized bi-exponential random walk, supported by prior work and learned-policy convergence.
- The collective strategy first uses unit-step random walking, then switches to ballistic motion when the walker reaches distance x < λ from a target.
- The analysis focuses on the dilute regime, ρ = N_t/L^2 → 0, where sparse rewards make learning more difficult.
- The model compares single and collective strategies by extracting their mean first-passage times as functions of depletion time τ and visual range r_v.
- The setup models replenishable targets in a two-dimensional periodic box, with target depletion lasting time τ after acquisition.
2. Single agent search
The blind single-agent search is environment-dependent and is modeled using an optimized bi-exponential walk. In the simulated regime, long flights exceed the mean free path, making search effectively ballistic before capture.
- The blind agent’s strategy depends fully on environmental properties, especially the depletion time τ.
- The optimized bi-exponential walk combines short intensive steps with a long relocating scale d_1 ≫ d_2.
- When d_1 ≳ T_∞, the searcher typically encounters a target during one flight and does not turn before capture, yielding effective walk dimension d_w = 1.
- For d_s = 2 > d_w = 1, exploration is non-compact and the MFPT is evaluated for a walker starting at distance x = τ from the nearest target.
- Figure 6 compares simulated single-agent MFPTs across depletion times and box sizes with Eq. (B4) fits and the mean path length from Eq. (B2).
- Numerical fits support Eq. (B4), with A ∼ 1 and B ≈ 0.64 in most cases, while the MFPT saturates proportionally to T_∞.
a. Deterministic absorption
The deterministic-absorption model treats search as diffusion within a target cell, with absorption at the detection radius and reflection at the cell boundary. In the dilute regime, the resulting mean first-passage time is logarithmic in the start-to-detection distance ratio.
- Absorption assumption: The deterministic model assumes that a ballistic flight initiated at x < λ encounters the target with probability 1.Under this assumption, the ballistic contribution is a fixed time λ, leaving the mean first-passage time related to the unbiased random-walk phase.
- Cell construction: The trapping problem replaces each target’s average area 1/ρ with a circular Wigner–Seitz cell of radius b.The searcher diffuses inside the cell, is absorbed at x = λ, and reflected at x = b.
- Boundary-value problem: The mean first-passage time T(x) obeys D∇2T = −1 with absorbing and reflecting radial boundaries.The absorbing boundary is the detection circle x = λ, while the cell edge x = b is reflecting.
- Diffusive limit: The microscopic walk gives a diffusion description by comparing ⟨R2⟩ = Nd2 and t = Nd with the two-dimensional diffusion law ⟨R2⟩ = 4Dt.This identifies the effective diffusion coefficient used in the trapping calculation.
- Dilute-regime result: In the dilute regime λ, τ ≪ b, the logarithm dominates and Tdet ≃ A ln(τ/λ).The approximation applies when the detection and starting scales are both much smaller than the cell size.
b. Probabilistic absorption
The probabilistic-absorption model relaxes perfect capture at the detection radius by representing missed encounters and delayed social signals through a partially absorbing boundary. Its reactive length rescales the effective detection radius and shifts the strategy crossover.
- Sources of imperfect capture: Agents can miss a target even within the detection radius, and rewarded-agent encounters represent a signal diffusing from the target.The visual cone has width θ = π/4 in the numerical results, while tagged agents provide the indirect target signal.
- Model simplification: The exact mean first-passage time for these effects is left for future work, so the analysis uses a partially absorbing boundary at x = λ.Only a fraction of agents reaching the boundary are assumed to capture the target.
- Reactive boundary: The reactive length ℓ interpolates between perfect absorption at ℓ → 0 and perfect reflection at ℓ → ∞.It bundles effects such as tag times and changes in agent density.
- Collective MFPT: The collective mean first-passage time is obtained by solving the diffusion equation with the Robin condition and retaining leading order in the dilute limit b ≫ λ.This produces the collective-strategy MFPT used for comparison with the single strategy.
- Effective-radius mapping: The reactive length is equivalent to replacing λ with an effective radius λeff = λ e−ℓ/λ.This maps the directional problem onto an isotropic perfectly absorbing problem with a wider trapping radius.
4. Dynamics phase diagram
The phase diagram compares single and collective search by equating their mean first-passage times. The single strategy dominates below the crossover, whereas the collective strategy dominates above it, with imperfect capture shifting the boundary.
- Crossover criterion: The crossover is defined by α = Tc(λ, τ)/Ts(λ) = 1, where the collective and single-search MFPTs are equal.The corresponding λ determines the boundary between the two search strategies.
- Strategy regions: For λ < λ∗, the single strategy has the lower MFPT, while for λ > λ∗ the collective strategy has the lower MFPT.The boundary is plotted for different reactive lengths ℓ and for the deterministic single-search expression.
- Effect of imperfect capture: Increasing ℓ moves the strategy boundary toward larger λ because lower capture probability makes single search dominant over a larger visual-range regime.The reactive length therefore shifts the predicted crossover in the phase diagram.
Appendix C: Sharp crossover and dynamical coexistence
Finite-size analysis shows a sharp but non-thermodynamic aggregation onset. Near the onset, dispersed and aggregated states coexist, while increasing system size makes aggregation more reliable without narrowing the crossover to zero width.
- Order parameter: The dense-phase order parameter m is the fraction of agents with at least Nℓ = 5 neighbours within radius rℓ = 2.5.This threshold is far above the Poisson expectation ρaπrℓ2 ≈ 0.4, so random populations have vanishing m.
- Aggregation onset: At fixed agent and target densities, both the mean ⟨m⟩ and aggregated-run fraction rise sharply at a well-defined visual-range onset r∗v.The rise steepens with system size as L ∝ √Na.
- Dynamical coexistence: Near r∗v, P(m) is bimodal, with one peak at m = 0 and another at finite m, indicating coexistence of dispersed and aggregated states.The coexistence is reproduced across four independently trained policies, supporting its origin in collective dynamics rather than learning variability.
- Finite-size reliability: Larger systems reach the aggregated state more reliably, sharpening the onset without eliminating the coexistence near r∗v.The effect is observed across independently trained policies and varied system sizes.
- Crossover classification: The transition-region width extrapolates to c ≈ 0.09 as 1/L2 → 0, identifying a sharp crossover rather than a thermodynamic phase transition.The width is defined as the visual-range interval over which the aggregated-run fraction rises from 0.1 to 0.9.