Source-linked AI summary

Win-stay-lose-learn promotes cooperation in the spatial prisoner's dilemma game

Yongkui Liu, Xiaojie Chen, Lin Zhang, Long Wang, Matjaz Perc

arXiv:1205.0802v1q-bio.PEcond-mat.stat-mechphysics.soc-ph

TL;DR

The paper asks whether unconditional strategy updating overlooks the role of satisfaction in cooperation. It introduces aspiration-based win-stay-lose-learn updating for the spatial prisoner’s dilemma and finds that it promotes cooperation robustly across initial conditions, supported by simulations and pair approximation.

  • Problem

    Standard evolutionary-game simulations commonly let players update strategies frequently, despite strategy changes being more likely when players feel unsuccessful or dissatisfied.

  • Method

    The model makes strategy updating conditional on aspiration-based satisfaction: satisfied players keep their strategies, whereas dissatisfied players attempt to imitate a nearest neighbor.

  • Results

    The win-stay-lose-learn rule promotes cooperation robustly, with highly cooperative states attainable from adverse initial conditions and pair approximation supporting the main simulation conclusions.

  • Takeaways & Limitations

    Aspiration-dependent strategy updating provides a way to sustain cooperation in the spatial prisoner’s dilemma, including when the initial cooperator fraction is very small.

Abstract

from arXiv · show

Holding on to one's strategy is natural and common if the later warrants success and satisfaction. This goes against widespread simulation practices of evolutionary games, where players frequently consider changing their strategy even though their payoffs may be marginally different than those of the other players. Inspired by this observation, we introduce an aspiration-based win-stay-lose-learn strategy updating rule into the spatial prisoner's dilemma game. The rule is simple and intuitive, foreseeing strategy changes only by dissatisfied players, who then attempt to adopt the strategy of one of their nearest neighbors, while the strategies of satisfied players are not subject to change. We find that the proposed win-stay-lose-learn rule promotes the evolution of cooperation, and it does so very robustly and independently of the initial conditions. In fact, we show that even a minute initial fraction of cooperators may be sufficient to eventually secure a highly cooperative final state. In addition to extensive simulation results that support our conclusions, we also present results obtained by means of the pair approximation of the studied game. Our findings continue the success story of related win-stay strategy updating rules, and by doing so reveal new ways of resolving the prisoner's dilemma.

Introduction

The paper addresses cooperation in the spatial prisoner’s dilemma by replacing frequent, unconditional strategy updating with aspiration-dependent changes made only by dissatisfied players. It introduces win-stay-lose-learn updating, in which satisfied players retain their strategies while dissatisfied players imitate neighbors.

  • Motivation: The prisoner’s dilemma pits individual incentives for defection against the socially optimal outcome of mutual cooperation.The payoff ordering T > R > P > S creates this tension.
  • Motivation: Spatial games study how cooperation can prevail over defection, with network topology identified as an important determinant of cooperative success.Scale-free topology has been especially beneficial, although some mechanisms can weaken that advantage.
  • Motivation: Conventional spatial-game models commonly let players consider updating their strategies in nearly every round, despite people being less prone to change strategies.The paper motivates aspirations as determinants of satisfaction and perceived personal success.
  • Proposed rule: Win-stay-lose-learn lets satisfied players keep their strategies, while dissatisfied players attempt to imitate one nearest neighbor.Whether a player is satisfied depends on the aspiration level A > 0, treated as a free parameter.

Results

Simulations show that aspiration-dependent updating can sustain or increase cooperation across temptation levels, while pair approximation reproduces the main qualitative conclusions. Robustness tests further examine adverse initial conditions and cooperative resistance to defectors.

  • Aspiration and temptation: ρC = 0.5 at A = 0 and ρC = 0.47 at A = 0.2, independently of b; at A = 0.6, ρC →1.0 for b ∈[0, 1.2].At A = 0.4, ρC = 0.7 for b ∈[0, 1.6].
  • Aspiration and temptation: Pair approximation qualitatively predicts cooperation levels, exactly matching simulations at A = 0.0 while differing more as A exceeds 0.5.For A > 0.5, dissatisfaction becomes increasingly common and updating approaches continuous updating.
  • Aspiration and temptation: The highest cooperation levels occur for 0.5 ≤ A < 0.75, while discontinuous transitions appear at A = 0.0, 0.25, 0.5, and 0.75.The phase-transition threshold can be calculated from local payoff conditions; for A = 0.4 and n3 = 1, b = 1.6.
  • Satisfaction: For A = 0.6 and b ∈[1.0, 1.2], the system combines high cooperation with a high satisfaction rate.At A = 0.2, more than 90% of individuals are satisfied despite ρC = 0.47, because satisfied defectors exploit cooperators.
  • Initial conditions: A single cooperator cannot survive when A > 0, whereas two neighboring cooperators require A ≤0.25 to survive.Under adverse initial conditions, dissatisfaction causes cooperators to imitate neighboring defectors.
  • Resistance to invasion: With a square block of four defectors, cooperators survive when A ≤0.75, invade when b/2 < A ≤0.75 and b < 1.5, and coexist with defectors when b > 1.5.When A > 0.75, cooperators are doomed to extinction in this configuration.

Discussion

The paper finds that satisfaction-dependent strategy updating facilitates cooperation, especially at intermediate aspiration levels, and distinguishes this mechanism from related models through adaptive evolutionary time scales.

  • Discussion: If satisfied players retain their strategies, cooperation is remarkably facilitated, with A = 0.6 supporting virtually complete cooperation even when temptation to defect exceeds 1.Dissatisfied players instead attempt to adopt a nearest neighbor’s strategy.
  • Discussion: Unlike unconditional updating, the rule changes strategies only when payoff falls below aspiration, making satisfaction the determinant of strategy updating.Satisfied players do not immediately need to change strategies, whereas dissatisfied players imitate a nearest neighbor.
  • Discussion: Related work includes stochastic win-stay-lose-shift, N-person aspiration models, inertia-based models, and early win-stay-lose-shift rules.These models differ in how dissatisfied players switch or how strategy changes are inhibited.
  • Discussion: Aspiration creates adaptive evolutionary time scales because players’ strategy-selection rates differ with satisfaction and can change after strategy updates.For small A, strategy transfers are rare; at intermediate aspiration levels, satisfied and dissatisfied players coexist.

Methods

The study models a spatial prisoner’s dilemma on a periodic square lattice, where aspiration-dependent imitation governs strategy adoption and pair approximation supplements simulation analysis.

  • Methods: The game uses a square lattice with periodic boundaries, payoffs T = b, R = 1, and P = S = 0, with 1 < b < 2.Each player has four neighbors and aspiration A_i = k_iA; on this lattice k_i = 4.
  • Methods: Player i keeps its strategy when P_i ≥ A_i; when P_i < A_i, it adopts randomly selected neighbor j’s strategy with probability f(P_i, P_j).The probability function includes noise amplitude κ and favors better-performing neighbors when dissatisfaction triggers updating.
  • Methods: The noise parameter κ represents imperfect information and decision errors, and the simulations use κ = 0.1.The paper states that the outcome is generally robust to variations in κ.
  • Methods: Pair approximation uses rate equations for cooperator-cooperator and cooperator-defector edges to estimate the cooperator density ρ_C.The approximation imposes p_c,d = p_d,c and p_c,c + p_c,d + p_d,c + p_d,d = 1.
  • Methods: The pair-approximation equations track edge configurations indexed by neighboring strategy triplets and transition probabilities between local states.The equations include cooperator counts n_c(x, y, z), where x, y, z are cooperators or defectors.
Loading 1205.0802v1…