Source-linked AI summary

Stochastic Models of Evolution in Genetics, Ecology and Linguistics

R. A. Blythe, A. J. McKane

arXiv:cond-mat/0703478v1cond-mat.stat-mechq-bio.PE

TL;DR

The paper surveys stochastic evolutionary models for genetics, ecology, and linguistics, emphasizing neutral models for nonspecialists. It develops ideal and non-ideal population models and reviews forward-time diffusion and backward-time coalescent approaches, finding that neutral theory can fit ecological abundance data while retaining important scope limits.

  • Problem

    The paper addresses how stochastic, especially neutral, models can describe evolutionary processes across genetics, ecology, and linguistics.

  • Method

    The paper synthesizes individual-based neutral models, their non-ideal extensions, diffusion approximations, coalescent analyses, and related ecological and linguistic models.

  • Results

    Neutral theory’s calculated distribution of species abundances gives a remarkably good fit to data, despite the model’s extreme simplicity and controversy.

  • Takeaways & Limitations

    Neutral stochastic models provide a shared framework whose predictions can be relevant to genetic, ecological, and linguistic systems.

  • Takeaways & Limitations

    Ideal-population predictions may require replacing actual population size with effective size, and multi-allele Fokker-Planck equations may be difficult to solve.

Abstract

from arXiv · show

We give a overview of stochastic models of evolution that have found applications in genetics, ecology and linguistics for an audience of nonspecialists, especially statistical physicists. In particular, we focus mostly on neutral models in which no intrinsic advantage is ascribed to a particular type of the variable unit, for example a gene, appearing in the theory. In many cases these models are exactly solvable and furthermore go some way to describing observed features of genetic, ecological and linguistic systems.

1. Introduction

The article introduces neutral stochastic models of evolution, centered on population genetics while connecting analogous processes in ecology and linguistics. It outlines ideal genetic drift, its non-ideal extensions, and complementary diffusion and coalescent analyses.

  • Statistical mechanics and evolutionary biology both use probabilistic models to describe typical system behavior and stochastic dynamics.
  • The Wright-Fisher model keeps population size constant, with each gene copy sampled randomly with replacement from the preceding generation.The model was introduced independently by Ronald Fisher and Sewall Wright in the 1930s.
  • Equal mean offspring numbers make the model neutral, while random sampling produces allele-frequency fluctuations known as genetic drift.
  • Without mutation, genetic drift eventually eliminates alleles until one allele fixes in the population.
  • Ideal-population predictions can sometimes describe non-ideal populations when actual population size is replaced by effective size.Non-random mating can arise from sex, geography, age, and other population structure.
  • Neutral models are applied beyond genetics: ecological species replacement resembles random genetic reproduction, while linguistic variants compete for use without prestige-driven differences.The ecological abundance distribution gives a remarkably good fit to data, although the model’s simplicity remains controversial.
  • The article proceeds from neutral genetic-drift definitions to diffusion and coalescent approaches, then reviews non-ideal and non-neutral evolution.

2. Genetic drift: the Wright-Fisher and Moran models

The Wright-Fisher and Moran models represent neutral genetic drift through stochastic reproduction, mutation, and migration processes. Their analysis yields fixation results and mappings to ecological and linguistic evolution models.

  • 2.1. The Wright-Fisher model: The Wright-Fisher model keeps population size constant by repeatedly sampling genes with replacement from one generation to construct the next.The process is equivalent to drawing allele-colored balls from an urn and copying each selected gene into the new generation.
  • 2.1. The Wright-Fisher model: At long times, the two-allele Wright-Fisher process reaches either loss or fixation of allele A, with fixation probability n0/N.The mean number of A alleles is conserved, and the eventual states are n = N or n = 0.
  • 2.2. The Moran Model: The Moran model is analyzed through transition-matrix eigenvalues and eigenvectors, including the slow relaxation mode with λ2 = 1 − (2/N^2).The two unit eigenvalues correspond to loss and fixation, while the third-largest eigenvalue governs long-time behavior of non-fixed states.
  • 2.3. Mutations in the Wright-Fisher and Moran Models: Mutation prevents either allele from becoming permanently fixed, while migration between demes drives mean allele frequencies toward a common stationary value.The deterministic mutation equation is at+1 = (1 − u)at + vbt, and island models represent partially isolated populations connected by migration.
  • 2.5. Mapping between population genetic, ecological and linguistic models: The same stochastic formalism maps genetic alleles, ecological individuals or species, and linguistic variants, while large populations admit Fokker-Planck approximations.The linguistic analogy includes innovation as mutation-like variation and migration as incorporation of tokens between speakers’ grammars.

3. Master and Fokker-Planck equations

The paper develops master-equation and diffusion descriptions of neutral genetic models, showing how large-population limits yield Fokker–Planck equations. These equations extend to mutation, migration, multiple islands, and multiple alleles, with exact solutions available under special restrictions.

  • Diffusion approximation: The diffusion approximation replaces discrete allele frequencies with continuous variables and rescales time for large populations, providing a mesoscopic description.Mutation and migration can be added as terms in the resulting equation.
  • Master equations: Moran-type models can be reformulated as master equations by replacing per-generation transition probabilities with continuous-time transition rates.The continuous-time interpretation uses sufficiently small intervals in which at most one sampling event occurs.
  • Fokker–Planck equations: Large-N limits of genetic drift, mutation, and migration produce Fokker–Planck equations that extend naturally to L islands and M alleles.The construction uses scaled time τ = 2t/N^2 and lets N approach infinity.
  • Fokker–Planck equations: The Wright–Fisher and Moran models yield the same mesoscopic diffusion equation, although their time scales differ.A Wright–Fisher generation is N/2 times as long as a Moran-model generation in this description.
  • M alleles: For mutation rates depending only on the destination allele, the multiallelic equation has a known stationary solution and, remarkably, solvable full-time dynamics.The dynamics can be obtained through orthogonal polynomials or a change of variables that makes the equation separable; the eigenfunctions are Jacobi polynomials.
  • M alleles: For general mutation rates depending on both initial and final allele states, the stationary solution is unknown, and detailed balance is only conjectured not to hold.The simplified destination-dependent case leads to a nonlinear partial differential equation in M − 1 variables that is not expected to be readily solved.

4. Backward-time formulation: coalescent theory

The coalescent reformulates neutral Wright–Fisher ancestry backward in time: lineages merge when they share parents, enabling efficient analysis of present-day samples and forward-time outcomes. In large populations, its dynamics extend across several neutral models under appropriate time rescaling.

  • 4.1. Reconstructing the ancestry: Backward-time analysis reconstructs present-day ancestry by tracing individuals to shared parents until their lineages coalesce into one ancestor.The number of ancestors decreases as one looks backward, and coalescence is represented as merging branches in an ancestral tree.
  • 4.1. Reconstructing the ancestry: The coalescent rate is proportional to n(n−1), causing expected waiting times between events to lengthen as the number of remaining lineages decreases backward in time.Most of the time to a common ancestor is spent when only a few ancestors remain; T(2 →1) exceeds half of T(m →1).
  • 4.1. Reconstructing the ancestry: The probability that any pair of ancestors shares a parent is 1/N, so coalescence events occur among possible ancestral pairs while the lineage count remains much smaller than N.The coalescence probability is based on random parent assignment in the Wright–Fisher model.
  • 4.1. Reconstructing the ancestry: The ancestry can also be represented as reaction-diffusion on a complete graph, where particles are lineages and coincident particles immediately coalesce.Each particle hops to a neighbour or stays in place, and the complete graph gives the corresponding ideal Wright–Fisher dynamics.
  • 4.1. Reconstructing the ancestry: Backward-time methods condition naturally on present-day genetic samples, simulate genealogical trees efficiently, and allow neutral mutations to be superimposed without altering population dynamics.A tree for n sampled individuals requires generating n−1 waiting times, and the approach generalizes more readily to non-ideal mating populations for selectively neutral alleles.
  • 4.2. Relationship between forward- and backward time dynamics: Backward-time quantities can compute forward-time probabilities from fixed initial conditions, including the probability that a mutant allele has become fixed by time t.The method uses lines of descent from an initial mutant set to relate ancestral and forward distributions; some relevant distributions are known analytically, while applications remain useful when they are not.
  • 4.3. Scaling with population size in the coalescent formulation: In the infinite-population limit, Wright–Fisher and Moran ancestries have the same dynamics under τ=t/N for Wright–Fisher and τ=2t/N^2 for Moran.The shared scaling is also the one previously used to obtain the same Fokker–Planck equation for both models.
  • 4.3. Scaling with population size in the coalescent formulation: Many neutral population models share these ancestral dynamics as N→∞ when time is appropriately rescaled and ternary coalescences vanish relative to binary coalescences.The stated conditions include exchange-invariant offspring-number distributions and a timescale on which binary coalescences dominate.

5. Non-ideal populations

Non-ideal populations are modeled by tracking nonuniform displacement and migration between subpopulations, while coalescent analysis quantifies structure and effective size. The results show how migration, finite population size, selfing, and sex-ratio imbalance affect FST and Ne.

  • Migration: Migration between demes is represented by probabilities µij that a sampled individual in deme i descends from a parent in deme j.These parameters support a backward-time formulation of non-ideal populations.
  • Migration: Under conservative migration, the common ancestor is in deme i with probability proportional to that deme’s size.This is a steady-state result for migration that balances expected entry and exit in every deme.
  • F-statistics: In infinite subpopulations, 0 ≤ FST ≤ 1, with the bounds corresponding to homogeneous and heterogeneous spatial allele distributions.Finite populations can produce slightly negative FST through sampling without replacement, so alternative definitions need not agree.
  • F-statistics: When migration remains finite as N →∞, FST →0, indicating vanishing spatial structure in the strong-migration limit.If Nµ remains finite instead, a distinct limiting regime results; finite-size corrections to the infinite-island result are of order 1/L.
  • Coalescent analysis: The mean coalescence times Tij satisfy linear equations that combine lineage migration with immediate coalescence when both lineages enter the same deme.For arbitrary migration rates, these equations can be cumbersome; slow-migration scaling removes simultaneous lineage hops and simplifies analysis.
  • Effective population size: Selfing reduces the effective haploid population size to Ne = L(2−S), while unequal male and female numbers also make Ne smaller than the actual gene count.For hermaphrodites, the selfing model requires a separation of timescales because the two parental choices are not independent.
  • Effective population size: Under strong migration, Ne is at most the total population size L N̄ and equals it only when the stationary distribution is proportional to deme size.Conservative migration is the condition for this size-proportional stationary distribution when it is unique.
  • Effective population size: Without conservative migration, no simple relation links Ne and inbreeding coefficients, and a single parameter may not capture all evolutionary dynamics.Weak migration may require a structured coalescent that tracks lineage numbers within demes at finite times.

6. Selection

The section extends neutral evolutionary models to selection, showing how selection interacts with drift, migration, population subdivision, and spatial structure. It derives deterministic and stochastic descriptions of allele-frequency change, including selective sweeps, travelling waves, and clines.

  • Selection in finite populations: Selection assigns different reproductive weights to alleles, while equal weights recover the neutral Wright-Fisher transition rule.The mean fitness normalizes parent-selection probabilities in the finite-population model.
  • Deterministic selection: When selection is weak and N is sufficiently large, drift can typically be neglected and mutant numbers follow logistic growth.The stated condition is s ≪1 with N ≫1/s at fixed s.
  • Selective sweeps: Once a beneficial allele exceeds a frequency threshold, logistic growth approximates a selective sweep, whereas early dynamics remain strongly affected by drift.The mutant is initially rare, so stochastic effects are expected to matter near the boundaries.
  • Genealogies under selection: Selection complicates backward-time genealogical calculations because ancestral lineages must track the numbers of alleles of each type.The added state information contrasts with the simpler neutral backward-time formulation.
  • Selection and subdivision: With population subdivision, strong migration produces effective-population-size behaviour, while slow migration separates selective sweeps from inter-island invasion dynamics.In the slow-migration limit, subpopulations are treated as fixed and switch states on the migration timescale.
  • Spatial selection: The Fisher-KPP equation supports travelling mutation fronts with velocities v = Dλ + s/λ ≥ 2√(Ds), with the smallest allowed velocity selected for sharp interfaces.Adding drift causes strong leading-edge fluctuations, a small velocity shift, and a more diffuse front.
  • Spatial selection: A spatial fitness change, with advantage +s for u > 0 and disadvantage −s for u < 0, produces a steady allele-frequency cline through migration-selection balance.The fitness discontinuity induces a discontinuity in the gradient at u = 0.

7. Conclusion

The conclusion presents the review as a bridge between stochastic evolutionary models and statistical physics across genetics, ecology, and linguistics. It emphasizes individual-based derivations, analytical master-equation methods, effective parameters, and unresolved limits of idealized models.

  • Scope and purpose: The review explains stochastic evolutionary models across population genetics, ecology, and linguistics for readers trained in statistical physics.It combines background motivation with mathematical formalism to address disciplinary barriers.
  • Neutral models: Neutral processes recur across all three fields and can serve as null models, although natural systems depart from idealized assumptions.The review treats neutrality as broadly useful without denying the role of selection.
  • Analytical methods: Master-equation formulations can yield stationary distributions analytically and are often more efficient than simulation-based treatments.The review contrasts this efficiency with formulations that required simulations to generate stationary distributions.
  • Analytical methods: Discrete-time transition-probability models and continuous-time rate-based master equations are related, but their correspondence is not always straightforward.Mutation implementations illustrate that the spontaneous-Poisson formulation imposes restrictions absent from the combined-event approach.
  • Microscopic-to-mesoscopic modelling: Starting from individual-based models and taking the large-N limit connects microscopic mechanisms to mesoscopic descriptions and supports simulations on complex interaction networks.The coalescent approach is especially useful for large-network questions such as fixation-time trends.
  • Limits and future work: Effective population size and inbreeding coefficients summarize aspects of non-ideal populations, but their adequacy and systematic use remain open questions.The conclusion calls for controlled approximation schemes for deviations from ideal behaviour.
  • Future directions: Exact solutions for the reaction-diffusion process A + A →A exist in some geometries, while heterogeneous graphs have so far required approximate methods.Extending analysis to heterogeneous graphs could reveal more detailed evolutionary dynamics.

Appendix A. The large N expansion of the master equation for M alleles

The appendix derives the large-N Fokker-Planck description for populations with multiple alleles and islands from master-equation transition rates. It separates mutation and migration contributions under the stated scaling and assembles them into the general subdivided-population equation.

  • Expansion procedure: The appendix begins the large-N expansion from master equations for populations with M > 2 alleles.It expands transition-rate terms in powers of 1/N.
  • Expansion procedure: Terms from transition rates are expanded, combined, and retained through the order contributing to the Fokker-Planck equation.Terms of order 1/N^3 and higher are neglected in the displayed derivations.
  • Mutation terms: For the first mutation scheme, algebraic combinations of rate terms produce the mutation contribution after introducing scaled mutation rates.The derivation treats allele pairs among A_1 through A_{M−1} and then handles A_M analogously.
  • Mutation terms: The second mutation scheme yields the same Fokker-Planck equation when its transition rates match the first scheme, subject to an additional restriction on the rates.The equivalence follows by setting u_αβ = u_α for β ≠ α.
  • Multiple islands: For L islands and M alleles, scaled mutation and migration rates make their products order 1/N^3, so the two contributions can be derived separately.The separate pieces are then combined in the general multi-island derivation.
  • Multiple islands: Combining the expansion terms gives equation (90), a general Fokker-Planck equation for M allele types distributed across L islands.The result incorporates mutation, migration, and island-specific population factors.
Loading cond-mat/0703478v1…