Source-linked AI summary
Is your phylogeny informative? Measuring the power of comparative methods
Carl Boettiger, Graham Coop, Peter Ralph
TL;DR
Phylogenetic comparative methods can be unreliable when models are inappropriate or phylogenies lack enough information to distinguish them, yet power and uncertainty are seldom quantified. The paper introduces Monte Carlo simulations for model comparison and parameter confidence intervals, finding that informativeness and model-discrimination power depend on phylogeny structure and that information criteria can perform poorly. It concludes that power analysis can identify when a tree is too small or unstructured for the proposed comparison.
Problem
Comparative methods need adequate data to fit increasingly complex evolutionary models and distinguish among them, but uncertainty and statistical power are often not quantified.
Method
The paper uses parametric bootstrapping and Monte Carlo simulations under specified models to estimate parameter uncertainty, power, false-positive rates, and model-comparison distributions.
Results
Information-criterion model choice can have remarkably high error rates, while the Monte Carlo approach quantifies and substantially reduces errors and produces meaningful parameter confidence intervals.
Takeaways & Limitations
Power analysis indicates when a phylogeny is too small or unstructured to resolve differences among proposed models and when added model complexity is unsupported.
Takeaways & Limitations
Power depends strongly on parameter values, so a simulation calibrated to different values may leave failures to support complex models attributable to insufficient power.
Abstract
from arXiv · showhide
Phylogenetic comparative methods may fail to produce meaningful results when either the underlying model is inappropriate or the data contain insufficient information to inform the inference. The ability to measure the statistical power of these methods has become crucial to ensure that data quantity keeps pace with growing model complexity. Through simulations, we show that commonly applied model choice methods based on information criteria can have remarkably high error rates; this can be a problem because methods to estimate the uncertainty or power are not widely known or applied. Furthermore, the power of comparative methods can depend significantly on the structure of the data. We describe a Monte Carlo based method which addresses both of these challenges, and show how this approach both quantifies and substantially reduces errors relative to information criteria. The method also produces meaningful confidence intervals for model parameters. We illustrate how the power to distinguish different models, such as varying levels of selection, varies both with number of taxa and structure of the phylogeny. We provide an open-source implementation in the pmc ("Phylogenetic Monte Carlo") package for the R programming language. We hope such power analysis becomes a routine part of model comparison in comparative methods.
1. Introduction
Phylogenetic comparative methods need both biologically appropriate models and enough informative data to distinguish among them. The paper motivates simulation-based uncertainty and power assessment as a response to unreliable inference from small or structurally limited phylogenies.
- Comparative-method concerns include whether evolutionary models reflect biological reality and whether data adequately fit and distinguish those models.
- Insufficient power can produce false biological conclusions because uncertainty, model fit, and power are rarely quantified in current comparative methods.
- Parametric bootstrapping uses simulations under specified comparative models to estimate parameter uncertainty, power, and false-positive rates.
- Continuous-trait models commonly assume a known ultrametric phylogeny, with trait evolution specified by stochastic processes along independently evolving branches.
- Brownian motion models trait evolution as branching random walks with initial state X0 and rate parameter σ, while Ornstein-Uhlenbeck evolution adds attraction toward an optimum θ controlled by α.
- Model extensions allow evolutionary rates or optima to vary across time, branches, or clades, including painted OU models and Pagel’s λ transformation of shared ancestry.
2. Methods
The paper estimates uncertainty and model-discrimination power by simulating trait datasets on a phylogeny and repeating parameter estimation and likelihood comparisons. This approach exposes how informativeness depends on tree size, topology, and the parameter being estimated.
- 2.1. Uncertainty in parameter estimates: On a small Geospiza phylogeny, 1000 simulations with true λ = 0.6 produced maximum-likelihood estimates most commonly at λ̂ = 0 or λ̂ = 1.These boundary estimates can misrepresent the phylogenetic signal in the generating process.
- 2.1. Uncertainty in parameter estimates: With a simulated tree of 281 tips, estimated λ values were closely centered around the true value, showing that the smaller-tree problem reflects insufficient data.
- 2.1. Uncertainty in parameter estimates: The data required for informativeness depend on tree size, topology, and the parameter or question being evaluated.Using the same 13-taxon Geospiza phylogeny, σ could be estimated more precisely than moderately different λ values.
- 2.1. Uncertainty in parameter estimates: Parametric bootstrap confidence intervals simulate datasets using the known phylogeny and estimated parameter, re-estimate the parameter, and use the resulting distribution’s percentiles.
- 2.2. The Monte Carlo approach: For two models, the Monte Carlo method simulates datasets under each model at its estimated parameters, re-estimates both models, and forms empirical distributions of the likelihood-ratio statistic δ.
- 2.2. The Monte Carlo approach: Larger δ supports model 1, while the proportion of simulated δ values exceeding the observed value estimates the p-value under model 0.Asymptotic χ2 approximations can be inadequate for phylogenetic comparisons.
3. An example using Anolis data
The Anolis example asks which evolutionary model best describes body-size data and whether the phylogeny can resolve that distinction. It compares homogeneous BM and OU models with heterogeneous OU models whose optima vary across the tree.
- The Anolis dataset contains mean body-size data for 23 Lesser Antilles lizard species and a phylogeny reconstructed from morphological and protein-electrophoretic evidence.
- The analysis evaluates both model fit and whether the available data are sufficient to resolve differences among candidate evolutionary models.
- Five candidate models include whole-tree Brownian motion and Ornstein-Uhlenbeck models, plus heterogeneous models allowing trait optima to vary across the phylogeny.
- The OU.3 model represents character displacement by assigning three body-size optima: an intermediate island-only optimum and larger and smaller optima on two-species islands.
- The heterogeneous models use phylogenetic paintings to specify where model parameters, including optima, may change across branches or clades.
4. Results
Monte Carlo model comparison quantifies both model-selection power and parameter uncertainty, while revealing that information criteria can misclassify models. Power varies substantially across comparisons, from strong discrimination to little information.
- Quantification of model choice: 93.6% power distinguished BM from OU.3 at a 5% false-positive rate, with only 2.5% of BM simulations exceeding the observed likelihood ratio.The observed likelihood ratio was 15 units.
- Model comparison: 98.8% of OU.15 simulations exceeded the 95% quantile of OU.3, yet the method did not select over-parameterized OU.15 from the observed data.OU.15 divided broader optima into finer, stronger-selection peaks that increased likelihood.
- Information criteria: 47.7% of simulations generated under OU.3 were incorrectly assigned to OU.15 by AIC, and the observed data were also assigned to OU.15.AIC falsely assigned 44% of OU.3 simulations to OU.4, while BIC and AICc could also prefer more complicated models.
- Model comparison: The OU.3–OU.4 comparison had relatively little power because the models’ likelihood-ratio distributions overlapped, despite the Monte Carlo method applying to non-nested models.The two models are not nested because OU.4 does not refine the OU.3 painting.
- Insufficient information: Only 7% power detected OU.1 selection against BM at α = 0.2, indicating that this phylogeny contained little information for distinguishing the models.At this α, trait correlations from the common ancestor decay to e^-0.2 = .81 of the BM expectation.
- Insufficient information: Power increased with the strength of selection α when OU.1 data were simulated across progressively larger α values.The analysis produced a power curve for detecting stabilizing selection.
5. Understanding the role of phylogeny shape and size on estimates of selection
The information available for estimating selection depends on phylogeny shape as well as taxon number. Trees with diversification farther in the past are less informative, while branching structure affects informativeness relative to common simulation trees.
- Understanding phylogeny shape and size: Phylogeny shape and size determine how much information leaf traits provide about evolutionary processes and selection strength.The analysis compares OU.1 with BM across trees of different shapes.
- Phylogeny shape: A 50-taxon tree became less informative as diversification occurred farther in the past and more evolutionary time was concentrated near the tips.This rescaling corresponds to the λ transformation; covariances from shared evolution are crucial for distinguishing models.
- Phylogeny shape: Pure-birth trees can be particularly informative because their many shallow nodes create highly correlated points, unlike branching patterns from density-dependent or niche-filling models.The paper cautions that pure-birth simulation trees often have more shallow nodes than trees generally observed in practice.
6. Discussion
The paper introduces simulation-based power analysis for phylogenetic model comparisons, showing when data cannot resolve proposed models and how this approach complements or improves information-criterion choices.
- Method and scope: Simulation-based model comparisons quantify power on a particular phylogeny and indicate whether a tree is too small or unstructured to resolve model differences.The approach provides a direct assessment of dataset informativeness for proposed character-evolution models.
- Empirical illustration: The anoles analysis selected OU.3, whereas AIC as used by Butler and King preferred the more complex OU.4 or OU.15 models.Power simulations can distinguish poor fit of a complex model from insufficient data when a simpler model is selected.
- Interpretation and limitations: Power depends on parameter values, so simulations under mismatched values can make failure to detect complex models reflect limited power.This dependence creates a practical interpretation boundary for power analyses applied to datasets with different parameter values.
- Comparison with information criteria: The method questions AIC use in phylogenetic model selection but retains a similar likelihood-ratio cutoff framework and permits explicit power–false positive tradeoffs.Unlike AIC, it compares models pairwise rather than selecting simultaneously among many models.
- Practical implication: Comparative datasets may contain limited information about evolutionary processes, motivating routine simulation-based assessment of model choice and power.This recommendation is framed as guidance for applying comparative methods to specific datasets.
- Implementation: The method is computationally heavier than fitting information-criterion models, requiring 2n simulations and 4n model fits for n replicates.The implementation addresses this burden through parallel computation in the pmc R package.