Source-linked AI summary

Predicting evolution from the shape of genealogical trees

Richard A. Neher, Colin A. Russell, Boris I. Shraiman

arXiv:1406.0789v2q-bio.PE

TL;DR

재구성된 계통수 형태에 표본 개체의 상대적 적합도와 미래 집단 구성을 예측할 정보가 포함되어 있는지는 여전히 핵심 예측 문제다. 저자들은 계통수의 분지 패턴에서 적합도를 추론하고, 주변 계통수 길이를 요약하는 local branching index로 서열의 순위를 매긴다. LBI로 순위를 매긴 서열은 인플루엔자 A/H3N2 선조 계통을 높은 정확도로 예측한다.

  • 문제

    재구성된 계통수 형태에 표본 개체의 상대적 적합도와 미래 집단 구성을 예측할 정보가 포함되어 있는지는 여전히 핵심 예측 문제다.

  • 방법

    저자들은 계통수의 분지 패턴에서 적합도를 추론하고, 주변 계통수 길이를 요약하는 local branching index를 사용해 서열의 순위를 매긴다.

  • 결과

    LBI로 순위를 매긴 서열은 인플루엔자 A/H3N2 선조 계통을 높은 정확도로 예측한다.

  • 핵심 해석 및 한계

    인플루엔자 계통수에서 도출한 의미 있는 예측은 순환하는 A/H3N2 집단에 지속적인 적합도 변이가 있음을 시사한다.

  • 핵심 해석 및 한계

    큰 효과의 돌연변이가 유발하는 항원 클러스터 전환기에는 예측 성능이 흔히 최적 수준에 미치지 못한다.

Abstract

from arXiv · show

Given a sample of genome sequences from an asexual population, can one predict its evolutionary future? Here we demonstrate that the branching patterns of reconstructed genealogical trees contains information about the relative fitness of the sampled sequences and that this information can be used to predict successful strains. Our approach is based on the assumption that evolution proceeds by accumulation of small effect mutations, does not require species specific input and can be applied to any asexual population under persistent selection pressure. We demonstrate its performance using historical data on seasonal influenza A/H3N2 virus. We predict the progenitor lineage of the upcoming influenza season with near optimal performance in 30% of cases and make informative predictions in 16 out of 19 years. Beyond providing a tool for prediction, our ability to make informative predictions implies persistent fitness variation among circulating influenza A/H3N2 viruses.

결과 · 계통수에서의 fitness 분포 · fitness 추론은 모델 가정에 민감하지 않다

이 연구는 selection-biased diffusion과 message passing을 이용한 계통수 기반 fitness 추론 방법을 개발한 뒤, 진화가 discrete mutation을 통해 일어날 때 그 강건성을 검증한다. Simulation에서 fitness ranking은 실제 fitness와 상관을 보이며 mutation rate가 증가할수록 향상되고, model parameter Γ에는 약하게만 의존한다.

  • 계통수에서의 fitness 분포: 이 방법은 branchwise propagator와 재구성된 genealogical tree 전체의 message passing을 결합해 node fitness를 추론한다.Tree를 따라 정보를 위아래로 전파하여 internal node와 external node의 marginal fitness distribution을 계산한다.
  • 계통수에서의 fitness 분포: Joint fitness distribution은 ancestral node와 sampled node에 relative fitness를 부여하고, tree branch를 따라 propagator의 곱으로 factorize되며 Z(T)로 normalization된다.각 node의 fitness는 sampling time에서 population mean을 기준으로 측정되며, 그 구조는 phylogenetic likelihood 계산과 유사하다.
  • 계통수에서의 fitness 분포: Branch propagator는 짧은 interval 동안 ancestral-fitness information을 유지하지만, 긴 interval에서는 population distribution에 가까워진다.Backward-time inference에서는 Bayesian inversion을 사용해 descendant로부터 ancestral fitness를 추정한다.
  • 계통수에서의 fitness 분포: 이 모델은 fitness evolution을 selection-biased random walk로 다루며, 여러 mutation이 fitness에 기여할 때 이를 diffusion으로 근사한다.Selection은 더 높은 fitness를 가진 lineage의 생존을 선호하도록 편향시키고, branch fitness diffusion constant는 확률적 fitness 변화를 포착한다.
  • 결과: 약 0.5 수준의 Spearman’s correlation coefficient는 추론된 fitness ranking이 실제 ranking을 잘 예측함을 보여주며, mutation rate가 증가할수록 ranking이 향상된다.이 결과는 전형적인 simulation에서 나타나며 adaptive-mutation rate 전반에서 일관되고, Γ에는 약하게만 의존한다.
  • Fitness 추론은 모델 가정에 민감하지 않다: Simulation은 infinitesimal-mutation SBD model을, 고정된 fitness variance와 beneficial mutation rate nA = 0.02, . . . , 0.16 per generation을 갖는 discrete-mutation evolution과 비교했다.그 밖의 simulated genome은 대부분 deleterious mutation으로 구성되어, 변화하는 환경에서 adaptive evolution이 일어날 때 모델의 강건성을 검증할 수 있었다.
  • Fitness 추론은 모델 가정에 민감하지 않다: Γ = 0.2와 0.5를 사용한 분석에서는 fitness diversity가 소수의 mutation에 의해 지배될 때 낮은 mutation rate에서 larger Γ performs better임을 확인했다.Mutation rate를 높이면 SBD approximation이 개선되지만, Γ의 선택은 전반적으로 약한 효과만 보인다. Γ는 stochastic diffusion과 selection의 상대적 중요성을 나타낸다.

높게 추론된 fitness는 progenitor sequence를 예측한다 · 휴리스틱 순위화로서의 국소 branching density

높게 추론된 fitness가 부여된 sequence는 미래 population의 progenitor lineage를 식별하는 경향이 있으며, model-independent한 local branching index는 주변 tree length를 이용해 이에 거의 견줄 만한 순위를 제공한다. 이 휴리스틱은 inference parameter에 비교적 둔감하고, 점진적으로 변하는 fitness와 연관된 branching pattern을 포착한다.

  • 높게 추론된 fitness는 progenitor sequence를 예측한다: 고 fitness 예측은 simulation에서 미래 progenitor sequence를 식별하며, Fig. 2D는 이들의 normalized sequence distance를 post-hoc optimal pick과 비교해 평가한다.비교에는 highest-fitness prediction과 200 generations 후 population 사이의 Hamming distance를 사용하며, 이를 present-to-future distance 평균으로 정규화한다.
  • 높게 추론된 fitness는 progenitor sequence를 예측한다: 100-generation interval에 걸쳐 sampling한 sequence는 한 generation에서 200 sequence를 sampling한 경우와 highly similar한 fitness-inference 결과를 낸다.
  • 휴리스틱 순위화로서의 국소 branching density: Fitness ranking과 progenitor prediction은 Γ와 ω/σ에 거의 의존하지 않지만, faithful posterior-fitness inference에는 numerical branch propagation과 parameter knowledge가 필요하다.이러한 parameter insensitivity는 ranking이 주로 더 보편적인 tree quantity에 의해 결정됨을 시사한다.
  • 휴리스틱 순위화로서의 국소 branching density: 짧은 time period에서는 downstream branch length가 늘어날수록 internal-node fitness estimate가 증가하는데, subtree length가 inferred fitness를 높은 값 쪽으로 편향시키기 때문이다.고정된 descendant 수에서는 star-like subtree가 length를 최대화하며, rapid branching 또는 multiple merger를 나타낸다.
  • 휴리스틱 순위화로서의 국소 branching density: Lineage를 따라 fitness가 점진적으로 변하므로, high-fitness node는 fitness decorrelation에 의해 크기가 정해지는 neighborhood 안에서 upstream and downstream branching을 모두 보일 것으로 예상된다.
  • 휴리스틱 순위화로서의 국소 branching density: Local branching index λi(τ)는 exponentially discounted surrounding tree length를 기준으로 internal node와 terminal node의 순위를 매기며, τ가 neighborhood scale을 정한다.SBD model에서 τ는 high-fitness lineage의 equilibration timescale에 해당하며, 그 크기는 Tc/√log N 정도다.
  • 휴리스틱 순위화로서의 국소 branching density: LBI ranking은 complex SBD fitness inference만큼 almost as accurate하며, Fig. 3은 pairwise diversity와 memory timescale에 걸쳐 Spearman correlation과 true fitness를 비교한다.LBI는 posterior fitness inference에 사용되는 것과 동일한 message-passing technique으로 효율적으로 계산할 수 있다.

계절성 인플루엔자 A/H3N2 progenitor lineage 예측

저자들은 1995–2013년 계절성 인플루엔자 A/H3N2 HA1 서열의 genealogical tree를 사용해 다음 북반구 겨울철의 progenitor를 예측했다. 이들의 ranking은 이후 순환한 lineage를 가변적인 정확도로 식별했으며, Łuksza and L¨assig (2014)의 influenza-specific predictor와 비슷한 성능을 보였다.

  • 예측 설정: 이 방법은 오월–이월에 지역별로 최대 100개의 HA1 서열을 분석해, 다음 십월–삼월 시즌에 유행할 가장 가까운 근연 계통을 예측했다.표본은 1995년부터 2013년까지 아시아와 북아메리카를 포괄했으며, 공개적으로 이용 가능한 Influenza Research Database 서열과 FastTree로 구축한 maximum-likelihood tree를 사용했다.
  • 예측 성능: 가장 높은 순위의 internal node는 1997–1999년, 2003년, 2006–2009년, 2013년을 상당히 잘 예측했지만, 1995년, 1996년, 2002년에는 실패했다.나머지 연도에서는 예측 정확도가 중간 수준이었으며, 가장 높은 순위의 external node도 1997년을 제외하면 비슷한 성능을 보였다.
  • 기존 predictor와의 비교: 연도별로 internal-node ranking은 Łuksza and L¨assig (2014)와 비슷했지만, external-node ranking은 약간 더 나빴다.그럼에도 두 접근법은 연도별 예측이 매우 유사했는데, 이는 이들의 model이 downstream synonymous mutation을 사용해 epistatic interaction을 포착하기 때문일 가능성이 있다.
  • 평가 지표: 예측 품질은 normalized genetic distance로 정량화했으며, d = 0은 optimal prediction, d = 1은 random pick을 의미한다.저자들은 연도별 d의 평균을 구하고 연도를 bootstrap해 internal-node 및 external-node ranking을 Łuksza and L¨assig (2014) 및 naive predictor와 비교했다.

추론된 fitness 증가와 epitope mutation의 연관성 … fitness inference algorithm의 유도

이 방법은 small-effect mutation model에서 genealogical branching pattern으로부터 relative fitness를 추론하고, 이를 이용해 미래 population을 예측한다. 높은 inferred fitness는 epitope substitution과 연관되며, genetic diversity가 낮을 때도 predictive performance가 유지된다.

  • 추론된 fitness 증가와 epitope mutation의 연관성: inferred fitness increase가 상위 사분위수에 속하는 branch에는 nonsynonymous substitution이 더 많이 나타나며, epitope A–D로 제한하면 enrichment가 약 2배로 증가한다.여기에 일곱 Koel loci까지 제한하면 enrichment가 약간 더 증가하지만, 해당 locus 수가 적어 추가 enrichment를 검출할 power가 제한된다.
  • 논의: 높은 inferred fitness와 epitope substitution의 연관성은 model이 sequence와 protein structure에 agnostic함에도 antigenic novelty가 influenza evolution을 주도한다는 해석과 일치한다.관련 epitope는 역사적으로 높은 d_n/d_s를 보여 positive selection을 시사한다.
  • 논의: 이 algorithm은 adaptive evolution을 genealogical tree 위의 fitness dynamics로 model하고 individual node의 fitness를 확률적으로 추론한다. simulation에서는 highest-fitness sequence가 미래 progenitor와 일치하는 경향을 보인다.이 framework는 selection-biased diffusion model을 사용하며, evolution은 다수의 small-effect mutation을 통해 진행된다.
  • 논의: predictive power는 nonneutral genetic diversity가 증가할수록 높아지지만, diffusion model이 poor approximation인 경우에도 pairwise distance가 상당히 낮은 수준에서 유지된다.이러한 지속성은 model의 best-fitting regime를 넘어 fitness와 genealogical structure 사이에 관계가 있음을 뒷받침한다.
  • 논의: 특정 large-effect mutation이 antigenic change를 급격히 일으키는 antigenic-cluster transition이 발생한 연도에는 prediction이 제한된다. sampling, migration, demographic structure도 genealogical pattern을 교란할 수 있다.reconstructed influenza genealogy가 여전히 informative하더라도 이러한 요인은 prediction을 저해할 수 있다.
  • fitness inference algorithm의 유도: reconstructed genealogy의 branching pattern에는 sampled individual의 relative fitness에 관한 정보가 담겨 있으며, fitness difference가 여러 mutation에 의존할 때 미래 population composition을 예측할 수 있다.이 algorithm은 input으로 reconstructed genealogy만 필요하며, RNA virus부터 cancer cell population까지 다양한 응용을 목적으로 한다.
  • fitness inference algorithm의 유도: 이 inference는 branching-process approximation에서 유도된다. offspring sampling probability로부터 branch propagator를 수치적으로 구하고, 이를 결합해 posterior fitness distribution을 얻는다.이 방법은 한 시점의 static sequence set에서 작동하며, influenza history는 validation에만 사용한다.

자손 수 분포

이 논문은 birth, death, mutation, environmental deterioration을 포함하는 P(n|x,t)에 대한 backward master equation을 통해 자손 수를 모델링한다. Mutational effect가 short-tailed이고 mutation rate가 높을 때, generating function과 reproductive value는 coalescence 이전의 lineage growth를 정량화한다.

  • 자손 수 분포: 자손 분포 P(n|x,t)는 rate 1+x의 birth, rate one의 death, mutation, velocity v의 environmental deterioration을 포함하는 backward master equation에서 유도된다.Mutation 항은 total rate u에서 fitness effect μ(s)에 대해 평균하며, environment가 시간에 따라 forward로 악화되므로 fitness는 backward로 Δtv만큼 증가한다.
  • 자손 수 분포: Fitness effect가 short-tailed이고 mutation rate u가 전형적인 effect보다 클 때, mutation dynamics는 D = u⟨s^2⟩/2와 σ^2 = v − u⟨s⟩로 특성화되는 directional 및 diffusive fitness change로 환원된다.Generating-function equation은 genealogical tree에서의 fitness distribution을 근사하기 위해 수치적으로 푼다.
  • 자손 수 분포: Reproductive value R(x,t)는 t generations 이후의 expected offspring number로 정의되며, generating function에서 얻어지고 linear equation을 만족한다.이 근사는 coalescence time Tc에 비해 짧은 시간에만 유효하다.
  • 자손 수 분포: 처음에 lineage는 rate x로 clonally grow하며, 나머지 population이 rate σ^2로 adapt함에 따라 growth가 느려지고, descendant mutation이 offspring fitness를 변화시킨다.이러한 효과가 coalescence 이전의 lineage dynamics를 결정한다.

계통 샘플링 확률

생성함수 φ_ω(x,t)는 계통이 샘플에 포함될 확률을 나타내며, φ_ω가 작거나 포화된 경우에 정확한 점근형을 갖는다. 이 근사는 초기 조건, 장시간 거동, 중립 극한을 만족한다.

  • 계통 샘플링 확률: φ_ω(x,t)는 비율 ω = M/N인 샘플에 계통이 포함될 확률을 나타낸다.여기서 M은 샘플 크기이고 N은 개체군 크기다.
  • 계통 샘플링 확률: 샘플링 확률은 n명의 자손 중 아무도 샘플링되지 않을 확률을 합산한 뒤, 그 합을 1에서 빼서 구한다.각 항 (1 − ω)^n은 n명의 자손 중 아무도 샘플에 들어가지 않을 확률이다.
  • 계통 샘플링 확률: φ_ω가 작거나 충분히 큰 x로 인해 포화될 때 근사가 정확하며, φ_ω(x,t) ≈ x이다.또한 φ_ω(x,0) = ω를 만족하고, 장시간에 x > 0에서 x에 접근하며, 중립 극한에서 φ_ω(0,t) = ω/(1 + ωt)를 복원한다.

Branch propagator

branch propagator는 표본에 기여하는 비분기 계통을 조건으로 자손의 fitness distribution을 기술한다. 그 dynamics에는 mutation에 의한 fitness diffusion, 계통 reproduction, sampling이 반영되며, 특징적인 조상 및 terminal-progenitor distribution을 산출한다.

  • 지배 방정식: propagator equation은 reproductive growth, sampling-dependent non-branching, mutation-driven fitness drift, fitness diffusion을 결합하며, delta-function initial condition을 갖는다.유도 과정에서는 unsampled branch가 현재 표본에 기여하지 않는 birth event를 다루고, 시간에 따라 mean fitness를 이동시킨다.
  • 가정: small-fitness-difference assumption은 fitness variance가 작은 population에 적절하며, 이를 위반하면 propagator의 정량적 거동은 달라지지만 정성적 거동은 달라지지 않는다.유도 과정에서는 y ≪1을 가정하며, 이는 작은 σ와 세대 간 작은 fitness difference에 해당한다.
  • Terminal branch propagator: positive-fitness ancestor의 경우 terminal progenitor probability는 처음에는 ancestral age와 함께 증가하지만, 계통의 survival 가능성이 낮아지면서 결국 감소한다.짧은 시간과 중간 정도의 parental fitness에서는 terminal propagator가 reproductive value로 단순화된다.
  • 수치적 거동: 수치해는 descendant fitness distribution이 시간에 따라 broaden되며, ancestral fitness distribution은 먼 과거로 갈수록 fit ancestor를 선호하는 공통 곡선으로 converge함을 보여준다.fit ancestor는 더 많은 offspring를 남기므로 표본에 포함될 가능성이 높지만, 지나치게 fit한 ancestor에는 계통이 생존하지 못하는 효과가 반대로 작용한다.

Tree 기반 추론

이 방법은 계보수 위에서 fitness probability를 인수분해하고 반복적 message passing을 통해 주변화하여 조상 fitness를 추론한다. downstream tree length가 더 긴 최근 노드는 더 높은 fitness를 갖는 쪽으로 추론되며, 별 모양 topology는 예외적으로 fitness가 높은 clone의 빠른 확장을 나타낸다.

  • Tree 기반 추론: 공동 조상 fitness distribution은 tree node와 branch에 대해 인수분해되므로, population size에 제약이 없고 branch 간 상호작용이 없는 경우 polytomy를 허용할 수 있다.이러한 근사는 selection이 지배적인 경우에 적절하다고 본다. coalescent property가 명시된 제약에 약하게만 의존하기 때문이다.
  • Tree 기반 추론: 반복적 message passing은 leaf와 조상 fitness variable을 적분해 제거하고, upstream 및 downstream branch message를 결합하여 각 node의 marginal fitness distribution을 구한다.이 방법은 node로 들어오는 모든 message를 곱한 뒤 그 결과를 정규화하여 marginal distribution을 계산한다.
  • Tree 기반 추론: Mean marginal fitness를 사용해 internal node와 external node의 순위를 정하며, downstream total tree length는 최근 node를 high-fitness edge 쪽으로 편향시킨다.descendant 수가 고정되면 downstream tree length는 star topology에서 최대화된다.
  • Tree 기반 추론: Star-like genealogy는 예외적으로 fitness가 높은 개체가 세운 clone의 빠른 확장과 연관된다.이 추론은 multiple merger와 확장된 downstream tree length를 founder의 예외적으로 높은 fitness와 연결한다.

Local branching index (LBI) 계산

Local branching index (LBI)는 각 node 주변의 tree length에 exponential discount를 적용해 계산한다. 계산에는 message passing이 사용되며, node의 discounted tree length를 평가하기 전에 자식으로부터의 up-messages와 부모로부터의 down-messages를 결합한다.

  • Local branching index (LBI) 계산: LBI는 node 주변의 integrated exponentially discounted tree length를 측정한다.계산은 fitness distribution을 평가하는 데 사용되는 message-passing framework와 유사하다.
  • Local branching index (LBI) 계산: Up-messages는 branch length를 사용해 각 node의 자식에서 부모를 향해 정보를 전파한다.node i의 경우 식은 branch length b_i와 자식 ij에 대한 합에 의존한다.
  • Local branching index (LBI) 계산: Down-messages는 각 부모에서 자식으로 정보를 전파해 message-passing 계산을 완성한다.모든 up-messages와 down-messages를 계산한 뒤, 이들이 exponentially discounted tree length를 결정한다.

추론 알고리즘 구현

추론 알고리즘은 이산화된 fitness grid에서 message passing을 수행해 재구성된 tree로부터 node-fitness marginal을 추정한다. 예측에서는 expected fitness에 따라 internal 및 external node의 순위를 매기며, branch length는 model time unit으로 변환하고 clade expansion rate는 longitudinal frequency에서 별도로 추정한다.

  • 구현: Message passing은 Python에서 SciPy와 NumPy를 사용해 구현된 discrete fitness grid에서 모든 external 및 internal tree node의 marginal fitness distribution을 계산한다.구현에는 `survival_gen_func`와 `fitness_inference` class가 사용된다.
  • 예측 절차: 알고리즘은 maximum-likelihood FastTree tree를 구축해 fitness inference에 전달하고, expected fitness가 가장 높은 node를 예측한다.FastTree는 짧은 branch를 더 잘 분해하도록 수정되었다.
  • Model parameterization: Branch propagation은 fitness diffusion D, fitness standard deviation σ, sampling fraction ω로 parameterize되며, time은 σ−1 unit으로 측정하고 selection strength는 σ unit으로 측정한다.차원 있는 diffusion constant는 Γ = Dσ−3이고, generating-function initial condition은 φω(x, 0) = ω/σ이다.
  • Time calibration: Branch length는 π ≈2µ⟨T2⟩ 및 SBD model의 경우 ⟨T2⟩σ ≈Γ−1을 사용해 nucleotide distance에서 σ−1 time unit으로 converted from nucleotide distance된다.Conversion factor β는 선택한 Γ, mutation rate µ, average pair coalescent time에 따라 달라진다.
  • Frequency analysis: Clade expansion rate는 pseudocount 5를 사용해 세 개의 동일한 May–February interval에서 측정한 log frequency에 line을 fitting하여 추정한다.frequency는 각 interval에서 각 internal node 아래에 있는 sequence의 fraction을 정량화한다.

시뮬레이션 · 인플루엔자 데이터

이 연구는 개체 기반 시뮬레이션과 전처리한 인플루엔자 A/H3N2 HA1 서열을 결합해 mutation과 selection 하에서 genealogical evolution을 조사한다. 시뮬레이션에서는 genomic mutation rate를 변화시키고, 인플루엔자 데이터는 1968년부터 2014년까지 수집된 human-host virus를 포괄한다.

  • 시뮬레이션: 시뮬레이션에는 σ = 0.03으로 고정된 fitness variance를 갖는 개체 기반 집단에 FFPopSim을 사용한다.시뮬레이션에서는 무작위 개체의 무작위 위치에 µ의 rate로 mutation을 도입한다.
  • 시뮬레이션: 전체 genomic mutation rate u = Lµ는 L = 2000개의 simulated site에 걸쳐 0.016에서 0.256까지 변화시킨다.이 parameterization은 simulated genome length를 고정한 채 mutation input을 변화시킨다.
  • 시뮬레이션: Mutation은 exponential distribution에서 추출한 deleterious effect를 갖도록 기본 설정한다.시뮬레이션 passage에서는 environment-changing setup도 설명하지만, 제공된 텍스트는 세부 사항이 나오기 전에 끝난다.
  • 인플루엔자 데이터: 인플루엔자 데이터셋에는 1968년부터 2014년까지의 전체 HA1 domain을 포괄하는 human-host A/H3N2 서열이 포함된다.서열은 IRD에서 다운로드했으며 IRD의 default alignment feature로 정렬했다.
  • 인플루엔자 데이터: HA1 alignment를 수동으로 검토하고 HA1 domain으로 trimmed했다.이 preprocessing step은 IRD의 default setting으로 alignment한 뒤 수행했다.
  • 인플루엔자 데이터: 데이터셋에서는 명백한 outlier, laboratory strain, indel이 있는 서열, 그리고 ambiguous nucleotide가 4개를 초과하는 서열을 제외한다.이러한 exclusion은 alignment를 검토한 뒤 수동으로 수행했다.

부록 A: Figure 2 – 보충자료 · 부록 B: Figure 3 – 보충자료 · 부록 D: Figure 4 – 보충자료

보충자료는 예측 성능이 genetic diversity와 memory scale에 좌우되며, 높은 LBI가 미래 population success와 연관된 sequences와 clades를 식별함을 보인다. LBI predictions는 대안적 forecasting method와 대체로 유사하고, 때로는 더 우수하다.

  • 부록 A: Figure 2 – 보충자료: Prediction rank correlation은 pairwise diversity가 증가할수록 높아지며, 거리가 작을 때는 더 큰 Γ가, 거리가 클 때는 더 작은 Γ가 선호된다.작은 거리 구간은 effect가 큰 mutation이 소수 존재하는 경우에 해당하고, 큰 거리는 여러 locus에 걸쳐 fitness variation이 분산된 경우에 해당한다.
  • 부록 A: Figure 2 – 보충자료: 중간 또는 높은 mutation rate에서는 continuous sampling이 rank correlation을 낮추지 않으며, predicted strain과 미래 population 사이의 distance도 유사하게 변화한다.이 보충자료에서는 한 시점이 아니라 100 generation에 걸쳐 sampling한 200 simulated sequence를 사용한다.
  • 부록 B: Figure 3 – 보충자료: LBI가 가장 높은 sequence는 200 generation 후 population의 progenitor에 가까이 위치하는 경향이 있다.이 보충자료에서는 predicted sequence와 sampled population 및 future population 사이의 평균 distance를 비교해 sequence distance를 측정한다.
  • 부록 B: Figure 3 – 보충자료: memory time scale τ가 2^-6에서 4로 변함에 따라 LBI-based prediction이 달라지며, internal node와 external node에 대해 별도의 trajectory가 나타난다.이 보충자료에서는 연도와 memory scale에 따른 prediction variation을 조사한다.
  • 부록 B: Figure 3 – 보충자료: 많은 연도에서 LBI가 가장 높은 sequence는 Łuksza and L¨assig (2014)가 예측한 sequence와 매우 유사하지만, 다른 연도에서는 두 method 중 하나가 미래에 더 가깝다.Łuksza and L¨assig는 epitope position에서 amino-acid distance를 최소화하는 것을 목표로 했지만, 여기서 비교하는 것은 nucleotide distance다.
  • 부록 B: Figure 3 – 보충자료: 높은 LBI는 clade expansion을 예측한다. LBI 순위가 높은 clade는 다음 연도로 확장되는 clade 가운데 과대표집된다.분석 대상은 May부터 February까지 수집한 샘플에서 빈도가 75% 미만인 clade다.
  • 부록 D: Figure 4 – 보충자료: LBI에 기반한 Influenza A/H3N2 prediction accuracy는 memory time scale τ가 증가할수록 향상된다.Accuracy는 future sample까지의 nucleotide distance로 측정하며, optimal pick은 d = 0, random pick은 d = 1이 되도록 scaling한다.
Loading 1406.0789v2…