Source-linked AI summary
Simulation-based model selection for dynamical systems in systems and population biology
Tina Toni, Michael P. H. Stumpf
TL;DR
희소하고 잡음이 많은 데이터와 계산적으로 다루기 어려운 likelihood 때문에 기계론적 생물학 모델을 비교하기 어렵다. 이 논문은 ABC SMC model selection을 개발하고 influenza 및 JAK-STAT 데이터를 포함한 다양한 생물학 모델에 적용 가능함을 보인다.
문제
관측이 희소하거나 잡음이 많고, parameter가 알려지지 않았으며, likelihood를 이용할 수 없거나 계산적으로 다루기 어려울 때 후보 기계론적 모델 중 하나를 선택하기 어렵다.
방법
이 논문은 approximate Bayesian computation과 sequential Monte Carlo sampling을 사용하는 계산 효율적인 Bayesian model-selection framework를 개발한다.
결과
ABC SMC는 chemical reaction, Gibbs random-field, influenza, JAK-STAT 예제 전반에서 model selection을 뒷받침했으며, non-nested ordinary 및 time-delay differential-equation models도 포함했다.
시사점 및 한계
이 framework는 exact likelihood evaluation이 계산적으로 다루기 어렵고 데이터가 부족한 경우, 효율적으로 시뮬레이션할 수 있는 생물학 시스템으로 Bayesian model selection을 확장한다.
시사점 및 한계
수백 또는 수천 개 parameter를 갖는 모델에 이 방법을 일상적으로 적용하려면 추가적인 수치적 발전이 필요하다. 반복적인 simulation이 computationally costly하기 때문이다.
Abstract
from arXiv · showhide
Computer simulations have become an important tool across the biomedical sciences and beyond. For many important problems several different models or hypotheses exist and choosing which one best describes reality or observed data is not straightforward. We therefore require suitable statistical tools that allow us to choose rationally between different mechanistic models of e.g. signal transduction or gene regulation networks. This is particularly challenging in systems biology where only a small number of molecular species can be assayed at any given time and all measurements are subject to measurement uncertainty. Here we develop such a model selection framework based on approximate Bayesian computation and employing sequential Monte Carlo sampling. We show that our approach can be applied across a wide range of biological scenarios, and we illustrate its use on real data describing influenza dynamics and the JAK-STAT signalling pathway. Bayesian model selection strikes a balance between the complexity of the simulation models and their ability to describe observed data. The present approach enables us to employ the whole formal apparatus to any system that can be (efficiently) simulated, even when exact likelihoods are computationally intractable.
1 서론
복잡한 생물학적 모델은 흔히 simulation에 의존하므로, 데이터가 희소하거나 잡음이 많거나 likelihood를 계산할 수 없을 때 경쟁 모델을 합리적으로 비교하기 어렵다. 이 논문은 sequential Monte Carlo를 사용하는 ABC model-selection framework를 개발해 Bayesian model-selection 체계를 유지하고, 이를 simulation 및 실제 생물학적 시스템 전반에 적용해 보인다.
- 관측값이 희소하고 잡음이 많은 경우가 빈번한 반면, 표준 model-selection 방법은 적절한 확률 모델을 필요로 하므로 simulation 기반 분석은 후보 생물학적 모델의 비교를 복잡하게 만든다.
- ABC는 likelihood 계산을 simulation과 데이터의 비교로 대체하며, d(D0, D∗) ≤ ϵ일 때 파라미터를 받아들이고 ϵ가 충분히 작아지면 posterior를 근사한다.
- 타당한 모델이 여러 개 존재할 때 Bayesian selection은 likelihood를 파라미터에 대해 적분해 모델의 marginal posterior probabilities를 비교하므로, 임의의 비중첩 모델도 순위를 매길 수 있다.
- 이 논문은 sequential Monte Carlo에 기반한 계산 효율적인 ABC model-selection formalism을 개발하고, 이를 chemical reaction dynamics, Gibbs random fields, influenza spread, JAK-STAT signalling에 적용한다.
2 모델 선택을 위한 ABC
이 절에서는 모델 지표와 매개변수를 함께 샘플링한 뒤 매개변수를 주변화하여 marginal model posterior를 추정한다. ABC SMC model selection을 제시하며, ABC rejection보다 효율적으로 target posterior를 근사하도록 simulation tolerance를 점진적으로 엄격하게 조정한다.
- Joint-space formulation: Joint-space model selection은 P(m|D0)를 추정하기 위해 P(θ,m|D0)를 구한 뒤 매개변수를 주변화한다.Particle은 model indicator와 parameter vector를 모두 포함하므로, 동일한 joint posterior approximation으로 marginal parameter distribution을 얻을 수 있다.
- ABC SMC model selection: ABC rejection과 비교하면 SMC extension은 두 model-selection approach를 모두 computationally more efficient하게 만들어, moderately parameterized model에서 rejection의 감당하기 어려운 비용을 해결한다.이 논문은 joint-space ABC SMC approach를 제시하고, 상세한 유도와 논의는 supplementary material로 안내한다.
- ABC SMC model selection: ABC SMC는 joint model–parameter particle을 감소하는 tolerance ϵ1 > … > ϵT를 거치며 전파하여 target posterior를 근사할 때까지 진행한다.알고리즘은 simulated-data distance를 조건으로 하는 intermediate distribution에서 샘플링하고, population 사이에서 model 및 parameter perturbation kernel을 사용한다.
- Algorithm specification: 알고리즘에는 prior, distance function, tolerance schedule, perturbation kernel이 필요하며, 예시에서는 uniform prior와 adaptive truncated-uniform or Gaussian kernel을 사용한다.Uniform prior는 사전에 candidate model을 동등하게 그럴듯한 것으로 만들면서 실현 가능한 parameter region을 정의하고, kernel은 이전 population의 parameter range를 이용해 조정된다.
3 결과
ABC SMC는 합성 데이터, 인플루엔자, JAK-STAT 응용 전반에서 확률적 생물학 모델을 선별했다. 알려진 모델을 정확히 식별하고 ABC rejection보다 계산 효율을 높였으며, 생물학적으로 해석 가능한 대안에 유리한 근거를 제시했다.
- 응용 사례: 결과는 autocatalytic 및 non-autocatalytic 반응, Gibbs random fields, 인플루엔자 유행, 그리고 경쟁하는 STAT5 shutoff 메커니즘에 대한 model selection을 포괄한다.응용 사례는 서로 다른 유행이 동일한 모델을 공유하는지 검토하고, signalling biology에서 대안적 mechanistic hypothesis를 비교한다.
- 확률적 반응 동역학: model 2에서 생성한 synthetic reaction-kinetics 데이터에 ABC SMC를 적용한 결과, 올바른 모델을 high confidence로 식별했다.데이터셋에는 20개 시점의 Y 측정값이 포함되었으며, k2 = 30, X0 = 40, Y0 = 3으로 시뮬레이션했다.
- Gibbs random fields: ABC SMC는 Gibbs random fields의 posterior model distribution을 정확히 추정했으며 ABC rejection과 비교해 considerable computational speed-up을 달성했다.비교에는 매개변수 값 전반에서 시뮬레이션한 1000개 데이터셋과 sequence length n = 100이 사용되었다.
- 인플루엔자 동역학: 두 H3N2 유행은 동일한 epidemiological characteristics를 공유하는 것으로 보인 반면, 서로 다른 인플루엔자 strain은 주로 지역사회 전반에서 서로 다른 spread dynamics를 보였다.서로 다른 strain에서는 two-parameter model의 posterior probability가 무시할 수 있을 정도로 작았으며, H3N2 데이터셋을 결합하면 model (1)보다 model (2)를 지지하는 근거가 제공되었다.
4 논의
ABC SMC model-selection methodology는 ODE 및 time-delay model 비교를 포함해 simulated biological systems에 폭넓게 적용할 수 있다. 이를 매우 큰 model에 일상적으로 적용하는 데에는 여전히 큰 계산 비용이 들므로, 수치적 개선과 병렬화가 필요하다.
- 방법론적 범위: ABC SMC는 실험 데이터가 부족하고 측정이 제한적인 systems를 위한 유용하고 폭넓게 적용 가능한 model-selection method를 제공한다.이 framework는 qualitative modelling을 포함해 simulation 및 modelling approaches 전반에 적용된다.
- 방법론적 범위: 이 procedure는 ODE 및 time-delay differential-equation models의 설명력을 비교할 수 있으며, 효율적인 simulation이 가능하다면 dynamical systems를 넘어 확장된다.그 범위는 효율적인 simulation approaches의 이용 가능성에 의해서만 제한된다고 설명된다.
- 한계와 향후 연구: 수백 개 또는 수천 개의 parameter 를 갖는 complex systems, computational, population-biology models에 일상적으로 사용하려면 반복 simulation에 큰 비용이 들기 때문에 추가적인 수치적 개발이 필요하다.저자들은 계산 효율을 높이는 핵심 경로로 SMC-based ABC methods의 병렬화를 제시한다.
5 결론
결론은 biomedical simulation 연구에서 경쟁 모델의 상대적 성능과 신뢰성을 평가할 수 있는 통계적으로 타당한 추론 방법이 시급히 필요하다는 점을 강조한다.
- 5 결론: Simulation-based biomedical 연구가 확대됨에 따라 모델 간 차이를 식별할 수 있는 통계적으로 타당한 model-selection 절차가 시급히 필요하다.이러한 방법은 모델의 상대적 성능과 신뢰성을 평가해야 한다.
보충 자료 A: ABC SMC 모델 선택 알고리즘의 유도
이 보충 자료는 sequential importance sampling과 ABC intermediate distributions로부터 ABC SMC 모델 선택 알고리즘을 유도한다. 실용적인 알고리즘으로 joint model–parameter selection을 제시하고, 계산 비용이 높거나 특수한 normalization 처리가 필요한 naive sequential 및 marginal-likelihood 대안도 유도한다.
- ABC SMC 구성 요소: ABC SMC는 관측값과의 거리가 tolerance ϵ_t보다 작을 때 채택되는 prior-weighted simulated datasets로부터 intermediate distributions를 구성하며, proposal은 prior에서 초기화한 뒤 perturbation한다.Intermediate distributions는 고정된 parameter를 조건으로 생성된 datasets를 사용하며, 이후 proposal은 perturbation kernels를 통해 이전 particles의 정보를 활용한다.
- II) Joint space에서의 ABC SMC model selection: Algorithm II는 perturbation kernels와 sequential importance weights를 사용해 model indicators와 parameters에 걸쳐 ABC SMC model selection을 jointly 수행한다. 이는 본 논문과 예시에서 사용한 실용적인 알고리즘이다.Algorithms I과 III은 계산 비용이 지나치게 높아 사용하기 어렵고 비실용적인 것으로 설명된다.
- I) Naive ABC SMC model selection: Algorithm I은 model probabilities를 sequential하게 갱신하지만 각 model의 parameters를 prior에서 반복적으로 sampling하므로, parameter integration의 계산 비용이 높다.이 보충 자료는 이전 particles에 포함된 parameter 정보를 활용하는 방법으로 joint-space sampling을 도입한다.
- III) Marginal likelihood의 ABC SMC 근사: Marginal-likelihood 접근법은 ABC acceptance rates 또는 ABC SMC normalization constants로부터 P(D0|m)을 추정한 뒤, 이 추정값을 posterior model probabilities로 변환한다.ABC SMC에서는 target distribution이 unnormalized이므로 통상적인 normalized output을 직접 사용할 수 없다. 대신 유도 과정에서 intermediate marginal likelihoods를 추적한다.