Source-linked AI summary
Automated versus do-it-yourself methods for causal inference: Lessons learned from a data analysis competition
Vincent Dorie, Jennifer Hill, Uri Shalit, Marc Scott, Dan Cervone
TL;DR
확장되는 causal inference 문헌은 다양한 방법을 제공하지만, 비교가 제한적이어서 응용 연구에 유용한 전략을 식별하기 어렵다. 저자들은 다양한 assignment mechanism과 response surface를 포괄하는 77개 simulated scenario에서 방법을 평가하는 대규모 competition을 진행했다. response surface를 유연하게 모델링한 방법이 일관되게 가장 우수한 성능을 보였으며, 특히 비선형 response surface와 treatment-effect heterogeneity가 존재하는 경우에 그러했다.
문제
확장되는 causal inference 문헌은 다양한 방법을 제공하지만, 제한된 비교로 인해 응용 연구에 유용한 전략을 식별하기 어렵다.
방법
저자들은 다양한 assignment mechanism과 response surface를 포괄하는 77개 simulated scenario에서 방법을 평가하는 대규모 competition을 진행했다.
결과
response surface를 유연하게 모델링한 방법이 일관되게 가장 우수한 성능을 보였으며, 특히 비선형 response surface와 treatment-effect heterogeneity가 존재하는 경우에 그러했다.
시사점 및 한계
여러 readily available method가 causal effect를 정확하게 추정할 수 있으며, 특히 bias 감소가 주요 목표인 경우에 그러하다.
시사점 및 한계
공유 클러스터 사용량이 주차별로 달랐고 통제된 환경에서의 성능과 동등하지 않았기 때문에, 실행 시간 비교는 대략적이었다.
Abstract
from arXiv · showhide
Statisticians have made great progress in creating methods that reduce our reliance on parametric assumptions. However this explosion in research has resulted in a breadth of inferential strategies that both create opportunities for more reliable inference as well as complicate the choices that an applied researcher has to make and defend. Relatedly, researchers advocating for new methods typically compare their method to at best 2 or 3 other causal inference strategies and test using simulations that may or may not be designed to equally tease out flaws in all the competing methods. The causal inference data analysis challenge, "Is Your SATT Where It's At?", launched as part of the 2016 Atlantic Causal Inference Conference, sought to make progress with respect to both of these issues. The researchers creating the data testing grounds were distinct from the researchers submitting methods whose efficacy would be evaluated. Results from 30 competitors across the two versions of the competition (black box algorithms and do-it-yourself analyses) are presented along with post-hoc analyses that reveal information about the characteristics of causal inference strategies and settings that affect performance. The most consistent conclusion was that methods that flexibly model the response surface perform better overall than methods that fail to do so. Finally new methods are proposed that combine features of several of the top-performing submitted methods.
1. 서론 … “현실”에 맞춰 보정되지 않은 testing ground
Causal inference에는 다양한 방법이 존재하지만 선택은 어렵고, 기존 평가는 방법 간 선택에 불완전하거나 편향되었거나 잘 보정되지 않은 근거를 제공하는 경우가 많다. 이러한 한계는 Causal inference 전략을 비교하기 위한 더 광범위하고 공정한 testing ground의 필요성을 뒷받침한다.
- 1. 서론: Randomized experiment이나 natural experiment가 없다면 Causal inference는 실질적이고 쉽게 드러나지 않는 집단 차이에도 불구하고 공정한 treatment-control 비교를 요구한다.따라서 연구자들은 많은 pre-treatment covariate를 통제하려는 유인을 받으며, 이는 더 강한 parametric assumption을 요구하거나 다른 우려를 낳을 수 있다.
- 2. Causal inference competition의 동기: Causal-inference methodology의 폭넓음은 많은 선택지를 제공하는 동시에 applied researcher가 가장 유용한 접근법을 식별하기 어렵게 만든다.몇 가지 추가적인 문제는 방법론 선택을 더욱 복잡하게 만든다.
- 2.1 Causal inference 방법의 성능을 비교하는 기존 문헌의 한계: Method paper는 일반적으로 새로운 방법을 소수의 경쟁 방법과만 비교하며, 더 정교한 대안보다 전통적 접근법을 선호하는 경우가 많다.더 광범위한 비교조차도 제안된 방법의 저자들이 이를 더 잘 이해하고 경쟁 방법을 덜 효과적으로 구현하거나 평가할 수 있기 때문에 해당 방법에 유리할 수 있다.
- 비교 방법이 적고 공정하지 않은 비교.: 비교는 순진하거나 tuning되지 않은 경쟁 방법 구현과 bias를 강조하면서 root mean squared error나 interval coverage를 무시하는 evaluation metric 때문에 편향될 수 있다.이러한 선택은 의도치 않게 제안된 방법에 유리하게 작용할 수 있다.
- 비교 방법이 적고 공정하지 않은 비교.: Simulation study는 소수의 data-generating mechanism만 검토하는 경우가 많아 실제 적용을 반영하지 못할 수 있으며, 실제 적용에서는 변수가 continuous, categorical, binary 유형으로 섞여 있고 복잡한 joint distribution에서 생성된다.실제 observational dataset에서는 방법들이 서로 다른 결과를 내놓을 때 어느 방법이 우수한지 식별할 수 없다.
- “현실”에 맞춰 보정되지 않은 testing ground.: 고도로 특화된 simulation은 협력자의 실제 데이터를 모방할 수 있지만 일반 연구자에게는 제한적인 지침만 제공하며, asymptotic theory는 더 작은 sample size나 가정이 위배된 상황에 적용되지 않을 수 있다.이론적 regularity condition과 distributional assumption은 실제 dataset에서 성립하지 않을 수 있다.
- “현실”에 맞춰 보정되지 않은 testing ground.: 구성된 observational study는 또 다른 testing ground를 제공하지만, unknown ignorability 때문에 불충분한 복원이 model failure와 가정 위배 중 어느 쪽 때문인지 모호해진다.또한 비교가 두 개의 estimate를 포함하므로 observational estimate가 적절하다고 인정되려면 얼마나 가까워야 하는지도 불분명하다.
- “현실”에 맞춰 보정되지 않은 testing ground.: File-drawer effect는 결론이 나지 않은 비교가 출판될 가능성이 낮기 때문에 상대적 방법 성능에 대한 지식을 더욱 왜곡한다.따라서 연구자들은 중요한 부정적 또는 모호한 근거에 접근하지 못할 수 있다.
파일 서랍 효과. · 3. 표기와 가정
이 competition은 폭넓은 제출과 다양하고 현실적으로 보정된 simulation을 통해 제한적이고 잠재적으로 불공정한 method 비교 문제를 다뤘다. 이 논문은 potential outcome을 사용해 causal effect를 정의하고, identification과 estimation에 필요한 ignorability, overlap, admissible adjustment, conditional-expectation 가정을 제시한다.
- 2.2 우리 competition을 통한 이러한 한계 극복 시도: 이 competition은 해당 method에 정통한 연구자 또는 team이 제출한 30 methods를 평가해 적고 불공정한 비교 문제를 다루고자 했다.참가자들은 method를 직접 구현하거나, 여러 setting에서 작동하도록 고안된 black-box version을 제출했다.
- 2.2 우리 competition을 통한 이러한 한계 극복 시도: 이 competition은 다양한 field와 data-structure 관행을 아우르도록 설계되었으며, effect size, nonlinearity, interaction, covariate structure, bias level을 변화시켰다.개발자들은 이를 data feature의 복잡성을 고려해 observational study에서 cause의 effect를 추정하는 데 초점을 둔 최초의 competition이라고 설명한다.
- 2.2 우리 competition을 통한 이러한 한계 극복 시도: Simulation에는 실제 data에서 선택한 covariate를 사용해 그럴듯한 observational-study 상관을 모사했으며, testing ground가 현실을 반영하지 못할 수 있다는 우려에 대응했다.동기를 제공한 예시는 출생 체중이 IQ에 미치는 effect를 연구하는 것이었다.
- 파일 서랍 효과.: Simulation과 evaluation code를 GitHub에 공개하면 연구자들이 선호하는 metric으로 추가 method를 검증할 수 있어 file-drawer effect에 대한 해독제가 된다.이 대목은 public code를 이 잠재적 문제에 대응하는 핵심 mechanism으로 규정한다.
- 3. 표기와 가정: Binary treatment Z에서 potential outcome Yi(0)과 Yi(1)은 individual causal effect를 Yi(1) − Yi(0)으로 정의하며, observed outcome은 Zi에 대응하는 potential outcome을 결합한다.Z = 0은 control을, Z = 1은 treatment를 뜻한다.
- 3.1 Estimands: Average treatment effects는 Y(1) − Y(0)의 expectation이며, sample version은 analysis sample 또는 제한된 treatment 및 control population에 대해 평균을 낸다.이러한 변형은 편의상 선택했거나 관심을 두는 subpopulation의 causal effect를 형식화한다.
- 3.2 Structural Assumptions: 각 개인의 untreated 또는 treated potential outcome은 관측되지 않으므로 causal effect를 추정하려면 ignorability가 필요하며, observational study에서는 potential outcome의 독립성을 covariate X에 조건부로 둔다.이 가정 아래에서는 conditional observed outcome이 대응하는 conditional potential-outcome mean을 식별한다.
- 3.3 Parametric assumptions: Identification에는 overlap과 admissible pre-treatment back-door adjustment set X가 추가로 필요하고, unbiased estimation에는 E[Y(1) | X] 및 E[Y(0) | X]와 같은 conditional expectations의 modeling이 필요하다.Overlap이 없으면 일부 observation에는 경험적 counterfactual이 없으며, 강한 parametric assumption 없이 고차원에서 이러한 expectation을 추정하기는 어려울 수 있다.
4. TESTING GROUNDS: DATA AND GENERATIVE MODELS … 처치와 결과 시뮬레이션
이 competition에서는 ignorability를 부과하고 SATT를 추정 대상으로 삼으면서 causal inference methods를 구분할 수 있도록 현실적이고 조정 가능한 simulation을 사용했다. 데이터는 실제 연구에 맞춰 보정했으며, 비선형 response surface, treatment assignment, overlap, treatment effect의 특성을 변화시켰다.
- 4. TESTING GROUNDS: DATA AND GENERATIVE MODELS: testing grounds는 실제 연구 데이터의 전형적 특성을 보이면서 methods를 구분하도록 설계되었다.competition은 실용적인 설계 선택과 현실적인 데이터 특성을 결합했다.
- 4.1 내재된 가정과 설계 선택: 모든 data-generating process에는 competition을 실용적으로 유지하고 지나치게 복잡해지지 않도록 소수의 가정만 부과했다.
- Ignorability.: treatment assignment는 ignorable하다고 가정하여 treatment, outcomes, covariates를 연결하는 기저의 과학 이론을 고안하거나 대응시킬 필요를 없앴다.
- Estimand.: estimand는 treated 집단의 sample average treatment effect였는데, 많은 causal methods가 자연스럽게 이를 추정 대상으로 삼고 population effect에 대한 variance estimator가 없기 때문이다.
- 추론 집단에서의 overlap: SATT에서는 overlap를 위해 treated units에 대해서만 empirical counterfactuals가 필요했으므로, 충분한 overlap를 보이는 treated observations에 추론을 집중할 수 있었다.
- 4.2 ‘실제 데이터’에 대한 calibration: Collaborative Perinatal Project의 covariates는 현실 데이터에 맞춰 simulation을 보정할 수 있는 그럴듯한 변수 유형과 자연스러운 연관성을 제공했다.가상의 연구에서는 4,802개의 complete-case observations와 58개의 covariates를 사용해 출생 체중이 아동 IQ에 미치는 영향을 조사했다.
- 4.3 Simulation procedure와 “knobs”: simulation은 potential outcomes와 treatment assignment를 covariates에 조건부인 response surface와 assignment mechanism으로 분해하여 ignorability를 반영했다.두 구성 요소 모두 조정 가능한 transformations와 interactions를 포함한 generalized additive functions를 사용했다.
- 처치와 결과 시뮬레이션: 77개의 black-box scenarios에서 각각 100회 replication을 수행했으며, 비선형성, treatment prevalence, overlap, alignment, effect heterogeneity, treatment-effect magnitude를 변화시켰다.이 framework에는 비선형 response surfaces와 assignment mechanisms를 의도적으로 포함했는데, 이러한 비선형성이 존재하면 단순한 methods가 실패할 수 있기 때문이다.
비선형성의 정도. … 정렬.
시뮬레이션에서는 causal inference methods에 서로 다른 난제를 만들기 위해 treatment prevalence, overlap, imbalance, alignment, treatment-effect heterogeneity를 변화시켰다. 이러한 요인들은 잠재적 bias, extrapolation, variable prioritization, heterogeneous effects 추정의 복잡성에 영향을 미쳤다.
- 비선형성의 정도.: Treatment prevalence는 expected treated proportions가 35%와 65%인 settings 사이에서 변화했으며, treatment effect on the treated 추정에 잠재적 난제를 만들었다.Low-treatment setting에서는 시뮬레이션의 95%에서 treatment proportions가 0.20에서 0.38 사이였다.
- Treatment group의 overlap.: Low overlap은 covariate-space corner의 observations가 treatment를 받지 못하게 구성하여 propensity scores를 0으로 만들었고, common support를 넘어 extrapolate하는 models에 난제를 제기했다.Excluded neighborhood의 더 복잡한 정의는 이를 overlap이 있는 regions와 근본적으로 다르다고 식별하기 어렵게 만들었다.
- Treatment group의 overlap.: Overlap이 perfect할 수 있는 경우에도 simulations는 imbalance를 만들었으며, treated와 control means 사이의 Euclidean distance는 settings에 따라 달라졌고 quartiles는 [0.78, 1.30, 2.68]였다.Lack of overlap은 항상 lack of balance를 의미하지만, imbalance는 lack of overlap 없이도 발생할 수 있다.
- Treatment group의 overlap.: Treatment assignment와 response surface 모두에 영향을 미치는 covariates만 bias를 일으킬 수 있으며, covariate functional form도 중요하다.Assignment에만 또는 response에만 영향을 미치는 covariates는 포함하면 efficiency를 높일 수 있지만 bias에는 영향을 미치지 않아야 한다.
- 정렬.: 많은 available covariates 중 true confounders가 적을 때, treatment에만 또는 response에만 대한 predictors를 우선시하는 methods는 양쪽의 predictors를 target하는 methods보다 성능이 낮을 수 있다.따라서 Alignment는 bias의 가능성을 만들고 어떤 variables에 priority를 부여해야 하는지 결정하기 어렵게 한다.
- 정렬.: Alignment는 assignment와 response models 모두에 terms가 나타나는 빈도를 변화시켜 조절했으며, true confounders의 비율이 서로 다른 복잡한 settings를 가능하게 했다.그 결과 scenarios는 true assignment score p(Z | X)의 logit과 outcome Y 사이의 correlation에서 큰 폭의 변이를 보였다.
- 정렬.: Treatment-effect heterogeneity는 computational 및 statistical difficulty를 높였다. nonparallel response surfaces는 parallel ones보다 fit하기 어렵기 때문이며, normalized heterogeneity는 0에서 2.06까지였고 quartiles는 [0.47, 0.73, 1.01]였다.7700 realizations에서 median SATT는 0.68이었고 interquartile range는 0.57에서 0.79였으며, outcome-standard-deviation units로 나타냈다.
치료 효과의 전반적 크기. · 5. CAUSAL INFERENCE 제출 방법과 주요 특징 · 층화, 매칭, 가중치 부여.
이 competition은 방법을 구분하는 특징, 특히 parametric assumption을 줄이기 위해 데이터를 전처리하는 방식을 중심으로 다양한 causal-inference 제출 방법을 구성했다. 범위는 binary treatment, continuous response, IID data, 고정된 data dimension, 측정된 covariate, ignorability와 overlap으로 제한되었다.
- 4.4 다루지 않은 문제: 이 competition은 non-binary treatment, non-continuous response, non-IID data, 변하는 sample 또는 covariate dimension, covariate measurement error, ignorability와 overlap의 위반을 다루지 않았다.이러한 문제는 향후 competition의 잠재적 주제로 식별되었으며, 주최자들은 이 competition의 범위가 제한적이라고 설명했다.
- 5. CAUSAL INFERENCE 제출 방법과 주요 특징: 이 competition에는 15개의 DIY 제출 방법과 15개의 black-box 제출 방법이 접수되었지만, 두 개의 DIY 제출 방법은 설명이 불충분하여 제외되었다.주최자들은 방법을 제출할 수 없었으며, 단순한 main-effects linear model이 black-box baseline으로 포함되었다.
- 5. CAUSAL INFERENCE 제출 방법과 주요 특징: 이 competition은 causal-inference method를 구분하는 특징의 taxonomy를 사용해 서로 상당히 다른 제출 접근법을 분류했다.이 taxonomy는 Table 1에 요약되었고, 추가 세부 사항은 Appendix A.2에 제시되었다.
- 층화, 매칭, 가중치 부여.: treatment와 control의 covariate를 balance하도록 데이터를 전처리하는 방법은 model-free estimation을 지원하거나, model-based estimate가 misspecification에 더 robust하도록 만들 수 있다.Stratification, matching, weighting은 비교 가능한 group 또는 pseudo-population을 구성하여 이러한 balance를 추구한다.
- 층화, 매칭, 가중치 부여.: Stratification은 covariate로 정의된 cell 내에서 treated와 control의 outcome을 비교하며, regression-tree leaf를 사용하는 변형도 있다.이 접근법은 subclassification이라고도 한다.
- 층화, 매칭, 가중치 부여.: Matching은 distance metric에 따라 treated unit에 가장 가까운 control을 선택하고, 유사성이 충분하지 않다고 판단된 control을 제외하며, 가장 일반적으로 propensity score를 사용한다.다른 distance metric도 사용할 수 있다.
- 층화, 매칭, 가중치 부여.: Weighting은 ATT를 추정할 때 covariate distribution이 inferential group의 distribution과 유사한 pseudo-population을 만들도록 control에 다시 가중치를 부여하며, 다른 target에서는 역할을 반대로 한다.이 접근법은 survey-sampling weighting과 밀접하게 관련된다.
- 층화, 매칭, 가중치 부여.: parametric reliance를 줄이는 방법은 propensity score를 balancing score로 포함하기 때문에 정확한 treatment-assignment modeling이 필요한 경우가 많다.동일한 propensity score를 조건으로 하면 treatment assignment는 ignorable하다.
할당 메커니즘 모델링. … Ensemble methods.
이 논문은 treatment assignment 또는 response surface를 모델링하는지에 따라 causal inference methods를 분류하고, flexibility와 robustness를 높이는 전략으로 variable selection과 ensembles를 강조한다. 제출 방법 전반에서 response-surface modeling과 정교한 nonparametric methods가 특히 두드러졌다.
- 할당 메커니즘 모델링.: Propensity scores는 stratification, matching, weighting 또는 TMLE를 통해 causal analyses에 포함될 수 있으며, balanced groups는 ignorability 하에서 unbiased treatment-effect estimation을 가능하게 한다.이에 대응하는 전략은 관측치를 balance를 위해 전처리하는 대신 response surface를 올바르게 모델링하는 것이다.
- Response surface 모델링.: Methods는 assignment mechanism 또는 response surface를 모델링하는지에 따라 분류되는데, formal double robustness를 판정하는 일은 이 논문의 범위를 벗어나기 때문이다.이 taxonomy는 모든 접근법에 formal double-robustness classification을 부여하지 않고 주요 modeling targets를 구분한다.
- Assignment mechanism 또는 response surface의 nonparametric modeling.: Flexible assignment-mechanism modeling은 역사적으로 충분히 활용되지 않았는데, propensity-score estimation이 주로 covariate balance를 달성하기 위한 도구로 간주되었기 때문이다.Nonparametric response-surface modeling은 response-surface misspecification에 대비하기 위해 고안된 balancing approaches의 대안으로 등장했다.
- Variable selection.: LASSO 와 elastic net 같은 Variable-selection methods는 후보 covariates가 많을 때 true confounders에 estimation을 집중할 수 있다.Estimation problem을 축소하면 중요한 변수에 더 복잡한 algorithms를 적용하기 쉬워진다.
- Ensemble methods.: Ensemble methods는 여러 methods를 적합하고, cross-validation 또는 model averaging으로 상대적 performance를 평가하며, estimates를 선택하거나 결합함으로써 settings 간 variation에 대응한다.Combined estimates는 relative performance를 반영하는 weights를 사용한 weighted averages이다.
- 5.2 제출 방법 개요.: 제출 방법 중 대부분은 response-surface models를 적합했고, 절반 이상이 weighting을 사용했으며, matching은 black-box methods에서 거의 없었고, black-box 참가자들은 정교한 nonparametric techniques를 선호했다.Table 1은 제출된 methods를 요약하고 do-it-yourself approaches와 black-box approaches를 구분한다.
- 5.3 최고 성능 방법.: BART는 작은 regression trees의 합으로 arbitrary functions를 적합하고, overfitting을 피하기 위해 priors를 사용하며, 두 potential outcomes에 대해 joint function f(x, z)를 모델링한다.Causal BART procedure는 y(1) = f(x, 1)과 y(0) = f(x, 0)의 posterior predictive distributions에서 추출한다 (Chipman, George and McCulloch, 2010).
Bayesian Additive Regression Trees (BART). · Super Learner plus Targeted Maximum Likelihood Estimation (SL+TMLE). · DR w/GBM + MDIA 1 및 2
이 방법들은 유연한 response-surface modeling을 ensemble learning, propensity-score adjustment, TMLE, calibrated weighting과 결합한다. 대회 후 변형에서는 BART, joint 대 separate outcome modeling, 추가 assignment-mechanism modeling, 대안적 BART fitting strategies를 검증했다.
- Super Learner plus Targeted Maximum Likelihood Estimation (SL+TMLE).: SL+TMLE는 cross-validated Super Learner predictions를 사용하고, squared-error-minimizing weights로 library fits를 결합한 뒤 TMLE correction을 적용했다.그 library에는 assignment와 response surfaces를 modeling하기 위한 glm, gbm, gam, glmnet, splines가 포함됐다.
- Super Learner plus Targeted Maximum Likelihood Estimation (SL+TMLE).: SL+TMLE ensemble은 assignment mechanism과 control response surface를 별도로 모델링하고, propensity-score weights를 통합해 ATT를 individual conditional treatment effects로 확장했다.구현에는 glm, random forest, deep learning, LASSO, ridge regression을 포함한 ensemble libraries도 사용됐다.
- Bayesian Additive Regression Trees (BART).: calCause ensemble은 out-of-sample prediction으로 random forests와 Gaussian processes 중에서 선택한 다음, treatment-effect estimation을 위해 treated units’ control responses를 impute했다.이 방법은 control response surface fitting에 초점을 맞췄다.
- DR w/GBM + MDIA 1 및 2: DR w/GBM + MDIA methods는 generalized boosted regression으로 assignment와 response surfaces를 추정하고, 최대 three-way interactions를 허용했으며, MDIA를 사용해 control treatment-on-treated weights를 calibrate했다.MDIA는 base weights를 최소한으로 perturb하면서 covariate means와 estimated response values를 정확히 balance했다.
- Bayesian Additive Regression Trees (BART).: 대회 후 analyses에서는 Super Learner library에 BART를 추가하고 BART IPTW를 만들어 assignment mechanism modeling이 response-surface modeling만 사용하는 것보다 개선되는지 검증했다.성능이 가장 높았던 두 ensemble submissions는 두 mechanism을 모두 modeling했다.
- Super Learner plus Targeted Maximum Likelihood Estimation (SL+TMLE).: Super Learner에서 TMLE를 제거하고 BART에 IPTW plus TMLE를 추가해 targeted correction과 propensity-score modeling의 기여를 분리했다.BART propensity score는 cross-validation으로 fit했다.
- Bayesian Additive Regression Trees (BART).: BART는 주요 stand-alone flexible response-surface approach를 제공했으며, ensemble performance에 rival하는 것으로 보고된 유일한 stand-alone method였다.대회 후 변형에는 cross-validated hyperparameters와 multiple chains가 포함됐다.
- Super Learner plus Targeted Maximum Likelihood Estimation (SL+TMLE).: BART는 treatment와 response surfaces를 jointly modeled한 반면 original Super Learner는 treatment- 및 control-condition models를 separate하게 fit했기 때문에, joint-response-surface SL+TMLE variant가 만들어졌다.joint implementation에는 ensemble에 BART가 포함됐다.
6. 제출 방법과 구성 방법의 성능 평가 · RMSE와 bias
20개 DIY 데이터 세트와 7,700개 black-box 데이터 세트에서 RMSE, bias, interval coverage, interval length, PEHE를 사용해 성능을 평가했다. Flexible method는 대체로 RMSE와 bias에서 가장 우수했지만, interval coverage와 length는 크게 달랐다.
- 6. 제출 방법과 구성 방법의 성능 평가: Global performance summary를 사용해 20개 DIY 데이터 세트와 7,700개 black-box method에서 제출 방법과 competition 후 방법을 비교했다.RMSE는 treatment-effect estimate가 true estimand에 평균적으로 얼마나 가까운지를 측정했고, bias는 SATT로부터의 평균 거리를 측정했다.
- 6.1 20개 DIY 데이터 세트에서 모든 method 비교: 20개 DIY 데이터 세트에서 두 DR w/GBM+MDIA 제출 방법이 bias와 RMSE에서 가장 우수했고, 그 뒤를 상위 black-box performer들이 이었다.DIY 제출 방법은 20개 데이터 세트에서만 평가되었으므로, 저자들은 상대적 성능에 대해 강한 결론을 내리는 데 신중했다.
- 6.2 Black-box 제출 방법 비교: 7,700개 데이터 세트에서 BART, SL + TMLE, calCause, h2o가 bias와 RMSE에 대해 최초 제출 method 중 가장 우수했다.Tree Strat, BalanceBoost, Adjusted Tree Strat, LASSO + CBPS가 그다음으로 우수한 group을 구성했다.
- RMSE와 bias: Treatment-effect distribution의 expected value가 양수이고 estimate가 대체로 0을 향해 shrink되었기 때문에 모든 method의 average bias는 negative였다.Black-box method의 보고된 bias summary에는 77개 setting과 100회 replication에 걸친 interquartile range도 포함되었다.
- RMSE와 bias: SL+TMLE library에 BART를 추가하면 BART가 없는 SL+TMLE보다 성능이 향상되었지만, BART 단독을 능가하지는 못했다.BART+TMLE, BART MChains, BART Xval은 bias와 RMSE를 소폭 개선했지만, BART IPTW는 개선하지 못했다.
- RMSE와 bias: 많은 automated algorithm이 RMSE와 bias에서 좋은 성능을 보였지만, interval coverage와 average interval length는 크게 달랐고 최초 제출 method는 다소 실망스러웠다.Coverage는 interval이 true SATT를 포함한 데이터 세트의 비율을 측정했고, interval length는 coverage와 precision 사이의 trade-off를 포착했다.
구간 coverage와 길이. … 7. PERFORMANCE 예측
방법 전반에서 유연한 response-surface modeling이 대체로 가장 우수했지만, 관찰된 데이터나 simulation setting만으로 상대적 성능을 예측하기는 어려웠다. BART augmentation으로 coverage를 크게 개선할 수 있었지만, 그 대가로 구간이 더 길어지고 방법별 성능 차이의 해석 가능성이 제한되었다.
- 구간 coverage와 길이.: 구간 길이 표시를 coverage로 해석해서는 안 된다: 95% 선 위의 삼각형은 구간 길이를 나타내므로 오해를 불러일으킬 수 있다.해당 부분은 구간 길이와 coverage를 명시적으로 구분한다.
- 구간 coverage와 길이.: BART + TMLE은 거의 nominal coverage를 달성한 반면 BART 단독은 약 82%였고, 평균 구간은 약 50% 더 길었다.augmentation을 적용한 구간은 다른 상위 성능 방법의 구간보다 약간 짧았다.
- 이질적 효과 추정의 정밀도.: 개별 treatment effect를 보고한 방법 중에서는 BART와 calCause가 다른 더 단순한 linear-model 옵션보다 PEHE에서 눈에 띄게 우수했다.PEHE를 계산하는 데 필요한 추정치를 산출한 방법은 일부에 불과했다.
- 계산 시간.: h2o는 24.8 seconds가 걸린 반면 BART는 29.4 seconds가 걸렸지만, h2o는 자주 실패하고 재시작이 필요했으며 계산을 외부로 넘겼기 때문에 시간 측정의 신뢰성이 낮았다.클러스터를 공유하고 background process가 자원 측정을 복잡하게 만들어 실행마다 computational usage가 달라졌다.
- 7. PERFORMANCE 예측: 상대적 성능은 대체로 context에 의존하지 않았다: 평균 성능을 넘어서는 수준에서, 여러 setting에 걸쳐 어떤 black-box method가 다른 방법보다 우수할지는 거의 예측할 수 없었다.가장 강력한 일반적 권고는 유연한 non-parametric response-surface modeling을 사용하는 것이었다.
- 7.2 설명된 성능 분산: 비오라클 측정치만 사용하는 예측 모델은 R2 = 0.10을 넘는 경우가 드물었지만, 오라클 측정치는 방법의 삼분의 일을 약간 웃도는 수준에서 0.40–0.50에 도달했으며, 이들은 대부분 성능이 낮은 방법이었다.더 성공적인 방법에서는 오라클을 포함한 R2가 0.10을 넘는 경우가 드물었다.
- 7. PERFORMANCE 예측: 유연한 non-parametric response-surface modeling은 방법 간 성능 차이의 대부분을 설명했지만, dataset 수준 변이의 절반 이상은 설명되지 않은 채 남았다.방법 특성은 방법 간 평균 성능 차이의 76%를 설명했지만, dataset 간 설명되지 않은 변이는 전체 변이의 절반을 넘었다.
- 7.3 데이터 및 모델 특성에 관한 방법 간 분석: setting은 전체 성능 변이의 5%만 설명했으며, setting 간 평균 차이는 response-surface nonlinearity와 assignment-response alignment로 설명되었다.두 predictor 모두 non-oracle data feature였다.
8. 논의
이 competition은 복잡한 observational-study 설정 전반에서 유연한 ensemble 방법이 우수한 성능을 보일 수 있음을 밝혔지만, bias가 낮은 경우에도 coverage는 여전히 어려웠다. 정확하고 쉽게 이용할 수 있는 방법이 여러 가지 존재하지만, 결론은 ignorability, overlap, i.i.d. data를 갖춘 testing ground로 제한된다.
- 8. 논의: 유연한 ensemble 방법은 data의 복잡성에 대응하기 위해 여러 model의 상대적 강점을 결합함으로써 강한 성능을 보였다.competition 결과는 대부분의 ensemble 방법이 우수한 성능을 보인 이유로 flexibility를 강조했다.
- 8. 논의: assignment mechanism과 response surface 사이의 불일치는 가장 까다로운 data 특성 중 하나였으며 causal-inference 문헌에서는 거의 논의되지 않았다.이러한 어려움은 이용 가능한 covariate가 true confounder 또는 관련 transformation의 일부만 포함할 때 발생한다.
- 8. 논의: bias가 낮은 경우에도 대부분의 방법에서 양호한 coverage를 얻기 어려웠으며, TMLE adjustment가 coverage를 일관되게 개선하지는 못했다.저자들은 coverage를 최적화하기 위한 강한 조언을 제시하지 않았다.
- 8. 논의: 여러 방법이 causal effect를 정확하게 추정하며, 특히 bias 감소가 주요 목표일 때 그러하고, 많은 방법에 readily available한 R package가 존재한다.이 결론은 inferential group에 대해 ignorability, adequate overlap, i.i.d. data를 갖춘 설정에만 적용되며, 이러한 조건은 실제로 성립하지 않을 수 있다.
- 8. 논의: 이 competition은 crowdsourced implementation을 사용하여 전형적인 methodological paper보다 더 많은 assignment-mechanism 및 response-surface data-generating process에 걸쳐 폭넓은 방법을 평가했다.저자들은 이 접근법이 다른 문제를 다루는 유사한 competition을 촉진하기를 기대한다.
부록 A: 부록 절 · A.1 Simulation Framework 세부사항 · A.2 제출 및 생성 방법 용어집
부록 A는 77개 설정으로 구성된 simulation framework를 상세히 설명하고, 제출 방법과 주최자 생성 방법을 기록한다. 이 framework는 treatment 및 response-model 구조를 변화시키며, 용어집에는 propensity weighting, flexible response modeling, matching, regression, targeted estimation을 아우르는 방법들이 수록되어 있다.
- A.1 Simulation Framework 세부사항: Simulation framework는 77개 설정으로 구성되며, 열거된 parameter 선택으로 각 black-box 및 DIY dataset을 재현하는 R package가 제공된다.부록은 조작된 각 simulation knob의 수준을 명시하고 재현 package를 제공한다.
- A.1 Simulation Framework 세부사항: Treatment-model complexity는 linear 및 polynomial term부터 assignment mechanism P(Z = 1 | X)에서 jump와 kink가 나타나는 step function까지 다양하다.이 대안들은 treatment assignment mechanism을 생성하는 데 사용되는 기본 function library를 정의한다.
- A.1 Simulation Framework 세부사항: 생성된 coefficient와 sub-function 위치는 그럴듯한 outcome과 propensity score를 산출하도록 scaling되며, coefficient는 Student-t 또는 beta-prime distribution에서 추출된다.Covariate는 대략 [−1, 1]로 scaling되고, functional term을 선택한 뒤 결합된 function을 다시 rescaling한다.
- A.2 제출 및 생성 방법 용어집: 용어집은 모든 competition 제출 방법과 주최자 생성 방법을 다루며, 가능한 경우 기여자가 제공한 설명을 사용하고 설명이 없는 방법은 제외한다.방법들은 do-it-yourself 접근법과 black-box 접근법에 대해 별도의 table로 정리되어 있다.
- A.2 제출 및 생성 방법 용어집: 문서화된 여러 방법은 flexible outcome modeling을 weighting 또는 targeting과 명시적으로 결합하며, GBM-plus-MDIA, BART-IPTW, Super Learner/TMLE 접근법이 이에 해당한다.이 설명들은 estimator의 구성요소로 cross-validation, propensity-score weighting, covariate balancing 또는 targeted correction을 명시한다.
- A.2 제출 및 생성 방법 용어집: Do-it-yourself 방법에는 propensity-score weighting, matching, stratification, regression adjustment, generalized additive model, neural network, Gaussian process가 포함된다.예를 들어 variable selection을 GenMatch 및 GAM과 결합하거나, tree 또는 neural network로 propensity score를 추정하거나, weighted Gaussian process로 covariate shifting을 다룬다.
- A.2 제출 및 생성 방법 용어집: Black-box 방법은 BART, cross-validated response-surface model, covariate-balancing propensity score, ensemble learner, LASSO, TMLE 또는 Super Learner 조합을 아우른다.용어집에는 conventional linear model, Stata treatment-effect estimator, tree-based stratification 변형도 포함된다.
A.3 추가 DIY 결과
보충 DIY 결과는 competition의 데이터셋 전반에서 이질적 효과를 추정할 때 coverage, 구간 길이, precision을 평가한다. Figures 4 and 5는 DIY 및 black-box methods에 대한 이러한 비교를 보고한다.
- A.3 추가 DIY 결과: 보충 DIY 분석은 이질적 효과를 추정할 때의 coverage, average interval length, PEHE를 검토한다.이 결과는 Figures 4 and 5에 보고된다.
- A.3 추가 DIY 결과: Figure 4는 20개 DIY data sets에서 모든 DIY 및 original black-box methods의 coverage와 average interval length를 비교한다.Methods는 coverage가 감소하는 순서로 배열되어 있으며, 회색 점은 plotting region을 벗어난 매우 낮은 coverage 또는 매우 긴 intervals를 나타낸다.
- A.3 추가 DIY 결과: Figure 5는 individual-level treatment-effect estimates를 제공한 DIY 및 black-box methods의 PEHE를 제시한다.
A.4 분산 설명: modeling 결과
modeling은 log absolute bias의 method 수준 및 setting 수준 변동을 상당 부분 설명하며, 특히 유연한 response-surface feature와 관측 가능한 setting metric을 통해 그러하다. Trial 수준 realization은 여전히 설명하기 어렵고, observables는 method 선택보다 맥락적 성능을 더 쉽게 예측한다.
- 다층 variance decomposition: Methods, settings, 그리고 이들의 interaction은 1.135 variance units, 즉 전체 변동의 46%를 설명했으며, 나머지 54%는 trial 수준 realization이 설명했다.Realization component는 trial 내 특이적 오차와 predictor 변동성을 반영한다.
- 다층 variance decomposition: Non-parametric response surfaces는 가장 중요한 method feature였으며, log scale에서 결과를 factor of -2만큼 크게 개선했다.Method indicator는 between-method main-effect variation의 76%를 설명했지만, 유의한 method feature는 거의 없었다.
- 다층 variance decomposition: Non-oracle metric은 between-setting main-effect variation의 81%를 설명했지만, 제한적인 setting-by-method interaction은 이 metric이 method 선택에 little guidance for choosing a method만 제공함을 시사한다.관측 가능한 quantity는 method 선택에 큰 영향을 주지 않으면서 research setting에 관한 상당한 정보를 제공할 수 있다.
- 다층 variance decomposition: Method-feature interaction은 setting-by-method variance component의 45%를 설명했지만, idiosyncratic realization variance에서는 little progress만 보였다.이 interaction은 setting-by-method variation을 겨냥했지만 realization variance도 설명할 수 있었다.
- Model-specific explanatory power: Full setting-indicator 및 metric model은 method 전반에서 0.06 to 0.50 범위의 R2 값을 산출했으며, between-method variation을 제외하면 전체 변동의 about 18%를 설명한다는 결과와 일치했다.분산 계산에서는 method 전체 평균으로 1.53 units 중 1.26이 설명되지 않은 채 남는다.
A.5 실험 설정을 설명하는 데 사용한 metric 전체 목록 · treatment effect heterogeneity. · A.6 제출 방법 및 감사의 말
이 논문은 비선형성, overlap, balance, alignment, treatment-effect heterogeneity를 포함해 실험 설정을 특성화하기 위한 oracle metric과 knob metric을 정의한다. 또한 제출자들에게 감사를 표하고, 제출된 방법이 반드시 best practice를 지지하는 것은 아니라고 밝힌다.
- A.5 실험 설정을 설명하는 데 사용한 metric 전체 목록: 이 competition은 일반적인 관찰연구에서는 이용할 수 없는 oracle metric과, 명시적으로 생성된 실험 설정을 나타내는 knob metric을 구분한다.이러한 label은 어떤 설정 특성이 data creator만 알고 있는 것인지, 어떤 특성이 experimental control에 해당하는지를 명확히 한다.
- A.5 실험 설정을 설명하는 데 사용한 metric 전체 목록: 실험 설정 metric에는 outcome과 treatment-assignment의 비선형성이 포함되며, 각각 linear, nonlinear, step-function model에 대해 0, 1, or 2로 점수화된다.관찰 data에서 쉽게 추정할 수 있음에도, treated 비율도 설정 metric으로 추적한다.
- treatment effect heterogeneity.: oracle metric은 propensity score와 treated unit 및 control unit 사이의 distance를 사용해 alignment, overlap, and balance를 평가한다.여기에는 propensity-score–outcome correlation, Mahalanobis nearest-neighbor distance, mean-design-matrix distance, Wasserstein distance가 포함된다.
- treatment effect heterogeneity.: 추가 metric은 true 또는 observable design information에 기반한 R2 value를 통해 treatment-assignment nonlinearity, treatment-effect heterogeneity, outcome predictability를 정량화한다.열거된 regression에는 true propensity score, true treatment effect, 그리고 design matrix에 대한 outcome regression이 포함된다.
- A.6 제출 방법 및 감사의 말: 감사의 말에서는 방법을 제출한 사람들에게 감사를 표하고, affiliation은 제출 당시 각 first author의 affiliation으로 보고한다.방법은 특정 순서 없이 나열된다.
- A.6 제출 방법 및 감사의 말: 일부 제출 방법은 해당 field를 대표하도록 구성되었으므로, 제출자 자신의 best practice에 대한 신념을 반영하지 may not reflect.이 단서는 제출 방법을 recommendation으로 해석하는 방식을 제한한다.
Do-It-Yourself Methods · Black Box Methods · calCause
이 논문은 weighting, regression, boosting, tree-based, ensemble, targeted-learning 접근법을 아우르는 do-it-yourself, black box, calCause 제출 방법을 정리한다. 나열된 방법은 대학, 정부 연구센터, 민간 기관의 연구자들이 기여했다.
- Do-It-Yourself Methods: Do-it-yourself 제출 방법에는 IPTW, Bayes LM, regression trees, calibrated IPW, DR w/GBM + MDIA와 Ad Hoc, LAS Gen GAM, weighted GP, GLM-Boost, manual RBD, TwoStepLM, ProxMatch, VarSel NN이 포함됐다.기여자는 Harvard, MIT, Mount Sinai, Helmholtz Zentrum München, the University of Florida, Seoul National University, the University of Alberta, Acumen, UC Berkeley, Columbia University의 연구자들이었다.
- Do-It-Yourself Methods: Do-it-yourself 목록은 inverse-probability weighting과 doubly robust estimators를 Bayesian linear models, generalized additive models, Gaussian processes, boosting, matching, neural-network variable selection과 결합했다.이 항목에는 IPTW, Bayes LM, calibrated IPW, DR w/GBM + MDIA, LAS Gen GAM, weighted GP, GLM-Boost, ProxMatch, VarSel NN이 포함됐다.
- Black Box Methods: Black box 제출 방법에는 teffects methods와 LASSO+CBPS가 포함됐다.나열된 기여자들은 the University of Maryland, the University of Wisconsin-Milwaukee, the University of Mississippi Medical Center, Brandeis University에 소속되어 있었다.
- calCause: calCause 제출 방법에는 BalanceBoost, Tree Strat and Adj. Tree Strat, h2o Ensemble, CBPS, SL+TMLE이 포함됐다.나열된 기여자는 Stanford University와 the University of Wisconsin Madison의 연구자들이었다.
- calCause: calCause 목록은 boosting, tree stratification, ensemble learning, covariate balancing, targeted maximum likelihood estimation을 포괄했다.이 접근법은 각각 BalanceBoost, Tree Strat and Adj. Tree Strat, h2o Ensemble, CBPS, SL+TMLE에 해당한다.