Source-linked AI summary
Google COVID-19 Community Mobility Reports: Anonymization Process Description (version 1.1)
Ahmet Aktay, Shailesh Bavadekar, Gwen Cossoul, John Davis, Damien Desfontaines, Alex Fabrikant, Evgeniy Gabrilovich, Krishna Gadepalli, Bryant Gipson, Miguel Guevara, Chaitanya Kamath, Mansi Kansal, Ali Lange, Chinmoy Mandayam, Andrew Oplinger, Christopher Pluntke, Thomas Roessler, Arran Schlosberg, Tomer Shekel, Swapnil Vispute, Mia Vu, Gregory Wellenius, Brian Williams, Royce J Wilson
TL;DR
개인 위치, 이동 또는 접촉 데이터가 도출되지 않도록 하면서 이동성 변화를 공개하는 방법을 다룬다. 옵트인한 사용자의 Location History 데이터를 differential privacy와 함께 집계하고, 과거 baseline 대비 신뢰할 수 있는 percentage changes를 공개하며, formal privacy guarantees를 제공한다.
문제
Location History 데이터에서 개인의 위치, 이동 또는 접촉을 노출하지 않고 이동성 변화를 공개하는 방법을 다룬다.
방법
옵트인한 사용자의 metric을 집계하고, differential privacy를 위해 Laplace noise를 추가하며, baseline percentage changes를 계산하고, 신뢰성이 낮거나 충분한 근거가 없는 metric을 제외한다.
결과
ε = 1.76의 일일 최대 사용자 기여량과 δ = 0은 공개된 metric set에 대해 명시된 differential privacy guarantee를 제공한다.
시사점 및 한계
이 과정은 private historical baseline과 mobility metric을 공개적으로 비교할 수 있게 하면서, 불확실성이 reporting threshold를 초과하는 변화는 공개하지 않는다.
시사점 및 한계
percentage change가 ±10 absolute percentage points를 초과해 잘못될 확률이 최소 5%가 되는 경우 metric을 공개하지 않는다.
Abstract
from arXiv · showhide
This document describes the aggregation and anonymization process applied to the initial version of Google COVID-19 Community Mobility Reports (published at http://google.com/covid19/mobility on April 2, 2020), a publicly available resource intended to help public health authorities understand what has changed in response to work-from-home, shelter-in-place, and other recommended policies aimed at flattening the curve of the COVID-19 pandemic. Our anonymization process is designed to ensure that no personal data, including an individual's location, movement, or contacts, can be derived from the resulting metrics. The high-level description of the procedure is as follows: we first generate a set of anonymized metrics from the data of Google users who opted in to Location History. Then, we compute percentage changes of these metrics from a baseline based on the historical part of the anonymized metrics. We then discard a subset which does not meet our bar for statistical reliability, and release the rest publicly in a format that compares the result to the private baseline.
1 정의
보고서는 기본값이 꺼져 있는 Location History에 참여하도록 선택한 사용자의 데이터를 사용한다. Metric은 세 가지 geographic granularity level에 걸쳐 매일 집계되며, 상위 level일수록 더 작은 영역을 나타내고 3 km^2 미만의 공개 지역은 없다.
- Location History 사용자: Metric은 기본값이 꺼져 있는 기능인 Location History에 참여하도록 선택한 Google 사용자를 기반으로 한다.
- Differential Privacy: Differential privacy는 randomized metric algorithm을 사용해 하루 동안 한 사용자의 데이터가 다른 데이터셋 간의 차이를 기준으로 정의된다.
- Granularity level: Metric은 세 가지 granularity level에 걸쳐 날짜와 geographic area별로 집계된다: 국가 또는 지역, 최상위 지정학적 하위 구역, 미국의 county와 같은 고해상도 영역이다.
- Granularity level: Granularity level 1과 2는 지역 공중보건 수요를 반영하도록 국가별로 달라지며, 일반적으로 상위 level일수록 더 작은 영역을 나타내고 3 km^2보다 작은 지역은 제외한다.
2 익명화된 metric 생성
metric은 differential privacy를 사용해 집계된 Location History 데이터에서 생성되며, 개인의 위치·이동·접촉 데이터를 추론하지 못하도록 Laplace noise와 contribution limit를 적용한다. 공공장소·주거지·직장 metric에는 사용자 contribution 상한과 명시적 privacy budget을 적용한다.
- Privacy framework: Differential privacy는 집계된 metric에 Laplace noise를 추가해 공개된 데이터에서 개인의 위치·이동·접촉을 추론하지 못하도록 보호한다.절차에는 Google의 오픈소스 differential privacy library 와 Laplace noise 를 사용한다.
- 공공장소 metric: 공공장소 metric은 7개 범주를 매일 방문한 고유 Location History 사용자 수를 세며, 각 사용자는 범주별로 최대 한 번, 그리고 날짜와 geographic level별로 최대 네 개의 category-location pair에만 기여한다.범주는 retail, recreation, eateries, groceries, pharmacies, transit, parks이며, 네 개를 초과하는 기여는 무작위로 줄인다.
- 공공장소 metric: 미국 county level 사용자의 99%는 평균적으로 하루에 category-place pair를 세 개 이하로 contribution하며, 각 일일 장소 방문에는 ε = 0.44, 사용자의 총 일일 contribution에는 ε = 1.76을 적용한다.이 privacy guarantee는 여러 metric이 동일한 dataset을 공유하므로 standard composition을 사용한다.
- 주거지 metric: 주거지 metric은 사용자별 시간 합계와 고유 사용자 수를 bounded한 뒤 noise를 추가하고, 조정된 비율을 계산해 범위를 제한함으로써 주거지에서의 daily average hours를 추정한다.개별 주거지 체류 시간 값은 [−12; 12]로 offset하고, 최종 추정치는 [0, 24] hours/day로 제한한다.
- 직장 metric: 직장 metric은 geographic area별로 하루에 직장 위치에서 more than one hour를 보낸 사용자를 집계하고, 각 사용자를 granularity별 최대 하나의 region으로 집계하면서 Laplace noise를 추가한다.이 metric은 ε = 0.44인 differential privacy로 보호한다.
3 익명화된 metric에서 보고서 생성
보고서는 일별 anonymized metric을 생성하고, 이를 요일별 baseline 대비 percentage change로 변환한 뒤, 지리적·사용자 수·통계적 신뢰성이 충분하지 않은 결과를 억제한다.
- 3 익명화된 metric에서 보고서 생성: 일별 metric은 2020-01-01부터 생성되며, 이후 모든 연산은 추가 privacy budget을 소모하지 않고 differential privacy output만 사용한다.
- 3 익명화된 metric에서 보고서 생성: 3km^2보다 작거나 differential privacy contributing user가 100명보다 적은 geographic region은 폐기한다. 단, 소규모 region은 국가 경계를 넘지 않는 범위에서 3km^2를 초과하도록 병합할 수 있으며, Vatican City–Italy는 예외다.
- 3 익명화된 metric에서 보고서 생성: 각 날짜의 metric은 2020-01-03부터 2020-02-06까지 고정된 5주 기간에 포함된 동일 요일의 5개 관측값에 대한 differential privacy metric의 median으로 정의한 baseline 대비 percentage ratio로 보고한다.
- 3 익명화된 metric에서 보고서 생성: differential privacy noise로 인해 오류가 ±10 absolute percentage points를 초과할 가능성이 있을 때, 구체적으로 해당 오류가 5% chance 이상일 때 percentage change를 공개하지 않는다.이 결정에는 metric과 baseline의 97.5% confidence interval을 사용한다. 그 결과 얻은 ratio bound가 private ratio와 10 absolute percentage points를 초과하여 다르면 해당 change를 공개하지 않는다.
4 δ에 관한 주석
이 과정은 고정된 metric 집합을 생성하고 값이 0인 metric에 noise를 추가하므로 δ = 0인 ε-differential privacy를 만족한다.
- 4 δ에 관한 주석: 이 절차는 geographic region, 기간 내 day, public place category의 모든 조합을 생성하고 값이 0인 metric에 noise를 추가함으로써 δ = 0인 ε-differential privacy를 달성한다.
5 시간 경과에 따른 metric 정확도 개선
metric computation은 지속적으로 개선되지만, update로 인해 값이 변하고 변경되지 않은 baseline과의 비교가 왜곡될 수 있다. Scaling factor는 metric을 그룹화하고 기존에 생성된 noisy metric을 재사용해 이러한 변화를 보정하면서 추가적인 privacy budget 사용을 제한한다.
- Scaling factor는 computation update로 인한 변화를 보정해, 그렇지 않으면 변경되지 않은 baseline period와의 비교를 왜곡할 수 있는 영향을 제거한다.데이터를 다시 공개하지 않기 위해 baseline은 재계산하지 않는다.
- metric을 grouped하여 각 group이 동일한 update effect를 공유하도록 한다. 예를 들어 한 period 내 날짜별 또는 요일별 metric이 이에 해당한다.예로 유월의 화요일별 Parks 또는 팔월의 주말별 Workplaces를 들 수 있다.
- 각 group에 대해 method는 previously generated noisy metric을 s_g로, smaller-budget noise를 적용해 recomputed한 metric을 s_n으로 합산한다.추가 noise는 일반적으로 corresponding region granularity에 비례하는 privacy budget의 10%를 사용한다.
- baseline을 scaling할 때는 future metric에 s_n/s_g를 곱하고, daily count를 scaling할 때는 s_g/s_n을 곱한다.Grouping을 통해 더 작은 step-3 privacy budget을 사용하고 이미 생성된 metric을 재사용할 수 있다.