Source-linked AI summary

Identifying Participants in the Personal Genome Project by Name (A Re-identification Experiment)

Latanya Sweeney, Akua Abu, Julia Winn

arXiv:1304.7605v1cs.CY

TL;DR

The paper examines whether publicly displayed PGP profiles that omit names can nevertheless be re-identified, a question that matters because the profiles contain sensitive medical, genomic, and demographic information. It links profile demographics to public records and extracts names embedded in attached documents, correctly identifying 84 to 97 percent of profiles for which names were provided. The authors conclude that demographic specificity, rather than DNA, enables the re-identification and propose technical remedies to help participants make better sharing decisions.

  • Problem

    The paper addresses whether apparently de-identified PGP profiles can reveal participants’ identities despite containing sensitive medical, genomic, and demographic information.

  • Method

    The study links PGP demographics with voter and public-record data and extracts names embedded in publicly downloadable profile documents.

  • Results

    84 percent of 241 submitted names were correctly matched to profiles, rising to 97 percent when possible nicknames were allowed.

  • Takeaways & Limitations

    Participants can make profiles harder to link by making birth dates or ZIP codes less specific and removing names from uploaded documents.

  • Takeaways & Limitations

    The observed matching rate was constrained by temporal mismatch, sampled voter data, data quality, and Sweeney’s 87 percent uniqueness estimate being an upper bound.

Abstract

from arXiv · show

We linked names and contact information to publicly available profiles in the Personal Genome Project. These profiles contain medical and genomic information, including details about medications, procedures and diseases, and demographic information, such as date of birth, gender, and postal code. By linking demographics to public records such as voter lists, and mining for names hidden in attached documents, we correctly identified 84 to 97 percent of the profiles for which we provided names. Our ability to learn their names is based on their demographics, not their DNA, thereby revisiting an old vulnerability that could be easily thwarted with minimal loss of research value. So, we propose technical remedies for people to learn about their demographics to make better decisions.

INTRODUCTION

The introduction argues that people need accurate information about the risks of publicly sharing medical and genomic data so they can make informed personal decisions. Apparently anonymous disclosures may still expose individuals to harms and threaten individual choice.

  • Personal control over medical and genomic data sharing matters because individuals can weigh risks and benefits against their own circumstances.
  • Sensitive disclosures, including sexual abuse, abortions, or depression medication, may benefit one person but harm another.
  • Removing names and addresses can create a false sense of anonymity that encourages people to share information publicly.
  • People need to understand actual risks and available technological protections to make smarter data-sharing decisions.

BACKGROUND

The PGP publicly displays extensive genetic, medical, behavioral, and demographic information under an open-consent model that omits direct names and addresses but may still permit identification. Its profiles also offer limited post-upload editing, while demographic specificity creates a linkage risk known from earlier work.

  • The PGP aims to publicly display genotypic and phenotypic information from 100,000 informed volunteers for research and personalized medicine.
  • Under “open consent,” participants may disclose identifying demographics while profiles appear de-identified through omission of names and addresses.
  • PGP consent warns that submitted data can identify participants or expose information previously shared under confidentiality.
  • Once uploaded, PGP profiles provide almost no way to amend or modify information, including files in complicated formats.
  • Unlike HIPAA-oriented public medical-data rules, PGP participants often disclose full birth dates and five-digit ZIP codes.
  • Sweeney’s earlier work showed that demographics without names can be linked to public registries to recover names and contact information.

METHODS

The study used publicly available PGP profiles and public-record sources to test whether demographic matching and embedded-name extraction could recover participant names. The main matching strategy compared profile demographics with voter and public-record data and counted uniquely resolved names.

  • The base Dataset contained 579 of 1,130 public PGP profiles with date of birth, gender, and five-digit ZIP code.
  • The experiment compared PGP demographics with a national sample of voter registrations and an online public-records source.
  • Demographic matching used date of birth, gender, and ZIP code, recording matches that yielded exactly one name.
  • Figure 1 depicts the demographic-linkage approach from a PGP profile to a voter list.
  • Programs downloaded public PGP-associated files and automatically extracted demographic information and names appearing in compressed-file filenames.

RESULTS

Across three strategies, the experiment produced 241 unique names matching PGP profiles, with 84 percent confirmed correct and up to 97 percent when plausible nicknames were allowed. The strategies contributed overlapping but also distinct names, while temporal mismatch and mobility limited correctness.

  • 84 percent of the 241 submitted names were correctly matched, rising to 97 percent when possible nicknames such as Jim for James were allowed.
  • Combining voter-data links, embedded names, and public-record searches yielded 241 unique names, or 42 percent of 579 profiles.
  • Embedded names found 74 otherwise-unaccounted names, voter data found 44 distinct names, and public records found 65 distinct names.
  • The strategies overlapped: embedded names shared 12 names with voter data and 17 with public records, while voter data and public records shared 74.
  • Public-record correctness was affected by temporal mismatch because profiles, 2011 voter data, and 2013 public records reflected different dates and people could change ZIP codes.

DISCUSSION

The discussion shows that re-identification can attach participants’ names to sensitive medical and genetic disclosures, creating privacy and economic risks. It proposes reducing demographic precision and removing names from uploaded documents, supported by tools that help participants assess and mitigate these risks.

  • Risks: PGP profiles can expose sensitive conditions, including abortions, sexual abuse, illegal drug use, alcoholism, and clinical depression, once participants are re-identified.The authors identify these disclosures as potential harms of profile re-identification.
  • Risks: Economic harm may arise when genetic predispositions become linked to named participants applying for life insurance.The paper notes that disclosure may lead to denied coverage or higher premiums, while nondisclosure may jeopardize claim payment; GINA does not cover life insurance.
  • Interpretation: The 42 percent unique-name matching result may reflect temporal mismatch, sampled voter data, data quality, and the fact that 87 percent was only an upper bound.These factors explain why the observed matching rate did not exceed the earlier demographic-uniqueness prediction.
  • Remedies: Participants can reduce re-identification risk by making date of birth or ZIP code less specific and removing names from uploaded documents.The proposed changes target both demographic linkage and explicit names embedded in files.
  • Remedies: The authors provide services that let people assess demographic uniqueness and edit CCR files to report only the birth year or omit date of birth.The CCR editor addresses a limitation of the PGP interface, which does not support editing the date-of-birth field directly.
Loading 1304.7605v1…