Source-linked AI summary

Routes for breaching and protecting genetic privacy

Yaniv Erlich, Arvind Narayanan

arXiv:1310.3197v1q-bio.GNcs.CRstat.AP

TL;DR

Broad genetic-data sharing is vital for biomedical discovery, but expanding privacy-breaching techniques threaten the ability to protect data originators. This review maps data-mining threats, evaluates their performance and limitations, and surveys privacy-preserving countermeasures; it concludes that technical and societal measures are both needed.

  • Problem

    Large-scale genetic-data sharing is needed for robust predictions and secondary discoveries, but it raises concerns about protecting the privacy of data originators.

  • Method

    The review categorizes genetic privacy-breaching techniques, examines their technical concepts, performance, and limitations, and groups potential privacy-preserving technologies by methodological approach.

  • Results

    Privacy-breaching techniques have rapidly expanded, and many can be executed by trained individuals with varying degrees of effort.

  • Takeaways & Limitations

    Protecting genetic privacy requires technical measures alongside informed consent, stakeholder engagement, social and ethical norms, legal frameworks, and education.

  • Takeaways & Limitations

    ADAD via expression data reaches maximal power when training and inference use eQTL from the same tissue, while cross-technology data causes significant accuracy loss.

Abstract

from arXiv · show

We are entering the era of ubiquitous genetic information for research, clinical care, and personal curiosity. Sharing these datasets is vital for rapid progress in understanding the genetic basis of human diseases. However, one growing concern is the ability to protect the genetic privacy of the data originators. Here, we technically map threats to genetic privacy and discuss potential mitigation strategies for privacy-preserving dissemination of genetic data.

Summary

Genetic data sharing accelerates research but creates privacy risks through cross-referencing information. The review identifies three main breach routes and surveys privacy-by-design responses as attacks become more capable.

  • Three main routes to genetic privacy breaches are identity tracing, attribute disclosure, and completion of sensitive DNA information.These routes respectively target identity, sensitive phenotypes, or masked genomic regions.
  • Identity tracing uses quasi-identifiers in DNA data or metadata to identify an unknown genetic dataset.
  • Attribute disclosure links an identified person’s DNA dataset to a sensitive phenotype.
  • Completion attacks seek genomic regions masked to protect participants.
  • Privacy-breaching techniques and tools have rapidly expanded, although most still require trained operators and varying degrees of effort.
  • Privacy-by-design responses include access control, differential privacy, and cryptographic techniques, while genetic databases have mainly adopted access control.

INTRODUCTION

Genetic information is being produced at rapidly increasing scale, and broad sharing is important for biomedical discovery and secondary analysis. The central challenge is whether genetic data can be de-identified well enough to balance privacy with dissemination.

  • Sequencing studies now involve thousands of individuals, while projects aim to sequence hundreds of thousands to millions and eventually everyone in routine care.
  • Broad sharing is needed to combine cohorts for analyses of complex traits and to support serendipitous secondary discoveries.
  • Privacy is a major participation concern and a regulatory requirement, motivating de-identification as a possible compromise between sharing and protection.
  • The review maps genetic privacy-breaching techniques and countermeasures while excluding general cybersecurity threats and privacy-loss implications dependent on social context.

PRIVACY BREACHING OF GENETIC DATA

The review organizes genetic privacy breaches by the sensitive information they reveal and summarizes their technical maturity, complexity, and supporting-data requirements. The categories are identity tracing, attribute disclosure attacks via DNA, and completion techniques.

  • Genetic privacy breaches fall into Identity Tracing, Attribute Disclosure Attacks via DNA, and Completion Techniques.
  • The categories share a strategy of exploiting DNA data to obtain new potentially sensitive information about a target or the target’s family.
  • Identity tracing links an unknown genome to the concealed identity of its data originator.
  • Table 1 categorizes the privacy-breaching techniques presented in the section.
  • The review evaluates techniques using maturation level, technical complexity, and auxiliary-information availability.

IDENTITY TRACING ATTACKS

Identity tracing attacks combine residual genetic or metadata clues to narrow an unknown dataset to its originator. Demographic, pedigree, genealogical, phenotypic, and side-channel information each create distinct tracing opportunities and limitations.

  • Identity tracing attacks: Identity tracing accumulates quasi-identifiers embedded in genetic datasets until one individual remains as the only match.Attack success depends on the information content available to the adversary relative to the base population size.
  • Identity tracing attacks: 60% of US individuals are uniquely identified by date of birth, sex, and 5-digit zip code, illustrating the power of demographic metadata.By contrast, Safe Harbor identifiers correctly identified 2 of 15,000 Hispanic patient records in one empirical study.
  • Identity tracing attacks: Pedigree structures can rapidly narrow searches because familial events create unique quasi-identifier combinations and enable linking relatives once one person is identified.Their low searchability generally limits pedigrees to manual verification, except where extensive population registries are available.
  • Identity tracing attacks: 10-14% of US Caucasian males from middle and upper classes were estimated to be subject to surname inference using two major Y-chromosome genealogy websites.Distant patrilineal relatives provide surname information, and surnames are highly searchable in public records and social networks.
  • Identity tracing attacks: Five surname inferences from Illumina datasets of three 1000 Genomes families exposed the identities of close to fifty research participants.Surname inference has therefore been demonstrated against whole-genome sequencing datasets, not only personal DNA searches.
  • Identity tracing attacks: Surname inference is constrained by reliance on Y-STRs, specialized processing, manual validation, false hits, spelling variants, and socio-ethnic variation.Most sequencing studies do not routinely report Y-STRs, and processing raw sequencing files is time- and resource-consuming.
  • Identity tracing attacks: Non-Y genealogical markers have uneven tracing potential: mitochondrial searches are expected to be weak, whereas autosomal matches may identify distant relatives.Autosomal matches can reduce the search space to no more than a few thousand individuals, but translating a match into candidate people remains challenging.
  • Identity tracing attacks: DNA-derived phenotypic quasi-identifiers have not yet been demonstrated for identity tracing, and their usefulness is limited by prediction error and weak searchability.Current genetic knowledge explains only a small share of variability in traits such as height, BMI, and facial morphology.

ATTRIBUTE DISCLOSURE ATTACKS VIA DNA (ADAD)

ADAD uses DNA data to connect an identified person with sensitive attributes, and can operate on individual genotypes, summary statistics, or expression profiles. These attacks can be powerful, though expression-based attacks face substantial practical barriers.

  • Attack model: ADAD creates a statistical bridge linking DNA from an identified target to DNA-derived data associated with sensitive attributes.The attributes may include disease, personality traits, or socioeconomic status.
  • Individual genotype data: 45 carefully chosen autosomal SNPs can match individuals with a TYPE I ERROR of 10-15 across most major populations.Random subsets of approximately 300 common SNPs are expected to provide sufficient information to uniquely match any person.
  • GWAS data: ADAD vulnerability motivated two-tier GWAS access, separating restricted individual genotypes and phenotypes from public allele-frequency summaries.The stated premise is that individual-level genotype–phenotype records are highly vulnerable because few SNPs are needed for matching.
  • GWAS data: Summary-statistic attacks integrate small allele-frequency biases across many SNPs to test whether a target contributed to a study.Later work improved the test statistic and exploited local LD structures; under common SNPs in LINKAGE EQUILIBRIUM, the improved statistic is guaranteed maximal POWER at any SPECIFICITY level.
  • Risk factors: Smaller studies produce stronger apparent summary-statistic biases, increasing ADAD discrimination power and specificity, while target participation is less probable a priori.Theoretical performance depends on both study size and the general population.
  • Expression data: Expression-profile ADAD perfectly matched 580 individuals in cross-dataset training and simulations predicted type I error of 1x10-5 with power of 85%.The prediction assumes an expression database covering the entire US population.
  • Expression data: Expression-based ADAD loses accuracy across technologies, reaches maximal power with tissue-matched training and inference, and requires large-scale data processing.These barriers led the NIH to make no policy changes regarding sharing human-subject expression data.

COMPLETION ATTACKS

Completion attacks reconstruct masked or unavailable genomic regions by exploiting linkage disequilibrium and reference panels. Genealogical information can extend this reconstruction to targets whose DNA is entirely unavailable.

  • Genotype completion: Genotype imputation restores missing genotype values using linkage disequilibrium between markers and reference panels with complete genetic information.The same strategy can expose genomic regions that were masked in partially accessible DNA data.
  • Genealogical completion: Completion attacks can predict genomic sequences without target DNA by combining reference panels with genealogical information about relatives.Shared DNA segments among relatives on a unique genealogical path indicate segments inherited by the target.
  • Governance: In May 2013, Iceland’s Data Protection Authority prohibited this completion technique until consent could be obtained from people outside the original reference panel.

MITIGATION TECHNIQUES

Mitigation strategies range from access control and anonymity to differential privacy and cryptographic computation. Each approach balances protection, data utility, computational feasibility, or administrative burden differently.

  • Access control: Access control secures sensitive data, screens applicants and projects, and requires secure storage, non-identification commitments, and periodic usage reports.A retrospective analysis found 8 data management incidents in close to 750 dbGAP studies, mostly non-adherence to technical regulations, with no reported participant privacy breaches.
  • Access control: Access control can create an illusion of security because custodians lack real oversight after applicants receive the data.The approach also requires constant resource management and creates administrative burden for custodians and users.
  • Anonymity: k-anonymity bins quasi-identifiers so each record matches at least k-1 others, trading greater anonymity for reduced data utility as k increases.A rule of thumb recommends k≥5, but k-anonymity remains vulnerable to attribute disclosure when adversaries have prior knowledge of target participation.
  • Anonymity: Differential privacy perturbs summary statistics so datasets differing by one individual produce extremely close outputs, limiting certainty about target participation.Its challenge is minimizing perturbation while preserving useful statistics.
  • Cryptography: Secure multiparty computation enables entities to compute on private genetic inputs without revealing those inputs to one another or third parties.Genetic applications include privacy-preserving matching and secure multi-center GWAS analysis using secret sharing.
  • Cryptography: Cryptographic performance varies sharply by task: few-locus analyses finish in under a second, whereas whole-genome analyses can take days and gigabytes of bandwidth.Whole-genome computation is therefore impractical at the current time.
  • Cryptography: Homomorphic encryption supports privacy-preserving outsourcing of computations on genetic information to interpretation services.The approach addresses repeated interactions with genetic services, which increase opportunities for privacy breaches.

CONCLUSION

Genetic privacy mitigation has shifted toward recognizing technically sophisticated attacks, but existing protections remain costly and incomplete. Durable protection also requires informed consent and social, legal, and educational measures.

  • Current status: Studies show that motivated, technically sophisticated adversaries can breach genetic privacy, while access control remains useful but resource- and time-consuming.
  • Current status: The privacy field still lacks methodologies with an impact comparable to communication security because of technical and human factors.
  • Broader measures: Balancing privacy and data sharing requires balanced informed consent plus stakeholder engagement to develop social and ethical norms, legal frameworks, and educational programs.These measures may reduce misuse despite the inability to theoretically prevent privacy breaches.

GLOSSARY

This glossary defines key concepts related to health-data de-identification, genetic variation, cryptographic protection, attacks, and statistical error.

  • Safe Harbor de-identifies protected health information by removing 18 types of quasi-identifiers.
  • Haplotypes are sets of alleles located along the same chromosome.
  • Cryptographic hashing produces a fixed-length output from any input and makes recovering the input difficult.
  • Dictionary attacks attempt to reverse cryptographic hashes by scanning a relatively small input space.
  • Type I error is introduced as a statistical concept, but its definition is incomplete in the passage.
Loading 1310.3197v1…