Source-linked AI summary
Technical Privacy Metrics: a Systematic Survey
Isabel Wagner, David Eckhoff
TL;DR
Privacy-metric diversity makes informed selection difficult and leaves privacy studies often incomparable. The survey systematizes over eighty metrics across multiple domains, classifies them by measured aspect, inputs, and protected data, and proposes nine selection questions. It concludes that informed choices benefit from multiple metrics while identifying unresolved issues in user-attitude integration and metric quality.
Problem
The diverse privacy-metrics literature lacks a structured overview, making informed metric selection difficult and studies often incomparable.
Method
The survey classifies and discusses over eighty metrics using output measures, required inputs, protected data, adversary models, and examples from multiple privacy domains.
Results
The survey provides a nine-question method for identifying suitable metrics and reports that attack and defense studies tend to emphasize different metric families.
Takeaways & Limitations
The authors recommend selecting multiple metrics to cover multiple privacy aspects and present the systematization as a reference toolbox for informed choices.
Takeaways & Limitations
Integrating user attitudes, behaviors, or perceptions with technical metrics remains an open question, including whether and where such integration is useful.
Abstract
from arXiv · showhide
The goal of privacy metrics is to measure the degree of privacy enjoyed by users in a system and the amount of protection offered by privacy-enhancing technologies. In this way, privacy metrics contribute to improving user privacy in the digital world. The diversity and complexity of privacy metrics in the literature makes an informed choice of metrics challenging. As a result, instead of using existing metrics, new metrics are proposed frequently, and privacy studies are often incomparable. In this survey we alleviate these problems by structuring the landscape of privacy metrics. To this end, we explain and discuss a selection of over eighty privacy metrics and introduce categorizations based on the aspect of privacy they measure, their required inputs, and the type of data that needs protection. In addition, we present a method on how to choose privacy metrics based on nine questions that help identify the right privacy metrics for a given scenario, and highlight topics where additional work on privacy metrics is needed. Our survey spans multiple privacy domains and can be understood as a general framework for privacy measurement.
1 INTRODUCTION
Privacy metrics quantify privacy levels and PET protection, but their diversity makes informed selection difficult. This survey structures the field and provides guidance for choosing metrics.
- Privacy and its measurement: Privacy is a fundamental right, but its meaning depends on control over personal information and the context governing its use.Contextual integrity treats privacy violations as uses of personal information that violate applicable social norms.
- Privacy and its measurement: Technical privacy metrics take system properties as inputs and produce numerical or canonical values that quantify privacy.Inputs may include leaked information, indistinguishability, adversary estimates, resources, true outcomes, prior knowledge, and parameters.
- Motivation: The literature lacks a structured, comprehensive overview, making metric selection for PET evaluation difficult.The authors identify this gap as a motivation for structuring technical metrics that measure system privacy or PET effectiveness.
- Survey approach: The survey classifies privacy metrics by adversary models, data sources, inputs, outputs, and protected data, including eight output categories.The categories are uncertainty, information gain or loss, data similarity, indistinguishability, adversary success probability, error, time, and accuracy/precision.
- Contributions: The authors review over eighty metrics, provide nine-question selection guidance, and identify future work on interdependent privacy, metric combinations, and metric quality.The survey is intended as a reference guide and framework for informed metric choices.
2 CONDITIONS FOR PRIVACY METRICS
The survey finds no consensus on universal conditions for privacy metrics. Proposed criteria instead provide guidelines for improving their strength, usability, and meaningfulness.
- General conditions: Privacy metrics lack a general consensus on the conditions they must fulfill.Many surveyed measures are called privacy metrics even when they do not satisfy all four mathematical metric properties.
- Existing recommendations: Proposed requirements include understandability, independence from cost and utility, probability-based identification bounds, and adversary-oriented interpretation.These criteria are drawn from recommendations by multiple authors.
- Existing recommendations: Other recommendations emphasize attack difficulty, predictable variables, required attack resources, privacy level, hidden sensitive data, and resulting data quality.These proposals contrast resource-based and probability- or cardinality-based approaches.
- Role of conditions: The authors treat proposed conditions as guidelines rather than strict qualification requirements for privacy metrics.They argue that such guidelines can improve the strength, usability, and meaningfulness of newly proposed measures.
3 PRIVACY DOMAINS
The survey illustrates privacy metrics across six domains, where adversaries and protected information differ substantially. Examples span communication, databases, location services, smart metering, social networks, and genomics.
- Overview: The survey describes six privacy domains to provide context and examples for privacy-enhancing technologies and metrics.The domains include communication systems, databases, location-based services, smart metering, social networks, and genomics.
- Communication systems: Communication-system privacy focuses on hiding sender, receiver, or sender-receiver relationships, rather than only concealing message contents.Anonymous communication is the central challenge in this domain.
- Databases: Database adversaries seek to identify individuals and reveal sensitive attributes in interactive queries or publicly released sanitized databases.Databases may contain microdata or aggregate data such as averages.
- Location-based services: Location-based services expose information that can reveal sensitive attributes such as home and work locations and enable movement profiling.The domain concerns mobile services that use location information to provide context-aware functionality.
- Smart metering: Smart-metering privacy concerns fine-grained household electricity data that providers use for billing and network optimization but may also exploit for behavioral inference.The energy provider can act as an adversary beyond the stated purpose of data collection.
- Social networks and genomics: Social-network and genomic privacy involve identifying users or inferring sensitive attributes, with genomic data also uniquely identifying individuals and revealing disease susceptibility.Genomic exposure can create risks such as genetic discrimination or blackmail.
4 CHARACTERISTICS OF PRIVACY METRICS
Privacy metrics can be classified by adversary model, data source, inputs, and output property to guide metric selection for specific scenarios. Because output categories capture different aspects of privacy, a complete estimate may require metrics from multiple categories.
- Four characteristics classify privacy metrics: adversary models, data sources, inputs, and output properties.These characteristics provide an initial guideline for choosing metrics for specific scenarios.
- Metrics can only compare privacy-enhancing technologies when they use the same adversary model; omitting adversary capabilities can overestimate privacy.Metrics without an explicit adversary model implicitly assume limited adversary capabilities, while attacks may exploit other data properties.
- Adversary models specify capabilities such as system access, activity, system role, adaptability, prior knowledge, and available resources.The survey distinguishes local or global, active or passive, internal or external, static or adaptive, knowledge-based, and resource-based adversaries.
- Data-source categories describe what information needs protection and how an adversary gains access, including published, observable, re-purposed, and other data.Published data is willingly and persistently public; observable data is transient; re-purposed data is used beyond its initial purpose; other data is neither public nor intended for adversary access.
- Input categories include the adversary’s estimate and available resources, and their availability or assumed conditions determine whether a metric fits a scenario.An adversary’s estimate may be a posterior probability distribution describing likely message senders or energy consumption.
- The survey’s eight output categories represent different privacy aspects, so a more complete estimate requires metrics from different categories.The output taxonomy is presented as an intuitive classification, although category boundaries can blur and some metrics could fit multiple categories.
5 PRIVACY METRICS
The survey organizes over eighty privacy metrics by measured output, while describing their inputs, applications, and important interpretive limitations.
- The survey groups over eighty metrics by measured output and unifies their notation where possible.Tables summarize classifications, value ranges, and whether higher or lower values indicate better privacy.
- Uncertainty: Uncertainty metrics measure an adversary’s uncertainty, often using entropy to assess associations between users, messages, or locations.Entropy can be computed over time with probabilities updated using Bayesian belief tables.
- Uncertainty: An anonymity set is the group of users indistinguishable from a target, but its size ignores adversary knowledge and target likelihood.The survey notes that anonymity-set size can be combined with normalized entropy.
- Uncertainty: Entropy captures uncertainty or effective anonymity-set size, but outliers, incomparable distributions, nearby locations, and incorrect estimates can make interpretation difficult.Normalized entropy uses Hartley entropy to support comparisons and can represent information leakage.
- Uncertainty: Rényi entropy generalizes Shannon entropy; Hartley entropy is the best-case number-based case, while min-entropy is the worst-case highest-probability case.Unlinkability instead applies entropy to possible relationships or partitions and can account for prior knowledge through a ratio in [0, 1].
5.2 Information Gain or Loss
Information gain or loss metrics quantify what an adversary can learn, often by comparing true data, observations, estimates, or prior knowledge.
- Information gain or loss metrics measure how much information an adversary obtains, explicitly accounting for prior information in many formulations.They are used across communication systems, databases, genome privacy, smart metering, and social networks.
- Counting disclosed information items identifies the amount disclosed but not the severity of a leak because sensitivity is ignored.Examples include compromised users and leaked DNA base pairs.
- Relative entropy measures probabilistic information revealed by comparing the true distribution with an adversary’s estimate or observations.The distributions must satisfy absolute continuity.
- Mutual information measures leakage between true data and observations, with normalizations representing dependence, average bits leaked per entry, or fractional privacy loss.Conditional mutual information incorporates prior knowledge.
- Maximum information leakage measures the greatest information about private data obtainable from a single output over all possible outputs.Loss of anonymity similarly evaluates worst-case leakage, while relative loss of anonymity additionally incorporates prior knowledge.
- In anonymous communication, system anonymity level estimates the additional information needed to reveal sender-receiver relationships using entropy over possible combinations.The estimate accounts for equivalence-class multiplicities and is normalized by the number of users.
5.3 Data Similarity
Data similarity metrics derive privacy levels from observable or published data, especially in database sanitization and publishing, but several anonymity variants have specific attack limitations.
- Data similarity metrics measure properties within one dataset or between private and public data, usually without modeling an adversary.Most originate in database sanitization and data publishing.
- k-Anonymity and extensions: k-Anonymity requires each equivalence class to contain at least k records after identifying columns are removed.It can fail against high-dimensional data, auxiliary-data correlation, attribute disclosure, repeated releases, and semantically close sensitive data.
- k-Anonymity and extensions: (α,k)-anonymity limits sensitive-value frequency within each equivalence class, but attribute linkage can still occur below α.
- k-Anonymity and extensions: ℓ-diversity requires multiple well-represented sensitive values per equivalence class, yet can fail under probabilistic inference, skewed distributions, semantic similarity, and repeated releases.Stronger instantiations constrain frequencies using entropy or recursion.
- k-Anonymity and extensions: m-Invariance protects multiple releases by requiring at least m rows and distinct sensitive values within each equivalence class.t-closeness instead bounds the Earth Mover Distance between class-level and table-wide sensitive-value distributions.
- Numerical and relational extensions: For numerical attributes, (k,e)-anonymity widens sensitive-value ranges, but uneven distributions can enable proximity attacks and attribute disclosure.(ϵ,m)-anonymity addresses this by bounding inference probability to at most 1/m.
- Numerical and relational extensions: Other extensions apply anonymity at record-owner level, bound sensitive-value inference for column groups, or support sequential releases using shared columns.
5.4 Indistinguishability
Indistinguishability metrics assess whether an adversary can distinguish outcomes, with formal variants adapting guarantees to databases, locations, distributed settings, and computational limits.
- Indistinguishability metrics indicate whether an adversary can distinguish two outcomes and often accompany formal privacy mechanisms.Applications include databases, communication, location-based systems, and smart metering.
- Semantic security uses a challenge-response game in which the adversary distinguishes alternative outcomes; computational privacy requires advantage below negligible ϵ(k).Unconditional privacy requires zero advantage.
- Differential privacy variants: Differential privacy makes query outputs similar whether or not an individual’s record is included, typically by adding random noise.Interactive guarantees weaken as permitted queries accumulate privacy parameter ϵ.
- Differential privacy variants: Approximate differential privacy adds δ, weakening guarantees while permitting higher-utility releases or broader query types.The survey states that δ should be smaller than the inverse of any polynomial in database size.
- Differential privacy variants: Distributed differential privacy conditions guarantees on compromised users’ randomness, while distributional privacy protects parameters governing data generation.These variants address distributed aggregation and settings such as smart metering.
- Domain-specific variants: Geo-indistinguishability applies planar noise to locations, making protection depend on distance d, while d-χ-privacy substitutes a domain-specific distinguishability metric.Euclidean distance supports location privacy, and maximum distance can preserve smart-metering trends while distorting accuracy.
- Domain-specific variants: Joint differential privacy protects an individual’s data from other individuals, and computational differential privacy limits the adversary to computationally bounded behavior.
- Domain-specific variants: Information privacy keeps posterior and prior probabilities close and additionally bounds maximum information leakage to ϵ/ln 2 bits.Event-source unobservability requires equality between prior and posterior probabilities for all possible observations.
5.5 Adversary’s Success Probability
Adversary’s success probability metrics quantify how likely an adversary is to achieve a scenario-specific privacy breach. They are broadly applicable but depend strongly on the adversary model and definition of success.
- These metrics can be applied across domains wherever an adversary and a meaningful success condition can be defined.
- Adversary’s success probability metrics measure the probability that an adversary succeeds, with success defined according to the application scenario.Examples include finding a similar target record in databases, identifying a message sender, or compromising a communication path.
- Degrees of anonymity classify communication privacy according to the adversary’s probability of identifying a sender or receiver.The categories range from absolute privacy and beyond suspicion to exposed and provably exposed, using thresholds on p(x).
- The degree of anonymity may not reflect the adversary’s real success probability because it ignores the anonymity-set cardinality.
- Privacy breach metrics compare posterior and prior probabilities, while d,γ-privacy adds bounds on these probabilities and their ratio.Related examples include transaction-item inference, networking message generation, and database membership inference; δ-presence assumes adversary and publisher access the same external tables.
- δ-presence’s shared-external-table assumption may not hold in practice, limiting its applicability.
5.6 Error
Error metrics quantify how inaccurately an adversary estimates a true outcome. They require knowledge of the true outcome for computation and can be applied across privacy domains.
- Error-based metrics quantify the error an adversary makes when creating an estimate.Because computing them requires the true outcome, the adversary cannot compute these metrics directly.
- These metrics are applicable to all domains, similarly to adversary’s success probability metrics.
- Adversary’s expected estimation error computes the expected distance between the true and estimated outcomes over the adversary’s posterior.The distance can be Euclidean or binary, in which case the metric reduces to the adversary’s probability of error.
- Distance error extends expected estimation error across multiple timesteps and location-assignment hypotheses.Each hypothesis assigns probabilities to locations, while its distance term measures deviation from the correct location.
- Mean squared error measures discrepancies between adversarial observations and the true outcome, including communication-relationship assignment and participatory-sensing reconstruction.
- Classification error measures the percentage of users or events incorrectly classified by the adversary.Examples include incorrect de-anonymization and incorrect classification of smart-metering events.
5.7 Time
Time-based metrics measure the time or distance an adversary needs to compromise privacy or become confused. They span communication, location, and smart-metering scenarios but depend on how success or confusion is defined.
- Time-based metrics treat time as a resource required by an adversary to compromise privacy.Some measure time until adversarial success, while others measure time until adversarial confusion.
- The general time-until-success metric assumes the adversary eventually succeeds and therefore represents a pessimistic privacy measure.Its value depends on the scenario-specific definition of success.
- Communication-path compromise is one success condition, occurring in Tor when the adversary controls all relays on a user’s onion-routing path.
- Maximum tracking time measures the cumulative time during which a target’s anonymity set remains of size 1.In smart metering, it instead describes the percentage of an interval during which the adversary correctly classifies load transitions.
- Maximum tracking time can overestimate privacy because it treats complete adversarial certainty as necessary for success.An adversary may continue tracking when only a small number of users remain in the target’s anonymity set.
- Mean time to confusion measures the duration for which adversarial uncertainty remains below a confusion threshold, while distance to confusion measures travel distance until uncertainty rises above it.Uncertainty is measured using entropy over the estimated probabilities of anonymity-set members.
5.8 Accuracy / Precision
Accuracy and precision metrics quantify properties of an adversary’s estimate, including geographic precision and uncertainty-region relationships. Their interpretation depends on the estimate, confidence level, and privacy scenario.
- Accuracy metrics quantify the accuracy of an adversary’s estimate, with many originating in location-based services and measuring geographic precision.The survey notes that inaccurate estimates can correspond to higher privacy, even though estimate accuracy may not capture correctness or certainty.
- Confidence interval width measures privacy at confidence level τ% using the width of the interval containing the true outcome within the adversary’s estimate.The metric is defined as |x2 − x1| for an interval whose probability is τ/100.
- Knowledge of confidence interval width can allow reconstruction of the original distribution when perturbed data are published.
- Statistically strong event unobservability compares message patterns across network locations to prevent event locations being revealed by sudden message bursts.Unobservability requires roughly similar inter-message-delay distributions across network parts.
- The size of the uncertainty region is the minimum region to which an adversary can narrow a target user’s position.
- Obfuscated-region accuracy indicates how relevant an enlarged reported region remains to a service provider after satisfying a minimum user requirement.
- Sensitive-region coverage evaluates overlap between a user’s sensitive region and the adversary’s uncertainty region.The adversary succeeds in linking the user to the sensitive region when the regions overlap; normalization reaches 1 when the uncertainty region is equal to or contained in the sensitive region.
6 HOW TO SELECT SUITABLE PRIVACY METRICS
The survey proposes nine questions for selecting privacy metrics according to the privacy aspects, adversary, data, inputs, audience, and related work relevant to a scenario. It recommends combining multiple metric outputs and considering both attack- and defense-oriented perspectives.
- Selection procedure: Nine questions guide metric selection by considering privacy aspects, adversary types, protected data sources, available inputs, target audience, and related work.The strategy was successfully applied in a genomic privacy case study.
- Output measures: Choose indistinguishability metrics when privacy properties must be proven; use other output categories when quantifying privacy levels.The survey classifies privacy-metric outputs into eight categories.
- Output measures: Multiple output categories capture additional privacy aspects because no metric measures privacy directly.The survey recommends measuring several outputs rather than fixing one output measure.
- Adversary models: Attack studies tend to use time, error, or adversary-success metrics, while PET studies tend toward accuracy, similarity, and indistinguishability metrics.The survey argues that both perspectives can benefit from selecting metrics associated with the other side.
- Adversary models: Time-based metrics differ across domains: communication systems measure time until adversary success, whereas location privacy measures time until adversary confusion.The survey notes that it is not obvious which adversary assumption applies across domains.
- Adversary models: Metrics without an adversary model can misstate privacy when adversaries possess relevant prior knowledge.The survey uses k-anonymity as an example and identifies resource-based metrics as an area needing further research.
7 FUTURE RESEARCH DIRECTIONS
The survey identifies open problems in adapting, combining, aggregating, and evaluating privacy metrics, including metrics for interdependent privacy and integration with user attitudes. It emphasizes that metric quality and interpretation remain insufficiently established across scenarios.
- Interdependent privacy: Interdependent privacy can be measured by tracking an existing metric as interdependency increases or by creating metrics that explicitly model consequences across users.The survey identifies social networks, location privacy, and genome privacy as example domains.
- Privacy attitudes and behaviors: Integrating user attitudes, behaviors, or perceptions with technical metrics remains open, including whether the integration is useful and which scenarios benefit.Some metrics already combine technical measures with user-specified preferences.
- Aggregation: Aggregating metrics across many entities and preserving distributional information in aggregate measures require further study.Box plots and violin plots reveal ranges, quartiles, and distributions that averages can hide, but it is unclear how to preserve these benefits in aggregate metrics.
- Combining metrics: Combining sensitivity scores with technical metrics has unclear interpretability, while normalization methods are not established for some metrics.The survey cites user-centric privacy, privacy score, normalized entropy, and normalized mutual information as examples.
- Combining metrics: Adapted metrics raise questions about whether their mechanisms generalize to other metrics and whether alternative adaptation mechanisms exist.Examples include computational differential privacy and entropy combined with Bayesian belief tables.
- Quality of metrics: Rigorous studies are needed to evaluate metric quality and meaningfulness beyond specific scenarios, indicators, and metric sets.A prior study of 22 genome-privacy metrics found substantial variation in consistency and monotonicity but had limited scope.
8 CONCLUSION
The survey reviews over eighty privacy metrics across six domains and organizes them by privacy aspect, required inputs, and protected data type. It also provides a nine-question selection method and identifies combination, aggregation, and interdependent privacy as research needs.
- Scope and review: The survey reviews over eighty privacy metrics using examples from six different privacy domains.It presents the work as a comprehensive review of privacy metrics.
- Systematization: Its categorizations organize metrics by the privacy aspect measured, required inputs, and type of data requiring protection.The survey also highlights combination and aggregation of metrics and interdependent privacy as areas needing additional work.
- Metric selection: A nine-question method helps identify suitable metrics for a given scenario and supports selecting multiple metrics to cover multiple privacy aspects.The survey presents the systematization as a reference guide and toolbox for privacy researchers.