Source-linked AI summary
The Privacy Policy Landscape After the GDPR
Thomas Linden, Rishabh Khandelwal, Hamza Harkous, Kassem Fawaz
TL;DR
The paper asks how the GDPR changed online privacy policies, addressing limited prior evidence with a longitudinal comparison of 6,278 English-language policies before and after enforcement. Its analyses find a major overhaul, stronger coverage and compliance, and improved EU presentation, but also longer policies, inconsistent specificity, and a landscape still in transition.
Problem
The paper investigates the GDPR’s impact on online privacy policies, a question for which previous studies had not provided a comprehensive answer.
Method
The authors compare pre-GDPR and post-GDPR versions of 6,278 English-language policies from inside and outside the EU across presentation, text, coverage, compliance, and specificity.
Results
The GDPR coincided with a major policy overhaul: policies became longer, EU presentation improved, coverage and compliance increased, and specificity changes were mixed.
Takeaways & Limitations
Privacy policies are incorporating more privacy rights and information, especially in the EU, but full disclosure and transparency have not yet stabilized.
Takeaways & Limitations
Automated policy selection and analysis are inherently error-prone, and machine-learning errors plus human disagreement may affect specificity and compliance findings.
Abstract
from arXiv · showhide
The EU General Data Protection Regulation (GDPR) is one of the most demanding and comprehensive privacy regulations of all time. A year after it went into effect, we study its impact on the landscape of privacy policies online. We conduct the first longitudinal, in-depth, and at-scale assessment of privacy policies before and after the GDPR. We gauge the complete consumption cycle of these policies, from the first user impressions until the compliance assessment. We create a diverse corpus of two sets of 6,278 unique English-language privacy policies from inside and outside the EU, covering their pre-GDPR and the post-GDPR versions. The results of our tests and analyses suggest that the GDPR has been a catalyst for a major overhaul of the privacy policies inside and outside the EU. This overhaul of the policies, manifesting in extensive textual changes, especially for the EU-based websites, comes at mixed benefits to the users. While the privacy policies have become considerably longer, our user study with 470 participants on Amazon MTurk indicates a significant improvement in the visual representation of privacy policies from the users' perspective for the EU websites. We further develop a new workflow for the automated assessment of requirements in privacy policies. Using this workflow, we show that privacy policies cover more data practices and are more consistent with seven compliance requirements post the GDPR. We also assess how transparent the organizations are with their privacy practices by performing specificity analysis. In this analysis, we find evidence for positive changes triggered by the GDPR, with the specificity level improving on average. Still, we find the landscape of privacy policies to be in a transitional phase; many policies still do not meet several key GDPR requirements or their improved coverage comes with reduced specificity.
1 Introduction
This paper examines how the GDPR changed online privacy policies through a longitudinal, large-scale comparison of pre- and post-GDPR versions. The analysis finds extensive changes across presentation, text, coverage, compliance, and specificity, with benefits that are stronger in the EU but not uniformly positive for users.
- Study scope: The study compares pre-GDPR and post-GDPR versions of 6,278 unique English-language privacy policies across regions and user-experience stages.The analysis covers presentation, textual features, coverage, compliance, and specificity.
- Presentation: 470 participants rated EU policies as more attractive and clearer after the GDPR, while Global policies showed no significant visual change.The user study was conducted on Amazon Mechanical Turk.
- Text-Features: +35% and +25% more words, and +33% and +22% more sentences, were found in EU and Global policies respectively, without major changes in sentence-structure metrics.The increases are reported as average changes for EU and Global policies.
- Coverage: Data-retention coverage improved by 52% for EU policies and 31% for Global policies, alongside gains for special audiences and user access.Special-audience coverage improved by 16% for EU and 12% for Global policies; user-access coverage improved by 21% and 13%, respectively.
- Compliance: 15.1% of EU policies and 10.66% of Global policies improved on compliance metrics, compared with 10.3% and 7.12% that worsened.These are averages across seven GDPR compliance metrics.
- Specificity: Specificity improved for 25.22% of EU and 19.4% of Global policies, but declined for 22.7% and 17.8%, respectively, when broader coverage came at its expense.The results indicate that increased topic coverage did not always preserve detail about data practices.
2 GDPR Background
The GDPR establishes detailed information and rights requirements for online privacy practices. Providers must explain how data is collected, shared, retained, and processed, while informing users about associated rights in clear and accessible language.
- Entities: The GDPR defines data subjects, controllers, processors, and third parties as distinct entities involved in personal-data collection and processing.A controller may employ a processor or authorize a third party to process user data.
- Information rights: Article 12 requires privacy information to be concise, transparent, intelligible, easily accessible, and written in clear, plain language.This requirement concerns users’ right to be informed about service providers’ privacy practices.
- Disclosure duties: Articles 13 and 14 require disclosure of the controller’s contact information, collection purposes, shared-data recipients, retention period, and data types.They distinguish between data collected directly from users and data obtained indirectly.
- User rights: Articles 15–22 cover user rights including access, rectification, erasure, restriction of processing, data portability, and objection.The paper focuses on GDPR requirements in Articles 12–22.
3 Creation of The Policies Dataset
The study constructs longitudinal EU and Global datasets of English-language privacy policies, pairing pre-GDPR and post-GDPR snapshots to measure policy change. The resulting corpus shows greater change for EU policies, with May 2018 serving as a major inflection point.
- Dataset construction: The researchers assembled EU and Global datasets of English-language privacy policies, selecting pre-GDPR and post-GDPR snapshots between January 2016 and May 2019.The pre-GDPR snapshot was the last stable version before GDPR enforcement, while the post-GDPR snapshot was the most recent version.
- Dataset construction: Candidate policy pages were identified by crawling websites and classified using language detection and a one-layer CNN, achieving 99.09% testing accuracy.The classifier prioritized high precision by rejecting unsuitable pages, including non-English pages and pages lacking a valid standalone privacy policy.
- Change analysis: Significant textual change was defined as a fuzzy similarity ratio of 95% or less between consecutive snapshots, with key-change dates identifying the nearest substantial change to GDPR enforcement.Stable pre-GDPR and post-GDPR versions were selected around each policy’s key-change date.
- Dataset construction: The final datasets contained 3,084 EU snapshot pairs and 3,592 Global snapshot pairs after removing policies without valid pre-GDPR or post-GDPR snapshots.Together, the datasets contained 6,676 snapshot pairs, while 398 policy URLs were shared between the sets.
- Change analysis: More than 45% of studied policies changed between March and July 2018, and June 2018 was the most frequent key-change date in both datasets.The EU set showed more signs of change than the Global set, indicating a higher impact inside the EU.
4 Presentation Analysis
A within-subjects user study evaluated screenshots of pre-GDPR and post-GDPR privacy policies for visual attractiveness, simplicity, and user experience. EU policies improved significantly in attractiveness and simplicity, whereas Global policies showed no significant visual change.
- Study design: 470 Amazon Mechanical Turk participants evaluated randomized screenshots of privacy policies without being primed about the comparison.The study used 800 screenshots from 400 randomly selected policies and analyzed images receiving at least five evaluations after quality filtering.
- Study design: The study measured whether policies had an attractive appearance, a clean and simple presentation, and a positive user experience.Respondents were instructed to assess presentation rather than read the screenshot content.
- Results: For EU policies, pre-GDPR versus post-GDPR differences were significant for attractiveness (p=0.039) and simplicity (p=0.040), but not positive experience (p=0.057).The analysis grouped responses into disagree, neutral, and agree categories before applying chi-squared tests.
- Results: For Global policies, all three presentation questions failed to show significant pre-GDPR versus post-GDPR differences, with all p>0.4.Thus, the study found no significant visual-interface change outside the EU.
- Results: Post-GDPR EU policies differed significantly from post-GDPR Global policies for attractiveness (p = 0.007), simplicity (p = 0.014), and positive experience (p=9.06e−5).The post-GDPR score distributions indicated a more positive user experience for EU policies than Global policies.
5 Text-Feature Analysis
The analysis compares pre-GDPR and post-GDPR policy text using five syntactic metrics. Policies became substantially longer in both datasets, but sentence structure showed little improvement.
- Method: The study uses five text metrics, compares paired pre-GDPR and post-GDPR policies, and applies Bonferroni-corrected statistical tests.Wilcoxon signed-rank tests were used because policy text showed high variability.
- Length: +50% syllables, +35% words, and +33% sentences in EU policies indicate substantial post-GDPR length increases.The null hypothesis was rejected for all three metrics.
- Length: +38% syllables, +25% words, and +21% sentences in Global policies show significant but smaller increases than in the EU set.The increase outside the EU was noticeably smaller than for EU policies.
- Sentence structure: Passive Voice Index did not change significantly in either dataset.The analysis therefore found no evidence of improvement on this sentence-structure metric.
- Sentence structure: Word-per-sentence changes were small: -5% for EU policies and +3% for Global policies.Average sentences exceeded 43 words in both datasets, so these changes did not establish a meaningful structural shift.
6 Automated Requirement Analysis
The paper develops a structured-querying workflow that translates privacy-policy text into machine-generated labels and applies logic queries to assess policy goals. This approach supports flexible semantic analysis beyond keyword matching.
- Methodology Overview: Structured querying adds two abstraction levels: labels for embedded privacy practices and logic queries over those labels.The workflow first represents text segments with labels, then reasons across segments using first-order logic.
- Polisis: Machine-learning classifiers assign labels to semantically coherent policy segments using human-annotated training data.Polisis produces high-level categories and lower-level attribute values for each segment.
- Methodology Overview: The querying workflow separates filtering relevant segments from scoring them against defined requirements or goals.This filtering-scoring approach is used for the paper’s in-depth analyses.
- Advantages: Structured querying covers semantically similar wording more flexibly than heuristics or keyword analysis.Queries can be adapted to new goals without creating new labeling data for each goal.
- Polisis: Polisis uses a taxonomy containing nine high-level privacy categories and 20 lower-level privacy attributes.Examples include purposes of data processing and categories of privacy practices.
7 Coverage Analysis
The coverage analysis tests whether privacy policies mention high-level privacy categories before and after the GDPR. Previously underrepresented GDPR-relevant categories improved, especially among EU policies.
- Findings: EU policies significantly improved in four categories: Data Retention, International & Specific Audiences, Privacy Contact Information, and User Access, Edit & Deletion.These improvements occurred among categories underrepresented before the GDPR.
- Methodology: Categories with already high coverage changed little in either dataset, while underrepresented categories showed significant change.Coverage was determined by whether at least one policy segment received the category label.
- Findings: User Access coverage improved 21% for EU policies and 13% for Global policies.The broader analysis links this category to users’ ability to access and rectify their information.
8 Compliance Analysis
The compliance analysis compares seven GDPR-related requirements in pre-GDPR and post-GDPR policies using ICO checklist queries. Coverage improved overall, but several requirements remained missing and performance varied by requirement.
- Scope: Compliance evidence was restricted to seven ICO-compatible requirements because Polisis’ pre-GDPR taxonomy could not quantify some newer checklist concepts.The selected items were manually compared with GDPR Articles 12–20.
- Requirements worsened: ICO-Q6 had the highest observed decline at 11.7% in both the EU and Global sets.ICO-Q6 concerns updating privacy policies when personal data is used for new purposes.
- Requirements still missing: ICO-Q2 remained absent from 47.1% of EU policies and 49.8% of Global policies.This requirement concerns specifying sources of personal data obtained by first parties.
- Requirements still covered: ICO-Q3 was maintained at score 0 in only 13.0% of EU policies and 13.9% of Global policies across both snapshots.ICO-Q3 concerns specifying third-party entities receiving the data.
- Requirements still covered: At least 38% of policies in both datasets remained compliant with the requirements, except for ICO-Q2.The covered requirements include data sources, recipients, and users’ rights to withdraw and access data.
- Overall compliance: 59.3% of EU and 58.2% of Global post-GDPR policies met the seven analyzed queries.These figures combine policies whose requirements remained covered with those whose coverage improved.
- Overall compliance: Average compliance improvement was 4.8% for EU policies and 3.5% for Global policies.The EU set had both the higher average compliance score and the larger improvement margin.
9 Specificity Analysis
The study measures whether privacy policies became more explicit after the GDPR by scoring coverage and specificity across eight data-practice queries. Post-GDPR policies show both improved specificity and reduced specificity, reflecting a tension between comprehensive coverage and precise disclosure.
- Measurement: Specificity analysis distinguishes whether policies mention practices from whether they describe them explicitly, using eight queries scored by explicitness.The queries cover collection methods, third-party acquisition, information types, collection and sharing purposes, and retention purposes.
- Measurement: The analysis compares each policy’s pre-GDPR and post-GDPR specificity scores across four outcomes: not covered, same, worse, or improved specificity.Policies can retain the same score even when their content changes, while lower scores indicate less explicit descriptions after the GDPR.
- Coverage: Approximately 10.6% of EU policies and 13.2% of Global policies failed to cover queries Q2, Q4, and Q5 because they lacked third-party sharing coverage.The purpose of data retention was also infrequently covered in both policy sets.
- Specificity outcomes: A large portion of policies retained the same specificity, while 40% of policies remained fully specified for Q2 across both versions and datasets.Fully specified policies described specific methods of collecting user data and sharing it.
- Specificity outcomes: The EU set had more improving policies and more policies with lower scores across all eight queries, indicating that greater coverage did not guarantee greater transparency.The authors connect lower specificity to attempts to describe more practices comprehensively and generally.
10 Limitations
The authors identify limits in measuring GDPR-driven changes automatically and in evaluating policy presentation and language coverage. These constraints affect the completeness, generalizability, and interpretation of the study’s findings.
- Measurement scope: The automated content analysis does not fully capture GDPR-related efforts such as layered privacy notices.Layered notices may distribute high-level and detailed explanations across varied formats, including multiple pages.
- Language scope: The study covers only English-language privacy policies, trading language coverage for deeper semantic analysis.The authors deliberately avoid keyword-based analysis in favor of existing advanced natural-language techniques.
- Automation: Automated policy selection and analysis are inherently error-prone, so machine-learning errors may affect the specificity and compliance results.The authors note that privacy policies are complex and that human annotators sometimes disagree about their interpretation.
- User evaluation: The user study evaluates policy appearance rather than content, leaving policy readability evolution for future research.The authors call for a more comprehensive user study focused on how much readability has changed.
11 Related Work
Earlier studies generally found more privacy-policy coverage and detail but declining readability and clarity, while post-GDPR research initially remained limited in scale. This paper extends that literature with a large longitudinal analysis of policy evolution.
- Earlier regulatory studies: Studies following HIPAA and GLBA reported more websites with privacy policies and more detailed descriptions, alongside worse readability and clarity.These studies used longitudinal comparisons, keyword investigations, readability metrics, and user studies across earlier regulatory settings.
- User perceptions: Users’ confidence in the label “privacy policy” has been studied as a factor affecting personal actions and responses to government and corporate activity.The cited survey research covered the United States from 2003 through 2015.
- Early GDPR research: Early GDPR studies examined company compliance and policy classification but used small samples, including 14 or 10 privacy policies.Reported issues included unclear language, problematic processing, and insufficient information.
- Large-scale GDPR research: Degeling et al. conducted a large longitudinal study of 6,759 EU websites and found 4.9% growth in sites with privacy policies and widespread updates before May 2018.Their work focused on cookie-consent notices and terminology-based policy analysis rather than the broader automated semantic analysis used here.
12 Takeaways and Conclusions
The GDPR prompted broad changes in privacy policies, especially within the EU, improving coverage, compliance, specificity, and user-facing presentation while also increasing length. Yet policies remain transitional: greater topic coverage does not always provide the detail or transparency required by the GDPR.
- 6,278 policies were compared across pre-GDPR and post-GDPR versions along five dimensions: presentation, textual features, coverage, compliance, and specificity.
- Presentation: User-perceived attractiveness and clarity improved significantly for EU policies, while Global policies showed no significant visual change.
- Coverage: Coverage of privacy categories increased in both sets, especially for previously underrepresented topics, with larger gains among EU policies.
- Compliance and Specificity: The average portion of compliant EU policies exceeded that of Global policies, although coverage did not consistently provide the detail required by the GDPR.In several cases, improved coverage coincided with reduced specificity.
- Conclusions: The GDPR positively affected incorporation of privacy rights and information, but policies remain longer, broader, and short of stable full disclosure and transparency.
A Policy Classifier Architecture
The study uses a convolutional neural-network classifier to distinguish valid privacy policies from invalid web pages. It tokenizes page text, embeds words, applies convolution, ReLU, max-pooling, and a dense layer, achieving high test accuracy on balanced labeled data.
- The classifier determines whether crawled web pages are valid privacy policies using tokenized page text and a neural architecture.
- Architecture: The architecture maps words to embeddings, then applies convolution, ReLU activation, max-pooling, and a fully connected layer.The figure specifies 100-dimensional embeddings, 250 filters, filter size 3, and a dense layer size of 250.
- Training Data: Training used 1,000 valid privacy policies and 1,000 invalid pages, split into 80% training and 20% testing data.
- Evaluation: 99.09% testing accuracy was achieved by the classifier on the labeled validation task.
- User Survey: Figure 13 presents an example step from the user survey used in the presentation analysis.