Source-linked AI summary

Personal Data: Thinking Inside the Box

Hamed Haddadi, Heidi Howard, Amir Chaudhry, Jon Crowcroft, Anil Madhavapeddy, Richard Mortier

arXiv:1501.04737v1cs.CY

TL;DR

Personal data is being widely accumulated with limited consideration for individuals, while existing legal and self-regulatory responses have struggled to protect them. The paper proposes Databox, a trusted personal platform for managing personal data and controlling access, while recognizing practical limits around revocation, device heterogeneity, and cloud dependence.

  • Problem

    Personal data accumulation gives individuals limited control and creates privacy threats, while legal and self-regulatory responses have not kept pace with the changing data ecosystem.

  • Method

    The paper proposes Databox as a trusted platform for personal data management, controlled access by other parties, and incentives for participants.

  • Results

    Databox is proposed as a personal, networked service that collates personal data and makes it available under the individual’s control.

  • Takeaways & Limitations

    A Databox must combine personal data management with controlled access and protections against misuse by software or repeated cross-dataset queries.

  • Takeaways & Limitations

    Revoking access is difficult when third parties copy data, because implementation requires their cooperation and the Databox cannot readily measure downstream information states.

Abstract

from arXiv · show

We propose there is a need for a technical platform enabling people to engage with the collection, management and consumption of personal data; and that this platform should itself be personal, under the direct control of the individual whose data it holds. In what follows, we refer to this platform as the Databox, a personal, networked service that collates personal data and can be used to make those data available. While your Databox is likely to be a virtual platform, in that it will involve multiple devices and services, at least one instance of it will exist in physical form such as on a physical form-factor computing device with associated storage and networking, such as a home hub.

INTRODUCTION

Personal data is being accumulated widely for advertising and other purposes, with limited consideration for individuals. Existing regulation, self-regulation, and user-control platforms have not adequately addressed the resulting trust and control problems.

  • Online services, advertisers, and governments are accumulating personal data at scale, largely with minimal consideration for the individuals concerned.
  • Regulatory frameworks have struggled to keep pace, while Do Not Track remained ineffective, with only 20 advertising services reportedly respecting its headers.
  • Personal-data startups seek to give users explicit control and enable metered access for advertisers and content providers.
  • Cloud-based personal-data services require users to trust providers, infrastructure companies, authorities, and possible collusion among cloud services.
  • Privacy attitudes vary substantially: Westin classified respondents as 16% unconcerned, 24% fundamentalist, and 60% pragmatic.

WHY DO WE NEED A DATABOX?

Personal data is fragmented across services, exposed to privacy and market-control risks, and difficult for individuals to govern. A Databox would provide a personal point of presence supporting legibility, agency, and negotiability while enabling decentralized applications.

  • Cloud data silos, leakages, opaque inferences, and third-party data trading create a need to index and control personal-data portfolios.
  • A Databox could address privacy threats from third-party websites and data aggregators, including governments, advertisers, and credit-scoring companies.
  • A decentralized Databox platform could let developers combine data from multiple silos while reducing users’ and software vendors’ dependence on dominant platforms.
  • Potential applications include privacy-preserving advertising, market research, health applications, Quantified Self, and personal archives.
  • The proposed point of presence supports legibility, agency, and negotiability: understanding data use, controlling access, and managing evolving data relationships.
  • Because users differ in their engagement with data, negotiable policies must support interaction and change over time.

WHAT IS A DATABOX?

A Databox is a trusted platform for managing personal data at rest, controlling access by other parties, and supporting incentives for participants.

  • A Databox must provide trusted personal-data management, controlled access for external users, and incentives for all parties.

Trusted Platform

Trust in a Databox depends on reliable infrastructure, user intervention over collection and sharing, and pervasive logging. These capabilities help users and auditors verify operation and investigate unforeseen events.

  • The platform’s usefulness requires reliable infrastructure alongside straightforward user control over its operations.
  • A Databox must remain consistently available while allowing users to intervene in data collection and sharing when automated policies have unforeseen consequences.
  • Pervasive logging and associated tools allow users and third-party auditors to build trust and track unforeseen events.

Controlled Access

The Databox is intended to provide fine-grained, selective access to personal data rather than merely collect it. It must also support per-case access expiry and revocation, although enforcing this is difficult when third parties copy data.

  • Controlled Access: The Databox must make personal data selectively query-able, giving users fine-grained control over what third parties can access.It may also support privacy-preserving analytics such as differential privacy and homomorphic encryption.
  • Controlled Access: Access periods must be controllable and previously granted access revocable on a per-case basis.Local processing makes revocation relatively straightforward, whereas copied data requires third-party cooperation.
  • Controlled Access: Measuring the impact of releasing a datum is difficult because the Databox cannot readily track all current and future information states of potential third parties.

Data Management

The Databox should let users interact with and reflect on stored data so they can make more informed decisions. It should also support correcting, deleting, and potentially forgetting data that are inaccurate or no longer relevant.

  • Data Management: Users must be able to interact with and reflect upon Databox data to make more informed decisions about their behaviour or delegated control.
  • Data Management: The Databox should let users edit and delete data when inaccurate data or inferred information has been uncovered and distributed.
  • Data Management: The Databox may need to forget data that are no longer relevant, countering the usual digital tendency toward perfect records.

Supporting Incentives

Controlled access creates a need for incentives that let users restrict data access without necessarily losing services. The Databox could also reduce organisations’ exposure to sensitive data while preserving their ability to query it.

  • Supporting Incentives: Users who deny third-party access to their data might otherwise lose access to services, creating a need for alternative payment arrangements.
  • Supporting Incentives: The Databox must support payments alongside data flows so users can pay with money instead of granting access to their data.This could be provided through an app-store-like mechanism.
  • Supporting Incentives: Organisations could reduce their exposure to sensitive data by letting data subjects retain control while continuing to access and query that data.This is particularly relevant to international organisations facing multiple legal frameworks.

WHAT’S IN THE DATABOX?

Personal data in a Databox is heterogeneous, distributed across many sources and devices, and difficult to assemble and maintain. A practical system must therefore handle varied formats, sizes, authentication standards, and device capabilities without copying everything everywhere.

  • WHAT’S IN THE DATABOX?: Personal data spans mixed formats, highly variable sizes, diverse authentication standards, and inconsistent processing tools, making assembly and maintenance time-consuming.
  • WHAT’S IN THE DATABOX?: A single partial personal footprint can exceed 55GB and include data collected from different sources over more than 10 years.
  • WHAT’S IN THE DATABOX?: The footprint can include communications and financial records, including email, messaging, phone calls, SMS, bank statements, credit cards, and housing contracts.
  • WHAT’S IN THE DATABOX?: It can also include family, individual, and social-network data, such as photographs, location traces, calendars, health records, sleep tracking, and accounts on current or defunct services.
  • WHAT’S IN THE DATABOX?: Copying all personal data to every device is not viable because devices differ in capacity and capability.A more privacy-minimising design gives complete indexes to only one or two strongly trusted devices while other devices issue limited queries.

WHERE IS MY DATABOX?

The Databox remains an unresolved proposal because existing approaches face technical and social barriers involving connectivity, trust, complexity, usability, cost, and adoption. Its design must manage personal data that is context-dependent and often shared among multiple individuals.

  • Open challenges: Existing personal-data systems have not successfully addressed the Databox’s fundamental technical and social barriers.The paper identifies unresolved challenges rather than presenting a completed system.
  • Connectivity: Cloud-based connectivity can mitigate firewall and middlebox problems, but introduces trust and cost issues.Earlier approaches use cloud servers because they are assumed to be more reachable than devices at the network edge.
  • Trust and availability: A Databox must protect against breaches and malicious software while remaining reliably available and allowing users to intervene in automated data-sharing actions.Trust includes protection against repeated queries or cross-dataset inference, alongside confidence in the software itself.
  • Complexity and usability: Personal-data management is difficult because user preferences are socially derived, context dependent, and difficult to express in machine-readable form.The system must make this complexity understandable and usable for end-users.
  • Shared data: Many data are inherently shared across individuals, making ownership and permission authority unclear.Examples include domestic energy data and email processed by a recipient’s cloud provider.
  • Adoption: Adoption also depends on acceptable operating and access costs, early-adopter uptake, and the development of trust that can spread to others.The paper notes that users may expect private records to be retained while becoming less visible to others over time.

WHEN CAN I HAVE MY DATABOX?

The Databox’s availability depends on both affordable technology and sufficient demand, but demand remains uncertain and may require policy support. The proposal therefore combines privacy-by-design with technological development and real-world studies of how people value privacy services.

  • Cost and demand: Databox adoption requires sufficiently high demand and sufficiently low cost.The authors pursue lower costs through associated technologies including Nymote, Mirage, Irmin, and Signpost.
  • Demand and regulation: Demand may remain low until people experience data-breach harm, while governments might regulate before clear popular demand emerges.The paper discusses shifting from informed consent toward a consumer-protection model as one possible response.
  • Design approach: The proposed design approach is privacy by design, while its successful implementation requires policy and technology to co-evolve and must include social considerations.The paper rejects assuming that either everything should be public or everything hidden is desirable.
  • Evaluation: Trial deployments and in-the-wild studies can examine willingness to pay for services and marginal willingness to pay for privacy.The paper links these studies to privacy’s negotiation through collective dynamics.
Loading 1501.04737v1…