Source-linked AI summary

Precision Health Data: Requirements, Challenges and Existing Techniques for Data Security and Privacy

Chandra Thapa, Seyit Camtepe

arXiv:2008.10733v1cs.CRcs.AI

TL;DR

Precision health depends on sensitive data collected across diverse and distributed sources, creating security, privacy, trust, ethical, and regulatory challenges. This paper surveys requirements and challenges, reviews privacy-preserving computation and machine-learning techniques with health projects, and proposes a conceptual compliance-oriented platform model. It identifies data-in-use protection as an active research area and presents combined technical and policy mechanisms for trustworthy handling.

  • Problem

    Precision-health data are sensitive, distributed, and difficult to protect during computation, while existing work does not jointly examine requirements, challenges, and evolving privacy-preserving techniques.

  • Method

    The paper surveys regulations, ethical guidelines, domain-specific requirements, challenges, privacy-preserving machine-learning methods, relevant health projects, and a conceptual platform model.

  • Results

    The survey finds data-at-rest and data-in-transit security comparatively mature, while data-in-use privacy and security remain an active research field addressed by hardware, cryptographic, and differential-privacy techniques.

  • Takeaways & Limitations

    Combining blockchain, cryptography, hardware-based techniques, differential privacy, consent policies, and model-to-data supports compliance, consent management, ethical clearance, and medical innovation.

  • Takeaways & Limitations

    The proposed distributed platform faces heterogeneous computing power, low speed, dropped connections, data heterogeneity, high communication cost, and malicious participation.

Abstract

from arXiv · show

Precision health leverages information from various sources, including omics, lifestyle, environment, social media, medical records, and medical insurance claims to enable personalized care, prevent and predict illness, and precise treatments. It extensively uses sensing technologies (e.g., electronic health monitoring devices), computations (e.g., machine learning), and communication (e.g., interaction between the health data centers). As health data contain sensitive private information, including the identity of patient and carer and medical conditions of the patient, proper care is required at all times. Leakage of these private information affects the personal life, including bullying, high insurance premium, and loss of job due to the medical history. Thus, the security, privacy of and trust on the information are of utmost importance. Moreover, government legislation and ethics committees demand the security and privacy of healthcare data. Herein, in the light of precision health data security, privacy, ethical and regulatory requirements, finding the best methods and techniques for the utilization of the health data, and thus precision health is essential. In this regard, firstly, this paper explores the regulations, ethical guidelines around the world, and domain-specific needs. Then it presents the requirements and investigates the associated challenges. Secondly, this paper investigates secure and privacy-preserving machine learning methods suitable for the computation of precision health data along with their usage in relevant health projects. Finally, it illustrates the best available techniques for precision health data security and privacy with a conceptual system model that enables compliance, ethics clearance, consent management, medical innovations, and developments in the health domain.

1 Introduction

Precision health combines diverse health data with sensing, computation, and communication to support personalized, predictive, prescriptive, and preventive care. Because these data are sensitive, fragmented, and increasingly analyzed across organizations, the paper frames security, privacy, trust, and compliance as central requirements.

  • Precision health integrates omics, lifestyle, environmental, social-media, IoMT, medical-history, pharmaceutical, and insurance-claims data for personalized and preventive care.Examples include longitudinal diabetes-risk prediction and restaurant-hygiene assessment from online reviews.
  • Health data are rapidly expanding through EHRs, medical images, and wearable devices, while remaining decentralized and non-iid.The paper estimates 2,314 exabytes of health data production in 2020.
  • The health-data life cycle spans generation, collection, processing, storage, management, analytics, and inference.Analytics transforms accumulated raw data into insights for evidence-based healthcare, including mutation prediction.
  • Security, privacy, ethical, and legal concerns affect data-at-rest, data-in-transit, and data-in-use because health data expose patient and carer identities and medical conditions.The paper emphasizes consent for data use and reuse.
  • The paper addresses a stated gap by jointly examining regulations, ethical requirements, challenges, privacy-preserving machine learning, health projects, and a conceptual platform model.Its scope includes techniques for computation on isolated and distributed precision-health data.

2 Requirements for precision health data privacy, security and trust

Precision-health data handling requires legally compliant, ethically grounded, secure, private, trustworthy, and user-controlled practices. The paper derives these requirements from regulations, ethical guidance, and domain-specific medical-data needs.

  • Requirements due to law: Legal requirements emphasize lawful, transparent, purpose-limited, minimized, accurate, confidential, integral, and available handling of personal information.The surveyed frameworks include HIPAA, GDPR, and Australian health-data legislation.
  • Requirements due to law: HIPAA combines privacy and security protections with authorized information flow, safeguards, breach notification, and written authorization for secondary data use.Its security standards apply to specified health plans, clearinghouses, and healthcare providers handling electronic health information.
  • Requirements due to law: GDPR grants rights to access, withdraw consent, erase data, restrict processing, and receive breach notification, while requiring explanations of algorithmic outcomes.The regulation applies to organizations processing personal information of EU residents.
  • Requirements due to ethics: Ethical requirements include awareness, user control, trust, and clear ownership of health data.Control includes removal of private data supplied to service providers and other parties.
  • Domain-specific needs: Because health data inform medical decisions, they must be accurate, complete, and precise to support trustworthy care.The paper warns that incorrect medical decisions can harm patients, including potentially causing death.

3 Major challenges in precision health data security, privacy and trust

Precision health faces intertwined challenges in computing security and privacy, consent management, data trustworthiness, and legal and ethical compliance. These difficulties arise from distributed, diverse data and evolving requirements for responsible use.

  • The paper identifies four major challenges: security and privacy during computing, consent management, data trustworthiness, and legal and ethical compliance.
  • Protecting data-in-use is difficult because computation usually requires decryption on potentially untrusted platforms.Trusted platforms, homomorphic encryption, and multi-party computation remain evolving approaches requiring further development or trusted vendors.
  • Static consent cannot accommodate changing environments, requirements, or later reuse of data beyond the originally approved project.Dynamic consent supports updates, varied permissions, revocation, data-linked consent, and communication with participants.
  • Trustworthiness is difficult to maintain because health data are complex, diverse, large, distributed, and generated by many sources.Faulty or improperly configured IoMT devices can produce unreliable health data.
  • Breaches can create trust problems or substantial penalties, including up to €20 million or 4% of annual global revenue under GDPR.
  • Compliance checking is difficult for large, multi-source datasets because laws may be vague and ethics highly conceptual.The paper proposes inspections, privacy-by-design and privacy-by-default, auditing, compliance analytics, and checklists across processing steps.

4 Techniques for PH data privacy and security

Security and privacy are mandatory requirements for precision health data, and no single technology provides a complete solution. The paper therefore frames protection as requiring combinations of techniques across the data lifecycle.

  • Precision health data security and privacy are mandatory because of legal provisions, financial reasons, and trust.
  • A complete security and privacy solution requires combining more than one technology.
  • Compliance analytics calculates and prioritizes risk factors to identify the highest-risk transactions for compliance-risk management.

Data security:

The paper organizes data security around cryptography, blockchain, access control and security analysis, and network security. These techniques address threats to data confidentiality, integrity, physical infrastructure, and transmission.

  • The paper identifies four primary data-security techniques: cryptographic security, blockchain-based security, access control and security analysis, and network security.
  • Cryptographic security: Cryptography protects against interception, tampering, and unauthorized reading while supporting authentication, integrity, confidentiality, and non-repudiation.Only 2.2% of worldwide data-breach incidents in the first half of 2018 involved encrypted data, while health data breaches accounted for 27%.
  • Blockchain-based security: Blockchain uses an immutable, time-stamped distributed ledger to support data sharing without centralized control and strengthen data integrity.
  • Access control and security analysis: Access control and security analysis must protect physical devices and infrastructures that hold sensitive private data.Physical actions, including theft of devices and paper documents, accounted for 11% of total breaches in Verizon’s 2018 report.
  • Network security: Network security protects data-in-transit through protocols such as SSL, TLS, Secure HTTP, IPsec, and SSH.Firewalls and intrusion-prevention systems monitor and control traffic between local networks and untrusted networks.

Data privacy:

The paper presents anonymization and pseudonymization as privacy techniques for reducing risks such as inference and linkability. They modify data or identifiers while supporting selected forms of data use and sharing.

  • Anonymization: Anonymization uses randomization and generalization to reduce links between data and individuals or to dilute identifying attributes.Examples include replacing a street with a region or a specific year with a range of years.
  • Pseudonymization: Pseudonymization replaces an original data-subject attribute with another to reduce linkability between identity and dataset.
  • Pseudonymization: Pseudonymization techniques include secret-key encryption, hashing, keyed hashing, deterministic encryption, tokenization, and masking.Tokenization and masking replace part of the data with random or semi-random tokens while retaining its format and data type.

4.2 Data security and privacy for data-in-use

This section introduces evolving security and privacy-preserving techniques for precision health data-in-use and connects them to healthcare implementations.

  • The section focuses on security and privacy-preserving techniques relevant to precision health data-in-use.It also discusses their implementation in the healthcare domain.
  • Trusted Execution Environment provides one approach for protecting sensitive computations from other processes.It uses isolation and cryptography to increase process security.

Trusted Execution Environment:

Trusted Execution Environments isolate sensitive computations and data using hardware or software mechanisms, while supporting attestation and secure execution. The paper describes TrustZone, Intel SGX, and Keystone Enclave, and notes that TEEs remain exposed to some attacks.

  • ARM TrustZone divides processors into trusted and non-trusted hardware-isolated zones for handling sensitive operations and data.A secure monitor manages context switching between the zones.
  • Intel SGX protects code and data inside encrypted-memory enclaves and supports remote software attestation.Each enclave uses a unique encryption key.
  • Keystone Enclave uses RISC-V hardware capabilities to provide an open-source alternative to proprietary TEE environments.Its openness supports research into enclave vulnerabilities and side-channel attacks, but its software stack is still developing.
  • TEE attacks and defenses continue to evolve, although attackers generally require privileged access or specific conditions.This qualification limits the scope of the stated security protection.
  • Intel SGX was used in KONFIDO to perform decryption, transformations, and encryption of patient summaries inside a TEE.KONFIDO was a Horizon 2020 health-data exchange project.

Homomorphic Encryption:

Homomorphic encryption enables computation over encrypted precision health data without decrypting either the data or results. The paper distinguishes encryption variants, highlights FHE's healthcare relevance, and identifies substantial implementation challenges.

  • Homomorphic encryption supports arbitrary computations on encrypted data without revealing the data or computation results.It enables secure computation on an untrusted platform.
  • Partially homomorphic encryption supports one operation, whereas somewhat homomorphic encryption supports multiple operations within bounded complexity and repetition.The passage identifies RSA, GM, and KTX as examples of partially homomorphic schemes.
  • Fully homomorphic encryption is important for protecting personal information but remains difficult to implement because of computational requirements and overhead.Optimization remains insufficient for general cases.
  • A lattice-based leveled FHE scheme based on RLWE was implemented to protect genomic-data privacy and security in i2b2.i2b2 supports collaborative sharing, integration, standardization, and analysis of clinical research data.

Multiparty Computation:

Multiparty computation enables distributed computation over encrypted shares without decryption or a central trusted party. It reduces computational cost relative to homomorphic encryption but introduces communication, availability, scalability, and residual-leakage challenges.

  • Multiparty computation distributes input shares among distrustful parties that jointly compute functions without revealing their inputs.The final result is shared among the parties, avoiding centralized storage and a trusted third party.
  • MPC has lower computational cost than homomorphic encryption but requires substantial communication and continuously online participants.These requirements arise because parties exchange encrypted data during joint computation.
  • MPC scalability is an issue, and its final output may still leak information about the inputs.The paper therefore states that MPC alone is insufficient for privacy and may require differential privacy or secure enclaves.
  • SODA uses MPC to preserve privacy while processing health big data from multiple distrusting parties, including hospitals and an insurance company.SODA is identified as a Horizon 2020 project.
  • In a leveled FHE scheme, parameters depend on the circuit depth that can be evaluated rather than its size.Leveled FHE computes functions only up to a fixed complexity or level.
  • Table 5 summarizes security and privacy-preserving techniques for data-in-use.
  • Sharemind provides a three-node secure infrastructure that processes privacy-preserving algorithms using additive secret sharing and secure MPC.It has been used for privacy-preserving patient linkage.

Differential Privacy:

Differential privacy (DP) protects health-data privacy by adding noise locally or globally, but its guarantees involve cumulative loss, group-size effects, trusted data holders, and sensitivity–utility trade-offs.

  • Local DP adds noise before collection or computation, whereas global DP adds noise to the final output after computation.
  • Post-processing cannot reduce differential privacy without additional private-database information, while composing mechanisms accumulates total privacy loss.
  • The privacy guarantee deteriorates linearly as group size increases, and group privacy differs from composition privacy.
  • DP assumes trusted initial data holders, although this assumption may fail in practice.
  • Computing global sensitivity while maintaining both privacy and acceptable noise remains difficult, and Laplace noise can alter original answers.
  • Healthcare applications combine DP with MPC, encryption, statistical testing, and distributed deep learning for research and clinical-data analytics.

5 Methods of data storage, computing, and learning

The section contrasts centralized and decentralized data arrangements, data-to-modeler with model-to-data computation, and several learning paradigms relevant to distributed precision-health data.

  • Methods of data storage, computing, and learning: Centralized storage simplifies computation but creates broad consequences if its server fails or is compromised and requires trust in that server.
  • Methods of data storage, computing, and learning: Data-to-modeler gives modelers direct access to sensitive data, conflicting with privacy-by-design and becoming infeasible across organizations.
  • Methods of data storage, computing, and learning: Model-to-data sends code or models to data contributors, where training and validation occur on unseen data before updated results return to the modeler.
  • Methods of data storage, computing, and learning: Centralized, distributed, and decentralized computing differ in where computation and process control are located; distributed systems parallelize processing across multiple servers.
  • Methods of data storage, computing, and learning: Transfer learning reuses pretrained models for related problems, helping when target data are insufficient; medical-image applications include 88.9% otitis-media detection accuracy.
  • Methods of data storage, computing, and learning: Multi-task learning uses related-task training signals as inductive bias to improve generalization across associated tasks.
  • Methods of data storage, computing, and learning: Continuous learning retains prior knowledge while incorporating new data streams or large datasets, but requires ongoing computation and risks catastrophic forgetting.
  • Methods of data storage, computing, and learning: Federated learning trains collaboratively on distributed devices and preserves privacy locally, while SplitNN reduces client computation and communication by transmitting cut-layer activations.

6 Related health projects, their privacy and security measures

Related healthcare projects combine cryptography, distributed computation, consent mechanisms, blockchain, and privacy techniques to support secure data sharing, analytics, and interoperability.

  • Related health projects, their privacy and security measures: MHMD uses encryption and blockchain to support patient-controlled medical-information sharing, with smart contracts, personal data accounts, and dynamic consent.
  • Related health projects, their privacy and security measures: MHMD combines semantic MPC and partial homomorphic encryption for secure computation, with k-anonymity and differential privacy for published datasets.
  • Related health projects, their privacy and security measures: SODA targets practical privacy-preserving analytics on large, multi-dataset healthcare data using multi-party computation and differential privacy.
  • Related health projects, their privacy and security measures: KONFIDO advances cross-border eHealth preservation, exchange, interoperability, and compliance through OpenNCP security components.
  • Related health projects, their privacy and security measures: KONFIDO combines Intel SGX, fully homomorphic encryption, secure exchange technologies, secure processing, identity support, and blockchain-based traceability.
  • Related health projects, their privacy and security measures: SPHN seeks harmonized information systems and data types across Swiss hospitals and research institutes to facilitate nationwide research-data exchange.
  • Related health projects, their privacy and security measures: MedCo enables regulation-compliant medical-data sharing through a hybrid decentralized approach that avoids limitations of fully centralized and fully decentralized systems.
  • Related health projects, their privacy and security measures: San-Shi supports encrypted aggregation and statistical processing for multi-facility clinical research, including joins, filtering, and aggregation.

7 Discussion

The conceptual model governs distributed precision-health computation through policy checks, consent mechanisms, encryption, and no-peek learning, while implementation still faces heterogeneous participants and system constraints.

  • 7 Discussion: The proposed conceptual system model presents techniques for precision-health data security and privacy after reviewing requirements and available methods.
  • 7 Discussion: Because health data reside across hospitals, pharmacies, devices, research centers, and social media, the model uses consent, ethics, and privacy policies plus dynamic and smart consent.
  • 7 Discussion: Policies cross-check program-code and data metadata before computation, including ethics, consent, and environmental privacy requirements.
  • 7 Discussion: Smart contracts describe data handling, secure consent storage, withdrawal, and renewal, while dynamic consent supports different permissions for different data uses.
  • 7 Discussion: No-peek learning uses federated, split, or splitfed learning to analyze siloed data without raw precision-health data leaving their sources.
  • 7 Discussion: The example model combines federated learning, differential privacy, Intel SGX, and encryption across end devices and communication channels.
  • 7 Discussion: Implementation challenges include unequal computing power, low speed, dropped connections, heterogeneous data, communication costs, stragglers, and malicious participation.

8 Conclusion

The paper synthesizes regulatory and ethical requirements for precision health data, identifies major challenges, and surveys techniques for secure, privacy-preserving computation. It highlights model-to-data learning and combined technical and policy measures as practical directions, while noting that data-in-use security remains an active research field.

  • Regulatory and ethical requirements include informed consent, secure processing and transfer, safeguards, confidentiality, integrity, transparency, fairness, ownership, and breach management.
  • Security and privacy for stored and transmitted precision health data present no significant difficulties, whereas protection during data use remains an active research field.
  • The survey discusses trusted execution environments, homomorphic encryption, multi-party computation, differential privacy, and No-peek learning for privacy-preserving precision health computation.No-peek learning is presented as suitable because it uses model-to-data computing on distributed data retained in silos.
  • Reviewed health projects commonly implement blockchain, cryptography, hardware-based techniques, differential privacy, and consent and privacy policies.The paper illustrates these techniques with a policy enforcer in a precision health platform to support compliant data handling.
Loading 2008.10733v1…