Pseudonymization for research data collection: is the juice worth the squeeze?

Background: The collection of data and biospecimens which characterize patients and probands in-depth is a core element of modern biomedical research. Relevant data must be considered highly sensitive and it needs to be protected from unauthorized use and re-identification. In this context, laws, re...

Full description

Bibliographic Details
Published in:BMC Medical Informatics & Decision Making Vol. 19; no. 1
Main Authors: Kohlmayer, Florian, Lautenschläger, Ronald, Prasser, Fabian
Format: Journal Article
Published: BioMed Central 9/4/2019
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=138430628&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 138430628
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14726947
        1CI0
      jtl: BMC Medical Informatics & Decision Making
      issn: 14726947
      maglogo: N
    pubinfo:
      dt: 9/4/2019
      vid: 19
      iid: 1
      pid: 24147
      pub: BioMed Central
    artinfo:
      ui:
        138430628
        138430628
        NLM31484555
        10.1186/s12911-019-0905-x
        NLM31484555
        138430628
      ppct: 1
      formats:
      tig:
        atl: Pseudonymization for research data collection: is the juice worth the squeeze?
      aug:
        au:
          Kohlmayer, Florian
          Lautenschläger, Ronald
          Prasser, Fabian
        affil: Institute of Medical Informatics, Statistics and Epidemiology, University Hospital rechts der Isar, Technical University of Munich, Munich, Germany
      sug:
        subj:
          Nomenclature
          Privacy and Confidentiality
          Research, Medical
          Data Security Standards
          Computer Communication Networks
          Clinical Assessment Tools
          Scales
      ab: Background: The collection of data and biospecimens which characterize patients and probands in-depth is a core element of modern biomedical research. Relevant data must be considered highly sensitive and it needs to be protected from unauthorized use and re-identification. In this context, laws, regulations, guidelines and best-practices often recommend or mandate pseudonymization, which means that directly identifying data of subjects (e.g. names and addresses) is stored separately from data which is primarily needed for scientific analyses.Discussion: When (authorized) re-identification of subjects is not an exceptional but a common procedure, e.g. due to longitudinal data collection, implementing pseudonymization can significantly increase the complexity of software solutions. For example, data stored in distributed databases, need to be dynamically combined with each other, which requires additional interfaces for communicating between the various subsystems. This increased complexity may lead to new attack vectors for intruders. Obviously, this is in contrast to the objective of improving data protection. What is lacking is a standardized process of evaluating and reporting risks, threats and countermeasures, which can be used to test whether integrating pseudonymization methods into data collection systems actually improves upon the degree of protection provided by system designs that simply follow common IT security best practices and implement fine-grained role-based access control models. To demonstrate that the methods used to describe systems employing pseudonymized data management are currently heterogeneous and ad-hoc, we examined the extent to which twelve recent studies address each of the six basic security properties defined by the International Organization for Standardization (ISO) standard 27,000. We show inconsistencies across the studies, with most of them failing to mention one or more security properties.Conclusion: We discuss the degree of privacy protection provided by implementing pseudonymization into research data collection processes. We conclude that (1) more research is needed on the interplay of pseudonymity, information security and data protection, (2) problem-specific guidelines for evaluating and reporting risks, threats and countermeasures should be developed and that (3) future work on pseudonymized research data collection should include the results of such structured and integrated analyses.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N