Is Multiclass Automatic Text De-Identification Worth the Effort?

Objectives: Automatic de-identification to remove protected health information (PHI) from clinical text can use a "binary" model that replaces redacted text with a generic tag (e.g., "<PHI>"), or can use a "multiclass" model that retains more class information (e.g., "<Phone Number>"). Binary models...

Descripción completa

Detalles Bibliográficos
Publicado en:Methods of Information in Medicine Vol. 57; no. 4; pp. 177 - 185
Autores principales: Duy Duc An Bui, Redden, David T., Cimino, James J., Bui, Duy Duc An
Formato: Journal Article
Publicado: Thieme Medical Publishing Inc. 2018
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=131965539&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 131965539
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        00261270
        W7M
      jtl: Methods of Information in Medicine
      issn: 00261270
      maglogo: N
    pubinfo:
      dt: 2018
      vid: 57
      iid: 4
      pid: 2811
      pub: Thieme Medical Publishing Inc.
      place: New York, New York
    artinfo:
      ui:
        131965539
        131965539
        NLM30919392
        10.3414/ME18-01-0017
        NLM30919392
        131965539
      ppf: 177
      ppct: 8
      formats:
      tig:
        atl: Is Multiclass Automatic Text De-Identification Worth the Effort?
      aug:
        au:
          Duy Duc An Bui
          Redden, David T.
          Cimino, James J.
          Bui, Duy Duc An
        affil: Informatics Institute, University of Alabama at Birmingham, Birmingham, AL, USA
      sug:
        subj:
          Algorithms
          False Positive Results
          Impact of Events Scale
          Scales
      ab: Objectives: Automatic de-identification to remove protected health information (PHI) from clinical text can use a "binary" model that replaces redacted text with a generic tag (e.g., "<PHI>"), or can use a "multiclass" model that retains more class information (e.g., "<Phone Number>"). Binary models are easier to develop, but result in text that is potentially less informative. We investigated whether building a multiclass de-identification is worth the extra effort.Methods: Using the 2014 i2b2 dataset, we compared the accuracy and impact on document readability of two models. In the first experiment, we generated one binary and two multiclass versions trained with the same machine-learning algorithm Conditional Random Field (CRF). Accuracy (recall, precision, f-score) and secondary metrics (e.g, training time, testing time, minimum memory required) were measured. In the second experiment, three reviewers accessed the readability of two redacted documents using the binary and multiclass methods. We estimated a pooled Kappa to estimate the inter-rater agreement.Results: The multiclass model did not demonstrate a clear accuracy advantage, with lower recall (-1.9%) and only slightly better precision (+0.6%), despite requiring additional computing resources. Three raters reached a very high agreement (Kappa = 0.975, 95% Confidence Interval (0.946, 1.00), p < 0.0001) that both binary and multiclass models have the same impact on document readability.Conclusions: This study suggests that the development of more sophisticated classification of PHI may not be worth the effort in terms of both system accuracy and the usefulness of the output.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N