Deep Learning Approaches Outperform Conventional Strategies in De-Identification of German Medical Reports.

One of the major obstacles for research on German medical reports is the lack of de-identified medical corpora. Previous de-identification tasks focused on non-German medical texts, which raised the demand for an in-depth evaluation of de-identification methods on German medical texts. Because of re...

Full description

Bibliographic Details
Published in:Studies in Health Technology & Informatics Vol. 267; pp. 101 - 110
Main Authors: RICHTER-PECHANSKI, Phillip, AMR, Ali, KATUS, Hugo A., DIETERICH, Christoph
Format: equations & formulas research tables/charts Journal Article
Published: Sage Publications Inc. 2019
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=138512463&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 138512463
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        09269630
        U1V
      jtl: Studies in Health Technology & Informatics
      issn: 09269630
      maglogo: N
    pubinfo:
      dt: 2019
      vid: 267
      pid: 344
      pub: Sage Publications Inc.
      place: Thousand Oaks, California
    artinfo:
      ui:
        138512463
        138512463
        138512463
        10.3233/SHTI190813
        138512463
      ppf: 101
      ppct: 9
      formats:
      tig:
        atl: Deep Learning Approaches Outperform Conventional Strategies in De-Identification of German Medical Reports.
      aug:
        au:
          RICHTER-PECHANSKI, Phillip
          AMR, Ali
          KATUS, Hugo A.
          DIETERICH, Christoph
        affil: Section of Bioinformatics and Systems Cardiology, Klaus Tschira Institute for Integrative Computational Cardiology, Heidelberg
      sug:
        subj:
          Deep Learning
          Electronic Health Records Germany
          Privacy and Confidentiality
          Human
          Germany
          DICOM
          Natural Language Processing
          Machine Learning Methods
          Data Management
          Funding Source
      ab: One of the major obstacles for research on German medical reports is the lack of de-identified medical corpora. Previous de-identification tasks focused on non-German medical texts, which raised the demand for an in-depth evaluation of de-identification methods on German medical texts. Because of remarkable advancements in natural language processing using supervised machine learning methods on limited training data, we evaluated them for the first time on German medical reports using our annotated data set consisting of 113 medical reports from the cardiology domain. We applied state-of-the-art deep learning methods using pre-trained models as input to a bidirectional LSTM network and wellestablished conditional random fields for de-identification of German medical reports. We performed an extensive evaluation for de-identification and multiclass named entity recognition. Using rule based and out of domain machine learning methods as a baseline, the conditional random field improved F2-score from 70 to 93% for de-identification, the neural approach reached 96% in F2-score while keeping balanced precision and recall rates. These results show, that state-of-theart machine learning methods can play a crucial role in de-identification of German medical reports.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N