Deep Learning Approaches Outperform Conventional Strategies in De-Identification of German Medical Reports.
One of the major obstacles for research on German medical reports is the lack of de-identified medical corpora. Previous de-identification tasks focused on non-German medical texts, which raised the demand for an in-depth evaluation of de-identification methods on German medical texts. Because of re...
| Published in: | Studies in Health Technology & Informatics Vol. 267; pp. 101 - 110 |
|---|---|
| Main Authors: | , , , |
| Format: | equations & formulas research tables/charts Journal Article |
| Published: |
Sage Publications Inc.
2019
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=138512463&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 138512463 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 09269630 U1V jtl: Studies in Health Technology & Informatics issn: 09269630 maglogo: N pubinfo: dt: 2019 vid: 267 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 138512463 138512463 138512463 10.3233/SHTI190813 138512463 ppf: 101 ppct: 9 formats: tig: atl: Deep Learning Approaches Outperform Conventional Strategies in De-Identification of German Medical Reports. aug: au: RICHTER-PECHANSKI, Phillip AMR, Ali KATUS, Hugo A. DIETERICH, Christoph affil: Section of Bioinformatics and Systems Cardiology, Klaus Tschira Institute for Integrative Computational Cardiology, Heidelberg sug: subj: Deep Learning Electronic Health Records Germany Privacy and Confidentiality Human Germany DICOM Natural Language Processing Machine Learning Methods Data Management Funding Source ab: One of the major obstacles for research on German medical reports is the lack of de-identified medical corpora. Previous de-identification tasks focused on non-German medical texts, which raised the demand for an in-depth evaluation of de-identification methods on German medical texts. Because of remarkable advancements in natural language processing using supervised machine learning methods on limited training data, we evaluated them for the first time on German medical reports using our annotated data set consisting of 113 medical reports from the cardiology domain. We applied state-of-the-art deep learning methods using pre-trained models as input to a bidirectional LSTM network and wellestablished conditional random fields for de-identification of German medical reports. We performed an extensive evaluation for de-identification and multiclass named entity recognition. Using rule based and out of domain machine learning methods as a baseline, the conditional random field improved F2-score from 70 to 93% for de-identification, the neural approach reached 96% in F2-score while keeping balanced precision and recall rates. These results show, that state-of-theart machine learning methods can play a crucial role in de-identification of German medical reports. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|