Comparing Different Methods for Named Entity Recognition in Portuguese Neurology Text.
Electronic Medical Records (EMRs) are written in an unstructured way, often using natural language. Information Extraction (IE) may be used for acquiring knowledge from such texts, including the automatic recognition of meaningful entities, through models for Named Entity Recognition (NER). However,...
| Publicado en: | Journal of Medical Systems Vol. 44; no. 4; pp. 1 - 21 |
|---|---|
| Autores principales: | , , |
| Formato: | algorithm equations & formulas research tables/charts Journal Article |
| Publicado: |
Springer Nature
Apr2020
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=142576170&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 142576170 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01485598 4N0 jtl: Journal of Medical Systems issn: 01485598 maglogo: N pubinfo: dt: Apr2020 vid: 44 iid: 4 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 142576170 142576170 142576170 10.1007/s10916-020-1542-8 142576170 ppf: 1 ppct: 20 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Comparing Different Methods for Named Entity Recognition in Portuguese Neurology Text. aug: au: Lopes, Fábio Teixeira, César Gonçalo Oliveira, Hugo affil: Center for Informatics and Systems, Department of Informatics Engineering, University of Coimbra, Coimbra, Portugal sug: subj: Natural Language Processing Machine Learning Neurology Electronic Health Records Language Memory, Short Term Deep Learning Data Mining Word Processing ab: Electronic Medical Records (EMRs) are written in an unstructured way, often using natural language. Information Extraction (IE) may be used for acquiring knowledge from such texts, including the automatic recognition of meaningful entities, through models for Named Entity Recognition (NER). However, while most work on the previous was made for English, this experience aimed at testing different methods in Portuguese text, more precisely, on the domain of Neurology, and take some conclusions. This paper comprised the comparison between Conditional Random Fields (CRF), bidirectional Long Short-term Memory - Conditional Random Fields (BiLSTM-CRF) and a BiLSTM-CRF with residual learning connections, using not only Portuguese texts from medical journals but also texts from the Coimbra Hospital and Universitary Centre (CHUC) Neurology Service. Furthermore, the performances of BiLSTM-CRF models using word embeddings (WEs) trained with clinical text and WEs trained with general language texts were compared. Deep learning models achieved F1-Scores of nearly 83% and 75%, respectively for relaxed and strict evaluation, on texts extracted from the medical journal. For texts collected from the Hospital, the same achieved F1-Scores of nearly 71% and 62%. This work concludes that deep learning models outperform the shallow learning models and that in-domain WEs get better results than general language WEs, even when the latter are trained with much more text than the former. Furthermore, the results show that it is possible to extract information from Hospital clinical texts with models trained with clinical cases extracted from medical journals, and thus openly available. Nevertheless, such results still require a healthcare technician to check if the information is well extracted. pubtype: Academic Journal doctype: algorithm equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|