Chinese clinical named entity recognition with variant neural structures based on BERT methods.
Clinical Named Entity Recognition (CNER) is a critical task which aims to identify and classify clinical terms in electronic medical records. In recent years, deep neural networks have achieved significant success in CNER. However, these methods require high-quality and large-scale labeled clinical...
| Publicado en: | Journal of Biomedical Informatics Vol. 107 |
|---|---|
| Autores principales: | , , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Academic Press Inc.
Jul2020
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=144627565&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 144627565 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Jul2020 vid: 107 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 144627565 144627565 NLM32353595 144627565 10.1016/j.jbi.2020.103422 NLM32353595 144627565 ppct: 1 formats: tig: atl: Chinese clinical named entity recognition with variant neural structures based on BERT methods. aug: au: Li, Xiangyang Zhang, Huan Zhou, Xiao-Hua affil: School of Mathematical Sciences, Peking University, Beijing 100871, China sug: subj: Text Messaging China Comparative Studies Multicenter Studies Evaluation Research Validation Studies ab: Clinical Named Entity Recognition (CNER) is a critical task which aims to identify and classify clinical terms in electronic medical records. In recent years, deep neural networks have achieved significant success in CNER. However, these methods require high-quality and large-scale labeled clinical data, which is challenging and expensive to obtain, especially data on Chinese clinical records. To tackle the Chinese CNER task, we pre-train BERT model on the unlabeled Chinese clinical records, which can leverage the unlabeled domain-specific knowledge. Different layers such as Long Short-Term Memory (LSTM) and Conditional Random Field (CRF) are used to extract the text features and decode the predicted tags respectively. In addition, we propose a new strategy to incorporate dictionary features into the model. Radical features of Chinese characters are used to improve the model performance as well. To the best of our knowledge, our ensemble model outperforms the state of the art models which achieves 89.56% strict F1 score on the CCKS-2018 dataset and 91.60% F1 score on CCKS-2017 dataset. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|