Biomedical named entity recognition using BERT in the machine reading comprehension framework.
Recognition of biomedical entities from literature is a challenging research focus, which is the foundation for extracting a large amount of biomedical knowledge existing in unstructured texts into structured formats. Using the sequence labeling framework to implement biomedical named entity recogni...
| Published in: | Journal of Biomedical Informatics Vol. 118 |
|---|---|
| Main Authors: | , , , , , |
| Format: | research Journal Article |
| Published: |
Academic Press Inc.
Jun2021
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=150696625&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 150696625 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Jun2021 vid: 118 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 150696625 150696625 NLM33965638 150696625 10.1016/j.jbi.2021.103799 NLM33965638 150696625 ppct: 1 formats: tig: atl: Biomedical named entity recognition using BERT in the machine reading comprehension framework. aug: au: Sun, Cong Yang, Zhihao Wang, Lei Zhang, Yin Lin, Hongfei Wang, Jian affil: School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China sug: subj: Semantics Readability Cognition Data Mining Comparative Studies Multicenter Studies Human ab: Recognition of biomedical entities from literature is a challenging research focus, which is the foundation for extracting a large amount of biomedical knowledge existing in unstructured texts into structured formats. Using the sequence labeling framework to implement biomedical named entity recognition (BioNER) is currently a conventional method. This method, however, often cannot take full advantage of the semantic information in the dataset, and the performance is not always satisfactory. In this work, instead of treating the BioNER task as a sequence labeling problem, we formulate it as a machine reading comprehension (MRC) problem. This formulation can introduce more prior knowledge utilizing well-designed queries, and no longer need decoding processes such as conditional random fields (CRF). We conduct experiments on six BioNER datasets, and the experimental results demonstrate the effectiveness of our method. Our method achieves state-of-the-art (SOTA) performance on the BC4CHEMD, BC5CDR-Chem, BC5CDR-Disease, NCBI-Disease, BC2GM and JNLPBA datasets, achieving F1-scores of 92.92%, 94.19%, 87.83%, 90.04%, 85.48% and 78.93%, respectively. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|