Biomedical named entity recognition using BERT in the machine reading comprehension framework.

Recognition of biomedical entities from literature is a challenging research focus, which is the foundation for extracting a large amount of biomedical knowledge existing in unstructured texts into structured formats. Using the sequence labeling framework to implement biomedical named entity recogni...

Full description

Bibliographic Details
Published in:Journal of Biomedical Informatics Vol. 118
Main Authors: Sun, Cong, Yang, Zhihao, Wang, Lei, Zhang, Yin, Lin, Hongfei, Wang, Jian
Format: research Journal Article
Published: Academic Press Inc. Jun2021
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=150696625&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 150696625
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Jun2021
      vid: 118
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        150696625
        150696625
        NLM33965638
        150696625
        10.1016/j.jbi.2021.103799
        NLM33965638
        150696625
      ppct: 1
      formats:
      tig:
        atl: Biomedical named entity recognition using BERT in the machine reading comprehension framework.
      aug:
        au:
          Sun, Cong
          Yang, Zhihao
          Wang, Lei
          Zhang, Yin
          Lin, Hongfei
          Wang, Jian
        affil: School of Computer Science and Technology, Dalian University of Technology, Dalian 116024, China
      sug:
        subj:
          Semantics
          Readability
          Cognition
          Data Mining
          Comparative Studies
          Multicenter Studies
          Human
      ab: Recognition of biomedical entities from literature is a challenging research focus, which is the foundation for extracting a large amount of biomedical knowledge existing in unstructured texts into structured formats. Using the sequence labeling framework to implement biomedical named entity recognition (BioNER) is currently a conventional method. This method, however, often cannot take full advantage of the semantic information in the dataset, and the performance is not always satisfactory. In this work, instead of treating the BioNER task as a sequence labeling problem, we formulate it as a machine reading comprehension (MRC) problem. This formulation can introduce more prior knowledge utilizing well-designed queries, and no longer need decoding processes such as conditional random fields (CRF). We conduct experiments on six BioNER datasets, and the experimental results demonstrate the effectiveness of our method. Our method achieves state-of-the-art (SOTA) performance on the BC4CHEMD, BC5CDR-Chem, BC5CDR-Disease, NCBI-Disease, BC2GM and JNLPBA datasets, achieving F1-scores of 92.92%, 94.19%, 87.83%, 90.04%, 85.48% and 78.93%, respectively.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N