Data augmentation and transfer learning for cross-lingual Named Entity Recognition in the biomedical domain.

Given the increase in production of data for the biomedical field and the unstoppable growth of the internet, the need for Information Extraction (IE) techniques has skyrocketed. Named Entity Recognition (NER) is one of such IE tasks useful for professionals in different areas. There are several set...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 59; no. 2; pp. 665 - 685
Main Authors: Lancheros, Brayan Stiven, Corpas Pastor, Gloria, Mitkov, Ruslan
Format: Article
Published: Springer Nature Jun2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240035&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 185240035
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 59
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        185240035
        10.1007/s10579-024-09738-8
      ppf: 665
      ppct: 20
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 937KB
      tig:
        atl: Data augmentation and transfer learning for cross-lingual Named Entity Recognition in the biomedical domain.
      aug:
        au:
          Lancheros, Brayan Stiven
          Corpas Pastor, Gloria
          Mitkov, Ruslan
        affil:
          https://ror.org/01k2y1055 University of Wolverhampton, Wolverhampton, UK
          https://ror.org/036b2ww28 Universidad de Malaga, IUITLM, Malaga, Spain
          https://ror.org/04f2nsd36 Lancaster University, Lancaster, UK
      su:
        Machine translating
        Artificial intelligence
        Data augmentation
        Data mining
        Spanish language
      sug:
        subj:
          Machine translating
          Artificial intelligence
          Data augmentation
          Data mining
          Spanish language
      keyword:
        Biomedical NER
        Information and Computing Sciences Artificial Intelligence and Image Processing
        Named entity recognition
        Spanish
      ab: Given the increase in production of data for the biomedical field and the unstoppable growth of the internet, the need for Information Extraction (IE) techniques has skyrocketed. Named Entity Recognition (NER) is one of such IE tasks useful for professionals in different areas. There are several settings where biomedical NER is needed, for instance, extraction and analysis of biomedical literature, relation extraction, organisation of biomedical documents, and knowledge-base completion. However, the computational treatment of entities in the biomedical domain has faced a number of challenges including its high cost of annotation, ambiguity, and lack of biomedical NER datasets in languages other than English. These difficulties have hampered data development, affecting both the domain itself and its multilingual coverage. The purpose of this study is to overcome the scarcity of biomedical data for NER in Spanish, for which only two datasets exist, by developing a robust bilingual NER model. Inspired by back-translation, this paper leverages the progress in Neural Machine Translation (NMT) to create a synthetic version of the Colorado Richly Annotated Full-Text (CRAFT) dataset in Spanish. Additionally, a new CRAFT dataset is constructed by replacing 20% of the entities in the original dataset generating a new augmented dataset. We evaluate two training methods: concatenation of datasets and continuous training to assess the transfer learning capabilities of transformers using the newly obtained datasets. The best performing NER system in the development set achieved an F-1 score of 86.39%. The novel methodology proposed in this paper presents the first bilingual NER system and it has the potential to improve applications across under-resourced languages.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N