UMLS-based data augmentation for natural language processing of clinical research literature.

Objective: The study sought to develop and evaluate a knowledge-based data augmentation method to improve the performance of deep learning models for biomedical natural language processing by overcoming training data scarcity.Materials and Methods: We extended the easy data augmentation (EDA) method...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Medical Informatics Association Vol. 28; no. 4; pp. 812 - 824
Autores principales: Kang, Tian, Perotte, Adler, Tang, Youlan, Ta, Casey, Weng, Chunhua
Formato: Journal Article
Publicado: Oxford University Press / USA Apr2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=149414964&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 149414964
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Apr2021
      vid: 28
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        149414964
        149414964
        NLM33367705
        10.1093/jamia/ocaa309
        NLM33367705
        149414964
      ppf: 812
      ppct: 12
      formats:
      tig:
        atl: UMLS-based data augmentation for natural language processing of clinical research literature.
      aug:
        au:
          Kang, Tian
          Perotte, Adler
          Tang, Youlan
          Ta, Casey
          Weng, Chunhua
        affil: Department of Biomedical Informatics, Columbia University , New York, New York, USA
      sug:
        subj:
          Research, Medical
          Information Retrieval Methods
          Natural Language Processing
          Unified Medical Language System
          Short Portable Mental Status Questionnaire
          Clinical Assessment Tools
          Scales
      ab: Objective: The study sought to develop and evaluate a knowledge-based data augmentation method to improve the performance of deep learning models for biomedical natural language processing by overcoming training data scarcity.Materials and Methods: We extended the easy data augmentation (EDA) method for biomedical named entity recognition (NER) by incorporating the Unified Medical Language System (UMLS) knowledge and called this method UMLS-EDA. We designed experiments to systematically evaluate the effect of UMLS-EDA on popular deep learning architectures for both NER and classification. We also compared UMLS-EDA to BERT.Results: UMLS-EDA enables substantial improvement for NER tasks from the original long short-term memory conditional random fields (LSTM-CRF) model (micro-F1 score: +5%, + 17%, and +15%), helps the LSTM-CRF model (micro-F1 score: 0.66) outperform LSTM-CRF with transfer learning by BERT (0.63), and improves the performance of the state-of-the-art sentence classification model. The largest gain on micro-F1 score is 9%, from 0.75 to 0.84, better than classifiers with BERT pretraining (0.82).Conclusions: This study presents a UMLS-based data augmentation method, UMLS-EDA. It is effective at improving deep learning models for both NER and sentence classification, and contributes original insights for designing new, superior deep learning approaches for low-resource biomedical domains.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N