A comparative study of pretrained language models for long clinical text.

Objective: Clinical knowledge-enriched transformer models (eg, ClinicalBERT) have state-of-the-art results on clinical natural language processing (NLP) tasks. One of the core limitations of these transformer models is the substantial memory consumption due to their full self-attention mechanism, wh...

Full description

Bibliographic Details
Published in:Journal of the American Medical Informatics Association Vol. 30; no. 2; pp. 340 - 348
Main Authors: Li, Yikuan, Wehbe, Ramsey M, Ahmad, Faraz S, Wang, Hanyin, Luo, Yuan
Format: Journal Article
Published: Oxford University Press / USA Feb2023
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=161440058&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 161440058
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Feb2023
      vid: 30
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        161440058
        161440058
        NLM36451266
        10.1093/jamia/ocac225
        NLM36451266
        161440058
      ppf: 340
      ppct: 8
      formats:
      tig:
        atl: A comparative study of pretrained language models for long clinical text.
      aug:
        au:
          Li, Yikuan
          Wehbe, Ramsey M
          Ahmad, Faraz S
          Wang, Hanyin
          Luo, Yuan
        affil: Division of Health and Biomedical Informatics, Department of Preventive Medicine, Feinberg School of Medicine, Northwestern University , Chicago, Illinois, USA
      sug:
      ab: Objective: Clinical knowledge-enriched transformer models (eg, ClinicalBERT) have state-of-the-art results on clinical natural language processing (NLP) tasks. One of the core limitations of these transformer models is the substantial memory consumption due to their full self-attention mechanism, which leads to the performance degradation in long clinical texts. To overcome this, we propose to leverage long-sequence transformer models (eg, Longformer and BigBird), which extend the maximum input sequence length from 512 to 4096, to enhance the ability to model long-term dependencies in long clinical texts.Materials and Methods: Inspired by the success of long-sequence transformer models and the fact that clinical notes are mostly long, we introduce 2 domain-enriched language models, Clinical-Longformer and Clinical-BigBird, which are pretrained on a large-scale clinical corpus. We evaluate both language models using 10 baseline tasks including named entity recognition, question answering, natural language inference, and document classification tasks.Results: The results demonstrate that Clinical-Longformer and Clinical-BigBird consistently and significantly outperform ClinicalBERT and other short-sequence transformers in all 10 downstream tasks and achieve new state-of-the-art results.Discussion: Our pretrained language models provide the bedrock for clinical NLP using long texts. We have made our source code available at https://github.com/luoyuanlab/Clinical-Longformer, and the pretrained models available for public download at: https://huggingface.co/yikuan8/Clinical-Longformer.Conclusion: This study demonstrates that clinical knowledge-enriched long-sequence transformers are able to learn long-term dependencies in long clinical text. Our methods can also inspire the development of other domain-enriched long-sequence transformers.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N