Comparing information extraction techniques for low-prevalence concepts: The case of insulin rejection by patients.

Objective: To comparatively evaluate a range of Natural Language Processing (NLP) approaches for Information Extraction (IE) of low-prevalence concepts in clinical notes on the example of decline of insulin therapy recommendation by patients.Materials and Methods: We evaluated the accuracy of detect...

Full description

Bibliographic Details
Published in:Journal of Biomedical Informatics Vol. 99
Main Authors: Malmasi, Shervin, Ge, Wendong, Hosomura, Naoshi, Turchin, Alexander
Format: Journal Article
Published: Academic Press Inc. Nov2019
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=139507190&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 139507190
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Nov2019
      vid: 99
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        139507190
        139507190
        NLM31618679
        10.1016/j.jbi.2019.103306
        NLM31618679
        139507190
      ppct: 1
      formats:
      tig:
        atl: Comparing information extraction techniques for low-prevalence concepts: The case of insulin rejection by patients.
      aug:
        au:
          Malmasi, Shervin
          Ge, Wendong
          Hosomura, Naoshi
          Turchin, Alexander
        affil: Division of Endocrinology, Brigham and Women's Hospital, Boston, MA, USA
      sug:
        subj:
          Treatment Refusal Statistics and Numerical Data
          Natural Language Processing
          Insulin Therapeutic Use
          Data Mining Methods
          User-Computer Interface
          Diabetes Mellitus Drug Therapy
          Hypoglycemic Agents Therapeutic Use
          Barthel Index
          Scales
      ab: Objective: To comparatively evaluate a range of Natural Language Processing (NLP) approaches for Information Extraction (IE) of low-prevalence concepts in clinical notes on the example of decline of insulin therapy recommendation by patients.Materials and Methods: We evaluated the accuracy of detection of documentation of decline of insulin therapy by patients using sentence-level naïve Bayes, logistic regression and support vector machine (SVM)-based classification (with and without SMOTE oversampling), token-level sequence labelling using conditional random fields (CRFs), uni- and bi-directional recurrent neural network (RNN) models with GRU and LSTM cells, and rule-based detection using Canary platform. All models were trained using the same manually annotated 50,046-document training set and evaluated on the same 1501-document held-out set. Hyperparameter optimization was performed using 10-fold cross-validation.Results: At the sentence level, prevalence of documentation of decline of insulin therapy by patients was 0.02% in both training and held-out sets. Naïve Bayes and logistic regression models did not achieve F1 score ≥ 0.5 on the training set and were not further evaluated. Among the other models, evaluation against the held-out test set showed that SVM identified decline of insulin therapy by patients with F1 score of 0.61, CRF with F1 of 0.51, RNN with F1 of 0.67 and Canary rule-based model with F1 of 0.97.Conclusions: Identification of low-prevalence concepts can present challenges in medical language processing. Rule-based systems that include the designer's background knowledge of language may be able to achieve higher accuracy under these circumstances.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N