Natural Language Processing for Imaging Protocol Assignment: Machine Learning for Multiclass Classification of Abdominal CT Protocols Using Indication Text Data.

A correct protocol assignment is critical to high-quality imaging examinations, and its automation can be amenable to natural language processing (NLP). Assigning protocols for abdominal imaging CT scans is particularly challenging given the multiple organ specific indications and parameters. We com...

Full description

Bibliographic Details
Published in:Journal of Digital Imaging Vol. 35; no. 5; pp. 1120 - 1131
Main Authors: Xavier, Brian Arun, Chen, Po-Hao
Format: research tables/charts Journal Article
Published: Springer Nature Oct2022
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=159758932&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 159758932
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Oct2022
      vid: 35
      iid: 5
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        159758932
        157203246
        159758932
        159758932
        10.1007/s10278-022-00633-8
        159758932
      ppf: 1120
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Natural Language Processing for Imaging Protocol Assignment: Machine Learning for Multiclass Classification of Abdominal CT Protocols Using Indication Text Data.
      aug:
        au:
          Xavier, Brian Arun
          Chen, Po-Hao
        affil: Imaging Institute, Cleveland Clinic Foundation, 9500 Euclid Ave., P34, 44195, Cleveland, OH, USA
      sug:
        subj:
          Natural Language Processing
          Tomography, X-Ray Computed
          Radiography, Abdominal
          Protocols Classification
          Digital Imaging
          Full-Text Databases Evaluation
          Deep Learning
          Learning Methods
          International Classification of Diseases
          Sample Size
          Workflow
          Automation
          Magnetic Resonance Imaging
          Machine Learning
          Algorithms
          Data Mining
          Data Analysis
          Random Forest
          Models, Theoretical
          Descriptive Statistics
      ab: A correct protocol assignment is critical to high-quality imaging examinations, and its automation can be amenable to natural language processing (NLP). Assigning protocols for abdominal imaging CT scans is particularly challenging given the multiple organ specific indications and parameters. We compared conventional machine learning, deep learning, and automated machine learning builder workflows for this multiclass text classification task. A total of 94,501 CT studies performed over 4 years and their assigned protocols were obtained. Text data associated with each study including the ordering provider generated free text study indication and ICD codes were used for NLP analysis and protocol class prediction. The data was classified into one of 11 abdominal CT protocol classes before and after augmentations used to account for imbalances in the class sample sizes. Four machine learning (ML) algorithms, one deep learning algorithm, and an automated machine learning (AutoML) builder were used for the multilabel classification task: Random Forest (RF), Tree Ensemble (TE), Gradient Boosted Tree (GBT), multi-layer perceptron (MLP), Universal Language Model Fine-tuning (ULMFiT), and Google's AutoML builder (Alphabet, Inc., Mountain View, CA), respectively. On the unbalanced dataset, the manually coded algorithms all performed similarly with F1 scores of 0.811 for RF, 0.813 for TE, 0.813 for GBT, 0.828 for MLP, and 0.847 for ULMFiT. The AutoML builder performed better with a F1 score of 0.854. On the balanced dataset, the tree ensemble machine learning algorithm performed the best with an F1 score of 0.803 and a Cohen's kappa of 0.612. AutoML methods took a longer time for completion of NLP model training and evaluation, 4 h and 45 min compared to an average of 51 min for manual methods. Machine learning and natural language processing can be used for the complex multiclass classification task of abdominal imaging CT scan protocol assignment.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N