Accelerated training of bootstrap aggregation-based deep information extraction systems from cancer pathology reports.

Objective: In machine learning, it is evident that the classification of the task performance increases if bootstrap aggregation (bagging) is applied. However, the bagging of deep neural networks takes tremendous amounts of computational resources and training time. The research question that we aim...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Biomedical Informatics Vol. 110
Autores principales: Yoon, Hong-Jun, Klasky, Hilda B., Gounley, John P., Alawad, Mohammed, Gao, Shang, Durbin, Eric B., Wu, Xiao-Cheng, Stroup, Antoinette, Doherty, Jennifer, Coyle, Linda, Penberthy, Lynne, Blair Christian, J., Tourassi, Georgia D.
Formato: research Journal Article
Publicado: Academic Press Inc. Oct2020
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=146481893&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 146481893
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Oct2020
      vid: 110
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        146481893
        146481893
        NLM32919043
        146481893
        10.1016/j.jbi.2020.103564
        NLM32919043
        146481893
      ppct: 1
      formats:
      tig:
        atl: Accelerated training of bootstrap aggregation-based deep information extraction systems from cancer pathology reports.
      aug:
        au:
          Yoon, Hong-Jun
          Klasky, Hilda B.
          Gounley, John P.
          Alawad, Mohammed
          Gao, Shang
          Durbin, Eric B.
          Wu, Xiao-Cheng
          Stroup, Antoinette
          Doherty, Jennifer
          Coyle, Linda
          Penberthy, Lynne
          Blair Christian, J.
          Tourassi, Georgia D.
        affil: Computational Sciences and Engineering Division, Oak Ridge National Laboratory, Oak Ridge, TN 37830, United States of America
      sug:
        subj:
          Neoplasms
          Information Retrieval
          Human
          Computing Methodologies
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Funding Source
      ab: Objective: In machine learning, it is evident that the classification of the task performance increases if bootstrap aggregation (bagging) is applied. However, the bagging of deep neural networks takes tremendous amounts of computational resources and training time. The research question that we aimed to answer in this research is whether we could achieve higher task performance scores and accelerate the training by dividing a problem into sub-problems.Materials and Methods: The data used in this study consist of free text from electronic cancer pathology reports. We applied bagging and partitioned data training using Multi-Task Convolutional Neural Network (MT-CNN) and Multi-Task Hierarchical Convolutional Attention Network (MT-HCAN) classifiers. We split a big problem into 20 sub-problems, resampled the training cases 2,000 times, and trained the deep learning model for each bootstrap sample and each sub-problem-thus, generating up to 40,000 models. We performed the training of many models concurrently in a high-performance computing environment at Oak Ridge National Laboratory (ORNL).Results: We demonstrated that aggregation of the models improves task performance compared with the single-model approach, which is consistent with other research studies; and we demonstrated that the two proposed partitioned bagging methods achieved higher classification accuracy scores on four tasks. Notably, the improvements were significant for the extraction of cancer histology data, which had more than 500 class labels in the task; these results show that data partition may alleviate the complexity of the task. On the contrary, the methods did not achieve superior scores for the tasks of site and subsite classification. Intrinsically, since data partitioning was based on the primary cancer site, the accuracy depended on the determination of the partitions, which needs further investigation and improvement.Conclusion: Results in this research demonstrate that 1. The data partitioning and bagging strategy achieved higher performance scores. 2. We achieved faster training leveraged by the high-performance Summit supercomputer at ORNL.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N