A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories.

This study examines the effectiveness of state-of-the-art supervised machine learning methods in conjunction with different feature types for the task of automatic annotation of fragments of clinical text based on codebooks with a large number of categories. We used a collection of motivational inte...

Full description

Bibliographic Details
Published in:Journal of Biomedical Informatics Vol. 62; pp. 21 - 32
Main Authors: Hasan, Mehedi, Kotov, Alexander, Idalski Carcone, April, Dong, Ming, Naar, Sylvie, Brogan Hartlieb, Kathryn, Carcone, April, Hartlieb, Kathryn Brogan
Format: research Journal Article
Published: Academic Press Inc. Aug2016
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=117443036&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 117443036
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Aug2016
      vid: 62
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        117443036
        117443036
        NLM27185608
        117443036
        10.1016/j.jbi.2016.05.004
        NLM27185608
        PMC4987168 [Available on 08/01/17]
        117443036
      ppf: 21
      ppct: 11
      formats:
      tig:
        atl: A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories.
      aug:
        au:
          Hasan, Mehedi
          Kotov, Alexander
          Idalski Carcone, April
          Dong, Ming
          Naar, Sylvie
          Brogan Hartlieb, Kathryn
          Carcone, April
          Hartlieb, Kathryn Brogan
        affil: Department of Computer Science, Wayne State University, 5057 Woodward Ave, Detroit, MI 48202, USA
      sug:
        subj:
          Decision Trees
          Data Curation Methods
          Probability
          Semantics
          Human
          Funding Source
      ab: This study examines the effectiveness of state-of-the-art supervised machine learning methods in conjunction with different feature types for the task of automatic annotation of fragments of clinical text based on codebooks with a large number of categories. We used a collection of motivational interview transcripts consisting of 11,353 utterances, which were manually annotated by two human coders as the gold standard, and experimented with state-of-art classifiers, including Naïve Bayes, J48 Decision Tree, Support Vector Machine (SVM), Random Forest (RF), AdaBoost, DiscLDA, Conditional Random Fields (CRF) and Convolutional Neural Network (CNN) in conjunction with lexical, contextual (label of the previous utterance) and semantic (distribution of words in the utterance across the Linguistic Inquiry and Word Count dictionaries) features. We found out that, when the number of classes is large, the performance of CNN and CRF is inferior to SVM. When only lexical features were used, interview transcripts were automatically annotated by SVM with the highest classification accuracy among all classifiers of 70.8%, 61% and 53.7% based on the codebooks consisting of 17, 20 and 41 codes, respectively. Using contextual and semantic features, as well as their combination, in addition to lexical ones, improved the accuracy of SVM for annotation of utterances in motivational interview transcripts with a codebook consisting of 17 classes to 71.5%, 74.2%, and 75.1%, respectively. Our results demonstrate the potential of using machine learning methods in conjunction with lexical, semantic and contextual features for automatic annotation of clinical interview transcripts with near-human accuracy.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N