A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories.
This study examines the effectiveness of state-of-the-art supervised machine learning methods in conjunction with different feature types for the task of automatic annotation of fragments of clinical text based on codebooks with a large number of categories. We used a collection of motivational inte...
| Published in: | Journal of Biomedical Informatics Vol. 62; pp. 21 - 32 |
|---|---|
| Main Authors: | , , , , , , , |
| Format: | research Journal Article |
| Published: |
Academic Press Inc.
Aug2016
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=117443036&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 117443036 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Aug2016 vid: 62 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 117443036 117443036 NLM27185608 117443036 10.1016/j.jbi.2016.05.004 NLM27185608 PMC4987168 [Available on 08/01/17] 117443036 ppf: 21 ppct: 11 formats: tig: atl: A study of the effectiveness of machine learning methods for classification of clinical interview fragments into a large number of categories. aug: au: Hasan, Mehedi Kotov, Alexander Idalski Carcone, April Dong, Ming Naar, Sylvie Brogan Hartlieb, Kathryn Carcone, April Hartlieb, Kathryn Brogan affil: Department of Computer Science, Wayne State University, 5057 Woodward Ave, Detroit, MI 48202, USA sug: subj: Decision Trees Data Curation Methods Probability Semantics Human Funding Source ab: This study examines the effectiveness of state-of-the-art supervised machine learning methods in conjunction with different feature types for the task of automatic annotation of fragments of clinical text based on codebooks with a large number of categories. We used a collection of motivational interview transcripts consisting of 11,353 utterances, which were manually annotated by two human coders as the gold standard, and experimented with state-of-art classifiers, including Naïve Bayes, J48 Decision Tree, Support Vector Machine (SVM), Random Forest (RF), AdaBoost, DiscLDA, Conditional Random Fields (CRF) and Convolutional Neural Network (CNN) in conjunction with lexical, contextual (label of the previous utterance) and semantic (distribution of words in the utterance across the Linguistic Inquiry and Word Count dictionaries) features. We found out that, when the number of classes is large, the performance of CNN and CRF is inferior to SVM. When only lexical features were used, interview transcripts were automatically annotated by SVM with the highest classification accuracy among all classifiers of 70.8%, 61% and 53.7% based on the codebooks consisting of 17, 20 and 41 codes, respectively. Using contextual and semantic features, as well as their combination, in addition to lexical ones, improved the accuracy of SVM for annotation of utterances in motivational interview transcripts with a codebook consisting of 17 classes to 71.5%, 74.2%, and 75.1%, respectively. Our results demonstrate the potential of using machine learning methods in conjunction with lexical, semantic and contextual features for automatic annotation of clinical interview transcripts with near-human accuracy. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|