Healthcare Text Classification System and its Performance Evaluation: A Source of Better Intelligence by Characterizing Healthcare Text.
A machine learning (ML)-based text classification system has several classifiers. The performance evaluation (PE) of the ML system is typically driven by the training data size and the partition protocols used. Such systems lead to low accuracy because the text classification systems lack the abilit...
| Published in: | Journal of Medical Systems Vol. 42; no. 5; pp. 1 - 2 |
|---|---|
| Main Authors: | , , |
| Format: | equations & formulas research tables/charts Journal Article |
| Published: |
Springer Nature
May2018
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=129370619&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 129370619 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01485598 4N0 jtl: Journal of Medical Systems issn: 01485598 maglogo: N pubinfo: dt: May2018 vid: 42 iid: 5 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 129370619 129370619 129370619 10.1007/s10916-018-0941-6 129370619 ppf: 1 ppct: 1 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Healthcare Text Classification System and its Performance Evaluation: A Source of Better Intelligence by Characterizing Healthcare Text. aug: au: Srivastava, Saurabh Kumar Singh, Sandeep Kumar Suri, Jasjit S. affil: Department of Computer Science & Engineering, JIIT, Noida, India sug: subj: Medical Literature Classification Machine Learning Human Neural Networks (Computer) ROC Curve Data Mining Experimental Studies Social Media ab: A machine learning (ML)-based text classification system has several classifiers. The performance evaluation (PE) of the ML system is typically driven by the training data size and the partition protocols used. Such systems lead to low accuracy because the text classification systems lack the ability to model the input text data in terms of noise characteristics. This research study proposes a concept of misrepresentation ratio (MRR) on input healthcare text data and models the PE criteria for validating the hypothesis. Further, such a novel system provides a platform to amalgamate several attributes of the ML system such as: data size, classifier type, partitioning protocol and percentage MRR. Our comprehensive data analysis consisted of five types of text data sets (TwitterA, WebKB4, Disease, Reuters (R8), and SMS); five kinds of classifiers (support vector machine with linear kernel (SVM-L), MLP-based neural network, AdaBoost, stochastic gradient descent and decision tree); and five types of training protocols (<italic>K2, K4, K5, K10</italic> and <italic>JK</italic>). Using the decreasing order of MRR, our ML system demonstrates the mean classification accuracies as: 70.13 ± 0.15%, 87.34 ± 0.06%, 93.73 ± 0.03%, 94.45 ± 0.03% and 97.83 ± 0.01%, respectively, using all the classifiers and protocols. The corresponding AUC is 0.98 for SMS data using Multi-Layer Perceptron (MLP) based neural network. All the classifiers, the best accuracy of 91.84 ± 0.04% is shown to be of MLP-based neural network and this is 6% better over previously published. Further we observed that as MRR decreases, the system robustness increases and validated by standard deviations. The overall text system accuracy using all data types, classifiers, protocols is 89%, thereby showing the entire ML system to be novel, robust and unique. The system is also tested for stability and reliability. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|