Prediction of cause of death from forensic autopsy reports using text classification techniques: A comparative study.
Objectives: Automatic text classification techniques are useful for classifying plaintext medical documents. This study aims to automatically predict the cause of death from free text forensic autopsy reports by comparing various schemes for feature extraction, term weighing or feature value represe...
| Publicado en: | Journal of Forensic & Legal Medicine Vol. 57; pp. 41 - 51 |
|---|---|
| Autores principales: | , , , , |
| Formato: | research Journal Article |
| Publicado: |
Elsevier B.V.
Jul2018
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=129752604&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 129752604 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 1752928X 30RS jtl: Journal of Forensic & Legal Medicine issn: 1752928X maglogo: N pubinfo: dt: Jul2018 vid: 57 pid: 467 pub: Elsevier B.V. place: New York, New York artinfo: ui: 129752604 129752604 NLM29801951 129752604 10.1016/j.jflm.2017.07.001 NLM29801951 129752604 ppf: 41 ppct: 10 formats: tig: atl: Prediction of cause of death from forensic autopsy reports using text classification techniques: A comparative study. aug: au: Mujtaba, Ghulam Shuib, Liyana Raj, Ram Gopal Rajandram, Retnagowri Shaikh, Khairunisa affil: Department of Information Systems, Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur, Malaysia sug: subj: Nomenclature Documentation Classification Autopsy Cause of Death Natural Language Processing Human Validation Studies Comparative Studies Evaluation Research Multicenter Studies ab: Objectives: Automatic text classification techniques are useful for classifying plaintext medical documents. This study aims to automatically predict the cause of death from free text forensic autopsy reports by comparing various schemes for feature extraction, term weighing or feature value representation, text classification, and feature reduction.Methods: For experiments, the autopsy reports belonging to eight different causes of death were collected, preprocessed and converted into 43 master feature vectors using various schemes for feature extraction, representation, and reduction. The six different text classification techniques were applied on these 43 master feature vectors to construct a classification model that can predict the cause of death. Finally, classification model performance was evaluated using four performance measures i.e. overall accuracy, macro precision, macro-F-measure, and macro recall.Results: From experiments, it was found that that unigram features obtained the highest performance compared to bigram, trigram, and hybrid-gram features. Furthermore, in feature representation schemes, term frequency, and term frequency with inverse document frequency obtained similar and better results when compared with binary frequency, and normalized term frequency with inverse document frequency. Furthermore, the chi-square feature reduction approach outperformed Pearson correlation, and information gain approaches. Finally, in text classification algorithms, support vector machine classifier outperforms random forest, Naive Bayes, k-nearest neighbor, decision tree, and ensemble-voted classifier.Conclusion: Our results and comparisons hold practical importance and serve as references for future works. Moreover, the comparison outputs will act as state-of-art techniques to compare future proposals with existing automated text classification techniques. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|