Lessons learned from data mining of WHO mortality database.
Objectives: The objectives of this research were to test the ability of classification algorithms to predict the cause of death in the mortality data with unknown causes, to find association between common causes of death, to identify groups of countries based on their common causes of death, and to...
| Publicado en: | Methods of Information in Medicine Vol. 50; no. 4; pp. 380 - 386 |
|---|---|
| Autores principales: | , |
| Formato: | research Journal Article |
| Publicado: |
Thieme Medical Publishing Inc.
2011
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=108192011&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 108192011 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 00261270 W7M jtl: Methods of Information in Medicine issn: 00261270 maglogo: N pubinfo: dt: 2011 vid: 50 iid: 4 pid: 2811 pub: Thieme Medical Publishing Inc. place: New York, New York artinfo: ui: 108192011 108192011 NLM21691674 2011241185 10.3414/ME10-02-0019 NLM21691674 108192011 ppf: 380 ppct: 6 formats: tig: atl: Lessons learned from data mining of WHO mortality database. aug: au: Paoin W Paoin, W affil: Faculty of Medicine, Thammasat University, Rangsit Campus, Paholyotin Road, Pathumthani 12120, Thailand sug: subj: Data Mining Methods Resource Databases Medical Informatics Administration Mortality World Health Organization Algorithms Cause of Death Chi Square Test Cluster Analysis Management Information Systems Decision Trees ab: Objectives: The objectives of this research were to test the ability of classification algorithms to predict the cause of death in the mortality data with unknown causes, to find association between common causes of death, to identify groups of countries based on their common causes of death, and to extract knowledge gained from data mining of the World Health Organization mortality database.Methods: The WEKA software version 3.5.3 was used for classification, clustering and association analysis of the World Health Organization mortality database which contained 1,109,537 records. Three major steps were performed: Step 1 - preprocessing of data to convert all records into suitable formats for each type of analysis algorithm; Step 2 - analyzing data using the C4.5 decision tree and Naïve Bayes classification algorithm, K-means clustering algorithm and Apriori association analysis algorithm; Step 3 - interpretation of results and hypothesis testing after clustering analysis.Results: Using a C4.5 decision tree classifier to predict cause of death, we obtained 440 leaf nodes that correctly classify death instances with an accuracy of 40.06%. Naïve Bayes classification algorithm calculated probability of death from each disease that correctly classify death instances with an accuracy of 28.13%. K means clustering divided the data into four clusters with 189, 59, 65, 144 country-years in each cluster. A Chi-square was used to test discriminate disease differences found in each cluster which had different diseases as predominant causes of death. Apriori association analysis produced association rules of linkage among cancer of the lung, hypertension and cerebrovascular diseases. These were found in the top five leading causes of death with 99-100% confidence level.Conclusion: Classification tools produced the poorest results in predicting cause of death. Given the inadequacy of variables in the WHO database, creation of a classification model to predict specific cause of death was impossible. Clustering and association tools yielded interesting results that could be used to identify new areas of interest in mortality data analysis. This can be used in data mining analysis to help solve some quality problems in mortality data. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|