Lessons learned from data mining of WHO mortality database.

Objectives: The objectives of this research were to test the ability of classification algorithms to predict the cause of death in the mortality data with unknown causes, to find association between common causes of death, to identify groups of countries based on their common causes of death, and to...

Descripción completa

Detalles Bibliográficos
Publicado en:Methods of Information in Medicine Vol. 50; no. 4; pp. 380 - 386
Autores principales: Paoin W, Paoin, W
Formato: research Journal Article
Publicado: Thieme Medical Publishing Inc. 2011
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=108192011&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 108192011
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        00261270
        W7M
      jtl: Methods of Information in Medicine
      issn: 00261270
      maglogo: N
    pubinfo:
      dt: 2011
      vid: 50
      iid: 4
      pid: 2811
      pub: Thieme Medical Publishing Inc.
      place: New York, New York
    artinfo:
      ui:
        108192011
        108192011
        NLM21691674
        2011241185
        10.3414/ME10-02-0019
        NLM21691674
        108192011
      ppf: 380
      ppct: 6
      formats:
      tig:
        atl: Lessons learned from data mining of WHO mortality database.
      aug:
        au:
          Paoin W
          Paoin, W
        affil: Faculty of Medicine, Thammasat University, Rangsit Campus, Paholyotin Road, Pathumthani 12120, Thailand
      sug:
        subj:
          Data Mining Methods
          Resource Databases
          Medical Informatics Administration
          Mortality
          World Health Organization
          Algorithms
          Cause of Death
          Chi Square Test
          Cluster Analysis
          Management Information Systems
          Decision Trees
      ab: Objectives: The objectives of this research were to test the ability of classification algorithms to predict the cause of death in the mortality data with unknown causes, to find association between common causes of death, to identify groups of countries based on their common causes of death, and to extract knowledge gained from data mining of the World Health Organization mortality database.Methods: The WEKA software version 3.5.3 was used for classification, clustering and association analysis of the World Health Organization mortality database which contained 1,109,537 records. Three major steps were performed: Step 1 - preprocessing of data to convert all records into suitable formats for each type of analysis algorithm; Step 2 - analyzing data using the C4.5 decision tree and Naïve Bayes classification algorithm, K-means clustering algorithm and Apriori association analysis algorithm; Step 3 - interpretation of results and hypothesis testing after clustering analysis.Results: Using a C4.5 decision tree classifier to predict cause of death, we obtained 440 leaf nodes that correctly classify death instances with an accuracy of 40.06%. Naïve Bayes classification algorithm calculated probability of death from each disease that correctly classify death instances with an accuracy of 28.13%. K means clustering divided the data into four clusters with 189, 59, 65, 144 country-years in each cluster. A Chi-square was used to test discriminate disease differences found in each cluster which had different diseases as predominant causes of death. Apriori association analysis produced association rules of linkage among cancer of the lung, hypertension and cerebrovascular diseases. These were found in the top five leading causes of death with 99-100% confidence level.Conclusion: Classification tools produced the poorest results in predicting cause of death. Given the inadequacy of variables in the WHO database, creation of a classification model to predict specific cause of death was impossible. Clustering and association tools yielded interesting results that could be used to identify new areas of interest in mortality data analysis. This can be used in data mining analysis to help solve some quality problems in mortality data.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N