An Efficient and Effective Model to Handle Missing Data in Classification.

Missing data is one of the most important causes in reduction of classification accuracy. Many real datasets suffer from missing values, especially in medical sciences. Imputation is a common way to deal with incomplete datasets. There are various imputation methods that can be applied, and the choi...

Full description

Bibliographic Details
Published in:BioMed Research International pp. 1 - 12
Main Authors: Mehrabani-Zeinabad, Kamran, Doostfatemeh, Marziyeh, Ayatollahi, Seyyed Mohammad Taghi
Format: research tables/charts Journal Article
Published: Wiley-Blackwell 11/25/2020
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=147200759&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 147200759
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 11/25/2020
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        147200759
        147200759
        147200759
        10.1155/2020/8810143
        147200759
      ppf: 1
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: An Efficient and Effective Model to Handle Missing Data in Classification.
      aug:
        au:
          Mehrabani-Zeinabad, Kamran
          Doostfatemeh, Marziyeh
          Ayatollahi, Seyyed Mohammad Taghi
        affil: Department of Biostatistics, Faculty of Medicine, Shiraz University of Medical Sciences, Shiraz, Iran
      sug:
        subj:
          Data Management
          Classification
          Human
          Regression
          Models, Statistical
          Decision Trees
      ab: Missing data is one of the most important causes in reduction of classification accuracy. Many real datasets suffer from missing values, especially in medical sciences. Imputation is a common way to deal with incomplete datasets. There are various imputation methods that can be applied, and the choice of the best method depends on the dataset conditions such as sample size, missing percent, and missing mechanism. Therefore, the better solution is to classify incomplete datasets without imputation and without any loss of information. The structure of the "Bayesian additive regression trees" (BART) model is improved with the "Missingness Incorporated in Attributes" approach to solve its inefficiency in handling the missingness problem. Implementation of MIA-within-BART is named "BART.m". As the abilities of BART.m are not investigated in classification of incomplete datasets, this simulation-based study aimed to provide such resource. The results indicate that BART.m can be used even for datasets with 90 missing present and more importantly, it diagnoses the irrelevant variables and removes them by its own. BART.m outperforms common models for classification with incomplete data, according to accuracy and computational time. Based on the revealed properties, it can be said that BART.m is a high accuracy model in classification of incomplete datasets which avoids any assumptions and preprocess steps.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N