Accurate Diabetes Risk Stratification Using Machine Learning: Role of Missing Value and Outliers.

Diabetes mellitus is a group of metabolic diseases in which blood sugar levels are too high. About 8.8% of the world was diabetic in 2017. It is projected that this will reach nearly 10% by 2045. The major challenge is that when machine learning-based classifiers are applied to such data sets for ri...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Medical Systems Vol. 42; no. 5; pp. 1 - 2
Autores principales: Maniruzzaman, Md., Rahman, Md. Jahanur, Al-MehediHasan, Md., Suri, Harman S., Abedin, Md. Menhazul, El-Baz, Ayman, Suri, Jasjit S.
Formato: algorithm equations & formulas research tables/charts Journal Article
Publicado: Springer Nature May2018
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=129370618&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 129370618
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: May2018
      vid: 42
      iid: 5
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        129370618
        129370618
        129370618
        10.1007/s10916-018-0940-7
        129370618
      ppf: 1
      ppct: 1
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Accurate Diabetes Risk Stratification Using Machine Learning: Role of Missing Value and Outliers.
      aug:
        au:
          Maniruzzaman, Md.
          Rahman, Md. Jahanur
          Al-MehediHasan, Md.
          Suri, Harman S.
          Abedin, Md. Menhazul
          El-Baz, Ayman
          Suri, Jasjit S.
        affil: Department of Statistics, University of Rajshahi, Rajshahi, Bangladesh
      sug:
        subj:
          Risk Assessment
          Diabetes Mellitus Classification
          Machine Learning
          Human
          Female
          Adult
          Native Americans
          Logistic Regression
          Neural Networks (Computer)
          Decision Trees
          Experimental Studies
          Factor Analysis
          Adult: 19-44 years
          Female
      ab: Diabetes mellitus is a group of metabolic diseases in which blood sugar levels are too high. About 8.8% of the world was diabetic in 2017. It is projected that this will reach nearly 10% by 2045. The major challenge is that when machine learning-based classifiers are applied to such data sets for risk stratification, leads to lower performance. Thus, our objective is to develop an optimized and robust machine learning (ML) system under the assumption that missing values or outliers if replaced by a median configuration will yield higher risk stratification accuracy. This ML-based risk stratification is designed, optimized and evaluated, where: (i) the features are extracted and optimized from the six feature selection techniques (random forest, logistic regression, mutual information, principal component analysis, analysis of variance, and Fisher discriminant ratio) and combined with ten different types of classifiers (linear discriminant analysis, quadratic discriminant analysis, naïve Bayes, Gaussian process classification, support vector machine, artificial neural network, Adaboost, logistic regression, decision tree, and random forest) under the hypothesis that both missing values and outliers when replaced by computed medians will improve the risk stratification accuracy. Pima Indian diabetic dataset (768 patients: 268 diabetic and 500 controls) was used. Our results demonstrate that on replacing the missing values and outliers by group median and median values, respectively and further using the combination of random forest feature selection and random forest classification technique yields an accuracy, sensitivity, specificity, positive predictive value, negative predictive value and area under the curve as: 92.26%, 95.96%, 79.72%, 91.14%, 91.20%, and 0.93, respectively. This is an improvement of 10% over previously developed techniques published in literature. The system was validated for its stability and reliability. RF-based model showed the best performance when outliers are replaced by median values.
      pubtype: Academic Journal
      doctype:
        algorithm
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N