Data mining: Potential applications in research on nutrition and health.

Aim: Data mining enables further insights from nutrition ‐ related research, but caution is required. The aim of this analysis was to demonstrate and compare the utility of data mining methods in classifying a categorical outcome derived from a nutrition ‐ related intervention. Methods: Baseline dat...

Full description

Bibliographic Details
Published in:Nutrition & Dietetics Vol. 74; no. 1; pp. 3 - 11
Main Authors: Batterham, Marijka, Neale, Elizabeth, Martin, Allison, Tapsell, Linda
Format: research tables/charts Journal Article
Published: Wiley-Blackwell Feb2017
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=121063185&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 121063185
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        14466368
        QMK
      jtl: Nutrition & Dietetics
      issn: 14466368
      maglogo: Y
    pubinfo:
      dt: Feb2017
      vid: 74
      iid: 1
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        121063185
        121063185
        121063185
        10.1111/1747-0080.12337
        121063185
      ppf: 3
      ppct: 8
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Data mining: Potential applications in research on nutrition and health.
      aug:
        au:
          Batterham, Marijka
          Neale, Elizabeth
          Martin, Allison
          Tapsell, Linda
        affil: Statistical Consulting Centre, National Institute for Applied Statistics Research Australia, University of Wollongong, Wollongong New South Wales, Australia
      sug:
        subj:
          Data Mining Methods
          Weight Loss
          Obesity Therapy
          Forecasting
          Algorithms
          Software
          Funding Source
          Intervention Trials Evaluation
          Logistic Regression
          Neural Networks (Computer)
          Human
          Adipose Tissue
          Quality of Life
          Cholesterol Blood
          Lipoproteins, HDL Cholesterol Blood
          Physical Activity
          Blood Glucose
          Educational Status
          Pedometers
          Data Analysis Software
          Decision Trees
          Clinical Assessment Tools
          Adult
          Middle Age
          Descriptive Statistics
          Male
          Female
          P-Value
          Confidence Intervals
          Linear Regression
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Male
          Female
      ab: Aim: Data mining enables further insights from nutrition ‐ related research, but caution is required. The aim of this analysis was to demonstrate and compare the utility of data mining methods in classifying a categorical outcome derived from a nutrition ‐ related intervention. Methods: Baseline data (23 variables, 8 categorical) on participants (n = 295) in an intervention trial were used to classify participants in terms of meeting the criteria of achieving 10 000 steps per day. Results from classification and regression trees (CARTs), random forests, adaptive boosting, logistic regression, support vector machines and neural networks were compared using area under the curve (AUC) and error assessments. Results: The CART produced the best model when considering the AUC (0.703), overall error (18%) and within class error (28%). Logistic regression also performed reasonably well compared to the other models (AUC 0.675, overall error 23%, within class error 36%). All the methods gave different rankings of variables’ importance. CART found that body fat, quality of life using the SF ‐ 12 Physical Component Summary (PCS) and the cholesterol: HDL ratio were the most important predictors of meeting the 10 000 steps criteria, while logistic regression showed the SF ‐ 12PCS, glucose levels and level of education to be the most significant predictors (P ≤ 0.01). Conclusions: Differing outcomes suggest caution is required with a single data mining method, particularly in a dataset with nonlinear relationships and outliers and when exploring relationships that were not the primary outcomes of the research.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N