A comparison of machine learning methods for classification using simulation with multiple real data examples from mental health studies.

Background: Recent literature on the comparison of machine learning methods has raised questions about the neutrality, unbiasedness and utility of many comparative studies. Reporting of results on favourable datasets and sampling error in the estimated performance measures based on single samples ar...

Descripción completa

Detalles Bibliográficos
Publicado en:Statistical Methods in Medical Research Vol. 25; no. 5; pp. 1804 - 1824
Autores principales: Khondoker, Mizanur, Dobson, Richard, Skirrow, Caroline, Simmons, Andrew, Stahl, Daniel
Formato: equations & formulas research tables/charts Journal Article
Publicado: Sage Publications Inc. Oct2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=118518694&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 118518694
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        09622802
        31F
      jtl: Statistical Methods in Medical Research
      issn: 09622802
      maglogo: Y
    pubinfo:
      dt: Oct2016
      vid: 25
      iid: 5
      pid: 344
      pub: Sage Publications Inc.
      place: Thousand Oaks, California
    artinfo:
      ui:
        118518694
        118518694
        NLM24047600
        118518694
        10.1177/0962280213502437
        NLM24047600
        118518694
      ppf: 1804
      ppct: 20
      formats:
      tig:
        atl: A comparison of machine learning methods for classification using simulation with multiple real data examples from mental health studies.
      aug:
        au:
          Khondoker, Mizanur
          Dobson, Richard
          Skirrow, Caroline
          Simmons, Andrew
          Stahl, Daniel
        affil: King's College London, Institute of Psychiatry, Department of Biostatistics, London, UK
      sug:
        subj:
          Behavioral Research Methods
          Mental Health
          Human
          Discriminant Analysis
          Time
          Sample Size
          Colorectal Neoplasms
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
      ab: Background: Recent literature on the comparison of machine learning methods has raised questions about the neutrality, unbiasedness and utility of many comparative studies. Reporting of results on favourable datasets and sampling error in the estimated performance measures based on single samples are thought to be the major sources of bias in such comparisons. Better performance in one or a few instances does not necessarily imply so on an average or on a population level and simulation studies may be a better alternative for objectively comparing the performances of machine learning algorithms.Methods: We compare the classification performance of a number of important and widely used machine learning algorithms, namely the Random Forests (RF), Support Vector Machines (SVM), Linear Discriminant Analysis (LDA) and k-Nearest Neighbour (kNN). Using massively parallel processing on high-performance supercomputers, we compare the generalisation errors at various combinations of levels of several factors: number of features, training sample size, biological variation, experimental variation, effect size, replication and correlation between features.Results: For smaller number of correlated features, number of features not exceeding approximately half the sample size, LDA was found to be the method of choice in terms of average generalisation errors as well as stability (precision) of error estimates. SVM (with RBF kernel) outperforms LDA as well as RF and kNN by a clear margin as the feature set gets larger provided the sample size is not too small (at least 20). The performance of kNN also improves as the number of features grows and outplays that of LDA and RF unless the data variability is too high and/or effect sizes are too small. RF was found to outperform only kNN in some instances where the data are more variable and have smaller effect sizes, in which cases it also provide more stable error estimates than kNN and LDA. Applications to a number of real datasets supported the findings from the simulation study.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N