Effects of pooling samples on the performance of classification algorithms: a comparative study.

A pooling design can be used as a powerful strategy to compensate for limited amounts of samples or high biological variation. In this paper, we perform a comparative study to model and quantify the effects of virtual pooling on the performance of the widely applied classifiers, support vector machi...

Descripción completa

Detalles Bibliográficos
Publicado en:Scientific World Journal pp. 278352 - 278353
Autores principales: Kusonmano, Kanthida, Netzer, Michael, Baumgartner, Christian, Dehmer, Matthias, Liedl, Klaus R, Graber, Armin
Formato: research Journal Article
Publicado: Wiley-Blackwell 2012
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104455867&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104455867
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        1537744X
        1BX5
      jtl: Scientific World Journal
      issn: 1537744X
      maglogo: N
    pubinfo:
      dt: 2012
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104455867
        NLM22654582
        2011572150
        10.1100/2012/278352
        NLM22654582
        PMC3361225
        104455867
      ppf: 278352
      ppct: 1
      formats:
      tig:
        atl: Effects of pooling samples on the performance of classification algorithms: a comparative study.
      aug:
        au:
          Kusonmano, Kanthida
          Netzer, Michael
          Baumgartner, Christian
          Dehmer, Matthias
          Liedl, Klaus R
          Graber, Armin
        affil: Institute for Bioinformatics and Translational Research, UMIT, 6060 Hall in Tyrol, Austria.
      sug:
        subj:
          Classification Algorithms
          Bioinformatics Methods
      ab: A pooling design can be used as a powerful strategy to compensate for limited amounts of samples or high biological variation. In this paper, we perform a comparative study to model and quantify the effects of virtual pooling on the performance of the widely applied classifiers, support vector machines (SVMs), random forest (RF), k-nearest neighbors (k-NN), penalized logistic regression (PLR), and prediction analysis for microarrays (PAMs). We evaluate a variety of experimental designs using mock omics datasets with varying levels of pool sizes and considering effects from feature selection. Our results show that feature selection significantly improves classifier performance for non-pooled and pooled data. All investigated classifiers yield lower misclassification rates with smaller pool sizes. RF mainly outperforms other investigated algorithms, while accuracy levels are comparable among all the remaining ones. Guidelines are derived to identify an optimal pooling scheme for obtaining adequate predictive power and, hence, to motivate a study design that meets best experimental objectives and budgetary conditions, including time constraints.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N