Effects of pooling samples on the performance of classification algorithms: a comparative study.
A pooling design can be used as a powerful strategy to compensate for limited amounts of samples or high biological variation. In this paper, we perform a comparative study to model and quantify the effects of virtual pooling on the performance of the widely applied classifiers, support vector machi...
| Publicado en: | Scientific World Journal pp. 278352 - 278353 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Wiley-Blackwell
2012
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104455867&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104455867 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 1537744X 1BX5 jtl: Scientific World Journal issn: 1537744X maglogo: N pubinfo: dt: 2012 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104455867 NLM22654582 2011572150 10.1100/2012/278352 NLM22654582 PMC3361225 104455867 ppf: 278352 ppct: 1 formats: tig: atl: Effects of pooling samples on the performance of classification algorithms: a comparative study. aug: au: Kusonmano, Kanthida Netzer, Michael Baumgartner, Christian Dehmer, Matthias Liedl, Klaus R Graber, Armin affil: Institute for Bioinformatics and Translational Research, UMIT, 6060 Hall in Tyrol, Austria. sug: subj: Classification Algorithms Bioinformatics Methods ab: A pooling design can be used as a powerful strategy to compensate for limited amounts of samples or high biological variation. In this paper, we perform a comparative study to model and quantify the effects of virtual pooling on the performance of the widely applied classifiers, support vector machines (SVMs), random forest (RF), k-nearest neighbors (k-NN), penalized logistic regression (PLR), and prediction analysis for microarrays (PAMs). We evaluate a variety of experimental designs using mock omics datasets with varying levels of pool sizes and considering effects from feature selection. Our results show that feature selection significantly improves classifier performance for non-pooled and pooled data. All investigated classifiers yield lower misclassification rates with smaller pool sizes. RF mainly outperforms other investigated algorithms, while accuracy levels are comparable among all the remaining ones. Guidelines are derived to identify an optimal pooling scheme for obtaining adequate predictive power and, hence, to motivate a study design that meets best experimental objectives and budgetary conditions, including time constraints. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|