Framework for classification of cancer gene expression data using Bayesian hyper-parameter optimization.
Computational classification of cancers is an important research problem. Gene expression data has 1000s of features, very few samples, and a class imbalance problem. In this paper, we have proposed a framework for the classification of cancer gene expression profiles. The framework consists of a pi...
| Publicado en: | Medical & Biological Engineering & Computing Vol. 59; no. 11/12; pp. 2353 - 2372 |
|---|---|
| Autores principales: | , |
| Formato: | Journal Article |
| Publicado: |
Springer Nature
Nov2021
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=153319074&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 153319074 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01400118 PO0 jtl: Medical & Biological Engineering & Computing issn: 01400118 maglogo: N pubinfo: dt: Nov2021 vid: 59 iid: 11/12 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 153319074 152812536 153319074 NLM34609687 10.1007/s11517-021-02442-7 NLM34609687 153319074 ppf: 2353 ppct: 19 formats: fmt: @attributes: type: P tig: atl: Framework for classification of cancer gene expression data using Bayesian hyper-parameter optimization. aug: au: Koul, Nimrita Manvi, Sunilkumar S. affil: School of Computer Science and Engineering, REVA University, 560064, Bangalore, Karnataka, India sug: subj: Neoplasms Probability Gene Expression Algorithms ab: Computational classification of cancers is an important research problem. Gene expression data has 1000s of features, very few samples, and a class imbalance problem. In this paper, we have proposed a framework for the classification of cancer gene expression profiles. The framework consists of a pipeline of methods for data pre-processing, feature selection, and classification. Data pre-processing is done by standard scaling and normalization of the features. The feature selection is performed in two steps. First, recursive feature elimination (RFE) is used; then, a genetic algorithm is applied only in case RFE results in a feature subset of size more than a specific threshold. Next, is a meta-pool of diverse, individual as well as ensemble classifiers. Hyper-parameters of each member in the meta-pool are optimized using Bayesian Optimization. An algorithm is developed to select the best classifier from the meta-pool based on classification accuracy and computation time taken. We evaluated the framework on 6 publicly available microarray datasets and the PAN-Cancer RNA Sequencing dataset. We found that the classifier selected by the proposed framework produced significant improvement in classification accuracy and computation time required to predict labels for test datasets. A detailed comparison with the state-of-the-art methods shows that the proposed framework outperforms all of them. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|