Framework for classification of cancer gene expression data using Bayesian hyper-parameter optimization.

Computational classification of cancers is an important research problem. Gene expression data has 1000s of features, very few samples, and a class imbalance problem. In this paper, we have proposed a framework for the classification of cancer gene expression profiles. The framework consists of a pi...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical & Biological Engineering & Computing Vol. 59; no. 11/12; pp. 2353 - 2372
Autores principales: Koul, Nimrita, Manvi, Sunilkumar S.
Formato: Journal Article
Publicado: Springer Nature Nov2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=153319074&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 153319074
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01400118
        PO0
      jtl: Medical & Biological Engineering & Computing
      issn: 01400118
      maglogo: N
    pubinfo:
      dt: Nov2021
      vid: 59
      iid: 11/12
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        153319074
        152812536
        153319074
        NLM34609687
        10.1007/s11517-021-02442-7
        NLM34609687
        153319074
      ppf: 2353
      ppct: 19
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Framework for classification of cancer gene expression data using Bayesian hyper-parameter optimization.
      aug:
        au:
          Koul, Nimrita
          Manvi, Sunilkumar S.
        affil: School of Computer Science and Engineering, REVA University, 560064, Bangalore, Karnataka, India
      sug:
        subj:
          Neoplasms
          Probability
          Gene Expression
          Algorithms
      ab: Computational classification of cancers is an important research problem. Gene expression data has 1000s of features, very few samples, and a class imbalance problem. In this paper, we have proposed a framework for the classification of cancer gene expression profiles. The framework consists of a pipeline of methods for data pre-processing, feature selection, and classification. Data pre-processing is done by standard scaling and normalization of the features. The feature selection is performed in two steps. First, recursive feature elimination (RFE) is used; then, a genetic algorithm is applied only in case RFE results in a feature subset of size more than a specific threshold. Next, is a meta-pool of diverse, individual as well as ensemble classifiers. Hyper-parameters of each member in the meta-pool are optimized using Bayesian Optimization. An algorithm is developed to select the best classifier from the meta-pool based on classification accuracy and computation time taken. We evaluated the framework on 6 publicly available microarray datasets and the PAN-Cancer RNA Sequencing dataset. We found that the classifier selected by the proposed framework produced significant improvement in classification accuracy and computation time required to predict labels for test datasets. A detailed comparison with the state-of-the-art methods shows that the proposed framework outperforms all of them.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N