Unsupervised learning assisted robust prediction of bioluminescent proteins.
Bioluminescence plays an important role in nature, for example, it is used for intracellular chemical signalling in bacteria. It is also used as a useful reagent for various analytical research methods ranging from cellular imaging to gene expression analysis. However, identification and annotation...
| Publicado en: | Computers in Biology & Medicine Vol. 68; pp. 27 - 37 |
|---|---|
| Autores principales: | , |
| Formato: | research Journal Article |
| Publicado: |
Elsevier B.V.
1/1/2016
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=115381124&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 115381124 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 00104825 JC2 jtl: Computers in Biology & Medicine issn: 00104825 maglogo: N pubinfo: dt: 1/1/2016 vid: 68 pid: 82545 pub: Elsevier B.V. place: Philadelphia, Pennsylvania artinfo: ui: 115381124 115381124 NLM26599828 115381124 10.1016/j.compbiomed.2015.10.013 NLM26599828 115381124 ppf: 27 ppct: 10 formats: tig: atl: Unsupervised learning assisted robust prediction of bioluminescent proteins. aug: au: Nath, Abhigyan Subbiah, Karthikeyan affil: Department of Computer Science, Banaras Hindu University, Varanasi 221005, India sug: subj: Algorithms Proteins Sequence Analysis Methods Predictive Value of Tests Human ab: Bioluminescence plays an important role in nature, for example, it is used for intracellular chemical signalling in bacteria. It is also used as a useful reagent for various analytical research methods ranging from cellular imaging to gene expression analysis. However, identification and annotation of bioluminescent proteins is a difficult task as they share poor sequence similarities among them. In this paper, we present a novel approach for within-class and between-class balancing as well as diversifying of a training dataset by effectively combining unsupervised K-Means algorithm with Synthetic Minority Oversampling Technique (SMOTE) in order to achieve the true performance of the prediction model. Further, we experimented by varying different levels of balancing ratio of positive data to negative data in the training dataset in order to probe for an optimal class distribution which produces the best prediction accuracy. The appropriately balanced and diversified training set resulted in near complete learning with greater generalization on the blind test datasets. The obtained results strongly justify the fact that optimal class distribution with a high degree of diversity is an essential factor to achieve near perfect learning. Using random forest as the weak learners in boosting and training it on the optimally balanced and diversified training dataset, we achieved an overall accuracy of 95.3% on a tenfold cross validation test, and an accuracy of 91.7%, sensitivity of 89. 3% and specificity of 91.8% on a holdout test set. It is quite possible that the general framework discussed in the current work can be successfully applied to other biological datasets to deal with imbalance and incomplete learning problems effectively. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|