A novel method of predicting protein disordered regions based on sequence features.
With a large number of disordered proteins and their important functions discovered, it is highly desired to develop effective methods to computationally predict protein disordered regions. In this study, based on Random Forest (RF), Maximum Relevancy Minimum Redundancy (mRMR), and Incremental Featu...
| Publicado en: | BioMed Research International Vol. 2013; pp. 414327 - 414328 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | research Journal Article |
| Publicado: |
Wiley-Blackwell
2013
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104074738&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104074738 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23146133 FT2T jtl: BioMed Research International issn: 23146133 maglogo: N pubinfo: dt: 2013 vid: 2013 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104074738 104074738 2012133645 NLM23710446 PMC3654632 104074738 ppf: 414327 ppct: 1 formats: fmt: @attributes: type: P tig: atl: A novel method of predicting protein disordered regions based on sequence features. aug: au: Zhao, Tong-Hui Jiang, Min Huang, Tao Li, Bi-Qing Zhang, Ning Li, Hai-Peng Cai, Yu-Dong affil: Institute of Systems Biology, Shanghai University, Shanghai 200444, China ; Department of Mathematics, College of Science, Shanghai University, Shanghai 200444, China. sug: subj: Algorithms Bioinformatics Proteins Sequence Analysis ab: With a large number of disordered proteins and their important functions discovered, it is highly desired to develop effective methods to computationally predict protein disordered regions. In this study, based on Random Forest (RF), Maximum Relevancy Minimum Redundancy (mRMR), and Incremental Feature Selection (IFS), we developed a new method to predict disordered regions in proteins. The mRMR criterion was used to rank the importance of all candidate features. Finally, top 128 features were selected from the ranked feature list to build the optimal model, including 92 Position Specific Scoring Matrix (PSSM) conservation score features and 36 secondary structure features. As a result, Matthews correlation coefficient (MCC) of 0.3895 was achieved on the training set by 10-fold cross-validation. On the basis of predicting results for each query sequence by using the method, we used the scanning and modification strategy to improve the performance. The accuracy (ACC) and MCC were increased by 4% and almost 0.2%, respectively, compared with other three popular predictors: DISOPRED, DISOclust, and OnD-CRF. The selected features may shed some light on the understanding of the formation mechanism of disordered structures, providing guidelines for experimental validation. pubtype: Academic Journal doctype: research Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|