Sequence-Based Prediction of RNA-Binding Proteins Using Random Forest with Minimum Redundancy Maximum Relevance Feature Selection.
The prediction of RNA-binding proteins is one of the most challenging problems in computation biology. Although some studies have investigated this problem, the accuracy of prediction is still not sufficient. In this study, a highly accurate method was developed to predict RNA-binding proteins from...
| Published in: | BioMed Research International Vol. 2015; pp. 1 - 11 |
|---|---|
| Main Authors: | , , |
| Format: | Journal Article |
| Published: |
Wiley-Blackwell
10/12/2015
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=110561856&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 110561856 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23146133 FT2T jtl: BioMed Research International issn: 23146133 maglogo: N pubinfo: dt: 10/12/2015 vid: 2015 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 110561856 10.1155/2015/425810 110561856 ppf: 1 ppct: 10 formats: fmt: @attributes: type: P tig: atl: Sequence-Based Prediction of RNA-Binding Proteins Using Random Forest with Minimum Redundancy Maximum Relevance Feature Selection. aug: au: Ma, Xin Guo, Jing Sun, Xiao affil: Golden Audit College, Nanjing Audit University, Nanjing 210029, China sug: ab: The prediction of RNA-binding proteins is one of the most challenging problems in computation biology. Although some studies have investigated this problem, the accuracy of prediction is still not sufficient. In this study, a highly accurate method was developed to predict RNA-binding proteins from amino acid sequences using random forests with the minimum redundancy maximum relevance (mRMR) method, followed by incremental feature selection (IFS). We incorporated features of conjoint triad features and three novel features: binding propensity (BP), nonbinding propensity (NBP), and evolutionary information combined with physicochemical properties (EIPP). The results showed that these novel features have important roles in improving the performance of the predictor. Using the mRMR-IFS method, our predictor achieved the best performance (86.62% accuracy and 0.737 Matthews correlation coefficient). High prediction accuracy and successful prediction performance suggested that our method can be a useful approach to identify RNA-binding proteins from sequence information. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|