A novel method of predicting protein disordered regions based on sequence features.

With a large number of disordered proteins and their important functions discovered, it is highly desired to develop effective methods to computationally predict protein disordered regions. In this study, based on Random Forest (RF), Maximum Relevancy Minimum Redundancy (mRMR), and Incremental Featu...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2013; pp. 414327 - 414328
Autores principales: Zhao, Tong-Hui, Jiang, Min, Huang, Tao, Li, Bi-Qing, Zhang, Ning, Li, Hai-Peng, Cai, Yu-Dong
Formato: research Journal Article
Publicado: Wiley-Blackwell 2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104074738&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104074738
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 2013
      vid: 2013
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104074738
        104074738
        2012133645
        NLM23710446
        PMC3654632
        104074738
      ppf: 414327
      ppct: 1
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: A novel method of predicting protein disordered regions based on sequence features.
      aug:
        au:
          Zhao, Tong-Hui
          Jiang, Min
          Huang, Tao
          Li, Bi-Qing
          Zhang, Ning
          Li, Hai-Peng
          Cai, Yu-Dong
        affil: Institute of Systems Biology, Shanghai University, Shanghai 200444, China ; Department of Mathematics, College of Science, Shanghai University, Shanghai 200444, China.
      sug:
        subj:
          Algorithms
          Bioinformatics
          Proteins
          Sequence Analysis
      ab: With a large number of disordered proteins and their important functions discovered, it is highly desired to develop effective methods to computationally predict protein disordered regions. In this study, based on Random Forest (RF), Maximum Relevancy Minimum Redundancy (mRMR), and Incremental Feature Selection (IFS), we developed a new method to predict disordered regions in proteins. The mRMR criterion was used to rank the importance of all candidate features. Finally, top 128 features were selected from the ranked feature list to build the optimal model, including 92 Position Specific Scoring Matrix (PSSM) conservation score features and 36 secondary structure features. As a result, Matthews correlation coefficient (MCC) of 0.3895 was achieved on the training set by 10-fold cross-validation. On the basis of predicting results for each query sequence by using the method, we used the scanning and modification strategy to improve the performance. The accuracy (ACC) and MCC were increased by 4% and almost 0.2%, respectively, compared with other three popular predictors: DISOPRED, DISOclust, and OnD-CRF. The selected features may shed some light on the understanding of the formation mechanism of disordered structures, providing guidelines for experimental validation.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N