Mining sequential patterns for protein fold recognition.
Abstract: Protein data contain discriminative patterns that can be used in many beneficial applications if they are defined correctly. In this work sequential pattern mining (SPM) is utilized for sequence-based fold recognition. Protein classification in terms of fold recognition plays an important...
| Publicado en: | Journal of Biomedical Informatics Vol. 41; no. 1; pp. 165 - 180 |
|---|---|
| Autores principales: | , , , |
| Formato: | Journal Article |
| Publicado: |
Academic Press Inc.
Feb2008
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=105867693&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 105867693 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Feb2008 vid: 41 iid: 1 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 105867693 2009820933 NLM17573243 105867693 ppf: 165 ppct: 15 formats: tig: atl: Mining sequential patterns for protein fold recognition. aug: au: Exarchos TP Papaloukas C Lampros C Fotiadis DI sug: subj: Biochemical Phenomena Genetic Techniques Methods Information Retrieval Methods Information Science Methods Proteins Resource Databases Sequence Analysis Methods Amino Acids Binding Sites ab: Abstract: Protein data contain discriminative patterns that can be used in many beneficial applications if they are defined correctly. In this work sequential pattern mining (SPM) is utilized for sequence-based fold recognition. Protein classification in terms of fold recognition plays an important role in computational protein analysis, since it can contribute to the determination of the function of a protein whose structure is unknown. Specifically, one of the most efficient SPM algorithms, cSPADE, is employed for the analysis of protein sequence. A classifier uses the extracted sequential patterns to classify proteins in the appropriate fold category. For training and evaluating the proposed method we used the protein sequences from the Protein Data Bank and the annotation of the SCOP database. The method exhibited an overall accuracy of 25% in a classification problem with 36 candidate categories. The classification performance reaches up to 56% when the five most probable protein folds are considered. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|