A novel method based on physicochemical properties of amino acids and one class classification algorithm for disease gene identification.
Identifying the genes that cause disease is one of the most challenging issues to establish the diagnosis and treatment quickly. Several interesting methods have been introduced for disease gene identification for a decade. In general, the main differences between these methods are the type of data...
| Publicado en: | Journal of Biomedical Informatics Vol. 56; pp. 300 - 307 |
|---|---|
| Autores principales: | , |
| Formato: | Journal Article |
| Publicado: |
Academic Press Inc.
Aug2015
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109616576&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 109616576 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15320464 OMB jtl: Journal of Biomedical Informatics issn: 15320464 maglogo: N pubinfo: dt: Aug2015 vid: 56 pid: 735 pub: Academic Press Inc. place: Burlington, Massachusetts artinfo: ui: 109616576 NLM26146156 2013116460 10.1016/j.jbi.2015.06.018 NLM26146156 109616576 ppf: 300 ppct: 7 formats: tig: atl: A novel method based on physicochemical properties of amino acids and one class classification algorithm for disease gene identification. aug: au: Yousef, Abdulaziz Charkari, Nasrollah Moghadam sug: ab: Identifying the genes that cause disease is one of the most challenging issues to establish the diagnosis and treatment quickly. Several interesting methods have been introduced for disease gene identification for a decade. In general, the main differences between these methods are the type of data used as a prior-knowledge, as well as machine learning (ML) methods used for identification. The disease gene identification task has been commonly viewed by ML methods as a binary classification problem (whether any gene is disease or not). However, the nature of the data (since there is no negative data available for training or leaners) creates a major problem which affect the results. In this paper, sequence-based, one class classification method is introduced to assign genes to disease class (yes, no). First, to generate feature vector, the sequences of proteins (genes) are initially transformed to numerical vector using physicochemical properties of amino acid. Second, as there is no definite approach to define non-disease genes (negative data); we have attempted to model solely disease genes (positive data) to make a prediction by employing Support Vector Data Description algorithm. The experimental results confirm the efficiency of the method with precision, recall and F-measure of 79.3%, 82.6% and 80.9%, respectively. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|