Efficient feature selection and classification of protein sequence data in bioinformatics.

Bioinformatics has been an emerging area of research for the last three decades. The ultimate aims of bioinformatics were to store and manage the biological data, and develop and analyze computational tools to enhance their understanding. The size of data accumulated under various sequencing project...

Descripción completa

Detalles Bibliográficos
Publicado en:Scientific World Journal pp. 173869 - 173870
Autores principales: Iqbal, Muhammad Javed, Faye, Ibrahima, Samir, Brahim Belhaouari, Md Said, Abas, Said, Abas Md
Formato: research Journal Article
Publicado: Wiley-Blackwell 2014
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=103835228&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 103835228
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        1537744X
        1BX5
      jtl: Scientific World Journal
      issn: 1537744X
      maglogo: N
    pubinfo:
      dt: 2014
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        103835228
        103835228
        NLM25045727
        2012655904
        10.1155/2014/173869
        NLM25045727
        PMC4089199
        103835228
      ppf: 173869
      ppct: 1
      formats:
      tig:
        atl: Efficient feature selection and classification of protein sequence data in bioinformatics.
      aug:
        au:
          Iqbal, Muhammad Javed
          Faye, Ibrahima
          Samir, Brahim Belhaouari
          Md Said, Abas
          Said, Abas Md
        affil: Computer and Information Sciences Department, Universiti Teknologi PETRONAS, Bandar Seri Iskandar, 31750 Tronoh, Perak, Malaysia
      sug:
        subj:
          Bioinformatics Methods
          Proteins
      ab: Bioinformatics has been an emerging area of research for the last three decades. The ultimate aims of bioinformatics were to store and manage the biological data, and develop and analyze computational tools to enhance their understanding. The size of data accumulated under various sequencing projects is increasing exponentially, which presents difficulties for the experimental methods. To reduce the gap between newly sequenced protein and proteins with known functions, many computational techniques involving classification and clustering algorithms were proposed in the past. The classification of protein sequences into existing superfamilies is helpful in predicting the structure and function of large amount of newly discovered proteins. The existing classification results are unsatisfactory due to a huge size of features obtained through various feature encoding methods. In this work, a statistical metric-based feature selection technique has been proposed in order to reduce the size of the extracted feature vector. The proposed method of protein classification shows significant improvement in terms of performance measure metrics: accuracy, sensitivity, specificity, recall, F-measure, and so forth.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N