A k-mer-based barcode DNA classification methodology based on spectral representation and a neural gas network.

Objectives: In this paper, an alignment-free method for DNA barcode classification that is based on both a spectral representation and a neural gas network for unsupervised clustering is proposed.Methods: In the proposed methodology, distinctive words are identified from a spectral representation of...

Full description

Bibliographic Details
Published in:Artificial Intelligence in Medicine Vol. 64; no. 3; pp. 173 - 185
Main Authors: Fiannaca, Antonino, La Rosa, Massimo, Rizzo, Riccardo, Urso, Alfonso
Format: research Journal Article
Published: Elsevier B.V. Jul2015
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109647283&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 109647283
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        09333657
        3HY
      jtl: Artificial Intelligence in Medicine
      issn: 09333657
      maglogo: N
    pubinfo:
      dt: Jul2015
      vid: 64
      iid: 3
      pid: 1004
      pub: Elsevier B.V.
    artinfo:
      ui:
        109647283
        NLM26170017
        2013165713
        10.1016/j.artmed.2015.06.002
        NLM26170017
        109647283
      ppf: 173
      ppct: 12
      formats:
      tig:
        atl: A k-mer-based barcode DNA classification methodology based on spectral representation and a neural gas network.
      aug:
        au:
          Fiannaca, Antonino
          La Rosa, Massimo
          Rizzo, Riccardo
          Urso, Alfonso
      sug:
      ab: Objectives: In this paper, an alignment-free method for DNA barcode classification that is based on both a spectral representation and a neural gas network for unsupervised clustering is proposed.Methods: In the proposed methodology, distinctive words are identified from a spectral representation of DNA sequences. A taxonomic classification of the DNA sequence is then performed using the sequence signature, i.e., the smallest set of k-mers that can assign a DNA sequence to its proper taxonomic category. Experiments were then performed to compare our method with other supervised machine learning classification algorithms, such as support vector machine, random forest, ripper, naïve Bayes, ridor, and classification tree, which also consider short DNA sequence fragments of 200 and 300 base pairs (bp). The experimental tests were conducted over 10 real barcode datasets belonging to different animal species, which were provided by the on-line resource "Barcode of Life Database".Results: The experimental results showed that our k-mer-based approach is directly comparable, in terms of accuracy, recall and precision metrics, with the other classifiers when considering full-length sequences. In addition, we demonstrate the robustness of our method when a classification is performed task with a set of short DNA sequences that were randomly extracted from the original data. For example, the proposed method can reach the accuracy of 64.8% at the species level with 200-bp fragments. Under the same conditions, the best other classifier (random forest) reaches the accuracy of 20.9%.Conclusions: Our results indicate that we obtained a clear improvement over the other classifiers for the study of short DNA barcode sequence fragments.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N