HBLAST: Parallelised sequence similarity--A Hadoop MapReducable basic local alignment search tool.

The recent exponential growth of genomic databases has resulted in the common task of sequence alignment becoming one of the major bottlenecks in the field of computational biology. It is typical for these large datasets and complex computations to require cost prohibitive High Performance Computing...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Biomedical Informatics Vol. 54; pp. 58 - 65
Autores principales: O'Driscoll, Aisling, Belogrudov, Vladislav, Carroll, John, Kropp, Kai, Walsh, Paul, Ghazal, Peter, Sleator, Roy D
Formato: Journal Article
Publicado: Academic Press Inc. Apr2015
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109726508&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 109726508
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: Apr2015
      vid: 54
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        109726508
        NLM25625550
        2012984846
        10.1016/j.jbi.2015.01.008
        NLM25625550
        109726508
      ppf: 58
      ppct: 7
      formats:
      tig:
        atl: HBLAST: Parallelised sequence similarity--A Hadoop MapReducable basic local alignment search tool.
      aug:
        au:
          O'Driscoll, Aisling
          Belogrudov, Vladislav
          Carroll, John
          Kropp, Kai
          Walsh, Paul
          Ghazal, Peter
          Sleator, Roy D
      sug:
      ab: The recent exponential growth of genomic databases has resulted in the common task of sequence alignment becoming one of the major bottlenecks in the field of computational biology. It is typical for these large datasets and complex computations to require cost prohibitive High Performance Computing (HPC) to function. As such, parallelised solutions have been proposed but many exhibit scalability limitations and are incapable of effectively processing "Big Data" - the name attributed to datasets that are extremely large, complex and require rapid processing. The Hadoop framework, comprised of distributed storage and a parallelised programming framework known as MapReduce, is specifically designed to work with such datasets but it is not trivial to efficiently redesign and implement bioinformatics algorithms according to this paradigm. The parallelisation strategy of "divide and conquer" for alignment algorithms can be applied to both data sets and input query sequences. However, scalability is still an issue due to memory constraints or large databases, with very large database segmentation leading to additional performance decline. Herein, we present Hadoop Blast (HBlast), a parallelised BLAST algorithm that proposes a flexible method to partition both databases and input query sequences using "virtual partitioning". HBlast presents improved scalability over existing solutions and well balanced computational work load while keeping database segmentation and recompilation to a minimum. Enhanced BLAST search performance on cheap memory constrained hardware has significant implications for in field clinical diagnostic testing; enabling faster and more accurate identification of pathogenic DNA in human blood or tissue samples.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N