Pattern Matching for DNA Sequencing Data Using Multiple Bloom Filters.

Storing and processing of large DNA sequences has always been a major problem due to increasing volume of DNA sequence data. However, a number of solutions have been proposed but they require significant computation and memory. Therefore, an efficient storage and pattern matching solution is require...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International pp. 1 - 10
Autores principales: Najam, Maleeha, Rasool, Raihan Ur, Ahmad, Hafiz Farooq, Ashraf, Usman, Malik, Asad Waqar
Formato: equations & formulas research tables/charts Journal Article
Publicado: Wiley-Blackwell 4/14/2019
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=135881494&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 135881494
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 4/14/2019
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        135881494
        135881494
        135881494
        10.1155/2019/7074387
        135881494
      ppf: 1
      ppct: 9
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Pattern Matching for DNA Sequencing Data Using Multiple Bloom Filters.
      aug:
        au:
          Najam, Maleeha
          Rasool, Raihan Ur
          Ahmad, Hafiz Farooq
          Ashraf, Usman
          Malik, Asad Waqar
        affil: Fatima Jinnah Women University, Rawalpindi, Pakistan
      sug:
        subj:
          Sequence Analysis Methods
          DNA Analysis
          DNA Classification
          Data Management
          Bioinformatics
          Human
      ab: Storing and processing of large DNA sequences has always been a major problem due to increasing volume of DNA sequence data. However, a number of solutions have been proposed but they require significant computation and memory. Therefore, an efficient storage and pattern matching solution is required for DNA sequencing data. Bloom filters (BFs) represent an efficient data structure, which is mostly used in the domain of bioinformatics for classification of DNA sequences. In this paper, we explore more dimensions where BFs can be used other than classification. A proposed solution is based on Multiple Bloom Filters (MBFs) that finds all the locations and number of repetitions of the specified pattern inside a DNA sequence. Both of these factors are extremely important in determining the type and intensity of any disease. This paper serves as a first effort towards optimizing the search for location and frequency of substrings in DNA sequences using MBFs. We expect that further optimizations in the proposed solution can bring remarkable results as this paper presents a proof of concept implementation for a given set of data using proposed MBFs technique. Performance evaluation shows improved accuracy and time efficiency of the proposed approach.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N