PairMotifChIP: A Fast Algorithm for Discovery of Patterns Conserved in Large ChIP-seq Data Sets.

Identifying conserved patterns in DNA sequences, namely, motif discovery, is an important and challenging computational task. With hundreds or more sequences contained, the high-throughput sequencing data set is helpful to improve the identification accuracy of motif discovery but requires an even h...

Full description

Bibliographic Details
Published in:BioMed Research International Vol. 2016; pp. 1 - 11
Main Authors: Yu, Qiang, Huo, Hongwei, Feng, Dazheng
Format: algorithm equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell 10/24/2016
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=119019460&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 119019460
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 10/24/2016
      vid: 2016
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        119019460
        119019460
        119019460
        10.1155/2016/4986707
        119019460
      ppf: 1
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: PairMotifChIP: A Fast Algorithm for Discovery of Patterns Conserved in Large ChIP-seq Data Sets.
      aug:
        au:
          Yu, Qiang
          Huo, Hongwei
          Feng, Dazheng
        affil: School of Computer Science and Technology, Xidian University, Xi’an 710071, China
      sug:
        subj:
          Algorithms Methods
          DNA Analysis
          Molecular Structure
          Nucleic Acids
          Simulations
          Data Mining
          Probability
          Algorithms Evaluation
          Nucleotides
          Time
          Descriptive Statistics
          Funding Source
      ab: Identifying conserved patterns in DNA sequences, namely, motif discovery, is an important and challenging computational task. With hundreds or more sequences contained, the high-throughput sequencing data set is helpful to improve the identification accuracy of motif discovery but requires an even higher computing performance. To efficiently identify motifs in large DNA data sets, a new algorithm called PairMotifChIP is proposed by extracting and combining pairs of l-mers in the input with relatively small Hamming distance. In particular, a method for rapidly extracting pairs of l-mers is designed, which can be used not only for PairMotifChIP, but also for other DNA data mining tasks with the same demand. Experimental results on the simulated data show that the proposed algorithm can find motifs successfully and runs faster than the state-of-the-art motif discovery algorithms. Furthermore, the validity of the proposed algorithm has been verified on real data.
      pubtype: Academic Journal
      doctype:
        algorithm
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N