Simpute: an efficient solution for dense genotypic data.

Single nucleotide polymorphism (SNP) data derived from array-based technology or massive parallel sequencing are often flawed with missing data. Missing SNPs can bias the results of association analyses. To maximize information usage, imputation is often adopted to compensate for the missing data by...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2013; pp. 813912 - 813913
Autores principales: Lin, Yen-Jen, Chang, Chun-Tien, Tang, Chuan Yi, Hsieh, Wen-Ping
Formato: Journal Article
Publicado: Wiley-Blackwell 2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104287154&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104287154
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 2013
      vid: 2013
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104287154
        2012116747
        NLM23509783
        PMC3581137
        104287154
      ppf: 813912
      ppct: 1
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Simpute: an efficient solution for dense genotypic data.
      aug:
        au:
          Lin, Yen-Jen
          Chang, Chun-Tien
          Tang, Chuan Yi
          Hsieh, Wen-Ping
        affil: Department of Computer Science, National Tsing Hua University, Hsinchu, Taiwan.
      sug:
        subj:
          Bioinformatics Methods
          Genotype
          Polymorphism, Genetic
          Software
          Algorithms
          Alleles
          Genome, Human
          Genomics
          Models, Biological
      ab: Single nucleotide polymorphism (SNP) data derived from array-based technology or massive parallel sequencing are often flawed with missing data. Missing SNPs can bias the results of association analyses. To maximize information usage, imputation is often adopted to compensate for the missing data by filling in the most probable values. To better understand the available tools for this purpose, we compare the imputation performances among BEAGLE, IMPUTE, BIMBAM, SNPMStat, MACH, and PLINK with data generated by randomly masking the genotype data from the International HapMap Phase III project. In addition, we propose a new algorithm called simple imputation (Simpute) that benefits from the high resolution of the SNPs in the array platform. Simpute does not require any reference data. The best feature of Simpute is its computational efficiency with complexity of order (mw + n), where n is the number of missing SNPs, w is the number of the positions of the missing SNPs, and m is the number of people considered. Simpute is suitable for regular screening of the large-scale SNP genotyping particularly when the sample size is large, and efficiency is a major concern in the analysis.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N