Sufficiency Revisited: Rethinking Statistical Algorithms in the Big Data Era.

The big data era demands new statistical analysis paradigms, since traditional methods often break down when datasets are too large to fit on a single desktop computer. Divide and Recombine (D&R) is becoming a popular approach for big data analysis, where results are combined over subanalyses perfor...

Full description

Bibliographic Details
Published in:American Statistician Vol. 71; no. 3; pp. 202 - 209
Main Authors: Lee, Jarod Y. L., Brown, James J., Ryan, Louise M.
Format: Article
Published: Taylor & Francis Ltd Aug2017
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=125746210&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 125746210
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00031305
        STT
      jtl: American Statistician
      issn: 00031305
      maglogo: Y
    pubinfo:
      dt: Aug2017
      vid: 71
      iid: 3
      pid: 377
      pub: Taylor & Francis Ltd
    artinfo:
      ui:
        125746210
        10.1080/00031305.2016.1255659
      ppf: 202
      ppct: 7
      formats:
      tig:
        atl: Sufficiency Revisited: Rethinking Statistical Algorithms in the Big Data Era.
      aug:
        au:
          Lee, Jarod Y. L.
          Brown, James J.
          Ryan, Louise M.
        affil:
          School of Mathematical and Physical Sciences, University of Technology Sydney, Ultimo, NSW, Australia
          Australian Research Council Centre of Excellence for Mathematical & Statistical Frontiers, The University of Melbourne, Parkville, VIC, Australia
          Department of Biostatistics, Harvard T.H. Chan School of Public Health, Boston, MA
      su:
        Electronic data processing
        Statistics
        Data analysis
        Big data
        Data mining
        Databases
      sug:
        subj:
          Electronic data processing
          Data Processing, Hosting, and Related Services
          Statistics
          Data analysis
          Big data
          Data mining
          Databases
      keyword:
        Distributed database
        Divide and recombine
        Generalized linear mixed model
        Multilevel model
        Privacy
        Distributed database
        Divide and recombine
        Generalized linear mixed model
        Multilevel model
        Privacy
      ab: The big data era demands new statistical analysis paradigms, since traditional methods often break down when datasets are too large to fit on a single desktop computer. Divide and Recombine (D&R) is becoming a popular approach for big data analysis, where results are combined over subanalyses performed in separate data subsets. In this article, we consider situations where unit record data cannot be made available by data custodians due to privacy concerns, and explore the concept of statistical sufficiency and summary statistics for model fitting. The resulting approach represents a type of D&R strategy, which we refer to assummary statistics D&R; as opposed to the standard approach, which we refer to ashorizontal D&R. We demonstrate the concept via an extended Gamma–Poisson model, where summary statistics are extracted from different databases and incorporated directly into the fitting algorithm without having to combine unit record data. By exploiting the natural hierarchy of data, our approach has major benefits in terms of privacy protection. Incorporating the proposed modelling framework into data extraction tools such as TableBuilder by the Australian Bureau of Statistics allows for potential analysis at a finer geographical level, which we illustrate with a multilevel analysis of the Australian unemployment data. Supplementary materials for this article are available online.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N