PlanAlyzer: Assessing Threats to the Validity of Online Experiments.

Online experiments are an integral part of the design and evaluation of software infrastructure at Internet firms. To handle the growing scale and complexity of these experiments, firms have developed software frameworks for their design and deployment. Ensuring that the results of experiments in th...

Descripción completa

Detalles Bibliográficos
Publicado en:Communications of the ACM Vol. 64; no. 9; pp. 108 - 117
Autores principales: Tosch, Emma, Bakshy, Eytan, Berger, Emery D., Jensen, David D., Moss, J. Eliot B.
Formato: Artículo
Publicado: Association for Computing Machinery Sep2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=152071866&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 152071866
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00010782
        ACM
      jtl: Communications of the ACM
      issn: 00010782
      maglogo: N
    pubinfo:
      dt: Sep2021
      vid: 64
      iid: 9
      pid: 68
      pub: Association for Computing Machinery
    artinfo:
      ui:
        152071866
        10.1145/3474385
      ppf: 108
      ppct: 9
      formats:
      tig:
        atl: PlanAlyzer: Assessing Threats to the Validity of Online Experiments.
      aug:
        au:
          Tosch, Emma
          Bakshy, Eytan
          Berger, Emery D.
          Jensen, David D.
          Moss, J. Eliot B.
        affil:
          University of Vermont, Burlington, VT, USA
          Facebook, Inc., Menlo Park, CA, USA
          University of Massachusetts, Amherst, MA, USA
      su:
        Experiments
        Test validity
        Statistical bias
        Programming languages
      sug:
        subj:
          Experiments
          Test validity
          Statistical bias
          Programming languages
      ab: Online experiments are an integral part of the design and evaluation of software infrastructure at Internet firms. To handle the growing scale and complexity of these experiments, firms have developed software frameworks for their design and deployment. Ensuring that the results of experiments in these frameworks are trustworthy--referred to as internal validity--can be difficult. Currently, verifying internal validity requires manual inspection by someone with substantial expertise in experimental design. We present the first approach for checking the internal validity of online experiments statically, that is, from code alone. We identify well-known problems that arise in experimental design and causal inference, which can take on unusual forms when expressed as computer programs: failures of randomization and treatment assignment, and causal sufficiency errors. Our analyses target PlanOut, a popular framework that features a domain-specific language (DSL) to specify and run complex experiments. We have built PlanAlyzer, a tool that checks PlanOut programs for threats to internal validity, before automatically generating important data for the statistical analyses of a large class of experimental designs. We demonstrate PlanAlyzer's utility on a corpus of PlanOut scripts deployed in production at Facebook, and we evaluate its ability to identify threats on a mutated subset of this corpus. PlanAlyzer has both precision and recall of 92% on the mutated corpus, and 82% of the contrasts it generates match hand-specified data.
      pubtype: Periodical
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N