More Data or Better Data? A Statistical Decision Problem.

When designing data collection, crucial questions arise regarding how much data to collect and how much effort to expend to enhance the quality of the collected data. To make choice of sample design a coherent subject of study, it is desirable to specify an explicit decision problem. We use theWald...

Descripción completa

Detalles Bibliográficos
Publicado en:Review of Economic Studies Vol. 84; no. 4; pp. 1583 - 1606
Autores principales: DOMINITZ, JEFF, MANSKI, CHARLES F.
Formato: Artículo
Publicado: Oxford University Press / USA Oct2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=125583555&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 125583555
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00346527
        REM
      jtl: Review of Economic Studies
      issn: 00346527
      maglogo: N
    pubinfo:
      dt: Oct2017
      vid: 84
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        125583555
        10.1093/restud/rdx005
      ppf: 1583
      ppct: 23
      formats:
      tig:
        atl: More Data or Better Data? A Statistical Decision Problem.
      aug:
        au:
          DOMINITZ, JEFF
          MANSKI, CHARLES F.
        affil:
          Resolution Economics
          Department of Economics and Institute for Policy Research, Northwestern University
      su:
        Acquisition of data
        Statistical decision making
        Sampling (Process)
        Data quality
        Mean square algorithms
      sug:
        subj:
          Acquisition of data
          Statistical decision making
          Sampling (Process)
          Data quality
          Mean square algorithms
      keyword:
        minimax regret
        missing data
        point prediction
        Sample design
        statistical decision theory
        minimax regret
        missing data
        point prediction
        Sample design
        statistical decision theory
      ab: When designing data collection, crucial questions arise regarding how much data to collect and how much effort to expend to enhance the quality of the collected data. To make choice of sample design a coherent subject of study, it is desirable to specify an explicit decision problem. We use theWald framework of statistical decision theory to study allocation of a budget between two or more sampling processes. These processes all draw random samples from a population of interest and aim to collect data that are informative about the sample realizations of an outcome. They differ in the cost of data collection and the quality of the data obtained. One may incur lower cost per sample member but yield lower data quality than another. Increasing the allocation of budget to a low-cost process yields more data, while increasing the allocation to a high-cost process yields better data. We initially view the concept of "better data" abstractly and then fix attention on two important cases. In both cases, a high-cost sampling process accurately measures the outcome of each sample member. The cases differ in the data yielded by a low-cost process. In one, the low-cost process has non-response and in the other it provides a low-resolution interval measure of each sample member's outcome. In these settings, we study minimax-regret sample design for prediction of a real-valued outcome under square loss; that is, design which minimizes maximum mean square error. The analysis imposes no assumptions that restrict the unobserved outcomes. Hence, the decision maker must cope with both the statistical imprecision of finite samples and the partial identification of the true state of nature.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N