Treating Words as Data with Error: Uncertainty in Text Statements of Policy Positions.

Political text offers extraordinary potential as a source of information about the policy positions of political actors. Despite recent advances in computational text analysis, human interpretative coding of text remains an important source of text-based data, ultimately required to validate more au...

Descripción completa

Detalles Bibliográficos
Publicado en:American Journal of Political Science Vol. 53; no. 2; pp. 495 - 514
Autores principales: Benoit, Kenneth, Laver, Michael, Mikhaylov, Slava
Formato: Artículo
Publicado: Wiley-Blackwell April 2009
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=511409014&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 511409014
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00925853
        APS
      jtl: American Journal of Political Science
      issn: 00925853
      maglogo: N
    pubinfo:
      dt: April 2009
      vid: 53
      iid: 2
      pid: 480
      pub: Wiley-Blackwell
    artinfo:
      ui:
        511409014
        10.1111/j.1540-5907.2009.00383.x
      ppf: 495
      ppct: 19
      formats:
      tig:
        atl: Treating Words as Data with Error: Uncertainty in Text Statements of Policy Positions.
      aug:
        au:
          Benoit, Kenneth
          Laver, Michael
          Mikhaylov, Slava
      su:
        Decision making
        Government policy
      sug:
        subj:
          Decision making
          Government policy
      ab: Political text offers extraordinary potential as a source of information about the policy positions of political actors. Despite recent advances in computational text analysis, human interpretative coding of text remains an important source of text-based data, ultimately required to validate more automatic techniques. The profession's main source of cross-national, time-series data on party policy positions comes from the human interpretative coding of party manifestos by the Comparative Manifesto Project (CMP). Despite widespread use of these data, the uncertainty associated with each point estimate has never been available, undermining the value of the dataset as a scientific resource. We propose a remedy. First, we characterize processes by which CMP data are generated. These include inherently stochastic processes of text authorship, as well as of the parsing and coding of observed text by humans. Second, we simulate these error-generating processes by bootstrapping analyses of coded quasi-sentences. This allows us to estimate precise levels of nonsystematic error for every category and scale reported by the CMP for its entire set of 3,000-plus manifestos. Using our estimates of these errors, we show how to correct biased inferences, in recent prominently published work, derived from statistical analyses of error-contaminated CMP data. Reprinted by permission of the publisher.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N