Processing Internet-derived Text—Creating a Corpus of Usenet Messages.

In recent years, linguists have become increasingly interested in the language of the Internet-both as an object of investigation as well as a source of authentic data to complement traditional electronic corpora. However, Internet-derived data is typically very messy data and a conversion process i...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 22; no. 2; pp. 151 - 166
Autor principal: Hoffmann, Sebastian
Formato: Artículo
Publicado: Oxford University Press / USA Jun2007
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=25583831&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 25583831
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Jun2007
      vid: 22
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        25583831
        10.1093/llc/fqm002
      ppf: 151
      ppct: 15
      formats:
        fmt:
          @attributes:
            type: P
            size: 256KB
      tig:
        atl: Processing Internet-derived Text—Creating a Corpus of Usenet Messages.
      aug:
        au: Hoffmann, Sebastian
        affil: Department of Linguistics and English Language, Bowland College, Lancaster University
      su:
        Telematics
        Written communication
        Corpora
        Language & languages
        Usenet (Computer network)
      sug:
        subj:
          Telematics
          Written communication
          Corpora
          Language & languages
          Usenet (Computer network)
      ab: In recent years, linguists have become increasingly interested in the language of the Internet-both as an object of investigation as well as a source of authentic data to complement traditional electronic corpora. However, Internet-derived data is typically very messy data and a conversion process is often required in order to enable researchers to carry out a reliable quantitative investigation of the patterns observed with the help of standard corpus tools. In this article, I discuss the technical and methodological aspects involved in creating a large corpus of asynchronous computer-mediated communication by downloading and post-processing hundreds of thousands messages posted in twelve Usenet newsgroups. After describing how messages can be arranged into hierarchically structured discussion threads, I focus at some length on the strategies that are required to correctly assign authorship to the different textual elements in individual messages. My algorithms have a success rate of well over 90% for most newsgroups and the resulting corpus can thus serve as a suitable basis for an investigation into the interactive strategies employed in this particular type of written communication.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2007
    holdings:
      @attributes:
        islocal: N