Processing Internet-derived Text—Creating a Corpus of Usenet Messages.
In recent years, linguists have become increasingly interested in the language of the Internet-both as an object of investigation as well as a source of authentic data to complement traditional electronic corpora. However, Internet-derived data is typically very messy data and a conversion process i...
| Publicado en: | Literary & Linguistic Computing Vol. 22; no. 2; pp. 151 - 166 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2007
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=25583831&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 25583831 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Jun2007 vid: 22 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 25583831 10.1093/llc/fqm002 ppf: 151 ppct: 15 formats: fmt: @attributes: type: P size: 256KB tig: atl: Processing Internet-derived Text—Creating a Corpus of Usenet Messages. aug: au: Hoffmann, Sebastian affil: Department of Linguistics and English Language, Bowland College, Lancaster University su: Telematics Written communication Corpora Language & languages Usenet (Computer network) sug: subj: Telematics Written communication Corpora Language & languages Usenet (Computer network) ab: In recent years, linguists have become increasingly interested in the language of the Internet-both as an object of investigation as well as a source of authentic data to complement traditional electronic corpora. However, Internet-derived data is typically very messy data and a conversion process is often required in order to enable researchers to carry out a reliable quantitative investigation of the patterns observed with the help of standard corpus tools. In this article, I discuss the technical and methodological aspects involved in creating a large corpus of asynchronous computer-mediated communication by downloading and post-processing hundreds of thousands messages posted in twelve Usenet newsgroups. After describing how messages can be arranged into hierarchically structured discussion threads, I focus at some length on the strategies that are required to correctly assign authorship to the different textual elements in individual messages. My algorithms have a success rate of well over 90% for most newsgroups and the resulting corpus can thus serve as a suitable basis for an investigation into the interactive strategies employed in this particular type of written communication. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2007 holdings: @attributes: islocal: N |
|---|