Processing Internet-derived Text—Creating a Corpus of Usenet Messages.

In recent years, linguists have become increasingly interested in the language of the Internet-both as an object of investigation as well as a source of authentic data to complement traditional electronic corpora. However, Internet-derived data is typically very messy data and a conversion process i...

Full description

Bibliographic Details
Published in:Literary & Linguistic Computing Vol. 22; no. 2; pp. 151 - 166
Main Author: Hoffmann, Sebastian
Format: Article
Published: Oxford University Press / USA Jun2007
Subjects:
Online Access:View this record in EBSCOhost