Pitfalls in Corpus Research.

This paper discusses some pitfalls in corpus research and suggests solutions on the basis of examples and computer simulations. We first address reliability problems in language transcriptions, agreement between transcribers, and how disagreements can be dealt with. We then show that the frequencies...

Full description

Bibliographic Details
Published in:Computers & the Humanities Vol. 38; no. 4; pp. 343 - 363
Main Authors: Rietveld, Toni, Van Hout, Roeland, Ernestus, Mirjam
Format: Article
Published: Springer Nature Nov2004
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:This paper discusses some pitfalls in corpus research and suggests solutions on the basis of examples and computer simulations. We first address reliability problems in language transcriptions, agreement between transcribers, and how disagreements can be dealt with. We then show that the frequencies of occurrence obtained from a corpus cannot always be analyzed with the traditional X² test, as corpus data are often not sequentially independent and unit independent. Next, we stress the relevance of the power of statistical tests, and the sizes of statistically significant effects. Finally, we point out that at-test based on log odds often provides a better alternative to a X² analysis based on frequency counts.