Language chunking, data sparseness, and the value of a long marker list: explorations with word n-grams and authorial attribution.

The frequencies of individual words have been the mainstay of computer-assisted authorial attribution over the past three decades. The usefulness of this sort of data is attested in many benchmark trials and in numerous studies of particular authorship problems. It is sometimes argued, however, that...

Full description

Bibliographic Details
Published in:Literary & Linguistic Computing Vol. 29; no. 2; pp. 147 - 164
Main Authors: Antonia, Alexis, Craig, Hugh, Elliott, Jack
Format: Article
Published: Oxford University Press / USA Jun2014
Subjects:
Online Access:View this record in EBSCOhost