Language chunking, data sparseness, and the value of a long marker list: explorations with word n-grams and authorial attribution.

The frequencies of individual words have been the mainstay of computer-assisted authorial attribution over the past three decades. The usefulness of this sort of data is attested in many benchmark trials and in numerous studies of particular authorship problems. It is sometimes argued, however, that...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 29; no. 2; pp. 147 - 164
Autores principales: Antonia, Alexis, Craig, Hugh, Elliott, Jack
Formato: Artículo
Publicado: Oxford University Press / USA Jun2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost