Word length frequency and distribution in English: Part I. Prose.

Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totalling a specified number of syllables) is given by dividing the total num...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 14; no. 3; pp. 339 - 359
Autores principales: Aoyama, H, Constable, J
Formato: Artículo
Publicado: Oxford University Press / USA 1999
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totalling a specified number of syllables) is given by dividing the total number of words by the mean number of syllables per word in the source text. This paper makes the latter point explicit mathematically, and in the course of this demonstration shows that the word length frequency totals in English prose output are distributed geometrically (previous researchers reported an adjusted Poisson distribution), and that the sequential distribution is random at the global level, with significant non-randomness in the fine structure. Data from a corpus of just under two million words and a syllable-count lexicon of 71,000 word forms is reported, together with some speculations concerning the relationship between the word length frequency distributions in output and in the lexicon. The pattern-matching theory is shown to be internally coherent, and it is observed that some of the analytical techniques described here form a satisfactory test for regular (isometric) lineation in a text.