Word length frequency and distribution in English: Part I. Prose.
Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totalling a specified number of syllables) is given by dividing the total num...
| Publicado en: | Literary & Linguistic Computing Vol. 14; no. 3; pp. 339 - 359 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
1999
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80066721&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 80066721 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: 1999 vid: 14 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 80066721 10.1093/llc/14.3.339 ppf: 339 ppct: 20 formats: fmt: @attributes: type: P size: 941KB tig: atl: Word length frequency and distribution in English: Part I. Prose. aug: au: Aoyama, H Constable, J affil: Faculty of Integrated Human Studies, Kyoto University, Kyoto 606-8501, Japan Corresponding author E-mail: aoyama@phys.h.kyoto-u.ac.jp su: Syllable (Grammar) Poisson distribution English prose literature Lexicon Verse satire sug: subj: Syllable (Grammar) Poisson distribution English prose literature Lexicon Verse satire ab: Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totalling a specified number of syllables) is given by dividing the total number of words by the mean number of syllables per word in the source text. This paper makes the latter point explicit mathematically, and in the course of this demonstration shows that the word length frequency totals in English prose output are distributed geometrically (previous researchers reported an adjusted Poisson distribution), and that the sequential distribution is random at the global level, with significant non-randomness in the fine structure. Data from a corpus of just under two million words and a syllable-count lexicon of 71,000 word forms is reported, together with some speculations concerning the relationship between the word length frequency distributions in output and in the lexicon. The pattern-matching theory is shown to be internally coherent, and it is observed that some of the analytical techniques described here form a satisfactory test for regular (isometric) lineation in a text. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 1999 holdings: @attributes: islocal: N |
|---|