Measuring vocabulary diversity using dedicated software.
This paper describes software (vocd) that implements a solution to problems encountered in quantifying vocabulary diversity. Researchers in various fields of linguistic enquiry have calculated vocabulary diversity using the ratio of different words (Types) to total words (Tokens) - the Type-Token Ra...
| Publicado en: | Literary & Linguistic Computing Vol. 15; no. 3; pp. 323 - 339 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
2000
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80079476&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 80079476 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: 2000 vid: 15 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 80079476 10.1093/llc/15.3.323 ppf: 323 ppct: 16 formats: fmt: @attributes: type: P size: 663KB tig: atl: Measuring vocabulary diversity using dedicated software. aug: au: McKee, G Malvern, D Richards, B affil: The University of Reading, School of Education, Bulmershe Court, Reading RG6 1HY, UK Corresponding author E-mail: b.j.richards@reading.ac.uk su: Vocabulary education Type & token (Linguistics) Language acquisition Language awareness in children Linguistic models sug: subj: Vocabulary education Type & token (Linguistics) Language acquisition Language awareness in children Linguistic models ab: This paper describes software (vocd) that implements a solution to problems encountered in quantifying vocabulary diversity. Researchers in various fields of linguistic enquiry have calculated vocabulary diversity using the ratio of different words (Types) to total words (Tokens) - the Type-Token Ratio (TTR) - or measures derived from it. Such measures are flawed, however, because the values obtained are related to the number of words in the sample. The paper shows how the relationship between TTR and sample size can be described by a new mathematical model, which in turn leads to an innovative method of measuring vocabulary diversity. The software automates measurement from transcripts prepared in a widely used computer-readable set of conventions: the CHAT format of the CHILDES project. Options in vocd are described to show how the user can determine which linguistic items will count as valid types and tokens in the analysis. The new measure is calculated by, first, randomly sampling words from the transcript to produce a curve of the TTR against Tokens for the empirical data. Then the software finds the best fit between this empirical curve and theoretical curves calculated from the model by adjusting the value of a parameter. The parameter, D, is shown to be a valid and reliable measure of vocabulary diversity without the problems of sample size found with previous methods. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2000 holdings: @attributes: islocal: N |
|---|