Measuring vocabulary diversity using dedicated software.

This paper describes software (vocd) that implements a solution to problems encountered in quantifying vocabulary diversity. Researchers in various fields of linguistic enquiry have calculated vocabulary diversity using the ratio of different words (Types) to total words (Tokens) - the Type-Token Ra...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 15; no. 3; pp. 323 - 339
Autores principales: McKee, G, Malvern, D, Richards, B
Formato: Artículo
Publicado: Oxford University Press / USA 2000
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80079476&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 80079476
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: 2000
      vid: 15
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        80079476
        10.1093/llc/15.3.323
      ppf: 323
      ppct: 16
      formats:
        fmt:
          @attributes:
            type: P
            size: 663KB
      tig:
        atl: Measuring vocabulary diversity using dedicated software.
      aug:
        au:
          McKee, G
          Malvern, D
          Richards, B
        affil: The University of Reading, School of Education, Bulmershe Court, Reading RG6 1HY, UK Corresponding author E-mail: b.j.richards@reading.ac.uk
      su:
        Vocabulary education
        Type & token (Linguistics)
        Language acquisition
        Language awareness in children
        Linguistic models
      sug:
        subj:
          Vocabulary education
          Type & token (Linguistics)
          Language acquisition
          Language awareness in children
          Linguistic models
      ab: This paper describes software (vocd) that implements a solution to problems encountered in quantifying vocabulary diversity. Researchers in various fields of linguistic enquiry have calculated vocabulary diversity using the ratio of different words (Types) to total words (Tokens) - the Type-Token Ratio (TTR) - or measures derived from it. Such measures are flawed, however, because the values obtained are related to the number of words in the sample. The paper shows how the relationship between TTR and sample size can be described by a new mathematical model, which in turn leads to an innovative method of measuring vocabulary diversity. The software automates measurement from transcripts prepared in a widely used computer-readable set of conventions: the CHAT format of the CHILDES project. Options in vocd are described to show how the user can determine which linguistic items will count as valid types and tokens in the analysis. The new measure is calculated by, first, randomly sampling words from the transcript to produce a curve of the TTR against Tokens for the empirical data. Then the software finds the best fit between this empirical curve and theoretical curves calculated from the model by adjusting the value of a parameter. The parameter, D, is shown to be a valid and reliable measure of vocabulary diversity without the problems of sample size found with previous methods.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2000
    holdings:
      @attributes:
        islocal: N