Representing variation in a spoken corpus of an endangered dialect: the case of Torlak.

The paper presents a spoken corpus of the endangered Torlak dialect from the Timok area of Southeast Serbia. This dialect expresses a great deal of variation in the use of non-standard features under the influence of standard Serbian (SSr). Accounting for this variation, a specific methodology has b...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 55; no. 3; pp. 731 - 757
Autor principal: Vuković, Teodora
Formato: Artículo
Publicado: Springer Nature Sep2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=151686298&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 151686298
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2021
      vid: 55
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        151686298
        10.1007/s10579-020-09522-4
      ppf: 731
      ppct: 26
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 976KB
      tig:
        atl: Representing variation in a spoken corpus of an endangered dialect: the case of Torlak.
      aug:
        au: Vuković, Teodora
        affil: Slavisches Seminar, University of Zurich, Plattenstrasse 43, 8032, Zurich, Switzerland
      su:
        Dialects
        Corpora
        Older people
        Linguistic change
        Semi-structured interviews
      sug:
        subj:
          Dialects
          Corpora
          Older people
          Linguistic change
          Semi-structured interviews
      keyword:
        Lemmatization
        Manual annotation
        Non-standard corpora
        Part-of-speech annotation
        Serbian
        Spoken corpora
        Torlak
      ab: The paper presents a spoken corpus of the endangered Torlak dialect from the Timok area of Southeast Serbia. This dialect expresses a great deal of variation in the use of non-standard features under the influence of standard Serbian (SSr). Accounting for this variation, a specific methodology has been selected for collection, sampling, transcription and annotation. Between 2015 and 2017, semi-structured interviews were conducted in the field eliciting spontaneous speech in the form of long narratives about traditional culture and history. The corpus comprises 500,697 tokens of semi-orthographic transcripts representing 80 h of recording from locations evenly distributed across the Timok area of the Torlak dialect zone, thus enabling a spatial contrastive analysis. The majority of speakers in the corpus are older people whose language represents the highly non-standard variety. In order to allow for analysis of language change under the influence of SSr, the corpus includes a number of younger people whose speech is closer to SSr. Tools for automatic PoS annotation and lemmatization that were lacking were developed based on the existing resources for SSr. For tagger training, a dialect sample of 27,000 manually verified tokens was merged with an existing training set for SSr.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N