Compilation, transcription and usage of a reference speech corpus: the case of the Slovene corpus GOS.

In recent years, building reference speech corpora was an important part of the activities which provided the necessary linguistic infrastructure in many European countries, for languages with many speakers (e.g., French, German, Spanish, Italian) as well as for those with smaller numbers of speaker...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 47; no. 4; pp. 1031 - 1049
Main Authors: Verdonik, Darinka, Kosem, Iztok, Vitez, Ana Zwitter, Krek, Simon, Stabej, Marko
Format: Article
Published: Springer Nature Dec2013
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=92719896&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 92719896
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2013
      vid: 47
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        92719896
        10.1007/s10579-013-9216-5
      ppf: 1031
      ppct: 18
      formats:
        fmt:
          @attributes:
            type: P
            size: 278KB
      tig:
        atl: Compilation, transcription and usage of a reference speech corpus: the case of the Slovene corpus GOS.
      aug:
        au:
          Verdonik, Darinka
          Kosem, Iztok
          Vitez, Ana Zwitter
          Krek, Simon
          Stabej, Marko
        affil:
          University of Maribor, Maribor, Slovenia
          Trojina, Institute for Applied Slovene Studies, Škofja Loka, Slovenia
          Amebis, d.o.o., Kamnik, Slovenia
          University of Ljubljana, Ljubljana, Slovenia
      su:
        Transcription (Linguistics)
        Reference (Linguistics)
        Oral communication
        Corpora
        Slovenes
      sug:
        subj:
          Transcription (Linguistics)
          Reference (Linguistics)
          Oral communication
          Corpora
          Slovenes
      keyword:
        Discourse
        Recordings
        Spoken language
        Transcription conventions
        Web concordancer
      ab: In recent years, building reference speech corpora was an important part of the activities which provided the necessary linguistic infrastructure in many European countries, for languages with many speakers (e.g., French, German, Spanish, Italian) as well as for those with smaller numbers of speakers (e.g., Swedish, Dutch, Czech, Slovak). This paper describes the process of the creation of a reference speech corpus and its distribution to potential users, as it was done in the case of the Slovene corpus GOS. The corpus structure and fieldwork experiences with recording, labelling system, and two levels of transcription (pronunciation-based and standardized) are described, as well as the main characteristics of the corpus interface (web concordancer) and the availability of the original corpus files.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N