The Tomsk Dialect Corpus: a comprehensively annotated database of a Siberian Russian dialect from material collected over the last 70 years.

The paper offers the first full description of the Tomsk Dialect Corpus – an electronic resource based on recordings of the Russian dialect speech of the Tomsk and Kemerovo regions (West Siberia), which has been collected since 1946. The corpus counts 3,350,272 tokens, which makes it the largest ele...

Full description

Bibliographic Details
Published in:Russian Linguistics Vol. 47; no. 2; pp. 231 - 253
Main Authors: Zemicheva, Svetlana, Gromov, Maxim, Dubtsova, Ludmila, Ugryumova, Maria, Vasilchenko, Anna, Zyuz'kova, Natalia
Format: Article
Published: Springer Nature Aug2023
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=169810710&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 169810710
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        03043487
        3W1
      jtl: Russian Linguistics
      issn: 03043487
      maglogo: N
    pubinfo:
      dt: Aug2023
      vid: 47
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        169810710
        10.1007/s11185-023-09277-w
      ppf: 231
      ppct: 22
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: The Tomsk Dialect Corpus: a comprehensively annotated database of a Siberian Russian dialect from material collected over the last 70 years.
      aug:
        au:
          Zemicheva, Svetlana
          Gromov, Maxim
          Dubtsova, Ludmila
          Ugryumova, Maria
          Vasilchenko, Anna
          Zyuz'kova, Natalia
        affil:
          Linguistic Convergence Laboratory, HSE University, Moscow, Russia
          Radiophysics faculty, National Research Tomsk State University, Tomsk, Russia
          Department of Russian as a Foreign Language, National Research Tomsk State University, Tomsk, Russia
          Institute of Education, National Research Tomsk State University, Tomsk, Russia
          Laboratory of General and Siberian Lexicography, Department of Russian as a Foreign Language, National Research Tomsk State University, Tomsk, Russia
      su:
        Databases
        Dialects
        Speech
        Corpora
        Educational attainment
        Tomsk (Russia)
        Kemerovo (Russia)
        Russia
      sug:
        subj:
          Tomsk (Russia)
          Kemerovo (Russia)
          Russia
          Databases
          Dialects
          Speech
          Corpora
          Educational attainment
      ab:
        The paper offers the first full description of the Tomsk Dialect Corpus – an electronic resource based on recordings of the Russian dialect speech of the Tomsk and Kemerovo regions (West Siberia), which has been collected since 1946. The corpus counts 3,350,272 tokens, which makes it the largest electronic collection of dialect speech in Russia. The originality of this resource consists in the uniqueness of the materials collected and their multifaceted annotation. Topic and pragmatic annotations were created manually. Topic annotation is available for the whole data, whereas pragmatic annotation is available for 45,445 speech acts. Grammatical annotation was performed automatically with the PhpMorphy parser, with additional manual correction for some dialect words. Metalinguistic annotation includes the recording's year and place, and the speakers' age, gender, and educational level. All annotated parameters are searchable. The corpus also includes a lexicographic component, i.e. definitions of dialect lexemes.
        Аннотация: В статье дается первое полное описание Томского диалектного корпуса – электронного ресурса на основе записей русской диалектной речи Томской и Кемеровской областей (Западная Сибирь), которые собирались с 1946 г. Корпус насчитывает 3 350 272 словоупотреблений и является крупнейшей электронной коллекцией диалектной речи в России. Оригинальность данного ресурса заключается в уникальности собранных материалов и их разносторонней разметке. Тематическая и прагматическая разметка были сделаны вручную. Тематическая разметка доступна для всего объёма материала, в рамках прагматической разметки выделено 45 445 речевых актов. Морфологическая аннотация сделана автоматически с помощью парсера PhpMorphy, дополнительно была выполнена ручная коррекция для некоторых диалектных слов. Металингвистическая аннотация включает год и место записи, возраст, пол и уровень образования говорящих. Все аннотированные параметры доступны для поиска. В состав корпуса также входит лексикографический компонент – толкования диалектных лексем.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Russian Linguistics is a copyright of Springer, 2023. All Rights Reserved.
      item: Russian Linguistics
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N