The Tomsk Dialect Corpus: a comprehensively annotated database of a Siberian Russian dialect from material collected over the last 70 years.

The paper offers the first full description of the Tomsk Dialect Corpus – an electronic resource based on recordings of the Russian dialect speech of the Tomsk and Kemerovo regions (West Siberia), which has been collected since 1946. The corpus counts 3,350,272 tokens, which makes it the largest ele...

Descripción completa

Detalles Bibliográficos
Publicado en:Russian Linguistics Vol. 47; no. 2; pp. 231 - 253
Autores principales: Zemicheva, Svetlana, Gromov, Maxim, Dubtsova, Ludmila, Ugryumova, Maria, Vasilchenko, Anna, Zyuz'kova, Natalia
Formato: Artículo
Publicado: Springer Nature Aug2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The paper offers the first full description of the Tomsk Dialect Corpus – an electronic resource based on recordings of the Russian dialect speech of the Tomsk and Kemerovo regions (West Siberia), which has been collected since 1946. The corpus counts 3,350,272 tokens, which makes it the largest electronic collection of dialect speech in Russia. The originality of this resource consists in the uniqueness of the materials collected and their multifaceted annotation. Topic and pragmatic annotations were created manually. Topic annotation is available for the whole data, whereas pragmatic annotation is available for 45,445 speech acts. Grammatical annotation was performed automatically with the PhpMorphy parser, with additional manual correction for some dialect words. Metalinguistic annotation includes the recording's year and place, and the speakers' age, gender, and educational level. All annotated parameters are searchable. The corpus also includes a lexicographic component, i.e. definitions of dialect lexemes.
Аннотация: В статье дается первое полное описание Томского диалектного корпуса – электронного ресурса на основе записей русской диалектной речи Томской и Кемеровской областей (Западная Сибирь), которые собирались с 1946 г. Корпус насчитывает 3 350 272 словоупотреблений и является крупнейшей электронной коллекцией диалектной речи в России. Оригинальность данного ресурса заключается в уникальности собранных материалов и их разносторонней разметке. Тематическая и прагматическая разметка были сделаны вручную. Тематическая разметка доступна для всего объёма материала, в рамках прагматической разметки выделено 45 445 речевых актов. Морфологическая аннотация сделана автоматически с помощью парсера PhpMorphy, дополнительно была выполнена ручная коррекция для некоторых диалектных слов. Металингвистическая аннотация включает год и место записи, возраст, пол и уровень образования говорящих. Все аннотированные параметры доступны для поиска. В состав корпуса также входит лексикографический компонент – толкования диалектных лексем.