The Electronic Corpus of 17th- and 18th-century Polish Texts.

The paper describes the process of building the electronic corpus of 17th- and 18th-century Polish texts, a relatively large, balanced, structurally and morphologically annotated resource of the Middle Polish language, available for searching at https://www.korba.edu.pl. The corpus consists of sampl...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 56; no. 1; pp. 309 - 333
Main Authors: Gruszczyński, Włodzimierz, Adamiec, Dorota, Bronikowska, Renata, Kieraś, Witold, Modrzejewski, Emanuel, Wieczorek, Aleksandra, Woliński, Marcin
Format: Article
Published: Springer Nature Mar2022
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:The paper describes the process of building the electronic corpus of 17th- and 18th-century Polish texts, a relatively large, balanced, structurally and morphologically annotated resource of the Middle Polish language, available for searching at https://www.korba.edu.pl. The corpus consists of samples extracted from over seven hundred texts written and published between 1601 and 1772, summing up to a total size of 13.5 million tokens which makes it one of the largest historical corpora for a Slavic language.