Toward Kurdish language processing: Experiments in collecting and processing the AsoSoft text corpus.
In this article, we introduce the first Kurdish text corpus for Central Kurdish (Sorani) branch, called AsoSoft text corpus. Kurdish language, which is spoken by more than 30 million people, has various dialects. As one of the two main branches of Kurdish, Central Kurdish is the formal dialect of Ku...
| Published in: | Digital Scholarship in the Humanities Vol. 35; no. 1; pp. 176 - 194 |
|---|---|
| Main Authors: | , , |
| Format: | Article |
| Published: |
Oxford University Press / USA
Apr2020
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142636780&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 142636780 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2020 vid: 35 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 142636780 10.1093/llc/fqy074 ppf: 176 ppct: 18 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2MB tig: atl: Toward Kurdish language processing: Experiments in collecting and processing the AsoSoft text corpus. aug: au: Veisi, Hadi MohammadAmini, Mohammad Hosseini, Hawre affil: Faculty of New Sciences and Technologies, University of Tehran, Iran IT Department, Faculty of Engineering, Tarbiat Modares University, Iran Electrical and Computer Engineering, Ryerson University, Canada su: Zipf's law Corpora Websites Publishing Standard language sug: subj: Zipf's law Corpora Websites Publishing Standard language ab: In this article, we introduce the first Kurdish text corpus for Central Kurdish (Sorani) branch, called AsoSoft text corpus. Kurdish language, which is spoken by more than 30 million people, has various dialects. As one of the two main branches of Kurdish, Central Kurdish is the formal dialect of Kurdish literature. AsoSoft text corpus is of size 188 million tokens and has been collected mostly from Web sites, published books, and magazines. The corpus has been normalized and converted into Text Encoding Initiative XML format. In both collecting and processing the text, we have faced several challenges and have proposed solutions to them. About 22% of the corpus is topic annotated with six topic tags, and a topic identification task has been done to evaluate the correctness of annotation. The computational experiments of the Central Kurdish text processing are also presented with the support of related supplementary statistics. For the first time, the validity of Zipf's law for Central Kurdish is presented and also perplexity of this language is calculated using standard N-gram language models. The perplexity of Central Kurdish is 276 for a tri-gram language model. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|