Toward Kurdish language processing: Experiments in collecting and processing the AsoSoft text corpus.

In this article, we introduce the first Kurdish text corpus for Central Kurdish (Sorani) branch, called AsoSoft text corpus. Kurdish language, which is spoken by more than 30 million people, has various dialects. As one of the two main branches of Kurdish, Central Kurdish is the formal dialect of Ku...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 35; no. 1; pp. 176 - 194
Main Authors: Veisi, Hadi, MohammadAmini, Mohammad, Hosseini, Hawre
Format: Article
Published: Oxford University Press / USA Apr2020
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142636780&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 142636780
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2020
      vid: 35
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        142636780
        10.1093/llc/fqy074
      ppf: 176
      ppct: 18
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2MB
      tig:
        atl: Toward Kurdish language processing: Experiments in collecting and processing the AsoSoft text corpus.
      aug:
        au:
          Veisi, Hadi
          MohammadAmini, Mohammad
          Hosseini, Hawre
        affil:
          Faculty of New Sciences and Technologies, University of Tehran, Iran
          IT Department, Faculty of Engineering, Tarbiat Modares University, Iran
          Electrical and Computer Engineering, Ryerson University, Canada
      su:
        Zipf's law
        Corpora
        Websites
        Publishing
        Standard language
      sug:
        subj:
          Zipf's law
          Corpora
          Websites
          Publishing
          Standard language
      ab: In this article, we introduce the first Kurdish text corpus for Central Kurdish (Sorani) branch, called AsoSoft text corpus. Kurdish language, which is spoken by more than 30 million people, has various dialects. As one of the two main branches of Kurdish, Central Kurdish is the formal dialect of Kurdish literature. AsoSoft text corpus is of size 188 million tokens and has been collected mostly from Web sites, published books, and magazines. The corpus has been normalized and converted into Text Encoding Initiative XML format. In both collecting and processing the text, we have faced several challenges and have proposed solutions to them. About 22% of the corpus is topic annotated with six topic tags, and a topic identification task has been done to evaluate the correctness of annotation. The computational experiments of the Central Kurdish text processing are also presented with the support of related supplementary statistics. For the first time, the validity of Zipf's law for Central Kurdish is presented and also perplexity of this language is calculated using standard N-gram language models. The perplexity of Central Kurdish is 276 for a tri-gram language model.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N