A hybrid part-of-speech tagger with annotated Kurdish corpus: advancements in POS tagging.

With the rapid growth of online content written in the Kurdish language, there is an increasing need to make it machine-readable and processable. Part of speech (POS) tagging is a critical aspect of natural language processing (NLP), playing a significant role in applications such as speech recognit...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 38; no. 4; pp. 1604 - 1613
Main Authors: Maulud, Dastan, Jacksi, Karwan, Ali, Ismael
Format: Article
Published: Oxford University Press / USA Dec2023
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=174444647&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 174444647
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Dec2023
      vid: 38
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        174444647
        10.1093/llc/fqad066
      ppf: 1604
      ppct: 9
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 507KB
      tig:
        atl: A hybrid part-of-speech tagger with annotated Kurdish corpus: advancements in POS tagging.
      aug:
        au:
          Maulud, Dastan
          Jacksi, Karwan
          Ali, Ismael
        affil:
          Department of Information Technology, Technical College of Informatics-Akre, Duhok Polytechnic University , Duhok, Kurdistan Region, Iraq
          Department of Computer Science, University of Zakho , Duhok, Kurdistan Region, Iraq
      su:
        Natural language processing
        Hidden Markov models
        Natural languages
        Parts of speech
        Corpora
      sug:
        subj:
          Natural language processing
          Hidden Markov models
          Natural languages
          Parts of speech
          Corpora
      keyword:
        bigram HMM
        machine readability
        natural language processing
        part of speech tagging
        rule-based approach
        speech recognition
        text corpus
      ab: With the rapid growth of online content written in the Kurdish language, there is an increasing need to make it machine-readable and processable. Part of speech (POS) tagging is a critical aspect of natural language processing (NLP), playing a significant role in applications such as speech recognition, natural language parsing, information retrieval, and multiword term extraction. This study details the creation of the DASTAN corpus, the first POS-annotated corpus for the Sorani Kurdish dialect. The corpus, containing 74,258 words and thirty-eight tags, employs a hybrid approach utilizing the bigram hidden Markov model in combination with the Kurdish rule-based approach to POS tagging. This approach addresses two key problems that arise with rule-based approaches, namely misclassified words and ambiguity-related unanalyzed words. The proposed approach's accuracy was assessed by training and testing it on the DASTAN corpus, yielding a 96% accuracy rate. Overall, this study's findings demonstrate the effectiveness of the proposed hybrid approach and its potential to enhance NLP applications for Sorani Kurdish.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N