A hybrid part-of-speech tagger with annotated Kurdish corpus: advancements in POS tagging.
With the rapid growth of online content written in the Kurdish language, there is an increasing need to make it machine-readable and processable. Part of speech (POS) tagging is a critical aspect of natural language processing (NLP), playing a significant role in applications such as speech recognit...
| Published in: | Digital Scholarship in the Humanities Vol. 38; no. 4; pp. 1604 - 1613 |
|---|---|
| Main Authors: | , , |
| Format: | Article |
| Published: |
Oxford University Press / USA
Dec2023
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=174444647&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 174444647 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Dec2023 vid: 38 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 174444647 10.1093/llc/fqad066 ppf: 1604 ppct: 9 formats: fmt: – @attributes: type: T – @attributes: type: P size: 507KB tig: atl: A hybrid part-of-speech tagger with annotated Kurdish corpus: advancements in POS tagging. aug: au: Maulud, Dastan Jacksi, Karwan Ali, Ismael affil: Department of Information Technology, Technical College of Informatics-Akre, Duhok Polytechnic University , Duhok, Kurdistan Region, Iraq Department of Computer Science, University of Zakho , Duhok, Kurdistan Region, Iraq su: Natural language processing Hidden Markov models Natural languages Parts of speech Corpora sug: subj: Natural language processing Hidden Markov models Natural languages Parts of speech Corpora keyword: bigram HMM machine readability natural language processing part of speech tagging rule-based approach speech recognition text corpus ab: With the rapid growth of online content written in the Kurdish language, there is an increasing need to make it machine-readable and processable. Part of speech (POS) tagging is a critical aspect of natural language processing (NLP), playing a significant role in applications such as speech recognition, natural language parsing, information retrieval, and multiword term extraction. This study details the creation of the DASTAN corpus, the first POS-annotated corpus for the Sorani Kurdish dialect. The corpus, containing 74,258 words and thirty-eight tags, employs a hybrid approach utilizing the bigram hidden Markov model in combination with the Kurdish rule-based approach to POS tagging. This approach addresses two key problems that arise with rule-based approaches, namely misclassified words and ambiguity-related unanalyzed words. The proposed approach's accuracy was assessed by training and testing it on the DASTAN corpus, yielding a 96% accuracy rate. Overall, this study's findings demonstrate the effectiveness of the proposed hybrid approach and its potential to enhance NLP applications for Sorani Kurdish. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|