Urdu part of speech tagging using conditional random fields.
Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the...
| Publicado en: | Language Resources & Evaluation Vol. 53; no. 3; pp. 331 - 363 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=138297593&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 138297593 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2019 vid: 53 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 138297593 10.1007/s10579-018-9439-6 ppf: 331 ppct: 32 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.4MB tig: atl: Urdu part of speech tagging using conditional random fields. aug: au: Khan, Wahab Daud, Ali Nasir, Jamal Abdul Amjad, Tehmina Arafat, Sachi Aljohani, Naif Alotaibi, Fahd S. affil: Department of Computer Science and Software Engineering, IIU, 44000, Islamabad, Pakistan Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah, Saudi Arabia su: Conditional random fields Parts of speech Natural language processing Support vector machines Data mining sug: subj: Conditional random fields Parts of speech Natural language processing Support vector machines Data mining keyword: Conditional random field (CRF) Part of speech (POS) Support vector machine (SVM) Urdu ab: Part of speech (POS) tagging, the assignment of syntactic categories for words in running text, is significant to natural language processing as a preliminary task in applications such as speech processing, information extraction, and others. Urdu language processing presents a challenge due to the dual behaviour of various Urdu POS tags in differing situations (morphosyntactic ambiguity). This paper addresses this challenge by developing a novel tagging approach using linear-chain conditional random fields (CRF). Our work is the first instance of a CRF approach for Urdu POS tagging. The proposed model employs a strong, stable and balanced language-independent as well as language dependent feature set. The language-dependent feature considered includes part-of-speech tag of the previous word and suffix of the current word while the language-independent features includes the 'context words window'. Our approach was evaluated against support vector machine techniques for Urdu POS—considered as state of the art—on two benchmark datasets. The results show our CRF approach to improve upon the F-measure of prior attempts by 8.3–8.5%. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|