POS tagging of low-resource Pashto language: annotated corpus and BERT-based model.
This paper presents the development of a comprehensive part-of-speech (POS) annotated corpus for the low-resource Pashto language, along with a deep learning model for automatic POS tagging. The corpus comprises approximately 700K words (30K sentences), labeled for word boundaries, considering Pasht...
| Published in: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3243 - 3266 |
|---|---|
| Main Authors: | , , |
| Format: | Conference Paper/Materials |
| Published: |
Springer Nature
Sep2025
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909090&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909090 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909090 10.1007/s10579-025-09834-3 ppf: 3243 ppct: 23 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.9MB tig: atl: POS tagging of low-resource Pashto language: annotated corpus and BERT-based model. aug: au: Haq, Ijazul Zhang, Yingjie Qadri, Intakhab Alam affil: https://ror.org/0530pts50 Shien-Ming Wu School of Intelligent Manufacturing, South China University of Technology, Guangzhou, China Artificial Intelligence Department, Guangdong CAS Angels Biotechnology Co. Ltd., Foshan, China https://ror.org/01vy4gh70 College of Computer Science and Software Engineering, Shenzhen University, 518060, Shenzhen, China su: Natural language processing Deep learning Grammatical categories Statistical accuracy Corpora Language models sug: subj: Natural language processing Deep learning Grammatical categories Statistical accuracy Corpora Language models keyword: Artificial Intelligence BERT Communication and Culture Linguistics Corpus Linguistics LLMs Low-resource Languages Machine Learning NLP Pashto POS Tagging Psychology and Cognitive Sciences Cognitive Sciences Language Transformers ab: This paper presents the development of a comprehensive part-of-speech (POS) annotated corpus for the low-resource Pashto language, along with a deep learning model for automatic POS tagging. The corpus comprises approximately 700K words (30K sentences), labeled for word boundaries, considering Pashto lacks explicit delimiters for word segmentation. The corpus was then annotated for POS information using a concise and pragmatic tagset of 36 grammatical categories. We utilized this corpus to train a supervised POS tagging model. For the model development, we employed a multilingual BERT (Bidirectional Encoder Representations from Transformers), fine-tuning it for this specific task. The BERT-based model performance was evaluated against an RNN-based model that employs the BiLSTM-CRF network and word embeddings (Word2Vec, fastText, and GloVe). Experimental results demonstrated that the BERT-based model achieved the best performance, attaining an accuracy of 96.24% and an F1-score (weighted average) of 96.22%. The model's performance is highly satisfactory, making it useful in practical applications. Furthermore, the applications of the annotated corpus are not limited to this study only; it can be employed in various NLP applications, including named entity recognition (NER), text proofing, and constituency and dependency parsing. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|