Part of speech (POS) tagging in Roman Urdu: datasets and models.
Roman Urdu is a prevalent medium of expression on social media, news websites, and text messages in the subcontinent, making it a valuable data source for social media and text analytics, particularly in the Indo-Pak perspective. However, despite the immense potential, limited efforts have been made...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 4; pp. 4285 - 4313 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912037&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 189912037 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2025 vid: 59 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 189912037 10.1007/s10579-025-09865-w ppf: 4285 ppct: 28 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.7MB tig: atl: Part of speech (POS) tagging in Roman Urdu: datasets and models. aug: au: Faheem, Ali Ullah, Faizad Azam, Ubaid Ayub, Muhammad Sohaib Karim, Asim affil: https://ror.org/05b5x4a35 Department of Computer Science, Lahore University of Management Science (LUMS), D.H.A, 54792, Lahore, Pakistan https://ror.org/01ryk1543 University of Southampton, Southampton, United Kingdom su: Natural language processing Language models Data libraries Text mining Categorization (Linguistics) Variation in language Urdu language Social media sug: subj: Natural language processing Language models Data libraries Text mining Categorization (Linguistics) Variation in language Urdu language Social media keyword: Communication and Culture Linguistics Language Low resource Part of speech Roman Urdu ab: Roman Urdu is a prevalent medium of expression on social media, news websites, and text messages in the subcontinent, making it a valuable data source for social media and text analytics, particularly in the Indo-Pak perspective. However, despite the immense potential, limited efforts have been made in the area of Roman Urdu text analytics due to various complexities, such as a lack of a standard lexicon, the informal nature of the text, and the lack of text processing tools. The development of the Roman Urdu Part-of-Speech (POS) dataset and the implementation of a robust tagger hold immense importance for text analytics in Roman Urdu. In this work, we created a comprehensive, large-scale Roman Urdu POS dataset and developed a Roman Urdu POS tagger, laying the foundation for future advancements in advanced text analysis. Our approach involved the utilization of Hidden Markov Models, Neural Networks, state-of-the-art transformer models, and Large Language Models as baselines. In our work, we curated two distinct test datasets: one with lexical variation and the other without such variation. This approach allowed us to test the model's robustness in handling different linguistic challenges posed by lexical variations. Our tagger yields high-quality output with an accuracy score of 96% without lexical variation and 86% on test data with lexical variations. We also evaluated state-of-the-art Large Language Models (GPT-4o and Llama-3-8B) in zero-shot and few-shot settings, with GPT-4o achieving up to 53.78% accuracy in the few-shot configuration, demonstrating a substantial performance gap compared to specialized models. This work establishes a comprehensive framework for Roman Urdu POS tagging that effectively addresses lexical variation challenges, providing essential resources and benchmarks for advancing Roman Urdu natural language processing research. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|