Part of speech (POS) tagging in Roman Urdu: datasets and models.

Roman Urdu is a prevalent medium of expression on social media, news websites, and text messages in the subcontinent, making it a valuable data source for social media and text analytics, particularly in the Indo-Pak perspective. However, despite the immense potential, limited efforts have been made...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 4; pp. 4285 - 4313
Autores principales: Faheem, Ali, Ullah, Faizad, Azam, Ubaid, Ayub, Muhammad Sohaib, Karim, Asim
Formato: Artículo
Publicado: Springer Nature Dec2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912037&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 189912037
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 59
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        189912037
        10.1007/s10579-025-09865-w
      ppf: 4285
      ppct: 28
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: Part of speech (POS) tagging in Roman Urdu: datasets and models.
      aug:
        au:
          Faheem, Ali
          Ullah, Faizad
          Azam, Ubaid
          Ayub, Muhammad Sohaib
          Karim, Asim
        affil:
          https://ror.org/05b5x4a35 Department of Computer Science, Lahore University of Management Science (LUMS), D.H.A, 54792, Lahore, Pakistan
          https://ror.org/01ryk1543 University of Southampton, Southampton, United Kingdom
      su:
        Natural language processing
        Language models
        Data libraries
        Text mining
        Categorization (Linguistics)
        Variation in language
        Urdu language
        Social media
      sug:
        subj:
          Natural language processing
          Language models
          Data libraries
          Text mining
          Categorization (Linguistics)
          Variation in language
          Urdu language
          Social media
      keyword:
        Communication and Culture Linguistics
        Language
        Low resource
        Part of speech
        Roman Urdu
      ab: Roman Urdu is a prevalent medium of expression on social media, news websites, and text messages in the subcontinent, making it a valuable data source for social media and text analytics, particularly in the Indo-Pak perspective. However, despite the immense potential, limited efforts have been made in the area of Roman Urdu text analytics due to various complexities, such as a lack of a standard lexicon, the informal nature of the text, and the lack of text processing tools. The development of the Roman Urdu Part-of-Speech (POS) dataset and the implementation of a robust tagger hold immense importance for text analytics in Roman Urdu. In this work, we created a comprehensive, large-scale Roman Urdu POS dataset and developed a Roman Urdu POS tagger, laying the foundation for future advancements in advanced text analysis. Our approach involved the utilization of Hidden Markov Models, Neural Networks, state-of-the-art transformer models, and Large Language Models as baselines. In our work, we curated two distinct test datasets: one with lexical variation and the other without such variation. This approach allowed us to test the model's robustness in handling different linguistic challenges posed by lexical variations. Our tagger yields high-quality output with an accuracy score of 96% without lexical variation and 86% on test data with lexical variations. We also evaluated state-of-the-art Large Language Models (GPT-4o and Llama-3-8B) in zero-shot and few-shot settings, with GPT-4o achieving up to 53.78% accuracy in the few-shot configuration, demonstrating a substantial performance gap compared to specialized models. This work establishes a comprehensive framework for Roman Urdu POS tagging that effectively addresses lexical variation challenges, providing essential resources and benchmarks for advancing Roman Urdu natural language processing research.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N