Multi-label document classification for Urdu: corpus and methods.
Automatically assigning various labels to language data is a core NLP task that has attracted the attention of researchers in multiple fields, for instance, medical records with diagnostic labels, products with categories, and so forth. All these tasks require an annotated multi-label corpus for sup...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 4; pp. 4253 - 4284 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912033&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 189912033 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2025 vid: 59 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 189912033 10.1007/s10579-025-09861-0 ppf: 4253 ppct: 31 formats: tig: atl: Multi-label document classification for Urdu: corpus and methods. aug: au: Shafi, Jawad Hussain, Amir Nawab, Rao M. Adeel Rasool, Madiha affil: https://ror.org/00nqqvk19 COMSATS University Islamabad, Lahore, Pakistan https://ror.org/03zjvnn91 Edinburgh Napier University, Edinburgh, UK su: Urdu language Natural language processing Corpora Machine learning Benchmark problems (Computer science) Deep learning Annotations sug: subj: Urdu language Natural language processing Corpora Machine learning Benchmark problems (Computer science) Deep learning Annotations keyword: DL Information and Computing Sciences Artificial Intelligence and Image Processing LLM Multi-label classification TL Urdu dataset ab: Automatically assigning various labels to language data is a core NLP task that has attracted the attention of researchers in multiple fields, for instance, medical records with diagnostic labels, products with categories, and so forth. All these tasks require an annotated multi-label corpus for supervised learning. Existing multi-label annotated corpora primarily focus on English, with a shortage in the Urdu language, which is widely spoken worldwide and whose digital text is increasing rapidly. To fill this gap, this research study presents the first large benchmark Urdu multi-label document corpus, which contains 600 documents from the field of journalism. The proposed corpus is manually annotated with 232 tags of a semantic classification scheme, with two to six tags per document to provide fine-grained categorization of the textual document. To evaluate the proposed corpus, we extracted content-based features and applied several multi-label classifiers. In addition, the results of the proposed content-based methods are compared with several deep and transfer learning-based methods. After conducting a comprehensive set of experiments, the best results are obtained using SOTA multi-label machine learning methods on word tri-gram (MicroPrecision = 0.62, MicroRecall = 0.48, MicroF1 = 0.62). To foster research in Urdu NLP, our proposed multi-label Urdu document classification corpus has been made publicly available for research and benchmarking purposes. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|