POS tagging of low-resource Pashto language: annotated corpus and BERT-based model.

This paper presents the development of a comprehensive part-of-speech (POS) annotated corpus for the low-resource Pashto language, along with a deep learning model for automatic POS tagging. The corpus comprises approximately 700K words (30K sentences), labeled for word boundaries, considering Pasht...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3243 - 3266
Autores principales: Haq, Ijazul, Zhang, Yingjie, Qadri, Intakhab Alam
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost