Towards neural named entity recognition system in Tigrinya with large-scale dataset.
The scarcity of annotated datasets remains a significant impediment to Natural Language Processing (NLP) advancement in low-resourced languages. In this work, we introduce a large-scale annotated Tigrinya Named Entity Recognition (NER) Corpus along with state-of-the-art models for Tigrinya NER. Our...
| Published in: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3311 - 3340 |
|---|---|
| Main Authors: | , , , , , , |
| Format: | Conference Paper/Materials |
| Published: |
Springer Nature
Sep2025
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909083&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909083 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909083 10.1007/s10579-025-09825-4 ppf: 3311 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 736KB tig: atl: Towards neural named entity recognition system in Tigrinya with large-scale dataset. aug: au: Berhane, Sham K. Beyene, Simon M. Teklit, Yoel G. Ibrahim, Ibrahim A. Teklu, Natnael A. Bereketeab, Sirak A. Gaim, Fitsum affil: Computer Science and Engineering, Mai Nefhi College of Engineering and Technology, Mai Nefhi, Eritrea https://ror.org/05apxxy63 School of Computing, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea su: Natural language processing Low-resource languages Data mining Morphology (Grammar) Artificial neural networks Transformer models Corpora sug: subj: Natural language processing Low-resource languages Data mining Morphology (Grammar) Artificial neural networks Transformer models Corpora keyword: Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences Language Named entity recognition Tigrinya Tigrinya NER Corpus TiRoBERTa ab: The scarcity of annotated datasets remains a significant impediment to Natural Language Processing (NLP) advancement in low-resourced languages. In this work, we introduce a large-scale annotated Tigrinya Named Entity Recognition (NER) Corpus along with state-of-the-art models for Tigrinya NER. Our human-labeled Tigrinya NER Corpus (TiNC24) comprises over 200K words tagged for NER with over 118K of the tokens also annotated with Parts-of-Speech (POS) tags, encompassing eight distinct classes of entities and multiple tagging schemes spanning ten diverse domains. We performed extensive experiments covering several recurrent neural networks and Transformer-based language models, achieving a highest performance of 90.18% weighted F1-score with the IO tagging scheme. These results are particularly notable given the unique challenges posed by Tigrinya's distinct grammatical structure and complex word morphology. This work establishes new benchmarks for Tigrinya NLP and provides essential resources for developing NER systems in other related morphologically rich, low-resourced languages. The dataset and models are made publicly available. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|