Towards neural named entity recognition system in Tigrinya with large-scale dataset.

The scarcity of annotated datasets remains a significant impediment to Natural Language Processing (NLP) advancement in low-resourced languages. In this work, we introduce a large-scale annotated Tigrinya Named Entity Recognition (NER) Corpus along with state-of-the-art models for Tigrinya NER. Our...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 59; no. 3; pp. 3311 - 3340
Main Authors: Berhane, Sham K., Beyene, Simon M., Teklit, Yoel G., Ibrahim, Ibrahim A., Teklu, Natnael A., Bereketeab, Sirak A., Gaim, Fitsum
Format: Conference Paper/Materials
Published: Springer Nature Sep2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909083&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909083
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909083
        10.1007/s10579-025-09825-4
      ppf: 3311
      ppct: 29
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 736KB
      tig:
        atl: Towards neural named entity recognition system in Tigrinya with large-scale dataset.
      aug:
        au:
          Berhane, Sham K.
          Beyene, Simon M.
          Teklit, Yoel G.
          Ibrahim, Ibrahim A.
          Teklu, Natnael A.
          Bereketeab, Sirak A.
          Gaim, Fitsum
        affil:
          Computer Science and Engineering, Mai Nefhi College of Engineering and Technology, Mai Nefhi, Eritrea
          https://ror.org/05apxxy63 School of Computing, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea
      su:
        Natural language processing
        Low-resource languages
        Data mining
        Morphology (Grammar)
        Artificial neural networks
        Transformer models
        Corpora
      sug:
        subj:
          Natural language processing
          Low-resource languages
          Data mining
          Morphology (Grammar)
          Artificial neural networks
          Transformer models
          Corpora
      keyword:
        Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Cognitive Sciences
        Language
        Named entity recognition
        Tigrinya
        Tigrinya NER Corpus
        TiRoBERTa
      ab: The scarcity of annotated datasets remains a significant impediment to Natural Language Processing (NLP) advancement in low-resourced languages. In this work, we introduce a large-scale annotated Tigrinya Named Entity Recognition (NER) Corpus along with state-of-the-art models for Tigrinya NER. Our human-labeled Tigrinya NER Corpus (TiNC24) comprises over 200K words tagged for NER with over 118K of the tokens also annotated with Parts-of-Speech (POS) tags, encompassing eight distinct classes of entities and multiple tagging schemes spanning ten diverse domains. We performed extensive experiments covering several recurrent neural networks and Transformer-based language models, achieving a highest performance of 90.18% weighted F1-score with the IO tagging scheme. These results are particularly notable given the unique challenges posed by Tigrinya's distinct grammatical structure and complex word morphology. This work establishes new benchmarks for Tigrinya NLP and provides essential resources for developing NER systems in other related morphologically rich, low-resourced languages. The dataset and models are made publicly available.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N