Uzbek news corpus for named entity recognition.

We have presented a corpus of Uzbek news articles containing manually annotated named entities. The corpus comprises 500 articles (222,536 tokens) and three entity classes (person, location, organization) sourced from Qalampir, an online news source in Uzbekistan. This corpus can be used for develop...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3139 - 3153
Autores principales: Yusufu, Aizihaierjiang, Aziz, Kamran, Yusufu, Aizierguli, Ainiwaer, Abidan, Li, Fei, Ji, Donghong
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909047&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909047
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909047
        10.1007/s10579-024-09786-0
      ppf: 3139
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 612KB
      tig:
        atl: Uzbek news corpus for named entity recognition.
      aug:
        au:
          Yusufu, Aizihaierjiang
          Aziz, Kamran
          Yusufu, Aizierguli
          Ainiwaer, Abidan
          Li, Fei
          Ji, Donghong
        affil:
          https://ror.org/033vjfk17 Key Laboratory of Aerospace Information Security and Trusted Computing, Wuhan University, No. 129 Luoyu Road, 430000, Wuhan, Hubei, China
          https://ror.org/00ndrvk93 School of Computer Science and Technology, Xinjiang Normal University, No. 102 Xinyi Road, 830054, Urumqi, Xinjiang, China
          https://ror.org/033vjfk17 School of Information Management, Wuhan University, No. 129 Luoyu Road, 430000, Wuhan, Hubei, China
      su:
        Natural language processing
        Corpora
        Turkic languages
        Data mining
        Uzbekistan
      sug:
        subj:
          Uzbekistan
          Natural language processing
          Corpora
          Turkic languages
          Data mining
      keyword:
        Entity naming recognition
        Newswire
        Uzbek
      ab: We have presented a corpus of Uzbek news articles containing manually annotated named entities. The corpus comprises 500 articles (222,536 tokens) and three entity classes (person, location, organization) sourced from Qalampir, an online news source in Uzbekistan. This corpus can be used for develop and evaluate natural language processing (NLP) models for Uzbek. We conducted a baseline experiment on the qalampir corpus using pre-trained models. The results showed that the pre-trained model CINO outperformed other multilingual models.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N