Uzbek news corpus for named entity recognition.
We have presented a corpus of Uzbek news articles containing manually annotated named entities. The corpus comprises 500 articles (222,536 tokens) and three entity classes (person, location, organization) sourced from Qalampir, an online news source in Uzbekistan. This corpus can be used for develop...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3139 - 3153 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909047&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909047 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909047 10.1007/s10579-024-09786-0 ppf: 3139 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 612KB tig: atl: Uzbek news corpus for named entity recognition. aug: au: Yusufu, Aizihaierjiang Aziz, Kamran Yusufu, Aizierguli Ainiwaer, Abidan Li, Fei Ji, Donghong affil: https://ror.org/033vjfk17 Key Laboratory of Aerospace Information Security and Trusted Computing, Wuhan University, No. 129 Luoyu Road, 430000, Wuhan, Hubei, China https://ror.org/00ndrvk93 School of Computer Science and Technology, Xinjiang Normal University, No. 102 Xinyi Road, 830054, Urumqi, Xinjiang, China https://ror.org/033vjfk17 School of Information Management, Wuhan University, No. 129 Luoyu Road, 430000, Wuhan, Hubei, China su: Natural language processing Corpora Turkic languages Data mining Uzbekistan sug: subj: Uzbekistan Natural language processing Corpora Turkic languages Data mining keyword: Entity naming recognition Newswire Uzbek ab: We have presented a corpus of Uzbek news articles containing manually annotated named entities. The corpus comprises 500 articles (222,536 tokens) and three entity classes (person, location, organization) sourced from Qalampir, an online news source in Uzbekistan. This corpus can be used for develop and evaluate natural language processing (NLP) models for Uzbek. We conducted a baseline experiment on the qalampir corpus using pre-trained models. The results showed that the pre-trained model CINO outperformed other multilingual models. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|