Construction of Amharic information retrieval resources and corpora.
The development of information retrieval systems and natural language processing tools has been made possible for many natural languages because of the availability of natural language resources and corpora. Although Amharic is the working language of Ethiopia, it is still an under-resourced languag...
| Published in: | Language Resources & Evaluation Vol. 58; no. 4; pp. 1157 - 1186 |
|---|---|
| Main Authors: | , , |
| Format: | Article |
| Published: |
Springer Nature
Dec2024
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=180627307&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 180627307 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2024 vid: 58 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 180627307 10.1007/s10579-024-09719-x ppf: 1157 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 4.9MB tig: atl: Construction of Amharic information retrieval resources and corpora. aug: au: Yeshambel, Tilahun Mothe, Josiane Assabie, Yaregal affil: https://ror.org/038b8e254 IT PhD Program, Addis Ababa University, Addis Ababa, Ethiopia INSPE, Univ. de Toulouse, IRIT, UMR5505, CNRS, Toulouse, France https://ror.org/038b8e254 Department of Computer Science, Addis Ababa University, Addis Ababa, Ethiopia su: Natural language processing Information storage & retrieval systems Natural languages Information resources Corpora Information retrieval sug: subj: Natural language processing Information storage & retrieval systems Natural languages Information resources Corpora Information retrieval keyword: Amharic language Evaluation Resources ab: The development of information retrieval systems and natural language processing tools has been made possible for many natural languages because of the availability of natural language resources and corpora. Although Amharic is the working language of Ethiopia, it is still an under-resourced language. There are no adequate resources and corpora for Amharic ad-hoc retrieval evaluation to date. The existing ones are not publicly accessible and are not suitable for making scientific evaluation of information retrieval systems. To promote the development of Amharic ad-hoc retrieval, we build an ad-hoc retrieval test collection that consists of raw text, morphologically annotated stem-based and root-based corpora, a stopword list, stem-based and root-based lexicons, and WordNet-like resources. We also created word embeddings using the raw text and morphologically segmented forms of the corpora. When building these resources and corpora, we heavily consider the morphological characteristics of the language. The aim of this paper is to present these Amharic resources and corpora that we made available to the research community for information retrieval tasks. These resources and corpora are also evaluated experimentally and by linguists. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|