Constructing Arabic Reading Comprehension Datasets: Arabic WikiReading and KaifLematha.
Neural machine reading comprehension models have gained immense popularity over the last decade given the availability of large-scale English datasets. A key limiting factor for neural model development and investigations of the Arabic language is the limitation of the currently available datasets....
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 3; pp. 729 - 765 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609440&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 158609440 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2022 vid: 56 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 158609440 10.1007/s10579-022-09577-5 ppf: 729 ppct: 36 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.1MB tig: atl: Constructing Arabic Reading Comprehension Datasets: Arabic WikiReading and KaifLematha. aug: au: Albilali, Eman Al-Twairesh, Nora Hosny, Manar affil: Computer Science Department, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia Information Technology Department, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia STC's Artificial Intelligence Research Chair, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia su: Wikipedia Natural language processing Reading comprehension Neural development Arabic language sug: subj: Wikipedia Natural language processing Reading comprehension Neural development Arabic language keyword: Deep learning Machine reading comprehension Natural language understanding Pre-trained language model Question answering ab: Neural machine reading comprehension models have gained immense popularity over the last decade given the availability of large-scale English datasets. A key limiting factor for neural model development and investigations of the Arabic language is the limitation of the currently available datasets. Current available datasets are either too small to train deep neural models or created by the automatic translation of the available English datasets, where the exact answer may not be found in the corresponding text. In this paper, we propose two high quality and large-scale Arabic reading comprehension datasets: Arabic WikiReading and KaifLematha with around +100 K instances. We followed two different methodologies to construct our datasets. First, we employed crowdworkers to collect non-factoid questions from paragraphs on Wikipedia. Then, we constructed Arabic WikiReading following a distant supervision strategy, utilizing the Wikidata knowledge base as a ground truth. We carried out both quantitative and qualitative analyses to investigate the level of reasoning required to answer the questions in the proposed datasets. We evaluated competitive pre-trained language model that attained F1 scores of 81.77 and 68.61 for the Arabic WikiReading and KaifLematha datasets, respectively, but struggled to extract a precise answer for the KaifLematha dataset. Human performance reported an F1 score of 82.54 for the KaifLematha development set, which leaves ample room for improvement. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|