Constructing Arabic Reading Comprehension Datasets: Arabic WikiReading and KaifLematha.

Neural machine reading comprehension models have gained immense popularity over the last decade given the availability of large-scale English datasets. A key limiting factor for neural model development and investigations of the Arabic language is the limitation of the currently available datasets....

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 56; no. 3; pp. 729 - 765
Autores principales: Albilali, Eman, Al-Twairesh, Nora, Hosny, Manar
Formato: Artículo
Publicado: Springer Nature Sep2022
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609440&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 158609440
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2022
      vid: 56
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        158609440
        10.1007/s10579-022-09577-5
      ppf: 729
      ppct: 36
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.1MB
      tig:
        atl: Constructing Arabic Reading Comprehension Datasets: Arabic WikiReading and KaifLematha.
      aug:
        au:
          Albilali, Eman
          Al-Twairesh, Nora
          Hosny, Manar
        affil:
          Computer Science Department, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia
          Information Technology Department, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia
          STC's Artificial Intelligence Research Chair, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia
      su:
        Wikipedia
        Natural language processing
        Reading comprehension
        Neural development
        Arabic language
      sug:
        subj:
          Wikipedia
          Natural language processing
          Reading comprehension
          Neural development
          Arabic language
      keyword:
        Deep learning
        Machine reading comprehension
        Natural language understanding
        Pre-trained language model
        Question answering
      ab: Neural machine reading comprehension models have gained immense popularity over the last decade given the availability of large-scale English datasets. A key limiting factor for neural model development and investigations of the Arabic language is the limitation of the currently available datasets. Current available datasets are either too small to train deep neural models or created by the automatic translation of the available English datasets, where the exact answer may not be found in the corresponding text. In this paper, we propose two high quality and large-scale Arabic reading comprehension datasets: Arabic WikiReading and KaifLematha with around +100 K instances. We followed two different methodologies to construct our datasets. First, we employed crowdworkers to collect non-factoid questions from paragraphs on Wikipedia. Then, we constructed Arabic WikiReading following a distant supervision strategy, utilizing the Wikidata knowledge base as a ground truth. We carried out both quantitative and qualitative analyses to investigate the level of reasoning required to answer the questions in the proposed datasets. We evaluated competitive pre-trained language model that attained F1 scores of 81.77 and 68.61 for the Arabic WikiReading and KaifLematha datasets, respectively, but struggled to extract a precise answer for the KaifLematha dataset. Human performance reported an F1 score of 82.54 for the KaifLematha development set, which leaves ample room for improvement.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N