Supervised collaboration for syntactic annotation of Quranic Arabic.

The Quranic Arabic Corpus () is a collaboratively constructed linguistic resource initiated at the University of Leeds, with multiple layers of annotation including part-of-speech tagging, morphological segmentation (Dukes and Habash ) and syntactic analysis using dependency grammar (Dukes and Buckw...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 47; no. 1; pp. 33 - 63
Autores principales: Dukes, Kais, Atwell, Eric, Habash, Nizar
Formato: Artículo
Publicado: Springer Nature Mar2013
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873228&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 85873228
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2013
      vid: 47
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        85873228
        10.1007/s10579-011-9167-7
      ppf: 33
      ppct: 30
      formats:
        fmt:
          @attributes:
            type: P
            size: 898KB
      tig:
        atl: Supervised collaboration for syntactic annotation of Quranic Arabic.
      aug:
        au:
          Dukes, Kais
          Atwell, Eric
          Habash, Nizar
        affil:
          University of Leeds, Leeds UK
          Columbia University, New York USA
      su:
        Arabic language
        Qur'an
        Linguistics
        Dependency grammar
        Islam
      sug:
        subj:
          Arabic language
          Qur'an
          Linguistics
          Dependency grammar
          Islam
      keyword:
        Arabic
        Collaborative annotation
        Corpus
        Quran
        Treebank
      ab: The Quranic Arabic Corpus () is a collaboratively constructed linguistic resource initiated at the University of Leeds, with multiple layers of annotation including part-of-speech tagging, morphological segmentation (Dukes and Habash ) and syntactic analysis using dependency grammar (Dukes and Buckwalter ). The motivation behind this work is to produce a resource that enables further analysis of the Quran, the 1,400 year-old central religious text of Islam. This project contrasts with other Arabic treebanks by providing a deep linguistic model based on the historical traditional grammar known as i′rāb (إعراب). By adapting this well-known canon of Quranic grammar into a familiar tagset, it is possible to encourage online annotation by Arabic linguists and Quranic experts. This article presents a new approach to linguistic annotation of an Arabic corpus: online supervised collaboration using a multi-stage approach. The different stages include automatic rule-based tagging, initial manual verification, and online supervised collaborative proofreading. A popular website attracting thousands of visitors per day, the Quranic Arabic Corpus has approximately 100 unpaid volunteer annotators each suggesting corrections to existing linguistic tagging. To ensure a high-quality resource, a small number of expert annotators are promoted to a supervisory role, allowing them to review or veto suggestions made by other collaborators. The Quran also benefits from a large body of existing historical grammatical analysis, which may be leveraged during this review. In this paper we evaluate and report on the effectiveness of the chosen annotation methodology. We also discuss the unique challenges of annotating Quranic Arabic online and describe the custom linguistic software used to aid collaborative annotation.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N