Crowdsourcing for web genre annotation.

Recently, genre collection and automatic genre identification for the web has attracted much attention. However, currently there is no genre-annotated corpus of web pages where inter-annotator reliability has been established, i.e. the corpora are either not tested for inter-annotator reliability or...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 50; no. 3; pp. 603 - 642
Main Authors: Asheghi, Noushin, Sharoff, Serge, Markert, Katja
Format: Article
Published: Springer Nature Sep2016
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=117418391&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 117418391
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2016
      vid: 50
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        117418391
        10.1007/s10579-015-9331-6
      ppf: 603
      ppct: 39
      formats:
        fmt:
          @attributes:
            type: P
            size: 1008KB
      tig:
        atl: Crowdsourcing for web genre annotation.
      aug:
        au:
          Asheghi, Noushin
          Sharoff, Serge
          Markert, Katja
        affil:
          School of Computing , University of Leeds , Leeds LS2 9JT UK
          School of Modern Languages and Cultures , University of Leeds , Leeds LS2 9JT UK
      su:
        Crowdsourcing
        Annotations
        Corpora
        Websites
        Cost effectiveness
      sug:
        subj:
          Crowdsourcing
          Annotations
          Corpora
          Websites
          Cost effectiveness
      keyword:
        Annotation guidelines
        Genres on the web
        Reliability testing
      ab: Recently, genre collection and automatic genre identification for the web has attracted much attention. However, currently there is no genre-annotated corpus of web pages where inter-annotator reliability has been established, i.e. the corpora are either not tested for inter-annotator reliability or exhibit low inter-coder agreement. Annotation has also mostly been carried out by a small number of experts, leading to concerns with regard to scalability of these annotation efforts and transferability of the schemes to annotators outside these small expert groups. In this paper, we tackle these problems by using crowd-sourcing for genre annotation, leading to the Leeds Web Genre Corpus-the first web corpus which is, demonstrably reliably annotated for genre and which can be easily and cost-effectively expanded using naive annotators. We also show that the corpus is source and topic diverse.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2016. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N