Crowdsourcing for web genre annotation.

Recently, genre collection and automatic genre identification for the web has attracted much attention. However, currently there is no genre-annotated corpus of web pages where inter-annotator reliability has been established, i.e. the corpora are either not tested for inter-annotator reliability or...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 50; no. 3; pp. 603 - 642
Autores principales: Asheghi, Noushin, Sharoff, Serge, Markert, Katja
Formato: Artículo
Publicado: Springer Nature Sep2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Recently, genre collection and automatic genre identification for the web has attracted much attention. However, currently there is no genre-annotated corpus of web pages where inter-annotator reliability has been established, i.e. the corpora are either not tested for inter-annotator reliability or exhibit low inter-coder agreement. Annotation has also mostly been carried out by a small number of experts, leading to concerns with regard to scalability of these annotation efforts and transferability of the schemes to annotators outside these small expert groups. In this paper, we tackle these problems by using crowd-sourcing for genre annotation, leading to the Leeds Web Genre Corpus-the first web corpus which is, demonstrably reliably annotated for genre and which can be easily and cost-effectively expanded using naive annotators. We also show that the corpus is source and topic diverse.