Cross-language transfer of semantic annotation via targeted crowdsourcing: task design and evaluation.

Modern data-driven spoken language systems (SLS) require manual semantic annotation for training spoken language understanding parsers. Multilingual porting of SLS demands significant manual effort and language resources, as this manual annotation has to be replicated. Crowdsourcing is an accessible...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 52; no. 1; pp. 341 - 365
Main Authors: Stepanov, Evgeny A., Chowdhury, Shammur Absar, Bayer, Ali Orkan, Ghosh, Arindam, Klasinas, Ioannis, Calvo, Marcos, Sanchis, Emilio, Riccardi, Giuseppe
Format: Article
Published: Springer Nature Mar2018
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=127930801&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 127930801
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2018
      vid: 52
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        127930801
        10.1007/s10579-017-9396-5
      ppf: 341
      ppct: 24
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 749KB
      tig:
        atl: Cross-language transfer of semantic annotation via targeted crowdsourcing: task design and evaluation.
      aug:
        au:
          Stepanov, Evgeny A.
          Chowdhury, Shammur Absar
          Bayer, Ali Orkan
          Ghosh, Arindam
          Klasinas, Ioannis
          Calvo, Marcos
          Sanchis, Emilio
          Riccardi, Giuseppe
        affil:
          Signals and Interactive Systems Lab, Department of Information Engineering and Computer Science, University of Trento, via Sommarive, 5, Trento, Italy
          Department of Electronics and Computer Engineering, Technical University of Crete, 731 00, Chania, Greece
          Google Switzerland, Brandschenkestrasse 110, 8002, Zurich, Switzerland
          Departamento de Sistemas Informáticos y Computación, Universitat Politècnica de València, Camino de Vera s/n, 46020, Valencia, Spain
      su:
        Semantics
        Multilingual computing
        Crowdsourcing
        Annotations
        English language
      sug:
        subj:
          Semantics
          Multilingual computing
          Crowdsourcing
          Annotations
          English language
      keyword:
        Cross-language transfer
        Evaluation
        Semantic annotation
      ab: Modern data-driven spoken language systems (SLS) require manual semantic annotation for training spoken language understanding parsers. Multilingual porting of SLS demands significant manual effort and language resources, as this manual annotation has to be replicated. Crowdsourcing is an accessible and cost-effective alternative to traditional methods of collecting and annotating data. The application of crowdsourcing to simple tasks has been well investigated. However, complex tasks, like cross-language semantic annotation transfer, may generate low judgment agreement and/or poor performance. The most serious issue in cross-language porting is the absence of reference annotations in the target language; thus, crowd quality control and the evaluation of the collected annotations is difficult. In this paper we investigate <italic>targeted</italic> crowdsourcing for semantic annotation transfer that delegates to crowds a complex task such as segmenting and labeling of concepts taken from a domain ontology; and evaluation using source language annotation. To test the applicability and effectiveness of the crowdsourced annotation transfer we have considered the case of close and distant language pairs: Italian–Spanish and Italian–Greek. The corpora annotated via crowdsourcing are evaluated against source and target language expert annotations. We demonstrate that the two evaluation references (source and target) highly correlate with each other; thus, drastically reduce the need for the target language reference annotations.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N