Creating a system for lexical substitutions from scratch using crowdsourcing.

This article describes the creation and application of the Turk Bootstrap Word Sense Inventory for 397 frequent nouns, which is a publicly available resource for lexical substitution. This resource was acquired using Amazon Mechanical Turk. In a bootstrapping process with massive collaborative input...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 47; no. 1; pp. 97 - 123
Main Author: Biemann, Chris
Format: Article
Published: Springer Nature Mar2013
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873226&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 85873226
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2013
      vid: 47
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        85873226
        10.1007/s10579-012-9180-5
      ppf: 97
      ppct: 26
      formats:
        fmt:
          @attributes:
            type: P
            size: 594KB
      tig:
        atl: Creating a system for lexical substitutions from scratch using crowdsourcing.
      aug:
        au: Biemann, Chris
        affil: Technische Universität Darmstadt, Hochschulstraße 10 64289 Darmstadt Germany
      su:
        Lexicology
        Substitution (Linguistics)
        Crowdsourcing
        Word (Linguistics)
        Linguistics
      sug:
        subj:
          Lexicology
          Substitution (Linguistics)
          Crowdsourcing
          Word (Linguistics)
          Linguistics
      keyword:
        Amazon Turk
        Language resource creation
        Lexical substitution
        Word sense disambiguation
      ab: This article describes the creation and application of the Turk Bootstrap Word Sense Inventory for 397 frequent nouns, which is a publicly available resource for lexical substitution. This resource was acquired using Amazon Mechanical Turk. In a bootstrapping process with massive collaborative input, substitutions for target words in context are elicited and clustered by sense; then, more contexts are collected. Contexts that cannot be assigned to a current target word's sense inventory re-enter the bootstrapping loop and get a supply of substitutions. This process yields a sense inventory with its granularity determined by substitutions as opposed to psychologically motivated concepts. It comes with a large number of sense-annotated target word contexts. Evaluation on data quality shows that the process is robust against noise from the crowd, produces a less fine-grained inventory than WordNet and provides a rich body of high precision substitution data at low cost. Using the data to train a system for lexical substitutions, we show that amount and quality of the data is sufficient for producing high quality substitutions automatically. In this system, co-occurrence cluster features are employed as a means to cheaply model topicality.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N