Creating a system for lexical substitutions from scratch using crowdsourcing.
This article describes the creation and application of the Turk Bootstrap Word Sense Inventory for 397 frequent nouns, which is a publicly available resource for lexical substitution. This resource was acquired using Amazon Mechanical Turk. In a bootstrapping process with massive collaborative input...
| Published in: | Language Resources & Evaluation Vol. 47; no. 1; pp. 97 - 123 |
|---|---|
| Main Author: | |
| Format: | Article |
| Published: |
Springer Nature
Mar2013
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=85873226&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 85873226 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2013 vid: 47 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 85873226 10.1007/s10579-012-9180-5 ppf: 97 ppct: 26 formats: fmt: @attributes: type: P size: 594KB tig: atl: Creating a system for lexical substitutions from scratch using crowdsourcing. aug: au: Biemann, Chris affil: Technische Universität Darmstadt, Hochschulstraße 10 64289 Darmstadt Germany su: Lexicology Substitution (Linguistics) Crowdsourcing Word (Linguistics) Linguistics sug: subj: Lexicology Substitution (Linguistics) Crowdsourcing Word (Linguistics) Linguistics keyword: Amazon Turk Language resource creation Lexical substitution Word sense disambiguation ab: This article describes the creation and application of the Turk Bootstrap Word Sense Inventory for 397 frequent nouns, which is a publicly available resource for lexical substitution. This resource was acquired using Amazon Mechanical Turk. In a bootstrapping process with massive collaborative input, substitutions for target words in context are elicited and clustered by sense; then, more contexts are collected. Contexts that cannot be assigned to a current target word's sense inventory re-enter the bootstrapping loop and get a supply of substitutions. This process yields a sense inventory with its granularity determined by substitutions as opposed to psychologically motivated concepts. It comes with a large number of sense-annotated target word contexts. Evaluation on data quality shows that the process is robust against noise from the crowd, produces a less fine-grained inventory than WordNet and provides a rich body of high precision substitution data at low cost. Using the data to train a system for lexical substitutions, we show that amount and quality of the data is sufficient for producing high quality substitutions automatically. In this system, co-occurrence cluster features are employed as a means to cheaply model topicality. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2013 holdings: @attributes: islocal: N |
|---|