Multiplicity and word sense: evaluating and learning from multiply labeled word sense annotations.
Supervised machine learning methods to model word sense often rely on human labelers to provide a single, ground truth label for each word in its context. We examine issues in establishing ground truth word sense labels using a fine-grained sense inventory from WordNet. Our data consist of a sentenc...
| Published in: | Language Resources & Evaluation Vol. 46; no. 2; pp. 219 - 253 |
|---|---|
| Main Authors: | , , , |
| Format: | Article |
| Published: |
Springer Nature
Jun2012
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=80202958&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 80202958 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2012 vid: 46 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 80202958 10.1007/s10579-012-9188-x ppf: 219 ppct: 34 formats: fmt: @attributes: type: P size: 731KB tig: atl: Multiplicity and word sense: evaluating and learning from multiply labeled word sense annotations. aug: au: Passonneau, Rebecca Bhardwaj, Vikas Salleb-Aouissi, Ansaf Ide, Nancy affil: Columbia University, New York USA Vassar College, Poughkeepsie USA su: Annotations Machine learning Senses Corpora Polysemy Sentences (Grammar) Vocabulary sug: subj: Annotations Machine learning Senses Corpora Polysemy Sentences (Grammar) Vocabulary keyword: Inter-annotator reliability Multilabel learning Word sense annotation ab: Supervised machine learning methods to model word sense often rely on human labelers to provide a single, ground truth label for each word in its context. We examine issues in establishing ground truth word sense labels using a fine-grained sense inventory from WordNet. Our data consist of a sentence corpus of 1,000 sentences: 100 for each of ten moderately polysemous words. Each word was given multiple sense labels-or a multilabel-from trained and untrained annotators. The multilabels give a nuanced representation of the degree of agreement on instances. A suite of assessment metrics is used to analyze the sets of multilabels, such as comparisons of sense distributions across annotators. Our assessment indicates that the general annotation procedure is reliable, but that words differ regarding how reliably annotators can assign WordNet sense labels, independent of the number of senses. We also investigate the performance of an unsupervised machine learning method to infer ground truth labels from various combinations of labels from the trained and untrained annotators. We find tentative support for the hypothesis that performance depends on the quality of the set of multilabels, independent of the number of labelers or their training. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2012. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2012 holdings: @attributes: islocal: N |
|---|