Improving English verb sense disambiguation performance with linguistically motivated features and clear sense distinction boundaries.
This paper presents a high-performance broad-coverage supervised word sense disambiguation (WSD) system for English verbs that uses linguistically motivated features and a smoothed maximum entropy machine learning model. We describe three specific enhancements to our system’s treatment of linguistic...
| Publicado en: | Language Resources & Evaluation Vol. 43; no. 2; pp. 181 - 209 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2009
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=39879143&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 39879143 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2009 vid: 43 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 39879143 10.1007/s10579-009-9085-0 ppf: 181 ppct: 28 formats: fmt: @attributes: type: P size: 580KB tig: atl: Improving English verb sense disambiguation performance with linguistically motivated features and clear sense distinction boundaries. aug: au: Jinying Chen Palmer, Martha affil: BBN Technologies, Cambridge USA. University of Colorado, Boulder USA. su: Verbs English language Linguistics Semantics Language & languages sug: subj: Verbs English language Linguistics Semantics Language & languages keyword: Linear regression Linguistically motivated features Maximum entropy Sense granularity Word sense disambiguation ab: This paper presents a high-performance broad-coverage supervised word sense disambiguation (WSD) system for English verbs that uses linguistically motivated features and a smoothed maximum entropy machine learning model. We describe three specific enhancements to our system’s treatment of linguistically motivated features which resulted in the best published results on SENSEVAL-2 verbs. We then present the results of training our system on OntoNotes data, both the SemEval-2007 task and additional data. OntoNotes data is designed to provide clear sense distinctions, based on using explicit syntactic and semantic criteria to group WordNet senses, with sufficient examples to constitute high quality, broad coverage training data. Using similar syntactic and semantic features for WSD, we achieve performance comparable to that of human taggers, and competitive with the top results for the SemEval-2007 task. Empirical analysis of our results suggests that clarifying sense boundaries and/or increasing the number of training instances for certain verbs could further improve system performance. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2009. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2009 holdings: @attributes: islocal: N |
|---|