Creation and evaluation of large keyphrase extraction collections with multiple opinions.
While several automatic keyphrase extraction (AKE) techniques have been developed and analyzed, there is little consensus on the definition of the task and a lack of overview of the effectiveness of different techniques. Proper evaluation of keyphrase extraction requires large test collections with...
| Publicado en: | Language Resources & Evaluation Vol. 52; no. 2; pp. 503 - 533 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=129593481&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 129593481 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2018 vid: 52 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 129593481 10.1007/s10579-017-9395-6 ppf: 503 ppct: 30 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.4MB tig: atl: Creation and evaluation of large keyphrase extraction collections with multiple opinions. aug: au: Sterckx, Lucas Demeester, Thomas Deleu, Johannes Develder, Chris affil: Department of Information Technology, Ghent University - imec, Technologiepark Zwijnaarde 15, 9052, Ghent, Belgium su: Machine learning Information resources management Data mining Support vector machines Data management sug: subj: Machine learning Information resources management Data mining Support vector machines Data management keyword: Annotator disagreement Automatic keyphrase extraction Test collections ab: While several automatic keyphrase extraction (AKE) techniques have been developed and analyzed, there is little consensus on the definition of the task and a lack of overview of the effectiveness of different techniques. Proper evaluation of keyphrase extraction requires large test collections with multiple opinions, currently not available for research. In this paper, we (i) present a set of test collections derived from various sources with multiple annotations (which we also refer to as <italic>opinions</italic> in the remained of the paper) for each document, (ii) systematically evaluate keyphrase extraction using several supervised and unsupervised AKE techniques, (iii) and experimentally analyze the effects of disagreement on AKE evaluation. Our newly created set of test collections spans different types of topical content from general news and magazines, and is annotated with multiple annotations per article by a large annotator panel. Our annotator study shows that for a given document there seems to be a large disagreement on the preferred keyphrases, suggesting the need for multiple opinions per document. A first systematic evaluation of ranking and classification of keyphrases using both unsupervised and supervised AKE techniques on the test collections shows a superior effectiveness of supervised models, even for a low annotation effort and with basic positional and frequency features, and highlights the importance of a suitable keyphrase candidate generation approach. We also study the influence of multiple opinions, training data and document length on evaluation of keyphrase extraction. Our new test collection for keyphrase extraction is one of the largest of its kind and will be made available to stimulate future work to improve reliable evaluation of new keyphrase extractors. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|