SENTiVENT: enabling supervised information extraction of company-specific events in economic and financial news.
We present SENTiVENT, a corpus of fine-grained company-specific events in English economic news articles. The domain of event processing is highly productive and various general domain, fine-grained event extraction corpora are freely available but economically-focused resources are lacking. This wo...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 1; pp. 225 - 258 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=155080021&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 155080021 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2022 vid: 56 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 155080021 10.1007/s10579-021-09562-4 ppf: 225 ppct: 33 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.4MB tig: atl: SENTiVENT: enabling supervised information extraction of company-specific events in economic and financial news. aug: au: Jacobs, Gilles Hoste, Véronique affil: Language & Translation Technology Team, Ghent University, UGent VTC Mercator A, Abdisstraat 1, 9000, Ghent, Belgium su: Data mining Modal logic Supervised learning Text mining Machine learning Source code sug: subj: Data mining Modal logic Supervised learning Text mining Machine learning Source code keyword: Annotation scheme Economic events English corpus Event detection Event extraction Financial information extraction ab: We present SENTiVENT, a corpus of fine-grained company-specific events in English economic news articles. The domain of event processing is highly productive and various general domain, fine-grained event extraction corpora are freely available but economically-focused resources are lacking. This work fills a large need for a manually annotated dataset for economic and financial text mining applications. A representative corpus of business news is crawled and an annotation scheme developed with an iteratively refined economic event typology. The annotations are compatible with benchmark datasets (ACE/ERE) so state-of-the-art event extraction systems can be readily applied. This results in a gold-standard dataset annotated with event triggers, participant arguments, event co-reference, and event attributes such as type, subtype, negation, and modality. An adjudicated reference test set is created for use in annotator and system evaluation. Agreement scores are substantial and annotator performance adequate, indicating that the annotation scheme produces consistent event annotations of high quality. In an event detection pilot study, satisfactory results were obtained with a macro-averaged F 1 -score of 59 % validating the dataset for machine learning purposes. This dataset thus provides a rich resource on events as training data for supervised machine learning for economic and financial applications. The dataset and related source code is made available at https://osf.io/8jec2/. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|