Fine-grained Dutch named entity recognition.
This paper describes the creation of a fine-grained named entity annotation scheme and corpus for Dutch, and experiments on automatic main type and subtype named entity recognition. We give an overview of existing named entity annotation schemes, and motivate our own, which describes six main types...
| Publicado en: | Language Resources & Evaluation Vol. 48; no. 2; pp. 307 - 344 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2014
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=95865798&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 95865798 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2014 vid: 48 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 95865798 10.1007/s10579-013-9255-y ppf: 307 ppct: 37 formats: fmt: @attributes: type: P size: 509KB tig: atl: Fine-grained Dutch named entity recognition. aug: au: Desmet, Bart Hoste, Véronique affil: LT3, Language and Translation Technology Team, Ghent University, Groot-Brittanniëlaan 45 9000 Ghent Belgium su: Corpora Dutch language Annotations Information theory Word recognition sug: subj: Corpora Dutch language Annotations Information theory Word recognition keyword: Annotation Classifier ensembles Named entity recognition Subtype classification ab: This paper describes the creation of a fine-grained named entity annotation scheme and corpus for Dutch, and experiments on automatic main type and subtype named entity recognition. We give an overview of existing named entity annotation schemes, and motivate our own, which describes six main types (persons, organizations, locations, products, events and miscellaneous named entities) and finer-grained information on subtypes and metonymic usage. This was applied to a one-million-word subset of the Dutch SoNaR reference corpus. The classifier for main type named entities achieves a micro-averaged F-score of 84.91 %, and is publicly available, along with the corpus and annotations. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2014. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2014 holdings: @attributes: islocal: N |
|---|