A Finnish news corpus for named entity recognition.
We present a corpus of Finnish news articles with a manually prepared named entity annotation. The corpus consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event, and date). The articles are extracted from the archives of Digitoday...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 1; pp. 247 - 273 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142203872&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 142203872 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2020 vid: 54 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 142203872 10.1007/s10579-019-09471-7 ppf: 247 ppct: 26 formats: fmt: – @attributes: type: T – @attributes: type: P size: 342KB tig: atl: A Finnish news corpus for named entity recognition. aug: au: Ruokolainen, Teemu Kauppinen, Pekka Silfverberg, Miikka Lindén, Krister affil: University of Helsinki, National Library of Finland, Helsinki, Finland Department of Modern Languages, University of Helsinki, Helsinki, Finland su: Corpora Attribution of news Deep learning Instructional systems sug: subj: Corpora Attribution of news Deep learning Instructional systems keyword: Finnish Named entity recognition Newswire Wikipedia ab: We present a corpus of Finnish news articles with a manually prepared named entity annotation. The corpus consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event, and date). The articles are extracted from the archives of Digitoday, a Finnish online technology news source. The corpus is available for research purposes. We present baseline experiments on the corpus using a rule-based and two deep learning systems on two, in-domain and out-of-domain, test sets. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|