From 0 to 10 million annotated words: part-of-speech tagging for Middle High German.
By building a part-of-speech (POS) tagger for Middle High German, we investigate strategies for dealing with a low resource, diverse and non-standard language in the domain of natural language processing. We highlight various aspects such as the data quantity needed for training and the influence of...
| Publicado en: | Language Resources & Evaluation Vol. 53; no. 4; pp. 837 - 864 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=139882013&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 139882013 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2019 vid: 53 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 139882013 10.1007/s10579-019-09462-8 ppf: 837 ppct: 27 formats: fmt: – @attributes: type: T – @attributes: type: P size: 551KB tig: atl: From 0 to 10 million annotated words: part-of-speech tagging for Middle High German. aug: au: Schulz, Sarah Ketschik, Nora affil: Institute for Natural Language Processing (IMS), University of Stuttgart, Pfaffenwaldring 5B, 70569, Stuttgart, Germany Institute for Literary Studies (ILW), University of Stuttgart, Keplerstraße 17, 70174, Stuttgart, Germany su: Natural language processing Data quality Training needs sug: subj: Natural language processing Data quality Training needs keyword: Digital Humanities Historical language Middle High German Non-standard text processing Part-of-speech tagging ab: By building a part-of-speech (POS) tagger for Middle High German, we investigate strategies for dealing with a low resource, diverse and non-standard language in the domain of natural language processing. We highlight various aspects such as the data quantity needed for training and the influence of data quality on tagger performance. Since the lack of annotated resources poses a problem for training a tagger, we exemplify how existing resources can be adapted fruitfully to serve as additional training data. The resulting POS model achieves a tagging accuracy of about 91% on a diverse test set representing the different genres, time periods and varieties of MHG. In order to verify its general applicability, we evaluate the performance on different genres, authors and varieties of MHG, separately. We explore self-learning techniques which yield the advantage that unannotated data can be utilized to improve tagging performance on specific subcorpora. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|