Lemmatization for variation-rich languages using deep learning.
In this article, we describe a novel approach to sequence tagging for languages that are rich in (e.g. orthographic) surface variation. We focus on lemmatization, a basic step in many processing pipelines in the Digital Humanities. While this task has long been considered solved for modern languages...
| Publicado en: | Digital Scholarship in the Humanities Vol. 32; no. 4; pp. 797 - 816 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Dec2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126069653&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 126069653 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Dec2017 vid: 32 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 126069653 10.1093/llc/fqw034 ppf: 797 ppct: 19 formats: fmt: @attributes: type: P size: 953KB tig: atl: Lemmatization for variation-rich languages using deep learning. aug: au: Kestemont, Mike de Pauw, Guy van Nie, Renske Daelemans, Walter affil: University of Antwerp, Belgium su: Language & languages Digital humanities Drama Orthography & spelling Terms & phrases sug: subj: Language & languages Digital humanities Drama Orthography & spelling Terms & phrases ab: In this article, we describe a novel approach to sequence tagging for languages that are rich in (e.g. orthographic) surface variation. We focus on lemmatization, a basic step in many processing pipelines in the Digital Humanities. While this task has long been considered solved for modern languages such as English, there exist many (e.g. historic) languages for which the problem is harder to solve, due to a lack of resources and unstable orthography. Our approach is based on recent advances in the field of 'deep' representation learning, where neural networks have led to a dramatic increase in performance across several domains. The proposed system combines two approaches: on the one hand, we apply temporal convolutions to model the orthography of input words at the character level; secondly, we use distributional word embeddings to represent the lexical context surrounding the input words. We demonstrate how this system reaches state-of the- art performance on a number of representative Middle Dutch data sets, even without corpus-specific parameter tuning. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|