Lemmatization for variation-rich languages using deep learning.

In this article, we describe a novel approach to sequence tagging for languages that are rich in (e.g. orthographic) surface variation. We focus on lemmatization, a basic step in many processing pipelines in the Digital Humanities. While this task has long been considered solved for modern languages...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 32; no. 4; pp. 797 - 816
Autores principales: Kestemont, Mike, de Pauw, Guy, van Nie, Renske, Daelemans, Walter
Formato: Artículo
Publicado: Oxford University Press / USA Dec2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126069653&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 126069653
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Dec2017
      vid: 32
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        126069653
        10.1093/llc/fqw034
      ppf: 797
      ppct: 19
      formats:
        fmt:
          @attributes:
            type: P
            size: 953KB
      tig:
        atl: Lemmatization for variation-rich languages using deep learning.
      aug:
        au:
          Kestemont, Mike
          de Pauw, Guy
          van Nie, Renske
          Daelemans, Walter
        affil: University of Antwerp, Belgium
      su:
        Language & languages
        Digital humanities
        Drama
        Orthography & spelling
        Terms & phrases
      sug:
        subj:
          Language & languages
          Digital humanities
          Drama
          Orthography & spelling
          Terms & phrases
      ab: In this article, we describe a novel approach to sequence tagging for languages that are rich in (e.g. orthographic) surface variation. We focus on lemmatization, a basic step in many processing pipelines in the Digital Humanities. While this task has long been considered solved for modern languages such as English, there exist many (e.g. historic) languages for which the problem is harder to solve, due to a lack of resources and unstable orthography. Our approach is based on recent advances in the field of 'deep' representation learning, where neural networks have led to a dramatic increase in performance across several domains. The proposed system combines two approaches: on the one hand, we apply temporal convolutions to model the orthography of input words at the character level; secondly, we use distributional word embeddings to represent the lexical context surrounding the input words. We demonstrate how this system reaches state-of the- art performance on a number of representative Middle Dutch data sets, even without corpus-specific parameter tuning.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N