From 0 to 10 million annotated words: part-of-speech tagging for Middle High German.

By building a part-of-speech (POS) tagger for Middle High German, we investigate strategies for dealing with a low resource, diverse and non-standard language in the domain of natural language processing. We highlight various aspects such as the data quantity needed for training and the influence of...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 53; no. 4; pp. 837 - 864
Autores principales: Schulz, Sarah, Ketschik, Nora
Formato: Artículo
Publicado: Springer Nature Dec2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=139882013&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 139882013
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2019
      vid: 53
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        139882013
        10.1007/s10579-019-09462-8
      ppf: 837
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 551KB
      tig:
        atl: From 0 to 10 million annotated words: part-of-speech tagging for Middle High German.
      aug:
        au:
          Schulz, Sarah
          Ketschik, Nora
        affil:
          Institute for Natural Language Processing (IMS), University of Stuttgart, Pfaffenwaldring 5B, 70569, Stuttgart, Germany
          Institute for Literary Studies (ILW), University of Stuttgart, Keplerstraße 17, 70174, Stuttgart, Germany
      su:
        Natural language processing
        Data quality
        Training needs
      sug:
        subj:
          Natural language processing
          Data quality
          Training needs
      keyword:
        Digital Humanities
        Historical language
        Middle High German
        Non-standard text processing
        Part-of-speech tagging
      ab: By building a part-of-speech (POS) tagger for Middle High German, we investigate strategies for dealing with a low resource, diverse and non-standard language in the domain of natural language processing. We highlight various aspects such as the data quantity needed for training and the influence of data quality on tagger performance. Since the lack of annotated resources poses a problem for training a tagger, we exemplify how existing resources can be adapted fruitfully to serve as additional training data. The resulting POS model achieves a tagging accuracy of about 91% on a diverse test set representing the different genres, time periods and varieties of MHG. In order to verify its general applicability, we evaluate the performance on different genres, authors and varieties of MHG, separately. We explore self-learning techniques which yield the advantage that unannotated data can be utilized to improve tagging performance on specific subcorpora.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N