Creating a richly annotated corpus of papyrological Greek: The possibilities of natural language processing approaches to a highly inflected historical language.

This article describes a first attempt to annotate the full Greek papyrus corpus automatically for linguistic information. It gives an overview of existing work on Ancient Greek and analyzes the typical problems one encounters when using natural language processing techniques on (1) a historical cor...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 35; no. 1; pp. 67 - 83
Autor principal: Keersmaekers, Alek
Formato: Artículo
Publicado: Oxford University Press / USA Apr2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142636789&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 142636789
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2020
      vid: 35
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        142636789
        10.1093/llc/fqz004
      ppf: 67
      ppct: 16
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 242KB
      tig:
        atl: Creating a richly annotated corpus of papyrological Greek: The possibilities of natural language processing approaches to a highly inflected historical language.
      aug:
        au: Keersmaekers, Alek
        affil: KU Leuven, Belgium
      su:
        Corpora
        Parsing (Computer grammar)
        Greeks
        Natural language processing
      sug:
        subj:
          Corpora
          Parsing (Computer grammar)
          Greeks
          Natural language processing
      ab: This article describes a first attempt to annotate the full Greek papyrus corpus automatically for linguistic information. It gives an overview of existing work on Ancient Greek and analyzes the typical problems one encounters when using natural language processing techniques on (1) a historical corpus of (2) a highly inflectional language (as opposed to the more analytic present-day English) and offers solutions to them, testing several different approaches. The focus is on part-of-speech/morphological tagging and lemmatization; some syntactic parsing experiments are also briefly discussed. The conclusion discusses the strengths and shortcomings of the examined techniques and suggests possible ways to further improve tagging and parsing accuracy.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N