Linguistic annotation of Byzantine book epigrams: Linguistic annotation of Byzantine book epigrams: C. Swaelens et al.

In this paper, we explore the feasibility of developing a part-of-speech tagger for not-normalised, Byzantine Greek epigrams. Hence, we compared three different transformer-based models with embedding representations, which are then fine-tuned on a fine-grained part-of-speech tagging task. To train...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 1; pp. 109 - 135
Autores principales: Swaelens, Colin, De Vos, Ilse, Lefever, Els
Formato: Artículo
Publicado: Springer Nature Mar2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=183750660&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 183750660
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2025
      vid: 59
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        183750660
        10.1007/s10579-023-09703-x
      ppf: 109
      ppct: 26
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.3MB
      tig:
        atl: Linguistic annotation of Byzantine book epigrams: Linguistic annotation of Byzantine book epigrams: C. Swaelens et al.
      aug:
        au:
          Swaelens, Colin
          De Vos, Ilse
          Lefever, Els
        affil:
          https://ror.org/00cv9y106 LT3, Ghent University, Groot-Brittanniëlaan 45, 9000, Ghent, Belgium
          VAIA, Flanders AI Academy, Kasteelpark Arenberg 10/2440, 3001, Leuven, Belgium
      su:
        Machine learning
        Language models
        Natural language processing
        Marginalia
        Computational linguistics
      sug:
        subj:
          Machine learning
          Language models
          Natural language processing
          Marginalia
          Computational linguistics
      keyword:
        Byzantine Greek
        Morphological analysis
        Neural networks
        Part-of-speech tagging
      ab: In this paper, we explore the feasibility of developing a part-of-speech tagger for not-normalised, Byzantine Greek epigrams. Hence, we compared three different transformer-based models with embedding representations, which are then fine-tuned on a fine-grained part-of-speech tagging task. To train the language models, we compiled two data sets: the first consisting of Ancient and Byzantine Greek texts, the second of Ancient, Byzantine and Modern Greek. This allowed us to ascertain whether Modern Greek contributes to the modelling of Byzantine Greek. For the supervised task of part-of-speech tagging, we collected a training set of existing, annotated (Ancient) Greek texts. For evaluation, a gold standard containing 10,000 tokens of unedited Byzantine Greek poems was manually annotated and validated through an inter-annotator agreement study. The experimental results look very promising, with the BERT model trained on all Greek data achieving the best performance for fine-grained part-of-speech tagging.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N