Linguistic annotation of Byzantine book epigrams: Linguistic annotation of Byzantine book epigrams: C. Swaelens et al.
In this paper, we explore the feasibility of developing a part-of-speech tagger for not-normalised, Byzantine Greek epigrams. Hence, we compared three different transformer-based models with embedding representations, which are then fine-tuned on a fine-grained part-of-speech tagging task. To train...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 1; pp. 109 - 135 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=183750660&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 183750660 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2025 vid: 59 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 183750660 10.1007/s10579-023-09703-x ppf: 109 ppct: 26 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: Linguistic annotation of Byzantine book epigrams: Linguistic annotation of Byzantine book epigrams: C. Swaelens et al. aug: au: Swaelens, Colin De Vos, Ilse Lefever, Els affil: https://ror.org/00cv9y106 LT3, Ghent University, Groot-Brittanniëlaan 45, 9000, Ghent, Belgium VAIA, Flanders AI Academy, Kasteelpark Arenberg 10/2440, 3001, Leuven, Belgium su: Machine learning Language models Natural language processing Marginalia Computational linguistics sug: subj: Machine learning Language models Natural language processing Marginalia Computational linguistics keyword: Byzantine Greek Morphological analysis Neural networks Part-of-speech tagging ab: In this paper, we explore the feasibility of developing a part-of-speech tagger for not-normalised, Byzantine Greek epigrams. Hence, we compared three different transformer-based models with embedding representations, which are then fine-tuned on a fine-grained part-of-speech tagging task. To train the language models, we compiled two data sets: the first consisting of Ancient and Byzantine Greek texts, the second of Ancient, Byzantine and Modern Greek. This allowed us to ascertain whether Modern Greek contributes to the modelling of Byzantine Greek. For the supervised task of part-of-speech tagging, we collected a training set of existing, annotated (Ancient) Greek texts. For evaluation, a gold standard containing 10,000 tokens of unedited Byzantine Greek poems was manually annotated and validated through an inter-annotator agreement study. The experimental results look very promising, with the BERT model trained on all Greek data achieving the best performance for fine-grained part-of-speech tagging. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|