Framework to build and lemmatise an Occitan historical corpus.
This paper presents a framework for building lemmatised Occitan corpora, focusing on early modern texts. Due to strong dialectal and diachronic variation, lemmatisation is essential for enabling cross-text and cross-period comparison. We adopt a semi-automatic approach based on the Pie neural model,...
| Publicado en: | Revue Romane Vol. 60; no. 1; pp. 30 - 43 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
John Benjamins Publishing Co.
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191107366&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 191107366 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00353906 HYU jtl: Revue Romane issn: 00353906 maglogo: N pubinfo: dt: 2025 vid: 60 iid: 1 pid: 11093 pub: John Benjamins Publishing Co. artinfo: ui: 191107366 10.1075/rro.25009.cou ppf: 30 ppct: 13 formats: tig: atl: Framework to build and lemmatise an Occitan historical corpus. aug: au: Couffignal, Gilles G. affil: Sorbonne Université, STIH su: Corpora Historical source material Deep learning Language & languages Variation in language Text processing (Computer science) sug: subj: Corpora Historical source material Deep learning Language & languages Variation in language Text processing (Computer science) keyword: corpus linguistics historical linguistics lemmatisation Occitan POS tagging ab: This paper presents a framework for building lemmatised Occitan corpora, focusing on early modern texts. Due to strong dialectal and diachronic variation, lemmatisation is essential for enabling cross-text and cross-period comparison. We adopt a semi-automatic approach based on the Pie neural model, combining tokenisation, super-lemma selection, and POS tagging aligned with Universal Dependencies. Initial experiments on 17th–18th century texts show promising results, particularly for frequent and grammatical words, while highlighting challenges with unknown lemmas. Despite its exploratory scope, the study demonstrates the feasibility of cost-effective corpus construction and lays the groundwork for a larger, more representative language model of Occitan. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|