Framework to build and lemmatise an Occitan historical corpus.

This paper presents a framework for building lemmatised Occitan corpora, focusing on early modern texts. Due to strong dialectal and diachronic variation, lemmatisation is essential for enabling cross-text and cross-period comparison. We adopt a semi-automatic approach based on the Pie neural model,...

Descripción completa

Detalles Bibliográficos
Publicado en:Revue Romane Vol. 60; no. 1; pp. 30 - 43
Autor principal: Couffignal, Gilles G.
Formato: Artículo
Publicado: John Benjamins Publishing Co. 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191107366&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 191107366
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00353906
        HYU
      jtl: Revue Romane
      issn: 00353906
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 60
      iid: 1
      pid: 11093
      pub: John Benjamins Publishing Co.
    artinfo:
      ui:
        191107366
        10.1075/rro.25009.cou
      ppf: 30
      ppct: 13
      formats:
      tig:
        atl: Framework to build and lemmatise an Occitan historical corpus.
      aug:
        au: Couffignal, Gilles G.
        affil: Sorbonne Université, STIH
      su:
        Corpora
        Historical source material
        Deep learning
        Language & languages
        Variation in language
        Text processing (Computer science)
      sug:
        subj:
          Corpora
          Historical source material
          Deep learning
          Language & languages
          Variation in language
          Text processing (Computer science)
      keyword:
        corpus linguistics
        historical linguistics
        lemmatisation
        Occitan
        POS tagging
      ab: This paper presents a framework for building lemmatised Occitan corpora, focusing on early modern texts. Due to strong dialectal and diachronic variation, lemmatisation is essential for enabling cross-text and cross-period comparison. We adopt a semi-automatic approach based on the Pie neural model, combining tokenisation, super-lemma selection, and POS tagging aligned with Universal Dependencies. Initial experiments on 17th–18th century texts show promising results, particularly for frequent and grammatical words, while highlighting challenges with unknown lemmas. Despite its exploratory scope, the study demonstrates the feasibility of cost-effective corpus construction and lays the groundwork for a larger, more representative language model of Occitan.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N