FullStop: punctuation and segmentation prediction for Dutch with transformers.

When applying automated speech recognition (ASR) for Belgian Dutch, the output consists of an unsegmented stream of words, without any punctuation. A next step is to perform segmentation and insert punctuation, making the ASR output more readable and easy to manually correct. We present the first (a...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 4; pp. 1335 - 1355
Autores principales: Vandeghinste, Vincent, Guhr, Oliver
Formato: Artículo
Publicado: Springer Nature Dec2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=180627303&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 180627303
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2024
      vid: 58
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        180627303
        10.1007/s10579-023-09676-x
      ppf: 1335
      ppct: 20
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2MB
      tig:
        atl: FullStop: punctuation and segmentation prediction for Dutch with transformers.
      aug:
        au:
          Vandeghinste, Vincent
          Guhr, Oliver
        affil:
          https://ror.org/04m5bjk54 Instituut voor de Nederlandse Taal, Rapenburg 61, 2311 GJ, Leiden, The Netherlands
          https://ror.org/05f950310 Centre for Computational Linguistics, Leuven.AI, KU Leuven, Blijde Inkomststraat 21, 3000, Leuven, Belgium
          https://ror.org/01xzwj424 Künstliche Intelligenz / Kognitive Robotik, Hochschule für Technik und Wirtschaft, Friedrich-List-Platz 1, 01069, Dresden, Germany
      su:
        Language models
        Text mining
        Speech perception
        Code switching (Linguistics)
        Dutch language
      sug:
        subj:
          Language models
          Text mining
          Speech perception
          Code switching (Linguistics)
          Dutch language
      keyword:
        Dutch
        Punctuation
        Speech recognition
        Transformers
      ab: When applying automated speech recognition (ASR) for Belgian Dutch, the output consists of an unsegmented stream of words, without any punctuation. A next step is to perform segmentation and insert punctuation, making the ASR output more readable and easy to manually correct. We present the first (as far as we know) publicly available punctuation insertion system for Dutch that functions at a usable level and that is publicly available. The model we present here is an extension of the approach of Guhr et al. (In: Swiss Text Analytics Conference. Shared task on Sentence End and Punctuation Prediction in NLG Text, 2021) for Dutch: we finetuned the Dutch language model RobBERT on a punctuation prediction sequence classification task. The model was finetuned on two datasets: the Dutch side of Europarl and the SoNaR corpus. For every word in the input sequence, the model predicts a punctuation marker that follows the word. In cases where the language is unknown or where code switching applies, we have extended an existing multilingual model with Dutch. Previous work showed that such a multilingual model, based on "xlm-roberta-base" performs on par or sometimes even better than the monolingual cases. The system was evaluated on in-domain data as a classifier and on out-of-domain data as a sentence segmentation system through full stop prediction. The evaluations on sentence segmentation on out of domain data show that models finetuned on SoNaR show the best results, which can be attributed to SoNaR being a reference corpus containing different language registers. The multilingual models show an even better precision (at the cost of a lower recall) compared to the monolingual models.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N