FullStop: punctuation and segmentation prediction for Dutch with transformers.
When applying automated speech recognition (ASR) for Belgian Dutch, the output consists of an unsegmented stream of words, without any punctuation. A next step is to perform segmentation and insert punctuation, making the ASR output more readable and easy to manually correct. We present the first (a...
| Publicado en: | Language Resources & Evaluation Vol. 58; no. 4; pp. 1335 - 1355 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=180627303&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 180627303 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2024 vid: 58 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 180627303 10.1007/s10579-023-09676-x ppf: 1335 ppct: 20 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2MB tig: atl: FullStop: punctuation and segmentation prediction for Dutch with transformers. aug: au: Vandeghinste, Vincent Guhr, Oliver affil: https://ror.org/04m5bjk54 Instituut voor de Nederlandse Taal, Rapenburg 61, 2311 GJ, Leiden, The Netherlands https://ror.org/05f950310 Centre for Computational Linguistics, Leuven.AI, KU Leuven, Blijde Inkomststraat 21, 3000, Leuven, Belgium https://ror.org/01xzwj424 Künstliche Intelligenz / Kognitive Robotik, Hochschule für Technik und Wirtschaft, Friedrich-List-Platz 1, 01069, Dresden, Germany su: Language models Text mining Speech perception Code switching (Linguistics) Dutch language sug: subj: Language models Text mining Speech perception Code switching (Linguistics) Dutch language keyword: Dutch Punctuation Speech recognition Transformers ab: When applying automated speech recognition (ASR) for Belgian Dutch, the output consists of an unsegmented stream of words, without any punctuation. A next step is to perform segmentation and insert punctuation, making the ASR output more readable and easy to manually correct. We present the first (as far as we know) publicly available punctuation insertion system for Dutch that functions at a usable level and that is publicly available. The model we present here is an extension of the approach of Guhr et al. (In: Swiss Text Analytics Conference. Shared task on Sentence End and Punctuation Prediction in NLG Text, 2021) for Dutch: we finetuned the Dutch language model RobBERT on a punctuation prediction sequence classification task. The model was finetuned on two datasets: the Dutch side of Europarl and the SoNaR corpus. For every word in the input sequence, the model predicts a punctuation marker that follows the word. In cases where the language is unknown or where code switching applies, we have extended an existing multilingual model with Dutch. Previous work showed that such a multilingual model, based on "xlm-roberta-base" performs on par or sometimes even better than the monolingual cases. The system was evaluated on in-domain data as a classifier and on out-of-domain data as a sentence segmentation system through full stop prediction. The evaluations on sentence segmentation on out of domain data show that models finetuned on SoNaR show the best results, which can be attributed to SoNaR being a reference corpus containing different language registers. The multilingual models show an even better precision (at the cost of a lower recall) compared to the monolingual models. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|