Automatic speech recognition system for Tunisian dialect.
Although Modern Standard Arabic is taught in schools and used in written communication and TV/radio broadcasts, all informal communication is typically carried out in dialectal Arabic. In this work, we focus on the design of speech tools and resources required for the development of an Automatic Sp...
| Publicado en: | Language Resources & Evaluation Vol. 52; no. 1; pp. 249 - 268 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=127930804&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 127930804 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2018 vid: 52 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 127930804 10.1007/s10579-017-9402-y ppf: 249 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.1MB tig: atl: Automatic speech recognition system for Tunisian dialect. aug: au: Masmoudi, Abir Bougares, Fethi Ellouze, Mariem Estève, Yannick Belguith, Lamia affil: LIUM, Le Mans University, Le Mans, France ANLP Research group, MIRACL Lab., University of Sfax, Sfax, Tunisia su: Written communication Speech perception Linguistics Frames (Linguistics) Dialects sug: subj: Written communication Speech perception Linguistics Frames (Linguistics) Dialects keyword: Automatic speech recognition Rule-based Grapheme-to-phoneme conversion Tunisian dialect Under-resourced language ab: Although Modern Standard Arabic is taught in schools and used in written communication and TV/radio broadcasts, all informal communication is typically carried out in dialectal Arabic. In this work, we focus on the design of speech tools and resources required for the development of an Automatic Speech Recognition system for the Tunisian dialect. The development of such a system faces the challenges of the lack of annotated resources and tools, apart from the lack of standardization at all linguistic levels (phonological, morphological, syntactic and lexical) together with the mispronunciation dictionary needed for ASR development. In this paper, we present a historical overview of the Tunisian dialect and its linguistic characteristics. We also describe and evaluate our rule-based phonetic tool. Next, we go deeper into the details of Tunisian dialect corpus creation. This corpus is finally approved and used to build the first ASR system for Tunisian dialect with a Word Error Rate of 22.6%. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|