DeepMine-multi-TTS: a Persian speech corpus for multi-speaker text-to-speech.
Speech synthesis has made significant progress in recent years thanks to deep neural networks (DNNs). However, one of the challenges of DNN-based models is the requirement for large and diverse data, which limits their applicability to many languages and domains. To date, no multi-speaker text-to-sp...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2245 - 2265 |
|---|---|
| Autores principales: | , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909066&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909066 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909066 10.1007/s10579-025-09807-6 ppf: 2245 ppct: 20 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: DeepMine-multi-TTS: a Persian speech corpus for multi-speaker text-to-speech. aug: au: Adibian, Majid Zeinali, Hossein Barmaki, Soroush affil: https://ror.org/04gzbav43 Department of Computer Engineering, Amirkabir University of Technology, Tehran, Iran Sharif DeepMine Ltd., Tehran, Iran su: Speech synthesis Databases Vocoder Artificial neural networks Corpora sug: subj: Speech synthesis Databases Vocoder Artificial neural networks Corpora keyword: Communication and Culture Linguistics Psychology and Cognitive Sciences Cognitive Sciences Information and Computing Sciences Artificial Intelligence and Image Processing Language Multi-speaker TTS dataset Persian speech corpus Text-to-speech (TTS) ab: Speech synthesis has made significant progress in recent years thanks to deep neural networks (DNNs). However, one of the challenges of DNN-based models is the requirement for large and diverse data, which limits their applicability to many languages and domains. To date, no multi-speaker text-to-speech (TTS) dataset has been available in Persian, which hinders the development of these models for this language. In this paper, we present a novel dataset for multi-speaker TTS in Persian, which consists of 120 h of high-quality speech from 67 speakers. We use this dataset to train two synthesizers and a vocoder and evaluate the quality of the synthesized speech. The results show that the naturalness of the generated samples, measured by the mean opinion score (MOS) criterion, is 3.94 and 4.12 for two trained multi-speaker synthesizers, which indicates that the dataset is suitable for training multi-speaker TTS models and can facilitate future research in this area for Persian. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|