DeepMine-multi-TTS: a Persian speech corpus for multi-speaker text-to-speech.

Speech synthesis has made significant progress in recent years thanks to deep neural networks (DNNs). However, one of the challenges of DNN-based models is the requirement for large and diverse data, which limits their applicability to many languages and domains. To date, no multi-speaker text-to-sp...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2245 - 2265
Autores principales: Adibian, Majid, Zeinali, Hossein, Barmaki, Soroush
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909066&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909066
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909066
        10.1007/s10579-025-09807-6
      ppf: 2245
      ppct: 20
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: DeepMine-multi-TTS: a Persian speech corpus for multi-speaker text-to-speech.
      aug:
        au:
          Adibian, Majid
          Zeinali, Hossein
          Barmaki, Soroush
        affil:
          https://ror.org/04gzbav43 Department of Computer Engineering, Amirkabir University of Technology, Tehran, Iran
          Sharif DeepMine Ltd., Tehran, Iran
      su:
        Speech synthesis
        Databases
        Vocoder
        Artificial neural networks
        Corpora
      sug:
        subj:
          Speech synthesis
          Databases
          Vocoder
          Artificial neural networks
          Corpora
      keyword:
        Communication and Culture Linguistics Psychology and Cognitive Sciences Cognitive Sciences Information and Computing Sciences Artificial Intelligence and Image Processing
        Language
        Multi-speaker TTS dataset
        Persian speech corpus
        Text-to-speech (TTS)
      ab: Speech synthesis has made significant progress in recent years thanks to deep neural networks (DNNs). However, one of the challenges of DNN-based models is the requirement for large and diverse data, which limits their applicability to many languages and domains. To date, no multi-speaker text-to-speech (TTS) dataset has been available in Persian, which hinders the development of these models for this language. In this paper, we present a novel dataset for multi-speaker TTS in Persian, which consists of 120 h of high-quality speech from 67 speakers. We use this dataset to train two synthesizers and a vocoder and evaluate the quality of the synthesized speech. The results show that the naturalness of the generated samples, measured by the mean opinion score (MOS) criterion, is 3.94 and 4.12 for two trained multi-speaker synthesizers, which indicates that the dataset is suitable for training multi-speaker TTS models and can facilitate future research in this area for Persian.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N