Benchmarking Hindi-to-English direct speech-to-speech translation with synthetic data.

Speech-to-speech translation (S2ST) tasks aim to translate speech from one language to another. Recent research focuses on direct S2ST models, which do not rely on intermediate text representation. This approach is useful for bridging the gap across multilingual communities. Towards such overarching...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2613 - 2652
Autores principales: Gupta, Mahendra, Dutta, Maitreyee, Maurya, Chandresh Kumar
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909084&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909084
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909084
        10.1007/s10579-025-09827-2
      ppf: 2613
      ppct: 39
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 4MB
      tig:
        atl: Benchmarking Hindi-to-English direct speech-to-speech translation with synthetic data.
      aug:
        au:
          Gupta, Mahendra
          Dutta, Maitreyee
          Maurya, Chandresh Kumar
        affil:
          https://ror.org/00rpc7w94 Department of Computer Science & Engineering, NITTTR, Sector-26, 160019, Chandigarh, India
          https://ror.org/00rpc7w94 Department of Information Management & Emerging Engineering, NITTTR, Sector-26, 160019, Chandigarh, India
          https://ror.org/01hhf7w52 Department of Computer Science & Engineering, Indian Institute of Technology (IIT) Indore, Simrol, 453552, Indore, M.P., India
      su:
        Hindi language
        Low-resource languages
        Corpora
        Oral communication
        Transformer models
        Translating & interpreting
      sug:
        subj:
          Hindi language
          Low-resource languages
          Corpora
          Oral communication
          Transformer models
          Translating & interpreting
      keyword:
        Communication and Culture Linguistics Psychology and Cognitive Sciences Cognitive Sciences
        Language
        Semantic similarity
        Speech-to-speech translation
      ab: Speech-to-speech translation (S2ST) tasks aim to translate speech from one language to another. Recent research focuses on direct S2ST models, which do not rely on intermediate text representation. This approach is useful for bridging the gap across multilingual communities. Towards such overarching goals, creating parallel speech corpora is a challenging and expensive process, resulting in the limited availability of datasets in various languages. Most of the available S2ST datasets are available only in high-resource languages. As such, direct S2ST models have not been tested in low-resource languages. Therefore, we present a Hindi–English S2ST dataset considered a low-resource language pair where raw speech and text are sourced from the TED talk platform. A cost-effective self-supervised pruning method, leveraging Cross-lingual Semantic Similarity and Word Error Rate (WER), is employed to enhance the quality of the developed dataset. Manual validation conducted by human evaluators on sampled data further confirms the high quality of the dataset. Further, existing S2ST models are evaluated through extensive experiments to establish a baseline for the Hindi–English language pair on the developed dataset. For pre-training and data augmentation, pseudo-labeled data is also used to improve the performance of baseline models. The performances of direct S2ST models are compared with the cascade baseline S2ST models. The results indicate that the Transformer-based direct S2ST model achieves a translation accuracy of 15.86 BLEU score after data augmentation, which lags the cascade model by a gap of 2.27 BLEU score. The dataset will be open-sourced after the acceptance of the paper.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N