Corpora compilation for prosody-informed speech processing.

Research on speech technologies necessitates spoken data, which is usually obtained through read recorded speech, and specifically adapted to the research needs. When the aim is to deal with the prosody involved in speech, the available data must reflect natural and conversational speech, which is u...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 55; no. 4; pp. 925 - 947
Autores principales: Öktem, Alp, Farrús, Mireia, Bonafonte, Antonio
Formato: Artículo
Publicado: Springer Nature Dec2021
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=152947522&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 152947522
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2021
      vid: 55
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        152947522
        10.1007/s10579-021-09556-2
      ppf: 925
      ppct: 22
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.5MB
      tig:
        atl: Corpora compilation for prosody-informed speech processing.
      aug:
        au:
          Öktem, Alp
          Farrús, Mireia
          Bonafonte, Antonio
        affil:
          Universitat Pompeu Fabra/Col·lectivaT, Barcelona, Spain
          Universitat Pompeu Fabra/Universitat de Barcelona, Barcelona, Spain
          Universitat Politècnica de Catalunya, Barcelona, Spain
      su:
        Corpora
        Machine translating
        Lipreading
        Scientific community
        Prosodic analysis (Linguistics)
      sug:
        subj:
          Corpora
          Machine translating
          Lipreading
          Scientific community
          Prosodic analysis (Linguistics)
      keyword:
        F0
        Intensity
        Parallel data
        Pause
        Punctuation
        Speech corpus
        Speech transcription
        Spoken machine translation
      ab: Research on speech technologies necessitates spoken data, which is usually obtained through read recorded speech, and specifically adapted to the research needs. When the aim is to deal with the prosody involved in speech, the available data must reflect natural and conversational speech, which is usually costly and difficult to get. This paper presents a machine learning-oriented toolkit for collecting, handling, and visualization of speech data, using prosodic heuristic. We present two corpora resulting from these methodologies: PANTED corpus, containing 250 h of English speech from TED Talks, and Heroes corpus containing 8 h of parallel English and Spanish movie speech. We demonstrate their use in two deep learning-based applications: punctuation restoration and machine translation. The presented corpora are freely available to the research community.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N