Corpora compilation for prosody-informed speech processing.
Research on speech technologies necessitates spoken data, which is usually obtained through read recorded speech, and specifically adapted to the research needs. When the aim is to deal with the prosody involved in speech, the available data must reflect natural and conversational speech, which is u...
| Publicado en: | Language Resources & Evaluation Vol. 55; no. 4; pp. 925 - 947 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2021
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=152947522&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 152947522 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2021 vid: 55 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 152947522 10.1007/s10579-021-09556-2 ppf: 925 ppct: 22 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.5MB tig: atl: Corpora compilation for prosody-informed speech processing. aug: au: Öktem, Alp Farrús, Mireia Bonafonte, Antonio affil: Universitat Pompeu Fabra/Col·lectivaT, Barcelona, Spain Universitat Pompeu Fabra/Universitat de Barcelona, Barcelona, Spain Universitat Politècnica de Catalunya, Barcelona, Spain su: Corpora Machine translating Lipreading Scientific community Prosodic analysis (Linguistics) sug: subj: Corpora Machine translating Lipreading Scientific community Prosodic analysis (Linguistics) keyword: F0 Intensity Parallel data Pause Punctuation Speech corpus Speech transcription Spoken machine translation ab: Research on speech technologies necessitates spoken data, which is usually obtained through read recorded speech, and specifically adapted to the research needs. When the aim is to deal with the prosody involved in speech, the available data must reflect natural and conversational speech, which is usually costly and difficult to get. This paper presents a machine learning-oriented toolkit for collecting, handling, and visualization of speech data, using prosodic heuristic. We present two corpora resulting from these methodologies: PANTED corpus, containing 250 h of English speech from TED Talks, and Heroes corpus containing 8 h of parallel English and Spanish movie speech. We demonstrate their use in two deep learning-based applications: punctuation restoration and machine translation. The presented corpora are freely available to the research community. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|