Speech emotion recognition for the Urdu language: Dataset and evaluation.
Crafting reliable Speech Emotion Recognition systems is an arduous task that inevitably requires large amounts of data for training purposes. Such voluminous datasets are currently obtainable in only a few languages, including English, German, and Italian. In this work, we present SEMOUR + : a Scrip...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 2; pp. 915 - 945 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=163826589&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 163826589 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2023 vid: 57 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 163826589 10.1007/s10579-022-09610-7 ppf: 915 ppct: 30 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.8MB tig: atl: Speech emotion recognition for the Urdu language: Dataset and evaluation. aug: au: Zaheer, Nimra Ahmad, Obaid Ullah Shabbir, Mudassir Raza, Agha Ali affil: Computer Science Department, Information Technology University, Lahore, Pakistan Computer Science Department, Vanderbilt University, Nashville, TN, USA Computer Science Department, Lahore University of Management Sciences, Lahore, Pakistan su: Automatic speech recognition Urdu language Affective forecasting (Psychology) Speech Native language Emotion recognition sug: subj: Automatic speech recognition Urdu language Affective forecasting (Psychology) Speech Native language Emotion recognition keyword: Accent diversity Deep learning Emotional speech dataset Speech emotion recognition ab: Crafting reliable Speech Emotion Recognition systems is an arduous task that inevitably requires large amounts of data for training purposes. Such voluminous datasets are currently obtainable in only a few languages, including English, German, and Italian. In this work, we present SEMOUR + : a Scripted EMOtional Speech Repository for Urdu, the first scripted database of emotion-tagged and diverse-accent speech in the Urdu language, to design an Urdu Speech Emotion Recognition system. Our gender-balanced 14-h repository contains 27, 640 unique instances recorded by 24 native speakers eliciting a syntactically complex script. The dataset is phonetically balanced, and reliably exhibits varied emotions, as marked by the high agreement scores among human raters in experiments. We also provide various baseline speech emotion prediction scores on SEMOUR + , which could be utilized for multiple applications like personalized robot assistants, diagnosis of psychological disorders, getting feedback from a low-tech-enabled population, etc. In a speaker-independent experimental setting, our ensemble model accurately predicts an emotion with a state-of-the-art 56 % accuracy. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|