Evaluation of end-to-end continuous spanish lipreading in different data conditions.
Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inter-personal variability among speakers, and the complex modeling of silence. Nonetheless, recent remarkable results have...
| Published in: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2365 - 2387 |
|---|---|
| Main Authors: | , |
| Format: | Conference Paper/Materials |
| Published: |
Springer Nature
Sep2025
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909068&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909068 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909068 10.1007/s10579-025-09809-4 ppf: 2365 ppct: 22 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: Evaluation of end-to-end continuous spanish lipreading in different data conditions. aug: au: Gimeno-Gómez, David Martínez-Hinarejos, Carlos-D. affil: https://ror.org/01460j859 Pattern Recognition and Human Language Technologies Research Center, Universitat Politècnica de València, Camino de Vera, s/n, 46022, Valencia, Comunitat Valenciana, Spain su: Lipreading Spanish language Error analysis in mathematics Artificial neural networks Benchmark problems (Computer science) sug: subj: Lipreading Spanish language Error analysis in mathematics Artificial neural networks Benchmark problems (Computer science) keyword: Benchmarking Error analysis Psychology and Cognitive Sciences Psychology Information and Computing Sciences Artificial Intelligence and Image Processing Visual speech recognition ab: Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inter-personal variability among speakers, and the complex modeling of silence. Nonetheless, recent remarkable results have been achieved in the field thanks to the availability of large-scale databases and the use of powerful attention mechanisms. Besides, multiple languages apart from English are nowadays a focus of interest. This paper presents noticeable advances in automatic continuous lipreading for Spanish. First, an end-to-end system based on the hybrid CTC/Attention architecture is presented. Experiments are conducted on two corpora of disparate nature, reaching state-of-the-art results that significantly improve the best performance obtained to date for both databases. In addition, a thorough ablation study is carried out, where it is studied how the different components that form the architecture influence the quality of speech recognition. Then, a rigorous error analysis is carried out to investigate the different factors that could affect the learning of the automatic system. Finally, a new Spanish lipreading benchmark is consolidated. Code and trained models are available at https://github.com/david-gimeno/evaluating-end2end-spanish-lipreading. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|