Audio-Visual Automatic Speech Recognition Towards Education for Disabilities.

Education is a fundamental right that enriches everyone's life. However, physically challenged people often debar from the general and advanced education system. Audio-Visual Automatic Speech Recognition (AV-ASR) based system is useful to improve the education of physically challenged people by prov...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Autism & Developmental Disorders Vol. 53; no. 9; pp. 3581 - 3595
Autores principales: Debnath, Saswati, Roy, Pinki, Namasudra, Suyel, Crespo, Ruben Gonzalez
Formato: algorithm equations & formulas review tables/charts Journal Article
Publicado: Springer Nature Sep2023
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Education is a fundamental right that enriches everyone's life. However, physically challenged people often debar from the general and advanced education system. Audio-Visual Automatic Speech Recognition (AV-ASR) based system is useful to improve the education of physically challenged people by providing hands-free computing. They can communicate to the learning system through AV-ASR. However, it is challenging to trace the lip correctly for visual modality. Thus, this paper addresses the appearance-based visual feature along with the co-occurrence statistical measure for visual speech recognition. Local Binary Pattern-Three Orthogonal Planes (LBP-TOP) and Grey-Level Co-occurrence Matrix (GLCM) is proposed for visual speech information. The experimental results show that the proposed system achieves 76.60 % accuracy for visual speech and 96.00 % accuracy for audio speech recognition.