Automatic Speech Recognition of Conversational Speech in Individuals With Disordered Speech.
Purpose: This study examines the effectiveness of automatic speech recognition (ASR) for individuals with speech disorders, addressing the gap in performance between read and conversational ASR. We analyze the factors influencing this disparity and the effect of speech mode-specific training on ASR...
| Publicado en: | Journal of Speech, Language & Hearing Research Vol. 67; no. 11; pp. 4176 - 4186 |
|---|---|
| Autores principales: | , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
American Speech-Language-Hearing Association
Nov2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=180765730&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 180765730 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 10924388 1SM jtl: Journal of Speech, Language & Hearing Research issn: 10924388 maglogo: N pubinfo: dt: Nov2024 vid: 67 iid: 11 pid: 42 pub: American Speech-Language-Hearing Association artinfo: ui: 180765730 10.1044/2024_JSLHR-24-00045 ppf: 4176 ppct: 10 formats: fmt: @attributes: type: P size: 741KB tig: atl: Automatic Speech Recognition of Conversational Speech in Individuals With Disordered Speech. aug: au: Tobin, Jimmy Nelson, Phillip MacDonald, Bob Heywood, Rus Cave, Richard Seaver, Katie Desjardins, Antoine Pan-Pan Jiang Green, Jordan R. affil: Google LLC, Mountain View, CA MND Association, Northampton, United Kingdom MGH Institute of Health Professions, Boston, MA Harvard University, Cambridge, MA su: United States United Kingdom Conversation Speech Artificial intelligence Cell phones Linguistics Communication Speech disorders Psychosocial factors Speech therapy Automatic speech recognition Speech therapists Scale analysis (Psychology) Multiple regression analysis Pilot projects Severity of illness index Descriptive statistics Physiological aspects of speech Metadata Artificial neural networks Automation Speech perception Comparative studies Confidence intervals Sensitivity & specificity (Statistics) sug: subj: Conversation Speech Artificial intelligence Cell phones Linguistics Communication Speech disorders Psychosocial factors United States United Kingdom Offices of Physical, Occupational and Speech Therapists, and Audiologists Electronics Stores Radio and Television Broadcasting and Wireless Communications Equipment Manufacturing Electronic components, navigational and communications equipment and supplies merchant wholesalers Wireless Telecommunications Carriers (except Satellite) Speech therapy Automatic speech recognition Speech therapists Scale analysis (Psychology) Multiple regression analysis Pilot projects Severity of illness index Descriptive statistics Physiological aspects of speech Metadata Artificial neural networks Automation Speech perception Comparative studies Confidence intervals Sensitivity & specificity (Statistics) ab: Purpose: This study examines the effectiveness of automatic speech recognition (ASR) for individuals with speech disorders, addressing the gap in performance between read and conversational ASR. We analyze the factors influencing this disparity and the effect of speech mode-specific training on ASR accuracy. Method: Recordings of read and conversational speech from 27 individuals with various speech disorders were analyzed using both (a) one speaker-independent ASR system trained and optimized for typical speech and (b) multiple ASR models that were personalized to the speech of the participants with disordered speech. Word error rates were calculated for each speech model, read versus conversational, and subject. Linear mixed-effects models were used to assess the impact of speech mode and disorder severity on ASR accuracy. We investigated nine variables, classified as technical, linguistic, or speech impairment factors, for their potential influence on the performance gap. Results: We found a significant performance gap between read and conversational speech in both personalized and unadapted ASR models. Speech impairment severity notably impacted recognition accuracy in unadapted models for both speech modes and in personalized models for read speech. Linguistic attributes of utterances were the most influential on accuracy, though atypical speech characteristics also played a role. Including conversational speech samples in model training notably improved recognition accuracy. Conclusions: We observed a significant performance gap in ASR accuracy between read and conversational speech for individuals with speech disorders. This gap was largely due to the linguistic complexity and unique characteristics of speech disorders in conversational speech. Training personalized ASR models using conversational speech significantly improved recognition accuracy, demonstrating the importance of domain-specific training and highlighting the need for further research into ASR systems capable of handling disordered conversational speech effectively. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|