Automatic Speech Recognition of Conversational Speech in Individuals With Disordered Speech.

Purpose: This study examines the effectiveness of automatic speech recognition (ASR) for individuals with speech disorders, addressing the gap in performance between read and conversational ASR. We analyze the factors influencing this disparity and the effect of speech mode-specific training on ASR...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 67; no. 11; pp. 4176 - 4186
Autores principales: Tobin, Jimmy, Nelson, Phillip, MacDonald, Bob, Heywood, Rus, Cave, Richard, Seaver, Katie, Desjardins, Antoine, Pan-Pan Jiang, Green, Jordan R.
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Nov2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=180765730&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 180765730
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Nov2024
      vid: 67
      iid: 11
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        180765730
        10.1044/2024_JSLHR-24-00045
      ppf: 4176
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
            size: 741KB
      tig:
        atl: Automatic Speech Recognition of Conversational Speech in Individuals With Disordered Speech.
      aug:
        au:
          Tobin, Jimmy
          Nelson, Phillip
          MacDonald, Bob
          Heywood, Rus
          Cave, Richard
          Seaver, Katie
          Desjardins, Antoine
          Pan-Pan Jiang
          Green, Jordan R.
        affil:
          Google LLC, Mountain View, CA
          MND Association, Northampton, United Kingdom
          MGH Institute of Health Professions, Boston, MA
          Harvard University, Cambridge, MA
      su:
        United States
        United Kingdom
        Conversation
        Speech
        Artificial intelligence
        Cell phones
        Linguistics
        Communication
        Speech disorders
        Psychosocial factors
        Speech therapy
        Automatic speech recognition
        Speech therapists
        Scale analysis (Psychology)
        Multiple regression analysis
        Pilot projects
        Severity of illness index
        Descriptive statistics
        Physiological aspects of speech
        Metadata
        Artificial neural networks
        Automation
        Speech perception
        Comparative studies
        Confidence intervals
        Sensitivity & specificity (Statistics)
      sug:
        subj:
          Conversation
          Speech
          Artificial intelligence
          Cell phones
          Linguistics
          Communication
          Speech disorders
          Psychosocial factors
          United States
          United Kingdom
          Offices of Physical, Occupational and Speech Therapists, and Audiologists
          Electronics Stores
          Radio and Television Broadcasting and Wireless Communications Equipment Manufacturing
          Electronic components, navigational and communications equipment and supplies merchant wholesalers
          Wireless Telecommunications Carriers (except Satellite)
          Speech therapy
          Automatic speech recognition
          Speech therapists
          Scale analysis (Psychology)
          Multiple regression analysis
          Pilot projects
          Severity of illness index
          Descriptive statistics
          Physiological aspects of speech
          Metadata
          Artificial neural networks
          Automation
          Speech perception
          Comparative studies
          Confidence intervals
          Sensitivity & specificity (Statistics)
      ab: Purpose: This study examines the effectiveness of automatic speech recognition (ASR) for individuals with speech disorders, addressing the gap in performance between read and conversational ASR. We analyze the factors influencing this disparity and the effect of speech mode-specific training on ASR accuracy. Method: Recordings of read and conversational speech from 27 individuals with various speech disorders were analyzed using both (a) one speaker-independent ASR system trained and optimized for typical speech and (b) multiple ASR models that were personalized to the speech of the participants with disordered speech. Word error rates were calculated for each speech model, read versus conversational, and subject. Linear mixed-effects models were used to assess the impact of speech mode and disorder severity on ASR accuracy. We investigated nine variables, classified as technical, linguistic, or speech impairment factors, for their potential influence on the performance gap. Results: We found a significant performance gap between read and conversational speech in both personalized and unadapted ASR models. Speech impairment severity notably impacted recognition accuracy in unadapted models for both speech modes and in personalized models for read speech. Linguistic attributes of utterances were the most influential on accuracy, though atypical speech characteristics also played a role. Including conversational speech samples in model training notably improved recognition accuracy. Conclusions: We observed a significant performance gap in ASR accuracy between read and conversational speech for individuals with speech disorders. This gap was largely due to the linguistic complexity and unique characteristics of speech disorders in conversational speech. Training personalized ASR models using conversational speech significantly improved recognition accuracy, demonstrating the importance of domain-specific training and highlighting the need for further research into ASR systems capable of handling disordered conversational speech effectively.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N