A large-scale comparison of two voice synthesis techniques on intelligibility, naturalness, preferences, and attitudes toward voices banked by individuals with amyotrophic lateral sclerosis.

Amyotrophic lateral sclerosis (ALS) commonly results in the inability to produce natural speech, making speech-generating devices (SGDs) important. Historically, synthetic voices generated by SGDs were neither unique, nor age- or dialect-appropriate, which depersonalized SGD use. Voices generated by...

Descripción completa

Detalles Bibliográficos
Publicado en:AAC: Augmentative & Alternative Communication Vol. 40; no. 1; pp. 31 - 46
Autores principales: Hyppa-Martin, Jolene, Lilley, Jason, Chen, Mo, Friese, Jaclyn, Schmidt, Corinne, Bunnell, H. Timothy
Formato: research tables/charts Journal Article
Publicado: Taylor & Francis Ltd Mar2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=175671430&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 175671430
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        07434618
        9OD
      jtl: AAC: Augmentative & Alternative Communication
      issn: 07434618
      maglogo: Y
    pubinfo:
      dt: Mar2024
      vid: 40
      iid: 1
      pid: 377
      pub: Taylor & Francis Ltd
      place: Philadelphia, Pennsylvania
    artinfo:
      ui:
        175671430
        172754246
        175671430
        175671430
        10.1080/07434618.2023.2262032
        175671430
      ppf: 31
      ppct: 15
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: A large-scale comparison of two voice synthesis techniques on intelligibility, naturalness, preferences, and attitudes toward voices banked by individuals with amyotrophic lateral sclerosis.
      aug:
        au:
          Hyppa-Martin, Jolene
          Lilley, Jason
          Chen, Mo
          Friese, Jaclyn
          Schmidt, Corinne
          Bunnell, H. Timothy
        affil: Department of Communication Sciences and Disorders, University of Minnesota Duluth, Duluth, MN, USA
      sug:
        subj:
          Amyotrophic Lateral Sclerosis
          Dysarthria
          Communication Aids for Persons with Disabilities
          Audiorecording
          Listening
          Speech Intelligibility Evaluation
          Voice Evaluation
          Alternative and Augmentative Communication
          Human
          Male
          Female
          Adolescence
          Adult
          Middle Age
          Aged
          Aged, 80 and Over
          Experimental Studies
          Severity of Illness
          Acoustic Stimulation
          Convenience Sample
          Data Analysis Software
          Analysis of Variance
          Chi Square Test
          T-Tests
          Descriptive Statistics
          Comparative Studies
          Funding Source
          Adolescent: 13-18 years
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Aged: 65+ years
          Aged, 80 & over
          Male
          Female
      ab: Amyotrophic lateral sclerosis (ALS) commonly results in the inability to produce natural speech, making speech-generating devices (SGDs) important. Historically, synthetic voices generated by SGDs were neither unique, nor age- or dialect-appropriate, which depersonalized SGD use. Voices generated by SGDs can now be customized via voice banking and should ideally sound uniquely like the individual's natural speech, be intelligible, and elicit positive reactions from communication partners. This large-scale 2 x 2 mixed between- and within-participants design examined perceptions of 831 adult listeners regarding custom synthetic voices created for two individuals diagnosed with ALS via two synthesis systems in common clinical use (waveform concatenation and statistical parametric synthesis). The study explored relationships among synthesis system, dysarthria severity, synthetic speech intelligibility, naturalness, and preferences, and also provided a preliminary examination of attitudes regarding the custom synthetic voices. Synthetic voices generated via statistical parametric synthesis trained on deep neural networks were more intelligible, natural, and preferred than voices produced via waveform concatenation, and were associated with more positive attitudes. The custom synthetic voice created from moderately dysarthric speech was more intelligible than the voice created from mildly dysarthric speech. Clinical implications and factors that may have contributed to the relative intelligibilities are discussed.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N