Demographic and Acoustic Factors Related to Automatic Speech Recognition Inaccuracies for Child African American English Speakers.

Purpose: This study investigated the relationship between acoustic measures and Google's Speech-to-Text inaccuracies in recognizing speech of children ages 4-9 years who speak African American English (AAE). Method: Audio recordings were collected from 11 AAE-speaking children with speech stimuli ta...

Full description

Bibliographic Details
Published in:Perspectives of the ASHA Special Interest Groups Vol. 10; no. 6; pp. 2278 - 2298
Main Authors: Fletcher, Brittany N., Hsu, Wei-Wen, Novak, Vesna D., Wilkens, Mary E., Hobek, Amy W., Pratt, Amy S., Leon, Michelle, Harrell, Kimmerly, McKenna, Victoria S.
Format: research tables/charts Journal Article
Published: American Speech-Language-Hearing Association Dec2025
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=190171850&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 190171850
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        2381473X
        KTSD
      jtl: Perspectives of the ASHA Special Interest Groups
      issn: 2381473X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 10
      iid: 6
      pid: 42
      pub: American Speech-Language-Hearing Association
      place: Rockville, Maryland
    artinfo:
      ui:
        190171850
        190171850
        190171850
        10.1044/2025_PERSP-25-00052
        190171850
      ppf: 2278
      ppct: 20
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Demographic and Acoustic Factors Related to Automatic Speech Recognition Inaccuracies for Child African American English Speakers.
      aug:
        au:
          Fletcher, Brittany N.
          Hsu, Wei-Wen
          Novak, Vesna D.
          Wilkens, Mary E.
          Hobek, Amy W.
          Pratt, Amy S.
          Leon, Michelle
          Harrell, Kimmerly
          McKenna, Victoria S.
        affil: Department of Communication Sciences and Disorders, University of Cincinnati, OH
      sug:
        subj:
          Voice Recognition Systems Methods
          Sociodemographic Factors
          Acoustics Evaluation
          African American English
          Human
          Funding Source
          Male
          Female
          Child, Preschool
          Child
          Exploratory Research
          Pilot Studies
          Descriptive Statistics
          Logistic Regression
          ROC Curve
          Confidence Intervals
          Speech Production Measurement
          Accents and Dialects
          Vowels
          Voice Evaluation
          Cues
          Speech Evaluation
          Child, Preschool: 2-5 years
          Child: 6-12 years
          Male
          Female
      ab: Purpose: This study investigated the relationship between acoustic measures and Google's Speech-to-Text inaccuracies in recognizing speech of children ages 4-9 years who speak African American English (AAE). Method: Audio recordings were collected from 11 AAE-speaking children with speech stimuli targeting final plosive variations observed within the AAE dialect. Dialectal density was measured using the Diagnostic Evaluation of Language Variation Language Screener. Recordings were transcribed using Google's Speech-to-Text application (Google Voice), and inaccuracies were determined through comparison to researcher-extracted transcriptions. Acoustic measures from vowels preceding final plosives (including vowel duration, fundamental frequency, average first formant) were extracted using Praat and a custom MATLAB algorithm. Individual mixed-effects logistic regression models were conducted to analyze the relationships between acoustic measures and transcription accuracy (accurate vs. inaccurate) for voiced and voiceless plosives separately. Results: There were no significant differences between inaccuracy rates for voiced and voiceless plosive productions, nor were acoustic measures predictive of automatic speech recognition inaccuracy. However, age and dialect density were significantly related to voiceless plosive accuracy. Conclusions: The complexities of voice, motor, and articulatory development within children can be characterized by acoustic measures. These measures inform acoustic algorithms created for speech technology. Research on acoustic measures in young child AAE speech, with considerations for dialect variability and age, will enhance speech recognition technology and clinical best practices.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N