A Tunable Forced Alignment System Based on Deep Learning: Applications to Child Speech.

Purpose: Phonetic forced alignment has a multitude of applications in automated analysis of speech, particularly in studying nonstandard speech such as children's speech. Manual alignment is tedious but serves as the gold standard for clinical-grade alignment. Current tools do not support direct tra...

Full description

Bibliographic Details
Published in:Journal of Speech, Language & Hearing Research Vol. 68; pp. 3583 - 3602
Main Authors: Kadambi, Prad, Mahr, Tristan J., Hustad, Katherine C., Berisha, Visar
Format: Article
Published: American Speech-Language-Hearing Association 2025 Supplement
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=187102507&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 187102507
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: 2025 Supplement
      vid: 68
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        187102507
        10.1044/2024_JSLHR-24-00347
      ppf: 3583
      ppct: 19
      formats:
        fmt:
          @attributes:
            type: P
            size: 8.4MB
      tig:
        atl: A Tunable Forced Alignment System Based on Deep Learning: Applications to Child Speech.
      aug:
        au:
          Kadambi, Prad
          Mahr, Tristan J.
          Hustad, Katherine C.
          Berisha, Visar
        affil:
          School of Electrical, Computer and Energy Engineering, Arizona State University, Tempe
          College of Health Solutions, Arizona State University, Tempe
          Waisman Center, University of Wisconsin--Madison
          Department of Communication Sciences and Disorders, University of Wisconsin--Madison
      su:
        Child development
        Phonetics
        Language acquisition
        Children
        Physiological adaptation
        Research funding
        Descriptive statistics
        Verbal behavior testing
        Audiometry
        Intelligibility of speech
        Physiological aspects of speech
        Deep learning
        Speech evaluation
        Speech audiometry
        Automation
        Comparative studies
        Articulation (Speech)
      sug:
        subj:
          Child development
          Phonetics
          Language acquisition
          Children
          Physiological adaptation
          Research funding
          Descriptive statistics
          Verbal behavior testing
          Audiometry
          Intelligibility of speech
          Physiological aspects of speech
          Deep learning
          Speech evaluation
          Speech audiometry
          Automation
          Comparative studies
          Articulation (Speech)
      ab: Purpose: Phonetic forced alignment has a multitude of applications in automated analysis of speech, particularly in studying nonstandard speech such as children's speech. Manual alignment is tedious but serves as the gold standard for clinical-grade alignment. Current tools do not support direct training on manual alignments. Thus, a trainable speaker adaptive phonetic forced alignment system, Wav2TextGrid, was developed for children's speech. The source code for the method is publicly available along with a graphical user interface at https://github.com/pkadambi/Wav2TextGrid. Method: We propose a trainable, speaker-adaptive, neural forced aligner developed using a corpus of 42 neurotypical children from 3 to 6 years of age. Evaluation on both child speech and on the TIMIT corpus was performed to demonstrate aligner performance across age and dialectal variations. Results: The trainable alignment tool markedly improved accuracy over baseline for several alignment quality metrics, for all phoneme categories. Accuracy for plosives and affricates in children's speech improved more than 40% over baseline. Performance matched existing methods using approximately 13 min of labeled data, while approximately 45--60 min of labeled alignments yielded significant improvement. Conclusion: The Wav2TextGrid tool allows alternate alignment workflows where the forced alignments, via training, are directly tailored to match clinical-grade, manually provided alignments.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N