FluencyBank Timestamped: An Updated Data Set for Disfluency Detection and Automatic Intended Speech Recognition.

Purpose: This work introduces updated transcripts, disfluency annotations, and word timings for FluencyBank, which we refer to as FluencyBank Timestamped. This data set will enable the thorough analysis of how speech processing models (such as speech recognition and disfluency detection models) perf...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 67; no. 11; pp. 4203 - 4216
Autores principales: Romana, Amrit, Minxue Niu, Perez, Matthew, Mower Provost, Emily
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Nov2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=180765732&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 180765732
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Nov2024
      vid: 67
      iid: 11
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        180765732
        10.1044/2024_JSLHR-24-00070
      ppf: 4203
      ppct: 13
      formats:
        fmt:
          @attributes:
            type: P
            size: 895KB
      tig:
        atl: FluencyBank Timestamped: An Updated Data Set for Disfluency Detection and Automatic Intended Speech Recognition.
      aug:
        au:
          Romana, Amrit
          Minxue Niu
          Perez, Matthew
          Mower Provost, Emily
        affil: University of Michigan
      su:
        Linguistics
        Semantics
        Automatic speech recognition
        Stuttering
        Descriptive statistics
        Natural language processing
        Speech evaluation
        Speech perception
        Automation
        Comparative studies
        Data analysis software
        Auditory perception
      sug:
        subj:
          Linguistics
          Semantics
          Automatic speech recognition
          Stuttering
          Descriptive statistics
          Natural language processing
          Speech evaluation
          Speech perception
          Automation
          Comparative studies
          Data analysis software
          Auditory perception
      ab: Purpose: This work introduces updated transcripts, disfluency annotations, and word timings for FluencyBank, which we refer to as FluencyBank Timestamped. This data set will enable the thorough analysis of how speech processing models (such as speech recognition and disfluency detection models) perform when evaluated with typical speech versus speech from people who stutter (PWS). Method: We update the FluencyBank data set, which includes audio recordings from adults who stutter, to explore the robustness of speech processing models. Our update (semi-automated with manual review) includes new transcripts with timestamps and disfluency labels corresponding to each token in the transcript. Our disfluency labels capture typical disfluencies (filled pauses, repetitions, revisions, and partial words), and we explore how speech model performance compares for Switchboard (typical speech) and FluencyBank Time-stamped. We present benchmarks for three speech tasks: intended speech recognition, text-based disfluency detection, and audio-based disfluency detection. For the first task, we evaluate how well Whisper performs for intended speech recognition (i.e., transcribing speech without disfluencies). For the next tasks, we evaluate how well a Bidirectional Embedding Representations from Transformers (BERT) text-based model and a Whisper audio-based model perform for disfluency detection. We select these models, BERT and Whisper, as they have shown high accuracies on a broad range of tasks in their language and audio domains, respectively. Results: For the transcription task, we calculate an intended speech word error rate (isWER) between the model's output and the speaker's intended speech (i.e., speech without disfluencies). We find isWER is comparable between Switchboard and FluencyBank Timestamped, but that Whisper transcribes filled pauses and partial words at higher rates in the latter data set. Within Fluency-Bank Timestamped, isWER increases with stuttering severity. For the disfluency detection tasks, we find the models detect filled pauses, revisions, and partial words relatively well in FluencyBank Timestamped, but performance drops substantially for repetitions because the models are unable to generalize to the different types of repetitions (e.g., multiple repetitions and sound repetitions) from PWS. We hope that FluencyBank Timestamped will allow researchers to explore closing performance gaps between typical speech and speech from PWS. Conclusions: Our analysis shows that there are gaps in speech recognition and disfluency detection performance between typical speech and speech from PWS. We hope that FluencyBank Timestamped will contribute to more advancements in training robust speech processing models.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N