Disfluency annotated corpora for Indian English in technical domains.

Disfluencies are common in spontaneous speech and can significantly affect the accuracy of automated systems that process spoken input. In this work, we tackled this issue for Indian English by developing a human-annotated disfluency corpus (DASIE (H)) comprising over 240K words for the technical le...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 1833 - 1865
Autores principales: Mujadia, Vandan, Mishra, Pruthwik, Sharma, Dipti Misra
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909045&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909045
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909045
        10.1007/s10579-024-09781-5
      ppf: 1833
      ppct: 32
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.3MB
      tig:
        atl: Disfluency annotated corpora for Indian English in technical domains.
      aug:
        au:
          Mujadia, Vandan
          Mishra, Pruthwik
          Sharma, Dipti Misra
        affil: https://ror.org/00qryer39 Language Technologies Research Centre (LTRC), International Institute of Information Technology, 500032, Hyderabad, Telangana, India
      su:
        Speech processing systems
        Corpora
        Speech
      sug:
        subj:
          Speech processing systems
          Corpora
          Speech
      keyword:
        Communication and Culture Linguistics
        Disfluency for Indian English
        Disfluency processing
        Indian languages
        Information and Computing Sciences Artificial Intelligence and Image Processing Language
        Speech translation
        Technical domain
      ab: Disfluencies are common in spontaneous speech and can significantly affect the accuracy of automated systems that process spoken input. In this work, we tackled this issue for Indian English by developing a human-annotated disfluency corpus (DASIE (H)) comprising over 240K words for the technical lecture domain. To have a larger disfluency dataset, we introduced a method to generate synthetic disfluency, employing contextual embeddings and shallow linguistic features such as part-of-speech patterns. This algorithm allowed us to generate a synthetic disfluency corpus (DASIE (S)) that exceeds 15.4 million words. We evaluate the efficacy of our disfluency-annotated corpora by developing models for disfluency identification. Our efforts result in achieving the highest F1 score of 0.93 on the Switchboard test set and 0.80 on the DASIE (H) test set with the coarser disfluency identifier. The resulting corpora and model can be utilized to effectively detect and process disfluencies in various speech-interfacing applications.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N