Disfluency annotated corpora for Indian English in technical domains.
Disfluencies are common in spontaneous speech and can significantly affect the accuracy of automated systems that process spoken input. In this work, we tackled this issue for Indian English by developing a human-annotated disfluency corpus (DASIE (H)) comprising over 240K words for the technical le...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 1833 - 1865 |
|---|---|
| Autores principales: | , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909045&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909045 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909045 10.1007/s10579-024-09781-5 ppf: 1833 ppct: 32 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: Disfluency annotated corpora for Indian English in technical domains. aug: au: Mujadia, Vandan Mishra, Pruthwik Sharma, Dipti Misra affil: https://ror.org/00qryer39 Language Technologies Research Centre (LTRC), International Institute of Information Technology, 500032, Hyderabad, Telangana, India su: Speech processing systems Corpora Speech sug: subj: Speech processing systems Corpora Speech keyword: Communication and Culture Linguistics Disfluency for Indian English Disfluency processing Indian languages Information and Computing Sciences Artificial Intelligence and Image Processing Language Speech translation Technical domain ab: Disfluencies are common in spontaneous speech and can significantly affect the accuracy of automated systems that process spoken input. In this work, we tackled this issue for Indian English by developing a human-annotated disfluency corpus (DASIE (H)) comprising over 240K words for the technical lecture domain. To have a larger disfluency dataset, we introduced a method to generate synthetic disfluency, employing contextual embeddings and shallow linguistic features such as part-of-speech patterns. This algorithm allowed us to generate a synthetic disfluency corpus (DASIE (S)) that exceeds 15.4 million words. We evaluate the efficacy of our disfluency-annotated corpora by developing models for disfluency identification. Our efforts result in achieving the highest F1 score of 0.93 on the Switchboard test set and 0.80 on the DASIE (H) test set with the coarser disfluency identifier. The resulting corpora and model can be utilized to effectively detect and process disfluencies in various speech-interfacing applications. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|