Community-Supported Shared Infrastructure in Support of Speech Accessibility.
Purpose: The Speech Accessibility Project (SAP) intends to facilitate research and development in automatic speech recognition (ASR) and other machine learning tasks for people with speech disabilities. The purpose of this article is to introduce this project as a resource for researchers, including...
| Publicado en: | Journal of Speech, Language & Hearing Research Vol. 67; no. 11; pp. 4162 - 4176 |
|---|---|
| Autores principales: | , , , , , , , , , , , , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
American Speech-Language-Hearing Association
Nov2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=180765729&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 180765729 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 10924388 1SM jtl: Journal of Speech, Language & Hearing Research issn: 10924388 maglogo: N pubinfo: dt: Nov2024 vid: 67 iid: 11 pid: 42 pub: American Speech-Language-Hearing Association artinfo: ui: 180765729 10.1044/2024_JSLHR-24-00122 ppf: 4162 ppct: 14 formats: fmt: @attributes: type: P size: 1.5MB tig: atl: Community-Supported Shared Infrastructure in Support of Speech Accessibility. aug: au: Hasegawa-Johnson, Mark Xiuwen Zheng Heejin Kim Mendes, Clarion Dickinson, Meg Hege, Erik Zwilling, Chris Moore Channell, Marie Mattie, Laura Hodges, Heather Ramig, Lorraine Bellard, Mary Shebanek, Mike Sari, Leda Kalgaonkar, Kaustubh Frerichs, David Bigham, Jeffrey P. Findlater, Leah Lea, Colin Herrlinger, Sarah affil: University of Illinois Urbana-Champaign LSVT Global, Tucson, AZ Microsoft, Redmond, WA Meta, Menlo Park, CA Media Tuners LLC, Los Altos, CA Apple, Cupertino, CA su: Illinois United States Community support Health services accessibility Cell phones Assistive technology Speech disorders Personal computers People with disabilities Automatic speech recognition Dysarthria Descriptive statistics Parkinson's disease Machine learning Data analysis software sug: subj: Community support Health services accessibility Cell phones Assistive technology Speech disorders Personal computers People with disabilities Illinois United States Computer and peripheral equipment manufacturing Electronic Computer Manufacturing Wireless Telecommunications Carriers (except Satellite) Electronics Stores Radio and Television Broadcasting and Wireless Communications Equipment Manufacturing Electronic components, navigational and communications equipment and supplies merchant wholesalers Automatic speech recognition Dysarthria Descriptive statistics Parkinson's disease Machine learning Data analysis software ab: Purpose: The Speech Accessibility Project (SAP) intends to facilitate research and development in automatic speech recognition (ASR) and other machine learning tasks for people with speech disabilities. The purpose of this article is to introduce this project as a resource for researchers, including baseline analysis of the first released data package. Method: The project aims to facilitate ASR research by collecting, curating, and distributing transcribed U.S. English speech from people with speech and/or language disabilities. Participants record speech from their place of residence by connecting their personal computer, cell phone, and assistive devices, if needed, to the SAP web portal. All samples are manually transcribed, and 30 per participant are annotated using differential diagnostic pattern dimensions. For purposes of ASR experiments, the participants have been randomly assigned to a training set, a development set for controlled testing of a trained ASR, and a test set to evaluate ASR error rate. Results: The SAP 2023-10-05 Data Package contains the speech of 211 people with dysarthria as a correlate of Parkinson's disease, and the associated test set contains 42 additional speakers. A baseline ASR, with a word error rate of 3.4% for typical speakers, transcribes test speech with a word error rate of 36.3%. Fine-tuning reduces the word error rate to 23.7%. Conclusions: Preliminary findings suggest that a large corpus of dysarthric and dysphonic speech has the potential to significantly improve speech technology for people with disabilities. By providing these data to researchers, the SAP intends to significantly accelerate research into accessible speech technology. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|