The Korean Speech Recognition Sentences: A Large Corpus for Evaluating Semantic Context and Language Experience in Speech Perception.
Purpose: The aim of this study was to develop and validate a large Korean sentence set with varying degrees of semantic predictability that can be used for testing speech recognition and lexical processing. Method: Sentences differing in the degree of final-word predictability (predictable, neutral,...
| Publicado en: | Journal of Speech, Language & Hearing Research Vol. 66; no. 9; pp. 3399 - 3413 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
American Speech-Language-Hearing Association
Sep2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=171950380&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 171950380 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 10924388 1SM jtl: Journal of Speech, Language & Hearing Research issn: 10924388 maglogo: N pubinfo: dt: Sep2023 vid: 66 iid: 9 pid: 42 pub: American Speech-Language-Hearing Association artinfo: ui: 171950380 10.1044/2023_JSLHR-23-00137 ppf: 3399 ppct: 14 formats: fmt: @attributes: type: P size: 1.2MB tig: atl: The Korean Speech Recognition Sentences: A Large Corpus for Evaluating Semantic Context and Language Experience in Speech Perception. aug: au: Jieun Song Byungjun Kim Minjeong Kim Iverson, Paul affil: School of Digital Humanities and Computational Social Sciences, Korea Advanced Institute of Science and Technology, Daejeon, South Korea Center for Digital Humanities and Computational Social Sciences, Korea Advanced Institute of Science and Technology, Daejeon, South Korea Graduate School of Culture Technology, Korea Advanced Institute of Science and Technology, Daejeon, South Korea Department of Speech, Hearing and Phonetic Sciences, University College London, United Kingdom su: South Korea Semantics Comparative grammar Language acquisition Language disorders Speech perception Computer software Deep learning Physiological aspects of speech Phonological awareness Research methodology evaluation Noise Research methodology Intelligibility of speech Comparative studies Descriptive statistics Chi-squared test Research funding sug: subj: Semantics Comparative grammar Language acquisition Language disorders South Korea Computer and software stores Software publishers (except video game publishers) Computer, computer peripheral and pre-packaged software merchant wholesalers Computer and Computer Peripheral Equipment and Software Merchant Wholesalers Speech perception Computer software Deep learning Physiological aspects of speech Phonological awareness Research methodology evaluation Noise Research methodology Intelligibility of speech Comparative studies Descriptive statistics Chi-squared test Research funding ab: Purpose: The aim of this study was to develop and validate a large Korean sentence set with varying degrees of semantic predictability that can be used for testing speech recognition and lexical processing. Method: Sentences differing in the degree of final-word predictability (predictable, neutral, and anomalous) were created with words selected to be suitable for both native and nonnative speakers of Korean. Semantic predictability was evaluated through a series of cloze tests in which native (n = 56) and nonnative (n = 19) speakers of Korean participated. This study also used a computer language model to evaluate final-word predictabilities; this is a novel approach that the current study adopted to reduce human effort in validating a large number of sentences, which produced results comparable to those of the cloze tests. In a speech recognition task, the sentences were presented to native (n = 23) and nonnative (n = 21) speakers of Korean in speech-shaped noise at two levels of noise. Results: The results of the speech-in-noise experiment demonstrated that the intelligibility of the sentences was similar to that of related English corpora. That is, intelligibility was significantly different depending on the semantic condition, and the sentences had the right degree of difficulty for assessing intelligibility differences depending on noise levels and language experience. Conclusions: This corpus (1,021 sentences in total) adds to the target languages available in speech research and will allow researchers to investigate a range of issues in speech perception in Korean. Supplemental Material: https://doi.org/10.23641/asha.24045582 pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|