The Korean Speech Recognition Sentences: A Large Corpus for Evaluating Semantic Context and Language Experience in Speech Perception.

Purpose: The aim of this study was to develop and validate a large Korean sentence set with varying degrees of semantic predictability that can be used for testing speech recognition and lexical processing. Method: Sentences differing in the degree of final-word predictability (predictable, neutral,...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Speech, Language & Hearing Research Vol. 66; no. 9; pp. 3399 - 3413
Autores principales: Jieun Song, Byungjun Kim, Minjeong Kim, Iverson, Paul
Formato: Artículo
Publicado: American Speech-Language-Hearing Association Sep2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=171950380&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 171950380
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        10924388
        1SM
      jtl: Journal of Speech, Language & Hearing Research
      issn: 10924388
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 66
      iid: 9
      pid: 42
      pub: American Speech-Language-Hearing Association
    artinfo:
      ui:
        171950380
        10.1044/2023_JSLHR-23-00137
      ppf: 3399
      ppct: 14
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.2MB
      tig:
        atl: The Korean Speech Recognition Sentences: A Large Corpus for Evaluating Semantic Context and Language Experience in Speech Perception.
      aug:
        au:
          Jieun Song
          Byungjun Kim
          Minjeong Kim
          Iverson, Paul
        affil:
          School of Digital Humanities and Computational Social Sciences, Korea Advanced Institute of Science and Technology, Daejeon, South Korea
          Center for Digital Humanities and Computational Social Sciences, Korea Advanced Institute of Science and Technology, Daejeon, South Korea
          Graduate School of Culture Technology, Korea Advanced Institute of Science and Technology, Daejeon, South Korea
          Department of Speech, Hearing and Phonetic Sciences, University College London, United Kingdom
      su:
        South Korea
        Semantics
        Comparative grammar
        Language acquisition
        Language disorders
        Speech perception
        Computer software
        Deep learning
        Physiological aspects of speech
        Phonological awareness
        Research methodology evaluation
        Noise
        Research methodology
        Intelligibility of speech
        Comparative studies
        Descriptive statistics
        Chi-squared test
        Research funding
      sug:
        subj:
          Semantics
          Comparative grammar
          Language acquisition
          Language disorders
          South Korea
          Computer and software stores
          Software publishers (except video game publishers)
          Computer, computer peripheral and pre-packaged software merchant wholesalers
          Computer and Computer Peripheral Equipment and Software Merchant Wholesalers
          Speech perception
          Computer software
          Deep learning
          Physiological aspects of speech
          Phonological awareness
          Research methodology evaluation
          Noise
          Research methodology
          Intelligibility of speech
          Comparative studies
          Descriptive statistics
          Chi-squared test
          Research funding
      ab: Purpose: The aim of this study was to develop and validate a large Korean sentence set with varying degrees of semantic predictability that can be used for testing speech recognition and lexical processing. Method: Sentences differing in the degree of final-word predictability (predictable, neutral, and anomalous) were created with words selected to be suitable for both native and nonnative speakers of Korean. Semantic predictability was evaluated through a series of cloze tests in which native (n = 56) and nonnative (n = 19) speakers of Korean participated. This study also used a computer language model to evaluate final-word predictabilities; this is a novel approach that the current study adopted to reduce human effort in validating a large number of sentences, which produced results comparable to those of the cloze tests. In a speech recognition task, the sentences were presented to native (n = 23) and nonnative (n = 21) speakers of Korean in speech-shaped noise at two levels of noise. Results: The results of the speech-in-noise experiment demonstrated that the intelligibility of the sentences was similar to that of related English corpora. That is, intelligibility was significantly different depending on the semantic condition, and the sentences had the right degree of difficulty for assessing intelligibility differences depending on noise levels and language experience. Conclusions: This corpus (1,021 sentences in total) adds to the target languages available in speech research and will allow researchers to investigate a range of issues in speech perception in Korean. Supplemental Material: https://doi.org/10.23641/asha.24045582
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N