Mandarin-English code-switching speech corpus in South-East Asia: SEAME.

This paper introduces the South East Asia Mandarin-English corpus, a 63-h spontaneous Mandarin-English code-switching transcribed speech corpus suitable for LVCSR and language change detection/identification research. The corpus is recorded under unscripted interview and conversational settings from...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 49; no. 3; pp. 581 - 601
Autores principales: Lyu, Dau-Cheng, Tan, Tien-Ping, Chng, Eng-Siong, Li, Haizhou
Formato: Artículo
Publicado: Springer Nature Sep2015
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108465724&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 108465724
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2015
      vid: 49
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        108465724
        10.1007/s10579-015-9303-x
      ppf: 581
      ppct: 20
      formats:
        fmt:
          @attributes:
            type: P
            size: 717KB
      tig:
        atl: Mandarin-English code-switching speech corpus in South-East Asia: SEAME.
      aug:
        au:
          Lyu, Dau-Cheng
          Tan, Tien-Ping
          Chng, Eng-Siong
          Li, Haizhou
        affil:
          Temasek Laboratories, Nanyang Technological University, Singapore 639798 Singapore
          School of Computer Sciences, Universiti Sains Malaysia, 11800 USM Malaysia
      su:
        Mandarin dialects -- Study & teaching
        Chinese dialects
        Phoneme (Linguistics)
        Lexical access
        Word recognition
        Chinese language
        China
      sug:
        subj:
          China
          Mandarin dialects -- Study & teaching
          Chinese dialects
          Phoneme (Linguistics)
          Lexical access
          Word recognition
          Chinese language
      keyword:
        Code-switching speech
        Language recognition
        Mandarin-English
        Speech recognition
        Spontaneous spoken corpus development
      ab: This paper introduces the South East Asia Mandarin-English corpus, a 63-h spontaneous Mandarin-English code-switching transcribed speech corpus suitable for LVCSR and language change detection/identification research. The corpus is recorded under unscripted interview and conversational settings from 157 Singaporean and Malaysian speakers who spoke a mixture of Mandarin and English within a single sentence. About 82 % of the transcribed utterances are intra-sentential code-switching speech and the corpus will be release by LDC in 2015. This paper presents an analysis of the code-switching statistics of the corpus, such as the duration of monolingual segments and the frequency of language turns in code-switch utterances. We also summarize the development effort, details such as the processing time for transcription, validation and language boundary labelling. Lastly, we present textual analyses of code-switch segments examining the word length of monolingual segments in code-switch utterances and the most common single word and two-word phrase of such segments.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2015. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2015
    holdings:
      @attributes:
        islocal: N