Mandarin-English code-switching speech corpus in South-East Asia: SEAME.
This paper introduces the South East Asia Mandarin-English corpus, a 63-h spontaneous Mandarin-English code-switching transcribed speech corpus suitable for LVCSR and language change detection/identification research. The corpus is recorded under unscripted interview and conversational settings from...
| Publicado en: | Language Resources & Evaluation Vol. 49; no. 3; pp. 581 - 601 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108465724&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 108465724 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2015 vid: 49 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 108465724 10.1007/s10579-015-9303-x ppf: 581 ppct: 20 formats: fmt: @attributes: type: P size: 717KB tig: atl: Mandarin-English code-switching speech corpus in South-East Asia: SEAME. aug: au: Lyu, Dau-Cheng Tan, Tien-Ping Chng, Eng-Siong Li, Haizhou affil: Temasek Laboratories, Nanyang Technological University, Singapore 639798 Singapore School of Computer Sciences, Universiti Sains Malaysia, 11800 USM Malaysia su: Mandarin dialects -- Study & teaching Chinese dialects Phoneme (Linguistics) Lexical access Word recognition Chinese language China sug: subj: China Mandarin dialects -- Study & teaching Chinese dialects Phoneme (Linguistics) Lexical access Word recognition Chinese language keyword: Code-switching speech Language recognition Mandarin-English Speech recognition Spontaneous spoken corpus development ab: This paper introduces the South East Asia Mandarin-English corpus, a 63-h spontaneous Mandarin-English code-switching transcribed speech corpus suitable for LVCSR and language change detection/identification research. The corpus is recorded under unscripted interview and conversational settings from 157 Singaporean and Malaysian speakers who spoke a mixture of Mandarin and English within a single sentence. About 82 % of the transcribed utterances are intra-sentential code-switching speech and the corpus will be release by LDC in 2015. This paper presents an analysis of the code-switching statistics of the corpus, such as the duration of monolingual segments and the frequency of language turns in code-switch utterances. We also summarize the development effort, details such as the processing time for transcription, validation and language boundary labelling. Lastly, we present textual analyses of code-switch segments examining the word length of monolingual segments in code-switch utterances and the most common single word and two-word phrase of such segments. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2015. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|