SCTB-V2: the 2nd version of the Chinese treebank in the scientific domain.
Word segmentation, part-of-speech (POS) tagging, and syntactic parsing are three fundamental Chinese analysis tasks for Chinese language processing, which are also crucial for various downstream tasks such as machine translation and information extraction. To achieve high accuracy for these tasks, t...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 3; pp. 1389 - 1404 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=170029255&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 170029255 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2023 vid: 57 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 170029255 10.1007/s10579-022-09615-2 ppf: 1389 ppct: 15 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: SCTB-V2: the 2nd version of the Chinese treebank in the scientific domain. aug: au: Chu, Chenhui Mao, Zhuoyuan Nakazawa, Toshiaki Kawahara, Daisuke Kurohashi, Sadao affil: Kyoto University, Kyoto, Japan The University of Tokyo, Tokyo, Japan Waseda University, Tokyo, Japan su: Chinese language Machine translating Language research Scientific language Data mining sug: subj: Chinese language Machine translating Language research Scientific language Data mining keyword: Chinese Scientific domain Treebank ab: Word segmentation, part-of-speech (POS) tagging, and syntactic parsing are three fundamental Chinese analysis tasks for Chinese language processing, which are also crucial for various downstream tasks such as machine translation and information extraction. To achieve high accuracy for these tasks, treebanks that contain sentences manually annotated with word segmentation, part-of-speech tags, and phrase structures are essential. Although there are large-scale Chinese treebanks in the news domain, such treebanks are unavailable in the scientific domain. This significantly limits the performance of Chinese language processing for scientific text. To address this problem, we annotate the 2nd version of the Chinese treebank in the scientific domain (SCTB-V2). SCTB-V2 contains 12,175 sentences annotated with word segmentation, part-of-speech tags, and phrase structures. We conducted Chinese analyses and machine translation experiments on SCTB-V2. The results show the effectiveness of SCTB-V2. We release this treebank to promote scientific Chinese language processing research http://nlp.ist.i.kyoto-u.ac.jp/EN/index.php?A%20Chinese%20Treebank%20 in%20Scientific%20Domain%20%28SCTB%29. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|