SCTB-V2: the 2nd version of the Chinese treebank in the scientific domain.

Word segmentation, part-of-speech (POS) tagging, and syntactic parsing are three fundamental Chinese analysis tasks for Chinese language processing, which are also crucial for various downstream tasks such as machine translation and information extraction. To achieve high accuracy for these tasks, t...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 3; pp. 1389 - 1404
Autores principales: Chu, Chenhui, Mao, Zhuoyuan, Nakazawa, Toshiaki, Kawahara, Daisuke, Kurohashi, Sadao
Formato: Artículo
Publicado: Springer Nature Sep2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=170029255&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 170029255
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 57
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        170029255
        10.1007/s10579-022-09615-2
      ppf: 1389
      ppct: 15
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: SCTB-V2: the 2nd version of the Chinese treebank in the scientific domain.
      aug:
        au:
          Chu, Chenhui
          Mao, Zhuoyuan
          Nakazawa, Toshiaki
          Kawahara, Daisuke
          Kurohashi, Sadao
        affil:
          Kyoto University, Kyoto, Japan
          The University of Tokyo, Tokyo, Japan
          Waseda University, Tokyo, Japan
      su:
        Chinese language
        Machine translating
        Language research
        Scientific language
        Data mining
      sug:
        subj:
          Chinese language
          Machine translating
          Language research
          Scientific language
          Data mining
      keyword:
        Chinese
        Scientific domain
        Treebank
      ab: Word segmentation, part-of-speech (POS) tagging, and syntactic parsing are three fundamental Chinese analysis tasks for Chinese language processing, which are also crucial for various downstream tasks such as machine translation and information extraction. To achieve high accuracy for these tasks, treebanks that contain sentences manually annotated with word segmentation, part-of-speech tags, and phrase structures are essential. Although there are large-scale Chinese treebanks in the news domain, such treebanks are unavailable in the scientific domain. This significantly limits the performance of Chinese language processing for scientific text. To address this problem, we annotate the 2nd version of the Chinese treebank in the scientific domain (SCTB-V2). SCTB-V2 contains 12,175 sentences annotated with word segmentation, part-of-speech tags, and phrase structures. We conducted Chinese analyses and machine translation experiments on SCTB-V2. The results show the effectiveness of SCTB-V2. We release this treebank to promote scientific Chinese language processing research http://nlp.ist.i.kyoto-u.ac.jp/EN/index.php?A%20Chinese%20Treebank%20 in%20Scientific%20Domain%20%28SCTB%29.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N