Automatic sentence segmentation for classical Chinese: The Spring and Autumn Annals as an example.

There exists no sentence boundary in most classical Chinese literature texts. Since it is difficult to read literature of this kind, experts in literature or linguistics would segment the sentence manually. This article explores the effectiveness of classical Chinese sentence segmentation method so...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 38; no. 3; pp. 1067 - 1078
Autores principales: Fan, Wenjie, Wang, Dongbo, Huang, Shuiqing
Formato: Artículo
Publicado: Oxford University Press / USA Sep2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=171389426&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 171389426
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 38
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        171389426
        10.1093/llc/fqad016
      ppf: 1067
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 603KB
      tig:
        atl: Automatic sentence segmentation for classical Chinese: The Spring and Autumn Annals as an example.
      aug:
        au:
          Fan, Wenjie
          Wang, Dongbo
          Huang, Shuiqing
        affil:
          College of Information Management, Nanjing Agricultural University , Nanjing, China
          Research Center for Humanities and Social Computing, Nanjing Agricultural University , Nanjing, China
      su:
        Chinese language
        Autumn
        Long-term memory
      sug:
        subj:
          Chinese language
          Autumn
          Long-term memory
      ab: There exists no sentence boundary in most classical Chinese literature texts. Since it is difficult to read literature of this kind, experts in literature or linguistics would segment the sentence manually. This article explores the effectiveness of classical Chinese sentence segmentation method so as to provide a reference for classical Chinese punctuation. On the basis of the machine learning methods, we chose three components of machine learning, namely models, tagging schemes, and features, to compare the learning results. The models include conditional random field (CRF) models, long short term memory (LSTM) models, BiLSTM–CRF models, and three Bidirectional Encoder Representation from Transformers (BERT) models. There are five tagging schemes in this article and three features including the statistical feature, Guangyun, and Fanqie. Finally, the performance of the combined feature template is evaluated by ten-fold cross-validation on four classical Chinese texts in different genres. The SikuBERT model is proved to be the most effective model for sentence segmentation at present. Different tagging schemes and various features are introduced. The results show that 5-tag-J tagging schemes can improve performance. Statistical feature, as an important clue for classical Chinese sentence segmentation, is useful in related tasks, but Guangyun and Fanqie have little impact. Other important factors of sentence segmentation are genres and writing styles.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N