Automatic sentence segmentation for classical Chinese: The Spring and Autumn Annals as an example.
There exists no sentence boundary in most classical Chinese literature texts. Since it is difficult to read literature of this kind, experts in literature or linguistics would segment the sentence manually. This article explores the effectiveness of classical Chinese sentence segmentation method so...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 3; pp. 1067 - 1078 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=171389426&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 171389426 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Sep2023 vid: 38 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 171389426 10.1093/llc/fqad016 ppf: 1067 ppct: 11 formats: fmt: – @attributes: type: T – @attributes: type: P size: 603KB tig: atl: Automatic sentence segmentation for classical Chinese: The Spring and Autumn Annals as an example. aug: au: Fan, Wenjie Wang, Dongbo Huang, Shuiqing affil: College of Information Management, Nanjing Agricultural University , Nanjing, China Research Center for Humanities and Social Computing, Nanjing Agricultural University , Nanjing, China su: Chinese language Autumn Long-term memory sug: subj: Chinese language Autumn Long-term memory ab: There exists no sentence boundary in most classical Chinese literature texts. Since it is difficult to read literature of this kind, experts in literature or linguistics would segment the sentence manually. This article explores the effectiveness of classical Chinese sentence segmentation method so as to provide a reference for classical Chinese punctuation. On the basis of the machine learning methods, we chose three components of machine learning, namely models, tagging schemes, and features, to compare the learning results. The models include conditional random field (CRF) models, long short term memory (LSTM) models, BiLSTM–CRF models, and three Bidirectional Encoder Representation from Transformers (BERT) models. There are five tagging schemes in this article and three features including the statistical feature, Guangyun, and Fanqie. Finally, the performance of the combined feature template is evaluated by ten-fold cross-validation on four classical Chinese texts in different genres. The SikuBERT model is proved to be the most effective model for sentence segmentation at present. Different tagging schemes and various features are introduced. The results show that 5-tag-J tagging schemes can improve performance. Statistical feature, as an important clue for classical Chinese sentence segmentation, is useful in related tasks, but Guangyun and Fanqie have little impact. Other important factors of sentence segmentation are genres and writing styles. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|