ChisNERE: a premodern Chinese corpus with named entity and relation annotation.
This work contributes to the digital humanities approach for studying premodern Chinese history and culture by creating a large-scale dataset annotated with named entities and relations. Through careful annotation guidelines and labeling of over 200,000 characters, we developed a dataset containing...
| Published in: | Digital Scholarship in the Humanities Vol. 40; no. 2; pp. 617 - 639 |
|---|---|
| Main Authors: | , , , |
| Format: | Article |
| Published: |
Oxford University Press / USA
Jun2025
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186085056&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186085056 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2025 vid: 40 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 186085056 10.1093/llc/fqaf001 ppf: 617 ppct: 22 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.5MB tig: atl: ChisNERE: a premodern Chinese corpus with named entity and relation annotation. aug: au: Tang, Xuemei Deng, Zekun Wang, Jun Su, Qi affil: Department of Information Management, Peking University, 100871 Beijing, China Research Center for Digital Humanities, Peking University, 100871 Beijing, China School of Foreign Languages, Peking University, 100871 Beijing, China su: Language models Knowledge graphs History of technology Chinese language Chinese history sug: subj: Language models Knowledge graphs History of technology Chinese language Chinese history keyword: ancient Chinese annotation dataset named entity recognition relation extraction ab: This work contributes to the digital humanities approach for studying premodern Chinese history and culture by creating a large-scale dataset annotated with named entities and relations. Through careful annotation guidelines and labeling of over 200,000 characters, we developed a dataset containing 30,000 named entities across six types and 7,000 relations spanning twenty categories. Experiments on named entity recognition (NER) using pre-trained language models and large language models on this dataset achieved an initial performance of NER (91.32 percent F1). In addition, relationship extraction (RE) on the pretrained language model achieves an 85.32 percent F1 score. While there is still room for improvement, our annotated dataset and models provide a useful starting point for extracting semantic information from premodern Chinese texts. It represents an effort to connect history and technology, increasing accessibility and preservation of premodern Chinese cultural treasures. Furthermore, our dataset can facilitate downstream tasks like culture analysis, knowledge graph construction, and computational understanding of premodern Chinese. Overall, this research represents a significant step toward digitally exploring premodern Chinese documents, providing a pathway for future work on knowledge organization and computational analysis of this valuable cultural legacy. Our code and data are available at: https://github.com/tangxuemei1995/AnChineseNERE pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|