ChisNERE: a premodern Chinese corpus with named entity and relation annotation.

This work contributes to the digital humanities approach for studying premodern Chinese history and culture by creating a large-scale dataset annotated with named entities and relations. Through careful annotation guidelines and labeling of over 200,000 characters, we developed a dataset containing...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 40; no. 2; pp. 617 - 639
Main Authors: Tang, Xuemei, Deng, Zekun, Wang, Jun, Su, Qi
Format: Article
Published: Oxford University Press / USA Jun2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186085056&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186085056
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 40
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        186085056
        10.1093/llc/fqaf001
      ppf: 617
      ppct: 22
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.5MB
      tig:
        atl: ChisNERE: a premodern Chinese corpus with named entity and relation annotation.
      aug:
        au:
          Tang, Xuemei
          Deng, Zekun
          Wang, Jun
          Su, Qi
        affil:
          Department of Information Management, Peking University, 100871 Beijing, China
          Research Center for Digital Humanities, Peking University, 100871 Beijing, China
          School of Foreign Languages, Peking University, 100871 Beijing, China
      su:
        Language models
        Knowledge graphs
        History of technology
        Chinese language
        Chinese history
      sug:
        subj:
          Language models
          Knowledge graphs
          History of technology
          Chinese language
          Chinese history
      keyword:
        ancient Chinese
        annotation
        dataset
        named entity recognition
        relation extraction
      ab: This work contributes to the digital humanities approach for studying premodern Chinese history and culture by creating a large-scale dataset annotated with named entities and relations. Through careful annotation guidelines and labeling of over 200,000 characters, we developed a dataset containing 30,000 named entities across six types and 7,000 relations spanning twenty categories. Experiments on named entity recognition (NER) using pre-trained language models and large language models on this dataset achieved an initial performance of NER (91.32 percent F1). In addition, relationship extraction (RE) on the pretrained language model achieves an 85.32 percent F1 score. While there is still room for improvement, our annotated dataset and models provide a useful starting point for extracting semantic information from premodern Chinese texts. It represents an effort to connect history and technology, increasing accessibility and preservation of premodern Chinese cultural treasures. Furthermore, our dataset can facilitate downstream tasks like culture analysis, knowledge graph construction, and computational understanding of premodern Chinese. Overall, this research represents a significant step toward digitally exploring premodern Chinese documents, providing a pathway for future work on knowledge organization and computational analysis of this valuable cultural legacy. Our code and data are available at: https://github.com/tangxuemei1995/AnChineseNERE
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N