XunZi-MLLM: a multimodal large language model for ancient text and image recognition.

Photocopies of ancient works, as valuable cultural heritage of China, can be digitized through the integration of multimodal large language models (MLLMs). This approach allows for more vivid representations of these historical documents, fostering the preservation and advancement of traditional cul...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 40; no. 2; pp. 709 - 723
Autores principales: Zhu, Dongmei, Liu, Chang, Zhao, Xue, Zhao, Zhixiao, Shen, Si, Wang, Dongbo
Formato: Artículo
Publicado: Oxford University Press / USA Jun2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186085068&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186085068
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 40
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        186085068
        10.1093/llc/fqaf026
      ppf: 709
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.7MB
      tig:
        atl: XunZi-MLLM: a multimodal large language model for ancient text and image recognition.
      aug:
        au:
          Zhu, Dongmei
          Liu, Chang
          Zhao, Xue
          Zhao, Zhixiao
          Shen, Si
          Wang, Dongbo
        affil:
          College of Information Management, Nanjing Agricultural University, Nanjing 210095, China
          School of Economics & Management, Nanjing University of Science and Technology, Nanjing 210094, China
      su:
        Language models
        Text recognition
        Image recognition (Computer vision)
        Data scrubbing
        Historical source material
      sug:
        subj:
          Language models
          Text recognition
          Image recognition (Computer vision)
          Data scrubbing
          Historical source material
      keyword:
        ancient text recognition
        digitization of ancient works
        image recognition
        multimodal large language model
        photocopies of ancient works
      ab: Photocopies of ancient works, as valuable cultural heritage of China, can be digitized through the integration of multimodal large language models (MLLMs). This approach allows for more vivid representations of these historical documents, fostering the preservation and advancement of traditional culture. We propose a MLLM specifically dedicated to the digitization of ancient works photocopies. Specifically, we first use web crawling technology to efficiently gather ancient works data, followed by data cleaning and format conversion to ensure the data were suitable for model training. Next, we construct an unlabeled pre-training dataset and an ancient text dialogue dataset for fine-tuning based on the collected data. Finally, we further pre-train and fine-tune the MiniCPM-v-2.6-chat baseline model, enhancing its understanding of ancient works and its conversational ability. Experimental results indicate that the Xunzi-MiniCPM-v-2.6-chat model surpasses existing models in metrics such as BLEU, ROUGE, accuracy, recall, and F1 score. This model demonstrates strong performance in recognizing ancient texts and images, effectively processing multimodal information.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N