XunZi-MLLM: a multimodal large language model for ancient text and image recognition.
Photocopies of ancient works, as valuable cultural heritage of China, can be digitized through the integration of multimodal large language models (MLLMs). This approach allows for more vivid representations of these historical documents, fostering the preservation and advancement of traditional cul...
| Publicado en: | Digital Scholarship in the Humanities Vol. 40; no. 2; pp. 709 - 723 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186085068&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186085068 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2025 vid: 40 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 186085068 10.1093/llc/fqaf026 ppf: 709 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.7MB tig: atl: XunZi-MLLM: a multimodal large language model for ancient text and image recognition. aug: au: Zhu, Dongmei Liu, Chang Zhao, Xue Zhao, Zhixiao Shen, Si Wang, Dongbo affil: College of Information Management, Nanjing Agricultural University, Nanjing 210095, China School of Economics & Management, Nanjing University of Science and Technology, Nanjing 210094, China su: Language models Text recognition Image recognition (Computer vision) Data scrubbing Historical source material sug: subj: Language models Text recognition Image recognition (Computer vision) Data scrubbing Historical source material keyword: ancient text recognition digitization of ancient works image recognition multimodal large language model photocopies of ancient works ab: Photocopies of ancient works, as valuable cultural heritage of China, can be digitized through the integration of multimodal large language models (MLLMs). This approach allows for more vivid representations of these historical documents, fostering the preservation and advancement of traditional culture. We propose a MLLM specifically dedicated to the digitization of ancient works photocopies. Specifically, we first use web crawling technology to efficiently gather ancient works data, followed by data cleaning and format conversion to ensure the data were suitable for model training. Next, we construct an unlabeled pre-training dataset and an ancient text dialogue dataset for fine-tuning based on the collected data. Finally, we further pre-train and fine-tune the MiniCPM-v-2.6-chat baseline model, enhancing its understanding of ancient works and its conversational ability. Experimental results indicate that the Xunzi-MiniCPM-v-2.6-chat model surpasses existing models in metrics such as BLEU, ROUGE, accuracy, recall, and F1 score. This model demonstrates strong performance in recognizing ancient texts and images, effectively processing multimodal information. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|