DIGITAL PALEOGRAPHY IN ARABIC MANUSCRIPTS: THE PERFORMANCE LEVELS OF THE GEMINI MODEL IN RECOGNIZING NASKH, TA'LIQ, AND RUQ'AH SCRIPTS.

The digitization of manuscripts remains a critical challenge in Digital Humanities due to the morphological complexity of Arabic scripts. This study evaluates the performance of Large Language Models, specifically Gemini, in digitizing Arabic manuscripts. Using manuscripts in naskh, ruq'ah, and ta'l...

Descripción completa

Detalles Bibliográficos
Publicado en:Dinbilimleri Journal Vol. 26; no. 2; pp. 909 - 941
Autores principales: ADIGÜZEL, Abdulcabbar, ÖZBEK, Muhammed Sadık
Formato: Artículo
Publicado: Dinbilimleri Journal haz2026
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The digitization of manuscripts remains a critical challenge in Digital Humanities due to the morphological complexity of Arabic scripts. This study evaluates the performance of Large Language Models, specifically Gemini, in digitizing Arabic manuscripts. Using manuscripts in naskh, ruq'ah, and ta'līq scripts, model outputs were analyzed via Levenshtein distance, Word Error Rate, and Character Error Rate metrics. Results indicate recognition accuracy varies by script, with raw character accuracy at 97.91% for naskh, 94.32% for ruq'ah, and 93.51% for ta'līq. A substantial portion of errors in ruq'ah and ta'līq arose from orthographic variations in the letter yā and the hamza. Excluding these anomalies, accuracy across all scripts consistently reached the 97-98% range. The study concludes that while AI models demonstrate robust performance in deciphering manuscript templates, their integration into classical text edition (tahqīq) requires careful oversight of paleographic nuances, despite their high potential to assist professional publication. The experiment was conducted with the Gemini 3.1 Pro model under a fixed-prompt protocol. Since only the opening page of a single manuscript was used per script type, the findings should be interpreted as preliminary indicators rather than definitive conclusions.