Multimodal Classification System for Hausa using LLMs and Vision Transformers.

This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoder...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 8
Autores principales: Mijiyawa, Ali, Sadat, Fatiha
Formato: Artículo
Publicado: Digital Humanities Association of Southern Africa (DHASA) 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoders, our system pre009 dicts answers from a fixed vocabulary. Exper010 iments conducted on the HaVQA dataset, un011 der offline text–image augmentation regimes, tailored to the specificity of Hausa as a low013 resource language, show that this augmentation strategy yields the best performance over the baseline, achieving 35.85% accuracy, 35.89% WuPalmer similarity, and 15.32% F1-score.