Multimodal Classification System for Hausa using LLMs and Vision Transformers.
This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoder...
| Publicado en: | Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 8 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Digital Humanities Association of Southern Africa (DHASA)
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoders, our system pre009 dicts answers from a fixed vocabulary. Exper010 iments conducted on the HaVQA dataset, un011 der offline text–image augmentation regimes, tailored to the specificity of Hausa as a low013 resource language, show that this augmentation strategy yields the best performance over the baseline, achieving 35.85% accuracy, 35.89% WuPalmer similarity, and 15.32% F1-score. |
|---|