Multimodal Classification System for Hausa using LLMs and Vision Transformers.
This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoder...
| Publicado en: | Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 8 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Digital Humanities Association of Southern Africa (DHASA)
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191564087&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 191564087 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: N79T jtl: Journal of the Digital Humanities Association of Southern Africa (DHASA) maglogo: N pubinfo: dt: 2025 vid: 6 iid: 2 pid: 74161 pub: Digital Humanities Association of Southern Africa (DHASA) artinfo: ui: 191564087 ppf: 1 ppct: 7 formats: tig: atl: Multimodal Classification System for Hausa using LLMs and Vision Transformers. aug: au: Mijiyawa, Ali Sadat, Fatiha affil: Université du Québec à Montréal, Montreal, QC, Canada. su: Language models Transformer models Low-resource languages Classification algorithms Data augmentation Language & languages sug: subj: Language models Transformer models Low-resource languages Classification algorithms Data augmentation Language & languages ab: This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoders, our system pre009 dicts answers from a fixed vocabulary. Exper010 iments conducted on the HaVQA dataset, un011 der offline text–image augmentation regimes, tailored to the specificity of Hausa as a low013 resource language, show that this augmentation strategy yields the best performance over the baseline, achieving 35.85% accuracy, 35.89% WuPalmer similarity, and 15.32% F1-score. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|