Multimodal Classification System for Hausa using LLMs and Vision Transformers.

This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoder...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 8
Autores principales: Mijiyawa, Ali, Sadat, Fatiha
Formato: Artículo
Publicado: Digital Humanities Association of Southern Africa (DHASA) 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191564087&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 191564087
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid: N79T
      jtl: Journal of the Digital Humanities Association of Southern Africa (DHASA)
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 6
      iid: 2
      pid: 74161
      pub: Digital Humanities Association of Southern Africa (DHASA)
    artinfo:
      ui: 191564087
      ppf: 1
      ppct: 7
      formats:
      tig:
        atl: Multimodal Classification System for Hausa using LLMs and Vision Transformers.
      aug:
        au:
          Mijiyawa, Ali
          Sadat, Fatiha
        affil: Université du Québec à Montréal, Montreal, QC, Canada.
      su:
        Language models
        Transformer models
        Low-resource languages
        Classification algorithms
        Data augmentation
        Language & languages
      sug:
        subj:
          Language models
          Transformer models
          Low-resource languages
          Classification algorithms
          Data augmentation
          Language & languages
      ab: This paper presents a classification-based Vi003 sual Question Answering (VQA) system for the Hausa language, integrating Large Language Models (LLMs) and vision transformers. By fine-tuning LLMs on monolingual Hausa text and fusing their representations with those of state-of-the-art vision encoders, our system pre009 dicts answers from a fixed vocabulary. Exper010 iments conducted on the HaVQA dataset, un011 der offline text–image augmentation regimes, tailored to the specificity of Hausa as a low013 resource language, show that this augmentation strategy yields the best performance over the baseline, achieving 35.85% accuracy, 35.89% WuPalmer similarity, and 15.32% F1-score.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N