SeSoDa: A Compact Context-Rich Sesotho-English Dataset for LoRA Fine-Tuning of SLMs.

We introduce SeSoDa, a multidomain Sesotho(Sa Lesotho)-English dataset of 1,966 prompt-completion pairs that span six categories (nouns, verbs, idioms, quantifiers, grammar rules, usage alerts). SeSoDa documents the morphosyntactic complexity, uncaptured Basotho cultural specificity, and orthographi...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 10
Autores principales: Mandla, Motaung, Hill, Graham, Mots’oehli, Moseli
Formato: Artículo
Publicado: Digital Humanities Association of Southern Africa (DHASA) 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191564078&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 191564078
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid: N79T
      jtl: Journal of the Digital Humanities Association of Southern Africa (DHASA)
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 6
      iid: 2
      pid: 74161
      pub: Digital Humanities Association of Southern Africa (DHASA)
    artinfo:
      ui: 191564078
      ppf: 1
      ppct: 9
      formats:
      tig:
        atl: SeSoDa: A Compact Context-Rich Sesotho-English Dataset for LoRA Fine-Tuning of SLMs.
      aug:
        au:
          Mandla, Motaung
          Hill, Graham
          Mots’oehli, Moseli
        affil:
          National University of Lesotho.
          MindForge AI.
          The Shard, South Africa.
          University of South Africa.
          University of Hawai‘i at Manoa.
      su:
        Linguistic complexity
        Machine translating
        Language models
        Culture
        Natural language processing
        Lesotho
      sug:
        subj:
          Lesotho
          Linguistic complexity
          Machine translating
          Language models
          Culture
          Natural language processing
      ab: We introduce SeSoDa, a multidomain Sesotho(Sa Lesotho)-English dataset of 1,966 prompt-completion pairs that span six categories (nouns, verbs, idioms, quantifiers, grammar rules, usage alerts). SeSoDa documents the morphosyntactic complexity, uncaptured Basotho cultural specificity, and orthographic/phonological differences between Lesotho and South African Sesotho. We created a user-friendly, JSON-style corpus with detailed metadata. This aims to lower the technical barrier for new researchers in Lesotho, helping them advance culture-aware machine translation, linguistic analysis, and cultural preservation using AI. As a proof of concept, we demonstrate SeSoDa’s utility by fine-tuning the TinyLlama-1.1B-Chat model using Low-Rank Adaptation (LoRA) on entirely free Google Colab GPUs and runtime limits. This parameterefficient fine-tuning approach is particularly vital for resource-constrained environments like Lesotho, making advanced NLP model adaptation feasible and accessible without requiring extensive computational resources. We open-source the code for the dataset creation, the baseline model, and the dataset itself. We hope to see both Basotho researchers and developers build on top of our effort.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N