SeSoDa: A Compact Context-Rich Sesotho-English Dataset for LoRA Fine-Tuning of SLMs.
We introduce SeSoDa, a multidomain Sesotho(Sa Lesotho)-English dataset of 1,966 prompt-completion pairs that span six categories (nouns, verbs, idioms, quantifiers, grammar rules, usage alerts). SeSoDa documents the morphosyntactic complexity, uncaptured Basotho cultural specificity, and orthographi...
| Publicado en: | Journal of the Digital Humanities Association of Southern Africa (DHASA) Vol. 6; no. 2; pp. 1 - 10 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Digital Humanities Association of Southern Africa (DHASA)
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=191564078&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 191564078 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: N79T jtl: Journal of the Digital Humanities Association of Southern Africa (DHASA) maglogo: N pubinfo: dt: 2025 vid: 6 iid: 2 pid: 74161 pub: Digital Humanities Association of Southern Africa (DHASA) artinfo: ui: 191564078 ppf: 1 ppct: 9 formats: tig: atl: SeSoDa: A Compact Context-Rich Sesotho-English Dataset for LoRA Fine-Tuning of SLMs. aug: au: Mandla, Motaung Hill, Graham Mots’oehli, Moseli affil: National University of Lesotho. MindForge AI. The Shard, South Africa. University of South Africa. University of Hawai‘i at Manoa. su: Linguistic complexity Machine translating Language models Culture Natural language processing Lesotho sug: subj: Lesotho Linguistic complexity Machine translating Language models Culture Natural language processing ab: We introduce SeSoDa, a multidomain Sesotho(Sa Lesotho)-English dataset of 1,966 prompt-completion pairs that span six categories (nouns, verbs, idioms, quantifiers, grammar rules, usage alerts). SeSoDa documents the morphosyntactic complexity, uncaptured Basotho cultural specificity, and orthographic/phonological differences between Lesotho and South African Sesotho. We created a user-friendly, JSON-style corpus with detailed metadata. This aims to lower the technical barrier for new researchers in Lesotho, helping them advance culture-aware machine translation, linguistic analysis, and cultural preservation using AI. As a proof of concept, we demonstrate SeSoDa’s utility by fine-tuning the TinyLlama-1.1B-Chat model using Low-Rank Adaptation (LoRA) on entirely free Google Colab GPUs and runtime limits. This parameterefficient fine-tuning approach is particularly vital for resource-constrained environments like Lesotho, making advanced NLP model adaptation feasible and accessible without requiring extensive computational resources. We open-source the code for the dataset creation, the baseline model, and the dataset itself. We hope to see both Basotho researchers and developers build on top of our effort. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|