Myanmar XNLI: building a dataset and exploring low-resource approaches to natural language inference with Myanmar.

Despite dramatic recent progress in NLP, it is still a major challenge to apply Large Language Models (LLM) to low-resource languages. This is made visible in benchmarks such as Cross-Lingual Natural Language Inference (XNLI), a key task that demonstrates cross-lingual capabilities of NLP systems ac...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3267 - 3311
Autores principales: Htet, Aung Kyaw, Dras, Mark
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909097&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909097
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909097
        10.1007/s10579-025-09844-1
      ppf: 3267
      ppct: 44
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.5MB
      tig:
        atl: Myanmar XNLI: building a dataset and exploring low-resource approaches to natural language inference with Myanmar.
      aug:
        au:
          Htet, Aung Kyaw
          Dras, Mark
        affil: https://ror.org/01sf06y89 Department of Computing, Macquarie University, Sydney, Australia
      su:
        Low-resource languages
        Data augmentation
        Acquisition of data
        Citizen science
        Natural languages
        Natural language processing
        Language models
        Myanmar
      sug:
        subj:
          Myanmar
          Low-resource languages
          Data augmentation
          Acquisition of data
          Citizen science
          Natural languages
          Natural language processing
          Language models
      keyword:
        Burmese
        Communication and Culture Linguistics Information and Computing Sciences Artificial Intelligence and Image Processing
        Language
        Low-resource
        Natural language inference
      ab: Despite dramatic recent progress in NLP, it is still a major challenge to apply Large Language Models (LLM) to low-resource languages. This is made visible in benchmarks such as Cross-Lingual Natural Language Inference (XNLI), a key task that demonstrates cross-lingual capabilities of NLP systems across a set of 15 languages. In this paper, we extend the XNLI task for one additional low-resource language, Myanmar, as a proxy challenge for broader low-resource languages, and make three core contributions. First, we build a dataset called Myanmar XNLI (myXNLI) using community crowd-sourced methods, as an extension to the existing XNLI corpus. This involves a two-stage process of community-based construction followed by expert verification; through an analysis, we demonstrate and quantify the value of the expert verification stage in the context of community-based construction for low-resource languages. We make the myXNLI dataset available to the community for future research. Second, we carry out evaluations of recent multilingual language models on the myXNLI benchmark, as well as explore data-augmentation methods to improve model performance. Our data-augmentation methods improve model accuracy by up to 2 percentage points for Myanmar, while uplifting other languages at the same time. Third, we investigate how well these data-augmentation methods generalise to other low-resource languages in the XNLI dataset.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N