Data augmentation strategies to improve text classification: a use case in smart cities.

Text classification is a very common and important task in Natural Language Processing. In many domains and real-world settings, a few labeled instances are the only resource available to train classifiers. Models trained on small datasets tend to overfit and produce inaccurate results – Data augmen...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 58; no. 2; pp. 659 - 695
Main Authors: Bencke, Luciana, Moreira, Viviane Pereira
Format: Article
Published: Springer Nature Jun2024
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=178064686&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 178064686
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2024
      vid: 58
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        178064686
        10.1007/s10579-023-09685-w
      ppf: 659
      ppct: 36
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: Data augmentation strategies to improve text classification: a use case in smart cities.
      aug:
        au:
          Bencke, Luciana
          Moreira, Viviane Pereira
        affil: https://ror.org/041yk2d64 Institute of Informatics, Federal University of Rio Grande do Sul (UFRGS), Porto Alegre, Rio Gande do Sul, Brasil
      su:
        Data augmentation
        Smart cities
        Natural language processing
        Language models
        Classification algorithms
      sug:
        subj:
          Data augmentation
          Smart cities
          Natural language processing
          Language models
          Classification algorithms
      keyword:
        Low-resources
        Text classification
      ab: Text classification is a very common and important task in Natural Language Processing. In many domains and real-world settings, a few labeled instances are the only resource available to train classifiers. Models trained on small datasets tend to overfit and produce inaccurate results – Data augmentation (DA) techniques come as an alternative to minimize this problem. DA generates synthetic instances that can be fed to the classification algorithm during training. In this article, we explore a variety of DA methods, including back translation, paraphrasing, and text generation. We assess the impact of the DA methods over simulated low-data scenarios using well-known public datasets in English with classifiers built fine-tuning BERT models. We describe the means to adapt these DA methods to augment a small Portuguese dataset containing tweets labeled with smart city dimensions (e.g., transportation, energy, water, etc.). Our experiments showed that some classes were noticeably improved by DA – with an improvement of 43% in terms of F1 compared to the baseline with no augmentation. In a qualitative analysis, we observed that the DA methods were able to preserve the label but failed to preserve the semantics in some cases and that generative models were able to produce high-quality synthetic instances.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N