HASTIKA: hate speech and target identification in Kannada-English code-mixed text.

In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2811 - 2857
Autores principales: Kavatagi, Sanjana, Rachh, Rashmi
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909092&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909092
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909092
        10.1007/s10579-025-09836-1
      ppf: 2811
      ppct: 46
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.6MB
      tig:
        atl: HASTIKA: hate speech and target identification in Kannada-English code-mixed text.
      aug:
        au:
          Kavatagi, Sanjana
          Rachh, Rashmi
        affil: https://ror.org/00ha14p11 Department of Computer Science and Engineering, Visvesvaraya Technological University, Belagavi, Karnataka, India
      su:
        Hate speech
        Code switching (Linguistics)
        Corpora
        Acquisition of data
        Classification algorithms
        Signal detection
        Social media
      sug:
        subj:
          Hate speech
          Code switching (Linguistics)
          Corpora
          Acquisition of data
          Classification algorithms
          Signal detection
          Social media
      keyword:
        Code-mixing
        Communication and Culture Linguistics
        Corpus building
        Kannada-English code-mixed text
        Language
        Low resource
        Social networks
      ab: In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional languages, leading to code-mixed text. While this linguistic flexibility fosters freedom of expression, it also contributes to the rampant spread of hate speech, posing significant societal challenges. Identifying hate speech in low-resource Kannada-English code-mixed text is challenging due to the scarcity of annotated corpora. To address this gap, the authors introduce HASTIKA (Hate Speech and Target Identification in Kannada-English Code-Mixed Text), a gold-standard corpus specifically designed for hate speech detection and target identification. It consists of 8,058 YouTube comments, annotated for binary classification "Hate" and "Non-hate" and fine-grained categorization into "Gender", "Political", "Religion", "Geo-political", "Violence", and "Others", marking a significant contribution as the first dataset exclusively tailored for this purpose. Topic modeling techniques were applied to uncover latent themes. Benchmark experiments show that fastText, which captures word-level features, achieves 0.7529 accuracy for binary classification and 0.6042 for multi-class classification, while BERT, excelling in sentence-level feature extraction, attains 0.8054 and 0.6819 accuracy, respectively. This research provides a foundational resource for hate speech detection in Kannada-English code-mixed text and underscores the importance of linguistic and contextual information in developing robust classification models.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N