Ngalawan Ujaran Sengit: hate speech detection in indonesian code-mixed social media data.

Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of ab...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 59; no. 3; pp. 2387 - 2415
Main Authors: Pamungkas, Endang Wahyu, Chiril, Patricia
Format: Article
Published: Springer Nature Sep2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909069&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909069
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909069
        10.1007/s10579-025-09810-x
      ppf: 2387
      ppct: 28
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: Ngalawan Ujaran Sengit: hate speech detection in indonesian code-mixed social media data.
      aug:
        au:
          Pamungkas, Endang Wahyu
          Chiril, Patricia
        affil:
          https://ror.org/03cnmz153 Department of Informatics Engineering, Universitas Muhammadiyah Surakarta, Surakarta, Central Java, Indonesia
          https://ror.org/024mw5h28 Data Science Institute, University of Chicago, Chicago, Illinois, USA
      su:
        Hate speech
        Social media
        Hypothesis
        Code switching (Linguistics)
        Detection algorithms
        Indonesians
        Natural language processing
        Invective
        Indonesia
      sug:
        subj:
          Indonesia
          Hate speech
          Social media
          Hypothesis
          Code switching (Linguistics)
          Detection algorithms
          Indonesians
          Natural language processing
          Invective
      keyword:
        Abusive language detection
        Code-mixed language
        Communication and Culture Language Studies Linguistics
        Hate speech detection
        Language
        Low-resource languages
      ab: Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of abusive language. Though most of the available resources created for handling this issue are in English, the problem of abusive speech is not restricted to any single language. As with other Natural Language Processing tasks, detecting abusive speech in low-resource languages poses significant challenges. In this study, we conduct a set of preliminary experiments for detecting hate speech in Indonesian social media. Geographically, Indonesia consists of several regions, each having its own language. The inhabitants tend to use a mix of their own local language and Bahasa (the official national language) to engage on social media, this posing significant challenges in detecting hate speech. Our contribution is twofold: (1) we manually filter available hate speech corpora to collect code-mixed data written in Indonesian-Javanese and Indonesian-Sundanese; and (2) we conduct an extensive set of experiments to determine the most robust model for detecting hate speech in this newly created corpus. Our results show that models trained on closely related languages perform better compared to those trained on languages with a higher linguistic distance. We also investigated the possibility of translating our code-mixed corpus to resource-rich languages and found that the machine translation models were not effective in handling the unique linguistic properties of code-mixed data.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N