Ngalawan Ujaran Sengit: hate speech detection in indonesian code-mixed social media data.

Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of ab...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2387 - 2415
Autores principales: Pamungkas, Endang Wahyu, Chiril, Patricia
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Social networking sites have become an important medium for online communication, enabling users from around the world to connect, share multimodal content, and express their feelings and opinions, regardless of the topic. As the number of users on these platforms increases, so does the amount of abusive language. Though most of the available resources created for handling this issue are in English, the problem of abusive speech is not restricted to any single language. As with other Natural Language Processing tasks, detecting abusive speech in low-resource languages poses significant challenges. In this study, we conduct a set of preliminary experiments for detecting hate speech in Indonesian social media. Geographically, Indonesia consists of several regions, each having its own language. The inhabitants tend to use a mix of their own local language and Bahasa (the official national language) to engage on social media, this posing significant challenges in detecting hate speech. Our contribution is twofold: (1) we manually filter available hate speech corpora to collect code-mixed data written in Indonesian-Javanese and Indonesian-Sundanese; and (2) we conduct an extensive set of experiments to determine the most robust model for detecting hate speech in this newly created corpus. Our results show that models trained on closely related languages perform better compared to those trained on languages with a higher linguistic distance. We also investigated the possibility of translating our code-mixed corpus to resource-rich languages and found that the machine translation models were not effective in handling the unique linguistic properties of code-mixed data.