HASTIKA: hate speech and target identification in Kannada-English code-mixed text.
In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2811 - 2857 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909092&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909092 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909092 10.1007/s10579-025-09836-1 ppf: 2811 ppct: 46 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.6MB tig: atl: HASTIKA: hate speech and target identification in Kannada-English code-mixed text. aug: au: Kavatagi, Sanjana Rachh, Rashmi affil: https://ror.org/00ha14p11 Department of Computer Science and Engineering, Visvesvaraya Technological University, Belagavi, Karnataka, India su: Hate speech Code switching (Linguistics) Corpora Acquisition of data Classification algorithms Signal detection Social media sug: subj: Hate speech Code switching (Linguistics) Corpora Acquisition of data Classification algorithms Signal detection Social media keyword: Code-mixing Communication and Culture Linguistics Corpus building Kannada-English code-mixed text Language Low resource Social networks ab: In the modern era, the widespread use of social media has facilitated connections among millions of people worldwide. However, these platforms have also been exploited for spreading hate speech, particularly in multilingual contexts. The informal nature of these platforms enables the use of regional languages, leading to code-mixed text. While this linguistic flexibility fosters freedom of expression, it also contributes to the rampant spread of hate speech, posing significant societal challenges. Identifying hate speech in low-resource Kannada-English code-mixed text is challenging due to the scarcity of annotated corpora. To address this gap, the authors introduce HASTIKA (Hate Speech and Target Identification in Kannada-English Code-Mixed Text), a gold-standard corpus specifically designed for hate speech detection and target identification. It consists of 8,058 YouTube comments, annotated for binary classification "Hate" and "Non-hate" and fine-grained categorization into "Gender", "Political", "Religion", "Geo-political", "Violence", and "Others", marking a significant contribution as the first dataset exclusively tailored for this purpose. Topic modeling techniques were applied to uncover latent themes. Benchmark experiments show that fastText, which captures word-level features, achieves 0.7529 accuracy for binary classification and 0.6042 for multi-class classification, while BERT, excelling in sentence-level feature extraction, attains 0.8054 and 0.6819 accuracy, respectively. This research provides a foundational resource for hate speech detection in Kannada-English code-mixed text and underscores the importance of linguistic and contextual information in developing robust classification models. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|