Detecting caste and migration hate speech in low-resource Tamil language.
The Indian constitution categorizes its population into groups such as Scheduled Castes, Scheduled Tribes, Other Backward Classes, and Forward Castes, reflecting historical inequalities that influence social dynamics and discrimination. Migrants who are relocating within the country for better oppor...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3051 - 3087 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909099&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909099 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909099 10.1007/s10579-025-09848-x ppf: 3051 ppct: 36 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.9MB tig: atl: Detecting caste and migration hate speech in low-resource Tamil language. aug: au: Chakravarthi, Bharathi Raja Rajiakodi, Saranya Ponnusamy, Rahul Sivagnanam, Bhuvaneswari Thakare, Sara Yogesh Thangasamy, Sathiyaraj affil: https://ror.org/03bea9k73 School of Computer Science, University of Galway, Galway, Ireland https://ror.org/03bea9k73 Data Science Institute, University of Galway, Galway, Ireland https://ror.org/03ytqnm28 Department of Computer Science, Central University of Tamil Nadu, Thiruvarur, Tamil Nadu, India https://ror.org/02v7trd43 Indian Institute of Technology Goa, Bandiwada, India Department of Tamil, Sri Krishna Adithya College of Arts and Science, Kovaipudur, Tamil Nadu, India su: Caste Hate speech Detection algorithms Internal migration Discrimination (Sociology) Tamil (Indic people) Acquisition of data Social media sug: subj: Caste Hate speech Detection algorithms Internal migration Discrimination (Sociology) Tamil (Indic people) Acquisition of data Social media keyword: Adapters Caste and migration hate speech Large language models Low-resourced dataset creation Model optimization Parameter-efficient fine-tuning Tamil ab: The Indian constitution categorizes its population into groups such as Scheduled Castes, Scheduled Tribes, Other Backward Classes, and Forward Castes, reflecting historical inequalities that influence social dynamics and discrimination. Migrants who are relocating within the country for better opportunities are often viewed as outsiders, leading to concerns about job security and crime, fostering hate and discrimination against them. Social media has exacerbated these issues, becoming hotspots for caste and migration-related hate speech, especially in low-resource languages. This study introduces a novel dataset specifically curated to detect hate speech related to caste and migration in the low-resourced Tamil language. Using this dataset, we benchmarked the dataset with baseline experiments with the highest macro-F1 score of 0.73. We also created custom-modified models integrating custom loss functions, adapter-based fine-tuning, and parameter-efficient fine-tuning techniques. To support further research, we released the dataset, conducted a shared task, and ranked participant systems. By publicly releasing this dataset, we aimed to facilitate further study and improve the detection and mitigation of hateful content related to caste and migration on social media. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|