Detecting caste and migration hate speech in low-resource Tamil language.

The Indian constitution categorizes its population into groups such as Scheduled Castes, Scheduled Tribes, Other Backward Classes, and Forward Castes, reflecting historical inequalities that influence social dynamics and discrimination. Migrants who are relocating within the country for better oppor...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3051 - 3087
Autores principales: Chakravarthi, Bharathi Raja, Rajiakodi, Saranya, Ponnusamy, Rahul, Sivagnanam, Bhuvaneswari, Thakare, Sara Yogesh, Thangasamy, Sathiyaraj
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909099&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909099
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909099
        10.1007/s10579-025-09848-x
      ppf: 3051
      ppct: 36
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.9MB
      tig:
        atl: Detecting caste and migration hate speech in low-resource Tamil language.
      aug:
        au:
          Chakravarthi, Bharathi Raja
          Rajiakodi, Saranya
          Ponnusamy, Rahul
          Sivagnanam, Bhuvaneswari
          Thakare, Sara Yogesh
          Thangasamy, Sathiyaraj
        affil:
          https://ror.org/03bea9k73 School of Computer Science, University of Galway, Galway, Ireland
          https://ror.org/03bea9k73 Data Science Institute, University of Galway, Galway, Ireland
          https://ror.org/03ytqnm28 Department of Computer Science, Central University of Tamil Nadu, Thiruvarur, Tamil Nadu, India
          https://ror.org/02v7trd43 Indian Institute of Technology Goa, Bandiwada, India
          Department of Tamil, Sri Krishna Adithya College of Arts and Science, Kovaipudur, Tamil Nadu, India
      su:
        Caste
        Hate speech
        Detection algorithms
        Internal migration
        Discrimination (Sociology)
        Tamil (Indic people)
        Acquisition of data
        Social media
      sug:
        subj:
          Caste
          Hate speech
          Detection algorithms
          Internal migration
          Discrimination (Sociology)
          Tamil (Indic people)
          Acquisition of data
          Social media
      keyword:
        Adapters
        Caste and migration hate speech
        Large language models
        Low-resourced dataset creation
        Model optimization
        Parameter-efficient fine-tuning
        Tamil
      ab: The Indian constitution categorizes its population into groups such as Scheduled Castes, Scheduled Tribes, Other Backward Classes, and Forward Castes, reflecting historical inequalities that influence social dynamics and discrimination. Migrants who are relocating within the country for better opportunities are often viewed as outsiders, leading to concerns about job security and crime, fostering hate and discrimination against them. Social media has exacerbated these issues, becoming hotspots for caste and migration-related hate speech, especially in low-resource languages. This study introduces a novel dataset specifically curated to detect hate speech related to caste and migration in the low-resourced Tamil language. Using this dataset, we benchmarked the dataset with baseline experiments with the highest macro-F1 score of 0.73. We also created custom-modified models integrating custom loss functions, adapter-based fine-tuning, and parameter-efficient fine-tuning techniques. To support further research, we released the dataset, conducted a shared task, and ranked participant systems. By publicly releasing this dataset, we aimed to facilitate further study and improve the detection and mitigation of hateful content related to caste and migration on social media.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N