Toxic comment classification and rationale extraction in code-mixed text leveraging co-attentive multi-task learning: Toxic comment classification...: K. B. Nelatoori and H. B. Kommanti.
Detecting toxic comments and rationale for the offensiveness of a social media post promotes moderation of social media content. For this purpose, we propose a Co-Attentive Multi-task Learning (CA-MTL) model through transfer learning for low-resource Hindi-English (commonly known as Hinglish) toxic...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 1; pp. 161 - 191 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=183750663&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 183750663 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2025 vid: 59 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 183750663 10.1007/s10579-023-09708-6 ppf: 161 ppct: 30 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: Toxic comment classification and rationale extraction in code-mixed text leveraging co-attentive multi-task learning: Toxic comment classification...: K. B. Nelatoori and H. B. Kommanti. aug: au: Nelatoori, Kiran Babu Kommanti, Hima Bindu affil: https://ror.org/0456pcg54 Department of CSE, National Institute of Technology Andhra Pradesh, 534101, Tadepalligudem, Andhra Pradesh, India su: Internet content moderation Social media Learning modules Classification sug: subj: Internet content moderation Social media Learning modules Classification keyword: Annotated code-mixed corpus Co-attention module Multi-task learning Rationale extraction Toxic comments Toxic span ab: Detecting toxic comments and rationale for the offensiveness of a social media post promotes moderation of social media content. For this purpose, we propose a Co-Attentive Multi-task Learning (CA-MTL) model through transfer learning for low-resource Hindi-English (commonly known as Hinglish) toxic texts. Together, the cooperative tasks of rationale/span detection and toxic comment classification create a strong multi-task learning objective. A task collaboration module is designed to leverage the bi-directional attention between the classification and span prediction tasks. The combined loss function of the model is constructed using the individual loss functions of these two tasks. Although an English toxic span detection dataset exists, one for Hinglish code-mixed text does not exist as of today. Hence, we developed a dataset with toxic span annotations for Hinglish code-mixed text. The proposed CA-MTL model is compared against single-task and multi-task learning models that lack the co-attention mechanism, using multilingual and Hinglish BERT variants. The F1 scores of the proposed CA-MTL model with HingRoBERTa encoder for both tasks are significantly higher than the baseline models. Caution: This paper may contain words disturbing to some readers. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|