Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization.
The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxificatio...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2705 - 2737 |
|---|---|
| Autores principales: | , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909086&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909086 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909086 10.1007/s10579-025-09830-7 ppf: 2705 ppct: 32 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.6MB tig: atl: Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization. aug: au: Mohan, G. Bharathi Gayathri, M. Kumar, R. Prasanna affil: https://ror.org/03am10p12 Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham, Chennai, India https://ror.org/03am10p12 Department of Electronics and Communication Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Chennai, India su: Language models Reinforcement learning Hate speech Invective Internet content moderation Multiagent systems Computational linguistics Text summarization sug: subj: Language models Reinforcement learning Hate speech Invective Internet content moderation Multiagent systems Computational linguistics Text summarization keyword: Adversarial attack Detoxification Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Psychology Large language model ab: The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxification of large language models. The proposed framework enhances model robustness by enabling iterative refinement through agent interaction, effectively reducing toxicity, especially in contexts involving hate speech. In addition, reinforcement learning is applied to fine tune sequence to sequence models for detoxifying dialogue summaries. Using models such as FLAN T5, BART, and GODEL, we evaluate the approach on RealToxicity Prompts and ParaDetox datasets. The results show consistent reductions in toxicity scores while maintaining content fidelity and coherence. These findings demonstrate the effectiveness of multi agent collaboration and learning based adaptation in mitigating toxic language and improving safety in real world applications such as content moderation and summarization. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|