Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization.
The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxificatio...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 2705 - 2737 |
|---|---|
| Autores principales: | , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |