Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization.

The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxificatio...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2705 - 2737
Autores principales: Mohan, G. Bharathi, Gayathri, M., Kumar, R. Prasanna
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909086&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909086
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909086
        10.1007/s10579-025-09830-7
      ppf: 2705
      ppct: 32
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.6MB
      tig:
        atl: Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization.
      aug:
        au:
          Mohan, G. Bharathi
          Gayathri, M.
          Kumar, R. Prasanna
        affil:
          https://ror.org/03am10p12 Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham, Chennai, India
          https://ror.org/03am10p12 Department of Electronics and Communication Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Chennai, India
      su:
        Language models
        Reinforcement learning
        Hate speech
        Invective
        Internet content moderation
        Multiagent systems
        Computational linguistics
        Text summarization
      sug:
        subj:
          Language models
          Reinforcement learning
          Hate speech
          Invective
          Internet content moderation
          Multiagent systems
          Computational linguistics
          Text summarization
      keyword:
        Adversarial attack
        Detoxification
        Information and Computing Sciences Artificial Intelligence and Image Processing Psychology and Cognitive Sciences Psychology
        Large language model
      ab: The increasing prevalence of online user generated content has raised serious concerns about toxic language, which reinforces societal biases and causes psychological harm. This study introduces a novel approach that combines multi agent debate and reinforcement learning to improve the detoxification of large language models. The proposed framework enhances model robustness by enabling iterative refinement through agent interaction, effectively reducing toxicity, especially in contexts involving hate speech. In addition, reinforcement learning is applied to fine tune sequence to sequence models for detoxifying dialogue summaries. Using models such as FLAN T5, BART, and GODEL, we evaluate the approach on RealToxicity Prompts and ParaDetox datasets. The results show consistent reductions in toxicity scores while maintaining content fidelity and coherence. These findings demonstrate the effectiveness of multi agent collaboration and learning based adaptation in mitigating toxic language and improving safety in real world applications such as content moderation and summarization.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N