Comparative performance of ensemble machine learning for Arabic cyberbullying and offensive language detection.

Since cyberbullying impacts both individual victims and entire society, research on abusive language and its detection has attracted attention in recent years. Because social media sites like Facebook, Instagram, Twitter, and others are so widely accessible, hate speech, bullying, sexism, racism, ag...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 2; pp. 695 - 713
Autores principales: Khairy, Marwa, Mahmoud, Tarek M., Omar, Ahmed, Abd El-Hafeez, Tarek
Formato: Artículo
Publicado: Springer Nature Jun2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=178064684&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 178064684
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2024
      vid: 58
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        178064684
        10.1007/s10579-023-09683-y
      ppf: 695
      ppct: 18
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1MB
      tig:
        atl: Comparative performance of ensemble machine learning for Arabic cyberbullying and offensive language detection.
      aug:
        au:
          Khairy, Marwa
          Mahmoud, Tarek M.
          Omar, Ahmed
          Abd El-Hafeez, Tarek
        affil:
          https://ror.org/02hcv4z63 Faculty of Computers and Information, Minia University, El Minya, Egypt
          https://ror.org/05p2q6194 Faculty of Computers and Artificial Intelligence, University of Sadat City, Sadat City, Egypt
          https://ror.org/02hcv4z63 Computer Science Department, Faculty of Science, Minia University, El Minya, Egypt
          Deraya University, El Minya, Egypt
      su:
        Cyberbullying
        Machine learning
        Social media
        Online social networks
        Machine performance
        Invective
      sug:
        subj:
          Cyberbullying
          Machine learning
          Social media
          Online social networks
          Machine performance
          Invective
      keyword:
        Arabic offensive language
        Ensemble machine learning
        NLP
        OSN
      ab: Since cyberbullying impacts both individual victims and entire society, research on abusive language and its detection has attracted attention in recent years. Because social media sites like Facebook, Instagram, Twitter, and others are so widely accessible, hate speech, bullying, sexism, racism, aggressive material, harassment, poisonous comments, and other types of abuse have all substantially increased. Due to the critical requirement to detect, regulate, and limit the spread of harmful content on social networking sites, we conducted this study to automate the detection of offensive language or cyberbullying. We created a new Arabic balanced data set to be used in the offensive detection process because having a balanced data set for a model would result in improved accuracy models. Recently, the performance of single classifiers has been improved using ensemble machine learning. The purpose of this study is to examine the effectiveness of several single and ensemble machine learning algorithms in identifying Arabic text that contains foul language and cyberbullying. Applying them to three Arabic datasets, we have selected three machine learning classifiers and three ensemble models for this aim. Two of them are offensive datasets that are readily accessible in the public, while the third one was created. The results showed that the single learner machine learning strategy is inferior to the ensemble machine learning methodology. Voting performs is the best performing trained ensemble machine learning classifier, outperforming the best single learner classifier (65.1%, 76.2%, and 98%) for the same datasets with accuracy scores of (71.1%, 76.7%, and 98.5%) for each of the three datasets used. Finally, we improve the voting technique's performance through hyperparameter tuning on the Arabic cyberbullying data set.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N