A multi-platform dataset for detecting cyberbullying in social media.

Recent work on cyberbullying detection relies on using machine learning models with text and metadata in small datasets, mostly drawn from single social media platforms. Such models have succeeded in predicting cyberbullying when dealing with posts containing the text and the metadata structure as f...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 4; pp. 851 - 875
Autores principales: Van Bruwaene, David, Huang, Qianjia, Inkpen, Diana
Formato: Artículo
Publicado: Springer Nature Dec2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146751848&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 146751848
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2020
      vid: 54
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        146751848
        10.1007/s10579-020-09488-3
      ppf: 851
      ppct: 24
      formats:
        fmt:
          @attributes:
            type: P
            size: 465KB
      tig:
        atl: A multi-platform dataset for detecting cyberbullying in social media.
      aug:
        au:
          Van Bruwaene, David
          Huang, Qianjia
          Inkpen, Diana
        affil:
          SafeToNet Ltd., 51 Breithaupt Street, Suite 100, N2H 5G5, Kitchener, ON, Canada
          School of Electrical Engineering and Computer Science, University of Ottawa, K1N 6N5, Ottawa, ON, Canada
      su:
        Cyberbullying
        Social media
        Machine learning
        Natural language processing
      sug:
        subj:
          Cyberbullying
          Social media
          Machine learning
          Natural language processing
      keyword:
        Bullying
        Cyberaggression
        Dataset
        Deep learning
      ab: Recent work on cyberbullying detection relies on using machine learning models with text and metadata in small datasets, mostly drawn from single social media platforms. Such models have succeeded in predicting cyberbullying when dealing with posts containing the text and the metadata structure as found on the platform. Instead, we develop a multi-platform dataset that consists purely of the text from posts gathered from seven social media platforms. We present a multi-stage and multi-technique annotation system that initially uses crowdsourcing for post and hashtag annotation and subsequently utilizes machine-learning methods to identify additional posts for annotation. This process has the benefit of selecting posts for annotation that have a significantly greater than chance likelihood of constituting clear cases of cyberbullying without limiting the range of samples to those containing predetermined features (as is the case when hashtags alone are used to select posts for annotation). We show that, despite the diversity of examples present in the dataset, good performance is possible for models trained on datasets produced in this manner. This becomes a clear advantage compared to traditional methods of post selection and labeling because it increases the number of positive examples that can be produced using the same resources and it enhances the diversity of communication media to which the models can be applied.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N