A multi-platform dataset for detecting cyberbullying in social media.
Recent work on cyberbullying detection relies on using machine learning models with text and metadata in small datasets, mostly drawn from single social media platforms. Such models have succeeded in predicting cyberbullying when dealing with posts containing the text and the metadata structure as f...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 4; pp. 851 - 875 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146751848&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 146751848 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2020 vid: 54 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 146751848 10.1007/s10579-020-09488-3 ppf: 851 ppct: 24 formats: fmt: @attributes: type: P size: 465KB tig: atl: A multi-platform dataset for detecting cyberbullying in social media. aug: au: Van Bruwaene, David Huang, Qianjia Inkpen, Diana affil: SafeToNet Ltd., 51 Breithaupt Street, Suite 100, N2H 5G5, Kitchener, ON, Canada School of Electrical Engineering and Computer Science, University of Ottawa, K1N 6N5, Ottawa, ON, Canada su: Cyberbullying Social media Machine learning Natural language processing sug: subj: Cyberbullying Social media Machine learning Natural language processing keyword: Bullying Cyberaggression Dataset Deep learning ab: Recent work on cyberbullying detection relies on using machine learning models with text and metadata in small datasets, mostly drawn from single social media platforms. Such models have succeeded in predicting cyberbullying when dealing with posts containing the text and the metadata structure as found on the platform. Instead, we develop a multi-platform dataset that consists purely of the text from posts gathered from seven social media platforms. We present a multi-stage and multi-technique annotation system that initially uses crowdsourcing for post and hashtag annotation and subsequently utilizes machine-learning methods to identify additional posts for annotation. This process has the benefit of selecting posts for annotation that have a significantly greater than chance likelihood of constituting clear cases of cyberbullying without limiting the range of samples to those containing predetermined features (as is the case when hashtags alone are used to select posts for annotation). We show that, despite the diversity of examples present in the dataset, good performance is possible for models trained on datasets produced in this manner. This becomes a clear advantage compared to traditional methods of post selection and labeling because it increases the number of positive examples that can be produced using the same resources and it enhances the diversity of communication media to which the models can be applied. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|