OMCD: Offensive Moroccan Comments Dataset.
Offensive content, such as verbal attacks, demeaning comments, or hate speech, has become widespread on social media. Automatic detection of this content is considered an important and challenging task. Although several research works have been proposed to address this challenge for high-resource la...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 4; pp. 1745 - 1766 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=173723397&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 173723397 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2023 vid: 57 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 173723397 10.1007/s10579-023-09663-2 ppf: 1745 ppct: 21 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.2MB tig: atl: OMCD: Offensive Moroccan Comments Dataset. aug: au: Essefar, Kabil Ait Baha, Hassan El Mahdaouy, Abdelkader El Mekki, Abdellah Berrada, Ismail affil: https://ror.org/03xc55g68 School of Computer Sciences, Mohammed VI Polytechnic University, Ben Guerir, Morocco https://ror.org/03xc55g68 Modeling, Simulation and Data Analysis (MSDA), Mohammed VI Polytechnic University, Ben Guerir, Morocco su: Natural language processing Freedom of speech Social media Deep learning Machine learning Hate speech sug: subj: Natural language processing Freedom of speech Social media Deep learning Machine learning Hate speech keyword: Arabic NLP Moroccan dialect Offensive language Social media platforms Text classification ab: Offensive content, such as verbal attacks, demeaning comments, or hate speech, has become widespread on social media. Automatic detection of this content is considered an important and challenging task. Although several research works have been proposed to address this challenge for high-resource languages, research on detecting offensive content in Dialectal Arabic (DA) remains under-explored. Recently, the detection of offensive language in DA has gained increasing interest among researchers in Natural Language Processing (NLP). However, only a limited number of annotated datasets have been introduced for single or multiple coarse-grained dialects. In this paper, we introduce Offensive Moroccan Comments Dataset (OMCD), the first dataset for offensive language detection for the Moroccan dialect. First, we present the data collection steps, the statistical analysis, and the annotation guidelines of the introduced dataset. Then, we evaluate several state-of-the-art Machine Learning (ML) and Deep Learning (DL) based models on the OMCD dataset. Finally, we highlight the impact of emojis on the evaluated models for offensive language detection. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|