DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text.
This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language identification for a total of more than 60,000 YouTube commen...
| Published in: | Language Resources & Evaluation Vol. 56; no. 3; pp. 765 - 807 |
|---|---|
| Main Authors: | , , , , , , |
| Format: | Article |
| Published: |
Springer Nature
Sep2022
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609442&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 158609442 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2022 vid: 56 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 158609442 10.1007/s10579-022-09583-7 ppf: 765 ppct: 42 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.2MB tig: atl: DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text. aug: au: Chakravarthi, Bharathi Raja Priyadharshini, Ruba Muralidaran, Vigneshwaran Jose, Navya Suryawanshi, Shardul Sherly, Elizabeth McCrae, John P. affil: Insight SFI Research Centre for Data Analytics, Data Science Institute, National University of Ireland Galway, Galway, Ireland ULTRA Arts and Science College, Madurai, Tamil Nadu, India School of Computer Science and Informatics, Cardiff University, Cardiff, UK Indian Institute of Information Technology and Management-Kerala, Kazhakkoottam, Kerala, India su: Google Inc. Sentiment analysis User-generated content Deep learning Machine learning Social media sug: subj: Google Inc. Sentiment analysis User-generated content Deep learning Machine learning Social media keyword: Code-mixed Corpora Dravidian languages Kannada Malayalam Offensive language identification Tamil ab: This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language identification for a total of more than 60,000 YouTube comments. The dataset consists of around 44,000 comments in Tamil-English, around 7000 comments in Kannada-English, and around 20,000 comments in Malayalam-English. The data was manually annotated by volunteer annotators and has a high inter-annotator agreement in Krippendorff's alpha. The dataset contains all types of code-mixing phenomena since it comprises user-generated content from a multilingual country. We also present baseline experiments to establish benchmarks on the dataset using machine learning and deep learning methods. The dataset is available on Github and Zenodo. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|