Low resource language specific pre-processing and features for sentiment analysis task.
Sentiment analysis is a classification task where polarity of textual data is identified, i.e. to analyze whether a sentence or document expresses a negative, positive or neutral sentiment. Manipuri is a less privileged, highly agglutinative and tonal language. Despite being a scheduled language of...
| Publicado en: | Language Resources & Evaluation Vol. 55; no. 4; pp. 947 - 970 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2021
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=152947518&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 152947518 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2021 vid: 55 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 152947518 10.1007/s10579-021-09541-9 ppf: 947 ppct: 23 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: Low resource language specific pre-processing and features for sentiment analysis task. aug: au: Meetei, Loitongbam Sanayai Singh, Thoudam Doren Borgohain, Samir Kumar Bandyopadhyay, Sivaji affil: Center for Natural Language Processing (CNLP), National Institute of Technology Silchar, Silchar, Assam, India Department of Computer Science and Engineering, National Institute of Technology Silchar, Silchar, Assam, India su: Sentiment analysis Task analysis Deep learning Machine learning Tone (Phonetics) sug: subj: Sentiment analysis Task analysis Deep learning Machine learning Tone (Phonetics) keyword: BM25 Ensembled classifier Low resource Manipuri Morphology Pre-processing TF-IDF ab: Sentiment analysis is a classification task where polarity of textual data is identified, i.e. to analyze whether a sentence or document expresses a negative, positive or neutral sentiment. Manipuri is a less privileged, highly agglutinative and tonal language. Despite being a scheduled language of Indian Constitution, it is also a resource constrained language. In this work, we report the sentiment analysis for Manipuri using different types of machine learning based approaches. The dataset used in our work is collected from local daily newspaper. The novelty of this work is that we carry out language specific pre-processing tasks such as transliteration, building negative morpheme based lexicon and filtering of noisy words. Using them as additional linguistic features in our models improves the classification result in terms of precision, recall and F-score. The ensemble voting of best three classifiers based on TF-IDF perform better than BM25 based classifiers and other stand-alone classifiers. Based on this result, we attempt to classify the sentiment of news articles during a certain period of time. Further, we report the finding of deep learning based approaches on the same dataset. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|