A semi-supervised method to generate a persian dataset for suggestion classification.
Suggestion mining has become a popular subject in the field of natural language processing (NLP) that is useful in areas like a service/product improvement. The purpose of this study is to provide an automated machine learning (ML) based approach to extract suggestions from Persian text. In this res...
| Published in: | Language Resources & Evaluation Vol. 58; no. 2; pp. 839 - 859 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Jun2024
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=178064687&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 178064687 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2024 vid: 58 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 178064687 10.1007/s10579-023-09688-7 ppf: 839 ppct: 20 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: A semi-supervised method to generate a persian dataset for suggestion classification. aug: au: Safari, Leila Mohammady, Zanyar affil: https://ror.org/05e34ej29 Department of Computer Engineering, University of Zanjan, 4537138791, Zanjan, Iran su: Automatic classification Language models Natural language processing Machine learning Long-term memory Convolutional neural networks sug: subj: Automatic classification Language models Natural language processing Machine learning Long-term memory Convolutional neural networks keyword: Annotator Automatic classification of suggestions Neural networks Pre-trained language model Transformers ab: Suggestion mining has become a popular subject in the field of natural language processing (NLP) that is useful in areas like a service/product improvement. The purpose of this study is to provide an automated machine learning (ML) based approach to extract suggestions from Persian text. In this research, first, a novel two-step semi-supervised method has been proposed to generate a Persian dataset called ParsSugg, which is then used in the automatic classification of the user's suggestions. The first step is manual labeling of data based on a proposed guideline, followed by a data augmentation phase. In the second step, using pre-trained Persian Bidirectional Encoder Representations from Transformers (ParsBERT) as a classifier and the data from the previous step, more data were labeled. The performance of various ML models, including Support Vector Machine (SVM), Random Forest (RF), Convolutional Neural Networks (CNN), Long Short Term Memory (LSTM), and the ParsBERT language model has been examined on the generated dataset. The F-score value of 97.27 for ParsBERT and about 94.5 for SVM and CNN classifiers were obtained for the suggestion class which is a promising result as the first research on suggestion classification on Persian texts. Also, the proposed guideline can be used for other NLP tasks, and the generated dataset can be used in other suggestion classification tasks. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|