A semi-supervised method to generate a persian dataset for suggestion classification.

Suggestion mining has become a popular subject in the field of natural language processing (NLP) that is useful in areas like a service/product improvement. The purpose of this study is to provide an automated machine learning (ML) based approach to extract suggestions from Persian text. In this res...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 58; no. 2; pp. 839 - 859
Main Authors: Safari, Leila, Mohammady, Zanyar
Format: Article
Published: Springer Nature Jun2024
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=178064687&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 178064687
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2024
      vid: 58
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        178064687
        10.1007/s10579-023-09688-7
      ppf: 839
      ppct: 20
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.3MB
      tig:
        atl: A semi-supervised method to generate a persian dataset for suggestion classification.
      aug:
        au:
          Safari, Leila
          Mohammady, Zanyar
        affil: https://ror.org/05e34ej29 Department of Computer Engineering, University of Zanjan, 4537138791, Zanjan, Iran
      su:
        Automatic classification
        Language models
        Natural language processing
        Machine learning
        Long-term memory
        Convolutional neural networks
      sug:
        subj:
          Automatic classification
          Language models
          Natural language processing
          Machine learning
          Long-term memory
          Convolutional neural networks
      keyword:
        Annotator
        Automatic classification of suggestions
        Neural networks
        Pre-trained language model
        Transformers
      ab: Suggestion mining has become a popular subject in the field of natural language processing (NLP) that is useful in areas like a service/product improvement. The purpose of this study is to provide an automated machine learning (ML) based approach to extract suggestions from Persian text. In this research, first, a novel two-step semi-supervised method has been proposed to generate a Persian dataset called ParsSugg, which is then used in the automatic classification of the user's suggestions. The first step is manual labeling of data based on a proposed guideline, followed by a data augmentation phase. In the second step, using pre-trained Persian Bidirectional Encoder Representations from Transformers (ParsBERT) as a classifier and the data from the previous step, more data were labeled. The performance of various ML models, including Support Vector Machine (SVM), Random Forest (RF), Convolutional Neural Networks (CNN), Long Short Term Memory (LSTM), and the ParsBERT language model has been examined on the generated dataset. The F-score value of 97.27 for ParsBERT and about 94.5 for SVM and CNN classifiers were obtained for the suggestion class which is a promising result as the first research on suggestion classification on Persian texts. Also, the proposed guideline can be used for other NLP tasks, and the generated dataset can be used in other suggestion classification tasks.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N