Sentiment analysis in Portuguese tweets: an evaluation of diverse word representation models.

During the past years, we have seen a steady increase in the number of social networks worldwide. Among them, Twitter has consolidated its position as one of the most influential social platforms, with Brazilian Portuguese speakers holding the fifth position in the number of users. Due to the inform...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 58; no. 1; pp. 223 - 273
Autores principales: Vianna, Daniela, Carneiro, Fernando, Carvalho, Jonnathan, Plastino, Alexandre, Paes, Aline
Formato: Artículo
Publicado: Springer Nature Mar2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=176079988&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 176079988
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2024
      vid: 58
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        176079988
        10.1007/s10579-023-09661-4
      ppf: 223
      ppct: 50
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.2MB
      tig:
        atl: Sentiment analysis in Portuguese tweets: an evaluation of diverse word representation models.
      aug:
        au:
          Vianna, Daniela
          Carneiro, Fernando
          Carvalho, Jonnathan
          Plastino, Alexandre
          Paes, Aline
        affil:
          https://ror.org/02263ky35 Institute of Computing, Universidade Federal do Amazonas (UFAM), Manaus, AM, Brazil
          Jusbrasil, Salvador, Brazil
          https://ror.org/02rjhbb08 Institute of Computing, Universidade Federal Fluminense (UFF), Niterói, RJ, Brazil
          https://ror.org/0009eqg37 Instituto Federal Fluminense (IFF), Itaperuna, RJ, Brazil
      su:
        X Corp.
        Sentiment analysis
        Natural language processing
        Language models
        Portuguese language
        Social media
      sug:
        subj:
          X Corp.
          Sentiment analysis
          Natural language processing
          Language models
          Portuguese language
          Social media
      keyword:
        Brazilian Portuguese tweets
        Word representation
      ab: During the past years, we have seen a steady increase in the number of social networks worldwide. Among them, Twitter has consolidated its position as one of the most influential social platforms, with Brazilian Portuguese speakers holding the fifth position in the number of users. Due to the informal linguistic style of tweets, the discovery of information in such an environment poses a challenge to Natural Language Processing (NLP) tasks such as sentiment analysis. In this work, we state sentiment analysis as a binary (positive and negative) and multiclass (positive, negative, and neutral) classification task at the Portuguese-written tweet level. Following a feature extraction approach, embeddings are initially gathered for a tweet and then given as input to learning a classifier. This study was designed to evaluate the effectiveness of different word representations, from the original pre-trained language model to continued pre-training strategies, to improve the predictive performance of sentiment classification, using three different classifier algorithms and eight Portuguese tweets datasets. Because of the lack of a language model specific to Brazilian Portuguese tweets, we have expanded our evaluation to consider six different embeddings: fastText, GloVe, Word2Vec, BERT-multilingual (mBERT), BERTweet, and BERTimbau. The experiments showed that embeddings trained from scratch solely using the target Portuguese language, BERTimbau, outperform the static representations, fastText, GloVe, and Word2Vec, and the Transformer-based models BERT multilingual and BERTweet. In addition, we show that extracting the contextualized embedding without any adjustment to the pre-trained language model is the best approach for most datasets.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2024. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N