Aspect-based multimodal sentiment analysis via employing visual-to-emotional-caption translation network using visual-caption pairs.

The growing interest in multimodal sentiment recognition has driven advancements in Aspect-Based Multimodal Sentiment Analysis (ABMSA), focusing on sentiment extraction from visual-caption pairs. Despite progress, integrating emotional cues from visual modalities, especially facial expressions, rema...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 2945 - 2973
Autores principales: Pandey, Ananya, Vishwakarma, Dinesh Kumar
Formato: Artículo
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909082&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909082
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909082
        10.1007/s10579-025-09824-5
      ppf: 2945
      ppct: 28
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.5MB
      tig:
        atl: Aspect-based multimodal sentiment analysis via employing visual-to-emotional-caption translation network using visual-caption pairs.
      aug:
        au:
          Pandey, Ananya
          Vishwakarma, Dinesh Kumar
        affil: https://ror.org/01ztcvt22 Multi-Modal Data Analytics Research Laboratory, Department of Information Technology, Delhi Technological University, Bawana Road, 110042, Delhi, India
      su:
        Facial expression
        Sentiment analysis
        Emotional conditioning
        Twitter (Web resource)
      sug:
        subj:
          Facial expression
          Sentiment analysis
          Emotional conditioning
          Twitter (Web resource)
      keyword:
        Aspect
        Caption
        Face descriptions
        Information and Computing Sciences Artificial Intelligence and Image Processing
        Multimodal
        Visual
      ab: The growing interest in multimodal sentiment recognition has driven advancements in Aspect-Based Multimodal Sentiment Analysis (ABMSA), focusing on sentiment extraction from visual-caption pairs. Despite progress, integrating emotional cues from visual modalities, especially facial expressions, remains an underexplored challenge. The key difficulty lies in efficiently extracting and aligning these visual and emotional cues with corresponding textual information. To address this gap, we introduce a novel approach, the Visual-to-Emotional-Caption Translation Network (VECTN), which captures visual sentiment cues from facial expressions and aligns them with target attributes in the caption. The effectiveness of the proposed model is validated through comparison with several state-of-the-art methods, including TomBERT [as reported by Yu and Jiang (in: Adapting BERT for target-oriented multimodal sentiment classification, 2019, https://github.com/jefferyYu/TomBERT)], ESAFN (Yu et al. in IEEE/ACM Trans Audio Speech Lang Process 28:429–439, 2020, https://doi.org/10.1109/TASLP.2019.2957872), EF-Net (Gu et al. in IEEE Access 9:157329–157336, 2021, https://doi.org/10.1109/ACCESS.2021.3126782), ModelNet (Zhang et al. in World Wide Web, 2021, https://doi.org/10.1007/s11280-021-00955-7), R-GCN [as reported by Zhao et al. (in: Proceedings of the 30th ACM international conference on multimedia, ACM, Lisboa, 2022, https://doi.org/10.1145/3503161.3548228)], HIMT (Yu et al. in IEEE Trans Affect Comput, 2022, https://doi.org/10.1109/TAFFC.2022.3171091), FGSN [as reported by Lu et al. (in: 2022 4th international conference on robotics and computer vision (ICRCV), 2022, https://doi.org/10.1109/ICRCV55858.2022.9953248), as reported by Gupta et al. (in: 2023 in International conference in advances in power, signal, and information technology (APSIT), IEEE, Bhubaneswar, 2023, https://doi.org/10.1109/APSIT58554.2023.10201711), EF-CaTrBERT [as reported by Khan et al. (in: MM 2021—proceedings of the 29th ACM international conference on multimedia, Association for Computing Machinery Inc, 2021, https://doi.org/10.1145/3474085.3475692)], CBMLNet (Li et al. in Signal Image Video Process 18:8403–8412, 2024, https://doi.org/10.1007/s11760-024-03482-w), REF (Chen et al. in Int J Mach Learn Cybern, 2024, https://doi.org/10.1007/s13042-024-02342-w), and FMCF (Du et al. in Appl Intell 54:12629–12643, 2024, https://doi.org/10.1007/s10489-024-05841-z). Experimental results show that the VECTN model outperforms existing approaches, achieving an accuracy of 81.23% and a macro-F1 score of 80.61% on the Twitter-2015 dataset, and 77.42% accuracy and 75.19% macro-F1 score on the Twitter-2017 dataset. Additionally, we utilized the Matthews Correlation Coefficient (MCC ) score to evaluate the model's robustness and reliability. The high MCC scores achieved by our model- 81.23% on Twitter-2015 dataset, 76.76% on Twitter-2017 dataset, and 80.04% on Tweet1517-Face dataset, demonstrate its exceptional ability to handle class imbalance and deliver accurate predictions in Aspect-Based Multimodal Sentiment Analysis (ABMSA). These results confirm that VECTN is highly effective in collecting target-level sentiment in multimodal data, particularly by utilizing facial expressions to enhance sentiment recognition. The performance improvement underscores the model's advantage over others in this domain.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N