A cross-validation study of Turkish sentiment analysis datasets and tools.

In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability and reuse of Turkish datasets across studies has yielded highly diverse outcomes. To address this, we conducted...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 4; pp. 4003 - 4042
Autores principales: Çakıcı, Şevval, Karaduman, Dilara, Çırlan, Mehmet Akif, Hürriyetoğlu, Ali
Formato: Artículo
Publicado: Springer Nature Dec2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912041&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 189912041
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 59
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        189912041
        10.1007/s10579-025-09869-6
      ppf: 4003
      ppct: 39
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.8MB
      tig:
        atl: A cross-validation study of Turkish sentiment analysis datasets and tools.
      aug:
        au:
          Çakıcı, Şevval
          Karaduman, Dilara
          Çırlan, Mehmet Akif
          Hürriyetoğlu, Ali
        affil:
          https://ror.org/01jjhfr75 Ozyegin University, İstanbul, Turkey
          https://ror.org/00jzwgz36 Koc University, Istanbul, Turkey
          https://ror.org/04qw24q55 Wageningen Food Safety Research, Wageningen, Netherlands
      su:
        Sentiment analysis
        Transformer models
        Language models
        Databases
        Model validation
      sug:
        subj:
          Sentiment analysis
          Transformer models
          Language models
          Databases
          Model validation
      keyword:
        Cross-validation
        Deep learning models
        Taxonomy
        Turkish dataset
      ab: In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability and reuse of Turkish datasets across studies has yielded highly diverse outcomes. To address this, we conducted a systematic review of sentiment analysis studies on Turkish text. Our search identified 78 relevant studies, from which we extracted over 80 datasets. These studies were labeled using a comprehensive sentiment analysis taxonomy, and the dataset details were compiled into a structured repository. Furthermore, we evaluated the performance of four state-of-the-art models-XLM-T, BERTurk (fine-tuned with the BounTi dataset), TSAM, and TurkishBERTweet-on four widely-used Turkish datasets. Among the models, XLM-T achieved the highest performance with an accuracy of 0.92 and F1 score of 0.95 on the Twt dataset, while TSAM reached 0.97 accuracy and F1 score on the Humir dataset. Our empirical results demonstrate that model performance varies significantly based on dataset characteristics such as domain, balance, and linguistic structure. Our review revealed key research gaps, including the limited application of emotion-based and concept-based sentiment analysis techniques and the lack of domain diversity in Turkish sentiment datasets. By highlighting such gaps and compiling a centralized repository, this study provides a comprehensive and publicly accessible resource to guide future research in Turkish sentiment analysis.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N